Skip to main content

multimodal-learning

Multi-modal learning, vision-language models, and cross-modal representation

Jump to install

Source facts

Repository
NeuralBlitz/Mito
Last source activity
March 22, 2026 at 13:29
Detected SKILL.md language
English
Stars
0
Forks
0

Install options

The review-first prompt is selected by default. You can switch to a direct command or download a local copy.

Review the source files

Read SKILL.md and any companion files shown by SkillsMP before deciding whether to install.

Showing SKILL.md

SKILL.md
Source instructions · Read-only preview
name
multimodal-learning
description
Multi-modal learning, vision-language models, and cross-modal representation
license
MIT
compatibility
opencode
metadata
{"audience":"researchers","category":"machine-learning"}
## What I do - Build models that process multiple modalities - Work with vision-language models - Align different data types - Fuse multimodal representations ## When to use me When working on multimodal AI or vision-language systems. ## Key Concepts - Vision-language models - Cross-modal attention - Representation learning - CLIP - Image captioning - Visual question answering - Audio-visual learning
View on GitHub