Skip to main content

qwen-mm-plugins-api

Cloud MCP tools for understanding media, by model family. VL model: vision_chat (caption/VQA), ocr, grounding (detect/locate objects). Omni model (reads frames + audio together): timestamped captioning, ASR (plain / controllable / multi-speaker diarized), temporal grounding, event counting, music captioning. Plus transcribe_audio (ASR) and segmentation (SAM3). Use when a question about an image/video/audio needs an external model, not just local reading.

Aller à l'installation

Informations de source

Dépôt
QwenLM/Qwen-MM-Plugins
Dernière activité de la source
11 août 2026 à 05:15
Langue détectée de SKILL.md
anglais
Étoiles
2 214
Forks
115

Options d'installation

Le prompt qui vérifie d'abord la source est sélectionné par défaut. Vous pouvez passer à une commande directe ou télécharger une copie locale.

Vérifiez les fichiers source

Lisez SKILL.md et les fichiers associés affichés par SkillsMP avant de décider de l'installer.