CoME-VL: Scaling Complementary Multi-Encoder Vision-Language Learning
-
Updated
Aug 25, 2026 - Python
CoME-VL: Scaling Complementary Multi-Encoder Vision-Language Learning
A Gradio-based demonstration for the AllenAI SAGE-MM-Qwen3-VL-4B-SFT_RL multimodal model, specialized in video reasoning tasks. Users upload MP4 videos, provide natural language prompts (e.g., "Describe this video in detail" or custom questions), and receive detailed textual analyses.
A Gradio-based demonstration for the AllenAI Molmo2-8B multimodal model, enabling image QA, multi-image pointing, video QA, and temporal tracking. Users upload images or videos, provide natural language prompts.
🎥 Analyze MP4 videos interactively, generating detailed text responses from custom prompts using the SAGE-MM visual reasoning model.
To associate your repository with the molmo2 topic, visit your repo's landing page and select "manage topics."