OmniBench: Towards The Future of Universal Omni-Language Models
Existing open-source multimodal large language models (MLLMs) exhibit significant deficiencies in joint visual-auditory-textual understanding and reasoning, achieving only ~50% instruction-following accuracy on trilingual multimodal tasks. Method: We introduce OmniBench—the first benchmark for trilingual multimodal collaborative reasoning—and formalize the omni-language model (OLM), a unified architecture capable of jointly processing visual, auditory, and textual (V-A-T) inputs. We construct OmniBench via expert human annotation across diverse trilingual multimodal tasks and curate OmniInstruct, a large-scale instruction-tuning dataset comprising 96K samples. Our methodology integrates cross-modal alignment modeling, trilingual multimodal instruction tuning, and a human-in-the-loop evaluation framework. Contribution/Results: Experiments reveal severe generalization limitations of current open-source OLMs on trilingual multimodal tasks; OmniInstruct substantially improves their reasoning performance. This work establishes a novel evaluation paradigm, provides high-quality resources, and outlines a technical pathway for advancing trilingual multimodal foundation models.