Giga-Embeddings: Mixture-of-Experts Encoders for High-Throughput Text Embeddings
本文提出Giga-Embeddings模型,采用Mixture-of-Experts编码器解决高效文本嵌入问题,实现高质量检索与高吞吐量处理。
本文提出Giga-Embeddings模型,采用Mixture-of-Experts编码器解决高效文本嵌入问题,实现高质量检索与高吞吐量处理。
This work addresses the performance gap in multilingual automatic speech recognition for low-resource Central Asian languages—such as Kazakh, Kyrgyz, and Uzbek—caused by severe data scarcity. The authors propose a robust foundation model trained on 2 million hours of audio using a HuBERT-style objective to pretrain a Conformer encoder. To mitigate dominance by high-resource languages, they introduce a novel cluster-level data balancing strategy and a domain-aware fine-tuning sampling approach. Experimental results demonstrate that the proposed model significantly outperforms strong open-source baselines like Whisper Large v3, particularly on spontaneous speech, while maintaining efficient inference capabilities.
This work addresses the challenge of precise temporal localization in long-form audio recordings—up to 120 minutes—by introducing a time-aware audio large language model. The proposed approach interleaves learnable time tokens periodically within continuous audio token sequences, creating a joint representation that tightly couples temporal information with acoustic content. The model is trained on large-scale synthetic supervision data generated via a cascaded pipeline and employs a mixed-duration training strategy to enhance robustness across varying audio lengths. It achieves high-precision performance on temporally grounded tasks such as time-anchored question answering, segment description, and summarization. Experimental results demonstrate significant improvements over existing methods on both short and long audio benchmarks. To foster further research in audio temporal understanding, the authors have publicly released the model weights and associated datasets.
To address the lag in developing large language models (LLMs) for Russian—largely due to prohibitive computational costs—this work introduces GigaChat, the first open-source family of Mixture-of-Experts (MoE) LLMs specifically designed for Russian. GigaChat leverages Russian-centric pretraining, supervised instruction fine-tuning, and comprehensive multi-stage evaluation on Russian benchmarks (e.g., RuEval, RACE-Ru, XGLUE), achieving high performance while substantially reducing training and inference overhead. The family comprises multiple base model sizes and corresponding instruction-tuned variants, consistently outperforming multilingual baselines such as mT5 and BLOOM on Russian understanding and generation tasks. To foster adoption, three models are publicly released on Hugging Face, accompanied by production-ready interfaces—including a REST API, Telegram bot, and web application—enabling scalable industrial deployment.
To address the limited generalization capability of self-supervised pretraining for low-resource languages (e.g., Russian) in automatic speech recognition (ASR), this paper proposes GigaAM: (1) a recognition-oriented joint self-supervised objective that unifies masked language modeling with knowledge distillation signals from ASR models; (2) a novel dynamic chunked attention mechanism, enabling variable chunk-size sampling to simultaneously support full-context offline inference and low-latency streaming fine-tuning; and (3) large-scale model scaling during training, yielding the open-source GigaAM model family released under the MIT License. On Russian ASR benchmarks, GigaAM reduces word error rate (WER) by 50% relative to Whisper-large-v3, demonstrating substantial improvements in robustness for low-resource languages and flexibility for diverse deployment scenarios.
本文提出Giga-Embeddings模型,采用Mixture-of-Experts编码器解决高效文本嵌入问题,实现高质量检索与高吞吐量处理。
This work addresses the performance gap in multilingual automatic speech recognition for low-resource Central Asian languages—such as Kazakh, Kyrgyz, and Uzbek—caused by severe data scarcity. The authors propose a robust foundation model trained on 2 million hours of audio using a HuBERT-style objective to pretrain a Conformer encoder. To mitigate dominance by high-resource languages, they introduce a novel cluster-level data balancing strategy and a domain-aware fine-tuning sampling approach. Experimental results demonstrate that the proposed model significantly outperforms strong open-source baselines like Whisper Large v3, particularly on spontaneous speech, while maintaining efficient inference capabilities.
This work addresses the challenge of precise temporal localization in long-form audio recordings—up to 120 minutes—by introducing a time-aware audio large language model. The proposed approach interleaves learnable time tokens periodically within continuous audio token sequences, creating a joint representation that tightly couples temporal information with acoustic content. The model is trained on large-scale synthetic supervision data generated via a cascaded pipeline and employs a mixed-duration training strategy to enhance robustness across varying audio lengths. It achieves high-precision performance on temporally grounded tasks such as time-anchored question answering, segment description, and summarization. Experimental results demonstrate significant improvements over existing methods on both short and long audio benchmarks. To foster further research in audio temporal understanding, the authors have publicly released the model weights and associated datasets.
To address the lag in developing large language models (LLMs) for Russian—largely due to prohibitive computational costs—this work introduces GigaChat, the first open-source family of Mixture-of-Experts (MoE) LLMs specifically designed for Russian. GigaChat leverages Russian-centric pretraining, supervised instruction fine-tuning, and comprehensive multi-stage evaluation on Russian benchmarks (e.g., RuEval, RACE-Ru, XGLUE), achieving high performance while substantially reducing training and inference overhead. The family comprises multiple base model sizes and corresponding instruction-tuned variants, consistently outperforming multilingual baselines such as mT5 and BLOOM on Russian understanding and generation tasks. To foster adoption, three models are publicly released on Hugging Face, accompanied by production-ready interfaces—including a REST API, Telegram bot, and web application—enabling scalable industrial deployment.
To address the limited generalization capability of self-supervised pretraining for low-resource languages (e.g., Russian) in automatic speech recognition (ASR), this paper proposes GigaAM: (1) a recognition-oriented joint self-supervised objective that unifies masked language modeling with knowledge distillation signals from ASR models; (2) a novel dynamic chunked attention mechanism, enabling variable chunk-size sampling to simultaneously support full-context offline inference and low-latency streaming fine-tuning; and (3) large-scale model scaling during training, yielding the open-source GigaAM model family released under the MIT License. On Russian ASR benchmarks, GigaAM reduces word error rate (WER) by 50% relative to Whisper-large-v3, demonstrating substantial improvements in robustness for low-resource languages and flexibility for diverse deployment scenarios.