Llama-Mobile: Efficient 2.7-Bit Quantization of VLMs
为解决视觉语言模型在移动设备上的部署问题,提出了一种结合自生成训练数据的量化管道和2.7位参数格式的高效量化框架。
为解决视觉语言模型在移动设备上的部署问题,提出了一种结合自生成训练数据的量化管道和2.7位参数格式的高效量化框架。
为解决生物测定活性预测数据有限的问题,本文提出Monroe模型,通过扩大预训练规模、改进图表示等方法提高分子基础模型性能。
本文针对药物分子属性改进问题,提出PGFS++方法,在保证合成可行性和多样性的同时,通过强化学习和前向合成路径提高了目标分子的属性。
This study addresses the lack of systematic evaluation of how quantization impacts inference efficiency and translation quality under orchestrated workloads in real-world server deployments of large language models, particularly given that mainstream benchmarks overlook long-context dynamics. We present the first controlled assessment of EuroLLM and Hy-MT2 (1.7B–22B parameters) on A100/H100 GPUs, combining W4A8/W8A8 quantization with document chunking strategies. Introducing a document-level evaluation based on WMT24++, we uncover quantization–long-context interactions invisible to sentence-level benchmarks. Our experiments reveal that Hy-MT2 maintains robustness under quantization, whereas EuroLLM suffers significant quality degradation. Jointly optimizing chunking and quantization substantially improves the latency–throughput Pareto frontier, demonstrating that the trade-off between translation quality and efficiency is highly dependent on model architecture, quantization format, and chunking strategy.
This work investigates how to accelerate large language model inference by relaxing the strict fidelity constraints of conventional speculative decoding without compromising generation quality. We present the first systematic evaluation of various training-free relaxed speculative decoding strategies, unifying existing frameworks and benchmarking them on modern large models. Our analysis reveals that most relaxation methods heavily rely on the draft model’s language modeling capabilities and struggle to generalize to lightweight, specialized predictors. Nevertheless, well-designed relaxation mechanisms can achieve a controllable trade-off between speed and capability—and may even yield modest performance gains. This study distills practical insights for practitioners and underscores the critical role of draft model capability assessment in effective relaxed decoding.
为解决视觉语言模型在移动设备上的部署问题,提出了一种结合自生成训练数据的量化管道和2.7位参数格式的高效量化框架。
为解决生物测定活性预测数据有限的问题,本文提出Monroe模型,通过扩大预训练规模、改进图表示等方法提高分子基础模型性能。
本文针对药物分子属性改进问题,提出PGFS++方法,在保证合成可行性和多样性的同时,通过强化学习和前向合成路径提高了目标分子的属性。
This study addresses the lack of systematic evaluation of how quantization impacts inference efficiency and translation quality under orchestrated workloads in real-world server deployments of large language models, particularly given that mainstream benchmarks overlook long-context dynamics. We present the first controlled assessment of EuroLLM and Hy-MT2 (1.7B–22B parameters) on A100/H100 GPUs, combining W4A8/W8A8 quantization with document chunking strategies. Introducing a document-level evaluation based on WMT24++, we uncover quantization–long-context interactions invisible to sentence-level benchmarks. Our experiments reveal that Hy-MT2 maintains robustness under quantization, whereas EuroLLM suffers significant quality degradation. Jointly optimizing chunking and quantization substantially improves the latency–throughput Pareto frontier, demonstrating that the trade-off between translation quality and efficiency is highly dependent on model architecture, quantization format, and chunking strategy.
This work investigates how to accelerate large language model inference by relaxing the strict fidelity constraints of conventional speculative decoding without compromising generation quality. We present the first systematic evaluation of various training-free relaxed speculative decoding strategies, unifying existing frameworks and benchmarking them on modern large models. Our analysis reveals that most relaxation methods heavily rely on the draft model’s language modeling capabilities and struggle to generalize to lightweight, specialized predictors. Nevertheless, well-designed relaxation mechanisms can achieve a controllable trade-off between speed and capability—and may even yield modest performance gains. This study distills practical insights for practitioners and underscores the critical role of draft model capability assessment in effective relaxed decoding.