Institution profile

Graphcore

Industry researcheurope · gb
Official website
Research library15linked papers
Opportunities0open roles
Selected work

Representative Papers

Studying quantization trade-offs for efficient inference deployment in machine translation

Jul 31, 2026

This study addresses the lack of systematic evaluation of how quantization impacts inference efficiency and translation quality under orchestrated workloads in real-world server deployments of large language models, particularly given that mainstream benchmarks overlook long-context dynamics. We present the first controlled assessment of EuroLLM and Hy-MT2 (1.7B–22B parameters) on A100/H100 GPUs, combining W4A8/W8A8 quantization with document chunking strategies. Introducing a document-level evaluation based on WMT24++, we uncover quantization–long-context interactions invisible to sentence-level benchmarks. Our experiments reveal that Hy-MT2 maintains robustness under quantization, whereas EuroLLM suffers significant quality degradation. Jointly optimizing chunking and quantization substantially improves the latency–throughput Pareto frontier, demonstrating that the trade-off between translation quality and efficiency is highly dependent on model architecture, quantization format, and chunking strategy.

0 citationsRead paper

A Practical Investigation of Training-free Relaxed Speculative Decoding

Jul 09, 2026

This work investigates how to accelerate large language model inference by relaxing the strict fidelity constraints of conventional speculative decoding without compromising generation quality. We present the first systematic evaluation of various training-free relaxed speculative decoding strategies, unifying existing frameworks and benchmarking them on modern large models. Our analysis reveals that most relaxation methods heavily rely on the draft model’s language modeling capabilities and struggle to generalize to lightweight, specialized predictors. Nevertheless, well-designed relaxation mechanisms can achieve a controllable trade-off between speed and capability—and may even yield modest performance gains. This study distills practical insights for practitioners and underscores the critical role of draft model capability assessment in effective relaxed decoding.

0 citationsRead paper
Recent publications

Latest Papers

Studying quantization trade-offs for efficient inference deployment in machine translation

Jul 31, 2026

This study addresses the lack of systematic evaluation of how quantization impacts inference efficiency and translation quality under orchestrated workloads in real-world server deployments of large language models, particularly given that mainstream benchmarks overlook long-context dynamics. We present the first controlled assessment of EuroLLM and Hy-MT2 (1.7B–22B parameters) on A100/H100 GPUs, combining W4A8/W8A8 quantization with document chunking strategies. Introducing a document-level evaluation based on WMT24++, we uncover quantization–long-context interactions invisible to sentence-level benchmarks. Our experiments reveal that Hy-MT2 maintains robustness under quantization, whereas EuroLLM suffers significant quality degradation. Jointly optimizing chunking and quantization substantially improves the latency–throughput Pareto frontier, demonstrating that the trade-off between translation quality and efficiency is highly dependent on model architecture, quantization format, and chunking strategy.

0 citationsRead paper

A Practical Investigation of Training-free Relaxed Speculative Decoding

Jul 09, 2026

This work investigates how to accelerate large language model inference by relaxing the strict fidelity constraints of conventional speculative decoding without compromising generation quality. We present the first systematic evaluation of various training-free relaxed speculative decoding strategies, unifying existing frameworks and benchmarking them on modern large models. Our analysis reveals that most relaxation methods heavily rely on the draft model’s language modeling capabilities and struggle to generalize to lightweight, specialized predictors. Nevertheless, well-designed relaxation mechanisms can achieve a controllable trade-off between speed and capability—and may even yield modest performance gains. This study distills practical insights for practitioners and underscores the critical role of draft model capability assessment in effective relaxed decoding.

0 citationsRead paper