Institution profile

Merantix-Momentum

Industry researcheurope · de
Official website
Research library3linked papers
Opportunities0open roles
Selected work

Representative Papers

Asymptotic e-processes

Apr 21, 2026

This work addresses the challenge in sequential hypothesis testing where model misspecification or estimation error prevents exact construction of e-variables, thereby lacking finite-sample guarantees. We introduce, for the first time, the notion of an asymptotic e-process, defined as a doubly indexed stochastic process $(E_{m,n})$, whose limiting behavior as $m \to \infty$ approximates a standard e-process. We establish its connection to asymptotic supermartingales, derive a corresponding variant of Ville’s inequality, and provide practical construction methods. This framework unifies the theoretical foundation for approximate e-variables, offering sequential inference guarantees under controllable approximation error and explicitly quantifying the trade-off between approximation accuracy and the effective monitoring horizon $r_m$.

0 citationsRead paper

Float8@2bits: Entropy Coding Enables Data-Free Model Compression

Jan 30, 2026

This work addresses the challenge of post-training model compression at extremely low bitrates (<4 bits), where existing data-free methods often suffer from catastrophic performance degradation, while calibration-based approaches incur high computational costs and are sensitive to distribution shifts. The authors propose EntQuant, a novel framework that uniquely integrates entropy coding with floating-point quantization (e.g., Float8@2bit), decoupling numerical precision from storage cost to enable efficient, calibration-free compression. EntQuant achieves both the universality of data-free methods and the high fidelity typically associated with data-dependent techniques. It compresses 70B-parameter models within 30 minutes while attaining state-of-the-art accuracy and preserving strong functional capabilities on complex instruction-tuned models, all with manageable inference overhead.

0 citationsRead paper

Choose Your Model Size: Any Compression by a Single Gradient Descent

Feb 03, 2025

To address the challenge of deploying foundation models under resource constraints, this paper proposes ACIP—a novel algorithm that generates a global parameter importance ranking via a single SGD pass, enabling zero-shot, on-the-fly instantiation of compressed models at arbitrary target sizes without fine-tuning. Methodologically, ACIP integrates SVD-based reparameterization, iterative singular-value pruning, and sparsity-inducing regularization to achieve efficient structured pruning. Unlike conventional compression paradigms requiring multiple training cycles or post-pruning fine-tuning, ACIP drastically reduces training overhead. Evaluated on multiple open-source LLMs, it achieves state-of-the-art compression performance—outperforming mainstream factorization-based methods—and natively supports quantization. Thus, ACIP establishes a new paradigm for lightweight deployment of large language models.

0 citationsRead paper
Recent publications

Latest Papers

Asymptotic e-processes

Apr 21, 2026

This work addresses the challenge in sequential hypothesis testing where model misspecification or estimation error prevents exact construction of e-variables, thereby lacking finite-sample guarantees. We introduce, for the first time, the notion of an asymptotic e-process, defined as a doubly indexed stochastic process $(E_{m,n})$, whose limiting behavior as $m \to \infty$ approximates a standard e-process. We establish its connection to asymptotic supermartingales, derive a corresponding variant of Ville’s inequality, and provide practical construction methods. This framework unifies the theoretical foundation for approximate e-variables, offering sequential inference guarantees under controllable approximation error and explicitly quantifying the trade-off between approximation accuracy and the effective monitoring horizon $r_m$.

0 citationsRead paper

Float8@2bits: Entropy Coding Enables Data-Free Model Compression

Jan 30, 2026

This work addresses the challenge of post-training model compression at extremely low bitrates (<4 bits), where existing data-free methods often suffer from catastrophic performance degradation, while calibration-based approaches incur high computational costs and are sensitive to distribution shifts. The authors propose EntQuant, a novel framework that uniquely integrates entropy coding with floating-point quantization (e.g., Float8@2bit), decoupling numerical precision from storage cost to enable efficient, calibration-free compression. EntQuant achieves both the universality of data-free methods and the high fidelity typically associated with data-dependent techniques. It compresses 70B-parameter models within 30 minutes while attaining state-of-the-art accuracy and preserving strong functional capabilities on complex instruction-tuned models, all with manageable inference overhead.

0 citationsRead paper

Choose Your Model Size: Any Compression by a Single Gradient Descent

Feb 03, 2025

To address the challenge of deploying foundation models under resource constraints, this paper proposes ACIP—a novel algorithm that generates a global parameter importance ranking via a single SGD pass, enabling zero-shot, on-the-fly instantiation of compressed models at arbitrary target sizes without fine-tuning. Methodologically, ACIP integrates SVD-based reparameterization, iterative singular-value pruning, and sparsity-inducing regularization to achieve efficient structured pruning. Unlike conventional compression paradigms requiring multiple training cycles or post-pruning fine-tuning, ACIP drastically reduces training overhead. Evaluated on multiple open-source LLMs, it achieves state-of-the-art compression performance—outperforming mainstream factorization-based methods—and natively supports quantization. Thus, ACIP establishes a new paradigm for lightweight deployment of large language models.

0 citationsRead paper