Institution profile

Intercom

Industry researchnorthamerica · us
Official website
Research library2linked papers
Opportunities0open roles
Selected work

Representative Papers

Tiny Recursive Reasoning with Mamba-2 Attention Hybrid

Feb 12, 2026

This work addresses the bottleneck in abstract reasoning capabilities of recursive reasoning models under parameter constraints by proposing a novel architecture that replaces the Transformer module in TRM with a Mamba-2 hybrid operator. The resulting model integrates state-space mechanisms with attention within an implicit recursive reasoning framework, maintaining a comparable parameter count (6.8M). This study presents the first validation of Mamba-2’s effectiveness in recursive reasoning, thereby expanding the design space for recursive operators and significantly improving candidate solution coverage. On the ARC-AGI-1 dataset, the model preserves pass@1 performance while achieving a 2.0% gain in pass@2 and a 4.75% improvement in pass@100, demonstrating substantially enhanced stability and diversity in generating correct solutions.

0 citationsRead paper

Low-Rank Key Value Attention

Jan 16, 2026

This work addresses the substantial memory and computational burden imposed by key-value (KV) caching in Transformer pretraining, which has become a bottleneck for both training and autoregressive decoding. The authors propose Low-Rank KV Adaptation (LRKV), a method that shares full-rank KV projections across attention heads while introducing head-specific low-rank residual components. This approach significantly compresses the KV cache without sacrificing token-level resolution or inter-head diversity. LRKV establishes a continuous trade-off between fully shared and fully independent attention mechanisms, subsuming query-sharing strategies such as MQA and GQA within a unified framework, and differs fundamentally from latent-variable-based compression methods like MLA. Evaluated on a 2.5B-parameter model, LRKV reduces KV cache size by approximately 50%, decreases training FLOPs by 20–25%, and achieves faster convergence, lower validation perplexity, and improved downstream performance.

0 citationsRead paper
Recent publications

Latest Papers

Tiny Recursive Reasoning with Mamba-2 Attention Hybrid

Feb 12, 2026

This work addresses the bottleneck in abstract reasoning capabilities of recursive reasoning models under parameter constraints by proposing a novel architecture that replaces the Transformer module in TRM with a Mamba-2 hybrid operator. The resulting model integrates state-space mechanisms with attention within an implicit recursive reasoning framework, maintaining a comparable parameter count (6.8M). This study presents the first validation of Mamba-2’s effectiveness in recursive reasoning, thereby expanding the design space for recursive operators and significantly improving candidate solution coverage. On the ARC-AGI-1 dataset, the model preserves pass@1 performance while achieving a 2.0% gain in pass@2 and a 4.75% improvement in pass@100, demonstrating substantially enhanced stability and diversity in generating correct solutions.

0 citationsRead paper

Low-Rank Key Value Attention

Jan 16, 2026

This work addresses the substantial memory and computational burden imposed by key-value (KV) caching in Transformer pretraining, which has become a bottleneck for both training and autoregressive decoding. The authors propose Low-Rank KV Adaptation (LRKV), a method that shares full-rank KV projections across attention heads while introducing head-specific low-rank residual components. This approach significantly compresses the KV cache without sacrificing token-level resolution or inter-head diversity. LRKV establishes a continuous trade-off between fully shared and fully independent attention mechanisms, subsuming query-sharing strategies such as MQA and GQA within a unified framework, and differs fundamentally from latent-variable-based compression methods like MLA. Evaluated on a 2.5B-parameter model, LRKV reduces KV cache size by approximately 50%, decreases training FLOPs by 20–25%, and achieves faster convergence, lower validation perplexity, and improved downstream performance.

0 citationsRead paper