Institution profile

x-humanoid

Industry researcheurope · fr
Official website
Research library17linked papers
Opportunities0open roles
Selected work

Representative Papers

Thinking Once Is Enough: Intermediate-Layer Evidence Routing for High-Resolution VQA

Jul 30, 2026

This work addresses the inefficiency of existing high-resolution visual question answering (HR-VQA) methods, which suffer from redundant image cropping or re-encoding and neglect the dilution of fine-grained intermediate evidence in subsequent processing. The authors propose Thinking-Once, a training-free, single-pass framework that preserves critical entities and compact background context through question-conditioned attention reshaping and token selection within a single visual forward pass. Crucially, it routes intermediate-layer evidence directly to higher layers without additional training or repeated visual processing. This approach reveals, for the first time, that the performance bottleneck in HR-VQA stems from evidence dilution rather than insufficient input resolution. Evaluated across five multimodal large language models, Thinking-Once improves average scores by 3.1, 3.0, and 2.7 points on V*Bench, HRBench-4K, and HRBench-8K, respectively, reduces peak memory usage by approximately 4 GB, accelerates inference by 97.2% over DeepScan, and achieves an average cross-benchmark score of 82.7.

0 citationsRead paper

Group-of-Latents: Perceptual Video Compression at Extreme Bitrates via Masked Latent Generative Modeling

Jul 21, 2026

This work addresses the challenge of achieving both high perceptual quality and temporal consistency in video compression at extremely low bitrates (<0.005 bpp), where existing methods struggle. The authors propose a unified generative framework that introduces a causal tokenizer to decompose latent representations into I-latents and P-latents, and employs a Group-of-Latents strategy for structured modeling. Key latents are efficiently encoded via an I-frame Deep Compression Module (I-DCM), while a unified latent denoising module (U-LDM), built upon a pre-trained Diffusion Transformer (DiT), reconstructs high-fidelity intra-frame textures and coherent temporal dynamics directly from noise. Notably, this approach incurs no additional bitrate overhead and significantly outperforms current state-of-the-art methods under extreme low-bitrate constraints, delivering spatially detailed and temporally stable visual quality.

0 citationsRead paper

GeoMathCode: Understanding Interleaved Math-Code Reasoning for Geometry Problem Solving

May 24, 2026

This work addresses the challenge of simulating human-like multi-step logical reasoning with auxiliary constructions in geometric problem solving by proposing a novel framework that integrates mathematical reasoning with procedural representations. The approach employs program code as an intermediate visual representation, decoupling discovery reasoning from code generation in a latent space and structuring the reasoning manifold through supervised fine-tuning. The study demonstrates that hierarchical syntactic code structures effectively encode rich mathematical semantics, offering greater expressiveness than purely visual representations. Experimental results show that the proposed method significantly enhances geometric reasoning performance while yielding clearer and more interpretable multi-step derivations.

0 citationsRead paper
Recent publications

Latest Papers

Thinking Once Is Enough: Intermediate-Layer Evidence Routing for High-Resolution VQA

Jul 30, 2026

This work addresses the inefficiency of existing high-resolution visual question answering (HR-VQA) methods, which suffer from redundant image cropping or re-encoding and neglect the dilution of fine-grained intermediate evidence in subsequent processing. The authors propose Thinking-Once, a training-free, single-pass framework that preserves critical entities and compact background context through question-conditioned attention reshaping and token selection within a single visual forward pass. Crucially, it routes intermediate-layer evidence directly to higher layers without additional training or repeated visual processing. This approach reveals, for the first time, that the performance bottleneck in HR-VQA stems from evidence dilution rather than insufficient input resolution. Evaluated across five multimodal large language models, Thinking-Once improves average scores by 3.1, 3.0, and 2.7 points on V*Bench, HRBench-4K, and HRBench-8K, respectively, reduces peak memory usage by approximately 4 GB, accelerates inference by 97.2% over DeepScan, and achieves an average cross-benchmark score of 82.7.

0 citationsRead paper

Group-of-Latents: Perceptual Video Compression at Extreme Bitrates via Masked Latent Generative Modeling

Jul 21, 2026

This work addresses the challenge of achieving both high perceptual quality and temporal consistency in video compression at extremely low bitrates (<0.005 bpp), where existing methods struggle. The authors propose a unified generative framework that introduces a causal tokenizer to decompose latent representations into I-latents and P-latents, and employs a Group-of-Latents strategy for structured modeling. Key latents are efficiently encoded via an I-frame Deep Compression Module (I-DCM), while a unified latent denoising module (U-LDM), built upon a pre-trained Diffusion Transformer (DiT), reconstructs high-fidelity intra-frame textures and coherent temporal dynamics directly from noise. Notably, this approach incurs no additional bitrate overhead and significantly outperforms current state-of-the-art methods under extreme low-bitrate constraints, delivering spatially detailed and temporally stable visual quality.

0 citationsRead paper

GeoMathCode: Understanding Interleaved Math-Code Reasoning for Geometry Problem Solving

May 24, 2026

This work addresses the challenge of simulating human-like multi-step logical reasoning with auxiliary constructions in geometric problem solving by proposing a novel framework that integrates mathematical reasoning with procedural representations. The approach employs program code as an intermediate visual representation, decoupling discovery reasoning from code generation in a latent space and structuring the reasoning manifold through supervised fine-tuning. The study demonstrates that hierarchical syntactic code structures effectively encode rich mathematical semantics, offering greater expressiveness than purely visual representations. Experimental results show that the proposed method significantly enhances geometric reasoning performance while yielding clearer and more interpretable multi-step derivations.

0 citationsRead paper