Institution profile

Pluralis Research

Research institution
Research library13linked papers
Opportunities0open roles
Selected work

Representative Papers

The Gate Always Closes: On Injecting Auxiliary Signals into Frozen Vision-Language Models

Jul 25, 2026

This work addresses the issue in frozen vision-language models where learnable gating mechanisms are inadvertently disabled by optimizers during training, thereby collapsing auxiliary signal pathways. The authors identify gradient vanishing and negative utility as the underlying causes and propose a fixed-scale injection strategy that eliminates the need for learnable gates. Innovatively integrating entailment cones and angular repulsion on hyperbolic manifolds into LoRA fine-tuning, they regularize the model with a geometric auxiliary loss. During inference, the geometric pathway is retained, preserving relational question-answering accuracy while enhancing attribute-based performance. Notably, on out-of-distribution Visual Spatial Relation (VSR) tasks, the geometric loss stabilizes spatial signals—its removal leads to a 4.6 percentage point drop in performance.

0 citationsRead paper

Agora: Collective and Permissionless Internet-Scale Pretraining of Large Language Models

Jul 14, 2026

This work addresses the longstanding limitation of large language model (LLM) training, which has been confined to high-performance data centers and unable to leverage the vast pool of heterogeneous, unreliable consumer-grade GPUs available across the internet. To overcome this, the authors introduce Agora, a novel system that pioneers a “protocol learning” paradigm, enabling permissionless, decentralized collective pretraining through communication-efficient pipeline parallelism, asynchronous optimization, and robust fault tolerance. In this framework, no participant holds the full model; instead, each retains only a shard of the parameters, ensuring collective ownership and inherent openness. The system successfully trained Pluralis-8B (8.6 billion parameters) on 500 billion FineWeb-Edu tokens using 330 dynamically joining and leaving consumer GPUs over 40 days, achieving 63% of the computational efficiency of a centralized baseline while matching its convergence performance closely.

0 citationsRead paper

Factored Gossip DiLoCo: Reducing Blocking Communication in DiLoCo

Jun 21, 2026

This work addresses the high communication overhead of DiLoCo in low-bandwidth environments and its sensitivity to stragglers and transient failures during large-scale distributed training. To mitigate these issues, the authors propose an approximate synchronization mechanism that, for the first time, integrates gossip-based hybrid communication into the DiLoCo framework. The synchronization process is factorized into two phases: a non-blocking phase that overlaps with computation to improve resource utilization, and a blocking phase that enhances inter-node consistency to ensure training stability. Evaluated on billion-parameter language model training, the proposed method achieves computational efficiency significantly higher than the original DiLoCo while maintaining comparable training progress and demonstrating markedly improved fault tolerance.

0 citationsRead paper

Mixtures of Subspaces for Bandwidth Efficient Context Parallel Training

Jun 15, 2026

This work addresses the high communication overhead in parallel training of large language models with long contexts under decentralized, low-bandwidth environments. The authors propose a dynamic subspace hybrid-based low-rank reparameterization method that leverages the intrinsic low-rank structure of activation outputs to enable highly efficient communication compression in context-parallel training. Without compromising convergence speed, the approach achieves over 95% communication compression, enabling successful training of billion-parameter models with context lengths exceeding 100,000 tokens on a 300 Mbps network. Remarkably, the resulting model performance matches that attained on a 100 Gbps high-speed cluster, demonstrating the method’s effectiveness in drastically reducing bandwidth requirements while maintaining scalability and training efficiency.

0 citationsRead paper
Recent publications

Latest Papers

The Gate Always Closes: On Injecting Auxiliary Signals into Frozen Vision-Language Models

Jul 25, 2026

This work addresses the issue in frozen vision-language models where learnable gating mechanisms are inadvertently disabled by optimizers during training, thereby collapsing auxiliary signal pathways. The authors identify gradient vanishing and negative utility as the underlying causes and propose a fixed-scale injection strategy that eliminates the need for learnable gates. Innovatively integrating entailment cones and angular repulsion on hyperbolic manifolds into LoRA fine-tuning, they regularize the model with a geometric auxiliary loss. During inference, the geometric pathway is retained, preserving relational question-answering accuracy while enhancing attribute-based performance. Notably, on out-of-distribution Visual Spatial Relation (VSR) tasks, the geometric loss stabilizes spatial signals—its removal leads to a 4.6 percentage point drop in performance.

0 citationsRead paper

Agora: Collective and Permissionless Internet-Scale Pretraining of Large Language Models

Jul 14, 2026

This work addresses the longstanding limitation of large language model (LLM) training, which has been confined to high-performance data centers and unable to leverage the vast pool of heterogeneous, unreliable consumer-grade GPUs available across the internet. To overcome this, the authors introduce Agora, a novel system that pioneers a “protocol learning” paradigm, enabling permissionless, decentralized collective pretraining through communication-efficient pipeline parallelism, asynchronous optimization, and robust fault tolerance. In this framework, no participant holds the full model; instead, each retains only a shard of the parameters, ensuring collective ownership and inherent openness. The system successfully trained Pluralis-8B (8.6 billion parameters) on 500 billion FineWeb-Edu tokens using 330 dynamically joining and leaving consumer GPUs over 40 days, achieving 63% of the computational efficiency of a centralized baseline while matching its convergence performance closely.

0 citationsRead paper

Factored Gossip DiLoCo: Reducing Blocking Communication in DiLoCo

Jun 21, 2026

This work addresses the high communication overhead of DiLoCo in low-bandwidth environments and its sensitivity to stragglers and transient failures during large-scale distributed training. To mitigate these issues, the authors propose an approximate synchronization mechanism that, for the first time, integrates gossip-based hybrid communication into the DiLoCo framework. The synchronization process is factorized into two phases: a non-blocking phase that overlaps with computation to improve resource utilization, and a blocking phase that enhances inter-node consistency to ensure training stability. Evaluated on billion-parameter language model training, the proposed method achieves computational efficiency significantly higher than the original DiLoCo while maintaining comparable training progress and demonstrating markedly improved fault tolerance.

0 citationsRead paper

Mixtures of Subspaces for Bandwidth Efficient Context Parallel Training

Jun 15, 2026

This work addresses the high communication overhead in parallel training of large language models with long contexts under decentralized, low-bandwidth environments. The authors propose a dynamic subspace hybrid-based low-rank reparameterization method that leverages the intrinsic low-rank structure of activation outputs to enable highly efficient communication compression in context-parallel training. Without compromising convergence speed, the approach achieves over 95% communication compression, enabling successful training of billion-parameter models with context lengths exceeding 100,000 tokens on a 300 Mbps network. Remarkably, the resulting model performance matches that attained on a 100 Gbps high-speed cluster, demonstrating the method’s effectiveness in drastically reducing bandwidth requirements while maintaining scalability and training efficiency.

0 citationsRead paper