Institution profile

DatologyAI

Industry research
Official website
Research library3linked papers
Opportunities0open roles
Selected work

Representative Papers

DatBench: Discriminative, Faithful, and Efficient VLM Evaluations

Jan 05, 2026arXiv.org

Current evaluation methods for vision-language models (VLMs) commonly suffer from modality unfaithfulness, insufficient discriminative power, and computational inefficiency. This work is the first to systematically articulate three core desiderata for VLM evaluation: faithfulness, discriminability, and efficiency. We establish a high-quality evaluation pipeline by reformulating multiple-choice tasks as generative ones, filtering out samples amenable to blind guessing (up to 70% of instances), and correcting mislabeled examples (42% of cases). Based on this framework, we introduce DatBench-Full, encompassing 33 datasets, along with its highly discriminative subset, DatBench. Our benchmarks maintain discriminative capacity comparable to original benchmarks while achieving an average 13× and up to 50× acceleration in evaluation speed.

1 citationsRead paper

Luxical: High-Speed Lexical-Dense Text Embeddings

Dec 09, 2025

To address the trade-off between speed and flexibility in web-scale text organization, this paper proposes a lightweight “lexical-dense” embedding paradigm. Leveraging knowledge distillation, it transfers semantic capabilities from large language models into a compact architecture that jointly encodes TF-IDF–based sparse lexical features and dense representations via a small ReLU network. The resulting embeddings retain the versatility of dense vectors—supporting retrieval, clustering, classification, and data cleaning—while achieving inference speeds comparable to FastText. Experiments demonstrate 3×–100× higher throughput than neural baselines on document retrieval and LLM data cleaning tasks, with quality matching state-of-the-art embedding models. This work is the first to systematically bridge the efficiency–expressiveness gap between traditional lexical models and Transformer-based embeddings, establishing a scalable, high-performance paradigm for large-scale text preprocessing.

0 citationsRead paper

BeyondWeb: Lessons from Scaling Synthetic Data for Trillion-scale Pretraining

Aug 14, 2025

To address the performance saturation in large language model (LLM) pretraining caused by data bottlenecks, this work proposes BeyondWeb—a framework that systematically investigates the joint impact of model scale, architecture family, and data rewriting strategies on synthetic data quality. It establishes a high-quality synthetic data generation paradigm integrating model feedback, rewrite filtering, and diversity control. Evaluated on trillion-token pretraining, BeyondWeb significantly enhances semantic richness and training efficiency of synthetic data. Experiments show it achieves +5.1 and +2.6 percentage points average accuracy over Cosmopedia and Nemotron-Synth across 14 benchmarks; accelerates training up to 7.7×; and enables a 3B model trained on 180B tokens to outperform an 8B baseline. This is the first work to realize a multi-factor co-optimized synthetic data generation system, providing a reproducible and scalable pathway to overcome pretraining data constraints.

0 citationsRead paper
Recent publications

Latest Papers

DatBench: Discriminative, Faithful, and Efficient VLM Evaluations

Jan 05, 2026arXiv.org

Current evaluation methods for vision-language models (VLMs) commonly suffer from modality unfaithfulness, insufficient discriminative power, and computational inefficiency. This work is the first to systematically articulate three core desiderata for VLM evaluation: faithfulness, discriminability, and efficiency. We establish a high-quality evaluation pipeline by reformulating multiple-choice tasks as generative ones, filtering out samples amenable to blind guessing (up to 70% of instances), and correcting mislabeled examples (42% of cases). Based on this framework, we introduce DatBench-Full, encompassing 33 datasets, along with its highly discriminative subset, DatBench. Our benchmarks maintain discriminative capacity comparable to original benchmarks while achieving an average 13× and up to 50× acceleration in evaluation speed.

1 citationsRead paper

Luxical: High-Speed Lexical-Dense Text Embeddings

Dec 09, 2025

To address the trade-off between speed and flexibility in web-scale text organization, this paper proposes a lightweight “lexical-dense” embedding paradigm. Leveraging knowledge distillation, it transfers semantic capabilities from large language models into a compact architecture that jointly encodes TF-IDF–based sparse lexical features and dense representations via a small ReLU network. The resulting embeddings retain the versatility of dense vectors—supporting retrieval, clustering, classification, and data cleaning—while achieving inference speeds comparable to FastText. Experiments demonstrate 3×–100× higher throughput than neural baselines on document retrieval and LLM data cleaning tasks, with quality matching state-of-the-art embedding models. This work is the first to systematically bridge the efficiency–expressiveness gap between traditional lexical models and Transformer-based embeddings, establishing a scalable, high-performance paradigm for large-scale text preprocessing.

0 citationsRead paper

BeyondWeb: Lessons from Scaling Synthetic Data for Trillion-scale Pretraining

Aug 14, 2025

To address the performance saturation in large language model (LLM) pretraining caused by data bottlenecks, this work proposes BeyondWeb—a framework that systematically investigates the joint impact of model scale, architecture family, and data rewriting strategies on synthetic data quality. It establishes a high-quality synthetic data generation paradigm integrating model feedback, rewrite filtering, and diversity control. Evaluated on trillion-token pretraining, BeyondWeb significantly enhances semantic richness and training efficiency of synthetic data. Experiments show it achieves +5.1 and +2.6 percentage points average accuracy over Cosmopedia and Nemotron-Synth across 14 benchmarks; accelerates training up to 7.7×; and enables a 3B model trained on 180B tokens to outperform an 8B baseline. This is the first work to realize a multi-factor co-optimized synthetic data generation system, providing a reproducible and scalable pathway to overcome pretraining data constraints.

0 citationsRead paper