Institution profile

Abacus.AI

Industry researchnorthamerica · us
Official website
Research library2linked papers
Opportunities0open roles
Selected work

Representative Papers

DSR-Bench: Evaluating the Structural Reasoning Abilities of LLMs via Data Structures

May 29, 2025

Existing benchmarks lack fine-grained evaluation of large language models’ (LLMs) structural reasoning capabilities at the data structure level. To address this, we propose DSR-Bench—the first automated, data-structure-centric benchmark—comprising 20 data structures, 35 operation types, and 4,140 synthetically generated questions. It establishes a hierarchical, fully automated, and subjectivity-free evaluation paradigm grounded in data structures. Leveraging structured prompt engineering, deterministic programmatic assessment, and multidimensional capability decomposition, we evaluate nine state-of-the-art models. Our analysis uncovers fundamental limitations in multi-attribute, multi-hop, and hybrid-structure reasoning: instruction-tuned models exhibit weak foundational structural reasoning; inference-optimized models achieve at most 47% accuracy on challenging subsets; and performance degrades significantly on tasks involving multidimensional data and natural-language descriptions.

0 citationsRead paper

Transformers Boost the Performance of Decision Trees on Tabular Data across Sample Sizes

Feb 04, 2025

This paper addresses the imbalanced generalization capability of tabular models across small- and large-sample regimes. We propose two lightweight ensemble frameworks—LLM-Boost and PFN-Boost—that synergistically integrate pretrained priors from large language models (LLMs) and TabPFN with the scalability of gradient-boosted decision trees (GBDTs), marking the first such integration for tabular data. Crucially, our approach incurs no additional training overhead: it fuses contextual learning and Transformer-based representations at the feature level to enhance GBDT’s adaptability to multi-scale tabular data. Experiments show that PFN-Boost achieves state-of-the-art average performance across all dataset sizes except the smallest; LLM-Boost consistently outperforms standalone LLM, TabPFN, and GBDT baselines on medium-sized datasets. Our core contribution is the novel bridging of pretrained foundation models with tree-based learners for tabular learning—enabling scalable, sample-robust modeling without architectural or training-cost penalties.

0 citationsRead paper
Recent publications

Latest Papers

DSR-Bench: Evaluating the Structural Reasoning Abilities of LLMs via Data Structures

May 29, 2025

Existing benchmarks lack fine-grained evaluation of large language models’ (LLMs) structural reasoning capabilities at the data structure level. To address this, we propose DSR-Bench—the first automated, data-structure-centric benchmark—comprising 20 data structures, 35 operation types, and 4,140 synthetically generated questions. It establishes a hierarchical, fully automated, and subjectivity-free evaluation paradigm grounded in data structures. Leveraging structured prompt engineering, deterministic programmatic assessment, and multidimensional capability decomposition, we evaluate nine state-of-the-art models. Our analysis uncovers fundamental limitations in multi-attribute, multi-hop, and hybrid-structure reasoning: instruction-tuned models exhibit weak foundational structural reasoning; inference-optimized models achieve at most 47% accuracy on challenging subsets; and performance degrades significantly on tasks involving multidimensional data and natural-language descriptions.

0 citationsRead paper

Transformers Boost the Performance of Decision Trees on Tabular Data across Sample Sizes

Feb 04, 2025

This paper addresses the imbalanced generalization capability of tabular models across small- and large-sample regimes. We propose two lightweight ensemble frameworks—LLM-Boost and PFN-Boost—that synergistically integrate pretrained priors from large language models (LLMs) and TabPFN with the scalability of gradient-boosted decision trees (GBDTs), marking the first such integration for tabular data. Crucially, our approach incurs no additional training overhead: it fuses contextual learning and Transformer-based representations at the feature level to enhance GBDT’s adaptability to multi-scale tabular data. Experiments show that PFN-Boost achieves state-of-the-art average performance across all dataset sizes except the smallest; LLM-Boost consistently outperforms standalone LLM, TabPFN, and GBDT baselines on medium-sized datasets. Our core contribution is the novel bridging of pretrained foundation models with tree-based learners for tabular learning—enabling scalable, sample-robust modeling without architectural or training-cost penalties.

0 citationsRead paper