Institution profile

Tabularis.AI

Industry research
Official website
Research library2linked papers
Opportunities0open roles
Selected work

Representative Papers

Do Chatbot LLMs Talk Too Much? The YapBench Benchmark

Jan 02, 2026arXiv.org

Large language models frequently produce verbose and redundant responses even to simple user requests, imposing unnecessary cognitive load. To address this issue, this work proposes YapBench—a lightweight benchmark that evaluates model over-generation through idealized prompts across three concise scenarios. The study introduces YapScore and YapIndex, novel character-level metrics that do not rely on tokenizers, and constructs an evaluation dataset based on human-annotated minimal sufficient answers. Redundancy is quantified by the character-length difference between model outputs and these minimal answers, enabling consistent cross-model comparisons. Evaluation of 76 assistant models reveals that median redundancy lengths differ by nearly an order of magnitude and uncovers characteristic over-generation patterns, particularly in ambiguous inputs and single-line code tasks.

1 citationsRead paper

Multilingual Multi-Label Emotion Classification at Scale with Synthetic Data

Apr 14, 2026

This study addresses the scarcity of annotated data in multilingual emotion classification, where existing corpora are predominantly English, single-label, and limited in linguistic coverage. The authors propose a culturally adapted generative approach combined with procedural quality filtering to construct, for the first time, a million-scale multilingual, multi-label synthetic dataset spanning 23 languages and 11 emotion categories—without requiring manual annotation. Multilingual Transformer models (from DistilBERT to XLM-R-Large) trained on this dataset achieve strong in-domain performance, with 0.868 F1-micro and 0.987 AUC-micro. In zero-shot transfer evaluations on GoEmotions and SemEval-2018 benchmarks, the models attain an AUC-micro of 0.810, rivaling or surpassing English-only counterparts. The best-performing base model has been publicly released.

0 citationsRead paper
Recent publications

Latest Papers

Multilingual Multi-Label Emotion Classification at Scale with Synthetic Data

Apr 14, 2026

This study addresses the scarcity of annotated data in multilingual emotion classification, where existing corpora are predominantly English, single-label, and limited in linguistic coverage. The authors propose a culturally adapted generative approach combined with procedural quality filtering to construct, for the first time, a million-scale multilingual, multi-label synthetic dataset spanning 23 languages and 11 emotion categories—without requiring manual annotation. Multilingual Transformer models (from DistilBERT to XLM-R-Large) trained on this dataset achieve strong in-domain performance, with 0.868 F1-micro and 0.987 AUC-micro. In zero-shot transfer evaluations on GoEmotions and SemEval-2018 benchmarks, the models attain an AUC-micro of 0.810, rivaling or surpassing English-only counterparts. The best-performing base model has been publicly released.

0 citationsRead paper

Do Chatbot LLMs Talk Too Much? The YapBench Benchmark

Jan 02, 2026arXiv.org

Large language models frequently produce verbose and redundant responses even to simple user requests, imposing unnecessary cognitive load. To address this issue, this work proposes YapBench—a lightweight benchmark that evaluates model over-generation through idealized prompts across three concise scenarios. The study introduces YapScore and YapIndex, novel character-level metrics that do not rely on tokenizers, and constructs an evaluation dataset based on human-annotated minimal sufficient answers. Redundancy is quantified by the character-length difference between model outputs and these minimal answers, enabling consistent cross-model comparisons. Evaluation of 76 assistant models reveals that median redundancy lengths differ by nearly an order of magnitude and uncovers characteristic over-generation patterns, particularly in ambiguous inputs and single-line code tasks.

1 citationsRead paper