Institution profile

Cleanlab

Industry researchnorthamerica · us
Official website
Research library2linked papers
Opportunities0open roles
Selected work

Representative Papers

Real-Time Trustworthiness Scoring for LLM Structured Outputs and Data Extraction

Feb 24, 2026

This work addresses the challenge of sporadic errors in structured outputs generated by large language models (LLMs), which hinder their reliable deployment in enterprise settings. The authors propose CONSTRUCT, a method that estimates real-time confidence scores based on output uncertainty, enabling assessment of both overall and field-level reliability for any LLM—including black-box APIs without access to log probabilities—without requiring labeled data or model customization. CONSTRUCT supports heterogeneous fields and nested JSON structures. The study introduces the first public benchmark for structured generation with reliable ground-truth annotations. Evaluated across four datasets involving models such as Gemini 3 and GPT-5, CONSTRUCT significantly outperforms existing approaches in precision and recall for error detection.

0 citationsRead paper

A General Test for Independent and Identically Distributed Hypothesis

Jun 27, 2025

This paper addresses the fundamental problem of testing the independent and identically distributed (IID) assumption in statistical modeling, particularly for random objects residing in general metric spaces. We propose a universal nonparametric test based on a novel off-diagonal sequential U-process as the test statistic. Our key theoretical contributions include: (i) establishing the first Gaussian approximation for the supremum of this process, accompanied by a non-asymptotic coupling error bound; and (ii) developing a jackknife multiplier bootstrap procedure for accurate inference without specifying alternative hypotheses. The method is sensitive to diverse violations of IID—such as temporal dependence, spatial correlation, and distributional drift—while requiring no parametric assumptions. Extensive simulations and real-data applications demonstrate superior detection power and broader applicability compared to existing approaches. Overall, the proposed framework provides a rigorous, robust, and broadly applicable theoretical tool for IID validation in complex, high-dimensional, and non-Euclidean data settings.

0 citationsRead paper
Recent publications

Latest Papers

Real-Time Trustworthiness Scoring for LLM Structured Outputs and Data Extraction

Feb 24, 2026

This work addresses the challenge of sporadic errors in structured outputs generated by large language models (LLMs), which hinder their reliable deployment in enterprise settings. The authors propose CONSTRUCT, a method that estimates real-time confidence scores based on output uncertainty, enabling assessment of both overall and field-level reliability for any LLM—including black-box APIs without access to log probabilities—without requiring labeled data or model customization. CONSTRUCT supports heterogeneous fields and nested JSON structures. The study introduces the first public benchmark for structured generation with reliable ground-truth annotations. Evaluated across four datasets involving models such as Gemini 3 and GPT-5, CONSTRUCT significantly outperforms existing approaches in precision and recall for error detection.

0 citationsRead paper

A General Test for Independent and Identically Distributed Hypothesis

Jun 27, 2025

This paper addresses the fundamental problem of testing the independent and identically distributed (IID) assumption in statistical modeling, particularly for random objects residing in general metric spaces. We propose a universal nonparametric test based on a novel off-diagonal sequential U-process as the test statistic. Our key theoretical contributions include: (i) establishing the first Gaussian approximation for the supremum of this process, accompanied by a non-asymptotic coupling error bound; and (ii) developing a jackknife multiplier bootstrap procedure for accurate inference without specifying alternative hypotheses. The method is sensitive to diverse violations of IID—such as temporal dependence, spatial correlation, and distributional drift—while requiring no parametric assumptions. Extensive simulations and real-data applications demonstrate superior detection power and broader applicability compared to existing approaches. Overall, the proposed framework provides a rigorous, robust, and broadly applicable theoretical tool for IID validation in complex, high-dimensional, and non-Euclidean data settings.

0 citationsRead paper