Institution profile

Gretel

Industry researchnorthamerica · us
Official website
Research library2linked papers
Opportunities0open roles
Selected work

Representative Papers

TinySQL: A Progressive Text-to-SQL Dataset for Mechanistic Interpretability Research

Mar 17, 2025

Mechanism interpretability research has long been hindered by the gap between synthetic “toy tasks” and real-world complexity. This paper proposes text-to-SQL generation as an ideal benchmark task—offering both formal syntactic structure and realistic semantic and compositional challenges. To this end, we introduce TinySQL, a progressively scaled synthetic dataset covering foundational to advanced SQL operations, and establish a dedicated evaluation platform spanning 33M–1B parameter models. We are the first to systematically apply mechanism interpretability methods to text-to-SQL. Our methodology integrates edge attribution patching, sparse autoencoders, circuit identification, and multi-scale comparative evaluation. This enables the precise localization of minimal functional circuits underlying SQL generation and reveals how such circuits dynamically reconfigure across query types. Critically, our analysis exposes fundamental limitations and biases of existing interpretability techniques, validates their utility in diagnosing model failure modes, and demonstrates their capacity to inform targeted dataset refinement.

0 citationsRead paper

Synthetic Data Privacy Metrics

Jan 07, 2025

Current privacy evaluation of synthetic data lacks standardized benchmarks, and existing metrics inadequately capture real-world adversarial risks. Method: This paper presents the first systematic empirical evaluation of mainstream privacy metrics—including adversarial attack simulation and membership inference success rates—in generative models. It comparatively analyzes practical efficacy of privacy-enhancing techniques such as differential privacy integration and privacy-aware generation, and quantifies the privacy–utility trade-off. Contribution/Results: We propose a deployment-oriented synthetic data privacy assessment framework featuring a reproducible evaluation pipeline, a standardized metric suite, and implementation guidelines. The study establishes a rigorous benchmarking methodology for academia and delivers an actionable, practice-driven privacy assurance evaluation paradigm—with concrete best practices—for industry adoption.

0 citationsRead paper
Recent publications

Latest Papers

TinySQL: A Progressive Text-to-SQL Dataset for Mechanistic Interpretability Research

Mar 17, 2025

Mechanism interpretability research has long been hindered by the gap between synthetic “toy tasks” and real-world complexity. This paper proposes text-to-SQL generation as an ideal benchmark task—offering both formal syntactic structure and realistic semantic and compositional challenges. To this end, we introduce TinySQL, a progressively scaled synthetic dataset covering foundational to advanced SQL operations, and establish a dedicated evaluation platform spanning 33M–1B parameter models. We are the first to systematically apply mechanism interpretability methods to text-to-SQL. Our methodology integrates edge attribution patching, sparse autoencoders, circuit identification, and multi-scale comparative evaluation. This enables the precise localization of minimal functional circuits underlying SQL generation and reveals how such circuits dynamically reconfigure across query types. Critically, our analysis exposes fundamental limitations and biases of existing interpretability techniques, validates their utility in diagnosing model failure modes, and demonstrates their capacity to inform targeted dataset refinement.

0 citationsRead paper

Synthetic Data Privacy Metrics

Jan 07, 2025

Current privacy evaluation of synthetic data lacks standardized benchmarks, and existing metrics inadequately capture real-world adversarial risks. Method: This paper presents the first systematic empirical evaluation of mainstream privacy metrics—including adversarial attack simulation and membership inference success rates—in generative models. It comparatively analyzes practical efficacy of privacy-enhancing techniques such as differential privacy integration and privacy-aware generation, and quantifies the privacy–utility trade-off. Contribution/Results: We propose a deployment-oriented synthetic data privacy assessment framework featuring a reproducible evaluation pipeline, a standardized metric suite, and implementation guidelines. The study establishes a rigorous benchmarking methodology for academia and delivers an actionable, practice-driven privacy assurance evaluation paradigm—with concrete best practices—for industry adoption.

0 citationsRead paper