Institution profile

Dialpad Inc.

Industry researchnorthamerica · us
Official website
Research library15linked papers
Opportunities0open roles
Selected work

Representative Papers

Beyond Text-to-SQL: An Agentic LLM System for Governed Enterprise Analytics APIs

May 20, 2026

This work addresses the challenge faced by non-technical users in enterprise environments who struggle to securely and compliantly access governed analytical data through natural language. Traditional text-to-SQL approaches fail to accommodate managed APIs that encapsulate complex business logic. To bridge this gap, the authors propose Analytic Agent—a large language model (LLM)-based agent system that translates natural language intents into secure invocations of governed analytical APIs through multi-step reasoning and policy-aware orchestration. The system integrates permission validation, query execution, and compliance-aware visualization, achieving the first deep integration of LLM agents with enterprise-grade governed APIs while ensuring auditability, consistency, and security. Evaluation on 90 real-world enterprise use cases demonstrates that the system accurately interprets user intent, executes compliant queries, and generates visual results, significantly enhancing self-service analytics capabilities for non-technical users.

0 citationsRead paper

From Text to Voice: A Reproducible and Verifiable Framework for Evaluating Tool Calling LLM Agents

May 14, 2026

Existing benchmarks for evaluating tool use in spoken-language agents lack realistic audio scenarios. This work proposes a dataset-agnostic framework that automatically converts existing text-based tool-use benchmarks into paired speech versions without requiring re-annotation, by leveraging text-to-speech synthesis, multi-speaker voice conversion, and environmental noise injection. The approach innovatively introduces a reference-free LLM-as-judge protocol coupled with ambiguity-aware stress testing, enabling verifiable evaluation under privacy-preserving conditions. Evaluations of seven models on spoken variants of Confetti and When2Call reveal strong dependencies of performance on both model architecture and task characteristics. An open-source Qwen3 judge model achieves over 80% agreement with commercial counterparts, demonstrating the framework’s validity and practical utility.

0 citationsRead paper
Recent publications

Latest Papers

Beyond Text-to-SQL: An Agentic LLM System for Governed Enterprise Analytics APIs

May 20, 2026

This work addresses the challenge faced by non-technical users in enterprise environments who struggle to securely and compliantly access governed analytical data through natural language. Traditional text-to-SQL approaches fail to accommodate managed APIs that encapsulate complex business logic. To bridge this gap, the authors propose Analytic Agent—a large language model (LLM)-based agent system that translates natural language intents into secure invocations of governed analytical APIs through multi-step reasoning and policy-aware orchestration. The system integrates permission validation, query execution, and compliance-aware visualization, achieving the first deep integration of LLM agents with enterprise-grade governed APIs while ensuring auditability, consistency, and security. Evaluation on 90 real-world enterprise use cases demonstrates that the system accurately interprets user intent, executes compliant queries, and generates visual results, significantly enhancing self-service analytics capabilities for non-technical users.

0 citationsRead paper

From Text to Voice: A Reproducible and Verifiable Framework for Evaluating Tool Calling LLM Agents

May 14, 2026

Existing benchmarks for evaluating tool use in spoken-language agents lack realistic audio scenarios. This work proposes a dataset-agnostic framework that automatically converts existing text-based tool-use benchmarks into paired speech versions without requiring re-annotation, by leveraging text-to-speech synthesis, multi-speaker voice conversion, and environmental noise injection. The approach innovatively introduces a reference-free LLM-as-judge protocol coupled with ambiguity-aware stress testing, enabling verifiable evaluation under privacy-preserving conditions. Evaluations of seven models on spoken variants of Confetti and When2Call reveal strong dependencies of performance on both model architecture and task characteristics. An open-source Qwen3 judge model achieves over 80% agreement with commercial counterparts, demonstrating the framework’s validity and practical utility.

0 citationsRead paper