Institution profile

VMware, Inc.

Industry researchnorthamerica · us
Official website
Research library11linked papers
Opportunities0open roles
Selected work

Representative Papers

An AI Approach to Verified Production Cryptographic Libraries

Aug 01, 2026

This work addresses the challenges of scale and complexity in formally verifying production-grade cryptographic libraries, where existing approaches fall short of end-to-end automation. We present CryptoProver, a system that achieves, for the first time, fully automated verification of real-world cryptographic implementations such as curve25519-dalek and RustCrypto’s chacha20. CryptoProver integrates large language models with the Verus verifier to automatically synthesize internal specifications and verifiable proofs from high-level API contracts, without requiring source code modifications. By leveraging a pre-defined trusted library, mechanical gating, and isolation mechanisms, the system ensures specification strength and cross-module consistency. In experiments, CryptoProver completed verification within 11.4 hours at an API cost of \$466.99, successfully covering core cryptographic components relied upon by widely deployed systems including Signal (with 218 million downloads) and Shadowsocks.

0 citationsRead paper

SpotIt+: Verification-based Text-to-SQL Evaluation with Database Constraints

Mar 04, 2026

This work addresses the limitations of existing Text-to-SQL evaluation methods, which often fail to capture semantic discrepancies between generated and reference SQL queries, particularly in the absence of real database constraints. To overcome this, the authors propose a bounded equivalence verification framework that actively searches for database instances capable of distinguishing the semantics of two queries. The core innovation lies in integrating rule-driven constraint mining with large language model–based validation, ensuring that the generated counterexamples are both semantically discriminative and realistic in practical deployment scenarios. Experiments on the BIRD dataset demonstrate that the proposed approach efficiently uncovers numerous semantic errors missed by conventional evaluation metrics, thereby substantially enhancing the validity and fidelity of Text-to-SQL system assessment.

0 citationsRead paper

Policy Compliance of User Requests in Natural Language for AI Systems

Feb 27, 2026

This work addresses the challenge of assessing compliance with organizational policies in natural language user requests by introducing the first annotated dataset tailored to industrial settings, along with a corresponding evaluation benchmark that encompasses diverse security and regulatory constraints prevalent in the technology sector. Leveraging large language models (LLMs) integrated with various prompting and reasoning strategies, the study systematically evaluates the performance of current approaches on this task. The findings reveal significant limitations of existing LLMs in accurately interpreting complex policy provisions and making reliable compliance judgments, thereby underscoring the inherent difficulty of the problem. This research establishes a reproducible benchmark and provides a foundational dataset to support future advancements in policy-aware language understanding.

0 citationsRead paper
Recent publications

Latest Papers

An AI Approach to Verified Production Cryptographic Libraries

Aug 01, 2026

This work addresses the challenges of scale and complexity in formally verifying production-grade cryptographic libraries, where existing approaches fall short of end-to-end automation. We present CryptoProver, a system that achieves, for the first time, fully automated verification of real-world cryptographic implementations such as curve25519-dalek and RustCrypto’s chacha20. CryptoProver integrates large language models with the Verus verifier to automatically synthesize internal specifications and verifiable proofs from high-level API contracts, without requiring source code modifications. By leveraging a pre-defined trusted library, mechanical gating, and isolation mechanisms, the system ensures specification strength and cross-module consistency. In experiments, CryptoProver completed verification within 11.4 hours at an API cost of \$466.99, successfully covering core cryptographic components relied upon by widely deployed systems including Signal (with 218 million downloads) and Shadowsocks.

0 citationsRead paper

SpotIt+: Verification-based Text-to-SQL Evaluation with Database Constraints

Mar 04, 2026

This work addresses the limitations of existing Text-to-SQL evaluation methods, which often fail to capture semantic discrepancies between generated and reference SQL queries, particularly in the absence of real database constraints. To overcome this, the authors propose a bounded equivalence verification framework that actively searches for database instances capable of distinguishing the semantics of two queries. The core innovation lies in integrating rule-driven constraint mining with large language model–based validation, ensuring that the generated counterexamples are both semantically discriminative and realistic in practical deployment scenarios. Experiments on the BIRD dataset demonstrate that the proposed approach efficiently uncovers numerous semantic errors missed by conventional evaluation metrics, thereby substantially enhancing the validity and fidelity of Text-to-SQL system assessment.

0 citationsRead paper

Policy Compliance of User Requests in Natural Language for AI Systems

Feb 27, 2026

This work addresses the challenge of assessing compliance with organizational policies in natural language user requests by introducing the first annotated dataset tailored to industrial settings, along with a corresponding evaluation benchmark that encompasses diverse security and regulatory constraints prevalent in the technology sector. Leveraging large language models (LLMs) integrated with various prompting and reasoning strategies, the study systematically evaluates the performance of current approaches on this task. The findings reveal significant limitations of existing LLMs in accurately interpreting complex policy provisions and making reliable compliance judgments, thereby underscoring the inherent difficulty of the problem. This research establishes a reproducible benchmark and provides a foundational dataset to support future advancements in policy-aware language understanding.

0 citationsRead paper