Institution profile

Snowflake Inc.

Industry researchnorthamerica · us
Official website
Research library85linked papers
Opportunities99open roles
Selected work

Representative Papers

ConvCodeWorld: Benchmarking Conversational Code Generation in Reproducible Feedback Environments

Feb 27, 2025

Existing code generation benchmarks fail to model diverse multi-turn feedback—such as compilation errors, execution outcomes, and natural-language critiques—limiting rigorous evaluation of LLMs in conversational programming. Method: We introduce ConvCodeBench (static) and ConvCodeWorld (dynamic), the first reproducible multi-turn feedback evaluation benchmarks, establishing a feedback-driven evaluation paradigm. Leveraging GPT-4o, we generate structured natural-language feedback; integrate compiler-based error simulation; and deploy a coverage-aware execution engine to emulate realistic developer interactions. Benchmark consistency is validated via Spearman correlation. Contribution/Results: Experiments reveal that feedback type and intensity critically affect model adaptability: weaker models can surpass stronger ones’ single-turn performance after multiple feedback rounds, yet feedback-combination specificity induces generalization bottlenecks. Moreover, a trade-off exists between Mean Reciprocal Rank (MRR) and Recall, highlighting inherent limitations in current feedback-integration strategies.

1 citationsRead paper
Recent publications

Latest Papers