Synthetic Semantic Supervision for Contrastive Code Representation Learning in Small Transformers: An Empirical Study

📅 2026-09-03
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究使用合成语义监督方法训练小型编码器,通过生成强调代码功能和意图的自然语言描述来解决代码表示学习中的标注问题。
📝 Abstract
General-purpose code embeddings power tools for code search, classification, and retrieval. Compact transformer encoders for code typically rely on either human-written docstrings (labor-intensive and inconsistent) or mined structural signals such as execution traces (setting-specific and costly to collect). We empirically study an alternative: contrastive pretraining of small encoders with synthetically generated natural-language descriptions emphasizing code functionality and intent, paired with code in a dual-encoder framework at training and discarded at inference. We benchmark this approach against pretraining-based baselines, generalist LLMs, and embedding-specific models on eight retrieval, classification, and generation tasks across C, C++, and Java. Synthetic semantic supervision yields statistically significant gains over pretraining baselines of the same inference-time size on five of eight tasks, with parity on two more; once fine-tuned, it matches or exceeds zero-shot models two orders of magnitude larger on classification, and it stays on par with execution-aware supervision at matched pretraining data, suggesting a scalable, effective alternative to existing code-representation paradigms.
Problem

Research questions and friction points this paper is trying to address.

code embeddings
contrastive learning
synthetic supervision
small transformers
code representation
Innovation

Methods, ideas, or system contributions that make the work stand out.

contrastive pretraining
synthetic semantic supervision
dual-encoder framework
🔎 Similar Papers
No similar papers found.