Amortizing Scaling Law Construction Costs
本文提出一种高效构建缩放定律的框架,通过贝叶斯优化问题和计算预算逐步扩展的方法,大幅降低计算成本并准确拟合缩放定律。
本文提出一种高效构建缩放定律的框架,通过贝叶斯优化问题和计算预算逐步扩展的方法,大幅降低计算成本并准确拟合缩放定律。
本文探讨了机器学习和法律界对'准确性'的不同理解和期望,分析了五个主要矛盾,并建议通过标准化和工具来解决这些差异。
研究通过引入基于Jensen-Shannon散度的解释一致性分数(ECS)来量化糖尿病视网膜病变筛查中不同人群归因图的一致性,以补充预测公平性的评估。
This work investigates how to reverse-engineer executable decision code of an agent solely from its behavioral trajectories in games and examines how actively designed adversarial experiments can enhance reconstruction fidelity. To this end, we introduce RevengeBench, a benchmark comprising 75 Elo-calibrated CodeClash strategies across five environments, where customized behavioral probes are generated by observing interactions between target and opponent agents to reconstruct underlying code logic. We formalize strategy reversal as a tractable inverse problem in code space and incorporate a mechanism for controlled experimentation. Experiments across 12 large language models demonstrate that our approach significantly reduces initial behavioral divergence (by 34%–72%) and yields reconstructed strategies that exhibit competitive performance in downstream adversarial settings, particularly bolstering the counterplay capabilities of weaker models.
This work addresses the challenge of quantifying semantic novelty in language model outputs, proposing *unattributability*—the inability to semantically retrieve any pretraining corpus sample as the source of an output—as a formal, interpretable metric. Methodologically, it introduces a two-stage retrieval pipeline: first, efficient coarse-grained indexing using GIST embeddings; second, fine-grained re-ranking via ColBERTv2, with attribution thresholds calibrated against human-written text. The paper provides the first formal definition and empirical evaluation of unattributability. Experiments on the SmolLM family reveal three key findings: (1) instruction tuning significantly improves output unattributability; (2) increased reliance on longer contexts enhances semantic novelty; and (3) domain-specific characteristics shape the distribution of unattributable outputs. By grounding novelty assessment in semantic retrieval fidelity rather than surface-level heuristics, this work establishes a scalable, principled framework for evaluating generative originality—offering both interpretability and practical applicability for safety, copyright, and alignment research.
本文提出一种高效构建缩放定律的框架,通过贝叶斯优化问题和计算预算逐步扩展的方法,大幅降低计算成本并准确拟合缩放定律。
本文探讨了机器学习和法律界对'准确性'的不同理解和期望,分析了五个主要矛盾,并建议通过标准化和工具来解决这些差异。
研究通过引入基于Jensen-Shannon散度的解释一致性分数(ECS)来量化糖尿病视网膜病变筛查中不同人群归因图的一致性,以补充预测公平性的评估。
This work investigates how to reverse-engineer executable decision code of an agent solely from its behavioral trajectories in games and examines how actively designed adversarial experiments can enhance reconstruction fidelity. To this end, we introduce RevengeBench, a benchmark comprising 75 Elo-calibrated CodeClash strategies across five environments, where customized behavioral probes are generated by observing interactions between target and opponent agents to reconstruct underlying code logic. We formalize strategy reversal as a tractable inverse problem in code space and incorporate a mechanism for controlled experimentation. Experiments across 12 large language models demonstrate that our approach significantly reduces initial behavioral divergence (by 34%–72%) and yields reconstructed strategies that exhibit competitive performance in downstream adversarial settings, particularly bolstering the counterplay capabilities of weaker models.
This work addresses the challenge of quantifying semantic novelty in language model outputs, proposing *unattributability*—the inability to semantically retrieve any pretraining corpus sample as the source of an output—as a formal, interpretable metric. Methodologically, it introduces a two-stage retrieval pipeline: first, efficient coarse-grained indexing using GIST embeddings; second, fine-grained re-ranking via ColBERTv2, with attribution thresholds calibrated against human-written text. The paper provides the first formal definition and empirical evaluation of unattributability. Experiments on the SmolLM family reveal three key findings: (1) instruction tuning significantly improves output unattributability; (2) increased reliance on longer contexts enhances semantic novelty; and (3) domain-specific characteristics shape the distribution of unattributable outputs. By grounding novelty assessment in semantic retrieval fidelity rather than surface-level heuristics, this work establishes a scalable, principled framework for evaluating generative originality—offering both interpretability and practical applicability for safety, copyright, and alignment research.