Calibrate What You SHIP: Post-Selection Risk Control for Verifier-Guided Text-to-Image Generation

๐Ÿ“… 2026-08-21
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
ๆœฌๆ–‡้’ˆๅฏน้ชŒ่ฏๅ™จๅผ•ๅฏผ็š„ๆ–‡ๆœฌๅˆฐๅ›พๅƒ็”Ÿๆˆ็ณป็ปŸไธญๅ€™้€‰ไธŽ็ญ–็•ฅๆ กๅ‡†ไธๅŒน้…็š„้—ฎ้ข˜๏ผŒๆๅ‡บSHIPๆ–นๆณ•่ฟ›่กŒ้€‰ๆ‹ฉๆ„Ÿ็Ÿฅ็š„ไฟ็•™ๆ กๅ‡†๏ผŒไปฅๆŽงๅˆถๅฎž้™…ๅ‘ๅธƒ่พ“ๅ‡บ็š„้ฃŽ้™ฉใ€‚
๐Ÿ“ Abstract
Verifier-guided text-to-image systems increasingly use test-time search to select, refine, or stop among multiple candidates, yet release thresholds are often calibrated on individual images. This creates a candidate-to-policy calibration mismatch: search changes both which prompts receive an output and which candidate is released, so candidate-level risk control need not imply control of released-output risk. We formalize this estimand shift through prompt reweighting and within-prompt selection, and introduce SHIP, Selection-aware Held-out calibration of Inference Policies. SHIP runs or replays the complete deployed policy on held-out prompts, evaluates the image it actually releases using an independent target judge, and selects the most permissive threshold whose risk upper bound satisfies a prescribed budget. For replayable policies with a prespecified threshold grid, simultaneous confidence control provides finite-sample validity. Experiments across fixed, sequential, and adaptive T2I inference procedures show that policy-level calibration recovers lower-risk operating points while exposing policy-dependent tradeoffs among risk, coverage, and compute. On GenEval2 with FLUX at N=16, a pooled-candidate threshold yields released risk 0.310, whereas SHIP reduces it to 0.162. Across 200 cached-stream splits, the fixed-grid certificate has no target crossing. Reliable inference-time scaling therefore requires calibrating the output distribution induced by the complete deployed policy.
Problem

Research questions and friction points this paper is trying to address.

Verifier-guided
Text-to-Image Generation
Post-Selection Risk Control
Calibration Mismatch
Inference Policies
Innovation

Methods, ideas, or system contributions that make the work stand out.

SHIP
calibration
text-to-image generation
risk control
policy-level