🤖 AI Summary
本文通过对比实验评估ChatGPT Images 2.5在伪造任务中的实际表现,特别是其OCR识别准确性和图像编辑能力是否有所提升。
📝 Abstract
We evaluate whether the improvements advertised for ChatGPT Images 2.5 translate into better performance on forgery tasks with predetermined answers. We compare its Flare and Sunburst API models with GPT-Image-2 re-run in the same week, using receipt-field edits, repeated editing, product placement and fine-print rendering. After image registration, Flare and Sunburst show fewer OCR-detected changes to surrounding receipt text (31.7% and 31.2% versus 44.2% for both GPT-Image-2 baselines), mainly on CORD receipts, without a detectable improvement in target-field correctness. Flare retains fewer earlier edits on CORD receipts, while photo-edit sequences provide little separation between models. Product codes are more often legible with Images 2.5, alongside larger product placement; the analyses do not establish a fidelity gain independent of size. Fine-print improvements remain unresolved below the OCR reliability limit. Refusals are rare and localisation is weak in both generations. At a fixed detection threshold, Community Forensics flags 68.6% of controlled Images 2.5 images averaged across cells, versus 35.9% of self-reported images posted online. These results motivate task-specific evaluation of advertised capabilities and defences, with explicit limits on what automatic checks can establish.