Nothing Changed but the Model: CellFill -- Bounded In-Cell Learning for Bit-Identical, Revocable Updates to Quantized LLMs

📅 2026-08-21
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究解决了更新量化语言模型时破坏原有评估和缓存的问题,通过在解量化间隙中学习新知识而不改变原权重位,确保更新可撤销且漂移受限。
📝 Abstract
Every way of teaching a deployed language model something new -- full fine-tuning, adapter merging, model editing -- replaces the released checkpoint, and with it every evaluation and cache that referred to those exact bits. We instead learn inside the dequantization gap: with the integer codes and scales of a 4-bit release frozen, new knowledge is written only into the per-weight residual that lives strictly inside each quantization decision cell. Re-quantization then returns the released artifact bit-for-bit, a machine-checkable guarantee; updates are exactly revocable by dropping the residual; and drift is bounded. We give six propositions and three training paths, including CellFill, a bounded reparameterization that makes invariance structural rather than enforced. Exact invariance turns out to be nearly free: across three paired seeds the constrained dense path matches an unconstrained reference whose weights provably escape the artifact (58.9 vs 59.3 percent fact recall; paired difference -0.5 points, 95% CI [-5.0,+4.0]), and is better on held-out cross-domain perplexity. Against the natural null hypothesis -- serving the same update as an unmerged adapter -- projecting into the cells reduces cross-domain forgetting in every run that converged, and a diverged control shows the boundary: projection is a trust region, not a repair. What no method escapes is the cost of knowledge itself, and the apparent free lunch of in-domain perplexity improving past the anchor is an artifact of rehearsal sharing a corpus with the metric. Methods differ threefold at matched rehearsal in knowledge bought per point of cross-domain perplexity, a ranking that is not the recall ranking. The method transfers to a 27B hybrid linear-attention model (2.4e10 constrained weights, verified bit-identical), where matched recall costs about half as much cross-domain perplexity as at 1.7B.
Problem

Research questions and friction points this paper is trying to address.

language model
quantization
in-cell learning
bit-identical
revocable updates
Innovation

Methods, ideas, or system contributions that make the work stand out.

CellFill
In-Cell Learning
Quantized LLMs
Bit-Identical Updates
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Z
Zifeng Liu
Big Data and Artificial Intelligence Center, The Third Affiliated Hospital of Sun Yat-sen University; Institute for Frontier Interdisciplinary Research in Health Sciences and Technology, Sun Yat-sen University; Sun Yat-sen University Institute of Artificial Intelligence; Guangdong Engineering Research Center of Medical Artificial Intelligence Multimodal System
Z
Zhiyong Du
School of Business, Sun Yat-sen University
Y
Yaxin Lu
Big Data and Artificial Intelligence Center, The Third Affiliated Hospital of Sun Yat-sen University
Y
Yiming Mao
School of Computer Science, China University of Geosciences (Wuhan)
Z
Zhenhe Wang
School of Public Health, Sun Yat-sen University
Wenqi Shi
Wenqi Shi
Assistant Professor, University of Texas Southwestern Medical Center
AI for HealthcareLLM AgentClinical Decision SupportClinical Informatics
Z
Zhengkun Jing
Hospital of Stomatology, Sun Yat-sen University