A Reproducible, License-Aware Distillation Recipe for CPUDeployable Safety Classification

📅 2026-08-21
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出了一种可再现、考虑许可的知识蒸馏方法,用于在CPU上部署大型语言模型的安全层,通过训练小型模型来复现强开放守护模型的安全分类信号。
📝 Abstract
Deploying a safety layer for large language models on commodity hardware is constrained by the guards available to do it: current open guard models hold between 1 and 9 billion parameters, are oriented toward the graphics processing unit, and answer in seconds per request on a central processing unit. This paper presents a reproducible, license-aware knowledge-distillation recipe addressing that constraint. A strong open guard labels a corpus of roughly 97,000 prompts, drawn from 24 public datasets, into seven safety categories aligned to a public hazard taxonomy, and a fleet of small students spanning lexical, shallow, encoder and generative architectures is trained to reproduce that signal. The corpus is partitioned at the license boundary, so that a deployable and a research model differ only in their training data and the cost of that restriction becomes measurable. Every model is scored against an independent gold benchmark of 6,361 rows over four slices, labeled apart from the teacher and including a slice of harmless prompts that makes over-defense measurable. The distilled students match the teachers on adversarial text within overlapping confidence intervals and reduce false alarms on harmless prompts, the smallest generative student reaching 3.8% against 4.8% for the 8-billion-parameter teacher, while the encoder classifies in roughly 24 ms per request on CPU. Per-class rebalancing is the only decisive ingredient of the recipe. No superiority over the distilled guards is claimed; on the clean reference slice they remain ahead.
Problem

Research questions and friction points this paper is trying to address.

safety layer
large language models
commodity hardware
parameter constraints
CPU deployment
Innovation

Methods, ideas, or system contributions that make the work stand out.

knowledge distillation
license-aware
safety classification
CPU deployable
rebalancing
E
Edson Rodrigues da Cruz Filho
Quickium Technology Ltd., Piracicaba, São Paulo, Brazil.
P
Paulo Ricardo Ferreira Neves
Quickium Technology Ltd., Piracicaba, São Paulo, Brazil.; University of São Paulo (USP), Escola Superior de Agricultura Luiz de Queiroz, Piracicaba Campus, São Paulo, Brazil.
P
Paulo Henrique Eleuterio Falsetti
Quickium Technology Ltd., Piracicaba, São Paulo, Brazil.; Federal University of São Carlos (UFSCar), Sorocaba Campus, São Paulo, Brazil.
J
João Vitor Pavan
Quickium Technology Ltd., Piracicaba, São Paulo, Brazil.
I
Ian Degaspari
Quickium Technology Ltd., Piracicaba, São Paulo, Brazil.
H
Henrique Vieira Laturrague
Quickium Technology Ltd., Piracicaba, São Paulo, Brazil.
P
Patrick Vieira Laturrague
Quickium Technology Ltd., Piracicaba, São Paulo, Brazil.
G
Guilherme Nielsen Dias
Quickium Technology Ltd., Piracicaba, São Paulo, Brazil.
M
Marccello Wilson Perez Berto
Quickium Technology Ltd., Piracicaba, São Paulo, Brazil.
G
Gustavo Voltani Von Atzingen
Quickium Technology Ltd., Piracicaba, São Paulo, Brazil.; Federal Institute of Education, Science and Technology of São Paulo (IFSP), Piracicaba Campus, São Paulo, Brazil.