Universal-2-TF: Robust All-Neural Text Formatting for ASR
This paper addresses critical post-processing challenges in commercial ASR outputs—namely, missing punctuation, inconsistent capitalization, and unnormalized numerals/abbreviations—by proposing the first end-to-end multi-objective text normalization framework. Methodologically, it departs from rule-based and hybrid approaches, introducing a lightweight two-stage fully neural architecture: (1) a multi-task token classifier jointly predicting punctuation, capitalization, and inverse text normalization labels; and (2) a sequence-to-sequence model for fine-grained correction. Both stages are jointly trained and seamlessly integrated into the Universal-2 ASR system. Experiments demonstrate significant improvements over strong baselines across objective metrics (e.g., +5.1 F1 points, −40% inference latency) and subjective listening evaluations. The framework further exhibits superior cross-domain generalization and enhanced hallucination suppression.