GALA: Generation-Aware Cross-Modal Alignment for Text-to-Time-Series Synthesis

📅 2026-08-13
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the lack of cross-modal alignment in text-conditioned time series generation by proposing the GALA framework. Employing a two-stage strategy, GALA utilizes an auxiliary generation loss to drive contrastive learning, thereby achieving explicit alignment between generation-aware text and time series encoders while integrating a flow matching generator to optimize synthesis quality. Evaluated on the TSFragment-600K benchmark, GALA effectively overcomes the trade-off between fidelity and semantic consistency, ranking first in 30 out of 36 metrics. The method achieves simultaneous improvements in FID, CTTP, and JFTSD scores, establishing a new state-of-the-art. These results demonstrate the critical role of generation-aware alignment mechanisms in enabling high-quality text-to-time-series synthesis.
📝 Abstract
Synthesizing time series from natural language is emerging as the most expressive form of controllable time series generation. However, existing text-conditioned generators either take caption embeddings frozen from off-the-shelf text encoders, or adapt the encoder end-to-end, letting the denoising loss shape the embeddings only as a by-product. In either case, the conditioning representation is never deliberately matched to the signal modality, leaving it ill-suited to guide generation. We address this by introducing GALA: Generation-Aware cross-modaL Alignment for text conditional time series generation. GALA is a two-stage approach that first contrastively couples a pretrained text encoder with a time-series foundation model into a shared embedding space with both encoders adapted to generation by an auxiliary generative loss, and then freezes the resulting caption embedding to drive a flow-matching generator. On TSFragment-600K, spanning four domains and three fragment lengths, GALA sets a new state of the art, ranking first in 30 of 36 metric columns and reaching an average rank of 1.08/1.08/1.42 at lengths 24/48/96 against 1.92/2.00/1.75 for the strongest baseline. We further find that generator-internal text encoders force a trade-off between fidelity and caption adherence, whereas conditioning on the aligned embedding breaks it: FID, CTTP, and JFTSD all improve at once. Ablating the auxiliary loss degrades FID, CTTP and JFTSD together, it indicates the generative term is a necessary component of the alignment rather than an add-on.
Problem

Research questions and friction points this paper is trying to address.

Text-to-Time-Series Synthesis
Cross-Modal Alignment
Conditioning Representation
Controllable Generation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Cross-Modal Alignment
Text-to-Time-Series Synthesis
Generation-Aware Learning
Flow Matching
Auxiliary Generative Loss
H
Haochen Zhang
UNITES Lab, University of North Carolina at Chapel Hill
Gengwei Zhang
Gengwei Zhang
University of Technology Sydney, Sun Yat-sen University
Computer VisionMachine Learning
L
Laura Yao
UNITES Lab, University of North Carolina at Chapel Hill
N
Nicholas Knoz
UNITES Lab, University of North Carolina at Chapel Hill
Tianlong Chen
Tianlong Chen
Assistant Professor, CS@UNC Chapel Hill; Chief AI Scientist, hireEZ
Machine LearningAI4ScienceComputer VisionSparsity