Transfer Learning for Evolving Domains
本文提出了一种针对数据可用性随时间演变的领域迁移学习问题(TrED),并探讨了现有方法在处理整个演化过程中的不足。
本文提出了一种针对数据可用性随时间演变的领域迁移学习问题(TrED),并探讨了现有方法在处理整个演化过程中的不足。
This study addresses the limitations of existing causal discovery methods, which rely on regular sampling and fixed lag structures and thus struggle with irregularly sampled time series commonly found in sensor, medical, and financial domains. To overcome this challenge, the work extends the PCMCI+ framework—originally designed for regularly spaced time series based on conditional independence tests—to irregular event streams by introducing a time-window aggregation mechanism. This mechanism replaces conventional fixed-lag modeling and enables time-aware causal inference. Empirical evaluations demonstrate that the proposed approach accurately recovers ground-truth causal graphs across synthetic irregular datasets under varying signal-to-noise ratios, significantly outperforming the original PCMCI+ and effectively eliminating the dependency on regular temporal structure.
This work addresses the high latency, resource contention, and operational overhead caused by frequent state updates in streaming machine learning. The authors propose a probabilistic sparsification strategy that decouples inference from state persistence: while all events contribute to inference scoring, only those deemed highly informative trigger persistence. This approach enables precise control over the persistence path without requiring high-frequency in-memory control planes or cross-node coordination, while preserving unbiasedness of time-aggregated statistics. By integrating approximate statistics from disk-based key-value stores with variance-aware temporal aggregation modeling, the method reduces persistence events by up to 90%, substantially lowering I/O and serialization costs while maintaining or even improving downstream task performance.
This study addresses the challenge of model performance degradation caused by temporal shifts in feature distributions within real-world time-series data. The authors propose a parameter-free, automated detection method that leverages a regression model to predict sample timestamps and integrates feature importance analysis to identify time-sensitive features. By unifying the quantification of both univariate and multivariate distributional changes across numerical and categorical features, the approach offers strong scalability. Experimental results demonstrate that the method effectively and comprehensively captures a wide range of fundamental drift patterns on both real-world and synthetic datasets, achieving high detection accuracy while maintaining computational efficiency.
Financial criminals often obscure fund flows through complex transactions, rendering traditional network visualizations ineffective for anti-money laundering (AML) analysis. This work proposes a tabular temporal graph visualization approach tailored for AML, introducing this representation to the domain for the first time. We design three aggregation strategies for nodes and edges—based on transaction amount, time, and their joint consideration—to balance detail preservation with analytical efficiency. Leveraging these strategies, we implement an interactive visualization system and validate its efficacy through expert user studies. The findings reveal a trade-off between graph simplification and analytical utility: finer-grained representations, while imposing higher cognitive load, significantly enhance analysts’ interest and depth of insight.
本文提出了一种针对数据可用性随时间演变的领域迁移学习问题(TrED),并探讨了现有方法在处理整个演化过程中的不足。
This study addresses the limitations of existing causal discovery methods, which rely on regular sampling and fixed lag structures and thus struggle with irregularly sampled time series commonly found in sensor, medical, and financial domains. To overcome this challenge, the work extends the PCMCI+ framework—originally designed for regularly spaced time series based on conditional independence tests—to irregular event streams by introducing a time-window aggregation mechanism. This mechanism replaces conventional fixed-lag modeling and enables time-aware causal inference. Empirical evaluations demonstrate that the proposed approach accurately recovers ground-truth causal graphs across synthetic irregular datasets under varying signal-to-noise ratios, significantly outperforming the original PCMCI+ and effectively eliminating the dependency on regular temporal structure.
This work addresses the high latency, resource contention, and operational overhead caused by frequent state updates in streaming machine learning. The authors propose a probabilistic sparsification strategy that decouples inference from state persistence: while all events contribute to inference scoring, only those deemed highly informative trigger persistence. This approach enables precise control over the persistence path without requiring high-frequency in-memory control planes or cross-node coordination, while preserving unbiasedness of time-aggregated statistics. By integrating approximate statistics from disk-based key-value stores with variance-aware temporal aggregation modeling, the method reduces persistence events by up to 90%, substantially lowering I/O and serialization costs while maintaining or even improving downstream task performance.
This study addresses the challenge of model performance degradation caused by temporal shifts in feature distributions within real-world time-series data. The authors propose a parameter-free, automated detection method that leverages a regression model to predict sample timestamps and integrates feature importance analysis to identify time-sensitive features. By unifying the quantification of both univariate and multivariate distributional changes across numerical and categorical features, the approach offers strong scalability. Experimental results demonstrate that the method effectively and comprehensively captures a wide range of fundamental drift patterns on both real-world and synthetic datasets, achieving high detection accuracy while maintaining computational efficiency.
Financial criminals often obscure fund flows through complex transactions, rendering traditional network visualizations ineffective for anti-money laundering (AML) analysis. This work proposes a tabular temporal graph visualization approach tailored for AML, introducing this representation to the domain for the first time. We design three aggregation strategies for nodes and edges—based on transaction amount, time, and their joint consideration—to balance detail preservation with analytical efficiency. Leveraging these strategies, we implement an interactive visualization system and validate its efficacy through expert user studies. The findings reveal a trade-off between graph simplification and analytical utility: finer-grained representations, while imposing higher cognitive load, significantly enhance analysts’ interest and depth of insight.