Institution profile

National Computer Network Emergency Response Technical Team/Coordination Center

Academic institutionasia · cn
Official website
Research library13linked papers
Opportunities0open roles
Selected work

Representative Papers

Cascading versus Joint Modeling for Hierarchical Offensive Language Detection

Jul 18, 2026

This study investigates the trade-offs between cascaded modeling and joint multi-task modeling for fine-grained offensive language detection, focusing on accuracy, parameter count, and inference latency, while optimizing class-imbalance handling strategies for each subtask. We construct a three-stage cascade system with tailored training protocols per stage and employ ablation studies to identify the optimal imbalance mitigation approach. A shared-encoder joint model serves as a baseline, enabling the first controlled quantitative comparison between the two paradigms. Results show that the cascade achieves macro F1 scores of 0.795, 0.716, and 0.557 across the three subtasks, outperforming the joint model by 7.1 points on the most imbalanced task, albeit with three times more parameters and 1.67× higher latency. Ablation analysis further reveals that approximately 20% of cascade errors originate in the first stage and are irrecoverable, underscoring the critical impact of pipeline design.

0 citationsRead paper

Finite-Blocklength Analysis for Noisy Permutation Channels

May 25, 2026

This work addresses the looseness of existing finite-blocklength performance bounds for noisy permutation channels whose achievable output polytope has an affine dimension \(d\) strictly lower than that of the output simplex. By projecting the empirical distribution onto the affine hull of achievable outputs and employing Euclidean nearest-neighbor decoding, the decoding error is geometrically reduced to a one-dimensional transition event. Leveraging a refined meta-converse argument, KL divergence covering, and local binary hypothesis testing, the paper establishes the first tight finite-blocklength achievability and converse bounds that depend on the affine dimension \(d\) rather than the ambient output space dimension. The derived achievability bound hinges on local coordinate var日消息 and relative volume ratios, while the converse bound features a blocklength-dependent term of order \(d \log\sqrt{n}\), substantially improving bound tightness.

0 citationsRead paper

ActiveFlowMark: Assessing Tor Anonymity under Active Bandwidth Watermarking

May 07, 2026

This work addresses infrastructure-level traffic analysis attacks against low-latency anonymous networks such as Tor, where adversaries exploit side-channel information from encrypted communications to perform traffic correlation. The paper proposes NATA, a novel algorithm that, for the first time, integrates active bandwidth watermarking with state-space learning to enable non-intrusive traffic correlation without endpoint compromise or packet modification. Specifically, NATA injects controllable bandwidth perturbations at upstream relays and observes them passively at exit relays. To achieve this, the authors design the BM-Net framework, which combines masked self-supervised pretraining with task-specific fine-tuning to efficiently learn traffic representations. Experimental results on real-world Tor traffic demonstrate a 99.65% F1 score for perturbation detection and a 97.5% macro F1 score for fine-grained modulation classification, with exit observation probabilities further evaluated through simulation.

0 citationsRead paper

P^2O: Joint Policy and Prompt Optimization

Mar 23, 2026

This work addresses the challenge of sparse rewards in reinforcement learning, where "hard examples" yield zero advantage estimates and thus lack effective supervision signals. The authors propose a novel co-optimization framework that jointly refines policies and prompts: by identifying hard examples during training, they employ a Genetic Evolutionary Pareto Algorithm (GEPA) to optimize prompt templates that guide large language models to generate successful trajectories, subsequently distilling the prompt-induced reasoning capabilities into the policy parameters. This approach represents the first method to integrate prompt optimization and policy learning within a unified training loop, moving beyond conventional prompt engineering paradigms that rely solely on input augmentation. Experimental results demonstrate state-of-the-art performance on in-distribution tasks and an average improvement of 4.7% on out-of-distribution benchmarks, significantly enhancing model generalization.

0 citationsRead paper

Rethinking LLM Watermark Detection in Black-Box Settings: A Non-Intrusive Third-Party Framework

Mar 16, 2026

This work addresses the limitation of existing large language model watermarking schemes, which rely on secret keys or service-provider-specific detectors and thus hinder independent third-party auditing. To overcome this, the authors propose TTP-Detect, a framework enabling key-free, non-invasive, black-box third-party watermark verification without access to model internals. The method enhances watermark signals using surrogate models and evaluates distribution alignment through multiple relative metrics, effectively decoupling watermark injection from detection. Experimental results demonstrate that TTP-Detect achieves strong detection performance and robustness against attacks across diverse watermarking schemes, datasets, and language models, establishing the first truly decentralized watermark auditing capability.

0 citationsRead paper
Recent publications

Latest Papers

Cascading versus Joint Modeling for Hierarchical Offensive Language Detection

Jul 18, 2026

This study investigates the trade-offs between cascaded modeling and joint multi-task modeling for fine-grained offensive language detection, focusing on accuracy, parameter count, and inference latency, while optimizing class-imbalance handling strategies for each subtask. We construct a three-stage cascade system with tailored training protocols per stage and employ ablation studies to identify the optimal imbalance mitigation approach. A shared-encoder joint model serves as a baseline, enabling the first controlled quantitative comparison between the two paradigms. Results show that the cascade achieves macro F1 scores of 0.795, 0.716, and 0.557 across the three subtasks, outperforming the joint model by 7.1 points on the most imbalanced task, albeit with three times more parameters and 1.67× higher latency. Ablation analysis further reveals that approximately 20% of cascade errors originate in the first stage and are irrecoverable, underscoring the critical impact of pipeline design.

0 citationsRead paper

Finite-Blocklength Analysis for Noisy Permutation Channels

May 25, 2026

This work addresses the looseness of existing finite-blocklength performance bounds for noisy permutation channels whose achievable output polytope has an affine dimension \(d\) strictly lower than that of the output simplex. By projecting the empirical distribution onto the affine hull of achievable outputs and employing Euclidean nearest-neighbor decoding, the decoding error is geometrically reduced to a one-dimensional transition event. Leveraging a refined meta-converse argument, KL divergence covering, and local binary hypothesis testing, the paper establishes the first tight finite-blocklength achievability and converse bounds that depend on the affine dimension \(d\) rather than the ambient output space dimension. The derived achievability bound hinges on local coordinate var日消息 and relative volume ratios, while the converse bound features a blocklength-dependent term of order \(d \log\sqrt{n}\), substantially improving bound tightness.

0 citationsRead paper

ActiveFlowMark: Assessing Tor Anonymity under Active Bandwidth Watermarking

May 07, 2026

This work addresses infrastructure-level traffic analysis attacks against low-latency anonymous networks such as Tor, where adversaries exploit side-channel information from encrypted communications to perform traffic correlation. The paper proposes NATA, a novel algorithm that, for the first time, integrates active bandwidth watermarking with state-space learning to enable non-intrusive traffic correlation without endpoint compromise or packet modification. Specifically, NATA injects controllable bandwidth perturbations at upstream relays and observes them passively at exit relays. To achieve this, the authors design the BM-Net framework, which combines masked self-supervised pretraining with task-specific fine-tuning to efficiently learn traffic representations. Experimental results on real-world Tor traffic demonstrate a 99.65% F1 score for perturbation detection and a 97.5% macro F1 score for fine-grained modulation classification, with exit observation probabilities further evaluated through simulation.

0 citationsRead paper

P^2O: Joint Policy and Prompt Optimization

Mar 23, 2026

This work addresses the challenge of sparse rewards in reinforcement learning, where "hard examples" yield zero advantage estimates and thus lack effective supervision signals. The authors propose a novel co-optimization framework that jointly refines policies and prompts: by identifying hard examples during training, they employ a Genetic Evolutionary Pareto Algorithm (GEPA) to optimize prompt templates that guide large language models to generate successful trajectories, subsequently distilling the prompt-induced reasoning capabilities into the policy parameters. This approach represents the first method to integrate prompt optimization and policy learning within a unified training loop, moving beyond conventional prompt engineering paradigms that rely solely on input augmentation. Experimental results demonstrate state-of-the-art performance on in-distribution tasks and an average improvement of 4.7% on out-of-distribution benchmarks, significantly enhancing model generalization.

0 citationsRead paper

Rethinking LLM Watermark Detection in Black-Box Settings: A Non-Intrusive Third-Party Framework

Mar 16, 2026

This work addresses the limitation of existing large language model watermarking schemes, which rely on secret keys or service-provider-specific detectors and thus hinder independent third-party auditing. To overcome this, the authors propose TTP-Detect, a framework enabling key-free, non-invasive, black-box third-party watermark verification without access to model internals. The method enhances watermark signals using surrogate models and evaluates distribution alignment through multiple relative metrics, effectively decoupling watermark injection from detection. Experimental results demonstrate that TTP-Detect achieves strong detection performance and robustness against attacks across diverse watermarking schemes, datasets, and language models, establishing the first truly decentralized watermark auditing capability.

0 citationsRead paper