SMTrap: Cost-Effective DoS Attacks Against Large Reasoning Models via SMT Conflict Guidance
本文提出SMTrap,一种基于SMT冲突指导的方法,生成计算密集型CSP查询以低成本实现对大型推理模型的DoS攻击。
本文提出SMTrap,一种基于SMT冲突指导的方法,生成计算密集型CSP查询以低成本实现对大型推理模型的DoS攻击。
Lightweight vision-language models (VLMs) suffer from a cross-modal alignment bottleneck due to the limited representational capacity of their language encoders. Method: This paper is the first to attribute this bottleneck to insufficient effective mutual information (EMI) between modalities. We propose TinyAlign—a retrieval-augmented generation (RAG)-based framework that constructs an updatable multimodal memory bank and employs a lightweight connector to enable dynamic contextual injection and alignment optimization. The method integrates EMI-theoretic analysis, efficient memory retrieval, and parameter-efficient fine-tuning. Contribution/Results: TinyAlign substantially reduces training loss and accelerates convergence. It achieves baseline performance using only 40% of the fine-tuning data, consistently improving alignment quality and generalization across multiple downstream tasks. The approach significantly enhances data efficiency and adaptability to resource-constrained settings, without increasing model size or inference latency.
This work exposes a critical security vulnerability in expressive human pose and shape (EHPS) estimation models widely used in digital human generation: while existing methods prioritize estimation accuracy, they largely neglect robustness and adversarial resilience. To address this gap, we propose Tangible Attack (TBA), a novel framework featuring a dual heterogeneous noise generator (DHNG) and a customized adversarial loss function, integrated with VAE-based latent modeling, ControlNet conditioning, and multi-gradient iterative optimization. TBA enables cross-model, highly controllable, and strongly disruptive targeted adversarial attacks. Experiments demonstrate that TBA increases EHPS estimation error by an average of 17.0% and up to 41.0%, providing the first systematic evidence of severe security risks in mainstream digital human generation systems. Our work establishes a vital benchmark for evaluating model robustness and offers concrete directions for enhancing reliability and trustworthiness in expressive human modeling.
Existing vision-language model (VLM) training relies heavily on large-scale, manually collected image-text pairs, resulting in high acquisition costs and poor scalability. To address this, this paper proposes the first purely text-driven, three-stage multimodal data synthesis framework: starting from sparse textual seeds, it employs LLM-guided caption expansion, iterative instruction construction, and cross-modal representation transfer—from textual to visual embeddings—to automatically generate 1.2 million high-quality image-text pairs (Unicorn-1.2M) and 471K multi-turn instruction-tuning samples (Unicorn-471K-Instruction). Crucially, this approach eliminates the need for real images while achieving high diversity and fidelity in synthetic data. Evaluated across multiple benchmarks, models trained exclusively on our synthetic data match or approach the performance of those trained on real-image datasets. This work establishes a cost-effective, scalable paradigm for VLM data curation, significantly reducing reliance on expensive image collection and manual annotation.
本文提出SMTrap,一种基于SMT冲突指导的方法,生成计算密集型CSP查询以低成本实现对大型推理模型的DoS攻击。
Lightweight vision-language models (VLMs) suffer from a cross-modal alignment bottleneck due to the limited representational capacity of their language encoders. Method: This paper is the first to attribute this bottleneck to insufficient effective mutual information (EMI) between modalities. We propose TinyAlign—a retrieval-augmented generation (RAG)-based framework that constructs an updatable multimodal memory bank and employs a lightweight connector to enable dynamic contextual injection and alignment optimization. The method integrates EMI-theoretic analysis, efficient memory retrieval, and parameter-efficient fine-tuning. Contribution/Results: TinyAlign substantially reduces training loss and accelerates convergence. It achieves baseline performance using only 40% of the fine-tuning data, consistently improving alignment quality and generalization across multiple downstream tasks. The approach significantly enhances data efficiency and adaptability to resource-constrained settings, without increasing model size or inference latency.
This work exposes a critical security vulnerability in expressive human pose and shape (EHPS) estimation models widely used in digital human generation: while existing methods prioritize estimation accuracy, they largely neglect robustness and adversarial resilience. To address this gap, we propose Tangible Attack (TBA), a novel framework featuring a dual heterogeneous noise generator (DHNG) and a customized adversarial loss function, integrated with VAE-based latent modeling, ControlNet conditioning, and multi-gradient iterative optimization. TBA enables cross-model, highly controllable, and strongly disruptive targeted adversarial attacks. Experiments demonstrate that TBA increases EHPS estimation error by an average of 17.0% and up to 41.0%, providing the first systematic evidence of severe security risks in mainstream digital human generation systems. Our work establishes a vital benchmark for evaluating model robustness and offers concrete directions for enhancing reliability and trustworthiness in expressive human modeling.
Existing vision-language model (VLM) training relies heavily on large-scale, manually collected image-text pairs, resulting in high acquisition costs and poor scalability. To address this, this paper proposes the first purely text-driven, three-stage multimodal data synthesis framework: starting from sparse textual seeds, it employs LLM-guided caption expansion, iterative instruction construction, and cross-modal representation transfer—from textual to visual embeddings—to automatically generate 1.2 million high-quality image-text pairs (Unicorn-1.2M) and 471K multi-turn instruction-tuning samples (Unicorn-471K-Instruction). Crucially, this approach eliminates the need for real images while achieving high diversity and fidelity in synthetic data. Evaluated across multiple benchmarks, models trained exclusively on our synthetic data match or approach the performance of those trained on real-image datasets. This work establishes a cost-effective, scalable paradigm for VLM data curation, significantly reducing reliance on expensive image collection and manual annotation.