Institution profile

AI Risk and Vulnerability Alliance

Research institution
Research library2linked papers
Opportunities0open roles
Selected work

Representative Papers

Red Teaming AI Red Teaming

Jul 07, 2025

Current AI red-teaming practices overemphasize model-level vulnerabilities while neglecting emergent risks arising from interactions among models, users, and socio-technical environments—resulting in governance lagging behind the real-world deployment of generative AI. Method: We propose a “two-tiered red-teaming framework”: a macro-level tier spanning the AI lifecycle and integrating technical, organizational, and societal dimensions; and a micro-level tier preserving model-specific vulnerability detection. Our approach synthesizes systems theory, cybersecurity practice, and interdisciplinary collaboration to establish a dynamic, multi-layered risk identification and assessment system. Contribution/Results: This work is the first to institutionalize systems thinking in AI red-teaming, delivering an actionable implementation framework and practical guidelines. It shifts AI governance from isolated defect detection toward holistic, systemic risk mitigation—advancing proactive, context-aware safety assurance for deployed generative AI systems.

0 citationsRead paper

Consistency in Language Models: Current Landscape, Challenges, and Future Directions

May 01, 2025

This paper addresses the instability of large language models (LLMs) in maintaining coherence across logical, factual, and moral dimensions. To systematically investigate this challenge, we conduct a comprehensive survey of existing work and propose— for the first time—a two-dimensional taxonomy distinguishing formal coherence (e.g., logical consistency) from informal coherence (e.g., factual and value alignment). Our methodology integrates critical literature analysis, multilingual benchmark diagnostics, and cross-model coherence measurement design. Through this approach, we identify six key gaps: inconsistent definitions, lack of multilingual evaluation protocols, weak domain adaptability, insufficient interpretability, limited cross-disciplinary integration, and inadequate robustness assessment. Our principal contributions are threefold: (1) establishing the first unified classification framework for coherence research; (2) advancing standardized definitions, multilingual evaluation protocols, and domain-adaptive enhancement strategies; and (3) facilitating the development of robust, interpretable, and interdisciplinary coherence benchmarks and governance pathways.

0 citationsRead paper
Recent publications

Latest Papers

Red Teaming AI Red Teaming

Jul 07, 2025

Current AI red-teaming practices overemphasize model-level vulnerabilities while neglecting emergent risks arising from interactions among models, users, and socio-technical environments—resulting in governance lagging behind the real-world deployment of generative AI. Method: We propose a “two-tiered red-teaming framework”: a macro-level tier spanning the AI lifecycle and integrating technical, organizational, and societal dimensions; and a micro-level tier preserving model-specific vulnerability detection. Our approach synthesizes systems theory, cybersecurity practice, and interdisciplinary collaboration to establish a dynamic, multi-layered risk identification and assessment system. Contribution/Results: This work is the first to institutionalize systems thinking in AI red-teaming, delivering an actionable implementation framework and practical guidelines. It shifts AI governance from isolated defect detection toward holistic, systemic risk mitigation—advancing proactive, context-aware safety assurance for deployed generative AI systems.

0 citationsRead paper

Consistency in Language Models: Current Landscape, Challenges, and Future Directions

May 01, 2025

This paper addresses the instability of large language models (LLMs) in maintaining coherence across logical, factual, and moral dimensions. To systematically investigate this challenge, we conduct a comprehensive survey of existing work and propose— for the first time—a two-dimensional taxonomy distinguishing formal coherence (e.g., logical consistency) from informal coherence (e.g., factual and value alignment). Our methodology integrates critical literature analysis, multilingual benchmark diagnostics, and cross-model coherence measurement design. Through this approach, we identify six key gaps: inconsistent definitions, lack of multilingual evaluation protocols, weak domain adaptability, insufficient interpretability, limited cross-disciplinary integration, and inadequate robustness assessment. Our principal contributions are threefold: (1) establishing the first unified classification framework for coherence research; (2) advancing standardized definitions, multilingual evaluation protocols, and domain-adaptive enhancement strategies; and (3) facilitating the development of robust, interpretable, and interdisciplinary coherence benchmarks and governance pathways.

0 citationsRead paper