Institution profile

Giskard AI

Industry researcheurope · fr
Official website
Research library2linked papers
Opportunities0open roles
Selected work

Representative Papers

Phare: A Safety Probe for Large Language Models

May 16, 2025

Existing LLM safety evaluations predominantly focus on performance ranking, neglecting systematic failure-mode identification. This paper introduces Phare—the first multilingual safety assessment framework centered on failure-mode diagnosis—systematically uncovering model vulnerabilities across three dimensions: hallucination and reliability, social bias, and harmful content generation. Phare innovatively integrates adversarial prompt engineering, controllable behavioral sampling, and fine-grained human annotation to identify concrete risk patterns—including sycophancy, prompt sensitivity, and stereotype reiteration—for the first time. Evaluated on 17 mainstream models, it reveals cross-lingual and cross-architectural common vulnerabilities. Furthermore, Phare provides actionable, model-agnostic improvement pathways. By shifting evaluation from aggregate scoring to diagnostic analysis, it advances the development of more robust, aligned, and trustworthy language systems.

0 citationsRead paper

RealHarm: A Collection of Real-World Language Model Application Failures

Apr 14, 2025

This paper addresses the lack of empirical failure analysis for language models deployed in consumer-facing applications. To bridge this gap, we introduce RealHarm—the first publicly available, real-world AI application failure dataset grounded in verified incidents. Our methodology involves systematic event mining from public sources and multi-dimensional human annotation, adopting a deployer-centric perspective to classify harm types (e.g., reputational damage, misinformation), root causes, and risk propagation pathways; we further empirically evaluate the efficacy of mainstream content safety guardrails against these authentic failures. Key contributions include: (1) establishing the first empirically grounded, organization-level AI failure repository; (2) revealing a substantial misalignment between regulatory frameworks and observed operational risks; (3) identifying reputational damage as the most prevalent organizational harm and misinformation as the dominant risk category; and (4) demonstrating critically low interception rates of existing safety systems against real-world failure instances.

0 citationsRead paper
Recent publications

Latest Papers

Phare: A Safety Probe for Large Language Models

May 16, 2025

Existing LLM safety evaluations predominantly focus on performance ranking, neglecting systematic failure-mode identification. This paper introduces Phare—the first multilingual safety assessment framework centered on failure-mode diagnosis—systematically uncovering model vulnerabilities across three dimensions: hallucination and reliability, social bias, and harmful content generation. Phare innovatively integrates adversarial prompt engineering, controllable behavioral sampling, and fine-grained human annotation to identify concrete risk patterns—including sycophancy, prompt sensitivity, and stereotype reiteration—for the first time. Evaluated on 17 mainstream models, it reveals cross-lingual and cross-architectural common vulnerabilities. Furthermore, Phare provides actionable, model-agnostic improvement pathways. By shifting evaluation from aggregate scoring to diagnostic analysis, it advances the development of more robust, aligned, and trustworthy language systems.

0 citationsRead paper

RealHarm: A Collection of Real-World Language Model Application Failures

Apr 14, 2025

This paper addresses the lack of empirical failure analysis for language models deployed in consumer-facing applications. To bridge this gap, we introduce RealHarm—the first publicly available, real-world AI application failure dataset grounded in verified incidents. Our methodology involves systematic event mining from public sources and multi-dimensional human annotation, adopting a deployer-centric perspective to classify harm types (e.g., reputational damage, misinformation), root causes, and risk propagation pathways; we further empirically evaluate the efficacy of mainstream content safety guardrails against these authentic failures. Key contributions include: (1) establishing the first empirically grounded, organization-level AI failure repository; (2) revealing a substantial misalignment between regulatory frameworks and observed operational risks; (3) identifying reputational damage as the most prevalent organizational harm and misinformation as the dominant risk category; and (4) demonstrating critically low interception rates of existing safety systems against real-world failure instances.

0 citationsRead paper