Institution profile

DATUMO Inc.

Industry researchnorthamerica · us
Research library9linked papers
Opportunities0open roles
Selected work

Representative Papers

FFinRED: An Expert-Guided Benchmark Generation and Evaluation Framework for Financial LLM Red-Teaming

Jun 18, 2026

This study addresses the inadequacy of existing safety evaluation benchmarks in capturing domain-specific financial risks—such as regulatory non-compliance, fraud inducement, and systemic trust erosion—by proposing the first two-tier threat taxonomy that integrates global financial regulatory standards (e.g., ISO/IEC 27001) with expert knowledge. Building upon this framework, the authors generate context-rich red-teaming prompt seeds derived from real-world financial documents to construct a scalable safety evaluation framework for large language models in finance. Deployed within the regulatory sandbox of the Korea Financial Security Institute, the approach employs expert-validated assessment rubrics that reduce critical false positive rates from 28% to 12%, substantially outperforming generic static rubrics and enabling high-fidelity, operationally viable AI safety evaluations in financial contexts.

0 citationsRead paper

Reliable to Expressive: A Curriculum for Rubric-Following Safety Judges

Jun 08, 2026

This work addresses the fragility of existing safety evaluation models when confronted with variations in scoring criteria and prompts, which undermines their ability to consistently adhere to diverse judgment standards. The authors frame safety assessment as a criterion-following problem and propose a curriculum learning framework that progresses from “reliable” to “expressive” behaviors. By integrating dynamically generated instance-conditional scoring rubrics with supervised fine-tuning, they train a 12B-parameter language model to robustly align with shifting evaluation guidelines. Their approach is the first to systematically resolve judgment instability under varying criteria, achieving accuracies of 94.12%–94.88% across three distinct scoring standards with a remarkably low cross-criterion performance variance of only 0.76—significantly outperforming general-purpose large language models, specialized safety classifiers, and reasoning-based evaluators with fewer than 30B parameters.

0 citationsRead paper

Culturally-Adapted Red-Teaming Across East and Southeast Asian Contexts: A Methodological and Comparative Analysis

Jun 08, 2026

This study addresses the limitations of current safety evaluations for multilingual large language models, which often rely on literal translations of English benchmarks and overlook cultural contextual differences, leading to inaccurate risk assessments. The authors present the first systematic construction of paired red-teaming datasets in Korean, Japanese, Thai, and Khmer, featuring both direct translation (DT) and culturally adapted (CA) prompts matched one-to-one by seed. Evaluation using attack success rate (ASR) and a novel Cultural Contextualization Score (C3) reveals that CA prompts increase ASR by an average of 9.3 percentage points across all 16 language–model combinations. Furthermore, DT significantly underestimates local threats in 44 out of 48 risk categories, while C3 scores rise from a mean of 0.17 to as high as 2.51, demonstrating that cultural adaptation is essential for accurately capturing localized safety risks.

0 citationsRead paper

STAR-Teaming: A Strategy-Response Multiplex Network Approach to Automated LLM Red Teaming

Apr 20, 2026

This work addresses the vulnerability of large language models to jailbreaking attacks that elicit harmful outputs. To counter this, the authors propose the STAR-Teaming framework, which introduces for the first time a strategy-response multiplex network to model red-teaming dynamics. By integrating multi-agent systems with network-driven optimization, the framework reconstructs interpretable semantic community structures in high-dimensional embedding spaces to guide efficient adversarial sampling. This approach substantially improves attack success rates while reducing computational overhead, achieving both high efficiency and strong interpretability.

0 citationsRead paper

CAC-CoT: Connector-Aware Compact Chain-of-Thought for Efficient Reasoning Data Synthesis Across Dual-System Cognitive Tasks

Aug 26, 2025

To address reasoning redundancy and efficiency degradation induced by chain-of-thought (CoT) prompting in “System 1” intuitive tasks, this paper proposes Connector-Aware Compact CoT (CAC-CoT). CAC-CoT introduces a connective-word constraint mechanism that retains only essential logical connectives (e.g., “therefore”, “because”) to generate structurally coherent, length-controllable compact reasoning chains. Unlike conventional free-form CoT generation, CAC-CoT constructs reasoning paths via fixed connective templates and leverages Gemini-2.0-Flash to synthesize high-quality training data. Experiments demonstrate that CAC-CoT achieves 85%, 40%, and 90% of the original performance on GSM8K, GPQA, and S1-Bench, respectively, while compressing average reasoning length to ~300 tokens and accelerating inference by 3×. This work marks the first approach enabling efficient joint modeling of System 1 (intuitive) and System 2 (analytic) reasoning tasks.

0 citationsRead paper
Recent publications

Latest Papers

FFinRED: An Expert-Guided Benchmark Generation and Evaluation Framework for Financial LLM Red-Teaming

Jun 18, 2026

This study addresses the inadequacy of existing safety evaluation benchmarks in capturing domain-specific financial risks—such as regulatory non-compliance, fraud inducement, and systemic trust erosion—by proposing the first two-tier threat taxonomy that integrates global financial regulatory standards (e.g., ISO/IEC 27001) with expert knowledge. Building upon this framework, the authors generate context-rich red-teaming prompt seeds derived from real-world financial documents to construct a scalable safety evaluation framework for large language models in finance. Deployed within the regulatory sandbox of the Korea Financial Security Institute, the approach employs expert-validated assessment rubrics that reduce critical false positive rates from 28% to 12%, substantially outperforming generic static rubrics and enabling high-fidelity, operationally viable AI safety evaluations in financial contexts.

0 citationsRead paper

Reliable to Expressive: A Curriculum for Rubric-Following Safety Judges

Jun 08, 2026

This work addresses the fragility of existing safety evaluation models when confronted with variations in scoring criteria and prompts, which undermines their ability to consistently adhere to diverse judgment standards. The authors frame safety assessment as a criterion-following problem and propose a curriculum learning framework that progresses from “reliable” to “expressive” behaviors. By integrating dynamically generated instance-conditional scoring rubrics with supervised fine-tuning, they train a 12B-parameter language model to robustly align with shifting evaluation guidelines. Their approach is the first to systematically resolve judgment instability under varying criteria, achieving accuracies of 94.12%–94.88% across three distinct scoring standards with a remarkably low cross-criterion performance variance of only 0.76—significantly outperforming general-purpose large language models, specialized safety classifiers, and reasoning-based evaluators with fewer than 30B parameters.

0 citationsRead paper

Culturally-Adapted Red-Teaming Across East and Southeast Asian Contexts: A Methodological and Comparative Analysis

Jun 08, 2026

This study addresses the limitations of current safety evaluations for multilingual large language models, which often rely on literal translations of English benchmarks and overlook cultural contextual differences, leading to inaccurate risk assessments. The authors present the first systematic construction of paired red-teaming datasets in Korean, Japanese, Thai, and Khmer, featuring both direct translation (DT) and culturally adapted (CA) prompts matched one-to-one by seed. Evaluation using attack success rate (ASR) and a novel Cultural Contextualization Score (C3) reveals that CA prompts increase ASR by an average of 9.3 percentage points across all 16 language–model combinations. Furthermore, DT significantly underestimates local threats in 44 out of 48 risk categories, while C3 scores rise from a mean of 0.17 to as high as 2.51, demonstrating that cultural adaptation is essential for accurately capturing localized safety risks.

0 citationsRead paper

STAR-Teaming: A Strategy-Response Multiplex Network Approach to Automated LLM Red Teaming

Apr 20, 2026

This work addresses the vulnerability of large language models to jailbreaking attacks that elicit harmful outputs. To counter this, the authors propose the STAR-Teaming framework, which introduces for the first time a strategy-response multiplex network to model red-teaming dynamics. By integrating multi-agent systems with network-driven optimization, the framework reconstructs interpretable semantic community structures in high-dimensional embedding spaces to guide efficient adversarial sampling. This approach substantially improves attack success rates while reducing computational overhead, achieving both high efficiency and strong interpretability.

0 citationsRead paper

CAC-CoT: Connector-Aware Compact Chain-of-Thought for Efficient Reasoning Data Synthesis Across Dual-System Cognitive Tasks

Aug 26, 2025

To address reasoning redundancy and efficiency degradation induced by chain-of-thought (CoT) prompting in “System 1” intuitive tasks, this paper proposes Connector-Aware Compact CoT (CAC-CoT). CAC-CoT introduces a connective-word constraint mechanism that retains only essential logical connectives (e.g., “therefore”, “because”) to generate structurally coherent, length-controllable compact reasoning chains. Unlike conventional free-form CoT generation, CAC-CoT constructs reasoning paths via fixed connective templates and leverages Gemini-2.0-Flash to synthesize high-quality training data. Experiments demonstrate that CAC-CoT achieves 85%, 40%, and 90% of the original performance on GSM8K, GPQA, and S1-Bench, respectively, while compressing average reasoning length to ~300 tokens and accelerating inference by 3×. This work marks the first approach enabling efficient joint modeling of System 1 (intuitive) and System 2 (analytic) reasoning tasks.

0 citationsRead paper