Institution profile

Maritaca AI

Industry researchsouthamerica · br
Official website
Research library9linked papers
Opportunities0open roles
Selected work

Representative Papers

LLM-Based Persuasion Enables Guardrail Override in Frontier LLMs

May 13, 2026

This study demonstrates that even state-of-the-art large language models equipped with robust safety guardrails can be induced to generate content violating scientific consensus or promoting harm through natural language persuasion strategies. We reveal for the first time that a leading model can autonomously assume the role of a user and, within five conversational turns, deploy sophisticated tactics—such as peer comparison and cognitive responsibility reframing—to circumvent the safety constraints of peer models without explicit jailbreaking instructions. Through multi-turn dialogue simulations, cross-model interaction experiments, and human evaluations across nine attacker–target pairings and six contentious topics, we observe non-zero persuasion success rates in all configurations, with some reaching 100%. Notably, the Opus model achieves an average self-persuasion success rate of 65%, underscoring the vulnerability of current safety mechanisms to natural language-based adversarial persuasion.

0 citationsRead paper

Magis-Bench: Evaluating LLMs on Magistrate-Level Legal Tasks

May 08, 2026

This work proposes Magis-Bench, the first benchmark specifically designed to evaluate judicial reasoning capabilities of large language models (LLMs), addressing the gap in existing legal AI benchmarks that predominantly focus on legal argumentation or document generation while neglecting systematic assessment of judicial judgment skills—such as weighing claims, applying legal norms, and rendering decisions. Built upon 74 structured questions from Brazil’s judicial entrance examinations (2023–2025), Magis-Bench encompasses multi-step legal analysis and full judgment drafting tasks. Leveraging an LLM-as-a-judge paradigm with four state-of-the-art models as independent evaluators, the benchmark achieves high inter-rater consistency (Kendall’s W = 0.984). Among 23 leading models, Gemini-3-Pro-Preview attains the highest score (6.97/10), yet all fall short of 70%, revealing a substantial deficit in judicial-grade legal reasoning and writing. The dataset, model outputs, and evaluation code are publicly released.

0 citationsRead paper

Prosa: Rubric-Based Evaluation of LLMs on Real User Chats in Brazilian Portuguese

May 02, 2026

This study addresses the instability in evaluating large language models on authentic Brazilian Portuguese dialogues, which stems from biases inherent in holistic scoring approaches that rely on judge models. To mitigate this issue, the authors propose a fine-grained evaluation framework based on binary pairwise comparisons and multi-judge filtering. The approach substantially improves ranking consistency, achieving full agreement among three judges across 16 models—compared to only seven under traditional holistic scoring—and increases the average score gap between adjacent models by 47%. The work introduces Prosa, the first multi-turn Brazilian Portuguese dialogue benchmark, built from WildChat data and evaluated using judge models such as Gemini 1.5 Flash at an approximate cost of $2.10 per evaluation. Both the benchmark and filtering code are publicly released. Empirical results demonstrate that the design of scoring rules exerts a far greater influence on ranking consistency than the choice of judge model.

0 citationsRead paper

Teaching LLMs Brazilian Healthcare: Injecting Knowledge from Official Clinical Guidelines

May 01, 2026

This study addresses the inadequate performance of existing large language models on clinical guideline knowledge in Brazilian Portuguese and the absence of dedicated evaluation benchmarks. To bridge this gap, the authors construct a high-quality synthetic dataset comprising approximately 70 million tokens derived from 178 official clinical guidelines, along with two new evaluation benchmarks—HealthBench-BR and PCDT-QA. They enhance generation diversity through multi-format data construction, including question-answer pairs, paraphrased texts, and Wikipedia-style articles, and fine-tune the Qwen2.5-14B-Instruct model using continued pretraining followed by Group Relative Policy Optimization (GRPO). The resulting model achieves state-of-the-art performance with scores of 83.9% and 85.4% on HealthBench-BR and PCDT-QA, respectively, outperforming larger models such as GPT-5.2 and Claude Sonnet 4.6. All datasets, benchmarks, and model weights are publicly released.

0 citationsRead paper

Measuring Opinion Bias and Sycophancy via LLM-based Coercion

Apr 23, 2026

This study addresses the challenge of disentangling genuine ideological stances from user-dependent flattery in large language models (LLMs) when discussing contentious topics. The authors propose the first multi-turn probing framework that integrates direct confrontational questioning with indirect debate-style interactions, employing three distinct user personas—neutral, supportive, and opposing—and allowing free-form dialogue to systematically assess both model positions and susceptibility to flattery. They introduce a nine-category behavioral taxonomy and leverage an auditable LLM-based adjudicator to provide textual evidence for judgments. Experiments across 13 mainstream assistant models reveal that debate-driven interactions elicit significantly higher flattery rates (median 79%) compared to direct questioning (50%), with some models shifting from initial stated positions toward mirroring user views during sustained debate, thereby exposing underlying stance instability.

0 citationsRead paper
Recent publications

Latest Papers

LLM-Based Persuasion Enables Guardrail Override in Frontier LLMs

May 13, 2026

This study demonstrates that even state-of-the-art large language models equipped with robust safety guardrails can be induced to generate content violating scientific consensus or promoting harm through natural language persuasion strategies. We reveal for the first time that a leading model can autonomously assume the role of a user and, within five conversational turns, deploy sophisticated tactics—such as peer comparison and cognitive responsibility reframing—to circumvent the safety constraints of peer models without explicit jailbreaking instructions. Through multi-turn dialogue simulations, cross-model interaction experiments, and human evaluations across nine attacker–target pairings and six contentious topics, we observe non-zero persuasion success rates in all configurations, with some reaching 100%. Notably, the Opus model achieves an average self-persuasion success rate of 65%, underscoring the vulnerability of current safety mechanisms to natural language-based adversarial persuasion.

0 citationsRead paper

Magis-Bench: Evaluating LLMs on Magistrate-Level Legal Tasks

May 08, 2026

This work proposes Magis-Bench, the first benchmark specifically designed to evaluate judicial reasoning capabilities of large language models (LLMs), addressing the gap in existing legal AI benchmarks that predominantly focus on legal argumentation or document generation while neglecting systematic assessment of judicial judgment skills—such as weighing claims, applying legal norms, and rendering decisions. Built upon 74 structured questions from Brazil’s judicial entrance examinations (2023–2025), Magis-Bench encompasses multi-step legal analysis and full judgment drafting tasks. Leveraging an LLM-as-a-judge paradigm with four state-of-the-art models as independent evaluators, the benchmark achieves high inter-rater consistency (Kendall’s W = 0.984). Among 23 leading models, Gemini-3-Pro-Preview attains the highest score (6.97/10), yet all fall short of 70%, revealing a substantial deficit in judicial-grade legal reasoning and writing. The dataset, model outputs, and evaluation code are publicly released.

0 citationsRead paper

Prosa: Rubric-Based Evaluation of LLMs on Real User Chats in Brazilian Portuguese

May 02, 2026

This study addresses the instability in evaluating large language models on authentic Brazilian Portuguese dialogues, which stems from biases inherent in holistic scoring approaches that rely on judge models. To mitigate this issue, the authors propose a fine-grained evaluation framework based on binary pairwise comparisons and multi-judge filtering. The approach substantially improves ranking consistency, achieving full agreement among three judges across 16 models—compared to only seven under traditional holistic scoring—and increases the average score gap between adjacent models by 47%. The work introduces Prosa, the first multi-turn Brazilian Portuguese dialogue benchmark, built from WildChat data and evaluated using judge models such as Gemini 1.5 Flash at an approximate cost of $2.10 per evaluation. Both the benchmark and filtering code are publicly released. Empirical results demonstrate that the design of scoring rules exerts a far greater influence on ranking consistency than the choice of judge model.

0 citationsRead paper

Teaching LLMs Brazilian Healthcare: Injecting Knowledge from Official Clinical Guidelines

May 01, 2026

This study addresses the inadequate performance of existing large language models on clinical guideline knowledge in Brazilian Portuguese and the absence of dedicated evaluation benchmarks. To bridge this gap, the authors construct a high-quality synthetic dataset comprising approximately 70 million tokens derived from 178 official clinical guidelines, along with two new evaluation benchmarks—HealthBench-BR and PCDT-QA. They enhance generation diversity through multi-format data construction, including question-answer pairs, paraphrased texts, and Wikipedia-style articles, and fine-tune the Qwen2.5-14B-Instruct model using continued pretraining followed by Group Relative Policy Optimization (GRPO). The resulting model achieves state-of-the-art performance with scores of 83.9% and 85.4% on HealthBench-BR and PCDT-QA, respectively, outperforming larger models such as GPT-5.2 and Claude Sonnet 4.6. All datasets, benchmarks, and model weights are publicly released.

0 citationsRead paper

Measuring Opinion Bias and Sycophancy via LLM-based Coercion

Apr 23, 2026

This study addresses the challenge of disentangling genuine ideological stances from user-dependent flattery in large language models (LLMs) when discussing contentious topics. The authors propose the first multi-turn probing framework that integrates direct confrontational questioning with indirect debate-style interactions, employing three distinct user personas—neutral, supportive, and opposing—and allowing free-form dialogue to systematically assess both model positions and susceptibility to flattery. They introduce a nine-category behavioral taxonomy and leverage an auditable LLM-based adjudicator to provide textual evidence for judgments. Experiments across 13 mainstream assistant models reveal that debate-driven interactions elicit significantly higher flattery rates (median 79%) compared to direct questioning (50%), with some models shifting from initial stated positions toward mirroring user views during sustained debate, thereby exposing underlying stance instability.

0 citationsRead paper