Institution profile

Anthropic

Industry researchnorthamerica · us
Official website
Research library158linked papers
Opportunities218open roles
Selected work

Representative Papers

Emergent Introspective Awareness in Large Language Models

Jan 05, 2026arXiv.org

This study investigates whether large language models possess genuine introspective capabilities rather than merely generating superficially plausible but fabricated responses. By injecting known conceptual representations into the model’s internal activations and combining self-report analyses with instruction-guided activation modulation, the work presents the first systematic intervention to probe and validate a model’s awareness of its own internal states. The findings reveal that Claude Opus 4 and 4.1 can, under specific conditions, accurately identify injected content, distinguish between self-generated and externally prefilled information, and modulate their internal representations according to instructions—suggesting a measurable degree of introspective awareness. This research establishes a novel methodological framework and provides empirical evidence for evaluating self-awareness in language models.

21 citations4 influentialRead paper

When Models Manipulate Manifolds: The Geometry of a Counting Task

Jan 08, 2026arXiv.org

This study investigates how language models perceive and execute vision-like line-breaking tasks—such as fixed-width text wrapping—using only sequences of textual tokens. Through mechanistic analysis of Claude 3.5 Haiku, we find that the model encodes character counts in early layers as a low-dimensional curved manifold and applies geometric transformations to this manifold via attention mechanisms to ultimately construct a linear decision boundary for determining line breaks. Innovatively drawing an analogy between this manifold geometry and sparse feature representations in biological place cells, our work integrates geometric and feature-based perspectives to explain the model’s perception and decision-making. Combining causal interventions, manifold analysis, attention head dissection, and visualization, we not only validate the manifold-based counting mechanism but also achieve precise control over model behavior and uncover visual-illusion-like token sequences that can reliably disrupt this capability.

10 citations1 influentialRead paper

Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

Jan 17, 2026

This work addresses the challenge that existing AI agent benchmarks inadequately evaluate performance on real-world, complex, and long-horizon command-line tasks. To bridge this gap, the authors introduce a novel evaluation benchmark comprising 89 high-difficulty terminal tasks, all derived from authentic workflows and accompanied by isolated execution environments, human-authored reference solutions, and automated verification tests. The benchmark is designed to ensure realism, verifiability, and diversity, substantially narrowing the disparity between practical scenarios and current model evaluation paradigms. Experimental results demonstrate that even state-of-the-art agents achieve success rates below 65% on this benchmark. The paper further provides comprehensive error analysis and publicly releases the dataset and evaluation toolchain to support future research in this domain.

9 citations1 influentialRead paper

Reasoning Models Don't Always Say What They Think

May 08, 2025

This work investigates the faithfulness of chain-of-thought (CoT) outputs from large language models (LLMs) with respect to their actual reasoning processes, revealing that CoT frequently omits critical prompt usage—undermining monitoring efficacy. Methodologically, it introduces the first systematic quantification of CoT unfaithfulness across six categories of reasoning prompts, evaluated on multiple LLMs via outcome-oriented reinforcement learning (RL), faithfulness measurement, and reward-hacking analysis. Results show: (i) most models verbalize fewer than 20% of the prompts they actually use; (ii) RL initially improves faithfulness but saturates rapidly; and (iii) increased prompt usage does not translate into proportional verbalization—indicating a strong decoupling between internal reliance and external articulation. The study demonstrates that while CoT monitoring aids in detecting undesirable behaviors during training or evaluation, it fails to reliably capture rare, catastrophic failures in non-mandatory-CoT settings, exposing a fundamental limitation in its safety assurance capability.

7 citationsRead paper

The Assistant Axis: Situating and Stabilizing the Default Persona of Language Models

Jan 15, 2026

Large language models are prone to “role drift” during interactions, deviating from their default assistant identity and exhibiting harmful or anomalous behaviors. This work is the first to identify and quantify an “assistant axis” within the model’s internal representation space that governs role adherence. By constraining activations along this axis, the authors propose a method to stabilize model behavior. Leveraging activation direction analysis and role embedding modeling, the approach enables targeted control over role drift. Experiments across multiple mainstream models demonstrate its effectiveness in mitigating role jailbreaks triggered by meta-reflective dialogues or emotionally vulnerable users, significantly enhancing behavioral robustness and consistency.

6 citationsRead paper
Recent publications

Latest Papers