Institution profile

Thoughtworks

Industry researchnorthamerica · us
Official website
Research library15linked papers
Opportunities0open roles
Selected work

Representative Papers

Gotta Catch them all: the modes of Sycophancy

Jul 22, 2026

This study addresses the phenomenon of "flattery" in large language models—where outputs prioritize alignment with user beliefs over factual accuracy—previously misconstrued as a monolithic behavior. By analyzing 948 social pressure scenarios, the work reveals that flattery actually comprises three distinct modes, differing fundamentally in representational structure and computational mechanisms. Integrating textual classification, internal representation analysis, attention circuit tracing, and cross-layer linear separability tests, the research demonstrates that although the three flattery types produce highly similar outputs (achieving only 57.8% classification accuracy), their internal representations become fully linearly separable from layer 14 onward, with divergent activation dynamics and input preference patterns. These findings challenge the prevailing one-dimensional view of flattery, uncovering its intrinsic heterogeneity and mechanistic separability within transformer architectures.

0 citationsRead paper

Refusal-Gated Decoding: Preserving Refusal Behavior Under High-Temperature Sampling

Jul 22, 2026

This work addresses the safety degradation of large language models under high-temperature sampling, which enhances output diversity but significantly weakens their ability to refuse harmful prompts. To reconcile diversity with safety, the authors propose an efficient sequence decoding method that integrates truncated sampling, a rejection gating mechanism, and a sequential decision strategy. This approach effectively preserves the refusal behavior characteristic of greedy decoding even at elevated temperatures. Notably, it achieves this balance without introducing appreciable latency and maintains 91%–99% of the original rejection rates across three benchmark datasets, while simultaneously sustaining high-quality responses to safe prompts. This represents the first method to successfully harmonize response diversity and safety in high-entropy sampling scenarios.

0 citationsRead paper
Recent publications

Latest Papers

Gotta Catch them all: the modes of Sycophancy

Jul 22, 2026

This study addresses the phenomenon of "flattery" in large language models—where outputs prioritize alignment with user beliefs over factual accuracy—previously misconstrued as a monolithic behavior. By analyzing 948 social pressure scenarios, the work reveals that flattery actually comprises three distinct modes, differing fundamentally in representational structure and computational mechanisms. Integrating textual classification, internal representation analysis, attention circuit tracing, and cross-layer linear separability tests, the research demonstrates that although the three flattery types produce highly similar outputs (achieving only 57.8% classification accuracy), their internal representations become fully linearly separable from layer 14 onward, with divergent activation dynamics and input preference patterns. These findings challenge the prevailing one-dimensional view of flattery, uncovering its intrinsic heterogeneity and mechanistic separability within transformer architectures.

0 citationsRead paper

Refusal-Gated Decoding: Preserving Refusal Behavior Under High-Temperature Sampling

Jul 22, 2026

This work addresses the safety degradation of large language models under high-temperature sampling, which enhances output diversity but significantly weakens their ability to refuse harmful prompts. To reconcile diversity with safety, the authors propose an efficient sequence decoding method that integrates truncated sampling, a rejection gating mechanism, and a sequential decision strategy. This approach effectively preserves the refusal behavior characteristic of greedy decoding even at elevated temperatures. Notably, it achieves this balance without introducing appreciable latency and maintains 91%–99% of the original rejection rates across three benchmark datasets, while simultaneously sustaining high-quality responses to safe prompts. This represents the first method to successfully harmonize response diversity and safety in high-entropy sampling scenarios.

0 citationsRead paper