AgoraResearch hub
ExploreLibraryProfile
Account
Sign In
Andy Arditi
Scholar

Andy Arditi

Google Scholar ID: NgyIgX4AAAAJ
Northeastern University
Interpretability
Homepage↗Google Scholar↗
Citations & Impact
All-time
Citations
367
 
H-index
6
 
i10-index
5
 
Publications
10
 
Co-authors
7
list available
Contact
Emailandyrdt@gmail.comTwitterOpen ↗
Publications
6 items
Synthetic Persona Pretraining: Alignment from Token Zero
2026
Cited
0
Weird Generalization and Inductive Backdoors: New Ways to Corrupt LLMs
2025
Cited
0
Persona Vectors: Monitoring and Controlling Character Traits in Language Models
2025
Cited
0
Inverse Scaling in Test-Time Compute
2025
Cited
0
Adversarial Manipulation of Reasoning Models using Internal Representations
2025
Cited
0
Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning
2025
Cited
0
Resume (English only)
Background
  • Research Interest: AI interpretability
Miscellany
  • Contact Information: andyrdt@gmail.com, andyarditi
Co-authors
7 total
Neel Nanda
Neel Nanda
Mechanistic Interpretability Team Lead, Google DeepMind
Co-author 2
Co-author 2
Wes Gurnee
Wes Gurnee
Anthropic
Nina Panickssery
Nina Panickssery
Anthropic
Daniel Paleka
Daniel Paleka
ETH Zurich
Runjin Chen
Runjin Chen
PHD student at UT Austin
Jack Lindsey
Jack Lindsey
Anthropic