⚡ mehyar.jobs · free job search, free forever

Research Engineer, Interpretability

Anthropic · San Francisco, CA

💰 $5 – $10 USD/yr 📅 2026-08-21

About this role

<div class="content-intro"><h2><strong>About Anthropic</strong></h2> <p>Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems.</p></div><h2>About the role:</h2> <p>When you see what modern language models are capable of, do you wonder, "How do these things work? How can we trust them?"</p> <p>The Interpretability team at Anthropic is working to reverse-engineer how trained models work because we believe that a mechanistic understanding is the most robust way to make advanced systems safe.</p> <p>Think of us as doing "neuroscience" of neural networks using "microscopes" we build - or reverse-engineering neural networks like binary programs.</p> <p>More resources to learn about our work:&nbsp;</p> <ul> <li><a href="https://transformer-circuits.pub/">Our research blog</a> - covering advances including <a href="https://transformer-circuits.pub/2024/scaling-monosemanticity/">Monosemantic Features</a> and <a href="https://transformer-circuits.pub/2025/attribution-graphs/methods.html">Circuits</a></li> <li><a href="https://www.youtube.com/watch?v=TxhhMTOTMDg">An Introduction to Interpretability</a> from our research lead, <a href="https://colah.github.io/about.html">Chris Olah</a></li> <li><a href="https://www.darioamodei.com/post/the-urgency-of-interpretability">The Urgency of Interpretability</a> from CEO Dario Amodei</li> <li><a href="https://www.anthropic.com/research/engineering-challenges-interpretability">Engineering Challenges Scaling Interpretability</a> - directly relevant to this role</li> <li><a href="https://www.youtube.com/watch?v=aAPpQC-3EyE">60 Minutes</a> segment - Around 8:07, see a demo of tooling our team built</li> <li><a href="https://www.newyorker.com/magazine/2026/02/16/what-is-claude-anthropic-doesnt-know-either">New Yorker</a> article - what it's like to work on one of AI's hardest open problems</li> </ul> <p>Even if you haven’t worked on interpretability before, the infrastructure expertise is similar to what's needed across the lifecycle of a production language model:</p> <ul> <li>Pretraining: Training dictionary learning models looks a lot like model pretraining - creating stable, performant training jobs for massively parameterized models across thousands of chips</li> <li>Inference: Interp runs a customized inference stack. Day-to-day analysis requires services that allow editing a model's internal activations mid-forward-pass - for example, adding a "steering vector"</li> <li>Performance: Like all LLM work, we push up against the limits of hardware and software. Rather than squeezing the last 0.1%, we are focused on finding bottlenecks, fixing them and moving ahead given rapidly evolving research and safety mission</li> </ul> <p>The science keeps scaling - and it's now applied directly in <a href="https://www-cdn.anthropic.com/bf10f64990cfda0ba858290be7b8cc6317685f47.pdf">safety audits</a> on frontier models, with real deadlines. As our research has matured, engineering and infrastructure have become a bottleneck. Your work will have a direct impact on one of the most important open problems in AI.</p> <h2>Responsibilities:</h2> <ul> <li>Build and maintain the specialized inference and training infrastructure that powers interpretability research - including instrumented forward/backward passes, activation extraction, and steering vector application</li> <li>Resolve scaling and efficiency bottlenecks through profiling, optimization, and close collaboration with peer infrastructure teams</li> <li>Design tools, abstractions, and platforms that enable researchers to rapidly experiment without hitting engineering barriers</li> <li>Help bring interpretability research into production safety audits - with real deadlines and high reliability expectations</li> <li>Work across the stack - from model internals and accelerator-level optimization to user-facing research tooling</li> </ul> <h2>You may be a good fit if you:</h2> <ul> <li>Have 5-10+ years of experience building software</li> <li>Are highly proficient in at least one programming language (e.g., Python, Rust, Go, Java) and productive with Python</li> <li>Are extremely curious about unfamiliar domains; can quickly learn and put that knowledge to work, e.g. diving into new layers of the stack to find bottlenecks</li> <li>Have a strong ability to prioritize the most impactful work and are comfortable operating with ambiguity and questioning assumptions</li> <li>Prefer fast-moving collaborative projects to extensive solo efforts</li> <li>Are curious about interpretability research and its role in AI safety (though no research experience is required!)</li> <li>Care about the societal impacts and ethics of your work</li> <li>Are comfortable working closely with researchers, translating research needs into engineering solutions.</li> </ul> <h2>Strong candidates may also have experience with:</h2> <ul> <li>Optimizing the performance of large-scale distributed systems</li> <li>Language modeling fundamentals with transformers</li> <li>High Performance LLM optimization: memory management, compute efficiency, parallelism strategies, inference throughput optimization</li> <li>Working hands-on in a mainstream ML stack - PyTorch/CUDA on GPUs or JAX/XLA on TPUs</li> <li>Collaborating closely with researchers and building tooling to support research teams; or directly performed research with complex engineering challenges</li> </ul> <h2>Representative Projects:</h2> <ul> <li>Building <a href="https://transformer-circuits.pub/2021/garcon/index.html">Garcon</a>, a tool that allows researchers to easily instrument LLMs to extract internal activations</li> <li>Designing and optimizing a pipeline to efficiently collect petabytes of transformer activations and shuffle them</li> <li>Profiling and optimizing ML training jobs, including multi-GPU parallelism and memory optimization</li> <li>Building a steered inference system that applies targeted interventions to model internals at scale (conceptually similar to <a href="https://www.anthropic.com/news/golden-gate-claude">Golden Gate Claude</a> but for safety research)</li> </ul> <h2>Role Specific Location Policy:</h2> <ul> <li>This role is based in the San Francisco office; however, we are open to considering exceptional candidates for remote work on a case-by-case basis.</li> </ul><div class="content-pay-transparency"><div class="pay-input"><div class="description"><p>The annual compensation range for this role is listed below.&nbsp;</p> <p>For sales roles, the range provided is the role’s On Target Earnings ("OTE") range, meaning that the range includes both the sales commissions/sales bonuses target and annual base salary for the role.</p></div><div class="title">Annual Salary:</div><div class="pay-range"><span>$315,000</span><span class="divider">&mdash;</span><span>$560,000 USD</span></div></div></div><div class="content-conclusion"><h2><strong>Logistics</strong></h2> <p><strong>Minimum education: </strong>Bachelor’s degree or an equivalent combination of education, training, and/or experience</p> <p><strong>Required field of study:&nbsp;</strong>A field relevant to the role as demonstrated through coursework, training, or professional experience</p> <p><strong>Minimum years of experience: </strong>Years of experience required will correlate with the internal job level requirements for the position</p> <p><strong>Location-based hybrid policy:</strong> Currently, we expect all staff to be in one of our offices at least 25% of the time. However, some roles may require more time in our offices.</p> <p><strong data-stringify-type="bold">Visa sponsorship:</strong>&nbsp;We do sponsor visas! However, we aren't able to successfully sponsor visas for every role and every candidate. But if

Apply on Anthropic →

Want jobs like this matched to your resume, free? Create a free account — daily alerts, AI resume review, zero cost, forever.