Atharv Naphade

Atharv Naphade

I am an undergrad studying Computer Science (AI track) at Carnegie Mellon University, where I maintain a 4.00 GPA. I work on post-training, alignment, reasoning, and model evaluation.

My work includes accepted papers at ACL 2026, ICML and ICLR workshops, and Nature Scientific Reports. I am currently a Research Intern at Applied Compute working on SoTA post-training, and an incoming Anthropic Research Fellow through MATS 11.0 for winter. Previously I completed public world model research at Roblox.

🔬 Research

I am interested in making frontier models more reliable: how they reason under uncertainty, how they aggregate evidence, and how alignment survives continued training. Recently this has meant three threads:

  1. Post-training: RLVR, test-time scaling, and world-model / multimodal post-training without collapsing diversity or safety.
  2. Alignment: Distillation, structural feedback, and methods that resist emergent misalignment.
  3. Evaluation: Introspection, hidden objectives, and better metrics than calibration for uncertainty.

🗞️ News

  • 💼 Incoming Anthropic Research Fellow through MATS 11.0.

  • 💼 Started as a Research Intern at Applied Compute, working on SoTA post-training.

  • 💼 Finished my internship at Roblox on public world model research, to be submitted to NeurIPS.

  • 📝 Paper accepted at ACL 2026: Rational Synthesizers or Heuristic Followers?

  • 📝 Spotlight (Top 3%) at ICML 2026 EIML for Rethinking Uncertainty Evaluation In Large Language Models.

  • 📝 Papers accepted at the ICML 2026 AI4GOOD and ICLR 2026 HCAIR workshops.

  • 🏆 Putnam 2025 — Top 270 among all students in North America.

📄 Selected Research

Uncertainty evaluation paper

Rethinking Uncertainty Evaluation In Large Language Models

Krishna Matta*, Atharv Naphade*, Andy Zou
ICML 2026 EIML Workshop

Reframes calibration as an incomplete uncertainty metric and introduces an exploitation-based view grounded in classical game theory.

Spotlight, Top 3% Evaluation Uncertainty
Conditioned diversity paper

Reinforcing Conditioned Diversity Optimizes Test Time Scaling

Atharv Naphade, Supriyo Chakraborty
COLM 2026   (under review)

Studies mode collapse during LLM post-training and introduces COLD, an RLVR algorithm rewarding conditional diversity.

Under review Post-training RLVR
Model diffing paper

Auditing LLMs for Hidden Behaviors via Model Diffing

Atharv Naphade*, Mukesh Ramanathan*, et al.
ICML 2026 AI4GOOD Workshop

Introduces adversarial decoding, a method for isolating unwanted behaviors in model organisms through low-probability tail distributions.

Accepted Safety Auditing
Mental states paper

Aligning Mental States in Large Language Models

Krishna Matta*, Atharv Naphade*, Andy Zou
NeurIPS 2026   (under review)

Develops a lightweight RL algorithm for aligning models to structural functions such as confidence, utility, and generalization.

Under review Alignment RL
Reasoning emergence paper

On the Emergence of Reasoning

Pratheek Humane, Supriyo Chakraborty, Atharv Naphade, et al.
MILA Institute

Proposes a conditional probabilistic framework for studying the importance of subthoughts in chain-of-thought reasoning.

Reasoning Theory

💼 Experience

  • Research Fellow — Anthropic ( MATS 11.0) Winter 2026 (Incoming)

    Incoming Anthropic Research Fellow through MATS 11.0.

  • Research Intern — Applied Compute August 2026 – Present

    SoTA post-training.

  • Intern — Roblox Summer 2026

    Public world model research; work to be submitted to NeurIPS.

  • Jane Street FTTP Spring 2026

    Highly selective 1-week Trading and Technology Program. 1 of 60 invitees out of thousands of applicants.

  • Research Fellow — SPAR Spring 2026

    Working on jailbreaks for the AI Safety stream.

  • Research Scientist Intern — CMU Robotics Department Fall 2025

    Scaling up RL post-training of Vision Language Models. Collaboration with NVIDIA Researchers.

  • Research Engineer — Refactor ( YC S24) Summer 2025

    Improved robustness of Lowe’s AI at scale by deploying novel RLVR environments. Implemented 11+ full-stack infrastructure features in SQL, Redis, and Next.js for scalable LLM evaluation including multi-turn evals, error tracking & mitigation, and efficient guardrails. First hire.

  • Machine Learning Engineer — Iowa State University 2024

    Built video-based deep learning models to detect and report risky driving behaviors in real-time using PyTorch & DeepStream. Algorithm deployed on 260+ highway cameras under Professor Anuj Sharma.

🏆 Awards

  • Putnam 2025 — Top 270 among all students in North America
  • USAMTS Medalist; 2× BAMO Award Winner
  • Stanford University Mathematics Camp Student Researcher (focus: Gradient Fields)
  • Stanford Math Tournament — 1st place / 2200 Individual
  • 5× AIME Qualifier; Top 250 USAMO Index
  • USACO Gold (Silver Perfect Score)
  • Math Kangaroo National Champion (1st in USA)

🔖 Misc

I create educational content explaining AI research for a general audience on @agi_atharv. 17k followers, 1M+ views.