Neel Rajani

Neel_Headshot.png

The website template wants an address here.

But I'm not sure if I want that to be public 😅

Hi! My name is Neel, and I do research on AI Safety. I’m currently doing a PhD at the CDT for Designing Responsible NLP at the University of Edinburgh. My work focuses on the alignment of LLMs, by researching how this is typically done during post-training, why it is so brittle, and what this means for safety more broadly. AI is clearly a transformational technology, and I think that it is critical for us to get a better understanding of its guardrails and when they break.

news

Sep 01, 2026 Our paper on CoT Faithfulness under cue effects is up on arXiv! :sparkles: Check out Aryo’s twitter thread for a pitch.
Jul 27, 2026 I’ve started a 3-month research visit at Mila in Montreal, where I’m joining Prof Siva Reddy’s lab and working on a project with the amazing Dr Verna Dankers.
Nov 07, 2015 This is a placeholder announcement left over from the template that I may repurpose later...

latest posts

selected publications

  1. Preprint
    FACE_eval.png
    Chain-of-Thought Faithfulness of Reasoning Models Varies with Where and How Preference Cues Are Delivered
    Aryo Pradipta Gema, Neel Rajani, Rohit Saxena, and 2 more authors
    2026
  2. DIS ’26
    SIBB_UI.png
    Stepping Into the Black Box: Opening Up LLMs to Public Exploration Through Discursive Design
    Sarah G Immel, Neel Rajani, Rayo Verweij, and 4 more authors
    In Proceedings of the 2026 Designing Interactive Systems Conference, Singapore, 2026
  3. AIW@ICML’25
    Scalpel_vs_Hammer.gif
    Scalpel vs. Hammer: GRPO Amplifies Existing Capabilities, SFT Replaces Them
    Neel Rajani, Aryo Pradipta Gema, Seraphina Goldfarb-Tarrant, and 1 more author
    In Actionable Interpretability Workshop at ICML 2025, 2025