Existing approaches to combat the opacity of large language models tend to centre around post-hoc explanations of model processes. We explore an alternative approach to fostering critical understanding through intuitive, experience-driven exploration of an LLM’s inner workings. Toward that end, we present “Stepping Into the Black Box” (SIBB): a critical artefact which enables participants to “see like an LLM” and take on an active role in its generative process. Through the box, pairs of participants enact a dynamic interaction with a chatbot presented as a local expert. Participants are invited to intervene in its algorithmic decisions, building responses in tandem with the model. Their interactions and reflections speak to the kinds of criticality developed by experiential learning, highlighting the particular ways that exploring an enclosed system from within affects trust. In this way, SIBB demonstrates the potential of collaborative, experiential research-through-design practices to foster community understanding of emerging technologies.
@inproceedings{10.1145/3800645.3812976,author={Immel, Sarah G and Rajani, Neel and Verweij, Rayo and Li, Jingjie and Nissen, Bettina and Soares, Luis and Taylor, Alex S},title={Stepping Into the Black Box: Opening Up LLMs to Public Exploration Through Discursive Design},year={2026},isbn={9798400725630},publisher={Association for Computing Machinery},address={New York, NY, USA},url={https://doi.org/10.1145/3800645.3812976},doi={10.1145/3800645.3812976},booktitle={Proceedings of the 2026 Designing Interactive Systems Conference},pages={3838–3855},numpages={18},keywords={Research through Design, Critical Artefacts, Large Language Models, Explainable AI, Interpretability, AI Literacy, Experiential Learning},location={Singapore},series={DIS '26},}
2025
AIW@ICML’25
Scalpel vs. Hammer: GRPO Amplifies Existing Capabilities, SFT Replaces Them
Neel Rajani, Aryo Pradipta Gema, Seraphina Goldfarb-Tarrant, and 1 more author
In Actionable Interpretability Workshop at ICML 2025, 2025
Training large language models (LLMs) for reasoning via maths and code datasets has become a major new focus in LLM post-training. Two particularly popular approaches are reinforcement learning (RL) and supervised fine-tuning (SFT), but their training dynamics are poorly understood. We present a comparative analysis of RL and SFT on the same maths problems with the same model and similar hyperparameters. We find that RL yields minor in-domain gains on maths and slight degradation on knowledge-intensive benchmarks like MMLU, while both trends are more pronounced in SFT. We also analyse model parameters across checkpoints, observing that both algorithms modify query and key weights the most. Meanwhile, SFT exhibits greater updates and also affects mid-layer MLPs more, leading us to hypothesise that this may have caused the out-of-domain degradation. We therefore investigate whether freezing parts of the model during training can mitigate the reduced performance on knowledge-intensive benchmarks. However, our results are inconclusive, with benefits on GPQA:Diamond and degradation on other benchmarks. Taken together, our observations provide a preliminary indication for why RL amplifies existing capabilities, while SFT replaces old skills with new ones.
@inproceedings{rajani2025scalpelvshammergrpo,title={Scalpel vs. Hammer: GRPO Amplifies Existing Capabilities, SFT Replaces Them},author={Rajani, Neel and Gema, Aryo Pradipta and Goldfarb-Tarrant, Seraphina and Titov, Ivan},year={2025},url={https://icml.cc/virtual/2025/49627},booktitle={Actionable Interpretability Workshop at ICML 2025},}
In a complete theory there is an element corresponding to each element of reality. A sufficient condition for the reality of a physical quantity is the possibility of predicting it with certainty, without disturbing the system. In quantum mechanics in the case of two physical quantities described by non-commuting operators, the knowledge of one precludes the knowledge of the other. Then either (1) the description of reality given by the wave function in quantum mechanics is not complete or (2) these two quantities cannot have simultaneous reality. Consideration of the problem of making predictions concerning a system on the basis of measurements made on another system that had previously interacted with it leads to the result that if (1) is false then (2) is also false. One is thus led to conclude that the description of reality as given by a wave function is not complete.