Daniele Potertì

A few of my favourite ideas in ML and interpretability. Hover to replay, click to read the paper.

GrokkingTrain accuracy saturates early; test accuracy jumps much later, as the embeddings snap into a circle.Power et al., 2022 ↗
Activation steeringAdding a direction v to the residual stream shifts what the model writes.Turner et al., 2023 ↗
Linear probingA linear classifier reads a concept, like truth, off the activations.Alain & Bengio, 2016 ↗
SuperpositionSparse features share dimensions: five features in two.Elhage et al., 2022 ↗
Daniele in front of Marina Bay Sands, Singapore

Daniele Potertì

I’m a final-year PhD candidate in Artificial Intelligence at the University of Milano-Bicocca, supervised by Fabio Mercorio. I’m currently a visiting PhD student at NTU Singapore in Erik Cambria’s group.

I work on understanding and evaluating large language models through their internals:

  • Steering and role vectors: activation directions that shape model behaviour (EMNLP Findings 2025)
  • Evaluating interpretability methods: do automatic explanations of sparse autoencoder features match what the features actually do? (EMNLP 2025)
  • Deception and safety: a theory-driven approach to studying deceptive behaviour in LLMs (in progress), and LLM fragility and malicious outputs (EMNLP 2026)
  • LLM evaluation and benchmarks, especially for Italian (ITALIC, NAACL 2025)

Before my PhD I spent four years as a freelance software developer (Go, Java, C++, Python).

News

September 2026

Diagnosing LLM Fragility accepted at EMNLP 2026 (main conference), Budapest.

May 2026

Started a research visit at NTU Singapore with Erik Cambria (until Dec 2026).

November 2025

Two papers at EMNLP 2025: role vectors (Findings) and SFAL (Industry Track).

April 2025

ITALIC presented at NAACL 2025.

Selected research

October 2026

Diagnosing LLM Fragility: An Empirical Study on Reasoning Entropy and Malicious Outputs

Accepted at EMNLP 2026 (main conference). An empirical study of reasoning entropy and malicious outputs in LLMs.

November 2025

Can Role Vectors Affect LLM Behaviour?

Role vectors: activation directions associated with roles that can steer LLM behaviour.