Diagnosing LLM Fragility: An Empirical Study on Reasoning Entropy and Malicious Outputs
Accepted at EMNLP 2026 (main conference). An empirical study of reasoning entropy and malicious outputs in LLMs.
A few of my favourite ideas in ML and interpretability. Hover to replay, click to read the paper.

I’m a final-year PhD candidate in Artificial Intelligence at the University of Milano-Bicocca, supervised by Fabio Mercorio. I’m currently a visiting PhD student at NTU Singapore in Erik Cambria’s group.
I work on understanding and evaluating large language models through their internals:
Before my PhD I spent four years as a freelance software developer (Go, Java, C++, Python).
Email Google Scholar Semantic Scholar GitHub LinkedIn LessWrong X
Diagnosing LLM Fragility accepted at EMNLP 2026 (main conference), Budapest.
Started a research visit at NTU Singapore with Erik Cambria (until Dec 2026).
Two papers at EMNLP 2025: role vectors (Findings) and SFAL (Industry Track).
ITALIC presented at NAACL 2025.
Accepted at EMNLP 2026 (main conference). An empirical study of reasoning entropy and malicious outputs in LLMs.
Role vectors: activation directions associated with roles that can steer LLM behaviour.
Authors in alphabetical order; I contributed to the technical work.
Scores for evaluating whether automatically generated explanations of sparse autoencoder features align with what the features actually do.
A benchmark for evaluating LLMs on Italian language and culture.