Diagnosing LLM Fragility: An Empirical Study on Reasoning Entropy and Malicious Outputs
Accepted at EMNLP 2026 (main conference). An empirical study of reasoning entropy and malicious outputs in LLMs.
Everything I've worked on, newest first. Also on Google Scholar.
Accepted at EMNLP 2026 (main conference). An empirical study of reasoning entropy and malicious outputs in LLMs.
Challenging the Abilities of Large Language Models in Italian: a Community Initiative The CALAMITA community initiative for evaluating LLMs in Italian. Italian Journal of Computational Linguistics (IJCoL)
Role vectors: activation directions associated with roles that can steer LLM behaviour.
Authors in alphabetical order; I contributed to the technical work.
Scores for evaluating whether automatically generated explanations of sparse autoencoder features align with what the features actually do.
A Benchmark to Evaluate LLMs' Proficiency on Italian Student Competencies Evaluating LLMs on the competencies measured in Italian student assessments. ECML PKDD 2025
Authors in alphabetical order; I contributed to the technical work.
A benchmark for evaluating LLMs on Italian language and culture.
Designing Role Vectors to Improve LLM Inference Behaviour Preprint on designing role vectors to improve LLM inference behaviour. arXiv preprint
BEEP - BEst DrivEr’s License Performer: A CALAMITA Challenge A CALAMITA challenge evaluating LLMs on the Italian driver’s license exam. Proceedings of CLiC-it 2024
Authors in alphabetical order; I contributed to the technical work.
Augmenting XAI with LLMs: A Case Study in Banking Marketing Recommendation Using LLMs to make explanations of a banking recommender system more usable. World Conference on Explainable Artificial Intelligence (xAI 2024)
Authors in alphabetical order; I contributed to the technical work.
Disce aut Deficere: Evaluating LLMs Proficiency on the INVALSI Italian Benchmark Evaluating LLMs on the INVALSI Italian benchmark. arXiv preprint
Authors in alphabetical order; I contributed to the technical work.
Marrying LLMs with Domain Expert Validation for Causal Graph Generation Combining LLMs and domain experts to build causal graphs. AIABI Workshop @ AI*IA 2023
Authors in alphabetical order; I contributed to the technical work.
Bridging Causal Theory and LLMs: A User Centric Approach to Causal Graph Generation Doctoral consortium paper on user-centric causal graph generation with LLMs. Doctoral Consortium @ AI*IA 2023