Sitemap
A list of all the posts and pages found on the site. For you robots out there is an XML version available for digesting as well.
Pages
Posts
portfolio
publications
Bridging Causal Theory and LLMs: A User Centric Approach to Causal Graph Generation
Doctoral Consortium @ AI*IA 2023 · 2023
Doctoral consortium paper on user-centric causal graph generation with LLMs.
Marrying LLMs with Domain Expert Validation for Causal Graph Generation
AIABI Workshop @ AI*IA 2023 · 2023
Combining LLMs and domain experts to build causal graphs.
Disce aut Deficere: Evaluating LLMs Proficiency on the INVALSI Italian Benchmark
arXiv preprint · 2024
Evaluating LLMs on the INVALSI Italian benchmark.
Augmenting XAI with LLMs: A Case Study in Banking Marketing Recommendation
World Conference on Explainable Artificial Intelligence (xAI 2024) · 2024
Using LLMs to make explanations of a banking recommender system more usable.
BEEP - BEst DrivEr’s License Performer: A CALAMITA Challenge
Proceedings of CLiC-it 2024 · 2024
A CALAMITA challenge evaluating LLMs on the Italian driver’s license exam.
Designing Role Vectors to Improve LLM Inference Behaviour
arXiv preprint · 2025
Preprint on designing role vectors to improve LLM inference behaviour.
ITALIC: An Italian Culture-Aware Natural Language Benchmark
Proceedings of NAACL 2025 (Long Papers) · 2025
A benchmark for evaluating LLMs on Italian language and culture.
A Benchmark to Evaluate LLMs’ Proficiency on Italian Student Competencies
ECML PKDD 2025 · 2025
Evaluating LLMs on the competencies measured in Italian student assessments.
SFAL: Semantic-Functional Alignment Scores for Distributional Evaluation of Auto-Interpretability in Sparse Autoencoders
Proceedings of EMNLP 2025: Industry Track · 2025
Scores for evaluating whether automatically generated explanations of sparse autoencoder features align with what the features actually do.
Can Role Vectors Affect LLM Behaviour?
Findings of the Association for Computational Linguistics: EMNLP 2025 · 2025
Role vectors: activation directions associated with roles that can steer LLM behaviour.
Challenging the Abilities of Large Language Models in Italian: a Community Initiative
Italian Journal of Computational Linguistics (IJCoL) · 2026
The CALAMITA community initiative for evaluating LLMs in Italian.
Diagnosing LLM Fragility: An Empirical Study on Reasoning Entropy and Malicious Outputs
Proceedings of EMNLP 2026 (Main Conference) · 2026
Accepted at EMNLP 2026 (main conference). An empirical study of reasoning entropy and malicious outputs in LLMs.
