Daniele Potertì
November 2025

SFAL: Semantic-Functional Alignment Scores for Distributional Evaluation of Auto-Interpretability in Sparse Autoencoders

Authors in alphabetical order; I contributed to the technical work.

Proceedings of EMNLP 2025: Industry Track

Paper ↗

Summary

Scores for evaluating whether automatically generated explanations of sparse autoencoder features align with what the features actually do.