November 2025
SFAL: Semantic-Functional Alignment Scores for Distributional Evaluation of Auto-Interpretability in Sparse Autoencoders
Authors in alphabetical order; I contributed to the technical work.
Proceedings of EMNLP 2025: Industry Track
Summary
Scores for evaluating whether automatically generated explanations of sparse autoencoder features align with what the features actually do.