Researchers at Tsinghua University have posted their preprint on arXiv addressing one of AI interpretability's most persistent scaling problems: the inability to label millions of sparse autoencoder ...