arXiv:2606. 18538v1 Announce Type: cross Abstract: One of the major difficulties in the mechanistic interpretability of neural networks is the occurrence of polysemanticity, which suggests that each neuron is typically responsible for multiple different tasks, impeding a clean interpretation of their function.
Paper
Effects of sparsity and superposition on loss in simple autoencoders
Unreadunread