Conference proceeding
- Saddle-to-saddle dynamics explains a simplicity bias across neural network architectures
Yedi Zhang, Andrew Saxe, Peter E. Latham
International Conference on Learning Representations (ICLR), 2026 - Training dynamics of in-context learning in linear attention
Yedi Zhang, Aaditya K. Singh, Peter E. Latham*, Andrew Saxe*
International Conference on Machine Learning (ICML), Spotlight, 2025 - Understanding unimodal bias in multimodal deep linear networks
Yedi Zhang, Peter E. Latham, Andrew Saxe
International Conference on Machine Learning (ICML), 2024
Journal article
- When are bias-free ReLU networks effectively linear networks?
Yedi Zhang, Andrew Saxe, Peter E. Latham
Transactions on Machine Learning Research (TMLR), 2025
Preprint
- To use or not to use muon: how simplicity bias in optimizers matters
Sara Dragutinović, Yedi Zhang, Rajesh Ranganath
2026
Non-archival conference
- Transitivity can redeem the reversal curse in attention models
Yedi Zhang, Andrew Lampinen*, James McClelland*
Annual Conference on Cognitive Computational Neuroscience (CCN), 2026 - Optimal learning rate scaling depends on data in deep scalar linear networks
Yedi Zhang, Peter E. Latham, Leena Chennuru Vankadara, Andrew Saxe
ICLR Workshop on Scientific Methods for Understanding Deep Learning, 2026
