publications

* denotes equal contribution

preprints

  1. PoLoRA: A Preconditioned Orthogonalized LoRA Optimizer
    Nikhil Ghosh, Tetiana Parshakova, and Robert M. Gower
    arXiv preprint, 2026
  2. There Will Be a Scientific Theory of Deep Learning
    Jamie Simon, Daniel Kunin, Alexander Atanasov, Enric Boix-Adserà, Blake Bordelon, Jeremy Cohen, Nikhil Ghosh, Florentin Guth, Arthur Jacot, Mason Kamb, Dhruva Karkada, Eric J. Michaud, Berkan Ottlik, and Joseph Turnbull
    arXiv preprint, 2026
  3. Does a Global Perspective Help Prune Sparse MoEs Elegantly?
    Zeliang Zhang, Nikhil Ghosh, Jiani Liu, Bin Yu, and Xiaodong Liu
    arXiv preprint, 2026
  4. GSM-Agent: Understanding Agentic Reasoning Using Controllable Environments
    Hanlin Zhu, Tianyu Guo, Song Mei, Stuart Russell, Nikhil Ghosh, Alberto Bietti, and Jiantao Jiao
    arXiv preprint, 2025

conference & journal articles

2026

  1. ICLR
    Understanding the Mechanisms of Fast Hyperparameter Transfer
    Nikhil Ghosh, Denny Wu, and Alberto Bietti
    In International Conference on Learning Representations (ICLR), 2026
  2. ICLR
    PLoP: Precise LoRA Placement for Efficient Finetuning of Large Models
    Soufiane Hayou, Nikhil Ghosh, and Bin Yu
    In International Conference on Learning Representations (ICLR), 2026

2025

  1. JMLR
    The Effect of SGD Batch Size on Autoencoder Learning: Sparsity, Sharpness, and Feature Learning
    Nikhil Ghosh, Spencer Frei, Wooseok Ha, and Bin Yu
    Journal of Machine Learning Research, 2025

2024

  1. ICML
    LoRA+: Efficient Low Rank Adaptation of Large Models
    Soufiane Hayou, Nikhil Ghosh, and Bin Yu
    In International Conference on Machine Learning (ICML), 2024
  2. NeurIPS
    The Impact of Initialization on LoRA Finetuning Dynamics
    Soufiane Hayou, Nikhil Ghosh, and Bin Yu
    2024
  3. ICLR
    More is Better in Modern Machine Learning: when Infinite Overparameterization is Optimal and Overfitting is Obligatory
    James B. Simon, Dhruva Karkada, Nikhil Ghosh, and Mikhail Belkin
    In International Conference on Learning Representations (ICLR), 2024

2023

  1. NeurIPS
    Alternating Updates for Efficient Transformers
    Cenk Baykal, Dylan Cutler, Nishanth Dikkala, Nikhil Ghosh, Rina Panigrahy, and Xin Wang
    In Advances in Neural Information Processing Systems (NeurIPS), 2023
  2. EMNLP
    On the Benefits of Learning to Route in Mixture-of-Experts Models
    Nishanth Dikkala, Nikhil Ghosh, Raghu Meka, Rina Panigrahy, Nikhil Vyas, and Xin Wang
    In Conference on Empirical Methods in Natural Language Processing (EMNLP), 2023
  3. SIMODS
    A Universal Trade-off Between the Model Size, Test Loss, and Training Loss of Linear Predictors
    Nikhil Ghosh and Mikhail Belkin
    SIAM Journal on Mathematics of Data Science (SIMODS), 2023
  4. ICLR
    Deconstructing Distributions: A Pointwise Framework of Learning
    Gal Kaplun*, Nikhil Ghosh*, Saurabh Garg, Boaz Barak, and Preetum Nakkiran
    In International Conference on Learning Representations (ICLR), 2023

2022

  1. ICLR
    The Three Stages of Learning Dynamics in High-Dimensional Kernel Methods
    Nikhil Ghosh, Song Mei, and Bin Yu
    In International Conference on Learning Representations (ICLR), 2022

2019

  1. NeurIPS
    Landmark Ordinal Embedding
    Nikhil Ghosh, Yuxin Chen, and Yisong Yue
    In Advances in Neural Information Processing Systems (NeurIPS), 2019