Publications

* denotes equal contribution

2026

DiffuMamba: High-Throughput Diffusion LMs with Mamba Backbone
Vaibhav Singh*, Oleksiy Ostapenko, Pierre-André Noël, Eugene Belilovsky, Torsten Scholak
International Conference on Machine Learning (ICML) (2026)
DiffuMamba: High-Throughput Diffusion LMs with Mamba Backbone
Dual-Phase Continual Learning: Supervised Adaptation Meets Unsupervised Retention
Vaibhav Singh*, Rahaf Aljundi, Eugene Belilovsky
Transactions on Machine Learning Research (TMLR) (2026)
Dual-Phase Continual Learning: Supervised Adaptation Meets Unsupervised Retention

2025

Beyond Cosine Decay: On the Effectiveness of Infinite Learning Rate Schedule for Continual Pre-training
Vaibhav Singh*, Paul Janson*, Paria Mehrbod, Adam Ibrahim, Irina Rish, Eugene Belilovsky, Benjamin Therien
Conference on Lifelong Learning Agents (CoLLAs) (2025) Oral presentation
Beyond Cosine Decay: On the Effectiveness of Infinite Learning Rate Schedule for Continual Pre-training
Model Parallelism With Subnetwork Data Parallelism
Vaibhav Singh*, Zafir Khalid, Edouard Oyallon, Eugene Belilovsky
ICML 2025 ESFOMO Workshop (2025)
Model Parallelism With Subnetwork Data Parallelism

2024

Wake-Sleep Energy-Based Models for Continual Learning
Vaibhav Singh*, Anna Choromanska, Shuang Li, Yilun Du
CLVISION Workshop, CVPR 2024 (2024)
Wake-Sleep Energy-Based Models for Continual Learning

2022

On Spectral and Temporal Feature Encoding Behaviour in Stacked Architectures
Vaibhav Singh*, Vinayak Abrol, Karan Nathwani
NeurIPS 2022 ENLSP Workshop (2022)
On Spectral and Temporal Feature Encoding Behaviour in Stacked Architectures