ECE Seminar Lecture Series
Training Task Diversity Shapes In-Context Learning via Low-Dimensional Subspaces
Qing Qu, Assistant Professor in EECS at the University of Michigan.
Friday, October 16, 2026
Noon1 p.m.
CSB 601
Abstract: The transformer's emergent ability to perform in-context learning (ICL) has sparked a wide range of studies designed to understand its underlying mechanism. Existing works often study how training task diversity, defined either as the number of ICL training task vectors or as the number of function classes from which the task vectors are drawn, shapes both the generalization capabilities and the learning dynamics of ICL. While both definitions have uncovered many interesting phenomena, many observations under the latter definition remain theoretically unexplained. This talk presents a minimal analytical model under which these phenomena provably emerge from the properties of the pre-training data. By modeling the pre-training task vectors as a mixture of low-rank Gaussians, we show that pre-training task diversity, defined by the number of non-overlapping columns between the subspaces that parameterize the covariance matrices, improves both the generalization and the optimization trajectory of ICL with linear attention. In particular, we show that our model can explain (i) why ICL can achieve out-of-distribution generalization, and (ii) why pre-training with multiple tasks can shorten the ICL training plateau. We conclude by showing how our results empirically extend to nonlinear transformers and nonlinear function classes. Overall, our work presents a mathematically tractable framework to unify existing observations.
Speaker Bio: Qing Qu is an Assistant Professor in EECS at the University of Michigan. He works at the intersection of the foundations of machine learning, numerical optimization, and signal/image processing, with a current focus on the theory of deep generative models and representation learning. Prior to joining Michigan in 2021, he was a Moore–Sloan Data Science Fellow at the Center for Data Science, New York University (2018–2020). He received his Ph.D. in Electrical Engineering from Columbia University in October 2018 and his B.Eng. in Electrical and Computer Engineering from Tsinghua University in July 2011. His work has been recognized with multiple honors, including the Best Student Paper Award at SPARS 2015, a Microsoft PhD Fellowship in Machine Learning (2016), the Best Paper Award at the NeurIPS Diffusion Models Workshop (2023), NSF CAREER Award (2022), Amazon Research Award (AWS AI, 2023), UM CHS Junior Faculty Award (2025), Google Research Scholar Award (2025), and the 1938E Award in Michigan Engineering (2026). He has led and delivered multiple tutorials at ICASSP, CPAL, CVPR, ICCV, and ICML. He was one of the founding organizers and Program Chair for the new Conference on Parsimony & Learning (CPAL), regularly serves as an Area Chair for NeurIPS, ICML, and ICLR, senior area chair for ICASSP’26, and is an Action Editor for TMLR and an Associate Editor for the IEEE TSP Journal.