Feature learning in Bayesian deep linear networks with multiple outputs and convolutional layers
Abstract: Deep linear networks have been extensively studied in recent literature. However, there is still limited understanding regarding finite-width architectures with multiple outputs and convolutional layers, which are fundamental aspects in deep learning. In this seminar, we start by providing a concise overview of notable findings and common challenges in this area, within a Bayesian framework. We then present few new insights, including:
(i) An exact, straightforward, and non-asymptotic integral representation for the joint prior distribution across outputs.
(ii) An analytical formula for the posterior distribution, specifically addressing the mean square loss function.
(iii) A quantitative analysis of feature learning within the infinite-width regime, using large deviation language.
From a physical perspective, architectures featuring multiple outputs or convolutional layers can be seen as different manifestations of shape kernel renormalization. Our work establishes a reliable connection, translating physical insights and terminology into rigorous mathematical statements.
The talk is based on a joint working paper with P. Rotondo (Dip. di Scienze Matematiche, Fisiche e Informatiche, Univ. Parma), M. Pastore (École normale supérieure, CNRS), and M. Gherardi (Dip. Fisica, Univ. Milano).
https://us02web.zoom.us/j/83344185446?pwd=RHRGai91RkZQTjg0eEVRMWQ5WXFjZz09
