an:07192417
Zbl 1441.90127
Palagi, Laura; Seccia, Ruggiero
Block layer decomposition schemes for training deep neural networks
EN
J. Glob. Optim. 77, No. 1, 97-124 (2020).
00448591
2020
j
90C26 90C06 68T05
deep feedforward neural networks; block coordinate decomposition; online optimization; large scale optimization
Summary: Deep feedforward neural networks' (DFNNs) weight estimation relies on the solution of a very large nonconvex optimization problem that may have many local (no global) minimizers, saddle points and large plateaus. Furthermore, the time needed to find good solutions of the training problem heavily depends on both the number of samples and the number of weights (variables). In this work, we show how block coordinate descent (BCD) methods can be fruitful applied to DFNN weight optimization problem and embedded in online frameworks possibly avoiding bad stationary points. We first describe a batch BCD method able to effectively tackle difficulties due to the network's depth; then we further extend the algorithm proposing an online BCD scheme able to scale with respect to both the number of variables and the number of samples. We perform extensive numerical results on standard datasets using various deep networks. We show that the application of BCD methods to the training problem of DFNNs improves over standard batch/online algorithms in the training phase guaranteeing good generalization performance as well.