Natural learning in NLDA networks

Authors:
Ana González;José R. Dorronsoro
Affiliations:
Depto. de Ingeniería Informática and Instituto de Ingeniería del Conocimiento, Universidad Autónoma de Madrid, 28049 Madrid, Spain;Depto. de Ingeniería Informática and Instituto de Ingeniería del Conocimiento, Universidad Autónoma de Madrid, 28049 Madrid, Spain
Venue:
Neural Networks
Year:
2007

Citing 6
Cited 2

Introduction to statistical pattern recognition (2nd ed.)

Introduction to statistical pattern recognition (2nd ed.)
Natural gradient works efficiently in learning

Neural Computation
Complexity issues in natural gradient descent method for training multilayer perceptrons

Neural Computation
Adaptive Method of Realizing Natural Gradient Learning for Multilayer Perceptrons

Neural Computation
On "Natural" Learning and Pruning in Multilayered Perceptrons

Neural Computation
A nonlinear discriminant algorithm for feature extraction and data classification

IEEE Transactions on Neural Networks

Machine Learning Techniques for the Automated Classification of Adhesin-Like Proteins in the Human Protozoan Parasite Trypanosoma cruzi

IEEE/ACM Transactions on Computational Biology and Bioinformatics (TCBB)
Intrinsic plasticity via natural gradient descent with application to drift compensation

Neurocomputing

Quantified Score

Hi-index	0.00

Visualization

Abstract

Non Linear Discriminant Analysis (NLDA) networks combine a standard Multilayer Perceptron (MLP) transfer function with the minimization of a Fisher analysis criterion. In this work we will define natural-like gradients for NLDA network training. Instead of a more principled approach, that would require the definition of an appropriate Riemannian structure on the NLDA weight space, we will follow a simpler procedure, based on the observation that the gradient of the NLDA criterion function J can be written as the expectation @?J(W)=E[Z(X,W)] of a certain random vector Z and defining then I=E[Z(X,W)Z(X,W)^t] as the Fisher information matrix in this case. This definition of I formally coincides with that of the information matrix for the MLP or other square error functions; the NLDA J criterion, however, does not have this structure. Although very simple, the proposed approach shows much faster convergence than that of standard gradient descent, even when its costlier complexity is taken into account. While the faster convergence of natural MLP batch training can be also explained in terms of its relationship with the Gauss-Newton minimization method, this is not the case for NLDA training, as we will see analytically and numerically that the hessian and information matrices are different.