Adaptive Importance Sampling to Accelerate Training of a Neural Probabilistic Language Model

Authors:
Y. Bengio;J. -S. Senecal
Affiliations:
Univ. de Montreal, Montreal;-
Venue:
IEEE Transactions on Neural Networks
Year:
2008

Citing 0
Cited 2

Importance Sampling for Objective Function Estimations in Neural Detector Training Driven by Genetic Algorithms

Neural Processing Letters
Efficient subsampling for training complex language models

EMNLP '11 Proceedings of the Conference on Empirical Methods in Natural Language Processing

Quantified Score

Hi-index	0.00

Visualization

Abstract

Previous work on statistical language modeling has shown that it is possible to train a feedforward neural network to approximate probabilities over sequences of words, resulting in significant error reduction when compared to standard baseline models based on n-grams. However, training the neural network model with the maximum-likelihood criterion requires computations proportional to the number of words in the vocabulary. In this paper, we introduce adaptive importance sampling as a way to accelerate training of the model. The idea is to use an adaptive n-gram model to track the conditional distributions produced by the neural network. We show that a very significant speedup can be obtained on standard problems.