Grammatical inference as a principal component analysis problem

Authors:
Raphaël Bailly;François Denis;Liva Ralaivola
Affiliations:
Aix-Marseille Université, Marseille, France;Aix-Marseille Université, Marseille, France;Aix-Marseille Université, Marseille, France
Venue:
ICML '09 Proceedings of the 26th Annual International Conference on Machine Learning
Year:
2009

Citing 6
Cited 3

Learning functions represented as multiplicity automata

Journal of the ACM (JACM)
Probabilistic DFA Inference using Kullback-Leibler Divergence and Minimality

ICML '00 Proceedings of the Seventeenth International Conference on Machine Learning
Learning Stochastic Regular Grammars by Means of a State Merging Method

ICGI '94 Proceedings of the Second International Colloquium on Grammatical Inference and Applications
On Rational Stochastic Languages

Fundamenta Informaticae
Languages as hyperplanes: grammatical inference with string kernels

ECML'06 Proceedings of the 17th European conference on Machine Learning
Learning rational stochastic languages

COLT'06 Proceedings of the 19th annual conference on Learning Theory

A spectral approach for probabilistic grammatical inference on trees

ALT'10 Proceedings of the 21st international conference on Algorithmic learning theory
Absolute convergence of rational series is semi-decidable

Information and Computation
A spectral learning algorithm for finite state transducers

ECML PKDD'11 Proceedings of the 2011 European conference on Machine learning and knowledge discovery in databases - Volume Part I

Quantified Score

Hi-index	0.01

Visualization

Abstract

One of the main problems in probabilistic grammatical inference consists in inferring a stochastic language, i.e. a probability distribution, in some class of probabilistic models, from a sample of strings independently drawn according to a fixed unknown target distribution p. Here, we consider the class of rational stochastic languages composed of stochastic languages that can be computed by multiplicity automata, which can be viewed as a generalization of probabilistic automata. Rational stochastic languages p have a useful algebraic characterization: all the mappings up: v → p(uv) lie in a finite dimensional vector subspace Vp* of the vector space ℝ 〈〈Σ〉〉 composed of all real-valued functions defined over Σ*. Hence, a first step in the grammatical inference process can consist in identifying the subspace Vp*. In this paper, we study the possibility of using Principal Component Analysis to achieve this task. We provide an inference algorithm which computes an estimate of this space and then build a multiplicity automaton which computes an estimate of the target distribution. We prove some theoretical properties of this algorithm and we provide results from numerical simulations that confirm the relevance of our approach.