A Speech/Music Discriminator of Radio Recordings Based on Dynamic Programming and Bayesian Networks

Authors:
A. Pikrakis;T. Giannakopoulos;S. Theodoridis
Affiliations:
Dept. of Inf. & Telecommun., Univ. of Athens, Athens;-;-
Venue:
IEEE Transactions on Multimedia
Year:
2008

Citing 0
Cited 4

Noise robust features for speech/music discrimination in real-time telecommunication

ICME'09 Proceedings of the 2009 IEEE international conference on Multimedia and Expo
Automatic speech segmentation based on acoustical clustering

SSPR&SPR'10 Proceedings of the 2010 joint IAPR international conference on Structural, syntactic, and statistical pattern recognition
Hierarchical audio content classification system using an optimal feature selection algorithm

Multimedia Tools and Applications
Spectral histogram of oriented gradients (SHOGs) for Tamil language male/female speaker classification

International Journal of Speech Technology

Quantified Score

Hi-index	0.00

Visualization

Abstract

This paper presents a multistage system for speech/music discrimination which is based on a three-step procedure. The first step is a computationally efficient scheme consisting of a region growing technique and operates on a 1-D feature sequence, which is extracted from the raw audio stream. This scheme is used as a preprocessing stage and yields segments with high music and speech precision at the expense of leaving certain parts of the audio recording unclassified. The unclassified parts of the audio stream are then fed as input to a more computationally demanding scheme. The latter treats speech/music discrimination of radio recordings as a probabilistic segmentation task, where the solution is obtained by means of dynamic programming. The proposed scheme seeks the sequence of segments and respective class labels (i.e., speech/music) that maximize the product of posterior class probabilities, given the data that form the segments. To this end, a Bayesian Network combiner is embedded as a posterior probability estimator. At a final stage, an algorithm that performs boundary correction is applied to remove possible errors at the boundaries of the segments (speech or music) that have been previously generated. The proposed system has been tested on radio recordings from various sources. The overall system accuracy is approximately 96%. Performance results are also reported on a musical genre basis and a comparison with existing methods is given.