Language model adaptation for video lectures transcription

Adria A. Martinez-Villaronga,Miguel A. del Agua,Jesús Andrés-Ferrer,Alfons Juan

Language model adaptation for video lectures transcription

2013

Adria A. Martinez-Villaronga
Miguel A. del Agua
Jesús Andrés-Ferrer
Alfons Juan

Videolectures are currently being digitised all over the world for its enormous value as reference resource. Many of these lectures are accompanied with slides. The slides offer a great opportunity for improving ASR systems performance. We propose a simple yet powerful extension to the linear interpolation of language models for adapting language models with slide information. Two types of slides are considered, correct slides, and slides automatic extracted from the videos with OCR. Furthermore, we compare both time aligned and unaligned slides. Results report an improvement of up to 3.8 % absolute WER points when using correct slides. Surprisingly, when using automatic slides obtained with poor OCR quality, the ASR system still improves up to 2.2 absolute WER points.

Keywords:

Speech recognition
Optical character recognition
Linear interpolation
Computer science
Language model
Interpolation

Correction
Source
Cite
Save
Machine Reading By IdeaReader

References

Citations