Multimodal Integration for Meeting Group Action Segmentation and Recognition

Marc Al-Hames,Alfred Dielmann,Daniel Gatica-Perez,Stephan Reiter,Steve Renals,Gerhard Rigoll,Dong Zhang

Multimodal Integration for Meeting Group Action Segmentation and Recognition

2006

Marc Al-Hames
Alfred Dielmann
Daniel Gatica-Perez
Stephan Reiter
Steve Renals
Gerhard Rigoll
Dong Zhang

We address the problem of segmentation and recognition of sequences of multimodal human interactions in meetings. These interactions can be seen as a rough structure of a meeting, and can be used either as input for a meeting browser or as a first step towards a higher semantic analysis of the meeting. A common lexicon of multimodal group meeting actions, a shared meeting data set, and a common evaluation procedure enable us to compare the different approaches. We compare three different multimodal feature sets and our modelling infrastructures: a higher semantic feature approach, multi-layer HMMs, a multi-stream DBN, as well as a multi-stream mixed-state DBN for disturbed data.

Keywords:

Semantics
Semantic feature
Lexicon
Computer science
Lexico
Natural language processing
Markov model
Segmentation
Machine learning
Artificial intelligence
Hidden Markov model
Shared memory

Correction
Source
Cite
Save
Machine Reading By IdeaReader

References

Citations