Combining Mask Estimates for Single Channel Audio Source Separation using Deep Neural Networks

Emad M. Grais,Gerard Roma,Andrew J. R. Simpson,Mark D. Plumbley

Combining Mask Estimates for Single Channel Audio Source Separation using Deep Neural Networks

2016

Emad M. Grais
Gerard Roma
Andrew J. R. Simpson
Mark D. Plumbley

Deep neural networks (DNNs) are usually used for single channel source separation to predict either soft or binary time frequency masks. The masks are used to separate the sources from the mixed signal. Binary masks produce separated sources with more distortion and less interference than soft masks. In this paper, we propose to use another DNN to combine the estimates of binary and soft masks to achieve the advantages and avoid the disadvantages of using each mask individually. We aim to achieve separated sources with low distortion and low interference between each other. Our experimental results show that combining the estimates of binary and soft masks using DNN achieves lower distortion than using each estimate individually and achieves as low interference as the binary mask.

Keywords:

Deep learning
Speech recognition
Binary number
Distortion
Interference (wave propagation)
Pattern recognition
Artificial neural network
Time–frequency analysis
Computer science
Communication channel
Source separation
Artificial intelligence
Mixed-signal integrated circuit

Correction
Source
Cite
Save
Machine Reading By IdeaReader

References

Citations