Semi-supervised URL Segmentation with Recurrent Neural Networks Pre-trained on Knowledge Graph Entities

Hao Zhang,Jae Ro,Richard Sproat

Semi-supervised URL Segmentation with Recurrent Neural Networks Pre-trained on Knowledge Graph Entities

2020

Hao Zhang
Jae Ro
Richard Sproat

Breaking domain names such as openresearch into component words open and research is important for applications like Text-to-Speech synthesis and web search. We link this problem to the classic problem of Chinese word segmentation and show the effectiveness of a tagging model based on Recurrent Neural Networks (RNNs) using characters as input. To compensate for the lack of training data, we propose a pre-training method on concatenated entity names in a large knowledge database. Pre-training improves the model by 33% and brings the sequence accuracy to 85%.

Keywords:

chinese word
Knowledge base
Recurrent neural network
Artificial intelligence
Natural language processing
Computer science
knowledge graph
Segmentation
Concatenation
Training set

Correction
Source
Cite
Save
Machine Reading By IdeaReader

References

Citations