IIRS: A Novel Framework of Identifying Commodity Entities on E-commerce Big Data

2015 
Identification of the same commodity entities is a major challenge in the heterogeneous multi-source e-commerce of big data. This paper introduces a framework based on Map-Reduce, called IIRS, which is made up of data index, data integration, entity recognition and data sorting. IIRS aims to form the unified model and high efficient commodity information with building an index model based on commodity’s attribute/value and constructing a global model map to record commodity’s attribute and value, identify the commodity entities in different e-commerce with measuring the similarity of the commodity’s identity, and then output the same identity commodity sets and their associated properties organized in the inverted index list. Through an extensive experimental study on real e-commerce dataset on Hadoop, IIRS significantly demonstrates its feasibility, accuracy, and high efficiency.
    • Correction
    • Source
    • Cite
    • Save
    • Machine Reading By IdeaReader
    8
    References
    1
    Citations
    NaN
    KQI
    []