Shao Sheng Cao

dblp:175/4823 · DBLP profile ↗
← Back
1ranked-venue papers
0as first author
0since 2021 · last 2016
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 1Databases, data management, data science and information retrieval · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
1 paper
Information retrieval · 91% Knowledge graphs · 9%

Topics — the 4 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Information retrieval › document processing › document analysis
document representation
0.212016
Semantic Documents Relatedness using Concept Graph Representation · WSDM 2016
Information retrieval › similarity measure
document similarity
0.212016
Semantic Documents Relatedness using Concept Graph Representation · WSDM 2016
Information retrieval › text analysis
semantic relatedness
0.212016
Semantic Documents Relatedness using Concept Graph Representation · WSDM 2016
Knowledge graphs
knowledge base linking
0.112016
Semantic Documents Relatedness using Concept Graph Representation · WSDM 2016

Methods — techniques the papers use, named apart from their topics

neural network embedding · 0.2graph similarity · 0.2closeness centrality · 0.2
YearPublicationVenuePosition
2016 Semantic Documents Relatedness using Concept Graph Representation
abstract
We deal with the problem of document representation for the task of measuring semantic relatedness between documents. A document is represented as a compact concept graph where nodes represent concepts extracted from the document through references to entities in a knowledge base such as DBpedia. Edges represent the semantic and structural relationships among the concepts. Several methods are presented to measure the strength of those relationships. Concepts are weighted through the concept graph using closeness centrality measure which reflects their relevance to the aspects of the document. A novel similarity measure between two concept graphs is presented. The similarity measure first represents concepts as continuous vectors by means of neural networks. Second, the continuous vectors are used to accumulate pairwise similarity between pairs of concepts while considering their assigned weights. We evaluate our method on a standard benchmark for document similarity. Our method outperforms state-of-the-art methods including ESA (Explicit Semantic Annotation) while our concept graphs are much smaller than the concept vectors generated by ESA. Moreover, we show that by combining our concept graph with ESA, we obtain an even further improvement.
Yuan Ni, Qiongkai Xu, Yosi Mass, Dafna Sheinwald, Huijia Zhu, Shao Sheng Cao
WSDM7