Yueji Yang

dblp:177/8972 · DBLP profile ↗
← Back
4ranked-venue papers
3as first author
2since 2021 · last 2021
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 4 · 3 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
4 papers
Information retrieval · 46% Knowledge graphs · 46% Graph data management · 9%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Parallel and multicore computing · 77% GPUs and heterogeneous computing · 23%

Topics — the 12 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Knowledge graphs
knowledge graph mining
0.512021
Context-aware Outstanding Fact Mining from Knowledge Graphs · KDD 2021
Knowledge graphs
knowledge graph querying
0.512021
NewsLink: Empowering Intuitive News Search with Knowledge Graphs · ICDE 2021
Knowledge graphs › relation learning
relation discovery
0.512021
Context-aware Outstanding Fact Mining from Knowledge Graphs · KDD 2021
Information retrieval
search engines
0.512021
NewsLink: Empowering Intuitive News Search with Knowledge Graphs · ICDE 2021
Information retrieval
keyword search
0.412019
An Efficient Parallel Keyword Search Engine on Knowledge Graphs · ICDE 2019
Graph data management › keyword search on graphs
keyword search over knowledge graphs
0.412019
An Efficient Parallel Keyword Search Engine on Knowledge Graphs · ICDE 2019
Parallel and multicore computing
parallel graph algorithms
0.412019
An Efficient Parallel Keyword Search Engine on Knowledge Graphs · ICDE 2019
Information retrieval › indexing
inverted index
0.312018
A Generic Inverted Index Framework for Similarity Search on the GPU · ICDE 2018
Information retrieval › hashing › hashing for nearest neighbor search
locality-sensitive hashing
0.312018
A Generic Inverted Index Framework for Similarity Search on the GPU · ICDE 2018
Information retrieval
similarity search
0.312018
A Generic Inverted Index Framework for Similarity Search on the GPU · ICDE 2018
Information retrieval
retrieval models
0.112021
NewsLink: Empowering Intuitive News Search with Knowledge Graphs · ICDE 2021
GPUs and heterogeneous computing
GPU computing
0.112019
An Efficient Parallel Keyword Search Engine on Knowledge Graphs · ICDE 2019

Methods — techniques the papers use, named apart from their topics

group steiner tree · 0.8central graph · 0.8top-k search · 0.5subgraph embedding · 0.5pruning · 0.5knowledge graph embedding · 0.5product quantization · 0.3locality-sensitive hashing · 0.3
YearPublicationVenuePosition
2021 NewsLink: Empowering Intuitive News Search with Knowledge Graphs
abstract
News search tools help end users to identify relevant news stories. However, existing search approaches often carry out in a "black-box" process. There is little intuition that helps users understand how the results are related to the query. In this paper, we propose a novel news search framework, called NEWSLINK, to empower intuitive news search by using relationship paths discovered from open Knowledge Graphs (KGs). Specifically, NEWSLINK embeds both a query and news documents to subgraphs, called subgraph embeddings, in the KG. Their embeddings' overlap induces relationship paths between the involving entities. Two major advantages are obtained by incorporating subgraph embeddings into search. First, they enrich the search context, leading to robust results. Second, the relationship paths linking entities inter and intra news documents can help users better understand and digest the results for the given query. Through both human and automatic evaluations, we verify that NEWSLINK can help users understand the result-to-query relatedness, while its search quality is robust and outperforms many established search approaches, including Apache Lucene and a KG-powered query expansion approach, as well as popular deep learning models, Sentence-BERT (SBERT) and DOC2VEC.
Yueji Yang, Yuchen Li 0001, Anthony K. H. Tung
ICDE1
2021 Context-aware Outstanding Fact Mining from Knowledge Graphs
abstract
An Outstanding Fact (OF) is an attribute that makes a target entity stand out from its peers. The mining of OFs has important applications, especially in Computational Journalism, such as news promotion, fact-checking, and news story finding. However, existing approaches to OF mining: (i) disregard the context in which the target entity appears, hence may report facts irrelevant to that context; and (ii) require relational data, which are often unavailable or incomplete in many application domains. In this paper, we introduce the novel problem of mining Context-aware Outstanding Facts (COFs) for a target entity under a given context specified by a context entity. We propose FMiner, a context-aware mining framework that leverages knowledge graphs (KGs) for COF mining. FMiner generates COFs in two steps. First, it discovers top-k relevant relationships between the target and the context entity from a KG. We propose novel optimizations and pruning techniques to expedite this operation, as this process is very expensive on large KGs due to its exponential complexity. Second, for each derived relationship, we find the attributes of the target entity that distinguish it from peer entities that have the same relationship with the context entity, yielding the top-l COFs. As such, the mining process is modeled as a top-(k,l) search problem. Context-awareness is ensured by relying on the relevant relationships with the context entity to derive peer entities for COF extraction. Consequently, FMiner can effectively navigate the search to obtain context-aware OFs by incorporating a context entity. We conduct extensive experiments, including a user study, to validate the efficiency and the effectiveness of FMiner.
Yueji Yang, Yuchen Li 0001, Panagiotis Karras, Anthony K. H. Tung
KDD1
2019 An Efficient Parallel Keyword Search Engine on Knowledge Graphs
abstract
Keyword search has recently become popular as a way to query relational databases, and even graphs, since it allows users to issue queries without learning a complex query language and data schema. Evaluating a keyword query is usually significantly more expensive than evaluating an equivalent selection query, since the query specification is less complete, and many alternative answers have to be considered by the system, requiring considerable effort to generate and compare. Current interest in big data and AI are putting even more demands on the efficiency of keyword search. In particular, searching of knowledge graphs is gaining popularity. As knowledge graphs often comprise many millions of nodes and edges, performing real-time search on graphs of this size is an open challenge. In this paper, we attempt to address this need by leveraging advances in hardware technologies, e.g. multi-core CPUs and GPUs. Specifically, we implement a parallel keyword search engine for Knowledge Bases (KB). To be able to do so, and to exploit parallelism, we devise a new approach to keyword search, based on a concept we introduce called Central Graph. Unlike the Group Steiner Tree (GST) model, widely used for keyword search, our approach can naturally work in parallel and still return compact answer graphs with rich information. Our approach can work in either multi-core CPUs or a single GPU. In particular, our GPU implementation is two to three orders of magnitudes faster than state-of-the-art keyword search method. We conduct extensive experiments to show that our approach is both efficient and effective.
Yueji Yang, Divyakant Agrawal, H. V. Jagadish, Anthony K. H. Tung, Shuang Wu 0002
ICDE1
2018 A Generic Inverted Index Framework for Similarity Search on the GPU
abstract
We propose a novel generic inverted index framework on the GPU (called GENIE), aiming to reduce the programming complexity of the GPU for parallel similarity search of different data types. Not every data type and similarity measure are supported by GENIE, but many popular ones are. We present the system design of GENIE, and demonstrate similarity search with GENIE on several data types along with a theoretical analysis of search results. A new concept of locality sensitive hashing (LSH) named tau-ANN search, and a novel data structure c-PQ on the GPU are also proposed for achieving this purpose. Extensive experiments on different real-life datasets demonstrate the efficiency and effectiveness of our framework. The implemented system has been released as open source: https://github.com/SeSaMe-NUS/genie.
H. V. Jagadish, Lubos Krcál, Wenhao Luan, Anthony K. H. Tung, Yueji Yang
ICDE8