Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Yang Li 0150

dblp:37/4190-150 · DBLP profile ↗
← Back
10ranked-venue papers
5as first author
0since 2021 · last 2019
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 8 · 3 first-authorArtificial intelligence and machine learning · 4 · 3 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
7 papers
Knowledge graphs · 46% Information retrieval · 18% Data integration and cleaning · 17%
Artificial intelligence
5 papers
Information extraction and text analysis · 55% Question answering and dialogue systems · 15% Knowledge representation and reasoning · 15%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Bioinformatics and computational biology · 100%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Storage systems · 100%

Topics — the 20 heaviest of 25, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Data integration and cleaning
entity disambiguation
0.422016
Entity Disambiguation with Linkless Knowledge Bases · WWW 2016
Mining evidences for named entity disambiguation · KDD 2013
Knowledge graphs › knowledge graph construction
knowledge extraction
0.412019
MIDAS: Finding the Right Web Sources to Fill Knowledge Gaps · ICDE 2019
Knowledge graphs
knowledge graph augmentation
0.412019
MIDAS: Finding the Right Web Sources to Fill Knowledge Gaps · ICDE 2019
Information retrieval › fact-checking
evidence verification
0.312017
Knowledge Verification for LongTail Verticals · Proc. VLDB Endow. 2017
Knowledge graphs
knowledge base integration
0.312017
Knowledge Verification for LongTail Verticals · Proc. VLDB Endow. 2017
Knowledge, reasoning and agents › Knowledge representation and reasoning
knowledge graph
0.212015
Answering Elementary Science Questions by Constructing Coherent Scenes using Background Knowledge · EMNLP 2015
Natural language and speech › Information extraction and text analysis
named entity recognition
0.212015
Leveraging Pattern Semantics for Extracting Entities in Enterprises · WWW 2015
Robotics › Robot navigation and mapping › environment representation
scene modeling
0.212015
Answering Elementary Science Questions by Constructing Coherent Scenes using Background Knowledge · EMNLP 2015
Natural language and speech › Question answering and dialogue systems › domain-specific question answering
science question answering
0.212015
Answering Elementary Science Questions by Constructing Coherent Scenes using Background Knowledge · EMNLP 2015
Web and social media mining › scholarly data mining
collaboration network analysis
0.212014
Analyzing expert behaviors in collaborative networks · KDD 2014
Information retrieval › question answering
question routing
0.212014
Analyzing expert behaviors in collaborative networks · KDD 2014
Data mining › text mining
sentiment analysis
0.212014
Interpreting the Public Sentiment Variations on Twitter · IEEE Trans. Knowl. Data Eng. 2014
Data mining › text mining
topic modeling
0.212014
Interpreting the Public Sentiment Variations on Twitter · IEEE Trans. Knowl. Data Eng. 2014
Natural language and speech › Information extraction and text analysis
entity linking
0.212013
Mining evidences for named entity disambiguation · KDD 2013
Bioinformatics and computational biology › sequence analysis › sequence assembly graph
de bruijn graph construction
0.212013
Memory Efficient Minimum Substring Partitioning · Proc. VLDB Endow. 2013
Bioinformatics and computational biology › sequence analysis › sequence assembly
genome assembly
0.212013
Memory Efficient Minimum Substring Partitioning · Proc. VLDB Endow. 2013
Data integration and cleaning
data quality
0.112019
MIDAS: Finding the Right Web Sources to Fill Knowledge Gaps · ICDE 2019
Information retrieval
query log analysis
0.112017
Knowledge Verification for LongTail Verticals · Proc. VLDB Endow. 2017
Natural language and speech › Information extraction and text analysis › entity linking
entity disambiguation
0.112016
Entity Disambiguation with Linkless Knowledge Bases · WWW 2016
Knowledge graphs
knowledge graph construction
0.112015
Leveraging Pattern Semantics for Extracting Entities in Enterprises · WWW 2015

Methods — techniques the papers use, named apart from their topics

generative model · 1.0semantic pattern graph · 0.4bootstrapping · 0.4web source slicing · 0.4profit function · 0.4minimum substring partitioning · 0.3k-mer overlap compression · 0.3end-to-end framework · 0.3distant supervised learning · 0.3knowledge fusion · 0.3evidence collection · 0.3knowledge graph · 0.2background knowledge · 0.2latent dirichlet allocation · 0.2generative modeling · 0.2cognitive process modeling · 0.2incremental algorithm · 0.2
YearPublicationVenuePosition
2019 MIDAS: Finding the Right Web Sources to Fill Knowledge Gaps
abstract
Knowledge bases, massive collections of facts (RDF triples) on diverse topics, support vital modern applications. However, existing knowledge bases contain very little data compared to the wealth of information on the Web. This is because the industry standard in knowledge base creation and augmentation suffers from a serious bottleneck: they rely on domain experts to identify appropriate web sources to extract data from. Efforts to fully automate knowledge extraction have failed to improve this standard: these automated systems are able to retrieve much more data and from a broader range of sources, but they suffer from very low precision and recall. As a result, these large-scale extractions remain unexploited. In this paper, we present MIDAS, a system that harnesses the results of automated knowledge extraction pipelines to repair the bottleneck in industrial knowledge creation and augmentation processes. MIDAS automates the suggestion of good-quality web sources and describes what to extract with respect to augmenting an existing knowledge base. We make three major contributions. First, we introduce a novel concept, web source slices, to describe the contents of a web source. Second, we define a profit function to quantify the value of a web source slice with respect to augmenting an existing knowledge base. Third, we develop effective and highly-scalable algorithms to derive high-profit web source slices. We demonstrate that MIDAS produces high-profit results and outperforms the baselines significantly on both real-world and synthetic datasets.
Xiaolan Wang 0001, Xin Dong 0001, Yang Li 0150, Alexandra Meliou
ICDE3
2018 Guess Me if You Can: Acronym Disambiguation for Enterprises
abstract
Acronyms are abbreviations formed from the initial components of words or phrases.In enterprises, people often use acronyms to make communications more efficient.However, acronyms could be difficult to understand for people who are not familiar with the subject matter (new employees, etc.), thereby affecting productivity.To alleviate such troubles, we study how to automatically resolve the true meanings of acronyms in a given context.Acronym disambiguation for enterprises is challenging for several reasons.First, acronyms may be highly ambiguous since an acronym used in the enterprise could have multiple internal and external meanings.Second, there are usually no comprehensive knowledge bases such as Wikipedia available in enterprises.Finally, the system should be generic to work for any enterprise.In this work we propose an end-to-end framework to tackle all these challenges.The framework takes the enterprise corpus as input and produces a high-quality acronym disambiguation system as output.Our disambiguation models are trained via distant supervised learning, without requiring any manually labeled training examples.Therefore, our proposed framework can be deployed to any enterprise to support highquality acronym disambiguation.Experimental results on real world data justified the effectiveness of our system.
Yang Li 0150, Bo Zhao 0001, Ariel Fuxman, Fangbo Tao
ACL (1)1
2017 Knowledge Verification for LongTail Verticals
abstract
Collecting structured knowledge for real-world entities has become a critical task for many applications. A big gap between the knowledge in existing knowledge repositories and the knowledge in the real world is the knowledge on tail verticals (i.e., less popular domains). Such knowledge, though not necessarily globally popular, can be personal hobbies to many people and thus collectively impactful. This paper studies the problem of knowledge verification for tail verticals ; that is, deciding the correctness of a given triple. Through comprehensive experimental study we answer the following questions. 1) Can we find evidence for tail knowledge from an extensive set of sources, including knowledge bases, the web, and query logs? 2) Can we judge correctness of the triples based on the collected evidence? 3) How can we further improve knowledge verification on tail verticals? Our empirical study suggests a new knowledge-verification framework, which we call F acty , that applies various kinds of evidence collection techniques followed by knowledge fusion. F acty can verify 50% of the (correct) tail knowledge with a precision of 84%, and it significantly outperforms state-of-the-art methods. Detailed error analysis on the obtained results suggests future research directions.
Xin Dong 0001, Anno Langen, Yang Li 0150
Proc. VLDB Endow.4
2016 Entity Disambiguation with Linkless Knowledge Bases
abstract
Named Entity Disambiguation is the task of disambiguating named entity mentions in natural language text and link them to their corresponding entries in a reference knowledge base (e.g. Wikipedia). Such disambiguation can help add semantics to plain text and distinguish homonymous entities. Previous research has tackled this problem by making use of two types of context-aware features derived from the reference knowledge base, namely, the context similarity and the semantic relatedness. Both features heavily rely on the cross-document hyperlinks within the knowledge base: the semantic relatedness feature is directly measured via those hyperlinks, while the context similarity feature implicitly makes use of those hyperlinks to expand entity candidates' descriptions and then compares them against the query context. Unfortunately, cross-document hyperlinks are rarely available in many closed domain knowledge bases and it is very expensive to manually add such links. Therefore few algorithms can work well on linkless knowledge bases. In this work, we propose the challenging Named Entity Disambiguation with Linkless Knowledge Bases (LNED) problem and tackle it by leveraging the useful disambiguation evidences scattered across the reference knowledge base. We propose a generative model to automatically mine such evidences out of noisy information. The mined evidences can mimic the role of the missing links and help boost the LNED performance. Experimental results show that our proposed method substantially improves the disambiguation accuracy over the baseline approaches.
Yang Li 0150, Shulong Tan, Huan Sun 0001, Jiawei Han 0001, Dan Roth 0001, Xifeng Yan
WWW1
2015 Answering Elementary Science Questions by Constructing Coherent Scenes using Background Knowledge
abstract
Much of what we understand from text is not explicitly stated.Rather, the reader uses his/her knowledge to fill in gaps and create a coherent, mental picture or "scene" depicting what text appears to convey.The scene constitutes an understanding of the text, and can be used to answer questions that go beyond the text.Our goal is to answer elementary science questions, where this requirement is pervasive; A question will often give a partial description of a scene and ask the student about implicit information.We show that by using a simple "knowledge graph" representation of the question, we can leverage several large-scale linguistic resources to provide missing background knowledge, somewhat alleviating the knowledge bottleneck in previous approaches.The coherence of the best resulting scene, built from a question/answer-candidate pair, reflects the confidence that the answer candidate is correct, and thus can be used to answer multiple choice questions.Our experiments show that this approach outperforms competitive algorithms on several datasets tested.The significance of this work is thus to show that a simple "knowledge graph" representation allows a version of "interpretation as scene construction" to be made viable.
Yang Li 0150, Peter Clark
EMNLP1
2015 Leveraging Pattern Semantics for Extracting Entities in Enterprises
abstract
Entity Extraction is a process of identifying meaningful entities from text documents. In enterprises, extracting entities improves enterprise efficiency by facilitating numerous applications, including search, recommendation, etc. However, the problem is particularly challenging on enterprise domains due to several reasons. First, the lack of redundancy of enterprise entities makes previous web-based systems like NELL and OpenIE not effective, since using only high-precision/low-recall patterns like those systems would miss the majority of sparse enterprise entities, while using more low-precision patterns in sparse setting also introduces noise drastically. Second, semantic drift is common in enterprises ("Blue" refers to "Windows Blue"), such that public signals from the web cannot be directly applied on entities. Moreover, many internal entities never appear on the web. Sparse internal signals are the only source for discovering them. To address these challenges, we propose an end-to-end framework for extracting entities in enterprises, taking the input of enterprise corpus and limited seeds to generate a high-quality entity collection as output. We introduce the novel concept of Semantic Pattern Graph to leverage public signals to understand the underlying semantics of lexical patterns, reinforce pattern evaluation using mined semantics, and yield more accurate and complete entities. Experiments on Microsoft enterprise data show the effectiveness of our approach.
Fangbo Tao, Bo Zhao 0001, Ariel Fuxman, Yang Li 0150, Jiawei Han 0001
WWW4
2014 Analyzing expert behaviors in collaborative networks
abstract
Collaborative networks are composed of experts who cooperate with each other to complete specific tasks, such as resolving problems reported by customers. A task is posted and subsequently routed in the network from an expert to another until being resolved. When an expert cannot solve a task, his routing decision (i.e., where to transfer a task) is critical since it can significantly affect the completion time of a task. In this work, we attempt to deduce the cognitive process of task routing, and model the decision making of experts as a generative process where a routing decision is made based on mixed routing patterns.
Huan Sun 0001, Mudhakar Srivatsa, Shulong Tan, Yang Li 0150, Lance M. Kaplan, Shu Tao, Xifeng Yan
KDD4
2014 Interpreting the Public Sentiment Variations on Twitter
abstract
Millions of users share their opinions on Twitter, making it a valuable platform for tracking and analyzing public sentiment. Such tracking and analysis can provide critical information for decision making in various domains. Therefore it has attracted attention in both academia and industry. Previous research mainly focused on modeling and tracking public sentiment. In this work, we move one step further to interpret sentiment variations. We observed that emerging topics (named foreground topics) within the sentiment variation periods are highly related to the genuine reasons behind the variations. Based on this observation, we propose a Latent Dirichlet Allocation (LDA) based model, Foreground and Background LDA (FB-LDA), to distill foreground topics and filter out longstanding background topics. These foreground topics can give potential interpretations of the sentiment variations. To further enhance the readability of the mined reasons, we select the most representative tweets for foreground topics and develop another generative model called Reason Candidate and Background LDA (RCB-LDA) to rank them with respect to their “popularity” within the variation period. Experimental results show that our methods can effectively find foreground topics and rank reason candidates. The proposed models can also be applied to other tasks such as finding topic differences between two sets of documents.
Shulong Tan, Yang Li 0150, Huan Sun 0001, Ziyu Guan, Xifeng Yan, Jiajun Bu, Chun Chen 0001, Xiaofei He 0001
IEEE Trans. Knowl. Data Eng.2
2013 Mining evidences for named entity disambiguation
abstract
Named entity disambiguation is the task of disambiguating named entity mentions in natural language text and link them to their corresponding entries in a knowledge base such as Wikipedia. Such disambiguation can help enhance readability and add semantics to plain text. It is also a central step in constructing high-quality information network or knowledge graph from unstructured text. Previous research has tackled this problem by making use of various textual and structural features from a knowledge base. Most of the proposed algorithms assume that a knowledge base can provide enough explicit and useful information to help disambiguate a mention to the right entity. However, the existing knowledge bases are rarely complete (likely will never be), thus leading to poor performance on short queries with not well-known contexts. In such cases, we need to collect additional evidences scattered in internal and external corpus to augment the knowledge bases and enhance their disambiguation power. In this work, we propose a generative model and an incremental algorithm to automatically mine useful evidences across documents. With a specific modeling of "background topic" and "unknown entities", our model is able to harvest useful evidences out of noisy information. Experimental results show that our proposed method outperforms the state-of-the-art approaches significantly: boosting the disambiguation accuracy from 43% (baseline) to 86% on short queries derived from tweets.
Yang Li 0150, Chi Wang 0001, Fangqiu Han, Jiawei Han 0001, Dan Roth 0001, Xifeng Yan
KDD1
2013 Memory Efficient Minimum Substring Partitioning
abstract
Massively parallel DNA sequencing technologies are revolutionizing genomics research. Billions of short reads generated at low costs can be assembled for reconstructing the whole genomes. Unfortunately, the large memory footprint of the existing de novo assembly algorithms makes it challenging to get the assembly done for higher eukaryotes like mammals. In this work, we investigate the memory issue of constructing de Bruijn graph, a core task in leading assembly algorithms, which often consumes several hundreds of gigabytes memory for large genomes. We propose a disk-based partition method, called Minimum Substring Partitioning (MSP), to complete the task using less than 10 gigabytes memory, without runtime slowdown. MSP breaks the short reads into multiple small disjoint partitions so that each partition can be loaded into memory, processed individually and later merged with others to form a de Bruijn graph. By leveraging the overlaps among the k-mers (substring of length k), MSP achieves astonishing compression ratio: The total size of partitions is reduced from Θ(kn) to Θ(n), wherenis the size of the short read database, andkis the length of ak-mer. Experimental results show that our method can build de Bruijn graphs using a commodity computer for any large-volume sequence dataset.
Yang Li 0150, Pegah Kamousi, Fangqiu Han, Shengqi Yang, Xifeng Yan, Subhash Suri
Proc. VLDB Endow.1