VLDB 2026 Research / reviewers in the wild / expert
Yang Li 0150
dblp:37/4190-150
· DBLP profile ↗
10ranked-venue papers
5as first author
0since 2021 · last 2019
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 8 · 3 first-authorArtificial intelligence and machine learning · 4 · 3 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
7 papers |
Knowledge graphs · 46% Information retrieval · 18% Data integration and cleaning · 17% | |
| Artificial intelligence
5 papers |
Information extraction and text analysis · 55% Question answering and dialogue systems · 15% Knowledge representation and reasoning · 15% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Bioinformatics and computational biology · 100% | |
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Storage systems · 100% |
Topics — the 20 heaviest of 25, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Data integration and cleaning
entity disambiguation |
0.4 | 2 | 2016 | Entity Disambiguation with Linkless Knowledge Bases · WWW 2016 Mining evidences for named entity disambiguation · KDD 2013 |
Knowledge graphs › knowledge graph construction
knowledge extraction |
0.4 | 1 | 2019 | MIDAS: Finding the Right Web Sources to Fill Knowledge Gaps · ICDE 2019 |
Knowledge graphs
knowledge graph augmentation |
0.4 | 1 | 2019 | MIDAS: Finding the Right Web Sources to Fill Knowledge Gaps · ICDE 2019 |
Information retrieval › fact-checking
evidence verification |
0.3 | 1 | 2017 | Knowledge Verification for LongTail Verticals · Proc. VLDB Endow. 2017 |
Knowledge graphs
knowledge base integration |
0.3 | 1 | 2017 | Knowledge Verification for LongTail Verticals · Proc. VLDB Endow. 2017 |
Knowledge, reasoning and agents › Knowledge representation and reasoning
knowledge graph |
0.2 | 1 | 2015 | Answering Elementary Science Questions by Constructing Coherent Scenes using Background Knowledge · EMNLP 2015 |
Natural language and speech › Information extraction and text analysis
named entity recognition |
0.2 | 1 | 2015 | Leveraging Pattern Semantics for Extracting Entities in Enterprises · WWW 2015 |
Robotics › Robot navigation and mapping › environment representation
scene modeling |
0.2 | 1 | 2015 | Answering Elementary Science Questions by Constructing Coherent Scenes using Background Knowledge · EMNLP 2015 |
Natural language and speech › Question answering and dialogue systems › domain-specific question answering
science question answering |
0.2 | 1 | 2015 | Answering Elementary Science Questions by Constructing Coherent Scenes using Background Knowledge · EMNLP 2015 |
Web and social media mining › scholarly data mining
collaboration network analysis |
0.2 | 1 | 2014 | Analyzing expert behaviors in collaborative networks · KDD 2014 |
Information retrieval › question answering
question routing |
0.2 | 1 | 2014 | Analyzing expert behaviors in collaborative networks · KDD 2014 |
Data mining › text mining
sentiment analysis |
0.2 | 1 | 2014 | Interpreting the Public Sentiment Variations on Twitter · IEEE Trans. Knowl. Data Eng. 2014 |
Data mining › text mining
topic modeling |
0.2 | 1 | 2014 | Interpreting the Public Sentiment Variations on Twitter · IEEE Trans. Knowl. Data Eng. 2014 |
Natural language and speech › Information extraction and text analysis
entity linking |
0.2 | 1 | 2013 | Mining evidences for named entity disambiguation · KDD 2013 |
Bioinformatics and computational biology › sequence analysis › sequence assembly graph
de bruijn graph construction |
0.2 | 1 | 2013 | Memory Efficient Minimum Substring Partitioning · Proc. VLDB Endow. 2013 |
Bioinformatics and computational biology › sequence analysis › sequence assembly
genome assembly |
0.2 | 1 | 2013 | Memory Efficient Minimum Substring Partitioning · Proc. VLDB Endow. 2013 |
Data integration and cleaning
data quality |
0.1 | 1 | 2019 | MIDAS: Finding the Right Web Sources to Fill Knowledge Gaps · ICDE 2019 |
Information retrieval
query log analysis |
0.1 | 1 | 2017 | Knowledge Verification for LongTail Verticals · Proc. VLDB Endow. 2017 |
Natural language and speech › Information extraction and text analysis › entity linking
entity disambiguation |
0.1 | 1 | 2016 | Entity Disambiguation with Linkless Knowledge Bases · WWW 2016 |
Knowledge graphs
knowledge graph construction |
0.1 | 1 | 2015 | Leveraging Pattern Semantics for Extracting Entities in Enterprises · WWW 2015 |
Methods — techniques the papers use, named apart from their topics
generative model · 1.0semantic pattern graph · 0.4bootstrapping · 0.4web source slicing · 0.4profit function · 0.4minimum substring partitioning · 0.3k-mer overlap compression · 0.3end-to-end framework · 0.3distant supervised learning · 0.3knowledge fusion · 0.3evidence collection · 0.3knowledge graph · 0.2background knowledge · 0.2latent dirichlet allocation · 0.2generative modeling · 0.2cognitive process modeling · 0.2incremental algorithm · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2019 | MIDAS: Finding the Right Web Sources to Fill Knowledge GapsabstractKnowledge bases, massive collections of facts (RDF triples) on diverse topics, support vital modern applications. However, existing knowledge bases contain very little data compared to the wealth of information on the Web. This is because the industry standard in knowledge base creation and augmentation suffers from a serious bottleneck: they rely on domain experts to identify appropriate web sources to extract data from. Efforts to fully automate knowledge extraction have failed to improve this standard: these automated systems are able to retrieve much more data and from a broader range of sources, but they suffer from very low precision and recall. As a result, these large-scale extractions remain unexploited. In this paper, we present MIDAS, a system that harnesses the results of automated knowledge extraction pipelines to repair the bottleneck in industrial knowledge creation and augmentation processes. MIDAS automates the suggestion of good-quality web sources and describes what to extract with respect to augmenting an existing knowledge base. We make three major contributions. First, we introduce a novel concept, web source slices, to describe the contents of a web source. Second, we define a profit function to quantify the value of a web source slice with respect to augmenting an existing knowledge base. Third, we develop effective and highly-scalable algorithms to derive high-profit web source slices. We demonstrate that MIDAS produces high-profit results and outperforms the baselines significantly on both real-world and synthetic datasets. Xiaolan Wang 0001, Xin Dong 0001, Yang Li 0150, Alexandra Meliou |
ICDE | 3 |
| 2018 | Guess Me if You Can: Acronym Disambiguation for EnterprisesabstractAcronyms are abbreviations formed from the initial components of words or phrases.In enterprises, people often use acronyms to make communications more efficient.However, acronyms could be difficult to understand for people who are not familiar with the subject matter (new employees, etc.), thereby affecting productivity.To alleviate such troubles, we study how to automatically resolve the true meanings of acronyms in a given context.Acronym disambiguation for enterprises is challenging for several reasons.First, acronyms may be highly ambiguous since an acronym used in the enterprise could have multiple internal and external meanings.Second, there are usually no comprehensive knowledge bases such as Wikipedia available in enterprises.Finally, the system should be generic to work for any enterprise.In this work we propose an end-to-end framework to tackle all these challenges.The framework takes the enterprise corpus as input and produces a high-quality acronym disambiguation system as output.Our disambiguation models are trained via distant supervised learning, without requiring any manually labeled training examples.Therefore, our proposed framework can be deployed to any enterprise to support highquality acronym disambiguation.Experimental results on real world data justified the effectiveness of our system. Yang Li 0150, Bo Zhao 0001, Ariel Fuxman, Fangbo Tao |
ACL (1) | 1 |
| 2017 | Knowledge Verification for LongTail VerticalsabstractCollecting structured knowledge for real-world entities has become a critical task for many applications. A big gap between the knowledge in existing knowledge repositories and the knowledge in the real world is the knowledge on tail verticals (i.e., less popular domains). Such knowledge, though not necessarily globally popular, can be personal hobbies to many people and thus collectively impactful. This paper studies the problem of knowledge verification for tail verticals ; that is, deciding the correctness of a given triple. Through comprehensive experimental study we answer the following questions. 1) Can we find evidence for tail knowledge from an extensive set of sources, including knowledge bases, the web, and query logs? 2) Can we judge correctness of the triples based on the collected evidence? 3) How can we further improve knowledge verification on tail verticals? Our empirical study suggests a new knowledge-verification framework, which we call F acty , that applies various kinds of evidence collection techniques followed by knowledge fusion. F acty can verify 50% of the (correct) tail knowledge with a precision of 84%, and it significantly outperforms state-of-the-art methods. Detailed error analysis on the obtained results suggests future research directions. Xin Dong 0001, Anno Langen, Yang Li 0150 |
Proc. VLDB Endow. | 4 |
| 2016 | Entity Disambiguation with Linkless Knowledge BasesabstractNamed Entity Disambiguation is the task of disambiguating named entity mentions in natural language text and link them to their corresponding entries in a reference knowledge base (e.g. Wikipedia). Such disambiguation can help add semantics to plain text and distinguish homonymous entities. Previous research has tackled this problem by making use of two types of context-aware features derived from the reference knowledge base, namely, the context similarity and the semantic relatedness. Both features heavily rely on the cross-document hyperlinks within the knowledge base: the semantic relatedness feature is directly measured via those hyperlinks, while the context similarity feature implicitly makes use of those hyperlinks to expand entity candidates' descriptions and then compares them against the query context. Unfortunately, cross-document hyperlinks are rarely available in many closed domain knowledge bases and it is very expensive to manually add such links. Therefore few algorithms can work well on linkless knowledge bases. In this work, we propose the challenging Named Entity Disambiguation with Linkless Knowledge Bases (LNED) problem and tackle it by leveraging the useful disambiguation evidences scattered across the reference knowledge base. We propose a generative model to automatically mine such evidences out of noisy information. The mined evidences can mimic the role of the missing links and help boost the LNED performance. Experimental results show that our proposed method substantially improves the disambiguation accuracy over the baseline approaches. Yang Li 0150, Shulong Tan, Huan Sun 0001, Jiawei Han 0001, Dan Roth 0001, Xifeng Yan |
WWW | 1 |
| 2015 | Answering Elementary Science Questions by Constructing Coherent Scenes using Background KnowledgeabstractMuch of what we understand from text is not explicitly stated.Rather, the reader uses his/her knowledge to fill in gaps and create a coherent, mental picture or "scene" depicting what text appears to convey.The scene constitutes an understanding of the text, and can be used to answer questions that go beyond the text.Our goal is to answer elementary science questions, where this requirement is pervasive; A question will often give a partial description of a scene and ask the student about implicit information.We show that by using a simple "knowledge graph" representation of the question, we can leverage several large-scale linguistic resources to provide missing background knowledge, somewhat alleviating the knowledge bottleneck in previous approaches.The coherence of the best resulting scene, built from a question/answer-candidate pair, reflects the confidence that the answer candidate is correct, and thus can be used to answer multiple choice questions.Our experiments show that this approach outperforms competitive algorithms on several datasets tested.The significance of this work is thus to show that a simple "knowledge graph" representation allows a version of "interpretation as scene construction" to be made viable. Yang Li 0150, Peter Clark |
EMNLP | 1 |
| 2015 | Leveraging Pattern Semantics for Extracting Entities in EnterprisesabstractEntity Extraction is a process of identifying meaningful entities from text documents. In enterprises, extracting entities improves enterprise efficiency by facilitating numerous applications, including search, recommendation, etc. However, the problem is particularly challenging on enterprise domains due to several reasons. First, the lack of redundancy of enterprise entities makes previous web-based systems like NELL and OpenIE not effective, since using only high-precision/low-recall patterns like those systems would miss the majority of sparse enterprise entities, while using more low-precision patterns in sparse setting also introduces noise drastically. Second, semantic drift is common in enterprises ("Blue" refers to "Windows Blue"), such that public signals from the web cannot be directly applied on entities. Moreover, many internal entities never appear on the web. Sparse internal signals are the only source for discovering them. To address these challenges, we propose an end-to-end framework for extracting entities in enterprises, taking the input of enterprise corpus and limited seeds to generate a high-quality entity collection as output. We introduce the novel concept of Semantic Pattern Graph to leverage public signals to understand the underlying semantics of lexical patterns, reinforce pattern evaluation using mined semantics, and yield more accurate and complete entities. Experiments on Microsoft enterprise data show the effectiveness of our approach. Fangbo Tao, Bo Zhao 0001, Ariel Fuxman, Yang Li 0150, Jiawei Han 0001 |
WWW | 4 |
| 2014 | Analyzing expert behaviors in collaborative networksabstractCollaborative networks are composed of experts who cooperate with each other to complete specific tasks, such as resolving problems reported by customers. A task is posted and subsequently routed in the network from an expert to another until being resolved. When an expert cannot solve a task, his routing decision (i.e., where to transfer a task) is critical since it can significantly affect the completion time of a task. In this work, we attempt to deduce the cognitive process of task routing, and model the decision making of experts as a generative process where a routing decision is made based on mixed routing patterns. Huan Sun 0001, Mudhakar Srivatsa, Shulong Tan, Yang Li 0150, Lance M. Kaplan, Shu Tao, Xifeng Yan |
KDD | 4 |
| 2014 | Interpreting the Public Sentiment Variations on TwitterabstractMillions of users share their opinions on Twitter, making it a valuable platform for tracking and analyzing public sentiment. Such tracking and analysis can provide critical information for decision making in various domains. Therefore it has attracted attention in both academia and industry. Previous research mainly focused on modeling and tracking public sentiment. In this work, we move one step further to interpret sentiment variations. We observed that emerging topics (named foreground topics) within the sentiment variation periods are highly related to the genuine reasons behind the variations. Based on this observation, we propose a Latent Dirichlet Allocation (LDA) based model, Foreground and Background LDA (FB-LDA), to distill foreground topics and filter out longstanding background topics. These foreground topics can give potential interpretations of the sentiment variations. To further enhance the readability of the mined reasons, we select the most representative tweets for foreground topics and develop another generative model called Reason Candidate and Background LDA (RCB-LDA) to rank them with respect to their “popularity” within the variation period. Experimental results show that our methods can effectively find foreground topics and rank reason candidates. The proposed models can also be applied to other tasks such as finding topic differences between two sets of documents. Shulong Tan, Yang Li 0150, Huan Sun 0001, Ziyu Guan, Xifeng Yan, Jiajun Bu, Chun Chen 0001, Xiaofei He 0001 |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2013 | Mining evidences for named entity disambiguationabstractNamed entity disambiguation is the task of disambiguating named entity mentions in natural language text and link them to their corresponding entries in a knowledge base such as Wikipedia. Such disambiguation can help enhance readability and add semantics to plain text. It is also a central step in constructing high-quality information network or knowledge graph from unstructured text. Previous research has tackled this problem by making use of various textual and structural features from a knowledge base. Most of the proposed algorithms assume that a knowledge base can provide enough explicit and useful information to help disambiguate a mention to the right entity. However, the existing knowledge bases are rarely complete (likely will never be), thus leading to poor performance on short queries with not well-known contexts. In such cases, we need to collect additional evidences scattered in internal and external corpus to augment the knowledge bases and enhance their disambiguation power. In this work, we propose a generative model and an incremental algorithm to automatically mine useful evidences across documents. With a specific modeling of "background topic" and "unknown entities", our model is able to harvest useful evidences out of noisy information. Experimental results show that our proposed method outperforms the state-of-the-art approaches significantly: boosting the disambiguation accuracy from 43% (baseline) to 86% on short queries derived from tweets. Yang Li 0150, Chi Wang 0001, Fangqiu Han, Jiawei Han 0001, Dan Roth 0001, Xifeng Yan |
KDD | 1 |
| 2013 | Memory Efficient Minimum Substring PartitioningabstractMassively parallel DNA sequencing technologies are revolutionizing genomics research. Billions of short reads generated at low costs can be assembled for reconstructing the whole genomes. Unfortunately, the large memory footprint of the existing de novo assembly algorithms makes it challenging to get the assembly done for higher eukaryotes like mammals. In this work, we investigate the memory issue of constructing de Bruijn graph, a core task in leading assembly algorithms, which often consumes several hundreds of gigabytes memory for large genomes. We propose a disk-based partition method, called Minimum Substring Partitioning (MSP), to complete the task using less than 10 gigabytes memory, without runtime slowdown. MSP breaks the short reads into multiple small disjoint partitions so that each partition can be loaded into memory, processed individually and later merged with others to form a de Bruijn graph. By leveraging the overlaps among the k-mers (substring of length k), MSP achieves astonishing compression ratio: The total size of partitions is reduced from Θ(kn) to Θ(n), wherenis the size of the short read database, andkis the length of ak-mer. Experimental results show that our method can build de Bruijn graphs using a commodity computer for any large-volume sequence dataset. Yang Li 0150, Pegah Kamousi, Fangqiu Han, Shengqi Yang, Xifeng Yan, Subhash Suri |
Proc. VLDB Endow. | 1 |