EDBT 2026 Demo / reviewers in the wild / expert
Cong Yu 0001
dblp:58/3771
· DBLP profile ↗
84ranked-venue papers
9as first author
5since 2021 · last 2022
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 78 · 9 first-author · 2 since 2021Artificial intelligence and machine learning · 14 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
63 papers |
Information retrieval · 24% Data mining · 14% Data integration and cleaning · 14% | |
| Artificial intelligence
17 papers |
Language models and text generation · 24% Representation and self-supervised learning · 16% Question answering and dialogue systems · 16% |
Topics — the 30 heaviest of 121, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Information retrieval
fact-checking |
1.3 | 5 | 2019 | AggChecker: A Fact-Checking System for Text Summaries of Relational Data Sets · Proc. VLDB Endow. 2019 Verifying Text Summaries of Relational Data Sets · SIGMOD Conference 2019 Computational Fact Checking through Query Perturbations · ACM Trans. Database Syst. 2017 |
Data mining
pattern mining |
1.3 | 7 | 2020 | On Detecting Cherry-picked Trendlines · Proc. VLDB Endow. 2020 An expressive framework and efficient algorithms for the analysis of collaborative tagging · VLDB J. 2014 Incremental discovery of prominent situational facts · ICDE 2014 |
Machine learning › Representation and self-supervised learning
pre-training |
0.9 | 2 | 2021 | ReasonBERT: Pre-trained to Reason with Distant Supervision · EMNLP (1) 2021 TURL: Table Understanding through Representation Learning · Proc. VLDB Endow. 2020 |
Data integration and cleaning
table understanding |
0.8 | 2 | 2020 | TURL: Table Understanding through Representation Learning · Proc. VLDB Endow. 2020 Generating Titles for Web Tables · WWW 2019 |
Knowledge graphs
knowledge graph construction |
0.6 | 2 | 2019 | Mining an "Anti-Knowledge Base" from Wikipedia Updates with Applications to Fact Checking and Beyond · Proc. VLDB Endow. 2019 LONLIES: Estimating Property Values for Long Tail Entities · SIGIR 2016 |
Information retrieval
ranking |
0.6 | 3 | 2019 | Contextual Fact Ranking and Its Applications in Table Synthesis and Compression · KDD 2019 On "one of the few" objects · KDD 2012 Incremental discovery of prominent situational facts · ICDE 2014 |
Machine learning › Transfer learning and domain adaptation
cross-modality alignment |
0.6 | 1 | 2022 | Scaling Multimodal Pre-Training via Cross-Modality Gradient Harmonization · NeurIPS 2022 |
Machine learning › Deep learning architectures and training › training optimization
gradient harmonization |
0.6 | 1 | 2022 | Scaling Multimodal Pre-Training via Cross-Modality Gradient Harmonization · NeurIPS 2022 |
Machine learning › Representation and self-supervised learning › pre-training
multimodal pretraining |
0.6 | 1 | 2022 | Scaling Multimodal Pre-Training via Cross-Modality Gradient Harmonization · NeurIPS 2022 |
Natural language and speech › Language models and text generation › tokenization
subword tokenization |
0.6 | 1 | 2022 | Charformer: Fast Character Transformers via Gradient-based Subword Tokenization · ICLR 2022 |
Natural language and speech › Language models and text generation
tokenization |
0.6 | 1 | 2022 | Charformer: Fast Character Transformers via Gradient-based Subword Tokenization · ICLR 2022 |
Machine learning › Deep learning architectures and training
transformer |
0.6 | 1 | 2022 | Charformer: Fast Character Transformers via Gradient-based Subword Tokenization · ICLR 2022 |
Natural language and speech › Language models and text generation
text generation |
0.5 | 2 | 2020 | Generating Titles for Web Tables · WWW 2019 Generating Representative Headlines for News Stories · WWW 2020 |
Natural language and speech › Question answering and dialogue systems › machine reading comprehension
extractive question answering |
0.5 | 1 | 2021 | ReasonBERT: Pre-trained to Reason with Distant Supervision · EMNLP (1) 2021 |
Natural language and speech › Question answering and dialogue systems › reasoning-based question answering
multi-hop question answering |
0.5 | 1 | 2021 | ReasonBERT: Pre-trained to Reason with Distant Supervision · EMNLP (1) 2021 |
Natural language and speech › Information extraction and text analysis › document analysis
news analysis |
0.5 | 1 | 2021 | Quiz-Style Question Generation for News Stories · WWW 2021 |
Natural language and speech › Language models and text generation
pre-trained language model |
0.5 | 1 | 2021 | NewsEmbed: Modeling News through Pre-trained Document Representations · KDD 2021 |
Natural language and speech › Question answering and dialogue systems
question generation |
0.5 | 1 | 2021 | Quiz-Style Question Generation for News Stories · WWW 2021 |
Information retrieval › document processing › document analysis
document representation |
0.5 | 1 | 2021 | NewsEmbed: Modeling News through Pre-trained Document Representations · KDD 2021 |
Information retrieval › fact-checking
claim verification |
0.5 | 2 | 2017 | Computational Fact Checking through Query Perturbations · ACM Trans. Database Syst. 2017 Toward Computational Fact-Checking · Proc. VLDB Endow. 2014 |
Knowledge, reasoning and agents › Knowledge representation and reasoning › knowledge acquisition › knowledge extraction
structured information extraction |
0.4 | 1 | 2020 | Factoring Fact-Checks: Structured Information Extraction from Fact-Checking Articles · WWW 2020 |
Data mining
anomaly detection |
0.4 | 1 | 2020 | On Detecting Cherry-picked Trendlines · Proc. VLDB Endow. 2020 |
Information retrieval › text summarization
multi-document summarization |
0.4 | 1 | 2020 | Generating Representative Headlines for News Stories · WWW 2020 |
Machine learning and data management
table representation learning |
0.4 | 1 | 2020 | TURL: Table Understanding through Representation Learning · Proc. VLDB Endow. 2020 |
Information retrieval
text summarization |
0.4 | 1 | 2020 | Generating Representative Headlines for News Stories · WWW 2020 |
Recommender systems
group recommendation |
0.4 | 3 | 2014 | Exploiting group recommendation functions for flexible preferences · ICDE 2014 Space efficiency in group recommendation · VLDB J. 2010 Group Recommendation: Semantics and Efficiency · Proc. VLDB Endow. 2009 |
Knowledge, reasoning and agents › Knowledge representation and reasoning › knowledge acquisition › knowledge extraction
fact extraction |
0.4 | 1 | 2019 | Automatically Generating Interesting Facts from Wikipedia Tables · SIGMOD Conference 2019 |
Query processing and optimization
aggregate query processing |
0.4 | 1 | 2019 | AggChecker: A Fact-Checking System for Text Summaries of Relational Data Sets · Proc. VLDB Endow. 2019 |
Data models and query languages
natural language interface |
0.4 | 1 | 2019 | Verifying Text Summaries of Relational Data Sets · SIGMOD Conference 2019 |
Query processing and optimization
parameterized queries |
0.3 | 3 | 2017 | iCheck: computationally combating "lies, d-ned lies, and statistics" · SIGMOD Conference 2014 Computational Fact Checking through Query Perturbations · ACM Trans. Database Syst. 2017 Toward Computational Fact-Checking · Proc. VLDB Endow. 2014 |
Methods — techniques the papers use, named apart from their topics
contrastive learning · 1.6probabilistic model · 1.5pre-training · 1.4multi-task learning · 1.0multi-label classification · 1.0BERT fine-tuning · 0.9gradient-based tokenization · 0.6gradient realignment · 0.6curriculum learning · 0.6distant supervision · 0.5meta-algorithm · 0.5support metric · 0.4sequence tagging · 0.4fine-tuning · 0.4computational social choice · 0.4approximate visualization algorithm · 0.2utility optimization · 0.1temporal constraint clustering · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | Charformer: Fast Character Transformers via Gradient-based Subword Tokenization
Yi Tay, Vinh Q. Tran 0002, Sebastian Ruder, Jai Gupta 0001, Hyung Won Chung, Dara Bahri, Zhen Qin 0001, Simon Baumgartner, Cong Yu 0001, Donald Metzler |
ICLR | 9 |
| 2022 | Scaling Multimodal Pre-Training via Cross-Modality Gradient HarmonizationabstractSelf-supervised pre-training recently demonstrates success on large-scale multimodal data, and state-of-the-art contrastive learning methods often enforce the feature consistency from cross-modality inputs, such as video/audio or video/text pairs. Despite its convenience to formulate and leverage in practice, such cross-modality alignment (CMA) is only a weak and noisy supervision, since two modalities can be semantically misaligned even they are temporally aligned. For example, even in the (often adopted) instructional videos, a speaker can sometimes refer to something that is not visually present in the current frame; and the semantic misalignment would only be more unpredictable for the raw videos collected from unconstrained internet sources. We conjecture that might cause conflicts and biases among modalities, and may hence prohibit CMA from scaling up to training with larger and more heterogeneous data. This paper first verifies our conjecture by observing that, even in the latest VATT pre-training using only narrated videos, there exist strong gradient conflicts between different CMA losses within the same sample triplet (video, audio, text), indicating them as the noisy source of supervision. We then propose to harmonize such gradients during pre-training, via two techniques: (i) cross-modality gradient realignment: modifying different CMA loss gradients for one sample triplet, so that their gradient directions are in more agreement; and (ii) gradient-based curriculum learning: leveraging the gradient conflict information on an indicator of sample noisiness, to develop a curriculum learning strategy to prioritize training with less noisy sample triplets. Applying those gradient harmonization techniques to pre-training VATT on the HowTo100M dataset, we consistently improve its performance on different downstream tasks. Moreover, we are able to scale VATT pre-training to more complicated non-narrative Youtube8M dataset to further improve the state-of-the-arts. Hassan Akbari, Zhangyang Wang, Cong Yu 0001 |
NeurIPS | 6 |
| 2021 | ReasonBERT: Pre-trained to Reason with Distant SupervisionabstractWe present ReasonBERT, a pre-training method that augments language models with the ability to reason over long-range relations and multiple, possibly hybrid, contexts.Unlike existing pre-training methods that only harvest learning signals from local contexts of naturally occurring texts, we propose a generalized notion of distant supervision to automatically connect multiple pieces of text and tables to create pre-training examples that require long-range reasoning.Different types of reasoning are simulated, including intersecting multiple pieces of evidence, bridging from one piece of evidence to another, and detecting unanswerable cases.We conduct a comprehensive evaluation on a variety of extractive question answering datasets ranging from single-hop to multi-hop and from text-only to table-only to hybrid that require various reasoning capabilities and show that ReasonBERT achieves remarkable improvement over an array of strong baselines.Fewshot experiments further demonstrate that our pre-training method substantially improves sample efficiency. 1 Xiang Deng 0001, Yu Su 0001, Alyssa Lees, You Wu 0001, Cong Yu 0001, Huan Sun 0001 |
EMNLP (1) | 5 |
| 2021 | NewsEmbed: Modeling News through Pre-trained Document RepresentationsabstractEffectively modeling text-rich fresh content such as news articles at document-level is a challenging problem. To ensure a content-based model generalize well to a broad range of applications, it is critical to have a training dataset that is large beyond the scale of human labels while achieving desired quality. In this work, we address those two challenges by proposing a novel approach to mine semantically-relevant fresh documents, and their topic labels, with little human supervision. Meanwhile, we design a multitask model called NewsEmbed that alternatively trains a contrastive learning with a multi-label classification to derive a universal document encoder. We show that the proposed approach can provide billions of high quality organic training examples and can be naturally extended to multilingual setting where texts in different languages are encoded in the same semantic space. We experimentally demonstrate NewsEmbed's competitive performance across multiple natural language understanding tasks, both supervised and unsupervised. Tianqi Liu 0002, Cong Yu 0001 |
KDD | 3 |
| 2021 | Quiz-Style Question Generation for News StoriesabstractA large majority of American adults get at least some of their news from the Internet. Even though many online news products have the goal of informing their users about the news, they lack scalable and reliable tools for measuring how well they are achieving this goal, and therefore have to resort to noisy proxy metrics (e.g., click-through rates or reading time) to track their performance. Ádám Dániel Lelkes, Vinh Q. Tran 0002, Cong Yu 0001 |
WWW | 3 |
| 2020 | Generating Representative Headlines for News StoriesabstractMillions of news articles are published online every day, which can be overwhelming for readers to follow. Grouping articles that are reporting the same event into news stories is a common way of assisting readers in their news consumption. However, it remains a challenging research problem to efficiently and effectively generate a representative headline for each story. Automatic summarization of a document set has been studied for decades, while few studies have focused on generating representative headlines for a set of articles. Unlike summaries, which aim to capture most information with least redundancy, headlines aim to capture information jointly shared by the story articles in short length and exclude information specific to each individual article. Xiaotao Gu, Yuning Mao, Jiawei Han 0001, You Wu 0001, Cong Yu 0001, Daniel Finnie, Hongkun Yu 0001, Jiaqi Zhai, Nicholas Zukoski |
WWW | 6 |
| 2020 | Factoring Fact-Checks: Structured Information Extraction from Fact-Checking ArticlesabstractFact-checking, which investigates claims made in public to arrive at a verdict supported by evidence and logical reasoning, has long been a significant form of journalism to combat misinformation in the news ecosystem. Most of the fact-checks share common structured information (called factors) such as claim, claimant, and verdict. In recent years, the emergence of ClaimReview as the standard schema for annotating those factors within fact-checking articles has led to wide adoption of fact-checking features by online platforms (e.g., Google, Bing). However, annotating fact-checks is a tedious process for fact-checkers and distracts them from their core job of investigating claims. As a result, less than half of the fact-checkers worldwide have adopted ClaimReview as of mid-2019. In this paper, we propose the task of factoring fact-checks for automatically extracting structured information from fact-checking articles. Exploring a public dataset of fact-checks, we empirically show that factoring fact-checks is a challenging task, especially for fact-checkers that are under-represented in the existing dataset. We then formulate the task as a sequence tagging problem and fine-tune the pre-trained BERT models with a modification made from our observations to approach the problem. Through extensive experiments, we demonstrate the performance of our models for well-known fact-checkers and promising initial results for under-represented fact-checkers. Shan Jiang 0008, Simon Baumgartner, Abraham Ittycheriah, Cong Yu 0001 |
WWW | 4 |
| 2020 | On Detecting Cherry-picked TrendlinesabstractPoorly supported stories can be told based on data by cherry-picking the data points included. While such stories may be technically accurate, they are misleading. In this paper, we build a system for detecting cherry-picking, with a focus on trendlines extracted from temporal data. We define a support metric for detecting such trendlines. Given a dataset and a statement made based on a trendline, we compute a support score that indicates how cherry-picked it is. Studying different types of trendlines and formalizing terms, we propose efficient and effective algorithms for computing the support measure. We also study the problem of discovering the most supported statements. Besides theoretical analysis, we conduct extensive experiments on real-world data, that demonstrate the validity of our proposed techniques. Abolfazl Asudeh, H. V. Jagadish, You Wu 0001, Cong Yu 0001 |
Proc. VLDB Endow. | 4 |
| 2020 | TURL: Table Understanding through Representation LearningabstractRelational tables on the Web store a vast amount of knowledge. Owing to the wealth of such tables, there has been tremendous progress on a variety of tasks in the area of table understanding. However, existing work generally relies on heavily-engineered task-specific features and model architectures. In this paper, we present TURL, a novel framework that introduces the pre-training/fine-tuning paradigm to relational Web tables. During pre-training, our framework learns deep contextualized representations on relational tables in an unsupervised manner. Its universal model design with pre-trained representations can be applied to a wide range of tasks with minimal task-specific fine-tuning. Specifically, we propose a structure-aware Transformer encoder to model the row-column structure of relational tables, and present a new Masked Entity Recovery (MER) objective for pre-training to capture the semantics and knowledge in large-scale unlabeled data. We systematically evaluate TURL with a benchmark consisting of 6 different tasks for table understanding (e.g., relation extraction, cell filling). We show that TURL generalizes well to all tasks and substantially outperforms existing methods in almost all instances. Xiang Deng 0001, Huan Sun 0001, Alyssa Lees, You Wu 0001, Cong Yu 0001 |
Proc. VLDB Endow. | 5 |
| 2019 | Contextual Fact Ranking and Its Applications in Table Synthesis and CompressionabstractModern search engines increasingly incorporate tabular content, which consists of a set of entities each augmented with a small set of facts. The facts can be obtained from multiple sources: an entity's knowledge base entry, the infobox on its Wikipedia page, or its row within a WebTable. Crucially, the informativeness of a fact depends not only on the entity but also the specific context(e.g., the query).To the best of our knowledge, this paper is the first to study the problem of contextual fact ranking: given some entities and a context (i.e., succinct natural language description), identify the most informative facts for the entities collectively within the context.We propose to contextually rank the facts by exploiting deep learning techniques. In particular, we develop pointwise and pair-wise ranking models, using textual and statistical information for the given entities and context derived from their sources. We enhance the models by incorporating entity type information from an IsA (hypernym) database. We demonstrate that our approaches achieve better performance than state-of-the-art baselines in terms of MAP, NDCG, and recall. We further conduct user studies for two specific applications of contextual fact ranking-table synthesis and table compression-and show that our models can identify more informative facts than the baselines. Silu Huang, Flip Korn, Xuezhi Wang 0002, You Wu 0001, Dale Markowitz, Cong Yu 0001 |
KDD | 7 |
| 2019 | Summarizing News Articles Using Question-and-Answer Pairs via Learning
Xuezhi Wang 0002, Cong Yu 0001 |
ISWC (1) | 2 |
| 2019 | Verifying Text Summaries of Relational Data SetsabstractWe present a novel natural language query interface, the AggChecker, aimed at text summaries of relational data sets. The tool focuses on natural language claims that translate into an SQL query and a claimed query result. Similar in spirit to a spell checker, the AggChecker marks up text passages that seem to be inconsistent with the actual data. At the heart of the system is a probabilistic model that reasons about the input document in a holistic fashion. Based on claim keywords and the document structure, it maps each text claim to a probability distribution over associated query translations. By efficiently executing tens to hundreds of thousands of candidate translations for a typical input document, the system maps text claims to correctness probabilities. This process becomes practical via a specialized processing backend, avoiding redundant work via query merging and result caching. Verification is an interactive process in which users are shown tentative results, enabling them to take corrective actions if necessary. We tested our system on 53 publicly available articles containing 392 claims. Our tool revealed erroneous claims in roughly a third of test cases. Also, AggChecker compares favorably against several automated and semi-automated fact checking baselines. Saehan Jo, Immanuel Trummer, Weicheng Yu, Xuezhi Wang 0002, Cong Yu 0001, Daniel Liu, Niyati Mehta |
SIGMOD Conference | 5 |
| 2019 | Automatically Generating Interesting Facts from Wikipedia TablesabstractModern search engines provide contextual information surrounding query entities beyond ten blue links in the form of information cards. Among the various attributes displayed about entities there has been recent interest in providing fun facts. Obtaining such trivia at a large scale is, however, non-trivial: hiring professional content creators is expensive and extracting statements from the Web is prone to uninteresting, out-of-context and/or unreliable facts. Flip Korn, Xuezhi Wang 0002, You Wu 0001, Cong Yu 0001 |
SIGMOD Conference | 4 |
| 2019 | Generating Titles for Web TablesabstractDescriptive titles provide crucial context for interpreting tables that are extracted from web pages and are a key component of search features such as tabular featured snippets from Google and Bing. Prior approaches have attempted to produce titles by selecting existing text snippets associated with the table. These approaches, however, are limited by their dependence on suitable titles existing a priori. In our user study, we observe that the relevant information for the title tends to be scattered across the page, and often-more than 80% of the time-does not appear verbatim anywhere in the page. We propose instead the application of a sequence-to-sequence neural network model as a more generalizable approach for generating high-quality table titles. This is accomplished by extracting many text snippets that have potentially relevant information to the table, encoding them into an input sequence, and using both copy and generation mechanisms in the decoder to balance relevance and readability of the generated title. We validate this approach with human evaluation on sample web tables and report that while sequence models with only a copy mechanism or only a generation mechanism are easily outperformed by simple selection-based baselines, the model with both capabilities performs the best, approaching the quality of crowdsourced titles while training on fewer than ten thousand examples. To the best of our knowledge, the proposed technique is the first to consider text-generation methods for table titles, and establishes a new state of the art. Braden Hancock, Hongrae Lee, Cong Yu 0001 |
WWW | 3 |
| 2019 | AggChecker: A Fact-Checking System for Text Summaries of Relational Data SetsabstractWe demonstrate AggChecker, a novel tool for verifying textual summaries of relational data sets. The system automatically verifies natural language claims about numerical aggregates against the underlying raw data. The system incorporates a combination of natural language processing, information retrieval, machine learning, and efficient query processing strategies. Each claim is translated into a semantically equivalent SQL query and evaluated against the database. Our primary goal is analogous to that of a spell-checker: to identify erroneous claims and provide guidance in correcting them. In this demonstration, we show that our system enables users to verify text summaries much more efficiently than a standard SQL interface. Saehan Jo, Immanuel Trummer, Weicheng Yu, Xuezhi Wang 0002, Cong Yu 0001, Daniel Liu, Niyati Mehta |
Proc. VLDB Endow. | 5 |
| 2019 | Mining an "Anti-Knowledge Base" from Wikipedia Updates with Applications to Fact Checking and BeyondabstractWe introduce the problem of anti-knowledge mining. Our goal is to create an "anti-knowledge base" that contains factual mistakes. The resulting data can be used for analysis, training, and benchmarking in the research domain of automated fact checking. Prior data sets feature manually generated fact checks of famous misclaims. Instead, we focus on the long tail of factual mistakes made by Web authors, ranging from erroneous sports results to incorrect capitals. We mine mistakes automatically, by an unsupervised approach, from Wikipedia updates that correct factual mistakes. Identifying such updates (only a small fraction of the total number of updates) is one of the primary challenges. We mine anti-knowledge by a multi-step pipeline. First, we filter out candidate updates via several simple heuristics. Next, we correlate Wikipedia updates with other statements made on the Web. Using claim occurrence frequencies as input to a probabilistic model, we infer the likelihood of corrections via an iterative expectation-maximization approach. Finally, we extract mistakes in the form of subject-predicate-object triples and rank them according to several criteria. Our end result is a data set containing over 110,000 ranked mistakes with a precision of 85% in the top 1% and a precision of over 60% in the top 25%. We demonstrate that baselines achieve significantly lower precision. Also, we exploit our data to verify several hypothesis on why users make mistakes. We finally show that the AKB can be used to find mistakes on the entire Web. Georgios Karagiannis, Immanuel Trummer, Saehan Jo, Shubham Khandelwal, Xuezhi Wang 0002, Cong Yu 0001 |
Proc. VLDB Endow. | 6 |
| 2018 | Investigating Rumor News Using Agreement-Aware SearchabstractRecent years have witnessed a widespread increase of rumor news generated by humans and machines. Therefore, tools for investigating rumor news have become an urgent necessity. One useful function of such tools is to see ways a specific topic or event is represented by presenting different points of view from multiple sources. In this paper, we propose Maester, a novel agreement-aware search framework for investigating rumor news. Given an investigative question, Maester will retrieve related articles to that question, assign and display top articles from agree, disagree, and discuss categories to users. Splitting the results into these three categories provides the user a holistic view towards the investigative question. We build Maester based on the following two key observations: (1) relatedness can commonly be determined by keywords and entities occurring in both questions and articles, and (2) the level of agreement between the investigative question and the related news article can often be decided by a few key sentences. Accordingly, we use gradient boosting tree models with keyword/entity matching features for relatedness detection, and leverage recurrent neural network to infer the level of agreement. Our experiments on the Fake News Challenge (FNC) dataset demonstrate up to an order of magnitude improvement of Maester over the original FNC winning solution, for agreement-aware search. Jingbo Shang, Tianhang Sun, Xingbang Liu, Anja Gruenheid, Flip Korn, Ádám Dániel Lelkes, Cong Yu 0001, Jiawei Han 0001 |
CIKM | 8 |
| 2018 | Ten Years of WebTablesabstractIn 2008, we wrote about WebTables, an effort to exploit the large and diverse set of structured databases casually published online in the form of HTML tables. The past decade has seen a flurry of research and commercial activities around the WebTables project itself, as well as the broad topic of informal online structured data. In this paper, we 1 will review the WebTables project, and try to place it in the broader context of the decade of work that followed. We will also show how the progress over the past ten years sets up an exciting agenda for the future, and will draw upon many corners of the data management community. Michael J. Cafarella, Alon Y. Halevy, Hongrae Lee, Jayant Madhavan, Cong Yu 0001, Daisy Zhe Wang, Eugene Wu 0002 |
Proc. VLDB Endow. | 5 |
| 2017 | Computational Fact Checking through Query PerturbationsabstractOur media is saturated with claims of “facts” made from data. Database research has in the past focused on how to answer queries, but has not devoted much attention to discerning more subtle qualities of the resulting claims, for example, is a claim “cherry-picking”? This article proposes a framework that models claims based on structured data as parameterized queries. Intuitively, with its choice of the parameter setting, a claim presents a particular (and potentially biased) view of the underlying data. A key insight is that we can learn a lot about a claim by “perturbing” its parameters and seeing how its conclusion changes. For example, a claim is not robust if small perturbations to its parameters can change its conclusions significantly. This framework allows us to formulate practical fact-checking tasks—reverse-engineering vague claims, and countering questionable claims—as computational problems. Along with the modeling framework, we develop an algorithmic framework that enables efficient instantiations of “meta” algorithms by supplying appropriate algorithmic building blocks. We present real-world examples and experiments that demonstrate the power of our model, efficiency of our algorithms, and usefulness of their results. You Wu 0001, Pankaj K. Agarwal, Chengkai Li 0001, Jun Yang 0001, Cong Yu 0001 |
ACM Trans. Database Syst. | 5 |
| 2016 | LONLIES: Estimating Property Values for Long Tail EntitiesabstractWeb search engines often retrieve answers for queries about popular entities from a growing knowledge base that is populated by a continuous information extraction process. However, less popular entities are not frequently mentioned on the web and are generally interesting to fewer users; these entities reside on the long tail of information. Traditional knowledge base construction techniques that rely on the high frequency of entity mentions to extract accurate facts about these mentions have little success with entities that have low textual support. We present Lonlies, a system for estimating property values of long tail entities by leveraging their relationships to head topics and entities. We demonstrate (1) how Lonlies builds communities of entities that are relevant to a long tail entity utilizing a text corpus and a knowledge base; (2) how Lonlies determines which communities to use in the estimation process; (3) how we aggregate estimates from community entities to produce final estimates, and (4) how users interact with Lonlies to provide feedback to improve the final estimation results. Mina H. Farid, Ihab F. Ilyas, Steven Euijong Whang, Cong Yu 0001 |
SIGIR | 4 |
| 2016 | Knowledge Exploration using Tables on the WebabstractThe increasing popularity of mobile device usage has ushered in many features in modern search engines that help users with various information needs. One of those needs is Knowledge Exploration, where related documents are returned in response to a user query, either directly through right-hand side knowledge panels or indirectly through navigable sections underneath individual search results. Existing knowledge exploration features have relied on a combination of Knowledge Bases and query logs. In this paper, we propose Knowledge Carousels of two modalities, namely sideways and downwards, that facilitate exploration of IS-A and HAS-A relationships, respectively, with regard to an entity-seeking query, based on leveraging the large corpus of tables on the Web. This brings many technical challenges, including associating correct carousels with the search entity, selecting the best carousel from the candidates, and finding titles that best describe the carousel. We describe how we address these challenges and also experimentally demonstrate through user studies that our approach produces better result sets than baseline approaches. Fernando Seabra Chirigati, Flip Korn, You Wu 0001, Cong Yu 0001, Hao Zhang 0010 |
Proc. VLDB Endow. | 5 |
| 2015 | Applying WebTables in Practice
Sreeram Balakrishnan, Alon Y. Halevy, Boulos Harb, Hongrae Lee, Jayant Madhavan, Afshin Rostamizadeh, Warren Shen, Kenneth Wilder, Fei Wu 0003, Cong Yu 0001 |
CIDR | 10 |
| 2015 | Inferencing in information extraction: Techniques and applicationsabstractInformation extraction at Web scale has become one of the most important research topics in data management since major commercial search engines started incorporating knowledge in their search results a couple of years ago [1]. Users increasingly expect structured knowledge as answers to their search needs. Using Bing as an example, the result page for “Lionel Messi” is full of structured knowledge facts, such as his birthday and awards. The research efforts towards improving the accuracy and coverage of such knowledge bases have led to significant advances in Information Extraction techniques [2], [3]. As the initial challenge of accurately extracting facts for popular entities are being addressed, more difficult challenges have emerged such as extending knowledge coverage to long tail entities and domains, understanding interestingness and usefulness of facts within a given context, and addressing information-seeking needs more directly and accurately. In this tutorial, we will survey the recent research efforts and provide an introduction to the techniques that address those challenges, and the applications that benefit from the adoption of those techniques. In particular, this tutorial will focus on a variety of techniques that can be broadly viewed as knowledge inferencing, i.e., combining multiple data sources and extraction techniques to verify existing knowledge and derive new knowledge. More specifically, we focus on four main categories of inferencing techniques: 1) deep natural language processing using machine learning techniques, 2) data cleaning using integrity constraints, 3) large-scale probabilistic reasoning, and 4) leveraging human expertise for domain knowledge extraction. Denilson Barbosa 0001, Haixun Wang, Cong Yu 0001 |
ICDE | 3 |
| 2015 | Efficient Evaluation of Object-Centric Exploration Queries for VisualizationabstractThe most effective way to explore data is through visualizing the results of exploration queries. For example, an exploration query could be an aggregate of some measures over time intervals, and a pattern or abnormality can be discovered through a time series plot of the query results. In this paper, we examine a special kind of exploration query, namely object-centric exploration query. Common examples include claims made about athletes in sports databases, such as "it is newsworthy that LeBron James has scored 35 or more points in nine consecutive games." We focus on one common type of visualization, i.e., 2d scatter plot with heatmap. Namely, we consider exploration queries whose results can be plotted on a two-dimensional space, possibly with colors indicating object densities in regions. While we model results as pairs of numbers, the types of the queries are limited only by the users' imagination. In the LeBron James example above, the two dimensions are minimum points scored per game and number of consecutive games, respectively. It is easy to find other equally interesting dimensions, such as minimum rebounds per game or number of playoff games. We formalize this problem and propose an efficient, interactive-speed algorithm that takes a user-provided exploration query (which can be a blackbox function) and produces an approximate visualization that preserves the two most important visual properties: the outliers and the overall distribution of all result points. You Wu 0001, Boulos Harb, Jun Yang 0001, Cong Yu 0001 |
Proc. VLDB Endow. | 4 |
| 2015 | Active learning in keyword search-based data integration
Zhepeng Yan, Zachary G. Ives, Partha P. Talukdar, Cong Yu 0001 |
VLDB J. | 5 |
| 2014 | Near neighbor joinabstractAn increasing number of Web applications such as friends recommendation depend on the ability to join objects at scale. The traditional approach taken is nearest neighbor join (also called similarity join), whose goal is to find, based on a given join function, the closest set of objects or all the objects within a distance threshold to each object in the input. The scalability of techniques utilizing this approach often depends on the characteristics of the objects and the join function. However, many real-world join functions are intricately engineered and constantly evolving, which makes the design of white-box methods that rely on understanding the join function impractical. Finding a technique that can join extremely large number of objects with complex join functions has always been a tough challenge. In this paper, we propose a practical alternative approach called near neighbor join that, although does not find the closest neighbors, finds close neighbors, and can do so at extremely large scale when the join functions are complex. In particular, we design and implement a super-scalable system we name SAJ that is capable of best-effort joining of billions of objects for complex functions. Extensive experimental analysis over real-world large datasets shows that SAJ is scalable and generates good results. Herald Kllapi, Boulos Harb, Cong Yu 0001 |
ICDE | 3 |
| 2014 | Exploiting group recommendation functions for flexible preferencesabstractWe examine the problem of enabling the flexibility of updating one's preferences in group recommendation. In our setting, any group member can provide a vector of preferences that, in addition to past preferences and other group members' preferences, will be accounted for in computing group recommendation. This functionality is essential in many group recommendation applications, such as travel planning, online games, book clubs, or strategic voting, as it has been previously shown that user preferences may vary depending on mood, context, and company (i.e., other people in the group). Preferences are enforced in an feedback box that replaces preferences provided by the users by a potentially different feedback vector that is better suited for maximizing the individual satisfaction when computing the group recommendation. The feedback box interacts with a traditional recommendation box that implements a group consensus semantics in the form of Aggregated Voting or Least Misery, two popular aggregation functions for group recommendation. We develop efficient algorithms to compute robust group recommendations that are appropriate in situations where users have changing preferences. Our extensive empirical study on real world data-sets validates our findings. Senjuti Basu Roy, Saravanan Thirumuruganathan, Sihem Amer-Yahia, Gautam Das 0001, Cong Yu 0001 |
ICDE | 5 |
| 2014 | Incremental discovery of prominent situational factsabstractWe study the novel problem of finding new, prominent situational facts, which are emerging statements about objects that stand out within certain contexts. Many such facts are newsworthy-e.g., an athlete's outstanding performance in a game, or a viral video's impressive popularity. Effective and efficient identification of these facts assists journalists in reporting, one of the main goals of computational journalism. Technically, we consider an ever-growing table of objects with dimension and measure attributes. A situational fact is a “contextual” skyline tuple that stands out against historical tuples in a context, specified by a conjunctive constraint involving dimension attributes, when a set of measure attributes are compared. New tuples are constantly added to the table, reflecting events happening in the real world. Our goal is to discover constraint-measure pairs that qualify a new tuple as a contextual skyline tuple, and discover them quickly before the event becomes yesterday's news. A brute-force approach requires exhaustive comparison with every tuple, under every constraint, and in every measure subspace. We design algorithms in response to these challenges using three corresponding ideas-tuple reduction, constraint pruning, and sharing computation across measure subspaces. We also adopt a simple prominence measure to rank the discovered facts when they are numerous. Experiments over two real datasets validate the effectiveness and efficiency of our techniques. Afroza Sultana, Naeemul Hassan, Chengkai Li 0001, Jun Yang 0001, Cong Yu 0001 |
ICDE | 5 |
| 2014 | On social event organizationabstractOnline platforms, such as Meetup and Plancast, have recently become popular for planning gatherings and event organization. However, there is a surprising lack of studies on how to effectively and efficiently organize social events for a large group of people through such platforms. In this paper, we study the key computational problem involved in organization of social events, to our best knowledge, for the first time. Keqian Li, Wei Lu 0002, Smriti Bhagat, Laks V. S. Lakshmanan, Cong Yu 0001 |
KDD | 5 |
| 2014 | iCheck: computationally combating "lies, d-ned lies, and statistics"abstractAre you fed up with "lies, d---ned lies, and statistics" made up from data in our media? For claims based on structured data, we present a system to automatically assess the quality of claims (beyond their correctness) and counter misleading claims that cherry-pick data to advance their conclusions. The key insight is to model such claims as parameterized queries and consider how parameter perturbations affect their results. We demonstrate our system on claims drawn from U.S. congressional voting records, sports statistics, and publication records of database researchers. You Wu 0001, Brett Walenz, Peggy Li, Andrew Shim, Emre Sonmez, Pankaj K. Agarwal, Chengkai Li 0001, Jun Yang 0001, Cong Yu 0001 |
SIGMOD Conference | 9 |
| 2014 | Data In, Fact Out: Automated Monitoring of Facts by FactWatcherabstractTowards computational journalism, we present FactWatcher, a system that helps journalists identify data-backed, attention-seizing facts which serve as leads to news stories. FactWatcher discovers three types of facts, including situational facts, one-of-the-few facts, and prominent streaks, through a unified suite of data model, algorithm framework, and fact ranking measure. Given an append-only database, upon the arrival of a new tuple, FactWatcher monitors if the tuple triggers any new facts. Its algorithms efficiently search for facts without exhaustively testing all possible ones. Furthermore, FactWatcher provides multiple features in striving for an end-to-end system, including fact ranking, fact-to-statement translation and keyword-based fact search. Naeemul Hassan, Afroza Sultana, You Wu 0001, Gensheng Zhang, Chengkai Li 0001, Jun Yang 0001, Cong Yu 0001 |
Proc. VLDB Endow. | 7 |
| 2014 | Toward Computational Fact-CheckingabstractOur news are saturated with claims of "facts" made from data. Database research has in the past focused on how to answer queries, but has not devoted much attention to discerning more subtle qualities of the resulting claims, e.g., is a claim "cherry-picking"? This paper proposes a framework that models claims based on structured data as parameterized queries. A key insight is that we can learn a lot about a claim by perturbing its parameters and seeing how its conclusion changes. This framework lets us formulate practical fact-checking tasks---reverse-engineering (often intentionally) vague claims, and countering questionable claims---as computational problems. Along with the modeling framework, we develop an algorithmic framework that enables efficient instantiations of "meta" algorithms by supplying appropriate algorithmic building blocks. We present real-world examples and experiments that demonstrate the power of our model, efficiency of our algorithms, and usefulness of their results. You Wu 0001, Pankaj K. Agarwal, Chengkai Li 0001, Jun Yang 0001, Cong Yu 0001 |
Proc. VLDB Endow. | 5 |
| 2014 | Front Matter
Li Xiong 0001, Cong Yu 0001 |
Proc. VLDB Endow. | 2 |
| 2014 | An expressive framework and efficient algorithms for the analysis of collaborative tagging
Mahashweta Das, Saravanan Thirumuruganathan, Sihem Amer-Yahia, Gautam Das 0001, Cong Yu 0001 |
VLDB J. | 5 |
| 2013 | CloudDB 2013: fifth international workshop on cloud data managementabstractThe fifth ACM international workshop on cloud data management is held in San Francisco, California, USA on October 28, 2013 and co-located with the ACM 22nd Conference on Information and Knowledge Management (CIKM). The main objective of the workshop is to address the challenges of large scale data management based on the cloud computing infrastructure. The workshop brings together researchers and practitioners from cloud computing, distributed storage, query processing, parallel algorithms, data mining, and system analysis, all attendees share common research interests in maximizing performance, reducing cost of cloud data management and enlarging the scale of their endeavors. We have constructed an exciting program of four refereed papers and an invited keynote talk that will give participants a full dose of emerging research. Feifei Li 0001, Xiaofeng Meng 0001, Fusheng Wang 0001, Cong Yu 0001 |
CIKM | 4 |
| 2013 | Recent progress towards an ecosystem of structured data on the WebabstractGoogle Fusion Tables aims to support an ecosystem of structured data on the Web by providing a tool for managing and visualizing data on the one hand, and for searching and exploring for data on the other. This paper describes a few recent developments in our efforts to further the ecosystem. Nitin Gupta 0003, Alon Y. Halevy, Boulos Harb, Heidi Lam, Hongrae Lee, Jayant Madhavan, Fei Wu 0003, Cong Yu 0001 |
ICDE | 8 |
| 2013 | Shallow Information Extraction for the knowledge WebabstractA new breed of Information Extraction tools has become popular and shown to be very effective in building massive-scale knowledge bases that fuel applications such as question answering and semantic search. These approaches rely on Web-scale probabilistic models populated through shallow language processing of the text, pre-existing knowledge, and structured data already on the Web. This tutorial provides an introduction to these techniques, starting from the foundations of information extraction, and covering some of its key applications. Denilson Barbosa 0001, Haixun Wang, Cong Yu 0001 |
ICDE | 3 |
| 2013 | Synthesizing Union Tables from the Web
Alon Y. Halevy, Fei Wu 0003, Cong Yu 0001 |
IJCAI | 4 |
| 2013 | Scalable Column Concept Determination for Web Tables Using Large Knowledge BasesabstractTabular data on the Web has become a rich source of structured data that is useful for ordinary users to explore. Due to its potential, tables on the Web have recently attracted a number of studies with the goals of understanding the semantics of those Web tables and providing effective search and exploration mechanisms over them. An important part of table understanding and search is column concept determination, i.e., identifying the most appropriate concepts associated with the columns of the tables. The problem becomes especially challenging with the availability of increasingly rich knowledge bases that contain hundreds of millions of entities. In this paper, we focus on an important instantiation of the column concept determination problem, namely, the concepts of a column are determined by fuzzy matching its cell values to the entities within a large knowledge base. We provide an efficient and scalable MapReduce-based solution that is scalable to both the number of tables and the size of the knowledge base and propose two novel techniques: knowledge concept aggregation and knowledge entity partition. We prove that both the problem of finding the optimal aggregation strategy and that of finding the optimal partition strategy are NP-hard, and propose efficient heuristic techniques by leveraging the hierarchy of the knowledge base. Experimental results on real-world datasets show that our method achieves high annotation quality and performance, and scales well. Dong Deng 0001, Guoliang Li 0001, Jian Li 0015, Cong Yu 0001 |
Proc. VLDB Endow. | 5 |
| 2013 | Front Matter
Cong Yu 0001 |
Proc. VLDB Endow. | 2 |
| 2013 | Actively Soliciting Feedback for Query Answers in Keyword Search-Based Data IntegrationabstractThe problem of scaling up data integration, such that new sources can be quickly utilized as they are discovered, remains elusive: global schemas for integrated data are difficult to develop and expand, and schema and record matching techniques are limited by the fact that data and metadata are often under-specified and must be disambiguated by data experts. One promising approach is to avoid using a global schema, and instead to develop keyword search-based data integration--where the system lazily discovers associations enabling it to join together matches to keywords, and return ranked results. The user is expected to understand the data domain and provide feedback about answers' quality. The system generalizes such feedback to learn how to correctly integrate data. A major open challenge is that under this model, the user only sees and offers feedback on a few "top- k " results: this result set must be carefully selected to include answers of high relevance and answers that are highly informative when feedback is given on them. Existing systems merely focus on predicting relevance, by composing the scores of various schema and record matching algorithms. In this paper we show how to predict the uncertainty associated with a query result's score, as well as how informative feedback is on a given result. We build upon these foundations to develop an active learning approach to keyword search-based data integration, and we validate the effectiveness of our solution over real data from several very different domains. Zhepeng Yan, Zachary G. Ives, Partha P. Talukdar, Cong Yu 0001 |
Proc. VLDB Endow. | 5 |
| 2012 | On "one of the few" objectsabstractObjects with multiple numeric attributes can be compared within any "subspace" (subset of attributes). In applications such as computational journalism, users are interested in claims of the form: Karl Malone is one of the only two players in NBA history with at least 25,000 points, 12,000 rebounds, and 5,000 assists in one's career. One challenge in identifying such "one-of-the-k" claims (k = 2 above) is ensuring their "interestingness". A small k is not a good indicator for interestingness, as one can often make such claims for many objects by increasing the dimensionality of the subspace considered. We propose a uniqueness-based interestingness measure for one-of-the-few claims that is intuitive for non-technical users, and we design algorithms for finding all interesting claims (across all subspaces) from a dataset. Sometimes, users are interested primarily in the objects appearing in these claims. Building on our notion of interesting claims, we propose a scheme for ranking objects and an algorithm for computing the top-ranked objects. Using real-world datasets, we evaluate the efficiency of our algorithms as well as the advantage of our object-ranking scheme over popular methods such as Kemeny optimal rank aggregation and weighted-sum ranking. You Wu 0001, Pankaj K. Agarwal, Chengkai Li 0001, Jun Yang 0001, Cong Yu 0001 |
KDD | 5 |
| 2012 | Finding related tablesabstractWe consider the problem of finding related tables in a large corpus of heterogenous tables. Detecting related tables provides users a powerful tool for enhancing their tables with additional data and enables effective reuse of available public data. Our first contribution is a framework that captures several types of relatedness, including tables that are candidates for joins and tables that are candidates for union. Our second contribution is a set of algorithms for detecting related tables that can be either unioned or joined. We describe a set of experiments that demonstrate that our algorithms produce highly related tables. We also show that we can often improve the results of table search by pulling up tables that are ranked much lower based on their relatedness to top-ranked tables. Finally, we describe how to scale up our algorithms and show the results of running it on a corpus of over a million tables extracted from Wikipedia. Anish Das Sarma, Lujun Fang, Nitin Gupta 0003, Alon Y. Halevy, Hongrae Lee, Fei Wu 0003, Reynold Xin, Cong Yu 0001 |
SIGMOD Conference | 8 |
| 2012 | Who Tags What? An Analysis FrameworkabstractThe rise of Web 2.0 is signaled by sites such as Flickr, del.icio.us, and YouTube, and social tagging is essential to their success. A typical tagging action involves three components, user, item (e.g., photos in Flickr), and tags (i.e., words or phrases). Analyzing how tags are assigned by certain users to certain items has important implications in helping users search for desired information. In this paper, we explore common analysis tasks and propose a dual mining framework for social tagging behavior mining. This framework is centered around two opposing measures, similarity and diversity , being applied to one or more tagging components, and therefore enables a wide range of analysis scenarios such as characterizing similar users tagging diverse items with similar tags, or diverse users tagging similar items with diverse tags, etc. By adopting different concrete measures for similarity and diversity in the framework, we show that a wide range of concrete analysis problems can be defined and they are NP-Complete in general. We design efficient algorithms for solving many of those problems and demonstrate, through comprehensive experiments over real data, that our algorithms significantly out-perform the exact brute-force approach without compromising analysis result quality. Mahashweta Das, Saravanan Thirumuruganathan, Sihem Amer-Yahia, Gautam Das 0001, Cong Yu 0001 |
Proc. VLDB Endow. | 5 |
| 2012 | MapRat: Meaningful Explanation, Interactive Exploration and Geo-Visualization of Collaborative RatingsabstractCollaborative rating sites such as IMDB and Yelp have become rich resources that users consult to form judgments about and choose from among competing items. Most of these sites either provide a plethora of information for users to interpret all by themselves or a simple overall aggregate information. Such aggregates (e.g., average rating over all users who have rated an item, aggregates along pre-defined dimensions, etc.) can not help a user quickly decide the desirability of an item. In this paper, we build a system MapRat that allows a user to explore multiple carefully chosen aggregate analytic details over a set of user demographics that meaningfully explain the ratings associated with item(s) of interest. MapRat allows a user to systematically explore, visualize and understand user rating patterns of input item(s) so as to make an informed decision quickly. In the demo, participants are invited to explore collaborative movie ratings for popular movies. Saravanan Thirumuruganathan, Mahashweta Das, Shrikant Desai, Sihem Amer-Yahia, Gautam Das 0001, Cong Yu 0001 |
Proc. VLDB Endow. | 6 |
| 2012 | Entity-Relationship Queries over WikipediaabstractWikipedia is the largest user-generated knowledge base. We propose a structured query mechanism, entity-relationship query , for searching entities in the Wikipedia corpus by their properties and interrelationships. An entity-relationship query consists of multiple predicates on desired entities. The semantics of each predicate is specified with keywords. Entity-relationship query searches entities directly over text instead of preextracted structured data stores. This characteristic brings two benefits: (1) Query semantics can be intuitively expressed by keywords; (2) It only requires rudimentary entity annotation, which is simpler than explicitly extracting and reasoning about complex semantic information before query-time. We present a ranking framework for general entity-relationship queries and a position-based Bounded Cumulative Model (BCM) for accurate ranking of query answers. We also explore various weighting schemes for further improving the accuracy of BCM. We test our ideas on a 2008 version of Wikipedia using a collection of 45 queries pooled from INEX entity ranking track and our own crafted queries. Experiments show that the ranking and weighting schemes are both effective, particularly on multipredicate queries. Chengkai Li 0001, Cong Yu 0001 |
ACM Trans. Intell. Syst. Technol. | 3 |
| 2012 | Data Cube Materialization and Mining over MapReduceabstractComputing interesting measures for data cubes and subsequent mining of interesting cube groups over massive data sets are critical for many important analyses done in the real world. Previous studies have focused on algebraic measures such as SUM that are amenable to parallel computation and can easily benefit from the recent advancement of parallel computing infrastructure such as MapReduce. Dealing with holistic measures such as TOP-K, however, is nontrivial. In this paper, we detail real-world challenges in cube materialization and mining tasks on web-scale data sets. Specifically, we identify an important subset of holistic measures and introduce MR-Cube, a MapReduce-based framework for efficient cube computation and identification of interesting cube groups on holistic measures. We provide extensive experimental analyses over both real and synthetic data. We demonstrate that, unlike existing techniques which cannot scale to the 100 million tuple mark for our data sets, MR-Cube successfully and efficiently computes cubes with holistic measures over billion-tuple data sets. Arnab Nandi 0001, Cong Yu 0001, Philip Bohannon, Raghu Ramakrishnan 0001 |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2011 | Computational Journalism: A Call to Arms to Database Researchers
Sarah Cohen, Chengkai Li 0001, Jun Yang 0001, Cong Yu 0001 |
CIDR | 4 |
| 2011 | Distributed cube materialization on holistic measuresabstractCube computation over massive datasets is critical for many important analyses done in the real world. Unlike commonly studied algebraic measures such as SUM that are amenable to parallel computation, efficient cube computation of holistic measures such as TOP-K is non-trivial and often impossible with current methods. In this paper we detail real-world challenges in cube materialization tasks on Web-scale datasets. Specifically, we identify an important subset of holistic measures and introduce MR-Cube, a MapReduce based framework for efficient cube computation on these measures. We provide extensive experimental analyses over both real and synthetic data. We demonstrate that, unlike existing techniques which cannot scale to the 100 million tuple mark for our datasets, MR-Cube successfully and efficiently computes cubes with holistic measures over billion-tuple datasets. Arnab Nandi 0001, Cong Yu 0001, Philip Bohannon, Raghu Ramakrishnan 0001 |
ICDE | 2 |
| 2011 | Interactive itinerary planningabstractPlanning an itinerary when traveling to a city involves substantial effort in choosing Points-of-Interest (POIs), deciding in which order to visit them, and accounting for the time it takes to visit each POI and transit between them. Several online services address different aspects of itinerary planning but none of them provides an interactive interface where users give feedbacks and iteratively construct their itineraries based on personal interests and time budget. In this paper, we formalize interactive itinerary planning as an iterative process where, at each step: (1) the user provides feedback on POIs selected by the system, (2) the system recommends the best itineraries based on all feedback so far, and (3) the system further selects a new set of POIs, with optimal utility, to solicit feedback for, at the next step. This iterative process stops when the user is satisfied with the recommended itinerary. We show that computing an itinerary is NP-complete even for simple itinerary scoring functions, and that POI selection is NP-complete. We develop heuristics and optimizations for a specific case where the score of an itinerary is proportional to the number of desired POIs it contains. Our extensive experiments show that our algorithms are efficient and return high quality itineraries. Senjuti Basu Roy, Gautam Das 0001, Sihem Amer-Yahia, Cong Yu 0001 |
ICDE | 4 |
| 2011 | Dynamic relationship and event discoveryabstractThis paper studies the problem of dynamic relationship and event discovery. A large body of previous work on relation extraction focuses on discovering predefined and static relationships between entities. In contrast, we aim to identify temporally defined (e.g., co-bursting) relationships that are not predefined by an existing schema, and we identify the underlying time constrained events that lead to these relationships. The key challenges in identifying such events include discovering and verifying dynamic connections among entities, and consolidating binary dynamic connections into events consisting of a set of entities that are connected at a given time period. We formalize this problem and introduce an efficient end-to-end pipeline as a solution. In particular, we introduce two formal notions, global temporal constraint cluster and local temporal constraint cluster, for detecting dynamic events. We further design efficient algorithms for discovering such events from a large graph of dynamic relationships. Finally, detailed experiments on real data show the Anish Das Sarma, Alpa Jain, Cong Yu 0001 |
WSDM | 3 |
| 2011 | MRI: Meaningful Interpretations of Collaborative Ratings
Mahashweta Das, Sihem Amer-Yahia, Gautam Das 0001, Cong Yu 0001 |
Proc. VLDB Endow. | 4 |
| 2011 | REX: Explaining Relationships between Entity PairsabstractKnowledge bases of entities and relations (either constructed manually or automatically) are behind many real world search engines, including those at Yahoo!, Microsoft, and Google. Those knowledge bases can be viewed as graphs with nodes representing entities and edges representing (primary) relationships, and various studies have been conducted on how to leverage them to answer entity seeking queries. Meanwhile, in a complementary direction, analyses over the query logs have enabled researchers to identify entity pairs that are statistically correlated. Such entity relationships are then presented to search users through the "related searches" feature in modern search engines. However, entity relationships thus discovered can often be "puzzling" to the users because why the entities are connected is often indescribable. In this paper, we propose a novel problem calledentity relationship explanation, which seeks to explain why a pair of entities are connected, and solve this challenging problem by integrating the above two complementary approaches, i.e., we leverage the knowledge base to "explain" the connections discovered between entity pairs. More specifically, we presentREX, a system that takes a pair of entities in a given knowledge base as input and efficiently identifies a ranked list of relationship explanations. We formally define relationship explanations and analyze their desirable properties. Furthermore, we design and implement algorithms to efficiently enumerate and rank all relationship explanations based on multiple measures of "interestingness." We perform extensive experiments over real web-scale data gathered from DBpedia and a commercial search engine, demonstrating the efficiency and scalability ofREX. We also perform user studies to corroborate the effectiveness of explanations generated byREX. Lujun Fang, Anish Das Sarma, Cong Yu 0001, Philip Bohannon |
Proc. VLDB Endow. | 3 |
| 2010 | Prioritization of Domain-Specific Web Information ExtractionabstractIt is often desirable to extract structured information from raw web pages for better information browsing, query answering, and pattern mining. many such Information Extraction (IE) technologies are costly and applying them at the web-scale is impractical. In this paper, we propose a novel prioritization approach where candidate pages from the corpus are ordered according to their expected contribution to the extraction results and those with higher estimated potential are extracted earlier. Systems employing this approach can stop the extraction process at any time when the resource gets scarce (i.e., not all pages in the corpus can be processed), without worrying about wasting extraction effort on unimportant pages. More specifically, we define a novel notion to measure the value of extraction results and design various mechanisms for estimating a candidate page’s contribution to this value. We further design and build the Extraction Prioritization (EP) system with efficient scoring and scheduling algorithms, and experimentally demonstrate that EP significantly outperforms the naive approach and is more flexible than the classifier approach. Cong Yu 0001 |
AAAI | 2 |
| 2010 | EntityEngine: answering entity-relationship queries using shallow semanticsabstractWe introduce EntityEngine, a system for answering entity-relationship queries over text. Such queries combine SQL-like structures with IR-style keyword constraints and therefore, can be expressive and flexible in querying about entities and their relationships. EntityEngine consists of various offline and online components, including a position-based ranking model for accurate ranking of query answers and a novel entity-centric index for efficient query evaluation. Chengkai Li 0001, Cong Yu 0001 |
CIKM | 3 |
| 2010 | Constructing and exploring composite itemsabstractNowadays, online shopping has become a daily activity. Web users purchase a variety of items ranging from books to electronics. The large supply of online products calls for sophisticated techniques to help users explore available items. We propose to build composite items which associate a central item with a set of packages, formed by satellite items, and help users explore them. For example, a user shopping for an iPhone (i.e., the Senjuti Basu Roy, Sihem Amer-Yahia, Ashish Chawla, Gautam Das 0001, Cong Yu 0001 |
SIGMOD Conference | 5 |
| 2010 | Popularity-Guided Top-k Extraction of Entity AttributesabstractRecent progress in information extraction technology has enabled a vast array of applications that rely on structured data that is embedded in natural-language text. In particular, the extraction of concepts from the Web---with their desired attributes---is important to provide applications with rich, structured access to information. In this paper, we focus on an important family of concepts, namely, entities (e.g., people or organizations) and their attributes, and study how to efficiently and effectively extract them from Web-accessible text documents. Unfortunately, information extraction over the Web is challenging for both quality and efficiency reasons. Regarding quality, many sources on the Web contain misleading or invalid information; furthermore, extraction systems often return incorrect data. Regarding efficiency, information extraction is a time-consuming process, often involving expensive text-processing steps. We present a top-k extraction processing approach that addresses both the quality and efficiency challenges: for each entity and attribute of interest, we return the top-k values of the attribute for the entity according to a scoring function for extracted attribute values. This scoring function weighs the extraction confidence from individual documents, as well as the "importance" of the documents where the information originates. We define the document importance in terms of entity-specific document "popularity" statistics from a major search engine. Overall, our top-k extraction processing approach manages to identify the top attribute values for the entities of interest efficiently, as we demonstrate with a large-scale experimental evaluation over real-life data. Matthew Solomon, Cong Yu 0001, Luis Gravano |
WebDB | 2 |
| 2010 | Constructing travel itineraries from tagged geo-temporal breadcrumbsabstractVacation planning is a frequent laborious task which requires skilled interaction with a multitude of resources. This paper develops an end-to-end approach for constructing intra-city travel itineraries automatically by tapping a latent source reflecting geo-temporal breadcrumbs left by millions of tourists. In particular, the popular rich media sharing site, Flickr, allows photos to be stamped by the date and time of when they were taken, and be mapped to Points Of Interest (POIs) by latitude-longitude information as well as semantic metadata (e.g., tags) that describe them. Munmun De Choudhury, Moran Feldman, Sihem Amer-Yahia, Nadav Golbandi, Ronny Lempel, Cong Yu 0001 |
WWW | 6 |
| 2010 | Space efficiency in group recommendation
Senjuti Basu Roy, Sihem Amer-Yahia, Ashish Chawla, Gautam Das 0001, Cong Yu 0001 |
VLDB J. | 5 |
| 2009 | SocialScope: Enabling Information Discovery on Social Content Sites
Sihem Amer-Yahia, Laks V. S. Lakshmanan, Cong Yu 0001 |
CIDR | 3 |
| 2009 | It takes variety to make a world: diversification in recommender systemsabstractRecommendations in collaborative tagging sites such as del.icio.us and Yahoo! Movies, are becoming increasingly important, due to the proliferation of general queries on those sites and the ineffectiveness of the traditional search paradigm to address those queries. Regardless of the underlying recommendation strategy, item-based or user-based, one of the key concerns in producing recommendations, is over-specialization, which results in returning items that are too homogeneous. Traditional solutions rely on post-processing returned items to identify those which differ in their attribute values (e.g., genre and actors for movies). Such approaches are not always applicable when intrinsic attributes are not available (e.g., URLs in del.icio.us). In a recent paper [20], we introduced the notion of explanation-based diversity and formalized the diversification problem as a compromise between accuracy and diversity. In this paper, we develop efficient diversification algorithms built upon this notion. The algorithms explore compromises between accuracy and diversity. We demonstrate their efficiency and effectiveness in diversification on two real life data sets: del.icio.us and Yahoo! Movies. Cong Yu 0001, Laks V. S. Lakshmanan, Sihem Amer-Yahia |
EDBT | 1 |
| 2009 | Jelly: A Language for Building Community-Centric Information Exploration ApplicationsabstractSocial content sites, which integrate traditional content sites (e.g., Yahoo! Travel) with social network features, have recently emerged as a significant new trend on the Web. Users on those sites share content and form various communities based on explicit friendships or shared interests. However, the existing information exploration mechanisms rarely leverage the rich community structure. In this work, we aim to unlock the value of social content sites by helping developers specify community-based information exploration strategies in a flexible and declarative way. Our solution makes use of two key notions, topics and communities, in order to identify socially and semantically relevant information for users. Specifically, we propose JELLY as a language for developing community-centric information exploration applications. JELLY provides several primitives which exploit both content and user behavior in social content sites in order to help users explore relevant content. The topic generation primitive is used to extract topics from tags. The community extraction primitive enables building different user communities. The information discovery primitive helps customize content relevance by combining a userpsilas query and profile, as well as insights from related communities. Finally, the information explanation primitive offers valuable social provenance to help users better understand the returned content. We describe JELLYpsilas data model and language, and its application to building a system for finding socially relevant travel destinations in Yahoo! Travel. Sihem Amer-Yahia, Cong Yu 0001 |
ICDE | 3 |
| 2009 | Recommendation Diversification Using ExplanationsabstractWe introduce the novel notion ofexplanation-baseddiversificationto address the well-known problem of over- specialization in item recommendations.Over-specializationin recommender systems leads to result sets with items that are too similar to one another, thus reducing the diversity of results and limiting user choices. Traditionally, the problem is addressed throughattribute-baseddiversification-grouping items in the result set that share many common attributes (e.g., genre for movies) and selecting only a limited number of items from each group. It is, however, not always applicable, especially for social content recommendations. For example, attributes may not be available as in the case of recommending URLs for users of del.icio.us. Explanation-based diversification provides a novel and complementary alternative-it leverages thereasonforwhichaparticularitemisbeingrecommended(i.e., explanation)-for diversifying the results, without the need to access the attributes of the items. In this paper, we formally define the problem ofexplanation-baseddiversificationand, without going into the details of the actual diversification process, demonstrate its effectiveness on a real world data set, Yahoo! Movies. Cong Yu 0001, Laks V. S. Lakshmanan, Sihem Amer-Yahia |
ICDE | 1 |
| 2009 | Getting recommender systems to think outside the boxabstractWe examine the case of over-specialization in recommender systems, which results from returning items that are too similar to those previously rated by the user. We propose Outside-The-Box (otb) recommendation, which takes some risk to help users make fresh discoveries, while maintaining high relevance. The proposed formalization relies on item regions and attempts to identify regions that are under-exposed to the user. We develop a recommendation algorithm which achieves a compromise between relevance and risk to find otb items. We evaluate this approach on the MovieLens data set and compare our otb recommendations against conventional recommendation strategies. Zeinab Abbassi, Sihem Amer-Yahia, Laks V. S. Lakshmanan, Sergei Vassilvitskii, Cong Yu 0001 |
RecSys | 5 |
| 2009 | Building community-centric information exploration applications on social content sitesabstractSocial content sites [4], which integrate traditional content sites with social networking features, have recently emerged as an exciting new trend on the Web. Users on those sites share content and form various communities based on explicit friendship, shared interest and common user properties. Recently, we proposed SOCIALSCOPE, a three-layered architecture to address the information management challenges in social content sites. In this paper, we focus on the information discovery and the information presentation layers, and describe how our previously proposed language, Jelly [3], is supported in SOCIALSCOPE to build community-centric information exploration applications on social content sites. Sihem Amer-Yahia, Cong Yu 0001 |
SIGMOD Conference | 3 |
| 2009 | Group Recommendation: Semantics and EfficiencyabstractWe study the problem of group recommendation. Recommendation is an important information exploration paradigm that retrieves interesting items for users based on their profiles and past activities. Single user recommendation has received significant attention in the past due to its extensive use in Amazon and Netflix. How to recommend to a group of users who may or may not share similar tastes, however, is still an open problem. The need for group recommendation arises in many scenarios: a movie for friends to watch together, a travel destination for a family to spend a holiday break, and a good restaurant for colleagues to have a working lunch. Intuitively, items that are ideal for recommendation to a group may be quite different from those for individual members. In this paper, we analyze the desiderata of group recommendation and propose a formal semantics that accounts for both item relevance to a group and disagreements among group members. We design and implement algorithms for efficiently computing group recommendations. We evaluate our group recommendation method through a comprehensive user study conducted on Amazon Mechanical Turk and demonstrate that incorporating disagreements is critical to the effectiveness of group recommendation. We further evaluate the efficiency and scalability of our algorithms on the MovieLens data set with 10M ratings. Sihem Amer-Yahia, Senjuti Basu Roy, Ashish Chawla, Gautam Das 0001, Cong Yu 0001 |
Proc. VLDB Endow. | 5 |
| 2009 | Data integration with uncertainty
Xin Dong 0001, Alon Y. Halevy, Cong Yu 0001 |
VLDB J. | 3 |
| 2008 | Yoopick: A Combinatorial Sports Prediction Market
Sharad Goel, David M. Pennock, Daniel M. Reeves, Cong Yu 0001 |
AAAI | 4 |
| 2008 | From del.icio.us to x.qui.site: recommendations in social tagging sitesabstractWe present X.QUI.SITE, a scalable system for managing recommendations for social tagging sites like del.icio.us. seamlessly incorporates various user behaviors into the recommendations and aims to recommend not only items of interest, but also other relevant information like interesting people and/or topics. Explanations are also provided so that users can obtain a better understanding of the recommendations and decide which recommendations to pursue further. We discuss the technical challenges involved in characterizing different user behaviors and in efficiently computing recommendation explanations. Sihem Amer-Yahia, Alban Galland, Julia Stoyanovich, Cong Yu 0001 |
SIGMOD Conference | 4 |
| 2008 | Enabling Schema-Free XQuery with meaningful query focus
Yunyao Li 0001, Cong Yu 0001, H. V. Jagadish |
VLDB J. | 2 |
| 2008 | XML schema refinement through redundancy detection and normalization
Cong Yu 0001, H. V. Jagadish |
VLDB J. | 1 |
| 2007 | Web-Scale Data Integration: You can afford to Pay as You Go
Jayant Madhavan, Shirley Cohen, Xin Dong 0001, Alon Y. Halevy, Shawn R. Jeffery, David Ko, Cong Yu 0001 |
CIDR | 7 |
| 2007 | Making database systems usableabstractDatabase researchers have striven to improve the capability of a database in terms of both performance and functionality. We assert that the usability of a database is as important as its capability. In this paper, we study why database systems today are so difficult to use. We identify a set of five pain points and propose a research agenda to address these. In particular, we introduce a presentation data model and recommend direct data manipulation with a schema later approach. We also stress the importance of provenance and of consistency across presentation models. H. V. Jagadish, Adriane Chapman, Aaron Elkiss, Magesh Jayapandian, Yunyao Li 0001, Arnab Nandi 0001, Cong Yu 0001 |
SIGMOD Conference | 7 |
| 2007 | Data Integration with Uncertainty
Xin Dong 0001, Alon Y. Halevy, Cong Yu 0001 |
VLDB | 3 |
| 2007 | Querying Complex Structured Databases
Cong Yu 0001, H. V. Jagadish |
VLDB | 1 |
| 2006 | Efficient Discovery of XML Data Redundancies
Cong Yu 0001, H. V. Jagadish |
VLDB | 1 |
| 2006 | Schema Summarization
Cong Yu 0001, H. V. Jagadish |
VLDB | 1 |
| 2005 | Semantic Adaptation of Schema Mappings when Schemas Evolve
Cong Yu 0001, Lucian Popa 0001 |
VLDB | 1 |
| 2004 | Constraint-Based XML Query Rewriting For Data IntegrationabstractWe study the problem of answering queries through a target schema, given a set of mappings between one or more source schemas and this target schema, and given that the data is at the sources. The schemas can be any combination of relational or XML schemas, and can be independently designed. In addition to the source-to-target mappings, we consider as part of the mapping scenario a set of target constraints specifying additional properties on the target schema. This becomes particularly important when integrating data from multiple data sources with overlapping data and when such constraints can express data merging rules at the target. We define the semantics of query answering in such an integration scenario, and design two novel algorithms, basic query rewrite and query resolution, to implement the semantics. The basic query rewrite algorithm reformulates target queries in terms of the source schemas, based on the mappings. The query resolution algorithm generates additional rewritings that merge related information from multiple sources and assemble a coherent view of the data, by incorporating target constraints. The algorithms are implemented and then evaluated using a comprehensive set of experiments based on both synthetic and real-life data integration scenarios. Cong Yu 0001, Lucian Popa 0001 |
SIGMOD Conference | 1 |
| 2004 | Schema-Free XQuery
Yunyao Li 0001, Cong Yu 0001, H. V. Jagadish |
VLDB | 2 |
| 2003 | Querying XML using structures and keywords in timberabstractThis demonstration will describe how Timber, a native XML database system, has been extended with the capability to answer XML-style structured queries (e.g., XQuery) with embedded IR-style keyword-based non-boolean conditions. With the original structured query processing engine and the IR extensions built into the system, Timber is well suited for efficiently and effectively processing queries with both structural and textual content constraints. Cong Yu 0001, H. V. Jagadish, Dragomir R. Radev |
SIGIR | 1 |
| 2003 | Querying Structured Text in an XML DatabaseabstractXML databases often contain documents comprising structured text. Therefore, it is important to integrate "information retrieval style" query evaluation, which is well-suited for natural language text, with standard "database style" query evaluation, which handles structured queries efficiently. Relevance scoring is central to information retrieval. In the case of XML, this operation becomes more complex because the data required for scoring could reside not directly in an element itself but also in its descendant elements.In this paper, we propose a bulk-algebra, TIX, and describe how it can be used as a basis for integrating information retrieval techniques into a standard pipelined database query evaluation engine. We develop new evaluation strategies essential to obtaining good performance, including a stack-based TermJoin algorithm for efficiently scoring composite elements. We report results from an extensive experimental evaluation, which show, among other things, that the new TermJoin access method outperforms a direct implementation of the same functionality using standard operators by a large factor. Shurug Al-Khalifa, Cong Yu 0001, H. V. Jagadish |
SIGMOD Conference | 2 |
| 2003 | TIMBER: A Native System for Querying XMLabstractXML has become ubiquitous, and XML data has to be managed in databases. The current industry standard is to map XML data into relational tables and store this information in a relational database. Such mappings create both expressive power problems and performance problems.In the TIMBER [7] project we are exploring the issues involved in storing XML in native format. We believe that the key intellectual contribution of this system is a comprehensive set-at-a-time query processing ability in a native XML store, with all the standard components of relational query processing, including algebraic rewriting and a cost-based optimizer. Stelios Paparizos, Shurug Al-Khalifa, Adriane Chapman, H. V. Jagadish, Laks V. S. Lakshmanan, Andrew Nierman, Jignesh M. Patel, Divesh Srivastava, Nuwee Wiwatwattana, Yuqing Wu, Cong Yu 0001 |
SIGMOD Conference | 11 |
| 2002 | TIMBER: A native XML database
H. V. Jagadish, Shurug Al-Khalifa, Adriane Chapman, Laks V. S. Lakshmanan, Andrew Nierman, Stelios Paparizos, Jignesh M. Patel, Divesh Srivastava, Nuwee Wiwatwattana, Yuqing Wu, Cong Yu 0001 |
VLDB J. | 11 |