EDBT 2026 Demo / reviewers in the wild / expert
Chengkai Li 0001
dblp:14/3692
· DBLP profile ↗
66ranked-venue papers in the field
9as first author
9since 2021 · last 2025
0000-0002-1724-8278ORCID · verified
Domains — venue-derived; a paper can count in several
Database Systems & Data Management · 36 (7 first)Information Retrieval & Web Search · 16 (2 first)Data Mining & Knowledge Discovery · 9Knowledge Engineering, Semantic Web & Information Systems · 4Business Process & Enterprise Data · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | TSD-CT: A Benchmark Dataset for Truthfulness Stance DetectionabstractWe present TSD-CT (Truthfulness Stance Detection-Claim and Tweet), a benchmark dataset designed to advance research in truthfulness stance detection. While prior stance detection datasets focus primarily on political figures, topics, or events, TSD-CT targets truthfulness stance of social media posts toward factual claims. Truthfulness stance reflects whether a post endorses a claim as true, rejects it as false, or expresses no clear position. This focus is particularly valuable for tracking public reactions to misinformation and for enabling applications that analyze belief dynamics in online discourse. TSD-CT comprises 5,331 claim-tweet pairs, each annotated into one of five classes: positive, negative, neutral/no stance, topically different, or problematic. To ensure annotation quality, we introduce a strategy that uses gold-standard labels to compute error scores, evaluate annotator performance, and filter out low-quality contributions. The resulting dataset achieves strong inter-annotator agreement. An error analysis further highlights frequent sources of confusion, particularly between neutral/no stance and other classes. The dataset, along with the annotation interface and codebase, is publicly released to facilitate further research. Zhengyuan Zhu, Chengkai Li 0001 |
CIKM | 4 |
| 2025 | TrustMap: Mapping Truthfulness Stance of Social Media Posts on Factual Claims for Geographical AnalysisabstractFactual claims and misinformation circulate widely on social media, shaping public opinion and decision-making. The concept of truthfulness stance refers to whether a text affirms a claim as true, rejects it as false, or takes no clear position. Capturing such stances is essential for understanding how the public engages with and propagates misinformation. We present TrustMap, an application that identifies and visualizes stances of tweets toward factual claims. Users may input factual claims or select claims from a curated set. For each claim, TrustMap retrieves relevant social media posts and applies a retrieval-augmented approach with fine-tuned language models to classify stance. Posts are classified as positive, negative, or neutral/no stance. These classifications are then aggregated by location to reveal regional variations in public opinion. To enhance interpretability, TrustMap uses large language models to generate stance explanations for individual posts and to produce regional stance summaries. By integrating retrieval-augmented truthfulness stance detection with geographical visualization, TrustMap provides the first tool of its kind for exploring how belief in factual claims varies across regions. Zhengyuan Zhu, Chengkai Li 0001 |
CIKM | 4 |
| 2025 | A Knowledge Graph Informing Soil Carbon Modeling
Nasim Shirvani-Mahdavi, Devin Wingfield, Juan Guajardo Gutierrez, Mai Tran, Zhengyuan Zhu, Abhishek Divakar Goudar, Chengkai Li 0001, Virginia L. Jin, Timothy Propst, Dan Roberts, Catherine Stewart, Jianzhong Su, Jennifer Woodward-Greene |
ICWE | 9 |
| 2025 | On Large-Scale Evaluation of Embedding Models for Knowledge Graph Completion
Nasim Shirvani-Mahdavi, Farahnaz Akrami, Chengkai Li 0001 |
WISE (2) | 3 |
| 2024 | Exploring Behavioral Tendencies on Social Media: A Perspective Through Claim Check-Worthiness
Zhengyuan Zhu, Chengkai Li 0001 |
ASONAM (1) | 4 |
| 2024 | Wildfire: A Twitter Social Sensing Platform for LaypersonabstractWe present Wildfire, an innovative social sensing platform designed for laypersons. The goal is to support users in conducting social sensing tasks using Twitter data without programming and data analytics skills. Existing open-source and commercial social sensing tools only support data collection using simple keyword-based or account-based search. On the contrary, Wildfire employs a heuristic graph exploration method to selectively expand the collected tweet-account graph in order to further retrieve more task-relevant tweets and accounts. This approach allows for the collection of data to support complex social sensing tasks that cannot be met with a simple keyword search. In addition, Wildfire provides a range of analytic tools, such as text classification, topic generation, and entity recognition, which can be crucial for tasks such as trend analysis. The platform also provides a web-based user interface for creating and monitoring tasks, exploring collected data, and performing analytics. Zhengyuan Zhu, Foram Patel, Josue Caraballo, Patrick Hennecke, Chengkai Li 0001 |
WSDM | 7 |
| 2024 | Gradient-Based Adversarial Training on Transformer Networks for Detecting Check-Worthy Factual ClaimsabstractThis article presents the latest developments to ClaimBuster’s claim-spotting model, which tackles the critical task of identifying check-worthy claims from large streams of information. We introduce the first adversarially regularized, transformer-based claim-spotting model, which achieves state-of-the-art results on several benchmark datasets. In addition to analyzing model performance metrics, we also quantitatively and qualitatively analyze the impact of ClaimBuster’s real-world deployment. Moreover, to help facilitate reproducibility and community engagement, we publicly release our codebase, dataset, data curation platform, API, Google Colab notebooks, and various ClaimBuster-based demo systems, at claimbuster.org . Kevin Meng, Damian Jimenez, Jacob Daniel Devasier, Sai Sandeep Naraparaju, Fatma Arslan, Daniel Obembe, Chengkai Li 0001 |
ACM Trans. Intell. Syst. Technol. | 7 |
| 2023 | Comprehensive Analysis of Freebase and Dataset Creation for Robust Evaluation of Knowledge Graph Link Prediction Models
Nasim Shirvani-Mahdavi, Farahnaz Akrami, Mohammed Samiul Saeef, Chengkai Li 0001 |
ISWC | 5 |
| 2022 | The CLEF-2022 CheckThat! Lab on Fighting the COVID-19 Infodemic and Fake News Detection
Preslav Nakov, Alberto Barrón-Cedeño, Giovanni Da San Martino, Firoj Alam, Julia Maria Struß, Thomas Mandl 0001, Rubén Míguez, Tommaso Caselli, Mucahid Kutlu, Wajdi Zaghouani, Chengkai Li 0001, Shaden Shaar, Gautam Kishore Shahi, Hamdy Mubarak, Alex Nikolov, Nikolay Babulkov, Yavuz Selim Kartal, Javier Beltrán |
ECIR (2) | 11 |
| 2020 | A Benchmark Dataset of Check-Worthy Factual Claims
Fatma Arslan, Naeemul Hassan, Chengkai Li 0001, Mark Tremayne |
ICWSM | 3 |
| 2020 | Generating Compact and Relaxable Answers to Keyword Queries over Knowledge Graphs
Gong Cheng 0001, Ke Zhang 0045, Chengkai Li 0001 |
ISWC (1) | 4 |
| 2020 | Realistic Re-evaluation of Knowledge Graph Completion Methods: An Experimental StudyabstractIn the active research area of employing embedding models for knowledge graph completion, particularly for the task of link prediction, most prior studies used two benchmark datasets FB15k and WN18 in evaluating such models. Most triples in these and other datasets in such studies belong to reverse and duplicate relations which exhibit high data redundancy due to semantic duplication, correlation or data incompleteness. This is a case of excessive data leakage---a model is trained using features that otherwise would not be available when the model needs to be applied for real prediction. There are also Cartesian product relations for which every triple formed by the Cartesian product of applicable subjects and objects is a true fact. Link prediction on the aforementioned relations is easy and can be achieved with even better accuracy using straightforward rules instead of sophisticated embedding models. A more fundamental defect of these models is that the link prediction scenario, given such data, is non-existent in the real-world. This paper is the first systematic study with the main objective of assessing the true effectiveness of embedding models when the unrealistic triples are removed. Our experiment results show these models are much less accurate than what we used to perceive. Their poor accuracy renders link prediction a task without truly effective automated solution. Hence, we call for re-investigation of possible effective approaches. Farahnaz Akrami, Mohammed Samiul Saeef, Qingheng Zhang, Wei Hu 0007, Chengkai Li 0001 |
SIGMOD Conference | 5 |
| 2020 | Corrigendum to "Discovering and learning sensational episodes of news events" [Inf. Syst. 78 (2018) 68-80]
Xiang Ao 0001, Ping Luo 0001, Chengkai Li 0001, Fuzhen Zhuang, Qing He 0003 |
Inf. Syst. | 3 |
| 2020 | A Benchmarking Study of Embedding-based Entity Alignment for Knowledge Graphs
Zequn Sun 0001, Qingheng Zhang, Wei Hu 0007, Muhao Chen 0001, Farahnaz Akrami, Chengkai Li 0001 |
Proc. VLDB Endow. | 7 |
| 2020 | Relaxing relationship queries on graph data
Gong Cheng 0001, Chengkai Li 0001 |
J. Web Semant. | 3 |
| 2019 | Corrigendum to "Discovering and learning sensational episodes of news events" [Information Systems 78 (2018) 68-80]
Xiang Ao 0001, Ping Luo 0001, Chengkai Li 0001, Fuzhen Zhuang, Qing He 0003 |
Inf. Syst. | 3 |
| 2018 | Re-evaluating Embedding-Based Knowledge Graph Completion MethodsabstractIncompleteness of large knowledge graphs (KG) has motivated many researchers to propose methods to automatically find missing edges in KGs. A promising approach for KG completion (link prediction) is embedding a KG into a continuous vector space. There are different methods in the literature that learn a continuous representation of KG (latent features of KG). The benchmark dataset FB15k has been widely employed to evaluate these methods. However, It has been noted that FB15k contains many pairs of edges in which a pair represents the same relationship in reverse directions. Therefore, the inverse of numerous test triples occurs in the training set. To address this problem, FB15k-237, a subset of FB15k, was created by removing those inverse-duplicate relations to form a more challenging, realistic dataset. There is not any study that investigates how the aforementioned bias in this widely used benchmark dataset affects the results of embedding-based knowledge graph completion methods and whether their promising results are largely due to the bias. Motivated by this question, we conducted extensive experiments and report the link prediction results on FB15K and FB15k-237 using several embedding-based methods. We compare the results of different methods to see how their performances change in absence of inverse relations. Our experiment results demonstrate that the performance of embedding models in link prediction task diminishes tremendously when the inverse relationships do not exist anymore. Farahnaz Akrami, Lingbing Guo, Wei Hu 0007, Chengkai Li 0001 |
CIKM | 4 |
| 2018 | Continuous Monitoring of Pareto Frontiers on Partially Ordered Attributes for Many UsersabstractWe study the problem of continuous object dissemination---given a large number of users and continuously arriving new objects, deliver an object to all users who prefer the object. Many real world applications analyze users' preferences for effective object dissemination. For continuously arriving objects, timely finding users who prefer a new object is challenging. In this paper, we consider an append-only table of objects with multiple attributes and users' preferences on individual attributes are modeled as strict partial orders. An object is preferred by a user if it belongs to the Pareto frontier with respect to the user's partial orders. Users' preferences can be similar. Exploiting shared computation across similar preferences of different users, we design algorithms to find target users of a new object. In order to find users of similar preferences, we study the novel problem of clustering users' preferences that are represented as partial orders. We also present an approximate solution of the problem of finding target users which is more efficient than the exact one while ensuring sufficient accuracy. Furthermore, we extend the algorithms to operate under the semantics of sliding window. We present the results from comprehensive experiments for evaluating the efficiency and effectiveness of the proposed techniques. Afroza Sultana, Chengkai Li 0001 |
EDBT | 2 |
| 2018 | TableView: A Visual Interface for Generating Preview Tables of Entity GraphsabstractThis paper introduces TableView, a tool that helps generate preview tables of massive heterogeneous entity graphs. Given the large number of datasets from many different sources, it is challenging to select entity graphs for a particular need. TableView produces preview tables to offer a compact presentation of important entity types and relationships in an entity graph. Users can quickly browse the preview in a limited space instead of spending precious time or resources to fetch and explore the complete dataset that turns out to be inappropriate for their needs. In this paper we present: 1) a detailed description of TableView's graphical user interface and its various features, 2) a brief description of TableView's preview scoring measures and algorithms, and 3) the demonstration scenarios we intend to present to the audience. Sona Hasani, Chengkai Li 0001 |
ICDE | 3 |
| 2018 | Maverick: Discovering Exceptional Facts from Knowledge GraphsabstractWe present Maverick, a general, extensible framework that discovers exceptional facts about entities in knowledge graphs. To the best of our knowledge, there was no previous study of the problem. We model an exceptional fact about an entity of interest as a context-subspace pair, in which a subspace is a set of attributes and a context is defined by a graph query pattern of which the entity is a match. The entity is exceptional among the entities in the context, with regard to the subspace. The search spaces of both patterns and subspaces are exponentially large. Maverick conducts beam search on the patterns which uses a match-based pattern construction method to evade the evaluation of invalid patterns. It applies two heuristics to select promising patterns to form the beam in each iteration. Maverick traverses and prunes the subspaces organized as a set enumeration tree by exploiting the upper bound properties of exceptionality scoring functions. Results of experiments and user studies using real-world datasets demonstrated substantial performance improvement of the proposed framework over the baselines as well as its effectiveness in discovering exceptional facts. Gensheng Zhang, Damian Jimenez, Chengkai Li 0001 |
SIGMOD Conference | 3 |
| 2018 | Discovering and learning sensational episodes of news events
Xiang Ao 0001, Ping Luo 0001, Chengkai Li 0001, Fuzhen Zhuang, Qing He 0003 |
Inf. Syst. | 3 |
| 2018 | Maverick: A System for Discovering Exceptional Facts from Knowledge GraphsabstractThis paper presents Maverick, a system for discovering exceptional facts about entities in knowledge graphs. Maverick is built upon a beam-search based algorithmic framework which we proposed in a research paper that is published in SIGMOD 2018. In this demonstration proposal, we showcase an end-to-end system that includes a user-facing portal and a cache server. In Maverick, an exceptional fact about an entity of interest is modeled as a context-subspace pair, in which the subspace is a set of attributes and the context is defined by a graph query pattern of which the entity is a match, together with other matching entities. The entity is exceptional among the entities in the context, with regard to the sub-space. The portal allows users to search entities in a knowledge graph and explores exceptional facts about the entities of interest. It presents exceptional facts to users in forms of natural language sentences and illustration charts, for better interpretability of the discovered exceptional facts. The cache server stores intermediate computation results, such as pattern evaluations, exceptionality calculations, and candidate patterns. It is built for sharing computation across entities, such that repetitive computation across entities can be avoided. Gensheng Zhang, Chengkai Li 0001 |
Proc. VLDB Endow. | 2 |
| 2017 | Toward Automated Fact-Checking: Detecting Check-worthy Factual Claims by ClaimBusterabstractThis paper introduces how ClaimBuster, a fact-checking platform, uses natural language processing and supervised learning to detect important factual claims in political discourses. The claim spotting model is built using a human-labeled dataset of check-worthy factual claims from the U.S. general election debate transcripts. The paper explains the architecture and the components of the system and the evaluation of the model. It presents a case study of how ClaimBuster live covers the 2016 U.S. presidential election debates and monitors social media and Australian Hansard for factual claims. It also describes the current status and the long-term goals of ClaimBuster as we keep developing and expanding it. Naeemul Hassan, Fatma Arslan, Chengkai Li 0001, Mark Tremayne |
KDD | 3 |
| 2017 | Cross-Lingual Entity Alignment via Joint Attribute-Preserving Embedding
Zequn Sun 0001, Wei Hu 0007, Chengkai Li 0001 |
ISWC (1) | 3 |
| 2017 | Graph Querying Meets HCI: State of the Art and Future DirectionsabstractQuerying graph databases has emerged as an important research problem for real-world applications that center on large graph data. Given the syntactic complexity of graph query languages (e.g., SPARQL, Cypher), visual graph query interfaces make it easy for non-expert users to query such graph data repositories. In this tutorial, we survey recent developments in the emerging area of visual graph querying paradigm that bridges traditional graph querying with human computer interaction (HCI). We discuss manual and data-driven visual graph query interfaces, various strategies and guidance for constructing graph queries visually, interleaving processing of graph queries and visual actions, and visual exploration of graph query results. In addition, the tutorial suggests open problems and new research directions. In summary, in this tutorial we review and summarize the research thus far into HCI and graph querying in the database community, giving researchers a snapshot of the current state of the art in this topic, and future research directions. Sourav S. Bhowmick, Byron Choi, Chengkai Li 0001 |
SIGMOD Conference | 3 |
| 2017 | ClaimBuster: The First-ever End-to-end Fact-checking SystemabstractOur society is struggling with an unprecedented amount of falsehoods, hyperboles, and half-truths. Politicians and organizations repeatedly make the same false claims. Fake news floods the cyberspace and even allegedly influenced the 2016 election. In fighting false information, the number of active fact-checking organizations has grown from 44 in 2014 to 114 in early 2017. 1 Fact-checkers vet claims by investigating relevant data and documents and publish their verdicts. For instance, PolitiFact.com, one of the earliest and most popular fact-checking projects, gives factual claims truthfulness ratings such as True, Mostly True, Half true, Mostly False, False, and even "Pants on Fire". In the U.S., the election year made fact-checking a part of household terminology. For example, during the first presidential debate on September 26, 2016, NPR.org's live fact-checking website drew 7.4 million page views and delivered its biggest traffic day ever. Naeemul Hassan, Gensheng Zhang, Fatma Arslan, Josue Caraballo, Damian Jimenez, Siddhant Gawsane, Shohedul Hasan, Minumol Joseph, Aaditya Kulkarni, Anil Kumar Nayak, Vikas Sable, Chengkai Li 0001, Mark Tremayne |
Proc. VLDB Endow. | 12 |
| 2017 | Computational Fact Checking through Query PerturbationsabstractOur media is saturated with claims of “facts” made from data. Database research has in the past focused on how to answer queries, but has not devoted much attention to discerning more subtle qualities of the resulting claims, for example, is a claim “cherry-picking”? This article proposes a framework that models claims based on structured data as parameterized queries. Intuitively, with its choice of the parameter setting, a claim presents a particular (and potentially biased) view of the underlying data. A key insight is that we can learn a lot about a claim by “perturbing” its parameters and seeing how its conclusion changes. For example, a claim is not robust if small perturbations to its parameters can change its conclusions significantly. This framework allows us to formulate practical fact-checking tasks—reverse-engineering vague claims, and countering questionable claims—as computational problems. Along with the modeling framework, we develop an algorithmic framework that enables efficient instantiations of “meta” algorithms by supplying appropriate algorithmic building blocks. We present real-world examples and experiments that demonstrate the power of our model, efficiency of our algorithms, and usefulness of their results. You Wu 0001, Pankaj K. Agarwal, Chengkai Li 0001, Jun Yang 0001, Cong Yu 0001 |
ACM Trans. Database Syst. | 3 |
| 2016 | Querying knowledge Graphs By Example entity tuplesabstractWe witness an unprecedented proliferation of knowledge graphs that record millions of entities and their relationships. While knowledge graphs are structure-flexible and content-rich, they are difficult to use. The challenge lies in the gap between their overwhelming complexity and the limited database knowledge of non-professional users. As an initial step toward improving the usability of knowledge graphs, we propose to query such data by example entity tuples, without requiring users to form complex graph queries. Our system, GQBE (Graph Query By Example), automatically discovers a weighted hidden maximum query graph based on input query tuples, to capture a user's query intent. It then efficiently finds top-ranked approximate answer graphs and answer tuples. Nandish Jayaram, Arijit Khan 0001, Chengkai Li 0001, Xifeng Yan, Ramez Elmasri |
ICDE | 3 |
| 2016 | Generating Preview Tables for Entity GraphsabstractUsers are tapping into massive, heterogeneous entity graphs for many applications. It is challenging to select entity graphs for a particular need, given abundant datasets from many sources and the oftentimes scarce information for them. We propose methods to produce preview tables for compact presentation of important entity types and relationships in entity graphs. The preview tables assist users in attaining a quick and rough preview of the data. They can be shown in a limited display space for a user to browse and explore, before she decides to spend time and resources to fetch and investigate the complete dataset. We formulate several optimization problems that look for previews with the highest scores according to intuitive goodness measures, under various constraints on preview size and distance between preview tables. The optimization problem under distance constraint is NP-hard. We design a dynamic-programming algorithm and an Apriori-style algorithm for finding optimal previews. Results from experiments, comparison with related work and user studies demonstrated the scoring measures' accuracy and the discovery algorithms' efficiency. Sona Hasani, Abolfazl Asudeh, Chengkai Li 0001 |
SIGMOD Conference | 4 |
| 2015 | Crowdsourcing Pareto-Optimal Object Finding By Pairwise ComparisonsabstractThis is the first study of crowdsourcing Pareto-optimal object finding over partial orders and by pairwise comparisons, which has applications in public opinion collection, group decision making, and information exploration. Departing from prior studies on crowdsourcing skyline and ranking queries, it considers the case where objects do not have explicit attributes and preference relations on objects are strict partial orders. The partial orders are derived by aggregating crowdsourcers' responses to pairwise comparison questions. The goal is to find all Pareto-optimal objects by the fewest possible questions. It employs an iterative question-selection framework. Guided by the principle of eagerly identifying non-Pareto optimal objects, the framework only chooses candidate questions which must satisfy three conditions. This design is both sufficient and efficient, as it is proven to find a short terminal question sequence. The framework is further steered by two ideas---macro-ordering and micro-ordering. By different micro-ordering heuristics, the framework is instantiated into several algorithms with varying power in pruning questions. Experiment results using both real crowdsourcing marketplace and simulations exhibited not only orders of magnitude reductions in questions when compared with a brute-force approach, but also close-to-optimal performance from the most efficient instantiation. Abolfazl Asudeh, Gensheng Zhang, Naeemul Hassan, Chengkai Li 0001, Gergely V. Záruba |
CIKM | 4 |
| 2015 | Detecting Check-worthy Factual Claims in Presidential DebatesabstractPublic figures such as politicians make claims about "facts" all the time. Journalists and citizens spend a good amount of time checking the veracity of such claims. Toward automatic fact checking, we developed tools to find check-worthy factual claims from natural language sentences. Specifically, we prepared a U.S. presidential debate dataset and built classification models to distinguish check-worthy factual claims from non-factual claims and unimportant factual claims. We also identified the most-effective features based on their impact on the classification models' accuracy. Naeemul Hassan, Chengkai Li 0001, Mark Tremayne |
CIKM | 2 |
| 2015 | Online Frequent Episode MiningabstractFrequent episode mining is a popular framework for discovering sequential patterns from sequence data. Previous studies on this topic usually process data offline in a batch mode. However, for fast-growing sequence data, old episodes may become obsolete while new useful episodes keep emerging. More importantly, in time-critical applications we need a fast solution to discovering the latest frequent episodes from growing data. To this end, we formulate the problem of Online Frequent Episode Mining (OFEM). By introducing the concept of last episode occurrence within a time window, our solution can detect new minimal episode occurrences efficiently, based on which all recent frequent episodes can be discovered directly. Additionally, a trie-based data structure, episode trie, is developed to store minimal episode occurrences in a compact way. We also formally prove the soundness and completeness of our solution and analyze its time as well as space complexity. Experiment results of both online and offline FEM on real data sets show the superiority of our solution. Xiang Ao 0001, Ping Luo 0001, Chengkai Li 0001, Fuzhen Zhuang, Qing He 0003 |
ICDE | 3 |
| 2015 | VIIQ: Auto-Suggestion Enabled Visual Interface for Interactive Graph Query FormulationabstractWe present VIIQ (pronounced as wick), an interactive and iterative visual query formulation interface that helps users construct query graphs specifying their exact query intent. Heterogeneous graphs are increasingly used to represent complex relationships in schemaless data, which are usually queried using query graphs. Existing graph query systems offer little help to users in easily choosing the exact labels of the edges and vertices in the query graph. VIIQ helps users easily specify their exact query intent by providing a visual interface that lets them graphically add various query graph components, backed by an edge suggestion mechanism that suggests edges relevant to the user's query intent. In this demo we present: 1) a detailed description of the various features and user-friendly graphical interface of VIIQ , 2) a brief description of the edge suggestion algorithm, and 3) a demonstration scenario that we intend to show the audience. Nandish Jayaram, Sidharth Goyal, Chengkai Li 0001 |
Proc. VLDB Endow. | 3 |
| 2015 | Querying Knowledge Graphs by Example Entity TuplesabstractWe witness an unprecedented proliferation of knowledge graphs that record millions of entities and their relationships. While knowledge graphs are structure-flexible and content-rich, they are difficult to use. The challenge lies in the gap between their overwhelming complexity and the limited database knowledge of non-professional users. If writing structured queries over “simple” tables is difficult, complex graphs are only harder to query. As an initial step toward improving the usability of knowledge graphs, we propose to query such data by example entity tuples, without requiring users to form complex graph queries. Our system, Graph Query By Example (GQBE), automatically discovers a weighted hidden maximum query graph based on input query tuples, to capture a user's query intent. It then efficiently finds and ranks the top approximate matching answer graphs and answer tuples. We conducted experiments and user studies on the large Freebase and DBpedia datasets and observed appealing accuracy and efficiency. Our system provides a complementary approach to the existing keyword-based methods, facilitating user-friendly graph querying. To the best of our knowledge, there was no such proposal in the past in the context of graphs. Nandish Jayaram, Arijit Khan 0001, Chengkai Li 0001, Xifeng Yan, Ramez Elmasri |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2014 | Anything You Can Do, I Can Do Better: Finding Expert Teams by CrewScoutabstractCrewScout is an expert-team finding system based on the concept of skyline teams and efficient algorithms for finding such teams. Given a set of experts, CrewScout finds all k-expert skyline teams, which are not dominated by any other k-expert teams. The dominance between teams is governed by comparing their aggregated expertise vectors. The need for finding expert teams prevails in applications such as question answering, crowdsourcing, panel selection, and project team formation. The new contributions of this paper include an end-to-end system with an interactive user interface that assists users in choosing teams and an demonstration of its application domains. Naeemul Hassan, Huadong Feng, Venkataraman Ramesh, Gautam Das 0001, Chengkai Li 0001, Nan Zhang 0004 |
CIKM | 5 |
| 2014 | GQBE: Querying knowledge graphs by example entity tuplesabstractWe present GQBE, a system that presents a simple and intuitive mechanism to query large knowledge graphs. Answers to tasks such as “list university professors who have designed some programming languages and also won an award in Computer Science” are best found in knowledge graphs that record entities and their relationships. Real-world knowledge graphs are difficult to use due to their sheer size and complexity and the challenging task of writing complex structured graph queries. Toward better usability of query systems over knowledge graphs, GQBE allows users to query knowledge graphs by example entity tuples without writing complex queries. In this demo we present: 1) a detailed description of the various features and user-friendly GUI of GQBE, 2) a brief description of the system architecture, and 3) a demonstration scenario that we intend to show the audience. Nandish Jayaram, Mahesh Gupta, Arijit Khan 0001, Chengkai Li 0001, Xifeng Yan, Ramez Elmasri |
ICDE | 4 |
| 2014 | Incremental discovery of prominent situational factsabstractWe study the novel problem of finding new, prominent situational facts, which are emerging statements about objects that stand out within certain contexts. Many such facts are newsworthy-e.g., an athlete's outstanding performance in a game, or a viral video's impressive popularity. Effective and efficient identification of these facts assists journalists in reporting, one of the main goals of computational journalism. Technically, we consider an ever-growing table of objects with dimension and measure attributes. A situational fact is a “contextual” skyline tuple that stands out against historical tuples in a context, specified by a conjunctive constraint involving dimension attributes, when a set of measure attributes are compared. New tuples are constantly added to the table, reflecting events happening in the real world. Our goal is to discover constraint-measure pairs that qualify a new tuple as a contextual skyline tuple, and discover them quickly before the event becomes yesterday's news. A brute-force approach requires exhaustive comparison with every tuple, under every constraint, and in every measure subspace. We design algorithms in response to these challenges using three corresponding ideas-tuple reduction, constraint pruning, and sharing computation across measure subspaces. We also adopt a simple prominence measure to rank the discovered facts when they are numerous. Experiments over two real datasets validate the effectiveness and efficiency of our techniques. Afroza Sultana, Naeemul Hassan, Chengkai Li 0001, Jun Yang 0001, Cong Yu 0001 |
ICDE | 3 |
| 2014 | iCheck: computationally combating "lies, d-ned lies, and statistics"abstractAre you fed up with "lies, d---ned lies, and statistics" made up from data in our media? For claims based on structured data, we present a system to automatically assess the quality of claims (beyond their correctness) and counter misleading claims that cherry-pick data to advance their conclusions. The key insight is to model such claims as parameterized queries and consider how parameter perturbations affect their results. We demonstrate our system on claims drawn from U.S. congressional voting records, sports statistics, and publication records of database researchers. You Wu 0001, Brett Walenz, Peggy Li, Andrew Shim, Emre Sonmez, Pankaj K. Agarwal, Chengkai Li 0001, Jun Yang 0001, Cong Yu 0001 |
SIGMOD Conference | 7 |
| 2014 | Data In, Fact Out: Automated Monitoring of Facts by FactWatcherabstractTowards computational journalism, we present FactWatcher, a system that helps journalists identify data-backed, attention-seizing facts which serve as leads to news stories. FactWatcher discovers three types of facts, including situational facts, one-of-the-few facts, and prominent streaks, through a unified suite of data model, algorithm framework, and fact ranking measure. Given an append-only database, upon the arrival of a new tuple, FactWatcher monitors if the tuple triggers any new facts. Its algorithms efficiently search for facts without exhaustively testing all possible ones. Furthermore, FactWatcher provides multiple features in striving for an end-to-end system, including fact ranking, fact-to-statement translation and keyword-based fact search. Naeemul Hassan, Afroza Sultana, You Wu 0001, Gensheng Zhang, Chengkai Li 0001, Jun Yang 0001, Cong Yu 0001 |
Proc. VLDB Endow. | 5 |
| 2014 | Toward Computational Fact-CheckingabstractOur news are saturated with claims of "facts" made from data. Database research has in the past focused on how to answer queries, but has not devoted much attention to discerning more subtle qualities of the resulting claims, e.g., is a claim "cherry-picking"? This paper proposes a framework that models claims based on structured data as parameterized queries. A key insight is that we can learn a lot about a claim by perturbing its parameters and seeing how its conclusion changes. This framework lets us formulate practical fact-checking tasks---reverse-engineering (often intentionally) vague claims, and countering questionable claims---as computational problems. Along with the modeling framework, we develop an algorithmic framework that enables efficient instantiations of "meta" algorithms by supplying appropriate algorithmic building blocks. We present real-world examples and experiments that demonstrate the power of our model, efficiency of our algorithms, and usefulness of their results. You Wu 0001, Pankaj K. Agarwal, Chengkai Li 0001, Jun Yang 0001, Cong Yu 0001 |
Proc. VLDB Endow. | 3 |
| 2014 | On Skyline GroupsabstractWe formulate and investigate the novel problem of finding the skyline k-tuple groups from an n-tuple data set-i.e., groups of k tuples which are not dominated by any other group of equal size, based on aggregate-based group dominance relationship. The major technical challenge is to identify effective anti-monotonic properties for pruning the search space of skyline groups. To this end, we first show that the anti-monotonic property in the well-known Apriori algorithm does not hold for skyline group pruning. Then, we identify two anti-monotonic properties with varying degrees of applicability: order-specific property which applies to SUM, MIN, and MAX as well as weak candidate-generation property which applies to MIN and MAX only. Experimental results on both real and synthetic data sets verify that the proposed algorithms achieve orders of magnitude performance gain over the baseline method. Nan Zhang 0004, Chengkai Li 0001, Naeemul Hassan, Sundaresan Rajasekaran, Gautam Das 0001 |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2014 | Set Predicates in SQL: Enabling Set-Level Comparisons for Dynamically Formed GroupsabstractIn data warehousing and OLAP applications, scalar-level predicates in SQL become increasingly inadequate to support a class of operations that require set-level comparison semantics, i.e., comparing a group of tuples with multiple values. Currently, complex SQL queries composed by scalar-level operations are often formed to obtain even very simple set-level semantics. Such queries are not only difficult to write but also challenging for a database engine to optimize, thus can result in costly evaluation. This paper proposes to augment SQL with set predicate, to bring out otherwise obscured set-level semantics. We studied two approaches to processing set predicates-an aggregate function-based approach and a bitmap index-based approach. Moreover, we designed a histogram-based probabilistic method of set predicate selectivity estimation, for optimizing queries with multiple predicates. The experiments verified its accuracy and effectiveness in optimizing queries. Chengkai Li 0001, Bin He 0001, Muhammad Assad Safiullah |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2013 | iPLUG: Personalized List Recommendation in Twitter
Lijiang Chen, Yibing Zhao, Shimin Chen, Hui Fang 0001, Chengkai Li 0001, Min Wang 0001 |
WISE (2) | 5 |
| 2013 | Wiki3C: exploiting wikipedia for context-aware concept categorizationabstractWikipedia is an important human generated knowledge base containing over 21 million articles organized by millions of categories. In this paper, we exploit Wikipedia for a new task of text mining: Context-aware Concept Categorization. In the task, we focus on categorizing concepts according to their context. We exploit article link feature and category structure in Wikipedia, followed by introducing Wiki3C, an unsupervised and domain independent concept categorization approach based on context. In the approach, we investigate two strategies to select and filter Wikipedia articles for the category representation. Besides, a probabilistic model is employed to compute the semantic relatedness between two concepts in Wikipedia. Experimental evaluation using manually labeled ground truth shows that our proposed Wiki3C can achieve a noticeable improvement over the baselines without considering contextual information. Peng Jiang 0002, Huiman Hou, Lijiang Chen, Shimin Chen, Conglei Yao, Chengkai Li 0001, Min Wang 0001 |
WSDM | 6 |
| 2013 | Personalized ranking in web databases: establishing and utilizing an appropriate workload
Aditya Telang, Sharma Chakravarthy, Chengkai Li 0001 |
Distributed Parallel Databases | 3 |
| 2013 | On contextual ranking queries in databases
Chengkai Li 0001 |
Inf. Syst. | 1 |
| 2013 | Discovering General Prominent Streaks in Sequence DataabstractThis article studies the problem of prominent streak discovery in sequence data. Given a sequence of values, a prominent streak is a long consecutive subsequence consisting of only large (small) values, such as consecutive games of outstanding performance in sports, consecutive hours of heavy network traffic, and consecutive days of frequent mentioning of a person in social media. Prominent streak discovery provides insightful data patterns for data analysis in many real-world applications and is an enabling technique for computational journalism. Given its real-world usefulness and complexity, the research on prominent streaks in sequence data opens a spectrum of challenging problems. A baseline approach to finding prominent streaks is a quadratic algorithm that exhaustively enumerates all possible streaks and performs pairwise streak dominance comparison. For more efficient methods, we make the observation that prominent streaks are in fact skyline points in two dimensions—streak interval length and minimum value in the interval. Our solution thus hinges on the idea to separate the two steps in prominent streak discovery: candidate streak generation and skyline operation over candidate streaks. For candidate generation, we propose the concept of local prominent streak (LPS). We prove that prominent streaks are a subset of LPSs and the number of LPSs is less than the length of a data sequence, in comparison with the quadratic number of candidates produced by the brute-force baseline method. We develop efficient algorithms based on the concept of LPS. The nonlinear local prominent streak (NLPS)-based method considers a superset of LPSs as candidates, and the linear local prominent streak (LLPS)-based method further guarantees to consider only LPSs. The proposed properties and algorithms are also extended for discovering general top- k , multisequence, and multidimensional prominent streaks. The results of experiments using multiple real datasets verified the effectiveness of the proposed methods and showed orders of magnitude performance improvement against the baseline method. Gensheng Zhang, Ping Luo 0001, Min Wang 0001, Chengkai Li 0001 |
ACM Trans. Knowl. Discov. Data | 5 |
| 2012 | On skyline groupsabstractWe formulate and investigate the novel problem of finding the skyline k-tuple groups from an n-tuple dataset - i.e., groups of k tuples which are not dominated by any other group of equal size, based on aggregate-based group dominance relationship. The major technical challenge is to identify effective anti-monotonic properties for pruning the search space of skyline groups. To this end, we show that the anti-monotonic property in the well-known Apriori algorithm does not hold for skyline group pruning. We then identify order-specific property which applies to SUM, MIN, and MAX and weak candidate-generation property which applies to MIN and MAX only. Experimental results on both real and synthetic datasets verify that the proposed algorithms achieve orders of magnitude performance gain over a baseline method. Chengkai Li 0001, Nan Zhang 0004, Naeemul Hassan, Sundaresan Rajasekaran, Gautam Das 0001 |
CIKM | 1 |
| 2012 | Infobox suggestion for Wikipedia entitiesabstractGiven the sheer amount of work and expertise required in authoring Wikipedia articles, automatic tools that help Wikipedia contributors in generating and improving content are valuable. This paper presents our initial step towards building a full-fledged author assistant, particularly for suggesting infobox templates for articles. We build SVM classifiers to suggest infobox template types, among a large number of possible types, to Wikipedia articles without infoboxes. Different from prior works on Wikipedia article classification which deal with only a few label classes for named entity recognition, the much larger 337-class setup in our study is geared towards realistic deployment of infobox suggestion tool. We also emphasize testing on articles without infoboxes, due to that labeled and unlabeled data exhibit different distributions of features, which departs from the typical assumption that they are drawn from the same underlying population. Afroza Sultana, Quazi Mainul Hasan, Ashis Kumer Biswas, Soumyava Das, Habibur Rahman 0001, Chris Ding, Chengkai Li 0001 |
CIKM | 7 |
| 2012 | An optimization framework for map-reduce queriesabstractWe present an effective optimization framework for general SQL-like map-reduce queries, which is based on a novel query algebra and uses a small number of higher-order physical operators that are directly implementable on existing map-reduce systems, such as Hadoop. Although our framework is applicable to any SQL-like map-reduce query language, we focus on a powerful query language, called MRQL. Current map-reduce query languages, such as HiveQL and PigLatin, enable users to plug-in custom map-reduce scripts into queries for those jobs that cannot be declaratively coded in the query language, which may result to suboptimal, error-prone, and hard-to-maintain code. In contrast to these languages, MRQL is expressive enough to capture most of these computations in declarative form and at the same time is amenable to optimization. We describe an optimization framework that maps the algebraic forms derived from the MRQL queries to efficient workflows of map-reduce operations that consist of our physical plan operators. We also describe many algebraic optimizations, such as fusing cascading map-reduce jobs into one job and synthesizing a combine function from the reduce function of a map-reduce job. Finally, we report on a prototype system implementation and we show some performance results of evaluating MRQL queries on a small cluster of computers. Leonidas Fegaras, Chengkai Li 0001, Upa Gupta |
EDBT | 2 |
| 2012 | On "one of the few" objectsabstractObjects with multiple numeric attributes can be compared within any "subspace" (subset of attributes). In applications such as computational journalism, users are interested in claims of the form: Karl Malone is one of the only two players in NBA history with at least 25,000 points, 12,000 rebounds, and 5,000 assists in one's career. One challenge in identifying such "one-of-the-k" claims (k = 2 above) is ensuring their "interestingness". A small k is not a good indicator for interestingness, as one can often make such claims for many objects by increasing the dimensionality of the subspace considered. We propose a uniqueness-based interestingness measure for one-of-the-few claims that is intuitive for non-technical users, and we design algorithms for finding all interesting claims (across all subspaces) from a dataset. Sometimes, users are interested primarily in the objects appearing in these claims. Building on our notion of interesting claims, we propose a scheme for ranking objects and an algorithm for computing the top-ranked objects. Using real-world datasets, we evaluate the efficiency of our algorithms as well as the advantage of our object-ranking scheme over popular methods such as Kemeny optimal rank aggregation and weighted-sum ranking. You Wu 0001, Pankaj K. Agarwal, Chengkai Li 0001, Jun Yang 0001, Cong Yu 0001 |
KDD | 3 |
| 2012 | Entity-Relationship Queries over WikipediaabstractWikipedia is the largest user-generated knowledge base. We propose a structured query mechanism, entity-relationship query , for searching entities in the Wikipedia corpus by their properties and interrelationships. An entity-relationship query consists of multiple predicates on desired entities. The semantics of each predicate is specified with keywords. Entity-relationship query searches entities directly over text instead of preextracted structured data stores. This characteristic brings two benefits: (1) Query semantics can be intuitively expressed by keywords; (2) It only requires rudimentary entity annotation, which is simpler than explicitly extracting and reasoning about complex semantic information before query-time. We present a ranking framework for general entity-relationship queries and a position-based Bounded Cumulative Model (BCM) for accurate ranking of query answers. We also explore various weighting schemes for further improving the accuracy of BCM. We test our ideas on a 2008 version of Wikipedia using a collection of 45 queries pooled from INEX entity ranking track and our own crafted queries. Experiments show that the ranking and weighting schemes are both effective, particularly on multipredicate queries. Chengkai Li 0001, Cong Yu 0001 |
ACM Trans. Intell. Syst. Technol. | 2 |
| 2012 | One Size Does Not Fit All: Toward User- and Query-Dependent Ranking for Web DatabasesabstractWith the emergence of the deep web, searching web databases in domains such as vehicles, real estate, etc., has become a routine task. One of the problems in this context is ranking the results of a user query. Earlier approaches for addressing this problem have used frequencies of database values, query logs, and user profiles. A common thread in most of these approaches is that ranking is done in a user- and/or query-independent manner. This paper proposes a novel query- and user-dependent approach for ranking query results in web databases. We present a ranking model, based on two complementary notions of user and query similarity, to derive a ranking function for a given user query. This function is acquired from a sparse workload comprising of several such ranking functions derived for various user-query pairs. The model is based on the intuition that similar users display comparable ranking preferences over the results of similar queries. We define these similarities formally in alternative ways and discuss their effectiveness analytically and experimentally over two distinct web databases. Aditya Telang, Chengkai Li 0001, Sharma Chakravarthy |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2011 | Computational Journalism: A Call to Arms to Database Researchers
Sarah Cohen, Chengkai Li 0001, Jun Yang 0001, Cong Yu 0001 |
CIDR | 2 |
| 2011 | Prominent streak discovery in sequence dataabstractThis paper studies the problem of prominent streak discovery in sequence data. Given a sequence of values, a prominent streak is a long consecutive subsequence consisting of only large (small) values. For finding prominent streaks, we make the observation that prominent streaks are skyline points in two dimensions- streak interval length and minimum value in the interval. Our solution thus hinges upon the idea to separate the two steps in prominent streak discovery' candidate streak generation and skyline operation over candidate streaks. For candidate generation, we propose the concept of local prominent streak (LPS). We prove that prominent streaks are a subset of LPSs and the number of LPSs is less than the length of a data sequence, in comparison with the quadratic number of candidates produced by a brute-force baseline method. We develop efficient algorithms based on the concept of LPS. The non-linear LPS-based method (NLPS) considers a superset of LPSs as candidates, and the linear LPS-based method (LLPS) further guarantees to consider only LPSs. The results of experiments using multiple real datasets verified the effectiveness of the proposed methods and showed orders of magnitude performance improvement against the baseline method. Chengkai Li 0001, Ping Luo 0001, Min Wang 0001, Yong Yu 0001 |
KDD | 2 |
| 2011 | XML Query Optimization in Map-Reduce
Leonidas Fegaras, Chengkai Li 0001, Upa Gupta, Jijo Philip |
WebDB | 2 |
| 2010 | EntityEngine: answering entity-relationship queries using shallow semanticsabstractWe introduce EntityEngine, a system for answering entity-relationship queries over text. Such queries combine SQL-like structures with IR-style keyword constraints and therefore, can be expressive and flexible in querying about entities and their relationships. EntityEngine consists of various offline and online components, including a position-based ranking model for accurate ranking of query answers and a novel entity-centric index for efficient query evaluation. Chengkai Li 0001, Cong Yu 0001 |
CIKM | 2 |
| 2010 | Facetedpedia: enabling query-dependent faceted search for wikipediaabstractFacetedpedia is a faceted search system that dynamically discovers query-dependent faceted interfaces for Wikipedia search result articles. In this paper, we give an overview of Facetedpedia, present the system architecture and implementation techniques, and elaborate on a demonstration scenario. Chengkai Li 0001, Senjuti Basu Roy, Rakesh Ramegowda, Gautam Das 0001 |
CIKM | 2 |
| 2010 | Facetedpedia: dynamic generation of query-dependent faceted interfaces for wikipediaabstractThis paper proposes Facetedpedia, a faceted retrieval system for information discovery and exploration in Wikipedia. Given the set of Wikipedia articles resulting from a keyword query, Facetedpedia generates a faceted interface for navigating the result articles. Compared with other faceted retrieval systems, Facetedpedia is fully automatic and dynamic in both facet generation and hierarchy construction, and the facets are based on the rich semantic information from Wikipedia. The essence of our approach is to build upon the collaborative vocabulary in Wikipedia, more specifically the intensive internal structures (hyperlinks) and folksonomy (category system). Given the sheer size and complexity of this corpus, the space of possible choices of faceted interfaces is prohibitively large. We propose metrics for ranking individual facet hierarchies by user's navigational cost, and metrics for ranking interfaces (each with k facets) by both their average pairwise similarities and average navigational costs. We thus develop faceted interface discovery algorithms that optimize the ranking metrics. Our experimental evaluation and user study verify the effectiveness of the system. Chengkai Li 0001, Senjuti Basu Roy, Lekhendro Lisham, Gautam Das 0001 |
WWW | 1 |
| 2009 | Query-By-Keywords (QBK): Query Formulation Using Semantics and Feedback
Aditya Telang, Sharma Chakravarthy, Chengkai Li 0001 |
ER | 3 |
| 2007 | Supporting ranking and clustering as generalized order-by and group-byabstractThe Boolean semantics of SQL queries cannot adequately capture the "fuzzy" preferences and "soft" criteria required in non-traditional data retrieval applications. One way to solve this problem is to add a flavor of "information retrieval" into database queries by allowing fuzzy query conditions and flexibly supporting grouping and ranking of the query results within the DBMS engine. While ranking is already supported by all major commercial DBMSs natively, support of flexibly grouping is still very limited (i.e., group-by). Chengkai Li 0001, Min Wang 0001, Lipyeow Lim, Haixun Wang, Kevin Chen-Chuan Chang |
SIGMOD Conference | 1 |
| 2006 | Supporting ad-hoc ranking aggregatesabstractThis paper presents a principled framework for efficient processing of ad-hoc top-k (ranking) aggregate queries, which provide the k groups with the highest aggregates as results. Essential support of such queries is lacking in current systems, which process the queries in a naïve materialize-group-sort scheme that can be prohibitively inefficient. Our framework is based on three fundamental principles. The Upper-Bound Principle dictates the requirements of early pruning, and the Group-Ranking and Tuple-Ranking Principles dictate group-ordering and tuple-ordering requirements. They together guide the query processor toward a provably optimal tuple schedule for aggregate query processing. We propose a new execution framework to apply the principles and requirements. We address the challenges in realizing the framework and implementing new query operators, enabling efficient group-aware and rank-aware query plans. The experimental study validates our framework by demonstrating orders of magnitude performance improvement in the new query plans, compared with the traditional plans. Chengkai Li 0001, Kevin Chen-Chuan Chang, Ihab F. Ilyas |
SIGMOD Conference | 1 |
| 2005 | RankSQL: Query Algebra and Optimization for Relational Top-k QueriesabstractThis paper introduces RankSQL, a system that provides a systematic and principled framework to support efficient evaluations of ranking (top-k) queries in relational database systems (RDBMS), by extending relational algebra and query optimization. Previously, top-k query processing is studied in the middleware scenario or in RDBMS in a piecemeal fashion, i.e., focusing on specific operator or sitting outside the core of query engines. In contrast, we aim to support ranking as a first-class database construct. As a key insight, the new ranking relationship can be viewed as another logical property of data, parallel to the property of relational data model. While membership is essentially supported in RDBMS, the same support for ranking is clearly lacking. We address the fundamental integration of ranking in RDBMS in a way similar to how membership, i.e., Boolean filtering, is supported. We extend relational algebra by proposing a rank-relational model to capture the ranking property, and introducing new and extended operators to support ranking as a first-class construct. Enabled by the extended algebra, we present a pipelined and incremental execution model of ranking query plans (that cannot be expressed traditionally) based on a fundamental ranking principle. To optimize top-k queries, we propose a dimensional enumeration algorithm to explore the extended plan space by enumerating plans along two dual dimensions: ranking and membership. We also propose a sampling-based method to estimate the cardinality of rank-aware operators, for costing plans. Our experiments show the validity of our framework and the accuracy of the proposed estimation model. Chengkai Li 0001, Kevin Chen-Chuan Chang, Ihab F. Ilyas, Sumin Song |
SIGMOD Conference | 1 |
| 2005 | RankSQL: Supporting Ranking Queries in Relational Database Management Systems
Chengkai Li 0001, Mohamed A. Soliman, Kevin Chen-Chuan Chang, Ihab F. Ilyas |
VLDB | 1 |
| 2003 | ROLEX: Relational On-Line Exchange with XMLabstractNo abstract available. Philip Bohannon, Xin Dong 0001, Sumit Ganguly, Henry F. Korth, Chengkai Li 0001, P. P. S. Narayan, Pradeep Shenoy |
SIGMOD Conference | 5 |
| 2003 | Composing XSL Transformations with XML Publishing ViewsabstractWhile the XML Stylesheet Language for Transformations (XSLT) was not designed as a query language, it is well-suited for many query-like operations on XML documents including selecting and restructuring data. Further, it actively fulfills the role of an XML query language in modern applications and is widely supported by application platform software. However, the use of database techniques to optimize and execute XSLT has only recently received attention in the research community. In this paper, we focus on the case where XSL transformations are to be run on XML documents defined as views of relational databases. For a subset of XSLT, we present an algorithm to compose a transformation with an XML view, eliminating the need for the XSLT execution. We then describe how to extend this algorithm to handle several additional features of XSLT, including a proposed approach for handling recursion. 1. Chengkai Li 0001, Philip Bohannon, Henry F. Korth, P. P. S. Narayan |
SIGMOD Conference | 1 |