EDBT 2026 Demo / reviewers in the wild / expert
Xiaolan Wang 0001
dblp:117/4879-1
· DBLP profile ↗
23ranked-venue papers
9as first author
7since 2021 · last 2024
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 18 · 9 first-author · 4 since 2021Artificial intelligence and machine learning · 5 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Computer networks · 1Graphics, computer vision, multimedia, augmented reality and games · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Noisy Pairing and Partial Supervision for Stylized Opinion SummarizationabstractOpinion summarization research has primarily focused on generating summaries reflecting important opinions from customer reviews without paying much attention to the writing style.In this paper, we propose the stylized opinion summarization task, which aims to generate a summary of customer reviews in the desired (e.g., professional) writing style.To tackle the difficulty in collecting customer and professional review pairs, we develop a non-parallel training framework, Noisy Pairing and Partial Supervision (Napa ), which trains a stylized opinion summarization system from non-parallel customer and professional review sets.We create a benchmark PRO-SUM by collecting customer and professional reviews from Yelp and Michelin.Experimental results on PROSUM and FewSum demonstrate that our non-parallel training framework consistently improves both automatic and human evaluations, successfully building a stylized opinion summarization model that can generate professionally-written summaries from customer reviews. 1 Hayate Iso, Xiaolan Wang 0001, Yoshi Suhara |
INLG | 2 |
| 2022 | Summarizing Community-based Question-Answer PairsabstractCommunity-based Question Answering (CQA), which allows users to acquire their desired information, has increasingly become an essential component of online services in various domains such as E-commerce, travel, and dining.However, an overwhelming number of CQA pairs makes it difficult for users without particular intent to find useful information spread over CQA pairs.To help users quickly digest the key information, we propose the novel CQA summarization task that aims to create a concise summary from CQA pairs.To this end, we first design a multi-stage data annotation process and create a benchmark dataset, CO-QASUM, based on the Amazon QA corpus.We then compare a collection of extractive and abstractive summarization methods and establish a strong baseline approach DedupLED for the CQA summarization task.Our experiment further confirms two key challenges, sentencetype transfer and deduplication removal, towards the CQA summarization task.Our data and code are publicly available.1 * Work done while at Megagon Labs. 1 https://github.com/megagonlabs/ qa-summarization Q: Is this actually a rigid board or more of a floppy mat?A: It is rigid.themain board is rigid,the two sides are semi.Q: Is this actually a rigid board or more of a floppy mat?A: The main area is very sturdy.Then there are two work area pads that are more flexible so when moving those I keep two hands on them.Q: how wide is each Side piece?"A: 16 inches wide (there are two).Q: will this mat hold 1000 piece puzzle?A: most certainly will.… … (omitted 27 QAs) Summary: This puzzle board comes with a rigid main board.You can arrange pieces in the middle and on two side pieces, and then pick up those side pieces to place them atop the middle area before folding the wings in.The dimension of the puzzle space is 32"x21.75".The closed unit is almost the same size as the puzzle workspace (32"x21.75").There are two 16" wide side inserts.The mat holds most 1000 pieces puzzles.It is too big to use on you lap and definitely needs a table.(a).QAs for a puzzle board product (Input) (b).Summary of QAs (Output) Q: what is the storage size when case is fully closed for storage?A: Closed size is 32.25 x 22.75".Q: What size is the closed unit?A: Closed is almost the same size as the puzzle workspace.32.25 x 22. Ting-Yao Hsu, Yoshi Suhara, Xiaolan Wang 0001 |
EMNLP | 3 |
| 2022 | Beyond Opinion Mining: Summarizing Opinions of Customer ReviewsabstractCustomer reviews are vital for making purchasing decisions in the Information Age. Such reviews can be automatically summarized to provide the user with an overview of opinions. In this tutorial, we present various aspects of opinion summarization that are useful for researchers and practitioners. First, we will introduce the task and major challenges. Then, we will present existing opinion summarization solutions, both pre-neural and neural. We will discuss how summarizers can be trained in the unsupervised, few-shot, and supervised regimes. Each regime has roots in different machine learning methods, such as auto-encoding, controllable text generation, and variational inference. Finally, we will discuss resources and evaluation methods and conclude with the future directions. This three-hour tutorial will provide a comprehensive overview over major advances in opinion summarization. The listeners will be well-equipped with the knowledge that is both useful for research and practical applications. Reinald Kim Amplayo, Arthur Brazinskas, Yoshi Suhara, Xiaolan Wang 0001, Bing Liu 0001 |
SIGIR | 4 |
| 2021 | Rotom: A Meta-Learned Data Augmentation Framework for Entity Matching, Data Cleaning, Text Classification, and BeyondabstractDeep Learning revolutionizes almost all fields of computer science including data management. However, the demand for high-quality training data is slowing down deep neural nets' wider adoption. To this end, data augmentation (DA), which generates more labeled examples from existing ones, becomes a common technique. Meanwhile, the risk of creating noisy examples and the large space of hyper-parameters make DA less attractive in practice. We introduce Rotom, a multi-purpose data augmentation framework for a range of data management and mining tasks including entity matching, data cleaning, and text classification. Rotom features InvDA, a new DA operator that generates natural yet diverse augmented examples by formulating DA as a seq2seq task. The key technical novelty of Rotom is a meta-learning framework that automatically learns a policy for combining examples from different DA operators, whereby combinatorially reduces the hyper-parameters space. Our experimental results show that Rotom effectively improves a model's performance by combining multiple DA operators, even when applying them individually does not yield performance improvement. With this strength, Rotom outperforms the state-of-the-art entity matching and data cleaning systems in the low-resource settings as well as two recently proposed DA techniques for text classification. Zhengjie Miao, Yuliang Li 0001, Xiaolan Wang 0001 |
SIGMOD Conference | 3 |
| 2021 | Constructing Explainable Opinion Graphs from ReviewsabstractThe Web is a major resource of both factual and subjective information. While there are significant efforts to organize factual information into knowledge bases, there is much less work on organizing opinions, which are abundant in subjective data, into a structured format. Nofar Carmeli, Xiaolan Wang 0001, Yoshihiko Suhara, Stefanos Angelidis, Yuliang Li 0001, Wang Chiew Tan |
WWW | 2 |
| 2021 | Data Augmentation for ML-driven Data Preparation and IntegrationabstractIn recent years, we have witnessed the development of novel data augmentation (DA) techniques for creating additional training data needed by machine learning based solutions. In this tutorial, we will provide a comprehensive overview of techniques developed by the data management community for data preparation and data integration. In addition to surveying task-specific DA operators that leverage rules, transformations, and external knowledge for creating additional training data, we also explore the advanced DA techniques such as interpolation, conditional generation, and DA policy learning. Finally, we describe the connection between DA and other machine learning paradigms such as active learning, pre-training, and weakly-supervised learning. We hope that this discussion can shed light on future research directions for a holistic data augmentation framework for high-quality dataset creation. Yuliang Li 0001, Xiaolan Wang 0001, Zhengjie Miao, Wang Chiew Tan |
Proc. VLDB Endow. | 2 |
| 2021 | Extractive Opinion Summarization in Quantized Transformer SpacesabstractAbstract We present the Quantized Transformer (QT), an unsupervised system for extractive opinion summarization. QT is inspired by Vector- Quantized Variational Autoencoders, which we repurpose for popularity-driven summarization. It uses a clustering interpretation of the quantized space and a novel extraction algorithm to discover popular opinions among hundreds of reviews, a significant step towards opinion summarization of practical scope. In addition, QT enables controllable summarization without further training, by utilizing properties of the quantized space to extract aspect-specific summaries. We also make publicly available Space, a large-scale evaluation benchmark for opinion summarizers, comprising general and aspect-specific summaries for 50 hotels. Experiments demonstrate the promise of our approach, which is validated by human studies where judges showed clear preference for our method over competitive baselines. Stefanos Angelidis, Reinald Kim Amplayo, Yoshihiko Suhara, Xiaolan Wang 0001, Mirella Lapata |
Trans. Assoc. Comput. Linguistics | 4 |
| 2020 | OpinionDigest: A Simple Framework for Opinion SummarizationabstractWe present OPINIONDIGEST, an abstractive opinion summarization framework, which does not rely on gold-standard summaries for training.The framework uses an Aspect-based Sentiment Analysis model to extract opinion phrases from reviews, and trains a Transformer model to reconstruct the original reviews from these extractions.At summarization time, we merge extractions from multiple reviews and select the most popular ones.The selected opinions are used as input to the trained Transformer model, which verbalizes them into an opinion summary.OPINIONDIGEST can also generate customized summaries, tailored to specific user needs, by filtering the selected opinions according to their aspect and/or sentiment.Automatic evaluation on YELP data shows that our framework outperforms competitive baselines.Human studies on two corpora verify that OPINIONDIGEST produces informative summaries and shows promising customization capabilities 1 . Yoshihiko Suhara, Xiaolan Wang 0001, Stefanos Angelidis, Wang Chiew Tan |
ACL | 2 |
| 2020 | Snippext: Semi-supervised Opinion Mining with Augmented DataabstractOnline services are interested in solutions to opinion mining, which is the problem of extracting aspects, opinions, and sentiments from text. One method to mine opinions is to leverage the recent success of pre-trained language models which can be fine-tuned to obtain high-quality extractions from reviews. However, fine-tuning language models still requires a non-trivial amount of training data. Zhengjie Miao, Yuliang Li 0001, Xiaolan Wang 0001, Wang Chiew Tan |
WWW | 3 |
| 2020 | Deep or Simple Models for Semantic Tagging? It Depends on your Data
Yuliang Li 0001, Xiaolan Wang 0001, Wang Chiew Tan |
Proc. VLDB Endow. | 3 |
| 2019 | MIDAS: Finding the Right Web Sources to Fill Knowledge GapsabstractKnowledge bases, massive collections of facts (RDF triples) on diverse topics, support vital modern applications. However, existing knowledge bases contain very little data compared to the wealth of information on the Web. This is because the industry standard in knowledge base creation and augmentation suffers from a serious bottleneck: they rely on domain experts to identify appropriate web sources to extract data from. Efforts to fully automate knowledge extraction have failed to improve this standard: these automated systems are able to retrieve much more data and from a broader range of sources, but they suffer from very low precision and recall. As a result, these large-scale extractions remain unexploited. In this paper, we present MIDAS, a system that harnesses the results of automated knowledge extraction pipelines to repair the bottleneck in industrial knowledge creation and augmentation processes. MIDAS automates the suggestion of good-quality web sources and describes what to extract with respect to augmenting an existing knowledge base. We make three major contributions. First, we introduce a novel concept, web source slices, to describe the contents of a web source. Second, we define a profit function to quantify the value of a web source slice with respect to augmenting an existing knowledge base. Third, we develop effective and highly-scalable algorithms to derive high-profit web source slices. We demonstrate that MIDAS produces high-profit results and outperforms the baselines significantly on both real-world and synthetic datasets. Xiaolan Wang 0001, Xin Dong 0001, Yang Li 0150, Alexandra Meliou |
ICDE | 1 |
| 2019 | Voyageur: An Experiential Travel Search EngineabstractWe describe Voyageur, which is an application of experiential search to the domain of travel. Unlike traditional search engines for online services, experiential search focuses on the experiential aspects of the service under consideration. In particular, Voyageur needs to handle queries for subjective aspects of the service (e.g., quiet hotel, friendly staff) and combine these with objective attributes, such as price and location. Voyageur also highlights interesting facts and tips about the services the user is considering to provide them with further insights into their choices. Sara Evensen, Aaron Feng, Alon Y. Halevy, Vivian Li, Yuliang Li 0001, Huining Liu, George A. Mihaila, John Morales, Natalie Nuno, Ekaterina Pavlovic, Wang Chiew Tan, Xiaolan Wang 0001 |
WWW | 13 |
| 2019 | Explain3D: Explaining Disagreements in Disjoint DatasetsabstractData plays an important role in applications, analytic processes, and many aspects of human activity. As data grows in size and complexity, we are met with an imperative need for tools that promote understanding and explanations over data-related operations. Data management research on explanations has focused on the assumption that data resides in a single dataset, under one common schema. But the reality of today's data is that it is frequently unintegrated, coming from different sources with different schemas. When different datasets provide different answers to semantically similar questions, understanding the reasons for the discrepancies is challenging and cannot be handled by the existing single-dataset solutions. In this paper, we propose explain3D, a framework for explaining the disagreements across disjoint datasets (3D). Explain3D focuses on identifying the reasons for the differences in the results of two semantically similar queries operating on two datasets with potentially different schemas. Our framework leverages the queries to perform a semantic mapping across the relevant parts of their provenance; discrepancies in this mapping point to causes of the queries' differences. Exploiting the queries gives explain3D an edge over traditional schema matching and record linkage techniques, which are query-agnostic. Our work makes the following contributions: (1) We formalize the problem of deriving optimal explanations for the differences of the results of semantically similar queries over disjoint datasets. Our optimization problem considers two types of explanations, provenance-based and value-based, defined over an evidence mapping, which makes our solution interpretable. (2) We design a 3-stage framework for solving the optimal explanation problem. (3) We develop a smart-partitioning optimizer that improves the efficiency of the framework by orders of magnitude. (4) We experiment with real-world and synthetic data to demonstrate that explain3D can derive precise explanations efficiently, and is superior to alternative methods based on integration techniques and single-dataset explanation frameworks. Xiaolan Wang 0001, Alexandra Meliou |
Proc. VLDB Endow. | 1 |
| 2018 | Scalable Semantic Querying of TextabstractWe present the Koko system that takes declarative information extraction to a new level by incorporating advances in natural language processing techniques in its extraction language. K oko is novel in that its extraction language simultaneously supports conditions on the surface of the text and on the structure of the dependency parse tree of sentences, thereby allowing for more refined extractions. K oko also supports conditions that are forgiving to linguistic variation of expressing concepts and allows to aggregate evidence from the entire document in order to filter extractions. To scale up, K oko exploits a multi-indexing scheme and heuristics for efficient extractions. We extensively evaluate K oko over publicly available text corpora. We show that K oko indices take up the smallest amount of space, are notably faster and more effective than a number of prior indexing schemes. Finally, we demonstrate K oko 's scalability on a corpus of 5 million Wikipedia articles. Xiaolan Wang 0001, Aaron Feng, Behzad Golshan, Alon Y. Halevy, George A. Mihaila, Hidekazu Oiwa, Wang Chiew Tan |
Proc. VLDB Endow. | 1 |
| 2018 | Koko: A System for Scalable Semantic Querying of TextabstractK oko is a declarative information extraction system that incorporates advances in natural language processing techniques in its extraction language. K oko 's extraction language supports simultaneous specification of conditions over the surface syntax and on the structure of the dependency parse tree of sentences, thereby allowing for more refined extractions. Furthermore, the K oko extraction language allows for aggregating evidence from an input document and supports conditions that are tolerant of linguistic variation of expressing concepts. In this demo, we outline the design of K oko , a system for extracting information and understanding the results of the extraction. K oko provides an interactive interface that allows participants to write queries, understand the input and results of the queries. In particular, the user can customize the input text, visualize the input text's dependency parse trees, and understand the correspondences between query components, dependency tree nodes, text tokens, and the computation and associated scores that led to an extraction. Xiaolan Wang 0001, Jiyu Komiya, Yoshihiko Suhara, Aaron Feng, Behzad Golshan, Alon Y. Halevy, Wang Chiew Tan |
Proc. VLDB Endow. | 1 |
| 2018 | Robust Multi-Variate Temporal Features of Multi-Variate Time SeriesabstractMany applications generate and/or consume multi-variate temporal data, and experts often lack the means to adequately and systematically search for and interpret multi-variate observations. In this article, we first observe that multi-variate time series often carry localized multi-variate temporal features that are robust against noise. We then argue that these multi-variate temporal features can be extracted by simultaneously considering, at multiple scales, temporal characteristics of the time series along with external knowledge , including variate relationships that are known a priori. Relying on these observations, we develop data models and algorithms to detect robust multi-variate temporal (RMT) features that can be indexed for efficient and accurate retrieval and can be used for supporting data exploration and analysis tasks. Experiments confirm that the proposed RMT algorithm is highly effective and efficient in identifying robust multi-scale temporal features of multi-variate time series. Silvestro Roberto Poccia, K. Selçuk Candan, Maria Luisa Sapino, Xiaolan Wang 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2017 | QFix: Diagnosing Errors through Query HistoriesabstractData-driven applications rely on the correctness of their data to function properly and effectively. Errors in data can be incredibly costly and disruptive, leading to loss of revenue, incorrect conclusions, and misguided policy decisions. While data cleaning tools can purge datasets of many errors before the data is used, applications and users interacting with the data can introduce new errors. Subsequent valid updates can obscure these errors and propagate them through the dataset causing more discrepancies. Even when some of these discrepancies are discovered, they are often corrected superficially, on a case-by-case basis, further obscuring the true underlying cause, and making detection of the remaining errors harder. Xiaolan Wang 0001, Alexandra Meliou, Eugene Wu 0002 |
SIGMOD Conference | 1 |
| 2016 | QFix: Demonstrating Error Diagnosis in Query HistoriesabstractAn increasing number of applications in all aspects of society rely on data. Despite the long line of research in data cleaning and repairs, data correctness has been an elusive goal. Errors in the data can be extremely disruptive, and are detrimental to the effectiveness and proper function of data-driven applications. Even when data is cleaned, new errors can be introduced by applications and users who interact with the data. Subsequent valid updates can obscure these errors and propagate them through the dataset causing more discrepancies. Any discovered errors tend to be corrected superficially, on a case-by-case basis, further obscuring the true underlying cause, and making detection of the remaining errors harder. In this demo proposal, we outline the design of QFix, a query-centric framework that derives explanations and repairs for discrepancies in relational data based on potential errors in the queries that operated on the data. This is a marked departure from traditional data-centric techniques that directly fix the data. We then describe how users will use QFix in a demonstration scenario. Participants will be able to select from a number of transactional benchmarks, introduce errors into the queries that are executed, and compare the fixes to the queries proposed by QFix as well as existing alternative algorithms such as decision trees. Xiaolan Wang 0001, Alexandra Meliou, Eugene Wu 0002 |
SIGMOD Conference | 1 |
| 2015 | Data X-Ray: A Diagnostic Tool for Data ErrorsabstractA lot of systems and applications are data-driven, and the correctness of their operation relies heavily on the correctness of their data. While existing data cleaning techniques can be quite effective at purging datasets of errors, they disregard the fact that a lot of errors are systematic, inherent to the process that produces the data, and thus will keep occurring unless the problem is corrected at its source. In contrast to traditional data cleaning, in this paper we focus on data diagnosis: explaining where and how the errors happen in a data generative process. Xiaolan Wang 0001, Xin Dong 0001, Alexandra Meliou |
SIGMOD Conference | 1 |
| 2015 | Error Diagnosis and Data Profiling with Data X-RayabstractThe problem of identifying and repairing data errors has been an area of persistent focus in data management research. However, while traditional data cleaning techniques can be effective at identifying several data discrepancies, they disregard the fact that many errors aresystematic, inherent to the process that produces the data, and thus will keep occurring unless the root cause is identified and corrected. In this demonstration, we will present a large-scale diagnostic framework called DataXRay. Like a medical X-ray that aids the diagnosis of medical conditions by revealing problems underneath the surface, DataXRayreveals hidden connections and common properties among data errors. Thus, in contrast to traditional cleaning methods, which treat the symptoms, our system investigates the underlying conditions that cause the errors. The core of DataXRaycombines an intuitive and principled cost model derived by Bayesian analysis, and an efficient, highly-parallelizable diagnostic algorithm that discovers common properties among erroneous data elements in a top-down fashion. Our system has a simple interface that allows users to load different datasets, to interactively adjust key diagnostic parameters, to explore the derived diagnoses, and to compare with solutions produced by alternative algorithms. Through this demonstration, participants will understand (1) the characteristics of good diagnoses, (2) how and why errors occur in real-world datasets, and (3) the distinctions with other related problems and approaches. Xiaolan Wang 0001, Mary Feng, Yue Wang 0070, Xin Dong 0001, Alexandra Meliou |
Proc. VLDB Endow. | 1 |
| 2014 | Leveraging metadata for identifying local, robust multi-variate temporal (RMT) featuresabstractMany applications generate and/or consume multi-variate temporal data, yet experts often lack the means to adequately and systematically search for and interpret multi-variate observations. In this paper, we first observe that multi-variate time series often carry localized multi-variate temporal features that are robust against noise. We then argue that these multi-variate temporal features can be extracted by simultaneously considering, at multiple scales, temporal characteristics of the time-series along with external knowledge, including variate relationships, known a priori. Relying on these observations, we develop algorithms to detect robust multi-variate temporal (RMT) features which can be indexed for efficient and accurate retrieval and can be used for supporting analysis tasks, such as classification. Experiments confirm that the proposed RMT algorithm is highly effective and efficient in identifying robust multi-scale temporal features of multi-variate time series. Xiaolan Wang 0001, K. Selçuk Candan, Maria Luisa Sapino |
ICDE | 1 |
| 2012 | STFMap: query- and feature-driven visualization of large time series data setsabstractSince many applications rely on time-based data, visualizing temporal data and helping experts explore large time series data sets are critical in many application domains. In this interactive system preview, we argue that time series often carry structural features that can, if efficiently identified and effectively visualized, help reduce visual overload and help the user quickly focus on the relevant portions of the data sets. Relying on this observation, we introduce a novel STFMap system, which includes four innovative query- and feature-driven time series data set visualization techniques: (a) segment-maps, (b) warp-maps, (c) stretch-maps, and (d) feature-maps. These rely on the salient temporal features of the time series and their alignments with respect to the given user query to help users explore the data set in a query-driven fashion. K. Selçuk Candan, Rosaria Rossini, Maria Luisa Sapino, Xiaolan Wang 0001 |
CIKM | 4 |
| 2012 | sDTW: Computing DTW Distances using Locally Relevant Constraints based on Salient Feature AlignmentsabstractMany applications generate and consume temporal data and retrieval of time series is a key processing step in many application domains. Dynamic time warping (DTW) distance between time series of size N and M is computed relying on a dynamic programming approach which creates and fills an N x M grid to search for an optimal warp path . Since this can be costly, various heuristics have been proposed to cut away the potentially unproductive portions of the DTW grid. In this paper, we argue that time series often carry structural features that can be used for identifying locally relevant constraints to eliminate redundant work. Relying on this observation, we propose salient feature based sDTW algorithms which first identify robust salient features in the given time series and then find a consistent alignment of these to establish the boundaries for the warp path search. More specifically, we propose alternative fixed core&adaptive width, adaptive core&fixed width , and adaptive core&adaptive width strategies which enforce different constraints reflecting the high level structural characteristics of the series in the data set. Experiment results show that the proposed sDTW algorithms help achieve much higher accuracy in DTW computation and time series retrieval than fixed core & fixed width algorithms that do not leverage local features of the given time series. K. Selçuk Candan, Rosaria Rossini, Maria Luisa Sapino, Xiaolan Wang 0001 |
Proc. VLDB Endow. | 4 |