VLDB 2026 Research / reviewers in the wild / expert
Qixu Chen
dblp:352/7194
· DBLP profile ↗
3ranked-venue papers
3as first author
3since 2021 · last 2025
0000-0003-3172-4039ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 3 · 3 first-author · 3 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
3 papers |
Query processing and optimization · 45% Data integration and cleaning · 38% Recommender systems · 17% |
Topics — the 5 heaviest of 6, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Data integration and cleaning › data quality
data quality rules |
0.9 | 1 | 2025 | Auto-Test: Learning Semantic-Domain Constraints for Unsupervised Error Detection in Tables · Proc. ACM Manag. Data 2025 |
Data integration and cleaning › data preprocessing › data cleaning
error detection |
0.9 | 1 | 2025 | Auto-Test: Learning Semantic-Domain Constraints for Unsupervised Error Detection in Tables · Proc. ACM Manag. Data 2025 |
Recommender systems
preference elicitation |
0.8 | 1 | 2024 | Robust Best Point Selection under Unreliable User Feedback · Proc. VLDB Endow. 2024 |
Query processing and optimization
top-k query processing |
0.8 | 1 | 2024 | Robust Best Point Selection under Unreliable User Feedback · Proc. VLDB Endow. 2024 |
Query processing and optimization
preference query |
0.7 | 1 | 2023 | Finding Best Tuple via Error-prone User Interaction · ICDE 2023 |
Methods — techniques the papers use, named apart from their topics
statistical tests · 0.9optimization · 0.9interaction algorithm design · 0.8asymptotic question complexity analysis · 0.8interactive comparison · 0.7approximation algorithm · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Auto-Test: Learning Semantic-Domain Constraints for Unsupervised Error Detection in TablesabstractData cleaning is a long-standing challenge in data management. While powerful logic and statistical algorithms have been developed to detect and repair data errors in tables, existing algorithms predominantly rely on domain-experts to first manually specify data-quality constraints specific to a given table, before data cleaning algorithms can be applied. In this work, we observe that there is an important class of data-quality constraints that we call Semantic-Domain Constraints, which can be reliably inferred and automatically applied to any tables, without requiring domain-experts to manually specify on a per-table basis. We develop a principled framework to systematically learn such constraints from table corpora using large-scale statistical tests, which can further be distilled into a core set of constraints using our optimization framework, with provable quality guarantees. Extensive evaluations show that this new class of constraints can be used to both (1) directly detect errors on real tables in the wild, and (2) augment existing expert-driven data-cleaning techniques as a new class of complementary constraints. Our code and data are available at https://github.com/qixuchen/AutoTest for future research. Qixu Chen, Yeye He, Raymond Chi-Wing Wong, Weiwei Cui 0001, Dongmei Zhang 0001, Surajit Chaudhuri |
Proc. ACM Manag. Data | 1 |
| 2024 | Robust Best Point Selection under Unreliable User FeedbackabstractThe task of finding a user's utility function (representing the user's preference) by asking them to compare pairs of points through a series of questions, each requiring him/her to compare 2 points for choosing a more preferred one, to find the best point in the database is a common problem in the database community. However, in real-world scenarios, users may provide unreliable answers due to two major types of errors, namely persistent errors and random errors. Existing interaction algorithms either simply assume that all answers provided by the user are reliable, or are capable of handling random errors only, which can lead to finding undesirable points, ignoring persistent errors. To address this challenge, we propose more generalized algorithms that are robust to both persistent and random errors made by the user. Specifically, we propose (1) an algorithm that asks an asymptotically optimal number of questions, and (2) an algorithm that asks an even smaller number of questions empirically, with provable performance guarantee. Our experiments on both real and synthetic datasets demonstrate that our algorithms outperform existing methods in terms of accuracy, even with a small number of questions asked. Qixu Chen, Raymond Chi-Wing Wong |
Proc. VLDB Endow. | 1 |
| 2023 | Finding Best Tuple via Error-prone User InteractionabstractIn the literature of the database community, there are a lot of studies about finding a utility function from a user (representing the user’s preference), via interaction with the user by asking a number of questions each requiring him/her to compare 2 points for choosing a more preferred point, in order to find the best tuple in the database containing a lot of tuples. In the real world, the user may make mistakes (carelessly), which means that s/he may answer some of the questions wrongly. Unfortunately, existing interaction algorithms may find the undesirable point based on the wrongly learnt utility function because they assume that all answers from the user are 100% correct. In particular, even if the user answers only 1 wrong answer, the output of the existing algorithms may be far away from the users’ real need. Motivated by this, in this paper, we propose a new problem of finding the most interesting point via interaction which is robust to possible mistakes made by a user. Besides, we propose (1) an algorithm that asks an asymptotically optimal number of questions when the dataset contains 2 dimensions and (2) two algorithms with provable performance guarantee when the dataset contains d dimensions where d≥ 2. Experiments on real and synthetic datasets show that our algorithms outperform the existing ones with a higher accuracy with only a small number of questions asked. Qixu Chen, Raymond Chi-Wing Wong |
ICDE | 1 |