EDBT 2026 Demo / reviewers in the wild / expert
Naiqing Guan
dblp:244/1368
· DBLP profile ↗
5ranked-venue papers in the field
4as first author
5since 2021 · last 2026
0000-0001-5170-8277ORCID · corroborated
Domains — venue-derived; a paper can count in several
Database Systems & Data Management · 5 (4 first)
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Relational Deep Dive: Error-Aware Queries Over Unstructured Data
Daren Chao, Naiqing Guan, Nick Koudas |
Proc. VLDB Endow. | 3 |
| 2025 | DataSculpt: Cost-Efficient Label Function Design via Prompting Large Language Models
Naiqing Guan, Nick Koudas |
EDBT | 1 |
| 2024 | ActiveDP: Bridging Active Learning and Data Programming
Naiqing Guan, Nick Koudas |
EDBT | 1 |
| 2024 | WeShap: Weak Supervision Source Evaluation with Shapley ValuesabstractEfficient data annotation stands as a significant bottleneck in training contemporary machine learning models. The Programmatic Weak Supervision (PWS) pipeline presents a solution by utilizing multiple weak supervision sources to automatically label data, thereby expediting the annotation process. Given the varied contributions of these weak supervision sources to the accuracy of PWS, it is imperative to employ a robust and efficient metric for their evaluation. This is crucial not only for understanding the behavior and performance of the PWS pipeline but also for facilitating corrective measures. In this paper, we introduce WeShap values as an evaluation metric. This metric quantifies the average contribution of weak supervision sources within a proxy PWS pipeline, leveraging the theoretical underpinnings of Shapley values. We demonstrate efficient computation of WeShap values using dynamic programming, achieving quadratic computational complexity relative to the number of weak supervision sources. Our experiments demonstrate the versatility of WeShap values across various applications, including the identification of beneficial or detrimental labeling functions, refinement of the PWS pipeline, comprehension of the pipeline's behavior, and scrutinizing specific instances of mislabeled data. Although initially derived from a specific proxy PWS pipeline, we empirically demonstrate the generalizability of WeShap values to other PWS pipeline configurations. Our findings indicate a noteworthy average improvement of 5.0 points in downstream model accuracy through the revision of the PWS pipeline compared to previous state-of-the-art methods, underscoring the efficacy of WeShap values in enhancing data quality for training machine learning models. Naiqing Guan, Nick Koudas |
Proc. VLDB Endow. | 1 |
| 2022 | FILA: Online Auditing of Machine Learning Model Accuracy under Finite Labelling BudgetabstractMachine learning (ML) is increasingly adopted in industrial applications. Typically, a ML pipeline is instantiated to automate the process of collecting training data, training a model, auditing the model accuracy and generating predictions. Naiqing Guan, Nick Koudas |
SIGMOD Conference | 1 |