VLDB 2026 Research / reviewers in the wild / expert
Su Feng
dblp:40/6521
· DBLP profile ↗
11ranked-venue papers
4as first author
7since 2021 · last 2025
0009-0009-8104-3128ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 7 · 3 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | FastPDB: Towards Bag-Probabilistic Queries at Interactive SpeedsabstractProbabilistic databases (PDBs) provide users with a principled way to query data that is incomplete or imprecise. In this work, we study computing expected multiplicities of query results over probabilistic databases under bag semantics which has PTIME data complexity. However, does this imply that bag probabilistic databases are practical? We strive to answer this question from both a theoretical as well as a systems perspective. We employ concepts from fine-grained complexity to demonstrate that exact bag probabilistic query processing is fundamentally less efficient than deterministic bag query evaluation, but that fast approximations are possible by sampling monomials from a circuit representation of a result tuple's lineage. A remaining issue, however, is that constructing such circuits, while in PTIME, can nonetheless have significant overhead. To avoid this cost, we utilize approximate query processing techniques to directly sample monomials without materializing lineage upfront. Our implementation in FastPDB provides accurate anytime approximation of probabilistic query answers and scales to datasets orders of magnitude larger than competing methods. Aaron Huber, Oliver Kennedy, Atri Rudra, Zhuoyue Zhao 0001, Su Feng, Boris Glavic |
Proc. ACM Manag. Data | 5 |
| 2024 | Learning from Uncertain Data: From Possible Worlds to Possible ModelsabstractWe introduce an efficient method for learning linear models from uncertain data, where uncertainty is represented as a set of possible variations in the data, leading to predictive multiplicity. Our approach leverages abstract interpretation and zonotopes, a type of convex polytope, to compactly represent these dataset variations, enabling the symbolic execution of gradient descent on all possible worlds simultaneously. We develop techniques to ensure that this process converges to a fixed point and derive closed-form solutions for this fixed point. Our method provides sound over-approximations of all possible optimal models and viable prediction ranges. We demonstrate the effectiveness of our approach through theoretical and empirical analysis, highlighting its potential to reason about model and prediction uncertainty due to data quality issues in training data. Jiongli Zhu, Su Feng, Boris Glavic, Babak Salimi |
NeurIPS | 2 |
| 2023 | Efficient Approximation of Certain and Possible Answers for Ranking and Window Queries over Uncertain DataabstractUncertainty arises naturally in many application domains due to, e.g., data entry errors and ambiguity in data cleaning. Prior work in incomplete and probabilistic databases has investigated the semantics and efficient evaluation of ranking and top-k queries over uncertain data. However, most approaches deal with top-k and ranking in isolation and do represent uncertain input data and query results using separate, incompatible data models. We present an efficient approach for under- and over-approximating results of ranking, top-k, and window queries over uncertain data. Our approach integrates well with existing techniques for querying uncertain data, is efficient, and is to the best of our knowledge the first to support windowed aggregation. We design algorithms for physical operators for uncertain sorting and windowed aggregation, and implement them in PostgreSQL. We evaluated our approach on synthetic and real world datasets, demonstrating that it outperforms all competitors, and often produces more accurate results. Su Feng, Boris Glavic, Oliver Kennedy |
Proc. VLDB Endow. | 1 |
| 2022 | LPGN: Language-Guided Proposal Generation Network for Referring Expression ComprehensionabstractReferring expression comprehension(REC) aims to ground the referring expression in an image. Many mainstream frameworks are implemented in a two-stage process: the first-stage model generates candidate region proposals, and the second-stage model locates the referent among the proposals. Existing proposal generators create proposals entirely based on images, so there is a gap between generated proposals and referring expression, which leads to a bottleneck limiting the performance of the whole model. In order to break the bottle-neck, we introduce a novel language-guided proposal generation network: LPGN. Moreover, we introduce an uncertainty-aware proposal generation strategy to tackle the vagueness of language, so as to improve the training effectiveness. LPGN is convenient to integrate it into the existing two-stage REC models, because it is agnostic to the second stage model. Through extensive experiments on benchmark datasets, we demonstrate that our LPGN can generate proposals of higher quality than existing proposal generators and effectively alleviate the proposal bottleneck of the existing two-stage REC model. Chao Yang 0015, Su Feng, Bin Jiang 0006 |
ICME | 3 |
| 2021 | DataSense: Display-Agnostic Data Documentation
Poonam Kumari, Mike Brachmann, Su Feng, Oliver Kennedy, Boris Glavic |
CIDR | 3 |
| 2021 | Learning Content and Context with Language Bias for Visual Question AnsweringabstractVisual Question Answering (VQA) is a challenging multi-modal task to answer questions about an image. Many works concentrate on how to reduce language bias which makes models answer questions ignoring visual content and language context. However, reducing language bias also weakens the ability of VQA models to learn context prior. To address this issue, we propose a novel learning strategy named CCB, which forces VQA models to answer questions relying on Content and Context with language Bias. Specifically, CCB establishes Content and Context branches on top of a base VQA model and forces them to focus on local key content and global effective context respectively. Moreover, a joint loss function is proposed to reduce the importance of biased samples and retain their beneficial influence on answering questions. Experiments show that CCB outperforms the state-of-the-art methods on VQA-CP v2. Chao Yang 0015, Su Feng, Dongsheng Li 0002, Huawei Shen, Bin Jiang 0006 |
ICME | 2 |
| 2021 | Efficient Uncertainty Tracking for Complex Queries with Attribute-level BoundsabstractIncomplete and probabilistic database techniques are principled methods for coping with uncertainty in data. Unfortunately, the class of queries that can be answered efficiently over such databases is severely limited, even when advanced approximation techniques are employed.We introduce attribute-annotated uncertain databases (AU-DBs), an uncertain data model that annotates tuples and attribute values with bounds to compactly approximate an incomplete database. AU-DBs are closed under relational algebra with aggregation using an efficient evaluation semantics. Using optimizations that trade accuracy for performance, our approach scales to complex queries and large datasets, and produces accurate results. Su Feng, Boris Glavic, Aaron Huber, Oliver Kennedy |
SIGMOD Conference | 1 |
| 2019 | Data Debugging and Exploration with VizierabstractWe present Vizier, a multi-modal data exploration and debugging tool. The system supports a wide range of operations by seamlessly integrating Python, SQL, and automated data curation and debugging methods. Using Spark as an execution backend, Vizier handles large datasets in multiple formats. Ease-of-use is attained through integration of a notebook with a spreadsheet-style interface and with visualizations that guide and support the user in the loop. In addition, native support for provenance and versioning enable collaboration and uncertainty management. In this demonstration we will illustrate the diverse features of the system using several realistic data science tasks based on real data. Mike Brachmann, Carlos Bautista, Sonia Castelo Quispe, Su Feng, Juliana Freire, Boris Glavic, Oliver Kennedy, Heiko Müller 0001, Rémi Rampin, William Spoth, Ying Yang 0005 |
SIGMOD Conference | 4 |
| 2019 | Uncertainty Annotated Databases - A Lightweight Approach for Approximating Certain AnswersabstractCertain answers are a principled method for coping with uncertainty that arises in many practical data management tasks. Unfortunately, this method is expensive and may ex- clude useful (if uncertain) answers. Thus, users frequently resort to less principled approaches to resolve uncertainty. In this paper, we propose Uncertainty Annotated Databases (UA-DBs), which combine an under- and over-approximation of certain answers to achieve the reliability of certain answers, with the performance of a classical database system. Furthermore, in contrast to prior work on certain answers, UA-DBs achieve a higher utility by including some (explicitly marked) answers that are not certain. UA-DBs are based on incomplete K-relations, which we introduce to generalize the classical set-based notion of incomplete databases and certain answers to a much larger class of data models. Using an implementation of our approach, we demonstrate experimentally that it efficiently produces tight approximations of certain answers that are of high utility. Su Feng, Aaron Huber, Boris Glavic, Oliver Kennedy |
SIGMOD Conference | 1 |
| 2017 | Debugging Transactions and Tracking their Provenance with ReenactmentabstractDebugging transactions and understanding their execution are of immense importance for developing OLAP applications, to trace causes of errors in production systems, and to audit the operations of a database. However, debugging transactions is hard for several reasons: 1) after the execution of a transaction, its input is no longer available for debugging, 2) internal states of a transaction are typically not accessible, and 3) the execution of a transaction may be affected by concurrently running transactions. We present a debugger for transactions that enables non-invasive, postmortem debugging of transactions with provenance tracking and supports what-if scenarios (changes to transaction code or data). Using reenactment , a declarative replay technique we have developed, a transaction is replayed over the state of the DB seen by its original execution including all its interactions with concurrently executed transactions from the history. Importantly, our approach uses the temporal database and audit logging capabilities available in many DBMS and does not require any modifications to the underlying database system nor transactional workload. Xing Niu 0002, Bahareh Arab, Seokki Lee, Su Feng, Xun Zou, Dieter Gawlick, Vasudha Krishnaswamy, Zhen Hua Liu, Boris Glavic |
Proc. VLDB Endow. | 4 |
| 2005 | Mechanizing Weakly Ground Termination Proving of Term Rewriting Systems by Structural and Cover-Set Inductions
Su Feng |
J. Comput. Sci. Technol. | 1 |