VLDB 2026 Research / reviewers in the wild / expert
Chuxuan Hu
dblp:332/5523
· DBLP profile ↗
5ranked-venue papers
3as first author
5since 2021 · last 2025
0009-0001-3746-2722ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 3 · 3 first-author · 3 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | UPP: Universal Predicate Pushdown to Smart StorageabstractIn large-scale analytics, in-storage processing (ISP) can significantly boost query performance by letting ISP engines (e.g., FPGAs) preselect only the relevant data before sending them to databases.This reduces the amount of not only data transfer between storage and host, but also database computation, facilitating faster query processing.However, existing ISP solutions cannot effectively support a wide range of modern analytical queries because they only support simple combinations of frequently used operators (e.g., =, <), particularly on fixed-length columns.As modern databases allow filter predicates to include numerous operators/functions (e.g., dateadd) compatible with diverse data formats (and their complex combinations), it becomes more challenging for existing approaches to accelerate such queries efficiently.To address the limitations, we propose a new ISP approach, called Universal Predicate Pushdown (UPP), that can accelerate modern analytical databases, leveraging hardware/software co-design for a high level of flexibility.Our core insight is that instead of programming for individual filter operators/functions, we should devise a compact instruction set architecture (ISA) tailored explicitly for predicate pushdown.The software (i.e., database) layer recognizes and compiles various general filters (called a universal predicate) to a set of UPP-compliant instructions, which are then processed efficiently by FPGA using bitwise comparisons, leveraging lightweight metadata.In our experiments with a 100 GB TPC-H dataset, UPP running on SmartSSD could speed up Spark's end-to-end query performance by 1.2×-7.9×without changing input data formats. Ipoom Jeong, Jinghan Huang 0001, Chuxuan Hu, Dohyun Park, Jaeyoung Kang 0004, Nam Sung Kim, Yongjoo Park |
ISCA | 3 |
| 2025 | Drama : Unifying Data Retrieval and Analysis for Open-Domain Analytic QueriesabstractManually conducting real-world data analyses is labor-intensive and inefficient. Despite numerous attempts to automate data science workflows, none of the existing paradigms or systems fully demonstrate all three key capabilities required to support them effectively: (1) open-domain data collection, (2) structured data transformation, and (3) analytic reasoning. To overcome these limitations, we propose Drama , an end-to-end paradigm that answers users' analytic queries in natural language on large-scale open-domain data. Drama unifies data collection, transformation, and analysis as a single pipeline. To quantitatively evaluate system performance on tasks representative of Drama , we construct a benchmark, DramaBench , consisting of two categories of tasks: claim verification and question answering, each comprising 100 instances. These tasks are derived from real-world applications that have gained significant public attention and require the retrieval and analysis of open-domain data. We develop DramaBot , a multi-agent system designed following Drama . It comprises a data retriever that collects and transforms data by coordinating the execution of sub-agents, and a data analyzer that performs structured reasoning over the retrieved data. We evaluate DramaBot on DramaBench together with five state-of-the-art baseline agents. DramaBot achieves 86.5% task accuracy at a cost of $0.05, outperforming all baselines with up to 6.9 times the accuracy and less than 1/6 of the cost. Drama is publicly available at https://github.com/uiuc-kang-lab/drama. Chuxuan Hu, Maxwell Yang, James Weiland, Yeji Lim, Suhas Palawala, Daniel Kang 0001 |
Proc. ACM Manag. Data | 1 |
| 2024 | Data Free Backdoor AttacksabstractBackdoor attacks aim to inject a backdoor into a classifier such that it predicts any input with an attacker-chosen backdoor trigger as an attacker-chosen target class.
Existing backdoor attacks require either retraining the classifier with some clean data or modifying the model's architecture.
As a result, they are 1) not applicable when clean data is unavailable, 2) less efficient when the model is large, and 3) less stealthy due to architecture changes.
In this work, we propose DFBA, a novel retraining-free and data-free backdoor attack without changing the model architecture.
Technically, our proposed method modifies a few parameters of a classifier to inject a backdoor.
Through theoretical analysis, we verify that our injected backdoor is provably undetectable and unremovable by various state-of-the-art defenses under mild assumptions.
Our evaluation on multiple datasets further demonstrates that our injected backdoor: 1) incurs negligible classification loss, 2) achieves 100\% attack success rates, and 3) bypasses six existing state-of-the-art defenses.
Moreover, our comparison with a state-of-the-art non-data-free backdoor attack shows our attack is more stealthy and effective against various defenses while achieving less classification accuracy loss.
We will release our code upon paper acceptance. Bochuan Cao, Jinyuan Jia 0001, Chuxuan Hu, Wenbo Guo 0002, Zhen Xiang, Bo Li 0026, Dawn Song |
NeurIPS | 3 |
| 2024 | Genius: Subteam Replacement with Clustering-based Graph Neural NetworksabstractThe state of the art for subteam replacement, based on random walk graph kernels, encounter the following limitations: (1) ineffective in capturing fine-grained node feature correlations, (2) inefficient without proper pruning mechanisms, and (3) limited applicability to single-member or equal-sized subteam replacements. In this paper, we address these limitations by proposing Genius, a clustering-based graph neural network (GNN) framework that (1) captures team social network knowledge for subteam replacement by deploying team-level attention GNNs (TAGs) and self-supervised positive team contrasting training scheme, (2) generates unsu-pervised team social network member clusters to prune candidates for fast computation, and (3) incorporates a subteam recommender that selects new subteams of flexible sizes. We demonstrate the efficacy of the proposed method in terms of (1) effectiveness: being able to select better subteam members that significantly increase the similarity between the new and original teams, and (2) efficiency: achieving more than 600× speed-up in average running time. Chuxuan Hu, Qinghai Zhou, Hanghang Tong |
SDM | 1 |
| 2024 | LEAP: LLM-powered End-to-end Automatic Library for Processing Social Science Queries on Unstructured DataabstractSocial scientists are increasingly interested in analyzing the semantic information (e.g., emotion) of unstructured data (e.g., Tweets), where the semantic information is not natively present. Performing this analysis in a cost-efficient manner requires using machine learning (ML) models to extract the semantic information and subsequently analyze the now structured data. However, this process remains challenging for domain experts. To demonstrate the challenges in social science analytics, we collect a dataset, QUIET-ML, of 120 real-world social science queries in natural language and their ground truth answers. Existing systems struggle with these queries since (1) they require selecting and applying ML models, and (2) more than a quarter of these queries are vague, making standard tools like natural language to SQL systems unsuited. To address these issues, we develop LEAP, an end-to-end library that answers social science queries in natural language with ML. LEAP filters vague queries to ensure that the answers are deterministic and selects from internally supported and user-defined ML functions to extend the unstructured data to structured tables with necessary annotations. LEAP further generates and executes code to respond to these natural language queries. LEAP achieves a 100% pass @ 3 and 92% pass @ 1 on QUIET-ML, with a $1.06 average end-to-end cost, of which code generation costs $0.02. Chuxuan Hu, Austin Peters, Daniel Kang 0001 |
Proc. VLDB Endow. | 1 |