VLDB 2026 Research / reviewers in the wild / expert
Longchao Da
dblp:334/1633
· DBLP profile ↗
14ranked-venue papers
7as first author
14since 2021 · last 2026
0009-0000-8631-9634ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 7 first-author · 11 since 2021Databases, data management, data science and information retrieval · 5 · 2 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 5 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Measuring What Matters: Scenario-Driven Evaluation for Trajectory Predictors in Autonomous DrivingabstractBeing able to anticipate the motion of surrounding agents is essential for the safe operation of autonomous driving systems in dynamic situations. While various methods have been proposed for trajectory prediction, the current evaluation practices still rely on error-based metrics (e.g., ADE, FDE), which reveal the accuracy from a post-hoc view but ignore the actual effect the predictor brings to the self-driving vehicles (SDVs), especially in complex interactive scenarios: a high-quality predictor not only chases accuracy, but should also captures all possible directions a neighbor agent might move, to support the SDVs' cautious decision-making. Given that the existing metrics hardly account for this standard, in our work, we propose a comprehensive pipeline that adaptively evaluates the predictor's performance by two dimensions: accuracy and diversity. Based on the criticality of the driving scenario, these two dimensions are dynamically combined and result in a final score for the predictor's performance. Extensive experiments on a closed-loop benchmark using a real-world dataset show that our pipeline yields a more reasonable evaluation than traditional metrics by better reflecting the correlation of the predictors' evaluation with the autonomous vehicles' driving performance. This evaluation pipeline shows a robust way to select a predictor that potentially contributes most to the SDV's driving performance. Longchao Da, David Isele, Hua Wei 0001, Manish Saroya |
AAAI | 1 |
| 2025 | Fully Heteroscedastic Count Regression with Deep Double Poisson NetworksabstractNeural networks capable of accurate, input-conditional uncertainty representation are essential for real-world AI systems. Deep ensembles of Gaussian networks have proven highly effective for continuous regression due to their ability to flexibly represent aleatoric uncertainty via unrestricted heteroscedastic variance, which in turn enables accurate epistemic uncertainty estimation. However, no analogous approach exists for $\textit{count}$ regression, despite many important applications. To address this gap, we propose the Deep Double Poisson Network (DDPN), a novel neural discrete count regression model that outputs the parameters of the Double Poisson distribution, enabling arbitrarily high or low predictive aleatoric uncertainty for count data and improving epistemic uncertainty estimation when ensembled. We formalize and prove that DDPN exhibits robust regression properties similar to heteroscedastic Gaussian models via learnable loss attenuation, and introduce a simple loss modification to control this behavior. Experiments on diverse datasets demonstrate that DDPN outperforms current baselines in accuracy, calibration, and out-of-distribution detection, establishing a new state-of-the-art in deep count regression. Spencer Young, Porter Jenkins, Longchao Da, Jeffrey Dotson, Hua Wei 0001 |
ICML | 3 |
| 2025 | DeepShade: Enable Shade Simulation by Text-conditioned Image GenerationabstractHeatwaves pose a significant threat to public health, especially as global warming intensifies. However, current routing systems (e.g., online maps) fail to incorporate shade information due to the difficulty of estimating shades directly from noisy satellite imagery and the limited availability of training data for generative models. In this paper, we address these challenges through two main contributions. First, we build an extensive dataset covering diverse longitude-latitude regions, varying levels of building density, and different urban layouts. Leveraging Blender-based 3D simulations alongside building outlines, we capture building shadows under various solar zenith angles throughout the year and at different times of day. These simulated shadows are aligned with satellite images, providing a rich resource for learning shade patterns. Second, we propose the DeepShade, a diffusion-based model designed to learn and synthesize shade variations over time. It emphasizes the nuance of edge features by jointly considering RGB with the Canny edge layer, and incorporates contrastive learning to capture the temporal change rules of shade. Then, by conditioning on textual descriptions of known conditions (e.g., time of day, solar angles), our framework provides improved performance in generating shade images. We demonstrate the utility of our approach by using our shade predictions to calculate shade ratios for real-world route planning in Tempe, Arizona. We believe this work will benefit society by providing a reference for urban planning in extreme heat weather and its potential practical applications in the environment. Longchao Da, Xiangrui Liu, Mithun Shivakoti, Thirulogasankar Pranav Kutralingam, Yezhou Yang, Hua Wei 0001 |
IJCAI | 1 |
| 2025 | GE-Chat: A Graph Enhanced RAG Framework for Evidential Response Generation of LLMsabstractLarge Language Models (LLMs) have become integral to human decision-making processes. However, their outputs are not always reliable, often requiring users to assess the accuracy of the information provided manually. This issue is exacerbated by hallucinated responses, which are frequently presented with convincing but incorrect explanations, leading to trust concerns among users. To address this challenge, we propose GE-Chat, a knowledge Graph-enhanced retrieval-augmented generation framework designed to deliver Evidence-based responses. Specifically, when users upload a document, GE-Chat constructs a knowledge graph to support a retrieval-augmented agent, enriching the agent's responses with external knowledge beyond its training data. We further incorporate Chain-of-Thought (CoT) reasoning, n-hop subgraph searching, and entailment-based sentence generation to ensure accurate evidence retrieval. Experimental results demonstrate that our approach improves the ability of existing models to identify precise evidence in free-form contexts, offering a reliable mechanism for verifying LLM-generated conclusions and enhancing trustworthiness. Longchao Da, Parth Mitesh Shah, Kuanru Liou, Jiaxing Zhang 0002, Hua Wei 0001 |
IJCAI | 1 |
| 2025 | FlanS: A Foundation Model for Free-Form Language-based Segmentation in Medical ImagesabstractKDD ’25, August 3–7, 2025, Toronto, ON, Canada Longchao Da, Rui Wang 0184, Xiaojian Xu 0002, Parminder Bhatia, Taha A. Kass-Hout, Hua Wei 0001, Cao Xiao |
KDD (2) | 1 |
| 2025 | Uncertainty Quantification and Confidence Calibration in Large Language Models: A SurveyabstractUncertainty quantification (UQ) enhances the reliability of Large Language Models (LLMs) by estimating confidence in outputs, enabling risk mitigation and selective prediction. However, traditional UQ methods struggle with LLMs due to computational constraints and decoding inconsistencies. Moreover, LLMs introduce unique uncertainty sources, such as input ambiguity, reasoning path divergence, and decoding stochasticity, that extend beyond classical aleatoric and epistemic uncertainty. To address this, we introduce a new taxonomy that categorizes UQ methods based on computational efficiency and uncertainty dimensions, including input, reasoning, parameter, and prediction uncertainty. We evaluate existing techniques, summarize existing benchmarks and metrics for UQ, assess their real-world applicability, and identify open challenges, emphasizing the need for scalable, interpretable, and robust UQ approaches to enhance LLM reliability. Xiaoou Liu, Tiejin Chen, Longchao Da, Chacha Chen, Zhen Lin 0001, Hua Wei 0001 |
KDD (2) | 3 |
| 2025 | Protecting Privacy against Membership Inference Attack with LLM Fine-tuning through FlatnessabstractThe privacy concerns associated with the use of Large Language Models (LLMs) have grown dramatically with the development of pioneer LLMs such as ChatGPT. Differential Privacy (DP) techniques that utilize DP-SGD are explored in existing work to mitigate their privacy risks at the cost of generalization degradation. Our paper reveals that the flatness of DP-SGD trained models’ loss landscape plays an essential role in the trade-off between their privacy and generalization. We further propose a holistic framework Privacy-Flat to enforce appropriate weight flatness, which substantially improves model generalization with promising privacy protection. It innovates from three coarse-to-grained levels: Perturbation-aware min-max optimization within a layer, flatness-guided sparse prefix-tuning across layers, and weight knowledge distillation between private & non-private weights copies. We empirically demonstrate that our framework Privacy-Flat outperforms vanilla private training baseline while protecting privacy from membership inference attacks (MIA). Comprehensive experiments of both black-box and white-box scenarios are conducted to demonstrate the effectiveness of our proposal in enhancing generalization. The code link is provided at https://github.com/tiejin98/Privacy_ Flatness. Tiejin Chen, Longchao Da, Huixue Zhou, Pingzhi Li, Kaixiong Zhou, Tianlong Chen 0001, Hua Wei 0001 |
SDM | 2 |
| 2025 | CoMAL: Collaborative Multi-Agent Large Language Models for Mixed-Autonomy TrafficabstractThe integration of autonomous vehicles into urban traffic has great potential to improve efficiency by reducing congestion and optimizing traffic flow systematically. In this paper, we introduce CoMAL (Collaborative Multi-Agent LLMs), a framework designed to address the mixed-autonomy traffic problem by collaboration among autonomous vehicles to optimize traffic flow. CoMAL is built upon large language models and operates in an interactive traffic simulation environment. Specifically, It utilizes a Perception Module to observe surrounding agents and a Memory Module to store strategies for each agent. The overall workflow includes a Collaboration Module that encourages autonomous vehicles to discuss the effective strategy and allocate roles, a reasoning engine to determine optimal behaviors based on assigned roles, and an Execution Module that controls vehicle actions using a hybrid approach combining rule-based models. Experimental results demonstrate that CoMAL achieves superior performance on the Flow benchmark. Additionally, we evaluate the impact of different language models and compare our framework with reinforcement learning approaches. It highlights the strong cooperative capability of LLM agents and presents a promising solution to the mixed-autonomy traffic challenge. The code is available at https://github.com/Hyan-Yao/CoMAL Huaiyuan Yao, Longchao Da, Vishnu Nandam, Justin Turnau, Zhiwei Liu 0001, Linsey Pang, Hua Wei 0001 |
SDM | 2 |
| 2025 | FM-LC: A Hierarchical Framework for Urban Flood Mapping by Land-Cover Identification ModelsabstractUrban flooding in arid regions threatens infrastructure and public safety. Fine-scale mapping of flood extents is vital for effective emergency response and resilience planning, but limited spectral contrast, rapid hydrological changes, and heterogeneous land covers make this task challenging. High-resolution, daily PlanetScope imagery provides the temporal and spatial detail needed. In this work, we introduce FM-LC, a hierarchical framework for Flood Mapping by Land Cover identification, for this challenging task. Through a three-stage process, it first uses an initial multi-class U-Net to segment imagery into water, vegetation, built area, and bare ground classes. We identify that this method has confusion between spectrally similar categories (e.g., water vs. vegetation). Second, by early checking, the class with the major misclassified area is flagged, and a lightweight binary ‘expert’ segmentation model is trained to distinguish the flagged class from the rest. Third, a Bayesian smoothing step refines boundaries and removes spurious noise by leveraging nearby pixel information. We validate the framework on the April 2024 Dubai storm event, demonstrating average F1-score improvements of up to 29% across all land-cover classes and notably sharper flood delineations. Compared to conventional single-stage U-Nets, FM-LC achieves over 12% higher mean F1, significant gains for vegetation classification, and more reliable temporal tracking of flood dynamics. These results highlight FM-LC as a practical and scalable solution for high-resolution flood mapping in complex urban and arid environments. Longchao Da, Hua Wei 0001 |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2024 | Prompt to Transfer: Sim-to-Real Transfer for Traffic Signal Control with Prompt LearningabstractNumerous solutions are proposed for the Traffic Signal Control (TSC) tasks aiming to provide efficient transportation and alleviate traffic congestion. Recently, promising results have been attained by Reinforcement Learning (RL) methods through trial and error in simulators, bringing confidence in solving cities' congestion problems. However, performance gaps still exist when simulator-trained policies are deployed to the real world. This issue is mainly introduced by the system dynamic difference between the training simulators and the real-world environments. In this work, we leverage the knowledge of Large Language Models (LLMs) to understand and profile the system dynamics by a prompt-based grounded action transformation to bridge the performance gap. Specifically, this paper exploits the pre-trained LLM's inference ability to understand how traffic dynamics change with weather conditions, traffic states, and road types. Being aware of the changes, the policies' action is taken and grounded based on realistic dynamics, thus helping the agent learn a more realistic policy. We conduct experiments on four different scenarios to show the effectiveness of the proposed PromptGAT's ability to mitigate the performance gap of reinforcement learning from simulation to reality (sim-to-real). Longchao Da, Minquan Gao, Hua Wei 0001 |
AAAI | 1 |
| 2024 | Probabilistic Offline Policy Ranking with Approximate Bayesian ComputationabstractIn practice, it is essential to compare and rank candidate policies offline before real-world deployment for safety and reliability. Prior work seeks to solve this offline policy ranking (OPR) problem through value-based methods, such as Off-policy evaluation (OPE). However, they fail to analyze special case performance (e.g., worst or best cases), due to the lack of holistic characterization of policies’ performance. It is even more difficult to estimate precise policy values when the reward is not fully accessible under sparse settings. In this paper, we present Probabilistic Offline Policy Ranking (POPR), a framework to address OPR problems by leveraging expert data to characterize the probability of a candidate policy behaving like experts, and approximating its entire performance posterior distribution to help with ranking. POPR does not rely on value estimation, and the derived performance posterior can be used to distinguish candidates in worst-, best-, and average-cases. To estimate the posterior, we propose POPR-EABC, an Energy-based Approximate Bayesian Computation (ABC) method conducting likelihood-free inference. POPR-EABC reduces the heuristic nature of ABC by a smooth energy function, and improves the sampling efficiency by a pseudo-likelihood. We empirically demonstrate that POPR-EABC is adequate for evaluating policies in both discrete and continuous action spaces across various experiment environments, and facilitates probabilistic comparisons of candidate policies before deployment. Longchao Da, Porter Jenkins, Trevor Schwantes, Jeffrey Dotson, Hua Wei 0001 |
AAAI | 1 |
| 2024 | Shaded Route Planning Using Active Segmentation and Identification of Satellite Images
Longchao Da, Rohan Chhibba, Rushabh Jaiswal, Ariane Middel, Hua Wei 0001 |
CIKM | 1 |
| 2024 | RegExplainer: Generating Explanations for Graph Neural Networks in Regression TasksabstractGraph regression is a fundamental task that has gained significant attention in
various graph learning tasks. However, the inference process is often not easily
interpretable. Current explanation techniques are limited to understanding Graph
Neural Network (GNN) behaviors in classification tasks, leaving an explanation gap
for graph regression models. In this work, we propose a novel explanation method
to interpret the graph regression models (XAIG-R). Our method addresses the
distribution shifting problem and continuously ordered decision boundary issues
that hinder existing methods away from being applied in regression tasks. We
introduce a novel objective based on the graph information bottleneck theory (GIB)
and a new mix-up framework, which can support various GNNs and explainers
in a model-agnostic manner. Additionally, we present a self-supervised learning
strategy to tackle the continuously ordered labels in regression tasks. We evaluate
our proposed method on three benchmark datasets and a real-life dataset introduced
by us, and extensive experiments demonstrate its effectiveness in interpreting GNN
models in regression tasks. Jiaxing Zhang 0002, Zhuomin Chen, Longchao Da, Hua Wei 0001 |
NeurIPS | 4 |
| 2024 | Libsignal: an open library for traffic signal controlabstractThis paper introduces a library for cross-simulator comparison of reinforcement learning models in traffic signal control tasks. This library is developed to implement recent state-of-the-art reinforcement learning models with extensible interfaces and unified cross-simulator evaluation metrics. It supports commonly-used simulators in traffic signal control tasks, including Simulation of Urban MObility(SUMO) and CityFlow, and multiple benchmark datasets for fair comparisons. We conducted experiments to validate our implementation of the models and to calibrate the simulators so that the experiments from one simulator could be referential to the other. Based on the validated models and calibrated environments, this paper compares and reports the performance of current state-of-the-art RL algorithms across different datasets and simulators. This is the first time that these methods have been compared fairly under the same datasets with different simulators. Xiaoliang Lei, Longchao Da, Bin Shi 0003, Hua Wei 0001 |
Mach. Learn. | 3 |