VLDB 2026 Research / reviewers in the wild / expert
Dequan Wang
dblp:168/4711
· DBLP profile ↗
28ranked-venue papers
5as first author
18since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 19 · 4 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 1 first-author · 5 since 2021Systems, architecture and hardware · 4 · 2 first-author · 2 since 2021Computer networks · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | AgencyBench: Benchmarking the Frontiers of Autonomous Agents in 1M-Token Real-World ContextsabstractKeyu Li, Junhao Shi, Yang Xiao, Mohan Jiang, Jie Sun, Yunze Wu, Dayuan Fu, Shijie Xia, Xiaojie Cai, Tianze Xu, Weiye Si, Wenjie Li, Dequan Wang, Pengfei Liu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Junhao Shi, Mohan Jiang, Jie Sun 0030, Yunze Wu, Dayuan Fu, Shijie Xia, Xiaojie Cai, Weiye Si, Wenjie Li 0002, Dequan Wang, Pengfei Liu 0003 |
ACL (1) | 13 |
| 2026 | FlowGait: Enabling Robust Long-Term Gait Recognition Across Real-World Covariates with mmWave RadarabstractGait recognition enables proactive and personalized smart home interactions, but its long-term reliability is challenged by the non-static nature of gait. Covariates like carrying items and clothing induce a persistent domain shift that degrades traditional, static models. To solve this, we introduce FlowGait, a mmWave-based framework designed for robust, long-term adaptation. It combines self-training with continual learning, allowing the model to daily align with a user’s evolving gait by learning from readily available unlabeled data. It features a specialized transformer network for radar spectrogram analysis and a novel two-stage labeling algorithm that leverages the gait’s hierarchical nature to assign pseudo-labels to the unlabeled data accurately. Evaluated on three challenging datasets from 47 volunteers (covering 12 gait-covariates, 11 routes, and two weeks), FlowGait achieves high accuracies of 94.8 (cross-covariate), 98.6% (cross-route), and 95.5% (cross-day). Notably, for the long-term dataset, it reduced performance decay from 13.6% to just 1.4%, demonstrating its real-world robustness. Dequan Wang, Chenming He, Chengzhen Meng, Xiaoran Fan, Yanyong Zhang |
CHI | 1 |
| 2026 | FeelWave: Enabling Emotion-Aware Voice Interaction through Noise-Robust mmWave Emotion SensingabstractVoice has been a primary interaction mode with LLM-powered assistants. Beyond semantics, voice carries emotional cues with potential to guide empathetic system responses. Yet, robust vocal emotion sensing in noise and its use in optimizing interactions remain underexplored. In response, we present FeelWave, which achieves empathetic voice interaction through noise-robust mmWave emotion sensing and structured LLM prompts. It extracts robust vocal information from mmWave signals, applies audio-to-mmWave transfer learning for efficient emotion recognition, and employs chain-of-thought-based query optimization to enable emotion-adaptive responses. Evaluations show that FeelWave achieves 92.3% emotion recognition accuracy and remains robust in noisy environments, yielding a 62.9 percentage-point gain over audio-based models. In voice interaction studies, 74.3% of users prefer FeelWave, reporting significantly higher satisfaction than a baseline without emotion sensing (4.37 vs. 3.22). A SUS score of 88.3 confirms FeelWave’s high usability in real-world deployment. We hope this work will inspire more empathetic, user-centered AI-driven assistants. You Zuo, Dequan Wang, Chenming He, Chengzhen Meng, Xiaoran Fan, Yanyong Zhang |
CHI | 3 |
| 2026 | Needle in a Haystack: Tracking UAVs from Massive Noise in Real-World 5G-A Base Station Data
Chengzhen Meng, Chenming He, Yidong Jiang, Xiaoran Fan, Dequan Wang, Jianmin Ji, Yanyong Zhang |
MobiSys | 5 |
| 2025 | Decentralized Vehicle Coordination: The Berkeley DeepDrive Drone Dataset and Consensus-Based ModelsabstractA significant portion of roads, particularly in densely populated developing countries, lacks explicitly defined right-of-way rules. These understructured roads pose substantial challenges for autonomous vehicle motion planning, where efficient and safe navigation relies on understanding decentralized human coordination for collision avoidance. This coordination, often termed “social driving etiquette,” remains underexplored due to limited open-source empirical data and suitable modeling frameworks. In this paper, we present a novel dataset and modeling framework designed to study motion planning in these understructured environments. The dataset includes 20 aerial videos of representative scenarios, an image dataset for training vehicle detection models, and a development kit for vehicle trajectory estimation. We demonstrate that a consensus-based modeling approach can effectively explain the emergence of priority orders observed in our dataset, and is therefore a viable framework for decentralized collision avoidance planning. Fangyu Wu 0003, Dequan Wang, Minjune Hwang, Chenhui Hao, Jiamu Zhang, Christopher Chou, Trevor Darrell, Alexandre M. Bayen |
ICRA | 2 |
| 2025 | Ghost Points Matter: Far-Range Vehicle Detection with a Single mmWave Radar in TunnelabstractVehicle detection in tunnels is crucial for traffic monitoring and accident response, yet remains underexplored. In this paper, we develop mmTunnel, a millimeter-wave radar system that achieves far-range vehicle detection in tunnels. The main challenge here is coping with ghost points caused by multi-path reflections, which lead to severe localization errors and false alarms. Instead of merely removing ghost points, we propose correcting them to true vehicle positions by recovering their signal reflection paths, thus reserving more data points and improving detection performance, even in occlusion scenarios. However, recovering complex 3D reflection paths from limited 2D radar points is highly challenging. To address this problem, we develop a multi-path ray tracing algorithm that leverages the ground plane constraint and identifies the most probable reflection path based on signal path loss and spatial distance. We also introduce a curve-to-plane segmentation method to simplify tunnel surface modeling such that we can significantly reduce the computational delay and achieve real-time processing. Chenming He, Chengzhen Meng, Xiaoran Fan, Dequan Wang, Haojie Ren, Jianmin Ji, Yanyong Zhang |
MobiCom | 5 |
| 2025 | Editorial for Special Issue on Foundation Models for Medical Image Analysis
Xiaosong Wang 0001, Dequan Wang, Jens Rittscher, Dimitris N. Metaxas, Shaoting Zhang 0001 |
Medical Image Anal. | 2 |
| 2024 | Dissecting Dissonance: Benchmarking Large Multimodal Models Against Self-Contradictory Instructions
Yuankai Li, Yixin Ye, Dequan Wang |
ECCV (57) | 5 |
| 2024 | Lost in Translation: Latent Concept Misalignment in Text-to-Image Diffusion Models
Juntu Zhao, Junyu Deng, Yixin Ye, Chongxuan Li, Zhijie Deng, Dequan Wang |
ECCV (69) | 6 |
| 2024 | BayesDiff: Estimating Pixel-wise Uncertainty in Diffusion via Bayesian InferenceabstractDiffusion models have impressive image generation capability, but low-quality generations still exist, and their identification remains challenging due to the lack of a proper sample-wise metric. To address this, we propose BayesDiff, a pixel-wise uncertainty estimator for generations from diffusion models based on Bayesian inference. In particular, we derive a novel uncertainty iteration principle to characterize the uncertainty dynamics in diffusion, and leverage the last-layer Laplace approximation for efficient Bayesian inference. The estimated pixel-wise uncertainty can not only be aggregated into a sample-wise metric to filter out low-fidelity images but also aids in augmenting successful generations and rectifying artifacts in failed generations in text-to-image tasks. Extensive experiments demonstrate the efficacy of BayesDiff and its promise for practical applications. Siqi Kou, Dequan Wang, Chongxuan Li, Zhijie Deng |
ICLR | 3 |
| 2024 | An Extensible Framework for Open Heterogeneous Collaborative PerceptionabstractCollaborative perception aims to mitigate the limitations of single-agent perception, such as occlusions, by facilitating data exchange among multiple agents. However, most current works consider a homogeneous scenario where all agents use identity sensors and perception models. In reality, heterogeneous agent types may continually emerge and inevitably face a domain gap when collaborating with existing agents. In this paper, we introduce a new open heterogeneous problem: how to accommodate continually emerging new heterogeneous agent types into collaborative perception, while ensuring high perception performance and low integration cost? To address this problem, we propose HEterogeneous ALliance (HEAL), a novel extensible collaborative perception framework. HEAL first establishes a unified feature space with initial agents via a novel multi-scale foreground-aware Pyramid Fusion network. When heterogeneous new agents emerge with previously unseen modalities or models, we align them to the established unified space with an innovative backward alignment. This step only involves individual training on the new agent type, thus presenting extremely low training costs and high extensibility. To enrich agents' data heterogeneity, we bring OPV2V-H, a new large-scale dataset with more diverse sensor types. Extensive experiments on OPV2V-H and DAIR-V2X datasets show that HEAL surpasses SOTA methods in performance while reducing the training parameters by 91.5\% when integrating 3 new agent types. We further implement a comprehensive codebase at: https://github.com/yifanlu0227/HEAL Yue Hu 0011, Yiqi Zhong, Dequan Wang, Yanfeng Wang 0001, Siheng Chen |
ICLR | 4 |
| 2023 | Back to the Source: Diffusion-Driven Adaptation to Test-Time CorruptionabstractTest-time adaptation harnesses test inputs to improve the accuracy of a model trained on source data when tested on shifted target data. Most methods update the source model by (re-)training on each target domain. While retraining can help, it is sensitive to the amount and order of the data and the hyperparameters for optimization. We update the target data instead, and project all test inputs toward the source domain with a generative diffusion model. Our diffusion-driven adaptation (DDA) method shares its models for classification and generation across all domains, training both on source then freezing them for all targets, to avoid expensive domain-wise retraining. We augment diffusion with image guidance and classifier self-ensembling to automatically decide how much to adapt. Input adaptation by DDA is more robust than model adaptation across a variety of corruptions, models, and data regimes on the ImageNet-C benchmark. With its input-wise updates, DDA succeeds where model adaptation degrades on too little data (small batches), on dependent data (correlated orders), or on mixed data (multiple corruptions). Jialing Zhang, Xihui Liu, Trevor Darrell, Evan Shelhamer, Dequan Wang |
CVPR | 6 |
| 2023 | Text-Guided Foundation Model Adaptation for Pathological Image Classification
Yunkun Zhang, Mu Zhou, Xiaosong Wang 0001, Yu Qiao 0001, Shaoting Zhang 0001, Dequan Wang |
MICCAI (5) | 7 |
| 2022 | Contrastive Test-Time AdaptationabstractTest-time adaptation is a special setting of unsupervised domain adaptation where a trained model on the source domain has to adapt to the target domain without accessing source data. We propose a novel way to leverage self-supervised contrastive learning to facilitate target feature learning, along with an online pseudo labeling scheme with refinement that significantly denoises pseudo labels. The contrastive learning task is applied jointly with pseudo labeling, contrasting positive and negative pairs constructed similarly as MoCo but with source-initialized encoder, and excluding same-class negative pairs indicated by pseudo labels. Meanwhile, we produce pseudo labels online and refine them via soft voting among their nearest neighbors in the target feature space, enabled by maintaining a memory queue. Our method, AdaContrast, achieves state-of-the-art performance on major benchmarks while having several desirable properties compared to existing works, including memory efficiency, insensitivity to hyper-parameters, and better model calibration. Code is released at https://github.com/DianCh/AdaContrast. Dian Chen 0001, Dequan Wang, Trevor Darrell, Sayna Ebrahimi |
CVPR | 2 |
| 2022 | GACT: Activation Compressed Training for Generic Network ArchitecturesabstractTraining large neural network (NN) models requires extensive memory resources, and Activation Compression Training (ACT) is a promising approach to reduce training memory footprint. This paper presents GACT, an ACT framework to support a broad range of machine learning tasks for generic NN architectures with limited domain knowledge. By analyzing a linearized version of ACT’s approximate gradient, we prove the convergence of GACT without prior knowledge on operator type or model architecture. To make training stable, we propose an algorithm that decides the compression ratio for each tensor by estimating its impact on the gradient at run time. We implement GACT as a PyTorch library that readily applies to any NN architecture. GACT reduces the activation memory for convolutional NNs, transformers, and graph NNs by up to 8.1x, enabling training with a 4.2x to 24.7x larger batch size, with negligible accuracy loss. Lianmin Zheng, Dequan Wang, Yukuo Cen, Weize Chen, Xu Han 0007, Jianfei Chen 0001, Zhiyuan Liu 0001, Jie Tang 0001, Joey Gonzalez, Michael W. Mahoney, Alvin Cheung |
ICML | 3 |
| 2021 | CoDeNet: Efficient Deployment of Input-Adaptive Object Detection on Embedded FPGAsabstractDeploying deep learning models on embedded systems for computer vision tasks has been challenging due to limited compute resources and strict energy budgets. The majority of existing work focuses on accelerating image classification, while other fundamental vision problems, such as object detection, have not been adequately addressed. Compared with image classification, detection problems are more sensitive to the spatial variance of objects, and therefore, require specialized convolutions to aggregate spatial information. To address this need, recent work introduces dynamic deformable convolution to augment regular convolutions. Regular convolutions process a fixed grid of pixels across all the spatial locations in an image, while dynamic deformable convolution may access arbitrary pixels in the image with the access pattern being input-dependent and varying with spatial location. These properties lead to inefficient memory accesses of inputs with existing hardware. Qijing Huang 0001, Dequan Wang, Zhen Dong 0003, Yizhao Gao 0002, Yaohui Cai, Bichen Wu, Kurt Keutzer, John Wawrzynek |
FPGA | 2 |
| 2021 | Tent: Fully Test-Time Adaptation by Entropy Minimization
Dequan Wang, Evan Shelhamer, Shaoteng Liu, Bruno A. Olshausen, Trevor Darrell |
ICLR | 1 |
| 2021 | ActNN: Reducing Training Memory Footprint via 2-Bit Activation Compressed TrainingabstractThe increasing size of neural network models has been critical for improvements in their accuracy, but device memory is not growing at the same rate. This creates fundamental challenges for training neural networks within limited memory environments. In this work, we propose ActNN, a memory-efficient training framework that stores randomly quantized activations for back propagation. We prove the convergence of ActNN for general network architectures, and we characterize the impact of quantization on the convergence via an exact expression for the gradient variance. Using our theory, we propose novel mixed-precision quantization strategies that exploit the activation’s heterogeneity across feature dimensions, samples, and layers. These techniques can be readily applied to existing dynamic graph frameworks, such as PyTorch, simply by substituting the layers. We evaluate ActNN on mainstream computer vision models for classification, detection, and segmentation tasks. On all these tasks, ActNN compresses the activation to 2 bits on average, with negligible accuracy loss. ActNN reduces the memory footprint of the activation by 12x, and it enables training with a 6.6x to 14x larger batch size. Jianfei Chen 0001, Lianmin Zheng, Zhewei Yao, Dequan Wang, Ion Stoica, Michael W. Mahoney, Joseph Gonzalez 0001 |
ICML | 4 |
| 2019 | Joint Monocular 3D Vehicle Detection and TrackingabstractVehicle 3D extents and trajectories are critical cues for predicting the future location of vehicles and planning future agent ego-motion based on those predictions. In this paper, we propose a novel online framework for 3D vehicle detection and tracking from monocular videos. The framework can not only associate detections of vehicles in motion over time, but also estimate their complete 3D bounding box information from a sequence of 2D images captured on a moving platform. Our method leverages 3D box depth-ordering matching for robust instance association and utilizes 3D trajectory prediction for re-identification of occluded vehicles. We also design a motion learning module based on an LSTM for more accurate long-term motion extrapolation. Our experiments on simulation, KITTI, and Argoverse datasets show that our 3D tracking pipeline offers robust data association and tracking. On Argoverse, our image-based method is significantly better for tracking 3D vehicles within 30 meters than the LiDAR-centric baseline methods. Hou-Ning Hu, Qi-Zhi Cai, Dequan Wang, Ji Lin 0002, Min Sun 0001, Philipp Krähenbühl, Trevor Darrell, Fisher Yu 0001 |
ICCV | 3 |
| 2019 | Convolutional Neural Networks on Non-uniform Geometrical Signals Using Euclidean Spectral Transformation
Chiyu Max Jiang, Dequan Wang, Jingwei Huang 0001, Philip Marcus, Matthias Nießner |
ICLR (Poster) | 2 |
| 2019 | Deep Object-Centric Policies for Autonomous DrivingabstractWhile learning visuomotor skills in an end-to-end manner is appealing, deep neural networks are often uninterpretable and fail in surprising ways. For robotics tasks, such as autonomous driving, models that explicitly represent objects may be more robust to new scenes and provide intuitive visualizations. We describe a taxonomy of “object-centric” models which leverage both object instances and end-to-end learning. In the Grand Theft Auto V simulator, we show that object-centric models outperform object-agnostic methods in scenes with other vehicles and pedestrians, even with an imperfect detector. We also demonstrate that our architectures perform well on real-world environments by evaluating on the Berkeley DeepDrive Video dataset, where an object-centric model outperforms object-agnostic models in the low-data regimes. Dequan Wang, Coline Devin, Qi-Zhi Cai, Fisher Yu 0001, Trevor Darrell |
ICRA | 1 |
| 2019 | Monocular Plan View Networks for Autonomous DrivingabstractConvolutions on monocular dash cam videos capture spatial invariances in the image plane but do not explicitly reason about distances and depth. We propose a simple transformation of observations into a bird's eye view, also known as plan view, for end-to-end control. We detect vehicles and pedestrians in the first person view and project them into an overhead plan view. This representation provides an abstraction of the environment from which a deep network can easily deduce the positions and directions of entities. Additionally, the plan view enables us to leverage advances in 3D object detection in conjunction with deep policy learning. We evaluate our monocular plan view network on the photo-realistic Grand Theft Auto V simulator. A network using both a plan view and front view causes less than half as many collisions as previous detection-based methods and an order of magnitude fewer collisions than pure pixel-based policies. Dequan Wang, Coline Devin, Qi-Zhi Cai, Philipp Krähenbühl, Trevor Darrell |
IROS | 1 |
| 2018 | Deep Layer AggregationabstractVisual recognition requires rich representations that span levels from low to high, scales from small to large, and resolutions from fine to coarse. Even with the depth of features in a convolutional network, a layer in isolation is not enough: compounding and aggregating these representations improves inference of what and where. Architectural efforts are exploring many dimensions for network backbones, designing deeper or wider architectures, but how to best aggregate layers and blocks across a network deserves further attention. Although skip connections have been incorporated to combine layers, these connections have been "shallow" themselves, and only fuse by simple, one-step operations. We augment standard architectures with deeper aggregation to better fuse information across layers. Our deep layer aggregation structures iteratively and hierarchically merge the feature hierarchy to make networks with better accuracy and fewer parameters. Experiments across architectures and tasks show that deep layer aggregation improves recognition and resolution compared to existing branching and merging schemes. Fisher Yu 0001, Dequan Wang, Evan Shelhamer, Trevor Darrell |
CVPR | 2 |
| 2017 | Iterative object and part transfer for fine-grained recognitionabstractThe aim of fine-grained recognition is to identify sub-ordinate categories in images like different species of birds. Existing works have confirmed that, in order to capture the subtle differences across the categories, automatic localization of objects and parts is critical. Most approaches for object and part localization rehed on the bottom-up pipeline, where thousands of region proposals are generated and then filtered by pre-trained object/part models. This is computationally expensive and not scalable once the number of objects/parts becomes large. In this paper, we propose a nonparametric data-driven method for object and part localization. Given an unlabeled test image, our approach transfers annotations from a few similar images retrieved in the training set. In particular, we propose an iterative transfer strategy that gradually refine the predicted bounding boxes. Based on the located objects and parts, deep convolutional features are extracted for recognition. We evaluate our approach on the widely-used CUB200-2011 dataset and a new and large dataset called Birdsnap. On both datasets, we achieve better results than many state-of-the-art approaches, including a few using oracle (manually annotated) bounding boxes in the test images. Yu-Gang Jiang 0001, Dequan Wang, Xiangyang Xue 0001 |
ICME | 3 |
| 2016 | Enhancing semi-supervised learning through label-aware base kernels
Qiaojun Wang, Kai Zhang 0001, Zhengzhang Chen, Dequan Wang, Guofei Jiang, Ivan Marsic |
Neurocomputing | 4 |
| 2015 | Matching Reviews to Object Based on 2-Stage CRF
Qingzhong Li, Dequan Wang, Yanhui Ding, Congli Liu, Zhongmin Yan |
APWeb | 3 |
| 2015 | Weakly supervised semantic segmentation for social imagesabstractImage semantic segmentation is the task of partitioning image into several regions based on semantic concepts. In this paper, we learn a weakly supervised semantic segmentation model from social images whose labels are not pixel-level but image-level; furthermore, these labels might be noisy. We present a joint conditional random field model leveraging various contexts to address this issue. More specifically, we extract global and local features in multiple scales by convolutional neural network and topic model. Inter-label correlations are captured by visual contextual cues and label co-occurrence statistics. The label consistency between image-level and pixel-level is finally achieved by iterative refinement. Experimental results on two real-world image datasets PASCAL VOC2007 and SIFT-Flow demonstrate that the proposed approach outperforms state-of-the-art weakly supervised methods and even achieves accuracy comparable with fully supervised methods. Wei Zhang 0016, Sheng Zeng, Dequan Wang, Xiangyang Xue 0001 |
CVPR | 3 |
| 2015 | Multiple Granularity Descriptors for Fine-Grained CategorizationabstractFine-grained categorization, which aims to distinguish subordinate-level categories such as bird species or dog breeds, is an extremely challenging task. This is due to two main issues: how to localize discriminative regions for recognition and how to learn sophisticated features for representation. Neither of them is easy to handle if there is insufficient labeled data. We leverage the fact that a subordinate-level object already has other labels in its ontology tree. These "free" labels can be used to train a series of CNN-based classifiers, each specialized at one grain level. The internal representations of these networks have different region of interests, allowing the construction of multi-grained descriptors that encode informative and discriminative features covering all the grain levels. Our multiple granularity framework can be learned with the weakest supervision, requiring only image-level label and avoiding the use of labor-intensive bounding box or part annotations. Experimental results on three challenging fine-grained image datasets demonstrate that our approach outperforms state-of-the-art algorithms, including those requiring strong labels. Dequan Wang, Jie Shao 0006, Wei Zhang 0016, Xiangyang Xue 0001, Zheng Zhang 0001 |
ICCV | 1 |