VLDB 2026 Research / reviewers in the wild / expert
Wufan Wang
dblp:194/9177
· DBLP profile ↗
23ranked-venue papers
8as first author
16since 2021 · last 2026
0000-0002-6838-3584ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 3 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 3 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 2 first-author · 5 since 2021Systems, architecture and hardware · 4 · 2 first-authorDatabases, data management, data science and information retrieval · 3 · 3 since 2021Computer networks · 2 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Learning to LEAP: Efficient Dense Point Tracking by Focusing Where It MattersabstractTracking Any Point (TAP) is a foundational task in computer vision with broad applicability. The state-of-the-art self-supervised TAP method leverages a global matching transformer and contrastive random walks to learn point correspondences. However, its dense all-pairs attention and correlation volume computation tend to introduce irrelevant features and produce less informative training signals, degrading both learning efficiency and tracking accuracy. To address these limitations, we introduce LEAP-Track, a self-supervised TAP approach that computes the attention matrices and correlation volume over adaptively selected sparse pairs. It consists of two core designs: (1) Curriculum-based Sparse Attention (CSA), which dynamically focuses on the most relevant keys, promoting the learning of discriminative features; and (2) Progressive k-NN Transition (PkT), which reformulates the contrastive random walk to operate on an increasingly sparse k-NN affinity graph to reinforce the learning of the most informative correspondences. By integrating the above two designs into a two-stage training paradigm, LEAP-Track is shown both theoretically and empirically to effectively boost learning efficiency, achieving superior tracking accuracy over existing self-supervised TAP methods. Chenzhi Zhao, Wufan Wang, Bo Zhang 0032, Wendong Wang 0003 |
AAAI | 2 |
| 2025 | SwinPose: A Unified Spatio-Temporal Transformer Network for Video-Based Human Pose Estimation
Ao Deng, Wufan Wang, Bo Zhang 0032, Xirong Que, Wendong Wang 0003 |
IEEE Big Data | 2 |
| 2025 | A Lightweight Real-Time Framework for Skeleton-Based Action Recognition on Mobile Devices
Qiujie Zhang, Wufan Wang, Bo Zhang 0032, Zheng Zhang 0038, Xirong Que, Wendong Wang 0003 |
IEEE Big Data | 2 |
| 2025 | Deblurring with Improved Video Diffusion Model
Haoyang Long, Bo Zhang 0032, Wufan Wang, Zheng Zhang 0038, Wendong Wang 0003 |
ICANN (2) | 3 |
| 2025 | Frame-Skeleton: A Dual-Stream Network for Action Events Sequence SpottingabstractWith the increasing popularity of golf, more and more amateurs are focusing on this sport. Deep learning based golf swing event detection becomes a meaningful task, yet current research relies on RGB images as the unique input, leading to insufficient accuracy in limited-sample scenarios. In this paper, we propose an improved dual-stream network framework, Frame-Skeleton, for detecting golf swing events. The framework combines RGB images and skeleton sequences, utilizing ST-GCN (Spatial-Temporal Graph Convolutional Network) to process joint features, significantly enhancing recognition accuracy and robustness in scenarios with limited data. Additionally, we construct a golf swing dataset, Frameflow, which contains down-the-line golf swing videos of professional and non-professional golfers, providing a new data source for research. Experimental validation shows that Frame-Skeleton outperforms traditional single-stream methods on the benchmark Golfdb dataset and demonstrates stronger generalization capabilities on the smaller Frameflow dataset. Tingyu Xie, Bo Zhang 0032, Wufan Wang, Xirong Que, Wendong Wang 0003 |
IJCNN | 3 |
| 2025 | AIGC-Enhanced UAV-Based 3D Mapping and Trajectory Planning for Rapid Disaster Response
Hui Gao 0002, Bo Zhang 0032, Kun Niu, Tan Yang, Wufan Wang, Wendong Wang 0003 |
ACM Multimedia | 7 |
| 2025 | FS-IoT: Fast Few Shot IoT Devices IdentificationabstractThe widespread deployment of Internet of Things (IoT) devices, coupled with their often limited security capabilities, has significantly increased the network attack surface. Consequently, network asset managers need to continuously monitor and assess vulnerable IoT devices to mitigate potential threats. However, existing passive IoT device identification approaches, which rely on network traffic analysis, are hindered by substantial labeling requirements and computational overhead, severely limiting their practicality in real-world environments. To address these challenges, we propose FS-IoT, a fast few-shot IoT device identification framework. FS-IoT introduces the concept of packet bursts as the fundamental unit of recognition and systematically explores their extraction, representation, classification, and aggregation for IoT device identification. Experimental evaluations on two public datasets demonstrate that FS-IoT achieves superior accuracy (99.94% and 98.15%) while requiring only 2% of the labeled training data needed by state-of-the-art methods. Furthermore, FS-IoT improves recognition speed by an order of magnitude, making it highly suitable for practical, large-scale deployments. Kunling Dai, Xirong Que, Wufan Wang |
SMC | 3 |
| 2024 | ADDG: An Adaptive Domain Generalization Framework for Cross-Plane MRI SegmentationabstractMulti-planar magnetic resonance imaging (MRI) can provide comprehensive 3D structural information for disease diagnosis. Compared to multi-source MRI, multi-planar MRI scans target areas in the human body from different directions. This atypical difference between directions may lead to poor performance of traditional domain generalization methods, especially when MRI from different planes also comes from different sources. In this paper, we propose ADDG, an Adaptive Domain Generalization framework for accurate cross-plane MRI segmentation. ADDG significantly mitigates the impact of information loss caused by slice spacing by injecting 3D shape prior to the segmentation target and capturing domain-agnostic feature differences from heterogeneous data sources through an adaptive data partitioning strategy. In addition, we propose a mesh deformation-based organ segmentation network to simultaneously delineate 2D boundary and 3D volume of organ, which could guide more accurate mesh deformation. We also develop an organ-specific mesh template and employ Loop subdivision for generating smoother 3D organ mesh. Furthermore, we design a flexible meta-learning paradigm to adaptively partition data domains based on invariant learning, which can learn domain-agnostic features from multi-source data to enhance the overall generalization ability. Experimental results show that ADDG outperforms several medical image segmentation, single-view 3D shape reconstruction, and domain generalization methods. Zibo Ma, Bo Zhang 0032, Zheng Zhang 0038, Wu Liu 0005, Wufan Wang, Hui Gao 0002, Wendong Wang 0003 |
ACM Multimedia | 5 |
| 2024 | Enhancing Low Latency Adaptive Live Streaming Through Precise Bandwidth PredictionabstractTo ensure high performance for HTTP adaptive streaming (HAS), it is critical to provide accurate prediction of end-to-end network bandwidth. Low Latency Live Streaming (LLLS), which has been gaining popularity, faces even greater challenges in this regard. Unlike Video-on-Demand (VOD) streaming, which only needs long-term bandwidth prediction and can tolerate some prediction errors, LLLS demands precise short-term bandwidth predictions. These challenges are amplified by the fact that short-term bandwidth experiences both large abrupt changes and uncertain fluctuations. Furthermore, obtaining valid bandwidth measurement samples in LLLS poses difficulties due to the on-off traffic pattern. In this work, we present DeeProphet, a system designed to enhance the performance of LLLS by achieving accurate bandwidth prediction. DeeProphet collects valid bandwidth samples by identifying intervals of packet continuous sending leveraging TCP state information, estimates the segment-level bandwidth robustly by filtering out noisy samples, and predicts both significant changes and uncertain fluctuations in future bandwidth by combining both time series and learning-based models. Experimental results demonstrate that DeeProphet effectively enhances the overall Quality of Experience (QoE) by 39.5% to 464.6% compared to state-of-the-art LLLS Adaptive Bitrate (ABR) algorithms. Bo Wang 0066, Muhan Su, Wufan Wang, Bingyang Liu, Fengyuan Ren, Mingwei Xu 0001, Jiangchuan Liu |
IEEE/ACM Trans. Netw. | 3 |
| 2023 | Revisiting Unsupervised Local Descriptor LearningabstractConstructing accurate training tuples is crucial for unsupervised local descriptor learning, yet challenging due to the absence of patch labels. The state-of-the-art approach constructs tuples with heuristic rules, which struggle to precisely depict real-world patch transformations, in spite of enabling fast model convergence. A possible solution to alleviate the problem is the clustering-based approach, which can capture realistic patch variations and learn more accurate class decision boundaries, but suffers from slow model convergence. This paper presents HybridDesc, an unsupervised approach that learns powerful local descriptor models with fast convergence speed by combining the rule-based and clustering-based approaches to construct training tuples. In addition, HybridDesc also contributes two concrete enhancing mechanisms: (1) a Differentiable Hyperparameter Search (DHS) strategy to find the optimal hyperparameter setting of the rule-based approach so as to provide accurate prior for the clustering-based approach, (2) an On-Demand Clustering (ODC) method to reduce the clustering overhead of the clustering-based approach without eroding its advantage. Extensive experimental results show that HybridDesc can efficiently learn local descriptors that surpass existing unsupervised local descriptors and even rival competitive supervised ones. Wufan Wang, Lei Zhang 0021, Hua Huang 0001 |
AAAI | 1 |
| 2023 | Semi-direct Sparse Odometry with Robust and Accurate Pose Estimation for Dynamic Scenes
Wufan Wang, Lei Zhang 0021 |
CAD/Graphics | 1 |
| 2023 | DeeProphet: Improving HTTP Adaptive Streaming for Low Latency Live Video by Meticulous Bandwidth PredictionabstractThe performance of HTTP adaptive streaming (HAS) depends heavily on the prediction of end-to-end network bandwidth. The increasingly popular low latency live streaming (LLLS) faces greater challenges since it requires accurate, short-term bandwidth prediction, compared with VOD streaming which needs long-term bandwidth prediction and has good tolerance against prediction error. Part of the challenges comes from the fact that short-term bandwidth experiences both large abrupt changes and uncertain fluctuations. Additionally, it is hard to obtain valid bandwidth measurement samples in LLLS due to its inter-chunk and intra-chunk sending idleness. In this work, we present DeeProphet, a system for accurate bandwidth prediction in LLLS to improve the performance of HAS. DeeProphet overcomes the above challenges by collecting valid measurement samples using fine-grained TCP state information to identify the packet bursting intervals, and by combining the time series model and learning-based model to predict both large change and uncertain fluctuations. Experiment results show that DeeProphet improves the overall QoE by 17.7%-359.2% compared with state-of-the-art LLLS ABR algorithms, and reduces the median bandwidth prediction error to 2.7%. Bo Wang 0066, Wufan Wang, Fengyuan Ren |
WWW | 3 |
| 2022 | Progressive Unsupervised Learning of Local DescriptorsabstractTraining tuple construction is a crucial step in unsupervised local descriptor learning. Existing approaches perform this step relying on heuristics, which suffer from inaccurate supervision signals and struggle to achieve the desired performance. To address the problem, this work presents DescPro, an unsupervised approach that progressively explores both accurate and informative training tuples for model optimization without using heuristics. Specifically, DescPro consists of a Robust Cluster Assignment (RCA) method to infer pairwise relationships by clustering reliable samples with the increasingly powerful CNN model, and a Similarity-weighted Positive Sampling (SPS) strategy to select informative positive pairs for training tuple construction. Extensive experimental results show that, with the collaboration of the above two modules, DescPro can outperform state-of-the-art unsupervised local descriptors and even rival competitive supervised ones on standard benchmarks. Wufan Wang, Lei Zhang 0021, Hua Huang 0001 |
ACM Multimedia | 1 |
| 2022 | Sensitivity-Aware Spatial Quality Adaptation for Live Video AnalyticsabstractTo address the conflict between the limited network bandwidth and high DNN inference accuracy, live video analytics desires a bandwidth-efficient streaming approach. To this end, more and more works study spatially variable quality streaming where high quality is only used for important regions. The key challenges are to accurately identify the important regions and select the right qualities for them to maximize accuracy. Existing approaches use either cheap analytics models or low-quality videos to locate important regions, and employ heuristic rules to make quality decisions, which struggle to address the above challenges. Our key insight is that the region’s accuracy “sensitivity” obtained by running the expensive DNN model on the high-quality video provides a reliable indication of the region’s importance and allows to allocate the available bandwidth optimally over regions by explicitly maximizing the frame accuracy. This work presents a sensitivity-aware algorithm Orchestra, which incorporates sensitivity into the design of spatial quality adaptation, including video zoning and quality selection. The design of Orchestra entails three main contributions: a feasible way of sensitivity estimation, sensitivity-aware zoning, and deduction-based accuracy estimation. Extensive experiments over realistic videos and network traces show that Orchestra improves accuracy by upto 14.1% with comparable bandwidth usage or reduces bandwidth usage by upto 44.2% while maintaining higher accuracy compared to baselines. Wufan Wang, Lei Zhang 0021, Hua Huang 0001 |
IEEE J. Sel. Areas Commun. | 1 |
| 2021 | Surface-to-air missile sites detection agent with remote sensing images
Jihong Zhu 0001, Wufan Wang, Minchi Kuang |
Sci. China Inf. Sci. | 3 |
| 2021 | Loop Closure Detection by Using Global and Local Features With Photometric and Viewpoint InvarianceabstractLoop closure detection plays an important role in many Simultaneous Localization and Mapping (SLAM) systems, while the main challenge lies in the photometric and viewpoint variance. This paper presents a novel loop closure detection algorithm that is more robust to the variance by using both global and local features. Specifically, the global feature with the consolidation of photometric and viewpoint invariance is learned by a Siamese Network from the intensity, depth, gradient and normal vectors distribution. The local feature with rotation invariance is based on the histogram of relative pixel intensity and geometric information like curvature and coplanarity. Then, these two types of features are jointly leveraged for the robust detection of loop closures. The extensive experiments have been conducted on the publicly available RGB-D benchmark datasets like TUM and KITTI. The results demonstrate that our algorithm can effectively address challenging scenarios with large photometric and viewpoint variance, which outperforms other state-of-the-art methods. Mingfei Yu, Lei Zhang 0021, Wufan Wang, Hua Huang 0001 |
IEEE Trans. Image Process. | 3 |
| 2020 | VALID: A Comprehensive Virtual Aerial Image DatasetabstractAerial imagery plays an important role in land-use planning, population analysis, precision agriculture, and unmanned aerial vehicle tasks. However, existing aerial image datasets generally suffer from the problem of inaccurate labeling, single ground truth type, and few category numbers. In this work, we implement a simulator that can simultaneously acquire diverse visual ground truth data in the virtual environment. Based on that, we collect a comprehensive Virtual AeriaL Image Dataset named VALID, consisting of 6690 high-resolution images, all annotated with panoptic segmentation on 30 categories, object detection with oriented bounding box, and binocular depth maps, collected in 6 different virtual scenes and 5 various ambient conditions (sunny, dusk, night, snow and fog). To our knowledge, VALID is the first aerial image dataset that can provide panoptic level segmentation and complete dense depth maps. We analyze the characteristics of VALID and evaluate state-of-the-art methods for multiple tasks to provide reference baselines. The experiment results demonstrate that VALID is well presented and challenging. The dataset is available at https://sites.google.com/view/valid-dataset/. Lyujie Chen, Wufan Wang, Xiaming Yuan, Jihong Zhu 0001 |
ICRA | 4 |
| 2019 | Design and hovering control of a twin rotor tail-sitter UAV
Wufan Wang, Jihong Zhu 0001, Minchi Kuang, Xiaming Yuan, Yunfei Tang, Yaqing Lai, Lyujie Chen, Yunjie Yang 0002 |
Sci. China Inf. Sci. | 1 |
| 2018 | Adaptive Attitude Control for a Tail-Sitter UAV with Single Thrust-Vectored PropellerabstractTail-sitter unmanned aerial vehicles (UAVs) have gained extensive popularity in recent years due to their inherent advantages of both fixed wing and rotary wing UAVs. However, these advantages are accompanied with control challenges because of two different flight regimes and drastically changing dynamics during transition flights. This paper focuses on the design of a unified controller free from cumbersome controller switchings and applicable in all attitude range for a tail-sitter with single thrust-vectored propeller. To achieve this, both thrust vectoring model and full-regime aerodynamics model are built first, after which a complete attitude dynamics model of the tail-sitter is established utilizing the quaternion attitude description to avoid the singularity problem. An adaptive controller is then derived based on a simplified model using the Lyapunov stability theory with unknown system parameters identified online by forgetting factor recursive least square (FF-RLS) method. Flight experiments are conducted to demonstrate the feasibility and effectiveness of the proposed control scheme. Wufan Wang, Jihong Zhu 0001, Minchi Kuang, Xufei Zhu |
ICRA | 1 |
| 2018 | Learning Transferable UAV for Forest Visual PerceptionabstractIn this paper, we propose a new pipeline of training a monocular UAV to fly a collision-free trajectory along the dense forest trail. As gathering high-precision images in the real world is expensive and the off-the-shelf dataset has some deficiencies, we collect a new dense forest trail dataset in a variety of simulated environment in Unreal Engine. Then we formulate visual perception of forests as a classification problem. A ResNet-18 model is trained to decide the moving direction frame by frame. To transfer the learned strategy to the real world, we construct a ResNet-18 adaptation model via multi-kernel maximum mean discrepancies to leverage the relevant labelled data and alleviate the discrepancy between simulated and real environment. Simulation and real-world flight with a variety of appearance and environment changes are both tested. The ResNet-18 adaptation and its variant model achieve the best result of 84.08% accuracy in reality. Lyujie Chen, Wufan Wang, Jihong Zhu 0001 |
IJCAI | 2 |
| 2017 | Flight controller design and demonstration of a thrust-vectored tailsitterabstractThis paper discusses the design and control methods of a thrust-vectored tailsitter that combines the advantages of both fixed wing and rotary wing systems. Separable takeoff bracket and controllable forward landing are implemented to reduce the flight weight and mitigate the effects of crosswinds. A six-degrees-of-freedom model especially for this tailsitter is then proposed to describe the dynamics of the whole system. Attitude representation based on horizontal /vertical Euler angles is presented to avoid the problem of singularity. Attitude and altitude controllers that switch between horizontal and vertical modes are used. In these controllers linear/constant acceleration approximation and filtered feed-forward acceleration algorithm are implemented. Effectiveness and reliability of the proposed control methods are demonstrated and evaluated by experimental results of the whole flight envelope. Minchi Kuang, Jihong Zhu 0001, Wufan Wang, Yunfei Tang |
ICRA | 3 |
| 2017 | Design, modelling and hovering control of a tail-sitter with single thrust-vectored propellerabstractThis paper focuses on the design, modelling and hovering control of a tail-sitter with single thrust-vectored propeller which possesses the inherent advantages of both fixed wing and rotary wing unmanned aerial vehicles (UAVs). The developed tail-sitter requires only the same number of actuators as a normal fixed wing aircraft and achieves attitude control through deflections of the thrust-vectored propeller and ailerons. Thrust vectoring is realized by mounting a simple gimbal mechanism beneath the propeller motor. Both the thrust vector model and aerodynamics model are established, which leads to a complete nonlinear model of the tail-sitter in hovering state. Quaternion is applied for attitude description to avoid the singularity problem and improve computation efficiency. Through reasonable assumptions, a simplified model of the tail-sitter is obtained, based on which a backstepping controller is designed using the Lyapunov stability theory. Experimental results are presented to demonstrate the effectiveness of the proposed control scheme. Wufan Wang, Jihong Zhu 0001, Minchi Kuang |
IROS | 1 |
| 2016 | Optimal state estimation for sampled-data systems with randomly sampled and delayed measurementsabstractThe optimal state estimation problem for sampled-data systems with randomly sampled and delayed measurements is addressed in this paper. An optimal filter is presented first for the sampled-data system with randomly sampled and delay-free measurements from multiple sensors. The filter, which has been proved to be optimal in the sense of minimum estimation variance, updates the state estimation once new measurements are available. The result is applicable to a wide range of sampling cases of which the corresponding state estimation procedures are formulated separately. Discrete-time equivalent of the filter is also derived rigorously which makes it feasible for computer implementation. Furthermore, we extend the optimal filter and develop a sliding time window estimator through the measurement reorganization technique to deal with the situation of delayed measurements. Monte-Carlo simulations are carried out to demonstrate the effectiveness of the proposed approach. Wufan Wang, Xiaming Yuan, Jihong Zhu 0001 |
SMC | 1 |