VLDB 2026 Research / reviewers in the wild / expert
Yunpeng Wu
dblp:80/2135
· DBLP profile ↗
30ranked-venue papers
7as first author
15since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 11 · 3 first-author · 5 since 2021Artificial intelligence and machine learning · 9 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 5 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Computer networks · 2 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Hyperbolic Adversarial Variational Embedding for item recommendation
Zhongchuan Sun, Yunpeng Wu, Yangdong Ye |
Eng. Appl. Artif. Intell. | 5 |
| 2026 | Open-Vocabulary Prior Guided Mamba Framework for Closed-Set Nighttime Object DetectionabstractNighttime object detection is essential for road, urban, and traffic perception, where detectors are typically required to recognize a predefined set of categories under a closed-set protocol. However, severe illumination degradation, noise amplification, and weak textures make visual evidence unreliable, leading to semantic ambiguity and missed detections. To address this issue, we propose Night-Mamba-World, a knowledge-guided framework that incorporates open-vocabulary priors into closed-set nighttime object detection, thereby effectively complementing degraded visual features with language-aligned semantic knowledge. Night-Mamba-World unifies language-aligned open-vocabulary priors, frequency adaptation, illumination-guided state-space aggregation, and inference-time prompt calibration within a single detection framework. Night-Mamba-World consists of three modules. Visual Fourier Prompt Tuning (VFPT) mitigates illumination-style mismatch while preserving phase-based structural information. Retinex-Guided Mamba-PAN (RG-Mamba-PAN) enhances long-range context aggregation in dark and weak texture regions through illumination-guided state-space modeling. ModPrompt further improves robustness to varying nighttime conditions via online prompt calibration during inference. Experiments on ExDark, LLVIP, and BDD100 K demonstrate that Night-Mamba-World achieves consistent improvements across diverse nighttime scenarios, reaching 0.831$mAP_{50}$on ExDark, 0.929$mAP_{50}$on LLVIP, and 0.545$mAP_{50}$on BDD100 K. It also maintains deployment-oriented efficiency with 49.564 M parameters, 91.451 G FLOPs, 13.972 ms latency, and 71.57 FPS. Yunpeng Wu, Zhibin Du |
IEEE Signal Process. Lett. | 1 |
| 2026 | Random Dense Knowledge Distillation for Continual LearningabstractContinual Learning (CL), involving sequential training on diverse tasks, often faces catastrophic forgetting. While knowledge distillation–based approaches exhibit notable success in preventing forgetting, we pinpoint a limitation in their ability to distill the cumulative knowledge of all the previous tasks. To remedy this, we propose Random Dense Knowledge Distillation (RDKD). RDKD uses a task pool to track the model’s capabilities. It partitions the output logits of the model into dense groups, each corresponding to a task in the task pool. It then distills all tasks’ knowledge using all groups. However, using all the groups can be computationally expensive, so we also suggest random group selection in each optimization step. Moreover, we propose an adaptive weighting scheme, which balances the learning of new classes and the retention of old classes, based on the count and similarity of the classes. Our RDKD outperforms recent state-of-the-art baselines across diverse benchmarks and scenarios. Empirical analysis underscores RDKD’s ability to enhance model stability, promotes flatter minima for improved generalization, and remains robust across various memory budgets and task orders. Moreover, it seamlessly integrates with other CL methods to boost performance and proves versatile in offline scenarios like model compression. Jie Chu, Yunpeng Wu, Zenglin Shi |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2025 | AirboardNet: A UAV onboard girder inspection approach for high-speed railroad bridge using multi-task knowledge distillation☆
Yunpeng Wu, Yong Qin 0002, Fengxiang Guo, Zheda Zhao |
Adv. Eng. Informatics | 1 |
| 2025 | Automatic risk level evaluation system for potential environmental hazards along high-speed railroad using UAV aerial photograph
Fanteng Meng, Yong Qin 0002, Yunpeng Wu, Changhong Shao, Huaizhi Yang, Limin Jia 0002 |
Expert Syst. Appl. | 3 |
| 2025 | Dual global information guidance for deep contrastive multi-modal clustering
Guoliang Zou, Shizhe Hu, Tongji Chen, Yunpeng Wu, Yangdong Ye |
Inf. Sci. | 4 |
| 2025 | SRLF: Sparse Representation Learning Framework for Railroad Surrounding Potential Risk Perception Using UAV ImageryabstractRegular inspection of potential risks in railroad surroundings is essential for operational safety. Uncrewed aerial vehicles (UAVs) offer an effective solution with aerial mobility and long-distance coverage. However, existing methods struggle with rare but extremely high risks characterized by limited samples and complex feature distributions. To address this, we propose SRLF (Sparse Representation Learning Framework), which decomposes sparse risks (SR) perception into three components: capture, excavation, and learning. First, Buffer Decouple Learning (BDL) decouples objectness from classification to capture and enhance foreground perception. Second, Feature Space Dynamic Sampling (FSDS) leverages adaptive quantity sampling from multivariate Gaussian distributions to excavate discriminative SR representations. Third, Triple Similarity Loss (TSL) constructs a triple comparison mechanism to contrastively shape uncertainty surfaces between SRs and common safety hazards (CSHs). Finally, extensive experiments conducted on the UAV-based railroad surroundings dataset demonstrate that SRLF can achieve a high detection rate of CSHs (95.6% mAP) while maintaining low miss-detection rate for SRs (81.9% Recall and 0.5% FPR95). Fanteng Meng, Yong Qin 0002, Yunpeng Wu, Mingyang Chen 0001, Ninghai Qiu, Zhipeng Wang 0002, Chongchong Yu, Huaizhi Yang |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2025 | Automatic Potential Safety Hazard Evaluation System for Environment Around High-Speed Railroad Using Hybrid U-Shape Learning ArchitectureabstractPotential safety hazards (PSHs) around the high-speed railroad need to be detected and evaluated regularly and timely to ensure high-speed railroad operation safety. Unmanned aerial vehicle (UAV)-based PSH evaluation has great potential to supplement the current manual visual inspection tasks by providing better overhead views and less man-made accidents. This study presents an evaluation system for PSHs along high-speed railroad tracks. First, a novel hybrid learning architecture named UYOLO (U-shape You Only Look Once) is designed, which integrate the CSP-based backbone and detection branch to produce three scale high-level features for the object detection. Then, an innovative parsing branch inserted behind high-level layer of the structure progressively transmits context information to the shallow layer to accurately accomplish the pixel-level parsing task. Second, a new loss function using minimum point distance IoU (MPD-IoU) is designed and incorporated into the architecture to optimize the coordinate regression process of the predicted bounding boxes. Conveniently, an image-based hazard evaluation model is also developed and integrated to rate the hazard level of the detected PSHs. Finally, extensive experiments conducted on the track environment dataset established with UAV imagery indicate the proposed system can achieve a high detection rate yet remains efficient and convenient. Zheda Zhao, Yong Qin 0002, Yunpeng Wu, Wenwen Qin, Xiaolei Wu |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2024 | A subtle defect recognition method for catenary fastener in high-speed railroad using destruction and reconstruction learning
Fanteng Meng, Yong Qin 0002, Yunpeng Wu, Changhong Shao, Limin Jia 0002 |
Adv. Eng. Informatics | 3 |
| 2024 | Correlation-attention guided regression network for efficient crowd counting
Huake Wang, Qiang Guo 0012, Yunpeng Wu |
J. Vis. Commun. Image Represent. | 4 |
| 2023 | UAV imagery based potential safety hazard evaluation for high-speed railroad using Real-time instance segmentation
Yunpeng Wu, Fanteng Meng, Yong Qin 0002, Limin Jia 0002 |
Adv. Eng. Informatics | 1 |
| 2023 | Joint face completion and super-resolution using multi-scale feature relation learning
Zhilei Liu, Chenggong Zhang, Yunpeng Wu, Cuicui Zhang |
J. Vis. Commun. Image Represent. | 3 |
| 2023 | Skip Connection YOLO Architecture for Noise Barrier Defect Detection Using UAV-Based Images in High-Speed RailwayabstractNoise barriers play a critical role in reducing noise and preventing foreign object from invading railway. Noise barrier structural defects such as rusted column, deteriorated mortar layer and other damages make its structure unstable, thereby threatening seriously railway operation safety. Unfortunately, existing noise barrier inspection methods still rely heavily on manual inspection, which are low-efficiency, subjective and difficult to detect the external structure of noise barriers. To solve these problems, this study proposes an automatic inspection manner for noise barrier using UAV images, and develops a fully convolutional network (FCN)-based noise barrier defect detection approach named skip connection YOLO detection network (SCYNet), which focuses on three aspects: network structure, loss function and data augmentation. First, a skip-connected feature structure Simi-BiFPN is incorporated into the network to fully fuse the features extracted from various scale layers without adding much computational overhead. Second, a NoiseIoU loss for bounding box regression is designed to improve existing IoU-based losses and get better performance on small dataset. Thirdly, a mixed sample data augmentation method named AutoFMix is proposed to eliminate the over-fitting issue caused by excessive similarity between samples, and further improve the detection accuracy. Finally, experiments conducted on the UAV railway noise barrier dataset show that the proposed SCYNet model achieves 92.2 mAP and 78.7 FPS, respectively, which outperform other models in terms of accuracy and processing speed. The fast-processing speed and high detection accuracy can quickly turn UAV images into useful information to assist railway maintenance, thereby improving the safety of train operation. Yong Qin 0002, Yunpeng Wu, Changhong Shao, Huaizhi Yang |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2021 | Multi-scale Feature Relation Modeling for Facial Expression RestorationabstractFacial expression analysis in the wild is easily vulnerable to the quality of facial images, such as low resolution or occlusion. Existing facial image restoration studies have mostly failed to take full advantage of facial expression prior information, which leads to loss of information related to facial expression in the restoration results. In this paper, we propose a multi-scale feature relation modeling GAN (MFRM-GAN) for facial expression restoration by exploring the multi-scale property and relationship of facial action units. Based on the GAN model, the MFRM-GAN integrates the graph convolution network (GCN) for feature relation modeling and feature pyramid network (FPN) for multi-scale feature extraction. Extensive qualitative and quantitative experiments on BP4D and DISFA datasets demonstrate that our proposed MFRM-GAN (i) can conduct facial expression in-painting and facial image super-resolution jointly, (ii) can recover better facial expression details comparing with state-of-the-art method in both visual effect and AU detection task. Zhilei Liu, Yunpeng Wu, Cuicui Zhang |
IJCNN | 2 |
| 2021 | Multi-level features extraction network with gating mechanism for crowd countingabstractAbstract Crowd counting is still a practical and challenging problem owing to scale variations and information loss. Most existing methods based on the straightforward fusion of different features from a deep neural network seem to eliminate this limitation. However, these features are difficult to be fused since they often differ significantly in modality and dimensionality. Unlike previous works, a multi‐level features extraction network with gating mechanism for crowd counting is proposed. Specifically, a multi‐channel gated unit to adaptively extract features in different levels of the network is proposed, which can avoid interference from confusing information. To fully aggregate features via multi‐level fusion, multi‐level features extraction scheme is presented. The multi‐level features extraction network learns to fuse features from multiple levels and reduce false predictions. Extensive experiments and evaluations clearly illustrate that the proposed approach achieves state‐of‐the‐art counting performance against other methods on four mainstream crowd counting benchmarks. Qiang Guo 0012, Haoran Duan 0001, Yunpeng Wu |
IET Image Process. | 4 |
| 2020 | Region Based Adversarial Synthesis of Facial Action Units
Zhilei Liu, Yunpeng Wu |
MMM (2) | 3 |
| 2020 | Facial Expression Restoration Based on Improved Graph Convolutional Networks
Zhilei Liu, Yunpeng Wu, Cuicui Zhang |
MMM (2) | 3 |
| 2020 | DSPNet: Deep scale purifier network for dense crowd counting
Yunpeng Wu, Shizhe Hu, Ruobin Wang, Yangdong Ye |
Expert Syst. Appl. | 2 |
| 2020 | Densely pyramidal residual network for UAV-based railway images dehazing
Yunpeng Wu, Yong Qin 0002, Zhipeng Wang 0002 |
Neurocomputing | 1 |
| 2020 | An Improved Faster R-CNN for UAV-Based Catenary Support Device InspectionabstractThe catenary support device inspection is of crucial importance for ensuring safety and reliability of railway systems. At present, visual detection tasks of catenary support devices defect are performed by trained personnel based on the images taken periodically by industrial cameras installed on inspection vehicle in a limited period of time at midnight. However, the inspection mean is inappropriate for low efficiency and high cost. This paper presents a novel network based on unmanned aerial vehicle (UAV) images for catenary support device inspection and focuses on small object detection and the imbalanced dataset. With regards to the first aspect, based on a pyramid network structure, the improved Faster R-CNN consists of a top-down-top feature pyramid fusion structure, which heavily fuses high-level semantic information and low-level detail information. The feature map fusions of three different pooling scales are employed for improving detection accuracy of predicted bounding boxes. With regards to the second, we copy and paste the small proportion objects of dataset for avoiding category imbalance. Finally, quantitative and qualitative evaluations illustrate that the improved Faster-RCNN achieves better performance over the classic methods, yet remains convenient and efficient. Zhipeng Wang 0002, Yunpeng Wu, Yong Qin 0002, Xianbin Cao 0003 |
Int. J. Softw. Eng. Knowl. Eng. | 3 |
| 2019 | Privacy Protection Sensing Data Aggregation for Crowd Sensing
Yunpeng Wu, Shukui Zhang, Yuren Yang, Yang Zhang 0122, Li Zhang 0052, Hao Long 0001 |
WASA | 1 |
| 2019 | APL: Adversarial Pairwise Learning for Recommender Systems
Zhongchuan Sun, Bin Wu 0019, Yunpeng Wu, Yangdong Ye |
Expert Syst. Appl. | 3 |
| 2018 | Collective Density Clustering for Coherent Motion DetectionabstractCoherent motion detection remains a challenging problem due to the inherent complexity and vast diversity found in crowded scenes. Inspired by divide-and-conquer strategy, we desire to detect coherent motion from both local and global level. In this study, a novel collective density clustering (CDC) method is proposed to detect local and global coherent motion. We creatively define a collective density to discover underlying ordered density estimation, and subsequently a novel collective clustering algorithm is introduced, which is able to identify collective subgroups rapidly. Considering the complex interaction among subgroups, we present a hierarchical Union-Find-based collective merging algorithm to recognize coherent motion by merging collective subgroups. Our method is very efficient and effective. Experimental results on several challenging video datasets demonstrate that the proposed CDC achieves better results than state-of-the-art works, and multiple times or even tens of times faster. The proposed framework shows potential to be further applied to other problems (e.g., affine motion segmentation), related to local and global clustering. Yunpeng Wu, Yangdong Ye, Zenglin Shi |
IEEE Trans. Multim. | 1 |
| 2016 | Rank-based pooling for deep convolutional neural networks
Zenglin Shi, Yangdong Ye, Yunpeng Wu |
Neural Networks | 3 |
| 2015 | Coherent Motion Detection with Collective Density ClusteringabstractDetecting coherent motion is significant for analysing the crowd motion in video applications. In this study, we propose the Collective Density Clustering(CDC) approach to recognize both local and global coherent motion having arbitrary shapes and varying densities. Firstly, the collective density is defined to reveal the underlying patterns with varying levels of density. Based on collective density, the collective clustering algorithm is further presented to recognize the local consistency, where density-based clustering is more adaptive to recognize clusters with arbitrary shapes. This algorithm has salient properties including single step of clustering process, automatical decision of clustering number and accurate identification of outliers. Finally, the collective merging algorithm is introduced to fully characterize the global consistency. Experiments on diverse crowd scenes, including pedestrians, traffic and bacterial colony, demonstrate the effectiveness for coherent motion detection. The comparisons show that our approach outperforms state-of-the-art coherent detection techniques. Yunpeng Wu, Yangdong Ye |
ACM Multimedia | 1 |
| 2015 | Collective Crowd Formation Transform with Mutual Information-Based Runtime FeedbackabstractAbstract This paper introduces a new crowd formation transform approach to achieve visually pleasing group formation transition and control. Its core idea is to transform crowd formation shapes with a least effort pair assignment using the Kuhn–Munkres algorithm, discover clusters of agent subgroups using affinity propagation and Delaunay triangulation algorithms and apply subgroup‐based social force model (SFM) to the agent subgroups to achieve alignment, cohesion and collision avoidance. Meanwhile, mutual information of the dynamic crowd is used to guide agents' movement at runtime. This approach combines both macroscopic (involving least effort position assignment and clustering) and microscopic (involving SFM) controls of the crowd transformation to maximally maintain subgroups' local stability and dynamic collective behaviour, while minimizing the overall effort (i.e. travelling distance) of the agents during the transformation. Through simulation experiments and comparisons, we demonstrate that this approach is efficient and effective to generate visually pleasing and smooth transformations and outperform several existing crowd simulation approaches including reciprocal velocity avoidances, optimal reciprocal collision avoidance and OpenSteer. Yunpeng Wu, Yangdong Ye, Illés J. Farkas, Hao Jiang 0013, Zhigang Deng 0001 |
Comput. Graph. Forum | 2 |
| 2015 | miSFM: On combination of Mutual Information and Social Force Model towards simulating crowd evacuation
Yunpeng Wu, Pei Lv, Hao Jiang 0013, Mingxuan Luo, Yangdong Ye |
Neurocomputing | 2 |
| 2012 | Smooth and efficient crowd transformationabstractCrowd transformation has been an important research field due to its diverse range of applications that include film production, computer games, robotics and performance training. We propose a novel approach with continuous space and continuous time for smooth crowd transformation. Most algorithms simulating the transformation of crowds focus on the trajectories of individual participants. As a contrast, the approach we proposed here focuses on the balanced assignment to maintain the collective behavior of the full crowd. We use a bitmap-based recognition of the starting and final formations and quantify the performance of the transformation with the mutual information. We demonstrate that our method is computationally efficient for several group sizes and formations. Mingliang Xu 0001, Yunpeng Wu, Yangdong Ye |
ACM Multimedia | 2 |
| 2010 | Knowledge acquisition method from domain text based on theme logic model and artificial neural network
Yunpeng Wu, Xuening Liu, Xiaoying Gao |
Expert Syst. Appl. | 2 |
| 2010 | Graph Pattern Matching: From Intractable to Polynomial TimeabstractGraph pattern matching is typically defined in terms of subgraph isomorphism, which makes it an np-complete problem. Moreover, it requires bijective functions, which are often too restrictive to characterize patterns in emerging applications. We propose a class of graph patterns, in which an edge denotes the connectivity in a data graph within a predefined number of hops. In addition, we define matching based on a notion of bounded simulation, an extension of graph simulation. We show that with this revision, graph pattern matching can be performed in cubic-time, by providing such an algorithm. We also develop algorithms for incrementally finding matches when data graphs are updated, with performance guarantees for dag patterns. We experimentally verify that these algorithms scale well, and that the revised notion of graph pattern matching allows us to identify communities commonly found in real-world networks. Wenfei Fan, Jianzhong Li 0001, Shuai Ma 0001, Nan Tang 0001, Yinghui Wu 0001, Yunpeng Wu |
Proc. VLDB Endow. | 6 |