VLDB 2026 Research / reviewers in the wild / expert
Wenrui Ding
dblp:08/10143
· DBLP profile ↗
50ranked-venue papers
0as first author
39since 2021 · last 2026
0000-0001-5490-4724ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 32 · 24 since 2021Artificial intelligence and machine learning · 21 · 17 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 4 since 2021Computer networks · 4 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Progressive Volume Distillation with Active Learning for Efficient NeRF Architecture Conversion
Shuangkang Fang, Yufeng Wang 0004, Yi Yang 0033, Wenrui Ding, Shuchang Zhou 0001 |
Int. J. Comput. Vis. | 6 |
| 2026 | Cross-attention relation network based on metric learning for few-shot specific emitter identification
Zeqi Shao, Wenrui Ding, Duona Zhang, Yikai Zheng |
Pattern Recognit. | 2 |
| 2026 | APDMs: Adversarial purification diffusion models for automatic modulation classification
Wenrui Ding, Duona Zhang, Zeqi Shao, Baihe Chen |
Signal Process. | 2 |
| 2026 | RAMP: Robust Adaptive Multimodal Perception for UAV Classification and 3D Localization
Henghui Guo, Qinglong Jiang, Wenrui Ding |
IEEE Signal Process. Lett. | 6 |
| 2026 | Editing 3D Scenes via Text Prompts Without RetrainingabstractNumerous diffusion models have been developed for 2D image synthesis and editing, and recently they are extended to 3D scene editing tasks. However, editing 3D scenes is still in its early stages, and the challenges of scene representations and multi-view consistency need to be addressed. A notable limitation of existing approaches is the need for specific modules for different edits and model retraining for each scene. To tackle these issues, we propose a novel and versatile text-driven 3D scene editing method, termed DN2N, which allows for the direct acquisition of the editing results without the requirement for retraining. Our method employs off-the-shelf text-based editing models of 2D images to modify the multi-view images of a 3D scene. A content filtering process is then applied to discard poorly edited images that disrupt 3D consistency. We consider the remaining inconsistency as a problem of removing noise perturbations and solve it by generating data with similar perturbation characteristics for training. We develop a versatile NeRF model structure and propose two novel cross-view regularization terms to help the DN2N mitigate these perturbations. Empirical results show that our method achieves multiple editing types based solely on text prompts, including but not limited to appearance editing, weather transition, object changing, and style transfer. Most importantly, DN2N exhibits a versatility of editing capabilities, eliminating the need to customize or retrain editing models for specific scenes or editing types. Namely, DN2N achieves comparable total editing time to the 3DGS-based editing method, enhancing its practical value. Shuangkang Fang, Yufeng Wang 0004, Yi-Hsuan Tsai, Wenrui Ding, Yi Yang 0033, Shuchang Zhou 0001, Ming-Hsuan Yang 0001 |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2025 | Open-world Radio Frequency Fingerprint Identification via Augmented Semi-supervised LearningabstractIn complex electromagnetic environments, the identification and differentiation of diverse radio frequency (RF) emitters become particularly crucial. Existing RF fingerprinting methods demonstrate limitations when dealing with numerous unknown emitters, making it challenging for accurate classification and recognition. These limitations hinder the effective handling of specific unknown emitters.To address this issue, we introduce a novel RF fingerprinting method suitable for open-world conditions for the first time. We develop a novel RF fingerprinting model, Roinformer, to extract signal features with positional attention. We then leverage data augmentation strategies such as noise jitter and signal frame rearrangement to construct an effective pre-training model. Moreover, by incorporating instance-level similarity loss and a novel local entropy regularization approach, we significantly enhance the accuracy of known class identification and mitigate the catastrophic forgetting of known signal samples. Experimental results on three temporal signal datasets demonstrate that our method effectively recognizes both the known and unknown classes, outperforming several state-of-the-art methods by a large margin. Zehua Han, Qirui Zhao, Zhexuan Cui, Yufeng Wang 0004, Duona Zhang, Wenrui Ding |
AAAI | 7 |
| 2025 | Graph Structure Refinement with Energy-based Contrastive LearningabstractGraph Neural Networks (GNNs) have recently gained widespread attention as a successful tool for analyzing graph-structured data. However, imperfect graph structure with noisy links lacks enough robustness and may damage graph representations, therefore limiting the GNNs' performance in practical tasks. Moreover, existing generative architectures fail to fit discriminative graph-related tasks. To tackle these issues, we introduce an unsupervised method based on a joint of generative training and discriminative training to learn graph structure and representation, aiming to improve the discriminative performance of generative models. We propose an Energy-based Contrastive Learning (ECL) guided Graph Structure Refinement (GSR) framework, denoted as ECL-GSR. To our knowledge, this is the first work to combine energy-based models with contrastive learning for GSR. Specifically, we leverage ECL to approximate the joint distribution of sample pairs, which increases the similarity between representations of positive pairs while reducing the similarity between negative ones. Refined structure is produced by augmenting and removing edges according to the similarity metrics among node representations. Extensive experiments demonstrate that ECL-GSR outperforms the state-of-the-art on eight benchmark datasets in node classification. ECL-GSR achieves faster training with fewer samples and memories against the leading baseline, highlighting its simplicity and efficiency in downstream tasks. Xianlin Zeng, Yufeng Wang 0004, Guodong Guo, Wenrui Ding, Baochang Zhang 0001 |
AAAI | 5 |
| 2025 | Cellular-Connected UAV Path Planning on Radio Maps: LLMs Enhanced ApproachabstractIn this paper, we propose a large language models (LLMs) enhanced path planning framework for a cellular-connected unmanned aerial vehicle (UAV), which can significantly improve the efficiency of path planning without sacrificing path quality. The ray tracing model is used to generate the radio map of UAV cellular connectivity from real-world geospatial data. By introducing quadtree-assisted representation, the radio map is formatted into prompts that can be understood by LLMs, bridging the gap between simplified obstacle environments and realistic electromagnetic environments. Meanwhile, a combined approach using LLMs and the graph-based search method is employed to find a path which can guarantee reliable connections between the UAV and its associated ground base station throughout the flight. Simulation results demonstrate that our proposed LLM-enhanced framework not only finds the path that guarantees communication quality, but also addresses computational and memory limitations of the conventional method, with a significant reduction in operations and storage. Xuetong Pei, Fuyuan Ma, Shutong Wang, Zesheng Wang 0002, Wenrui Ding, Yufeng Wang 0004 |
GLOBECOM | 5 |
| 2025 | ASFC-NeRF: Large-Scale Scene Rendering with Adaptive Sampling and Feature-aware CompressionabstractWhile significant progress has been made in large-scale scene representation using Neural Radiance Fields (NeRF), several limitations remain. For instance, most methods still rely on the original coarse-to-fine sampling strategy, leading to an inefficient rendering process. Additionally, to model larger scenes, these methods often use complex network models, resulting in redundant model parameters. To address these issues, we propose a novel model with adaptive sampling and feature-aware compression for large-scale scene rendering, named ASFC-NeRF. We first introduce a weight prediction network to replace the original coarse sampling strategy, then employ a teacher network and depth constraints for knowledge distillation in the early stages of training to enhance the high-fidelity of the scene. Furthermore, we optimize the number of Grids and the channels of Planes and prune the network to efficiently compress model parameters. Experimental results demonstrate that our method significantly accelerates the rendering process and greatly reduces parameter quantity while maintaining or only slightly lowering image quality. Therefore, ASFC-NeRF exhibits advantages in comprehensive performance and practicality. Yufeng Wang 0004, Shuangkang Fang, Zesheng Wang 0002, Dacheng Qi, Wenrui Ding |
ICASSP | 7 |
| 2025 | NeRF is a Valuable Assistant for 3D Gaussian SplattingabstractWe introduce NeRF-GS, a novel framework that jointly optimizes Neural Radiance Fields (NeRF) and 3D Gaussian Splatting (3DGS). This framework leverages the inherent continuous spatial representation of NeRF to mitigate several limitations of 3DGS, including sensitivity to Gaussian initialization, limited spatial awareness, and weak inter-Gaussian correlations, thereby enhancing its performance. In NeRF-GS, we revisit the design of 3DGS and progressively align its spatial features with NeRF, enabling both representations to be optimized within the same scene through shared 3D spatial information. We further address the formal distinctions between the two approaches by optimizing residual vectors for both implicit features and Gaussian positions to enhance the personalized capabilities of 3DGS. Experimental results on benchmark datasets show that NeRF-GS surpasses existing methods and achieves state-of-the-art performance. This outcome confirms that NeRF and 3DGS are complementary rather than competing, offering new insights into hybrid approaches that combine 3DGS and NeRF for efficient 3D scene representation. Shuangkang Fang, I-Chao Shen, Takeo Igarashi, Yufeng Wang 0004, Zesheng Wang 0002, Yi Yang 0033, Wenrui Ding, Shuchang Zhou 0001 |
ICCV | 7 |
| 2025 | MeshLLM: Empowering Large Language Models to Progressively Understand and Generate 3D Mesh
Shuangkang Fang, I-Chao Shen, Yufeng Wang 0004, Yi-Hsuan Tsai, Yi Yang 0033, Shuchang Zhou 0001, Wenrui Ding, Takeo Igarashi, Ming-Hsuan Yang 0001 |
ICCV | 7 |
| 2025 | Guiding Yourself with Your Own Insights: Student-Driven Knowledge DistillationabstractKnowledge distillation (KD) stands as an efficient technique for compressing models, typically employing a teacher-student framework. Nevertheless, optimizing KD to yield models with reduced parameters and enhanced performance remains an area warranting deeper investigation. In this article, we recognize the significance of preserving structural consistency to enhance knowledge transfer efficiency between networks. Leveraging this insight, we introduce a novel approach termed Student-Driven Knowledge Distillation (SDKD), which integrates a proxy teacher intermediary between the primary teacher and student model. Specifically, we construct the architecture of the proxy teacher entirely based on the student network to generate logits that closely align with the distribution of the student network. Besides, we propose a Feature Fusion Block (FFB) to integrate features from the teacher network into the proxy teacher. FFB can not only provide high-quality feature-based knowledge for distillation but also impart response-based knowledge to facilitate the learning process. Extensive experiments illustrate that SDKD outperforms 29 state-of-the-art methods on several tasks, including image classification, semantic segmentation, and depth estimation. Dacheng Qi, Yufeng Wang 0004, Shuangkang Fang, Zehao Zhang, Zesheng Wang 0002, Wenrui Ding |
ICME | 7 |
| 2025 | SANE: Enhancing Large-scale Scene Representation with Semantic-aware NeRF ExpertsabstractWe propose the Semantic-aware NeRF Experts (SANE), which fully exploits the intrinsic characteristics of large-scale scenes, including semantics and material features, to achieve high-quality novel view synthesis results and provide accurate 3D semantic information. SANE begins by building a semantic Mixture of Experts (MoE), utilizing a learnable gating network to semantically partition the scene into blocks for corresponding NeRF experts. We then develop a semantic volume rendering scheme that integrates discrete semantics into the end-to-end differentiable process of NeRF, enabling refined semantic labeling of each scene point. Additionally, we implement a dual-implicit encoding strategy: intra-block encoding captures lighting variations across viewpoints, while inter-block one captures texture features among different semantic objects. Experiments on benchmark datasets show that SANE delivers higher-quality scene representations and effective semantic decomposition for downstream tasks, such as precise editing of large-scale scenes based on semantics. Zesheng Wang 0002, Yufeng Wang 0004, Shuangkang Fang, Dacheng Qi, Shengxi Li, Mai Xu, Wenrui Ding |
ICME | 8 |
| 2025 | Measurement-Based Channel Characterization of Air-to-Ground Propagation for UAV in Coastal BeachesabstractIn this paper, a miniaturization-and-lightweighting air-to-ground (A2G) channel measurement system is developed based on the software-defined radio (SDR) platform. Utilizing the system, low-altitude A2G channel measurements are conducted in a coastal beaches environment over a 1-kilometer flight distance at the carrier frequency of 1.5 GHz. The collected measurement data is statistically analyzed, and a fitted channel model is proposed. The fitted results are then compared with the free space path loss (FSPL) model and the flat earth two ray (FETR) model, which considers the electromagnetic characteristics of ground reflection points. The error analysis is performed from the perspectives of link distance and the pitch angle between the UAV and the ground station. The results demonstrate that the fitted model provides greater reliability than the FSPL model and the FETR model. Moreover, in the medium pitch angle region, the two-ray model that accounts for actual electromagnetic characteristics exhibits higher accuracy. Fuyuan Ma, Wenrui Ding, Yufeng Wang 0004, Shutong Wang, Xuetong Pei |
VTC2025-Spring | 2 |
| 2025 | Frequency-learning adversarial networks based on transfer learning for cross-scenario signal modulation classificationabstractAutomatic modulation classification (AMC) serves a challenging yet crucial role in wireless communications. Despite deep learning-based approaches being widely used in signal processing, they are challenged by signal distribution variations, especially in various channel conditions. In this paper, we introduce an adversarial transfer framework named frequency-learning adversarial networks (FLANs) based on transfer learning for cross-scenario signal classification. This method uses the stability in the frequency spectrum by introducing a frequency adaptation (FA) technique to incorporate target channel information into source-domain signals. To address the unpredictable interference in the channel, a fitting channel adaptation (FCA) module is used to reduce the difference between the source and target domains caused by variations in the channel environment. Experimental results illustrate that FLANs outperforms state-of-the-art transfer approaches, demonstrating an improved top-1 classification accuracy by about 5.2 percentage points in high signal-to-noise ratio (SNR) scenes on a cross-scenario real collected dataset CSRC2023. Qinyan Ma, Zeqi Shao, Duona Zhang, Yufeng Wang 0004, Wenrui Ding |
Frontiers Inf. Technol. Electron. Eng. | 6 |
| 2025 | Arch-Net: Model conversion and quantization for architecture agnostic model deployment
Shuangkang Fang, Zipeng Feng, Song Yuan, Yufeng Wang 0004, Yi Yang 0033, Wenrui Ding, Shuchang Zhou 0001 |
Neural Networks | 7 |
| 2025 | Finite-Time Robust Distributed Estimate for Nonlinear Systems With Heterogeneous SensorsabstractThis article proposes a finite-time distributed state estimation (DSE) algorithm for discrete-time stochastic nonlinear systems with heterogeneous sensors. Considering the network with heterogeneous sensors, the distributed estimate framework is designed by three phases, namely, priori prediction, measurement update, and consensus fusion. To obtain the accurate priori prediction results, the interactive multiple model (IMM) method is adopted to calculate the priori state value in the priori prediction phase. By introducing the measurement probability matrix, a novel heterogeneous measurement information fusion algorithm is designed. Then the measurement information of each sensor is used to update the priori prediction estimates to calculate the estimate results in the measurement update phase. Based on the consensus method, the estimate results of each sensor are fused with consensus weight to calculate the distributed state estimates of nonlinear systems in the consensus fusion phase. Besides, with finite consensus fusion steps, the bounds of the proposed distributed estimate algorithm are proved to be existed. Finally, distributed state estimate simulation example for nonlinear system is set to validate the performance. Zheng Zhang 0032, Xiwang Dong, Wenrui Ding, Zhang Ren |
IEEE Trans. Cybern. | 3 |
| 2025 | Binary Lightweight Neural Networks for Arbitrary Scale Super-Resolution of Remote Sensing ImagesabstractSuper-resolution (SR) of remote sensing images (RSIs) has been improved significantly with the development of deep learning. However, better performances usually come from complex network architectures and require a substantial number of parameters. Moreover, many methods can only deal with SR of single and fixed-scale factors. As such, we propose a binary lightweight SR (BLiSR) method to decrease the computation and storage burden and increase the practicality of SR, where we employ a binary neural network (BNN) as the backbone and leverage a binary continuous up-sampling module (BCUM) to achieve arbitrary scale RSI SR. Specifically, we introduce an adaptive binary convolution (ABConv) as the basic unit of BLiSR, which can adaptively adjust the learnable parameters to fit the distribution of full-precision weights and activations. Then, a scalable hyperbolic tangent function is presented to approximate the Sign function in backpropagation and increase the learning capability of BNN. Furthermore, we design a lightweight SR network that considers the full-precision information flow of BNN. The network comprises several basic binary units and a multilayer group fusion block (MGFB), which can extract and fuse the multilevel information from LR images, respectively. Finally, BCUM can predict the pixel values of HR images based on the frequency implicit representation network (IRN) and reconstruct the LR images at arbitrary scales. Extensive experiments on four RSI datasets demonstrate that the proposed BLiSR is superior to several lightweight state-of-the-art (SOTA) methods on both fixed and arbitrary scale SR settings, with a better balance of complexity and performance. Yufeng Wang 0004, Xianlin Zeng, Wei Li 0022, Wenrui Ding |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2025 | HALO: Hierarchical Adaptive Feature Learning and Cross-View Interaction for Object-Level Geo-LocalizationabstractCross-view object localization (CVOGL) offers fine-grained geographic information regarding the specified object of interest, in comparison to typical cross-view geo-location at the image level. However, it faces more challenges, primarily due to significant visual appearance changes induced by viewpoint discrepancies and the difficulty of accurately identifying corresponding objects within reference images containing multiple targets. To address these challenges, we propose a novel CVOGL model termed HALO. First, we introduce a cascade cross-view feature interaction module that enables effective local-to-global feature fusion across different viewpoints, thereby enhancing the feature representation of objects in the reference image. Second, to mitigate the scale variations and feature distribution discrepancies between the query and reference images, we propose an adaptive feature hierarchization and aggregate module. It extracts more representative global descriptors by adaptively hierarchizing and aggregating features. Lastly, utilizing the learned global descriptors, we perform cross-view CL and introduce a hard sample mining strategy to further enhance the discriminative ability of the network. Through these advancements, we significantly improve the discriminative representations of the query objects at both views. Extensive experiments on the CVOGL dataset demonstrate the effectiveness and robustness of the proposed method. In the CVOGL task of drone→satellite, HALO improves 2.36% in [email protected] on the test sets. Similarly, HALO shows an increase of 2.49% in [email protected] on the validation set for the CVOGL task of ground→satellite. Our codes are publicly available at https://github.com/ZehaoZhang-Uestc/HALO. Zehao Zhang, Lei Ding 0008, Yufeng Wang 0004, Wenrui Ding |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2024 | Efficient Implicit SDF and Color Reconstruction via Shared Feature Field
Shuangkang Fang, Dacheng Qi, Yufeng Wang 0004, Zehao Zhang, Zeqi Shao, Wenrui Ding |
ACCV (10) | 9 |
| 2024 | Chat-Edit-3D: Interactive 3D Scene Editing via Text Prompts
Shuangkang Fang, Yufeng Wang 0004, Yi-Hsuan Tsai, Yi Yang 0033, Wenrui Ding, Shuchang Zhou 0001, Ming-Hsuan Yang 0001 |
ECCV (42) | 5 |
| 2024 | Progressive prediction: Video anomaly detection via multi-grained predictionabstractAbstract Video Anomaly Detection (VAD) has been an active research field for several decades. However, most existing approaches merely extract a single type of feature from videos and define a single paradigm to indicate the extent of abnormalities. A coarse‐to‐fine three‐level prediction is built by integrating different levels of spatio‐temporal representations, better highlighting the difference between normal and abnormal behaviors. First, an object‐level trajectory prediction is proposed to model human historical position using a graph transformer network. Subsequently, skeleton‐level prediction is achieved by incorporating the positional information from the trajectory prediction. More importantly, based on the predicted skeleton, a skeleton‐guided pixel‐level region prediction is performed. A novel Skeleton Conditioned Generative Adversarial Network (SCGAN) is designed to explore the correlation between skeleton‐level and pixel‐level motion prediction. Benefiting from SCGAN, the prediction of human regions is contributed by both coarse‐grained and fine‐grained motion features. This three‐level prediction, namely Progressive Prediction Video Anomaly Detection (P 3 VAD), enlarges the prediction error on irregular motion patterns. Besides, a pixel‐level analysis method is proposed to achieve Background‐bias Elimination (BE) and denoise the predicted region. Experimental results validate the effectiveness of P 3 VAD on the four benchmark datasets (ShanghaiTech, CUHK Avenue, IITB‐Corridor, and ADOC). Xianlin Zeng, Yalong Jiang, Yufeng Wang 0004, Wenrui Ding |
IET Image Process. | 5 |
| 2024 | Knowledge Embedding Networks Based on Deep Learning for Automatic Modulation Classification in Cognitive RadioabstractDeep learning has shown remarkable success in cognitive radio. However, popular approaches mainly focus on the purely data-driven architecture design, and fail to explore the professional knowledge of wireless communication which is particularly significant for radio signal essential feature extraction. Inspired by digital signal processing theories, we propose knowledge embedding networks (KENs) to introduce the high-order statistics and radio spectral information into deep neural networks, based on the high-order convolution and multi-spectral attention mechanism. KENs are the general form of attention mechanism from frequency perspective and leverage identical structures of the conventional convolution layer for high-order statistics. To extract efficient representations of radio signal, frequency loss with generative adversarial networks is proposed to enhance the discrimination and richness of attention knowledge in an explicit manner. KENs enhance interpretability concerning essential radio properties while achieving the state-of-the-art accuracy by a significant margin compared to traditional automatic modulation classification approaches on RADIOML 2018.01A dataset. Duona Zhang, Yuanyao Lu, Wenrui Ding, Yundong Li |
IEEE Trans. Commun. | 3 |
| 2024 | Reasonable Anomaly Detection Based on Long-Term Sequence ModelingabstractVideo anomaly detection is a challenging task due to the unpredictable nature of abnormal actions, sophisticated semantics and a lack in training data. The visual representations of most existing approaches are limited by short-term sequences which cannot provide necessary clues for achieving reasonable detections. In this paper, we propose to comprehensively represent the motion patterns in human actions by learning from long-term sequences. Firstly, a Stacked State Machine (SSM) model with distinctive basis functions is proposed to represent the temporal dependencies which are consistent across long-term observations. Secondly, the dependencies are leveraged in filtering out problematic motion estimations which are influenced by short-term observation noises, plausible motion parameters are obtained in this way. Finally, SSM model predicts future states based on past ones, the divergence between the predictions with inherent normal patterns and observed ones determines anomalies which violate normal motion patterns. To address the challenges in drone-based surveillance, a dataset which is more diversified than existing ones is built. Extensive experiments are carried out to evaluate the proposed approach on the dataset and existing ones. Improvements over state-of-the-art methods can be observed. The proposed dataset will be made publicly available. Code is available athttps://github.com/AllenYLJiang/Anomaly-Detection-in-Sequences. Yalong Jiang, Changkang Li, Wenrui Ding, Jinzhi Xiang, Zheru Chi |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | UAV-ENeRF: Text-Driven UAV Scene Editing With Neural Radiance Fieldsabstract3D reconstruction of Unmanned Aerial Vehicle (UAV) scenes is vital for agriculture, environmental protection, urban planning, and disaster response, to name a few. However, data acquisition can be constrained and hazardous under hostile environments, which limits the image data available in real-world applications. In this work, we propose a text-driven online editing framework for UAV scenes, which can generate novel views of existing scenes with abundant editing types. Compared with small single-object scenes, large-scale UAV scene editing suffers from several particular challenges: 1) broader capturing scope exhibits illumination variation and complicated objects that reduce the 3D scene consistency after editing; and 2) high-resolution 2D editing and 3D reconstruction can be computationally expensive with tremendous GPU memory. To tackle these issues, we first design a dual-branch compact NeRF structure to reduce memory usage and enhance accuracy for 3D reconstruction. We then introduce a sub-pixel sampling scheme to expedite the generation of low-resolution images for 2D editing, followed by a super-resolution module that restores the fine details of rendered images. Additionally, we develop a grouped content filtering mechanism to improve the 3D scene consistency of the model by matching the rendering images and text descriptions, which also significantly reduces memory usage during editing. Extensive experiments demonstrate that the proposed method can achieve various editing effects, including different seasons, weather conditions, times of the day, disaster scenarios, etc. Our technique is computationally efficient and conveniently expandable for large-scale UAV scenes, alleviating data scarcity in harsh scenarios. Yufeng Wang 0004, Shuangkang Fang, Zehao Zhang, Xianlin Zeng, Wenrui Ding |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2024 | Multi-Person Pose Tracking With Sparse Key-Point Flow Estimation and Hierarchical Graph Distance MinimizationabstractIn this paper, we propose a novel framework for multi-person pose estimation and tracking on challenging scenarios. In view of occlusions and motion blurs which hinder the performance of pose tracking, we proposed to model humans as graphs and perform pose estimation and tracking by concentrating on the visible parts of human bodies which are informative about complete skeletons under incomplete observations. Specifically, the proposed framework involves three parts: (i) A Sparse Key-point Flow Estimating Module (SKFEM) and a Hierarchical Graph Distance Minimizing Module (HGMM) for estimating pixel-level and human-level motion, respectively; (ii) Pixel-level appearance consistency and human-level structural consistency are combined in measuring the visibility scores of body joints. The scores guide the pose estimator to predict complete skeletons by observing high-visibility parts, under the assumption that visible and invisible parts are inherently correlated in human part graphs. The pose estimator is iteratively fine-tuned to achieve this capability; (iii) Multiple historical frames are combined to benefit tracking which is implemented using HGMM. The proposed approach not only achieves state-of-the-art performance on PoseTrack datasets but also contributes to significant improvements in other tasks such as human-related anomaly detection. Yalong Jiang, Wenrui Ding, Zheru Chi |
IEEE Trans. Image Process. | 2 |
| 2023 | DFFG: Fast Gradient Iteration for Data-free Quantization
Huixing Leng, Shuangkang Fang, Yufeng Wang 0004, Zehao Zhang, Dacheng Qi, Wenrui Ding |
BMVC | 6 |
| 2023 | Frequency learning attention networks based on deep learning for automatic modulation classification in wireless communication
Duona Zhang, Yuanyao Lu, Yundong Li, Wenrui Ding, Baochang Zhang 0001 |
Pattern Recognit. | 4 |
| 2023 | Multiscale Correlation Networks Based on Deep Learning for Automatic Modulation ClassificationabstractAutomatic Modulation Classification (AMC) is a challenging yet significant technique for communication systems. Deep learning methods, though widely employed for AMC, are challenged by the poor representation in noisy scenarios. In this letter, we propose a novel Multiscale Correlation Networks (MSCNs) approach to enhance noise suppression and bolster representation power for AMC. MSCNs leverage learning correlation wavelet transform to redistribute radio signal and noise features across various scales, and incorporate multiscale correlation and attention mechanisms to characterize feature properties in terms of frequency. Experiments reveal that MSCNs achieve an overall classification rate of 91.36% at 10dB using 152k training samples on the public RadioML 2018.01A benchmark. Yufeng Wang 0004, Duona Zhang, Qinyan Ma, Wenrui Ding |
IEEE Signal Process. Lett. | 5 |
| 2023 | A Hierarchical Spatio-Temporal Graph Convolutional Neural Network for Anomaly Detection in VideosabstractDeep learning models have been widely used for anomaly detection in surveillance videos. Typical models are equipped with the capability to reconstruct normal videos and evaluate the reconstruction errors on anomalous videos to indicate the extent of abnormalities. However, existing approaches suffer from two disadvantages. Firstly, they can only encode the movements of each identity independently, without considering the interactions among identities which may also indicate anomalies. Secondly, they leverage inflexible models whose structures are fixed under different scenes, this configuration disables the understanding of scenes. In this paper, we propose a Hierarchical Spatio-Temporal Graph Convolutional Neural Network (HSTGCNN) to address these problems, the HSTGCNN is composed of multiple branches that correspond to different levels of graph representations. High-level graph representations encode the trajectories of people and the interactions among multiple identities while low-level graph representations encode the local body postures of each person. Furthermore, we propose to weightedly combine multiple branches that are better at different scenes. An improvement over single-level graph representations is achieved in this way. An understanding of scenes is achieved and serves anomaly detection. High-level graph representations are assigned higher weights to encode moving speed and directions of people in low-resolution videos while low-level graph representations are assigned higher weights to encode human skeletons in high-resolution videos. Experimental results show that the proposed HSTGCNN significantly outperforms current state-of-the-art models on four benchmark datasets (UCSD Pedestrian, ShanghaiTech, CUHK Avenue and IITB-Corridor) by using much less learnable parameters. Xianlin Zeng, Yalong Jiang, Wenrui Ding, Yafeng Hao, Zifeng Qiu |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2023 | High-Order Convolutional Attention Networks for Automatic Modulation Classification in CommunicationabstractAutomatic modulation classification is a challenging and critical task in the field of communication. Deep convolutional networks (ConvNets) have been recently applied in cognitive radio and achieved remarkable performance. However, existing ConvNet-based methods mainly focus on the first-order architecture design, while fail to explore feature correlations of radio signal, which are particularly significant for useful information extraction in low signal-to-noise ratios (SNRs). In this paper, we propose high-order convolutional attention networks (HoCANs) for radio signal expression and feature correlation learning, based on a novel high-order attention mechanism to rescale the convolutional features along channel and sequence dimensions. High-order convolutional layer and covariance matrix after nonlinear transformation are led for tenser filtering with more discriminative representations of radio signals. Experiments have been conducted to validate the superiority of HoCANs which achieve state-of-the-art accuracy for automatic modulation classification on RADIOML 2018.01A dataset. Duona Zhang, Yuanyao Lu, Yundong Li, Wenrui Ding, Baochang Zhang 0001 |
IEEE Trans. Wirel. Commun. | 4 |
| 2022 | Towards Accurate Binary Neural Networks via Modeling Contextual Dependencies
Xingrun Xing, Yangguang Li 0001, Wei Li 0022, Wenrui Ding, Yalong Jiang, Yufeng Wang 0004, Chunlei Liu 0001, Xianglong Liu 0001 |
ECCV (11) | 4 |
| 2022 | Binary Dense Predictors for Human Pose Estimation Based on Dynamic Thresholds and FilteringabstractBinary neural networks (BNNs) contribute a lot to the efficiency of image classification models. However, in dense predication tasks such as human pose estimation, predictions in different locations are coupled and rely on the extraction of features across entire images. As a result, more robust and adaptive binarization is required to bridge the performance gap between binarized and full precision models. We propose two approaches to conduct image-aware and pixel-aware dynamic binarization in a model for human pose estimation. Firstly, a simplified dynamic thresholding is leveraged in the backbone to determine unique binarization thresholds for each image. Secondly, in the decoder, we decouple binarization for each pixel according to the activations surrounding the pixel. Dynamic filtering modules are proposed to determine a different binarization strategy for each pixel. Compared with the strong baselines, the proposed framework improves 5.2% and 3.6% mAP on the COCO test-dev benchmark for ResNet-18/34 architectures respectively. Xingrun Xing, Yalong Jiang, Baochang Zhang 0001, Wenrui Ding, Huan Peng |
ICASSP | 4 |
| 2022 | Semi-supervised Multi-task Learning for Semantics and DepthabstractMulti-Task Learning (MTL) aims to enhance the model generalization by sharing representations between related tasks for better performance. Typical MTL methods are jointly trained with the complete multitude of ground-truths for all tasks simultaneously. However, one single dataset may not contain the annotations for each task of interest. To address this issue, we propose the Semi-supervised Multi-Task Learning (SemiMTL) method to leverage the available supervisory signals from different datasets, particularly for semantic segmentation and depth estimation tasks. To this end, we design an adversarial learning scheme in our semi-supervised training by leveraging unlabeled data to optimize all the task branches simultaneously and accomplish all tasks across datasets with partial annotations. We further present a domain-aware discriminator structure with various alignment formulations to mitigate the domain discrepancy issue among datasets. Finally, we demonstrate the effectiveness of the proposed method to learn across different datasets on challenging street view and remote sensing benchmarks. Yufeng Wang 0004, Yi-Hsuan Tsai, Wei-Chih Hung, Wenrui Ding, Ming-Hsuan Yang 0001 |
WACV | 4 |
| 2022 | RB-Net: Training Highly Accurate and Efficient Binary Neural Networks With Reshaped Point-Wise Convolution and Balanced ActivationabstractIn this paper, we find that the conventional convolution operation becomes the bottleneck for extremely efficient binary neural networks (BNNs). To address this issue, we open up a new direction by introducing a reshaped point-wise convolution (RPC) to replace the conventional one to build BNNs. Specifically, we conduct a point-wise convolution after rearranging the spatial information into depth, with which at least$2.25\times $computation reduction can be achieved. Such an efficient RPC allows us to explore more powerful representational capacity of BNNs under a given computation complexity budget. Moreover, we propose to use a balanced activation (BA) to adjust the distribution of the scaled activations after binarization, which enables significant performance improvement of BNNs. After integrating RPC and BA, the proposed network, dubbed as RB-Net, strikes a good trade-off between accuracy and efficiency, achieving superior performance with lower computational cost against the state-of-the-art BNN methods. Specifically, our RB-Net achieves 66.8% Top-1 accuracy with ResNet-18 backbone on ImageNet, exceeding the state-of-the-art Real-to-Binary Net (65.4%) by 1.4% while achieving more than$3\times $reduction (52M vs. 165M) in computational complexity. Chunlei Liu 0001, Wenrui Ding, Peng Chen 0037, Bohan Zhuang, Yufeng Wang 0004, Yang Zhao 0019, Baochang Zhang 0001, Yuqi Han |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2021 | SA-BNN: State-Aware Binary Neural NetworkabstractBinary Neural Networks (BNNs) have received significant attention due to the memory and computation efficiency recently. However, the considerable accuracy gap between BNNs and their full-precision counterparts hinders BNNs to be deployed to resource-constrained platforms. One of the main reasons for the performance gap can be attributed to the frequent weight flip, which is caused by the misleading weight update in BNNs. To address this issue, we propose a state-aware binary neural network (SA-BNN) equipped with the well designed state-aware gradient. Our SA-BNN is inspired by the observation that the frequent weight flip is more likely to occur, when the gradient magnitude for all quantization states {-1,1} is identical. Accordingly, we propose to employ independent gradient coefficients for different states when updating the weights. Furthermore, we also analyze the effectiveness of the state-aware gradient on suppressing the frequent weight flip problem. Experiments on ImageNet show that the proposed SA-BNN outperforms the current state-of-the-arts (e.g., Bi-Real Net) by more than 3% when using a ResNet architecture. Specifically, we achieve 61.7%, 65.5% and 68.7% Top-1 accuracy with ResNet-18, ResNet-34 and ResNet-50 on ImageNet, respectively. Chunlei Liu 0001, Peng Chen 0037, Bohan Zhuang, Chunhua Shen, Baochang Zhang 0001, Wenrui Ding |
AAAI | 6 |
| 2021 | TRQ: Ternary Neural Networks With Residual QuantizationabstractTernary neural networks (TNNs) are potential for network acceleration by reducing the full-precision weights in network to ternary ones, e.g., {-1,0,1}. However, existing TNNs are mostly calculated based on rule-of-thumb quantization methods by simply thresholding operations, which causes a significant accuracy loss. In this paper, we introduce a stem-residual framework which provides new insight into Ternary quantization, termed Residual Quantization (TRQ), to achieve more powerful TNNs. Rather than directly thresholding operations, TRQ recursively performs quantization on full-precision weights for a refined reconstruction by combining the binarized stem and residual parts. With such a unique quantization process, TRQ endows the quantizer with high flexibility and precision. Our TRQ is generic, which can be easily extended to multiple bits through recursively encoded residual for a better recognition accuracy. Extensive experimental results demonstrate that the proposed method yields great recognition accuracy while being accelerated. Wenrui Ding, Chunlei Liu 0001, Baochang Zhang 0001, Guodong Guo |
AAAI | 2 |
| 2021 | Rectified Binary Convolutional Networks with Generative Adversarial Learning
Chunlei Liu 0001, Wenrui Ding, Baochang Zhang 0001, Jianzhuang Liu, Guodong Guo, David S. Doermann |
Int. J. Comput. Vis. | 2 |
| 2021 | Learning modulation filter networks for weak signal detection in noise
Duona Zhang, Wenrui Ding, Baochang Zhang 0001, Chunhui Liu 0004, Jungong Han, David S. Doermann |
Pattern Recognit. | 2 |
| 2020 | Adaptive Mixture Regression Network with Local Counting Map for Crowd Counting
Wenrui Ding, Tieqiang Wang, Zhijin Wang, Junjun Xiong |
ECCV (24) | 3 |
| 2020 | Cam-Net: Compressed Attentive Multi-Granularity Network For Dynamic Scene ClassificationabstractDynamic scene classification on portable platforms is extremely challenging due to the contradiction between model complexity and computing resources. To resolve this long-standing dilemma, we propose the compressed attentive multi-granularity network (CAM-Net) in a two-step manner. First, we present a novel AM-Net based on multi-granularity attention units to boost the performance of the full-precision model. It captures and enhances both coarse and fine target-related information. Then, we introduce an efficient binary approximation to AM-Net to improve computing efficiency, leading to CAM-Net. Particularly, a grouping guidance approach is adopted to guide the reconstruction of full-precision weights from binary ones. With this guidance, CAM-Net can significantly reduce memory usage as well as CPU consumption, yet only cause a slight decline in accuracy. Extensive experiments have been conducted on three benchmark datasets, i.e., Maryland, YUPENN ++ and ActivityNet, demonstrating the effectiveness and superiority of the proposed method on scene classification. Wenrui Ding, Yanjun Zhu, Yuanjun Huang, Yalong Jiang, Baochang Zhang 0001 |
ICIP | 2 |
| 2020 | Superpixel Labeling Priors and MRF for Aerial Video SegmentationabstractVideo segmentation is a task of partitioning pixels that exhibit homogeneous appearance and motion into coherent spatial-temporal groups, which is still challenging for aerial applications. In this paper, a principled combination of superpixel labeling priors and Markov random field (S-MRF) is proposed for aerial video segmentation. The proposed approach has several contributions: 1) we develop a metadata-based global projection model with coordinate transformation to estimate motion information between frames; 2) the superpixel labeling priors from previous frames are incorporated into the segmentation of the current frame, leading to a highly efficient probabilistic label propagation algorithm; and 3) we perform an MRF optimization on the initial segments with propagated labeling priors to improve the temporal coherency. In addition, a new video dataset is collected and will be made publicly available to evaluate the performance of aerial video segmentation algorithms. The experimental results show that the proposed approach outperforms the state-of-the-art video segmentation methods. Yufeng Wang 0004, Wenrui Ding, Baochang Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2020 | Aggregation Signature for Small Object TrackingabstractSmall object tracking becomes an increasingly important task, which however has been largely unexplored in computer vision. The great challenges stem from the facts that: 1) small objects show extreme vague and variable appearances, and 2) they tend to be lost easier as compared to normal-sized ones due to the shaking of lens. In this paper, we propose a novel aggregation signature suitable for small object tracking, especially aiming for the challenge of sudden and large drift. We make three-fold contributions in this work. First, technically, we propose a new descriptor, named aggregation signature, based on saliency, able to represent highly distinctive features for small objects. Second, theoretically, we prove that the proposed signature matches the foreground object more accurately with a high probability. Third, experimentally, the aggregation signature achieves a high performance on multiple datasets, outperforming the state-of-the-art methods by large margins. Moreover, we contribute with two newly collected benchmark datasets, i.e., small90 and small112, for visually small object tracking. The datasets will be available in https://github.com/bczhangbczhang/. Chunlei Liu 0001, Wenrui Ding, Vittorio Murino, Baochang Zhang 0001, Jungong Han, Guodong Guo |
IEEE Trans. Image Process. | 2 |
| 2019 | Circulant Binary Convolutional Networks: Enhancing the Performance of 1-Bit DCNNs With Circulant Back PropagationabstractThe rapidly decreasing computation and memory cost has recently driven the success of many applications in the field of deep learning. Practical applications of deep learning in resource-limited hardware, such as embedded devices and smart phones, however, remain challenging. For binary convolutional networks, the reason lies in the degraded representation caused by binarizing full-precision filters. To address this problem, we propose new circulant filters (CiFs) and a circulant binary convolution (CBConv) to enhance the capacity of binarized convolutional features via our circulant back propagation (CBP). The CiFs can be easily incorporated into existing deep convolutional neural networks (DCNNs), which leads to new Circulant Binary Convolutional Networks (CBCNs). Extensive experiments confirm that the performance gap between the 1-bit and full-precision DCNNs is minimized by increasing the filter diversity, which further increases the representational ability in our networks. Our experiments on ImageNet show that CBCNs achieve 61.4% top-1 accuracy with ResNet18. Compared to the state-of-the-art such as XNOR, CBCNs can achieve up to 10% higher top-1 accuracy with more powerful representational ability. Chunlei Liu 0001, Wenrui Ding, Xin Xia 0005, Baochang Zhang 0001, Jiaxin Gu, Jianzhuang Liu, Rongrong Ji, David S. Doermann |
CVPR | 2 |
| 2019 | Rectified Binary Convolutional Networks for Enhancing the Performance of 1-bit DCNNsabstractBinarized convolutional neural networks (BCNNs) are widely used to improve memory and computation efficiency of deep convolutional neural networks (DCNNs) for mobile and AI chips based applications. However, current BCNNs are not able to fully explore their corresponding full-precision models, causing a significant performance gap between them. In this paper, we propose rectified binary convolutional networks (RBCNs), towards optimized BCNNs, by combining full-precision kernels and feature maps to rectify the binarization process in a unified framework. In particular, we use a GAN to train the 1-bit binary network with the guidance of its corresponding full-precision model, which significantly improves the performance of BCNNs. The rectified convolutional layers are generic and flexible, and can be easily incorporated into existing DCNNs such as WideResNets and ResNets. Extensive experiments demonstrate the superior performance of the proposed RBCNs over state-of-the-art BCNNs. In particular, our method shows strong generalization on the object tracking task. Chunlei Liu 0001, Wenrui Ding, Xin Xia 0005, Baochang Zhang 0001, Jianzhuang Liu, Bohan Zhuang, Guodong Guo |
IJCAI | 2 |
| 2018 | A Saliency-Based Object Tracking Method for UAV Application
Wenrui Ding, Chunlei Liu 0001, Zechen Ha |
PRCV (4) | 2 |
| 2018 | Haze removal for unmanned aerial vehicle aerial video based on spatial-temporal coherence optimisationabstractHaze removal is a non‐trivial work for unmanned aerial vehicle (UAV) aerial video processing, and challenges are mainly attributed to spatial‐temporal coherence and computational efficiency. The authors propose a novel dehazing algorithm for hazy UAV aerial video and improve the classical dark channel prior approach with a bright region filling process to alleviate colour distortion in the recovered video, which enhances the spatial consistency. To achieve better temporal coherence, the authors constrain the atmospheric light estimation between adjacent frames by using a temporal filter. The authors also optimise the transmission calculation to reduce computational complexity. Experimental findings show that the proposed algorithm yields results superior to those obtained from previous methods. Compared with that in frame‐by‐frame dehazing, the processing time in the proposed method is reduced by 74.5%. Xintao Zhao, Wenrui Ding, Chunhui Liu 0004 |
IET Image Process. | 2 |
| 2018 | Hybrid Gabor Convolutional Networks
Chunlei Liu 0001, Wenrui Ding, Baochang Zhang 0001 |
Pattern Recognit. Lett. | 2 |
| 2018 | Adaptive Equalizer Design for Unmanned Aircraft Vehicle Image Transmission over Relay ChannelsabstractA novel length adaptive method is proposed for time domain equalizer by taking the channel attenuation ratio between different multipath components into account in UAV‐UAV and UAV‐ground channels. Then, considering received image quality, the minimum bit error ratio (MBER) criterion is exploited to design adaptive equalizers for both amplify‐and‐forward (AF) and decode‐and‐forward (DF) relaying systems by the proposed length adaptive method. Results show that proposed MBER adaptive equalizers outperform the traditional ones in both AF relaying and DF relaying as channel attenuation ratio in UAV‐ground channel increases. Moreover, DF outperforms AF as channel attenuation ratio in UAV‐UAV channel increases. Furthermore, bit error ratio (BER) and peak signal‐to‐noise ratio (PSNR) performances in both AF and DF are evaluated to show the enhancement by the proposed MBER adaptive equalizers. Wenqian Huang, Wenrui Ding |
Wirel. Commun. Mob. Comput. | 2 |
| 2016 | Airborne SAR and optical image fusion based on IHS transform and joint non-negative sparse representationabstractIn this paper, a novel airborne synthetic aperture radar (SAR) and optical image fusion method based on Intensity-Hue-Saturation (IHS) and joint non-negative sparse representation (JNNSR) is proposed. Firstly, the color optical image is transformed into IHS space. Then, the intensity component of the optical image and the SAR image are decomposed by JNNSR into common component and innovation components. Based on the non-negative property of the sparse coefficients, the innovation component of the SAR image is fused with the intensity component of the optical image. Finally, the fused result is obtained by performing inverse IHS transform. The experimental result shows that our method is superior to the traditional methods in terms of several universal quality evaluation indexes, as well as in the visual quality. Chunhui Liu 0004, Wenrui Ding |
IGARSS | 3 |