Yufeng Wang 0004

dblp:90/6339-4 · DBLP profile ↗
← Back
34ranked-venue papers
5as first author
33since 2021 · last 2026
0000-0001-8713-3153ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 21 · 3 first-author · 20 since 2021Artificial intelligence and machine learning · 17 · 1 first-author · 17 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 first-author · 4 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Progressive Volume Distillation with Active Learning for Efficient NeRF Architecture Conversion
Shuangkang Fang, Yufeng Wang 0004, Yi Yang 0033, Wenrui Ding, Shuchang Zhou 0001
Int. J. Comput. Vis.2
2026 Dynamic Prompt Compression for Efficient Inference of Large Language Models
abstract
Large language models (LLMs) have shown outstanding performance across a variety of tasks, partly due to advanced prompting techniques. However, these techniques often require lengthy prompts, which increase computational costs and can hinder performance because of the limited context windows of LLMs. While prompt compression is a straightforward solution, existing methods confront the challenges of retaining essential information, adapting to context changes, and remaining effective across different tasks. To tackle these issues, we propose a task-agnostic method called Dynamic Prompt Compression (LLM-DPC). Our method reduces the number of prompt tokens while minimizing any degradation in LLM performance. We model prompt compression as a Markov Decision Process (MDP), enabling the DPC-Agent to sequentially remove redundant tokens by adapting to dynamic contexts and retaining crucial content. We develop a reward function for training the DPC-Agent that balances the compression ratio, the quality of the LLM output, and the retention of key information. This allows for prompt token reduction without needing an external black-box LLM. Inspired by the progressive difficulty adjustment in curriculum learning, we introduce a Hierarchical Prompt Compression (HPC) training strategy that gradually increases the compression difficulty, enabling the DPC-Agent to learn an effective compression method that maintains information integrity. Experiments demonstrate that our method outperforms state-of-the-art techniques, especially at higher compression ratio.
Jinwu Hu, Wei Zhang 0098, Yufeng Wang 0004, Yu Hu 0004, Bin Xiao 0002, Mingkui Tan
IEEE Trans. Knowl. Data Eng.3
2026 Editing 3D Scenes via Text Prompts Without Retraining
abstract
Numerous diffusion models have been developed for 2D image synthesis and editing, and recently they are extended to 3D scene editing tasks. However, editing 3D scenes is still in its early stages, and the challenges of scene representations and multi-view consistency need to be addressed. A notable limitation of existing approaches is the need for specific modules for different edits and model retraining for each scene. To tackle these issues, we propose a novel and versatile text-driven 3D scene editing method, termed DN2N, which allows for the direct acquisition of the editing results without the requirement for retraining. Our method employs off-the-shelf text-based editing models of 2D images to modify the multi-view images of a 3D scene. A content filtering process is then applied to discard poorly edited images that disrupt 3D consistency. We consider the remaining inconsistency as a problem of removing noise perturbations and solve it by generating data with similar perturbation characteristics for training. We develop a versatile NeRF model structure and propose two novel cross-view regularization terms to help the DN2N mitigate these perturbations. Empirical results show that our method achieves multiple editing types based solely on text prompts, including but not limited to appearance editing, weather transition, object changing, and style transfer. Most importantly, DN2N exhibits a versatility of editing capabilities, eliminating the need to customize or retrain editing models for specific scenes or editing types. Namely, DN2N achieves comparable total editing time to the 3DGS-based editing method, enhancing its practical value.
Shuangkang Fang, Yufeng Wang 0004, Yi-Hsuan Tsai, Wenrui Ding, Yi Yang 0033, Shuchang Zhou 0001, Ming-Hsuan Yang 0001
IEEE Trans. Vis. Comput. Graph.2
2025 Open-world Radio Frequency Fingerprint Identification via Augmented Semi-supervised Learning
abstract
In complex electromagnetic environments, the identification and differentiation of diverse radio frequency (RF) emitters become particularly crucial. Existing RF fingerprinting methods demonstrate limitations when dealing with numerous unknown emitters, making it challenging for accurate classification and recognition. These limitations hinder the effective handling of specific unknown emitters.To address this issue, we introduce a novel RF fingerprinting method suitable for open-world conditions for the first time. We develop a novel RF fingerprinting model, Roinformer, to extract signal features with positional attention. We then leverage data augmentation strategies such as noise jitter and signal frame rearrangement to construct an effective pre-training model. Moreover, by incorporating instance-level similarity loss and a novel local entropy regularization approach, we significantly enhance the accuracy of known class identification and mitigate the catastrophic forgetting of known signal samples. Experimental results on three temporal signal datasets demonstrate that our method effectively recognizes both the known and unknown classes, outperforming several state-of-the-art methods by a large margin.
Zehua Han, Qirui Zhao, Zhexuan Cui, Yufeng Wang 0004, Duona Zhang, Wenrui Ding
AAAI5
2025 Graph Structure Refinement with Energy-based Contrastive Learning
abstract
Graph Neural Networks (GNNs) have recently gained widespread attention as a successful tool for analyzing graph-structured data. However, imperfect graph structure with noisy links lacks enough robustness and may damage graph representations, therefore limiting the GNNs' performance in practical tasks. Moreover, existing generative architectures fail to fit discriminative graph-related tasks. To tackle these issues, we introduce an unsupervised method based on a joint of generative training and discriminative training to learn graph structure and representation, aiming to improve the discriminative performance of generative models. We propose an Energy-based Contrastive Learning (ECL) guided Graph Structure Refinement (GSR) framework, denoted as ECL-GSR. To our knowledge, this is the first work to combine energy-based models with contrastive learning for GSR. Specifically, we leverage ECL to approximate the joint distribution of sample pairs, which increases the similarity between representations of positive pairs while reducing the similarity between negative ones. Refined structure is produced by augmenting and removing edges according to the similarity metrics among node representations. Extensive experiments demonstrate that ECL-GSR outperforms the state-of-the-art on eight benchmark datasets in node classification. ECL-GSR achieves faster training with fewer samples and memories against the leading baseline, highlighting its simplicity and efficiency in downstream tasks.
Xianlin Zeng, Yufeng Wang 0004, Guodong Guo, Wenrui Ding, Baochang Zhang 0001
AAAI2
2025 Cellular-Connected UAV Path Planning on Radio Maps: LLMs Enhanced Approach
abstract
In this paper, we propose a large language models (LLMs) enhanced path planning framework for a cellular-connected unmanned aerial vehicle (UAV), which can significantly improve the efficiency of path planning without sacrificing path quality. The ray tracing model is used to generate the radio map of UAV cellular connectivity from real-world geospatial data. By introducing quadtree-assisted representation, the radio map is formatted into prompts that can be understood by LLMs, bridging the gap between simplified obstacle environments and realistic electromagnetic environments. Meanwhile, a combined approach using LLMs and the graph-based search method is employed to find a path which can guarantee reliable connections between the UAV and its associated ground base station throughout the flight. Simulation results demonstrate that our proposed LLM-enhanced framework not only finds the path that guarantees communication quality, but also addresses computational and memory limitations of the conventional method, with a significant reduction in operations and storage.
Xuetong Pei, Fuyuan Ma, Shutong Wang, Zesheng Wang 0002, Wenrui Ding, Yufeng Wang 0004
GLOBECOM6
2025 ASFC-NeRF: Large-Scale Scene Rendering with Adaptive Sampling and Feature-aware Compression
abstract
While significant progress has been made in large-scale scene representation using Neural Radiance Fields (NeRF), several limitations remain. For instance, most methods still rely on the original coarse-to-fine sampling strategy, leading to an inefficient rendering process. Additionally, to model larger scenes, these methods often use complex network models, resulting in redundant model parameters. To address these issues, we propose a novel model with adaptive sampling and feature-aware compression for large-scale scene rendering, named ASFC-NeRF. We first introduce a weight prediction network to replace the original coarse sampling strategy, then employ a teacher network and depth constraints for knowledge distillation in the early stages of training to enhance the high-fidelity of the scene. Furthermore, we optimize the number of Grids and the channels of Planes and prune the network to efficiently compress model parameters. Experimental results demonstrate that our method significantly accelerates the rendering process and greatly reduces parameter quantity while maintaining or only slightly lowering image quality. Therefore, ASFC-NeRF exhibits advantages in comprehensive performance and practicality.
Yufeng Wang 0004, Shuangkang Fang, Zesheng Wang 0002, Dacheng Qi, Wenrui Ding
ICASSP2
2025 NeRF is a Valuable Assistant for 3D Gaussian Splatting
abstract
We introduce NeRF-GS, a novel framework that jointly optimizes Neural Radiance Fields (NeRF) and 3D Gaussian Splatting (3DGS). This framework leverages the inherent continuous spatial representation of NeRF to mitigate several limitations of 3DGS, including sensitivity to Gaussian initialization, limited spatial awareness, and weak inter-Gaussian correlations, thereby enhancing its performance. In NeRF-GS, we revisit the design of 3DGS and progressively align its spatial features with NeRF, enabling both representations to be optimized within the same scene through shared 3D spatial information. We further address the formal distinctions between the two approaches by optimizing residual vectors for both implicit features and Gaussian positions to enhance the personalized capabilities of 3DGS. Experimental results on benchmark datasets show that NeRF-GS surpasses existing methods and achieves state-of-the-art performance. This outcome confirms that NeRF and 3DGS are complementary rather than competing, offering new insights into hybrid approaches that combine 3DGS and NeRF for efficient 3D scene representation.
Shuangkang Fang, I-Chao Shen, Takeo Igarashi, Yufeng Wang 0004, Zesheng Wang 0002, Yi Yang 0033, Wenrui Ding, Shuchang Zhou 0001
ICCV4
2025 MeshLLM: Empowering Large Language Models to Progressively Understand and Generate 3D Mesh
Shuangkang Fang, I-Chao Shen, Yufeng Wang 0004, Yi-Hsuan Tsai, Yi Yang 0033, Shuchang Zhou 0001, Wenrui Ding, Takeo Igarashi, Ming-Hsuan Yang 0001
ICCV3
2025 Deformable Spherical Geometry Transformer For Panoramic Semantic Segmentation
abstract
The increasing availability of 360° images has created a demand for effective Panoramic Semantic Segmentation (PASS) to enable comprehensive scene understanding. However, the spherical nature of 360° image introduces significant spatial distortions due to Equirectangular Projection (ERP), making it challenging for traditional 2D methods, which are designed for Euclidean spaces. Existing PASS methods typically mitigate these distortions through developing spherical-to-tangent polyhedron transformations or special-ized convolutional structures. Nevertheless, these approaches still struggle to preserve the spherical geometry and fail to adequately capture the semantic context of 360° images. In this paper, we propose a Deformable Spherical Geometry Transformer (DSGT) network that adapts to spherical distortions through a local-global self-attention mechanism. The local self-attention module captures local semantic information to alleviate distortions, while the global self-attention module integrates spherical geometric priors to enhance predictions. Experimental results on the Stanford2D3D panoramic dataset demonstrate that DSGT outperforms state-of-the-art PASS methods.
Boyang Lan, Li Yang 0014, Mai Xu, Lai Jiang 0004, Yufeng Wang 0004
ICIP5
2025 Guiding Yourself with Your Own Insights: Student-Driven Knowledge Distillation
abstract
Knowledge distillation (KD) stands as an efficient technique for compressing models, typically employing a teacher-student framework. Nevertheless, optimizing KD to yield models with reduced parameters and enhanced performance remains an area warranting deeper investigation. In this article, we recognize the significance of preserving structural consistency to enhance knowledge transfer efficiency between networks. Leveraging this insight, we introduce a novel approach termed Student-Driven Knowledge Distillation (SDKD), which integrates a proxy teacher intermediary between the primary teacher and student model. Specifically, we construct the architecture of the proxy teacher entirely based on the student network to generate logits that closely align with the distribution of the student network. Besides, we propose a Feature Fusion Block (FFB) to integrate features from the teacher network into the proxy teacher. FFB can not only provide high-quality feature-based knowledge for distillation but also impart response-based knowledge to facilitate the learning process. Extensive experiments illustrate that SDKD outperforms 29 state-of-the-art methods on several tasks, including image classification, semantic segmentation, and depth estimation.
Dacheng Qi, Yufeng Wang 0004, Shuangkang Fang, Zehao Zhang, Zesheng Wang 0002, Wenrui Ding
ICME3
2025 SANE: Enhancing Large-scale Scene Representation with Semantic-aware NeRF Experts
abstract
We propose the Semantic-aware NeRF Experts (SANE), which fully exploits the intrinsic characteristics of large-scale scenes, including semantics and material features, to achieve high-quality novel view synthesis results and provide accurate 3D semantic information. SANE begins by building a semantic Mixture of Experts (MoE), utilizing a learnable gating network to semantically partition the scene into blocks for corresponding NeRF experts. We then develop a semantic volume rendering scheme that integrates discrete semantics into the end-to-end differentiable process of NeRF, enabling refined semantic labeling of each scene point. Additionally, we implement a dual-implicit encoding strategy: intra-block encoding captures lighting variations across viewpoints, while inter-block one captures texture features among different semantic objects. Experiments on benchmark datasets show that SANE delivers higher-quality scene representations and effective semantic decomposition for downstream tasks, such as precise editing of large-scale scenes based on semantics.
Zesheng Wang 0002, Yufeng Wang 0004, Shuangkang Fang, Dacheng Qi, Shengxi Li, Mai Xu, Wenrui Ding
ICME2
2025 Enhancing User-Oriented Proactivity in Open-Domain Dialogues with Critic Guidance
abstract
Open-domain dialogue systems aim to generate natural and engaging conversations, providing significant practical value in real applications such as social robotics and personal assistants. The advent of large language models (LLMs) has greatly advanced this field by improving context understanding and conversational fluency. However, existing LLM-based dialogue systems often fall short in proactively understanding the user's chatting preferences and guiding conversations toward user-centered topics. This lack of user-oriented proactivity can lead users to feel unappreciated, reducing their satisfaction and willingness to continue the conversation in human-computer interactions. To address this issue, we propose a User-oriented Proactive Chatbot (UPC) to enhance the user-oriented proactivity. Specifically, we first construct a critic to evaluate this proactivity inspired by the LLM-as-a-judge strategy. Given the scarcity of high-quality training data, we then employ the critic to guide dialogues between the chatbot and user agents, generating a corpus with enhanced user-oriented proactivity. To ensure the diversity of the user backgrounds, we introduce the ISCO-800, a diverse user background dataset for constructing user agents. Moreover, considering the communication difficulty varies among users, we propose an iterative curriculum learning method that trains the chatbot from easy-to-communicate users to more challenging ones, thereby gradually enhancing its performance. Experiments demonstrate that our proposed training method is applicable to different LLMs, improving user-oriented proactivity and attractiveness in open-domain dialogues. Code and appendix are available at github.com/wang678/LLM-UPC.
Yufeng Wang 0004, Jinwu Hu, Ziteng Huang, Kunyang Lin, Zitian Zhang, Peihao Chen, Yu Hu 0004, Qianyue Wang, Zhu Liang Yu, Bin Sun 0001, Xiaofen Xing, Mingkui Tan
IJCAI1
2025 Efficient Dynamic Ensembling for Multiple LLM Experts
abstract
LLMs have demonstrated impressive performance across various language tasks. However, the strengths of LLMs can vary due to different architectures, model sizes, areas of training data, etc. Therefore, ensemble reasoning for the strengths of different LLM experts is critical to achieving consistent and satisfactory performance on diverse inputs across a wide range of tasks. However, existing LLM ensemble methods are either computationally intensive or incapable of leveraging complementary knowledge among LLM experts for various inputs. In this paper, we propose an efficient Dynamic Ensemble Reasoning paradigm, called DER to integrate the strengths of multiple LLM experts conditioned on dynamic inputs. Specifically, we model the LLM ensemble reasoning problem as a Markov Decision Process, wherein an agent sequentially takes inputs to request knowledge from an LLM candidate and passes the output to a subsequent LLM candidate. Moreover, we devise a reward function to train a DER-Agent to dynamically select an optimal answering route given the input questions, aiming to achieve the highest performance with as few computational resources as possible. Last, to fully transfer the expert knowledge from the prior LLMs, we develop a Knowledge Transfer Prompt that enables the subsequent LLM candidates to transfer complementary knowledge effectively. Experiments demonstrate that our method uses fewer computational resources to achieve better performance compared to state-of-the-art baselines. Code and appendix are available at https://github.com/Fhujinwu/DER.
Jinwu Hu, Yufeng Wang 0004, Shuhai Zhang, Yu Hu 0004, Bin Xiao 0004, Mingkui Tan
IJCAI2
2025 Generating Long-form Story Using Dynamic Hierarchical Outlining with Memory-Enhancement
abstract
Qianyue Wang, Jinwu Hu, Zhengping Li, Yufeng Wang, Daiyuan Li, Yu Hu, Mingkui Tan. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Qianyue Wang, Jinwu Hu, Zhengping Li, Yufeng Wang 0004, Daiyuan Li, Yu Hu 0004, Mingkui Tan
NAACL (Long Papers)4
2025 Open-World Drone Active Tracking with Goal-Centered Rewards
abstract
Drone Visual Active Tracking aims to autonomously follow a target object by controlling the motion system based on visual observations, providing a more practical solution for effective tracking in dynamic environments. However, accurate Drone Visual Active Tracking using reinforcement learning remains challenging due to the absence of a unified benchmark and the complexity of open-world environments with frequent interference. To address these issues, we pioneer a systematic solution. First, we propose DAT, the first open-world drone active air-to-ground tracking benchmark. It encompasses 24 city-scale scenes, featuring targets with human-like behaviors and high-fidelity dynamics simulation. DAT also provides a digital twin tool for unlimited scene generation. Additionally, we propose a novel reinforcement learning method called GC-VAT, which aims to improve the performance of drone tracking targets in complex scenarios. Specifically, we design a Goal-Centered Reward to provide precise feedback across viewpoints to the agent, enabling it to expand perception and movement range through unrestricted perspectives. Inspired by curriculum learning, we introduce a Curriculum-Based Training strategy that progressively enhances the tracking performance in complex environments. Besides, experiments on simulator and real-world images demonstrate the superior performance of GC-VAT, achieving a Tracking Success Rate of approximately 72% on the simulator. The benchmark and code are available at https://github.com/SHWplus/DAT_Benchmark.
Haowei Sun, Jinwu Hu, Zhirui Zhang, Haoyuan Tian, Xinze Xie, Yufeng Wang 0004, Xiaohua Xie, Zhu Liang Yu, Mingkui Tan
NeurIPS6
2025 Measurement-Based Channel Characterization of Air-to-Ground Propagation for UAV in Coastal Beaches
abstract
In this paper, a miniaturization-and-lightweighting air-to-ground (A2G) channel measurement system is developed based on the software-defined radio (SDR) platform. Utilizing the system, low-altitude A2G channel measurements are conducted in a coastal beaches environment over a 1-kilometer flight distance at the carrier frequency of 1.5 GHz. The collected measurement data is statistically analyzed, and a fitted channel model is proposed. The fitted results are then compared with the free space path loss (FSPL) model and the flat earth two ray (FETR) model, which considers the electromagnetic characteristics of ground reflection points. The error analysis is performed from the perspectives of link distance and the pitch angle between the UAV and the ground station. The results demonstrate that the fitted model provides greater reliability than the FSPL model and the FETR model. Moreover, in the medium pitch angle region, the two-ray model that accounts for actual electromagnetic characteristics exhibits higher accuracy.
Fuyuan Ma, Wenrui Ding, Yufeng Wang 0004, Shutong Wang, Xuetong Pei
VTC2025-Spring3
2025 Frequency-learning adversarial networks based on transfer learning for cross-scenario signal modulation classification
abstract
Automatic modulation classification (AMC) serves a challenging yet crucial role in wireless communications. Despite deep learning-based approaches being widely used in signal processing, they are challenged by signal distribution variations, especially in various channel conditions. In this paper, we introduce an adversarial transfer framework named frequency-learning adversarial networks (FLANs) based on transfer learning for cross-scenario signal classification. This method uses the stability in the frequency spectrum by introducing a frequency adaptation (FA) technique to incorporate target channel information into source-domain signals. To address the unpredictable interference in the channel, a fitting channel adaptation (FCA) module is used to reduce the difference between the source and target domains caused by variations in the channel environment. Experimental results illustrate that FLANs outperforms state-of-the-art transfer approaches, demonstrating an improved top-1 classification accuracy by about 5.2 percentage points in high signal-to-noise ratio (SNR) scenes on a cross-scenario real collected dataset CSRC2023.
Qinyan Ma, Zeqi Shao, Duona Zhang, Yufeng Wang 0004, Wenrui Ding
Frontiers Inf. Technol. Electron. Eng.5
2025 Arch-Net: Model conversion and quantization for architecture agnostic model deployment
Shuangkang Fang, Zipeng Feng, Song Yuan, Yufeng Wang 0004, Yi Yang 0033, Wenrui Ding, Shuchang Zhou 0001
Neural Networks5
2025 Binary Lightweight Neural Networks for Arbitrary Scale Super-Resolution of Remote Sensing Images
abstract
Super-resolution (SR) of remote sensing images (RSIs) has been improved significantly with the development of deep learning. However, better performances usually come from complex network architectures and require a substantial number of parameters. Moreover, many methods can only deal with SR of single and fixed-scale factors. As such, we propose a binary lightweight SR (BLiSR) method to decrease the computation and storage burden and increase the practicality of SR, where we employ a binary neural network (BNN) as the backbone and leverage a binary continuous up-sampling module (BCUM) to achieve arbitrary scale RSI SR. Specifically, we introduce an adaptive binary convolution (ABConv) as the basic unit of BLiSR, which can adaptively adjust the learnable parameters to fit the distribution of full-precision weights and activations. Then, a scalable hyperbolic tangent function is presented to approximate the Sign function in backpropagation and increase the learning capability of BNN. Furthermore, we design a lightweight SR network that considers the full-precision information flow of BNN. The network comprises several basic binary units and a multilayer group fusion block (MGFB), which can extract and fuse the multilevel information from LR images, respectively. Finally, BCUM can predict the pixel values of HR images based on the frequency implicit representation network (IRN) and reconstruct the LR images at arbitrary scales. Extensive experiments on four RSI datasets demonstrate that the proposed BLiSR is superior to several lightweight state-of-the-art (SOTA) methods on both fixed and arbitrary scale SR settings, with a better balance of complexity and performance.
Yufeng Wang 0004, Xianlin Zeng, Wei Li 0022, Wenrui Ding
IEEE Trans. Geosci. Remote. Sens.1
2025 HALO: Hierarchical Adaptive Feature Learning and Cross-View Interaction for Object-Level Geo-Localization
abstract
Cross-view object localization (CVOGL) offers fine-grained geographic information regarding the specified object of interest, in comparison to typical cross-view geo-location at the image level. However, it faces more challenges, primarily due to significant visual appearance changes induced by viewpoint discrepancies and the difficulty of accurately identifying corresponding objects within reference images containing multiple targets. To address these challenges, we propose a novel CVOGL model termed HALO. First, we introduce a cascade cross-view feature interaction module that enables effective local-to-global feature fusion across different viewpoints, thereby enhancing the feature representation of objects in the reference image. Second, to mitigate the scale variations and feature distribution discrepancies between the query and reference images, we propose an adaptive feature hierarchization and aggregate module. It extracts more representative global descriptors by adaptively hierarchizing and aggregating features. Lastly, utilizing the learned global descriptors, we perform cross-view CL and introduce a hard sample mining strategy to further enhance the discriminative ability of the network. Through these advancements, we significantly improve the discriminative representations of the query objects at both views. Extensive experiments on the CVOGL dataset demonstrate the effectiveness and robustness of the proposed method. In the CVOGL task of drone→satellite, HALO improves 2.36% in [email protected] on the test sets. Similarly, HALO shows an increase of 2.49% in [email protected] on the validation set for the CVOGL task of ground→satellite. Our codes are publicly available at https://github.com/ZehaoZhang-Uestc/HALO.
Zehao Zhang, Lei Ding 0008, Yufeng Wang 0004, Wenrui Ding
IEEE Trans. Geosci. Remote. Sens.5
2025 When to Align: Dynamic Behavior Consistency for Multiagent Systems via Intrinsic Rewards
abstract
In multiagent systems, learning optimal behavior policies for individual agents remains a challenging yet crucial task. While recent research has made strides in this area, the issue of when agents should maintain consistent behaviors with one another is still not adequately addressed. This article proposes a novel approach to enable agents to autonomously decide whether their behaviors should align with those of their peers by leveraging intrinsic rewards to optimize their policies. We define behavior consistency as the divergence between the actions taken by two agents given the same observations. To encourage agents to be aware of each other's behaviors, we propose dynamic consistency-based intrinsic reward (DCIR), which guides agents in determining when to synchronize their behaviors. In addition, we introduce a dynamic scaling network (DSN) that provides learnable scaling factors at each time step, enabling agents to dynamically decide the extent of rewarding consistent behavior. Our method is evaluated on environments including Multiagent Particle, Google Research Football, and StarCraft II Micromanagement. Experimental results demonstrate its effectiveness in learning optimal policies.
Kunyang Lin, Yufeng Wang 0004, Peihao Chen, Runhao Zeng, Yinjie Lei, Mingkui Tan, Chuang Gan 0001
IEEE Trans. Neural Networks Learn. Syst.2
2024 Efficient Implicit SDF and Color Reconstruction via Shared Feature Field
Shuangkang Fang, Dacheng Qi, Yufeng Wang 0004, Zehao Zhang, Zeqi Shao, Wenrui Ding
ACCV (10)4
2024 Chat-Edit-3D: Interactive 3D Scene Editing via Text Prompts
Shuangkang Fang, Yufeng Wang 0004, Yi-Hsuan Tsai, Yi Yang 0033, Wenrui Ding, Shuchang Zhou 0001, Ming-Hsuan Yang 0001
ECCV (42)2
2024 Progressive prediction: Video anomaly detection via multi-grained prediction
abstract
Abstract Video Anomaly Detection (VAD) has been an active research field for several decades. However, most existing approaches merely extract a single type of feature from videos and define a single paradigm to indicate the extent of abnormalities. A coarse‐to‐fine three‐level prediction is built by integrating different levels of spatio‐temporal representations, better highlighting the difference between normal and abnormal behaviors. First, an object‐level trajectory prediction is proposed to model human historical position using a graph transformer network. Subsequently, skeleton‐level prediction is achieved by incorporating the positional information from the trajectory prediction. More importantly, based on the predicted skeleton, a skeleton‐guided pixel‐level region prediction is performed. A novel Skeleton Conditioned Generative Adversarial Network (SCGAN) is designed to explore the correlation between skeleton‐level and pixel‐level motion prediction. Benefiting from SCGAN, the prediction of human regions is contributed by both coarse‐grained and fine‐grained motion features. This three‐level prediction, namely Progressive Prediction Video Anomaly Detection (P 3 VAD), enlarges the prediction error on irregular motion patterns. Besides, a pixel‐level analysis method is proposed to achieve Background‐bias Elimination (BE) and denoise the predicted region. Experimental results validate the effectiveness of P 3 VAD on the four benchmark datasets (ShanghaiTech, CUHK Avenue, IITB‐Corridor, and ADOC).
Xianlin Zeng, Yalong Jiang, Yufeng Wang 0004, Wenrui Ding
IET Image Process.3
2024 UAV-ENeRF: Text-Driven UAV Scene Editing With Neural Radiance Fields
abstract
3D reconstruction of Unmanned Aerial Vehicle (UAV) scenes is vital for agriculture, environmental protection, urban planning, and disaster response, to name a few. However, data acquisition can be constrained and hazardous under hostile environments, which limits the image data available in real-world applications. In this work, we propose a text-driven online editing framework for UAV scenes, which can generate novel views of existing scenes with abundant editing types. Compared with small single-object scenes, large-scale UAV scene editing suffers from several particular challenges: 1) broader capturing scope exhibits illumination variation and complicated objects that reduce the 3D scene consistency after editing; and 2) high-resolution 2D editing and 3D reconstruction can be computationally expensive with tremendous GPU memory. To tackle these issues, we first design a dual-branch compact NeRF structure to reduce memory usage and enhance accuracy for 3D reconstruction. We then introduce a sub-pixel sampling scheme to expedite the generation of low-resolution images for 2D editing, followed by a super-resolution module that restores the fine details of rendered images. Additionally, we develop a grouped content filtering mechanism to improve the 3D scene consistency of the model by matching the rendering images and text descriptions, which also significantly reduces memory usage during editing. Extensive experiments demonstrate that the proposed method can achieve various editing effects, including different seasons, weather conditions, times of the day, disaster scenarios, etc. Our technique is computationally efficient and conveniently expandable for large-scale UAV scenes, alleviating data scarcity in harsh scenarios.
Yufeng Wang 0004, Shuangkang Fang, Zehao Zhang, Xianlin Zeng, Wenrui Ding
IEEE Trans. Geosci. Remote. Sens.1
2023 One Is All: Bridging the Gap between Neural Radiance Fields Architectures with Progressive Volume Distillation
abstract
Neural Radiance Fields (NeRF) methods have proved effective as compact, high-quality and versatile representations for 3D scenes, and enable downstream tasks such as editing, retrieval, navigation, etc. Various neural architectures are vying for the core structure of NeRF, including the plain Multi-Layer Perceptron (MLP), sparse tensors, low-rank tensors, hashtables and their compositions. Each of these representations has its particular set of trade-offs. For example, the hashtable-based representations admit faster training and rendering but their lack of clear geometric meaning hampers downstream tasks like spatial-relation-aware editing. In this paper, we propose Progressive Volume Distillation (PVD), a systematic distillation method that allows any-to-any conversions between different architectures, including MLP, sparse or low-rank tensors, hashtables and their compositions. PVD consequently empowers downstream applications to optimally adapt the neural representations for the task at hand in a post hoc fashion. The conversions are fast, as distillation is progressively performed on different levels of volume representations, from shallower to deeper. We also employ special treatment of density to deal with its specific numerical instability problem. Empirical evidence is presented to validate our method on the NeRF-Synthetic, LLFF and TanksAndTemples datasets. For example, with PVD, an MLP-based NeRF model can be distilled from a hashtable-based Instant-NGP model at a 10~20X faster speed than being trained the original NeRF from scratch, while achieving a superior level of synthesis quality. Code is available at https://github.com/megvii-research/AAAI2023-PVD.
Shuangkang Fang, Yi Yang 0033, Yufeng Wang 0004, Shuchang Zhou 0001
AAAI5
2023 DFFG: Fast Gradient Iteration for Data-free Quantization
Huixing Leng, Shuangkang Fang, Yufeng Wang 0004, Zehao Zhang, Dacheng Qi, Wenrui Ding
BMVC3
2023 Multiscale Correlation Networks Based on Deep Learning for Automatic Modulation Classification
abstract
Automatic Modulation Classification (AMC) is a challenging yet significant technique for communication systems. Deep learning methods, though widely employed for AMC, are challenged by the poor representation in noisy scenarios. In this letter, we propose a novel Multiscale Correlation Networks (MSCNs) approach to enhance noise suppression and bolster representation power for AMC. MSCNs leverage learning correlation wavelet transform to redistribute radio signal and noise features across various scales, and incorporate multiscale correlation and attention mechanisms to characterize feature properties in terms of frequency. Experiments reveal that MSCNs achieve an overall classification rate of 91.36% at 10dB using 152k training samples on the public RadioML 2018.01A benchmark.
Yufeng Wang 0004, Duona Zhang, Qinyan Ma, Wenrui Ding
IEEE Signal Process. Lett.2
2022 Towards Accurate Binary Neural Networks via Modeling Contextual Dependencies
Xingrun Xing, Yangguang Li 0001, Wei Li 0022, Wenrui Ding, Yalong Jiang, Yufeng Wang 0004, Chunlei Liu 0001, Xianglong Liu 0001
ECCV (11)6
2022 Semi-supervised Multi-task Learning for Semantics and Depth
abstract
Multi-Task Learning (MTL) aims to enhance the model generalization by sharing representations between related tasks for better performance. Typical MTL methods are jointly trained with the complete multitude of ground-truths for all tasks simultaneously. However, one single dataset may not contain the annotations for each task of interest. To address this issue, we propose the Semi-supervised Multi-Task Learning (SemiMTL) method to leverage the available supervisory signals from different datasets, particularly for semantic segmentation and depth estimation tasks. To this end, we design an adversarial learning scheme in our semi-supervised training by leveraging unlabeled data to optimize all the task branches simultaneously and accomplish all tasks across datasets with partial annotations. We further present a domain-aware discriminator structure with various alignment formulations to mitigate the domain discrepancy issue among datasets. Finally, we demonstrate the effectiveness of the proposed method to learn across different datasets on challenging street view and remote sensing benchmarks.
Yufeng Wang 0004, Yi-Hsuan Tsai, Wei-Chih Hung, Wenrui Ding, Ming-Hsuan Yang 0001
WACV1
2022 Adaptive dense pyramid network for object detection in UAV imagery
Ruiqian Zhang, Xiao Huang 0003, Jiaming Wang 0001, Yufeng Wang 0004, DeRen Li
Neurocomputing5
2022 RB-Net: Training Highly Accurate and Efficient Binary Neural Networks With Reshaped Point-Wise Convolution and Balanced Activation
abstract
In this paper, we find that the conventional convolution operation becomes the bottleneck for extremely efficient binary neural networks (BNNs). To address this issue, we open up a new direction by introducing a reshaped point-wise convolution (RPC) to replace the conventional one to build BNNs. Specifically, we conduct a point-wise convolution after rearranging the spatial information into depth, with which at least$2.25\times $computation reduction can be achieved. Such an efficient RPC allows us to explore more powerful representational capacity of BNNs under a given computation complexity budget. Moreover, we propose to use a balanced activation (BA) to adjust the distribution of the scaled activations after binarization, which enables significant performance improvement of BNNs. After integrating RPC and BA, the proposed network, dubbed as RB-Net, strikes a good trade-off between accuracy and efficiency, achieving superior performance with lower computational cost against the state-of-the-art BNN methods. Specifically, our RB-Net achieves 66.8% Top-1 accuracy with ResNet-18 backbone on ImageNet, exceeding the state-of-the-art Real-to-Binary Net (65.4%) by 1.4% while achieving more than$3\times $reduction (52M vs. 165M) in computational complexity.
Chunlei Liu 0001, Wenrui Ding, Peng Chen 0037, Bohan Zhuang, Yufeng Wang 0004, Yang Zhao 0019, Baochang Zhang 0001, Yuqi Han
IEEE Trans. Circuits Syst. Video Technol.5
2020 Superpixel Labeling Priors and MRF for Aerial Video Segmentation
abstract
Video segmentation is a task of partitioning pixels that exhibit homogeneous appearance and motion into coherent spatial-temporal groups, which is still challenging for aerial applications. In this paper, a principled combination of superpixel labeling priors and Markov random field (S-MRF) is proposed for aerial video segmentation. The proposed approach has several contributions: 1) we develop a metadata-based global projection model with coordinate transformation to estimate motion information between frames; 2) the superpixel labeling priors from previous frames are incorporated into the segmentation of the current frame, leading to a highly efficient probabilistic label propagation algorithm; and 3) we perform an MRF optimization on the initial segments with propagated labeling priors to improve the temporal coherency. In addition, a new video dataset is collected and will be made publicly available to evaluate the performance of aerial video segmentation algorithms. The experimental results show that the proposed approach outperforms the state-of-the-art video segmentation methods.
Yufeng Wang 0004, Wenrui Ding, Baochang Zhang 0001
IEEE Trans. Circuits Syst. Video Technol.1