Rongyu Zhang

dblp:270/4249 · DBLP profile ↗
← Back
23ranked-venue papers
13as first author
22since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 14 · 9 first-author · 13 since 2021Artificial intelligence and machine learning · 13 · 4 first-author · 13 since 2021Computer networks · 5 · 4 first-author · 5 since 2021Systems, architecture and hardware · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Decomposing the Neurons: Activation Sparsity via Mixture of Experts for Continual Test Time Adaptation
abstract
Continual Test-Time Adaptation (CTTA), which aims to adapt the pre-trained model to ever-evolving target domains, emerges as an important task for vision models. As current vision models appear to be heavily biased towards texture, continuously adapting the model from one domain distribution to another can result in serious catastrophic forgetting. Drawing inspiration from the the encoding characteristics of neuron activation in neural networks, we propose the Mixture-of-Activation-Sparsity-Experts (MoASE) for the CTTA task. Given the distinct reaction of neurons with low and high activation to domain-specific and agnostic features, MoASE decomposes the neural activation into high-activation and low-activation components in each expert with a Spatial Differentiable Dropout (SDD). Based on the decomposition, we devise a Domain-Aware Router (DAR) that utilizes domain information to adaptively weight experts that process the post-SDD sparse activations, and the Activation Sparsity Gate (ASG) that adaptively assigns feature selection thresholds of the SDD for different experts for more precise feature decomposition. Finally, we introduce a Homeostatic-Proximal (HP) loss to maintain update consistency between the teacher and student experts to prevent error accumulation. Extensive experiments substantiate that MoASE achieves state-of-the-art performance in both classification and segmentation tasks.
Rongyu Zhang, Aosong Cheng, Yulin Luo, Gaole Dai, Huanrui Yang, Jiaming Liu 0003, Ran Xu 0013, Dan Wang 0002, Yuan Du
AAAI1
2026 MoLe-VLA: Dynamic Layer-skipping Vision Language Action Model via Mixture-of-Layers for Efficient Robot Manipulation
abstract
Vision-Language-Action (VLA) models enable robotic systems to perform embodied tasks but face deployment challenges due to the high computational demands of the dense Large Language Models (LLMs), with existing early-exit-based sparsification methods often overlooking the critical semantic role of final layers in downstream tasks. Aligning with the recent breakthrough of the Shallow Brain Hypothesis (SBH) in neuroscience and the mixture of experts in model sparsification, we conceptualize each LLM layer as an expert and propose a Mixture-of-LayEr Vision Language Action model (MoLe-VLA or simply MoLe) architecture for dynamic LLM layer activation. Specifically, we introduce a Spatial-Temporal Aware Router (STAR) for MoLe to selectively activate only parts of the layers based on the robot’s current state, mimicking the brain's distinct signal pathways specialized for cognition and causal reasoning. Additionally, to compensate for the cognition ability of LLM lost during the layer-skipping, we devise a Cognitive self-Knowledge Distillation (CogKD) to enhance the understanding of task demands and generate task-relevant action sequences by leveraging cognition features. Extensive experiments in RLBench simulations and real-world environments demonstrate the superiority of MoLe-VLA in both efficiency and performance, improving the mean success rate by 9.7% across ten simulation tasks while accelerating inference by 36.8% over OpenVLA.
Rongyu Zhang, Menghang Dong, Yuan Zhang 0020, Liang Heng, Xiaowei Chi, Gaole Dai, Dan Wang 0002, Yuan Du, Shanghang Zhang
AAAI1
2026 PANDA: Empowering Small Language Models for Proactive Dialogue Through Agent-Based Synthesis (Student Abstract)
abstract
Proactive dialogue systems, which are designed to guide conversations toward predetermined goals. However, contemporary LLMs predominantly function as passive assistants, mechanically executing human instructions. A key challenge contributing to this limitation is the inherent difficulty in acquiring and annotating high-quality training data for proactive dialogue. Consequently, the scarcity of such data results in a notable deficiency in the proactive conversational capabilities of current LLMs.In this paper, we introduce PANDA (Proactive Agent-based Negotiation Dialogue Augmentation), a method designed to generate accurate, complex, and diverse proactive dialogue data for a challenging task—financial dispute mediation—where a LLM acts as the mediator. PANDA leverages a novel self-evolving synthesis process to manage a pool of user profiles and generate dialogues through structured interactions between multiple LLM-driven agents. To ensure data fidelity, we propose a comprehensive evaluation framework and build a two-level validation system combining automated and expert human verification. Our experiments demonstrate that an 8B-parameter model, trained on our synthesized dataset, achieves state-of-the-art results in the task's evaluation framework. Its performance rivals top closed-source models guided by heavily engineered prompts, even when provided with only essential information.
Rongyu Zhang, Dingyuan Zhang
AAAI1
2025 PAT: Pruning-Aware Tuning for Large Language Models
abstract
Large language models (LLMs) excel in language tasks, especially with supervised fine-tuning after pre-training. However, their substantial memory and computational requirements hinder practical applications. Structural pruning, which reduces less significant weight dimensions, is one solution. Yet, traditional post-hoc pruning often leads to significant performance loss, with limited recovery from further fine-tuning due to reduced capacity. Since the model fine-tuning refines the general and chaotic knowledge in pre-trained models, we aim to incorporate structural pruning with the fine-tuning, and propose the Pruning-Aware Tuning (PAT) paradigm to eliminate model redundancy while preserving the model performance to the maximum extend. Specifically, we insert the innovative Hybrid Sparsification Modules (HSMs) between the Attention and FFN components to accordingly sparsify the upstream and downstream linear modules. The HSM comprises a lightweight operator and a globally shared trainable mask. The lightweight operator maintains a training overhead comparable to that of LoRA, while the trainable mask unifies the channels to be sparsified, ensuring structural pruning. Additionally, we propose the Identity Loss which decouples the transformation and scaling properties of the HSMs to enhance training robustness. Extensive experiments demonstrate that PAT excels in both performance and efficiency. For example, our Llama2-7b model with a 25% pruning ratio achieves 1.33x speedup while outperforming the LoRA-finetuned model by up to 1.26% in accuracy with a similar training cost.
Yijiang Liu, Huanrui Yang, Youxin Chen, Rongyu Zhang, Yuan Du
AAAI4
2025 Empowering World Models with Reflection for Embodied Video Prediction
abstract
Video generation models have made significant progress in simulating future states, showcasing their potential as world simulators in embodied scenarios. However, existing models often lack robust understanding, limiting their ability to perform multi-step predictions or handle Out-of-Distribution (OOD) scenarios. To address this challenge, we propose the Reflection of Generation (RoG), a set of intermediate reasoning strategies designed to enhance video prediction. It leverages the complementary strengths of pre-trained vision-language and video generation models, enabling them to function as a world model in embodied scenarios. To support RoG, we introduce Embodied Video Anticipation Benchmark(EVA-Bench), a comprehensive benchmark that evaluates embodied world models across diverse tasks and scenarios, utilizing both in-domain and OOD datasets. Building on this foundation, we devise a world model, Embodied Video Anticipator (EVA), that follows a multistage training paradigm to generate high-fidelity video frames and apply an autoregressive strategy to enable adaptive generalization for longer video sequences. Extensive experiments demonstrate the efficacy of EVA in various downstream tasks like video generation and robotics, thereby paving the way for large-scale pre-trained models in real-world video prediction applications. The video demos are available at https://sites.google.com/view/icml-eva.
Xiaowei Chi, Chun-Kai Fan, Xingqun Qi, Rongyu Zhang, Anthony Chen, Chi-Min Chan, Wei Xue 0002, Shanghang Zhang, Yike Guo
ICML5
2025 FBQuant: FeedBack Quantization for Large Language Models
abstract
Deploying Large Language Models (LLMs) on edge devices is increasingly important, as it eliminates reliance on network connections, reduces expensive API calls, and enhances user privacy. However, on-device deployment is challenging due to the limited computational resources of edge devices. In particular, the key bottleneck stems from memory bandwidth constraints related to weight loading. Weight-only quantization effectively reduces memory access, yet often induces significant accuracy degradation. Recent efforts to incorporate sub-branches have shown promise for mitigating quantization errors, but these methods either lack robust optimization strategies or rely on suboptimal objectives. To address these gaps, we propose FeedBack Quantization (FBQuant), a novel approach inspired by negative feedback mechanisms in automatic control. FBQuant inherently ensures that the reconstructed weights remain bounded by the quantization process, thereby reducing the risk of overfitting. To further offset the additional latency introduced by sub-branches, we develop an efficient CUDA kernel that decreases 60% of extra inference time. Comprehensive experiments demonstrate the efficiency and effectiveness of FBQuant across various LLMs. Notably, for 3-bit Llama2-7B, FBQuant improves zero-shot accuracy by 1.2%.
Yijiang Liu, Hengyu Fang, Liulu He, Rongyu Zhang, Yichuan Bai, Yuan Du
IJCAI4
2025 Omni-LLaMA-AD: A Unified Model for Open-Set Visual Anomaly Detection
abstract
Visual anomaly detection (VAD) aims to identify image regions that deviate from established normal patterns. Existing methods often rely on domain-specific training and follow a ''one-class-one-model'' paradigm, limiting scalability. We propose Omni-LLaMA-AD, the first unified multimodal large language model for open-set anomaly detection, capable of handling diverse domains with minimal supervision. Built on a pretrained LLaMA backbone, the model uses a VQGAN-based tokenizer and supports joint vision-language generation. Trained via vision-language alignment and instruction tuning, it achieves effective anomaly detection with only a few normal samples and no domain-specific fine-tuning. Our demo showcases the model's ability to generate high-quality anomaly masks across industrial, medical, and logical datasets, highlighting its strong cross-domain generalization and interactive dialogue-based user experience.
Rongyu Zhang, Zhanbin Hu, Jiamu Wang
ACM Multimedia1
2025 Orochi: Versatile Biomedical Image Processor
abstract
Deep learning has emerged as a pivotal tool for accelerating research in the life sciences, with the low-level processing of biomedical images (e.g., registration, fusion, restoration, super-resolution) being one of its most critical applications. Platforms such as ImageJ (Fiji) and napari have enabled the development of customized plugins for various models. However, these plugins are typically based on models that are limited to specific tasks and datasets, making them less practical for biologists. To address this challenge, we introduce **Orochi**, the first application-oriented, efficient, and versatile image processor designed to overcome these limitations. Orochi is pre-trained on patches/volumes extracted from the raw data of over 100 publicly available studies using our Random Multi-scale Sampling strategy. We further propose Task-related Joint-embedding Pre-Training (TJP), which employs biomedical task-related degradation for self-supervision rather than relying on Masked Image Modelling (MIM), which performs poorly in downstream tasks such as registration. To ensure computational efficiency, we leverage Mamba's linear computational complexity and construct Multi-head Hierarchy Mamba. Additionally, we provide a three-tier fine-tuning framework (Full, Normal, and Light) and demonstrate that Orochi achieves comparable or superior performance to current state-of-the-art specialist models, even with lightweight parameter-efficient options. We hope that our study contributes to the development of an all-in-one workflow, thereby relieving biologists from the overwhelming task of selecting among numerous models. Our pre-trained weights and code will be released.
Gaole Dai, Chenghao Zhou, Rongyu Zhang, Yuan Zhang 0020, Chengkai Hou, Tiejun Huang 0001, Jianxu Chen 0001, Shanghang Zhang
NeurIPS4
2025 BEVUDA++: Geometric-Aware Unsupervised Domain Adaptation for Multi-View 3D Object Detection
abstract
Vision-centric Bird’s Eye View (BEV) perception holds considerable promise for autonomous driving. Recent studies have prioritized efficiency or accuracy enhancements, yet the issue of domain shift has been overlooked, leading to substantial performance degradation upon transfer. We identify major domain gaps in real-world cross-domain scenarios and initiate the first effort to address the Domain Adaptation (DA) challenge in multi-view 3D object detection for BEV perception. Given the complexity of BEV perception approaches with their multiple components, domain shift accumulation across multi-geometric spaces (e.g., 2D, 3D Voxel, BEV) poses a significant challenge for BEV domain adaptation. In this paper, we introduce an innovative geometric-aware teacher-student framework, BEVUDA++, to diminish this issue, comprising a Reliable Depth Teacher (RDT) and a Geometric Consistent Student (GCS) model. Specifically, RDT effectively blends target LiDAR with dependable depth predictions to generate depth-aware information based on uncertainty estimation, enhancing the extraction of Voxel and BEV features that are essential for understanding the target domain. To collaboratively reduce the domain shift, GCS maps features from multiple spaces into a unified geometric embedding space, thereby narrowing the gap in data distribution between the two domains. Additionally, we introduce a novel Uncertainty-guided Exponential Moving Average (UEMA) to further reduce error accumulation due to domain shifts informed by previously obtained uncertainty guidance. To demonstrate the superiority of our proposed method, we execute comprehensive experiments in four cross-domain scenarios, securing state-of-the-art performance in BEV 3D object detection tasks, e.g., 12.9% NDS and 9.5% mAP enhancement on Day-Night adaptation.
Rongyu Zhang, Jiaming Liu 0003, Xiaoqi Li 0009, Xiaowei Chi, Dan Wang 0002, Yuan Du, Shanghang Zhang
IEEE Trans. Circuits Syst. Video Technol.1
2025 Unimodal Training-Multimodal Prediction: Cross-Modal Federated Learning With Hierarchical Aggregation
abstract
Multimodal learning has significantly advanced the extraction of features from varied data sources, enhancing model performance. Federated learning (FL) complements this by enabling collaborative training while maintaining data privacy. The fusion of these two fields, multimodal federated learning, offers considerable promise. Yet, standard methods often incorrectly assume that each node in the FL network has a full complement of multimodal data, which is rare in real-world applications. In our study, we present a novel architecture designed to surmount these challenges, termed the Unimodal Training - Multimodal Prediction (UTMP) framework, positioned within the multimodal federated learning paradigm. Our proposed model, the HA-Fedformer, is a transformer-based model crafted to facilitate unimodal training on the client-side using exclusively unimodal datasets and to execute multimodal inference by synthesizing insights from multiple clients. Our HA-Fedformer model effectively handles non-IID data through a novel uncertainty-aware aggregation technique and layer-wise Markov Chain Monte Carlo sampling in local encoders. It also resolves misaligned language sequences via cross-modal decoder aggregation, capturing correlations between decoders trained on different modalities. Our comprehensive evaluations conducted on widely recognized sentiment analysis benchmarks demonstrate the superiority of the HA-Fedformer. The results show that our model achieves a substantial uplift in performance.
Rongyu Zhang, Xiaowei Chi, Guiliang Liu, Dan Wang 0002, Fangxin Wang 0001
IEEE Trans. Mob. Comput.1
2025 RepCaM++: Exploring Transparent Visual Prompt With Inference-Time Re-Parameterization for Neural Video Delivery
abstract
Recently, content-aware methods have been employed to reduce bandwidth and enhance the quality of Internet video delivery. These methods involve training distinct content-aware super-resolution (SR) models for each video chunk on the server, subsequently streaming the low-resolution (LR) video chunks with the SR models to the client. Prior research has incorporated additional partial parameters to customize the models for individual video chunks. However, this leads to parameter accumulation and can fail to adapt appropriately as video lengths increase, resulting in increased delivery costs and reduced performance. In this paper, we introduce RepCaM++, an innovative framework based on a novel Re- parameterization Content-aware Modulation (RepCaM) module that uniformly modulates video chunks. The RepCaM framework integrates extra parallel-cascade parameters during training to accommodate multiple chunks, subsequently eliminating these additional parameters through re- parameterization during inference. Furthermore, to enhance RepCaM's performance, we propose the Transparent Visual Prompt (TVP), which includes a minimal set of zero-initialized image-level parameters (e.g., less than 0.1%) to capture fine details within video chunks. We conduct extensive experiments on the VSD4K dataset, encompassing six different video scenes, and achieve state-of-the-art results in video restoration quality and delivery bandwidth compression.
Rongyu Zhang, Xize Duan, Jiaming Liu 0003, Yuan Du, Dan Wang 0002, Shanghang Zhang, Fangxin Wang 0001
IEEE Trans. Mob. Comput.1
2024 Efficient Deweahter Mixture-of-Experts with Uncertainty-Aware Feature-Wise Linear Modulation
abstract
The Mixture-of-Experts (MoE) approach has demonstrated outstanding scalability in multi-task learning including low-level upstream tasks such as concurrent removal of multiple adverse weather effects. However, the conventional MoE architecture with parallel Feed Forward Network (FFN) experts leads to significant parameter and computational overheads that hinder its efficient deployment. In addition, the naive MoE linear router is suboptimal in assigning task-specific features to multiple experts which limits its further scalability. In this work, we propose an efficient MoE architecture with weight sharing across the experts. Inspired by the idea of linear feature modulation (FM), our architecture implicitly instantiates multiple experts via learnable activation modulations on a single shared expert block. The proposed Feature Modulated Expert (FME) serves as a building block for the novel Mixture-of-Feature-Modulation-Experts (MoFME) architecture, which can scale up the number of experts with low overhead. We further propose an Uncertainty-aware Router (UaR) to assign task-specific features to different FM modules with well-calibrated weights. This enables MoFME to effectively learn diverse expert functions for multiple tasks. The conducted experiments on the multi-deweather task show that our MoFME outperforms the state-of-the-art in the image restoration quality by 0.1-0.2 dB while saving more than 74% of parameters and 20% inference time over the conventional MoE counterpart. Experiments on the downstream segmentation and classification tasks further demonstrate the generalizability of MoFME to real open-world applications.
Rongyu Zhang, Yulin Luo, Jiaming Liu 0003, Huanrui Yang, Zhen Dong 0003, Denis A. Gudovskiy, Tomoyuki Okuno, Yohei Nakata, Kurt Keutzer, Yuan Du, Shanghang Zhang
AAAI1
2024 BEVUDA: Multi-geometric Space Alignments for Domain Adaptive BEV 3D Object Detection
abstract
Vision-centric bird-eye-view (BEV) perception has shown promising potential in autonomous driving. Recent works mainly focus on improving efficiency or accuracy but neglect the challenges when facing environment changing, resulting in severe degradation of transfer performance. For BEV perception, we figure out the significant domain gaps existing in typical real-world cross-domain scenarios and comprehensively solve the Domain Adaption (DA) problem for multi-view 3D object detection. Since BEV perception approaches are complicated and contain several components, the domain shift accumulation on multiple geometric spaces (i.e., 2D, 3D Voxel, BEV) makes BEV DA even challenging. In this paper, we propose a Multi-space Alignment Teacher-Student (MATS) framework to ease the domain shift accumulation, which consists of a Depth-Aware Teacher (DAT) and a Geometric-space Aligned Student (GAS) model. DAT tactfully combines target lidar and reliable depth prediction to construct depth-aware information, extracting target domain-specific knowledge in Voxel and BEV feature spaces. It then transfers the sufficient domain knowledge of multiple spaces to the student model. In order to jointly alleviate the domain shift, GAS projects multi-geometric space features to a shared geometric embedding space and decreases data distribution distance between two domains. To verify the effectiveness of our method, we conduct BEV 3D object detection experiments on three cross-domain scenarios and achieve state-of-the-art performance. Code: https://github.com/liujiaming1996/BEVUDA.
Jiaming Liu 0003, Rongyu Zhang, Xiaoqi Li 0020, Xiaowei Chi, Ming Lu 0002, Yandong Guo, Shanghang Zhang
ICRA2
2024 VeCAF: Vision-language Collaborative Active Finetuning with Training Objective Awareness
abstract
Finetuning a pretrained vision model (PVM) is a common technique for learning downstream vision tasks. The conventional finetuning process with the randomly sampled data points results in diminished training efficiency. To address this drawback, we propose a novel approach, Vision- languag e C ollaborative A ctive F inetuning (VeCAF). VeCAF optimizes a parametric data selection model by incorporating the training objective of the model being tuned. Effectively, this guides the PVM towards the performance goal with improved data and computational efficiency.With the ever-growing feasibility of acquiring labels and natural language annotations of image data through web-scale crawling, we exploit the inherent semantic richness of the text embedding space and utilize text embeddings of image annotations to augment PVM image features for better data selection and finetuning. Furthermore, the flexibility of text-domain augmentation gives VeCAF the unique ability to handle out-of-distribution scenarios without external augmented data. Extensive experiments show the leading performance and high efficiency of VeCAF that is superior to baselines in both in-distribution and out-of-distribution image classification tasks. On ImageNet, VeCAF needs up to 3.3× less training batches to reach the target performance compared to full fine-tuning and achieves an accuracy improvement of 2.8% over active SOTA fine-tuning methods with the same number of batches. Our code is now available at https://github.com/RoyZry98/VeCAF-Pytorch.
Rongyu Zhang, Zefan Cai, Huanrui Yang, Denis A. Gudovskiy, Tomoyuki Okuno, Yohei Nakata, Kurt Keutzer, Baobao Chang, Yuan Du, Shanghang Zhang
ACM Multimedia1
2024 Multi-Level Personalized Federated Learning on Heterogeneous and Long-Tailed Data
abstract
Federated learning (FL) offers a privacy-centric distributed learning framework, enabling model training on individual clients and central aggregation without necessitating data exchange. Nonetheless, FL implementations often suffer from non-i.i.d. and long-tailed class distributions across mobile applications, e.g., autonomous vehicles, which leads models to overfitting as local training may converge to sub-optimal. In our study, we explore the impact of data heterogeneity on model bias and introduce an innovative personalized FL framework, Multi-level Personalized Federated Learning (MuPFL), which leverages the hierarchical architecture of FL to fully harness computational resources at various levels. This framework integrates three pivotal modules: Biased Activation Value Dropout (BAVD) to mitigate overfitting and accelerate training; Adaptive Cluster-based Model Update (ACMU) to refine local models ensuring coherent global aggregation; and Prior Knowledge-assisted Classifier Fine-tuning (PKCF) to bolster classification and personalize models in accord with skewed local data with shared knowledge. Extensive experiments on diverse real-world datasets for image classification and semantic segmentation validate thatMuPFLconsistently outperforms state-of-the-art baselines, even under extreme non-i.i.d. and long-tail conditions, which enhances accuracy by as much as 7.39% and accelerates training by up to 80% at most, marking significant advancements in both efficiency and effectiveness.
Rongyu Zhang, Chenrui Wu 0002, Fangxin Wang 0001, Bo Li 0001
IEEE Trans. Mob. Comput.1
2023 BEV-SAN: Accurate BEV 3D Object Detection via Slice Attention Networks
abstract
Bird'View (BEV) 3D Object Detection is a crucial multi-view technique for autonomous driving systems. Recently, plenty of works are proposed, following a similar paradigm consisting of three essential components, i.e., camera feature extraction, BEV feature construction, and task heads. Among the three components, BEV feature construction is BEV-specific compared with 2D tasks. Existing methods aggregate the multi-view camera features to the flattened grid in order to construct the BEV feature. However, flattening the BEV space along the height dimension fails to emphasize the informative features of different heights. For example, the barrier is located at a low height while the truck is located at a high height. In this paper, we propose a novel method named BEV Slice Attention Network (BEV-SAN) for exploiting the intrinsic characteristics of different heights. Instead of flattening the BEV space, we first sample along the height dimension to build the global and local BEV slices. Then, the features of BEV slices are aggregated from the camera features and merged by the attention mechanism. Finally, we fuse the merged local and global BEV features by a transformer to generate the final feature map for task heads. The purpose of local BEV slices is to emphasize informative heights. In order to find them, we further propose a LiDAR-guided sampling strategy to leverage the statistical distribution of LiDAR to determine the heights of local slices. Compared with uniform sampling, LiDAR-guided sampling can determine more informative heights. We conduct detailed experiments to demonstrate the effectiveness of BEV-SAN. Code will be released.
Xiaowei Chi, Jiaming Liu 0003, Ming Lu 0002, Rongyu Zhang, Zhaoqing Wang, Yandong Guo, Shanghang Zhang
CVPR4
2023 Cloud-Device Collaborative Adaptation to Continual Changing Environments in the Real-World
abstract
When facing changing environments in the real world, the lightweight model on client devices suffers from severe performance drops under distribution shifts. The main limitations of the existing device model lie in (1) unable to update due to the computation limit of the device, (2) the limited generalization ability of the lightweight model. Meanwhile, recent large models have shown strong generalization capability on the cloud while they can not be deployed on client devices due to poor computation constraints. To enable the device model to deal with changing environments, we propose a new learning paradigm of Cloud-Device Collaborative Continual Adaptation, which encourages collaboration between cloud and device and improves the generalization of the device model. Based on this paradigm, we further propose an Uncertainty-based Visual Prompt Adapted (U-VPA) teacher-student model to transfer the generalization capability of the large model on the cloud to the device model. Specifically, we first design the Uncertainty Guided Sampling (UGS) to screen out challenging data continuously and transmit the most out-of-distribution samples from the device to the cloud. Then we propose a Visual Prompt Learning Strategy with Uncertainty guided updating (VPLU) to specifically deal with the selected samples with more distribution shifts. We transmit the visual prompts to the device and concatenate them with the incoming data to pull the device testing distribution closer to the cloud training distribution. We conduct extensive experiments on two object detection datasets with continually changing environments. Our proposed U-VPA teacher-student framework outperforms previous state-of-the-art test time adaptation and device-cloud collaboration methods. The code and datasets will be released.
Yulu Gan, Mingjie Pan, Rongyu Zhang, Zijian Ling, Lingran Zhao, Jiaming Liu 0003, Shanghang Zhang
CVPR3
2023 DFGC-VRA: DeepFake Game Competition on Visual Realism Assessment
abstract
This paper presents the summary report on the DeepFake Game Competition on Visual Realism Assessment (DFGC-VRA). Deep-learning based face-swap videos, also known as deepfakes, are becoming more and more realistic and deceiving. The malicious usage of these face-swap videos has caused wide concerns. There is a ongoing deepfake game between its creators and detectors, with the human in the loop. The research community has been focusing on the automatic detection of these fake videos, but the assessment of their visual realism, as perceived by human eyes, is still an unexplored dimension. Visual realism assessment, or VRA, is essential for assessing the potential impact that may be brought by a specific face-swap video, and it is also useful as a quality metric to compare different face-swap methods. This is the third edition of DFGC competitions, which focuses on the new visual realism assessment topic, different from previous ones that compete creators versus detectors. With this competition, we conduct a comprehensive study of the SOTA performance on the new task. We also release our MindSpore codes to further facilitate research in this field (https://github.com/bomb2peng/DFGC-VRA-benckmark).
Bo Peng 0002, Xianyun Sun, Caiyong Wang, Wei Wang 0025, Jing Dong 0003, Zhenan Sun, Rongyu Zhang, Heng Cong, Lingzhi Fu, Yusheng Zhang, Boyuan Liu, Luka Dragar, Borut Batagelj, Peter Peer, Vitomir Struc, Xinghui Zhou, Kunlin Liu, Wenxiu Diao
IJCB7
2023 Cluster-driven GNN-based Federated Recommendation with Biased Message Dropout
abstract
Due to the remarkable ability to model the high-order links within user-item relations, the graph neural network (GNN) is gradually applied to personalized recommendations in many online services. Besides, federated learning (FL) recently emerged as a powerful framework that enables collaborative training while protecting user data privacy. However, the integration of GNN and FL still exists with vital challenges unsolved, e.g., learning from non-IID local sub-graphs with only low-order user-item interactions jointly and overcoming the over-fitting problems with high training efficiency. In this paper, we propose CdFed, a Cluster-driven GNN-based Federated Learning framework, to address the GNN+FL challenges. CdFedhas two major components. First, to learn from non-IID sub-graphs, we design an Adaptive Model Clustering (AMC) strategy that takes advantage of the similarity across the uploaded model weights and updates clusters adaptively in each communication round. Second, we develop a Biased Message Dropout (BMD) strategy to combat the overfitting problem and accelerate the training process of federated learning. Together with AMC and BMD, CdFedcan implicitly complete the missing links between sub-graphs more efficiently and greatly improve the model’s generalization ability in non-IID scenarios. We have conducted extensive evaluations and the results reveal that our proposed approaches can outperform the SOTA solution by 24% in model performance and 8x in training speed.
Rongyu Zhang, Chenrui Wu 0002, Fangxin Wang 0001
ICME1
2023 RepCaM: Re-parameterization Content-aware Modulation for Neural Video Delivery
abstract
Recently, content-aware methods have been utilized to reduce the bandwidth and improve the quality of Internet video delivery. Existing methods train corresponding content-aware super-resolution (SR) models for each video chunk on the server and stream low-resolution (LR) video chunks along with SR models to the client. Previous works introduce additional partial parameters to privatize the models of different video chunks. However, this still leads to the accumulation of parameters and even fails to modulate when the length of video increases, bringing extra delivery costs and performance degradation. In this paper, we introduce a novel Re-parameterization Content-aware Modulation (RepCaM) method to modulate all the video chunks with an end-to-end training strategy. Our method adopts extra parallel-cascade parameters during training to fit multiple chunks while removing the additional parameters through re-parameterization during inference. Therefore, RepCaM increases no extra model size compared with the original SR model. Moreover, in order to improve the training efficiency on servers, we propose an online Video Patch Sampling (VPS) method to speed up the training convergence. We conduct extensive experiments on VSD4K and newly collected dataset (VSD4K-2022), achieving state-of-the-art results in video restoration quality and delivery bandwidth compression. Code is available at: https://github.com/Neural-video-delivery/RepCaM-Pytorch-NOSSDAV2023.
Rongyu Zhang, Lixuan Du, Jiaming Liu 0003, Congcong Song, Fangxin Wang 0001, Xiaoqi Li 0009, Ming Lu 0002, Yandong Guo, Shanghang Zhang
NOSSDAV1
2023 FedAB: Truthful Federated Learning With Auction-Based Combinatorial Multi-Armed Bandit
abstract
Federated learning (FL) emerges as a new distributed machine learning (ML) paradigm that enables thousands of mobile devices to collaboratively train ML models using local data without compromising user privacy. However, the FL learning quality highly relies on the data contribution from the distributed mobile devices. Therefore, a well-designed incentive mechanism with effectiveness, fairness, and reciprocity is in urgent need to guarantee the stable participation of users. In this article, we propose federated auction bandit (FedAB), an incentive and client selection strategy based on a novel multiattribute reverse auction mechanism and a combinatorial multi-armed bandit (CMAB) algorithm. First, we develop a local contribution evaluation method based on importance sampling in the FL context. We then design a novel payment mechanism that is able to preserve individual rationality and incentive compatibility (truthfulness). At last, we design a UCB-based winner selection algorithm that is proven to achieve the server’s utility maximization with fairness and reciprocity. We have conducted extensive experiments on real data sets. The results demonstrate the superiority ofFedAB, with a 10%–50% improvement in total reward, final accuracy, and convergence speed compared to state-of-the-art solutions.
Chenrui Wu 0002, Yifei Zhu 0001, Rongyu Zhang, Fangxin Wang 0001, Shuguang Cui
IEEE Internet Things J.3
2022 Multi-Frames Temporal Abnormal Clues Learning Method for Face Anti-Spoofing
abstract
Face anti-spoofing researches are widely used in face recognition and has received more attention from industry and academics.In this paper, we propose the EulerNet, a new temporal feature fusion network in which the differential filter and residual pyramid are used to extract and amplify abnormal clues from continuous frames, respectively.A lightweight sample labeling method based on face landmarks is designed to label large-scale samples at a lower cost and has better results than other methods such as 3D camera.Finally, we collect 30,000 live and spoofing samples using various mobile ends to create a dataset that replicates various forms of attacks in a real-world setting.Extensive experiments on public OULU-NPU show that our algorithm is superior to the state of art and our solution has already been deployed in real-world systems servicing millions of users.
Heng Cong, Rongyu Zhang, Jiarong He
SEKE2
2020 A Dense U-Net with Cross-Layer Intersection for Detection and Localization of Image Forgery
abstract
In this paper, we apply cross-layer intersection mechanism to dense u-net for image forgery detection and localization. We first train DenseNet for binary classification. Spatial rich model (SRM) filters are adopted for capturing residual signals in the detected images. Then we propose a new approach to preserve complete feature maps of fully connected layer and consider them as the spatial decision information for image segmentation. In addition, these features in downsampling path are transferred more effectively and densely to upsampling path through multiscale upsampling and concatenation. A multi-stage training scheme is then applied to improve the convergence of the network. The experimental results show that the proposed method works well on several standard datasets.
Rongyu Zhang, Jiangqun Ni
ICASSP1