Jiafei Wu

dblp:227/7227 · DBLP profile ↗
← Back
39ranked-venue papers
5as first author
37since 2021 · last 2026
0009-0001-8125-1586ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 26 · 1 first-author · 25 since 2021Artificial intelligence and machine learning · 14 · 1 first-author · 14 since 2021Security and privacy · 4 · 4 since 2021Computer networks · 3 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Robust-R1: Degradation-Aware Reasoning for Robust Visual Understanding
abstract
Multimodal Large Language Models struggle to maintain reliable performance under extreme real-world visual degradations, which impede their practical robustness. Existing robust MLLMs predominantly rely on implicit training/adaptation that focuses solely on visual encoder generalization, suffering from limited interpretability and isolated optimization. To overcome these limitations, we propose Robust-R1, a novel framework that explicitly models visual degradations through structured reasoning chains. Our approach integrates: (i) supervised fine-tuning for degradation-aware reasoning foundations, (ii) reward-driven alignment for accurately perceiving degradation parameters, and (iii) dynamic reasoning depth scaling adapted to degradation intensity. To facilitate this approach, we introduce a specialized 11K dataset featuring realistic degradations synthesized across four critical real-world visual processing stages, each annotated with structured chains connecting degradation parameters, perceptual influence, pristine semantic reasoning chain, and conclusion. Comprehensive evaluations demonstrate state-of-theart robustness: Robust-R1 outperforms all general and robust baselines on the real-world degradation benchmark R-Bench, while maintaining superior anti-degradation performance under multi-intensity adversarial degradations on MMMB, MMStar, and RealWorldQA.
Jiaqi Tang 0005, Jianmin Chen, Wei Wei 0008, Xiaogang Xu 0002, Runtao Liu, Qipeng Xie, Jiafei Wu, Lei Zhang 0001, Qifeng Chen 0001
AAAI8
2026 Class Incremental Medical Image Segmentation via Prototype-Guided Calibration and Dual-Aligned Distillation
abstract
Class incremental medical image segmentation (CIMIS) aims to preserve knowledge of previously learned classes while learning new ones without relying on old-class annotations. However, existing methods 1) either adopt one-size-fits-all strategies that treat all spatial regions and feature channels equally, which may hinder the preservation of accurate old knowledge, 2) or focus solely on aligning local prototypes with global ones for old classes while overlooking their local representations in new data, leading to knowledge degradation. To mitigate the above issues, we propose Prototype-Guided Calibration Distillation (PGCD) and Dual-Aligned Prototype Distillation (DAPD) for CIMIS in this paper. Specifically, PGCD exploits prototype-to-feature similarity to calibrate class-specific distillation intensity in different spatial regions, effectively reinforcing reliable old knowledge and suppressing misleading cues from old classes. Complementarily, DAPD aligns the local prototypes of old classes extracted from the current model with both global historical prototypes and local prototypes, further enhancing segmentation performance on old categories. Comprehensive evaluations on two widely used multi-organ segmentation benchmarks demonstrate that our method outperforms current state-of-the-art methods, highlighting its robustness and generalization capabilities.
Shengqian Zhu, Chengrong Yu, Guangjun Li, Jiafei Wu, Xiaogang Xu 0002, Zhang Yi 0001, Junjie Hu 0004
AAAI6
2026 GADT: Enhancing transferable adversarial attacks through gradient-guided adversarial data transformation
Yating Ma, Xiaogang Xu 0002, Liming Fang 0001, Jiafei Wu, Lu Zhou 0002
Neurocomputing4
2026 Attack-agnostic robust decentralized federated learning
Jiafei Wu, Puning Zhao, Haoyi Yuan, Chunhua Su, Lu Zhou 0002
Knowl. Based Syst.1
2026 Semantic Boosting via Knowledge Sharing and Feedback for Video Anomaly Detection
abstract
Vision-language models have the potential to enrich purely visual tasks by utilizing the combined representation of images/videos and corresponding textual descriptions. Recent advances in video anomaly detection have also integrated textual information to enhance the understanding of abnormal events. However, existing approaches often merge visual and textual modalities in a straightforward, bottom-up manner, failing to fully explore their interconnections. Moreover, textual captions themselves do not inherently convey “abnormal” attributes. Consequently, these joint representations tend to highlight all salient input features without adequately focusing on high-level tasks such as video anomaly detection. To direct the model’s attention towards anomalies more effectively, we propose incorporating a top-down mechanism into weakly supervised video anomaly detection tasks. A new Knowledge Sharing and Feedback (KSF) framework is designed to unify the representation of anomalies across both video and text. Specifically, we develop a category pattern sharing module that performs knowledge matching, acting as an alignment bridge between abnormal events and their corresponding descriptions. This ensures consistent representations for identical anomalies while maintaining distinct representations for different ones. Following this alignment process, matched high-level semantic priors are fed back into the forward path to enhance differentiation between abnormal and normal patterns. Comprehensive experiments on three benchmark datasets demonstrate the superiority of our proposed method in learning the implicit definition of anomaly patterns. The code is available at https://github.com/XJ-Cai/KSF.
Xiaojie Cai, Yucheng Qian, Chong Wang 0001, Xiaohao Peng, Yuanbin Qian, Jiafei Wu
IEEE Trans. Circuits Syst. Video Technol.6
2026 Restoration-Oriented Video Frame Interpolation With Region-Distinguishable Priors From SAM
Xiaogang Xu 0002, Yingqi Lin, Jiafei Wu, Zhe Liu 0001, Ming-Hsuan Yang 0001
IEEE Trans. Circuits Syst. Video Technol.4
2026 Learning Suspected Anomalies from Event Prompts for Video Anomaly Detection
abstract
Most models for Weakly Supervised Video Anomaly Detection (WS-VAD) rely on multiple instance learning, aiming to distinguish normal and abnormal snippets without specifying the type of anomaly. However, the ambiguous nature of anomaly definitions across contexts may introduce inaccuracy in discriminating abnormal and normal events. To show the model what is anomalous, a novel framework is proposed to guide the learning of suspected anomalies from event prompts. Given a textual prompt dictionary of potential anomaly events and the captions generated from anomaly videos, the semantic anomaly similarity between them could be calculated to identify the suspected events for each video snippet. It enables a new multi-prompt learning process to constrain the visual-semantic features across all videos, as well as provides a new way to label pseudo anomalies for self-training. To demonstrate its effectiveness, comprehensive experiments and detailed ablation studies are conducted on four datasets, namely XD-Violence, UCF-Crime, TAD, and ShanghaiTech. Our proposed model outperforms most state-of-the-art methods in terms of AP or AUC (86.5%, 90.4%, 94.4%, and 97.4%). Furthermore, it shows promising performance in open-set and cross-dataset cases. The data, code, and models can be found at: https://github.com/shiwoaz/lap .
Chenchen Tao, Xiaohao Peng, Chong Wang 0001, Jiafei Wu, Puning Zhao, Jun Wang 0071, Jiangbo Qian
ACM Trans. Multim. Comput. Commun. Appl.4
2025 UCF-Crime-DVS: A Novel Event-Based Dataset for Video Anomaly Detection with Spiking Neural Networks
abstract
Video anomaly detection plays a significant role in intelligent surveillance systems. To enhance model's anomaly recognition ability, previous works have typically involved RGB, optical flow, and text features. Recently, dynamic vision sensors (DVS) have emerged as a promising technology, which capture visual information as discrete events with a very high dynamic range and temporal resolution. It reduces data redundancy and enhances the capture capacity of moving objects compared to conventional camera. To introduce this rich dynamic information into the surveillance field, we created the first DVS video anomaly detection benchmark, namely UCF-Crime-DVS. To fully utilize this new data modality, a multi-scale spiking fusion network (MSF) is designed based on spiking neural networks (SNNs). This work explores the potential application of dynamic information from event data in video anomaly detection. Our experiments demonstrate the effectiveness of our framework on UCF-Crime-DVS and its superior performance compared to other models, establishing a new baseline for SNN-based weakly supervised video anomaly detection.
Yuanbin Qian, Shuhan Ye, Chong Wang 0001, Xiaojie Cai, Jiangbo Qian, Jiafei Wu
AAAI6
2025 DR-Encoder: Encode Low-rank Gradients with Random Prior for Large Language Models Differentially Privately
abstract
The emergence of the large language model (LLM) has shown its superiority in a wide range of disciplines, including language understanding and translation, relational logic reasoning, and even partial differential equations solving. The transformer is the pervasive backbone architecture for the foundation model construction. It is vital to research how to adjust the Transformer architecture to achieve an end-to-end privacy guarantee in LLM fine-tuning. This paper investigates three potential information leaks during a federated fine-tuning procedure for LLM (FedLLM). Based on the potential information leakage, we insert two-stage randomness into FedLLM to provide an end-to-end privacy guarantee solution. The first stage is to train a gradient auto-encoder with a Gaussian random prior based on the statistical information of the gradients generated by local clients. The second stage is fine-tuning the overall LLM with a differential privacy guarantee by adopting appropriate Gaussian noises. We show our proposed method's efficiency and accuracy gains with several foundation models and two popular evaluation benchmarks. Furthermore, we present a comprehensive privacy analysis with Gaussian Differential Privacy (GDP) and Renyi Differential Privacy (RDP).
Huiwen Wu, Deyi Zhang, Xiaogang Xu 0002, Jiafei Wu, Zhe Liu 0001
AAAI5
2025 Differential Private Stochastic Optimization with Heavy-tailed Data: Towards Optimal Rates
abstract
We study convex optimization problems under differential privacy (DP). With heavy-tailed gradients, existing works achieve suboptimal rates. The main obstacle is that existing gradient estimators have suboptimal tail property, resulting in a superfluous factor of d in the union bound. In this paper, we explore algorithms achieving optimal rates of DP optimization with heavy-tailed gradients. Our first method is a simple clipping approach. Under bounded p-th order moments of gradients, with n samples, it achieves minimax optimal population risk with epsilon less than 1/d. We then propose an iterative updating method, which is more complex but achieves this rate for all epsilon smaller than 1. The results significantly improve over existing methods. Such improvement relies on a careful treatment of the tail behavior of gradient estimators. Our results match the minimax lower bound, indicating that the theoretical limit of stochastic convex optimization under DP is achievable.
Puning Zhao, Jiafei Wu, Zhe Liu 0001, Chong Wang 0001, Rongfei Fan, Qingming Li
AAAI2
2025 Comparing and Improving Frequency Estimation Perturbation Mechanisms Under Local Differential Privacy
She Sun, Jiafei Wu, Huiwen Wu
ACISP (3)2
2025 CG-FedLLM: How to Compress Gradients in Federated Fine-Tuning for Large Language Models
abstract
The success of current Large-Language Models (LLMs) hinges on extensive training data that are collected and stored centrally, called Centralized Learning (CL). However, such a collection manner poses a privacy threat, and one potential solution is Federated Learning (FL), which transfers gradients, not raw data, among clients. Unlike traditional networks, FL for LLMs incurs significant communication costs due to their tremendous parameters. In this study, we introduce an innovative approach to compress gradients to improve communication efficiency during LLM FL, formulating the new FL pipeline named CG-FedLLM. This approach integrates an encoder on the client side to acquire the compressed gradient features and a decoder on the server side to reconstruct the gradients. We also develop a novel training strategy that comprises Temporal-ensemble Gradient-Aware Pre-training (TGAP) to identify characteristic gradients of the target model and Federated AutoEncoder-Involved Fine-tuning (FAF) to compress gradients adaptively. Extensive experiments confirm that our approach reduces communication costs and improves performance (e.g., average 3 points increment compared with traditional CL- and FL-based fine-tuning with several foundation models on well-recognized benchmarks, MMLU and C-Eval). This is because our encoder-decoder, trained via TGAP and FAF, can filter gradients while selectively preserving critical features. Furthermore, we present a series of experimental analyses that focus on the communication efficiency, accuracy, and generalization ability within this privacy-centric framework, providing insights into the development of more efficient and private LLMs fine-tuning.
Huiwen Wu, Xiaogang Xu 0002, Deyi Zhang, Jiafei Wu, Zhe Liu 0001
ECAI5
2025 An Optimized GPU-based Acceleration of CRYSTALS-Dilithium
abstract
CRYSTALS-Dilithium has recently been selected as one of the next generation post-quantum signature algorithm standards. However, due to the extensive volume of data elements and the high complexity of operations, post-quantum cryptographic algorithms commonly face significant performance challenges, for which GPU acceleration has proven to be an effective hardware solution. This paper presents an optimized GPU-based implementation of Dilithium, accelerating the algorithm across three dimensions: inter-thread, intra-thread, and inter-block. Specifically, we propose shared memory and loop unrolling techniques to optimize the most time consuming number theoretic transform operation. Meanwhile, we employ a PTX assembly implementation of Montgomery modular multiplication for frequent dot product computations. Experimental results indicate that, compared to the NIST standard implementation, our implementations of Dilithium across all three parameter sets achieve speedups of up to 143x for Gen, 255x for Sign, and 152x for Verify on the GPU 4090. On the GPU A100, the speedups reach 114x for Gen, 158x for Sign, and 124x for Verify.
Weimin He, Jiafei Wu, Jingjie Liu, Boqin Xu, Xiaoning Bian, Zhe Liu 0001
ICASSP3
2025 Learnable Feature Patches and Vectors for Boosting Low-Light Image Enhancement Without External Knowledge
Xiaogang Xu 0002, Jiafei Wu, Qingsen Yan, Jiequan Cui, Richang Hong, Bei Yu 0001
ICCV2
2025 Enhancing Learning with Label Differential Privacy by Vector Approximation
abstract
Label differential privacy (DP) is a framework that protects the privacy of labels in training datasets, while the feature vectors are public. Existing approaches protect the privacy of labels by flipping them randomly, and then train a model to make the output approximate the privatized label. However, as the number of classes K increases, stronger randomization is needed, thus the performances of these methods become significantly worse. In this paper, we propose a vector approximation approach for learning with label local differential privacy, which is easy to implement and introduces little additional computational overhead. Instead of flipping each label into a single scalar, our method converts each label into a random vector with K components, whose expectations reflect class conditional probabilities. Intuitively, vector approximation retains more information than scalar labels. A brief theoretical analysis shows that the performance of our method only decays slightly with K. Finally, we conduct experiments on both synthesized and real datasets, which validate our theoretical analysis as well as the practical performance of our method.
Puning Zhao, Jiafei Wu, Zhe Liu 0001, Li Shen 0008, Zhikun Zhang 0001, Rongfei Fan, Qingming Li
ICLR2
2025 Low-Light Video Enhancement via Spatial-Temporal Consistent Decomposition
abstract
Low-Light Video Enhancement (LLVE) seeks to restore dynamic or static scenes plagued by severe invisibility and noise. In this paper, we present an innovative video decomposition strategy that incorporates view-independent and view-dependent components to enhance the performance of LLVE. We leverage dynamic cross-frame correspondences for the view-independent term (which primarily captures intrinsic appearance) and impose a scene-level continuity constraint on the view-dependent term (which mainly describes the shading condition) to achieve consistent and satisfactory decomposition results. To further ensure consistent decomposition, we introduce a dual-structure enhancement network featuring a cross-frame interaction mechanism. By supervising different frames simultaneously, this network encourages them to exhibit matching decomposition features. This mechanism can seamlessly integrate with encoder-decoder single-frame networks, incurring minimal additional parameter costs. Extensive experiments are conducted on widely recognized LLVE benchmarks, covering diverse scenarios. Our framework consistently outperforms existing methods, establishing a new SOTA performance.
Xiaogang Xu 0002, Kun Zhou 0001, Tao Hu 0011, Jiafei Wu, Ruixing Wang, Hao Peng 0002, Bei Yu 0001
IJCAI4
2025 HARMONY: A Privacy-preserving and Sensor-agnostic Tele-monitoring system
abstract
Global aging necessitates tele-monitoring systems to provide real-time tracking and timely assistance for older adults living independently. While pervasive wireless devices (e.g., CSI, IMU, UWB) enable cost-effective, non-intrusive monitoring, existing systems lack flexibility, limiting their adaptability to different environments. In this work, we posit that the motion dynamics of human movement are invariant across sensing modalities, inspiring the design of HARMONY—a privacy-preserving, sensor-agnostic system that supports multi-modal inputs and diverse tele-monitoring tasks. HARMONY incorporates Modality-agnostic Data Processing to uniformly encrypt multi-modal signals and Task-specific Activity Recognition for seamless tasks adaptation. A novel Encrypted-processing Engine then significantly accelerates computations on encrypted data by optimizing matrix and convolution operations. Evaluations across five different sensing modalities show that HARMONY consistently achieves high accuracy while delivering 3.5 × to 130 × speedups over state-of-the-art baselines. Our results demonstrate that HARMONY is a practical, scalable, and privacy-centric prototype for next-generation remote healthcare.
Qipeng Xie, Weizheng Wang 0001, Yongzhi Huang 0002, Linshan Jiang, Jiafei Wu, Shuxin Zhong, Lu Wang 0002, Kaishun Wu
IJCAI6
2025 PRIME: Prototype-Driven Class Incremental Learning for Medical Image Segmentation
abstract
Class incremental medical segmentation (CIMS) aims to sequentially learn new classes while preserving knowledge of previously learned categories in the absence of old-class labels. Current methods suffer from performance degradation under class imbalance and require additional segmentation heads to accommodate new categories. Inspired by recent prototype learning that leverages prototypes to achieve robust recognition of new categories under limited-data regimes, we introduce a Prototype-dRIven class increMEntal (PRIME) method. PRIME replaces the incremental segmentation heads with prototypes to mitigate class imbalance, allowing new class learning with the simple addition of new prototypes. Based on prototype learning, PRIME further involves three tailored techniques. First, prototype structure alignment imposes structural constraints on inter-prototype relations to maintain consistent relative distances in the feature space, improving the model's ability to distinguish distinct classes. Second, pixel-wise contrastive loss term groups embeddings of similar samples while separating those of different classes, enhancing segmentation accuracy across all categories. Finally, the consensus-based prototype update mechanism refines the old prototypes during the learning of new classes, preventing performance degradation on the old classes. Extensive experiments on two public multi-organ segmentation datasets demonstrate that our approach significantly outperforms state-of-the-art methods, validating the effectiveness of the proposed PRIME.
Shengqian Zhu, Chengrong Yu, Wenbo Qi, Jiafei Wu, Guangjun Li, Zhang Yi 0001, Xiaogang Xu 0002, Junjie Hu 0004
ACM Multimedia4
2025 Privacy-Preserving LLM Agent for Multi-modal Health Monitoring
Qipeng Xie, Jiafei Wu, Zhuotao Lian, Mu Yuan, Xian Shuai, Weizheng Wang 0001, Yuan Haoyi, Haibo Hu 0001, Kaishun Wu
ProvSec2
2025 An Attack-Agnostic Defense Framework Against Manipulation Attacks Under Local Differential Privacy
abstract
Protection of local differential privacy (LDP) proto-cols against manipulation attacks is an important and challenging problem. We hope to design an attack-agnostic framework, which does not rely on any knowledge of attackers. An early work [1] restricts the attacker's capability by converting each sample into a binary signal. However, the compression of signal leads to severe loss of information, and thus results in unnecessary sacrifice of utility, especially when$\epsilon > 1$. In this paper, we propose a general estimation framework RobustLDP for robust estimation under LDP. The general idea is to send carefully crafted pre-defined information to all users, and then aggregate the feedback at the server. We strike a better tradeoff between preserving information and restricting the attacker's capability. We instantiate RobustLDP for frequency estimation and mean estimation in$\ell_{1}$and$\ell_{2}$support, which serve as building blocks for more advanced tasks. We also establish theoretical guarantees for all possible attacks. The result shows that our method significantly outperforms the existing one for$\epsilon > 1$. Extensive experiments on multiple real-world datasets validate the effectiveness of our method.
Puning Zhao, Zhikun Zhang 0001, Jiafei Wu, Zhe Liu 0001, Shaowei Wang 0003, Yunjun Gao
SP4
2025 CBPSPX: A CUDA-Based Batch Parallel Optimization of Post-Quantum Signature SPHINCS+
abstract
Security and privacy are critical in cloud-based Internet of Things (IoT) and Artificial Intelligence of Things (AIoT) applications. As quantum computing advances, Post-Quantum Cryptography (PQC) has emerged as a key technology for ensuring security in future IoT and AIoT architectures. SPHINCS+, a leading post-quantum signature algorithm, has been selected by the National Institute of Standards and Technology (NIST) as one of the next-generation signature standards. However, due to its complex structure and extensive hash operations, SPHINCS+ suffers from slower signature generation and verification compared to other post-quantum algorithms. Consequently, accelerating SPHINCS+ is essential for adapting it to IoT environments. This paper presents a CUDA-based Batch Parallel optimization of SPHINCS+ (CBPSPX), which fully utilizes the computing resources of NVIDIA Graphics Processing Units (GPUs) to enhance the performance of SPHINCS+. Specifically, we propose the Thread Utilization Efficiency Index (TUEI), which can be used to theoretically evaluate the effectiveness of various parallel methods. Then, we propose an intra-block batch processing model that dynamically adjusts parallel task scales within a block to optimize throughput, making it particularly suitable for IoT scenarios requiring high-throughput large-scale device authentication. Meanwhile, we divide the signature generation process into three sub-processes and adopt different parallel strategies based on the thread requirements of each sub-process to maximize the value of TUEI. For the signature verification process, we propose a columnar storage strategy to replace the traditional row storage structure, which significantly improves the performance of batch signature verification. Experimental results indicate that our SPHINCS+ implementations across all three parameter sets are better than previous optimized GPU-based implementations and achieve speedups of 1.4× to 2.5× for signature generation and 4.6× to 11.3× for signature verification on GPU RTX 3090.
Jiafei Wu, Hao Yang 0062, Zhe Liu 0001
IEEE Internet Things J.1
2025 Robust Federated Learning Under Realistic Corruption: An Iterative Filtering Approach
abstract
Robustness is one of the critical concerns in federated learning. Existing research focuses primarily on the worst case, typically modeled as the Byzantine attack, which alters the gradients in an optimal way. However, in practice, the corruption usually happens randomly, and is much weaker than the Byzantine attack. Therefore, existing methods overestimate the power of corruption, resulting in unnecessary sacrifice of performance. In this article, we build practical algorithms that can withstand realistic corruption, which is weaker than the Byzantine attack, in a better way. Toward this goal, we propose a new iterative filtering approach. In each iteration, it calculates the geometric median of all gradient vectors uploaded from clients and remove the gradients that are far away from the geometric median. A theoretical analysis is then provided, showing that under suitable parameter regimes, gradient vectors from corrupted clients are filtered if the noise is large, while those from benign clients are never filtered throughout the training process. For realistic gradient noise, our approach significantly outperforms existing methods, while the performance under the worst-case attack (i.e., the Byzantine attack) remains nearly the same. Experiments on both synthesized and real data validate our theoretical results, as well as the practical performance of our approach. In particular, we have achieved 3%–10% increase in MNIST and CIFAR10 datasets.
Jiafei Wu, Puning Zhao, Chong Wang 0001, Zhe Liu 0001
IEEE Internet Things J.1
2025 Geometric-Aware Low-Light Image and Video Enhancement via Depth Guidance
abstract
Low-Light Enhancement (LLE) is aimed at improving the quality of photos/videos captured under low-light conditions. It is worth noting that most existing LLE methods do not take advantage of geometric modeling. We believe that incorporating geometric information can enhance LLE performance, as it provides insights into the physical structure of the scene that influences illumination conditions. To address this, we propose a Geometry-Guided Low-Light Enhancement Refine Framework (GG-LLERF) designed to assist low-light enhancement models in learning improved features by integrating geometric priors into the feature representation space. In this paper, we employ depth priors as the geometric representation. Our approach focuses on the integration of depth priors into various LLE frameworks using a unified methodology. This methodology comprises two key novel modules. First, a depth-aware feature extraction module is designed to inject depth priors into the image representation. Then, the Hierarchical Depth-Guided Feature Fusion Module (HDGFFM) is formulated with a cross-domain attention mechanism, which combines depth-aware features with the original image features within LLE models. We conducted extensive experiments on public low-light image and video enhancement benchmarks. The results illustrate that our framework significantly enhances existing LLE methods. The source code and pre-trained models are available at https://github.com/Estheryingqi/GG-LLERF.
Yingqi Lin, Xiaogang Xu 0002, Jiafei Wu, Zhe Liu 0001
IEEE Trans. Image Process.3
2025 Thinking on Context: Inductive Relation Prediction Guided by the Reasoning Ability of Large Language Models
abstract
Inductive relation prediction aims to predict missing connections between entities unseen during training. Recent approaches adopt binary (positive or negative) training labels, which indicate whether the query relation exists between the entities, as supervision to teach models recognizing the entity-independent relation patterns in the context (enclosed subgraph or connective path). However, we argue that in this kind of method, the trained models are guided to make relation predictions by remembering whether the query relation and its contextual relational pattern co-occur more frequently in positive or negative samples. This solution could introduce two major limitations: 1) the model struggles with long-tail combinations, i.e., the combination between query relation and the relational pattern rarely occurs during training; 2) when noisy relational patterns, which fail to provide evidence for predicting the query relation, frequently occur with the query relation in positive training samples, the model will be misled into considering the noisy relational patterns as a feature supporting the existence of the query relation. To solve these problems, we propose ToC (Thinking on Context). ToC first utilizes large language models (LLMs) to incorporate a chain of thought as an additional supervisory constraint, guiding the model to make relational predictions based on logical reasoning instead of co-occurrence frequency. Additionally, ToC employs the reasoning capabilities of LLMs to construct context-level negative samples, aiding the model in identifying and disregarding noisy relational patterns. Extensive experiments show that ToC significantly outperforms state-of-the-art methods across three widely used datasets in multiple inductive settingshttps://github.com/AI-Chen/ToC_KGC.
Xiaoshu Chen, Sihang Zhou 0001, Ke Liang 0006, Jiafei Wu, Xinwang Liu 0002, Dongsheng Li 0001, Kai Lu 0001
IEEE Trans. Knowl. Data Eng.4
2024 Gradient-Aware for Class-Imbalanced Semi-supervised Medical Image Segmentation
Wenbo Qi, Jiafei Wu, S. C. Chan 0001
ECCV (55)2
2024 Efficient Large-Scale Multi-party Computation Based on Garbled Circuit
Zhusen Liu, Jiafei Wu, Zhe Liu 0001
ISPEC2
2024 A Huber Loss Minimization Approach to Mean Estimation under User-level Differential Privacy
abstract
Privacy protection of users' entire contribution of samples is important in distributed systems. The most effective approach is the two-stage scheme, which finds a small interval first and then gets a refined estimate by clipping samples into the interval. However, the clipping operation induces bias, which is serious if the sample distribution is heavy-tailed. Besides, users with large local sample sizes can make the sensitivity much larger, thus the method is not suitable for imbalanced users. Motivated by these challenges, we propose a Huber loss minimization approach to mean estimation under user-level differential privacy. The connecting points of Huber loss can be adaptively adjusted to deal with imbalanced users. Moreover, it avoids the clipping operation, thus significantly reducing the bias compared with the two-stage approach. We provide a theoretical analysis of our approach, which gives the noise strength needed for privacy protection, as well as the bound of mean squared error. The result shows that the new method is much less sensitive to the imbalance of user-wise sample sizes and the tail of sample distributions. Finally, we perform numerical experiments to validate our theoretical analysis.
Puning Zhao, Lifeng Lai, Li Shen 0008, Qingming Li, Jiafei Wu, Zhe Liu 0001
NeurIPS5
2024 scCaT: An explainable capsulating architecture for sepsis diagnosis transferring from single-cell RNA sequencing
abstract
Sepsis is a life-threatening condition characterized by an exaggerated immune response to pathogens, leading to organ damage and high mortality rates in the intensive care unit. Although deep learning has achieved impressive performance on prediction and classification tasks in medicine, it requires large amounts of data and lacks explainability, which hinder its application to sepsis diagnosis. We introduce a deep learning framework, called scCaT, which blends the capsulating architecture with Transformer to develop a sepsis diagnostic model using single-cell RNA sequencing data and transfers it to bulk RNA data. The capsulating architecture effectively groups genes into capsules based on biological functions, which provides explainability in encoding gene expressions. The Transformer serves as a decoder to classify sepsis patients and controls. Our model achieves high accuracy with an AUROC of 0.93 on the single-cell test set and an average AUROC of 0.98 on seven bulk RNA cohorts. Additionally, the capsules can recognize different cell types and distinguish sepsis from control samples based on their biological pathways. This study presents a novel approach for learning gene modules and transferring the model to other data types, offering potential benefits in diagnosing rare diseases with limited subjects.
Xubin Zheng, Dian Meng, Wan-Ki Wong, Ka-Ho To, Lei Zhu 0016, Jiafei Wu, Yining Liang, Kwong-Sak Leung, Man Hon Wong 0001, Lixin Cheng
PLoS Comput. Biol.7
2024 Restructuring the Teacher and Student in Self-Distillation
abstract
Knowledge distillation aims to achieve model compression by transferring knowledge from complex teacher models to lightweight student models. To reduce reliance on pre-trained teacher models, self-distillation methods utilize knowledge from the model itself as additional supervision. However, their performance is limited by the same or similar network architecture between the teacher and student. In order to increase architecture variety, we propose a new self-distillation framework called restructured self-distillation (RSD), which involves restructuring both the teacher and student networks. The self-distilled model is expanded into a multi-branch topology to create a more powerful teacher. During training, diverse student sub-networks are generated by randomly discarding the teacher's branches. Additionally, the teacher and student models are linked by a randomly inserted feature mixture block, introducing additional knowledge distillation in the mixed feature space. To avoid extra inference costs, the branches of the teacher model are then converted back to its original structure equivalently. Comprehensive experiments have demonstrated the effectiveness of our proposed framework for most architectures on CIFAR-10/100 and ImageNet datasets. Code is available at https://github.com/YujieZheng99/RSD.
Chong Wang 0001, Chenchen Tao, Sunqi Lin, Jiangbo Qian, Jiafei Wu
IEEE Trans. Image Process.6
2023 Enlightening the Student in Knowledge Distillation
abstract
Knowledge distillation is a common method of model compression, which uses large models (teacher networks) to guide the training of small models (student networks). However, the student may find a hard time absorbing the knowledge from a sophisticated teacher due to the capacity and confidence gaps between them. To address this issue, a new knowledge distillation and refinement (KDrefine) framework is proposed to enlighten the student by expending and refining its network structure. In addition, a confidence refinement strategy is utilized to generate adaptive soften logits for efficient distillation. The experiments show that the proposed framework outperforms state-of-the-art methods on both CIFAR-100 and Tiny-ImageNet datasets. The code is available at https://github.com/YujieZheng99/KDrefine.
Chong Wang 0001, Yi Chen 0001, Jiangbo Qian, Jun Wang 0071, Jiafei Wu
ICASSP6
2023 Synthetic Feature Assessment for Zero-Shot Object Detection
abstract
Zero-shot object detection aims to simultaneously identify and localize classes that were not presented during training. Many generative model-based methods have shown promising performance by synthesizing the visual features of unseen classes from semantic embeddings. However, these synthetic features are inevitably of varied quality, which may be far from the ground truth. It degrades the performance of trained unseen classifier. Instead of tweaking the generative model, a new idea of feature quality assessment is proposed to utilize both the good and bad features to optimize the classifier in the right direction. Moreover, contrastive learning is also introduced to enhance the feature uniqueness between unseen and seen classes, which helps the feature assessment implicitly. To demonstrate the effectiveness of the proposed algorithm, comprehensive experiments are conducted on the MS COCO dataset and PASCAL VOC dataset, the state-of-the-art performance is achieved. Our code is available at: https://github.com/Dai1029/SFA-ZSD.
Xinmiao Dai, Chong Wang 0001, Haohe Li, Sunqi Lin, Li Dong 0006, Jiafei Wu, Jun Wang 0071
ICME6
2022 Balanced Stripe-Wise Pruning In The Filter
abstract
Neural network pruning offers a promising prospect to compress and accelerate modern deep convolution networks. The stripe-wise pruning method with a finer granularity than traditional methods has become the focus of research. Inspired by the previous work, a new balanced stripe-wise pruning, including the balanced pruning strategy and dynamic pruning threshold, is proposed to achieve higher performance. Specifically, the survived inter-filter stripes and the intra-filter stripes are redistributed by a balanced pruning strategy. Meanwhile the dynamic pruning threshold method makes survival rates will be further balanced across all layers. Comprehensive experiments are conducted on two public datasets (CIFAR-10 and TinyImageNet-200) for different models (ResNet and VGG). The experimental results show that the proposed model is capable of reducing the most parameters, yet achieving the highest accuracy. Our code is available at: https://github.com/ajdt1111/BSWP.
Zheng Huo, Chong Wang 0001, Jun Wang 0071, Jiafei Wu
ICASSP6
2022 Novel Instance Mining with Pseudo-Margin Evaluation for Few-Shot Object Detection
abstract
Few-shot object detection (FSOD) enables the detector to recognize novel objects only using limited training samples, which could greatly alleviate model’s dependency on data. Most existing methods include two training stages, namely base training and fine-tuning. However, the unlabeled novel instances in the base set were untouched in previous works, which can be re-used to enhance the FSOD performance. Thus, a new instance mining model is proposed in this paper to excavate the novel samples from the base set. The detector is thus fine-tuned again by these additional free novel instances. Meanwhile, a novel pseudo-margin evaluation algorithm is designed to address the quality problem of pseudo-labels brought by those new novel instances. The experimental results on MS-COCO dataset show the effectiveness of the proposed model, which does not require any additional training samples or parameters. Our code is available at: https://github.com/liuweijie19980216/NimPme.
Chong Wang 0001, Shenghao Yu, Chenchen Tao, Jun Wang 0071, Jiafei Wu
ICASSP6
2022 Pruning Dynamic Group Convolution with Static Substitute
abstract
Deep learning networks are gradually deployed in edge applications, such as phones and cameras, which has a restriction of the computational resources. Thus, to improve the computational efficiency, numerous types of group convolution-based frameworks have been studied, including the dynamic group convolution (DGC). However, it is worth noting that sometimes the same channels are selected in different dynamic heads in DGC, which violates its original intention. It indicates the number of dynamic heads exceeds what is really needed. In this paper, a novel model by pruning dynamic group convolution with static substitute is proposed. Specifically, those channels pruned in the dynamic head can be reselected in the static head of the group convolution in the proposed model. In addition, a new polarization regularization is introduced to prune more useless channels with less accuracy loss. The experiment results on two image classification benchmarks (CIFAR-I00 and TinylmageNet-200) show promising performance.
Jieyong Che, Chong Wang 0001, Xinmiao Dai, Jun Wang 0071, Jiafei Wu
ICME6
2022 Multi-Scale Continuity-Aware Refinement Network for Weakly Supervised Video Anomaly Detection
abstract
In many previous work, weakly supervised video anomaly detection is formulated as a multiple instance learning (MIL) problem, which represents the video as a bag of multiple instances. However, most MIL-based frameworks only focused on identifying anomalous events from the given instances, without considering the event continuity. Motivated by the fact that abnormal events tend to be more continuous in real-world videos, a Multi-scale Continuity-aware Refinement Network (MCR) is proposed in this paper. It utilizes the property of multi-scale continuity to refine anomaly scores by introducing differential contextual information of instances. At the same time, multi-scale attention is designed to produce a video-level weights in order to select the proper scale and fuse all scores at different scales. Experimental results of MCR show noticeable improvement on two public datasets, specifically obtaining a frame-level AUC 94.92% on ShanghaiTech dataset.
Yiling Gong, Chong Wang 0001, Xinmiao Dai, Shenghao Yu, Lehong Xiang, Jiafei Wu
ICME6
2022 TCA-VAD: Temporal Context Alignment Network for Weakly Supervised Video Anomly Detection
abstract
Video Anomaly detection (VAD) with weakly supervised is usually formulated as a multiple instance learning (MIL) problem. Although the current MIL-based methods have achieved promising detection performance, the temporal dependencies in videos are not well exploited. There may multiple abnormal clips in a given anomaly video, while the previous work only focused on the most abnormal one. To address above issues, a temporal context alignment (TCA) network for video anomaly detection is proposed in this work. Its merits are three-fold, 1) a sparse continuous sampling strategy is proposed to adapt the varying length of untrimmed videos; 2) a multi-scale attention module is used to establish the video temporal dependencies; 3) a top-k loss strategy is used to enlarge the distance between the top-k normal and abnormal clips. Extensive experiments demonstrate the noticeable anomaly discriminability of the proposed network on two public datasets (ShanghaiTech and UCF-Crime).
Shenghao Yu, Chong Wang 0001, Lehong Xiang, Jiafei Wu
ICME4
2021 Cross-Epoch Learning for Weakly Supervised Anomaly Detection in Surveillance Videos
abstract
Weakly Supervised Anomaly Detection (WSAD) in surveillance videos is a complex task since usually only video-level annotations are available. Previous work treated it as a regression problem by giving different scores on normal and anomaly events. However, the widely used mini-batch training strategy may suffer from the data imbalance between these two types of events, which limits the model's performance. In this work, a cross-epoch learning (XEL) strategy associated with a hard instance bank (HIB) is proposed to introduce additional information from previous training epochs. Two new losses are proposed for XEL to achieve a higher detection rate as well as a lower false alarm rate of anomaly events. Moreover, the proposed XEL can be directly integrated into any existing WSAD framework. Experimental results of three XEL embedded models have shown promising AUC improvement (3%~7%) on two public datasets, surpassing the state-of-the-art methods. Our code is available at: https://github.com/sdjsngs/XEL-WSAD.
Shenghao Yu, Chong Wang 0001, Qiao-mei Ma, Jiafei Wu
IEEE Signal Process. Lett.5
2019 Efficient multiplier-less inference of deep autoencoders on wearable healthcare systems
abstract
This paper presents an efficient multiplier-less inference (MLI) approach of deep autoencoders (DAE) for wearable healthcare systems. It employs a novel grouped multiplier block (GMB) module to reduce computational/hardwired complexity of DAE during inference process. First, the fixed weights of DAE are transformed into sum-of-powers-of-two (SOPOT) representations so that multiplications in DAE can be realized as limited adds and shifts only. Further, a GMB is designed to reuse the partial sums in generating the products from the same inputs, which can greatly reduce the adds required. Experimental results show that our proposed MLI method is effective and efficient for wearable healthcare systems to reduce computational/hardwired complexity as well as to offer a faster software implementation.
Jiafei Wu, S. C. Chan 0001, Shuai Zhang 0004
UbiComp1
2018 An Improved Guided Filtering Algorithm for Image Enhancement
abstract
Guided image filter (GIF) is popular in image processing and computer vision for the properties of edge-preserving and low computational complexity, but GIF may suffer from over-smoothing (halo artifacts) near sharp edges and under-smoothing at flat regions. There is a tradeoff between them in the original cost function of GIF. In this paper, an improved guided filter (IGIF) is proposed by incorporating an adaptive structure aware constraint. The adaptive structure aware constraint can well preserve edges and smooth details through assigning different weights to different local structure. Simultaneously, thanks to the L1penalty, the proposed IGIF can exactly remove small details at the flat regions. To illustrate the effectiveness of the proposed IGIF, we apply it to image enhancement. Experimental results show that the proposed filter can produce enhanced images with better visual quality as well as quantitative performance.
Jiafei Wu, Chong Wang 0001, Yongze Xu
ICME1