VLDB 2026 Research / reviewers in the wild / expert
Huiping Zhuang
dblp:194/5829
· DBLP profile ↗
62ranked-venue papers
13as first author
58since 2021 · last 2026
0000-0002-4612-5445ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 40 · 10 first-author · 36 since 2021Graphics, computer vision, multimedia, augmented reality and games · 22 · 3 first-author · 20 since 2021Systems, architecture and hardware · 5 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | RCP-Merging: Merging Long Chain-of-Thought Models with Domain-Specific Models by Considering Reasoning Capability as PriorabstractLarge Language Models (LLMs) with long chain-of-thought (CoT) capability, termed Reasoning Models, demonstrate superior intricate problem-solving abilities through multi-step long CoT reasoning. To create a dual-capability model with long CoT capability and domain-specific knowledge without substantial computational and data costs, model merging emerges as a highly resource-efficient method. However, significant challenges lie in merging domain-specific LLMs with long CoT ones since nowadays merging methods suffer from reasoning capability degradation, even gibberish output and output collapse. To overcome this, we introduce RCP-Merging: Merging Long Chain-of-Thought Models with Domain-Specific Models by Considering Reasoning Capability as Prior, a novel merging framework designed to integrate domain-specific LLMs with long CoT capability, meanwhile maintaining model performance in the original domain. Treating reasoning model weights as foundational prior, our method utilizes a reasoning capability indicator to preserve core long CoT capability model weights while selectively merging essential domain-specific weights. We conducted extensive experiments on Qwen2.5-7B, Llama3.1-8B, and Qwen2.5-1.5B models in BioMedicine and Finance domains. Our results show that RCP-Merging successfully merges a reasoning model with domain-specific ones, improving domain task performance by 9.5% and 9.2% over state-of-the-art methods, without significantly harming the original long CoT reasoning capability. Junyao Yang, Huiping Zhuang, Cen Chen 0002, Ziqian Zeng |
AAAI | 3 |
| 2026 | Rep Deep & Machine Learning: Exemplar-Free Continual Video Action Recognition via Slow-Fast Collaborative LearningabstractIn real-world applications, video action recognition models must continuously learn new action categories while retaining previously acquired knowledge. However, most existing approaches rely on storing historical data for replay, which introduces storage burdens and raises data privacy concerns. To address these challenges, we investigate the problem of Exemplar-Free Continual Video Action Recognition (EF-CVAR) and propose a novel framework named Slow-Fast Collaborative Learning (SFCL). SFCL integrates two complementary learning paradigms: a slow branch based on gradient-driven deep learning, which provides strong adaptability to new tasks, and a fast branch based on analytic learning (e.g., Recursive Least Squares), which efficiently preserves old knowledge without requiring access to past samples. To enable effective collaboration between the two branches, we design the Slow-Fast Dynamic Re-parameterization (SFDR) mechanism for adaptive fusion, and the Knowledge Reflection Mechanism (KRM), which mitigates forgetting and task-recency bias via pseudo-feature generation and dual-level knowledge distillation. Extensive experiments on UCF101, HMDB51, and Something-Something V2 demonstrate that SFCL achieves superior performance compared to existing replay-based methods, despite being exemplar-free. Notably, in long-duration continual learning scenarios, SFCL exhibits remarkable robustness, achieving up to a 30.39\% improvement in accuracy over baselines while maintaining a low forgetting rate, highlighting its scalability and effectiveness in real-world video recognition tasks. Xueyi Zhang 0001, Siqi Cai 0002, Mingrui Lao, Yanming Guo, Huiping Zhuang |
AAAI | 8 |
| 2026 | MixKVQ: Query-Aware Mixed-Precision KV Cache Quantization for Long-Context ReasoningabstractLong Chain-of-Thought (CoT) reasoning has significantly advanced the capabilities of Large Language Models (LLMs), but this progress is accompanied by substantial memory and latency overhead from the extensive Key-Value (KV) cache.Although KV cache quantization is a promising compression technique, existing low-bit quantization methods often exhibit severe performance degradation on complex reasoning tasks.Fixed-precision quantization struggles to handle outlier channels in the key cache, while current mixed-precision strategies fail to accurately identify components requiring high-precision representation.We find that an effective low-bit KV cache quantization strategy must consider two factors: a key channel's intrinsic quantization difficulty and its relevance to the query.Based on this insight, we propose MixKVQ, a novel quantization method that introduces a lightweight, query-aware algorithm to identify and preserve critical key channels that need higher precision, while applying per-token quantization for value cache.Experiments on complex reasoning datasets demonstrate that our approach significantly outperforms existing low-bit methods, achieving performance comparable to a fullprecision baseline at a substantially reduced memory footprint.The source code is available at https://github.com/ZeroNLP/MixKVQ. Tao Zhang 0019, Ziqian Zeng, Huiping Zhuang, Cen Chen 0002 |
ACL (1) | 4 |
| 2026 | LP-VFedNN: A Lightweight and Lossless Privacy-Preserving Vertical Federated Learning Framework For Heterogeneous Neural Network Via Homomorphic Encryption and Intel SGXabstractVertical federated learning (VFL) enhances model performance by jointly leveraging features from multiple parties. However, its intensive interactions increase the risk of privacy leakage. Existing privacy-preserving VFL solutions based on cryptographic primitives suffer from high computation and communication costs, limited algorithmic support, and poor scalability. We propose LP-VFedNN, a lightweight and lossless heterogeneous neural network framework for VFL that integrates CKKS fully homomorphic encryption (FHE) with a trusted execution environment (TEE) and requires no trusted third party. LP-VFedNN adopts a linear-nonlinear separation design: linear operations are executed in the CKKS ciphertext domain, while nonlinear functions are decrypted and accelerated inside the TEE. This avoids the high cost of high-order homomorphic computations and eliminates accuracy degradation caused by polynomial approximations, overcoming the limitations of generalized linear models (GLMs). We further introduce enhanced remote attestation and key agreement to support bidirectional authentication and secure key delivery. For multi-party settings, we propose an adaptive strategy that operates in an efficiency-oriented, restricted-leakage execution mode, improving scalability by offloading substantial TEE computation to secure plaintext computation after decryption. Experiments show that, compared to Paillier-based and pure-TEE schemes, LP-VFedNN significantly reduces communication and computation costs while maintaining accuracy, demonstrating robust scalability with controlled complexity growth as the number of parties increases. Additional experiments on a real-world multimedia dataset further validate its applicability for privacy-aware multimedia retrieval systems. Liji Wu, Baisong Li, Huiping Zhuang, Weiping Wang 0007, Yaoyi Deng, Hailong Zhang 0001 |
ICMR | 5 |
| 2026 | FACT: Feature Adaptive Continual-learning Tracker for multiple object tracking
Rongzihan Song, Zhenyu Weng, Huiping Zhuang, Jinchang Ren, Yongming Chen, Zhiping Lin 0001 |
Knowl. Based Syst. | 3 |
| 2026 | PAL: Prompting analytic learning with missing modality for multi-modal class-incremental learning
Xianghu Yue, Yiming Chen 0010, Xueyi Zhang 0001, Xiaoxue Gao, Mengling Feng, Mingrui Lao, Huiping Zhuang, Haizhou Li 0001 |
Pattern Recognit. | 7 |
| 2026 | Rethinking Obscured Sub-Optimality in Analytic Learning for Exemplar-Free Class-Incremental LearningabstractExemplar-free Class-Incremental Learning (EFCIL) poses a significant challenge in mitigating catastrophic forgetting, due to the absence of exemplars. Recently, analytic learning-based methods propose a recursive alignment procedure to execute EFCIL in a phase-invariant manner and show state-of-the-art performance. However, they heavily rely on a frozen feature extractor trained with the initial dataset to avoid the misalignment between feature and label spaces, ignoring the importance of acquiring generalizable features across incremental tasks for performance improvement. To tackle this, we rethink the obscured sub-optimality of analytic learning-based methods, particularly through empirical reevaluation, and then introduce the Multi-head analytic learning (Muheal) approach. Muheal forms the multi-head model with a delicate feature extractor, thereby introducing a feature optimization procedure and a forgetting compensation module to balance the learning and forgetting. Specifically, within the feature optimization procedure, the feature extractor seeks to learn more generalizable features in a self-supervised manner using the fully-connected classification head. An analytic learning-based classification head follows to align the feature-label space. Additionally, we employ the compensation module to generate and align pseudo-features with a replicated analytic head, thus preventing overfitting and testing. Comprehensive experiments on several benchmark datasets have demonstrated that Muheal significantly outperforms existing state-of-the-art EFCIL methods and is comparable, if not superior, to methods that use replay techniques. Zijian Gao, Kele Xu, Xingxing Zhang 0001, Huiping Zhuang, Tianjiao Wan, Bo Ding 0001, Xinjun Mao, Huaimin Wang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | SEA: Low-Resource Safety Alignment for Multimodal Large Language Models via Synthetic EmbeddingsabstractMultimodal Large Language Models (MLLMs) have serious security vulnerabilities.While safety alignment using multimodal datasets consisting of text and data of additional modalities can effectively enhance MLLM's security, it is costly to construct these datasets.Existing low-resource security alignment methods, including textual alignment, have been found to struggle with the security risks posed by additional modalities.To address this, we propose Synthetic Embedding augmented safety Alignment (SEA), which optimizes embeddings of additional modality through gradient updates to expand textual datasets.This enables multimodal safety alignment training even when only textual data is available.Extensive experiments on image, video, and audio-based MLLMs demonstrate that SEA can synthesize a high-quality embedding on a single RTX3090 GPU within 24 seconds.SEA significantly improves the security of MLLMs when faced with threats from additional modalities.To assess the security risks introduced by video and audio, we also introduced a new benchmark called VA-SafetyBench.High attack success rates across multiple MLLMs validate its challenge.Our code and data will be available at https://github.com/ZeroNLP/SEA.This paper contains harmful data and modelgenerated content that can be offensive in nature. Weikai Lu, Huiping Zhuang, Cen Chen 0002, Ziqian Zeng |
ACL (1) | 3 |
| 2025 | PrivacyRestore: Privacy-Preserving Inference in Large Language Models via Privacy Removal and RestorationabstractZiqian Zeng, Jianwei Wang, Junyao Yang, Zhengdong Lu, Haoran Li, Huiping Zhuang, Cen Chen. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Ziqian Zeng, Junyao Yang, Zhengdong Lu, Huiping Zhuang, Cen Chen 0002 |
ACL (1) | 6 |
| 2025 | GenderAlign: An Alignment Dataset for Mitigating Gender Bias in Large Language ModelsabstractLarge Language Models (LLMs) are prone to generating content that exhibits gender biases, raising significant ethical concerns. Alignment, the process of fine-tuning LLMs to better align with desired behaviors, is recognized as an effective approach to mitigate gender biases. Although proprietary LLMs have made significant strides in mitigating gender bias, their alignment datasets are not publicly available. The commonly used and publicly available alignment dataset, HH-RLHF, still exhibits gender bias to some extent. There is a lack of publicly available alignment datasets specifically designed to address gender bias. Hence, we developed a new dataset named GenderAlign, aiming at mitigating a comprehensive set of gender biases in LLMs. This dataset comprises 8k single-turn dialogues, each paired with a “chosen” and a “rejected” response. Compared to the “rejected” responses, the “chosen” responses demonstrate lower levels of gender bias and higher quality. Furthermore, we categorized the gender biases in the “rejected” responses of GenderAlign into 4 principal categories. The experimental results show the effectiveness of GenderAlign in reducing gender bias in LLMs. Tao Zhang 0019, Ziqian Zeng, YuxiangXiao YuxiangXiao, Huiping Zhuang, Cen Chen 0002, James R. Foulds, Shimei Pan |
ACL (1) | 4 |
| 2025 | AFL: A Single-Round Analytic Approach for Federated Learning with Pre-trained ModelsabstractIn this paper, we introduce analytic federated learning (AFL), a new training paradigm that brings analytical (i.e., closed- form) solutions to the federated learning (FL) with pre-trained models. Our AFL draws inspiration from analytic learning—a gradient-free technique that trains neural networks with analytical solutions in one epoch. In the local client training stage, the AFL facilitates a one-epoch training, eliminating the necessity for multi-epoch updates. In the aggregation stage, we derive an absolute aggregation (AA) law. This AA law allows a single-round aggregation, reducing heavy communication overhead and achieving fast convergence by removing the need for multiple aggregation rounds. More importantly, the AFL exhibits a property that invariance to data partitioning, meaning that regardless of how the full dataset is distributed among clients, the aggregated result remains identical. This could spawn various potentials, such as data heterogeneity invariance and client-number invariance. We conduct experiments across various FL settings including extremely non-IID ones, and scenarios with a large number of clients (e.g., ≥ 1000). In all these settings, our AFL constantly performs competitively while existing FL techniques encounter various obstacles. Our codes are available at https://github.com/ZHUANGHP/Analytic-federated-learning. Run He, Kai Tong, Di Fang 0004, Ziqian Zeng, Huiping Zhuang |
CVPR | 8 |
| 2025 | C-Adapter: Adapting Deep Classifiers for Efficient Conformal Prediction SetsabstractConformal prediction, as an emerging uncertainty quantification technique, typically functions as post-hoc processing for the outputs of trained classifiers. To optimize the classifier for maximum predictive efficiency, Conformal Training rectifies the training objective of base classifiers with a regularization that minimizes the average prediction set size at a specific error rate. However, the regularization term inevitably deteriorates the classification accuracy of classifiers, thereby leading to suboptimal efficiency of conformal predictors. To address this issue, we introduce Conformal Adapter (C-Adapter), an adapter-based tuning method to enhance the efficiency of conformal predictors without sacrificing accuracy. In particular, we implement the adapter as a class of intra order-preserving functions and tune it with our proposed loss that maximizes the discriminability of non-conformity scores between correctly and randomly matched data-label pairs. Using C-Adapter, the model tends to produce higher non-conformity scores for incorrect labels than for correct ones, thereby enhancing predictive efficiency across different coverage rates. Extensive experiments demonstrate that C-Adapter can effectively adapt various classifiers for efficient conformal prediction sets, as well as enhance the conformal training method. Kangdao Liu, Hao Zeng 0005, Jianguo Huang, Huiping Zhuang, Chi-Man Vong, Hongxin Wei |
ECAI | 4 |
| 2025 | RewardDS: Privacy-Preserving Fine-Tuning for Large Language Models via Reward Driven Data SynthesisabstractThe success of large language models (LLMs) has attracted many individuals to fine-tune them for domain-specific tasks by uploading their data.However, in sensitive areas like healthcare and finance, privacy concerns often arise.One promising solution is to generate synthetic data with Differential Privacy (DP) guarantees to replace private data.However, these synthetic data contain significant flawed data, which are considered as noise.Existing solutions typically rely on naive filtering by comparing ROUGE-L scores or embedding similarities, which are ineffective in addressing the noise.To address this issue, we propose RewardDS, a novel privacy-preserving framework that fine-tunes a reward proxy model and uses reward signals to guide the synthetic data generation.Our RewardDS introduces two key modules, Reward Guided Filtering and Self-Optimizing Refinement, to both filter and refine the synthetic data, effectively mitigating the noise.Extensive experiments across medical, financial, and code generation domains demonstrate the effectiveness of our method. Chengming Shi, Junyao Yang, Huiping Zhuang, Cen Chen 0002, Ziqian Zeng |
EMNLP | 6 |
| 2025 | Any-SSR: How Recursive Least Squares Works in Continual Learning of Large Language Models
Kai Tong, Kang Pan, Xiao Zhang 0053, Erli Meng, Run He, Yawen Cui, Nuoyan Guo, Huiping Zhuang |
ICCV | 8 |
| 2025 | Semantic Shift Estimation via Dual-Projection and Classifier Reconstruction for Exemplar-Free Class-Incremental LearningabstractExemplar-Free Class-Incremental Learning (EFCIL) aims to sequentially learn from distinct categories without retaining exemplars but easily suffers from catastrophic forgetting of learned knowledge. While existing EFCIL methods leverage knowledge distillation to alleviate forgetting, they still face two critical challenges: semantic shift and decision bias. Specifically, the embeddings of old tasks shift in the embedding space after learning new tasks, and the classifier becomes biased towards new tasks due to training solely with new data, hindering the balance between old and new knowledge. To address these issues, we propose the Dual-Projection Shift Estimation and Classifier Reconstruction (DPCR) approach for EFCIL. DPCR effectively estimates semantic shift through a dual-projection, which combines a learnable transformation with a row-space projection to capture both task-wise and category-wise shifts. Furthermore, to mitigate decision bias, DPCR employs ridge regression to reformulate a classifier reconstruction process. This reconstruction exploits previous in covariance and prototype of each class after calibration with estimated shift, thereby reducing decision bias. Extensive experiments demonstrate that, on various datasets, DPCR effectively balances old and new tasks, outperforming state-of-the-art EFCIL methods. Our codes are available at https://github.com/RHe502/ICML25-DPCR. Run He, Di Fang 0004, Yawen Cui, Ming Li 0011, Cen Chen 0002, Ziqian Zeng, Huiping Zhuang |
ICML | 8 |
| 2025 | WMarkGPT: Watermarked Image Understanding via Multimodal Large Language ModelsabstractInvisible watermarking is widely used to protect digital images from unauthorized use. Accurate assessment of watermarking efficacy is crucial for advancing algorithmic development. However, existing statistical metrics, such as PSNR, rely on access to original images, which are often unavailable in text-driven generative watermarking and fail to capture critical aspects of watermarking, particularly visibility. More importantly, these metrics fail to account for potential corruption of image content. To address these limitations, we propose WMarkGPT, the first multimodal large language model (MLLM) specifically designed for comprehensive watermarked image understanding, without accessing original images. WMarkGPT not only predicts watermark visibility but also generates detailed textual descriptions of its location, content, and impact on image semantics, enabling a more nuanced interpretation of watermarked images. Tackling the challenge of precise location description and understanding images with vastly different content, we construct three visual question-answering (VQA) datasets: an object location-aware dataset, a synthetic watermarking dataset, and a real watermarking dataset. We introduce a meticulously designed three-stage learning pipeline to progressively equip WMarkGPT with the necessary abilities. Extensive experiments on synthetic and real watermarking QA datasets demonstrate that WMarkGPT outperforms existing MLLMs, achieving significant improvements in visibility prediction and content description. The datasets and code are released at https://github.com/TanSongBai/WMarkGPT. Songbai Tan, Xuerui Qiu, Yao Shu, Linrui Xu, Huiping Zhuang, Ming Li 0011, F. Richard Yu |
ICML | 7 |
| 2025 | L3A: Label-Augmented Analytic Adaptation for Multi-Label Class Incremental LearningabstractClass-incremental learning (CIL) enables models to learn new classes continually without forgetting previously acquired knowledge. Multi-label CIL (MLCIL) extends CIL to a real-world scenario where each sample may belong to multiple classes, introducing several challenges: label absence, which leads to incomplete historical information due to missing labels, and class imbalance, which results in the model bias toward majority classes. To address these challenges, we propose Label-Augmented Analytic Adaptation (L3A), an exemplar-free approach without storing past samples. L3A integrates two key modules. The pseudo-label (PL) module implements label augmentation by generating pseudo-labels for current phase samples, addressing the label absence problem. The weighted analytic classifier (WAC) derives a closed-form solution for neural networks. It introduces sample-specific weights to adaptively balance the class contribution and mitigate class imbalance. Experiments on MS-COCO and PASCAL VOC datasets demonstrate that L3A outperforms existing methods in MLCIL tasks. Our code is available at https://github.com/scut-zx/L3A. Run He, Chen Jiao, Di Fang 0004, Ming Li 0073, Ziqian Zeng, Cen Chen 0002, Huiping Zhuang |
ICML | 8 |
| 2025 | A Bio-inspired Robotic Electric Ray Design of Multimodal Locomotion with Grasping FunctionabstractIn nature, fish locomotion is primarily classified into the BCF (body and caudal fin) propulsion mode and the MPF (median and paired fin) propulsion mode. This paper presents a bio-inspired robotic electric ray that integrates a BCF-mode caudal fin with MPF-mode pectoral fins. The caudal fin consists of a set of wire-driven, multi-joint active segments coupled with a soft, compliant segment, while each symmetrical pectoral fin incorporates two sets of wire-driven joints and a soft fin structure. The undulatory motion of the MPF-mode pectoral fins enables the robotic ray to execute maneuvers such as forward swimming, backward swimming, and in-place turning, whereas the BCF-mode caudal fin enhances linear swimming and turning capabilities. Experimental results demonstrate that MPF-mode swimming achieves a maximum speed of 0.190 m/s (0.358 BL/s), while the cooperative propulsion of MPF and BCF modes enables speeds of up to 0.352 m/s (0.664 BL/s). Notably, the robotic electric ray's large pectoral fins can function as grippers, allowing it to grasp and transport objects using caudal fin propulsion, thereby facilitating object manipulation tasks. Yuyang Mo, Zicun Hong, Huiping Zhuang, Yong Zhong |
IROS | 4 |
| 2025 | Probabilistic Mixture of Hyperbolic Mamba for Few-Shot Class-Incremental LearningabstractFew-shot class-incremental learning (FSCIL) grapples with the dual challenge of learning new classes from minimal labeled training data while alleviating catastrophic forgetting of previous learned classes. Compared with previous methods employing static adaptation on specific parameters, current works verify that dynamic weights and sequence modeling in Selective State Space Models (SSMs) can capture distinctive feature drifts in FSCIL. However, the flattening operation in SSMs fragments the latent semantic relationship, where the resulting task isolation and representation degeneration are detrimental to FSCIL. Toward this issue, this paper presents a novel framework named Probabilistic Mixture of Hyperbolic State Space Experts (PmH-SSE) for FSCIL. First, since SSMs rely on scanning as an alternative to self-attention, the Hyperbolic state space model with multi-scale hybrid scan is built to facilitate few-shot learning by providing an extra Hyperbolic geometry that encodes hierarchical relationships. Moreover, we propose the probabilistic mixture of Mamba to increase the model's flexibility in handling non-stationary data streams in FSCIL and enhance the stability of high-parameter models in few-shot conditions. Finally, under the same experimental conditions, the proposed PmH-SSE demonstrates superior performance in comprehensive experiments. The codes are available at https://github.com/yawencui/PmH-SSE. Yawen Cui, Wenbin Zou, Huiping Zhuang, Yi Wang 0068, Lap-Pui Chau |
ACM Multimedia | 3 |
| 2025 | CFSSeg: Closed-Form Solution for Class-Incremental Semantic Segmentation of 2D Images and 3D Point Cloudsabstract2D images and 3D point clouds are foundational data types for multimedia applications, including real-time video analysis, augmented reality (AR), and 3D scene understanding. Class-incremental semantic segmentation (CSS) requires incrementally learning new semantic categories while retaining prior knowledge. Existing methods typically rely on computationally expensive training based on stochastic gradient descent, employing complex regularization or exemplar replay. However, stochastic gradient descent-based approaches inevitably update the model's weights for past knowledge, leading to catastrophic forgetting, a problem exacerbated by pixel/point-level granularity. To address these challenges, we propose CFSSeg, a novel exemplar-free approach that leverages a closed-form solution, offering a practical and theoretically grounded solution for continual semantic segmentation tasks. This eliminates the need for iterative gradient-based optimization and storage of past data, requiring only a single pass through new samples per step. It not only enhances computational efficiency but also provides a practical solution for dynamic, privacy-sensitive multimedia environments. Extensive experiments on 2D and 3D benchmark datasets such as Pascal VOC2012, S3DIS, and ScanNet demonstrate CFSSeg's superior performance. Jianyu Qi, Songning Lai, Linpu Lv, Kejia Fan, Jianheng Tang 0001, Yutao Yue, Dongzhan Zhou, Yunhuai Liu, Huiping Zhuang |
ACM Multimedia | 11 |
| 2025 | Analytic Continual Test-Time Adaptation for Multi-Modality Corruption
Hongxin Wei, Zhiping Lin 0001, Xiaofeng Zou, Cen Chen 0002, Huiping Zhuang |
ACM Multimedia | 7 |
| 2025 | GraphKeeper: Graph Domain-Incremental Learning via Knowledge Disentanglement and PreservationabstractGraph incremental learning (GIL), which continuously updates graph models by sequential knowledge acquisition, has garnered significant interest recently. However, existing GIL approaches focus on task-incremental and class-incremental scenarios within a single domain. Graph domain-incremental learning (Domain-IL), aiming at updating models across multiple graph domains, has become critical with the development of graph foundation models (GFMs), but remains unexplored in the literature. In this paper, we propose **Graph** Domain-Incremental Learning via **K**nowledge Dis**e**ntangl**e**ment and **P**res**er**vation (**GraphKeeper**), to address catastrophic forgetting in Domain-IL scenario from the perspectives of embedding shifts and decision boundary deviations. Specifically, to prevent embedding shifts and confusion across incremental graph domains, we first propose the domain-specific parameter-efficient fine-tuning together with intra- and inter-domain disentanglement objectives. Consequently, to maintain a stable decision boundary, we introduce deviation-free knowledge preservation to continuously fit incremental domains. Additionally, for graphs with unobservable domains, we perform domain-aware distribution discrimination to obtain precise embeddings. Extensive experiments demonstrate the proposed GraphKeeper achieves state-of-the-art results with 6.5%\~16.6% improvement over the runner-up with negligible forgetting. Moreover, we show GraphKeeper can be seamlessly integrated with various representative GFMs, highlighting its broad applicative potential. Qingyun Sun, Ziwei Zhang 0001, Haonan Yuan, Huiping Zhuang, Xingcheng Fu, Jianxin Li 0002 |
NeurIPS | 5 |
| 2025 | ReFu: Recursive Fusion for Exemplar-Free 3D Class-Incremental LearningabstractWe introduce a novel Recursive Fusion model, dubbed ReFu, designed to integrate point clouds and meshes for exemplar-free 3D Class-Incremental Learning, where the model learns new 3D classes while retaining knowledge of previously learned ones. Unlike existing methods that either rely on storing historical data to mitigate forgetting or focus on single data modalities, ReFu eliminates the need for exemplar storage while utilizing the complementary strengths of both point clouds and meshes. To achieve this, we introduce a recursive method which continuously accumulates knowledge by updating the regularized auto-correlation matrix. Furthermore, we propose a fusion module, featuring a Pointcloud-guided Mesh Attention Layer that learns correlations between the two modalities. This mechanism effectively integrates point cloud and mesh features, leading to more robust and stable continual learning. Experiments across various datasets demonstrate that our proposed framework outperforms existing methods in 3D class-incremental learning. Project Page: https://arlo-yang.github.io/ReFu/ Huiping Zhuang |
WACV | 3 |
| 2025 | PEAR: privacy-preserving and effective aggregation for byzantine-robust federated learning in real-world scenariosabstractAbstract Federated learning (FL) enables collaborative training of global models among distributed clients without sharing local data. Secure aggregation, a new security primitive of FL, enhances the confidentiality of data and model parameters. Unfortunately, privacy-preserving (PP) FL is vulnerable to common poisoning attacks by Byzantine adversaries. Existing defense strategies mainly focus on identifying abnormal local gradients over plaintexts, which provides a weak privacy guarantee. In PPFL, adversaries can escape existing defenses by uploading encrypted poisonous gradients. In addition, most mainstream aggregation algorithms assume that clients’ local training data is uniformly distributed, Independent and Identically Distributed (IID), which is unrealistic for real-world FL scenarios where data are only stored on large-scale terminal devices. To address these issues, we propose PEAR, a PP aggregation strategy based on single key-dual server CKKS full homomorphic encryption in real-world distributed scenarios, which can resist encrypted poisoning attacks. Specifically, we use cosine similarity to measure the distance between encrypted gradients. Then, we propose a novel Byzantine-tolerance aggregation mechanism using cosine similarity, which includes trust score generation that can tolerate differentiated local gradients and a two-step weight generation method that considers both the degree of gradient deviation in direction and training data size. This mechanism can achieve robustness for both IID and non-IID data without compromising privacy. Our extensive evaluations for two typical poisoning attacks on different datasets show that PEAR is robust and effective in IID and non-IID data and outperforms existing mainstream Byzantine-robust algorithms, especially achieving 16.4% to 53.2% testing error rate reduction in non-IID settings with significant label distribution and quantity skew while maintaining the same efficiency as FedAvg. Yan Zhang 0014, Huiping Zhuang, Zhen Xu 0009, Liji Wu |
Comput. J. | 3 |
| 2025 | An analytic formulation of convolutional neural network learning for pattern recognition
Huiping Zhuang, Zhiping Lin 0001, Yimin Yang 0001, Kar-Ann Toh |
Inf. Sci. | 1 |
| 2025 | DRViT: A dynamic redundancy-aware vision transformer accelerator via algorithm and architecture co-design on FPGA
Xiangfeng Sun, Yuan-Ting Zhang, Xiaofeng Zou, Ziqian Zeng, Huiping Zhuang |
J. Parallel Distributed Comput. | 7 |
| 2025 | REAL: Representation enhanced analytic learning for exemplar-free class-incremental learning
Run He, Di Fang 0004, Yizhu Chen, Kai Tong, Cen Chen 0002, Yi Wang 0068, Lap-Pui Chau, Huiping Zhuang |
Knowl. Based Syst. | 8 |
| 2025 | Multi-modality integrated class incremental learning networks for 3D object recognition
Dongyun Lin, Xiao Zhang 0053, Erli Meng, Zhiping Lin 0001, Huiping Zhuang |
Knowl. Based Syst. | 6 |
| 2025 | CrossACL: Analytic Continual Learning via Feature Cross for Hyperspectral Image ClassificationabstractRapidly developing remote sensing technologies expand the volume and variety of hyperspectral images (HSIs). An HSI classification (HSIC) model should be able to adapt to new classes continually while retaining knowledge of previously learned classes to reduce training resources. However, popular HSIC models based on deep neural networks exhibit a significant performance decline in previously learned classes, known as the catastrophic forgetting phenomenon. To efficiently address this issue in HSIC, we propose an analytic continual learning method based on feature cross (CrossACL). CrossACL introduces a novel and training-free feature cross module (FCM) to better adapt to the increasingly complex feature space as the number of HSI classes increases. Furthermore, it utilizes an analytic recursive ridge regression classifier with a closed-form solution. This formulation achieves conditional equivalence between continual learning and joint training on all data seen so far, providing a theoretical guarantee against catastrophic forgetting. In addition, CrossACL introduces a simple but effective oversampling strategy to mitigate classification discrimination due to class-imbalanced HSI samples. Experiments on HSIC datasets demonstrate that CrossACL achieves competitive results compared with state-of-the-art methods at significantly lower computational consumption. Jianan Ji, Yuxuan Cheng, Peiting Xiong, Huiping Zhuang |
IEEE Geosci. Remote. Sens. Lett. | 6 |
| 2025 | Online weighted hashing for cross-modal retrieval
Zining Jiang, Zhenyu Weng, Runhao Li, Huiping Zhuang, Zhiping Lin 0001 |
Pattern Recognit. | 4 |
| 2025 | Subkv: Quantizing Long Context KV Cache for Sub-Billion Parameter Language Models on Edge DevicesabstractABSTRACT Background Large Language Models (LLMs) have demonstrated remarkable capabilities across various tasks. However, their substantial computational and memory requirements present significant challenges for widespread deployment on edge devices. Motivation In long‐context scenarios, even sub‐billion parameter LLMs face unavoidable memory and performance bottlenecks due to inefficient KV Cache utilization. Existing quantization methods fail to address these challenges effectively. Method This paper addresses these challenges by introducing advanced quantization techniques tailored for sub‐billion parameter LLMs. It specifically targets reducing memory consumption through the conversion of the model's KV Cache to lower‐bit integers. We present SubKV, a quantization method specifically designed to optimize the KV Cache in sub‐billion parameter LLMs. Our analysis reveals distinct distributional differences in the magnitude of key and value caches. Leveraging this insight, we apply Per‐Channel Quantization to the key cache and Per‐Token Quantization to the value cache. Furthermore, we introduce the Dynamic Window Quantization method to enhance attention computations. To mitigate the extreme sensitivity of the first token, we also introduce Attention Sink‐Aware Quantization. Results Experimental results demonstrate that SubKV significantly reduces the KV Cache size during long context inference while maintaining model performance, offering superior results to existing KV Cache quantization methods. Ziqian Zeng, Tao Zhang 0019, Zhengdong Lu, Huiping Zhuang, Hongen Shao, Sin G. Teo, Xiaofeng Zou |
Softw. Pract. Exp. | 5 |
| 2025 | Analytic Class Incremental Learning for Sound Source Localization With Privacy ProtectionabstractSound Source Localization (SSL) enabling technology for applications such as surveillance and robotics. While traditional Signal Processing (SP)-based Sound Source Localization (SSL) methods provide analytic solutions under specific signal and noise assumptions, recent Deep Learning (DL)-based methods have significantly outperformed them. However, their success depends on extensive training data and substantial computational resources. Moreover, they often rely on large-scale annotated spatial data and may struggle to adapt to evolving sound classes. To mitigate these challenges, we propose a novel Class Incremental Learning (CIL) approach, termed SSL-CIL, which avoids serious accuracy degradation due tocatastrophic forgettingby incrementally updating the DL-based SSL model through a closed-form analytic solution. In particular, data privacy is ensured since the learning process does not revisit any historical data (exemplar-free), which is more suitable for smart home scenarios. Empirical results in the public SSLR dataset demonstrate the superior performance of our proposal, achieving a localization accuracy of 90.9%, surpassing other competitive methods. Xinyuan Qian 0001, Xianghu Yue, Huiping Zhuang, Haizhou Li 0001 |
IEEE Signal Process. Lett. | 4 |
| 2024 | ConsistentEE: A Consistent and Hardness-Guided Early Exiting Method for Accelerating Language Models InferenceabstractEarly Exiting is one of the most popular methods to achieve efficient inference. Current early exiting methods adopt the (weighted) sum of the cross entropy loss of all internal classifiers as the objective function during training, imposing all these classifiers to predict all instances correctly. However, during inference, as long as one internal classifier predicts an instance correctly, it can accelerate without losing accuracy. Thus, there is a notable gap between training and inference. We propose ConsistentEE, an early exiting method that is consistent in training and inference. ConsistentEE formulates the early exiting process as a reinforcement learning problem. A policy network is added to decide whether an instance should exit or continue. The training objective of ConsistentEE only requires each instance to be predicted correctly by one internal classifier. Additionally, we introduce the concept "Memorized Layer" to measure the hardness of an instance. We incorporate the memorized layer into reward function design, which allows "easy'' instances to focus more on acceleration while ``hard'' instances to focus more on accuracy. Experimental results show that our method outperforms other baselines on various natural language understanding and generation tasks using PLMs and LLMs as backbones respectively. Ziqian Zeng, Yihuai Hong, Huiping Zhuang, Cen Chen 0002 |
AAAI | 4 |
| 2024 | DS-AL: A Dual-Stream Analytic Learning for Exemplar-Free Class-Incremental LearningabstractClass-incremental learning (CIL) under an exemplar-free constraint has presented a significant challenge. Existing methods adhering to this constraint are prone to catastrophic forgetting, far more so than replay-based techniques that retain access to past samples. In this paper, to solve the exemplar-free CIL problem, we propose a Dual-Stream Analytic Learning (DS-AL) approach. The DS-AL contains a main stream offering an analytical (i.e., closed-form) linear solution, and a compensation stream improving the inherent under-fitting limitation due to adopting linear mapping. The main stream redefines the CIL problem into a Concatenated Recursive Least Squares (C-RLS) task, allowing an equivalence between the CIL and its joint-learning counterpart. The compensation stream is governed by a Dual-Activation Compensation (DAC) module. This module re-activates the embedding with a different activation function from the main stream one, and seeks fitting compensation by projecting the embedding to the null space of the main stream's linear mapping. Empirical results demonstrate that the DS-AL, despite being an exemplar-free technique, delivers performance comparable with or better than that of replay-based methods across various datasets, including CIFAR-100, ImageNet-100 and ImageNet-Full. Additionally, the C-RLS' equivalent property allows the DS-AL to execute CIL in a phase-invariant manner. This is evidenced by a never-before-seen 500-phase CIL ImageNet task, which performs on a level identical to a 5-phase one. Our codes are available at https://github.com/ZHUANGHP/Analytic-continual-learning. Huiping Zhuang, Run He, Kai Tong, Ziqian Zeng, Cen Chen 0002, Zhiping Lin 0001 |
AAAI | 1 |
| 2024 | Zero-shot Event Detection Using a Textual Entailment Model as an Enhanced AnnotatorabstractZero-shot event detection is a challenging task. Recent research work proposed to use a pre-trained textual entailment (TE) model on this task. However, those methods treated the TE model as a frozen annotator. We treat the TE model as an annotator that can be enhanced. We propose to use TE models to annotate large-scale unlabeled text and use annotated data to finetune the TE model, yielding an improved TE model. Finally, the improved TE model is used for inference on the test set. To improve the efficiency, we propose to use keywords to filter out sentences with a low probability of expressing event(s). To improve the coverage of keywords, we expand limited number of seed keywords using WordNet, so that we can use the TE model to annotate unlabeled text efficiently. The experimental results show that our method can outperform other baselines by 15% on the ACE05 dataset. Ziqian Zeng, Runyu Wu, Yuxiang Xiao, Xiaoda Zhong, Zhengdong Lu, Huiping Zhuang |
LREC/COLING | 7 |
| 2024 | Mitigating Privacy Risk in Membership Inference by Convex-Concave LossabstractMachine learning models are susceptible to membership inference attacks (MIAs), which aim to infer whether a sample is in the training set. Existing work utilizes gradient ascent to enlarge the loss variance of training data, alleviating the privacy risk. However, optimizing toward a reverse direction may cause the model parameters to oscillate near local minima, leading to instability and suboptimal performance. In this work, we propose a novel method – Convex Concave Loss (CCL), which enables a high variance of training loss distribution by gradient descent. Our method is motivated by the theoretical analysis that convex losses tend to decrease the loss variance during training. Thus, our key idea behind CCL is to reduce the convexity of loss functions with a concave term. Trained with CCL, neural networks produce losses with high variance for training data, reinforcing the defense against MIAs. Extensive experiments demonstrate the superiority of CCL, achieving a state-of-the-art balance in the privacy-utility trade-off. Zhenlong Liu, Lei Feng 0006, Huiping Zhuang, Hongxin Wei |
ICML | 3 |
| 2024 | Complex Motion Planning for Quadruped Robots Using Large Language ModelsabstractLarge language models (LLMs) have shown dominant performance in various language tasks, including code-writing, machine translation, and semantic comprehension. With prompt engineering, LLMs can also comprehend complex tasks and translate them into executable code. These powers offer great potential for controlling the motion of robots. In this paper, we focus on leveraging the ability of LLMs, prompt engineering, and predefined robot action APIs to facilitate high-level motion planning for quadruped robots. With LLMs, we enable the robot to autonomously plan and execute sophisticated actions based on the comprehension of effective prompts. Through various experiments and evaluations, we demonstrate the effectiveness and adaptability of our approach in handling intricate motion tasks. Our research contributes to the advancement of intelligent robotics and paves the way for more versatile quadruped robots in real-world scenarios. Run He, Kai Tong, Shuquan Man, Jingyu Tong, Huiping Zhuang |
ISCAS | 7 |
| 2024 | MMAL: Multi-Modal Analytic Learning for Exemplar-Free Audio-Visual Class Incremental TasksabstractClass-incremental learning poses a significant challenge under an exemplar-free constraint, leading to catastrophic forgetting and sub-par incremental accuracy. Previous attempts have focused primarily on single-modality tasks, such as image classification or audio event classification. However, in the context of Audio-Visual Class-Incremental Learning (AVCIL), the effective integration and utilization of heterogeneous modalities, with their complementary and enhancing characteristics, remains largely unexplored. To bridge this gap, we propose the Multi-Modal Analytic Learning (MMAL) framework, an exemplar-free solution for AVCIL that employs a closed-form, linear approach. To be specific, MMAL introduces a modality fusion module that re-formulates the AVCIL problem through a Recursive Least-Square (RLS) perspective. Complementing this, a Modality-Specific Knowledge Compensation (MSKC) module is designed to further alleviate the under-fitting limitation intrinsic to analytic learning by harnessing individual knowledge from audio and visual modality in tandem. Comprehensive experimental comparisons with existing methods show that our proposed MMAL demonstrates superior performance with the accuracy of 76.71%, 78.98%, and 76.19% on AVE, Kinetics-Sounds, and VGGSounds100 datasets, respectively, setting new state-of-the-art AVCIL performance. Notably, compared to those memory-based methods, our MMAL, being an exemplar-free approach, provides good data privacy and can better leverage multi-modal information for improved incremental accuracy. Xianghu Yue, Xueyi Zhang 0001, Yiming Chen 0010, Mingrui Lao, Huiping Zhuang, Xinyuan Qian 0001, Haizhou Li 0001 |
ACM Multimedia | 6 |
| 2024 | Advancing Cross-domain Discriminability in Continual Learning of Vision-Language ModelsabstractContinual learning (CL) with Vision-Language Models (VLMs) has overcome the constraints of traditional CL, which only focuses on previously encountered classes. During the CL of VLMs, we need not only to prevent the catastrophic forgetting on incrementally learned knowledge but also to preserve the zero-shot ability of VLMs. However, existing methods require additional reference datasets to maintain such zero-shot ability and rely on domain-identity hints to classify images across different domains. In this study, we propose Regression-based Analytic Incremental Learning (RAIL), which utilizes a recursive ridge regression-based adapter to learn from a sequence of domains in a non-forgetting manner and decouple the cross-domain correlations by projecting features to a higher-dimensional space. Cooperating with a training-free fusion module, RAIL absolutely preserves the VLM's zero-shot ability on unseen domains without any reference data.
Additionally, we introduce Cross-domain Task-Agnostic Incremental Learning (X-TAIL) setting. In this setting, a CL learner is required to incrementally learn from multiple domains and classify test images from both seen and unseen domains without any domain-identity hint.
We theoretically prove RAIL's absolute memorization on incrementally learned domains. Experiment results affirm RAIL's state-of-the-art performance in both X-TAIL and existing Multi-domain Task-Incremental Learning settings. The code is released at https://github.com/linghan1997/Regression-based-Analytic-Incremental-Learning. Jiahao Nie 0002, Huiping Zhuang, Manabu Okumura |
NeurIPS | 5 |
| 2024 | GACL: Exemplar-Free Generalized Analytic Continual LearningabstractClass incremental learning (CIL) trains a network on sequential tasks with separated categories in each task but suffers from catastrophic forgetting, where models quickly lose previously learned knowledge when acquiring new tasks. The generalized CIL (GCIL) aims to address the CIL problem in a more real-world scenario, where incoming data have mixed data categories and unknown sample size distribution. Existing attempts for the GCIL either have poor performance or invade data privacy by saving exemplars. In this paper, we propose a new exemplar-free GCIL technique named generalized analytic continual learning (GACL). The GACL adopts analytic learning (a gradient-free training technique) and delivers an analytical (i.e., closed-form) solution to the GCIL scenario. This solution is derived via decomposing the incoming data into exposed and unexposed classes, thereby attaining a weight-invariant property, a rare yet valuable property supporting an equivalence between incremental learning and its joint training. Such an equivalence is crucial in GCIL settings as data distributions among different tasks no longer pose challenges to adopting our GACL. Theoretically, this equivalence property is validated through matrix analysis tools. Empirically, we conduct extensive experiments where, compared with existing GCIL methods, our GACL exhibits a consistently leading performance across various datasets and GCIL settings. Source code is available at https://github.com/CHEN-YIZHU/GACL. Huiping Zhuang, Yizhu Chen, Di Fang 0004, Run He, Kai Tong, Hongxin Wei, Ziqian Zeng, Cen Chen 0002 |
NeurIPS | 1 |
| 2024 | F-OAL: Forward-only Online Analytic Learning with Fast Training and Low Memory Footprint in Class Incremental LearningabstractOnline Class Incremental Learning (OCIL) aims to train models incrementally, where data arrive in mini-batches, and previous data are not accessible. A major challenge in OCIL is Catastrophic Forgetting, i.e., the loss of previously learned knowledge. Among existing baselines, replay-based methods show competitive results but requires extra memory for storing exemplars, while exemplar-free (i.e., data need not be stored for replay in production) methods are resource friendly but often lack accuracy. In this paper, we propose an exemplar-free approach—Forward-only Online Analytic Learning (F-OAL). Unlike traditional methods, F-OAL does not rely on back-propagation and is forward-only, significantly reducing memory usage and computational time. Cooperating with a pre-trained frozen encoder with Feature Fusion, F-OAL only needs to update a linear classifier by recursive least square. This approach simultaneously achieves high accuracy and low resource consumption. Extensive experiments on bench mark datasets demonstrate F-OAL’s robust performance in OCIL scenarios. Code is available at: https://github.com/liuyuchen-cz/F-OAL Huiping Zhuang, Yuchen Liu 0001, Run He, Kai Tong, Ziqian Zeng, Cen Chen 0002, Yi Wang 0068, Lap-Pui Chau |
NeurIPS | 1 |
| 2024 | Joint-Neighborhood Product Quantization for Unsupervised Cross-Modal RetrievalabstractProduct quantization (PQ) is a technique that transforms high-dimensional data into compact binary codes to reduce data storage and improve search efficiency. However, existing PQ methods separate the learning of modality-specific features from the learning of quantization codewords, resulting in suboptimal performance in cross-modal retrieval tasks. In this paper, we propose a joint-neighborhood product quantization (JNPQ) method to simultaneously learn modality-specific features and quantization codewords. To achieve this, we first introduce a cross-modal quantization contrastive learning module that preserves the inter-modal neighborhood of the original data and reduces the quantization error. Then, we design a self-neighbor contrastive learning module that enhances the intra-modal neighborhood within individual modalities. Extensive experiments demonstrate that JNPQ achieves state-of-the-art results in crossmodal retrieval when compared with other unsupervised crossmodal quantization methods. Runhao Li, Zhenyu Weng, Yongming Chen, Huiping Zhuang, Yap-Peng Tan, Zhiping Lin 0001 |
VCIP | 4 |
| 2024 | Explored seeds generation for weakly supervised semantic segmentation
Terence Chow, Haojin Deng, Yimin Yang 0001, Zhiping Lin 0001, Huiping Zhuang, Shan Du 0001 |
Neural Comput. Appl. | 5 |
| 2024 | Less confidence, less forgetting: Learning with a humbler teacher in exemplar-free Class-Incremental learning
Zijian Gao, Kele Xu, Huiping Zhuang, Li Liu 0036, Xinjun Mao, Bo Ding 0001, Huaimin Wang 0001 |
Neural Networks | 3 |
| 2024 | A three-stream fusion and self-differential attention network for multi-modal crowd counting
Haihan Tang, Yi Wang 0068, Zhiping Lin 0001, Lap-Pui Chau, Huiping Zhuang |
Pattern Recognit. Lett. | 5 |
| 2024 | Few-Shot Contrastive Transfer Learning With Pretrained Model for Masked Face VerificationabstractFace verification has seen remarkable progress that benefits from large-scale publicly available databases. However, it remains a challenge how to generalize a pretrained face verification model to a new scenario with a limited amount of data. In many real-world applications, the training database only contains a limited number of identities with two images for each identity due to the privacy concern. In this article, we propose to transfer knowledge from a pretrained unmasked face verification model to a new model for verification between masked and unmasked faces, to meet the application requirements during the COVID-19 pandemic. To overcome the lack of intra-class diversity resulting from only a pair of masked and unmasked faces for each identity ($\text{i.e.},$two shots for each identity), a static prototype classification function is designed to learn features for masked faces by utilizing unmasked face knowledge from the pretrained model. Meanwhile, a contrastive constrained embedding function is designed to preserve unmasked face knowledge of the pretrained model during the transfer learning process. By combining these two functions, our method uses knowledge acquired from the pretrained unmasked face verification model to proceed with verification between masked and unmasked faces with a limited amount of training data. Extensive experiments demonstrate that our method can perform better than state-of-the-art methods for verification between masked and unmasked faces in the few-shot transfer learning setting. Zhenyu Weng, Huiping Zhuang, Fulin Luo, Haizhou Li 0001, Zhiping Lin 0001 |
IEEE Trans. Multim. | 2 |
| 2023 | GKEAL: Gaussian Kernel Embedded Analytic Learning for Few-Shot Class Incremental TaskabstractFew-shot class incremental learning (FSCIL) aims to address catastrophic forgetting during class incremental learning in a few-shot learning setting. In this paper, we approach the FSCIL by adopting analytic learning, a technique that converts network training into linear problems. This is inspired by the fact that the recursive implementation (batch-by-batch learning) of analytic learning gives identical weights to that produced by training on the entire dataset at once. The recursive implementation and the weight-identical property highly resemble the FSCIL setting (phase-by-phase learning) and its goal of avoiding catastrophic forgetting. By bridging the FSCIL with the analytic learning, we propose a Gaussian kernel embedded analytic learning (GKEAL) for FSCIL. The key components of GKEAL include the kernel analytic module which allows the GKEAL to conduct FSCIL in a recursive manner, and the augmented feature concatenation module that balances the preference between old and new tasks especially effectively under the few-shot setting. Our experiments show that the GKEAL gives state-of-the-art performance on several benchmark datasets. Huiping Zhuang, Zhenyu Weng, Run He, Zhiping Lin 0001, Ziqian Zeng |
CVPR | 1 |
| 2023 | Mitigating Memorization of Noisy Labels by Clipping the Model PredictionabstractIn the presence of noisy labels, designing robust loss functions is critical for securing the generalization performance of deep neural networks. Cross Entropy (CE) loss has been shown to be not robust to noisy labels due to its unboundedness. To alleviate this issue, existing works typically design specialized robust losses with the symmetric condition, which usually lead to the underfitting issue. In this paper, our key idea is to induce a loss bound at the logit level, thus universally enhancing the noise robustness of existing losses. Specifically, we propose logit clipping (LogitClip), which clamps the norm of the logit vector to ensure that it is upper bounded by a constant. In this manner, CE loss equipped with our LogitClip method is effectively bounded, mitigating the overfitting to examples with noisy labels. Moreover, we present theoretical analyses to certify the noise-tolerant ability of LogitClip. Extensive experiments show that LogitClip not only significantly improves the noise robustness of CE loss, but also broadly enhances the generalization performance of popular robust losses. Hongxin Wei, Huiping Zhuang, Renchunzi Xie, Lei Feng 0006, Gang Niu 0001, Bo An 0001, Yixuan Li 0001 |
ICML | 2 |
| 2023 | Neighborhood Learning from Noisy Labels for Cross-Modal RetrievalabstractCross-modal retrieval methods are developed to retrieve relevant data across different modalities. Usually, super-vised cross-modal retrieval methods can achieve higher accuracy than unsupervised methods because they can utilize the semantic information provided by clean labels. However, training data with noisy labels will lead to the performance degradation of supervised cross-modal retrieval methods. In this work, we present a novel framework called Neighborhood Learning for Cross-Modal Retrieval (NLCMR) that is robust against noisy labels by exploiting the information contained in the neighbor-hood. Our NLCMR contains two main components: Clustering with Neighborhood Alignment and Neighborhood Contrastive Learning. The first component focuses on reducing the impact of noisy labels and improving clustering robustness, and the second component learns from noisy data by exploring pairwise and neighborhood information. Extensive experiments are conducted on three multi-modal datasets to demonstrate the effectiveness of NLCMR. Runhao Li, Zhenyu Weng, Huiping Zhuang, Yongming Chen, Zhiping Lin 0001 |
ISCAS | 3 |
| 2023 | Online Multi-Face Tracking With Multi-Modality Cascaded MatchingabstractTracking multiple faces online in unconstrained videos is a challenging problem as faces may appear drastically different over time and identities can be inferred only based on information available from past frames. Previous tracking methods focus on face information without reference to other modality information such as a person’s overall body appearance, leading to suboptimal performance. In this paper, we propose a new online multi-face tracking method, called online multi-face tracking with multi-modality cascaded matching (OMTMCM), to improve the tracking performance by using both face and body information. The proposed OMTMCM consists of two stages, namely detection alignment and detection association. In the first stage, a detection alignment module is designed to align face detection with body detection from the same person for the subsequent detection association. In the second stage, a cascaded matching module is designed to associate face detections across frames to locate trajectory of each target face by using both face and body information. Specifically, aligned face-body detections in the current frame are matched in a cascade manner with body and face features that are selected from past frames and stored in the designed feature memory. In this way, our method can track multiple faces online with both face and body information while eliminating the possibility of face detection and body detection from the same person being separately assigned with different identities. Experimental results demonstrate our method is on par with or better than other online tracking methods for multi-face tracking. Zhenyu Weng, Huiping Zhuang, Haizhou Li 0001, Balakrishnan Ramalingam, Mohan Rajesh Elara, Zhiping Lin 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2023 | Attention Multihop Graph and Multiscale Convolutional Fusion Network for Hyperspectral Image ClassificationabstractConvolutional neural networks (CNNs) for hyperspectral image (HSI) classification have generated good progress. Meanwhile, graph convolutional networks (GCNs) have also attracted considerable attention by using unlabeled data, broadly and explicitly exploiting correlations between adjacent parcels. However, the CNN with a fixed square convolution kernel is not flexible enough to deal with irregular patterns, while the GCN using the superpixel to reduce the number of nodes will lose the pixel-level features, and the features from the two networks are always partial. In this paper, to make good use of the advantages of CNN and GCN, we propose a novel multiple feature fusion model termed attention multi-hop graph and multi-scale convolutional fusion network (AMGCFN), which includes two sub-networks of multi-scale fully CNN and multi-hop GCN to extract the multi-level information of HSI. Specifically, the multi-scale fully CNN aims to comprehensively capture pixel-level features with different kernel sizes, and a multi-head attention fusion module is used to fuse the multi-scale pixel-level features. The multi-hop GCN systematically aggregates the multi-hop contextual information by applying multi-hop graphs on different layers to transform the relationships between nodes, and a multi-head attention fusion module is adopted to combine the multi-hop features. Finally, we design a cross attention fusion module to adaptively fuse the features of two sub-networks. AMGCFN makes full use of multi-scale convolution and multi-hop graph features, which is conducive to the learning of multi-level contextual semantic features. Experimental results on three benchmark HSI datasets show that AMGCFN has better performance than a few state-of-the-art methods. Fulin Luo, Huiping Zhuang, Zhenyu Weng, Xiuwen Gong, Zhiping Lin 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | Multiple Temporal Context Embedding Networks for Unsupervised time Series Anomaly DetectionabstractUnsupervised anomaly detection for time series signals is challenging, due to the imbalanced distribution of data and the lack of ground-truth labels. Current methods on this topic are mainly based on deep neural networks, which are optimized by heuristic constraints or empirical priors. However, various patterns of anomalous data, especially those lasting for varying periods, are hard to be captured by plain networks. To tackle this problem, we propose a multiple temporal context embedding method. The core of our method is to construct a unified representation of the multiple temporal contexts of data, which is achieved by learning a set of base features to reconstruct the hidden features within existing anomaly detection networks. The proposed method can be implemented as a convenient plug-in module, and be combined with various network architectures, such as autoencoders and graph neural networks. Extensive experiments on multiple datasets demonstrate that the proposed method can boost the performance of baseline networks significantly. Xinggan Peng, Huiping Zhuang, Zhiping Lin 0001 |
ICASSP | 3 |
| 2022 | ACIL: Analytic Class-Incremental Learning with Absolute Memorization and Privacy ProtectionabstractClass-incremental learning (CIL) learns a classification model with training data of different classes arising progressively. Existing CIL either suffers from serious accuracy loss due to catastrophic forgetting, or invades data privacy by revisiting used exemplars. Inspired by learning of linear problems, we propose an analytic class-incremental learning (ACIL) with absolute memorization of past knowledge while avoiding breaching of data privacy (i.e., without storing historical data). The absolute memorization is demonstrated in the sense that the CIL using ACIL given present data would give identical results to that from its joint-learning counterpart that consumes both present and historical samples. This equality is theoretically validated. The data privacy is ensured by showing that no historical data are involved during the learning process. Empirical validations demonstrate ACIL's competitive accuracy performance with near-identical results for various incremental task settings (e.g., 5-50 phases). This also allows ACIL to outperform the state-of-the-art methods for large-phase scenarios (e.g., 25 and 50 phases). Huiping Zhuang, Zhenyu Weng, Hongxin Wei, Renchunzi Xie, Kar-Ann Toh, Zhiping Lin 0001 |
NeurIPS | 1 |
| 2022 | Fully Decoupled Neural Network Learning Using Delayed GradientsabstractTraining neural networks with backpropagation (BP) requires a sequential passing of activations and gradients. This has been recognized as the lockings (i.e., the forward, backward, and update lockings) among modules (each module contains a stack of layers) inherited from the BP. In this brief, we propose a fully decoupled training scheme using delayed gradients (FDG) to break all these lockings. The FDG splits a neural network into multiple modules and trains them independently and asynchronously using different workers (e.g., GPUs). We also introduce a gradient shrinking process to reduce the stale gradient effect caused by the delayed gradients. Our theoretical proofs show that the FDG can converge to critical points under certain conditions. Experiments are conducted by training deep convolutional neural networks to perform classification tasks on several benchmark data sets. These experiments show comparable or better results of our approach compared with the state-of-the-art methods in terms of generalization and acceleration. We also show that the FDG is able to train various networks, including extremely deep ones (e.g., ResNet-1202), in a decoupled fashion. Huiping Zhuang, Yi Wang 0068, Qinglai Liu, Zhiping Lin 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2022 | Blockwise Recursive Moore-Penrose Inverse for Network LearningabstractTraining neural networks with the Moore–Penrose (MP) inverse has recently gained attention in view of its noniterative training nature. However, a significant drawback of learning based on the MP inverse is that the computational memory consumption grows along with the size of a dataset. In this article, based on the partitioning of the MP inverse, we propose a blockwise recursive MP inverse formulation (BRMP) for network learning with low-memory property while preserving its training effectiveness. The BRMP is an equivalent formulation to its batchwise counterpart since neither approximation nor assumption is made in the derivation process. Our further exploration of this recursive method leads to a switching structure among three different scenarios. This structure also reveals that the well-known recursive least squares method is a special case of our proposed technique. Subsequently, we apply BRMP to the training of radial basis function networks as well as multilayer perceptrons. The experimental validation covers both regression and classification tasks. Huiping Zhuang, Zhiping Lin 0001, Kar-Ann Toh |
IEEE Trans. Syst. Man Cybern. Syst. | 1 |
| 2021 | Accumulated Decoupled Learning with Gradient Staleness Mitigation for Convolutional Neural NetworksabstractGradient staleness is a major side effect in decoupled learning when training convolutional neural networks asynchronously. Existing methods that ignore this effect might result in reduced generalization and even divergence. In this paper, we propose an accumulated decoupled learning (ADL), which includes a module-wise gradient accumulation in order to mitigate the gradient staleness. Unlike prior arts ignoring the gradient staleness, we quantify the staleness in such a way that its mitigation can be quantitatively visualized. As a new learning scheme, the proposed ADL is theoretically shown to converge to critical points in spite of its asynchronism. Extensive experiments on CIFAR-10 and ImageNet datasets are conducted, demonstrating that ADL gives promising generalization results while the state-of-the-art methods experience reduced generalization and divergence. In addition, our ADL is shown to have the fastest training speed among the compared methods. Huiping Zhuang, Zhenyu Weng, Fulin Luo, Kar-Ann Toj, Haizhou Li 0001, Zhiping Lin 0001 |
ICML | 1 |
| 2021 | Training Multilayer Neural Networks Analytically Using Kernel ProjectionabstractThis paper proposes a kernel projection (KP) neural network that analytically determines its network parameters. The proposed network is composed of cascaded modules of 2-layer sub-networks. A technique which encodes the label information into each module has been introduced to enable a locally supervised learning. Such a supervised learning in the 2-layer module begins with a kernel projection in the first layer and determines its parameters analytically via solving a least squares problem in the second layer. We show that the analytic nature of the proposed network allows a learning process significantly faster than that of the traditional backpropagation method as it only needs to visit the dataset once. Experiments of classification tasks on various datasets are carried out, showing comparable or better results compared with several competing methods. Huiping Zhuang, Zhiping Lin 0001, Kar-Ann Toh |
ISCAS | 1 |
| 2021 | Correlation Projection for Analytic Learning of a Classification Network
Huiping Zhuang, Zhiping Lin 0001, Kar-Ann Toh |
Neural Process. Lett. | 1 |
| 2020 | Robust Real-time Face Tracking for People Wearing Face MasksabstractDue to the outbreak of the novel coronavirus (or known as COVID-19), people are advised to wear masks when they stay outdoors in many countries. This could result in difficulty for some public safety surveillance systems involving face detection or tracking. Therefore, the development of face detection and tracking algorithms for people wearing face masks is particularly important. In this paper, a real-time tracking algorithm for people with or without face masks is proposed. This algorithm is trained on public face datasets with faces without masks. Although the training does not involve face images of people wearing face masks, we show that the proposed algorithm is robust as it is able to perform well in face tracking for people wearing face masks. We also discuss the possible scenarios where the algorithm could lose track of the target when experimenting in tracking masked faces. This can motivate future research in this area. Xinggan Peng, Huiping Zhuang, Guang-Bin Huang, Haizhou Li 0001, Zhiping Lin 0001 |
ICARCV | 2 |
| 2019 | A Low-Memory Learning Formulation for a Kernel-and-Range NetworkabstractRecently, a learning method based on the kernel and the range space projections has been introduced. This method has been applied to learn the multilayer network analytically with interpretable relationships among the weight matrices. However, the learning method carries a high-memory demand during training. In this study, a low-memory formulation is proposed to address this issue of high-memory demand. The developed method is inspired by a recursive implementation of the Moore-Penrose inverse and is shown to be mathematically equivalent to the original batch learning. Next, we further improved our proposed low-memory formulation to annul the potential divergence caused by rounding errors. The regression and classification behaviors of the proposed learning method are demonstrated using both synthetic and benchmark datasets. Our experiments confirm that the proposed formulation consumes significantly lower memory. Huiping Zhuang, Zhiping Lin 0001, Kar-Ann Toh |
IJCNN | 1 |
| 2019 | Augmented EMD for complex-valued univariate signalsabstractIn this study, the authors propose an efficient extension of the standard empirical mode decomposition (EMD) for complex‐valued univariate signal decomposition. The key idea of the extension is to convert a complex‐valued univariate signal into a longer real‐valued signal by augmenting the real part with the flipped imaginary part, and then to decompose it into intrinsic mode functions (IMFs) using the EMD once only. The bivariate IMFs are then retrieved from the obtained IMFs. Their empirical results on synthetic data show that the proposed method significantly outperforms the traditional bivariate EMD (BEMD) method in terms of computational efficiency while producing a comparable extraction error. Moreover, the proposed method shows better micro‐Doppler signature analysis performance on physically measured continuous‐wave radar data than that of the BEMD. Beom-Seok Oh, Huiping Zhuang, Kar-Ann Toh, Zhiping Lin 0001 |
IET Signal Process. | 2 |
| 2016 | Robust two-stage Kalman filtering in presence of autoregressive inputabstractThe two-stage Kalman filter was proposed with the objective to avoid bias when the system is in presence of input signals. A common technical difficulty in this technique is that the dynamics of input signal is always unknown whereas the optimality of such filter can only be achieved with sufficient priori knowledge (i.e., known dynamics and statistics). Unbiased minimum-variance filter is capable of obtaining unbiased state estimates even in presence of an unknown input but the price is paid and it loses the capability to gain access to more accurate state estimates. This paper takes advantages of both estimators to propose a new estimator that combines their merits and discuss an estimation problem when the input signal displays autoregressive property. We manage to simultaneously estimate the input signal from unbiased minimum-variance filter ahead of the parameter identification procedure, which is of significance as the information is required to procure input dynamics. The un-biasedness of the input estimator is also proved in the paper. The identification step is then completed by converting it into solving an eigenvector problem. This proposed filter builds a bridge connecting unbiased minimum-variance filter and two-stage Kalman filter and the validity of the proposed method is justified via simulation results. Huiping Zhuang |
ICARCV | 1 |