Ziqian Zeng

dblp:155/0168 · DBLP profile ↗
← Back
34ranked-venue papers
8as first author
31since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 28 · 6 first-author · 26 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 3 first-author · 6 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Neural Architecture for Fast and Reliable Coagulation Assessment in Clinical Settings: Leveraging Thromboelastography
abstract
In an ideal medical environment, real-time coagulation monitoring can enable early detection and prompt remediation of risks. However, traditional Thromboelastography (TEG), a widely employed diagnostic modality, can only provide such outputs after nearly 1 hour of measurement. The delay might lead to elevated mortality rates. These issues clearly point out one of the key challenges for medical AI development: Making reasonable predictions based on very small data sets and accounting for variation between different patient populations, a task where conventional deep learning methods typically perform poorly. We present Physiological State Reconstruction (PSR), a new algorithm specifically designed to take advantage of dynamic changes between individuals and to maximize useful information produced by small amounts of clinical data through mapping to reliable predictions and diagnosis. We develop MDFE to facilitate integration of varied temporal signals using multi-domain learning, and jointly learn high-level temporal interactions together with attentions via HLA; furthermore, the parameterized DAM we designed maintains the stability of the computed vital signs. PSR evaluates with 4 TEG-specialized data sets and establishes remarkable performance -- predictions of R^2 > 0.98 for coagulation traits and error reduction around half compared to the state-of-the-art methods, and halving the inferencing time too. Drift-aware learning suggests a new future, with potential uses well beyond thrombophilia discovery towards medical AI applications with data scarcity.
Yulu Wang, Ziqian Zeng, Zhifeng Tang
AAAI2
2026 RCP-Merging: Merging Long Chain-of-Thought Models with Domain-Specific Models by Considering Reasoning Capability as Prior
abstract
Large Language Models (LLMs) with long chain-of-thought (CoT) capability, termed Reasoning Models, demonstrate superior intricate problem-solving abilities through multi-step long CoT reasoning. To create a dual-capability model with long CoT capability and domain-specific knowledge without substantial computational and data costs, model merging emerges as a highly resource-efficient method. However, significant challenges lie in merging domain-specific LLMs with long CoT ones since nowadays merging methods suffer from reasoning capability degradation, even gibberish output and output collapse. To overcome this, we introduce RCP-Merging: Merging Long Chain-of-Thought Models with Domain-Specific Models by Considering Reasoning Capability as Prior, a novel merging framework designed to integrate domain-specific LLMs with long CoT capability, meanwhile maintaining model performance in the original domain. Treating reasoning model weights as foundational prior, our method utilizes a reasoning capability indicator to preserve core long CoT capability model weights while selectively merging essential domain-specific weights. We conducted extensive experiments on Qwen2.5-7B, Llama3.1-8B, and Qwen2.5-1.5B models in BioMedicine and Finance domains. Our results show that RCP-Merging successfully merges a reasoning model with domain-specific ones, improving domain task performance by 9.5% and 9.2% over state-of-the-art methods, without significantly harming the original long CoT reasoning capability.
Junyao Yang, Huiping Zhuang, Cen Chen 0002, Ziqian Zeng
AAAI5
2026 Activation-Guided Local Editing for Jailbreaking Attacks
abstract
Jiecong Wang, Haoran Li, Hao Peng, Ziqian Zeng, Zihao Wang, Haohua Du, Zhengtao Yu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Jiecong Wang, Haoran Li 0003, Hao Peng 0001, Ziqian Zeng, Zihao Wang 0001, Haohua Du, Zhengtao Yu 0001
ACL (1)4
2026 MixKVQ: Query-Aware Mixed-Precision KV Cache Quantization for Long-Context Reasoning
abstract
Long Chain-of-Thought (CoT) reasoning has significantly advanced the capabilities of Large Language Models (LLMs), but this progress is accompanied by substantial memory and latency overhead from the extensive Key-Value (KV) cache.Although KV cache quantization is a promising compression technique, existing low-bit quantization methods often exhibit severe performance degradation on complex reasoning tasks.Fixed-precision quantization struggles to handle outlier channels in the key cache, while current mixed-precision strategies fail to accurately identify components requiring high-precision representation.We find that an effective low-bit KV cache quantization strategy must consider two factors: a key channel's intrinsic quantization difficulty and its relevance to the query.Based on this insight, we propose MixKVQ, a novel quantization method that introduces a lightweight, query-aware algorithm to identify and preserve critical key channels that need higher precision, while applying per-token quantization for value cache.Experiments on complex reasoning datasets demonstrate that our approach significantly outperforms existing low-bit methods, achieving performance comparable to a fullprecision baseline at a substantially reduced memory footprint.The source code is available at https://github.com/ZeroNLP/MixKVQ.
Tao Zhang 0019, Ziqian Zeng, Huiping Zhuang, Cen Chen 0002
ACL (1)2
2025 SDD: Self-Degraded Defense against Malicious Fine-tuning
abstract
Open-source Large Language Models (LLMs) often employ safety alignment methods to resist harmful instructions.However, recent research shows that maliciously fine-tuning these LLMs on harmful data can easily bypass these safeguards.To counter this, we theoretically uncover why malicious fine-tuning succeeds and identify potential defense strategies.Building on the theoretical analysis, we introduce the Self-Degraded Defense (SDD) framework.SDD encourages LLMs to produce high-quality but irrelevant responses to harmful prompts.When attackers attempt malicious fine-tuning, the general capability of the LLM aligned by SDD will significantly decrease, rendering it incapable of following harmful instructions.Our experimental results confirm SDD's effectiveness against such attacks.Our code is available at https://github.com/ZeroNLP/SDD.
Weikai Lu, Ziqian Zeng
ACL (1)4
2025 SEA: Low-Resource Safety Alignment for Multimodal Large Language Models via Synthetic Embeddings
abstract
Multimodal Large Language Models (MLLMs) have serious security vulnerabilities.While safety alignment using multimodal datasets consisting of text and data of additional modalities can effectively enhance MLLM's security, it is costly to construct these datasets.Existing low-resource security alignment methods, including textual alignment, have been found to struggle with the security risks posed by additional modalities.To address this, we propose Synthetic Embedding augmented safety Alignment (SEA), which optimizes embeddings of additional modality through gradient updates to expand textual datasets.This enables multimodal safety alignment training even when only textual data is available.Extensive experiments on image, video, and audio-based MLLMs demonstrate that SEA can synthesize a high-quality embedding on a single RTX3090 GPU within 24 seconds.SEA significantly improves the security of MLLMs when faced with threats from additional modalities.To assess the security risks introduced by video and audio, we also introduced a new benchmark called VA-SafetyBench.High attack success rates across multiple MLLMs validate its challenge.Our code and data will be available at https://github.com/ZeroNLP/SEA.This paper contains harmful data and modelgenerated content that can be offensive in nature.
Weikai Lu, Huiping Zhuang, Cen Chen 0002, Ziqian Zeng
ACL (1)5
2025 PrivacyRestore: Privacy-Preserving Inference in Large Language Models via Privacy Removal and Restoration
abstract
Ziqian Zeng, Jianwei Wang, Junyao Yang, Zhengdong Lu, Haoran Li, Huiping Zhuang, Cen Chen. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Ziqian Zeng, Junyao Yang, Zhengdong Lu, Huiping Zhuang, Cen Chen 0002
ACL (1)1
2025 GenderAlign: An Alignment Dataset for Mitigating Gender Bias in Large Language Models
abstract
Large Language Models (LLMs) are prone to generating content that exhibits gender biases, raising significant ethical concerns. Alignment, the process of fine-tuning LLMs to better align with desired behaviors, is recognized as an effective approach to mitigate gender biases. Although proprietary LLMs have made significant strides in mitigating gender bias, their alignment datasets are not publicly available. The commonly used and publicly available alignment dataset, HH-RLHF, still exhibits gender bias to some extent. There is a lack of publicly available alignment datasets specifically designed to address gender bias. Hence, we developed a new dataset named GenderAlign, aiming at mitigating a comprehensive set of gender biases in LLMs. This dataset comprises 8k single-turn dialogues, each paired with a “chosen” and a “rejected” response. Compared to the “rejected” responses, the “chosen” responses demonstrate lower levels of gender bias and higher quality. Furthermore, we categorized the gender biases in the “rejected” responses of GenderAlign into 4 principal categories. The experimental results show the effectiveness of GenderAlign in reducing gender bias in LLMs.
Tao Zhang 0019, Ziqian Zeng, YuxiangXiao YuxiangXiao, Huiping Zhuang, Cen Chen 0002, James R. Foulds, Shimei Pan
ACL (1)2
2025 AFL: A Single-Round Analytic Approach for Federated Learning with Pre-trained Models
abstract
In this paper, we introduce analytic federated learning (AFL), a new training paradigm that brings analytical (i.e., closed- form) solutions to the federated learning (FL) with pre-trained models. Our AFL draws inspiration from analytic learning—a gradient-free technique that trains neural networks with analytical solutions in one epoch. In the local client training stage, the AFL facilitates a one-epoch training, eliminating the necessity for multi-epoch updates. In the aggregation stage, we derive an absolute aggregation (AA) law. This AA law allows a single-round aggregation, reducing heavy communication overhead and achieving fast convergence by removing the need for multiple aggregation rounds. More importantly, the AFL exhibits a property that invariance to data partitioning, meaning that regardless of how the full dataset is distributed among clients, the aggregated result remains identical. This could spawn various potentials, such as data heterogeneity invariance and client-number invariance. We conduct experiments across various FL settings including extremely non-IID ones, and scenarios with a large number of clients (e.g., ≥ 1000). In all these settings, our AFL constantly performs competitively while existing FL techniques encounter various obstacles. Our codes are available at https://github.com/ZHUANGHP/Analytic-federated-learning.
Run He, Kai Tong, Di Fang 0004, Ziqian Zeng, Huiping Zhuang
CVPR5
2025 Context Reasoner: Incentivizing Reasoning Capability for Contextualized Privacy and Safety Compliance via Reinforcement Learning
abstract
Wenbin Hu, Haoran Li, Huihao Jing, Qi Hu, Ziqian Zeng, Sirui Han, Xu Heli, Tianshu Chu, Peizhao Hu, Yangqiu Song. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Wenbin Hu 0001, Haoran Li 0003, Huihao Jing, Ziqian Zeng, Sirui Han, Heli Xu, Peizhao Hu, Yangqiu Song
EMNLP5
2025 RewardDS: Privacy-Preserving Fine-Tuning for Large Language Models via Reward Driven Data Synthesis
abstract
The success of large language models (LLMs) has attracted many individuals to fine-tune them for domain-specific tasks by uploading their data.However, in sensitive areas like healthcare and finance, privacy concerns often arise.One promising solution is to generate synthetic data with Differential Privacy (DP) guarantees to replace private data.However, these synthetic data contain significant flawed data, which are considered as noise.Existing solutions typically rely on naive filtering by comparing ROUGE-L scores or embedding similarities, which are ineffective in addressing the noise.To address this issue, we propose RewardDS, a novel privacy-preserving framework that fine-tunes a reward proxy model and uses reward signals to guide the synthetic data generation.Our RewardDS introduces two key modules, Reward Guided Filtering and Self-Optimizing Refinement, to both filter and refine the synthetic data, effectively mitigating the noise.Extensive experiments across medical, financial, and code generation domains demonstrate the effectiveness of our method.
Chengming Shi, Junyao Yang, Huiping Zhuang, Cen Chen 0002, Ziqian Zeng
EMNLP8
2025 DrFrattn: Directly Learn Adaptive Policy from Attention for Simultaneous Machine Translation
abstract
Simultaneous machine translation (SiMT) necessitates a robust read/write (R/W) policy to determine the optimal moments for translation, thereby balancing translation quality and latency.Effective timing in translation can align source and target tokens accurately.The attention mechanism within translation models inherently provides valuable alignment information.Building on this, previous research has attempted to modify the attention mechanism's structure to leverage its alignment properties during training, employing multi-task learning to derive the read/write policy.However, this multi-task learning approach may compromise the efficacy of the attention mechanism itself.This raises a natural question: why not directly learn the read/write policy from the well-trained attention mechanism?In this study, we propose DrFrattn, a method that directly learns adaptive policies from the attention mechanism.Experimental results across various benchmarks demonstrate that our approach achieves an improved balance between translation accuracy and latency.
Ziqian Zeng
EMNLP3
2025 Semantic Shift Estimation via Dual-Projection and Classifier Reconstruction for Exemplar-Free Class-Incremental Learning
abstract
Exemplar-Free Class-Incremental Learning (EFCIL) aims to sequentially learn from distinct categories without retaining exemplars but easily suffers from catastrophic forgetting of learned knowledge. While existing EFCIL methods leverage knowledge distillation to alleviate forgetting, they still face two critical challenges: semantic shift and decision bias. Specifically, the embeddings of old tasks shift in the embedding space after learning new tasks, and the classifier becomes biased towards new tasks due to training solely with new data, hindering the balance between old and new knowledge. To address these issues, we propose the Dual-Projection Shift Estimation and Classifier Reconstruction (DPCR) approach for EFCIL. DPCR effectively estimates semantic shift through a dual-projection, which combines a learnable transformation with a row-space projection to capture both task-wise and category-wise shifts. Furthermore, to mitigate decision bias, DPCR employs ridge regression to reformulate a classifier reconstruction process. This reconstruction exploits previous in covariance and prototype of each class after calibration with estimated shift, thereby reducing decision bias. Extensive experiments demonstrate that, on various datasets, DPCR effectively balances old and new tasks, outperforming state-of-the-art EFCIL methods. Our codes are available at https://github.com/RHe502/ICML25-DPCR.
Run He, Di Fang 0004, Yawen Cui, Ming Li 0011, Cen Chen 0002, Ziqian Zeng, Huiping Zhuang
ICML7
2025 L3A: Label-Augmented Analytic Adaptation for Multi-Label Class Incremental Learning
abstract
Class-incremental learning (CIL) enables models to learn new classes continually without forgetting previously acquired knowledge. Multi-label CIL (MLCIL) extends CIL to a real-world scenario where each sample may belong to multiple classes, introducing several challenges: label absence, which leads to incomplete historical information due to missing labels, and class imbalance, which results in the model bias toward majority classes. To address these challenges, we propose Label-Augmented Analytic Adaptation (L3A), an exemplar-free approach without storing past samples. L3A integrates two key modules. The pseudo-label (PL) module implements label augmentation by generating pseudo-labels for current phase samples, addressing the label absence problem. The weighted analytic classifier (WAC) derives a closed-form solution for neural networks. It introduces sample-specific weights to adaptively balance the class contribution and mitigate class imbalance. Experiments on MS-COCO and PASCAL VOC datasets demonstrate that L3A outperforms existing methods in MLCIL tasks. Our code is available at https://github.com/scut-zx/L3A.
Run He, Chen Jiao, Di Fang 0004, Ming Li 0073, Ziqian Zeng, Cen Chen 0002, Huiping Zhuang
ICML6
2025 DRViT: A dynamic redundancy-aware vision transformer accelerator via algorithm and architecture co-design on FPGA
Xiangfeng Sun, Yuan-Ting Zhang, Xiaofeng Zou, Ziqian Zeng, Huiping Zhuang
J. Parallel Distributed Comput.6
2025 Subkv: Quantizing Long Context KV Cache for Sub-Billion Parameter Language Models on Edge Devices
abstract
ABSTRACT Background Large Language Models (LLMs) have demonstrated remarkable capabilities across various tasks. However, their substantial computational and memory requirements present significant challenges for widespread deployment on edge devices. Motivation In long‐context scenarios, even sub‐billion parameter LLMs face unavoidable memory and performance bottlenecks due to inefficient KV Cache utilization. Existing quantization methods fail to address these challenges effectively. Method This paper addresses these challenges by introducing advanced quantization techniques tailored for sub‐billion parameter LLMs. It specifically targets reducing memory consumption through the conversion of the model's KV Cache to lower‐bit integers. We present SubKV, a quantization method specifically designed to optimize the KV Cache in sub‐billion parameter LLMs. Our analysis reveals distinct distributional differences in the magnitude of key and value caches. Leveraging this insight, we apply Per‐Channel Quantization to the key cache and Per‐Token Quantization to the value cache. Furthermore, we introduce the Dynamic Window Quantization method to enhance attention computations. To mitigate the extreme sensitivity of the first token, we also introduce Attention Sink‐Aware Quantization. Results Experimental results demonstrate that SubKV significantly reduces the KV Cache size during long context inference while maintaining model performance, offering superior results to existing KV Cache quantization methods.
Ziqian Zeng, Tao Zhang 0019, Zhengdong Lu, Huiping Zhuang, Hongen Shao, Sin G. Teo, Xiaofeng Zou
Softw. Pract. Exp.1
2024 ConsistentEE: A Consistent and Hardness-Guided Early Exiting Method for Accelerating Language Models Inference
abstract
Early Exiting is one of the most popular methods to achieve efficient inference. Current early exiting methods adopt the (weighted) sum of the cross entropy loss of all internal classifiers as the objective function during training, imposing all these classifiers to predict all instances correctly. However, during inference, as long as one internal classifier predicts an instance correctly, it can accelerate without losing accuracy. Thus, there is a notable gap between training and inference. We propose ConsistentEE, an early exiting method that is consistent in training and inference. ConsistentEE formulates the early exiting process as a reinforcement learning problem. A policy network is added to decide whether an instance should exit or continue. The training objective of ConsistentEE only requires each instance to be predicted correctly by one internal classifier. Additionally, we introduce the concept "Memorized Layer" to measure the hardness of an instance. We incorporate the memorized layer into reward function design, which allows "easy'' instances to focus more on acceleration while ``hard'' instances to focus more on accuracy. Experimental results show that our method outperforms other baselines on various natural language understanding and generation tasks using PLMs and LLMs as backbones respectively.
Ziqian Zeng, Yihuai Hong, Huiping Zhuang, Cen Chen 0002
AAAI1
2024 DS-AL: A Dual-Stream Analytic Learning for Exemplar-Free Class-Incremental Learning
abstract
Class-incremental learning (CIL) under an exemplar-free constraint has presented a significant challenge. Existing methods adhering to this constraint are prone to catastrophic forgetting, far more so than replay-based techniques that retain access to past samples. In this paper, to solve the exemplar-free CIL problem, we propose a Dual-Stream Analytic Learning (DS-AL) approach. The DS-AL contains a main stream offering an analytical (i.e., closed-form) linear solution, and a compensation stream improving the inherent under-fitting limitation due to adopting linear mapping. The main stream redefines the CIL problem into a Concatenated Recursive Least Squares (C-RLS) task, allowing an equivalence between the CIL and its joint-learning counterpart. The compensation stream is governed by a Dual-Activation Compensation (DAC) module. This module re-activates the embedding with a different activation function from the main stream one, and seeks fitting compensation by projecting the embedding to the null space of the main stream's linear mapping. Empirical results demonstrate that the DS-AL, despite being an exemplar-free technique, delivers performance comparable with or better than that of replay-based methods across various datasets, including CIFAR-100, ImageNet-100 and ImageNet-Full. Additionally, the C-RLS' equivalent property allows the DS-AL to execute CIL in a phase-invariant manner. This is evidenced by a never-before-seen 500-phase CIL ImageNet task, which performs on a level identical to a 5-phase one. Our codes are available at https://github.com/ZHUANGHP/Analytic-continual-learning.
Huiping Zhuang, Run He, Kai Tong, Ziqian Zeng, Cen Chen 0002, Zhiping Lin 0001
AAAI4
2024 On the Use of Silver Standard Data for Zero-shot Classification Tasks in Information Extraction
abstract
The superior performance of supervised classification methods in the information extraction (IE) area heavily relies on a large amount of gold standard data. Recent zero-shot classification methods converted the task to other NLP tasks (e.g., textual entailment) and used off-the-shelf models of these NLP tasks to directly perform inference on the test data without using a large amount of IE annotation data. A potentially valuable by-product of these methods is the large-scale silver standard data, i.e., pseudo-labeled data by the off-the-shelf models of other NLP tasks. However, there is no further investigation into the use of these data. In this paper, we propose a new framework, Clean-LaVe, which aims to utilize silver standard data to enhance the zero-shot performance. Clean-LaVe includes four phases: (1) Obtaining silver data; (2) Identifying relatively clean data from silver data; (3) Finetuning the off-the-shelf model using clean data; (4) Inference on the test data. The experimental results show that Clean-LaVe can outperform the baseline by 5% and 6% on TACRED and Wiki80 dataset in the zero-shot relation classification task, and by 3% ~7 % on Smile (Korean and Polish) in the zero-shot cross-lingual relation classification task, and by 8% on ACE05-E+ in the zero-shot event argument classification task.
Ziqian Zeng
LREC/COLING3
2024 Zero-shot Event Detection Using a Textual Entailment Model as an Enhanced Annotator
abstract
Zero-shot event detection is a challenging task. Recent research work proposed to use a pre-trained textual entailment (TE) model on this task. However, those methods treated the TE model as a frozen annotator. We treat the TE model as an annotator that can be enhanced. We propose to use TE models to annotate large-scale unlabeled text and use annotated data to finetune the TE model, yielding an improved TE model. Finally, the improved TE model is used for inference on the test set. To improve the efficiency, we propose to use keywords to filter out sentences with a low probability of expressing event(s). To improve the coverage of keywords, we expand limited number of seed keywords using WordNet, so that we can use the TE model to annotate unlabeled text efficiently. The experimental results show that our method can outperform other baselines by 15% on the ACE05 dataset.
Ziqian Zeng, Runyu Wu, Yuxiang Xiao, Xiaoda Zhong, Zhengdong Lu, Huiping Zhuang
LREC/COLING1
2024 Dissecting Fine-Tuning Unlearning in Large Language Models
abstract
Fine-tuning-based unlearning methods prevail for preventing targeted harmful, sensitive, or copyrighted information within large language models while preserving overall capabilities.However, the true effectiveness of these methods is unclear.In this work, we delve into the limitations of fine-tuning-based unlearning through activation patching and parameter restoration experiments.Our findings reveal that these methods alter the model's knowledge retrieval process, providing further evidence that they do not genuinely erase the problematic knowledge embedded in the model parameters.Instead, the coefficients generated by the MLP components in the model's final layer are the primary contributors to these seemingly positive unlearning effects, playing a crucial role in controlling the model's behaviors.Furthermore, behavioral tests demonstrate that this unlearning mechanism inevitably impacts the global behavior of the models, affecting unrelated knowledge or capabilities.The code is released at https://github.com/yihuaihong/ Dissecting-FT-Unlearning.
Yihuai Hong, Yuelin Zou, Lijie Hu, Ziqian Zeng, Di Wang 0015, Haiqin Yang
EMNLP4
2024 PsFuture: A Pseudo-Future-based Zero-Shot Adaptive Policy for Simultaneous Machine Translation
abstract
Simultaneous Machine Translation (SiMT) requires target tokens to be generated in realtime as streaming source tokens are consumed.Traditional approaches to SiMT typically require sophisticated architectures and extensive parameter configurations for training adaptive read/write policies, which in turn demand considerable computational power and memory.We propose PsFuture, the first zero-shot adaptive read/write policy for SiMT, enabling the translation model to independently determine read/write actions without the necessity for additional training.Furthermore, we introduce a novel training strategy, Prefix-to-Full (P2F), specifically tailored to adjust offline translation models for SiMT applications, exploiting the advantages of the bidirectional attention mechanism inherent in offline models.Experiments across multiple benchmarks demonstrate that our zero-shot policy attains performance on par with strong baselines and the P2F method can further enhance performance, achieving an outstanding trade-off between translation quality and latency. 1
Ziqian Zeng
EMNLP3
2024 GACL: Exemplar-Free Generalized Analytic Continual Learning
abstract
Class incremental learning (CIL) trains a network on sequential tasks with separated categories in each task but suffers from catastrophic forgetting, where models quickly lose previously learned knowledge when acquiring new tasks. The generalized CIL (GCIL) aims to address the CIL problem in a more real-world scenario, where incoming data have mixed data categories and unknown sample size distribution. Existing attempts for the GCIL either have poor performance or invade data privacy by saving exemplars. In this paper, we propose a new exemplar-free GCIL technique named generalized analytic continual learning (GACL). The GACL adopts analytic learning (a gradient-free training technique) and delivers an analytical (i.e., closed-form) solution to the GCIL scenario. This solution is derived via decomposing the incoming data into exposed and unexposed classes, thereby attaining a weight-invariant property, a rare yet valuable property supporting an equivalence between incremental learning and its joint training. Such an equivalence is crucial in GCIL settings as data distributions among different tasks no longer pose challenges to adopting our GACL. Theoretically, this equivalence property is validated through matrix analysis tools. Empirically, we conduct extensive experiments where, compared with existing GCIL methods, our GACL exhibits a consistently leading performance across various datasets and GCIL settings. Source code is available at https://github.com/CHEN-YIZHU/GACL.
Huiping Zhuang, Yizhu Chen, Di Fang 0004, Run He, Kai Tong, Hongxin Wei, Ziqian Zeng, Cen Chen 0002
NeurIPS7
2024 F-OAL: Forward-only Online Analytic Learning with Fast Training and Low Memory Footprint in Class Incremental Learning
abstract
Online Class Incremental Learning (OCIL) aims to train models incrementally, where data arrive in mini-batches, and previous data are not accessible. A major challenge in OCIL is Catastrophic Forgetting, i.e., the loss of previously learned knowledge. Among existing baselines, replay-based methods show competitive results but requires extra memory for storing exemplars, while exemplar-free (i.e., data need not be stored for replay in production) methods are resource friendly but often lack accuracy. In this paper, we propose an exemplar-free approach—Forward-only Online Analytic Learning (F-OAL). Unlike traditional methods, F-OAL does not rely on back-propagation and is forward-only, significantly reducing memory usage and computational time. Cooperating with a pre-trained frozen encoder with Feature Fusion, F-OAL only needs to update a linear classifier by recursive least square. This approach simultaneously achieves high accuracy and low resource consumption. Extensive experiments on bench mark datasets demonstrate F-OAL’s robust performance in OCIL scenarios. Code is available at: https://github.com/liuyuchen-cz/F-OAL
Huiping Zhuang, Yuchen Liu 0001, Run He, Kai Tong, Ziqian Zeng, Cen Chen 0002, Yi Wang 0068, Lap-Pui Chau
NeurIPS5
2024 Online network traffic classification based on external attention and convolution by IP packet header
Yahui Hu, Ziqian Zeng, Junping Song, Luyang Xu
Comput. Networks2
2023 From Ultra-Fine to Fine: Fine-tuning Ultra-Fine Entity Typing Models to Fine-grained
abstract
For the task of fine-grained entity typing (FET), due to the use of a large number of entity types, it is usually considered too costly to manually annotate a training dataset that contains an ample number of examples for each type.A common way to address this problem is to use distantly annotated training examples that contains incorrect labels.But the errors in the automatic annotation may limit the performance of trained models.Recently, there are a few approaches that no longer depend on such weak training data.However, without using sufficient direct entity typing supervision may also cause them to yield inferior performance.In this paper, we propose a new approach that can avoid the need of creating distantly labeled data.We first train an entity typing model that have an extremely broad type coverage by using the ultrafine entity typing data.Then, when there is a need to produce a model for a newly designed fine-grained entity type schema, we can simply fine-tune the previously trained model with a small number of corresponding annotated examples.Experimental results show that our approach achieves outstanding performance for FET under the few-shot setting.It can also outperform state-of-the-art weak supervision based methods after fine-tuning the model with only a small-size manually annotated training set.[ FedEx ] ( Type : [MASK] ) is a major player in the package delivery market .
Ziqian Zeng
ACL (1)2
2023 GKEAL: Gaussian Kernel Embedded Analytic Learning for Few-Shot Class Incremental Task
abstract
Few-shot class incremental learning (FSCIL) aims to address catastrophic forgetting during class incremental learning in a few-shot learning setting. In this paper, we approach the FSCIL by adopting analytic learning, a technique that converts network training into linear problems. This is inspired by the fact that the recursive implementation (batch-by-batch learning) of analytic learning gives identical weights to that produced by training on the entire dataset at once. The recursive implementation and the weight-identical property highly resemble the FSCIL setting (phase-by-phase learning) and its goal of avoiding catastrophic forgetting. By bridging the FSCIL with the analytic learning, we propose a Gaussian kernel embedded analytic learning (GKEAL) for FSCIL. The key components of GKEAL include the kernel analytic module which allows the GKEAL to conduct FSCIL in a recursive manner, and the augmented feature concatenation module that balances the preference between old and new tasks especially effectively under the few-shot setting. Our experiments show that the GKEAL gives state-of-the-art performance on several benchmark datasets.
Huiping Zhuang, Zhenyu Weng, Run He, Zhiping Lin 0001, Ziqian Zeng
CVPR5
2023 Adaptive Policy with Wait-k Model for Simultaneous Translation
abstract
Simultaneous machine translation (SiMT) requires a robust read/write (R/W) policy in conjunction with a high-quality translation model.Traditional methods rely on either a fixed waitk policy coupled with a standalone wait-k translation model, or an adaptive policy jointly trained with the translation model.In this study, we propose a more flexible approach by decoupling the adaptive policy model from the translation model.Our motivation stems from the observation that a standalone multi-path waitk model performs competitively with adaptive policies utilized in state-of-the-art SiMT approaches.Specifically, we introduce DaP, a divergence-based adaptive policy, that makes read/write decisions for any translation model based on the potential divergence in translation distributions resulting from future information.DaP extends a frozen wait-k model with lightweight parameters, and is both memory and computation efficient.Experimental results across various benchmarks demonstrate that our approach offers an improved trade-off between translation accuracy and latency, outperforming strong baselines.1
Kai Fan 0002, Shushu Wang, Ziqian Zeng, Zhongqiang Huang
EMNLP6
2021 Variational Weakly Supervised Sentiment Analysis with Posterior Regularization
abstract
Sentiment analysis is an important task in natural language processing (NLP). Most of existing state-of-the-art methods are under the supervised learning paradigm. However, human annotations can be scarce. Thus, we should leverage more weak supervision for sentiment analysis. In this paper, we propose a posterior regularization framework for the variational approach to the weakly supervised sentiment analysis to better control the posterior distribution of the label assignment. The intuition behind the posterior regularization is that if extracted opinion words from two documents are semantically similar, the posterior distributions of two documents should be similar. Our experimental results show that the posterior regularization can improve the original variational approach to the weakly supervised sentiment analysis and the performance is more stable with smaller prediction variance.
Ziqian Zeng, Yangqiu Song
EACL1
2021 Fair Representation Learning for Heterogeneous Information Networks
Ziqian Zeng, Rashidul Islam, Kamrun Keya, James R. Foulds, Yangqiu Song, Shimei Pan
ICWSM1
2021 Debiasing Career Recommendations with Neural Fair Collaborative Filtering
abstract
A growing proportion of human interactions are digitized on social media platforms and subjected to algorithmic decision-making, and it has become increasingly important to ensure fair treatment from these algorithms. In this work, we investigate gender bias in collaborative-filtering recommender systems trained on social media data. We develop neural fair collaborative filtering (NFCF), a practical framework for mitigating gender bias in recommending career-related sensitive items (e.g. jobs, academic concentrations, or courses of study) using a pre-training and fine-tuning approach to neural collaborative filtering, augmented with bias correction techniques. We show the utility of our methods for gender de-biased career and college major recommendations on the MovieLens dataset and a Facebook dataset, respectively, and achieve better performance and fairer behavior than several state-of-the-art models.
Rashidul Islam, Kamrun Keya, Ziqian Zeng, Shimei Pan, James R. Foulds
WWW3
2018 Biased Random Walk based Social Regularization for Word Embeddings
abstract
Nowadays, people publish a lot of natural language texts on social media. Socialized word embeddings (SWE) has been proposed to deal with two phenomena of language use: everyone has his/her own personal characteristics of language use and socially connected users are likely to use language in similar ways. We observe that the spread of language use is transitive. Namely, one user can affect his/her friends and the friends can also affect their friends. However, SWE modeled the transitivity implicitly. The social regularization in SWE only applies to one-hop neighbors and thus users outside the one-hop social circle will not be affected directly. In this work, we adopt random walk methods to generate paths on the social graph to model the transitivity explicitly. Each user on a path will be affected by his/her adjacent user(s) on the path. Moreover, according to the update mechanism of SWE, fewer friends a user has, fewer update opportunities he/she can get. Hence, we propose a biased random walk method to provide these users with more update opportunities. Experiments show that our random walk based social regularizations perform better on sentiment classification.
Ziqian Zeng, Xin Liu 0039, Yangqiu Song
IJCAI1
2017 Socialized Word Embeddings
abstract
Word embeddings have attracted a lot of attention. On social media, each user’s language use can be significantly affected by the user’s friends. In this paper, we propose a socialized word embedding algorithm which can consider both user’s personal characteristics of language use and the user’s social relationship on social media. To incorporate personal characteristics, we propose to use a user vector to represent each user. Then for each user, the word embeddings are trained based on each user’s corpus by combining the global word vectors and local user vector. To incorporate social relationship, we add a regularization term to impose similarity between two friends. In this way, we can train the global word vectors and user vectors jointly. To demonstrate the effectiveness, we used the latest large-scale Yelp data to train our vectors, and designed several experiments to show how user vectors affect the results.
Ziqian Zeng, Yichun Yin, Yangqiu Song, Ming Zhang 0004
IJCAI1
2015 Asymmetric Cyclical Hashing for Large Scale Image Retrieval
abstract
This paper addresses a problem in the hashing technique for large scale image retrieval: learn a compact hash code to reduce the storage cost with performance comparable to that of the long hash code. A longer hash code yields a better precision rate of retrieved images. However, it also requires a larger storage, which limits the number of stored images. Current hashing methods employ the same code length for both queries and stored images. We propose a new hashing scheme using two hash codes with different lengths for queries and stored images, i.e., the asymmetric cyclical hashing. A compact hash code is used to reduce the storage requirement, while a long hash code is used for the query image. The image retrieval is performed by computing the Hamming distance of the long hash code of the query and the cyclically concatenated compact hash code of the stored image to yield a high precision and recall rate. Experiments on benchmarking databases consisting up to one million images show the effectiveness of the proposed method.
Yueming Lv, Wing W. Y. Ng, Ziqian Zeng, Daniel S. Yeung, Patrick P. K. Chan
IEEE Trans. Multim.3