Zijian Feng

dblp:45/10114 · DBLP profile ↗
← Back
17ranked-venue papers
6as first author
17since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 15 · 5 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Beyond the Next Token: Towards Prompt-Robust Zero-Shot Classification via Efficient Multi-Token Prediction
abstract
Junlang Qian, Zixiao Zhu, Hanzhang Zhou, Zijian Feng, Zepeng Zhai, Kezhi Mao. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Junlang Qian, Zixiao Zhu, Hanzhang Zhou, Zijian Feng, Zepeng Zhai, Kezhi Mao
NAACL (Long Papers)4
2025 Logit Separability-Driven Samples and Multiple Class-Related Words Selection for Advancing In-Context Learning
abstract
Zixiao Zhu, Zijian Feng, Hanzhang Zhou, Junlang Qian, Kezhi Mao. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Zixiao Zhu, Zijian Feng, Hanzhang Zhou, Junlang Qian, Kezhi Mao
NAACL (Long Papers)2
2025 Restoring Pruned Large Language Models via Lost Component Compensation
abstract
Pruning is a widely used technique to reduce the size and inference cost of large language models (LLMs), but it often causes performance degradation. To mitigate this, existing restoration methods typically employ parameter-efficient fine-tuning (PEFT), such as LoRA, to recover the pruned model's performance. However, most PEFT methods are designed for dense models and overlook the distinct properties of pruned models, often resulting in suboptimal recovery. In this work, we propose a targeted restoration strategy for pruned models that restores performance while preserving their low cost and high efficiency. We observe that pruning-induced information loss is reflected in attention activations, and selectively reintroducing components of this information can significantly recover model performance. Based on this insight, we introduce RestoreLCC (Restoring Pruned LLMs via Lost Component Compensation), a plug-and-play method that contrastively probes critical attention heads via activation editing, extracts lost components from activation differences, and finally injects them back into the corresponding pruned heads for compensation and recovery. RestoreLCC is compatible with structured, semi-structured, and unstructured pruning schemes. Extensive experiments demonstrate that RestoreLCC consistently outperforms state-of-the-art baselines in both general and task-specific performance recovery, without compromising the sparsity or inference efficiency of pruned models.
Zijian Feng, Hanzhang Zhou, Zixiao Zhu, Chua Jia Jim Deryl, Lee Onn Mak, Gee Wah Ng, Kezhi Mao
NeurIPS1
2025 Using homologous network to identify reassortment risk in H5Nx avian influenza viruses
abstract
The resurgence of H5Nx reassortment has caused multiple epidemics resulting in severe disease even death in wild birds and poultry. Assessing H5Nx reassortment risk is crucial for designing targeted interventions and enhancing preparedness efforts to manage H5Nx outbreaks effectively. However, the complexity in H5Nx reassortment, driven by the diversity of influenza A viruses (IAVs) and wide range of hosts, has hindered the effective quantification of reassortment risk. In this study, we utilized a network approach to explore the reassortment history using a large-scale dataset. By inferring genomic homogeneity among IAVs, we constructed an IAVs homologous network with reassortment history embedded within it. We estimated the communities within the IAVs homologous network to represent the reassortment risk of various viruses, revealing diverse reassortment risks across different H5Nx viruses. Our analysis also identified the primary hosts contributing to reassortment: domestic poultry in China, and wild birds in North America and Europe. These primary hosts are critical targets for future H5Nx reassortment interventions. Our study provides a framework for quantifying and ranking H5Nx reassortment risk, contributing to enhanced preparedness and prevention efforts.
Ruihao Gong, Zijian Feng
PLoS Comput. Biol.2
2024 FreeCtrl: Constructing Control Centers with Feedforward Layers for Learning-Free Controllable Text Generation
abstract
Controllable text generation (CTG) seeks to craft texts adhering to specific attributes, traditionally employing learning-based techniques such as training, fine-tuning, or prefix-tuning with attribute-specific datasets.These approaches, while effective, demand extensive computational and data resources.In contrast, some proposed learning-free alternatives circumvent learning but often yield inferior results, exemplifying the fundamental machine learning trade-off between computational expense and model efficacy.To overcome these limitations, we propose FreeCtrl, a learningfree approach that dynamically adjusts the weights of selected feedforward neural network (FFN) vectors to steer the outputs of large language models (LLMs).FreeCtrl hinges on the principle that the weights of different FFN vectors influence the likelihood of different tokens appearing in the output.By identifying and adaptively adjusting the weights of attributerelated FFN vectors, FreeCtrl can control the output likelihood of attribute keywords in the generated content.Extensive experiments on single-and multi-attribute control reveal that the learning-free FreeCtrl outperforms other learning-free and learning-based methods, successfully resolving the dilemma between learning costs and model performance 1 .
Zijian Feng, Hanzhang Zhou, Kezhi Mao, Zixiao Zhu
ACL (1)1
2024 LLMs Learn Task Heuristics from Demonstrations: A Heuristic-Driven Prompting Strategy for Document-Level Event Argument Extraction
abstract
In this study, we explore in-context learning (ICL) in document-level event argument extraction (EAE) to alleviate the dependency on large-scale labeled data for this task.We introduce the Heuristic-Driven Link-of-Analogy (HD-LoA) prompting tailored for the EAE task.Specifically, we hypothesize and validate that LLMs learn task-specific heuristics from demonstrations in ICL.Building upon this hypothesis, we introduce an explicit heuristicdriven demonstration construction approach, which transforms the haphazard example selection process into a systematic method that emphasizes task heuristics.Additionally, inspired by the analogical reasoning of human, we propose the link-of-analogy prompting, which enables LLMs to process new situations by drawing analogies to known situations, enhancing their performance on unseen classes beyond limited ICL examples.Experiments show that our method outperforms existing prompting methods and few-shot supervised learning methods on document-level EAE datasets.Additionally, the HD-LoA prompting shows effectiveness in other tasks like sentiment analysis and natural language inference, demonstrating its broad adaptability 1 .
Hanzhang Zhou, Junlang Qian, Zijian Feng, Zixiao Zhu, Kezhi Mao
ACL (1)3
2024 Unveiling and Manipulating Prompt Influence in Large Language Models
abstract
Prompts play a crucial role in guiding the responses of Large Language Models (LLMs). However, the intricate role of individual tokens in prompts, known as input saliency, in shaping the responses remains largely underexplored. Existing saliency methods either misalign with LLM generation objectives or rely heavily on linearity assumptions, leading to potential inaccuracies. To address this, we propose Token Distribution Dynamics (TDD), an elegantly simple yet remarkably effective approach to unveil and manipulate the role of prompts in generating LLM outputs. TDD leverages the robust interpreting capabilities of the language model head (LM head) to assess input saliency. It projects input tokens into the embedding space and then estimates their significance based on distribution dynamics over the vocabulary. We introduce three TDD variants: forward, backward, and bidirectional, each offering unique insights into token relevance. Extensive experiments reveal that the TDD surpasses state-of-the-art baselines with a big margin in elucidating the causal relationships between prompts and LLM outputs. Beyond mere interpretation, we apply TDD to two prompt manipulation tasks for controlled text generation: zero-shot toxic language suppression and sentiment steering. Empirical results underscore TDD's proficiency in identifying both toxic and sentimental cues in prompts, subsequently mitigating toxicity or modulating sentiment in the generated content.
Zijian Feng, Hanzhang Zhou, Zixiao Zhu, Junlang Qian, Kezhi Mao
ICLR1
2024 UniBias: Unveiling and Mitigating LLM Bias through Internal Attention and FFN Manipulation
abstract
Large language models (LLMs) have demonstrated impressive capabilities in various tasks using the in-context learning (ICL) paradigm. However, their effectiveness is often compromised by inherent bias, leading to prompt brittleness—sensitivity to design settings such as example selection, order, and prompt formatting. Previous studies have addressed LLM bias through external adjustment of model outputs, but the internal mechanisms that lead to such bias remain unexplored. Our work delves into these mechanisms, particularly investigating how feedforward neural networks (FFNs) and attention heads result in the bias of LLMs. By Interpreting the contribution of individual FFN vectors and attention heads, we identify the biased LLM components that skew LLMs' prediction toward specific labels. To mitigate these biases, we introduce UniBias, an inference-only method that effectively identifies and eliminates biased FFN vectors and attention heads. Extensive experiments across 12 NLP datasets demonstrate that UniBias significantly enhances ICL performance and alleviates prompt brittleness of LLMs.
Hanzhang Zhou, Zijian Feng, Zixiao Zhu, Junlang Qian, Kezhi Mao
NeurIPS2
2024 DGCPPISP: a PPI site prediction model based on dynamic graph convolutional network and two-stage transfer learning
abstract
BACKGROUND: Proteins play a pivotal role in the diverse array of biological processes, making the precise prediction of protein-protein interaction (PPI) sites critical to numerous disciplines including biology, medicine and pharmacy. While deep learning methods have progressively been implemented for the prediction of PPI sites within proteins, the task of enhancing their predictive performance remains an arduous challenge. RESULTS: In this paper, we propose a novel PPI site prediction model (DGCPPISP) based on a dynamic graph convolutional neural network and a two-stage transfer learning strategy. Initially, we implement the transfer learning from dual perspectives, namely feature input and model training that serve to supply efficacious prior knowledge for our model. Subsequently, we construct a network designed for the second stage of training, which is built on the foundation of dynamic graph convolution. CONCLUSIONS: To evaluate its effectiveness, the performance of the DGCPPISP model is scrutinized using two benchmark datasets. The ensuing results demonstrate that DGCPPISP outshines competing methods in terms of performance. Specifically, DGCPPISP surpasses the second-best method, EGRET, by margins of 5.9%, 10.1%, and 13.3% for F1-measure, AUPRC, and MCC metrics respectively on Dset_186_72_PDB164. Similarly, on Dset_331, it eclipses the performance of the runner-up method, HN-PPISP, by 14.5%, 19.8%, and 29.9% respectively.
Zijian Feng, Haohao Li, Hancan Zhu, Yanlei Kang
BMC Bioinform.1
2024 Adaptive micro- and macro-knowledge incorporation for hierarchical text classification
Zijian Feng, Kezhi Mao, Hanzhang Zhou
Expert Syst. Appl.1
2023 Feature-aware conditional GAN for category text generation
Kezhi Mao, Fanfan Lin, Zijian Feng
Neurocomputing4
2022 A Deep-Learning-based System for Indoor Active Cleaning
abstract
Cleaning public areas like commercial complexes is challenging due to their sophisticated surroundings and the vast kinds of real-life dirt. Robots are required to distinguish dirts and apply corresponding cleaning strategies. In this work, we proposed an active-cleaning framework by utilizing deep-learning methods for both solid wastes detection and liquid stains segmentation. Our system consists of 4 components: a Perception module integrated with deep-learning models, a Post-processing module for projection, a Tracking module for map localization, and a Planning and Control module for cleaning strategies. Compared with classic approaches, our vision-based system significantly improves cleaning efficiency. Besides, we released the largest real-world indoor hybrid dirt cleaning dataset (HD10K) containing 10K labeled images, together with a track-level evaluation metric for better cleaning performance measurement. The proposed deep-learning based system is verified with extensive experiments on our dataset, and deployed to Gaussian Robotics's robots operating globally. Dataset is available at: https://gaussianopensource.github.io/projects/active_cleaning.
Yike Yun, Linjie Hou, Zijian Feng, Ruonan He, Weitao Guo, Baoxing Qin
IROS3
2022 Tailored text augmentation for sentiment analysis
Zijian Feng, Hanzhang Zhou, Zixiao Zhu, Kezhi Mao
Expert Syst. Appl.1
2021 Interaction via Bi-directional Graph of Semantic Region Affinity for Scene Parsing
abstract
In this work, we devote to address the challenging problem of scene parsing. It is well known that pixels in an image are highly correlated with each other, especially those from the same semantic region, while treating pixels independently fails to take advantage of such correlations. In this work, we treat each respective region in an image as a whole, and capture the structure topology as well as the affinity among different regions. To this end, we first divide the entire feature maps to different regions and extract respective global features from them. Next, we construct a directed graph whose nodes are regional features, and the bi-directional edges connecting every two nodes are the affinities between the regional features they represent. After that, we transfer the affinity-aware nodes in the directed graph back to corresponding regions of the image, which helps to model the region dependencies and mitigate unrealistic results. In addition, to further boost the correlation among pixels, we propose a region-level loss that evaluates all pixels in a region as a whole and motivates the network to learn the exclusive regional feature per class. With the proposed approach, we achieves new state-of-the-art segmentation results on PASCAL-Context, ADE20K, and COCO-Stuff consistently.
Henghui Ding, Hui Zhang 0100, Jun Liu 0036, Zijian Feng, Xudong Jiang 0001
ICCV5
2021 MT-ORL: Multi-Task Occlusion Relationship Learning
abstract
Retrieving occlusion relation among objects in a single image is challenging due to sparsity of boundaries in image. We observe two key issues in existing works: firstly, lack of an architecture which can exploit the limited amount of coupling in the decoder stage between the two subtasks, namely occlusion boundary extraction and occlusion orientation prediction, and secondly, improper representation of occlusion orientation. In this paper, we propose a novel architecture called Occlusion-shared and Path-separated Network (OPNet), which solves the first issue by exploiting rich occlusion cues in shared high-level features and structured spatial information in task-specific low-level features. We then design a simple but effective orthogonal occlusion representation (OOR) to tackle the second issue. Our method surpasses the state-of-the-art methods by 6.1%/8.3% Boundary-AP and 6.5%/10% Orientation-AP on standard PIOD/BSDS ownership datasets. Code is available at https://github.com/fengpanhe/MT-ORL.
Panhe Feng, Qi She, Lei Zhu 0012, Lin Zhang 0040, Zijian Feng, Changhu Wang, Chunpeng Li, Xuejing Kang, Anlong Ming
ICCV6
2021 MINE: Towards Continuous Depth MPI with NeRF for Novel View Synthesis
abstract
In this paper, we propose MINE to perform novel view synthesis and depth estimation via dense 3D reconstruction from a single image. Our approach is a continuous depth generalization of the Multiplane Images (MPI) by introducing the NEural radiance fields (NeRF). Given a single image as input, MINE predicts a 4-channel image (RGB and volume density) at arbitrary depth values to jointly reconstruct the camera frustum and fill in occluded contents. The reconstructed and inpainted frustum can then be easily rendered into novel RGB or depth views using differentiable rendering. Extensive experiments on RealEstate10K, KITTI and Flowers Light Fields show that our MINE outperforms state-of-the-art by a large margin in novel view synthesis. We also achieve competitive results in depth estimation on iBims-1 and NYU-v2 without annotated depth supervision. Our source code is available at https://github.com/vincentfung13/MINE.
Zijian Feng, Qi She, Henghui Ding, Changhu Wang, Gim Hee Lee
ICCV2
2021 Manifold Learning Based on Straight-Like Geodesics and Local Coordinates
abstract
In this article, a manifold learning algorithm based on straight-like geodesics and local coordinates is proposed, called SGLC-ML for short. The contribution and innovation of SGLC-ML lie in that; first, SGLC-ML divides the manifold data into a number of straight-like geodesics, instead of a number of local areas like many manifold learning algorithms do. Figuratively speaking, SGLC-ML covers manifold data set with a sparse net woven with threads (straight-like geodesics), while other manifold learning algorithms with a tight roof made of titles (local areas). Second, SGLC-ML maps all straight-like geodesics into straight lines of a low-dimensional Euclidean space. All these straight lines start from the same point and extend along the same coordinate axis. These straight lines are exactly the local coordinates of straight-like geodesics as described in the mathematical definition of the manifold. With the help of local coordinates, dimensionality reduction can be divided into two relatively simple processes: calculation and alignment of local coordinates. However, many manifold learning algorithms seem to ignore the advantages of local coordinates. The experimental results between SGLC-ML and other state-of-the-art algorithms are presented to verify the good performance of SGLC-ML.
Zhengming Ma, Zengrong Zhan, Zijian Feng, Jiajing Guo
IEEE Trans. Neural Networks Learn. Syst.3