Zhixuan Chen

dblp:280/8963 · DBLP profile ↗
← Back
17ranked-venue papers
3as first author
17since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 1 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 first-author · 4 since 2021
YearPublicationVenuePosition
2026 Knowledge-Enhanced Explainable Prompting for Vision-Language Models
abstract
Large-scale vision-language models (VLMs) embedded with expansive representations and visual concepts have showcased significant potential in image and text understanding. Efficiently adapting VLMs such as CLIP to downstream tasks like few-shot image classification has garnered growing attention, with prompt learning emerging as a representative approach. However, most existing prompt-based adaptation methods, which rely solely on coarse-grained textual prompts, suffer from limited performance and interpretability when handling domain tasks that require specific knowledge. This results in a failure to satisfy the stringent trustworthiness requirements of Explainable Artificial Intelligence (XAI) in high-risk scenarios like healthcare. To address this issue, we propose a Knowledge-Enhanced Explainable Prompting (KEEP) framework that leverages fine-grained domain-specific knowledge to enhance the adaptation process of VLMs across various domains and image modalities. By incorporating retrieval augmented generation and domain foundation models, our framework can provide more reliable image-wise knowledge for prompt learning in various domains, alleviating the lack of fine-grained annotations, while offering both visual and textual explanations. Extensive experiments and explainability analyses conducted on eight datasets of different domains and image modalities demonstrate that our method simultaneously achieves superior performance and interpretability, highlighting the effectiveness of the collaboration between foundation models and XAI.
Yequan Bie, Andong Tan, Zhixuan Chen, Zhiyuan Cai, Luyang Luo, Hao Chen 0011
AAAI3
2026 OTARo: Once Tuning for All Precisions Toward Robust On-Device LLMs
abstract
Large Language Models (LLMs) fine-tuning techniques not only improve the adaptability to diverse downstream tasks, but also mitigate adverse effects of model quantization. Despite this, conventional quantization suffers from its structural limitation that hinders flexibility during the fine-tuning and deployment stages. Practical on-device tasks demand different quantization precisions (i.e. different bit-widths), e.g., understanding tasks tend to exhibit higher tolerance to reduced precision compared to generation tasks. Conventional quantization, typically relying on scaling factors that are incompatible across bit-widths, fails to support the on-device switching of precisions when confronted with complex real‑world scenarios. To overcome the dilemma, we propose OTARo, a novel method that enables on-device LLMs to flexibly switch quantization precisions while maintaining performance robustness through once fine-tuning. OTARo introduces Shared Exponent Floating Point (SEFP), a distinct quantization mechanism, to produce different bit-widths through simple mantissa truncations of a single model. Moreover, to achieve bit-width robustness in downstream applications, OTARo performs a learning process toward losses induced by different bit-widths. The method involves two critical strategies: (1) Exploitation-Exploration Bit-Width Path Search (BPS), which iteratively updates the search path via a designed scoring mechanism; (2) Low-Precision Asynchronous Accumulation (LAA), which performs asynchronous gradient accumulations and delayed updates under low bit-widths. Experiments on popular LLMs, e.g., LLaMA3.2-1B, LLaMA3-8B, demonstrate that OTARo achieves consistently strong and robust performance for all precisions.
Shaoyuan Chen, Zhixuan Chen, Zhihang Yuan, Qiang Wu 0012
AAAI2
2026 Optimizing potential-based reward automata in partially observable reinforcement learning using genetic local search
Zhixuan Chen, Chenyang Zhu 0006, Wen Si
Eng. Appl. Artif. Intell.2
2026 Breaking error coupling via divergent-convergent coordination for semi-supervised medical image segmentation
Zhixuan Chen, Yuquan Xu, Mingfeng Li, Yuefei Wang
Medical Image Anal.2
2026 A carving hierarchical information integration network for medical image segmentation
abstract
Semantic segmentation techniques are widely applied in various image analysis tasks. However, compared with natural images, medical image segmentation presents greater challenges. For instance, lesions often vary significantly in morphology, size, and structure, and are frequently accompanied by low contrast and blurred boundaries. To simultaneously preserve fine tissue structures when handling large-scale lesions and ensure the coherence of the divergent structures of vessels, tumors, and other organs while accurately segmenting adjacent cells, this paper proposes the concept of “Global Capture and Local Carving”. It introduces a model that integrates a hierarchical information fusion strategy, named CarveNet. CarveNet incorporates a carving mechanism at three levels: downsampling, feature transmission, and bottleneck processing. Structural Carving Pooling Module underpins the downsampling carving, deeply optimizing the information structure and morphology at different levels to maximize detail retention and minimize downsampling loss. Multi-window Carving ViT is employed for transmission carving, enhancing global information modeling while refining local feature representation. The bottleneck carving integrates a long-distance recurrent communication mechanism with grid-like spatial random shuffling to strengthen the robustness and diversity of feature extraction. Experiments conducted on eight medical image datasets demonstrate that CarveNet consistently delivers outstanding performance across all tasks, surpassing the second-best method in Dice coefficient by 1.136 %. This fully validates its effectiveness in terms of multi-lesion adaptability, accuracy, and generalization capability. The code is available at https://github.com/YF-W/CarveNet .
Yuefei Wang, Qinyu Zhao, Liangyan Zhao, Binxiong Li, Zhixuan Chen
Pattern Recognit.8
2026 Effective and Efficient Temporal Graph Neural Networks via Polynomial Spectral Sparsification
abstract
Temporal graph neural networks (T-GNNs) have emerged as an effective paradigm for learning dynamic node representations by modeling the temporal evolution of graph structures. Among them, personalized PageRank (PPR)-based T-GNNs have recently gained increasing attention due to their strong empirical performance and theoretical grounding. Nevertheless, existing PPR-based T-GNN methods suffer from two fundamental limitations. First, their aggregation mechanisms are typically restricted to lower order neighborhoods, which hinders the capture of higher order temporal dependences and leads to degraded performance. Second, feature propagation and feature transformation are often tightly coupled, resulting in limited scalability on large temporal graphs. To overcome these challenges, we propose temporal personalized PageRank with spectral sparsification (TPPSS), a novel framework that seamlessly incorporates temporal information into the PPR formulation while decoupling propagation from transformation through a polynomial spectral sparsification. A key technical contribution ofTPPSSlies in a unified matrix-based formulation that replaces the conventional rowwise computation of PPR-like scores, enabling efficient and scalable. Moreover, by leveraging spectral sparsification,TPPSSconstructs a sparse computational graph that effectively suppresses noise and redundancy. Extensive experiments on six datasets showcase the effectiveness, efficiency, and scalability of our solutions compared to eight competitors.
Zhixuan Chen, Longlong Lin, Tao Jia 0001
IEEE Trans. Ind. Informatics1
2026 Fast Track Anything With Sparse Spatio-Temporal Propagation for Unified Video Segmentation
abstract
Recent advances in "track-anything" models have significantly improved fine-grained video understanding by simultaneously handling multiple video segmentation and tracking tasks. However, existing models often struggle with robust and efficient temporal propagation. To address these challenges, we propose the Sparse Spatio-Temporal Propagation (SSTP) method, which achieves robust and efficient unified video segmentation by selectively leveraging key spatio-temporal features in videos. Specifically, we design a dynamic 3D spatio-temporal convolution to aggregate global multi-frame spatio-temporal information into memory frames during memory construction. Additionally, we introduce a spatio-temporal aggregation reading strategy to efficiently aggregate the relevant spatio-temporal features from multiple memory frames during memory retrieval. By combining SSTP with an image segmentation foundation model, such as the segment anything model, our method effectively addresses multiple data-scarce video segmentation tasks. Our experimental results demonstrate state-of-the-art performance on five video segmentation tasks across eleven datasets, outperforming both task-specific and unified methods. Notably, SSTP exhibits strong robustness in handling sparse, low-frame-rate videos, making it well-suited for real-world applications.
Jisheng Dang, Huicheng Zheng, Zhixuan Chen, Yulan Guo, Tat-Seng Chua
IEEE Trans. Image Process.3
2025 OSTQuant: Refining Large Language Model Quantization with Orthogonal and Scaling Transformations for Better Distribution Fitting
abstract
Post-training quantization (PTQ) has emerged as a widely adopted technique for compressing and accelerating Large Language Models (LLMs). The major challenge in LLM quantization is that uneven and heavy-tailed data distributions can expand the quantization range, thereby reducing bit precision for most values. Recent methods attempt to eliminate outliers and balance inter-channel differences by employing linear transformations; however, they remain heuristic and are often overlook optimizing the data distribution across the entire quantization space. In this paper, we introduce Quantization Space Utilization Rate (QSUR), a novel metric that effectively assesses the quantizability of transformed data by measuring the space utilization of the data in the quantization space. We complement QSUR with mathematical derivations that examine the effects and limitations of various transformations, guiding our development of Orthogonal and Scaling Transformation-based Quantization (OSTQuant). OSTQuant employs a learnable equivalent transformation, consisting of an orthogonal transformation and a scaling transformation, to optimize the distributions of weights and activations across the entire quantization space. Futhermore, we propose the KL-Top loss function, designed to mitigate noise during optimization while retaining richer semantic information within the limited calibration data imposed by PTQ. OSTQuant outperforms existing work on various LLMs and benchmarks. In the W4-only setting, it retains 99.5\% of the floating-point accuracy. In the more challenging W4A4KV4 configuration, OSTQuant reduces the performance gap by 32\% on the LLaMA-3-8B model compared to state-of-the-art methods. Code will be available.
Xing Hu 0010, Zhixuan Chen, Zukang Xu, Jiangyong Yu, Zhihang Yuan, Zhe Jiang 0004, Sifan Zhou
ICLR4
2025 MambaQuant: Quantizing the Mamba Family with Variance Aligned Rotation Methods
abstract
Mamba is an efficient sequence model that rivals Transformers and demonstrates significant potential as a foundational architecture for various tasks. Quantization is commonly used in neural networks to reduce model size and computational latency. However, applying quantization to Mamba remains underexplored, and existing quantization methods, which have been effective for CNN and Transformer models, appear inadequate for Mamba models (e.g., Quarot suffers a 21% accuracy drop on Vim-T$\dagger$ even under W8A8). We have pioneered the exploration of this issue and identified several key challenges. First, significant outliers arepresent in gate projections, output projections, and matrix multiplications. Second, Mamba’s unique parallel scan further amplifies these outliers, leading to uneven and heavy-tailed data distributions. Third, even with the application of the Hadamard transform, the variance across channels in weights and activations still remains inconsistent. To these ends, we propose MambaQuant, a post-training quantization (PTQ) framework consisting of: 1) Karhunen-Lo`eve Transformation (KLT) enhanced rotation, rendering the rotation matrix adaptable to diverse channel distributions. 2) Smooth-Fused rotation, which equalizes channel variances and can merge additional parameters into model weights. Experiments show that MambaQuant can quantize both weights and activations into 8-bit with less than 1% accuracy loss for Mamba-based vision and language tasks. To our knowledge, MambaQuant is the first comprehensive PTQ design for the Mamba family, paving the way for further advancements in its application.
Zukang Xu, Yuxuan Yue, Xing Hu 0010, Zhihang Yuan, Zixu Jiang, Zhixuan Chen, Jiangyong Yu, Sifan Zhou
ICLR7
2025 Quality-Guided Dynamic Memory for LLMs-based Long-Term Video Understanding
abstract
Using the impressive learning representation capacity of large language models (LLMs), LLM-based video understanding methods have made significant strides recently. However, most existing methods overlook the crucial importance discrepancy of frames, which often include massive low-quality frames, leading to limited performance and inferior inference efficiency, particularly for long-term videos. To this end, this paper proposes a new video understanding method called quality- guided dynamic memory network (QDM-Net). First, we design a memory quality evolution module (MQEM), which dynamically assigns weights to each frame according to contextual relationships between adjacent frames. Second, we devise a high- level quality memory bank updating mechanism (HQMBU), which selectively maintains high-quality frames in the memory bank, avoiding the negative influences of redundant frames and ensuring that the model focuses on the most informative visual cues. Extensive experiments on long-term video understanding benchmarks demonstrate that our QDM-Net consistently outperforms state-of-the-art methods, showcasing its potential in real-world applications. Our code and model will be publicly available.
Bimei Wang, Jingmei Jiao, Jisheng Dang, Qingrun Jiang, Jiyuan Lin, Zhixuan Chen, Teng Wang 0007
ICME6
2025 Instruction-aware Memory Network for Video Recognition
abstract
The rapid development of multimodal large language models (MLLMs) has highlighted their potential in video understanding. However, challenges remain in long video tasks, particularly in integrating visual features with prompt texts. Existing methods naively store processed video frames in a long-term memory bank, but neglect simple yet effective cross-modal integration. To address this, we introduce the instruction-aware memory construction (IaMC) model for long-term video understanding. By integrating visual and textual information, our model can obtain cross-modal features with robust understanding capabilities. These features are stored in a text-visual memory bank, enabling efficient long-term aggregation without surpassing LLM context or GPU memory limits. Experiments on the LVU dataset demonstrate state-of-the-art performance in video understanding and question answering, showcasing the IaMC model’s effectiveness and setting a new benchmark for long-term video analysis. The source code and trained models will be released publicly.
Bimei Wang, Haijiang Li, Jisheng Dang, Yun Wang 0053, Zhixuan Chen, Jiyuan Lin, Teng Wang 0007
ICME5
2025 MoEQuant: Enhancing Quantization for Mixture-of-Experts Large Language Models via Expert-Balanced Sampling and Affinity Guidance
abstract
Mixture-of-Experts (MoE) large language models (LLMs), which leverage dynamic routing and sparse activation to enhance efficiency and scalability, have achieved higher performance while reducing computational costs. However, these models face significant memory overheads, limiting their practical deployment and broader adoption. Post-training quantization (PTQ), a widely used method for compressing LLMs, encounters severe accuracy degradation and diminished generalization performance when applied to MoE models. This paper investigates the impact of MoE’s sparse and dynamic characteristics on quantization and identifies two primary challenges: (1) Inter-expert imbalance, referring to the uneven distribution of samples across experts, which leads to insufficient and biased calibration for less frequently utilized experts; (2) Intra-expert imbalance, arising from MoE’s unique aggregation mechanism, which leads to varying degrees of correlation between different samples and their assigned experts. To address these challenges, we propose MoEQuant, a novel quantization framework tailored for MoE LLMs. MoEQuant includes two novel techniques: 1) Expert-Balanced Self-Sampling (EBSS) is an efficient sampling method that efficiently constructs a calibration set with balanced expert distributions by leveraging the cumulative probabilities of tokens and expert balance metrics as guiding factors. 2) Affinity-Guided Quantization (AGQ), which incorporates affinities between experts and samples into the quantization process, thereby accurately assessing the impact of individual samples on different experts within the MoE layer. Experiments demonstrate that MoEQuant achieves substantial performance gains (more than 10 points accuracy gain in the HumanEval for DeepSeekMoE-16B under 4-bit quantization) and boosts efficiency.
Zhixuan Chen, Xing Hu 0010, Zukang Xu, Zhihang Yuan, Sifan Zhou, Jiangyong Yu
ICML1
2025 RWKVQuant: Quantizing the RWKV Family with Proxy Guided Hybrid of Scalar and Vector Quantization
abstract
RWKV is a modern RNN architecture with comparable performance to Transformer, but still faces challenges when deployed to resource-constrained devices. Post Training Quantization (PTQ), which is a an essential technique to reduce model size and inference latency, has been widely used in Transformer models. However, it suffers significant degradation of performance when applied to RWKV. This paper investigates and identifies two key constraints inherent in the properties of RWKV: (1) Non-linear operators hinder the parameter-fusion of both smooth- and rotation-based quantization, introducing extra computation overhead. (2) The larger amount of uniformly distributed weights poses challenges for cluster-based quantization, leading to reduced accuracy. To this end, we propose RWKVQuant, a PTQ framework tailored for RWKV models, consisting of two novel techniques: (1) a coarse-to-fine proxy capable of adaptively selecting different quantization approaches by assessing the uniformity and identifying outliers in the weights, and (2) a codebook optimization algorithm that enhances the performance of cluster-based quantization methods for element-wise multiplication in RWKV. Experiments show that RWKVQuant can quantize RWKV-6-14B into about 3-bit with less than 1% accuracy loss and 2.14$\times$ speed up.
Yuxuan Yue, Zukang Xu, Xing Hu 0010, Jiangyong Yu, Zhixuan Chen, Sifan Zhou, Zhihang Yuan
ICML6
2025 Dia-LLaMA: Towards Large Language Model-Driven CT Report Generation
Zhixuan Chen, Luyang Luo, Yequan Bie, Hao Chen 0011
MICCAI (7)1
2025 A segmentation network for generalized lesion extraction with semantic fusion of transformer with value vector enhancement
Yuefei Wang, Yuanhong Wei, Zhixuan Chen
Expert Syst. Appl.8
2025 A feature enhancement network based on image partitioning in a multi-branch encoder-decoder architecture
Yuefei Wang, Zhixuan Chen, Yuquan Xu, Ruixin Cao, Liangyan Zhao, Yixi Yang
Knowl. Based Syst.5
2024 XCoOp: Explainable Prompt Learning for Computer-Aided Diagnosis via Concept-Guided Context Optimization
Yequan Bie, Luyang Luo, Zhixuan Chen, Hao Chen 0011
MICCAI (12)3