Zixun Zhang

dblp:274/0310 · DBLP profile ↗
← Back
15ranked-venue papers
3as first author
14since 2021 · last 2027
0000-0003-3141-1278ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 2 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 5 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2027 MedDATP: Adapting CLIP for few-shot medical image classification via domain adapter and task prompts
Zixun Zhang, Yuncheng Jiang 0002, Jun Wei 0006, Huazhu Fu, Shuguang Cui, Tao Luo 0014, Zhen Li 0026
Expert Syst. Appl.1
2025 DCP: Dual-Cue Pruning for Efficient Large Vision-Language Models
abstract
Large Vision-Language Models (LVLMs) achieve remarkable performance in multimodal tasks but suffer from high computational costs due to the large number of visual tokens.Existing pruning methods either apply after visual tokens enter the LLM or perform pre-pruning based solely on visual attention.Both fail to balance efficiency and semantic alignment, as post-pruning incurs redundant computation, while visual-only pre-pruning overlooks multimodal relevance.To address this limitation, we propose Dual-Cue Pruning (DCP), a novel cross-modal pruning framework that jointly considers textual semantics and visual selfattention.DCP consists of a text-aware computation module, which employs a gradientweighted attention mechanism to enhance textvisual alignment, and an image-aware computation module, which utilizes deep-layer selfattention distributions to retain essential structural information.By integrating both cues, DCP adaptively selects the most informative visual tokens, achieving efficient inference acceleration while maintaining strong task performance.Experimental results show that DCP can retain only 25% of the visual tokens, with a minimal performance degradation of 0.063% on LLaVA-1.5-13B,demonstrating its effectiveness in balancing efficiency and accuracy.
Zixun Zhang, Yuting Zeng, Chunzhao Xie, Tongxuan Liu, Lechao Cheng
EMNLP2
2025 Self-distillation with model averaging
Xiaozhe Gu, Zixun Zhang, Rick Siow Mong Goh, Tao Luo 0014
Inf. Sci.2
2024 Let Video Teaches You More: Video-to-Image Knowledge Distillation using Detection TRansformer for Medical Video Lesion Detection
abstract
AI-assisted lesion detection models play a crucial role in the early screening of cancer. However, previous image-based models ignore the inter-frame contextual information present in videos. On the other hand, video-based models capture the inter-frame context but are computationally expensive. To mitigate this contradiction, we delve into Video-to-Image knowledge distillation leveraging DEtection TRansformer (V2I-DETR) for the task of medical video lesion detection. V2I-DETR adopts a teacher-student network paradigm. The teacher network aims at extracting temporal contexts from multiple frames and transferring them to the student network, and the student network is an image-based model dedicated to fast prediction in inference. By distilling multi-frame contexts into a single frame, the proposed V2I-DETR combines the advantages of utilizing temporal contexts from video-based models and the inference speed of image-based models. Through extensive experiments, V2I-DETR outperforms previous state-of-the-art methods by a large margin while achieving the real-time inference speed (30 FPS) as the image-based model.
Yuncheng Jiang 0002, Zixun Zhang, Jun Wei 0006, Chun-Mei Feng 0001, Guanbin Li, Shuguang Cui, Zhen Li 0026
BIBM2
2024 Flexible Graph Neural Diffusion with Latent Class Representation Learning
abstract
In existing graph data, the connection relationships often exhibit uniform weights, leading to the model aggregating neighboring nodes with equal weights across various connection types. However, this uniform aggregation of diverse information diminishes the discriminability of node representations, contributing significantly to the over-smoothing issue in models. In this paper, we propose the Flexible Graph Neural Diffusion (FGND) model, incorporating latent class representation to address the misalignment between graph topology and node features. In particular, we combine latent class representation learning with the inherent graph topology to reconstruct the diffusion matrix during the graph diffusion process. We introduce the sim metric to quantify the degree of mismatch between graph topology and node features. By flexibly adjusting the dependency level on node features through the hyperparameter, we accommodate diverse adjacency relationships. The effective filtering of noise in the topology also allows the model to capture higher order information, significantly alleviating the over-smoothing problem. Meanwhile, we model the graphical diffusion process as a set of differential equations and employ advanced partial differential equation tools to obtain more accurate solutions. Empirical evaluations on five benchmarks reveal that our FGND model outperforms existing popular GNN methods in terms of both overall performance and stability under data perturbations. Meanwhile, our model exhibits superior performance in comparison to models tailored for heterogeneous graphs and those designed to address oversmoothing issues.
Liangtian Wan, Huijin Han, Lu Sun 0004, Zixun Zhang, Zhaolong Ning, Xiaoran Yan, Feng Xia 0001
KDD4
2024 Towards a Benchmark for Colorectal Cancer Segmentation in Endorectal Ultrasound Videos: Dataset and Model Development
Yuncheng Jiang 0002, Yiwen Hu 0001, Zixun Zhang, Jun Wei 0006, Chun-Mei Feng 0001, Xuemei Tang, Yong Liu 0026, Shuguang Cui, Zhen Li 0026
MICCAI (8)3
2024 ECC-PolypDet: Enhanced CenterNet With Contrastive Learning for Automatic Polyp Detection
abstract
Accurate polyp detection is critical for early colorectal cancer diagnosis. Although remarkable progress has been achieved in recent years, the complex colon environment and concealed polyps with unclear boundaries still pose severe challenges in this area. Existing methods either involve computationally expensive context aggregation or lack prior modeling of polyps, resulting in poor performance in challenging cases. In this paper, we propose the Enhanced CenterNet with Contrastive Learning (ECC-PolypDet), a two-stage training & end-to-end inference framework that leverages images and bounding box annotations to train a general model and fine-tune it based on the inference score to obtain a final robust model. Specifically, we conduct Box-assisted Contrastive Learning (BCL) during training to minimize the intra-class difference and maximize the inter-class difference between foreground polyps and backgrounds, enabling our model to capture concealed polyps. Moreover, to enhance the recognition of small polyps, we design the Semantic Flow-guided Feature Pyramid Network (SFFPN) to aggregate multi-scale features and the Heatmap Propagation (HP) module to boost the model's attention on polyp targets. In the fine-tuning stage, we introduce the IoU-guided Sample Re-weighting (ISR) mechanism to prioritize hard samples by adaptively adjusting the loss weight for each sample during fine-tuning. Extensive experiments on six large-scale colonoscopy datasets demonstrate the superiority of our model compared with previous state-of-the-art detectors.
Yuncheng Jiang 0002, Zixun Zhang, Yiwen Hu 0001, Guanbin Li, Shuguang Cui, Silin Huang, Zhen Li 0026
IEEE J. Biomed. Health Informatics2
2024 Hierarchical Weight Averaging for Deep Neural Networks
abstract
Despite simplicity, stochastic gradient descent (SGD)-like algorithms are successful in training deep neural networks (DNNs). Among various attempts to improve SGD, weight averaging (WA), which averages the weights of multiple models, has recently received much attention in the literature. Broadly, WA falls into two categories: 1) online WA, which averages the weights of multiple models trained in parallel, is designed for reducing the gradient communication overhead of parallel mini-batch SGD and 2) offline WA, which averages the weights of one model at different checkpoints, is typically used to improve the generalization ability of DNNs. Though online and offline WA are similar in form, they are seldom associated with each other. Besides, these methods typically perform either offline parameter averaging or online parameter averaging, but not both. In this work, we first attempt to incorporate online and offline WA into a general training framework termed hierarchical WA (HWA). By leveraging both the online and offline averaging manners, HWA is able to achieve both faster convergence speed and superior generalization performance without any fancy learning rate adjustment. Besides, we also analyze the issues faced by the existing WA methods, and how our HWA addresses them, empirically. Finally, extensive experiments verify that HWA outperforms the state-of-the-art methods significantly.
Xiaozhe Gu, Zixun Zhang, Yuncheng Jiang 0002, Tao Luo 0014, Ruimao Zhang, Shuguang Cui, Zhen Li 0026
IEEE Trans. Neural Networks Learn. Syst.2
2024 Joint Signal Detection and Automatic Modulation Classification via Deep Learning
abstract
Signal detection and modulation classification are two crucial tasks in various wireless communication systems. Different from prior works that investigate them independently, this paper studies the joint signal detection and automatic modulation classification (AMC) by considering a realistic and complex scenario, in which multiple signals with different modulation schemes coexist at different carrier frequencies. We first generate a coexisting RADIOML dataset (CRML23) to facilitate the joint design. Different from the publicly available AMC dataset, ignoring the signal detection step and containing only one signal, our synthetic dataset covers the more realistic multiple-signal coexisting scenario. Then, we present a joint framework for detection and classification (JDM) for such a multiple-signal coexisting environment, which consists of two modules for signal detection and AMC, respectively. In particular, these two modules are interconnected using a designated data structure called “proposal”. Finally, we conduct extensive simulations over the newly developed dataset, which demonstrate the effectiveness of our designs. Our code and dataset are now available as open-source resources athttps://github.com/Singingkettle/ChangShuoRadioData.
Huijun Xing, Shuo Chang, Jinke Ren, Zixun Zhang, Jie Xu 0002, Shuguang Cui
IEEE Trans. Wirel. Commun.5
2023 ScribblePolyp: Scribble-Supervised Polyp Segmentation through Dual Consistency Alignment
abstract
Automatic polyp segmentation models play a pivotal role in the clinical diagnosis of gastrointestinal diseases. In previous studies, most methods relied on fully supervised approaches, necessitating pixel-level annotations for model training. However, the creation of pixel-level annotations is both expensive and time-consuming, impeding the development of model generalization. In response to this challenge, we introduce ScribblePolyp, a novel scribble-supervised polyp segmentation framework. Unlike fully-supervised models, ScribblePolyp only requires the annotation of two lines (scribble labels) for each image, significantly reducing the labeling cost. Despite the coarse nature of scribble labels, which leave a substantial portion of pixels unlabeled, we propose a two-branch consistency alignment approach to provide supervision for these unlabeled pixels. The first branch employs transformation consistency alignment to narrow the gap between predictions under different transformations of the same input image. The second branch leverages affinity propagation to refine predictions into a soft version, extending additional supervision to unlabeled pixels. In summary, ScribblePolyp is an efficient model that does not rely on teacher models or moving average pseudo labels during training. Extensive experiments on the SUN-SEG dataset underscore the effectiveness of ScribblePolyp, achieving a Dice score of 0.8155, with the potential for a 1.8% improvement in the Dice score through a straightforward self-training strategy.
Zixun Zhang, Yuncheng Jiang 0002, Jun Wei 0006, Hannah Cui, Zhen Li 0026
BIBM1
2023 YONA: You Only Need One Adjacent Reference-Frame for Accurate and Fast Video Polyp Detection
Yuncheng Jiang 0002, Zixun Zhang, Ruimao Zhang, Guanbin Li, Shuguang Cui, Zhen Li 0026
MICCAI (5)2
2023 The "rebirth" of traditional musical instrument: An interactive installation based on augmented reality and somatosensory technology to empower the exhibition of chimes
abstract
Abstract Tangible cultural heritage is rich in historical, artistic and scientific value, but due to its own characteristics and the constraints of museum displays, the key issue facing us today is how to utilize and activate cultural heritage to enhance the experience and learning interest of the audience. Focusing on the Chinese cultural treasure, the chimes, this article innovatively proposes an interaction system based on digital augmented projection and somatosensory technology to empower the aesthetic expression and knowledge dissemination of chimes. On this basis, a series of evaluations were carried out and the results showed that the system effectively provides a satisfying, personalized and fun way for users to interact with chimes and learn the traditional culture, successfully bringing this ancient instrument back to life. Moreover, this article shares in detail the research, design and development process, which provides inspiration and examples for other museums to present their tangible cultural heritage through mixed reality and interactive technology.
Wenchen Guo, Yiyuan Huang, Zixun Zhang, Guoyu Sun, Qingxiang Zeng
Comput. Animat. Virtual Worlds4
2023 A "magic world" for children: Design and development of a serious game to improve spatial ability
abstract
Abstract Research has shown that spatial perception is not only one of the essential abilities for success in science, technology, engineering, and mathematics (STEM), but is also closely related to the quality of human existence. However, for a variety of reasons, many students' spatial skills are less than ideal. In recent years, various video games are showing great potential as low‐cost but effective training tools to improve the educational level and cognitive skills. This paper presented a novel serious strategy game named Magic World. The game was designed to enhance children's spatial perception and motivation by using narrative, virtual contexts, and game mechanics that incorporate educational content with entertainment as a powerful extra‐curricular aid. A pilot study and evaluation experiment were conducted with primary school students (N = 68) and the results showed that the training had a measurable positive impact on students' spatial ability. In addition, user experience surveys of the game showed that Magic World was considered a fun, challenging, and popular game for children.
Wenchen Guo, Shucheng Li, Zixun Zhang, KuoHsiang Chang
Comput. Animat. Virtual Worlds3
2022 APAUNet: Axis Projection Attention UNet for Small Target in 3D Medical Segmentation
Yuncheng Jiang 0002, Zixun Zhang, Shixi Qin, Yao Guo 0002, Zhen Li 0026, Shuguang Cui
ACCV (6)2
2020 MetaSelection: Metaheuristic Sub-Structure Selection for Neural Network Pruning Using Evolutionary Algorithm
abstract
Neural network pruning is widely applied to various mobile applications. Previous pruning methods mainly leverage ad-hoc criteria to evaluate channel importance. In this paper, we propose an effective metaheuristic sub-structure selection (MetaSelection) method for neural network pruning. MetaSelection exploits evolutionary algorithm (EA) to search the proper sub-structure satisfying the resource constraints. In comparison with previous AutoML based methods, MetaSelection can automatically achieve the pruning rate and channel selection at the same time instead of hand-crafted criteria in a cascaded way. Regarding the tremendous search space of channel selection as a combinatorial optimization problem, we further utilize a coarse-to-fine strategy and the novel probability distribution crossover (PDC) to speed up the search procedure. Besides, MetaSelection prunes the network globally rather than in a layer-by-layer way. We evaluate MetaSelection on several appealing deep neural networks, achieving superior results with adaptive depth and width. Concretely, on ImageNet, MetaSelection achieves a top-1 accuracy of 71.5% on MobileNetV2 under 70% FLOPs constraint and a FLOPs reduction of 30% with 76.4% top-1 accuracy for ResNet50.
Zixun Zhang, Zhen Li 0026, Lin Lin 0008, Na Lei, Guanbin Li, Shuguang Cui
ECAI1