Xinyi Zeng

dblp:260/7181 · DBLP profile ↗
← Back
23ranked-venue papers
9as first author
23since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 14 · 4 first-author · 14 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 3 first-author · 10 since 2021Artificial intelligence and machine learning · 8 · 5 first-author · 8 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 PromptEmo: Learning Emotion with Bilateral Textual Prompts in Multi-Domain Open-set Scenarios
abstract
Facial Expression Recognition (FER) is crucial to human-computer interaction. Existing cross-domain FER (CD-FER) methods mainly focus on single-source closed-set scenarios, transferring knowledge from a single source domain to a target domain with identical class sets. However, CD-FER faces two real-world challenges: 1) the need to leverage information from multiple sources, leading to multi-domain shift, and 2) the necessity to recognize unseen target classes, resulting in class shift. These issues give rise to a novel and challenging task, which we define as Multi-domain Open-set FER (MO-FER). In this paper, we propose PromptEmo, a novel CLIP-based framework that leverages bilateral textual prompts to address both shifts in the MO-FER task. Leveraging the generalizability of LLM, PromptEmo constructs trainable positive prompts with LLM-generated emotion descriptions for seen classes, as well as template-derived negative prompts to enhance the reasoning for unseen classes. Then, we introduce a modal-task optimization paradigm organized from two perspectives: textual semantics and visual domains, yielding Intra-modal Space-specific Optimization (ISO) and Cross-modal Emotion-aware Interaction (CEI) strategies. ISO refines the CLIP-based textual space to ensure semantic separation between bilateral prompts and improves the latent visual space by promoting inter-domain alignment. Founded on ISO, CEI facilitates effective vision-language interactions, resulting in four joint loss terms that improve emotion recognition by shaping a domain-invariant, discriminative feature space. PromptEmo surpasses the current SOTA method by 7.7% AUC on unseen classes across four FER datasets, serving as a strong baseline for the MO-FER task.
Xinyi Zeng, Yuxiang Yang 0009, Pinxian Zeng, Wenxia Yin, Bo Liu 0113, Xi Wu 0004, Yan Wang 0015
AAAI1
2026 Rethinking the Reliability of Multi-agent System: A Perspective from Byzantine Fault Tolerance
abstract
Ensuring the reliability of agent architectures and effectively identifying problematic agents when failures occur are crucial challenges in multi-agent systems (MAS). Advances in large language models (LLMs) have established LLM-based agents as a major branch of MAS, enabling major breakthroughs in complex problem solving and world modeling. However, the reliability implications of this shift remain largely unexplored. i.e., whether substituting traditional agents with LLM-based agents can effectively enhance the reliability of MAS. In this work, we investigate and quantify the reliability of LLM-based agents from the perspective of Byzantine fault tolerance. We observe that LLM-based agents demonstrate stronger skepticism when processing erroneous message flows, a characteristic that enables them to outperform traditional agents across different topological structures. Motivated by the results of the pilot experiment, we design CP-WBFT, a confidence probe-based weighted Byzantine Fault Tolerant consensus mechanism to enhance the stability of MAS with different topologies. It capitalizes on the intrinsic reflective and discriminative capabilities of LLMs by employing a probe-based, weighted information flow transmission method to improve the reliability of LLM-based agents. Extensive experiments demonstrate that CP-WBFT achieves superior performance across diverse network topologies under extreme Byzantine conditions (85.7 % fault rate). Notably, our approach surpasses traditional methods by attaining remarkable accuracy on various topologies and maintaining strong reliability in both mathematical reasoning and safety assessment tasks.
Lifan Zheng, Qinghong Yin, Xinyi Zeng
AAAI5
2026 Beyond shortcuts: Mitigating spurious correlations in radiological diagnosis with causal intervention
Xinyi Zeng, Jia Wang 0009, Yi Dong 0002, Wei Wang 0042, Yanji Jiang, Haiyang Zhang 0004
Knowl. Based Syst.1
2026 MGTP: Multi-Granularity Textual Prompts for Low-Dose Brain PET Image Denoising via Adversarial Diffusion Model
abstract
Positron emission tomography (PET) is an advanced nuclear imaging technique and has been widely applied in clinic. However, radiation risks associated with standard-dose PET imaging raise health concerns, whereas the quality of low-dose PET images fails to meet clinical requirements. To reduce the tracer dose while maintaining image quality, it is of great interest to estimate high-quality PET images from low-dose images. However, existing low-dose PET image denoising methods primarily focus on image data, overlooking crucial information in non-image textual data such as patients' clinical tabular and textual descriptions of general image quality. This neglect can lead to subpar denoising quality with inaccurate contexts and poor details. To address these problems, in this paper, we propose Multi-Granularity Textual Prompts, namely MGTP, to denoise low-dose PET images via an adversarial diffusion model. Different from prior methods that rely solely on image conditioning, our MGTP innovatively introduces textual prompts spanning diverse granularities to capture both high-level semantic-related contexts and low-level degradation-related details. To harmonize multi-granularity textual prompts with low-dose PET images, we design a Cross-Modality Selective Conditioning (CMSC) module, which prioritizes semantic- and detail-relevant information while eliminating irrelevant components. The resulting features are fed into diffusion model as conditions, enforcing a more controlled diffusion process. In addition, we develop a Masked Prompt Reconstruction Network (MPR-Net) to enhance the preservation of semantics and details in denoised images, mitigating distortions brought by the random noise in the diffusion process. Experiments on clinical PET data show that our method achieves the state-of-the-art performance.
Xinyi Zeng, Pinxian Zeng, Bo Liu 0113, Xi Wu 0004, Deng Xiong, Jiliu Zhou, Yan Wang 0015, Dinggang Shen
IEEE J. Biomed. Health Informatics2
2025 Root Defense Strategies: Ensuring Safety of LLM at the Decoding Level
abstract
Large language models (LLMs) have demonstrated immense utility across various industries.However, as LLMs advance, the risk of harmful outputs increases due to incorrect or malicious prompts.While current methods effectively address jailbreak risks, they share common limitations: 1) Judging harmful outputs from the prefill-level lacks utilization of the model's decoding outputs, leading to relatively lower effectiveness and robustness.2) Rejecting potentially harmful outputs based on a single evaluation can significantly impair the model's helpfulness.To address the above issues, we examine LLMs' capability to recognize harmful outputs, revealing and quantifying their proficiency in assessing the danger of previous tokens.Motivated by pilot experiment results, we design a robust defense mechanism at the decoding level.Our novel decoder-oriented, step-by-step defense architecture corrects the outputs of harmful queries directly rather than rejecting them outright.We introduce speculative decoding to enhance usability and facilitate deployment to boost safe decoding speed.Extensive experiments demonstrate that our method improves model security without compromising reasoning speed.Notably, our method leverages the model's ability to discern hazardous information, maintaining its helpfulness compared to existing methods 1 .
Xinyi Zeng, Yuying Shang
ACL (1)1
2025 Skin Disease Classification with LVLMs: An Empirical Study
abstract
Skin diseases pose significant challenges to accurate and efficient diagnosis, often due to their diverse and complex representations. This study investigates the capabilities and limitations of Large Vision-Language Models (LVLMs) in addressing these challenges through skin disease classification tasks. We evaluated LVLMs in zero-shot, few-shot, and finetuning scenarios, exploring their performance, bias, and potential for improvement. Results show that LVLMs lack perceptual granularity in skin disease, though positive signals are also observed. Our findings underscore the necessity for domain- specific optimisation and highlight opportunities for advancing LVLMs in medical diagnostics through innovative strategies and collaborative efforts.
Xinyi Zeng, Haiyang Zhang 0004, Wei Wang 0042
CSCWD1
2025 MAK-GAN: Multi-level Adaptive Convolutional Kernels for Asymmetric Multi-modal PET Reconstruction
Xinyi Zeng, Pinxian Zeng, Yan Wang 0015, Luping Zhou, Caiwen Jiang, Han Zhang 0002, Dinggang Shen
MICCAI (2)1
2025 Incorporating the Refractory Period into Spiking Neural Networks through Spike-Triggered Threshold Dynamics
abstract
As the third generation of neural networks, spiking neural networks (SNNs) have recently gained widespread attention for their biological plausibility, energy efficiency, and effectiveness in processing neuromorphic datasets. To better emulate biological neurons, various models such as Integrate-and-Fire (IF) and Leaky Integrate-and-Fire (LIF) have been widely adopted in SNNs. However, these neuron models overlook the refractory period, a fundamental characteristic of biological neurons. Research on excitable neurons reveal that after firing, neurons enter a refractory period during which they are temporarily unresponsive to subsequent stimuli. This mechanism is critical for preventing over-excitation and mitigating interference from aberrant signals. Therefore, we propose a simple yet effective method to incorporate the refractory period into spiking LIF neurons through spike-triggered threshold dynamics, termed RPLIF. Our method ensures that each spike accurately encodes neural information, effectively preventing neuron over-excitation under continuous inputs and interference from anomalous inputs. Incorporating the refractory period into LIF neurons is seamless and computationally efficient, enhancing robustness and efficiency while yielding better performance with negligible overhead. To the best of our knowledge, RPLIF achieves state-of-the-art performance on Cifar10-DVS(82.40%) and N-Caltech101(83.35%) with fewer timesteps and demonstrates superior performance on DVS128 Gesture(97.22%) at low latency.
Xinyi Zeng, Zhe Xue, Pinxian Zeng, Yan Wang 0015
ACM Multimedia2
2025 From Pixels to Tokens: Revisiting Object Hallucinations in Large Vision-Language Models
abstract
Hallucination in large vision-language models (LVLMs) is a significant challenge, i.e., generating objects that are not present in the visual input, which significantly compromises the reliability of models. Recent studies often attribute hallucinations to a lack of visual understanding, yet ignore a more fundamental issue: the model's inability to effectively extract or decouple visual features. In this paper, we revisit the hallucinations in LVLMs from an architectural perspective, investigating whether the primary cause lies in the visual encoder (feature extraction) or the modal alignment module (feature decoupling). Motivated by our preliminary findings, we propose a parameter-efficient fine-tuning strategy, PATCH, to mitigate hallucinations in LVLMs. This plug-and-play method can be integrated into various LVLMs, leveraging adaptive virtual tokens to extract object features from bounding boxes, thereby addressing hallucinations stemming from inadequate feature decoupling. PATCH achieves state-of-the-art performance across multiple multi-modal hallucination datasets and demonstrates significant improvements in general capabilities. We hope this work provides deeper insights into the underlying causes of hallucinations in LVLMs, fostering further advancements and innovation in this field. The code will be available at https://github.com/YuyingShang/PATCH.
Yuying Shang, Xinyi Zeng, Yutao Zhu 0001, Xiao Yang 0028, Zhengwei Fang
ACM Multimedia2
2025 RAIN: Reconstructed-aware in-context enhancement with graph denoising for session-based recommendation
Xinyi Zeng, Shuchao Li, Zequn Zhang, Li Jin 0001, Zhi Guo, Kaiwen Wei
Neural Networks1
2025 Multi-Modal Long-Short Distance Attention-Based Transformer-GAN for PET Reconstruction With Auxiliary MRI
abstract
To obtain high-quality PET scans while minimizing potential radiation hazards for patients, various GAN-based methods have been developed to reconstruct high-quality standard-count PET (SPET) images from low-count PET (LPET) ones. While recent efforts try to integrate MRI or CT to enhance reconstruction in a multi-modal way, current architectures mainly face two limitations: 1) CNN backbones or simple Transformer bottleneck layers are insufficient for robust semantic understanding; and 2) the identical strategies for multi-modal feature extraction and fusion overlook each modality’s respective importance for the reconstruction task. In this work, we propose the Multi-modal Long-Short Distance Attention-based Transformer-GAN (MLSDA-GAN), a novel network combining 3D transformer and CNN architecture for PET image reconstruction. Specifically, to extract fine-grained features with a small number of parameters, our MLSDA-GAN integrates multi-scale convolution into the embedding part of the transformer. As for our multi-modal design, given the strong correlation between LPET and SPET in structural characteristics, we treat MRI as an auxiliary modality to LPET and achieve effective multi-modal extraction and fusion strategies. These strategies include 1) a PET-specific Self-attention Extraction (PSE) block for comprehensive feature extraction of the primary LPET and 2) a Multi-modality Cross-attention Fusion (MCF) block for effective multi-modal interaction and fusion, enabling us to more efficiently model both long- and short-range relationships in the corresponding feature extraction and fusion processes. Experiments demonstrate superiority of our method quantitatively and qualitatively. Code is available athttps://github.com/Aru321/MLSDA-GAN.
Pinxian Zeng, Xinyi Zeng, Yan Wang 0015, Luping Zhou, Chen Zu, Xi Wu 0004, Jiliu Zhou, Dinggang Shen
IEEE Trans. Circuits Syst. Video Technol.2
2025 Adaptive Hardness-Driven Augmentation and Alignment Strategies for Multisource Domain Adaptations
abstract
Multisource domain adaptation (MDA) aims to transfer knowledge from multiple labeled source domains to an unlabeled target domain. Nevertheless, traditional methods primarily focus on achieving interdomain alignment through sample-level constraints, such as maximum mean discrepancy (MMD), neglecting three pivotal aspects: 1) the potential of data augmentation; 2) the significance of intradomain alignment; and 3) the design of cluster-level constraints. In this article, we introduce a novel hardness-driven strategy for MDA tasks, named $\mathrm {A}^{3}\mathrm {MDA}$ , which collectively considers these three aspects through adaptive hardness quantification and utilization in both data augmentation and domain alignment. To achieve this, $\mathrm {A}^{3}\mathrm {MDA}$ progressively proposes three adaptive hardness measurements (AHMs), i.e., basic, smooth, and comparative AHMs, each incorporating distinct mechanisms for diverse scenarios. Specifically, basic AHM aims to gauge the instantaneous hardness for each source/target sample. Then, hardness values measured by smooth AHM will adaptively adjust the intensity level of strong data augmentation to maintain compatibility with the model's generalization capacity. In contrast, comparative AHM is designed to facilitate cluster-level constraints. By leveraging hardness values as sample-specific weights, the traditional MMD is enhanced into a weighted-clustered variant, strengthening the robustness and precision of interdomain alignment. As for the often-neglected intradomain alignment, we adaptively construct a pseudo-contrastive matrix (PCM) by selecting harder samples based on the hardness rankings, enhancing the quality of pseudo-labels, and shaping a well-clustered target feature space. Experiments on multiple MDA benchmarks show that $\mathrm {A}^{3}\mathrm {MDA}$ outperforms other methods.
Yuxiang Yang 0009, Xinyi Zeng, Pinxian Zeng, Chen Zu, Binyu Yan, Jiliu Zhou, Yan Wang 0015
IEEE Trans. Neural Networks Learn. Syst.2
2024 MCAD: Multi-modal Conditioned Adversarial Diffusion Model for High-Quality PET Image Reconstruction
Xinyi Zeng, Pinxian Zeng, Bo Liu 0113, Xi Wu 0004, Jiliu Zhou, Yan Wang 0015
MICCAI (7)2
2024 Common Vision-Language Attention for Text-Guided Medical Image Segmentation of Pneumonia
Yunpeng Guo, Xinyi Zeng, Pinxian Zeng, Yuchen Fei, Lu Wen, Jiliu Zhou, Yan Wang 0015
MICCAI (9)2
2024 Textmatch: Using Text Prompts to Improve Semi-supervised Medical Image Segmentation
Aibing Li, Xinyi Zeng, Pinxian Zeng, Sixian Ding, Chengdi Wang, Yan Wang 0015
MICCAI (8)2
2024 ABP: Asymmetric Bilateral Prompting for Text-Guided Medical Image Segmentation
Xinyi Zeng, Pinxian Zeng, Aibing Li, Bo Liu 0113, Chengdi Wang, Yan Wang 0015
MICCAI (9)1
2024 Learning with Alignments: Tackling the Inter- and Intra-domain Shifts for Cross-multidomain Facial Expression Recognition
abstract
Facial Expression Recognition (FER) holds significant importance in human-computer interactions. Existing cross-domain FER methods often transfer knowledge solely from a single labeled source domain to an unlabeled target domain, neglecting the comprehensive information across multiple sources. Nevertheless, cross-multidomain FER (CMFER) is very challenging for (i) the inherent inter-domain shifts across multiple domains and (ii) the intra-domain shifts stemming from the ambiguous expressions and low inter-class distinctions. In this paper, we propose a novel Learning with Alignments CMFER framework, named LA-CMFER, to handle both inter- and intra-domain shifts. Specifically, LA-CMFER is constructed with a global branch and a local branch to extract features from the full images and local subtle expressions, respectively. Based on this, LA-CMFER presents a dual-level inter-domain alignment method to force the model to prioritize hard-to-align samples in knowledge transfer at a sample level while gradually generating a well-clustered feature space with the guidance of class attributes at a cluster level, thus narrowing the inter-domain shifts. To address the intra-domain shifts, LA-CMFER introduces a multi-view intra-domain alignment method with a multi-view clustering consistency constraint where a prediction similarity matrix is built to pursue consistency between the global and local views, thus refining pseudo labels and eliminating latent noise. Extensive experiments on six benchmark datasets have validated the superiority of our LA-CMFER.
Yuxiang Yang 0009, Lu Wen, Xinyi Zeng, Xi Wu 0004, Jiliu Zhou, Yan Wang 0015
ACM Multimedia3
2024 Graph-enhanced context aware framework for session-based recommendation
Xinyi Zeng, Zequn Zhang, Shuchao Li, Zhi Guo, Li Jin 0001, Xian Sun 0001
Neurocomputing1
2024 Semi-supervised medical image segmentation via hard positives oriented contrastive learning
Cheng Tang 0003, Xinyi Zeng, Luping Zhou, Qizheng Zhou, Xi Wu 0004, Hongping Ren, Jiliu Zhou, Yan Wang 0015
Pattern Recognit.2
2024 Prior Knowledge-Guided Triple-Domain Transformer-GAN for Direct PET Reconstruction From Low-Count Sinograms
abstract
To obtain high-quality positron emission tomography (PET) images while minimizing radiation exposure, numerous methods have been dedicated to acquiring standard-count PET (SPET) from low-count PET (LPET). However, current methods have failed to take full advantage of the different emphasized information from multiple domains, i.e., the sinogram, image, and frequency domains, resulting in the loss of crucial details. Meanwhile, they overlook the unique inner-structure of the sinograms, thereby failing to fully capture its structural characteristics and relationships. To alleviate these problems, in this paper, we proposed a prior knowledge-guided transformer-GAN that unites triple domains of sinogram, image, and frequency to directly reconstruct SPET images from LPET sinograms, namely PK-TriDo. Our PK-TriDo consists of a Sinogram Inner-Structure-based Denoising Transformer (SISD-Former) to denoise the input LPET sinogram, a Frequency-adapted Image Reconstruction Transformer (FaIR-Former) to reconstruct high-quality SPET images from the denoised sinograms guided by the image domain prior knowledge, and an Adversarial Network (AdvNet) to further enhance the reconstruction quality via adversarial training. Specifically tailored for the PET imaging mechanism, we injected a sinogram embedding module that partitions the sinograms by rows and columns to obtain 1D sequences of angles and distances to faithfully preserve the inner-structure of the sinograms. Moreover, to mitigate high-frequency distortions and enhance reconstruction details, we integrated global-local frequency parsers (GLFPs) into FaIR-Former to calibrate the distributions and proportions of different frequency bands, thus compelling the network to preserve high-frequency details. Evaluations on three datasets with different dose levels and imaging scenarios demonstrated that our PK-TriDo outperforms the state-of-the-art methods.
Pinxian Zeng, Xinyi Zeng, Jiliu Zhou, Yan Wang 0015, Dinggang Shen
IEEE Trans. Medical Imaging3
2023 TriDo-Former: A Triple-Domain Transformer for Direct PET Reconstruction from Low-Dose Sinograms
Pinxian Zeng, Xinyi Zeng, Xi Wu 0004, Jiliu Zhou, Yan Wang 0015, Dinggang Shen
MICCAI (10)3
2023 DBTrans: A Dual-Branch Vision Transformer for Multi-Modal Brain Tumor Segmentation
Xinyi Zeng, Pinxian Zeng, Cheng Tang 0003, Binyu Yan, Yan Wang 0015
MICCAI (4)1
2022 3D CVT-GAN: A 3D Convolutional Vision Transformer-GAN for PET Reconstruction
Pinxian Zeng, Luping Zhou, Chen Zu, Xinyi Zeng, Zhengyang Jiao, Xi Wu 0004, Jiliu Zhou, Dinggang Shen, Yan Wang 0015
MICCAI (6)4