Xueyi Zhang 0001

dblp:90/10666-1 · DBLP profile ↗
← Back
20ranked-venue papers
10as first author
20since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 5 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 7 first-author · 9 since 2021Systems, architecture and hardware · 3 · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Rep Deep & Machine Learning: Exemplar-Free Continual Video Action Recognition via Slow-Fast Collaborative Learning
abstract
In real-world applications, video action recognition models must continuously learn new action categories while retaining previously acquired knowledge. However, most existing approaches rely on storing historical data for replay, which introduces storage burdens and raises data privacy concerns. To address these challenges, we investigate the problem of Exemplar-Free Continual Video Action Recognition (EF-CVAR) and propose a novel framework named Slow-Fast Collaborative Learning (SFCL). SFCL integrates two complementary learning paradigms: a slow branch based on gradient-driven deep learning, which provides strong adaptability to new tasks, and a fast branch based on analytic learning (e.g., Recursive Least Squares), which efficiently preserves old knowledge without requiring access to past samples. To enable effective collaboration between the two branches, we design the Slow-Fast Dynamic Re-parameterization (SFDR) mechanism for adaptive fusion, and the Knowledge Reflection Mechanism (KRM), which mitigates forgetting and task-recency bias via pseudo-feature generation and dual-level knowledge distillation. Extensive experiments on UCF101, HMDB51, and Something-Something V2 demonstrate that SFCL achieves superior performance compared to existing replay-based methods, despite being exemplar-free. Notably, in long-duration continual learning scenarios, SFCL exhibits remarkable robustness, achieving up to a 30.39\% improvement in accuracy over baselines while maintaining a low forgetting rate, highlighting its scalability and effectiveness in real-world video recognition tasks.
Xueyi Zhang 0001, Siqi Cai 0002, Mingrui Lao, Yanming Guo, Huiping Zhuang
AAAI1
2026 Explainable spatial-temporal-spectral meta-learning for subject-independent EEG-based emotion recognition
Siqi Cai 0002, Xueyi Zhang 0001, Xian Tang
Expert Syst. Appl.3
2026 Towards few-shot deepfake detection with an enhanced CLIP model
Yumin Yang, Xueyi Zhang 0001, Yanming Guo
Neural Networks2
2026 Closed-loop correction reprogramming for fine-grained visual prompting
Xueyi Zhang 0001, Siqi Cai 0002, Mingrui Lao, Haizhou Li 0001
Neural Networks1
2026 PAL: Prompting analytic learning with missing modality for multi-modal class-incremental learning
Xianghu Yue, Yiming Chen 0010, Xueyi Zhang 0001, Xiaoxue Gao, Mengling Feng, Mingrui Lao, Huiping Zhuang, Haizhou Li 0001
Pattern Recognit.3
2025 HeadTile: A Scalable and Efficient Accelerator for Large Language Model Inference with 3D Memory Integration
abstract
Large Language Models (LLMs) inference have gained popularity over the past two years, driven by their high performance achieved through rapid increases in the number of parameters. The inference process of LLMs consists of two distinct stages: prefill and decode, each with unique computational characteristics. While existing neural network inference platforms, such as Google's TPU, perform well during the prefill stage, they often suffer from poor resource utilization during the decode stage. To address this challenge, we propose the scalable Headtile architecture, specifically designed to improve hardware resource utilization. By analyzing the inference behavior of LLMs, we examine how each layer executes on TPUv3 and introduce the Maarg paradigm in Headtile for inter-layer scheduling and mapping. Experimental results show that Headtile can achieve up to 24 × higher throughput in the decode stage compared to TPUv3. In addition, the Maarg paradigm reduces memory accesses by up to 60 % during the prefill stage.
Qingshan Xue, Yihao Shi, Shengbai Luo, Xueyi Zhang 0001, Sheng Ma
HPCC5
2025 Generalization-Preserved Learning: Closing the Backdoor to Catastrophic Forgetting in Continual Deepfake Detection
Xueyi Zhang 0001, Peiyin Zhu, Zhiyuan Yan 0002, Jikang Cheng, Mingrui Lao, Siqi Cai 0002, Yanming Guo
ICCV1
2025 Choose Your Expert: Uncertainty-Guided Expert Selection for Continual Deepfake Detection
abstract
The rapid evolution of deepfake techniques presents dual challenges for detection models: adapting to continuously shifting attack distributions while retaining previously learned knowledge. Although recent continual deepfake detection methods have made progress, they often rely on replay-based training, which limits scalability and deployment. Meanwhile, the task structure of deepfake detection offers a unique opportunity that remains under-explored: it is inherently a binary classification problem with a fixed label space, where the main difficulty lies in distributional drift rather than class expansion. This insight enables the modeling of each incremental distribution shift as a dedicated expert, focusing on specific forgery patterns. To this end, we propose a novel analytically driven, replay-free continual detection framework that eliminates the need for iterative gradient updates. In this framework, task-specific experts are constructed via closed-form ridge regression, requiring only a single forward pass and ensuring non-interference with previous tasks. To enhance the model's capacity for fine-grained forgery recognition, we introduce a lightweight Forgery-Aware Residual Enhancer (FARE). At inference, an Uncertainty-Guided Expert Selection module (UGES) dynamically routes each sample to the most confident expert, which does not require prior knowledge of the attack type. The proposed framework achieves a favorable trade-off between efficiency, privacy, and generalization. It achieves state-of-the-art performance across four benchmark datasets, with an average accuracy of 91.82% and only 1.78% forgetting. Notably, it improves cross-forgery generalization by 9.28% on unseen forgery types, demonstrating strong generalization.
Xueyi Zhang 0001, Peiyin Zhu, Jinping Sui, Xiaoda Yang, Mingrui Lao, Siqi Cai 0002, Yanming Guo, Jun Tang 0001
ACM Multimedia1
2025 EventLip: Enhancing Event-Based Lip Reading via Frequency-Aware Spatiotemporal Hypergraph Modeling
abstract
Event cameras, with their microsecond-level temporal resolution and sparse visual encoding, provide a transformative paradigm for automatic lip reading (ALR). However, event data inherently lack explicit spatial structure and exhibit a pronounced frequency-domain bias. The low-frequency components fail to capture crucial lip structural information, which fundamentally impedes the modeling of intra-frame topological dependencies and inter-frame semantic evolution-both of which are critical for robust lip reading. To this end, we propose FAST-HG, a Frequency-Aware SpatioTemporal HyperGraph framework specifically designed for event-based lip reading. First, we apply low-frequency perturbation to improve the model's robustness for capturing discriminative features, and integrate adaptive high-frequency filtering to enhance edge-aware representations. Then, we construct a Spatial Region Hypergraph (SRH) and a Temporal Semantic Hypergraph (TSH). The former captures intra-frame topological dependencies among lip regions, while the latter explicitly models inter-frame structural associations throughout the lip movement process, enabling the model to capture discriminative patterns in lip dynamics. Furthermore, we propose a viseme-aware label smoothing strategy, where a novel viseme-level edit distance is designed to quantify visual similarities between classes and guide the construction of soft labels. FAST-HG achieves 79.85% and 84.03% accuracy on the DVS-Lip and DVS-LRW100 datasets, respectively, significantly outperforming prior methods and establishing a new benchmark for event-based lip reading.
Xueyi Zhang 0001, Jialu Sun, Xianghu Yue, Tianfang Xiao, Siqi Cai 0002, Mingrui Lao, Haizhou Li 0001
ACM Multimedia1
2025 TrustCLIP: Learning from Noisy Labels via Semantic Label Verification and Trust-aligned Gradient Projection
abstract
Prompt learning has emerged as an efficient adaptation paradigm for vision-language models (VLMs), yet it remains highly vulnerable to label noise, which limits its real-world applicability. We propose TrustCLIP, a noise-robust prompt tuning framework that leverages the inherent semantic structure of CLIP through two key components: Semantic Label Verification (SLV) and Trust-aligned Gradient Projection (TGP). SLV defines a semantic trust boundary based on CLIP's zero-shot predictions to identify reliable samples for standard supervised training. For uncertain samples, TGP projects their gradients into a trust-aligned subspace constructed from the gradients of clean samples, thereby preserving semantically aligned learning signals while suppressing noise-induced optimization drift. Unlike prior approaches, TrustCLIP doesn't require additional parameters, loss reweighting, or uncertainty estimation. Extensive experiments on 7 benchmark datasets with both synthetic and real-world noisy labels demonstrate that TrustCLIP consistently outperforms state-of-the-art methods in terms of both robustness and transferability.
Xueyi Zhang 0001, Peiyin Zhu, Mingrui Lao, Siqi Cai 0002, Yanming Guo, Haizhou Li 0001
ACM Multimedia1
2025 Listening to the Brain: Multi-Band sEEG Auditory Reconstruction via Dynamic Spatio-Temporal Hypergraphs
abstract
Speech is a fundamental form of human communication, and speech perception constitutes the initial stage of language comprehension. Although brain-to-speech interface technologies have made significant progress in recent years, most existing studies focus on neural decoding during speech production. Such approaches heavily rely on articulatory motor regions, rendering them unsuitable for individuals with speech motor impairments, such as those with aphasia or locked-in syndrome. To address this limitation, we construct and release NeuroListen, the first publicly available stereo-electroencephalography (sEEG) dataset specifically designed for auditory reconstruction. It contains over 10 hours of neural–speech paired recordings from 5 clinical participants, covering a wide range of semantic categories. Building on this dataset, we propose HyperSpeech, a multi-band neural decoding framework that employs dynamic spatio-temporal hypergraph neural networks to capture high-order dependencies across frequency, spatial, and temporal dimensions. Experimental results demonstrate that HyperSpeech significantly outperforms existing methods across multiple objective speech quality metrics, and achieves superior performance in human subjective evaluations, validating its effectiveness and advancement. This study provides a dedicated dataset and modeling framework for auditory speech decoding, offering foundations for neural language processing and assistive communication systems.
Xueyi Zhang 0001, Ruicong Wang, Jialu Sun, Siqi Cai 0002, Haizhou Li 0001
NeurIPS1
2025 Boosting Discriminability for Robust Multimodal Entity Linking with Visual Modality Missing
abstract
Multimodal Entity Linking (MEL) aims to retrieve ambiguous mentions within multimodal contexts to the referent entities in a multimodal knowledge base, typically based on the assumption of modality completeness. However, when deployed in open-world applications, MEL systems may encounter uncertainly missing of visual modalities from user-proposed mentions. In this paper, we propose a novel setting dubbed MEL-MM to simulate the practical challenge, and reveal that the semantic discriminability is a crucial factor to enhance the anti-missingness resilience. To this end, we introduce an innovative yet efficient approach termed Cross-View Introspective Ranking Distillation (CVIRD), which seeks to sufficiently align the linking similarities between teacher and student models trained from modality-complete and incomplete data. To be specific, as the first concept in CVIRD, Missing-Aware Ranking Distillation (MARD) focuses on modeling the discriminability by formulating the similarity rankings between mention and entities in a missing-sensitive and differentiable manner. Moreover, the second concept of Cross-View Distillation with Introspection (CVDI) aims to improve discriminability extraction in MARD through multi-level distillation, considering both cross-view retrieval and self-consistency. Experiments verify the effectiveness and model-agnostic ability of our method, which achieves superior performance in contrast to competitive missingness-resilient strategies.
Mingrui Lao, Yanming Guo, Xueyi Zhang 0001, Siqi Cai 0002, Zhaoyun Ding, Haizhou Li 0001
SIGIR4
2024 Sparm: A Sparse Matrix Multiplication Accelerator Supporting Multiple Dataflows
abstract
As the main workload of many scientific and machine learning applications, sparse matrix-matrix multiplication (spGEMM) has become a hot research field. The current spG EMM workloads exhibit sparsity and irregularity, leading to computational inefficiencies on traditional hardware platforms and motivating numerous customized accelerators. These spe-cialized hardware designs typically accelerate only one type of spGEMM dataflow (such as Inner-Product, Outer-Product, or Gustavson), yet the computational efficiency of the same spG EMM kernel can vary significantly under different dataflows. Flexagon is the first spGEMM accelerator to support multiple dataflows, but its MRN (Merger-Reduction Network) design causes a lot of data blocking and load imbalance, which limits its performance. In this work, we propose Sparm, which achieves efficient merging of psums (partial sums) for different dataflows through a specialized indexing unit. Sparm addresses the performance bottlenecks encountered by Flexagon when facing highly sparse matrices. Furthermore, Sparm employs a row/column prefetcher to load the streaming matrices proactively and thus significantly reduces the amount of DRAM access. We conduct simulations using a cycle-accurate simulator on workloads from various application domains, and the results demonstrate that Sparm achieves average performance gains of 2.62× and 1.35× compared to state-of-the-art spGEMM accelerators SIGMA and Flexagon. Meanwhile, Sparm brings only a small amount of additional hardware overhead over Flexagon.
Shengbai Luo, Yihao Shi, Xueyi Zhang 0001, Qingshan Xue, Sheng Ma
ASAP4
2024 Understanding and Mitigating the Soft Error of Contrastive Language-Image Pre-training Models
abstract
In recent years, MultiModal Large Language Models (MM-LLMs), based on the Contrastive Language-Image Pretraining models (CLIP), have achieved the best results in many fields. CLIP breaks through the gaps between language models and image models, realizes zero-shot image classification, and achieves excellent performance in tasks such as text-to-image generation, image style transformation, and long video generation. However, there are few studies on the fault tolerance of CLIP with soft errors, which hinders the application of multimodal large models in the field of security. Based on the analysis of the fault tolerance of common multimodal large models, we proposes a soft error mitigation framework. According to the experiments in this paper, the framework can effectively detect soft errors and mitigate the errors.
Yihao Shi, Shengbai Luo, Qingshan Xue, Xueyi Zhang 0001, Sheng Ma
ITC-Asia5
2024 MMAL: Multi-Modal Analytic Learning for Exemplar-Free Audio-Visual Class Incremental Tasks
abstract
Class-incremental learning poses a significant challenge under an exemplar-free constraint, leading to catastrophic forgetting and sub-par incremental accuracy. Previous attempts have focused primarily on single-modality tasks, such as image classification or audio event classification. However, in the context of Audio-Visual Class-Incremental Learning (AVCIL), the effective integration and utilization of heterogeneous modalities, with their complementary and enhancing characteristics, remains largely unexplored. To bridge this gap, we propose the Multi-Modal Analytic Learning (MMAL) framework, an exemplar-free solution for AVCIL that employs a closed-form, linear approach. To be specific, MMAL introduces a modality fusion module that re-formulates the AVCIL problem through a Recursive Least-Square (RLS) perspective. Complementing this, a Modality-Specific Knowledge Compensation (MSKC) module is designed to further alleviate the under-fitting limitation intrinsic to analytic learning by harnessing individual knowledge from audio and visual modality in tandem. Comprehensive experimental comparisons with existing methods show that our proposed MMAL demonstrates superior performance with the accuracy of 76.71%, 78.98%, and 76.19% on AVE, Kinetics-Sounds, and VGGSounds100 datasets, respectively, setting new state-of-the-art AVCIL performance. Notably, compared to those memory-based methods, our MMAL, being an exemplar-free approach, provides good data privacy and can better leverage multi-modal information for improved incremental accuracy.
Xianghu Yue, Xueyi Zhang 0001, Yiming Chen 0010, Mingrui Lao, Huiping Zhuang, Xinyuan Qian 0001, Haizhou Li 0001
ACM Multimedia2
2024 PD-Refiner: An Underlying Surface Inheritance Refiner with Adaptive Edge-Aware Supervision for Point Cloud Denoising
abstract
Point clouds from real-world scenarios inevitably contain complex noise, significantly impairing the accuracy of downstream tasks. To tackle this challenge, cascading encoder-decoder architecture has become a conventional technical route to iterative denoise. However, circularly feeding the output of denoiser as its input again involves the re-extraction of underlying surface, leading to unstable denoising process and over-smoothed geometric details. To address these issues, we propose a novel denoising paradigm dubbed PD-Refiner that employs a single encoder to model the underlying surface. Then, we leverage several lightweight hierarchical Underlying Surface Inheritance Refiners (USIRs) to inherit and strengthen it, thereby avoiding the re-extraction from the intermediate point cloud. Furthermore, we design adaptive edge-aware supervision to improve the edge awareness of the USIRs, allowing for the adjustment of the denoising preferences from global structure to local details. The results demonstrate that our method not only achieves state-of-the-art performance in terms of denoising stability and efficacy, but also enhances edge clarity and point cloud uniformity.
Xueyi Zhang 0001, Xianghu Yue, Mingrui Lao, Tao Jiang 0062, Fubo Zhang, Longyong Chen
ACM Multimedia2
2024 Language Without Borders: A Dataset and Benchmark for Code-Switching Lip Reading
abstract
Lip reading aims at transforming the videos of continuous lip movement into textual contents, and has achieved significant progress over the past decade. It serves as a critical yet practical assistance for speech-impaired individuals, with more practicability than speech recognition in noisy environments. With the increasing interpersonal communications in social media owing to globalization, the existing monolingual datasets for lip reading may not be sufficient to meet the exponential proliferation of bilingual and even multilingual users. However, to our best knowledge, research on code-switching is only explored in speech recognition, while the attempts in lip reading are seriously neglected. To bridge this gap, we have collected a bilingual code-switching lip reading benchmark composed of Chinese and English, dubbed CSLR. As the pioneering work, we recruited 62 speakers with proficient foundations in bothspoken Chinese and English to express sentences containing both involved languages. Through rigorous criteria in data selection, CSLR benchmark has accumulated 85,560 video samples with a resolution of 1080x1920, totaling over 71.3 hours of high-quality code-switching lip movement data. To systematically evaluate the technical challenges in CSLR, we implement commonly-used lip reading backbones, as well as competitive solutions in code-switching speech for benchmark testing. Experiments show CSLR to be a challenging and under-explored lip reading task. We hope our proposed benchmark will extend the applicability of code-switching lip reading, and further contribute to the communities of cross-lingual communication and collaboration. Our dataset and benchmark are accessible at https://github.com/cslr-lipreading/CSLR.
Xueyi Zhang 0001, Mingrui Lao, Jun Tang 0001, Yanming Guo, Siqi Cai 0002, Xianghu Yue, Haizhou Li 0001
NeurIPS1
2024 Audio-Visual Kinship Verification: A New Dataset and a Unified Adaptive Adversarial Multimodal Learning Approach
abstract
Facial kinship verification refers to automatically determining whether two people have a kin relation from their faces. It has become a popular research topic due to potential practical applications. Over the past decade, many efforts have been devoted to improving the verification performance from human faces only while lacking other biometric information, for example, speaking voice. In this article, to interpret and benefit from multiple modalities, we propose for the first time to combine human faces and voices to verify kinship, which we refer it as the audio-visual kinship verification study. We first establish a comprehensive audio-visual kinship dataset that consists of familial talking facial videos under various scenarios, called TALKIN-Family. Based on the dataset, we present the extensive evaluation of kinship verification from faces and voices. In particular, we propose a deep-learning-based fusion method, called unified adaptive adversarial multimodal learning (UAAML). It consists of the adversarial network and the attention module on the basis of unified multimodal features. Experiments show that audio (voice) information is complementary to facial features and useful for the kinship verification problem. Furthermore, the proposed fusion method outperforms baseline methods. In addition, we also evaluate the human verification ability on a subset of TALKIN-Family. It indicates that humans have higher accuracy when they have access to both faces and voices. The machine-learning methods could effectively and efficiently outperform the human ability. Finally, we include the future work and research opportunities with the TALKIN-Family dataset.
Xiaoting Wu, Xueyi Zhang 0001, Xiaoyi Feng, Miguel Bordallo López, Li Liu 0002
IEEE Trans. Cybern.2
2023 Slow-Fast Time Parameter Aggregation Network for Class-Incremental Lip Reading
abstract
Class incremental learning has yet to be explored in the field of lip-reading, which can circumvent data privacy issues and avoid the high training costs associated with joint training. In this paper, we introduce a benchmark for Class-Incremental Lip-Reading (CILR). To simultaneously improve the plasticity for new classes and stability for old classes in incremental learning, we propose a Slow-Fast Time Parameter Aggregation Network (TPAN) that decouples representation learning of new and old knowledge, taking into account the task characteristics of lip-reading. The TPAN comprises two dynamically evolving branches: one that uses fast gradient descent and the other employs slow momentum updates to retain old knowledge while adapting to new knowledge. Additionally, to achieve efficient knowledge transfer of the incremental model, we design a Hybrid Sequence-Distribution Distillation (HSDD) strategy to transfer knowledge in temporal feature view and classification probability view. We present a comprehensive comparison of the proposed method and previous state-of-the-art class incremental learning methods on the most commonly used lip-reading datasets LRW and LRW1000. The experimental result show that the proposed method can reduce the effect of catastrophic forgetting and improve the incremental accuracy.
Xueyi Zhang 0001, Tao Wang 0074, Jun Tang 0001, Songyang Lao, Haizhou Li 0001
ACM Multimedia1
2022 Boosting Lip Reading with a Multi-View Fusion Network
abstract
Lip reading aims to decode speech information by analyzing lip movement without involving audio. Numerous deep learning based methods are proposed to address this task. Generally, most existing methods extract visual features only based on the lip appearance, while ignoring the shape dynamic information of the lip region. Motivated by this, we propose a Multi-View Fusion Network (MVFN), which can extract more discriminative visual representations by incorporating appearance and shape information. Besides, a novel adaptive graph convolutional network model called Adaptive Spatial Graph Model(ASGM) is proposed to learn lip spatial topology and lip shape dynamics automatically. Experiments on LRW (word-level) and OuluVS2 (phrase-level) clearly show that the proposed method significantly outperforms the baseline methods by a large margin and achieves state-of-the-art performance.
Xueyi Zhang 0001, Jinping Sui, Changchong Sheng, Wanxia Deng, Li Liu 0002
ICME1