Xuan Tong

dblp:347/0254 · DBLP profile ↗
← Back
14ranked-venue papers
2as first author
14since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 2 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 1 first-author · 9 since 2021Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
YearPublicationVenuePosition
2026 Commonality in Few: Few-Shot Multimodal Anomaly Detection via Hypergraph-Enhanced Memory
abstract
Few-shot multimodal industrial anomaly detection is a critical yet underexplored task, offering the ability to quickly adapt to complex industrial scenarios. In few-shot settings, insufficient training samples often fail to cover the diverse patterns present in test samples. This challenge can be mitigated by extracting structural commonality from a small number of training samples. In this paper, we propose a novel few-shot unsupervised multimodal industrial anomaly detection method based on structural commonality, CIF (Commonality In Few). To extract intra-class structural information, we employ hypergraphs, which are capable of modeling higher-order correlations, to capture the structural commonality within training samples, and use a memory bank to store this intra-class structural prior. Firstly, we design a semantic-aware hypergraph construction module tailored for single-semantic industrial images, from which we extract common structures to guide the construction of the memory bank. Secondly, we use a training-free hypergraph message passing module to update the visual features of test samples, reducing the distribution gap between test features and features in the memory bank. We further propose a hyperedge-guided memory search module, which utilizes structural information to assist the memory search process and reduce the false positive rate. Experimental results on the MVTec 3D-AD dataset and the Eyecandies dataset show that our method outperforms the state-of-the-art (SOTA) methods in few-shot settings.
Yuxuan Lin 0001, Hanjing Yan, Xuan Tong, Yang Chang, Huanzhen Wang, Ziheng Zhou 0005, Shuyong Gao, Yan Wang 0068
AAAI3
2026 CADiff: Context-Aware Diffusion for Controllable Anomaly Generation in Anomaly Detection
abstract
Generating anomalies is a crucial method to enhance detection and classification performance by expanding anomalous data repository. However, existing anomaly generation methods overlook the intrinsic entanglement between diverse anomaly types and product structures, leading to semantic ambiguity. We propose CADiff, a context-aware generation framework that reframes anomalies as compositional perturbations. Firstly, we propose Context-aware Text Prompt (CTP), a mechanism which contains multiple tokens that characterize anomalies and products separately to enhance the contextual consistency of generated images and refine the local variability of anomalies. Secondly, we develop Self-adaptive Spatial Control (SSC), a self-adaptive interaction design that mitigates anomaly leakage or missing phenomena. Thirdly, we introduce Intensity-controllable Attention Re-weighting (IAR), an inference scheduling scheme with the ability to amplify or attenuate abnormal semantic effects to improve generation diversity. Extensive experiments on MVTec AD and VisA datasets demonstrate the superiority of our proposed method over state-of-the-art methods in both realism and diversity of the generated results, and significantly improve the performance of downstream tasks, including anomaly detection, anomaly localization, and anomaly classification tasks.
Xuan Tong, Yuxuan Lin 0001, Junxiong Lin, Xinji Mai, Haoran Wang 0006, Zeng Tao
AAAI1
2026 Hi-EF: Benchmarking Emotion Forecasting in Human-interaction
abstract
Affective Forecasting is an psychology task that involves predicting an individual's future emotional responses, often hampered by reliance on external factors leading to inaccuracies, and typically remains at a qualitative analysis stage. To address these challenges, we narrows the scope of Affective Forecasting by introducing the concept of Human-interaction-based Emotion Forecasting (EF). This task is set within the context of a two-party interaction, positing that an individual's emotions are significantly influenced by their interaction partner's emotional expressions and informational cues. This dynamic provides a structured perspective for exploring the patterns of emotional change, thereby enhancing the feasibility of emotion forecasting.
Haoran Wang 0006, Xinji Mai, Zeng Tao, Junxiong Lin, Xuan Tong, Ivy Pan, Shaoqi Yan, Yan Wang 0068, Shuyong Gao
AAAI5
2025 OUS: Bridging Scene Context and Facial Features to Overcome the Rigid Cognitive Problem
abstract
Dynamic Facial Expression Recognition (DFER) is crucial for affective computing but often overlooks the impact of scene context. We have identified a significant issue in current DFER tasks: human annotators typically integrate emotions from various angles, including environmental cues and body language, whereas existing DFER methods tend to consider the scene as noise that needs to be filtered out, focusing solely on facial information. We refer to this as the Rigid Cognitive Problem. The Rigid Cognitive Problem can lead to discrepancies between the cognition of annotators and models in some samples. To align more closely with the human cognitive paradigm of emotions, we propose an Overall Understanding of the Scene DFER method (OUS). OUS effectively integrates scene and facial features, combining scene-specific emotional knowledge for DFER. Extensive experiments on the two largest datasets in the DFER field, DFEW and FERV39k, demonstrate that OUS significantly outperforms existing methods. By analyzing the Rigid Cognitive Problem, OUS successfully understands the complex relationship between scene context and emotional expression, closely aligning with human emotional understanding in real-world scenarios.
Xinji Mai, Haoran Wang 0006, Zeng Tao, Junxiong Lin, Shaoqi Yan, Yan Wang 0068, Jiawen Yu, Xuan Tong
AAAI8
2025 D2SP: Dynamic Dual-Stage Purification Framework for Dual Noise Mitigation in Vision-based Affective Recognition
abstract
The current advancements in Dynamic Facial Expression Recognition (DFER) methods mainly focus on better capturing the spatial and temporal features of facial expressions. However, DFER datasets contain a substantial amount of noisy samples, and few have addressed the issue of handling this noise. We identified two types of noise: one is caused by low-quality data resulting from factors such as occlusion, dim lighting, and blurriness; the other arises from mislabeled data due to annotation bias by annotators. Addressing the two types of noise, we have meticulously crafted a Dynamic Dual-Stage Purification (D2SP) Framework. This initiative aims to dynamically purify the DFER datasets of these two types of noise, ensuring that only high-quality and correctly labeled data is used in the training process. To mitigate low-quality samples, we introduce the Coarse-Grained Pruning (CGP) stage, which computes sample weights and prunes those low-weight samples. After CGP, the Fine-Grained Correction (FGC) stage evaluates prediction stability to correct mislabeled data. Moreover, D2SP is conceived as a general, plug-and-play framework, tailored to integrate seamlessly with prevailing DFER methods. Extensive experiments covering prevalent DFER datasets and deploying multiple benchmark methods have substantiated D2SP’s ability to enhance performance metrics.
Haoran Wang 0006, Xinji Mai, Zeng Tao, Xuan Tong, Junxiong Lin, Yan Wang 0068, Jiawen Yu, Shaoqi Yan, Ziheng Zhou 0005
CVPR4
2025 HSS-IAD: A Heterogeneous Same-Sort Industrial Anomaly Detection Dataset
abstract
Multi-class Unsupervised Anomaly Detection algorithms (MUAD) are receiving increasing attention due to their relatively low deployment costs and improved training efficiency. However, the real-world effectiveness of MUAD methods is questioned due to limitations in current Industrial Anomaly Detection (IAD) datasets. These datasets contain numerous classes that are unlikely to be produced by the same factory and fail to cover multiple structures or appearances. Additionally, the defects do not reflect real-world characteristics. Therefore, we introduce the Heterogeneous Same-Sort Industrial Anomaly Detection (HSS-IAD) dataset, which contains 8,580 images of metallic-like industrial parts and precise anomaly annotations. These parts exhibit variations in structure and appearance, with subtle defects that closely resemble the base materials. We also provide foreground images for synthetic anomaly generation. Finally, we evaluate popular IAD methods on this dataset under multi-class and class-separated settings, demonstrating its potential to bridge the gap between existing datasets and real factory conditions. The dataset is available at https://github.com/Qiqigeww/HSS-IAD-Dataset.
Qishan Wang 0002, Shuyong Gao, Jiawen Yu, Xuan Tong
ICME5
2025 Component-Aware Unsupervised Logical Anomaly Generation for Industrial Anomaly Detection
abstract
Anomaly detection is critical in industrial manufacturing for ensuring product quality and improving efficiency in automated processes. The scarcity of anomalous samples limits traditional detection methods, making anomaly generation essential for expanding the data repository. However, recent generative models often produce unrealistic anomalies increasing false positives, or require real-world anomaly samples for training. In this work, we treat anomaly generation as a compositional problem and propose ComGEN, a component-aware and unsupervised framework that addresses the gap in logical anomaly generation. Our method comprises a multi-component learning strategy to disentangle visual components, followed by subsequent generation editing procedures. Disentangled text-to-component pairs, revealing intrinsic logical constraints, conduct attention-guided residual mapping and model training with iteratively matched references across multiple scales. Experiments on the MVTecLOCO dataset confirm the efficacy of ComGEN, achieving the best AUROC score of$\mathbf{9 1. 2 \%}$. Additional experiments on the real-world scenario of Diesel Engine and widelyused MVTecAD dataset demonstrate significant performance improvements when integrating simulated anomalies generated by ComGEN into automated production workflows.
Xuan Tong, Yang Chang, Qing Zhao 0007, Jiawen Yu, Boyang Wang 0003, Junxiong Lin, Yuxuan Lin 0001, Xinji Mai, Haoran Wang 0006, Zeng Tao, Yan Wang 0068
ICRA1
2025 Noise Fusion-based Distillation Learning for Anomaly Detection in Complex Industrial Environments
abstract
Anomaly detection and localization in automated industrial manufacturing can significantly enhance production efficiency and product quality. Existing methods are capable of detecting surface defects in pre-defined or controlled imaging environments. However, accurately detecting workpiece defects in complex and unstructured industrial environments with varying views, poses and illumination remains challenging. We propose a novel anomaly detection and localization method specifically designed to handle inputs with perturbative patterns. Our approach introduces a new framework based on a collaborative distillation heterogeneous teacher network (HetNet), an adaptive local-global feature fusion module, and a local multivariate Gaussian noise generation module. HetNet can learn to model the complex feature distribution of normal patterns using limited information about local disruptive changes. We conducted extensive experiments on mainstream benchmarks. HetNet demonstrates superior performance with approximately 10% improvement across all evaluation metrics on MSC-AD under industrial conditions, while achieving state-of-the-art results on other datasets, validating its resilience to environmental fluctuations and its capability to enhance the reliability of industrial anomaly detection systems across diverse scenarios. Tests in real-world environments further confirm that HetNet can be effectively integrated into production lines to achieve robust and real-time anomaly detection. Codes, images and videos are published on the project website at: https://zihuatanejoyu.github.io/HetNet/
Jiawen Yu, Jieji Ren, Yang Chang, Qiaojun Yu, Xuan Tong, Boyang Wang 0003, Xinji Mai
IROS5
2025 JailBound: Jailbreaking Internal Safety Boundaries of Vision-Language Models
abstract
Vision-Language Models (VLMs) exhibit impressive performance, yet the integration of powerful vision encoders has significantly broadened their attack surface, rendering them increasingly susceptible to jailbreak attacks. However, lacking well-defined attack objectives, existing jailbreak methods often struggle with gradient-based strategies prone to local optima and lacking precise directional guidance, and typically decouple visual and textual modalities, thereby limiting their effectiveness by neglecting crucial cross-modal interactions. Inspired by the Eliciting Latent Knowledge (ELK) framework, we posit that VLMs encode safety-relevant information within their internal fusion-layer representations, revealing an implicit safety decision boundary in the latent space. This motivates exploiting boundary to steer model behavior. Accordingly, we propose \textbf{JailBound}, a novel latent space jailbreak framework comprising two stages: (1) \textbf{Safety Boundary Probing}, which addresses the guidance issue by approximating decision boundary within fusion layer's latent space, thereby identifying optimal perturbation directions towards the target region; and (2) \textbf{Safety Boundary Crossing}, which overcomes the limitations of decoupled approaches by jointly optimizing adversarial perturbations across both image and text inputs. This latter stage employs an innovative mechanism to steer the model's internal state towards policy-violating outputs while maintaining cross-modal semantic consistency. Extensive experiments on six diverse VLMs demonstrate JailBound's efficacy, achieves 94.32\% white-box and 67.28\% black-box attack success averagely, which are 6.17\% and 21.13\% higher than SOTA methods, respectively. Our findings expose a overlooked safety risk in VLMs and highlight the urgent need for more robust defenses. \textcolor{red}{Warning: This paper contains potentially sensitive, harmful and offensive content.}
Yixu Wang, Jie Li 0052, Xuan Tong, Yan Teng 0002, Xingjun Ma, Yingchun Wang 0004
NeurIPS4
2024 Adaptive Multi-modal Fusion of Spatially Variant Kernel Refinement with Diffusion Model for Blind Image Super-Resolution
Junxiong Lin, Yan Wang 0068, Zeng Tao, Boyang Wang 0003, Qing Zhao 0007, Haorang Wang, Xuan Tong, Xinji Mai, Yuxuan Lin 0001, Wei Song 0007, Jiawen Yu, Shaoqi Yan
ECCV (52)7
2024 Suppressing Uncertainties in Degradation Estimation for Blind Super-Resolution
Junxiong Lin, Zen Tao, Xuan Tong, Xinji Mai, Haoran Wang 0006, Boyang Wang 0003, Yan Wang 0068, Qing Zhao 0007, Jiawen Yu, Yuxuan Lin 0001, Shaoqi Yan, Shuyong Gao
ACM Multimedia3
2024 All rivers run into the sea: Unified Modality Brain-Inspired Emotional Central Mechanism
abstract
In the field of affective computing, fully leveraging information from a variety of sensory modalities is essential for the comprehensive understanding and processing of human emotions. Inspired by the process through which the human brain handles emotions and the theory of cross-modal plasticity, we propose UMBEnet, a brain-like unified modal affective processing network. The primary design of UMBEnet includes a Dual-Stream (DS) structure that fuses inherent prompts with a Prompt Pool and a Sparse Feature Fusion (SFF) module. The design of the Prompt Pool is aimed at integrating information from different modalities, while inherent prompts are intended to enhance the system's predictive guidance capabilities and effectively manage knowledge related to emotion classification. Moreover, considering the sparsity of effective information across different modalities, the SSF module aims to make full use of all available sensory data through the sparse integration of modality fusion prompts and inherent prompts, maintaining high adaptability and sensitivity to complex emotional states. Extensive experiments on the largest benchmark datasets in the Dynamic Facial Expression Recognition (DFER) field, including DFEW, FERV39k, and MAFW, have proven that UMBEnet consistently outperforms the current state-of-the-art methods. Notably, in scenarios of Modality Missingness and multimodal contexts, UMBEnet significantly surpasses the leading current methods, demonstrating outstanding performance and adaptability in tasks that involve complex emotional understanding with rich multimodal information. Code can be obtained at https://github.com/Xinji-Mai/UMBEnet.
Xinji Mai, Junxiong Lin, Haoran Wang 0006, Zeng Tao, Yan Wang 0068, Shaoqi Yan, Xuan Tong, Jiawen Yu, Boyang Wang 0003, Ziheng Zhou 0005, Qing Zhao 0007, Shuyong Gao
ACM Multimedia7
2023 Transfer-Learning-Based Approach to Retrieve the Cloud Properties Using Diverse Remote Sensing Datasets
abstract
Clouds play an important role in the Earth’s climate system; however, various observational methods describe clouds differently, leading to cloud products being described with different characteristics, and affecting our understanding of cloud effects. To address this problem, this study integrates different cloud products into the transfer-learning procedure of a deep learning model and determined the Cloud Effective Radius (CER), Cloud Optical Thickness (COT), and Cloud Top Height (CTH) from Himawari-8 thermal infrared measurements. The retrieval results were independently evaluated against the Moderate-resolution Imaging Spectroradiometer cloud products and further compared with Himawari-8 cloud products during the day. The Root Mean Squared Errors (RMSE) of the model for the CER, COT, and CTH were 4.490 μm, 11.198, and 1.904 km, respectively, which are lower than those of Himawari-8 cloud products (RmSe:11.172 μm, 14.755, and 2.860 km). Moreover, validation results against active sensors show that the model performs slightly better during the day than at night, and both are generally better than the Himawari-8 cloud product. Overall, the model maintains stable performance during both day and night, and its accuracy is higher than that of Himawari-8 cloud products.
Feng Zhang 0041, Xuan Tong, Baoxiang Pan, Jun Li 0026, Husi Letu, Farhan Mustafa
IEEE Trans. Geosci. Remote. Sens.4
2023 Cloud Identification and Properties Retrieval of the Fengyun-4A Satellite Using a ResUnet Model
abstract
The Advanced Geostationary Radiation Imager (AGRI) onboard the Fengyun-4A (FY4A) satellite has good cloud observation ability, but it still absents all-weather and high-precision official cloud products. This study develops a deep-learning ResUnet model for all-weather retrieval of cloud phase (CLP) and cloud properties using the brightness temperature from water vapor and longwave infrared channels of AGRI. The ResUnet model is trained with the Himawari-8 satellite Level-2 (H8-L2) cloud products as true targets, and adopts image-by-image way to learn the spatial structure information of clouds, which compensates for the difficulty of retrieving thick clouds by thermal infrared radiation at night to some extent. On an independent testing dataset, the model has an overall accuracy of 90.64% for CLP identification and performs well at retrieving cloud top height (CTH). Even without using visible and near-infrared radiation, the root mean square error of cloud effective radius (CER) and cloud optical thickness (COT) estimations still reaches 7.14 μm and 9.01 in the range of 0–60. To further illustrate the reliability and applicability, CLP and cloud properties provided by the CALIPSO and MODIS are used as benchmarks to assess the quality of cloud products from FY4A satellite Level-2 (FY4A-L2), H8-L2 and ResUnet model retrieval. The ResUnet model provides a significant improvement over FY4A-L2 for the accuracy of cloud identification and in the quality of CTH products. In the range of 0–40 μm (0–60), the CER (COT) product of ResUnet model retrieval has a reliable and higher precision that is comparable with H8-L2.
Zhijun Zhao, Feng Zhang 0041, Zhengqiang Li, Xuan Tong
IEEE Trans. Geosci. Remote. Sens.5