Yixuan Zhou 0001

dblp:281/6539-1 · DBLP profile ↗
← Back
12ranked-venue papers
8as first author
12since 2021 · last 2025
0000-0003-1397-9396ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 9 · 5 first-author · 9 since 2021Artificial intelligence and machine learning · 4 · 4 first-author · 4 since 2021
YearPublicationVenuePosition
2025 TAU-106K: A New Dataset for Comprehensive Understanding of Traffic Accident
abstract
Multimodal Large Language Models (MLLMs) have demonstrated impressive performance in general visual understanding tasks. However, their potential for high-level, fine-grained comprehension, such as anomaly understanding, remains unexplored. Focusing on traffic accidents, a critical and practical scenario within anomaly understanding, we investigate the advanced capabilities of MLLMs and propose TABot, a multimodal MLLM specialized for accident-related tasks. To facilitate this, we first construct TAU-106K, a large-scale multimodal dataset containing 106K traffic accident videos and images collected from academic benchmarks and public platforms. The dataset is meticulously annotated through a video-to-image annotation pipeline to ensure comprehensive and high-quality labels. Building upon TAU-106K, we train TABot using a two-step approach designed to integrate multi-granularity tasks, including accident recognition, spatial-temporal grounding, and an auxiliary description task to enhance the model's understanding of accident elements. Extensive experiments demonstrate TABot's superior performance in traffic accident understanding, highlighting not only its capabilities in high-level anomaly comprehension but also the robustness of the TAU-106K benchmark. Our code and data will be available at https://github.com/cool-xuan/TABot.
Yixuan Zhou 0001, Long Bai 0012, Sijia Cai, Bing Deng, Xing Xu 0001, Heng Tao Shen
ICLR1
2025 SaP-Bot: A Multimodal Large-Language Model for End-to-End Same-Product Identification
abstract
Same-product identification serves as a critical infrastructure in e-commerce systems, enabling accurate product matching across heterogeneous marketing representations for key applications such as price comparison and personalized recommendation. Conventional approaches typically depend on manual feature engineering and extensive rule tuning, which limits their adaptability to varying identification criteria across different product categories and inconsistent business scenarios. To overcome these challenges, we propose an end-to-end same-product identification model powered by multimodal large language models (MLLMs) that inherently support multimodal alignment and exhibit strong generalization across diverse real-world settings. We first introduce a novel group-wise annotation pipeline to construct a high-quality dataset, consisting of diverse product pairs with multimodal presentations and labeled at the SKU level. Building on this dataset, we incorporate task-specific training recipes from the perspective of data augmentation, resulting in our SaP-Bot, which demonstrates advanced performance and generalization capabilities. Moreover, we identify a strong correlation between the output logits of MLLMs and the product similarity, enabling interpretable confidence estimation that benefits both data annotation and downstream applications.
Yixuan Zhou 0001, Yulu Tian, Leon Wenliang Zhong, Xingbin Yu, Heng Tao Shen, Xing Xu 0001
ACM Multimedia1
2025 AnoOnly: Semi-supervised anomaly detection with the only loss on anomalies
Yixuan Zhou 0001, Peiyu Yang, Xing Xu 0001, Zhe Sun 0009, Andrzej Cichocki
Expert Syst. Appl.1
2025 VQ-Flow: Taming Normalizing Flows for Multi-Class Anomaly Detection via Hierarchical Vector Quantization
abstract
Normalizing flows, a category of probabilistic models famed for their capabilities in modeling complex data distributions, have exhibited remarkable efficacy in unsupervised anomaly detection. This paper explores the potential of normalizing flows in multi-class anomaly detection, wherein the normal data is compounded with multiple classes without providing class labels. Through the integration of vector quantization (VQ), we empower the flow models to distinguish different concepts of multi-class normal data in an unsupervised manner, resulting in a novel flow-based unified method, named VQ-Flow. Specifically, our VQ-Flow leverages hierarchical vector quantization to estimate two relative codebooks: a Conceptual Prototype Codebook (CPC) for concept distinction and its concomitant Concept-Specific Pattern Codebook (CSPC) to capture concept-specific normal patterns. The flow models in VQ-Flow are conditioned on the concept-specific patterns captured in CSPC, capable of modeling specific normal patterns associated with different concepts. Moreover, CPC further enables our VQ-Flow for concept-aware distribution modeling, faithfully mimicking the intricate multi-class normal distribution through a mixed Gaussian distribution reparametrized on the conceptual prototypes. Through the introduction of vector quantization, the proposed VQ-Flow advances the state-of-the-art in multi-class anomaly detection within a unified training scheme, yielding the Det./Loc. AUROC of 99.5%/98.3% on MVTec AD.
Yixuan Zhou 0001, Xing Xu 0001, Zhe Sun 0009, Jingkuan Song, Andrzej Cichocki, Heng Tao Shen
IEEE Trans. Multim.1
2025 MSFlow: Multiscale Flow-Based Framework for Unsupervised Anomaly Detection
abstract
Unsupervised anomaly detection (UAD) attracts a lot of research interest and drives widespread applications, where only anomaly-free samples are available for training. Some UAD applications intend to locate the anomalous regions further even without any anomaly information. Although the absence of anomalous samples and annotations deteriorates the UAD performance, an inconspicuous, yet powerful statistics model, the normalizing flows, is appropriate for anomaly detection (AD) and localization in an unsupervised fashion. The flow-based probabilistic models, only trained on anomaly-free data, can efficiently distinguish unpredictable anomalies by assigning them much lower likelihoods than normal data. Nevertheless, the size variation of unpredictable anomalies introduces another inconvenience to the flow-based methods for high-precision AD and localization. To generalize the anomaly size variation, we propose a novel multiscale flow-based framework (MSFlow) composed of asymmetrical parallel flows followed by a fusion flow to exchange multiscale perceptions. Moreover, different multiscale aggregation strategies are adopted for image-wise AD and pixel-wise anomaly localization according to the discrepancy between them. The proposed MSFlow is evaluated on three AD datasets, significantly outperforming existing methods. Notably, on the challenging MVTec AD benchmark, our MSFlow achieves a new state-of-the-art (SOTA) with a detection AUORC score of up to 99.7%, localization AUCROC score of 98.8% and PRO score of 97.1%.
Yixuan Zhou 0001, Xing Xu 0001, Jingkuan Song, Fumin Shen, Heng Tao Shen
IEEE Trans. Neural Networks Learn. Syst.1
2024 BatchNorm-Based Weakly Supervised Video Anomaly Detection
abstract
In weakly supervised video anomaly detection (WVAD), where only video-level labels indicating the presence or absence of abnormal events are available, the primary challenge arises from the inherent ambiguity in temporal annotations of abnormal occurrences. Inspired by the statistical insight that temporal features of abnormal events often exhibit outlier characteristics, we propose a novel method, BN-WVAD, which incorporates BatchNorm into WVAD. In the proposed BN-WVAD, we leverage the Divergence of Feature from the Mean vector (DFM) of BatchNorm as a reliable abnormality criterion to discern potential abnormal snippets in abnormal videos. The proposed DFM criterion is also discriminative for anomaly recognition and more resilient to label noise, serving as the additional anomaly score to amend the prediction of the anomaly classifier that is susceptible to noisy labels. Moreover, a batch-level selection strategy is devised to filter more abnormal snippets in videos where more abnormal events occur. The proposed BN-WVAD model demonstrates state-of-the-art performance on UCF-Crime with an AUC of 87.24%, and XD-Violence, where AP reaches up to 84.93%. Our code implementation is accessible athttps://github.com/cool-xuan/BN-WVAD.
Yixuan Zhou 0001, Xing Xu 0001, Fumin Shen, Jingkuan Song, Heng Tao Shen
IEEE Trans. Circuits Syst. Video Technol.1
2023 ImbSAM: A Closer Look at Sharpness-Aware Minimization in Class-Imbalanced Recognition
abstract
Class imbalance is a common challenge in real-world recognition tasks, where the majority of classes have few samples, also known as tail classes. We address this challenge with the perspective of generalization and empirically find that the promising Sharpness-Aware Minimization (SAM) fails to address generalization issues under the class-imbalanced setting. Through investigating this specific type of task, we identify that its generalization bottleneck primarily lies in the severe overfitting for tail classes with limited training data. To overcome this bottleneck, we leverage class priors to restrict the generalization scope of the class-agnostic SAM and propose a class-aware smoothness optimization algorithm named Imbalanced-SAM (ImbSAM). With the guidance of class priors, our ImbSAM specifically improves generalization targeting tail classes. We also verify the efficacy of ImbSAM on two prototypical applications of class-imbalanced recognition: long-tailed classification and semi-supervised anomaly detection, where our ImbSAM demonstrates remarkable performance improvements for tail classes and anomaly. Our code implementation is available at https://github.com/cool-xuan/Imbalanced_SAM.
Yixuan Zhou 0001, Xing Xu 0001, Heng Tao Shen
ICCV1
2023 Region-Aware Semantic Consistency for Unsupervised Domain-Adaptive Semantic Segmentation
abstract
As acquiring pixel-wise labels for semantic segmentation is labor-intensive, unsupervised domain adaptation (UDA) techniques aim to transfer knowledge from synthetic data to real-scene data. To overcome the distribution misalignment between the source domain and the target domain, Teacher-Student (TS) methods are widely-used and promising. In TS methods, the student resorts to the one-hot pseudo labels generated by the teacher. However, the generated one-hot pseudo labels are dubious and ignore the semantic correlation among classes. Besides, in the same position of the same image, the output distributions between the student and the teacher should be consistent. Such prediction consistency is defined as Region-Aware Semantic Consistency (RASC). Correspondingly, we propose an RASC module to assimilate the output distributions of the teacher and the student. Our RASC module is flexible and easily plugged into TS state-of-the-arts (SOTAs) based on either CNNs or Transformers.
Yixuan Zhou 0001, Xing Xu 0001, Guoqing Wang 0001, Fumin Shen, Yang Yang 0002
ICME2
2023 Towards Boosting Black-Box Attack Via Sharpness-Aware
abstract
For black-box attacks, we utilize the transferability of adversarial examples to attack the unseen model successfully. However, existing attack algorithms are easily trapped into a sharp maximum, where a small perturbation changes its value significantly, leading to the failure of the attack. To tackle this issue, we propose a novel Sharpness-Aware Attack seeking the adversarial example with a flat maximum. Specifically, we sample several poor adversarial examples from the neighborhoods of the current point and alleviate the sharpness between them. We also sample anticipatory neighborhoods examples to make the attack algorithm converge quickly to an excellent starting point. Extensive experiments on the ImageNet dataset show the effectiveness of our method, combined with existing gradient-based attacks, our method yields an average attack success rate of 70.0% for nine defense models.
Shengming Yuan, Jingkuan Song, Yixuan Zhou 0001, Yulan He 0001
ICME4
2022 X-HRNet: Towards Lightweight Human Pose Estimation with Spatially Unidimensional Self-Attention
abstract
High-resolution representation is necessary for human pose estimation to achieve high performance, and the ensuing problem is high computational complexity. In particular, predominant pose estimation methods estimate human joints by 2D single-peak heatmaps. Each 2D heatmap can be hori-zontally and vertically projected to and reconstructed by a pair of 1D heat vectors. Inspired by this observation, we introduce a lightweight and powerful alternative, Spatially Unidimensional Self-Attention (SUSA), to the pointwise (1 x 1) convolution that is the main computational bottleneck in the depthwise separable 3 x 3 convolution. Our SUSA reduces the computational complexity of the pointwise (1 x 1) convolution by 96% without sacrificing accuracy. Furthermore, we use the SUSA as the main module to build our lightweight pose estimation backbone X-HRNet, where$X$represents the estimated cross-shape attention vectors. Extensive experiments on the COCO benchmark demonstrate the superiority of our X-HRNet, and comprehensive ablation studies show the effectiveness of the SUSA modules. The code is publicly available at https://github.com/cool-xuan/x-hrnet.
Yixuan Zhou 0001, Xuanhan Wang, Xing Xu 0001, Lei Zhao 0017, Jingkuan Song
ICME1
2022 KTN: Knowledge Transfer Network for Learning Multiperson 2D-3D Correspondences
abstract
Human densepose estimation, aiming at establishing dense correspondences between 2D pixels of human body and 3D human body template, is a key technique in enabling machines to have an understanding of people in images. It still poses several challenges due to practical scenarios where real-world scenes are complex and only partial annotations are available, leading to incompelete or false estimations. In this work, we present a novel framework to detect the densepose of multiple people in an image. The proposed method, which we refer to Knowledge Transfer Network (KTN), tackles two main problems: 1) how to refine image representation for alleviating incomplete estimations, and 2) how to reduce false estimation caused by the low-quality training labels (i.e., limited annotations and class-imbalance labels). Unlike existing works directly propagating the pyramidal features of regions for densepose estimation, the KTN uses a refinement of pyramidal representation, where it simultaneously maintains feature resolution and suppresses background pixels, and this strategy results in a substantial increase in accuracy. Moreover, the KTN enhances the ability of 3D based body parsing with external knowledges, where it casts 2D based body parsers trained from sufficient annotations as a 3D based body parser through a structural body knowledge graph. In this way, it significantly reduces the adverse effects caused by the low-quality annotations. The effectiveness of KTN is demonstrated by its superior performance to the state-of-the-art methods on DensePose-COCO dataset. Extensive ablation studies and experimental results on representative tasks (e.g., human body segmentation, human part segmentation and keypoints detection) and two popular densepose estimation pipelines (i.e., RCNN and fully-convolutional frameworks), further indicate the generalizability of the proposed method.
Xuanhan Wang, Lianli Gao, Yixuan Zhou 0001, Jingkuan Song, Meng Wang 0001
IEEE Trans. Circuits Syst. Video Technol.3
2021 Semantic-aware Transfer with Instance-adaptive Parsing for Crowded Scenes Pose Estimation
abstract
Crowded scenes human pose estimation remains challenging, which requires joint comprehension of multi-persons and their keypoints in a highly complex scenario. The top-down mechanism, which is a detect-then-estimate pipeline, has become the mainstream solution for general pose estimation and obtained impressive progress. However, simply applying this mechanism to crowded scenes pose estimation results in unsatisfactory performance due to several issues, in particular involving missing keypoints in crowds and ambiguously labeling during training. To tackle above two issues, we introduce a novel method named Semantic-aware Transfer with Instance-adaptive Parsing (STIP). Specifically, our STIP first enhances the discriminative power of pixel-level representations with a semantic-aware mechanism, where it smartly decides which pixels to enhance and what semantic embeddings to add. In this way, the missing keypoints detection can be alleviated.Secondly, instead of adopting a standard regressor with fixed parameters, we propose a new instance-adaptive parsing method, where it dynamically generates instance-specific parameters for reducing adverse effects caused by ambiguously labeling. Notably, STIP is designed in a plugin fashion and it can be integrated into any top-down models, such as HRNet. Extensive experiments on two challenging benchmarks, i.e., CrowdPose and MS-COCO, demonstrate the superiority and generalizability of our approach.
Xuanhan Wang, Lianli Gao, Yan Dai 0001, Yixuan Zhou 0001, Jingkuan Song
ACM Multimedia4