Haote Xu

dblp:251/4038 · DBLP profile ↗
← Back
9ranked-venue papers
1as first author
9since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Segmentation and scene understanding · 40% Time series and sequential data · 30% Vision and language · 17%
Databases, data mining, and information retrieval
1 paper
Data mining · 100%

Topics — the 13 heaviest of 14, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Data mining
anomaly detection
0.912025
Demeaned Sparse: Efficient Anomaly Detection by Residual Estimate · ICML 2025
Data mining › anomaly detection
anomaly score
0.912025
Demeaned Sparse: Efficient Anomaly Detection by Residual Estimate · ICML 2025
Data mining › anomaly detection
frequency-domain anomaly detection
0.912025
Demeaned Sparse: Efficient Anomaly Detection by Residual Estimate · ICML 2025
Data mining › anomaly detection › deep anomaly detection
reconstruction-based anomaly detection
0.912025
Demeaned Sparse: Efficient Anomaly Detection by Residual Estimate · ICML 2025
Computer vision › Segmentation and scene understanding › medical image segmentation
ambiguous medical image segmentation
0.812024
P2SAM: Probabilistically Prompted SAMs Are Efficient Segmentator for Ambiguous Medical Images · ACM Multimedia 2024
Machine learning › Time series and sequential data
anomaly detection
0.812024
SimCLIP: Refining Image-Text Alignment with Simple Prompts for Zero-/Few-shot Anomaly Detection · ACM Multimedia 2024
Computer vision › Segmentation and scene understanding
medical image segmentation
0.812024
P2SAM: Probabilistically Prompted SAMs Are Efficient Segmentator for Ambiguous Medical Images · ACM Multimedia 2024
Computer vision › Segmentation and scene understanding › image segmentation
probabilistic segmentation
0.812024
P2SAM: Probabilistically Prompted SAMs Are Efficient Segmentator for Ambiguous Medical Images · ACM Multimedia 2024
Natural language and speech › Language models and text generation
prompt tuning
0.812024
SimCLIP: Refining Image-Text Alignment with Simple Prompts for Zero-/Few-shot Anomaly Detection · ACM Multimedia 2024
Computer vision › Vision and language › vision-language model
vision-language model adaptation
0.812024
SimCLIP: Refining Image-Text Alignment with Simple Prompts for Zero-/Few-shot Anomaly Detection · ACM Multimedia 2024
Machine learning › Time series and sequential data › anomaly detection
zero-shot anomaly detection
0.812024
SimCLIP: Refining Image-Text Alignment with Simple Prompts for Zero-/Few-shot Anomaly Detection · ACM Multimedia 2024
Machine learning › Time series and sequential data › anomaly detection
anomaly segmentation
0.212024
SimCLIP: Refining Image-Text Alignment with Simple Prompts for Zero-/Few-shot Anomaly Detection · ACM Multimedia 2024
Computer vision › Vision and language › vision-language model › prompt learning
prompt-based adaptation
0.212024
P2SAM: Probabilistically Prompted SAMs Are Efficient Segmentator for Ambiguous Medical Images · ACM Multimedia 2024

Methods — techniques the papers use, named apart from their topics

fourier transform · 1.7bootstrap · 1.7factor models · 0.9factor model · 0.9vision adapter · 0.8segment anything model · 0.8probabilistic prompt space · 0.8implicit prompt tuning · 0.8contrastive learning · 0.8
YearPublicationVenuePosition
2026 Exploiting point-language models with dual-prompts for 3D anomaly detection
Haote Xu, Xiaolu Chen, Haodi Xu, Yue Huang 0001, Xinghao Ding, Xiaotong Tu
Expert Syst. Appl.2
2025 Demeaned Sparse: Efficient Anomaly Detection by Residual Estimate
abstract
Frequency-domain image anomaly detection methods can substantially enhance anomaly detection performance, however, they still lack an interpretable theoretical framework to guarantee the effectiveness of the detection process. We propose a novel test to detect anomalies in structural image via a Demeaned Fourier transform (DFT) under factor model framework, and we proof its effectiveness. We also briefly give the asymptotic theories of our test, the asymptotic theory explains why the test can detect anomalies at both the image and pixel levels within the theoretical lower bound. Based on our test, we derive a module called Demeaned Fourier Sparse (DFS) that effectively enhances detection performance in unsupervised anomaly detection tasks, which can construct masks in the Fourier domain and utilize a distribution-free sampling method similar to the bootstrap method. The experimental results indicate that this module can accurately and efficiently generate effective masks for reconstruction-based anomaly detection tasks, thereby enhancing the performance of anomaly detection methods and validating the effectiveness of the theoretical framework.
Yifan Fang, Yifei Fang, Ruizhe Chen, Haote Xu, Xinghao Ding, Yue Huang 0001
ICML4
2025 DiRNet: A Domain-Invariant Reconstruction Framework for Unsupervised Anomaly Detection Under Distribution Shift
Haodi Xu, Haote Xu, Guoliang Hu
PRCV (1)3
2025 CVC: Further aligning LLMs via cross-view correction for time series forecasting
Haote Xu, Haodi Xu, Yinhao Liu, Xinghao Ding
Knowl. Based Syst.2
2024 Implicit Foreground-Guided Network for Anomaly Detection and Localization
abstract
Anomaly detection plays an essential role in large-scale industrial manufacturing. However, reconstruction-based anomaly detection methods, as one of the mainstream methods, are prone to incorrectly detecting background noise as anomalous regions. Therefore, inspired by multi-task learning, we propose an Implicit Foreground-guided Network (IFgNet), which consists of a Multi-Task Attention Shared (MTAS) sub-network and a discriminative sub-network. Specifically, the MTAS sub-network implements the foreground detection and reconstruction tasks within the shared network, while the discriminative sub-network performs the final anomaly detection. In the MTAS sub-network, multiple task-specific attention blocks are applied to learn task-specific features while allowing features to be shared between different tasks. Consequently, the features that contain both semantic and edge structure information are learned through the foreground detection task, which also facilitates the reconstruction task. Furthermore, the outputs of foreground detection can be utilized to refine the anomaly detection results. In this way, IFgNet effectively mitigates the influence of background noise and achieves competitive performance on the VisA and BTAD datasets with existing methods.
Xiaolu Chen, Haote Xu, Chenghao Deng, Xiaotong Tu, Xinghao Ding, Yue Huang 0001
ICASSP2
2024 SimCLIP: Refining Image-Text Alignment with Simple Prompts for Zero-/Few-shot Anomaly Detection
abstract
Recently, large pre-trained vision-language models, such as CLIP, have demonstrated significant potential in zero-/few-shot anomaly detection tasks. However, existing methods not only rely on expert knowledge to manually craft extensive text prompts but also suffer from a misalignment of high-level language features with fine-level vision features in anomaly segmentation tasks. In this paper, we propose a method, named SimCLIP, which focuses on refining the aforementioned misalignment problem through bidirectional adaptation of both Multi-Hierarchy Vision Adapter (MHVA) and Implicit Prompt Tuning (IPT). In this way, our approach requires only a simple binary prompt to efficiently accomplish anomaly classification and segmentation tasks in zero-shot scenarios. Furthermore, we introduce its few-shot extension, SimCLIP+, integrating the relational information among vision embeddings and skillfully merging the cross-modal synergy information between vision and language to address downstream anomaly detection tasks. Extensive experiments on two challenging datasets prove the more remarkable generalization capacity of our method compared to the current SOTA approaches. Our code is available at https://github.com/CH-ORGI/SimCLIP.
Chenghao Deng, Haote Xu, Xiaolu Chen, Haodi Xu, Xiaotong Tu, Xinghao Ding, Yue Huang 0001
ACM Multimedia2
2024 P2SAM: Probabilistically Prompted SAMs Are Efficient Segmentator for Ambiguous Medical Images
abstract
Generating diverse plausible outputs from a single input is crucial for addressing visual ambiguities, exemplified in medical imaging where experts may provide varying semantic segmentation annotations for the same image.Existing methods handles ambiguous segmentation relying on probabilistic modeling and extensive multi-output annotated data while often struggles with limited ambiguously labeled datasets common in real-world applications.To surmount the challenge, we propose P²SAM, a novel framework that leverages the Segment Anything Model (SAM)'s prior knowledge for ambiguous object segmentation. By transforming SAM's sensitivity to prompts into an advantage, we introduce a prior probabilistic space for prompts.Experimental results show that P²SAM significantly enhances medical segmentation precision and diversity using minimal ambiguously annotated samples. Benchmarking against state-of-the-art methods demonstrates superior performance with just 5.5% of the training data (+12% Dmax). This approach marks a significant advancement towards deploying probabilistic models in data-limited real-world scenarios.
Yuzhi Huang, Chenxin Li, Zixu Lin, Hengyu Liu 0007, Haote Xu, Yifan Liu 0010, Yue Huang 0001, Xinghao Ding, Xiaotong Tu, Yixuan Yuan
ACM Multimedia5
2024 AFSC: Adaptive Fourier Space Compression for Anomaly Detection
abstract
The primary challenge faced by reconstruction-based anomaly detection (AD) methods is that neural networks exhibit strong generalization, resulting in a high probability and accuracy of anomaly reconstruction. Several existing methods attempt to alleviate this problem by randomly masking partial image regions and reconstructing the image from partial inpaintings. However, local masking in spatial space is not guaranteed to remove anomalous regions during the testing phase and poses the risk of normal regions being inaccurately reconstructed. Hence, we explore an approach to compress the global information of the image while ensuring the loss of partial anomaly information renders it difficult to reconstruct. Inspired by the fact that each Fourier coefficient contains global information of the image, we propose an adaptive Fourier space compression (AFSC) method. Specifically, the Fourier coefficients of the input image are sparsely sampled by binary masks obtained from the AFSC module (AFSCm). In AFSCm, the masks are jointly optimized with the reconstruction network subject to sparsity constraint. The learned masks are forced to selectively retain part of the global information that is favourable to recovering normal images. In addition, we introduce an efficient Fourier convolution module that enables the network to accurately reconstruct normal regions under conditions of losing partial information. Experimental results on three benchmarks of industrial scenarios demonstrate our method (without external prior) achieves competitive results compared with recent methods.
Haote Xu, Xiaolu Chen, Changxing Jing, Liyan Sun, Yue Huang 0001, Xinghao Ding
IEEE Trans. Ind. Informatics1
2022 Unsupervised Anomaly Segmentation for Brain Lesions Using Dual Semantic-Manifold Reconstruction
Zhiyuan Ding, Haote Xu, Chenxin Li, Xinghao Ding, Yue Huang 0001
ICONIP (3)3