EDBT 2026 Demo / reviewers in the wild / expert
Xian Tao
dblp:166/8979
· DBLP profile ↗
17ranked-venue papers
2as first author
16since 2021 · last 2026
0000-0001-5834-5181ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 8 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ONIR: Object-Noted Tagging for Aerial Image Captioning generationabstractAutomated captioning for remote sensing imagery often struggles to balance the high descriptive power of large models with the deployment feasibility of smaller ones. To bridge this gap, this paper introduces ONIR, a LLM-efficient, tag-guided framework that empowers compact language models (1-3B parameters) to achieve state-of-the-art captioning accuracy. Specifically, the proposed approach synthesizes a large-scale pseudo-caption dataset by leveraging GPT-4O on existing segmentation benchmarks. Explicit semantic tags are then extracted to train a multi-label Contrastive Language-Image Pre-Training (CLIP) encoder, providing interpretable visual guidance. To maintain parameter efficiency, the architecture incorporates a simple Multilayer Perceptron (MLP) bridge and a two-stage LoRA fine-tuning strategy. Extensive experiments on standard benchmark dataset, such as UCM and Sydney Captions, demonstrate that ONIR significantly outperforms models up to four times its size (7-13B). By combining superior performance with computational efficiency and tag-based controllability, ONIR offers a highly practical solution for real-world remote sensing applications. Xing Zi, Tengjun Ni, Xianjing Fan, Xian Tao, Xinyi Gong, Jun Li 0010, Ali Braytee, Mukesh Prasad |
J. Vis. Commun. Image Represent. | 4 |
| 2026 | SAM-IAD: Injecting specific knowledge into SAM for industrial anomaly detection
Yichi Chen 0002, Bin Chen 0022, Weizhi Xian, Xinyi Gong, Jianwen Han, Xian Tao |
Knowl. Based Syst. | 7 |
| 2026 | MPFR: Memory prompt feature reconstruction for continual anomaly detection and segmentation
Yichi Chen 0002, Xian Tao, Bin Chen 0022, Pang-jo Chun, Xinmiao Zhou |
Pattern Recognit. | 2 |
| 2025 | Bayesian Prompt Flow Learning for Zero-Shot Anomaly DetectionabstractRecently, vision-language models (e.g. CLIP) have demonstrated remarkable performance in zero-shot anomaly detection (ZSAD). By leveraging auxiliary data during training, these models can directly perform cross-category anomaly detection on target datasets, such as detecting defects on industrial product surfaces or identifying tumors in organ tissues. Existing approaches typically construct text prompts through either manual design or the optimization of learnable prompt vectors. However, these methods face several challenges: 1) handcrafted prompts require extensive expert knowledge and trial-and-error; 2) single-form learnable prompts struggle to capture complex anomaly semantics; and 3) an unconstrained prompt space limits generalization to unseen categories. To address these issues, we propose Bayesian Prompt Flow Learning (Bayes-PFL), which models the prompt space as a learnable probability distribution from a Bayesian perspective. Specifically, a prompt flow module is designed to learn both image-specific and image-agnostic distributions, which are jointly utilized to regularize the text prompt space and improve the model's generalization on unseen categories. These learned distributions are then sampled to generate diverse text prompts, effectively covering the prompt space. Additionally, a residual cross-model attention (RCA) module is introduced to better align dynamic text embeddings with fine-grained image features. Extensive experiments on 15 industrial and medical datasets demonstrate our method's superior performance. The code is available at https://github.com/xiaozhen228/Bayes-PFL. Zhen Qu, Xian Tao, Xinyi Gong, Shichen Qu, Qiyu Chen 0002, Zhengtao Zhang, Guiguang Ding |
CVPR | 2 |
| 2025 | DictAS: A Framework for Class-Generalizable Few-Shot Anomaly Segmentation via Dictionary LookupabstractRecent vision-language models (e.g., CLIP) have demonstrated remarkable class-generalizable ability to unseen classes in few-shot anomaly segmentation (FSAS), leveraging supervised prompt learning or fine-tuning on seen classes. However, their cross-category generalization largely depends on prior knowledge of real seen anomaly samples. In this paper, we propose a novel framework, namely DictAS, which enables a unified model to detect visual anomalies in unseen object categories without any retraining on the target data, only employing a few normal reference images as visual prompts. The insight behind DictAS is to transfer dictionary lookup capabilities to the FSAS task for unseen classes via self-supervised learning, instead of merely memorizing the normal and abnormal feature patterns from the training set. Specifically, DictAS mainly consists of three components: (1) Dictionary Construction - to simulate the index and content of a real dictionary using features from normal reference images. (2) Dictionary Lookup - to retrieve queried region features from the dictionary via a sparse lookup strategy. When a query feature cannot be retrieved, it is classified as an anomaly. (3) Query Discrimination Regularization - to enhance anomaly discrimination by making abnormal features harder to retrieve from the dictionary. To achieve this, Contrastive Query Constraint and Text Alignment Constraint are further proposed. Extensive experiments on seven public industrial and medical datasets demonstrate that DictAS consistently outperforms state-of-the-art FSAS methods. Zhen Qu, Xian Tao, Xinyi Gong, Shichen Qu, Fei Shen 0002, Zhengtao Zhang, Mukesh Prasad, Guiguang Ding |
ICCV | 2 |
| 2025 | RSVLM-QA: A Benchmark Dataset for Remote Sensing Vision Language Model-based Question AnsweringabstractVisual Question Answering (VQA) in remote sensing (RS) is pivotal for interpreting Earth observation data. However, existing RS VQA datasets are constrained by limitations in annotation richness, question diversity, and the assessment of specific reasoning capabilities. This paper introduces Remote Sensing Vision Language Model Question Answering (RSVLM-QA) dataset, a new large-scale, content-rich VQA dataset for the RS domain. RSVLM-QA is constructed by integrating data from several prominent RS segmentation and detection datasets: WHU, LoveDA, INRIA, and iSAID. We employ an innovative dual-track annotation generation pipeline. Firstly, we leverage Large Language Models (LLMs), specifically GPT-4.1, with meticulously designed prompts to automatically generate a suite of detailed annotations including image captions, spatial relations, and semantic tags, alongside complex caption-based VQA pairs. Secondly, to address the challenging task of object counting in RS imagery, we have developed a specialized automated process that extracts object counts directly from the original segmentation data; GPT-4.1 then formulates natural language answers from these counts, which are paired with preset question templates to create counting QA pairs. RSVLM-QA comprises 13,820 images and 162,373 VQA pairs, featuring extensive annotations and diverse question types. We provide a detailed statistical analysis of the dataset and a comparison with existing RS VQA benchmarks, highlighting the superior depth and breadth of RSVLM-QA's annotations. Furthermore, we conduct benchmark experiments on Six mainstream Vision Language Models (VLMs), demonstrating that RSVLM-QA effectively evaluates and challenges the understanding and reasoning abilities of current VLMs in the RS domain. We believe RSVLM-QA will serve as a pivotal resource for the RS VQA and VLM research communities, poised to catalyze advancements in the field. The dataset, generation code, and benchmark models are publicly available at https://github.com/StarZi0213/RSVLM-QA. Xing Zi, Jinghao Xiao, Yunxiao Shi, Xian Tao, Jun Li 0010, Ali Braytee, Mukesh Prasad |
ACM Multimedia | 4 |
| 2025 | A surface defect detection instrument for large aperture spherical optical elements
Yali Shi, Zhengtao Zhang, Xian Tao, Xiuqin Shang |
Neural Comput. Appl. | 4 |
| 2025 | Adaptive Meta Policy Learning With Virtual Model for Multi-Category Peg-in-Hole Assembly SkillsabstractThe generalization model for multicategory peg-in-hole assembly (MPHA) skills is hard to acquire. An adaptive meta policy learning (AMPL) algorithm with virtual model set is proposed to deal with the difficulties of low learning efficiency and low adaptability for the multicategory assembly skill learning. First, the AMPL framework incorporates meta-reinforcement learning and can obtain generalization model of multicategory assembly skills. It has higher learning efficiency compared to single-category skill learning algorithms. Second, the AMPL algorithm introduces a similarity function constructed from the demonstration learning algorithm in the state value function. It has stronger adaptability compared to the multicategory skills learning algorithms. Finally, a simulation environment set for MPHA is constructed, consisting of the mathematical force contact models and the physical simulation models. The simulations and experiments are well conducted with the proposed algorithm. The results demonstrate the efficacy of the proposed algorithm. Shaohua Yan, Xian Tao, Xumiao Ma, Tiantian Hao, De Xu |
IEEE Trans. Ind. Informatics | 2 |
| 2024 | BDC Dataset: A Comprehensive Dataset for Automated Build Damage Classification
Xing Zi, Yunxiao Shi, Taoyuan Zhu, Kairui Jin, Xian Tao, Jun Li 0010, Karthick Thiyagarajan, Mukesh Prasad |
ADMA (1) | 5 |
| 2024 | VCP-CLIP: A Visual Context Prompting Model for Zero-Shot Anomaly Segmentation
Zhen Qu, Xian Tao, Mukesh Prasad, Fei Shen 0002, Zhengtao Zhang, Xinyi Gong, Guiguang Ding |
ECCV (69) | 2 |
| 2024 | ALMRR: Anomaly Localization Mamba on Industrial Textured Surface with Feature Reconstruction and Refinement
Shichen Qu, Xian Tao, Zhen Qu, Xinyi Gong, Zhengtao Zhang, Mukesh Prasad |
PRCV (9) | 2 |
| 2023 | Hierarchical Policy Learning With Demonstration Learning for Robotic Multiple Peg-in-Hole Assembly TasksabstractThe force-based control algorithm of robotic multiple peg-in-hole assembly is a challenge. For the difficulty of low adaptability of model-based control algorithms and low learning efficiency of model-free control algorithms, a goal-based hierarchical policy learning (HPL) algorithm that combines conventional control algorithm and demonstration learning (DL) algorithm is proposed to learn the assembly skill. First, the goal-based HPL algorithm adds goal as a new variable to the action value function. Multiple states reached in each episode are randomly selected as subgoals to improve the distribution of positive rewards. Second, an initial policy that combines conventional control algorithm and DL algorithm is designed. The combined coefficient of these two algorithms is learned by HPL algorithm. Finally, a conical surface is used to compute the forces and moments of simplified assembly simulation model. Our algorithm is well implemented in both simulation and real-world environments. The experimental results verify the effectiveness of the proposed method. Shaohua Yan, De Xu, Xian Tao |
IEEE Trans. Ind. Informatics | 3 |
| 2022 | Unsupervised Anomaly Detection for Surface Defects With Dual-Siamese NetworkabstractUnsupervised anomaly detection in real industrial scenarios is challenging since the small amount of defect-free images contain limited discriminative information, and anomaly defects are unpredictable. Although nowadays image reconstruction-based methods are widely being used in various anomaly detection applications, they cannot effectively learn semantic representation, which leads to imperfect reconstruction. In this article, anomaly detection is formulated as a joint problem of feature reconstruction and inpainting in the dual-siamese framework. The proposed approach forces the network to model the feature distribution from the normal area and capture the semantic context for discriminating normal and abnormal areas. It first uses a Siamese architecture to capture discriminative features of defect-free samples and its corresponding defective samples generated by the defect random generation module. A dense feature fusion module is then employed to obtain the dense feature representation of dual input. The second Siamese network is proposed to reconstruct and inpaint the dual-dense features of the previous stage. Compared to the existing methods that mostly employ single image reconstruction, it is beneficial to simultaneously reconstruct and inpaint the information of dense discriminative features. The experimental results on the MVTec AD datasets and some major real industrial datasets demonstrate that our method achieves state-of-the-art inspection accuracy. Xian Tao, Wenzhi Ma, Zhanxin Hou, Zhenfeng Lu, Chandranath Adak |
IEEE Trans. Ind. Informatics | 1 |
| 2021 | BigyaPAn: Deep Analysis of Old Paper AdvertisementabstractIn this paper, we work on analyzing old paper advertisement (Ad). An Ad usually contains various types of textual and non-textual objects, which may also be in different orientations. We attempt to detect such objects from an early Indian-print paper Ad database comprising 1500 Ad images. The past major object detectors did not perform well on this database. We propose a deep reinforcement learning-based orientation-aware object detector. Our system learns by itself where to look and what to look of an Ad image. Therefore, it can bypass the impeding zone due to degraded image quality. To find the looking spot, we come up with a foveal transformation. In reinforcement learning, we present a scheme for shaping an internal reward with a top-up. For oriented object detection, we also propose a generic loss function. Our system obtained encouraging results from the experiments performed on the Ad database. Chandranath Adak, Xian Tao |
IJCNN | 2 |
| 2021 | A novel online self-learning system with automatic object detection model for multimedia applications
Eric-Juwei Cheng, Mukesh Prasad, Jie Yang 0052, Ding-Rong Zheng, Xian Tao, Domingo Mery, Kuu-Young Young, Chin-Teng Lin |
Multim. Tools Appl. | 5 |
| 2021 | A robust real-time facial alignment system with facial landmarks detection and rectification for multimedia applications
Kuang-Pen Chou, Mukesh Prasad, Jie Yang 0052, Sheng-Yao Su, Xian Tao, Amit Saxena 0001, Wen-Chieh Lin, Chin-Teng Lin |
Multim. Tools Appl. | 5 |
| 2020 | Detection of Power Line Insulator Defects Using Aerial Images Analyzed With Convolutional Neural NetworksabstractAs the failure of power line insulators leads to the failure of power transmission systems, an insulator inspection system based on an aerial platform is widely used. Insulator defect detection is performed against complex backgrounds in aerial images, presenting an interesting but challenging problem. Traditional methods, based on handcrafted features or shallow learning techniques, can only localize insulators and detect faults under specific detection conditions, such as when sufficient prior knowledge is available, with low background interference, at certain object scales, or under specific illumination conditions. This paper discusses the automatic detection of insulator defects using aerial images, accurately localizing insulator defects appearing in input images captured from real inspection environments. We propose a novel deep convolutional neural network (CNN) cascading architecture for performing localization and detecting defects in insulators. The cascading network uses a CNN based on a region proposal network to transform defect inspection into a two-level object detection problem. To address the scarcity of defect images in a real inspection environment, a data augmentation method is also proposed that includes four operations: 1) affine transformation; 2) insulator segmentation and background fusion; 3) Gaussian blur; and 4) brightness transformation. Defect detection precision and recall of the proposed method are 0.91 and 0.96 using a standard insulator dataset, and insulator defects under various conditions can be successfully detected. Experimental results demonstrate that this method meets the robustness and accuracy requirements for insulator defect detection. Xian Tao, Xilong Liu, Hongyan Zhang 0005, De Xu |
IEEE Trans. Syst. Man Cybern. Syst. | 1 |