VLDB 2026 Research / reviewers in the wild / expert
Xun Xu 0002
dblp:47/3944-2
· DBLP profile ↗
53ranked-venue papers
9as first author
43since 2021 · last 2026
0000-0002-5220-2240ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 34 · 6 first-author · 25 since 2021Artificial intelligence and machine learning · 30 · 5 first-author · 25 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Enhancing Generalization of Depth Estimation Foundation Model via Weakly-Supervised Adaptation with RegularizationabstractThe emergence of foundation models has substantially advanced zero-shot generalization in monocular depth estimation (MDE), as exemplified by the Depth Anything series. However, given access to some data from downstream tasks, a natural question arises: can the performance of these models be further improved? To this end, we propose WeSTAR, a parameter-efficient framework that performs \textbf{We}akly supervised \textbf{S}elf-\textbf{T}raining \textbf{A}daptation with \textbf{R}egularization, designed to enhance the robustness of MDE foundation models in unseen and diverse domains. We first adopt a dense self-training objective as the primary source of structural self-supervision. To further improve robustness, we introduce semantically-aware hierarchical normalization, which exploits instance-level segmentation maps to perform more stable and multi-scale structural normalization. Beyond dense supervision, we introduce a cost-efficient weak supervision in the form of pairwise ordinal depth annotations to further guide the adaptation process, which enforces informative ordinal constraints to mitigate local topological errors. Finally, a weight regularization loss is employed to anchor the LoRA updates, ensuring training stability and preserving the model's generalizable knowledge. Extensive experiments on both realistic and corrupted out-of-distribution datasets under diverse and challenging scenarios demonstrate that WeSTAR consistently improves generalization and achieves state-of-the-art performance across a wide range of benchmarks. Yongyi Su, Le Zhang 0001, Xun Xu 0002 |
AAAI | 5 |
| 2026 | AD-FM: Multimodal LLMs for Anomaly Detection via Multi-Stage Reasoning and Fine-Grained Reward OptimizationabstractWhile Multimodal Large Language Models (MLLMs) demonstrate remarkable capabilities across diverse domains, their application to specialized anomaly detection (AD) remains constrained by domain adaptation challenges. Existing Group Relative Policy Optimization (GRPO) based approaches suffer from two critical limitations: inadequate training data utilization when models produce uniform responses, and insufficient supervision over reasoning processes that encourage immediate binary decisions without deliberative analysis. We propose a comprehensive framework addressing these limitations through two synergistic innovations. First, we introduce a multi-stage deliberative reasoning process that guides models from region identification to focused examination, generating diverse response patterns essential for GRPO optimization while enabling structured supervision over analytical workflows. Second, we develop a fine-grained reward mechanism incorporating classification accuracy and localization supervision, transforming binary feedback into continuous signals that distinguish genuine analytical insight from spurious correctness. Comprehensive evaluation across multiple industrial datasets shows that our method achieves superior accuracy by enabling general-purpose MLLMs to acquire fine-grained visual discrimination for detecting subtle manufacturing defects. Jingyi Liao, Yongyi Su, Rongcheng Tu, Xun Xu 0002, Dacheng Tao, Xulei Yang |
AAAI | 7 |
| 2026 | STFAR: Test-time adaptive object detection through self-training and feature alignment regularization
Nanqing Liu, Yongyi Su, Lile Cai, Heng-Chao Li 0001, Kui Jia, Tianrui Li 0001, Xun Xu 0002, Chuan-Sheng Foo |
Expert Syst. Appl. | 9 |
| 2026 | Investigate Interactive Semantic Segmentation via an Uncertainty Mining ViewabstractWith the rapid development of intelligence media, traditional semantic segmentation has shown excellent potential in application scenarios like autonomous driving. However, due to limited performance, traditional segmentation models usually lead to poor user experiences in applications that require high segmentation precision. Therefore, interactive semantic segmentation (ISS) is gaining the attention is gaining attention due to its capability to generate high-precision semantic segmentation results through a few user-provided clicks for experience improvement, which thus has a promising development prospect in fine-grained application scenarios,e.g., virtual reality, smart medical, data annotation,etc.. For good interaction efficiency, most existing interactive methods make efforts to conduct suitable click simulation strategies and reasonable click encoding methods, aiming at the robust understanding of diverse user clicks and translating comprehensible user intent,i.e., assign the correct category to the clicked area, for the neural network. Though proved effective, their designs ignore the uncertainty hiding in the extracted interaction features, which reflects the interaction difficulty and the user clicking intents. This can lead to inappropriate click simulation and click encoding, limiting the interaction efficiency. Hence we focus on exploring a reasonable ISS scheme via an uncertainty mining view. Specifically, we propose an uncertainty-based class-balanced click sampling (UCCS) simulation strategy by considering both the uncertainty of the click simulation region and its semantic imbalance, to form a reasonable click distribution. Furthermore, we propose a semantic uncertainty residual encoding (SURE) method to better embed the user's intention into the localization maps, by mining semantic confusion between the click and misprediction classes. We prove the effectiveness of our design through extensive experiments and initially analyze the importance of uncertainty mining for the ISS. Our model can achieve state-of-the-art performance on three semantic segmentation benchmarks. Yutong Gao 0001, Congyan Lang, Fayao Liu, Xun Xu 0002, Yuanzhouhan Cao, Yunchao Wei |
IEEE Trans. Multim. | 4 |
| 2025 | Distribution Alignment Informed Thresholding for Semi-Supervised Curvilinear Structure SegmentationabstractCurvilinear structure segmentation using deep neural networks is often limited by the high cost of annotation. Semi-supervised learning (SSL) helps mitigate this dependency on extensive annotated data. State-of-the-art SSL approaches generate pseudo-labels for unlabeled data, which are then used for further model training. These methods primarily focus on calibrating thresholds to binarize the predictions. In this work, we assume that when labeled and unlabeled data are similar, the foreground-to-background ratio should be consistent between them. To leverage this assumption, we calibrate the threshold by minimizing the distribution gap between labeled ground truth and pseudo-labels on unlabeled data. Our proposed threshold calibration can be integrated with existing SSL methods. We evaluate its effectiveness on four datasets, demonstrating that our method outperforms current state-of-the-art SSL techniques, especially in scenarios with very low labeled data. Yuhao Mo, Bihan Wen, Xulei Yang, Ce Zhu, Xun Xu 0002 |
ICASSP | 6 |
| 2025 | Exploiting Vision Language Model for Training-Free 3D Point Cloud OOD Detection via Graph Score PropagationabstractOut-of-distribution (OOD) detection in 3D point cloud data remains a challenge, particularly in applications where safe and robust perception is critical. While existing OOD detection methods have shown progress for 2D image data, extending these to 3D environments involves unique obstacles. This paper introduces a training-free framework that leverages Vision-Language Models (VLMs) for effective OOD detection in 3D point clouds. By constructing a graph based on class prototypes and testing data, we exploit the data manifold structure to enhancing the effectiveness of VLMs for 3D OOD detection. We propose a novel Graph Score Propagation (GSP) method that incorporates prompt clustering and self-training negative prompting to improve OOD scoring with VLM. Our method is also adaptable to few-shot scenarios, providing options for practical applications. We demonstrate that GSP consistently outperforms state-of-the-art methods across synthetic and real-world datasets 3D point cloud OOD detection. Tiankai Chen, Yushu Li, Adam Goodge, Fei Teng 0001, Xulei Yang, Tianrui Li 0001, Xun Xu 0002 |
ICCV | 7 |
| 2025 | Global-Aware Monocular Semantic Scene Completion with State Space Models
Shijie Li 0006, Zhongyao Cheng, Juergen Gall, Xun Xu 0002, Xulei Yang |
ICCV | 6 |
| 2025 | Future-Aware Interaction Network for Motion Forecasting
Shijie Li 0006, Xun Xu 0002, Si Yong Yeo, Xulei Yang |
ICCV | 3 |
| 2025 | Evidential Learning-based Certainty Estimation for Robust Dense Feature MatchingabstractDense feature matching methods aim to estimate a dense correspondence field between images. Inaccurate correspondence can occur due to the presence of unmatchable region, necessitating the need for certainty measurement. This is typically addressed by training a binary classifier to decide whether each predicted correspondence is reliable. However, deep neural network-based classifiers can be vulnerable to image corruptions or perturbations, making it difficult to obtain reliable matching pairs in corrupted scenario. In this work, we propose an evidential deep learning framework to enhance the robustness of dense matching against corruptions. We modify the certainty prediction branch in dense matching models to generate appropriate belief masses and compute the certainty score by taking expectation over the resulting Dirichlet distribution. We evaluate our method on a wide range of benchmarks and show that our method leads to improved robustness against common corruptions and adversarial attacks, achieving up to 10.1\% improvement under severe corruptions. Lile Cai, Chuan-Sheng Foo, Xun Xu 0002, Zaiwang Gu, Jun Cheng 0003, Xulei Yang |
ICLR | 3 |
| 2025 | Efficient and Context-Aware Label Propagation for Zero-/Few-Shot Training-Free Adaptation of Vision-Language ModelabstractVision-language models (VLMs) have revolutionized machine learning by leveraging large pre-trained models to tackle various downstream tasks. Although label, training, and data efficiency have improved, many state-of-the-art VLMs still require task-specific hyperparameter tuning and fail to fully exploit test samples. To overcome these challenges, we propose a graph-based approach for label-efficient adaptation and inference. Our method dynamically constructs a graph over text prompts, few-shot examples, and test samples, using label propagation for inference without task-specific tuning. Unlike existing zero-shot label propagation techniques, our approach requires no additional unlabeled support set and effectively leverages the test sample manifold through dynamic graph expansion. We further introduce a context-aware feature re-weighting mechanism to improve task adaptation accuracy. Additionally, our method supports efficient graph expansion, enabling real-time inductive inference. Extensive evaluations on downstream tasks, such as fine-grained categorization and out-of-distribution generalization, demonstrate the effectiveness of our approach. The source code is available at https://github.com/Yushu-Li/ECALP. Yushu Li, Yongyi Su, Adam Goodge, Kui Jia, Xun Xu 0002 |
ICLR | 5 |
| 2025 | On the Adversarial Risk of Test Time Adaptation: An Investigation into Realistic Test-Time Data PoisoningabstractTest-time adaptation (TTA) updates the model weights during the inference stage using testing data to enhance generalization. However, this practice exposes TTA to adversarial risks. Existing studies have shown that when TTA is updated with crafted adversarial test samples, also known as test-time poisoned data, the performance on benign samples can deteriorate. Nonetheless, the perceived adversarial risk may be overstated if the poisoned data is generated under overly strong assumptions. In this work, we first review realistic assumptions for test-time data poisoning, including white-box versus grey-box attacks, access to benign data, attack order, and more. We then propose an effective and realistic attack method that better produces poisoned samples without access to benign samples, and derive an effective in-distribution attack objective. We also design two TTA-aware attack objectives. Our benchmarks of existing attack methods reveal that the TTA methods are more robust than previously believed. In addition, we analyze effective defense strategies to help develop adversarially robust TTA methods. The source code is available at https://github.com/Gorilla-Lab-SCUT/RTTDP. Yongyi Su, Yushu Li, Nanqing Liu, Kui Jia, Xulei Yang, Chuan-Sheng Foo, Xun Xu 0002 |
ICLR | 7 |
| 2025 | SODA: Out-of-Distribution Detection in Domain-Shifted Point Clouds via Neighborhood Propagation
Adam Goodge, Bryan Hooi, Jingyi Liao, Yongyi Su, Wee Siong Ng, Xun Xu 0002, Xulei Yang |
ECML/PKDD (1) | 6 |
| 2025 | Correction to: Deep negative correlation classification
Le Zhang 0001, Qibin Hou, Yun Liu 0011, Jiawang Bian, Xun Xu 0002, Joey Tianyi Zhou, Ce Zhu |
Mach. Learn. | 5 |
| 2025 | Mitigating Missing Feature Channels at Inference Stage: Test-Time Adaptation Through Self-Training With Data ImputationabstractThe robustness of deep learning model can be compromised by out-of-distribution (OOD) testing data. Test-time adaptation (TTA) emerges as an efficient method to mitigate the distribution gap by tuning model weights at inference stage. TTA are mainly demonstrated on robustifying model on additive visual corruptions or adversarial attacks. In this work, we specify an overlooked type of OOD where feature channels could be missing in testing data, potentially due to sensor fault. We reveal that self-training and data imputation can improve the model’s generalization to data with missing feature channel. To address the uncertainty associated with imputed samples, we fuse predictions from imputed and weakly-augmented samples for more reliable pseudo labels. We evaluate the effectiveness on multiple image classification benchmarks with synthesized and realistic missing feature channels, and our proposed method outperforms state-of-the-art TTA methods on all benchmarks. Yongyi Su, Xulei Yang, Xun Xu 0002 |
IEEE Signal Process. Lett. | 5 |
| 2025 | PointSAM: Pointly-Supervised Segment Anything Model for Remote Sensing ImagesabstractSegment anything model (SAM) is an advanced foundational model for image segmentation, which is gradually being applied to remote sensing images (RSIs). Due to the domain gap between RSIs and natural images, traditional methods typically use SAM as a source pretrained model and fine-tune it with fully supervised masks. Unlike these methods, our work focuses on fine-tuning SAM using more convenient and challenging point annotations. Leveraging SAM’s zero-shot capability, we adopt a self-training framework that iteratively generates pseudolabels. However, noisy labels in pseudolabels can cause error accumulation. To address this, we introduce prototype-based regularization (PBR), where target prototypes are extracted from the dataset and matched to predicted prototypes using the Hungarian algorithm to guide learning in the correct direction. In addition, RSIs have complex backgrounds and densely packed objects, making it possible for point prompts to mistakenly group multiple objects as one. To resolve this, we propose a negative prompt calibration (NPC) method based on the nonoverlapping nature of instance masks, where overlapping masks are used as negative signals to refine segmentation. Combining these techniques, we present a novel pointly-supervised SAM (PointSAM). We conduct experiments on three RSI datasets, including WHU, HRSID, and NWPU VHR-10, showing that our method significantly outperforms direct testing with SAM, SAM2, and other comparison methods. In addition, PointSAM can act as a point-to-box converter for oriented object detection, achieving promising results and indicating its potential for other point-supervised tasks. The code is available athttps://github.com/Lans1ng/PointSAM. Nanqing Liu, Xun Xu 0002, Yongyi Su, Heng-Chao Li 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2025 | SDCoT++: Improved Static-Dynamic Co-Teaching for Class-Incremental 3D Object DetectionabstractDeep learning approaches have demonstrated high effectiveness in 3D object detection tasks. However, they often suffer from a notable drop in performance on the previously trained classes when learning new classes incrementally without revisiting the old data. This is the "catastrophic forgetting" phenomenon which impedes 3D object detection in real-world scenarios, where intelligent machines must continuously learn to detect previously unseen categories. Furthermore, frequent co-occurrences of old and new classes in scenes exacerbate catastrophic forgetting and cause model confusion. To address these challenges, we propose a novel static-dynamic co-teaching approach. Our framework involves a student model and two teacher models: a static teacher with fixed weights which imparts preserved old knowledge to the student, and a dynamic teacher with continuously updated weights which transfers underlying knowledge from new data to the student. To mitigate the issue of co-occurrence, we generate pseudo labels for base (i.e. old) classes from both static and dynamic sources during incremental learning. Additionally, to mitigate the negative impact of varying occurrence frequencies of classes on fixed thresholding during the selection of pseudo labels, we calibrate the probabilities of base classes to attain more balanced class probabilities. Moreover, our static-dynamic co-teaching framework is backbone-agnostic, making it compatible with different detection architectures. We demonstrate its backbone-agnostic nature by adapting three representative 3D object detectors: VoteNet, 3DETR and CAGroup3D. Extensive experiments showcase the superior performance of our proposed method compared to baseline approaches across indoor and outdoor benchmark datasets and applicability with different backbone models. Na Zhao 0004, Peisheng Qian, Fang Wu 0009, Xun Xu 0002, Xulei Yang, Gim Hee Lee |
IEEE Trans. Image Process. | 4 |
| 2024 | Towards Real-World Test-Time Adaptation: Tri-net Self-Training with Balanced NormalizationabstractTest-Time Adaptation aims to adapt source domain model to testing data at inference stage with success demonstrated in adapting to unseen corruptions. However, these attempts may fail under more challenging real-world scenarios. Existing works mainly consider real-world test-time adaptation under non-i.i.d. data stream and continual domain shift. In this work, we first complement the existing real-world TTA protocol with a globally class imbalanced testing set. We demonstrate that combining all settings together poses new challenges to existing methods. We argue the failure of state-of-the-art methods is first caused by indiscriminately adapting normalization layers to imbalanced testing data. To remedy this shortcoming, we propose a balanced batchnorm layer to swap out the regular batchnorm at inference stage. The new batchnorm layer is capable of adapting without biasing towards majority classes. We are further inspired by the success of self-training (ST) in learning from unlabeled data and adapt ST for test-time adaptation. However, ST alone is prone to over adaption which is responsible for the poor performance under continual domain shift. Hence, we propose to improve self-training under continual domain shift by regularizing model updates with an anchored loss. The final TTA model, termed as TRIBE, is built upon a tri-net architecture with balanced batchnorm layers. We evaluate TRIBE on four datasets representing real-world TTA settings. TRIBE consistently achieves the state-of-the-art performance across multiple evaluation protocols. The code is available at https://github.com/Gorilla-Lab-SCUT/TRIBE. Yongyi Su, Xun Xu 0002, Kui Jia |
AAAI | 2 |
| 2024 | Improving the Generalization of Segmentation Foundation Model under Distribution Shift via Weakly Supervised AdaptationabstractThe success of large language models has inspired the computer vision community to explore image segmentation foundation model that is able to zero/few-shot generalize through prompt engineering. Segment-Anything (SAM), among others, is the state-of-the-art image segmentation foundation model demonstrating strong zero/few-shot generalization. Despite the success, recent studies reveal the weakness of SAM under strong distribution shift. In particular, SAM performs awkwardly on corrupted natural images, camouflaged images, medical images, etc. Motivated by the observations, we aim to develop a self-training based strategy to adapt SAM to target distribution. Given the unique challenges of large source dataset, high computation cost and incorrect pseudo label, we propose a weakly supervised self-training architecture with anchor regularization and low-rank finetuning to improve the robustness and computation efficiency of adaptation. We validate the effectiveness on 5 types of downstream segmentation tasks including natural clean/corrupted images, medical images, camouflaged images and robotic images. Our proposed method is task-agnostic in nature and outperforms pre-trained SAM and state-of-the-art domain adaptation methods on almost all downstream tasks with the same testing prompt inputs. Yongyi Su, Xun Xu 0002, Kui Jia |
CVPR | 3 |
| 2024 | Box-Level Class-Balanced Sampling For Active Object DetectionabstractTraining deep object detectors demands expensive bounding box annotation. Active learning (AL) is a promising technique to alleviate the annotation burden. Performing AL at box-level for object detection, i.e., selecting the most informative boxes to label and supplementing the sparsely-labelled image with pseudo labels, has been shown to be more cost-effective than selecting and labelling the entire image. In box-level AL for object detection, we observe that models at early stage can only perform well on majority classes, making the pseudo labels severely class-imbalanced. We propose a class-balanced sampling strategy to select more objects from minority classes for labelling, so as to make the final training data, i.e., ground truth labels obtained by AL and pseudo labels, more class-balanced to train a better model. We also propose a task-aware soft pseudo labelling strategy to increase the accuracy of pseudo labels. We evaluate our method on public benchmarking datasets and show that our method achieves state-of-the-art performance. Jingyi Liao, Xun Xu 0002, Chuan-Sheng Foo, Lile Cai |
ICIP | 2 |
| 2024 | Clip-Guided Source-Free Object Detection in Aerial ImagesabstractDomain adaptation is crucial in aerial imagery, as the visual representation of these images can significantly vary based on factors such as geographic location, time, and weather conditions. Additionally, high-resolution aerial images often require substantial storage space and may not be readily accessible to the public. To address these challenges, we propose a novel Source-Free Object Detection (SFOD) method. Specifically, our approach begins with a self-training framework, which significantly enhances the performance of baseline methods. To alleviate the noisy labels in self-training, we utilize Contrastive Language-Image Pre-training (CLIP) to guide the generation of pseudo-labels, termed CLIP-guided Aggregation (CGA). By leveraging CLIP’s zero-shot classification capability, we aggregate its scores with the original predicted bounding boxes, enabling us to obtain refined scores for the pseudo-labels. To validate the effectiveness of our method, we constructed two new datasets from different domains based on the DIOR dataset, named DIOR-C and DIOR-Cloudy. Experimental results demonstrate that our method outperforms other comparative algorithms. The code is available at https://github.com/Lans1ng/SFOD-RS. Nanqing Liu, Xun Xu 0002, Yongyi Su, Peiliang Gong, Heng-Chao Li 0001 |
IGARSS | 2 |
| 2024 | A Novel SegNet Model for Crack Image Semantic Segmentation in Bridge Inspection
Rong Pang, Xun Xu 0002, Nanqing Liu |
PAKDD (3) | 4 |
| 2024 | Revisiting pretraining for semi-supervised learning in the low-label regime
Xun Xu 0002, Jingyi Liao, Lile Cai, Kangkang Lu 0001, Wanyue Zhang, Yasin Yazici, Chuan-Sheng Foo |
Neurocomputing | 1 |
| 2024 | Deep negative correlation classification
Le Zhang 0001, Qibin Hou, Yun Liu 0011, Jiawang Bian, Xun Xu 0002, Joey Tianyi Zhou, Ce Zhu |
Mach. Learn. | 5 |
| 2024 | Revisiting Realistic Test-Time Training: Sequential Inference and Adaptation by Anchored Clustering Regularized Self-TrainingabstractDeploying models on target domain data subject to distribution shift requires adaptation. Test-time training (TTT) emerges as a solution to this adaptation under a realistic scenario where access to full source domain data is not available, and instant inference on the target domain is required. Despite many efforts into TTT, there is a confusion over the experimental settings, thus leading to unfair comparisons. In this work, we first revisit TTT assumptions and categorize TTT protocols by two key factors, i.e., whether testing data is sequentially streamed and whether source model is allowed to be trained with modified loss function. Among the multiple protocols, we adopt a realistic sequential test-time training (sTTT) protocol, under which we develop a test-time anchored clustering (TTAC) approach to enable stronger test-time feature learning. TTAC discovers clusters in both source and target domains and matches the target clusters to the source ones to improve adaptation. When source domain information is strictly absent (i.e., source-free) we further develop an efficient method to infer source domain distributions for anchored clustering. Finally, self-training (ST) has demonstrated great success in learning from unlabeled data and we empirically figure out that applying ST alone to TTT is prone to confirmation bias. Therefore, a more effective TTT approach is introduced by regularizing self-training with anchored clustering, and the improved model is referred to as TTAC++. We demonstrate that, under all TTT protocols, TTAC++ consistently outperforms the state-of-the-art methods on five TTT datasets, including corrupted target domain, selected hard samples, synthetic-to-real adaptation and adversarially attacked target domain. We hope this work will provide a fair benchmarking of TTT methods, and future research should be compared within respective protocols. Yongyi Su, Xun Xu 0002, Tianrui Li 0001, Kui Jia |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2024 | COFT-AD: COntrastive Fine-Tuning for Few-Shot Anomaly DetectionabstractExisting approaches towards anomaly detection (AD) often rely on a substantial amount of anomaly-free data to train representation and density models. However, large anomaly-free datasets may not always be available before the inference stage; in which case an anomaly detection model must be trained with only a handful of normal samples, a.k.a. few-shot anomaly detection (FSAD). In this paper, we propose a novel methodology to address the challenge of FSAD which incorporates two important techniques. Firstly, we employ a model pre-trained on a large source dataset to initialize model weights. Secondly, to ameliorate the covariate shift between source and target domains, we adopt contrastive training to fine-tune on the few-shot target domain data. To learn suitable representations for the downstream AD task, we additionally incorporate cross-instance positive pairs to encourage a tight cluster of the normal samples, and negative pairs for better separation between normal and synthesized negative samples. We evaluate few-shot anomaly detection on 3 controlled AD tasks and 4 real-world AD tasks to demonstrate the effectiveness of the proposed method. Jingyi Liao, Xun Xu 0002, Adam Goodge, Chuan-Sheng Foo |
IEEE Trans. Image Process. | 2 |
| 2024 | Exploring Diversity-Based Active Learning for 3D Object Detection in Autonomous Drivingabstract3D object detection has recently received much attention due to its great potential in autonomous vehicle (AV). The success of deep learning based object detectors relies on the availability of large-scale annotated datasets, which is time-consuming and expensive to compile, especially for 3D bounding box annotation. In this work, we investigate diversity-based active learning (AL) as a potential solution to alleviate the annotation burden. Given limited annotation budget, only the most informative frames and objects are automatically selected for human to annotate. Technically, we take the advantage of the multimodal information provided in an AV dataset, and propose a novel acquisition function that enforces spatial and temporal diversity in the selected samples. We benchmark the proposed method against other AL strategies under realistic annotation cost measurements, where the realistic costs for annotating a frame and a 3D bounding box are both taken into consideration. We demonstrate the effectiveness of the proposed method on the nuScenes dataset and show that it outperforms existing AL strategies significantly. Jinpeng Lin, Zhihao Liang 0002, Shengheng Deng, Lile Cai, Tao Jiang 0014, Tianrui Li 0001, Kui Jia, Xun Xu 0002 |
IEEE Trans. Intell. Transp. Syst. | 8 |
| 2023 | On Adversarial Robustness of Audio ClassifiersabstractWe make three contributions to improve adversarial robustness of audio classifiers. First, most existing works focus on ℓp-norm bounded adversarial perturbations. Instead, we consider signal-to-noise ratio (SNR) as a more natural measure of adversarial perturbations for audio data. We show that perturbed examples with a particular SNR can be generated using a corresponding ℓ2-norm perturbation, and establish the equivalence of these two metrics in assessing adversarial perturbations. This connection enables direct control of the SNR quality of perturbed examples and allows comparison using perturbations with different ℓp-norm constraints. Second, we are among the first to introduce APGD attack for adversarial training on audio data. In our experiments, APGD adversarial training is robust to adversarial attacks without compromising clean accuracy. Last, we improve adversarial robustness by adapting CutMix to audio - cutting and mixing two audio clips together - in conjunction with adversarial training, and observe improvements in robustness on US8K. Kangkang Lu 0001, Xun Xu 0002, Chuan-Sheng Foo |
ICASSP | 3 |
| 2023 | On the Robustness of Open-World Test-Time Training: Self-Training with Dynamic Prototype ExpansionabstractGeneralizing deep learning models to unknown target domain distribution with low latency has motivated research into test-time training/adaptation (TTT/TTA). Existing approaches often focus on improving test-time training performance under well-curated target domain data. As figured out in this work, many state-of-the-art methods fail to maintain the performance when the target domain is contaminated with strong out-of-distribution (OOD) data, a.k.a. open-world test-time training (OWTTT). The failure is mainly due to the inability to distinguish strong OOD samples from regular weak OOD samples. To improve the robustness of OWTTT we first develop an adaptive strong OOD pruning which improves the efficacy of the self-training TTT method. We further propose a way to dynamically expand the prototypes to represent strong OOD samples for an improved weak/strong OOD data separation. Finally, we regularize self-training with distribution alignment and the combination yields the state-of-the-art performance on 5 OWTTT benchmarks. The code is available at https://github.com/Yushu-Li/OWTTT. Yushu Li, Xun Xu 0002, Yongyi Su, Kui Jia |
ICCV | 2 |
| 2023 | Diverse and consistent multi-view networks for semi-supervised regression
Cuong Manh Nguyen, Arun Raja, Le Zhang 0001, Xun Xu 0002, Balagopal Unnikrishnan, Mohamed Ragab 0002, Kangkang Lu 0001, Chuan-Sheng Foo |
Mach. Learn. | 4 |
| 2023 | Weakly Supervised 3D Point Cloud Segmentation via Multi-Prototype LearningabstractAddressing the annotation challenge in 3D Point Cloud segmentation has inspired research into weakly supervised learning. Existing approaches mainly focus on exploiting manifold and pseudo-labeling to make use of large unlabeled data points. A fundamental challenge here lies in the large intra-class variations of local geometric structure, resulting in subclasses within a semantic class. In this work, we leverage this intuition and opt for maintaining an individual classifier for each subclass. Technically, we design a multi-prototype classifier, each prototype serves as the classifier weights for one subclass. To enable effective updating of multi-prototype classifier weights, we propose two constraints respectively for updating the prototypes w.r.t. all point features and for encouraging the learning of diverse prototypes. Experiments on weakly supervised 3D point cloud segmentation tasks validate the efficacy of proposed method in particular at low-label regime. Our hypothesis is also verified given the consistent discovery of semantic subclasses at no cost of additional annotations. Yongyi Su, Xun Xu 0002, Kui Jia |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2023 | Transformation-Invariant Network for Few-Shot Object Detection in Remote-Sensing ImagesabstractObject detection in remote sensing images relies on a large amount of labeled data for training. However, the increasing number of new categories and class imbalance make exhaustive annotation impractical. Few-shot object detection (FSOD) addresses this issue by leveraging meta-learning on seen base classes and fine-tuning on novel classes with limited labeled samples. Nonetheless, the substantial scale and orientation variations of objects in remote sensing images pose significant challenges to existing few-shot object detection methods. To overcome these challenges, we propose integrating a feature pyramid network and utilizing prototype features to enhance query features, thereby improving existing FSOD methods. We refer to this modified FSOD approach as a Strong Baseline, which has demonstrated significant performance improvements compared to the original baselines. Furthermore, we tackle the issue of spatial misalignment caused by orientation variations between the query and support images by introducing a Transformation-Invariant Network (TINet). TINet ensures geometric invariance and explicitly aligns the features of the query and support branches, resulting in additional performance gains while maintaining the same inference speed as the Strong Baseline. Extensive experiments on three widely used remote sensing object detection datasets, i.e., NWPU VHR-10.v2, DIOR, and HRRSD demonstrated the effectiveness of the proposed method. Nanqing Liu, Xun Xu 0002, Turgay Çelik 0001, Zongxin Gan, Heng-Chao Li 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | Exploring Active Learning for Semiconductor Defect SegmentationabstractThe development of X-Ray microscopy (XRM) technology has enabled non-destructive inspection of semiconductor structures for defect identification. Deep learning is widely used as the state-of-the-art approach to perform visual analysis tasks. However, deep learning based models require large amount of annotated data to train. This can be time-consuming and expensive to obtain especially for dense prediction tasks like semantic segmentation. In this work, we explore active learning (AL) as a potential solution to alleviate the annotation burden. We identify two unique challenges when applying AL on semiconductor XRM scans: large domain shift and severe class-imbalance. To address these challenges, we propose to perform contrastive pretraining on the unlabelled data to obtain the initialization weights for each AL cycle, and a rareness-aware acquisition function that favors the selection of samples containing rare classes. We evaluate our method on a semiconductor dataset that is compiled from XRM scans of high bandwidth memory structures composed of logic and memory dies, and demonstrate that our method achieves state-of-the-art performance. Lile Cai, Ramanpreet Singh Pahwa, Xun Xu 0002, Jie Wang 0042, Richard Chang 0002, Lining Zhang, Chuan-Sheng Foo |
ICIP | 3 |
| 2022 | Open-Set Semi-Supervised Learning for 3D Point Cloud UnderstandingabstractSemantic understanding of 3D point cloud relies on learning models with massively annotated data, which, in many cases, are expensive or difficult to collect. This has led to an emerging research interest in semi-supervised learning (SSL) for 3D point cloud. It is commonly assumed in SSL that the unlabeled data are drawn from the same distribution as that of the labeled ones; This assumption, however, rarely holds true in realistic environments. Blindly using out-of-distribution (OOD) unlabeled data could harm SSL performance. In this work, we propose to selectively utilize unlabeled data through sample weighting, so that only conducive unlabeled data would be prioritized. To estimate the weights, we adopt a bi-level optimization framework which iteratively optimizes a meta-objective on a held-out validation set and a task-objective on a training set. Faced with the instability of efficient bi-level optimizers, we further propose three regularization techniques to enhance the training stability. Extensive experiments on 3D point cloud classification and segmentation tasks verify the effectiveness of our proposed method. We also demonstrate the feasibility of a more efficient training strategy. Our code is released on Github1. Xian Shi, Xun Xu 0002, Wanyue Zhang, Xiatian Zhu, Chuan-Sheng Foo, Kui Jia |
ICPR | 2 |
| 2022 | Revisiting Realistic Test-Time Training: Sequential Inference and Adaptation by Anchored ClusteringabstractDeploying models on target domain data subject to distribution shift requires adaptation. Test-time training (TTT) emerges as a solution to this adaptation under a realistic scenario where access to full source domain data is not available and instant inference on target domain is required. Despite many efforts into TTT, there is a confusion over the experimental settings, thus leading to unfair comparisons. In this work, we first revisit TTT assumptions and categorize TTT protocols by two key factors. Among the multiple protocols, we adopt a realistic sequential test-time training (sTTT) protocol, under which we further develop a test-time anchored clustering (TTAC) approach to enable stronger test-time feature learning. TTAC discovers clusters in both source and target domain and match the target clusters to the source ones to improve generalization. Pseudo label filtering and iterative updating are developed to improve the effectiveness and efficiency of anchored clustering. We demonstrate that under all TTT protocols TTAC consistently outperforms the state-of-the-art methods on six TTT datasets. We hope this work will provide a fair benchmarking of TTT methods and future research should be compared within respective protocols. A demo code is available at https://github.com/Gorilla-Lab-SCUT/TTAC. Yongyi Su, Xun Xu 0002, Kui Jia |
NeurIPS | 2 |
| 2022 | Learning Clustering for Motion SegmentationabstractSubspace clustering has been extensively studied from the hypothesis-and-test, algebraic, and spectral clustering-based perspectives. Most assume that only a single type/class of subspace is present. Generalizations to multiple types are non-trivial, plagued by challenges such as choice of types and numbers of models, sampling imbalance and parameter tuning. In many real world problems, data may not lie perfectly on a linear subspace and hand designed linear subspace models may not fit into these situations. In this work, we formulate the multi-type subspace clustering problem as one of learning non-linear subspace filters via deep multi-layer perceptrons (mlps). The response to the learnt subspace filters serve as the feature embedding that is clustering-friendly, i.e., points of the same clusters will be embedded closer together through the network. For inference, we apply K-means to the network output to cluster the data. Experiments are carried out on synthetic data and real world motion segmentation problems, producing state-of-the-art results. Xun Xu 0002, Le Zhang 0001, Loong Fah Cheong, Zhuwen Li, Ce Zhu |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2022 | MA-GANet: A Multi-Attention Generative Adversarial Network for Defocus Blur DetectionabstractBackground clutters pose challenges to defocus blur detection. Existing approaches often produce artifact predictions in background areas with clutter and relatively low confident predictions in boundary areas. In this work, we tackle the above issues from two perspectives. Firstly, inspired by the recent success of self-attention mechanism, we introduce channel-wise and spatial-wise attention modules to attentively aggregate features at different channels and spatial locations to obtain more discriminative features. Secondly, we propose a generative adversarial training strategy to suppress spurious and low reliable predictions. This is achieved by utilizing a discriminator to identify predicted defocus map from ground-truth ones. As such, the defocus network (generator) needs to produce ‘realistic’ defocus map to minimize discriminator loss. We further demonstrate that the generative adversarial training allows exploiting additional unlabeled data to improve performance, a.k.a. semi-supervised learning, and we provide the first benchmark on semi-supervised defocus detection. Finally, we demonstrate that the existing evaluation metrics for defocus detection generally fail to quantify the robustness with respect to thresholding. For a fair and practical evaluation, we introduce an effective yet efficient$AUF_\beta $metric. Extensive experiments on three public datasets verify the superiority of the proposed methods compared against state-of-the-art approaches. Xun Xu 0002, Le Zhang 0001, Chao Zhang 0072, Chuan-Sheng Foo, Ce Zhu |
IEEE Trans. Image Process. | 2 |
| 2022 | SemiCurv: Semi-Supervised Curvilinear Structure SegmentationabstractRecent work on curvilinear structure segmentation has mostly focused on backbone network design and loss engineering. The challenge of collecting labelled data, an expensive and labor intensive process, has been overlooked. While labelled data is expensive to obtain, unlabelled data is often readily available. In this work, we propose SemiCurv, a semi-supervised learning (SSL) framework for curvilinear structure segmentation that is able to utilize such unlabelled data to reduce the labelling burden. Our framework addresses two key challenges in formulating curvilinear segmentation in a semi-supervised manner. First, to fully exploit the power of consistency based SSL, we introduce a geometric transformation as strong data augmentation and then align segmentation predictions via a differentiable inverse transformation to enable the computation of pixel-wise consistency. Second, the traditional mean square error (MSE) on unlabelled data is prone to collapsed predictions and this issue exacerbates with severe class imbalance (significantly more background pixels). We propose a N-pair consistency loss to avoid trivial predictions on unlabelled data. We evaluate SemiCurv on six curvilinear segmentation datasets, and find that with no more than 5% of the labelled data, it achieves close to 95% of the performance relative to its fully supervised counterpart. Xun Xu 0002, Cuong Manh Nguyen, Yasin Yazici, Kangkang Lu 0001, Hlaing Min, Chuan-Sheng Foo |
IEEE Trans. Image Process. | 1 |
| 2021 | On Automatic Data Augmentation for 3D Point Cloud Classification
Wanyue Zhang, Xun Xu 0002, Fayao Liu, Le Zhang 0001, Chuan-Sheng Foo |
BMVC | 2 |
| 2021 | Revisiting Superpixels for Active Learning in Semantic Segmentation With Realistic Annotation CostsabstractState-of-the-art methods for semantic segmentation are based on deep neural networks that are known to be data-hungry. Region-based active learning has shown to be a promising method for reducing data annotation costs. A key design choice for region-based AL is whether to use regularly-shaped regions (e.g., rectangles) or irregularly-shaped region (e.g., superpixels). In this work, we address this question under realistic, click-based measurement of annotation costs. In particular, we revisit the use of super-pixels and demonstrate that the inappropriate choice of cost measure (e.g., the percentage of labeled pixels), may cause the effectiveness of the superpixel-based approach to be under-estimated. We benchmark the superpixel-based approach against the traditional "rectangle+polygon"-based approach with annotation cost measured in clicks, and show that the former outperforms on both Cityscapes and PASCAL VOC. We further propose a class-balanced acquisition function to boost the performance of the superpixel-based approach and demonstrate its effectiveness on the evaluation datasets. Our results strongly argue for the use of superpixel-based AL for semantic segmentation and highlight the importance of using realistic annotation costs in evaluating such methods. Lile Cai, Xun Xu 0002, Jun Hao Liew, Chuan-Sheng Foo |
CVPR | 2 |
| 2021 | 3D AffordanceNet: A Benchmark for Visual Object Affordance UnderstandingabstractThe ability to understand the ways to interact with objects from visual cues, a.k.a. visual affordance, is essential to vision-guided robotic research. This involves categorizing, segmenting and reasoning of visual affordance. Relevant studies in 2D and 2.5D image domains have been made previously, however, a truly functional understanding of object affordance requires learning and prediction in the 3D physical domain, which is still absent in the community. In this work, we present a 3D AffordanceNet dataset, a bench-mark of 23k shapes from 23 semantic object categories, annotated with 18 visual affordance categories. Based on this dataset, we provide three benchmarking tasks for evaluating visual affordance understanding, including full-shape, partial-view and rotation-invariant affordance estimations. Three state-of-the-art point cloud deep learning networks are evaluated on all tasks. In addition we also investigate a semi-supervised learning setup to explore the possibility to benefit from unlabeled data. Comprehensive results on our contributed dataset show the promise of visual affordance understanding as a valuable yet challenging benchmark. Shengheng Deng, Xun Xu 0002, Chaozheng Wu, Ke Chen 0004, Kui Jia |
CVPR | 2 |
| 2021 | ARMOURED: Adversarially Robust MOdels using Unlabeled data by REgularizing Diversity
Kangkang Lu 0001, Cuong Manh Nguyen, Xun Xu 0002, Kiran Chari, Yu Jing Goh, Chuan-Sheng Foo |
ICLR | 3 |
| 2021 | 3D Rigid Motion Segmentation with Mixed and Unknown Number of ModelsabstractMany real-world video sequences cannot be conveniently categorized as general or degenerate; in such cases, imposing a false dichotomy in using the fundamental matrix or homography model for motion segmentation on video sequences would lead to difficulty. Even when we are confronted with a general scene-motion, the fundamental matrix approach as a model for motion segmentation still suffers from several defects, which we discuss in this paper. The full potential of the fundamental matrix approach could only be realized if we judiciously harness information from the simpler homography model. From these considerations, we propose a multi-model spectral clustering framework that synergistically combines multiple models (homography and fundamental matrix) together. We show that the performance can be substantially improved in this way. For general motion segmentation tasks, the number of independently moving objects is often unknown a priori and needs to be estimated from the observations. This is referred to as model selection and it is essentially still an open research problem. In this work, we propose a set of model selection criteria balancing data fidelity and model complexity. We perform extensive testing on existing motion segmentation datasets with both segmentation and model selection tasks, achieving state-of-the-art performance on all of them; we also put forth a more realistic and challenging dataset adapted from the KITTI benchmark, containing real-world effects such as strong perspectives and strong forward translations not seen in the traditional datasets. Xun Xu 0002, Loong Fah Cheong, Zhuwen Li |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2021 | Exploring Spatial Diversity for Region-Based Active LearningabstractState-of-the-art methods for semantic segmentation are based on deep neural networks trained on large-scale labeled datasets. Acquiring such datasets would incur large annotation costs, especially for dense pixel-level prediction tasks like semantic segmentation. We consider region-based active learning as a strategy to reduce annotation costs while maintaining high performance. In this setting, batches of informative image regions instead of entire images are selected for labeling. Importantly, we propose that enforcing local spatial diversity is beneficial for active learning in this case, and to incorporate spatial diversity along with the traditional active selection criterion, e.g., data sample uncertainty, in a unified optimization framework for region-based active learning. We apply this framework to the Cityscapes and PASCAL VOC datasets and demonstrate that the inclusion of spatial diversity effectively improves the performance of uncertainty-based and feature diversity-based active learning methods. Our framework achieves 95% performance of fully supervised methods with only 5 - 9% of the labeled pixels, outperforming all state-of-the-art region-based active learning methods for semantic segmentation. Lile Cai, Xun Xu 0002, Lining Zhang, Chuan-Sheng Foo |
IEEE Trans. Image Process. | 2 |
| 2020 | MultiANet: a Multi-Attention Network for Defocus Blur DetectionabstractDefocus blur detection is a challenging task because of obscure homogenous regions and interferences of background clutter. Most existing deep learning-based methods mainly focus on building wider or deeper network to capture multi-level features, neglecting to extract the feature relationships of intermediate layers, thus hindering the discriminative ability of network. Moreover, fusing features at different levels have been demonstrated to be effective. However, direct integrating without distinction is not optimal because low-level features focus on fine details only and could be distracted by background clutters. To address these issues, we propose the Multi-Attention Network for stronger discriminative learning and spatial guided low-level feature learning. Specifically, a channel-wise attention module is applied to both high-level and low-level feature maps to capture channel-wise global dependencies. In addition, a spatial attention module is employed to low-level features maps to emphasize effective detailed information. Experimental results show the performance of our network is superior to the state-of-the-art algorithms. Xun Xu 0002, Chao Zhang 0072, Ce Zhu |
MMSP | 2 |
| 2019 | C3AE: Exploring the Limits of Compact Model for Age EstimationabstractAge estimation is a classic learning problem in computer vision. Many larger and deeper CNNs have been proposed with promising performance, such as AlexNet, VggNet, GoogLeNet and ResNet. However, these models are not practical for the embedded/mobile devices. Recently, MobileNets and ShuffleNets have been proposed to reduce the number of parameters, yielding lightweight models. However, their representation has been weakened because of the adoption of depth-wise separable convolution. In this work, we investigate the limits of compact model for small-scale image and propose an extremely Compact yet efficient Cascade Context-based Age Estimation model(C3AE). This model possesses only 1/9 and 1/2000 parameters compared with MobileNets/ShuffleNets and VggNet, while achieves competitive performance. In particular, we re-define age estimation problem by two-points representation, which is implemented by a cascade model. Moreover, to fully utilize the facial context information, multi-branch CNN network is proposed to aggregate multi-scale context. Experiments are carried out on three age estimation datasets. The state-of-the-art performance on compact model has been achieved with a relatively large margin. Chao Zhang 0072, Shuaicheng Liu, Xun Xu 0002, Ce Zhu |
CVPR | 3 |
| 2018 | Robust Video Background Identification by Dominant Rigid Motion Estimation
Kaimo Lin, Nianjuan Jiang, Loong Fah Cheong, Jiangbo Lu, Xun Xu 0002 |
ACCV (2) | 5 |
| 2018 | Motion Segmentation by Exploiting Complementary Geometric ModelsabstractMany real-world sequences cannot be conveniently categorized as general or degenerate; in such cases, imposing a false dichotomy in using the fundamental matrix or homography model for motion segmentation would lead to difficulty. Even when we are confronted with a general scene-motion, the fundamental matrix approach as a model for motion segmentation still suffers from several defects, which we discuss in this paper. The full potential of the fundamental matrix approach could only be realized if we judiciously harness information from the simpler homography model. From these considerations, we propose a multi-view spectral clustering framework that synergistically combines multiple models together. We show that the performance can be substantially improved in this way. We perform extensive testing on existing motion segmentation datasets, achieving state-of-the-art performance on all of them; we also put forth a more realistic and challenging dataset adapted from the KITTI benchmark, containing real-world effects such as strong perspectives and strong forward translations not seen in the traditional datasets. Xun Xu 0002, Loong Fah Cheong, Zhuwen Li |
CVPR | 1 |
| 2018 | Image Ordinal Classification and Understanding: Grid Dropout with Masking LabelabstractImage ordinal classification refers to predicting a discrete target value which carries ordering correlation among image categories. The limited size of labeled ordinal data renders modern deep learning approaches easy to overfit. To tackle this issue, neuron dropout and data augmentation were proposed which, however, still suffer from over-parameterization and breaking spatial structure, respectively. To address the issues, we first propose a grid dropout method that randomly dropout/blackout some areas of the training image. Then we combine the objective of predicting the blackout patches with classification to take advantage of the spatial information. Finally we demonstrate the effectiveness of both approaches by visualizing the Class Activation Map (CAM) and discover that grid dropout is more aware of the whole facial areas and more robust than neuron dropout for small training dataset. Experiments are conducted on a challenging age estimation dataset-Adience dataset with very competitive results compared with state-of-the-art methods. Chao Zhang 0072, Ce Zhu, Jimin Xiao, Xun Xu 0002, Yipeng Liu 0001 |
ICME | 4 |
| 2018 | Visual aesthetic understanding: Sample-specific aesthetic classification and deep activation map visualization
Chao Zhang 0072, Ce Zhu, Xun Xu 0002, Yipeng Liu 0001, Jimin Xiao, Tammam Tillo |
Signal Process. Image Commun. | 3 |
| 2017 | Transductive Zero-Shot Action Recognition by Word-Vector Embedding
Xun Xu 0002, Timothy M. Hospedales, Shaogang Gong |
Int. J. Comput. Vis. | 1 |
| 2017 | Discovery of Shared Semantic Spaces for Multiscene Video Query and SummarizationabstractThe growing rate of public space closed-circuit television (CCTV) installations has generated a need for automated methods for exploiting video surveillance data, including scene understanding, query, behavior annotation, and summarization. For this reason, extensive research has been performed on surveillance scene understanding and analysis. However, most studies have considered single scenes or groups of adjacent scenes. The semantic similarity between different but related scenes (e.g., many different traffic scenes of a similar layout) is not generally exploited to improve any automated surveillance tasks and reduce manual effort. Exploiting commonality and sharing any supervised annotations between different scenes is, however, challenging due to the following reason: some scenes are totally unrelated and thus any information sharing between them would be detrimental, whereas others may share only a subset of common activities and thus information sharing is only useful if it is selective. Moreover, semantically similar activities that should be modeled together and shared across scenes may have quite different pixel-level appearances in each scene. To address these issues, we develop a new framework for distributed multiple-scene global understanding that clusters surveillance scenes by their ability to explain each other's behaviors and further discovers which subset of activities are shared versus scene specific within each cluster. We show how to use this structured representation of multiple scenes to improve common surveillance tasks, including scene activity understanding, cross-scene query-by-example, behavior classification with reduced supervised labeling requirements, and video summarization. In each case, we demonstrate how our multiscene model improves on a collection of standard single-scene models and a flat model of all scenes. Xun Xu 0002, Timothy M. Hospedales, Shaogang Gong |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2016 | Multi-Task Zero-Shot Action Recognition with Prioritised Data Augmentation
Xun Xu 0002, Timothy M. Hospedales, Shaogang Gong |
ECCV (2) | 1 |
| 2015 | Semantic embedding space for zero-shot action recognitionabstractThe number of categories for action recognition is growing rapidly. It is thus becoming increasingly hard to collect sufficient training data to learn conventional models for each category. This issue may be ameliorated by the increasingly popular “zero-shot learning” (ZSL) paradigm. In this framework a mapping is constructed between visual features and a human interpretable semantic description of each category, allowing categories to be recognised in the absence of any training data. Existing ZSL studies focus primarily on image data, and attribute-based semantic representations. In this paper, we address zero-shot recognition in contemporary video action recognition tasks, using semantic word vector space as the common space to embed videos and category labels. This is more challenging because the mapping between the semantic space and space-time features of videos containing complex actions is more complex and harder to learn. We demonstrate that a simple self-training and data augmentation strategy can significantly improve the efficacy of this mapping. Experiments on human action datasets including HMDB51 and UCF101 demonstrate that our approach achieves the state-of-the-art zero-shot action recognition performance. Xun Xu 0002, Timothy M. Hospedales, Shaogang Gong |
ICIP | 1 |