EDBT 2026 Demo / reviewers in the wild / expert
Qian Wang 0017
dblp:75/5723-17
· DBLP profile ↗
30ranked-venue papers
16as first author
22since 2021 · last 2026
0000-0002-5906-1890ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 24 · 13 first-author · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 3 first-author · 3 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | TranSAC: An unsupervised transferability metric based on task speciality and domain commonality
Qianshan Zhan, Xiaojun Zeng, Qian Wang 0017 |
Pattern Recognit. | 3 |
| 2026 | To Transfer or Not to Transfer: Unified Transferability Metric and AnalysisabstractTransferability estimation is a fundamental problem in transfer learning, which aims to predict whether transferring knowledge from a source domain will improve performance on a target task. Existing research focuses on classification and neglects domain/task differences, as well as only very limited research for regression. Most importantly, there is a lack of research to determine whether to transfer or not. To address these gaps, we propose Wasserstein distance-based joint estimation (WDJE), a unified transferability metric for both classification and regression under domain and task differences. WDJE facilitates decision-making on whether to transfer by comparing the target risk with and without transfer. To enable this comparison, we estimate the unobservable post-transfer risk using a nonsymmetric, interpretable, and easy-to-calculate upper bound that remains applicable even with limited target labels. The proposed bound relates the target transfer risk to source model performance, domain, and task differences based on the Wasserstein distance. We further extend the proposed bound to the unsupervised setting and establish a generalization bound from finite empirical samples. We evaluate WDJE and the proposed risk bound across 42 transfer scenarios, including CIFAR-100 (CF100) and Office-Home image classification and C-MAPSS remaining-useful-life regression prediction. WDJE achieves a perfect consistency index ( $CI$ ) of 1 in 25 cases and an overall mean $CI$ of 0.89, accurately suggesting when transfer should (or should not) be performed. The proposed bound achieves the average Pearson correlations of 0.99 on CF100, 0.72 on Office-Home, and 0.96 on C-MAPSS, illustrating state-of-the-art performance in approximating the true post-transfer risk. Qianshan Zhan, Xiaojun Zeng, Qian Wang 0017 |
IEEE Trans. Cybern. | 3 |
| 2025 | Progressive Dual-Space Discovering of Unknowns for Source-Free Open-Set Domain Adaptation
Qianshan Zhan, Qian Wang 0017, Xiaojun Zeng |
ECML/PKDD (8) | 2 |
| 2025 | Enhanced cross-domain lithology classification in imbalanced datasets using an unsupervised domain Adversarial Network
Yunxin Xie, Liangyu Jin, Chenyang Zhu 0001, Weibin Luo, Qian Wang 0017 |
Eng. Appl. Artif. Intell. | 5 |
| 2025 | Reducing bias in source-free unsupervised domain adaptation for regressionabstractDue to data privacy and storage concerns, Source-Free Unsupervised Domain Adaptation (SFUDA) focuses on improving an unlabelled target domain by leveraging a pre-trained source model without access to source data. While existing studies attempt to train target models by mitigating biases induced by noisy pseudo labels, they often lack theoretical guarantees for fully reducing biases and have predominantly addressed classification tasks rather than regression ones. To address these gaps, our analysis delves into the generalisation error bound of the target model, aiming to understand the intrinsic limitations of pseudo-label-based SFUDA methods. Theoretical results reveal that biases influencing generalisation error extend beyond the commonly highlighted label inconsistency bias, which denotes the mismatch between pseudo labels and ground truths, and the feature-label mapping bias, which represents the difference between the proxy target regressor and the real target regressor. Equally significant is the feature misalignment bias, indicating the misalignment between the estimated and real target feature distributions. This factor is frequently neglected or not explicitly addressed in current studies. Additionally, the label inconsistency bias can be unbounded in regression due to the continuous label space, further complicating SFUDA for regression tasks. Guided by these theoretical insights, we propose a Bias-Reduced Regression (BRR) method for SFUDA in regression. This method incorporates Feature Distribution Alignment (FDA) to reduce the feature misalignment bias, Hybrid Reliability Evaluation (HRE) to reduce the feature-label mapping bias and pseudo label updating to mitigate the label inconsistency bias. Experiments demonstrate the superior performance of the proposed BRR, and the effectiveness of FDA and HRE in reducing biases for regression tasks in SFUDA. Qianshan Zhan, Xiaojun Zeng, Qian Wang 0017 |
Neural Networks | 3 |
| 2025 | Tensorial multiview low-rank high-order graph learning for context-enhanced domain adaptation
Chenyang Zhu 0001, Lanlan Zhang, Weibin Luo, Guangqi Jiang, Qian Wang 0017 |
Neural Networks | 5 |
| 2024 | Single-Step Support Set Mining for Realistic Few-Shot Image ClassificationabstractTraditional few-shot learning (FSL) methods, often based on N-way K-shot classification, typically assume access to a large amount of labelled base classes and a class-balanced support set, which are not always feasible in real-world applications. This assumption limits the applicability of these methods in scenarios where data is scarce or imbalanced. The lack of base classes prevents the meta-training for generalized image feature extraction. We investigate the efficacy of using pre-trained models for feature extraction in practical FSL tasks, exploring how these models can compensate for the limited data availability. On the other hand, annotating K samples for each class to form the class-balanced support set is non-trivial, given that the ground truth label is unknown before annotating actually happens. This challenge highlights the need for efficient annotation strategies in FSL. This raises an overlooked research question how we can efficiently select samples to annotate to form the support set for FSL. To address this, we introduce and compare various Single-Step Support Set Mining (S4M) strategies for efficient data labeling. The proposed S4M enhances the relevance and representativeness of the selected samples, offering a more data-driven strategy for optimizing sample selection. This approach is particularly beneficial in reducing the annotation burden and improving the quality of the support set. To validate the effectiveness of proposed S4M strategies, we conduct extensive empirical studies on the transductive few-shot image classification tasks by incorporating S4M into a state-of-the-art transductive FSL framework. Our experimental results on benchmark and realistic datasets demonstrate the effectiveness of S4M under various scenarios, marking it as a practical and efficient choice for few-shot image classification in real-world applications, especially in situations where data is limited or imbalanced. Qian Wang 0017, Lanfang Dong |
IJCNN | 2 |
| 2024 | A Hybrid Few-Shot Image Classification Framework Combining Gaussian Modeling and Label PropagationabstractHumans possess the remarkable ability to recognize new objects with merely a handful of labeled examples, whereas contemporary deep learning models continue to face challenges in few-shot learning scenarios, primarily due to the scarcity of training data. In this study, we concentrate on addressing the challenges associated with transductive and semi-supervised few-shot image classification, both methods permitting the incorporation of unlabeled data during the training phase. To fully leverage the potential of unlabeled data, we explore a variety of unsupervised and semi-supervised learning approaches, including manifold learning, aimed at uncovering the intrinsic properties of the data. Specifically, we employ the locality preserving projection method as a powerful enabling technique for discriminative feature learning. The features learned are integrated into our proposed hybrid few-shot learning (FSL) framework, collaboratively augmenting the performance of few-shot image classification. Our proposed hybrid FSL framework capitalizes on the synergistic capabilities of both the parametric Gaussian model and the non-parametric label propagation model through a straightforward score-level ensemble learning approach. Consequently, our methodology yields superior outcomes on four benchmark datasets (miniImageNet, tieredImageNet, CUB, and CIFAR-FS), for both transductive and semi-supervised few-shot image classification tasks. Qian Wang 0017, Lanfang Dong |
ICMR | 2 |
| 2024 | Multiview latent space learning with progressively fine-tuned deep features for unsupervised domain adaptation
Chenyang Zhu 0001, Qian Wang 0017, Yunxin Xie, Shoukun Xu |
Inf. Sci. | 2 |
| 2023 | FederatedNILM: A Distributed and Privacy-Preserving Framework for Non-Intrusive Load Monitoring Based on Federated Deep LearningabstractNon-intrusive load monitoring (NILM), which usually utilizes machine learning methods and is effective in disaggregating smart meter readings from the household-level into appliance-level consumption, can help to analyze electricity consumption behaviours of users and enable practical smart energy and smart grid applications. However, smart meters are privately owned and distributed, which make real-world applications of NILM challenging. To this end, this paper develops a distributed and privacy-preserving federated deep learning framework for NILM (FederatedNILM), which combines federated learning with a state-of-the-art deep learning architecture to conduct NILM for the classification of typical states of household appliances. Through extensive comparative experiments, the effectiveness of the proposed FederatedNILM framework is demonstrated. Shuang Dai, Fan-Lin Meng, Qian Wang 0017, Xizhong Chen |
IJCNN | 3 |
| 2023 | On Fine-Tuned Deep Features for Unsupervised Domain AdaptationabstractPrior feature transformation based approaches to Unsupervised Domain Adaptation (UDA) employ the deep features extracted by pre-trained deep models without fine-tuning them on the specific source or target domain data for a particular domain adaptation task. In contrast, end-to-end learning based approaches optimise the pre-trained backbones and the customised adaptation modules simultaneously to learn domain-invariant features for UDA. In this work, we explore the potential of combining fine-tuned features and feature transformation based UDA methods for improved domain adaptation performance. Specifically, we integrate the prevalent progressive pseudo-labelling techniques into the fine-tuning framework to extract fine-tuned features which are subsequently used in a state-of-the-art feature transformation based domain adaptation method SPL (Selective Pseudo-Labeling). Thorough experiments with multiple deep models including ResNet-50/101 and DeiT-small/base are conducted to demonstrate the combination of fine-tuned features and SPL can achieve state-of-the-art performance on several benchmark datasets. Qian Wang 0017, Fan-Lin Meng, Toby P. Breckon |
IJCNN | 1 |
| 2023 | Generalized zero-shot domain adaptation via coupled conditional variational autoencodersabstractDomain adaptation aims to exploit useful information from the source domain where annotated training data are easier to obtain to address a learning problem in the target domain where only limited or even no annotated data are available. In classification problems, domain adaptation has been studied under the assumption all classes are available in the target domain regardless of the annotations. However, a common situation where only a subset of classes in the target domain are available has not attracted much attention. In this paper, we formulate this particular domain adaptation problem within a generalized zero-shot learning framework by treating the labelled source-domain samples as semantic representations for zero-shot learning. For this novel problem, neither conventional domain adaptation approaches nor zero-shot learning algorithms directly apply. To solve this problem, we present a novel Coupled Conditional Variational Autoencoder (CCVAE) which can generate synthetic target-domain image features for unseen classes from real images in the source domain. Extensive experiments have been conducted on three domain adaptation datasets including a bespoke X-ray security checkpoint dataset to simulate a real-world application in aviation security. The results demonstrate the effectiveness of our proposed approach both against established benchmarks and in terms of real-world applicability. Qian Wang 0017, Toby P. Breckon |
Neural Networks | 1 |
| 2023 | Data augmentation with norm-AE and selective pseudo-labelling for unsupervised domain adaptationabstractWe address the Unsupervised Domain Adaptation (UDA) problem in image classification from a new perspective. In contrast to most existing works which either align the data distributions or learn domain-invariant features, we directly learn a unified classifier for both the source and target domains in the high-dimensional homogeneous feature space without explicit domain alignment. To this end, we employ the effective Selective Pseudo-Labelling (SPL) technique to take advantage of the unlabelled samples in the target domain. Surprisingly, data distribution discrepancy across the source and target domains can be well handled by a computationally simple classifier (e.g., a shallow Multi-Layer Perceptron) trained in the original feature space. Besides, we propose a novel generative model norm-AE to generate synthetic features for the target domain as a data augmentation strategy to enhance the classifier training. Experimental results on several benchmark datasets demonstrate the pseudo-labelling strategy itself can lead to comparable performance to many state-of-the-art methods whilst the use of norm-AE for feature augmentation can further improve the performance in most cases. As a result, our proposed methods (i.e. naive-SPL and norm-AE-SPL) can achieve comparable performance with state-of-the-art methods with the average accuracy of 93.4% and 90.4% on Office-Caltech and ImageCLEF-DA datasets, and achieve competitive performance on Digits, Office31 and Office-Home datasets with the average accuracy of 97.2%, 87.6% and 68.6% respectively. Qian Wang 0017, Fan-Lin Meng, Toby P. Breckon |
Neural Networks | 1 |
| 2022 | On Data Annotation Efficiency for Image Based Crowd CountingabstractCrowd counting aims at automatically estimating the number of persons in still images. It has attracted much attention due to its potential usage in surveillance, intelligent transportation and many other scenarios. In the recent decade, most researchers have been focusing on the design of novel deep learning models for improved crowd counting performance. Such attempts include proposing advanced architectures of deep neural networks, using different training strategies and loss functions. Other than the capabilities of models, the crowd counting performance is also determined by the quantity and the quality of training data. Whilst the deep models are data-hungry and better performance can usually be expected with more training data, annotating images for training is time-consuming and expensive in real-world applications. In this work, we focus on the efficiency of data annotation for crowd counting. By varying the number of annotated images and the number of annotated points (one point is annotated per person head) for training, our experimental results demonstrate it is more efficient to annotate a small number of points per image across a large number of images for training. Based on this conclusion, we present a novel adaptive scaling mechanism for data augmentation to diversify the training images without extra annotation cost. The mechanism is proved effective via thorough experiments. Tianfang Ma, Shuoyan Liu, Qian Wang 0017 |
VCIP | 3 |
| 2022 | A benchmark for multi-class object counting and size estimation using deep convolutional neural networks
Zixu Liu, Qian Wang 0017, Fan-Lin Meng |
Eng. Appl. Artif. Intell. | 2 |
| 2022 | Cross-domain structure preserving projection for heterogeneous domain adaptation
Qian Wang 0017, Toby P. Breckon |
Pattern Recognit. | 1 |
| 2022 | Crowd Counting via Segmentation Guided Attention Networks and Curriculum LossabstractAutomatic crowd behaviour analysis is an important task for intelligent transportation systems to enable effective flow control and dynamic route planning for varying road participants. Crowd counting is one of the keys to automatic crowd behaviour analysis. Crowd counting using deep convolutional neural networks (CNN) has achieved encouraging progress in recent years. Researchers have devoted much effort to the design of variant CNN architectures and most of them are based on the pre-trained VGG16 model. Due to the insufficient expressive capacity, the backbone network of VGG16 is usually followed by another cumbersome network specially designed for good counting performance. Although VGG models have been outperformed by Inception models in image classification tasks, the existing crowd counting networks built with Inception modules still only have a small number of layers with basic types of Inception modules. To fill in this gap, in this paper, we firstly benchmark the baseline Inception-v3 model on commonly used crowd counting datasets and achieve surprisingly good performance comparable with or better than most existing crowd counting models. Subsequently, we push the boundary of this disruptive work further by proposing a Segmentation Guided Attention Network (SGANet) with Inception-v3 as the backbone and a novel curriculum loss for crowd counting. We conduct thorough experiments to compare the performance of our SGANet with prior arts and the proposed model can achieve state-of-the-art performance with MAE of 57.6, 6.3 and 87.6 on ShanghaiTechA, ShanghaiTechB and UCF_QNRF, respectively. Qian Wang 0017, Toby P. Breckon |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2021 | Source Class Selection With Label Propagation For Partial Domain AdaptationabstractIn traditional unsupervised domain adaptation problems, the target domain is assumed to share the same set of classes as the source domain. In practice, there exist situations where target-domain data are from only a subset of source-domain classes and it is not known which classes the target-domain data belong to since they are unlabeled. This problem has been formulated as Partial Domain Adaptation (PDA) in the literature and is a challenging task due to the negative transfer issue (i.e. source-domain data belonging to the irrelevant classes harm the domain adaptation). We address the PDA problem by detecting the outlier classes in the source domain progressively. As a result, the PDA is boiled down to an easier unsupervised domain adaptation problem which can be solved without the issue of negative transfer. Specifically, we employ the locality preserving projection to learn a latent common subspace in which a label propagation algorithm is used to label the target-domain data. The outlier classes can be detected if no target-domain data are labeled as these classes. We remove the detected outlier classes from the source domain and repeat the process for multiple iterations until convergence. Experimental results on commonly used datasets Office31 and Office-Home demonstrate our proposed method achieves state-of-the-art performance with an average accuracy of 98.1% and 75.4% respectively. Qian Wang 0017, Toby P. Breckon |
ICIP | 1 |
| 2021 | Contraband Materials Detection Within Volumetric 3D Computed Tomography Baggage Security Screening ImageryabstractAutomatic prohibited object detection within 2D/3D X-ray Computed Tomography (CT) has been studied in literature to enhance the aviation security screening at checkpoints. Deep Convolutional Neural Networks (CNN) have demonstrated superior performance in 2D X-ray imagery. However, there exists very limited proof of how deep neural networks perform in materials detection within volumetric 3D CT baggage screening imagery. We attempt to close this gap by applying Deep Neural Networks in 3D contraband substance detection based on their material signatures. Specifically, we formulate it as a 3D semantic segmentation problem to identify material types for all voxels based on which contraband materials can be detected. To this end, we firstly investigate 3D CNN based semantic segmentation algorithms such as 3D U-Net and its variants. In contrast to the original dense representation form of volumetric 3D CT data, we propose to convert the CT volumes into sparse point clouds which allows the use of point cloud processing approaches such as PointNet++ towards more efficient processing. Experimental results on a publicly available dataset (NEU ATR) demonstrate the effectiveness of both 3D U-Net and PointNet++ in materials detection in 3D CT imagery for baggage security screening. Qian Wang 0017, Toby P. Breckon |
ICMLA | 1 |
| 2021 | A telehealth framework for dementia care: an ADLs patterns recognition model for patients based on NILMabstractThe ageing of the population and the increasing number of patients with dementia in modern society undoubtedly put tremendous pressure on the medical system. Providing telehealth care for potential patients and patients with dementia can reduce the burden on both the health system and care-givers. This paper describes a telehealth framework for dementia early detection and dementia care. Specifically, we propose an improved deep neural network model for Non-Intrusive Load Monitoring (NILM), which disaggregates the household's overall energy usage into those of individual appliances based on the sequence-to-point model and transfer learning. The daily behaviour regularities of patients are then inferred by combining principal component analysis and$K$-means clustering based on the disaggregated appliance-level consumptions. Experiments show that the proposed model can significantly improve training efficiency and maintain load disaggregation accuracy, and the inferred behaviour regularities have great potential to be used as useful inputs and prior knowledge to the dementia condition detection platform for early detection and real-time monitoring of patient's conditions. Shuang Dai, Qian Wang 0017, Fan-Lin Meng |
IJCNN | 2 |
| 2021 | On the Evaluation of Semi-Supervised 2D Segmentation for Volumetric 3D Computed Tomography Baggage Security ScreeningabstractWe address the automatic contraband material detection problem within volumetric 3D Computed Tomography (CT) data for baggage security screening. Distinct from the prohibited item detection using object detection techniques, contraband material detection is usually formulated as a segmentation problem due to the variations of their potential appearances and shapes. Previous studies have employed either morphological operation based traditional methods or 3D Convolutional Neural Networks (CNN) for 3D segmentation towards target material detection within volumetric 3D CT baggage security screening imagery. In this work, we investigate the effectiveness of 2D semantic segmentation techniques in this 3D CT segmentation problem. Specifically, we extract 2D slices from three planes of the 3D CT volumes and train a 2D segmentation model which is subsequently used to predict segmentation results for all the slices from a given test CT volume. Moreover, we also evaluate how the performance is affected when using a reduced number of annotated slices for training. As a result, it is demonstrated reasonable performance can be achieved with very limited annotated slices (1–2) per CT volume during training. Finally, we propose a semi-supervised learning framework for 3D CT segmentation. Using only 1/128 of the total number of annotated slices, our framework can achieve comparable performance with full supervision. Qian Wang 0017, Toby P. Breckon |
IJCNN | 1 |
| 2021 | Dynamic clustering analysis for driving styles identification
Maria Valentina Niño de Zepeda, Fan-Lin Meng, Jinya Su, Xiaojun Zeng, Qian Wang 0017 |
Eng. Appl. Artif. Intell. | 5 |
| 2020 | Unsupervised Domain Adaptation via Structured Prediction Based Selective Pseudo-LabelingabstractUnsupervised domain adaptation aims to address the problem of classifying unlabeled samples from the target domain whilst labeled samples are only available from the source domain and the data distributions are different in these two domains. As a result, classifiers trained from labeled samples in the source domain suffer from significant performance drop when directly applied to the samples from the target domain. To address this issue, different approaches have been proposed to learn domain-invariant features or domain-specific classifiers. In either case, the lack of labeled samples in the target domain can be an issue which is usually overcome by pseudo-labeling. Inaccurate pseudo-labeling, however, could result in catastrophic error accumulation during learning. In this paper, we propose a novel selective pseudo-labeling strategy based on structured prediction. The idea of structured prediction is inspired by the fact that samples in the target domain are well clustered within the deep feature space so that unsupervised clustering analysis can be used to facilitate accurate pseudo-labeling. Experimental results on four datasets (i.e. Office-Caltech, Office31, ImageCLEF-DA and Office-Home) validate our approach outperforms contemporary state-of-the-art methods. Qian Wang 0017, Toby P. Breckon |
AAAI | 1 |
| 2020 | Multi-Class 3D Object Detection Within Volumetric 3D Computed Tomography Baggage Security Screening ImageryabstractAutomatic detection of prohibited objects within passenger baggage is important for aviation security. X-ray Computed Tomography (CT) based 3D imaging is widely used in airports for aviation security screening whilst prior work on automatic prohibited item detection focus primarily on 2D X-ray imagery. Whilst some prior work has proven the possibility of extending deep convolutional neural networks (CNN) based automatic prohibited item detection from 2D X-ray imagery to volumetric 3D CT baggage security screening imagery, it focuses on the detection of one specific type of objects (e.g., either bottles or handguns). As a result, multiple models are needed if more than one type of prohibited item is required to be detected in practice. In this paper, we consider the detection of multiple object categories of interest using one unified framework. To this end, we formulate a more challenging multi-class 3D object detection problem within 3D CT imagery and propose a viable solution (3D RetinaNet) to tackle this problem. To enhance the performance of detection we investigate a variety of strategies including data augmentation and varying backbone networks. Experimentation carried out to provide both quantitative and qualitative evaluations of the proposed approach to multi-class 3D object detection within 3D CT baggage security screening imagery. Experimental results demonstrate the combination of the 3D RetinaNet and a series of favorable strategies can achieve a mean Average Precision (mAP) of 65.3% over five object classes (i.e. bottles, handguns, binoculars, glock frames, iPods). The overall performance is affected by the poor performance on glock frames and iPods due to the lack of data and their resemblance with the baggage clutter. Qian Wang 0017, Neelanjan Bhowmik, Toby P. Breckon |
ICMLA | 1 |
| 2020 | On the Evaluation of Prohibited Item Classification and Detection in Volumetric 3D Computed Tomography Baggage Security Screening ImageryabstractX-ray Computed Tomography (CT) based 3D imaging is widely used in airports for aviation security screening whilst prior work on prohibited item detection focuses primarily on 2D X-ray imagery. In this paper, we aim to evaluate the possibility of extending the automatic prohibited item detection from 2D X-ray imagery to volumetric 3D CT baggage security screening imagery. To these ends, we take advantage of 3D Convolutional Neural Networks (CNN) and popular object detection frameworks such as RetinaNet and Faster R-CNN in our work. As the first attempt to use 3D CNN for volumetric 3D CT baggage security screening, we first evaluate different CNN architectures on the classification of isolated prohibited item volumes and compare against traditional methods which use hand-crafted features. Subsequently, we evaluate object detection performance of different architectures on volumetric 3D CT baggage images. The results of our experiments on Bottle and Handgun datasets demonstrate that 3D CNN models can achieve comparable performance (~ 98% true positive rate and ~1.5% false positive rate) to traditional methods but require significantly less time for inference (0.014s per volume). Furthermore, the extended 3D object detection models achieve promising performance in detecting prohibited items within volumetric 3D CT baggage imagery with ~76% mAP for bottles and ~88% mAP for handguns, which shows both the challenge and promise of such threat detection within 3D CT X-ray security imagery. Qian Wang 0017, Neelanjan Bhowmik, Toby P. Breckon |
IJCNN | 1 |
| 2020 | Multi-label zero-shot human action recognition via joint latent ranking embedding
Qian Wang 0017, Ke Chen 0001 |
Neural Networks | 1 |
| 2019 | A Baseline for Multi-Label Image Classification Using an Ensemble of Deep Convolutional Neural NetworksabstractRecent studies on multi-label image classification have focused on designing more complex architectures of deep neural networks such as the use of attention mechanisms and region proposal networks. Although performance gains have been reported, the backbone deep models of the proposed approaches and the evaluation metrics employed in different works vary, making it difficult to compare fairly. Moreover, due to the lack of properly investigated baselines, the advantage introduced by the proposed techniques are often ambiguous. To address these issues, we make a thorough investigation of the mainstream deep convolutional neural network architectures for multi-label image classification and present a strong baseline. With the use of proper data augmentation techniques and model ensembles, the basic deep architectures can achieve better performance than many existing more complex ones on three benchmark datasets, providing great insight for the future studies on multi-label image classification. Qian Wang 0017, Toby P. Breckon |
ICIP | 1 |
| 2019 | Unifying Unsupervised Domain Adaptation and Zero-Shot Visual RecognitionabstractUnsupervised domain adaptation aims to transfer knowledge from a source domain to a target domain so that the target domain data can be recognized without any explicit labelling information for this domain. One limitation of the problem setting is that testing data (despite no labels) from the target domain is needed during training, which prevents the trained model being directly applied to classify newly arrived test instances. We formulate a new cross-domain classification problem arising from real-world scenarios where labelled data are available for a subset of classes (known classes) in the target domain, and we expect to recognize new samples belonging to any class (known and unseen classes) once the model is learned. This is a generalized zero-shot learning problem where the side information comes from the source domain in the form of labelled samples instead of class-level semantic representations commonly used in traditional zero-shot learning. We present a unified domain adaptation framework for both unsupervised and zero-shot learning conditions. Our approach learns a joint subspace from source and target domains so that the projections of both data in the subspace can be domain invariant and easily separable. We use the supervised locality preserving projection (SLPP) as the enabling technique and conduct experiments under both unsupervised and zero-shot learning conditions, achieving state-of-the-art results on three domain adaptation benchmark datasets: Office-Caltech, Office31 and Office-Home. Qian Wang 0017, Penghui Bu, Toby P. Breckon |
IJCNN | 1 |
| 2017 | Alternative Semantic Representations for Zero-Shot Human Action Recognition
Qian Wang 0017, Ke Chen 0001 |
ECML/PKDD (1) | 1 |
| 2017 | Zero-Shot Visual Recognition via Bidirectional Latent Embedding
Qian Wang 0017, Ke Chen 0001 |
Int. J. Comput. Vis. | 1 |