VLDB 2026 Research / reviewers in the wild / expert
Dongliang Chang
dblp:236/3116
· DBLP profile ↗
41ranked-venue papers
6as first author
34since 2021 · last 2026
0000-0002-4081-3001ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 31 · 5 first-author · 24 since 2021Artificial intelligence and machine learning · 21 · 3 first-author · 18 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MIX-based Foreground and Background Patch Augmentation Guided by Physics and Material Properties for X-ray DetectionabstractThe performance of deep learning-based models for X-ray prohibited item detection heavily relies on large-scale, diverse datasets, which are often unavailable. While data augmentation offers a promising solution, prevalent methods ignore the fundamental principles of X-ray imaging, leading to artifacts such as distorted material properties and unnatural thickness perturbations. To bridge this gap, we present MIX, a physics-grounded data augmentation pipeline. The core idea of MIX is to manipulate image attributes in a way that reflects real-world physical variations. Our contributions are twofold: (1) To address material ambiguity, MIX modulates foreground pseudo-colors by directly manipulating hue and saturation, informed by the relationship between color and effective atomic number. This forces the model to learn more robust material representations. (2) To simulate variations in object density and thickness, MIX introduces a novel thickness perturbation technique based on X-ray attenuation principles. This significantly improves the model’s adaptability to geometric changes. Our proposed method seamlessly integrates with existing detectors and yields substantial performance gains across multiple benchmarks. Our work not only provides an effective augmentation solution but also highlights the critical need for domain-specific approaches in X-ray computer vision. Dongliang Chang, Yujun Tong, Zhanyu Ma |
WACV | 2 |
| 2026 | APENet: Task-aware adaptation prototype evolution network for few-shot semantic segmentation
Zhaobin Chang, Xiong Gao, Dongliang Chang, Yande Li, Yonggang Lu |
Expert Syst. Appl. | 3 |
| 2026 | FCNet: Extracting undistorted images for fine-grained image classification
Junhan Chen, Dongliang Chang, Yujun Tong, Ruoyi Du, Yingqing Wang, Zhanyu Ma, Yi-Zhe Song |
Neurocomputing | 2 |
| 2026 | Toward Generalizable Forgery Detection and ReasoningabstractAccurate and interpretable detection of AI-generated images is essential for mitigating risks associated with AI misuse. However, the substantial domain gap among generative models makes it challenging to develop a generalizable forgery detection model. Moreover, since every pixel in an AI-generated image is synthesized, traditional saliency-based forgery explanation methods are not well suited for this task. To address these challenges, we formulate detection and explanation as a unified Forgery Detection and Reasoning task (FDR-Task), leveraging Multi-Modal Large Language Models (MLLMs) to provide accurate detection through reliable reasoning over forgery attributes. To facilitate this task, we introduce the Multi-Modal Forgery Reasoning dataset (MMFR-Dataset), a large-scale dataset containing 120K images across 10 generative models, with 378K reasoning annotations on forgery attributes, enabling comprehensive evaluation of the FDR-Task. Furthermore, we propose FakeReasoning, a forgery detection and reasoning framework with three key components: 1) a dual-branch visual encoder that integrates CLIP and DINO to capture both high-level semantics and low-level artifacts; 2) a Forgery-Aware Feature Fusion Module that leverages DINO's attention maps and cross-attention mechanisms to guide MLLMs toward forgery-related clues; 3) a Classification Probability Mapper that couples language modeling and forgery detection, enhancing overall performance. Experiments across multiple generative models demonstrate that FakeReasoning not only achieves robust generalization but also outperforms state-of-the-art methods on both detection and reasoning tasks. The code is available at: https://github.com/PRIS-CV/FakeReasoning. Yueying Gao, Dongliang Chang, Bingyao Yu, Haotian Qin, Muxi Diao, Lei Chen 0069, Kongming Liang, Zhanyu Ma |
IEEE Trans. Image Process. | 2 |
| 2025 | Self-randomized focuses effectively boost metric-based few-shot classifiers
Zhen Li 0026, Zhongyuan Liu, Dongliang Chang, Aneeshan Sain, Zhanyu Ma, Jing-Hao Xue, Yi-Zhe Song |
Pattern Recognit. | 3 |
| 2025 | Adaptive Multi-Resolution Feature Fusion for Fine-Grained Visual ClassificationabstractDespite significant progress, the shortage of labeled data and expert knowledge remains a challenge for Fine-grained Visual Classification (FGVC). Some multi-source approaches that incorporate additional modalities, such as sound or bounding boxes, show promise for data enrichment but introduce added complexity to data collection. In this paper, we pose the question: can multi-source capabilities be achieved solely with existing images? The answer, confirmed by a pilot study, is affirmative. By analyzing the probability distribution of model output with different resolutions image, we find that complementary information beneficial to FGVC exists among images of different resolutions. Although the classification accuracy of low-resolution images is lower than high-resolution images, it can provide additional information for high-resolution input images. We designed a naive baseline that uses mixed training of multi-resolution images. Through the experimental results of the baseline, we find that i) not all low-resolution images are beneficial, and ii) adaptively selecting low-resolution images is what we need. Therefore, we proposed a meta-learning-based adaptive “resolution” pooling layer. Through the pooling operation, the features of low-resolution images are obtained from high-resolution images, and the most appropriate complementary features are selected for the features of high-resolution images through the gating mechanism, which enables the model to fully and autonomously exploit the complementary information. Experimental results on three FGVC datasets validate the effectiveness of our proposed method. Our code is available athttps://github.com/PRIS-CV/Adaptive-Multi-Resolution-Feature-Fusion. Dongliang Chang, Ruoyi Du, Yi-Zhe Song, Zhanyu Ma |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | Query-Aware Cross-Mixup and Cross-Reconstruction for Few-Shot Fine-Grained Image ClassificationabstractFew-shot fine-grained image classification is prominent but challenging in computer vision, aiming to distinguish sub-classes under the same parent class but with only a few labeled support samples. Data augmentation techniques were explored to address the few-shot issue, but they often fail to mitigate the bias between support and query samples. Therefore, in this paper we propose a query-aware cross-mixup and cross-reconstruction method to address both few-shot and fine-grained issues. Specifically, in the training phase, we randomly select query samples and mix them with the support samples from the same class to augment the support set. This first strategy ensures the augmented support set query-aware within each sub-class. Then, we reconstruct both query samples and support samples from both original and cross-mixed support samples, thus leveraging both cross-reconstruction and self-reconstruction to enhance classification. This second strategy, enabling the reconstruction also query-aware, further mitigates the bias between support and query samples, leading to more reliable generalization. We evaluate our proposed method on four widely used few-shot fine-grained image classification datasets, and experimental results demonstrate its effectiveness in achieving the state-of-the-art classification performance. Dongliang Chang, Rui Zhu 0006, Zhanyu Ma, Jing-Hao Xue |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | Reserve to Adapt: Mining Inter-Class Relations for Open-Set Domain AdaptationabstractOpen-Set Domain Adaptation (OSDA) aims at adapting a model trained on a labelled source domain, to an unlabeled target domain that is corrupted with unknown classes. The key challenge inherent to this open-set setting is therefore how best to avoid the negative transfer incurred by unknown classes during model adaptation. Most existing works tackle this challenge by simply pushing the entire unknown classes away. In this paper, we take a different stance - instead of addressing these unknown classes as a single entity, we "reserve" in-between spaces for their subsets in the learned embedding. Our key finding is that the inter-class relations learned off the source domain, can help to enforce class separations in the target domain - thereby reserving spaces for unknown classes. More specifically, we first prep the "reservation" by tightening the known-class representations while enlarging their inter-class margin. We then learn soft-label prototypes in the source domain to facilitate the discrimination of known and unknown samples in the target domain. It follows that these two steps are iterated at each epoch in a mutually beneficial manner - better discrimination of unknown samples helps with space reservation, and vice versa. We show state-of-the-art results on four standard OSDA datasets, Office-31, Office-Home, VisDA and ImageCLEF, and conduct further analysis to help understand our method. Codes are available at: https://github.com/PRIS-CV/Reserve_to_Adapt. Yujun Tong, Dongliang Chang, Da Li 0001, Kongming Liang, Zhongjiang He, Yi-Zhe Song, Zhanyu Ma |
IEEE Trans. Image Process. | 2 |
| 2024 | Class-Aware Contrastive Learning for Fine-Grained Skeleton-Based Action Recognition
Xinyu Bian, Dongliang Chang, Zhongjiang He, Kongming Liang, Zhanyu Ma |
ACCV (1) | 2 |
| 2024 | Hierarchical Prompting for Diffusion Classifiers
Wenxin Ning, Dongliang Chang, Yujun Tong, Zhongjiang He, Kongming Liang, Zhanyu Ma |
ACCV (8) | 2 |
| 2024 | DemoFusion: Democratising High-Resolution Image Generation With No $$$abstractHigh-resolution image generation with Generative Artificial Intelligence (GenAl) has immense potential but, due to the enormous capital investment required for training, it is increasimgly centralised to a few large corporations, and hidden behind paywalls. This paper aims to democratise high-resolution GenAl by advancing the frontier of high-resolution generation while remaining accessible to a broad audience. We demonstrate that existing Latent Diffusion Models (LDMs) possess untapped potential for higher-resolution image generation. Our novel DemoFusion framework seamlessly extends open-source GenAl models, employing Progressive Upscaling, Skip Residual, and Di-lated Sampling mechanisms to achieve higher-resolution image generation. The progressive nature of DemoFusion requires more passes, but the intermediate results can serve as “previews”, facilitating rapid prompt iteration. Ruoyi Du, Dongliang Chang, Timothy M. Hospedales, Yi-Zhe Song, Zhanyu Ma |
CVPR | 2 |
| 2024 | Channel-Spatial Support-Query Cross-Attention for Fine-Grained Few-Shot Image ClassificationabstractFew-shot fine-grained image classification aims to use only few labelled samples to successfully recognize subtle sub-classes within the same parent class. This task is extremely challenging, due to the co-occurrence of large inter-class similarity, low intra-class similarity, and only few labelled samples. In this paper, to address these challenges, we propose a new Channel-Spatial Cross-Attention Module (CSCAM), which can effectively drive a model to extract discriminative fine-grained feature representations with only few shots. CSCAM collaboratively integrates a channel cross-attention module and a spatial cross-attention module, for the attentions across support and query samples. In addition, to fit for the characteristics of fine-grained images, a support averaging method is proposed in CSCAM to reduce the intra-class distance and increase the inter-class distance. Extensive experiments on four few-shot fine-grained classification datasets validate the effectiveness of CSCAM. Furthermore, CSCAM is a plug-and-play module, conveniently enabling effective improvement of state-of-the-art methods for few-shot fine-grained image classification. Shicheng Yang, Dongliang Chang, Zhanyu Ma, Jing-Hao Xue |
ACM Multimedia | 3 |
| 2024 | Semi-Supervised Learning for FGVC With Out-of-Category DataabstractDespite great strides made on fine-grained visual classification (FGVC), current methods are still heavily reliant on fully-supervised paradigms where ample expert labels are called for. Semi-supervised learning (SSL) techniques, acquiring knowledge from unlabeled data, provide a considerable means forward and have shown great promise for coarse-grained problems. However, exiting SSL paradigms mostly assume in-category (i.e., category-aligned) unlabeled data, which hinders their effectiveness when re-proposed on FGVC. In this paper, we put forward a novel design specifically aimed at making out-of-category data work for semi-supervised FGVC. We work off an important assumption that all fine-grained categories naturally follow a hierarchical structure (e.g., the phylogenetic tree of "Aves" that covers all bird species). It follows that, instead of operating on individual samples, we can instead predict sample relations within this tree structure as the optimization goal of SSL. Beyond this, we further introduced two strategies uniquely brought by these tree structures to achieve inter-sample consistency regularization and reliable pseudo-relation. Our experimental results reveal that (i) the proposed method yields good robustness against out-of-category data, and (ii) it can be equipped with prior arts, boosting their performance thus yielding state-of-the-art results. Ruoyi Du, Dongliang Chang, Zhanyu Ma, Kongming Liang, Yi-Zhe Song, Jun Guo 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2024 | Bi-Directional Ensemble Feature Reconstruction Network for Few-Shot Fine-Grained ClassificationabstractThe main challenge for fine-grained few-shot image classification is to learn feature representations with higher inter-class and lower intra-class variations, with a mere few labelled samples. Conventional few-shot learning methods however cannot be naively adopted for this fine-grained setting - a quick pilot study reveals that they in fact push for the opposite (i.e., lower inter-class variations and higher intra-class variations). To alleviate this problem, prior works predominately use a support set to reconstruct the query image and then utilize metric learning to determine its category. Upon careful inspection, we further reveal that such unidirectional reconstruction methods only help to increase inter-class variations and are not effective in tackling intra-class variations. In this paper, we introduce a bi-reconstruction mechanism that can simultaneously accommodate for inter-class and intra-class variations. In addition to using the support set to reconstruct the query set for increasing inter-class variations, we further use the query set to reconstruct the support set for reducing intra-class variations. This design effectively helps the model to explore more subtle and discriminative features which is key for the fine-grained problem in hand. Furthermore, we also construct a self-reconstruction module to work alongside the bi-directional module to make the features even more discriminative. We introduce the snapshot ensemble method in the episodic learning strategy - a simple trick to further improve model performance without increasing training costs. Experimental results on three widely used fine-grained image classification datasets, as well as general and cross-domain few-shot image datasets, consistently show considerable improvements compared with other methods. Jijie Wu, Dongliang Chang, Aneeshan Sain, Zhanyu Ma, Jie Cao 0014, Jun Guo 0002, Yi-Zhe Song |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2024 | Mind the Gap: Open Set Domain Adaptation via Mutual-to-Separate FrameworkabstractUnsupervised domain adaptation aims to leverage labeled data from a source domain to learn a classifier for an unlabeled target domain. Amongst its many variants, open set domain adaptation (OSDA) is perhaps the most challenging one, as it further assumes the presence of unknown classes in the target domain. In this paper, we study OSDA with a particular focus on enriching its ability to traverse across larger domain gaps, and we show that existing state-of-the-art methods suffer a considerable performance drop in the presence of larger domain gaps, especially on a new dataset (PACS) that we re-purposed for OSDA. Exploring this is pivotal for OSDA as with increasing domain shift, identifying unknown samples in the target domain becomes harder for the model, thus making negative transfer between source and target domains more challenging. Accordingly, we propose a Mutual-to-Separate (MTS) framework to address the larger domain gaps. Essentially we design two networks – (a) Sample Separation Network (SSN): which is trained to learn a hyperplane for separating unknown samples from known ones, and (b) Distribution Matching Network (DMN): which is trained to maximise domain confusion between source and target domains without unknown samples under the guidance of the SSN. The key insight lies in how we exploit the mutually beneficial information between these two networks. On closer observation, we see that SSN can reveal which samples in the target domain belong to the unknown class by instance weighting whereas, DMN pushes apart the samples that most likely belong to the unknown class in the target domain, which in turn reduces the difficulty of SSN in identifying unknown samples. It follows that (a) and (b) will mutually supervise each other and alternate until convergence, which can better align the source and target domains in the shared label space. Extensive experiments on five datasets (Office-31, Office-Home, PACS, VisDA, andmini_DomainNet) demonstrate the efficiency of the proposed method. Detailed ablation experiments also validate the effectiveness of each component and the generality of the proposed framework. Codes are available at: https://github.com/PRIS-CV/Mutual-to-Separate. Dongliang Chang, Aneeshan Sain, Zhanyu Ma, Yi-Zhe Song, Ruiping Wang 0001, Jun Guo 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2024 | CreativeSeg: Semantic Segmentation of Creative SketchesabstractThe problem of sketch semantic segmentation is far from being solved. Despite existing methods exhibiting near-saturating performances on simple sketches with high recognisability, they suffer serious setbacks when the target sketches are products of an imaginative process with high degree of creativity. We hypothesise that human creativity, being highly individualistic, induces a significant shift in distribution of sketches, leading to poor model generalisation. Such hypothesis, backed by empirical evidences, opens the door for a solution that explicitly disentangles creativity while learning sketch representations. We materialise this by crafting a learnable creativity estimator that assigns a scalar score of creativity to each sketch. It follows that we introduce CreativeSeg, a learning-to-learn framework that leverages the estimator in order to learn creativity-agnostic representation, and eventually the downstream semantic segmentation task. We empirically verify the superiority of CreativeSeg on the recent "Creative Birds" and "Creative Creatures" creative sketch datasets. Through a human study, we further strengthen the case that the learned creativity score does indeed have a positive correlation with the subjective creativity of human. Codes are available at https://github.com/PRIS-CV/Sketch-CS. Yixiao Zheng, Kaiyue Pang, Ayan Das 0003, Dongliang Chang, Yi-Zhe Song, Zhanyu Ma |
IEEE Trans. Image Process. | 4 |
| 2023 | Bi-directional Feature Reconstruction Network for Fine-Grained Few-Shot Image ClassificationabstractThe main challenge for fine-grained few-shot image classification is to learn feature representations with higher inter-class and lower intra-class variations, with a mere few labelled samples. Conventional few-shot learning methods however cannot be naively adopted for this fine-grained setting -- a quick pilot study reveals that they in fact push for the opposite (i.e., lower inter-class variations and higher intra-class variations). To alleviate this problem, prior works predominately use a support set to reconstruct the query image and then utilize metric learning to determine its category. Upon careful inspection, we further reveal that such unidirectional reconstruction methods only help to increase inter-class variations and are not effective in tackling intra-class variations. In this paper, we for the first time introduce a bi-reconstruction mechanism that can simultaneously accommodate for inter-class and intra-class variations. In addition to using the support set to reconstruct the query set for increasing inter-class variations, we further use the query set to reconstruct the support set for reducing intra-class variations. This design effectively helps the model to explore more subtle and discriminative features which is key for the fine-grained problem in hand. Furthermore, we also construct a self-reconstruction module to work alongside the bi-directional module to make the features even more discriminative. Experimental results on three widely used fine-grained image classification datasets consistently show considerable improvements compared with other methods. Codes are available at: https://github.com/PRIS-CV/Bi-FRN. Jijie Wu, Dongliang Chang, Aneeshan Sain, Zhanyu Ma, Jie Cao 0014, Jun Guo 0002, Yi-Zhe Song |
AAAI | 2 |
| 2023 | An Erudite Fine-Grained Visual Classification ModelabstractCurrent fine-grained visual classification (FGVC) models are isolated. In practice, we first need to identify the coarse-grained label of an object, then select the corresponding FGVC model for recognition. This hinders the application of FGVC algorithms in real-life scenarios. In this paper, we propose an erudite FGVC model jointly trained by several different datasets11In this paper, different datasets mean different fine-grained visual classification datasets., which can efficiently and accurately predict an object's fine-grained label across the combined label space. We found through a pilot study that positive and negative transfers co-occur when different datasets are mixed for training, i.e., the knowledge from other datasets is not always useful. Therefore, we first propose a feature disentanglement module and a feature re-fusion module to reduce negative transfer and boost positive transfer between different datasets. In detail, we reduce negative transfer by decoupling the deep features through many dataset-specific feature extractors. Subsequently, these are channel-wise re-fused to facilitate positive transfer. Finally, we propose a meta-learning based dataset-agnostic spatial attention layer to take full advantage of the multi-dataset training data, given that localisation is dataset-agnostic between different datasets. Experimental results across 11 different mixed-datasets built on four different FGVC datasets demonstrate the effectiveness of the proposed method. Furthermore, the proposed method can be easily combined with existing FGVC methods to obtain state-of-the-art results. Our code is available at https://github.com/PRIS-CV/An-Erudite-FGVC-Model. Dongliang Chang, Yujun Tong, Ruoyi Du, Timothy M. Hospedales, Yi-Zhe Song, Zhanyu Ma |
CVPR | 1 |
| 2023 | On-the-Fly Category DiscoveryabstractAlthough machines have surpassed humans on visual recognition problems, they are still limited to providing closed-set answers. Unlike machines, humans can cognize novel categories at the first observation. Novel category discovery (NCD) techniques, transferring knowledge from seen categories to distinguish unseen categories, aim to bridge the gap. However, current NCD methods assume a transductive learning and offline inference paradigm, which restricts them to a predefined query set and renders them unable to deliver instant feedback. In this paper, we study on-the-fly category discovery (OCD) aimed at making the model instantaneously aware of novel category samples (i.e., enabling inductive learning and streaming inference). We first design a hash coding-based expandable recognition model as a practical baseline. Afterwards, noticing the sensitivity of hash codes to intra-category variance, we further propose a novel Sign-Magnitude dIsentangLEment (SMILE) architecture to alleviate the disturbance it brings. Our experimental results demonstrate the superiority of SMILE against our baseline model and prior art. Our code is available at https://github.com/PRIS-CV/On-the-fly-Category-Discovery. Ruoyi Du, Dongliang Chang, Kongming Liang, Timothy M. Hospedales, Yi-Zhe Song, Zhanyu Ma |
CVPR | 2 |
| 2023 | Multi-View Active Fine-Grained Visual RecognitionabstractDespite the remarkable progress of Fine-grained visual classification (FGVC) with years of history, it is still limited to recognizing 2D images. Recognizing objects in the physical world (i.e., 3D environment) poses a unique challenge – discriminative information is not only present in visible local regions but also in other unseen views. Therefore, in addition to finding the distinguishable part from the current view, efficient and accurate recognition requires inferring the critical perspective with minimal glances. E.g., a person might recognize a "Ford sedan" with a glance at its side and then know that looking at the front can help tell which model it is. In this paper, towards FGVC in the real physical world, we put forward the problem of multi-view active fine-grained visual recognition (MAFR) and complete this study in three steps: (i) a multi-view, fine-grained vehicle dataset is collected as the testbed, (ii) a pilot experiment is designed to validate the need and research value of MAFR, (iii) a policy-gradient-based framework along with a dynamic exiting strategy is proposed to achieve efficient recognition with active view selection. Our comprehensive experiments demonstrate that the proposed method outperforms previous multi-view recognition works and can extend existing state-of-the-art FGVC methods and advanced neural networks to become "FGVC experts" in the 3D environment. Our code is available at https://github.com/PRIS-CV/MAFR. Ruoyi Du, Wenqing Yu, Heqing Wang, Ting-En Lin, Dongliang Chang, Zhanyu Ma |
ICCV | 5 |
| 2023 | Making a Bird AI Expert Work for You and MeabstractAs powerful as fine-grained visual classification (FGVC) is, responding your query with a bird name of “Whip-poor-will” or “Mallard” probably does not make much sense. This however commonly accepted in the literature, underlines a fundamental question interfacing AI and human – what constitutes transferable knowledge for human to learn from AI? This paper sets out to answer this very question using FGVC as a test bed. Specifically, we envisage a scenario where a trained FGVC model (the AI expert) functions as a knowledge provider in enabling average people (you and me) to become better domain experts ourselves,i.e.,those capable in distinguishing between “Whip-poor-will” and “Mallard”. Fig. 1 lays out our approach in answering this question. Assuming an AI expert trained using expert human labels, we ask (i) what is the best transferable knowledge we can extract from AI, and (ii) what is the most practical means to measure the gains in expertise given that knowledge? On the former, we propose to represent knowledge as highly discriminative visual regions that are expert-exclusive. For that, we devise a multi-stage learning framework, which starts with modelling visual attention of domain experts and novices separately, before discriminatively distilling their differences to acquire those exclusive to experts. For the latter, we simulate the evaluation process as a book guide to best accommodate the learning practice of that is accustomed to humans. A comprehensive human study of 15,000 trials shows our method is able to consistently improve people of divergent bird expertise to recognise once unrecognisable birds. To counter the lack of reproducibility of perceptual studies, and in turn to make a sustainable direction out of our “AI for Human” effort, we further propose a quantitative metric, namely Transferable Effective Model Attention (TEMI). TEMI acts as a crude but benchmarkable metric to replace large-scale human studies, and therefore allows future efforts in this direction to be comparable to ours. We attest to the integrity of TEMI by (i) empirically showing a strong correlation between TEMI scores and raw human study data, and (ii) its expected behaviour holds for a large body of attention models. Last but not least, our approach also leads to improved FGVC performance in the conventional benchmarking sense, when the extracted knowledge defined is utilised as means to achieve discriminative localisation. Codes and all details on the human study are available at:https://github.com/PRIS-CV/Making-a-Bird-AI-Expert-Work-for-You-and-Me. Dongliang Chang, Kaiyue Pang, Ruoyi Du, Yujun Tong, Yi-Zhe Song, Zhanyu Ma, Jun Guo 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2022 | Domain Generalization via Frequency-domain-based Feature Disentanglement and InteractionabstractAdaptation to out-of-distribution data is a meta-challenge for all statistical learning algorithms that strongly rely on the i.i.d. assumption. It leads to unavoidable labor costs and confidence crises in realistic applications. For that, domain generalization aims at mining domain-irrelevant knowledge from multiple source domains that can generalize to unseen target domains. In this paper, by leveraging the frequency domain of an image, we uniquely work with two key observations: (i) the high-frequency information of an image depicts object edge structure, which preserves high-level semantic information of the object is naturally consistent across different domains, and (ii) the low-frequency component retains object smooth structure, while this information is susceptible to domain shifts. Motivated by the above observations, we introduce (i) an encoder-decoder structure to disentangle high- and low-frequency features of an image, (ii) an information interaction mechanism to ensure the helpful knowledge from both two parts can cooperate effectively, and (iii) a novel data augmentation technique that works on the frequency domain to encourage the robustness of frequency-wise feature disentangling. The proposed method obtains state-of-the-art performance on three widely used domain generalization benchmarks (Digit-DG, Office-Home, and PACS). Jingye Wang, Ruoyi Du, Dongliang Chang, Kongming Liang, Zhanyu Ma |
ACM Multimedia | 3 |
| 2022 | Complex Scenario-Oriented Fine-Grained Visual Classification PlatformabstractIn recent years, fine-grained visual classification (FGVC) algorithms have achieved excellent performance across a variety of datasets. However, it is still rare to see these algorithms applied in daily life. The main reasons for this are i) the algorithms are developed based on different design guidelines and cannot be deployed in the same environment; ii) there is not a simple and efficient platform to present the algorithm's results to the user - the accuracy is meaningless to the users. To address the above problem, we built a complex scenario-oriented fine-grained visual classification platform. The platform consists of a PyTorch-based fine-grained visual recognition algorithm library (FGL) and a WeChat applet-based user interaction module (WEM). We can quickly develop new algorithms or readily apply existing algorithms in the same environment through FGL. Driven by FGL, the WEM enables users to achieve fine-grained recognition of complex scenes interactively. In addition to showing the user the fine-grained labels of objects, we will also show how the model makes decisions to help the user master the ability to recognise the fine-grained object so that everyone can become a domain expert. A video demo shows an example of the proposed platform in a real-world scenario: https://reurl.cc/rRZE7O. Dongliang Chang, Junhan Chen, Ruoyi Du, Wenqing Yu, Yujun Tong, Kongming Liang, Yi-Zhe Song, Zhanyu Ma |
MMSP | 1 |
| 2022 | Multi-modal Human-machine Conversation System for Real Physical WorldabstractEnabling machines to process multi-modal information and understand the real physical world is an important step to achieving free human-machine conversation. However, previous human-machine conversation systems are mostly limited to single-modal (e.g., chat robot), single-round (e.g., visual Q & A), and static visual information (e.g., visual dialogue). To address the above problem, we develop a multi-modal human-machine conversation system for specific application scenarios. The system includes two modules: (i) an interactive visual grounding module that can actively disambiguate user's queries, and (ii) an interactive fine-grained recognition module that can model objects in the 3D environment and actively ask for missing visual information. A video demo of our system under the automobile sales scenario can be found here11https://drive.google.com/file/d/1IfBsMKq55ryLOZchIT6G-R5CyWnC3Skiew?usp=sharing. Shibo Nie, Mandan Guan, Ruoyi Du, Dongliang Chang, Kongming Liang, Zhanyu Ma |
MMSP | 6 |
| 2022 | Cross-Layer Feature based Multi-Granularity Visual ClassificationabstractIn contrast to traditional fine-grained visual clas-sification, multi-granularity visual classification is no longer limited to identifying the different sub-classes belonging to the same super-class (e.g., bird species, cars, and aircraft models). Instead, it gives a sequence of labels from coarse to fine (e.g., Passeriformes → Corvidae → Fish Crow), which is more convenient in practice. The key to solving this task is how to use the relationships between the different levels of labels to learn feature representations that contain different levels of granularity. Interestingly, the feature pyramid structure naturally implies different granularity of feature representation, with the shallow layers representing coarse-grained features and the deep layers representing fine-grained features. Therefore, in this paper, we exploit this property of the feature pyramid structure to decouple features and obtain feature representations corre-sponding to different granularities. Specifically, we use shallow features for coarse-grained classification and deep features for fine-grained classification. In addition, to enable fine-grained features to enhance the coarse-grained classification, we propose a feature reinforcement module based on the feature pyramid structure, where deep features are first upsampled and then combined with shallow features to make decisions. Experimental results on three widely used fine-grained image classification datasets such as CUB-200-2011, Stanford Cars, and FGVC-Aircraft validate the method's effectiveness. Code available at https://github.com/PRIS-CV/CGVC. Junhan Chen, Dongliang Chang, Jiyang Xie 0001, Ruoyi Du, Zhanyu Ma |
VCIP | 2 |
| 2022 | Progressive Learning of Category-Consistent Multi-Granularity Features for Fine-Grained Visual ClassificationabstractFine-grained visual classification (FGVC) is much more challenging than traditional classification tasks due to the inherently subtle intra-class object variations. Recent works are mainly part-driven (either explicitly or implicitly), with the assumption that fine-grained information naturally rests within the parts. In this paper, we take a different stance, and show that part operations are not strictly necessary - the key lies with encouraging the network to learn at different granularities and progressively fusing multi-granularity features together. In particular, we propose: (i) a progressive training strategy that effectively fuses features from different granularities, and (ii) a consistent block convolution that encourages the network to learn the category-consistent features at specific granularities. We evaluate on several standard FGVC benchmark datasets, and demonstrate the proposed method consistently outperforms existing alternatives or delivers competitive results. Codes are available at https://github.com/PRIS-CV/PMG-V2. Ruoyi Du, Jiyang Xie 0001, Zhanyu Ma, Dongliang Chang, Yi-Zhe Song, Jun Guo 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2022 | GPCA: A Probabilistic Framework for Gaussian Process Embedded Channel AttentionabstractChannel attention mechanisms have been commonly applied in many visual tasks for effective performance improvement. It is able to reinforce the informative channels as well as to suppress the useless channels. Recently, different channel attention modules have been proposed and implemented in various ways. Generally speaking, they are mainly based on convolution and pooling operations. In this paper, we propose Gaussian process embedded channel attention (GPCA) module and further interpret the channel attention schemes in a probabilistic way. The GPCA module intends to model the correlations among the channels, which are assumed to be captured by beta distributed variables. As the beta distribution cannot be integrated into the end-to-end training of convolutional neural networks (CNNs) with a mathematically tractable solution, we utilize an approximation of the beta distribution to solve this problem. To specify, we adapt a Sigmoid-Gaussian approximation, in which the Gaussian distributed variables are transferred into the interval [0,1]. The Gaussian process is then utilized to model the correlations among different channels. In this case, a mathematically tractable solution is derived. The GPCA module can be efficiently implemented and integrated into the end-to-end training of the CNNs. Experimental results demonstrate the promising performance of the proposed GPCA module. Codes are available at https://github.com/PRIS-CV/GPCA. Jiyang Xie 0001, Zhanyu Ma, Dongliang Chang, Guoqiang Zhang 0003, Jun Guo 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2021 | Your "Flamingo" is My "Bird": Fine-Grained, or NotabstractWhether what you see in Figure 1 is a "flamingo" or a "bird", is the question we ask in this paper. While fine-grained visual classification (FGVC) strives to arrive at the former, for the majority of us non-experts just "bird" would probably suffice. The real question is therefore – how can we tailor for different fine-grained definitions under divergent levels of expertise. For that, we re-envisage the traditional setting of FGVC, from single-label classification, to that of top-down traversal of a pre-defined coarse-to-fine label hierarchy – so that our answer becomes "bird" ⇒ "Phoenicopteriformes" ⇒ "Phoenicopteridae" ⇒ "flamingo".To approach this new problem, we first conduct a comprehensive human study where we confirm that most participants prefer multi-granularity labels, regardless whether they consider themselves experts. We then discover the key intuition that: coarse-level label prediction exacerbates fine-grained feature learning, yet fine-level feature betters the learning of coarse-level classifier. This discovery enables us to design a very simple albeit surprisingly effective solution to our new problem, where we (i) leverage level-specific classification heads to disentangle coarse-level features with fine-grained ones, and (ii) allow finer-grained features to participate in coarser-grained label predictions, which in turn helps with better disentanglement. Experiments show that our method achieves superior performance in the new FGVC setting, and performs better than state-of-the-art on the traditional single-label FGVC problem as well. Thanks to its simplicity, our method can be easily implemented on top of any existing FGVC frameworks and is parameter-free. Dongliang Chang, Kaiyue Pang, Yixiao Zheng, Zhanyu Ma, Yi-Zhe Song, Jun Guo 0002 |
CVPR | 1 |
| 2021 | Knowledge Transfer Based Fine-Grained Visual ClassificationabstractFine-grained visual classification (FGVC) aims to distinguish the sub-classes of the same category and its essential solution is to mine the subtle and discriminative regions. Convolution neural networks (CNNs), which employ the cross entropy loss (CE-loss) as the loss function, show poor performance since the model can only learn the most discriminative part and ignore other meaningful regions. Some existing works try to solve this problem by mining more discriminative regions by some detection techniques or attention mechanisms. However, most of them will meet the background noise problem when trying to find more discriminative regions. In this paper, we address it in a knowledge transfer learning manner. Multiple models are trained one by one, and all previously trained models are regarded as teacher models to supervise the training of the current one. Specifically, a orthogonal loss (OR-loss) is proposed to encourage the network to find diverse and meaningful regions. In addition, the first model is trained with only CE-Loss. Finally, all models’ outputs with complementary knowledge are combined together for the final prediction result. We demonstrate the superiority of the proposed method and obtain state-of-the-art (SOTA) performances on three popular FGVC datasets. Siqing Zhang 0001, Ruoyi Du, Dongliang Chang, Zhanyu Ma, Jun Guo 0002 |
ICME | 3 |
| 2021 | Exploring Category-Shared and Category-Specific Features for Fine-Grained Image Classification
Dongliang Chang, Bo Xiao 0002, Zhanyu Ma, Jun Guo 0002, Yaning Chang |
PRCV (1) | 2 |
| 2021 | Progressive Co-Attention Network for Fine-Grained Visual ClassificationabstractFine-grained visual classification aims to recognize images belonging to multiple sub-categories within a same category. It is a challenging task due to the inherently subtle variations among highly-confused categories. Most existing methods only take an individual image as input, which may limit the ability of models to recognize contrastive clues from different images. In this paper, we propose an effective method called progressive co-attention network (PCA-Net) to tackle this problem. Specifically, we calculate the channel-wise similarity by encouraging interaction between the feature channels within same-category image pairs to capture the common discriminative features. Considering that complementary information is also crucial for recognition, we erase the prominent areas enhanced by the channel interaction to force the network to focus on other discriminative regions. The proposed model has achieved competitive results on three fine-grained visual classification benchmark datasets: CUB-200-2011, Stanford Cars, and FGVC Aircraft. Tian Zhang 0029, Dongliang Chang, Zhanyu Ma, Jun Guo 0002 |
VCIP | 2 |
| 2021 | Deep InterBoost networks for small-sample image classification
Dongliang Chang, Zhanyu Ma, Zheng-Hua Tan, Jing-Hao Xue, Jie Cao 0014, Jun Guo 0002 |
Neurocomputing | 2 |
| 2021 | Competing ratio loss for discriminative multi-class image classification
Ke Zhang 0005, Yurong Guo 0001, Dongliang Chang, Zhenbing Zhao, Zhanyu Ma, Tony X. Han |
Neurocomputing | 4 |
| 2021 | AP-CNN: Weakly Supervised Attention Pyramid Convolutional Neural Network for Fine-Grained Visual ClassificationabstractClassifying the sub-categories of an object from the same super-category (e.g., bird species and cars) in fine-grained visual classification (FGVC) highly relies on discriminative feature representation and accurate region localization. Existing approaches mainly focus on distilling information from high-level features. In this article, by contrast, we show that by integrating low-level information (e.g., color, edge junctions, texture patterns), performance can be improved with enhanced feature representation and accurately located discriminative regions. Our solution, named Attention Pyramid Convolutional Neural Network (AP-CNN), consists of 1) a dual pathway hierarchy structure with a top-down feature pathway and a bottom-up attention pathway, hence learning both high-level semantic and low-level detailed feature representation, and 2) an ROI-guided refinement strategy with ROI-guided dropblock and ROI-guided zoom-in operation, which refines features with discriminative local regions enhanced and background noises eliminated. The proposed AP-CNN can be trained end-to-end, without the need of any additional bounding box/part annotation. Extensive experiments on three popularly tested FGVC datasets (CUB-200-2011, Stanford Cars, and FGVC-Aircraft) demonstrate that our approach achieves state-of-the-art performance. Models and code are available at https://github.com/PRIS-CV/AP-CNN_Pytorch-master. Zhanyu Ma, Shaoguo Wen, Jiyang Xie 0001, Dongliang Chang, Zhongwei Si, Ming Wu 0001, Haibin Ling |
IEEE Trans. Image Process. | 5 |
| 2020 | Fine-Grained Visual Classification via Progressive Multi-granularity Training of Jigsaw Patches
Ruoyi Du, Dongliang Chang, Ayan Kumar Bhunia, Jiyang Xie 0001, Zhanyu Ma, Yi-Zhe Song, Jun Guo 0002 |
ECCV (20) | 2 |
| 2020 | IU-Module: Intersection and Union Module for Fine-Grained Visual ClassificationabstractA predominant viewpoint in previous works of fine-grained visual classification (FGVC) is to the localize discriminative parts by auxiliary networks and extract the part-based finegrained features for classification. In this paper, we propose a simple yet effective approach by introducing an intersection and union module (IU-Module). The IU-Module aims to capture more discriminative features by 1) dividing features into distinct groups, 2) sharing parts of interests within each group, and 3) adding a differentiation loss to reduce the similarity among those grouped feature channels. Without adding any new learnable parameters, the proposed approach imposes two straightforward operations, namely channel intersection (CI) and channel union (CU) operations, on the convolutional features and achieves competitive results compared with the state-of-the-art methods. Experimental results on three publicly available FGVC datasets show the effectiveness of the IU-Module. Ablation studies and visualizations are also provided to make further demonstrations. Yixiao Zheng, Dongliang Chang, Jiyang Xie 0001, Zhanyu Ma |
ICME | 2 |
| 2020 | CC-Loss: Channel Correlation Loss for Image ClassificationabstractThe loss function is a key component in deep learning models. A commonly used loss function for classification is the cross entropy loss, which is a simple yet effective application of information theory for classification problems. Based on this loss, many other loss functions have been proposed, e.g., by adding intra-class and inter-class constraints to enhance the discriminative ability of the learned features. However, these loss functions fail to consider the connections between the feature distribution and the model structure. Aiming at addressing this problem, we propose a channel correlation loss (CC-Loss) that is able to constrain the specific relations between classes and channels as well as maintain the intra-class and the inter-class separability. CC-Loss uses a channel attention module to generate channel attention of features for each sample in the training stage. Next, an Euclidean distance matrix is calculated to make the channel attention vectors associated with the same class become identical and to increase the difference between different classes. Finally, we obtain a feature embedding with good intra-class compactness and inter-class separability. Experimental results show that two different backbone models trained with the proposed CC-Loss outperform the state-of-the-art loss functions on three image classification datasets. Dongliang Chang, Zhanyu Ma, Zheng-Hua Tan |
ICPR | 2 |
| 2020 | Dual-attention Guided Dropblock Module for Weakly Supervised Object LocalizationabstractAttention mechanisms is frequently used to learn the discriminative features for better feature representations. In this paper, we extend the attention mechanism to the task of weakly supervised object localization (WSOL) and propose the dual-attention guided dropblock module (DGDM), which aims at learning the informative and complementary visual patterns for WSOL. This module contains two key components, the channel attention guided dropout (CAGD) and the spatial attention guided dropblock (SAGD). To model channel interdependencies, the CAGD ranks the channel attentions and treats the top-k attentions with the largest magnitudes as the important ones. It also keeps some low-valued elements to increase their value if they become important during training. The SAGD can efficiently remove the most discriminative information by erasing the contiguous regions of feature maps rather than individual pixels. This guides the model to capture the less discriminative parts for classification. Furthermore, it can also distinguish the foreground objects from the background regions to alleviate the attention misdirection. Experimental results demonstrate that the proposed method achieves new state-of-the-art localization performance. Junhui Yin, Siqing Zhang 0001, Dongliang Chang, Zhanyu Ma, Jun Guo 0002 |
ICPR | 3 |
| 2020 | The Devil is in the Channels: Mutual-Channel Loss for Fine-Grained Image ClassificationabstractThe key to solving fine-grained image categorization is finding discriminate and local regions that correspond to subtle visual traits. Great strides have been made, with complex networks designed specifically to learn part-level discriminate feature representations. In this paper, we show that it is possible to cultivate subtle details without the need for overly complicated network designs or training mechanisms - a single loss is all it takes. The main trick lies with how we delve into individual feature channels early on, as opposed to the convention of starting from a consolidated feature map. The proposed loss function, termed as mutual-channel loss (MC-Loss), consists of two channel-specific components: a discriminality component and a diversity component. The discriminality component forces all feature channels belonging to the same class to be discriminative, through a novel channel-wise attention mechanism. The diversity component additionally constraints channels so that they become mutually exclusive across the spatial dimension. The end result is therefore a set of feature channels, each of which reflects different locally discriminative regions for a specific class. The MC-Loss can be trained end-to-end, without the need for any bounding-box/part annotations, and yields highly discriminative regions during inference. Experimental results show our MC-Loss when implemented on top of common base networks can achieve state-of-the-art performance on all four fine-grained categorization datasets (CUB-Birds, FGVC-Aircraft, Flowers-102, and Stanford Cars). Ablative studies further demonstrate the superiority of the MC-Loss when compared with other recently proposed general-purpose losses for visual classification, on two different base networks. Dongliang Chang, Jiyang Xie 0001, Ayan Kumar Bhunia, Zhanyu Ma, Ming Wu 0001, Jun Guo 0002, Yi-Zhe Song |
IEEE Trans. Image Process. | 1 |
| 2020 | OSLNet: Deep Small-Sample Classification With an Orthogonal Softmax LayerabstractA deep neural network of multiple nonlinear layers forms a large function space, which can easily lead to overfitting when it encounters small-sample data. To mitigate overfitting in small-sample classification, learning more discriminative features from small-sample data is becoming a new trend. To this end, this paper aims to find a subspace of neural networks that can facilitate a large decision margin. Specifically, we propose the Orthogonal Softmax Layer (OSL), which makes the weight vectors in the classification layer remain orthogonal during both the training and test processes. The Rademacher complexity of a network using the OSL is only 1/K, where K is the number of classes, of that of a network using the fully connected classification layer, leading to a tighter generalization error bound. Experimental results demonstrate that the proposed OSL has better performance than the methods used for comparison on four small-sample benchmark datasets, as well as its applicability to large-sample datasets. Codes are available at: https://github.com/dongliangchang/OSLNet. Dongliang Chang, Zhanyu Ma, Zheng-Hua Tan, Jing-Hao Xue, Jie Cao 0014, Jingyi Yu 0001, Jun Guo 0002 |
IEEE Trans. Image Process. | 2 |
| 2019 | FICAL: Focal Inter-Class Angular Loss for Image ClassificationabstractConvolutional Neural Networks (CNNs) have been successfully applied in various image analysis tasks and gradually become one of the most powerful machine learning approaches. In order to improve the capability of the model generalization and performance in image classification, a new trend is to learn more discriminative features via CNNs. The main contribution of this paper is to increase the angles between the categories to extract discriminative features and enlarge the inter-class variance. To this end, we propose a loss function named focal inter-class angular loss (FICAL) which introduces the confusion rate-weighted cosine distance as the similarity measurement between categories. This measurement is dynamically evaluated during each iteration to adapt the model. Compared with other loss functions, experimental results demonstrate that the proposed FICAL achieved best performance among the referred loss functions on two image classificaton datasets. Xinran Wei, Dongliang Chang, Jiyang Xie 0001, Yixiao Zheng, Chen Gong 0002, Zhanyu Ma |
VCIP | 2 |