Kongming Liang

dblp:161/1948 · DBLP profile ↗
← Back
57ranked-venue papers
11as first author
49since 2021 · last 2026
0000-0002-4726-093XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 42 · 7 first-author · 37 since 2021Artificial intelligence and machine learning · 26 · 8 first-author · 19 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 first-author · 5 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 MedReasoner: Reinforcement Learning Drives Reasoning Grounding from Clinical Thought to Pixel-Level Precision
abstract
Accurately grounding regions of interest (ROIs) is critical for diagnosis and treatment planning in medical imaging. While multimodal large language models (MLLMs) combine visual perception with natural language, current medical-grounding pipelines still rely on supervised fine-tuning with explicit spatial hints, making them ill-equipped to handle the implicit queries common in clinical practice. This work makes three core contributions. We first define Unified Medical Reasoning Grounding (UMRG), a novel vision–language task that demands clinical reasoning and pixel-level grounding. Second, we release U-MRG-14K, a dataset of 14K samples featuring pixel-level masks alongside implicit clinical queries and reasoning traces, spanning 10 modalities, 15 super-categories, and 108 specific categories. Finally, we introduce MedReasoner, a modular framework that distinctly separates reasoning from segmentation: an MLLM reasoner is optimized with reinforcement learning, while a frozen segmentation expert converts spatial prompts into masks, with alignment achieved through format and accuracy rewards. MedReasoner achieves state-of-the-art performance on U-MRG-14K and demonstrates strong generalization to unseen clinical queries, underscoring the significant promise of reinforcement learning for interpretable medical grounding.
Zhonghao Yan, Muxi Diao, Ruoyan Jing, Jiayuan Xu, Kaizhou Zhang, Lele Yang, Yanxi Liu 0006, Kongming Liang, Zhanyu Ma
AAAI9
2026 Benchmarking Semantic Segmentation Models via Appearance and Geometry Attribute Editing
abstract
Semantic segmentation takes a pivotal role in various applications such as autonomous driving and medical image analysis. When deploying segmentation models in practice, it is critical to test their behaviors in varied and complex scenes in advance. In this paper, we construct an automatic data generation pipeline Gen4Seg to stress-test semantic segmentation models by generating various challenging samples with different attribute changes. Beyond previous evaluation paradigms focusing solely on global weather and style transfer, we investigate variations in both appearance and geometry attributes at the object and image level. These include object color, material, size, and position, as well as image-level variations such as weather and style. To achieve this, we propose to edit visual attributes of existing real images with precise control of structural information, empowered by diffusion models. In this way, the existing segmentation labels can be reused for the edited images, which greatly reduces the labor costs of constructing datasets. Using our pipeline, we construct two new benchmarks, Pascal-EA and COCO-EA. We benchmark a broad variety of semantic segmentation models, spanning from conventional close-set models to recent open-vocabulary large models. We have several key findings: 1) advanced open-vocabulary models do not exhibit greater robustness compared to closed-set methods under geometric variations; 2) traditional data augmentation techniques, such as CutOut and CutMix, are limited in enhancing robustness against appearance variations; 3) our generation pipeline can also be employed as a data augmentation tool and improve both in-distribution and out-of-distribution performances. Our work suggests the potential of generative models as effective tools for automatically analyzing segmentation models, and we hope our findings will assist practitioners and researchers in developing more robust and reliable segmentation models.
Zijin Yin, Bing Li 0015, Kongming Liang, Hao Sun 0015, Zhongjiang He, Zhanyu Ma, Jun Guo 0002
IEEE Trans. Pattern Anal. Mach. Intell.3
2026 Toward Generalizable Forgery Detection and Reasoning
abstract
Accurate and interpretable detection of AI-generated images is essential for mitigating risks associated with AI misuse. However, the substantial domain gap among generative models makes it challenging to develop a generalizable forgery detection model. Moreover, since every pixel in an AI-generated image is synthesized, traditional saliency-based forgery explanation methods are not well suited for this task. To address these challenges, we formulate detection and explanation as a unified Forgery Detection and Reasoning task (FDR-Task), leveraging Multi-Modal Large Language Models (MLLMs) to provide accurate detection through reliable reasoning over forgery attributes. To facilitate this task, we introduce the Multi-Modal Forgery Reasoning dataset (MMFR-Dataset), a large-scale dataset containing 120K images across 10 generative models, with 378K reasoning annotations on forgery attributes, enabling comprehensive evaluation of the FDR-Task. Furthermore, we propose FakeReasoning, a forgery detection and reasoning framework with three key components: 1) a dual-branch visual encoder that integrates CLIP and DINO to capture both high-level semantics and low-level artifacts; 2) a Forgery-Aware Feature Fusion Module that leverages DINO's attention maps and cross-attention mechanisms to guide MLLMs toward forgery-related clues; 3) a Classification Probability Mapper that couples language modeling and forgery detection, enhancing overall performance. Experiments across multiple generative models demonstrate that FakeReasoning not only achieves robust generalization but also outperforms state-of-the-art methods on both detection and reasoning tasks. The code is available at: https://github.com/PRIS-CV/FakeReasoning.
Yueying Gao, Dongliang Chang, Bingyao Yu, Haotian Qin, Muxi Diao, Lei Chen 0069, Kongming Liang, Zhanyu Ma
IEEE Trans. Image Process.7
2025 ConMo: Controllable Motion Disentanglement and Recomposition for Zero-Shot Motion Transfer
abstract
The development of Text-to-Video (T2V) generation has made motion transfer possible, enabling the control of video motion based on existing footage. However, current methods have two limitations: 1) struggle to handle multi-subjects videos, failing to transfer specific subject motion; 2) struggle to preserve the diversity and accuracy of motion as transferring to subjects with varying shapes. To overcome these, we introduce ConMo, a zero-shot framework that disentangle and recompose the motions of subjects and camera movements. ConMo isolates individual subject and background motion cues from complex trajectories in source videos using only subject masks, and reassembles them for target video generation. This approach enables more accurate motion control across diverse subjects and improves performance in multi-subject scenarios. Additionally, we propose soft guidance in the recomposition stage which controls the retention of original motion to adjust shape constraints, aiding subject shape adaptation and semantic transformation. Unlike previous methods, ConMo unlocks a wide range of applications, including subject size and position editing, subject removal, semantic modifications, and camera motion simulation. Extensive experiments demonstrate that ConMo significantly outperforms state-of-the-art methods in motion fidelity and semantic consistency. The code is available at https://github.com/Andyplus1/ConMo.
Jiayi Gao, Zijin Yin, Changcheng Hua, Yuxin Peng 0001, Kongming Liang, Zhanyu Ma, Jun Guo 0002, Yang Liu 0105
CVPR5
2025 FairHuman: Boosting Hand and Face Quality in Human Image Generation with Minimum Potential Delay Fairness in Diffusion Models
Tianwei Cao, Zhongjiang He, Kongming Liang, Zhanyu Ma
ICCV5
2025 SelectVision: Adaptive Vision Resolution Selection for Visual Document Understanding
Zhongjiang He, Han Fang 0002, Hao Sun 0015, Kongming Liang, Zhanyu Ma
ICDAR (4)6
2025 A Debiasing Framework For Attribute Binding In Diffusion-Based Text-To-Image Generation
abstract
Despite the impressive generative capabilities of text-to-image (T2I) models, accurate binding of objects and attributes specified in input prompts remains a significant challenge. Existing approaches often fail to address inherent biases in text embeddings, where objects tend to associate with their frequent attributes in the training data. This work identifies these biases and introduces a novel debiasing framework centered on alternating prompt vector binding, dynamically improving semantic alignment during image generation. First, we quantify bias strength in object-attribute pairs using large language models (LLMs), revealing problematic associations. Next, biased object terms are replaced with broader concepts during diffusion sampling to reduce bias. Finally, object embeddings are refined by pulling target attributes while suppressing irrelevant ones, avoiding attribute confusion. Extensive experiments on public datasets demonstrate the framework’s effectiveness in addressing attribute binding challenges through bias mitigation.
Yueheng Luo, Tianwei Cao, Ling Jin 0004, Donghui Gao, Kongming Liang, Zhanyu Ma
ICIP7
2025 RotatedMVPS: Multi-view Photometric Stereo with Rotated Natural Light
abstract
Multiview photometric stereo (MVPS) seeks to recover high-fidelity surface shapes and reflectances from images captured under varying views and illuminations. However, existing MVPS methods often require controlled darkroom settings for varying illuminations or overlook the recovery of reflectances and illuminations properties, limiting their applicability in natural illumination scenarios and downstream inverse rendering tasks. In this paper, we propose RotatedMVPS to solve shape and reflectance recovery under rotated natural light, achievable with a practical rotation stage. By ensuring light consistency across different camera and object poses, our method reduces the unknowns associated with complex environment light. Furthermore, we integrate data priors from off-the-shelf learning-based single-view photometric stereo methods into our MVPS framework, significantly enhancing the accuracy of shape and reflectance recovery. Experimental results on both synthetic and real-world datasets demonstrate the effectiveness of our approach.
Songyun Yang, Yufei Han 0002, Kongming Liang, Peng Yu 0001, Zhaowei Qu, Heng Guo 0003
ICME4
2025 CineTechBench: A Benchmark for Cinematographic Technique Understanding and Generation
abstract
Cinematography is a cornerstone of film production and appreciation, shaping mood, emotion, and narrative through visual elements such as camera movement, shot composition, and lighting. Despite recent progress in multimodal large language models (MLLMs) and video generation models, the capacity of current models to grasp and reproduce cinematographic techniques remains largely uncharted, hindered by the scarcity of expert-annotated data. To bridge this gap, we present CineTechBench, a pioneering benchmark founded on precise, manual annotation by seasoned cinematography experts across key cinematography dimensions. Our benchmark covers seven essential aspects—shot scale, shot angle, composition, camera movement, lighting, color, and focal length—and includes over 600 annotated movie images and 120 movie clips with clear cinematographic techniques. For the understanding task, we design question–answer pairs and annotated descriptions to assess MLLMs’ ability to interpret and explain cinematographic techniques. For the generation task, we assess advanced video generation models on their capacity to reconstruct cinema-quality camera movements given conditions such as textual prompts or keyframes. We conduct a large-scale evaluation on 15+ MLLMs and 5+ video generation models. Our results offer insights into the limitations of current models and future directions for cinematography understanding and generation in automatical film production and appreciation. The code and benchmark can be accessed at \url{https://github.com/PRIS-CV/CineTechBench}.
Songyu Xu, Xiangxuan Shan, Muxi Diao, Xueyan Duan, Yanhua Huang, Kongming Liang, Zhanyu Ma
NeurIPS8
2025 Animal-CLIP: A Dual-Prompt Enhanced Vision-Language Model for Animal Action Recognition
Yinuo Jing, Kongming Liang, Ruxu Zhang, Hao Sun 0015, Zhongjiang He, Zhanyu Ma
Int. J. Comput. Vis.2
2025 Understanding Episode Hardness in Few-Shot Learning
abstract
Achieving generalization for deep learning models has usually suffered from the bottleneck of annotated sample scarcity. As a common way of tackling this issue, few-shot learning focuses on "episodes", i.e., sampled tasks that help the model acquire generalizable knowledge onto unseen categories - better the episodes, the higher a model's generalisability. Despite extensive research, the characteristics of episodes and their potential effects are relatively less explored. A recent paper discussed that different episodes exhibit different prediction difficulties, and coined a new metric "hardness" to quantify episodes, which however is too wide-range for an arbitrary dataset and thus remains impractical for realistic applications. In this paper therefore, we for the first time conduct an algebraic analysis of the critical factors influencing episode hardness supported by experimental demonstrations, that reveal episode hardness to largely depend on classes within an episode, and importantly propose an efficient pre-sampling hardness assessment technique named Inverse-Fisher Discriminant Ratio (IFDR). This enables sampling hard episodes at the class level via class-level (CL) sampling scheme that drastically decreases quantification cost. Delving deeper, we also develop a variant called class-pair-level (CPL) sampling, which further reduces the sampling cost while guaranteeing the sampled distribution. Finally, comprehensive experiments conducted on benchmark datasets verify the efficacy of our proposed method.
Yurong Guo 0001, Ruoyi Du, Aneeshan Sain, Kongming Liang, Yi-Zhe Song, Zhanyu Ma
IEEE Trans. Pattern Anal. Mach. Intell.4
2025 Disentangling Before Composing: Learning Invariant Disentangled Features for Compositional Zero-Shot Learning
abstract
Compositional Zero-Shot Learning (CZSL) aims to recognize novel compositions using knowledge learned from seen attribute-object compositions in the training set. Previous works mainly project an image and its corresponding composition into a common embedding space to measure their compatibility score. However, both attributes and objects share the visual representations learned above, leading the model to exploit spurious correlations and bias towards seen compositions. Instead, we reconsider CZSL as an out-of-distribution generalization problem. If an object is treated as a domain, we can learn object-invariant features to recognize attributes attached to any object reliably, and vice versa. Specifically, we propose an invariant feature learning framework to align different domains at the representation and gradient levels to capture the intrinsic characteristics associated with the tasks. To further facilitate and encourage the disentanglement of attributes and objects, we propose an "encoding-reshuffling-decoding" process to help the model avoid spurious correlations by randomly regrouping the disentangled features into synthetic features. Ultimately, our method improves generalization by learning to disentangle features that represent two independent factors of attributes and objects. Experiments demonstrate that the proposed method achieves state-of-the-art or competitive performance in both closed-world and open-world scenarios.
Tian Zhang 0029, Kongming Liang, Ruoyi Du, Wei Chen 0071, Zhanyu Ma
IEEE Trans. Pattern Anal. Mach. Intell.2
2025 Reserve to Adapt: Mining Inter-Class Relations for Open-Set Domain Adaptation
abstract
Open-Set Domain Adaptation (OSDA) aims at adapting a model trained on a labelled source domain, to an unlabeled target domain that is corrupted with unknown classes. The key challenge inherent to this open-set setting is therefore how best to avoid the negative transfer incurred by unknown classes during model adaptation. Most existing works tackle this challenge by simply pushing the entire unknown classes away. In this paper, we take a different stance - instead of addressing these unknown classes as a single entity, we "reserve" in-between spaces for their subsets in the learned embedding. Our key finding is that the inter-class relations learned off the source domain, can help to enforce class separations in the target domain - thereby reserving spaces for unknown classes. More specifically, we first prep the "reservation" by tightening the known-class representations while enlarging their inter-class margin. We then learn soft-label prototypes in the source domain to facilitate the discrimination of known and unknown samples in the target domain. It follows that these two steps are iterated at each epoch in a mutually beneficial manner - better discrimination of unknown samples helps with space reservation, and vice versa. We show state-of-the-art results on four standard OSDA datasets, Office-31, Office-Home, VisDA and ImageCLEF, and conduct further analysis to help understand our method. Codes are available at: https://github.com/PRIS-CV/Reserve_to_Adapt.
Yujun Tong, Dongliang Chang, Da Li 0001, Kongming Liang, Zhongjiang He, Yi-Zhe Song, Zhanyu Ma
IEEE Trans. Image Process.5
2025 Detailed Object Description With Controllable Dimensions
abstract
Object description plays an important role for visually impaired individuals to understand and compare the differences between objects. Recent multimodal large language models (MLLMs) exhibit powerful perceptual abilities and demonstrate impressive potential for generating object-centric descriptions. However, the descriptions generated by such models may still usually contain a lot of content that is not relevant to the user intent or miss some important object dimension details. Under special scenarios, users may only need the details of certain dimensions of an object. In this paper, we propose a training-free object description refinement pipeline,Dimension Tailor, designed to enhance user-specified details in object descriptions. This pipeline includes three steps: dimension extracting, erasing, and supplementing, which decompose the description into user-specified dimensions. Dimension Tailor can not only improve the quality of object details but also offer flexibility in including or excluding specific dimensions based on user preferences. We conducted extensive experiments to demonstrate the effectiveness of Dimension Tailor on controllable object descriptions. Notably, the proposed pipeline can consistently improve the performance of the recent MLLMs. The code is currently accessible athttps://github.com/PRIS-CV/ControllableObjectDescription.
Haiwen Zhang, Baoteng Li, Kongming Liang, Hao Sun 0015, Zhongjiang He, Zhanyu Ma, Jun Guo 0002
IEEE Trans. Multim.4
2024 Dual-Prior Augmented Decoding Network for Long Tail Distribution in HOI Detection
abstract
Human object interaction detection aims at localizing human-object pairs and recognizing their interactions. Trapped by the long-tailed distribution of the data, existing HOI detection methods often have difficulty recognizing the tail categories. Many approaches try to improve the recognition of HOI tasks by utilizing external knowledge (e.g. pre-trained visual-language models). However, these approaches mainly utilize external knowledge at the HOI combination level and achieve limited improvement in the tail categories. In this paper, we propose a dual-prior augmented decoding network by decomposing the HOI task into two sub-tasks: human-object pair detection and interaction recognition. For each subtask, we leverage external knowledge to enhance the model's ability at a finer granularity. Specifically, we acquire the prior candidates from an external classifier and embed them to assist the subsequent decoding process. Thus, the long-tail problem is mitigated from a coarse-to-fine level with the corresponding external knowledge. Our approach outperforms existing state-of-the-art models in various settings and significantly boosts the performance on the tail HOI categories. The source code is available at https://github.com/PRIS-CV/DP-ADN.
Jiayi Gao, Kongming Liang, Wei Chen 0071, Zhanyu Ma, Jun Guo 0002
AAAI2
2024 Class-Aware Contrastive Learning for Fine-Grained Skeleton-Based Action Recognition
Xinyu Bian, Dongliang Chang, Zhongjiang He, Kongming Liang, Zhanyu Ma
ACCV (1)5
2024 Hierarchical Prompting for Diffusion Classifiers
Wenxin Ning, Dongliang Chang, Yujun Tong, Zhongjiang He, Kongming Liang, Zhanyu Ma
ACCV (8)5
2024 Polyp-E: Benchmarking the Robustness of Deep Segmentation Models via Polyp Editing
abstract
In daily clinical practice, clinicians exhibit robustness in identifying polyps with both location and size variations. It is uncertain if deep segmentation models can achieve comparable robustness in automated colonoscopic analysis. To benchmark the model robustness, we focus on evaluating the segmentation models on the polyps with various attributes (e.g. location and size) and healthy samples. Based on the Latent Diffusion Model, we perform attribute editing on real polyps and build a new dataset named Polyp-E. Our synthetic dataset boasts exceptional realism, to the extent that clinical experts find it challenging to discern them from real data. We evaluate various existing polyp segmentation models on the proposed benchmark. The results reveal most of the models are highly sensitive to attribute variations. As a novel data augmentation technique, the proposed editing pipeline can improve both in-distribution and out-ofdistribution generalization ability. The code and datasets has been released at https://github.com/RunpuWei/Polyp-E-Benchmark.
Runpu Wei, Zijin Yin, Kongming Liang, Min Min, Chengwei Pan, Haonan Huang, Zhanyu Ma
BIBM3
2024 Benchmarking Segmentation Models with Mask-Preserved Attribute Editing
abstract
When deploying segmentation models in practice, it is critical to evaluate their behaviors in varied and complex scenes. Different from the previous evaluation paradigms only in consideration of global attribute variations (e.g. adverse weather), we investigate both local and global attribute variations for robustness evaluation. To achieve this, we construct a mask-preserved attribute editing pipeline to edit visual attributes of real images with precise control of structural information. Therefore, the original segmentation labels can be reused for the edited images. Using our pipeline, we construct a benchmark covering both object and image attributes (e.g. color, material, pattern, style). We evaluate a broad variety of semantic segmentation models, spanning from conventional close-set models to recent open-vocabulary large models on their robustness to different types of variations. We find that both local and global attribute variations affect segmentation performances, and the sensitivity of models diverges across different variation types. We argue that local attributes have the same importance as global attributes, and should be considered in the robustness evaluation of segmentation models. Code: https://github.com/PRIS-CV/Pascal-EA.
Zijin Yin, Kongming Liang, Bing Li 0015, Zhanyu Ma, Jun Guo 0002
CVPR2
2024 SLNL: Soft Label Regularization For Semi-Supervised Facial Expression Recognition With Negative Label Learning
abstract
Semi-supervised learning (SSL) methods have been widely employed in facial expression recognition (FER) to eliminate the substantial cost of acquiring well-labeled samples. These methods typically involve assigning pseudo labels to unlabeled samples beyond a certain confidence threshold for calculating cross-entropy loss. However, they overlook the correctness issue with pseudo labels, potentially resulting in network over-fitting. In this paper, we introduce an SLNL model that incorporates a Soft Label Regularization (SLR) module and a Negative Label Learning (NLL) module, preventing network from over-fitting while enhancing the efficiency of utilizing unlabeled data. For each unlabeled sample whose pseudo-label confidence surpasses a threshold, SLR concurrently uses the pseudo label and the previously recorded soft label for supervised learning. Additionally, NLL dynamically explores negative labels by calculating top-k accuracy, further removing irrelevant information. The SLNL model achieves state-of-the-art performance across several widely used datasets, notably surpassing the fully-supervised baseline on AffectNet with fewer labeled data.
Yuying Zhao, Kongming Liang
ICIP4
2024 Learning Conditional Prompt for Compositional Zero-Shot Learning
abstract
Compositional zero-shot learning (CZSL) strives to learn attributes and objects from seen compositions and transfer the acquired knowledge to unseen compositions. Existing methods either learn primitive concepts in an entangled manner, leading to the model relying on spurious correlations between attributes and objects. Alternatively, they adopt a decoupled approach, causing the model to overlook relationships between attributes and objects. In this paper, we propose a conditional prompting (CoP) method to enhance the performance of vision-language models (e.g., CLIP) in CZSL. Specifically, we utilize two image-to-word mapping networks to learn pseudo attribute and object word embeddings that can represent the corresponding semantics of input images. Subsequently, the model recognizes one concept based on another generated pseudo word embeddings, enabling the recognition of individual sub-concepts while leveraging the image-specific correlations between attributes and objects. The experimental results on three CZSL benchmarks indicate that the proposed method achieves competitive performance compared to previous state-of-the-art methods.
Tian Zhang 0029, Kongming Liang, Ke Zhang 0005, Zhanyu Ma
ICME2
2024 Efficient Face Super-Resolution via Wavelet-based Feature Enhancement Network
abstract
Face super-resolution aims to reconstruct a high-resolution face image from a low-resolution face image. Previous methods typically employ an encoder-decoder structure to extract facial structural features, where the direct downsampling inevitably introduces distortions, especially to high-frequency features such as edges. To address this issue, we propose a wavelet-based feature enhancement network, which mitigates feature distortion by losslessly decomposing the input feature into high and low-frequency components using the wavelet transform and processing them separately. To improve the efficiency of facial feature extraction, a full domain Transformer is further proposed to enhance local, regional, and global facial features. Such designs allow our method to perform better without stacking many modules as previous methods did. Experiments show that our method effectively balances performance, model size, and speed. Code link: https://github.com/PRIS-CV/WFEN.
Heng Guo 0003, Xuannan Liu, Kongming Liang, Jiani Hu, Zhanyu Ma, Jun Guo 0002
ACM Multimedia4
2024 Animal-Bench: Benchmarking Multimodal Video Models for Animal-centric Video Understanding
abstract
With the emergence of large pre-trained multimodal video models, multiple benchmarks have been proposed to evaluate model capabilities. However, most of the benchmarks are human-centric, with evaluation data and tasks centered around human applications. Animals are an integral part of the natural world, and animal-centric video understanding is crucial for animal welfare and conservation efforts. Yet, existing benchmarks overlook evaluations focused on animals, limiting the application of the models. To address this limitation, our work established an animal-centric benchmark, namely Animal-Bench, to allow for a comprehensive evaluation of model capabilities in real-world contexts, overcoming agent-bias in previous benchmarks. Animal-Bench includes 13 tasks encompassing both common tasks shared with humans and special tasks relevant to animal conservation, spanning 7 major animal categories and 819 species, comprising a total of 41,839 data entries. To generate this benchmark, we defined a task system centered on animals and proposed an automated pipeline for animal-centric data processing. To further validate the robustness of models against real-world challenges, we utilized a video editing approach to simulate realistic scenarios like weather changes and shooting parameters due to animal movements. We evaluated 8 current multimodal video models on our benchmark and found considerable room for improvement. We hope our work provides insights for the community and opens up new avenues for research in multimodal video models. Our data and code will be released at https://github.com/PRIS-CV/Animal-Bench.
Yinuo Jing, Ruxu Zhang, Kongming Liang, Zhongjiang He, Zhanyu Ma, Jun Guo 0002
NeurIPS3
2024 Mixture-of-Hand-Experts: Repainting the Deformed Hand Images Generated by Diffusion Models
Tianwei Cao, Kongming Liang, Zhongjiang He, Hao Sun 0015, Zhanyu Ma
PRCV (5)3
2024 Evaluating Attribute Comprehension in Large Vision-Language Models
Haiwen Zhang, Zixi Yang, Zheqi He, Kongming Liang, Zhanyu Ma
PRCV (5)6
2024 Learning Dynamic Prototypes for Visual Pattern Debiasing
abstract
Abstract Deep learning has achieved great success in academic benchmarks but fails to work effectively in the real world due to the potential dataset bias. The current learning methods are prone to inheriting or even amplifying the bias present in a training dataset and under-represent specific demographic groups. More recently, some dataset debiasing methods have been developed to address the above challenges based on the awareness of protected or sensitive attribute labels. However, the number of protected or sensitive attributes may be considerably large, making it laborious and costly to acquire sufficient manual annotation. To this end, we propose a prototype-based network to dynamically balance the learning of different subgroups for a given dataset. First, an object pattern embedding mechanism is presented to make the network focus on the foreground region. Then we design a prototype learning method to discover and extract the visual patterns from the training data in an unsupervised way. The number of prototypes is dynamic depending on the pattern structure of the feature space. We evaluate the proposed prototype-based network on three widely used polyp segmentation datasets with abundant qualitative and quantitative experiments. Experimental results show that our proposed method outperforms the CNN-based and transformer-based state-of-the-art methods in terms of both effectiveness and fairness metrics. Moreover, extensive ablation studies are conducted to show the effectiveness of each proposed component and various parameter values. Lastly, we analyze how the number of prototypes grows during the training process and visualize the associated subgroups for each learned prototype. The code and data will be released at https://github.com/zijinY/dynamic-prototype-debiasing .
Kongming Liang, Zijin Yin, Min Min, Zhanyu Ma, Jun Guo 0002
Int. J. Comput. Vis.1
2024 Semi-Supervised Learning for FGVC With Out-of-Category Data
abstract
Despite great strides made on fine-grained visual classification (FGVC), current methods are still heavily reliant on fully-supervised paradigms where ample expert labels are called for. Semi-supervised learning (SSL) techniques, acquiring knowledge from unlabeled data, provide a considerable means forward and have shown great promise for coarse-grained problems. However, exiting SSL paradigms mostly assume in-category (i.e., category-aligned) unlabeled data, which hinders their effectiveness when re-proposed on FGVC. In this paper, we put forward a novel design specifically aimed at making out-of-category data work for semi-supervised FGVC. We work off an important assumption that all fine-grained categories naturally follow a hierarchical structure (e.g., the phylogenetic tree of "Aves" that covers all bird species). It follows that, instead of operating on individual samples, we can instead predict sample relations within this tree structure as the optimization goal of SSL. Beyond this, we further introduced two strategies uniquely brought by these tree structures to achieve inter-sample consistency regularization and reliable pseudo-relation. Our experimental results reveal that (i) the proposed method yields good robustness against out-of-category data, and (ii) it can be equipped with prior arts, boosting their performance thus yielding state-of-the-art results.
Ruoyi Du, Dongliang Chang, Zhanyu Ma, Kongming Liang, Yi-Zhe Song, Jun Guo 0002
IEEE Trans. Pattern Anal. Mach. Intell.4
2023 On-the-Fly Category Discovery
abstract
Although machines have surpassed humans on visual recognition problems, they are still limited to providing closed-set answers. Unlike machines, humans can cognize novel categories at the first observation. Novel category discovery (NCD) techniques, transferring knowledge from seen categories to distinguish unseen categories, aim to bridge the gap. However, current NCD methods assume a transductive learning and offline inference paradigm, which restricts them to a predefined query set and renders them unable to deliver instant feedback. In this paper, we study on-the-fly category discovery (OCD) aimed at making the model instantaneously aware of novel category samples (i.e., enabling inductive learning and streaming inference). We first design a hash coding-based expandable recognition model as a practical baseline. Afterwards, noticing the sensitivity of hash codes to intra-category variance, we further propose a novel Sign-Magnitude dIsentangLEment (SMILE) architecture to alleviate the disturbance it brings. Our experimental results demonstrate the superiority of SMILE against our baseline model and prior art. Our code is available at https://github.com/PRIS-CV/On-the-fly-Category-Discovery.
Ruoyi Du, Dongliang Chang, Kongming Liang, Timothy M. Hospedales, Yi-Zhe Song, Zhanyu Ma
CVPR3
2023 Semantic Centralized Contrastive Learning for Unsupervised Hashing
abstract
Contrastive learning has shown its potential in many unsupervised tasks, including hashing. However, the representations obtained by contrastive learning generally fail to produce no-table margins between semantic classes. Different semantic samples around the boundary are likely to collide into the same hash code. In this paper, we propose a novel Semantic Centralized Contrastive Hashing (SCCH) to allow the learned features closer to their semantic centers and more applicable to hashing. Specifically, a semantic centralization strategy is proposed by pulling strongly augmented samples towards weakly augmented ones since the weak are closer to semantic centers than the strong. Moreover, quantization directly after contrastive learning would damage the learned similarity relationship. We provide a solution to eliminate the mismatch of similarity metrics between contrastive learning and hashing mapping. Extensive experiments on three benchmark datasets demonstrate that SCCH outperforms the existing state-of-the-art methods.
Fengming Liang, Changlin Fan, Kongming Liang
ICASSP4
2023 Super-Resolution Information Enhancement for Crowd Counting
abstract
Crowd counting is a challenging task due to the heavy occlusions, scales, and density variations. Existing methods handle these challenges effectively while ignoring low-resolution (LR) circumstances. The LR circumstances weaken the counting performance deeply for two crucial reasons: 1) limited detail information; 2) overlapping head regions accumulate in density maps and result in extreme ground-truth values. An intuitive solution is to employ super-resolution (SR) pre-processes for the input LR images. However, it complicates the inference steps and thus limits application potentials when requiring real-time. We propose a more elegant method termed Multi-Scale Super-Resolution Module (MSSRM). It guides the network to estimate the lost details and enhances the detailed information in the feature space. Noteworthy that the MSSRM is plug-in plug-out and deals with the LR problems with no inference cost. As the proposed method requires SR labels, we further propose a Super-Resolution Crowd Counting dataset (SR-Crowd). Extensive experiments on three datasets demonstrate the superiority of our method. The code will be available at https://github.com/PRIS-CV/MSSRM.git.
Wei Xu 0037, Dingkang Liang, Zhanyu Ma, Kongming Liang, Ling Jin 0004
ICASSP5
2023 Multi-Head Uncertainty Inference for Adversarial Attack Detection
abstract
Deep neural networks (DNNs) are sensitive and susceptible to tiny perturbations by adversarial attacks which cause erroneous predictions. Various methods, including adversarial defense and uncertainty inference (UI), have been developed to overcome adversarial attacks in recent years. In this paper, we propose a multi-head uncertainty inference (MH-UI) framework for detecting adversarial attack examples. We adopt a multi-head architecture with multiple prediction heads (i.e., classifiers) to obtain predictions from different depths in the DNNs and introduce shallow information for the UI. Using independent heads at different depths, the normalized predictions are assumed to follow the same Dirichlet distribution, and we estimate the distribution parameter of it by moment matching. Cognitive uncertainty brought by the adversarial attacks will be reflected and amplified in the distribution. Experimental results show that the proposed MH-UI framework has good performance in different settings of adversarial attack detection tasks.
Songyun Yang, Jiyang Xie 0001, Zhongwei Si, Ke Zhang 0005, Kongming Liang
ICASSP7
2023 Semantic Memory Guided Image Representation for Polyp Segmentation
abstract
Polyp segmentation is important in the early diagnosis and treatment of colorectal cancer. Since polyps vary in shape, size, color, and texture, accurate polyp segmentation is very challenging. One promising solution is to model the contextual relation for each pixel. However, previous methods only focus on learning the dependencies between the position within an individual image and ignore the contextual relation across different images. In this paper, we propose a memory-based feature enhancement module to capture the cross-image contextual relations. Specifically, we first present a polyp-centric representation. Then a semantic memory is designed to extract the polyp prototypes across different images. The feature at one position can be further enhanced by the contextual embeddings stored in the semantic memory. The enhanced feature is propagated into the features of the previous levels as the multi-scale guidance. The experimental results show that our method achieves better performance than other state-of-the-art methods.
Zijin Yin, Runpu Wei, Kongming Liang, Yiyang Lin, Zhanyu Ma, Min Min, Jun Guo 0002
ICASSP3
2023 Self-Enhanced Training Framework for Referring Expression Grounding
abstract
Weakly-supervised referring expression grounding (REG) aims at locating the image region described by a query sentence, where the mapping between the referential region and query is not available during the training stage. Noticing the significant gap between the fully- and weakly-supervised approaches, we develop a Self-Enhanced Training(SET) framework in this paper. Specifically, we first train the network under a weakly-supervised setting. Then, the model outputs are collected and filtered according to the confidence score and serve as pseudo-labels. Finally, with the help of these pseudo-labels, we tune the model under a fully-supervised setting. The SET framework provides a simple way of generating pseudo-labels that build a bridge between weak and full supervision. Experimental results demonstrate that model trained through our SET framework outperforms existing traditional methods on RefCOCO, RefCOCO+, and RefCOCOg datasets. The code is available at https://github.com/HTDL98/SET-framework.
Ruoyi Du, Kongming Liang, Zhanyu Ma
ICIP3
2023 Attribute Learning with Knowledge Enhanced Partial Annotations
abstract
Under limited annotation cost, large-scale attribute learning datasets only contain partial labels for each image. The conventional methods treat the un-annotated attributes as negative or ignore their loss without considering the associated knowledge. In this paper, we present a knowledge enhanced selective loss for partially labeled attribute learning. Given a visual instance, we investigate the object-attribute co-occurrence as internal knowledge to subdivide the unannotated attributes into feasible and infeasible sets. Based on that, we can enhance the model to focus on the learning of feasible un-annotated attributes and remove the distraction from the infeasible ones. Besides the internal knowledge, we adopt external knowledge to excavate the unseen object-attribute pairs. Experimental results show that our proposed loss can achieve state-of-the-art performance on the newly cleaned VAW2 dataset that contains 170,407 instances, 1763 objects, and 591 attributes. The code and VAW2 dataset are available at https://github.com/GriffinLiang/seal.
Kongming Liang, Wei Chen 0071, Zhanyu Ma, Jun Guo 0002
ICIP1
2023 Ariadne's Thread: Using Text Prompts to Improve Segmentation of Infected Areas from Chest X-ray Images
Mengqiu Xu, Kongming Liang, Kaixin Chen 0001, Ming Wu 0001
MICCAI (4)3
2023 Category-Specific Prompts for Animal Action Recognition with Pretrained Vision-Language Models
abstract
Animal action recognition has a wide range of applications. However, the field largely remains unexplored due to the greater challenges compared to human action recognition, such as lack of annotated training data, large intra-class variation, and interference of cluttered background. Most of the existing methods directly apply human action recognition techniques, which essentially require a large amount of annotated data. In recent years, contrastive vision-language pretraining has demonstrated strong zero-shot generalization ability and has been used for human action recognition. Inspired by the success, we develop a highly performant action recognition framework based on the CLIP model. Our model addresses the above challenges via a novel category-specific prompting module to generate adaptive prompts for both text and video based on the animal category detected in input videos. On one hand, it can generate more precise and customized textual descriptions for each action and animal category pair, being helpful in the alignment of textual and visual space. On the other hand, it allows the model to focus on video features of the target animal in the video and reduce the interference of video background noise. Experimental results demonstrate that our method outperforms five previous action recognition methods on the Animal Kingdom dataset and has shown best generalization ability on unseen animals.
Yinuo Jing, Chunyu Wang 0001, Ruxu Zhang, Kongming Liang, Zhanyu Ma
ACM Multimedia4
2023 Hierarchical Visual Attribute Learning in the Wild
abstract
Observing objects' attributes at different levels of detail is a fundamental aspect of how humans perceive and understand the world around them. Existing studies focused on attribute prediction in a flat way, but they overlook the underlying attribute hierarchy, e.g., navy blue is a subcategory of blue. In recent years, large language models, e.g., ChatGPT, have emerged with the ability to perform an extensive range of natural language processing tasks like text generation and classification. The factual knowledge learned by LLM can assist us build the hierarchical relations of visual attributes in the wild. Based on that, we propose a model called the object-specific attribute relation net, which takes advantage of three types of relations among attributes - positive, negative, and hierarchical - to better facilitate attribute recognition in images. Guided by the extracted hierarchical relations, our model can predict attributes from coarse to fine. Additionally, we introduce several evaluation metrics for attribute hierarchy to comprehensively assess the model's ability to comprehend hierarchical relations. Our extensive experiments demonstrate that our proposed hierarchical annotation brings improvements to the model's understanding of hierarchical relations of attributes, and the object-specific attribute relation net can recognize visual attributes more accurately.
Kongming Liang, Haiwen Zhang, Zhanyu Ma, Jun Guo 0002
ACM Multimedia1
2023 Focus the Overlapping Problem on Few-Shot Object Detection via Multiple Predictions
Mandan Guan, Wenqing Yu, Yurong Guo 0001, Keyan Huang, Jiaxun Zhang, Kongming Liang, Zhanyu Ma
PRCV (2)6
2023 Plugging Stylized Controls in Open-Stylized Image Captioning
Yixiao Zheng, Ruoyi Du, Yiming Zhang 0025, Kongming Liang, Zhanyu Ma
PRCV (1)5
2023 Image Generation Based Intra-class Variance Smoothing for Fine-Grained Visual Classification
Ruoyi Du, Kongming Liang, Wei Chen 0071, Zhanyu Ma
PRCV (6)3
2023 Graph Convolution Based Cross-Network Multiscale Feature Fusion for Deep Vessel Segmentation
abstract
Vessel segmentation is widely used to help with vascular disease diagnosis. Vessels reconstructed using existing methods are often not sufficiently accurate to meet clinical use standards. This is because 3D vessel structures are highly complicated and exhibit unique characteristics, including sparsity and anisotropy. In this paper, we propose a novel hybrid deep neural network for vessel segmentation. Our network consists of two cascaded subnetworks performing initial and refined segmentation respectively. The second subnetwork further has two tightly coupled components, a traditional CNN-based U-Net and a graph U-Net. Cross-network multi-scale feature fusion is performed between these two U-shaped networks to effectively support high-quality vessel segmentation. The entire cascaded network can be trained from end to end. The graph in the second subnetwork is constructed according to a vessel probability map as well as appearance and semantic similarities in the original CT volume. To tackle the challenges caused by the sparsity and anisotropy of vessels, a higher percentage of graph nodes are distributed in areas that potentially contain vessels while a higher percentage of edges follow the orientation of potential nearby vessels. Extensive experiments demonstrate our deep network achieves state-of-the-art 3D vessel segmentation performance on multiple public and in-house datasets.
Gangming Zhao, Kongming Liang, Chengwei Pan, Fandong Zhang, Xianpeng Wu, Xinyang Hu, Yizhou Yu
IEEE Trans. Medical Imaging2
2022 Learning Invariant Visual Representations for Compositional Zero-Shot Learning
Tian Zhang 0029, Kongming Liang, Ruoyi Du, Zhanyu Ma, Jun Guo 0002
ECCV (24)2
2022 Domain Generalization via Frequency-domain-based Feature Disentanglement and Interaction
abstract
Adaptation to out-of-distribution data is a meta-challenge for all statistical learning algorithms that strongly rely on the i.i.d. assumption. It leads to unavoidable labor costs and confidence crises in realistic applications. For that, domain generalization aims at mining domain-irrelevant knowledge from multiple source domains that can generalize to unseen target domains. In this paper, by leveraging the frequency domain of an image, we uniquely work with two key observations: (i) the high-frequency information of an image depicts object edge structure, which preserves high-level semantic information of the object is naturally consistent across different domains, and (ii) the low-frequency component retains object smooth structure, while this information is susceptible to domain shifts. Motivated by the above observations, we introduce (i) an encoder-decoder structure to disentangle high- and low-frequency features of an image, (ii) an information interaction mechanism to ensure the helpful knowledge from both two parts can cooperate effectively, and (iii) a novel data augmentation technique that works on the frequency domain to encourage the robustness of frequency-wise feature disentangling. The proposed method obtains state-of-the-art performance on three widely used domain generalization benchmarks (Digit-DG, Office-Home, and PACS).
Jingye Wang, Ruoyi Du, Dongliang Chang, Kongming Liang, Zhanyu Ma
ACM Multimedia4
2022 Complex Scenario-Oriented Fine-Grained Visual Classification Platform
abstract
In recent years, fine-grained visual classification (FGVC) algorithms have achieved excellent performance across a variety of datasets. However, it is still rare to see these algorithms applied in daily life. The main reasons for this are i) the algorithms are developed based on different design guidelines and cannot be deployed in the same environment; ii) there is not a simple and efficient platform to present the algorithm's results to the user - the accuracy is meaningless to the users. To address the above problem, we built a complex scenario-oriented fine-grained visual classification platform. The platform consists of a PyTorch-based fine-grained visual recognition algorithm library (FGL) and a WeChat applet-based user interaction module (WEM). We can quickly develop new algorithms or readily apply existing algorithms in the same environment through FGL. Driven by FGL, the WEM enables users to achieve fine-grained recognition of complex scenes interactively. In addition to showing the user the fine-grained labels of objects, we will also show how the model makes decisions to help the user master the ability to recognise the fine-grained object so that everyone can become a domain expert. A video demo shows an example of the proposed platform in a real-world scenario: https://reurl.cc/rRZE7O.
Dongliang Chang, Junhan Chen, Ruoyi Du, Wenqing Yu, Yujun Tong, Kongming Liang, Yi-Zhe Song, Zhanyu Ma
MMSP8
2022 Multi-modal Human-machine Conversation System for Real Physical World
abstract
Enabling machines to process multi-modal information and understand the real physical world is an important step to achieving free human-machine conversation. However, previous human-machine conversation systems are mostly limited to single-modal (e.g., chat robot), single-round (e.g., visual Q & A), and static visual information (e.g., visual dialogue). To address the above problem, we develop a multi-modal human-machine conversation system for specific application scenarios. The system includes two modules: (i) an interactive visual grounding module that can actively disambiguate user's queries, and (ii) an interactive fine-grained recognition module that can model objects in the 3D environment and actively ask for missing visual information. A video demo of our system under the automobile sales scenario can be found here11https://drive.google.com/file/d/1IfBsMKq55ryLOZchIT6G-R5CyWnC3Skiew?usp=sharing.
Shibo Nie, Mandan Guan, Ruoyi Du, Dongliang Chang, Kongming Liang, Zhanyu Ma
MMSP7
2022 Dual-granularity feature alignment for cross-modality person re-identification
Junhui Yin, Zhanyu Ma, Jiyang Xie 0001, Shibo Nie, Kongming Liang, Jun Guo 0002
Neurocomputing5
2021 Symmetry-Enhanced Attention Network for Acute Ischemic Infarct Segmentation with Non-contrast CT Images
Kongming Liang, Kai Han 0010, Xiuli Li, Xiaoqing Cheng, Yizhou Wang 0001, Yizhou Yu
MICCAI (7)1
2021 Improved Brain Lesion Segmentation with Anatomical Priors from Healthy Subjects
Xiangzhu Zeng, Kongming Liang, Yizhou Yu, Chuyang Ye
MICCAI (1)3
2021 Cross-layer Navigation Convolutional Neural Network for Fine-grained Visual Classification
abstract
Fine-grained visual classification (FGVC) aims to classify sub-classes of objects in the same super-class (e.g., species of birds, models of cars). For the FGVC tasks, the essential solution is to find discriminative subtle information of the target from local regions. Traditional FGVC models preferred to use the refined features, i.e., high-level semantic information for recognition and rarely use low-level information. However, it turns out that low-level information which contains rich detail information also has effect on improving performance. Therefore, in this paper, we propose cross-layer navigation convolutional neural network for feature fusion. First, the feature maps extracted by the backbone network are fed into a convolutional long short-term memory model sequentially from high-level to low-level to perform feature aggregation. Then, attention mechanisms are used after feature fusion to extract spatial and channel information while linking the high-level semantic information and the low-level texture features, which can better locate the discriminative regions for the FGVC. In the experiments, three commonly used FGVC datasets, including CUB-200-2011, Stanford-Cars, and FGVC-Aircraft datasets, are used for evaluation and we demonstrate the superiority of the proposed method by comparing it with other referred FGVC methods to show that this method achieves superior results. https://github.com/PRIS-CV/CN-CNN.git
Chenyu Guo, Jiyang Xie 0001, Kongming Liang, Zhanyu Ma
MMAsia3
2020 Context-Aware Refinement Network Incorporating Structural Connectivity Prior for Brain Midline Delineation
Kongming Liang, Yizhou Yu, Yizhou Wang 0001
MICCAI (7)2
2020 Visual concept conjunction learning with recurrent neural networks
Kongming Liang, Hong Chang 0001, Shiguang Shan, Xilin Chen 0001
Neurocomputing1
2019 Unifying Visual Attribute Learning with Object Recognition in a Multiplicative Framework
abstract
Attributes are mid-level semantic properties of objects. Recent research has shown that visual attributes can benefit many typical learning problems in computer vision community. However, attribute learning is still a challenging problem as the attributes may not always be predictable directly from input images and the variation of visual attributes is sometimes large across categories. In this paper, we propose a unified multiplicative framework for attribute learning, which tackles the key problems. Specifically, images and category information are jointly projected into a shared feature space, where the latent factors are disentangled and multiplied to fulfil attribute prediction. The resulting attribute classifier is category-specific instead of being shared by all categories. Moreover, our model can leverage auxiliary data to enhance the predictive ability of attribute classifiers, which can reduce the effort of instance-level attribute annotation to some extent. By integrated into an existing deep learning framework, our model can both accurately predict attributes and learn efficient image representations. Experimental results show that our method achieves superior performance on both instance-level and category-level attribute prediction. For zero-shot learning based on visual attributes and human-object interaction recognition, our method can improve the state-of-the-art performance on several widely used datasets.
Kongming Liang, Hong Chang 0001, Bingpeng Ma, Shiguang Shan, Xilin Chen 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2018 Visual Relationship Detection With Deep Structural Ranking
abstract
Visual relationship detection aims to describe the interactions between pairs of objects. Different from individual object learning tasks, the number of possible relationships are much larger, which makes it hard to explore only based on the visual appearance of objects. In addition, due to the limited human effort, the annotations for visual relationships are usually incomplete which increases the difficulty of model training and evaluation. In this paper, we propose a novel framework, called Deep Structural Ranking, for visual relationship detection. To complement the representation ability of visual appearance, we integrate multiple cues for predicting the relationships contained in an input image. Moreover, we design a new ranking objective function by enforcing the annotated relationships to have higher relevance scores. Unlike previous works, our proposed method can both facilitate the co-occurrence of relationships and mitigate the incompleteness problem. Experimental results show that our proposed method outperforms the state-of-the-art on the two widely used datasets. We also demonstrate its superiority in detecting zero-shot relationships.
Kongming Liang, Yuhong Guo, Hong Chang 0001, Xilin Chen 0001
AAAI1
2017 Incomplete Attribute Learning with auxiliary labels
abstract
Visual attribute learning is a fundamental and challenging problem for image understanding. Considering the huge semantic space of attributes, it is economically impossible to annotate all their presence or absence for a natural image via crowd-sourcing. In this paper, we tackle the incompleteness nature of visual attributes by introducing auxiliary labels into a novel transductive learning framework. By jointly predicting the attributes from the input images and modeling the relationship of attributes and auxiliary labels, the missing attributes can be recovered effectively. In addition, the proposed model can be solved efficiently in an alternative way by optimizing quadratic programming problems and updating parameters in closed-form solutions. Moreover, we propose and investigate different methods for acquiring auxiliary labels. We conduct experiments on three widely used attribute prediction datasets. The experimental results show that our proposed method can achieve the state-of-the-art performance with access to partially observed attribute annotations.
Kongming Liang, Yuhong Guo, Hong Chang 0001, Xilin Chen 0001
IJCAI1
2016 Attribute Conjunction Learning with Recurrent Neural Network
Kongming Liang, Hong Chang 0001, Shiguang Shan, Xilin Chen 0001
ECML/PKDD (1)1
2015 A Unified Multiplicative Framework for Attribute Learning
abstract
Attributes are mid-level semantic properties of objects. Recent research has shown that visual attributes can benefit many traditional learning problems in computer vision community. However, attribute learning is still a challenging problem as the attributes may not always be predictable directly from input images and the variation of visual attributes is sometimes large across categories. In this paper, we propose a unified multiplicative framework for attribute learning, which tackles the key problems. Specifically, images and category information are jointly projected into a shared feature space, where the latent factors are disentangled and multiplied for attribute prediction. The resulting attribute classifier is category-specific instead of being shared by all categories. Moreover, our method can leverage auxiliary data to enhance the predictive ability of attribute classifiers, reducing the effort of instance-level attribute annotation to some extent. Experimental results show that our method achieves superior performance on both instance-level and category-level attribute prediction. For zero-shot learning based on attributes, our method significantly improves the state-of-the-art performance on AwA dataset and achieves comparable performance on CUB dataset.
Kongming Liang, Hong Chang 0001, Shiguang Shan, Xilin Chen 0001
ICCV1
2014 Representation Learning with Smooth Autoencoder
Kongming Liang, Hong Chang 0001, Zhen Cui 0001, Shiguang Shan, Xilin Chen 0001
ACCV (2)1