EDBT 2026 Demo / reviewers in the wild / expert
Shaobo Min
dblp:174/2964
· DBLP profile ↗
26ranked-venue papers
9as first author
18since 2021 · last 2026
0000-0002-7700-2149ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 23 · 8 first-author · 16 since 2021Artificial intelligence and machine learning · 11 · 3 first-author · 8 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | S²Flow: Towards Fast and Authentic Training-Free High-Resolution Video GenerationabstractRectified flow models have shown strong potential in high-fidelity video generation, yet extending them to high-resolution remains challenging due to the high cost of full attention and error accumulation in the ODE-solving process. In this paper, we propose S^2Flow, a training-free framework that enables efficient and authentic high-resolution video generation by jointly exploring Flow-guided Sparse attention and Second-order ODE solution. Specifically, S^2Flow exploits and transfers the semantic and structural information from the low-resolution flow trajectory to guide the high-resolution flow in two aspects. First, S^2Flow dynamically captures the sparse patterns of the spatio-temporal attention maps from low-resolution videos to construct localized 3D windows, enabling efficient window attention in high-resolution inference. This can significantly reduce redundant computation while preserving contextual dependencies. Second, S^2Flow adopts a second-order ODE solver based on Taylor expansion, where the high-order derivative is approximated via central difference from the low-resolution flow, facilitating accurate high-resolution denoising. Extensive experiments on VBench dataset demonstrate that S^2Flow outperforms prior methods in both visual quality and inference speed, enabling 4x acceleration on 2560x1536 video generation. Chaoqun Wang 0011, Shaobo Min, Xu Yang 0019 |
AAAI | 2 |
| 2025 | Infinite-Canvas: Higher-Resolution Video Outpainting with Extensive Content GenerationabstractThis paper explores higher-resolution video outpainting with extensive content generation. We point out common issues faced by existing methods when attempting to largely outpaint videos: the generation of low-quality content and limitations imposed by GPU memory. To address these challenges, we propose a diffusion-based method called Infinite-Canvas. It builds upon two core designs. First, instead of employing the common practice of "single-shot" outpainting, we distribute the task across spatial windows and seamlessly merge them. It allows us to outpaint videos of any size and resolution without being constrained by GPU memory. Second, the source video and its relative positional relation are injected into the generation process of each window. It makes the generated spatial layout within each window harmonize with the source video. Coupling with these two designs enables us to generate higher-resolution outpainting videos with rich content while keeping spatial and temporal consistency. Infinite-Canvas excels in large-scale video outpainting, e.g., from 512 × 512 to 1152 × 2048 (9×), while producing high-quality and aesthetically pleasing results. It achieves the best quantitative results across various resolution and scale setups. The code is available at https://github.com/mayuelala/FollowYourCanvas. Qihua Chen, Yue Ma 0016, Hongfa Wang, Junkun Yuan, Qi Tian 0003, Shaobo Min, Qifeng Chen 0001, Wei Liu 0005 |
AAAI | 8 |
| 2025 | Towards Multiple Character Image Animation Through Enhancing Implicit DecouplingabstractControllable character image animation has a wide range of applications. Although existing studies have consistently improved performance, challenges persist in the field of character image animation, particularly concerning stability in complex backgrounds and tasks involving multiple characters. To address these challenges, we propose a novel multi-condition guided framework for character image animation, employing several well-designed input modules to enhance the implicit decoupling capability of the model. First, the optical flow guider calculates the background optical flow map as guidance information, which enables the model to implicitly learn to decouple the background motion into background constants and background momentum during training, and generate a stable background by setting zero background momentum during inference. Second, the depth order guider calculates the order map of the characters, which transforms the depth information into the positional information of multiple characters. This facilitates the implicit learning of decoupling different characters, especially in accurately separating the occluded body parts of multiple characters. Third, the reference pose map is input to enhance the ability to decouple character texture and pose information in the reference image. Furthermore, to fill the gap of fair evaluation of multi-character image animation, we propose a new benchmark comprising about 4,000 frames. Extensive qualitative and quantitative evaluations demonstrate that our method excels in generating high-quality character animations, especially in scenarios of complex backgrounds and multiple characters. Jingyun Xue, Hongfa Wang, Qi Tian 0003, Yue Ma 0016, Andong Wang, Zhiyuan Zhao 0002, Shaobo Min, Kaihao Zhang, Harry Shum, Wei Liu 0005, Mengyang Liu, Wenhan Luo |
ICLR | 7 |
| 2025 | CausalCtrl: Causality-Aware Control Framework for Text-Guided Visual EditingabstractText-guided visual editing aims to modify visual content according to a target prompt while faithfully preserving the structure and identity of the source image or video. However, existing methods ignore confounding effects brought from the pretrained model, i.e., harmful biases learned from the pretraining datasets, leading to spurious correlations during the editing processing. To address this issue, we introduce CausalCtrl, a novel training-free framework that reformulates text-guided visual editing from a causal inference perspective. The core idea is to leverage frontdoor adjustment to estimate the interventional distribution of the output, effectively blocking the influence of hidden confounders introduced by the pretrained model. Specifically, we first design a dual-branch inversion mechanism that disentangles the source content and target semantics into two separate latent embeddings to simplify the sampling space of interventional operation, and perform unbiased denoising through their controlled interaction. Besides, we propose a Structured Attention Injection Module (SAIM) that adaptively identifies and amplifies dominant attention heads using a lightweight SVD-based top-K selection strategy. Extensive experiments on several challenging image and video editing benchmarks demonstrate that CausalCtrl consistently outperforms existing methods in both target semantic alignment and source content preservation, validating the effectiveness of causal intervention in this task. Haoxiang Cao, Chaoqun Wang 0011, Yongwen Lai, Shaobo Min, Xuejin Chen |
ACM Multimedia | 4 |
| 2025 | Robust visual place recognition with adaptive deformable token aggregation
Chaoqun Wang 0011, Shaobo Min, Xuejin Chen |
Comput. Graph. | 3 |
| 2024 | BadCLIP: Trigger-Aware Prompt Learning for Backdoor Attacks on CLIPabstractContrastive Vision-Language Pre-training, known as CLIP, has shown promising effectiveness in addressing downstream image recognition tasks. However, recent works revealed that the CLIP model can be implanted with a downstream-oriented backdoor. On downstream tasks, one victim model performs well on clean samples but predicts a specific target class whenever a specific trigger is present. For injecting a backdoor, existing attacks depend on a large amount of additional data to maliciously fine-tune the entire pre-trained CLIP model, which makes them inapplicable to data-limited scenarios. In this work, motivated by the recent success of learnable prompts, we address this problem by injecting a backdoor into the CLIP model in the prompt learning stage. Our method named BadCLIP is built on a novel and effective mechanism in backdoor attacks on CLIP, i.e., influencing both the image and text encoders with the trigger. It consists of a learnable trigger applied to images and a trigger-aware context generator, such that the trigger can change text features via trigger-aware prompts, resulting in a powerful and generalizable attack. Extensive experiments conducted on 11 datasets verify that the clean accuracy of BadCLIP is similar to those of advanced prompt learning methods and the attack success rate is higher than 99% in most cases. BadCLIP is also generalizable to unseen classes, and shows a strong generalization capability under cross-dataset and cross-domain settings. The code is available at https://github.com/jiawangbai/BadCLIP. Jiawang Bai, Kuofeng Gao, Shaobo Min, Shutao Xia, Zhifeng Li 0001, Wei Liu 0005 |
CVPR | 3 |
| 2024 | Towards Discriminative Feature Generation for Generalized Zero-Shot LearningabstractGeneralized Zero-Shot Learning (GZSL) aims to recognize both seen and unseen categories by establishing visual and semantic relations. Recently, generation-based methods that focus on synthesizing fictitious visual features from corresponding attributes have gained significant attention. However, these generated features often lack discriminative capabilities due to inadequate training of the generative model. To address this issue, we propose a novel Discriminative Enhanced Network (DENet) to harness the potential of the generative model by adapting the training features and imposing constraints on the generated features. Our approach incorporates three pivotal modules: (1) Before the generative network training, we implement a Pre-Tuning Module (PTM) to eliminate irrelevant background noise in the raw features extracted from a fixed CNN backbone. Therefore, PTM can provide tuned training features without redundant noise for generative model. (2) During the generative network training, we propose an Asymmetry Cross-authenticity Contrastive (AC2) loss to group visual features of the same category while repel features from different categories by optimizing a large number of sample pairs. Additionally, we incorporate intra-class and relation-specific inter-class boundaries within the AC2 loss to enrich sample diversity and preserve valid semantic information. (3) Also within the generative network training, a Dual-semantic Alignment Module (DAM) is designed to align visual features with both attributes and label embeddings, enabling the model to learn attribute-related information and discriminative extended semantics. Experiments on four standard benchmarks demonstrate that our approach learns more discriminative features and surpasses the existing methods. Jiannan Ge, Hongtao Xie 0001, Pandeng Li, Lingxi Xie, Shaobo Min, Yongdong Zhang 0001 |
IEEE Trans. Multim. | 5 |
| 2022 | Dual-Stream Knowledge-Preserving Hashing for Unsupervised Video Retrieval
Pandeng Li, Hongtao Xie 0001, Jiannan Ge, Lei Zhang 0119, Shaobo Min, Yongdong Zhang 0001 |
ECCV (14) | 5 |
| 2022 | Dual Part Discovery Network for Zero-Shot LearningabstractZero-Shot Learning (ZSL) aims to recognize unseen classes by transferring knowledge from seen classes. Recent methods focus on learning a common semantic space to align visual and attribute information. However, they always over-relied on provided attributes and ignored the category discriminative information that contributes to accurate unseen class recognition, resulting in weak transferability. To this end, we propose a novel Dual Part Discovery Network (DPDN) that considers both attribute and category discriminative information by discovering attribute-guided parts and category-guided parts simultaneously to improve knowledge transfer. Specifically, for attribute-guided parts discovery, DPDN can localize the regions with specific attribute information and significantly bridge the gap between visual and semantic information guided by the given attributes. For category-guided parts discovery, the local parts are explored to discover other important regions that bring latent crucial details ignored by attributes, with the guidance of adaptive category prototypes. To better mine the transferable knowledge, we impose class correlations constraints to regularize the category prototypes. Finally, attribute- and category-guided parts complement each other and provide adequate discriminative subtle information for more accurate unseen class recognition. Extensive experimental results demonstrate that DPDN can discover discriminative parts and outperform state-of-the-art methods on three standard benchmarks. Jiannan Ge, Hongtao Xie 0001, Shaobo Min, Pandeng Li, Yongdong Zhang 0001 |
ACM Multimedia | 3 |
| 2022 | Deep Fourier Ranking Quantization for Semi-Supervised Image RetrievalabstractTo reduce the extreme label dependence of supervised product quantization methods, the semi-supervised paradigm usually employs massive unlabeled data to assist in regularizing deep networks, thereby improving model performance. However, the existing method focuses on the overall distribution consistency between unlabeled data and class prototypes, while ignoring subtle individual variances between unlabeled instances. Therefore, the local neighborhood structure is not fully explored, which will cause the model to easily overfit in the training set. In this paper, we introduce a new Fourier perspective to alleviate this issue by exploring the semantic relations between unlabeled instances in a self-supervised manner. Specifically, based on Fourier Transform, we first design a Phase Mixing (PM) strategy, which can manipulate the mixing area and values of the phase component between two images to control the proportion of semantic information. In this way, we can construct multi-level similarity neighbors naturally for unlabeled data. Then, a ranking quantization loss is formulated to perceive multi-level semantic variances in neighbor instances, which improves the robustness and generalization of the model. Extensive experiments in three different semi-supervised settings show that our method outperforms existing state-of-the-art methods by averaged 3.95% improvement on four datasets. Pandeng Li, Hongtao Xie 0001, Shaobo Min, Jiannan Ge, Xun Chen 0001, Yongdong Zhang 0001 |
IEEE Trans. Image Process. | 3 |
| 2022 | Online Residual Quantization Via Streaming Data Correlation PreservingabstractRecently, the online retrieval task has been receiving widespread attention, which is closely related to many real-world applications. However, existing online retrieval methods based on hashing suffer from two main problems: a) the models tend to be biased towards the current streaming data due to unavailable history streaming data; b) when new streaming data comes in and the hashing functions have been updated, all history binary codes should be recomputed, which takes much computation burden. To address the above two issues, we propose a novel Online Residual Quantization (ORQ) method that can achieve efficient streaming data quantization via the small-scale residual quantization codebooks. For the first problem, we design a residual quantization module by learning multiple residual codebooks to quantize the float streaming data, which effectively reduces the quantization error and enables the binary codes to be easily reconstructed back to original float data. Then, with the reconstructed history data, a balanced affinity matrix is developed to model the semantic relationship,e.g.,similarity and difference, between the history and current data distributions, which can prevent the model from being biased towards the current data distribution. For the second problem, when inputting current streaming data, only the residual codebooks should be updated, instead of the whole history binary codes in hashing-based methods, which significantly reduces the computation burden. Comprehensive experiments on six benchmarks demonstrate that ORQ yields significant improvements (i.e.,1.2%$\sim$4.9% in average mAP) compared to the state-of-the-art methods. Pandeng Li, Hongtao Xie 0001, Shaobo Min, Zhengjun Zha, Yongdong Zhang 0001 |
IEEE Trans. Multim. | 3 |
| 2022 | Attribute-Induced Bias Eliminating for Transductive Zero-Shot LearningabstractTransductive zero-shot learning is designed to recognize unseen categories by aligning both visual and semantic information in a joint embedding space. Four types of domain biases exist in Transductive ZSL,i.e.,visual biasandsemantic biasin two domains, and twovisual-semantic biasesexist in the seen and unseen domains. However, the existing work has only focused on specific components of these topics, leading to severe semantic ambiguity during knowledge transfer. To solve this problem, we propose a novel attribute-induced bias eliminating (AIBE) module for Transductive ZSL. Specifically, for thevisual biasbetween the two domains, the mean-teacher module is first used to bridge the visual representation discrepancy between the two domains using unsupervised learning and unlabeled images. Then, an attentional graph attribute embedding process is proposed to reduce thesemantic biasbetween seen and unseen categories using a graph operation to describe the semantic relationship between categories. To reduce semantic-visual bias in the seen domain, we align the visual center of each category with the corresponding semantic attributes instead of with the individual visual data point, which preserves the semantic relationship in the embedding space. Finally, for the semantic-visual bias in the unseen domain, an unseen semantic alignment constraint is designed to align visual and semantic space using an unsupervised process. The evaluations on several benchmarks demonstrate the effectiveness of the proposed method,e.g.,82.8%/75.5%, 97.1%/82.5%, and 73.2%/52.1% for Conventional/Generalized ZSL settings for CUB, AwA2, and SUN datasets, respectively. Hantao Yao, Shaobo Min, Yongdong Zhang 0001, Changsheng Xu |
IEEE Trans. Multim. | 2 |
| 2021 | Semantic-guided Reinforced Region Embedding for Generalized Zero-Shot LearningabstractGeneralized zero-shot Learning (GZSL) aims to recognize images from either seen or unseen domain, mainly by learning a joint embedding space to associate image features with the corresponding category descriptions. Recent methods have proved that localizing important object regions can effectively bridge the semantic-visual gap. However, these are all based on one-off visual localizers, lacking of interpretability and flexibility. In this paper, we propose a novel Semantic-guided Reinforced Region Embedding (SR2E) network that can localize important objects in the long-term interests to construct semantic-visual embedding space. SR2E consists of Reinforced Region Module (R2M) and Semantic Alignment Module (SAM). First, without the annotated bounding box as supervision, R2M encodes the semantic category guidance into the reward and punishment criteria to teach the localizer serialized region searching. Besides, R2M explores different action spaces during the serialized searching path to avoid local optimal localization, which thereby generates discriminative visual features with less redundancy. Second, SAM preserves the semantic relationship into visual features via semantic-visual alignment and designs a domain detector to alleviate the domain confusion. Experiments on four public benchmarks demonstrate that the proposed SR2E is an effective GZSL method with reinforced embedding space, which obtains averaged 6.1% improvements. Jiannan Ge, Hongtao Xie 0001, Shaobo Min, Yongdong Zhang 0001 |
AAAI | 3 |
| 2021 | Task-Independent Knowledge Makes for Transferable Representations for Generalized Zero-Shot LearningabstractGeneralized Zero-Shot Learning (GZSL) targets recognizing new categories by learning transferable image representations. Existing methods find that, by aligning image representations with corresponding semantic labels, the semantic-aligned representations can be transferred to unseen categories. However, supervised by only seen category labels, the learned semantic knowledge is highly task-specific, which makes image representations biased towards seen categories. In this paper, we propose a novel Dual-Contrastive Embedding Network (DCEN) that simultaneously learns task-specific and task-independent knowledge via semantic alignment and instance discrimination. First, DCEN leverages task labels to cluster representations of the same semantic category by cross-modal contrastive learning and exploring semantic-visual complementarity. Besides task-specific knowledge, DCEN then introduces task-independent knowledge by attracting representations of different views of the same image and repelling representations of different images. Compared to high-level seen category supervision, this instance discrimination supervision encourages DCEN to capture low-level visual knowledge, which is less biased toward seen categories and alleviates the representation bias. Consequently, the task-specific and task-independent knowledge jointly make for transferable representations of DCEN, which obtains averaged 4.1% improvement on four public benchmarks. Chaoqun Wang 0011, Xuejin Chen, Shaobo Min, Xiaoyan Sun 0001, Houqiang Li |
AAAI | 3 |
| 2021 | Dual Progressive Prototype Network for Generalized Zero-Shot LearningabstractGeneralized Zero-Shot Learning (GZSL) aims to recognize new categories with auxiliary semantic information, e.g., category attributes. In this paper, we handle the critical issue of domain shift problem, i.e., confusion between seen and unseen categories, by progressively improving cross-domain transferability and category discriminability of visual representations. Our approach, named Dual Progressive Prototype Network (DPPN), constructs two types of prototypes that record prototypical visual patterns for attributes and categories, respectively. With attribute prototypes, DPPN alternately searches attribute-related local regions and updates corresponding attribute prototypes to progressively explore accurate attribute-region correspondence. This enables DPPN to produce visual representations with accurate attribute localization ability, which benefits the semantic-visual alignment and representation transferability. Besides, along with progressive attribute localization, DPPN further projects category prototypes into multiple spaces to progressively repel visual representations from different categories, which boosts category discriminability. Both attribute and category prototypes are collaboratively learned in a unified framework, which makes visual representations of DPPN transferable and distinctive.Experiments on four benchmarks prove that DPPN effectively alleviates the domain shift problem in GZSL. Chaoqun Wang 0011, Shaobo Min, Xuejin Chen, Xiaoyan Sun 0001, Houqiang Li |
NeurIPS | 2 |
| 2021 | Structure-Guided Deep Video InpaintingabstractA fundamental challenge in video inpainting is the difficulty of generating video contents with fine details, while keeping spatio-temporal coherence in the missing region. Recent studies focus on synthesizing temporally smooth pixels by exploiting the flow information, while ignoring maintaining the semantic structural coherence between frames. This makes them suffer from over-smoothing and blurry contours, which significantly reduce the visual quality of inpainting results. To address this issue, we present a novel structure-guided video inpainting approach that enhances temporal structure coherence to improve video inpainting results. In contrast to directly synthesizing the missing pixel colors, we first complete edges in the missing regions to depict scene structures and object shapes via an edge inpainting network with 3D convolutions. Then, we replenish textures using a coarse-to-fine synthesis network with a structure attention module (SAM), under the guidance of the synthesized edges. Specifically, our SAM is designed to model the semantic correlation between video textures and structural edges to generate more realistic content. Besides, motion flows between neighboring frames are employed to enhance temporal consistency for self-supervision during training the edge inpainting and texture inpainting modules. Consequently, the inpainting results using our approach are visually pleasing with fine details and temporal coherence. Experiments on the YouTubeVOS, DAVIS, and 300VW datasets show that our method obtains state-of-the-art performance under diverse video inpainting settings. Chaoqun Wang 0011, Xuejin Chen, Shaobo Min, Jiaping Wang, Zhengjun Zha |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2021 | A Mutually Attentive Co-Training Framework for Semi-Supervised RecognitionabstractSelf-training plays an important role in practical recognition applications where sufficient clean labels are unavailable. Existing methods focus on generating reliable pseudo labels to retrain a model, while ignoring the importance of improving model reliability to those inevitably mislabeled data. In this paper, we propose a novel Mutually Attentive Co-training Framework (MACF) that can effectively alleviate the negative impacts of incorrect labels on model retraining by exploring deep model disagreements. Specifically, MACF trains two symmetrical sub-networks that have the same input and are connected by several attention modules at different layers. Each attention module analyzes the inferred features from two sub-networks for the same input and feedback attention maps for them to indicate noisy gradients. This is realized by exploring the back-propagation process of incorrect labels at different layers to design attention modules. By multi-layer interception, the noisy gradients caused by incorrect labels can be effectively reduced for both sub-networks, leading to robust training to potential incorrect labels. In addition, a hierarchical distillation strategy is developed to improve the pseudo labels by aggregating the predictions from multi-models and data transformations. The experiments on six general benchmarks, including classification and biomedical segmentation, demonstrate that MACF is much robust to noisy labels than previous methods. Shaobo Min, Xuejin Chen, Hongtao Xie 0001, Zhengjun Zha, Yongdong Zhang 0001 |
IEEE Trans. Multim. | 1 |
| 2021 | Domain-Oriented Semantic Embedding for Zero-Shot LearningabstractZero-Shot Learning (ZSL) targets to recognize images from new classes. Existing methods focus on learning a projection function to associate the visual features and category descriptions in the seen domain, which is directly transferred to the unseen domain. However, due to the inherent domain shift, a single shared projection cannot fully capture the domain difference and similarity, thereby making the unseen samples tend to be recognized as seen categories. In this paper, we propose a novel Domain-Oriented Semantic Embedding (DOSE) network that learns specific projections for different domains to better capture the domain characteristics for unbiased ZSL. Besides a domain-shared projection, DOSE learns two auxiliary domain-specific sub-projections to model the semantic-visual association in respective seen and unseen domains. Specifically, the domain-specific projections are learned in a cycle consistency way to capture domain characteristics, and a domain division constraint is developed to penalize the margin between two domain embeddings. Furthermore, to boost semantic-visual association, a semantic-visual dual attention module is designed to automatically remove trivial information in both visual and semantic embeddings under a co-guidance learning manner. Experiments on four public benchmarks prove that the proposed DOSE is robust to the domain shift problem in ZSL and obtains an averaged 5.6% improvement in terms of harmonic mean. Shaobo Min, Hantao Yao, Hongtao Xie 0001, Zhengjun Zha, Yongdong Zhang 0001 |
IEEE Trans. Multim. | 1 |
| 2020 | Domain-Aware Visual Bias Eliminating for Generalized Zero-Shot LearningabstractGeneralized zero-shot learning aims to recognize images from seen and unseen domains. Recent methods focus on learning a unified semantic-aligned visual representation to transfer knowledge between two domains, while ignoring the effect of semantic-free visual representation in alleviating the biased recognition problem. In this paper, we propose a novel Domain-aware Visual Bias Eliminating (DVBE) network that constructs two complementary visual representations, i.e., semantic-free and semantic-aligned, to treat seen and unseen domains separately. Specifically, we explore cross-attentive second-order visual statistics to compact the semantic-free representation, and design an adaptive margin Softmax to maximize inter-class divergences. Thus, the semantic-free representation becomes discriminative enough to not only predict seen class accurately but also filter out unseen images, i.e., domain detection, based on the predicted class entropy. For unseen images, we automatically search an optimal semantic-visual alignment architecture, rather than manual designs, to predict unseen classes. With accurate domain detection, the biased recognition problem towards the seen domain is significantly reduced. Experiments on five benchmarks for classification and segmentation show that DVBE outperforms existing methods by averaged 5.7% improvement. Shaobo Min, Hantao Yao, Hongtao Xie 0001, Chaoqun Wang 0011, Zhengjun Zha, Yongdong Zhang 0001 |
CVPR | 1 |
| 2020 | Hierarchical Granularity Transfer LearningabstractIn the real world, object categories usually have a hierarchical granularity tree. Nowadays, most researchers focus on recognizing categories in a specific granularity, \emph{e.g.,} basic-level or sub(ordinate)-level. Compared with basic-level categories, the sub-level categories provide more valuable information, but its training annotations are harder to acquire. Therefore, an attractive problem is how to transfer the knowledge learned from basic-level annotations to sub-level recognition. In this paper, we introduce a new task, named Hierarchical Granularity Transfer Learning (HGTL), to recognize sub-level categories with basic-level annotations and semantic descriptions for hierarchical categories. Different from other recognition tasks, HGTL has a serious granularity gap,~\emph{i.e.,} the two granularities share an image space but have different category domains, which impede the knowledge transfer. To this end, we propose a novel Bi-granularity Semantic Preserving Network (BigSPN) to bridge the granularity gap for robust knowledge transfer. Explicitly, BigSPN constructs specific visual encoders for different granularities, which are aligned with a shared semantic interpreter via a novel subordinate entropy loss. Experiments on three benchmarks with hierarchical granularities show that BigSPN is an effective framework for Hierarchical Granularity Transfer Learning. Shaobo Min, Hongtao Xie 0001, Hantao Yao, Xuran Deng, Zhengjun Zha, Yongdong Zhang 0001 |
NeurIPS | 1 |
| 2020 | Multi-Objective Matrix Normalization for Fine-Grained Visual RecognitionabstractBilinear pooling achieves great success in fine-grained visual recognition (FGVC). Recent methods have shown that the matrix power normalization can stabilize the second-order information in bilinear features, but some problems, e.g., redundant information and over-fitting, remain to be resolved. In this paper, we propose an efficient Multi-Objective Matrix Normalization (MOMN) method that can simultaneously normalize a bilinear representation in terms of square-root, low-rank, and sparsity. These three regularizers can not only stabilize the second-order information, but also compact the bilinear features and promote model generalization. In MOMN, a core challenge is how to jointly optimize three non-smooth regularizers of different convex properties. To this end, MOMN first formulates them into an augmented Lagrange formula with approximated regularizer constraints. Then, auxiliary variables are introduced to relax different constraints, which allow each regularizer to be solved alternately. Finally, several updating strategies based on gradient descent are designed to obtain consistent convergence and efficient implementation. Consequently, MOMN is implemented with only matrix multiplication, which is well-compatible with GPU acceleration, and the normalized bilinear features are stabilized and discriminative. Experiments on five public benchmarks for FGVC demonstrate that the proposed MOMN is superior to existing normalization-based methods in terms of both accuracy and efficiency. The code is available: https://github.com/mboboGO/MOMN. Shaobo Min, Hantao Yao, Hongtao Xie 0001, Zhengjun Zha, Yongdong Zhang 0001 |
IEEE Trans. Image Process. | 1 |
| 2019 | A Two-Stream Mutual Attention Network for Semi-Supervised Biomedical Segmentation with Noisy LabelsabstractLearning-based methods suffer from a deficiency of clean annotations, especially in biomedical segmentation. Although many semi-supervised methods have been proposed to provide extra training data, automatically generated labels are usually too noisy to retrain models effectively. In this paper, we propose a Two-Stream Mutual Attention Network (TSMAN) that weakens the influence of back-propagated gradients caused by incorrect labels, thereby rendering the network robust to unclean data. The proposed TSMAN consists of two sub-networks that are connected by three types of attention models in different layers. The target of each attention model is to indicate potentially incorrect gradients in a certain layer for both sub-networks by analyzing their inferred features using the same input. In order to achieve this purpose, the attention models are designed based on the propagation analysis of noisy gradients at different layers. This allows the attention models to effectively discover incorrect labels and weaken their influence during parameter updating process. By exchanging multi-level features within two-stream architecture, the effects of noisy labels in each sub-network are reduced by decreasing the noisy gradients. Furthermore, a hierarchical distillation is developed to provide reliable pseudo labels for unlabelded data, which further boosts the performance of TSMAN. The experiments using both HVSMR 2016 and BRATS 2015 benchmarks demonstrate that our semi-supervised learning framework surpasses the state-of-the-art fully-supervised results. Shaobo Min, Xuejin Chen, Zhengjun Zha, Feng Wu 0001, Yongdong Zhang 0001 |
AAAI | 1 |
| 2019 | Accurate Segmentation of Synaptic Cleft with Contour Growing Concatenated with a ConvnetabstractSynaptic cleft is an important area for neuroscientists to analyze the macromolecular complexes related to neurotransmitter transmission. However, the large amount of noise and low signal-to-noise ratio in raw electron micrographs make it challenging to extract this region automatically. In this paper, we propose a simple but effective framework to automatically extract accurate boundaries of synaptic cleft regions. Our approach concatenates a novel contour growing algorithm to a fully convolutional network (FCN), so that it takes both advantages of large receptive field of FCNs and fine-level localization of contour evolution. The contour growing algorithm is based on the flexible evolving tension and synchronous growing controlling to localize the opening contour of clef region. With consideration of both global localization and local segmentation, our approach is more robust to noisy electron micrographs and outperforms all existing single-model FCNs on accurate segmentation of synaptic clefts. Shaobo Min, Xuejin Chen, Hongtao Xie 0001, Zhengjun Zha, Guoqiang Bi, Feng Wu 0001, Yongdong Zhang 0001 |
ICIP | 1 |
| 2019 | Structure Generation and Guidance Network for Unsupervised Monocular Depth EstimationabstractStructure information is important to unsupervised depth learning from monocular videos. However, most existing methods focus on depth smoothing on planar regions, while other structure information, such as object shape and surface curvature, is ignored. In this work, we propose SGGN, a novel Structure Generation and Guidance Network to refine depth estimation under the guidance of extracted image structure. We introduce second-order Domain Transform filtering, which explores spatial depth variation by gradient propagation, to capture long-range dependence in the extracted structure for depth refinement. Then, several structure-aware constraints, as well as an attention mechanism, are applied to guide the training of SGGN, which leads to better depth estimation with structural guidance. Notably, our structure-aware constraints are designed in terms of different characteristics. Experiments on three benchmarks demonstrate the effectiveness of our structure-guided model and its state-of-the-art performance for unsupervised depth estimation. Chaoqun Wang 0011, Xuejin Chen, Shaobo Min, Feng Wu 0001 |
ICME | 3 |
| 2019 | Domain-Specific Embedding Network for Zero-Shot RecognitionabstractZero-Shot Learning (ZSL) seeks to recognize a sample from either seen or unseen domain by projecting the image data and semantic labels into a joint embedding space. However, most existing methods directly adapt a well-trained projection from one domain to another, thereby ignoring the serious bias problem caused by domain differences. To address this issue, we propose a novel Domain-Specific Embedding Network (DSEN) that can apply specific projections to different domains for unbiased embedding, as well as several domain constraints. In contrast to previous methods, the DSEN decomposes the domain-shared projection function into one domain-invariant and two domain-specific sub-functions to explore the similarities and differences between two domains. To prevent the two specific projections from breaking the semantic relationship, a semantic reconstruction constraint is proposed by applying the same decoder function to them in a cycle consistency way. Furthermore, a domain division constraint is developed to directly penalize the margin between real and pseudo image features in respective seen and unseen domains, which can enlarge the inter-domain difference of visual features. Extensive experiments on four public benchmarks demonstrate the effectiveness of DSEN with an average of $9.2%$ improvement in terms of harmonic mean. The code is available in \urlhttps://github.com/mboboGO/DSEN-for-GZSL. Shaobo Min, Hantao Yao, Hongtao Xie 0001, Zhengjun Zha, Yongdong Zhang 0001 |
ACM Multimedia | 1 |
| 2019 | Adaptive Bilinear Pooling for Fine-grained Representation LearningabstractFine-grained representation learning targets to generate discriminative description for fine-grained visual objects. Recently, the bilinear feature interaction has been proved effective in generating powerful high-order representation with spatially invariant information. However, the existing methods apply a fixed feature interaction strategy to all samples, which ignore the image and region heterogeneity in a dataset. To this end, we propose a generalized feature interaction method, named Adaptive Bilinear Pooling (ABP), which can adaptively infer a suitable pooling strategy for a given sample based on image content. Specifically, ABP consists of two learning strategies: p-order learning (P-net) and spatial attention learning (S-net). The p-order learning predicts an optimal exponential coefficient rather than a fixed order number to extract moderate visual information from an image. The spatial attention learning aims to infer a weighted score that measures the importance of each local region, which can compact the image representations. To make ABP compatible with kernelized bilinear feature interaction, a crossed two-branch structure is utilized to combine the P-net and S-net. This structure can facilitate complementary information exchange between two different visual branches. The experiments on three widely used benchmarks, including fine-grained object classification and action recognition, demonstrate the effectiveness of the proposed method. Shaobo Min, Hongtao Xie 0001, Youliang Tian, Hantao Yao, Yongdong Zhang 0001 |
MMAsia | 1 |