EDBT 2026 Demo / reviewers in the wild / expert
Hao Tang 0007
dblp:07/5751-7
· DBLP profile ↗
29ranked-venue papers
6as first author
28since 2021 · last 2026
0000-0002-6973-8121ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 18 · 4 first-author · 17 since 2021Artificial intelligence and machine learning · 17 · 4 first-author · 17 since 2021Security and privacy · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Fine-Grained Image Retrieval via Dual-Vision AdaptationabstractFine-Grained Image Retrieval~(FGIR) faces challenges in learning discriminative visual representations to retrieve images with similar fine-grained features. Current leading FGIR solutions typically follow two regimes: enforce pairwise similarity constraints in the semantic embedding space, or incorporate a localization sub-network to fine-tune the entire model. However, such two regimes tend to overfit the training data while forgetting the knowledge gained from large-scale pre-training, thus reducing their generalization ability. In this paper, we propose a Dual-Vision Adaptation (DVA) approach for FGIR, which guides the frozen pre-trained model to perform FGIR through collaborative sample and feature adaptation. Specifically, we design Object-Perceptual Adaptation, which modifies input samples to help the pre-trained model perceive critical objects and elements within objects that are helpful for category prediction. Meanwhile, we propose In-Context Adaptation, which introduces a small set of parameters for feature adaptation without modifying the pre-trained parameters. This makes the FGIR task using these adapted features closer to the task solved during the pre-training. Additionally, to balance retrieval efficiency and performance, we propose Discrimination Perception Transfer to transfer the discriminative knowledge in the object-perceptual adaptation to the image encoder using the knowledge distillation mechanism. Extensive experiments show that DVA performs well on three fine-grained datasets. Xin Jiang 0010, Meiqi Cao, Hao Tang 0007, Fei Shen 0004, Zechao Li |
AAAI | 3 |
| 2026 | Cross-modal Proxy Evolving for OOD Detection with Vision-Language ModelsabstractReliable zero-shot detection of out-of-distribution (OOD) inputs is critical for deploying vision-language models in open-world settings. However, the lack of labeled negatives in zero-shot OOD detection necessitates proxy signals that remain effective under distribution shift. Existing negative-label methods rely on a fixed set of textual proxies, which (i) sparsely sample the semantic space beyond in-distribution (ID) classes and (ii) remain static while only visual features drift, leading to cross-modal misalignment and unstable predictions. In this paper, we propose CoEvo, a training- and annotation-free test-time framework that performs bidirectional, sample-conditioned adaptation of both textual and visual proxies. Specifically, CoEvo introduces a proxy-aligned co-evolution mechanism to maintain two evolving proxy caches, which dynamically mines contextual textual negatives guided by test images and iteratively refines visual proxies, progressively realigning cross-modal similarities and enlarging local OOD margins. Finally, we dynamically re-weight the contributions of dual-modal proxies to obtain a calibrated OOD score that is robust to distribution shift. Extensive experiments on standard benchmarks demonstrate that CoEvo achieves state-of-the-art performance, improving AUROC by 1.33% and reducing FPR95 by 45.98% on ImageNet-1K compared to strong negative-label baselines. Hao Tang 0007, Yu Liu 0158, Shuanglin Yan, Fei Shen 0004, Shengfeng He, Harry Qin |
AAAI | 1 |
| 2026 | IMAGGarment+: Efficient Attribute-Wise Diffusion for Garment GenerationabstractDiffusion models have advanced fine-grained garment generation, yet balancing controllability, efficiency, and texture fidelity remains challenging. Adapter-based methods often yield incoherent details, while full fine-tuning is computationally expensive and prone to overwriting pretrained priors. To address these limitations, we propose IMAGGarment+, an efficient diffusion framework for controllable and high-quality garment synthesis. It comprises two key modules designed for efficient and attribute-aware conditioning. First, we introduce an attribute-wise feature extractor (AFE) that disentangles key garment attributes, silhouette, logo, position, and color, into parallel latent streams. Each stream is optimized independently via LoRA, ensuring minimal parameter overhead while retaining expressive capacity. Second, we develop an attribute-adaptive attention (AA) module to inject attribute-specific cues into the generative process through a selective, layer-wise injection strategy. Specifically, silhouette and color features are injected into early decoder layers to guide structural and appearance formation, while logo features are propagated across all layers to ensure cross-scale consistency. Extensive experiments on fine-grained garment benchmarks demonstrate that IMAGGarment+ outperforms state-of-the-art baselines with less than 20% additional parameters, validating its effectiveness and efficiency. Fei Shen 0004, Cong Wang 0034, Yanpeng Sun, Hao Tang 0007, Xiaoyu Du 0002 |
AAAI | 5 |
| 2026 | FourierPET: Deep Fourier-based Unrolled Network for Low-count PET ReconstructionabstractLow-count positron emission tomography (PET) reconstruction is a challenging inverse problem due to severe degradations arising from Poisson noise, photon scarcity, and attenuation correction errors. Existing deep learning methods typically address these in the spatial domain with an undifferentiated optimization objective, making it difficult to disentangle overlapping artifacts and limiting correction effectiveness. In this work, we perform a Fourier-domain analysis and reveal that these degradations are spectrally separable: Poisson noise and photon scarcity cause high-frequency phase perturbations, while attenuation errors suppress low-frequency amplitude components. Leveraging this insight, we propose FourierPET, a Fourier-based unrolled reconstruction framework grounded in the Alternating Direction Method of Multipliers. It consists of three tailored modules: a spectral consistency module that enforces global frequency alignment to maintain data fidelity, an amplitude–phase correction module that decouples and compensates for high-frequency phase distortions and low-frequency amplitude suppression, and a dual adjustment module that accelerates convergence during iterative reconstruction. Extensive experiments demonstrate that FourierPET achieves state-of-the-art performance with significantly fewer parameters, while offering enhanced interpretability through frequency-aware correction. Hao Tang 0007, Zhanli Hu, Harry Qin |
AAAI | 2 |
| 2026 | GradAlign: Detecting Out-of-Distribution Samples via Gradient Concentration
Jiawei Gu, Yanpeng Sun, Hao Tang 0007, Zechao Li |
Int. J. Comput. Vis. | 3 |
| 2026 | CylindFormer: Image-to-Point Cloud Registration with Cylindrical Transformer
Hao Tang 0007, Yanpeng Sun, Shengfeng He, Zechao Li |
Int. J. Comput. Vis. | 2 |
| 2026 | Trajectory-enhanced transferable attacks for vision-language pre-trained models
Haiqi Zhang 0001, Ziqiang Li 0001, Hao Tang 0007, Zechao Li |
Pattern Recognit. | 3 |
| 2026 | Gradient Pruning Interactive Attack for Vision-Language Pre-Training ModelsabstractVision-Language Pre-training (VLP) models exhibit pronounced vulnerability to multimodal adversarial examples, necessitating rigorous robustness research, particularly for transferable attacks in black-box scenarios. Current research predominantly enhances attack transferability across VLP models by diversifying image and text inputs. However, during adversarial example generation, these methods often prioritize amplifying inter-modal semantic discrepancies (i.e., modality-discrepancy features) while overlooking model-specific semantic features critical to transferable attacks. To address this limitation, we pro pose a transferable Gradient Pruning Interactive Attack (GPI Attack), which integrates gradient-pruned image perturbations with semantic-oriented text perturbations through modality interaction. For image attacks, extreme backpropagated gradients may cause adversarial examples to highlight certain model specific features, leading to poor transferability. To suppress this feature, the textual modality guides the pruning of extreme gradients within intermediate VLP blocks, and these pruned gradients are subsequently employed to direct the generation of adversarial images. For text attacks, we consolidate the perturbation process solely at the embedding level, which reduces semantic discrepancies across hierarchical structures and significantly enhances the generalizability of adversarial texts. Experimental results demonstrate the effectiveness of GPI-Attack in image text retrieval tasks on multimodal datasets such as Flickr30K and MSCOCO. Additionally, the proposed gradient pruning technique is plug-and-play, showing performance improvements even when applied to baseline methods, indicating its potential as a valuable enhancement for attack performance. Haiqi Zhang 0001, Hao Tang 0007, Yanpeng Sun, Zechao Li |
IEEE Trans. Dependable Secur. Comput. | 2 |
| 2026 | Rethinking Vision Transformer for Large-Scale Fine-Grained Image RetrievalabstractLarge-scale fine-grained image retrieval (FGIR) aims to retrieve images belonging to the same subcategory as a given query by capturing subtle differences in a large-scale setting. Recently, Vision Transformers (ViT) have been employed in FGIR due to their powerful self-attention mechanism for modeling long-range dependencies. However, most Transformer-based methods focus primarily on leveraging self-attention to distinguish fine-grained details, while overlooking the high computational complexity and redundant dependencies inherent to these models, limiting their scalability and effectiveness in large-scale FGIR. In this paper, we propose an Efficient and Effective ViT-based framework, termedEET, which integrates token pruning module with a discriminative transfer strategy to address these limitations. Specifically, we introduce a content-based token pruning scheme to enhance the efficiency of the vanilla ViT, progressively removing background or low-discriminative tokens at different stages by exploiting feature responses and self-attention mechanism. To ensure the resulting efficient ViT retains strong discriminative power, we further present a discriminative transfer strategy comprising bothdiscriminative knowledge transferanddiscriminative region guidance. Using a distillation paradigm, these components transfer knowledge from a larger “teacher” ViT to a more efficient “student” model, guiding the latter to focus on subtle yet crucial regions in a cost-free manner. Extensive experiments on two widely-used fine-grained datasets and four large-scale fine-grained datasets demonstrate the effectiveness of our method. Specifically, EET reduces the inference latency of ViT-Small by 42.7% and boosts the retrieval performance of 16-bit hash codes by 5.15% on the challenging NABirds dataset. The code is publicly available at:https://github.com/WhiteJiang/EET. Xin Jiang 0010, Hao Tang 0007, Yonghua Pan, Zechao Li |
IEEE Trans. Multim. | 2 |
| 2025 | Multi-scale Activation, Selection, and Aggregation: Exploring Diverse Cues for Fine-Grained Bird RecognitionabstractGiven the critical role of birds in ecosystems, Fine-Grained Bird Recognition (FGBR) has gained increasing attention, particularly in distinguishing birds within similar subcategories. Although Vision Transformer (ViT)-based methods often outperform Convolutional Neural Network (CNN)-based methods in FGBR, recent studies reveal that the limited receptive field of plain ViT model hinders representational richness and makes them vulnerable to scale variance. Thus, enhancing the multi-scale capabilities of existing ViT-based models to overcome this bottleneck in FGBR is a worthwhile pursuit. In this paper, we propose a novel framework for FGBR, namely Multi-scale Diverse Cues Modeling (MDCM), which explores diverse cues at different scales across various stages of a multi-scale Vision Transformer (MS-ViT) in an ``Activation-Selection-Aggregation'' paradigm. Specifically, we first propose a multi-scale cue activation module to ensure the discriminative cues learned at different stage are mutually different. Subsequently, a multi-scale token selection mechanism is proposed to remove redundant noise and highlight discriminative, scale-specific cues at each stage. Finally, the selected tokens from each stage are independently utilized for bird recognition, and the recognition results from multiple stages are adaptively fused through a multi-scale dynamic aggregation mechanism for final model decisions. Both qualitative and quantitative results demonstrate the effectiveness of our proposed MDCM, which outperforms CNN- and ViT-based models on several widely-used FGBR benchmarks. Hao Tang 0007, Jinhui Tang 0001 |
AAAI | 2 |
| 2025 | OT-DETECTOR: Delving into Optimal Transport for Zero-shot Out-of-Distribution DetectionabstractOut-of-distribution (OOD) detection is crucial for ensuring the reliability and safety of machine learning models in real-world applications. While zero-shot OOD detection, which requires no training on in-distribution (ID) data, has become feasible with the emergence of vision-language models like CLIP, existing methods primarily focus on semantic matching and fail to fully capture distributional discrepancies. To address these limitations, we propose OT-DETECTOR, a novel framework that employs Optimal Transport (OT) to quantify both semantic and distributional discrepancies between test samples and ID labels. Specifically, we introduce cross-modal transport mass and transport cost as semantic-wise and distribution-wise OOD scores, respectively, enabling more robust detection of OOD samples. Additionally, we present a semantic-aware content refinement (SaCR) module, which utilizes semantic cues from ID labels to amplify the distributional discrepancy between ID and hard OOD samples. Extensive experiments on several benchmarks demonstrate that OT-DETECTOR achieves state-of-the-art performance across various OOD detection tasks, particularly in challenging hard-OOD scenarios. Yu Liu 0158, Hao Tang 0007, Haiqi Zhang 0001, Harry Qin, Zechao Li |
IJCAI | 2 |
| 2025 | Connecting Giants: Synergistic Knowledge Transfer of Large Multimodal Models for Few-Shot LearningabstractFew-shot learning (FSL) addresses the challenge of classifying novel classes with limited training samples. While some methods leverage semantic knowledge from smaller-scale models to mitigate data scarcity, these approaches often introduce noise and bias due to the data’s inherent simplicity. In this paper, we propose a novel framework, Synergistic Knowledge Transfer (SynTrans), which effectively transfers diverse and complementary knowledge from large multimodal models to empower the off-the-shelf few-shot learner. Specifically, SynTrans employs CLIP as a robust teacher and uses a few-shot vision encoder as a weak student, distilling semantic-aligned visual knowledge via an unsupervised proxy task. Subsequently, a training-free synergistic knowledge mining module facilitates collaboration among large multimodal models to extract high-quality semantic knowledge. Building upon this, a visual-semantic bridging module enables bi-directional knowledge transfer between visual and semantic spaces, transforming explicit visual and implicit semantic knowledge into category-specific classifier weights. Finally, SynTrans introduces a visual weight generator and a semantic weight reconstructor to adaptively construct optimal multimodal FSL classifiers. Experimental results on four FSL datasets demonstrate that SynTrans, even when paired with a simple few-shot vision encoder, significantly outperforms current state-of-the-art methods. Hao Tang 0007, Shengfeng He, Harry Qin |
IJCAI | 1 |
| 2025 | Divide-and-Conquer: Confluent Triple-Flow Network for RGB-T Salient Object DetectionabstractRGB-Thermal Salient Object Detection (RGB-T SOD) aims to pinpoint prominent objects within aligned pairs of visible and thermal infrared images. A key challenge lies in bridging the inherent disparities between RGB and Thermal modalities for effective saliency map prediction. Traditional encoder-decoder architectures, while designed for cross-modality feature interactions, may not have adequately considered the robustness against noise originating from defective modalities, thereby leading to suboptimal performance in complex scenarios. Inspired by hierarchical human visual systems, we propose the ConTriNet, a robust Confluent Triple-Flow Network employing a "Divide-and-Conquer" strategy. This framework utilizes a unified encoder with specialized decoders, each addressing different subtasks of exploring modality-specific and modality-complementary information for RGB-T SOD, thereby enhancing the final saliency map prediction. Specifically, ConTriNet comprises three flows: two modality-specific flows explore cues from RGB and Thermal modalities, and a third modality-complementary flow integrates cues from both modalities. ConTriNet presents several notable advantages. It incorporates a Modality-induced Feature Modulator (MFM) in the modality-shared union encoder to minimize inter-modality discrepancies and mitigate the impact of defective samples. Additionally, a foundational Residual Atrous Spatial Pyramid Module (RASPM) in the separated flows enlarges the receptive field, allowing for the capture of multi-scale contextual information. Furthermore, a Modality-aware Dynamic Aggregation Module (MDAM) in the modality-complementary flow dynamically aggregates saliency-related cues from both modality-specific flows. Leveraging the proposed parallel triple-flow framework, we further refine saliency maps derived from different flows through a flow-cooperative fusion strategy, yielding a high-quality, full-resolution saliency map for the final prediction. To evaluate the robustness and stability of our approach, we collect a comprehensive RGB-T SOD benchmark, VT-IMAG, covering various real-world challenging scenarios. Extensive experiments on public benchmarks and our VT-IMAG dataset demonstrate that ConTriNet consistently outperforms state-of-the-art competitors in both common and challenging scenarios, even when dealing with incomplete modality data. The code and VT-IMAG will be available at: https://cser-tang-hao.github.io/contrinet.html. Hao Tang 0007, Zechao Li, Shengfeng He, Jinhui Tang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2025 | Modality-Specific Interactive Attack for Vision-Language Pre-Training ModelsabstractRecent advances have heightened the interest in the adversarial transferability of Vision-Language Pre-training (VLP) models. However, most existing strategies constrained by two persistent limitations: suboptimal utilization of cross-modal interactive information, and inherent discrepancies across hierarchical textual representation. To address these challenges, we propose the Modality-Specific Interactive Attack (MSI-Attack), a novel approach that integrates semantic-level image perturbations with embedding-level text perturbations, all while maintaining minimal inter-modal constraints. In our image attack methodology, we introduce Multi-modal Integrated Gradients (MIG) to guide perturbations toward the core semantics of images, enriched by their associated deeply text information. This technique enhances transferability by capturing consistent features across various models, thereby effectively misleading similar-model perception areas. Additionally, we employ a momentum iteration strategy in conjunction with MIG, which amalgamates current and historical gradients to expedite the perturbation updates. For text attacks, we streamline the perturbation process by operating exclusively at the embedding level. This reduces semantic gaps across hierarchical structures and significantly enhances the generalizability of adversarial text. Moreover, we delve deeper into how semantic perturbations with varying degrees of similarity affect the overall attack effectiveness. Our experimental results on image-text retrieval tasks using the multi-modal datasets Flickr30K and MSCOCO underscore the efficacy of MSI-Attack. Our method achieves superior performance, setting a new state-of-the-art benchmark, all without the need for additional mechanisms. Haiqi Zhang 0001, Hao Tang 0007, Yanpeng Sun, Shengfeng He, Zechao Li |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2025 | Knowledge-Guided Semantic Transfer Network for Few-Shot Image RecognitionabstractDeep learning-based models have been shown to outperform human beings in many computer vision tasks with massive available labeled training data in learning. However, humans have an amazing ability to easily recognize images of novel categories by browsing only a few examples of these categories. In this case, few-shot learning comes into being to make machines learn from extremely limited labeled examples. One possible reason why human beings can well learn novel concepts quickly and efficiently is that they have sufficient visual and semantic prior knowledge. Toward this end, this work proposes a novel knowledge-guided semantic transfer network (KSTNet) for few-shot image recognition from a supplementary perspective by introducing auxiliary prior knowledge. The proposed network jointly incorporates vision inferring, knowledge transferring, and classifier learning into one unified framework for optimal compatibility. A category-guided visual learning module is developed in which a visual classifier is learned based on the feature extractor along with the cosine similarity and contrastive loss optimization. To fully explore prior knowledge of category correlations, a knowledge transfer network is then developed to propagate knowledge information among all categories to learn the semantic-visual mapping, thus inferring a knowledge-based classifier for novel categories from base categories. Finally, we design an adaptive fusion scheme to infer the desired classifiers by effectively integrating the above knowledge and visual information. Extensive experiments are conducted on two widely used Mini-ImageNet and Tiered-ImageNet benchmarks to validate the effectiveness of KSTNet. Compared with the state of the art, the results show that the proposed method achieves favorable performance with minimal bells and whistles, especially in the case of one-shot learning. Zechao Li, Hao Tang 0007, Zhimao Peng, Guo-Jun Qi, Jinhui Tang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2025 | ADPS: Asymmetric Distillation Postsegmentation for Image Anomaly DetectionabstractKnowledge distillation-based anomaly detection (KDAD) methods rely on the teacher-student paradigm to detect and segment anomalous regions by contrasting the unique features extracted by both networks. However, existing KDAD methods suffer from two main limitations: 1) the student network can effortlessly replicate the teacher network's representations and 2) the features of the teacher network serve solely as a "reference standard" and are not fully leveraged. Toward this end, we depart from the established paradigm and instead propose an innovative approach called asymmetric distillation postsegmentation (ADPS). Our ADPS employs an asymmetric distillation paradigm that takes distinct forms of the same image as the input of the teacher-student networks, driving the student network to learn discriminating representations for anomalous regions. Meanwhile, a customized Weight Mask Block (WMB) is proposed to generate a coarse anomaly localization mask that transfers the distilled knowledge acquired from the asymmetric paradigm to the teacher network. Equipped with WMB, the proposed postsegmentation module (PSM) can effectively detect and segment abnormal regions with fine structures and clear boundaries. Experimental results demonstrate that the proposed ADPS outperforms the state-of-the-art methods in detecting and segmenting anomalies. Surprisingly, ADPS significantly improves average precision (AP) metric by $\mathbf {9}\%$ and $\mathbf {20}\%$ on the MVTec anomaly detection (AD) and KolektorSDD2 datasets, respectively. Peng Xing, Hao Tang 0007, Jinhui Tang 0001, Zechao Li |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2024 | Delving into Multimodal Prompting for Fine-Grained Visual ClassificationabstractFine-grained visual classification (FGVC) involves categorizing fine subdivisions within a broader category, which poses challenges due to subtle inter-class discrepancies and large intra-class variations. However, prevailing approaches primarily focus on uni-modal visual concepts. Recent advancements in pre-trained vision-language models have demonstrated remarkable performance in various high-level vision tasks, yet the applicability of such models to FGVC tasks remains uncertain. In this paper, we aim to fully exploit the capabilities of cross-modal description to tackle FGVC tasks and propose a novel multimodal prompting solution, denoted as MP-FGVC, based on the contrastive language-image pertaining (CLIP) model. Our MP-FGVC comprises a multimodal prompts scheme and a multimodal adaptation scheme. The former includes Subcategory-specific Vision Prompt (SsVP) and Discrepancy-aware Text Prompt (DaTP), which explicitly highlights the subcategory-specific discrepancies from the perspectives of both vision and language. The latter aligns the vision and text prompting elements in a common semantic space, facilitating cross-modal collaborative reasoning through a Vision-Language Fusion Module (VLFM) for further improvement on FGVC. Moreover, we tailor a two-stage optimization strategy for MP-FGVC to fully leverage the pre-trained CLIP model and expedite efficient adaptation for FGVC. Extensive experiments conducted on four FGVC datasets demonstrate the effectiveness of our MP-FGVC. Xin Jiang 0010, Hao Tang 0007, Junyao Gao 0002, Xiaoyu Du 0002, Shengfeng He, Zechao Li |
AAAI | 2 |
| 2024 | Learning with Unreliability: Fast Few-Shot Voxel Radiance Fields with Relative Geometric ConsistencyabstractWe propose a voxel-based optimization framework, Re VoRF, for few-shot radiance fields that strategically ad-dress the unreliability in pseudo novel view synthesis. Our method pivots on the insight that relative depth relationships within neighboring regions are more reliable than the ab-solute color values in disoccluded areas. Consequently, we devise a bilateral geometric consistency loss that carefully navigates the trade-off between color fidelity and geometric accuracy in the context of depth consistency for uncertain regions. Moreover, we present a reliability-guided learning strategy to discern and utilize the variable quality across syn-thesized views, complemented by a reliability-aware voxel smoothing algorithm that smoothens the transition between reliable and unreliable data patches. Our approach allows for a more nuanced use of all available data, promoting en-hanced learning from regions previously considered unsuit-able for high-quality reconstruction. Extensive experiments across diverse datasets reveal that our approach attains significant gains in efficiency and accuracy, delivering ren-dering speeds of 3 FPS, 7 mins to train a 360° scene, and a 5% improvement in PSNR over existing few-shot methods. Code is available at https://github.com/HKCLynn/ReVoRF. Bangzhen Liu, Hao Tang 0007, Bailin Deng, Shengfeng He |
CVPR | 3 |
| 2024 | DVF: Advancing Robust and Accurate Fine-Grained Image Retrieval with Retrieval GuidelinesabstractFine-grained image retrieval (FGIR) is to learn visual representations that distinguish visually similar objects while maintaining generalization. Existing methods propose to generate discriminative features, but rarely consider the particularity of the FGIR task itself. This paper presents a meticulous analysis leading to the proposal of practical guidelines to identify subcategory-specific discrepancies and generate discriminative features to design effective FGIR models. These guidelines include emphasizing the object (G1), highlighting subcategory-specific discrepancies (G2), and employing effective training strategy (G3). Following G1 and G2, we design a novel Dual Visual Filtering mechanism for the plain visual transformer, denoted as DVF, to capture subcategory-specific discrepancies. Specifically, the dual visual filtering mechanism comprises an object-oriented module and a semantic-oriented module. These components serve to magnify objects and identify discriminative regions, respectively. Following G3, we implement a discriminative model training strategy to improve the discriminability and generalization ability of DVF. Extensive analysis and ablation studies confirm the efficacy of our proposed guidelines. Without bells and whistles, the proposed DVF achieves state-of-the-art performance on three widely-used fine-grained datasets in closed-set and open-set settings. Xin Jiang 0010, Hao Tang 0007, Rui Yan 0010, Jinhui Tang 0001, Zechao Li |
ACM Multimedia | 2 |
| 2024 | Erasing, Transforming, and Noising Defense Network for Occluded Person Re-IdentificationabstractOcclusion perturbation presents a significant challenge in person re-identification (re-ID), and existing methods that rely on external visual cues require additional computational resources and only consider the issue of missing information caused by occlusion. In this paper, we propose a simple yet effective framework, termed Erasing, Transforming, and Noising Defense Network (ETNDNet), which treats occlusion as a noise disturbance and solves occluded person re-ID from the perspective of adversarial defense. In the proposed ETNDNet, we introduce three strategies: Firstly, we randomly erase the feature map to create an adversarial representation with incomplete information, enabling adversarial learning of identity loss to protect the re-ID system from the disturbance of missing information. Secondly, we introduce random transformations to simulate the position misalignment caused by occlusion, training the extractor and classifier adversarially to learn robust representations immune to misaligned information. Thirdly, we perturb the feature map with random values to address noisy information introduced by obstacles and non-target pedestrians, and employ adversarial gaming in the re-ID system to enhance its resistance to occlusion noise. Without bells and whistles, ETNDNet has three key highlights: (i) it does not require any external modules with parameters, (ii) it effectively handles various issues caused by occlusion from obstacles and non-target pedestrians, and (iii) it designs the first GAN-based adversarial defense paradigm for occluded person re-ID. Extensive experiments on six public datasets fully demonstrate the effectiveness, superiority, and practicality of the proposed ETNDNet. The code will be released at https://github.com/nengdong96/ETNDNet. Neng Dong, Liyan Zhang 0001, Shuanglin Yan, Hao Tang 0007, Jinhui Tang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | Learning Contrastive Self-Distillation for Ultra-Fine-Grained Visual Categorization Targeting Limited SamplesabstractIn the field of intelligent multimedia analysis, ultra-fine-grained visual categorization (Ultra-FGVC) plays a vital role in distinguishing intricate subcategories within broader categories. However, this task is inherently challenging due to the complex granularity of category subdivisions and the limited availability of data for each category. To address these challenges, this work proposes CSDNet, a pioneering framework that effectively explores contrastive learning and self-distillation to learn discriminative representations specifically designed for Ultra-FGVC tasks. CSDNet comprises three main modules: Subcategory-Specific Discrepancy Parsing (SSDP), Dynamic Discrepancy Learning (DDL), and Subcategory-Specific Discrepancy Transfer (SSDT), which collectively enhance the generalization of deep models across instance, feature, and logit prediction levels. To increase the diversity of training samples, the SSDP module introduces adaptive augmented samples to spotlight subcategory-specific discrepancies. Simultaneously, the proposed DDL module stores historical intermediate features by a dynamic memory queue, which optimizes the feature learning space through iterative contrastive learning. Furthermore, the SSDT module effectively distills subcategory-specific discrepancies knowledge from the inherent structure of limited training data using a self-distillation paradigm at the logit prediction level. Experimental results demonstrate that CSDNet outperforms current state-of-the-art Ultra-FGVC methods, emphasizing its powerful efficacy and adaptability in addressing Ultra-FGVC tasks. Ziye Fang, Xin Jiang 0010, Hao Tang 0007, Zechao Li |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | Global Meets Local: Dual Activation Hashing Network for Large-Scale Fine-Grained Image RetrievalabstractIn the Internet era, the exponential growth of fine-grained image databases poses a considerable challenge for efficient information retrieval. Hashing-based approaches gained traction for their computational and storage efficiency, yet fine-grained hashing retrieval presents unique challenges due to small inter-class and large intra-class variations inherent to fine-grained entities. Thus, traditional hashing algorithms falter in discerning these subtle, yet critical, visual differences and fail to generate compact yet semantically rich hash codes. To address this, we introduce a Dual Activation Hashing Network (DAHNet) designed to convert high-dimensional image data into optimized binary codes via an innovative feature activation paradigm. The architecture consists of dual branches specifically tailored for global and local semantic activation, thereby establishing direct correspondences between hash codes and distinguishable object parts through a hierarchical activation pipeline. Specifically, our spatial-oriented semantic activation module modulates dominant visual regions while amplifying the activations of subtle yet semantically rich areas in a controlled manner. Building on these activated visual representations, the proposed inter-region semantic enrichment module further enriches them by unearthing semantically complementary cues. Concurrently,DAHNetintegrates a channel-oriented semantic activation module that exploits channel-specific correlations to distill contextual cues from spatially-activated visual features, thereby reinforcing robust learning to hash. To maintain the similarity of the original entities, we amalgamate final hash codes from both activation branches, capturing both local textural details and global structural information. Comprehensive evaluations on five fine-grained image retrieval benchmarks demonstrateDAHNet's superior performance over existing state-of-the-art hashing solutions, especially on 12-bit, improving performance by 4%-15% compared to the current best results on the five benchmarks. Moreover, generalization studies validate the efficacy of our dual-activation framework in the domain of content-based fine-grained image retrieval. The code is publicly available at:https://github.com/WhiteJiang/DAHNet. Xin Jiang 0010, Hao Tang 0007, Zechao Li |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2024 | Image-Specific Information Suppression and Implicit Local Alignment for Text-Based Person SearchabstractText-based person search (TBPS) is a challenging task that aims to search pedestrian images with the same identity from an image gallery given a query text. In recent years, TBPS has made remarkable progress, and state-of-the-art (SOTA) methods achieve superior performance by learning local fine-grained correspondence between images and texts. However, most existing methods rely on explicitly generated local parts to model fine-grained correspondence between modalities, which is unreliable due to the lack of contextual information or the potential introduction of noise. Moreover, the existing methods seldom consider the information inequality problem between modalities caused by image-specific information. To address these limitations, we propose an efficient joint multilevel alignment network (MANet) for TBPS, which can learn aligned image/text feature representations between modalities at multiple levels, and realize fast and effective person search. Specifically, we first design an image-specific information suppression (ISS) module, which suppresses image background and environmental factors by relation-guided localization (RGL) and channel attention filtration (CAF), respectively. This module effectively alleviates the information inequality problem and realizes the alignment of information volume between images and texts. Second, we propose an implicit local alignment (ILA) module to adaptively aggregate all pixel/word features of image/text to a set of modality-shared semantic topic centers and implicitly learn the local fine-grained correspondence between modalities without additional supervision and cross-modal interactions. Also, a global alignment (GA) is introduced as a supplement to the local perspective. The cooperation of global and local alignment modules enables better semantic alignment between modalities. Extensive experiments on multiple databases demonstrate the effectiveness and superiority of our MANet. Shuanglin Yan, Hao Tang 0007, Liyan Zhang 0001, Jinhui Tang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2023 | M3Net: Multi-view Encoding, Matching, and Fusion for Few-shot Fine-grained Action RecognitionabstractDue to the scarcity of manually annotated data required for fine-grained video understanding, few-shot fine-grained (FS-FG) action recognition has gained significant attention, with the aim of classifying novel fine-grained action categories with only a few labeled instances. Despite the progress made in FS coarse-grained action recognition, current approaches encounter two challenges when dealing with the fine-grained action categories: the inability to capture subtle action details and the insufficiency of learning from limited data that exhibit high intra-class variance and inter-class similarity. To address these limitations, we propose M3Net, a matching-based framework for FS-FG action recognition, which incorporates multi-view encoding, multi-view matching, and multi-view fusion to facilitate embedding encoding, similarity matching, and decision making across multiple viewpoints.Multi-view encoding captures rich contextual details from the intra-frame, intra-video, and intra-episode perspectives, generating customized higher-order embeddings for fine-grained data.Multi-view matching integrates various matching functions enabling flexible relation modeling within limited samples to handle multi-scale spatio-temporal variations by leveraging the instance-specific, category-specific, and task-specific perspectives. Multi-view fusion consists of matching-predictions fusion and matching-losses fusion over the above views, where the former promotes mutual complementarity and the latter enhances embedding generalizability by employing multi-task collaborative learning. Explainable visualizations and experimental results on three challenging benchmarks demonstrate the superiority of M3Net in capturing fine-grained action details and achieving state-of-the-art performance for FS-FG action recognition. Hao Tang 0007, Jun Liu 0036, Shuanglin Yan, Rui Yan 0010, Zechao Li, Jinhui Tang 0001 |
ACM Multimedia | 1 |
| 2023 | Boosting Few-Shot Fine-Grained Recognition With Background Suppression and Foreground AlignmentabstractFew-shot fine-grained recognition (FS-FGR) aims to recognize novel fine-grained categories with the help of limited available samples. Undoubtedly, this task inherits the main challenges from both few-shot learning and fine-grained recognition. First, the lack of labeled samples makes the learned model easy to overfit. Second, it also suffers from high intra-class variance and low inter-class differences in the datasets. To address this challenging task, we propose a two-stage background suppression and foreground alignment framework, which is composed of a background activation suppression (BAS) module, a foreground object alignment (FOA) module, and a local-to-local (L2L) similarity metric. Specifically, the BAS is introduced to generate a foreground mask for localization to weaken background disturbance and enhance dominative foreground objects. The FOA then reconstructs the feature map of each support sample according to its correction to the query ones, which addresses the problem of misalignment between support-query image pairs. To enable the proposed method to have the ability to capture subtle differences in confused samples, we present a novel L2L similarity metric to further measure the local similarity between a pair of aligned spatial features in the embedding space. What’s more, considering that background interference brings poor robustness, we infer the pairwise similarity of feature maps using both the raw image and the refined image. Extensive experiments conducted on multiple popular fine-grained benchmarks demonstrate that our method outperforms the existing state of the art by a large margin. The source codes are available at:https://github.com/CSer-Tang-hao/BSFA-FSFG. Zican Zha, Hao Tang 0007, Yunlian Sun, Jinhui Tang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2022 | Learning attention-guided pyramidal features for few-shot fine-grained recognition
Hao Tang 0007, Chengcheng Yuan, Zechao Li, Jinhui Tang 0001 |
Pattern Recognit. | 1 |
| 2021 | Coupled Patch Similarity Network FOR One-Shot Fine-Grained Image RecognitionabstractOne-shot fine-grained image recognition (OSFG) aims to distinguish different fine-grained categories with only one training sample per category. Previous works mainly focus on learning a global feature representation through only a using single similarity metric branch, which is unsuitable for OSFG to effectively capture subtle and local differences under limited supervision. In this work, we propose a Coupled Patch Similarity Network (CPSN) for OSFG. Firstly, we propose a Feature Enhancement Module (FEM) to extract more discriminative features of the fine-grained samples. Then, we develop two coupled and symmetrical branches to capture discriminative parts of the samples and reduce the deviation of the distance metric. For each branch, we design a Patch Similarity Module (PSM) to calculate the patch similarity map for the sample pair. Especially, a Patch Weight Generator (PWG) is proposed to generate the patch weight map, which indicates the degree of importance for each position in the patch similarity map, so that the model can focus on diverse and informative parts. We analyze the effect of the different components in the proposed network, and extensive experimental results demonstrate the effectiveness and superiority of the proposed method on two fine-grained benchmark datasets. Hao Tang 0007, Longquan Dai |
ICIP | 2 |
| 2021 | Learning a Tree-Structured Channel-Wise Refinement Network for Efficient Image DerainingabstractSignificant advances have been made in image deraining due to the use of kinds of deep neural networks. However, existing deep neural network-based methods usually contain significant abundant network parameters and thus leads to expensive computation cost, which limits the application of deraining technology in high-level vision tasks. In this paper, we propose a compact and flexible Tree-structured Channel-wise Refinement Block (TCRB) for efficient image deraining, which contains augmentation, refinement, and aggregation modules to better explore features. Specifically, the refinement module can progressively extract groups of more discriminative features from the channel augmented inputs, and then the aggregation module adaptively fuses features from the refinement module to preserve image details by leveraging the Enhanced Channel Attention (ECA) method. Moreover, we present a Tree-structured Channel-wise Refinement Network (TCRN) by stacking multiple TCRBs, which could achieve competitive performance as the complicated networks. We embed the TCRB into a Multi-scale Tree-structured Channel-wise Refinement Network (MTCRN) based on an encoder and decoder network architecture and show that it performs favorably against state-of-the-art deraining algorithms on both synthetic datasets and real-world rainy images, while reaching a better trade-off in terms of model parameters and inference time. Di Wang 0018, Hao Tang 0007, Jinshan Pan, Jinhui Tang 0001 |
ICME | 2 |
| 2020 | BlockMix: Meta Regularization and Self-Calibrated Inference for Metric-Based Meta-LearningabstractMost metric-based meta-learning methods learn only the sophisticated similarity metric for few-shot classification, which may lead to the feature deterioration and unreliable prediction. Toward this end, we propose new mechanisms to learn generalized and discriminative feature embeddings as well as improve the robustness of classifiers against prediction corruptions for meta-learning. For this purpose, a new generation operator BlockMix is proposed by integrating interpolation on the images and labels within metric learning. Based on the above BlockMix, we propose a novel regularization method Meta Regularization as an auxiliary task branch with its own classifier to better constraint the feature embedding module and stabilize the meta-learning process. Furthermore, a novel inference scheme Self-Calibrated Inference is proposed to alleviate the unreliable prediction problem by calibrating the prototype of each category with the confidence-weighted average of the support and generated samples. The proposed mechanisms can be used as supplementary techniques alongside standard metric-based meta-learning algorithms without any pre-training. Experimental results demonstrate the insights and the efficiency of the proposed mechanisms respectively, compared with the state-of-the-art methods on the prevalent few-shot benchmarks. Hao Tang 0007, Zechao Li, Zhimao Peng, Jinhui Tang 0001 |
ACM Multimedia | 1 |