EDBT 2026 Demo / reviewers in the wild / expert
Xin Jiang 0010
dblp:42/4142-10
· DBLP profile ↗
13ranked-venue papers
6as first author
13since 2021 · last 2026
0000-0002-4522-0096ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 11 · 5 first-author · 11 since 2021Artificial intelligence and machine learning · 6 · 2 first-author · 6 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Fine-Grained Image Retrieval via Dual-Vision AdaptationabstractFine-Grained Image Retrieval~(FGIR) faces challenges in learning discriminative visual representations to retrieve images with similar fine-grained features. Current leading FGIR solutions typically follow two regimes: enforce pairwise similarity constraints in the semantic embedding space, or incorporate a localization sub-network to fine-tune the entire model. However, such two regimes tend to overfit the training data while forgetting the knowledge gained from large-scale pre-training, thus reducing their generalization ability. In this paper, we propose a Dual-Vision Adaptation (DVA) approach for FGIR, which guides the frozen pre-trained model to perform FGIR through collaborative sample and feature adaptation. Specifically, we design Object-Perceptual Adaptation, which modifies input samples to help the pre-trained model perceive critical objects and elements within objects that are helpful for category prediction. Meanwhile, we propose In-Context Adaptation, which introduces a small set of parameters for feature adaptation without modifying the pre-trained parameters. This makes the FGIR task using these adapted features closer to the task solved during the pre-training. Additionally, to balance retrieval efficiency and performance, we propose Discrimination Perception Transfer to transfer the discriminative knowledge in the object-perceptual adaptation to the image encoder using the knowledge distillation mechanism. Extensive experiments show that DVA performs well on three fine-grained datasets. Xin Jiang 0010, Meiqi Cao, Hao Tang 0007, Fei Shen 0004, Zechao Li |
AAAI | 1 |
| 2026 | Progressive Feature Encoding With Background Perturbation Learning for Ultra-Fine-Grained Visual CategorizationabstractUltra-Fine-Grained Visual Categorization (Ultra-FGVC) aims to classify objects into sub-granular categories, presenting the challenge of distinguishing visually similar objects with limited data. Existing methods primarily address sample scarcity but often overlook the importance of leveraging intrinsic object features to construct highly discriminative representations. This limitation significantly constrains their effectiveness in Ultra-FGVC tasks. To address these challenges, we propose SV-Transformer that progressively encodes object features while incorporating background perturbation modeling to generate robust and discriminative representations. At the core of our approach is a progressive feature encoder, which hierarchically extracts global semantic structures and local discriminative details from backbone-generated representations. This design enhances inter-class separability while ensuring resilience to intra-class variations. Furthermore, our background perturbation learning mechanism introduces controlled variations in the feature space, effectively mitigating the impact of sample limitations and improving the model's capacity to capture fine-grained distinctions. Comprehensive experiments demonstrate that SV-Transformer achieves state-of-the-art performance on benchmark Ultra-FGVC datasets, showcasing its efficacy in addressing the challenges of Ultra-FGVC task. Xin Jiang 0010, Ziye Fang, Fei Shen 0004, Junyao Gao 0002, Zechao Li |
IEEE Trans. Image Process. | 1 |
| 2026 | Rethinking Vision Transformer for Large-Scale Fine-Grained Image RetrievalabstractLarge-scale fine-grained image retrieval (FGIR) aims to retrieve images belonging to the same subcategory as a given query by capturing subtle differences in a large-scale setting. Recently, Vision Transformers (ViT) have been employed in FGIR due to their powerful self-attention mechanism for modeling long-range dependencies. However, most Transformer-based methods focus primarily on leveraging self-attention to distinguish fine-grained details, while overlooking the high computational complexity and redundant dependencies inherent to these models, limiting their scalability and effectiveness in large-scale FGIR. In this paper, we propose an Efficient and Effective ViT-based framework, termedEET, which integrates token pruning module with a discriminative transfer strategy to address these limitations. Specifically, we introduce a content-based token pruning scheme to enhance the efficiency of the vanilla ViT, progressively removing background or low-discriminative tokens at different stages by exploiting feature responses and self-attention mechanism. To ensure the resulting efficient ViT retains strong discriminative power, we further present a discriminative transfer strategy comprising bothdiscriminative knowledge transferanddiscriminative region guidance. Using a distillation paradigm, these components transfer knowledge from a larger “teacher” ViT to a more efficient “student” model, guiding the latter to focus on subtle yet crucial regions in a cost-free manner. Extensive experiments on two widely-used fine-grained datasets and four large-scale fine-grained datasets demonstrate the effectiveness of our method. Specifically, EET reduces the inference latency of ViT-Small by 42.7% and boosts the retrieval performance of 16-bit hash codes by 5.15% on the challenging NABirds dataset. The code is publicly available at:https://github.com/WhiteJiang/EET. Xin Jiang 0010, Hao Tang 0007, Yonghua Pan, Zechao Li |
IEEE Trans. Multim. | 1 |
| 2026 | IMAGGarment: Fine-Grained Garment Generation for Controllable Fashion DesignabstractThis paper presents IMAGGarment, a fine-grained garment generation (FGG) framework that enables high-fidelity garment synthesis with precise control over silhouette, color, and logo placement. Unlike existing methods that are limited to single-condition inputs, IMAGGarment addresses the challenges of multi-conditional controllability in personalized fashion design and digital apparel applications. Specifically, IMAGGarment employs a two-stage training strategy to separately model global appearance and local details, while enabling unified and controllable generation through end-to-end inference. In the first stage, we propose a global appearance model that jointly encodes silhouette and color using a mixed attention module and a color adapter. In the second stage, we present a local enhancement model with an adaptive appearance-aware module to inject user-defined logos and spatial constraints, enabling accurate placement and visual consistency. To support this task, we release GarmentBench, a large-scale dataset comprising over 180 K garment samples paired with multi-level design conditions, including sketches, color references, logo placements, and textual prompts. Extensive experiments demonstrate that our method outperforms existing baselines, achieving superior structural stability, color fidelity, and local controllability performance. Fei Shen 0004, Cong Wang 0034, Xin Jiang 0010, Xiaoyu Du 0002, Jinhui Tang 0001 |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2025 | IMAGDressing-v1: Customizable Virtual DressingabstractExisting virtual try-on (VTON) methods provide only limited user control over garment attributes and generally overlook essential factors such as face, pose, and scene context. To address these limitations, we introduce the virtual dressing (VD) task, which aims to synthesize freely editable human images conditioned on fixed garments and optional user-defined inputs. We further propose a comprehensive affinity metric index (CAMI) to quantify the consistency between generated outputs and reference garments. We present IMAGDressing-v1, which leverages a garment-specific U-Net to integrate semantic features from CLIP and texture features from a VAE. To incorporate these garment features into a frozen denoising U-Net for flexible text-driven scene control, we employ a hybrid attention mechanism composed of frozen self-attention and trainable cross-attention layers. IMAGDressing-v1 seamlessly integrates with extension modules, such as ControlNet and IP-Adapter, enabling enhanced diversity and controllability. To alleviate data constraints, we introduce the Interactive Garment Pairing (IGPair) dataset, comprising over 300,000 garment–image pairs and a standardized data assembly pipeline. Extensive experiments demonstrate that IMAGDressing-v1 achieves state-of-the-art performance in controlled human image synthesis. The code and model will be available at https://github.com/muzishen/IMAGDressing. Fei Shen 0004, Xin Jiang 0010, Hu Ye, Cong Wang 0034, Xiaoyu Du 0002, Zechao Li, Jinhui Tang 0001 |
AAAI | 2 |
| 2025 | Exploiting Frequency Dynamics for Enhanced Multimodal Event-Based Action Recognition
Meiqi Cao, Xiangbo Shu, Xin Jiang 0010, Rui Yan 0010, Yazhou Yao, Jinhui Tang 0001 |
ICCV | 3 |
| 2025 | A Unified Interpretation of Training-Time Out-Of-Distribution Detection
Xin Jiang 0010, Zechao Li |
ICCV | 2 |
| 2025 | FaceShot: Bring Any Character into LifeabstractIn this paper, we present ***FaceShot***, a novel training-free portrait animation framework designed to bring any character into life from any driven video without fine-tuning or retraining.
We achieve this by offering precise and robust reposed landmark sequences from an appearance-guided landmark matching module and a coordinate-based landmark retargeting module.
Together, these components harness the robust semantic correspondences of latent diffusion models to produce facial motion sequence across a wide range of character types.
After that, we input the landmark sequences into a pre-trained landmark-driven animation model to generate animated video.
With this powerful generalization capability, FaceShot can significantly extend the application of portrait animation by breaking the limitation of realistic portrait landmark detection for any stylized character and driven video.
Also, FaceShot is compatible with any landmark-driven animation model, significantly improving overall performance.
Extensive experiments on our newly constructed character benchmark CharacBench confirm that FaceShot consistently surpasses state-of-the-art (SOTA) approaches across any character domain.
More results are available at our project website https://faceshot2024.github.io/faceshot/. Junyao Gao 0002, Yanan Sun 0005, Fei Shen 0004, Xin Jiang 0010, Zhening Xing, Kai Chen 0026, Cairong Zhao |
ICLR | 4 |
| 2025 | ReDet: Effective Real-time Object Detection via Efficient Multi-scale Extraction AggregationabstractReal-time object detection demands detectors that excel in both speed and accuracy. However, existing methods rely on complex Feature Pyramid Networks and computationally intensive post-processing to boost performance, often struggling to balance efficiency and accuracy. In this paper, we propose ReDet, an efficient real-time end-to-end object detection framework that improves detection capability while maintaining low computational cost. Specifically, we propose a Multi-scale Extraction Aggregation module to enhance feature fusion ability with minimal overhead, boosting feature representation across scales. Additionally, a Regression Enhancement Module is incorporated to mitigate the assignment inconsistency between localization and classification, further enhancing detection accuracy with negligible computational cost. Moreover, we introduce a Multi-label Auxiliary Strategy to eliminate the reliance on post-processing by enabling one-to-one label assignment. The experimental results demonstrate that ReDet-T achieves 39.2% AP and 500 FPS on the COCO val2017 dataset, while ReDet-S achieves 45.3% AP and 294 FPS, outperforming many other detectors in both speed and accuracy. Xin Jiang 0010, Lu Jin 0001, Zechao Li |
ICME | 2 |
| 2024 | Delving into Multimodal Prompting for Fine-Grained Visual ClassificationabstractFine-grained visual classification (FGVC) involves categorizing fine subdivisions within a broader category, which poses challenges due to subtle inter-class discrepancies and large intra-class variations. However, prevailing approaches primarily focus on uni-modal visual concepts. Recent advancements in pre-trained vision-language models have demonstrated remarkable performance in various high-level vision tasks, yet the applicability of such models to FGVC tasks remains uncertain. In this paper, we aim to fully exploit the capabilities of cross-modal description to tackle FGVC tasks and propose a novel multimodal prompting solution, denoted as MP-FGVC, based on the contrastive language-image pertaining (CLIP) model. Our MP-FGVC comprises a multimodal prompts scheme and a multimodal adaptation scheme. The former includes Subcategory-specific Vision Prompt (SsVP) and Discrepancy-aware Text Prompt (DaTP), which explicitly highlights the subcategory-specific discrepancies from the perspectives of both vision and language. The latter aligns the vision and text prompting elements in a common semantic space, facilitating cross-modal collaborative reasoning through a Vision-Language Fusion Module (VLFM) for further improvement on FGVC. Moreover, we tailor a two-stage optimization strategy for MP-FGVC to fully leverage the pre-trained CLIP model and expedite efficient adaptation for FGVC. Extensive experiments conducted on four FGVC datasets demonstrate the effectiveness of our MP-FGVC. Xin Jiang 0010, Hao Tang 0007, Junyao Gao 0002, Xiaoyu Du 0002, Shengfeng He, Zechao Li |
AAAI | 1 |
| 2024 | DVF: Advancing Robust and Accurate Fine-Grained Image Retrieval with Retrieval GuidelinesabstractFine-grained image retrieval (FGIR) is to learn visual representations that distinguish visually similar objects while maintaining generalization. Existing methods propose to generate discriminative features, but rarely consider the particularity of the FGIR task itself. This paper presents a meticulous analysis leading to the proposal of practical guidelines to identify subcategory-specific discrepancies and generate discriminative features to design effective FGIR models. These guidelines include emphasizing the object (G1), highlighting subcategory-specific discrepancies (G2), and employing effective training strategy (G3). Following G1 and G2, we design a novel Dual Visual Filtering mechanism for the plain visual transformer, denoted as DVF, to capture subcategory-specific discrepancies. Specifically, the dual visual filtering mechanism comprises an object-oriented module and a semantic-oriented module. These components serve to magnify objects and identify discriminative regions, respectively. Following G3, we implement a discriminative model training strategy to improve the discriminability and generalization ability of DVF. Extensive analysis and ablation studies confirm the efficacy of our proposed guidelines. Without bells and whistles, the proposed DVF achieves state-of-the-art performance on three widely-used fine-grained datasets in closed-set and open-set settings. Xin Jiang 0010, Hao Tang 0007, Rui Yan 0010, Jinhui Tang 0001, Zechao Li |
ACM Multimedia | 1 |
| 2024 | Learning Contrastive Self-Distillation for Ultra-Fine-Grained Visual Categorization Targeting Limited SamplesabstractIn the field of intelligent multimedia analysis, ultra-fine-grained visual categorization (Ultra-FGVC) plays a vital role in distinguishing intricate subcategories within broader categories. However, this task is inherently challenging due to the complex granularity of category subdivisions and the limited availability of data for each category. To address these challenges, this work proposes CSDNet, a pioneering framework that effectively explores contrastive learning and self-distillation to learn discriminative representations specifically designed for Ultra-FGVC tasks. CSDNet comprises three main modules: Subcategory-Specific Discrepancy Parsing (SSDP), Dynamic Discrepancy Learning (DDL), and Subcategory-Specific Discrepancy Transfer (SSDT), which collectively enhance the generalization of deep models across instance, feature, and logit prediction levels. To increase the diversity of training samples, the SSDP module introduces adaptive augmented samples to spotlight subcategory-specific discrepancies. Simultaneously, the proposed DDL module stores historical intermediate features by a dynamic memory queue, which optimizes the feature learning space through iterative contrastive learning. Furthermore, the SSDT module effectively distills subcategory-specific discrepancies knowledge from the inherent structure of limited training data using a self-distillation paradigm at the logit prediction level. Experimental results demonstrate that CSDNet outperforms current state-of-the-art Ultra-FGVC methods, emphasizing its powerful efficacy and adaptability in addressing Ultra-FGVC tasks. Ziye Fang, Xin Jiang 0010, Hao Tang 0007, Zechao Li |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | Global Meets Local: Dual Activation Hashing Network for Large-Scale Fine-Grained Image RetrievalabstractIn the Internet era, the exponential growth of fine-grained image databases poses a considerable challenge for efficient information retrieval. Hashing-based approaches gained traction for their computational and storage efficiency, yet fine-grained hashing retrieval presents unique challenges due to small inter-class and large intra-class variations inherent to fine-grained entities. Thus, traditional hashing algorithms falter in discerning these subtle, yet critical, visual differences and fail to generate compact yet semantically rich hash codes. To address this, we introduce a Dual Activation Hashing Network (DAHNet) designed to convert high-dimensional image data into optimized binary codes via an innovative feature activation paradigm. The architecture consists of dual branches specifically tailored for global and local semantic activation, thereby establishing direct correspondences between hash codes and distinguishable object parts through a hierarchical activation pipeline. Specifically, our spatial-oriented semantic activation module modulates dominant visual regions while amplifying the activations of subtle yet semantically rich areas in a controlled manner. Building on these activated visual representations, the proposed inter-region semantic enrichment module further enriches them by unearthing semantically complementary cues. Concurrently,DAHNetintegrates a channel-oriented semantic activation module that exploits channel-specific correlations to distill contextual cues from spatially-activated visual features, thereby reinforcing robust learning to hash. To maintain the similarity of the original entities, we amalgamate final hash codes from both activation branches, capturing both local textural details and global structural information. Comprehensive evaluations on five fine-grained image retrieval benchmarks demonstrateDAHNet's superior performance over existing state-of-the-art hashing solutions, especially on 12-bit, improving performance by 4%-15% compared to the current best results on the five benchmarks. Moreover, generalization studies validate the efficacy of our dual-activation framework in the domain of content-based fine-grained image retrieval. The code is publicly available at:https://github.com/WhiteJiang/DAHNet. Xin Jiang 0010, Hao Tang 0007, Zechao Li |
IEEE Trans. Knowl. Data Eng. | 1 |