VLDB 2026 Research / reviewers in the wild / expert
Fei Shen 0004
dblp:99/6100-4
· DBLP profile ↗
54ranked-venue papers
12as first author
54since 2021 · last 2026
0000-0001-9749-0234ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 29 · 7 first-author · 29 since 2021Artificial intelligence and machine learning · 28 · 5 first-author · 28 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 1 first-author · 9 since 2021Computer networks · 2 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Fine-Grained Image Retrieval via Dual-Vision AdaptationabstractFine-Grained Image Retrieval~(FGIR) faces challenges in learning discriminative visual representations to retrieve images with similar fine-grained features. Current leading FGIR solutions typically follow two regimes: enforce pairwise similarity constraints in the semantic embedding space, or incorporate a localization sub-network to fine-tune the entire model. However, such two regimes tend to overfit the training data while forgetting the knowledge gained from large-scale pre-training, thus reducing their generalization ability. In this paper, we propose a Dual-Vision Adaptation (DVA) approach for FGIR, which guides the frozen pre-trained model to perform FGIR through collaborative sample and feature adaptation. Specifically, we design Object-Perceptual Adaptation, which modifies input samples to help the pre-trained model perceive critical objects and elements within objects that are helpful for category prediction. Meanwhile, we propose In-Context Adaptation, which introduces a small set of parameters for feature adaptation without modifying the pre-trained parameters. This makes the FGIR task using these adapted features closer to the task solved during the pre-training. Additionally, to balance retrieval efficiency and performance, we propose Discrimination Perception Transfer to transfer the discriminative knowledge in the object-perceptual adaptation to the image encoder using the knowledge distillation mechanism. Extensive experiments show that DVA performs well on three fine-grained datasets. Xin Jiang 0010, Meiqi Cao, Hao Tang 0007, Fei Shen 0004, Zechao Li |
AAAI | 4 |
| 2026 | StyleTailor: Towards Personalized Fashion Styling via Hierarchical Negative FeedbackabstractThe advancement of intelligent agents has revolutionized problem-solving across diverse domains, yet solutions for personalized fashion styling remain underexplored, which holds immense promise for promoting shopping experiences. In this work, we present StyleTailor, the first collaborative agent framework that seamlessly unifies personalized apparel design, shopping recommendation, virtual try-on, and systematic evaluation into a cohesive workflow. To this end, StyleTailor pioneers an iterative visual refinement paradigm driven by multi-level negative feedback, enabling adaptive and precise user alignment. Specifically, our framework features two core agents, i.e., Designer for personalized garment selection and Consultant for virtual try-on, whose outputs are progressively refined via hierarchical vision-language model feedback spanning individual items, complete outfits, and try-on efficacy. Counterexamples are aggregated into negative prompts, forming a closed-loop mechanism that enhances recommendation quality. To assess the performance, we introduce a comprehensive evaluation suite encompassing style consistency, visual quality, face similarity, and artistic appraisal. Extensive experiments demonstrate StyleTailor's superior performance in delivering personalized designs and recommendations, outperforming strong baselines without negative feedback and establishing a new benchmark for intelligent fashion systems. Hongbo Ma, Fei Shen 0004, Xiaoce Wang, Jinkai Zheng, Liangqiong Qu, Ming Li 0073 |
AAAI | 2 |
| 2026 | SGMHand: Structure-Guided Modulation for Structure-Aware Hand InpaintingabstractDiffusion-based generative models have demonstrated remarkable capabilities in image synthesis, yet realistic hand generation remains a persistent challenge due to complex articulations, self-occlusion, and the lack of explicit structural guidance. To address these issues, we present SGMHand, a novel structure-guided hand inpainting framework that explicitly injects topological priors to enhance structural fidelity and spatial precision. Specifically, we present a structure-guided modulation (SGM) module that synergistically combines structure spatial attention with global feature calibration, enabling fine-grained geometric control over the generative process. Then, we devise a keypoint-aware (KA) loss that enforces topological coherence by aligning attention activations with structures, thereby bridging the gap between high-level semantics and low-level geometry. By jointly optimizing over structural constraints in both representation and learning objectives, SGMHand achieves semantically consistent and geometrically plausible hand synthesis, even under severe occlusion. Extensive experiments demonstrate the effectiveness and strong generalization ability of SGMHand across various foundation models, significantly enhancing the quality and realism of human image synthesis in diverse scenarios. Chuancheng Shi, Shiming Guo, Ke Shui, Fei Shen 0004 |
AAAI | 5 |
| 2026 | Cross-modal Proxy Evolving for OOD Detection with Vision-Language ModelsabstractReliable zero-shot detection of out-of-distribution (OOD) inputs is critical for deploying vision-language models in open-world settings. However, the lack of labeled negatives in zero-shot OOD detection necessitates proxy signals that remain effective under distribution shift. Existing negative-label methods rely on a fixed set of textual proxies, which (i) sparsely sample the semantic space beyond in-distribution (ID) classes and (ii) remain static while only visual features drift, leading to cross-modal misalignment and unstable predictions. In this paper, we propose CoEvo, a training- and annotation-free test-time framework that performs bidirectional, sample-conditioned adaptation of both textual and visual proxies. Specifically, CoEvo introduces a proxy-aligned co-evolution mechanism to maintain two evolving proxy caches, which dynamically mines contextual textual negatives guided by test images and iteratively refines visual proxies, progressively realigning cross-modal similarities and enlarging local OOD margins. Finally, we dynamically re-weight the contributions of dual-modal proxies to obtain a calibrated OOD score that is robust to distribution shift. Extensive experiments on standard benchmarks demonstrate that CoEvo achieves state-of-the-art performance, improving AUROC by 1.33% and reducing FPR95 by 45.98% on ImageNet-1K compared to strong negative-label baselines. Hao Tang 0007, Yu Liu 0158, Shuanglin Yan, Fei Shen 0004, Shengfeng He, Harry Qin |
AAAI | 4 |
| 2026 | HiMo-CLIP: Modeling Semantic Hierarchy and Monotonicity in Vision-Language AlignmentabstractContrastive vision-language models like CLIP have achieved impressive results in image-text retrieval by aligning image and text representations in a shared embedding space. However, these models often treat text as flat sequences, limiting their ability to handle complex, compositional, and long-form descriptions. In particular, they fail to capture two essential properties of language: semantic hierarchy, which reflects the multi-level compositional structure of text, and semantic monotonicity, where richer descriptions should result in stronger alignment with visual content. To address these limitations, we propose HiMo-CLIP, a representation-level framework that enhances CLIP-style models without modifying the encoder architecture. HiMo-CLIP introduces two key components: a hierarchical decomposition (HiDe) module that extracts latent semantic components from long-form text via in-batch PCA, enabling flexible, batch-aware alignment across different semantic granularities, and a monotonicity-aware contrastive loss (MoLo) that jointly aligns global and component-level representations, encouraging the model to internalize semantic ordering and alignment strength as a function of textual completeness. These components work together to produce structured, cognitively aligned cross-modal representations. Experiments on multiple image-text retrieval benchmarks show that HiMo-CLIP consistently outperforms strong baselines, particularly under long or compositional descriptions. Ruijia Wu, Fei Shen 0004, Shaoan Zhao, Qiang Hui, Huanlin Gao, Zhaoxiang Liu, Kai Wang 0012, Shiguo Lian |
AAAI | 3 |
| 2026 | Seeing through the Conflict: Transparent Knowledge Conflict Handling in Retrieval-Augmented GenerationabstractLarge language models (LLMs) equipped with retrieval—the Retrieval-Augmented Generation (RAG) paradigm—should combine their parametric knowledge with external evidence, yet in practice they often hallucinate, over-trust noisy snippets, or ignore vital context. We introduce TCR (Transparent Conflict Resolution), a plug-and-play framework that makes this decision process observable and controllable. TCR (i) disentangles semantic match and factual consistency via dual contrastive encoders, (ii) estimates self-answerability to gauge confidence in internal memory, and (iii) feeds the three scalar signals to the generator through a lightweight soft-prompt with SNR-based weighting. Across seven benchmarks TCR improves conflict detection (+5–18 F₁), raises knowledge-gap recovery by +21.4 percentage points and cuts misleading-context overrides by –29.3 percentage points, while adding only 0.3% parameters. The signals align with human judgements and expose temporal decision patterns. Ziqi Zhong, Canran Xiao, Haoliang Zhang, Fei Shen 0004 |
AAAI | 7 |
| 2026 | IMAGGarment+: Efficient Attribute-Wise Diffusion for Garment GenerationabstractDiffusion models have advanced fine-grained garment generation, yet balancing controllability, efficiency, and texture fidelity remains challenging. Adapter-based methods often yield incoherent details, while full fine-tuning is computationally expensive and prone to overwriting pretrained priors. To address these limitations, we propose IMAGGarment+, an efficient diffusion framework for controllable and high-quality garment synthesis. It comprises two key modules designed for efficient and attribute-aware conditioning. First, we introduce an attribute-wise feature extractor (AFE) that disentangles key garment attributes, silhouette, logo, position, and color, into parallel latent streams. Each stream is optimized independently via LoRA, ensuring minimal parameter overhead while retaining expressive capacity. Second, we develop an attribute-adaptive attention (AA) module to inject attribute-specific cues into the generative process through a selective, layer-wise injection strategy. Specifically, silhouette and color features are injected into early decoder layers to guide structural and appearance formation, while logo features are propagated across all layers to ensure cross-scale consistency. Extensive experiments on fine-grained garment benchmarks demonstrate that IMAGGarment+ outperforms state-of-the-art baselines with less than 20% additional parameters, validating its effectiveness and efficiency. Fei Shen 0004, Cong Wang 0034, Yanpeng Sun, Hao Tang 0007, Xiaoyu Du 0002 |
AAAI | 2 |
| 2026 | Focusing Condition: Inference-Time Self-Contrastive Steering Elicits Better Conditional Text Embeddings in LLMsabstractZifeng Cheng, Lingyun Qian, Zhiwei Jiang, Cong Wang, Yafeng Yin, Fei Shen, Ao Zhou, Qing Gu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Zifeng Cheng, Lingyun Qian, Zhiwei Jiang 0001, Cong Wang 0034, Yafeng Yin 0002, Fei Shen 0004, Qing Gu 0001 |
ACL (1) | 6 |
| 2026 | Maximizing Local Entropy Where It Matters: Prefix-Aware Localized LLM UnlearningabstractNaixin Zhai, Pengyang Shao, Binbin Zheng, Yonghui Yang, Fei Shen, Long Bai, Xun Yang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Naixin Zhai, Pengyang Shao, Yonghui Yang 0001, Fei Shen 0004, Long Bai 0015, Xun Yang 0001 |
ACL (1) | 5 |
| 2026 | Curiosity-driven cooperation for long-tailed multi-label learning
Canran Xiao, Chuangxin Zhao, Zong Ke, Fei Shen 0004, Li Shen 0008 |
Neural Networks | 4 |
| 2026 | Progressive local self-attention for content-aligned super-resolution
Detian Huang, Xiancheng Zhu, Fei Shen 0004, Taotao Lai, Huanqiang Zeng, Junhui Hou |
Pattern Recognit. | 4 |
| 2026 | Selection, Aggregation, and Enhancement: Trajectory Consistent Diffusion Model for Image Super-ResolutionabstractDiffusion models have shown strong promise for image super-resolution (ISR). However, current approaches often underuse pretrained diffusion backbones and lack constraints on the sampling trajectory, which degrades structural consistency and fine details. For that, we introduce the trajectory consistent diffusion model (TCDM) for super-resolution, which jointly optimizes the sampling process through lightweight components and inference-time strategies while keeping the diffusion backbone frozen, yielding high-fidelity, detail-rich reconstructions. First, we propose a dynamic semantic selection (DSS) mechanism that records early intermediates, matches them to upsampled low-resolution features, and reconditions sampling with the best match to reduce the mismatch between conditioning and noise scale. Next, we design a cross-step aggregation guidance (CAG) strategy that aggregates features from the current state with the selected intermediate to enforce trajectory-level consistency in noise prediction. Finally, we present a plug-and-play frequency enhancement adapter (FE-Adapter) that injects different frequency-domain cues into the encoder during training, strengthening high-frequency perception while preserving global structures. Extensive experiments on multiple ISR benchmarks show that TCDM achieves strong structural fidelity and competitive no-reference perceptual quality, offering a favorable fidelity-perception trade-off. Detian Huang, Yaohui Guo, Luanyuan Dai, Fei Shen 0004, Huanqiang Zeng |
IEEE Trans. Image Process. | 5 |
| 2026 | Progressive Feature Encoding With Background Perturbation Learning for Ultra-Fine-Grained Visual CategorizationabstractUltra-Fine-Grained Visual Categorization (Ultra-FGVC) aims to classify objects into sub-granular categories, presenting the challenge of distinguishing visually similar objects with limited data. Existing methods primarily address sample scarcity but often overlook the importance of leveraging intrinsic object features to construct highly discriminative representations. This limitation significantly constrains their effectiveness in Ultra-FGVC tasks. To address these challenges, we propose SV-Transformer that progressively encodes object features while incorporating background perturbation modeling to generate robust and discriminative representations. At the core of our approach is a progressive feature encoder, which hierarchically extracts global semantic structures and local discriminative details from backbone-generated representations. This design enhances inter-class separability while ensuring resilience to intra-class variations. Furthermore, our background perturbation learning mechanism introduces controlled variations in the feature space, effectively mitigating the impact of sample limitations and improving the model's capacity to capture fine-grained distinctions. Comprehensive experiments demonstrate that SV-Transformer achieves state-of-the-art performance on benchmark Ultra-FGVC datasets, showcasing its efficacy in addressing the challenges of Ultra-FGVC task. Xin Jiang 0010, Ziye Fang, Fei Shen 0004, Junyao Gao 0002, Zechao Li |
IEEE Trans. Image Process. | 3 |
| 2026 | IMAGGarment: Fine-Grained Garment Generation for Controllable Fashion DesignabstractThis paper presents IMAGGarment, a fine-grained garment generation (FGG) framework that enables high-fidelity garment synthesis with precise control over silhouette, color, and logo placement. Unlike existing methods that are limited to single-condition inputs, IMAGGarment addresses the challenges of multi-conditional controllability in personalized fashion design and digital apparel applications. Specifically, IMAGGarment employs a two-stage training strategy to separately model global appearance and local details, while enabling unified and controllable generation through end-to-end inference. In the first stage, we propose a global appearance model that jointly encodes silhouette and color using a mixed attention module and a color adapter. In the second stage, we present a local enhancement model with an adaptive appearance-aware module to inject user-defined logos and spatial constraints, enabling accurate placement and visual consistency. To support this task, we release GarmentBench, a large-scale dataset comprising over 180 K garment samples paired with multi-level design conditions, including sketches, color references, logo placements, and textual prompts. Extensive experiments demonstrate that our method outperforms existing baselines, achieving superior structural stability, color fidelity, and local controllability performance. Fei Shen 0004, Cong Wang 0034, Xin Jiang 0010, Xiaoyu Du 0002, Jinhui Tang 0001 |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2025 | Exploring the Potential of Large Vision-Language Models for Unsupervised Text-Based Person RetrievalabstractThe aim of text-based person retrieval is to identify pedestrians using natural language descriptions within a large-scale image gallery. Traditional methods rely heavily on manually annotated image-text pairs, which are resource-intensive to obtain. With the emergence of Large Vision-Language Models (LVLMs), the advanced capabilities of contemporary models in image understanding have led to the generation of highly accurate captions. Therefore, this paper explores the potential of employing Large Vision-Language Models for unsupervised text-based pedestrian image retrieval and proposes a Multi-grained Uncertainty Modeling and Alignment framework (MUMA). Initially, multiple Large Vision-Language Models are employed to generate diverse and hierarchically structured pedestrian descriptions across different styles and granularities. However, the generated captions inevitably introduce noise. To address this issue, an uncertainty-guided sample filtration module is proposed to estimate and filter out unreliable image-text pairs. Additionally, to simulate the diversity of styles and granularities in captions, a multi-grained uncertainty modeling approach is applied to model the distributions of captions, with each caption represented as a multivariate Gaussian distribution. Finally, a multi-level consistency distillation loss is employed to integrate and align the multi-grained captions, aiming to transfer knowledge across different granularities. Experimental evaluations conducted on three widely-used datasets demonstrate the significant advancements achieved by our approach. Zongyi Li, Yuxuan Shi 0002, Jiazhong Chen, Shijuan Huang, Linnan Tu, Fei Shen 0004 |
AAAI | 7 |
| 2025 | IMAGDressing-v1: Customizable Virtual DressingabstractExisting virtual try-on (VTON) methods provide only limited user control over garment attributes and generally overlook essential factors such as face, pose, and scene context. To address these limitations, we introduce the virtual dressing (VD) task, which aims to synthesize freely editable human images conditioned on fixed garments and optional user-defined inputs. We further propose a comprehensive affinity metric index (CAMI) to quantify the consistency between generated outputs and reference garments. We present IMAGDressing-v1, which leverages a garment-specific U-Net to integrate semantic features from CLIP and texture features from a VAE. To incorporate these garment features into a frozen denoising U-Net for flexible text-driven scene control, we employ a hybrid attention mechanism composed of frozen self-attention and trainable cross-attention layers. IMAGDressing-v1 seamlessly integrates with extension modules, such as ControlNet and IP-Adapter, enabling enhanced diversity and controllability. To alleviate data constraints, we introduce the Interactive Garment Pairing (IGPair) dataset, comprising over 300,000 garment–image pairs and a standardized data assembly pipeline. Extensive experiments demonstrate that IMAGDressing-v1 achieves state-of-the-art performance in controlled human image synthesis. The code and model will be available at https://github.com/muzishen/IMAGDressing. Fei Shen 0004, Xin Jiang 0010, Hu Ye, Cong Wang 0034, Xiaoyu Du 0002, Zechao Li, Jinhui Tang 0001 |
AAAI | 1 |
| 2025 | Boosting Consistency in Story Visualization with Rich-Contextual Conditional Diffusion ModelsabstractRecent research showcases the considerable potential of conditional diffusion models for generating consistent stories. However, current methods, which primarily generate stories in a caption-dependent manner, often overlook the importance of contextual consistency and the relevance of frames during sequential generation. To address this, we propose a novel Rich-contextual Conditional Diffusion Models (RCDMs), a two-stage approach designed to enhance story generation's semantic consistency and temporal consistency. Specifically, in the first stage, the frame-prior transformer diffusion model is presented to predict the frame semantic embedding of the unknown clip by aligning the semantic correlations between the captions and frames of the known clip. The second stage establishes a robust model with rich contextual conditions, including reference images of the known clip, the predicted frame semantic embedding of the unknown clip, and text embeddings of all captions. By jointly injecting these rich contextual conditions at the image and feature levels, RCDMs can generate semantic and temporal consistency stories. Moreover, RCDMs can generate consistent stories with a single forward inference compared to autoregressive models. Our qualitative and quantitative results demonstrate that our proposed RCDMs outperform in challenging scenarios. Fei Shen 0004, Hu Ye, Sibo Liu, Jun Zhang 0018, Cong Wang 0034, Xiao Han 0011 |
AAAI | 1 |
| 2025 | DVP-MVS: Synergize Depth-Edge and Visibility Prior for Multi-View StereoabstractPatch deformation-based methods have recently exhibited substantial effectiveness in multi-view stereo, due to the incorporation of deformable and expandable perception to reconstruct textureless areas. However, such approaches typically focus on exploring correlative reliable pixels to alleviate match ambiguity during patch deformation, but ignore the deformation instability caused by mistaken edge-skipping and visibility occlusion, leading to potential estimation deviation. To remedy the above issues, we propose DVP-MVS, which innovatively synergizes depth-edge aligned and cross-view prior for robust and visibility-aware patch deformation. Specifically, to avoid unexpected edge-skipping, we first utilize Depth Anything V2 followed by the Roberts operator to initialize coarse depth and edge maps respectively, both of which are further aligned through an erosion-dilation strategy to generate fine-grained homogeneous boundaries for guiding patch deformation. In addition, we reform view selection weights as visibility maps and restore visible areas by cross-view depth reprojection, then regard them as cross-view prior to facilitate visibility-aware patch deformation. Finally, we improve propagation and refinement with multi-view geometry consistency by introducing aggregated visible hemispherical normals based on view selection and local projection depth differences based on epipolar lines, respectively. Extensive evaluations on ETH3D and Tanks & Temples benchmarks demonstrate that our method can achieve state-of-the-art performance with excellent robustness and generalization. Zhenlong Yuan, Jinguo Luo, Fei Shen 0004, Zhaoxin Li, Tianlu Mao |
AAAI | 3 |
| 2025 | MSP-MVS: Multi-Granularity Segmentation Prior Guided Multi-View StereoabstractRecently, patch deformation-based methods have demonstrated significant strength in multi-view stereo by adaptively expanding the reception field of patches to help reconstruct textureless areas. However, such methods mainly concentrate on searching for pixels without matching ambiguity (i.e., reliable pixels) when constructing deformed patches, while neglecting the deformation instability caused by unexpected edge-skipping, resulting in potential matching distortions. Addressing this, we propose MSP-MVS, a method introducing multi-granularity segmentation prior for edge-confined patch deformation. Specifically, to avoid unexpected edge-skipping, we first aggregate and further refine multi-granularity depth edges gained from Semantic-SAM as prior to guide patch deformation within depth-continuous (i.e., homogeneous) areas. Moreover, to address attention imbalance caused by edge-confined patch deformation, we implement adaptive equidistribution and disassemble-clustering of correlative reliable pixels (i.e., anchors), thereby promoting attention-consistent patch deformation. Finally, to prevent deformed patches from falling into local-minimum matching costs caused by the fixed sampling pattern, we introduce disparity-sampling synergistic 3D optimization to help identify global-minimum matching costs. Evaluations on ETH3D and Tanks & Temples benchmarks prove our method obtains state-of-the-art performance with remarkable generalization. Zhenlong Yuan, Fei Shen 0004, Zhaoxin Li, Jinguo Luo, Tianlu Mao |
AAAI | 3 |
| 2025 | SPEAK: Speech-Driven Pose and Emotion-Adjustable Talking Head GenerationabstractMost earlier researches on talking face generation have focused on the synchronization of lip motion and speech content. However, head pose and facial emotions are equally important characteristics of natural faces. While audio-driven talking face generation has seen notable advancements, existing methods either overlook facial emotions or are limited to specific individuals and cannot be applied to arbitrary subjects. In this paper, we propose a novel one-shot Talking Head Generation framework (SPEAK) that distinguishes itself from the general Talking Face Generation by enabling emotional and postural control. Specifically, we introduce Inter-Reconstructed Feature Disentanglement (IRFD) module to decouple facial features into three latent spaces. Then we design a face editing module that modifies speech content and facial latent codes into a single latent space. Subsequently, we present a novel generator that employs modified latent codes derived from the editing module to regulate emotional expression, head poses, and speech content in synthesizing facial animations. Extensive trials demonstrate that our method ensures lip synchronization with the audio while enabling decoupled control of facial features, it can generate realistic talking head with coordinated lip motions, authentic facial emotions, and smooth head movements. The demo video is available: https://anonymous.4open.science/r/SPEAK-8A22. Changpeng Cai, Guinan Guo, Junhao Su, Fei Shen 0004, Chenghao He, Yuanxu Chen |
ICASSP | 5 |
| 2025 | PrimeK-Net: Multi-scale Spectral Learning via Group Prime-Kernel Convolutional Neural Networks for Single Channel Speech EnhancementabstractSingle-channel speech enhancement is a challenging ill-posed problem focused on estimating clean speech from degraded signals. Existing studies have demonstrated the competitive performance of combining convolutional neural networks (CNNs) with Transformers in speech enhancement tasks. However, existing frameworks have not sufficiently addressed computational efficiency and have overlooked the natural multi-scale distribution of the spectrum. Additionally, the potential of CNNs in speech enhancement has yet to be fully realized. To address these issues, this study proposes a Deep Separable Dilated Dense Block (DSDDB) and a Group Prime Kernel Feedforward Channel Attention (GPFCA) module. Specifically, the DSDDB introduces higher parameter and computational efficiency to the Encoder/Decoder of existing frameworks. The GPFCA module replaces the position of the Conformer, extracting deep temporal and frequency features of the spectrum with linear complexity. The GPFCA leverages the proposed Group Prime Kernel Feed-forward Network (GPFN) to integrate multi-granularity long-range, medium-range, and short-range receptive fields, while utilizing the properties of prime numbers to avoid periodic overlap effects. Experimental results demonstrate that PrimeK-Net, proposed in this study, achieves state-of-the-art (SOTA) performance on the VoiceBank+Demand dataset, reaching a PESQ score of 3.61 with only 1.41M parameters. Zizhen Lin, Ruili Li, Fei Shen 0004, Xi Xuan |
ICASSP | 4 |
| 2025 | Towards Maximizing Semantic Coverage for Image-Text RetrievalabstractTo establish semantic associations between images and texts, existing Image-Text Retrieval (ITR) methods primarily focus on fixed-scale fragments, which only identify explicit semantic categories. Consequently, semantic coverage is constrained, leading to the omission of certain semantic associations. To enlarge the semantic coverage, we propose the Semantic Coverage-Aware Network (SCA-Net). First, explicit semantic categories are identified by SCA-Net through analyzing the semantic membership of visual and textual fragments, thereby establishing more precise explicit semantic associations. Second, implicit semantic categories are identified by SCA-Net via adaptively aggregating visual and textual fragments across various scales using a co-occurrence-aware router, thereby significantly expanding the semantic coverage and establishing complete semantic associations. Third, image-text similarity is calculated using the attention mechanism over a broader range of semantic coverage. Extensive experiments demonstrate that SCA-Net significantly enhances ITR performance compared to state-of-the-art methods by maximizing semantic coverage. Zhumin Chen, Fei Shen 0004 |
ICASSP | 4 |
| 2025 | FaceShot: Bring Any Character into LifeabstractIn this paper, we present ***FaceShot***, a novel training-free portrait animation framework designed to bring any character into life from any driven video without fine-tuning or retraining.
We achieve this by offering precise and robust reposed landmark sequences from an appearance-guided landmark matching module and a coordinate-based landmark retargeting module.
Together, these components harness the robust semantic correspondences of latent diffusion models to produce facial motion sequence across a wide range of character types.
After that, we input the landmark sequences into a pre-trained landmark-driven animation model to generate animated video.
With this powerful generalization capability, FaceShot can significantly extend the application of portrait animation by breaking the limitation of realistic portrait landmark detection for any stylized character and driven video.
Also, FaceShot is compatible with any landmark-driven animation model, significantly improving overall performance.
Extensive experiments on our newly constructed character benchmark CharacBench confirm that FaceShot consistently surpasses state-of-the-art (SOTA) approaches across any character domain.
More results are available at our project website https://faceshot2024.github.io/faceshot/. Junyao Gao 0002, Yanan Sun 0005, Fei Shen 0004, Xin Jiang 0010, Zhening Xing, Kai Chen 0026, Cairong Zhao |
ICLR | 3 |
| 2025 | Ensembling Diffusion Models via Adaptive Feature AggregationabstractThe success of the text-guided diffusion model has inspired the development and release of numerous powerful diffusion models within the open-source community. These models are typically fine-tuned on various expert datasets, showcasing diverse denoising capabilities. Leveraging multiple high-quality models to produce stronger generation ability is valuable, but has not been extensively studied. Existing methods primarily adopt parameter merging strategies to produce a new static model. However, they overlook the fact that the divergent denoising capabilities of the models may dynamically change across different states, such as when experiencing different prompts, initial noises, denoising steps, and spatial locations. In this paper, we propose a novel ensembling method, Adaptive Feature Aggregation (AFA), which dynamically adjusts the contributions of multiple models at the feature level according to various states (i.e., prompts, initial noises, denoising steps, and spatial locations), thereby keeping the advantages of multiple diffusion models, while suppressing their disadvantages. Specifically, we design a lightweight Spatial-Aware Block-Wise (SABW) feature aggregator that adaptive aggregates the block-wise intermediate features from multiple U-Net denoisers into a unified one. The core idea lies in dynamically producing an individual attention map for each model's features by comprehensively considering various states. It is worth noting that only SABW is trainable with about 50 million parameters, while other models are frozen. Both the quantitative and qualitative experiments demonstrate the effectiveness of our proposed method. Cong Wang 0034, Kuan Tian, Yonghang Guan, Fei Shen 0004, Zhiwei Jiang 0001, Qing Gu 0001, Jun Zhang 0018 |
ICLR | 4 |
| 2025 | AS-Memory: Adaptive Sparse Memory Meeting Video-Language ModelsabstractLong-term video understanding in intelligent transportation systems (ITS) has advanced significantly with the integration of large language models (LLMs) and vision foundation models. However, existing LLM-based multimodal approaches are limited by context length and memory constraints, restricting their effectiveness to short video scenarios. To address these challenges, we propose AS-Memory, a novel framework that combines adaptive sparse memory with LLMs for efficient and scalable long-term video understanding. AS-Memory introduces a plug-and-play memory bank, a lightweight module designed to seamlessly integrate with existing multimodal LLMs. This memory bank stores and retrieves historical video content, enabling long-term analysis while mitigating context length and GPU memory limitations. To further enhance efficiency, we propose a sparse adaptive mechanism that dynamically compresses redundant features and retains critical information, ensuring effective management of streaming video data. Comprehensive evaluations on long-term video understanding benchmarks demonstrate that AS-Memory consistently outperforms state-of-the-art methods in terms of accuracy. The source code and trained models will be made available to the public. Bimei Wang, Huilin Song, Jisheng Dang, Fei Shen 0004, Mangang Xie, Jizhao Liu, Jia-Si Weng 0001 |
ICME | 4 |
| 2025 | UniMC: Taming Diffusion Transformer for Unified Keypoint-Guided Multi-Class Image GenerationabstractAlthough significant advancements have been achieved in the progress of keypoint-guided Text-to-Image diffusion models, existing mainstream keypoint-guided models encounter challenges in controlling the generation of more general non-rigid objects beyond humans (e.g., animals). Moreover, it is difficult to generate multiple overlapping humans and animals based on keypoint controls solely. These challenges arise from two main aspects: the inherent limitations of existing controllable methods and the lack of suitable datasets.
First, we design a DiT-based framework, named UniMC, to explore unifying controllable multi-class image generation. UniMC integrates instance- and keypoint-level conditions into compact tokens, incorporating attributes such as class, bounding box, and keypoint coordinates. This approach overcomes the limitations of previous methods that struggled to distinguish instances and classes due to their reliance on skeleton images as conditions.
Second, we propose HAIG-2.9M, a large-scale, high-quality, and diverse dataset designed for keypoint-guided human and animal image generation. HAIG-2.9M includes 786K images with 2.9M instances. This dataset features extensive annotations such as keypoints, bounding boxes, and fine-grained captions for both humans and animals, along with rigorous manual inspection to ensure annotation accuracy.
Extensive experiments demonstrate the high quality of HAIG-2.9M and the effectiveness of UniMC, particularly in heavy occlusions and multi-class scenarios. Ailing Zeng, Dongxu Yue, Ceyuan Yang, Yang Cao 0017, Hanzhong Guo, Fei Shen 0004, Wei Liu 0005, Xihui Liu, Dan Xu 0002 |
ICML | 7 |
| 2025 | Long-Term TalkingFace Generation via Motion-Prior Conditional Diffusion ModelabstractRecent advances in conditional diffusion models have shown promise for generating realistic TalkingFace videos, yet challenges persist in achieving consistent head movement, synchronized facial expressions, and accurate lip synchronization over extended generations. To address these, we introduce the Motion-priors Conditional Diffusion Model (MCDM), which utilizes both archived and current clip motion priors to enhance motion prediction and ensure temporal consistency. The model consists of three key elements: (1) an archived-clip motion-prior that incorporates historical frames and a reference frame to preserve identity and context; (2) a present-clip motion-prior diffusion model that captures multimodal causality for accurate predictions of head movements, lip sync, and expressions; and (3) a memory-efficient temporal attention mechanism that mitigates error accumulation by dynamically storing and updating motion features. We also introduce the TalkingFace-Wild dataset, a multilingual collection of over 200 hours of footage across 10 languages. Experimental results demonstrate the effectiveness of MCDM in maintaining identity and motion continuity for long-term TalkingFace generation. Fei Shen 0004, Cong Wang 0018, Junyao Gao 0002, Jisheng Dang, Jinhui Tang 0001, Tat-Seng Chua |
ICML | 1 |
| 2025 | Visual Content Generation in the Era of Large Foundation ModelsabstractThe rapid advancements in large foundation models have significantly transformed the field of visual content generation, impacting domains such as image synthesis, video generation, and 3D modeling. This tutorial will provide an in-depth exploration of the state-of-the-art techniques and methodologies used in visual content generation, emphasizing the role of large-scale generative models. The tutorial will cover fundamental principles, model architectures, recent breakthroughs, and practical applications. We will discuss various generative paradigms, including diffusion models, autoregressive models, and large multimodal models, highlighting their strengths and limitations. Additionally, we will delve into the challenges of controllability, personalization, and realism in generated content, along with open research problems and future directions. By the end of the tutorial, attendees will gain a comprehensive understanding of contemporary visual content generation techniques and their applications, equipping them with the knowledge to leverage these models in their research and projects. Leigang Qu, Fei Shen 0004, Zhenglin Zhou, Jiayi Lyu, Wenjie Wang 0007, Lu Jiang 0004 |
ICMR | 2 |
| 2025 | SceneDecorator: Towards Scene-Oriented Story Generation with Scene Planning and Scene ConsistencyabstractRecent text-to-image models have revolutionized image generation, but they still struggle with maintaining concept consistency across generated images. While existing works focus on character consistency, they often overlook the crucial role of scenes in storytelling, which restricts their creativity in practice. This paper introduces scene-oriented story generation, addressing two key challenges: (i) scene planning, where current methods fail to ensure scene-level narrative coherence by relying solely on text descriptions, and (ii) scene consistency, which remains largely unexplored in terms of maintaining scene consistency across multiple stories. We propose SceneDecorator, a training-free framework that employs VLM-Guided Scene Planning to ensure narrative coherence across different scenes in a ``global-to-local'' manner, and Long-Term Scene-Sharing Attention to maintain long-term scene consistency and subject diversity across generated stories. Extensive experiments demonstrate the superior performance of SceneDecorator, highlighting its potential to unleash creativity in the fields of arts, films, and games. Quanjian Song, Fei Shen 0004, Xiaowei Hu 0001, Cunjian Chen, Pheng-Ann Heng |
NeurIPS | 4 |
| 2025 | Advanced Sign Language Video Generation with Compressed and Quantized Multi-Condition TokenizationabstractSign Language Video Generation (SLVG) seeks to generate identity-preserving sign language videos from spoken language texts. Existing methods primarily rely on the single coarse condition (e.g., skeleton sequences) as the intermediary to bridge the translation model and the video generation model, which limits both the naturalness and expressiveness of the generated videos. To overcome these limitations, we propose SignViP, a novel SLVG framework that incorporate multiple fine-grained conditions for improved generation fidelity. Rather than directly translating error-prone high-dimensional conditions, SignViP adopts a discrete tokenization paradigm to integrate and represent fine-grained conditions (i.e., fine-grained poses and 3D hands). SignViP contains three core components. (1) Sign Video Diffusion Model is jointly trained with a multi-condition encoder to learn continuous embeddings that encapsulate fine-grained motion and appearance. (2) Finite Scalar Quantization (FSQ) Autoencoder is further trained to compress and quantize these embeddings into discrete tokens for compact representation of the conditions. (3) Multi-Condition Token Translator is trained to translate spoken language text to discrete multi-condition tokens. During inference, Multi-Condition Token Translator first translates the spoken language text into discrete multi-condition tokens. These tokens are then decoded to continuous embeddings by FSQ Autoencoder, which are subsequently injected into Sign Video Diffusion Model to guide video generation. Experimental results show that SignViP achieves state-of-the-art performance across metrics, including video quality, temporal coherence, and semantic fidelity. The code is available at https://github.com/umnooob/signvip/. Cong Wang 0034, Zexuan Deng, Zhiwei Jiang 0001, Yafeng Yin 0002, Fei Shen 0004, Zifeng Cheng, Shiping Ge, Shiwei Gan, Qing Gu 0001 |
NeurIPS | 5 |
| 2025 | CoC-VLA: Delving into Adversarial Domain Transfer for Explainable Autonomous Driving via Chain-of-Causality Visual-Language-Action ModelabstractAutonomous driving represents a prominent application of artificial intelligence. Recent approaches have shifted from focusing solely on common scenarios to addressing complex, long-tail situations such as subtle human behaviors, traffic accidents, and non-compliant driving patterns. Given the demonstrated capabilities of large language models (LLMs) in understanding visual and natural language inputs and following instructions, recent methods have integrated LLMs into autonomous driving systems to enhance reasoning, interpretability, and performance across diverse scenarios. However, existing methods typically rely either on real-world data, which is suitable for industrial deployment, or on simulation data tailored to rare or hard case scenarios. Few approaches effectively integrate the complementary advantages of both data sources. To address this limitation, we propose a novel VLM-guided, end-to-end adversarial transfer framework for autonomous driving that transfers long-tail handling capabilities from simulation to real-world deployment, named CoC-VLA. The framework comprises a teacher VLM model, a student VLM model, and a discriminator. Both the teacher and student VLM models utilize a shared base architecture, termed the Chain-of-Causality Visual–Language Model (CoC VLM), which integrates temporal information via an end-to-end text adapter. This architecture supports chain-of-thought reasoning to infer complex driving logic. The teacher and student VLM models are pre-trained separately on simulated and real-world datasets. The discriminator is trained adversarially to facilitate the transfer of long-tail handling capabilities from simulated to real-world environments by the student VLM model, using a novel backpropagation strategy. Experimental results show that our method effectively bridges the gap between simulation and real-world autonomous driving, indicating a promising direction for future research. Fei Shen 0004, Yinda Chen, Peng Zhi, Rui Zhou 0005, Qingguo Zhou |
NeurIPS | 2 |
| 2025 | Triplet Contrastive Representation Learning for Unsupervised Vehicle Re-IdentificationabstractPart feature learning plays a crucial role in achieving fine-grained semantic understanding in unsupervised vehicle re-identification. However, existing approaches directly model part and global features, which can easily lead to severe gradient vanishing issues due to their unequal feature information and unreliable pseudo-labels. To address this problem, in this article, we propose a triplet contrastive representation learning (TCRL) framework, which leverages cluster features to bridge the part features and global features for unsupervised vehicle re-identification. Specifically, TCRL devises three memory banks to store the instance/cluster features and proposes a proxy contrastive loss (PCL) to make contrastive learning between adjacent memory banks, thus presenting the associations between the part and global features as a transition of the part-cluster and cluster-global associations. Since the cluster memory bank copes with all the vehicle features, it can summarize them into a discriminative feature representation. To deeply exploit the instance/cluster information, TCRL proposes two additional loss functions. For the instance-level feature, a hybrid contrastive loss (HCL) re-defines the sample correlations by approaching the positive instance features and pushing all negative instance features away. For the cluster-level feature, a weighted regularization cluster contrastive loss (WRCCL) refines the pseudo labels by penalizing the mislabeled images according to the instance similarity. Extensive experiments show that TCRL outperforms many state-of-the-art unsupervised vehicle re-identification approaches. Fei Shen 0004, Xiaoyu Du 0002, Liyan Zhang 0002, Xiangbo Shu, Jinhui Tang 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2024 | Synthesizing High-Quality Construction Segmentation Datasets Through Pre-trained Diffusion Model
Jiahao Huo, Zhengyao Wang, Fei Shen 0004 |
ICIC (10) | 5 |
| 2024 | Feature Pyramid Full Granularity Attention Network for Object Detection in Remote Sensing Imagery
Chang Liu 0183, Bowei Song, Fei Shen 0004 |
ICIC (10) | 6 |
| 2024 | Fourier-FPN: Fourier Improves Multi-scale Feature Learning for Oriented Tiny Object Detection
Hongan Pan, Fei Shen 0004, Zhengzhou Zhu, Honghui Jia |
ICIC (10) | 4 |
| 2024 | RSSRDiff: An Effective Diffusion Probability Model with Attention for Single Remote Sensing Image Super-Resolution
Tian Wei, Hanyi Zhang, Fei Shen 0004 |
ICIC (10) | 5 |
| 2024 | Training-Free Diffusion Models for Content-Style Synthesis
Ruipeng Xu, Fei Shen 0004, Zongyi Li |
ICIC (10) | 2 |
| 2024 | Enhancing Landslide Segmentation with Guide Attention Mechanism and Fast Fourier Transformer
Kaiyu Yan, Fei Shen 0004, Zongyi Li |
ICIC (10) | 2 |
| 2024 | Advancing Pose-Guided Image Synthesis with Progressive Conditional Diffusion ModelsabstractRecent work has showcased the significant potential of diffusion models in pose-guided person image synthesis.
However, owing to the inconsistency in pose between the source and target images, synthesizing an image with a distinct pose, relying exclusively on the source image and target pose information, remains a formidable challenge.
This paper presents Progressive Conditional Diffusion Models (PCDMs) that incrementally bridge the gap between person images under the target and source poses through three stages.
Specifically, in the first stage, we design a simple prior conditional diffusion model that predicts the global features of the target image by mining the global alignment relationship between pose coordinates and image appearance.
Then, the second stage establishes a dense correspondence between the source and target images using the global features from the previous stage, and an inpainting conditional diffusion model is proposed to further align and enhance the contextual features, generating a coarse-grained person image.
In the third stage, we propose a refining conditional diffusion model to utilize the coarsely generated image from the previous stage as a condition, achieving texture restoration and enhancing fine-detail consistency.
The three-stage PCDMs work progressively to generate the final high-quality and high-fidelity synthesized image.
Both qualitative and quantitative results demonstrate the consistency and photorealism of our proposed PCDMs under challenging scenarios.
The code and model will be available at https://github.com/tencent-ailab/PCDMs. Fei Shen 0004, Hu Ye, Jun Zhang 0018, Cong Wang 0034, Xiao Han 0011 |
ICLR | 1 |
| 2024 | Exploring Warping-Guided Features via Adaptive Latent Diffusion Model for Virtual try-onabstractVirtual try-on (VTON) aims to generate a target image that aligns with the reference garment using the source image and the reference clothing. The main challenge lies in accurately generating the warped details to fit the person while maintaining the original patterns of the clothing. However, most methods face significant challenges when directly capturing the complex structure of spatial transformations, especially when it’s necessary to infer warping features from the given source clothing, which often results in noticeable visual artifacts. For that, this paper proposes a novel adaptive latent diffusion model (ALDM) to implement warping-guided before generating target images, which contains two modules: prior warping module (PWM) and adaptive alignment module (AAM). Specifically, we first present the PWM to explore the features of warping-guided and align them with the target person’s posture. This extraction process is much simpler than directly generating the target image, as PWM focuses solely on this task. Then, we devise an AAM that utilizes the warping-guided features and reference clothing features mined in the previous stage to establish an adaptive alignment between them. Lastly, our extensive results on two large-scale datasets and a user study demonstrate the photorealism of our proposed ALDM under challenging scenarios. The code and model will be available at https://github.com/gaogao2002/ALDM. Bo Gao 0004, Junchi Ren, Fei Shen 0004, Mengwan Wei |
ICME | 3 |
| 2024 | LR-FPN: Enhancing Remote Sensing Object Detection with Location Refined Feature Pyramid NetworkabstractRemote sensing target detection aims to identify and locate critical targets within remote sensing images, finding extensive applications in agriculture and urban planning. Feature pyramid networks (FPNs) are commonly used to extract multi-scale features. However, existing FPNs often overlook extracting low-level positional information and fine-grained context interaction. To address this, we propose a novel location refined feature pyramid network (LR-FPN) to enhance the extraction of shallow positional information and facilitate fine-grained context interaction. The LR-FPN consists of two primary modules: the shallow position information extraction module (SPIEM) and the contextual interaction module (CIM). Specifically, SPIEM first maximizes the retention of solid location information of the target by simultaneously extracting positional and saliency information from the low-level feature map. Subsequently, CIM injects this robust location information into different layers of the original FPN through spatial and channel interaction, explicitly enhancing the object area. Moreover, in spatial interaction, we introduce a simple local and non-local interaction strategy to learn and retain the saliency information of the object. Lastly, the LR-FPN can be readily integrated into common object detection frameworks to improve performance significantly. Extensive experiments on two large-scale remote sensing datasets (i.e., DOTAV1.0 and HRSC2016) demonstrate that the proposed LR-FPN is superior to state-of-the-art object detection approaches. Our code and models will be publicly available. Hanqian Li, Ruinan Zhang, Junchi Ren, Fei Shen 0004 |
IJCNN | 5 |
| 2024 | IMAGPose: A Unified Conditional Framework for Pose-Guided Person GenerationabstractDiffusion models represent a promising avenue for image generation, having demonstrated competitive performance in pose-guided person image generation.
However, existing methods are limited to generating target images from a source image and a target pose, overlooking two critical user scenarios: generating multiple target images with different poses simultaneously and generating target images from multi-view source images.
To overcome these limitations, we propose IMAGPose, a unified conditional framework for pose-guided image generation, which incorporates three pivotal modules: a feature-level conditioning (FLC) module, an image-level conditioning (ILC) module, and a cross-view attention (CVA) module.
Firstly, the FLC module combines the low-level texture feature from the VAE encoder with the high-level semantic feature from the image encoder, addressing the issue of missing detail information due to the absence of a dedicated person image feature extractor.
Then, the ILC module achieves an alignment of images and poses to adapt to flexible and diverse user scenarios by injecting a variable number of source image conditions and introducing a masking strategy.
Finally, the CVA module introduces decomposing global and local cross-attention, ensuring local fidelity and global consistency of the person image when multiple source image prompts.
The three modules of IMAGPose work together to unify the task of person image generation under various user scenarios.
Extensive experiment results demonstrate the consistency and photorealism of our proposed IMAGPose under challenging user scenarios.
The code and model will be available at https://github.com/muzishen/IMAGPose. Fei Shen 0004, Jinhui Tang 0001 |
NeurIPS | 1 |
| 2023 | A Bag of Tricks for Fine-Grained Roof ExtractionabstractIn this article, we introduce the method we used in the 2023 IEEE GRSS Data Fusion Contest Track 1. The task demands a fine-grained classification method for semantic urban reconstruction. Our experiments are based on Swin transformer, combined with Double-Head module and RFLA (Gaussian Receptive Field based Lable Assignment) strategy, which can effectively improve model's performance on small objects. Experimental results show that our method can bring significant improvement. We achieved the 4thplace in the final leader board. Jiarui Hu 0001, Fei Shen 0004, Dian He, Qingyu Xian |
IGARSS | 3 |
| 2023 | A Rubust Method for Roof Extraction and Height EstimationabstractIn this article, we introduce the solution we used in the 2023 IEEE GRSS Data Fusion Contest Track 2. The task demands a roof type classification method and a building height estimation method. For roof type classification, our experiments are based on Swin transformer, combined with DoubleHead module and RFLA strategy, which can effectively improve model’s performance on small objects. For building height estimation, our experiments are based on SegFormer. Our experiments use a part of train set as validation set. Jiarui Hu 0001, Fei Shen 0004, Dian He, Qingyu Xian |
IGARSS | 3 |
| 2023 | Pedestrian-specific Bipartite-aware Similarity Learning for Text-based Person RetrievalabstractText-based person retrieval is a challenging task that aims to search pedestrian images with the same identity according to language descriptions. Current methods usually indiscriminately measure the similarity between text and image by matching global visual-textual features and matched local region-word features. However, these methods underestimate the key cue role of mismatched region-word pairs and ignore the problem of low similarity between matched region-word pairs. To alleviate these issues, we propose a novel Pedestrian-specific Bipartite-aware Similarity Learning (PBSL) framework that efficiently reveals the plausible and credible levels of contribution of pedestrian-specific mismatched and matched region-word pairs towards overall similarity. Specifically, to focus on mismatched region-word pairs, we first develop a new co-interactive attention that utilizes cross-modal information to guide the extraction of pedestrian-specific information in a single modality. We then design a negative similarity regularization mechanism to use the negative similarity score as a bias to correct the overall similarity. Additionally, to enhance the contribution of matched region-word pairs, we introduce graph networks to aggregate and propagate local information of pedestrian-specific, using overall visual-textual similarity to evaluate locally matched region-word pairs for weight refinement. Finally, extensive experiments are conducted on the CUHK-PEDES, ICFG-PEDES, and RSTPReid datasets to demonstrate the competitive performance of the proposed PBSL in the text-based person retrieval task. Fei Shen 0004, Xiangbo Shu, Xiaoyu Du 0002, Jinhui Tang 0001 |
ACM Multimedia | 1 |
| 2023 | A Novel Cross Frequency-Domain Interaction Learning for Aerial Oriented Object Detection
Weijie Weng, Weiming Lin, Junchi Ren, Fei Shen 0004 |
PRCV (4) | 5 |
| 2023 | GiT: Graph Interactive Transformer for Vehicle Re-IdentificationabstractTransformers are more and more popular in computer vision, which treat an image as a sequence of patches and learn robust global features from the sequence. However, pure transformers are not entirely suitable for vehicle re-identification because vehicle re-identification requires both robust global features and discriminative local features. For that, a graph interactive transformer (GiT) is proposed in this paper. In the macro view, a list of GiT blocks are stacked to build a vehicle re-identification model, in where graphs are to extract discriminative local features within patches and transformers are to extract robust global features among patches. In the micro view, graphs and transformers are in an interactive status, bringing effective cooperation between local and global features. Specifically, one current graph is embedded after the former level's graph and transformer, while the current transform is embedded after the current graph and the former level's transformer. In addition to the interaction between graphs and transforms, the graph is a newly-designed local correction graph, which learns discriminative local features within a patch by exploring nodes' relationships. Extensive experiments on three large-scale vehicle re-identification datasets demonstrate that our GiT method is superior to state-of-the-art vehicle re-identification approaches. Fei Shen 0004, Jianqing Zhu, Xiaobin Zhu 0001, Huanqiang Zeng |
IEEE Trans. Image Process. | 1 |
| 2022 | Enhancing Part Features via Contrastive Attention Module for Vehicle Re-identificationabstractVehicle re-identification’s methods usually exploit the spatial uniform partition strategy via dividing deep feature maps into several parts. Then each of them is further independently processed by the multi-network branch to obtain refined part features. However, the cooperation among those part features is underestimated. This paper proposes a contrastive attention module (CAM) to assess one part feature’s importance based on all parts. Practical cooperation among part features is derived by re-weighting the part feature. Furthermore, a flexible CAM network (CAMNet) compatible with contrastive attention module is proposed to enhance part features for vehicle re-identification. Extensive experiments show that the proposed CAMNet method outperforms many state-of-the-art vehicle re-identification approaches. Manyu Li, Mengwan Wei, Fei Shen 0004 |
ICIP | 4 |
| 2022 | HSGM: A Hierarchical Similarity Graph Module for Object Re-IdentificationabstractExisting object re-identification methods usually utilize backbone networks developed based on classification tasks to obtain the final object features. However, these backbone networks lack a unique mechanism to explore discriminative feature representation and handle rich scale changes. For that, a novel hierarchical similarity graph module (HSGM) is proposed to relieve the conflict of backbone networks and mine the discriminative features. Specifically, the proposed HSGM constructs a rich hierarchical graph to explore the pairwise relationships among global-local and local-local. Then, in each hierarchical graph, the HSGM regards local features extracted from different locations as nodes and utilizes the similarity scores between nodes to construct a similarity graph. During the HSGM's propagation, a learnable parameter is reweighted at each spatial position to optimize the correlation between adjacent nodes. Besides, the HSGM can be readily inserted into backbone networks at any depth to improve object discrimination. Extensive experiments on two large-scale object datasets (i.e., VeRi776 and Market-1501) demonstrate that the proposed HSGM is superior to state-of-the-art object re-identification approaches. Fei Shen 0004, Xiaoxiao Peng, Lisheng Wang, Xingmeng Hao, Mei Shu |
ICME | 1 |
| 2022 | A Novel Multi-Frequency Coordinated Module for SAR Ship DetectionabstractSynthetic aperture radar (SAR) ships have rich multi-frequency information, however, existing SAR ship detection methods mostly only consider high-frequency information, ignoring other frequency features and structured relationships. For that, a novel plug-and-play Multi-Frequency Coordinated (MFC) module is developed for SAR ship detection. Specifically, the proposed MFC consists of the two key submodules, i.e., Multi-Frequency Aggregate (MFA) and Frequency Response (FR). First, MFA is used to refine the different frequency feature maps along the channel dimension and Discrete Cosine Transform (DCT) bases to solve the problem of dense multi-target SAR ship detection. Then, FR is introduced to select the one with a better response from multi-frequency and boost the significant features of ship targets and suppress interference of surroundings. Lastly, we develop a YOLOv5s-MFC by embedding the MFC for ship detection. Extensive experiments on three large-scale ship datasets (SSDD, HRSID, and LS-SSDD-v1.0) demonstrate that the proposed YOLOv5s-MFC is superior to state-of-the-art ship detection approaches. Chenchen Qiao, Fei Shen 0004, Sixian Zhao |
ICTAI | 2 |
| 2022 | A sample-proxy dual triplet loss function for object re-identificationabstractAbstract Object re‐identification, such as vehicle re‐identification or pedestrian re‐identification, plays a significant role in intelligent video surveillance systems for public security. Due to viewpoint variations and appearance changes, both pedestrians and vehicles usually have complex intra‐class variations. However, most existing object re‐identification methods often use a sample‐level triplet loss function cooperating with a single‐proxy softmax loss function, which could not handle complex intra‐class variations well. In this paper, a sample‐proxy dual triplet (SPDT) loss function is proposed, which works with a multi‐proxy softmax (MPS) loss function. The MPS loss function is in charge of learning multiple proxies to represent a class. The SPDT loss function is responsible for enlarging inter‐class distances as well as shrinking intra‐class distances on both sample and proxy levels. Therefore, the method not only handles multi‐proxy intra‐class variations but also fully learns discrimination on samples and proxies. Experiments on two large datasets, that is, VeRi776 and DukeMTMC‐reID, demonstrate that the method is superior to state‐of‐the‐art object re‐identification approaches. Hanxiao Wu, Fei Shen 0004, Jianqing Zhu, Huanqiang Zeng, Xiaobin Zhu 0001, Zhen Lei 0001 |
IET Image Process. | 2 |
| 2022 | An Efficient Multiresolution Network for Vehicle ReidentificationabstractIn general, vehicle images have varying resolutions due to vehicles’ movements and different camera settings. However, most existing vehicle reidentification models are single-resolution deep networks trained with preuniformly resizing vehicle images, which underestimate adverse effects of varying resolutions and lead to unsatisfactory performance. A straightforward solution for dealing with varying resolutions is to train multiple vehicle reidentification models. Each model is independently trained with images of a specific resolution. However, this straightforward solution requires significant overhead and ignores intrinsic associations among different resolution images. For that, an efficient multiresolution network (EMRN) is proposed for vehicle reidentification in this article. First, EMRN embeds a newly designed multiresolution feature dimension uniform module (MR-FDUM) behind a traditional backbone network (i.e., ResNet-50). As a result, the whole model can extract fixed dimensional features from different resolution images so that it can be trained with one loss function of fixed dimensional parameters rather than training multiple models. Second, a multiresolution image randomly feeding strategy is designed to train EMRN, making each minibatch data of a random resolution during the training process. Consequently, EMRN can implicitly learn collaborative multiresolution features via only a unitary deep network. The experiments on three large-scale data sets, i.e., VeRi776, VehicleID, and VRIC, demonstrate that EMRN is superior to state-of-the-art vehicle reidentification methods. Fei Shen 0004, Jianqing Zhu, Xiaobin Zhu 0001, Jingchang Huang, Huanqiang Zeng, Zhen Lei 0001, Canhui Cai |
IEEE Internet Things J. | 1 |
| 2022 | Exploring Spatial Significance via Hybrid Pyramidal Graph Network for Vehicle Re-IdentificationabstractExisting vehicle re-identification methods commonly use spatial pooling operations to aggregate feature maps extracted via off-the-shelf backbone networks, such as visual geometry group network (VGGNet), Google network (GoogLeNet) and residual network (ResNet). They ignore exploring the spatial significance of feature maps, eventually degrading the vehicle re-identification performance. In this paper, firstly, an innovative spatial graph network (SGN) is proposed to elaborately explore the spatial significance of feature maps. The SGN stacks multiple spatial graphs (SGs). Each SG assigns feature map’s elements as nodes and utilizes spatial neighborhood relationships to determine edges among nodes. During the SGN’s propagation, each node and its spatial neighbors on an SG are aggregated to the next SG. On the next SG, each aggregated node is re-weighted with a learnable parameter to find the significance at the corresponding location. Secondly, a novel pyramidal graph network (PGN) is designed to comprehensively explore the spatial significance of feature maps at multiple scales. The PGN organizes multiple SGNs in a pyramidal manner and makes each SGN handles feature maps of a specific scale. Finally, a hybrid pyramidal graph network (HPGN) is developed by embedding the PGN behind a ResNet-50 based backbone network. Extensive experiments on three large scale vehicle databases (i.e., VeRi776, VehicleID, and VeRi-Wild) demonstrate that the proposed HPGN is superior to state-of-the-art vehicle re-identification approaches in terms of accuracy, parameter cost, and computation cost. In addition, experiments show that the proposed PGN is universal to various backbone networks. Fei Shen 0004, Jianqing Zhu, Xiaobin Zhu 0001, Jingchang Huang |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2021 | Object Re-identification Using Teacher-Like and Light Students
Hanxiao Wu, Fei Shen 0004, Jianqing Zhu, Huanqiang Zeng |
BMVC | 3 |