Haoran Duan 0001

dblp:82/4334-1 · DBLP profile ↗
← Back
53ranked-venue papers
6as first author
52since 2021 · last 2026
0000-0001-9956-7020ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 29 · 5 first-author · 29 since 2021Graphics, computer vision, multimedia, augmented reality and games · 21 · 20 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 8 since 2021Systems, architecture and hardware · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Debiasing Diffusion Priors via 3D Attention for Consistent Gaussian Splatting
abstract
Versatile 3D tasks (e.g., generation or editing) distilling Text-to-Image (T2I) diffusion models have attracted significant research interest for not relying on extensive 3D training data. However, T2I models exhibit limitations resulting from prior view bias, which produces conflicting appearances between different views of an object. This bias causes subject-words to preferentially activate prior view features during cross-attention (CA) computation, regardless of the target view condition. To overcome this limitation, we conduct a comprehensive mathematical analysis to reveal the root cause of the prior view bias in T2I models. Moreover, we find different UNet-Layers show different effects of prior view in CA. Therefore, we propose a novel framework, TD-Attn, which addresses multi-view inconsistency via two key components: (1) the 3D-Aware Attention Guidance Module 3D-AAG constructs a view-consistent 3D attention Gaussian for subject-words to enforce spatial consistency across attention-focused regions, thereby compensating for the limited spatial information in 2D individual view CA maps; (2) the Hierarchical Attention Modulation Module (HAM) utilizes a semantic guidance tree to direct the Semantic Response Profiler (SRP) in localizing and modulating CA layers that are highly responsive to view conditions, where the enhanced CA maps further support the construction of more consistent 3D attention Gaussians. Notably, HAM facilitates semantic-specific interventions, enabling controllable and precise 3D editing. Extensive experiments firmly establish that TD-Attn has the potential to serve as a transformative, universal plugin, significantly enhancing multi-view consistency across a wide range of 3D tasks.
Shilong Jin, Haoran Duan 0001, Litao Hua, Yuan Zhou 0023
AAAI2
2026 vMFCoOp: Towards Equilibrium on a Unified Hyperspherical Manifold for Prompting Biomedical VLMs
abstract
Recent advances in context optimization (CoOp) guided by large language model (LLM)–distilled medical semantic priors offer a scalable alternative to manual prompt engineering and full fine-tuning for adapting biomedical CLIP-based vision-language models (VLMs). However, prompt learning in this context is challenged by semantic misalignment between LLMs and CLIP variants due to divergent training corpora and model architectures; it further lacks scalability across continuously evolving families of foundation models. More critically, pairwise multimodal alignment via conventional Euclidean-space optimization lacks the capacity to model unified representations or apply localized geometric constraints, which tends to amplify modality gaps in complex biomedical imaging and destabilize few-shot adaptation. To address these challenges, we propose vMFCoOp, a framework that inversely estimates von Mises–Fisher (vMF) distributions on a shared Hyperspherical Manifold, aligning semantic biases between arbitrary LLMs and CLIP backbones via Unified Semantic Anchors to achieve robust biomedical prompting and superior few-shot classification. Grounded in three complementary constraints, vMFCoOp demonstrates consistent improvements across 14 medical datasets, 12 medical imaging modalities, and 13 anatomical regions, outperforming state-of-the-art methods in accuracy, generalization, and clinical applicability.
Minye Shao, Sihan Guo, Xinrun Li, Xingyu Miao, Haoran Duan 0001, Yang Long 0001
AAAI5
2026 ReaSon: Reinforced Causal Search with Information Bottleneck for Video Understanding
abstract
Keyframe selection has become essential for video understanding with vision-language models (VLMs) due to limited input tokens and the temporal sparsity of relevant information across video frames. Video understanding often relies on effective keyframes that are not only informative but also causally decisive. To this end, we propose Reinforced Causal Search with Information Bottleneck (ReaSon), a framework that formulates keyframe selection as an optimization problem with the help of a novel Causal Information Bottleneck (CIB), which explicitly defines keyframes as those satisfying both predictive sufficiency and causal necessity. Specifically, ReaSon employs a learnable policy network to select keyframes from a visually relevant pool of candidate frames to capture predictive sufficiency, and then assesses causal necessity via counterfactual interventions. Finally, a composite reward aligned with the CIB principle is designed to guide the selection policy through reinforcement learning. Extensive experiments on NExT-QA, EgoSchema, and Video-MME demonstrate that ReaSon consistently outperforms existing state-of-the-art methods under limited-frame settings, validating its effectiveness and generalization ability.
Yuan Zhou 0023, Litao Hua, Shilong Jin, Haoran Duan 0001
AAAI5
2026 Decoding visual neural representations by multimodal with dynamic balancing
Kaili Sun, Xingyu Miao, Bing Zhai, Haoran Duan 0001, Yang Long 0001
Expert Syst. Appl.4
2026 Cross-modal progressive modeling for neuro-visual representation learning
abstract
Neural decoding from scalp signals requires models that respect spatial, temporal, and spectral structure while leveraging strong visual priors. In this paper, we introduce CFT-NET for disentangled neural visual representation together with a progressive visual–semantic adaptation (PVSA) framework that aligns EEG embeddings to pretrained visual backbones under a contrastive objective followed by pairwise matching. CFT-NET integrates Frequency-Separated Weights (FSW), Spatial-Context Aggregation (SCA), and Adaptive Temporal Filtering (ATF) to explicitly extract spectral, spatial, and temporal factors. PVSA consists of an instance-guided visual encoder and a visual-guided semantic decoder linked by cross attention, enabling fine-grained neuro–image interaction. On THINGS-EEG and THINGS-MEG dataset, the approach consistently outperforms state-of-the-art baselines in both subject-dependent and subject-independent zero-shot classification. By aligning model architecture with visual cognition principles and coupling it to strong visual priors, our methods narrows the gap between neural activity and visual cognition, provides novel cross-modal neural decoding method that achieves competitive performance against recent state-of-the-art baselines.
Jiyao Pu, Kaili Sun, Zeyu Fu, Haoran Duan 0001, Yang Long 0001
Neurocomputing5
2026 Semi-supervised crowd counting from unlabeled data
Haoran Duan 0001, Yawen Huang, Yang Long 0001, Xian Wu 0001, Feiyue Huang, Shaoxin Li 0001
Pattern Recognit.1
2026 ConRF: Zero-shot stylization of 3D scenes with conditioned radiation fields
abstract
• We propose a novel method that leverages CLIP for zero-shot 3D scene artistic style transfer by a single condition (i.e. image or text). • We introduce a mapping network to alleviate the ambiguity in CLIP features related to style. • We present a 3D selection volume that allows for localized style manipulation within 3D scenes, expanding the possibilities in scene stylization and manipulation. Most of the existing works on arbitrary 3D NeRF style transfer required retraining on each single style condition. This work aims to achieve zero-shot controlled stylization in 3D scenes utilizing text or visual input as conditioning factors. We introduce ConRF, a novel method of zero-shot stylization. Specifically, due to the ambiguity of CLIP features, we employ a conversion process that maps the CLIP feature space to the style space of a pre-trained VGG network and then refine the CLIP multi-modal knowledge into a style transfer neural radiation field. Additionally, we use a 3D volumetric representation to perform local style transfer. By combining these operations, ConRF offers the capability to utilize either text or images as references, resulting in the generation of sequences with novel views enhanced by global or local stylization. Our experiment demonstrates that ConRF outperforms other existing methods for 3D scene and single-text stylization in terms of visual quality. Code is available: https://xingy038.github.io/ConRF/ .
Xingyu Miao, Yang Bai 0011, Haoran Duan 0001, Fan Wan, Yawen Huang, Yang Long 0001, Yefeng Zheng 0001
Pattern Recognit.3
2026 M3GStyler: Enhancing consistency across multi-view in multi-modality 3D Gaussian style transfer
abstract
• Unified text- and image-guided 3D style transfer with consistent views across a scene, while preserving fine details. • Flow-based feature matching aligns text and image cues for faithful style. • Parallel frequency branches balance structure (low frequency) and detail (high frequency). • Cross-dimension and texture losses improve color and texture fidelity.
Xueqi Qiu, Xingyu Miao, Haoran Duan 0001, Minye Shao, Bing Zhai, Gongjie Zhang, Jingjing Deng 0001, Yang Long 0001
Pattern Recognit.3
2026 Motion2Motion: Learning human pose refining in videos without ground truth label
Zhenyu Wen, Zhen Hong, Haoran Duan 0001, Xiaoqin Zhang 0002
Pattern Recognit.6
2026 Multimodal Feature Interaction and High-Quality Pseudolabel Generation With Self-Training for Cognitive State Detection
abstract
Cognitive state detection holds significant research value in the field of human–computer interaction and neural engineering. However, existing works are insufficient in modeling the temporal dynamics of multimodal physiological signals, which leads to heterogeneous distribution differences in cross-modal feature interactions. In addition, domain shift issues under cross-subject and few-sample conditions restrict the model generalization performance. To cope with these problems, this work proposes a cognitive state detection framework that integrates Transformer-based multimodal feature interaction and self-training of pseudolabel optimization. First, the multihead attention mechanism is introduced to model the temporal evolution patterns across modalities, dynamically harmonizing cross-modal contributions to extract cognitive state-related shared features. Then, a dual-model cross-validation strategy is designed to filter high-quality pseudolabeled samples from the target domain for subsequent self-training, effectively avoiding the dependency on auxiliary modules in domain adaptation. Finally, Extensive experiments show that the proposed work significantly improves the recognition accuracy, and the designed pseudolabel optimization mechanism can be transferred to related tasks without increasing model complexity.
Kevin W. Tong, Xuefeng Men, Haoran Duan 0001, Shaojun Cai, Changyu Li, Ping Li 0044, Guangyu Zhu 0001, Qi Wu 0003, Limin Zhu 0001
IEEE Trans. Ind. Informatics4
2026 Rethinking Multi-Focus Image Fusion: An Input Space Optimization View
abstract
Multi-focus image fusion (MFIF) addresses the challenge of partial focus by integrating multiple source images taken at different focal depths. Unlike most existing methods that rely on complex loss functions or large-scale synthetic datasets, this study approaches MFIF from a novel perspective: optimizing the input space. The core idea is to construct a high-quality MFIF input space in a cost-effective manner by using intermediate features from well-trained, non-MFIF networks. To this end, we propose a cascaded framework comprising two feature extractors, a Feature Distillation and Fusion Module (FDFM), and a focus segmentation network Y ${}^{U}$ Net. Based on our observation that discrepancy and edge features are essential for MFIF, we select a image deblurring network and a salient object detection network as feature extractors. To transform these extracted features into an MFIF-suitable input space, we propose FDFM as a training-free feature adapter. To make FDFM compatible with high-dimensional feature maps, we extend the manifold theory from the edge-preserving field and design a novel isometric domain transformation. Extensive experiments on six benchmark datasets show that 1) our model consistently outperforms 13 state-of-the-art methods in both qualitative and quantitative evaluations, and 2) the constructed input space can directly enhance the performance of many MFIF models without additional requirements.
Zeyu Wang 0009, Haoran Duan 0001, Yang Long 0001, Ling Shao 0001
IEEE Trans. Image Process.3
2026 ConsDreamer: Advancing Multi-View Consistency for Zero-Shot Text-to-3D Generation
abstract
Recent advances in zero-shot text-to-3D generation have revolutionised 3D content creation by enabling direct synthesis from textual descriptions. While state-of-the-art methods leverage 3D Gaussian Splatting with score distillation to enhance multi-view rendering through pre-trained text-to-image (T2I) models, they suffer from inherent prior view biases in T2I Models. These biases lead to inconsistent 3D generation, particularly manifesting as the multi-face Janus problem, where objects exhibit conflicting features across views. To address this fundamental challenge, we propose ConsDreamer, a novel method that mitigates view bias by refining both the conditional and unconditional terms in the score distillation process: (1) a View Disentanglement Module (VDM) that eliminates viewpoint biases in conditional prompts by decoupling irrelevant view components and injecting precise view control; and (2) a similarity-based partial order loss that enforces geometric consistency in the unconditional term by aligning cosine similarities with azimuth relationships. Extensive experiments demonstrate that ConsDreamer can be seamlessly integrated into various 3D representations and score distillation paradigms, effectively mitigating the multi-face Janus problem.
Yuan Zhou 0023, Shilong Jin, Litao Hua, Wanjun Lv, Haoran Duan 0001, Jungong Han
IEEE Trans. Image Process.5
2026 From Gaze to Insight: Bridging Human Visual Attention and Vision Language Model Explanation for Weakly-Supervised Medical Image Segmentation
abstract
Medical image segmentation remains challenging due to the high cost of pixel-level annotations for training. In the context of weak supervision, clinician gaze data captures regions of diagnostic interest; however, its sparsity limits its use for segmentation. In contrast, vision-language models (VLMs) provide semantic context through textual descriptions but lack the explanation precision required. Recognizing that neither source alone suffices, we propose a teacher-student framework that integrates both gaze and language supervision, leveraging their complementary strengths. Our key insight is that gaze data indicates "where" clinicians focus during diagnosis, while VLMs explain "why" those regions are significant. To implement this, the teacher model first learns from gaze points enhanced by VLM-generated descriptions of lesion morphology, establishing a foundation for guiding the student model. The teacher then directs the student through three strategies: 1) Multi-scale feature alignment to fuse visual cues with textual semantics; 2) Confidence-weighted consistency constraints to focus on reliable predictions; 3) Adaptive masking to limit error propagation in uncertain areas. Experiments on the Kvasir-SEG, NCI-ISBI, and ISIC datasets show that our method achieves Dice scores of 80.78%, 80.53%, and 84.22%, respectively-improving 3-5% over gaze baselines without increasing the annotation burden. By preserving correlations among predictions, gaze data, and lesion descriptions, our framework also maintains clinical interpretability. This work illustrates how integrating human visual attention with AI-generated semantic context can effectively overcome the limitations of individual weak supervision signals, thereby advancing the development of deployable, annotation-efficient medical AI systems. Code is available at: https://github.com/jingkunchen/FGI.
Jingkun Chen, Haoran Duan 0001, Xiao Zhang 0028, Boyan Gao, Vicente Grau, Jungong Han
IEEE Trans. Medical Imaging2
2026 SNH-SLAM: Implicit Dense SLAM Based on Scalable Neural-Hash Representation
abstract
We present SNH-SLAM, a novel expandable dense neural simultaneous localization and mapping (SLAM) method that constructs a neural field in real-time based on run-time observation. To reach this challenging goal without any scene prior, we utilize instant depth supervision to drive the extension of planar convex hulls, where a single hash table maintains multi-level feature units embedded in the planar convex hulls. This design facilitates high-fidelity, hole-free, and low-memory map reconstruction while adding only a tiny time burden to the training process. Our approach performs mapping by minimizing both RGBD-based re-rendering loss and Truncated Signed Distance Field (TSDF) loss. In addition, for camera tracking, our optimization strategy allows SNH-SLAM to converge faster on the pose estimation and maintain robustness. We evaluate our method on common benchmarks and compare it with existing dense neural RGB-D SLAM methods. The evaluation results show the competitiveness of the SNH-SLAM in tracking accuracy, reconstruction quality, memory usage, and frame processing speed. Project page:https://xiaoshumiao123.github.io.
Zhenyu Wen, Zhanshuo Dong, Haoran Duan 0001, Tianrun Chen, Zhen Hong
IEEE Trans. Multim.4
2025 LAGD: Local Topological-Alignment and Global Semantic-Deconstruction for Incremental 3D Semantic Segmentation
abstract
Numerous deep learning-based works focusing on 3D semantic segmentation have been proposed and have achieved impressive performance. However, due to the catastrophic forgetting, existing methods will degrade dramatically in a real-world scenario where new 3D semantic categories are arriving continually. Straightforwardly applying typical class-incremental learning methods on 3D data even aggravates forgetting due to the irregular and noisy geometric structure. Aiming to address this realistic challenge, from the perspective of capturing local topological characteristics and mitigating global semantic shift, we propose a unified framework named Local topological Alignment and Global semantic Deconstruction (LAGD) to incrementally learn semantic knowledge of novel 3D categories while maintaining performance on previously learned knowledge. Specifically, we develop a novel Interaction Topological-aware Alignment (ITA) to maintain the learned knowledge efficiently by capturing the local geometric characteristics with interacted adjacent state-specific knowledge. Besides, to mitigate the forgetting caused by the global semantic shift, we deconstruct the logits into positive and negative parts which are distilled separately, achieving an elaborate distillation process in terms of Semantic-knowledge Deconstruction Distillation (SDD). With the cooperation of ITA and SDD, LAGD achieves a sota performance, especially in the long-term incremental learning scenario. Extensive experimental results illustrate the superiority of our proposed LAGD.
Haoran Duan 0001, Rui Sun 0010, Tejal Shah, Rajiv Ranjan 0001, Bo Wei 0003
AAAI2
2025 Multi-Modal Medical Image Fusion via 3D Manifold Fitting and Dual-Domain Cross-Attention
abstract
Medical image fusion (MIF) aims to extract complementary features from multi-modal source images and fuse them into a single image to assist in clinical diagnostics. Despite its importance, MIF faces two primary challenges: the lack of tailored paradigms for CMSF extraction and insufficient dual exploration of multi-modality and multi-frequency domains. To address these challenges, we propose a novel MIF model in this study. From the perspective of image manifolds, we reformulate CMSF extraction as a 3D manifold fitting problem and introduce a paradigm that uses mathematical fitting methods to generate CMSF. This approach achieves accurate feature extraction without the need for carefully designed loss functions as constraints, significantly reducing the number of parameters. Additionally, we introduce Cross-Modality Co-Frequency (CM-CoF) and Cross-Frequency Co-Modality (CF-CoM) attention modules, which explore implicit relationships between modalities and frequency domains. Experimental results demonstrate that the proposed model outperforms many state-of-the-art MIF algorithms.
Zeyu Wang 0009, Haiyu Song 0002, Haoran Duan 0001
ICASSP5
2025 Towards Scalable Spatial Intelligence Via 2D-To-3D Data Lifting
abstract
Spatial intelligence is emerging as a transformative frontier in AI, yet it remains constrained by the scarcity of largescale 3D datasets. Unlike the abundant 2D imagery, acquiring 3D data typically requires specialized sensors and laborious annotation. In this work, we present a scalable pipeline that converts single-view images into comprehensive, scale- and appearance-realistic 3D representations - including point clouds, camera poses, depth maps, and pseudo-RGBD - via integrated depth estimation, camera calibration, and scale calibration. Our method bridges the gap between the vast repository of imagery and the increasing demand for spatial scene understanding. By automatically generating authentic, scale-aware 3D data from images, we significantly reduce data collection costs and open new avenues for advancing spatial intelligence. We release two generated spatial datasets, i.e., COCO-3D and Objects365-v2-3D, and demonstrate through extensive experiments that our generated data can benefit various 3D tasks, ranging from fundamental perception to MLLMbased reasoning. These results validate our pipeline as an effective solution for developing AI systems capable of perceiving, understanding, and interacting with physical environments.
Xingyu Miao, Haoran Duan 0001, Quanhao Qian, Jiuniu Wang, Yang Long 0001, Ling Shao 0001, Deli Zhao, Gongjie Zhang
ICCV2
2025 Highlight What You Want: Weakly-Supervised Instance-Level Controllable Infrared-Visible Image Fusion
Zeyu Wang 0009, Jizheng Zhang, Haiyu Song 0002, Mingyu Ge, Haoran Duan 0001
ICCV6
2025 Rehearsal-free Federated Domain-incremental Learning
abstract
We introduce a rehearsal-free federated domain incremental learning framework, RefFiL, based on a global prompt-sharing paradigm to alleviate catastrophic forgetting challenges in federated domain-incremental learning, where unseen domains are continually learned. Typical methods for mitigating forgetting, such as the use of additional datasets and the retention of private data from earlier tasks, are not viable in federated learning (FL) due to devices’ limited resources. Our method, RefFiL, addresses this by learning domain-invariant knowledge and incorporating various domain-specific prompts from the domains represented by different FL participants. A key feature of RefFiL is the generation of local finegrained prompts by our domain adaptive prompt generator, which effectively learns from local domain knowledge while maintaining distinctive boundaries on a global scale. We also introduce a domain-specific prompt contrastive learning loss that differentiates between locally generated prompts and those from other domains, enhancing RefFiL’s precision and effectiveness. Compared to existing methods, RefFiL significantly alleviates catastrophic forgetting without requiring extra memory space, making it ideal for privacy-sensitive and resource-constrained devices.
Rui Sun 0010, Haoran Duan 0001, Jiahua Dong 0001, Varun Ojha 0001, Tejal Shah, Rajiv Ranjan 0001
ICDCS2
2025 Rethinking Score Distilling Sampling for 3D Editing and Generation
abstract
Score Distillation Sampling (SDS) has emerged as a prominent method for text-to-3D generation by leveraging the strengths of 2D diffusion models. However, SDS is limited to generation tasks and lacks the capability to edit existing 3D assets. Conversely, variants of SDS that introduce editing capabilities often can not generate new 3D assets effectively. In this work, we observe that the processes of generation and editing within SDS and its variants have unified underlying gradient terms. Building on this insight, we propose Unified Distillation Sampling (UDS), a method that seamlessly integrates both the generation and editing of 3D assets. Essentially, UDS refines the gradient terms used in vanilla SDS methods, unifying them to support both tasks. Extensive experiments demonstrate that UDS not only outperforms baseline methods in generating 3D assets with richer details but also excels in editing tasks, thereby bridging the gap between 3D generation and editing.
Xingyu Miao, Haoran Duan 0001, Yang Long 0001, Jungong Han
ICML2
2025 TRACE: Temporally Reliable Anatomically-Conditioned 3D CT Generation with Enhanced Efficiency
Minye Shao, Xingyu Miao, Haoran Duan 0001, Zeyu Wang 0009, Jingkun Chen, Yawen Huang, Xian Wu 0001, Jingjing Deng 0001, Yang Long 0001, Yefeng Zheng 0001
MICCAI (4)3
2025 Parameter Efficient Fine-Tuning for Multi-modal Generative Vision Models with Möbius-Inspired Transformation
abstract
Abstract The rapid development of multimodal generative vision models has drawn scientific curiosity. Notable advancements, such as OpenAI’s ChatGPT and Stable Diffusion, demonstrate the potential of combining multimodal data for generative content. Nonetheless, customising these models to specific domains or tasks is challenging due to computational costs and data requirements. Conventional fine-tuning methods take redundant processing resources, motivating the development of parameter-efficient fine-tuning technologies such as adapter module, low-rank factorization and orthogonal fine-tuning. These solutions selectively change a subset of model parameters, reducing learning needs while maintaining high-quality results. Orthogonal fine-tuning, regarded as a reliable technique, preserves semantic linkages in weight space but has limitations in its expressive powers. To better overcome these constraints, we provide a simple but innovative and effective transformation method inspired by Möbius geometry, which replaces conventional orthogonal transformations in parameter-efficient fine-tuning. This strategy improved fine-tuning’s adaptability and expressiveness, allowing it to capture more data patterns. Our strategy, which is supported by theoretical understanding and empirical validation, outperforms existing approaches, demonstrating competitive improvements in generation quality for key generative tasks.
Haoran Duan 0001, Bing Zhai, Tejal Shah, Jungong Han, Rajiv Ranjan 0001
Int. J. Comput. Vis.1
2025 Correction: Parameter Efficient Fine-Tuning for Multi-modal Generative Vision Models with Möbius-Inspired Transformation
Haoran Duan 0001, Bing Zhai, Tejal Shah, Jungong Han, Rajiv Ranjan 0001
Int. J. Comput. Vis.1
2025 DFAN++: Enhanced triple-branch network for generalized zero-shot image classification
Yuan Zhou 0023, Haoran Duan 0001, Yang Long 0001
Neurocomputing4
2025 FMDConv: Fast multi-attention dynamic convolution via speed-accuracy trade-off
abstract
Spatial convolution is fundamental in constructing deep Convolutional Neural Networks (CNNs) for visual recognition. While dynamic convolution enhances model accuracy by adaptively combining static kernels, it incurs significant computational overhead, limiting its deployment in resource-constrained environments such as federated edge computing. To address this, we propose Fast Multi-Attention Dynamic Convolution (FMDConv), which integrates input attention, temperature-degraded kernel attention, and output attention to optimize the speed-accuracy trade-off. FMDConv achieves a better balance between accuracy and efficiency by selectively enhancing feature extraction with lower complexity. Furthermore, we introduce two novel quantitative metrics, the Inverse Efficiency Score and Rate-Correct Score, to systematically evaluate this trade-off. Extensive experiments on CIFAR-10, CIFAR-100, and ImageNet demonstrate that FMDConv reduces the computational cost by up to 49.8% on ResNet-18 and 42.2% on ResNet-50 compared to prior multi-attention dynamic convolution methods while maintaining competitive accuracy. These advantages make FMDConv highly suitable for real-world, resource-constrained applications. This figure presents an overview of the proposed Fast Multi-Attention Dynamic Convolution (FMDConv) framework, which integrates Input Attention, Temperature-Degraded Kernel Attention, and Output Attention to optimize the speed-accuracy trade-off in convolutional neural networks. The diagram illustrates how these mechanisms enhance feature selection at different stages, significantly reducing computational cost while maintaining competitive accuracy, making FMDConv suitable for resource-constrained applications. • Introduces IES and RCS to quantify speed-accuracy trade-off in CNNs. • Evaluates channel, kernel, and filter attention for effective structures. • Develops a new CNN with kernel attention to enhance efficiency and accuracy. • Demonstrates FMDConv’s advantages via standard benchmark testing.
Fan Wan, Haoran Duan 0001, Kevin W. Tong, Jingjing Deng 0001, Yang Long 0001
Knowl. Based Syst.3
2025 Laser: Efficient Language-Guided Segmentation in Neural Radiance Fields
abstract
In this work, we propose a method that leverages CLIP feature distillation, achieving efficient 3D segmentation through language guidance. Unlike previous methods that rely on multi-scale CLIP features and are limited by processing speed and storage requirements, our approach aims to streamline the workflow by directly and effectively distilling dense CLIP features, thereby achieving precise segmentation of 3D scenes using text. To achieve this, we introduce an adapter module and mitigate the noise issue in the dense CLIP feature distillation process through a self-cross-training strategy. Moreover, to enhance the accuracy of segmentation edges, this work presents a low-rank transient query attention mechanism. To ensure the consistency of segmentation for similar colors under different viewpoints, we convert the segmentation task into a classification task through label volume, which significantly improves the consistency of segmentation in color-similar areas. We also propose a simplified text augmentation strategy to alleviate the issue of ambiguity in the correspondence between CLIP features and text. Extensive experimental results show that our method surpasses current state-of-the-art technologies in both training speed and performance.
Xingyu Miao, Haoran Duan 0001, Yang Bai 0011, Tejal Shah, Jun Song 0003, Yang Long 0001, Rajiv Ranjan 0001, Ling Shao 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2025 Attention-driven acoustic properties learning for underwater target ranging
Xiaohui Chu, Hantao Zhou, Yan Zhang 0109, Yachao Zhang 0001, Runze Hu, Haoran Duan 0001, Yawen Huang, Yefeng Zheng 0001, Rongrong Ji
Pattern Recognit.6
2025 SP-SLAM: Neural Real-Time Dense SLAM With Scene Priors
abstract
Neural implicit representations have recently shown promising progress in dense Simultaneous Localization And Mapping (SLAM). However, existing works have shortcomings in terms of reconstruction quality and real-time performance, mainly due to inflexible scene representation strategy without leveraging any prior information. In this paper, we introduce SP-SLAM, a novel neural RGB-D SLAM system that performs tracking and mapping in real-time. SP-SLAM computes depth images and establishes sparse voxel-encoded scene priors near the surface reconstruction. Simultaneously, we employ triplanes to store scene appearance information, striking a balance between achieving high-quality geometric texture mapping and minimizing memory consumption. Furthermore, in SP-SLAM, we introduce an effective optimization strategy for mapping, allowing the system to continuously optimize the poses of all historical input frames during runtime without increasing computational overhead. We conduct extensive evaluations on five benchmark datasets (Replica, ScanNet, TUM RGB-D, Synthetic RGB-D, 7-Scenes). The results demonstrate that, compared to existing methods, we achieve superior tracking accuracy and reconstruction quality, while running at a significantly faster speed.
Zhen Hong, Haoran Duan 0001, Yawen Huang, Zhenyu Wen, Xiang Wu 0012, Wei Xiang 0001, Yefeng Zheng 0001
IEEE Trans. Circuits Syst. Video Technol.3
2025 Union-Domain Knowledge Distillation for Underwater Acoustic Target Recognition
abstract
Underwater acoustic target recognition (UATR) can be significantly empowered by advancements in deep learning (DL). However, the effectiveness of DL-based UATR methods is often constrained by the limited computing resources available on underwater platforms. Most of the existing knowledge distillation (KD) strategies try to build lightweight DL models, but these strategies rarely consider the acoustic properties of underwater environments, making them less efficient for UATR tasks. Thus, fully harnessing the potential of DL techniques while ensuring the model’s practicality, is one of the urgent problems to be solved in UATR research. In this work, we introduce the union-domain KD (UDKD) to establish an accurate and lightweight UATR model. UDKD integrates two KD strategies: dual-frequency band distillation (DBD) and cross-domain masked distillation (CMD). DBD improves the learning process for a simple student model by decoupling the knowledge of spectrograms into the local structural (i.e., line spectra) and global composition (i.e., propagation patterns) aspects. CMD reduces redundant information from the Fourier Transform process, enabling the student model to concentrate on essential signal elements and to learn underlying time–frequency distribution. Extensive experiments on two real-world oceanic datasets confirm the superior performance of UDKD compared to existing KD methods, i.e., achieving an accuracy of 94.81% ($\uparrow ~3.19$% versus 91.62%). Notably, UDKD showcases a 10.5% improvement in the prediction accuracy of the lightweight student model.
Xiaohui Chu, Haoran Duan 0001, Zhenyu Wen, Runze Hu, Wei Xiang 0001
IEEE Trans. Geosci. Remote. Sens.2
2025 DSleepNet: Disentanglement Learning for Personal Attribute-Agnostic Three-Stage Sleep Classification Using Wearable Sensing Data
abstract
Long-term non-invasive sleep stage monitoring is instrumental in comprehending the progression of sleep disorders, cardiovascular diseases, and the interplay between sleep, type 2 diabetes, and neurodegenerative diseases. However, the conventional deep learning approach is susceptible to personal attributes (PAs) such as age, Body Mass Index, and severity of sleep apnea existing in the training dataset, potentially hindering its generalisation capacity to unseen cohorts. This paper introduces DSleepNet, a novel approach that disentangles the feature space into PA-specific and PA-agnostic components using two probabilistic encoders. The PA-agnostic features, designed to remain unaffected by personal attributes, outperformed the baseline CNN, improving the mean F1 score by up to 8.7% (baseline: 60.3) and Cohen's Kappa by 4.7% (baseline: 55.5), especially in reducing the impact of sleep apnea. DSleepNet functions without the need for target cohort data during training. It operates without the need to acquire PA data during inference, nor does it require fine-tuning. A novel Independent Excitation mechanism is incorporated into the latent feature space to remove correlations between the two types of features. Comprehensive testing in various PA settings has demonstrated its efficacy in improving the model's robustness.
Bing Zhai, Haoran Duan 0001, Yu Guan 0001, Huy Phan, Wai Lok Woo
IEEE J. Biomed. Health Informatics2
2025 Rethinking Brain Tumor Segmentation From the Frequency Domain Perspective
abstract
Precise segmentation of brain tumors, particularly contrast-enhancing regions visible in post-contrast MRI (areas highlighted by contrast agent injection), is crucial for accurate clinical diagnosis and treatment planning but remains challenging. However, current methods exhibit notable performance degradation in segmenting these enhancing brain tumor areas, largely due to insufficient consideration of MRI-specific tumor features such as complex textures and directional variations. To address this, we propose the Harmonized Frequency Fusion Network (HFF-Net), which rethinks brain tumor segmentation from a frequency-domain perspective. To comprehensively characterize tumor regions, we develop a Frequency Domain Decomposition (FDD) module that separates MRI images into low-frequency components, capturing smooth tumor contours and high-frequency components, highlighting detailed textures and directional edges. To further enhance sensitivity to tumor boundaries, we introduce an Adaptive Laplacian Convolution (ALC) module that adaptively emphasizes critical high-frequency details using dynamically updated convolution kernels. To effectively fuse tumor features across multiple scales, we design a Frequency Domain Cross-Attention (FDCA) integrating semantic, positional, and slice-specific information. We further validate and interpret frequency-domain improvements through visualization, theoretical reasoning, and experimental analyses. Extensive experiments on four public datasets demonstrate that HFF-Net achieves an average relative improvement of 4.48% (ranging from 2.39% to 7.72%) in the mean Dice scores across the three major subregions, and an average relative improvement of 7.33% (ranging from 5.96% to 8.64%) in the segmentation of contrast-enhancing tumor regions, while maintaining favorable computational efficiency and clinical applicability. Our code is available at: https://github.com/VinyehShaw/HFF.
Minye Shao, Zeyu Wang 0009, Haoran Duan 0001, Yawen Huang, Bing Zhai, Shizheng Wang, Yang Long 0001, Yefeng Zheng 0001
IEEE Trans. Medical Imaging3
2025 A Semantic-Consistent Few-Shot Modulation Recognition Framework for IoT Applications
abstract
The rapid growth of the Internet of Things (IoT) has led to the widespread adoption of the IoT networks in numerous digital applications. To counter physical threats in these systems, automatic modulation classification (AMC) has emerged as an effective approach for identifying the modulation format of signals in noisy environments. However, identifying those threats can be particularly challenging due to the scarcity of labeled data, which is a common issue in various IoT applications, such as anomaly detection for unmanned aerial vehicles (UAVs) and intrusion detection in the IoT networks. Few-shot learning (FSL) offers a promising solution by enabling models to grasp the concepts of new classes using only a limited number of labeled samples. However, prevalent FSL techniques are primarily tailored for tasks in the computer vision domain and are not suitable for the wireless signal domain. Instead of designing a new FSL model, this work suggests a novel approach that enhances wireless signals to be more efficiently processed by the existing state-of-the-art (SOTA) FSL models. We present the semantic-consistent signal pretransformation (ScSP), a parameterized transformation architecture that ensures signals with identical semantics exhibit similar representations. ScSP is designed to integrate seamlessly with various SOTA FSL models for signal modulation recognition and supports commonly used deep learning backbones. Our evaluation indicates that ScSP boosts the performance of numerous SOTA FSL models, while preserving flexibility.
Jie Su 0001, Zhenyu Wen, Fangda Guo, Yiming Wu 0009, Zhen Hong, Haoran Duan 0001, Yawen Huang, Rajiv Ranjan 0001, Yefeng Zheng 0001
IEEE Trans. Neural Networks Learn. Syst.8
2025 UniHead: Unifying Multi-Perception for Detection Heads
abstract
The detection head constitutes a pivotal component within object detectors, tasked with executing both classification and localization functions. Regrettably, the commonly used parallel head often lacks omni perceptual capabilities, such as deformation perception (DP), global perception (GP), and cross-task perception (CTP). Despite numerous methods attempting to enhance these abilities from a single aspect, achieving a comprehensive and unified solution remains a significant challenge. In response to this challenge, we develop an innovative detection head, termed UniHead, to unify three perceptual abilities simultaneously. More precisely, our approach: 1) introduces DP, enabling the model to adaptively sample object features; 2) proposes a dual-axial aggregation transformer (DAT) to adeptly model long-range dependencies, thereby achieving GP; and 3) devises a cross-task interaction transformer (CIT) that facilitates interaction between the classification and localization branches, thus aligning the two tasks. As a plug-and-play method, the proposed UniHead can be conveniently integrated with existing detectors. Extensive experiments on the COCO dataset demonstrate that our UniHead can bring significant improvements to many detectors. For instance, the UniHead can obtain +2.7 AP gains in RetinaNet, +2.9 AP gains in FreeAnchor, and +2.1 AP gains in GFL. The code is available at https://github.com/zht8506/UniHead.
Hantao Zhou, Rui Yang 0040, Yachao Zhang 0001, Haoran Duan 0001, Yawen Huang, Runze Hu, Xiu Li 0001, Yefeng Zheng 0001
IEEE Trans. Neural Networks Learn. Syst.4
2024 Dual Variational Knowledge Attention for Class Incremental Vision Transformer
abstract
Class incremental learning (CIL) strives to emulate the human cognitive process of continuously learning and adapting to new tasks while retaining knowledge from past experiences. Despite significant advancements in this field, Transformer-based models have not fully leveraged the potential of attention mechanisms to balance the transferable knowledge between tokens and the associated information. This paper addresses this gap by using a dual variational knowledge attention (DVKA) mechanism within a Transformer-based encoder-decoder framework, tailored for CIL. DVKA mechanism aims to manage the information flow through the attention maps, ensuring a balanced representation of all classes, and mitigating the risk of information dilution as new classes are incrementally introduced. This method, leverage the information bottleneck and mutual information principle, selectively filters less relevant information, directing the model’s focus towards the most significant details for each class. The DVKA is designed with two distinct attentions: one focused on the feature level and the other on the token dimension. The feature-focused attention aims to purify the complex nature of various classification tasks, ensuring a comprehensive representation of both old and new tasks. The token-focused attention mechanism highlights specific tokens, facilitating local discrimination among disparate patches and fostering global coordination for a spectrum of task tokens. Our work is a major stride towards improving transformer models for class incremental learning, presenting a theoretical rationale and effective experimental results on three widely-used datasets.
Haoran Duan 0001, Rui Sun 0010, Varun Ojha 0001, Tejal Shah, Zhuoxu Huang, Zizhou Ouyang, Yawen Huang, Yang Long 0001, Rajiv Ranjan 0001
IJCNN1
2024 Wearable-based behaviour interpolation for semi-supervised human activity recognition
abstract
While traditional feature engineering for Human Activity Recognition (HAR) involves a trial-and-error process, deep learning has emerged as a preferred method for high-level representations of sensor-based human activities. However, most deep learning-based HAR requires a large amount of labelled data and extracting HAR features from unlabelled data for effective deep learning training remains challenging. We, therefore, introduce a deep semi-supervised HAR approach, MixHAR, which concurrently uses labelled and unlabelled activities. Our MixHAR employs a linear interpolation mechanism to blend labelled and unlabelled activities while addressing both inter- and intra-activity variability. A unique challenge identified is the activity-intrusion problem during mixing, for which we propose a mixing calibration mechanism to mitigate it in the feature embedding space. Additionally, we rigorously explored and evaluated the five conventional/popular deep semi-supervised technologies on HAR, acting as the benchmark of deep semi-supervised HAR. Our results demonstrate that MixHAR significantly improves performance, underscoring the potential of deep semi-supervised techniques in HAR.
Haoran Duan 0001, Varun Ojha 0001, Shizheng Wang, Yawen Huang, Yang Long 0001, Rajiv Ranjan 0001, Yefeng Zheng 0001
Inf. Sci.1
2024 CTNeRF: Cross-time Transformer for dynamic neural radiance field from monocular video
Xingyu Miao, Yang Bai 0011, Haoran Duan 0001, Fan Wan, Yawen Huang, Yang Long 0001, Yefeng Zheng 0001
Pattern Recognit.3
2024 DS-Depth: Dynamic and Static Depth Estimation via a Fusion Cost Volume
abstract
Self-supervised monocular depth estimation methods typically rely on the reprojection error to capture geometric relationships between successive frames in static environments. However, this assumption does not hold in dynamic objects in scenarios, leading to errors during the view synthesis stage, such as feature mismatch and occlusion, which can significantly reduce the accuracy of the generated depth maps. To address this problem, we propose a novel dynamic cost volume that exploits residual optical flow to describe moving objects, improving incorrectly occluded regions in static cost volumes used in previous work. Nevertheless, the dynamic cost volume inevitably generates extra occlusions and noise, thus we alleviate this by designing a fusion module that makes static and dynamic cost volumes compensate for each other. In other words, occlusion from the static volume is refined by the dynamic volume, and incorrect information from the dynamic volume is eliminated by the static volume. Furthermore, we propose a pyramid distillation loss to reduce photometric error inaccuracy at low resolutions and an adaptive photometric error loss to alleviate the flow direction of the large gradient in the occlusion regions. We conducted extensive experiments on the KITTI and Cityscapes datasets, and the results demonstrate that our model outperforms previously published baselines for self-supervised monocular depth estimation.
Xingyu Miao, Yang Bai 0011, Haoran Duan 0001, Yawen Huang, Fan Wan, Xinxing Xu, Yang Long 0001, Yefeng Zheng 0001
IEEE Trans. Circuits Syst. Video Technol.3
2024 Rules for Expectation: Learning to Generate Rules via Social Environment Modeling
abstract
The evolution of natural life is guided by a perpetually adaptive set of rules, encompassing natural laws, human policies, and game mechanics. Automated game design, through the creation of simulated environments populated by AI agents, embodies these rules, aligning with the objectives of artificial life research that seeks to replicate the dynamics of biological life through computational models. This paper presents a comprehensive framework, the Rule Generation Networks (RGN), devised for automated rule design, evaluation, and evolution in line with controllable expectations. We refine and formalize three cardinal elements - rules, strategies, and evaluation - to elucidate the intricate relationships inherent in rule generation tasks. The RGN integrates generative neural networks for rule design and a suite of reinforcement learning models for rule evaluation. To exemplify rule evolution and adaptation across varying environments, we introduce a controllability metric to gauge game dynamics and evolve the rule designer accordingly. Furthermore, we develop two game environments, Maze Run and Trust Evolution, modelling human exploration and societal trade dynamics, to gamify and evaluate the generated rules.
Jiyao Pu, Haoran Duan 0001, Junzhe Zhao, Yang Long 0001
IEEE Trans. Circuits Syst. Video Technol.2
2024 Sentinel-Guided Zero-Shot Learning: A Collaborative Paradigm Without Real Data Exposure
abstract
With increasing concerns over data privacy and model copyrights, especially in the context of collaborations between AI service providers and data owners, an innovative Sentinel-Guided Zero-Shot Learning (SG-ZSL) paradigm is proposed in this work. SG-ZSL is designed to foster efficient collaboration without the need to exchange models or sensitive data. It consists of a teacher model, a student model and a generator that links both model entities. The teacher model serves as a sentinel on behalf of the data owner, replacing real data, to guide the student model at the AI service provider’s end during training. Considering the disparity of knowledge space between the teacher and student, we introduce two variants of the teacher model: the omniscient and the quasi-omniscient teachers. Under these teachers’ guidance, the student model seeks to match the teacher model’s performance and explores domains that the teacher has not covered. To trade-off between privacy and performance, we further introduce two distinct security-level training protocols: white-box and black-box, enhancing the paradigm’s adaptability. Despite the inherent challenges of real data absence in the SG-ZSL paradigm, it consistently outperforms in ZSL and GZSL tasks, notably in the white-box protocol. Our comprehensive evaluation further attests to its robustness and efficiency across various setups, including stringent black-box training protocol.
Fan Wan, Xingyu Miao, Haoran Duan 0001, Jingjing Deng 0001, Yang Long 0001
IEEE Trans. Circuits Syst. Video Technol.3
2024 MRL-Seg: Overcoming Imbalance in Medical Image Segmentation With Multi-Step Reinforcement Learning
abstract
Medical image segmentation is a critical task for clinical diagnosis and research. However, dealing with highly imbalanced data remains a significant challenge in this domain, where the region of interest (ROI) may exhibit substantial variations across different slices. This presents a significant hurdle to medical image segmentation, as conventional segmentation methods may either overlook the minority class or overly emphasize the majority class, ultimately leading to a decrease in the overall generalization ability of the segmentation results. To overcome this, we propose a novel approach based on multi-step reinforcement learning, which integrates prior knowledge of medical images and pixel-wise segmentation difficulty into the reward function. Our method treats each pixel as an individual agent, utilizing diverse actions to evaluate its relevance for segmentation. To validate the effectiveness of our approach, we conduct experiments on four imbalanced medical datasets, and the results show that our approach surpasses other state-of-the-art methods in highly imbalanced scenarios. These findings hold substantial implications for clinical diagnosis and research.
Feiyang Yang, Haoran Duan 0001, Feilong Xu, Yawen Huang, Xiaoli Zhang 0001, Yang Long 0001, Yefeng Zheng 0001
IEEE J. Biomed. Health Informatics3
2024 Dynamic visual-guided selection for zero-shot learning
Yuan Zhou 0023, Haoran Duan 0001, Yang Long 0001
J. Supercomput.4
2024 Prototype Correlation Matching and Class- Relation Reasoning for Few-Shot Medical Image Segmentation
abstract
Few-shot medical image segmentation has achieved great progress in improving accuracy and efficiency of medical analysis in the biomedical imaging field. However, most existing methods cannot explore inter-class relations among base and novel medical classes to reason unseen novel classes. Moreover, the same kind of medical class has large intra-class variations brought by diverse appearances, shapes and scales, thus causing ambiguous visual characterization to degrade generalization performance of these existing methods on unseen novel classes. To address the above challenges, in this paper, we propose a Prototype correlation Matching and Class-relation Reasoning (i.e., PMCR) model. The proposed model can effectively mitigate false pixel correlation matches caused by large intra-class variations while reasoning inter-class relations among different medical classes. Specifically, in order to address false pixel correlation match brought by large intra-class variations, we propose a prototype correlation matching module to mine representative prototypes that can characterize diverse visual information of different appearances well. We aim to explore prototypelevel rather than pixel-level correlation matching between support and query features via optimal transport algorithm to tackle false matches caused by intra-class variations. Meanwhile, in order to explore inter-class relations, we design a class-relation reasoning module to segment unseen novel medical objects via reasoning inter-class relations between base and novel classes. Such inter-class relations can be well propagated to semantic encoding of local query features to improve few-shot segmentation performance. Quantitative comparisons illustrates the large performance improvement of our model over other baseline methods.
Hongliu Li, Yajun Gao, Haoran Duan 0001, Yawen Huang, Yefeng Zheng 0001
IEEE Trans. Medical Imaging4
2023 Dual Feature Augmentation Network for Generalized Zero-shot Learning
Yuan Zhou 0023, Haoran Duan 0001, Yang Long 0001
BMVC3
2023 Privacy-Enhanced Zero-Shot Learning via Data-Free Knowledge Transfer
abstract
Considering the increasing concerns about data copyright and sensitivity issues, we present a novel Privacy-Enhanced Zero-Shot Learning (PE-ZSL) paradigm. The key innovation is to involve a teacher model as the data safeguard to guide the PE-ZSL model training without data sharing. The PE-ZSL model consists of a generator and student network, which can achieve data-free knowledge transfer while maintaining the performance of teacher model. We investigate ‘black-’ and ‘white-box’ scenarios in PE-ZSL task as different levels of framework privacy. Besides, we provide the discussion of teacher model in both omniscient and quasi-omniscient settings according to the knowledge space. Despite simple implementations and data-missing disadvantages, our PE-ZSL framework can retain state-of-the-art ZSL and GZSL performance under the ‘white-box’ scenario. Extensive qualitative and quantitative analysis also demonstrates promising results when deploying the model under ‘black-box’ scenario.
Fan Wan, Daniel Organisciak, Jiyao Pu, Haoran Duan 0001, Peng Zhang 0058, Xingsong Hou, Yang Long 0001
ICME5
2023 Community-Aware Federated Video Summarization
abstract
Video summarization aims to extract representative frames to retain high-level information. Increasing concerns about privacy issues have been raised because conventional large-scale training requires users to upload video samples that may inevitably release sensitive information. In this paper, we thoroughly discuss the Federated Video Summarization problem, i.e., how to obtain a robust video summarization model when video data is distributed on private data islands. Our key contribution includes 1) We propose a fundamental Frame-Based aggregation method to video-related tasks, which differs from the sample-based aggregation in conventional FedAvg. 2) To mitigate the heterogeneous distribution due to community diversity, we propose the Community-Aware Clustering Federated Video Summarization Framework (CFed-VS) that clusters clients via a novel data-driven clustering approach. 3) We further tackle the challenging non-IID setting with a proposed Mixture Transformer, which manifests state-of-the-art performance via extensive quantitative and qualitative experiments on TVSum and SumMe datasets.
Fan Wan, Junyan Wang 0001, Haoran Duan 0001, Yang Song 0001, Maurice Pagnucco, Yang Long 0001
IJCNN3
2023 When Multi-Focus Image Fusion Networks Meet Traditional Edge-Preservation Technology
Zeyu Wang 0009, Haoran Duan 0001, Xiaoli Zhang 0001
Int. J. Comput. Vis.4
2023 Dynamic Unary Convolution in Transformers
abstract
It is uncertain whether the power of transformer architectures can complement existing convolutional neural networks. A few recent attempts have combined convolution with transformer design through a range of structures in series, where the main contribution of this paper is to explore a parallel design approach. While previous transformed-based approaches need to segment the image into patch-wise tokens, we observe that the multi-head self-attention conducted on convolutional features is mainly sensitive to global correlations and that the performance degrades when these correlations are not exhibited. We propose two parallel modules along with multi-head self-attention to enhance the transformer. For local information, a dynamic local enhancement module leverages convolution to dynamically and explicitly enhance positive local patches and suppress the response to less informative ones. For mid-level structure, a novel unary co-occurrence excitation module utilizes convolution to actively search the local co-occurrence between patches. The parallel-designed Dynamic Unary Convolution in Transformer (DUCT) blocks are aggregated into a deep architecture, which is comprehensively evaluated across essential computer vision tasks in image-based classification, segmentation, retrieval and density estimation. Both qualitative and quantitative results show our parallel convolutional-transformer approach with dynamic and unary convolution outperforms existing series-designed structures.
Haoran Duan 0001, Yang Long 0001, Haofeng Zhang 0001, Chris G. Willcocks, Ling Shao 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2023 VSP-Fuse: Multifocus Image Fusion Model Using the Knowledge Transferred From Visual Salience Priors
abstract
Multifocus image fusion (MFIF), as an efficient way to improve the visual effect of images with partial focus defects, is of great significance in the field of image enhancement. According to the imaging principle of the lens, we summarize the visual salience priors (VSP) from the daily photo scene and two relationships from MFIF. Thereby, an edge-sensitive model for MFIF is presented in this study. Supported by VSP, we consider the correlation between salience object detection (SOD) and MFIF, and select the former as a pre-training task. SOD provides the network with realistic depth of field and bokeh effects to learn, and enhances the network’s ability to extract and express the edges of focused objects. Meanwhile, given the scarcity of real multifocus training sets, we propose a randomized approach to generate massive training sets and pseudo-labels based on limited unlabeled data. Besides, two attention modules are designed based on isometric domain transformation (IDT) in the traditional edge-preservation field. IDT removes interference information from feature maps in a low-cost manner, thereby facilitating channel-wise and spatial-wise weight assignments. Experimental results on four datasets show that the performance of our model is superior to that of many supervised models, without the need of any real MFIF training set.
Zeyu Wang 0009, Haoran Duan 0001, Xiaoli Zhang 0001, Jizheng Zhang, Shiping Chen 0001
IEEE Trans. Circuits Syst. Video Technol.4
2022 EfficientTDNN: Efficient Architecture Search for Speaker Recognition
abstract
Convolutional neural networks (CNNs), such as the time-delay neural network (TDNN), have shown their remarkable capability in learning speaker embedding. However, they meanwhile bring a huge computational cost in storage size, processing, and memory. Discovering the specialized CNN that meets a specific constraint requires a substantial effort of human experts. Compared with hand-designed approaches, neural architecture search (NAS) appears as a practical technique in automating the manual architecture design process and has attracted increasing interest in spoken language processing tasks such as speaker recognition. In this paper, we propose EfficientTDNN, an efficient architecture search framework consisting of a TDNN-based supernet and a TDNN-NAS algorithm. The proposed supernet introduces temporal convolution of different ranges of the receptive field and feature aggregation of various resolutions from different layers to TDNN. On top of it, the TDNN-NAS algorithm quickly searches for the desired TDNN architecture via weight-sharing subnets, which surprisingly reduces computation while handling the vast number of devices with various resources requirements. Experimental results on the VoxCeleb dataset show the proposed EfficientTDNN enables approximate$10^{13}$architectures concerning depth, kernel, and width. Considering different computation constraints, it achieves a 2.20% equal error rate (EER) with 204 M multiply-accumulate operations (MACs), 1.41% EER with 571 M MACs as well as 0.94% EER with 1.45 G MACs. Comprehensive investigations suggest that the trained supernet generalizes subnets not sampled during training and obtains a favorable trade-off between accuracy and efficiency.
Rui Wang 0073, Zhihua Wei 0001, Haoran Duan 0001, Shouling Ji, Yang Long 0001, Zhen Hong
IEEE ACM Trans. Audio Speech Lang. Process.3
2022 A Self-Supervised Residual Feature Learning Model for Multifocus Image Fusion
abstract
Multi-focus image fusion (MFIF) attempts to achieve an "all-focused" image from multiple source images with the same scene but different focused objects. Given the lack of multi-focus image sets for network training, we propose a self-supervised residual feature learning model in this paper. The model consists of a feature extraction network and a fusion module. We select image super-resolution as a pretext task in the MFIF field, which is supported by a new residual gradient prior discovered by our theoretical study for low- and high-resolution (LR-HR) image pairs, as well as for multi-focus images. In the pretext task, our network's training set is LR-HR image pairs generated from natural images, and HR images can be regarded as pseudo-labels of LR images. In the fusion task, the trained network extracts residual features of multi-focus images firstly. Secondly, the fusion module, consisting of an activity level measurement and a new boundary refinement method, is leveraged for the features to generated decision maps. Experimental results, both subjective evaluations and objective evaluations, demonstrate that our approach outperforms other state-of-the-art fusion algorithms.
Zeyu Wang 0009, Haoran Duan 0001, Xiaoli Zhang 0001
IEEE Trans. Image Process.3
2021 Medical image fusion based on convolutional neural networks and non-subsampled contourlet transform
Zeyu Wang 0009, Haoran Duan 0001, Yanchi Su, Xiaoli Zhang 0001, Xinjiang Guan
Expert Syst. Appl.3
2021 Multi-level features extraction network with gating mechanism for crowd counting
abstract
Abstract Crowd counting is still a practical and challenging problem owing to scale variations and information loss. Most existing methods based on the straightforward fusion of different features from a deep neural network seem to eliminate this limitation. However, these features are difficult to be fused since they often differ significantly in modality and dimensionality. Unlike previous works, a multi‐level features extraction network with gating mechanism for crowd counting is proposed. Specifically, a multi‐channel gated unit to adaptively extract features in different levels of the network is proposed, which can avoid interference from confusing information. To fully aggregate features via multi‐level fusion, multi‐level features extraction scheme is presented. The multi‐level features extraction network learns to fuse features from multiple levels and reduce false predictions. Extensive experiments and evaluations clearly illustrate that the proposed approach achieves state‐of‐the‐art counting performance against other methods on four mainstream crowd counting benchmarks.
Qiang Guo 0012, Haoran Duan 0001, Yunpeng Wu
IET Image Process.3
2019 Multifocus image fusion using convolutional neural networks in the discrete wavelet transform domain
Zeyu Wang 0009, Haoran Duan 0001, Xiaoli Zhang 0001, Hancheng Wang
Multim. Tools Appl.3