VLDB 2026 Research / reviewers in the wild / expert
Jinming Duan 0001
dblp:125/3221-1
· DBLP profile ↗
64ranked-venue papers
8as first author
39since 2021 · last 2027
0000-0002-5108-2128ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 31 · 4 first-author · 15 since 2021Artificial intelligence and machine learning · 27 · 2 first-author · 20 since 2021Applied, interdisciplinary, general and emerging computing · 24 · 3 first-author · 15 since 2021Databases, data management, data science and information retrieval · 1Theory of computation · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | ASMFNet: Anatomical symmetry-guided multi-modal fusion network for glioma segmentation in MRI
Tongxue Zhou, Su Ruan, Weiping Ding, Haigen Hu, Jinming Duan 0001, Maël Balluet, Bai Ying Lei |
Expert Syst. Appl. | 7 |
| 2026 | FRMF-Net: Feature rectification and adaptive modality fusion guided multi-modal brain tumor segmentation networkabstractBrain tumor segmentation from multi-modal magnetic resonance imaging (MRI) is crucial for computer-assisted diagnosis and treatment planning. However, this task remains highly challenging due to substantial image heterogeneity, modality-inherent variability, and severe class imbalance among tumor sub-regions. To address these issues, we propose FRMF-Net , a F eature R ectification and adaptive M odality F usion guided multi-modal brain tumor segmentation Net work, which consists of three key components: a Modality-Specific Feature Rectification (MSFR) module, an Adaptive Modality Fusion (AMF) module, and a Region-Adaptive Loss (RAL). Specifically, MSFR enhances modality-specific representations by jointly modeling shared and private information, thereby mitigating inter-modality noise and reducing feature discrepancies across modalities. Building on this, AMF performs voxel-wise adaptive fusion through modality-, channel-, and spatial-wise attention, enabling the network to dynamically emphasize the most informative features for accurate tumor delineation. In addition, RAL alleviates the class imbalance issue by adaptively reweighting the contribution of each tumor sub-region according to its spatial extent in each sample. Extensive experiments on the BraTS 2019 and BraTS 2020 datasets demonstrate that FRMF-Net consistently outperforms the state-of-the-art methods, achieving superior Dice score and lower Hausdorff distance, particularly in small and challenging tumor regions. These results confirm that FRMF-Net provides a robust and effective solution for multi-modal brain tumor segmentation. Tongxue Zhou, Su Ruan, Jinming Duan 0001, Yanda Meng, Zhiwei Ji, Bangli Liu, Maël Balluet, Bai Ying Lei |
Expert Syst. Appl. | 4 |
| 2026 | A hierarchical teacher-student learning framework with adaptive cross-modal fusion for brain tumor segmentationabstractAccurate brain tumor segmentation plays an important role in clinical diagnosis, treatment planning, and therapeutic response monitoring. Multi-modal MRI provides complementary structural and functional information, but existing methods remain limited by their inadequate exploitation of cross-modal complementarity and their inability to effectively handle modality-specific disparities and redundant information. To address these challenges, this paper proposes a novel hierarchical teacher-student learning framework with adaptive cross-modal fusion. MRI modalities are grouped into teacher modalities (Flair and T1c) and student modalities (T2 and T1) based on their intrinsic tumor-related characteristics. Central to this framework is the Modality Guidance Module (MGM), which consists of two key components designed to achieve multi-modal feature distillation. Within MGM, the Modality Enhancement Module (MEM) extracts highly discriminative features from teacher modalities. While the Modality Fusion Module (MFM) leverages these features to guide and refine the learning of student modalities. To further capture inter-modal dependencies, a Cross-Modal Fusion Module (CMFM) is introduced to adaptively integrate complementary information across all modalities. Extensive experiments on the BraTS 2018, 2019 and 2020 datasets demonstrate that the proposed method achieves superior performance compared with state-of-the-art approaches. Beyond brain tumor segmentation, the hierarchical teacher-student paradigm and adaptive fusion strategy also hold potential for broader multi-modal image analysis tasks. Tongxue Zhou, Su Ruan, Jinming Duan 0001, Haigen Hu, Yanda Meng, Ling Huang 0003, Defu Yang, Bingbing Jiang 0001, Tingjin Luo, Zhiwei Ji, Bai Ying Lei |
Expert Syst. Appl. | 3 |
| 2026 | UTriGate-Net : Uncertainty-aware brain tumor segmentation via triaxial context encoding and gated modality fusionabstractAccurate segmentation of brain tumors from multi-modal MRI is crucial for diagnosis and treatment planning. However, challenges such as severe class imbalance, modality-specific feature heterogeneity, and predictive uncertainty hinder reliable performance. In this work, we propose UTriGate-Net, a novel uncertainty-aware multi-modal brain tumor segmentation framework. First, we design a Triaxial Context Encoding (TCE) block that extracts anisotropic spatial features by applying directional convolutions along the axial, coronal, and sagittal planes, thereby enhancing 3D contextual representation. Second, we introduce a Gated Modality Fusion (GMF) module, which adaptively integrates complementary information across modalities through modality-specific gating weights that suppress redundancy while retaining salient features. Finally, to improve segmentation reliability, we develop an Uncertainty-Regularized Weighted Loss (URWL) that combines dynamic class-specific weighting to mitigate class imbalance with an entropy-based uncertainty penalty to encourage well-calibrated predictions. Experiments on the BraTS 2019 and 2020 datasets demonstrate that UTriGate-Net achieves superior segmentation accuracy and robustness, particularly in challenging subregions. Overall, the proposed framework offers a promising solution for reliable and precise brain tumor delineation in clinical practice. Tongxue Zhou, Su Ruan, Yanda Meng, Jinming Duan 0001, Haigen Hu, Bingbing Jiang 0001, Zhiwei Ji, Bangli Liu, Tingjin Luo, Bai Ying Lei |
Expert Syst. Appl. | 4 |
| 2026 | IMH-Net: Importance-aware Mamba and cross-modal hypergraph modeling for precise PET/CT tumor segmentationabstract• IMH-Net boosts segmentation accuracy via salient-first modeling and hypergraph fusion • IA-Mamba prioritizes salient regions and preserves fine-grained local details • CSCEM boosts channel spatial complementarity and suppresses crossmodal noise • CHB capture high-order skip connection dependencies to recover details in decoding • Experiments show IMH-Net beats prior methods and stays SOTAcompetitive Precise multimodal tumor segmentation is essential for radiotherapy target contouring, surgical planning, and therapeutic efficacy evaluation. PET provides metabolic activity information, whereas CT offers detailed anatomical structures; their complementarity improves segmentation reliability in complex cases. However, existing sequence-modeling schemes are susceptible to order bias induced by a fixed scanning order, and cross-modal fusion and skip-connection interactions often remain at low-order, coarse-grained levels, making it difficult to jointly achieve salient-region–prioritized modeling, noise suppression, and high-order semantic coupling. To address this, we propose IMH-Net, an automatic multimodal tumor segmentation network based on importance-aware Mamba and hypergraph modeling. The proposed network includes three core components: (1) importance-aware Mamba (IA-Mamba), which estimates patch importance in the encoder stage and dynamically reshuffles the scan order to model salient regions first. (2) The Cross-modal Spatial Channel Enhancement Module (CSCEM) performs cross-modal collaborative enhancement in both channel and spatial dimensions at the bottleneck, emphasizing complementary semantics while suppressing redundant conflicts. (3) The Cross-modal Hypergraph Bridge (CHB) constructs intra- and inter-modality hyperedges at skip connections and leverages hypergraph convolution and hypergraph attention to enable stable high-order interactions and feature coupling. Comprehensive experiments on the public STS, Hecktor 2022, and ECPC datasets validate both the effectiveness of the proposed modules and their complementary synergy. IMH-Net achieves Dice scores of 81.82%, 80.86%, and 91.40% on STS, Hecktor 2022, and ECPC datasets, respectively, outperforming state-of-the-art (SOTA) multimodal segmentation methods in overall performance. Ziwei Zou, Wenqi Lu 0001, Qiongyao Liu, Tongxue Zhou, Jinming Duan 0001 |
Expert Syst. Appl. | 5 |
| 2026 | Long-term stabilized iris tracking with unsupervised constraints on dynamic AS-OCT
Lingxi Hu, Risa Higashita, Xiaoli Xing, Menglan Zhou, Xiaorong Li, Zunjie Xiao, Yinglin Zhang, Chenglin Yao, Jinming Duan 0001, Jiang Liu 0001 |
Medical Image Anal. | 12 |
| 2026 | From pixels to polygons: A survey of deep learning approaches for medical image-to-mesh reconstructionabstractDeep learning-based medical image-to-mesh reconstruction has rapidly evolved, enabling the transformation of medical imaging data into three-dimensional mesh models that are critical in computational medicine and in silico trials for advancing our understanding of disease mechanisms, and diagnostic and therapeutic techniques in modern medicine. This survey systematically categorizes existing approaches into four main categories: template models, statistical models, generative models, and implicit models. Each category is analysed in detail, examining their methodological foundations, strengths, limitations, and applicability to different anatomical structures and imaging modalities. We provide an extensive evaluation of these methods across various anatomical applications, from cardiac imaging to neurological studies, supported by quantitative comparisons using standard metrics. Additionally, we compile and analyse major public datasets available for medical mesh reconstruction tasks and discuss commonly used evaluation metrics and loss functions. The survey identifies current challenges in the field, including requirements for topological correctness, geometric accuracy, and multi-modality integration. Finally, we present promising future research directions in this domain. This systematic review aims to serve as a comprehensive reference for researchers and practitioners in medical image analysis and computational medicine. Fengming Lin, Arezoo Zakeri, Yidan Xue, Michael MacRaild, Haoran Dou, Zherui Zhou, Ziwei Zou, Ali Sarrami-Foroushani, Jinming Duan 0001, Alejandro F. Frangi |
Medical Image Anal. | 9 |
| 2026 | DFuse-Net: Disentangled feature fusion with uncertainty-aware learning for reliable multi-modal brain tumor segmentation
Tongxue Zhou, Su Ruan, Yanda Meng, Jinming Duan 0001, Bai Ying Lei |
Medical Image Anal. | 5 |
| 2026 | EsurvFusion: An Evidential Multimodal Survival Fusion Model Based on Epistemic Random Fuzzy SetsabstractMultimodal survival analysis aims to combine heterogeneous data sources to improve the prediction quality of survival outcomes. However, this task is particularly challenging due to high heterogeneity and noise across data sources. Additionally, the exact survival time is often censored (partially known) due to incomplete event observation. To address the above challenges, we propose a novel interpretable evidential multimodal survival fusion model, EsurvFusion. This model is designed to combine multimodal data at the decision level using Epistemic Random Fuzzy Sets that jointly handle both data and model uncertainty while incorporating modality-level reliability. Specifically, EsurvFusion first models unimodal data with newly introduced Gaussian random fuzzy numbers, producing possible unimodal survival predictions along with corresponding aleatory and epistemic uncertainty. It then estimates modality-level reliability through a reliability discounting layer to correct the misleading impact of noisy data modalities. Finally, a multimodal evidence fusion layer is introduced to combine the discounted predictions, revealing modality-level influence based on the learned reliability coefficients. Extensive experiments on four multimodal cancer survival datasets demonstrate the effectiveness of our model in handling highly heterogeneous data, establishing a new state-of-the-art performance on several benchmarks. Ling Huang 0003, Yucheng Xing, Qika Lin, Jinming Duan 0001, Su Ruan, Mengling Feng |
IEEE Trans. Fuzzy Syst. | 4 |
| 2025 | Aligning Medical Images and Language Through Multimodal Medical RationalesabstractLarge vision-language models (LVLMs) have gained widespread attention in the medical field for their outstanding capability in handling image-text representations. However, the misalignment between medical images and clinical text presents factuality challenges for Medical LVLMs (Med-LVLMs), often resulting in hallucinations. Multimodal Chain-of-Thought (MCoT) can reduce factual errors of Med-LVLMs by encouraging explicit step-by-step reasoning, but it poses two major challenges. First, factually accurate medical rationales are crucial for aligning medical images with the corresponding clinical texts, yet existing Med-LVLMs struggle to generate such rationales. Second, if the model's initial prediction is correct, its inherent knowledge can be disrupted by an over-reliance on the generated medical rationale, resulting in an incorrect answer. To address these challenges, we propose MMReT, a novel multimodal medical reasoning tuning approach designed to improve the factual accuracy of Med-LVLMs. First, we employ carefully designed prompts to guide GPT-4o in generating high-quality medical rationales, which are then used to fine-tune the original Med-LVLM. Second, to address errors stemming from excessive reliance on generated rationales, we introduce medical reasoning preference fine-tuning, which encourages the model to maintain an appropriate balance between leveraging its inherent knowledge and incorporating generated medical rationale. Experimental results show that MMReT substantially enhances the factuality of Med-LVLMs, outperforming previous methods with average improvements of 8.0% on VQA-RAD and 11.6% on SLAKE in factual accuracy. Zhi Chen 0015, Beiji Zou 0001, Xiaoyan Kui, Ziwei Zou, Jinming Duan 0001 |
BIBM | 5 |
| 2025 | SACB-Net: Spatial-awareness Convolutions for Medical Image RegistrationabstractDeep learning-based image registration methods have shown state-of-the-art performance and rapid inference speeds. Despite these advances, many existing approaches fall short in capturing spatially varying information in non-local regions of feature maps due to the reliance on spatially-shared convolution kernels. This limitation leads to suboptimal estimation of deformation fields. In this paper, we propose a 3D Spatial-Awareness Convolution Block (SACB) to enhance the spatial information within feature representations. Our SACB estimates the spatial clusters within feature maps by leveraging feature similarity and subsequently parameterizes the adaptive convolution kernels across diverse regions. This adaptive mechanism generates the convolution kernels (weights and biases) tailored to spatial variations, thereby enabling the network to effectively capture spatially varying information. Building on SACB, we introduce a pyramid flow estimator (named SACB-Net) that integrates SACBs to facilitate multi-scale flow composition, particularly addressing large deformations. Experimental results on the brain IXI and LPBA datasets as well as Abdomen CT datasets demonstrate the effectiveness of SACB and the superiority of SACB-Net over the state-of-the-art learning-based registration methods. The code is available at https://github.com/x-xc/SACB_Net. Xinxing Cheng, Tianyang Miller, Wenqi Lu 0001, Qingjie Meng, Alejandro F. Frangi, Jinming Duan 0001 |
CVPR | 6 |
| 2025 | LongVALE: Vision-Audio-Language-Event Benchmark Towards Time-Aware Omni-Modal Perception of Long VideosabstractDespite impressive advancements in video understanding, most efforts remain limited to coarse-grained or visual-only video tasks. However, real-world videos encompass omni-modal information (vision, audio, and speech) with a series of events forming a cohesive storyline. The lack of multi-modal video data with fine-grained event annotations and the high cost of manual labeling are major obstacles to comprehensive omni-modality video perception. To address this gap, we propose an automatic pipeline consisting of high-quality multi-modal video filtering, semantically coherent omni-modal event boundary detection, and cross-modal correlation-aware event captioning. In this way, we present LongVALE, the first- ever Vision-Audio-Language Event understanding benchmark comprising 105K omni-modal events with precise temporal boundaries and detailed relation-aware captions within 8.4K high-quality long videos. Further, we build a baseline that leverages LongVALE to enable video large language models (LLMs) for omni-modality fine-grained temporal video understanding for the first time. Extensive experiments demonstrate the effectiveness and great potential of LongVALE in advancing comprehensive multi-modal video understanding. The dataset and code are available at https://ttgeng233.github.io/LongVALE/. Tiantian Geng, Qingni Wang, Teng Wang 0007, Jinming Duan 0001, Feng Zheng 0001 |
CVPR | 5 |
| 2025 | Exploring Temporal Constraints for Unsupervised Iris Motion Tracking in AS-OCT VideosabstractIris motion tracking is critical for discriminating the iris stiffness and developmental stage of primary angle-closure disease (PACD). Anterior segment optical coherence tomography (AS-OCT) video is a highly efficient approach to observe the morphological determinant in iris motion. However, the iris exhibits inconsistent elastic changes during movement, accompanied by changes in local features after long-term frames. Currently, iris tracking methods have not yet been studied in AS-OCT videos. In this paper, we propose a Temporal Constraint-based Tracking Morph (TCTMorph) for estimating iris trajectory in long-term AS-OCT videos. We first estimate the deformation fields between three interrelated frames by a multi-frame diffeomorphic registration network. Then, we estimate iris trajectory from these results in long-term AS-OCT video sequences by leveraging temporal constraints among the consecutive flows. Our experiments on multi-center AS-OCT glaucoma datasets demonstrate that our method outperforms conventional motion tracking methods for long-term iris trajectory tracking. Lingxi Hu, Risa Higashita, Xiaoli Xing, Menglan Zhou, Xiaorong Li, Jinming Duan 0001, Jiang Liu 0001 |
ICASSP | 9 |
| 2025 | 4D CardioSynth: Synthesising Dynamic Virtual Heart Populations Through Spatiotemporal Disentanglement
Haoran Dou, Jinghan Huang 0003, Arezoo Zakeri, Zherui Zhou, Tingting Mu, Jinming Duan 0001, Alejandro F. Frangi |
MICCAI (3) | 6 |
| 2025 | Accelerating cardiac radial-MRI: Fully polar based technique using compressed sensing and deep learningabstractFast radial-MRI approaches based on compressed sensing (CS) and deep learning (DL) often use non-uniform fast Fourier transform (NUFFT) as the forward imaging operator, which might introduce interpolation errors and reduce image quality. Using the polar Fourier transform (PFT), we developed fully polar CS and DL algorithms for fast 2D cardiac radial-MRI. Our methods directly reconstruct images in polar spatial space from polar k-space data, eliminating frequency interpolation and ensuring an easy-to-compute data consistency term for the DL framework via the variable splitting (VS) scheme. Furthermore, PFT reconstruction produces initial images with fewer artifacts in a reduced field of view, making it a better starting point for CS and DL algorithms, especially for dynamic imaging, where information from a small region of interest is critical, as opposed to NUFFT, which often results in global streaking artifacts. In the cardiac region, PFT-based CS technique outperformed NUFFT-based CS at acceleration rates of 5x (mean SSIM: 0.8831 vs. 0.8526), 10x (0.8195 vs. 0.7981), and 15x (0.7720 vs. 0.7503). Our PFT(VS)-DL technique outperformed the NUFFT(GD)-based DL method, which used unrolled gradient descent with the NUFFT as the forward imaging operator, with mean SSIM scores of 0.8914 versus 0.8617 at 10x and 0.8470 versus 0.8301 at 15x. Radiological assessments revealed that PFT(VS)-based DL scored 2.9±0.30 and 2.73±0.45 at 5x and 10x, whereas NUFFT(GD)-based DL scored 2.7±0.47 and 2.40±0.50, respectively. Our methods suggest a promising alternative to NUFFT-based fast radial-MRI for dynamic imaging, prioritizing reconstruction quality in a small region of interest over whole image quality. Vahid Ghodrati, Jinming Duan 0001, Fadil Ali, Arash Bedayat, Ashley Prosper, Mark Bydder |
Medical Image Anal. | 2 |
| 2025 | UniAV: Unified Audio-Visual Perception for Multi-Task Video Event LocalizationabstractVideo event localization tasks include temporal action localization (TAL), sound event detection (SED) and audio-visual event localization (AVEL). Existing methods tend to over-specialize on individual tasks, neglecting the equal importance of these different events for a complete understanding of video content. In this work, we aim to develop a unified framework to solve TAL, SED and AVEL tasks together to facilitate holistic video understanding. However, it is challenging since different tasks emphasize distinct event characteristics and there are substantial disparities in existing task-specific datasets (size/domain/duration). It leads to unsatisfactory results when applying a naive multi-task strategy. To tackle the problem, we introduce UniAV, a Unified Audio-Visual perception network to effectively learn and share mutually beneficial knowledge across tasks and modalities. Concretely, we propose a unified audio-visual encoder to derive generic representations from multiple temporal scales for videos from all tasks. Meanwhile, task-specific experts are designed to capture the unique knowledge specific to each task. Besides, instead of using separate prediction heads, we develop a novel unified language-aware classifier by utilizing semantic-aligned task prompts, enabling our model to flexibly localize various instances across tasks with an impressive open-set ability to localize novel categories. Extensive experiments demonstrate that UniAV, with its unified architecture, significantly outperforms both single-task models and the naive multi-task baseline across all three tasks. It achieves superior or on-par performances compared to the state-of-the-art task-specific methods on ActivityNet 1.3, DESED and UnAV-100 benchmarks. Tiantian Geng, Teng Wang 0007, Jinming Duan 0001, Yanfu Zhang, Weili Guan, Feng Zheng 0001, Ling Shao 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2025 | Hierarchical Context Transformer for Multi-Level Semantic Scene UnderstandingabstractA comprehensive and explicit understanding of surgical scenes plays a vital role in developing context-aware computer-assisted systems in the operating theatre. However, few works provide systematical analysis to enable hierarchical surgical scene understanding. In this work, we propose to represent the tasks set [phase recognition$\rightarrow $step recognition$\rightarrow $action and instrument detection] as multi-level semantic scene understanding (MSSU). For this target, we propose a novel hierarchical context transformer (HCT) network and thoroughly explore the relations across the different level tasks. Specifically, a hierarchical relation aggregation module (HRAM) is designed to concurrently relate entries inside multi-level interaction information and then augment task-specific features. To further boost the representation learning of the different tasks, inter-task contrastive learning (ICL) is presented to guide the model to learn task-wise features via absorbing complementary information from other tasks. Furthermore, considering the computational costs of the transformer, we propose HCT+ to integrate the spatial and temporal adapter to access competitive performance on substantially fewer tunable parameters. Extensive experiments on our cataract dataset and a publicly available endoscopic PSI-AVA dataset demonstrate the outstanding performance of our method, consistently exceeding the state-of-the-art methods by a large margin. The code is available athttps://github.com/Aurora-hao/HCT. Luoying Hao, Huazhu Fu, Jinming Duan 0001, Jiang Liu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2025 | TFNet: A Temporal-Frequency Domain Model for Gait Biomechanical Signal PredictionabstractPrecise prediction of upcoming gait signals, especially over an extended time scale like the entire gait cycle, is crucial. During this period, devices or gait retraining programs can respond to dynamic changes while considering multiple factors in the neurological and musculoskeletal systems. This enables effective adjustments, ultimately optimising outcomes based on the unique rehabilitation goals. However, current state-of-the-art models, whether driven by physical modelling or data modelling approaches, are constrained by short prediction time scales, limited accuracy, and high computational costs, which hinder their use on edge devices. We developed TFNet, a dual-stream neural network model that integrates temporal and frequency domain analyses to accurately predict biomechanical signals across the entire gait cycle. TFNet predicted lower limb joint angles and ground reaction forces with high precision, within 5 degrees and 0.1 body weight, respectively. The model demonstrated the feasibility for deployment on edge devices and adaptability to patients with gait impairments. Explainability analysis highlighted key biomechanical features throughout the gait cycle, improving interpretability and clinical relevance. These comprehensive validations demonstrate the potential of TFNet as a reliable and cost-effective solution for clinical applications aimed at restoring and enhancing gait function. Qingyao Bian, Weida Wang, Jinming Duan 0001, Ziyun Ding |
IEEE J. Biomed. Health Informatics | 3 |
| 2025 | Decoder-Only Image RegistrationabstractIn unsupervised medical image registration, encoder-decoder architectures are widely used to predict dense, full-resolution displacement fields from paired images. Despite their popularity, we question the necessity of making both the encoder and decoder learnable. To address this, we propose LessNet, a simplified network architecture with only a learnable decoder, while completely omitting a learnable encoder. Instead, LessNet replaces the encoder with simple, handcrafted features, eliminating the need to optimize encoder parameters. This results in a compact, efficient, and decoder-only architecture for 3D medical image registration. We evaluate our decoder-only LessNet on five registration tasks: 1) inter-subject brain registration using the OASIS-1 dataset, 2) atlas-based brain registration using the IXI dataset, 3) cardiac ES-ED registration using the ACDC dataset, 4) inter-subject abdominal MR registration using the CHAOS dataset, and 5) multi-study, multi-site brain registration using images from 13 public datasets. Our results demonstrate that LessNet can effectively and efficiently learn both dense displacement and diffeomorphic deformation fields. Furthermore, our decoder-only LessNet can achieve comparable registration performance to benchmarking methods such as VoxelMorph and TransMorph, while requiring significantly fewer computational resources. Our code and pre-trained models are available at https://github.com/xi-jia/LessNet. Xi Jia, Wenqi Lu 0001, Xinxing Cheng, Jinming Duan 0001 |
IEEE Trans. Medical Imaging | 4 |
| 2024 | Optimizing ADMM and Over-Relaxed ADMM Parameters for Linear Quadratic ProblemsabstractThe Alternating Direction Method of Multipliers (ADMM) has gained significant attention across a broad spectrum of machine learning applications. Incorporating the over-relaxation technique shows potential for enhancing the convergence rate of ADMM. However, determining optimal algorithmic parameters, including both the associated penalty and relaxation parameters, often relies on empirical approaches tailored to specific problem domains and contextual scenarios. Incorrect parameter selection can significantly hinder ADMM's convergence rate. To address this challenge, in this paper we first propose a general approach to optimize the value of penalty parameter, followed by a novel closed-form formula to compute the optimal relaxation parameter in the context of linear quadratic problems (LQPs). We then experimentally validate our parameter selection methods through random instantiations and diverse imaging applications, encompassing diffeomorphic image registration, image deblurring, and MRI reconstruction. Jintao Song, Wenqi Lu 0001, Yunwen Lei, Yuchao Tang, Zhenkuan Pan 0001, Jinming Duan 0001 |
AAAI | 6 |
| 2024 | WiNet: Wavelet-Based Incremental Learning for Efficient Medical Image Registration
Xinxing Cheng, Xi Jia, Wenqi Lu 0001, Qiufu Li, LinLin Shen, Alexander Krull, Jinming Duan 0001 |
MICCAI (2) | 7 |
| 2024 | Multi-granularity learning of explicit geometric constraint and contrast for label-efficient medical image segmentation and differentiable clinical function assessmentabstractAutomated segmentation is a challenging task in medical image analysis that usually requires a large amount of manually labeled data. However, most current supervised learning based algorithms suffer from insufficient manual annotations, posing a significant difficulty for accurate and robust segmentation. In addition, most current semi-supervised methods lack explicit representations of geometric structure and semantic information, restricting segmentation accuracy. In this work, we propose a hybrid framework to learn polygon vertices, region masks, and their boundaries in a weakly/semi-supervised manner that significantly advances geometric and semantic representations. Firstly, we propose multi-granularity learning of explicit geometric structure constraints via polygon vertices (PolyV) and pixel-wise region (PixelR) segmentation masks in a semi-supervised manner. Secondly, we propose eliminating boundary ambiguity by using an explicit contrastive objective to learn a discriminative feature space of boundary contours at the pixel level with limited annotations. Thirdly, we exploit the task-specific clinical domain knowledge to differentiate the clinical function assessment end-to-end. The ground truth of clinical function assessment, on the other hand, can serve as auxiliary weak supervision for PolyV and PixelR learning. We evaluate the proposed framework on two tasks, including optic disc (OD) and cup (OC) segmentation along with vertical cup-to-disc ratio (vCDR) estimation in fundus images; left ventricle (LV) segmentation at end-diastolic and end-systolic frames along with ejection fraction (LVEF) estimation in two-dimensional echocardiography images. Experiments on nine large-scale datasets of the two tasks under different label settings demonstrate our model’s superior performance on segmentation and clinical function assessment. Yanda Meng, Jianyang Xie, Jinming Duan 0001, Martha Joddrell, Savita Madhusudhan, Tunde Peto, Yitian Zhao, Yalin Zheng |
Medical Image Anal. | 4 |
| 2024 | Structure and Intensity Unbiased Translation for 2D Medical Image SegmentationabstractData distribution gaps often pose significant challenges to the use of deep segmentation models. However, retraining models for each distribution is expensive and time-consuming. In clinical contexts, device-embedded algorithms and networks, typically unretrainable and unaccessable post-manufacture, exacerbate this issue. Generative translation methods offer a solution to mitigate the gap by transferring data across domains. However, existing methods mainly focus on intensity distributions while ignoring the gaps due to structure disparities. In this paper, we formulate a new image-to-image translation task to reduce structural gaps. We propose a simple, yet powerful Structure-Unbiased Adversarial (SUA) network which accounts for both intensity and structural differences between the training and test sets for segmentation. It consists of a spatial transformation block followed by an intensity distribution rendering module. The spatial transformation block is proposed to reduce the structural gaps between the two images. The intensity distribution rendering module then renders the deformed structure to an image with the target intensity distribution. Experimental results show that the proposed SUA method has the capability to transfer both intensity distribution and structural content between multiple pairs of datasets and is superior to prior arts in closing the gaps for improving segmentation. Tianyang Miller, Shaoming Zheng, Jun Cheng 0003, Xi Jia, Joseph Bartlett, Xinxing Cheng, Zhaowen Qiu, Huazhu Fu, Jiang Liu 0001, Ales Leonardis, Jinming Duan 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 11 |
| 2023 | Fourier-Net: Fast Image Registration with Band-Limited DeformationabstractUnsupervised image registration commonly adopts U-Net style networks to predict dense displacement fields in the full-resolution spatial domain. For high-resolution volumetric image data, this process is however resource-intensive and time-consuming. To tackle this problem, we propose the Fourier-Net, replacing the expansive path in a U-Net style network with a parameter-free model-driven decoder. Specifically, instead of our Fourier-Net learning to output a full-resolution displacement field in the spatial domain, we learn its low-dimensional representation in a band-limited Fourier domain. This representation is then decoded by our devised model-driven decoder (consisting of a zero padding layer and an inverse discrete Fourier transform layer) to the dense, full-resolution displacement field in the spatial domain. These changes allow our unsupervised Fourier-Net to contain fewer parameters and computational operations, resulting in faster inference speeds. Fourier-Net is then evaluated on two public 3D brain datasets against various state-of-the-art approaches. For example, when compared to a recent transformer-based method, named TransMorph, our Fourier-Net, which only uses 2.2% of its parameters and 6.66% of the multiply-add operations, achieves a 0.5% higher Dice score and an 11.48 times faster inference speed. Code is available at https://github.com/xi-jia/Fourier-Net. Xi Jia, Joseph Bartlett, Wei Chen 0092, Siyang Song, Tianyang Miller, Xinxing Cheng, Wenqi Lu 0001, Zhaowen Qiu, Jinming Duan 0001 |
AAAI | 9 |
| 2023 | AVS-Net: Attention-based Variable Splitting Network for P-MRI AccelerationabstractReconstructing undersampled medical images from k-space magnetic resonance imaging (MRI) data with multi-coil acquisitions is a major challenge for deep learning algorithms. Applying convolutional architecture has become a widely used framework for reconstructing MR images. It offers excellent performance but causes signal synthesis and artifact compensation concerns due to the limitations of local receptive fields. Inspired by Transformers, we propose a novel Attention-based Variable MRI reconstruction Splitting Network termed AVS-Net, whose features are expressed as keys and queries, capturing a wider range of relations among parallel k-space data for the compressed sensing reconstruction problem. AVSNet overcomes two challenges in visual tokenization and parallel reconstruction pixel consistency. The model was trained with 400 knee MRI data in 40 slices and 15 channels and tested with 400 knee MR images. As a quantitative result with respect to PSNR, the training achieved an average of 39.46 with default parameters, compared to a 6.8% increase with SOTA. Combining the NMSE, PSNR, SSIM, and MSE metrics, the numerical results show that our method provides better reconstruction perceptual quality compared to SOTA. The code is publicly available at https://github.com/AVS-Net/AVS-Net. Jing Li 0114, Zigan Wang, Jinming Duan 0001 |
BIBM | 4 |
| 2023 | Dense-Localizing Audio-Visual Events in Untrimmed Videos: A Large-Scale Benchmark and BaselineabstractExisting audio-visual event localization (AVE) handles manually trimmed videos with only a single instance in each of them. However, this setting is unrealistic as natural videos often contain numerous audio-visual events with different categories. To better adapt to real-life applications, in this paper we focus on the task of dense-localizing audio-visual events, which aims to jointly localize and recognize all audio-visual events occurring in an untrimmed video. The problem is challenging as it requires fine-grained audio-visual scene and context understanding. To tackle this problem, we introduce the first Untrimmed Audio-Visual (UnAV-J 00) dataset, which contains 10K untrimmed videos with over 30K audio-visual events. Each video has 2.8 audio-visual events on average, and the events are usually related to each other and might co-occur as in real-life scenes. Next, we formulate the task using a new learning-based framework, which is capable of fully integrating audio and visual modalities to localize audio-visual events with various lengths and capture dependencies between them in a single pass. Extensive experiments demonstrate the effectiveness of our method as well as the significance of multi-scale cross-modal perception and dependency modeling for this task. The dataset and code are available at https://unav100.github.io. Tiantian Geng, Teng Wang 0007, Jinming Duan 0001, Runmin Cong, Feng Zheng 0001 |
CVPR | 3 |
| 2023 | UniFace: Unified Cross-Entropy Loss for Deep Face RecognitionabstractAs a widely used loss function in deep face recognition, the softmax loss cannot guarantee that the minimum positive sample-to-class similarity is larger than the maximum negative sample-to-class similarity. As a result, no unified threshold is available to separate positive sample-to-class pairs from negative sample-to-class pairs. To bridge this gap, we design a UCE (Unified Cross-Entropy) loss for face recognition model training, which is built on the vital constraint that all the positive sample-to-class similarities shall be larger than the negative ones. Our UCE loss can be integrated with margins for a further performance boost. The face recognition model trained with the proposed UCE loss, UniFace, was intensively evaluated using a number of popular public datasets like MFR, IJB-C, LFW, CFP-FP, AgeDB, and MegaFace. Experimental results show that our approach outperforms SOTA methods like SphereFace, CosFace, ArcFace, Partial FC, etc. Especially, till the submission of this work (Mar. 8, 2023), the proposed UniFace achieves the highest TAR@MR-All on the academic track of the MFR-ongoing challenge. $\color{Blue}{\mathbf{Code}}$ is publicly available. Jiancan Zhou, Xi Jia, Qiufu Li, LinLin Shen, Jinming Duan 0001 |
ICCV | 5 |
| 2023 | ACT-Net: Anchor-Context Action Detection in Surgery Videos
Luoying Hao, Heng Li 0010, Huazhu Fu, Jinming Duan 0001, Jiang Liu 0001 |
MICCAI (9) | 7 |
| 2023 | UniTSFace: Unified Threshold Integrated Sample-to-Sample Loss for Face RecognitionabstractSample-to-class-based face recognition models can not fully explore the cross-sample relationship among large amounts of facial images, while sample-to-sample-based models require sophisticated pairing processes for training. Furthermore, neither method satisfies the requirements of real-world face verification applications, which expect a unified threshold separating positive from negative facial pairs. In this paper, we propose a unified threshold integrated sample-to-sample based loss (USS loss), which features an explicit unified threshold for distinguishing positive from negative pairs. Inspired by our USS loss, we also derive the sample-to-sample based softmax and BCE losses, and discuss their relationship. Extensive evaluation on multiple benchmark datasets, including MFR, IJB-C, LFW, CFP-FP, AgeDB, and MegaFace, demonstrates that the proposed USS loss is highly efficient and can work seamlessly with sample-to-class-based losses. The embedded loss (USS and sample-to-class Softmax loss) overcomes the pitfalls of previous approaches and the trained facial model UniTSFace exhibits exceptional performance, outperforming state-of-the-art methods, such as CosFace, ArcFace, VPL, AnchorFace, and UNPG. Our code is available at https://github.com/CVI-SZU/UniTSFace. Qiufu Li, Xi Jia, Jiancan Zhou, LinLin Shen, Jinming Duan 0001 |
NeurIPS | 5 |
| 2023 | Weakly/Semi-supervised Left Ventricle Segmentation in 2D Echocardiography with Uncertain Region-Aware Contrastive Learning
Yanda Meng, Jianyang Xie, Jinming Duan 0001, Yitian Zhao, Yalin Zheng |
PRCV (13) | 4 |
| 2023 | Arbitrary Order Total Variation for Deformable Image RegistrationabstractIn this work, we investigate image registration in a variational framework and focus on regularization generality and solver efficiency. We first propose a variational model combining the state-of-the-art sum of absolute differences (SAD) and a new arbitrary order total variation regularization term. The main advantage is that this variational model preserves discontinuities in the resultant deformation while being robust to outlier noise. It is however non-trivial to optimize the model due to its non-convexity, non-differentiabilities, and generality in the derivative order. To tackle these, we propose to first apply linearization to the problem to formulate a convex objective function and then break down the resultant convex optimization into several point-wise, closed-form subproblems using a fast, over-relaxed alternating direction method of multipliers (ADMM). With our proposed algorithm, we show that solving higher-order variational formulations is similar to solving their lower-order counterparts. Extensive experiments show that our ADMM is significantly more efficient than both the subgradient and primal-dual algorithms particularly when higher-order derivatives are used, and that our new models outperform state-of-the-art methods based on deep learning and free-form deformation. Our code implemented in both Matlab and Pytorch is publicly available at https://github.com/j-duan/AOTV. Jinming Duan 0001, Xi Jia, Joseph Bartlett, Wenqi Lu 0001, Zhaowen Qiu |
Pattern Recognit. | 1 |
| 2023 | Adversarial Learning of Object-Aware Activation Map for Weakly-Supervised Semantic SegmentationabstractRecent years have witnessed impressive advances in the area of weakly-supervised semantic segmentation (WSSS). However, most of existing approaches are based on class activation maps (CAMs), which suffer from the under-segmentation problem (i.e., objects of interest are segmented partially). Although a number of literature works have been proposed to tackle this under-segmentation problem, we argue that these solutions built on CAMs may not be optimal for the WSSS task. Instead, in this paper we propose a network based on the object-aware activation map (OAM). The proposed network, termed OAM-Net, consists of four loss functions (foreground loss, background loss, average pixel and consistency loss) which ensure exactness, completeness, compactness and consistency of segmented objects via adversarial training. Compared to conventional CAM-based methods, our OAM-Net overcomes the under-segmentation drawback and significantly improves segmentation accuracy with negligible computational cost. A thorough comparison between OAM-Net and CAM-based approaches is carried out on the PASCAL VOC2012 dataset, and experimental results show that our network outperforms state-of-the-art approaches by a large margin. The code will be available soon. Junliang Chen 0002, Weizeng Lu, Yuexiang Li, LinLin Shen, Jinming Duan 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2023 | Learn2Reg: Comprehensive Multi-Task Medical Image Registration Challenge, Dataset and Evaluation in the Era of Deep LearningabstractImage registration is a fundamental medical image analysis task, and a wide variety of approaches have been proposed. However, only a few studies have comprehensively compared medical image registration approaches on a wide range of clinically relevant tasks. This limits the development of registration methods, the adoption of research advances into practice, and a fair benchmark across competing approaches. The Learn2Reg challenge addresses these limitations by providing a multi-task medical image registration data set for comprehensive characterisation of deformable registration algorithms. A continuous evaluation will be possible at https://learn2reg.grand-challenge.org. Learn2Reg covers a wide range of anatomies (brain, abdomen, and thorax), modalities (ultrasound, CT, MR), availability of annotations, as well as intra- and inter-patient registration evaluation. We established an easily accessible framework for training and validation of 3D registration methods, which enabled the compilation of results of over 65 individual method submissions from more than 20 unique teams. We used a complementary set of metrics, including robustness, accuracy, plausibility, and runtime, enabling unique insight into the current state-of-the-art of medical image registration. This paper describes datasets, tasks, evaluation methods and results of the challenge, as well as results of further analysis of transferability to new datasets, the importance of label supervision, and resulting bias. While no single approach worked best across all tasks, many methodological aspects could be identified that push the performance of medical image registration to new state-of-the-art performance. Furthermore, we demystified the common belief that conventional registration methods have to be much slower than deep-learning-based methods. Alessa Hering, Lasse Hansen, Tony C. W. Mok, Albert C. S. Chung, Hanna Siebert, Stephanie Häger, Annkristin Lange, Sven Kuckertz, Stefan Heldmann, Wei Shao 0008, Sulaiman Vesal, Mirabela Rusu, Geoffrey A. Sonn, Théo Estienne, Maria Vakalopoulou, Luyi Han, Yunzhi Huang, Pew-Thian Yap, Mikael Brudfors, Yaël Balbastre, Samuel Joutard, Marc Modat, Gal Lifshitz, Dan Raviv, Jinxin Lv, Qiang Li 0018, Vincent Jaouen, Dimitris Visvikis, Constance Fourcade, Mathieu Rubeaux, Wentao Pan 0001, Zhe Xu 0012, Bailiang Jian, Francesca De Benetti, Marek Wodzinski, Niklas Gunnarsson, Jens Sjölund, Daniel Grzech, Huaqi Qiu, Zeju Li, Alexander Thorley, Jinming Duan 0001, Christoph Großbröhmer, Andrew Hoopes, Ingerid Reinertsen, Yiming Xiao 0001, Bennett A. Landman, Yuankai Huo, Keelin Murphy, Nikolas Leßmann, Bram van Ginneken, Adrian V. Dalca, Mattias P. Heinrich |
IEEE Trans. Medical Imaging | 42 |
| 2022 | Learning a Model-Driven Variational Network for Deformable Image RegistrationabstractData-driven deep learning approaches to image registration can be less accurate than conventional iterative approaches, especially when training data is limited. To address this issue and meanwhile retain the fast inference speed of deep learning, we propose VR-Net, a novel cascaded variational network for unsupervised deformable image registration. Using a variable splitting optimization scheme, we first convert the image registration problem, established in a generic variational framework, into two sub-problems, one with a point-wise, closed-form solution and the other one being a denoising problem. We then propose two neural layers (i.e. warping layer and intensity consistency layer) to model the analytical solution and a residual U-Net (termed generalized denoising layer) to formulate the denoising problem. Finally, we cascade the three neural layers multiple times to form our VR-Net. Extensive experiments on three (two 2D and one 3D) cardiac magnetic resonance imaging datasets show that VR-Net outperforms state-of-the-art deep learning methods on registration accuracy, whilst maintaining the fast inference speed of deep learning and the data-efficiency of variational models. Xi Jia, Alexander Thorley, Wei Chen 0092, Huaqi Qiu, LinLin Shen, Iain B. Styles, Hyung Jin Chang, Ales Leonardis, Antonio M. Simoes Monteiro de Marvao, Declan P. O'Regan, Daniel Rueckert, Jinming Duan 0001 |
IEEE Trans. Medical Imaging | 12 |
| 2021 | FS-Net: Fast Shape-Based Network for Category-Level 6D Object Pose Estimation With Decoupled Rotation MechanismabstractIn this paper, we focus on category-level 6D pose and size estimation from a monocular RGB-D image. Previous methods suffer from inefficient category-level pose feature extraction, which leads to low accuracy and inference speed. To tackle this problem, we propose a fast shape-based network (FS-Net) with efficient category-level feature extraction for 6D pose estimation. First, we design an orientation aware autoencoder with 3D graph convolution for latent feature extraction. Thanks to the shift and scale-invariance properties of 3D graph convolution, the learned latent feature is insensitive to point shift and object size. Then, to efficiently decode category-level rotation information from the latent feature, we propose a novel decoupled rotation mechanism that employs two decoders to complementarily access the rotation information. For translation and size, we estimate them by two residuals: the difference between the mean of object points and ground truth translation, and the difference between the mean size of the category and ground truth size, respectively. Finally, to increase the generalization ability of the FS-Net, we propose an on-line box-cage based 3D deformation mechanism to augment the training data. Extensive experiments on two benchmark datasets show that the proposed method achieves state-of-the-art performance in both category- and instance-level 6D object pose estimation. Especially in category-level pose estimation, without extra synthetic data, our method outperforms existing methods by 6.3% on the NOCS-REAL dataset1. Wei Chen 0092, Xi Jia, Hyung Jin Chang, Jinming Duan 0001, LinLin Shen, Ales Leonardis |
CVPR | 4 |
| 2021 | Nesterov Accelerated ADMM for Fast Diffeomorphic Image Registration
Alexander Thorley, Xi Jia, Hyung Jin Chang, Karina Bunting, Victoria Stoll, Antonio M. Simoes Monteiro de Marvao, Declan P. O'Regan, Georgios V. Gkoutos, Dipak Kotecha, Jinming Duan 0001 |
MICCAI (4) | 11 |
| 2021 | Nonlocal graph theory based transductive learning for hyperspectral image classification
Baoxiang Huang, Linyao Ge, Ge Chen 0002, Milena Radenkovic 0001, Jinming Duan 0001, Zhenkuan Pan 0001 |
Pattern Recognit. | 6 |
| 2021 | Surrogate network-based sparseness hyper-parameter optimization for deep expression recognition
Weicheng Xie 0001, Wenting Chen, LinLin Shen, Jinming Duan 0001, Meng Yang 0001 |
Pattern Recognit. | 4 |
| 2021 | Adaptive Weighting of Handcrafted Feature Losses for Facial Expression RecognitionabstractDue to the importance of facial expressions in human-machine interaction, a number of handcrafted features and deep neural networks have been developed for facial expression recognition. While a few studies have shown the similarity between the handcrafted features and the features learned by deep network, a new feature loss is proposed to use feature bias constraint of handcrafted and deep features to guide the deep feature learning during the early training of network. The feature maps learned with and without the proposed feature loss for a toy network suggest that our approach can fully explore the complementarity between handcrafted features and deep features. Based on the feature loss, a general framework for embedding the traditional feature information into deep network training was developed and tested using the FER2013, CK+, Oulu-CASIA, and MMI datasets. Moreover, adaptive loss weighting strategies are proposed to balance the influence of different losses for different expression databases. The experimental results show that the proposed feature loss with adaptive weighting achieves much better accuracy than the original handcrafted feature and the network trained without using our feature loss. Meanwhile, the feature loss with adaptive weighting can provide complementary information to compensate for the deficiency of a single feature. Weicheng Xie 0001, LinLin Shen, Jinming Duan 0001 |
IEEE Trans. Cybern. | 3 |
| 2020 | G2L-Net: Global to Local Network for Real-Time 6D Pose Estimation With Embedding Vector FeaturesabstractIn this paper, we propose a novel real-time 6D object pose estimation framework, named G2L-Net. Our network operates on point clouds from RGB-D detection in a divide-and-conquer fashion. Specifically, our network consists of three steps. First, we extract the coarse object point cloud from the RGB-D image by 2D detection. Second, we feed the coarse object point cloud to a translation localization network to perform 3D segmentation and object translation prediction. Third, via the predicted segmentation and translation, we transfer the fine object point cloud into a local canonical coordinate, in which we train a rotation localization network to estimate initial object rotation. In the third step, we define point-wise embedding vector features to capture viewpoint-aware information. To calculate more accurate rotation, we adopt a rotation residual estimator to estimate the residual between initial rotation and ground truth, which can boost initial pose estimation performance. Our proposed G2L-Net is real-time despite the fact multiple steps are stacked via the proposed coarse-to-fine framework. Extensive experiments on two benchmark datasets show that G2L-Net achieves state-of-the-art performance in terms of both accuracy and speed. Wei Chen 0092, Xi Jia, Hyung Jin Chang, Jinming Duan 0001, Ales Leonardis |
CVPR | 4 |
| 2020 | Geometry Constrained Weakly Supervised Object Localization
Weizeng Lu, Xi Jia, Weicheng Xie 0001, LinLin Shen, Yicong Zhou, Jinming Duan 0001 |
ECCV (26) | 6 |
| 2020 | One-stage Multi-task Detector for 3D Cardiac MR ImagingabstractFast and accurate landmark location and bounding box detection are important steps in 3D medical imaging. In this paper, we propose a novel multi-task learning framework, for real-time, simultaneous landmark location and bounding box detection in 3D space. Our method extends the famous single-shot multibox detector (SSD) from single-task learning to multitask learning and from 2D to 3D. Furthermore, we propose a post-processing approach to refine the network landmark output, by averaging the candidate landmarks. Owing to these settings, the proposed framework is fast and accurate. For 3D cardiac magnetic resonance (MR) images with size 224×224×64, our framework runs ~128 volumes per second (VPS) on GPU and achieves 6.75mm average point-to-point distance error for landmark location, which outperforms both state-of-the-art and baseline methods. We also show that segmenting the 3D image cropped with the bounding box results in both improved performance and efficiency. Weizeng Lu, Xi Jia, Wei Chen 0092, Nicolò Savioli, Antonio M. Simoes Monteiro de Marvao, LinLin Shen, Declan P. O'Regan, Jinming Duan 0001 |
ICPR | 8 |
| 2020 | PointPoseNet: Point Pose Network for Robust 6D Object Pose EstimationabstractIn this paper, we propose a novel pipeline to estimate 6D object pose from RGB-D images of known objects present in complex scenes. The pipeline directly operates on raw point clouds extracted from RGB-D scans. Specifically, our method takes the point cloud as input and regresses the point-wise unit vectors pointing to the 3D keypoints. We then use these vectors to generate keypoint hypotheses from which the 6D object pose hypotheses are computed. Finally, we select the best 6D object pose from the hypotheses based on a proposed scoring mechanism with geometry constraints. Extensive experiments show that the proposed method is robust against the variety in object shape and appearance as well as occlusions between objects, and that our method outperforms the state-of-the-art methods on the LINEMOD and Occlusion LINEMOD datasets. Wei Chen 0092, Jinming Duan 0001, Hector Basevi, Hyung Jin Chang, Ales Leonardis |
WACV | 2 |
| 2020 | Efficient image structural similarity quality assessment method using image regularised featureabstractImage regularised features play a critical role in image processing domain, by integrating regularised feature and structural similarity, a new full‐reference image assessment method (IRF_SSIM) is proposed in this study. As well known, the gradient operator always be used to capture the edge information of the image, while the total variational regularised features can be adopted to calculate the detailed change information of image contrast and texture, as well as noise removal and edge retention. Therefore, the IRF_SSIM method extends the gradient features into the image regularised features to measure the structural changes in the image. In addition, image quality is also affected by variations of luminance and contrast. For a more comprehensive image quality assessment, the IRF_SSIM method considers the changes in structure, luminance and contrast simultaneously. In other words, the total image quality is estimated by structural similarity calculated by integrating the effects of image structure, luminance and contrast changes. Comparing with the representative methods, the experimental results illustrate that the IRF_SSIM method is highly consistent with the subjective assessment results. Baoxiang Huang, Huan Yang 0001, Guojia Hou, Jinming Duan 0001 |
IET Image Process. | 6 |
| 2020 | Explainable Anatomical Shape Analysis Through Deep Hierarchical Generative ModelsabstractQuantification of anatomical shape changes currently relies on scalar global indexes which are largely insensitive to regional or asymmetric modifications. Accurate assessment of pathology-driven anatomical remodeling is a crucial step for the diagnosis and treatment of many conditions. Deep learning approaches have recently achieved wide success in the analysis of medical images, but they lack interpretability in the feature extraction and decision processes. In this work, we propose a new interpretable deep learning model for shape analysis. In particular, we exploit deep generative networks to model a population of anatomical segmentations through a hierarchy of conditional latent variables. At the highest level of this hierarchy, a two-dimensional latent space is simultaneously optimised to discriminate distinct clinical conditions, enabling the direct visualisation of the classification space. Moreover, the anatomical variability encoded by this discriminative latent space can be visualised in the segmentation space thanks to the generative properties of the model, making the classification task transparent. This approach yielded high accuracy in the categorisation of healthy and remodelled left ventricles when tested on unseen segmentations from our own multi-centre dataset as well as in an external validation set, and on hippocampi from healthy controls and patients with Alzheimer's disease when tested on ADNI data. More importantly, it enabled the visualisation in three-dimensions of both global and regional anatomical features which better discriminate between the conditions under exam. The proposed approach scales effectively to large populations, facilitating high-throughput analysis of normal anatomy and pathology in large-scale studies of volumetric imaging. Carlo Biffi, Juan J. Cerrolaza, Giacomo Tarroni, Wenjia Bai, Antonio M. Simoes Monteiro de Marvao, Ozan Oktay, Christian Ledig, Loïc Le Folgoc, Konstantinos Kamnitsas, Georgia Doumou, Jinming Duan 0001, Sanjay K. Prasad, Stuart A. Cook, Declan P. O'Regan, Daniel Rueckert |
IEEE Trans. Medical Imaging | 11 |
| 2019 | Local Normalization Based BN Layer Pruning
Xi Jia, LinLin Shen, Zhong Ming 0001, Jinming Duan 0001 |
ICANN (2) | 5 |
| 2019 | Outlier-Suppressed Triplet Loss with Adaptive Class-Aware Margins for Facial Expression RecognitionabstractTriplet loss has been proposed to increase the inter-class distance and decrease the intra-class distance for various tasks of image recognition. However, for facial expression recognition (FER) problem, the fixed margin parameter does not fit the diversity of scales between different expressions. Meanwhile, the strategy of selecting the hardest triplets can introduce noisy guidance information since various persons may present significantly different expressions. In this work, we propose a new triplet loss based on class-aware margins and outlier-suppressed triplet for FER, where each pair of expressions, e.g. 'happy' and 'fear', is assigned with an adaptive margin parameter and the abnormal hard triplets are discarded according to the feature distance distribution. Experimental results of the proposed triplet loss on the FER2013 and CK+ expression databases show that the proposed network achieves much better accuracy than the original triplet loss and the network without using the proposed strategies, and competitive performance compared with the state-of-the-art algorithms. Zhiwei Wen, Weicheng Xie 0001, LinLin Shen, Jinming Duan 0001 |
ICIP | 6 |
| 2019 | VS-Net: Variable Splitting Network for Accelerated Parallel MRI Reconstruction
Jinming Duan 0001, Jo Schlemper, Chen Qin, Cheng Ouyang, Wenjia Bai, Carlo Biffi, Ghalib Bello, Ben Statton, Declan P. O'Regan, Daniel Rueckert |
MICCAI (4) | 1 |
| 2019 | Self-Supervised Learning for Cardiac MR Image Segmentation by Anatomical Position Prediction
Wenjia Bai, Chen Chen 0042, Giacomo Tarroni, Jinming Duan 0001, Florian Guitton, Steffen E. Petersen, Yike Guo, Paul M. Matthews, Daniel Rueckert |
MICCAI (2) | 4 |
| 2019 | Data Efficient Unsupervised Domain Adaptation For Cross-modality Image Segmentation
Cheng Ouyang, Konstantinos Kamnitsas, Carlo Biffi, Jinming Duan 0001, Daniel Rueckert |
MICCAI (2) | 4 |
| 2019 | k-t NEXT: Dynamic MR Image Reconstruction Exploiting Spatio-Temporal Correlations
Chen Qin, Jo Schlemper, Jinming Duan 0001, Gavin Seegoolam, Anthony N. Price, Joseph V. Hajnal, Daniel Rueckert |
MICCAI (2) | 3 |
| 2019 | An efficient nonlocal variational method with application to underwater image restoration
Guojia Hou, Zhenkuan Pan 0001, Guodong Wang 0001, Huan Yang 0001, Jinming Duan 0001 |
Neurocomputing | 5 |
| 2019 | Undersampled CS image reconstruction using nonconvex nonsmooth mixed constraints
Ryan Wen Liu, Lin Shi 0001, Jinming Duan 0001, Simon C. H. Yu, Defeng Wang |
Multim. Tools Appl. | 4 |
| 2019 | Parameter-Free Selective Segmentation With Convex Variational MethodsabstractSelective segmentation methods involve incorporating user input to partition an image into a foreground and background. These methods are often sensitive to some aspect of the user input in a counter intuitive manner, making their use in practice difficult. The most robust methods often involve laborious refinement on the part of the user, and sometimes editing/supervision. The proposed method reduces the burden of the user by simplifying the requirements in the input. Specifically, the fitting term does not depend on a distance function, and so no selection parameter is introduced. Instead, we consider how the user input relates to some general intensity fitting term to ensure the approach is less sensitive to the decisions or intuition of the user. We give comparisons to existing approaches to show the advantages of the new selective segmentation model. Jack A. Spencer, Ke Chen 0002, Jinming Duan 0001 |
IEEE Trans. Image Process. | 3 |
| 2019 | Automatic 3D Bi-Ventricular Segmentation of Cardiac Images by a Shape-Refined Multi- Task Deep Learning ApproachabstractDeep learning approaches have achieved state-of-the-art performance in cardiac magnetic resonance (CMR) image segmentation. However, most approaches have focused on learning image intensity features for segmentation, whereas the incorporation of anatomical shape priors has received less attention. In this paper, we combine a multi-task deep learning approach with atlas propagation to develop a shape-refined bi-ventricular segmentation pipeline for short-axis CMR volumetric images. The pipeline first employs a fully convolutional network (FCN) that learns segmentation and landmark localization tasks simultaneously. The architecture of the proposed FCN uses a 2.5D representation, thus combining the computational advantage of 2D FCNs networks and the capability of addressing 3D spatial consistency without compromising segmentation accuracy. Moreover, a refinement step is designed to explicitly impose shape prior knowledge and improve segmentation quality. This step is effective for overcoming image artifacts (e.g., due to different breath-hold positions and large slice thickness), which preclude the creation of anatomically meaningful 3D cardiac shapes. The pipeline is fully automated, due to network's ability to infer landmarks, which are then used downstream in the pipeline to initialize atlas propagation. We validate the pipeline on 1831 healthy subjects and 649 subjects with pulmonary hypertension. Extensive numerical experiments on the two datasets demonstrate that our proposed method is robust and capable of producing accurate, high-resolution, and anatomically smooth bi-ventricular 3D models, despite the presence of artifacts in input CMR volumes. Jinming Duan 0001, Ghalib Bello, Jo Schlemper, Wenjia Bai, Timothy Dawes, Carlo Biffi, Antonio M. Simoes Monteiro de Marvao, Georgia Doumou, Declan P. O'Regan, Daniel Rueckert |
IEEE Trans. Medical Imaging | 1 |
| 2018 | Deep Nested Level Sets: Fully Automated Segmentation of Cardiac MR Images in Patients with Pulmonary Hypertension
Jinming Duan 0001, Jo Schlemper, Wenjia Bai, Timothy Dawes, Ghalib Bello, Georgia Doumou, Antonio M. Simoes Monteiro de Marvao, Declan P. O'Regan, Daniel Rueckert |
MICCAI (4) | 1 |
| 2018 | Cardiac MR Segmentation from Undersampled k-space Using Deep Latent Representation Learning
Jo Schlemper, Ozan Oktay, Wenjia Bai, Daniel C. Castro, Jinming Duan 0001, Chen Qin, Joseph V. Hajnal, Daniel Rueckert |
MICCAI (1) | 5 |
| 2018 | Open snake model based on global guidance field for embryo vessel locationabstractThe development of vessels can provide important information about the growth status of animal embryos. It is, therefore, important to automatically locate the deformed vessel branches from the embryo images. However, very few vessel detectors can accurately locate all vessel branches when the captured images are low quality and the implied vessel shapes are complex. In this study, a new framework consisting of vessel region extraction and snake shape optimisation is proposed. The main contribution in this detector is a novel open snake model based on the global guidance field and deformation template initialisation. Experimental results on a specific application of an embryo vessel database [Database and source codes: https://github.com/wcxie/Egg‐embryro‐vessel‐location/ .] demonstrate that the proposed algorithm not only locates the vessel shape properly but also obtains the orientations of embryo vessel branches accurately. Comparison to traditional guidance fields and the active appearance model illustrates the effectiveness and competitiveness of the proposed model. Weicheng Xie 0001, Jinming Duan 0001, LinLin Shen, Yuexiang Li, Meng Yang 0001, Guojun Lin |
IET Comput. Vis. | 2 |
| 2017 | Single-image blind deblurring with hybrid sparsity regularizationabstractSingle-image blind deblurring could be considered as an important preprocessing step in imaging information fusion. Its purpose is to simultaneously estimate blur kernel and latent sharp image from only one observed blurred image. Blind deblurring has been attracting increasing attention in the fields of image processing, computer vision, computational photography, etc. However, it is a typically ill-posed inverse problem, which requires regularization methods to guarantee stable image restoration results. We first proposed to robustly estimate the blur kernels by exploiting non-convex sparsity constraints on image gradients and blur kernels. The corresponding combined non-convex regularization term has the capacity of enhancing estimation accuracy. To guarantee the high-quality non-blind deblurring with estimated blur kernels, the hybrid non-convex first- and second-order TV regularizer was then introduced to stabilize the final image restoration process. The hybrid non-convex regularizer is able to achieve a good balance between sharp edges preservation and undesirable artifacts suppression. The resulting non-convex minimization problems related to blur kernel estimation and non-blind deblurring were handled using efficient numerical optimization algorithms in this paper. Numerous experiments on both synthetic and realistic images have demonstrated the good performance of the proposed blind deblurring method. Ryan Wen Liu, Jinming Duan 0001, Tian Xu 0001, Jingxian Liu |
FUSION | 4 |
| 2017 | Automated segmentation of retinal layers from optical coherence tomography images using geodesic distance
Jinming Duan 0001, Christopher R. Tench, Irene Gottlob, Frank Proudlock, Li Bai 0001 |
Pattern Recognit. | 1 |
| 2017 | Novel Methods for Microglia Segmentation, Feature Extraction, and ClassificationabstractSegmentation and analysis of histological images provides a valuable tool to gain insight into the biology and function of microglial cells in health and disease. Common image segmentation methods are not suitable for inhomogeneous histology image analysis and accurate classification of microglial activation states has remained a challenge. In this paper, we introduce an automated image analysis framework capable of efficiently segmenting microglial cells from histology images and analyzing their morphology. The framework makes use of variational methods and the fast-split Bregman algorithm for image denoising and segmentation, and of multifractal analysis for feature extraction to classify microglia by their activation states. Experiments show that the proposed framework is accurate and scalable to large datasets and provides a useful tool for the study of microglial biology. Yuchun Ding, Marie Christine Pardon, Alessandra Agostini, Henryk Faas, Jinming Duan 0001, Wil O. C. Ward, Felicity Easton, Dorothee Auer, Li Bai 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 5 |
| 2015 | Second order Mumford-Shah model for image denoisingabstractA second order Mumford-Shah model is proposed for image denoising. Unlike the original Mumford-Shah model, the proposed new model uses second order derivatives defined in bounded Hessian space as its regulariser. This model is capable of eliminating the undesirable staircase effect associated with the original Mumford-Shah model with a total variation regulariser. Unlike other second order models that use bounded Hessian regulariser, the proposed new model does not blur the edges in the restored image. To improve computational efficiency, the implementation of the proposed model does not directly solve the high order nonlinear partial differential equations and instead exploit the efficient split Bregman algorithm, which uses the fast Fourier transform. Numerical experiments are conducted to compare the performance of the new model in image denoising with those of the original Mumford-Shah model and the pure second order model. Jinming Duan 0001, Yuchun Ding, Zhenkuan Pan 0001, Jie Yang 0002, Li Bai 0001 |
ICIP | 1 |
| 2015 | Optical coherence tomography image segmentationabstractOptical coherence tomography (OCT) is a three-dimensional non-invasive imaging technique that can generate images of the eye at microscopic level to help diagnosis of eye diseases. However, OCT images often suffer from inhomogeneity and are corrupted by speckle noise, posing challenges to automated OCT image segmentation and analysis. In this paper, a novel method is proposed to segment retinal layers in OCT images. The proposed method uses a coarse-to-fine approach to segmentation and includes: (1) A variational retinex model that can enhance the details as well as correct the intensity in-homogeneities; (2) An anisotropic coherent enhancing diffusion that can remove speckle noise and simultaneously connect the interrupted retinal layers; (3) A nonlinear isotropic filter that smooths the processed OCT image and leads to the initial coarse segmentation; (4) A high-pass unsharp masking filter that highlights the remaining layers and gives the fine segmentation. Afterwards, all the retinal layers that can be seen by human eyes are segmented accurately using common edge detection of region based segmentation algorithms. Extensive experiments results validate the effectiveness and performance of the proposed method. Jinming Duan 0001, Christopher R. Tench, Irene Gottlob, Frank Proudlock, Li Bai 0001 |
ICIP | 1 |
| 2015 | Fast algorithm for color texture image inpainting using the non-local CTV model
Jinming Duan 0001, Zhenkuan Pan 0001, Baochang Zhang 0001, Wanquan Liu, Xue-Cheng Tai |
J. Glob. Optim. | 1 |