EDBT 2026 Demo / reviewers in the wild / expert
Min Liu 0008
dblp:99/76-8
· DBLP profile ↗
106ranked-venue papers
27as first author
73since 2021 · last 2026
0000-0001-6406-4896ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 45 · 10 first-author · 29 since 2021Applied, interdisciplinary, general and emerging computing · 36 · 9 first-author · 27 since 2021Artificial intelligence and machine learning · 27 · 9 first-author · 16 since 2021Databases, data management, data science and information retrieval · 4 · 4 since 2021Security and privacy · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Mono3DVG-EnSD: Enhanced Spatial-aware and Dimension-decoupled Text Encoding for Monocular 3D Visual GroundingabstractMonocular 3D Visual Grounding (Mono3DVG) is an emerging task that locates 3D objects in RGB images using text descriptions with geometric cues. However, existing methods face two key limitations. Firstly, they often over-rely on high-certainty keywords that explicitly identify the target object while neglecting critical spatial descriptions. Secondly, generalized textual features contain both 2D and 3D descriptive information, thereby capturing an additional dimension of details compared to singular 2D or 3D visual features. This characteristic leads to cross-dimensional interference when refining visual features under text guidance. To overcome these challenges, we propose Mono3DVG-EnSD, a novel framework that integrates two key components: the CLIP-Guided Lexical Certainty Adapter (CLIP-LCA) and the Dimension-Decoupled Module (D2M). The CLIP-LCA dynamically masks high-certainty keywords while retaining low-certainty implicit spatial descriptions, thereby forcing the model to develop a deeper understanding of spatial relationships in captions for object localization. Meanwhile, the D2M decouples dimension-specific (2D/3D) textual features from generalized textual features to guide corresponding visual features at same dimension, which mitigates cross-dimensional interference by ensuring dimensionally-consistent cross-modal interactions. Through comprehensive comparisons and ablation studies on the Mono3DRefer dataset, our method achieves state-of-the-art (SOTA) performance across all metrics. Notably, it improves the challenging Far([email protected]) scenario by a significant +13.54%. Min Liu 0008, Zhaoyang Li 0011, Yuan Bian 0002, Erbo Zhai, Yaonan Wang 0001 |
AAAI | 2 |
| 2026 | Beyond the LUMIR challenge: The pathway to foundational registration models
Junyu Chen 0002, Shuwen Wei, Joel Honkamaa, Pekka Marttinen, Hang Zhang 0010, Min Liu 0008, Yichao Zhou 0002, Zuopeng Tan, Yi Wang 0028, Hongchao Zhou, Shunbo Hu, Yi Zhang 0120, Lukas Förner, Thomas Wendler 0001, Bailiang Jian, Benedikt Wiestler, Tim Hable, Dan Ruan, Frederic Madesta, Thilo Sentker, Wiebke Heyer, Lianrui Zuo, Yuwei Dai, Jerry L. Prince, Harrison X. Bai, Yong Du 0002, Yihao Liu 0003, Alessa Hering, Reuben Dorent, Lasse Hansen, Mattias P. Heinrich, Aaron Carass |
Medical Image Anal. | 6 |
| 2026 | ZUMA: Training-Free Zero-Shot Unified Multimodal Anomaly DetectionabstractMultimodal anomaly detection (MAD) aims to exploit both texture and spatial attributes to identify deviations from normal patterns in complex scenarios. However, zero-shot (ZS) settings arising from privacy concerns or confidentiality constraints present significant challenges to existing MAD methods. To address this issue, we introduce ZUMA, a training-free, Zero-shot Unified Multimodal Anomaly detection framework that unleashes CLIP's cross-modal potential to perform ZS MAD. To mitigate the domain gap between CLIP's pretraining space and point clouds, we propose cross-domain calibration (CDC), which efficiently bridges the manifold misalignment through source-domain semantic transfer and establishes a hybrid semantic space, enabling a joint embedding of 2D and 3D representations. Subsequently, ZUMA performs dynamic semantic interaction (DSI) to enable structural decoupling of anomaly regions in the high-dimensional embedding space constructed by CDC, where natural languages serve as semantic anchors to help DSI establish discriminative hyperplanes within hybrid modality representations. Within this framework, ZUMA enables plug-and-play detection of 2D, 3D or multimodal anomalies, without training or fine-tuning even for cross-dataset or incomplete-modality scenarios. Additionally, to further investigate the potential of the training-free ZUMA within the training-based paradigm, we develop ZUMA-FT, a fine-tuned variant that achieves notable improvements with minimal parameter trade-off. Extensive experiments are conducted on two MAD benchmarks, MVTec 3D-AD and Eyecandies. Notably, the training-free ZUMA achieves state-of-the-art (SOTA) performance on both datasets, outperforming existing ZS MAD methods, including training-based approaches. Moreover, ZUMA-FT further extends the performance boundary of ZUMA with only 6.75 M learnable parameters. Yunfeng Ma, Min Liu 0008, Jingyu Zhou, Yuan Bian 0002, Yaonan Wang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2026 | A multimodal fusion model based on graph convolutional network for 3D neural morphology optimization
Hongji Qiu, Zhao Yao, Yaonan Wang 0001, Min Liu 0008 |
Pattern Recognit. | 5 |
| 2026 | Encoder-Only Image RegistrationabstractLearning-based techniques have significantly improved the accuracy and speed of deformable image registration. However, challenges such as reducing computational complexity and handling large deformations persist. To address these challenges, we analyze how convolutional neural networks (ConvNets) influence registration performance using the Horn-Schunck optical flow equation. Supported by prior studies and our empirical experiments, we observe that ConvNets play two key roles in registration: linearizing local intensities and harmonizing global contrast variations. Guided by these insights, we propose the Encoder-Only Image Registration (EOIR) framework comprising five modifications to existing approaches, to achieve a better accuracy-efficiency trade-off. EOIR separates feature learning from flow estimation, employing only a 3-layer ConvNet for feature extraction and a set of 3-layer flow estimators to construct a Laplacian feature pyramid, progressively composing diffeomorphic deformations under a large-deformation model. Results on six datasets across different modalities and anatomical regions demonstrate EOIR’s effectiveness, achieving superior accuracy-efficiency and accuracy-smoothness trade-offs. With comparable accuracy, EOIR provides better efficiency and smoothness, and vice versa. The source code of EOIR is available on Github. Xiang Chen 0008, Renjiu Hu, Min Liu 0008, Yaonan Wang 0001, Hang Zhang 0010 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2026 | Exploring Volume Representation Similarity in Long-Tail Biased Stereo Matching
Renjie Ding, Yaonan Wang 0001, Min Liu 0008, Jiazheng Wang 0001, Wenting Shen, Zhe Zhang 0022, Xiang Chen 0008 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2026 | Beijing Institute of TechnologyCMANet: A TCN-RMamba-Attention Network for Surgical Phase Online Recognition
Wenpei Fan, Yaonan Wang 0001, Licheng Liu, Min Liu 0008 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2026 | UniSurg: A Unified Multitask Framework for Robotic Surgical Scene UnderstandingabstractSurgical scene understanding is a vital intelligent technique in robot-assisted surgery, including surgical instrument detection, segmentation, and instrument–tissue interaction detection. Existing methods typically address these tasks in isolation, neglecting the intrinsic correlations among them. In this work, we innovatively propose a unified multitask framework named UniSurg, being the first to jointly address these three critical aspects of surgical scene understanding, thereby providing the robot with multidimensional perceptual capabilities. By exploring the inter-task correlations and reusing shared features, UniSurg has been demonstrated to significantly enhance the scene analysis performance. To address pose variability of the instruments under the constrained field of view in laparoscopic surgery, we design an Attention Enhanced Conditional Convolution (AEC-Conv) that dynamically adjusts kernels based on pose-specific features for improved adaptability. To further enhance interaction detection, we propose the Temporal Difference Enhancement module (TDE), which captures motion cues by amplifying inter-frame differences, and the Pyramid Global Feature Enhancement module (PGFE), which leverages graph-based hierarchical context to model global relational dependencies. Experiments on the Endovis2018 dataset and a clinical multitask dataset MILVis demonstrate the superior multitask performance of UniSurg. Wenting Shen, Yaonan Wang 0001, Min Liu 0008, Jiazheng Wang 0001, Renjie Ding |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2026 | High-Precision Multi-Instance Registration for Stacked Objects in Bin-Picking ScenesabstractIn industrial bin-picking, robotic systems must estimate the poses of multiple object instances, where accurate pose estimation is essential for reliable downstream manipulation and grasping. Most existing multi-instance registration methods primarily establish point correspondences based on local features to alleviate the challenges posed by occlusion and clutter. However, local features are easily disturbed by neighboring instances and lack global context, leading to unreliable correspondences and degraded registration accuracy. In addition, the absence of rotational invariance further reduces correspondence accuracy in scenes with stacked instances and highly varying object orientations. To address these challenges, we present a one-stage multi-instance point cloud registration framework for stacked-object scenes. Our framework incorporates a rotation-invariant operator to enhance the robustness of feature representations under arbitrary orientations. Then, we propose a Center-Aware Res-Masked Transformer module, which incorporates an object center embedding to enrich global instance-level context and a center-aware residual mask prediction module to balance weight distribution across objects of varying sizes during training. Extensive experiments on the challenging ROBI dataset demonstrate that our method outperforms the competitive baseline MIRETR by more than 10% in mean precision, highlighting its effectiveness in complex bin-picking scenes. Furthermore, evaluations on the unstacked Scan2CAD dataset confirm the generalizability of the proposed framework across different application scenarios. Jiawen Zhao, Qing Zhu 0003, Yaonan Wang 0001, Weixing Peng, Jianxu Mao, Min Liu 0008, Xuebing Liu, Hui Zhang 0023 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2026 | SCAP: Semantic Prototype Alignment for Robust Point Cloud RegistrationabstractPoint cloud rigid registration is a fundamental problem in robotics, 3D reconstruction, and augmented reality. However, existing methods predominantly rely on local geometric neighborhoods, which fail to capture higher-order semantic structures and thus degrade performance under noisy or complex geometry conditions. To address these limitations, we propose SCAP, a new point cloud registration paradigm that transforms feature interaction from geometry-driven to semantics–geometric co-driven. Specifically, a semantic prototype extractor is devised to abstract high-level semantic prototypes through graph embedding and clustering, thereby mitigating sensitivity to local feature noise. Since semantic abstraction alone cannot guarantee consistent correspondences across point clouds, SCAP performs a prototype alignment path learning to infer reliable semantic mappings through optimal transport. To enhance cross-layer feature integration and prevent redundant attention, an alignment-driven cross-layer transformer is proposed to incorporate the learned priors into the attention mechanism, thereby enabling feature aggregation with improved semantic coherence and local precision. Extensive experiments on ModelNet, ModelLoNet, 3DMatch, and 3DLoMatch demonstrate that our SCAP consistently surpasses state-of-the-art approaches, showing superior robustness and generalization in challenging scenarios with noise and partial overlap. The code will be available at https://github.com/Zhou-111jy/SCAP.git. Jingyu Zhou, Yunfeng Ma, Yaonan Wang 0001, Min Liu 0008 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2026 | Cross-View Dynamic Learning-Based Multi-Class Industrial Anomaly DetectionabstractIndustrial anomaly detection plays a crucial role in smart manufacturing. Traditional methods typically train separate models for each category, leading to substantial memory demands and computational cost. Moreover, relying solely on single-view images is prone to detection blind spots and poor sensitivity to subtle defects. To address these problems, this study proposes CVDL, a cross-view dynamic learning-based multi-class industrial anomaly detection method. Specifically, the CVDL leverages a proposed cross-view dynamic attention in conjunction with intra-view self-attention to dynamically modulate the model’s attention on multi-view information, thereby enhancing the detection performance of subtle defects. Furthermore, a category-guided prompt is developed to utilize object category information, which improves the model’s class-aware detection accuracy. To enhance the model’s robustness, we introduce a structured noise injection strategy and a region-wise mask into the CVDL, mitigating the “identity shortcut” that preserves anomalies during reconstruction. Extensive experiments on the authentic multi-view industrial datasets (Real-IAD) and well-known datasets (MVTec-AD and VisA) confirm the superior detection capability and robustness of the proposed CVDL, and the overall performance of CVDL is superior to all advanced approaches on Real-IAD, achieving SoTA performance of 90.1% image-level and 99.0% pixel-level AUROC. The code will be available at https://github.com/zfinn1/CVDL.git. Jingyu Zhou, Yunfeng Ma, Yaonan Wang 0001, Min Liu 0008 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2026 | PANDA: Progressive Adaptive Network for Defect-Aware Few-Shot SegmentationabstractFew-shot semantic segmentation aims to reduce reliance on dense annotations, while enhancing generalization to unseen categories. However, most methods are constrained by static global prototypes, which fail to represent subtle defect details, resulting in pronounced support–query misalignment. To address this issue, we propose progressive adaptive network for defect-aware few-shot segmentation (PANDA), a few-shot segmentation framework that integrates representational modulation with semantic consistency constraints. Specifically, to capture subtle and scale-sensitive variations in defect patterns, we design anchored representational modulation (ARM), which overcomes the rigidity of static prototypes by dynamically adjusting representations. In addition, we develop hierarchical semantic coherence (HSC), which enforces consistency across representation hierarchies to suppress the accumulation of semantic drift as depth increases. Collectively, ARM and HSC mitigate support–query misalignment and stabilize representations in few-shot defect segmentation. PANDA achieves state-of-the-art performance on MetFS-18, with 55.1% and 56.3% mean intersection over union under the one-shot and five-shot settings. Moreover, PANDA has been integrated into a real-time industrial inspection platform, where it delivers accurate segmentation across diverse defect types, highlighting its robustness in practical application. Yunfeng Ma, Min Liu 0008, Xiangfei Meng, Yaonan Wang 0001 |
IEEE Trans. Ind. Informatics | 4 |
| 2026 | Unified Multimodal Industrial Anomaly Detection via Few Normal SamplesabstractMultimodal industrial anomaly detection (MIAD) is the process of integrating multiple sensor data and utilizing visual intelligence to identify abnormal states in industrial production. In this article, we focus on two main practical but challenging issues in MIAD, i.e., a unified model for multiclass anomaly detection, and model training with only few normal samples. The current mainstream “one-for-one” paradigm requires training time that grows exponentially, and it relies on a sufficient number of samples (even just normal samples), which cannot adapt to practical industrial scenarios with rich abnormal classes. To this end, we offer aUnifiedMIAD model that trained using onlyFew (e.g., 1, 2, and 4) normal samples, termed UniMF. Specifically, we propose a fusion-guided prompt engineering process that generates paired antithetical instance-specific prompts with the assistance of multimodal fusion at both query and token levels. To enable cross-modal prompt learning under multimodal conditions, UniMF performs multi-proxy pairwise matching that involves alignment among multimodal feature patches, embeddings, and tokens of antithetical prompts. Experimental results show that UniMF stands state-of-the-art performance while remaining “one-for-all” paradigm, and even outperforms “one-for-one” methods under certain settings. Cross-dataset evaluation between MVTec 3D-AD and Eyecandies datasets also shows the transferability of UniMF. Yunfeng Ma, Jingyu Zhou, Yaonan Wang 0001, Min Liu 0008 |
IEEE Trans. Ind. Informatics | 5 |
| 2026 | Micro Surface Defect Inspection of Aero-Engine Blades via Dynamic Cross-Scale Semantic Aggregation
Kaijie Li, Jingyu Zhou, Xiangfei Meng, Yaonan Wang 0001, Min Liu 0008 |
IEEE Trans. Ind. Informatics | 8 |
| 2026 | Information-Bottleneck-Guided Hybrid Neural Architecture Search for Temporal Action Detection in Untrimmed VideosabstractTemporal Action Detection (TAD) in untrimmed videos requires effective spatial feature extraction for precise action classification and temporal feature modeling for accurate boundary localization. To achieve effective spatio-temporal feature integration, several works manually design rule-based (i.e., sequential or parallel) hybrid Mamba-Transformer networks for TAD. However, few studies explore diverse integration strategies and network topologies due to the inherent limitations of manual design. Therefore, we propose NAS-TAD, the first Neural Architecture Search framework for TAD, systematically exploring this untouched problem. Specifically, we develop a spatio-temporal NAS objective function based on information-bottleneck theory to quantify task-relevant spatio-temporal features, providing interpretable guidance for the network search and optimization process. Furthermore, we reformulate Transformer self-attention as a state-space model, thereby enabling seamless switching between Mamba and Transformer blocks in a unified weight-sharing search space. Consequently, comprehensive experiments on ActivityNet, THUMOS14, HACS and FineAction demonstrate the effectiveness of the searched hybrid architectures, providing new insights into temporal and spatial feature fusion for TAD. Code is available for reproduction at https://github.com/tyhnu/nastad.git. Mansen Chen, Lepeng Chen, Min Liu 0008, Yaonan Wang 0001 |
IEEE Trans. Image Process. | 5 |
| 2026 | Semantic-Aware Multimodal Collaborative Learning for Unsupervised Visible-Infrared Person Re-IdentificationabstractUnsupervised visible-infrared person re-identification (VI-ReID) is challenging due to the significant modality gap between visible and infrared images. Most existing methods rely on one-hot clustering pseudo-labels as supervision signals, which often fail to capture the full semantic relationships among samples and are highly susceptible to noise. To address these limitations, we propose a Semantic-aware Multimodal Collaborative Learning (SAMCL) framework for unsupervised VI-ReID. Specifically, a Modality-aware Semantic Fusion (MSF) module is designed to bridge the inter-modality gap by integrating complementary semantic details from both visible and infrared modalities, generating enriched cross-modal supervision signals, for cross-modal collaborative learning. Meanwhile, we present a Dynamic Contrastive Learning (DCL) module to refine intra-modality feature learning by dynamically aligning samples with their neighboring centroids in the feature space, improving clustering reliability and intra-modality feature discrimination. By combining the two modules, SAMCL harnesses multimodal collaboration, minimizes dependence on noisy pseudo-labels, and provides a robust approach to unsupervised VI-ReID. Extensive experiments demonstrate the superiority of our proposed method. For instance, on the SYSU-MM01 dataset, our model achieves a Rank-1 accuracy of 68.68% in the All Search setting, surpassing the state-of-the-art (SOTA) by 3.48%. On the RegDB dataset, it achieves a Rank-1 accuracy of 94.47% in the Visible-to-Infrared setting, outperforming the SOTA by 3.57%. On the LLCM dataset, it achieves a Rank-1 accuracy of 50.6% in the Visible-to-Infrared setting, outperforming the SOTA by 3.7%. The code is available at https://github.com/luoshixi123/SAMCL. Shixi Luo, Min Liu 0008, Gautam Srivastava 0001, Shuai Liu 0002, Yaonan Wang 0001 |
IEEE Trans. Image Process. | 3 |
| 2026 | Neural Optimization for Image Registration via Joint Modeling of Global Affine and Local Deformation TransformationsabstractConventional registration approaches frequently underperform when applied to sparse feature alignment (e.g., retinal vessels and filamentous collagen fibers in second-harmonic generation (SHG) and bright-field (BF) images), as these tasks demand simultaneous handling of global affine registration and local deformation correction. End-to-end learning-based approaches struggle with minimal effective gradients from loss back-propagation of these sparse features, while descriptor matching methods, though helpful, lack fidelity loss and fail to adapt to local deformation. To address these issues, we propose Neural Affine Optimization (NeOn), which implicitly approximates discrete optimization using a few neural network layers, combined with a sampling-regression layer to handle affine transformations. NeOn allows iterative refinement with fidelity loss and provides a flexible transition between a purely affine configuration and a linear weighted blend of affine and deformation fields. NeOn's performance was validated on four public datasets. In multi-modal SHG-BF microscopy registration, NeOn achieved top rankings on the validation leaderboard for Task 3 of the Learn2Reg Challenge 2024. For retinal image registration, NeOn outperformed existing methods on both mono-modal and multi-modal datasets, reducing target registration error from 6.3 to 2.1 pixels in mono-modal and from 2.6 to 1.8 pixels in multi-modal registration. Furthermore, NeOn demonstrates strong generalization and can be effectively extended to 3D multi-modality image registration scenarios. Xiang Chen 0008, Renjiu Hu, Jiacheng Wang 0001, Min Liu 0008, Yaonan Wang 0001, Jiazheng Wang 0001, Rongguang Wang, Gaolei Li, Hang Zhang 0010 |
IEEE Trans. Medical Imaging | 4 |
| 2026 | CiSeg: Unsupervised Cross-Modality Adaptation for 3D Medical Image Segmentation via Causal InterventionabstractUnsupervised domain adaptation (UDA) addresses the domain shift problem by transferring knowledge from labeled source domain data (e.g. CT) to unlabeled target domain data (e.g. MRI). While state-of-the-art methods reduce domain gaps via image- or feature-level alignment, their reliance on spurious correlations in the training data often limits generalization across domains. To overcome this limitation, we propose the Causal Intervention Segmentation Network (CiSeg), a novel framework that first integrates causal inference into UDA. A Structural Causal Model (SCM) is first constructed for the source domain to disentangle causal variables from bias variables, alleviating the impact of spurious correlations. Based on this SCM, we introduce a Counterfactual Disentanglement (CD) module to decompose the source domain's latent features into distinct causal and bias components, effectively eliminating their mutual dependencies. To enhance cross-domain consistency, two auxiliary components are introduced: Prototype-guided Contrastive Learning (PCL) and Causal-bias Residual Alignment (CBRA). PCL aligns pixel-level representations with their corresponding semantic prototypes, promoting stronger intra-class consistency and clearer inter-class separability. CBRA employs adversarial learning to align causal and bias residual features across domains, further enhancing feature-level invariance. Extensive experiments on cardiac, abdominal multi-organ, and BraTS18 segmentation tasks demonstrate that CiSeg outperforms state-of-the-art methods, achieving superior segmentation performance and robust cross-domain generalization. Code and models are available at https://github.com/lvpeiqing/CiSeg. Peiqing Lv, Yaonan Wang 0001, Min Liu 0008, Zhe Zhang 0022, Yunfeng Ma, Licheng Liu, Erik Meijering |
IEEE Trans. Medical Imaging | 3 |
| 2026 | EPDiff: Erasure Perception Diffusion Model for Unsupervised Anomaly Detection in Preoperative Multimodal ImagesabstractUnsupervised anomaly detection (UAD) methods typically detect anomalies by learning and reconstructing the normative distribution. However, since anomalies constantly invade and affect their surroundings, sub-healthy areas in the junction present structural deformations that could be easily misidentified as anomalies, posing difficulties for UAD methods that solely learn the normative distribution. The use of multimodal images can facilitate to address the above challenges, as they can provide complementary information of anomalies. Therefore, this paper propose a novel method for UAD in preoperative multimodal images, called Erasure Perception Diffusion model (EPDiff). First, the Local Erasure Progressive Training (LEPT) framework is designed to better rebuild sub-healthy structures around anomalies through the diffusion model with a two-phase process. Initially, healthy images are used to capture deviation features labeled as potential anomalies. Then, these anomalies are locally erased in multimodal images to progressively learn sub-healthy structures, obtaining a more detailed reconstruction around anomalies. Second, the Global Structural Perception (GSP) module is developed in the diffusion model to realize global structural representation and correlation within images and between modalities through interactions of high-level semantic information. In addition, a training-free module, named Multimodal Attention Fusion (MAF) module, is presented for weighted fusion of anomaly maps between different modalities and obtaining binary anomaly outputs. Experimental results show that EPDiff improves the AUPRC and mDice scores by 2% and 3.9% on BraTS2021, and by 5.2% and 4.5% on Shifts over the state-of-the-art methods, which proves the applicability of EPDiff in diverse anomaly diagnosis. The code is available at https://github.com/wjiazheng/EPDiff. Jiazheng Wang 0001, Min Liu 0008, Wenting Shen, Renjie Ding, Yaonan Wang 0001, Erik Meijering |
IEEE Trans. Medical Imaging | 2 |
| 2025 | Prompt-Driven Transferable Adversarial Attack on Person Re-identification with Attribute-Aware Textual InversionabstractPerson re-identification (re-id) models are vital in security surveillance systems, requiring transferable adversarial attacks to explore the vulnerabilities of them. Recently, vision-language models (VLM) based attacks have shown superior transferability by attacking generalized image and textual features of VLM, but they lack comprehensive feature disruption due to the overemphasis on discriminative semantics in integral representation. In this paper, we introduce the Attribute-aware Prompt Attack (AP-Attack), a novel method that leverages VLM's image-text alignment capability to explicitly disrupt fine-grained semantic features of pedestrian images by destroying attribute-specific textual embeddings. To obtain personalized textual descriptions for individual attributes, textual inversion networks are designed to map pedestrian images to pseudo tokens that represent semantic embeddings, trained in the contrastive learning manner with images and a predefined prompt template that explicitly describes the pedestrian attributes. Inverted benign and adversarial fine-grained textual semantics facilitate attacker in effectively conducting thorough disruptions, enhancing the transferability of adversarial examples. Extensive experiments show that AP-Attack achieves state-of-the-art transferability, significantly outperforming previous methods by 22.9% on mean Drop Rate in cross-model&dataset attack scenarios. Yuan Bian 0002, Min Liu 0008, Yunqi Yi, Yaonan Wang 0001 |
ICCV | 2 |
| 2025 | Gaussian Primitive Optimized Deformable Retinal Image Registration
Jiazheng Wang 0001, Xiang Chen 0008, Renjiu Hu, Gaolei Li, Min Liu 0008, Hang Zhang 0010 |
MICCAI (4) | 7 |
| 2025 | VoxelOpt: Voxel-Adaptive Message Passing for Discrete Optimization in Deformable Abdominal CT Registration
Hang Zhang 0010, Jiazheng Wang 0001, Xiang Chen 0008, Renjiu Hu, Gaolei Li, Min Liu 0008 |
MICCAI (4) | 8 |
| 2025 | Dual Enhancement on 3D Vision-Language Perception for Monocular 3D Visual GroundingabstractMonocular 3D visual grounding is a novel task that aims to locate 3D objects in RGB images using text descriptions with explicit geometry information. Despite the inclusion of geometry details in the text, we observe that the text embeddings are sensitive to the magnitude of numerical values but largely ignore the associated measurement units. For example, simply equidistant mapping the length with unit 'meters' to 'decimeters' or 'centimeters' leads to severe performance degradation, even though the physical length remains equivalent. This observation signifies the weak 3D comprehension of pre-trained language model, which generates misguiding text features to hinder 3D perception. Therefore, we propose to enhance the 3D perception of model on text embeddings and geometry features with two simple and effective methods. Firstly, we introduce a pre-processing method named 3D-text Enhancement (3DTE), which enhances the comprehension of mapping relationships between different units by augmenting the diversity of distance descriptors in text queries. Next, we propose a Text-Guided Geometry Enhancement (TGE) module to further enhance the 3D-text information by projecting the basic text features into geometrically consistent space. These 3D-enhanced text features are then leveraged to precisely guide the attention of geometry features. We evaluate the proposed method through extensive comparisons and ablation studies on the Mono3DRefer dataset. Experimental results demonstrate substantial improvements over previous methods, achieving new state-of-the-art results with a notable accuracy gain of 11.94% in the 'Far' scenario. Our code will be made publicly available. Min Liu 0008, Yuan Bian 0002, Zhaoyang Li 0011, Gen Li 0008, Yaonan Wang 0001 |
ACM Multimedia | 2 |
| 2025 | Decoupled Identity and Attribute Tokenization for Person Re-IdentificationabstractVision-language models like CLIP have revolutionized person re-identification (ReID) by enabling cross-modal semantic alignment. However, most of the existing CLIP-based ReID methods suffer from a critical limitation: semantic entanglement, where identity and attribute features are indiscriminately compressed into a single, undifferentiated token representation. This oversight fails to account for their inherently distinct roles in characterizing individuals.To address this limitation, we propose an Identity-Attribute-Decoupled Tokenization (IADT) method, a hierarchical framework with two synergistic components:Subject-oriented tokens that model identity through a cross-modality feature inverse mapping paradigm, preserving invariant biometric features;Attribute-aware tokens that capture localized characteristics through the cross-interaction of local features and learnable prototype vectors, dynamically focusing on discriminative regions without manual supervision.The hierarchical tokenization enables disentangled yet complementary representation learning: Identity and attribute semantics are encoded into distinct embedding subspaces, while cross-token contrastive learning establishes semantic reinforcement through attention-guided feature interaction. Crucially, this process does not require part-level annotations, making it directly applicable to real-world deployment. Extensive experiments validate effectiveness of the proposed method. For example, on the Market-1501 dataset, IADT achieves 97.1% mAP (+2.5% over SOTA) and 98.2% Rank-1 accuracy. For the challenging MSMT benchmark, it attains 88.9% mAP (+1.7% improvement) with 93.1% Rank-1 accuracy, demonstrating consistent superiority. The code will be available at https://github.com/llraay/IADT. Min Liu 0008, Yuan Bian 0002, Yaonan Wang 0001 |
ACM Multimedia | 2 |
| 2025 | Searching Efficient Semantic Segmentation Architectures via Dynamic Path SelectionabstractExisting NAS methods for semantic segmentation typically apply uniform optimization to all candidate networks (paths) within a one-shot supernet. However, the concurrent existence of both promising and suboptimal paths often results in inefficient weight updates and gradient conflicts. This issue is particularly severe in semantic segmentation due to its complex multi-branch architectures and large search space, which further degrade the supernet's ability to accurately evaluate individual paths and identify high-quality candidates. To address this issue, we propose Dynamic Path Selection (DPS), a selective training strategy that leverages multiple performance proxies to guide path optimization. DPS follows a stage-wise paradigm, where each phase emphasizes a different objective: early stages prioritize convergence, the middle stage focuses on expressiveness, and the final stage emphasizes a balanced combination of expressiveness and generalization. At each stage, paths are selected based on these criteria, concentrating optimization efforts on promising paths, thus facilitating targeted and efficient model updates. Additionally, DPS integrates a dynamic stage scheduler and a diversity-driven exploration strategy, which jointly enable adaptive stage transitions and maintain structural diversity among selected paths. Extensive experiments demonstrate that, under the same search space, DPS can discover efficient models with strong generalization and superior performance. Yuxi Liu 0019, Min Liu 0008, Yaonan Wang 0001 |
NeurIPS | 2 |
| 2025 | NAS-PED: Neural Architecture Search for Pedestrian DetectionabstractPedestrian detection currently suffers from two issues in crowded scenes: occlusion and dense boundary prediction, making it still challenging in complex real-world scenarios. In recent years, Convolutional Neural Networks (CNN) and Vision Transformers (ViT) have shown their superiorities in addressing these issues, where ViTs capture global feature dependency to infer occlusion parts and CNNs make accurate dense predictions by local detailed features. Nevertheless, limited by the narrow receptive field, CNNs fail to infer occlusion parts, while ViTs tend to ignore local features that are vital to distinguish different pedestrians in the crowd. Therefore, it is essential to combine the advantages of CNN and ViT for pedestrian detection. However, manually designing a specific CNN and ViT hybrid network requires enormous time and resources for trial and error. To address this issue, we propose the first Neural Architecture Search (NAS) framework specifically designed for pedestrian detection named NAS-PED, which automatically designs an appropriate CNNs and ViTs hybrid backbone for the crowded pedestrian detection task. Specifically, we formulate transformers and convolutions with various kernel sizes in the same format, which provides an unconstrained space for diverse hybrid network search. Furthermore, to search for a suitable backbone, we propose an information bottleneck based NAS objective function, which treats the process of NAS as an information extraction process, preserving relevant information and suppressing redundant information from the dense pedestrians in crowd scenes Extensive experiments on CrowdHuman, CityPersons and EuroCity Persons datasets demonstrate the effectiveness of the proposed method. Our NAS-PED obtains absolute gains of 4.0% MR and 1.9% AP over the state-of-the-art (SOTA) pedestrian detection framework on CrowdHuman datasets. For the CityPersons and EuroCity Persons datasets, the searched backbone achieves stable improvement across all three subsets, outperforming some large language-image pre-trained models. Code will be released after acceptance. Min Liu 0008, Baopu Li, Yaonan Wang 0001, Wanli Ouyang |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2025 | DAUNet: A deformable aggregation UNet for multi-organ 3D medical image segmentation
Qinghao Liu, Min Liu 0008, Yuehao Zhu, Licheng Liu, Zhe Zhang 0022, Yaonan Wang 0001 |
Pattern Recognit. Lett. | 2 |
| 2025 | Altering Query Prompting With Contrastive Learning for Multimodal Intent RecognitionabstractMultimodal intent recognition utilizes heterogeneous modalities such as visual, auditory, and textual cues to infer user intent, serving as a pivotal component in human-machine interaction. Existing approaches, however, often rely on unimodal paradigms or shallow multimodal fusion, failing to model cross-modal semantic dependencies and struggling to extract discriminative features from non-verbal modalities, limiting their robustness in complex scenarios. To mitigate these limitations, we propose an Altering Query Prompting with Contrastive Learning framework (AQP-CL) that dynamically aligns and refines multimodal representations. Specifically, the Altering Query Prompting (AQP) module introduces a tri-modality rotation attention mechanism, where textual, visual, and acoustic modalities cyclically alternate as queries in cross-attention operations. This approach addresses modality bias while strengthening interdependencies between modalities, ultimately yielding intent-aware fused feature representations that preserve discriminative cues. The Label-semantic Augmented Contrastive Learning (LACL) strategy generates augmented samples through the intent-aware query prompt and enhances feature discrimination via NT-Xent loss on label tokens. By integrating high-confidence textual semantics from intent labels, LACL refines auxiliary modality features through contrastive alignment, ensuring robust cross-modal representation learning. Evaluations on IEMOCAP and MIntRec validate AQP-CL's superiority, achieving state-of-the-art precision of 77.78% on IEMOCAP, a 3.41% improvement over existing methods. Yuxin Jia, Zhanpeng Shao, Min Liu 0008 |
IEEE Signal Process. Lett. | 4 |
| 2025 | Multi-Context Aggregation Network With Foreground Correction for Automated Few-Shot Defect SegmentationabstractState-of-the-art defect segmentation methods rely on sufficient training data and struggle to generalize to unseen categories. Few-Shot Semantic Segmentation (FSS) is introduced to specifically address these issues. However, existing FSS models still face two challenges in the industry. 1) Defects usually present as weak features, resulting in incomplete segmentation; 2) Severe background interference often leads to incorrect segmentation. To tackle these problems, we propose the Multi-Context Aggregation Network (MCANet). Specifically, we design a Cross-Layer Multi-Level Feature Aggregation Module (CMAM). CMAM effectively aggregates discretely distributed multi-level defect features across different layers and guides the query image to perceive defects from the pixel level, which avoids incomplete segmentation caused by weak features. Additionally, a Foreground Correction Module (FCM) is developed, which is equipped with a dedicated background predictor (BP) and a foreground corrector (FC). BP places more emphasis on learning features from backgrounds rather than defects. FC achieves efficient feature ensemble and further suppresses the backgrounds misidentified as defects in CMAM. They collaborate to prevent incorrect segmentation caused by background interference. Extensive experiments demonstrate the effectiveness of our method. We achieve state-of-the-art results on both FSSD-12, a public benchmark FSS dataset for strip steel, and FSS-AEB, an FSS dataset for aero-engine blades. Specifically, with 1/5 support images, we achieve 64.6%/65.6% mIoU on FSSD-12 and 55.0%/57.8% mIoU on FSS-AEB. Note to Practitioners—Surface defect segmentation has always been a hot topic in the industry. However, existing methods rely on sufficient training data and struggle to generalize to unseen categories, which significantly hinders the automation of defect segmentation. To address this problem, we propose MCANet for automated few-shot defect segmentation. It achieves effective segmentation for surface defects with limited data, even for unseen categories. Furthermore, MCANet achieves state-of-the-art results on two datasets from real-world industrial scenarios and also delivers significant improvements over the widely concerned large vision models. Finally, we integrate MCANet into an automated surface defect inspection platform consisting of an imaging system and a high-performance computing server for real-world performance validation. Yunfeng Ma, Min Liu 0008, Yuan Bian 0002, Yaonan Wang 0001 |
IEEE Trans Autom. Sci. Eng. | 2 |
| 2025 | SPDP-Net: A Semantic Prior Guided Defect Perception Network for Automated Aero-Engine Blades Surface Visual InspectionabstractAutomated surface defect detection is essential to manufacturing automation. However, automated inspection of aero-engine blades remains challenging due to tiny defects and weak features. To address this issue, we propose a semantic prior guided defect perception network, named SPDP-Net, which is ultimately integrated into an automated system to achieve efficient detection of defects. Firstly, a semantic prior mining module (SPM) is developed to capture finer-grained pixel-level location priors of defects by leveraging image feature mapping relations which facilitates the precise perception of tiny defects. Subsequently, we propose a defect enhancement perception module (DEP) to separate weak defects from complex backgrounds by utilizing defect location priors provided by SPM to enhance the features of defects while suppressing the values of non-defect regions, which makes the weak defects present as more obvious outliers. Finally, the global information extraction module (GIE) extracts the global features of defects, which helps to further improve the predicted results. When equipped with SPM, DEP and GIE, SPDP-Net can accurately identify and locate defects, exhibiting more competitive recognition and feature extraction capabilities for tiny defects and weak defects. To evaluate the effectiveness of our method, we construct an aero-engine blade surface defect detection dataset from real industrial scenarios called ABSDD with the collaboration of senior engineers. We achieve 95.9% precision, 94.0% recall and 94.9% F1 score on ABSDD. In addition, we also achieve state-of-the-art results on two public benchmark datasets, KSDD2 and DAGM. Finally, we have applied the developed SPDP-Net to an automated system and have conducted actual tests in collaboration with a well-known aero-engine production company.Note to Practitioners—At present, the automated detection system for surface defects in industrial manufacturing has not been well developed, which is particularly trailing behind in the field of aero-engine manufacturing. To the best of our knowledge, the surface defect detection of aero-engine blades is still carried out manually. To address this problem, we propose SPDP-Net for the automated detection of surface defects in aero-engine blades. It can accurately perceive and capture tiny and weak defects with excellent performance. SPDP-Net achieves state-of-the-art results on three tasks from different industrial manufacturing fields, which demonstrates its good transfer application capabilities. In addition, we integrate SPDP-Net into an automated detection system consisting of an autonomous imaging system and a high-performance computing server and conduct actual tests in an aero-engine production company. The test achieves remarkable results, indicating the good application prospects of the automated detection system. Yunfeng Ma, Min Liu 0008, Yiqiong Zhang, Yaonan Wang 0001 |
IEEE Trans Autom. Sci. Eng. | 2 |
| 2025 | Modality Unified Attack for Omni-Modality Person Re-IdentificationabstractDeep learning based person re-identification (re-id) models have been widely employed in surveillance systems. Recent studies have demonstrated that black-box single-modality and cross-modality re-id models are vulnerable to adversarial examples (AEs), leaving the robustness of multi-modality re-id models unexplored. Due to the lack of knowledge about the specific type of model deployed in the target black-box surveillance system, we aim to generate modality unified AEs for omni-modality (single-, cross- and multi-modality) re-id models. Specifically, we propose a novel Modality Unified Attack method to train modality-specific adversarial generators to generate AEs that effectively attack different omni-modality models. A multi-modality model is adopted as the surrogate model, wherein the features of each modality are perturbed by metric disruption loss before fusion. To collapse the common features of omnimodality models, Cross Modality Simulated Disruption approach is introduced to mimic the cross-modality feature embeddings by intentionally feeding images to non-corresponding modality-specific subnetworks of the surrogate model. Moreover, Multi Modality Collaborative Disruption strategy is devised to facilitate the attacker to comprehensively corrupt the informative content of person images by leveraging a multi modality feature collaborative metric disruption loss. Extensive experiments show that our MUA method can effectively attack the omni-modality re-id models, achieving 55.9%, 24.4%, 49.0% and 62.7% mean mAP Drop Rate, respectively. Yuan Bian 0002, Min Liu 0008, Yunqi Yi, Yunfeng Ma, Yaonan Wang 0001 |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2025 | Learnable Prompts With Neighbor-Aware Correction for Text-Based Person Search
Min Liu 0008, Yaonan Wang 0001 |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2025 | Learning to Learn Transferable Generative Attack for Person Re-IdentificationabstractDeep learning-based person re-identification (re-id) models are widely employed in surveillance systems and inevitably inherit the vulnerability of deep networks to adversarial attacks. Existing attacks merely consider cross-dataset and cross-model transferability, ignoring the cross-test capability to perturb models trained in different domains. To powerfully examine the robustness of real-world re-id models, the Meta Transferable Generative Attack (MTGA) method is proposed, which adopts meta-learning optimization to promote the generative attacker producing highly transferable adversarial examples by learning comprehensively simulated transfer-based cross-model&dataset&test black-box meta attack tasks. Specifically, cross-model&dataset black-box attack tasks are first mimicked by selecting different re-id models and datasets for meta-train and meta-test attack processes. As different models may focus on different feature regions, the Perturbation Random Erasing module is further devised to prevent the attacker from learning to only corrupt model-specific features. To boost the attacker learning to possess cross-test transferability, the Normalization Mix strategy is introduced to imitate diverse feature embedding spaces by mixing multi-domain statistics of target models. Extensive experiments show the superiority of MTGA, especially in cross-model&dataset and cross-model&dataset&test attacks, our MTGA outperforms the SOTA methods by 20.0% and 11.3% on mean mAP drop rate, respectively. The source codes are available at https://github.com/yuanbianGit/MTGA. Yuan Bian 0002, Min Liu 0008, Yunfeng Ma, Yaonan Wang 0001 |
IEEE Trans. Image Process. | 2 |
| 2025 | Cross-Modality Semantic Consistency Learning for Visible-Infrared Person Re-IdentificationabstractVisible-infrared person re-identification (VI-ReID) seeks to identify and match individuals across visible and infrared ranges within intelligent monitoring environments. Most current approaches predominantly explore a two-stream network structure that extract global or rigidly split part features and introduce an extra modality for image compensation to guide networks reducing the huge differences between the two modalities. However, these methods are sensitive to misalignment caused by pose/viewpoint variations and additional noises produced by extra modality generating. Within the confines of this articles, we clearly consider addresses above issues and propose a Cross-modality Semantic Consistency Learning (CSCL) network to excavate the semantic consistent features in different modalities by utilizing human semantic information. Specifically, a Parsing-aligned Attention Module (PAM) is introduced to filter out the irrelevant noises with channel-wise attention and dynamically highlight the semantic-aware representations across modalities in different stages of the network. Then, a Semantic-guided Part Alignment Module (SPAM) is introduced, aimed at efficiently producing a collection of semantic-aligned fine-grained features. This is achieved by incorporating parsing loss and division loss constraints, ultimately enhancing the overall person representation. Finally, an Identity-aware Center Mining (ICM) loss is presented to reduce the distribution between modality centers within classes, thereby further alleviating intra-class modality discrepancies. Extensive experiments indicate that CSCL outperforms the state-of-the-art methods on the SYSU-MM01 and RegDB datasets. Notably, the Rank-1/mAP accuracy on the SYSU-MM01 dataset can achieve 75.72%/72.08%. Min Liu 0008, Yuan Bian 0002, Yeqing Sun, Baida Zhang, Yaonan Wang 0001 |
IEEE Trans. Multim. | 1 |
| 2025 | Spatially Covariant Image Registration With Text PromptsabstractMedical images are often characterized by their structured anatomical representations and spatially inhomogeneous contrasts. Leveraging anatomical priors in neural networks can greatly enhance their utility in resource-constrained clinical settings. Prior research has harnessed such information for image segmentation, yet progress in deformable image registration has been modest. Our work introduces textSCF, a novel method that integrates spatially covariant filters and textual anatomical prompts encoded by visual-language models, to fill this gap. This approach optimizes an implicit function that correlates text embeddings of anatomical regions to filter weights. textSCF not only boosts computational efficiency but can also retain or improve registration accuracy. By capturing the contextual interplay between anatomical regions, it offers impressive interregional transferability and the ability to preserve structural discontinuities during registration. textSCF's performance has been rigorously tested on intersubject brain magnetic resonance imaging (MRI) and abdominal computerized tomography (CT) registration tasks, outperforming existing state-of-the-art models in the MICCAI Learn2Reg 2021 challenge and leading the leaderboard. In abdominal registrations, textSCF's larger model variant improved the Dice score by 11.3% over the second-best model, while its smaller variant maintained similar accuracy but with an 89.13% reduction in network parameters and a 98.34% decrease in computational operations. Xiang Chen 0008, Min Liu 0008, Rongguang Wang, Renjiu Hu, Gaolei Li, Yaonan Wang 0001, Hang Zhang 0010 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2025 | MBUNeXt: Multibranch Encoder Aggregation Network Based on Layer-Fusion Strategy for Multimodal Brain Tumor SegmentationabstractMultimodal brain tumor segmentation (BraTS), integrated with surgical robots and navigation systems, enables accurate surgical interventions while maximizing the preservation of surrounding healthy brain tissue. However, multimodal brain scans suffer from large interclass differences in brain tumor subregions and information redundancy, leading to inadequate fusion of multimodal information and significantly affecting the accuracy of BraTS. To address the above problems, we propose a multibranch encoder aggregation (MEA) network based on a layer-fusion strategy called multibranch UNeXt (MBUNeXt). The network comprises three well-designed modules: the multimodal feature attention (MFA) module, the MEA module, and the large-kernel convolution skip (LCS)-connection module. These modules work together to achieve precise segmentation of brain tumors. Specifically, the MFA module preserves the intermodality similarity structure through attention mechanisms and Gaussian modulation functions, thereby filtering redundant information. Then, the MEA module exploits the correlations among multiple modalities to effectively integrate multimodal hybrid feature representation and optimize multimodal information fusion. In addition, the LCS module constructs multiple groups of depthwise separable convolutions with large kernel, which can guide the network to attend to features at different scales, thereby addressing the issue of significant interclass differences in brain tumor subregions. The experimental results on the large-scale public datasets, BraTS2019 and BraTS2021, which consist of approximately 5000 3-D brain scans, demonstrate that our proposed method has achieved SOTA performance, with average Dice scores of 85.84% and 91.11%, respectively. It also performs well on the BraTS-Africa2024 dataset with low imaging quality, confirming its robustness. The code is available at https://github.com/liuqinghao2018/MBUNeXt. Qinghao Liu, Yuehao Zhu, Min Liu 0008, Zhao Yao, Yaonan Wang 0001, Erik Meijering |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2024 | Structured serialization semantic transfer network for unsupervised cross-domain recognition and retrieval
Dan Song 0006, Yuanxiang Yang, Wenhui Li 0001, Xuanya Li, Min Liu 0008, Anan Liu |
Inf. Process. Manag. | 5 |
| 2024 | Noise-robust re-identification with triple-consistency perception
Zhanpeng Shao, Shixi Luo, Jiazheng Wang 0001, Min Liu 0008, Jianhua Dai 0003 |
Image Vis. Comput. | 5 |
| 2024 | A Two-Stage Noise-Tolerant Paradigm for Label Corrupted Person Re-IdentificationabstractSupervised person re-identification (Re-ID) approaches are sensitive to label corrupted data, which is inevitable and generally ignored in the field of person Re-ID. In this paper, we propose a two-stage noise-tolerant paradigm (TSNT) for labeling corrupted person Re-ID. Specifically, at stage one, we present a self-refining strategy to separately train each network in TSNT by concentrating more on pure samples. These pure samples are progressively refurbished via mining the consistency between annotations and predictions. To enhance the tolerance of TSNT to noisy labels, at stage two, we employ a co-training strategy to collaboratively supervise the learning of the two networks. Concretely, a rectified cross-entropy loss is proposed to learn the mutual information from the peer network by assigning large weights to the refurbished reliable samples. Moreover, a noise-robust triplet loss is formulated for further improving the robustness of TSNT by increasing inter-class distances and reducing intra-class distances in the label-corrupted dataset, where a constraint condition for reliability discrimination is carefully designed to select reliable triplets. Extensive experiments demonstrate the superiority of TSNT, for instance, on the Market1501 dataset, our paradigm achieves 90.3% rank-1 accuracy (6.2% improvement over the state-of-the-art method) under noise ratio 20%. Min Liu 0008, Fei Wang 0124, Yaonan Wang 0001, Amit K. Roy-Chowdhury |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2024 | Weakly Supervised Tracklet Association Learning With Video Labels for Person Re-IdentificationabstractSupervised person re-identification (re-id) methods require expensive manual labeling costs. Although unsupervised re-id methods can reduce the requirement of the labeled datasets, the performance of these methods is lower than the supervised alternatives. Recently, some weakly supervised learning-based person re-id methods have been proposed, which is a balance between supervised and unsupervised learning. Nevertheless, most of these models require another auxiliary fully supervised datasets or ignore the interference of noisy tracklets. To address this problem, in this work, we formulate a weakly supervised tracklet association learning (WS-TAL) model only leveraging the video labels. Specifically, we first propose an intra-bag tracklet discrimination learning (ITDL) term. It can capture the associations between person identities and images by assigning pseudo labels to each person image in a bag. And then, the discriminative feature for each person is learned by utilizing the obtained associations after filtering the noisy tracklets. Based on that, a cross-bag tracklet association learning (CTAL) term is presented to explore the potential tracklet associations between bags by mining reliable positive tracklet pairs and hard negative pairs. Finally, these two complementary terms are jointly optimized to train our re-id model. Extensive experiments on the weakly labeled datasets demonstrate that WS-TAL achieves 88.1% and 90.3% rank-1 accuracy on the MARS and DukeMTMC-VideoReID datasets respectively. The performance of our model surpasses the state-of-the-art weakly supervised models by a large margin, even outperforms some fully supervised re-id models. Min Liu 0008, Yuan Bian 0002, Qing Liu 0035, Yaonan Wang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2024 | Fashion Customization: Image Generation Based on Editing ClueabstractFashion image generation attracts increasing attentions with wide applications in fashion design, virtual try-on, cosmetic industry, etc. Editing clues such as segmentation masks, keypoints and sketches are usually taken to guide the desired transformation of a reference image. However, spatial manipulation of the reference image remains a challenge, especially facing large-scale deformations and multiple editing requirements. In this paper, we propose a general model for multiple fashion editing tasks such as facial editing, pose transformation and clothes design based on user-defined editing instructions like semantic segmentation masks, keypoints, and sketches. With diverse editing requirements and various deformation scales, it is hard to learn the corresponding relationship between the editing clue and reference image with a uniform framework. Accordingly, we design a feature flow estimation network, which can adaptively adjust the feature flow according to the editing clue and the reference image, and generate a coarsely aligned image. Then we propose an image generative network to enrich the texture details of the transformed reference image. Experiments on three tasks verify the effectiveness of the proposed method and the adaptability to multiple tasks. The code and pretrained models will be available at https://github.com/zengjianhao/Fashion-Image-Generation-Based-on-Editing-Clue. Dan Song 0006, Jianhao Zeng, Min Liu 0008, Xuanya Li, Anan Liu |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | LSKANet: Long Strip Kernel Attention Network for Robotic Surgical Scene SegmentationabstractSurgical scene segmentation is a critical task in Robotic-assisted surgery. However, the complexity of the surgical scene, which mainly includes local feature similarity (e.g., between different anatomical tissues), intraoperative complex artifacts, and indistinguishable boundaries, poses significant challenges to accurate segmentation. To tackle these problems, we propose the Long Strip Kernel Attention network (LSKANet), including two well-designed modules named Dual-block Large Kernel Attention module (DLKA) and Multiscale Affinity Feature Fusion module (MAFF), which can implement precise segmentation of surgical images. Specifically, by introducing strip convolutions with different topologies (cascaded and parallel) in two blocks and a large kernel design, DLKA can make full use of region- and strip-like surgical features and extract both visual and structural information to reduce the false segmentation caused by local feature similarity. In MAFF, affinity matrices calculated from multiscale feature maps are applied as feature fusion weights, which helps to address the interference of artifacts by suppressing the activations of irrelevant regions. Besides, the hybrid loss with Boundary Guided Head (BGH) is proposed to help the network segment indistinguishable boundaries effectively. We evaluate the proposed LSKANet on three datasets with different surgical scenes. The experimental results show that our method achieves new state-of-the-art results on all three datasets with improvements of 2.6%, 1.4%, and 3.4% mIoU, respectively. Furthermore, our method is compatible with different backbones and can significantly increase their segmentation accuracy. Code is available at https://github.com/YubinHan73/LSKANet. Min Liu 0008, Yubin Han, Jiazheng Wang 0001, Can Wang 0011, Yaonan Wang 0001, Erik Meijering |
IEEE Trans. Medical Imaging | 1 |
| 2024 | Brain Image Segmentation for Ultrascale Neuron Reconstruction via an Adaptive Dual-Task Learning NetworkabstractAccurate morphological reconstruction of neurons in whole brain images is critical for brain science research. However, due to the wide range of whole brain imaging, uneven staining, and optical system fluctuations, there are significant differences in image properties between different regions of the ultrascale brain image, such as dramatically varying voxel intensities and inhomogeneous distribution of background noise, posing an enormous challenge to neuron reconstruction from whole brain images. In this paper, we propose an adaptive dual-task learning network (ADTL-Net) to quickly and accurately extract neuronal structures from ultrascale brain images. Specifically, this framework includes an External Features Classifier (EFC) and a Parameter Adaptive Segmentation Decoder (PASD), which share the same Multi-Scale Feature Encoder (MSFE). MSFE introduces an attention module named Channel Space Fusion Module (CSFM) to extract structure and intensity distribution features of neurons at different scales for addressing the problem of anisotropy in 3D space. Then, EFC is designed to classify these feature maps based on external features, such as foreground intensity distributions and image smoothness, and select specific PASD parameters to decode them of different classes to obtain accurate segmentation results. PASD contains multiple sets of parameters trained by different representative complex signal-to-noise distribution image blocks to handle various images more robustly. Experimental results prove that compared with other advanced segmentation methods for neuron reconstruction, the proposed method achieves state-of-the-art results in the task of neuron reconstruction from ultrascale brain images, with an improvement of about 49% in speed and 12% in F1 score. Min Liu 0008, Shuhan Wu, Zhuangdian Lin, Yaonan Wang 0001, Erik Meijering |
IEEE Trans. Medical Imaging | 1 |
| 2024 | Occlusion-Aware Feature Recover Model for Occluded Person Re-IdentificationabstractOccluded person re-identification (Re-ID) is a challenging task, as various object-to-person (OTP) and person-to-person (PTP) occlusion scenarios cause diverse occlusion interference and target person feature loss problems in person matching. Most existing methods, which utilize auxiliary models to evaluate the unoccluded person parts for occlusion feature elimination, are inefficient and cannot handle the PTP occlusion scenarios and person feature loss problems. To solve these issues, we propose a novel Occlusion-Aware Feature Recover (OAFR) model. OAFR simulates diverse occlusions to facilitate the model perceiving OTP, PTP occlusions and recovers occluded query features with unoccluded retrieved gallery features. Concretely, the Prior Knowledge-based Occlusion Simulation method is firstly introduced to synthesize OTP, PTP occlusions and corresponding occlusion labels, empowering model target person perception and occlusion-aware capability through self-supervised learning. Afterward, the feature recovery module reconstructs occluded query features with corresponding unoccluded local features of the top-$K$retrieved images by the visibility weighted average scheme, thus recovering the occluded query features to maintain more comprehensive features for better retrieval. Extensive experiments demonstrate that the proposed OAFR achieves superior performance to the state-of-the-art for both holistic and occluded Re-ID. Especially for Occluded-DukeMTMC dataset, OAFR outperforms the state-of-the-art by 6.0% for Rank-1 accuracy and 2.2% for mAP score. The source codes are available athttps://github.com/yuanbianGit/OAFR. Yuan Bian 0002, Min Liu 0008, Yaonan Wang 0001 |
IEEE Trans. Multim. | 2 |
| 2024 | Relation-Preserving Feature Embedding for Unsupervised Person Re-IdentificationabstractSome unsupervised approaches have been proposed recently for the person re-identification (ReID) problem since annotations of samples across cameras are time-consuming. However, most of these methods focus on the appearance content of the sample itself, and thus seldom take the structure relations among samples into account when learning the feature representation, which would provide a valuable guide for learning the representations of the samples. Thus hard samples may not be well solved due to the limited or even misleading information of the sample itself. To address this issue, in this article, we propose a Relation-Preserving Feature Embedding (RPE) model that leverages structure relations among samples to boost the performance of the unsupervised person ReID methods without requiring any sample annotations. RPE aims at integrating the sample content and the neighborhood structure relations among samples into the learning of feature embeddings by combining the advantages of the autoencoder and graph autoencoder. Specifically, a relation and content information fusion (RCIF) module is proposed to dynamically merge the information from both perspectives of content and relation levels for feature embedding learning. Also, due to the lack of the identity labels of samples, we adopt an adaptive optimization strategy to update the affinity relations among samples instead of the reconstruction of the whole affinity matrix for optimizing the RPE model, which is more suitable for the unsupervised ReID task. Rigorous experiments on three widely-used large-scale benchmarks for person ReID demonstrate the superiority of the proposed method over current state-of-the-art unsupervised methods. Min Liu 0008, Fei Wang 0124, Jianhua Dai 0003, Anan Liu, Yaonan Wang 0001 |
IEEE Trans. Multim. | 2 |
| 2024 | SwinPA-Net: Swin Transformer-Based Multiscale Feature Pyramid Aggregation Network for Medical Image SegmentationabstractThe precise segmentation of medical images is one of the key challenges in pathology research and clinical practice. However, many medical image segmentation tasks have problems such as large differences between different types of lesions and similar shapes as well as colors between lesions and surrounding tissues, which seriously affects the improvement of segmentation accuracy. In this article, a novel method called Swin Pyramid Aggregation network (SwinPA-Net) is proposed by combining two designed modules with Swin Transformer to learn more powerful and robust features. The two modules, named dense multiplicative connection (DMC) module and local pyramid attention (LPA) module, are proposed to aggregate the multiscale context information of medical images. The DMC module cascades the multiscale semantic feature information through dense multiplicative feature fusion, which minimizes the interference of shallow background noise to improve the feature expression and solves the problem of excessive variation in lesion size and type. Moreover, the LPA module guides the network to focus on the region of interest by merging the global attention and the local attention, which helps to solve similar problems. The proposed network is evaluated on two public benchmark datasets for polyp segmentation task and skin lesion segmentation task as well as a clinical private dataset for laparoscopic image segmentation task. Compared with existing state-of-the-art (SOTA) methods, the SwinPA-Net achieves the most advanced performance and can outperform the second-best method on the mean Dice score by 1.68%, 0.8%, and 1.2% on the three tasks, respectively. Jiazheng Wang 0001, Min Liu 0008, Yaonan Wang 0001, Erik Meijering |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2024 | Design and Analysis of a Novel Distributed Gradient Neural Network for Solving Consensus Problems in a Predefined TimeabstractIn this article, a novel distributed gradient neural network (DGNN) with predefined-time convergence (PTC) is proposed to solve consensus problems widely existing in multiagent systems (MASs). Compared with previous gradient neural networks (GNNs) for optimization and computation, the proposed DGNN model works in a nonfully connected way, in which each neuron only needs the information of neighbor neurons to converge to the equilibrium point. The convergence and asymptotic stability of the DGNN model are proved according to the Lyapunov theory. In addition, based on a relatively loose condition, three novel nonlinear activation functions are designed to speedup the DGNN model to PTC, which is proved by rigorous theory. Computer numerical results further verify the effectiveness, especially the PTC, of the proposed nonlinearly activated DGNN model to solve various consensus problems of MASs. Finally, a practical case of the directional consensus is presented to show the feasibility of the DGNN model and a corresponding connectivity-testing example is given to verify the influence on the convergence speed. Lin Xiao 0002, Lei Jia 0001, Jianhua Dai 0003, Yingkun Cao, Yiwei Li 0006, Quanxin Zhu, Jichun Li 0002, Min Liu 0008 |
IEEE Trans. Neural Networks Learn. Syst. | 8 |
| 2024 | Anchor Association Learning for Unsupervised Video Person Re-IdentificationabstractVideo-based person re-identification (re-id) has attracted a significant attention in recent years due to the increasing demand of video surveillance. However, existing methods are usually based on the supervised learning, which requires vast labeled identities across cameras and is not suitable for real scenes. Although some unsupervised approaches have been proposed for video re-id, their performance is far from satisfactory. In this article, we propose an unsupervised anchor association learning (UAAL) framework to address the video-based person re-id task, in which the feature representation of each sampled tracklet is regarded as an anchor. Specifically, we first propose an intracamera anchor association learning (IAAL) term that learns the discriminative anchor by utilizing the affiliation relations between an image and the anchors in each camera. Then, the exponential moving average (EMA) strategy is employed to update the anchor and the updated anchors are stored into an anchor memory module. On top of that, a cross-camera anchor association learning (CAAL) term is introduced to mine potential positive anchor pairs across cameras by presenting a cyclic ranking anchor alignment and threshold filtering method. Extensive experiments conducted on two public datasets show the superiority of the proposed method; for example, our method achieves 73.2% for rank-1 accuracy and 60.1% for mean average precision (mAP) score, respectively, on MARS, similarly 89.7% and 87.0% on DukeMTMC-VideoReID. Shujun Zeng, Min Liu 0008, Qing Liu 0035, Yaonan Wang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2024 | Multi-stage reasoning on introspecting and revising bias for visual question answeringabstractVisual Question Answering (VQA) is a task that involves predicting an answer to a question depending on the content of an image. However, recent VQA methods have relied more on language priors between the question and answer rather than the image content. To address this issue, many debiasing methods have been proposed to reduce language bias in model reasoning. However, the bias can be divided into two categories: good bias and bad bias. Good bias can benefit to the answer prediction, while the bad bias may associate the models with the unrelated information. Therefore, instead of excluding good and bad bias indiscriminately in existing debiasing methods, we proposed a bias discrimination module to distinguish them. Additionally, bad bias may reduce the model’s reliance on image content during answer reasoning and thus attend little on image features updating. To tackle this, we leverage Markov theory to construct a Markov field with image regions and question words as nodes. This helps with feature updating for both image regions and question words, thereby facilitating more accurate and comprehensive reasoning about both the image content and question. To verify the effectiveness of our network, we evaluate our network on VQA v2 and VQA cp v2 datasets and conduct extensive quantity and quality studies to verify the effectiveness of our proposed network. Experimental resu- lts show that our network achieves significant performance against the previous state-of-the-art methods. Anan Liu, Zimu Lu, Ning Xu 0003, Min Liu 0008, Chenggang Yan 0001, Bolun Zheng, Yulong Duan, Xuanya Li |
ACM Trans. Web | 4 |
| 2024 | Robust object recognition via context-driven reliability assessment
Jiazheng Wang 0001, Min Liu 0008 |
Vis. Comput. | 4 |
| 2023 | External Knowledge Dynamic Modeling for Image-text RetrievalabstractImage-text retrieval is a fundamental branch in cross-modal retrieval. The core is to explore the semantic correspondence to align relevant image-text pairs. Some existing methods rely on global semantics and co-occurrence frequency to design knowledge introduction patterns for consistent representations. However, they lack flexibility due to the limitations of fixed information and empirical feedback. To address these issues, we develop an External Knowledge Dynamic Modeling~(EKDM) architecture based on the filtering mechanism, which dynamically explores different knowledge towards varied image-text pairs. Specially, we first capture abundant concepts and relationships from external knowledge to construct visual and textual corpus sets. Then, we progressively explores concepts related to images and texts by dynamic global representations. To endow the model with the capability of relationship decision, we integrate the variable spatial locations between objects for association exploration. Since the filtering mechanism is conditioned on dynamic semantics and variable spatial locations, our model can dynamically model different knowledge for different image-text pairs. Extensive experimental results on two benchmark datasets demonstrate the effectiveness of our proposed method. Qiang Li 0048, Wenhui Li 0001, Min Liu 0008, Xuanya Li, Anan Liu |
ACM Multimedia | 4 |
| 2023 | DITN: User's indirect side-information involved domain-invariant feature transfer network for cross-domain recommendation
Jie Nie, Zijie Zuo, Huaxin Xie, Mingxing Jiang, Jianliang Xu, Shusong Yu, Min Liu 0008 |
Inf. Process. Manag. | 9 |
| 2023 | Unsupervised self-training correction learning for 2D image-based 3D model retrieval
Yaqian Zhou 0002, Yu Liu 0004, Jun Xiao 0001, Min Liu 0008, Xuanya Li, Anan Liu |
Inf. Process. Manag. | 4 |
| 2023 | OTP-NMS: Toward Optimal Threshold Prediction of NMS for Crowded Pedestrian DetectionabstractPedestrian detection is still a challenging task for computer vision, especially in crowded scenes where the overlaps between pedestrians tend to be large. The non-maximum suppression (NMS) plays an important role in removing the redundant false positive detection proposals while retaining the true positive detection proposals. However, the highly overlapped results may be suppressed if the threshold of NMS is lower. Meanwhile, a higher threshold of NMS will introduce a larger number of false positive results. To solve this problem, we propose an optimal threshold prediction (OTP) based NMS method that predicts a suitable threshold of NMS for each human instance. First, a visibility estimation module is designed to obtain the visibility ratio. Then, we propose a threshold prediction subnet to determine the optimal threshold of NMS automatically according to the visibility ratio and classification score. Finally, we re-formulate the objective function of the subnet and utilize the reward-guided gradient estimation algorithm to update the subnet. Comprehensive experiments on CrowdHuman and CityPersons show the superior performance of the proposed method in pedestrian detection, especially in crowded scenes. Min Liu 0008, Baopu Li, Yaonan Wang 0001, Wanli Ouyang |
IEEE Trans. Image Process. | 2 |
| 2023 | Branch Aggregation Attention Network for Robotic Surgical Instrument SegmentationabstractSurgical instrument segmentation is of great significance to robot-assisted surgery, but the noise caused by reflection, water mist, and motion blur during the surgery as well as the different forms of surgical instruments would greatly increase the difficulty of precise segmentation. A novel method called Branch Aggregation Attention network (BAANet) is proposed to address these challenges, which adopts a lightweight encoder and two designed modules, named Branch Balance Aggregation module (BBA) and Block Attention Fusion module (BAF), for efficient feature localization and denoising. By introducing the unique BBA module, features from multiple branches are balanced and optimized through a combination of addition and multiplication to complement strengths and effectively suppress noise. Furthermore, to fully integrate the contextual information and capture the region of interest, the BAF module is proposed in the decoder, which receives adjacent feature maps from the BBA module and localizes the surgical instruments from both global and local perspectives by utilizing a dual branch attention mechanism. According to the experimental results, the proposed method has the advantage of being lightweight while outperforming the second-best method by 4.03%, 1.53%, and 1.34% in mIoU scores on three challenging surgical instrument datasets, respectively, compared to the existing state-of-the-art methods. Code is available at https://github.com/SWT-1014/BAANet. Wenting Shen, Yaonan Wang 0001, Min Liu 0008, Jiazheng Wang 0001, Renjie Ding, Zhe Zhang 0022, Erik Meijering |
IEEE Trans. Medical Imaging | 3 |
| 2023 | 3D Soma Detection in Large-Scale Whole Brain Images via a Two-Stage Neural Networkabstract3D soma detection in whole brain images is a critical step for neuron reconstruction. However, existing soma detection methods are not suitable for whole mouse brain images with large amounts of data and complex structure. In this paper, we propose a two-stage deep neural network to achieve fast and accurate soma detection in large-scale and high-resolution whole mouse brain images (more than 1TB). For the first stage, a lightweight Multi-level Cross Classification Network (MCC-Net) is proposed to filter out images without somas and generate coarse candidate images by combining the advantages of the multi convolution layer's feature extraction ability. It can speed up the detection of somas and reduce the computational complexity. For the second stage, to further obtain the accurate locations of somas in the whole mouse brain images, the Scale Fusion Segmentation Network (SFS-Net) is developed to segment soma regions from candidate images. Specifically, the SFS-Net captures multi-scale context information and establishes a complementary relationship between encoder and decoder by combining the encoder-decoder structure and a 3D Scale-Aware Pyramid Fusion (SAPF) module for better segmentation performance. The experimental results on three whole mouse brain images verify that the proposed method can achieve excellent performance and provide the reconstruction of neurons with beneficial information. Additionally, we have established a public dataset named WBMSD, including 798 high-resolution and representative images ( 256 ×256 ×256 voxels) from three whole mouse brain images, dedicated to the research of soma detection, which will be released along with this paper. Xiaodan Wei, Qinghao Liu, Min Liu 0008, Yaonan Wang 0001, Erik Meijering |
IEEE Trans. Medical Imaging | 3 |
| 2023 | Refining Noisy Labels With Label Reliability Perception for Person Re-IdentificationabstractMost person re-identification (Re-ID) approaches rely excessively on a great quantity of annotated training data. However, due to sampling errors or annotated errors, the label noise is unavoidable, which usually causes a dramatic decrease in the performance of existing Re-ID methods. To address this problem, we propose the label reliability perception (LRP) for person Re-ID by refining noisy labels. Specifically, a feature-fusion block (FFB) is proposed to enhance the discrimen- ability of pedestrians' features by expanding the network's attention span due to the fused feature, which is generated by overlapping the coarse-grained feature obtained by global average pooling and fine-grained features obtained by evenly dividing the feature map in the height dimension and performing global max pooling. In addition, the label dual perception (LDP) is proposed to refine noisy labels instead of filtering samples by evaluating the reliability of each training sample's label. Specifically, we meticulously design five evaluation modes for each sample to perceive the reliability of the labels of thek-nearest neighbor images. Finally, we utilize the most reliable label to replace the noisy label and optimize the network. Extensive experiments prove the superiority of the proposed model over the competing methods; for instance, on Market1501, our method achieves 88.8% rank-1 accuracy and 70.5% mAP (4.7% and 4.3% improvements over the state-of-the-arts) under noise ratio 20%, and similarly on DukeMTMC-ReID, our method achieves 77.7% and 60.3%. Yongchun Chen, Min Liu 0008, Fei Wang 0124, Anan Liu, Yaonan Wang 0001 |
IEEE Trans. Multim. | 2 |
| 2022 | A Deep Local Patch Matching Network for Cell Tracking in Microscopy Image Sequences Without RegistrationabstractCell tracking is critical for the modeling of plant cell growth patterns. A local graph matching algorithm is proposed to track cells by exploiting the tight spatial topology of cells. However, the local graph matching approach lacks robustness in the unregistered images because the feature descriptors are handcrafted. In this paper, we propose a Deep Local Patch Matching Network (DLPM-Net) to track cells robustly, by exploiting local patches' deep similarity information and cells' spatial-temporal contextual information. Furthermore, to reduce the time consumption during the matching process and enhance tracking accuracy, we take two steps to realize the tracking of non-division cells and the detection of cell divisions. In the first step, the DLPM-Net is employed to match the non-division cells by exploiting the cell pair candidates' local patch contextual information, then the non-matched cells are recorded as the cell division candidates. In the second step, the DLPM-Net is used to detect cell divisions from these non-matched cells, by exploiting the local patch contextual similarity between the mother cell's local patch and daughter cells' local patch. Compared with the existing local graph matching method, the experimental results show that the proposed method gains 29.1% improvement in the tracking accuracy. Yulian Xie, Min Liu 0008, Shirui Zhou, Yaonan Wang 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2022 | DeepRayburst for Automatic Shape Analysis of Tree-Like Structures in Biomedical ImagesabstractPrecise quantification of tree-like structures from biomedical images, such as neuronal shape reconstruction and retinal blood vessel caliber estimation, is increasingly important in understanding normal function and pathologic processes in biology. Some handcrafted methods have been proposed for this purpose in recent years. However, they are designed only for a specific application. In this paper, we propose a shape analysis algorithm, DeepRayburst, that can be applied to many different applications based on a Multi-Feature Rayburst Sampling (MFRS) and a Dual Channel Temporal Convolutional Network (DC-TCN). Specifically, we first generate a Rayburst Sampling (RS) core containing a set of multidirectional rays. Then the MFRS is designed by extending each ray of the RS to multiple parallel rays which extract a set of feature sequences. A Gaussian kernel is then used to fuse these feature sequences and outputs one feature sequence. Furthermore, we design a DC-TCN to make the rays terminate on the surface of tree-like structures according to the fused feature sequence. Finally, by analyzing the distribution patterns of the terminated rays, the algorithm can serve multiple shape analysis applications of tree-like structures. Experiments on three different applications, including soma shape reconstruction, neuronal shape reconstruction, and vessel caliber estimation, confirm that the proposed method outperforms other state-of-the-art shape analysis methods, which demonstrate its flexibility and robustness. Weixun Chen, Min Liu 0008, Yaonan Wang 0001, Erik Meijering |
IEEE J. Biomed. Health Informatics | 3 |
| 2022 | An O-Shape Neural Network With Attention Modules to Detect Junctions in Biomedical Images Without SegmentationabstractJunction plays an important role in biomedical research such as retinal biometric identification, retinal image registration, eye-related disease diagnosis and neuron reconstruction. However, junction detection in original biomedical images is extremely challenging. For example, retinal images contain many tiny blood vessels with complicated structures and low contrast, which makes it challenging to detect junctions. In this paper, we propose an O-shape Network architecture with Attention modules (Attention O-Net), which includes Junction Detection Branch (JDB) and Local Enhancement Branch (LEB) to detect junctions in biomedical images without segmentation. In JDB, the heatmap indicating the probabilities of junctions is estimated and followed by choosing the positions with the local highest value as the junctions, whereas it is challenging to detect junctions when the images contain weak filament signals. Therefore, LEB is constructed to enhance the thin branch foreground and make the network pay more attention to the regions with low contrast, which is helpful to alleviate the imbalance of the foreground between thin and thick branches and to detect the junctions of the thin branch. Furthermore, attention modules are utilized to introduce the feature maps of LEB to JDB, which can establish a complementary relationship and further integrate local features and contextual information between these two branches. The proposed method achieves the highest average F1-scores of 0.82, 0.73 and 0.94 in two retinal datasets and one neuron dataset, respectively. The experimental results confirm that Attention O-Net outperforms other state-of-the-art detection methods, and is helpful for retinal biometric identification. Min Liu 0008, Fuhao Yu, Tieyong Zeng, Yaonan Wang 0001 |
IEEE J. Biomed. Health Informatics | 2 |
| 2022 | Deep-Learning-Based Automated Neuron Reconstruction From 3D Microscopy Images Using Synthetic Training ImagesabstractDigital reconstruction of neuronal structures from 3D microscopy images is critical for the quantitative investigation of brain circuits and functions. It is a challenging task that would greatly benefit from automatic neuron reconstruction methods. In this paper, we propose a novel method called SPE-DNR that combines spherical-patches extraction (SPE) and deep-learning for neuron reconstruction (DNR). Based on 2D Convolutional Neural Networks (CNNs) and the intensity distribution features extracted by SPE, it determines the tracing directions and classifies voxels into foreground or background. This way, starting from a set of seed points, it automatically traces the neurite centerlines and determines when to stop tracing. To avoid errors caused by imperfect manual reconstructions, we develop an image synthesizing scheme to generate synthetic training images with exact reconstructions. This scheme simulates 3D microscopy imaging conditions as well as structural defects, such as gaps and abrupt radii changes, to improve the visual realism of the synthetic images. To demonstrate the applicability and generalizability of SPE-DNR, we test it on 67 real 3D neuron microscopy images from three datasets. The experimental results show that the proposed SPE-DNR method is robust and competitive compared with other state-of-the-art neuron reconstruction methods. Weixun Chen, Min Liu 0008, Miroslav Radojevic, Yaonan Wang 0001, Erik Meijering |
IEEE Trans. Medical Imaging | 2 |
| 2022 | A 3D Tubular Flux Model for Centerline Extraction in Neuron Volumetric ImagesabstractDigital morphology reconstruction from neuron volumetric images is essential for computational neuroscience. The centerline of the axonal and dendritic tree provides an effective shape representation and serves as a basis for further neuron reconstruction. However, it is still a challenge to directly extract the accurate centerline from the complex neuron structure with poor image quality. In this paper, we propose a neuron centerline extraction method based on a 3D tubular flux model via a two-stage CNN framework. In the first stage, a 3D CNN is used to learn the latent neuron structure features, namely flux features, from neuron images. In the second stage, a light-weight U-Net takes the learned flux features as input to extract the centerline with a spatial weighted average strategy to constrain the multi-voxel width response. Specifically, the labels of flux features in the first stage are generated by the 3D tubular model which calculates the geometric representations of the flux between each voxel in the tubular region and the nearest point on the centerline ground truth. Compared with self-learned features by networks, flux features, as a kind of prior knowledge, explicitly take advantage of the contextual distance and direction distribution information around the centerline, which is beneficial for the precise centerline extraction. Experiments on two challenging datasets demonstrate that the proposed method outperforms other state-of-the-art methods by 18% and 35.1% in F1-measurement and average distance scores at the most, and the extracted centerline is helpful to improve the neuron reconstruction performance. Min Liu 0008, Yaonan Wang 0001, Jiawang Fan, Erik Meijering |
IEEE Trans. Medical Imaging | 2 |
| 2022 | Structure-Guided Segmentation for 3D Neuron ReconstructionabstractDigital reconstruction of neuronal morphologies in 3D microscopy images is critical in the field of neuroscience. However, most existing automatic tracing algorithms cannot obtain accurate neuron reconstruction when processing 3D neuron images contaminated by strong background noises or containing weak filament signals. In this paper, we present a 3D neuron segmentation network named Structure-Guided Segmentation Network (SGSNet) to enhance weak neuronal structures and remove background noises. The network contains a shared encoding path but utilizes two decoding paths called Main Segmentation Branch (MSB) and Structure-Detection Branch (SDB), respectively. MSB is trained on binary labels to acquire the 3D neuron image segmentation maps. However, the segmentation results in challenging datasets often contain structural errors, such as discontinued segments of the weak-signal neuronal structures and missing filaments due to low signal-to-noise ratio (SNR). Therefore, SDB is presented to detect the neuronal structures by regressing neuron distance transform maps. Furthermore, a Structure Attention Module (SAM) is designed to integrate the multi-scale feature maps of the two decoding paths, and provide contextual guidance of structural features from SDB to MSB to improve the final segmentation performance. In the experiments, we evaluate our model in two challenging 3D neuron image datasets, the BigNeuron dataset and the Extended Whole Mouse Brain Sub-image (EWMBS) dataset. When using different tracing methods on the segmented images produced by our method rather than other state-of-the-art segmentation methods, the distance scores gain 42.48% and 35.83% improvement in the BigNeuron dataset and 37.75% and 23.13% in the EWMBS dataset. Bo Yang 0065, Min Liu 0008, Yaonan Wang 0001, Erik Meijering |
IEEE Trans. Medical Imaging | 2 |
| 2021 | Multi-Expert Adversarial Attack Detection in Person Re-identification Using Context InconsistencyabstractThe success of deep neural networks (DNNs) has promoted the widespread applications of person re-identification (ReID). However, ReID systems inherit the vulnerability of DNNs to malicious attacks of visually in-conspicuous adversarial perturbations. Detection of adversarial attacks is, therefore, a fundamental requirement for robust ReID systems. In this work, we propose a Multi-Expert Adversarial Attack Detection (MEAAD) approach to achieve this goal by checking context inconsistency, which is suitable for any DNN-based ReID systems. Specifically, three kinds of context inconsistencies caused by adversarial attacks are employed to learn a detector for distinguishing the perturbed examples, i.e., a) the embedding distances between a perturbed query person image and its top-K retrievals are generally larger than those between a benign query image and its top-K retrievals, b) the embedding distances among the top-K retrievals of a perturbed query image are larger than those of a benign query image, c) the top-K retrievals of a benign query image obtained with multiple expert ReID models tend to be consistent, which is not preserved when attacks are present. Extensive experiments on the Market1501 and DukeMTMC-ReID datasets show that, as the first adversarial attack detection approach for ReID, MEAAD effectively detects various adversarial attacks and achieves high ROC-AUC (over 97.5%). Shasha Li 0001, Min Liu 0008, Yaonan Wang 0001, Amit K. Roy-Chowdhury |
ICCV | 3 |
| 2021 | DeepSeed Local Graph Matching for Densely Packed Cells TrackingabstractThe tracking of densely packed plant cells across microscopy image sequences is very challenging, because their appearance change greatly over time. A local graph matching algorithm was proposed to track such cells by exploiting the tight spatial topology of neighboring cells, and then an iterative searching strategy was used to grow the correspondence from a seed cell pair. Thus, the performance of the existing tracking approach heavily relies on the robustness of finding seed cell pair. However, the existing local graph matching algorithm cannot guarantee the correctness of the seed cell pair, especially in unregistered image sequences or image sequences with large time intervals. In this paper, we propose a DeepSeed local graph matching model to find seed cell pair robustly, by combining local graph matching and CNN-based similarity learning, which uses cells' spatial-temporal contextual information and cell pairs' similarity information. The CNN-based similarity learning is designed to learn cells' deep feature and measure cell pairs' similarity. Compared with the existing plant cell matching methods, the experimental results show that the DeepSeed local graph matching method can track most cells in unregistered image sequences. Moreover, the DeepSeed tracking algorithm can accurately track cells across image sequences with large time intervals. Min Liu 0008, Yalan Liu, Weili Qian, Yaonan Wang 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2021 | Exploiting Global Camera Network Constraints for Unsupervised Video Person Re-IdentificationabstractMany unsupervised approaches have been proposed recently for the video-based re-identification problem since annotations of samples across cameras are time-consuming. However, higher-order relationships across the entire camera network are ignored by these methods, leading to contradictory outputs when matching results from different camera pairs are combined. In this paper, we address the problem of unsupervised video-based re-identification by proposing a consistent cross-view matching (CCM) framework, in which global camera network constraints are exploited to guarantee the matched pairs are with consistency. Specifically, we first propose to utilize the first neighbor of each sample to discover relations among samples and find the groups in each camera. Additionally, a cross-view matching strategy followed by global camera network constraints is proposed to explore the matching relationships across the entire camera network. Finally, we learn metric models for camera pairs progressively by alternatively mining consistent cross-view matching pairs and updating metric models using these obtained matches. Rigorous experiments on two widely-used benchmarks for video re-identification demonstrate the superiority of the proposed method over current state-of-the-art unsupervised methods; for example, on the MARS dataset, our method achieves an improvement of 4.2% over unsupervised methods, and even 2.5% over one-shot supervision-based methods for rank-1 accuracy. Rameswar Panda, Min Liu 0008, Yaonan Wang 0001, Amit K. Roy-Chowdhury |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2021 | HDCB-Net: A Neural Network With the Hybrid Dilated Convolution for Pixel-Level Crack Detection on Concrete BridgesabstractCrack detection on concrete bridges is a critical task to ensure bridge safety. However, many cracks on concrete bridges show low contrast and blurry edges in practice, which brings challenges to image-based crack detection. In this article, to improve the detection accuracy of blurred cracks, we propose the HDCB-Net-a deep learning-based network with the hybrid dilated convolutional block (HDCB) for the pixel-level crack detection. Specifically, HDCB is employed to expand the receptive field of the convolution kernel without increasing the computational complexity and to avoid the gridding effect generated by the dilated convolution. Meanwhile, to achieve a reasonable efficiency/accuracy tradeoff, the HDCB-Net only contains a few downsampling stages, which can avoid the loss of blurred crack pixels due to excessive downsampling. Furthermore, a two-stage strategy is proposed to realize the fast crack detection in a massive number of images (more than 100 000) with the high resolution (5120 × 5120 pixels). At the first stage, YOLOv4 is employed to filter out images without cracks and generate coarse region proposals. At the second stage, to achieve refined damage analysis, the HDCB-Net is used to detect pixel-level cracks from the coarse region proposals. The experimental results demonstrate that the proposed HDCB-Net is genetic and able to improve the detection accuracy of blurred cracks, and our two-stage strategy is efficient for fast crack detection. The whole detection process takes only 0.64 s to handle a single image. Additionally, we have established a public dataset, including 150 632 high-resolution images, dedicated to the research of crack detection, which have been released along with this article. Wenbo Jiang 0002, Min Liu 0008, Yunuo Peng, Lehui Wu, Yaonan Wang 0001 |
IEEE Trans. Ind. Informatics | 2 |
| 2021 | AutoPedestrian: An Automatic Data Augmentation and Loss Function Search Scheme for Pedestrian DetectionabstractPedestrian detection is a challenging and hot research topic in the field of computer vision, especially for the crowded scenes where occlusion happens frequently. In this paper, we propose a novel AutoPedestrian scheme that automatically augments the pedestrian data and searches for suitable loss functions, aiming for better performance of pedestrian detection especially in crowded scenes. To our best knowledge, it is the first work to automatically search the optimal policy of data augmentation and loss function jointly for the pedestrian detection. To achieve the goal of searching the optimal augmentation scheme and loss function jointly, we first formulate the data augmentation policy and loss function as probability distributions based on different hyper-parameters. Then, we apply a double-loop scheme with importance-sampling to solve the optimization problem of data augmentation and loss function types efficiently. Comprehensive experiments on two popular benchmarks of CrowdHuman and CityPersons show the effectiveness of our proposed method. In particular, we achieve 40.58% in MR on CrowdHuman datasets and 11.3% in MR on CityPersons reasonable subset, yielding new state-of-the-art results on these two datasets. Baopu Li, Min Liu 0008, Yaonan Wang 0001, Wanli Ouyang |
IEEE Trans. Image Process. | 3 |
| 2021 | Learning Person Re-Identification Models From Videos With Weak SupervisionabstractMost person re-identification methods, being supervised techniques, suffer from the burden of massive annotation requirement. Unsupervised methods overcome this need for labeled data, but perform poorly compared to the supervised alternatives. In order to cope with this issue, we introduce the problem of learning person re-identification models from videos with weak supervision. The weak nature of the supervision arises from the requirement of video-level labels, i.e. person identities who appear in the video, in contrast to the more precise frame-level annotations. Towards this goal, we propose a multiple instance attention learning framework for person re-identification using such video-level labels. Specifically, we first cast the video person re-identification task into a multiple instance learning setting, in which person images in a video are collected into a bag. The relations between videos with similar labels can be utilized to identify persons, on top of that, we introduce a co-person attention mechanism which mines the similarity correlations between videos with person identities in common. The attention weights are obtained based on all person images instead of person tracklets in a video, making our learned model less affected by noisy annotations. Extensive experiments demonstrate the superiority of the proposed method over the related methods on two weakly labeled person re-identification datasets. Min Liu 0008, Dripta S. Raychaudhuri, Sujoy Paul, Yaonan Wang 0001, Amit K. Roy-Chowdhury |
IEEE Trans. Image Process. | 2 |
| 2021 | Efficient 3D Junction Detection in Biomedical Images Based on a Circular Sampling Model and Reverse MappingabstractDetection and localization of terminations and junctions is a key step in the morphological reconstruction of tree-like structures in images. Previously, a ray-shooting model was proposed to detect termination points automatically. In this paper, we propose an automatic method for 3D junction points detection in biomedical images, relying on a circular sampling model and a 2D-to-3D reverse mapping approach. First, the existing ray-shooting model is improved to a circular sampling model to extract the pixel intensity distribution feature across the potential branches around the point of interest. The computation cost can be reduced dramatically compared to the existing ray-shooting model. Then, the Density-Based Spatial Clustering of Applications with Noise (DBSCAN) algorithm is employed to detect 2D junction points in maximum intensity projections (MIPs) of sub-volume images in a given 3D image, by determining the number of branches in the candidate junction region. Further, a 2D-to-3D reverse mapping approach is used to map these detected 2D junction points in MIPs to the 3D junction points in the original 3D images. The proposed 3D junction point detection method is implemented as a build-in tool in the Vaa3D platform. Experiments on multiple 2D images and 3D images show average precision and recall rates of 87.11% and 88.33% respectively. In addition, the proposed algorithm is dozens of times faster than the existing deep-learning based model. The proposed method has excellent performance in both detection precision and computation efficiency for junction detection even in large-scale biomedical images. Lan Shen, Min Liu 0008, Chao Wang 0072, Changhao Guo, Erik Meijering, Yaonan Wang 0001 |
IEEE J. Biomed. Health Informatics | 2 |
| 2021 | Neuron Image Segmentation via Learning Deep Features and Enhancing Weak Neuronal StructuresabstractNeuron morphology reconstruction (tracing) in 3D volumetric images is critical for neuronal research. However, most existing neuron tracing methods are not applicable in challenging datasets where the neuron images are contaminated by noises or containing weak filament signals. In this paper, we present a two-stage 3D neuron segmentation approach via learning deep features and enhancing weak neuronal structures, to reduce the impact of image noise in the data and enhance the weak-signal neuronal structures. In the first stage, we train a voxel-wise multi-level fully convolutional network (FCN), which specializes in learning deep features, to obtain the 3D neuron image segmentation maps in an end-to-end manner. In the second stage, a ray-shooting model is employed to detect the discontinued segments in segmentation results of the first-stage, and the local neuron diameter of the broken point is estimated and direction of the filamentary fragment is detected by rayburst sampling algorithm. Then, a Hessian-repair model is built to repair the broken structures, by enhancing weak neuronal structures in a fibrous structure determined by the estimated local neuron diameter and the filamentary fragment direction. Experimental results demonstrate that our proposed segmentation approach achieves better segmentation performance than other state-of-the-art methods for 3D neuron segmentation. Compared with the neuron reconstruction results on the segmented images produced by other segmentation methods, the proposed approach gains 47.83% and 34.83% improvement in the average distance scores. The average Precision and Recall rates of the branch point detection with our proposed method are 38.74% and 22.53% higher than the detection results without segmentation. Bo Yang 0065, Weixun Chen, Huiqiong Luo, Yinghui Tan, Min Liu 0008, Yaonan Wang 0001 |
IEEE J. Biomed. Health Informatics | 5 |
| 2021 | Spherical-Patches Extraction for Deep-Learning-Based Critical Points Detection in 3D Neuron Microscopy ImagesabstractDigital reconstruction of neuronal structures is very important to neuroscience research. Many existing reconstruction algorithms require a set of good seed points. 3D neuron critical points, including terminations, branch points and cross-over points, are good candidates for such seed points. However, a method that can simultaneously detect all types of critical points has barely been explored. In this work, we present a method to simultaneously detect all 3 types of 3D critical points in neuron microscopy images, based on a spherical-patches extraction (SPE) method and a 2D multi-stream convolutional neural network (CNN). SPE uses a set of concentric spherical surfaces centered at a given critical point candidate to extract intensity distribution features around the point. Then, a group of 2D spherical patches is generated by projecting the surfaces into 2D rectangular image patches according to the orders of the azimuth and the polar angles. Finally, a 2D multi-stream CNN, in which each stream receives one spherical patch as input, is designed to learn the intensity distribution features from those spherical patches and classify the given critical point candidate into one of four classes: termination, branch point, cross-over point or non-critical point. Experimental results confirm that the proposed method outperforms other state-of-the-art critical points detection methods. The critical points based neuron reconstruction results demonstrate the potential of the detected neuron critical points to be good seed points for neuron reconstruction. Additionally, we have established a public dataset dedicated for neuron critical points detection, which has been released along with this article. Weixun Chen, Min Liu 0008, Qi Zhan, Yinghui Tan, Erik Meijering, Miroslav Radojevic, Yaonan Wang 0001 |
IEEE Trans. Medical Imaging | 2 |
| 2021 | 3D Neuron Microscopy Image Segmentation via the Ray-Shooting Model and a DC-BLSTM NetworkabstractThe morphology reconstruction (tracing) of neurons in 3D microscopy images is important to neuroscience research. However, this task remains very challenging because of the low signal-to-noise ratio (SNR) and the discontinued segments of neurite patterns in the images. In this paper, we present a neuronal structure segmentation method based on the ray-shooting model and the Long Short-Term Memory (LSTM)-based network to enhance the weak-signal neuronal structures and remove background noise in 3D neuron microscopy images. Specifically, the ray-shooting model is used to extract the intensity distribution features within a local region of the image. And we design a neural network based on the dual channel bidirectional LSTM (DC-BLSTM) to detect the foreground voxels according to the voxel-intensity features and boundary-response features extracted by multiple ray-shooting models that are generated in the whole image. This way, we transform the 3D image segmentation task into multiple 1D ray/sequence segmentation tasks, which makes it much easier to label the training samples than many existing Convolutional Neural Network (CNN) based 3D neuron image segmentation methods. In the experiments, we evaluate the performance of our method on the challenging 3D neuron images from two datasets, the BigNeuron dataset and the Whole Mouse Brain Sub-image (WMBS) dataset. Compared with the neuron tracing results on the segmented images produced by other state-of-the-art neuron segmentation methods, our method improves the distance scores by about 32% and 27% in the BigNeuron dataset, and about 38% and 27% in the WMBS dataset. Weixun Chen, Min Liu 0008, Yaonan Wang 0001, Erik Meijering |
IEEE Trans. Medical Imaging | 3 |
| 2020 | Exact and Convergent Iterative Methods to Compute the Orthogonal Point-to-Ellipse DistanceabstractComputation of the orthogonal distance from a given point to an ellipse is the basis of orthogonal distance based ellipse fitting methods. The problem of this orthogonal distance and the corresponding orthogonal contacting point on the ellipse is investigated, and two algorithms, the exact one and the convergent iterative one, are proposed. The exact algorithm utilizes the closed form solution of quartic equations, but is numerically unstable. The iterative algorithm, however, uses Newton's method to solve the equation, and starts from an initial solution that is proven to lead to a convergent iteration. The proposed algorithms are compared in experiments with an existing rival. Although the rival algorithm is slightly faster and more accurate in realistic scenarios, divergence is likely to occur. On the other hand, both our exact and iterative algorithms can reliably produce the solution needed. While the exact algorithm encounters numeric instability, the iterative algorithm is only slightly outperformed by the existing rival in speed and accuracy, but at the same time provides more reliable computation process, thus making it a preferable method for the task. Pingping Hu, Zhigang Ling, He Wen 0003, Min Liu 0008, Lu Tang 0002 |
ICPR | 5 |
| 2020 | DeepBranch: Deep Neural Networks for Branch Point Detection in Biomedical ImagesabstractMorphology reconstruction of tree-like structures in volumetric images, such as neurons, retinal blood vessels, and bronchi, is of fundamental interest for biomedical research. 3D branch points play an important role in many reconstruction applications, especially for graph-based or seed-based reconstruction methods and can help to visualize the morphology structures. There are a few hand-crafted models proposed to detect the branch points. However, they are highly dependent on the empirical setting of the parameters for different images. In this paper, we propose a DeepBranch model for branch point detection with two-level designed convolutional networks, a candidate region segmenter and a false positive reducer. On the first level, an improved 3D U-Net model with anisotropic convolution kernels is employed to detect initial candidates. Compared with the traditional sliding window strategy, the improved 3D U-Net can avoid massive redundant computations and dramatically speed up the detection process by employing dense-inference with fully convolutional neural networks (FCN). On the second level, a method based on multi-scale multi-view convolutional neural networks (MSMV-Net) is proposed for false positive reduction by feeding multi-scale views of 3D volumes into multiple streams of 2D convolution neural networks (CNNs), which can take full advantage of spatial contextual information as well as fit different sizes. Experiments on multiple 3D biomedical images of neurons, retinal blood vessels and bronchi confirm that the proposed 3D branch point detection method outperforms other state-of-the-art detection methods, and is helpful for graph-based or seed-based reconstruction methods. Yinghui Tan, Min Liu 0008, Weixun Chen, Hanchuan Peng, Yaonan Wang 0001 |
IEEE Trans. Medical Imaging | 2 |
| 2019 | CNN-based two-stage cell segmentation improves plant cell tracking
Wenbo Jiang 0002, Lehui Wu, Min Liu 0008 |
Pattern Recognit. Lett. | 4 |
| 2019 | A Multiscale Ray-Shooting Model for Termination Detection of Tree-Like Structures in Biomedical ImagesabstractDigital reconstruction (tracing) of tree-like structures, such as neurons, retinal blood vessels, and bronchi, from volumetric images and 2D images is very important to biomedical research. Many existing reconstruction algorithms rely on a set of good seed points. The 2D or 3D terminations are good candidates for such seed points. In this paper, we propose an automatic method to detect terminations for tree-like structures based on a multiscale ray-shooting model and a termination visual prior. The multiscale ray-shooting model detects 2D terminations by extracting and analyzing the multiscale intensity distribution features around a termination candidate. The range of scale is adaptively determined according to the local neurite diameter estimated by the Rayburst sampling algorithm in combination with the gray-weighted distance transform. The termination visual prior is based on a key observation-when observing a 3D termination from three orthogonal directions without occlusion, we can recognize it in at least two views. Using this prior with the multiscale ray-shooting model, we can detect 3D terminations with high accuracies. Experiments on 3D neuron image stacks, 2D neuron images, 3D bronchus image stacks, and 2D retinal blood vessel images exhibit average precision and recall rates of 87.50% and 90.54%. The experimental results confirm that the proposed method outperforms other the state-of-the-art termination detection methods. Min Liu 0008, Weixun Chen, Chao Wang 0072, Hanchuan Peng |
IEEE Trans. Medical Imaging | 1 |
| 2018 | An Adaptive Ray-Shooting Model for Terminations Detection: Applications in Neuron and Retinal Blood Vessel Images
Weixun Chen, Min Liu 0008, Keran Liu, Zhigang Ling |
BIBM | 2 |
| 2018 | Improved V-Net Based Image Segmentation for 3D Neuron Reconstruction
Min Liu 0008, Huiqiong Luo, Yinghui Tan, Weixun Chen |
BIBM | 1 |
| 2018 | 3D Neuron Branch Points Detection in Microscopy Images
Min Liu 0008, Chao Wang 0072, Weixun Chen |
BIBM | 1 |
| 2018 | Cell Tracking Across Noisy Image Sequences Via Faster R-CNN and Dynamic Local Graph Matching
Min Liu 0008, Lehui Wu, Weili Qian, Yalan Liu |
BIBM | 1 |
| 2018 | Automatic 3D Neuron Tracing Based on Terminations Detection
Chao Wang 0072, Weixun Chen, Min Liu 0008, Zhi Zhou 0006 |
BIBM | 3 |
| 2018 | A Multi-Seed 3D Local Graph Matching Model for Tracking of Densely Packed CellsabstractAutomated tracking of cells in time-lapse live-imaging datasets of developing multicellular tissues is required for high throughput spatio-temporal quantitative measurements of a range of cell behaviors. The tracking of shoot apical meristems (SAM) cells in large-scale microscopy image sequences is challenging, because plant cells are densely packed within a specific honeycomb structure and share very similar physical features. In this paper, we propose a 3D local graph matching model to track the plant SAM cells, by exploiting the cells' tight spatial and temporal contextual information. The proposed 3D local graph matching model is further combined with a multi-seed based majority voting scheme to rectify possible matching errors in the cell correspondence growing process. Compared with the existing 2D local graph matching model, the experimental results show that the proposed method can greatly improve the tracking accuracy for plant cells. Min Liu 0008, Yalan Liu, Weili Qian |
ICASSP | 1 |
| 2018 | Local Neuron Radius Estimation in Volumetric Microscopy ImagesabstractAccurate estimation of local neuron size of complex 3D neuron structures imaged by microscopy is very important for quantification of neuronal morphology, such as neuron tracing applications. Previously, the Rayburst sampling algorithm was developed to estimate the neuron size in volumetric microscopy images. However, it requires an accurate detection of the neuron centerline, which is not easy and time-consuming. In this paper, we propose to estimate the neuron local radius of any region of interest in a given neuron, by first detecting its corresponding neuron centerline point using Multistencils Fast Marching (MSFM) method and then estimating the local neuron radius by the Rayburst sampling algorithm. The experimental results on different datasets show that the proposed method could obtain very accurate local neuron radius efficiently. Min Liu 0008, Keran Liu, Chao Wang 0072, Huiqiong Luo |
ICIP | 1 |
| 2018 | Convolutional Neural Network Cascade Based Neuron Termination Detection in 3D Image StacksabstractFull reconstruction of neuron morphology in volumetric images is of fundamental interest for the analysis and understanding of neuron function. Termination points could be very good candidates of seeding points for neuron reconstructing (tracing) applications. Previously, some hand-crafted models were proposed to detect the neuron terminations. However, they are highly depending on empirical setting of the parameters for different images. In this paper, we propose a neuron termination detection approach with two-level designed convolutional networks. At first level, a Triple-Crossing 2.5D convolutional neural network with inception block and residual block is used to generate the termination candidate points by early identification of `many' volumetric patches. At second level, those termination candidates are further evaluated by using the cubic patches across the adjacent slices of those candidates. It is shown that the convolutional neural network cascade based detection approach outperforms the current top-performing neuron termination detection methods in many challenging datasets. Yinghui Tan, Huiqiong Luo, Min Liu 0008 |
ICIP | 4 |
| 2018 | Multi-View Deep Metric Learning for Volumetric Image RecognitionabstractThis paper presents a multi-view deep metric learning (MVDML) architecture for the recognition of volumetric image stacks. Different from existing metric learning methods which aim to learn a Mahalanobis distance metric to maximize the inter-class variations and minimize the intra-class variations, the proposed multi-view deep metric learning approach learns a function that maps input volumetric images into a compact Euclidean space where distances approximate the “semantic” distances in the input space. The learning process minimizes a contrastive loss function that drives the similarity metric to be small for pairs of samples from same class, and large for pairs from different classes. The mapping from input to the target space is a multi-view convolutional neural network (MVCNN) which combines information from multiple views of a volumetric image into a single and compact feature descriptor. The experimental results on the nematode volumetric image database show that our proposed method outperforms models based on hand-crafted visual features, conventional metric learning methods and deep classification models. Min Liu 0008 |
ICME | 2 |
| 2018 | Convolutional Features-Based CRF Graph Matching for Tracking of Densely Packed CellsabstractThe tracking of plant cells across large-scale microscopy image sequences is very challenging, because plant cells are densely packed in a specific honeycomb structure, and the microscopy images can be randomly translated, rotated and scaled in the imaging process. This paper proposes a convolutional features-based conditional random field (CRF) graph matching method to track plant cells in unregistered image sequences, by exploiting deep features extracted from deep convolutional neural networks and tight spatial topology feature of neighboring cells as contextual information. Because the extracted convolutional feature and spatial topology feature are resilient to image translation, rotation and scaling, the proposed CRF matching approach is able to track plant cells across unregistered image sequences. Compared with other plant cell tracking methods, the experimental results show that the proposed method improves the tracking accuracy rate by about 30% in the unregistered cell image sequences. Weili Qian, Yangliu Wei, Min Liu 0008 |
ICPR | 4 |
| 2018 | A multi-seed dynamic local graph matching model for tracking of densely packed cells across unregistered microscopy image sequences
Min Liu 0008, Jieqin Li, Weili Qian |
Mach. Vis. Appl. | 1 |
| 2018 | Multi-focal nematode image stack classification using a projection-based multi-linear method
Min Liu 0008, Keran Liu |
Mach. Vis. Appl. | 1 |
| 2018 | 3D neuron tip detection in volumetric microscopy images using an adaptive ray-shooting model
Min Liu 0008, Rong Gong, Weixun Chen, Hanchuan Peng |
Pattern Recognit. | 1 |
| 2018 | Cell Population Tracking in a Honeycomb Structure Using an IMM Filter Based 3D Local Graph Matching ModelabstractDeveloping algorithms for plant cell population tracking is very critical for the modeling of plant cell growth pattern and gene expression dynamics. The tracking of plant cells in microscopic image stacks is very challenging for several reasons: (1) plant cells are densely packed in a specific honeycomb structure; (2) they are frequently dividing; and (3) they are imaged in different layers within 3D image stacks. Based on an existing 2D local graph matching algorithm, this paper focuses on building a 3D plant cell matching model, by exploiting the cells' 3D spatiotemporal context. Furthermore, the Interacting Multi-Model filter (IMM) is combined with the 3D local graph matching model to track the plant cell population simultaneously. Because our tracking algorithm does not require the identification of "tracking seeds", the tracking stability and efficiency are greatly enhanced. Last, the plant cell lineages are achieved by associating the cell tracklets, using a maximum-a-posteriori (MAP) method. Compared with the 2D matching method, the experimental results on multiple datasets show that our proposed approach does not only greatly improve the tracking accuracy by 18 percent, but also successfully tracks the plant cells located at the high curvature primordial region, which is not addressed in previous work. Min Liu 0008, Yue He 0003, Weili Qian, Yangliu Wei |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2017 | IMM filter based local graph matching for plant cell lineage estimationabstractIn this paper, an interacting multiple models (IMM) motion filter based local graph matching method is proposed to track the plant cells, by exploiting the tight spatial topology of neighboring cells in a multicellular field as contextual information. The IMM filter is used to predict the movement of the cells, and then the local graph matching approach is used to search the target cells in the local neighborhood of the predicted position. The combination of the IMM filter and local graph matching greatly reduces the size of the searching region in the matching process and enhances the tracking stability as well. Furthermore, the cells' lineages are generated by using a maximum-a-posteriori (MAP) lineage association method. The effectiveness and efficiency of the proposed tracking method are validated by experiments on real plant cell datasets. Min Liu 0008, Yue He 0003, Jieqin Li, Hongzhong Zhang |
ICIP | 1 |
| 2017 | Classification of multi-focal nematode image stacks using a projection based multilinear approachabstractIn this paper, we propose to use projection methods such as coefficient of variation projection (COV) to exploit the entire information of Digital Multi-focal Images (DMI) using its projection images along different directions. The COV projection takes into account the intensity distribution feature of multi-focal images, so it overcomes the limitation of poor contrast of the projection images from the 3D X-Ray Transform, which is used in a previous work. Because the DMI stacks represent the effect of different factors — texture, projection directions, different instances within the same class and different classes of objects, we embed the projection method within a multilinear classification framework. The experimental results on the nematode data show that the image projection based multilinear classifier can achieve very reliable recognition rate (95.5%), even we only use the texture feature instead of the combination of texture and shape features as in the previous work. Min Liu 0008, Hongzhong Zhang |
ICIP | 1 |
| 2017 | A multi-direction image fusion based approach for classification of multi-focal nematode image stacksabstractIn this paper, we present to use a multi-direction image fusion based feature extraction approach to classify multi-focal image stacks. The discrete wavelet transform sparse representation (DWTSR) image fusion technique is used to combine relevant information from a given image stack into a single image, which is more informative and complete than any single individual image within the given stack. Besides, multi-focal images within a multi-focal stack are fused along 3 orthogonal directions, and multiple features extracted from the fused images along different directions are combined by using canonical correlation analysis (CCA). The experimental results on the nematode multi-focal images show that our proposed multi-direction image fusion based feature extraction method can improve the recognition rate from 83.8% in the previous work to 96% by using texture feature only. Min Liu 0008, Hongzhong Zhang |
ICIP | 1 |
| 2017 | Plant cell tracking using Kalman filter based local graph matching
Min Liu 0008, Yue He 0003, Yangliu Wei |
Image Vis. Comput. | 1 |
| 2017 | Classification of nematode image stacks by an information fusion based multilinear approach
Min Liu 0008, Hongzhong Zhang |
Pattern Recognit. Lett. | 1 |
| 2017 | Robust Plant Cell Tracking in Noisy Image Sequences Using Optimal CRF Graph MatchingabstractIn time-lapse live-imaging datasets of developing multicellular tissues, automated tracking of cells is required for high-throughput spatiotemporal quantitative measurements of a range of cell behaviors. This letter proposes a conditional random field (CRF) graph matching method to track plant cells in noisy images by exploiting the tight spatial topology of neighboring cells in a multicellular field as contextual information. The CRF potential of cells dynamically changes during the cell correspondence growing process, because the cells that have been matched already are not included in the calculation of the second-order potential. Therefore, the proposed CRF-based tracker tends to reduce tracking errors, while the previous local graph matching method tends to accumulate errors during the cell correspondence growing process. The CRF graph matching method greatly improves the tracking accuracy in noisy images and enhances the tracking stability because it always matches the most reliable cell pairs with the least CRF potential in the neighboring system. Compared with the previous method, the experimental results show that the proposed method can improve the tracking accuracy rate by 10% in noisy image sequences. Min Liu 0008, Yangliu Wei, Weili Qian, Hongzhong Zhang |
IEEE Signal Process. Lett. | 1 |
| 2016 | Robust plant cell tracking using local spatio-temporal context
Min Liu 0008, Guocai Liu |
Neurocomputing | 1 |
| 2015 | Progressive Learning Machine: A New Approach for General Hybrid System ApproximationabstractAs the most important property of neural networks (NNs), the universal approximation capability of NNs is widely used in many applications. However, this property is generally proven for continuous systems. Most industrial systems are hybrid systems (e.g., piecewise continuous), which is a significant limitation for real applications. Recently, many identification methods have been proposed for hybrid system approximation; however, these methods only operate in linear hybrid systems. In this paper, the progressive learning machine-a new learning algorithm based on multi-NNs-is proposed for general hybrid nonlinear/linear system approximation. This algorithm classifies hybrid systems into several continuous systems and can approximate any hybrid system with zero output error. The performance of the proposed learning method is demonstrated via numerical examples and with experimental data from real applications. Yimin Yang 0001, Yaonan Wang 0001, Q. M. Jonathan Wu, Min Liu 0008 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2011 | 3D Neuron Tip Detection in Volumetric Microscopy ImagesabstractThis paper addresses the problem of 3D neuron tips detection in volumetric microscopy image stacks. We focus particularly on neuron tracing applications, where the detected 3D tips could be used as the seeding points. Most of the existing neuron tracing methods require a good choice of seeding points. In this paper, we propose an automated neuron tips detection method for volumetric microscopy image stacks. Our method is based on first detecting 2D tips using curvature information and a ray-shooting intensity distribution model, and then extending it to the 3D stack by rejecting false positives. We tested this method based on the V3D platform, which can reconstruct a neuron based on automated searching of the optimal 'paths' connecting those detected 3D tips. The experiments demonstrate the effectiveness of the proposed method in building a fully automatic neuron tracing system. Min Liu 0008, Hanchuan Peng, Amit K. Roy-Chowdhury, Eugene W. Myers |
BIBM | 1 |
| 2011 | Efficient cell segmentation and tracking of developing plant meristemabstractAnalysis of Confocal Laser Scanning Microscopy (CLSM) images is gaining popularity in developmental biology for understanding growth dynamics. The automated analysis of such images is highly desirable for efficiency and accuracy. The first step in this process is segmentation and tracking leading to computation of cell lineages. In this paper, we present efficient, accurate, and robust segmentation and tracking algorithms for cells and detection of cell divisions in a 4D spatio-temporal image stack of a growing plant meristem. We show how to optimally choose the parameters in the watershed algorithm for high quality segmentation results. This yields high quality tracking results using cell correspondence evaluation functions. We show segmentation and tracking results on Confocal laser scanning microscopy data captured for 72 hours at every 3 hour intervals. Compared to recent results in this area, the proposed algorithms provide significantly longer cell lineages and more comprehensive identification of cell divisions. Katya Mkrtchyan, Damanpreet Singh, Min Liu 0008, G. Venugopala Reddy, Amit K. Roy-Chowdhury, Meenakshisundaram Gopi |
ICIP | 3 |
| 2010 | Multilinear feature extraction and classification of multi-focal images, with applications in nematode taxonomyabstractIn this paper, we present a 3D X-Ray Transform based multilinear feature extraction and classification method for Digital Multi-focal Images (DMI). In such images, morphological information for a transparent specimen can be captured in the form of a stack of high-quality images, representing individual focal planes through the specimen's body. We present a method that can effectively exploit the entire information in the stack using the 3D X-Ray projections at different viewing angles. These DMI stacks represent the effect of different factors - shape, texture, viewpoint, different instances within the same class and different classes of specimens. For this purpose, we embed the 3D X-Ray Transform within a multilinear framework and propose a Multilinear X-Ray Transform (MXRT) feature representation. By combining the tensor texture and shape information we can get better recognition rates than just relying on the original or key frames of DMI stacks. The experimental results on the nematode DMI data show that the 3D X-Ray Transform based multilinear analysis method can effectively give 100% recognition rate on a real-life database. Min Liu 0008, Amit K. Roy-Chowdhury |
CVPR | 1 |
| 2010 | Multi-target tracking using long-term stochastic associationsabstractMaintaining the stability of tracks on multiple targets in video over extended time periods remains a challenging problem. A few methods which have recently shown encouraging results in this direction rely on learning context models or the availability of training data. However, this may not be feasible in many application scenarios. Moreover, tracking methods should be able to work across multiple resolutions of the video. In this paper, we consider the problem of long-term tracking in video in application domains where context information is not available a priori, nor can it be learned online. We build our solution on the hypothesis that most existing trackers can obtain reasonable short-term tracks (tracklets). By analyzing the statistical properties of these tracklets, we develop associations between them so as to come up with longer tracks. On multiple real-life video sequences spanning low and high resolution data, we show the ability to accurately track over extended time periods. Ting-Yueh Jeng, Bi Song, Elliot Staudt, Min Liu 0008, Amit K. Roy-Chowdhury, Ashis SenGupta |
ICIP | 4 |
| 2010 | Multi-focal nematode image classification using the 3D X-Ray TransformabstractIn this paper, we present a 3D X-Ray Transform based feature extraction and classification method for Digital Multi-focal Images (DMI). In such images, morphological information for a transparent specimen can be captured in the form of a stack of high-quality images, representing individual focal planes through the specimen's body. We present a method that can effectively exploit the entire information in the stack using the 3D X-Ray Transform at different angle views. By combining the texture and shape information from different angles, we can get better recognition rates than just relying on the original or key frames of DMI stacks. The experimental results on the nematode DMI data show that the 3D X-Ray Transform based classification method can effectively improve the recognition rate from 60% (PCA) to 96.8%. Min Liu 0008, Amit K. Roy-Chowdhury, Melissa Yoder, Paul De Ley |
ICIP | 1 |
| 2010 | Pattern analysis of stem cell growth dynamics in the shoot apex of arabidopsisabstractThe Shoot Apical Meristem (SAM) is made of stem cells that are responsible for all above ground plant structures. Differentiating cells in the development of SAM form primordia. Primordia develop to become various plant organs. Understanding the growth dynamics of primordia is critical to understanding the developmental dynamics of the entire SAM. We present a method for performing quantitative analysis of primordia development in model plant Arabidopsis thaliana. A contour based approach is used to detect and isolate individual primordia from 3D live imaging data. Regions of growth are detected by analyzing eigenvalues of curvature covariance matrices. After primordia detection and isolation, a Dynamic Time Warping (DTW) Algorithm is applied to compute the rate of growth. Results show the successful use of our method to quantitatively analyze primordial growth. Oben M. Tataw, Min Liu 0008, Amit K. Roy-Chowdhury, Ram Kishor Yadav, G. Venugopala Reddy |
ICIP | 2 |
| 2009 | Exploiting local structure for tracking plant cells in noisy imagesabstractIn this paper, we present a local graph matching based method for tracking cells and cell divisions in noisy images. We work with plant cells, where the cells are tightly clustered in space and computing correspondences across time can be very challenging. The local graph matching method is able to track the cells and cell divisions even when significant portions of the images are corrupted due to sensor noise in the imaging process. The geometric structure and topology of the cells' relative positions are efficiently exploited to solve the tracking problem using the local graph matching technique. Using this method we can track almost all of the properly segmented cells, even when some of those images are highly noisy. Min Liu 0008, Amit K. Roy-Chowdhury, G. Venugopala Reddy |
ICIP | 1 |