Min Xu 0009

dblp:09/0-9 · DBLP profile ↗
← Back
82ranked-venue papers
5as first author
48since 2021 · last 2026
0000-0002-0881-5891ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 39 · 5 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 33 · 24 since 2021Artificial intelligence and machine learning · 32 · 25 since 2021
YearPublicationVenuePosition
2026 DiLO: Disentangled Latent Optimization for Learning Shape and Deformation in Grouped Deforming 3D Objects
abstract
In this work, we propose a disentangled latent optimization-based method for parameterizing grouped deforming 3D objects into shape and deformation factors in an unsupervised manner. Our approach involves the joint optimization of a generator network along with the shape and deformation factors, supported by specific regularization techniques. For efficient amortized inference of disentangled shape and deformation codes, we train two order-invariant PointNet-based encoder networks in the second stage of our method. We demonstrate several significant downstream applications of our method, including unsupervised deformation transfer, deformation classification, and explainability analyses. Extensive experiments conducted on 3D human, animal, and facial expression datasets demonstrate that our simple approach is highly effective in these downstream tasks, comparable or superior to existing methods with much higher complexity.
Mostofa Rafid Uddin, Jana Armouti, Umong Sain, Md. Asib Rahman, Xingjian Li 0002, Min Xu 0009
AAAI6
2026 From Unsupervised to Zero-Shot: 3D Domain Adaptation for Cryo-Electron Tomography Segmentation
abstract
Abstract With the refinement of 3D imaging modality, cryo-electron tomography (cryo-ET) has emerged as a powerful technique for the structural analysis of macromolecular complexes at near-atomic resolution. Recent advancements in volumetric segmentation methods applied to cryo-ET datasets have garnered significant attention within the biomedical sector. However, existing methods rely heavily on manually labeled data, which demands highly specialized expertise, making fully supervised approaches less feasible for cryo-ET images. To address this, a number of unsupervised domain adaptation (UDA) techniques have been developed to improve segmentation network performance using unlabeled data. Nevertheless, directly applying these methods to cryo-ET image segmentation presents two major challenges: 1) the source dataset, usually obtained through simulation, contains a fixed level of noise, while the target dataset, being directly collected from raw-data from the real-world scenario, has unpredictable noise levels; 2) the source data used for training typically consists of known macromolecules, in contrast, the target domain data are often unknown, causing the model to be biased towards those known macromolecules, leading to a domain shift problem. To address such challenges, in this paper, we introduce a voxel-wise unsupervised domain adaptation approach, termed Vox-UDA, specifically for cryo-ET subtomogram segmentation. Vox-UDA incorporates a noise generation module to simulate target-like noises in the source dataset for cross-noise level adaptation, and a denoised pseudo-labeling strategy based on the improved bilateral filter to alleviate the domain shift problem. Additionally, we further consider a scenario that is more in line with the real world, where the target (experimental) dataset might not be accessible during training or has a very small sample size, and we present a voxel-wise zero-shot domain adaptation (ZSDA) approach, named Vox-ZSDA. In Vox-ZSDA, we introduce a self-supervised graph learning strategy to eliminate any dependency on the target data, accompanied by a dynamic graph contrastive learning technique to enhance the model’s sensitivity to macromolecular structures to boost the segmentation performance. More importantly, we construct the first UDA and ZSDA cryo-ET subtomogram segmentation benchmark on three experimental datasets. Extensive experimental results on multiple benchmarks and newly curated real-world datasets demonstrate the superiority of our proposed approach compared to state-of-the-art UDA and ZSDA methods.
Haoran Li 0024, Xingjian Li 0002, Jiahua Shi, Huaming Chen, Bo Du 0004, Johan Barthélemy, Daisuke Kihara, Jun Shen 0001, Min Xu 0009
Int. J. Comput. Vis.10
2026 Enhancing weakly supervised 3D medical image segmentation through probabilistic-aware learning
abstract
3D medical image segmentation is a challenging task with crucial implications for disease diagnosis and treatment planning. Recent advances in deep learning have significantly enhanced fully supervised medical image segmentation. However, this approach heavily relies on labor-intensive and time-consuming fully annotated ground-truth labels, particularly for 3D volumes. To overcome this limitation, we propose a novel probabilistic-aware weakly supervised learning pipeline, specifically designed for 3D medical imaging. Our pipeline integrates three innovative components: a Probability-based Pseudo-Label Generation technique for synthesizing dense segmentation masks from sparse annotations, a Probabilistic Multi-head Self-Attention network for robust feature extraction within our Probabilistic Transformer Network, and a Probability-informed Segmentation Loss Function to enhance training with annotation confidence. Demonstrating significant advances, our approach not only rivals the performance of fully supervised methods but also surpasses existing weakly supervised methods in CT and MRI datasets, achieving up to 18.1% improvement in Dice scores for certain organs. The code is available at https://github.com/runminjiang/PW4MedSeg .
Runmin Jiang, Zhaoxin Fan, Junhao Wu 0003, Lenghan Zhu, Tianyang Wang 0004, Heng Huang 0001, Min Xu 0009
Pattern Recognit.8
2026 Representation-centric survey of supervised skeletal action recognition and the new benchmark
abstract
3D skeletal action recognition has emerged as a powerful alternative to traditional RGB and depth-based approaches, offering robustness to environmental variations, computational efficiency, and enhanced privacy. Despite remarkable progress, current research remains fragmented across diverse input representations and lacks evaluation under scenarios that reflect real-world challenges. This paper presents a representation-centric review of supervised skeletal action recognition, systematically categorizing state-of-the-art methods by their input feature types: joint coordinates, bone vectors, motion flows, and extended representations, and analyzing how these choices influence spatiotemporal modeling strategies. Building on the insights from this review, we introduce ANUBIS, a large-scale, challenging dataset designed to address critical gaps in existing benchmarks. ANUBIS incorporates multi-view recordings with back-view perspectives, complex multi-person interactions, fine-grained and violent actions, and contemporary social behaviors. We benchmark a diverse set of state-of-the-art models on ANUBIS and conduct an in-depth analysis of how different feature types affect recognition performance across 102 action categories. Our results show strong action-feature dependencies, highlight the limitations of naïve multi-representational fusion, and point toward the need for task-aware, semantically aligned integration strategies. This work offers both a comprehensive foundation and a practical benchmarking resource, aiming to guide the next generation of robust, generalizable skeleton-based action recognition systems for complex real-world scenarios. The dataset, benchmarking framework, and code are available at https://yliu1082.github.io/ANUBIS/ .
Yang Liu 0249, Jiyao Yang, Madhawa Perera, Pan Ji, Dongwoo Kim 0002, Min Xu 0009, Tianyang Wang 0004, Saeed Anwar, Tom Gedeon, Lei Wang 0108, Zhenyue Qin
Pattern Recognit.6
2026 Adaptive knowledge transferring with switching dual-student framework for semi-supervised medical image segmentation
abstract
Teacher-student frameworks have emerged as a leading approach in semi-supervised medical image segmentation, demonstrating strong performance across various tasks. However, the learning effects are still limited by the strong correlation and unreliable knowledge transfer process between teacher and student networks. To overcome this limitation, we introduce a novel switching Dual-Student architecture that strategically selects the most reliable student at each iteration to enhance dual-student collaboration and prevent error reinforcement. We also introduce a strategy of Loss-Aware Exponential Moving Average to dynamically ensure that the teacher absorbs meaningful information from students, improving the quality of pseudo-labels. Our plug-and-play framework is extensively evaluated on 3D medical image segmentation datasets, where it outperforms state-of-the-art semi-supervised methods, demonstrating its effectiveness in improving segmentation accuracy under limited supervision.
Hoang-Thien Nguyen, Thanh-Huy Nguyen, Ba-Thinh Lam, Vi Vu, Bach X. Nguyen, Jianhua Xing, Tianyang Wang 0004, Xingjian Li 0002, Min Xu 0009
Pattern Recognit.9
2026 FUGC: Benchmarking Semi-Supervised Learning Methods for Cervical Segmentation
abstract
Accurate segmentation of cervical structures in transvaginal ultrasound (TVS) is critical for assessing the risk of spontaneous preterm birth (PTB), yet the scarcity of labeled data limits the performance of supervised learning approaches. This paper introduces the Fetal Ultrasound Grand Challenge (FUGC), the first benchmark for semi-supervised learning in cervical segmentation, hosted at ISBI 2025. FUGC provides a dataset of 890 TVS images, including 500 training images, 90 validation images, and 300 test images. Methods were evaluated using the Dice Similarity Coefficient (DSC), Hausdorff Distance (HD), and runtime (RT), with a weighted combination of 0.4/0.4/0.2. The challenge attracted 10 teams with 82 participants submitting innovative solutions. The best-performing methods for each individual metric achieved 90.26% mDSC, 38.88 mHD, and 32.85 ms RT, respectively. FUGC establishes a standardized benchmark for cervical segmentation, demonstrates the efficacy of semi-supervised methods with limited labeled data, and provides a foundation for AI-assisted clinical PTB risk assessment.
Jieyun Bai, Yitong Tang, Mahdi Islam, Musarrat Tabassum, Enrique Almar-Munoz, Nianjiang Lv, Yu Chen 0099, Zilun Peng, Yusong Xiao, Li Xiao 0002, Nam-Khanh Tran, Dac-Phu Phan-Le, Hai-Dang Nguyen, Xiao Liu 0037, Jiale Hu, Mingxu Huang, Jitao Liang, Chaolu Feng, Xuezhi Zhang, Lyuyang Tong, Bo Du 0001, Ha-Hieu Pham, Thanh-Huy Nguyen, Min Xu 0009, Juntao Jiang, Jiangning Zhang, Yong Liu 0007, Md. Kamrul Hasan 0002, Zhuonan Liang, Tom Weidong Cai, Gongning Luo, Mohammad Yaqub, Karim Lekadir
IEEE Trans. Medical Imaging28
2025 Vox-UDA: Voxel-wise Unsupervised Domain Adaptation for Cryo-Electron Subtomogram Segmentation with Denoised Pseudo-Labeling
abstract
Cryo-Electron Tomography (cryo-ET) is a 3D imaging technology that facilitates the study of macromolecular structures at near-atomic resolution. Recent volumetric segmentation approaches on cryo-ET images have drawn widespread interest in the biological sector. However, existing methods heavily rely on manually labeled data, which requires highly professional skills, thereby hindering the adoption of fully-supervised approaches for cryo-ET images. Some unsupervised domain adaptation (UDA) approaches have been designed to enhance the segmentation network performance using unlabeled data. However, applying these methods directly to cryo-ET image segmentation tasks remains challenging due to two main issues: 1) the source dataset, usually obtained through simulation, contains a fixed level of noise, while the target dataset, directly collected from raw-data from the real-world scenario, have unpredictable noise levels. 2) the source data used for training typically consists of known macromoleculars. In contrast, the target domain data are often unknown, causing the model to be biased towards those known macromolecules, leading to a domain shift problem. To address such challenges, in this work, we introduce a voxel-wise unsupervised domain adaptation approach, termed Vox-UDA, specifically for cryo-ET subtomogram segmentation. Vox-UDA incorporates a noise generation module to simulate target-like noises in the source dataset for cross-noise level adaptation. Additionally, we propose a denoised pseudo-labeling strategy based on the improved Bilateral Filter to alleviate the domain shift problem. More importantly, we construct the first UDA cryo-ET subtomogram segmentation benchmark on three experimental datasets. Extensive experimental results on multiple benchmarks and newly curated real-world datasets demonstrate the superiority of our proposed approach compared to state-of-the-art UDA methods.
Haoran Li 0024, Xingjian Li 0002, Jiahua Shi, Huaming Chen, Bo Du 0004, Daisuke Kihara, Johan Barthélemy, Jun Shen 0001, Min Xu 0009
AAAI9
2025 Agentic Reasoning: A Streamlined Framework for Enhancing LLM Reasoning with Agentic Tools
abstract
We introduce Agentic Reasoning, a framework that enhances large language model (LLM) reasoning by integrating external tool-using agents.Agentic Reasoning dynamically leverages web search, code execution, and structured memory to address complex problems requiring deep research.A key innovation in our framework is the Mind-Map agent, which constructs a structured knowledge graph to store reasoning context and track logical relationships, ensuring coherence in long reasoning chains with extensive tool usage.Additionally, we conduct a comprehensive exploration of the Web-Search agent, leading to a highly effective search mechanism that surpasses all prior approaches.When deployed on DeepSeek-R1, our method achieves a new state-of-the-art (SOTA) among public models and delivers performance comparable to OpenAI Deep Research, the leading proprietary model in this domain.Extensive ablation studies validate the optimal selection of agentic tools and confirm the effectiveness of our Mind-Map and Web-Search agents in enhancing LLM reasoning.Our code and data are publicly available.
Jiayuan Zhu, Yuyuan Liu, Min Xu 0009, Yueming Jin
ACL (1)4
2025 Medical Graph RAG: Evidence-based Medical Large Language Model via Graph Retrieval-Augmented Generation
abstract
Junde Wu, Jiayuan Zhu, Yunli Qi, Jingkun Chen, Min Xu, Filippo Menolascina, Yueming Jin, Vicente Grau. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Jiayuan Zhu, Yunli Qi, Jingkun Chen, Min Xu 0009, Filippo Menolascina, Yueming Jin, Vicente Grau
ACL (1)5
2025 BOE-ViT: Boosting Orientation Estimation with Equivariance in Self-Supervised 3D Subtomogram Alignment
abstract
Subtomogram alignment is a critical task in cryo-electron tomography (cryo-ET) analysis, essential for achieving high-resolution reconstructions of macromolecular complexes. However, learning effective positional representations remains challenging due to limited labels and high noise levels inherent in cryo-ET data. In this work, we address this challenge by proposing a self-supervised learning approach that leverages intrinsic geometric transformations as implicit supervisory signals, enabling robust representation learning despite data scarcity. We introduce BOE-ViT, the first Vision Transformer (ViT) framework for 3D subtomogram alignment. Recognizing that traditional ViTs lack equivariance and are therefore suboptimal for orientation estimation, we enhance the model with two innovative modules that introduce equivariance include 1) the Polyshift module for improved shift estimation and 2) Multi-Axis Rotation Encoding (MARE) for enhanced rotation estimation. Experimental results demonstrate that BOE-ViT significantly outperforms state-of-the-art methods. Notably, at SNR 0.01 dataset, our approach achieves a 77.3% reduction in rotation estimation error and a 62.5% reduction in translation estimation error, effectively overcoming the challenges in cryo-ET subtomogram alignment.
Runmin Jiang, Jackson Daggett, Shriya Pingulkar, Priyanshu Dhingra, Qifeng Wu, Xingjian Li 0002, Min Xu 0009
CVPR10
2025 DiffCAM: Data-Driven Saliency Maps by Capturing Feature Differences
abstract
In recent years, the interpretability of Deep Neural Networks (DNNs) has garnered significant attention, particularly due to their widespread deployment in critical domains like healthcare, finance, and autonomous systems. To address the challenge of understanding how DNNs make decisions, Explainable AI (XAI) methods, such as saliency maps, have been developed to provide insights into the inner workings of these models. This paper introduces DiffCAM, a novel XAI method designed to overcome limitations in existing Class Activation Map (CAM)-based techniques, which often rely on decision boundary gradients to estimate feature importance. DiffCAM differentiates itself by considering the actual data distribution of the reference class, identifying feature importance based on how a target example differs from reference examples. This approach captures the most discriminative features without relying on decision boundaries or prediction results, making DiffCAM applicable to a broader range of models, including foundation models. Through extensive experiments, we demonstrate the superior performance and flexibility of DiffCAM in providing meaningful explanations across diverse datasets and scenarios.
Xingjian Li 0002, Qiming Zhao, Neelesh Bisht, Mostofa Rafid Uddin, Jin Yu Kim, Bryan Zhang, Min Xu 0009
CVPR7
2025 TD-RD: A Top-Down Benchmark with Real-Time Framework for Road Damage Detection
abstract
Object detection has witnessed remarkable advancements over the past decade, largely driven by breakthroughs in deep learning and the proliferation of large-scale datasets. However, the domain of road damage detection remains relatively underexplored, despite its critical significance for applications such as infrastructure maintenance and road safety. This paper addresses this gap by introducing a novel top-down benchmark that offers a complementary perspective to existing datasets, specifically tailored for road damage detection. Our proposed Top-Down Road Damage Detection Dataset (TD-RD) includes three primary categories of road damage—cracks, potholes, and patches—captured from an top-down viewpoint. The dataset consists of 7,088 high-resolution images, encompassing 12,882 annotated instances of road damage. Additionally, we present a novel real-time object detection framework, TD-YOLOV10, designed to handle the unique challenges posed by the TD-RD dataset. Comparative studies with state-of-the-art models demonstrate competitive baseline results. By releasing TD-RD, we aim to accelerate research in this crucial area. A sample of the dataset will be made publicly available upon the paper’s acceptance.
Xi Xiao 0003, Zhengji Li, Houjie Lin, Swalpa Kumar Roy, Tianyang Wang 0004, Min Xu 0009
ICASSP8
2025 Unsupervised Identification of Protein Compositions and Conformations Via Implicit Content-Transformation Disentanglement
abstract
, a novel contrastive learning-based method that implicitly parameterizes and disentangles both transformation and content. DualContrast achieves this by generating positive and negative pairs for content and transformation in both data and latent spaces. We demonstrate that existing self-supervised approaches fail under similar implicit parameterization, underscoring the necessity of our method. Through extensive experiments on 3D microscopic images of protein mixtures and additional shape-focused datasets beyond microscopy, we validate our claims and demonstrate the first in-principle fully unsupervised identification of different protein compositions and conformations in 3D microscopic images.
Mostofa Rafid Uddin, Jana Armouti, Min Xu 0009
ICCV3
2025 Toward Material-Agnostic System Identification From Videos
Chunjiang Liu, Charles Herrmann, Junhwa Hur, Yinxiao Li, Ming-Hsuan Yang 0001, Bhiksha Raj, Min Xu 0009
ICCV10
2025 Visual Instance-aware Prompt Tuning
abstract
Visual Prompt Tuning (VPT) has emerged as a parameter-efficient fine-tuning paradigm for vision transformers, with conventional approaches utilizing dataset-level prompts that remain the same across all input instances. We observe that this strategy results in sub-optimal performance due to high variance in downstream datasets. To address this challenge, we propose Visual Instance-aware Prompt Tuning (ViaPT), which generates instance-aware prompts based on each individual input and fuses them with dataset-level prompts, leveraging Principal Component Analysis (PCA) to retain important prompting information. Moreover, we reveal that VPT-Deep and VPT-Shallow represent two corner cases based on a conceptual understanding, in which they fail to effectively capture instance-specific information, while random dimension reduction on prompts only yields performance between the two extremes. Instead, ViaPT overcomes these limitations by balancing dataset-level and instance-level knowledge, while reducing the amount of learnable parameters compared to VPT-Deep. Extensive experiments across 34 diverse datasets demonstrate that our method consistently outperforms state-of-the-art baselines, establishing a new paradigm for analyzing and optimizing visual prompts for vision transformers.
Xi Xiao 0003, Yunbei Zhang, Xingjian Li 0002, Tianyang Wang 0004, Xiao Wang 0004, Yuxiang Wei 0004, Jihun Hamm, Min Xu 0009
ACM Multimedia8
2025 Distilling Aggregated Knowledge for Weakly-Supervised Video Anomaly Detection
abstract
Video anomaly detection aims to develop automated models capable of identifying abnormal events in surveillance videos. The benchmark setup for this task is extremely challenging due to: i) the limited size of the training sets, ii) weak supervision provided in terms of video-level labels, and iii) intrinsic class imbalance induced by the scarcity of abnormal events. In this work, we show that distilling knowledge from aggregated representations of multiple backbones into a single-backbone Student model achieves state-of-the-art performance. In particular, we develop a bi-level distillation approach along with a novel disen-tangled cross-attention-based feature aggregation network. Our proposed approach, DAKD (Distilling Aggregated Knowledge with Disentangled Attention), demonstrates superior performance compared to existing methods across multiple benchmark datasets. Notably, we achieve significant improvements of 1.36%, 0.78%, and 7.02% on the UCF-Crime, ShanghaiTech, and XD-Violence datasets, respectively.
Jash Dalvi, Ali Dabouei, Gunjan Dhanuka, Min Xu 0009
WACV4
2025 J-Invariant Volume Shuffle for Self-Supervised Cryo-Electron Tomogram Denoising on Single Noisy Volume
Xiwei Liu, Mohamad Kassab, Min Xu 0009, Qirong Ho
WACV3
2025 CryoMAE: Few-Shot Cryo-EM Particle Picking with Masked Autoencoders
abstract
Cryo-electron microscopy (cryo-EM) emerges as a pivotal technology for determining the architecture of cells, viruses, and protein assemblies at near-atomic resolution. Traditional particle picking, a key step in cryo-EM, struggles with manual effort and automated methods' sensitivity to low signal-to-noise ratio (SNR) and varied particle orientations. Furthermore, existing neural network (NN)-based approaches often require extensive labeled datasets, limiting their practicality. To overcome these obstacles, we introduce cryoMAE, a novel approach based on few-shot learning that harnesses the capabilities of Masked Autoencoders (MAE) to enable efficient selection of single particles in cryo-EM images. Contrary to conventional NN-based techniques, cryoMAE requires only a minimal set of positive particle images for training yet demonstrates high performance in particle detection. Furthermore, the imple-mentation of a self-cross similarity loss ensures distinct features for particle and background regions, thereby enhancing the discrimination capability of cryoMAE. Experiments on large-scale cryo-EM datasets show that cryoMAE outperforms existing state-of-the-art (SOTA) methods, improving 3D reconstruction resolution by up to 22.4%. Our code is available at: https://github.com/xulabs/aitom.
Chentianye Xu, Xueying Zhan, Min Xu 0009
WACV3
2025 Towards molecular structure discovery from cryo-ET density volumes via modelling auxiliary semantic prototypes
abstract
Cryo-electron tomography (cryo-ET) is confronted with the intricate task of unveiling novel structures. General class discovery (GCD) seeks to identify new classes by learning a model that can pseudo-label unannotated (novel) instances solely using supervision from labeled (base) classes. While 2D GCD for image data has made strides, its 3D counterpart remains unexplored. Traditional methods encounter challenges due to model bias and limited feature transferability when clustering unlabeled 2D images into known and potentially novel categories based on labeled data. To address this limitation and extend GCD to 3D structures, we propose an innovative approach that harnesses a pretrained 2D transformer, enriched by an effective weight inflation strategy tailored for 3D adaptation, followed by a decoupled prototypical network. Incorporating the power of pretrained weight-inflated Transformers, we further integrate CLIP, a vision-language model to incorporate textual information. Our method synergizes a graph convolutional network with CLIP's frozen text encoder, preserving class neighborhood structure. In order to effectively represent unlabeled samples, we devise semantic distance distributions, by formulating a bipartite matching problem for category prototypes using a decoupled prototypical network. Empirical results unequivocally highlight our method's potential in unveiling hitherto unknown structures in cryo-ET. By bridging the gap between 2D GCD and the distinctive challenges of 3D cryo-ET data, our approach paves novel avenues for exploration and discovery in this domain.
Ashwin R. Nair, Xingjian Li 0002, Bhupendra Solanki, Souradeep Mukhopadhyay, Ankit Jha, Mostofa Rafid Uddin, Mainak Singha, Biplab Banerjee, Min Xu 0009
Briefings Bioinform.9
2025 Localization of macromolecules in crowded cellular cryo-electron tomograms from extremely sparse labels
abstract
Localizing macromolecules in crowded cellular cryo-electron tomography (cryo-ET) images or tomograms is crucial for determining their in situ structures. Traditional template matching-based approaches for this task suffer from template-specific biases and have low throughput. Given these problems, learning-based solutions are necessary. However, the paucity of annotated data for training poses substantial challenges for such learning-based methods. Moreover, preparing extensively annotated cellular tomograms for training macromolecule localization methods is extremely time-consuming and burdensome due to the large volume and low signal-to-noise ratio of the tomograms. In this work, we developed TomoPicker, an annotation-efficient macromolecule localization method for tomograms. To achieve such annotation-efficiency, TomoPicker regards macromolecule localization as a voxel classification problem and solves it with two different positive-unlabeled learning approaches. We evaluated TomoPicker on two experimental cryo-ET datasets of crowded eukaryotic cells and one experimental dataset of relatively less crowded prokaryotic cell. We observed that, with only 10 annotated macromolecule locations, TomoPicker with positive unlabeled learning achieved a performance comparable to that of state-of-the-art supervised methods trained with several hundred annotations. In other words, TomoPicker achieved plausible segmentation with up to 98% less data compared with supervised learning-based methods. Furthermore, it demonstrated substantial improvements over existing learning-based macromolecule localization methods under sparse annotation scenarios.
Mostofa Rafid Uddin, Ajmain Yasar Ahmed, H. M. Shadman Tabib, Md Toki Tahmid, Md. Zarif Ul Alam, Zachary Freyberg, Min Xu 0009
Briefings Bioinform.7
2025 Deep video anomaly detection in automated laboratory setting
abstract
Laboratory automation integrates robotics, machine learning, and computer vision to enhance precision and efficiency while reducing costs. Despite its pivotal importance in fully automated experimentation, automatic monitoring of procedures is an overlooked task. In this paper, we aim to address this shortcoming by developing a learning method for detecting anomalies in the laboratory setting. To formalize the problem, we focus on the liquid transfer task as a major task in laboratory automation, given the frequent need for reagent mixing and solution preparation. We introduce a novel video anomaly detection framework that leverages the robust CLIP features in conjunction with a transformer encoder. Through an array of experiments and ablation studies, we demonstrate that our proposed method surpasses current state-of-the-art anomaly detection techniques adapted to laboratory automation by a notable margin of over 11%, achieving an impressive AUC of 98.79% for video-level anomaly detection.
Ali Dabouei, Jishnu Parayil Shibu, Vibhu Dalal, Chengzhi Cao, Andy Macwilliams, Joshua Kangas, Min Xu 0009
Expert Syst. Appl.7
2025 MVGFormer: Multi-view perspective with graph-guided transformer for cryo-ET segmentation
abstract
• We propose MVGFormer, a multi-view fusion framework with a dual-stream encoder guided by a visual graph. • We design two decoder variants: MF for hierarchical feature fusion and P3DA for multi-scale representation. • We introduce a view-masked self-supervised strategy that reconstructs masked views from remaining ones. • MVGFormer achieves superior performance over existing state-of-the-art cryo-ET segmentation methods. Cryo-Electron Tomography (cryo-ET) is a cutting-edge 3D imaging technology that enables detailed examination of biological macromolecular structures at near-atomic resolution. Recent deep learning applications on cryo-ET, such as cryo-ET segmentation, have drawn widespread interest for their potential to improve particle alignment, classification, and other tasks. However, current methods heavily rely on convolutional architectures, which prioritize local information while neglecting the global structural information inherent in cryo-ET data. Transformer-based models, known for their large receptive field, have become the de-facto design for 2D vision tasks due to their ability to effectively capture global information. This approach is also well-suited for 3D tasks, given the complex nature of 3D objects. Based on this, we extend 2D vision transformers into 3D and propose a novel transformer-based framework for cryo-ET segmentation, named MVGFormer. MVGFormer introduces a multi-view perspective fusion transformer encoder, which captures rich global structural information from multiple perspectives using unique positional embeddings. To enhance contextual awareness, we design a parallel context encoder that builds a visual graph to guide attention. We further introduce two complementary 3D decoders: multi-level feature fusion (MF) and parallel atrous convolutions (P3DA), which together capture multi-scale structural cues for precise segmentation. Furthermore, we introduce a view-masked self-supervised learning strategy to reinforce the effectiveness of the multi-view design and improve the model’s representation capability. To our knowledge, MVGFormer is the first transformer-based model for cryo-ET segmentation. We empirically evaluate MVGFormer on six cryo-ET datasets across three different tasks. Extensive experimental results demonstrate its superiority over state-of-the-art 3D segmentation methods.
Haoran Li 0024, Xingjian Li 0002, Jiahua Shi, Huaming Chen, Bo Du 0004, Johan Barthélemy, Daisuke Kihara, Jun Shen 0001, Min Xu 0009
Knowl. Based Syst.11
2025 Pro-NeXt: An All-in-One Unified Model for General Fine-Grained Visual Recognition
abstract
Unlike general visual classification (CLS) tasks, certain CLS problems are significantly more challenging as they involve recognizing professionally categorized or highly specialized images. Fine-Grained Visual Classification (FGVC) has emerged as a broad solution to address this complexity. However, most existing methods have been predominantly evaluated on a limited set of homogeneous benchmarks, such as bird species or vehicle brands. Moreover, these approaches often train separate models for each specific task, which restricts their generalizability. This paper proposes a scalable and explainable foundational model designed to tackle a wide range of FGVC tasks from a unified and generalizable perspective. We introduce a novel architecture named Pro-NeXt and reveal that Pro-NeXt exhibits substantial generalizability across diverse professional fields such as fashion, medicine, and art areas, previously considered disparate. Our basic-sized Pro-NeXt-B surpasses all preceding task-specific models across 12 distinct datasets within 5 diverse domains. Furthermore, we find its good scaling property that scaling up Pro-NeXt in depth and width with increasing GFlops can consistently enhance its accuracy. Beyond scalability and adaptability, the intermediate features of Pro-NeXt achieve reliable object detection and segmentation performance without extra training, highlighting its solid explainability. We will release the code to promote further research in this area.
Jiayuan Zhu, Min Xu 0009, Yueming Jin
IEEE Trans. Pattern Anal. Mach. Intell.3
2024 Deep Active Learning with Noise Stability
abstract
Uncertainty estimation for unlabeled data is crucial to active learning. With a deep neural network employed as the backbone model, the data selection process is highly challenging due to the potential over-confidence of the model inference. Existing methods resort to special learning fashions (e.g. adversarial) or auxiliary models to address this challenge. This tends to result in complex and inefficient pipelines, which would render the methods impractical. In this work, we propose a novel algorithm that leverages noise stability to estimate data uncertainty. The key idea is to measure the output derivation from the original observation when the model parameters are randomly perturbed by noise. We provide theoretical analyses by leveraging the small Gaussian noise theory and demonstrate that our method favors a subset with large and diverse gradients. Our method is generally applicable in various tasks, including computer vision, natural language processing, and structural data analysis. It achieves competitive performance compared against state-of-the-art active learning baselines.
Xingjian Li 0002, Pengkun Yang, Yangcheng Gu, Xueying Zhan, Tianyang Wang 0004, Min Xu 0009, Cheng-Zhong Xu 0001
AAAI6
2024 MedSegDiff-V2: Diffusion-Based Medical Image Segmentation with Transformer
abstract
The Diffusion Probabilistic Model (DPM) has recently gained popularity in the field of computer vision, thanks to its image generation applications, such as Imagen, Latent Diffusion Models, and Stable Diffusion, which have demonstrated impressive capabilities and sparked much discussion within the community. Recent investigations have further unveiled the utility of DPM in the domain of medical image analysis, as underscored by the commendable performance exhibited by the medical image segmentation model across various tasks. Although these models were originally underpinned by a UNet architecture, there exists a potential avenue for enhancing their performance through the integration of vision transformer mechanisms. However, we discovered that simply combining these two models resulted in subpar performance. To effectively integrate these two cutting-edge techniques for the Medical image segmentation, we propose a novel Transformer-based Diffusion framework, called MedSegDiff-V2. We verify its effectiveness on 20 medical image segmentation tasks with different image modalities. Through comprehensive evaluation, our approach demonstrates superiority over prior state-of-the-art (SOTA) methodologies. Code is released at https://github.com/KidsWithTokens/MedSegDiff.
Wei Ji 0011, Huazhu Fu, Min Xu 0009, Yueming Jin, Yanwu Xu 0001
AAAI4
2024 Synergistic Global-Space Camera and Human Reconstruction from Videos
abstract
Remarkable strides have been made in reconstructing static scenes or human bodies from monocular videos. Yet, the two problems have largely been approached independently, without much synergy. Most visual SLAM methods can only reconstruct camera trajectories and scene structures up to scale, while most HMR methods reconstruct human meshes in metric scale but fall short in reasoning with cameras and scenes. This work introduces Synergistic Camera and Human Reconstruction (SynCHMR) to marry the best of both worlds. Specifically, we design Human-aware Metric SLAM to reconstruct metric-scale camera poses and scene point clouds using camera-frame HMR as a strong prior, addressing depth, scale, and dynamic ambiguities. Conditioning on the dense scene recovered, we further learn a Scene-aware SMPL Denoiser to enhance world-frame HMR by incorporating spatio-temporal coherency and dynamic scene constraints. Together, they lead to consistent reconstructions of camera trajectories, human meshes, and dense scene point clouds in a common world frame.
Tuanfeng Y. Wang, Bhiksha Raj, Min Xu 0009, Jimei Yang, Chun-Hao Paul Huang
CVPR4
2024 HGTDP-DTA: Hybrid Graph-Transformer with Dynamic Prompt for Drug-Target Binding Affinity Prediction
Xi Xiao 0003, Lijing Zhu, Gaofei Chen, Zhengji Li, Tianyang Wang 0004, Min Xu 0009
ICONIP (4)8
2024 CryoSAM: Training-Free CryoET Tomogram Segmentation with Foundation Models
Hengwei Bian, Michael Mu, Mostofa Rafid Uddin, Tianyang Wang 0004, Min Xu 0009
MICCAI (8)8
2024 Multi-dimensional Fusion and Consistency for Semi-supervised Medical Image Segmentation
Yixing Lu, Zhaoxin Fan, Min Xu 0009
MMM (2)3
2024 Metric from Human: Zero-shot Monocular Metric Depth Estimation via Test-time Adaptation
abstract
Monocular depth estimation (MDE) is fundamental for deriving 3D scene structures from 2D images. While state-of-the-art monocular relative depth estimation (MRDE) excels in estimating relative depths for in-the-wild images, current monocular metric depth estimation (MMDE) approaches still face challenges in handling unseen scenes. Since MMDE can be viewed as the composition of MRDE and metric scale recovery, we attribute this difficulty to scene dependency, where MMDE models rely on scenes observed during supervised training for predicting scene scales during inference. To address this issue, we propose to use humans as landmarks for distilling scene-independent metric scale priors from generative painting models. Our approach, Metric from Human (MfH), bridges from generalizable MRDE to zero-shot MMDE in a generate-and-estimate manner. Specifically, MfH generates humans on the input image with generative painting and estimates human dimensions with an off-the-shelf human mesh recovery (HMR) model. Based on MRDE predictions, it propagates the metric information from painted humans to the contexts, resulting in metric depth estimations for the original input. Through this annotation-free test-time adaptation, MfH achieves superior zero-shot performance in MMDE, demonstrating its strong generalization ability.
Hengwei Bian, Kaihua Chen, Pengliang Ji, Liao Qu, Shao-yu Lin, Weichen Yu, Haoran Li 0024, Hao Chen 0102, Jun Shen 0001, Bhiksha Raj, Min Xu 0009
NeurIPS12
2024 Knowledge transfer from macro-world to micro-world: enhancing 3D Cryo-ET classification through fine-tuning video-based deep models
abstract
MOTIVATION: Deep learning models have achieved remarkable success in a wide range of natural-world tasks, such as vision, language, and speech recognition. These accomplishments are largely attributed to the availability of open-source large-scale datasets. More importantly, pre-trained foundational modellearnings exhibit a surprising degree of transferability to downstream tasks, enabling efficient learning even with limited training examples. However, the application of such natural-domain models to the domain of tiny Cryo-Electron Tomography (Cryo-ET) images has been a relatively unexplored frontier. This research is motivated by the intuition that 3D Cryo-ET voxel data can be conceptually viewed as a sequence of progressively evolving video frames. RESULTS: Leveraging the above insight, we propose a novel approach that involves the utilization of 3D models pre-trained on large-scale video datasets to enhance Cryo-ET subtomogram classification. Our experiments, conducted on both simulated and real Cryo-ET datasets, reveal compelling results. The use of video initialization not only demonstrates improvements in classification accuracy but also substantially reduces training costs. Further analyses provide additional evidence of the value of video initialization in enhancing subtomogram feature extraction. Additionally, we observe that video initialization yields similar positive effects when applied to medical 3D classification tasks, underscoring the potential of cross-domain knowledge transfer from video-based models to advance the state-of-the-art in a wide range of biological and medical data types. AVAILABILITY AND IMPLEMENTATION: https://github.com/xulabs/aitom.
Sabhay Jain, Xingjian Li 0002, Min Xu 0009
Bioinform.3
2024 CoNIC Challenge: Pushing the frontiers of nuclear detection, segmentation, classification and counting
abstract
Nuclear detection, segmentation and morphometric profiling are essential in helping us further understand the relationship between histology and patient outcome. To drive innovation in this area, we setup a community-wide challenge using the largest available dataset of its kind to assess nuclear segmentation and cellular composition. Our challenge, named CoNIC, stimulated the development of reproducible algorithms for cellular recognition with real-time result inspection on public leaderboards. We conducted an extensive post-challenge analysis based on the top-performing models using 1,658 whole-slide images of colon tissue. With around 700 million detected nuclei per model, associated features were used for dysplasia grading and survival analysis, where we demonstrated that the challenge's improvement over the previous state-of-the-art led to significant boosts in downstream performance. Our findings also suggest that eosinophils and neutrophils play an important role in the tumour microevironment. We release challenge models and WSI-level results to foster the development of further methods for biomarker discovery.
Simon Graham, Quoc Dang Vu, Mostafa Jahanifar, Martin Weigert 0001, Jun Zhang 0018, Sen Yang 0006, Jinxi Xiang, Josef Lorenz Rumberger, Elias Baumann, Peter Hirsch 0001, Chenyang Hong, Angelica I. Avilés-Rivero, Ayushi Jain, Heeyoung Ahn, Yiyu Hong, Hussam Azzuni, Min Xu 0009, Mohammad Yaqub, Marie-Claire Blache, Benoît Piégu, Bertrand Vernay, Tim Scherr, Moritz Böhland, Katharina Löffler, Weiqin Ying, Chixin Wang, David R. J. Snead, Shan E Ahmed Raza, Fayyaz ul Amir Afsar Minhas, Nasir M. Rajpoot
Medical Image Anal.21
2023 iHerd: an integrative hierarchical graph representation learning framework to quantify network changes and prioritize risk genes in disease
abstract
Different genes form complex networks within cells to carry out critical cellular functions, while network alterations in this process can potentially introduce downstream transcriptome perturbations and phenotypic variations. Therefore, developing efficient and interpretable methods to quantify network changes and pinpoint driver genes across conditions is crucial. We propose a hierarchical graph representation learning method, called iHerd. Given a set of networks, iHerd first hierarchically generates a series of coarsened sub-graphs in a data-driven manner, representing network modules at different resolutions (e.g., the level of signaling pathways). Then, it sequentially learns low-dimensional node representations at all hierarchical levels via efficient graph embedding. Lastly, iHerd projects separate gene embeddings onto the same latent space in its graph alignment module to calculate a rewiring index for driver gene prioritization. To demonstrate its effectiveness, we applied iHerd on a tumor-to-normal GRN rewiring analysis and cell-type-specific GCN analysis using single-cell multiome data of the brain. We showed that iHerd can effectively pinpoint novel and well-known risk genes in different diseases. Distinct from existing models, iHerd's graph coarsening for hierarchical learning allows us to successfully classify network driver genes into early and late divergent genes (EDGs and LDGs), emphasizing genes with extensive network changes across and within signaling pathway levels. This unique approach for driver gene classification can provide us with deeper molecular insights. The code is freely available at https://github.com/aicb-ZhangLabs/iHerd. All other relevant data are within the manuscript and supporting information files.
Ziheng Duan, Ahyeon Hwang, Cheyu Lee, Kaichi Xie, Chutong Xiao, Min Xu 0009, Matthew J. Girgenti, Jing Zhang 0062
PLoS Comput. Biol.7
2022 Boosting Active Learning via Improving Test Performance
abstract
Central to active learning (AL) is what data should be selected for annotation. Existing works attempt to select highly uncertain or informative data for annotation. Nevertheless, it remains unclear how selected data impacts the test performance of the task model used in AL. In this work, we explore such an impact by theoretically proving that selecting unlabeled data of higher gradient norm leads to a lower upper-bound of test loss, resulting in a better test performance. However, due to the lack of label information, directly computing gradient norm for unlabeled data is infeasible. To address this challenge, we propose two schemes, namely expected-gradnorm and entropy-gradnorm. The former computes the gradient norm by constructing an expected empirical loss while the latter constructs an unsupervised loss with entropy. Furthermore, we integrate the two schemes in a universal AL framework. We evaluate our method on classical image classification and semantic segmentation tasks. To demonstrate its competency in domain applications and its robustness to noise, we also validate our method on a cellular imaging analysis task, namely cryo-Electron Tomography subtomogram classification. Results demonstrate that our method achieves superior performance against the state of the art. We refer readers to https://arxiv.org/pdf/2112.05683.pdf for the full version of this paper which includes the appendix and source code link.
Tianyang Wang 0004, Xingjian Li 0002, Pengkun Yang, Guosheng Hu, Siyu Huang, Cheng-Zhong Xu 0001, Min Xu 0009
AAAI8
2022 TransResNet: Integrating the Strengths of ViTs and CNNs for High Resolution Medical Image Segmentation via Feature Grafting
Muhammad Hamza Sharif, Dmitry Demidov, Asif Hanif, Mohammad Yaqub, Min Xu 0009
BMVC5
2022 Harmony: A Generic Unsupervised Approach for Disentangling Semantic Content from Parameterized Transformations
abstract
In many real-life image analysis applications, particularly in biomedical research domains, the objects of interest undergo multiple transformations that alters their visual properties while keeping the semantic content unchanged. Disentangling images into semantic content factors and transformations can provide significant benefits into many domain-specific image analysis tasks. To this end, we propose a generic unsupervised framework, Harmony, that simultaneously and explicitly disentangles semantic content from multiple parameterized transformations. Harmony leverages a simple cross-contrastive learning framework with multiple explicitly parameterized latent representations to disentangle content from transformations. To demonstrate the efficacy of Harmony, we apply it to disentangle image semantic content from several parameterized transformations (rotation, translation, scaling, and contrast). Harmony achieves significantly improved disentanglement over the baseline models on several image datasets of diverse domains. With such disentanglement, Harmony is demonstrated to incentivize bioimage analysis research by modeling structural heterogeneity of macromolecules from cryo-ET images and learning transformation-invariant representations of protein particles from single-particle cryo-EM images. Harmony also performs very well in disentangling content from 3D transformations and can perform coarse and fast alignment of 3D cryo-ET subtomograms. Therefore, Harmony is generalizable to many other imaging domains and can potentially be extended to domains beyond imaging as well.
Mostofa Rafid Uddin, Gregory Howe, Min Xu 0009
CVPR4
2022 Deep Active Learning for Cryo-Electron Tomography Classification
abstract
Cryo-Electron Tomography (cryo-ET) is an emerging 3D imaging technique which shows great potentials in structural biology research. One of the main challenges is to perform classification of macromolecules captured by cryo-ET. Re-cent efforts exploit deep learning to address this challenge. However, training reliable deep models usually requires a huge amount of labeled data in supervised fashion. Annotating cryo-ET data is arguably very expensive. Deep Active Learning (DAL) can be used to reduce labeling cost while not sacrificing the task performance too much. Nevertheless, most existing methods resort to auxiliary models or complex fashions (e.g. adversarial learning) for uncertainty estimation, the core of DAL. These models need to be highly customized for cryo-ET tasks which require 3D networks, and extra efforts are also indispensable for tuning these models, rendering a difficulty of deployment on cryo-ET tasks. To address these challenges, we propose a novel metric for data selection in DAL, which can also be leveraged as a regularizer of the empirical loss, further boosting the task model. We demonstrate the superiority of our method via extensive experiments on both simulated and real cryo-ET datasets. Our source Code and Appendix can be found at this URL.
Tianyang Wang 0004, Bo Li 0013, Jing Zhang 0062, Mostofa Rafid Uddin, Min Xu 0009
ICIP7
2022 Unsupervised Multi-Task Learning for 3D Subtomogram Image Alignment, Clustering and Segmentation
abstract
3D subtomogram image alignment, clustering, and segmentation are vital to macromolecular structure recognition in cryo-electron tomography (cryo-ET). However, acquiring ground-truth labels to train a unified deep learning model that can simultaneously deal with these tasks is unaffordable. To this end, we propose an end-to-end unified multi-task learning framework to simultaneously complete the three tasks, where models are trained in an unsupervised manner without using any labels. In particular, we have three parallel branches. In the alignment branch, we adopt a two-stage training scheme, i.e., self-supervised pretraining and constrained unsupervised training using our proposed skip correlation attention layer and constrained loss. Synchronously, in the clustering branch, the learned deep cluster features are utilized to iteratively cluster subtomograms into groups using pseudo-labels from an image-wise Gaussian Mixture Model (GMM). Meanwhile, in the segmentation branch, we use rough pseudo-labels generated from a voxel-wise GMM as supervision signals, and prior knowledge from humans is utilized to jointly learn how to correct these labels as well as predict reliable segmentation results. Benefiting from the end-to-end unified network architecture, our method achieves overall state-of-the-art performance on both simulated and real subtomogram processing benchmarks.
Haoyi Zhu, Chuting Wang, Yuanxin Wang 0001, Zhaoxin Fan, Mostofa Rafid Uddin, Xin Gao 0001, Jing Zhang 0062, Min Xu 0009
ICIP9
2022 Cryo-shift: reducing domain shift in cryo-electron subtomograms with unsupervised domain adaptation and randomization
abstract
MOTIVATION: Cryo-Electron Tomography (cryo-ET) is a 3D imaging technology that enables the visualization of subcellular structures in situ at near-atomic resolution. Cellular cryo-ET images help in resolving the structures of macromolecules and determining their spatial relationship in a single cell, which has broad significance in cell and structural biology. Subtomogram classification and recognition constitute a primary step in the systematic recovery of these macromolecular structures. Supervised deep learning methods have been proven to be highly accurate and efficient for subtomogram classification, but suffer from limited applicability due to scarcity of annotated data. While generating simulated data for training supervised models is a potential solution, a sizeable difference in the image intensity distribution in generated data as compared with real experimental data will cause the trained models to perform poorly in predicting classes on real subtomograms. RESULTS: In this work, we present Cryo-Shift, a fully unsupervised domain adaptation and randomization framework for deep learning-based cross-domain subtomogram classification. We use unsupervised multi-adversarial domain adaption to reduce the domain shift between features of simulated and experimental data. We develop a network-driven domain randomization procedure with 'warp' modules to alter the simulated data and help the classifier generalize better on experimental data. We do not use any labeled experimental data to train our model, whereas some of the existing alternative approaches require labeled experimental samples for cross-domain classification. Nevertheless, Cryo-Shift outperforms the existing alternative approaches in cross-domain subtomogram classification in extensive evaluation studies demonstrated herein using both simulated and experimental data. AVAILABILITYAND IMPLEMENTATION: https://github.com/xulabs/aitom. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Hmrishav Bandyopadhyay, Leiting Ding, Sinuo Liu, Mostofa Rafid Uddin, Sima Behpour, Min Xu 0009
Bioinform.8
2022 Venus: An efficient virus infection detection and fusion site discovery method using single-cell and bulk RNA-seq data
abstract
Early and accurate detection of viruses in clinical and environmental samples is essential for effective public healthcare, treatment, and therapeutics. While PCR detects potential pathogens with high sensitivity, it is difficult to scale and requires knowledge of the exact sequence of the pathogen. With the advent of next-gen single-cell sequencing, it is now possible to scrutinize viral transcriptomics at the finest possible resolution-cells. This newfound ability to investigate individual cells opens new avenues to understand viral pathophysiology with unprecedented resolution. To leverage this ability, we propose an efficient and accurate computational pipeline, named Venus, for virus detection and integration site discovery in both single-cell and bulk-tissue RNA-seq data. Specifically, Venus addresses two main questions: whether a tissue/cell type is infected by viruses or a virus of interest? And if infected, whether and where has the virus inserted itself into the human genome? Our analysis can be broken into two parts-validation and discovery. Firstly, for validation, we applied Venus on well-studied viral datasets, such as HBV- hepatocellular carcinoma and HIV-infection treated with antiretroviral therapy. Secondly, for discovery, we analyzed datasets such as HIV-infected neurological patients and deeply sequenced T-cells. We detected viral transcripts in the novel target of the brain and high-confidence integration sites in immune cells. In conclusion, here we describe Venus, a publicly available software which we believe will be a valuable virus investigation tool for the scientific community at large.
Cheyu Lee, Ziheng Duan, Min Xu 0009, Matthew J. Girgenti, Ke Xu 0010, Mark Gerstein, Jing Zhang 0062
PLoS Comput. Biol.4
2022 Macromolecules Structural Classification With a 3D Dilated Dense Network in Cryo-Electron Tomography
abstract
Cryo-electron tomography, combined with subtomogram averaging (STA), can reveal three-dimensional (3D) macromolecule structures in the near-native state from cells and other biological samples. In STA, to get a high-resolution 3D view of macromolecule structures, diverse macromolecules captured by the cellular tomograms need to be accurately classified. However, due to the poor signal-to-noise-ratio (SNR) and severe ray artifacts in the tomogram, it remains a major challenge to classify macromolecules with high accuracy. In this paper, we propose a new convolutional neural network, named 3D-Dilated-DenseNet, to improve the performance of macromolecule classification. In 3D-Dilated-DenseNet, there are two key strategies to guarantee macromolecule classification accuracy: 1) Using dense connections to enhance feature map utilization (corresponding to the baseline 3D-C-DenseNet); 2) Adopting dilated convolution to enrich multi-level information in feature maps. We tested 3D-Dilated-DenseNet and 3D-C-DenseNet both on synthetic data and experimental data. The results show that, on synthetic data, compared with the state-of-the-art method in the SHREC contest (SHREC-CNN), both 3D-C-DenseNet and 3D-Dilated-DenseNet outperform SHREC-CNN. In particular, 3D-Dilated-DenseNet improves 0.393 of F1 metric on tiny-size macromolecules and 0.213 on small-size macromolecules. On experimental data, compared with 3D-C-DenseNet, 3D-Dilated-DenseNet can increase classification performance by 2.1 percent.
Renmin Han, Zhiyong Liu 0002, Min Xu 0009, Fa Zhang 0001
IEEE ACM Trans. Comput. Biol. Bioinform.5
2021 End-to-end robust joint unsupervised image alignment and clustering
abstract
Computing dense pixel-to-pixel image correspondences is a fundamental task of computer vision. Often, the objective is to align image pairs from the same semantic category for manipulation or segmentation purposes. Despite achieving superior performance, existing deep learning alignment methods cannot cluster images; consequently, clustering and pairing images needed to be a separate laborious and expensive step. Given a dataset with diverse semantic categories, we propose a multi-task model, Jim-Net, that can directly learn to cluster and align images without any pixel-level or image-level annotations. We design a pair-matching alignment unsupervised training algorithm that selectively matches and aligns image pairs from the clustering branch. Our unsupervised Jim-Net achieves comparable accuracy with state-of-the-art supervised methods on benchmark 2D image alignment dataset PF-PASCAL. Specifically, we apply Jim-Net to cryo-electron tomography, a revolutionary 3D microscopy imaging technique of native subcellular structures. After extensive evaluation on seven datasets, we demonstrate that Jim-Net enables systematic discovery and recovery of representative macromolecular structures in situ, which is essential for revealing molecular mechanisms underlying cellular functions. To our knowledge, Jim-Net is the first end-to-end model that can simultaneously align and cluster images, which significantly improves the performance as compared to performing each task alone.
Gregory Howe, Min Xu 0009
ICCV3
2021 Weakly Supervised 3D Semantic Segmentation Using Cross-Image Consensus and Inter-Voxel Affinity Relations
abstract
We propose a novel weakly supervised approach for 3D semantic segmentation on volumetric images. Unlike most existing methods that require voxel-wise densely labeled training data, our weakly-supervised CIVA-Net is the first model that only needs image-level class labels as guidance to learn accurate volumetric segmentation. Our model learns from cross-image co-occurrence for integral region generation, and explores inter-voxel affinity relations to predict segmentation with accurate boundaries. We empirically validate our model on both simulated and real cryo-ET datasets. Our experiments show that CIVA-Net achieves comparable performance to the state-of-the-art models trained with stronger supervision.
Jeffrey Chen, Junwei Liang 0001, Chengqi Li, Sinuo Liu, Sima Behpour, Min Xu 0009
ICCV8
2021 Unsupervised Domain Alignment Based Open Set Structural Recognition of Macromolecules Captured By Cryo-Electron Tomography
abstract
Cellular cryo-Electron Tomography (cryo-ET) provides three-dimensional views of structural and spatial information of various macromolecules in cells in a near-native state. Subtomogram classification is a key step for recognizing and differentiating these macromolecular structures. In recent years, deep learning methods have been developed for high-throughput subtomogram classification tasks; however, conventional supervised deep learning methods cannot recognize macromolecular structural classes that do not exist in the training data. This imposes a major weakness since most native macromolecular structures in cells are unknown and consequently, cannot be included in the training data. Therefore, open set learning which can recognize unknown macromolecular structures is necessary for boosting the power of automatic subtomogram classification. In this paper, we propose a method called Margin-based Loss for Unsupervised Domain Alignment (MLUDA) for open set recognition problems where only a few categories of interest are shared between cross-domain data. Through extensive experiments, we demonstrate that MLUDA performs well at cross-domain open-set classification on both public datasets and medical imaging datasets. So our method is of practical importance.
Gregory Howe, Kai Yi, Jing Zhang 0062, Yi-Wei Chang, Min Xu 0009
ICIP7
2021 SCAN-ATAC-Sim: a scalable and efficient method for simulating single-cell ATAC-seq data from bulk-tissue experiments
abstract
SUMMARY: scATAC-seq is a powerful approach for characterizing cell-type-specific regulatory landscapes. However, it is difficult to benchmark the performance of various scATAC-seq analysis techniques (such as clustering and deconvolution) without having a priori a known set of gold-standard cell types. To simulate scATAC-seq experiments with known cell-type labels, we introduce an efficient and scalable scATAC-seq simulation method (SCAN-ATAC-Sim) that down-samples bulk ATAC-seq data (e.g. from representative cell lines or tissues). Our protocol uses a consistent but tunable signal-to-noise ratio across cell types in a scATAC-seq simulation for integrating bulk experiments with different levels of background noise, and it independently samples twice without replacement to account for the diploid genome. Because it uses an efficient weighted reservoir sampling algorithm and is highly parallelizable with OpenMP, our implementation in C++ allows millions of cells to be simulated in less than an hour on a laptop computer. AVAILABILITY AND IMPLEMENTATION: SCAN-ATAC-Sim is available at scan-atac-sim.gersteinlab.org. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Zhanlin Chen, Jing Zhang 0062, Jason Liu 0003, Jiangqi Zhu, Donghoon Lee 0006, Min Xu 0009, Mark Gerstein
Bioinform.7
2021 DECODE: a Deep-learning framework for Condensing enhancers and refining boundaries with large-scale functional assays
abstract
MOTIVATION: Mapping distal regulatory elements, such as enhancers, is a cornerstone for elucidating how genetic variations may influence diseases. Previous enhancer-prediction methods have used either unsupervised approaches or supervised methods with limited training data. Moreover, past approaches have implemented enhancer discovery as a binary classification problem without accurate boundary detection, producing low-resolution annotations with superfluous regions and reducing the statistical power for downstream analyses (e.g. causal variant mapping and functional validations). Here, we addressed these challenges via a two-step model called Deep-learning framework for Condensing enhancers and refining boundaries with large-scale functional assays (DECODE). First, we employed direct enhancer-activity readouts from novel functional characterization assays, such as STARR-seq, to train a deep neural network for accurate cell-type-specific enhancer prediction. Second, to improve the annotation resolution, we implemented a weakly supervised object detection framework for enhancer localization with precise boundary detection (to a 10 bp resolution) using Gradient-weighted Class Activation Mapping. RESULTS: Our DECODE binary classifier outperformed a state-of-the-art enhancer prediction method by 24% in transgenic mouse validation. Furthermore, the object detection framework can condense enhancer annotations to only 13% of their original size, and these compact annotations have significantly higher conservation scores and genome-wide association study variant enrichments than the original predictions. Overall, DECODE is an effective tool for enhancer classification and precise localization. AVAILABILITY AND IMPLEMENTATION: DECODE source code and pre-processing scripts are available at decode.gersteinlab.org. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Zhanlin Chen, Jing Zhang 0062, Jason Liu 0003, Donghoon Lee 0006, Martin Renqiang Min, Min Xu 0009, Mark Gerstein
Bioinform.7
2021 Active learning to classify macromolecular structures in situ for less supervision in cryo-electron tomography
abstract
MOTIVATION: Cryo-Electron Tomography (cryo-ET) is a 3D bioimaging tool that visualizes the structural and spatial organization of macromolecules at a near-native state in single cells, which has broad applications in life science. However, the systematic structural recognition and recovery of macromolecules captured by cryo-ET are difficult due to high structural complexity and imaging limits. Deep learning-based subtomogram classification has played critical roles for such tasks. As supervised approaches, however, their performance relies on sufficient and laborious annotation on a large training dataset. RESULTS: To alleviate this major labeling burden, we proposed a Hybrid Active Learning (HAL) framework for querying subtomograms for labeling from a large unlabeled subtomogram pool. Firstly, HAL adopts uncertainty sampling to select the subtomograms that have the most uncertain predictions. This strategy enforces the model to be aware of the inductive bias during classification and subtomogram selection, which satisfies the discriminativeness principle in AL literature. Moreover, to mitigate the sampling bias caused by such strategy, a discriminator is introduced to judge if a certain subtomogram is labeled or unlabeled and subsequently the model queries the subtomogram that have higher probabilities to be unlabeled. Such query strategy encourages to match the data distribution between the labeled and unlabeled subtomogram samples, which essentially encodes the representativeness criterion into the subtomogram selection process. Additionally, HAL introduces a subset sampling strategy to improve the diversity of the query set, so that the information overlap is decreased between the queried batches and the algorithmic efficiency is improved. Our experiments on subtomogram classification tasks using both simulated and real data demonstrate that we can achieve comparable testing performance (on average only 3% accuracy drop) by using less than 30% of the labeled subtomograms, which shows a very promising result for subtomogram classification task with limited labeling resources. AVAILABILITY AND IMPLEMENTATION: https://github.com/xulabs/aitom. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Xuefeng Du, Haohan Wang, Zhenxi Zhu, Yi-Wei Chang, Jing Zhang 0062, Eric P. Xing, Min Xu 0009
Bioinform.8
2021 Few shot domain adaptation for in situ macromolecule structural classification in cryoelectron tomograms
abstract
MOTIVATION: Cryoelectron tomography (cryo-ET) visualizes structure and spatial organization of macromolecules and their interactions with other subcellular components inside single cells in the close-to-native state at submolecular resolution. Such information is critical for the accurate understanding of cellular processes. However, subtomogram classification remains one of the major challenges for the systematic recognition and recovery of the macromolecule structures in cryo-ET because of imaging limits and data quantity. Recently, deep learning has significantly improved the throughput and accuracy of large-scale subtomogram classification. However, often it is difficult to get enough high-quality annotated subtomogram data for supervised training due to the enormous expense of labeling. To tackle this problem, it is beneficial to utilize another already annotated dataset to assist the training process. However, due to the discrepancy of image intensity distribution between source domain and target domain, the model trained on subtomograms in source domain may perform poorly in predicting subtomogram classes in the target domain. RESULTS: In this article, we adapt a few shot domain adaptation method for deep learning-based cross-domain subtomogram classification. The essential idea of our method consists of two parts: (i) take full advantage of the distribution of plentiful unlabeled target domain data, and (ii) exploit the correlation between the whole source domain dataset and few labeled target domain data. Experiments conducted on simulated and real datasets show that our method achieves significant improvement on cross domain subtomogram classification compared with baseline methods. AVAILABILITY AND IMPLEMENTATION: Software is available online https://github.com/xulabs/aitom. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Liangyong Yu, Ge Yang 0002, Rui Jiang 0001, Min Xu 0009
Bioinform.8
2020 Efficient Cryo-Electron Tomogram Simulation of Macromolecular Crowding with Application to SARS-CoV-2
abstract
We propose an efficient method for simulating a cryo-Electron Tomography (cryo-ET) image of a target macromolecule with several neighbor macromolecules packed to achieve a realistic crowded cytoplasm content. The simulated results are subtomograms with corresponding noise-free 3D density maps and pre-specified labels (PDB ID, center locations, and orientations) to assist bioimage analysis. They can serve as benchmark datasets for testing developing cryo-ET analysis algorithms and as training datasets with readily available ground truth labels for learning neural network models. The COVID-19 pandemic has sparked a global health crisis that severely impacting lives worldwide. As an important application, we simulated the scene of SARS-CoV-2 interacting with the host cell. The simulated cryo-ET images clearly showed the binding domain of the virus and the host cell to facilitate the research of SARS-CoV-2' infection. We also trained two different classification models to demonstrate that our simulated cryo-ET data is able to assist the cryo-ET analysis task and to validate the performance between different methods.
Sinuo Liu, Mohan Vamsi Nallapareddy, Ajinkya Chaudhari, Min Xu 0009
BIBM7
2020 Gum-Net: Unsupervised Geometric Matching for Fast and Accurate 3D Subtomogram Image Alignment and Averaging
abstract
We propose a Geometric unsupervised matching Network (Gum-Net) for finding the geometric correspondence between two images with application to 3D subtomogram alignment and averaging. Subtomogram alignment is the most important task in cryo-electron tomography (cryo-ET), a revolutionary 3D imaging technique for visualizing the molecular organization of unperturbed cellular landscapes in single cells. However, subtomogram alignment and averaging are very challenging due to severe imaging limits such as noise and missing wedge effects. We introduce an end-to-end trainable architecture with three novel modules specifically designed for preserving feature spatial information and propagating feature matching information. The training is performed in a fully unsupervised fashion to optimize a matching metric. No ground truth transformation information nor category-level or instance-level matching supervision information is needed. After systematic assessments on six real and nine simulated datasets, we demonstrate that Gum-Net reduced the alignment error by 40 to 50% and improved the averaging resolution by 10%. Gum-Net also achieved 70 to 110 times speedup in practice with GPU acceleration compared to state-of-the-art subtomogram alignment methods. Our work is the first 3D unsupervised geometric matching method for images of strong transformation variation and high noise level. The training code, trained model, and datasets are available in our open-source software AITom.
Min Xu 0009
CVPR2
2020 Dilated-DenseNet for Macromolecule Classification in Cryo-electron Tomography
Renmin Han, Xuefeng Cui, Zhiyong Liu 0002, Min Xu 0009, Fa Zhang 0001
ISBRA6
2020 A unified framework for packing deformable and non-deformable subcellular structures in crowded cryo-electron tomogram simulation
abstract
BACKGROUND: Cryo-electron tomography is an important and powerful technique to explore the structure, abundance, and location of ultrastructure in a near-native state. It contains detailed information of all macromolecular complexes in a sample cell. However, due to the compact and crowded status, the missing edge effect, and low signal to noise ratio (SNR), it is extremely challenging to recover such information with existing image processing methods. Cryo-electron tomogram simulation is an effective solution to test and optimize the performance of the above image processing methods. The simulated images could be regarded as the labeled data which covers a wide range of macromolecular complexes and ultrastructure. To approximate the crowded cellular environment, it is very important to pack these heterogeneous structures as tightly as possible. Besides, simulating non-deformable and deformable components under a unified framework also need to be achieved. RESULT: In this paper, we proposed a unified framework for simulating crowded cryo-electron tomogram images including non-deformable macromolecular complexes and deformable ultrastructures. A macromolecule was approximated using multiple balls with fixed relative positions to reduce the vacuum volume. A ultrastructure, such as membrane and filament, was approximated using multiple balls with flexible relative positions so that this structure could deform under force field. In the experiment, 400 macromolecules of 20 representative types were packed into simulated cytoplasm by our framework, and numerical verification proved that our method has a smaller volume and higher compression ratio than the baseline single-ball model. We also packed filaments, membranes and macromolecules together, to obtain a simulated cryo-electron tomogram image with deformable structures. The simulated results are closer to the real Cryo-ET, making the analysis more difficult. The DOG particle picking method and the image segmentation method are tested on our simulation data, and the experimental results show that these methods still have much room for improvement. CONCLUSION: The proposed multi-ball model can achieve more crowded packaging results and contains richer elements with different properties to obtain more realistic cryo-electron tomogram simulation. This enables users to simulate cryo-electron tomogram images with non-deformable macromolecular complexes and deformable ultrastructures under a unified framework. To illustrate the advantages of our framework in improving the compression ratio, we calculated the volume of simulated macromolecular under our multi-ball method and traditional single-ball method. We also performed the packing experiment of filaments and membranes to demonstrate the simulation ability of deformable structures. Our method can be used to do a benchmark by generating large labeled cryo-ET dataset and evaluating existing image processing methods. Since the content of the simulated cryo-ET is more complex and crowded compared with previous ones, it will pose a greater challenge to existing image processing methods.
Sinuo Liu, Fengnian Zhao, Hongpan Zhang, Thomas Hall, Xin Gao 0001, Min Xu 0009
BMC Bioinform.11
2020 Spark-based parallel calculation of 3D fourier shell correlation for macromolecule structure local resolution estimation
abstract
BACKGROUND: Resolution estimation is the main evaluation criteria for the reconstruction of macromolecular 3D structure in the field of cryoelectron microscopy (cryo-EM). At present, there are many methods to evaluate the 3D resolution for reconstructed macromolecular structures from Single Particle Analysis (SPA) in cryo-EM and subtomogram averaging (SA) in electron cryotomography (cryo-ET). As global methods, they measure the resolution of the structure as a whole, but they are inaccurate in detecting subtle local changes of reconstruction. In order to detect the subtle changes of reconstruction of SPA and SA, a few local resolution methods are proposed. The mainstream local resolution evaluation methods are based on local Fourier shell correlation (FSC), which is computationally intensive. However, the existing resolution evaluation methods are based on multi-threading implementation on a single computer with very poor scalability. RESULTS: This paper proposes a new fine-grained 3D array partition method by key-value format in Spark. Our method first converts 3D images to key-value data (K-V). Then the K-V data is used for 3D array partitioning and data exchange in parallel. So Spark-based distributed parallel computing framework can solve the above scalability problem. In this distributed computing framework, all 3D local FSC tasks are simultaneously calculated across multiple nodes in a computer cluster. Through the calculation of experimental data, 3D local resolution evaluation algorithm based on Spark fine-grained 3D array partition has a magnitude change in computing speed compared with the mainstream FSC algorithm under the condition that the accuracy remains unchanged, and has better fault tolerance and scalability. CONCLUSIONS: In this paper, we proposed a K-V format based fine-grained 3D array partition method in Spark to parallel calculating 3D FSC for getting a 3D local resolution density map. 3D local resolution density map evaluates the three-dimensional density maps reconstructed from single particle analysis and subtomogram averaging. Our proposed method can significantly increase the speed of the 3D local resolution evaluation, which is important for the efficient detection of subtle variations among reconstructed macromolecular structures.
Yongchun Lü, Xinhui Tian, Xiao Shi 0003, Xiaohui Zheng, Xin Gao 0001, Min Xu 0009
BMC Bioinform.10
2020 SHREC 2020: Classification in cryo-electron tomograms
Ilja Gubins, Marten L. Chaillet, Gijs van der Schot, Remco C. Veltkamp, Friedrich Förster, Xiaohua Wan 0001, Xuefeng Cui, Fa Zhang 0001, Emmanuel Moebel, Xiao Wang 0004, Daisuke Kihara, Min Xu 0009, Nguyen P. Nguyen, Tommi A. White, Filiz Bunyak
Comput. Graph.14
2020 Few-shot learning for classification of novel macromolecular structures in cryo-electron tomograms
abstract
Cryo-electron tomography (cryo-ET) provides 3D visualization of subcellular components in the near-native state and at sub-molecular resolutions in single cells, demonstrating an increasingly important role in structural biology in situ. However, systematic recognition and recovery of macromolecular structures in cryo-ET data remain challenging as a result of low signal-to-noise ratio (SNR), small sizes of macromolecules, and high complexity of the cellular environment. Subtomogram structural classification is an essential step for such task. Although acquisition of large amounts of subtomograms is no longer an obstacle due to advances in automation of data collection, obtaining the same number of structural labels is both computation and labor intensive. On the other hand, existing deep learning based supervised classification approaches are highly demanding on labeled data and have limited ability to learn about new structures rapidly from data containing very few labels of such new structures. In this work, we propose a novel approach for subtomogram classification based on few-shot learning. With our approach, classification of unseen structures in the training data can be conducted given few labeled samples in test data through instance embedding. Experiments were performed on both simulated and real datasets. Our experimental results show that we can make inference on new structures given only five labeled samples for each class with a competitive accuracy (> 0.86 on the simulated dataset with SNR = 0.1), or even one sample with an accuracy of 0.7644. The results on real datasets are also promising with accuracy > 0.9 on both conditions and even up to 1 on one of the real datasets. Our approach achieves significant improvement compared with the baseline method and has strong capabilities of generalizing to other cellular components.
Liangyong Yu, Bo Zhou 0009, Jing Zhang 0062, Xin Gao 0001, Rui Jiang 0001, Min Xu 0009
PLoS Comput. Biol.10
2020 Poly(A)-DG: A deep-learning-based domain generalization method to identify cross-species Poly(A) signal without prior knowledge from target species
abstract
In eukaryotes, polyadenylation (poly(A)) is an essential process during mRNA maturation. Identifying the cis-determinants of poly(A) signal (PAS) on the DNA sequence is the key to understand the mechanism of translation regulation and mRNA metabolism. Although machine learning methods were widely used in computationally identifying PAS, the need for tremendous amounts of annotation data hinder applications of existing methods in species without experimental data on PAS. Therefore, cross-species PAS identification, which enables the possibility to predict PAS from untrained species, naturally becomes a promising direction. In our works, we propose a novel deep learning method named Poly(A)-DG for cross-species PAS identification. Poly(A)-DG consists of a Convolution Neural Network-Multilayer Perceptron (CNN-MLP) network and a domain generalization technique. It learns PAS patterns from the training species and identifies PAS in target species without re-training. To test our method, we use four species and build cross-species training sets with two of them and evaluate the performance of the remaining ones. Moreover, we test our method against insufficient data and imbalanced data issues and demonstrate that Poly(A)-DG not only outperforms state-of-the-art methods but also maintains relatively high accuracy when it comes to a smaller or imbalanced training set.
Yumin Zheng, Haohan Wang, Yang Zhang 0042, Xin Gao 0001, Eric P. Xing, Min Xu 0009
PLoS Comput. Biol.6
2019 Domain Randomization for Macromolecule Structure Classification and Segmentation in Electron Cyro-tomograms
abstract
It is crucial to study and understand cellular processes. In recent years, Cellular Electron CryoTomography (CECT) serves as a powerful 3D imaging tool to visualize spatial structure of macromolecules inside the cell. However, it is challenging to analyze the macromolecular structures in a systematic way due to nature of the structural complexity of subcellular components. Existing computational and deep learning based approaches suffer from limited scalability, discrimination ability and lack of accurate annotated CECT data. Training with cheap simulated data can alleviate this problem while facing new challenges of bridging the “reality gap” between synthetic training data and real testing data. In this paper, we tackle the tasks of macromolecule structure classification and segmentation in CECT images by adapting a simple but effective technique, domain randomization. We show that by combining deep neural models and domain randomization, we are able to achieve significant improvements of 35.21% and 46.34% in tasks of classification and semantic segmentation for real CECT data, comparing to the model trained only on syhthetic data that aims to faithfully reproduce real-world data distribution.
Chengqian Che, Zhou Xian, Xin Gao 0001, Min Xu 0009
BIBM5
2019 Regularized Adversarial Training (RAT) for Robust Cellular Electron Cryo Tomograms Classification
abstract
Cellular Electron Cryo Tomography (CECT) 3D imaging has permitted biomedical community to study macromolecule structures inside single cells with deep learning approaches. Many deep learning-based methods have since been developed to classify macromolecule structures from tomograms with high accuracy. However, several recent studies have demonstrated the lack of robustness in these models against often-imperceptible, designed changes of input. Therefore, making existing subtomogram-classification models robust remains a serious challenge. In this paper, we study the robustness of the state-of-the-art subtomogram classifier on CECT images and propose a method called Regularized Adversarial Training (RAT) to defend the classifier against a wide range of designed threats. Our results show that RAT improves robustness for CECT image classification over the previous methods.
Xindi Wu, Yijun Mao, Haohan Wang, Xin Gao 0001, Eric P. Xing, Min Xu 0009
BIBM7
2019 Open-set Recognition of Unseen Macromolecules in Cellular Electron Cryo-Tomograms by Soft Large Margin Centralized Cosine Loss
Xuefeng Du, Bo Zhou 0009, Alex Singh, Min Xu 0009
BMVC5
2019 Semi-supervised Macromolecule Structural Classification in Cellular Electron Cryo-Tomograms using 3D Autoencoding Classifier
Xuefeng Du, Rong Xi, Fuya Xu, Bo Zhou 0009, Min Xu 0009
BMVC7
2019 Deep Learning-Based Strategy For Macromolecules Classification with Imbalanced Data from Cellular Electron Cryotomography
abstract
Deep learning model trained by imbalanced data may not work satisfactorily since it could be determined by major classes and thus may ignore the classes with small amount of data. In this paper, we apply deep learning based imbalanced data classification for the first time to cellular macromolecular complexes captured by Cryo-electron tomography (Cryo-ET). We adopt a range of strategies to cope with imbalanced data, including data sampling, bagging, boosting, Genetic Programming based method and. Particularly, inspired from Inception 3D network, we propose a multi-path CNN model combining focal loss and mixup on the Cryo-ET dataset to expand the dataset, where each path had its best performance corresponding to each type of data and let the network learn the combinations of the paths to improve the classification performance. In addition, extensive experiments have been conducted to show our proposed method is flexible enough to cope with different number of classes by adjusting the number of paths in our multi-path model. To our knowledge, this work is the first application of deep learning methods of dealing with imbalanced data to the internal tissue classification of cell macromolecular complexes, which opened up a new path for cell classification in the field of computational biology.
Ziqian Luo, Zhipeng Bao, Min Xu 0009
IJCNN4
2019 Neural Architecture Search for Adversarial Medical Image Segmentation
Nanqing Dong, Min Xu 0009, Xiaodan Liang, Yiliang Jiang, Wei Dai 0003, Eric P. Xing
MICCAI (6)2
2019 A joint method for marker-free alignment of tilt series in electron tomography
abstract
MOTIVATION: Electron tomography (ET) is a widely used technology for 3D macro-molecular structure reconstruction. To obtain a satisfiable tomogram reconstruction, several key processes are involved, one of which is the calibration of projection parameters of the tilt series. Although fiducial marker-based alignment for tilt series has been well studied, marker-free alignment remains a challenge, which requires identifying and tracking the identical objects (landmarks) through different projections. However, the tracking of these landmarks is usually affected by the pixel density (intensity) change caused by the geometry difference in different views. The tracked landmarks will be used to determine the projection parameters. Meanwhile, different projection parameters will also affect the localization of landmarks. Currently, there is no alignment method that takes interrelationship between the projection parameters and the landmarks. RESULTS: Here, we propose a novel, joint method for marker-free alignment of tilt series in ET, by utilizing the information underlying the interrelationship between the projection model and the landmarks. The proposed method is the first joint solution that combines the extrinsic (track-based) alignment and the intrinsic (intensity-based) alignment, in which the localization of landmarks and projection parameters keep refining each other until convergence. This iterative approach makes our solution robust to different initial parameters and extreme geometric changes, which ensures a better reconstruction for marker-free ET. Comprehensive experimental results on three real datasets show that our new method achieved a significant improvement in alignment accuracy and reconstruction quality, compared to the state-of-the-art methods. AVAILABILITY AND IMPLEMENTATION: The main program is available at https://github.com/icthrm/joint-marker-free-alignment. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Renmin Han, Zhipeng Bao, Tongxin Niu, Fa Zhang 0001, Min Xu 0009, Xin Gao 0001
Bioinform.6
2019 Adversarial domain adaptation for cross data source macromolecule in situ structural classification in cellular electron cryo-tomograms
abstract
MOTIVATION: Since 2017, an increasing amount of attention has been paid to the supervised deep learning-based macromolecule in situ structural classification (i.e. subtomogram classification) in cellular electron cryo-tomography (CECT) due to the substantially higher scalability of deep learning. However, the success of such supervised approach relies heavily on the availability of large amounts of labeled training data. For CECT, creating valid training data from the same data source as prediction data is usually laborious and computationally intensive. It would be beneficial to have training data from a separate data source where the annotation is readily available or can be performed in a high-throughput fashion. However, the cross data source prediction is often biased due to the different image intensity distributions (a.k.a. domain shift). RESULTS: We adapt a deep learning-based adversarial domain adaptation (3D-ADA) method to timely address the domain shift problem in CECT data analysis. 3D-ADA first uses a source domain feature extractor to extract discriminative features from the training data as the input to a classifier. Then it adversarially trains a target domain feature extractor to reduce the distribution differences of the extracted features between training and prediction data. As a result, the same classifier can be directly applied to the prediction data. We tested 3D-ADA on both experimental and realistically simulated subtomogram datasets under different imaging conditions. 3D-ADA stably improved the cross data source prediction, as well as outperformed two popular domain adaptation methods. Furthermore, we demonstrate that 3D-ADA can improve cross data source recovery of novel macromolecular structures. AVAILABILITY AND IMPLEMENTATION: https://github.com/xulabs/projects. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Ruogu Lin, Kris Makoto Kitani, Min Xu 0009
Bioinform.4
2019 Automatic localization and identification of mitochondria in cellular electron cryo-tomography using faster-RCNN
abstract
BACKGROUND: Cryo-electron tomography (cryo-ET) enables the 3D visualization of cellular organization in near-native state which plays important roles in the field of structural cell biology. However, due to the low signal-to-noise ratio (SNR), large volume and high content complexity within cells, it remains difficult and time-consuming to localize and identify different components in cellular cryo-ET. To automatically localize and recognize in situ cellular structures of interest captured by cryo-ET, we proposed a simple yet effective automatic image analysis approach based on Faster-RCNN. RESULTS: Our experimental results were validated using in situ cyro-ET-imaged mitochondria data. Our experimental results show that our algorithm can accurately localize and identify important cellular structures on both the 2D tilt images and the reconstructed 2D slices of cryo-ET. When ran on the mitochondria cryo-ET dataset, our algorithm achieved Average Precision >0.95. Moreover, our study demonstrated that our customized pre-processing steps can further improve the robustness of our model performance. CONCLUSIONS: In this paper, we proposed an automatic Cryo-ET image analysis algorithm for localization and identification of different structure of interest in cells, which is the first Faster-RCNN based method for localizing an cellular organelle in Cryo-ET images and demonstrated the high accuracy and robustness of detection and classification tasks of intracellular mitochondria. Furthermore, our approach can be easily applied to detection tasks of other cellular structures as well.
Stephanie E. Sigmund, Ruogu Lin, Bo Zhou 0009, Chang Liu 0031, Rui Jiang 0001, Zachary Freyberg, Hairong Lv, Min Xu 0009
BMC Bioinform.11
2019 Fine-grained alignment of cryo-electron subtomograms based on MPI parallel optimization
abstract
Cryo-electron tomography (Cryo-ET) is an imaging technique used to generate three-dimensional structures of cellular macromolecule complexes in their native environment. Due to developing cryo-electron microscopy technology, the image quality of three-dimensional reconstruction of cryo-electron tomography has greatly improved. However, cryo-ET images are characterized by low resolution, partial data loss and low signal-to-noise ratio (SNR). In order to tackle these challenges and improve resolution, a large number of subtomograms containing the same structure needs to be aligned and averaged. Existing methods for refining and aligning subtomograms are still highly time-consuming, requiring many computationally intensive processing steps (i.e. the rotations and translations of subtomograms in three-dimensional space). In this article, we propose a Stochastic Average Gradient (SAG) fine-grained alignment method for optimizing the sum of dissimilarity measure in real space. We introduce a Message Passing Interface (MPI) parallel programming model in order to explore further speedup. We compare our stochastic average gradient fine-grained alignment algorithm with two baseline methods, high-precision alignment and fast alignment. Our SAG fine-grained alignment algorithm is much faster than the two baseline methods. Results on simulated data of GroEL from the Protein Data Bank (PDB ID:1KP8) showed that our parallel SAG-based fine-grained alignment method could achieve close-to-optimal rigid transformations with higher precision than both high-precision alignment and fast alignment at a low SNR (SNR=0.003) with tilt angle range ±60 ∘ or ±40 ∘ . For the experimental subtomograms data structures of GroEL and GroEL/GroES complexes, our parallel SAG-based fine-grained alignment can achieve higher precision and fewer iterations to converge than the two baseline methods.
Yongchun Lü, Shirui Li, Hua Li 0009, Xin Gao 0001, Min Xu 0009
BMC Bioinform.7
2018 Feature Decomposition Based Saliency Detection in Electron Cryo-Tomograms
Bo Zhou 0009, Xin Gao 0001, Min Xu 0009
BIBM6
2018 Multi-task Learning for Macromolecule Classification, Segmentation and Coarse Structural Recovery in Cryo-Tomography
Chang Liu 0031, Min Xu 0009
BMVC5
2018 Image-derived generative modeling of pseudo-macromolecular structures - towards the statistical assessment of Electron CryoTomography template matching
Xiaodan Liang, Zhiguang Huo, Eric P. Xing, Min Xu 0009
BMVC6
2018 Deep Learning Based Supervised Semantic Segmentation of Electron Cryo-Subtomograms
abstract
Cellular Electron Cryo-Tomography (CECT) is a powerful imaging technique for the 3D visualization of cellular structure and organization at submolecular resolution. It enables analyzing the native structures of macromolecular complexes and their spatial organization inside single cells. However, due to the high degree of structural complexity and practical imaging limitations, systematic macromolecular structural recovery inside CECT images remains challenging. Particularly, the recovery of a macromolecule is likely to be biased by its neighbor structures due to the high molecular crowding. To reduce the bias, here we introduce a novel 3D convolutional neural network inspired by Fully Convolutional Network and Encoder-Decoder Architecture for the supervised segmentation of macromolecules of interest in subtomograms. The tests of our models on realistically simulated CECT data demonstrate that our new approach has significantly improved segmentation performance compared to our baseline approach. Also, we demonstrate that the proposed model has generalization ability to segment new structures that do not exist in training data.
Chang Liu 0031, Ruogu Lin, Xiaodan Liang, Zachary Freyberg, Eric P. Xing, Min Xu 0009
ICIP7
2018 Respond-CAM: Analyzing Deep Models for 3D Imaging Data by Visualizations
Guannan Zhao, Bo Zhou 0009, Rui Jiang 0001, Min Xu 0009
MICCAI (1)5
2018 An integration of fast alignment and maximum-likelihood methods for electron subtomogram averaging and classification
abstract
Motivation: Cellular Electron CryoTomography (CECT) is an emerging 3D imaging technique that visualizes subcellular organization of single cells at sub-molecular resolution and in near-native state. CECT captures large numbers of macromolecular complexes of highly diverse structures and abundances. However, the structural complexity and imaging limits complicate the systematic de novo structural recovery and recognition of these macromolecular complexes. Efficient and accurate reference-free subtomogram averaging and classification represent the most critical tasks for such analysis. Existing subtomogram alignment based methods are prone to the missing wedge effects and low signal-to-noise ratio (SNR). Moreover, existing maximum-likelihood based methods rely on integration operations, which are in principle computationally infeasible for accurate calculation. Results: Built on existing works, we propose an integrated method, Fast Alignment Maximum Likelihood method (FAML), which uses fast subtomogram alignment to sample sub-optimal rigid transformations. The transformations are then used to approximate integrals for maximum-likelihood update of subtomogram averages through expectation-maximization algorithm. Our tests on simulated and experimental subtomograms showed that, compared to our previously developed fast alignment method (FA), FAML is significantly more robust to noise and missing wedge effects with moderate increases of computation cost. Besides, FAML performs well with significantly fewer input subtomograms when the FA method fails. Therefore, FAML can serve as a key component for improved construction of initial structural models from macromolecules captured by CECT. Availability and implementation: http://www.cs.cmu.edu/mxu1.
Yixiu Zhao, Min Xu 0009
Bioinform.4
2018 Improved deep learning-based macromolecules structure classification from electron cryo-tomograms
Chengqian Che, Ruogu Lin, Karim Elmaaroufi, John M. Galeotti, Min Xu 0009
Mach. Vis. Appl.6
2017 Deep learning-based subdivision approach for large scale macromolecules structure recovery from electron cryo tomograms
abstract
MOTIVATION: Cellular Electron CryoTomography (CECT) enables 3D visualization of cellular organization at near-native state and in sub-molecular resolution, making it a powerful tool for analyzing structures of macromolecular complexes and their spatial organizations inside single cells. However, high degree of structural complexity together with practical imaging limitations makes the systematic de novo discovery of structures within cells challenging. It would likely require averaging and classifying millions of subtomograms potentially containing hundreds of highly heterogeneous structural classes. Although it is no longer difficult to acquire CECT data containing such amount of subtomograms due to advances in data acquisition automation, existing computational approaches have very limited scalability or discrimination ability, making them incapable of processing such amount of data. RESULTS: To complement existing approaches, in this article we propose a new approach for subdividing subtomograms into smaller but relatively homogeneous subsets. The structures in these subsets can then be separately recovered using existing computation intensive methods. Our approach is based on supervised structural feature extraction using deep learning, in combination with unsupervised clustering and reference-free classification. Our experiments show that, compared with existing unsupervised rotation invariant feature and pose-normalization based approaches, our new approach achieves significant improvements in both discrimination ability and scalability. More importantly, our new approach is able to discover new structural classes and recover structures that do not exist in training data. AVAILABILITY AND IMPLEMENTATION: Source code freely available at http://www.cs.cmu.edu/∼mxu1/software . CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Min Xu 0009, Xiaoqi Chai, Hariank Muthakana, Xiaodan Liang, Ge Yang 0002, Tzviya Zeev-Ben-Mordehai, Eric P. Xing
Bioinform.1
2016 Simulating cryo electron tomograms of crowded cell cytoplasm for assessment of automated particle picking
abstract
BACKGROUND: Cryo-electron tomography is an important tool to study structures of macromolecular complexes in close to native states. A whole cell cryo electron tomogram contains structural information of all its macromolecular complexes. However, extracting this information remains challenging, and relies on sophisticated image processing, in particular for template-free particle extraction, classification and averaging. To develop these methods it is crucial to realistically simulate tomograms of crowded cellular environments, which can then serve as ground truth models for assessing and optimizing methods for detection of complexes in cell tomograms. RESULTS: We present a framework to generate crowded mixtures of macromolecular complexes for realistically simulating cryo electron tomograms including noise and image distortions due to the missing-wedge effects. Simulated tomograms are then used for assessing the template-free Difference-of-Gaussian (DoG) particle-picking method to detect complexes of different shapes and sizes under various crowding and noise levels. We identified DoG parameter settings that maximize precision and recall for detecting particles over a wide range of sizes and shapes. We observed that medium sized DoG scaling factors showed the overall best performance. To further improve performance, we propose a combination strategy for integrating results from multiple parameter settings. With increasing macromolecular crowding levels, the precision of particle picking remained relatively high, while the recall was dramatically reduced, which limits the detection of sufficient copy numbers of complexes in a crowded environment. Over a wide range of increasing noise levels, the DoG particle picking performance remained stable, but dramatically reduced beyond a specific noise threshold. CONCLUSIONS: Automatic and reference-free particle picking is an important first step in a visual proteomics analysis of cell tomograms. However, cell cytoplasm is highly crowded, which makes particle detection challenging. It is therefore important to test particle-picking methods in a realistic crowded setting. Here, we present a framework for simulating tomograms of cellular environments at high crowding levels and assess the DoG particle picking method. We determined optimal parameter settings to maximize the performance of the DoG particle-picking method.
Long Pei, Min Xu 0009, Zachary Frazier, Frank Alber
BMC Bioinform.2
2013 Automated target segmentation and real space fast alignment methods for high-throughput classification and averaging of crowded cryo-electron subtomograms
abstract
MOTIVATION: Cryo-electron tomography allows the imaging of macromolecular complexes in near living conditions. To enhance the nominal resolution of a structure it is necessary to align and average individual subtomograms each containing identical complexes. However, if the sample of complexes is heterogeneous, it is necessary to first classify subtomograms into groups of identical complexes. This task becomes challenging when tomograms contain mixtures of unknown complexes extracted from a crowded environment. Two main challenges must be overcomed: First, classification of subtomograms must be performed without knowledge of template structures. However, most alignment methods are too slow to perform reference-free classification of a large number of (e.g. tens of thousands) of subtomograms. Second, subtomograms extracted from crowded cellular environments, contain often fragments of other structures besides the target complex. However, alignment methods generally assume that each subtomogram only contains one complex. Automatic methods are needed to identify the target complexes in a subtomogram even when its shape is unknown. RESULTS: In this article, we propose an automatic and systematic method for the isolation and masking of target complexes in subtomograms extracted from crowded environments. Moreover, we also propose a fast alignment method using fast rotational matching in real space. Our experiments show that, compared with our previously proposed fast alignment method in reciprocal space, our new method significantly improves the alignment accuracy for highly distorted and especially crowded subtomograms. Such improvements are important for achieving successful and unbiased high-throughput reference-free structural classification of complexes inside whole-cell tomograms. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Min Xu 0009, Frank Alber
Bioinform.1
2011 Template-free detection of macromolecular complexes in cryo electron tomograms
abstract
MOTIVATION: Cryo electron tomography (CryoET) produces 3D density maps of biological specimen in its near native states. Applied to small cells, cryoET produces 3D snapshots of the cellular distributions of large complexes. However, retrieving this information is non-trivial due to the low resolution and low signal-to-noise ratio in tomograms. Current pattern recognition methods identify complexes by matching known structures to the cryo electron tomogram. However, so far only a small fraction of all protein complexes have been structurally resolved. It is, therefore, of great importance to develop template-free methods for the discovery of previously unknown protein complexes in cryo electron tomograms. RESULTS: Here, we have developed an inference method for the template-free discovery of frequently occurring protein complexes in cryo electron tomograms. We provide a first proof-of-principle of the approach and assess its applicability using realistically simulated tomograms, allowing for the inclusion of noise and distortions due to missing wedge and electron optical factors. Our method is a step toward the template-free discovery of the shapes, abundance and spatial distributions of previously unknown macromolecular complexes in whole cell tomograms. CONTACT: [email protected]
Min Xu 0009, Martin Beck, Frank Alber
Bioinform.1
2010 A fast mathematical programming procedure for simultaneous fitting of assembly components into cryoEM density maps
abstract
MOTIVATION: Single-particle cryo electron microscopy (cryoEM) typically produces density maps of macromolecular assemblies at intermediate to low resolution (approximately 5-30 A). By fitting high-resolution structures of assembly components into these maps, pseudo-atomic models can be obtained. Optimizing the quality-of-fit of all components simultaneously is challenging due to the large search space that makes the exhaustive search over all possible component configurations computationally unfeasible. RESULTS: We developed an efficient mathematical programming algorithm that simultaneously fits all component structures into an assembly density map. The fitting is formulated as a point set matching problem involving several point sets that represent component and assembly densities at a reduced complexity level. In contrast to other point matching algorithms, our algorithm is able to match multiple point sets simultaneously and not only based on their geometrical equivalence, but also based on the similarity of the density in the immediate point neighborhood. In addition, we present an efficient refinement method based on the Iterative Closest Point registration algorithm. The integer quadratic programming method generates an assembly configuration in a few seconds. This efficiency allows the generation of an ensemble of candidate solutions that can be assessed by an independent scoring function. We benchmarked the method using simulated density maps of 11 protein assemblies at 20 A, and an experimental cryoEM map at 23.5 A resolution. Our method was able to generate assembly structures with root-mean-square errors <6.5 A, which have been further reduced to <1.8 A by the local refinement procedure. AVAILABILITY: The program is available upon request as a Matlab code package. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics Online.
Daven Vasishtan, Min Xu 0009, Maya Topf, Frank Alber
Bioinform.3
2010 Unraveling complex temporal associations in cellular systems across multiple time-series microarray datasets
Wenyuan Li 0006, Min Xu 0009, Xianghong Jasmine Zhou
J. Biomed. Informatics2
2009 3D Rotation Invariant Features for the Characterization of Molecular Density Maps
abstract
Cryo electron microscopy produces 3D density maps of large macromolecular structures. An important task lies in the efficient identification of structural similarities between different molecular density maps. Here we construct and test three different types of 3D rotation invariant features for template free similarity detection in molecular density maps. The density map comparison is based on feature vectors that describe the surrounding density distribution for a given map position. We propose Fast Fourier Transform based methods to speed up the computation of feature vectors. Previously, little is known about the discriminative power of rotation invariant features for noisy maps. Here, we test the three feature types with density maps at different noise levels. We assess the performance of our feature vectors by a classification experiment of protein density maps. Our results show that at low noise levels the three types of features perform equally well. However, at high noise levels the features that are constructed by a spherical harmonics decomposition of the density neighborhood is significantly more reliable and outperform the other two feature types, which are based on the moments and intensity histograms of the density distribution.
Min Xu 0009, Frank Alber
BIBM1
2007 A Robust Method for Generating Discriminative Gene Clusters
abstract
Microarray technology is often used to identify the genes that are differentially expressed between two biological conditions. Since microarray datasets contain a small number of samples and a large number of genes, it is not difficult to find small gene subsets which are highly discriminative. However, such identified classifiers tend to have poor generalization properties on the test samples due to overfitting. We propose a novel approach for generating discriminative gene clusters. Our experiments on both simulated and real datasets show that our method can generate a series of robust gene clusters with good classification performance.
Min Xu 0009, Louxin Zhang, Pei Li Joe Zhou
BIBE1
2006 Integrative Array Analyzer: a software package for analysis of cross-platform and cross-species microarray data
abstract
The rapid accumulation of microarray data translates into an urgent need for tools to perform integrative microarray analysis. Integrative Array Analyzer is a comprehensive analysis and visualization software toolkit, which aims to facilitate the reuse of the large amount of cross-platform and cross-species microarray data. It is composed of the data preprocess module, the co-expression analysis module, the differential expression analysis module, the functional and transcriptional annotation module and the graph visualization module.
Kiran Kamath, Kangyu Zhang, Sudip Pulapura, Avinash Achar, Juan Nunez-Iglesias, Yu Huang 0003, Xifeng Yan, Jiawei Han 0001, Haiyan Hu 0004, Min Xu 0009, Jianjun Hu, Xianghong Jasmine Zhou
Bioinform.11