VLDB 2026 Research / reviewers in the wild / expert
Fa Zhang 0001
dblp:10/1849-1
· DBLP profile ↗
135ranked-venue papers
6as first author
66since 2021 · last 2026
0000-0002-2081-9369ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 90 · 5 first-author · 56 since 2021Computer networks · 17 · 1 since 2021Systems, architecture and hardware · 15 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 6 since 2021Artificial intelligence and machine learning · 5 · 5 since 2021Theory of computation · 3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ST-LLM: Spatial Transcriptomics Embedding with Large Language ModelsabstractSpatial transcriptomics provides unprecedented opportunities to analyze gene patterns while preserving spatial tissue architecture. However, traditional deep learning methods for spatial transcriptomics analysis face significant challenges in multi-modal data integration, spatial dependency modeling, and biological knowledge incorporation, while existing large language models lack explicit spatial modeling capabilities for transcriptomic data. So we first present a Spatial Transcriptomics Embedding with Large Language Models (ST-LLM), a novel simple and effective approach that transforms intricate spatial graph structures into structured textual representations suitable for large language models (LLMs). ST-LLM dynamically constructs graph adjacency construction using reinforcement learning paradigms to adaptively optimize spatial relationships, converts the resulting graphs into hierarchical textual descriptions with spatial context, and leverages pre-trained semantic understanding to generate high-dimensional spatial-aware representations. Comprehensive experiments on 14 datasets demonstrate that ST-LLM achieves comparable or better performance than traditional model. ST-LLM shows that LLMs embeddings provide a new simple and effective path to encoding spatial transcriptomics biological knowledge. Zhetao Xu, Fa Zhang 0001, Bin Hu 0001 |
AAAI | 6 |
| 2026 | Cyto-SSL: A Self-Supervised Pretraining Framework for Cytology Foundation ModelabstractCytological images originate from exfoliated cells, collected via liquid-based slides and digitized into whole slide images (WSIs). Unlike histological WSIs that exhibit continuous and well-structured tissue, cytological WSIs are sparse in spatial distribution and unstructured in cellular relationships. Typically, the nucleus serves as the primary diagnostic feature, while surrounding cytoplasmic information plays a supportive role. These unique characteristics limit the development of effective foundation models and hinder the transferability of histology-based models for cytopathology. To address this, we propose **Cyto-SSL**, the first self-supervised pretraining framework for cytological images. It introduces **Nuclei-Centered Perturbation**, which highlights individual nuclei by perturbing non-nuclear regions. We also design an SR-Transformer module, which complements this by using sparse attention to concentrate on diagnostically relevant scattered cells, while iRPE helps model to capture local spatial relationships and avoids unnecessary attention to irrelevant global structures. Experimental results show that **Cyto-SSL** enhances performance across diverse cytological datasets and Multiple Instance Learning (MIL) methods. On a WSI-level dataset, it achieved 95.67% accuracy and outperformed ImageNet-pretrained ResNet-50 by 11.33%, demonstrating superior feature representation for cytological analysis. Additionally, **Cyto-SSL** modules are plug-and-play, easily integrated into other pretraining frameworks, yielding a 2.6% accuracy gain across different SSL methods. Rui Yan 0009, Zhetao Xu, Ying Wang 0043, Fa Zhang 0001, Bin Hu 0001 |
AAAI | 8 |
| 2026 | A Unified Benchmark and Conditional Sequence Modeling Framework for Future Influenza Evolution Prediction
Zhuolun Li, Xulinyi Huang, Qianyao Lin, Yuxiao Cui, Shabir Madhi, Fa Zhang 0001 |
ISBRA (1) | 8 |
| 2026 | CryoDETR: a Deformable DETR-Based Method for Particle Picking in Cryo-EM Micrographs
Xuan Wang 0002, Fa Zhang 0001 |
ISBRA (2) | 4 |
| 2026 | DyMamba: dynamic Mamba for microscopy image semantic segmentationabstractMOTIVATION: Segmentation of cell bodies and organelles in microscopy images is critical for biological research, particularly in scenarios with multiple regions of interest where spatial continuity is essential. The Mamba architecture, derived from State Space Models (SSMs), has recently gained attention for efficiently modeling long-range dependencies in sequences, achieving excellent results in both natural and medical image segmentation. However, in vision tasks, current Mamba scanning strategies mainly focus on raster-scanning and local-scanning, which introduce spatial discontinuities, severely affecting the effectiveness of segmentation at the pixel level, especially in dense segmentation tasks. RESULTS: In this article, we propose DyMamba, a Mamba-based model featuring a dynamic scanning strategy that adaptively plans scanning paths based on local features and complexity. In addition, to address the challenges of detail prediction and small object detection, we introduce a local aware module that performs pixel-level regional processing on images. DyMamba achieves robust segmentation across diverse microscopy image types, including cell-, organelle- and tissue-scale images. Experiments on six datasets and multiple scanning strategies demonstrate the excellent performance of our method in segmenting microscopy images, achieving an average improvement of 6.9% in mDice and 4.3% in mIoU over state-of-the-art methods across all datasets. AVAILABILITY: The code is released at https://github.com/cbqBit/dymamba. Buqing Cai, Xingsheng Wang, Zhuo Jia, Fa Zhang 0001, Bin Hu 0001 |
Bioinform. | 4 |
| 2026 | A variational framework with composite sparse regularization for cryo-electron tomography reconstruction
Chenyun Yu, Zihe Xu, Qiong Zeng, Haythem El-Messiry, Fa Zhang 0001, Renmin Han |
Bioinform. | 6 |
| 2026 | Point cloud deformation modeling for particle selection following cryo-EM 2D classificationabstractBACKGROUND: Cryo-electron microscopy (cryo-EM) has emerged as a powerful technique for high-resolution structural determination of macromolecules. However, accurately classifying single-particle cryo-EM images remains challenging, especially when dealing with deformed particles. In traditional 2D classification methods, clustering algorithms are used for classification. This assumption leads to some deformed particles being misclassified in 2D images, which adversely affects downstream tasks. To address this challenge, we propose a point cloud-based deformation measurement model that integrates a Variational Autoencoder (VAE) with a heuristic point cloud matching algorithm to calculate particle deformation values. RESULTS: This model enables the identification and removal of particles with large deformations. Our experiments on simulated and real cryo-EM datasets, including Tobacco Mosaic Virus (TMV) and mixed capsids of MS2 virions (MS2). The model achieves robust classification (F1: 0.85-0.88) while preserving 93-95% of structural details, and can effectively filter out deformed particles after 2D classification. CONCLUSION: The model identifies and removes deformed or misclassified particles to improve classification quality. It serves as a data-filtering post-processing step following 2D classification. By improving the quality of particle datasets, it enhances the reliability of subsequent analysis in cryo-EM. Xuan Wang 0042, Zhengao Mo, Fuwei Li, Fa Zhang 0001 |
BMC Bioinform. | 4 |
| 2025 | CoDiST: Combining Unimodal Contrastive Denoising and Crossmodal Disentanglement for Spatial TranscriptomicsabstractIntegrating spatial transcriptomics (ST) with histology enables precise delineation of spatial domains in complex tissues. However, effective multimodal integration faces two primary challenges: (1) High sparsity and frequent dropout events in transcriptomic data and artifacts in histological images together obscure true biological signals, resulting in intra-modal noise. (2) Excessive emphasis on modality alignment in current fusion methods allows redundant information to dominate over modality -specific features, leading to inter-modal redundancy. To address these challenges, we propose CoDiST, a novel multi-modal representation learning framework. Specifically, CoDiST uses histological information to strengthen spatial neighborhood relationships and employs unimodal contrastive learning to enhance robustness against technical noise in both modalities. Furthermore, to overcome inter-modal redundancy, CoDiST introduces the Gene Image Disentanglement Network (GIDNet). It disentangles representations into a shared subspace and specific subspaces, achieving crossmodal semantic alignment while preserving complementary features unique to each modality. In benchmarking on multiple ST datasets, CoDiST delivers clear gains over existing methods. On the Human Breast Cancer dataset, CoDiST achieves a 0.66 Adjusted Rand Index (ARI), representing a 3.6% improvement over state-of-the-art methods. Kai Hu 0002, Xuefeng Cui, Fa Zhang 0001 |
BIBM | 4 |
| 2025 | MambaST: Hexagonal State Space Modeling for Spatial Domain Identification
Kai Hu 0002, Xuefeng Cui, Fa Zhang 0001 |
ISBRA (1) | 4 |
| 2025 | UPicker: a semi-supervised particle picking transformer method for cryo-EM micrographsabstractAutomatic single particle picking is a critical step in the data processing pipeline of cryo-electron microscopy structure reconstruction. In recent years, several deep learning-based algorithms have been developed, demonstrating their potential to solve this challenge. However, current methods highly depend on manually labeled training data, which is labor-intensive and prone to biases especially for high-noise and low-contrast micrographs, resulting in suboptimal precision and recall. To address these problems, we propose UPicker, a semi-supervised transformer-based particle-picking method with a two-stage training process: unsupervised pretraining and supervised fine-tuning. During the unsupervised pretraining, an Adaptive Laplacian of Gaussian region proposal generator is proposed to obtain pseudo-labels from unlabeled data for initial feature learning. For the supervised fine-tuning, UPicker only needs a small amount of labeled data to achieve high accuracy in particle picking. To further enhance model performance, UPicker employs a contrastive denoising training strategy to reduce redundant detections and accelerate convergence, along with a hybrid data augmentation strategy to deal with limited labeled data. Comprehensive experiments on both simulated and experimental datasets demonstrate that UPicker outperforms state-of-the-art particle-picking methods in terms of accuracy and robustness while requiring fewer labeled data than other transformer-based models. Furthermore, ablation studies demonstrate the effectiveness and necessity of each component of UPicker. The source code and data are available at https://github.com/JachyLikeCoding/UPicker. Chi Zhang 0110, Yiran Cheng, Kaiwen Feng, Fa Zhang 0001, Renmin Han, Jieqing Feng |
Briefings Bioinform. | 4 |
| 2025 | TiltRec: an ultra-fast and open-source toolkit for cryo-electron tomographic reconstructionabstractMOTIVATION: Cryo-electron tomography (cryo-ET) has revolutionized our ability to observe structures from the subcellular to the atomic level in their native states. Achieving high-resolution reconstruction involves collecting tilt series at different angles and subsequently backprojecting them into 3D space or iteratively reconstructing them to build a 3D volume of the specimen. However, the intricate computational demands of tomographic reconstruction pose significant challenges, requiring extensive calculation times that hinder efficiency, especially with large and complex datasets. RESULTS: We present TiltRec, an open-source toolkit that leverages the parallel capabilities of Central Processing Units and Graphics Processing Units to enhance tomographic reconstruction. TiltRec implements six classical tomographic reconstruction algorithms, utilizing optimized parallel computation strategies and advanced memory management techniques. Performance evaluations across multiple datasets of varying sizes demonstrate that TiltRec significantly improves efficiency, reducing computational times while maintaining reconstruction resolution. SUMMARY: TiltRec effectively addresses the computational challenges associated with cryo-ET reconstruction by fully exploiting parallel acceleration. As an open-source tool, TiltRec not only facilitates extensive applications by the research community but also supports further algorithm modifications and extensions, enabling the continued development of novel algorithms. AVAILABILITY AND IMPLEMENTATION: The source code, documentation, and sample data can be downloaded at https://github.com/icthrm/TiltRec. Yanxin Jiao, Hongjia Li 0001, Fa Zhang 0001, Dawei Zang, Renmin Han |
Bioinform. | 6 |
| 2025 | CryoAlign2: efficient global and local Cryo-EM map retrieval based on parallel-accelerated local spatial structural featuresabstractMOTIVATION: With the rapid advancements in Cryo-Electron Microscopy (Cryo-EM), an increasing number of high-resolution 3D density maps are being made publicly available, highlighting the urgent need for efficient structure similarity retrieval. Exploring map similarity at various levels is critical for fully utilizing these valuable resources. Our previously proposed CryoAlign can provide more accurate density map alignment while maintaining a low failure rate. However, CryoAlign only offers a method for aligning density maps, with low efficiency in local alignment, and has not yet been applied to the retrieval of Cryo-EM density maps. RESULTS: We have developed an alignment-based retrieval tool to perform both global and local retrieval. Our approach adopts parallel-accelerated CryoAlign for high-precision 3D alignment and transforms density maps into point clouds for efficient retrieval and storage. Additionally, a multi-dimension scoring function is introduced to accurately assess structural similarities between superimposed density maps. To demonstrate its applicability, we conducted thorough testing across different retrieval tasks, such as global, local or hybrid similarity retrieval. Our tool achieves up to a 7-fold speedup while supporting precise local alignments. Comprehensive experiments demonstrate that even when one density map is entirely contained within another, our tool performs exceptionally well in high-resolution density map retrieval. It provides researchers with an efficient and accurate solution for density map similarity search. AVAILABILITY AND IMPLEMENTATION: The source code, documentation, and sample data can be downloaded at https://github.com/JokerL2/CryoAlign2. Bintao He, Chenjie Feng, Fa Zhang 0001, Zhongjun Yang, Renmin Han |
Bioinform. | 5 |
| 2025 | Central Feature Network Enables Accurate Detection of Both Small and Large Particles in Cryo-Electron Tomography
Yaoyu Wang, Fa Zhang 0001, Xuefeng Cui |
J. Comput. Sci. Technol. | 4 |
| 2025 | Similarity-guided multi-view functional brain network fusion
Jingyu Liu 0002, Mengkai Sun, Fa Zhang 0001, Bin Hu 0001, Qunxi Dong |
Medical Image Anal. | 4 |
| 2025 | Particle Restoration: A Novel Image Processing Framework for Improving Real Cryo-EM Image Quality in Single Particle AnalysisabstractCryo-electron microscopy single particle analysis (cryo-EM SPA) is the most powerful technique for biomacromolecule structure determination. However, many factors such as complicated noise and radiation damage make the quality of cryo-EM images extremely poor, where high-frequency structure details are submerged, limiting the application of deep learning and suppressing the resolution of reconstruction. Thus, image restoration is of vital importance. Some related works explore micrograph restoration, but the particles in restored micrographs are still of poor quality. Moreover, the training approach of existing methods uses noisy observations or simulated data as supervision, leading to reduced performance on real cryo-EM data. In this paper, we define the task of particle restoration and propose a novel 4-step framework to this end. Labels are created for each particle image and paired data is collected within our framework, compensating for the absence of ground truth. A deep neural network with encoder-decoder architecture is designed to learn the mapping from degraded particles to high-quality ones, while other networks can also be employed as a plug-and-play module. Three datasets are constructed from real cryo-EM data and extensive experiments are carried out. Both quantitative metrics and qualitative visualization indicate that our framework is effective for cryo-EM particle restoration. It becomes easier to extract particle features after restoration, aiding in SPA and the effective application of deep learning on cryo-EM images. The downstream task experiments of cryo-EM SPA are also conducted, showing that the proposed framework has the potential to improve cryo-EM SPA performance. Bin Hu 0001, Shiqi Liu 0004, Xiaoliang Xie, Xiao-Hu Zhou, Hong-Jia Li, Qing-Bing Zheng, Fa Zhang 0001, Zeng-Guang Hou, Ning-Shao Xia |
IEEE Trans. Comput. Biol. Bioinform. | 8 |
| 2025 | Enhancing Colorectal Lesion Segmentation Through Internal Feature Extraction and Computational Modeling InsightsabstractColorectal cancer (CRC) is a prevalent malignancy with significant social and healthcare implications, ranking among the top causes of cancer-related mortality worldwide. In this computational social systems context, we address the challenge of accurately segmenting CRC lesions from colonoscopy images, which is pivotal for early cancer detection and treatment. The complexity of intestinal environments and variations in medical expertise contribute to the high rate of undetected or misdiagnosed lesions, underscoring the need for advanced computational models. This study presents ColoSegNet, a novel self-supervised deep learning framework designed to enhance the accuracy and effectiveness of CRC diagnosis. By leveraging a comprehensive and annotated colorectal lesion segmentation dataset (CLSD), ColoSegNet incorporates a temporal correlation module to extract critical features from colonoscopy video frames, significantly improving the segmentation of colorectal lesions. Furthermore, ColoSegNet employs a masked autoencoder (MAE) module for self-supervised image reconstruction, preserving the original image integrity and facilitating precise segmentation. Comparative assessments against established models such as UNet, PraNet, and Deeplab V3 demonstrate ColoSegNet’s superior performance in detailed feature representation and overall segmentation accuracy. This research not only contributes to the field of medical imaging but also to computational social systems, by capturing inherent data patterns and integrating specialized modules for feature representation in a healthcare context. Our findings provide valuable insights into the early and accurate detection of CRC, a critical issue given the disease’s high incidence and mortality rates, and its impact on social systems. Yulong Hu, Dehui Qiu, Rui Li 0115, Liguo Deng, Tinghui Ye, Shengtao Zhu, Xiujing Sun, Weilong Yao, Fa Zhang 0001 |
IEEE Trans. Comput. Soc. Syst. | 11 |
| 2025 | Toward a New Fusion Paradigm: Integrating Specificity and Heterogeneity in Multimodal Biomedicine Data
Fa Zhang 0001, Bin Hu 0001 |
IEEE Trans. Comput. Soc. Syst. | 1 |
| 2025 | Harmonic Wavelet Neural Network for Discovering Neuropathological Propagation Patterns in Alzheimer's DiseaseabstractEmerging researchindicates that the degenerative biomarkers associated with Alzheimer's disease (AD) exhibit a non-random distribution within the cerebral cortex, instead following the structural brain network. The alterations in brain networks occur much earlier than the onset of clinical symptoms, thereby affecting the progression of brain disease. In this context, the utilization of computational methods to ascertain the propagation patterns of neuropathological events would contribute to the comprehension of the pathophysiological mechanism involved in the evolution of AD. Despite the encouraging findings achieved by existing graph-based deep learning approaches in analyzing irregular graph data, their applications in identifying the spreading pathway of neuropathology are limited due to two disadvantages. They include (1) lack of a common brain network as an unbiased reference basis for group comparison, and (2) lack of an appropriate mechanism for the identification of propagation patterns. To this end, we propose a proof-of-concept harmonic wavelet neural network (HWNN) to predict the early stage of AD and localize disease-related significant wavelets, which can be used to characterize the spreading pathways of neuropathological events across the brain network. The extensive experiments constructed on both synthetic and real datasets demonstrate that our proposed method achieves superior performance in classification accuracy and statistical power of identifying propagation patterns, compared with other representative approaches. Hongmin Cai, Ranran Deng, Defu Yang, Fa Zhang 0001, Guorong Wu 0001, Jiazhou Chen 0001 |
IEEE J. Biomed. Health Informatics | 4 |
| 2024 | SelfAlign: Achieving Subtomogram Alignment with Self-Supervised Deep LearningabstractCryo-Electron Tomography (Cryo-ET) and subtomogram averaging (STA) have been instrumental in advancing the analysis of high-resolution structural biology, enabling detailed insights into macromolecular complexes. However, due to limitations in sample thickness and electronic metrology, there are inherent issues with missing wedge artifacts and low signal-to-noise ratio in Cryo-ET. Researchers use STA to align and average subtomograms to address these two issues. Traditional STA methods, reliant on cross-correlation, are computationally expensive and not scalable for large datasets. The emerging method of using deep learning for STA has low accuracy and unstable performance at low signal-to-noise ratios. To address these issues, we proposed SelfAlign, a self-supervised deep learning approach for subtomogram alignment. To improve alignment accuracy, we introduce a rotation and translation method effectively reducing translation errors. Further, we present a self-labeling mechanism optimized for end-to-end processes,thereby abolishing the need for manual labeling. Additionally, we design a concise and efficient loss function to uphold stable training in scenarios with low signal-to-noise ratios. We demonstrate the efficacy of SelfAlign using four datasets, showcasing its superior performance in terms of alignment accuracy compared to existing methods. SelfAlign offers a robust and scalable solution for subtomogram analysis. Xuan Wang 0002, Haofan Cao, Fa Zhang 0001 |
BIBM | 5 |
| 2024 | Advancing Template-Based Flexible Docking of P450-Heme Complexes via JAX MDabstractAlphaFold2 has advanced protein structure prediction but often overlooks complexes with essential ligands, crucial for understanding protein-ligand interactions relevant to drug design. AlphaFill was developed to integrate ligands and ions from experimental structures into predicted models. However, many resulting complexes lack precise spatial constraints, leading to clashes between filled heme atoms and the P450 structure, as well as discrepancies in the key Fe-S bond.So, we propose FlexFill, a rapid flexible docking algorithm based on JAX MD. Specifically, FlexFill takes the predicted P450 structure as input to generate the P450-heme complex structure. During the docking process, we leverage JAX MD to account for the structural flexibility of P450. Unlike traditional molecular dynamics methods, we developed a customized energy function that focuses on the atoms near the pocket, significantly reducing runtime while satisfying spatial constraints.The results indicate that AlphaFill experienced almost 50% structural crashes, while FlexFill exhibited a substantially lower rate at less than 5%. Furthermore, FlexFill is 25% faster than AlphaFill, with further speed improvements by removing redundant entries from the template library. These observations suggest that FlexFill is capable of generating high-fidelity P450heme complexes and has the ability to flexibly dock a diverse array of protein-ligand complexes. Code and pre-trained model are released at https://github.com/xfcui/FlexFill. Rongxu Guo, Jiaheng Sang, Guishan Cui, Fa Zhang 0001, Xuefeng Cui, Shengying Li |
BIBM | 5 |
| 2024 | Transformer-Based Multi-Scale Fusion for Robust Predicting Microsatellite Instability from Pathological ImagesabstractMicrosatellite instability (MSI) is a crucial biomarker for guiding the efficacy of immunotherapy and adjuvant chemotherapy, making its detection essential for effective cancer treatment and prognosis. Traditional MSI prediction methods encounter challenges including high costs and limited accuracy under low tumor purity conditions. Recent advancements have explored deep learning for MSI prediction from pathological images, yet these approaches often overlook the multi-scale nature of pathological images and specific pathological features critical for MSI diagnosis. In this study, we proposed MSIscope, a novel Transformer-based method for detecting MSI from pathological images by fusing multi-scale pathological image information. Our approach consists of three key components: 1) ROI selection: we design a region of interest (ROI) selector based on convolutional neural networks and attention mechanisms, selecting tumor regions and important non-tumor regions as our focus; 2) Multi-scale vision expansion and feature extraction: we develop an algorithm that captures a broader view centered on a specified area to obtain a multi-scale field of view. The CTransPath feature extractor is then used to extract features from the image; 3) Multi-scale fusion Transformer: we propose a multi-scale feature aggregator (MS-Transformer) to aggregate contextual features across regions and scales. Our method was experimentally validated on public datasets, achieving an AU-ROC of 0.911 on the TCGA pan-cancer dataset and 0.887 on the TCGA-CRC dataset, surpassing existing methods. Additionally, it maintains high AUROC on datasets with lower tumor purity, outperforming current approaches. These results highlight the potential of MSIscope as an robust method for MSI prediction. Taiyuan Hu, Haijing Luan, Rui Yan 0009, Jifang Hu, Kaixing Yang, Xinyin Han, Weier Liu, Jiayin He, Xiaohong Duan, Fa Zhang 0001, Beifang Niu |
BIBM | 11 |
| 2024 | SS-SwinUnet: A Distillation Method of Swin Transformer for Superior Ocular Image SegmentationabstractPrecise segmentation of the pupil, iris, and sclera is critical for diagnosing and treating ocular diseases such as glaucoma, strabismus, and retinal disorders. However, the fine structural differences within the eye and the interference of complex backgrounds, especially with VR devices prone to reflections, tilts, distortions, and occlusions, present significant challenges. In this paper, we introduce SS-SwinUnet, a novel segmentation method that integrates Swin Transformer and knowledge distillation to achieve superior performance. Specifically, SS-SwinUnet balances feature transfer between the encoder and decoder, reducing redundancy and enhancing representation. Additionally, we incorporate a Boundary Difference over Union Loss to improve boundary segmentation accuracy. We also propose an eye modeling method that parameterizes segmentation results to optimize the semantic segmentation of ocular structures. We constructed the TongRenD dataset, comprising 400 VR-captured videos and 4,100 images, which, along with the TEyeD dataset, was used in our experiments. Results demonstrate that SS-SwinUnet significantly outperforms existing medical image segmentation methods across multiple datasets. Bowei Ma, Dehui Qiu, Ze Xiong, Yulong Hu, Liguo Deng, Huimei Yuan, Fa Zhang 0001 |
BIBM | 8 |
| 2024 | Innovative 3D CTF Correction Techniques for Cryo-ET of Individual ParticlesabstractCryo-electron tomography (Cryo-ET) and subtomogram averaging techniques are highly effective in revealing high-resolution molecular structures. In this technique, accurate Contrast Transfer Function (CTF) correction is essential, as it profoundly affects image quality. This impact arises from various factors, including electron beam characteristics, sample thickness, diffraction phenomena, and lens aberrations. Traditional two-dimensional CTF(2D-CTF) correction methods become inadequate when Cryo-ET is increasingly applied to thicker samples and multi-angle imaging scenarios, since variations in defocus along the tilt axis lead to resolution deterioration. To address this challenge, a novel methodology, termed the Focus Defocus Changes Contrast Transfer Function (FDC-CTF), is put forward as a new type of three-dimensional Contrast Transfer Function (3D-CTF) method. This method geometrically determines the degree of particle defocus based on their coordinates within tomographic slices, effectively tackling the defocus gradient issues induced by sample thickness and tilting. Furthermore, a new workflow is proposed where, after obtaining the particle coordinates, each particle undergoes individual 3D-CTF correction, thereby enhancing the overall accuracy and efficiency of the process. We use two datasets to demonstrate the effectiveness of our method, showcasing its superior performance in alignment accuracy compared to existing methods. It provides a powerful solution for tomographic reconstruction. Xuan Wang 0002, Fa Zhang 0001 |
BIBM | 5 |
| 2024 | IG-GRD: A Model Based on Disentangled Graph Representation Learning for Imaging Genetic Data Fusion
Fa Zhang 0001, Bin Hu 0001 |
ICIC (2) | 5 |
| 2024 | Central Feature Network Enables Accurate Detection of Both Small and Large Particles in Cryo-Electron Tomography
Yaoyu Wang, Fa Zhang 0001, Xuefeng Cui |
ISBRA (1) | 4 |
| 2024 | Serial Section Microscopy Image Inpainting Guided by Axial Optical FlowabstractVolume electron microscopy (vEM) is becoming a prominent technique in three-dimensional (3D) cellular visualization. vEM collects a series of two-dimensional (2D) images and reconstructs ultrastructures at the nanometer scale by rational axial interpolation between neighboring sections. However, section damage inevitably occurs in the sample preparation and imaging process, suffering from manual operational errors or occasional mechanical failures. The damaged regions present blurry and contaminated structure information, even local blank holes. Despite significant progress in single-image inpainting, it is still a great challenge to recover missing biological structures, that satisfy 3D structural continuity among sections. In this paper, we propose an optical flow-based serial section inpainting architecture to effectively combine the 3D structure information from neighboring sections and 2D image features from surrounding regions. We design a two-stage reference generation strategy to predict a rational and detailed intermediate state image from coarse to fine. Then, a GAN-based inpainting network is adopted to integrate all reference information and guide the restoration of missing structures, while ensuring consistent distribution of pixel values across the 2D image. Extensive experimental results well demonstrate the superiority of our method over existing inpainting tools. Our code is available at https://github.com/chengyr1999/FlowInpaint/. Yiran Cheng, Bintao He, Fa Zhang 0001, Renmin Han |
ACM Multimedia | 3 |
| 2024 | Realize Generative Yet Complete Latent Representation for Incomplete Multi-View LearningabstractIn multi-view environment, it would yield missing observations due to the limitation of the observation process. The most current representation learning methods struggle to explore complete information by lacking either cross-generative via simply filling in missing view data, or solidative via inferring a consistent representation among the existing views. To address this problem, we propose a deep generative model to learn a complete generative latent representation, namely Complete Multi-view Variational Auto-Encoders (CMVAE), which models the generation of the multiple views from a complete latent variable represented by a mixture of Gaussian distributions. Thus, the missing view can be fully characterized by the latent variables and is resolved by estimating its posterior distribution. Accordingly, a novel variational lower bound is introduced to integrate view-invariant information into posterior inference to enhance the solidative of the learned latent representation. The intrinsic correlations between views are mined to seek cross-view generality, and information leading to missing views is fused by view weights to reach solidity. Benchmark experimental results in clustering, classification, and cross-view image generation tasks demonstrate the superiority of CMVAE, while time complexity and parameter sensitivity analyses illustrate the efficiency and robustness. Additionally, application to bioinformatics data exemplifies its practical significance. Hongmin Cai, Weitian Huang, Sirui Yang, Siqi Ding, Yue Zhang 0045, Bin Hu 0001, Fa Zhang 0001, Yiu-Ming Cheung |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2024 | Guest Editorial Computational Mathematics Modeling in Cancer AnalysisabstractCancer is a complex disease that can affect any body part. One key feature of cancer is the rapid production of abnormal cells that grow beyond their usual borders and can invade adjoining parts of the body and spread/metastasized to other organs. The process of metastasis is the crucial cause of cancer death. Environmental factors are a significant contributor to cancer initiation [1]. Numerous studies investigate various aspects of cancer, including pathogenesis, prevention, diagnosis, and treatment methods, with the goal of improving patient quality of life and increasing survival rates. Despite significant advances in the field, cancer continues to represent a global challenge for prevention and treatment. Wenjian Qin, Tianming Liu 0001, Fa Zhang 0001 |
IEEE J. Biomed. Health Informatics | 3 |
| 2024 | Sparse and Hierarchical Transformer for Survival Analysis on Whole Slide ImagesabstractThe Transformer-based methods provide a good opportunity for modeling the global context of gigapixel whole slide image (WSI), however, there are still two main problems in applying Transformer to WSI-based survival analysis task. First, the training data for survival analysis is limited, which makes the model prone to overfitting. This problem is even worse for Transformer-based models which require large-scale data to train. Second, WSI is of extremely high resolution (up to 150,000 x 150,000 pixels) and is typically organized as a multi-resolution pyramid. Vanilla Transformer cannot model the hierarchical structure of WSI (such as patch cluster-level relationships), which makes it incapable of learning hierarchical WSI representation. To address these problems, in this paper, we propose a novel Sparse and Hierarchical Transformer (SH-Transformer) for survival analysis. Specifically, we introduce sparse self-attention to alleviate the overfitting problem, and propose a hierarchical Transformer structure to learn the hierarchical WSI representation. Experimental results based on three WSI datasets show that the proposed framework outperforms the state-of-the-art methods. Rui Yan 0009, Zhilong Lv, Zhidong Yang, Senlin Lin, Chun-Hou Zheng 0001, Fa Zhang 0001 |
IEEE J. Biomed. Health Informatics | 6 |
| 2023 | A Tongue Feature Extraction Method Based on a Sublingual Vein SegmentationabstractSublingual vein features including swelling, varicose and cyanosis are essential for the symptoms differentiation and treatment selection in Traditional Chinese Medicine (TCM) tongue diagnosis, especially reflecting the state of human blood circulation. However, automatic and accurate extraction of sublingual vein features remains a great challenge, limited by both the lack of datasets for sublingual images and the influence of noise from non-tongue and non-sublingual vein components. In this paper, we propose a novel tongue features extraction method based on segmenting the sublingual vein instead of the whole tongue bottom, in which a sublingual vein segmentation framework based on a Polyp-PVT network is developed to eliminate the noise from the surrounding part of the sublingual vein. Meanwhile, we first adopt a transformer-based method such as Swin-Transformer network to extract sublingual vein features by virtue of the awesome capability of the transformer network. In addition, we construct a large dataset including 4018 sublingual vein images for the segmentation and classification of sublingual veins. Experimental results have shown that the tongue feature extraction method combined with a sublingual vein segmentation can greatly outperform the existing tongue feature extracting methods. Yulong Hu, Dehui Qiu, Fa Zhang 0001, Bin Hu 0001 |
BIBM | 6 |
| 2023 | Integrative Drug Discovery Platform: A Modular Approach for Efficient and Automated Virtual ScreeningabstractThis paper presents a drug development platform based on virtual screening technology. The platform integrates key components such as pocket prediction, molecular docking, molecular dynamics simulation, and ADMET evaluation to achieve an efficient and automated drug virtual screening process. The platform utilizes Docker for modular encapsulation, ensuring environment isolation and convenient deployment. It also provides standardized input-output formats and a task allocation system, enabling users to quickly deploy and customize the workflow. Experimental results demonstrate the effectiveness of the platform in identifying real drugs and evaluating virtual screening results, providing an efficient and reliable solution for drug development. The platform features easy deployment and migration, independent module execution, automated workflow implementation, personalized customization and replacement, task allocation for computationally intensive steps, and complex operations in molecular dynamics simulation. Lulu Xie, Zhonghai Zhang, Bo Duan, Gang Niu 0008, Shiwei Sun, Fa Zhang 0001, Runting Zhang, Guangming Tan |
BIBM | 8 |
| 2023 | Variational Clustering and Denoising of Spatial TranscriptomicsabstractSpatial transcriptomics data provides a unique opportunity to investigate both gene expression and spatial structure in tissues at the same time. However, incorporating spatial information to accurately identify spatial domains is difficult due to factors such as high-dimensionality, sparsity, noise, and dropout events. To address these issues, we introduce vGraphST, a novel graph-based deep learning approach tailored for spatial transcriptomics data. Our method combines auto-encoder and contrastive learning techniques to process high-dimensional data and generate meaningful low-dimensional embeddings. Additionally, we use continuous distributions instead of discrete values in both the latent space and the denoised gene expression space. Specifically, Gaussian distributions are used to model the latent space, while zero-inflated Poisson distributions are used to model the denoised gene expression space. Experimental results demonstrate the effectiveness of vGraphST in accurately representing and analyzing spatial transcriptomics data. When compared to other methods using the DLPFC dataset, vGraphST achieves an average Adjusted Rand Index (ARI) of 0.58, demonstrating its superiority in segmenting spatial domains and recognizing biologically relevant spatiotemporal patterns. Cuiyuan Li, Fa Zhang 0001, Kai Hu 0002, Xuefeng Cui |
BIBM | 2 |
| 2023 | GPU Optimization of Biological Macromolecule Multi-tilt Electron Tomography Reconstruction Algorithm
Zi-Ang Fu, Fa Zhang 0001 |
ICIC (3) | 3 |
| 2023 | A Novel Impervious Surface Extraction Method Based on TransformerabstractThe amount of impervious surface is an important indicator to measure the degree of urbanization and the urban ecological environment. However, the objects in the low-density impervious surface areas are small and scattered, which are easily confused with the background. Therefore, the extraction of the small and scattered impervious surfaces is still challenging. In this study, we propose a dual-branch network combing transformer and CNN with attention mechanism. In this model, transformer branch is first used to extract impervious surface to capture long-distance and large-scale dependencies. In addition, another UNet branch embedded the coordinate attention mechanism can capture detailed information and meanwhile reduce information redundancy. Experiments show that our proposed method performs better than the traditional CNN methods. Dehui Qiu, Fa Zhang 0001, Huimei Yuan |
IGARSS | 4 |
| 2023 | SaID: Simulation-Aware Image Denoising Pre-trained Model for Cryo-EM Micrographs
Zhidong Yang, Hongjia Li 0001, Dawei Zang, Renmin Han, Fa Zhang 0001 |
ISBRA | 5 |
| 2023 | Cooperation of local features and global representations by a dual-branch network for transcription factor binding sites predictionabstractInteractions between DNA and transcription factors (TFs) play an essential role in understanding transcriptional regulation mechanisms and gene expression. Due to the large accumulation of training data and low expense, deep learning methods have shown huge potential in determining the specificity of TFs-DNA interactions. Convolutional network-based and self-attention network-based methods have been proposed for transcription factor binding sites (TFBSs) prediction. Convolutional operations are efficient to extract local features but easy to ignore global information, while self-attention mechanisms are expert in capturing long-distance dependencies but difficult to pay attention to local feature details. To discover comprehensive features for a given sequence as far as possible, we propose a Dual-branch model combining Self-Attention and Convolution, dubbed as DSAC, which fuses local features and global representations in an interactive way. In terms of features, convolution and self-attention contribute to feature extraction collaboratively, enhancing the representation learning. In terms of structure, a lightweight but efficient architecture of network is designed for the prediction, in particular, the dual-branch structure makes the convolution and the self-attention mechanism can be fully utilized to improve the predictive ability of our model. The experiment results on 165 ChIP-seq datasets show that DSAC obviously outperforms other five deep learning based methods and demonstrate that our model can effectively predict TFBSs based on sequence feature alone. The source code of DSAC is available at https://github.com/YuBinLab-QUST/DSAC/. Yutong Yu, Pengju Ding, Hongli Gao, Guozhu Liu, Fa Zhang 0001, Bin Yu 0007 |
Briefings Bioinform. | 5 |
| 2023 | A strategy combining denoising and cryo-EM single particle analysisabstractIn cryogenic electron microscopy (cryo-EM) single particle analysis (SPA), high-resolution three-dimensional structures of biological macromolecules are determined by iteratively aligning and averaging a large number of two-dimensional projections of molecules. Since the correlation measures are sensitive to the signal-to-noise ratio, various parameter estimation steps in SPA will be disturbed by the high-intensity noise in cryo-EM. However, denoising algorithms tend to damage high frequencies and suppress mid- and high-frequency contrast of micrographs, which exactly the precise parameter estimation relies on, therefore, limiting their application in SPA. In this study, we suggest combining a cryo-EM image processing pipeline with denoising and maximizing the signal's contribution in various parameter estimation steps. To solve the inherent flaws of denoising algorithms, we design an algorithm named MScale to correct the amplitude distortion caused by denoising and propose a new orientation determination strategy to compensate for the high-frequency loss. In the experiments on several real datasets, the denoised particles are successfully applied in the class assignment estimation and orientation determination tasks, ultimately enhancing the quality of biomacromolecule reconstruction. The case study on classification indicates that our strategy not only improves the resolution of difficult classes (up to 5 Å) but also resolves an additional class. In the case study on orientation determination, our strategy improves the resolution of the final reconstructed density map by 0.34 Å compared with conventional strategy. The code is available at https://github.com/zhanghui186/Mscale. Hongjia Li 0001, Fa Zhang 0001 |
Briefings Bioinform. | 3 |
| 2023 | TransSurv: Transformer-Based Survival Analysis Model Integrating Histopathological Images and Genomic Data for Colorectal CancerabstractSurvival analysis is a significant study in cancer prognosis, and the multi-modal data, including histopathological images, genomic data, and clinical information, provides unprecedented opportunities for its development. However, because of the high dimensionality and the heterogeneity of histopathological images and genomic data, acquiring effective predictive characters from these multi-modal data has always been a challenge for survival analysis. In this article, we propose a transformer-based survival analysis model (TransSurv) for colorectal cancer that can effectively integrate intra-modality and inter-modality features of histopathological images, genomic data, and clinical information. Specifically, to integrate the intra-modality relationship of image patches, we develop a multi-scale histopathological features fusion transformer (MS-Trans). Furthermore, we provide a cross-modal fusion transformer based on cross attention for multi-scale pathological representation and multi-omics representation, which includes RNA-seq expression and copy number alteration (CNA). At the output layer of the TransSurv, we adopt the Cox layer to integrate multi-modal fusion representation with clinical information for end-to-end survival analysis. The experimental results on the Cancer Genome Atlas (TCGA) colorectal cancer cohort demonstrate that the proposed TransSurv outperforms the existing methods and improves the prognosis prediction of colorectal cancer. Zhilong Lv, Yuexiao Lin, Rui Yan 0009, Ying Wang 0043, Fa Zhang 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 5 |
| 2022 | Multi-site MRI classification using Weighted federated learning based on Mixture of Experts domain adaptationabstractDeep learning often requires large amounts of data from different institutions. Federated learning, as a distributed training framework, enables multiple participants to collaboratively train models without collecting data together and hence protecting data privacy, but the datasets from different institutions usually bring the problem of domain shift, which affects the performance of the model. When addressing domain shift, previous works often use a single global model to share parameters. Therefore, we propose a novel method to train multiple public models with different structures under the federated framework to improve the reliability and robustness of the public models. And each participant keeps its own domain-tuned private model, the private model does not share parameters with other participants. We use Mixture of Experts (MoE) domain adaptation to dynamically combine different public models and private model, which utilizes the similarity between different datasets to update the parameters of the public models. We apply the proposed method to the multi-site Magnetic resonance imaging (MRI) end-to-end classification, and the experiments demonstrate its effectiveness. Tian Bai 0002, Yingfang Zhang, Yuzhao Wang, Yanguo Qin, Fa Zhang 0001 |
BIBM | 5 |
| 2022 | Hexagonal Convolutional Neural Network for Spatial Transcriptomics ClassificationabstractRecent advances in spatial transcriptomics have enabled the comprehensive measurement of transcriptional profiles while retaining the spatial contextual information. Identifying spatial domains is a critical step in the analysis of spatially resolved transcriptomics. Existing unsupervised methods perform poorly on this task owing to the large amount of noise and dropout events in the transcriptomic profiles. To address this problem, we first extend an unsupervised algorithm to a supervised learning method that can identify useful features and reduce noise hindrance. Second, inspired by the classical convolution in convolutional neural networks (CNNs), we designed a regular hexagonal convolution to compensate for the missing gene expression patterns from adjacent nodes. Compared with the graph convolution in graph neural networks (GNNs), our hexagonal convolution can preserve the relative spatial location information of different nodes in graph-structured data. Third, based on the hexagonal convolution, a novel hexagonal Convolutional Neural Network (hexCNN) is proposed for spatial transcriptomics classification. Finally, we compared the proposed hexCNN with existing methods on the DLPFC dataset. The results show that hexCNN achieves a classification accuracy of 87.2% and an average Rand index (ARI) of 78.2% (1.9% and 3.3% higher than those of GNNs). Fa Zhang 0001, Kai Hu 0002, Xuefeng Cui |
BIBM | 2 |
| 2022 | A Segmentation-aware Synergy Network for Single Particle Recognition in Cryo-EMabstractCryo-electron microscopy (cryo-EM) single particle analysis (SPA) has been an indispensable technology to reconstruct three-dimensional (3D) structures of biomolecules at near-atomic resolution. Tens of thousands of particles are required to obtain high-resolution 3D reconstructions, nevertheless, it is rather challenging due to the extremely noisy microscopy images and the diversity of particles. Recently, while deep learning-based methods have been devoted into the improvement of particle feature extraction and location estimation, most of them are plagued with vulnerable feature representation, inexact supervised ground truth. Furthermore, these DL-methods usually adopt denoising and particle picking as two-stage operations in the existing pipeline, which is inadequate to achieve accurate estimation for location. In this paper, we propose a segmentation-aware synergy framework to automatically select particles in which two tightly-coupled networks are designed including a multiple output convolution subnet for denoise to jointly learn strong object representation and pixel representation simultaneously and a deep convolution subnet for particle location. Furthermore, joint learning of the two networks can effectively enhance the synergy relationship between denoising and downstream recognition, thus leading to accurate and reliable location estimations for SPA. When applied with various EMPAIR real-world datasets, our model improves the performance of particle detection and exaction, especially intersection over union metric, and this strength has important implications for the next 2D alignment, 2D classification averaging, and high-resolution 3D refinement steps in SPA. Hongjia Li 0001, Chi Zhang 0110, Fa Zhang 0001 |
BIBM | 4 |
| 2022 | Temporal Correlation Network for Video Polyp SegmentationabstractAccurate polyp segmentation from colonoscopy images is essential for identifying colorectal cancer. Recently, segmentation methods based on convolutional neural networks and transformers have represented excellent performance for image polyp segmentation. However, these methods are mostly designed for individual images rather than the entire video datasets, which results in the absence of sequential relationships among lesion images and neglects the significant intrinsic property of continuous video. In this work, we propose a temporal correlation network (TC-Net) for video polyp segmentation. In TC-Net, the temporal correlation is unprecedentedly modeled based on the relationship between the original video and the captured frames to be adaptable for video polyp segmentation, and the network is also calibrated for the corresponding time correlation output. Furthermore, we design a dual-track learning strategy for the optimization method in TC-Net to ensure the independence of TC-Net during the learning process to adequately exploit the optimization effect of temporal correlation. The network’s effectiveness is demonstrated by extensive experiments on five publicly available biomedical datasets, and TC-Net achieves state-of-the-art (SOTA) performance. Dehui Qiu, Senlin Lin, Sheng Shi, Shengtao Zhu, Fa Zhang 0001 |
BIBM | 7 |
| 2022 | Method for Preoperative Prediction of Microvascular Invasion of Hepatocellular CarcinomaabstractMicrovascular invasion (MVI) in hepatocellular carcinoma (HCC) is of great guiding significance for the formulating treatment strategies and accessing the prognosis before the surgery. However, in traditional medicine, the gold standard for the diagnosis of MVI is obtained by examining pathological images which can only be obtained by sampling and sectioning tumors after surgery. At this time, MVI results have lost the timeliness of guiding tumor resection surgery. In order to solve this problem, existing studies began to use deep learning-based methods for preoperative prediction of MVI using non-invasive imaging. Most of these methods adopt the fusion methods of multi-sequence images to predict MVI, but fail to make full use of the characteristics of multiply sequences as prior knowledge to combine into the model, resulting in no further improvement of prediction performance. So we propose a multi-sequence image difference and correlation deep learning model. The model can extract the difference and correlation information between sequences from different scales and combine them into the model. To validate proposed model, we collected a data set consists of 120 HCC patients, including 50 MVI-positive patients. Compared with existing studies, our method has greatly improved in all evaluation metrics. Tian Bai 0002, Tongjia Chu, Fa Zhang 0001 |
BIBM | 5 |
| 2022 | Predicting Drug-Disease Associations by Self-topological Generalized Matrix Factorization with Neighborhood Constraints
Zonglan Zuo, Rui Yan 0009, Chun-Hou Zheng 0001, Fa Zhang 0001 |
ICIC (2) | 6 |
| 2022 | CSAM: A Channel and Spatial Attention Mechanism for Impervious Surface Extraction in Difficult AreasabstractImpervious surface extraction from remote sensing images has become a promising technology to measure the urban ecological environment and monitor human activity. However, due to the complex characteristics of impervious landscapes, most researches on impervious surface extraction hardly identify the scattered and small objects especially in difficult areas, which severely affect the accuracy of mapping impervious surface. In this work, we propose a channel and spatial attention mechanism (CSAM) to extract impervious surface in difficult areas, which includes a channel attention module to learn the relationship in the multi-channel remote sensing images and a spatial attention module to capture the features of the inconspicuous objects. Experiments with the Sentinel-2 dataset in South Africa demonstrate that CSAM can outperform the state-of-the-art methods. Fangyuan Zhao, Zhongchang Sun, Dehui Qiu, Fa Zhang 0001, Xinyu Liu 0008, Guangming Tan |
IGARSS | 7 |
| 2022 | Joint Region-Attention and Multi-scale Transformer for Microsatellite Instability Detection from Whole Slide Images in Gastrointestinal Cancer
Zhilong Lv, Rui Yan 0009, Yuexiao Lin, Ying Wang 0043, Fa Zhang 0001 |
MICCAI (2) | 5 |
| 2022 | Correction of image distortion in large-field ssEM stitching by an unsupervised intermediate-space solving networkabstractMOTIVATION: Serial-section electron microscopy (ssEM) is a powerful technique for cellular visualization, especially for large-scale specimens. Limited by the field of view, a megapixel image of whole-specimen is regularly captured by stitching several overlapping images. However, suffering from distortion by manual operations, lens distortion or electron impact, simple rigid transformations are not adequate for perfect mosaic generation. Non-linear deformation usually causes 'ghosting' phenomenon, especially with high magnification. To date, existing microscope image processing tools provide mature rigid stitching methods but have no idea with local distortion correction. RESULTS: In this article, following the development of unsupervised deep learning, we present a multi-scale network to predict the dense deformation fields of image pairs in ssEM and blend these images into a clear and seamless montage. The model is composed of two pyramidal backbones, sharing parameters and interacting with a set of registration modules, in which the pyramidal architecture could effectively capture large deformation according to multi-scale decomposition. A novel 'intermediate-space solving' paradigm is adopted in our model to treat inputted images equally and ensure nearly perfect stitching of the overlapping regions. Combining with the existing rigid transformation method, our model further improves the accuracy of sequential image stitching. Extensive experimental results well demonstrate the superiority of our method over the other traditional methods. AVAILABILITY AND IMPLEMENTATION: The code is available at https://github.com/HeracleBT/ssEM_stitching. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Bintao He, Fa Zhang 0001, Renmin Han |
Bioinform. | 3 |
| 2022 | Noise-Transfer2Clean: denoising cryo-EM images based on noise modeling and transferabstractMOTIVATION: Cryo-electron microscopy (cryo-EM) is a widely used technology for ultrastructure determination, which constructs the 3D structures of protein and macromolecular complex from a set of 2D micrographs. However, limited by the electron beam dose, the micrographs in cryo-EM generally suffer from the extremely low signal-to-noise ratio (SNR), which hampers the efficiency and effectiveness of downstream analysis. Especially, the noise in cryo-EM is not simple additive or multiplicative noise whose statistical characteristics are quite different from the ones in natural image, extremely shackling the performance of conventional denoising methods. RESULTS: Here, we introduce the Noise-Transfer2Clean (NT2C), a denoising deep neural network (DNN) for cryo-EM to enhance image contrast and restore specimen signal, whose main idea is to improve the denoising performance by correctly learning the noise distribution of cryo-EM images and transferring the statistical nature of noise into the denoiser. Especially, to cope with the complex noise model in cryo-EM, we design a contrast-guided noise and signal re-weighted algorithm to achieve clean-noisy data synthesis and data augmentation, making our method authentically achieve signal restoration based on noise's true properties. Our work verifies the feasibility of denoising based on mining the complex cryo-EM noise patterns directly from the noise patches. Comprehensive experimental results on simulated datasets and real datasets show that NT2C achieved a notable improvement in image denoising, especially in background noise removal, compared with the commonly used methods. Moreover, a case study on the real dataset demonstrates that NT2C can greatly alleviate the obstacles caused by the SNR to particle picking and simplify the identifying of particles. AVAILABILITYAND IMPLEMENTATION: The code is available at https://github.com/Lihongjia-ict/NoiseTransfer2Clean/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Hongjia Li 0001, Zhidong Yang, Chengmin Li, Jintao Li 0001, Renmin Han, Fa Zhang 0001 |
Bioinform. | 9 |
| 2022 | Macromolecules Structural Classification With a 3D Dilated Dense Network in Cryo-Electron TomographyabstractCryo-electron tomography, combined with subtomogram averaging (STA), can reveal three-dimensional (3D) macromolecule structures in the near-native state from cells and other biological samples. In STA, to get a high-resolution 3D view of macromolecule structures, diverse macromolecules captured by the cellular tomograms need to be accurately classified. However, due to the poor signal-to-noise-ratio (SNR) and severe ray artifacts in the tomogram, it remains a major challenge to classify macromolecules with high accuracy. In this paper, we propose a new convolutional neural network, named 3D-Dilated-DenseNet, to improve the performance of macromolecule classification. In 3D-Dilated-DenseNet, there are two key strategies to guarantee macromolecule classification accuracy: 1) Using dense connections to enhance feature map utilization (corresponding to the baseline 3D-C-DenseNet); 2) Adopting dilated convolution to enrich multi-level information in feature maps. We tested 3D-Dilated-DenseNet and 3D-C-DenseNet both on synthetic data and experimental data. The results show that, on synthetic data, compared with the state-of-the-art method in the SHREC contest (SHREC-CNN), both 3D-C-DenseNet and 3D-Dilated-DenseNet outperform SHREC-CNN. In particular, 3D-Dilated-DenseNet improves 0.393 of F1 metric on tiny-size macromolecules and 0.213 on small-size macromolecules. On experimental data, compared with 3D-C-DenseNet, 3D-Dilated-DenseNet can increase classification performance by 2.1 percent. Renmin Han, Zhiyong Liu 0002, Min Xu 0009, Fa Zhang 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 6 |
| 2021 | Moment Invariants with Data Augmentation for Tongue Image SegmentationabstractTongue diagnosis plays an essential role in diagnosing the syndrome types, pathological types, lesion location and clinical stages of cancers in Traditional Chinese Medicine (TCM). The quality of the tongue image datasets is crucial to tongue image segmentation in modern tongue diagnosis. However, the tongue image dataset maintains a challenging problem because of the lack of datasets for the sublingual image, the complexity and scarcity of the tongue images. In this paper, we propose an effective segmentation framework for tongue images, called Moment Invariants with Data Augmentation (DAMI), which can be flexibly applied to different segmentation models. To overcome over-fitting caused by excessive data augmentation, a data augmentation module with random transformations is designed to achieve appropriate data augmentation. Meanwhile, we develop a new moment invariants module to optimize data augmentation in image segmentation. In addition, a novel tongue image dataset, Lingual-Sublingual Image Dataset (LSID), has been established for the classification and segmentation of tongue or sublingual veins. Experimental results confirm that the models using DAMI can remarkably outperform the existing methods on LID, SID and BioHit for the tongue image segmentation. Senlin Lin, Xuekun Song, Yingqing Lin, Fa Zhang 0001, Dehui Qiu, Yuling Zheng |
BIBM | 8 |
| 2021 | PG-TFNet: Transformer-based Fusion Network Integrating Pathological Images and Genomic Data for Cancer Survival AnalysisabstractSurvival analysis is crucial to the evaluation of cancer treatment options and deep learning-based methods integrating pathological images and genomic data have been used for prognosis prediction. However, the most methods are based on the analysis of pathological image patches, thus ignoring the morphological structure information at larger field-of-view and intrinsic relationships between patches. Meanwhile, the existing models fail to exploit the powerful representation learning capabilities of the neural networks for effective multimodal feature fusion of pathological images and genomic data. In this paper, we propose a novel transformer-based fusion network integrating pathological images and genomic data (PGTFNet) for cancer survival analysis. Specifically, we present a transformer-based feature fusion module for multi-scale pathological slides to fully exploit the intra-modality relationships between image patches at various fields of view. Moreover, in order to make effective inter-modality feature fusion of pathological images and genomic data, we introduce a cross-attention transformer module that can exchange feature representations of different modalities between two transformers branches. The PG-TFNet is performed on the colorectal cancer dataset from the Cancer Genome Atlas (TCGA), which contains paired whole-slide images and genomic data with ground truth survival data. The experimental results from a 10-fold cross validation demonstrate that the proposed PG-TFNet facilitates the prognosis prediction of colorectal cancer and shows superiority over the existing methods. Zhilong Lv, Yuexiao Lin, Rui Yan 0009, Zhenghe Yang, Ying Wang 0043, Fa Zhang 0001 |
BIBM | 6 |
| 2021 | TransPicker: a Transformer-based Framework for Particle Picking in cryoEM MicrographsabstractSingle-particle cryo-electron microscopy (cryoEM) methods are powerful for solving high-resolution structures of biological macromolecules. Locating numerous particles from micrographs is essential for three-dimensional reconstruction but challenging due to the extremely low signal-to-noise ratio and various particle shapes in micrographs. In this study, we devise the TransPicker, a two-dimensional particle picking framework based on a novel end-to-end transformer-based detective method named crDETR (cryoEM DEtection TRansformer). crDETR applies an improved deformable Transformer to perform inference in parallel on particle relocation and global context, without hand-crafted components like anchors, non-maximum suppression procedure, or sliding windows. Also, it uses a combined loss function to guarantee fast convergence. It uses divide-and-conquer to overcome the limitations of the object query number. Moreover, we develop a series of optimizations, including denoising, enhancing, bad particle filtering, adding masks on carbon areas and ice contaminants to decrease the false-positive ratio and improve the accuracy of particle picking. Experimental results on various datasets demonstrate that TransPicker can select particles with more accuracy, especially in high noise compared with other methods. To our knowledge, TransPicker is the first application of the transformer technique in cryoEM particle picking. Chi Zhang 0110, Hongjia Li 0001, Zhenghe Yang, Jieqing Feng, Fa Zhang 0001 |
BIBM | 7 |
| 2021 | A Deep Reinforcement Learning-based Task Scheduling Algorithm for Energy Efficiency in Data CentersabstractCloud data centers provide end-users with a wide range of application scenarios, including scientific computing, smart grids, etc. The number and size of data centers have rapidly increased in recent years, which causes severe environmental problems and colossal power demand. Therefore, it is desirable to use a proper scheduling method to optimize resource usage and reduce energy consumption in a data center. However, it is rather difficult to design an effective and efficient task scheduling algorithm because of the dynamic and complex environment of data centers. This paper proposes a task scheduling algorithm, WSS, to optimize resource usage and reduce energy consumption based on a model-free deep reinforcement learning framework inspired by the Wolpertinger architecture. The proposed algorithm can handle the scheduling problem on a sizeable discrete action space, improve decision efficiency, and save the training convergence time. Meanwhile, the proposed algorithm based on Soft Actor-Critic is designed to improve the stability and exploration capability of WSS. Experiments based on real-world traces prove that WSS can reduce energy consumption by nearly 25% compared with the Deep Q-network task scheduling algorithm. Moreover, WSS can provide a short time of training convergence without increasing the average waiting time of tasks and achieve stable performance. Penglei Song, Ce Chi, Kaixuan Ji, Zhiyong Liu 0002, Fa Zhang 0001, Shikui Zhang, Dehui Qiu |
ICCCN | 5 |
| 2021 | A Hybrid Frequency-Spatial Domain Model for Sparse Image Reconstruction in Scanning Transmission Electron MicroscopyabstractScanning transmission electron microscopy (STEM) is a powerful technique in high-resolution atomic imaging of materials. Decreasing scanning time and reducing electron beam exposure with an acceptable signal-to-noise ratio are two popular research aspects when applying STEM to beam-sensitive materials. Specifically, partially sampling with fixed electron doses is one of the most important solutions, and then the lost information is restored by computational methods. Following successful applications of deep learning in image in-painting, we have developed an encoder-decoder network to reconstruct STEM images in extremely sparse sampling cases. In our model, we combine both local pixel information from convolution operators and global texture features, by applying specific filter operations on the frequency domain to acquire initial reconstruction and global structure prior. Our method can effectively restore texture structures and be robust in different sampling ratios with Poisson noise. A comprehensive study demonstrates that our method gains about 50% performance enhancement in comparison with the state-of-art methods. Code is available at https://github.com/icthrm/Sparse-Sampling-Reconstruction. Bintao He, Fa Zhang 0001, Huanshui Zhang, Renmin Han |
ICCV | 2 |
| 2021 | Self-Supervised Cryo-Electron Tomography Volumetric Image Restoration from Single Noisy Volume with Sparsity ConstraintabstractCryo-Electron Tomography (cryo-ET) is a powerful tool for 3D cellular visualization. Due to instrumental limitations, cryo-ET images and their volumetric reconstruction suffer from extremely low signal-to-noise ratio. In this paper, we propose a novel end-to-end self-supervised learning model, the Sparsity Constrained Network (SC-Net), to restore volumetric image from single noisy data in cryo-ET. The proposed method only requires a single noisy data as training input and no ground-truth is needed in the whole training procedure. A new target function is proposed to preserve both local smoothness and detailed structure. Additionally, a novel procedure for the simulation of electron tomographic photographing is designed to help the evaluation of methods. Experiments are done on three simulated data and four real-world data. The results show that our method could produce a strong enhancement for a single very noisy cryo-ET volumetric data, which is much better than the state-of-the-art Noise2Void, and with a competitive performance comparing with Noise2Noise. Code is available at https://github.com/icthrm/SC-Net. Zhidong Yang, Fa Zhang 0001, Renmin Han |
ICCV | 2 |
| 2021 | Decomposition-and-Fusion Network for HE-Stained Pathological Image Classification
Rui Yan 0009, Jintao Li 0001, Shaohua Kevin Zhou, Zhilong Lv, Xueyuan Zhang, Xiaosong Rao, Chun-Hou Zheng 0001, Fa Zhang 0001 |
ICIC (3) | 9 |
| 2021 | Predicting Drug-Disease Associations Based on Network Consistency Projection
Zonglan Zuo, Rui Yan 0009, Chun-Hou Zheng 0001, Fa Zhang 0001 |
ICIC (3) | 5 |
| 2021 | Spatio-Temporal Features Processing Network for Change Detection in Remote Sensing ImagesabstractChange detection is a significant remote sensing challenge, which can capture changes in land use and land cover. Recently, deep learning achieves great performance in change detection, but most of the existing methods based on deep learning only process local spatial relationships and single directional temporal relationships, which severely affect the accuracy of change detection. In this paper, we present a novel end-to-end spatio-temporal processing network (STPNet) for precise change detection in remote sensing images. In our network, we design a spatial processing module which can learn long-range relationship and rich features and a temporal processing module capturing bidirectional rich contextual information, respectively. Also, we combine the two modules into a building block named spatio-temporal processing module (STPM) which can be easily incorporated into the existing siamese architectures. Experiments with the WHU building change detection dataset demonstrate that STPNet can obtain better performance than state-of-the-art methods. Zhaobin Cao, Fa Zhang 0001, Guangming Tan |
IGARSS | 4 |
| 2021 | A Multi-GPU Design for Large Size Cryo-EM 3D ReconstructionabstractThree-dimensional (3D) reconstruction of cryo-electron microscopy (cryo-EM) is a powerful method to determine the structures of macromolecules at near-atomic resolution. Recently, larger size with finer resolution 2D images has been collected, which can improve the reconstruction resolution. However, large size data incurs high computation and huge memory overhead. Current implementations fail to perform the complete reconstruction workflow on a multi-GPU cluster for large size data. Because of no effective parallel method for 3D convolution and the huge memory demanding, large size data can not be efficiently reconstructed, which impede the resolution improving 3D reconstruction. To enable cryo-EM 3D reconstruction with large size data on multi-GPU, in this work, we propose a new parallel framework called OML-Relion. In OML-Relion, we first adopt a stride based Fourier transform and eliminate data dependence to parallelize the 3D convolution on multi-GPU. Considering the input size varying in each iteration, we next use an auto-tuning model to optimize 3D convolution performance. Finally, guaranteeing the whole reconstruction on a multi-GPU cluster for large size data, we design a novel lossless data compression algorithm to reduce memory overhead on each GPU further. The experiment shows that OML-Relion can efficiently handle large size cryo-EM 3D reconstruction on multi-GPU. The reconstruction module, including 3D convolution operation, achieves 225-330x times speedup for 200-800 pixel size particles. The compression algorithm significantly reduces memory overhead approaching 70%. Moreover, the whole workflow with OMLRelion can achieve 54-65x speedup compared with Relion using two large size datasets. Zhiyong Liu 0002, Qianshuo Fan, Fa Zhang 0001, Guangming Tan |
IPDPS | 5 |
| 2021 | PickerOptimizer: A Deep Learning-Based Particle Optimizer for Cryo-Electron Microscopy Particle-Picking Algorithms
Hongjia Li 0001, Jintao Li 0001, Fa Zhang 0001 |
ISBRA | 5 |
| 2021 | A novel constrained reconstruction model towards high-resolution subtomogram averagingabstractMOTIVATION: Electron tomography (ET) offers a unique capacity to image biological structures in situ. However, the resolution of ET reconstructed tomograms is not comparable to that of the single-particle cryo-EM. If many copies of the object of interest are present in the tomograms, their structures can be reconstructed in the tomogram, picked, aligned and averaged to increase the signal-to-noise ratio and improve the resolution, which is known as the subtomogram averaging. To date, the resolution improvement of the subtomogram averaging is still limited because each reconstructed subtomogram is of low reconstruction quality due to the missing wedge issue. RESULTS: In this article, we propose a novel computational model, the constrained reconstruction model (CRM), to better recover the information from the multiple subtomograms and compensate for the missing wedge issue in each of them. CRM is supposed to produce a refined reconstruction in the final turn of subtomogram averaging after alignment, instead of directly taking the average. We first formulate the averaging method and our CRM as linear systems, and prove that the solution space of CRM is no larger, and in practice much smaller, than that of the averaging method. We then propose a sparse Kaczmarz algorithm to solve the formulated CRM, and further extend the solution to the simultaneous algebraic reconstruction technique (SART). Experimental results demonstrate that CRM can significantly alleviate the missing wedge issue and improve the final reconstruction quality. In addition, our model is robust to the number of images in each tilt series, the tilt range and the noise level. AVAILABILITY AND IMPLEMENTATION: The codes of CRM-SIRT and CRM-SART are available at https://github.com/icthrm/CRM. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Renmin Han, Peng Yang 0010, Fa Zhang 0001, Xin Gao 0001 |
Bioinform. | 4 |
| 2021 | Multi-labelled proteins recognition for high-throughput microscopy images using deep convolutional neural networksabstractBACKGROUND: Proteins are of extremely vital importance in the human body, and no movement or activity can be performed without proteins. Currently, microscopy imaging technologies developed rapidly are employed to observe proteins in various cells and tissues. In addition, due to the complex and crowded cellular environments as well as various types and sizes of proteins, a considerable number of protein images are generated every day and cannot be classified manually. Therefore, an automatic and accurate method should be designed to properly solve and analyse protein images with mixed patterns. RESULTS: In this paper, we first propose a novel customized architecture with adaptive concatenate pooling and "buffering" layers in the classifier part, which could make the networks more adaptive to training and testing datasets, and develop a novel hard sampler at the end of our network to effectively mine the samples from small classes. Furthermore, a new loss is presented to handle the label imbalance based on the effectiveness of samples. In addition, in our method, several novel and effective optimization strategies are adopted to solve the difficult training-time optimization problem and further increase the accuracy by post-processing. CONCLUSION: Our methods outperformed the SOTA method of multi-labelled protein classification on the HPA dataset, GapNet-PL, by above 2% in the F1 score. Therefore, experimental results based on the test set split from the Human Protein Atlas dataset show that our methods have good performance in automatically classifying multi-class and multi-labelled high-throughput microscopy protein images. Boheng Zhang, Shaohan Hu, Fa Zhang 0001, Zhiyong Liu 0002 |
BMC Bioinform. | 4 |
| 2021 | Preface
Yi Pan 0001, De-Shuang Huang, Jianxin Wang 0001, Fa Zhang 0001 |
J. Comput. Sci. Technol. | 4 |
| 2021 | Improve the Resolution and Parallel Performance of the Three-Dimensional Refine Algorithm in RELION Using CUDA and MPIabstractIn cryo-electron microscopy, RELION is a powerful tool for high-resolution reconstruction. Due to the complicated imaging procedure and the heterogeneity of particles, some of the selected particle images offer more disturbing information than others. However, in the current RELION, all these particle images are treated equally. In our work, we extend RELION's model with one scalar parameter to score the contribution of a particle depending on the error between the experimental particle and the corresponding reprojection. This scores down weight potentially poor particles, hence accelerating the convergence. Besides, by now there is no sophisticated memory management system for RELION, fragmentation on GPU will increase with iterations, eventually crashing the program. In our work, we designed the stack-based memory management system to guarantee the stability of RELION and to optimize the memory usage condition. Also, to reduce memory usage, we developed a customized compressed data structure for the memory-demanding weight array. In addition, to speed up the GPU version of RELION, we proposed two highly efficient parallel algorithms for weight calculation algorithm and weight selection algorithm. Experiments show that compared with RELION, the optimized three-dimensional refine algorithm can speed up the converge procedure, the memory system can avoid memory fragmentation, and a better speed-up ratio can be obtained. Jingrong Zhang, Zhiyong Liu 0002, Fa Zhang 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 4 |
| 2021 | Classification-Based and Energy-Efficient Dynamic Task Scheduling Scheme for Virtualized Cloud Data CenterabstractThe size and number of cloud data centers (CDCs) have grown rapidly with the increasing popularity of cloud computing and high-performance computing. This has the unintended consequences of creating new challenges due to inefficient use of resources and high energy consumption. Hence, this necessitates the need to maximize resource utilization and ensure energy efficiency in CDCs. One viable approach to achieve energy efficiency and resource utilization in CDC is task scheduling. While several task scheduling approaches have been proposed in the literature, there appears to be a lack of classification-based merging concept for real-time tasks in these existing approaches. Thus, an energy-efficient dynamic scheduling scheme (EDS) of real-time tasks for virtualized CDC is presented in this paper. In the scheduling scheme, the heterogeneous tasks and virtual machines are first classified based on a historical scheduling record. Then, similar type of tasks are merged and scheduled to maximally utilize an operational state of the host. In addition, energy efficiencies and optimal operating frequencies of heterogeneous physical hosts are employed to attain energy preservation while creating and deleting the virtual machines. Experimental results show that, in comparison with existing techniques, EDS significantly improves overall scheduling performance, achieves a higher CDC resource utilization, increases task guarantee ratio, minimizes the mean response time, and reduces energy consumption. Avinab Marahatta, Sandeep Pirbhulal, Fa Zhang 0001, Reza M. Parizi, Kim-Kwang Raymond Choo, Zhiyong Liu 0002 |
IEEE Trans. Cloud Comput. | 3 |
| 2021 | PEFS: AI-Driven Prediction Based Energy-Aware Fault-Tolerant Scheduling Scheme for Cloud Data CenterabstractCloud data centers (CDCs) have become increasingly popular and widespread in recent years with the growing popularity of cloud computing and high-performance computing. Due to the multi-step computation of data streams and heterogeneous task dependencies, task failure frequently occurs, resulting in poor user experience and additional energy consumption. To reduce task execution failure as well as energy consumption, we propose a novel AI-driven energy-aware proactive fault-tolerant scheduling scheme for CDCs in this paper. First, a prediction model based on the machine learning approach is trained to classify the arriving tasks into “failure-prone tasks” and “non-failure-prone tasks” according to the predicted failure rate. Then, two efficient scheduling mechanisms are proposed to allocate two types of tasks to the most appropriate hosts in a CDC. The vector reconstruction method is developed to construct super tasks from failure-prone tasks and separately schedule these super tasks and non-failure-prone tasks to the most suitable physical host. All the tasks are scheduled in an earliest-deadline-first manner. Our evaluation results show that the proposed scheme can intelligently predict task failure and achieves better fault tolerance and reduces total energy consumption better than the existing schemes. Avinab Marahatta, Qin Xin 0001, Ce Chi, Fa Zhang 0001, Zhiyong Liu 0002 |
IEEE Trans. Sustain. Comput. | 4 |
| 2020 | LR-Net: A Multi-task Model Using Relationship-based Contour Information to Enhance the Semantic Segmentation of Cancer RegionsabstractThe segmentation of cancer regions is a key step in pathological image analysis. Although traditional methods (such as U-Net) have achieved good results in general medical image segmentation, the segmentation performance of the tumor region is still unsatisfactory because the boundary of the tumor is too blurred. Moreover, most tumor region segmentation methods focus on the learning of image content features while ignoring learning relationship among pixels on tumor contours. In this paper, we developed a multi-task learning technique to enhance the importance of contours and increase the weight of pixels relationship learning for the tumor segmentation. Different from the traditional single-decoder network, a parallel contour decoder with LRLM (location relationship learning module) is introduced as an auxiliary decoder to learn the relationship-based features of tumor contours, which forms a two-decoder network. To promote the information fusion of the two tasks, the two decoders share a same encoder with bidirectional skip connections between the auxiliary contour decoder and the main content decoder. Experimental results show that LR-Net is superior to many popular approaches, such as CE-Net and U-Net. Baorong Shi, Rui Yan 0009, Wang Jing, Jinfeng Zang, Fa Zhang 0001 |
BIBM | 6 |
| 2020 | NANet: Nuclei-Aware Network for Grading of Breast Cancer in HE Stained Pathological ImagesabstractAutomatic breast cancer grading methods based on HE stained pathological images can be summarized into two categories. The first category is to use learning-based methods to directly extract the features of the pathological image for breast cancer grading. However, unlike the coarse-grained problem of breast cancer classification, grading of breast Invasive Ductal Carcinoma (IDC) is a fine-grained classification problem. Only using general methods cannot classify IDC well. The second category is to conduct the three evaluation criteria of Nottingham Grading System (NGS) separately, and then integrate the results of the three criteria to obtain the final IDC grading result. However, NGS is only a semi-quantitative evaluation method. The inherent medical motivation of NGS is to grade IDC with the help of nuclei-related features. In this paper, we proposed a nuclei-aware network for IDC grading in pathological images. The entire network achieves an effect similar to the attention mechanism in end-to-end learning, so as to learn fine-grained and nuclei-related feature representations for IDC grading. It should to be pointed out that our method can emphasize custom areas, thus providing a way to model medical knowledge into the network structure. This is different from the general attention mechanism that cannot artificially control the area of attention. Experimental results show that the performance of proposed method is better than the state-of-the-art. Rui Yan 0009, Jintao Li 0001, Xiaosong Rao, Zhilong Lv, Chun-Hou Zheng 0001, Jinjin Dou, Fa Zhang 0001 |
BIBM | 9 |
| 2020 | An Energy Saving-Oriented Incentive Mechanism in Colocation Data CentersabstractThe size and amount of colocation data centers (colocations, for short) have been growing rapidly with the increasing popularity of cloud services, which ultimately causes a heavy burden on the power grid and the environment. However, even a colocation owner wishes to reduce its energy consumption, it may not be able to apply some effective energy saving techniques directly to the servers, since the servers belong to and are operated by its tenants. To solve the "uncoordinated relationship" issue between owners and tenants and achieve the energy reduction with a limited cost budget, an energy saving-oriented incentive mechanism, ESCo is proposed in this paper. Different from existing mechanisms that emphasize on minimizing the cost for the owners, our mechanism emphasizes on the maximization of the amount of energy saved by the tenants given a limited budget that the owner wants to pay for the energy saving. An algorithm is developed to realize a stable assignment in the mechanism. Trace-driven simulations based on real-world data are performed to verify the effectiveness of ESCo. The results show that ESCo can achieve 14.47% more energy saving than existing colocation incentive mechanisms. Ce Chi, Kaixuan Ji, Avinab Marahatta, Fa Zhang 0001, Youshi Wang, Zhiyong Liu 0002 |
ICCCN | 4 |
| 2020 | PStream: Priority-Based Stream Scheduling for Heterogeneous Paths in Multipath-QUICabstractWeb latency remains the main obstacle to improving user experience with the continuous development of the web. A lot of works have been made in this course. Quick UDP Internet Connection (QUIC) embeds stream multiplexing to solve the head-of-line blocking caused by the in-order requirement of TCP. Multipath-QUIC (MPQUIC) brings further improvements by utilizing multiple paths, as is done in MultiPath TCP (MPTCP). Different from MPTCP schedulers, MPQUIC schedulers are stream-aware and thus can provide finer granularity of multipath scheduling. As streams are with different features based on their contents, the resource preferences of a stream are highly related to its feature. We find that scheduling without the recognition of the stream features can aggravate inter-stream blocking when sharing paths. We fill this gap and propose PStream - a priority-based online stream scheduling mechanism for MPQUIC, which performs path scheduling based on the stream features. We examine the effectiveness of PStream under different path heterogeneity comparing to the original and the latest scheduler of MPQUIC. Our evaluation shows that our scheduler can reduce up to 25.4% of page load time in high path heterogeneity. Lin Wang 0015, Fa Zhang 0001, Biyu Zhou, Zhiyong Liu 0002 |
ICCCN | 3 |
| 2020 | An Integration Framework for Liver Cancer Subtype Classification and Survival Prediction Based on Multi-omics Data
Zhonglie Wang, Rui Yan 0009, Chun-Hou Zheng 0001, Fa Zhang 0001 |
ICIC (3) | 7 |
| 2020 | Dilated-DenseNet for Macromolecule Classification in Cryo-electron Tomography
Renmin Han, Xuefeng Cui, Zhiyong Liu 0002, Min Xu 0009, Fa Zhang 0001 |
ISBRA | 7 |
| 2020 | An Incentive Mechanism for Improving Energy Efficiency of Colocation Data Centers Based on Power PredictionabstractColocation data centers (colocations, for short) are developing rapidly in recent years, resulting in a heavy burden on the power grid and the environment. Due to the special management mode of colocations, even the colocation operators wish to reduce their power demand, they have no authority to control the servers because the servers belong to and are operated by the tenants themselves. To solve the "uncoordinated relationship" issue between operators and tenants, a truthful and feasible incentive mechanism MesPP is proposed in this paper. Different from existing works, MesPP aims at maximizing the energy reduction of colocations with a limited cost budget and can be applied even there is no demand response (DR) program. Meanwhile, power prediction is integrated into MesPP to further improve the energy efficiency of colocations and fairness of the mechanism. To solve the optimization problem, we develop a (1−ε)-approximation algorithm. Simulations are performed and show that MesPP can achieve 12.23% more energy saving compared with existing incentive mechanisms. Ce Chi, Kaixuan Ji, Avinab Marahatta, Fa Zhang 0001, Youshi Wang, Zhiyong Liu 0002 |
ISCC | 4 |
| 2020 | Compressed sensing improved iterative reconstruction-reprojection algorithm for electron tomographyabstractAbstract Background Electron tomography (ET) is an important technique for the study of complex biological structures and their functions. Electron tomography reconstructs the interior of a three-dimensional object from its projections at different orientations. However, due to the instrument limitation, the angular tilt range of the projections is limited within +70∘ to −70∘. The missing angle range is known as the missing wedge and will cause artifacts. Results In this paper, we proposed a novel algorithm, compressed sensing improved iterative reconstruction-reprojection (CSIIRR), which follows the schedule of improved iterative reconstruction-reprojection but further considers the sparsity of the biological ultra-structural content in specimen. The proposed algorithm keeps both the merits of the improved iterative reconstruction-reprojection (IIRR) and compressed sensing, resulting in an estimation of the electron tomography with faster execution speed and better reconstruction result. A comprehensive experiment has been carried out, in which CSIIRR was challenged on both simulated and real-world datasets as well as compared with a number of classical methods. The experimental results prove the effectiveness and efficiency of CSIIRR, and further show its advantages over the other methods. Conclusions The proposed algorithm has an obvious advance in the suppression of missing wedge effects and the restoration of missing information, which provides an option to the structural biologist for clear and accurate tomographic reconstruction. Renmin Han, Zhaotian Zhang, Tiande Guo, Zhiyong Liu 0002, Fa Zhang 0001 |
BMC Bioinform. | 6 |
| 2020 | SHREC 2020: Classification in cryo-electron tomograms
Ilja Gubins, Marten L. Chaillet, Gijs van der Schot, Remco C. Veltkamp, Friedrich Förster, Xiaohua Wan 0001, Xuefeng Cui, Fa Zhang 0001, Emmanuel Moebel, Xiao Wang 0004, Daisuke Kihara, Min Xu 0009, Nguyen P. Nguyen, Tommi A. White, Filiz Bunyak |
Comput. Graph. | 9 |
| 2019 | Cerebrovascular Segmentation Algorithm Based on Focused Multi-Gaussians Model and Weighted 3D Markov Random FieldabstractSegmenting the cerebral vessels precisely from the time-of-flight magnetic resonance angiography (TOF-MRA) images is important for the diagnosis and therapy of the cerebrovascular diseases. Since the complex structures of cerebral vessels, the current cerebrovascular segmentation algorithms based on statistical model have less accuracy for stenotic vessels and are quite time-consuming. In this paper, we propose a novel automatic cerebrovascular segmentation algorithm based on focused Multi-Gaussians (FMG) model and weighted 3D Markov Random Field. As far as our knowledge, this is the first time to adopt multi-Gaussians distributions as vascular model with the purpose of modeling the vascular tissue more accurately. Furthermore, the fitting range is narrowed to local region related to vessels in order to make the model focus on the vascular tissue and simplify the finite mixture model. To incorporate precise local character of images to the model, we design a new weighted 3D MRF by a weighted neighborhood system (W-NBS). Finally, the particle swarm optimization (PSO) algorithm of parameter estimation has been implemented parallelly based on GPUs and the execution speed was improved by about 70 times. The experimental results show that the algorithm can produce detailed segmentation results especially for stenotic vessels. Zhilong Lv, Rui Yan 0009, Xinyu Liu 0008, Zhongke Wu, Yicheng Zhu, Shiwei Sun, Fa Zhang 0001, Xingce Wang |
BIBM | 7 |
| 2019 | Predicting Tumor Mutational Burden from Liver Cancer Pathological Images Using Convolutional Neural NetworkabstractTumor mutational burden (TMB) is the most important and most promising biomarker in the era of tumor immunotherapy, and it can predict the immunotherapy efficiency of patients in various cancers including liver cancer. TMB is mainly obtained by next generation sequencing technology such as whole exome sequencing (WES). However, conditions such as excessive testing costs, lengthy detection cycles, and tissue sample dependence severely limit the clinical application of TMB. Inspired by the inner link between the intrinsic characteristics of the tumor cell genome and the pathological features of tumor cells and their microenvironment-related cells, we propose a deep learning method for predicting the level of TMB (high or low) directly from pathological images. This study found that the feature scale (receptive field) is the biggest factor affecting the classification effect of TMB prediction, and further determined the best receptive field through a series of experiments. Experimental results show that our method is far more out performance of the commonly used panel sequencing (99.7% VS 79.2%). To the best of our knowledge, this is the first research to predict TMB and the highest level of accuracy of genomic characteristic predicted by pathological images. The proposed method has the potential to provide immunotherapy to a much broader subset of patients with liver cancer. Fa Zhang 0001, Zhonglie Wang, Xiaosong Rao, Junbo Hao, Rui Yan 0009, Jiancheng Luo |
BIBM | 2 |
| 2019 | Integration of Multimodal Data for Breast Cancer Classification Using a Hybrid Deep Learning Method
Rui Yan 0009, Xiaosong Rao, Baorong Shi, Tiange Xiang, Chun-Hou Zheng 0001, Fa Zhang 0001 |
ICIC (1) | 10 |
| 2019 | Classifying Mixed Patterns of Proteins in High-Throughput Microscopy Images Using Deep Neural Networks
Boheng Zhang, Shaohan Hu, Fa Zhang 0001 |
ICIC (1) | 4 |
| 2019 | FStream: Flexible Stream Scheduling and Prioritizing in Multipath-QUICabstractWhile the web keeps evolving, web latency remains a major obstacle to improving user experience. In the past, many efforts have been made in this course. SPDY achieves reduced latency through multiplexing and prioritization by manipulating HTTP. Quick UDP Internet Connection (QUIC) generalizes the idea and embeds multiplexing in the transport layer by introducing application-oriented streams. Multipath-QUIC brings further improvements by utilizing multiple paths as is done in MultiPath TCP (MPTCP). However, failing to account for stream priorities in the transport layer can result in suboptimal performance for time-critical streams. We fill this gap and propose FStream - a flexible stream scheduling mechanism for Multipath-QUIC, which provides stream prioritization down to the transport layer. We implement FStream in Multipath-QUIC and demonstrate its effectiveness in reducing the completion time of time-critical streams ( 3x) through extensive experiments under different path dissimilarity conditions. Lin Wang 0015, Fa Zhang 0001, Zhiyong Liu 0002 |
ICPADS | 3 |
| 2019 | DM-SIRT: A Distributed Method for Multi-tilt Reconstruction in Electron Tomography
Jingrong Zhang, Zhiyong Liu 0002, Fa Zhang 0001 |
ISBRA | 6 |
| 2019 | A joint method for marker-free alignment of tilt series in electron tomographyabstractMOTIVATION: Electron tomography (ET) is a widely used technology for 3D macro-molecular structure reconstruction. To obtain a satisfiable tomogram reconstruction, several key processes are involved, one of which is the calibration of projection parameters of the tilt series. Although fiducial marker-based alignment for tilt series has been well studied, marker-free alignment remains a challenge, which requires identifying and tracking the identical objects (landmarks) through different projections. However, the tracking of these landmarks is usually affected by the pixel density (intensity) change caused by the geometry difference in different views. The tracked landmarks will be used to determine the projection parameters. Meanwhile, different projection parameters will also affect the localization of landmarks. Currently, there is no alignment method that takes interrelationship between the projection parameters and the landmarks. RESULTS: Here, we propose a novel, joint method for marker-free alignment of tilt series in ET, by utilizing the information underlying the interrelationship between the projection model and the landmarks. The proposed method is the first joint solution that combines the extrinsic (track-based) alignment and the intrinsic (intensity-based) alignment, in which the localization of landmarks and projection parameters keep refining each other until convergence. This iterative approach makes our solution robust to different initial parameters and extreme geometric changes, which ensures a better reconstruction for marker-free ET. Comprehensive experimental results on three real datasets show that our new method achieved a significant improvement in alignment accuracy and reconstruction quality, compared to the state-of-the-art methods. AVAILABILITY AND IMPLEMENTATION: The main program is available at https://github.com/icthrm/joint-marker-free-alignment. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Renmin Han, Zhipeng Bao, Tongxin Niu, Fa Zhang 0001, Min Xu 0009, Xin Gao 0001 |
Bioinform. | 5 |
| 2019 | AuTom-dualx: a toolkit for fully automatic fiducial marker-based alignment of dual-axis tilt series with simultaneous reconstructionabstractMotivation: Dual-axis electron tomography is an important 3 D macro-molecular structure reconstruction technology, which can reduce artifacts and suppress the effect of missing wedge. However, the fully automatic data process for dual-axis electron tomography still remains a challenge due to three difficulties: (i) how to track the mass of fiducial markers automatically; (ii) how to integrate the information from the two different tilt series; and (iii) how to cope with the inconsistency between the two different tilt series. Results: Here we develop a toolkit for fully automatic alignment of dual-axis electron tomography, with a simultaneous reconstruction procedure. The proposed toolkit and its workflow carries out the following solutions: (i) fully automatic detection and tracking of fiducial markers under large-field datasets; (ii) automatic combination of two different tilt series and global calibration of projection parameters; and (iii) inconsistency correction based on distortion correction parameters and the consequently simultaneous reconstruction. With all of these features, the presented toolkit can achieve accurate alignment and reconstruction simultaneously and conveniently under a single global coordinate system. Availability and implementation: The toolkit AuTom-dualx (alignment module dualxmauto and reconstruction module volrec_mltm) are accessible for general application at http://ear.ict.ac.cn, and the key source code is freely available under request. Supplementary information: Supplementary data are available at Bioinformatics online. Renmin Han, Albert F. Lawrence, Peng Yang 0010, Yu Li 0006, Sheng Wang 0001, Zhiyong Liu 0002, Xin Gao 0001, Fa Zhang 0001 |
Bioinform. | 11 |
| 2019 | PIXER: an automated particle-selection method based on segmentation using a deep neural networkabstractBACKGROUND: Cryo-electron microscopy (cryo-EM) has become a widely used tool for determining the structures of proteins and macromolecular complexes. To acquire the input for single-particle cryo-EM reconstruction, researchers must select hundreds of thousands of particles from micrographs. As the signal-to-noise ratio (SNR) of micrographs is extremely low, the performance of automated particle-selection methods is still unable to meet research requirements. To free researchers from this laborious work and to acquire a large number of high-quality particles, we propose an automated particle-selection method (PIXER) based on the idea of segmentation using a deep neural network. RESULTS: First, to accommodate low-SNR conditions, we convert micrographs into probability density maps using a segmentation network. These probability density maps indicate the likelihood that each pixel of a micrograph is part of a particle instead of just background noise. Particles selected from density maps have a more robust signal than do those directly selected from the original noisy micrographs. Second, at present, there is no segmentation-training dataset for cryo-EM. To enable our plan, we present an automated method to generate a training dataset for segmentation using real-world data. Third, we propose a grid-based, local-maximum method to locate the particles from the probability density maps. We tested our method on simulated and real-world experimental datasets and compared PIXER with the mainstream methods RELION, DeepEM and DeepPicker to demonstrate its performance. The results indicate that, as a fully automated method, PIXER can acquire results as good as the semi-automated methods RELION and DeepEM. CONCLUSION: To our knowledge, our work is the first to address the particle-selection problem using the segmentation network concept. As a fully automated particle-selection method, PIXER can free researchers from laborious particle-selection work. Based on the results of experiments, PIXER can acquire accurate results under low-SNR conditions within minutes. Jingrong Zhang, Renmin Han, Zhiyong Liu 0002, Fa Zhang 0001 |
BMC Bioinform. | 7 |
| 2019 | PABO: Mitigating congestion via packet bounce in data center networks
Lin Wang 0015, Fa Zhang 0001, Kai Zheng 0003, Max Mühlhäuser, Zhiyong Liu 0002 |
Comput. Commun. | 3 |
| 2019 | Energy-Aware Fault-Tolerant Dynamic Task Scheduling Scheme for Virtualized Cloud Data Centers
Avinab Marahatta, Youshi Wang, Fa Zhang 0001, Arun Kumar Sangaiah, Sumarga Kumar Sah Tyagi, Zhiyong Liu 0002 |
Mob. Networks Appl. | 3 |
| 2018 | Fiducial marker detection via deep learning approach for electron tomography
Renmin Han, Fa Zhang 0001, Shiwei Sun |
BIBM | 4 |
| 2018 | Mmalloc: A Dynamic Memory Management on Many-core Coprocessor for the Acceleration of Storage-intensive Bioinformatics Application
Mingzhe Zhang 0005, Jingrong Zhang, Rui Yan 0009, Zhiyong Liu 0002, Fa Zhang 0001, Xuefeng Cui |
BIBM | 7 |
| 2018 | A Hybrid Convolutional and Recurrent Deep Neural Network for Breast Cancer Pathological Image Classification
Rui Yan 0009, Yubo Ren, Xiaosong Rao, Chun-Hou Zheng 0001, Fa Zhang 0001 |
BIBM | 9 |
| 2018 | Memory-Efficient and Stabilizing Management System and Parallel Methods for RELION Using CUDA and MPI
Jingrong Zhang, Zhiyong Liu 0002, Fa Zhang 0001 |
ISBRA | 5 |
| 2018 | A fast fiducial marker tracking model for fully automatic alignment in electron tomographyabstractMotivation: Automatic alignment, especially fiducial marker-based alignment, has become increasingly important due to the high demand of subtomogram averaging and the rapid development of large-field electron microscopy. Among the alignment steps, fiducial marker tracking is a crucial one that determines the quality of the final alignment. Yet, it is still a challenging problem to track the fiducial markers accurately and effectively in a fully automatic manner. Results: In this paper, we propose a robust and efficient scheme for fiducial marker tracking. Firstly, we theoretically prove the upper bound of the transformation deviation of aligning the positions of fiducial markers on two micrographs by affine transformation. Secondly, we design an automatic algorithm based on the Gaussian mixture model to accelerate the procedure of fiducial marker tracking. Thirdly, we propose a divide-and-conquer strategy against lens distortions to ensure the reliability of our scheme. To our knowledge, this is the first attempt that theoretically relates the projection model with the tracking model. The real-world experimental results further support our theoretical bound and demonstrate the effectiveness of our algorithm. This work facilitates the fully automatic tracking for datasets with a massive number of fiducial markers. Availability and implementation: The C/C ++ source code that implements the fast fiducial marker tracking is available at https://github.com/icthrm/gmm-marker-tracking. Markerauto 1.6 version or later (also integrated in the AuTom platform at http://ear.ict.ac.cn/) offers a complete implementation for fast alignment, in which fast fiducial marker tracking is available by the '-t' option. Contact: [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online. Renmin Han, Fa Zhang 0001, Xin Gao 0001 |
Bioinform. | 2 |
| 2018 | DLBI: deep learning guided Bayesian inference for structure reconstruction of super-resolution fluorescence microscopyabstractMotivation: Super-resolution fluorescence microscopy with a resolution beyond the diffraction limit of light, has become an indispensable tool to directly visualize biological structures in living cells at a nanometer-scale resolution. Despite advances in high-density super-resolution fluorescent techniques, existing methods still have bottlenecks, including extremely long execution time, artificial thinning and thickening of structures, and lack of ability to capture latent structures. Results: Here, we propose a novel deep learning guided Bayesian inference (DLBI) approach, for the time-series analysis of high-density fluorescent images. Our method combines the strength of deep learning and statistical inference, where deep learning captures the underlying distribution of the fluorophores that are consistent with the observed time-series fluorescent images by exploring local features and correlation along time-axis, and statistical inference further refines the ultrastructure extracted by deep learning and endues physical meaning to the final image. In particular, our method contains three main components. The first one is a simulator that takes a high-resolution image as the input, and simulates time-series low-resolution fluorescent images based on experimentally calibrated parameters, which provides supervised training data to the deep learning model. The second one is a multi-scale deep learning module to capture both spatial information in each input low-resolution image as well as temporal information among the time-series images. And the third one is a Bayesian inference module that takes the image from the deep learning module as the initial localization of fluorophores and removes artifacts by statistical inference. Comprehensive experimental results on both real and simulated datasets demonstrate that our method provides more accurate and realistic local patch and large-field reconstruction than the state-of-the-art method, the 3B analysis, while our method is more than two orders of magnitude faster. Availability and implementation: The main program is available at https://github.com/lykaust15/DLBI. Supplementary information: Supplementary data are available at Bioinformatics online. Yu Li 0006, Fa Zhang 0001, Pingyong Xu, Mingshu Zhang, Ming Fan 0003, Lihua Li 0002, Xin Gao 0001, Renmin Han |
Bioinform. | 3 |
| 2017 | Joint Optimization of Server and Network Resource Utilization in Cloud Data CentersabstractVirtual machine placement is a key component of cloud resource management, which may affect network bandwidth allocation. In this paper, we revisit the virtual machine placement problem in cloud data centers and aim to maximize the overall resource utilization in multiple dimensions, while ensuring that the resource constraints on both the server such as CPU capacity and the network such as bandwidth are not violated. We model the bandwidth-guaranteed virtual machine placement problem and prove its NP-hardness, and design offline and online algorithms to solve the problem. We first consider the offline version and develop approximation algorithms with bounded performance ratios for both the homogeneous and the heterogeneous cases. Then, for the online version, we propose simple and efficient heuristics based on the insights from the offline algorithm design. Comprehensive experimental results verify that the overall resource utilization can be significantly improved by applying our proposals. Biyu Zhou, Jie Wu 0001, Lin Wang 0015, Fa Zhang 0001, Zhiyong Liu 0002 |
GLOBECOM | 4 |
| 2017 | PABO: Congestion mitigation via packet bounceabstractToday's data center applications can generate a diverse mix of short and long flows. However, switches used in a typical data center network are usually shallow buffered in order to reduce queueing delay and deployment cost. As a result, the buildup of the queues by long flows can block short flows, leading to frequent packet losses and retransmissions, which translates to crucial performance degradation. While multiple end-to-end TCP-based solutions have been proposed, none of them have tackled the real challenge: reliable transmission in the network. In this paper, we fill this gap by presenting PABO — a novel link-layer design that can mitigate congestion by temporarily bouncing packets to upstream switches. PABO's design fulfills the following demands: i) providing per-flow based flow control on the link layer, ii) handling transient congestion without the intervention of end devices, and iii) gradually back propagating the congestion signal to the source when the network is not capable to handle the congestion. We complete a proof-of-concept implementation, and experiments under different severities of congestion show that PABO outperforms the standard unreliable link-layer protocol by guaranteeing zero packet loss while introducing only a reasonable stretch on packet delay. Lin Wang 0015, Fa Zhang 0001, Kai Zheng 0003, Zhiyong Liu 0002 |
ICC | 3 |
| 2017 | Online Flow Scheduling with Deadline for Energy Conservation in Data Center NetworksabstractWe study the problem of flow scheduling in data center networks. Using speed scaling, our aim is to find an online scheduling algorithm that minimizes the total energy consumption of the network by determining both the transmission order and rates of the arriving flows while providing a strict flow deadline guarantee. Observing the superlinear property of link power consumption, the key challenge is in constantly determining the minimum transmission rate for “delay-tolerable” flows without any priori knowledge. To leverage the flow arrival pattern, we propose a probability-based flow prediction model to capture the uncertainty of the network flows. Based on the prediction model, we propose a tunable online flow scheduling algorithm to solve the online flow scheduling problem effectively. By introducing a scaling factor on bandwidth allocation, this algorithm allows us to conduct arbitrary trade-offs between the conservative and aggressive behaviors in terms of energy conser- vation. The effectiveness of the proposed algorithm is validated through rigorous theoretical analysis and further confirmed by extensive numerical simulations. Biyu Zhou, Jie Wu 0001, Lin Wang 0015, Fa Zhang 0001, Zhiyong Liu 0002 |
ICPADS | 4 |
| 2017 | Resource optimization for survivable embedding of virtual clusters in cloud data centersabstractWith the popularity of cloud computing, optimizing cloud resource consumption while providing predictable cloud service has become one of the focuses of research in recent years. In order to ensure a predictable performance, the requests from tenants are abstracted as Virtual Clusters, which not only specify the computing demands, but also establish the communication requirements among virtual machines. While much work has been done on virtual cluster embedding under a variety of goals, very few people have studied this issue in consideration of service survivability, which also plays a vital role in ensuring the performance in cloud data centers. In this paper, we study the resource optimization for survivable embedding of virtual clusters and aim to minimize the consumption of cloud resources in terms of server and bandwidth, while ensuring that both the resource constraints and the survivability constraints are not violated. We formally define this problem and analyze its complexity, and design efficient algorithms to solve the problem. Comprehensive experimental results verify that the overall resource consumption can be significantly reduced by applying our proposals. Biyu Zhou, Jie Wu 0001, Fa Zhang 0001, Zhiyong Liu 0002 |
IPCCC | 3 |
| 2017 | A Fully Automatic Geometric Parameters Determining Method for Electron Tomography
Fa Zhang 0001 |
ISBRA | 6 |
| 2017 | Accelerating Electron Tomography Reconstruction Algorithm ICON Using the Intel Xeon Phi Coprocessor on Tianhe-2 Supercomputer
Jingrong Zhang, Zhiyong Liu 0002, Fa Zhang 0001 |
ISBRA | 8 |
| 2017 | Real-time Task Scheduling for joint energy efficiency optimization in data centersabstractThe high energy consumption has become one bottleneck in the development of the data centers (DCs), where the main energy consumers are the cooling system and the servers. Therefore, the joint optimization for the energy efficiency of the cooling system and the servers is a crucial problem, while most of previous works on energy saving only studies one of these two components in an isolated manner. In this paper, we propose a real-time strategy, rTCS (real-time Task Classification and Scheduling strategy), to jointly optimize the energy efficiency of these two components in the scenario where the tasks arrive dynamically. Strategy rTCS first labels the tasks to classify them according to their run time and end time with a time complexity of O(1) and a bounded space complexity. Then, rTCS schedules the tasks in real time based on their labels and the energy consumption model of the DC. Simulation results show that rTCS can effectively improve the energy efficiency of DCs. Youshi Wang, Fa Zhang 0001, Rui Wang 0028, Yangguang Shi, Zhiyong Liu 0002 |
ISCC | 2 |
| 2017 | Hardness of Routing for Minimizing Superlinear Polynomial Cost in Directed Graphs
Yangguang Shi, Fa Zhang 0001, Zhiyong Liu 0002 |
TAMC | 2 |
| 2017 | A Two-Phase Improved Correlation Method for Automatic Particle Selection in Cryo-EMabstractParticle selection from cryo-electron microscopy (Cryo-EM) images is very important for high-resolution reconstruction of macromolecular structure. The methods of particle selection can be roughly grouped into two classes, template-matching methods and feature-based methods. In general, template-matching methods usually generate better results than feature-based methods. However, the accuracy of template-matching methods is restricted by the noise and low contrast of Cryo-EM images. Moreover, the processing speed of template-matching methods, restricted by the random orientation of particles, further limits their practical applications. In this paper, combining the advantages of feature-based methods and template-matching methods, we present a two-phase improved correlation method for automatic, fast particle selection. In Phase I, we generate a preliminary particle set using rotation-invariant features of particles. In Phase II, we filter the preliminary particle set using a correlation method to reduce the interference of the high noise background and improve the precision of particle selection. We apply several optimization strategies, including a modified adaboost algorithm, Divide and Conquer technique, cascade strategy and graphics processing unit parallel technique, to improve feature recognition ability and reduce processing time. In addition, we developed two correlation score functions for different correlation situations. Experimental results on the benchmark of Cryo-EM images show that our method can improve the accuracy and processing speed of particle selection significantly. Fa Zhang 0001, Xuan Wang 0002, Zhiyong Liu 0002 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2017 | Joint Optimization of Operational Cost and Performance Interference in Cloud Data CentersabstractVirtual machine (VM) scheduling is an important technique for the efficient operation of the computing resources in a data center. Previous work has mainly focused on consolidating VMs to improve resource utilization and to optimize energy consumption. However, the interference between collocated VMs is usually ignored, which can result in much worse performance degradation of the applications running on the VMs due to the contention of the shared resources. Based on this observation, we aim at designing efficient VM assignment and scheduling strategies in which we consider optimizing both the operational cost of the data center and the performance degradation of the running applications. We then propose a general model that captures the tradeoff between the two contradictory objectives. We present offline and online solutions for this problem by exploiting the spatial and temporal information of performance interference of VM collocation, where VM scheduling is performed by jointly considering the combinations and the life-cycle overlap of the VMs. Evaluation results show that the proposed methods can generate efficient schedules for VMs, achieving low operational cost while significantly reducing the performance degradation of applications in cloud data centers. Xibo Jin, Fa Zhang 0001, Lin Wang 0015, Songlin Hu 0001, Biyu Zhou, Zhiyong Liu 0002 |
IEEE Trans. Cloud Comput. | 2 |
| 2017 | Towards the Tradeoffs in Designing Data Center Network ArchitecturesabstractExisting Data Center Network (DCN) architectures are classified into two categories: switch-centric and server-centric architectures. In switch-centric DCNs, routing intelligence is placed on switches; each server usually uses only one port of the Network Interface Card (NIC) to connect to the network. In server-centric DCNs, switches are only used as cross-bars, and routing intelligence is placed on servers, where multiple NIC ports may be used. In this paper, we formally introduce a new category of DCN architectures: thedual-centricDCN architectures, where routing intelligence can be placed on both switches and servers. The dual-centric philosophy can achieve various tradeoffs in designing DCN architectures. We propose three novel dual-centric DCN architectures: FCell, FRectangle, and FSquare, all of which are based on the folded Clos topology. FCell is a power-efficient DCN architecture, with a larger diameter and lower bisection bandwidth than FSquare and FRectangle. FSquare is a high performance DCN architecture, in which the diameter is small and the bisection bandwidth is large; however, the DCN power consumption per server in FSquare is high. FRectangle significantly reduces the DCN power consumption per server, compared to FSquare, at the sacrifice of some networking performances. By investigating FCell, FRectangle and FSquare, and by comparing them with existing architectures, we demonstrate that, the three novel dual-centric architectures enjoy the advantages of both switch-centric designs and server-centric designs, have various nice properties for practical data centers, and provide flexible tradeoff choices in designing DCN architectures. Dawei Li 0002, Jie Wu 0001, Zhiyong Liu 0002, Fa Zhang 0001 |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2016 | aWGRS: Automates paired-end whole genome re-sequencing data analysis frameworkabstractIn order to enable people to avoid too many cumbersome and complex operations of the command line and repeated parameter adjustments, automates pair-end whole genome re-sequence (aWGRS) data processing whereby pre-installed dependencies are presented in this paper, which are used to map reads to a reference and realign variations. This method presents aWGRS which is a method that takes as input paired-end reads and a reference genome and returns re-sequencing information. The concept behind the development of this tool is that re-sequencing requires several steps: alignment to the reference, single nucleotide polymorphisms (SNPs) calling, Insertion / Deletion (InDels) calling, structure variant (SVs) calling, and annotation. By introducing and adjusting a new concept called the recall rate, the coverage rate and accuracy rate can be met at the same time. Within the range of recall rate, a variation is evaluated by two criteria: the quality value and the number of reads that support it, and one read with higher quality value and larger supported number will be picked out finally. Genome-wide genetic variations between precocious trifoliate orange and its wild type are identified in [1], and empirical results show that there is a big reduction in the amount of variation and great improvement of accuracy between the results of aWGRS and [1] which offered by the Beijing Genomics Institute (BGI). Overall, the adjustable parameters adopted in aWGRS can affect the results of the experiment and the default filtering strategy using the mutation recall rate also can attain good results automatically. Xiujuan Sun, Fa Zhang 0001, Jinzhi Zhang |
BIBM | 2 |
| 2016 | HDEER: A Distributed Routing Scheme for Energy-Efficient NetworkingabstractThe proliferation of new online Internet services has substantially increased the energy consumption in wired networks, which has become a critical issue for Internet service providers. In this paper, we target the network-wide energy-saving problem by leveraging speed scaling as the energy-saving strategy. We propose a distributed routing scheme-HDEER-to improve network energy efficiency in a distributed manner without significantly compromising traffic delay. HDEER is a two-stage routing scheme where a simple distributed multipath finding algorithm is firstly performed to guarantee loop-free routing, and then a distributed routing algorithm is executed for energy-efficient routing in each node among the multiple loop-free paths. We conduct extensive experiments on the NS3 simulator and simulations with real network topologies in different scales under different traffic scenarios. Experiment results show that HDEER can reduce network energy consumption with a fair tradeoff between network energy consumption and traffic delay. Biyu Zhou, Fa Zhang 0001, Lin Wang 0015, Chenying Hou, Antonio Fernández 0001, Athanasios V. Vasilakos, Youshi Wang, Jie Wu 0001, Zhiyong Liu 0002 |
IEEE J. Sel. Areas Commun. | 2 |
| 2015 | Dual-centric Data Center Network ArchitecturesabstractExisting Data Center Network (DCN) architectures are classified into two categories: switch-centric and server-centric architectures. In switch-centric DCNs, routing intelligence is placed on switches, each server usually uses only one port of the Network Interface Card (NIC) to connect to the network. In server-centric DCNs, switches are only used as cross-bars, and routing intelligence is placed on servers, where multiple NIC ports may be used. In this paper, we formally introduce a new category of DCN architectures: the dual-centric DCN architectures, where routing intelligence can be placed on both switches and servers. We propose two typical dual-centric DCN architectures: FSquare and Rectangle, both of which are based on the folded Clos topology. FSquare is a high performance DCN architecture, in which the diameter is small and the bisection bandwidth is large, however, the DCN power consumption per server in FSquare is high. Rectangle significantly reduces the DCN power consumption per server, compared to FSquare, at the sacrifice of some performances, thus, Rectangle has a larger diameter and a smaller bisection bandwidth. By investigating FSquare and Rectangle, and by comparing them with existing architectures, we demonstrate that, these two novel dual-centric architectures enjoy the advantages of both switch-centric designs and server-centric designs, have various nice properties for practical data centers, and provide flexible choices in designing DCN architectures. Dawei Li 0002, Jie Wu 0001, Zhiyong Liu 0002, Fa Zhang 0001 |
ICPP | 4 |
| 2015 | Multi-resource energy-efficient routing in cloud data centers with network-as-a-serviceabstractWith the rapid development of software defined networking and network function virtualization, researchers have proposed a new cloud networking model called Network-as-a-Service (NaaS) which enables both in-network packet processing and application-specific network control. In this paper, we revisit the problem of achieving network energy efficiency in data centers and identify some new optimization challenges under the NaaS model. Particularly, we extend the energy-efficient routing optimization from single-resource to multi-resource settings. We characterize the problem through a detailed model and provide a formal problem definition. Due to the high complexity of direct solutions, we propose a greedy routing scheme to approximate the optimum, where flows are selected progressively to exhaust residual capacities of active nodes, and routing paths are assigned based on the distributions of both node residual capacities and flow demands. By leveraging the structural regularity of data center networks, we also provide a fast topology-aware heuristic method based on hierarchically solving a series of vector bin packing instances. Extensive simulations show that the proposed routing scheme can achieve significant gain on energy savings and the topology-aware heuristic can produce comparably good results while reducing the computation time to a large extent. Lin Wang 0015, Antonio Fernández 0001, Fa Zhang 0001, Jie Wu 0001, Zhiyong Liu 0002 |
ISCC | 3 |
| 2015 | Scheduling for energy minimization on restricted parallel processors
Xibo Jin, Fa Zhang 0001, Liya Fan, Zhiyong Liu 0002 |
J. Parallel Distributed Comput. | 2 |
| 2015 | Randomized oblivious integral routing for minimizing power cost
Yangguang Shi, Fa Zhang 0001, Jie Wu 0001, Zhiyong Liu 0002 |
Theor. Comput. Sci. | 2 |
| 2014 | Energy-Efficient Flow Scheduling and Routing with Hard Deadlines in Data Center NetworksabstractThe power consumption of enormous network devices in data centers has emerged as a big concern to data center operators. Despite many traffic-engineering-based solutions, very little attention has been paid on performance-guaranteed energy saving schemes. In this paper, we propose a novel energy-saving model for data center networks by scheduling and routing "deadline-constrained flows" where the transmission of every flow has to be accomplished before a rigorous deadline, being the most critical requirement in production data center networks. Based on speed scaling and power-down energy saving strategies for network devices, we aim to explore the most energy efficient way of scheduling and routing flows on the network, as well as determining the transmission speed for every flow. We consider two general versions of the problem. For the version of only flow scheduling where routes of flows are pre-given, we show that it can be solved polynomially and we develop an optimal combinatorial algorithm for it. For the version of joint flow scheduling and routing, we prove that it is strongly NP-hard and cannot have a Fully Polynomial-Time Approximation Scheme (FPTAS) unless P=NP. Based on a relaxation and randomized rounding technique, we provide an efficient approximation algorithm which can guarantee a provable performance ratio with respect to a polynomial of the total number of flows. Lin Wang 0015, Fa Zhang 0001, Kai Zheng 0003, Athanasios V. Vasilakos, Shaolei Ren, Zhiyong Liu 0002 |
ICDCS | 2 |
| 2014 | Polylogarithmic Competitive Algorithm for Energy Minimization in Optical WDM NetworksabstractWe study the energy minimization problem (EMP) in the optical WDM networks with arbitrary topologies. It is assumed that the traffic requests can arrive at and depart from the network arbitrarily, and idle network devices can be dynamically switched off to save energy. For each traffic request R, we need to specify a wavelength λRand a fiber in each link along its path to carry λR. The objective is to minimize the energy consumption incurred by the active devices over the entire network for any time period [0, t]. In this paper, a randomized online algorithm is proposed for EMP. Particularly, for each traffic request, our algorithm only needs O(1)-time to determine the wavelength, and the fiber allocation procedure can be performed in a fully distributed manner in each link with polynomial time. The competitive ratio of our algorithm is bounded by O(log μ · log hmax), where μ represents the number of wavelengths carried by each fiber and hmaxrepresents the holding time of the longest traffic request. Yangguang Shi, Fa Zhang 0001, Zhiyong Liu 0002 |
ICNP | 2 |
| 2014 | Joint power optimization through VM placement and flow scheduling in data centersabstractTwo important components that consume the majority of IT power in data centers are the servers and the Data Center Network (DCN). Existing works fail to fully utilize power management techniques on the servers and in the DCN at the same time. In this paper, we jointly consider VM placement on servers with scalable frequencies and flow scheduling in the DCN, to minimize the overall system's power consumption. Due to the convex relation between a server's power consumption and its operating frequency, we prove that, given the number of servers to be used, computation workloads should be allocated to severs in a balanced way, to minimize the power consumption on servers. To reduce the power consumption of the DCN, we further consider the flow requirements among the VMs during VM allocation and assignment. Also, after VM placement, flow consolidation is conducted to reduce the number of active switches and ports. We notice that, choosing the minimum number of servers to accommodate the VMs may result in high power consumption on servers, due to servers' increased operating frequencies. Choosing the optimal number of servers purely based on servers' power consumption leads to reduced power consumption on servers, but may increase power consumption of the DCN. We propose to choose the optimal number of servers to be used, based on the overall system's power consumption. Simulations show that, our joint power optimization method helps to reduce the overall power consumption significantly, and outperforms various existing state-of-the-art methods in terms of reducing the overall system's power consumption. Dawei Li 0002, Jie Wu 0001, Zhiyong Liu 0002, Fa Zhang 0001 |
IPCCC | 4 |
| 2014 | An Improved Correlation Method Based on Rotation Invariant Feature for Automatic Particle Selection
Xuan Wang 0002, Fa Zhang 0001 |
ISBRA | 5 |
| 2014 | A Parallel Scheme for Three-Dimensional Reconstruction in Large-Field Electron Tomography
Jingrong Zhang, Fa Zhang 0001, Xuan Wang 0002, Zhiyong Liu 0002 |
ISBRA | 3 |
| 2014 | DEER: A distributed routing scheme for achieving network energy efficiencyabstractThe rapid growth of Internet services has brought emergent concerns over network energy efficiency. This study aims to improve network energy efficiency using power-down technique. We propose DEER, a fully distributed routing scheme. The main concept of DEER is to dynamically allocate the traffic demands in the nodes so that some links connected to the nodes can be put into sleep mode, thus reducing the energy consumption. The special features of DEER include that it does not need global traffic matrix of the network and that it uses only the local information of link loads, making DEER be able to be implemented in a distributed manner without centralized control. We develop algorithms in DEER to dynamically change the link state (into active or sleep mode) according to link utilization and to balance the loads of links by adjusting the link weights. With the traffic load varying over time, the link state transformation is triggered when any of the pre-defined thresholds is violated. Extensive simulations with the network topology and real traffic traces from the GÉANT network confirm that by involving DEER, up to 50% of the links can be put into sleep while the frequency of chaining the state of a link stays fairly low. Biyu Zhou, Lin Wang 0015, Fa Zhang 0001, Xibo Jin, Zhiyong Liu 0002 |
LANMAN | 3 |
| 2014 | GreenDCN: A General Framework for Achieving Energy Efficiency in Data Center NetworksabstractThe popularization of cloud computing has raised concerns over the energy consumption that takes place in data centers. In addition to the energy consumed by servers, the energy consumed by large numbers of network devices emerges as a significant problem. Existing work on energy-efficient data center networking primarily focuses on traffic engineering, which is usually adapted from traditional networks. We propose a new framework to embrace the new opportunities brought by combining some special features of data centers with traffic engineering. Based on this framework, we characterize the problem of achieving energy efficiency with a time-aware model, and we prove its NP-hardness with a solution that has two steps. First, we solve the problem of assigning virtual machines (VM) to servers to reduce the amount of traffic and to generate favorable conditions for traffic engineering. The solution reached for this problem is based on three essential principles that we propose. Second, we reduce the number of active switches and balance traffic flows, depending on the relation between power consumption and routing, to achieve energy conservation. Experimental results confirm that, by using this framework, we can achieve up to 50 percent energy savings. We also provide a comprehensive discussion on the scalability and practicability of the framework. Lin Wang 0015, Fa Zhang 0001, Jordi Arjona Aroca, Athanasios V. Vasilakos, Kai Zheng 0003, Chenying Hou, Dan Li 0001, Zhiyong Liu 0002 |
IEEE J. Sel. Areas Commun. | 2 |
| 2013 | Energy-Efficient Scheduling with Time and Processors Eligibility Restrictions
Xibo Jin, Fa Zhang 0001, Liya Fan, Zhiyong Liu 0002 |
Euro-Par | 2 |
| 2013 | Risk management for virtual machines consolidation in data centersabstractVirtual machines (VMs) consolidation has emerged as an important method for the design of energy-efficient data centers. The purpose is to aggregate VMs to fewer physical machines and put the idle servers into power-saving mode. Existing researches mainly focus on transforming the VMs consolidation into various bin packing problems. However, VMs consolidation may cost Service Level Agreement (SLA) violations just after the migration due to the uncertainty of applications' demands. In this paper, we provide a SLA risk management framework, involving a stochastic program to solve the resource allocation for VMs and an algorithm for dynamic VMs consolidation at runtime, to optimize both the energy consumption saving and SLA violations. We validate the proposed algorithm using workloads from a real world system. The results compare with other VMs consolidation algorithms that without considering risk, and show that our SLA violations is reduced by four times from 25% to 2% - 5% while only losing little energy consumption saving. Xibo Jin, Fa Zhang 0001, Songlin Hu 0001, Zhiyong Liu 0002 |
GLOBECOM | 2 |
| 2013 | Improving the Network Energy Efficiency in MapReduce SystemsabstractApart from servers, the energy consumed by enormous amount of network devices in data centers also emerges as a big problem. Existing work on energy- efficient data center networking primarily focuses on traffic engineering to consolidate flows and shut down unused devices, not considering another important factor, virtual machine assignment, which has been shown to have a big influence on traffic engineering. Moreover, the lack of information about upper layer applications leads to misunderstand the traffic patterns of the network. This may result in poor effectiveness in the traffic-based optimization in practice. In this paper, we aim to achieve better network energy efficiency in MapReduce systems by combining virtual machine assignment and traffic engineering. By exploiting the characteristics of MapReduce applications, we provide a unified model to describe this problem. Due to its NP-hardness, a general framework is proposed to solve it, where virtual machines are first clustered and then different virtual machine assignments are generated greedily and a local search procedure is used to improve them. The local search procedure depends on the results of an energy-efficient routing provided by GEERA. GEERA is an approximate algorithm designed to select routing paths for flows. Experimental results confirm the efficiency of GEERA, as well as the overall framework. By using this framework, up to $20\%$ more energy savings can be achieved compared with sole traffic engineering solutions. Lin Wang 0015, Fa Zhang 0001, Zhiyong Liu 0002 |
ICCCN | 2 |
| 2013 | A Universally Stable and Energy-Efficient Scheduling Protocol for Packet Switching NetworkabstractEnergy efficiency is becoming an important issue in networks. Although some research works have been devoted to this topic, only a little attention has been paid to the stability of the network equipped with the energy conservation mechanisms. In fact, we find that the stability of networks can be undermined in the worst case if it isn't considered with care by the energy conservation mechanism.In this paper, we propose an energy-efficient scheduling protocol which can guarantee the stability of the network in all cases. We start by building a new model which can be used to verify the stability of the network equipped with the energy conservation mechanisms. With this model, we transform the packet scheduling problem to a Job Shop Scheduling problem. For this problem, we propose a time-stamp based work-conserving scheduling algorithm - G-FSA. Compared with existing methods, this scheduling algorithm guarantees a tighter bound on the make span of the jobs. Then, we integrate G-FSA with a time partition approach to generate our energy-efficient packet scheduling protocol. It is proved that the obtained protocol can guarantee the stability of network in all cases. And its approximation ratio in terms of energy efficiency can be bounded by O((1+∈ )α) for any ∈ > 0, where is an input parameter depending on the hardware infrastructure. Typically, 1 <; α ≤ 3. Yangguang Shi, Fa Zhang 0001, Zhiyong Liu 0002 |
NCA | 2 |
| 2013 | Incorporating Rate Adaptation Into Green Networking for Future Data CentersabstractDespite some proposals for energy-efficient topologies, most of the studies for saving energy in data center networks are focused on traffic engineering, i.e., consolidating flows and switching off unnecessary network devices. The major weakness of this approach is network oscillation brought by the frequent change of network topology when traffic fluctuates very fast. In this paper, we propose to incorporate rate adaptation into green data center networks. With rate adaptive network devices, we aim at approaching network-wide energy proportionality by routing optimization. We formalize the problem with an integer program and propose an efficient approximation algorithm - TSRR, solving the problem quickly while guaranteeing a constant performance ratio. Extensive range of simulations confirm that more than 40% of the energy can be saved while introducing very slight stretch on network delay. Lin Wang 0015, Fa Zhang 0001, Chenying Hou, Jordi Arjona Aroca, Zhiyong Liu 0002 |
NCA | 2 |
| 2013 | A fast calculation strategy of density function in ISAF reconstruction algorithm
Gongming Wang, Fa Zhang 0001, Qi Chu 0002, Liya Fan, Zhiyong Liu 0002 |
Sci. China Inf. Sci. | 2 |
| 2012 | Energy-Efficient Network Routing with Discrete Cost Functions
Lin Wang 0015, Antonio Fernández 0001, Fa Zhang 0001, Chenying Hou, Zhiyong Liu 0002 |
TAMC | 3 |
| 2012 | High-performance blob-based iterative three-dimensional reconstruction in electron tomography using multi-GPUsabstractBACKGROUND: Three-dimensional (3D) reconstruction in electron tomography (ET) has emerged as a leading technique to elucidate the molecular structures of complex biological specimens. Blob-based iterative methods are advantageous reconstruction methods for 3D reconstruction in ET, but demand huge computational costs. Multiple graphic processing units (multi-GPUs) offer an affordable platform to meet these demands. However, a synchronous communication scheme between multi-GPUs leads to idle GPU time, and a weighted matrix involved in iterative methods cannot be loaded into GPUs especially for large images due to the limited available memory of GPUs. RESULTS: In this paper we propose a multilevel parallel strategy combined with an asynchronous communication scheme and a blob-ELLR data structure to efficiently perform blob-based iterative reconstructions on multi-GPUs. The asynchronous communication scheme is used to minimize the idle GPU time so as to asynchronously overlap communications with computations. The blob-ELLR data structure only needs nearly 1/16 of the storage space in comparison with ELLPACK-R (ELLR) data structure and yields significant acceleration. CONCLUSIONS: Experimental results indicate that the multilevel parallel scheme combined with the asynchronous communication scheme and the blob-ELLR data structure allows efficient implementations of 3D reconstruction in ET on multi-GPUs. Fa Zhang 0001, Qi Chu 0002, Zhiyong Liu 0002 |
BMC Bioinform. | 2 |
| 2012 | An effective approximation algorithm for the Malleable Parallel Task Scheduling problem
Liya Fan, Fa Zhang 0001, Gongming Wang, Zhiyong Liu 0002 |
J. Parallel Distributed Comput. | 2 |
| 2011 | High-Performance Blob-Based Iterative Reconstruction of Electron Tomography on Multi-GPUs
Fa Zhang 0001, Qi Chu 0002, Zhiyong Liu 0002 |
ISBRA | 2 |
| 2010 | An accurate, automatic method for markerless alignment of electron tomographic imagesabstractAccurate alignment of electron tomographic images without using embedded gold particles as fiducial markers is still a challenge. Here we propose a new markerless alignment method that employs Scale Invariant Feature Transform features (SIFT) as virtual markers. It differs from other types of feature in a way the sufficient and distinctive information it represents. This characteristic makes the following feature matching and tracking steps automatic and more reliable, which allows for estimating alignment parameters accurately. Furthermore, we use Sparse Bundle Adjustment (SPA) with M-estimation to estimate alignment parameters for each image. Experiments show that our method can achieve a reprojection residual less than 0.4 pixel and can approach the same accuracy of marker alignment. Besides, our method can apply to adjusting typical misalignments such as magnitude divergences or in-plane rotation and can detect bad images. Qi Chu 0002, Fa Zhang 0001, Zhiyong Liu 0002 |
BIBM | 2 |
| 2009 | An effective scheduling algorithm for Linear Makespan Minimization on Unrelated Parallel MachinesabstractA simple yet common scheduling problem is identified, as a special case of the R||Cmaxproblem. We name it Linear Makespan Minimization on Unrelated Parallel Machines (LMMUPM). A novel algorithm, MOBSA (Multi-Objective Based Scheduling Algorithm), is presented to solve it. Two auxiliary problems are introduced as the basis of our algorithm. The first one can be reduced to a Multi-Objective Integer Program, while the second is constructed based on the solution of the first one. Results on random datasets revealed that MOBSA produced smaller and more stable makespans than other scheduling algorithms. Additionally, the makespan produced by MOBSA was within 1% of the optimum for every case. Presently, MOBSA has been applied to parallelize EMAN, one of the most popular software packages for cryo-electron microscopy single particle reconstruction. High speedups and ideal load balancing have been obtained. It is expected that MOBSA is also applicable to other similar applications. Liya Fan, Fa Zhang 0001, Gongming Wang, Zhiyong Liu 0002 |
HiPC | 2 |
| 2009 | Modified Simultaneous Algebraic Reconstruction Technique and its Parallelization in Cryo-electron TomographyabstractThree-dimensional reconstruction of cryo-electron tomography (cryo-ET) has emerged as the leading technique in analyzing structures of complex pleomorphic cellulars. A classical iterative method, simultaneous algebraic reconstruction technique (SART), has been employed to reconstruct volume images in cryo-ET. However, SART starts with an arbitrary approximation and takes into account only a weighted factor when updating density value in every error-correction iterative procedure, thus limits the improvement of the reconstruction resolution. Facing these problems, we present a modified simultaneous algebraic reconstruction technique (MSART) which applies several key techniques, a back projection technique (BPT) and an adaptive adjustment of corrections. Experimental results show that MSART can improve significantly the quality of reconstruction. Additionally, in order to address the computational requirements demanded by the reconstruction of large volumes, we have presented and implanted a strategy to parallel the MSART algorithm on DAWNING 4000H cluster system, and obtained a good computational performance. Fa Zhang 0001, Zhiyong Liu 0002 |
ICPADS | 2 |
| 2009 | A framework to refine particle clusters produced by EMANabstractMOTIVATION: EMAN is one of the most popular software packages for single particle reconstruction. But the particle clusters produced during its model refining stage are of low qualities. We attempt to refine the particle clusters by more accurately determining orientations of particles, and thereby achieving higher resolutions of consequent 3D structures. RESULTS: A particle reclustering framework (PRF) is introduced, which consists of three components. Each of them is responsible for one of the basic tasks of PRF: normalization, threshold determination and reclustering. Our implementation is also described and proved to meet the constraints proposed by PRF. Experiments revealed that our implementation improved resolutions of consequent structures for most cases, but only a little extra execution time was incurred. Therefore, it is practical to incorporate PRF in EMAN to improve qualities of generated 3D structures. AVAILABILITY AND IMPLEMENTATION: Implementation of our algorithm is available upon request from the authors. Liya Fan, Fa Zhang 0001, Gongming Wang, Zhiyong Liu 0002 |
Bioinform. | 2 |
| 2007 | Using Domain-Based Structural Ensemble to Improve Structure ModelingabstractIn this paper, we presented a method to improve structural modeling based on conserved domain clusters and structure-anchored alignment. First we mapped all the InterPro domains in the entire PDB, partitioned and clustered homologous domains into the domain-based template library. This aimed at expanding structural coverage to more protein sequences. For each cluster, we generated a multiple structural alignment based only on the 3 D information. Then we extracted a core-structure and built a position-specific profile from the structure and sequence information for each of cluster. Based on the multiple structural alignments, core-structures and the profiles, we developed a structure-anchored alignment method to increase the alignment accuracy between a query and its templates. Preliminary results show that our template library and the structure-anchored alignment method can be used for the prediction for a majority of known protein sequences with better qualities. Fa Zhang 0001, Zhaoyun Ma, Zhiyong Liu 0002 |
BIBE | 1 |
| 2006 | A profile-based protein sequence alignment algorithm for a domain clustering databaseabstractAiming at the two main shortcomings in Homology Modeling, we have designed and established a domain clustering database. Searching the database is a fundamental work for it. However, current alignment algorithms are mainly based on the sequences, ignoring the structure conservation in domain. This paper proposed a profile-based alignment which considers the structure information into the profile, based on the character of our domain database. We designed an experiment within the database. The results show that both the quality and sensitivity of our scheme are better than pure Smith-Waterman and sequence-based profile algorithms. We strongly believe that this work can help to improve the protein structure prediction Fa Zhang 0001, Zhiyong Liu 0002 |
CIBCB | 2 |
| 2006 | A method to integrate, assess and characterize the protein-protein interactionsabstractRecently, large-scale protein-protein interactions were recovered using the similar two-hybrid system for the model systems. This information allows us to investigate the protein interaction network from a systematic point of view. However, experimentally determined interactions are susceptible to errors. A previous assessment estimated that only ~10% of the interactions can be supported by more than one independent experiment, and about half of the interactions may be false positives. These false positives might unnecessarily link unrelated proteins, resulting in huge apparent interaction clusters, which complicate elucidation for the biological importance of these interactions. Address this problem, we present an approach to integrate, assess and characterize all available protein-protein interactions in model organisms yeast and fly. We first integrate all available protein-protein interaction databases of yeast and fly, and merge all the datasets. We then use machine learning techniques to score the reliability for each interaction, and to rigorously validate the scoring scheme of yeast protein-protein interactions from different aspects. Our results show that this scoring scheme provides a good basis for selecting reliable protein-protein interaction dataset Fa Zhang 0001, Jingchun Chen, Zhiyong Liu 0002 |
CIBCB | 1 |
| 2006 | A method to improve structural modeling based on conserved domain clustersabstractHomology modeling requires an accurate alignment between a query sequence and its homologs with known three-dimensional (3D) information. Current structural modeling techniques largely use entire protein chains as templates, which are selected based only on their sequence alignments with the queries. Protein can be largely described as combinations of conserved domains, and already more than two-third of the known protein domains can be found in the protein data bank (PDB). We presented a method to improve structural modeling based on conserved domain clusters. First, we searched and mapped all the inter pro domains in the entire PDB, partitioned and clustered homologous domains into the domain-based template library. For each of the resulting clusters created, a multiple structural alignment was generated based only on the 3D coordinates of all the residues involved. Then we used the structural alignments as anchors to increase the alignment accuracy between a query and its templates, and consequently improve the quality of predicted structure for query protein. We implemented the method on DAWNING 4000A cluster system. The preliminary results show that our domain-based template library and the structure-anchored alignment protocol can be used for the partial prediction for a majority of known protein sequences with better qualities. Fa Zhang 0001 |
IPDPS | 1 |
| 2004 | Parallel divide and conquer bio-sequence comparison based on smith-waterman algorithm
Fa Zhang 0001, Xiangzhen Qiao, Zhiyong Liu 0002 |
Sci. China Ser. F Inf. Sci. | 1 |