Ming Hu 0001

dblp:82/378-1 · DBLP profile ↗
← Back
14ranked-venue papers
3as first author
4since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 14 · 3 first-author · 4 since 2021
YearPublicationVenuePosition
2025 MCAMamba: Multilevel Cross-Modal Attention-Guided State-Space Model for Multisource Remote Sensing Image Classification
abstract
Effective fusion of multi-source remote sensing data remains a fundamental challenge for Earth observation, as CNN and Transformer models suffer from limited receptive fields and high computational complexity. While State Space Models (SSM) like Mamba show promise in sequence modeling, they face three critical challenges in multi-source remote sensing: insufficient spatial-spectral coordination, cross-modal heterogeneity, and inadequate multi-scale feature integration. To address these limitations, this paper proposes MCAMamba: a Multi-Level Cross-Modal Attention-Guided Mamba framework for joint classification of hyperspectral images (HSI) and Light Detection and Ranging (LiDAR)/Synthetic Aperture Radar (SAR) data. MCAMamba introduces a novel three-stage feature fusion pipeline: 1) The FExt-Attention module enhances spatial structure and spectral information through parallel spatial-channel attention mechanisms. 2) The SSM-Attention module achieves deep cross-modal fusion by combining attention mechanisms with SSM for parametric interaction. 3) The FFus-Attention module performs adaptive multi-scale feature integration through global context modeling and cascaded attention. This hierarchical design enables superior feature representation with enhanced computational efficiency. Experiments on four public benchmark datasets (Houston2013, Houston2018, Augsburg, and Berlin) show that MCAMamba achieves Overall Accuracy (OA) of 94.75%, 93.35%, 92.46%, and 79.18%. The code will be available at https://github.com/Dmygithub/MCAMamba.
Mingyu Dou, Shi Qiu 0002, Ming Hu 0001, Xiaozhen Qiao, Huping Ye, Xiaohan Liao, Zhe Sun 0007
IEEE Trans. Geosci. Remote. Sens.3
2024 SnapHiC-G: identifying long-range enhancer-promoter interactions from single-cell Hi-C data via a global background model
abstract
Harnessing the power of single-cell genomics technologies, single-cell Hi-C (scHi-C) and its derived technologies provide powerful tools to measure spatial proximity between regulatory elements and their target genes in individual cells. Using a global background model, we propose SnapHiC-G, a computational method, to identify long-range enhancer-promoter interactions from scHi-C data. We applied SnapHiC-G to scHi-C datasets generated from mouse embryonic stem cells and human brain cortical cells. SnapHiC-G achieved high sensitivity in identifying long-range enhancer-promoter interactions. Moreover, SnapHiC-G can identify putative target genes for noncoding genome-wide association study (GWAS) variants, and the genetic heritability of neuropsychiatric diseases is enriched for single-nucleotide polymorphisms (SNPs) within SnapHiC-G-identified interactions in a cell-type-specific manner. In sum, SnapHiC-G is a powerful tool for characterizing cell-type-specific enhancer-promoter interactions from complex tissues and can facilitate the discovery of chromatin interactions important for gene regulation in biologically relevant cell types.
Weifang Liu, Wujuan Zhong, Paola Giusti-Rodríguez, Zhiyun Jiang, Geoffery W. Wang, Huaigu Sun, Ming Hu 0001
Briefings Bioinform.7
2023 SnapHiC-D: a computational pipeline to identify differential chromatin contacts from single-cell Hi-C data
abstract
Single-cell high-throughput chromatin conformation capture technologies (scHi-C) has been used to map chromatin spatial organization in complex tissues. However, computational tools to detect differential chromatin contacts (DCCs) from scHi-C datasets in development and through disease pathogenesis are still lacking. Here, we present SnapHiC-D, a computational pipeline to identify DCCs between two scHi-C datasets. Compared to methods designed for bulk Hi-C data, SnapHiC-D detects DCCs with high sensitivity and accuracy. We used SnapHiC-D to identify cell-type-specific chromatin contacts at 10 Kb resolution in mouse hippocampal and human prefrontal cortical tissues, demonstrating that DCCs detected in the hippocampal and cortical cell types are generally associated with cell-type-specific gene expression patterns and epigenomic features. SnapHiC-D is freely available at https://github.com/HuMingLab/SnapHiC-D.
Lindsay Lee, Xiaoqi Li 0018, Chenxu Zhu, Yanxiao Zhang, Hongyu Yu, Ziyin Chen, Shreya Mishra, Ming Hu 0001
Briefings Bioinform.11
2022 A systematic evaluation of Hi-C data enhancement methods for enhancing PLAC-seq and HiChIP data
abstract
The three-dimensional organization of chromatin plays a critical role in gene regulation. Recently developed technologies, such as HiChIP and proximity ligation-assisted ChIP-Seq (PLAC-seq) (hereafter referred to as HP for brevity), can measure chromosome spatial organization by interrogating chromatin interactions mediated by a protein of interest. While offering cost-efficiency over genome-wide unbiased high-throughput chromosome conformation capture (Hi-C) data, HP data remain sparse at kilobase (Kb) resolution with the current sequencing depth in the order of 108 reads per sample. Deep learning models, including HiCPlus, HiCNN, HiCNN2, DeepHiC and Variationally Encoded Hi-C Loss Enhancer (VEHiCLE), have been developed to enhance the sequencing depth of Hi-C data, but their performance on HP data has not been benchmarked. Here, we performed a comprehensive evaluation of HP data sequencing depth enhancement using models developed for Hi-C data. Specifically, we analyzed various HP data, including Smc1a HiChIP data of the human lymphoblastoid cell line GM12878, H3K4me3 PLAC-seq data of four human neural cell types as well as of mouse embryonic stem cells (mESC), and mESC CCCTC-binding factor (CTCF) PLAC-seq data. Our evaluations lead to the following three findings: (i) most models developed for Hi-C data achieve reasonable performance when applied to HP data (e.g. with Pearson correlation ranging 0.76-0.95 for pairs of loci within 300 Kb), and the enhanced datasets lead to improved statistical power for detecting long-range chromatin interactions, (ii) models trained on HP data outperform those trained on Hi-C data and (iii) most models are transferable across cell types. Our results provide a general guideline for HP data enhancement using existing methods designed for Hi-C data.
Gang Li 0034, Minzhi Jiang, Armen Abnousi, Jonathan D. Rosen, Ming Hu 0001
Briefings Bioinform.8
2020 Inferring Spatial Organization of Individual Topologically Associated Domains via Piecewise Helical Model
abstract
The recently developed Hi-C technology enables a genome-wide view of chromosome spatial organizations, and has shed deep insights into genome structure and genome function. However, multiple sources of uncertainties make downstream data analysis and interpretation challenging. Specifically, statistical models for inferring three-dimensional (3D) chromosomal structure from Hi-C data are far from their maturity. Most existing methods are highly over-parameterized, lacking clear interpretations, and sensitive to outliers. In this study, we propose a parsimonious, easy to interpret, and robust piecewise helical model for the inference of 3D chromosomal structure of individual topologically associated domain from Hi-C data. When applied to a real Hi-C dataset, the piecewise helical model not only achieves much better model fitting than existing models, but also reveals that geometric properties of chromatin spatial organization are closely related to genome function.
Ming Hu 0001, Zhaohui S. Qin, Jun S. Liu
IEEE ACM Trans. Comput. Biol. Bioinform.2
2019 MAPS: Model-based analysis of long-range chromatin interactions from PLAC-seq and HiChIP experiments
abstract
Hi-C and chromatin immunoprecipitation (ChIP) have been combined to identify long-range chromatin interactions genome-wide at reduced cost and enhanced resolution, but extracting information from the resulting datasets has been challenging. Here we describe a computational method, MAPS, Model-based Analysis of PLAC-seq and HiChIP, to process the data from such experiments and identify long-range chromatin interactions. MAPS adopts a zero-truncated Poisson regression framework to explicitly remove systematic biases in the PLAC-seq and HiChIP datasets, and then uses the normalized chromatin contact frequencies to identify significant chromatin interactions anchored at genomic regions bound by the protein of interest. MAPS shows superior performance over existing software tools in the analysis of chromatin interactions from multiple PLAC-seq and HiChIP datasets centered on different transcriptional factors and histone marks. MAPS is freely available at https://github.com/ijuric/MAPS.
Ivan Juric, Armen Abnousi, Ramya Raviram, Rongxin Fang, Yan-Xiao Zhang, Yunjiang Qiu, Ming Hu 0001
PLoS Comput. Biol.12
2018 DIMM-SC: a Dirichlet mixture model for clustering droplet-based single cell transcriptomic data
abstract
Motivation: Single cell transcriptome sequencing (scRNA-Seq) has become a revolutionary tool to study cellular and molecular processes at single cell resolution. Among existing technologies, the recently developed droplet-based platform enables efficient parallel processing of thousands of single cells with direct counting of transcript copies using Unique Molecular Identifier (UMI). Despite the technology advances, statistical methods and computational tools are still lacking for analyzing droplet-based scRNA-Seq data. Particularly, model-based approaches for clustering large-scale single cell transcriptomic data are still under-explored. Results: We developed DIMM-SC, a Dirichlet Mixture Model for clustering droplet-based Single Cell transcriptomic data. This approach explicitly models UMI count data from scRNA-Seq experiments and characterizes variations across different cell clusters via a Dirichlet mixture prior. We performed comprehensive simulations to evaluate DIMM-SC and compared it with existing clustering methods such as K-means, CellTree and Seurat. In addition, we analyzed public scRNA-Seq datasets with known cluster labels and in-house scRNA-Seq datasets from a study of systemic sclerosis with prior biological knowledge to benchmark and validate DIMM-SC. Both simulation studies and real data applications demonstrated that overall, DIMM-SC achieves substantially improved clustering accuracy and much lower clustering variability compared to other existing clustering methods. More importantly, as a model-based approach, DIMM-SC is able to quantify the clustering uncertainty for each single cell, facilitating rigorous statistical inference and biological interpretations, which are typically unavailable from existing clustering methods. Availability and implementation: DIMM-SC has been implemented in a user-friendly R package with a detailed tutorial available on www.pitt.edu/∼wec47/singlecell.html. Contact: [email protected] or [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online.
Zhe Sun 0007, Ting Wang 0003, Robert Lafyatis, Ying Ding 0003, Ming Hu 0001, Wei Chen 0074
Bioinform.7
2017 HUGIn: Hi-C Unifying Genomic Interrogator
abstract
MOTIVATION: High throughput chromatin conformation capture (3C) technologies, such as Hi-C and ChIA-PET, have the potential to elucidate the functional roles of non-coding variants. However, most of published genome-wide unbiased chromatin organization studies have used cultured cell lines, limiting their generalizability. RESULTS: We developed a web browser, HUGIn, to visualize Hi-C data generated from 21 human primary tissues and cell lines. HUGIn enables assessment of chromatin contacts both constitutive across and specific to tissue(s) and/or cell line(s) at any genomic loci, including GWAS SNPs, eQTLs and cis-regulatory elements, facilitating the understanding of both GWAS and eQTL results and functional genomics data. AVAILABILITY AND IMPLEMENTATION: HUGIn is available at http://yunliweb.its.unc.edu/HUGIn. CONTACT: [email protected] or [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Joshua S. Martin, Zheng Xu 0010, Alex P. Reiner, Karen L. Mohlke, Patrick F. Sullivan, Ming Hu 0001
Bioinform.7
2016 A hidden Markov random field-based Bayesian method for the detection of long-range chromosomal interactions in Hi-C data
abstract
MOTIVATION: Advances in chromosome conformation capture and next-generation sequencing technologies are enabling genome-wide investigation of dynamic chromatin interactions. For example, Hi-C experiments generate genome-wide contact frequencies between pairs of loci by sequencing DNA segments ligated from loci in close spatial proximity. One essential task in such studies is peak calling, that is, detecting non-random interactions between loci from the two-dimensional contact frequency matrix. Successful fulfillment of this task has many important implications including identifying long-range interactions that assist interpreting a sizable fraction of the results from genome-wide association studies. The task - distinguishing biologically meaningful chromatin interactions from massive numbers of random interactions - poses great challenges both statistically and computationally. Model-based methods to address this challenge are still lacking. In particular, no statistical model exists that takes the underlying dependency structure into consideration. RESULTS: In this paper, we propose a hidden Markov random field (HMRF) based Bayesian method to rigorously model interaction probabilities in the two-dimensional space based on the contact frequency matrix. By borrowing information from neighboring loci pairs, our method demonstrates superior reproducibility and statistical power in both simulation studies and real data analysis. AVAILABILITY AND IMPLEMENTATION: The Source codes can be downloaded at: http://www.unc.edu/∼yunmli/HMRFBayesHiC CONTACT: [email protected] or [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Zheng Xu 0010, Fulai Jin, Mengjie Chen, Terrence S. Furey, Patrick F. Sullivan, Zhaohui S. Qin, Ming Hu 0001
Bioinform.8
2016 FastHiC: a fast and accurate algorithm to detect long-range chromosomal interactions from Hi-C data
abstract
MOTIVATION: How chromatin folds in three-dimensional (3D) space is closely related to transcription regulation. As powerful tools to study such 3D chromatin conformation, the recently developed Hi-C technologies enable a genome-wide measurement of pair-wise chromatin interaction. However, methods for the detection of biologically meaningful chromatin interactions, i.e. peak calling, from Hi-C data, are still under development. In our previous work, we have developed a novel hidden Markov random field (HMRF) based Bayesian method, which through explicitly modeling the non-negligible spatial dependency among adjacent pairs of loci manifesting in high resolution Hi-C data, achieves substantially improved robustness and enhanced statistical power in peak calling. Superior to peak callers that ignore spatial dependency both methodologically and in performance, our previous Bayesian framework suffers from heavy computational costs due to intensive computation incurred by modeling the correlated peak status of neighboring loci pairs and the inference of hidden dependency structure. RESULTS: In this work, we have developed FastHiC, a novel approach based on simulated field approximation, which approximates the joint distribution of the hidden peak status by a set of independent random variables, leading to more tractable computation. Performance comparisons in real data analysis showed that FastHiC not only speeds up our original Bayesian method by more than five times, bus also achieves higher peak calling accuracy. AVAILABILITY AND IMPLEMENTATION: FastHiC is freely accessible at:http://www.unc.edu/∼yunmli/FastHiC/ CONTACTS: : [email protected] or [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Zheng Xu 0010, Ming Hu 0001
Bioinform.5
2013 Bayesian Inference of Spatial Organizations of Chromosomes
abstract
Knowledge of spatial chromosomal organizations is critical for the study of transcriptional regulation and other nuclear processes in the cell. Recently, chromosome conformation capture (3C) based technologies, such as Hi-C and TCC, have been developed to provide a genome-wide, three-dimensional (3D) view of chromatin organization. Appropriate methods for analyzing these data and fully characterizing the 3D chromosomal structure and its structural variations are still under development. Here we describe a novel Bayesian probabilistic approach, denoted as "Bayesian 3D constructor for Hi-C data" (BACH), to infer the consensus 3D chromosomal structure. In addition, we describe a variant algorithm BACH-MIX to study the structural variations of chromatin in a cell population. Applying BACH and BACH-MIX to a high resolution Hi-C dataset generated from mouse embryonic stem cells, we found that most local genomic regions exhibit homogeneous 3D chromosomal structures. We further constructed a model for the spatial arrangement of chromatin, which reveals structural properties associated with euchromatic and heterochromatic regions in the genome. We observed strong associations between structural properties and several genomic and epigenetic features of the chromosome. Using BACH-MIX, we further found that the structural variations of chromatin are correlated with these genomic and epigenetic features. Our results demonstrate that BACH and BACH-MIX have the potential to provide new insights into the chromosomal architecture of mammalian cells.
Ming Hu 0001, Zhaohui S. Qin, Jesse R. Dixon, Siddarth Selvaraj, Jennifer Fang, Jun S. Liu
PLoS Comput. Biol.1
2012 HiCNorm: removing biases in Hi-C data via Poisson regression
abstract
SUMMARY: We propose a parametric model, HiCNorm, to remove systematic biases in the raw Hi-C contact maps, resulting in a simple, fast, yet accurate normalization procedure. Compared with the existing Hi-C normalization method developed by Yaffe and Tanay, HiCNorm has fewer parameters, runs >1000 times faster and achieves higher reproducibility. AVAILABILITY: Freely available on the web at: http://www.people.fas.harvard.edu/∼junliu/HiCNorm/. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Ming Hu 0001, Siddarth Selvaraj, Zhaohui S. Qin, Jun S. Liu
Bioinform.1
2012 Using Poisson mixed-effects model to quantify transcript-level gene expression in RNA-Seq
abstract
MOTIVATION: RNA sequencing (RNA-Seq) is a powerful new technology for mapping and quantifying transcriptomes using ultra high-throughput next-generation sequencing technologies. Using deep sequencing, gene expression levels of all transcripts including novel ones can be quantified digitally. Although extremely promising, the massive amounts of data generated by RNA-Seq, substantial biases and uncertainty in short read alignment pose challenges for data analysis. In particular, large base-specific variation and between-base dependence make simple approaches, such as those that use averaging to normalize RNA-Seq data and quantify gene expressions, ineffective. RESULTS: In this study, we propose a Poisson mixed-effects (POME) model to characterize base-level read coverage within each transcript. The underlying expression level is included as a key parameter in this model. Since the proposed model is capable of incorporating base-specific variation as well as between-base dependence that affect read coverage profile throughout the transcript, it can lead to improved quantification of the true underlying expression level. AVAILABILITY AND IMPLEMENTATION: POME can be freely downloaded at http://www.stat.purdue.edu/~yuzhu/pome.html. CONTACT: [email protected]; [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Ming Hu 0001, Jeremy M. G. Taylor, Jun S. Liu, Zhaohui S. Qin
Bioinform.1
2010 HPeak: an HMM-based algorithm for defining read-enriched regions in ChIP-Seq data
abstract
BACKGROUND: Protein-DNA interaction constitutes a basic mechanism for the genetic regulation of target gene expression. Deciphering this mechanism has been a daunting task due to the difficulty in characterizing protein-bound DNA on a large scale. A powerful technique has recently emerged that couples chromatin immunoprecipitation (ChIP) with next-generation sequencing, (ChIP-Seq). This technique provides a direct survey of the cistrom of transcription factors and other chromatin-associated proteins. In order to realize the full potential of this technique, increasingly sophisticated statistical algorithms have been developed to analyze the massive amount of data generated by this method. RESULTS: Here we introduce HPeak, a Hidden Markov model (HMM)-based Peak-finding algorithm for analyzing ChIP-Seq data to identify protein-interacting genomic regions. In contrast to the majority of available ChIP-Seq analysis software packages, HPeak is a model-based approach allowing for rigorous statistical inference. This approach enables HPeak to accurately infer genomic regions enriched with sequence reads by assuming realistic probability distributions, in conjunction with a novel weighting scheme on the sequencing read coverage. CONCLUSIONS: Using biologically relevant data collections, we found that HPeak showed a higher prevalence of the expected transcription factor binding motifs in ChIP-enriched sequences relative to the control sequences when compared to other currently available ChIP-Seq analysis approaches. Additionally, in comparison to the ChIP-chip assay, ChIP-Seq provides higher resolution along with improved sensitivity and specificity of binding site detection. Additional file and the HPeak program are freely available at http://www.sph.umich.edu/csg/qin/HPeak.
Zhaohui S. Qin, Jianjun Yu, Jincheng Shen, Christopher A. Maher, Ming Hu 0001, Shanker Kalyana-Sundaram, Jindan Yu, Arul M. Chinnaiyan
BMC Bioinform.5