Bai Zhang

dblp:46/2754 · DBLP profile ↗
← Back
35ranked-venue papers
9as first author
6since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 31 · 6 first-author · 5 since 2021Artificial intelligence and machine learning · 3 · 3 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
YearPublicationVenuePosition
2025 Microbe-Drug Association Prediction Model Based on Adaptive Network Fusion of Structural-Topological Information With Integration Strategy
abstract
Currently, microbial drug resistance has become a serious global problem. Investigating the relationship between microbes and drugs can facilitate understanding of microbial resistance to certain drugs and help develop more effective therapeutic strategies. Therefore, an efficient and rapid computational method for identifying microbe-drug associations is very important for optimizing antimicrobial therapy, monitoring drug resistance, and addressing bacterial resistance challenges. In this paper, a model called ANFAISMDA for identifying potential microbe-drug associations is proposed. First, Structural information of microbes and drugs is obtained based on 16S rRNA gene sequences of microbes and SMILES structures of drugs, respectively. Topological information of microbes and drugs is obtained based on symmetric matrix completion algorithm, respectively. Then, the structural and topological information is effectively fused according to our proposed adaptive network fusion algorithm. The algorithm combines network node and path information while also adaptively updating weights based on inter-network transmissibility. Finally, an integration strategy is proposed to significantly improve model performance and thereby more accurately identify novel microbe-drug associations. Overall, the reliability and validity of the model are fully verified by experiments. Further, we also have visualized the top 50 microbe-drug associations and have found some interesting association patterns. The results of the study show the importance of our approach in optimizing antimicrobial therapy, monitoring drug resistance, and addressing the challenges of bacterial resistance. In addition, the study confirms the usefulness of the model in gaining a deeper understanding of the complex mechanisms of microbe-drug associations.
Liugen Wang, Bai Zhang, Hanwen Wu, Yongle Shi, Yibing Ma
IEEE Trans. Comput. Biol. Bioinform.2
2024 DDN3.0: determining significant rewiring of biological network structure with differential dependency networks
abstract
MOTIVATION: Complex diseases are often caused and characterized by misregulation of multiple biological pathways. Differential network analysis aims to detect significant rewiring of biological network structures under different conditions and has become an important tool for understanding the molecular etiology of disease progression and therapeutic response. With few exceptions, most existing differential network analysis tools perform differential tests on separately learned network structures that are computationally expensive and prone to collapse when grouped samples are limited or less consistent. RESULTS: We previously developed an accurate differential network analysis method-differential dependency networks (DDN), that enables joint learning of common and rewired network structures under different conditions. We now introduce the DDN3.0 tool that improves this framework with three new and highly efficient algorithms, namely, unbiased model estimation with a weighted error measure applicable to imbalance sample groups, multiple acceleration strategies to improve learning efficiency, and data-driven determination of proper hyperparameters. The comparative experimental results obtained from both realistic simulations and case studies show that DDN3.0 can help biologists more accurately identify, in a study-specific and often unknown conserved regulatory circuitry, a network of significantly rewired molecular players potentially responsible for phenotypic transitions. AVAILABILITY AND IMPLEMENTATION: The Python package of DDN3.0 is freely available at https://github.com/cbil-vt/DDN3. A user's guide and a vignette are provided at https://ddn-30.readthedocs.io/.
Yingzhou Lu, Yizhi Wang 0009, Bai Zhang, Guoqiang Yu, Chunyu Liu 0001, Robert Clarke, David M. Herrington, Yue Joseph Wang
Bioinform.4
2023 Predicting potential microbe-disease associations based on multi-source features and deep learning
abstract
Studies have confirmed that the occurrence of many complex diseases in the human body is closely related to the microbial community, and microbes can affect tumorigenesis and metastasis by regulating the tumor microenvironment. However, there are still large gaps in the clinical observation of the microbiota in disease. Although biological experiments are accurate in identifying disease-associated microbes, they are also time-consuming and expensive. The computational models for effective identification of diseases related microbes can shorten this process, and reduce capital and time costs. Based on this, in the paper, a model named DSAE_RF is presented to predict latent microbe-disease associations by combining multi-source features and deep learning. DSAE_RF calculates four similarities between microbes and diseases, which are then used as feature vectors for the disease-microbe pairs. Later, reliable negative samples are screened by k-means clustering, and a deep sparse autoencoder neural network is further used to extract effective features of the disease-microbe pairs. In this foundation, a random forest classifier is presented to predict the associations between microbes and diseases. To assess the performance of the model in this paper, 10-fold cross-validation is implemented on the same dataset. As a result, the AUC and AUPR of the model are 0.9448 and 0.9431, respectively. Furthermore, we also conduct a variety of experiments, including comparison of negative sample selection methods, comparison with different models and classifiers, Kolmogorov-Smirnov test and t-test, ablation experiments, robustness analysis, and case studies on Covid-19 and colorectal cancer. The results fully demonstrate the reliability and availability of our model.
Liugen Wang, Yan Wang 0092, Chenxu Xuan, Bai Zhang, Hanwen Wu
Briefings Bioinform.4
2023 P-CSN: single-cell RNA sequencing data analysis by partial cell-specific network
abstract
Although many single-cell computational methods proposed use gene expression as input, recent studies show that replacing 'unstable' gene expression with 'stable' gene-gene associations can greatly improve the performance of downstream analysis. To obtain accurate gene-gene associations, conditional cell-specific network method (c-CSN) filters out the indirect associations of cell-specific network method (CSN) based on the conditional independence of statistics. However, when there are strong connections in networks, the c-CSN suffers from false negative problem in network construction. To overcome this problem, a new partial cell-specific network method (p-CSN) based on the partial independence of statistics is proposed in this paper, which eliminates the singularity of the c-CSN by implicitly including direct associations among estimated variables. Based on the p-CSN, single-cell network entropy (scNEntropy) is further proposed to quantify cell state. The superiorities of our method are verified on several datasets. (i) Compared with traditional gene regulatory network construction methods, the p-CSN constructs partial cell-specific networks, namely, one cell to one network. (ii) When there are strong connections in networks, the p-CSN reduces the false negative probability of the c-CSN. (iii) The input of more accurate gene-gene associations further optimizes the performance of downstream analyses. (iv) The scNEntropy effectively quantifies cell state and reconstructs cell pseudo-time.
Chenxu Xuan, Hanwen Wu, Bai Zhang
Briefings Bioinform.4
2023 Towards fidelity of graph data augmentation via equivariance
Bai Zhang, Yixing Gao 0001, Linbo Xie, Xiaofeng Cao 0002, Yixiang Shan, Jielong Yang
Knowl. Based Syst.1
2023 scGAMNN: Graph Antoencoder-Based Single-Cell RNA Sequencing Data Integration Algorithm Using Mutual Nearest Neighbors
abstract
It is critical to correctly assemble high-dimensional single-cell RNA sequencing (scRNA-seq) datasets and downscale them for downstream analysis. However, given the complex relationships between cells, it remains a challenge to simultaneously eliminate batch effects between datasets and maintain the topology between cells within each dataset. Here, we propose scGAMNN, a deep learning model based on graph autoencoder, to simultaneously achieve batch correction and topology-preserving dimensionality reduction. The low-dimensional integrated data obtained by scGAMNN can be used for visualization, clustering and trajectory inference.By comparing it with the other five methods, multiple tasks show that scGAMNN consistently has comparable data integration performance in clustering and trajectory conservation.
Bai Zhang, Hanwen Wu, Yan Wang 0092, Chenxu Xuan
IEEE J. Biomed. Health Informatics1
2015 KDDN: an open-source Cytoscape app for constructing differential dependency networks with significant rewiring
abstract
UNLABELLED: We have developed an integrated molecular network learning method, within a well-grounded mathematical framework, to construct differential dependency networks with significant rewiring. This knowledge-fused differential dependency networks (KDDN) method, implemented as a Java Cytoscape app, can be used to optimally integrate prior biological knowledge with measured data to simultaneously construct both common and differential networks, to quantitatively assign model parameters and significant rewiring p-values and to provide user-friendly graphical results. The KDDN algorithm is computationally efficient and provides users with parallel computing capability using ubiquitous multi-core machines. We demonstrate the performance of KDDN on various simulations and real gene expression datasets, and further compare the results with those obtained by the most relevant peer methods. The acquired biologically plausible results provide new insights into network rewiring as a mechanistic principle and illustrate KDDN's ability to detect them efficiently and correctly. Although the principal application here involves microarray gene expressions, our methodology can be readily applied to other types of quantitative molecular profiling data. AVAILABILITY: Source code and compiled package are freely available for download at http://apps.cytoscape.org/apps/kddn. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Bai Zhang, Eric P. Hoffman, Robert Clarke, Ie-Ming Shih, Jianhua Xuan, David M. Herrington, Yue Joseph Wang
Bioinform.2
2014 Ensemble random projection for multi-label classification with application to protein subcellular localization
abstract
The curse of dimensionality severely restricts the predictive power of multi-label classification systems. High-dimensional feature vectors may contain redundant or irrelevant information, causing the classification systems suffer from overfitting. To address this problem, this paper proposes a dimensionality-reduction method that applies random projection (RP) to construct an ensemble of multilabel classifiers. The merits of the proposed method are demonstrated through a multi-label protein classification task. Specifically, high-dimensional feature vectors are extracted from protein sequences using the gene ontology (GO) and Swiss-Prot databases. The feature vectors are then projected onto lower-dimensional spaces by random projection matrices whose elements conform to a distribution with zero mean and unit variance. The transformed low-dimensional vectors are classified by an ensemble of one-vs-rest multi-label support vector machine (SVM) classifiers, each corresponding to one of the RP matrices. The scores obtained from the ensemble are then fused for predicting the subcellular localization of proteins. Experimental results suggest that the proposed method can reduce the dimensions by seven folds and impressively improve the classification performance.
Shibiao Wan, Man-Wai Mak, Bai Zhang, Yue Joseph Wang, Sun-Yuan Kung
ICASSP3
2014 AISAIC: a software suite for accurate identification of significant aberrations in cancers
abstract
UNLABELLED: Accurate identification of significant aberrations in cancers (AISAIC) is a systematic effort to discover potential cancer-driving genes such as oncogenes and tumor suppressors. Two major confounding factors against this goal are the normal cell contamination and random background aberrations in tumor samples. We describe a Java AISAIC package that provides comprehensive analytic functions and graphic user interface for integrating two statistically principled in silico approaches to address the aforementioned challenges in DNA copy number analyses. In addition, the package provides a command-line interface for users with scripting and programming needs to incorporate or extend AISAIC to their customized analysis pipelines. This open-source multiplatform software offers several attractive features: (i) it implements a user friendly complete pipeline from processing raw data to reporting analytic results; (ii) it detects deletion types directly from copy number signals using a Bayes hypothesis test; (iii) it estimates the fraction of normal contamination for each sample; (iv) it produces unbiased null distribution of random background alterations by iterative aberration-exclusive permutations; and (v) it identifies significant consensus regions and the percentage of homozygous/hemizygous deletions across multiple samples. AISAIC also provides users with a parallel computing option to leverage ubiquitous multicore machines. AVAILABILITY AND IMPLEMENTATION: AISAIC is available as a Java application, with a user's guide and source code, at https://code.google.com/p/aisaic/.
Bai Zhang, Xuchu Hou, Xiguo Yuan, Ie-Ming Shih, Robert Clarke, Roger R. Wang, Subha Madhavan, Yue Joseph Wang, Guoqiang Yu
Bioinform.1
2014 Integration of Network Biology and Imaging to Study Cancer Phenotypes and Responses
abstract
Ever growing "omics" data and continuously accumulated biological knowledge provide an unprecedented opportunity to identify molecular biomarkers and their interactions that are responsible for cancer phenotypes that can be accurately defined by clinical measurements such as in vivo imaging. Since signaling or regulatory networks are dynamic and context-specific, systematic efforts to characterize such structural alterations must effectively distinguish significant network rewiring from random background fluctuations. Here we introduced a novel integration of network biology and imaging to study cancer phenotypes and responses to treatments at the molecular systems level. Specifically, Differential Dependence Network (DDN) analysis was used to detect statistically significant topological rewiring in molecular networks between two phenotypic conditions, and in vivo Magnetic Resonance Imaging (MRI) was used to more accurately define phenotypic sample groups for such differential analysis. We applied DDN to analyze two distinct phenotypic groups of breast cancer and study how genomic instability affects the molecular network topologies in high-grade ovarian cancer. Further, FDA-approved arsenic trioxide (ATO) and the ND2-SmoA1 mouse model of Medulloblastoma (MB) were used to extend our analyses of combined MRI and Reverse Phase Protein Microarray (RPMA) data to assess tumor responses to ATO and to uncover the complexity of therapeutic molecular biology.
Sean S. Wang, Olga C. Rodriguez, Emanuel Petricoin III, Ie-Ming Shih, Daniel Chan, Maria Avantaggiati, Guoqiang Yu, Shaozhen Ye, Robert Clarke, Chao Wang 0005, Bai Zhang, Yue Joseph Wang, Chris Albanese
IEEE ACM Trans. Comput. Biol. Bioinform.13
2013 An ensemble classifier with random projection for predicting multi-label protein subcellular localization
abstract
In protein subcellular localization prediction, a predominant scenario is that the number of available features is much larger than the number of data samples. Among the large number of features, many of them may contain redundant or irrelevant information, causing the prediction systems suffer from overfitting. To address this problem, this paper proposes a dimensionality-reduction method that applies random projection (RP) to construct an ensemble multi-label classifier for predicting protein subcellular localization. Specifically, the frequencies of occurrences of gene-ontology terms are used as feature vectors, which are projected onto lower-dimensional spaces by random projection matrices whose elements conform to a distribution with zero mean and unit variance. The transformed low-dimensional vectors are classified by an ensemble of one-vs-rest multi-label support vector machine (SVM) classifiers, each corresponding to one of the RP matrices. The scores obtained from the ensemble are then fused for making the final decision. Experimental results on two recent datasets suggest that the proposed method can reduce the dimensions by six folds and remarkably improve the classification performance.
Shibiao Wan, Man-Wai Mak, Bai Zhang, Yue Joseph Wang, Sun-Yuan Kung
BIBM3
2012 Nonlinear System Modeling With Random Matrices: Echo State Networks Revisited
abstract
Echo state networks (ESNs) are a novel form of recurrent neural networks (RNNs) that provide an efficient and powerful computational model approximating nonlinear dynamical systems. A unique feature of an ESN is that a large number of neurons (the "reservoir") are used, whose synaptic connections are generated randomly, with only the connections from the reservoir to the output modified by learning. Why a large randomly generated fixed RNN gives such excellent performance in approximating nonlinear systems is still not well understood. In this brief, we apply random matrix theory to examine the properties of random reservoirs in ESNs under different topologies (sparse or fully connected) and connection weights (Bernoulli or Gaussian). We quantify the asymptotic gap between the scaling factor bounds for the necessary and sufficient conditions previously proposed for the echo state property. We then show that the state transition mapping is contractive with high probability when only the necessary condition is satisfied, which corroborates and thus analytically explains the observation that in practice one obtains echo states when the spectral radius of the reservoir weight matrix is smaller than 1.
Bai Zhang, David J. Miller 0001, Yue Joseph Wang
IEEE Trans. Neural Networks Learn. Syst.1
2011 BACOM: in silico detection of genomic deletion types and correction of normal cell contamination in copy number data
abstract
MOTIVATION: Identification of somatic DNA copy number alterations (CNAs) and significant consensus events (SCEs) in cancer genomes is a main task in discovering potential cancer-driving genes such as oncogenes and tumor suppressors. The recent development of SNP array technology has facilitated studies on copy number changes at a genome-wide scale with high resolution. However, existing copy number analysis methods are oblivious to normal cell contamination and cannot distinguish between contributions of cancerous and normal cells to the measured copy number signals. This contamination could significantly confound downstream analysis of CNAs and affect the power to detect SCEs in clinical samples. RESULTS: We report here a statistically principled in silico approach, Bayesian Analysis of COpy number Mixtures (BACOM), to accurately estimate genomic deletion type and normal tissue contamination, and accordingly recover the true copy number profile in cancer cells. We tested the proposed method on two simulated datasets, two prostate cancer datasets and The Cancer Genome Atlas high-grade ovarian dataset, and obtained very promising results supported by the ground truth and biological plausibility. Moreover, based on a large number of comparative simulation studies, the proposed method gives significantly improved power to detect SCEs after in silico correction of normal tissue contamination. We develop a cross-platform open-source Java application that implements the whole pipeline of copy number analysis of heterogeneous cancer tissues including relevant processing steps. We also provide an R interface, bacomR, for running BACOM within the R environment, making it straightforward to include in existing data pipelines. AVAILABILITY: The cross-platform, stand-alone Java application, BACOM, the R interface, bacomR, all source code and the simulation data used in this article are freely available at authors' web site: http://www.cbil.ece.vt.edu/software.htm.
Guoqiang Yu, Bai Zhang, G. Steven Bova, Ie-Ming Shih, Yue Joseph Wang
Bioinform.2
2011 DDN: a caBIG® analytical tool for differential network analysis
abstract
UNLABELLED: Differential dependency network (DDN) is a caBIG® (cancer Biomedical Informatics Grid) analytical tool for detecting and visualizing statistically significant topological changes in transcriptional networks representing two biological conditions. Developed under caBIG®'s In Silico Research Centers of Excellence (ISRCE) Program, DDN enables differential network analysis and provides an alternative way for defining network biomarkers predictive of phenotypes. DDN also serves as a useful systems biology tool for users across biomedical research communities to infer how genetic, epigenetic or environment variables may affect biological networks and clinical phenotypes. Besides the standalone Java application, we have also developed a Cytoscape plug-in, CytoDDN, to integrate network analysis and visualization seamlessly. AVAILABILITY: The Java and MATLAB source code can be downloaded at the authors' web site http://www.cbil.ece.vt.edu/software.htm.
Bai Zhang, Huai Li, Ie-Ming Shih, Subha Madhavan, Robert Clarke, Eric P. Hoffman, Jianhua Xuan, Leena Hilakivi-Clarke, Yue Joseph Wang
Bioinform.1
2010 Spatial distribution and the possible source of CDOM for inland water in summer in the Northeast China
abstract
The availability of underwater light is a critical factor in the growth and abundance of primary producers in shallow embayments. The aim of this study was to examine the spatial distribution and the possible source of CDOM for inland water in summer in the northeast china. Absorption spectra of inland water samples were measured from 200nm to 800nm. Highest mean-value of a(375) occurred in Chagan Lake. A significant spatial difference was found among four different inland waters, and evident spatial variation was in Chagan Lake. A consistent negative non-linear relationship was recorded between S value and CDOM absorption coefficient. Furthermore, S value was used as a proxy for CDOM composition and source. Fulvic acids is primary contribution for CDOM absorption in Songhua Lake and Shitoukoumen Reservoir, but humic acids in Nanhu Lake and Chagan Lake. The relationships between CDOM absorption and total suspended matter concentration and chlorophyll-a concentration were analyzed. It demonstrated the biological processes source for Nanhu Lake, Shitoukoumen Reservoir and Chagan Lake. But for Songhua Lake, the dominating source is from river inputs, but biological process was also an important portion for CDOM concentration.
Guangjia Jiang, Dianwei Liu, Kaishan Song, Bai Zhang
IGARSS4
2010 Analysis on the spectral reflectance response to snow contaminants in northeast China
abstract
By simulating atmospheric deposition experiment, this paper analyzed the relationship between the measured spectral reflectance and the concentrations of contaminants in the snow. It is found that the visible spectrum is sensitive to snow contaminants. From 350nm to 850nm, with the increase concentrations of contaminants in snow, snow reflectivity dramatically decreases. We get the conclusion that the most sensitive bands to snow contaminants are 384nm, 450nm and 1495nm.Using the non-linear regression method to analyze the relationship between spectral reflectance and the contaminants. The results showed the reflectivity of snow at visible bands logarithmically decreases with the snow contaminants increasing; the R2can reach 0.9.To the contrary, the spectral reflectance at nearinfrared increases with the snow contaminants increasing. Therefore, this method can be combined satellite image to forecast the contaminants in the snow at large-scale.
Xiaochun Lei, Kaishan Song, Zongming Wang, Jia Du, Yanqing Wu, Xuguang Tang, Lihong Zeng, Guangjia Jiang, Dianwei Liu, Bai Zhang
IGARSS11
2010 Retrival of total suspended matter (TSM) using remotely sensed images in Shitoukoumen Reservior, Northeast China
abstract
The Shitoukoumen Reservoir is the major drinking water resources for the Changchun metropolitan region. The concentration of total suspended matter (TSM) is major water quality parameter that could be retrieved using remotely sensed data. 225 samples were analyzed during 12 times of field works from April 2006 to September 2008, in which the field work conducted on 17thJuly 2007 and 13thSeptember 2008 were concurrent with IRS-P6 satellite over pass. Empirical regression models were established to analysis the relationship between TSM and satellite-received radiances. It was found that the regression model performed well on the TSM concentration estimation with higher accuracy (R2= 0.94, 0.91) with IRS-P6 visible and near infrared bands as inputs. The high concentration of TSM made it the dominant upwelling signature captured by IRS-P6 satellite data, which may explain why the regression model relatively accurate. The RMSE for the TSM was less than 15%. Future work will need to be undertaken to refine model for Shitoukoumen Reservoir water quality monitoring with remotely sensed images or hyperspectral imaging data is to be launched in the near future.
Kaishan Song, Dongmei Lu, Dianwei Liu, Zongming Wang, Lin Li 0006, Bai Zhang
IGARSS6
2010 Application of PCA and canopy near, shortwave-infrared bands for soybean and corn FPAR estimation in the Songnen Plain, China
abstract
The fraction of photosynthetically active radiation (FPAR) absorbed by global vegetation is a key state variable in most ecosystem productivity models and in global models of climate, hydrology, biogeochemistry, and ecology. Therefore, how accurately retrieve FPAR will directly influence the estimation of many models and requires special attention. In this paper, based on the ground truth data in the Songnen Plain of China, we studied the correlations between FPAR and the corresponding vegetation indices. Comparing with NDVI and RVI (calculated by visible and near-infrared band), NDSI and RSI were also constructed (calculated by near, shortwave infrared band) to estimate FPAR. All vegetation indices were under the best wavelength combinations. PCA approach was also introduced for extracting hyperspectral reflectance information and estimating FPAR. The research results indicated that NDSI and RSI calculated by the near, shortwave infrared bands (R2of the validating models were 0.74 and 0.69 and RMSE were 0.108 and 0.171, respectively) showed better performance than NDVI and RVI that was computed by the visible and near-infrared bands (R2of the validating models were 0.71 and 0.65 and RMSE were 0.187 and 0.213, respectively). PCA approach could compress the hyperspectral reflectance information effectively, and showed better performance for FPAR estimating. From the above study, it also suggested that shortwave infrared bands had great potential for the estimation of FPAR.
Xuguang Tang, Kaishan Song, Dianwei Liu, Zongming Wang, Bai Zhang
IGARSS5
2010 Improved models for Chla estimation by considering the effect of phytoplankton specific absorption
abstract
In remote chlorophyll-a (Chla) retrieval in Case-II waters, the uncertainties from Chla specific absorption coefficient a*phhas a significant effect on the retrieving accuracy. In this paper, we presented newly improved three-band model and four-band model by a case study in Shitoukoumen Reservoir. The proposed three-band model can be expressed as [Rrs-1(λ1)-Rrs-1(λ2)]×Rrs(λ3)×a*ph-1(λ1), while the improved four-band model is [Rrs-1(λ1)-Rrs-1(λ2)]×[Rrs-1(λ4)×Rrs-1(λ3)]-1×a*ph-1(λ1). Results showed that the latter one was slightly superior to the former one. Comparison with original models without correction of a*ph, the improved models achieved higher precision and stability. The findings underlined the rationale behind the proposed models and demonstrated a potentially use for assessing Chla in Case-II waters.
Jingping Xu, Xingfa Gu, Bai Zhang, Tao Yu 0001
IGARSS3
2010 Spatial mapping of actual evapotranspiration and water deficit with MODIS products in the Songnen Plain, northeast China
abstract
Analysis of spatial patterns of evapotranspiration (ET) and water deficit (WD) is significant in the evaluation of crop growth status and water use efficiency for the Songnen Plain, an important commodity grain product base of China. Spatial patterns of ET and WD in the Songnen Plain of 2008 growing season (from May to September) were mapped by using MODIS products and meteorological data. The results indicated that ET and WD exhibited obvious spatial variation and gradually increased from southwest to east and northeast. Total ET over the Songnen Plain during the 2008 growing season ranged from 182.7 mm to 1002.4 mm with the mean value of 591.1 mm, and WD ranged from -163.0 mm to 645.9 mm with the mean value of 195.9 mm. Average ET and WD for different land covers varied significantly, water-body and wetlands obtained the highest ET and WD values, while grassland got the lowest ET and WD values. Through this study, it would provide some supports for the assessment of crop growth in arid environments of Songnen Plain.
Lihong Zeng, Kaishan Song, Bai Zhang, Zongming Wang
IGARSS3
2010 Learning Structural Changes of Gaussian Graphical Models in Controlled Experiments
Bai Zhang, Yue Joseph Wang
UAI1
2009 Accurate Estimation of Genomic Deletions and Normal Cell Contamination by Bayesian Analysis of Mixtures
abstract
Copy number change is an important form of structural variation in human genomes. Somatic copy number alterations can cause the acquisition of oncogenes and loss of tumor suppressor genes in tumorigenesis. Recent development of SNP array technology facilitates studies on copy number changes in a genome-wide scale with high resolution. However, tumor samples often consist of mixed cancer and normal cells. Such tissue heterogeneity poses as a serious hurdle to analyzing copy number changes and could confound subsequent marker identification and diagnostic classification rooted in specific cells. We report here a statistically-principled in silico approach to accurately estimate genomic deletions and normal tissue contamination, and accordingly recover the true copy number profile in cancer cells. We tested the proposed method on three simulation and one real datasets and obtained highly promising results validated by the ground truth and figure of merit. We expect this newly developed method to be a useful tool in routine copy number analysis of heterogeneous tissues.
Guoqiang Yu, Bai Zhang, Ie-Ming Shih, Yue Joseph Wang
BIBM2
2009 Land Use/Cover Characterization with MODIS Time Series Data with Hybrid Classification Mothed over Australia for 2001 and 2003
abstract
Improved and up-to-date land use/land cover (LULC) data sets are needed over the whole country of Australia to support science and policy applications focused on understanding the role and response of the LULC to environmental change. The main goal of this study was to map LULC in Australia using MODIS 250 m Normalized Difference Vegetation Index (NDVI), Land Surface Vegetation Index (LSWI) and reflectance time series data of 2000 and 2003. NDVI time-series were filtered by the Savitzky-Golay algorithm in the present study to smooth out noise. A combination of unsupervised ISODATA and a hierarchical decision tree classification were performed on 2 years 12-month time-series MODIS data. Also, Australian Vegetation Map and other land use/land cover data set were used as labeling reference during the classification process. The MODIS land cover products were evaluated using existing land use/cover data derived from Landsat TM as reference data (AUS-2000), also LULC information derived from 11 scenes of Landsat-5 TM data were used as validation data source. The overall classification accuracy was 76.4%. It turned out that our result is acceptable because the relative high resolution of MODIS data and more prior knowledge was applied.
Kaishan Song, Mohsin Hafeez, Zongming Wang, Dongmei Lu, Lihong Zeng, Dianwei Liu, Bai Zhang, Jia Du, Qingfeng Liu
IGARSS (3)8
2009 Land Use/land Cover (LULC) Characterizaitoin with MODIS Time Series Data in the Amu River Basin
abstract
Improved and up-to-date land use/land cover (LULC) data sets are needed over intensively land use/cover change area in the Amur River Basin (ARB) to support science and policy applications focused on understanding of the role and response of the LULC to environmental change issues. The main goal of this study was to map LULC in the Amur River Basin using MODIS 250 m Normalized Difference Vegetation Index (NDVI) and Land Surface Water Index (LSWI) time series data in 2001 and 2007. A combination of unsupervised ISODATA and hierarchical decision tree classification were performed on 12-month time-series of MODIS NDVI data over the study region. The MODIS land cover result of Northeast China was evaluated using existing land use/cover data, and the rest part was evaluated by LULC information derived from LANDSAT-TM. MODIS 250m NDVI, LSWI and reflectance datasets were found to have sufficient spatial, spectral, and temporal resolutions to detect unique multitemporal signatures of the major land cover types over the region. The overall classification accuracy was 0.81 and the kappa coefficient is 0.64. In conclusion, this method has been used successively for LULC change monitoring in the year 2001 and 2007. The result indicate that MODIS 250 NDVI time series data can derive relatively accurate LULC information for hydrological and climate modeling.
Kaishan Song, Zongming Wang, Qingfeng Liu, Dongmei Lu, Lihong Zeng, Dianwei Liu, Bai Zhang, Jia Du
IGARSS (4)8
2009 Differential dependency network analysis to identify condition-specific topological changes in biological networks
abstract
MOTIVATION: Significant efforts have been made to acquire data under different conditions and to construct static networks that can explain various gene regulation mechanisms. However, gene regulatory networks are dynamic and condition-specific; under different conditions, networks exhibit different regulation patterns accompanied by different transcriptional network topologies. Thus, an investigation on the topological changes in transcriptional networks can facilitate the understanding of cell development or provide novel insights into the pathophysiology of certain diseases, and help identify the key genetic players that could serve as biomarkers or drug targets. RESULTS: Here, we report a differential dependency network (DDN) analysis to detect statistically significant topological changes in the transcriptional networks between two biological conditions. We propose a local dependency model to represent the local structures of a network by a set of conditional probabilities. We develop an efficient learning algorithm to learn the local dependency model using the Lasso technique. A permutation test is subsequently performed to estimate the statistical significance of each learned local structure. In testing on a simulation dataset, the proposed algorithm accurately detected all the genes with network topological changes. The method was then applied to the estrogen-dependent T-47D estrogen receptor-positive (ER+) breast cancer cell line datasets and human and mouse embryonic stem cell datasets. In both experiments using real microarray datasets, the proposed method produced biologically meaningful results. We expect DDN to emerge as an important bioinformatics tool in transcriptional network analyses. While we focus specifically on transcriptional networks, the DDN method we introduce here is generally applicable to other biological networks with similar characteristics. AVAILABILITY: The DDN MATLAB toolbox and experiment data are available at http://www.cbil.ece.vt.edu/software.htm.
Bai Zhang, Huai Li, Rebecca B. Riggins, Ming Zhan, Jianhua Xuan, Eric P. Hoffman, Robert Clarke, Yue Joseph Wang
Bioinform.1
2006 An Algorithm on Extraction of Saline-Alkalized Land by Image Segmentation Based on ETM+ Image
Jianping Li 0002, Bai Zhang, Shuqing Zhang, Zongming Wang
ICCSA (1)2
2006 Study on Automatic Extraction of Corn Fields Information on Remotely Sensed Imagery Based on Multi-characters Space
abstract
Many researches have focused on the automatic extraction of thematic information from Landsat/TM remotely sensed imagery. A method was put forward for automatic thematic information extraction based on the multi-characters space in remote sensing images. According to the analysis of the result of classification, it was concluded that the new method adopted in the paper could improve the efficiency of thematic information extraction from remotely sensed images; and then supervised classification was adopted in the Landsat/TM image classification, corn land in the study area was extracted from the Landsat/TM images with the precision of up to 85.5%. The extraction of corn fields in the study area from the images was performed again on the basis of the expert database, and it was found that the interpretation was notably improved with the precision of 92.9%. Comparing this classification result with the traditional visual interpretation, it was concluded that the new method adopted in the paper could improve efficiency of thematic information extraction from the remotely sensed images. The new method was also theoretically significant with providing new thoughts for intellectualized interpretation of remote-sensing images.
Yang Guang, Xianghua Yang, Bai Zhang, Kaishan Song, Zongming Wang
IGARSS3
2005 A GIS-based study on the model for evaluation of direct submerging damage of flood disaster
abstract
IEEE, IEEE Geosci & Remote Sensing Soc, NASA, NOAA, USN Off Res, Japan Aerosp Explorat Agcy, Natl Polar orbiting Operat Environm Satellite Syst, Ball Aerosp & Technologies Corp, Int Union Radio Sci, Elect & Telecommun Res Inst, Korea Sci & Engn Fdn, Korea Natl Tourism Org, Korea Telecommun
Jianping Li 0002, Bai Zhang
IGARSS2
2005 Hyperspectral approaches for detecting the roadside tree chlorophyll content with BP neural networks
Dianwei Liu, Kaishan Song, Bai Zhang
IGARSS3
2005 Establishing a ann model with in-situ hyperspectral data for estimation chlorophyll-a concentrations in Nanhu Lake of Changchun, China
abstract
IEEE, IEEE Geosci & Remote Sensing Soc, NASA, NOAA, USN Off Res, Japan Aerosp Explorat Agcy, Natl Polar orbiting Operat Environm Satellite Syst, Ball Aerosp & Technologies Corp, Int Union Radio Sci, Elect & Telecommun Res Inst, Korea Sci & Engn Fdn, Korea Natl Tourism Org, Korea Telecommun
Kaishan Song, Bai Zhang, Hongtao Duan 0001, Zongming Wang
IGARSS2
2005 Corn chlorophyll estimation with in situ collected hyperspectral reflectance data
abstract
IEEE, IEEE Geosci & Remote Sensing Soc, NASA, NOAA, USN Off Res, Japan Aerosp Explorat Agcy, Natl Polar orbiting Operat Environm Satellite Syst, Ball Aerosp & Technologies Corp, Int Union Radio Sci, Elect & Telecommun Res Inst, Korea Sci & Engn Fdn, Korea Natl Tourism Org, Korea Telecommun
Zongming Wang, Bai Zhang, Kaishan Song, Hongtao Duan 0001
IGARSS2
2005 Study on the relationship between hyperspectral reflectance and soybean LAI, aboveground biomass
abstract
IEEE, IEEE Geosci & Remote Sensing Soc, NASA, NOAA, USN Off Res, Japan Aerosp Explorat Agcy, Natl Polar orbiting Operat Environm Satellite Syst, Ball Aerosp & Technologies Corp, Int Union Radio Sci, Elect & Telecommun Res Inst, Korea Sci & Engn Fdn, Korea Natl Tourism Org, Korea Telecommun
Bai Zhang, Kaishan Song, Yuanzhi Zhang 0003, Zongming Wang
IGARSS1
2003 Analysis on agro-ecological landscape pattern in Phaeozem land area in Northeast China
abstract
The spatial structure of landscape is the result of complicated interaction among nature conditions, living organisms and human activities, which changes reflect the interrelationship among them. With remote sensing (RS) images in the same season of different years, we can get the generalized and visualized information of landscape and its changes. With the help of Geographic Information System (GIS) software and RS image processing software, we interpreted the images and quantified the information of the landscape and its changes. By overlaying the land-use maps, we identified the location of changes and conversion between different land-use types. Phaeozem (black soil) is the main type of cultivated soil, distributed in the shape of belt in the mesa and terrace in the east of Songnen Plain in Northeast China, Phaeozem is rich in organic matters and very fertile, and the reclamation rate is high, which makes Songnen Plain be the main commodities food base, and contribute a lot to the social and economical development of Northeast China and even the whole country. The continuous reclamation for the forest land and grassland to the cultivated land made the landscape simplify, and expansion of the cities and transportation makes the space for the agriculture development limited, since "market economy" implemented at the beginning in 20th century, the comparative profit of paddy field and dry land drives the conversion to paddy field.
Yanfen He, Bai Zhang
IGARSS2
2003 Study on polarizing reflectance characteristics of leaves from four species of deciduous tree at various phenological phase in northeast of China
abstract
Polarized data of four primary deciduous tree leaves in northeast of China are collected with a bi-directional reflectance detection instrument in four phenological phases during 2002. It was found that those tree leaves' polarized reflectance characteristics have direct relation with tree species, leaf property and sensor viewing geometry. This study could provide the theoretical basis for further researches on the remote sensing of polarized light.
Kaishan Song, Bai Zhang, Yunsheng Zhao
IGARSS2
2003 The monitoring of land degradation: the change of saline-alkaline land in Jilin Province of China
abstract
Jilin Province is a major distribution area of saline-alkaline land in China, which mainly distributes in the middle and west part of Jilin Province. By comparing the record in 1958 and in 1981, the area of saline-alkaline land increased greatly. The remote sensing image provided very useful information for the research of the change of saline-alkaline land. Through comparing remote sensing images in same season of different times, the change of saline-alkaline land can be identified. In the research, the Landsat TM images of 1980's and 1990's were used. Taking Landsat remote sensing information as the main source and use multi-time remote sensing information to analyze the change of the saline-alkaline land cover in Jilin Province of China in recent twenty years, we can get both quantity change and quality change of saline-alkaline land, which shows that the increasing trend of the saline-alkaline land cover in Jilin Province of China is obvious. By analyze the driving forces, we can see that natural conditions, such as climate change has part of role, but human activities, such as unreasonable land reclamation and over-pasturing are also the major factors.
Bai Zhang, Haishang Cui, Yanfen He
IGARSS1