EDBT 2026 Demo / reviewers in the wild / expert
Geng Tian
dblp:15/9003
· DBLP profile ↗
20ranked-venue papers
5as first author
14since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 13 · 11 since 2021Computer networks · 2 · 2 first-authorSecurity and privacy · 2 · 2 first-author · 1 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | PDGCL-DTI: Parallel Dual-Channel Graph Contrastive Learning for Drug-Target Binding Prediction in Heterogeneous NetworksabstractPredicting drug-target interactions (DTI) is critical for advancing drug discovery. However, existing DTI approaches struggle with data imbalance and heterogeneous information. This study presents a novel framework called PDGCL-DTI, which leverages two graph contrastive learning frameworks in parallel to capture both local and global features from drug-target heterogeneous networks. First, PDGCL-DTI effectively handles data imbalance through the AdaL-GCL module, which dynamically adjusts the weights of minority class samples to mitigate the impact of the imbalance. Second, by combining local and global contrastive learning, it extracts features from both local node information and global structural information, improving its adaptability to complex heterogeneous networks. This dual strategy enables PDGCL-DTI to exhibit greater robustness and higher prediction accuracy when handling complex DTI data. Experimental results on the ChEMBL, DrugBank, and DAVIS datasets show that PDGCL-DTI outperforms existing DTI methods, achieving an average AUC of 0.958 and an average accuracy of 0.95 across the three datasets. Additionally, case studies demonstrate that PDGCL-DTI successfully predicts interactions between Enasidenib and GABA-AT, as well as Sorafenib and Caspase-3, underscoring its practical applicability in the visualization workflow on the ChEMBL dataset. Qihui Zheng, Xianfang Tang, Yajie Meng, Junlin Xu, Xueying Zeng 0001, Geng Tian, Jialiang Yang |
IEEE J. Biomed. Health Informatics | 6 |
| 2025 | MMsurv: a multimodal multi-instance multi-cancer survival prediction model integrating pathological images, clinical information, and sequencing dataabstractAccurate prediction of patient survival rates in cancer treatment is essential for effective therapeutic planning. Unfortunately, current models often underutilize the extensive multimodal data available, affecting confidence in predictions. This study presents MMSurv, an interpretable multimodal deep learning model to predict survival in different types of cancer. MMSurv integrates clinical information, sequencing data, and hematoxylin and eosin-stained whole-slide images (WSIs) to forecast patient survival. Specifically, we segment tumor regions from WSIs into image tiles and employ neural networks to encode each tile into one-dimensional feature vectors. We then optimize clinical features by applying word embedding techniques, inspired by natural language processing, to the clinical data. To better utilize the complementarity of multimodal data, this study proposes a novel fusion method, multimodal fusion method based on compact bilinear pooling and transformer, which integrates bilinear pooling with Transformer architecture. The fused features are then processed through a dual-layer multi-instance learning model to remove prognosis-irrelevant image patches and predict each patient's survival risk. Furthermore, we employ cell segmentation to investigate the cellular composition within the tiles that received high attention from the model, thereby enhancing its interpretive capacity. We evaluate our approach on six cancer types from The Cancer Genome Atlas. The results demonstrate that utilizing multimodal data leads to higher predictive accuracy compared to using single-modal image data, with an average C-index increase from 0.6750 to 0.7283. Additionally, we compare our proposed baseline model with state-of-the-art methods using the C-index and five-fold cross-validation approach, revealing a significant average improvement of nearly 10% in our model's performance. Shufang Shi, Lijing Liu, Yuhua Yao, Geng Tian, Peizhen Wang, Jialiang Yang |
Briefings Bioinform. | 7 |
| 2025 | Deep learning-based fusion of nuclear segmentation features for microsatellite instability and tumor mutational burden prediction in digestive tract cancers: a multicenter validation studyabstractMicrosatellite instability (MSI) and tumor mutational burden (TMB) are crucial biomarkers in gastric (GC) and colorectal cancer (CRC), yet their conventional sequencing-based detection is costly and time-consuming. Since only ~20% of patients are MSI-high or TMB-high and likely to benefit from immunotherapy, expensive genomic testing is often unjustified. This study developed a deep learning framework to predict MSI and TMB status directly from routinely available Hematoxylin and Eosin (H&E)-stained whole-slide images, leveraging fused nuclear segmentation features to improve accuracy. Using samples from TCGA (350 GC and 376 CRC for MSI; 400 GC and 387 CRC for TMB), image features were extracted with CLAM and nuclear features with Hover-Net. These features were combined via Multimodal Compact Bilinear Pooling and utilized in six distinct deep learning models. By fusing the nucleus segmentation features, the model increased area under the receiver operating characteristic curve (AUC) by 1%-3% and recall by 5%-11% in five-fold cross-validation, significantly outperforming models that relied solely on image features. External validation on a CRC dataset from the China-Japan Friendship hospital further validated the model's robustness, achieving an AUC of 0.81 and a recall of 0.80 for MSI prediction. Additionally, notable differences in cellular composition were observed across cancer types and clinical groups, emphasizing the pivotal role of cellular features in cancer development. These findings highlight the advantages of integrating H&E-stained image features with nuclear segmentation data and advanced deep learning techniques to improve predictive accuracy and reduce the cost of MSI/TMB testing, potentially advancing personalized cancer treatment strategies. Jiaying Han, Fengyuan Hu, Geng Tian, Dingrong Zhong, Jialiang Yang |
Briefings Bioinform. | 6 |
| 2025 | CAMIL: channel attention-based multiple instance learning for whole slide image classificationabstractMOTIVATION: The classification task based on whole-slide images (WSIs) is a classic problem in computational pathology. Multiple instance learning (MIL) provides a robust framework for analyzing whole slide images with slide-level labels at gigapixel resolution. However, existing MIL models typically focus on modeling the relationships between instances while neglecting the variability across the channel dimensions of instances, which prevents the model from fully capturing critical information in the channel dimension. RESULTS: To address this issue, we propose a plug-and-play module called Multi-scale Channel Attention Block (MCAB), which models the interdependencies between channels by leveraging local features with different receptive fields. By alternately stacking four layers of Transformer and MCAB, we designed a channel attention-based MIL model (CAMIL) capable of simultaneously modeling both inter-instance relationships and intra-channel dependencies. To verify the performance of the proposed CAMIL in classification tasks, several comprehensive experiments were conducted across three datasets: Camelyon16, TCGA-NSCLC, and TCGA-RCC. Empirical results demonstrate that, whether the feature extractor is pretrained on natural images or on WSIs, our CAMIL surpasses current state-of-the-art MIL models across multiple evaluation metrics. AVAILABILITY AND IMPLEMENTATION: All implementation code is available at https://github.com/maojy0914/CAMIL. Jinyang Mao, Junlin Xu, Xianfang Tang, Heaven Zhao, Geng Tian, Jialiang Yang |
Bioinform. | 6 |
| 2025 | SWMA-UNet: Multi-Path Attention Network for Improved Medical Image SegmentationabstractIn recent years, deep learning achieves significant advancements in medical image segmentation. Research finds that integrating Transformers and CNNs effectively addresses the limitations of CNNs in managing long-distance dependencies and understanding global information.However, existing models typically employ a serial approach to combine Transformers and CNNs, which complicates the simultaneous processing of global and local information. To address this, our study proposes a parallel multi-path attention architecture, SWMA-UNET, that integrates Transformers and CNNs. This architecture deeply mines features through parallel strategies while capturing both local details and global context information, thereby enhancing the accuracy of medical image segmentation. Experimental results indicate that our method surpasses all previously reported methods in the literature on the Synapse, ACDC, ISIC 2018 and MoNuSeg datasets. Xianfang Tang, Jincan Li, Qianrui Liu, Chang Zhou 0007, Pan Zeng, Yajie Meng, Junlin Xu, Geng Tian, Jialiang Yang |
IEEE J. Biomed. Health Informatics | 8 |
| 2025 | Enhancing Drug Repositioning Through Local Interactive Learning With Bilinear Attention NetworksabstractDrug repositioning has emerged as a promising strategy for identifying new therapeutic applications for existing drugs. In this study, we present DRGBCN, a novel computational method that integrates heterogeneous information through a deep bilinear attention network to infer potential drugs for specific diseases. DRGBCN involves constructing a comprehensive drug-disease network by incorporating multiple similarity networks for drugs and diseases. Firstly, we introduce a layer attention mechanism to effectively learn the embeddings of graph convolutional layers from these networks. Subsequently, a bilinear attention network is constructed to capture pairwise local interactions between drugs and diseases. This combined approach enhances the accuracy and reliability of predictions. Finally, a multi-layer perceptron module is employed to evaluate potential drugs. Through extensive experiments on three publicly available datasets, DRGBCN demonstrates better performance over baseline methods in 10-fold cross-validation, achieving an average area under the receiver operating characteristic curve (AUROC) of 0.9399. Furthermore, case studies on bladder cancer and acute lymphoblastic leukemia confirm the practical application of DRGBCN in real-world drug repositioning scenarios. Importantly, our experimental results from the drug-disease network analysis reveal the successful clustering of similar drugs within the same community, providing valuable insights into drug-disease interactions. In conclusion, DRGBCN holds significant promise for uncovering new therapeutic applications of existing drugs, thereby contributing to the advancement of precision medicine. Xianfang Tang, Chang Zhou 0007, Changcheng Lu, Yajie Meng, Junlin Xu, Xinrong Hu, Geng Tian, Jialiang Yang |
IEEE J. Biomed. Health Informatics | 7 |
| 2024 | Convergence Analysis of Split Federated Learning on Heterogeneous DataabstractSplit federated learning (SFL) is a recent distributed approach for collaborative model training among multiple clients. In SFL, a global model is typically split into two parts, where clients train one part in a parallel federated manner, and a main server trains the other. Despite the recent research on SFL algorithm development, the convergence analysis of SFL is missing in the literature, and this paper aims to fill this gap. The analysis of SFL can be more challenging than that of federated learning (FL), due to the potential dual-paced updates at the clients and the main server. We provide convergence analysis of SFL for strongly convex and general convex objectives on heterogeneous data. The convergence rates are $O(1/T)$ and $O(1/\sqrt[3]{T})$, respectively, where $T$ denotes the total number of rounds for SFL training. We further extend the analysis to non-convex objectives and where some clients may be unavailable during training. Numerical experiments validate our theoretical results and show that SFL outperforms FL and split learning (SL) when data is highly heterogeneous across a large number of clients. Pengchao Han, Geng Tian |
NeurIPS | 3 |
| 2024 | Drug repositioning based on weighted local information augmented graph neural networkabstractDrug repositioning, the strategy of redirecting existing drugs to new therapeutic purposes, is pivotal in accelerating drug discovery. While many studies have engaged in modeling complex drug-disease associations, they often overlook the relevance between different node embeddings. Consequently, we propose a novel weighted local information augmented graph neural network model, termed DRAGNN, for drug repositioning. Specifically, DRAGNN firstly incorporates a graph attention mechanism to dynamically allocate attention coefficients to drug and disease heterogeneous nodes, enhancing the effectiveness of target node information collection. To prevent excessive embedding of information in a limited vector space, we omit self-node information aggregation, thereby emphasizing valuable heterogeneous and homogeneous information. Additionally, average pooling in neighbor information aggregation is introduced to enhance local information while maintaining simplicity. A multi-layer perceptron is then employed to generate the final association predictions. The model's effectiveness for drug repositioning is supported by a 10-times 10-fold cross-validation on three benchmark datasets. Further validation is provided through analysis of the predicted associations using multiple authoritative data sources, molecular docking experiments and drug-disease network analysis, laying a solid foundation for future drug discovery. Yajie Meng, Junlin Xu, Changcheng Lu, Xianfang Tang, Ben-gong Zhang, Geng Tian, Jialiang Yang |
Briefings Bioinform. | 8 |
| 2024 | LDA-VGHB: identifying potential lncRNA-disease associations with singular value decomposition, variational graph auto-encoder and heterogeneous Newton boosting machineabstractLong noncoding RNAs (lncRNAs) participate in various biological processes and have close linkages with diseases. In vivo and in vitro experiments have validated many associations between lncRNAs and diseases. However, biological experiments are time-consuming and expensive. Here, we introduce LDA-VGHB, an lncRNA-disease association (LDA) identification framework, by incorporating feature extraction based on singular value decomposition and variational graph autoencoder and LDA classification based on heterogeneous Newton boosting machine. LDA-VGHB was compared with four classical LDA prediction methods (i.e. SDLDA, LDNFSGB, IPCARF and LDASR) and four popular boosting models (XGBoost, AdaBoost, CatBoost and LightGBM) under 5-fold cross-validations on lncRNAs, diseases, lncRNA-disease pairs and independent lncRNAs and independent diseases, respectively. It greatly outperformed the other methods with its prominent performance under four different cross-validations on the lncRNADisease and MNDR databases. We further investigated potential lncRNAs for lung cancer, breast cancer, colorectal cancer and kidney neoplasms and inferred the top 20 lncRNAs associated with them among all their unobserved lncRNAs. The results showed that most of the predicted top 20 lncRNAs have been verified by biomedical experiments provided by the Lnc2Cancer 3.0, lncRNADisease v2.0 and RNADisease databases as well as publications. We found that HAR1A, KCNQ1DN, ZFAT-AS1 and HAR1B could associate with lung cancer, breast cancer, colorectal cancer and kidney neoplasms, respectively. The results need further biological experimental validation. We foresee that LDA-VGHB was capable of identifying possible lncRNAs for complex diseases. LDA-VGHB is publicly available at https://github.com/plhhnu/LDA-VGHB. Lihong Peng, Liangliang Huang, Qiongli Su, Geng Tian, Min Chen 0028 |
Briefings Bioinform. | 4 |
| 2023 | Automatic Intelligent Chronic Kidney Disease Detection in Healthcare 5.0abstractHealth systems worldwide have an unprecedented opportunity to enhance healthcare service delivery due to the rapid development of emerging digital technologies. Many advancements have been made in the medical field, with deep learning proving particularly useful when applied to a large enough number of well-defined samples. Although, this aspect may make deep learning harder to implement in settings with limited-size datasets. In this study, we present a new method of chronic kidney disease detection (CKDD) by combining Generative Adversarial Networks (GAN) with Convolutional Neural Networks (CNN). Afterward, synthetic sample data was created using GAN, which enlarged the dataset. Subsequently, processing these synthetic samples, the CNN classifier was applied. According to experimental assessments, the suggested CKDD-GAN methodology accuracy is superior to without the GAN technique. Moreover, the proposed CKDD-GAN-based model outperformed with an accuracy of 98.10%. Even though standard synthetic data samples seemed to improve classification performance, GAN-based enhancements resulted in a 2.91% improvement. GAN implementations for detecting chronic kidney disease are highly beneficial since they also increase awareness about its possible uses in various other diseases. Geng Tian, Amir Rehman, Huanlai Xing, Nighat Gulzar, Abid Hussain 0002 |
TrustCom | 1 |
| 2023 | Spatial-Temporal Graph Network for Video Crowd CountingabstractIn recent years, researchers have developed many deep-learning-based methods to count crowd numbers in static images. However, much fewer works focus on video-based crowd counting, in which the critical challenge of temporal correlation has not been well explored. This paper proposes a Spatial-Temporal Graph Network (STGN) to achieve efficient and accurate crowd counting in videos via learning pixel-wise and patch-wise relations in local spatial-temporal domains. Specifically, we design a pyramid graph module to leverage multi-scale features. In each scale, we sequentially construct three graphs: spatial-temporal pixel graph, temporal patch graph, and spatial pixel graph, in which we apply the self-attention mechanism to capture pixel-wise relation, learn structure-aware relation, and aggregate local features, respectively. Furthermore, we propose spatial-aware channel-wise attention to effectively fuse multi-scale features. To demonstrate the effectiveness of the proposed method, we conduct experiments on five crowd counting datasets, including a large-scale video crowd dataset (FDST). Moreover, the proposed model is also applied in the vehicle counting dataset (TRANCOS). The results show that the proposed model outperforms existing spatial-temporal crowd counting models and achieves state-of-the-art. The code is available athttps://github.com/wuzhe71/STGN Zhe Wu 0006, Xinfeng Zhang 0001, Geng Tian, Yaowei Wang 0001, Qingming Huang |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2022 | DGHNE: network enhancement-based method in identifying disease-causing genes through a heterogeneous biomedical networkabstractThe identification of disease-causing genes is critical for mechanistic understanding of disease etiology and clinical manipulation in disease prevention and treatment. Yet the existing approaches in tackling this question are inadequate in accuracy and efficiency, demanding computational methods with higher identification power. Here, we proposed a new method called DGHNE to identify disease-causing genes through a heterogeneous biomedical network empowered by network enhancement. First, a disease-disease association network was constructed by the cosine similarity scores between phenotype annotation vectors of diseases, and a new heterogeneous biomedical network was constructed by using disease-gene associations to connect the disease-disease network and gene-gene network. Then, the heterogeneous biomedical network was further enhanced by using network embedding based on the Gaussian random projection. Finally, network propagation was used to identify candidate genes in the enhanced network. We applied DGHNE together with five other methods into the most updated disease-gene association database termed DisGeNet. Compared with all other methods, DGHNE displayed the highest area under the receiver operating characteristic curve and the precision-recall curve, as well as the highest precision and recall, in both the global 5-fold cross-validation and predicting new disease-gene associations. We further performed DGHNE in identifying the candidate causal genes of Parkinson's disease and diabetes mellitus, and the genes connecting hyperglycemia and diabetes mellitus. In all cases, the predicted causing genes were enriched in disease-associated gene ontology terms and Kyoto Encyclopedia of Genes and Genomes pathways, and the gene-disease associations were highly evidenced by independent experimental studies. Binsheng He, Ju Xiang, Pingping Bing, Geng Tian, Cheng Guo 0010, Jialiang Yang |
Briefings Bioinform. | 6 |
| 2022 | ICSDA: a multi-modal deep learning model to predict breast cancer recurrence and metastasis risk by integrating pathological, clinical and gene expression dataabstractBreast cancer patients often have recurrence and metastasis after surgery. Predicting the risk of recurrence and metastasis for a breast cancer patient is essential for the development of precision treatment. In this study, we proposed a novel multi-modal deep learning prediction model by integrating hematoxylin & eosin (H&E)-stained histopathological images, clinical information and gene expression data. Specifically, we segmented tumor regions in H&E into image blocks (256 × 256 pixels) and encoded each image block into a 1D feature vector using a deep neural network. Then, the attention module scored each area of the H&E-stained images and combined image features with clinical and gene expression data to predict the risk of recurrence and metastasis for each patient. To test the model, we downloaded all 196 breast cancer samples from the Cancer Genome Atlas with clinical, gene expression and H&E information simultaneously available. The samples were then divided into the training and testing sets with a ratio of 7: 3, in which the distributions of the samples were kept between the two datasets by hierarchical sampling. The multi-modal model achieved an area-under-the-curve value of 0.75 on the testing set better than those based solely on H&E image, sequencing data and clinical data, respectively. This study might have clinical significance in identifying high-risk breast cancer patients, who may benefit from postoperative adjuvant treatment. Yuhua Yao, Yaping Lv, Yuebin Liang, Shuxue Xi, Binbin Ji, Guanglu Zhang, Geng Tian, Xiyue Hu, Jialiang Yang |
Briefings Bioinform. | 9 |
| 2022 | Predicting colorectal cancer tumor mutational burden from histopathological images and clinical information using multi-modal deep learningabstractMOTIVATION: Tumor mutational burden (TMB) is an indicator of the efficacy and prognosis of immune checkpoint therapy in colorectal cancer (CRC). In general, patients with higher TMB values are more likely to benefit from immunotherapy. Though whole-exome sequencing is considered the gold standard for determining TMB, it is difficult to be applied in clinical practice due to its high cost. There are also a few DNA panel-based methods to estimate TMB; however, their detection cost is also high, and the associated wet-lab experiments usually take days, which emphasize the need for faster and cheaper alternatives. RESULTS: In this study, we propose a multi-modal deep learning model based on a residual network (ResNet) and multi-modal compact bilinear pooling to predict TMB status (i.e. TMB high (TMB_H) or TMB low(TMB_L)) directly from histopathological images and clinical data. We applied the model to CRC data from The Cancer Genome Atlas and compared it with four other popular methods, namely, ResNet18, ResNet50, VGG19 and AlexNet. We tested different TMB thresholds, namely, percentiles of 10%, 14.3%, 15%, 16.3%, 20%, 30% and 50%, to differentiate TMB_H and TMB_L.For the percentile of 14.3% (i.e. TMB value 20) and ResNet18, our model achieved an area under the receiver operating characteristic curve of 0.817 after 5-fold cross-validation, which was better than that of other compared models. In addition, we also found that TMB values were significantly associated with the tumor stage and N and M stages. Our study shows that deep learning models can predict TMB status from histopathological images and clinical information only, which is worth clinical application. Kaimei Huang, Binghu Lin, Yankun Liu, Jingwu Li, Geng Tian, Jialiang Yang |
Bioinform. | 6 |
| 2017 | CEFF: An efficient approach for traffic anomaly detection and classificationabstractNowadays, there are two major challenges to detect traffic anomalies in a large scale network. One is how to handle huge amounts of traffic data when we detect traffic anomalies in a network, and the other is how to carry out fast and detailed detection and classification. To address these two challenges, we propose a Change based Effective Frequent flow Features approach (CEFF), which can quickly obtain the anomaly detection and classification results by scanning the flow data only once. We implement CEFF for both offline and online detection and classification in Spark, a popular big data processing platform. Besides, we evaluate CEFF using China Telecom NetFlow format data in experiments, and make comparisons between CEFF and Shannon entropy based method, which has been proved to be effective for traffic anomaly detection. The experiment results show that CEFF has excellent performance in traffic anomaly detection and classification. Geng Tian, Xia Yin 0001, Xingang Shi, Zimu Li, Yingya Guo |
ISCC | 1 |
| 2015 | Mining network traffic anomaly based on adjustable piecewise entropyabstractToday network traffic anomaly detection is very challenging in a big and constantly changing network, because there are millions of flows being transferred in a network at the same time, and the flow numbers change all the time. Although traditional information entropy has been proved to be an effective metric on network traffic anomaly detection, such a metric shows some limitations in large scale networks with constantly changing flow numbers, and it makes the traditional entropy inefficient for traffic anomaly detection. Another challenge is how to process large-scale traffic data in a scalable way. In this paper, we propose Adjustable Piecewise Entropy for traffic anomaly detection, and implement Adjustable Piecewise Shannon entropy in Hadoop platform with a cluster of five servers in Tsinghua University Campus Network. Furthermore, we analyze and validate Adjustable Piecewise Entropy in both mathematics and experiments. The experiment results show that Adjustable Piecewise Entropy has better performance for traffic anomaly detection. Geng Tian, Xia Yin 0001, Zimu Li, Xingang Shi, Ziyi Lu, Yingya Guo |
IWQoS | 1 |
| 2015 | TADOOP: Mining Network Traffic Anomalies with Hadoop
Geng Tian, Xia Yin 0001, Zimu Li, Xingang Shi, Ziyi Lu |
SecureComm | 1 |
| 2013 | Recommending scientific articles using bi-relational graph-based iterative RWRabstractThe overabundance of scientific article information has created much inconvenience to researchers seeking interesting articles online. In this paper, we provide a Bi-Relational graph to represent the heterogenous information of scientific article recommendation system, which includes three parts: the article content similarity, researcher interest correlation, and researcher-article readership. Meanwhile, an iterative random walk with restarts learning method is proposed on the Bi-Relational graph to recommend a researcher rating for each article by making use of the known information. The proposed method has ability to perform both old and new article recommendation. A series of experiments on CiteULike dataset have shown that our method is more effective than other testing methods in the paper. Geng Tian, Liping Jing |
RecSys | 1 |
| 2011 | Estimation of allele frequency and association mapping using next-generation sequencing dataabstractBACKGROUND: Estimation of allele frequency is of fundamental importance in population genetic analyses and in association mapping. In most studies using next-generation sequencing, a cost effective approach is to use medium or low-coverage data (e.g., < 15X). However, SNP calling and allele frequency estimation in such studies is associated with substantial statistical uncertainty because of varying coverage and high error rates. RESULTS: We evaluate a new maximum likelihood method for estimating allele frequencies in low and medium coverage next-generation sequencing data. The method is based on integrating over uncertainty in the data for each individual rather than first calling genotypes. This method can be applied to directly test for associations in case/control studies. We use simulations to compare the likelihood method to methods based on genotype calling, and show that the likelihood method outperforms the genotype calling methods in terms of: (1) accuracy of allele frequency estimation, (2) accuracy of the estimation of the distribution of allele frequencies across neutrally evolving sites, and (3) statistical power in association mapping studies. Using real re-sequencing data from 200 individuals obtained from an exon-capture experiment, we show that the patterns observed in the simulations are also found in real data. CONCLUSIONS: Overall, our results suggest that association mapping and estimation of allele frequencies should not be based on genotype calling in low to medium coverage data. Furthermore, if genotype calling methods are used, it is usually better not to filter genotypes based on the call confidence score. Su Yeon Kim, Kirk E. Lohmueller, Anders Albrechtsen, Yingrui Li, Thorfinn Korneliussen, Geng Tian, Niels Grarup, Gitte Andersen, Daniel Witte, Torben Jørgensen, Torben Hansen, Oluf Pedersen, Jun Wang 0004, Rasmus Nielsen |
BMC Bioinform. | 6 |
| 2010 | Analysis of the correlation properties of digital satellite signals and their applicability in bistatic remote sensingabstractThis paper presents a study of relevant correlation properties of signal transmitted from commercial communication satellites in order to evaluate their potential use as “signal of opportunity” for bistatic remote sensing. The ambiguity function for the XM radio satellites was computed analytically from published information on the modulation schemes and bandwidth, under the assumption that the data modulation is random. The model was then experimentally tested by recording the received signals from these satellites. Next, a cross-correlation waveform for digital signal reflected from random rough surface was simulated. Scattering model that were originally developed for Global Navigation Satellite System (GNSS-R) signals was applied to the modified simulator to incorporate the derived ambiguity function. The simulator was then used to generate synthetic waveform with a realistic signal to noise ratio (SNR). Retrieval algorithms for ocean surface roughness and reflectivity that were derived originally for GNSS-R, were applied to these simulated signals. Non-linear least square methods were applied to invert a scattering model and estimate the slope variances of the probability density function (PDF), which best fits the measurements of the reflected XM signal waveform. The SNR for the experimental data was found to be within 0.5dB of the theoretically calculated SNR. Rashmi Shah, James L. Garrison, Michael S. Grant, Stephen J. Katzberg, Geng Tian |
IGARSS | 5 |