EDBT 2026 Demo / reviewers in the wild / expert
Shaolong Cao
dblp:137/7744
· DBLP profile ↗
5ranked-venue papers
1as first author
1since 2021 · last 2026
0000-0002-4443-9424ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Interdisciplinary, comprehensive, and emerging computing
2 papers |
Bioinformatics and computational biology · 100% |
Topics — the 5 heaviest of 5, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Bioinformatics and computational biology › transcriptomics
RNA splicing analysis |
1.0 | 1 | 2026 | SpliceHarmonization: an integrated method for identifying RNA splicing events in therapeutics for splicing modulation · Bioinform. 2026 |
Bioinformatics and computational biology › transcriptomics › RNA splicing analysis
splice variant prediction |
1.0 | 1 | 2026 | SpliceHarmonization: an integrated method for identifying RNA splicing events in therapeutics for splicing modulation · Bioinform. 2026 |
Bioinformatics and computational biology
transcriptomics |
1.0 | 1 | 2026 | SpliceHarmonization: an integrated method for identifying RNA splicing events in therapeutics for splicing modulation · Bioinform. 2026 |
Bioinformatics and computational biology › statistical genetics
fine-mapping |
0.2 | 1 | 2016 | Unified tests for fine-scale mapping and identifying sparse high-dimensional sequence associations · Bioinform. 2016 |
Bioinformatics and computational biology
statistical genetics |
0.2 | 1 | 2016 | Unified tests for fine-scale mapping and identifying sparse high-dimensional sequence associations · Bioinform. 2016 |
Methods — techniques the papers use, named apart from their topics
ensemble integration · 1.0sparse linear mixed regression · 0.2lp norm regularization · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SpliceHarmonization: an integrated method for identifying RNA splicing events in therapeutics for splicing modulationabstractMOTIVATION: Splicing, a critical co-transcriptional process in eukaryotes, enhances transcriptome diversity by generating isoforms specific to cell types, tissues, or developmental stages. Recent advancements in splicing modulators have opened new avenues for targeting previously undruggable genes by inducing significant perturbations in splicing events. These developments underscore the need for comprehensive methods to accurately identify and compare splicing events. While several tools have been developed to detect local splice variants, inconsistencies across methods remain a significant challenge. To address this, we present SpliceHarmonization, an integrated approach that combines the strengths of rMATS, LeafCutter, and MAJIQ, enabling robust and reliable splicing analysis with event type annotations. RESULTS: In a comprehensive evaluation using diverse simulated datasets, SpliceHarmonization streamlined and standardized the outputs from three detection methods into a unified format, thereby improving splicing detection with event type annotation and outperforming individual methods. By integrating the outputs from rMATS, LeafCutter, and MAJIQ, our approach not only enhanced identification of a wide range of splicing events but also effectively mitigated method-specific discrepancies. This integration led to an accuracy exceeding 0.8 and a recall of up to 0.5, with an observed increase in AUC of up to 10%. Furthermore, SpliceHarmonization demonstrated high sensitivity in detecting low-abundance and complex splicing events, providing annotations including genomic coordinates and event type. AVAILABILITY AND IMPLEMENTATION: SpliceHarmonization is available at https://github.com/interactivereport/SpliceHarmonization. Yirui Chen, Yu H. Sun, Soumya Negi, Shaolong Cao, Zhengyu Ouyang, Baohong Zhang, Jessica Hurt, Dann Huh |
Bioinform. | 5 |
| 2018 | A Sparse Regression Method for Group-Wise Feature Selection with False Discovery Rate ControlabstractThe method of Sorted L-One Penalized Estimation, or SLOPE, is a sparse regression method recently introduced by Bogdan et. al. [1] . It can be used to identify significant predictor variables in a linear model that may have more unknown parameters than observations. When the correlations between predictor variables are small, the SLOPE method is shown to successfully control the false discovery rate (the expected proportion of the irrelevant among all selected predictors) at a user specified level. However, the requirement for nearly uncorrelated predictors is too restrictive for genomic data, as demonstrated in our recent study [2] by an application of SLOPE to realistic simulated DNA sequence data. A possible solution is to divide the predictor variables into nearly uncorrelated groups, and to modify the procedure to select entire groups with an overall significant group effect, rather than individual predictors. Following this motivation, we extend SLOPE in the spirit of Group LASSO to Group SLOPE, a method that can handle group structures between the predictor variables, which are ubiquitous in real genomic data. Our theoretical results show that Group SLOPE controls the group-wise false discovery rate (gFDR), when groups are orthogonal to each other. For use in non-orthogonal settings, we propose two types of Monte Carlo based heuristics, which lead to gFDR control with Group SLOPE in simulations based on real SNP data. As an illustration of the merits of this method, an application of Group SLOPE to a dataset from the Framingham Heart Study results in the identification of some known DNA sequence regions associated with bone health, as well as some new candidate regions. The novel methods are implemented in the R package grpSLOPEMC , which is publicly available at https://github.com/agisga/grpSLOPEMC. Alexej Gossmann, Shaolong Cao, Damian Brzyski, Hong-Wen Deng, Yu-Ping Wang 0002 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2017 | Identifying Stages of Kidney Renal Cell Carcinoma by Combining Gene Expression and DNA Methylation DataabstractIn this study, in order to take advantage of complementary information from different types of data for better disease status diagnosis, we combined gene expression with DNA methylation data and generated a fused network, based on which the stages of Kidney Renal Cell Carcinoma (KIRC) can be better identified. It is well recognized that a network is important for investigating the connectivity of disease groups. We exploited the potential of the network's features to identify the KIRC stage. We first constructed a patient network from each type of data. We then built a fused network based on network fusion method. Based on the link weights of patients, we used a generalized linear model to predict the group of KIRC subjects. Finally, the group prediction method was applied to test the power of network-based features. The performance (e.g., the accuracy of identifying cancer stages) when using the fused network from two types of data is shown to be superior to that when using two patient networks from only one data type. The work provides a good example for using network based features from multiple data types for a more comprehensive diagnosis. Su-Ping Deng, Shaolong Cao, De-Shuang Huang, Yu-Ping Wang 0002 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2016 | Unified tests for fine-scale mapping and identifying sparse high-dimensional sequence associationsabstractMOTIVATION: In searching for genetic variants for complex diseases with deep sequencing data, genomic marker sets of high-dimensional genotypic data and sparse functional variants are quite common. Existing sequence association tests are incapable of identifying such marker sets or individual causal loci, although they appeared powerful to identify small marker sets with dense functional variants. In sequence association studies of admixed individuals, cryptic relatedness and population structure are known to confound the association analyses. METHOD: We here propose a unified marker wise test (uFineMap) to accurately localize causal loci and a unified high-dimensional set based test (uHDSet) to identify high-dimensional sparse associations in deep sequencing genomic data of multi-ethnic individuals with random relatedness. These two novel tests are based on scaled sparse linear mixed regressions with Lp (0 < p < 1) norm regularization. They jointly adjust for cryptic relatedness, population structure and other confounders to prevent false discoveries and improve statistical power for identifying promising individual markers and marker sets that harbor functional genetic variants of a complex trait. RESULTS: With large scale simulation data and real data analyses, the proposed tests appropriately controlled Type I error rates and appeared to be more powerful than several prominent methods. We illustrated their practical utilities by the applications to DNA sequence data of Framingham Heart Study for osteoporosis. The proposed tests identified 11 novel significant genes that were missed by the prominent famSKAT and GEMMA. In particular, four out of six most significant pathways identified by the uHDSet but missed by famSKAT have been reported to be related to BMD or osteoporosis in the literature. AVAILABILITY AND IMPLEMENTATION: The computational toolkit is available for academic use: https://sites.google.com/site/shaolongscode/home/uhdset CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Shaolong Cao, Huaizhen Qin, Alexej Gossmann, Hong-Wen Deng, Yu-Ping Wang 0002 |
Bioinform. | 1 |
| 2015 | The effective diagnosis of schizophrenia by using multi-layer RBMs deep networksabstractSchizophrenia is one of the most prevalent mental diseases, and is considered to be caused by the interplay of a number of genetic factors. In this paper, by constructing a multilayer restricted Boltzmann machines (RBMs) deep network, we use the genomic data (i.e., SNP data) for unsupervised feature learning and disease diagnosis of schizophrenia. In order to obtain some more accurate diagnosis results by RBMs, firstly, we transform the SNP data into binary sequences, and then by training the multi-layer RBMs deep network on unlabeled data, the multi-level abstract features of the genomic data are obtained and stored in the network. Finally, by adding a linear classifier to the top of the multi-layer RBMs deep network, the classification results on the testing data are gained. The results show that the average performance of this method is better than that of other methods, e.g., SVM (including linear SVM as well as SVM with multilayer perceptron kernel), sparse representations based classifier and k-nearest neighbors method. It is indicated that the multi-layer RBMs deep network can extract deep hierarchical representations of the genomic data, and then promises a more comprehensive approach for the mental disease diagnosis. Chen Qiao, Dongdong Lin, Shaolong Cao, Yu-Ping Wang 0002 |
BIBM | 3 |