EDBT 2026 Demo / reviewers in the wild / expert
Qinghua Cui
dblp:63/2331
· DBLP profile ↗
26ranked-venue papers
2as first author
10since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 22 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | FreeCG: Free the Design Space of Clebsch-Gordan Transform for Machine Learning Force FieldsabstractMachine Learning Force Fields (MLFFs) are of great importance for chemistry, physics, materials science, and many other related fields. The Clebsch–Gordan transform (CG transform) effectively encodes many-body interactions and is thus an important building block for many models of MLFFs. However, the permutation-equivariance requirement of MLFFs limits the design space of CG transform, that is, intensive CG transform has to be conducted for each neighboring edge and the operations should be performed in the same manner for all edges. Freeing up the design space can greatly improve the model's expressiveness while simultaneously decreasing computational demands. To reach this goal, we utilize a mathematical proposition, invariance transitivity, to show that implementing the CG transform layer on the permutation-invariant abstract edges allows complete freedom in the design of the layer without compromising the overall permutation equivariance. Developing on this free design space, we further propose group CG transform with sparse path, abstract edges shuffling, and attention enhancer to form a powerful and efficient CG transform layer. Our method, known as FreeCG, achieves state-of-the-art (SOTA) results in force prediction for MD17, rMD17, MD22, and is well extended to property prediction in QM9 datasets with several improvements greater than 15% and the maximum beyond 20%. The extensive real-world applications showcase high practicality. FreeCG introduces a novel paradigm for carrying out efficient and expressive CG transform in future geometric network designs. To demonstrate this, the recent SOTA, QuinNet, is also enhanced under our paradigm. Code: https://github.com/ShihaoShao-GH/FreeCG. Shihao Shao, Qinghua Cui |
ICLR | 4 |
| 2025 | PIVNO: Particle Image Velocimetry Neural OperatorabstractParticle Image Velocimetry (PIV) aims to infer underlying velocity fields from time-separated particle images, forming a PDE-constrained inverse problem governed by advection dynamics. Traditional cross-correlation methods and deep learning-based feature matching approaches often struggle with ambiguity, limited resolution, and generalization to real-world conditions. To address these challenges, we propose a PIV Neural Operator (PIVNO) framework that directly approximates the inverse mapping from paired particle images to flow fields within a function space. Leveraging a position informed Galerkin-style attention operator, PIVNO captures global flow structures while supporting resolution-adaptive inference across arbitrary subdomains. Moreover, to enhance real-world adaptability, we introduce a self-supervised fine-tuning scheme based on physical divergence constraints, enabling the model to generalize from synthetic to real experiments without requiring labeled data. Extensive evaluations demonstrate the accuracy, flexibility, and robustness of our approach across both simulated and experimental PIV datasets. Our code is at https://github.com/ZXS-Labs/PIVNO. Xu Jie, Xuesong Zhang 0001, Qinghua Cui |
NeurIPS | 4 |
| 2025 | High-Rank Irreducible Cartesian Tensor Decomposition and Bases of Equivariant SpacesabstractIrreducible Cartesian tensors (ICTs) play a crucial role in the design of equivariant graph neural networks, as well as in theoretical chemistry and chemical physics. Meanwhile, the design space of available linear operations on tensors that preserve symmetry presents a significant challenge. The ICT decomposition and a basis of this equivariant space are difficult to obtain for high-rank tensors. After decades of research, Bonvicini (2024) has recently achieved an explicit ICT decomposition for $n=5$ with factorial time/space complexity. In this work we, for the first time, obtain decomposition matrices for ICTs up to rank $n=9$ with reduced and affordable complexity, by constructing what we call path matrices. The path matrices are obtained via performing chain-like contractions with Clebsch-Gordan matrices following the parentage scheme. We prove and leverage that the concatenation of path matrices is an orthonormal change-of-basis matrix between the Cartesian tensor product space and the spherical direct sum spaces. Furthermore, we identify a complete orthogonal basis for the equivariant space, rather than a spanning set (Pearce-Crump, 2023b), through this path matrices technique. Our method avoids the RREF algorithm and maintains a fully analytical derivation of each ICT decomposition matrix, thereby significantly improving the algorithm’s speed to obtain arbitrary rank orthogonal ICT decomposition matrices and orthogonal equivariant bases. We further extend our result to the arbitrary tensor product and direct sum spaces, enabling free design between different spaces while keeping symmetry. The Python code is available at https://github.com/ShihaoShao-GH/ICT-decomposition-and-equivariant-bases, where the $n=6,\dots,9$ ICT decomposition matrices are obtained in 1s, 3s, 11s, and 4m32s on 28-core Intel Xeon Gold 6330 CPU @ 2.00GHz, respectively. Shihao Shao, Zhouchen Lin, Qinghua Cui |
J. Mach. Learn. Res. | 4 |
| 2023 | Global Features are All You Need for Image Retrieval and RerankingabstractImage retrieval systems conventionally use a two-stage paradigm, leveraging global features for initial retrieval and local features for reranking. However, the scalability of this method is often limited due to the significant storage and computation cost incurred by local feature matching in the reranking stage. In this paper, we present SuperGlobal, a novel approach that exclusively employs global features for both stages, improving efficiency without sacrificing accuracy. SuperGlobal introduces key enhancements to the retrieval system, specifically focusing on the global feature extraction and reranking processes. For extraction, we identify sub-optimal performance when the widely-used ArcFace loss and Generalized Mean (GeM) pooling methods are combined and propose several new modules to improve GeM pooling. In the reranking stage, we introduce a novel method to update the global features of the query and top-ranked images by only considering feature refinement with a small set of images, thus being very compute and memory efficient. Our experiments demonstrate substantial improvements compared to the state of the art in standard benchmarks. Notably, on the Revisited Oxford+1M Hard dataset, our single-stage results improve by 7.1%, while our two-stage gain reaches 3.7% with a strong 64,865× speedup. Our two-stage system surpasses the current single-stage state-of-the-art by 16.3%, offering a scalable, accurate alternative for high-performing image retrieval systems with minimal time overhead.Code: https://github.com/ShihaoShao-GH/SuperGlobal. Shihao Shao, Kaifeng Chen, Arjun Karpur, Qinghua Cui, André Araújo 0001, Bingyi Cao |
ICCV | 4 |
| 2023 | Defining the single base importance of human mRNAs and lncRNAsabstractAs the fundamental unit of a gene and its transcripts, nucleotides have enormous impacts on the gene function and evolution, and thus on phenotypes and diseases. In order to identify the key nucleotides of one specific gene, it is quite crucial to quantitatively measure the importance of each base on the gene. However, there are still no sequence-based methods of doing that. Here, we proposed Base Importance Calculator (BIC), an algorithm to calculate the importance score of each single base based on sequence information of human mRNAs and long noncoding RNAs (lncRNAs). We then confirmed its power by applying BIC to three different tasks. Firstly, we revealed that BIC can effectively evaluate the pathogenicity of both genes and single bases through single nucleotide variations. Moreover, the BIC score in The Cancer Genome Atlas somatic mutations is able to predict the prognosis of some cancers. Finally, we show that BIC can also precisely predict the transmissibility of SARS-CoV-2. The above results indicate that BIC is a useful tool for evaluating the single base importance of human mRNAs and lncRNAs. Xiangwen Ji, Qinghua Cui, Chunmei Cui |
Briefings Bioinform. | 4 |
| 2023 | Deciphering gene contributions and etiologies of somatic mutational signatures of cancerabstractSomatic mutational signatures (MSs) identified by genome sequencing play important roles in exploring the cause and development of cancer. Thus far, many such signatures have been identified, and some of them do imply causes of cancer. However, a major bottleneck is that we do not know the potential meanings (i.e. carcinogenesis or biological functions) and contributing genes for most of them. Here, we presented a computational framework, Gene Somatic Genome Pattern (GSGP), which can decipher the molecular mechanisms of the MSs. More importantly, it is the first time that the GSGP is able to process MSs from ribonucleic acid (RNA) sequencing, which greatly extended the applications of both MS analysis and RNA sequencing (RNAseq). As a result, GSGP analyses match consistently with previous reports and identify the etiologies for a number of novel signatures. Notably, we applied GSGP to RNAseq data and revealed an RNA-derived MS involved in deficient deoxyribonucleic acid mismatch repair and microsatellite instability in colorectal cancer. Researchers can perform customized GSGP analysis using the web tools or scripts we provide. Xiangwen Ji, Edwin Wang, Qinghua Cui |
Briefings Bioinform. | 3 |
| 2021 | Defining the functional divergence of orthologous genes between human and mouse in the context of miRNA regulationabstractAnimal models have a certain degree of similarity with human in genes and physiological processes, which leads them to be valuable tools for studying human diseases and for assisting drug development. However, translational researches adopting animal models are largely restricted by the species heterogeneity, which is also a major reason for the failure of drug research. Currently, computational method for exploring the functional differences between orthologous genes is still insufficient. For this purpose, here, we presented an algorithm, functional divergence score (FDS), by comprehensively evaluating the functional differences between the microRNAs regulating the paired orthologous genes. Given that mouse is one of the most popular model animals, currently, FDS was designed to dissect the functional divergence of orthologous genes between human and mouse. The results showed that gene FDS value is significantly associated with gene evolutionary characteristics and can discover expression divergence of human-mouse orthologous genes. Moreover, FDS performed well in distinguishing the targets of approved drugs and the failed ones. These results suggest that FDS is a valuable tool to evaluate the functional divergence of paired human and mouse orthologous genes. In addition, for each orthologous gene pair, FDS can provide detailed differences in functions and phenotypes. Our study provided a useful tool for quantifying the functional difference between human and mouse, and the presented framework is easily to be extended to the orthologous genes between human and other species. An online server of FDS is available at http://www.cuilab.cn/fds/. Chunmei Cui, Qinghua Cui |
Briefings Bioinform. | 3 |
| 2021 | Toward comprehensive functional analysis of gene lists weighted by gene essentiality scoresabstractMOTIVATION: Gene functional enrichment analysis represents one of the most popular bioinformatics methods for annotating the pathways and function categories of a given gene list. Current algorithms for enrichment computation such as Fisher's exact test and hypergeometric test totally depend on the category count numbers of the gene list and one gene set. In this case, whatever the genes are, they were treated equally. However, actually genes show different scores in their essentiality in a gene list and in a gene set. It is thus hypothesized that the essentiality scores could be important and should be considered in gene functional analysis. RESULTS: For this purpose, here, we proposed weighted enrichment analysis tool (WEAT) (https://www.cuilab.cn/weat/), a weighted gene set enrichment algorithm and online tool by weighting genes using essentiality scores. We confirmed the usefulness of WEAT using three case studies, the functional analysis of one aging-related gene list, one gene list involved in Lung Squamous Cell Carcinoma and one cardiomyopathy gene list from Drosophila model. Finally, we believe that the WEAT method and tool could provide more possibilities for further exploring the functions of given gene lists. AVAILABILITY AND IMPLEMENTATION: The datasets generated and analyzed during the current study are available on our website at https://www.cuilab.cn/weat/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Qinghua Cui |
Bioinform. | 2 |
| 2021 | Comparative analysis of aneurysm subtypes associated genes based on protein-protein interaction networkabstractThe arterial aneurysm refers to localized dilation of blood vessel wall and is common in general population. The majority of aneurysm cases remains asymptomatic until a sudden rupture which is usually fatal and of extremely high mortality (~ 50-60%). Therefore, early diagnosis, prevention and management of aneurysm are in urgent need. Unfortunately, current understanding of disease driver genes of various aneurysm subtypes is still limited, and without appropriate biomarkers and drug targets no specialized drug has been developed for aneurysm treatment. In this research, aneurysm subtypes were analyzed based on protein-protein interaction network to better understand aneurysm pathogenesis. By measuring network-based proximity of aneurysm subtypes, we identified a relevant closest relationship between aortic aneurysm and aortic dissection. An improved random walk method was performed to prioritize candidate driver genes of each aneurysm subtype. Thereafter, transcriptomes of 6 human aneurysm subtypes were collected and differential expression genes were identified to further filter potential driver genes. Functional enrichment of above driver genes indicated a general role of ubiquitination and programmed cell death in aneurysm pathogenesis. Especially, we further observed participation of BCL-2-mediated apoptosis pathway and caspase-1 related pyroptosis in the development of cerebral aneurysm and aneurysmal subarachnoid hemorrhage in corresponding transcriptomes. Ruya Sun, Qinghua Cui |
BMC Bioinform. | 3 |
| 2021 | AGTR2, One Possible Novel Key Gene for the Entry of SARS-CoV-2 Into Human CellsabstractRecently, it was confirmed that ACE2 is the receptor of SARS-CoV-2, the pathogen causing the recent outbreak of severe pneumonia around the world. It is confused that ACE2 is widely expressed across a variety of organs and is expressed moderately but not highly in lung, which, however, is the major infected organ. Therefore, we hypothesized that there could be some other genes playing key roles in the entry of SARS-CoV-2 into human cells. Here we found that AGTR2 (angiotensin II receptor type 2), a G-protein coupled receptor, has interaction with ACE2 and is highly expressed in lung with a high tissue specificity. More importantly, simulation of 3D structure based protein-protein interaction reveals that AGTR2 shows a higher binding affinity with the Spike protein of SARS-CoV-2 than ACE2 (energy: -8.2 vs. -5.1 [kcal/mol]). A number of compounds, biologics and traditional Chinese medicine that could decrease the expression level of AGTR2 were predicted. Finally, we suggest that AGTR2 could be a putative novel gene for the entry of SARS-CoV-2 into human cells, which could provide different insight for the research of SARS-CoV-2 proteins with their receptors. Chunmei Cui, Chuanbo Huang, Wanlu Zhou, Xiangwen Ji, Fenghong Zhang, Qinghua Cui |
IEEE ACM Trans. Comput. Biol. Bioinform. | 8 |
| 2020 | smORFunction: a tool for predicting functions of small open reading frames and microproteinsabstractBACKGROUND: Small open reading frame (smORF) is open reading frame with a length of less than 100 codons. Microproteins, translated from smORFs, have been found to participate in a variety of biological processes such as muscle formation and contraction, cell proliferation, and immune activation. Although previous studies have collected and annotated a large abundance of smORFs, functions of the vast majority of smORFs are still unknown. It is thus increasingly important to develop computational methods to annotate the functions of these smORFs. RESULTS: In this study, we collected 617,462 unique smORFs from three studies. The expression of smORF RNAs was estimated by reannotated microarray probes. Using a speed-optimized correlation algorism, the functions of smORFs were predicted by their correlated genes with known functional annotations. After applying our method to 5 known microproteins from literatures, our method successfully predicted their functions. Further validation from the UniProt database showed that at least one function of 202 out of 270 microproteins was predicted. CONCLUSIONS: We developed a method, smORFunction, to provide function predictions of smORFs/microproteins in at most 265 models generated from 173 datasets, including 48 tissues/cells, 82 diseases (and normal). The tool can be available at https://www.cuilab.cn/smorfunction . Xiangwen Ji, Chunmei Cui, Qinghua Cui |
BMC Bioinform. | 3 |
| 2020 | m6Acorr: an online tool for the correction and comparison of m6A methylation profilesabstractAbstract Background The analysis and comparison of RNA m6A methylation profiles have become increasingly important for understanding the post-transcriptional regulations of gene expression. However, current m6A profiles in public databases are not readily intercomparable, where heterogeneous profiles from the same experimental report but different cell types showed unwanted high correlations. Results Several normalizing or correcting methods were tested to remove such laboratory bias. And m6Acorr, an effective pipeline for correcting m6A profiles, was presented on the basis of quantile normalization and empirical Bayes batch regression method. m6Acorr could efficiently correct laboratory bias in the simulated dataset and real m6A profiles in public databases. The preservation of biological signals was examined after correction, and m6Acorr was found to better preserve differential methylation signals, m6A regulated targets, and m6A-related biological features than alternative methods. Finally, the m6Acorr server was established. This server could eliminate the potential laboratory bias in m6A methylation profiles and perform profile–profile comparisons and functional analysis of hyper- (hypo-) methylated genes based on corrected methylation profiles. Conclusion m6Acorr was established to correct the existing laboratory bias in RNA m6A methylation profiles and perform profile comparisons on the corrected datasets. The m6Acorr server is available at http://www.rnanut.net/m6Acorr . A stand-alone version with the correction function is also available in GitHub at https://github.com/emersON106/m6Acorr . Qinghua Cui, Yuan Zhou 0018 |
BMC Bioinform. | 3 |
| 2019 | miES: predicting the essentiality of miRNAs with machine learning and sequence featuresabstractMOTIVATION: MicroRNAs (miRNAs) are one class of small noncoding RNA molecules, which regulate gene expression at the post-transcriptional level and play important roles in health and disease. To dissect the critical miRNAs in miRNAome, it is needed to predict the essentiality of miRNAs, however, bioinformatics methods for this purpose are limited. RESULTS: Here we propose miES, a novel algorithm, for the prioritization of miRNA essentiality. miES implements a machine learning strategy based on learning from positive and unlabeled samples. miES uses sequence features of known essential miRNAs and performs miRNAome-wide searching for new essential miRNAs. miES achieves an AUC of 0.9 for 5-fold cross validation. Moreover, experiments further show that the miES score is significantly correlated with some established biological metrics for miRNA importance, such as miRNA conservation, miRNA disease spectrum width (DSW) and expression level. AVAILABILITY AND IMPLEMENTATION: The R source code is available at the download page of the web server, http://www.cuilab.cn/mies. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Chunmei Cui, Lin Gao 0006, Qinghua Cui |
Bioinform. | 4 |
| 2019 | NmSEER V2.0: a prediction tool for 2′-O-methylation sites based on random forest and multi-encoding combinationabstractBACKGROUND: 2'-O-methylation (2'-O-me or Nm) is a post-transcriptional RNA methylation modified at 2'-hydroxy, which is common in mRNAs and various non-coding RNAs. Previous studies revealed the significance of Nm in multiple biological processes. With Nm getting more and more attention, a revolutionary technique termed Nm-seq, was developed to profile Nm sites mainly in mRNA with single nucleotide resolution and high sensitivity. In a recent work, supported by the Nm-seq data, we have reported a method in silico for predicting Nm sites, which relies on nucleotide sequence information, and established an online server named NmSEER. More recently, a more confident dataset produced by refined Nm-seq was available. Therefore, in this work, we redesigned the prediction model to achieve a more robust performance on the new data. RESULTS: We redesigned the prediction model from two perspectives, including machine learning algorithm and multi-encoding scheme combination. With optimization by 5-fold cross-validation tests and evaluation by independent test respectively, random forest was selected as the most robust algorithm. Meanwhile, one-hot encoding, together with position-specific dinucleotide sequence profile and K-nucleotide frequency encoding were collectively applied to build the final predictor. CONCLUSIONS: The predictor of updated version, named NmSEER V2.0, achieves an accurate prediction performance (AUROC = 0.862) and has been settled into a brand-new server, which is available at http://www.rnanut.net/nmseer-v2/ for free. Yiran Zhou, Qinghua Cui |
BMC Bioinform. | 2 |
| 2018 | NmSEER: A Prediction Tool for 2'-O-Methylation (Nm) Sites Based on Random Forest
Yiran Zhou, Qinghua Cui |
ICIC (1) | 2 |
| 2018 | Identification and analysis of the human sex-biased genesabstractTremendous differences between human sexes are universally observed. Therefore, identifying and analyzing the sex-biased genes are becoming basically important for uncovering the mystery of sex differences and personalized medicine. Here, we presented a computational method to identify sex-biased genes from public gene expression databases. We obtained 1407 female-biased genes (FGs) and 1096 male-biased genes (MGs) across 14 different tissues. Bioinformatics analysis revealed that compared with MGs, FGs have higher evolutionary rate, higher single-nucleotide polymorphism density, less homologous gene numbers and smaller phyletic age. FGs have lower expression level, higher tissue specificity and later expressed stage in body development. Moreover, FGs are highly involved in immune-related functions, whereas MGs are more enriched in metabolic process. In addition, cellular network analysis revealed that MGs have higher degree, more cellular activating signaling and tend to be located in cellular inner space, whereas FGs have lower degree, more cellular repressing signaling and tend to be located in cellular outer space. Finally, the identified sex-biased genes and the discovered biological insights together can be a valuable resource helpful for investigating sex-biased physiology and medicine, for example sex-biased disease diagnosis and therapy, which represents one important aspect of personalized and precision medicine. Sisi Guo, Pan Zeng, Guoheng Xu, Qinghua Cui |
Briefings Bioinform. | 6 |
| 2017 | An analysis of human microbe-disease associationsabstractThe microbiota living in the human body has critical impacts on our health and disease, but a systems understanding of its relationships with disease remains limited. Here, we use a large-scale text mining-based manually curated microbe-disease association data set to construct a microbe-based human disease network and investigate the relationships between microbes and disease genes, symptoms, chemical fragments and drugs. We reveal that microbe-based disease loops are significantly coherent. Microbe-based disease connections have strong overlaps with those constructed by disease genes, symptoms, chemical fragments and drugs. Moreover, we confirm that the microbe-based disease analysis is able to predict novel connections and mechanisms for disease, microbes, genes and drugs. The presented network, methods and findings can be a resource helpful for addressing some issues in medicine, for example, the discovery of bench knowledge and bedside clinical solutions for disease mechanism understanding, diagnosis and therapy. Wei Ma 0010, Pan Zeng, Chuanbo Huang, Bin Geng, Jichun Yang, Xuezhong Zhou, Qinghua Cui |
Briefings Bioinform. | 10 |
| 2015 | LncTar: a tool for predicting the RNA targets of long noncoding RNAsabstractLong noncoding RNAs (lncRNAs) represent a big category of noncoding RNA molecules, and increasing studies have shown that they play important roles in various critical biological processes. They show a diversity of functions through diverse mechanisms, among which regulating RNA molecules is one of the most popular ones. Given the big number of lncRNAs, it becomes urgent and important to predict the RNA targets of lncRNAs in a large scale for the comprehensive understanding of lncRNA functions and action mechanisms. Although several methods have been developed to predict RNA-RNA interactions, none of them can be used to predict the RNA targets of lncRNAs in a large scale. Here we presented a tool, LncTar, which shows the ability to efficiently predict the RNA targets of lncRNAs in a large scale. To test the accuracy of LncTar, we applied it to 10 experimentally supported lncRNA-mRNA interactions. As a result, LncTar successfully predicted 8 (80%) of the 10 lncRNA-mRNA pairs, suggesting that LncTar has a reliable accuracy. Finally, we believe that LncTar could be an efficient tool for the fast identification of the RNA targets of lncRNAs. LncTar is freely available at http://www.cuilab.cn/lnctar. Wei Ma 0010, Pan Zeng, Bin Geng, Jichun Yang, Qinghua Cui |
Briefings Bioinform. | 7 |
| 2015 | PPUS: a web server to predict PUS-specific pseudouridine sitesabstractMOTIVATION: Pseudouridine (Ψ), catalyzed by pseudouridine synthase (PUS), is the most abundant RNA modification and has important cellular functions. Developing an algorithm to identify Ψ sites is an important work. And it is better if the algorithm could assign which PUS modifies the Ψ sites. Here, we developed PPUS (http://lyh.pkmu.cn/ppus/), the first web server to predict PUS-specific Ψ sites. PPUS: employed support vector machine as the classifier and used nucleotides around Ψ sites as the features. Currently, PPUS: could accurately predict new Ψ sites for PUS1, PUS4 and PUS7 in yeast and PUS4 in human. PPUS: is well designed and friendly to user. AVAILABILITY AND IMPLEMENTATION: Our web server is available freely for non-commercial purposes at: http://lyh.pkmu.cn/ppus/ CONTACT: [email protected] or [email protected]. Gaigai Zhang, Qinghua Cui |
Bioinform. | 3 |
| 2012 | The relationship between rational drug design and drug side effectsabstractPrevious analysis of systems pharmacology has revealed a tendency of rational drug design in the pharmaceutical industry. The targets of new drugs tend to be close with the corresponding disease genes in the biological networks. However, it remains unclear whether the rational drug design introduces disadvantages, i.e. side effects. Therefore, it is important to dissect the relationship between rational drug design and drug side effects. Based on a recently released drug side effect database, SIDER, here we analyzed the relationship between drug side effects and the rational drug design. We revealed that the incidence drug side effect is significantly associated with the network distance of drug targets and diseases genes. Drugs with the distances of three or four have the smallest incidence of side effects, whereas drugs with the distances of more than four or smaller than three show significantly greater incidence of side effects. Furthermore, protein drugs and small molecule drugs show significant differences. Drugs hitting membrane targets and drugs hitting cytoplasm targets also show differences. Failure drugs because of severe side effects show smaller network distances than approved drugs. These results suggest that researchers should be prudent on rationalizing the drug design. Too small distances between drug targets and diseases genes may not always be advantageous for rational design for drug discovery. Juan Wang 0018, Zhi-xin Li, Chengxiang Qiu, Qinghua Cui |
Briefings Bioinform. | 5 |
| 2011 | miREnvironment Database: providing a bridge for microRNAs, environmental factors and phenotypesabstractUNLABELLED: The interaction between genetic factors and environmental factors has critical roles in determining the phenotype of an organism. In recent years, a number of studies have reported that the dysfunctions on microRNA (miRNAs), environmental factors and their interactions have strong effects on phenotypes and even may result in abnormal phenotypes and diseases, whereas there has been no a database linking miRNAs, environmental factors and phenotypes. Such a resource platform is believed to be of great value in the understanding of miRNAs, environmental factors, especially drugs and diseases. In this study, we constructed the miREnvironment database, which contains a comprehensive collection and curation of experimentally supported interactions among miRNAs, environmental factors and phenotypes. The names of miRNAs, phenotypes, environmental factors, conditions of environmental factors, samples, species, evidence and references were further annotated. miREnvironment represents a biomedical resource for researches on miRNAs, environmental factors and diseases. AVAILABILITY: http://cmbi.bjmu.edu.cn/miren. CONTACT: [email protected]. Chengxiang Qiu, Qinghua Cui |
Bioinform. | 5 |
| 2010 | Inferring the human microRNA functional similarity and functional network based on microRNA-associated diseasesabstractMOTIVATION: It is popular to explore meaningful molecular targets and infer new functions of genes through gene functional similarity measuring and gene functional network construction. However, little work is available in this field for microRNA (miRNA) genes due to limited miRNA functional annotations. With the rapid accumulation of miRNAs, it is increasingly needed to uncover their functional relationships in a systems level. RESULTS: It is known that genes with similar functions are often associated with similar diseases, and the relationship of different diseases can be represented by a structure of directed acyclic graph (DAG). This is also true for miRNA genes. Therefore, it is feasible to infer miRNA functional similarity by measuring the similarity of their associated disease DAG. Based on the above observations and the rapidly accumulated human miRNA-disease association data, we presented a method to infer the pairwise functional similarity and functional network for human miRNAs based on the structures of their disease relationships. Comparisons showed that the calculated miRNA functional similarity is well associated with prior knowledge of miRNA functional relationship. More importantly, this method can also be used to predict novel miRNA biomarkers and to infer novel potential functions or associated diseases for miRNAs. In addition, this method can be easily extended to other species when sufficient miRNA-associated disease data are available for specific species. AVAILABILITY: The online tool is available at http://cmbi.bjmu.edu.cn/misim CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Juan Wang 0018, Qinghua Cui |
Bioinform. | 5 |
| 2010 | TAM: A method for enrichment and depletion analysis of a microRNA category in a list of microRNAsabstractBACKGROUND: MicroRNAs (miRNAs) are a class of important gene regulators. The number of identified miRNAs has been increasing dramatically in recent years. An emerging major challenge is the interpretation of the genome-scale miRNA datasets, including those derived from microarray and deep-sequencing. It is interesting and important to know the common rules or patterns behind a list of miRNAs, (i.e. the deregulated miRNAs resulted from an experiment of miRNA microarray or deep-sequencing). RESULTS: For the above purpose, this study presents a method and develops a tool (TAM) for annotations of meaningful human miRNAs categories. We first integrated miRNAs into various meaningful categories according to prior knowledge, such as miRNA family, miRNA cluster, miRNA function, miRNA associated diseases, and tissue specificity. Using TAM, given lists of miRNAs can be rapidly annotated and summarized according to the integrated miRNA categorical data. Moreover, given a list of miRNAs, TAM can be used to predict novel related miRNAs. Finally, we confirmed the usefulness and reliability of TAM by applying it to deregulated miRNAs in acute myocardial infarction (AMI) from two independent experiments. CONCLUSION: TAM can efficiently identify meaningful categories for given miRNAs. In addition, TAM can be used to identify novel miRNA biomarkers. TAM tool, source codes, and miRNA category data are freely available at http://cmbi.bjmu.edu.cn/tam. Juan Wang 0018, Qun Cao, Qinghua Cui |
BMC Bioinform. | 5 |
| 2005 | Characterizing the dynamic connectivity between genes by variable parameter regression and Kalman filtering based on temporal gene expression dataabstractMOTIVATION: One popular method for analyzing functional connectivity between genes is to cluster genes with similar expression profiles. The most popular metrics measuring the similarity (or dissimilarity) among genes include Pearson's correlation, linear regression coefficient and Euclidean distance. As these metrics only give some constant values, they can only depict a stationary connectivity between genes. However, the functional connectivity between genes usually changes with time. Here, we introduce a novel insight for characterizing the relationship between genes and find out a proper mathematical model, variable parameter regression and Kalman filtering to model it. RESULTS: We applied our algorithm to some simulated data and two pairs of real gene expression data. The changes of connectivity in simulated data are closely identical with the truth and the results of two pairs of gene expression data show that our method has successfully demonstrated the dynamic connectivity between genes. CONTACT: [email protected]. Qinghua Cui, Bing Liu 0008, Tianzi Jiang, Songde Ma |
Bioinform. | 1 |
| 2004 | Esub8: A novel tool to predict protein subcellular localizations in eukaryotic organismsabstractBACKGROUND: Subcellular localization of a new protein sequence is very important and fruitful for understanding its function. As the number of new genomes has dramatically increased over recent years, a reliable and efficient system to predict protein subcellular location is urgently needed. RESULTS: Esub8 was developed to predict protein subcellular localizations for eukaryotic proteins based on amino acid composition. In this research, the proteins are classified into the following eight groups: chloroplast, cytoplasm, extracellular, Golgi apparatus, lysosome, mitochondria, nucleus and peroxisome. We know subcellular localization is a typical classification problem; consequently, a one-against-one (1-v-1) multi-class support vector machine was introduced to construct the classifier. Unlike previous methods, ours considers the order information of protein sequences by a different method. Our method is tested in three subcellular localization predictions for prokaryotic proteins and four subcellular localization predictions for eukaryotic proteins on Reinhardt's dataset. The results are then compared to several other methods. The total prediction accuracies of two tests are both 100% by a self-consistency test, and are 92.9% and 84.14% by the jackknife test, respectively. Esub8 also provides excellent results: the total prediction accuracies are 100% by a self-consistency test and 87% by the jackknife test. CONCLUSIONS: Our method represents a different approach for predicting protein subcellular localization and achieved a satisfactory result; furthermore, we believe Esub8 will be a useful tool for predicting protein subcellular localizations in eukaryotic organisms. Qinghua Cui, Tianzi Jiang, Bing Liu 0008, Songde Ma |
BMC Bioinform. | 1 |
| 2004 | A combinational feature selection and ensemble neural network method for classification of gene expression dataabstractBACKGROUND: Microarray experiments are becoming a powerful tool for clinical diagnosis, as they have the potential to discover gene expression patterns that are characteristic for a particular disease. To date, this problem has received most attention in the context of cancer research, especially in tumor classification. Various feature selection methods and classifier design strategies also have been generally used and compared. However, most published articles on tumor classification have applied a certain technique to a certain dataset, and recently several researchers compared these techniques based on several public datasets. But, it has been verified that differently selected features reflect different aspects of the dataset and some selected features can obtain better solutions on some certain problems. At the same time, faced with a large amount of microarray data with little knowledge, it is difficult to find the intrinsic characteristics using traditional methods. In this paper, we attempt to introduce a combinational feature selection method in conjunction with ensemble neural networks to generally improve the accuracy and robustness of sample classification. RESULTS: We validate our new method on several recent publicly available datasets both with predictive accuracy of testing samples and through cross validation. Compared with the best performance of other current methods, remarkably improved results can be obtained using our new strategy on a wide range of different datasets. CONCLUSIONS: Thus, we conclude that our methods can obtain more information in microarray data to get more accurate classification and also can help to extract the latent marker genes of the diseases for better diagnosis and treatment. Bing Liu 0008, Qinghua Cui, Tianzi Jiang, Songde Ma |
BMC Bioinform. | 2 |