Ruijiang Li

dblp:77/1298 · DBLP profile ↗
← Back
21ranked-venue papers
5as first author
6since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 12 · 1 first-author · 6 since 2021Artificial intelligence and machine learning · 8 · 3 first-authorGraphics, computer vision, multimedia, augmented reality and games · 5 · 3 first-author · 1 since 2021
YearPublicationVenuePosition
2025 Active Learning for Protein Structure Prediction
Zexin Xue, Michael Bailey, Ruijiang Li, Alejandro Corrochano-Navarro, Sizhen Li, Lorenzo Kogler-Anele, Qui Yu, Ziv Bar-Joseph, Sven Jager
RECOMB4
2024 Single-Cell Heterogeneity-Aware Transformer-Guided Multiple Instance Learning for Cancer Aneuploidy Prediction From Whole Slide Histopathology Images
abstract
Aneuploidy is a hallmark of aggressive malignancies associated with therapeutic resistance and poor survival. Measuring aneuploidy requires expensive specialized techniques that are not clinically applicable. Deep learning analysis of routine histopathology slides has revealed associations with genetic mutations. However, existing studies focus on image patches or tiles, and there is no prior work that predicts aneuploidy using single-cell analysis. Here, we present a single-cell heterogeneity-aware and transformer-guided deep learning framework to predict aneuploidy from whole slide histopathology images. First, we perform nuclei segmentation and classification to obtain individual cancer cells, which are clustered into multiple subtypes. The cell subtype distributions are computed to measure cancer cell heterogeneity. Additionally, morphological features of different cell subtypes are extracted. Further, we leverage a multiple instance learning module with Transformer, which encourages the network to focus on the most informative cancer cells. Lastly, a hybrid network is built to unify cell heterogeneity, morphology, and deep features for aneuploidy prediction. We train and validate our method on two public datasets from TCGA: lung adenocarcinoma (LUAD) and head and neck squamous cell carcinoma (HNSC), with 339 and 245 patients. Our model achieves promising performance with AUC of 0.818 (95% CI: 0.718-0.919) and 0.827 (95% CI: 0.704-0.949) on the LUAD and HNSC test sets, respectively. Through extensive ablation and comparison studies, we demonstrate the effectiveness of each component of the model and superior performance over alternative networks. In conclusion, we present a novel deep learning approach to predict aneuploidy from histopathology images, which could inform personalized cancer treatment.
Feiyang Yu, Rasoul Sali, Ruijiang Li
IEEE J. Biomed. Health Informatics4
2023 Deep Learning for Cell Classification in Histopathology Images Using Large-Scale Manually Annotated Datasets
abstract
Accurate cell classification in histopathology images is critical for investigating biological mechanisms behind disease progression and discovering interpretable biomarkers for precision medicine. However, major challenges remain including a limited amount of annotated data and inconsistency in the annotated cell types across different datasets. Here, we combine two manually annotated datasets to curate a single large dataset consisting of approximately 206,000 nuclei images for five common cell types in the tumor microenvironment. We compare three deep learning models (ResNet18, ConvNeXt, and Efficient-net) trained with three different loss functions for cell classification. The best-performing model achieves an overall accuracy of 91.5% (range 63-98%) with an average one vs rest micro area under curve (AUC) of 0.98. This study provides a comprehensive evaluation of the state-of-the-art deep learning models using a benchmark dataset and demonstrates the need for further improvement for accurately classifying specific cell types.
Leon Liu, Ruijiang Li
BIBE3
2023 Semi-supervised Class Imbalanced Deep Learning for Cardiac MRI Segmentation
Yuchen Yuan, Xi Wang 0013, Xikai Yang, Ruijiang Li, Pheng-Ann Heng
MICCAI (4)4
2023 Survival Prediction via Hierarchical Multimodal Co-Attention Transformer: A Computational Histology-Radiology Solution
abstract
The rapid advances in deep learning-based computational pathology and radiology have demonstrated the promise of using whole slide images (WSIs) and radiology images for survival prediction in cancer patients. However, most image-based survival prediction methods are limited to using either histology or radiology alone, leaving integrated approaches across histology and radiology relatively underdeveloped. There are two main challenges in integrating WSIs and radiology images: (1) the gigapixel nature of WSIs and (2) the vast difference in spatial scales between WSIs and radiology images. To address these challenges, in this work, we propose an interpretable, weakly-supervised, multimodal learning framework, called Hierarchical Multimodal Co-Attention Transformer (HMCAT), to integrate WSIs and radiology images for survival prediction. Our approach first uses hierarchical feature extractors to capture various information including cellular features, cellular organization, and tissue phenotypes in WSIs. Then the hierarchical radiology-guided co- attention (HRCA) in HMCAT characterizes the multimodal interactions between hierarchical histology-based visual concepts and radiology features and learns hierarchical co- attention mappings for two modalities. Finally, HMCAT combines their complementary information into a multimodal risk score and discovers prognostic features from two modalities by multimodal interpretability. We apply our approach to two cancer datasets (365 WSIs with matched magnetic resonance [MR] images and 213 WSIs with matched computed tomography [CT] images). Our results demonstrate that the proposed HMCAT consistently achieves superior performance over the unimodal approaches trained on either histology or radiology data alone, as well as other state-of-the-art methods.
Zhe Li 0006, Yuming Jiang 0005, Mengkang Lu, Ruijiang Li, Yong Xia 0001
IEEE Trans. Medical Imaging4
2021 Triple attention learning for classification of 14 thoracic diseases using chest radiography
Hongyu Wang 0011, Shanshan Wang 0002, Zibo Qin, Yanning Zhang 0001, Ruijiang Li, Yong Xia 0001
Medical Image Anal.5
2020 New insights on human essential genes based on integrated analysis and the construction of the HEGIAP web-based platform
abstract
Essential genes are those whose loss of function compromises organism viability or results in profound loss of fitness. Recent gene-editing technologies have provided new opportunities to characterize essential genes. Here, we present an integrated analysis that comprehensively and systematically elucidates the genetic and regulatory characteristics of human essential genes. First, we found that essential genes act as 'hubs' in protein-protein interaction networks, chromatin structure and epigenetic modification. Second, essential genes represent conserved biological processes across species, although gene essentiality changes differently among species. Third, essential genes are important for cell development due to their discriminate transcription activity in embryo development and oncogenesis. In addition, we developed an interactive web server, the Human Essential Genes Interactive Analysis Platform (http://sysomics.com/HEGIAP/), which integrates abundant analytical tools to enable global, multidimensional interpretation of gene essentiality. Our study provides new insights that improve the understanding of human essential genes.
Hebing Chen, Ruijiang Li, Chenghui Zhao, Hao Hong, Xin Huang 0005, Hao Li 0035, Xiaochen Bo
Briefings Bioinform.4
2020 DeepHiC: A generative adversarial network for enhancing Hi-C data resolution
abstract
Hi-C is commonly used to study three-dimensional genome organization. However, due to the high sequencing cost and technical constraints, the resolution of most Hi-C datasets is coarse, resulting in a loss of information and biological interpretability. Here we develop DeepHiC, a generative adversarial network, to predict high-resolution Hi-C contact maps from low-coverage sequencing data. We demonstrated that DeepHiC is capable of reproducing high-resolution Hi-C data from as few as 1% downsampled reads. Empowered by adversarial training, our method can restore fine-grained details similar to those in high-resolution Hi-C matrices, boosting accuracy in chromatin loops identification and TADs detection, and outperforms the state-of-the-art methods in accuracy of prediction. Finally, application of DeepHiC to Hi-C data on mouse embryonic development can facilitate chromatin loop detection. We develop a web-based tool (DeepHiC, http://sysomics.com/deephic) that allows researchers to enhance their own Hi-C data with just a few clicks.
Hao Hong, Hao Li 0035, Guifang Du, Yu Sun 0070, Cheng Quan, Chenghui Zhao, Ruijiang Li, Xiaoyao Yin, Yangchen Huang, Hebing Chen, Xiaochen Bo
PLoS Comput. Biol.9
2019 Relation Embedding with Dihedral Group in Knowledge Graph
abstract
Link prediction is critical for the application of incomplete knowledge graph (KG) in the downstream tasks.As a family of effective approaches for link predictions, embedding methods try to learn low-rank representations for both entities and relations such that the bilinear form defined therein is a well-behaved scoring function.Despite of their successful performances, existing bilinear forms overlook the modeling of relation compositions, resulting in lacks of interpretability for reasoning on KG.To fulfill this gap, we propose a new model called DihEdral, named after dihedral symmetry group.This new model learns knowledge graph embeddings that can capture relation compositions by nature.Furthermore, our approach models the relation embeddings parametrized by discrete values, thereby decrease the solution space drastically.Our experiments show that DihEdral is able to capture all desired properties such as (skew-) symmetry, inversion and (non-) Abelian composition, and outperforms existing bilinear form based approach and is comparable to or better than deep learning models such as ConvE (Dettmers et al., 2018).
Canran Xu, Ruijiang Li
ACL (1)2
2019 A survey and evaluation of Web-based tools/databases for variant analysis of TCGA data
abstract
The Cancer Genome Atlas (TCGA) is a publicly funded project that aims to catalog and discover major cancer-causing genomic alterations with the goal of creating a comprehensive 'atlas' of cancer genomic profiles. The availability of this genome-wide information provides an unprecedented opportunity to expand our knowledge of tumourigenesis. Computational analytics and mining are frequently used as effective tools for exploring this byzantine series of biological and biomedical data. However, some of the more advanced computational tools are often difficult to understand or use, thereby limiting their application by scientists who do not have a strong computational background. Hence, it is of great importance to build user-friendly interfaces that allow both computational scientists and life scientists without a computational background to gain greater biological and medical insights. To that end, this survey was designed to systematically present available Web-based tools and facilitate the use TCGA data for cancer research.
Hao Li 0035, Ruijiang Li, Hebing Chen, Xiaochen Bo
Briefings Bioinform.4
2019 Stable H3K4me3 is associated with transcription initiation during early embryo development
abstract
MOTIVATION: During development of the mammalian embryo, histone modification H3K4me3 plays an important role in regulating gene expression and exhibits extensive reprograming on the parental genomes. In addition to these dramatic epigenetic changes, certain unchanging regulatory elements are also essential for embryonic development. RESULTS: Using large-scale H3K4me3 chromatin immunoprecipitation sequencing data, we identified a form of H3K4me3 that was present during all eight stages of the mouse embryo before implantation. This 'stable H3K4me3' was highly accessible and much longer than normal H3K4me3. Moreover, most of the stable H3K4me3 was in the promoter region and was enriched in higher chromatin architecture. Using in-depth analysis, we demonstrated that stable H3K4me3 was related to higher gene expression levels and transcriptional initiation during embryonic development. Furthermore, stable H3K4me3 was much more active in blood tumor cells than in normal blood cells, suggesting a potential mechanism of cancer progression. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Xin Huang 0005, Ruijiang Li, Hao Hong, Chenghui Zhao, Pingkun Zhou, Hebing Chen, Xiaochen Bo, Hao Li 0035
Bioinform.5
2017 Incorporating prior biological knowledge for network-based differential gene expression analysis using differentially weighted graphical LASSO
abstract
BACKGROUND: Conventional differential gene expression analysis by methods such as student's t-test, SAM, and Empirical Bayes often searches for statistically significant genes without considering the interactions among them. Network-based approaches provide a natural way to study these interactions and to investigate the rewiring interactions in disease versus control groups. In this paper, we apply weighted graphical LASSO (wgLASSO) algorithm to integrate a data-driven network model with prior biological knowledge (i.e., protein-protein interactions) for biological network inference. We propose a novel differentially weighted graphical LASSO (dwgLASSO) algorithm that builds group-specific networks and perform network-based differential gene expression analysis to select biomarker candidates by considering their topological differences between the groups. RESULTS: Through simulation, we showed that wgLASSO can achieve better performance in building biologically relevant networks than purely data-driven models (e.g., neighbor selection, graphical LASSO), even when only a moderate level of information is available as prior biological knowledge. We evaluated the performance of dwgLASSO for survival time prediction using two microarray breast cancer datasets previously reported by Bild et al. and van de Vijver et al. Compared with the top 10 significant genes selected by conventional differential gene expression analysis method, the top 10 significant genes selected by dwgLASSO in the dataset from Bild et al. led to a significantly improved survival time prediction in the independent dataset from van de Vijver et al. Among the 10 genes selected by dwgLASSO, UBE2S, SALL2, XBP1 and KIAA0922 have been confirmed by literature survey to be highly relevant in breast cancer biomarker discovery study. Additionally, we tested dwgLASSO on TCGA RNA-seq data acquired from patients with hepatocellular carcinoma (HCC) on tumors samples and their corresponding non-tumorous liver tissues. Improved sensitivity, specificity and area under curve (AUC) were observed when comparing dwgLASSO with conventional differential gene expression analysis method. CONCLUSIONS: The proposed network-based differential gene expression analysis algorithm dwgLASSO can achieve better performance than conventional differential gene expression analysis methods by integrating information at both gene expression and network topology levels. The incorporation of prior biological knowledge can lead to the identification of biologically meaningful genes in cancer biomarker studies.
Yiming Zuo 0002, Guoqiang Yu, Ruijiang Li, Habtom W. Ressom
BMC Bioinform.4
2015 Rating Knowledge Sharing in Cross-Domain Collaborative Filtering
abstract
Cross-domain collaborative filtering (CF) aims to share common rating knowledge across multiple related CF domains to boost the CF performance. In this paper, we view CF domains as a 2-D site-time coordinate system, on which multiple related domains, such as similar recommender sites or successive time-slices, can share group-level rating patterns. We propose a unified framework for cross-domain CF over the site-time coordinate system by sharing group-level rating patterns and imposing user/item dependence across domains. A generative model, say ratings over site-time (ROST), which can generate and predict ratings for multiple related CF domains, is developed as the basic model for the framework. We further introduce cross-domain user/item dependence into ROST and extend it to two real-world cross-domain CF scenarios: 1) ROST (sites) for alleviating rating sparsity in the target domain, where multiple similar sites are viewed as related CF domains and some items in the target domain depend on their correspondences in the related ones; and 2) ROST (time) for modeling user-interest drift over time, where a series of time-slices are viewed as related CF domains and a user at current time-slice depends on herself in the previous time-slice. All these ROST models are instances of the proposed unified framework. The experimental results show that ROST (sites) can effectively alleviate the sparsity problem to improve rating prediction performance and ROST (time) can clearly track and visualize user-interest drift over time.
Bin Li 0015, Xingquan Zhu 0001, Ruijiang Li, Chengqi Zhang
IEEE Trans. Cybern.3
2013 Multi-Stage Non-Negative Matrix Factorization for Monaural Singing Voice Separation
abstract
Separating singing voice from music accompaniment can be of interest for many applications such as melody extraction, singer identification, lyrics alignment and recognition, and content-based music retrieval. In this paper, a novel algorithm for singing voice separation in monaural mixtures is proposed. The algorithm consists of two stages, where non-negative matrix factorization (NMF) is applied to decompose the mixture spectrograms with long and short windows respectively. A spectral discontinuity thresholding method is devised for the long-window NMF to select out NMF components originating from pitched instrumental sounds, and a temporal discontinuity thresholding method is designed for the short-window NMF to pick out NMF components that are from percussive sounds. By eliminating the selected components, most pitched and percussive elements of the music accompaniment are filtered out from the input sound mixture, with little effect on the singing voice. Extensive testing on the MIR-1K public dataset of 1000 short audio clips and the Beach-Boys dataset of 14 full-track real-world songs showed that the proposed algorithm is both effective and efficient.
Bilei Zhu, Wei Li 0012, Ruijiang Li, Xiangyang Xue 0001
IEEE Trans. Speech Audio Process.3
2012 Groupwise Constrained Reconstruction for Subspace Clustering
Ruijiang Li, Bin Li 0015, Cheng Jin 0001, Xiangyang Xue 0001
ICML1
2011 Tracking User-Preference Varying Speed in Collaborative Filtering
abstract
In real-world recommender systems, some users are easily influenced by new products and whereas others are unwilling to change their minds. So the preference varying speeds for users are different. Based on this observation, we propose a dynamic nonlinear matrix factorization model for collaborative filtering, aimed to improve the rating prediction performance as well as track the preference varying speeds for different users. We assume that user-preference changes smoothly over time, and the preference varying speeds for users are different. These two assumptions are incorporated into the proposed model as prior knowledge on user feature vectors, which can be learned efficiently by MAP estimation. The experimental results show that our method not only achieves state-of-the-art performance in the rating prediction task, but also provides an effective way to track user-preference varying speed.
Ruijiang Li, Bin Li 0015, Cheng Jin 0001, Xiangyang Xue 0001, Xingquan Zhu 0001
AAAI1
2011 Cross-Domain Collaborative Filtering over Time
abstract
Collaborative filtering (CF) techniques recommend items to users based on their historical ratings. In real-world scenarios, user interests may drift over time since they are affected by moods, contexts, and pop culture trends. This leads to the fact that a user's historical ratings comprise many aspects of user interests spanning a long time period. However, at a certain time slice, one user's interest may only focus on one or a couple of aspects. Thus, CF techniques based on the entire historical ratings may recommend inappropriate items. In this paper, we consider modeling user-interest drift over time based on the assumption that each user has multiple counterparts over temporal domains and successive counterparts are closely related. We adopt the cross-domain CF framework to share the static group-level rating matrix across temporal domains, and let user-interest distribution over item groups drift slightly between successive temporal domains. The derived method is based on a Bayesian latent factor model which can be inferred using Gibbs sampling. Our experimental results show that our method can achieve state-of-the-art recommendation performance as well as explicitly track and visualize user-interest drift over time.
Bin Li 0015, Xingquan Zhu 0001, Ruijiang Li, Chengqi Zhang, Xiangyang Xue 0001, Xindong Wu 0001
IJCAI3
2010 Single-Projection Based Volumetric Image Reconstruction and 3D Tumor Localization in Real Time for Lung Cancer Radiotherapy
Ruijiang Li, Xun Jia, John H. Lewis, Xuejun Gu, Michael Folkerts, Chunhua Men, Steve B. Jiang
MICCAI (3)1
2009 A Dynamic CT Image Reconstruction Method by Inducing Prior Information from PCA Analysis
abstract
Under-sampling and insufficient data result in a big challenge in the reconstruction of x-ray computed tomographic (CT) images. In addition, patient's respiratory motion also deteriorates this reconstruction process as it normally leads to blurred outputs. In this work, we propose an iterative method with a combination of total variation (TV) regularization and principle component analysis (PCA) regularization. Partial prior knowledge of the CT images, obtained through PCA analysis of training images is incorporated in the reconstruction process. Numerical experiments are performed in the context of a fan-beam CT reconstruction, which shows advantages of our method over the ones with just TV regularization or just PCA regularization.
Xun Jia, Yifei Lou, Ruijiang Li, Xuejun Gu, John Levis, Steve B. Jiang
ICMLA3
2009 Markerless Fluoroscopic Gating for Lung Cancer Radiotherapy Using Generalized Linear Discriminant Analysis
abstract
Respiratory gated radiotherapy for lung cancer allows for more precise delivery of prescribed radiation dose to the tumor, while minimizing normal tissue complications. Techniques for fluoroscopic gating without implanted fiducial markers have been developed in a classification framework. Due to the high-dimensionality nature of the images, dimensionality reduction techniques such as principal component analysis (PCA) were used to preprocess the data. In this work, we have applied generalized linear discriminant analysis (GLDA) to the respiratory gating problem. The fundamental difference from conventional dimensionality reduction techniques is that GLDA explicitly takes into account the label information available in the training set and therefore is efficient for discrimination among classes. On average, GLDA was demonstrated to perform similarly with PCA trained with SVM at high nominal duty cycles and outperform PCA in terms of classification accuracy (CA) and target coverage (TC) at lower nominal duty cycle (20%). A major advantage of GLDA is its robustness, while CA and TC using PCA can be reduced by up to 10% depending on the data dimensionality. With only 1-dimensional feature vectors, GLDA is much more computationally efficient than PCA. Therefore, GLDA is an effective and efficient method for respiratory gating with markerless fluoroscopic images.
Ruijiang Li, John H. Lewis, Steve B. Jiang
ICMLA1
2007 A unifying criterion for instantaneous blind source separation based on correntropy
Ruijiang Li, Weifeng Liu 0016, José C. Príncipe
Signal Process.1