Jifan Shi

dblp:262/8360 · DBLP profile ↗
← Back
8ranked-venue papers
5as first author
5since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 4 · 3 first-author · 1 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
1 paper
Graph data management · 67% Indexing and storage engines · 33%
Artificial intelligence
1 paper
Probabilistic and Bayesian machine learning · 100%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Bioinformatics and computational biology · 100%
Theoretical computer science
1 paper
Information theory · 100%

Topics — the 5 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Probabilistic and Bayesian machine learning
causal inference
1.012026
Dynamical Causality Under Latent Confounders for Biological Network Reconstruction · IEEE Trans. Pattern Anal. Mach. Intell. 2026
Bioinformatics and computational biology › network bioinformatics › biological network analysis
biological network inference
1.012026
Dynamical Causality Under Latent Confounders for Biological Network Reconstruction · IEEE Trans. Pattern Anal. Mach. Intell. 2026
Graph data management
dynamic graph
0.812024
Spruce: a Fast yet Space-saving Structure for Dynamic Graph Storage · Proc. ACM Manag. Data 2024
Graph data management › graph storage
dynamic graph storage
0.812024
Spruce: a Fast yet Space-saving Structure for Dynamic Graph Storage · Proc. ACM Manag. Data 2024
Indexing and storage engines
in-memory index
0.812024
Spruce: a Fast yet Space-saving Structure for Dynamic Graph Storage · Proc. ACM Manag. Data 2024

Methods — techniques the papers use, named apart from their topics

orthogonal decomposition · 3.0delay embedding · 3.0van emde boas tree · 0.8optimistic locking · 0.8
YearPublicationVenuePosition
2026 Dynamical Causality Under Latent Confounders for Biological Network Reconstruction
abstract
Causal interaction inference is prone to spurious causal interactions, due to the substantial confounders in a biological system. While many existing methods attempt to address misidentification challenges, there remains a notable lack of effective methods to infer causal interaction under latent/unobserved confounders. In this work, we propose a method to overcome such challenges to infer dynamical causality under invisible confounders (CIC) and further reconstruct the latent confounders from time-series data by developing an orthogonal decomposition theorem in a delay embedding space. This theoretical foundation ensures the causal detection for any high-dimensional system even with only two observed variables under many latent confounders, which is a long-standing problem in the field. In addition to the latent confounder problem, such a decomposition makes the coupled variables separable in the embedding space, thus also solving the non-separability problem of causal inference. Extensive validation of the CIC method is carried out using various real datasets, which all demonstrates its effectiveness to reconstruct real biological networks and unobserved confounders.
Jinling Yan, Shaowu Zhang 0001, Chihao Zhang 0002, Weitian Huang, Jifan Shi, Luonan Chen
IEEE Trans. Pattern Anal. Mach. Intell.5
2025 HiCHT: High-Performance Compact Hash Table
Jifan Shi
DASFAA (4)3
2025 diffMIN: Reconstructing Sample-Specific Differential Gene Regulatory Networks Based on Mutual Information
abstract
Gene regulatory mechanisms are pivotal in the study of biological systems. Sample-specific networks (SSNs) have recently proven highly effective in scrutinizing diverse biological processes more precisely. Utilizing SSNs as features for clustering, personalized diagnosis, and treatment has gained increasing significance, surpassing the conventional reliance on differential gene expression. Several computational methods have been developed for the inference of SSNs. However, the existing methods are currently limited to constructing gene regulatory networks for individual samples. We proposed a new method for reconstructing sample-specific differential gene regulatory networks based on mutual information (diffMIN). Analysis of various single-cell datasets demonstrated that diffMIN outperforms existing SSNs algorithms regarding the cell clustering. We utilized diffMIN to analyze transverse aortic constriction (TAC) single-cell data, enabling us to generate differential networks highlighting the changes in pressure-overload, ultimately identifying a monotonic change in regulatory networks during heart failure. Additionally, we applied diffMIN on breast cancer samples from The Cancer Genome Atlas (TCGA). We found that, diffMIN can effectively identify important genes that may not exhibit significant expressional differences but are highly correlated with patient prognosis. This discovery highlights the potential utility of diffMIN in identifying crucial genes in cancer. In conclusion, our method is widely applicable and performs well in cell clustering, identification of key genes, investigation of biological processes, and disease diagnosis by constructing sample-specific differential networks.
Songwei Ai, Yiwei Zhou, Wenkui Zou, Emeli Chatterjee, Jifan Shi
IEEE Trans. Comput. Biol. Bioinform.6
2024 Spruce: a Fast yet Space-saving Structure for Dynamic Graph Storage
abstract
Dynamic graphs have been gaining increasing popularity across various application domains. With the growing size of these graphs, the update performance as well as space occupancy is becoming a crucial aspect of dynamic graph storage. Although existing dynamic graph systems can handle massive streaming updates (e.g., insertions and deletions), they cannot achieve both high throughput and low memory footprint. Drawing inspiration from the basic operations of the van Emde Boas (vEB) tree in double-logarithmic time, we designed Spruce, a high-performance yet space-saving in-memory structure to store dynamic graphs. Spruce uses a compact representation to construct the tree-like multilevel structure, which shares the common prefixes of vertices and has no merging or splitting of nodes to achieve the requirements of low memory consumption and high-efficiency dynamic operations. Furthermore, Spruce incorporates a read-optimized concurrency protocol, which refines ROWEX and Optimistic Locking, to facilitate efficient simultaneous read/write operations. Our experiment demonstrates that compared to Sortledton (the best of competitors), Spruce is up to 2.4X faster in ingesting graph updates, while saving up to 38.5% of memory space. As for graph analytics, Spruce shows high adaptability to different analytical workloads, and achieves comparable performance to other state-of-the-art dynamic graph structures.
Jifan Shi
Proc. ACM Manag. Data1
2023 ABi-BFS: A High-performance Parallel Breadth-First Search on Shared-memory Systems
abstract
Breadth-first search (BFS) is a cornerstone in graph traversal, widely employed in areas such as social network analysis, routing algorithms, and biological network exploration. As the size of these graphs increases, the performance of BFS becomes significant. There are two key factors that influence BFS performance: synchronization overhead and memory access efficiency. To address them, we propose the ABi-BFS, an asynchronous and bidirectional optimized breadth-first search algorithm. ABi-BFS novelly designs an asynchronous mode for node-expanding and task-assignment processes to decrease the synchronization overhead and introduces an optimized strategy to select the frontier nodes in the bottom-up step of bidirectional search to lower the number of node accesses. Our experiment demonstrates that ABi-BFS outperforms other implementations of state-of-the-art PBFS algorithms on various graphs and achieves the highest scalability among them on shared-memory systems.
Jifan Shi
ICPADS1
2020 Quantifying Waddington's epigenetic landscape: a comparison of single-cell potency measures
abstract
MOTIVATION: Estimating differentiation potency of single cells is a task of great biological and clinical significance, as it may allow identification of normal and cancer stem cell phenotypes. However, very few single-cell potency models have been proposed, and their robustness and reliability across independent studies have not yet been fully assessed. RESULTS: Using nine independent single-cell RNA-Seq experiments, we here compare four different single-cell potency models to each other, in their ability to discriminate cells that ought to differ in terms of differentiation potency. Two of the potency models approximate potency via network entropy measures that integrate the single-cell RNA-Seq profile of a cell with a protein interaction network. The comparison between the four models reveals that integration of RNA-Seq data with a protein interaction network dramatically improves the robustness and reliability of single-cell potency estimates. We demonstrate that underlying this robustness is a correlation relationship, according to which high differentiation potency is positively associated with overexpression of network hubs. We further show that overexpressed network hubs are strongly enriched for ribosomal mitochondrial proteins, suggesting that their mRNA levels may provide a universal marker of a cell's potency. Thus, this study provides novel systems-biological insight into cellular potency and may provide a foundation for improved models of differentiation potency with far-reaching implications for the discovery of novel stem cell or progenitor cell phenotypes.
Jifan Shi, Andrew E. Teschendorff, Luonan Chen
Briefings Bioinform.1
2020 Quantifying Direct Dependencies in Biological Networks by Multiscale Association Analysis
abstract
Partial correlation (PC) or conditional mutual information (CMI) is widely used in detecting direct dependencies between the observed variables in biological networks by eliminating indirect correlations/associations, but it fails whenever there are some strong correlations in a network. In this paper, we theoretically develop a multiscale association analysis to overcome this flaw. We propose a new measure, partial association (PA), based on the multiscale conditional mutual information. We show that linear PA and nonlinear PA have clear advantages over PC and CMI from both theoretical and computational aspects. Both simulated models and real omics datasets demonstrate that PA is superior to PC and CMI in terms of accuracy, and is a powerful tool to identify the direct associations or reconstruct molecular networks based on the observed data. Survival and functional analyses of the hub genes in the gene networks reconstructed from TCGA data for different cancers also validated the effectiveness of our method.
Jifan Shi, Xiaoping Liu 0002, Luonan Chen
IEEE ACM Trans. Comput. Biol. Bioinform.1
2019 Quantifying pluripotency landscape of cell differentiation from scRNA-seq data by continuous birth-death process
abstract
Modeling cell differentiation from omics data is an essential problem in systems biology research. Although many algorithms have been established to analyze scRNA-seq data, approaches to infer the pseudo-time of cells or quantify their potency have not yet been satisfactorily solved. Here, we propose the Landscape of Differentiation Dynamics (LDD) method, which calculates cell potentials and constructs their differentiation landscape by a continuous birth-death process from scRNA-seq data. From the viewpoint of stochastic dynamics, we exploited the features of the differentiation process and quantified the differentiation landscape based on the source-sink diffusion process. In comparison with other scRNA-seq methods in seven benchmark datasets, we found that LDD could accurately and efficiently build the evolution tree of cells with pseudo-time, in particular quantifying their differentiation landscape in terms of potency. This study provides not only a computational tool to quantify cell potency or the Waddington potential landscape based on scRNA-seq data, but also novel insights to understand the cell differentiation process from a dynamic perspective.
Jifan Shi, Luonan Chen, Kazuyuki Aihara
PLoS Comput. Biol.1