Ruiming Li

dblp:50/416 · DBLP profile ↗
← Back
12ranked-venue papers
6as first author
4since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 4 first-authorApplied, interdisciplinary, general and emerging computing · 4 · 2 first-author · 2 since 2021Theory of computation · 2 · 1 since 2021Software engineering, systems software and programming languages · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Visual analysis of LLM-based entity resolution from scientific papers
abstract
This paper focuses on the visual analytics support for extracting domain-specific entities from extensive scientific literature, a task with inherent limitations using traditional named entity resolution methods. With the advent of large language models (LLMs) such as GPT-4, significant improvements over conventional machine learning approaches have been achieved due to LLM’s capability on entity resolution integrate abilities such as understanding multiple types of text. This research introduces a new visual analysis pipeline that integrates these advanced LLMs with versatile visualization and interaction designs to support batch entity resolution. Specifically, we focus on a specific material science field of Metal-Organic Frameworks (MOFs) and a large data collection namely CSD-MOFs. Through collaboration with domain experts in material science, we obtain well-labeled synthesis paragraphs. We propose human-in-the-loop refinement over the entity resolution process using visual analytics techniques, which allows domain experts to interactively integrate insights into LLM intelligence, including error analysis and interpretation of the retrieval-augmented generation (RAG) algorithm. Our evaluation through the case study of example selection for RAG demonstrates that this visual analysis approach effectively improves the accuracy of single-document entity resolution.
Weize Wu, Ruiming Li, Huobin Tan, Zipeng Liu
Vis. Informatics4
2022 Densest subgraph-based methods for protein-protein interaction hot spot prediction
abstract
BACKGROUND: Hot spots play an important role in protein binding analysis. The residue interaction network is a key point in hot spot prediction, and several graph theory-based methods have been proposed to detect hot spots. Although the existing methods can yield some interesting residues by network analysis, low recall has limited their abilities in finding more potential hot spots. RESULT: In this study, we develop three graph theory-based methods to predict hot spots from only a single residue interaction network. We detect the important residues by finding subgraphs with high densities, i.e., high average degrees. Generally, a high degree implies a high binding possibility between protein chains, and thus a subgraph with high density usually relates to binding sites that have a high rate of hot spots. By evaluating the results on 67 complexes from the SKEMPI database, our methods clearly outperform existing graph theory-based methods on recall and F-score. In particular, our main method, Min-SDS, has an average recall of over 0.665 and an f2-score of over 0.364, while the recall and f2-score of the existing methods are less than 0.400 and 0.224, respectively. CONCLUSION: The Min-SDS method performs best among all tested methods on the hot spot prediction problem, and all three of our methods provide useful approaches for analyzing bionetworks. In addition, the densest subgraph-based methods predict hot spots with only one residue interaction network, which is constructed from spatial atomic coordinate data to mitigate the shortage of data from wet-lab experiments.
Ruiming Li, Jung-Yu Lee 0001, Jinn-Moon Yang, Tatsuya Akutsu
BMC Bioinform.1
2021 Weighted minimum feedback vertex sets and implementation in human cancer genes detection
abstract
BACKGROUND: Recently, many computational methods have been proposed to predict cancer genes. One typical kind of method is to find the differentially expressed genes between tumour and normal samples. However, there are also some genes, for example, 'dark' genes, that play important roles at the network level but are difficult to find by traditional differential gene expression analysis. In addition, network controllability methods, such as the minimum feedback vertex set (MFVS) method, have been used frequently in cancer gene prediction. However, the weights of vertices (or genes) are ignored in the traditional MFVS methods, leading to difficulty in finding the optimal solution because of the existence of many possible MFVSs. RESULTS: Here, we introduce a novel method, called weighted MFVS (WMFVS), which integrates the gene differential expression value with MFVS to select the maximum-weighted MFVS from all possible MFVSs in a protein interaction network. Our experimental results show that WMFVS achieves better performance than using traditional bio-data or network-data analyses alone. CONCLUSION: This method balances the advantage of differential gene expression analyses and network analyses, improves the low accuracy of differential gene expression analyses and decreases the instability of pure network analyses. Furthermore, WMFVS can be easily applied to various kinds of networks, providing a useful framework for data analysis and prediction.
Ruiming Li, Chun-Yu Lin 0003, Weifeng Guo, Tatsuya Akutsu
BMC Bioinform.1
2021 New and improved algorithms for unordered tree inclusion
Tatsuya Akutsu, Jesper Jansson 0001, Ruiming Li, Atsuhiro Takasu, Takeyuki Tamura
Theor. Comput. Sci.3
2018 Deep Learning with Evolutionary and Genomic Profiles for Identifying Cancer Subtypes
abstract
Cancer subtype identification is an unmet need in precision diagnosis. Recently, evolutionary conservation has been indicated containing understandable signatures for functional significance in cancers. However, the importance of evolutionary conservation in distinguishing cancer subtypes remains unclear. Here, we identified the evolutionarily conserved genes (i.e., core gene) and observed that they are mainly involved in the pathways relevant to cell growth and metabolisms. By using these core genes, we integrated their evolutionary and genomic profiles with deep learning to develop a feature-based strategy (FES) and an image-based strategy (IMS). In comparison with FES using the random set and the strategy using the PAM50 classifier, core gene set-based FES has higher accuracy for identifying breast cancer subtypes. Moreover, the IMS with data augmentation yields better performance than the other strategies. Comprehensive analysis of eight TCGA cancer data demonstrates that our evolutionary conservation-based models provide a valid and helpful approach to identify cancer subtypes and the core gene set offers distinguishable clues of cancer subtypes.
Chun-Yu Lin 0003, Peiying Ruan, Ruiming Li, Jinn-Moon Yang, Simon See, Tatsuya Akutsu
BIBE3
2018 New and Improved Algorithms for Unordered Tree Inclusion
abstract
The tree inclusion problem is, given two node-labeled trees P and T (the "pattern tree" and the "text tree"), to locate every minimal subtree in T (if any) that can be obtained by applying a sequence of node insertion operations to P. Although the ordered tree inclusion problem is solvable in polynomial time, the unordered tree inclusion problem is NP-hard. The currently fastest algorithm for the latter is from 1995 and runs in O(poly(m,n) * 2^{2d}) = O^*(2^{2d}) time, where m and n are the sizes of the pattern and text trees, respectively, and d is the maximum outdegree of the pattern tree. Here, we develop a new algorithm that improves the exponent 2d to d by considering a particular type of ancestor-descendant relationships and applying dynamic programming, thus reducing the time complexity to O^*(2^d). We then study restricted variants of the unordered tree inclusion problem where the number of occurrences of different node labels and/or the input trees' heights are bounded. We show that although the problem remains NP-hard in many such cases, it can be solved in polynomial time for c = 2 and in O^*(1.8^d) time for c = 3 if the leaves of P are distinctly labeled and each label occurs at most c times in T. We also present a randomized O^*(1.883^d)-time algorithm for the case that the heights of P and T are one and two, respectively.
Tatsuya Akutsu, Jesper Jansson 0001, Ruiming Li, Atsuhiro Takasu, Takeyuki Tamura
ISAAC3
2008 Incorporating logic exclusivity (LE) constraints in noise analysis using gain guided backtracking method
abstract
Crosstalk noise becomes one of the critical issues gating design closure for nano-meter designs. Pessimism in noise analysis can lead to significant additional time spent addressing false violations. Taking logic correlation into consideration, noise analysis can reduce pessimism significantly by eliminating false noise signals [1]-[3][5]-[7][10]-[13]. Eliminating the aggressors from the aggressor candidate set that can not switch simultaneously restricted by the logic exclusivity (LE) relationship among them can save simulation time as well. The LE problem, being proved as NP-complete, is basically to determine the subset (possibly multiple equivalent subsets) of a given aggressor candidate set which has the largest combined weight out of all possible subsets governed by logic exclusivity constraints. This paper presents a new approach in resolving the LE problem, which employs a gain guided backtrack search technique that does not require exhaustive search of all the binary paths to reach an optimal solution. We first prove that under certain conditions, if the gain at each level is non-negative, then the result will be optimal. Based on this theorem, a new algorithm is developed. The experimental results demonstrate the efficiency and accuracy of this approach. The algorithm can quickly find the optimal solutions for most cases from industry designs and outperforms other methods.
Ruiming Li, An-Jui Shey, Michel Laudes
ICCAD1
2005 Design and Verification of High-Speed VLSI Physical Design
Dian Zhou, Ruiming Li
J. Comput. Sci. Technol.2
2005 Power-optimal simultaneous buffer insertion/sizing and wire sizing for two-pin nets
abstract
This paper studies the problems of optimizing power dissipation for simultaneous buffer insertion/sizing and uniform wire sizing (BISUWS), and simultaneous buffer insertion/sizing and tapered wire sizing (BISTWS). For BISUWS, we analyze the optimal total power dissipation under the delay constraints as well as the power-delay tradeoff. For BISTWS, we study the problems of minimizing power dissipation with optimal delay constraints or with a given delay penalty. We derive optimal solutions for both cases. These solutions can be used to efficiently estimate the power dissipation for long single wires in the interconnect designs.
Ruiming Li, Dian Zhou, Jin Liu 0004, Xuan Zeng 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2004 Satisfiability and integer programming as complementary tools
Ruiming Li, Dian Zhou, Donglei Du
ASP-DAC1
2004 Steady-State Analysis of Nonlinear Circuits Using Discrete Singular Convolution Method
abstract
In this paper, we propose a novel time-domain based method, discrete singular convolution algorithm, for computing steady-state response in nonlinear circuit. Properties and advantages of discrete singular convolution method are discussed, compared with some other approaches. The accuracy and efficiency of this method are tested by the numerical experiments.
Dian Zhou, Jin Liu 0004, Ruiming Li, Xuan Zeng 0001, Charles C. Chiang
DATE4
2003 Power-Optimal Simultaneous Buffer Insertion/Sizing and Wire Sizing
Ruiming Li, Dian Zhou, Jin Liu 0004, Xuan Zeng 0001
ICCAD1