VLDB 2026 Research / reviewers in the wild / expert
Hsueh-Fen Juan
dblp:49/2453
· DBLP profile ↗
13ranked-venue papers
0as first author
4since 2021 · last 2026
0000-0003-4876-3309ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 13 · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | scDock: streamlining drug discovery targeting cell-cell communication via scRNA-seq analysis and molecular dockingabstractSUMMARY: Identifying drugs that target intercellular communication networks represents a promising therapeutic strategy, yet linking single-cell RNA sequencing (scRNA-seq) analysis to structure-based drug screening remains technically challenging and requires substantial bioinformatics expertise. We present scDock, an integrated and user-friendly pipeline that seamlessly connects scRNA-seq data processing, cell-cell communication inference, and molecular docking-based drug discovery. Through a single configuration file, users can execute the complete workflow, from raw scRNA-seq data to ranked drug candidates, without programming skills. scDock automates the identification of disease-relevant ligand-receptor interactions from scRNA-seq data and performs structure-based virtual screening against these communication targets using Protein Data Bank (PDB) or AlphaFold-predicted protein structures. The pipeline generates comprehensive outputs at each stage, enabling users to explore intercellular signaling alterations and discover therapeutic compounds targeting specific cell-cell communications. scDock addresses a critical gap by providing an accessible end-to-end solution for communication-targeted drug discovery from single-cell data. AVAILABILITY AND IMPLEMENTATION: scDock is freely available at https://doi.org/10.6084/m9.figshare.31370368 and https://github.com/Andrewneteye4343/scDock. It is implemented in R, Python, shell scripts, and supports Linux systems, including Ubuntu and Debian. Chen-Hao Huang 0003, Yen-Jen Oyang, Hsuan-Cheng Huang, Hsueh-Fen Juan |
Bioinform. | 4 |
| 2025 | A large language model framework for literature-based disease-gene association predictionabstractWith the exponential growth of biomedical literature, leveraging Large Language Models (LLMs) for automated medical knowledge understanding has become increasingly critical for advancing precision medicine. However, current approaches face significant challenges in reliability, verifiability, and scalability when extracting complex biological relationships from scientific literature using LLMs. To overcome the obstacles of LLM development in biomedical literature understating, we propose LORE, a novel unsupervised two-stage reading methodology with LLM that models literature as a knowledge graph of verifiable factual statements and, in turn, as semantic embeddings in Euclidean space. LORE captured essential gene pathogenicity information when applied to PubMed abstracts for large-scale understanding of disease-gene relationships. We demonstrated that modeling a latent pathogenic flow in the semantic embedding with supervision from the ClinVar database led to a 90% mean average precision in identifying relevant genes across 2097 diseases. This work provides a scalable and reproducible approach for leveraging LLMs in biomedical literature analysis, offering new opportunities for researchers to identify therapeutic targets efficiently. Peng-Hsuan Li, Yih-Yun Sun, Hsueh-Fen Juan, Chien-Yu Chen 0001, Huai-Kuang Tsai, Jia-Hsin Huang |
Briefings Bioinform. | 3 |
| 2023 | Context-dependent gene regulatory network reveals regulation dynamics and cell trajectories using unspliced transcriptsabstractGene regulatory networks govern complex gene expression programs in various biological phenomena, including embryonic development, cell fate decisions and oncogenesis. Single-cell techniques are increasingly being used to study gene expression, providing higher resolution than traditional approaches. However, inferring a comprehensive gene regulatory network across different cell types remains a challenge. Here, we propose to construct context-dependent gene regulatory networks (CDGRNs) from single-cell RNA sequencing data utilizing both spliced and unspliced transcript expression levels. A gene regulatory network is decomposed into subnetworks corresponding to different transcriptomic contexts. Each subnetwork comprises the consensus active regulation pairs of transcription factors and their target genes shared by a group of cells, inferred by a Gaussian mixture model. We find that the union of gene regulation pairs in all contexts is sufficient to reconstruct differentiation trajectories. Functions specific to the cell cycle, cell differentiation or tissue-specific functions are enriched throughout the developmental process in each context. Surprisingly, we also observe that the network entropy of CDGRNs decreases along differentiation trajectories, indicating directionality in differentiation. Overall, CDGRN allows us to establish the connection between gene regulation at the molecular level and cell differentiation at the macroscopic level. Yueh-Hua Tu, Hsueh-Fen Juan, Hsuan-Cheng Huang |
Briefings Bioinform. | 2 |
| 2022 | Phylotranscriptomic patterns of network stochasticity and pathway dynamics during embryogenesisabstractMOTIVATION: The hourglass model is a popular evo-devo model depicting that the developmental constraints in the middle of a developmental process are higher, and hence the phenotypes are evolutionarily more conserved, than those that occur in early and late ontogeny stages. Although this model has been supported by studies analyzing developmental gene expression data, the evolutionary explanation and molecular mechanism behind this phenomenon are not fully understood yet. To approach this problem, Raff proposed a hypothesis and claimed that higher interconnectivity among elements in an organism during organogenesis resulted in the larger constraints at the mid-developmental stage. By employing stochastic network analysis and gene-set pathway analysis, we aim to demonstrate such changes of interconnectivity claimed in Raff's hypothesis. RESULTS: We first compared the changes of network randomness among developmental processes in different species by measuring the stochasticity within the biological network in each developmental stage. By tracking the network entropy along each developmental process, we found that the network stochasticity follows an anti-hourglass trajectory, and such a pattern supports Raff's hypothesis in dynamic changes of interconnections among biological modules during development. To understand which biological functions change during the transition of network stochasticity, we sketched out the pathway dynamics along the developmental stages and found that species may activate similar groups of biological processes across different stages. Moreover, higher interspecies correlations are found at the mid-developmental stages. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Kuei-Yueh Ko, Cho-Yi Chen, Hsueh-Fen Juan, Hsuan-Cheng Huang |
Bioinform. | 3 |
| 2017 | DynaPho: a web platform for inferring the dynamics of time-series phosphoproteomicsabstractSummary: Large-scale phosphoproteomics studies have improved our understanding of dynamic cellular signaling, but the downstream analysis of phosphoproteomics data is still a bottleneck. We develop DynaPho, a useful web-based tool providing comprehensive and in-depth analyses of time-course phosphoproteomics data, making analysis intuitive and accessible to non-bioinformatics experts. The tool currently implements five analytic modules, which reveal the transition of biological pathways, kinase activity, dynamics of interaction networks and the predicted kinase-substrate associations. These features can assist users in translating their larger-scale time-course phosphoproteomics data into valuable biological discoveries. Availability and implementation: DynaPho is freely available at http://dynapho.jhlab.tw/ . Contact: [email protected] or [email protected] . Supplementary information: Supplementary data are available at Bioinformatics online. Chia-Lang Hsu, Jiankai Wang, Pei-Chun Lu, Hsuan-Cheng Huang, Hsueh-Fen Juan |
Bioinform. | 5 |
| 2015 | Circadian systems biology in MetazoaabstractSystems biology, which can be defined as integrative biology, comprises multistage processes that can be used to understand components of complex biological systems of living organisms and provides hierarchical information to decoding life. Using systems biology approaches such as genomics, transcriptomics and proteomics, it is now possible to delineate more complicated interactions between circadian control systems and diseases. The circadian rhythm is a multiscale phenomenon existing within the body that influences numerous physiological activities such as changes in gene expression, protein turnover, metabolism and human behavior. In this review, we describe the relationships between the circadian control system and its related genes or proteins, and circadian rhythm disorders in systems biology studies. To maintain and modulate circadian oscillation, cells possess elaborative feedback loops composed of circadian core proteins that regulate the expression of other genes through their transcriptional activities. The disruption of these rhythms has been reported to be associated with diseases such as arrhythmia, obesity, insulin resistance, carcinogenesis and disruptions in natural oscillations in the control of cell growth. This review demonstrates that lifestyle is considered as a fundamental factor that modifies circadian rhythm, and the development of dysfunctions and diseases could be regulated by an underlying expression network with multiple circadian-associated signals. Li-Ling Lin, Hsuan-Cheng Huang, Hsueh-Fen Juan |
Briefings Bioinform. | 3 |
| 2014 | Mirin: identifying microRNA regulatory modules in protein-protein interaction networksabstractUNLABELLED: Exploring microRNA (miRNA) regulations and protein-protein interactions could reveal the molecular mechanisms responsible for complex biological processes. Mirin is a web-based application suitable for identifying functional modules from protein-protein interaction networks regulated by aberrant miRNAs under user-defined biological conditions such as cancers. The analysis involves combining miRNA regulations, protein-protein interactions between target genes, as well as mRNA and miRNA expression profiles provided by users. Mirin has successfully uncovered oncomirs and their regulatory networks in various cancers, such as gastric and breast cancer. AVAILABILITY AND IMPLEMENTATION: Mirin is freely available at http://mirin.ym.edu.tw/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Ken-Chi Yang, Chia-Lang Hsu, Chen-Ching Lin, Hsueh-Fen Juan, Hsuan-Cheng Huang |
Bioinform. | 4 |
| 2012 | Lengthening of 3′UTR increases with morphological complexity in animal evolutionabstractMOTIVATION: Evolutionary expansion of gene regulatory circuits seems to boost morphological complexity. However, the expansion patterns and the quantification relationships have not yet been identified. In this study, we focus on the regulatory circuits at the post-transcriptional level, investigating whether and how this principle may apply. RESULTS: By analysing the structure of mRNA transcripts in multiple metazoan species, we observed a striking exponential correlation between the length of 3' untranslated regions (3'UTR) and morphological complexity as measured by the number of cell types in each organism. Cellular diversity was similarly associated with the accumulation of microRNA genes and their putative targets. We propose that the lengthening of 3'UTRs together with a commensurate exponential expansion in post-transcriptional regulatory circuits can contribute to the emergence of new cell types during animal evolution. Cho-Yi Chen, Shui-Tein Chen, Hsueh-Fen Juan, Hsuan-Cheng Huang |
Bioinform. | 3 |
| 2011 | Coregulation of transcription factors and microRNAs in human transcriptional regulatory networkabstractBACKGROUND: MicroRNAs (miRNAs) are small RNA molecules that regulate gene expression at the post-transcriptional level. Recent studies have suggested that miRNAs and transcription factors are primary metazoan gene regulators; however, the crosstalk between them still remains unclear. METHODS: We proposed a novel model utilizing functional annotation information to identify significant coregulation between transcriptional and post-transcriptional layers. Based on this model, function-enriched coregulation relationships were discovered and combined into different kinds of functional coregulation networks. RESULTS: We found that miRNAs may engage in a wider diversity of biological processes by coordinating with transcription factors, and this kind of cross-layer coregulation may have higher specificity than intra-layer coregulation. In addition, the coregulation networks reveal several types of network motifs, including feed-forward loops and massive upstream crosstalk. Finally, the expression patterns of these coregulation pairs in normal and tumour tissues were analyzed. Different coregulation types show unique expression correlation trends. More importantly, the disruption of coregulation may be associated with cancers. CONCLUSION: Our findings elucidate the combinatorial and cooperative properties of transcription factors and miRNAs regulation, and we proposes that the coordinated regulation may play an important role in many biological processes. Cho-Yi Chen, Shui-Tein Chen, Chiou-Shann Fuh, Hsueh-Fen Juan, Hsuan-Cheng Huang |
BMC Bioinform. | 4 |
| 2008 | MeInfoText: associated gene methylation and cancer information from text miningabstractBACKGROUND: DNA methylation is an important epigenetic modification of the genome. Abnormal DNA methylation may result in silencing of tumor suppressor genes and is common in a variety of human cancer cells. As more epigenetics research is published electronically, it is desirable to extract relevant information from biological literature. To facilitate epigenetics research, we have developed a database called MeInfoText to provide gene methylation information from text mining. DESCRIPTION: MeInfoText presents comprehensive association information about gene methylation and cancer, the profile of gene methylation among human cancer types and the gene methylation profile of a specific cancer type, based on association mining from large amounts of literature. In addition, MeInfoText offers integrated protein-protein interaction and biological pathway information collected from the Internet. MeInfoText also provides pathway cluster information regarding to a set of genes which may contribute the development of cancer due to aberrant methylation. The extracted evidence with highlighted keywords and the gene names identified from each methylation-related abstract is also retrieved. The database is now available at http://mit.lifescience.ntu.edu.tw/. CONCLUSION: MeInfoText is a unique database that provides comprehensive gene methylation and cancer association information. It will complement existing DNA methylation information and will be useful in epigenetics research and the prevention of cancer. Yu-Ching Fang, Hsuan-Cheng Huang, Hsueh-Fen Juan |
BMC Bioinform. | 3 |
| 2004 | An Efficient Mechanism for Prediction of Protein-Ligand Interactions Based on Analysis of Protein Tertiary SubstructuresabstractAnalysis of protein-ligand interactions is a fundamental issue in drug design. As the detailed and accurate analysis of protein-ligand interactions involves calculation of binding free energy based on thermodynamics and even quantum mechanics, which is highly expensive in terms of computing time, conformational and structural analysis of proteins and ligands has been widely employed as a screening process in computer-aided drug design. In this paper, an efficient mechanism for identifying possible protein-ligand interactions based on analysis of protein tertiary substructures is proposed. In one experiment reported in this paper, the proposed prediction mechanism has been exploited to obtain some clues about a hypothesis that the biochemists have been speculating. The main distinction in the design of the prediction mechanism is the filtering process incorporated to expedite the analysis. The filtering process extracts the residues located in a cave of the protein tertiary structure for analysis and operates with O(nlogn) time complexity, where n is the number of residues in the protein. In comparison, the /spl alpha/hull algorithm, which is a widely used algorithm in computer graphics for identifying those instances that are on the contour of a 3-dimensional object, features O(n/sup 2/) time complexity. Experimental results show that the filtering process presented in this paper is able to speed up the analysis by a factor ranging from 3.11 to 9.79 times. Darby Tien-Hao Chang, Chien-Yu Chen 0001, Yen-Jen Oyang, Hsueh-Fen Juan, Hsuan-Cheng Huang |
BIBE | 4 |
| 2004 | Incremental generation of summarized clustering hierarchy for protein family analysisabstractMOTIVATION: Protein sequence clustering has been widely exploited to facilitate in-depth analysis of protein functions and families. For some applications of protein sequence clustering, it is highly desirable that a hierarchical structure, also referred to as dendrogram, which shows how proteins are clustered at various levels, is generated. However, as the sizes of contemporary protein databases continue to grow at rapid rates, it is of great interest to develop some summarization mechanisms so that the users can browse the dendrogram and/or search for the desired information more effectively. RESULTS: In this paper, the design of a novel incremental clustering algorithm aimed at generating summarized dendrograms for analysis of protein databases is described. The proposed incremental clustering algorithm employs a statistics-based model to summarize the distributions of the similarity scores among the proteins in the database and to control formation of clusters. Experimental results reveal that, due to the summarization mechanism incorporated, the proposed incremental clustering algorithm offers the users highly concise dendrograms for analysis of protein clusters with biological significance. Another distinction of the proposed algorithm is its incremental nature. As the sizes of the contemporary protein databases continue to grow at fast rates, due to the concern of efficiency, it is desirable that cluster analysis of a protein database can be carried out incrementally, when the protein database is updated. Experimental results with the Swiss-Prot protein database reveal that the time complexity for carrying out incremental clustering with k new proteins added into the database containing n proteins is O(n2betalogn), where beta congruent with 0.865, provided that k << n. AVAILABILITY: The Linux executable is available on the following supplementary page. Chien-Yu Chen 0001, Yen-Jen Oyang, Hsueh-Fen Juan |
Bioinform. | 3 |
| 2004 | GeneNetwork: an interactive tool for reconstruction of genetic networks using microarray dataabstractUNLABELLED: Inferring genetic network architecture from time series data generated from high-throughput experimental technologies, such as cDNA microarray, can help us to understand the system behavior of living organisms. We have developed an interactive tool, GeneNetwork, which provides four reverse engineering models and three data interpolation approaches to infer relationships between genes. GeneNetwork enables a user to readily reconstruct genetic networks based on microarray data without having intimate knowledge of the mathematical models. A simple graphical user interface enables rapid, intuitive mapping and analysis of the reconstructed network allowing biologists to explore gene relationships at the system level. AVAILABILITY: Download from http://genenetwork.sbl.bc.sinica.edu.tw/. SUPPLEMENTARY INFORMATION: Supplement documentation of algorithms for the four approaches is downloadable at the above location. Chia-Chin Wu, Hsuan-Cheng Huang, Hsueh-Fen Juan, Shui-Tein Chen |
Bioinform. | 3 |