EDBT 2026 Demo / reviewers in the wild / expert
Honglin Wang
dblp:157/2794
· DBLP profile ↗
13ranked-venue papers
2as first author
10since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 8 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Fuzzy symbolic convergent cross mapping: A causal coupling measure for EEG signals in disorders of consciousness patients
Xingwei An, Yang Di, Honglin Wang, Yujia Yan, Shuang Liu 0004, Yueqing Dong, Dong Ming |
Neural Networks | 4 |
| 2024 | FuncNet: A Machine Learning Framework Capable of Enhancing Neural Cell Type Classification via Functional Subtype ClusteringabstractCell type identification in single cell RNA sequencing (scRNA-seq) remains an essential yet challenging task. Existing algorithms assign cell type nomenclature typically by (i) subgrouping cells by computing proximity of cellular gene expression patterns at the single cell level, (ii) identifying marker genes in each cluster, and (iii) performing a lookup of these genes in the literature and databases to assign cell subtype labels. Unfortunately, this approach falls short when classifying cells based on functions, because biological function(s) of a group of cells cannot be determined based on a limited set of marker genes. Here we propose a new modeling formalism, namely FuncNet, which is a form of neural network-based method capable of discerning the functional profiles of individual cells. The key idea of FuncNet is to incorporate "a secondary learning arm" into the otherwise classical multi-layer neural network in order to bring in "prior knowledge" into the overall learning process. The prior knowledge that is brought into the FuncNet learning is Gene Ontology Biological Processes that should have been captured in the cells being studied. We applied FuncNet to three publicly available single cell transcriptomics datasets from human and mouse brains. The results demonstrate that FuncNet can successfully reproduce the functional nomenclatures associated with brain cells in the literature. FuncNet suggests a new way of subtyping cells, specifically for the functional subtyping of single cells—an area which has not been well established in the current single cell data analysis practices. Chenyu Zhang 0007, Honglin Wang, Seung-Hyun Hong, Riqiang Yan, Dong-Guk Shin |
BIBM | 2 |
| 2024 | vSPACE: exploring virtual spatial representation of articular chondrocytes at the single-cell levelabstractSUMMARY: vSPACE is a web-based application presenting a spatial representation of scRNAseq data obtained from human articular cartilage by emulating the concept of spatial transcriptomics technology, but virtually. This virtual 2D plot presentation of human articular cartage cells generates several zonal distribution patterns, for one or multiple genes at a time, revealing patterns that scientists can appreciate as imputed spatial distribution patterns along the zonal axis. AVAILABILITY AND IMPLEMENTATION: vSPACE is implemented in Python Dash as a web-based toolbox designed for data visualization of zonal gene expression patterns in articular cartilage chondrocytes. This tool is freely accessible at: https://vspace.cse.uconn.edu/The source code and extra materials for this service can be downloaded from: https://github.com/zhacheny/vSPACE. Chenyu Zhang 0007, Honglin Wang, Yeonsoo Chung, Seung-Hyun Hong, Merissa Olmer, Hannah Swahn, Martin Lotz, Peter F. Maye, David W. Rowe, Dong-Guk Shin |
Bioinform. | 2 |
| 2024 | A survey on intelligent management of alerts and incidents in IT services
Qingyang Yu, Nengwen Zhao, Mingjie Li 0005, Zeyan Li 0001, Honglin Wang, Wenchi Zhang, Kaixin Sui, Dan Pei |
J. Netw. Comput. Appl. | 5 |
| 2023 | Using Biological Processes as Prior Knowledge Identifies New Microglial Immune Signatures at Single Cell Level in Alzheimer's DiseaseabstractThe brain’s resident immune cells, microglia, play important roles in the pathological process of Alzheimer's disease (AD). Upon encountering amyloid-β (Aβ) plaque accumulation, microglia change its state to produce proinflammatory cytokines to uptake and clear Aβ but during the chronic state of neuroinflammation they become responsible for neurodegeneration. In this paper we propose a new method of analyzing microglia undergoing multi-stage state changes from homeostasis (HOM) to disease associated microglia (DAM). Our single cell gene expression data analysis method differs from the conventional marker based subtyping methods in the sense that it uses prior knowledge, namely, Gene Ontology’s Biological Processes known for immune functions and other related ones. In this regard, our method can be thought of supervised clustering rather than UMAP/tSNE style unbiased subgrouping of cells. The strengths of this "prior knowledge" using method are multi-faceted: (i) one can assign meaningful functional nomenclatures to identified subtypes of cells, (ii) one can refine GO Biological Process genes to be context-specific (e.g., phagocytosis responsible by "microglia" as opposed to "general immune cells"), and (iii) additional context-specific biological processes can be identified in addition to the ones used as "prior knowledge". We illustrate the advantages of our semi-supervised cell clustering method using a set of publicly available human AD and mouse AD model gene expression datasets. Chenyu Zhang 0007, Honglin Wang, Seung-Hyun Hong, Riqiang Yan, Dong-Guk Shin |
BIBM | 2 |
| 2022 | Pola Viz Reveals Microglia Polarization at Single Cell Level in Alzheimer's DiseaseabstractMicroglia are macrophages residing in the brain and spinal cord and responsible for neurological immune defense. A growing number of studies point out the importance of microglia’s role in Alzheimer’s disease (AD) pathology. Like macrophages in other tissues, microglia are polarized to M1/M2 states, Ml inhibiting cell inflammatory and causing tissue damage and M2 resolving inflammation and repairing damaged tissue. Thanks to scRNA-seq technology, single cell gene expression data sets for microglia have become available. However, examining the polarization of microglia at single cell level to study their association with AD is difficult for reasons such as complexity (too many genes, too many cells), expression dropouts, etc. We propose a scRNA-seq data visualization framework, namely, Pola Viz that can associate the polarization states of microglia with gene regulatory pathways that are potentially responsible for defining M1/M2 states or transition between them. For its two-dimensional placement of single cells, the x-axis depicts each cell’s M1/M2 or its transition score and the y-axis denotes what we call, the pathway “route” scores designed to assign the degree of contribution of the concerned pathway routes to each cell’s M1/M2 state. We conducted two case studies involving publicly available AD single cell mRNA datasets, one from human brain and one from mouse brain. Our method identified pathway routes that are highly correlated with known functional roles of microglia. It also discovers similarities between mice and human AD microglia activation such as T cell receptor signaling pathway, cell cycle, NF-kappa B signaling pathways, all highly disturbed at late stage of AD. Our method provides a new way of examining the M1/M2 polarization of microglia at single cell level Chenyu Zhang 0007, Pujan Joshi, Honglin Wang, Seung-Hyun Hong, Riqiang Yan, Dong-Guk Shin |
BIBM | 3 |
| 2021 | Identification of Crosstalk between Biological Pathway Routes in Cancer CohortsabstractSignal transduction pathways can affect each other through crosstalk between one or more shared components. In this paper, we present a framework that uses a route-based approach to identify crosstalk between signaling pathway routes in different cancer cohorts. The framework identifies all possible routes originating from a ligand/receptor (LR) of one pathway to a transcription factor (TF) of another pathway. Then, the downstream crosstalk routes, which originate from the TF of one pathway to ligand of another pathway, are identified. For each crosstalk route, activity scores and p-values are computed. Overall route activity in each cancer cohort was assessed in terms of two summary metrics, “Proportion of Significance” (PS) and “Average Route Score” (ARS). Case studies of five human cancer cohorts from The Cancer Genome Atlas (TCGA) repository are presented to demonstrate the discovery value of this approach. Pujan Joshi, Honglin Wang, Salvatore Jaramillo, Seung-Hyun Hong, Charles Giardina, Dong-Guk Shin |
BIBM | 2 |
| 2021 | ctBuilder: A framework for building pathway crosstalks by combining single cell data with bulk cell dataabstractBiomedical research scientists routinely use publicly available gene regulatory pathway databases to extract interpretation from high-throughput multi-omics studies. However, deficiencies of these existing pathway resources include: (i) the databases store “individualized” pathways as a collection of networks thus making discovery of crosstalk across the boundaries of curated pathways difficult, and (ii) curated pathways themselves are incomplete, particularly from the disease perspective. We propose a computational pathway extension framework, called ctBuilder, which aims to tackle the deficiencies. It uses single cell gene expression data to identify potential candidates to tune the pathway extension “context-specific” and uses bulk cell gene expression data sets to further refine and improve likelihoods of the regulatory relationships estimated to interlink constituents involved between pathways. For demonstration, we applied ctBuilder to 5 non-alcoholic steatohepatis related studies from GEO. We show how a pair of known pathways, TNF signaling vs. NF-kB signaling, can be extended to include crosstalk subnetworks specifically related to the liver disease. Discovered relationships among them merit wetlab experiments for validation. Honglin Wang, Pujan Joshi, Seung-Hyun Hong, Dong-Ju Shin, Dong-Guk Shin |
BIBM | 1 |
| 2021 | Identifying bad software changes via multimodal anomaly detection for online service systemsabstractIn large-scale online service systems, software changes are inevitable and frequent. Due to importing new code or configurations, changes are likely to incur incidents and destroy user experience. Thus it is essential for engineers to identify bad software changes, so as to reduce the influence of incidents and improve system re- liability. To better understand bad software changes, we perform the first empirical study based on large-scale real-world data from a large commercial bank. Our quantitative analyses indicate that about 50.4% of incidents are caused by bad changes, mainly be- cause of code defect, configuration error, resource contention, and software version. Besides, our qualitative analyses show that the current practice of detecting bad software changes performs not well to handle heterogeneous multi-source data involved in soft- ware changes. Based on the findings and motivation obtained from the empirical study, we propose a novel approach named SCWarn aiming to identify bad changes and produce interpretable alerts accurately and timely. The key idea of SCWarn is drawing support from multimodal learning to identify anomalies from heterogeneous multi-source data. An extensive study on two datasets with various bad software changes demonstrates our approach significantly outperforms all the compared approaches, achieving 0.95 F1-score on average and reducing MTTD (mean time to detect) by 20.4%∼60.7%. In particular, we shared some success stories and lessons learned from the practical usage. Nengwen Zhao, Junjie Chen 0003, Zhaoyang Yu 0002, Honglin Wang, Jiesong Li, Bin Qiu, Hongyu Xu, Wenchi Zhang, Kaixin Sui, Dan Pei |
ESEC/SIGSOFT FSE | 4 |
| 2021 | An empirical investigation of practical log anomaly detection for online service systemsabstractLog data is an essential and valuable resource of online service systems, which records detailed information of system running status and user behavior. Log anomaly detection is vital for service reliability engineering, which has been extensively studied. However, we find that existing approaches suffer from several limitations when deploying them into practice, including 1) inability to deal with various logs and complex log abnormal patterns; 2) poor interpretability; 3) lack of domain knowledge. To help understand these practical challenges and investigate the practical performance of existing work quantitatively, we conduct the first empirical study and an experimental study based on large-scale real-world data. We find that logs with rich information indeed exhibit diverse abnormal patterns (e.g., keywords, template count, template sequence, variable value, and variable distribution). However, existing approaches fail to tackle such complex abnormal patterns, producing unsatisfactory performance. Motivated by obtained findings, we propose a generic log anomaly detection system named LogAD based on ensemble learning, which integrates multiple anomaly detection approaches and domain knowledge, so as to handle complex situations in practice. About the effectiveness of LogAD, the average F1-score achieves 0.83, outperforming all baselines. Besides, we also share some success cases and lessons learned during our study. To our best knowledge, we are the first to investigate practical log anomaly detection in the real world deeply. Our work is helpful for practitioners and researchers to apply log anomaly detection to practice to enhance service reliability. Nengwen Zhao, Honglin Wang, Zeyan Li 0001, Zhu Pan, Xidao Wen, Wenchi Zhang, Kaixin Sui, Dan Pei |
ESEC/SIGSOFT FSE | 2 |
| 2020 | Identification of Key Biological Pathway Routes in Cancer CohortsabstractOver the last two decades, various pathway analysis methods have been proposed to investigate complex biological interactions with omics data. Topology based (TB) pathway analysis techniques are generally considered to have better performance than non-topology-based methods. However, these methods score an entire pathway as a unit, where the relevance of individual routes within the pathway is lost. In this paper, a novel route-based pathway analysis framework is discussed that can effectively process entire cohorts of gene expression data sets and identify significant pathway routes in the given cohort. The framework begins with identifying all possible transcription factor (TF) centric routes from KEGG signaling pathways. For each route in a pathway, activity scores and p-values are calculated for samples in the given cohort. Overall route activity in a cohort is assessed in terms of two summary metrics, “Proportion of Significance” (PS) and “Average Route Score” (ARS). Case studies of two human cancer cohorts from The Cancer Genome Atlas (TCGA) repository are presented. Pujan Joshi, Brent Basso, Honglin Wang, Seung-Hyun Hong, Charles Giardina, Dong-Guk Shin |
BIBM | 3 |
| 2020 | cTAP: A Machine Learning Framework for Predicting Target Genes of a Transcription Factor using a Cohort of Gene Expression Data SetsabstractIdentifying target genes of a transcription factor is crucial in biomedical research. Thanks to ChIP-seq technology, scientists can estimate potential genome-wide target genes of a transcription factor. However, finding the consistently behaving Up/Down targets of a transcription factor in a given biological context is difficult because it requires analysis of a large number of studies under the same or comparable context. We present a transcription target prediction method, called Cohort-based TF target prediction system (cTAP). This method assumes that the pathway involving the transcription factor of interest is featured with multiple functional groups of marker genes pertaining to the concerned biological process. It uses the notion of gene-presence and gene-absence in addition to log2 ratios of gene expression values for the prediction. Target prediction is made by applying multiple machine-learning models that learn the patterns of genepresence and gene-absence from log2 ratio and four types of Z scores from the normalized cohort's gene expression data. The learned patterns are then associated with the putative targets of the concerned transcription factor to elicit genes exhibiting Up/Down gene regulation patterns “consistently” within the cohort. Totally 11 publicly available GEO data sets related to osteoclastogenesis are used in our experiment. The learned models using gene-presence and gene-absence produce target genes different from using only log2 ratios such as CASP1, BID, and IRF5. Our literature survey reveals that all these predicted targets have known roles in bone remodeling, specifically related to immune and osteoclasts, suggesting confidence in our method and potential merit for a wet-lab experiment for validation. Honglin Wang, Pujan Joshi, Seung-Hyun Hong, Peter F. Maye, David W. Rowe, Dong-Guk Shin |
BIBM | 1 |
| 2008 | Dependency Tree-based SRL with Proper Pruning and Extensive Feature Engineering
Hongling Wang, Honglin Wang, Guodong Zhou 0001, Qiaoming Zhu |
CoNLL | 2 |