EDBT 2026 Demo / reviewers in the wild / expert
Zhiming Dai
dblp:79/993
· DBLP profile ↗
20ranked-venue papers
5as first author
12since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 15 · 4 first-author · 8 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Informal Embodied Auditing: Exploring Facial Emotion AI (FEAI) through Community WorkshopsabstractEmotion AI (EAI) is increasingly deployed and ethically controversial—motivating a need for greater public understanding, critique, and ethical discussions. Facial Emotion AI (FEAI) is a common type of EAI that infers emotions from facial expressions. We developed Explore-FEAI, an FEAI model and accompanying interactive website that offers open-ended exploration with FEAI firsthand. We designed a workshop wherein participants learn about FEAI using Explore-FEAI and discuss societal implications, partnering with local organizations to host community workshops (N=30). Our findings analyze participants’ growing critical AI literacy through exploring inputs/outputs, mechanistic reasoning, data critiques, sociocultural critiques, ethical concerns, and embodied and material exploration of FEAI. Our discussion offers informal embodied auditing as an approach for critical engagement with AI through embodied and material exploration, as well as reflections on informal auditing for supporting AI literacy, informal auditing for questioning EAI ethics, and expanding participation roles for more holistic EAI training. Alexandra Teixeira Riggs, Zhiming Dai, Crystal Byrd Farmer, Kalia G. Morrison, Noura Howell |
CHI | 3 |
| 2026 | PLNK: Prompt Learning With Neutral Knowledge for Few-Shot Out-of-Distribution DetectionabstractRecent developments in few-shot out-of-distribution (OOD) detection have yielded remarkable performance, benefiting from large pre-trained vision-language models (VLMs). Our prior work focuses on using in-distribution (ID) knowledge as references to learn richer knowledge beyond the textual semantics of class labels, which is prone to cause the model overconfidence and results in a limited score gap between ID and OOD data. In this paper, rather than treating ID knowledge as references, we propose Prompt Learning with Neutral Knowledge (PLNK) to better differentiate ID from OOD data. Our key insight lies in leveraging diverse neutral knowledge to improve ID discrimination while alleviating the inherent model overconfidence on OOD data induced by ID knowledge, thereby capturing the notable discrepancy between ID and OOD data. By introducing neutral knowledge with a balanced degree of similarity to both ID and OOD data, we amplify the discrepancy between the learnable prompt and the references (i.e., diverse neutral knowledge) for ID data, while reducing it for OOD data. In this way, the simple yet effective PLNK framework brings a notable score gap between ID and OOD data, thereby improving OOD detection. Moreover, we incorporate the visual neutral prompt with richer semantics alongside the original text-only reference. Comprehensive experiments show that our method consistently surpasses current state-of-the-art methods. The codes will be released publicly. Xinhua Lu, Runhe Lai, Yanqi Wu, Kanghao Chen, Zhiming Dai, Wei-Shi Zheng 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2025 | Augmenting Continual Learning of Diseases with LLM-Generated Visual ConceptsabstractContinual learning is essential for medical image classification systems to adapt to dynamically evolving clinical environments. The integration of multimodal information can significantly enhance continual learning of image classes. However, while existing approaches do utilize textual modality information, they solely rely on simplistic templates with a class name, thereby neglecting richer semantic information. To address these limitations, we propose a novel framework that harnesses visual concepts generated by large language models (LLMs) as discriminative semantic guidance. Our method dynamically constructs a visual concept pool with a similarity-based filtering mechanism to prevent redundancy. Then, to integrate the concepts into the continual learning process, we employ a crossmodal image-concept attention module, coupled with an attention loss. Through attention, the module can leverage the semantic knowledge from relevant visual concepts and produce classrepresentative fused features for classification. Experiments on medical and natural image datasets show our method achieves state-of-the-art performance. Jiantao Tan, Peixian Ma, Zhiming Dai, Kanghao Chen |
BIBM | 3 |
| 2025 | ACID: Improving Ambiguous Cell Type Annotation in Spatial TranscriptomicsabstractCurrent spatial transcriptomics(ST) technologies have reached cellular or even subcellular resolution, providing deeper insights into single cells and histological analysis. Cell type annotation serves as the cornerstone for downstream analysis of ST data. Most methods are committed to integrate both spatial location relationships in ST and scRNA-seq data to improve the accuracy of cell type annotation. However, some cells are annotated with low confidence, and these cells are referred to as ambiguous cells. This paper develops a new method, ACID, based on our proposed refined annotation module, to improve ambiguous cells type annotation in ST data. Extensive experiments on datasets from different sequencing technologies demonstrate that the number of ambiguous cells can be effectively reduced by ACID. ACID outperforms state-of-the-art methods on cell type annotation. In addition, experiments on various ST annotation methods show that our developed refined annotation module can improve cell type annotation when the module is integrated into these methods. Zhiming Dai |
BIBM | 3 |
| 2024 | Dynamic Knowledge Prompt for Chest X-ray Report GenerationabstractAutomatic generation of radiology reports can relieve the burden of radiologist. In the radiology library, the biased dataset and the sparse features of chest X-ray image make it difficult to generate reports. Many approaches strive to integrate prior information to enhance generation, but they fail to dynamically utilize pulmonary lesion knowledge at the instance-level. To alleviate above problem, we propose a novel Dynamic Knowledge Prompt (DKP) framework for chest X-ray report generation. The DKP can dynamically incorporate the pulmonary lesion information at the instance-level to facilitate report generation. Initially, we design a knowledge prompt for each pulmonary lesion using numerous radiology reports. After that, the DKP using an anomaly detector generates the dynamic knowledge prompt by extracting discriminative lesion features in the corresponding X-ray image. Finally, the knowledge prompt is encoded and fused with hidden states extracted from decoder, to form multi-modal features that guide visual features to generate reports. Extensive experiments on the public datasets MIMIC-CXR and IU X-Ray show that our approach achieves state-of-the-art performance. Shenshen Bu, Taiji Li, Zhiming Dai |
LREC/COLING | 4 |
| 2024 | Instance-level Expert Knowledge and Aggregate Discriminative Attention for Radiology Report GenerationabstractAutomatic radiology report generation can provide sub-stantial advantages to clinical physicians by effectively re-ducing their workload and improving efficiency. Despite the promising potential of current methods, challenges persist in effectively extracting and preventing degradation of prominent features, as well as enhancing attention on piv-otal regions. In this paper, we propose an Instance-level Expert Knowledge and Aggregate Discriminative Attention framework (EKAGen11https://github.com/hnjzbss/EKAGen) for radiology report generation. We convert expert reports into an embedding space and gener-ate comprehensive representations for each disease, which serve as Preliminary Knowledge Support (PKS). To prevent feature disruption, we select the representations in the em-bedding space with the smallest distances to P KS as Rec-tified Knowledge Support (RKS). Then, EKAGen diagnoses the diseases and retrieves knowledge from RKS, creating Instance-level Expert Knowledge (IEK) for each query image, boosting generation. Additionally, we introduce Ag-gregate Discriminative Attention Map (ADM), which uses weak supervision to create maps of discriminative regions that highlight pivotal regions. For training, we propose a Global Information Self-Distillation (GID) strategy, using an iteratively optimized model to distill global knowledge into EKAGen. Extensive experiments and analyses on IU X-Ray and MIMIC-CXR datasets demonstrate that EKAGen outperforms previous state-of-the-art methods. Shenshen Bu, Taiji Li, Yuedong Yang, Zhiming Dai |
CVPR | 4 |
| 2024 | Comprehensive single-cell RNA-seq analysis using deep interpretable generative modeling guided by biological hierarchy knowledgeabstractRecent advances in microfluidics and sequencing technologies allow researchers to explore cellular heterogeneity at single-cell resolution. In recent years, deep learning frameworks, such as generative models, have brought great changes to the analysis of transcriptomic data. Nevertheless, relying on the potential space of these generative models alone is insufficient to generate biological explanations. In addition, most of the previous work based on generative models is limited to shallow neural networks with one to three layers of latent variables, which may limit the capabilities of the models. Here, we propose a deep interpretable generative model called d-scIGM for single-cell data analysis. d-scIGM combines sawtooth connectivity techniques and residual networks, thereby constructing a deep generative framework. In addition, d-scIGM incorporates hierarchical prior knowledge of biological domains to enhance the interpretability of the model. We show that d-scIGM achieves excellent performance in a variety of fundamental tasks, including clustering, visualization, and pseudo-temporal inference. Through topic pathway studies, we found that d-scIGM-learned topics are better enriched for biologically meaningful pathways compared to the baseline models. Furthermore, the analysis of drug response data shows that d-scIGM can capture drug response patterns in large-scale experiments, which provides a promising way to elucidate the underlying biological mechanisms. Lastly, in the melanoma dataset, d-scIGM accurately identified different cell types and revealed multiple melanin-related driver genes and key pathways, which are critical for understanding disease mechanisms and drug development. Hegang Chen, Yuyin Lu, Zhiming Dai, Yuedong Yang, Qing Li 0001, Yanghui Rao |
Briefings Bioinform. | 3 |
| 2024 | Detecting novel cell type in single-cell chromatin accessibility data via open-set domain adaptationabstractRecent advances in single-cell technologies enable the rapid growth of multi-omics data. Cell type annotation is one common task in analyzing single-cell data. It is a challenge that some cell types in the testing set are not present in the training set (i.e. unknown cell types). Most scATAC-seq cell type annotation methods generally assign each cell in the testing set to one known type in the training set but neglect unknown cell types. Here, we present OVAAnno, an automatic cell types annotation method which utilizes open-set domain adaptation to detect unknown cell types in scATAC-seq data. Comprehensive experiments show that OVAAnno successfully identifies known and unknown cell types. Further experiments demonstrate that OVAAnno also performs well on scRNA-seq data. Our codes are available online at https://github.com/lisaber/OVAAnno/tree/master. Yuefan Lin, Zixiang Pan, Yuansong Zeng, Yuedong Yang, Zhiming Dai |
Briefings Bioinform. | 5 |
| 2024 | Subgraph extraction and graph representation learning for single cell Hi-C imputation and clusteringabstractSingle-cell Hi-C (scHi-C) technology enables the investigation of 3D chromatin structure variability across individual cells. However, the analysis of scHi-C data is challenged by a large number of missing values. Here, we present a scHi-C data imputation model HiC-SGL, based on Subgraph extraction and graph representation learning. HiC-SGL can also learn informative low-dimensional embeddings of cells. We demonstrate that our method surpasses existing methods in terms of imputation accuracy and clustering performance by various metrics. Jiahao Zheng 0005, Yuedong Yang, Zhiming Dai |
Briefings Bioinform. | 3 |
| 2023 | Enhancing Medical Report Generation in Multi-Slice Fusion ScenariosabstractAutomatic medical report generation has garnered considerable research attention owing to its practical significance in alleviating the workload burden on radiologists. Despite the promise shown by existing methods, limitations still persist due to their tendency to overlook the multi-view or high-resolution characteristics of medical images, which contain a wealth of rich and diverse information. To tackle this issue, we propose a Multi-Slice Fusion report generation framework (ReFuGen). In detail, ReFuGen utilizes a Multi-slice Feature Extractor to extract features from multiple image slices. Then, we introduce the Exemplary Slice Amplifier, which adaptively identifies important slices, applies weighting based on their contribution, and generates enhanced features by integrating spatial features from all slices. Next, the Adaptive Receptive Field Seq-Enhancer is proposed to enhance feature interactions across different channels by adaptively using a series of 1D convolutions. Lastly, we introduce the Text Knowledge Integrator, which integrates textual knowledge to address the issue of sparse features in medical images. Experimentally, we outperform state-of-the-art methods on three widely used benchmark datasets: IU X-Ray, MIMIC-CXR and FFA-IR. Shenshen Bu, Taiji Li, Zhiming Dai |
BIBM | 3 |
| 2023 | CSIFNet: Deep multimodal network based on cellular spatial information fusion for survival predictionabstractUsing multimodal data to construct models for cancer survival prediction is essential for the prognosis of patients. Most of existing methods do not consider the cellular spatial organization of tumors at the single-cell level. In this work, we present CSIFNet, a deep multimodal network based on cellular spatial information fusion. The multimodal data used includes genomics data, clinical variables data and cellular spatial data. Specifically, we construct a spatial information fusion module (SIFM) to learn the local interactions between tumor cells and the tumor microenvironment to obtain a more comprehensive high-level representation of tumor cellular spatial information. Experimental results show that our proposed method has better performance than the state-of-the-art survival prediction methods. The source code is available at https://github.com/heyhola/CSIFNet. Siqi Xiao, Zhiming Dai |
BIBM | 2 |
| 2023 | Benchmarking deep learning methods for predicting CRISPR/Cas9 sgRNA on- and off-target activitiesabstractIn silico design of single guide RNA (sgRNA) plays a critical role in clustered regularly interspaced, short palindromic repeats/CRISPR-associated protein 9 (CRISPR/Cas9) system. Continuous efforts are aimed at improving sgRNA design with efficient on-target activity and reduced off-target mutations. In the last 5 years, an increasing number of deep learning-based methods have achieved breakthrough performance in predicting sgRNA on- and off-target activities. Nevertheless, it is worthwhile to systematically evaluate these methods for their predictive abilities. In this review, we conducted a systematic survey on the progress in prediction of on- and off-target editing. We investigated the performances of 10 mainstream deep learning-based on-target predictors using nine public datasets with different sample sizes. We found that in most scenarios, these methods showed superior predictive power on large- and medium-scale datasets than on small-scale datasets. In addition, we performed unbiased experiments to provide in-depth comparison of eight representative approaches for off-target prediction on 12 publicly available datasets with various imbalanced ratios of positive/negative samples. Most methods showed excellent performance on balanced datasets but have much room for improvement on moderate- and severe-imbalanced datasets. This study provides comprehensive perspectives on CRISPR/Cas9 sgRNA on- and off-target activity prediction and improvement for method development. Guishan Zhang, Xianhua Dai, Zhiming Dai |
Briefings Bioinform. | 4 |
| 2020 | SANPolyA: a deep learning method for identifying Poly(A) signalsabstractMOTIVATION: Polyadenylation plays a regulatory role in transcription. The recognition of polyadenylation signal (PAS) motif sequence is an important step in polyadenylation. In the past few years, some statistical machine learning-based and deep learning-based methods have been proposed for PAS identification. Although these methods predict PAS with success, there is room for their improvement on PAS identification. RESULTS: In this study, we proposed a deep neural network-based computational method, called SANPolyA, for identifying PAS in human and mouse genomes. SANPolyA requires no manually crafted sequence features. We compared our method SANPolyA with several previous PAS identification methods on several PAS benchmark datasets. Our results showed that SANPolyA outperforms the state-of-art methods. SANPolyA also showed good performance on leave-one-motif-out evaluation. AVAILABILITY AND IMPLEMENTATION: https://github.com/yuht4/SANPolyA. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Zhiming Dai |
Bioinform. | 2 |
| 2017 | pgRNAFinder: a web-based tool to design distance independent paired-gRNAabstractSUMMARY: The CRISPR/Cas System has been shown to be an efficient and accurate genome-editing technique. There exist a number of tools to design the guide RNA sequences and predict potential off-target sites. However, most of the existing computational tools on gRNA design are restricted to small deletions. To address this issue, we present pgRNAFinder, with an easy-to-use web interface, which enables researchers to design single or distance-free paired-gRNA sequences. The web interface of pgRNAFinder contains both gRNA search and scoring system. After users input query sequences, it searches gRNA by 3' protospacer-adjacent motif (PAM), and possible off-targets, and scores the conservation of the deleted sequences rapidly. Filters can be applied to identify high-quality CRISPR sites. PgRNAFinder offers gRNA design functionality for 8 vertebrate genomes. Furthermore, to keep pgRNAFinder open, extensible to any organism, we provide the source package for local use. AVAILABILITY AND IMPLEMENTATION: The pgRNAFinder is freely available at http://songyanglab.sysu.edu.cn/wangwebs/pgRNAFinder/, and the source code and user manual can be obtained from https://github.com/xiexiaowei/pgRNAFinder. CONTACT: [email protected] or [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Yuanyan Xiong, Xiaowei Xie, Wenbin Ma, Puping Liang, Songyang Zhou, Zhiming Dai |
Bioinform. | 7 |
| 2012 | Antisense transcription is coupled to nucleosome occupancy in sense promotersabstractMOTIVATION: Genome-wide pervasive transcription is widespread in eukaryotes, revealing an extensive array of antisense transcription that involves hundreds of previously unknown non-coding RNAs. Individual cases have shown that antisense transcription influences sense transcription, however, genome-wide mechanisms of how antisense transcription regulates sense transcription remain to be elucidated. RESULTS: Here, we performed a systematic analysis of sense-antisense transcription and nucleosome occupancy in yeast. We found that antisense transcription is associated with nucleosome occupancy in sense promoters. Using RNA polymerase II inactivation data as a reasonable approximation to antisense transcription inactivation data, we further showed that antisense transcripts increase nucleosome occupancy in sense promoter regions they overlap, and reduce nucleosome occupancy in sense promoter regions around their transcription termination sites. These results reveal the previously unappreciated roles of antisense transcription in directing nucleosome occupancy in sense promoters. Our findings will have implications in understanding regulatory functions of antisense transcription. Zhiming Dai, Xianhua Dai |
Bioinform. | 1 |
| 2011 | Genome-wide DNA sequence polymorphisms facilitate nucleosome positioning in yeastabstractMOTIVATION: The intrinsic DNA sequence is an important determinant of nucleosome positioning. Some DNA sequence patterns can facilitate nucleosome formation, while others can inhibit nucleosome formation. Nucleosome positioning influences the overall rate of sequence evolution. However, its impacts on specific patterns of sequence evolution are still poorly understood. RESULTS: Here, we examined whether nucleosomal DNA and nucleosome-depleted DNA show distinct polymorphism patterns to maintain adequate nucleosome architecture on a genome scale in yeast. We found that sequence polymorphisms in nucleosomal DNA tend to facilitate nucleosome formation, whereas polymorphisms in nucleosome-depleted DNA tend to inhibit nucleosome formation, which is especially evident at nucleosome-disfavored sequences in nucleosome-free regions at both ends of genes. Sequence polymorphisms facilitating nucleosome positioning correspond to stable nucleosome positioning. These results reveal that sequence polymorphisms are under selective constraints to maintain nucleosome positioning. CONTACT: [email protected]; [email protected] Zhiming Dai, Xianhua Dai |
Bioinform. | 1 |
| 2011 | Gene Expression Divergence is Coupled to Evolution of DNA Structure in Coding RegionsabstractSequence changes in coding region and regulatory region of the gene itself (cis) determine most of gene expression divergence between closely related species. But gene expression divergence between yeast species is not correlated with evolution of primary nucleotide sequence. This indicates that other factors in cis direct gene expression divergence. Here, we studied the contribution of DNA three-dimensional structural evolution as cis to gene expression divergence. We found that the evolution of DNA structure in coding regions and gene expression divergence are correlated in yeast. Similar result was also observed between Drosophila species. DNA structure is associated with the binding of chromatin remodelers and histone modifiers to DNA sequences in coding regions, which influence RNA polymerase II occupancy that controls gene expression level. We also found that genes with similar DNA structures are involved in the same biological process and function. These results reveal the previously unappreciated roles of DNA structure as cis-effects in gene expression. Zhiming Dai, Xianhua Dai |
PLoS Comput. Biol. | 1 |
| 2009 | Transcriptional interaction-assisted identification of dynamic nucleosome positioningabstractBACKGROUND: Nucleosomes regulate DNA accessibility and therefore play a central role in transcription control. Computational methods have been developed to predict static nucleosome positions from DNA sequences, but nucleosomes are dynamic in vivo. RESULTS: Motivated by our observation that transcriptional interaction is discriminative information for nucleosome occupancy, we developed a novel computational approach to identify dynamic nucleosome positions at promoters by combining transcriptional interaction and genomic sequence information. Our approach successfully identified experimentally determined nucleosome positioning dynamics available in three cellular conditions, and significantly improved the prediction accuracy which is based on sequence information alone. We then applied our approach to various cellular conditions and established a comprehensive landscape of dynamic nucleosome positioning in yeast. CONCLUSION: Analysis of this landscape revealed that the majority of nucleosome positions are maintained during most conditions. However, nucleosome occupancy at most promoters fluctuates with the corresponding gene expression level and is reduced specifically at the phase of peak expression. Further investigation into properties of nucleosome occupancy identified two gene groups associated with distinct modes of nucleosome modulation. Our results suggest that both the intrinsic sequence and regulatory proteins modulate nucleosomes in an altered manner. Zhiming Dai, Xianhua Dai, Jihua Feng, Yangyang Deng, Jiang Wang 0014, Caisheng He |
BMC Bioinform. | 1 |
| 2008 | Missing value imputation for microarray gene expression data using histone acetylation informationabstractBACKGROUND: It is an important pre-processing step to accurately estimate missing values in microarray data, because complete datasets are required in numerous expression profile analysis in bioinformatics. Although several methods have been suggested, their performances are not satisfactory for datasets with high missing percentages. RESULTS: The paper explores the feasibility of doing missing value imputation with the help of gene regulatory mechanism. An imputation framework called histone acetylation information aided imputation method (HAIimpute method) is presented. It incorporates the histone acetylation information into the conventional KNN(k-nearest neighbor) and LLS(local least square) imputation algorithms for final prediction of the missing values. The experimental results indicated that the use of acetylation information can provide significant improvements in microarray imputation accuracy. The HAIimpute methods consistently improve the widely used methods such as KNN and LLS in terms of normalized root mean squared error (NRMSE). Meanwhile, the genes imputed by HAIimpute methods are more correlated with the original complete genes in terms of Pearson correlation coefficients. Furthermore, the proposed methods also outperform GOimpute, which is one of the existing related methods that use the functional similarity as the external information. CONCLUSION: We demonstrated that the using of histone acetylation information could greatly improve the performance of the imputation especially at high missing percentages. This idea can be generalized to various imputation methods to facilitate the performance. Moreover, with more knowledge accumulated on gene regulatory mechanism in addition to histone acetylation, the performance of our approach can be further improved and verified. Xianhua Dai, Yangyang Deng, Caisheng He, Jiang Wang 0014, Jihua Feng, Zhiming Dai |
BMC Bioinform. | 7 |
| 2006 | Identifying Transcription Factor Binding Sites Based on a Neural Network
Zhiming Dai, Xianhua Dai, Jiang Wang 0014 |
ISNN (2) | 1 |