EDBT 2026 Demo / reviewers in the wild / expert
Zhongliang Zhou
dblp:169/4613
· DBLP profile ↗
16ranked-venue papers
4as first author
16since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 1 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 3 first-author · 6 since 2021Databases, data management, data science and information retrieval · 5 · 2 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Reducing Token Redundancy in LVLMs: A Systematic Review of Token Pruning MethodsabstractHanzhang Yuan, Mengxuan Hu, Wenhao Zhang, Tianlong Wang, Zhongliang Zhou, Jiasen Lu, Sheng Li. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Hanzhang Yuan, Mengxuan Hu, Tianlong Wang, Zhongliang Zhou, Jiasen Lu |
ACL (1) | 5 |
| 2026 | SpatialCausal : a spatially-aware causal inference deep learning model for out-of-hospital cardiac arrest survival predictionabstractRecently, numerous machine learning methods have been effectively applied to uncover spatial relationships between health risk factors and health outcomes. However, traditional machine learning methods often fail to address confounding bias, which arises when a common factor simultaneously influences both the treatment and the outcome – a challenge frequently encountered in observational studies. Deep learning-based causal inference models seek to mitigate confounding bias by learning balanced representations of covariates between treated and control groups, thereby reducing the dependence of treatment assignment on covariates. This enables accurate estimation of causal effects on health outcomes. Moreover, distinct geospatial patterns of risk exposure and health outcomes are common in many chronic diseases. Therefore, developing a spatially-aware causal inference model is essential for guiding geospatial health interventions. Here, we propose SpatialCausal, a spatially-aware deep learning-based causal inference model that explicitly integrates spatial, non-spatial, and unmeasured confounders, enabling accurate estimation of spatially-aware causal effects. We demonstrate the effectiveness of our approach through an application to Out-of-Hospital Cardiac Arrest survival outcome prediction. Our method surpasses state-of-the-art approaches and exhibits robust adaptability to various geospatial disease scenarios, making it a valuable tool for spatially-aware causal effect estimation in health geography. Jielu Zhang, Lan Mu, Gengchen Mai, Andrew Grundstein, Zhongliang Zhou, Donglan Zhang |
Int. J. Geogr. Inf. Sci. | 5 |
| 2026 | SpaCE: a spatial counterfactual explainable deep learning model for predicting out-of-hospital cardiac arrest survival outcomeabstractUnderstanding the relationship between risk factors, geospatial patterns, and disease outcomes is essential in health geography research. These relationships can inform the implementation of healthcare and public health strategies to improve health outcomes. To accurately uncover such complex relationships, it is necessary to have a predictive model capable of integrating both health variables and spatial information to forecast health outcomes, along with a tool to interpret and reveal the patterns identified by this model. We developed a Spatial Counterfactual Explainable Deep Learning model (SpaCE), comprising a spatially explicit health outcome predictor and a prototype-guided counterfactual explanation. The SpaCE model unifies geospatial and health variables to improve predictions and generates hypothetical examples with minimal changes but opposite outcomes. Using these counterfactuals, SpaCE assesses the impact of each variable in different spatial contexts. We evaluated the model for predicting cardiac arrest survival outcomes. With a 0.682 AUCROC score, the SpaCE exceeds baseline models by 10.2%. Further analysis also reveals that the geospatial context significantly affects how various risk factors affect the survival outcomes of patients. Overall, the SpaCE model significantly improves predictive accuracy and explainability. It provides targeted interventions at both individual and geographic levels, and the cardiac arrest case study shows its high adaptability to various disease scenarios. Jielu Zhang, Lan Mu, Donglan Zhang, Zhuo Chen 0012, Janani Rajbhandari-Thapa, José A. Pagán, Yan Li 0017, Gengchen Mai, Zhongliang Zhou |
Int. J. Geogr. Inf. Sci. | 9 |
| 2026 | Multiscale Spatio-Temporal Graph Convolutional Network for UAV Anomaly Detection
Gang Hu 0018, Zhongliang Zhou, Shi-tao Chen |
IEEE Internet Things J. | 2 |
| 2026 | M2SC2-AD: Multiscale Missing-aware Anomaly Detection framework with scale consistency constraints for sparse multi-sensor
Gang Hu 0018, Zhongliang Zhou |
Inf. Process. Manag. | 2 |
| 2025 | ClinicalRAG: Automating Pharmaceutical Label Quality Control with Hierarchical RAG and Large Language ModelsabstractEvery pharmaceutical product must be accompanied by a comprehensive label that delineates its indications, usage, dosages, and side effects, essential for safe medication practices. Traditionally, creating drug labels is labor-intensive and dependent on manual quality checks. Recent advancements in Large Language Models (LLMs) offer a promising avenue to streamline this process. In this paper we introduce ClinicalRAG, an automated labeling quality control pipeline that integrates LLM with hierarchical Retrieval Augmented Generation that allows to cross-check every statement in the drug label document. ClinicalRAG enhances the reliability of automated drug labeling by systematically reducing hallucination risks, achieving an accuracy of 96.1% in internal validation. With user-friendly interface, our pipeline aims to support pharmaceutical company in drug approval and expedite patients' access to new treatments. Qiaohui Zhou, Zhongliang Zhou, Michelle Ngo, Federico Ferrari, Junshui Ma |
AAAI | 2 |
| 2025 | Mind Control through Causal Inference: Predicting Clean Images from Poisoned DataabstractAnti-backdoor learning, aiming to train clean models directly from poisoned datasets, serves as an important defense method for backdoor attack. However, existing methods usually fail to recover backdoored samples to their original, correct labels and suffer from poor generalization to large pre-trained models due to its non end-to end training, making them unsuitable for protecting the increasingly prevalent large pre-trained models. To bridge the gap, we first revisit the anti-backdoor learning problem from a causal perspective. Our theoretical causal analysis reveals that incorporating \emph{\textbf{both}} images and the associated attack indicators preserves the model's integrity. Building on the theoretical analysis, we introduce an end-to-end method, Mind Control through Causal Inference (MCCI), to train clean models directly from poisoned datasets. This approach leverages both the image and the attack indicator to train the model. Based on this training paradigm, the model’s perception of whether an input is clean or backdoored can be controlled. Typically, by introducing fake non-attack indicators, the model perceives all inputs as clean and makes correct predictions, even for poisoned samples. Extensive experiments demonstrate that our method achieves state-of-the-art performance, efficiently recovering the original correct predictions for poisoned samples and enhancing accuracy on clean samples. Mengxuan Hu, Zihan Guan 0001, Yi Zeng 0005, Zhongliang Zhou, Jielu Zhang, Ruoxi Jia 0001, Anil Vullikanti, Sheng Li 0001 |
ICLR | 5 |
| 2025 | LocDiff: Identifying Locations on Earth by Diffusing in the Hilbert SpaceabstractImage geolocalization is a fundamental yet challenging task, aiming at inferring the geolocation on Earth where an image is taken. State-of-the-art methods employ either grid-based classification or gallery-based image-location retrieval, whose spatial generalizability significantly suffers if the spatial distribution of test images does not align with the choices of grids and galleries. Recently emerging generative approaches, while getting rid of grids and galleries, use raw geographical coordinates and suffer quality losses due to their lack of multi-scale information. To address these limitations, we propose a multi-scale latent diffusion model called LocDiff for image geolocalization. We developed a novel positional encoding-decoding framework called Spherical Harmonics Dirac Delta (SHDD) Representations, which encodes points on a spherical surface (e.g., geolocations on Earth) into a Hilbert space of Spherical Harmonics coefficients and decodes points (geolocations) by mode-seeking on spherical probability distributions. We also propose a novel SirenNet-based architecture (CS-UNet) to learn an image-based conditional backward process in the latent SHDD space by minimizing a latent KL-divergence loss. To the best of our knowledge, LocDiff is the first image geolocalization model that performs latent diffusion in a multi-scale location encoding space and generates geolocations under the guidance of images. Experimental results show that LocDiff can outperform all state-of-the-art grid-based, retrieval-based, and diffusion-based baselines across 5 challenging global-scale image geolocalization datasets, and demonstrates significantly stronger generalizability to unseen geolocations. Zeping Liu, Jielu Zhang, Zhongliang Zhou, Nemin Wu, Lan Mu, Yiqun Xie, Ni Lao, Gengchen Mai |
NeurIPS | 4 |
| 2025 | Determination of the Optimal Channel Configuration for Land Surface Temperature Retrieval Using Split Window AlgorithmabstractCurrently, various algorithms have been developed to retrieve regional and global Land surface temperature (LST) from satellite thermal infrared (TIR) observations, among which, the split window (SW) algorithm is the most widely used one. However, the LST retrieval accuracy would be affected by the channel centers and channel widths owing to the vast atmospheric conditions and land surface types around the world. The theoretical channel configuration leading to the best performance of the SW algorithm is still not well investigated currently. In this study, the LST retrieval accuracies of the SW algorithm using different channel configurations were studied iteratively through the whole TIR atmospheric window. Consequently, the two channels centered at 10.3 μm and 11.5 μm with the widths of 0.3 μm and 0.4 μm were found to be the optimal channel configuration for applying the SW algorithm. Based on the global atmospheric profiles provided in the ERA5 and SeeBor V5.0 database and the emissivity spectra provided in the ECOSTRESS library, the performance of the SW algorithm using the determined channel configuration was delicately evaluated. Results show that the LST retrieval Root Mean Square Error (RMSE) of the determined channel configuration was 1.09 K, better than that of the MODIS (1.28 K), Landsat-9 (1.24 K), and Sentinel-3A (1.24 K) instruments regarding the global atmospheric profiles provided in the ERA5 database. Similar results were obtained corresponding to SeeBor V5.0 atmospheric profiles with the LST retrieval RMSE of 1.28K (determined channel configuration), 1.49 K (MODIS), 1.96 K (Landsat-9), and 1.94 K (Sentinel-3A). Youying Guo, Xiaopo Zheng, Zhongliang Zhou, Dahui Li |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2024 | BadSAM: Exploring Security Vulnerabilities of SAM via Backdoor Attacks (Student Abstract)abstractImage segmentation is foundational to computer vision applications, and the Segment Anything Model (SAM) has become a leading base model for these tasks. However, SAM falters in specialized downstream challenges, leading to various customized SAM models. We introduce BadSAM, a backdoor attack tailored for SAM, revealing that customized models can harbor malicious behaviors. Using the CAMO dataset, we confirm BadSAM's efficacy and identify SAM vulnerabilities. This study paves the way for the development of more secure and customizable vision foundation models. Zihan Guan 0001, Mengxuan Hu, Zhongliang Zhou, Jielu Zhang, Sheng Li 0001, Ninghao Liu 0001 |
AAAI | 3 |
| 2024 | Img2Loc: Revisiting Image Geolocalization using Multi-modality Foundation Models and Image-based Retrieval-Augmented GenerationabstractGeolocating precise locations from images presents a challenging problem in computer vision and information retrieval. Traditional methods typically employ either classification-dividing the Earth's surface into grid cells and classifying images accordingly, or retrieval-identifying locations by matching images with a database of image-location pairs. However, classification-based approaches are limited by the cell size and cannot yield precise predictions, while retrieval-based systems usually suffer from poor search quality and inadequate coverage of the global landscape at varied scale and aggregation levels. To overcome these drawbacks, we present Img2Loc, a novel system that redefines image geolocalization as a text generation task. This is achieved using cutting-edge large multi-modality models (LMMs) like GPT-4V or LLaVA with retrieval augmented generation. Img2Loc first employs CLIP-based representations to generate an image-based coordinate query database. It then uniquely combines query results with images itself, forming elaborate prompts customized for LMMs. When tested on benchmark datasets such as Im2GPS3k and YFCC4k, Img2Loc not only surpasses the performance of previous state-of-the-art models but does so without any model training. A video demonstration of the system can be accessed via this link https://drive.google.com/file/d/16A6A-mc7AyUoKHRH3_WBRToRC13sn7tU/view?usp=sharing Zhongliang Zhou, Jielu Zhang, Zihan Guan 0001, Mengxuan Hu, Ni Lao, Lan Mu, Sheng Li 0001, Gengchen Mai |
SIGIR | 1 |
| 2024 | Using explainable machine learning to uncover the kinase-substrate interaction landscapeabstractMOTIVATION: Phosphorylation, a post-translational modification regulated by protein kinase enzymes, plays an essential role in almost all cellular processes. Understanding how each of the nearly 500 human protein kinases selectively phosphorylates their substrates is a foundational challenge in bioinformatics and cell signaling. Although deep learning models have been a popular means to predict kinase-substrate relationships, existing models often lack interpretability and are trained on datasets skewed toward a subset of well-studied kinases. RESULTS: Here we leverage recent peptide library datasets generated to determine substrate specificity profiles of 300 serine/threonine kinases to develop an explainable Transformer model for kinase-peptide interaction prediction. The model, trained solely on primary sequences, achieved state-of-the-art performance. Its unique multitask learning paradigm built within the model enables predictions on virtually any kinase-peptide pair, including predictions on 139 kinases not used in peptide library screens. Furthermore, we employed explainable machine learning methods to elucidate the model's inner workings. Through analysis of learned embeddings at different training stages, we demonstrate that the model employs a unique strategy of substrate prediction considering both substrate motif patterns and kinase evolutionary features. SHapley Additive exPlanation (SHAP) analysis reveals key specificity determining residues in the peptide sequence. Finally, we provide a web interface for predicting kinase-substrate associations for user-defined sequences and a resource for visualizing the learned kinase-substrate associations. AVAILABILITY AND IMPLEMENTATION: All code and data are available at https://github.com/esbgkannan/Phosformer-ST. Web server is available at https://phosformer.netlify.app. Zhongliang Zhou, Wayland Yeung, Saber Soleymani, Nathan Gravel, Mariah Salcedo, Sheng Li 0001, Natarajan Kannan |
Bioinform. | 1 |
| 2023 | Alignment-free estimation of sequence conservation for identifying functional sites using protein sequence embeddingsabstractProtein language modeling is a fast-emerging deep learning method in bioinformatics with diverse applications such as structure prediction and protein design. However, application toward estimating sequence conservation for functional site prediction has not been systematically explored. Here, we present a method for the alignment-free estimation of sequence conservation using sequence embeddings generated from protein language models. Comprehensive benchmarks across publicly available protein language models reveal that ESM2 models provide the best performance to computational cost ratio for conservation estimation. Applying our method to full-length protein sequences, we demonstrate that embedding-based methods are not sensitive to the order of conserved elements-conservation scores can be calculated for multidomain proteins in a single run, without the need to separate individual domains. Our method can also identify conserved functional sites within fast-evolving sequence regions (such as domain inserts), which we demonstrate through the identification of conserved phosphorylation motifs in variable insert segments in protein kinases. Overall, embedding-based conservation analysis is a broadly applicable method for identifying potential functional sites in any full-length protein sequence and estimating conservation in an alignment-free manner. To run this on your protein sequence of interest, try our scripts at https://github.com/esbgkannan/kibby. Wayland Yeung, Zhongliang Zhou, Sheng Li 0001, Natarajan Kannan |
Briefings Bioinform. | 2 |
| 2023 | Tree visualizations of protein sequence embedding space enable improved functional clustering of diverse protein superfamiliesabstractProtein language models, trained on millions of biologically observed sequences, generate feature-rich numerical representations of protein sequences. These representations, called sequence embeddings, can infer structure-functional properties, despite protein language models being trained on primary sequence alone. While sequence embeddings have been applied toward tasks such as structure and function prediction, applications toward alignment-free sequence classification have been hindered by the lack of studies to derive, quantify and evaluate relationships between protein sequence embeddings. Here, we develop workflows and visualization methods for the classification of protein families using sequence embedding derived from protein language models. A benchmark of manifold visualization methods reveals that Neighbor Joining (NJ) embedding trees are highly effective in capturing global structure while achieving similar performance in capturing local structure compared with popular dimensionality reduction techniques such as t-SNE and UMAP. The statistical significance of hierarchical clusters on a tree is evaluated by resampling embeddings using a variational autoencoder (VAE). We demonstrate the application of our methods in the classification of two well-studied enzyme superfamilies, phosphatases and protein kinases. Our embedding-based classifications remain consistent with and extend upon previously published sequence alignment-based classifications. We also propose a new hierarchical classification for the S-Adenosyl-L-Methionine (SAM) enzyme superfamily which has been difficult to classify using traditional alignment-based approaches. Beyond applications in sequence classification, our results further suggest NJ trees are a promising general method for visualizing high-dimensional data sets. Wayland Yeung, Zhongliang Zhou, Liju Mathew, Nathan Gravel, Rahil Taujale, Brady O'boyle, Mariah Salcedo, Aarya Venkat, William Lanzilotta, Sheng Li 0001, Natarajan Kannan |
Briefings Bioinform. | 2 |
| 2023 | Phosformer: an explainable transformer model for protein kinase-specific phosphorylation predictionsabstractMOTIVATION: The human genome encodes over 500 distinct protein kinases which regulate nearly all cellular processes by the specific phosphorylation of protein substrates. While advances in mass spectrometry and proteomics studies have identified thousands of phosphorylation sites across species, information on the specific kinases that phosphorylate these sites is currently lacking for the vast majority of phosphosites. Recently, there has been a major focus on the development of computational models for predicting kinase-substrate associations. However, most current models only allow predictions on a subset of well-studied kinases. Furthermore, the utilization of hand-curated features and imbalances in training and testing datasets pose unique challenges in the development of accurate predictive models for kinase-specific phosphorylation prediction. Motivated by the recent development of universal protein language models which automatically generate context-aware features from primary sequence information, we sought to develop a unified framework for kinase-specific phosphosite prediction, allowing for greater investigative utility and enabling substrate predictions at the whole kinome level. RESULTS: We present a deep learning model for kinase-specific phosphosite prediction, termed Phosformer, which predicts the probability of phosphorylation given an arbitrary pair of unaligned kinase and substrate peptide sequences. We demonstrate that Phosformer implicitly learns evolutionary and functional features during training, removing the need for feature curation and engineering. Further analyses reveal that Phosformer also learns substrate specificity motifs and is able to distinguish between functionally distinct kinase families. Benchmarks indicate that Phosformer exhibits significant improvements compared to the state-of-the-art models, while also presenting a more generalized, unified, and interpretable predictive framework. AVAILABILITY AND IMPLEMENTATION: Code and data are available at https://github.com/esbgkannan/phosformer. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Zhongliang Zhou, Wayland Yeung, Nathan Gravel, Mariah Salcedo, Saber Soleymani, Sheng Li 0001, Natarajan Kannan |
Bioinform. | 1 |
| 2022 | Pigmentation-based Visual Learning for Salvelinus fontinalis Individual Re-identificationabstractBrook trout (Salvelinus fontinalis) is a freshwater fish species of ecological, economic, and cultural importance in eastern North America. Estimating the abundance, movement, and survival of brook trout in the wild is an important task for environmental management, and current methods often involve physical tagging or collection of DNA samples for each of the fish as their unique identifier. However, this process is expensive and inefficient. Meanwhile, although deep learning methods have proven effective for individual recognition of humans, it remains challenging to apply this system to wildlife biology due to fewer available images, different biometric patterns, and relatively poor image quality. In this paper, we develop a framework to automate the process of individual recognition of brook trout. Distinguished from simply adopting traditional feature descriptors (e.g., SIFT and HOG) or using deep neural networks on the raw images, our framework utilizes multiple modalities consisting of the region of interest and gray-scaled pigmentation patterns. We use these multiple modality features in a Convolutional Neural network to generate feature vectors as fish descriptors. These descriptors are then used to distinguish individual brook trout by ranking their relative distance in latent space. Our experimental framework demonstrates better results than baseline methods such as SIFT and HOG while being more robust to distortions characteristic of large imagery datasets collected through crowdsourcing and citizen science. Zhongliang Zhou, Nathaniel P. Hitt, Benjamin H. Letcher, Weili Shi, Sheng Li 0001 |
IEEE Big Data | 1 |