VLDB 2026 Research / reviewers in the wild / expert
Jielu Zhang
dblp:345/7809
· DBLP profile ↗
7ranked-venue papers
2as first author
7since 2021 · last 2026
0000-0002-4321-0580ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4 · 4 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
Trustworthy machine learning · 49% Segmentation and scene understanding · 16% Representation and self-supervised learning · 16% | |
| Network and information security
2 papers |
Security and privacy of machine learning · 100% | |
| Databases, data mining, and information retrieval
1 paper |
Information retrieval · 100% |
Topics — the 11 heaviest of 12, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Security and privacy of machine learning › adversarial attack
backdoor attack |
1.6 | 2 | 2025 | Mind Control through Causal Inference: Predicting Clean Images from Poisoned Data · ICLR 2025 BadSAM: Exploring Security Vulnerabilities of SAM via Backdoor Attacks (Student Abstract) · AAAI 2024 |
Machine learning › Trustworthy machine learning
robustness |
0.9 | 1 | 2025 | Mind Control through Causal Inference: Predicting Clean Images from Poisoned Data · ICLR 2025 |
Security and privacy of machine learning › adversarial attack › backdoor attack
backdoor defense |
0.9 | 1 | 2025 | Mind Control through Causal Inference: Predicting Clean Images from Poisoned Data · ICLR 2025 |
Machine learning › Trustworthy machine learning
fairness |
0.8 | 1 | 2024 | TorchSpatial: A Location Encoding Framework and Benchmark for Spatial Representation Learning · NeurIPS 2024 |
Machine learning › Trustworthy machine learning › dataset bias
geographic bias |
0.8 | 1 | 2024 | TorchSpatial: A Location Encoding Framework and Benchmark for Spatial Representation Learning · NeurIPS 2024 |
Computer vision › Segmentation and scene understanding
image segmentation |
0.8 | 1 | 2024 | BadSAM: Exploring Security Vulnerabilities of SAM via Backdoor Attacks (Student Abstract) · AAAI 2024 |
Machine learning › Deep learning architectures and training
positional encoding |
0.8 | 1 | 2024 | TorchSpatial: A Location Encoding Framework and Benchmark for Spatial Representation Learning · NeurIPS 2024 |
Machine learning › Representation and self-supervised learning
spatial representation learning |
0.8 | 1 | 2024 | TorchSpatial: A Location Encoding Framework and Benchmark for Spatial Representation Learning · NeurIPS 2024 |
Information retrieval
image retrieval |
0.8 | 1 | 2024 | Img2Loc: Revisiting Image Geolocalization using Multi-modality Foundation Models and Image-based Retrieval-Augmented Generation · SIGIR 2024 |
Information retrieval
retrieval-augmented generation |
0.8 | 1 | 2024 | Img2Loc: Revisiting Image Geolocalization using Multi-modality Foundation Models and Image-based Retrieval-Augmented Generation · SIGIR 2024 |
Information retrieval › multimedia analysis and retrieval
visual geolocalization |
0.8 | 1 | 2024 | Img2Loc: Revisiting Image Geolocalization using Multi-modality Foundation Models and Image-based Retrieval-Augmented Generation · SIGIR 2024 |
Methods — techniques the papers use, named apart from their topics
causal inference · 1.7attack indicator · 1.7backdoor attack · 1.5self-supervised learning · 0.8retrieval-augmented generation · 0.8multimodal large language model prompting · 0.8diffusion transformer · 0.8CLIP representation · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SpatialCausal : a spatially-aware causal inference deep learning model for out-of-hospital cardiac arrest survival predictionabstractRecently, numerous machine learning methods have been effectively applied to uncover spatial relationships between health risk factors and health outcomes. However, traditional machine learning methods often fail to address confounding bias, which arises when a common factor simultaneously influences both the treatment and the outcome – a challenge frequently encountered in observational studies. Deep learning-based causal inference models seek to mitigate confounding bias by learning balanced representations of covariates between treated and control groups, thereby reducing the dependence of treatment assignment on covariates. This enables accurate estimation of causal effects on health outcomes. Moreover, distinct geospatial patterns of risk exposure and health outcomes are common in many chronic diseases. Therefore, developing a spatially-aware causal inference model is essential for guiding geospatial health interventions. Here, we propose SpatialCausal, a spatially-aware deep learning-based causal inference model that explicitly integrates spatial, non-spatial, and unmeasured confounders, enabling accurate estimation of spatially-aware causal effects. We demonstrate the effectiveness of our approach through an application to Out-of-Hospital Cardiac Arrest survival outcome prediction. Our method surpasses state-of-the-art approaches and exhibits robust adaptability to various geospatial disease scenarios, making it a valuable tool for spatially-aware causal effect estimation in health geography. Jielu Zhang, Lan Mu, Gengchen Mai, Andrew Grundstein, Zhongliang Zhou, Donglan Zhang |
Int. J. Geogr. Inf. Sci. | 1 |
| 2026 | SpaCE: a spatial counterfactual explainable deep learning model for predicting out-of-hospital cardiac arrest survival outcomeabstractUnderstanding the relationship between risk factors, geospatial patterns, and disease outcomes is essential in health geography research. These relationships can inform the implementation of healthcare and public health strategies to improve health outcomes. To accurately uncover such complex relationships, it is necessary to have a predictive model capable of integrating both health variables and spatial information to forecast health outcomes, along with a tool to interpret and reveal the patterns identified by this model. We developed a Spatial Counterfactual Explainable Deep Learning model (SpaCE), comprising a spatially explicit health outcome predictor and a prototype-guided counterfactual explanation. The SpaCE model unifies geospatial and health variables to improve predictions and generates hypothetical examples with minimal changes but opposite outcomes. Using these counterfactuals, SpaCE assesses the impact of each variable in different spatial contexts. We evaluated the model for predicting cardiac arrest survival outcomes. With a 0.682 AUCROC score, the SpaCE exceeds baseline models by 10.2%. Further analysis also reveals that the geospatial context significantly affects how various risk factors affect the survival outcomes of patients. Overall, the SpaCE model significantly improves predictive accuracy and explainability. It provides targeted interventions at both individual and geographic levels, and the cardiac arrest case study shows its high adaptability to various disease scenarios. Jielu Zhang, Lan Mu, Donglan Zhang, Zhuo Chen 0012, Janani Rajbhandari-Thapa, José A. Pagán, Yan Li 0017, Gengchen Mai, Zhongliang Zhou |
Int. J. Geogr. Inf. Sci. | 1 |
| 2025 | Mind Control through Causal Inference: Predicting Clean Images from Poisoned DataabstractAnti-backdoor learning, aiming to train clean models directly from poisoned datasets, serves as an important defense method for backdoor attack. However, existing methods usually fail to recover backdoored samples to their original, correct labels and suffer from poor generalization to large pre-trained models due to its non end-to end training, making them unsuitable for protecting the increasingly prevalent large pre-trained models. To bridge the gap, we first revisit the anti-backdoor learning problem from a causal perspective. Our theoretical causal analysis reveals that incorporating \emph{\textbf{both}} images and the associated attack indicators preserves the model's integrity. Building on the theoretical analysis, we introduce an end-to-end method, Mind Control through Causal Inference (MCCI), to train clean models directly from poisoned datasets. This approach leverages both the image and the attack indicator to train the model. Based on this training paradigm, the model’s perception of whether an input is clean or backdoored can be controlled. Typically, by introducing fake non-attack indicators, the model perceives all inputs as clean and makes correct predictions, even for poisoned samples. Extensive experiments demonstrate that our method achieves state-of-the-art performance, efficiently recovering the original correct predictions for poisoned samples and enhancing accuracy on clean samples. Mengxuan Hu, Zihan Guan 0001, Yi Zeng 0005, Zhongliang Zhou, Jielu Zhang, Ruoxi Jia 0001, Anil Vullikanti, Sheng Li 0001 |
ICLR | 6 |
| 2025 | LocDiff: Identifying Locations on Earth by Diffusing in the Hilbert SpaceabstractImage geolocalization is a fundamental yet challenging task, aiming at inferring the geolocation on Earth where an image is taken. State-of-the-art methods employ either grid-based classification or gallery-based image-location retrieval, whose spatial generalizability significantly suffers if the spatial distribution of test images does not align with the choices of grids and galleries. Recently emerging generative approaches, while getting rid of grids and galleries, use raw geographical coordinates and suffer quality losses due to their lack of multi-scale information. To address these limitations, we propose a multi-scale latent diffusion model called LocDiff for image geolocalization. We developed a novel positional encoding-decoding framework called Spherical Harmonics Dirac Delta (SHDD) Representations, which encodes points on a spherical surface (e.g., geolocations on Earth) into a Hilbert space of Spherical Harmonics coefficients and decodes points (geolocations) by mode-seeking on spherical probability distributions. We also propose a novel SirenNet-based architecture (CS-UNet) to learn an image-based conditional backward process in the latent SHDD space by minimizing a latent KL-divergence loss. To the best of our knowledge, LocDiff is the first image geolocalization model that performs latent diffusion in a multi-scale location encoding space and generates geolocations under the guidance of images. Experimental results show that LocDiff can outperform all state-of-the-art grid-based, retrieval-based, and diffusion-based baselines across 5 challenging global-scale image geolocalization datasets, and demonstrates significantly stronger generalizability to unseen geolocations. Zeping Liu, Jielu Zhang, Zhongliang Zhou, Nemin Wu, Lan Mu, Yiqun Xie, Ni Lao, Gengchen Mai |
NeurIPS | 3 |
| 2024 | BadSAM: Exploring Security Vulnerabilities of SAM via Backdoor Attacks (Student Abstract)abstractImage segmentation is foundational to computer vision applications, and the Segment Anything Model (SAM) has become a leading base model for these tasks. However, SAM falters in specialized downstream challenges, leading to various customized SAM models. We introduce BadSAM, a backdoor attack tailored for SAM, revealing that customized models can harbor malicious behaviors. Using the CAMO dataset, we confirm BadSAM's efficacy and identify SAM vulnerabilities. This study paves the way for the development of more secure and customizable vision foundation models. Zihan Guan 0001, Mengxuan Hu, Zhongliang Zhou, Jielu Zhang, Sheng Li 0001, Ninghao Liu 0001 |
AAAI | 4 |
| 2024 | TorchSpatial: A Location Encoding Framework and Benchmark for Spatial Representation LearningabstractSpatial representation learning (SRL) aims at learning general-purpose neural network representations from various types of spatial data (e.g., points, polylines, polygons, networks, images, etc.) in their native formats. Learning good spatial representations is a fundamental problem for various downstream applications such as species distribution modeling, weather forecasting, trajectory generation, geographic question answering, etc. Even though SRL has become the foundation of almost all geospatial artificial intelligence (GeoAI) research, we have not yet seen significant efforts to develop an extensive deep learning framework and benchmark to support SRL model development and evaluation. To fill this gap, we propose TorchSpatial, a learning framework and benchmark for location (point) encoding,which is one of the most fundamental data types of spatial representation learning. TorchSpatial contains three key components: 1) a unified location encoding framework that consolidates 15 commonly recognized location encoders, ensuring scalability and reproducibility of the implementations; 2) the LocBench benchmark tasks encompassing 7 geo-aware image classification and 10 geo-aware imageregression datasets; 3) a comprehensive suite of evaluation metrics to quantify geo-aware models’ overall performance as well as their geographic bias, with a novel Geo-Bias Score metric. Finally, we provide a detailed analysis and insights into the model performance and geographic bias of different location encoders. We believe TorchSpatial will foster future advancement of spatial representationlearning and spatial fairness in GeoAI research. The TorchSpatial model framework and LocBench benchmark are available at https://github.com/seai-lab/TorchSpatial, and the Geo-Bias Score evaluation framework is available at https://github.com/seai-lab/PyGBS. Nemin Wu, Zeping Liu, Yanlin Qi, Jielu Zhang, Joshua Ni, Xiaobai Angela Yao, Lan Mu, Stefano Ermon, Tanuja Ganu, Akshay Uttama Nambi, Ni Lao, Gengchen Mai |
NeurIPS | 6 |
| 2024 | Img2Loc: Revisiting Image Geolocalization using Multi-modality Foundation Models and Image-based Retrieval-Augmented GenerationabstractGeolocating precise locations from images presents a challenging problem in computer vision and information retrieval. Traditional methods typically employ either classification-dividing the Earth's surface into grid cells and classifying images accordingly, or retrieval-identifying locations by matching images with a database of image-location pairs. However, classification-based approaches are limited by the cell size and cannot yield precise predictions, while retrieval-based systems usually suffer from poor search quality and inadequate coverage of the global landscape at varied scale and aggregation levels. To overcome these drawbacks, we present Img2Loc, a novel system that redefines image geolocalization as a text generation task. This is achieved using cutting-edge large multi-modality models (LMMs) like GPT-4V or LLaVA with retrieval augmented generation. Img2Loc first employs CLIP-based representations to generate an image-based coordinate query database. It then uniquely combines query results with images itself, forming elaborate prompts customized for LMMs. When tested on benchmark datasets such as Im2GPS3k and YFCC4k, Img2Loc not only surpasses the performance of previous state-of-the-art models but does so without any model training. A video demonstration of the system can be accessed via this link https://drive.google.com/file/d/16A6A-mc7AyUoKHRH3_WBRToRC13sn7tU/view?usp=sharing Zhongliang Zhou, Jielu Zhang, Zihan Guan 0001, Mengxuan Hu, Ni Lao, Lan Mu, Sheng Li 0001, Gengchen Mai |
SIGIR | 2 |