VLDB 2026 Research / reviewers in the wild / expert
Zekun Li 0007
dblp:150/2008-7
· DBLP profile ↗
14ranked-venue papers
6as first author
11since 2021 · last 2025
0000-0001-9603-9329ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 6 first-author · 9 since 2021Databases, data management, data science and information retrieval · 8 · 4 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 2 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 1 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Rethinking the Role of Prompting Strategies in LLM Test-Time Scaling: A Perspective of Probability TheoryabstractRecently, scaling test-time compute on Large Language Models (LLM) has garnered wide attention. However, there has been limited investigation of how various reasoning prompting strategies perform as scaling. In this paper, we focus on a standard and realistic scaling setting: majority voting. We systematically conduct experiments on 6 LLMs $\times$ 8 prompting strategies $\times$ 6 benchmarks. Experiment results consistently show that as the sampling time and computational overhead increase, complicated prompting strategies with superior initial performance gradually fall behind simple Chain-of-Thought. We analyze this phenomenon and provide theoretical proofs. Additionally, we propose a probabilistic method to efficiently predict scaling performance and identify the best prompting strategy under large sampling times, eliminating the need for resource-intensive inference processes in practical applications. Furthermore, we introduce two ways derived from our theoretical analysis to significantly improve the scaling performance. We hope that our research can promote to re-examine the role of complicated prompting, unleash the potential of simple prompting strategies, and provide new insights for enhancing test-time scaling performance. Code is available at https://github.com/MraDonkey/rethinking_prompting. Yexiang Liu, Zekun Li 0007, Zhi Fang, Nan Xu 0014, Ran He 0001, Tieniu Tan |
ACL (1) | 2 |
| 2025 | Benchmarking Geospatial Question Answering with MapQAabstractGeospatial question answering (QA) is a fundamental task in navigation and point of interest (POI) searches, yet existing datasets are limited in scale, diversity, and they rely on text-only descriptions without incorporating geometries. We introduce MapQA, a dataset that couples question-answer pairs with geo-entity geometries from OpenStreetMap (OSM) across two regions (Southern California and Illinois). MapQA contains 3,154 QA pairs covering nine geospatial reasoning types, including neighborhood inference and type identification, expanding both the quantity and variety of existing resources. To evaluate methods, we compare (1) a retrieval-based model that ranks geo-entities by embedding similarity and (2) large language models (LLMs) that translate questions into SQL queries executed on OSM. Retrieval-based models capture spatial relations like closeness and direction but fail on explicit distance computations, while LLMs excel at one-hop reasoning yet struggle with multi-hop tasks, revealing a key challenge for future systems. MapQA is publicly available at https://github.com/knowledge-computing/MapQA-dataset. Zekun Li 0007, Malcolm Grossman, Ehsan Qasemi, Mihir Kulkarni, Muhao Chen 0001, Yao-Yi Chiang |
SIGSPATIAL/GIS | 1 |
| 2025 | DIGMAPPER: A Modular System for Automated Geologic Map DigitizationabstractHistorical geologic maps contain rich geospatial information—such as rock units, faults, folds, and bedding planes—that is critical for assessing mineral resources essential to renewable energy, electric vehicles, and national security. However, digitizing maps remains a labor-intensive and time-consuming task. We present DIGMAPPER, a modular, scalable system developed in collaboration with the United States Geological Survey (USGS) to automate the digitization of geologic maps. DIGMAPPER features a fully dockerized, workflow-orchestrated architecture that integrates state-of-the-art deep learning models for map layout analysis, feature extraction, and georeferencing. To overcome challenges such as limited training data and complex visual content, our system employs innovative techniques, including in-context learning with large language models, synthetic data generation, and transformer-based models. Evaluations on over 100 annotated maps from the DARPA-USGS dataset demonstrate high accuracy across polygon, line, and point feature extraction, and reliable georeferencing performance. Deployed at USGS, DIGMAPPER significantly accelerates the creation of analysis-ready geospatial datasets, supporting national-scale critical mineral assessments and broader geoscientific applications. Yao-Yi Chiang, Theresa Chen, Michael P. Gerlek, Leeje Jang, Sofia Kirsanova, Craig A. Knoblock, Fandel Lin, Yijun Lin 0001, Zekun Li 0007, Steven N. Minton |
SIGSPATIAL/GIS | 10 |
| 2025 | ICDAR 2025 Competition on Historical Map Text Detection, Recognition, and Linking
Yijun Lin 0001, Solenn Tual, Zekun Li 0007, Leeje Jang, Yao-Yi Chiang, Jerod J. Weinman, Joseph Chazalon, Edwin Carlinet, Julien Perret, Nathalie Abadie, Bertrand Dumenieu, Ta-Chien Chan, Hsiung-Ming Liao, Wen-Rong Su, Mengjie Zou, Tianhao Dai, Rémi Petitpierre, Beatrice Vaienti, Frédéric Kaplan, Isabella diLenardo, Youngmin Baek, Michael Hentschel, Yu Nakagome, Ichimura Shuta, Jeongtae Lee, Chankyu Choi |
ICDAR (5) | 3 |
| 2025 | MuirBench: A Comprehensive Benchmark for Robust Multi-image UnderstandingabstractWe introduce MuirBench, a comprehensive benchmark that focuses on robust multi-image understanding capabilities of multimodal LLMs. MuirBench consists of 12 diverse multi-image tasks (e.g., scene understanding, ordering) that involve 10 categories of multi-image relations (e.g., multiview, temporal relations). Comprising 11,264 images and 2,600 multiple-choice questions, MuirBench is created in a pairwise manner, where each standard instance is paired with an unanswerable variant that has minimal semantic differences, in order for a reliable assessment. Evaluated upon 20 recent multi-modal LLMs, our results reveal that even the best-performing models like GPT-4o and Gemini Pro find it challenging to solve MuirBench, achieving 68.0% and 49.3% in accuracy. Open-source multimodal LLMs trained on single images can hardly generalize to multi-image questions, hovering below 33.3% in accuracy. These results highlight the importance of MuirBench in encouraging the community to develop multimodal LLMs that can look beyond a single image, suggesting potential pathways for future improvements. Fei Wang 0060, James Y. Huang, Zekun Li 0007, Qin Liu 0010, Xiaogeng Liu, Mingyu Derek Ma, Nan Xu 0014, Wenxuan Zhou 0002, Kai Zhang 0008, Tianyi Lorena Yan, Wenjie Mo 0001, Hsiang-Hui Liu, Pan Lu, Chunyuan Li, Chaowei Xiao, Kai-Wei Chang 0001, Dan Roth 0001, Sheng Zhang 0012, Hoifung Poon, Muhao Chen 0001 |
ICLR | 4 |
| 2024 | ICDAR 2024 Competition on Historical Map Text Detection, Recognition, and Linking
Zekun Li 0007, Yijun Lin 0001, Yao-Yi Chiang, Jerod J. Weinman, Solenn Tual, Joseph Chazalon, Julien Perret, Bertrand Dumenieu, Nathalie Abadie |
ICDAR (6) | 1 |
| 2023 | GeoLM: Empowering Language Models for Geospatially Grounded Language UnderstandingabstractHumans subconsciously engage in geospatial reasoning when reading articles.We recognize place names and their spatial relations in text and mentally associate them with their physical locations on Earth.Although pretrained language models can mimic this cognitive process using linguistic context, they do not utilize valuable geospatial information in large, widely available geographical databases, e.g., OpenStreetMap.This paper introduces GEOLM ( ), a geospatially grounded language model that enhances the understanding of geo-entities in natural language.GEOLM leverages geo-entity mentions as anchors to connect linguistic information in text corpora with geospatial information extracted from geographical databases.GEOLM connects the two types of context through contrastive learning and masked language modeling.It also incorporates a spatial coordinate embedding mechanism to encode distance and direction relations to capture geospatial context.In the experiment, we demonstrate that GEOLM exhibits promising capabilities in supporting toponym recognition, toponym linking, relation extraction, and geo-entity typing, which bridge the gap between natural language processing and geospatial sciences.The code is publicly available at https://github.com/ knowledge-computing/geolm. Zekun Li 0007, Wenxuan Zhou 0002, Yao-Yi Chiang, Muhao Chen 0001 |
EMNLP | 1 |
| 2023 | The mapKurator System: A Complete Pipeline for Extracting and Linking Text from Historical MapsabstractScanned historical maps in libraries and archives are valuable repositories of geographic data that often do not exist elsewhere. Despite the potential of machine learning tools like the Google Vision APIs for automatically transcribing text from these maps into machine-readable formats, they do not work well with large-sized images (e.g., high-resolution scanned documents), cannot infer the relation between the recognized text and other datasets, and are challenging to integrate with post-processing tools. This paper introduces the mapKurator system, an end-to-end system integrating machine learning models with a comprehensive data processing pipeline. mapKurator empowers automated extraction, post-processing, and linkage of text labels from large numbers of large-dimension historical map scans. The output data, comprising bounding polygons and recognized text, is in the standard GeoJSON format, making it easily modifiable within Geographic Information Systems (GIS). The proposed system allows users to quickly generate valuable data from large numbers of historical maps for in-depth analysis of the map content and, in turn, encourages map findability, accessibility, interoperability, and reusability (FAIR principles). We deployed the mapKurator system and enabled the processing of over 60,000 maps and over 100 million text/place names in the David Rumsey Historical Map collection. We also demonstrated a seamless integration of mapKurator with a collaborative web platform to enable accessing automated approaches for extracting and linking text labels from historical map scans and collective work to improve the results. Zekun Li 0007, Yijun Lin 0001, Min Namgung, Leeje Jang, Yao-Yi Chiang |
SIGSPATIAL/GIS | 2 |
| 2023 | Exploiting Polygon Metadata to Understand Raster Maps - Accurate Polygonal Feature ExtractionabstractLocating undiscovered deposits of critical minerals requires accurate geological data. However, most of the 100,000 historical geological maps of the United States Geological Survey (USGS) are in raster format. This hinders critical mineral assessment. We target the problem of extracting geological features represented as polygons from raster maps. We exploit the polygon metadata that provides information on the geological features, such as the map keys indicating how the polygon features are represented, to extract the features. We present a metadata-driven machine-learning approach that encodes the raster map and map key into a series of bitmaps and uses a convolutional model to learn to recognize the polygon features. We evaluated our approach on USGS geological maps; our approach achieves a median F1 score of 0.809 and outperforms state-of-the-art methods by 4.52%. Fandel Lin, Craig A. Knoblock, Basel Shbita, Zekun Li 0007, Yao-Yi Chiang |
SIGSPATIAL/GIS | 5 |
| 2023 | The Best Protection is Attack: Fooling Scene Text Recognition With Minimal PixelsabstractScene text recognition (STR) has witnessed tremendous progress in the era of deep learning, but it also raises concerns about privacy infringement as scene texts usually contain valuable or sensitive information. Previous works in privacy protection of scene texts mainly focus on masking out the texts from the image/video. In this work, we learn from the idea of adversarial examples and use minimal pixel perturbation to protect the privacy of text information. Although there are well-established attacking methods on non-sequential vision tasks (e.g., classification), the attack on sequential tasks (e.g., scene text recognition) has not received sufficient attention yet. Moreover, existing works mainly focus on the white-box setting, which requires complete knowledge of the target model (e.g., architecture, parameters, or gradients). These requirements limit the scope of applications for the white-box adversarial attack. Therefore, we propose a novel black-box attacking approach for the STR models, only requiring prior knowledge of the model output. Besides, instead of disturbing most pixels as in existing STR attack methods, our proposed approach only manipulates a few pixels, meaning the perturbation is more inconspicuous. To determine the location and value of the manipulated pixels, we also provide an efficient Adaptive-Discrete Differential Evolution (AD$^{2}\text{E}$) by narrowing down the continuous searching space to a discrete space. It can greatly reduce the queries to the target model. Experiments on several real-world benchmarks show the effectiveness of our proposed approach. Especially, when attacking the commercial STR engine, Baidu-OCR, our method achieves higher attack success rates by a large margin than existing approaches. Our work establishes an important step towards using the black-box adversarial attack with minimal pixels to protect the privacy of text information from being easily obtained by STR models. Yikun Xu, Pengwen Dai, Zekun Li 0007, Hongjun Wang 0005, Xiaochun Cao |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2022 | ACE: Anchor-Free Corner Evolution for Real-Time Arbitrarily-Oriented Object DetectionabstractObjects with different orientations are ubiquitous in the real world (e.g., texts/hands in the scene image, objects in the aerial image, etc.), and the widely-used axis-aligned bounding box does not compactly enclose the oriented objects. Thus arbitrarily-oriented object detection has attracted rising attention in recent years. In this paper, we propose a novel and effective model to detect arbitrarily-oriented objects. Instead of directly predicting the angles of oriented bounding boxes like most existing methods, we evolve the axis-aligned bounding box to the oriented quadrilateral box with the assistance of dynamically gathering contour information. More specifically, we first obtain the axis-aligned bounding box in an anchor-free manner. After that, we set the key points based on the sampled contour points of the axis-aligned bounding box. To improve the localization performance, we enrich the feature representations of these key points by exploiting a dynamic information gathering mechanism. This technique propagates the geometrical and semantic information along the sampled contour points, and fuses the information from the semantic neighbors of each sampled point, which varies for different locations. Finally, we estimate the offsets between the axis-aligned bounding box key points and the oriented quadrilateral box corner points. Extensive experiments on two frequently-used aerial image benchmarks HRSC2016 and DOTA, as well as scene text/hand datasets ICDAR2015, TD500, and Oxford-Hand, demonstrate the effectiveness and advantage of our proposed model. Pengwen Dai, Siyuan Yao, Zekun Li 0007, Sanyi Zhang, Xiaochun Cao |
IEEE Trans. Image Process. | 3 |
| 2020 | An Automatic Approach for Generating Rich, Linked Geo-Metadata from Historical Map ImagesabstractHistorical maps contain detailed geographic information difficult to find elsewhere covering long-periods of time (e.g., 125 years for the historical topographic maps in the US). However, these maps typically exist as scanned images without searchable metadata. Existing approaches making historical maps searchable rely on tedious manual work (including crowd-sourcing) to generate the metadata (e.g., geolocations and keywords). Optical character recognition (OCR) software could alleviate the required manual work, but the recognition results are individual words instead of location phrases (e.g., "Black'' and "Mountain'' vs. "Black Mountain''). This paper presents an end-to-end approach to address the real-world problem of finding and indexing historical map images. This approach automatically processes historical map images to extract their text content and generates a set of metadata that is linked to large external geospatial knowledge bases. The linked metadata in the RDF (Resource Description Framework) format support complex queries for finding and indexing historical maps, such as retrieving all historical maps covering mountain peaks higher than 1,000 meters in California. We have implemented the approach in a system called mapKurator. We have evaluated mapKurator using historical maps from several sources with various map styles, scales, and coverage. Our results show significant improvement over the state-of-the-art methods. The code has been made publicly available as modules of the Kartta Labs project at https://github.com/kartta-labs/Project. Zekun Li 0007, Yao-Yi Chiang, Sasan Tavakkol, Basel Shbita, Johannes H. Uhl, Stefan Leyk, Craig A. Knoblock |
KDD | 1 |
| 2019 | Generating Historical Maps from Online MapsabstractThis paper proposes an automatic system to generate a large amount of data for the training of text detection systems for historical maps. The system takes online maps as input and learns a conditional GAN model, to generate realistic historical map images from existing geographic datasets. Then the system uses the generated images as the base map and inserts synthetic text. Since the system has the control of text content, font style, and location, the system can obtain ground truth information (minimum bounding boxes) of the synthetic text. To overcome the challenge of content mismatch, the proposed system uses a novel loss function to encourage the generation of historical cartographic symbols in the foreground areas and discourage the generation in the background. The final output is a set of images resembling historical maps and the minimum bounding boxes around text regions on the images as annotations. Zekun Li 0007 |
SIGSPATIAL/GIS | 1 |
| 2018 | Weighted Feature Pooling Network in Template-Based Recognition
Zekun Li 0007, Yue Wu 0001, Wael Abd-Almageed, Premkumar Natarajan |
ACCV (5) | 1 |