VLDB 2026 Research / reviewers in the wild / expert
Lilin Yu
dblp:357/5838
· DBLP profile ↗
5ranked-venue papers
4as first author
5since 2021 · last 2026
0009-0008-1729-083XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 4 · 3 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 3 first-author · 4 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Mapping Change: A Temporal and Semantic Knowledge Base of Scottish Gazetteers
Lilin Yu, Angelo A. Salatino, Rosa Filgueira |
ESWC (2) | 1 |
| 2025 | ImpactLens: Using AI and Data to Communicate Research Impact EffectivelyabstractAlthough the impact science has on society is a dimension that has been receiving increasing attention in recent years, surprisingly little work has been done on developing software tools that track it systematically by integrating information from different data sources and trying to capture all of its dimensions explicitly (e.g. impact on the economy, health and wellbeing, public understanding, the environment etc).In this paper, we present the prototypical ImpactLens system, which aims to leverage modern-day data-driven and AI technologies to extract ‘impact stories’ from rich, yet diverse data sources. We argue that the development of such tools can significantly enhance the ability of research communities to communicate the value of their work to society, while also providing opportunities to benefit researcher development, strategic research planning, and better exploitation of research results. Rosa Filgueira, Lilin Yu, Xingran Ruan, Sypro Nita, Laura Moran, Michael Rovatsos |
eScience | 2 |
| 2025 | Frances++: LLM-Based Semantic Enrichment and Spatial Graphs for Digitized Historical CollectionsabstractThis paper presents major enhancements to frances, a platform for exploring digitized historical collections using LLM-based semantic and spatial enrichments. We focus on three corpora from the National Library of Scotland (NLS): the Gazetteers of Scotland, the Broadsides, and the Encyclopaedia Britannica. Using a flexible, prompt-based LLM pipeline, we extract over 50,000 structured articles from the Gazetteers, adapting to diverse typographic layouts. For Broadsides, we align OCR-derived text (extracted using defoe) with manually corrected transcriptions published by NLS, linking both sources for improved access and reuse. Historical text is further improved using a fine-tuned Llama2 model selected after evaluating several LLMs for post-OCR correction. To enrich geographic information, we combine Stanza NER, in-text coordinate parsing, and the Edinburgh Geoparser, enabling both modern georesolution and preservation of historical geography. We extend the Heritage Textual Ontology (HTO) to support article-level records, spatiotemporal entities, and in-text annotations via CRMgeo and Web Annotation standards. Resulting RDF knowledge graphs are deployed on a GeoSPARQL-enabled Fuseki server and indexed in Elasticsearch to support full-text, semantic, and spatial search. The upgraded frances interface offers entity highlighting, historical maps, and provenance tracking, demonstrating a scalable pipeline for structured, enriched access to OCRed heritage texts. Lilin Yu, Rosa Filgueira |
eScience | 1 |
| 2024 | Advancing frances: New Heritage Textual Ontology, Enhanced Knowledge Graphs, and Refined Search CapabilitiesabstractThis paper presents significant enhancements to the frances platform, incorporating the Heritage Textual Ontology (HTO), advanced knowledge graphs, and sophisticated search capabilities, along with innovative data visualization methods. The HTO integrates diverse historical collections and unifies various sources. Leveraging this ontology, the new knowledge graphs connect data across different sources and editions, linking to external resources like Wikipedia and Dbpedia to enrich semantic relationships. We employed deep-learning-based spell correction for OCR error correction. Enhanced search functionalities, powered by Elasticsearch and semantic technologies, enable precise retrieval and analysis. Additionally, new data visualization approaches offer multifaceted interpretations of search results. A case study tracking slavery references in historical editions of the Encyclopaedia Britannica 1768-1860 demonstrates the platform’s effectiveness in analyzing historical text and validating frances’s capabilities. Lilin Yu, Ash Charlton, Melissa Terras, Rosa Filgueira |
e-Science | 1 |
| 2023 | frances: Cloud-Based Historical Text Mining with Deep Learning and Parallel ProcessingabstractFrances is an advanced cloud-based text mining digital platform that leverages information extraction, knowledge graphs, natural language processing (NLP), deep learning, and parallel processing techniques. It has been specifically designed to unlock the full potential of historical digital textual collections, such as those from the National Library of Scotland, offering cloud-based capabilities and extended support for complex NLP analyses and data visualizations. frances enables realtime recurrent operational text mining and provides robust capabilities for temporal analysis, accompanied by automatic visualizations for easy result inspection. In this paper, we present the motivation behind the development of frances, emphasizing its innovative design and novel implementation aspects. We also outline future development directions, and we evaluate the platform through two comprehensive case studies in history and publishing history. Lilin Yu, Ash Charlton, Wilfrid Askins, Melissa Terras, Rosa Filgueira |
e-Science | 1 |