Yi-Yun Cheng

dblp:222/6722 · DBLP profile ↗
← Back
4ranked-venue papers
4as first author
3since 2021 · last 2026
0000-0001-6123-7595ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 4 · 4 first-author · 3 since 2021
YearPublicationVenuePosition
2026 Extracting geographic relations from large social media text data
abstract
Extracting geographical information from a large corpus of social media text is useful for monitoring live events such as natural disasters and public health crises. However, the noisy nature of texts creates challenges for reliable geoparsing, geocoding, and geotagging. These challenges can be remedied with relation extraction (RE) to efficiently identify geographic relations, but existing RE solutions are not domain-specific for geographic contexts. In this study, we domain-adapt and validate existing RE models to identify variations of geographic relations used in unstructured texts. We analyze 163,037 tweets containing counties and state names of the United States with 1,672 annotated tweets as training data. We apply five RE models with domain-adaptation: (1) our own heuristics-based model; (2) OpenIE; (3) BERT; (4) GPT-3.5-turbo; and (5) Mistral. We identify a list of lexically similar but semantically different relations and categorized them into meronymic, prepositional, ‘include’, ‘locate’, geographic noun, and other spatial relations. The contributions of this study are twofold: (1) we provide a list of geographic relations that can be used for GIS tasks that require more accurate detection of location from text; (2) we show that the performance of large language models (LLMs) improved with domain-specific training for geographic relation extraction.
Yi-Yun Cheng, Ly Dinh
Int. J. Geogr. Inf. Sci.1
2025 An experiment on the impact of relation types towards taxonomy alignment problems
abstract
This study investigates the impact of five relation types towards taxonomy alignment problems and finds that the presence and prevalence of each relation type have profound impact on the resulting merged solutions. The five relation types used in this study are: equal , include , included-in , overlap , and disjoint . We take a logic-based approach to work with the taxonomy alignment problem, and evaluate (1) the presence of a relation type and (2) the prevalence of a relation type towards the number of possible worlds (merged solutions) produced. We find that equal relation type has a minimal effect on the number of possible worlds, whereas include , included-in , overlap , and disjoint are all significantly impacting the number of possible worlds. We also find that more than one of a relation type from either of the four ( include , included-in , overlap , disjoint ) can substantially increase the number of possible worlds, with overlap and disjoint contributing to a profound exponential growth of that number. This study demonstrates how choices of relation types are important in taxonomy alignment problems, and that aligning taxonomies beyond equivalence is necessary to accurately capture the nuances of real-world taxonomies. • This study is the first to examine different relation types on taxonomy alignments. • This study finds that only using equals for mapping is not enough. • This study demonstrates how combinations of relation types on alignment matters.
Yi-Yun Cheng, Ly Dinh
Inf. Process. Manag.1
2025 Under whose wings? A conceptual model for incorporating historical sovereignty information in biodiversity data
abstract
Abstract Linking historical and contemporary geographic information in biodiversity data is a useful approach to approximate species population. However, one of the prominent factors that causes ambiguity in geographic information, and hinders the linking process, is the way sovereignty information is used. While historical biodiversity records often use sovereignties as proxies for geographic information about a species, contemporary records do not. This study proposes a conceptual model that incorporates sovereignty information in biodiversity data to foster the linkage between historical and contemporary geographical information. The model comprises two phases: the first phase relates tangible data sources and core components needed to construct historical sovereignty taxonomies; and the second phase is a process model to align historical sovereignty taxonomies with contemporary taxonomies in four phases. The output of the model presents all possible sovereignties that a geographic entity belongs to based on the degree of congruence between the historical and contemporary taxonomies. The contributions of this work are threefold: (1) making all possible ambiguities in historical geographic information explicit in biodiversity data; (2) bringing attention to the modeling choices that domain experts have to make when deciding which sovereignty a place name belongs to; and (3) extending and improving current geo‐referencing practices.
Yi-Yun Cheng
J. Assoc. Inf. Sci. Technol.1
2016 A Study on the Best Practice for Constructing a Cross-lingual Ontology
Yi-Yun Cheng, Hsueh-Hua Chen
Dublin Core Conference1