EDBT 2026 Demo / reviewers in the wild / expert
HyeongSik Kim 0001
dblp:83/9728 · also Hyeongsik Kim 0001
· DBLP profile ↗
14ranked-venue papers
6as first author
3since 2021 · last 2024
0000-0002-3002-9838ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 9 · 4 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 4 first-author · 1 since 2021Artificial intelligence and machine learning · 5 · 2 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
3 papers |
Query processing and optimization · 44% Database theory · 18% Data models and query languages · 12% | |
| Artificial intelligence
1 paper |
Question answering and dialogue systems · 100% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Bioinformatics and computational biology · 100% |
Topics — the 12 heaviest of 15, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Question answering and dialogue systems › dialogue dataset
dialogue dataset construction |
0.7 | 1 | 2023 | A Textual Dataset for Situated Proactive Response Selection · ACL (1) 2023 |
Database theory
ontology-mediated queries |
0.4 | 1 | 2019 | Semantic query transformations for increased parallelization in distributed knowledge graph query processing · SC 2019 |
Query processing and optimization
parallel query processing |
0.4 | 1 | 2019 | Semantic query transformations for increased parallelization in distributed knowledge graph query processing · SC 2019 |
Query processing and optimization
query rewriting |
0.3 | 1 | 2017 | Type-based Semantic Optimization for Scalable RDF Graph Pattern Matching · WWW 2017 |
Bioinformatics and computational biology › knowledge representation in biology
biomedical ontology |
0.2 | 1 | 2024 | Structured Prompt Interrogation and Recursive Extraction of Semantics (SPIRES): a method for populating knowledge bases using zero-shot learning · Bioinform. 2024 |
Distributed and cloud data management
mapreduce |
0.1 | 1 | 2011 | From SPARQL to MapReduce: The Journey Using a Nested TripleGroup Algebra · Proc. VLDB Endow. 2011 |
Data models and query languages
query language |
0.1 | 1 | 2011 | From SPARQL to MapReduce: The Journey Using a Nested TripleGroup Algebra · Proc. VLDB Endow. 2011 |
Information retrieval › cross-language information retrieval
query translation |
0.1 | 1 | 2011 | From SPARQL to MapReduce: The Journey Using a Nested TripleGroup Algebra · Proc. VLDB Endow. 2011 |
Graph data management › RDF data management
RDF query processing |
0.1 | 1 | 2011 | From SPARQL to MapReduce: The Journey Using a Nested TripleGroup Algebra · Proc. VLDB Endow. 2011 |
Data models and query languages › RDF query language
SPARQL |
0.1 | 1 | 2011 | From SPARQL to MapReduce: The Journey Using a Nested TripleGroup Algebra · Proc. VLDB Endow. 2011 |
Information retrieval › query reformulation
query expansion |
0.1 | 1 | 2019 | Semantic query transformations for increased parallelization in distributed knowledge graph query processing · SC 2019 |
Graph data management › graph data model
RDF data model |
0.1 | 1 | 2017 | Type-based Semantic Optimization for Scalable RDF Graph Pattern Matching · WWW 2017 |
Methods — techniques the papers use, named apart from their topics
zero-shot learning · 0.8relation extraction · 0.8prompt interrogation · 0.8large language model · 0.8neural conversation model · 0.7semantic query transformation · 0.4ontology inference · 0.4integrity constraints · 0.3indexing · 0.3nested triplegroup algebra · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Structured Prompt Interrogation and Recursive Extraction of Semantics (SPIRES): a method for populating knowledge bases using zero-shot learningabstractMOTIVATION: Creating knowledge bases and ontologies is a time consuming task that relies on manual curation. AI/NLP approaches can assist expert curators in populating these knowledge bases, but current approaches rely on extensive training data, and are not able to populate arbitrarily complex nested knowledge schemas. RESULTS: Here we present Structured Prompt Interrogation and Recursive Extraction of Semantics (SPIRES), a Knowledge Extraction approach that relies on the ability of Large Language Models (LLMs) to perform zero-shot learning and general-purpose query answering from flexible prompts and return information conforming to a specified schema. Given a detailed, user-defined knowledge schema and an input text, SPIRES recursively performs prompt interrogation against an LLM to obtain a set of responses matching the provided schema. SPIRES uses existing ontologies and vocabularies to provide identifiers for matched elements. We present examples of applying SPIRES in different domains, including extraction of food recipes, multi-species cellular signaling pathways, disease treatments, multi-step drug mechanisms, and chemical to disease relationships. Current SPIRES accuracy is comparable to the mid-range of existing Relation Extraction methods, but greatly surpasses an LLM's native capability of grounding entities with unique identifiers. SPIRES has the advantage of easy customization, flexibility, and, crucially, the ability to perform new tasks in the absence of any new training data. This method supports a general strategy of leveraging the language interpreting capabilities of LLMs to assemble knowledge bases, assisting manual knowledge curation and acquisition while supporting validation with publicly-available databases and ontologies external to the LLM. AVAILABILITY AND IMPLEMENTATION: SPIRES is available as part of the open source OntoGPT package: https://github.com/monarch-initiative/ontogpt. J. Harry Caufield, Harshad Hegde, Vincent Emonet, Nomi L. Harris, Marcin P. Joachimiak, Nicolas Matentzoglu, HyeongSik Kim 0001, Sierra A. T. Moxon, Justin T. Reese, Melissa A. Haendel, Peter N. Robinson, Chris Mungall |
Bioinform. | 7 |
| 2023 | A Textual Dataset for Situated Proactive Response SelectionabstractRecent data-driven conversational models are able to return fluent, consistent, and informative responses to many kinds of requests and utterances in task-oriented scenarios.However, these responses are typically limited to just the immediate local topic instead of being widerranging and proactively taking the conversation further, for example making suggestions to help customers achieve their goals.This inadequacy reflects a lack of understanding of the interlocutor's situation and implicit goal.To address the problem, we introduce a task of proactive response selection based on situational information.We present a manuallycurated dataset of 1.7k English conversation examples that include situational background information plus for each conversation a set of responses, only some of which are acceptable in the situation.A responsive and informed conversation system should select the appropriate responses and avoid inappropriate ones; doing so demonstrates the ability to adequately understand the initiating request and situation.Our benchmark experiments show that this is not an easy task even for strong neural models, offering opportunities for future research. Naoki Otani, Jun Araki, HyeongSik Kim 0001, Eduard H. Hovy |
ACL (1) | 3 |
| 2022 | EaT-PIM: Substituting Entities in Procedural Instructions Using Flow Graphs and Embeddings
Sola S. Shirai, HyeongSik Kim 0001 |
ISWC | 2 |
| 2019 | Semantic query transformations for increased parallelization in distributed knowledge graph query processingabstractOntologies have become an increasingly popular semantic layer for integrating multiple heterogeneous datasets. However, significant challenges remain with supporting efficient and scalable processing of queries with data linked with ontologies (ontological queries). Ontological query processing queries requires explicitly defined query patterns be expanded to capture implicit ones, based on available ontology inference axioms. However, in practice such as in the biomedical domain, the complexity of the ontological axioms results in significantly large query expansions which present day query processing infrastructure cannot support. In particular, it remains unclear how to effectively parallelize such queries. HyeongSik Kim 0001, Abhisha Bhattacharyya, Kemafor Anyanwu |
SC | 1 |
| 2018 | Scalable Exploratory Search on Knowledge Graphs Using Apache SparkabstractFaceted search is a popular exploratory search paradigm on Big Knowledge Graphs. Translating exploration steps into database queries for processing leads to several joins when dealing with knowledge graphs as opposed to filter conditions when dealing with structured data. Further, existing engines handle each exploration step as independent queries in spite of data dependencies that often exist between steps. In this work, we propose an incremental query execution model RAPIDFacet, that exploits the iterative nature of faceted search and reuses intermediate results. The approach is built on top of Apache Spark which naturally supports iterative models and the Nested Triplegroup Data Model and Algebra (NTGA) which uses a coarse grained data model to avoid joins. Evaluations showed up to 150× faster execution than existing approaches. Avimanyu Mukhopadhyay, HyeongSik Kim 0001, Kemafor Anyanwu |
WETICE | 2 |
| 2017 | A semantics-aware storage framework for scalable processing of knowledge graphs on HadoopabstractKnowledge graphs are graph-based data models which employ named nodes and edges to capture differentiation among entities and relationships in richly diverse data collections such as in the biomedical domain. The flexibility of knowledge graphs allows for heterogeneous collections to be linked and integrated in precise ways. However, resulting data models often have irregular structures which are not easy to manage using platforms for structured, schema-first data models like the relational model. To facilitate exchange, inter-operability and reuse of data, standards such as Resource Description Framework (RDF) have been increasingly adopted for representation. Domains such as the biomedical now have large collections of publicly available RDF graphs as well as benchmark workloads. To achieve scalability in data processing, some efforts are being made to build on distributed processing platforms such as Hadoop and Spark. However, while some distributed graph platforms have emerged for certain classes of mining workloads for non-semantic graphs (without typed edges and nodes), knowledge graph processing, which often involves ontological inferencing, continues to be plagued by scalability and efficiency challenges. In this paper, we present the design of a Hadoop-based storage architecture for knowledge graphs that overcomes some of the challenges of big RDF data processing. The rationale of the design strategy is to go beyond the traditional approach of exploiting structural properties of graphs while storing to include exploitation of semantic properties of knowledge graphs. Our system SemStorm is a Hadoop-based indexed, polymorphic, signatured file organization that supports efficient storage of data collections with significant data heterogeneity. Naive storage models for such data place more demands for meta-data management than traditional systems can support. The polymorphic file organization is further coupled with a nested, column-oriented file format to enable discriminatory data access based on queries. A major hallmark of SemStorm is the enabling of semantic-awareness in storage framework. The idea is to exploit the knowledge represented in ontologies that accompany data for optimizing data storage models such as identifying and managing data (sometimes implicit) redundancies. Another major advantage of SemStorm is that it derives optimized storage models for data autonomically, i.e., without user input. Extensive experiments conducted on real-world and synthetic benchmark datasets show that SemStorm is up to 10X faster than existing approaches. HyeongSik Kim 0001, Padmashree Ravindra, Kemafor Anyanwu |
IEEE BigData | 1 |
| 2017 | Type-based Semantic Optimization for Scalable RDF Graph Pattern MatchingabstractScalable query processing relies on early and aggressive determination and pruning of query-irrelevant data. Besides the traditional space-pruning techniques such as indexing, type-based optimizations that exploit integrity constraints defined on the types can be used to rewrite queries into more efficient ones. However, such optimizations are only applicable in strongly-typed data and query models which make it a challenge for semi-structured models such as RDF. Consequently, developing techniques for enabling typebased query optimizations will contribute new insight to improving the scalability of RDF processing systems. HyeongSik Kim 0001, Padmashree Ravindra, Kemafor Anyanwu |
WWW | 1 |
| 2016 | Optimization of Complex SPARQL Analytical Queries
Padmashree Ravindra, HyeongSik Kim 0001, Kemafor Anyanwu |
EDBT | 2 |
| 2015 | Rewriting complex SPARQL analytical queries for efficient cloud-based processingabstractMany emerging Semantic Web applications combine and aggregate data across domains for analysis. Such analytical queries compute aggregates over multiple groupings of data, resulting in query plans with complex grouping-aggregation constraints. In the context of an RDF analytical query, each such grouping maps to a graph pattern subquery with multiple join operations, and related groups often result in overlapping graph patterns within the same query. In this paper, we propose a holistic approach to optimize RDF analytical queries by refactoring queries to achieve shared execution of common subexpressions that enables parallel evaluation of groupings as well as aggregations. Such a rewriting enables shorter execution workflows, particularly beneficial for scale-out processing on distributed Cloud systems with multiple I/O phases. Experiments on real-world and synthetic benchmarks confirm that such a rewriting can achieve more efficient execution plans when compared to relational-style SPARQL query plans executed on popular Cloud systems. Padmashree Ravindra, HyeongSik Kim 0001, Kemafor Anyanwu |
IEEE BigData | 2 |
| 2013 | Scaling concurrency of personalized Semantic search over Large RDF dataabstractRecent keyword search techniques on Semantic Web are moving away from shallow, information retrieval-style approaches that merely find “keyword matches” towards more interpretive approaches that attempt to induce structure from keyword queries. The process of query interpretation is usually guided by structures in data, and schema and is often supported by a graph exploration procedure. However, graph exploration-based interpretive techniques are impractical for multi-tenant scenarios for large databases because separate expensive graph exploration states need to be maintained for different user queries. This leads to significant memory overhead in situations of large numbers of concurrent requests. This limitation could negatively impact the possibility of achieving the ultimate goal of personalizing search. In this paper, we propose a lightweight interpretation approach that employs indexing to improve throughput and concurrency with much less memory overhead. It is also more amenable to distributed or partitioned execution. The approach is implemented in a system called “SKI” and an experimental evaluation of SKI's performance on the DBPedia and Billion Triple Challenge datasets shows orders-of-magnitude performance improvement over existing techniques. Haizhou Fu, HyeongSik Kim 0001, Kemafor Anyanwu |
IEEE BigData | 2 |
| 2013 | Optimizing queries over semantically integrated datasets on MapReduce platformsabstractLife science databases generally consist of multiple heterogeneous datasets that have been integrated using complex ontologies. Querying such databases typically involves complex graph patterns, and evaluating such patterns poses challenges when MapReduce-based platforms are used to scale up processing, translating to long execution workflows with large amount of disk and network I/O costs. In this poster, we focus on optimizing UNION queries (e.g., unions of conjunctives for inference) and present an algebraic interpretation of the query rewritings which are more amenable to efficient processing on MapReduce. HyeongSik Kim 0001, Kemafor Anyanwu |
IEEE BigData | 1 |
| 2012 | Scan-Sharing for Optimizing RDF Graph Pattern Matching on MapReduceabstractRecently, the number and size of RDF data collections has increased rapidly making the issue of scalable processing techniques crucial. The MapReduce model has become a de facto standard for large scale data processing using a cluster of machines in the cloud. Generally, RDF query processing creates join-intensive workloads, resulting in lengthy MapReduce workflows with expensive I/O, data transfer, and sorting costs. However, the MapReduce computation model provides limited static optimization techniques used in relational databases (e.g., indexing and cost-based optimization). Consequently, dynamic optimization techniques for such join-intensive tasks on MapReduce need to be investigated. In some previous efforts, we propose a Nested Triple Group data model and Algebra (NTGA) for efficient graph pattern query processing in the cloud. Here, we extend this work with a scan-sharing technique that is used to optimize the processing of graph patterns with repeated properties. Specifically, our scan-sharing technique eliminates the need for repeated scanning of input relations when properties are used repeatedly in graph patterns. A formal foundation underlying this scan sharing technique is discussed as well as an implementation strategy that has been integrated in the Apache Pig framework is presented. We also present a comprehensive evaluation demonstrating performance benefits of our NTGA plus scan-sharing approach. HyeongSik Kim 0001, Padmashree Ravindra, Kemafor Anyanwu |
IEEE CLOUD | 1 |
| 2011 | An Intermediate Algebra for Optimizing RDF Graph Pattern Matching on MapReduce
Padmashree Ravindra, HyeongSik Kim 0001, Kemafor Anyanwu |
ESWC (2) | 2 |
| 2011 | From SPARQL to MapReduce: The Journey Using a Nested TripleGroup Algebra
HyeongSik Kim 0001, Padmashree Ravindra, Kemafor Anyanwu |
Proc. VLDB Endow. | 1 |