Jing Ao

dblp:75/8496 · DBLP profile ↗
← Back
4ranked-venue papers in the field
2as first author
3since 2021 · last 2023
0009-0000-4104-1368ORCID · corroborated

Domains — venue-derived; a paper can count in several

Big Data, Cloud & Distributed Data Systems · 3 (2 first)Information Retrieval & Web Search · 1
YearPublicationVenuePosition
2023 Provenance-Aware Data Integration and Summarization Querying for Knowledge Graphs
Pei-Yu Hou, Jing Ao, Kara Schatz, Alexey V. Gulyuk, Yaroslava G. Yingling, Rada Chirkova
iiWAS2
2022 KGIQ: Scalable Translation of User-Specified Examples into Knowledge-Graph Queries
abstract
Querying large-scale knowledge graphs (KGs) can be difficult for users in real-life scenarios, in which formal graph query languages potentially present usability barriers. Query-By-Example (QBE) approaches, which allow users to specify their query intent with examples, have become an emerging trend to address this issue. However, existing QBE approaches either require user-specified examples to provide values of all the KG attributes, or may return incorrect formalizations of user query intent due to the lack of user interaction. In this paper we propose an approach called Knowledge Graph Intuitive Querier (KGIQ) that addresses both challenges on large-scale KGs by using novel scalable algorithms. Unlike existing approaches, KGIQ ensures correct translation of user-specified examples into formal executable graph queries by interacting with users in a lightweight manner. Our experimental results suggest that KGIQ can correctly formalize graph queries on large-scale KGs with the help of uncomplicated user interactions and a feedback loop to improve user experience, while consistently outperforming the state of the art in terms of efficiency and outcome quality.
Jing Ao, Rada Chirkova
IEEE Big Data1
2021 Trustworthy Knowledge Graph Population From Texts for Domain Query Answering
abstract
Obtaining answers to domain-specific questions over large-scale unstructured (text) data is an important component of data analytics in many application domains. As manual question answering does not scale to large text corpora, it is common to use information extraction (IE) to preprocess the texts of interest prior to posing the questions. This is often done by transforming text corpora into the knowledge-graph (KG) triple format that is suitable for efficient processing of the user questions in graph-oriented data-intensive systems.In a number of real-life scenarios, trustworthiness of the answers obtained from domain-specific texts is vital for downstream decision making. In this paper we focus on one critical aspect of trustworthiness, which concerns aligning with the given domain vocabularies (ontologies) those KG triples that are obtained from the source texts via IE solutions. To address this problem, we introduce a scalable domain-independent text-to-KG approach that adapts to specific domains by using domain ontologies, without having to consult external triple repositories. Our IE solution builds on the power of neural-based learning models and leverages feature engineering to distinguish ontology-aligned data from generic data in the source texts. Our experimental results indicate that the proposed approach could be more dependable than a state-of-the-art IE baseline in constructing KGs that are suitable for trustworthy domain question answering on text data.
Jing Ao, Swathi Dinakaran, Hongjian Yang, David R. Wright 0001, Rada Chirkova
IEEE BigData1
2019 Collaborative Workflow for Analyzing Large-Scale Data for Antimicrobial Resistance: An Experience Report
abstract
In real-life analytics-oriented information-integration projects, the processes of information curation and integration cannot be completely automated. Rather, in each large-scale project the key objectives include maximizing scalability and throughput, while at the same time keeping the processes manageable and productive for the human experts in the loop. In this paper, we describe our experience with addressing these major objectives in the process of building a scalable end-to-end data-extraction, integration, and analytics workflow in the domain of antimicrobial resistance (AMR). The workflow is built using open-source tools, with the aims of enhancing the efficiency and accuracy of data collection and integration, while involving an acceptable level of efforts by collaborative multidisciplinary teams of humans-in-the-loop. We present the components of the proposed workflow, outline the challenges encountered in its development and testing, and discuss the experiences and lessons learned in enabling AMR experts and data analysts to interact with the workflow, with some of the lessons potentially applicable to other application domains.
Pei-Yu Hou, Jing Ao, Andrew J. Rindos, Shivaramu Keelara, Paula J. Fedorka-Cray, Rada Chirkova
IEEE BigData2