EDBT 2026 Demo / reviewers in the wild / expert
Kalpa Gunaratna
dblp:00/9080
· DBLP profile ↗
14ranked-venue papers
6as first author
5since 2021 · last 2024
0000-0002-9058-4577ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 4 first-author · 4 since 2021Databases, data management, data science and information retrieval · 7 · 3 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | AlpaGasus: Training a Better Alpaca with Fewer DataabstractLarge language models~(LLMs) strengthen instruction-following capability through instruction-finetuning (IFT) on supervised instruction/response data. However, widely used IFT datasets (e.g., Alpaca's 52k data) surprisingly contain many low-quality instances with incorrect or irrelevant responses, which are misleading and detrimental to IFT. In this paper, we propose a simple and effective data selection strategy that automatically identifies and removes low-quality data using a strong LLM (e.g., ChatGPT). To this end, we introduce Alpagasus, which is finetuned on only 9k high-quality data filtered from the 52k Alpaca data. Alpagasus significantly outperforms the original Alpaca as evaluated by GPT-4 on multiple test sets and the controlled human study. Its 13B variant matches $>90\%$ performance of its teacher LLM (i.e., Text-Davinci-003) on test tasks. It also provides 5.7x faster training, reducing the training time for a 7B variant from 80 minutes (for Alpaca) to 14 minutes \footnote{We apply IFT for the same number of epochs as Alpaca(7B) but on fewer data, using 4$\times$NVIDIA A100 (80GB) GPUs and following the original Alpaca setting and hyperparameters.}. In the experiment, we also demonstrate that our method can work not only for machine-generated datasets but also for human-written datasets. Overall, Alpagasus demonstrates a novel data-centric IFT paradigm that can be generally applied to instruction-tuning data, leading to faster training and better instruction-following models. Lichang Chen, Kalpa Gunaratna, Vikas Yadav, Vijay Srinivasan, Tianyi Zhou 0001, Heng Huang 0001, Hongxia Jin |
ICLR | 5 |
| 2023 | Explainable and Accurate Natural Language Understanding for Voice Assistants and BeyondabstractJoint intent detection and slot filling, which is also termed as joint NLU (Natural Language Understanding) is invaluable for smart voice assistants. Recent advancements in this area have been heavily focusing on improving accuracy using various techniques. Explainability is undoubtedly an important aspect for deep learning-based models including joint NLU models. Without explainability, their decisions are opaque to the outside world and hence, have tendency to lack user trust. Therefore to bridge this gap, we transform the full joint NLU model to be 'inherently' explainable at granular levels without compromising on accuracy. Further, as we enable the full joint NLU model explainable, we show that our extension can be successfully used in other general classification tasks. We demonstrate this using sentiment analysis and named entity recognition. Kalpa Gunaratna, Vijay Srinivasan, Hongxia Jin |
CIKM | 1 |
| 2022 | ISEEQ: Information Seeking Question Generation Using Dynamic Meta-Information Retrieval and Knowledge GraphsabstractConversational Information Seeking (CIS) is a relatively new research area within conversational AI that attempts to seek information from end-users in order to understand and satisfy the users' needs. If realized, such a CIS system has far-reaching benefits in the real world; for example, CIS systems can assist clinicians in pre-screening or triaging patients in healthcare. A key open sub-problem in CIS that remains unaddressed in the literature is generating Information Seeking Questions (ISQs) based on a short initial query from the end-user. To address this open problem, we propose Information SEEking Question generator (ISEEQ), a novel approach for generating ISQs from just a short user query, given a large text corpus relevant to the user query. Firstly, ISEEQ uses a knowledge graph to enrich the user query. Secondly, ISEEQ uses the knowledge-enriched query to retrieve relevant context passages to ask coherent ISQs adhering to a conceptual flow. Thirdly, ISEEQ introduces a new deep generative-adversarial reinforcement learning-based approach for generating ISQs. We show that ISEEQ can generate high-quality ISQs to promote the development of CIS agents. ISEEQ significantly outperforms comparable baselines on five ISQ evaluation metrics across four datasets having user queries from diverse domains. Further, we argue that ISEEQ is transferable across domains for generating ISQs, as it shows the acceptable performance when trained and tested on different pairs of domains. A qualitative human evaluation confirms that ISEEQ generated ISQs are comparable in quality to human-generated questions, and it outperformed the best comparable baseline. Manas Gaur, Kalpa Gunaratna, Vijay Srinivasan, Hongxia Jin |
AAAI | 2 |
| 2021 | Using Neighborhood Context to Improve Information Extraction from Visual Documents Captured on Mobile PhonesabstractInformation Extraction from visual documents enables convenient and intelligent assistance to end users. We present a Neighborhood-based Information Extraction (NIE) approach that uses contextual language models and pays attention to the local neighborhood context in the visual documents to improve information extraction accuracy. We collect two different visual document datasets and show that our approach outperforms the state-of-the-art global context-based IE technique. In fact, NIE outperforms existing approaches in both small and large model sizes. Our on-device implementation of NIE on a mobile platform that generally requires small models showcases NIE's usefulness in practical real-world applications. Kalpa Gunaratna, Vijay Srinivasan, Sandeep Nama, Hongxia Jin |
CIKM | 1 |
| 2021 | Entity summarization: State of the art and future challenges
Qingxia Liu, Gong Cheng 0001, Kalpa Gunaratna, Yuzhong Qu |
J. Web Semant. | 3 |
| 2020 | 3rd International Workshop on EntitY Retrieval and lEarning (EYRE 2020)abstractEntity retrieval has received increasing research attention. The recent progress in deep and machine learning techniques provides powerful tools for developing effective entity-centered solutions. This workshop series provides a platform where interdisciplinary studies of entity retrieval and learning can be presented, and focused discussions can take place. We also organize a shared task related to entity retrieval. The 3rd International Workshop on EntitY Retrieval and lEarning (EYRE 2020) was a half-day workshop co-located with the 29th ACM International Conference on Information and Knowledge Management (CIKM 2020) as a virtual event in Ireland. Gong Cheng 0001, Kalpa Gunaratna, Jun Wang 0012 |
CIKM | 2 |
| 2020 | ESBM: An Entity Summarization BenchMark
Qingxia Liu, Gong Cheng 0001, Kalpa Gunaratna, Yuzhong Qu |
ESWC | 3 |
| 2020 | Neural Entity Summarization with Joint Encoding and Weak SupervisionabstractIn a large-scale knowledge graph (KG), an entity is often described by a large number of triple-structured facts. Many applications require abridged versions of entity descriptions, called entity summaries. Existing solutions to entity summarization are mainly unsupervised. In this paper, we present a supervised approach NEST that is based on our novel neural model to jointly encode graph structure and text in KGs and generate high-quality diversified summaries. Since it is costly to obtain manually labeled summaries for training, our supervision is weak as we train with programmatically labeled data which may contain noise but is free of manual work. Evaluation results show that our approach significantly outperforms the state of the art on two public benchmarks. Junyou Li, Gong Cheng 0001, Qingxia Liu, Wen Zhang 0015, Evgeny Kharlamov, Kalpa Gunaratna, Huajun Chen |
IJCAI | 6 |
| 2020 | Enriching Documents with Compact, Representative, Relevant Knowledge GraphsabstractA prominent application of knowledge graph (KG) is document enrichment. Existing methods identify mentions of entities in a background KG and enrich documents with entity types and direct relations. We compute an entity relation subgraph (ERG) that can more expressively represent indirect relations among a set of mentioned entities. To find compact, representative, and relevant ERGs for effective enrichment, we propose an efficient best-first search algorithm to solve a new combinatorial optimization problem that achieves a trade-off between representativeness and compactness, and then we exploit ontological knowledge to rank ERGs by entity-based document-KG and intra-KG relevance. Extensive experiments and user studies show the promising performance of our approach. Zixian Huang, Gong Cheng 0001, Evgeny Kharlamov, Kalpa Gunaratna |
IJCAI | 5 |
| 2019 | EYRE 2019: 2nd International Workshop on EntitY REtrievalabstractEntity retrieval has received increasing research attention from both the Information Retrieval (IR) and Semantic Web communities. This workshop series provides a platform where interdisciplinary studies of entity retrieval can be presented, and focused discussions can take place. We also organize two shared tasks related to entity retrieval. The 2nd International Workshop on EntitY REtrieval (EYRE 2019) was a half-day workshop co-located with the 28th ACM International Conference on Information and Knowledge Management (CIKM 2019) in Beijing, China. Gong Cheng 0001, Kalpa Gunaratna, Jun Wang 0012 |
CIKM | 2 |
| 2017 | Relatedness-based Multi-Entity SummarizationabstractRepresenting world knowledge in a machine processable format is important as entities and their descriptions have fueled tremendous growth in knowledge-rich information processing platforms, services, and systems. Prominent applications of knowledge graphs include search engines (e.g., Google Search and Microsoft Bing), email clients (e.g., Gmail), and intelligent personal assistants (e.g., Google Now, Amazon Echo, and Apple's Siri). In this paper, we present an approach that can summarize facts about a collection of entities by analyzing their relatedness in preference to summarizing each entity in isolation. Specifically, we generate informative entity summaries by selecting: (i) inter-entity facts that are similar and (ii) intra-entity facts that are important and diverse. We employ a constrained knapsack problem solving approach to efficiently compute entity summaries. We perform both qualitative and quantitative experiments and demonstrate that our approach yields promising results compared to two other stand-alone state-of-the-art entity summarization approaches. Kalpa Gunaratna, Amir Hossein Yazdavar, Krishnaprasad Thirunarayan, Amit P. Sheth, Gong Cheng 0001 |
IJCAI | 1 |
| 2016 | Gleaning Types for Literals in RDF Triples with Application to Entity Summarization
Kalpa Gunaratna, Krishnaprasad Thirunarayan, Amit P. Sheth, Gong Cheng 0001 |
ESWC | 1 |
| 2015 | FACES: Diversity-Aware Entity Summarization Using Incremental Hierarchical Conceptual ClusteringabstractSemantic Web documents that encode facts about entities on the Web have been growing rapidly in size and evolving over time. Creating summaries on lengthy Semantic Web documents for quick identification of the corresponding entity has been of great contemporary interest. In this paper, we explore automatic summarization techniques that characterize and enable identification of an entity and create summaries that are human friendly. Specifically, we highlight the importance of diversified (faceted) summaries by combining three dimensions: diversity, uniqueness, and popularity. Our novel diversity-aware entity summarization approach mimics human conceptual clustering techniques to group facts and picks representative facts from each group to form concise (i.e., short) and comprehensive (i.e., improved coverage through diversity) summaries. We evaluate our approach against the state-of-the-art techniques and show that our work improves both the quality and the efficiency of entity summarization. Kalpa Gunaratna, Krishnaprasad Thirunarayan, Amit P. Sheth |
AAAI | 1 |
| 2010 | A Study in Hadoop Streaming with Matlab for NMR Data ProcessingabstractApplying Cloud computing techniques for analyzing large data sets has shown promise in many data-driven scientific applications. Our approach presented here is to use Cloud computing for Nuclear Magnetic Resonance (NMR)data analysis which normally consists of large amounts of data. Biologists often use third party or commercial software for ease of use. Enabling the capability to use this kind of software in a Cloud will be highly advantageous in many ways. Scripting languages especially designed for clouds may not have the flexibility biologists need for their purposes. Although this is true, they are familiar with special software packages that allow them to write complex calculations with minimum effort, but are often not compatible with a Cloud environment. Therefore, biologists who are trying to perform analysis on NMR data, acquire many advantages due to our proposed solution. Our solution gives them the flexibility to Cloud-enable their familiar software and it also enables them to perform calculations on a significant amount of data that was not previously possible. Our study is also applicable to any other environment in need of similar flexibility. We are currently in the initial stage of developing a framework for NMR data analysis. Kalpa Gunaratna, Ajith Ranabahu, Amit P. Sheth |
CloudCom | 1 |