Koninika Pal

dblp:151/6417 · DBLP profile ↗
← Back
12ranked-venue papers
4as first author
4since 2021 · last 2025
0000-0002-8829-5661ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 12 · 4 first-author · 4 since 2021Artificial intelligence and machine learning · 3 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 NuFact: Validating Numerical Assertions for Knowledge Graphs
abstract
Validating extracted assertions is one of the crucial steps for curating knowledge graphs (KGs).While existing research extensively explores methods for validating KG assertions, specifically for the categorical facts -where both the Subject and Object in triples are entities, there remains a significant gap in the validation of numerical assertions, where the Object represents a quantity.Moreover, general fact-validation methods are inefficient for validating numerical claims due to their limited coverage in KGs.Furthermore, large-language models (LLMs) exhibit limitations in quantitative reasoning, further exacerbating the challenge.These gaps compromise the reliability of KGs in applications that require precise numerical accuracy.Addressing these challenges, we propose Nu-Fact, a framework for validating numerical assertions using evidence gathered from the web.NuFact combines the rich contextual understanding of LLMs with manually crafted quantity-focused and temporal features derived from extracted evidence to assess numerical claims.Experimental evaluations show that NuFact significantly outperforms existing fact-checking baselines and popular LLM-powered agents.
Mohammad Taufeeq, Koninika Pal
CIKM2
2025 Understanding Numerical Context by Asking Quantitative Questions
Koninika Pal
ECIR (5)2
2021 QuTE: Answering Quantity Queries from Web Tables
abstract
Quantities are financial, technological, physical and other measures that denote relevant properties of entities, such as revenue of companies, energy efficiency of cars or distance and brightness of stars and galaxies. Queries with filter conditions on quantities are an important building block for downstream analytics, and pose challenges when the content of interest is spread across a huge number of web tables and other ad-hoc datasets. Search engines support quantity lookups, but largely fail on quantity filters. The QuTE system presented in this paper aims to overcome these problems. It comprises methods for automatically extracting entity-quantity facts from web tables, as well as methods for online query processing, with new techniques for query matching and answer ranking.
Vinh Thinh Ho, Koninika Pal, Gerhard Weikum
SIGMOD Conference2
2021 Extracting Contextualized Quantity Facts from Web Tables
abstract
Quantity queries, with filter conditions on quantitative measures of entities, are beyond the functionality of search engines and QA assistants. To enable such queries over web contents, this paper develops a novel method for automatically extracting quantity facts from ad-hoc web tables. This involves recognizing quantities, with normalized values and units, aligning them with the proper entities, and contextualizing these pairs with informative cues to match sophisticated queries with modifiers. Our method includes a new approach to aligning quantity columns to entity columns. Prior works assumed a single subject-column per table, whereas our approach is geared for complex tables and leverages external corpora as evidence. For contextualization, we identify informative cues from text and structural markup that surrounds a table. For query-time fact ranking, we devise a new scoring technique that exploits both context similarity and inter-fact consistency. Comparisons of our building blocks against state-of-the-art baselines and extrinsic experiments with two query benchmarks demonstrate the benefits of our method.
Vinh Thinh Ho, Koninika Pal, Simon Razniewski, Klaus Berberich, Gerhard Weikum
WWW2
2020 Entities with Quantities: Extraction, Search, and Ranking
abstract
Quantities are more than numeric values. They represent measures for entities, expressed in numbers with associated units. Search queries often include quantities, such as athletes who ran 200m under 20 seconds or companies with quarterly revenue above $2 Billion. Processing such queries requires understanding the quantities, where capturing the surrounding context is an essential part of it. Although modern search engines or QA systems handle entity-centric queries well, they consider numbers and units as simple keywords, and therefore fail to understand the condition (less than, above, etc.), the unit of interest (seconds, dollar, etc.), and the context of the quantity (200m race, quarterly revenue, etc.) As a result, they cannot generate the correct candidate answers. In this work, we demonstrate a prototype QA system, called Qsearch, that can handle advanced queries with quantity constraints using the common cues present in both query and the text sources.
Vinh Thinh Ho, Koninika Pal, Niko Kleer, Klaus Berberich, Gerhard Weikum
WSDM2
2019 Commonsense Properties from Query Logs and Question Answering Forums
abstract
Commonsense knowledge about object properties, human behavior and general concepts is crucial for robust AI applications. However, automatic acquisition of this knowledge is challenging because of sparseness and bias in online sources. This paper presents Quasimodo, a methodology and tool suite for distilling commonsense properties from non-standard web sources. We devise novel ways of tapping into search-engine query logs and QA forums, and combining the resulting candidate assertions with statistical cues from encyclopedias, books and image tags in a corroboration step. Unlike prior work on commonsense knowledge bases, Quasimodo focuses on salient properties that are typically associated with certain objects or concepts. Extensive evaluations, including extrinsic use-case studies, show that Quasimodo provides better coverage than state-of-the-art baselines with comparable quality.
Julien Romero, Simon Razniewski, Koninika Pal, Jeff Z. Pan, Archit Sakhadeo, Gerhard Weikum
CIKM3
2019 Qsearch: Answering Quantity Queries from Text
Vinh Thinh Ho, Yusra Ibrahim, Koninika Pal, Klaus Berberich, Gerhard Weikum
ISWC (1)3
2018 Learning interesting attributes for automated data categorization
abstract
This work proposes and evaluates a novel approach to determining interesting attributes, in order to categorize entities accordingly. Once identified, such categories are of immense value to allow constraining (filtering) a user's current view to subsets of entities. We show how a classifier is trained that is able to tell whether or not a categorical attribute can act as a constraint, in the sense of human-perceived interestingness. The training data is harnessed from Wikipedia tables, treating the presence or absence of a table as an indication that the attribute used as a filter constraint is reasonable or not. For learning the classification model, we review four well-known statistical measures (features) for categorical attributes---entropy, unalikeability, peculiarity, and coverage. We additionally propose three new statistical measures to capture the distribution of data, tailored to our main objective. The learned model is evaluated by relevance assessments obtained through a user study, reflecting the applicability of the approach as a whole and, further, demonstrates the superiority of the proposed diversity measures over existing measures like information entropy.
Koninika Pal, Sebastian Michel 0001
SSDBM1
2017 LSH-Based Probabilistic Pruning of Inverted Indices for Sets and Ranked Lists
abstract
We address the problem of index pruning without compromising the quality of ad-hoc similarity search among sets and ranked lists. We discuss three different ways to prune the index structure and, by linking the index structure with the concept of Locality Sensitive Hashing (LSH), we introduce two solutions to query processing over the pruned index. Through a probabilistic analysis we ensure that a user-defined recall goal is still guaranteed. We are able to formulate an optimization problem that can determine the optimal pruning factor for all three pruning methods. The experimental evaluations over real-world data validate that the optimal pruning factor indeed ensures the recall goal without any significant effect on the quality of similarity search on a much smaller index.
Koninika Pal, Sebastian Michel 0001
WebDB1
2016 A Data Mining Approach to Choosing Categorical Attributes for Ranked Lists
Koninika Pal, Sebastian Michel 0001
EDBT1
2016 Efficient Similarity Search across Top-k Lists under the Kendall's Tau Distance
abstract
We consider the problem of similarity search in a set of top-k lists under the generalized Kendall's Tau distance. This distance describes how related two rankings are in terms of discordantly ordered items. We consider pair- and triplets-based indices to counter the shortcomings of naive inverted indices and derive efficient query schemes by relating the proposed index structures to the concept of locality sensitive hashing (LSH). Specifically, we devise four different LSH schemes for Kendall's Tau using two generic hash families over individual elements or pairs of them. We show that each of these functions has the desired property of being locality sensitive. Further, we discuss the selection of hash functions for the proposed LSH schemes for a given query ranking, called query-driven LSH and derive bounds for the required number of hash functions to use in order to achieve a predefined recall goal. Experimental results, using two real-world datasets, show that the devised methods outperform the SimJoin method---the state of the art method to query for similar sets---and are far superior to a plain inverted-index--based approach.
Koninika Pal, Sebastian Michel 0001
SSDBM1
2016 Exploring Databases via Reverse Engineering Ranking Queries with PALEO
abstract
A novel approach to explore databases using ranked lists is demonstrated. Working with ranked lists, capturing the relative performance of entities, is a very intuitive and widely applicable concept. Users can post lists of entities for which explanatory SQL queries and full result lists are returned. By refining the input, the results, or the queries, user can interactively explore the database content. The demonstrated system is centered around our PALEO framework for reverse engineering OLAP-style database queries and novel work on mining interesting categorical attributes.
Kiril Panev, Sebastian Michel 0001, Evica Milchevski, Koninika Pal
Proc. VLDB Endow.4