Vaishali Pal

dblp:258/0544 · DBLP profile ↗
← Back
6ranked-venue papers
4as first author
5since 2021 · last 2024
0000-0002-1493-3659ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 3 first-author · 3 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Question answering and dialogue systems · 73% Language models and text generation · 27%
Databases, data mining, and information retrieval
1 paper
Information retrieval · 100%

Topics — the 4 heaviest of 5, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Question answering and dialogue systems
table question answering
1.422024
Table Question Answering for Low-resourced Indic Languages · EMNLP 2024
MultiTabQA: Generating Tabular Answers for Multi-Table Question Answering · ACL (1) 2023
Natural language and speech › Language models and text generation
low-resource languages
0.812024
Table Question Answering for Low-resourced Indic Languages · EMNLP 2024
Natural language and speech › Question answering and dialogue systems › table question answering
multi-table question answering
0.712023
MultiTabQA: Generating Tabular Answers for Multi-Table Question Answering · ACL (1) 2023
Information retrieval › interactive information retrieval › conversational information seeking
conversational search
0.612022
The Seventh Workshop on Search-Oriented Conversational Artificial Intelligence (SCAI'22) · SIGIR 2022

Methods — techniques the papers use, named apart from their topics

mathematical reasoning · 0.8large language model · 0.8cross-lingual transfer · 0.8pre-training · 0.7fine-tuning · 0.7natural language processing · 0.6machine learning · 0.6
YearPublicationVenuePosition
2024 QFMTS: Generating Query-Focused Summaries over Multi-Table Inputs
abstract
Table summarization is a crucial task aimed at condensing information from tabular data into concise and comprehensible textual summaries. However, existing approaches often fall short of adequately meeting users’ information and quality requirements and tend to overlook the complexities of real-world queries. In this paper, we propose a novel method to address these limitations by introducing query-focused multi-table summarization. Our approach, which comprises a table serialization module, a summarization controller, and a large language model (LLM), utilizes textual queries and multiple tables to generate query-dependent table summaries tailored to users’ information needs. To facilitate research in this area, we present a comprehensive dataset specifically tailored for this task, consisting of 4,909 query-summary pairs, each associated with multiple tables. Through extensive experiments using our curated dataset, we demonstrate the effectiveness of our proposed method compared to baseline approaches. Our findings offer insights into the challenges of complex table reasoning for precise summarization, contributing to the advancement of research in query-focused multi-table summarization.
Weijia Zhang 0004, Vaishali Pal, Jia-Hong Huang, Evangelos Kanoulas, Maarten de Rijke
ECAI2
2024 Table Question Answering for Low-resourced Indic Languages
abstract
TableQA is the task of answering questions over tables of structured information, returning individual cells or tables as output.TableQA research has focused primarily on high-resource languages, leaving medium-and low-resource languages with little progress due to scarcity of annotated data and neural models.We address this gap by introducing a fully automatic large-scale table question answering (tableQA) data generation process for low-resource languages with limited budget.We incorporate our data generation method on two Indic languages, Bengali and Hindi, which have no tableQA datasets or models.TableQA models trained on our large-scale datasets outperform stateof-the-art LLMs.We further study the trained models on different aspects, including mathematical reasoning capabilities and zero-shot cross-lingual transfer.Our work is the first on low-resource tableQA focusing on scalable data generation and evaluation procedures.Our proposed data generation method can be applied to any low-resource language with a web presence.We release datasets, models, and code.
Vaishali Pal, Evangelos Kanoulas, Andrew Yates, Maarten de Rijke
EMNLP1
2023 MultiTabQA: Generating Tabular Answers for Multi-Table Question Answering
abstract
Recent advances in tabular question answering (QA) with large language models are constrained in their coverage and only answer questions over a single table.However, real-world queries are complex in nature, often over multiple tables in a relational database or web page.Single table questions do not involve common table operations such as set operations, Cartesian products (joins), or nested queries.Furthermore, multi-table operations often result in a tabular output, which necessitates table generation capabilities of tabular QA models.To fill this gap, we propose a new task of answering questions over multiple tables.Our model, MultiTabQA, not only answers questions over multiple tables, but also generalizes to generate tabular answers.To enable effective training, we build a pre-training dataset comprising of 132,645 SQL queries and tabular answers.Further, we evaluate the generated tables by introducing table-specific metrics of varying strictness assessing various levels of granularity of the table structure.MultiTabQA outperforms state-of-the-art single table QA models adapted to a multi-table QA setting by finetuning on three datasets: Spider, Atis and GeoQuery.
Vaishali Pal, Andrew Yates, Evangelos Kanoulas, Maarten de Rijke
ACL (1)1
2023 Parameter-Efficient Sparse Retrievers and Rerankers Using Adapters
Vaishali Pal, Carlos Eduardo Rosar Kós Lassance, Hervé Déjean, Stéphane Clinchant
ECIR (2)1
2022 The Seventh Workshop on Search-Oriented Conversational Artificial Intelligence (SCAI'22)
abstract
The goal of the seventh edition of SCAI (https://scai.info) is to bring together and further grow a community of researchers and practitioners interested in conversational systems for information access. The previous iterations of the workshop already demonstrated the breadth and multidisciplinarity inherent in the design and development of conversational search agents. The proposed shift from traditional web search to search interfaces enabled via human-like dialogue leads to a number of challenges, and although such challenges have received more attention in the recent years, there are many pending research questions that should be addressed by the information retrieval community and can largely benefit from a collaboration with other research fields, such as natural language processing, machine learning, human-computer interaction and dialogue systems. This workshop is intended as a platform enabling a continuous discussion of the major research challenges that surround the design of search-oriented conversational systems. This year, participants have the opportunity to meet in person and have more in-depth interactive discussions with a full-day onsite workshop.
Gustavo Penha, Svitlana Vakulenko, Ondrej Dusek, Leigh Clark, Vaishali Pal, Vaibhav Adlakha
SIGIR5
2020 Modeling ASR Ambiguity for Neural Dialogue State Tracking
abstract
Spoken dialogue systems typically use a list of top-N ASR hypotheses for inferring the semantic meaning and tracking the state of the dialogue. However ASR graphs, such as confusion networks (confnets), provide a compact representation of a richer hypothesis space than a top-N ASR list. In this paper, we study the benefits of using confusion networks with a state-of-the-art neural dialogue state tracker (DST). We encode the 2-dimensional confnet into a 1-dimensional sequence of embeddings using an attentional confusion network encoder which can be used with any DST system. Our confnet encoder is plugged into the state-of-the-art 'Global-locally Self-Attentive Dialogue State Tacker' (GLAD) model for DST and obtains significant improvements in both accuracy and inference time compared to using top-N ASR hypotheses.
Vaishali Pal, Fabien Guillot, Manish Shrivastava 0001, Jean-Michel Renders, Laurent Besacier
INTERSPEECH1