VLDB 2026 Research / reviewers in the wild / expert
Pranava Swaroop Madhyastha
dblp:151/8501 · also Pranava Madhyastha
· DBLP profile ↗
31ranked-venue papers
8as first author
15since 2021 · last 2026
0000-0002-4438-8161ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 27 · 8 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | "Although Powerful, it's not Infallible": Investigating Academic Researchers' Verification Challenges with LLMsabstractLLMs have great potential for shaping how people find and understand information. However, current tools can struggle to provide authoritative sources, fabricate plausible references, and present obstacles to assessing truthfulness of their outputs. Understanding how users verify LLM outputs is particularly important in scholarly disciplines where information produced becomes the foundation of future knowledge. We investigated the factors that influence academic researchers’ decisions to verify LLM responses, their verification strategies, and the effectiveness of those strategies. We conducted a naturalistic think-aloud study, followed by a semi-structured interview, where we observed 16 researchers across disciplines using LLMs of their choice to conduct a research information-seeking task. Our findings highlight that prevailing LLM design can hamper users’ ability to satisfy their information needs for several reasons, such as lack of transparency about sources used in LLM outputs and lack of faithfulness of LLM outputs to the source. Based on these findings, we discuss how future LLMs can better support users in effective verification. Monica Visani Scozzi, Stephann Makri, Pranava Swaroop Madhyastha |
CHIIR | 3 |
| 2025 | Foundation model assisted visual analytics: Opportunities and ChallengesabstractWe explore the integration of foundation models, such as large language models (LLMs) and multimodal LLMs (MLLMs), into visual analytics (VA) systems through intuitive natural language interactions. We survey current research directions in this emerging field, examining how foundation models have already been integrated into key visualisation-related processes in VA: visual mapping, the creation of data visualisations; visualisation observation, the process of generating a finding through visualisation; and visualisation manipulation, changing the viewport or highlighting areas of interest within a visualisation. We also highlight new possibilities that foundation models bring to VA, in particular, the opportunities to use MLLMs to interpret visualisations directly, to integrate multimodal interactions, and to provide guidance to users. We finally conclude with a vision of future VA systems as collaborative partners in analysis and address the prominent challenges in realising this vision through foundation models. Our discussions in this paper aim to guide future researchers working on foundation model assisted VA systems and help them navigate common obstacles when developing these systems. Maeve Hutchinson, Radu Jianu, Aidan Slingsby, Pranava Swaroop Madhyastha |
Comput. Graph. | 4 |
| 2024 | Towards Holistic, Pragmatic and Multimodal Conversational SystemsabstractLanguage acquisition and utilization transcend the mere exchange of lexical units. Visual cues, prosody, gestures, body movements, and context play an undeniably crucial role. Humans naturally communicate multimodally, employing multiple channels and synthesizing information from diverse modalities. My research delves into the characterization and construction of multimodal models that seamlessly integrate data from multiple independent modalities. I will cover recent work that highlights the challenges, achievements, and opportunities towards developing capable multimodal discursive models. Pranava Swaroop Madhyastha |
AAAI | 1 |
| 2023 | Are words equally surprising in audio and audio-visual comprehension?
Pranava Swaroop Madhyastha, Gabriella Vigliocco |
CogSci | 1 |
| 2023 | Towards preserving word order importance through Forced InvalidationabstractLarge pre-trained language models such as BERT have been widely used as a framework for natural language understanding (NLU) tasks.However, recent findings have revealed that pre-trained language models are insensitive to word order.The performance on NLU tasks remains unchanged even after randomly permuting the word of a sentence, where crucial syntactic information is destroyed.To help preserve the importance of word order, we propose a simple approach called FORCED INVALIDA-TION (FI): forcing the model to identify permuted sequences as invalid samples.We perform an extensive evaluation of our approach on various English NLU and QA based tasks over BERT-based and attention-based models over word embeddings.Our experiments demonstrate that FI significantly improves the sensitivity of the models to word order. 1 Hadeel Al-Negheimish, Pranava Swaroop Madhyastha, Alessandra Russo |
EACL | 2 |
| 2023 | A study towards contextual understanding of toxicity in online conversationsabstractAbstract Identifying and annotating toxic online content on social media platforms is an extremely challenging problem. Work that studies toxicity in online content has predominantly focused on comments as independent entities. However, comments on social media are inherently conversational, and therefore, understanding and judging the comments fundamentally requires access to the context in which they are made. We introduce a study and resulting annotated dataset where we devise a number of controlled experiments on the importance of context and other observable confounders – namely gender, age and political orientation – towards the perception of toxicity in online content. Our analysis clearly shows the significance of context and the effect of observable confounders on annotations. Namely, we observe that the ratio of toxic to non-toxic judgements can be very different for each control group, and a higher proportion of samples are judged toxic in the presence of contextual information. Pranava Swaroop Madhyastha, Antigoni Founta, Lucia Specia |
Nat. Lang. Eng. | 1 |
| 2022 | Belief Revision Based Caption Re-ranker with Visual Semantic InformationabstractIn this work, we focus on improving the captions generated by image-caption generation systems. We propose a novel re-ranking approach that leverages visual-semantic measures to identify the ideal caption that maximally captures the visual information in the image. Our re-ranker utilizes the Belief Revision framework (Blok et al., 2003) to calibrate the original likelihood of the top-n captions by explicitly exploiting semantic relatedness between the depicted caption and the visual context. Our experiments demonstrate the utility of our approach, where we observe that our re-ranker can enhance the performance of a typical image-captioning system without necessity of any additional training or fine-tuning. Ahmed Sabir, Francesc Moreno-Noguer, Pranava Swaroop Madhyastha, Lluís Padró 0001 |
COLING | 3 |
| 2022 | Evaluation of Fake News Detection with Knowledge-Enhanced Language Models
Chenxi Whitehouse, Tillman Weyde, Pranava Swaroop Madhyastha, Nikos Komninos |
ICWSM | 3 |
| 2021 | BERTGen: Multi-task Generation through BERTabstractFaidon Mitzalis, Ozan Caglayan, Pranava Madhyastha, Lucia Specia. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Faidon Mitzalis, Ozan Caglayan, Pranava Swaroop Madhyastha, Lucia Specia |
ACL/IJCNLP (1) | 3 |
| 2021 | Cross-lingual Visual Pre-training for Multimodal Machine TranslationabstractOzan Caglayan, Menekse Kuyu, Mustafa Sercan Amac, Pranava Madhyastha, Erkut Erdem, Aykut Erdem, Lucia Specia. Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. 2021. Ozan Caglayan, Menekse Kuyu, Mustafa Sercan Amac, Pranava Swaroop Madhyastha, Erkut Erdem, Aykut Erdem, Lucia Specia |
EACL | 4 |
| 2021 | Exploiting Multimodal Reinforcement Learning for Simultaneous Machine TranslationabstractJulia Ive, Andy Mingren Li, Yishu Miao, Ozan Caglayan, Pranava Madhyastha, Lucia Specia. Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. 2021. Julia Ive, Andy Mingren Li, Yishu Miao, Ozan Caglayan, Pranava Swaroop Madhyastha, Lucia Specia |
EACL | 5 |
| 2021 | Numerical reasoning in machine reading comprehension tasks: are we there yet?abstractNumerical reasoning based machine reading comprehension is a task that involves reading comprehension along with using arithmetic operations such as addition, subtraction, sorting, and counting.The DROP benchmark (Dua et al., 2019) is a recent dataset that has inspired the design of NLP models aimed at solving this task.The current standings of these models in the DROP leaderboard, over standard metrics, suggest that the models have achieved near-human performance.However, does this mean that these models have learned to reason?In this paper, we present a controlled study on some of the top-performing model architectures for the task of numerical reasoning.Our observations suggest that the standard metrics are incapable of measuring progress towards such tasks. Hadeel Al-Negheimish, Pranava Swaroop Madhyastha, Alessandra Russo |
EMNLP (1) | 2 |
| 2021 | MSVD-Turkish: a comprehensive multimodal video dataset for integrated vision and language research in Turkish
Begüm Çitamak Erdinç, Ozan Caglayan, Menekse Kuyu, Erkut Erdem, Aykut Erdem, Pranava Swaroop Madhyastha, Lucia Specia |
Mach. Transl. | 6 |
| 2021 | Read, spot and translateabstractAbstract We propose multimodal machine translation (MMT) approaches that exploit the correspondences between words and image regions. In contrast to existing work, our referential grounding method considers objects as the visual unit for grounding, rather than whole images or abstract image regions, and performs visual grounding in the source language, rather than at the decoding stage via attention. We explore two referential grounding approaches: (i) implicit grounding, where the model jointly learns how to ground the source language in the visual representation and to translate; and (ii) explicit grounding, where grounding is performed independent of the translation model, and is subsequently used to guide machine translation. We performed experiments on the Multi30K dataset for three language pairs: English–German, English–French and English–Czech. Our referential grounding models outperform existing MMT models according to automatic and human evaluation metrics. Lucia Specia, Josiah Wang, Sun Jae Lee, Alissa Ostapenko, Pranava Swaroop Madhyastha |
Mach. Transl. | 5 |
| 2021 | Leveraging auxiliary image descriptions for dense video captioning
Emre Boran, Aykut Erdem, Nazli Ikizler-Cinbis, Erkut Erdem, Pranava Swaroop Madhyastha, Lucia Specia |
Pattern Recognit. Lett. | 5 |
| 2020 | Curious Case of Language Generation Evaluation Metrics: A Cautionary TaleabstractAutomatic evaluation of language generation systems is a well-studied problem in Natural Language Processing.While novel metrics are proposed every year, a few popular metrics remain as the de facto metrics to evaluate tasks such as image captioning and machine translation, despite their known limitations.This is partly due to ease of use, and partly because researchers expect to see them and know how to interpret them.In this paper, we urge the community for more careful consideration of how they automatically evaluate their models by demonstrating important failure cases on multiple datasets, language pairs and tasks.Our experiments show that metrics (i) usually prefer system outputs to human-authored texts, (ii) can be insensitive to correct translations of rare words, (iii) can yield surprisingly high scores when given a single sentence as system output for the entire test set. Ozan Caglayan, Pranava Swaroop Madhyastha, Lucia Specia |
COLING | 2 |
| 2020 | Deciding When, How and for Whom to SimplifyabstractCurrent Automatic Text Simplification (TS) work relies on sequence-to-sequence neural models that learn simplification operations from parallel complex-simple corpora. In this paper we address three open challenges in these approaches: (i) avoiding unnecessary transformations, (ii) determining which operations to perform, and (iii) generating simplifications that are suitable for a given target audience. For (i), we propose joint and two-stage approaches where instances are marked or classified as simple or complex. For (ii) and (iii), we propose fusion-based approaches to incorporate information on the target grade level as well as the types of operation to perform in the models. While grade-level information is provided as metadata, we devise predictors for the type of operation. We study different representations for this information as well as different ways in which it is used in the models. Our approach outperforms previous work on neural TS, with our best model following the two-stage approach and using the information about grade level and type of operation to initialise the encoder and the decoder, respectively. Carolina Scarton, Pranava Swaroop Madhyastha, Lucia Specia |
ECAI | 2 |
| 2020 | Simultaneous Machine Translation with Visual ContextabstractSimultaneous machine translation (SiMT) aims to translate a continuous input text stream into another language with the lowest latency and highest quality possible.The translation thus has to start with an incomplete source text, which is read progressively, creating the need for anticipation.In this paper, we seek to understand whether the addition of visual information can compensate for the missing source context.To this end, we analyse the impact of different multimodal approaches and visual features on state-of-the-art SiMT frameworks.Our results show that visual context is helpful and that visually-grounded models based on explicit object region information are much better than commonly used global features, reaching up to 3 BLEU points improvement under low latency scenarios.Our qualitative analysis illustrates cases where only the multimodal systems are able to translate correctly from English into gender-marked languages, as well as deal with differences in word order, such as adjective-noun placement between English and French. Ozan Caglayan, Julia Ive, Veneta Haralampieva, Pranava Swaroop Madhyastha, Loïc Barrault, Lucia Specia |
EMNLP (1) | 4 |
| 2019 | Distilling Translations with Visual AwarenessabstractPrevious work on multimodal machine translation has shown that visual information is only needed in very specific cases, for example in the presence of ambiguous words where the textual context is not sufficient.As a consequence, models tend to learn to ignore this information.We propose a translate-and-refine approach to this problem where images are only used by a second stage decoder.This approach is trained jointly to generate a good first draft translation and to improve over this draft by (i) making better use of the target language textual context (both left and right-side contexts) and (ii) making use of visual context.This approach leads to the state of the art results.Additionally, we show that it has the ability to recover from erroneous or missing words in the source language.EN: Three children in football uniforms are playing football.DE: Drei Kinder in Fußballtrikots spielen Fußball.PE: Drei Kinder in Footballtrikots spielen Football.(a) Ambiguous word football translated as soccer (Fußball) EN: A baseball player in a black shirt just tagged a player in a white shirt.DE: Ein Baseballspieler in einem schwarzen Shirt fängt einen Spieler in einem weißen Shirt.PE: Eine Baseballspielerin in einem schwarzen Shirt fängt eine Spielerin in einem weißen Shirt.(b) Gender-neutral word player translated as male player (Spieler) EN: A woman wearing a white shirt works out on an elliptical machine. Julia Ive, Pranava Swaroop Madhyastha, Lucia Specia |
ACL (1) | 2 |
| 2019 | VIFIDEL: Evaluating the Visual Fidelity of Image DescriptionsabstractWe address the task of evaluating image description generation systems.We propose a novel image-aware metric for this task: VIFIDEL.It estimates the faithfulness of a generated caption with respect to the content of the actual image, based on the semantic similarity between labels of objects depicted in images and words in the description.The metric is also able to take into account the relative importance of objects mentioned in human reference descriptions during evaluation.Even if these human reference descriptions are not available, VIFIDEL can still reliably evaluate system descriptions.The metric achieves high correlation with human judgments on two well-known datasets and is competitive with metrics that depend on and rely exclusively on human references. Pranava Swaroop Madhyastha, Josiah Wang, Lucia Specia |
ACL (1) | 1 |
| 2019 | On Model Stability as a Function of Random SeedabstractIn this paper, we focus on quantifying model stability as a function of random seed by investigating the effects of the induced randomness on model performance and the robustness of the model in general.We specifically perform a controlled study on the effect of random seeds on the behaviour of attention, gradientbased and surrogate model based (LIME) interpretations.Our analysis suggests that random seeds can adversely affect the consistency of models resulting in counterfactual interpretations.We propose a technique called Aggressive Stochastic Weight Averaging (ASWA) and an extension called Norm-filtered Aggressive Stochastic Weight Averaging (NASWA) which improves the stability of models over random seeds.With our ASWA and NASWA based optimization, we are able to improve the robustness of the original model, on average reducing the standard deviation of the model's performance by 72%. Pranava Swaroop Madhyastha |
CoNLL | 1 |
| 2019 | Deep Copycat Networks for Text-to-Text GenerationabstractJulia Ive, Pranava Madhyastha, Lucia Specia. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Julia Ive, Pranava Swaroop Madhyastha, Lucia Specia |
EMNLP/IJCNLP (1) | 2 |
| 2019 | Learning from Multiview Correlations in Open-domain VideosabstractAn increasing number of datasets contain multiple views, such as video, sound and automatic captions. A basic challenge in representation learning is how to leverage multiple views to learn better representations. This is further complicated by the existence of a latent alignment between views, such as between speech and its transcription, and by the multitude of choices for the learning objective. We explore an advanced, correlation-based representation learning method on a 4-way parallel, multimodal dataset, and assess the quality of the learned representations on retrieval-based tasks. We show that the proposed approach produces rich representations that capture most of the information shared across views. Our best models for speech and textual modalities achieve retrieval rates from 70.7% to 96.9% on open-domain, user-generated instructional videos. This shows it is possible to learn reliable representations across disparate, unaligned and noisy modalities, and encourages using the proposed approach on larger datasets. Nils Holzenberger, Shruti Palaskar, Pranava Swaroop Madhyastha, Florian Metze, Raman Arora |
ICASSP | 3 |
| 2018 | End-to-end Image Captioning Exploits Distributional Similarity in Multimodal Space
Pranava Swaroop Madhyastha, Josiah Wang, Lucia Specia |
BMVC | 1 |
| 2018 | Object Counts! Bringing Explicit Detections Back into Image CaptioningabstractJosiah Wang, Pranava Swaroop Madhyastha, Lucia Specia. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018. Josiah Wang, Pranava Swaroop Madhyastha, Lucia Specia |
NAACL-HLT | 2 |
| 2018 | The role of image representations in vision to language tasksabstractAbstract Tasks that require modeling of both language and visual information, such as image captioning, have become very popular in recent years. Most state-of-the-art approaches make use of image representations obtained from a deep neural network, which are used to generate language information in a variety of ways with end-to-end neural-network-based models. However, it is not clear how different image representations contribute to language generation tasks. In this paper, we probe the representational contribution of the image features in an end-to-end neural modeling framework and study the properties of different types of image representations. We focus on two popular vision to language problems: The task of image captioning and the task of multimodal machine translation. Our analysis provides interesting insights into the representational properties and suggests that end-to-end approaches implicitly learn a visual-semantic subspace and exploit the subspace to generate captions. Pranava Swaroop Madhyastha, Josiah Wang, Lucia Specia |
Nat. Lang. Eng. | 1 |
| 2017 | Exploring the use of acoustic embeddings in neural machine translationabstractNeural Machine Translation (NMT) has recently demonstrated improved performance over statistical machine translation and relies on an encoder-decoder framework for translating text from source to target. The structure of NMT makes it amenable to add auxiliary features, which can provide complementary information to that present in the source text. In this paper, auxiliary features derived from accompanying audio, are investigated for NMT and are compared and combined with text-derived features. These acoustic embeddings can help resolve ambiguity in the translation, thus improving the output. The following features are experimented with: Latent Dirichlet Allocation (LDA) topic vectors and GMM subspace i-vectors derived from audio. These are contrasted against: skip-gram/Word2Vec features and LDA features derived from text. The results are encouraging and show that acoustic information does help with NMT, leading to an overall 3.3% relative improvement in BLEU scores. Salil Deena, Raymond W. M. Ng, Pranava Swaroop Madhyastha, Lucia Specia, Thomas Hain |
ASRU | 3 |
| 2017 | Semi-Supervised Adaptation of RNNLMs by Fine-Tuning with Domain-Specific Auxiliary FeaturesabstractRecurrent neural network language models (RNNLMs) can be augmented with auxiliary features, which can provide an extra modality on top of the words. It has been found that RNNLMs perform best when trained on a large corpus of generic text and then fine-tuned on text corresponding to the sub-domain for which it is to be applied. However, in many cases the auxiliary features are available for the sub-domain text but not for the generic text. In such cases, semi-supervised techniques can be used to infer such features for the generic text data such that the RNNLM can be trained and then fine-tuned on the available in-domain data with corresponding auxiliary features. \n \nIn this paper, several novel approaches are investigated for dealing with the semi-supervised adaptation of RNNLMs with auxiliary features as input. These approaches include: using zero features during training to mask the weights of the feature sub-network; adding the feature sub-network only at the time of fine-tuning; deriving the features using a parametric model and; back-propagating to infer the features on the generic text. These approaches are investigated and results are reported both in terms of PPL and WER on a multi-genre broadcast ASR task. Salil Deena, Raymond W. M. Ng, Pranava Swaroop Madhyastha, Lucia Specia, Thomas Hain |
INTERSPEECH | 3 |
| 2017 | Exploring Hypotheses Spaces in Neural Machine Translation
Frédéric Blain, Lucia Specia, Pranava Swaroop Madhyastha |
MTSummit (1) | 3 |
| 2016 | Structured Prediction with Output Embeddings for Semantic Image AnnotationabstractAriadna Quattoni, Arnau Ramisa, Pranava Swaroop Madhyastha, Edgar Simo-Serra, Francesc Moreno-Noguer. Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2016. Ariadna Quattoni, Arnau Ramisa, Pranava Swaroop Madhyastha, Edgar Simo-Serra, Francesc Moreno-Noguer |
HLT-NAACL | 3 |
| 2014 | Learning Task-specific Bilexical Embeddings
Pranava Swaroop Madhyastha, Xavier Carreras, Ariadna Quattoni |
COLING | 1 |