Pranava Swaroop Madhyastha

dblp:151/8501 · also Pranava Madhyastha · DBLP profile ↗
← Back
31ranked-venue papers
8as first author
15since 2021 · last 2026
0000-0002-4438-8161ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 27 · 8 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 "Although Powerful, it's not Infallible": Investigating Academic Researchers' Verification Challenges with LLMs
abstract
LLMs have great potential for shaping how people find and understand information. However, current tools can struggle to provide authoritative sources, fabricate plausible references, and present obstacles to assessing truthfulness of their outputs. Understanding how users verify LLM outputs is particularly important in scholarly disciplines where information produced becomes the foundation of future knowledge. We investigated the factors that influence academic researchers’ decisions to verify LLM responses, their verification strategies, and the effectiveness of those strategies. We conducted a naturalistic think-aloud study, followed by a semi-structured interview, where we observed 16 researchers across disciplines using LLMs of their choice to conduct a research information-seeking task. Our findings highlight that prevailing LLM design can hamper users’ ability to satisfy their information needs for several reasons, such as lack of transparency about sources used in LLM outputs and lack of faithfulness of LLM outputs to the source. Based on these findings, we discuss how future LLMs can better support users in effective verification.
Monica Visani Scozzi, Stephann Makri, Pranava Swaroop Madhyastha
CHIIR3
2025 Foundation model assisted visual analytics: Opportunities and Challenges
abstract
We explore the integration of foundation models, such as large language models (LLMs) and multimodal LLMs (MLLMs), into visual analytics (VA) systems through intuitive natural language interactions. We survey current research directions in this emerging field, examining how foundation models have already been integrated into key visualisation-related processes in VA: visual mapping, the creation of data visualisations; visualisation observation, the process of generating a finding through visualisation; and visualisation manipulation, changing the viewport or highlighting areas of interest within a visualisation. We also highlight new possibilities that foundation models bring to VA, in particular, the opportunities to use MLLMs to interpret visualisations directly, to integrate multimodal interactions, and to provide guidance to users. We finally conclude with a vision of future VA systems as collaborative partners in analysis and address the prominent challenges in realising this vision through foundation models. Our discussions in this paper aim to guide future researchers working on foundation model assisted VA systems and help them navigate common obstacles when developing these systems.
Maeve Hutchinson, Radu Jianu, Aidan Slingsby, Pranava Swaroop Madhyastha
Comput. Graph.4
2024 Towards Holistic, Pragmatic and Multimodal Conversational Systems
abstract
Language acquisition and utilization transcend the mere exchange of lexical units. Visual cues, prosody, gestures, body movements, and context play an undeniably crucial role. Humans naturally communicate multimodally, employing multiple channels and synthesizing information from diverse modalities. My research delves into the characterization and construction of multimodal models that seamlessly integrate data from multiple independent modalities. I will cover recent work that highlights the challenges, achievements, and opportunities towards developing capable multimodal discursive models.
Pranava Swaroop Madhyastha
AAAI1
2023 Are words equally surprising in audio and audio-visual comprehension?
Pranava Swaroop Madhyastha, Gabriella Vigliocco
CogSci1
2023 Towards preserving word order importance through Forced Invalidation
abstract
Large pre-trained language models such as BERT have been widely used as a framework for natural language understanding (NLU) tasks.However, recent findings have revealed that pre-trained language models are insensitive to word order.The performance on NLU tasks remains unchanged even after randomly permuting the word of a sentence, where crucial syntactic information is destroyed.To help preserve the importance of word order, we propose a simple approach called FORCED INVALIDA-TION (FI): forcing the model to identify permuted sequences as invalid samples.We perform an extensive evaluation of our approach on various English NLU and QA based tasks over BERT-based and attention-based models over word embeddings.Our experiments demonstrate that FI significantly improves the sensitivity of the models to word order. 1
Hadeel Al-Negheimish, Pranava Swaroop Madhyastha, Alessandra Russo
EACL2
2023 A study towards contextual understanding of toxicity in online conversations
abstract
Abstract Identifying and annotating toxic online content on social media platforms is an extremely challenging problem. Work that studies toxicity in online content has predominantly focused on comments as independent entities. However, comments on social media are inherently conversational, and therefore, understanding and judging the comments fundamentally requires access to the context in which they are made. We introduce a study and resulting annotated dataset where we devise a number of controlled experiments on the importance of context and other observable confounders – namely gender, age and political orientation – towards the perception of toxicity in online content. Our analysis clearly shows the significance of context and the effect of observable confounders on annotations. Namely, we observe that the ratio of toxic to non-toxic judgements can be very different for each control group, and a higher proportion of samples are judged toxic in the presence of contextual information.
Pranava Swaroop Madhyastha, Antigoni Founta, Lucia Specia
Nat. Lang. Eng.1
2022 Belief Revision Based Caption Re-ranker with Visual Semantic Information
abstract
In this work, we focus on improving the captions generated by image-caption generation systems. We propose a novel re-ranking approach that leverages visual-semantic measures to identify the ideal caption that maximally captures the visual information in the image. Our re-ranker utilizes the Belief Revision framework (Blok et al., 2003) to calibrate the original likelihood of the top-n captions by explicitly exploiting semantic relatedness between the depicted caption and the visual context. Our experiments demonstrate the utility of our approach, where we observe that our re-ranker can enhance the performance of a typical image-captioning system without necessity of any additional training or fine-tuning.
Ahmed Sabir, Francesc Moreno-Noguer, Pranava Swaroop Madhyastha, Lluís Padró 0001
COLING3
2022 Evaluation of Fake News Detection with Knowledge-Enhanced Language Models
Chenxi Whitehouse, Tillman Weyde, Pranava Swaroop Madhyastha, Nikos Komninos
ICWSM3
2021 BERTGen: Multi-task Generation through BERT
abstract
Faidon Mitzalis, Ozan Caglayan, Pranava Madhyastha, Lucia Specia. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Faidon Mitzalis, Ozan Caglayan, Pranava Swaroop Madhyastha, Lucia Specia
ACL/IJCNLP (1)3
2021 Cross-lingual Visual Pre-training for Multimodal Machine Translation
abstract
Ozan Caglayan, Menekse Kuyu, Mustafa Sercan Amac, Pranava Madhyastha, Erkut Erdem, Aykut Erdem, Lucia Specia. Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. 2021.
Ozan Caglayan, Menekse Kuyu, Mustafa Sercan Amac, Pranava Swaroop Madhyastha, Erkut Erdem, Aykut Erdem, Lucia Specia
EACL4
2021 Exploiting Multimodal Reinforcement Learning for Simultaneous Machine Translation
abstract
Julia Ive, Andy Mingren Li, Yishu Miao, Ozan Caglayan, Pranava Madhyastha, Lucia Specia. Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. 2021.
Julia Ive, Andy Mingren Li, Yishu Miao, Ozan Caglayan, Pranava Swaroop Madhyastha, Lucia Specia
EACL5
2021 Numerical reasoning in machine reading comprehension tasks: are we there yet?
abstract
Numerical reasoning based machine reading comprehension is a task that involves reading comprehension along with using arithmetic operations such as addition, subtraction, sorting, and counting.The DROP benchmark (Dua et al., 2019) is a recent dataset that has inspired the design of NLP models aimed at solving this task.The current standings of these models in the DROP leaderboard, over standard metrics, suggest that the models have achieved near-human performance.However, does this mean that these models have learned to reason?In this paper, we present a controlled study on some of the top-performing model architectures for the task of numerical reasoning.Our observations suggest that the standard metrics are incapable of measuring progress towards such tasks.
Hadeel Al-Negheimish, Pranava Swaroop Madhyastha, Alessandra Russo
EMNLP (1)2
2021 MSVD-Turkish: a comprehensive multimodal video dataset for integrated vision and language research in Turkish
Begüm Çitamak Erdinç, Ozan Caglayan, Menekse Kuyu, Erkut Erdem, Aykut Erdem, Pranava Swaroop Madhyastha, Lucia Specia
Mach. Transl.6
2021 Read, spot and translate
abstract
Abstract We propose multimodal machine translation (MMT) approaches that exploit the correspondences between words and image regions. In contrast to existing work, our referential grounding method considers objects as the visual unit for grounding, rather than whole images or abstract image regions, and performs visual grounding in the source language, rather than at the decoding stage via attention. We explore two referential grounding approaches: (i) implicit grounding, where the model jointly learns how to ground the source language in the visual representation and to translate; and (ii) explicit grounding, where grounding is performed independent of the translation model, and is subsequently used to guide machine translation. We performed experiments on the Multi30K dataset for three language pairs: English–German, English–French and English–Czech. Our referential grounding models outperform existing MMT models according to automatic and human evaluation metrics.
Lucia Specia, Josiah Wang, Sun Jae Lee, Alissa Ostapenko, Pranava Swaroop Madhyastha
Mach. Transl.5
2021 Leveraging auxiliary image descriptions for dense video captioning
Emre Boran, Aykut Erdem, Nazli Ikizler-Cinbis, Erkut Erdem, Pranava Swaroop Madhyastha, Lucia Specia
Pattern Recognit. Lett.5
2020 Curious Case of Language Generation Evaluation Metrics: A Cautionary Tale
abstract
Automatic evaluation of language generation systems is a well-studied problem in Natural Language Processing.While novel metrics are proposed every year, a few popular metrics remain as the de facto metrics to evaluate tasks such as image captioning and machine translation, despite their known limitations.This is partly due to ease of use, and partly because researchers expect to see them and know how to interpret them.In this paper, we urge the community for more careful consideration of how they automatically evaluate their models by demonstrating important failure cases on multiple datasets, language pairs and tasks.Our experiments show that metrics (i) usually prefer system outputs to human-authored texts, (ii) can be insensitive to correct translations of rare words, (iii) can yield surprisingly high scores when given a single sentence as system output for the entire test set.
Ozan Caglayan, Pranava Swaroop Madhyastha, Lucia Specia
COLING2
2020 Deciding When, How and for Whom to Simplify
abstract
Current Automatic Text Simplification (TS) work relies on sequence-to-sequence neural models that learn simplification operations from parallel complex-simple corpora. In this paper we address three open challenges in these approaches: (i) avoiding unnecessary transformations, (ii) determining which operations to perform, and (iii) generating simplifications that are suitable for a given target audience. For (i), we propose joint and two-stage approaches where instances are marked or classified as simple or complex. For (ii) and (iii), we propose fusion-based approaches to incorporate information on the target grade level as well as the types of operation to perform in the models. While grade-level information is provided as metadata, we devise predictors for the type of operation. We study different representations for this information as well as different ways in which it is used in the models. Our approach outperforms previous work on neural TS, with our best model following the two-stage approach and using the information about grade level and type of operation to initialise the encoder and the decoder, respectively.
Carolina Scarton, Pranava Swaroop Madhyastha, Lucia Specia
ECAI2
2020 Simultaneous Machine Translation with Visual Context
abstract
Simultaneous machine translation (SiMT) aims to translate a continuous input text stream into another language with the lowest latency and highest quality possible.The translation thus has to start with an incomplete source text, which is read progressively, creating the need for anticipation.In this paper, we seek to understand whether the addition of visual information can compensate for the missing source context.To this end, we analyse the impact of different multimodal approaches and visual features on state-of-the-art SiMT frameworks.Our results show that visual context is helpful and that visually-grounded models based on explicit object region information are much better than commonly used global features, reaching up to 3 BLEU points improvement under low latency scenarios.Our qualitative analysis illustrates cases where only the multimodal systems are able to translate correctly from English into gender-marked languages, as well as deal with differences in word order, such as adjective-noun placement between English and French.
Ozan Caglayan, Julia Ive, Veneta Haralampieva, Pranava Swaroop Madhyastha, Loïc Barrault, Lucia Specia
EMNLP (1)4
2019 Distilling Translations with Visual Awareness
abstract
Previous work on multimodal machine translation has shown that visual information is only needed in very specific cases, for example in the presence of ambiguous words where the textual context is not sufficient.As a consequence, models tend to learn to ignore this information.We propose a translate-and-refine approach to this problem where images are only used by a second stage decoder.This approach is trained jointly to generate a good first draft translation and to improve over this draft by (i) making better use of the target language textual context (both left and right-side contexts) and (ii) making use of visual context.This approach leads to the state of the art results.Additionally, we show that it has the ability to recover from erroneous or missing words in the source language.EN: Three children in football uniforms are playing football.DE: Drei Kinder in Fußballtrikots spielen Fußball.PE: Drei Kinder in Footballtrikots spielen Football.(a) Ambiguous word football translated as soccer (Fußball) EN: A baseball player in a black shirt just tagged a player in a white shirt.DE: Ein Baseballspieler in einem schwarzen Shirt fängt einen Spieler in einem weißen Shirt.PE: Eine Baseballspielerin in einem schwarzen Shirt fängt eine Spielerin in einem weißen Shirt.(b) Gender-neutral word player translated as male player (Spieler) EN: A woman wearing a white shirt works out on an elliptical machine.
Julia Ive, Pranava Swaroop Madhyastha, Lucia Specia
ACL (1)2
2019 VIFIDEL: Evaluating the Visual Fidelity of Image Descriptions
abstract
We address the task of evaluating image description generation systems.We propose a novel image-aware metric for this task: VIFIDEL.It estimates the faithfulness of a generated caption with respect to the content of the actual image, based on the semantic similarity between labels of objects depicted in images and words in the description.The metric is also able to take into account the relative importance of objects mentioned in human reference descriptions during evaluation.Even if these human reference descriptions are not available, VIFIDEL can still reliably evaluate system descriptions.The metric achieves high correlation with human judgments on two well-known datasets and is competitive with metrics that depend on and rely exclusively on human references.
Pranava Swaroop Madhyastha, Josiah Wang, Lucia Specia
ACL (1)1
2019 On Model Stability as a Function of Random Seed
abstract
In this paper, we focus on quantifying model stability as a function of random seed by investigating the effects of the induced randomness on model performance and the robustness of the model in general.We specifically perform a controlled study on the effect of random seeds on the behaviour of attention, gradientbased and surrogate model based (LIME) interpretations.Our analysis suggests that random seeds can adversely affect the consistency of models resulting in counterfactual interpretations.We propose a technique called Aggressive Stochastic Weight Averaging (ASWA) and an extension called Norm-filtered Aggressive Stochastic Weight Averaging (NASWA) which improves the stability of models over random seeds.With our ASWA and NASWA based optimization, we are able to improve the robustness of the original model, on average reducing the standard deviation of the model's performance by 72%.
Pranava Swaroop Madhyastha
CoNLL1
2019 Deep Copycat Networks for Text-to-Text Generation
abstract
Julia Ive, Pranava Madhyastha, Lucia Specia. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Julia Ive, Pranava Swaroop Madhyastha, Lucia Specia
EMNLP/IJCNLP (1)2
2019 Learning from Multiview Correlations in Open-domain Videos
abstract
An increasing number of datasets contain multiple views, such as video, sound and automatic captions. A basic challenge in representation learning is how to leverage multiple views to learn better representations. This is further complicated by the existence of a latent alignment between views, such as between speech and its transcription, and by the multitude of choices for the learning objective. We explore an advanced, correlation-based representation learning method on a 4-way parallel, multimodal dataset, and assess the quality of the learned representations on retrieval-based tasks. We show that the proposed approach produces rich representations that capture most of the information shared across views. Our best models for speech and textual modalities achieve retrieval rates from 70.7% to 96.9% on open-domain, user-generated instructional videos. This shows it is possible to learn reliable representations across disparate, unaligned and noisy modalities, and encourages using the proposed approach on larger datasets.
Nils Holzenberger, Shruti Palaskar, Pranava Swaroop Madhyastha, Florian Metze, Raman Arora
ICASSP3
2018 End-to-end Image Captioning Exploits Distributional Similarity in Multimodal Space
Pranava Swaroop Madhyastha, Josiah Wang, Lucia Specia
BMVC1
2018 Object Counts! Bringing Explicit Detections Back into Image Captioning
abstract
Josiah Wang, Pranava Swaroop Madhyastha, Lucia Specia. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018.
Josiah Wang, Pranava Swaroop Madhyastha, Lucia Specia
NAACL-HLT2
2018 The role of image representations in vision to language tasks
abstract
Abstract Tasks that require modeling of both language and visual information, such as image captioning, have become very popular in recent years. Most state-of-the-art approaches make use of image representations obtained from a deep neural network, which are used to generate language information in a variety of ways with end-to-end neural-network-based models. However, it is not clear how different image representations contribute to language generation tasks. In this paper, we probe the representational contribution of the image features in an end-to-end neural modeling framework and study the properties of different types of image representations. We focus on two popular vision to language problems: The task of image captioning and the task of multimodal machine translation. Our analysis provides interesting insights into the representational properties and suggests that end-to-end approaches implicitly learn a visual-semantic subspace and exploit the subspace to generate captions.
Pranava Swaroop Madhyastha, Josiah Wang, Lucia Specia
Nat. Lang. Eng.1
2017 Exploring the use of acoustic embeddings in neural machine translation
abstract
Neural Machine Translation (NMT) has recently demonstrated improved performance over statistical machine translation and relies on an encoder-decoder framework for translating text from source to target. The structure of NMT makes it amenable to add auxiliary features, which can provide complementary information to that present in the source text. In this paper, auxiliary features derived from accompanying audio, are investigated for NMT and are compared and combined with text-derived features. These acoustic embeddings can help resolve ambiguity in the translation, thus improving the output. The following features are experimented with: Latent Dirichlet Allocation (LDA) topic vectors and GMM subspace i-vectors derived from audio. These are contrasted against: skip-gram/Word2Vec features and LDA features derived from text. The results are encouraging and show that acoustic information does help with NMT, leading to an overall 3.3% relative improvement in BLEU scores.
Salil Deena, Raymond W. M. Ng, Pranava Swaroop Madhyastha, Lucia Specia, Thomas Hain
ASRU3
2017 Semi-Supervised Adaptation of RNNLMs by Fine-Tuning with Domain-Specific Auxiliary Features
abstract
Recurrent neural network language models (RNNLMs) can be augmented with auxiliary features, which can provide an extra modality on top of the words. It has been found that RNNLMs perform best when trained on a large corpus of generic text and then fine-tuned on text corresponding to the sub-domain for which it is to be applied. However, in many cases the auxiliary features are available for the sub-domain text but not for the generic text. In such cases, semi-supervised techniques can be used to infer such features for the generic text data such that the RNNLM can be trained and then fine-tuned on the available in-domain data with corresponding auxiliary features. \n \nIn this paper, several novel approaches are investigated for dealing with the semi-supervised adaptation of RNNLMs with auxiliary features as input. These approaches include: using zero features during training to mask the weights of the feature sub-network; adding the feature sub-network only at the time of fine-tuning; deriving the features using a parametric model and; back-propagating to infer the features on the generic text. These approaches are investigated and results are reported both in terms of PPL and WER on a multi-genre broadcast ASR task.
Salil Deena, Raymond W. M. Ng, Pranava Swaroop Madhyastha, Lucia Specia, Thomas Hain
INTERSPEECH3
2017 Exploring Hypotheses Spaces in Neural Machine Translation
Frédéric Blain, Lucia Specia, Pranava Swaroop Madhyastha
MTSummit (1)3
2016 Structured Prediction with Output Embeddings for Semantic Image Annotation
abstract
Ariadna Quattoni, Arnau Ramisa, Pranava Swaroop Madhyastha, Edgar Simo-Serra, Francesc Moreno-Noguer. Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2016.
Ariadna Quattoni, Arnau Ramisa, Pranava Swaroop Madhyastha, Edgar Simo-Serra, Francesc Moreno-Noguer
HLT-NAACL3
2014 Learning Task-specific Bilexical Embeddings
Pranava Swaroop Madhyastha, Xavier Carreras, Ariadna Quattoni
COLING1