Shravan Nayak

dblp:258/1105 · DBLP profile ↗
← Back
9ranked-venue papers
3as first author
8since 2021 · last 2026
0000-0002-5298-7121ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 2 first-author · 7 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Vision and language · 31% Language models and text generation · 25% Trustworthy machine learning · 14%
Human-computer interaction and pervasive computing
2 papers
Interaction techniques and input · 44% Human-AI interaction · 44% Design research and methods · 13%
Software engineering, system software, and programming languages
1 paper
Program synthesis and code generation · 100%

Topics — the 12 heaviest of 16, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Knowledge, reasoning and agents › Multi-agent systems › multi-agent systems engineering
multi-agent system design
1.012026
Grammar Search for Multi-Agent Systems · ACL (1) 2026
Machine learning › Trustworthy machine learning
fairness
0.912025
LIVS: A Pluralistic Alignment Dataset for Inclusive Public Spaces · ICML 2025
Computer vision › Vision and language › multimodal understanding
multimodal document understanding
0.912025
BigDocs: An Open Dataset for Training Multimodal Models on Document and Code Tasks · ICLR 2025
Natural language and speech › Language models and text generation › alignment
multi-objective alignment
0.912025
LIVS: A Pluralistic Alignment Dataset for Inclusive Public Spaces · ICML 2025
Natural language and speech › Language models and text generation › alignment
pluralistic alignment
0.912025
LIVS: A Pluralistic Alignment Dataset for Inclusive Public Spaces · ICML 2025
Machine learning › Generative modeling › diffusion model
text-to-image generation
0.912025
LIVS: A Pluralistic Alignment Dataset for Inclusive Public Spaces · ICML 2025
Human-AI interaction › GUI agent
computer use agents
0.912025
UI-Vision: A Desktop-centric GUI Benchmark for Visual Perception and Interaction · ICML 2025
Interaction techniques and input
GUI interaction
0.912025
UI-Vision: A Desktop-centric GUI Benchmark for Visual Perception and Interaction · ICML 2025
Program synthesis and code generation › code generation with language models
image-to-code generation
0.912025
BigDocs: An Open Dataset for Training Multimodal Models on Document and Code Tasks · ICLR 2025
Knowledge, reasoning and agents › Knowledge representation and reasoning
spatial reasoning
0.312025
UI-Vision: A Desktop-centric GUI Benchmark for Visual Perception and Interaction · ICML 2025
Design research and methods
participatory design
0.312025
LIVS: A Pluralistic Alignment Dataset for Inclusive Public Spaces · ICML 2025
Machine learning › Trustworthy machine learning › fairness › bias evaluation
cultural bias evaluation
0.212024
Benchmarking Vision Language Models for Cultural Understanding · EMNLP 2024

Methods — techniques the papers use, named apart from their topics

vision-language model · 1.7direct preference optimization · 1.7dataset curation · 1.7benchmark construction · 1.7Stable Diffusion XL · 1.7grammar search · 1.0vision language model benchmarking · 0.8
YearPublicationVenuePosition
2026 Grammar Search for Multi-Agent Systems
abstract
Mayank Singh, Vikas Yadav, Shiva Krishna Reddy Malay, Shravan Nayak, Sai Rajeswar, Sathwik Tejaswi Madhusudhan, Eduardo Blanco. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Vikas Yadav, Shiva Krishna Reddy Malay, Shravan Nayak, Sai Rajeswar, Sathwik Tejaswi Madhusudhan, Eduardo Blanco 0002
ACL (1)4
2026 Exploiting Domain-Specific Parallel Data on Multilingual Language Models for Low-Resource Language Translation
abstract
Neural Machine Translation (NMT) systems built on multilingual sequence-to-sequence Language Models (msLMs) fail to deliver expected results when the amount of parallel data for a language, as well as the language’s representation in the model are limited. This restricts the capabilities of domain-specific NMT systems for low-resource languages (LRLs). As a solution, parallel data from auxiliary domains can be used either to fine-tune or to further pre-train the msLM. We present an evaluation of the effectiveness of these two techniques in the context of domain-specific LRL-NMT. We also explore the impact of domain divergence on NMT model performance. We recommend several strategies for utilizing auxiliary parallel data in building domain-specific NMT models for LRLs.
Surangika Ranathunga, Shravan Nayak, Annie En-Shiun Lee, Shih-Ting Cindy Huang, Yuchen Zeng 0001, Yanke Mao, Yun-Hsiang Ray Chan, Songchen Yuan, Anthony Rinaldi
ACM Trans. Asian Low Resour. Lang. Inf. Process.2
2025 BigDocs: An Open Dataset for Training Multimodal Models on Document and Code Tasks
abstract
Multimodal AI has the potential to significantly enhance document-understanding tasks, such as processing receipts, understanding workflows, extracting data from documents, and summarizing reports. Code generation tasks that require long-structured outputs can also be enhanced by multimodality. Despite this, their use in commercial applications is often limited due to limited access to relevant training data and restrictive licensing, which hinders open access. To address these limitations, we introduce BigDocs-7.5M, a high-quality, open-access dataset comprising 7.5 million multimodal documents across 30 tasks. We use an efficient data curation process to ensure that our data is high quality and license-permissive. Our process emphasizes accountability, responsibility, and transparency through filtering rules, traceable metadata, and careful content analysis. Additionally, we introduce BigDocs-Bench,, a benchmark suite with 10 novel tasks where we carefully create datasets that reflect real-world use cases involving reasoning over Graphical User Interfaces (GUI) and code generation from images. Our experiments show that training with BigDocs-Bench, improves average performance up to 25.8% over closed-source GPT-4o in document reasoning and structured output tasks such as Screenshot2HTML or Image2Latex generation. Finally, human evaluations revealed that participants preferred the outputs from models trained with BigDocs over those from GPT-4o. This suggests that BigDocs can help both academics and the open-source community utilize and improve AI tools to enhance multimodal capabilities and document reasoning.
Juan A. Rodríguez, Xiangru Jian, Siba Smarak Panigrahi, Aarash Feizi, Abhay Puri, Akshay Kalkunte Suresh, François Savard, Ahmed Masry, Shravan Nayak, Rabiul Awal, Mahsa Massoud, Amirhossein Abaskohi, Suyuchen Wang, Pierre-André Noël, Mats Leon Richter, Saverio Vadacchino, Sanket Biswas
ICLR10
2025 LIVS: A Pluralistic Alignment Dataset for Inclusive Public Spaces
abstract
We introduce the Local Intersectional Visual Spaces (LIVS) dataset, a benchmark for multi-criteria alignment, developed through a two-year participatory process with 30 community organizations to support the pluralistic alignment of text-to-image (T2I) models in inclusive urban planning. The dataset encodes 37,710 pairwise comparisons across 13,462 images, structured along six criteria—Accessibility, Safety, Comfort, Invitingness, Inclusivity, and Diversity—derived from 634 community-defined concepts. Using Direct Preference Optimization (DPO), we fine-tune Stable Diffusion XL to reflect multi-criteria spatial preferences and evaluate the LIVS dataset and the fine-tuned model through four case studies: (1) DPO increases alignment with annotated preferences, particularly when annotation volume is high; (2) preference patterns vary across participant identities, underscoring the need for intersectional data; (3) human-authored prompts generate more distinctive visual outputs than LLM-generated ones, influencing annotation decisiveness; and (4) intersectional groups assign systematically different ratings across criteria, revealing the limitations of single-objective alignment. While DPO improves alignment under specific conditions, the prevalence of neutral ratings indicates that community values are heterogeneous and often ambiguous. LIVS provides a benchmark for developing T2I models that incorporate local, stakeholder-driven preferences, offering a foundation for context-aware alignment in spatial design.
Rashid Mushkani, Shravan Nayak, Hugo Berard, Allison Cohen, Shin Koseki, Hadrien Bertrand
ICML2
2025 UI-Vision: A Desktop-centric GUI Benchmark for Visual Perception and Interaction
abstract
Autonomous agents that navigate Graphical User Interfaces (GUIs) to automate tasks like document editing and file management can greatly enhance computer workflows. While existing research focuses on online settings, desktop environments, critical for many professional and everyday tasks, remain underexplored due to data collection challenges and licensing issues. We introduce UI-Vision, the first comprehensive, license-permissive benchmark for offline, fine-grained evaluation of computer use agents in real-world desktop environments. Unlike online benchmarks, UI-Vision provides: (i) dense, high-quality annotations of human demonstrations, including bounding boxes, UI labels, and action trajectories (clicks, drags, and keyboard inputs) across 83 software applications, and (ii) three fine-to-coarse grained tasks—Element Grounding, Layout Grounding, and Action Prediction—with well-defined metrics to rigorously evaluate agents’ performance in desktop environments. Our evaluation reveals critical limitations in state-of-the-art models like UI-TARS-72B, including issues with understanding professional software, spatial reasoning, and complex actions like drag-and-drop. These findings highlight the challenges in developing fully autonomous computer-use agents. With UI-Vision, we aim to advance the development of more capable agents for real-world desktop tasks.
Shravan Nayak, Xiangru Jian, Qinghong Lin, Juan A. Rodríguez, Montek Kalsi, Nicolas Chapados, M. Tamer Özsu, Aishwarya Agrawal, David Vázquez 0001, Christopher Joseph Pal, Perouz Taslakian, Spandana Gella, Sai Rajeswar
ICML1
2024 Benchmarking Vision Language Models for Cultural Understanding
abstract
Shravan Nayak, Kanishk Jain, Rabiul Awal, Siva Reddy, Sjoerd Van Steenkiste, Lisa Anne Hendricks, Karolina Stanczak, Aishwarya Agrawal. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Shravan Nayak, Kanishk Jain, Rabiul Awal, Siva Reddy, Sjoerd van Steenkiste, Lisa Anne Hendricks, Karolina Stanczak, Aishwarya Agrawal
EMNLP1
2022 A Deep Dive Into Neural Synchrony Evaluation for Audio-visual Translation
abstract
We present a comprehensive analysis of the neural audio-visual synchrony evaluation tool SyncNet. We assess the agreement of SyncNet scores vis-a-vis human perception and whether we can use these as a reliable metric for evaluating audio-visual lip-synchrony in generation tasks with no ground truth reference audio-video pair. We further look into the underlying elements in audio and video which vitally affect synchrony using interpretable explanations from SyncNet predictions and analyse its susceptibility by introducing adversarial noise. SyncNet has been used in numerous papers on visually-grounded text-to-speech for scenarios such as dubbing. We focus on this scenario which features many local asynchronies (something that SyncNet isn’t made for).
Shravan Nayak, Christian Schuler, Debjoy Saha, Timo Baumann
ICMI1
2022 Merkel Podcast Corpus: A Multimodal Dataset Compiled from 16 Years of Angela Merkel's Weekly Video Podcasts
abstract
We introduce the Merkel Podcast Corpus, an audio-visual-text corpus in German collected from 16 years of (almost) weekly Internet podcasts of former German chancellor Angela Merkel. To the best of our knowledge, this is the first single speaker corpus in the German language consisting of audio, visual and text modalities of comparable size and temporal extent. We describe the methods used with which we have collected and edited the data which involves downloading the videos, transcripts and other metadata, forced alignment, performing active speaker recognition and face detection to finally curate the single speaker dataset consisting of utterances spoken by Angela Merkel. The proposed pipeline is general and can be used to curate other datasets of similar nature, such as talk show contents. Through various statistical analyses and applications of the dataset in talking face generation and TTS, we show the utility of the dataset. We argue that it is a valuable contribution to the research community, in particular, due to its realistic and challenging material at the boundary between prepared and spontaneous speech.
Debjoy Saha, Shravan Nayak, Timo Baumann
LREC2
2020 The Two Shades of Dubbing in Neural Machine Translation
abstract
Dubbing has two shades; synchronisation constraints are applied only when the actor's mouth is visible on screen, while the translation is unconstrained for off-screen dubbing.Consequently, different synchronisation requirements, and therefore translation strategies, are applied depending on the type of dubbing.In this work, we manually annotate an existing dubbing corpus (Heroes) for this dichotomy.We show that, even though we did not observe distinctive features between on-and off-screen dubbing at the textual level, on-screen dubbing is more difficult for MT (-4 BLEU points).Moreover, synchronisation constraints dramatically decrease translation quality for off-screen dubbing.We conclude that, distinguishing between on-screen and off-screen dubbing is necessary for determining successful strategies for dubbing-customised Machine Translation.
Alina Karakanta, Supratik Bhattacharya, Shravan Nayak, Timo Baumann, Matteo Negri, Marco Turchi
COLING3