Srinivas Sunkara

dblp:248/9019 · DBLP profile ↗
← Back
8ranked-venue papers
1as first author
7since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 1 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Question answering and dialogue systems · 37% Representation and self-supervised learning · 36% Vision and language · 23%
Human-computer interaction and pervasive computing
2 papers
User interface design and tools · 68% Accessibility and assistive technology · 32%

Topics — the 11 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Representation and self-supervised learning › pre-training
multimodal pretraining
1.022021
UIBert: Learning Generic Multimodal Representations for UI Understanding · IJCAI 2021
ActionBert: Leveraging User Actions for Semantic Understanding of User Interfaces · AAAI 2021
Natural language and speech › Question answering and dialogue systems › task-oriented dialogue
dialogue state tracking
0.922021
Overview of the Eighth Dialog System Technology Challenge: DSTC8 · IEEE ACM Trans. Audio Speech Lang. Process. 2021
Towards Scalable Multi-Domain Conversational Agents: The Schema-Guided Dialogue Dataset · AAAI 2020
Computer vision › Vision and language
vision-language model
0.812024
ScreenAI: A Vision-Language Model for UI and Infographics Understanding · IJCAI 2024
User interface design and tools
UI understanding
0.722021
ActionBert: Leveraging User Actions for Semantic Understanding of User Interfaces · AAAI 2021
UIBert: Learning Generic Multimodal Representations for UI Understanding · IJCAI 2021
Computer vision › Vision and language › multimodal understanding
GUI understanding
0.512021
UIBert: Learning Generic Multimodal Representations for UI Understanding · IJCAI 2021
Machine learning › Representation and self-supervised learning
multimodal representation learning
0.512021
UIBert: Learning Generic Multimodal Representations for UI Understanding · IJCAI 2021
Machine learning › Representation and self-supervised learning
pre-training
0.512021
UIBert: Learning Generic Multimodal Representations for UI Understanding · IJCAI 2021
Natural language and speech › Question answering and dialogue systems
task-oriented dialogue
0.412020
Towards Scalable Multi-Domain Conversational Agents: The Schema-Guided Dialogue Dataset · AAAI 2020
Natural language and speech › Language models and text generation
multimodal language model
0.212024
ScreenAI: A Vision-Language Model for UI and Infographics Understanding · IJCAI 2024
Natural language and speech › Question answering and dialogue systems › multimodal dialogue system
audio-visual scene-aware dialog
0.112021
Overview of the Eighth Dialog System Technology Challenge: DSTC8 · IEEE ACM Trans. Audio Speech Lang. Process. 2021
Accessibility and assistive technology
accessibility
0.112021
UIBert: Learning Generic Multimodal Representations for UI Understanding · IJCAI 2021

Methods — techniques the papers use, named apart from their topics

transformer · 1.0pre-training tasks · 1.0pre-training · 1.0multimodal representation learning · 1.0vision-language pretraining · 0.8end-to-end dialog modeling · 0.5zero-shot generalization · 0.4
YearPublicationVenuePosition
2025 ScreenQA: Large-Scale Question-Answer Pairs Over Mobile App Screenshots
abstract
Yu-Chung Hsiao, Fedir Zubach, Gilles Baechler, Srinivas Sunkara, Victor Carbune, Jason Lin, Maria Wang, Yun Zhu, Jindong Chen. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Yu-Chung Hsiao, Fedir Zubach, Gilles Baechler, Srinivas Sunkara, Victor Carbune, Maria Wang, Jindong Chen
NAACL (Long Papers)4
2024 ScreenAI: A Vision-Language Model for UI and Infographics Understanding
Gilles Baechler, Srinivas Sunkara, Maria Wang, Fedir Zubach, Hassan Mansoor, Vincent Etter, Victor Carbune, Jindong Chen, Abhanshu Sharma
IJCAI2
2022 Towards Better Semantic Understanding of Mobile Interfaces
abstract
Improving the accessibility and automation capabilities of mobile devices can have a significant positive impact on the daily lives of countless users. To stimulate research in this direction, we release a human-annotated dataset with approximately 500k unique annotations aimed at increasing the understanding of the functionality of UI elements. This dataset augments images and view hierarchies from RICO, a large dataset of mobile UIs, with annotations for icons based on their shapes and semantics, and associations between different elements and their corresponding text labels, resulting in a significant increase in the number of UI elements and the categories assigned to them. We also release models using image-only and multimodal inputs; we experiment with various architectures and study the benefits of using multimodal inputs on the new dataset. Our models demonstrate strong performance on an evaluation set of unseen apps, indicating their generalizability to newer screens. These models, combined with the new dataset, can enable innovative functionalities like referring to UI elements by their labels, improved coverage and better semantics for icons etc., which would go a long way in making UIs more usable for everyone.
Srinivas Sunkara, Maria Wang, Gilles Baechler, Yu-Chung Hsiao, Jindong Chen, Abhanshu Sharma, James W. Stout
COLING1
2022 A Unified Approach to Entity-Centric Context Tracking in Social Conversations
abstract
In human-human conversations, Context Tracking deals with identifying important entities and keeping track of their properties and relationships. This is a challenging problem that encompasses several subtasks such as slot tagging, coreference resolution, resolving plural mentions and entity linking. We approach this problem as an end-to-end modeling task where the conversational context is represented by an entity repository containing the entity references mentioned so far, their properties and the relationships between them. The repository is updated turn-by-turn, thus making training and inference computationally efficient even for long conversations. This paper lays the groundwork for an investigation of this framework in two ways. First, we release Contrack, a large scale human-human conversation corpus for context tracking with people and location annotations. It contains over 7000 conversations with an average of 11.8 turns, 5.8 entities and 15.2 references per conversation. Second, we open-source a neural network architecture for context tracking. Finally we compare this network to state-of-the-art approaches for the subtasks it subsumes and report results on the involved tradeoffs.
Ulrich Rückert 0003, Srinivas Sunkara, Abhinav Rastogi, Sushant Prakash, Pranav Khaitan
LREC2
2021 ActionBert: Leveraging User Actions for Semantic Understanding of User Interfaces
abstract
As mobile devices are becoming ubiquitous, regularly interacting with a variety of user interfaces (UIs) is a common aspect of daily life for many people. To improve the accessibility of these devices and to enable their usage in a variety of settings, building models that can assist users and accomplish tasks through the UI is vitally important. However, there are several challenges to achieve this. First, UI components of similar appearance can have different functionalities, making understanding their function more important than just analyzing their appearance. Second, domain-specific features like Document Object Model (DOM) in web pages and View Hierarchy (VH) in mobile applications provide important signals about the semantics of UI elements, but these features are not in a natural language format. Third, owing to a large diversity in UIs and absence of standard DOM or VH representations, building a UI understanding model with high coverage requires large amounts of training data. Inspired by the success of pre-training based approaches in NLP for tackling a variety of problems in a data-efficient way, we introduce a new pre-trained UI representation model called ActionBert. Our methodology is designed to leverage visual, linguistic and domain-specific features in user interaction traces to pre-train generic feature representations of UIs and their components. Our key intuition is that user actions, e.g., a sequence of clicks on different UI components, reveals important information about their functionality. We evaluate the proposed model on a wide variety of downstream tasks, ranging from icon classification to UI component retrieval based on its natural language description. Experiments show that the proposed ActionBert model outperforms multi-modal baselines across all downstream tasks by up to 15.5%.
Zecheng He, Srinivas Sunkara, Xiaoxue Zang, Nevan Wichers, Gabriel Schubiner, Ruby B. Lee, Jindong Chen
AAAI2
2021 UIBert: Learning Generic Multimodal Representations for UI Understanding
abstract
To improve the accessibility of smart devices and to simplify their usage, building models which understand user interfaces (UIs) and assist users to complete their tasks is critical. However, unique challenges are proposed by UI-specific characteristics, such as how to effectively leverage multimodal UI features that involve image, text, and structural metadata and how to achieve good performance when high-quality labeled data is unavailable. To address such challenges we introduce UIBert, a transformer-based joint image-text model trained through novel pre-training tasks on large-scale unlabeled UI data to learn generic feature representations for a UI and its components. Our key intuition is that the heterogeneous features in a UI are self-aligned, i.e., the image and text features of UI components, are predictive of each other. We propose five pretraining tasks utilizing this self-alignment among different features of a UI component and across various components in the same UI. We evaluate our method on nine real-world downstream UI tasks where UIBert outperforms strong multimodal baselines by up to 9.26% accuracy.
Chongyang Bai, Xiaoxue Zang, Srinivas Sunkara, Abhinav Rastogi, Jindong Chen, Blaise Agüera y Arcas
IJCAI4
2021 Overview of the Eighth Dialog System Technology Challenge: DSTC8
abstract
This paper introduces the Eighth Dialog System Technology Challenge. In line with recent challenges, the eighth edition focuses on applying end-to-end dialog technologies in a pragmatic way for multi-domain task-completion, noetic response selection, audio visual scene-aware dialog, and schema-guided dialog state tracking tasks. This paper describes the task definition, provided datasets, baselines and evaluation set-up for each track. We also summarize the results of the submitted systems to highlight the overall trends of the state-of-the-art technologies for the tasks.
Seokhwan Kim, Michel Galley, R. Chulaka Gunasekara, Adam Atkinson, Baolin Peng, Hannes Schulz, Jianfeng Gao 0001, Jinchao Li, Mahmoud Adada, Minlie Huang, Luis A. Lastras, Jonathan K. Kummerfeld, Walter S. Lasecki, Chiori Hori, Anoop Cherian, Tim K. Marks, Abhinav Rastogi, Xiaoxue Zang, Srinivas Sunkara
IEEE ACM Trans. Audio Speech Lang. Process.20
2020 Towards Scalable Multi-Domain Conversational Agents: The Schema-Guided Dialogue Dataset
abstract
Virtual assistants such as Google Assistant, Alexa and Siri provide a conversational interface to a large number of services and APIs spanning multiple domains. Such systems need to support an ever-increasing number of services with possibly overlapping functionality. Furthermore, some of these services have little to no training data available. Existing public datasets for task-oriented dialogue do not sufficiently capture these challenges since they cover few domains and assume a single static ontology per domain. In this work, we introduce the the Schema-Guided Dialogue (SGD) dataset, containing over 16k multi-domain conversations spanning 16 domains. Our dataset exceeds the existing task-oriented dialogue corpora in scale, while also highlighting the challenges associated with building large-scale virtual assistants. It provides a challenging testbed for a number of tasks including language understanding, slot filling, dialogue state tracking and response generation. Along the same lines, we present a schema-guided paradigm for task-oriented dialogue, in which predictions are made over a dynamic set of intents and slots, provided as input, using their natural language descriptions. This allows a single dialogue system to easily support a large number of services and facilitates simple integration of new services without requiring additional training data. Building upon the proposed paradigm, we release a model for dialogue state tracking capable of zero-shot generalization to new APIs, while remaining competitive in the regular setting.
Abhinav Rastogi, Xiaoxue Zang, Srinivas Sunkara, Pranav Khaitan
AAAI3