Varsha Suresh

dblp:271/0325 · DBLP profile ↗
← Back
9ranked-venue papers
5as first author
9since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 3 · 2 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Language models and text generation · 46% Speech recognition and synthesis · 23% Information extraction and text analysis · 17%
Human-computer interaction and pervasive computing
1 paper
Human-AI interaction · 77% Interaction techniques and input · 23%
Computer graphics and multimedia
1 paper
Audio and music processing · 100%

Topics — the 8 heaviest of 9, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation › natural language understanding
discourse modeling
0.912025
Enhancing Spoken Discourse Modeling in Language Models Using Gestural Cues · ACL (1) 2025
Natural language and speech › Language models and text generation › language modeling
multimodal language modeling
0.912025
Enhancing Spoken Discourse Modeling in Language Models Using Gestural Cues · ACL (1) 2025
Natural language and speech › Speech recognition and synthesis
speech language model
0.912025
Enhancing Spoken Discourse Modeling in Language Models Using Gestural Cues · ACL (1) 2025
Machine learning › Representation and self-supervised learning
contrastive learning
0.512021
Not All Negatives are Equal: Label-Aware Contrastive Loss for Fine-grained Text Classification · EMNLP (1) 2021
Natural language and speech › Information extraction and text analysis › text classification
fine-grained text classification
0.512021
Not All Negatives are Equal: Label-Aware Contrastive Loss for Fine-grained Text Classification · EMNLP (1) 2021
Audio and music processing
speech processing
0.312025
Enhancing Spoken Discourse Modeling in Language Models Using Gestural Cues · ACL (1) 2025
Interaction techniques and input
gesture
0.312025
GestureCoach: Rehearsing for Engaging Talks with LLM-Driven Gesture Recommendations · UIST 2025
Natural language and speech › Information extraction and text analysis
emotion recognition
0.112021
Not All Negatives are Equal: Label-Aware Contrastive Loss for Fine-grained Text Classification · EMNLP (1) 2021

Methods — techniques the papers use, named apart from their topics

text infilling · 1.7feature alignment · 1.7VQ-VAE · 1.7large language model · 0.9gesture retrieval · 0.9label-aware contrastive learning · 0.5contrastive loss · 0.5
YearPublicationVenuePosition
2026 MUStReason: A Benchmark for Diagnosing Pragmatic Reasoning in VideoLMs for Multimodal Sarcasm Detection
Anisha Saha, Varsha Suresh, Timothy M. Hospedales, Vera Demberg
LREC2
2025 Enhancing Spoken Discourse Modeling in Language Models Using Gestural Cues
abstract
Research in linguistics shows that non-verbal cues, such as gestures, play a crucial role in spoken discourse.For example, speakers perform hand gestures to indicate topic shifts, helping listeners identify transitions in discourse.In this work, we investigate whether the joint modeling of gestures using human motion sequences and language can improve spoken discourse modeling in language models.To integrate gestures into language models, we first encode 3D human motion sequences into discrete gesture tokens using a VQ-VAE.These gesture token embeddings are then aligned with text embeddings through feature alignment, mapping them into the text embedding space.To evaluate the gesture-aligned language model on spoken discourse, we construct text infilling tasks targeting three key discourse cues grounded in linguistic research: discourse connectives, stance markers, and quantifiers.Results show that incorporating gestures enhances marker prediction accuracy across the three tasks, highlighting the complementary information that gestures can offer in modeling spoken discourse.We view this work as an initial step toward leveraging non-verbal cues to advance spoken language modeling in language models.
Varsha Suresh, Muhammad Hamza Mughal, Christian Theobalt, Vera Demberg
ACL (1)1
2025 Hybrid Multi-View Approach Towards Augmenting Large Language Models for Human Activity Recognition
abstract
Human Activity Recognition (HAR) is pivotal for behavior monitoring in public healthcare, supporting tasks like medical rehabilitation and targeted wellness campaigns. Initially utilizing hand-crafted features and traditional machine learning models such as support vector machines and decision trees, HAR has since evolved to include deep learning techniques, especially convolutional neural networks and Transformers, which excel at modeling temporal and task-specific information from sensor data. Despite these advancements, challenges related to high task dependency and data scarcity persist. To address these issues, there have been efforts to harness the vast pre-trained knowledge of Large Language Models (LLMs) for HAR. Yet, LLMs often fail to fully capture the temporal dynamics inherent in sensor data. We introduce HyMv, a novel hybrid approach that combines an auxiliary HAR model with soft-prompt tuning of LLMs. This approach leverages the auxiliary model’s proficiency in processing sensor data to guide the parameter optimization of LLMs during prompt tuning. Importantly, the auxiliary HAR model is only active during training, augmenting the LLM’s parameters without adding computational overhead during testing. We evaluate the performance of HyMv for various HAR tasks on diverse datasets and demonstrate its adaptability to different sensor modalities, further showcasing its broad applicability.
Suman Bhoi, Varsha Suresh, Wynne Hsu, Mong-Li Lee
ECAI2
2025 Synthetic Data Augmentation for Cross-domain Implicit Discourse Relation Recognition
abstract
Implicit discourse relation recognition (IDRR) – the task of identifying the implicit coherence relation between two text spans – requires deep semantic understanding. Recent studies have shown that zero-/few-shot approaches significantly lag behind supervised models. However, LLMs may be useful for synthetic data augmentation, where LLMs generate a second argument following a specified coherence relation. We applied this approach in a cross-domain setting, generating discourse continuations using unlabelled target-domain data to adapt a base model which was trained on source-domain labelled data. Evaluations conducted on a large-scale test set revealed that different variations of the approach did not result in any significant improvements. We conclude that LLMs often fail to generate useful samples for IDRR, and emphasize the importance of considering both statistical significance and comparability when evaluating IDRR models.
Frances Yung, Varsha Suresh, Zaynab Reza, Mansoor Ahmad, Vera Demberg
SIGDIAL2
2025 GestureCoach: Rehearsing for Engaging Talks with LLM-Driven Gesture Recommendations
abstract
Rehearse 3/10Why are we here, why are we alive?One key reason is that our ancestors on the savannas of Africa were really good at one thing…They weren't bigger than the animals they took down a lot of the time, they weren't faster than the animals they took down a lot of the time, but they were much better at banding together into groups and cooperating.…One key reason… Proactive Gesture Cues Current Slide Hover to Preview and Modify Gestures Presenter Notes with Gesture highlights Edit DeleteWhy were humans at hunting ?• Not bigger or faster than animals • Good at cooperationFigure 1: GestureCoach guides speakers to perform gestures while rehearsing their talk.A gesture recommendation model predicts text segments in the presenter notes that should be emphasized with gestures and retrieves relevant semantic gestures for each segment.During rehearsal, the system highlights the segments and proactively cues a video clip of the gesture by tracking users' speech, allowing them to integrate the gesture smoothly into their talk.Hovering over a segment allows users to preview and modify the associated gesture with alternate suggestions from the model.
Ashwin Ram 0004, Varsha Suresh, Artin Saberpour, Vera Demberg, Jürgen Steimle
UIST2
2024 An Adapter-Based Unified Model for Multiple Spoken Language Processing Tasks
abstract
Self-supervised learning models have revolutionized the field of speech processing. However, the process of fine-tuning these models on downstream tasks requires substantial computational resources, particularly when dealing with multiple speech-processing tasks. In this paper, we explore the potential of adapter-based fine-tuning in developing a unified model capable of effectively handling multiple spoken language processing tasks. The tasks we investigate are Automatic Speech Recognition, Phoneme Recognition, Intent Classification, Slot Filling, and Spoken Emotion Recognition. We validate our approach through a series of experiments on the SUPERB benchmark, and our results indicate that adapter-based fine-tuning enables a single encoder-decoder model to perform multiple speech processing tasks with an average improvement of 18.4 % across the five target tasks while staying efficient in terms of parameter updates.
Varsha Suresh, Salah Ait-Mokhtar, Caroline Brun, Ioan Calapodescu
ICASSP1
2022 Using Positive Matching Contrastive Loss with Facial Action Units to mitigate bias in Facial Expression Recognition
abstract
Machine learning models automatically learn dis-criminative features from the data, and are therefore susceptible to learn strongly-correlated biases, such as using protected attributes like gender and race. Most existing bias mitigation approaches aim to explicitly reduce the model's focus on these protected features. In this work, we propose to mitigate bias by explicitly guiding the model's focus towards task-relevant features using domain knowledge, and we hypothesize that this can indirectly reduce the dependence of the model on spurious correlations it learns from the data. We explore bias mitigation in facial expression recognition systems using facial Action Units (AUs) as the task-relevant feature. To this end, we introduce Feature-based Positive Matching Contrastive Loss which learns the distances between the positives of a sample based on the similarity between their corresponding AU embeddings. We compare our approach with representative baselines and show that incorporating task-relevant features via our method can improve model fairness at minimal cost to classification performance.
Varsha Suresh, Desmond C. Ong
ACII1
2021 Using Knowledge-Embedded Attention to Augment Pre-trained Language Models for Fine-Grained Emotion Recognition
abstract
Modern emotion recognition systems are trained to recognize only a small set of emotions, and hence fail to capture the broad spectrum of emotions people experience and express in daily life. In order to engage in more empathetic interactions, future AI has to perform fine-grained emotion recognition, distinguishing between many more varied emotions. Here, we focus on improving fine-grained emotion recognition by introducing external knowledge into a pre-trained self-attention model. We propose Knowledge-Embedded Attention (KEA) to use knowledge from emotion lexicons to augment the contextual representations from pre-trained ELECTRA and BERT models. Our results and error analyses outperform previous models on several datasets, and is better able to differentiate closely-confusable emotions, such as afraid and terrified.
Varsha Suresh, Desmond C. Ong
ACII1
2021 Not All Negatives are Equal: Label-Aware Contrastive Loss for Fine-grained Text Classification
abstract
Fine-grained classification involves dealing with datasets with larger number of classes with subtle differences between them.Guiding the model to focus on differentiating dimensions between these commonly confusable classes is key to improving performance on fine-grained tasks.In this work, we analyse the contrastive fine-tuning of pre-trained language models on two fine-grained text classification tasks, emotion classification and sentiment analysis.We adaptively embed class relationships into a contrastive objective function to help differently weigh the positives and negatives, and in particular, weighting closely confusable negatives more than less similar negative examples.We find that Label-aware Contrastive Loss outperforms previous contrastive methods, in the presence of larger number and/or more confusable classes, and helps models to produce output distributions that are more differentiated.
Varsha Suresh, Desmond C. Ong
EMNLP (1)1