VLDB 2026 Research / reviewers in the wild / expert
Chitralekha Gupta
dblp:116/5545
· DBLP profile ↗
22ranked-venue papers
10as first author
16since 2021 · last 2026
0000-0003-1350-9095ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 6 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 5 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 5 · 2 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Feeling the Facts: Real-time wearable fact-checkers can use nudges to reduce user belief in false informationabstractMisinformation can spread rapidly in everyday conversation, where pausing to verify is not always possible. We envision a wearable system that bridges the timing gap between hearing a claim and forming a judgment. It uses ambient listening to detect verifiable claims, performs rapid web verification, and provides a subtle haptic nudge with a glanceable overview. A controlled study (N=34) simulated this approach and tested against a no-support baseline. Results show that instant, body-integrated feedback significantly improved real-time truth discernment and increased verification activity compared to unsupported fact-checking. However, it also introduced over-reliance when the system made errors, i.e. failed to flag false claims or flagged true claims as false. We contribute empirical evidence of improved discernment alongside insights into trust, effort, and user–system tensions in verification wearables. Chitralekha Gupta, Nadia Victoria Aritonang, Dixon Prem Daniel Rajendran, Valdemar Danry, Pattie Maes, Suranga Nanayakkara |
CHI | 1 |
| 2026 | Beyond Descriptions: A Generative Scene2Audio Framework for Blind and Low-Vision Users to Experience Vista LandscapesabstractCurrent scene perception tools for Blind and Low Vision (BLV) individuals rely on spoken descriptions but lack engaging representations of visually pleasing distant environmental landscapes (Vista spaces). Our proposed Scene2Audio framework generates comprehensible and enjoyable nonverbal audio using generative models informed by psychoacoustics, and principles of scene audio composition. Through a user study with 11 BLV participants, we found that combining the Scene2Audio sounds with speech creates a better experience than speech alone, as the sound effects complement the speech making the scene easier to imagine. A mobile app “in-the-wild” study with 7 BLV users for more than a week further showed the potential of Scene2Audio in enhancing outdoor scene experiences. Our work bridges the gap between visual and auditory scene perception by moving beyond purely descriptive aids, addressing the aesthetic needs of BLV users. Chitralekha Gupta, Ashwin Ram 0002, Shreyas Sridhar, Christophe Jouffrais, Suranga Nanayakkara |
CHI | 1 |
| 2025 | Human Robot Interaction for Blind and Low Vision People: A Systematic Literature ReviewabstractInternational audience Yize Wei, Nathan Rocher, Chitralekha Gupta, Mia Huong Nguyen, Roger Zimmermann, Wei Tsang Ooi, Christophe Jouffrais, Suranga Nanayakkara |
CHI | 3 |
| 2025 | MorphFader: Enabling Fine-grained Controllable Morphing with Text-to-Audio ModelsabstractSound morphing is the process of gradually and smoothly transforming one sound into another to generate novel and perceptually hybrid sounds that simultaneously resemble both. Recently, diffusion-based text-to-audio models have produced high-quality sounds using text prompts. However, granularly controlling the semantics of the sound, which is necessary for morphing, can be challenging using text. In this paper, we propose MorphFader, a controllable method for morphing sounds generated by disparate prompts using text-to-audio models. By intercepting and interpolating the components of the cross-attention layers within the diffusion process, we can create smooth morphs between sounds generated by different text prompts. Using both objective metrics and perceptual listening tests, we demonstrate the ability of our method to granularly control the semantics in the sound and generate smooth morphs. Purnima Kamath, Chitralekha Gupta, Suranga Nanayakkara |
ICASSP | 2 |
| 2025 | Towards Temporally Explainable Dysarthric Speech Clarity Assessment
Seohyun Park, Chitralekha Gupta, Michelle Kah Yian Kwan, Xinhui Fung, Alexander Wenjun Yip, Suranga Nanayakkara |
INTERSPEECH | 2 |
| 2025 | DroneAudioset: An Audio Dataset for Drone-based Search and RescueabstractUnmanned Aerial Vehicles (UAVs) or drones, are increasingly used in search and rescue missions to detect human presence. Existing systems primarily leverage vision-based methods which are prone to fail under low-visibility or occlusion. Drone-based audio perception offers promise but suffers from extreme ego-noise that masks sounds indicating human presence. Existing datasets are either limited in diversity or synthetic, lacking real acoustic interactions, and there are no standardized setups for drone audition. To this end, we present DroneAudioset (The dataset is publicly available at https://huggingface.co/datasets/ahlab-drone-project/DroneAudioSet/ under the MIT license), a comprehensive drone audition dataset featuring 23.5 hours of annotated recordings, covering a wide range of signal-to-noise ratios (SNRs) from -57.2 dB to -2.5 dB, across various drone types, throttles, microphone configurations as well as environments. The dataset enables development and systematic evaluation of noise suppression and classification methods for human-presence detection under challenging conditions, while also informing practical design considerations for drone audition systems, such as microphone placement trade-offs, and development of drone noise-aware audio processing. This dataset is an important step towards enabling design and deployment of drone-audition systems. Chitralekha Gupta, Soundarya Ramesh, Praveen Sasikumar, Kian Peen Yeo, Suranga Nanayakkara |
NeurIPS | 1 |
| 2024 | Can AI Models Summarize Your Diary Entries? Investigating Utility of Abstractive Summarization for Autobiographical TextabstractJournaling is a widely adopted technique, known to improve mental health and well-being by enabling reflection on past events. Large amounts of text in digital journaling applications could hinder the reflection process due to information overload. Abstractive summarization can solve this problem by generating short summaries to quickly glance at and reminisce. In this paper, we present an investigation of the utility of large language models in the context of autobiographical text summarization. We study two approaches to adapt a self-supervised learning (SSL) model to the domain of autobiographical text. One model employs transfer learning using our new autobiographical text summary dataset to fine-tune the SSL model. The second model leverages existing news datasets for high-quality text summarization mixed with our autobiographical summary dataset. We conducted mixed methods research to analyze the performance of these two models. Through objective evaluation using ROUGE and BART scores, we find that both these approaches perform significantly better than the SSL model fine-tuned with only high-quality news datasets, showing the importance of domain adaptation and autobiographical text summary dataset for this task. Secondly, through a subjective evaluation on a crowd-sourcing platform, we evaluated the summaries generated from these models on various quality criteria such as grammar, non-redundancy, structure, and coherence. We found that on all criteria, these summaries score >4 out of 5, and the two models show comparable results. We deployed a proof-of-concept web-based journaling application to assess the practical real-world implications of incorporating abstractive summarization in a digital journaling context. We found that the participants showed a high consensus that the summaries generated by the system captured the main idea of their journal entry (80% of the 75 participants gave a Likert scale rating of ≥5.0 out of 7.0, with the overall mean rating of 5.56 ± 1.32) while being factually correct, and they found it to be a useful feature of a journaling application. Finally, we conducted human evaluation studies to compare the quality of the summaries generated from a commercial tool ChatGPT and mixed distribution fine-tuned SSL model, and present insights into these systems in the context of autobiographical abstractive text summarization. We have made our model, dataset, and subjective evaluation questionnaire openly available to the research community. Shamane Siriwardhana, Chitralekha Gupta, Tharindu Kaluarachchi, Vipula Dissanayake, Suveen Ellawela, Suranga Nanayakkara |
Int. J. Hum. Comput. Interact. | 2 |
| 2024 | Example-Based Framework for Perceptually Guided Audio Texture GenerationabstractControllable generation in StyleGANs is usually achieved by training the model using labeled data. For audio textures, however, there is currently a lack of large semantically labeled datasets. Therefore, to control generation, we develop a method for semantic control over an unconditionally trained StyleGAN in the absence of such labeled datasets. In this paper, we propose an example-based framework to determine guidance vectors for audio texture generation based on user-defined semantic attributes. Our approach leverages the semantically disentangled latent space of an unconditionally trained StyleGAN. By using a few synthetic examples to indicate the presence or absence of a semantic attribute, we infer the guidance vectors in the latent space of the StyleGAN to control that attribute during generation. Our results show that our framework can find user-defined and perceptually relevant guidance vectors for controllable generation for audio textures. Furthermore, we demonstrate an application of our framework to other tasks, such as selective semantic attribute transfer. Purnima Kamath, Chitralekha Gupta, Lonce L. Wyse, Suranga Nanayakkara |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2024 | VR.net: A Real-world Dataset for Virtual Reality Motion Sickness ResearchabstractResearchers have used machine learning approaches to identify motion sickness in VR experience. These approaches would certainly benefit from an accurately labeled, real-world, diverse dataset that enables the development of generalizable ML models. We introduce 'VR.net', a dataset comprising 165-hour gameplay videos from 100 real-world games spanning ten diverse genres, evaluated by 500 participants. VR.net accurately assigns 24 motion sickness-related labels for each video frame, such as camera/object movement, depth of field, and motion flow. Building such a dataset is challenging since manual labeling would require an infeasible amount of time. Instead, we implement a tool to automatically and precisely extract ground truth data from 3D engines' rendering pipelines without accessing VR games' source code. We illustrate the utility of VR.net through several applications, such as risk factor detection and sickness level prediction. We believe that the scale, accuracy, and diversity of VR.net can offer unparalleled opportunities for VR motion sickness research and beyond.We also provide access to our data collection tool, enabling researchers to contribute to the expansion of VR.net. Elliott Wen, Chitralekha Gupta, Prasanth Sasikumar, Mark Billinghurst, James Wilmott, Emily Skow, Arindam Dey 0001, Suranga Nanayakkara |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2023 | Towards Controllable Audio Texture MorphingabstractIn this paper, we propose a data-driven approach to train a Generative Adversarial Network (GAN) conditioned on "soft-labels" distilled from the penultimate layer of an audio classifier trained on a target set of audio texture classes. We demonstrate that interpolation between such conditions or control vectors provide smooth morphing between the generated audio textures, and show similar or better audio texture morphing capability compared to the state-of-the-art methods. The proposed approach results in a well-organized latent space that generates novel audio outputs while remaining consistent with the semantics of the conditioning parameters. This is a step towards a general data-driven approach to designing generative audio models with customized controls capable of traversing out-of-distribution regions for novel sound synthesis. Chitralekha Gupta, Purnima Kamath, Yize Wei, Zhuoyao Li, Suranga Nanayakkara, Lonce L. Wyse |
ICASSP | 1 |
| 2023 | Evaluating Descriptive Quality of AI-Generated Audio Using Image-SchemasabstractNovel AI-generated audio samples are evaluated for descriptive qualities such as the smoothness of a morph using crowdsourced human listening tests. However, the methods to design interfaces for such experiments and to effectively articulate the descriptive audio quality under test receive very little attention in the evaluation metrics literature. In this paper, we explore the use of visual metaphors of image-schema to design interfaces to evaluate AI-generated audio. Furthermore, we highlight the importance of framing and contextualizing a descriptive audio quality under measurement using such constructs. Using both pitched sounds and textures, we conduct two sets of experiments to investigate how the quality of responses vary with audio and task complexities. Our results show that, in both cases, by using image-schemas we can improve the quality and consensus of AI-generated audio evaluations. Our findings reinforce the importance of interface design for listening tests and stationary visual constructs to communicate temporal qualities of AI-generated audio samples, especially to naïve listeners on crowdsourced platforms. Purnima Kamath, Zhuoyao Li, Chitralekha Gupta, Kokil Jaidka, Suranga Nanayakkara, Lonce L. Wyse |
IUI | 3 |
| 2023 | PoLyScriber: Integrated Fine-Tuning of Extractor and Lyrics Transcriber for Polyphonic MusicabstractLyrics transcription of polyphonic music is challenging as the background music affects lyrics intelligibility. Typically, lyrics transcription can be performed by a two-step pipeline, i.e. a singing vocal extraction front end, followed by a lyrics transcriber back end, where the front end and back end are trained separately. Such a two-step pipeline suffers from both imperfect vocal extraction and mismatch between front end and back end. In this work, we propose a novel end-to-end integrated fine-tuning framework, that we call PoLyScriber, to globally optimize the vocal extractor front end and lyrics transcriber back end for lyrics transcription in polyphonic music. The experimental results show that our proposed PoLyScriber achieves substantial improvements over the existing approaches on publicly available test datasets. Xiaoxue Gao, Chitralekha Gupta, Haizhou Li 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2022 | Genre-Conditioned Acoustic Models for Automatic Lyrics Transcription of Polyphonic MusicabstractLyrics transcription of polyphonic music is challenging not only because the singing vocals are corrupted by the background music, but also because the background music and the singing style vary across music genres, such as pop, metal, and hip hop, which affects lyrics intelligibility of the song in different ways. In this work, we propose to transcribe the lyrics of polyphonic music using a novel genre-conditioned network. The proposed network adopts pre-trained model parameters, and incorporates the genre adapters between layers to capture different genre peculiarities for lyrics-genre pairs, thereby only requiring lightweight genre-specific parameters for training. Our experiments show that the proposed genre-conditioned network outperforms the existing lyrics transcription systems. Xiaoxue Gao, Chitralekha Gupta, Haizhou Li 0001 |
ICASSP | 2 |
| 2022 | Automatic Lyrics Transcription of Polyphonic Music With Lyrics-Chord Multi-Task LearningabstractLyrics are the words that make up a song, while chords are harmonic sets of multiple notes in music. Lyrics and chords are generally essential information in music, i.e. unaccompanied singing vocals mixed with instrumental music, representing important components in polyphonic music. In a traditional lyrics transcription task, we first extract the singing vocals from the polyphonic music and then transcribe the resulting singing vocals, where the two steps are optimized independently. In this paper, we propose novel end-to-end network architectures that are designed to disentangle lyrics from chords in polyphonic music for effective lyrics transcription in a single step, where we consider chords as musical words, analogously to lexical words as lyrics intuitively. We start by studying a single-task lyrics transcriber as the reference baseline and the initial model to develop the multi-task lyrics transcription solutions. The main idea is to take advantage of chord transcription available in the training data through multi-task training to improve lyrics transcription. The experiments show that the proposed multitask lyrics transcriber significantly outperforms other competing solutions, with a word error rate (WER) of 31.82% on a standard test dataset. Xiaoxue Gao, Chitralekha Gupta, Haizhou Li 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2022 | Deep Learning Approaches in Topics of Singing Information ProcessingabstractSinging, the vocal production of musical tones, is one of the most important elements of music. Addressing the needs of real-world applications, the study of technologies related to singing voices has become an increasingly active area of research. In this paper, we provide a comprehensive overview of the recent developments in the field of singing information processing, specifically in the topics of singing skill evaluation, singing voice synthesis, singing voice separation, and lyrics synchronization and transcription. We will especially focus on deep learning approaches including modern representation learning techniques for singing voices. We will also provide an overview of contributions in public datasets for singing voice research. Chitralekha Gupta, Haizhou Li 0001, Masataka Goto |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2021 | Proceedings of the 22nd Annual Meeting of the Special Interest Group on Discourse and Dialogue
Haizhou Li 0001, Gina-Anne Levow, Chitralekha Gupta, Berrak Sisman, Siqi Cai 0002, David Vandyke, Nina Dethlefs, Yan Wu 0002, Junyi Jessy Li |
SIGDIAL | 4 |
| 2020 | Automatic Lyrics Alignment and Transcription in Polyphonic Music: Does Background Music Help?abstractAutomatic lyrics alignment and transcription in polyphonic music are challenging tasks because the singing vocals are corrupted by the background music. In this work, we propose to learn music genre-specific characteristics to train polyphonic acoustic models. We first compare several automatic speech recognition pipelines for the application of lyrics transcription. We then present the lyrics alignment and transcription performance of music-informed acoustic models for the best-performing pipeline, and systematically study the impact of music genre and language model on the performance. With such genre-based approach, we explicitly model the music without removing it during acoustic modeling. The proposed approach outperforms all competing systems in the lyrics alignment and transcription tasks on well-known polyphonic test datasets. Chitralekha Gupta, Emre Yilmaz 0001, Haizhou Li 0001 |
ICASSP | 1 |
| 2020 | Automatic Leaderboard: Evaluation of Singing Quality Without a Standard ReferenceabstractAutomatic evaluation of singing quality can be done with the help of a reference singing or the digital sheet music of the song. However, such a standard reference is not always available. In this article, we propose a framework to rank a large pool of singers according to their singing quality without any standard reference. We define musically motivated absolute measures based on pitch histogram, and relative measures based on inter-singer statistics to evaluate the quality of singing attributes such as intonation, and rhythm. The absolute measures evaluate the goodness of pitch histogram specific to a singer, while the relative measures use the similarity between singers in terms of pitch, rhythm, and timbre as an indicator of singing quality. With the relative measures, we formulate the concept of veracity or truth-finding for the ranking of singing quality. We successfully validate a self-organizing approach to rank-ordering a large pool of singers. The fusion of absolute and relative measures results in an average Spearman's rank correlation of 0.71 with human judgments in a 10-fold cross-validation experiment, which is close to the inter-judge correlation. Chitralekha Gupta, Haizhou Li 0001, Ye Wang 0007 |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2019 | Automatic Lyrics-to-audio Alignment on Polyphonic Music Using Singing-adapted Acoustic ModelsabstractLyrics-to-audio alignment is to automatically align the lyrical words with the mixed singing audio (singing voice+musical accompaniment). Such alignment can be achieved with an automatic speech recognition (ASR) system. We propose to adapt the acoustic model of a speech recognizer towards solo singing voice. This avoids the hurdles of annotating a large polyphonic music training dataset. Moreover, a lexicon-modification based duration modelling has been incorporated to account for the long duration vowels in singing. As practical application demand the alignment on polyphonic music, we study the effect of different singing vocal separation methods in the task of lyrics-to-audio alignment in polyphonic music. The extracted vocals are forced-aligned with the singing-adapted models. We demonstrate that the use of audio source separation method and effective end-pointing of the songs has a high impact on the alignment performance through the experiments. We report a mean average absolute error of 3.87 seconds, which is comparable with the state-of-the-art lyrics-to-audio alignment system that is trained on a large polyphonic music database. Bidisha Sharma, Chitralekha Gupta, Haizhou Li 0001, Ye Wang 0007 |
ICASSP | 2 |
| 2019 | NUS Speak-to-Sing: A Web Platform for Personalized Speech-to-Singing Conversion
Chitralekha Gupta, Karthika Vijayan, Bidisha Sharma, Xiaoxue Gao, Haizhou Li 0001 |
INTERSPEECH | 1 |
| 2019 | Acoustic Modeling for Automatic Lyrics-to-Audio AlignmentabstractAutomatic lyrics to polyphonic audio alignment is a challenging task not only because the vocals are corrupted by background music, but also there is a lack of annotated polyphonic corpus for effective acoustic modeling. In this work, we propose (1) using additional speech and music-informed features and (2) adapting the acoustic models trained on a large amount of solo singing vocals towards polyphonic music using a small amount of in-domain data. Incorporating additional information such as voicing and auditory features together with conventional acoustic features aims to bring robustness against the increased spectro-temporal variations in singing vocals. By adapting the acoustic model using a small amount of polyphonic audio data, we reduce the domain mismatch between training and testing data. We perform several alignment experiments and present an in-depth alignment error analysis on acoustic features, and model adaptation techniques. The results demonstrate that the proposed strategy provides a significant error reduction of word boundary alignment over comparable existing systems, especially on more challenging polyphonic data with long-duration musical interludes. Chitralekha Gupta, Emre Yilmaz 0001, Haizhou Li 0001 |
INTERSPEECH | 1 |
| 2018 | Automatic Pronunciation Evaluation of Singing
Chitralekha Gupta, Haizhou Li 0001, Ye Wang 0007 |
INTERSPEECH | 1 |