Laxmi Pandey

dblp:170/0699 · DBLP profile ↗
← Back
12ranked-venue papers
9as first author
6since 2021 · last 2025
0000-0001-8523-7741ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Human-computer interaction and ubiquitous computing · 8 · 7 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-author · 2 since 2021Artificial intelligence and machine learning · 3 · 2 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2025 A Domain Adaptation Framework for Speech Recognition Systems with Only Synthetic data
abstract
We introduce DAS (Domain Adaptation with Synthetic data), a novel domain adaptation framework for pre-trained ASR model, designed to efficiently adapt to various language-defined domains without requiring any real data. In particular, DAS first prompts large language models (LLMs) to generate domain-specific texts before converting these texts to speech via text-to-speech technology. The synthetic data is used to finetune Whisper with Low-Rank Adapters (LoRAs) for targeted domains such as music, weather, and sports. We introduce a novel one-pass decoding strategy that merges predictions from multiple LoRA adapters efficiently during the auto-regressive text generation process. Experimental results show significant improvements, reducing the Word Error Rate (WER) by 10% to 17% across all target domains compared to the original model, with minimal performance regression in out-of-domain settings −1% on Librispeech test sets. We also demonstrate that DAS operates efficiently during inference, introducing an additional 9% increase in Real Time Factor (RTF) compared to the original model when inferring with three LoRA adapters.
Yutong Pang, Debjyoti Paul, Laxmi Pandey, Kevin Jiang, Jinxi Guo, Ke Li 0023
ICASSP4
2025 R2S: Real-to-Synthetic Representation Learning for Training Speech Recognition Models on Synthetic Data
Debjyoti Paul, Yutong Pang, Laxmi Pandey, Jinxi Guo, Ke Li 0023
INTERSPEECH4
2024 MELDER: The Design and Evaluation of a Real-time Silent Speech Recognizer for Mobile Devices
abstract
Silent speech is unaffected by ambient noise, increases accessibility, and enhances privacy and security. Yet current silent speech recognizers operate in a phrase-in/phrase-out manner, thus are slow, error prone, and impractical for mobile devices. We present MELDER, a Mobile Lip Reader that operates in real-time by splitting the input video into smaller temporal segments to process them individually. An experiment revealed that this substantially improves computation time, making it suitable for mobile devices. We further optimize the model for everyday use by exploiting the knowledge from a high-resource vocabulary using a transfer learning model. We then compare MELDER in both stationary and mobile settings with two state-of-the-art silent speech recognizers, where MELDER demonstrated superior overall performance. Finally, we compare two visual feedback methods of MELDER with the visual feedback method of Google Assistant. The outcomes shed light on how these proposed feedback methods influence users’ perceptions of the model’s performance.
Laxmi Pandey, Ahmed Sabbir Arif
CHI1
2022 Design and Evaluation of a Silent Speech-Based Selection Method for Eye-Gaze Pointing
abstract
We investigate silent speech as a hands-free selection method in eye-gaze pointing. We first propose a stripped-down image-based model that can recognize a small number of silent commands almost as fast as state-of-the-art speech recognition models. We then compare it with other hands-free selection methods (dwell, speech) in a Fitts' law study. Results revealed that speech and silent speech are comparable in throughput and selection time, but the latter is significantly more accurate than the other methods. A follow-up study revealed that target selection around the center of a display is significantly faster and more accurate, while around the top corners and the bottom are slower and error prone. We then present a method for selecting menu items with eye-gaze and silent speech. A study revealed that it significantly reduces task completion time and error rate.
Laxmi Pandey, Ahmed Sabbir Arif
Proc. ACM Hum. Comput. Interact.1
2021 LipType: A Silent Speech Recognizer Augmented with an Independent Repair Model
abstract
Speech recognition is unreliable in noisy places, compromises privacy and security when around strangers, and inaccessible to people with speech disorders. Lip reading can mitigate many of these challenges but the existing silent speech recognizers for lip reading are error prone. Developing new recognizers and acquiring new datasets is impractical for many since it requires enormous amount of time, effort, and other resources. To address these, first, we develop LipType, an optimized version of LipNet for improved speed and accuracy. We then develop an independent repair model that processes video input for poor lighting conditions, when applicable, and corrects potential errors in output for increased accuracy. We then test this model with LipType and other speech and silent speech recognizers to demonstrate its effectiveness.
Laxmi Pandey, Ahmed Sabbir Arif
CHI1
2021 Acceptability of Speech and Silent Speech Input Methods in Private and Public
abstract
Silent speech input converts non-acoustic features like tongue and lip movements into text. It has been demonstrated as a promising input method on mobile devices and has been explored for a variety of audiences and contexts where the acoustic signal is unavailable (e.g., people with speech disorders) or unreliable (e.g., noisy environment). Though the method shows promise, very little is known about peoples’ perceptions regarding using it. In this work, first, we conduct two user studies to explore users’ attitudes towards the method with a particular focus on social acceptance and error tolerance. Results show that people perceive silent speech as more socially acceptable than speech input and are willing to tolerate more errors with it to uphold privacy and security. We then conduct a third study to identify a suitable method for providing real-time feedback on silent speech input. Results show users find an abstract feedback method effective and significantly more private and secure than a commonly used video feedback method.
Laxmi Pandey, Khalad Hasan, Ahmed Sabbir Arif
CHI1
2020 Enabling Predictive Number Entry and Editing on Touchscreen-Based Mobile Devices
abstract
We propose a method for accommodating predictive number entry and editing on mobile devices via the suggestion bar of a virtual keyboard. We developed a simple predictive system to demonstrate the benefit of this method. It utilizes text-based querying and regular expression to suggest the most probable next numeric actions in the suggestion bar along with word suggestions. We evaluated this method in two user studies. The first explored number entry and the second explored editing. Results revealed that the proposed method significantly increases number entry and editing speed and accuracy. It reduces the number of actions needed per task. It also significantly reduces the time and effort needed to fix errors. Subjective analysis revealed that almost all participants found the method faster, more reliable, and easier to use than the conventional method, thus wanted to keep using it on their mobile devices.
Laxmi Pandey, Azar Alizadeh, Ahmed Sabbir Arif
CHIIR1
2020 Enabling Text Translation Using the Suggestion Bar of a Virtual Keyboard
abstract
This work augments novel text translation features to the suggestion bar of a virtual keyboard to facilitate fast and easy translation on mobile devices. The method was evaluated in two user studies. In the first study, native Hindi and Mandarin speakers exchanged text messages in each other's language using the method. All participants found the method fast and easy, the quality and the flow of the conversation satisfactory, and wanted to use it frequently for multilingual and polyglot texting. In the second study, participants performed various translation tasks using the proposed method and the default Google keyboard's translation feature. Results revealed that the proposed method was significantly faster and required fewer actions for translation tasks. Further, most participants found the method more effective and user-friendly.
Laxmi Pandey, Ahmed Sabbir Arif
SMC1
2019 TiltWriter: design and evaluation of a no-touch tilt-based text entry method for handheld devices
abstract
Touch is the predominant method of text entry on mobile devices. However, these devices also facilitate tilt-based input using accelerometers and gyroscopes. This paper presents the design and evaluation of TiltWriter, a non-touch, tilt-based text entry technique. TiltWriter aims to supplement conventional techniques when precise touch is not convenient or possible (e.g., the user lacks sufficient motor skills). Two keyboard layouts were designed and evaluated: Qwerty, and "Custom", a layout inspired by the telephone keypad. Novice participants in a longitudinal study achieved speeds of 12.1 wpm for Qwerty, 10.7 wpm for Custom. Error rate averaged 0.76% for Qwerty, 0.62% for Custom. A post-study extended session yielded 15.2 wpm with Custom, versus 11.5 wpm with Qwerty. Results and participant feedback suggest that a selection dwell time between 700 - 800 ms benefits accuracy.
Steven J. Castellucci, I. Scott MacKenzie, Mudit Misra, Laxmi Pandey, Ahmed Sabbir Arif
MUM4
2019 Context-sensitive app prediction on the suggestion bar of a mobile keyboard
abstract
This work augments context-sensitive app prediction feature to the suggestion bar of a mobile virtual keyboard to accommodate fast and easy information acquisition and sharing in textual conversations. The purpose is to eliminate the need for switching between apps while typing. A user study revealed that the proposed method improves performance both in terms of speed and effort for common tasks, such as playing a song or finding and sharing the address of a restaurant. Post-study questionnaire revealed that all participants found the method fast, easy, and likely to facilitate more engaging and meaningful textual conversations. All wanted to keep using it on their mobile devices.
Laxmi Pandey, Ahmed Sabbir Arif
MUM1
2018 Monoaural Audio Source Separation Using Variational Autoencoders
Laxmi Pandey, Anurendra Kumar, Vinay P. Namboodiri
INTERSPEECH1
2018 LSTM Based Attentive Fusion of Spectral and Prosodic Information for Keyword Spotting in Hindi Language
Laxmi Pandey, Karan Nathwani
INTERSPEECH1