Tanmay Srivastava

dblp:274/6319 · DBLP profile ↗
← Back
8ranked-venue papers
6as first author
8since 2021 · last 2024
0000-0003-0144-7931ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 5 · 5 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2024 Whispering Wearables: Multimodal Approach to Silent Speech Recognition with Head-Worn Devices
abstract
Silent speech recognition has emerged as a promising approach for enabling hands-free and discreet interaction with head-worn devices. In this paper, we present QuietSync, a multimodal system that combines inertial measurement unit (IMU) and contact electrode (ExG) signals to achieve accurate silent speech recognition using off-the-shelf devices. QuietSync utilizes an IMU attached to the lower part of the headphones near the ear and strategically places ExG electrodes on the headphones, glasses (nose and behind the ear), and face (for VR applications) to capture subtle movements and muscle activity associated with silent speech production. We conducted a user study with 9 participants and successfully recognized 12 commands with an accuracy of 94.2%. Our system leverages the complementary nature of IMU and ExG signals to enhance the robustness and reliability of silent speech recognition. The IMU captures subtle movements of the jaw and facial muscles, while the ExG electrodes detect low-amplitude surface muscle activity associated with speech production. We show that our system is not affected by the length and speech mannerisms of the commands, and can be fine-tuned for users of varied native languages with only 5 samples. Our findings demonstrate the feasibility of using off-the-shelf head-worn devices to enable silent speech recognition, opening up new possibilities for seamless and discreet interaction with devices such as VR/AR headsets and earables. To the best of our knowledge, QuietSync is the first system to enable silent speech interaction for multiple form factors.
Tanmay Srivastava, R. Michael Winters, Thomas M. Gable, Yu-Te Wang, Teresa LaScala, Ivan Tashev
ICMI1
2024 Enabling Accessible and Ubiquitous Interaction in Next-Generation Wearables: An Unvoiced Speech Approach
abstract
As wearable devices increase, there's a growing need for intuitive, private, and accessible interaction methods. This position paper builds on the research on unvoiced speech interaction and authentication to propose a vision for interaction in next-generation wearables. This paper draws upon our previous work on unvoiced speech interfaces that leverage jaw movements and facial vibrations for command recognition and user authentication. We argue that unvoiced speech interaction can provide a robust, privacy-preserving, and noise-resistant alternative to traditional interfaces, enhancing accessibility and offering discrete interaction in public spaces. We discuss the potential integration of these systems into commercial devices and explore gesture-based interactions as an alternative to touch. Additionally, we discuss the future direction of unvoiced speech interfaces. This paper sets the stage for implementing unvoiced speech and gesture-based interaction in mainstream wearables in our daily interactions with technology.
Tanmay Srivastava, Prerna Khanna, Shijia Pan, V. P. Nguyen, Shubham Jain 0003
MobiCom1
2024 Unvoiced: Designing an LLM-assisted Unvoiced User Interface using Earables
abstract
We present Unvoiced, a novel unvoiced user interface that leverages jaw motion to enable users to silently interact with their devices using earables. The core idea is to translate low-frequency jaw motion signals into high-frequency information-rich mel spectrograms. Our proposed cross-modal translation incorporates phonetic, contextual, and syntactic information, while the specialized loss function optimizes for these linguistic features. This ensures that the generated spectrograms capture nuanced speech characteristics. Evaluated for 19 users across four tasks, Unvoiced demonstrates >94% task completion rate and <9% word error rate for over 90% of phrases. Further, Unvoiced maintains >90% task completion rate in noisy conditions.
Tanmay Srivastava, Prerna Khanna, Shijia Pan, Phuc Nguyen 0002, Shubham Jain 0003
SenSys1
2024 Poster Unvoiced: Designing an Unvoiced User Interface using Earables and LLMs
abstract
This poster presents the design and implementation of Unvoiced, a silent speech interaction system. Unvoiced transforms subtle jaw movements into rich speech spectrograms, enabling seamless and private device interaction. Our system captures low-frequency jaw motion signals using ear-worn IMUs and translates them into high-fidelity mel-spectrograms through cross-modal translation techniques. By incorporating phonetic, contextual, and syntactic information, Unvoiced generates high-fidelity spectrograms that existing speech recognition systems can process. In our evaluation with 19 users across four common tasks, Unvoiced achieved a remarkable >94% task completion rate and <9% Word Error Rate (WER) for over 90% of phrases, maintaining robust performance even in noisy conditions.
Tanmay Srivastava, Prerna Khanna, Shijia Pan, Phuc Nguyen 0002, Shubham Jain 0003
SenSys1
2023 Jawthenticate: Microphone-free Speech-based Authentication using Jaw Motion and Facial Vibrations
abstract
In this paper, we present Jawthenticate, an earable system that authenticates a user using audible or inaudible speech without using a microphone. This system can overcome the shortcomings of traditional voice-based authentication systems like unreliability in noisy conditions and spoofing using microphone-based replay attacks. Jawthenticate derives distinctive speech-related features from the jaw motion and associated facial vibrations. This combination of features makes Jawthenticate resilient to vocal imitations as well as camera-based spoofing. We use these features to train a two-class SVM classifier for each user. Our system is invariant to the content and language of speech. In a study conducted with 41 subjects, who speak different native languages, Jawthenticate achieves a Balanced Accuracy (BAC) of 97.07%, True Positive Rate (TPR) of 97.75%, and True Negative Rate (TNR) of 96.4% with just 3 seconds of speech data.
Tanmay Srivastava, Shijia Pan, Phuc Nguyen 0002, Shubham Jain 0003
SenSys1
2023 SpiroMask: Measuring Lung Function Using Consumer-Grade Masks
abstract
According to the World Health Organisation (WHO), 235 million people suffer from respiratory illnesses which causes four million deaths annually. Regular lung health monitoring can lead to prognoses about deteriorating lung health conditions. This article presents our system SpiroMask that retrofits a microphone in consumer-grade masks (N95 and cloth masks) for continuous lung health monitoring. We evaluate our approach on 48 participants (including 14 with lung health issues) and find that we can estimate parameters such as lung volume and respiration rate within the approved error range by the American Thoracic Society (ATS). Further, we show that our approach is robust to sensor placement inside the mask.
Rishiraj Adhikary, Dhruvi Lodhavia, Chris Francis, Rohit Patil, Tanmay Srivastava, Prerna Khanna, Nipun Batra 0001, Joseph Breda, Jacob Peplinski, Shwetak N. Patel
ACM Trans. Comput. Heal.5
2022 Leveraging earables for unvoiced command recognition
abstract
We demonstrate an ear-worn technology that recognizes unvoiced human commands by tracking jaw motion. The ear-worn system is designed to achieve continual unvoiced command recognition for robust human-computer interaction (HCI) applications. First, the system reliably extracts the jaw motion signals buried under the noise caused by head motion, walking, and other motion artifacts to track single secondary voice articulator (i.e., word). Then, learning from linguistics and human speech anatomy, we design a novel algorithm that localizes the phonemes in the command, and reconstructs the word. We evaluate the proposed system in real-world experiments with 15 volunteers. Our preliminary results show that the proposed system obtains a word recognition accuracy of 95.6% in noise-free conditions and 93.2% and 91.6%, while head nodding and walking.
Tanmay Srivastava, Prerna Khanna, Shijia Pan, Phuc Nguyen 0002, Shubham Jain 0003
MobiSys1
2021 Vartalaap: What Drives #AirQuality Discussions: Politics, Pollution or Pseudo-science?
abstract
Air pollution is a global challenge for cities across the globe. Understanding the public perception of air pollution can help policymakers engage better with the public and appropriately introduce policies. Accurate public perception can also help people to identify the health risks of air pollution and act accordingly. Unfortunately, current techniques for determining perception are not scalable: it involves surveying few hundred people with questionnaire-based surveys. Using the advances in natural language processing (NLP), we propose a more scalable solution called Vartalaap to gauge public perception of air pollution via the microblogging social network Twitter. We curated a dataset of more than 1.2M tweets discussing Delhi-specific air pollution. We find that (unfortunately) the public is supportive of unproven mitigation strategies to reduce pollution, thus risking their health due to a false sense of security. We also find that air quality is a year-long problem, but the discussions are not proportional to the level of pollution and spike up when pollution is more visible. The information required by Vartalaap is publicly available and, as such, it can be immediately applied to study different societal issues across the world.
Rishiraj Adhikary, Zeel B. Patel, Tanmay Srivastava, Nipun Batra 0001, Mayank Singh 0001, Udit Bhatia, Sarath Guttikunda
Proc. ACM Hum. Comput. Interact.3