EDBT 2026 Demo / reviewers in the wild / expert
Rahul Gupta 0001
dblp:24/3213-1
· DBLP profile ↗
63ranked-venue papers
17as first author
32since 2021 · last 2026
0000-0002-9277-3718ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 45 · 12 first-author · 25 since 2021Graphics, computer vision, multimedia, augmented reality and games · 36 · 13 first-author · 8 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ARES: Adaptive Red-Teaming and End-to-End Repair of Policy-Reward SystemabstractJiacheng Liang, Yao Ma, Tharindu Kumarage, Satyapriya Krishna, Rahul Gupta, Kai-Wei Chang, Aram Galstyan, Charith Peris. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Jiacheng Liang, Tharindu Kumarage, Satyapriya Krishna, Rahul Gupta 0001, Kai-Wei Chang 0001, Aram Galstyan, Charith Peris |
ACL (1) | 5 |
| 2026 | SWAN: Semantic Watermarking with Abstract Meaning RepresentationabstractZiping Ye, Gourab Dey, Christos Christodoulopoulos, Charith Peris, Anil Ramakrishna, Weitong Ruan, Aram Galstyan, Kai-Wei Chang, Rahul Gupta, Ninareh Mehrabi. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Ziping Ye, Gourab Dey, Christos Christodoulopoulos 0001, Charith Peris, Anil Ramakrishna, Weitong Ruan, Aram Galstyan, Kai-Wei Chang 0001, Rahul Gupta 0001, Ninareh Mehrabi |
ACL (1) | 9 |
| 2025 | Diagnosing Memorization in Chain-of-Thought Reasoning, One Token at a TimeabstractLarge Language Models (LLMs) perform well on reasoning benchmarks but often fail when inputs alter slightly, raising concerns about the extent to which their success relies on memorization.This issue is especially acute in Chain-of-Thought (CoT) reasoning, where spurious memorized patterns can trigger intermediate errors that cascade into incorrect final answers.We introduce STIM, a novel framework for Source-aware Token-level Identification of Memorization, which attributes each token in a reasoning chain to one of multiple memorization sources -local, mid-range, or long-rangebased on their statistical co-occurrence with the token in the pretraining corpus.Our token-level analysis across tasks and distributional settings reveals that models rely more on memorization in complex or long-tail cases, and that local memorization is often the dominant driver of errors, leading to up to 67% of wrong tokens.We also show that memorization scores from STIM can be effective in predicting the wrong tokens in the wrong reasoning step.STIM offers a powerful tool for diagnosing and improving model reasoning and can generalize to other structured step-wise generation tasks. 1 * Equal Contribution † Work done in Amazon AGI 1 https://github.com/INK-USC/STIMModel Reasoning Steps: Rick killed 10 + 15 = 25 animals last night.Today Rick killed 3 * 3 = 9 wolves.Today Rick killed 15 -3 = 12 cougars.Today Rick killed 9 + 12 = 21 animals.Total animals killed by Rick is 25 + 21 = 46.So the answer is 46. Examples of Simultaneous Memorization Sources of Different StrengthQuestion: Here is a list: [pitanga, pitanga, yumberry, yumberry, pitanga, yumberry, pitanga, yumberry, pitanga, pitanga].How many times does 'pitanga' appear on it?Model Output: 1.The first element is 'pitanga'.... 10.The tenth element is 'pitanga'.Now, counting the occurrences of 'pitanga': We have 'pitanga' appearing 5 times at positions 1, 2, 5, 7, and 10.So the answer is 5.Analysis: Token 5 is incorrect and has the highest influence from local context, where the n-gram "appearing 5" has a high pretraining frequency.STIM scores are: local = 0.81 , mid = 0.24, long = 0.03, suggesting that the model was influenced by locally frequent patterns rather than actual counting.Question: What is ( 100101100 -100101100 * 0.1) * 111110100 + 100101100 * 1111 / 1100100 * 111110100 * 0.1 + (100101100 -100101100 * 0.1 -100101100 * 1111 / 1100100) * 111110100 equal to? (base-2) Model Output:To solve this, we first handle the multiplication and division in base-2, converting to base-10 as necessary:1. (100101100 × 0.1base-2 = 10010110base-2) (which is 180 in base-10) ... So the answer is 26807.536.Analysis: Token 180 is incorrect and has the highest influence from input tokens, having higher pretraining co-occurrence frequency with "111110100", "*" and "100101100" in the input.STIM scores are local=-0.19,mid=0.09,long=0.156, indicating long-range memorization being the primary influence. Huihan Li 0001, Ninareh Mehrabi, Rahul Gupta 0001, Xiang Ren 0001 |
EMNLP | 6 |
| 2025 | Certifying Counterfactual Bias in LLMsabstractLarge Language Models (LLMs) can produce biased responses that can cause representational harms. However, conventional studies are insufficient to thoroughly
evaluate biases across LLM responses for different demographic groups (a.k.a.
counterfactual bias), as they do not scale to large number of inputs and do not
provide guarantees. Therefore, we propose the first framework, LLMCert-B that
certifies LLMs for counterfactual bias on distributions of prompts. A certificate
consists of high-confidence bounds on the probability of unbiased LLM responses
for any set of counterfactual prompts - prompts differing by demographic groups,
sampled from a distribution. We illustrate counterfactual bias certification for
distributions of counterfactual prompts created by applying prefixes sampled from
prefix distributions, to a given set of prompts. We consider prefix distributions consisting random token sequences, mixtures of manual jailbreaks, and perturbations
of jailbreaks in LLM’s embedding space. We generate non-trivial certificates for
SOTA LLMs, exposing their vulnerabilities over distributions of prompts generated
from computationally inexpensive prefix distributions. Isha Chaudhary, Manoj Kumar 0007, Morteza Ziyadi, Rahul Gupta 0001, Gagandeep Singh 0001 |
ICLR | 5 |
| 2025 | VMDT: Decoding the Trustworthiness of Video Foundation ModelsabstractAs foundation models become more sophisticated, ensuring their trustworthiness becomes increasingly critical; yet, unlike text and image, the video modality still lacks comprehensive trustworthiness benchmarks. We introduce VMDT (Video-Modal DecodingTrust), the first unified platform for evaluating text-to-video (T2V) and video-to-text (V2T) models across five key trustworthiness dimensions: safety, hallucination, fairness, privacy, and adversarial robustness. Through our extensive evaluation of 7 T2V models and 19 V2T models using VMDT, we uncover several significant insights. For instance, all open-source T2V models evaluated fail to recognize harmful queries and often generate harmful videos, while exhibiting higher levels of unfairness compared to image modality models. In V2T models, unfairness and privacy risks rise with scale, whereas hallucination and adversarial robustness improve---though overall performance remains low. Uniquely, safety shows no correlation with model size, implying that factors other than scale govern current safety levels. Our findings highlight the urgent need for developing more robust and trustworthy video foundation models, and VMDT provides a systematic framework for measuring and tracking progress toward this goal. The code is available at https://sunblaze-ucb.github.io/VMDT-page/. Yujin Potter, Zhun Wang, Nicholas Crispino, Kyle Montgomery, Alexander Xiong, Ethan Y. Chang, Francesco Pinto, Rahul Gupta 0001, Morteza Ziyadi, Christos Christodoulopoulos 0001, Bo Li 0026, Chenguang Wang 0001, Dawn Song |
NeurIPS | 9 |
| 2025 | Establishing Best Practices in Building Rigorous Agentic BenchmarksabstractBenchmarks are essential for quantitatively tracking progress in AI. As AI agents become increasingly capable, researchers and practitioners have introduced agentic benchmarks to evaluate agents on complex, real-world tasks. These benchmarks typically measure agent capabilities by evaluating task outcomes via specific reward designs. However, we show that many agentic benchmarks have issues in task setup or reward design. For example, SWE-bench-Verified uses insufficient test cases, while $\tau$-bench counts empty responses as successes. Such issues can lead to under- or overestimation of agents’ performance by up to 100% in relative terms. To make agentic evaluation rigorous, we introduce the Agentic Benchmark Checklist (ABC), a set of guidelines that we synthesized from our benchmark-building experience, a survey of best practices, and previously reported issues. When applied to CVE-Bench, a benchmark with a particularly complex evaluation design, ABC reduces performance overestimation by 33%. Yuxuan Zhu 0003, Tengjun Jin, Yada Pruksachatkun, Andy Zhang, Sasha Cui, Sayash Kapoor, Shayne Longpre, Kevin Meng, Rebecca Weiss, Fazl Barez, Rahul Gupta 0001, Jwala Dhamala, Jacob Merizian, Mario Giulianelli, Harry Coppock, Cozmin Ududec, Antony Kellermann, Jasjeet S. Sekhon, Jacob Steinhardt, Sarah Schwettmann, Arvind Narayanan, Matei Zaharia, Ion Stoica, Percy Liang, Daniel Kang 0001 |
NeurIPS | 12 |
| 2024 | Tree-of-Traversals: A Zero-Shot Reasoning Algorithm for Augmenting Black-box Language Models with Knowledge GraphsabstractElan Markowitz, Anil Ramakrishna, Jwala Dhamala, Ninareh Mehrabi, Charith Peris, Rahul Gupta, Kai-Wei Chang, Aram Galstyan. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Elan Markowitz, Anil Ramakrishna, Jwala Dhamala, Ninareh Mehrabi, Charith Peris, Rahul Gupta 0001, Kai-Wei Chang 0001, Aram Galstyan |
ACL (1) | 6 |
| 2024 | FLIRT: Feedback Loop In-context Red TeamingabstractNinareh Mehrabi, Palash Goyal, Christophe Dupuy, Qian Hu, Shalini Ghosh, Richard Zemel, Kai-Wei Chang, Aram Galstyan, Rahul Gupta. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Ninareh Mehrabi, Palash Goyal, Christophe Dupuy, Shalini Ghosh, Richard S. Zemel, Kai-Wei Chang 0001, Aram Galstyan, Rahul Gupta 0001 |
EMNLP | 9 |
| 2024 | Data Advisor: Dynamic Data Curation for Safety Alignment of Large Language ModelsabstractData is a crucial element in large language model (LLM) alignment.Recent studies have explored using LLMs for efficient data collection.However, LLM-generated data often suffers from quality issues, with underrepresented or absent aspects and low-quality datapoints.To address these problems, we propose DATA ADVISOR, an enhanced LLMbased method for generating data that takes into account the characteristics of the desired dataset.Starting from a set of pre-defined principles in hand, DATA ADVISOR monitors the status of the generated data, identifies weaknesses in the current dataset, and advises the next iteration of data generation accordingly.DATA ADVISOR can be easily integrated into existing data generation methods to enhance data quality and coverage.Experiments on safety alignment of three representative LLMs (i.e., Mistral, Llama2, and Falcon) demonstrate the effectiveness of DATA ADVISOR in enhancing model safety against various fine-grained safety issues without sacrificing model utility.Warning: this paper contains example data that may be offensive or harmful. Fei Wang 0060, Ninareh Mehrabi, Palash Goyal, Rahul Gupta 0001, Kai-Wei Chang 0001, Aram Galstyan |
EMNLP | 4 |
| 2024 | The steerability of large language models toward data-driven personasabstractJunyi Li, Charith Peris, Ninareh Mehrabi, Palash Goyal, Kai-Wei Chang, Aram Galstyan, Richard Zemel, Rahul Gupta. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Junyi Li 0002, Charith Peris, Ninareh Mehrabi, Palash Goyal, Kai-Wei Chang 0001, Aram Galstyan, Richard S. Zemel, Rahul Gupta 0001 |
NAACL-HLT | 8 |
| 2024 | Toward Informal Language Processing: Knowledge of Slang in Large Language ModelsabstractZhewei Sun, Qian Hu, Rahul Gupta, Richard Zemel, Yang Xu. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Zhewei Sun, Rahul Gupta 0001, Richard S. Zemel, Yang Xu 0023 |
NAACL-HLT | 3 |
| 2023 | Resolving Ambiguities in Text-to-Image Generative ModelsabstractNinareh Mehrabi, Palash Goyal, Apurv Verma, Jwala Dhamala, Varun Kumar, Qian Hu, Kai-Wei Chang, Richard Zemel, Aram Galstyan, Rahul Gupta. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Ninareh Mehrabi, Palash Goyal, Apurv Verma, Jwala Dhamala, Kai-Wei Chang 0001, Richard S. Zemel, Aram Galstyan, Rahul Gupta 0001 |
ACL (1) | 10 |
| 2023 | Multi-VALUE: A Framework for Cross-Dialectal English NLPabstractDialect differences caused by regional, social, and economic factors cause performance discrepancies for many groups of language technology users.Inclusive and equitable language technology must critically be dialect invariant, meaning that performance remains constant over dialectal shifts.Current systems often fall short of this ideal since they are designed and tested on a single dialect: Standard American English (SAE).We introduce a suite of resources for evaluating and achieving English dialect invariance.The resource is called Multi-VALUE, a controllable rule-based translation system spanning 50 English dialects and 189 unique linguistic features.Multi-VALUE maps SAE to synthetic forms of each dialect.First, we use this system to stress tests question answering, machine translation, and semantic parsing.Stress tests reveal significant performance disparities for leading models on nonstandard dialects.Second, we use this system as a data augmentation technique to improve the dialect robustness of existing systems.Finally, we partner with native speakers of Chicano and Indian English to release new goldstandard variants of the popular CoQA task.To execute the transformation code, run model checkpoints, and download both synthetic and gold-standard dialectal benchmark datasets, see http://value-nlp.org/. Caleb Ziems, William Barr Held, Jingfeng Yang 0001, Jwala Dhamala, Rahul Gupta 0001, Diyi Yang |
ACL (1) | 5 |
| 2023 | Faithful Model Evaluation for Model-Based MetricsabstractStatistical significance testing is used in natural language processing (NLP) to determine whether the results of a study or experiment are likely to be due to chance or if they reflect a genuine relationship.A key step in significance testing is the estimation of confidence interval which is a function of sample variance.Sample variance calculation is straightforward when evaluating against ground truth.However, in many cases, a metric model is often used for evaluation.For example, to compare toxicity of two large language models, a toxicity classifier is used for evaluation.Existing works usually do not consider the variance change due to metric model errors, which can lead to wrong conclusions.In this work, we establish the mathematical foundation of significance testing for model-based metrics.With experiments on public benchmark datasets and a production system, we show that considering metric model errors to calculate sample variances for model-based metrics changes the conclusions in certain experiments. Palash Goyal, Rahul Gupta 0001 |
EMNLP | 3 |
| 2023 | Evaluating Large Language Models on Controlled Generation TasksabstractJiao Sun, Yufei Tian, Wangchunshu Zhou, Nan Xu, Qian Hu, Rahul Gupta, John Wieting, Nanyun Peng, Xuezhe Ma. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023. Jiao Sun, Yufei Tian, Wangchunshu Zhou, Rahul Gupta 0001, John Wieting, Nanyun Peng 0001, Xuezhe Ma |
EMNLP | 6 |
| 2023 | Quantifying Catastrophic Forgetting in Continual Federated LearningabstractThe deployment of Federated Learning (FL) systems poses various challenges such as data heterogeneity and communication efficiency. We focus on a practical FL setup that has recently drawn attention, where the data distribution on each device is not static but dynamically evolves over time. This setup, referred to as Continual Federated Learning (CFL), suffers from catastrophic forgetting, i.e., the undesired forgetting of previous knowledge after learning on new data, an issue not encountered with vanilla FL. In this work, we formally quantify catastrophic forgetting in a CFL setup, establish links to training optimization and evaluate different episodic replay approaches for CFL on a large scale real-world NLP dataset. To the best of our knowledge, this is the first such study of episodic replay for CFL. We show that storing a small set of past data boosts performance and significantly reduce forgetting, providing evidence that carefully designed sampling strategies can lead to further improvements. Christophe Dupuy, Jimit Majmudar, Jixuan Wang, Tanya G. Roosta, Rahul Gupta 0001, Clement Chung, Jie Ding 0002, Amir Salman Avestimehr |
ICASSP | 5 |
| 2023 | Self-Healing Through Error Detection, Attribution, and RetrainingabstractNegative feedback received from users of voice agents can provide valuable training signal to their underlying ML systems. However, such systems tend to have complex inference pipelines consisting of multiple model-based and deterministic components. Therefore, when negative feedback is received, it can be difficult to attribute the system error to a specific sub-component. In this work, we address this challenge by building a system for error attribution and correction. We prototype attributing errors to the ML models used for do-main classification (DC) in the NLU component of an assistant’s pipeline, using a combination of a model and rule based system. We propose a simple method to add these detected errors directly to offline DC model training, and study our system’s effectiveness on a challenging test set of low-frequency utterances. Our experiments on nine domains suggest that augmenting DC training data with our method significantly improves performance on a majority of them. Ansel MacLaughlin, Anna Rumshisky, Rinat Khaziev, Anil Ramakrishna, Yuval Merhav, Rahul Gupta 0001 |
ICASSP | 6 |
| 2023 | Sampling bias in NLU models: Impact and Mitigation
Zefei Li, Anil Ramakrishna, Anna Rumshisky, Andy Rosenbaum, Saleh Soltan, Rahul Gupta 0001 |
INTERSPEECH | 6 |
| 2023 | FedMultimodal: A Benchmark for Multimodal Federated LearningabstractOver the past few years, Federated Learning (FL) has become an emerging machine learning technique to tackle data privacy challenges through collaborative training. In the Federated Learning algorithm, the clients submit a locally trained model, and the server aggregates these parameters until convergence. Despite significant efforts that have been made to FL in fields like computer vision, audio, and natural language processing, the FL applications utilizing multimodal data streams remain largely unexplored. It is known that multimodal learning has broad real-world applications in emotion recognition, healthcare, multimedia, and social media, while user privacy persists as a critical concern. Specifically, there are no existing FL benchmarks targeting multimodal applications or related tasks. In order to facilitate the research in multimodal FL, we introduce FedMultimodal, the first FL benchmark for multimodal learning covering five representative multimodal applications from ten commonly used datasets with a total of eight unique modalities. FedMultimodal offers a systematic FL pipeline, enabling end-to-end modeling framework ranging from data partition and feature extraction to FL benchmark algorithms and model evaluation. Unlike existing FL benchmarks, FedMultimodal provides a standardized approach to assess the robustness of FL against three common data corruptions in real-life multimodal applications: missing modalities, missing labels, and erroneous labels. We hope that FedMultimodal can accelerate numerous future research directions, including designing multimodal FL algorithms toward extreme data heterogeneity, robustness multimodal FL, and efficient multimodal FL. The datasets and benchmark results can be accessed at: https://github.com/usc-sail/fed-multimodal. Tiantian Feng, Digbalay Bose, Rajat Hebbar, Anil Ramakrishna, Rahul Gupta 0001, Mi Zhang 0002, Amir Salman Avestimehr, Shri Narayanan |
KDD | 6 |
| 2023 | Incorporating Fairness in Large Scale NLU SystemsabstractNLU models power several user facing experiences such as conversations agents and chat bots. Building NLU models typically consist of 3 stages: a) building or finetuning a pre-trained model b) distilling or fine-tuning the pre-trained model to build task specific models and, c) deploying the task-specific model to production. In this presentation, we will identify fairness considerations that can be incorporated in the aforementioned three stages in the life-cycle of NLU model building: (i) selection/building of a large scale language model, (ii) distillation/fine-tuning the large model into task specific model and, (iii) deployment of the task specific model. We will present select metrics that can be used to quantify fairness in NLU models and fairness enhancement techniques that can be deployed in each of these stages. Finally, we will share some recommendations to successfully implement fairness considerations when building an industrial scale NLU system. Rahul Gupta 0001, Lisa Bauer, Kai-Wei Chang 0001, Jwala Dhamala, Aram Galstyan, Palash Goyal, Avni Khatri, Rohit Parimi, Charith Peris, Apurv Verma, Richard S. Zemel, Premkumar Natarajan |
WSDM | 1 |
| 2023 | Privacy in the Time of Language ModelsabstractPretrained large language models (LLMs) have consistently shown state-of-the-art performance across multiple natural language processing (NLP) tasks. These models are of much interest for a variety of industrial applications that use NLP as a core component. However, LLMs have also been shown to memorize portions of their training data, which can contain private information. Therefore, when building and deploying LLMs, it is of value to apply privacy-preserving techniques that protect sensitive data. Charith Peris, Christophe Dupuy, Jimit Majmudar, Rahil Parikh, Sami Smaili, Richard S. Zemel, Rahul Gupta 0001 |
WSDM | 7 |
| 2022 | Measuring Fairness of Text Classifiers via Prediction SensitivityabstractSatyapriya Krishna, Rahul Gupta, Apurv Verma, Jwala Dhamala, Yada Pruksachatkun, Kai-Wei Chang. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Satyapriya Krishna, Rahul Gupta 0001, Apurv Verma, Jwala Dhamala, Yada Pruksachatkun, Kai-Wei Chang 0001 |
ACL (1) | 2 |
| 2022 | An Efficient DP-SGD Mechanism for Large Scale NLU ModelsabstractRecent advances in deep learning have drastically improved performance on many Natural Language Understanding (NLU) tasks. However, the data used to train NLU models may contain private information such as addresses or phone numbers, particularly when drawn from human subjects. It is desirable that underlying models do not expose private information contained in the training data. Differentially Private Stochastic Gradient Descent (DP-SGD) has been proposed as a mechanism to build privacy-preserving models. However, DP-SGD can be prohibitively slow to train. In this work, we propose a more efficient DP-SGD for training using a GPU infrastructure and apply it to fine-tuning models based on LSTM and transformer architectures. We report faster training times, alongside accuracy, theoretical privacy guarantees and success of Membership inference attacks for our models and observe that fine-tuning with proposed variant of DP-SGD can yield competitive models without significant degradation in training time and improvement in privacy protection. We also make observations such as looser theoretical ϵ, δ can translate into significant practical privacy gains. Christophe Dupuy, Radhika Arava, Rahul Gupta 0001, Anna Rumshisky |
ICASSP | 3 |
| 2022 | Learnings from Federated Learning in The Real WorldabstractFederated Learning (FL) applied to real world data may suffer from several idiosyncrasies. One such idiosyncrasy is the data distribution across devices. Data across devices could be distributed such that there are some "heavy devices" with large amounts of data while there are many "light users" with only a handful of data points. There also exists heterogeneity of data across devices. In this study, we evaluate the impact of such idiosyncrasies on Natural Language Understanding (NLU) models trained using FL. We conduct experiments on data obtained from a large scale NLU system serving thousands of devices and show that simple non-uniform device selection based on the number of interactions at each round of FL training boosts the performance of the model. This benefit is further amplified in continual FL on consecutive time periods, where non-uniform sampling manages to swiftly catch up with FL methods using all data at once. Christophe Dupuy, Tanya G. Roosta, Leo Long, Clement Chung, Rahul Gupta 0001, Amir Salman Avestimehr |
ICASSP | 5 |
| 2022 | Advin: Automatically Discovering Novel Domains and Intents from User Text UtterancesabstractRecognizing the intents and domains of users’ spoken and written language is a key component of Natural Language Understanding (NLU) systems. Real applications however encounter dynamic, rapidly evolving environments with newly emerging intents and domains, for which no labeled data or prior information is available. For such a setting, we propose a novel framework, ADVIN, to automatically discover novel domains and intents from large volumes of unlabeled text. We first employ an open classification model to discriminate all utterances potentially consisting of a novel intent. Next, we train a deep learning model with a pairwise margin loss function and knowledge transfer, to discover multiple latent intent categories in an unsupervised manner. We finally form a hierarchical intent-domain taxonomy by linking mutually related novel intents into novel domains. ADVIN significantly outperforms strong baselines on four benchmark datasets, and data from a real-world voice agent. Nikhita Vedula, Rahul Gupta 0001, Aman Alok, Mukund Sridhar, Shankar Ananthakrishnan |
ICASSP | 2 |
| 2022 | Training Mixed-Domain Translation Models via Federated LearningabstractPeyman Passban, Tanya Roosta, Rahul Gupta, Ankit Chadha, Clement Chung. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Peyman Passban, Tanya G. Roosta, Rahul Gupta 0001, Ankit Chadha, Clement Chung |
NAACL-HLT | 3 |
| 2022 | Federated Learning with Noisy User FeedbackabstractRahul Sharma, Anil Ramakrishna, Ansel MacLaughlin, Anna Rumshisky, Jimit Majmudar, Clement Chung, Salman Avestimehr, Rahul Gupta. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Anil Ramakrishna, Ansel MacLaughlin, Anna Rumshisky, Jimit Majmudar, Clement Chung, Amir Salman Avestimehr, Rahul Gupta 0001 |
NAACL-HLT | 8 |
| 2022 | An Analysis of The Effects of Decoding Algorithms on Fairness in Open-Ended Language GenerationabstractSeveral prior works have shown that language models (LMs) can generate text containing harmful social biases and stereotypes. While decoding algorithms play a central role in determining properties of LM generated text, their impact on the fairness of the generations has not been studied. We present a systematic analysis of the impact of decoding algorithms on LM fairness, and analyze the trade-off between fairness, diversity and quality. Our experiments with top-p, top-k and temperature decoding algorithms, in open-ended language generation, show that fairness across demographic groups changes significantly with change in decoding algorithm's hyper-parameters. Notably, decoding algorithms that output more diverse text also output more texts with negative sentiment and regard. We present several findings and provide recommendations on standardized reporting of decoding details in fairness evaluations and optimization of decoding algorithms for fairness alongside quality and diversity. Jwala Dhamala, Rahul Gupta 0001, Kai-Wei Chang 0001, Aram Galstyan |
SLT | 3 |
| 2022 | Joint Multi-Dimensional Model for Global and Time-Series Annotations
Anil Ramakrishna, Rahul Gupta 0001, Shri Narayanan |
IEEE Trans. Affect. Comput. | 2 |
| 2022 | Modeling Feature Representations for Affective Speech Using Generative Adversarial NetworksabstractEmotion recognition is a classic field of research with a typical setup extracting features and feeding them through a classifier for prediction. On the other hand, generative models jointly capture the distributional relationship between emotions and the feature profiles. Recently, Generative Adversarial Networks (GANs) have surfaced as a new class of generative models and have shown considerable success in modeling distributions in the fields of computer vision and natural language understanding. In this article, we experiment with variants of GAN architectures to generate feature vectors corresponding to an emotion in two ways: (i) A generator is trained with samples from a mixture prior. Each mixture component corresponds to an emotional class and can be sampled to generate features from the corresponding emotion. (ii) A one-hot vector corresponding to an emotion can be explicitly used to generate the features. We perform analysis on such models and also propose different metrics used to measure the performance of the GAN models in their ability to generate realistic synthetic samples. Apart from evaluation on a given dataset of interest, we perform a cross-corpus study where we study the utility of the synthetic samples as additional training data in low resource conditions. Saurabh Sahu, Rahul Gupta 0001, Carol Y. Espy-Wilson |
IEEE Trans. Affect. Comput. | 2 |
| 2021 | ADePT: Auto-encoder based Differentially Private Text TransformationabstractPrivacy is an important concern when building statistical models on data containing personal information.Differential privacy offers a strong definition of privacy and can be used to solve several privacy concerns (Dwork et al., 2014).Multiple solutions have been proposed for the differentially-private transformation of datasets containing sensitive information.However, such transformation algorithms offer poor utility in Natural Language Processing (NLP) tasks due to noise added in the process.In this paper, we address this issue by providing a utility-preserving differentially private text transformation algorithm using auto-encoders.Our algorithm transforms text to offer robustness against attacks and produces transformations with high semantic quality that perform well on downstream NLP tasks.We prove the theoretical privacy guarantee of our algorithm and assess its privacy leakage under Membership Inference Attacks (MIA) (Shokri et al., 2017) on models trained with transformed data.Our results show that the proposed model performs better against MIA attacks while offering lower to no degradation in the utility of the underlying transformation process compared to existing baselines. Satyapriya Krishna, Rahul Gupta 0001, Christophe Dupuy |
EACL | 2 |
| 2021 | Protoda: Efficient Transfer Learning for Few-Shot Intent ClassificationabstractPractical sequence classification tasks in natural language processing often suffer from low training data availability for target classes. Recent works towards mitigating this problem have focused on transfer learning using embeddings pre-trained on often unrelated tasks, for instance, language modeling. We adopt an alternative approach by transfer learning on an ensemble of related tasks using prototypical networks under the meta-learning paradigm. Using intent classification as a case study, we demonstrate that increasing variability in training tasks can significantly improve classification performance. Further, we apply data augmentation in conjunction with meta-learning to reduce sampling bias. We make use of a conditional generator for data augmentation that is trained directly using the meta-learning objective and simultaneously with prototypical networks, hence ensuring that data augmentation is customized to the task. We explore augmentation in the sentence embedding space as well as prototypical embedding space. Combining meta-learning with augmentation provides upto 6.49% and 8.53% relative F1-score improvements over the best performing systems in the 5-shot and 10-shot learning, respectively. Manoj Kumar 0007, Hadrien Glaude, Cyprien de Lichy, Aman Alok, Rahul Gupta 0001 |
SLT | 6 |
| 2020 | Design Considerations for Hypothesis Rejection Modules in Spoken Language Understanding SystemsabstractSpoken Language Understanding (SLU) systems typically consist of a set of machine learning models that operate in conjunction to produce an SLU hypothesis. The generated hypothesis is then sent to downstream components for further action. However, it is desirable to discard an incorrect hypothesis before sending it downstream. In this work, we present two designs for SLU hypothesis rejection modules: (i) scheme R1 that performs rejection on domain specific SLU hypothesis and, (ii) scheme R2 that performs rejection on hypothesis generated from the overall SLU system. Hypothesis rejection modules in both schemes reject/accept a hypothesis based on features drawn from the utterance directed to the SLU system, the associated SLU hypothesis and SLU confidence score. Our experiments suggest that both the schemes yield similar results (scheme R1: 2.5% FRR @ 4.5% FAR, scheme R2: 2.5% FRR @ 4.6% FAR), with the best performing systems using all the available features. We argue that while either of the rejection schemes can be chosen over the other, they carry some inherent differences which need to be considered while making this choice. Additionally, we incorporate ASR features in the rejection module (obtaining an 1.9% FRR @ 3.8% FAR) and analyze the improvements. Aman Alok, Rahul Gupta 0001, Shankar Ananthakrishnan |
ICASSP | 2 |
| 2019 | On Evaluating CNN Representations for Low Resource Medical Image ClassificationabstractConvolutional Neural Networks (CNNs) have revolutionized performances in several machine learning tasks such as image classification, object tracking, and keyword spotting. However, given that they contain a large number of parameters, their direct applicability into low resource tasks is not straightforward. In this work, we experiment with an application of CNN models to gastrointestinal landmark classification with only a few thousands of training samples through transfer learning. As in a standard transfer learning approach, we train CNNs on a large external corpus, followed by representation extraction for the medical images. Finally, a classifier is trained on these CNN representations. However, given that several variants of CNNs exist, the choice of CNN is not obvious. To address this, we develop a novel metric that can be used to predict test performances, given CNN representations on the training set. Not only we demonstrate the superiority of the CNN based transfer learning approach against an assembly of knowledge driven features, but the proposed metric also carries an 87% correlation with the test set performances as obtained using various CNN representations. Taruna Agrawal, Rahul Gupta 0001, Shri Narayanan |
ICASSP | 2 |
| 2019 | One-vs-All Models for Asynchronous Training: An Empirical AnalysisabstractAny given classification problem can be modeled using multi-class or One-vs-All (OVA) architecture. An OVA system consists of as many OVA models as the number of classes, providing the advantage of asynchrony, where each OVA model can be re-trained independent of other models. This is particularly advantageous in settings where scalable model training is a consideration (for instance in an industrial environment where multiple and frequent updates need to be made to the classification system). In this paper, we conduct empirical analysis on realizing independent updates to OVA models and its impact on the accuracy of the overall OVA system. Given that asynchronous updates lead to differences in training datasets for OVA models, we first define a metric to quantify the differences in datasets. Thereafter, using Natural Language Understanding as a task of interest, we estimate the impact of three factors: (i) number of classes, (ii) number of data points and, (iii) divergences in training datasets across OVA models; on the OVA system accuracy. Finally, we observe the accuracy impact of increased asynchrony in a Spoken Language Understanding system. We analyze the results and establish that the proposed metric correlates strongly with the model performances in both the experimental settings. Rahul Gupta 0001, Aman Alok, Shankar Ananthakrishnan |
INTERSPEECH | 1 |
| 2018 | Semi-Supervised and Transfer Learning Approaches for Low Resource Sentiment ClassificationabstractSentiment classification involves quantifying the affective reaction of a human to a document, media item or an event. Although researchers have investigated several methods to reliably infer sentiment from lexical, speech and body language cues, training a model with a small set of labeled datasets is still a challenge. For instance, in expanding sentiment analysis to new languages and cultures, it may not always be possible to obtain comprehensive labeled datasets. In this paper, we investigate the application of semi- supervised and transfer learning methods to improve performances on low resource sentiment classification tasks. We experiment with extracting dense feature representations, pre-training and manifold regularization in enhancing the performance of sentiment classification systems. Our goal is a coherent implementation of these methods and we evaluate the gains achieved by these methods in matched setting involving training and testing on a single corpus setting as well as two cross corpora settings. In both the cases, our experiments demonstrate that the proposed methods can significantly enhance the model performance against a purely supervised approach, particularly in cases involving a handful of training data. Rahul Gupta 0001, Saurabh Sahu, Carol Y. Espy-Wilson, Shri Narayanan |
ICASSP | 1 |
| 2018 | Smoothing Model Predictions Using Adversarial Training Procedures for Speech Based Emotion RecognitionabstractTraining discriminative classifiers involves learning a conditional distribution p(yi|xi), given a set of feature vectors xiand the corresponding labels yi, i=1...N. For a classifier to be generalizable and not overfit to training data, the resulting conditional distribution p(yi|xi) is desired to be smoothly varying over the inputs xi. Adversarial training procedures enforce this smoothness using manifold regularization techniques. Manifold regularization makes the model's output distribution more robust to local perturbation added to a datapoint xi. In this paper, we experiment with the application of adversarial training procedures to increase the accuracy of a deep neural network based emotion recognition system using speech cues. Specifically, we investigate two training procedures: (i) adversarial training where we determine the adversarial direction based on the given labels for the training data and, (ii) virtual adversarial training where we determine the adversarial direction based only on the output distribution of the training data. We demonstrate the efficacy of adversarial training procedures by performing a k-fold cross validation experiment on the Interactive Emotional Dyadic Motion Capture (IEMOCAP) and a cross-corpus performance analysis on three separate corpora. Results show improvement over a purely supervised approach, as well as better generalization capability to cross-corpus settings. Saurabh Sahu, Rahul Gupta 0001, Ganesh Sivaraman, Carol Y. Espy-Wilson |
ICASSP | 2 |
| 2018 | On Enhancing Speech Emotion Recognition Using Generative Adversarial NetworksabstractGenerative Adversarial Networks (GANs) have gained a lot of attention from machine learning community due to their ability to learn and mimic an input data distribution. GANs consist of a discriminator and a generator working in tandem playing a min-max game to learn a target underlying data distribution; when fed with data-points sampled from a simpler distribution (like uniform or Gaussian distribution). Once trained, they allow synthetic generation of examples sampled from the target distribution. We investigate the application of GANs to generate synthetic feature vectors used for speech emotion recognition. Specifically, we investigate two set ups: (i) a vanilla GAN that learns the distribution of a lower dimensional representation of the actual higher dimensional feature vector and, (ii) a conditional GAN that learns the distribution of the higher dimensional feature vectors conditioned on the labels or the emotional class to which it belongs. As a potential practical application of these synthetically generated samples, we measure any improvement in a classifier's performance when the synthetic data is used along with real data for training. We perform cross-validation analyses followed by a cross-corpus study. Saurabh Sahu, Rahul Gupta 0001, Carol Y. Espy-Wilson |
INTERSPEECH | 2 |
| 2018 | A Re-Ranker Scheme For Integrating Large Scale NLU ModelsabstractLarge scale Natural Language Understanding (NLU) systems are typically trained on large quantities of data, requiring a fast and scalable training strategy. A typical design for NLU systems consists of domain-level NLU modules (domain classifier, intent classifier and named entity recognizer). Hypotheses (NLU interpretations consisting of various intent+slot combinations) from these domain specific modules are typically aggregated with another downstream component. The re-ranker integrates outputs from domain-level recognizers, returning a scored list of cross domain hypotheses. An ideal re-ranker will exhibit the following two properties: (a) it should prefer the most relevant hypothesis for the given input as the top hypothesis and, (b) the interpretation scores corresponding to each hypothesis produced by the re-ranker should be calibrated. Calibration allows the final NLU interpretation score to be comparable across domains. We propose a novel re-ranker strategy that addresses these aspects, while also maintaining domain specific modularity. We design optimization loss functions for such a modularized re-ranker and present results on decreasing the top hypothesis error rate as well as maintaining the model calibration. We also experiment with an extension involving training the domain specific re-rankers on datasets curated independently by each domain to allow further asynchronization. Chengwei Su, Rahul Gupta 0001, Shankar Ananthakrishnan, Spyridon Matsoukas |
SLT | 2 |
| 2018 | Modeling Multiple Time Series Annotations as Noisy Distortions of the Ground Truth: An Expectation-Maximization ApproachabstractStudies of time-continuous human behavioral phenomena often rely on ratings from multiple annotators. Since the ground truth of the target construct is often latent, the standard practice is to use ad-hoc metrics (such as averaging annotator ratings). Despite being easy to compute, such metrics may not provide accurate representations of the underlying construct. In this paper, we present a novel method for modeling multiple time series annotations over a continuous variable that computes the ground truth by modeling annotator specific distortions. We condition the ground truth on a set of features extracted from the data and further assume that the annotators provide their ratings as modification of the ground truth, with each annotator having specific distortion tendencies. We train the model using an Expectation-Maximization based algorithm and evaluate it on a study involving natural interaction between a child and a psychologist, to predict confidence ratings of the children's smiles. We compare and analyze the model against two baselines where: (i) the ground truth in considered to be framewise mean of ratings from various annotators and, (ii) each annotator is assumed to bear a distinct time delay in annotation and their annotations are aligned before computing the framewise mean. Rahul Gupta 0001, Kartik Audhkhasi, Zach Jacokes, Agata Rozga, Shri Narayanan |
IEEE Trans. Affect. Comput. | 1 |
| 2017 | A knowledge transfer and boosting approach to the prediction of affect in moviesabstractAffect prediction is a classical problem and has recently garnered special interest in multimedia applications. Affect prediction in movies is one such domain, potentially aiding the design as well as the impact analysis of movies. Given the large diversity in movies (such as different genres and languages), obtaining a comprehensive movie dataset for modeling affect is challenging while models trained on smaller datasets may not generalize. In this paper, we address the problem of continuous affect ratings with the availability of limited in-domain data resources. We initially setup several baseline models trained on in-domain data, followed by a proposal of a Knowledge Transfer (KT) + Gradient Boosting (GB) approach. KT learns models on a larger (mismatched) data which are then adapted to make predictions on the data of interest. GB further updates these predictions based on models learnt from the in-domain data. We observe that the KT + GB models provide Concordance Correlation Coefficient values of 0.13 and 0.27 for valence and affect prediction on the continuous LIRIS ACCEDE dataset against best baseline prediction values of 0.12 and 0.11. Not only the KT + GB models improve the overall performance metrics, we also observe a more consistent model performance across movies of various genres. Sabyasachee Baruah, Rahul Gupta 0001, Shri Narayanan |
ICASSP | 2 |
| 2017 | An Affect Prediction Approach Through Depression Severity Parameter Incorporation in Neural Networks
Rahul Gupta 0001, Saurabh Sahu, Carol Y. Espy-Wilson, Shri Narayanan |
INTERSPEECH | 1 |
| 2017 | Transfer Learning Between Concepts for Human Behavior Modeling: An Application to Sincerity and Deception Prediction
Qinyi Luo, Rahul Gupta 0001, Shri Narayanan |
INTERSPEECH | 2 |
| 2017 | Adversarial Auto-Encoders for Speech Based Emotion RecognitionabstractRecently, generative adversarial networks and adversarial autoencoders have gained a lot of attention in machine learning community due to their exceptional performance in tasks such as digit classification and face recognition. They map the autoencoder's bottleneck layer output (termed as code vectors) to different noise Probability Distribution Functions (PDFs), that can be further regularized to cluster based on class information. In addition, they also allow a generation of synthetic samples by sampling the code vectors from the mapped PDFs. Inspired by these properties, we investigate the application of adversarial autoencoders to the domain of emotion recognition. Specifically, we conduct experiments on the following two aspects: (i) their ability to encode high dimensional feature vector representations for emotional utterances into a compressed space (with a minimal loss of emotion class discriminability in the compressed space), and (ii) their ability to regenerate synthetic samples in the original feature space, to be later used for purposes such as training emotion recognition classifiers. We demonstrate the promise of adversarial autoencoders with regards to these aspects on the Interactive Emotional Dyadic Motion Capture (IEMOCAP) corpus and present our analysis. Saurabh Sahu, Rahul Gupta 0001, Ganesh Sivaraman, Wael Abd-Almageed, Carol Y. Espy-Wilson |
INTERSPEECH | 2 |
| 2016 | Pathological speech processing: State-of-the-art, current challenges, and future directionsabstractThe study of speech pathology involves evaluation and treatment of speech production related disorders affecting phonation, fluency, intonation and aeromechanical components of respiration. Recently, speech pathology has garnered special interest amongst machine learning and signal processing (ML-SP) scientists. This growth in interest is led by advances in novel data collection technology, data science, speech processing and computational modeling. These in turn have enabled scientists in better understanding both the causes and effects of pathological speech conditions. In this paper, we review the application of machine learning and signal processing techniques to speech pathology and specifically focus on three different aspects. First, we list challenges such as controlling subjectivity in pathological speech assessments and patient variability in the application of ML-SP tools to the domain. Second, we discuss feature design methods and machine learning algorithms using a combination of domain knowledge and data driven methods. Finally, we present some case studies related to analysis of pathological speech and discuss their design. Rahul Gupta 0001, Theodora Chaspari, Jangwon Kim, Naveen Kumar 0004, Daniel Bone, Shri Narayanan |
ICASSP | 1 |
| 2016 | Acoustic-Prosodic and Turn-Taking Features in Interactions with Children with Neurodevelopmental Disorders
Daniel Bone, Somer Bishop, Rahul Gupta 0001, Sungbok Lee, Shri Narayanan |
INTERSPEECH | 3 |
| 2016 | Automatic Estimation of Perceived Sincerity from Spoken Language
Brandon M. Booth, Rahul Gupta 0001, Pavlos Papadopoulos, Ruchir Travadi, Shri Narayanan |
INTERSPEECH | 2 |
| 2016 | Predicting Affective Dimensions Based on Self Assessed Depression Severity
Rahul Gupta 0001, Shri Narayanan |
INTERSPEECH | 1 |
| 2016 | Laughter Valence Prediction in Motivational Interviewing Based on Lexical and Acoustic Cues
Rahul Gupta 0001, Nishant Nath, Taruna Agrawal, Panayiotis G. Georgiou, David C. Atkins, Shri Narayanan |
INTERSPEECH | 1 |
| 2016 | Objective Language Feature Analysis in Children with Neurodevelopmental Disorders During Autism Assessment
Manoj Kumar 0007, Rahul Gupta 0001, Daniel Bone, Nikos Malandrakis, Somer Bishop, Shri Narayanan |
INTERSPEECH | 2 |
| 2016 | An Expectation Maximization Approach to Joint Modeling of Multidimensional Ratings Derived from Multiple AnnotatorsabstractRatings from multiple human annotators are often pooled in applications where the ground truth is hidden. Examples include annotating perceived emotions and assessing quality metrics for speech and image. These ratings are not restricted to a single dimension and can be multidimensional. In this paper, we propose an Expectation-Maximization based algorithm to model such ratings. Our model assumes that there exists a latent multidimensional ground truth that can be determined from the observation features and that the ratings provided by the annotators are noisy versions of the ground truth. We test our model on a study conducted on children with autism to predict a four dimensional rating of expressivity, naturalness, pronunciation goodness and engagement. Our goal in this application is to reliably predict the individual annotator ratings which can be used to address issues of cognitive load on the annotators as well as the rating cost. We initially train a baseline directly predicting annotator ratings from the features and compare it to our model under three different settings assuming: (i) each entry in the multidimensional rating is independent of others, (ii) a joint distribution among rating dimensions exists, (iii) a partial set of ratings to predict the remaining entries is available. Anil Ramakrishna, Rahul Gupta 0001, Ruth B. Grossman, Shri Narayanan |
INTERSPEECH | 2 |
| 2016 | Detecting paralinguistic events in audio stream using context in features and probabilistic decisions
Rahul Gupta 0001, Kartik Audhkhasi, Sungbok Lee, Shri Narayanan |
Comput. Speech Lang. | 1 |
| 2016 | Analysis of engagement behavior in children during dyadic interactions using prosodic cues
Rahul Gupta 0001, Daniel Bone, Sungbok Lee, Shri Narayanan |
Comput. Speech Lang. | 1 |
| 2015 | A language-based generative model framework for behavioral analysis of couples' therapyabstractObservational studies for psychological evaluations rely on careful assessment of multiple behavioral cues. Recent studies have made good progress in automating the psychological evaluation, which often involved tedious manual annotation of a set of behavioral codes. However, the current methods impose strict and often unnatural assumptions for evaluation. In this work, we specifically investigate two goals: (1) Human behavior changes throughout an interaction and better models of this evolution can improve automated behavioral annotation and (2) Human perception of this evolution can be quite complex and non-linear and better techniques than averaging need to be investigated. For this purpose, we propose a Dynamic Behavior Modeling (DBM) scheme, which models a spouse as undergoing changes in behavioral state within a session, and contrast it against a Static Behavior Model (SBM) which allows only a constant session-long behavioral state. We use Negativity in a couples therapy task as our case study. We present results and analysis on both models for capturing the local behavior information and predicting the session level negativity label. Sandeep Nallan Chakravarthula, Rahul Gupta 0001, Brian R. Baucom, Panayiotis G. Georgiou |
ICASSP | 2 |
| 2015 | A mixture of experts approach towards intelligibility classification of pathological speechabstractPathological speech involves atypical speech production which may result from several factors including oral diseases, physical disabilities in the voice production system and atypical anatomy. Automatic evaluation of intelligibility in patients with pathological speech can assist accurate diagnosis of pathological conditions. Loss of intelligibility may be associated with one of the several pathological conditions, making automatic evaluation a challenging computational problem. A Mixture of Experts (MoE) models class boundaries using a weighted combination of several experts and can characterize the complex class boundaries arising due to pathological variability. We train an MoE for intelligibility evaluation using a modified Expectation Maximization (EM) algorithm based on joint simulated annealing-gradient ascent procedure. Our algorithm optimizes the expert parameters and simultaneously obtains the feature subsets for each expert. We observe that the MoE trained using the new EM algorithm not only outperforms a single classifier baseline but also the vanilla MoE. We perform further data analysis and interpret the weights assigned to each expert during inference. Also, we obtain a different feature subset per expert in the mixture. This illustrates feature use based on location of the data point in the feature space. Rahul Gupta 0001, Kartik Audhkhasi, Shri Narayanan |
ICASSP | 1 |
| 2015 | Automated evaluation of non-native English pronunciation quality: combining knowledge- and data-driven features at multiple time scalesabstractAutomatically evaluating pronunciation quality of non-native speech has seen tremendous success in both research and com-mercial settings, with applications in L2 learning. In this paper, submitted for the INTERSPEECH 2015 Degree of Nativeness Sub-Challenge, this problem is posed under a challenging cross-corpora setting using speech data drawn from multiple speakers from a variety of language backgrounds (L1) reading different English sentences. Since the perception of non-nativeness is re-alized at the segmental and suprasegmental linguistic levels, we explore a number of acoustic cues at multiple time scales. We experiment with both data-driven and knowledge-inspired fea-tures that capture degree of nativeness from pauses in speech, speaking rate, rhythm/stress, and goodness of phone pronunci-ation. One promising finding is that highly accurate automated assessment can be attained using a small diverse set of intuitive and interpretable features. Performance is further boosted by smoothing scores across utterances from the same speaker; our best system significantly outperforms the challenge baseline. Matthew Black, Daniel Bone, Z.-I. Skordilis, Rahul Gupta 0001, Pavlos Papadopoulos, Sandeep Nallan Chakravarthula, Bo Xiao 0003, Maarten Van Segbroeck, Jangwon Kim, Panayiotis G. Georgiou, Shri Narayanan |
INTERSPEECH | 4 |
| 2015 | Analysis and modeling of the role of laughter in motivational interviewing based psychotherapy conversations
Rahul Gupta 0001, Theodora Chaspari, Panayiotis G. Georgiou, David C. Atkins, Shri Narayanan |
INTERSPEECH | 1 |
| 2015 | Automatic estimation of parkinson's disease severity from diverse speech tasksabstractThe need for reliable, scalable and efficient diagnosis of Parkin-son’s Disease (PD) is a major clinical need. Automating the diagnosis can lead to more accurate and objective predictions as well as provide insights regarding the nature of Parkinson’s condition. This paper proposes a fully automated system to rate the severity (UPDRS-III scale) of PD from patients ’ speech. Specifically, the system captures atypicalities in an individ-ual’s voice when performing multiple diverse speaking tasks and makes a unified prediction of the PD severity. The perfor-mance is tested in a cross-data setting, with different subjects and dissimilar recording conditions. Results indicate that (i) effective features vary depending on the nature of the specific speech task, (ii) additional novel feature sets to detect distor-tions in Parkinson’s speech significantly improve the prediction accuracy from the Interspeech15 Challenge baseline system and (iii) our fusion system based on an unsupervised clustering tech-nique also improves the accuracy. Our system incorporates i-vector and functionals for segmental features, non-linear time series features, speech rhythm and automatic speech recogni-tion decoding based features. By its application on the Inter-speech15 eating condition challenge, the system also shows its potential for detecting other sources of speech variability. Jangwon Kim, Md. Nasir, Rahul Gupta 0001, Maarten Van Segbroeck, Daniel Bone, Matthew Black, Z.-I. Skordilis, Panayiotis G. Georgiou, Shri Narayanan |
INTERSPEECH | 3 |
| 2014 | Training ensemble of diverse classifiers on feature subsetsabstractEnsembles of diverse classifiers often out-perform single classifiers as has been well-demonstrated across several applications. Existing training algorithms either learn a classifier ensemble on pre-defined feature sets or independently perform classifier training and feature selection. Neither of these schemes is optimal. We pose feature subset selection and training of diverse classifiers on selected subsets as a joint optimization problem. We propose a novel greedy algorithm to solve this problem. We sequentially learn an ensemble of classifiers where each subsequent classifier is encouraged to learn data instances misclassified by previous classifiers on a concurrently selected feature set. Our experiments on synthetic and real-world data sets show the effectiveness of our algorithm. We observe that ensembles trained by our algorithm performs better than both a single classifier and an ensemble of classifiers learnt on pre-defined feature sets. We also test our algorithm as a feature selector on a synthetic dataset to filter out irrelevant features. Rahul Gupta 0001, Kartik Audhkhasi, Shri Narayanan |
ICASSP | 1 |
| 2014 | Variable Span disfluency detection in ASR transcripts
Rahul Gupta 0001, Sankaranarayanan Ananthakrishnan, Shri Narayanan |
INTERSPEECH | 1 |
| 2014 | Predicting client's inclination towards target behavior change in motivational interviewing and investigating the role of laughter
Rahul Gupta 0001, Panayiotis G. Georgiou, David C. Atkins, Shri Narayanan |
INTERSPEECH | 1 |
| 2013 | Paralinguistic event detection from speech using probabilistic time-series smoothing and masking
Rahul Gupta 0001, Kartik Audhkhasi, Sungbok Lee, Shri Narayanan |
INTERSPEECH | 1 |
| 2012 | Classification of emotional content of sighs in dyadic human interactionsabstractEmotions are an important part of human communication and are expressed both verbally and non-verbally. Common nonverbal vocalizations such as laughter, cries and sighs carry important emotional content in conversations. Sighs often are associated with negative emotion. In this work, we show that emotional sighs exist along both ends of the valence axis (positive-emotion vs. negative-emotion sighs) in spontaneous affective dialogs and that they have certain distinct multimodal characteristics. Classification results show that it is possible to differentiate between the two types of emotionally valenced sighs, using a combination of acoustic and gestural features with an overall unweighted accuracy of 58.26%. Rahul Gupta 0001, Chi-Chun Lee, Shri Narayanan |
ICASSP | 1 |