Hema Swetha Koppula

dblp:41/7536 · also Hema Koppula · DBLP profile ↗
← Back
17ranked-venue papers
6as first author
5since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 13 · 6 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-authorSystems, architecture and hardware · 2 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
9 papers
3D vision · 16% Generative modeling · 15% Video understanding and tracking · 14%
Databases, data mining, and information retrieval
1 paper
Information retrieval · 64% Web and social media mining · 28% Data integration and cleaning · 8%

Topics — the 28 heaviest of 32, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Speech recognition and synthesis
text-to-speech synthesis
0.612022
Style Equalization: Unsupervised Learning of Controllable Generative Sequence Models · ICML 2022
Computer vision › 3D vision
3d scene understanding
0.532016
Modeling 3D Environments through Hidden Human Context · IEEE Trans. Pattern Anal. Mach. Intell. 2016
Hallucinated Humans as the Hidden Context for Labeling 3D Scenes · CVPR 2013
Semantic Labeling of 3D Point Clouds for Indoor Scenes · NIPS 2011
Robotics › Autonomous driving
maneuver anticipation
0.522016
Recurrent Neural Networks for driver activity anticipation via sensory-fusion architecture · ICRA 2016
Car that Knows Before You Do: Anticipating Maneuvers via Learning Temporal Driving Models · ICCV 2015
Computer vision › Segmentation and scene understanding
semantic segmentation
0.422016
Modeling 3D Environments through Hidden Human Context · IEEE Trans. Pattern Anal. Mach. Intell. 2016
Semantic Labeling of 3D Point Clouds for Indoor Scenes · NIPS 2011
Computer vision › Video understanding and tracking › activity recognition
activity detection
0.212016
Anticipating Human Activities Using Object Affordances for Reactive Robotic Response · IEEE Trans. Pattern Anal. Mach. Intell. 2016
Robotics › Robot manipulation › service robot
assistive robotics
0.212016
Anticipating Human Activities Using Object Affordances for Reactive Robotic Response · IEEE Trans. Pattern Anal. Mach. Intell. 2016
Computer vision › Video understanding and tracking › action anticipation
driver action prediction
0.212016
Recurrent Neural Networks for driver activity anticipation via sensory-fusion architecture · ICRA 2016
Machine learning › Deep learning architectures and training › recurrent neural network
LSTM
0.212016
Recurrent Neural Networks for driver activity anticipation via sensory-fusion architecture · ICRA 2016
Machine learning › Deep learning architectures and training
recurrent neural network
0.212016
Recurrent Neural Networks for driver activity anticipation via sensory-fusion architecture · ICRA 2016
Robotics › Autonomous driving
driver behavior modeling
0.212015
Car that Knows Before You Do: Anticipating Maneuvers via Learning Temporal Driving Models · ICCV 2015
Knowledge, reasoning and agents › Knowledge representation and reasoning
commonsense reasoning
0.212014
Physically Grounded Spatio-temporal Object Affordances · ECCV (3) 2014
Computer vision › 3D vision › embodied vision
object affordance
0.212014
Physically Grounded Spatio-temporal Object Affordances · ECCV (3) 2014
Machine learning › Generative modeling
autoregressive model
0.212022
Style Equalization: Unsupervised Learning of Controllable Generative Sequence Models · ICML 2022
Machine learning › Generative modeling › generative model
generative sequence model
0.212022
Style Equalization: Unsupervised Learning of Controllable Generative Sequence Models · ICML 2022
Computer vision › 3D vision › 3d scene understanding
3d scene labeling
0.212013
Hallucinated Humans as the Hidden Context for Labeling 3D Scenes · CVPR 2013
Computer vision › Video understanding and tracking
spatio-temporal modeling
0.212013
Learning Spatio-Temporal Structure from RGB-D Videos for Human Activity Detection and Anticipation · ICML (3) 2013
Computer vision › Image recognition and object detection
object detection
0.112011
Semantic Labeling of 3D Point Clouds for Indoor Scenes · NIPS 2011
Computer vision › 3D vision › point cloud segmentation
point cloud semantic segmentation
0.112011
Semantic Labeling of 3D Point Clouds for Indoor Scenes · NIPS 2011
Information retrieval › search engines
web crawling
0.112010
Learning URL patterns for webpage de-duplication · WSDM 2010
Web and social media mining › web mining
web page deduplication
0.112010
Learning URL patterns for webpage de-duplication · WSDM 2010
Information retrieval
web search
0.112010
Learning URL patterns for webpage de-duplication · WSDM 2010
Machine learning › Probabilistic and Bayesian machine learning › structured prediction
conditional random field
0.112016
Modeling 3D Environments through Hidden Human Context · IEEE Trans. Pattern Anal. Mach. Intell. 2016
Machine learning › Probabilistic and Bayesian machine learning › structured models
graphical models
0.112016
Modeling 3D Environments through Hidden Human Context · IEEE Trans. Pattern Anal. Mach. Intell. 2016
Computer vision › Video understanding and tracking
temporal modeling
0.112015
Car that Knows Before You Do: Anticipating Maneuvers via Learning Temporal Driving Models · ICCV 2015
Natural language and speech › Information extraction and text analysis
topic model
0.012013
Hallucinated Humans as the Hidden Context for Labeling 3D Scenes · CVPR 2013
Robotics › Robot navigation and mapping
mobile robot perception
0.012011
Semantic Labeling of 3D Point Clouds for Indoor Scenes · NIPS 2011
Data integration and cleaning › entity resolution
duplicate detection
0.012010
Learning URL patterns for webpage de-duplication · WSDM 2010
Information retrieval
indexing
0.012010
Learning URL patterns for webpage de-duplication · WSDM 2010

Methods — techniques the papers use, named apart from their topics

style transformation module · 0.6style equalization · 0.6sequence-to-sequence prediction · 0.2sensory fusion · 0.2particle filtering · 0.2infinite latent conditional random field · 0.2human-object interaction modeling · 0.2dirichlet process · 0.2anticipatory temporal conditional random field · 0.2autoregressive input-output HMM · 0.2rule mining · 0.1mapreduce · 0.1machine learning · 0.1
YearPublicationVenuePosition
2024 Corpus Synthesis for Zero-Shot ASR Domain Adaptation Using Large Language Models
abstract
While Automatic Speech Recognition (ASR) systems are widely used in many real-world applications, they often do not generalize well to new domains and need to be fine-tuned on data from these domains. However, target-domain data usually are not readily available in many scenarios. In this paper, we propose a new strategy for adapting ASR models to new target domains without any text or speech from those domains. To accomplish this, we propose a novel data synthesis pipeline that uses a Large Language Model (LLM) to generate a target domain text corpus, and a state-of-the-art controllable speech synthesis model to generate the corresponding speech. We propose a simple yet effective in-context instruction fine-tuning strategy to increase the effectiveness of LLM in generating text corpora for new domains. Experiments on the SLURP dataset show that the proposed method achieves an average relative word error rate improvement of 28% on unseen target domains without any performance drop in source domains.
Hsuan Su, Ting-Yao Hu, Hema Swetha Koppula, Raviteja Vemulapalli, Jen-Hao Rick Chang, Karren D. Yang, Gautam Varma Mantena, Oncel Tuzel
ICASSP3
2023 Text is all You Need: Personalizing ASR Models Using Controllable Speech Synthesis
abstract
Adapting generic speech recognition models to specific individuals is a challenging problem due to the scarcity of personalized data. Recent works have proposed boosting the amount of training data using personalized text-to-speech synthesis. Here, we ask two fundamental questions about this strategy: when is synthetic data effective for personalization, and why is it effective in those cases? To address the first question, we adapt a state-of-the-art automatic speech recognition (ASR) model to target speakers from four benchmark datasets representative of different speaker types. We show that ASR personalization with synthetic data is effective in all cases, but particularly when (i) the target speaker is underrepresented in the global data, and (ii) the capacity of the global model is limited. To address the second question of why personalized synthetic data is effective, we use controllable speech synthesis (CSS) to generate speech with varied styles and content. Surprisingly, we find that the text content of the synthetic data, rather than style, is important for speaker adaptation. These results lead us to propose a data selection strategy for ASR personalization based on speech content.
Karren D. Yang, Ting-Yao Hu, Jen-Hao Rick Chang, Hema Swetha Koppula, Oncel Tuzel
ICASSP4
2022 SYNT++: Utilizing Imperfect Synthetic Data to Improve Speech Recognition
abstract
With recent advances in speech synthesis, synthetic data is becoming a viable alternative to real data for training speech recognition models. However, machine learning with synthetic data is not trivial due to the gap between the synthetic and the real data distributions. Synthetic datasets may contain artifacts that do not exist in real data such as structured noise, content errors, or unrealistic speaking styles. Moreover, the synthesis process may introduce a bias due to uneven sampling of the data manifold. We propose two novel techniques during training to mitigate the problems due to the distribution gap: (i) a rejection sampling algorithm and (ii) using separate batch normalization statistics for the real and the synthetic samples. We show that these methods significantly improve the training of speech recognition models using synthetic data. We evaluate the proposed approach on keyword detection and Automatic Speech Recognition (ASR) tasks, and observe up to 18% and 13% relative error reduction, respectively, compared to naively using the synthetic data.
Ting-Yao Hu, Mohammadreza Armandpour, Ashish Shrivastava 0001, Jen-Hao Rick Chang, Hema Swetha Koppula, Oncel Tuzel
ICASSP5
2022 Style Equalization: Unsupervised Learning of Controllable Generative Sequence Models
abstract
Controllable generative sequence models with the capability to extract and replicate the style of specific examples enable many applications, including narrating audiobooks in different voices, auto-completing and auto-correcting written handwriting, and generating missing training samples for downstream recognition tasks. However, under an unsupervised-style setting, typical training algorithms for controllable sequence generative models suffer from the training-inference mismatch, where the same sample is used as content and style input during training but unpaired samples are given during inference. In this paper, we tackle the training-inference mismatch encountered during unsupervised learning of controllable generative sequence models. The proposed method is simple yet effective, where we use a style transformation module to transfer target style information into an unrelated style input. This method enables training using unpaired content and style samples and thereby mitigate the training-inference mismatch. We apply style equalization to text-to-speech and text-to-handwriting synthesis on three datasets. We conduct thorough evaluation, including both quantitative and qualitative user studies. Our results show that by mitigating the training-inference mismatch with the proposed style equalization, we achieve style replication scores comparable to real data in our user studies.
Jen-Hao Rick Chang, Ashish Shrivastava 0001, Hema Swetha Koppula, Xiaoshuai Zhang, Oncel Tuzel
ICML3
2021 SapAugment: Learning A Sample Adaptive Policy for Data Augmentation
abstract
Data augmentation methods usually apply the same augmentation (or a mix of them) to all the training samples. For example, to perturb data with noise, the noise is sampled from a Normal distribution with a fixed standard deviation, for all samples. We hypothesize that a hard sample with high training loss already provides strong training signal to update the model parameters and should be perturbed with mild or no augmentation. Perturbing a hard sample with a strong augmentation may also make it too hard to learn from. Furthermore, a sample with low training loss should be perturbed by a stronger augmentation to provide more robustness to a variety of conditions. To formalize these intuitions, we propose a novel method to learn a Sample-Adaptive Policy for Augmentation – SapAugment. Our policy adapts the augmentation parameters based on the training loss of the data samples. In the example of Gaussian noise, a hard sample will be perturbed with a low variance noise and an easy sample with a high variance noise. Furthermore, the proposed method combines multiple augmentation methods into a methodical policy learning framework and obviates hand-crafting augmentation parameters by trial-and-error. We apply our method on an automatic speech recognition (ASR) task, and combine existing and novel augmentations using the proposed framework. We show substantial improvement, up to 21% relative reduction in word error rate on LibriSpeech dataset, over the state-of-the-art speech augmentation method.
Ting-Yao Hu, Ashish Shrivastava 0001, Jen-Hao Rick Chang, Hema Swetha Koppula, Kyuyeon Hwang, Ozlem Kalinli, Oncel Tuzel
ICASSP4
2016 Recurrent Neural Networks for driver activity anticipation via sensory-fusion architecture
abstract
Anticipating the future actions of a human is a widely studied problem in robotics that requires spatio-temporal reasoning. In this work we propose a deep learning approach for anticipation in sensory-rich robotics applications. We introduce a sensory-fusion architecture which jointly learns to anticipate and fuse information from multiple sensory streams. Our architecture consists of Recurrent Neural Networks (RNNs) that use Long Short-Term Memory (LSTM) units to capture long temporal dependencies. We train our architecture in a sequence-to-sequence prediction manner, and it explicitly learns to predict the future given only a partial temporal context. We further introduce a novel loss layer for anticipation which prevents over-fitting and encourages early anticipation. We use our architecture to anticipate driving maneuvers several seconds before they happen on a natural driving data set of 1180 miles. The context for maneuver anticipation comes from multiple sensors installed on the vehicle. Our approach shows significant improvement over the state-of-the-art in maneuver anticipation by increasing the precision from 77.4% to 90.5% and recall from 71.2% to 87.4%.
Ashesh Jain, Avi Singh, Hema Swetha Koppula, Shane Soh, Ashutosh Saxena
ICRA3
2016 Modeling 3D Environments through Hidden Human Context
abstract
The idea of modeling object-object relations has been widely leveraged in many scene understanding applications. However, as the objects are designed by humans and for human usage, when we reason about a human environment, we reason about it through an interplay between the environment, objects and humans. In this paper, we model environments not only through objects, but also through latent human poses and human-object interactions. In order to handle the large number of latent human poses and a large variety of their interactions with objects, we present Infinite Latent Conditional Random Field (ILCRF) that models a scene as a mixture of CRFs generated from Dirichlet processes. In each CRF, we model objects and object-object relations as existing nodes and edges, and hidden human poses and human-object relations as latent nodes and edges. ILCRF generatively models the distribution of different CRF structures over these latent nodes and edges. We apply the model to the challenging applications of 3D scene labeling and robotic scene arrangement. In extensive experiments, we show that our model significantly outperforms the state-of-the-art results in both applications. We further use our algorithm on a robot for arranging objects in a new scene using the two applications aforementioned.
Hema Swetha Koppula, Ashutosh Saxena
IEEE Trans. Pattern Anal. Mach. Intell.2
2016 Anticipating Human Activities Using Object Affordances for Reactive Robotic Response
abstract
An important aspect of human perception is anticipation, which we use extensively in our day-to-day activities when interacting with other humans as well as with our surroundings. Anticipating which activities will a human do next (and how) can enable an assistive robot to plan ahead for reactive responses. Furthermore, anticipation can even improve the detection accuracy of past activities. The challenge, however, is two-fold: We need to capture the rich context for modeling the activities and object affordances, and we need to anticipate the distribution over a large space of future human activities. In this work, we represent each possible future using an anticipatory temporal conditional random field (ATCRF) that models the rich spatial-temporal relations through object affordances. We then consider each ATCRF as a particle and represent the distribution over the potential futures using a set of particles. In extensive evaluation on CAD-120 human activity RGB-D dataset, we first show that anticipation improves the state-of-the-art detection results. We then show that for new subjects (not seen in the training set), we obtain an activity anticipation accuracy (defined as whether one of top three predictions actually happened) of 84.1, 74.4 and 62.2 percent for an anticipation time of 1, 3 and 10 seconds respectively. Finally, we also show a robot using our algorithm for performing a few reactive responses.
Hema Swetha Koppula, Ashutosh Saxena
IEEE Trans. Pattern Anal. Mach. Intell.1
2015 Car that Knows Before You Do: Anticipating Maneuvers via Learning Temporal Driving Models
abstract
Advanced Driver Assistance Systems (ADAS) have made driving safer over the last decade. They prepare vehicles for unsafe road conditions and alert drivers if they perform a dangerous maneuver. However, many accidents are unavoidable because by the time drivers are alerted, it is already too late. Anticipating maneuvers beforehand can alert drivers before they perform the maneuver and also give ADAS more time to avoid or prepare for the danger. In this work we anticipate driving maneuvers a few seconds before they occur. For this purpose we equip a car with cameras and a computing device to capture the driving context from both inside and outside of the car. We propose an Autoregressive Input-Output HMM to model the contextual information alongwith the maneuvers. We evaluate our approach on a diverse data set with 1180 miles of natural freeway and city driving and show that we can anticipate maneuvers 3.5 seconds before they occur with over 80% F1-score in real-time.
Ashesh Jain, Hema Swetha Koppula, Bharad Raghavan, Shane Soh, Ashutosh Saxena
ICCV2
2014 Physically Grounded Spatio-temporal Object Affordances
Hema Swetha Koppula, Ashutosh Saxena
ECCV (3)1
2013 Hallucinated Humans as the Hidden Context for Labeling 3D Scenes
abstract
For scene understanding, one popular approach has been to model the object-object relationships. In this paper, we hypothesize that such relationships are only an artifact of certain hidden factors, such as humans. For example, the objects, monitor and keyboard, are strongly spatially correlated only because a human types on the keyboard while watching the monitor. Our goal is to learn this hidden human context (i.e., the human-object relationships), and also use it as a cue for labeling the scenes. We present Infinite Factored Topic Model (IFTM), where we consider a scene as being generated from two types of topics: human configurations and human-object relationships. This enables our algorithm to hallucinate the possible configurations of the humans in the scene parsimoniously. Given only a dataset of scenes containing objects but not humans, we show that our algorithm can recover the human object relationships. We then test our algorithm on the task of attribute and object labeling in 3D scenes and show consistent improvements over the state-of-the-art.
Hema Swetha Koppula, Ashutosh Saxena
CVPR2
2013 Learning Spatio-Temporal Structure from RGB-D Videos for Human Activity Detection and Anticipation
abstract
We consider the problem of detecting past activities as well as anticipating which activity will happen in the future and how. We start by modeling the rich spatio-temporal relations between human poses and objects (called affordances) using a conditional random field (CRF). However, because of the ambiguity in the temporal segmentation of the sub-activities that constitute an activity, in the past as well as in the future, multiple graph structures are possible. In this paper, we reason about these alternate possibilities by reasoning over multiple possible graph structures. We obtain them by approximating the graph with only additive features, which lends to efficient dynamic programming. Starting with this proposal graph structure, we then design moves to obtain several other likely graph structures. We then show that our approach improves the state-of-the-art significantly for detecting past activities as well as for anticipating future activities, on a dataset of 120 activity videos collected from four subjects.
Hema Swetha Koppula, Ashutosh Saxena
ICML (3)1
2013 Anticipating human activities for reactive robotic response
abstract
An important aspect of human perception is anticipation, which we use extensively in our day-to-day activities when interacting with other humans as well as with our surroundings. Anticipating which activities will a human do next (and how to do them) can enable an assistive robot to plan ahead for reactive responses in the human environments. In this work, our goal is to enable robots to predict the future activities as well as the details of how a human is going to perform them in short-term (e.g., 1-10 seconds). For example, if a robot has seen a person move his hand to a coffee mug, it is possible he would move the coffee mug to a few potential places such as his mouth, to a kitchen sink or just move it to a different location on the table. If a robot can anticipate this, then it would rather not start pouring milk into the coffee when the person is moving his hand towards the mug, thus avoiding a spill. We represent each possible future using an anticipatory temporal conditional random field (ATCRF) that models the rich spatial-temporal relations through object affordances. We then consider each ATCRF as a particle and represent the distribution over the potential futures using a set of particles. We evaluate our anticipation approach extensively on CAD-120 human activity dataset, which contains 120 RGB-D videos of daily human activities, such as microwaving food, taking medicine, etc. For robotic evaluation, we measure how many times the robot anticipates and performs the correct reactive response. The accompanying video shows a PR2 robot performing assistive tasks based on the anticipations generated by our proposed method.
Hema Swetha Koppula, Ashutosh Saxena
IROS1
2011 Semantic Labeling of 3D Point Clouds for Indoor Scenes
abstract
Inexpensive RGB-D cameras that give an RGB image together with depth data have become widely available. In this paper, we use this data to build 3D point clouds of full indoor scenes such as an office and address the task of semantic labeling of these 3D point clouds. We propose a graphical model that captures various features and contextual relations, including the local visual appearance and shape cues, object co-occurence relationships and geometric relationships. With a large number of object classes and relations, the model’s parsimony becomes important and we address that by using multiple types of edge potentials. The model admits efficient approximate inference, and we train it using a maximum-margin learning approach. In our experiments over a total of 52 3D scenes of homes and offices (composed from about 550 views, having 2495 segments labeled with 27 object classes), we get a performance of 84.06% in labeling 17 object classes for offices, and 73.38% in labeling 17 object classes for home scenes. Finally, we applied these algorithms successfully on a mobile robot for the task of finding objects in large cluttered rooms.
Hema Swetha Koppula, Abhishek Anand, Thorsten Joachims, Ashutosh Saxena
NIPS1
2010 Alignment of short length parallel corpora with an application to web search
abstract
With evolving Web, short length parallel corpora is becoming very common and some of these include user queries, web snippets etc. This paper concerns situations where short length parallel corpora has to be analyzed in order to find meaningful unit-alignment. This is similar to dealing with parallel corpora where a sentence level alignment of translations is required, but differs in that the alignment is to be inferred at unit (word or phrase) level. A Conditional Random Field (CRF) based approach is proposed to discover this unit alignment. Given pairs of semantically or syntactically similar entities, the problem is formulated as that of mutual segmentation and sequence alignment problem. The mutual segmentation refers to the process of segmenting the first entity based on units (or labels) in the second entity and vice-versa. The process of optimizing this mutual segmentation also results in optimal unit alignment. Since our training data is not segmented and unit-aligned, we modify the CRF objective function to accommodate unsupervised data and iterative learning. We have applied this framework to Web Search domain and specifically for query reformulation task. Finally, our experiments suggest that the proposed approach indeed results in meaningful alternatives of the original query.
Jitendra Ajmera, Hema Swetha Koppula, Krishna P. Leela, Shibnath Mukherjee, Mehul Parsana
CIKM2
2010 Learning URL patterns for webpage de-duplication
abstract
Presence of duplicate documents in the World Wide Web adversely affects crawling, indexing and relevance, which are the core building blocks of web search. In this paper, we present a set of techniques to mine rules from URLs and utilize these rules for de-duplication using just URL strings without fetching the content explicitly. Our technique is composed of mining the crawl logs and utilizing clusters of similar pages to extract transformation rules, which are used to normalize URLs belonging to each cluster. Preserving each mined rule for de-duplication is not efficient due to the large number of such rules. We present a machine learning technique to generalize the set of rules, which reduces the resource footprint to be usable at web-scale. The rule extraction techniques are robust against web-site specific URL conventions. We compare the precision and scalability of our approach with recent efforts in using URLs for de-duplication. Experimental results demonstrate that our approach achieves 2 times more reduction in duplicates with only half the rules compared to the most recent previous approach. Scalability of the framework is demonstrated by performing a large scale evaluation on a set of 3 Billion URLs, implemented using the MapReduce framework.
Hema Swetha Koppula, Krishna P. Leela, Krishna Prasad Chitrapura, Sachin Garg, Amit Sasturkar
WSDM1
2009 URL normalization for de-duplication of web pages
abstract
Presence of duplicate documents in the World Wide Web adversely affects crawling, indexing and relevance, which are the core building blocks of web search. In this paper, we present a set of techniques to mine rules from URLs and utilize these learnt rules for de-duplication using just URL strings without fetching the content explicitly. Our technique is composed of mining the crawl logs and utilizing clusters of similar pages to extract specific rules from URLs belonging to each cluster. Preserving each mined rules for de-duplication is not efficient due to the large number of specific rules. We present a machine learning technique to generalize the set of rules, which reduces the resource footprint to be usable at web-scale. The rule extraction techniques are robust against web-site specific URL conventions. We demonstrate the effectiveness of our techniques through experimental evaluation.
Hema Swetha Koppula, Krishna P. Leela, Krishna Prasad Chitrapura, Sachin Garg, Pavan Kumar GM, Chittaranjan Haty, Amit Sasturkar
CIKM2