Deepak Ramachandran

dblp:80/703 · DBLP profile ↗
← Back
36ranked-venue papers
7as first author
21since 2021 · last 2025
0000-0001-5412-6133ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 33 · 6 first-author · 19 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 4 first-author · 5 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Focus-N-Fix: Region-Aware Fine-Tuning for Text-to-Image Generation
abstract
Text-to-image (T2I) generation has made significant advances in recent years, but challenges still remain in the generation of perceptual artifacts, misalignment with complex prompts, and safety. The prevailing approach to address these issues involves collecting human feedback on generated images, training reward models to estimate human feedback, and then fine-tuning T2I models based on the reward models to align them with human preferences. However, while existing reward fine-tuning methods can produce images with higher rewards, they may change model behavior in unexpected ways. For example, fine-tuning for one quality aspect (e.g., safety) may degrade other aspects (e.g., prompt alignment), or may lead to reward hacking (e.g., finding a way to increase rewards without having the intended effect). In this paper, we propose Focus-N-Fix, the first region-aware fine-tuning method that trains models to correct only previously problematic image regions. The resulting fine-tuned model generates images with the same high-level structure as the original model but shows significant improvements in regions where the original model was deficient in safety (over-sexualization and violence), plausibility, or other criteria. Our experiments demonstrate that Focus-N-Fix improves these localized quality aspects with little or no degradation to others and typically imperceptible changes in the rest of the image. Disclaimer: This paper contains images that may be overly sexual, violent, offensive or harmful.
Xiaoying Xing, Avinab Saha, Junfeng He, Susan Hao, Paul Vicol, Moonkyung Ryu, Gang Li 0021, Sahil Singla 0005, Sarah Young, Yinxiao Li, Feng Yang 0008, Deepak Ramachandran
CVPR12
2025 Erasing More Than Intended? How Concept Erasure Degrades the Generation of Non-Target Concepts
Ibtihel Amara, Ahmed Imtiaz Humayun, Ivana Kajic, Zarana Parekh, Natalie Harris, Sarah Young, Chirag Nagpal, Najoung Kim, Junfeng He, Cristina Nader Vasconcelos, Deepak Ramachandran, Golnoosh Farnadi, Katherine A. Heller, Mohammad Havaei, Negar Rostamzadeh
ICCV11
2025 What Secrets Do Your Manifolds Hold? Understanding the Local Geometry of Generative Models
abstract
Deep Generative Models are frequently used to learn continuous representations of complex data distributions by training on a finite number of samples. For any generative model, including pre-trained foundation models with Diffusion or Transformer architectures, generation performance can significantly vary across the learned data manifold. In this paper, we study the local geometry of the learned manifold and its relationship to generation outcomes for a wide range of generative models, including DDPM, Diffusion Transformer (DiT), and Stable Diffusion 1.4. Building on the theory of continuous piecewise-linear (CPWL) generators, we characterize the local geometry in terms of three geometric descriptors - scaling ($\psi$), rank ($\nu$), and complexity/un-smoothness ($\delta$). We provide quantitative and qualitative evidence showing that for a given latent vector, the local descriptors are indicative of post-generation aesthetics, generation diversity, and memorization by the generative model. Finally, we demonstrate that by training a reward model on the 'local scaling' for Stable Diffusion, we can self-improve both generation aesthetics and diversity using geometry sensitive guidance during denoising. Website: https://imtiazhumayun.github.io/generative_geometry.
Ahmed Imtiaz Humayun, Ibtihel Amara, Cristina Nader Vasconcelos, Deepak Ramachandran, Candice Schumann, Junfeng He, Katherine A. Heller, Golnoosh Farnadi, Negar Rostamzadeh, Mohammad Havaei
ICLR4
2025 Preference Adaptive and Sequential Text-to-Image Generation
abstract
We address the problem of interactive text-to-image (T2I) generation, designing a reinforcement learning (RL) agent which iteratively improves a set of generated images for a user through a sequence of prompt expansions. Using human raters, we create a novel dataset of sequential preferences, which we leverage, together with large-scale open-source (non-sequential) datasets. We construct user-preference and user-choice models using an EM strategy and identify varying user preference types. We then leverage a large multimodal language model (LMM) and a value-based RL approach to suggest an adaptive and diverse slate of prompt expansions to the user. Our Preference Adaptive and Sequential Text-to-image Agent (PASTA) extends T2I models with adaptive multi-turn capabilities, fostering collaborative co-creation and addressing uncertainty or underspecification in a user’s intent. We evaluate PASTA using human raters, showing significant improvement compared to baseline methods. We also open-source our sequential rater dataset and simulated user-rater interactions to support future research in user-centric multi-turn T2I systems.
Ofir Nabati, Guy Tennenholtz, Moonkyung Ryu, Deepak Ramachandran, Yinlam Chow, Craig Boutilier
ICML5
2025 Robust LLM Alignment via Distributionally Robust Direct Preference Optimization
abstract
A major challenge in aligning large language models (LLMs) with human preferences is the issue of distribution shift. LLM alignment algorithms rely on static preference datasets, assuming that they accurately represent real-world user preferences. However, user preferences vary significantly across geographical regions, demographics, linguistic patterns, and evolving cultural trends. This preference distribution shift leads to catastrophic alignment failures in many real-world applications. We address this problem using the principled framework of distributionally robust optimization, and develop two novel distributionally robust direct preference optimization (DPO) algorithms, namely, Wasserstein DPO (WDPO) and Kullback–Leibler DPO (KLDPO). We characterize the sample complexity of learning the optimal policy parameters for WDPO and KLDPO. Moreover, we propose scalable gradient descent-style learning algorithms by developing suitable approximations for the challenging minimax loss functions of WDPO and KLDPO. Our empirical experiments using benchmark data sets and LLMs demonstrate the superior performance of WDPO and KLDPO in substantially improving the alignment when there is a preference distribution shift.
Zaiyan Xu, Sushil Vemuri, Kishan Panaganti, Dileep M. Kalathil, Rahul Jain 0002, Deepak Ramachandran
NeurIPS6
2024 TaskLAMA: Probing the Complex Task Understanding of Language Models
abstract
Structured Complex Task Decomposition (SCTD) is the problem of breaking down a complex real-world task (such as planning a wedding) into a directed acyclic graph over individual steps that contribute to achieving the task, with edges specifying temporal dependencies between steps. SCTD is an important component of assistive planning tools, and a challenge for commonsense reasoning systems. We probe how accurately SCTD can be done with the knowledge extracted from pre-trained Large Language Models (LLMs). We introduce a new high-quality human-annotated dataset for this problem and novel metrics to fairly assess performance of LLMs against several baselines. Our experiments reveal that LLMs are able to decompose complex tasks into individual steps effectively, with a relative improvement of 15% to 280% over the best baseline. We also propose a number of approaches to further improve their performance, with a relative improvement of 7% to 37%. However, we find that LLMs still struggle to predict pairwise temporal dependencies, which reveals a gap in their understanding of complex tasks.
Quan Yuan 0001, Mehran Kazemi, Isaac Noble, Vaiva Imbrasaite, Deepak Ramachandran
AAAI6
2024 Prompt Expansion for Adaptive Text-to-Image Generation
abstract
Text-to-image generation models are powerful but difficult to use. Users craft specific prompts to get better images, though the images can be repetitive. This paper proposes the Prompt Expansion framework that helps users generate high-quality, diverse images with less effort. The Prompt Expansion model takes a text query as input and outputs a set of expanded text prompts that are optimized such that when passed to a text-to-image model, they generate a wider variety of appealing images. We conduct a human evaluation study that shows that images generated through Prompt Expansion are more aesthetically pleasing and diverse than those generated by baseline methods. Overall, this paper presents a novel and effective approach to improving the text-to-image generation experience.
Siddhartha Datta, Alexander Ku, Deepak Ramachandran
ACL (1)3
2024 Beyond Thumbs Up/Down: Untangling Challenges of Fine-Grained Feedback for Text-to-Image Generation
abstract
Human feedback plays a critical role in learning and refining reward models for text-to-image generation, but the optimal form the feedback should take for learning an accurate reward function has not been conclusively established. This paper investigates the effectiveness of fine-grained feedback which captures nuanced distinctions in image quality and prompt-alignment, compared to traditional coarse-grained feedback (for example, thumbs up/down or ranking between a set of options). While fine-grained feedback holds promise, particularly for systems catering to diverse societal preferences, we show that demonstrating its superiority to coarse-grained feedback is not automatic. Through experiments on real and synthetic preference data, we surface the complexities of building effective models due to the interplay of model choice, feedback type, and the alignment between human judgment and computational interpretation. We identify key challenges in eliciting and utilizing fine-grained feedback, prompting a reassessment of its assumed benefits and practicality. Our findings -- e.g., that fine-grained feedback can lead to worse models for a fixed budget, in some settings; however, in controlled settings with known attributes, fine grained rewards can indeed be more helpful -- call for careful consideration of feedback attributes and potentially beckon novel modeling approaches to appropriately unlock the potential value of fine-grained feedback in-the-wild.
Katie Collins, Najoung Kim, Yonatan Bitton, Verena Rieser, Shayegan Omidshafiei, Yushi Hu, Sherol Chen, Senjuti Dutta, Minsuk Chang, Kimin Lee, Youwei Liang, Georgina Evans, Sahil Singla 0005, Gang Li 0021, Adrian Weller, Junfeng He, Deepak Ramachandran, Krishnamurthy Dvijotham
AIES (1)17
2024 Rich Human Feedback for Text-to-Image Generation
abstract
Recent Text-to-Image (T2I) generation models such as Stable Diffusion and Imagen have made significant progress in generating high-resolution images based on text descriptions. However, many generated images still suffer from issues such as artifacts/implausibility, misalignment with text descriptions, and low aesthetic quality. Inspired by the success of Reinforcement Learning with Human Feedback (RLHF) for large language models, prior works collected human-provided scores as feedback on generated images and trained a reward model to improve the T2I generation. In this paper, we enrich the feedback signal by (i) marking image regions that are implausible or misaligned with the text, and (ii) annotating which words in the text prompt are misrepresented or missing on the image. We collect such rich human feedback on 18K generated images (RichHF-18K) and train a multimodal transformer to predict the rich feedback automatically. We show that the predicted rich human feedback can be leveraged to improve image generation, for example, by selecting high-quality training data to finetune and improve the generative models, or by creating masks with predicted heatmaps to inpaint the problematic regions. Notably, the improvements generalize to models (Muse) beyond those used to generate the images on which human feedback data were collected (Stable Diffusion variants). The RichHF-18K data set will be released in our GitHub repository: https://github.com/google-research/google-research/tree/master/richhf_18k.
Youwei Liang, Junfeng He, Gang Li 0021, Peizhao Li, Arseniy Klimovskiy, Nicholas Carolan, Jiao Sun, Jordi Pont-Tuset, Sarah Young, Feng Yang 0008, Junjie Ke, Krishnamurthy Dvijotham, Katie Collins, Yiwen Luo, Yang Li 0058, Kai Kohlhoff, Deepak Ramachandran, Vidhya Navalpakkam
CVPR17
2024 Demystifying Embedding Spaces using Large Language Models
abstract
Embeddings have become a pivotal means to represent complex, multi-faceted information about entities, concepts, and relationships in a condensed and useful format. Nevertheless, they often preclude direct interpretation. While downstream tasks make use of these compressed representations, meaningful interpretation usually requires visualization using dimensionality reduction or specialized machine learning interpretability methods. This paper addresses the challenge of making such embeddings more interpretable and broadly useful, by employing large language models (LLMs) to directly interact with embeddings -- transforming abstract vectors into understandable narratives. By injecting embeddings into LLMs, we enable querying and exploration of complex embedding data. We demonstrate our approach on a variety of diverse tasks, including: enhancing concept activation vectors (CAVs), communicating novel embedded entities, and decoding user preferences in recommender systems. Our work couples the immense information potential of embeddings with the interpretative power of LLMs.
Guy Tennenholtz, Yinlam Chow, Jihwan Jeong, Lior Shani, Aza Tulepbergenov, Deepak Ramachandran, Martin Mladenov, Craig Boutilier
ICLR7
2024 e-COP : Episodic Constrained Optimization of Policies
abstract
In this paper, we present the e-COP algorithm, the first policy optimization algorithm for constrained Reinforcement Learning (RL) in episodic (finite horizon) settings. Such formulations are applicable when there are separate sets of optimization criteria and constraints on a system's behavior. We approach this problem by first establishing a policy difference lemma for the episodic setting, which provides the theoretical foundation for the algorithm. Then, we propose to combine a set of established and novel solution ideas to yield the e-COP algorithm that is easy to implement and numerically stable, and provide a theoretical guarantee on optimality under certain scaling assumptions. Through extensive empirical analysis using benchmarks in the Safety Gym suite, we show that our algorithm has similar or better performance than SoTA (non-episodic) algorithms adapted for the episodic setting. The scalability of the algorithm opens the door to its application in safety-constrained Reinforcement Learning from Human Feedback for Large Language or Diffusion Models.
Akhil Agnihotri, Rahul Jain 0002, Deepak Ramachandran, Sahil Singla 0005
NeurIPS3
2024 Discovering Personalized Semantics for Soft Attributes in Recommender Systems Using Concept Activation Vectors
abstract
Interactive recommender systems have emerged as a promising paradigm to overcome the limitations of the primitive user feedback used by traditional recommender systems (e.g., clicks, item consumption, ratings). They allow users to express intent, preferences, constraints, and contexts in a richer fashion, often using natural language (including faceted search and dialogue). Yet more research is needed to find the most effective ways to use this feedback. One challenge is inferring a user’s semantic intent from the open-ended terms or attributes often used to describe a desired item. This is critical for recommender systems that wish to support users in their everyday, intuitive use of natural language to refine recommendation results. Leveraging concept activation vectors (CAVs) [ 26 ], a recently developed approach for model interpretability in machine learning, we develop a framework to learn a representation that captures the semantics of such attributes and connects them to user preferences and behaviors in recommender systems. One novel feature of our approach is its ability to distinguish objective and subjective attributes (both subjectivity of degree and of sense ) and associate different senses of subjective attributes with different users. We demonstrate on both synthetic and real-world datasets that our CAV representation not only accurately interprets users’ subjective semantics but also can be used to improve recommendations through interactive item critiquing .
Christina Göpfert, Alex Haig, Yinlam Chow, Ivan Vendrov, Tyler Lu, Deepak Ramachandran, Hubert Pham, Mohammad Ghavamzadeh, Craig Boutilier
Trans. Recomm. Syst.7
2023 LAMBADA: Backward Chaining for Automated Reasoning in Natural Language
abstract
Remarkable progress has been made on automated reasoning with natural text, by using Language Models (LMs) and methods such as Chain-of-Thought and Selection-Inference.These techniques search for proofs in the forward direction from axioms to the conclusion, which suffers from a combinatorial explosion of the search space, and thus high failure rates for problems requiring longer chains of reasoning.The classical automated reasoning literature has shown that reasoning in the backward direction (i.e. from the intended conclusion to supporting axioms) is significantly more efficient at proof-finding.Importing this intuition into the LM setting, we develop a Backward Chaining algorithm, called LAM-BADA, that decomposes reasoning into four sub-modules.These sub-modules are simply implemented by few-shot prompted LM inference.We show that LAMBADA achieves sizable accuracy boosts over state-of-the-art forward reasoning methods on two challenging logical reasoning datasets, particularly when deep and accurate proof chains are required. Facts:1. Rough and cold that is what they say about Blue Bob. 2. Eric, who is relatively young, is also pretty big and tends to be cold.3. Fred is green and cold too.4. For being so cold, it's good Harry can remain nice.Rules: 1. Rough, cold people are blue.2. Big, kind folks are green ones.3.If a person is big, rough, and cold, they are also red. 4. Most round and cold people are often rough.5. Cold, young people are also certain to be rough people.6.An individual who is big, red and young is also a nice individual.
Mehran Kazemi, Najoung Kim, Deepti Bhatia, Deepak Ramachandran
ACL (1)5
2023 Using Domain Knowledge to Guide Dialog Structure Induction via Neural Probabilistic Soft Logic
abstract
Connor Pryor, Quan Yuan, Jeremiah Liu, Mehran Kazemi, Deepak Ramachandran, Tania Bedrax-Weiss, Lise Getoor. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Connor Pryor, Quan Yuan 0001, Jeremiah Z. Liu, Mehran Kazemi, Deepak Ramachandran, Tania Bedrax-Weiss, Lise Getoor
ACL (1)5
2023 Pushing the Accuracy-Group Robustness Frontier with Introspective Self-play
Jeremiah Z. Liu, Krishnamurthy Dvijotham, Jihyeon Lee, Quan Yuan 0001, Balaji Lakshminarayanan, Deepak Ramachandran
ICLR6
2023 KwikBucks: Correlation Clustering with Cheap-Weak and Expensive-Strong Signals
Sandeep Silwal, Sara Ahmadian, Andrew Nystrom, Andrew McCallum, Deepak Ramachandran, Mehran Kazemi
ICLR5
2023 BoardgameQA: A Dataset for Natural Language Reasoning with Contradictory Information
abstract
Automated reasoning with unstructured natural text is a key requirement for many potential applications of NLP and for developing robust AI systems. Recently, Language Models (LMs) have demonstrated complex reasoning capacities even without any finetuning. However, existing evaluation for automated reasoning assumes access to a consistent and coherent set of information over which models reason. When reasoning in the real-world, the available information is frequently inconsistent or contradictory, and therefore models need to be equipped with a strategy to resolve such conflicts when they arise. One widely-applicable way of resolving conflicts is to impose preferences over information sources (e.g., based on source credibility or information recency) and adopt the source with higher preference. In this paper, we formulate the problem of reasoning with contradictory information guided by preferences over sources as the classical problem of defeasible reasoning, and develop a dataset called BoardgameQA for measuring the reasoning capacity of LMs in this setting. BoardgameQA also incorporates reasoning with implicit background knowledge, to better reflect reasoning problems in downstream applications. We benchmark various LMs on BoardgameQA and the results reveal a significant gap in the reasoning capacity of state-of-the-art LMs on this problem, showing that reasoning with conflicting information does not surface out-of-the-box in LMs. While performance can be improved with finetuning, it nevertheless remains poor.
Mehran Kazemi, Quan Yuan 0001, Deepti Bhatia, Najoung Kim, Vaiva Imbrasaite, Deepak Ramachandran
NeurIPS7
2022 Subjective Attributes in Conversational Recommendation Systems: Challenges and Opportunities
abstract
The ubiquity of recommender systems has increased the need for higher-bandwidth, natural and efficient communication with users. This need is increasingly filled by recommenders that support natural language interaction, often conversationally. Given the inherent semantic subjectivity present in natural language, we argue that modeling subjective attributes in recommenders is a critical, yet understudied, avenue of AI research. We propose a novel framework for understanding different forms of subjectivity, examine various recommender tasks that will benefit from a systematic treatment of subjective attributes, and outline a number of research challenges.
Filip Radlinski, Craig Boutilier, Deepak Ramachandran, Ivan Vendrov
AAAI3
2022 FETA: A Benchmark for Few-Sample Task Transfer in Open-Domain Dialogue
abstract
Alon Albalak, Yi-Lin Tuan, Pegah Jandaghi, Connor Pryor, Luke Yoffe, Deepak Ramachandran, Lise Getoor, Jay Pujara, William Yang Wang. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022.
Alon Albalak, Yi-Lin Tuan, Pegah Jandaghi, Connor Pryor, Luke Yoffe, Deepak Ramachandran, Lise Getoor, Jay Pujara, William Yang Wang
EMNLP6
2022 Discovering Personalized Semantics for Soft Attributes in Recommender Systems using Concept Activation Vectors
abstract
Interactive recommender systems (RSs) allow users to express intent, preferences and contexts in a rich fashion, often using natural language. One challenge in using such feedback is inferring a user’s semantic intent from the open-ended terms used to describe an item, and using it to refine recommendation results. Leveraging concept activation vectors (CAVs) [21], we develop a framework to learn a representation that captures the semantics of such attributes and connects them to user preferences and behaviors in RSs. A novel feature of our approach is its ability to distinguish objective and subjective attributes and associate different senses with different users. Using synthetic and real-world datasets, we show that our CAV representation accurately interprets users’ subjective semantics, and can improve recommendations via interactive critiquing.
Christina Göpfert, Yinlam Chow, Ivan Vendrov, Tyler Lu, Deepak Ramachandran, Craig Boutilier
WWW6
2021 Which Linguist Invented the Lightbulb? Presupposition Verification for Question-Answering
abstract
Najoung Kim, Ellie Pavlick, Burcu Karagol Ayan, Deepak Ramachandran. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Najoung Kim, Ellie Pavlick, Burcu Karagol Ayan, Deepak Ramachandran
ACL/IJCNLP (1)4
2019 How Large Are Lions? Inducing Distributions over Quantitative Attributes
abstract
Most current NLP systems have little knowledge about quantitative attributes of objects and events.We propose an unsupervised method for collecting quantitative information from large amounts of web data, and use it to create a new, very large resource consisting of distributions over physical quantities associated with objects, adjectives, and verbs which we call Distribution over Quantities (DOQ) 1 .This contrasts with recent work in this area which has focused on making only relative comparisons such as "Is a lion bigger than a wolf?".Our evaluation shows that DOQ compares favorably with state of the art results on existing datasets for relative comparisons of nouns and adjectives, and on a new dataset we introduce.* Work carried out during an internship at Google.† Work carried out during employment at Google. 1 The resource is available at https:// github.com/google-research-datasets/ distribution-over-quantities
Yanai Elazar, Abhijit Mahabal, Deepak Ramachandran, Tania Bedrax-Weiss, Dan Roth 0001
ACL (1)3
2015 A TV Program Discovery Dialog System using recommendations
abstract
Deepak Ramachandran, Mark Fanty, Ronald Provine, Peter Yeh, William Jarrold, Adwait Ratnaparkhi, Benjamin Douglas. Proceedings of the 16th Annual Meeting of the Special Interest Group on Discourse and Dialogue. 2015.
Deepak Ramachandran, Mark A. Fanty, Ronald Provine, Peter Z. Yeh, William Jarrold, Adwait Ratnaparkhi, Benjamin Douglas
SIGDIAL Conference1
2015 Belief Tracking with Stacked Relational Trees
abstract
We describe a new model for Dialog State Tracking called a Stacked Relational Tree, which naturally models complex relationships between entities across user utterances.It can represent multiple conversational intents and the change of focus between them.Updates to the model are made by a rule-based system in the language of tree regular expressions.We also introduce a probabilistic version that can handle ASR/NLU uncertainty.We show how the parameters can be trained from log data, showing gains on a variety of standard Belief Tracker metrics, and a measurable impact on the success rate of an end-to-end dialog system for TV program discovery.
Deepak Ramachandran, Adwait Ratnaparkhi
SIGDIAL Conference1
2014 A Speech-Driven Second Screen Application for TV Program Discovery
abstract
In this paper, we present a speech-driven second screen application for TV program discovery. We give an overview of the application and its architecture. We also present a user study along with a failure analysis. The results from the study are encouraging, and demonstrate our application’s effectiveness in the target domain. We conclude with a discussion of follow-on efforts to further enhance our application.
Peter Z. Yeh, Benjamin Douglas, William Jarrold, Adwait Ratnaparkhi, Deepak Ramachandran, Peter F. Patel-Schneider, Stephen Laverty, Nirvana Tikku, Sean Brown 0001, Jeremy Mendel
AAAI5
2014 An end-to-end dialog system for TV program discovery
abstract
In this paper, we present an end-to-end dialog system for TV program discovery that uniquely combines several technologies such as trainable relation extraction, belief tracking over relational structures, mixed-initiative dialog management, and inference over large-scale knowledge graphs. We present an evaluation of our end-to-end system with real users, and found that our system performed well along several dimensions such as usability and task success rate. These results demonstrate the effectiveness of our system in the target domain.
Deepak Ramachandran, Peter Z. Yeh, William Jarrold, Benjamin Douglas, Adwait Ratnaparkhi, Ronald Provine, Jeremy Mendel, Adam Emfield
SLT1
2013 The Dialog State Tracking Challenge
Jason D. Williams, Antoine Raux, Deepak Ramachandran, Alan W. Black
SIGDIAL Conference3
2012 Improving Hybrid Vehicle Fuel Efficiency Using Inverse Reinforcement Learning
abstract
Deciding what mix of engine and battery power to use is critical to hybrid vehicles' fuel efficiency. Current solutions consider several factors such as the charge of the battery and how efficient the engine operates at a given speed. Previous research has shown that by taking into account the future power requirements of the vehicle, a more efficient balance of engine vs. battery power can be attained. In this paper, we utilize a probabilistic driving route prediction system, trained using Inverse Reinforcement Learning, to optimize the hybrid control policy. Our approach considers routes that the driver is likely to be taking, computing an optimal mix of engine and battery power. This approach has the potential to increase vehicle power efficiency while not requiring any hardware modification or change in driver behavior. Our method outperforms a standard hybrid control policy, yielding an average of 1.22% fuel savings.
Adam Vogel, Deepak Ramachandran, Rakesh Gupta 0001, Antoine Raux
AAAI2
2012 Landmark-Based Location Belief Tracking in a Spoken Dialog System
Antoine Raux, Deepak Ramachandran, Rakesh Gupta 0001
SIGDIAL Conference3
2010 Dynamic language modeling using Bayesian networks for spoken dialog systems
Antoine Raux, Neville Mehta, Deepak Ramachandran, Rakesh Gupta 0001
INTERSPEECH3
2010 Probabilistic Ontology Trees for Belief Tracking in Dialog Systems
Neville Mehta, Rakesh Gupta 0001, Antoine Raux, Deepak Ramachandran, Stefan Krawczyk
SIGDIAL Conference4
2009 Smoothed Sarsa: Reinforcement learning for robot delivery tasks
abstract
Our goal in this work is to make high level decisions for mobile robots. In particular, given a queue of prioritized object delivery tasks, we wish to find a sequence of actions in real time to accomplish these tasks efficiently. We introduce a novel reinforcement learning algorithm called Smoothed Sarsa that learns a good policy for these delivery tasks by delaying the backup reinforcement step until the uncertainty in the state estimate improves. The state space is modeled by a Dynamic Bayesian Network and updated using a Region-based Particle Filter. We take advantage of the fact that only discrete (topological) representations of entity locations are needed for decision-making, to make the tracking and decision making more efficient. Our experiments show that policy search leads to faster task completion times as well as higher total reward compared to a manually crafted policy. Smoothed Sarsa learns a policy orders of magnitude faster than previous policy search algorithms. We demonstrate our results on the Player/Stage simulator and on the Pioneer robot.
Deepak Ramachandran, Rakesh Gupta 0001
ICRA1
2008 ManifoldBoost: stagewise function approximation for fully-, semi- and un-supervised learning
abstract
We describe a manifold learning framewor that naturally accommodates supervised learning, partially supervised learning and unsupervised clustering as particular cases. Our method chooses a function by minimizing loss subject to a manifold regularization penalty. This augmented cost is minimized using a greedy, stagewise, functional minimization procedure, as in Gradientboost. Each stage of boosting is fast and efficient. We demonstrate our approach using both radial basis function approximations and trees. The performance of our method is at the state of the art on many standard semi-supervised learning benchmarks, and we produce results for large scale datasets.
Nicolas Loeff, David A. Forsyth, Deepak Ramachandran
ICML3
2007 Bayesian Inverse Reinforcement Learning
Deepak Ramachandran, Eyal Amir
IJCAI1
2005 Compact Propositional Encodings of First-Order Theories
Deepak Ramachandran, Eyal Amir
AAAI1
2005 Compact Propositional Encodings of First-Order Theories
Deepak Ramachandran, Eyal Amir
IJCAI1