Alborz Geramifard

dblp:39/3250 · DBLP profile ↗
← Back
23ranked-venue papers
5as first author
12since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 23 · 5 first-author · 12 since 2021Databases, data management, data science and information retrieval · 2Systems, architecture and hardware · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
YearPublicationVenuePosition
2024 When should we prefer Decision Transformers for Offline Reinforcement Learning?
abstract
Offline reinforcement learning (RL) allows agents to learn effective, return-maximizing policies from a static dataset. Three popular algorithms for offline RL are Conservative Q-Learning (CQL), Behavior Cloning (BC), and Decision Transformer (DT), from the class of Q-Learning, Imitation Learning, and Sequence Modeling respectively. A key open question is: which algorithm is preferred under what conditions? We study this question empirically by exploring the performance of these algorithms across the commonly used D4RL and Robomimic benchmarks. We design targeted experiments to understand their behavior concerning data suboptimality, task complexity, and stochasticity. Our key findings are: (1) DT requires more data than CQL to learn competitive policies but is more robust; (2) DT is a substantially better choice than both CQL and BC in sparse-reward and low-quality data settings; (3) DT and BC are preferable as task horizon increases, or when data is obtained from human demonstrators; and (4) CQL excels in situations characterized by the combination of high stochasticity and low data quality. We also investigate architectural choices and scaling trends for DT on \textsc{atari} and D4RL and make design/scaling recommendations. We find that scaling the amount of data for DT by 5x gives a 2.5x average score improvement on Atari.
Prajjwal Bhargava, Rohan Chitnis, Alborz Geramifard, Shagun Sodhani, Amy Zhang 0001
ICLR3
2024 Score Models for Offline Goal-Conditioned Reinforcement Learning
abstract
Offline Goal-Conditioned Reinforcement Learning (GCRL) is tasked with learning to achieve multiple goals in an environment purely from offline datasets using sparse reward functions. Offline GCRL is pivotal for developing generalist agents capable of leveraging pre-existing datasets to learn diverse and reusable skills without hand-engineering reward functions. However, contemporary approaches to GCRL based on supervised learning and contrastive learning are often suboptimal in the offline setting. An alternative perspective on GCRL optimizes for occupancy matching, but necessitates learning a discriminator, which subsequently serves as a pseudo-reward for downstream RL. Inaccuracies in the learned discriminator can cascade, negatively influencing the resulting policy. We present a novel approach to GCRL under a new lens of mixture-distribution matching, leading to our discriminator-free method: SMORe. The key insight is combining the occupancy matching perspective of GCRL with a convex dual formulation to derive a learning objective that can better leverage suboptimal offline data. SMORe learns *scores* or unnormalized densities representing the importance of taking an action at a state for reaching a particular goal. SMORe is principled and our extensive experiments on the fully offline GCRL benchmark composed of robot manipulation and locomotion tasks, including high-dimensional observations, show that SMORe can outperform state-of-the-art baselines by a significant margin.
Harshit Sikchi, Rohan Chitnis, Ahmed Touati, Alborz Geramifard, Amy Zhang 0001, Scott Niekum
ICLR4
2024 Overview of the Ninth Dialog System Technology Challenge: DSTC9
abstract
This paper introduces the Ninth Dialog System Technology Challenge (DSTC-9). This edition of the DSTC focuses on applying end-to-end dialog technologies for four distinct tasks in dialog systems, namely, 1. Task-oriented dialog Modeling with Unstructured Knowledge Access, 2. Multi-domain task-oriented dialog, 3. Interactive evaluation of dialog and 4. Situated interactive multimodal dialog. This paper describes the task definition, provided datasets, baselines, and evaluation setup for each track. We also summarize the results of the submitted systems to highlight the general trends of the state-of-the-art technologies for the tasks.
R. Chulaka Gunasekara, Seokhwan Kim, Luis Fernando D'Haro, Abhinav Rastogi, Yun-Nung Chen, Mihail Eric, Behnam Hedayatnia, Karthik Gopalakrishnan 0001, Yang Liu 0004, Chao-Wei Huang, Dilek Hakkani-Tür, Jinchao Li, Qi Zhu 0007, Lingxiao Luo, Lars Liden, Kaili Huang, Shahin Shayandeh, Runze Liang, Baolin Peng, Zheng Zhang 0020, Swadheen Shukla, Minlie Huang, Jianfeng Gao 0001, Shikib Mehri, Yulan Feng, Carla Gordon, Seyed Hossein Alavi, David R. Traum, Maxine Eskénazi, Ahmad Beirami, Eunjoon Cho, Paul A. Crook, Ankita De, Alborz Geramifard, Satwik Kottur, Seungwhan Moon, Shivani Poddar, Rajen Subba
IEEE ACM Trans. Audio Speech Lang. Process.34
2024 Overview of the Tenth Dialog System Technology Challenge: DSTC10
abstract
This article introduces the Tenth Dialog System Technology Challenge (DSTC-10). This edition of the DSTC focuses on applying end-to-end dialog technologies for five distinct tasks in dialog systems, namely 1. Incorporation of Meme images into open domain dialogs, 2. Knowledge-grounded Task-oriented Dialogue Modeling on Spoken Conversations, 3. Situated Interactive Multimodal dialogs, 4. Reasoning for Audio Visual Scene-Aware Dialog, and 5. Automatic Evaluation and Moderation of Open-domainDialogue Systems. This article describes the task definition, provided datasets, baselines, and evaluation setup for each track. We also summarize the results of the submitted systems to highlight the general trends of the state-of-the-art technologies for the tasks.
Koichiro Yoshino, Yun-Nung Chen, Paul A. Crook, Satwik Kottur, Jinchao Li, Behnam Hedayatnia, Seungwhan Moon, Zhengcong Fei, Zekang Li, Jinchao Zhang 0001, Yang Feng 0004, Jie Zhou 0016, Seokhwan Kim, Yang Liu 0004, Di Jin 0005, Alexandros Papangelis, Karthik Gopalakrishnan 0001, Dilek Hakkani-Tür, Babak Damavandi, Alborz Geramifard, Chiori Hori, Chen Zhang 0020, Haizhou Li 0001, João Sedoc, Luis Fernando D'Haro, Rafael E. Banchs, Alexander I. Rudnicky
IEEE ACM Trans. Audio Speech Lang. Process.20
2022 Navigating Connected Memories with a Task-oriented Dialog System
abstract
Recent years have seen an increasing trend in the volume of personal media captured by users, thanks to the advent of smartphones and smart glasses, resulting in large media collections.Despite conversation being an intuitive humancomputer interface, current efforts focus mostly on single-shot natural language based media retrieval to aid users query their media and re-live their memories.This severely limits the search functionality as users can neither ask followup queries nor obtain information without first formulating a single-turn query.In this work, we propose dialogs for connected memories as a powerful tool to empower users to search their media collection through a multiturn, interactive conversation.Towards this, we collect a new task-oriented dialog dataset COMET, which contains 11.5k user↔assistant dialogs (totalling 103k utterances), grounded in simulated personal memory graphs.We employ a resource-efficient, two-phase data collection pipeline that uses: (1) a novel multimodal dialog simulator that generates synthetic dialog flows grounded in memory graphs, and, (2) manual paraphrasing to obtain natural language utterances.We analyze COMET, formulate four main tasks to benchmark meaningful progress, and adopt state-of-the-art language models as strong baselines, in order to highlight the multimodal challenges captured by our dataset 1 .
Satwik Kottur, Seungwhan Moon, Alborz Geramifard, Babak Damavandi
EMNLP3
2022 Database Search Results Disambiguation for Task-Oriented Dialog Systems
abstract
Kun Qian, Satwik Kottur, Ahmad Beirami, Shahin Shayandeh, Paul Crook, Alborz Geramifard, Zhou Yu, Chinnadhurai Sankar. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Kun Qian 0016, Satwik Kottur, Ahmad Beirami, Shahin Shayandeh, Paul A. Crook, Alborz Geramifard, Zhou Yu 0005, Chinnadhurai Sankar
NAACL-HLT6
2021 DVD: A Diagnostic Dataset for Multi-step Reasoning in Video Grounded Dialogue
abstract
Hung Le, Chinnadhurai Sankar, Seungwhan Moon, Ahmad Beirami, Alborz Geramifard, Satwik Kottur. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Hung Le 0003, Chinnadhurai Sankar, Seungwhan Moon, Ahmad Beirami, Alborz Geramifard, Satwik Kottur
ACL/IJCNLP (1)5
2021 SIMMC 2.0: A Task-oriented Dialog Dataset for Immersive Multimodal Conversations
abstract
Next generation task-oriented dialog systems need to understand conversational contexts with their perceived surroundings, to effectively help users in the real-world multimodal environment.Existing task-oriented dialog datasets aimed towards virtual assistance fall short and do not situate the dialog in the user's multimodal context.To overcome, we present a new dataset for Situated and Interactive Multimodal Conversations, SIMMC 2.0, which includes 11K taskoriented user$assistant dialogs (117K utterances) in the shopping domain, grounded in immersive and photo-realistic scenes.The dialogs are collected using a two-phase pipeline: (1) A novel multimodal dialog simulator generates simulated dialog flows, with an emphasis on diversity and richness of interactions, (2) Manual paraphrasing of the generated utterances to collect diverse referring expressions.We provide an in-depth analysis of the collected dataset, and describe in detail the four main benchmark tasks we propose.Our baseline model, powered by the state-of-theart language model, shows promising results, and highlights new challenges and directions for the community to study 1 .
Satwik Kottur, Seungwhan Moon, Alborz Geramifard, Babak Damavandi
EMNLP (1)3
2021 An Analysis of State-of-the-Art Models for Situated Interactive MultiModal Conversations (SIMMC)
abstract
Satwik Kottur, Paul Crook, Seungwhan Moon, Ahmad Beirami, Eunjoon Cho, Rajen Subba, Alborz Geramifard. Proceedings of the 22nd Annual Meeting of the Special Interest Group on Discourse and Dialogue. 2021.
Satwik Kottur, Paul A. Crook, Seungwhan Moon, Ahmad Beirami, Eunjoon Cho, Rajen Subba, Alborz Geramifard
SIGDIAL7
2021 DialogStitch: Synthetic Deeper and Multi-Context Task-Oriented Dialogs
abstract
Real-world conversational agents must effectively handle long conversations that span multiple contexts.Such context can be interspersed with chitchat (dialog turns not directly related to the task at hand), and potentially grounded in a multimodal setting.While prior work focused on the above aspects in isolation, there is a lack of a unified framework that studies them together.To overcome this, we propose DialogStitch, a novel framework to seamlessly 'stitch' multiple conversations and highlight these desirable traits in a task-oriented dialog.After stitching, our dialogs are provably deeper, contain longer-term dependencies, and span multiple contexts, when compared with the source dialogs-all by leveraging existing human annotations!Though our framework generalizes to a variety of combinations, we demonstrate its benefits in two settings: (a) multimodal, image-grounded conversations, and, (b) task-oriented dialogs fused with chit-chat conversations.We benchmark state-of-the-art dialog models on our datasets and find accuracy drops of (a) 12% and (b) 45% respectively, indicating the additional challenges in the stitched dialogs.Our code and data are publicly available 1 .
Satwik Kottur, Chinnadhurai Sankar, Alborz Geramifard
SIGDIAL4
2021 Annotation Inconsistency and Entity Bias in MultiWOZ
abstract
MultiWOZ (Budzianowski et al., 2018) is one of the most popular multi-domain taskoriented dialog datasets, containing 10K+ annotated dialogs covering eight domains.It has been widely accepted as a benchmark for various dialog tasks, e.g., dialog state tracking (DST), natural language generation (NLG) and end-to-end (E2E) dialog modeling.In this work, we identify an overlooked issue with dialog state annotation inconsistencies in the dataset, where a slot type is tagged inconsistently across similar dialogs leading to confusion for DST modeling.We propose an automated correction for this issue, which is present in 70% of the dialogs.Additionally, we notice that there is significant entity bias in the dataset (e.g., "cambridge" appears in 50% of the destination cities in the train domain).The entity bias can potentially lead to named entity memorization in generative models, which may go unnoticed as the test set suffers from a similar entity bias as well.We release a new test set with all entities replaced with unseen entities.Finally, we benchmark joint goal accuracy (JGA) of the state-of-theart DST baselines on these modified versions of the data.Our experiments show that the annotation inconsistency corrections lead to 7-10% improvement in JGA.On the other hand, we observe a 29% drop in JGA when models are evaluated on the new test set with unseen entities.* The work of KQ and ZY was done as a research intern and a visiting research scientist at Facebook AI.
Kun Qian 0016, Ahmad Beirami, Zhouhan Lin, Ankita De, Alborz Geramifard, Zhou Yu 0005, Chinnadhurai Sankar
SIGDIAL5
2021 Guest editorial: special issue on reinforcement learning for real life
Alborz Geramifard, Lihong Li 0001, Csaba Szepesvári
Mach. Learn.2
2020 Situated and Interactive Multimodal Conversations
abstract
Seungwhan Moon, Satwik Kottur, Paul Crook, Ankita De, Shivani Poddar, Theodore Levin, David Whitney, Daniel Difranco, Ahmad Beirami, Eunjoon Cho, Rajen Subba, Alborz Geramifard. Proceedings of the 28th International Conference on Computational Linguistics. 2020.
Seungwhan Moon, Satwik Kottur, Paul A. Crook, Ankita De, Shivani Poddar, Theodore Levin, David Whitney, Daniel Difranco, Ahmad Beirami, Eunjoon Cho, Rajen Subba, Alborz Geramifard
COLING12
2020 Resource Constrained Dialog Policy Learning Via Differentiable Inductive Logic Programming
abstract
Motivated by the needs of resource constrained dialog policy learning, we introduce dialog policy via differentiable inductive logic (DILOG).We explore the tasks of one-shot learning and zero-shot domain transfer with DILOG on SimDial and MultiWoZ.Using a single representative dialog from the restaurant domain, we train DILOG on the SimDial dataset and obtain 99+% in-domain test accuracy.We also show that the trained DILOG zero-shot transfers to all other domains with 99+% accuracy, proving the suitability of DILOG to slot-filling dialogs.We further extend our study to the MultiWoZ dataset achieving 90+% inform and success metrics.We also observe that these metrics are not capturing some of the shortcomings of DILOG in terms of false positives, prompting us to measure an auxiliary Action F1 score.We show that DILOG is 100x more data efficient than state-of-the-art neural approaches on MultiWoZ while achieving similar performance metrics.We conclude with a discussion on the strengths and weaknesses of DILOG.
Zhenpeng Zhou, Ahmad Beirami, Paul A. Crook, Pararth Shah, Rajen Subba, Alborz Geramifard
COLING6
2017 The Future of Artificially Intelligent Assistants
abstract
Artificial Intelligence has been present in literature at least since the ancient Greeks. Depictions present a wide range of perspectives of AI ranging from malefic overlords to depressive androids. Perhaps the most common recurring theme is the AI Assistant: C3PO from Star Wars; the Jetson's Rosie the Robot; the benign hyper-efficient Minds of Iain M. Banks's Culture novels; the eerie HAL 9000 of Arthur C. Clarke's 2001: A Space Odyssey. Today, artificially intelligent assistants are actual products in the marketplace, based on startling recent progress in technologies like speaker-independent speech recognition. These products are in their infancy, but are improving rapidly. In this panel, we will address the product and technology landscape, and will ask a series of experts in the field plus the members of the audience to take a stance on what the future of artificially intelligent assistants will look like.
Muthu Muthukrishnan, Andrew Tomkins, Larry Heck, Alborz Geramifard, Deepak Agarwal
KDD4
2015 RLPy: a value-function-based reinforcement learning framework for education and research
Alborz Geramifard, Christoph Dann, Robert H. Klein, William Dabney, Jonathan P. How
J. Mach. Learn. Res.1
2013 Reinforcement learning with misspecified model classes
abstract
Real-world robots commonly have to act in complex, poorly understood environments where the true world dynamics are unknown. To compensate for the unknown world dynamics, we often provide a class of models to a learner so it may select a model, typically using a minimum prediction error metric over a set of training data. Often in real-world domains the model class is unable to capture the true dynamics, due to either limited domain knowledge or a desire to use a small model. In these cases we call the model class misspecified, and an unfortunate consequence of misspecification is that even with unlimited data and computation there is no guarantee the model with minimum prediction error leads to the best performing policy. In this work, our approach improves upon the standard maximum likelihood model selection metric by explicitly selecting the model which achieves the highest expected reward, rather than the most likely model. We present an algorithm for which the highest performing model from the model class is guaranteed to be found given unlimited data and computation. Empirically, we demonstrate that our algorithm is often superior to the maximum likelihood learner in a batch learning setting for two common RL benchmark problems and a third real-world system, the hydrodynamic cart-pole, a domain whose complex dynamics cannot be known exactly.
Joshua Mason Joseph, Alborz Geramifard, John W. Roberts, Jonathan P. How, Nicholas Roy
ICRA2
2013 Batch-iFDD for Representation Expansion in Large MDPs
Alborz Geramifard, Thomas J. Walsh 0001, Nicholas Roy, Jonathan P. How
UAI1
2012 Adaptive Planning for Markov Decision Processes with Uncertain Transition Models via Incremental Feature Dependency Discovery
N. Kemal Ure, Alborz Geramifard, Girish Chowdhary 0001, Jonathan P. How
ECML/PKDD (2)2
2011 Online Discovery of Feature Dependencies
Alborz Geramifard, Finale Doshi-Velez, Joshua D. Redding, Nicholas Roy, Jonathan P. How
ICML1
2008 Dyna-Style Planning with Linear Function Approximation and Prioritized Sweeping
Richard S. Sutton, Csaba Szepesvári, Alborz Geramifard, Michael H. Bowling
UAI3
2006 Incremental Least-Squares Temporal Difference Learning
Alborz Geramifard, Michael H. Bowling, Richard S. Sutton
AAAI1
2006 iLSTD: Eligibility Traces and Convergence Analysis
abstract
We present new theoretical and empirical results with the iLSTD algorithm for policy evaluation in reinforcement learning with linear function approximation. iLSTD is an incremental method for achieving results similar to LSTD, the dataefficient, least-squares version of temporal difference learning, without incurring the full cost of the LSTD computation. LSTD is O(n2 ), where n is the number of parameters in the linear function approximator, while iLSTD is O(n). In this paper, we generalize the previous iLSTD algorithm and present three new results: (1) the first convergence proof for an iLSTD algorithm; (2) an extension to incorporate eligibility traces without changing the asymptotic computational complexity; and (3) the first empirical results with an iLSTD algorithm for a problem (mountain car) with feature vectors large enough (n = 10, 000) to show substantial computational advantages over LSTD.
Alborz Geramifard, Michael H. Bowling, Martin Zinkevich, Richard S. Sutton
NIPS1