Charles L. Isbell Jr.

dblp:i/LCIsbell · also Charles Isbell, Charles Lee Isbell Jr. · DBLP profile ↗
← Back
57ranked-venue papers
5as first author
0since 2021 · last 2020
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 46 · 4 first-authorGraphics, computer vision, multimedia, augmented reality and games · 18Human-computer interaction and ubiquitous computing · 7 · 1 first-authorSystems, architecture and hardware · 6Applied, interdisciplinary, general and emerging computing · 4Software engineering, systems software and programming languages · 1Databases, data management, data science and information retrieval · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
29 papers
Reinforcement learning · 68% Motion planning and robot control · 10% Planning, search and constraint satisfaction · 5%
Theoretical computer science
6 papers
Algorithmic game theory and mechanism design · 50% Algorithms and data structures · 23% Mathematical optimization · 16%
Databases, data mining, and information retrieval
5 papers
Data mining · 74% Recommender systems · 24% Machine learning and data management · 1%

Topics — the 30 heaviest of 87, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning
imitation learning
0.722019
Imitating Latent Policies from Observation · ICML 2019
State Aware Imitation Learning · NIPS 2017
Machine learning › Reinforcement learning
model-based reinforcement learning
0.622020
Estimating Q(s,s') with Deep Deterministic Dynamics Gradients · ICML 2020
A Physics-Based Model Prior for Object-Oriented MDPs · ICML 2014
Machine learning › Reinforcement learning › hierarchical reinforcement learning
modular reinforcement learning
0.422019
Composable Modular Reinforcement Learning · AAAI 2019
On the Difficulty of Modular Reinforcement Learning for Real-World Partial Programming · AAAI 2006
Robotics › Motion planning and robot control
dynamic modeling
0.412020
Estimating Q(s,s') with Deep Deterministic Dynamics Gradients · ICML 2020
Machine learning › Reinforcement learning
value-based reinforcement learning
0.412020
Estimating Q(s,s') with Deep Deterministic Dynamics Gradients · ICML 2020
Machine learning › Reinforcement learning
value function estimation
0.412020
Estimating Q(s,s') with Deep Deterministic Dynamics Gradients · ICML 2020
Machine learning › Reinforcement learning › human-in-the-loop reinforcement learning
interactive reinforcement learning
0.422015
Policy Shaping with Human Teachers · IJCAI 2015
Policy Shaping: Integrating Human Feedback with Reinforcement Learning · NIPS 2013
Machine learning › Reinforcement learning
multi-objective reinforcement learning
0.412019
Composable Modular Reinforcement Learning · AAAI 2019
Algorithmic game theory and mechanism design
stochastic games
0.322012
Computing Optimal Strategies to Commit to in Stochastic Games · AAAI 2012
Quick Polytope Approximation of All Correlated Equilibria in Stochastic Games · AAAI 2011
Robotics › Motion planning and robot control
robot learning
0.322015
A Physics-Based Model Prior for Object-Oriented MDPs · ICML 2014
Learning non-holonomic object models for mobile manipulation · ICRA 2015
Machine learning › Reinforcement learning › multi-agent reinforcement learning
markov games
0.222012
Computing Optimal Strategies to Commit to in Stochastic Games · AAAI 2012
Solving Stochastic Games · NIPS 2009
Robotics › Robot manipulation
mobile manipulation
0.212015
Learning non-holonomic object models for mobile manipulation · ICRA 2015
Machine learning › Reinforcement learning
efficient reinforcement learning
0.212014
Abstraction from demonstration for efficient reinforcement learning in high-dimensional domains · Artif. Intell. 2014
Machine learning › Reinforcement learning › markov decision process › structured markov decision process
object-oriented MDPs
0.212014
A Physics-Based Model Prior for Object-Oriented MDPs · ICML 2014
Machine learning › Reinforcement learning › policy search
bayesian policy search
0.212013
Policy Shaping: Integrating Human Feedback with Reinforcement Learning · NIPS 2013
Machine learning › Reinforcement learning › multi-agent reinforcement learning › markov games
decentralized partially observable markov decision process
0.212013
Point Based Value Iteration with Optimal Belief Compression for Dec-POMDPs · NIPS 2013
Machine learning › Reinforcement learning
human feedback
0.212013
Policy Shaping: Integrating Human Feedback with Reinforcement Learning · NIPS 2013
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › constraint optimization
integer programming
0.212013
Point Based Value Iteration with Optimal Belief Compression for Dec-POMDPs · NIPS 2013
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › planning under uncertainty
point-based value iteration
0.212013
Point Based Value Iteration with Optimal Belief Compression for Dec-POMDPs · NIPS 2013
Mathematical optimization › integer programming
branch-and-bound
0.212013
Tree-Independent Dual-Tree Algorithms · ICML (3) 2013
Algorithms and data structures › similarity search
nearest neighbor search
0.212013
Tree-Independent Dual-Tree Algorithms · ICML (3) 2013
Machine learning › Reinforcement learning › multi-agent reinforcement learning › equilibrium learning
correlated equilibrium
0.112012
Computing Optimal Strategies to Commit to in Stochastic Games · AAAI 2012
Data mining › structured data mining › graph mining
motif discovery
0.122007
Detecting Subdimensional Motifs: An Efficient Algorithm for Generalized Multivariate Pattern Discovery · ICDM 2007
Discovering Multivariate Motifs using Subsequence Density Estimation and Greedy Mixture Learning · AAAI 2007
Data mining
pattern mining
0.122007
Detecting Subdimensional Motifs: An Efficient Algorithm for Generalized Multivariate Pattern Discovery · ICDM 2007
Discovering Multivariate Motifs using Subsequence Density Estimation and Greedy Mixture Learning · AAAI 2007
Algorithmic game theory and mechanism design › stackelberg game
stackelberg strategies
0.112012
Computing Optimal Strategies to Commit to in Stochastic Games · AAAI 2012
Machine learning › Reinforcement learning
markov decision process
0.122007
Authorial Idioms for Target Distributions in TTD-MDPs · AAAI 2007
Targeting Specific Distributions of Trajectories in MDPs · AAAI 2006
Machine learning › Reinforcement learning › function approximation › representation learning for reinforcement learning
state abstraction
0.112011
Automatic State Abstraction from Demonstration · IJCAI 2011
Algorithmic game theory and mechanism design › equilibrium computation
correlated equilibrium
0.112011
Quick Polytope Approximation of All Correlated Equilibria in Stochastic Games · AAAI 2011
Computer vision › Video understanding and tracking
activity recognition
0.112009
A novel sequence representation for unsupervised analysis of human activities · Artif. Intell. 2009
Machine learning › Reinforcement learning
hierarchical reinforcement learning
0.112009
Discovering options from example trajectories · ICML 2009

Methods — techniques the papers use, named apart from their topics

reinforcement learning · 0.5off-policy learning · 0.4deep deterministic dynamics gradients · 0.4q-learning · 0.4causal effect characterization · 0.4action alignment · 0.4temporal difference learning · 0.3gradient estimation · 0.3system identification · 0.2physics-based estimation · 0.2human feedback · 0.2pruning bound · 0.2meta-algorithm · 0.2user context modeling · 0.2QPACE · 0.1polytope approximation · 0.1modified bellman equation · 0.1game-theoretic solution · 0.1
YearPublicationVenuePosition
2020 Estimating Q(s,s') with Deep Deterministic Dynamics Gradients
abstract
In this paper, we introduce a novel form of value function, $Q(s, s’)$, that expresses the utility of transitioning from a state $s$ to a neighboring state $s’$ and then acting optimally thereafter. In order to derive an optimal policy, we develop a forward dynamics model that learns to make next-state predictions that maximize this value. This formulation decouples actions from values while still learning off-policy. We highlight the benefits of this approach in terms of value function transfer, learning within redundant action spaces, and learning off-policy from state observations generated by sub-optimal or completely random policies. Code and videos are available at http://sites.google.com/view/qss-paper.
Ashley D. Edwards, Himanshu Sahni, Rosanne Liu, Jane Hung, Rui Wang 0052, Adrien Ecoffet, Thomas Miconi, Charles L. Isbell Jr., Jason Yosinski
ICML9
2020 Supportive Actions for Manipulation in Human-Robot Coworker Teams
abstract
The increasing presence of robots alongside humans, such as in human-robot teams in manufacturing, gives rise to research questions about the kind of behaviors people prefer in their robot counterparts. We term actions that support interaction by reducing future interference with others as supportive robot actions and investigate their utility in a co-located manipulation scenario. We compare two robot modes in a shared table pick-and-place task: (1) Task-oriented: the robot only takes actions to further its task objective and (2) Supportive: the robot sometimes prefers supportive actions to task-oriented ones when they reduce future goal-conflicts. Our experiments in simulation, using a simplified human model, reveal that supportive actions reduce the interference between agents, especially in more difficult tasks, but also cause the robot to take longer to complete the task. We implemented these modes on a physical robot in a user study where a human and a robot perform object placement on a shared table. Our results show that a supportive robot was perceived more favorably as a coworker and also reduced interference with the human in one of two scenarios. However, it also took longer to complete the task highlighting an interesting trade-off between task-efficiency and human-preference that needs to be considered before designing robot behavior for close-proximity manipulation scenarios.
Shray Bansal, Rhys Newbury, Wesley P. Chan, Akansel Cosgun, Aimee Allen, Dana Kulic, Tom Drummond, Charles L. Isbell Jr.
IROS8
2019 Composable Modular Reinforcement Learning
abstract
Modular reinforcement learning (MRL) decomposes a monolithic multiple-goal problem into modules that solve a portion of the original problem. The modules’ action preferences are arbitrated to determine the action taken by the agent. Truly modular reinforcement learning would support not only decomposition into modules, but composability of separately written modules in new modular reinforcement learning agents. However, the performance of MRL agents that arbitrate module preferences using additive reward schemes degrades when the modules have incomparable reward scales. This performance degradation means that separately written modules cannot be composed in new modular reinforcement learning agents as-is – they may need to be modified to align their reward scales. We solve this problem with a Q-learningbased command arbitration algorithm and demonstrate that it does not exhibit the same performance degradation as existing approaches to MRL, thereby supporting composability.
Christopher L. Simpkins, Charles L. Isbell Jr.
AAAI2
2019 Imitating Latent Policies from Observation
abstract
In this paper, we describe a novel approach to imitation learning that infers latent policies directly from state observations. We introduce a method that characterizes the causal effects of latent actions on observations while simultaneously predicting their likelihood. We then outline an action alignment procedure that leverages a small amount of environment interactions to determine a mapping between the latent and real-world actions. We show that this corrected labeling can be used for imitating the observed behavior, even though no expert actions are given. We evaluate our approach within classic control environments and a platform game and demonstrate that it performs better than standard approaches. Code for this work is available at https://github.com/ashedwards/ILPO.
Ashley D. Edwards, Himanshu Sahni, Yannick Schroecker, Charles L. Isbell Jr.
ICML4
2019 Master's at Scale: Five Years in a Scalable Online Graduate Degree
abstract
In 2014, Georgia Tech launched the first for-credit MOOC-based graduate degree program. In the five years since, the program has proven generally successful, enrolling over 14,000 unique students, and several other similar programs have followed in its footsteps. Existing research on the program has focused largely on details of individual classes; program-level research, however, has been scarce. In this paper, we delve into the program-level details of an at-scale Master's degree, from the story of its creation through the data generated by the program, including the numbers of applications, admissions, matriculations, and graduations; enrollment details including demographic information and retention patterns; trends in student grades and experience as compared to the on-campus student body; and alumni perceptions. Among our findings, we note that the program has stabilized at a retention rate of around 70%; that the program's growth has not slowed; that the program has not cannibalized its on-campus counterpart; and that the program has seen an upward trend in the number of women enrolled as well as a persistently higher number of underrepresented minorities than the on-campus program. Throughout this analysis, we abstract out distinct lessons that should inform the development and growth of similar programs.
David A. Joyner, Charles L. Isbell Jr.
L@S2
2018 Rising CS Enrollments: Meeting the Challenges
abstract
No abstract available.
Eric Roberts 0001, Tracy Camp, David E. Culler, Charles L. Isbell Jr., Jodi L. Tims
SIGCSE4
2017 State Aware Imitation Learning
abstract
Imitation learning is the study of learning how to act given a set of demonstrations provided by a human expert. It is intuitively apparent that learning to take optimal actions is a simpler undertaking in situations that are similar to the ones shown by the teacher. However, imitation learning approaches do not tend to use this insight directly. In this paper, we introduce State Aware Imitation Learning (SAIL), an imitation learning algorithm that allows an agent to learn how to remain in states where it can confidently take the correct action and how to recover if it is lead astray. Key to this algorithm is a gradient learned using a temporal difference update rule which leads the agent to prefer states similar to the demonstrated states. We show that estimating a linear approximation of this gradient yields similar theoretical guarantees to online temporal difference learning approaches and empirically show that SAIL can effectively be used for imitation learning in continuous domains with non-linear function approximators used for both the policy representation and the gradient estimate.
Yannick Schroecker, Charles L. Isbell Jr.
NIPS2
2016 Navigation Among Movable Obstacles with learned dynamic constraints
abstract
In this paper we present the first planner for the problem of Navigation Among Movable Obstacles (NAMO) on a real robot that can handle environments with under-specified object dynamics. This result makes use of recent progress from two threads of the Reinforcement Learning literature. The first is a hierarchical Markov-Decision Process formulation of the NAMO problem designed to handle dynamics uncertainty. The second is a physics-based Reinforcement Learning framework which offers a way to ground this uncertainty in a compact model space that can be efficiently updated from data received by the robot online. Our results demonstrate the ability of a robot to adapt to unexpected object behavior in a real office scenario.
Jonathan Scholz, Nehchal Jindal, Martin Levihn, Charles L. Isbell Jr., Henrik I. Christensen
IROS4
2016 The Unexpected Pedagogical Benefits of Making Higher Education Accessible
abstract
Many ongoing efforts in online education aim to increase accessibility through affordability and flexibility, but some critics have noted that pedagogy often suffers during these efforts. In contrast, in the low-cost for-credit Georgia Tech Online Masters of Science in Computer Science (OMSCS) program, we have observed that the features that make the program accessible also lead to pedagogical benefits. In this paper, we discuss the pedagogical benefits, and draw a causal link between those benefits and the factors that increase the program's accessibility.
David A. Joyner, Ashok K. Goel 0001, Charles L. Isbell Jr.
L@S3
2016 Peer Reviewing Short Answers using Comparative Judgement
abstract
We propose a comparative judgement scheme for grading short answer questions in an online class. The scheme works by asking students to answer short answer questions. Then a multiple choice question is created whose choices are the answers given by students. We show that we can formulate a probabilistic graphical model for this scheme which lets us infer each students proficiency for answering and grading questions.
Pushkar Kolhe, Michael L. Littman, Charles L. Isbell Jr.
L@S3
2015 Learning non-holonomic object models for mobile manipulation
abstract
For a mobile manipulator to interact with large everyday objects, such as office tables, it is often important to have dynamic models of these objects. However, as it is infeasible to provide the robot with models for every possible object it may encounter, it is desirable that the robot can identify common object models autonomously. Existing methods for addressing this challenge are limited by being either purely kinematic, or inefficient due to a lack of physical structure. In this paper, we present a physics-based method for estimating the dynamics of common non-holonomic objects using a mobile manipulator, and demonstrate its efficiency compared to existing approaches.
Jonathan Scholz, Martin Levihn, Charles L. Isbell Jr., Henrik I. Christensen, Mike Stilman
ICRA3
2015 Policy Shaping with Human Teachers
Thomas Cederborg, Ishaan Grover, Charles L. Isbell Jr., Andrea Thomaz
IJCAI3
2014 A Physics-Based Model Prior for Object-Oriented MDPs
abstract
One of the key challenges in using reinforcement learning in robotics is the need for models that capture natural world structure. There are, methods that formalize multi-object dynamics using relational representations, but these methods are not sufficiently compact for real-world robotics. We present a physics-based approach that exploits modern simulation tools to efficiently parameterize physical dynamics. Our results show that this representation can result in much faster learning, by virtue of its strong but appropriate inductive bias in physical environments.
Jonathan Scholz, Martin Levihn, Charles L. Isbell Jr., David Wingate
ICML3
2014 Abstraction from demonstration for efficient reinforcement learning in high-dimensional domains
Luis C. Cobo, Kaushik Subramanian, Charles L. Isbell Jr., Aaron D. Lanterman, Andrea Thomaz
Artif. Intell.3
2014 Lessons on Using Computationally Generated Influence for Shaping Narrative Experiences
abstract
In this paper, we present computational models for generating influence that allow story managers to shape players' decisions in interactive narrative experiences. Our approach uses concepts from social psychology, discourse analysis, and natural language generation. We describe an abstract formalism to operationalize tools of social psychological influence described by Cialdini (Influence: The Psychology of Persuasion, New York, NY, USA: Harper-Collins, 1998) and evaluate two example implementations that enable a storytelling system to generate influence on the fly (with varying degrees of success), thereby adapting stories to realize goals specified by authors. These implementations are used in an interactive story where influence is generated dynamically as players' experiences unfold. We present the results of a user study to characterize the effectiveness of these models. Results did not indicate the presence of any significant differences in players' sense of control over the story with, or without, the use of influence. Further, the use of influence resulted in a set of stories experienced by players that more closely matched the author's goals.
David L. Roberts 0001, Charles L. Isbell Jr.
IEEE Trans. Comput. Intell. AI Games2
2013 Tree-Independent Dual-Tree Algorithms
abstract
Dual-tree algorithms are a widely used class of branch-and-bound algorithms. Unfortunately, developing dual-tree algorithms for use with different trees and problems is often complex and burdensome. We introduce a four-part logical split: the tree, the traversal, the point-to-point base case, and the pruning rule. We provide a meta-algorithm which allows development of dual-tree algorithms in a tree-independent manner and easy extension to entirely new types of trees. Representations are provided for five common algorithms; for k-nearest neighbor search, this leads to a novel, tighter pruning bound. The meta-algorithm also allows straightforward extensions to massively parallel settings.
Ryan R. Curtin, William B. March, Parikshit Ram, David V. Anderson, Alexander G. Gray, Charles L. Isbell Jr.
ICML (3)6
2013 Policy Shaping: Integrating Human Feedback with Reinforcement Learning
abstract
A long term goal of Interactive Reinforcement Learning is to incorporate non-expert human feedback to solve complex tasks. State-of-the-art methods have approached this problem by mapping human information to reward and value signals to indicate preferences and then iterating over them to compute the necessary control policy. In this paper we argue for an alternate, more effective characterization of human feedback: Policy Shaping. We introduce Advise, a Bayesian approach that attempts to maximize the information gained from human feedback by utilizing it as direct labels on the policy. We compare Advise to state-of-the-art approaches and highlight scenarios where it outperforms them and importantly is robust to infrequent and inconsistent human feedback.
Shane Griffith, Kaushik Subramanian, Jonathan Scholz, Charles L. Isbell Jr., Andrea Thomaz
NIPS4
2013 Point Based Value Iteration with Optimal Belief Compression for Dec-POMDPs
abstract
This paper presents four major results towards solving decentralized partially observable Markov decision problems (DecPOMDPs) culminating in an algorithm that outperforms all existing algorithms on all but one standard infinite-horizon benchmark problems. (1) We give an integer program that solves collaborative Bayesian games (CBGs). The program is notable because its linear relaxation is very often integral. (2) We show that a DecPOMDP with bounded belief can be converted to a POMDP (albeit with actions exponential in the number of beliefs). These actions correspond to strategies of a CBG. (3) We present a method to transform any DecPOMDP into a DecPOMDP with bounded beliefs (the number of beliefs is a free parameter) using optimal (not lossless) belief compression. (4) We show that the combination of these results opens the door for new classes of DecPOMDP algorithms based on previous POMDP algorithms. We choose one such algorithm, point-based valued iteration, and modify it to produce the first tractable value iteration method for DecPOMDPs which outperforms existing algorithms.
Liam MacDermed, Charles L. Isbell Jr.
NIPS2
2012 Computing Optimal Strategies to Commit to in Stochastic Games
abstract
Significant progress has been made recently in the following two lines of research in the intersection of AI and game theory: (1) the computation of optimal strategies to commit to (Stackelberg strategies), and (2) the computation of correlated equilibria of stochastic games. In this paper, we unite these two lines of research by studying the computation of Stackelberg strategies in stochastic games. We provide theoretical results on the value of being able to commit and the value of being able to correlate, as well as complexity results about computing Stackelberg strategies in stochastic games. We then modify the QPACE algorithm (MacDermed et al. 2011) to compute Stackelberg strategies, and provide experimental results.
Joshua Letchford, Liam MacDermed, Vincent Conitzer, Ronald Parr, Charles L. Isbell Jr.
AAAI5
2011 Quick Polytope Approximation of All Correlated Equilibria in Stochastic Games
abstract
Stochastic or Markov games serve as reasonable models for a variety of domains from biology to computer security, and are appealing due to their versatility. In this paper we address the problem of finding the complete set of correlated equilibria for general-sum stochastic games with perfect information. We present QPACE — an algorithm orders of magnitude more efficient than previous approaches while maintaining a guarantee of convergence and bounded error. Finally, we validate our claims and demonstrate the limits of our algorithm with extensive empirical tests.
Liam MacDermed, Karthik S. Narayan, Charles L. Isbell Jr., Lora Weiss
AAAI3
2011 Automatic State Abstraction from Demonstration
Luis C. Cobo, Peng Zang, Charles L. Isbell Jr., Andrea Thomaz
IJCAI3
2009 Beyond Adversarial: The Case for Game AI as Storytelling
David L. Roberts 0001, Mark O. Riedl, Charles L. Isbell Jr.
DiGRA Conference3
2009 Discovering options from example trajectories
abstract
We present a novel technique for automated problem decomposition to address the problem of scalability in reinforcement learning. Our technique makes use of a set of near-optimal trajectories to discover options and incorporates them into the learning process, dramatically reducing the time it takes to solve the underlying problem. We run a series of experiments in two different domains and show that our method offers up to 30 fold speedup over the baseline.
Peng Zang, David Minnen, Charles L. Isbell Jr.
ICML4
2009 Solving Stochastic Games
abstract
Solving multi-agent reinforcement learning problems has proven difficult because of the lack of tractable algorithms. We provide the first approximation algorithm which solves stochastic games to within $\epsilon$ relative error of the optimal game-theoretic solution, in time polynomial in $1/\epsilon$. Our algorithm extends Murrays and Gordon’s (2007) modified Bellman equation which determines the \emph{set} of all possible achievable utilities; this provides us a truly general framework for multi-agent learning. Further, we empirically validate our algorithm and find the computational cost to be orders of magnitude less than what the theory predicts.
Liam MacDermed, Charles L. Isbell Jr.
NIPS2
2009 A novel sequence representation for unsupervised analysis of human activities
Raffay Hamid, Siddhartha Maddi, Amos Y. Johnson, Aaron F. Bobick, Irfan A. Essa, Charles L. Isbell Jr.
Artif. Intell.6
2008 On the Use of Computational Models of Influence for Managing Interactive Virtual Experiences
David L. Roberts 0001, Charles L. Isbell Jr., Mark O. Riedl, Ian Bogost, Merrick L. Furst
ICIDS2
2008 Impact of user context on song selection
abstract
The rise of digital music has led to a parallel rise in the need to manage music collections of several thousands of songs on a single device. Manual selection of songs for a music listening experience is a cumbersome task. In this paper, we present an initial exploration of the feasibility of using song signal properties and user context information to assist in automatic song selection. Users listened to music over the course of a month while their context and song selections were tracked. Initial results suggest the use of context information can improve automated song selection when patterns are learned for each individual.
Olufisayo Omojokun, Michael Genovese, Charles L. Isbell Jr.
ACM Multimedia3
2008 Partial signal extraction for mobile media players
abstract
Audio signal properties can provide a media player with highly descriptive feature sets in order to intelligently select similar songs for a music stream. A well-known problem among researchers in music information retrieval, however, is that extracting signal properties requires a significant amount of computational resources, thus making it impractical for even the most advanced mobile media players. Although other approaches to retrieving data are possible, local extraction still has unique benefits. Using a combination of machine learning and profiling techniques, this paper presents an initial evaluation of partial signal extraction, which reduces resource requirements by locally collecting signals from parts of a song rather than all. Our preliminary experiments suggest that this idea can offer significantly lower resource requirements while losing marginal song information.
Olufisayo Omojokun, Michael Genovese, Charles L. Isbell Jr.
MoMM3
2008 QUIC-SVD: Fast SVD Using Cosine Trees
abstract
The Singular Value Decomposition is a key operation in many machine learning methods. Its computational cost, however, makes it unscalable and impractical for the massive-sized datasets becoming common in applications. We present a new method, QUIC-SVD, for fast approximation of the full SVD with automatic sample size minimization and empirical relative error control. Previous Monte Carlo approaches have not addressed the full SVD nor benefited from the efficiency of automatic, empirically-driven sample sizing. Our empirical tests show speedups of several orders of magnitude over exact SVD. Such scalability should enable QUIC-SVD to meet the needs of a wide array of methods and applications.
Michael P. Holmes, Alexander G. Gray, Charles L. Isbell Jr.
NIPS3
2008 Towards adaptive programming: integrating reinforcement learning into a programming language
abstract
Current programming languages and software engineering paradigms are proving insufficient for building intelligent multi-agent systems--such as interactive games and narratives--where developers are called upon to write increasingly complex behavior for agents in dynamic environments. A promising solution is to build adaptive systems; that is, to develop software written specifically to adapt to its environment by changing its behavior in response to what it observes in the world. In this paper we describe a new programming language, An Adaptive Behavior Language (A2BL), that implements adaptive programming primitives to support partial programming, a paradigm in which a programmer need only specify the details of behavior known at code-writing time, leaving the run-time system to learn the rest. Partial programming enables programmers to more easily encode software agents that are difficult to write in existing languages that do not offer language-level support for adaptivity. We motivate the use of partial programming with an example agent coded in a cutting-edge, but non-adaptive agent programming language (ABL), and show how A2BL can encode the same agent much more naturally.
Christopher L. Simpkins, Sooraj Bhat, Charles L. Isbell Jr., Michael Mateas
OOPSLA3
2007 Discovering Multivariate Motifs using Subsequence Density Estimation and Greedy Mixture Learning
David Minnen, Charles L. Isbell Jr., Irfan A. Essa, Thad Starner
AAAI2
2007 Authorial Idioms for Target Distributions in TTD-MDPs
David L. Roberts 0001, Sooraj Bhat, Kenneth St. Clair, Charles L. Isbell Jr.
AAAI4
2007 Detecting Subdimensional Motifs: An Efficient Algorithm for Generalized Multivariate Pattern Discovery
abstract
Discovering recurring patterns in time series data is a fundamental problem for temporal data mining. This paper addresses the problem of locating subdimensional motifs in real-valued, multivariate time series, which requires the simultaneous discovery of sets of recurring patterns along with the corresponding relevant dimensions. While many approaches to motif discovery have been developed, most are restricted to categorical data, univariate time series, or multivariate data in which the temporal patterns span all of the dimensions. In this paper, we present an expected linear-time algorithm that addresses a generalization of multivariate pattern discovery in which each motif may span only a subset of the dimensions. To validate our algorithm, we discuss its theoretical properties and empirically evaluate it using several data sets including synthetic data and motion capture data collected by an on-body iner- tial sensor.
David Minnen, Charles L. Isbell Jr., Irfan A. Essa, Thad Starner
ICDM2
2007 Improving Activity Discovery with Automatic Neighborhood Estimation
David Minnen, Thad Starner, Irfan A. Essa, Charles L. Isbell Jr.
IJCAI4
2007 Transfer Learning in Real-Time Strategy Games Using Hybrid CBR/RL
Manu Sharma, Michael P. Holmes, Juan Carlos Santamaría, Arya Irani, Charles L. Isbell Jr., Ashwin Ram 0001
IJCAI5
2007 Managing Domain Knowledge and Multiple Models with Boosting
Peng Zang, Charles L. Isbell Jr.
IJCAI2
2007 Ultrafast Monte Carlo for Statistical Summations
abstract
Machine learning contains many computational bottlenecks in the form of nested summations over datasets. Kernel estimators and other methods are burdened by these expensive computations. Exact evaluation is typically O(n2 ) or higher, which severely limits application to large datasets. We present a multi-stage stratified Monte Carlo method for approximating such summations with probabilistic relative error control. The essential idea is fast approximation by sampling in trees. This method differs from many previous scalability techniques (such as standard multi-tree methods) in that its error is stochastic, but we derive conditions for error control and demonstrate that they work. Further, we give a theoretical sample complexity for the method that is independent of dataset size, and show that this appears to hold in experiments, where speedups reach as high as 1014 , many orders of magnitude beyond the previous state of the art.
Michael P. Holmes, Alexander G. Gray, Charles L. Isbell Jr.
NIPS3
2007 ThreadsTM: how to restructure a computer science curriculum for a flat world
abstract
In his book The World is Flat, Thomas Friedman convincingly explains the challenges of a global marketplace [4]. One implication is that software development can be out-sourced, as can any narrow, skills-based occupation; however, as Friedman also points out, leadership, innovation, and insight are always in demand. We have recently created and are implementing threadstm, a new structuring principle for computing curricula. Threads provides one clear path for scientists seeking to reinvent and re-invigorate science degree programs. Threads form a cohesive, coordinated set of contexts for understanding computing. The union of all threads covers the breadth computer science. The union of any two threads is sufficient to cover a science degree. In this paper, we describe Threads, our process, the impact so far, and some of our future plans. We close with recommendations for other schools, especially schools with smaller programs.
Merrick L. Furst, Charles L. Isbell Jr., Mark Guzdial
SIGCSE2
2007 Fast Nonparametric Conditional Density Estimation
Michael P. Holmes, Alexander G. Gray, Charles L. Isbell Jr.
UAI3
2006 On the Difficulty of Modular Reinforcement Learning for Real-World Partial Programming
Sooraj Bhat, Charles L. Isbell Jr., Michael Mateas
AAAI2
2006 Wavelet Statistics for Human Motion Classification
Kevin Quennesson, Elias Ioup, Charles L. Isbell Jr.
AAAI3
2006 Targeting Specific Distributions of Trajectories in MDPs
David L. Roberts 0001, Mark J. Nelson, Charles L. Isbell Jr., Michael Mateas, Michael L. Littman
AAAI3
2006 Looping suffix tree-based inference of partially observable hidden state
abstract
We present a solution for inferring hidden state from sensorimotor experience when the environment takes the form of a POMDP with deterministic transition and observation functions. Such environments can appear to be arbitrarily complex and non-deterministic on the surface, but are actually deterministic with respect to the unobserved underlying state. We show that there always exists a finite history-based representation that fully captures the unobserved world state, allowing for perfect prediction of action effects. This representation takes the form of a looping prediction suffix tree (PST). We derive a sound and complete algorithm for learning a looping PST from a sufficient sample of sensorimotor experience. We also give empirical illustrations of the advantages conferred by this approach, and characterize the approximations to the looping PST that are made by existing algorithms such as Variable Length Markov Models, Utile Suffix Memory and Causal State Splitting Reconstruction.
Michael P. Holmes, Charles L. Isbell Jr.
ICML2
2006 Cobot in LambdaMOO: An Adaptive Social Statistics Agent
Charles L. Isbell Jr., Michael Kearns, Satinder Singh 0001, Christian R. Shelton, Peter Stone 0001, David P. Kormann
Auton. Agents Multi Agent Syst.1
2006 How Multirobot Systems Research will Accelerate our Understanding of Social Animal Behavior
abstract
Our understanding of social insect behavior has significantly influenced artificial intelligence (AI) and multirobot systems' research (e.g., ant algorithms and swarm robotics). In this work, however, we focus on the opposite question: "How can multirobot systems research contribute to the understanding of social animal behavior?" As we show, we are able to contribute at several levels. First, using algorithms that originated in the robotics community, we can track animals under observation to provide essential quantitative data for animal behavior research. Second, by developing and applying algorithms originating in speech recognition and computer vision, we can automatically label the behavior of animals under observation. In some cases the automatic labeling is more accurate and consistent than manual behavior identification. Our ultimate goal, however, is to automatically create, from observation, executable models of behavior. An executable model is a control program for an agent that can run in simulation (or on a robot). The representation for these executable models is drawn from research in multirobot systems programming. In this paper we present the algorithms we have developed for tracking, recognizing, and learning models of social animal behavior, details of their implementation, and quantitative experimental results using them to study social insects
Tucker R. Balch, Frank Dellaert, Adam Feldman, Andrew Guillory, Charles L. Isbell Jr., Zia Khan, Stephen Pratt, Andrew N. Stein, Hank Wilde
Proc. IEEE5
2006 Comparing end-user and intelligent remote control interface generation
Olufisayo Omojokun, Jeffrey S. Pierce, Charles L. Isbell Jr., Prasun Dewan
Pers. Ubiquitous Comput.3
2005 Detection and Explanation of Anomalous Activities: Representing Activities as Bags of Event n-Grams
abstract
We present a novel representation and method for detecting and explaining anomalous activities in a video stream. Drawing from natural language processing, we introduce a representation of activities as bags of event n-grams, where we analyze the global structural information of activities using their local event statistics. We demonstrate how maximal cliques in an undirected edge-weighted graph of activities, can be used in an unsupervised manner, to discover regular sub-classes of an activity class. Based on these discovered sub-classes, we formulate a definition of anomalous activities and present a way to detect them. Finally, we characterize each discovered sub-class in terms of its "most representative member" and present an information-theoretic method to explain the detected anomalies in a human-interpretable form.
Raffay Hamid, Amos Y. Johnson, Samir Batta, Aaron F. Bobick, Charles L. Isbell Jr., Graham Coleman
CVPR (1)5
2005 Unsupervised Activity Discovery and Characterization From Event-Streams
Rafay Hammid, Siddhartha Maddi, Amos Y. Johnson, Aaron F. Bobick, Irfan A. Essa, Charles L. Isbell Jr.
UAI6
2004 Schema Learning: Experience-Based Construction of Predictive Action Models
abstract
Schema learning is a way to discover probabilistic, constructivist, pre- dictive action models (schemas) from experience. It includes meth- ods for finding and using hidden state to make predictions more accu- rate. We extend the original schema mechanism [1] to handle arbitrary discrete-valued sensors, improve the original learning criteria to handle POMDP domains, and better maintain hidden state by using schema pre- dictions. These extensions show large improvement over the original schema mechanism in several rewardless POMDPs, and achieve very low prediction error in a difficult speech modeling task. Further, we compare extended schema learning to the recently introduced predictive state rep- resentations [2], and find their predictions of next-step action effects to be approximately equal in accuracy. This work lays the foundation for a schema-based system of integrated learning and planning.
Michael P. Holmes, Charles L. Isbell Jr.
NIPS2
2004 From devices to tasks: automatic task prediction for personalized appliance control
Charles L. Isbell Jr., Olufisayo Omojokun, Jeffrey S. Pierce
Pers. Ubiquitous Comput.1
2004 Supporting routine decision-making with a next-generation alarm clock
Brian M. Landry, Jeffrey S. Pierce, Charles L. Isbell Jr.
Pers. Ubiquitous Comput.3
2002 Real Time Voice Processing with Audiovisual Feedback: Toward Autonomous Agents with Perfect Pitch
abstract
We have implemented a real time front end for detecting voiced speech and estimating its fundamental frequency. The front end performs the signal processing for voice-driven agents that attend to the pitch contours of human speech and provide continuous audiovisual feedback. The al- gorithm we use for pitch tracking has several distinguishing features: it makes no use of FFTs or autocorrelation at the pitch period; it updates the pitch incrementally on a sample-by-sample basis; it avoids peak picking and does not require interpolation in time or frequency to obtain high res- olution estimates; and it works reliably over a four octave range, in real time, without the need for postprocessing to produce smooth contours. The algorithm is based on two simple ideas in neural computation: the introduction of a purposeful nonlinearity, and the error signal of a least squares fit. The pitch tracker is used in two real time multimedia applica- tions: a voice-to-MIDI player that synthesizes electronic music from vo- calized melodies, and an audiovisual Karaoke machine with multimodal feedback. Both applications run on a laptop and display the user’s pitch scrolling across the screen as he or she sings into the computer.
Lawrence K. Saul, Daniel D. Lee, Charles L. Isbell Jr., Yann LeCun
NIPS3
2001 Cobot: A Social Reinforcement Learning Agent
abstract
We report on the use of reinforcement learning with Cobot, a software agent residing in the well-known online community LambdaMOO. Our initial work on Cobot (Isbell et al.2000) provided him with the ability to collect social statistics and report them to users. Here we describe an application of RL allowing Cobot to take proactive actions in this complex social environment, and adapt behavior from multiple sources of human reward. After 5 months of training, and 3171 reward and punishment events from 254 different LambdaMOO users, Cobot learned nontrivial preferences for a number of users, modifing his behavior based on his current state. Here we describe LambdaMOO and the state and action spaces of Cobot, and report the statistical results of the learning experiment.
Charles L. Isbell Jr., Christian R. Shelton, Michael Kearns, Satinder Singh 0001, Peter Stone 0001
NIPS1
1999 The Parallel Problems Server: an Interactive Tool for Large Scale Machine Learning
Charles L. Isbell Jr., Parry Husbands
NIPS1
1998 Restructuring Sparse High Dimensional Data for Effective Retrieval
Charles L. Isbell Jr., Paul A. Viola
NIPS1
1996 MIMIC: Finding Optima by Estimating Probability Densities
Jeremy S. De Bonet, Charles L. Isbell Jr., Paul A. Viola
NIPS2
1995 Description Logic in Practice: A CLASSIC Application
Deborah L. McGuinness, Lori Alperin Resnick, Charles L. Isbell Jr.
IJCAI3