Gita Reese Sukthankar

dblp:54/1919 · also Gita Sukthankar · DBLP profile ↗
← Back
46ranked-venue papers
4as first author
12since 2021 · last 2025
0000-0002-6863-6609ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 37 · 2 first-author · 12 since 2021Databases, data management, data science and information retrieval · 13 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 2 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 12 · 2 since 2021Systems, architecture and hardware · 6 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-authorComputer networks · 1
YearPublicationVenuePosition
2025 Scaling Effects on Latent Representation Edits in GPT Models (Student Abstract)
abstract
Probing classifiers are a technique for understanding and modifying the operation of neural networks in which a smaller classifier is trained to use the model's internal representation to learn a related probing task. Similar to a neural electrode array, training probing classifiers can help researchers both discern and edit the internal representation of a neural network. This paper presents an evaluation of the use of probing classifiers to modify the internal hidden state of a chess-playing transformer. We demonstrate that intervention vector scaling should follow a negative exponential according to the length of the input to ensure model outputs remain semantically valid after editing the residual stream activations.
Austin L. Davis, Gita Reese Sukthankar
AAAI2
2024 Hidden Pieces: An Analysis of Linear Probes for GPT Representation Edits
abstract
Probing classifiers are a technique for understanding and modifying the operation of neural networks in which a smaller classifier is trained to use the model's internal representation to learn a probing task. Similar to a neural electrode array, probing classifiers help both discern and edit the internal representation of a neural network. This paper evaluates the use of probing classifiers to modify the internal hidden state of a chess-playing transformer. The weights of the learned linear classifiers are very informative and can be used to reliably delete pieces from the board showing that the model internally maintains an editable emergent representation of game state.
Austin L. Davis, Gita Reese Sukthankar
ICMLA2
2024 Multimodal Fusion Networks for Workload Modeling
abstract
The advent of low cost sensors for measuring gaze, heart rate, EEG, and galvanic skin response have made it feasible to cheaply collect physiological data from human operators. However, leveraging this data for machine learning problems requires a good multimodal fusion architecture. When dealing with multimodal features, uncovering the correlations between different modalities is as crucial as identifying effective unimodal features. This paper proposes a hybrid multimodal tensor fusion network that is effective at learning both unimodal and bimodal dynamics for cognitive workload modeling. Our architecture comprises two parts: (1) intra-modality for learning high-level representations of each signal modality (2) inter-modality for modeling bimodal interactions using a tensor fusion layer created from the Cartesian product of modality embeddings. We compare this architecture to the usage of a cross-modal transformer fusion module that learns an inter-modality embedding. Experimental results conducted on the HP Omnicept Cognitive Load Database (HPO-CLD) show that both techniques outperform the most commonly used techniques used for multimodal fusion of physio-logical data and that the cross-modal transformer fusion module is especially effective.
Shengnan Hu, Gita Reese Sukthankar
ICMLA2
2023 How Popularity Shapes User Interactions in Tech-Related Online Communities
abstract
Tech-related online communities on GitHub, Reddit, and Stack Overflow are an invaluable resource for software engineers, allowing them to find solutions to problems and connect with other professionals. Much of the discourse on these platforms is conducted using commenting mechanisms in which one user responds to content posted by another user. Even though these communities lack formal organizational structures, these technologists are often followed by other software developers who monitor their posts; users who regularly post useful solutions are recognized using platform-specific mechanisms such as stars or karma points. This paper investigates the relationship between popularity and discourse in tech-related online communities. To do this, we create comment timelines from sequences of user interactions and extract commenting networks from comment response patterns. Although there are some commonalities, there are distinct differences between the commenting behavior of GitHub users vs. Reddit and Stack Overflow. By understanding how popularity affects user interactions, we can design communities that are more effective at supporting learning, collaboration, and knowledge sharing.
Abduljaleel Al-Rubaye, Gita Reese Sukthankar
ASONAM2
2023 Improving the Generalizability of Collaborative Dialogue Analysis With Multi-Feature Embeddings
abstract
Conflict prediction in communication is integral to the design of virtual agents that support successful teamwork by providing timely assistance.The aim of our research is to analyze discourse to predict collaboration success.Unfortunately, resource scarcity is a problem that teamwork researchers commonly face since it is hard to gather a large number of training examples.To alleviate this problem, this paper introduces a multi-feature embedding (MFeEmb) that improves the generalizability of conflict prediction models trained on dialogue sequences.MFeEmb leverages textual, structural, and semantic information from the dialogues by incorporating lexical, dialogue acts, and sentiment features.The use of dialogue acts and sentiment features reduces performance loss from natural distribution shifts caused mainly by changes in vocabulary.This paper demonstrates the performance of MFeEmb on domain adaptation problems in which the model is trained on discourse from one task domain and applied to predict team performance in a different domain.The generalizability of MFeEmb is quantified using the similarity measure proposed by Bontonou et al. (2021).Our results show that MFeEmb serves as an excellent domain-agnostic representation for meta-pretraining a few-shot model on collaborative multiparty dialogues.
Ayesha Enayet, Gita Reese Sukthankar
EACL2
2023 The Potential of Vision-Language Models for Content Moderation of Children's Videos
abstract
Natural language supervision has been shown to be effective for zero-shot learning in many computer vision tasks, such as object detection and activity recognition. However, generating informative prompts can be challenging for more subtle tasks, such as video content moderation. This can be difficult, as there are many reasons why a video might be inappropriate, beyond violence and obscenity. For example, scammers may attempt to create junk content that is similar to popular educational videos but with no meaningful information. This paper evaluates the performance of several CLIP variations for content moderation of children's cartoons in both the supervised and zero-shot setting. We show that our proposed model (Vanilla CLIP with Projection Layer) outperforms previous work conducted on the Malicious or Benign (MOB) benchmark for video content moderation. This paper presents an in depth analysis of how context-specific language prompts affect content moderation performance. Our results indicate that it is important to include more context in content moderation prompts, particularly for cartoon videos as they are not well represented in the CLIP training data.
Syed Hammad Ahmed, Shengnan Hu, Gita Reese Sukthankar
ICMLA3
2023 LAMP: Leveraging Language Prompts for Multi-Person Pose Estimation
abstract
Human-centric visual understanding is an important desideratum for effective human-robot interaction. In order to navigate crowded public places, social robots must be able to interpret the activity of the surrounding humans. This paper addresses one key aspect of human-centric visual understanding, multi-person pose estimation. Achieving good performance on multi-person pose estimation in crowded scenes is difficult due to the challenges of occluded joints and instance separation. In order to tackle these challenges and overcome the limitations of image features in representing invisible body parts, we propose a novel prompt-based pose inference strategy called LAMP (Language Assisted Multi-person Pose estimation). By utilizing the text representations generated by a well-trained language model (CLIP), LAMP can facilitate the understanding of poses on the instance and joint levels, and learn more robust visual representations that are less susceptible to occlusion. This paper demonstrates that language-supervised training boosts the performance of single-stage multi-person pose estimation, and both instance-level and joint-level prompts are valuable for training. The code is available at https://github.com/shengnanh20/LAMP.
Shengnan Hu, Chen Chen 0001, Gita Reese Sukthankar
IROS5
2022 Improving Code Review with GitHub Issue Tracking
abstract
Software quality is an important problem for technology companies, since it substantially impacts the efficiency, usefulness, and maintainability of the final product; hence, code review is a must-do activity for software developers. During the code review process, senior engineers monitor other developers' work to spot possible problems and enforce coding standards. One of the most widely used open-source software platforms, GitHub, attracts millions of developers who use it to store their projects. This study aims to analyze code quality on GitHub from the standpoint of code reviews. We examined the code review process using GitHub's Issues Tracker, which allows team members to evaluate, discuss, and share their opinions on the proposed code before it is approved. Based on our analysis, we present a novel approach for improving the code review process by promoting regularity and community involvement.
Abduljaleel Al-Rubaye, Gita Reese Sukthankar
ASONAM2
2022 Predicting Team Performance with Spatial Temporal Graph Convolutional Networks
abstract
This paper presents a new approach for predicting team performance from the behavioral traces of a set of agents. This spatiotemporal forecasting problem is very relevant to sports analytics challenges such as coaching and opponent modeling. We demonstrate that our proposed model, Spatial Temporal Graph Convolutional Networks (ST-GCN), outperforms other classification techniques at predicting game score from a short segment of player movement and game features. Our proposed architecture uses a graph convolutional network to capture the spatial relationships between team members and Gated Recurrent Units to analyze dynamic motion information. An ablative evaluation was performed to demonstrate the contributions of different aspects of our architecture.
Shengnan Hu, Gita Reese Sukthankar
ICPR2
2022 An Analysis of Dialogue Act Sequence Similarity Across Multiple Domains
abstract
This paper presents an analysis of how dialogue act sequences vary across different datasets in order to anticipate the potential degradation in the performance of learned models during domain adaptation. We hypothesize the following: 1) dialogue sequences from related domains will exhibit similar n-gram frequency distributions 2) this similarity can be expressed by measuring the average Hamming distance between subsequences drawn from different datasets. Our experiments confirm that when dialogue acts sequences from two datasets are dissimilar they lie further away in embedding space, making it possible to train a classifier to discriminate between them even when the datasets are corrupted with noise. We present results from eight different datasets: SwDA, AMI (DialSum), GitHub, Hate Speech, Teams, Diplomacy Betrayal, SAMsum, and Military (Army). Our datasets were collected from many types of human communication including strategic planning, informal discussion, and social media exchanges. Our methodology provides intuition on the generalizability of dialogue models trained on different datasets. Based on our analysis, it is problematic to assume that machine learning models trained on one type of discourse will generalize well to other settings, due to contextual differences.
Ayesha Enayet, Gita Reese Sukthankar
LREC2
2021 Leveraging Transformers for StarCraft Macromanagement Prediction
abstract
Inspired by the recent success of transformers in natural language processing and computer vision applications, we introduce a transformer-based neural architecture for two key StarCraft II (SC2) macromanagement tasks: global state and build order prediction. Unlike recurrent neural networks which suffer from a recency bias, transformers are able to capture patterns across very long time horizons, making them well suited for full game analysis. Our model utilizes the MSC (Macromanagement in StarCraft II) dataset and improves on the top performing gated recurrent unit (GRU) architecture in predicting global state and build order as measured by mean accuracy over multiple time horizons. We present ablation studies on our proposed architecture that support our design decisions.One key advantage of transformers is their ability to generalize well, and we demonstrate that our model achieves an even better accuracy when used in a transfer learning setting in which models trained on games with one racial matchup (e.g., Terran vs. Protoss) are transferred to a different one. We believe that transformers’ ability to model long games, potential for parallelization, and generalization performance make them an excellent choice for StarCraft agents.
Muhammad Junaid Khan, Shah Hassan, Gita Reese Sukthankar
ICMLA3
2021 Estimating the Variance of Return Sequences for Exploration
abstract
This paper introduces a method for estimating an upper bound for an exploration policy using either the weighted variance of return sequences or the weighted temporal difference (TD) error. We demonstrate that the variance of the return sequence for a specific state-action pair is an important information source that can be leveraged to guide exploration in reinforcement learning. The intuition is that fluctuation in the return sequence indicates greater uncertainty in the near future returns. This divergence occurs because of the cyclic nature of value-based reinforcement learning; improved estimates of the value function result in policy changes which in turn modify the value function. Although both variance and TD errors capture different aspects of this uncertainty, our analysis shows that both can be valuable to guide exploration. We propose a two-stream network architecture to estimate weighted variance/TD errors within DQN agents for our exploration method and show that it outperforms the baseline on a wide range of Atari games.
Zerong Xi, Gita Reese Sukthankar
ICMLA2
2020 SelfieDroneStick: A Natural Interface for Quadcopter Photography
abstract
A physical selfie stick extends the user's reach, enabling the acquisition of personal photos that include more of the background scene. Similarly, a quadcopter can capture photos from vantage points unattainable by the user; but teleoperating a quadcopter to good viewpoints is a difficult task. This paper presents a natural interface for quadcopter photography, the SelfieDroneStick that allows the user to guide the quadcopter to the optimal vantage point based on the phone's sensors. Users specify the composition of their desired long-range selfies using their smartphone, and the quadcopter autonomously flies to a sequence of vantage points from where the desired shots can be taken. The robot controller is trained from a combination of real-world images and simulated flight data. This paper describes two key innovations required to deploy deep reinforcement learning models on a real robot: 1) an abstract state representation for transferring learning from simulation to the hardware platform, and 2) reward shaping and staging paradigms for training the controller. Both of these improvements were found to be essential in learning a robot controller from simulation that transfers successfully to the real robot.
Saif Alabachi, Gita Reese Sukthankar, Rahul Sukthankar
IROS2
2019 The Benefits of Immersive Demonstrations for Teaching Robots
abstract
One of the advantages of teaching robots by demonstration is that it can be more intuitive for users to demonstrate rather than describe the desired robot behavior. However, when the human demonstrates the task through an interface, the training data may inadvertently acquire artifacts unique to the interface, not the desired execution of the task. Being able to use one's own body usually leads to more natural demonstrations, but those examples can be more difficult to translate to robot control policies. This paper quantifies the benefits of using a virtual reality system that allows human demonstrators to use their own body to perform complex manipulation tasks. We show that our system generates superior demonstrations for a deep neural network without introducing a correspondence problem. The effectiveness of this approach is validated by comparing the learned policy to that of a policy learned from data collected via a conventional gaming system, where the user views the environment on a monitor screen, using a Sony Play Station 3 (PS3) DualShock 3 wireless controller as input.
Astrid Jackson, Brandon D. Northcutt, Gita Reese Sukthankar
HRI3
2019 Customizing Object Detectors for Indoor Robots
abstract
Object detection models based on convolutional neural networks (CNNs) demonstrate impressive performance when trained on large-scale labeled datasets. While a generic object detector trained on such a dataset performs adequately in applications where the input data is similar to user photographs, the detector performs poorly on small objects, particularly ones with limited training data or imaged from uncommon viewpoints. Also, a specific room will have many objects that are missed by standard object detectors, frustrating a robot that continually operates in the same indoor environment.This paper describes a system for rapidly creating customized object detectors. Data is collected from a quadcopter that is teleoperated with an interactive interface. Once an object is selected, the quadcopter autonomously photographs the object from multiple viewpoints to collect data to train DUNet (Dense Upscaled Network), our proposed model for learning customized object detectors from scratch given limited data. Our experiments compare the performance of learning models from scratch with DUNet vs. fine tuning existing state of the art object detectors, both on our indoor robotics domain and on standard datasets.
Saif Alabachi, Gita Reese Sukthankar, Rahul Sukthankar
ICRA2
2018 Joint Value of Information and Energy Aware Sleep Scheduling in Wireless Sensor Networks: A Linear Programming Approach
abstract
We consider wireless sensor networks that nodes offload data to a central collector node (sink) via wireless communication. Sensed data are associated with a value, decaying in time. In this scenario, we address the problem of finding the path of sensed data so that the Value of Information (VoI) of the data delivered to a sink is maximized while keeping energy usage as low as possible. Sleep scheduling is a widely used technique in MAC-layer to reduce unnecessary idle energy consumption in WSN; however, when it is carried out without paying attention to network-layer routing, it may adversely affect sensed data value of information. In this paper, we employ linear programming (LP) to establish a paradigm of cross-layer formulation to capture the interplay between scheduling and routing. We propose a biobjective model of data value of information maximization and energy cost minimization in a WSN. Compared to existing work, our formulation is not only bi-objective which considers both data value of information and energy consumption jointly, but also is more realistic given that it explicitly accounts for different types of signal interference that may affect a wireless transmission.
Neda Hajiakhoond Bidoki, Masoud Baghbahari Baghdadabad, Gita Reese Sukthankar, Damla Turgut
ICC3
2016 A holistic approach for predicting links in coevolving multiplex networks
abstract
Networks extracted from social media platforms frequently include multiple types of links that dynamically change over time; these links can be used to represent dyadic interactions such as economic transactions, communications, and shared activities. Organizing this data into a dynamic multiplex network, where each layer is composed of a single edge type linking the same underlying vertices, can reveal interesting cross-layer interaction patterns. In coevolving networks, links in one layer result in an increased probability of other types of links forming between the same node pair. Hence we believe that a holistic approach in which all the layers are simultaneously considered can outperform a factored approach in which link prediction is performed separately in each layer. This paper introduces a comprehensive framework, MLP (Multiplex Link Prediction), in which link existence likelihoods for the target layer are learned from the other network layers. These likelihoods are used to reweight the output of a single layer link prediction method that uses rank aggregation to combine a set of topological metrics. Our experiments show that our reweighting procedure outperforms other methods for fusing information across network layers.
Alireza Hajibagheri, Gita Reese Sukthankar, Kiran Lakkaraju
ASONAM2
2015 Cognitive Social Learners: An Architecture for Modeling Normative Behavior
abstract
In many cases, creating long-term solutions to sustainability issues requires not only innovative technology, but also large-scale public adoption of the proposed solutions. Social simulations are a valuable but underutilized tool that can help public policy researchers understand when sustainable practices are likely to make the delicate transition from being an individual choice to becoming a social norm. In this paper, we introduce a new normative multi-agent architecture, Cognitive Social Learners (CSL), that models bottom-up norm emergence through a social learning mechanism, while using BDI (Belief/Desire/Intention) reasoning to handle adoption and compliance. CSL preserves a greater sense of cognitive realism than influence propagation or infectious transmission approaches, enabling the modeling of complex beliefs and contradictory objectives within an agent-based simulation. In this paper, we demonstrate the use of CSL for modeling norm emergence of recycling practices and public participation in a smoke-free campus initiative.
Rahmatollah Beheshti, Awrad Mohammed Ali, Gita Reese Sukthankar
AAAI3
2014 Community detection in dynamic social networks: A game-theoretic approach
abstract
Most real-world social networks are inherently dynamic and composed of communities that are constantly changing in membership. As a result, recent years have witnessed increased attention toward the challenging problem of detecting evolving communities. This paper presents a game-theoretic approach for community detection in dynamic social networks in which each node is treated as a rational agent who periodically chooses from a set of predefined actions in order to maximize its utility function. The community structure of a snapshot emerges after the game reaches Nash equilibrium; the partitions and agent information are then transferred to the next snapshot. An evaluation of our method on two real world dynamic datasets (AS-Internet Routers Graph and AS-Oregon Graph) demonstrates that we are able to report more stable and accurate communities over time compared to the benchmark methods.
Hamidreza Alvari, Alireza Hajibagheri, Gita Reese Sukthankar
ASONAM3
2014 Know thy user: Designing human-robot interaction paradigms for multi-robot manipulation
abstract
This paper tackles the problem of designing an effective user interface for a multi-robot delivery system, composed of robots with wheeled bases and two 3 DOF arms. There are several proven paradigms for increasing the efficacy of human-robot interaction: 1) multimodal interfaces in which the user controls the robots using voice and gesture; 2) configurable interfaces which allow the user to create new commands by demonstrating them; 3) adaptive interfaces which reduce the operator's workload as necessary through increasing robot autonomy. Here we study the relative benefits of configurable vs. adaptive interfaces for multi-robot manipulation. User expertise was measured along three axes (navigation, manipulation, and coordination), and users who performed above threshold on two out of three dimensions on a calibration task were rated as expert. Our experiments reveal that the relative expertise of the user was the key determinant of the best performing interface paradigm for that user, indicating that good user modeling is essential for designing a human-robot interaction system meant to be used for an extended period of time.
Bennie Lewis, Gita Reese Sukthankar
IROS2
2013 Modeling information diffusion and community membership using stochastic optimization
abstract
Communities are vehicles for efficiently disseminating news, rumors, and opinions in human social networks. Modeling information diffusion through a network can enable us to reach a superior functional understanding of the effect of network structures such as communities on information propagation. The intrinsic assumption is that form follows function---rational actors exercise social choice mechanisms to join communities that best serve their information needs. Particle Swarm Optimization (PSO) was originally designed to simulate aggregate social behavior; our proposed diffusion model, PSODM (Particle Swarm Optimization Diffusion Model) models information flow in a network by creating particle swarms for local network neighborhoods that optimize a continuous version of Holland's hyperplane-defined objective functions. In this paper, we show how our approach differs from prior modeling work in the area and demonstrate that it outperforms existing model-based community detection methods on several social network datasets.
Alireza Hajibagheri, Ali Hamzeh, Gita Reese Sukthankar
ASONAM3
2013 Hierarchical influence maximization for advertising in multi-agent markets
abstract
Maximizing product adoption within a customer social network under a constrained advertising budget is an important special case of the general influence maximization problem. Specialized optimization techniques that account for product correlations and community effects can outperform network-based techniques that do not model interactions that arise from marketing multiple products to the same consumer base. However, it can be infeasible to use exact optimization methods that utilize expensive matrix operations on larger networks without parallel computation techniques. In this paper, we present a hierarchical influence maximization approach for product marketing that constructs an abstraction hierarchy for scaling optimization techniques to larger networks. An exact solution is computed on smaller partitions of the network, and a candidate set of influential nodes is propagated upward to an abstract representation of the original network that maintains distance information. This process of abstraction, solution, and propagation is repeated until the resulting abstract network is small enough to be solved exactly. Our proposed method scales to much larger networks and outperforms other influence maximization techniques on marketing products.
Mahsa Maghami, Gita Reese Sukthankar
ASONAM2
2013 Link prediction in multi-relational collaboration networks
abstract
Traditional link prediction techniques primarily focus on the effect of potential linkages on the local network neighborhood or the paths between nodes. In this paper, we study the problem of link prediction in networks where instances can simultaneously belong to multiple communities, engendering different types of collaborations. Links in these networks arise from heterogeneous causes, limiting the performance of predictors that treat all links homogeneously. To solve this problem, we introduce a new link prediction framework, Link Prediction using Social Features (LPSF), which weights the network using a similarity function based on features extracted from patterns of prominent interactions across the network.
Gita Reese Sukthankar
ASONAM2
2013 An adjustable autonomy paradigm for adapting to expert-novice differences
abstract
Multi-robot manipulation tasks are challenging for robots to complete in an entirely autonomous way due to the perceptual and cognitive requirements of grasp planning, necessitating the development of specialized user interfaces. Yet even for humans, the task is sufficiently complex that a high level of performance variability exists between a novice and an expert's ability to teleoperate the robots in a sufficiently tightly coupled fashion to manipulate objects without dropping them. The ultimate success of the task relies on the skill level of the human operator to manage and coordinate the robot team. Although most systems focus their effort on forging a unified connection between the robots and the operator, less attention has been spent on the problem of identifying and adapting to the human operator's skill level. In this paper, we present a method for modeling the human operator and adjusting the autonomy levels of the robots based on the operator's skill level. This added functionality serves as a crucial mechanism toward making human operators of any skill level a vital asset to the team even when their teleoperation performance is uneven.
Bennie Lewis, Bulent Tastan, Gita Reese Sukthankar
IROS3
2013 Multi-label relational neighbor classification using social context features
abstract
Networked data, extracted from social media, web pages, and bibliographic databases, can contain entities of multiple classes, interconnected through different types of links. In this paper, we focus on the problem of performing multi-label classification on networked data, where the instances in the network can be assigned multiple labels. In contrast to traditional content-only classification methods, relational learning succeeds in improving classification performance by leveraging the correlation of the labels between linked instances. However, instances in a network can be linked for various causal reasons, hence treating all links in a homogeneous way can limit the performance of relational classifiers.
Gita Reese Sukthankar
KDD2
2013 Tractable POMDP representations for intelligent tutoring systems
abstract
With Partially Observable Markov Decision Processes (POMDPs), Intelligent Tutoring Systems (ITSs) can model individual learners from limited evidence and plan ahead despite uncertainty. However, POMDPs need appropriate representations to become tractable in ITSs that model many learner features, such as mastery of individual skills or the presence of specific misconceptions. This article describes two POMDP representations— state queues and observation chains —that take advantage of ITS task properties and let POMDPs scale to represent over 100 independent learner features. A real-world military training problem is given as one example. A human study ( n = 14) provides initial validation for the model construction. Finally, evaluating the experimental representations with simulated students helps predict their impact on ITS performance. The compressed representations can model a wide range of simulated problems with instructional efficacy equal to lossless representations. With improved tractability, POMDP ITSs can accommodate more numerous or more detailed learner states and inputs.
Jeremiah T. Folsom-Kovarik, Gita Reese Sukthankar, Sae Lynne Schatz
ACM Trans. Intell. Syst. Technol.2
2012 Integrating Learner Help Requests Using a POMDP in an Adaptive Training System
abstract
This paper describes the development and empirical testing of an intelligent tutoring system (ITS) with two emerging methodologies: (1) a partially observable Markov decision process (POMDP) for representing the learner model and (2) inquiry modeling, which informs the learner model with questions learners ask during instruction. POMDPs have been successfully applied to non-ITS domains but, until recently, have seemed intractable for large-scale intelligent tutoring challenges. New, ITS-specific representations leverage common regularities in intelligent tutoring to make a POMDP practical as a learner model. Inquiry modeling is a novel paradigm for informing learner models by observing rich features of learners’ help requests such as categorical content, context, and timing. The experiment described in this paper demonstrates that inquiry modeling and planning with POMDPs can yield significant and substantive learning improvements in a realistic, scenario-based training task.
Jeremiah T. Folsom-Kovarik, Gita Reese Sukthankar, Sae Lynne Schatz
IAAI2
2012 Importance-weighted label prediction for active learning with noisy annotations
Liyue Zhao, Gita Reese Sukthankar, Rahul Sukthankar
ICPR2
2011 Using Network Structure to Identify Groups in Virtual Worlds
Fahad Shah, Gita Reese Sukthankar
ICWSM2
2011 A Real-Time Opponent Modeling System for Rush Football
abstract
One drawback with using plan recognition in adversarial games is that often players must commit to a plan before it is possible to infer the opponent's intentions. In such cases, it is valuable to couple plan recognition with plan repair, particularly in multi-agent domains where complete replanning is not computationally feasible. This paper presents a method for learning plan repair policies in realtime using Upper Confidence Bounds applied to Trees (UCT). We demonstrate how these policies can be coupled with plan recognition in an American football game (Rush 2008) to create an autonomous offensive team capable of responding to unexpected changes in defensive strategy. Our realtime version of UCT learns play modifications that result in a significantly higher average yardage and fewer interceptions than either the baseline game or domain-specific heuristics. Although it is possible to use the actual game simulator to measure reward offline, to execute UCT in real-time demands a different approach; here we describe two modules for reusing data from offline UCT searches to learn accurate state and reward estimators.
Kennard R. Laviers, Gita Reese Sukthankar
IJCAI2
2011 Two hands are better than one: Assisting users with multi-robot manipulation tasks
abstract
Multi-robot manipulation, where two or more robots cooperatively grasp and move objects, is extremely challenging due to the necessity of tightly coupled temporal coordination between the robots. Unfortunately, introducing a human operator does not necessarily ameliorate performance due to the complexity of teleoperating mobile robots with high degrees of freedom. The human operator's attention is divided not only among multiple robots but also between controlling a robot arm and its mobile base. This complexity substantially increases the potential neglect time, since the operator's inability to effectively attend to each robot during a critical phase of the task leads to a significant degradation in task performance. In this paper, we propose an approach for semi-autonomously performing multi-robot manipulation tasks and demonstrate how our user interface reduces both task completion time and the number of dropped items over a fully teleoperated robotic system. Propagating the user's commands from the actively-controlled robot to the neglected robot allows the neglected robot to leverage this control information and position itself effectively without direct human supervision.
Bennie Lewis, Gita Reese Sukthankar
IROS2
2011 Leveraging human behavior models to predict paths in indoor environments
Bulent Tastan, Gita Reese Sukthankar
Pervasive Mob. Comput.2
2011 Activity Recognition for Dynamic Multi-Agent Teams
abstract
This article addresses the problem of activity recognition for dynamic, physically embodied agent teams. We define team activity recognition as the process of identifying team behaviors from traces of agent positions over time; for many physical domains, military or athletic, coordinated team behaviors create distinctive spatio-temporal patterns that can be used to identify low-level action sequences. This article focuses on the novel problem of recovering agent-to-team assignments for complex team tasks where team composition, the mapping of agents into teams, changes over time. We suggest two methods for improving the computational efficiency of the multi-agent plan recognition process in these cases of changing team composition; our proposed approach is robust to sensor observation noise and errors in behavior classification.
Gita Reese Sukthankar, Katia P. Sycara
ACM Trans. Intell. Syst. Technol.1
2010 Motif Discovery and Feature Selection for CRF-based Activity Recognition
abstract
Due to their ability to model sequential data without making unnecessary independence assumptions, conditional random fields (CRFs) have become an increasingly popular discriminative model for human activity recognition. However, how to represent signal sensor data to achieve the best classification performance within a CRF model is not obvious. This paper presents a framework for extracting motif features for CRF-based classification of IMU (inertial measurement unit) data. To do this, we convert the signal data into a set of motifs, approximately repeated symbolic sub sequences, for each dimension of IMU data. These motifs leverage structure in the data and serve as the basis to generate a large candidate set of features from the multi-dimensional raw data. By measuring reductions in the conditional log-likelihood error of the training samples, we can select features and train a CRF classifier to recognize human activities. An evaluation of our classifier on the CMU Multi-Modal Activity Database reveals that it outperforms the CRF-classifier trained on the raw features as well as other standard classifiers used in prior work.
Liyue Zhao, Gita Reese Sukthankar, Rahul Sukthankar
ICPR3
2010 Modeling Group Dynamics in Virtual Worlds
Fahad Shah, Gita Reese Sukthankar, Chris Usher
ICWSM2
2010 Analyzing Team Decision-Making in Tactical Scenarios
abstract
Team decision-making is a bundle of interdependent activities that involve gathering, interpreting and exchanging information; creating and identifying alternative courses of action; choosing among alternatives by integrating the often different perspectives of team members and implementing a choice and monitoring its consequences. To accomplish joint tasks, human team members often assume distinctive roles in task completion. We believe that to design and build software agents that can assist human teams, we need develop automated techniques to identify the roles of the human decision-makers. If the supporting agents are insensitive to shifts in the team's roles, they cannot effectively monitor the team's activities. This article addresses the problem of doing offline role analysis of battle scenarios from multi-player team games. The ability to identify team roles from observations is important for a wide range of applications including automated commentary generation, game coaching and opponent modeling. We define a role as a preference model over possible actions based on the game state. This article explores two promising approaches for automated role analysis: (1) a model-based system for combining evidence from observed events using the Dempster–Shafer theory and (2) a data-driven discriminative classifier using support vector machines.
Gita Reese Sukthankar, Katia P. Sycara
Comput. J.1
2009 Exploiting human steering models for path prediction
Bulent Tastan, Gita Reese Sukthankar
FUSION2
2009 Case-Based Reasoning in Transfer Learning
David W. Aha, Matthew Molineaux, Gita Reese Sukthankar
ICCBR3
2009 Agent-Assisted Navigation for Virtual Worlds
Fahad Shah, Philip Bell, Gita Reese Sukthankar
IVA3
2009 An active learning approach for segmenting human activity datasets
abstract
Human activity datasets collected under natural conditions are an important source of data. Since these contain multiple activities in unscripted sequence, temporal segmentation of multimodal datasets is an important precursor to recognition and analysis. Manual segmentation is prohibitively time consuming and unsupervised approaches for segmentation are unreliable since they fail to exploit the semantic context of the data. Gathering labels for supervised learning places a large workload on the human user since it is relatively easy to gather a mass of unlabeled data but expensive to annotate. This paper proposes an active learning approach for segmenting large motion capture datasets with both small training sets and working sets. Support Vector Machines (SVMs) are learned using an active learning paradigm; after the classifiers are initialized with a small set of labeled data, the users are iteratively queried for labels as needed. We propose a novel method for initializing the classifiers, based on unsupervised segmentation and clustering of the dataset. By identifying and training the SVM with points from pure clusters, we can improve upon a random sampling strategy for creating the query set. Our active learning approach improves upon the initial unsupervised segmentation used to initialize the classifier, while requiring substantially less data than a fully supervised method; the resulting segmentation is comparable to the latter while requiring significantly less effort from the user.
Liyue Zhao, Gita Reese Sukthankar
ACM Multimedia2
2008 Hypothesis Pruning and Ranking for Large Plan Recognition Problems
Gita Reese Sukthankar, Katia P. Sycara
AAAI1
2006 Simultaneous Team Assignment and Behavior Recognition from Spatio-Temporal Agent Traces
Gita Reese Sukthankar, Katia P. Sycara
AAAI1
2003 Shadow Elimination and Occluder Light Suppression for Multi-Projector Displays
abstract
Two related problems of front projection displays, which occur when users obscure a projector, are: (i) undesirable shadows cast on the display by the users, and (ii) projected light falling on and distracting the users. This paper provides a computational framework for solving these two problems based on multiple overlapping projectors and cameras. The overlapping projectors are automatically aligned to display the same dekeystoned image. The system detects when and where shadows are cast by occluders and is able to determine the pixels, which are occluded in different projectors. Through a feedback control loop, the contributions of unoccluded pixels from other projectors are boosted in the shadowed regions, thereby eliminating the shadows. In addition, pixels, which are being occluded, are blanked, thereby preventing the projected light from falling on a user when they occlude the display. This can be accomplished even when the occluders are not visible to the camera. The paper presents results from a number of experiments demonstrating that the system converges rapidly with low steady-state errors.
Tat-Jen Cham, James M. Rehg, Rahul Sukthankar, Gita Reese Sukthankar
CVPR (2)4
2002 Projected light displays using visual feedback
abstract
A system of coordinated projectors and cameras enables the creation of projected light displays that are robust to environmental disturbances. This paper describes approaches for tackling both geometric and photometric aspects of the problem: (1) the projected image remains stable even when the system components (projector, camera or screen) are moved; (2) the display automatically removes shadows caused by users moving between a projector and the screen, while simultaneously suppressing projected light on the user. The former can be accomplished without knowing the positions of the system components. The latter can be achieved without direct observation of the occluder. We demonstrate that the system responds quickly to environmental disturbances and achieves low steady-state errors.
James M. Rehg, Matthew Flagg, Tat-Jen Cham, Rahul Sukthankar, Gita Reese Sukthankar
ICARCV5
2001 Dynamic Shadow Elimination for Multi-Projector Displays
abstract
A major problem with interactive displays based on front-projection is that users cast undesirable shadows on the display surface. This situation is only partially addressed by mounting a single projector at an extreme angle and pre-warping the projected image to undo keystoning distortions. This paper demonstrates that shadows can be muted by redundantly illuminating the display surface using multiple projectors, all mounted at different locations. However, this technique alone does not eliminate shadows: multiple projectors create multiple dark regions on the surface (penumbral occlusions). We solve the problem by using cameras to automatically identify occlusions as they occur and dynamically adjust each projector's output so that additional light is projected onto each partially-occluded patch. The system is self-calibrating: relevant homographies relating projectors, cameras and the display surface are recovered by observing the distortions induced in projected calibration patterns. The resulting redundantly-projected display retains the high image quality of a single-projector system while dynamically correcting for all penumbral occlusions. Our initial two-projector implementation operates at 3 Hz.
Rahul Sukthankar, Tat-Jen Cham, Gita Reese Sukthankar
CVPR (2)3
2001 Self-Calibrating Camera Projector Systems for Interactive Displays and Presentations
abstract
The authors demonstrate a self-calibrating system that employs uncalibrated cameras and microportable projectors to create novel interactive displays and presentations. Three benefits of ther system are detailed.
Rahul Sukthankar, Tat-Jen Cham, Gita Reese Sukthankar, James M. Rehg, David Hsu, Thomas K. Leung
ICCV3