Manfred Huber

dblp:81/682 · DBLP profile ↗
← Back
61ranked-venue papers
5as first author
17since 2021 · last 2025
0009-0007-0294-9147ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Human-computer interaction and ubiquitous computing · 30 · 7 since 2021Applied, interdisciplinary, general and emerging computing · 30 · 1 first-author · 8 since 2021Artificial intelligence and machine learning · 25 · 4 first-author · 8 since 2021Systems, architecture and hardware · 15 · 3 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 1 since 2021Computer networks · 3
YearPublicationVenuePosition
2025 Label-Efficient Human Activity Recognition from Wearables via Self-Supervised Representation Learning
Taoran Sheng, Manfred Huber
ICMLA2
2025 Modeling The States of Liquid Phase Change Pouch Actuators by Reservoir Computing
abstract
Liquid phase change pouch actuators (liquid pouch motors) hold great promise for a wide range of robotic applications, from artificial organs to pneumatic manipulators for dexterous manipulation. However, the usability of liquid pouch motors remains challenging due to the nonlinear intrinsic properties of liquids and their highly dynamic implications for liquid-gas phase changes, which complicate state modeling and estimation. To address these issues, we propose a reservoir computing-based method for modeling the inflation states of a customized liquid pouch motor, which serves as an actuator, featuring four Peltier heating junctions. We use a motion capture system to track the landmark movements on the pouch as a proxy for its volumetric profile. These movements represent the internal liquid-gas phase changes of the pouch at stable room temperature, atmospheric pressure, and in the presence of electrical noise. The motion coordinates are thus learned by our reservoir computing framework, PhysRes, to model the states based on prior observations. Through training, our model achieves excellent results on the test set, with a normalized root mean squared error of 0.0041 in estimating the states and a corresponding volumetric error of 0.0160%. To further demonstrate how such actuators could be implemented in the future, we also design a dual-pouch actuator-based robotic gripper to control the grasping of soft objects. Our design and source code are available at: https://github.com/tatung/liquidpouch_reservoir.
Cedric Caremel, Anh Nguyen 0003, Manfred Huber, Yoshihiro Kawahara, Tung D. Ta
IROS4
2025 Bio-Inspired Hybrid Map: Spatial Implicit Local Frames and Topological Map for Mobile Cobot Navigation
abstract
Navigation is a fundamental capacity for mobile robots, enabling them to operate autonomously in complex and dynamic environments. Conventional approaches use probabilistic models to localize robots and build maps simultaneously using sensor observations. Recent approaches employ human-inspired learning, such as imitation and reinforcement learning, to navigate robots more effectively. However, these methods suffer from high computational costs, global map inconsistency, and poor generalization to unseen environments. This paper presents a novel method inspired by how humans perceive and navigate themselves effectively in novel environments. Specifically, we first build local frames that mimic how humans represent essential spatial information in the short term. Points in local frames are hybrid representations, including spatial information and learned features, so-called spatial-implicit local frames. Then, we integrate spatial-implicit local frames into the global topological map represented as a factor graph. Lastly, we developed a novel navigation algorithm based on Rapid-Exploring Random Tree Star (RRT*) that leverages spatial-implicit local frames and the topological map to navigate effectively in environments. To validate our approach, we conduct extensive experiments in real-world datasets and in-lab environments. We open our source code at https://github.com/tuantdang/simn.
Tuan Dang, Manfred Huber
IROS2
2025 FlowMP: Learning Motion Fields for Robot Planning with Conditional Flow Matching
abstract
Prior flow matching methods in robotics have primarily learned velocity fields to morph one distribution of trajectories into another. In this work, we extend flow matching to capture second-order trajectory dynamics, incorporating acceleration effects either explicitly in the model or implicitly through the learning objective. Unlike diffusion models, which rely on a noisy forward process and iterative denoising steps, flow matching trains a continuous transformation (flow) that directly maps a simple prior distribution to the target trajectory distribution without any denoising procedure. By modeling trajectories with second-order dynamics, our approach ensures that the generated robot motions are smooth and physically executable, avoiding the jerky or dynamically infeasible trajectories that first-order models might produce. We empirically demonstrate that this second-order conditional flow matching yields superior performance on motion planning benchmarks, achieving smoother trajectories and higher success rates than baseline planners. These findings highlight the advantage of learning acceleration-aware motion fields, as our method outperforms existing motion planning methods in terms of trajectory quality and planning success. Our source code is available at: https://github.com/mkhangg/flow_mp.
Khang Nguyen 0003, An T. Le 0001, Tien Pham, Manfred Huber, Jan Peters 0001, Minh Nhat Vu
IROS4
2025 Categorical Policies: Multimodal Policy Learning and Exploration in Continuous Control
abstract
A policy in deep reinforcement learning (RL), either deterministic or stochastic, is commonly parameterized as a Gaussian distribution alone, limiting the learned behavior to be unimodal. However, the nature of many practical decision-making problems favors a multimodal policy that facilitates robust exploration of the environment and thus to address learning challenges arising from sparse rewards, complex dynamics, or the need for strategic adaptation to varying contexts. This issue is exacerbated in continuous control domains where exploration usually takes place in the vicinity of the predicted optimal action, either through an additive Gaussian noise or the sampling process of a stochastic policy. In this paper, we introduce Categorical Policies to model multimodal behavior modes with an intermediate categorical distribution, and then generate output action that is conditioned on the sampled mode. We explore two sampling schemes that ensure differentiable discrete latent structure while maintaining efficient gradient-based optimization. By utilizing a latent categorical distribution to select the behavior mode, our approach naturally expresses multimodality while remaining fully differentiable via the sampling tricks. We evaluate our multimodal policy on a set of DeepMind Control Suite environments, demonstrating that through better exploration, our learned policies converge faster and outperform standard Gaussian policies. Our results indicate that the Categorical distribution serves as a powerful tool for structured exploration and multimodal behavior representation in continuous control.
SM Mazharul Islam, Manfred Huber
SMC2
2025 A Dual-Agent Learning Framework for Emotion-Aware Personalized Game Level Generation
abstract
A dual-agent learning framework is proposed for emotion-aware, personalized game level generation, minimizing real user interaction while maximizing engagement. The Inner Agent models player behavior using a Siamese Network to generate user embeddings, a Gaussian Mixture Model (GMM) to capture user-type distributions, and a Conditional GAN (CGAN) to simulate performance data conditioned on game state, emotion, and user embeddings. The Outer Agent, a Q-learning-based Reinforcement Learning (RL) agent, selects from 10 predefined game states based on player performance and facial emotion data. In experiments with 100 real-user interactions and 100 simulated interactions per iteration, the framework reduced the number of real interactions required to reach peak performance compared to a baseline without the Inner Agent. These results demonstrate the framework’s effectiveness in scalable, emotion-sensitive game personalization with reduced user burden.
Subharag Sarkar, Manfred Huber
SMC2
2024 V3D-SLAM: Robust RGB-D SLAM in Dynamic Environments with 3D Semantic Geometry Voting
abstract
Simultaneous localization and mapping (SLAM) in highly dynamic environments is challenging due to the correlation complexity between moving objects and the camera pose. Many methods have been proposed to deal with this problem; however, the moving properties of dynamic objects with a moving camera remain unclear. Therefore, to improve SLAM’s performance, minimizing disruptive events of moving objects with a physical understanding of 3D shapes and dynamics of objects is needed. In this paper, we propose a robust method, V3D-SLAM, to remove moving objects via two lightweight reevaluation stages, including identifying potentially moving and static objects using a spatial-reasoned Hough voting mechanism and refining static objects by detecting dynamic noise caused by intra-object motions using Chamfer distances as similarity measurements. Through our experiment on the TUM RGB-D benchmark on dynamic sequences with ground-truth camera trajectories, the results show that our methods outperform most other recent state-of-the-art SLAM methods. Our source code is available at https://github.com/tuantdang/v3d-slam.
Tuan Dang, Khang Nguyen 0003, Manfred Huber
IROS3
2024 Volumetric Mapping with Panoptic Refinement using Kernel Density Estimation for Mobile Robots
abstract
Reconstructing three-dimensional (3D) scenes with semantic understanding is vital in many robotic applications. Robots need to identify which objects, along with their positions and shapes, to manipulate them precisely with given tasks. Mobile robots, especially, usually use lightweight networks to segment objects on RGB images and then localize them via depth maps; however, they often encounter out-of-distribution scenarios where masks over-cover the objects. In this paper, we address the problem of panoptic segmentation quality in 3D scene reconstruction by refining segmentation errors using non-parametric statistical methods. To enhance mask precision, we map the predicted masks into a depth frame to estimate their distribution via kernel densities. The outliers in depth perception are then rejected without the need for additional parameters in an adaptive manner to out-of-distribution scenarios, followed by 3D reconstruction using projective signed distance functions (SDFs). We validate our method on a synthetic dataset, which shows improvements in both quantitative and qualitative results for panoptic mapping. Through real-world testing, the results furthermore show our method’s capability to be deployed on a real-robot system. Our source code is available at: https://github.com/mkhangg/refined_panoptic_mapping.
Khang Nguyen 0003, Tuan Dang, Manfred Huber
IROS3
2024 Domain Knowledge Based Weakly Self-Supervised Human Activity Recognition With Wearables
Taoran Sheng, Manfred Huber
SMC2
2023 A SSIM Guided cGAN Architecture For Clinically Driven Generative Image Synthesis of Multiplexed Spatial Proteomics Channels
abstract
Histopathological work in clinical labs often relies on immunostaining of proteins, which can be time-consuming and costly. Multiplexed spatial proteomics imaging can increase interpretive power, but current methods cannot cost-effectively sample the entire proteomic retinue important to diagnostic medicine or drug development. To address this challenge, we developed a conditional generative adversarial network (cGAN) that performs image-to-image (i2i) synthesis to generate accurate biomarker channels in multiplexed spatial proteomics images1. We approached this problem as missing biomarker expression generation, where we assumed that a given n-channel multiplexed image has p channels (biomarkers) present and q channels (biomarkers) absent, with p+q=n, and we aimed to generate the missing q channels. To improve accuracy, we selected p and q channels based on their structural similarity, as measured by a structural similarity index measure (SSIM). We demonstrated the effectiveness of our approach using spatial proteomic data from the Human BioMolecular Atlas Program (HuBMAP)2, which we used to generate spatial representations of missing proteins through a U-Net based image synthesis pipeline. Channels were hierarchically clustered by SSIM to obtain the minimal set needed to recapitulate the underlying biology represented by the spatial landscape of proteins. We also assessed the scalability of our algorithm using regression slope analysis, which showed that it can generate increasing numbers of missing biomarkers in multiplexed spatial proteomics images. Furthermore, we validated our approach by generating a new spatial proteomics data set from human lung adenocarcinoma tissue sections and showed that our model could accurately synthesize the missing channels from this new data set. Overall, our approach provides a cost-effective and time-efficient alternative to traditional immunostaining methods for generating missing biomarker channels, while also increasing the amount of data that can be generated through experiments. This has important implications for the future of medical diagnostics and drug development, and raises important questions about the ethical implications of utilizing data produced by generative image synthesis in the clinical setting.1https://github.com/aauthors131/mu1tip1exed-image-synthesis2https://portal.hubmapconsortium.org
Jillur Rahman Saurav, Mohammad Sadegh Nasr 0001, Helen H. Shang, Paul Koomey, Michael Robben, Manfred Huber, Jon Weidanz, Bríd Ryan, Eytan Ruppin, Jacob M. Luber
CIBCB6
2023 Multiplanar Self-Calibration for Mobile Cobot 3D Object Manipulation Using 2D Detectors and Depth Estimation
abstract
Calibration is the first and foremost step in dealing with sensor displacement errors that can appear during extended operation and off-time periods to enable robot object manipulation with precision. In this paper, we present a novel multiplanar self-calibration between the camera system and the robot's end-effector for 3D object manipulation. Our approach first takes the robot end-effector as ground truth to calibrate the camera's position and orientation while the robot arm moves the object in multiple planes in 3D space, and a 2D state-of-the-art vision detector identifies the object's center in the image coordinates system. The transformation between world coordinates and image coordinates is then computed using 2D pixels from the detector and 3D known points obtained by robot kinematics. Next, an integrated stereo-vision system estimates the distance between the camera and the object, resulting in 3D object localization. We test our proposed method on the Baxter robot with two 7-DOF arms and a 2D detector that can run in real time on an onboard GPU. After self-calibrating, our robot can localize objects in 3D using an RGB camera and depth image. The source code is available at https://github.com/tuantdang/calib_cobot.
Tuan Dang, Khang Nguyen 0003, Manfred Huber
IROS3
2022 Learning Hierarchical Traversability Representations for Efficient Multi-Resolution Path Planning*
abstract
Path planning on grid-based obstacle maps is an essential and much-studied problem with applications in robotics and autonomy. Traditionally, in the AI community, heuristic search methods (e.g., based on Dijkstra, $\text{A}^{*}$, or random trees) are used to solve this problem. This search, however, incurs a high computational cost that grows with the size and resolution of the obstacle grid and has to be mitigated with effective heuristics to allow path planning in real-time. This work introduces a learning framework using a deep neural network with a stackable convolution kernel to establish a hierarchy of directional traversability representations with decreasing resolution that can serve as an efficient heuristic to guide a multi-resolution path planner. This path planner finds paths efficiently, starting on the lowest resolution traversability representation and then refining the path incrementally through the hierarchy until it addresses the original obstacle constraints. We demonstrate the benefits and applicability of this approach on datasets of maps created to represent both indoor and outdoor environments to represent different real-world applications. The conducted experiments show that our method can accelerate path planning by 40% in indoor environments and 65% in outdoor environments compared to the same heuristic search method applied to the original obstacle map, which demonstrates the effectiveness of this method.
Reza Etemadi Idgahi, Manfred Huber
SMC2
2022 An AI-based Approach for Improved Sign Language Recognition using Multiple Videos
Cameron Dignan, Eliud Perez, Ishfaq Ahmad 0001, Manfred Huber, Addison Clark
Multim. Tools Appl.4
2021 Learning the Next Best View for 3D Point Clouds via Topological Features
abstract
In this paper, we introduce a reinforcement learning approach utilizing a novel topology-based information gain metric for directing the next best view of a noisy 3D sensor. The metric combines the disjoint sections of an observed surface to focus on high-detail features such as holes and concave sections. Experimental results show that our approach can aid in establishing the placement of a robotic sensor to optimize the information provided by its streaming point cloud data. Furthermore, a labeled dataset of 3D objects, a CAD design for a custom robotic manipulator, and software for the transformation, union, and registration of point clouds has been publicly released to the research community.
Christopher Collander, William J. Beksi, Manfred Huber
ICRA3
2021 An Automatic Calibration Technique for Force Sensors in a Dynamic Smart Floor Environment
abstract
Pressure-sensitive smart floors deployed within homes can give great insight to the health and activity level of individuals through gait and location information. Due to the ever-changing dynamic nature of household deployments involving furniture movement, floor tile shifts, and sensor drift, challenges arise in ensuring the constant reliability of floor sensor readings over time. This paper presents a procedure to automatically calibrate a smart floor’s force sensors without specialized physical effort. The calibration algorithm automatically filters out non-human static weight while retaining weight generated by human activity. This technique is designed to correctly translate sensor values to weight units even when direct access to the force sensors is not available and when a shared tile floor sits above the sensor grid. These calibrated sensor values can then feed machine learning techniques used to extract individual contact points generated by a person’s walking cycle. Using known human weights but no knowledge of the human’s location or walking trajectory, this calibration technique resulted in small percentage differences of -7.8%, -4.8%, and -1.6% for the mean, median, and mode of calibrated smart floor walking sequences, respectively.
Nicholas Brent Burns, Kathryn Daniel, Manfred Huber, Gergely V. Záruba
SMC3
2021 Personalized Learning Path Generation in E-Learning Systems using Reinforcement Learning and Generative Adversarial Networks
abstract
Accelerated by the pandemic and the resulting increased adoption of online classes by universities, E-learning and the challenge to increase its efficiency in terms of conveying content effectively to the individual learner has gained interest and importance. To achieve this customization, personalization of the learning path and the learning objects based on the characteristics of the learner could form an important component that is currently largely absent from the most used e-learning modalities. This paper introduces an approach to personalization of the learning path that attempts to optimize the learner’s performance using Reinforcement Learning based on implicit feedback while minimizing the need for actual interaction with the learning system during training. The proposed system adopts the concepts of the Felder and Silverman Learning Style Model and Differentiated Pedagogy and builds an architecture using two interacting machine learning components, one using Reinforcement Learning to optimize the learning path and learning objects, and one using Conditional Generative Adversarial Networks to rapidly adapt a model of the learner’s characteristics. The model strives to minimize student interactions to learn about the learner’s performance characteristics and to generate personalized learning paths, thus reducing negative experiences during training. For this, a Conditional Generative Adversarial Networks is used that can rapidly adapt to provide realistic simulations of the student’s performance to use in the training of the e-learning strategy and thus to help reduce the need for actual learner feedback. In a set of base experiments with synthetic data, the model is shown to be effective at learning personalization for different learner types and efficient compared to models without Generative Adversarial Networks.
Subharag Sarkar, Manfred Huber
SMC2
2021 A Deep Reinforcement Learning Approach to Learning Tree Edit Policies on Binary Trees
abstract
Modeling sequential transformations between two trees is a fundamental task in domains such as bioinformatics. Traditional methods to address this usually rely on hand-coded heuristics and algorithms which are often hard to derive. The large advances in deep learning technologies and computational power have recently opened up new potential avenues. In particular, representation and policy learning provides an interesting opportunity to model the dynamic evolution of two trees, where each tree can be embedded in a Euclidean space and its evolution can be modeled by an embedding trajectory in this space.In this paper, we propose a representation and policy learning framework that learns a representation for arbitrary sized binary tree pairs using recurrent LSTM networks and a policy to transfer one tree in to the corresponding target tree using Reinforcement Learning. Here, the representation is pre-trained on tree transfer similarity to transform pairs of tree-structured data into an approximate numerical multidimensional vector which encodes the original structure information. This model, used with a deep reinforcement learning approach, yields a constructive method for generating basis functions for approximating value functions and permits to learn an efficient, general tree transfer policy that incrementally transforms the source tree into the target tree.
Shirin Shirvani, Manfred Huber
SMC2
2020 Evolutionary Feature Scaling in K-Nearest Neighbors Based on Label Dispersion Minimization
abstract
K-Nearest Neighbors (KNN) has remained one of the most popular methods for supervised machine learning tasks. However, its performance often depends on the characteristics of the dataset and on appropriate feature scaling. In this paper, we explore characteristics of a dataset that make it suitable for being used within KNN. As part of this, two new measures for dataset dispersion, called mean neighborhood target standard deviation (MNTSD), and mean neighborhood target entropy (MNTE) are formulated to determine the expeced performance while using KNN regressors and classifiers, respectively. It is empirically demonstrated that these measures of dispersion can be indicative of the performance of KNN regression and classification. This idea is further used to learn feature weights that help improve the accuracy of KNN classification and regression. For this, it is argued that the MNTSD and MNTE, when used to learn feature weights, cannot be optimized using gradient-based optimization methods and we develop optimization strategies based on metaheuristic methods, namely genetic algorithms and particle swarm optimization. The feature-weighting method is tried in both regression and classification contexts on publicly available datasets, and the performance is compared to KNN without feature weighting. The results indicate that the performance of KNN with appropriate feature weighting leads to better performance.
Suryoday Basak, Manfred Huber
SMC2
2019 Show, Infer and Tell: Contextual Inference for Creative Captioning
Ankit Khare, Manfred Huber
BMVC2
2019 MDP Autoencoder
abstract
This paper proposes a novel deep reinforcement learning (RL) architecture, which learns a dynamics model in latent space that is behaviorally grounded to the observed space and applies the framework of MDP homomorphisms to provide bounds for the loss in performance. In contrast to traditional model based reinforcement learning algorithms, this approach models the next latent state while predicting the rewards and chance of termination after current step, instead of reconstructing the future observations. This results in a much simpler model that can encode the underlying dynamics of the observed system. Experiments show that the proposed approach learns concise latent Markov models that can encode policies which achieve scores close to optimal.
Sourabh Bose, Manfred Huber
SMC2
2019 A Utility-Based Path Planning for Safe UAS Operations with a Task-Level Decision-Making Capability
abstract
Unmanned aircraft systems (UAS) are being used more and more every day in almost any area to solve challenging real-life problems. Increased autonomy and advancements in low-cost high-computing technologies made these compact autonomous solutions accessible to any party with ease. However, this ease of use brings its own challenges that need to be addressed. In an autonomous flight scenario over a public space, an autonomous operation plan has to consider the public safety and regulations as well as the task specific objectives. In this work, we propose a generic utility function for the path planning of UAS operations that includes the benefits of accomplishing the goals as well as the safety risks incurred along the flight trajectories, with the purpose of making task-level decisions through the optimization of the carefully constructed utility function for a given scenario. As an optimizer, we benefited from a multi-tree variant of the optimal T-RRT*(Multi-T-RRT*path planning algorithm. To illustrate its operation, results of simulation of a UAS scenario are presented.
Uluhan C. Kaya, Atilla Dogan, Manfred Huber
SMC3
2019 Siamese Networks for Weakly Supervised Human Activity Recognition
abstract
Deep learning has been successfully applied to human activity recognition. However, training deep neural networks requires explicitly labeled data which is difficult to acquire. In this paper, we present a model with multiple siamese networks that are trained by using only the information about the similarity between pairs of data samples without knowing the explicit labels. The trained model maps the activity data samples into fixed size representation vectors such that the distance between the vectors in the representation space approximates the similarity of the data samples in the input space. Thus, the trained model can work as a metric for a wide range of different clustering algorithms. The training process minimizes a similarity loss function that forces the distance metric to be small for pairs of samples from the same kind of activity, and large for pairs of samples from different kinds of activities. We evaluate the model on three datasets to verify its effectiveness in segmentation and recognition of continuous human activity sequences.
Taoran Sheng, Manfred Huber
SMC2
2017 Training neural networks with policy gradient
abstract
Neural networks are a powerful function approximation tool which has the ability to model any function with arbitrary precision. For any function as a black box, it is able to reconstruct the function given the target and the input data. However, there are problems where the target is at least partially unknown. In such cases it is impossible for a traditional neural network to compute the gradient of the system. This problem is evident in sparse autoencoder systems where lateral inhibitions are required. Lateral inhibitions are usually imposed with extra lateral connections among nodes in the immediate vicinity of the hidden layer. However, it is computationally expensive, and results in a very complex architecture which is not scalable to real world problems. Thus it is often necessary to achieve structural constraints without such complexity limitations. Similar situation arises in case of multi-label classification problems without a known target. In such cases only an evaluation is accessible to the network, which states whether the sample was correctly classified or not. In this problem, it is necessary to derive the gradient for this partial information to solve the problem. The proposed model solves such problems using policy gradient algorithms. This approach allows for imposing any arbitrary non-differentiable constraints on the neural network system by deriving the required gradient from a system of rewards. Furthermore a novel form of sparsity with lateral inhibitions over the entire hidden layer is proposed. It is shown that the proposed model is not only able to achieve a much tighter and sparser representation, it also achieves similar or better results than traditional forms of sparsity.
Sourabh Bose, Manfred Huber
IJCNN2
2016 Incremental learning of neural network classifiers using reinforcement learning
abstract
With the availability of more data, classification is increasingly important. However, traditional classification algorithms do not scale well to large data sets and are often not suited when only limited samples of the dataset are available at any point in time. The latter arises, for example, in streaming data when the accumulation of data a priori is infeasible either due to limitations in memory or computation, or due to privacy and data ownership limitations. In these situations, traditional classification algorithms are difficult to apply since they are generally not incrementally trainable on changing data sets. To address this, this paper presents a novel approach that first uses Reinforcement Learning to learn a policy to incrementally build neural network classifiers for a broad distribution of problems and subsequently applies it to new data to learn a classifier for this specific problem. In both phases, learning operates on a sequence of small, randomly drawn subsets of the data, thus making it suitable for streaming data and for very large data sets where processing the entire set is not feasible. Experiments comparing this approach with kernel SVMs and large neural networks applied to the complete dataset show that this approach achieves comparable performance. Additional experiments were done to evaluate the performance of this approach for real world, streaming datasets and datasets with concept drift properties.
Sourabh Bose, Manfred Huber
SMC2
2016 Temporal and agent abstractions in multiagent reinforcement learning
abstract
A major challenge in the area of multiagent reinforcement learning has been addressing the problem of scale, more specifically the fact that increasing the number of agents in a system dramatically increases both the cost of representing the problem and the cost of calculating a solution. In single agent systems, temporal abstractions in the form of options have been used to address part of the scaling problem, but only limited work exists for multiagent systems, largely limited to cooperative games. This paper presents a formalization of options for multiagent systems and introduces a framework for agent abstraction that treats coalitions executing options analogously to agents with policies, resulting in a lower-dimensional game whose equilibria approximately correspond to equilibria in the higher dimensional game.
Danielle M. Clement, Manfred Huber
SMC2
2016 Dynamic heuristic planner selection
abstract
Heuristic search is considered state-of-the-art for classical planning. However, the performance of search heuristics varies significantly from problem to problem and no single heuristic is superior to all others. As a result, it is highly desirable to identify and utilize the best available heuristic for a particular planning problem. This paper presents a novel approach for planning that monitors the search dynamics of a heuristic planner over time in order to recognize whether the planner is making progress toward a solution. It then dynamically selects from a set of heuristic planners during the planning process so that planners that appear to be making progress are allocated more processor time. Experimental results show this approach is more effective than static approaches of dividing processor time equally between planners or selecting any one planner a priori.
Brian Cook, Manfred Huber
SMC2
2013 Deep Belief Network for Modeling Hierarchical Reinforcement Learning Policies
abstract
Intelligent agents over their lifetime face multiple tasks that require simultaneous modeling and control of complex, initially unknown environments, observed via incomplete and uncertain observations. In such scenarios, policy learning is subject to the curse of dimensionality, leading to scaling problems for traditional Reinforcement Learning (RL). To address this, the agent has to efficiently acquire and reuse latent knowledge. One way is through Hierarchical Reinforcement Learning (HRL), which embellishes RL with a hierarchical, model-based approach to state, reward and policy representation. This paper presents a novel learning approach for HRL based on Conditional Restricted Boltzmann Machines (CRBMs). The proposed model provides a uniform means to simultaneously learn policies and associated abstract state features, and allows learning and executing hierarchical skills within a consistent, uniform network structure. In this model, learning is performed incrementally from basic grounded features to complex abstract policies based on automatically extracted latent states and rewards.
Predrag Djurdjevic, Manfred Huber
SMC2
2013 Data Modeling Using Channel-Remapped Generalized Features
abstract
Sparse coding is a very powerful method to learn high-level features from raw data input. It is able to learn an over complete basis that has the potential to capture robust and discriminative patterns within the data. However, like many other feature learning algorithms, it is unable to detect very similar features or stimuli on different input channels. In this paper, we propose a novel method to build general features that can be applicable to different sets of channels. This succinct representational model will express the stimuli independent of the locality in which they appeared. As a result, it prepares the groundwork for transferring the learned features from a set of input channels to other possible sets of input channels.
Houtan Rahmanian, Manfred Huber
SMC2
2012 A Sampling-Based Approach to Reducing the Complexity of Continuous State Space POMDPs by Decomposition Into Coupled Perceptual and Decision Processes
abstract
In this paper, we propose a method to reduce the complexity of solving POMDPs in continuous state spaces by decomposing them into separate, coupled perceptual and decision processes which leads to a reduction of the state space size of the decision learning problem. In our method, we reduce the state space of the POMDP by handling some aspects of the state space outside of the decision POMDP. To achieve this, the whole problem state space is decomposed into separate state spaces for the decision and perceptual process. The Perceptual process just serves to estimate aspects of the belief state while the decision process estimates the remainder and determines a policy. As a result, the decision process is modeled as a reduced state space POMDP. To allow the application of this method to continuous state spaces, the decision and the perceptual processes are here both handled by a sampling method within which this separation makes it possible to represent the POMDP with a smaller state space which leads to smaller sample sets for the decision POMDP and as a result to reduced representational and decision learning complexity. The goal here is to focus decision learning on the aspects of the space that are important for decision making while the observations and attributes that are important for estimating the state of the decision process are handled separately by the perceptual process. In this way, the separation into different processes can significantly reduce the complexity of decision learning. In the proposed framework and algorithm, Monte Carlo based sampling methods and corresponding sample set representations are used for both the perceptual and decision processes to be able to deal efficiently with continuous domains. We show analytically and experimentally how much the complexity of solving a POMDP can be reduced to increase the range of decision learning tasks that can be addressed.
Rasool Fakoor, Manfred Huber
ICMLA (1)2
2012 A Game Theoretic Framework for Communication in Fully Observable Multiagent Systems
abstract
Communication is an important element of multiagent systems (MAS). In fully decentralized systems it is needed to allow the agents to coordinate their actions to achieve certain goals. When the agents have no means to coordinate their actions, they generally choose actions that minimize their chance of losses. If the agents were allowed to coordinate, on the other hand, they can choose actions that allow them to get higher rewards instead. In game theory, these concepts are known as risk dominance and payoff dominance. In this paper, we model communication between agents in stochastic games as an extensive form game in which the agents can numerically evaluate the benefit of communicating a piece of information to the other agents. This allows the agent to answer two important questions about the communication process, namely when should an agent communicate, and what should an agent communicate. We also present a more stable way to select between Nash equilibria in multiagent reinforcement learning using NashQ which is used in the calculation of the values for the communication game.
Tummalapalli Sudhamsh Reddy, Gergely V. Záruba, Manfred Huber
ICMLA (1)3
2012 Improving tractability of POMDPs by separation of decision and perceptual processes
abstract
Markov Decision Processes (MDPs) and Partially Observable Markov Decision Processes (POMDPs) are very powerful frameworks to model decision and decision learning tasks in a wide range of problem domains. Thus, they are used widely in complex and real-world situations such as robot control tasks. However, this modeling power and generality of the framework comes at a cost in that the complexity of the underlying model and corresponding algorithms grows dramatically as the complexity of the task domain increases. To address this issue in the context of tasks where raw sensory features are used as a basis for complex decision making, this paper presents an integrated and adaptive approach that attempts to reduce the complexity of the decision learning problem by separating the POMDP model into separate decision and perceptual processes. In the proposed framework, a sampling method is used for the perceptual process and reinforcement learning serves to address the decision process. Handling the perceptual and decision processes separately here promises the potential to make it easier to extract relevant perceptual information and concentrate the decision process on relevant state attributes. This, in turn, promises to allow the framework to scale to problems in which traditional POMDP methods are intractable. We show and discuss the effectiveness of our method analytically and empirically.
Rasool Fakoor, Manfred Huber
SMC2
2012 Symbol generation and feature selection for reinforcement learning agents using affordances and U-Trees
abstract
One of the challenges for artificial agents is managing the complexity of their environment and task domain as they learn increasingly difficult tasks. This is especially true of agents that are grounded in the physical world, which contains a vast number of features and potentially very complex dynamics. A scalable solution to this problem in terms of forming, managing, and re-using compact, grounded representations in order to address the state explosion problem is thus a prerequisite of physically grounded, agent-based systems that can apply their past experience to new tasks and communicate that experience with other agents. To achieve this, it is essential that agents can form conceptual features that are relevant for and re-usable in their task domain without outside intervention and that these agents can effectively focus their attention on only the relevant features and concepts for the task at hand. This paper presents a framework for managing state complexity by automatically constructing abstract, symbolic features which encode important, task and domain-relevant properties and partition the raw feature space such that the agent need only consider a compressed view of the environment when learning new tasks. To exploit this, the framework during learning of new tasks uses U-Trees to construct minimal feature sets and thus compact state representations for these new tasks, allowing for potentially significant improvements in learning times.
Marcus Carlos Oladell, Manfred Huber
SMC2
2012 Inverse reinforcement learning for decentralized non-cooperative multiagent systems
abstract
The objective of inverse reinforcement learning (IRL) is to learn an agent's reward function based on either the agent's policies or the observations of the policy. In this paper we address the issue of using inverse reinforcement learning to learn the reward function in a multi agent setting, where the agents can either cooperate or be strictly non-cooperative. The case of cooperataing agents is a subcase of the non-cooperative setting, where the agents collectively try to maximize a common reward function, instead of maximizing their individual reward functions. Here we present an IRL algorithm that considers the case where the policies of the agents are known. We use the framework that was described by Ng and Russell [2001] and extend it for a Multiagent setting. We assume that the agents are rational and follow an optimal policy in the sense of the Nash Equilibrium. These assumptions are very common in Multiagent systems. We show that in the case of known policies we can reduce the Multiagent problem to a distributed solution where the reward function for each agent can be solved independently using a very similar formulation as for the single agent case.
Tummalapalli Sudhamsh Reddy, Vamsikrishna Gopikrishna, Gergely V. Záruba, Manfred Huber
SMC4
2011 Reinforcement field
abstract
Complex control tasks involving varying or evolving system dynamics often pose a great challenge to mainstream reinforcement learning algorithms. Specifically, in most standard methods, actions are often assumed to be a concrete and fixed set that applies to the state space in a predefined manner. Consequently, without resorting to a substantial re-learning procedure, the derived policy lacks the ability to adapt to variations in action outcomes or shifts in the action set. In addition, the standard action representation and its attendant state transition mechanism limit the applicability of the RL framework in complex domains primarily due to the intractability of the resulting large state space and lack of the facility to generalize the learned policy to the unknown parts of the state space. This paper proposes an alternative view of reinforcement learning by establishing the notion of the reinforcement field through a collection of policy-embedded particles gathered during the policy learning process. The reinforcement field serves as a policy generalization mechanism through the use of kernel functions as a state correlation hypothesis in combination with Gaussian process regression as a value function approximator.
Po-Hsiang Chiu, Manfred Huber
SMC2
2011 Generalized reinforcement learning with concept-driven abstract actions
abstract
The standard reinforcement learning framework often faces challenges in a varying or evolving environment due to an inherent limitation in its representation. In particular, useful actions for decision making are often assumed to be a prefixed set prior to the learning process. Consequently, the derived policy in general lacks the ability to adapt to possible variations in the action outcomes or the action set itself without resorting to a substantial re-learning process. In addition, complexity in the state space modeling is often a bottleneck for standard learning methods. This paper proposes a new framework of reinforcement learning that enables the agent to formulate an action-oriented conceptual model while deriving the decision policy simultaneously. The new framework, Concept-Driven Learning Architecture (CDLA), formulates the abstract actions based on associating the correlated past decision history. Specifically, the kernel function, Gaussian process and spectral clustering mechanisms are combined into a functional clustering method to identify a set of coherent, concept-driven abstract actions using which the agent derives a control policy.
Po-Hsiang Chiu, Manfred Huber
SMC2
2011 Autonomous identification, categorization and generalization of policies based on task type
abstract
A life-long learning agent must have the ability to learn new tasks, adapt the policies of already learned tasks, and extract and reuse knowledge from previous tasks for future use. To do the latter, it needs methods that can autonomously identify, categorize and generalize control and representational knowledge. This paper presents a novel approach to achieve this by combining the policy homomorphism framework with a utility criterion to autonomously identify task types, categorize situation-specific policy instances into these types, and generalize the policies into a single abstract policy for each identified task type. The capabilities of this approach to identify, categorize, and generalize skills, as well as the potential benefit of reuse of the abstracted policies for the learning of new tasks is demonstrated in a grid world domain.
Srividhya Rajendran, Manfred Huber
SMC2
2011 Building Bayesian Network based expert systems from rules
abstract
Combining expert knowledge and user explanation with automated reasoning in domains with uncertain information poses significant challenges in terms of representation and reasoning mechanisms. In particular, reasoning structures understandable and usable by humans are often different from the ones used for automated reasoning and data mining systems. Rules with certainty factors represent one possible way to express domain knowledge and build expert system that can deal with uncertainty. Although convenient to humans, this approach has limitations in accurately modeling the domain. Alternatively, a Bayesian Network allows accurate modeling of a domain and automated reasoning but its inference is less intuitive to humans. In this paper, we propose a method to combine these two frameworks to build Bayesian Networks from rules and derive user understandable explanations in terms of these rules. Expert specified rules are augmented with importance parameters for antecedents and are used to derive probabilistic bounds for the Bayesian Network's conditional probability table. The partial structure constructed from the rules is fully learned from the data. The paper also discusses methods for using the rules to provide user understandable explanations, identify incorrect rules, suggest new rules and perform incremental learning.
Saravanan Thirumuruganathan, Manfred Huber
SMC2
2010 Pseudo-hierarchical ant-based clustering using a heterogeneous agent hierarchy and automatic boundary formation
abstract
The behavior and self-organization of ant colonies provides a promising model to address distributed clustering. However, most ant-based clustering approaches suffer from inefficiencies due to large numbers of unproductive ant movements and inefficient cluster merging, leading them to produce too many clusters and to converge too slowly. To address these issues, this paper presents a new ant-based clustering algorithm in which ants are organized in a loose two-level hierarchy with worker ants maintaining movement zone boundaries around each cluster and organizing its internal structure while a single queen ant in each cluster is responsible for moving items between clusters by directly handing them to other queens. This provides an infrastructure that avoids excessive ant movements between cluster regions while allowing for efficient long distance cluster merging. Comparison of this approach with traditional ant-based clustering shows its promise to significantly improve performance and scalability.
Jeremy B. Brown, Manfred Huber
GECCO2
2010 Pseudo-Hierarchical Ant-Based Clustering
abstract
The behavior and self-organization of ant colonies has been widely studied to address distributed clustering. However, most models that directly mimic ants produce too many clusters and converge too slowly. A wide range of research has attempted to address this through various means, but a number of sources of inefficiency remain, including: i) ants must physically move from one cluster to another through intermediate locations, ii) patterns in movement among clusters is not considered, and iii) while some approaches have included bulk item movement, they do not provide efficient movement while still maintaining the self-organizing nature of ant-based clustering. To address these issues, this paper presents a new algorithm for ant-based clustering. Here ants maintain a movement zone around each cluster, keeping ants close to data items. These movement zones are used to elect representatives that are responsible for all long distance movement. Representatives can, probabilistically, pass an object it has to any other representative. Since each cluster has approximately one representative at any given time, the search space for placing items over a long distance is reduce to the number of clusters. This provides an infrastructure that allows bulk movement and efficient long distance merging.
Jeremy B. Brown, Manfred Huber
SMC2
2009 Clustering Similar Actions in Sequential Decision Processes
abstract
The presence of a large number of available actions in the context of an automated, adaptive decision process can lead to an excessively large search space and thus significantly increase the overhead for the policy learning process. This issue occurs particularly in problem domains such as path planning or grid scheduling where the number of decision points is large and irreducible. The learning algorithm developed in this paper attempts to create a more compact representation of the state and action space by grouping similar actions that are likely leading to very similar future results. Actions are considered similar if they, with high probability, lead to future results with sufficient commonality. This paper develops this action clustering framework within the MDP formalism where actions in any given state are grouped if they result in similar reinforcement feedback based on the past learning experience. The resulting action sets are then considered as a whole in the decision process.
Po-Hsiang Chiu, Manfred Huber
ICMLA2
2009 Reactive grasping using optical proximity sensors
abstract
We propose a system for improving grasping using fingertip optical proximity sensors that allows us to perform online grasp adjustments to an initial grasp point without requiring premature object contact or regrasping strategies. We present novel optical proximity sensors that fit inside the fingertips of a Barrett Hand, and demonstrate their use alongside a probabilistic model for robustly combining sensor readings and a hierarchical reactive controller for improving grasps online. This system can be used to complement existing grasp planning algorithms, or be used in more interactive settings where a human indicates the location of objects. Finally, we perform a series of experiments using a Barrett hand equipped with our sensors to grasp a variety of common objects with mixed geometries and surface textures.
Kaijen Hsiao, Paul Nangeroni, Manfred Huber, Ashutosh Saxena, Andrew Y. Ng
ICRA3
2009 Encoding user motion preferences in harmonic function path planning
abstract
Humans have unique motion preferences when pursuing a given task. These motion preferences could be expressed as moving in a straight line, following the wall, avoiding sharp turns, avoiding damp surfaces or choosing the shortest path. While it would be very useful for a range of applications to allow robot systems or artificial agents to generate paths with similar specific characteristics, it is generally very difficult to capture and reproduce them from observed information since user trajectories can not be easily generalized. To address this, this paper introduces an approach that modifies a harmonic function path planner to model the user's motion preferences as parameters which could then be used to generate new paths in similar environments without the risk of collisions or incorrect paths. Given a small set of user-specific trajectories and starting from an initial, generic parameter configuration, this approach incrementally minimizes the difference between the direction of the user trajectory segments and the gradient of the parametric harmonic function by modifying its underlying parameters, thus capturing the trajectory preferences. Subsequently, these parameters could be transferred to new, locally similar environments and used to generate new paths. The use of harmonic function parameters to represent the user preferences not only facilitates customization of the path planner but also assures that the customized planner remains complete and correct.
Giles D'Silva, Manfred Huber
IROS2
2009 A Temporal Potential Function Approach For Path Planning in Dynamic Environments
abstract
A dynamic environment is one in which either the obstacles or the goal or both are in motion. In most of the current research, robots attempting to navigate in dynamic environments use reactive systems. Although reactive systems have the advantage of fast execution and low overheads, the tradeoff is in performance in terms of the path optimality. Often, the robot ends up tracking the goal, thus following the path taken by the goal, and deviates from this strategy only to avoid a collision with an obstacle it may encounter. In a path planner, the path from the start to the goal is calculated before the robot sets off. This path has to be recalculated if the goal or the obstacles change positions. In the case of a dynamic environment this happens often. One method to compensate for this is to take the velocity of the goal and obstacles into account when planning the path. So instead of following the goal, the robot can estimate where the best position to reach the goal is and plan a path to that location. In this paper, we propose a method for path planning in dynamic environments that uses a potential function which indicates the probability that a robot will collide with an obstacle, assuming that the robot executes a random walk from that location and that time onwards. The robot plans a path by extrapolating the object's motion using current velocities and by calculating the potential values up to a look-ahead limit that is determined by calculating the minimum path length using connectivity evaluation and then determining the utility of expanding the look-ahead limit beyond the minimum path length. This paper will discuss how the potential values are calculated and how a suitable look-ahead limit is decided. Finally the performance of the proposed method is demonstrated in a simulated environment.
Vamsikrishna Gopikrishna, Manfred Huber
SMC2
2009 Learning to Generalize and Reuse Skills Using Approximate Partial Policy Homomorphisms
abstract
A reinforcement learning (RL) agent that performs successfully in a complex and dynamic environment has to continuously learn and adapt to perform new tasks. This necessitates for them to not only extract control and representation knowledge from the tasks learned, but also to reuse the extracted knowledge to learn new tasks. This paper presents a new method to extract this control and representational knowledge. Here we present a policy generalization approach that uses the novel concept of policy homomorphism to derive these abstractions. The paper further extends the policy homomorphism framework to an approximate policy. The extension allows policy generalization framework to efficiently address more realistic tasks and environments in non-deterministic domains. The approximate policy homomorphism derives an abstract policy for a set of similar tasks (a task type) from a set of basic policies learned for previously seen task instances. The resulting generalized policy is then applied in new contexts to address new instances of related tasks. The approach also allows to identify similar tasks based on the functional characteristics of the corresponding skills and provides a means of transferring the learned knowledge to new situations without the need for complete knowledge of the state space and the system dynamics in the new environment. We demonstrate the working of policy abstraction using approximate policy homomorphism and illustrate policy reuse to learn new tasks in novel situations using a set of grid world examples.
Srividhya Rajendran, Manfred Huber
SMC2
2008 Learning task decomposition and exploration shaping for reinforcement learning agents
abstract
For situated reinforcement learning agents to succeed in complex real world environments they have to be able to efficiently acquire and reuse control knowledge in order to accomplish new tasks faster and to accelerate the learning of new policies. While hierarchical learning approaches which transfer previously acquired skills and representations to model and control new tasks have the potential to significantly improve learning times, they also pose the risk of ldquobehavior proliferationrdquo where the growing set of available actions makes it increasingly difficult to determine a strategy for a new task. To overcome this problem and to further improve knowledge reuse, the learning agent should thus also have the ability to predict the utility of an action or reusable skill in a new context and to analyze new tasks in order to decompose them into known subtasks. This paper presents a novel approach for learning task decomposition by learning to predict the utility of subgoals and subgoal types in the context of a new task, as well as for exploration shaping by predicting the likelihood with which each available action is useful in the given task context. This information, encoded as a set of utility functions, is then used to focus the exploration and learning process of the agent to increase performance both in terms of the time spent to reach the new task's goal the first time and of the time required to learn an optimal policy. This ability is demonstrated here in the context of navigation and manipulation tasks in a feature enhanced grid world domain.
Predrag Djurdjevic, Manfred Huber
SMC2
2008 An approach for behavior discovery using clustering of dynamics
abstract
As robots enter more complex application domains and start to interact autonomously with their surroundings and with humans, it becomes essential that they can efficiently represent and interpret their streams of sensor data and model the behavior of objects in their environment. To do this automatically and without the need for extensive a priori models of the environment, these systems have to be able to autonomously discover the different dynamic behaviors of entities in their environment as well as to predict points at which objects' behaviors and interactions change. This paper presents a technique to simultaneously learn to identify segmentation points and to extract a discrete set of models of the behaviors of objects from a stream of sensor observations. The approach presented here uses unsupervised learning techniques to learn these models in the form of representative sensor signatures which, once acquired, are used to interpret the robot's observations and to represent the actual sensor data more compactly as a sequence of behaviors. To enable the system to simultaneously learn to segment the continuous data stream and to build appropriate models for the data segments, a set of similarity metrics for sensor streams is derived and used in an expectation maximization algorithm. This algorithm alternates between segmenting the sensor data based on the existing behavior models and clustering the data segments to derive better models, thus iteratively improving both the quality of the segmentation and of the behavior models. To illustrate the approach, experiments are performed using simulated observations derived from different types of sensors which observe the dynamic interactions of objects.
Ashokkumaar P. Loganathan, Manfred Huber
SMC2
2007 Complexity and Error Propagation of Localization Using Interferometric Ranging
abstract
An interferometric ranging technique has been recently proposed as a possible way to localize ad hoc and sensor networks. Compared to the more common techniques such as received signal strength, time of arrival, and angle of arrival ranging, interferometric ranging has the advantage that the measurement could be highly precise. However, localization using interferometric ranging is difficult as it requires a large number of measurement readings. In this paper, we provide a formal proof of this difficulty in terms of algorithmic complexity. Furthermore, we propose an iterative algorithm that calculates node locations from a set of seeding anchors, gradually building a more global localization solution. Compared to previous localization algorithms, which treat localization as a global optimization problem, the iterative algorithm is a distributed algorithm that is simple to implement in larger networks. More importantly, the iterative algorithm allows us to study the error propagation behavior of localization using interferometric ranging. Using simulations, we validate the performance of the iterative algorithm in terms of localization error and coverage.
Gergely V. Záruba, Manfred Huber
ICC3
2007 Effective Control Knowledge Transfer through Learning Skill and Representation Hierarchies
Mehran Asadi, Manfred Huber
IJCAI2
2007 A particle filter approach for multi-target tracking
abstract
The problem of tracking multiple objects poses a number of challenges due to the ambiguity of the observations and the presence of partial or complete occlusions. This paper introduces a novel extension to the Particle Filter algorithm for tracking multiple objects with a vision system. The presented approach instantiates separate particle filters for each object and explicitly handles partial and complete occlusion for non- transparent objects, as well as the instantiation and removal of filters in case new objects enter the scene or previously tracked objects are removed. As opposed to single particle filters or mixture particle filter approaches which estimate a single multi-modal distribution, the proposed filter extension allows the continued tracking of objects through occlusion situations as well as the tracking of multiple objects of different types. To allow for the handling of occlusions without an increase in computational complexity beyond the one of the Mixture Particle Filter, the approach presented here addresses occlusions by projecting particles into the image space and back into the particle space, thus avoiding the use of a joint distribution. To present qualitative results, experiments were performed using color-based tracking of multiple objects of different and identical colors. The experiments demonstrate that the Particle filters implemented using the proposed method effectively and precisely track multiple targets and can successfully instantiate and remove filters of objects that enter or leave the image area.
Hwang Ryol Ryu, Manfred Huber
IROS2
2007 Indoor location tracking using RSSI readings from a single Wi-Fi access point
Gergely V. Záruba, Manfred Huber, Farhad Kamangar, Imrich Chlamtac
Wirel. Networks2
2006 Particle Filter Based Object Tracking in a Stereo Vision System
abstract
Tracking objects based on their visual features is of major importance for robotic systems performing autonomous tasks but also poses many challenges. This paper presents and compares two methods for tracking objects in a stereo camera system using particle filters which differ in the way they address the problem of stereo correspondence during the filtering process. In the first approach, two particle sets, one for each of the left and right stereo image frames are maintained and a mapping between the two sets is established by soft-stereo method, the particles are tracked in three-dimensional space and mapped back into the image frames to make the observations. The effectiveness of the approaches is demonstrated through experiments tracking a moving object and the two approaches are compared in terms of their convergence times and tracking errors
Anup S. Sabbi, Manfred Huber
ICRA2
2005 Autonomous Subgoal Discovery and Hierarchical Abstraction for Reinforcement Learning Using Monte Carlo Method
Mehran Asadi, Manfred Huber
AAAI2
2005 A Bayesian Sampling Approach to In-Door Localization of Wireless Devices Using Received Signal Strength Indication
abstract
This paper describes a probabilistic approach to global localization within an in-door environment with minimum infrastructure requirements. Global localization is a flavor of localization in which the device is unaware of its initial position and has to determine the same from scratch. Localization is performed based on the received signal strength indication (RSSI) as the only sensor reading, which is provided by most off-the-shelf wireless network interface cards. Location and orientation estimates are computed using Bayesian filtering on a sample set derived using Monte-Carlo sampling. Research leading to the proposed method is outlined along with results and conclusions from simulations and real life experiments.
Vinay Seshadri, Gergely V. Záruba, Manfred Huber
PerCom3
2005 Research to classroom: experiences from a multi-institutional course in smart home technologies
Charles J. Hannon, Manfred Huber, Lisa J. Burnell
SIGCSE2
2004 Monte Carlo sampling based in-home location tracking with minimal RF infrastructure requirements
abstract
The paper describes research towards a system for locating users in a home environment requiring only a minimal wireless infrastructure. The only sensor reading used for the location estimation is the radiofrequency received signal strength indication (RSSI) measured by an RF interface (e.g., Wi-Fi). Location estimates are computed using Bayesian filtering on sample sets derived by Monte Carlo sampling. Wireless signal strength maps for the filter are obtained by a two-step parametric and measurement driven ray-tracing approach to account for absorption and reflection characteristics of various obstacles. Our trace driven simulations indicate that RSSI readings from a single access point in an indoor environment are sufficient to derive good location estimates of users.
Gergely V. Záruba, Manfred Huber, Farhad A. Karnangar, Imrich Chlarntac
GLOBECOM2
2003 User-guided reinforcement learning of robot assistive tasks for an intelligent environment
abstract
Autonomous robots hold the possibility of performing a variety of assistive tasks in intelligent environments. However, widespread use of robot assistants in these environments requires ease of use by individuals who are generally not skilled robot operators. In this paper we present a method of training robots that bridges the gap between user programming of a robot and autonomous learning of a robot task. With our approach to variable autonomy, we integrate user commands at varying levels of abstraction into a reinforcement learner to permit faster policy acquisition. We illustrate the ideas using a robot assistant task, that of retrieving medicine for an inhabitant of a smart home.
Manfred Huber, Vinay N. Papudesi, Diane J. Cook
IROS2
2002 Robust finger gaits from closed-loop controllers
abstract
Object manipulation using multi-fingered dextrous robot hands poses formidable control challenges due to the complexity of the manipulator, its workspace limitations. To address large-scale manipulation tasks, finger gaits are often used to reposition fingers. In this paper, an approach to finger gaiting is presented which constructs behavior on-line by activating combinations of reusable feedback control laws with formal stability and convergence properties drawn from a control basis. Finger gaits are constructed as finite state control strategies in a DEDS framework, while actual contact locations and object motions are computed reactively based on local contact information. This leads to manipulation strategies that are robust with respect to a limited range of perturbations, object geometries, and manipulator kinematics. To demonstrate this, two finger gaits are constructed and applied to different object geometries and robot hands.
Manfred Huber, Roderic A. Grupen
IROS1
2000 A Hybrid Architecture for Hierarchical Reinforcement Learning
abstract
Autonomous robot systems operating in the real world have to be able to learn new tasks and environmental conditions without the need for an outside teacher. While reinforcement learning represents a good formalism to achieve this, its long learning times and need for extensive exploration often make it impracticable for online learning on complex systems. The hybrid architecture presented in this paper addresses this issue by applying reinforcement learning on top of an automatically derived abstract discrete event dynamic system (DEDS) supervisor. This reduces the problem of policy acquisition within this approach to learning to coordinate a set of closed-loop control strategies in order to perform a given task. Besides dramatically reducing the complexity of the learning task this framework also permits the incorporation of a priori knowledge and facilitates the inclusion of learned policies as actions in order to transfer skills to new task domains. To demonstrate the applicability of this approach, the architecture is used to learn locomotion gaits on a four-legged robot platform.
Manfred Huber
ICRA1
1997 Learning to Coordinate Controllers - Reinforcement Learning on a Control Basis
Manfred Huber, Roderic A. Grupen
IJCAI1
1996 A control basis for multilegged walking
abstract
This paper presents a distributed control approach to legged locomotion that constructs behavior online by activating combinations of reusable feedback control laws drawn from a control basis. Sequences of such controller activation result in flexible aperiodic step sequences based on local sensory information. Different tasks are achieved by varying the composition functions over the same basis controllers, rather than by geometric planning of leg placements or the design of new task-specific behaviors. In addition, the device-independent nature of the control basis allows its generalization not only over task domains, but also over different hardware platforms. To show the applicability of this approach, a control basis and two generic control gaits for four-legged walking are introduced and tested on an even terrain walking task in an unknown environment.
Manfred Huber, Willard S. MacDonald, Roderic A. Grupen
ICRA1
1994 2-D contact detection and localization using proprioceptive information
abstract
This paper employs proprioceptive information (joint angles and torques) to estimate properties of the contact between a planar robot and an unknown object without specifically requiring strategic manipulator motions. The algorithm presented tackles this task in two stages; a contact localization analysis is followed by a force domain contact detection analysis. In the former, the Cartesian endpoint velocities are used for each link to obtain an estimate of the location of a hypothetical contact point on the link surface. A second observer, based on the displacement between two consecutive postures of the manipulator, provides an estimate of the error associated with this location. This data is fused over time by tracking the contact location using a linear observer and results in a hypothetical contact location, an associated uncertainty region, and a surface normal estimate. The detection phase uses torque domain evidence and the location estimates to verify the existence of each of the contacts. This process allows the detection of one contact per link and provides estimates of contact location, velocity, surface normal, and contact force.>
Manfred Huber, Roderic A. Grupen
IEEE Trans. Robotics Autom.1