VLDB 2026 Research / reviewers in the wild / expert
Daniel Nikovski
dblp:75/2522 · also Daniel Nikolaev Nikovski
· DBLP profile ↗
42ranked-venue papers
14as first author
14since 2021 · last 2025
0000-0003-2919-645XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 32 · 9 first-author · 9 since 2021Systems, architecture and hardware · 9 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 6 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 6 · 4 first-author · 5 since 2021Software engineering, systems software and programming languages · 5 · 4 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 1 since 2021Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Disentangled Object-Centric Configuration Representation Learning for Articulated Robot ArmsabstractThe paper proposes a method for learning compact representations of the configuration (joint positions) of articulated mechanisms consisting of interconnected rigid bodies, from collected sequences of keypoint positions observed and tracked in camera images. The method analyzes the variations in pairwise distances between keypoints over time to deduce which of the keypoints must belong to the same rigid body and then computes the relative pose of all rigid bodies with respect to a reference image representing an initial or target configuration of the mechanism. By analyzing the rank of data matrices representing the translational and rotational components of the relative poses between the rigid bodies over time, the algorithm infers the order of the kinematic chain of the mechanism and the type of joints used in it, allowing the construction of a configuration vector as compact as the true joint positions of the mechanism. Daniel Nikovski |
CoDIT | 1 |
| 2025 | Observation-Based Inverse Kinematics for Visual Servo Control
Daniel Nikovski |
ICINCO (1) | 1 |
| 2024 | Adaptive Velocity Estimators for Learning ControlabstractThe paper proposes a method for learning velocity estimators, in the form of finite impulse response (FIR) filters, from data collected from a system equipped with quantizing position encoders that is to be controlled by means of a full-state feedback controller making use of the velocity estimates. The resulting estimators are tailored to the properties of the controlled system and show empirically superior performance in comparison with commonly used baseline velocity estimators, both in terms of velocity estimation error as well as in terms of reduced regulation cost when tested on control problems. The proposed adaptive estimators are resistant to overfitting the training data, are easy to implement on embedded controller devices, and can be used in conjunction with various learning control methods. Daniel Nikovski, William Yerazunis |
CoDIT | 1 |
| 2024 | Memory-Based Global Iterative Linear Quadratic ControlabstractWe propose a method for designing global nonlinear controllers based on the application of memory-based learning schemes for the purpose of aggregating multiple solutions produced by optimal control algorithms based on differential dynamic programming. The method leverages the fact that these optimal control algorithms produce not only nominal state and control trajectories, but entire full-state feedback (FSF) controllers, and the combined controller effectively switches between these multiple FSF controllers. Empirical verification demonstrates that it can be very effective in solving difficult benchmark control problems at high control rates. Daniel Nikovski, Junmin Zhong, William Yerazunis |
CoDIT | 1 |
| 2024 | Memory-Based Learning of Global Control Policies from Local Controllers
Daniel Nikovski, Junmin Zhong, William Yerazunis |
ICINCO (1) | 1 |
| 2024 | Learning Time-Optimal Control of Gantry CranesabstractThe paper presents an experimental study on the application of deep reinforcement learning (DRL) methods to the problem of optimally transporting cargo loads by an overhead gantry crane in minimal time. Experiments in simulation using a physics engine on two versions of the problem, with two and four degrees of freedom and employing reward functions that reflect the objective of load stabilization in minimal time, demonstrate that policies trained with the Stochastic Actor Critic (SAC) DRL method achieve up to 20% shorter transport time in comparison with controllers designed by means of more traditional methods from the field of control engineering. Junmin Zhong, Daniel Nikovski, William Yerazunis, Taishi Ando |
ICMLA | 2 |
| 2023 | Model-Based Learning Controller Design for a Furuta PendulumabstractWe present a method for designing and tuning controllers for the problem of swing-up and stabilization of a Furuta pendulum. The method is based on suitable parameterization of a family of controllers and the application of Bayesian optimization to their tuning with minimal interaction with the physical system. Unlike traditional controller design methodologies, the method does not require the derivation of an exact physical model of the controlled plant, thus saving significant design time and effort. Furthermore, the method has much more favorable sample complexity than most policy optimization methods proposed in the field of reinforcement learning. Daniel Nikovski, William Yerazunis, Abraham Goldsmith |
CoDIT | 1 |
| 2023 | Stochastic Learning Manipulation of Object Pose With Under-Actuated Impulse Generator ArraysabstractRobotic assembly systems are common in modern industry and a fixture of commerce. However, the robots themselves lack the adaptability of humans in terms of singulating and grasping parts with uncontrolled pose. To this end, vibratory bowl feeder (VBF) devices are often employed to pre-orient the part for robot grasping. Unfortunately, VBFs themselves are inflexible (usually bespoken for one specific part), noisy, and very expensive to design and tune. We consider an alternative to the VBF - an array of impulse-generating solenoids positioned under a semi-rigid part-carrying platform that uses computer vision and self-supervised machine learning to generate a policy implementing a closed-loop controller to orient randomly positioned parts into a pose acceptable for robot grasping. Using a flat square wooden nut from a child's assembly toy as a test object, we were able to flip the nut into the desired orientation (standing vertically on the narrow edge) 21.1% of the time with a single impulse, and 35.4% of the time with two impulses, versus just 10.2% and 19.2% (respectively) of the time for a baseline policy of random choice of solenoid position and impulse duration, thus demonstrating black-box control of a process commonly considered too difficult to physically model. Chuizheng Kong, William Yerazunis, Daniel Nikovski |
ICMLA | 3 |
| 2023 | Constrained Dynamic Movement Primitives for Collision Avoidance in Novel EnvironmentsabstractDynamic movement primitives are widely used for learning skills that can be demonstrated to a robot by a skilled human or controller. While their generalization capabilities and simple formulation make them very appealing to use, they possess no strong guarantees to satisfy operational safety constraints for a task. We present constrained dynamic movement primitives (CDMPs), which can allow for positional constraint satisfaction in the robot workspace. Our method solves a non-linear optimization to perturb an existing DMP's forcing weights to admit a Zeroing Barrier Function (ZBF), which certifies positional workspace constraint satisfaction. We demonstrate our approach under different positional constraints on the end-effector movement on multiple physical robots, such as obstacle avoidance and workspace limitations. Seiji Shaw, Devesh K. Jha, Arvind U. Raghunathan, Radu Corcodel, Diego Romeres, George Dimitri Konidaris, Daniel Nikovski |
IROS | 7 |
| 2022 | Transfer Learning for Bayesian Optimization with Principal Component AnalysisabstractBayesian Optimization has been widely used for black-box optimization. Especially in the field of machine learning, BO has obtained remarkable results in hyperparameters optimization. However, the best hyperparameters depend on the specific task and traditionally the BO algorithm needs to be repeated for each task. On the other hand, the relationship between hyperparameters and objectives has similar tendency among tasks. Therefore, transfer learning is an important technology to accelerate the optimization of novel task by leveraging the knowledge acquired in prior tasks. In this work, we propose a new transfer learning strategy for BO. We use information geometry based principal component analysis (PCA) to extract a low-dimension manifold from a set of Gaussian process (GP) posteriors that models the objective functions of the prior tasks. Then, the low dimensional parameters of this manifold can be optimized to adapt to a new task and set a prior distribution for the objective function of the novel task. Experiments on hyperparameters optimization benchmarks show that our proposed algorithm, called BO-PCA, accelerates the learning of an unseen task (less data are required) while having low computational cost. Hideyuki Masui, Diego Romeres, Daniel Nikovski |
ICMLA | 3 |
| 2022 | Deep Reinforcement Learning for Optimal Sailing UpwindabstractWe describe the application of deep reinforcement learning (DRL) methods to determine the optimal decision policy when sailing a sailboat towards a target point located upwind from the boat's current position, under the conditions of wind direction and speed that vary according to an unknown stochastic process, as is typical in real sailing races. A model of the dynamics of the sailboat is described together with a suitable choice of actions, in the form of a Markov decision process (MDP), which allows the application of a wide variety of DRL algorithms. Empirical results show that the learned policy outperforms baseline control algorithms that do not take into consideration the variability in wind strength and direction, and instead assume that the current wind conditions will persist indefinitely. Takumi Suda, Daniel Nikovski |
IJCNN | 2 |
| 2022 | Model-Based Policy Search Using Monte Carlo Gradient Estimation With Real Systems ApplicationabstractIn this article, we present a model-based reinforcement learning (MBRL) algorithm namedMonte Carlo probabilistic inference for learning control(MC-PILCO). This algorithm relies on Gaussian processes (GPs) to model the system dynamics and on a Monte Carlo approach to estimate the policy gradient. This defines a framework in which we ablate the choice of the components, which are the selection of the cost function, the optimization of policies using dropout, and an improved data efficiency through the use of structured kernels in the GP models. The combination of the aforementioned aspects affects dramatically the performance of MC-PILCO. Numerical comparisons in a simulated cart–pole environment show that MC-PILCO exhibits better data efficiency and control performance w.r.t. state-of-the-art GP-based MBRL algorithms. Finally, we apply MC-PILCO to real systems, considering, in particular, systems with partially measurable states. We discuss the importance of modeling both the measurement system and the state estimators during policy optimization. The effectiveness of the proposed solutions has been tested in simulation and on two real systems, which are the Furuta pendulum and the ball-and-plate rig. Fabio Amadio, Alberto Dalla Libera, Riccardo Antonello, Daniel Nikovski, Ruggero Carli, Diego Romeres |
IEEE Trans. Robotics | 4 |
| 2021 | Personalizing Individual Comfort in the Group SettingabstractMaintaining individual thermal comfort in indoor spaces shared by multiple occupants is difficult because it requires both intuition about the thermal properties of the room, as well as an understanding of the thermal comfort preferences of each individual. We explore an approach to optimizing individual thermal comfort within a group through temperature set-point optimization of HVAC equipment. We propose a weakly-supervised algorithm to learn the individual thermal comfort preferences and an autoencoding framework to learn static approximations of room thermodynamics. We further propose two approaches to learn a control law that sets the HVAC set-points subject to the preferred user temperatures. The proposed method is tested on a real data-set obtained from workers in an open office. The results show that, on average, the temperature in the room at each user's location can be regulated to within 0.5C of the user's desired temperature. Emil Laftchiev, Diego Romeres, Daniel Nikovski |
AAAI | 3 |
| 2021 | Tactile-RL for Insertion: Generalization to Objects of Unknown GeometryabstractObject insertion is a classic contact-rich manipulation task. The task remains challenging, especially when considering general objects of unknown geometry, which significantly limits the ability to understand the contact configuration between the object and the environment. We study the problem of aligning the object and environment with a tactile-based feedback insertion policy. The insertion process is modeled as an episodic policy that iterates between insertion attempts followed by pose corrections. We explore different mechanisms to learn such a policy based on Reinforcement Learning. The key contribution of this paper is to demonstrate that it is possible to learn a tactile insertion policy that generalizes across different object geometries, and an ablation study of the key design choices for the learning agent: 1) the type of learning scheme: supervised vs. reinforcement learning; 2) the type of learning schedule: unguided vs. curriculum learning ; 3) the type of sensing modality: force/torque vs. tactile; and 4) the type of tactile representation: tactile RGB vs. tactile flow. We show that the optimal configuration of the learning agent (RL + curriculum + tactile flow) exposed to 4 training objects yields an closed-loop insertion policy that inserts 4 novel objects with over 85.0% success rate and within 3~4 consecutive attempts. Comparisons between F/T and tactile sensing, shows that while an F/T-based policy learns more efficiently, a tactile-based policy provides better generalization. See supplementary video and results at https://sites.google.com/view/tactileinsertion. Siyuan Dong, Devesh K. Jha, Diego Romeres, Sangwoon Kim, Daniel Nikovski, Alberto Rodriguez 0003 |
ICRA | 5 |
| 2020 | The Missing Input ProblemabstractRapid advances in information and communications technologies (ICT) have made it possible to deploy large collections of sensors to be used in traditional Supervisory Control and Data Acquisition (SCADA) systems and modern Internet of Things (IoT) installations. These sensors are intended for use in analytical formulas, AI algorithms, and traditional rule- based monitoring that determine optimal operation parameters, maintain smooth operation, or detect operation anomalies. Yet, advances in ICT do not always improve the reliability of data collection. Instead, the frequent use of consumer-grade sensors in IoT deployments, and an increasing array of customer choices often lead to inaccessibility of sensors or systematic absence of sensor readings. The lack of reliability in data collection leads to failures in the algorithms responsible for monitoring the system operation. We term this "the missing input problem", and discuss several state-of-the-art solutions. We specifically focus on the straightforward approach using standard imputation methods, as well as recent deep learning imputation methods. We show that none of the existing algorithms today perform very well in the face of missing sensors, and we outline several research directions that can lead to improvements. Emil Laftchiev, Daniel Nikovski |
IEEE BigData | 3 |
| 2020 | Can Increasing Input Dimensionality Improve Deep Reinforcement Learning?abstractDeep reinforcement learning (RL) algorithms have recently achieved remarkable successes in various sequential decision making tasks, leveraging advances in methods for training large deep networks. However, these methods usually require large amounts of training data, which is often a big problem for real-world applications. One natural question to ask is whether learning good representations for states and using larger networks helps in learning better policies. In this paper, we try to study if increasing input dimensionality helps improve performance and sample efficiency of model-free deep RL algorithms. To do so, we propose an online feature extractor network (OFENet) that uses neural nets to produce \emph{good} representations to be used as inputs to an off-policy RL algorithm. Even though the high dimensionality of input is usually thought to make learning of RL agents more difficult, we show that the RL agents in fact learn more efficiently with the high-dimensional representation than with the lower-dimensional state observations. We believe that stronger feature propagation together with larger networks allows RL agents to learn more complex functions of states and thus improves the sample efficiency. Through numerical experiments, we show that the proposed method achieves much higher sample efficiency and better performance. Codes for the proposed method are available at http://www.merl.com/research/license/OFENet Kei Ota, Tomoaki Oiki, Devesh K. Jha, Toshisada Mariyama, Daniel Nikovski |
ICML | 5 |
| 2020 | Local Policy Optimization for Trajectory-Centric Reinforcement LearningabstractThe goal of this paper is to present a method for simultaneous trajectory and local stabilizing policy optimization to generate local policies for trajectory-centric model-based reinforcement learning (MBRL). This is motivated by the fact that global policy optimization for non-linear systems could be a very challenging problem both algorithmically and numerically. However, a lot of robotic manipulation tasks are trajectory-centric, and thus do not require a global model or policy. Due to inaccuracies in the learned model estimates, an open-loop trajectory optimization process mostly results in very poor performance when used on the real system. Motivated by these problems, we try to formulate the problem of trajectory optimization and local policy synthesis as a single optimization problem. It is then solved simultaneously as an instance of nonlinear programming. We provide some results for analysis as well as achieved performance of the proposed technique under some simplifying assumptions. Patrik Kolaric, Devesh K. Jha, Arvind U. Raghunathan, Frank L. Lewis, Mouhacine Benosman, Diego Romeres, Daniel Nikovski |
ICRA | 7 |
| 2020 | Personalized Destination Prediction Using Transformers in a Contextless Data SettingabstractDestination prediction is an important task where the primary goal is to correctly predict a user's destination given an input movement trajectory. Intelligent machine learning models that learn from observed movement data and can automatically forecast destinations from partial query trajectories are of high interest as they can provide a plethora of benefits to both creators and consumers in various markets. In this work, we present a novel framework for tackling the problem of destination prediction in a contextless data setting where we solely learn from trajectory coordinate information. We propose a Transformer model to predict destinations from partial trajectories and we demonstrate its use on two datasets from different domains, including a simulated indoor dataset and an outdoor taxi trajectory dataset. Our proposed method improves upon the previous state-of-the-art LSTM and BiLSTM deep learning approaches in terms of accuracy and distance from true destinations. Athanasios Tsiligkaridis, Jing Zhang 0030, Hiroshi Taguchi, Daniel Nikovski |
IJCNN | 4 |
| 2019 | Sim-to-Real Transfer Learning using Robustified Controllers in Robotic Tasks involving Complex DynamicsabstractLearning robot tasks or controllers using deep reinforcement learning has been proven effective in simulations. Learning in simulation has several advantages. For example, one can fully control the simulated environment, including halting motions while performing computations. Another advantage when robots are involved, is that the amount of time a robot is occupied learning a task-rather than being productive-can be reduced by transferring the learned task to the real robot. Transfer learning requires some amount of fine-tuning on the real robot. For tasks which involve complex (non-linear) dynamics, the fine-tuning itself may take a substantial amount of time. In order to reduce the amount of fine-tuning we propose to learn robustified controllers in simulation. Robustified controllers are learned by exploiting the ability to change simulation parameters (both appearance and dynamics) for successive training episodes. An additional benefit for this approach is that it alleviates the precise determination of physics parameters for the simulator, which is a non-trivial task. We demonstrate our proposed approach on a real setup in which a robot aims to solve a maze game, which involves complex dynamics due to static friction and potentially large accelerations. We show that the amount of fine-tuning in transfer learning for a robustified controller is substantially reduced compared to a non-robustified controller. Jeroen van Baar, Alan Sullivan, Radu Cordorel, Devesh K. Jha, Diego Romeres, Daniel Nikovski |
ICRA | 6 |
| 2019 | Semiparametrical Gaussian Processes Learning of Forward Dynamical Models for Navigating in a Circular MazeabstractThis paper presents a problem of model learning for the purpose of learning how to navigate a ball to a goal state in a circular maze environment with two degrees of freedom. The motion of the ball in the maze environment is influenced by several non-linear effects such as dry friction and contacts, which are difficult to model physically. We propose a semiparametric model to estimate the motion dynamics of the ball based on Gaussian Process Regression equipped with basis functions obtained from physics first principles. The accuracy of this semiparametric model is shown not only in estimation but also in prediction at n-steps ahead and its compared with standard algorithms for model learning. The learned model is then used in a trajectory optimization algorithm to compute ball trajectories. We propose the system presented in the paper as a benchmark problem for reinforcement and robot learning, for its interesting and challenging dynamics and its relative ease of reproducibility. Diego Romeres, Devesh K. Jha, Alberto Dalla Libera, William Yerazunis, Daniel Nikovski |
ICRA | 5 |
| 2019 | Trajectory Optimization for Unknown Constrained Systems using Reinforcement LearningabstractIn this paper, we propose a reinforcement learning-based algorithm for trajectory optimization for constrained dynamical systems. This problem is motivated by the fact that for most robotic systems, the dynamics may not always be known. Generating smooth, dynamically feasible trajectories could be difficult for such systems. Using sampling-based algorithms for motion planning may result in trajectories that are prone to undesirable control jumps. However, they can usually provide a good reference trajectory which a model-free reinforcement learning algorithm can then exploit by limiting the search domain and quickly finding a dynamically smooth trajectory. We use this idea to train a reinforcement learning agent to learn a dynamically smooth trajectory in a curriculum learning setting. Furthermore, for generalization, we parameterize the policies with goal locations, so that the agent can be trained for multiple goals simultaneously. We show result in both simulated environments as well as real experiments, for a 6-DoF manipulator arm operated in position-controlled mode to validate the proposed idea. We compare the proposed ideas against a PID controller which is used to track a designed trajectory in configuration space. Our experiments show that our RL agent trained with a reference path outperformed a model-free PID controller of the type commonly used on many robotic platforms for trajectory tracking. Kei Ota, Devesh K. Jha, Tomoaki Oiki, Mamoru Miura, Takashi Nammoto, Daniel Nikovski, Toshisada Mariyama |
IROS | 6 |
| 2019 | Introducing time series chains: a new primitive for time series data mining
Yan Zhu 0014, Makoto Imamura, Daniel Nikovski, Eamonn J. Keogh |
Knowl. Inf. Syst. | 3 |
| 2018 | Reinforcement Learning with Function-Valued Action Spaces for Partial Differential Equation ControlabstractRecent work has shown that reinforcement learning (RL) is a promising approach to control dynamical systems described by partial differential equations (PDE). This paper shows how to use RL to tackle more general PDE control problems that have continuous high-dimensional action spaces with spatial relationship among action dimensions. In particular, we propose the concept of action descriptors, which encode regularities among spatially-extended action dimensions and enable the agent to control high-dimensional action PDEs. We provide theoretical evidence suggesting that this approach can be more sample efficient compared to a conventional approach that treats each action dimension separately and does not explicitly exploit the spatial regularity of the action space. The action descriptor approach is then used within the deep deterministic policy gradient algorithm. Experiments on two PDE control problems, with up to 256-dimensional continuous actions, show the advantage of the proposed approach over the conventional one. Yangchen Pan, Amir-massoud Farahmand, Martha White, Saleh Nabi, Piyush Grover, Daniel Nikovski |
ICML | 6 |
| 2018 | Time Series Chains: A Novel Tool for Time Series Data MiningabstractSince their introduction over a decade ago, time se-ries motifs have become a fundamental tool for time series analytics, finding diverse uses in dozens of domains. In this work we introduce Time Series Chains, which are related to, but distinct from, time series motifs. Informally, time series chains are a temporally ordered set of subsequence patterns, such that each pattern is similar to the pattern that preceded it, but the first and last patterns are arbi-trarily dissimilar. In the discrete space, this is simi-lar to extracting the text chain “hit, hot, dot, dog” from a paragraph. The first and last words have nothing in common, yet they are connected by a chain of words with a small mutual difference. Time Series Chains can capture the evolution of systems, and help predict the future. As such, they potentially have implications for prognostics. In this work, we introduce a robust definition of time series chains, and a scalable algorithm that allows us to discover them in massive datasets. Yan Zhu 0014, Makoto Imamura, Daniel Nikovski, Eamonn J. Keogh |
IJCAI | 3 |
| 2018 | Anomaly Detection in Discrete Manufacturing Systems using Event Relationship Tables
Emil Laftchiev, Xinmaio Sun, Hoang Anh Dau, Daniel Nikovski |
DX | 4 |
| 2017 | Value-Aware Loss Function for Model-based Reinforcement LearningabstractWe consider the problem of estimating the transition probability kernel to be used by a model-based reinforcement learning (RL) algorithm. We argue that estimating a generative model that minimizes a probabilistic loss, such as the log-loss, is an overkill because it does not take into account the underlying structure of decision problem and the RL algorithm that intends to solve it. We introduce a loss function that takes the structure of the value function into account. We provide a finite-sample upper bound for the loss function showing the dependence of the error on model approximation error, number of samples, and the complexity of the model space. We also empirically compare the method with the maximum likelihood estimator on a simple problem. Amir-massoud Farahmand, André Barreto 0001, Daniel Nikovski |
AISTATS | 3 |
| 2017 | Matrix Profile VII: Time Series Chains: A New Primitive for Time Series Data Mining (Best Student Paper Award)abstractSince their introduction over a decade ago, time series motifs have become a fundamental tool for time series analytics, finding diverse uses in dozens of domains. In this work we introduce Time Series Chains, which are related to, but distinct from, time series motifs. Informally, time series chains are a temporally ordered set of subsequence patterns, such that each pattern is similar to the pattern that preceded it, but the first and last patterns are arbitrarily dissimilar. In the discrete space, this is similar to extracting the text chain "hit, hot, dot, dog" from a paragraph. The first and last words have nothing in common, yet they are connected by a chain of words with a small mutual difference. Time series chains can capture the evolution of systems, and help predict the future. As such, they potentially have implications for prognostics. In this work, we introduce a robust definition of time series chains, and a scalable algorithm that allows us to discover them in massive datasets. Yan Zhu 0014, Makoto Imamura, Daniel Nikovski, Eamonn J. Keogh |
ICDM | 3 |
| 2017 | Random Projection Filter Bank for Time Series DataabstractWe propose Random Projection Filter Bank (RPFB) as a generic and simple approach to extract features from time series data. RPFB is a set of randomly generated stable autoregressive filters that are convolved with the input time series to generate the features. These features can be used by any conventional machine learning algorithm for solving tasks such as time series prediction, classification with time series data, etc. Different filters in RPFB extract different aspects of the time series, and together they provide a reasonably good summary of the time series. RPFB is easy to implement, fast to compute, and parallelizable. We provide an error upper bound indicating that RPFB provides a reasonable approximation to a class of dynamical systems. The empirical results in a series of synthetic and real-world problems show that RPFB is an effective method to extract features from time series. Amir-massoud Farahmand, Sepideh Pourazarm, Daniel Nikovski |
NIPS | 3 |
| 2016 | Truncated Approximate Dynamic Programming with Task-Dependent Terminal ValueabstractWe propose a new class of computationally fast algorithms to find close to optimal policy for Markov Decision Processes (MDP) with large finite horizon T.The main idea is that instead of planning until the time horizon T, we plan only up to a truncated horizon H << T and use an estimate of the true optimal value function as the terminal value. Our approach of finding the terminal value function is to learn a mapping from an MDP to its value function by solving many similar MDPs during a training phase and fit a regression estimator. We analyze the method by providing an error propagation theorem that shows the effect of various sources of errors to the quality of the solution. We also empirically validate this approach in a real-world application of designing an energy management system for Hybrid Electric Vehicles with promising results. Amir-massoud Farahmand, Daniel Nikovski, Yuji Igarashi, Hiroki Konaka |
AAAI | 2 |
| 2016 | Regularized covariance matrix estimation with high dimensional data for supervised anomaly detection problemsabstractWe address the problem of estimating high-dimensional covariance matrices (CM) for the explicit purpose of supervised anomaly detection, in the case when the number n of data points is lower than their dimensionality p. This is increasingly common with the emergence of the Internet of Things that makes it possible to collect data from many sensors simultaneously, resulting in very high-dimensional data points. When we attempt to perform anomaly detection for such data by modeling the normal behavior of the system by means of a multivariate Gaussian distribution, and n <; p, the sample CM is singular, and cannot be used directly without some form of regularization. In contrast to existing methods for CM regularization that aim to fit the training data accurately, we propose a regularization algorithm for CM estimation that directly aims to maximize the area under the resulting receiver-operator characteristic (AUROC) for the ultimate decision problem that needs to be solved: anomaly detection. Experiments on test problems demonstrate the ability of the proposed algorithm to find CM estimates significantly better at anomaly detection than existing estimation methods that are unaware of the decision task that the CMs they produce will be used in. Daniel Nikovski, Kiran Byadarhaly |
IJCNN | 1 |
| 2016 | Exemplar learning for extremely efficient anomaly detection in real-valued time series
Michael J. Jones 0001, Daniel Nikovski, Makoto Imamura, Takahisa Hirata |
Data Min. Knowl. Discov. | 2 |
| 2010 | Fast adaptive algorithms for abrupt change detection
Daniel Nikovski |
Mach. Learn. | 1 |
| 2008 | Incremental Exemplar Learning Schemes for Classification on Embedded Devices
Daniel Nikovski |
ECML/PKDD (1) | 2 |
| 2008 | Incremental exemplar learning schemes for classification on embedded devices
Daniel Nikovski |
Mach. Learn. | 2 |
| 2004 | Optimal Parking in Group Elevator ControlabstractWe consider the problem of optimally parking empty cars in an elevator group so as to anticipate and intercept the arrival of new passengers and minimize their waiting times. Two solutions are proposed, for the down-peak and up-peak traffic patterns. We demonstrate that matching the distribution of free cars to the arrival distribution of passengers is sufficient to produce savings of up to 80% in down-peak traffic. Since this approach Is not useful for the much harder case of up-peak traffic, we propose a solution based on the representation of the elevator system as a Markov decision process (MDP) model with relatively few aggregated states, and determination of the optimal parking policy by means of dynamic programming on the MDP model. Matthew Brand, Daniel Nikovski |
ICRA | 2 |
| 2004 | Theory and Applied Computing: Observations and Anecdotes
Matthew Brand, Sarah F. Frisken, Neal Lesh, Joe Marks, Daniel Nikovski, Ronald N. Perry, Jonathan S. Yedidia |
MFCS | 5 |
| 2003 | Marginalizing Out Future Passengers in Group Elevator Control
Daniel Nikovski, Matthew Brand |
UAI | 1 |
| 2002 | Learning probabilistic models for state tracking of mobile robotsabstractWe propose a learning algorithm for acquiring a stochastic model of the behavior of a mobile robot, which allows the robot to localize itself along the outer boundary of its environment while traversing it. Compared to previously suggested solutions based on learning self-organizing neural nets, our approach achieves much higher spatial resolution which is limited only by the control time-step of the robot. We demonstrate the successful work of the algorithm on a small robot with only three infrared range sensors and a digital compass, and suggest how this algorithm can be extended to learn probabilistic models for full decision-theoretic reasoning and planning. Daniel Nikovski, Illah R. Nourbakhsh |
IROS | 1 |
| 2002 | Learning probabilistic models for optimal visual servo control of dynamic manipulationabstractWe present an experiment in sequential visual servo control of a dynamic manipulation task with unknown equations of motion and feedback from an uncalibrated camera. Our algorithm constructs a model of a Markov decision process (MDP) by means of grounding states in observed trajectories, and uses the model to find a control policy based on visual input, which maximizes a prespecified optimal control criterion balancing performance and control effort. Daniel Nikovski, Illah R. Nourbakhsh |
IROS | 1 |
| 2000 | Learning Probabilistic Models for Decision-Theoretic Navigation of Mobile Robots
Daniel Nikovski, Illah R. Nourbakhsh |
ICML | 1 |
| 2000 | Constructing Bayesian Networks for Medical Diagnosis from Incomplete and Partially Correct StatisticsabstractThe paper discusses several knowledge engineering techniques for the construction of Bayesian networks for medical diagnostics when the available numerical probabilistic information is incomplete or partially correct. This situation occurs often when epidemiological studies publish only indirect statistics and when significant unmodeled conditional dependence exists in the problem domain. While nothing can replace precise and complete probabilistic information, still a useful diagnostic system can be built with imperfect data by introducing domain-dependent constraints. We propose a solution to the problem of determining the combined influences of several diseases on a single test result from specificity and sensitivity data for individual diseases. We also demonstrate two techniques for dealing with unmodeled conditional dependencies in a diagnostic network. These techniques are discussed in the context of an effort to design a portable device for cardiac diagnosis and monitoring from multimodal signals. Daniel Nikovski |
IEEE Trans. Knowl. Data Eng. | 1 |
| 1996 | Comparison of Two Learning Networks for Time Series Prediction
Daniel Nikovski, Mehdi Zargham |
IEA/AIE | 1 |