R. Bhushan Gopaluni

dblp:41/1931 · DBLP profile ↗
← Back
9ranked-venue papers
0as first author
5since 2021 · last 2026
0000-0002-4321-0468ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1Graphics, computer vision, multimedia, augmented reality and games · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Motion planning and robot control · 75% Deep learning architectures and training · 12% Reinforcement learning · 12%

Topics — the 5 heaviest of 5, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Robotics › Motion planning and robot control › robot control › optimal control
linear quadratic regulator
0.912025
DiLQR: Differentiable Iterative Linear Quadratic Regulator via Implicit Differentiation · ICML 2025
Robotics › Motion planning and robot control
robot learning
0.912025
DiLQR: Differentiable Iterative Linear Quadratic Regulator via Implicit Differentiation · ICML 2025
Robotics › Motion planning and robot control
trajectory optimization
0.912025
DiLQR: Differentiable Iterative Linear Quadratic Regulator via Implicit Differentiation · ICML 2025
Machine learning › Deep learning architectures and training › sequence modeling
deep dynamics model
0.412020
Almost Surely Stable Deep Dynamics · NeurIPS 2020
Machine learning › Reinforcement learning
model-based reinforcement learning
0.412020
Almost Surely Stable Deep Dynamics · NeurIPS 2020

Methods — techniques the papers use, named apart from their topics

iterative LQR · 0.9implicit differentiation · 0.9lyapunov neural network · 0.4implicit output layer · 0.4convex lyapunov function · 0.4
YearPublicationVenuePosition
2026 Integrating transfer learning into multi-agent reinforcement learning for energy-efficient adaptive control of heating systems in thermoforming
abstract
This study proposes a framework that integrates Transfer Learning (TL) into Multi-Agent Reinforcement Learning (MARL) for real-time, energy-efficient thermal control of the heating stage in thermoforming processes, while concurrently minimizing energy consumption and adapting to varying manufacturing conditions. To enable a scalable framework under these conditions, a Fully Connected Neural Network-based Multi-Agent Proximal Policy Optimization (FCNN-MAPPO) architecture is developed using a multi-objective reward function for concurrently minimizing control error, thermal energy consumption, and instability, while preserving temporal dynamics through state augmentation. The resulting Multi-Agent Transfer Reinforcement Learning (MATRL) framework combines direct parameter transfer for within-scenario learning with experience sharing for cross-scenario adaptation, enabling faster convergence and improved generalization. Results show that under high convective heat transfer conditions, MATRL can reduce training time by 29.3%, decrease energy consumption by 28.1%, and improve average error by 2.02 °C compared to baseline MARL (i.e., without TL). Under elevated ambient temperature conditions, energy usage and settling time were reduced by 42.6% and 41.7%, respectively. Under synchronous multi-parameter variations (extreme conditions of convective heat transfer, sheet conductivity, and ambient temperature), MATRL maintained energy efficiency within 0.5% of nominal conditions while reducing temperature dispersion by 47%, demonstrating robust multidimensional adaptability without retraining. Statistical validation across multiple runs with random seeds showed stable performance, with a coefficient of variation of 7.3% and no divergence.
Iman Jalilvand, Amir M. Soufi Enayati, Hadi Hosseinionari, Rudolf J. Seethaler, Apurva Narayan, R. Bhushan Gopaluni, Abbas S. Milani
Eng. Appl. Artif. Intell.6
2025 DiLQR: Differentiable Iterative Linear Quadratic Regulator via Implicit Differentiation
abstract
While differentiable control has emerged as a powerful paradigm combining model-free flexibility with model-based efficiency, the iterative Linear Quadratic Regulator (iLQR) remains underexplored as a differentiable component. The scalability of differentiating through extended iterations and horizons poses significant challenges, hindering iLQR from being an effective differentiable controller. This paper introduces DiLQR, a framework that facilitates differentiation through iLQR, allowing it to serve as a trainable and differentiable module, either as or within a neural network. A novel aspect of this framework is the analytical solution that it provides for the gradient of an iLQR controller through implicit differentiation, which ensures a constant backward cost regardless of iteration, while producing an accurate gradient. We evaluate our framework on imitation tasks on famous control benchmarks. Our analytical method demonstrates superior computational performance, achieving up to $\textbf{128x}$ speedup and a minimum of $\textbf{21x}$ speedup compared to automatic differentiation. Our method also demonstrates superior learning performance ($\mathbf{10^6x}$) compared to traditional neural network policies and better model loss with differentiable controllers that lack exact analytical gradients. Furthermore, we integrate our module into a larger network with visual inputs to demonstrate the capacity of our method for high-dimensional, fully end-to-end tasks. Codes can be found on the project homepage https://sites.google.com/view/dilqr/.
Shuyuan Wang, Philip D. Loewen, Michael G. Forbes, R. Bhushan Gopaluni
ICML4
2025 Time series representation learning via cross-domain predictive and contextual contrasting: Application to fault detection
abstract
Data-driven methods for fault detection increasingly rely on large historical datasets, yet annotations are costly and time-consuming. As a result, learning approaches that minimize the need for extensive labeling, such as self-supervised learning (SSL), are becoming more popular. Contrastive learning , a subset of SSL, has shown promise in fields like computer vision and natural language processing (NLP), yet its application in fault detection is not fully explored. In this paper, we introduce Cross-Domain Predictive and Contextual Contrasting (CDPCC), a novel contrastive learning framework that integrates temporal and spectral information to capture informative time-frequency features from time series data. CDPCC consists of two key components: cross-domain predictive contrasting, which predicts future embeddings across time and frequency domains, and cross-domain contextual contrasting, which aligns time- and frequency-based representations in a shared latent space. We evaluate CDPCC on fault detection tasks using both simulated and industrial datasets. Our results show that a linear classifier trained on features learned by CDPCC performs comparably to fully supervised models. Moreover, CDPCC proves highly effective in scenarios with limited labeled data , achieving superior performance with only 50% of the labeled data compared to fully supervised training on the entire dataset. The source code is publicly available at https://github.com/iy641/CDPCC.git .
Ibrahim Yousef, Sirish L. Shah, R. Bhushan Gopaluni
Eng. Appl. Artif. Intell.3
2024 Guiding Reinforcement Learning with Incomplete System Dynamics
abstract
Model-free reinforcement learning (RL) is inherently a reactive method, operating under the assumption that it starts with no prior knowledge of the system and entirely depends on trial-and-error for learning. This approach faces several challenges, such as poor sample efficiency, generalization, and the need for well-designed reward functions to guide learning effectively. On the other hand, controllers based on complete system dynamics do not require data. This paper addresses the intermediate situation where there is not enough model information for complete controller design, but there is enough to suggest that a model-free approach is not the best approach either. By carefully decoupling known and unknown information about the system dynamics, we obtain an embedded controller guided by our partial model and thus improve the learning efficiency of an RL-enhanced approach. A modular design allows us to deploy mainstream RL algorithms to refine the policy. Simulation results show that our method significantly improves sample efficiency compared with standard RL methods on continuous control tasks, and also offers enhanced performance over traditional control approaches. Experiments on a real ground vehicle also validate the performance of our method, including generalization and robustness.
Shuyuan Wang, Jingliang Duan, Nathan P. Lawrence, Philip D. Loewen, Michael G. Forbes, R. Bhushan Gopaluni, Lixian Zhang 0001
IROS6
2023 Process Monitoring Using Domain-Adversarial Probabilistic Principal Component Analysis: A Transfer Learning Framework
abstract
Probabilistic principal component analysis (PPCA) is a feature extraction method that has been widely used in the field of process monitoring. However, PPCA assumes that training and testing data are drawn from the same input feature space with the same distributions. This assumption is not valid for complex processes that exhibit multiple operating modes and generate data with different distributions. In this article, we propose a novel transfer learning approach to monitoring processes with data from multiple distributions. To this end, we introduce a novel extension of PPCA, which is we refer to as the domain adversarial probabilistic principal component analysis (DAPPCA). DAPPCA algorithm automatically learns feature representations that are relevant across different operational modes. The algorithm extracts the most informative shared fault features and improves the accuracy of the fault detection model in a new operating mode using the knowledge transferred from previously known modes. The parameters of DAPPCA are estimated using a variational inference approach, and the monitoring statistics are calculated using the proposed model. We demonstrate the efficacy and real-time applicability of the proposed method with simulated and industrial examples.
Atefeh Daemi, R. Bhushan Gopaluni, Biao Huang 0001
IEEE Trans. Ind. Informatics2
2020 Almost Surely Stable Deep Dynamics
abstract
We introduce a method for learning provably stable deep neural network based dynamic models from observed data. Specifically, we consider discrete-time stochastic dynamic models, as they are of particular interest in practical applications such as estimation and control. However, these aspects exacerbate the challenge of guaranteeing stability. Our method works by embedding a Lyapunov neural network into the dynamic model, thereby inherently satisfying the stability criterion. To this end, we propose two approaches and apply them in both the deterministic and stochastic settings: one exploits convexity of the Lyapunov function, while the other enforces stability through an implicit output layer. We demonstrate the utility of each approach through numerical examples.
Nathan P. Lawrence, Philip D. Loewen, Michael G. Forbes, Johan U. Backström, R. Bhushan Gopaluni
NeurIPS5
2020 Deep Learning of Complex Batch Process Data and Its Application on Quality Prediction
abstract
Batch process quality prediction is an important application in manufacturing and chemical industries. The complexity of batch processes is characterized by multiphase, nonlinearity, dynamics, and uneven durations so that modeling of these batch processes is rather difficult. Moreover, there are other challenges in the face of quality prediction. Specifically, the process trajectories over the whole running duration potentially make specific contributions to the final targets so that the prediction issue embraces tremendously high-dimensional inputs but very low-dimensional outputs. This means that the prediction suffers from a severe dimensional imbalance between inputs and outputs. Motivated by these difficulties, this paper proposes a new deep learning-based framework for complex feature representative and quality prediction. Long short-term memory (LSTM) is used to extract comprehensive quality-relevant hidden features from a long-time sequence in each phase, significantly reducing the predictor dimensions. And these features from different phases are further integrated and compressed by a stacked auto-encoder (SAE). A practical industrial example testifies to the efficacy of the proposed framework.
Kai Wang 0024, R. Bhushan Gopaluni, Junghui Chen
IEEE Trans. Ind. Informatics2
2018 A switching strategy for adaptive state estimation
Aditya Tulsyan, Swanand R. Khare, Biao Huang 0001, R. Bhushan Gopaluni, J. Fraser Forbes
Signal Process.4
2008 Adaptive signal processing of asset price dynamics with predictability analysis
Rogemar S. Mamon, Christina Erlwein-Sayer, R. Bhushan Gopaluni
Inf. Sci.3