Harshad Khadilkar

dblp:144/2689 · also Harshad D. Khadilkar · DBLP profile ↗
← Back
11ranked-venue papers
1as first author
8since 2021 · last 2025
0000-0003-3601-778XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 8 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Efficiency Boost in Decentralized Optimization: Reimagining Neighborhood Aggregation with Minimal Overhead
abstract
In today's data-sensitive landscape, distributed learning emerges as a vital tool, not only fortifying privacy measures but also streamlining computational operations. This becomes especially crucial within fully decentralized infrastructures where local processing is imperative due to the absence of centralized aggregation. Here, we introduce DYNAWEIGHT, a novel framework to information aggregation in multi-agent networks. DYNAWEIGHT offers substantial acceleration in decentralized learning with minimal additional communication and memory overhead. Unlike traditional static weight assignments, such as Metropolis weights, DYNAWEIGHT dynamically allocates weights to neighboring servers based on their relative losses on local datasets. Consequently, it favors servers possessing diverse information, particularly in scenarios of substantial data heterogeneity. Our experiments on various datasets MNIST, CIFAR10, and CIFAR100 incorporating various server counts and graph topologies, demonstrate notable enhancements in training speeds. Notably, DYNAWEIGHT functions as an aggregation scheme compatible with any underlying server-level optimization algorithm, underscoring its versatility and potential for widespread integration.
Durgesh Kalwar, Mayank Baranwal, Harshad Khadilkar
CIKM3
2024 Linear-Time Optimal Deadlock Detection for Efficient Scheduling in Multi-Track Railway Networks
Hastyn Doshi, Ayush Tripathi, Keshav Agarwal, Harshad Khadilkar, Shivaram Kalyanakrishnan
IJCAI4
2023 Reinforcement Replaces Supervision: Query focused Summarization using Deep Reinforcement Learning
abstract
Query-focused Summarization (QfS) deals with systems that generate summaries from document(s) based on a query.Motivated by the insight that Reinforcement Learning (RL) provides a generalization to Supervised Learning (SL) for Natural Language Generation, and thereby performs better (empirically) than SL, we use an RL-based approach for this task of QfS.Additionally, we also resolve the conflict of employing RL in Transformers with Teacher Forcing.We develop multiple Policy Gradient networks, trained on various reward signals: ROUGE, BLEU, and Semantic Similarity, which lead to a 10 -point improvement over the State-of-the-Art approach on the ROUGE-L metric for a benchmark dataset (ELI5).We also show performance of our approach in zero-shot setting for another benchmark dataset (Debate-Pedia) -our approach leads to results comparable to baselines, which were specifically trained on DebatePedia.To aid the RL training, we propose a better semantic similarity reward, enabled by a novel Passage Embedding scheme developed using Cluster Hypothesis.Lastly, we contribute a gold-standard test dataset to further research in QfS and Long-form Question Answering (LfQA).
Swaroop Nath, Pushpak Bhattacharyya, Harshad Khadilkar
EMNLP3
2022 Gatekeeper: A deep reinforcement learning-cum-heuristic based algorithm for scheduling and routing trains in complex environments
abstract
The problem of optimal and efficient scheduling and navigation of trains in large railway networks has attracted attention from both operations research (OR) and artificial intelligence (AI) communities. At its core, this problem is comprised of two inter-linked sub-problems: the vehicle re-scheduling problem (VRSP) and the multi-agent path-finding problem (MAPF). In this paper, we propose Gatekeeper: a reinforcement-learning-cum-heuristic based approach for scheduling and path planning of trains in complex environments. By extensive experiments on the Flatland (a public customisable environment for multi-train scheduling and path planning), we show that Gatekeeper outperforms top RL baselines both in terms of normalized scores and makespan, while remaining competitive against pure heuristic algorithms.
Deepak Mohapatra, Ankush Ojha, Harshad Khadilkar, Supratim Ghosh
IJCNN3
2022 A Novel Data Augmentation Technique for Out-of-Distribution Sample Detection Using Compounded Corruptions
Ramya Hebbalaguppe, Soumya Suvra Ghosal, Jatin Prakash, Harshad Khadilkar, Chetan Arora 0001
ECML/PKDD (3)4
2022 Scalable multi-product inventory control with lead time constraints using reinforcement learning
Hardik Meisheri, Nazneen N. Sultana, Mayank Baranwal, Vinita Baniwal, Somjit Nath, Satyam Verma, Balaraman Ravindran, Harshad Khadilkar
Neural Comput. Appl.8
2021 Revisiting State Augmentation methods for Reinforcement Learning with Stochastic Delays
abstract
Several real-world scenarios, such as remote control and sensing, are comprised of action and observation delays. The presence of delays degrades the performance of reinforcement learning (RL) algorithms, often to such an extent that algorithms fail to learn anything substantial. This paper formally describes the notion of Markov Decision Processes (MDPs) with stochastic delays and shows that delayed MDPs can be transformed into equivalent standard MDPs (without delays) with significantly simplified cost structure. We employ this equivalence to derive a model-free Delay-Resolved RL framework and show that even a simple RL algorithm built upon this framework achieves near-optimal rewards in environments with stochastic delays in actions and observations. The delay-resolved deep Q-network (DRDQN) algorithm is bench-marked on a variety of environments comprising of multi-step and stochastic delays and results in better performance, both in terms of achieving near-optimal rewards and minimizing the computational overhead thereof, with respect to the currently established algorithms.
Somjit Nath, Mayank Baranwal, Harshad Khadilkar
CIKM3
2021 FoLaR: Foggy Latent Representations for Reinforcement Learning with Partial Observability
abstract
We propose a novel methodology for improving the rate and consistency of reinforcement learning in partially observable (foggy) environments, under the broader umbrella of robust latent representations. The present work addresses partially observable environments, which violate the canonical Markov assumptions. We propose adaptations for any on-policy model-free deep reinforcement learning algorithm, in order to improve training in partially observable situations: (i) recurrent layers for including information from previous observations, (ii) predicting the step reward and the next latent representation as auxiliary outputs from the same latent space as used for inferring the action, and (iii) modification of the loss function to penalise errors in the two auxiliary outputs, in addition to the reward-based gradients used for policy training. We show that the proposed changes substantially improve learning in several environments over vanilla Proximal Policy Optimisation (PPO) and other baselines in literature, especially in known challenging environments with hard exploration.
Hardik Meisheri, Harshad Khadilkar
IJCNN2
2019 Reinforcement Learning of Supply Chain Control Policy Using Closed Loop Multi-agent Simulation
Souvik Barat, Monika Gajrani, Harshad Khadilkar, Hardik Meisheri, Vinita Baniwal, Vinay Kulkarni 0001
MABS4
2019 A Scalable Reinforcement Learning Algorithm for Scheduling Railway Lines
abstract
This paper describes an algorithm for scheduling bidirectional railway lines (both single- and multi-track) using a reinforcement learning (RL) approach. The goal is to define the track allocations and arrival/departure times for all trains on the line, given their initial positions, priority, halt times, and traversal times, while minimizing the total priority-weighted delay. The primary advantage of the proposed algorithm compared to exact approaches is its scalability, and compared to heuristic approaches is its solution quality. Efficient scaling is ensured by decoupling the size of the state-action space from the size of the problem instance. Improved solution quality is obtained because of the inherent adaptability of reinforcement learning to specific problem instances. An additional advantage is that the learning from one instance can be transferred with minimal re-learning to another instance with different infrastructure resources and traffic mix. It is shown that the solution quality of the RL algorithm exceeds that of two prior heuristic-based approaches while having comparable computation times. Two lines from the Indian rail network are used for demonstrating the applicability of the proposed algorithm in the real world.
Harshad Khadilkar
IEEE Trans. Intell. Transp. Syst.1
2014 Hybrid Communication Protocols and Control Algorithms for NextGen Aircraft Arrivals
abstract
Capacity constraints imposed by current air traffic management technologies and protocols could severely limit the performance of the Next Generation Air Transportation System (NextGen). A fundamental design decision in the development of this system is the level of decentralization that balances system safety and efficiency. A new surveillance technology called automatic dependent surveillance-broadcast (ADS-B) can be potentially used to shift air traffic control to a more distributed architecture; however, channel variations and interference with existing secondary radar replies can affect ADS-B systems. This paper presents a framework for managing arrivals at an airport by using a hybrid centralized/distributed algorithm for communication and control. The algorithm combines the centralized control that is used in congested regions with the distributed control that is used in lower traffic density regions. The hybrid algorithm is evaluated through realistic simulations of operations around a major airport. The proposed strategy is shown to significantly improve air traffic control performance under various operating conditions by adapting to the underlying communication, navigation, and surveillance systems. The performance of the proposed strategy is found to be comparable to fully centralized strategies, despite requiring significantly less ground infrastructure.
Pan Gun Park, Harshad Khadilkar, Hamsa Balakrishnan, Claire J. Tomlin
IEEE Trans. Intell. Transp. Syst.2