Siddharth Verma

dblp:99/5437 · DBLP profile ↗
← Back
6ranked-venue papers
1as first author
2since 2021 · last 2023
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Reinforcement learning · 47% Language models and text generation · 31% Robot manipulation · 22%

Topics — the 6 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning
exploration
0.412020
Continual Learning of Control Primitives : Skill Discovery via Reset-Games · NeurIPS 2020
Machine learning › Reinforcement learning › non-stationary reinforcement learning › continual reinforcement learning
reset-free reinforcement learning
0.412020
Continual Learning of Control Primitives : Skill Discovery via Reset-Games · NeurIPS 2020
Machine learning › Reinforcement learning › hierarchical reinforcement learning › skill learning
skill discovery
0.412020
Continual Learning of Control Primitives : Skill Discovery via Reset-Games · NeurIPS 2020
Natural language and speech › Language models and text generation
evaluation of language models
0.212023
ALERT: Adapt Language Models to Reasoning Tasks · ACL (1) 2023
Robotics › Robot manipulation › continuum robot
concentric tube robot
0.112019
TREE: A Variable Topology, Branching Continuum Robot · ICRA 2019
Robotics › Robot manipulation › cable-driven robot
tendon-driven robot
0.112019
TREE: A Variable Topology, Branching Continuum Robot · ICRA 2019

Methods — techniques the papers use, named apart from their topics

reasoning task evaluation · 0.7language model adaptation · 0.7reset-games · 0.4general-sum game formulation · 0.4prototype design · 0.4hybrid concentric-tube/tendon actuation · 0.4
YearPublicationVenuePosition
2023 ALERT: Adapt Language Models to Reasoning Tasks
abstract
Ping Yu, Tianlu Wang, Olga Golovneva, Badr AlKhamissi, Siddharth Verma, Zhijing Jin, Gargi Ghosh, Mona Diab, Asli Celikyilmaz. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Olga Golovneva, Badr AlKhamissi, Siddharth Verma, Zhijing Jin 0001, Gargi Ghosh, Mona T. Diab, Asli Celikyilmaz
ACL (1)5
2022 CHAI: A CHatbot AI for Task-Oriented Dialogue with Offline Reinforcement Learning
abstract
Siddharth Verma, Justin Fu, Sherry Yang, Sergey Levine. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Siddharth Verma, Justin Fu, Sherry Yang 0001, Sergey Levine
NAACL-HLT1
2020 A Guided Learning Approach for Generative Adversarial Networks
abstract
In this paper, we propose a novel technique for training Generative Adversarial Networks (GANs) using autoencoders. GANs, in recent years, have emerged as one of the most popular generative models. Despite their success, there are several challenges in maintaining the trade-off between diversity and quality of the generated distribution. Our idea stems from the fact that deeper layers of an autoencoder contain high-level feature representation of the input data distribution. Reusing these layers provides GAN with information about the representative characteristics of real data and hence can guide its adversarial training. We call our model Guided GAN since the autoencoder (guiding network) provides a direction to train the GAN (generative network). Guided GAN also minimizes both the forward and reverse Kullback-Leibler (KL) divergence in a single model, exploiting the complementary statistical properties of the two. We conduct extensive experiments and use various metrics for assessing the quality, diversity of generated images and convergence of the model. Our model is evaluated on two standard datasets: CIFAR-10 and CelebA demonstrating either superior or competitive performance compared to baseline GANs, especially in the earlier training stages. Our guided training procedure has been tested on different baseline GANs without any changes to their hyper-parameter configuration or architecture.
Sidhant Nagpal, Siddharth Verma, Shikhar Gupta, Swati Aggarwal
IJCNN2
2020 Continual Learning of Control Primitives : Skill Discovery via Reset-Games
abstract
Reinforcement learning has the potential to automate the acquisition of behavior in complex settings, but in order for it to be successfully deployed, a number of practical challenges must be addressed. First, in real world settings, when an agent attempts a tasks and fails, the environment must somehow "reset" so that the agent can attempt the task again. While easy in simulation, this could require considerable human effort in the real world, especially if the number of trials is very large. Second, real world learning is often limited by challenges in exploration, as complex, temporally extended behavior is often times difficult to acquire with random exploration. In this work, we show how a single method can allow an agent to acquire skills with minimal supervision while removing the need for resets. We do this by exploiting the insight that the need to reset" an agent to a broad set of initial states for a learning task provides a natural setting to learn a diverse set ofreset-skills." We propose a general-sum game formulation that naturally balances the objective of resetting and learning skills, and demonstrate that this approach improves performance on reset-free tasks, and additionally show that the skills we obtain can be used to significantly accelerate downstream learning.
Kelvin Xu, Siddharth Verma, Chelsea Finn, Sergey Levine
NeurIPS2
2019 TREE: A Variable Topology, Branching Continuum Robot
abstract
We describe the design and physical realization of a novel branching continuum robot, aimed at inspection and cleaning operations in hard-to-reach environments at depths greater than human arm lengths. The design, based on a hybrid concentric-tube/tendon actuated continuum trunk core, features two pairs of fully retractable continuum branches. The retractable nature of the branches allows the robot to actively change its topology, allowing it to penetrate narrow openings and expand to adaptively engage complex environmental geometries. We detail and discuss the realization of a physical prototype of the design, and its testing in a simulated glove box environment.
Michael C. Lastinger, Siddharth Verma, Apoorva Kapadia, Ian D. Walker
ICRA2
2018 An Evolutionary Learning Approach to Play Othello Using XCS
abstract
Due to the multifarious challenges that emerge when developing an artificial intelligent (AI) agent that can compete with human players, the classic game of Othello has received a lot of attention from the Computational-Intelligence community. This paper proposes an AI agent that learns a winning strategy for the game of Othello using the eXtended Classifier System (XCS) algorithm which is a popular variant of the Learning Classifier System (LCS) algorithm. Othello has been a favourite in the study of AI due to its simple set of rules, low branching factor and well defined strategic concepts. The LCS system consists of a rule-set which is made to evolve using a combination of Reinforcement Learning (RL) and Genetic Algorithm (GA) such that the evolved rule-set learns an optimal action for each board state. A 6×6 Othello board will be used for this experiment in order to evaluate the applicability of the proposed agent in learning a winning game-playing strategy. The performance of the proposed agent was evaluated against three categories of agents: minimax, human and random agent. The XCS agent was able to outperform the above-mentioned agents showing the effectiveness of rule-based evolutionary learning in Othello. This work demonstrates the possibility of using the XCS algorithm in other strategy-based combinatorial games.
Satvik Jain, Siddharth Verma, Swaraj Kumar, Swati Aggarwal
CEC2