Viacheslav Sinii

dblp:351/7957 · DBLP profile ↗
← Back
4ranked-venue papers
2as first author
4since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Reinforcement learning · 46% Language models and text generation · 27% Learning paradigms · 12%

Topics — the 6 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning › meta-reinforcement learning
in-context reinforcement learning
1.522024
Emergence of In-Context Reinforcement Learning from Noise Distillation · ICML 2024
In-Context Reinforcement Learning for Variable Action Spaces · ICML 2024
Machine learning › Reinforcement learning
meta-reinforcement learning
1.522024
XLand-MiniGrid: Scalable Meta-Reinforcement Learning Environments in JAX · NeurIPS 2024
In-Context Reinforcement Learning for Variable Action Spaces · ICML 2024
Natural language and speech › Language models and text generation
alignment
0.912025
Steering LLM Reasoning Through Bias-Only Adaptation · EMNLP 2025
Machine learning › Learning paradigms
curriculum learning
0.812024
Emergence of In-Context Reinforcement Learning from Noise Distillation · ICML 2024
Machine learning › Efficient and distributed learning › large-scale learning
scalable training
0.812024
XLand-MiniGrid: Scalable Meta-Reinforcement Learning Environments in JAX · NeurIPS 2024
Machine learning › Learning theory
generalization
0.212024
XLand-MiniGrid: Scalable Meta-Reinforcement Learning Environments in JAX · NeurIPS 2024

Methods — techniques the papers use, named apart from their topics

transformer · 1.5parameter-efficient fine-tuning · 0.9bias-only adaptation · 0.9noise distillation · 0.8multi-episode context · 0.8curriculum learning · 0.8JAX · 0.8Headless-AD · 0.8GPU acceleration · 0.8
YearPublicationVenuePosition
2025 Steering LLM Reasoning Through Bias-Only Adaptation
abstract
Viacheslav Sinii, Alexey Gorbatovski, Artem Cherepanov, Boris Shaposhnikov, Nikita Balagansky, Daniil Gavrilov. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Viacheslav Sinii, Alexey Gorbatovski, Artem Cherepanov, Boris Shaposhnikov, Nikita Balagansky, Daniil Gavrilov
EMNLP1
2024 In-Context Reinforcement Learning for Variable Action Spaces
abstract
Recently, it has been shown that transformers pre-trained on diverse datasets with multi-episode contexts can generalize to new reinforcement learning tasks in-context. A key limitation of previously proposed models is their reliance on a predefined action space size and structure. The introduction of a new action space often requires data re-collection and model re-training, which can be costly for some applications. In our work, we show that it is possible to mitigate this issue by proposing the Headless-AD model that, despite being trained only once, is capable of generalizing to discrete action spaces of variable size, semantic content and order. By experimenting with Bernoulli and contextual bandits, as well as a gridworld environment, we show that Headless-AD exhibits significant capability to generalize to action spaces it has never encountered, even outperforming specialized models trained for a specific set of actions on several environment configurations.
Viacheslav Sinii, Alexander Nikulin, Vladislav Kurenkov, Ilya Zisman, Sergey Kolesnikov
ICML1
2024 Emergence of In-Context Reinforcement Learning from Noise Distillation
abstract
Recently, extensive studies in Reinforcement Learning have been carried out on the ability of transformers to adapt in-context to various environments and tasks. Current in-context RL methods are limited by their strict requirements for data, which needs to be generated by RL agents or labeled with actions from an optimal policy. In order to address this prevalent problem, we propose AD$^\varepsilon$, a new data acquisition approach that enables in-context Reinforcement Learning from noise-induced curriculum. We show that it is viable to construct a synthetic noise injection curriculum which helps to obtain learning histories. Moreover, we experimentally demonstrate that it is possible to alleviate the need for generation using optimal policies, with in-context RL still able to outperform the best suboptimal policy in a learning dataset by a 2x margin.
Ilya Zisman, Vladislav Kurenkov, Alexander Nikulin, Viacheslav Sinii, Sergey Kolesnikov
ICML4
2024 XLand-MiniGrid: Scalable Meta-Reinforcement Learning Environments in JAX
abstract
Inspired by the diversity and depth of XLand and the simplicity and minimalism of MiniGrid, we present XLand-MiniGrid, a suite of tools and grid-world environments for meta-reinforcement learning research. Written in JAX, XLand-MiniGrid is designed to be highly scalable and can potentially run on GPU or TPU accelerators, democratizing large-scale experimentation with limited resources. Along with the environments, XLand-MiniGrid provides pre-sampled benchmarks with millions of unique tasks of varying difficulty and easy-to-use baselines that allow users to quickly start training adaptive agents. In addition, we have conducted a preliminary analysis of scaling and generalization, showing that our baselines are capable of reaching millions of steps per second during training and validating that the proposed benchmarks are challenging. XLand-MiniGrid is open-source and available at \url{https://github.com/corl-team/xland-minigrid}.
Alexander Nikulin, Vladislav Kurenkov, Ilya Zisman, Artem Agarkov, Viacheslav Sinii, Sergey Kolesnikov
NeurIPS5