VLDB 2026 Research / reviewers in the wild / expert
Viacheslav Sinii
dblp:351/7957
· DBLP profile ↗
4ranked-venue papers
2as first author
4since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
4 papers |
Reinforcement learning · 46% Language models and text generation · 27% Learning paradigms · 12% |
Topics — the 6 heaviest of 7, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Reinforcement learning › meta-reinforcement learning
in-context reinforcement learning |
1.5 | 2 | 2024 | Emergence of In-Context Reinforcement Learning from Noise Distillation · ICML 2024 In-Context Reinforcement Learning for Variable Action Spaces · ICML 2024 |
Machine learning › Reinforcement learning
meta-reinforcement learning |
1.5 | 2 | 2024 | XLand-MiniGrid: Scalable Meta-Reinforcement Learning Environments in JAX · NeurIPS 2024 In-Context Reinforcement Learning for Variable Action Spaces · ICML 2024 |
Natural language and speech › Language models and text generation
alignment |
0.9 | 1 | 2025 | Steering LLM Reasoning Through Bias-Only Adaptation · EMNLP 2025 |
Machine learning › Learning paradigms
curriculum learning |
0.8 | 1 | 2024 | Emergence of In-Context Reinforcement Learning from Noise Distillation · ICML 2024 |
Machine learning › Efficient and distributed learning › large-scale learning
scalable training |
0.8 | 1 | 2024 | XLand-MiniGrid: Scalable Meta-Reinforcement Learning Environments in JAX · NeurIPS 2024 |
Machine learning › Learning theory
generalization |
0.2 | 1 | 2024 | XLand-MiniGrid: Scalable Meta-Reinforcement Learning Environments in JAX · NeurIPS 2024 |
Methods — techniques the papers use, named apart from their topics
transformer · 1.5parameter-efficient fine-tuning · 0.9bias-only adaptation · 0.9noise distillation · 0.8multi-episode context · 0.8curriculum learning · 0.8JAX · 0.8Headless-AD · 0.8GPU acceleration · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Steering LLM Reasoning Through Bias-Only AdaptationabstractViacheslav Sinii, Alexey Gorbatovski, Artem Cherepanov, Boris Shaposhnikov, Nikita Balagansky, Daniil Gavrilov. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Viacheslav Sinii, Alexey Gorbatovski, Artem Cherepanov, Boris Shaposhnikov, Nikita Balagansky, Daniil Gavrilov |
EMNLP | 1 |
| 2024 | In-Context Reinforcement Learning for Variable Action SpacesabstractRecently, it has been shown that transformers pre-trained on diverse datasets with multi-episode contexts can generalize to new reinforcement learning tasks in-context. A key limitation of previously proposed models is their reliance on a predefined action space size and structure. The introduction of a new action space often requires data re-collection and model re-training, which can be costly for some applications. In our work, we show that it is possible to mitigate this issue by proposing the Headless-AD model that, despite being trained only once, is capable of generalizing to discrete action spaces of variable size, semantic content and order. By experimenting with Bernoulli and contextual bandits, as well as a gridworld environment, we show that Headless-AD exhibits significant capability to generalize to action spaces it has never encountered, even outperforming specialized models trained for a specific set of actions on several environment configurations. Viacheslav Sinii, Alexander Nikulin, Vladislav Kurenkov, Ilya Zisman, Sergey Kolesnikov |
ICML | 1 |
| 2024 | Emergence of In-Context Reinforcement Learning from Noise DistillationabstractRecently, extensive studies in Reinforcement Learning have been carried out on the ability of transformers to adapt in-context to various environments and tasks. Current in-context RL methods are limited by their strict requirements for data, which needs to be generated by RL agents or labeled with actions from an optimal policy. In order to address this prevalent problem, we propose AD$^\varepsilon$, a new data acquisition approach that enables in-context Reinforcement Learning from noise-induced curriculum. We show that it is viable to construct a synthetic noise injection curriculum which helps to obtain learning histories. Moreover, we experimentally demonstrate that it is possible to alleviate the need for generation using optimal policies, with in-context RL still able to outperform the best suboptimal policy in a learning dataset by a 2x margin. Ilya Zisman, Vladislav Kurenkov, Alexander Nikulin, Viacheslav Sinii, Sergey Kolesnikov |
ICML | 4 |
| 2024 | XLand-MiniGrid: Scalable Meta-Reinforcement Learning Environments in JAXabstractInspired by the diversity and depth of XLand and the simplicity and minimalism of MiniGrid, we present XLand-MiniGrid, a suite of tools and grid-world environments for meta-reinforcement learning research. Written in JAX, XLand-MiniGrid is designed to be highly scalable and can potentially run on GPU or TPU accelerators, democratizing large-scale experimentation with limited resources. Along with the environments, XLand-MiniGrid provides pre-sampled benchmarks with millions of unique tasks of varying difficulty and easy-to-use baselines that allow users to quickly start training adaptive agents. In addition, we have conducted a preliminary analysis of scaling and generalization, showing that our baselines are capable of reaching millions of steps per second during training and validating that the proposed benchmarks are challenging. XLand-MiniGrid is open-source and available at \url{https://github.com/corl-team/xland-minigrid}. Alexander Nikulin, Vladislav Kurenkov, Ilya Zisman, Artem Agarkov, Viacheslav Sinii, Sergey Kolesnikov |
NeurIPS | 5 |