VLDB 2026 Research / reviewers in the wild / expert
Roman Sultimov
dblp:423/7302
· DBLP profile ↗
3ranked-venue papers
2as first author
3since 2021 · last 2026
0000-0001-9081-2616ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 3 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
Multi-agent systems · 80% Reinforcement learning · 20% | |
| Interdisciplinary, comprehensive, and emerging computing
2 papers |
Smart cities and intelligent transportation · 50% Computational social science and digital humanities · 50% | |
| Theoretical computer science
1 paper |
Mathematical optimization · 100% |
Topics — the 4 heaviest of 6, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Knowledge, reasoning and agents › Multi-agent systems
agent-based simulation |
2.0 | 2 | 2026 | RESPOND: Realistic Environment Simulation of Population and Natural Disasters with LLM-Driven Agents (Student Abstract) · AAAI 2026 RESPOND: Realistic Environment Simulation of Population and Natural Disasters with LLM-Driven Agents · AAAI 2026 |
Machine learning › Reinforcement learning › bandit
contextual bandit |
1.0 | 1 | 2026 | Efficient Contextual Bandit Learning via Reward-Space Sampling and Online Optimization (Student Abstract) · AAAI 2026 |
Mathematical optimization
stochastic optimization |
1.0 | 1 | 2026 | Efficient Contextual Bandit Learning via Reward-Space Sampling and Online Optimization (Student Abstract) · AAAI 2026 |
Smart cities and intelligent transportation
disaster management |
0.6 | 2 | 2026 | RESPOND: Realistic Environment Simulation of Population and Natural Disasters with LLM-Driven Agents (Student Abstract) · AAAI 2026 RESPOND: Realistic Environment Simulation of Population and Natural Disasters with LLM-Driven Agents · AAAI 2026 |
Methods — techniques the papers use, named apart from their topics
large language model · 4.0flood forecasting · 4.0agent-based modeling · 4.0upper confidence bound · 2.0thompson sampling · 2.0reward-space sampling · 2.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | RESPOND: Realistic Environment Simulation of Population and Natural Disasters with LLM-Driven AgentsabstractClimate change is driving more frequent and severe disasters, putting people and infrastructure at risk. Protecting communities requires models that capture both natural disasters dynamics and how people behave under extreme conditions. This demo presents RESPOND, a multi-agent LLM-enhanced platform that jointly simulates natural hazards and human response. RESPOND couples high-fidelity flood AI forecasting an agent-based model of human behavior. LLM modules improve each agent decision-making, enabling context-aware reasoning over alerts, road closures, social signals, and changing water levels. The system simulates evacuation flows, resource seeking, and communication patterns producing actionable outputs for emergency management, urban planning, and policy. In the live demo one can run what-if or predicted scenarios, adjust assumptions, and observe emergent population behavior and risk hot spots in real time. By tightly coupling dynamic hazards with LLM-driven multi-agent behavior, RESPOND moves beyond fragmented tools and offers a practical, integrated platform for disaster preparedness and response. Roman Sultimov, Mikhail Mozikov, Dmitrii Abramov, Mariia Kovalchuk, Maksim Malykh, Ilya Makarov, Andrei Osiptsov, Aleksandr Volkov, Yury Maximov |
AAAI | 1 |
| 2026 | RESPOND: Realistic Environment Simulation of Population and Natural Disasters with LLM-Driven Agents (Student Abstract)abstractClimate change is driving more frequent and severe disasters, putting people and infrastructure at risk. Protecting communities requires models that capture both natural disasters dynamics and how people behave under extreme conditions. This demo presents RESPOND, a multi-agent LLM-enhanced platform that jointly simulates natural hazards and human response. RESPOND couples high-fidelity flood AI forecasting an agent-based model of human behavior. LLM modules improve each agent decision-making, enabling context-aware reasoning over alerts, road closures, social signals, and changing water levels. The system simulates evacuation flows, resource seeking, and communication patterns producing actionable outputs for emergency management, urban planning, and policy. In the live demo one can run what-if or predicted scenarios, adjust assumptions, and observe emergent population behavior and risk hot spots in real time. By tightly coupling dynamic hazards with LLM-driven multi-agent behavior, RESPOND moves beyond fragmented tools and offers a practical, integrated platform for disaster preparedness and response. Roman Sultimov, Mikhail Mozikov, Dmitrii Abramov, Mariia Kovalchuk, Maksim Malykh, Aleksandr Volkov, Ilya Makarov, Andrei Osiptsov, Yury Maximov |
AAAI | 1 |
| 2026 | Efficient Contextual Bandit Learning via Reward-Space Sampling and Online Optimization (Student Abstract)abstractThe contextual multi-armed bandit problem underlies applications in recommendations, e-commerce, finance, and healthcare, where balancing exploration and exploitation is critical. While algorithms such as Upper Confidence Bound (UCB) and Thompson Sampling (TS) achieve strong theoretical guarantees, they often incur heavy computational cost from high-dimensional parameter estimation. We propose a new approach that combines reward sampling with online stochastic optimization. At each round, the algorithm samples hypothetical rewards for all actions and selects the action with the largest draw; the observed reward then updates the model via stochastic optimization. This design is both simple and efficient, preserving exploration while avoiding the pitfalls of greedy behavior on near-duplicate arms. Across synthetic and real-world datasets, our method attains near-optimal reward more quickly and with substantially lower computation than TS and UCB, demonstrating that sampling directly in reward space can improve both statistical efficiency and scalability. Egor Suraveikin, Dastan Omirzak, Roman Sultimov, Yury Maximov |
AAAI | 3 |