Roman Sultimov

dblp:423/7302 · DBLP profile ↗
← Back
3ranked-venue papers
2as first author
3since 2021 · last 2026
0000-0001-9081-2616ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Multi-agent systems · 80% Reinforcement learning · 20%
Interdisciplinary, comprehensive, and emerging computing
2 papers
Smart cities and intelligent transportation · 50% Computational social science and digital humanities · 50%
Theoretical computer science
1 paper
Mathematical optimization · 100%

Topics — the 4 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Knowledge, reasoning and agents › Multi-agent systems
agent-based simulation
2.022026
RESPOND: Realistic Environment Simulation of Population and Natural Disasters with LLM-Driven Agents (Student Abstract) · AAAI 2026
RESPOND: Realistic Environment Simulation of Population and Natural Disasters with LLM-Driven Agents · AAAI 2026
Machine learning › Reinforcement learning › bandit
contextual bandit
1.012026
Efficient Contextual Bandit Learning via Reward-Space Sampling and Online Optimization (Student Abstract) · AAAI 2026
Mathematical optimization
stochastic optimization
1.012026
Efficient Contextual Bandit Learning via Reward-Space Sampling and Online Optimization (Student Abstract) · AAAI 2026
Smart cities and intelligent transportation
disaster management
0.622026
RESPOND: Realistic Environment Simulation of Population and Natural Disasters with LLM-Driven Agents (Student Abstract) · AAAI 2026
RESPOND: Realistic Environment Simulation of Population and Natural Disasters with LLM-Driven Agents · AAAI 2026

Methods — techniques the papers use, named apart from their topics

large language model · 4.0flood forecasting · 4.0agent-based modeling · 4.0upper confidence bound · 2.0thompson sampling · 2.0reward-space sampling · 2.0
YearPublicationVenuePosition
2026 RESPOND: Realistic Environment Simulation of Population and Natural Disasters with LLM-Driven Agents
abstract
Climate change is driving more frequent and severe disasters, putting people and infrastructure at risk. Protecting communities requires models that capture both natural disasters dynamics and how people behave under extreme conditions. This demo presents RESPOND, a multi-agent LLM-enhanced platform that jointly simulates natural hazards and human response. RESPOND couples high-fidelity flood AI forecasting an agent-based model of human behavior. LLM modules improve each agent decision-making, enabling context-aware reasoning over alerts, road closures, social signals, and changing water levels. The system simulates evacuation flows, resource seeking, and communication patterns producing actionable outputs for emergency management, urban planning, and policy. In the live demo one can run what-if or predicted scenarios, adjust assumptions, and observe emergent population behavior and risk hot spots in real time. By tightly coupling dynamic hazards with LLM-driven multi-agent behavior, RESPOND moves beyond fragmented tools and offers a practical, integrated platform for disaster preparedness and response.
Roman Sultimov, Mikhail Mozikov, Dmitrii Abramov, Mariia Kovalchuk, Maksim Malykh, Ilya Makarov, Andrei Osiptsov, Aleksandr Volkov, Yury Maximov
AAAI1
2026 RESPOND: Realistic Environment Simulation of Population and Natural Disasters with LLM-Driven Agents (Student Abstract)
abstract
Climate change is driving more frequent and severe disasters, putting people and infrastructure at risk. Protecting communities requires models that capture both natural disasters dynamics and how people behave under extreme conditions. This demo presents RESPOND, a multi-agent LLM-enhanced platform that jointly simulates natural hazards and human response. RESPOND couples high-fidelity flood AI forecasting an agent-based model of human behavior. LLM modules improve each agent decision-making, enabling context-aware reasoning over alerts, road closures, social signals, and changing water levels. The system simulates evacuation flows, resource seeking, and communication patterns producing actionable outputs for emergency management, urban planning, and policy. In the live demo one can run what-if or predicted scenarios, adjust assumptions, and observe emergent population behavior and risk hot spots in real time. By tightly coupling dynamic hazards with LLM-driven multi-agent behavior, RESPOND moves beyond fragmented tools and offers a practical, integrated platform for disaster preparedness and response.
Roman Sultimov, Mikhail Mozikov, Dmitrii Abramov, Mariia Kovalchuk, Maksim Malykh, Aleksandr Volkov, Ilya Makarov, Andrei Osiptsov, Yury Maximov
AAAI1
2026 Efficient Contextual Bandit Learning via Reward-Space Sampling and Online Optimization (Student Abstract)
abstract
The contextual multi-armed bandit problem underlies applications in recommendations, e-commerce, finance, and healthcare, where balancing exploration and exploitation is critical. While algorithms such as Upper Confidence Bound (UCB) and Thompson Sampling (TS) achieve strong theoretical guarantees, they often incur heavy computational cost from high-dimensional parameter estimation. We propose a new approach that combines reward sampling with online stochastic optimization. At each round, the algorithm samples hypothetical rewards for all actions and selects the action with the largest draw; the observed reward then updates the model via stochastic optimization. This design is both simple and efficient, preserving exploration while avoiding the pitfalls of greedy behavior on near-duplicate arms. Across synthetic and real-world datasets, our method attains near-optimal reward more quickly and with substantially lower computation than TS and UCB, demonstrating that sampling directly in reward space can improve both statistical efficiency and scalability.
Egor Suraveikin, Dastan Omirzak, Roman Sultimov, Yury Maximov
AAAI3