Dmitriy Rivkin

dblp:279/1174 · DBLP profile ↗
← Back
9ranked-venue papers
4as first author
9since 2021 · last 2025
0000-0003-4136-4831ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021Systems, architecture and hardware · 5 · 2 first-author · 5 since 2021Computer networks · 4 · 2 first-author · 4 since 2021
YearPublicationVenuePosition
2025 AIoT Smart Home via Autonomous LLM Agents
abstract
The common-sense reasoning abilities and vast general knowledge of large language models (LLMs) make them a natural fit for interpreting user requests in a smart home assistant context. LLMs, however, lack specific knowledge about the user and their home, which limits their potential impact. Smart home agent with grounded execution (SAGE), overcomes these and other limitations by using a scheme in which a user request triggers an LLM-controlled sequence of discrete actions. These actions can be used to retrieve information, interact with the user, or manipulate device states. SAGE controls this process through a dynamically constructed tree of LLM prompts, which help it decide which action to take next, whether an action was successful, and when to terminate the process. The SAGE action set augments an LLM’s capabilities to support some of the most critical requirements for a smart home assistant. These include: flexible and scalable user preference management (“Is my team playing tonight?”), access to any smart device’s full functionality without device-specific code via API reading (“Turn down the screen brightness on my dryer”), persistent device state monitoring (“Remind me to throw out the milk when I open the fridge”), natural device references using only a photo of the room (“Turn on the lamp on the dresser”), and more. We introduce a benchmark of 50 new and challenging smart home tasks where SAGE achieves a 76% success rate, significantly outperforming existing LLM-enabled baselines (30% success rate).
Dmitriy Rivkin, Francois Robert Hogan, Amal Feriani, Abhisek Konar, Adam Sigal, Xue (Steve) Liu, Gregory Dudek
IEEE Internet Things J.1
2024 CARTIER: Cartographic lAnguage Reasoning Targeted at Instruction Execution for Robots
abstract
This work explores the capacity of large language models (LLMs) to address problems at the intersection of spatial planning and natural language interfaces for navigation. We focus on following complex instructions that are more akin to natural conversation than traditional explicit procedural directives typically seen in robotics. Unlike most prior work where navigation directives are provided as simple imperative commands (e.g., "go to the fridge"), we examine implicit directives obtained through conversational interactions.We leverage the 3D simulator AI2Thor to create household query scenarios at scale, and augment it by adding complex language queries for 40 object types. We demonstrate that a robot using our method CARTIER (Cartographic lAnguage Reasoning Targeted at Instruction Execution for Robots) can parse descriptive language queries up to 42% more reliably than existing LLM-enabled methods by exploiting the ability of LLMs to interpret the user interaction in the context of the objects in the scenario.
Dmitriy Rivkin, Nikhil Kakodkar, Francois Robert Hogan, Bobak H. Baghi, Gregory Dudek
ICRA1
2024 PhotoBot: Reference-Guided Interactive Photography via Natural Language
abstract
We introduce PhotoBot, a framework for fully automated photo acquisition based on an interplay between high-level human language guidance and a robot photographer. We propose to communicate photography suggestions to the user via reference images that are selected from a curated gallery. We leverage a visual language model (VLM) and an object detector to characterize the reference images via textual descriptions and then use a large language model (LLM) to retrieve relevant reference images based on a user’s language query through text-based reasoning. To correspond the reference image and the observed scene, we exploit pretrained features from a vision transformer capable of capturing semantic similarity across marked appearance variations. Using these features, we compute suggested pose adjustments for an RGB-D camera by solving a perspective-n-point (PnP) problem. We demonstrate our approach using a manipulator equipped with a wrist camera. Our user studies show that photos taken by PhotoBot are often more aesthetically pleasing than those taken by users themselves, as measured by human feedback. We also show that PhotoBot can generalize to other reference sources such as paintings.
Oliver Limoyo, Jimmy Li 0001, Dmitriy Rivkin, Jonathan Kelly, Gregory Dudek
IROS3
2023 Self-Supervised Transformer Architecture for Change Detection in Radio Access Networks
abstract
Radio Access Networks (RANs) for telecommunications represent large agglomerations of interconnected hardware consisting of hundreds of thousands of transmitting devices (cells). Such networks undergo frequent and often heterogeneous changes caused by network operators, who are seeking to tune their system parameters for optimal performance. The effects of such changes are challenging to predict and will become even more so with the adoption of fifth-generation/sixth-generation (5G/6G) networks. Therefore, RAN monitoring is vital for network operators. We propose a self-supervised learning framework that leverages self-attention and self-distillation for this task. It works by detecting changes in Performance Measurement data, a collection of time-varying metrics which reflect a set of diverse measurements of the network performance at the cell level. Experimental results show that our approach outperforms the state of the art by 4% on a real-world based dataset consisting of about hundred thousands time series. It also has the merits of being scalable and generalizable. This allows it to provide deep insight into the specifics of mode of operation changes while relying minimally on expert knowledge.
Igor Kozlov, Dmitriy Rivkin, Wei-Di Chang, Di Wu 0044, Xue Liu 0004, Gregory Dudek
ICC2
2023 ANSEL Photobot: A Robot Event Photographer with Semantic Intelligence
abstract
Our work examines the way in which large language models can be used for robotic planning and sampling in the context of automated photographic documentation. Specifically, we illustrate how to produce a photo-taking robot with an exceptional level of semantic awareness by leveraging recent advances in general purpose language (LM) and vision-language (VLM) models. Given a high-level description of an event we use an LM to generate a natural-language list of photo descriptions that one would expect a photographer to capture at the event. We then use a VLM to identify the best matches to these descriptions in the robot's video stream. The photo portfolios generated by our method are consistently rated as more appropriate to the event by human evaluators than those generated by existing methods.
Dmitriy Rivkin, Gregory Dudek, Nikhil Kakodkar, David Meger, Oliver Limoyo, Michael R. M. Jenkin, Xue Liu 0004, Francois Robert Hogan
ICRA1
2022 Visuotactile-RL: Learning Multimodal Manipulation Policies with Deep Reinforcement Learning
abstract
Manipulating objects with dexterity requires timely feedback that simultaneously leverages the senses of vision and touch. In this paper, we focus on the problem setting where both visual and tactile sensors provide pixel-level feedback for Visuotactile reinforcement learning agents. We investigate the challenges associated with multimodal learning and propose several improvements to existing RL methods; including tactile gating, tactile data augmentation, and visual degradation. When compared with visual-only and tactile-only baselines, our Visuotactile-RL agents showcase (1) significant improvements in contact-rich tasks; (2) improved robustness to visual changes (lighting/camera view) in the workspace; and (3) resilience to physical changes in the task environment (weight/friction of objects).
Johanna Hansen, Francois Robert Hogan, Dmitriy Rivkin, David Meger, Michael R. M. Jenkin, Gregory Dudek
ICRA3
2021 Learning Assisted Identification of Scenarios Where Network Optimization Algorithms Under-Perform
abstract
We present a generative adversarial method that uses deep learning to identify network load traffic conditions in which network optimization algorithms under-perform other known algorithms: the Deep Convolutional Failure Generator (DCFG). The spatial distribution of network load presents challenges for network operators for tasks such as load balancing, in which a network optimizer attempts to maintain high quality communication while at the same time abiding capacity constraints. Testing a network optimizer for all possible load distributions is challenging if not impossible. We propose a novel method that searches for load situations where a target network optimization method underperforms baseline, which are key test cases that can be used for future refinement and performance optimization. By modeling a realistic network simulator's quality assessments with a deep network and, in parallel, optimizing a load generation network, our method efficiently searches the high dimensional space of load patterns and reliably finds cases in which a target network optimization method under-performs a baseline by a significant margin.
Dmitriy Rivkin, David Meger, Di Wu 0044, Xi Chen 0009, Xue Liu 0004, Gregory Dudek
GLOBECOM1
2021 Load Balancing for Communication Networks via Data-Efficient Deep Reinforcement Learning
abstract
Within a cellular network, load balancing between different cells is of critical importance to network performance and quality of service. Most existing load balancing algorithms are manually designed and tuned rule-based methods where near-optimality is almost impossible to achieve. These rule-based meth-ods are difficult to adapt quickly to traffic changes in real-world environments. Given the success of Reinforcement Learning (RL) algorithms in many application domains, there have been a number of efforts to tackle load balancing for communication systems using RL-based methods. To our knowledge, none of these efforts have addressed the need for data efficiency within the RL framework, which is one of the main obstacles in applying RL to wireless network load balancing. In this paper, we formulate the communication load balancing problem as a Markov Decision Process and propose a data-efficient transfer deep reinforcement learning algorithm to address it. Experimental results show that the proposed method can significantly improve the system performance over other baselines and is more robust to environmental changes.
Di Wu 0044, Jikun Kang, Yi Tian Xu, Jimmy Li 0001, Xi Chen 0009, Dmitriy Rivkin, Michael R. M. Jenkin, Taeseop Lee, Intaik Park, Xue Liu 0004, Gregory Dudek
GLOBECOM7
2021 Optimizing Cellular Networks via Continuously Moving Base Stations on Road Networks
abstract
Although existing cellular network base stations are typically immobile, the recent development of small form factor base stations and self driving cars has enabled the possibility of deploying a team of continuously moving base stations that can reorganize the network infrastructure to adapt to changing network traffic usage patterns. Given such a system of mobile base stations (MBSes) that can freely move on the road, how should their path be planned in an effort to optimize the experience of the users? This paper addresses this question by modeling the problem as a Markov Decision Process where the actions correspond to the MBSes deciding which direction to go at traffic intersections; states corresponds to the position of MBSes; and rewards correspond to minimization of packet loss in the network. A Monte Carlo Tree Search (MCTS)-based anytime algorithm that produces path plans for multiple base stations while optimizing expected packet loss is proposed. Simulated experiments in the city of Verdun, QC, Canada with varying user equipment (UE) densities and random initial conditions show that the proposed approach consistently outperforms myopic planners, and is able to achieve near-optimal performance.
Yogesh A. Girdhar, Dmitriy Rivkin, Di Wu 0044, Michael R. M. Jenkin, Xue Liu 0004, Gregory Dudek
ICRA2