Jimmy Li 0001

dblp:83/9965-1 · DBLP profile ↗
← Back
15ranked-venue papers
6as first author
7since 2021 · last 2024
0000-0001-5472-4563ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 5 first-author · 1 since 2021Systems, architecture and hardware · 9 · 5 first-author · 1 since 2021Computer networks · 6 · 1 first-author · 6 since 2021
YearPublicationVenuePosition
2024 PhotoBot: Reference-Guided Interactive Photography via Natural Language
abstract
We introduce PhotoBot, a framework for fully automated photo acquisition based on an interplay between high-level human language guidance and a robot photographer. We propose to communicate photography suggestions to the user via reference images that are selected from a curated gallery. We leverage a visual language model (VLM) and an object detector to characterize the reference images via textual descriptions and then use a large language model (LLM) to retrieve relevant reference images based on a user’s language query through text-based reasoning. To correspond the reference image and the observed scene, we exploit pretrained features from a vision transformer capable of capturing semantic similarity across marked appearance variations. Using these features, we compute suggested pose adjustments for an RGB-D camera by solving a perspective-n-point (PnP) problem. We demonstrate our approach using a manipulator equipped with a wrist camera. Our user studies show that photos taken by PhotoBot are often more aesthetically pleasing than those taken by users themselves, as measured by human feedback. We also show that PhotoBot can generalize to other reference sources such as paintings.
Oliver Limoyo, Jimmy Li 0001, Dmitriy Rivkin, Jonathan Kelly, Gregory Dudek
IROS2
2023 Learning to Adapt: Communication Load Balancing via Adaptive Deep Reinforcement Learning
abstract
The association of mobile devices with network resources (e.g., base stations, frequency bands/channels), known as load balancing, is critical to reduce communication traffic congestion and network performance. Reinforcement learning (RL) has shown to be effective for communication load balancing and achieves better performance than currently used rule-based methods, especially when the traffic load changes quickly. However, RL-based methods usually need to interact with the environment for a large number of time steps to learn an effective policy and can be difficult to tune. In this work, we aim to improve the data efficiency of RL-based solutions to make them more suitable and applicable for real-world applications. Specifically, we propose a simple, yet efficient and effective deep RL-based wireless network load balancing framework. In this solution, a set of good initialization values for control actions are selected with some cost-efficient approach to center the training of the RL agent. Then, a deep RL-based agent is trained to find offsets from the initialization values that optimize the load balancing problem. Experimental evaluation on a set of dynamic traffic scenarios demonstrates the effectiveness and efficiency of the proposed method.
Di Wu 0044, Yi Tian Xu, Jimmy Li 0001, Michael R. M. Jenkin, Ekram Hossain 0001, Seowoo Jang, Jianzhong Zhang 0002, Xue Liu 0004, Gregory Dudek
GLOBECOM3
2023 Policy Reuse for Communication Load Balancing in Unseen Traffic Scenarios
abstract
With the continuous growth in communication network complexity and traffic volume, communication load balancing solutions are receiving increasing attention. Specifically, reinforcement learning (RL)-based methods have shown impressive performance compared with traditional rule-based methods. However, standard RL methods generally require an enormous amount of data to train, and generalize poorly to scenarios that are not encountered during training. We propose a policy reuse framework in which a policy selector chooses the most suitable pre-trained RL policy to execute based on the current traffic condition. Our method hinges on a policy bank composed of policies trained on a diverse set of traffic scenarios. When deploying to an unknown traffic scenario, we select a policy from the policy bank based on the similarity between the previous-day traffic of the current scenario and the traffic observed during training. Experiments demonstrate that this framework can outperform classical and adaptive rule-based methods by a large margin.
Jimmy Li 0001, Di Wu 0044, Michael R. M. Jenkin, Seowoo Jang, Xue Liu 0004, Gregory Dudek
ICC2
2022 Traffic Scenario Clustering and Load Balancing with Distilled Reinforcement Learning Policies
abstract
Due to the rapid increase in wireless communication traffic in recent years, load balancing is becoming increasingly important for ensuring the quality of service. However, variations in traffic patterns near different serving base stations make this task challenging. On one hand, crafting a single control policy that performs well across all base station sectors is often difficult. On the other hand, maintaining separate controllers for every sector introduces overhead, and leads to redundancy if some of the sectors experience similar traffic patterns. In this paper, we propose to construct a concise set of controllers that cover a wide range of traffic scenarios, allowing the operator to select a suitable controller for each sector based on local traffic conditions. To construct these controllers, we present a method that clusters similar scenarios and learns a general control policy for each cluster. We use deep reinforcement learning (RL) to first train separate control policies on diverse traffic scenarios, and then incrementally merge together similar RL policies via knowledge distillation. Experimental results show that our concise policy set reduces redundancy with very minor performance degradation compared to policies trained separately on each traffic scenario. Our method also outperforms handcrafted control parameters, joint learning on all tasks, and two popular clustering methods.
Jimmy Li 0001, Di Wu 0044, Yi Tian Xu, Tianyu Li 0008, Seowoo Jang, Xue Liu 0004, Gregory Dudek
ICC1
2022 Coordinated Load Balancing in Mobile Edge Computing Network: a Multi-Agent DRL Approach
abstract
Mobile edge computing (MEC) networks have been recently adopted to accommodate the fast-growing number of mobile devices performing complicated tasks with limited hardware capability. Recently, edge nodes with communication, computation, and caching capacities are starting to be deployed in MEC networks. Due to the physical separation of these resources, efficient coordination and scheduling are important for efficient resource utilization and optimal network performance. In this paper, we study mobility load balancing for communication, computation, and caching-enabled heterogeneous MEC networks. Specifically, we propose to tackle this problem via a multi-agent deep reinforcement learning-based framework. Users served by overloaded edge nodes are handed over to less loaded ones, to minimize the load in the most loaded base station in the network. In this framework, the handover decision for each user is made based on the user’s own observation which comprises the user’s task at hand and the load status of the MEC network. Simulation results show that our proposed multi-agent deep reinforcement learning-based approach can reduce the time-average maximum load by up to 30% and the end-to-end delay by 50% compared to baseline algorithms.
Manyou Ma, Di Wu 0044, Yi Tian Xu, Jimmy Li 0001, Seowoo Jang, Xue Liu 0004, Gregory Dudek
ICC4
2022 Multiobjective Load Balancing for Multiband Downlink Cellular Networks: A Meta- Reinforcement Learning Approach
abstract
Load balancing has become a key technique to handle the increasing traffic demand and improve the user experience. It evenly distributes the traffic across network resources by offloading users from overloaded base stations or channels to less crowded ones. Load balancing is a multi-objective optimization problem involving the automatic adjustment of several parameters to simultaneously maximize multiple network performance indicators. However, the existing methods mostly rely on single-objective approaches which lead to sub-optimal solutions. In this paper, we introduce the first multi-objective reinforcement learning (MORL) framework for load balancing. Specifically, we propose a solution based on meta-reinforcement learning (meta-RL) to learn a general policy capable of quickly adapting to new trade-offs between the objectives. We further enhance the generalization of our proposed solution using policy distillation techniques. To showcase the effectiveness of our framework, experiments are conducted based on real-world traffic scenarios. Our results show that our load balancing framework can (i) significantly outperform the existing rule-based and single-objective solutions, (ii) compute better Pareto front approximations compared to MORL baselines, and (iii) quickly adapt to new objective trade-offs.
Amal Feriani, Di Wu 0044, Yi Tian Xu, Jimmy Li 0001, Seowoo Jang, Ekram Hossain 0001, Xue Liu 0004, Gregory Dudek
IEEE J. Sel. Areas Commun.4
2021 Load Balancing for Communication Networks via Data-Efficient Deep Reinforcement Learning
abstract
Within a cellular network, load balancing between different cells is of critical importance to network performance and quality of service. Most existing load balancing algorithms are manually designed and tuned rule-based methods where near-optimality is almost impossible to achieve. These rule-based meth-ods are difficult to adapt quickly to traffic changes in real-world environments. Given the success of Reinforcement Learning (RL) algorithms in many application domains, there have been a number of efforts to tackle load balancing for communication systems using RL-based methods. To our knowledge, none of these efforts have addressed the need for data efficiency within the RL framework, which is one of the main obstacles in applying RL to wireless network load balancing. In this paper, we formulate the communication load balancing problem as a Markov Decision Process and propose a data-efficient transfer deep reinforcement learning algorithm to address it. Experimental results show that the proposed method can significantly improve the system performance over other baselines and is more robust to environmental changes.
Di Wu 0044, Jikun Kang, Yi Tian Xu, Jimmy Li 0001, Xi Chen 0009, Dmitriy Rivkin, Michael R. M. Jenkin, Taeseop Lee, Intaik Park, Xue Liu 0004, Gregory Dudek
GLOBECOM5
2020 View-Invariant Loop Closure with Oriented Semantic Landmarks
abstract
Recent work on semantic simultaneous localization and mapping (SLAM) have shown the utility of natural objects as landmarks for improving localization accuracy and robustness. In this paper we present a monocular semantic SLAM system that uses object identity and inter-object geometry for view-invariant loop detection and drift correction. Our system's ability to recognize an area of the scene even under large changes in viewing direction allows it to surpass the mapping accuracy of ORB-SLAM, which uses only local appearance-based features that are not robust to large viewpoint changes. Experiments on real indoor scenes show that our method achieves mean drift reduction of 70% when compared directly to ORB-SLAM. Additionally, we propose a method for object orientation estimation, where we leverage the tracked pose of a moving camera under the SLAM setting to overcome ambiguities caused by object symmetry. This allows our SLAM system to produce geometrically detailed semantic maps with object orientation, translation, and scale.
Jimmy Li 0001, Karim Koreitem, David Meger, Gregory Dudek
ICRA1
2019 Underwater Communication Using Full-Body Gestures and Optimal Variable-Length Prefix Codes
abstract
In this paper we consider inter-robot communication in the context of joint activities. In particular, we focus on convoying and passive communication for radio-denied environments by using whole-body gestures to provide cues regarding future actions. We develop a communication protocol whereby information described by codewords is transmitted by a series of actions executed by a swimming robot. These action sequences are chosen to optimize robustness and transmission duration given the observability, natural activity of the robot and the frequency of different messages. Our approach uses a convolutional network to make core observations of the pose of the robot being tracked, which is sending messages. The observer robot then uses an adaptation of classical decoding methods to infer a message that is being transmitted. The system is trained and validated using simulated data, tested in the pool and is targeted for deployment in the open ocean. Our decoder achieves.94 precision and.66 recall on real footage of robot gesture execution recorded in a swimming pool.
Karim Koreitem, Jimmy Li 0001, Ian Karp, Travis Manderson, Gregory Dudek
ICRA2
2019 Semantic Mapping for View-Invariant Relocalization
abstract
We propose a system for visual simultaneous localization and mapping (SLAM) that combines traditional local appearance-based features with semantically meaningful object landmarks to achieve both accurate local tracking and highly view-invariant object-driven relocalization. Our mapping process uses a sampling-based approach to efficiently infer the 3D pose of object landmarks from 2D bounding box object detections. These 3D landmarks then serve as a view-invariant representation which we leverage to achieve camera relocalization even when the viewing angle changes by more than 125 degrees. This level of view-invariance cannot be attained by local appearance-based features (e.g. SIFT) since the same set of surfaces are not even visible when the viewpoint changes significantly. Our experiments show that even when existing methods fail completely for viewpoint changes of more than 70 degrees, our method continues to achieve a relocalization rate of around 90%, with a mean rotational error of around 8 degrees.
Jimmy Li 0001, David Meger, Gregory Dudek
ICRA1
2017 Context-coherent scenes of objects for camera pose estimation
abstract
We propose an approach to vision-based pose estimation using object recognition and identity. Whereas feature based scene recognition and pose estimation methods are well established as effective means for estimating motion and recognizing locations, feature-based methods depend critically on the detection of common local features from one view of a scene to another. We focus on place recognition and pose change estimation in the context of large changes in viewing position, even to the extent that no common surfaces are seen between the two views. Our approach is based on using object identities and their inter-relationship to compute pose change. An important secondary outcome of our method is that it simultaneously infers the 3D poses of objects in the scene that are used as features. Such an object-based approach is inspired by a vast literature on human perception and has the potential for great robustness, albeit at the expense of accuracy. We propose a formulation of the problem using pairwise contextual constraints and develop an efficient algorithmic solution. We validate the approach and quantify its performance using the publicly available TUM SLAM dataset [1].
Jimmy Li 0001, David Meger, Gregory Dudek
IROS1
2017 Underwater multi-robot convoying using visual tracking by detection
abstract
We present a robust multi-robot convoying approach that relies on visual detection of the leading agent, thus enabling target following in unstructured 3-D environments. Our method is based on the idea of tracking-by-detection, which interleaves efficient model-based object detection with temporal filtering of image-based bounding box estimation. This approach has the important advantage of mitigating tracking drift (i.e. drifting away from the target object), which is a common symptom of model-free trackers and is detrimental to sustained convoying in practice. To illustrate our solution, we collected extensive footage of an underwater robot in ocean settings, and hand-annotated its location in each frame. Based on this dataset, we present an empirical comparison of multiple tracker variants, including the use of several convolutional neural networks, both with and without recurrent connections, as well as frequency-based model-free trackers. We also demonstrate the practicality of this tracking-by-detection strategy in real-world scenarios by successfully controlling a legged underwater robot in five degrees of freedom to follow another robot's independent motion.
Florian Shkurti, Wei-Di Chang, Peter Henderson 0002, Md Jahidul Islam, Juan Camilo Gamboa Higuera, Jimmy Li 0001, Travis Manderson, Anqi Xu 0003, Gregory Dudek, Junaed Sattar
IROS6
2016 Learning to generalize 3D spatial relationships
abstract
This paper presents an approach to learn meaningful spatial relationships in an unsupervised fashion from the distribution of 3D object poses in the real world. Our approach begins by extracting an over-complete set of features to describe the relative geometry of two objects. Each relationship type is modeled using a relevance-weighted distance over this feature space. This effectively ignores irrelevant feature dimensions. Our algorithm RANSEM for determining subsets of data that share a relationship as well as the model to describe each relationship is based on robust sample-based clustering. This approach combines the search for consistent groups of data with the extraction of models that precisely capture the geometry of those groups. An iterative refinement scheme has shown to be an effective approach for finding concepts of differing degrees of geometric specificity. Our results show that the models learned by our approach correlate strongly with the English labels that have been given by a human annotator to a set of validation data drawn from the NYUv2 real-world Kinect dataset, demonstrating that these concepts can be automatically acquired given sufficient experience. Additionally, the results of our method significantly out-perform K-means, a standard baseline for unsupervised cluster extraction.
Jimmy Li 0001, David Meger, Gregory Dudek
ICRA1
2012 Multi-domain monitoring of marine environments using a heterogeneous robot team
abstract
In this paper we describe a heterogeneous multi-robot system for assisting scientists in environmental monitoring tasks, such as the inspection of marine ecosystems. This team of robots is comprised of a fixed-wing aerial vehicle, an autonomous airboat, and an agile legged underwater robot. These robots interact with off-site scientists and operate in a hierarchical structure to autonomously collect visual footage of interesting underwater regions, from multiple scales and mediums. We discuss organizational and scheduling complexities associated with multi-robot experiments in a field robotics setting. We also present results from our field trials, where we demonstrated the use of this heterogeneous robot team to achieve multi-domain monitoring of coral reefs, based on real-time interaction with a remotely-located marine biologist.
Florian Shkurti, Anqi Xu 0003, Malika Meghjani, Juan Camilo Gamboa Higuera, Yogesh A. Girdhar, Philippe Giguère, Bir Bikram Dey, Jimmy Li 0001, Arnold Kalmbach, Chris Prahacs, Katrine Turgeon, Ioannis M. Rekleitis, Gregory Dudek
IROS8
2011 Graphical State Space Programming: A visual programming paradigm for robot task specification
abstract
We describe a framework that combines a software development paradigm, a software visualization technique, and a tool for robot programming. This infrastructure is called "Graphical State Space Programming" (GSSP), and allows robot application programs to be decomposed and visualized within state-dependent views. Our approach simplifies and expedites the programming process for robot routines and behaviors, and we examine the performance improvement that ensues through a set of controlled user studies. The usability and effectiveness of GSSP are also illustrated using a field demonstration with an aerial robotic vehicle.
Jimmy Li 0001, Anqi Xu 0003, Gregory Dudek
ICRA1