Ryan Cheng

dblp:333/2957 · DBLP profile ↗
← Back
2ranked-venue papers
0as first author
2since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 1 · 1 since 2021Computer networks · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Language models and text generation · 33% Reinforcement learning · 33% Generative modeling · 33%

Topics — the 2 heaviest of 3, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning
multi-turn reinforcement learning
0.912025
Consistently Simulating Human Personas with Multi-Turn Reinforcement Learning · NeurIPS 2025
Machine learning › Generative modeling › diffusion model › diffusion model adaptation
reward fine-tuning
0.912025
Consistently Simulating Human Personas with Multi-Turn Reinforcement Learning · NeurIPS 2025

Methods — techniques the papers use, named apart from their topics

large language model · 0.9human annotation · 0.9automatic metrics · 0.9
YearPublicationVenuePosition
2025 Consistently Simulating Human Personas with Multi-Turn Reinforcement Learning
abstract
Large Language Models (LLMs) are increasingly used to simulate human users in interactive settings such as therapy, education, and social role-play. While these simulations enable scalable training and evaluation of AI agents, off-the-shelf LLMs often drift from their assigned personas, contradict earlier statements, or abandon role-appropriate behavior. We introduce a unified framework for evaluating and improving persona consistency in LLM-generated dialogue. We define three automatic metrics—prompt-to-line consistency, line-to-line consistency, and Q\&A consistency—that capture different types of persona drift and validate each against human annotations. Using these metrics as reward signals, we apply multi-turn reinforcement learning to fine-tune LLMs for three user roles: a patient, a student, and a social chat partner. Our method reduces inconsistency by over 55%, resulting in more coherent, faithful, and trustworthy simulated users.
Marwa Abdulhai, Ryan Cheng, Donovan Clay, Tim Althoff, Sergey Levine, Natasha Jaques
NeurIPS2
2023 Analysis and Design of an Edge Computing Enabled Real-Time Object Detection Platform for Drone-as-a-Service Using Network Calculus
abstract
Numerous Drone-as-a-Service (DaaS) applications, such as surveillance, search and rescue, and infrastructure inspection, may employ realtime object detection to achieve computer vision-based autonomous functions. However, running object detection algorithms, e.g., YOLO, locally on a drone requires extensive computational power, which is expensive in terms of cost and energy consumption. Conversely, edge computing facil-itates the implementation of an affordable and efficient platform where drones compress and transmit images to an edge server for realtime object detection. Nevertheless, DaaS designers applying Edge Computing Enabled Real-Time Object Detection (ECOD) must be cognizant of the network design and performance of the ECOD platform to ensure object detection in realtime. In our research, we propose an approach to analyzing the delay performance of an ECOD platform utilizing network calculus. A testbed was implemented to evaluate the effectiveness of this approach. The analysis result provides principled guidance for the ECOD platform design lacking in previous studies. Examples are provided in this paper to illustrate how to apply the guidance to the ECOD platform design in terms of traffic profile, network capacity, and delay requirements in DaaS.
Ryan Cheng, Unmesh Khanolkar, Liang Cheng 0001
ICC2