EDBT 2026 Demo / reviewers in the wild / expert
Fangkai Yang
dblp:91/3135
· DBLP profile ↗
46ranked-venue papers
12as first author
20since 2021 · last 2026
0000-0002-3089-0345ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 35 · 8 first-author · 16 since 2021Human-computer interaction and ubiquitous computing · 10 · 5 first-authorSoftware engineering, systems software and programming languages · 8 · 3 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 5 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 first-author · 1 since 2021Theory of computation · 3Systems, architecture and hardware · 2 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | RepoGenesis: Benchmarking End-to-End Microservice Generation from Readme to RepositoryabstractZhiyuan Peng, Xin Yin, Pu Zhao, Fangkai Yang, Lu Wang, Ran Jia, Xu Chen, Qingwei Lin, Saravan Rajmohan, Dongmei Zhang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Pu Zhao 0004, Fangkai Yang, Lu Wang 0029, Ran Jia, Xu Chen 0022, Qingwei Lin, Saravan Rajmohan, Dongmei Zhang 0001 |
ACL (1) | 4 |
| 2025 | WarriorCoder: Learning from Expert Battles to Augment Code Large Language ModelsabstractHuawen Feng, Pu Zhao, Qingfeng Sun, Can Xu, Fangkai Yang, Lu Wang, Qianli Ma, Qingwei Lin, Saravan Rajmohan, Dongmei Zhang, Qi Zhang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Huawen Feng, Pu Zhao 0004, Qingfeng Sun, Can Xu 0002, Fangkai Yang, Lu Wang 0029, Qianli Ma 0001, Qingwei Lin, Saravan Rajmohan, Dongmei Zhang 0001, Qi Zhang 0066 |
ACL (1) | 5 |
| 2025 | AXIS: Efficient Human-Agent-Computer Interaction with API-First LLM-Based AgentsabstractJunting Lu, Zhiyang Zhang, Fangkai Yang, Jue Zhang, Lu Wang, Chao Du, Qingwei Lin, Saravan Rajmohan, Dongmei Zhang, Qi Zhang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Junting Lu, Fangkai Yang, Lu Wang 0029, Qingwei Lin, Saravan Rajmohan, Dongmei Zhang 0001, Qi Zhang 0066 |
ACL (1) | 3 |
| 2025 | Thread: A Logic-Based Data Organization Paradigm for How-To Question Answering with Retrieval Augmented GenerationabstractKaikai An, Fangkai Yang, Liqun Li, Junting Lu, Sitao Cheng, Shuzheng Si, Lu Wang, Pu Zhao, Lele Cao, Qingwei Lin, Saravan Rajmohan, Dongmei Zhang, Baobao Chang. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Kaikai An, Fangkai Yang, Liqun Li, Junting Lu, Sitao Cheng, Shuzheng Si, Lu Wang 0029, Pu Zhao 0004, Le-le Cao, Qingwei Lin, Saravan Rajmohan, Dongmei Zhang 0001, Baobao Chang |
EMNLP | 2 |
| 2025 | ExeCoder: Empowering Large Language Models with Executability Representation for Code TranslationabstractMinghua He, Yue Chen, Fangkai Yang, Pu Zhao, Wenjie Yin, Yu Kang, Qingwei Lin, Saravan Rajmohan, Dongmei Zhang. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Minghua He, Yue Chen 0014, Fangkai Yang, Pu Zhao 0004, Yu Kang 0006, Qingwei Lin, Saravan Rajmohan, Dongmei Zhang 0001 |
EMNLP | 3 |
| 2025 | Token-level Proximal Policy Optimization for Query GenerationabstractYichen Ouyang, Lu Wang, Fangkai Yang, Pu Zhao, Chenghua Huang, Jianfeng Liu, Bochen Pang, Yaming Yang, Yuefeng Zhan, Hao Sun, Qingwei Lin, Saravan Rajmohan, Weiwei Deng, Dongmei Zhang, Feng Sun. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Yichen Ouyang, Lu Wang 0029, Fangkai Yang, Pu Zhao 0004, Chenghua Huang, Bochen Pang, Yaming Yang 0001, Yuefeng Zhan, Hao Sun 0015, Qingwei Lin, Saravan Rajmohan, Dongmei Zhang 0001, Feng Sun 0008 |
EMNLP | 3 |
| 2025 | Self-Evolved Reward Learning for LLMSabstractReinforcement Learning from Human Feedback (RLHF) is a crucial technique for aligning language models with human preferences and is a key factor in the success of modern conversational models like GPT-4, ChatGPT, and Llama 2. A significant challenge in employing RLHF lies in training a reliable RM, which relies on high-quality labels. Typically, these labels are provided by human experts or a stronger AI, both of which can be costly and introduce bias that may affect the language model's responses. As models improve, human input may become less effective in enhancing their performance. This paper explores the potential of using the RM itself to generate additional training data for a more robust RM. Our experiments demonstrate that reinforcement learning from self-feedback outperforms baseline approaches.
We conducted extensive experiments with our approach on multiple datasets, such as HH-RLHF and UltraFeedback, and models including Mistral and Llama 3, comparing it against various baselines. Our results indicate that, even with a limited amount of human-labeled data, learning from self-feedback can robustly enhance the performance of the RM, thereby improving the capabilities of large language models. Chenghua Huang, Zhizhen Fan, Lu Wang 0029, Fangkai Yang, Pu Zhao 0004, Zeqi Lin, Qingwei Lin, Dongmei Zhang 0001, Saravan Rajmohan, Qi Zhang 0066 |
ICLR | 4 |
| 2025 | LettinGo: Explore User Profile Generation for Recommendation SystemabstractUser profiling is pivotal for recommendation systems, as it transforms raw user interaction data into concise and structured representations that drive personalized recommendations. While traditional embedding-based profiles lack interpretability and adaptability, recent advances with large language models (LLMs) enable text-based profiles that are semantically richer and more transparent. However, existing methods often adhere to fixed formats that limit their ability to capture the full diversity of user behaviors. In this paper, we introduce LettinGo, a novel framework for generating diverse and adaptive user profiles. By leveraging the expressive power of LLMs and incorporating direct feedback from downstream recommendation tasks, our approach avoids the rigid constraints imposed by supervised fine-tuning (SFT). Instead, we employ Direct Preference Optimization (DPO) to align the profile generator with task-specific performance, ensuring that the profiles remain adaptive and effective. LettinGo operates in three stages: (1) exploring diverse user profiles via multiple LLMs(2) evaluating profile quality based on their impact in recommendation systems, and (3) aligning the profile generation through pairwise preference data derived from task performance. Experimental results demonstrate that our framework significantly enhances recommendation accuracy, flexibility, and contextual awareness. This work enhances profile generation as a key innovation for next-generation recommendation systems. Lu Wang 0029, Fangkai Yang, Pu Zhao 0004, Yuefeng Zhan, Hao Sun 0015, Qingwei Lin, Dongmei Zhang 0001, Feng Sun 0008, Qi Zhang 0066 |
KDD (2) | 3 |
| 2025 | GenCeption: Evaluate vision LLMs with unlabeled unimodal data
Le-le Cao, Valentin Leonhard Buchner, Zineb Senane, Fangkai Yang |
Comput. Speech Lang. | 4 |
| 2024 | COIN: Chance-Constrained Imitation Learning for Safe and Adaptive Resource Oversubscription under UncertaintyabstractWe address the real problem of safe, robust, adaptive resource oversubscription in uncertain environments with our proposed novel technique of chance-constrained imitation learning. Our objective is to enhance resource efficiency while ensuring safety against congestion risk. Traditional supervised or forecasting models are ineffective in learning adaptive oversubscription policies, and conventional online optimization or reinforcement learning is difficult to deploy on real systems. Offline policy learning methods, such as Imitation Learning (IL) can leverage historical resource utilization telemetry data to learn effective policies if we can ensure robustness and safety from the underlying uncertainty in the domain, and thus the data. Our work investigates the nature of this uncertainty, how it can be quantified and proposes a novel chance-constrained IL that implicitly models such uncertainty in a principled manner via additional knowledge in the form of stochastic constraints on the associated risk, to learn provably safe and robust policies. We show empirically a substantial improvement (~ 3-4×) in capacity efficiency and congestion safety in test as well as real deployments. Lu Wang 0029, Mayukh Das, Fangkai Yang, Bo Qiao 0001, Hang Dong 0004, Chetan Bansal, Si Qin, Saravan Rajmohan, Qingwei Lin, Dongmei Zhang 0001, Qi Zhang 0066 |
CIKM | 3 |
| 2024 | Nissist: An Incident Mitigation Copilot based on Troubleshooting GuidesabstractEffective incident management is pivotal for the smooth operation of Microsoft cloud services. In order to expedite incident mitigation, service teams gather troubleshooting knowledge into Troubleshooting Guides (TSGs) accessible to On-Call Engineers (OCEs). While automated pipelines are enabled to resolve the most frequent and easy incidents, there still exist complex incidents that require OCEs’ intervention. In addition, TSGs are often unstructured and incomplete, which requires manual interpretation by OCEs, leading to on-call fatigue and decreased productivity, especially among new-hire OCEs. In this work, we propose Nissist which leverages unstructured TSGs and incident mitigation history to provide proactive incident mitigation suggestions, reducing human intervention. Leveraging Large Language Models (LLM), Nissist extracts knowledge from unstructured TSGs and incident mitigation history, forming a comprehensive knowledge base. Its multi-agent system design enhances proficiency in precisely discerning OCE intents, retrieving relevant information, and delivering systematic plans consecutively. Through our user experiments, we demonstrate that Nissist significantly reduce Time to Mitigate (TTM) in incident mitigation, alleviating operational burdens on OCEs and improving service reliability. Our webpage is available at https://aka.ms/nissist. Kaikai An, Fangkai Yang, Junting Lu, Liqun Li, Zhixing Ren, Lu Wang 0029, Pu Zhao 0004, Yu Kang 0006, Qingwei Lin, Saravan Rajmohan, Dongmei Zhang 0001, Qi Zhang 0066 |
ECAI | 2 |
| 2024 | EfficientRAG: Efficient Retriever for Multi-Hop Question AnsweringabstractZiyuan Zhuang, Zhiyang Zhang, Sitao Cheng, Fangkai Yang, Jia Liu, Shujian Huang, Qingwei Lin, Saravan Rajmohan, Dongmei Zhang, Qi Zhang. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Ziyuan Zhuang, Sitao Cheng, Fangkai Yang, Shujian Huang, Qingwei Lin, Saravan Rajmohan, Dongmei Zhang 0001, Qi Zhang 0066 |
EMNLP | 4 |
| 2024 | SELF-GUARD: Empower the LLM to Safeguard ItselfabstractZezhong Wang, Fangkai Yang, Lu Wang, Pu Zhao, Hongru Wang, Liang Chen, Qingwei Lin, Kam-Fai Wong. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Zezhong Wang 0004, Fangkai Yang, Lu Wang 0029, Pu Zhao 0004, Hongru Wang 0003, Liang Chen 0001, Qingwei Lin, Kam-Fai Wong |
NAACL-HLT | 2 |
| 2023 | Snape: Reliable and Low-Cost Computing with Mixture of Spot and On-Demand VMsabstractCloud providers often have resources that are not being fully utilized, and they may offer them at a lower cost to make up for the reduced availability of these resources. However, customers may be hesitant to use such offerings (such as spot VMs) as making trade-offs between cost and resource availability is not always straightforward. In this work, we propose Snape (Spot On-demand Perfect Mixture), an intelligent framework to optimize the cost and resource availability by dynamically mixing on-demand VMs with spot VMs. Through a detailed characterization based on real production traces, we verify that the eviction of spot VMs is predictable to some extent. Snape also leverages constrained reinforcement learning to adjust the mixture policy online. Experiments across different configurations show that Snape achieves 44% savings compared to using only on-demand VMs while maintaining 99.96% availability, which is 2.77% higher than using only spot VMs. Fangkai Yang, Lu Wang 0029, Zhenyu Xu 0003, Liqun Li, Bo Qiao 0001, Camille Couturier, Chetan Bansal, Soumya Ram, Si Qin, Íñigo Goiri, Eli Cortez, Terry Yang, Victor Rühle, Saravan Rajmohan, Qingwei Lin, Dongmei Zhang 0001 |
ASPLOS (3) | 1 |
| 2023 | Measuring Acoustics with Collaborative Multiple AgentsabstractAs humans, we hear sound every second of our life. The sound we hear is often affected by the acoustics of the environment surrounding us. For example, a spacious hall leads to more reverberation. Room Impulse Responses (RIR) are commonly used to characterize environment acoustics as a function of the scene geometry, materials, and source/receiver locations. Traditionally, RIRs are measured by setting up a loudspeaker and microphone in the environment for all source/receiver locations, which is time-consuming and inefficient. We propose to let two robots measure the environment's acoustics by actively moving and emitting/receiving sweep signals. We also devise a collaborative multi-agent policy where these two robots are trained to explore the environment's acoustics while being rewarded for wide exploration and accurate prediction. We show that the robots learn to collaborate and move to explore environment acoustics while minimizing the prediction error. To the best of our knowledge, we present the very first problem formulation and solution to the task of collaborative environment acoustics measurements with multiple agents. Yinfeng Yu, Changan Chen, Le-le Cao, Fangkai Yang, Fuchun Sun 0001 |
IJCAI | 4 |
| 2023 | Contextual Self-attentive Temporal Point Process for Physical Decommissioning Prediction of Cloud AssetsabstractAs cloud computing continues to expand globally, the need for effective management of decommissioned cloud assets in data centers becomes increasingly important. This work focuses on predicting the physical decommissioning date of cloud assets as a crucial component in reverse cloud supply chain management and data center warehouse operation. The decommissioning process is modeled as a contextual self-attentive temporal point process, which incorporates contextual information to model sequences with parallel events and provides more accurate predictions with more seen historical data. We conducted extensive offline and online experiments in 20 sampled data centers. The results show that the proposed methodology achieves the best performance compared with baselines and improves remarkable 94% prediction accuracy in online experiments. This modeling methodology can be extended to other domains with similar workflow-like processes. Fangkai Yang, Lu Wang 0029, Bo Qiao 0001, Di Weng, Xiaoting Qin, Gregory Weber, Durgesh Nandini Das, Srinivasan Rakhunathan, Ranganathan Srikanth, Qingwei Lin, Dongmei Zhang 0001 |
KDD | 1 |
| 2023 | Diffusion-Based Time Series Data Imputation for Cloud Failure Prediction at Microsoft 365abstractEnsuring reliability in large-scale cloud systems like Microsoft 365 is crucial. Cloud failures, such as disk and node failure, threaten service reliability, causing service interruptions and financial loss. Existing works focus on failure prediction and proactively taking action before failures happen. However, they suffer from poor data quality, like data missing in model training and prediction, which limits performance. In this paper, we focus on enhancing data quality through data imputation by the proposed Diffusion+, a sample-efficient diffusion model, to impute the missing data efficiently conditioned on the observed data. Experiments with industrial datasets and application practice show that our model contributes to improving the performance of downstream failure prediction. Fangkai Yang, Lu Wang 0029, Pu Zhao 0004, Bo Liu 0006, Bo Qiao 0001, Mårten Björkman, Saravan Rajmohan, Qingwei Lin, Dongmei Zhang 0001 |
ESEC/SIGSOFT FSE | 1 |
| 2023 | Learning Cooperative Oversubscription for Cloud by Chance-Constrained Multi-Agent Reinforcement LearningabstractOversubscription is a common practice for improving cloud resource utilization. It allows the cloud service provider to sell more resources than the physical limit, assuming not all users would fully utilize the resources simultaneously. However, how to design an oversubscription policy that improves utilization while satisfying some safety constraints remains an open problem. Existing methods and industrial practices are over-conservative, ignoring the coordination of diverse resource usage patterns and probabilistic constraints. To address these two limitations, this paper formulates the oversubscription for cloud as a chance-constrained optimization problem and proposes an effective Chance-Constrained Multi-Agent Reinforcement Learning (C2MARL) method to solve this problem. Specifically, C2MARL reduces the number of constraints by considering their upper bounds and leverages a multi-agent reinforcement learning paradigm to learn a safe and optimal coordination policy. We evaluate our C2MARL on an internal cloud platform and public cloud datasets. Experiments show that our C2MARL outperforms existing methods in improving utilization () under different levels of safety constraints. Junjie Sheng, Lu Wang 0029, Fangkai Yang, Bo Qiao 0001, Hang Dong 0004, Xiangfeng Wang 0001, Bo Jin 0003, Jun Wang 0006, Si Qin, Saravan Rajmohan, Qingwei Lin, Dongmei Zhang 0001 |
WWW | 3 |
| 2022 | NENYA: Cascade Reinforcement Learning for Cost-Aware Failure Mitigation at Microsoft 365abstractLarge-scale distributed systems, such as Microsoft 365's database system, require timely mitigation solutions to address failures and improve service availability and reliability. Still, mitigation actions can be costly as they may cause temporal performance degradation and even incur monetary expenses. Mitigation actions can be either administrated in a reactive fashion to contain detected failures or a proactive fashion to reduce potential failures. The proactive mitigation approach typically relies on a two-stage strategy: the prediction model will firstly identify instances (such as databases or disks) with high failure risk, then appropriate mitigation actions chosen by engineers or an automatic bandit learning model can be applied. As information is not fully shared across those two stages, important factors such as mitigation costs and states of instances are often ignored in one of those two stages. To address these issues, we propose NENYA, an end-to-end mitigation solution for a large-scale database system powered by a novel cascade reinforcement learning model. By taking the states of databases as input, NENYA directly outputs mitigation actions and is optimized based on jointly cumulative feedback on mitigation costs and failure rates. As the overwhelming majority of databases do not require mitigation actions, NENYA utilizes a novel cascade decision structure to firstly reliably filter out such databases and then focus on choosing appropriate mitigation actions for the rest. Extensive offline and online experiments have shown that our methods can outperform existing practices in reducing both failure rates of databases and mitigation costs. NENYA has been integrated into Microsoft 365, a productive platform, with sounding success. Lu Wang 0029, Pu Zhao 0004, Chuan Luo 0002, Mengna Su, Fangkai Yang, Qingwei Lin, Yingnong Dang, Hongyu Zhang 0002, Saravan Rajmohan, Dongmei Zhang 0001 |
KDD | 6 |
| 2021 | Intelligent container reallocation at Microsoft 365abstractThe use of containers in microservices has gained popularity as it facilitates agile development, resource governance, and software maintenance. Container reallocation aims to achieve workload balance via reallocating containers over physical machines. It affects the overall performance of microservice-based systems. However, container scheduling and reallocation remain an open issue due to their complexity in real-world scenarios. In this paper, we propose a novel Multi-Phase Local Search (MPLS) algorithm to optimize container reallocation. The experimental results show that our optimization algorithm outperforms state-of-the-art methods. In practice, it has been successfully applied to Microsoft 365 system to mitigate hotspot machines and balance workloads across the entire system. Bo Qiao 0001, Fangkai Yang, Chuan Luo 0002, Johnny Li, Qingwei Lin, Hongyu Zhang 0002, Mohit Datta, Andrew Zhou, Thomas Moscibroda, Saravanakumar Rajmohan, Dongmei Zhang 0001 |
ESEC/SIGSOFT FSE | 2 |
| 2020 | Group Behavior Recognition Using Attention- and Graph-Based Neural Networks
Fangkai Yang, Tetsunari Inamura, Mårten Björkman, Christopher Peters 0001 |
ECAI | 1 |
| 2020 | Impact of Trajectory Generation Methods on Viewer Perception of Robot Approaching Group BehaviorsabstractMobile robots that approach free-standing conversational groups to join them should behave in a safe and socially-acceptable way. Existing trajectory generation methods focus on collision avoidance with pedestrians, and the models that generate approach behaviors into groups are evaluated in simulation. However, it is challenging to generate approach and join trajectories that avoid collisions with group members while also ensuring that they do not invoke feelings of discomfort. In this paper, we conducted an experiment to examine the impact of three trajectory generation methods for a mobile robot to approach groups from multiple directions: a Wizard-of-Oz (WoZ) method, a procedural social-aware navigation model (PM) and a novel generative adversarial model imitating human approach behaviors (IL). Measures also compared two camera viewpoints and static versus quasi-dynamic groups. The latter refers to a group whose members change orientation and position throughout the approach task, even though the group entity remains static in the environment. This represents a more realistic but challenging scenario for the robot. We evaluate three methods with objective measurements and subjective measurements from viewer perception, and results show that WoZ and IL have comparable performance, and both perform better than PM under most conditions. Fangkai Yang, Mårten Björkman, Christopher Peters 0001 |
RO-MAN | 1 |
| 2019 | SDRL: Interpretable and Data-Efficient Deep Reinforcement Learning Leveraging Symbolic PlanningabstractDeep reinforcement learning (DRL) has gained great success by learning directly from high-dimensional sensory inputs, yet is notorious for the lack of interpretability. Interpretability of the subtasks is critical in hierarchical decision-making as it increases the transparency of black-box-style DRL approach and helps the RL practitioners to understand the high-level behavior of the system better. In this paper, we introduce symbolic planning into DRL and propose a framework of Symbolic Deep Reinforcement Learning (SDRL) that can handle both high-dimensional sensory inputs and symbolic planning. The task-level interpretability is enabled by relating symbolic actions to options.This framework features a planner – controller – meta-controller architecture, which takes charge of subtask scheduling, data-driven subtask learning, and subtask evaluation, respectively. The three components cross-fertilize each other and eventually converge to an optimal symbolic plan along with the learned subtasks, bringing together the advantages of long-term planning capability with symbolic knowledge and end-to-end reinforcement learning directly from a high-dimensional sensory input. Experimental results validate the interpretability of subtasks, along with improved data efficiency compared with state-of-the-art approaches. Daoming Lyu, Fangkai Yang, Bo Liu 0006, Steven Gustafson |
AAAI | 2 |
| 2019 | Logic-Based Sequential Decision-MakingabstractDeep reinforcement learning (DRL) has gained great success by learning directly from high-dimensional sensory inputs, yet is notorious for the lack of interpretability. Interpretability of the subtasks is critical in hierarchical decision-making as it increases the transparency of black-box-style DRL approach and helps the RL practitioners to understand the high-level behavior of the system better. In this paper, we introduce symbolic planning into DRL and propose a framework of Symbolic Deep Reinforcement Learning (SDRL) that can handle both high-dimensional sensory inputs and symbolic planning. The task-level interpretability is enabled by relating symbolic actions to options. This framework features a planner – controller – meta-controller architecture, which takes charge of subtask scheduling, data-driven subtask learning, and subtask evaluation, respectively. The three components cross-fertilize each other and eventually converge to an optimal symbolic plan along with the learned subtasks, bringing together the advantages of long-term planning capability with symbolic knowledge and end-to-end reinforcement learning directly from a high-dimensional sensory input. Experimental results validate the interpretability of subtasks, along with improved data efficiency compared with state-of-the-art approaches. Daoming Lyu, Fangkai Yang, Bo Liu 0006, Daesub Yoon |
AAAI | 2 |
| 2019 | Criticality-based Collision Avoidance Prioritization for Crowd NavigationabstractGoal directed agent navigation in crowd simulations involves a complex decision making process. An agent must avoid all collisions with static or dynamic obstacles (such as other agents) and keep a trajectory faithful to its target at the same time. This seemingly global optimization problem can be broken down into smaller local optimization problems by looking at a concept of criticality. Our method resolves critical agents - agents that are likely to come within collision range of each other - in order of priority using a Particle Swarm Optimization scheme. The resolution involves altering the velocities of agents to avoid criticality. Results from our method show that the navigation problem can be solved in several important test cases with minimal number of collisions and minimal deviation to the target direction. We prove the efficiency and correctness of our method by comparing it to four other well-known algorithms, and performing evaluations on them based on various quality measures. Himangshu Saikia, Fangkai Yang, Christopher Peters 0001 |
HAI | 2 |
| 2019 | App-LSTM: Data-driven Generation of Socially Acceptable Trajectories for Approaching Small Groups of AgentsabstractWhile many works involving human-agent interactions have focused on individuals or crowds, modelling interactions on the group scale has not been considered in depth. Simulation of interactions with groups of agents is vital in many applications, enabling more comprehensive and realistic behavior encompassing all possibilities between crowd and individual levels. In this paper, we propose a novel neural network App-LSTM to generate the approach trajectory of an agent towards a small free-standing conversational group of agents. The App-LSTM model is trained on a dataset of approach behaviors towards the group. Since current publicly available datasets for these encounters are limited, we develop a social-aware navigation method as a basis for creating a semi-synthetic dataset composed of a mixture of real and simulated data representing safe and socially-acceptable approach trajectories. Via a group interaction module, App-LSTM then captures the position and orientation features of the group and refines the current state of the approaching agent iteratively to better focus on the current intention of group members. We show our App-LSTM outperforms baseline methods in generating approaching group trajectories. Fangkai Yang, Christopher Peters 0001 |
HAI | 1 |
| 2019 | Task-Motion Planning with Reinforcement Learning for Adaptable Mobile Service RobotsabstractTask-motion planning (TMP) addresses the problem of efficiently generating executable and low-cost task plans in a discrete space such that the (initially unknown) action costs are determined by motion plans in a corresponding continuous space. A task-motion plan for a mobile service robot that behaves in a highly dynamic domain can be sensitive to domain uncertainty and changes, leading to suboptimal behaviors or execution failures. In this paper, we propose a novel framework, TMP-RL, which is an integration of TMP and reinforcement learning (RL), to solve the problem of robust TMP in dynamic and uncertain domains. The robot first generates a low-cost, feasible task-motion plan by iteratively planning in the discrete space and updating relevant action costs evaluated by the motion planner in continuous space. During execution, the robot learns via model-free RL to further improve its task-motion plans. RL enables adaptability to the current domain, but can be costly with regards to experience; using TMP, which does not rely on experience, can jump-start the learning process before executing in the real world. TMP-RL is evaluated in a mobile service robot domain where the robot navigates in an office area, showing significantly improved adaptability to unseen domain dynamics over TMP and task planning (TP)-RL methods. Yuqian Jiang, Fangkai Yang, Shiqi Zhang 0001, Peter Stone 0001 |
IROS | 2 |
| 2019 | Learning Socially Appropriate Robot Approaching Behavior Toward Groups using Deep Reinforcement LearningabstractDeep reinforcement learning has recently been widely applied in robotics to study tasks such as locomotion and grasping, but its application to social human-robot interaction (HRI) remains a challenge. In this paper, we present a deep learning scheme that acquires a prior model of robot approaching behavior in simulation and applies it to real-world interaction with a physical robot approaching groups of humans. The scheme, which we refer to as Staged Social Behavior Learning (SSBL), considers different stages of learning in social scenarios. We learn robot approaching behaviors towards small groups in simulation and evaluate the performance of the model using objective and subjective measures in a perceptual study and a HRI user study with human participants. Results show that our model generates more socially appropriate behavior compared to a state-of-the-art model. Alex Yuan Gao, Fangkai Yang, Martin Frisk, Daniel Hemandez, Christopher Peters 0001, Ginevra Castellano |
RO-MAN | 2 |
| 2019 | AppGAN: Generative Adversarial Networks for Generating Robot Approach Behaviors into Small Groups of PeopleabstractRobots that navigate to approach free-standing conversational groups should do so in a safe and socially acceptable manner. This is challenging since it not only requires the robot to plot trajectories that avoid collisions with members of the group, but also to do so without making those in the group feel uncomfortable, for example, by moving too close to them or approaching them from behind. Previous trajectory prediction models focus primarily on formations of walking pedestrians, and those models that do consider approach behaviours into free-standing conversational groups typically have handcrafted features and are only evaluated via simulation methods, limiting their effectiveness. In this paper, we propose AppGAN, a novel trajectory prediction model capable of generating trajectories into free-standing conversational groups trained on a dataset of safe and socially acceptable paths. We evaluate the performance of our model with state-of-the-art trajectory prediction methods on a semi-synthetic dataset. We show that our model outperforms baselines by taking advantage of the GAN framework and our novel group interaction module. Fangkai Yang, Christopher Peters 0001 |
RO-MAN | 1 |
| 2019 | Introduction to the 35th International Conference on Logic Programming Special Issue
Esra Erdem 0001, Andrea Formisano 0001, Germán Vidal, Fangkai Yang |
Theory Pract. Log. Program. | 4 |
| 2018 | PEORL: Integrating Symbolic Planning and Hierarchical Reinforcement Learning for Robust Decision-MakingabstractReinforcement learning and symbolic planning have both been used to build intelligent autonomous agents. Reinforcement learning relies on learning from interactions with real world, which often requires an unfeasibly large amount of experience. Symbolic planning relies on manually crafted symbolic knowledge, which may not be robust to domain uncertainties and changes. In this paper we present a unified framework PEORL that integrates symbolic planning with hierarchical reinforcement learning (HRL) to cope with decision-making in dynamic environment with uncertainties. Symbolic plans are used to guide the agent's task execution and learning, and the learned experience is fed back to symbolic knowledge to improve planning. This method leads to rapid policy search and robust symbolic plans in complex domains. The framework is tested on benchmark domains of HRL. Fangkai Yang, Daoming Lyu, Bo Liu 0006, Steven Gustafson |
IJCAI | 1 |
| 2018 | Effects of Posture and Embodiment on Social Distance in Human-Agent Interaction in Mixed RealityabstractMixed reality offers new potentials for social interaction experiences with virtual agents. In addition, it can be used to experiment with the design of physical robots. However, while previous studies have investigated comfortable social distances between humans and artificial agents in real and virtual environments, there is little data with regards to mixed reality environments. In this paper, we conducted an experiment in which participants were asked to walk up to an agent to ask a question, in order to investigate the social distances maintained, as well as the subject's experience of the interaction. We manipulated both the embodiment of the agent (robot vs. human and virtual vs. physical) as well as closed vs. open posture of the agent. The virtual agent was displayed using a mixed reality headset. Our experiment involved 35 participants in a within-subject design. We show that, in the context of social interactions, mixed reality fares well against physical environments, and robots fare well against humans, barring a few technical challenges. Theofronia Androulakaki, Alex Yuan Gao, Fangkai Yang, Himangshu Saikia, Christopher Peters 0001, Gabriel Skantze |
IVA | 4 |
| 2018 | Pedestrian simulation as multi-objective reinforcement learningabstractModelling and simulation of pedestrian crowds require agents to reach pre-determined goals and avoid collisions with static obstacles and dynamic pedestrians, while maintaining natural gait behaviour. We model pedestrians as autonomous, learning, and reactive agents employing Reinforcement Learning (RL). Typical RL-based agent simulations suffer poor generalization due to handcrafted reward function to ensure realistic behaviour. In this work, we model pedestrians in a modular framework integrating navigation and collision-avoidance tasks as separate modules. Each such module consists of independent state-spaces and rewards, but with shared action-spaces. Empirical results suggest that such modular framework learning models can show satisfactory performance without tuning parameters, and we compare it with the state-of-art crowd simulation methods. Naresh Balaji Ravichandran, Fangkai Yang, Christopher Peters 0001, Anders Lansner, Pawel Andrzej Herman |
IVA | 2 |
| 2018 | Who are my neighbors?: A perception model for selecting neighbors of pedestrians in crowdsabstractPedestrian trajectory prediction is a challenging problem. One of the aspects that makes it so challenging is the fact that the future positions of an agent are not only determined by its previous positions, but also by the interaction of the agent with its neighbors. Previous methods, like Social Attention have considered the interactions with all agents as neighbors. However, this ends up assigning high attention weights to agents who are far away from the queried agent and/or moving in the opposite direction, even though, such agents might have little to no impact on the queried agent's trajectory. Furthermore, trajectory prediction of a queried agent involving all agents in a large crowded scenario is not efficient. In this paper, we propose a novel approach for selecting neighbors of an agent by modeling its perception as a combination of a location and a locomotion model. We demonstrate the performance of our method by comparing it with the existing state-of-the-art method on publicly available datasets. The results show that our neighbor selection model overall improves the accuracy of trajectory prediction and enables prediction in scenarios with large numbers of agents in which other methods do not scale well. Fangkai Yang, Himangshu Saikia, Christopher Peters 0001 |
IVA | 1 |
| 2018 | Do you see groups?: The impact of crowd density and viewpoint on the perception of groupsabstractAgent-based crowd simulation in virtual environments is of great utility in a variety of domains, from the entertainment industry to serious applications including mobile robots and swarms. Many studies of crowd behavior simulations do not consider the fact that people tend to congregate in smaller social gatherings, such as friends, or families, rather than walking alone. Based on a real-time crowd simulator which has been implemented as a unilateral incompressible fluid and augmented with group behaviors, a perceptual study was conducted to determine the impact of groups on the perception of the crowds at various densities from different camera views. If it is not possible to see groups under certain circumstances, then it may not be necessary to simulate them, to reduce the amount of calculations, an important issue in real-time simulations. This study provides researchers with a proper reference to design better algorithms to simulate realistic behaviors. Fangkai Yang, Jack Shabo, Adam Qureshi, Christopher Peters 0001 |
IVA | 1 |
| 2017 | A Virtual Poster Presenter Using Mixed Reality
Vanya Avramova, Fangkai Yang, Christopher Peters 0001, Gabriel Skantze |
IVA | 2 |
| 2016 | Planning with Task-Oriented Knowledge Acquisition for a Service Robot
Fangkai Yang |
IJCAI | 2 |
| 2015 | Mobile Robot Planning Using Action Language BC with an Abstraction Hierarchy
Shiqi Zhang 0001, Fangkai Yang, Piyush Khandelwal, Peter Stone 0001 |
LPNMR | 2 |
| 2014 | The Semantics of Gringo and Infinitary Propositional Formulas
Amelia Harrison, Vladimir Lifschitz, Fangkai Yang |
KR | 3 |
| 2013 | Action Language BC: Preliminary Report
Joohyung Lee 0002, Vladimir Lifschitz, Fangkai Yang |
IJCAI | 3 |
| 2013 | Lloyd-Topor completion and general stable modelsabstractAbstract We investigate the relationship between the generalization of program completion defined in 1984 by Lloyd and Topor and the generalization of the stable model semantics introduced recently by Ferraris et al. The main theorem can be used to characterize, in some cases, the general stable models of a logic program by a first-order formula. The proof uses Truszczynski's stable model semantics of infinitary propositional formulas. Vladimir Lifschitz, Fangkai Yang |
Theory Pract. Log. Program. | 2 |
| 2013 | Representing Actions in Logic-based Languages
Fangkai Yang |
Theory Pract. Log. Program. | 1 |
| 2012 | Representing first-order causal theories by logic programsabstractAbstract Nonmonotonic causal logic, introduced by McCain and Turner (McCain, N. and Turner, H. 1997. Causal theories of action and change. In Proceedings of National Conference on Artificial Intelligence (AAAI), Stanford, CA, 460–465) became the basis for the semantics of several expressive action languages. McCain's embedding of definite propositional causal theories into logic programming paved the way to the use of answer set solvers for answering queries about actions described in such languages. In this paper we extend this embedding to nondefinite theories and to the first-order causal logic. Paolo Ferraris, Joohyung Lee 0002, Yuliya Lierler, Vladimir Lifschitz, Fangkai Yang |
Theory Pract. Log. Program. | 5 |
| 2012 | Relational theories with null values and non-herbrand stable modelsabstractAbstract Generalized relational theories with null values in the sense of Reiter are first-order theories that provide a semantics for relational databases with incomplete information. In this paper we show that any such theory can be turned into an equivalent logic program, so that models of the theory can be generated using computational methods of answer set programming. As a step towards this goal, we develop a general method for calculating stable models under the domain closure assumption but without the unique name assumption. Vladimir Lifschitz, Karl Pichotta, Fangkai Yang |
Theory Pract. Log. Program. | 3 |
| 2010 | Translating First-Order Causal Theories into Answer Set Programming
Vladimir Lifschitz, Fangkai Yang |
JELIA | 2 |
| 2009 | A Distance-Based Operator to Revising Ontologies in DL SHOQ
Fangkai Yang, Guilin Qi, Zhisheng Huang |
ECSQARU | 1 |