EDBT 2026 Demo / reviewers in the wild / expert
Yan Ding 0002
dblp:57/4533-2
· DBLP profile ↗
23ranked-venue papers
5as first author
18since 2021 · last 2025
0000-0002-7949-4351ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 14 · 2 first-author · 13 since 2021Systems, architecture and hardware · 8 · 2 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021Computer networks · 2 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 first-authorSoftware engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Learning 2D Invariant Affordance Knowledge for 3D Affordance Groundingabstract3D Object Affordance Grounding aims to predict the functional regions on a 3D object and has laid the foundation for a wide range of applications in robotics. Recent advances tackle this problem via learning a mapping between 3D regions and a single human-object interaction image. However, the geometric structure of the 3D object and the object in the human-object interaction image are not always consistent, leading to poor generalization. To address this issue, we propose to learn generalizable invariant affordance knowledge from multiple human-object interaction images within the same affordance category. Specifically, we introduce the Multi-Image Guided Invariant-Feature-Aware 3D Affordance Grounding (MIFAG) framework. It grounds 3D object affordance regions by identifying common interaction patterns across multiple human-object interaction images. First, the Invariant Affordance Knowledge Extraction Module (IAM) utilizes an iterative updating strategy to gradually extract aligned affordance knowledge from multiple images and integrate it into an affordance dictionary. Then, the Affordance Dictionary Adaptive Fusion Module (ADM) learns comprehensive point cloud representations that consider all affordance candidates in multiple images. Besides, the Multi-Image and Point Affordance (MIPA) benchmark is constructed and our method outperforms existing state-of-the-art methods on various experimental comparisons. Xianqiang Gao 0001, Pingrui Zhang, Delin Qu, Dong Wang 0028, Zhigang Wang 0002, Yan Ding 0002, Bin Zhao 0001 |
AAAI | 6 |
| 2025 | SKE-Layout: Spatial Knowledge Enhanced Layout Generation with LLMsabstractGenerating layouts from textual descriptions by large language models (LLMs) plays a crucial role in precise spatial reasoning-induced domains such as robotic object rearrangement and text-to-image generation. However, current methods face challenges in limited real-world examples, handling diverse layout descriptions and varying levels of granularity. To address these issues, a novel framework named Spatial Knowledge Enhanced Layout (SKE-Layout), is introduced. SKE-Layout integrates mixed spatial knowledge sources, leveraging both real and synthetic data to enhance spatial contexts. It utilizes diverse representations tailored to specific tasks and employs contrastive learning and multitask learning techniques for accurate spatial knowledge retrieval. This framework generates more accurate and fine-grained visual layouts for object rearrangement and text-to-image generation tasks, achieving improvements of 5%-30% compared to existing methods. Nieqing Cao, Yan Ding 0002, Mengying Xie, Fuqiang Gu, Chao Chen 0004 |
CVPR | 3 |
| 2025 | Think Small, Act Big: Primitive Prompt Learning for Lifelong Robot ManipulationabstractBuilding a lifelong robot that can effectively leverage prior knowledge for continuous skill acquisition remains significantly challenging. Despite the success of experience replay and parameter-efficient methods in alleviating catastrophic forgetting problem, naively applying these methods causes a failure to leverage the shared primitives between skills. To tackle these issues, we propose Primitive Prompt Learning (PPL), to achieve lifelong robot manipulation via reusable and extensible primitives. Within our two stage learning scheme, we first learn a set of primitive prompts to represent shared primitives through multi-skills pre-training stage, where motion-aware prompts are learned to capture semantic and motion shared primitives across different skills. Secondly, when acquiring new skills in lifelong span, new prompts are concatenated and optimized with frozen pretrained prompts, boosting the learning via knowledge transfer from old skills to new ones. For evaluation, we construct a large-scale skill dataset and conduct extensive experiments in both simulation and real-world tasks, demonstrating PPL’s superior performance over state-of-the-art methods. Yuanqi Yao, Siao Liu, Haoming Song, Delin Qu, Yan Ding 0002, Bin Zhao 0001, Zhigang Wang 0002, Xuelong Li 0001, Dong Wang 0028 |
CVPR | 6 |
| 2025 | MoMa-Pos: An Efficient Object-Kinematic-Aware Base Placement Determination Framework for Mobile Manipulation
Beichen Shao, Nieqing Cao, Yan Ding 0002, Fuqiang Gu, Chao Chen 0004 |
ICA3PP (3) | 3 |
| 2025 | OVA-Fields: Weakly Supervised Open-Vocabulary Affordance Fields for Robot Operational Part Detection
Heng Su, Mengying Xie, Nieqing Cao, Yan Ding 0002, Beichen Shao, Xianlei Long, Fuqiang Gu, Chao Chen 0004 |
ICCV | 4 |
| 2025 | MoMa-Kitchen: A 100K+ Benchmark for Affordance-Grounded Last-Mile Navigation in Mobile ManipulationabstractIn mobile manipulation, navigation and manipulation are often treated as separate problems, resulting in a significant gap between merely approaching an object and engaging with it effectively. Many navigation approaches primarily define success by proximity to the target, often overlooking the necessity for optimal positioning that facilitates subsequent manipulation. To address this, we introduce MoMa-Kitchen, a benchmark dataset comprising over 100k samples that provide training data for models to learn optimal final navigation positions for seamless transition to manipulation. Our dataset includes affordance-grounded floor labels collected from diverse kitchen environments, in which robotic mobile manipulators of different models attempt to grasp target objects amidst clutter. Using a fully automated pipeline, we simulate diverse real-world scenarios and generate affordance labels for optimal manipulation positions. Visual data are collected from RGB-D inputs captured by a first-person view camera mounted on the robotic arm, ensuring consistency in viewpoint during data collection. We also develop a lightweight baseline model, NavAff, for navigation affordance grounding that demonstrates promising performance on the MoMa-Kitchen benchmark. Our approach enables models to learn affordance-based final positioning that accommodates different arm types and platform heights, thereby paving the way for more robust and generalizable integration of navigation and manipulation in embodied AI. Project page: \href{https://momakitchen.github.io/}{https://momakitchen.github.io/}. Pingrui Zhang, Xianqiang Gao 0001, Kehui Liu, Dong Wang 0028, Zhigang Wang 0002, Bin Zhao 0001, Yan Ding 0002, Xuelong Li 0001 |
ICCV | 8 |
| 2025 | ORLA*: Mobile Manipulator-Based Object Rearrangement with Lazy AabstractEffectively performing object rearrangement is an essential skill for mobile manipulators, e.g., setting up a dinner table. A key challenge in such problems is deciding an appropriate ordering to effectively untangle object-object dependencies while considering the necessary motions for realizing manipulation tasks (e.g., pick and place). Computing time-optimal multi-object rearrangement solutions for mobile manipulators remains a largely untapped research direction. In this work, we propose ORLA*, which leverages delayed/lazy evaluation in searching for a high-quality object pick-n-place sequence that considers both end-effector and mobile robot base travel. ORLA* readily handles multi-layered rearrangement tasks powered by learning-based stability predictions. Employing an optimal solver for finding temporary locations for displacing objects, ORLA* can achieve global optimality. Through extensive simulation and ablation study, we confirm the effectiveness of ORLA* delivering quality solutions for challenging rearrangement instances. Supplementary materials are available at: gaokai15.github.io/ORLA-Star/ Zhaxizhuoma, Yan Ding 0002, Shiqi Zhang 0001, Jingjin Yu |
ICRA | 3 |
| 2025 | AlignBot: Aligning VLM-Powered Customized Task Planning with User Reminders Through Fine-Tuning for Household RobotsabstractThis paper presents AlignBot, a novel framework designed to optimize VLM-powered customized task planning for household robots by effectively aligning with user reminders. In domestic settings, aligning task planning with user reminders poses significant challenges due to the limited quantity, diversity, and multimodal nature of the reminders. To address these challenges, AlignBot employs a fine-tuned LLaVA-7B model, functioning as an adapter for GPT-40. This adapter model internalizes diverse forms of user reminders-such as personalized preferences, corrective guidance, and contextual assistance-into structured instruction-formatted cues that prompt GPT-40 in generating customized task plans. Additionally, AlignBot integrates a dynamic retrieval mechanism that selects task-relevant historical successes as prompts for GPT-40, further enhancing task planning accuracy. To validate the effectiveness of AlignBot, experiments are conducted in real-world household environments, which are constructed within the laboratory to replicate typical household settings. A multimodal dataset with over 1,500 entries derived from volunteer reminders is used for training and evaluation. The results demonstrate that AlignBot significantly improves customized task planning, outperforming existing LLM- and VLM-powered planners by interpreting and aligning with user reminders, achieving 86.8 % success rate compared to the vanilla GPT-40 baseline at 21.6%, reflecting a 65% improvement and over four times greater effectiveness. Supplementary materials are available at: https://yding25.com/AlignBot/ Zhaxizhuoma, Pengan Chen, Ziniu Wu, Dong Wang 0028, Peng Zhou 0018, Nieqing Cao, Yan Ding 0002, Bin Zhao 0001, Xuelong Li 0001 |
ICRA | 8 |
| 2025 | DualCLIP: Bridging 3D Geometry and Multimodal Semantics for Robotic PerceptionabstractCurrent approaches to integrating CLIP into language-driven robotics face a fundamental dilemma: While robotic implementations overlook cutting-edge 3D classification adaptations of CLIP, existing 3D-oriented CLIP methods prove inadequate for interpreting color-critical instructions prevalent in manipulation tasks. We resolve this through DualCLIP, a contrastive multimodal fusion framework that hierarchically integrates depth-aligned CLIP encoders. Our approach first aligns depth and CLIP RGB encoders using synthetic RGB-D pairs, then performs multimodal fusion via contrastive learning with language-triplet optimization. This joint training preserves 3D geometric coherence and color semantics. Evaluations demonstrate DualCLIP’s combined strength — surpassing CLIP2Point in 3D classification while showing promising improvements for CLIPORT in color-sensitive robotic manipulation. This work establishes a paradigm for translating vision-language models into 3D-aware robotic systems without compromising task-specific modality sensitivity. Yinghao Liu, Penglin Dai, Yan Ding 0002, Nieqing Cao |
IROS | 3 |
| 2025 | BestMan: a modular mobile manipulator platform for embodied AI with unified simulation-hardware APIs
Kui Yang, Nieqing Cao, Beichen Shao, Yan Ding 0002, Chao Chen 0004 |
Frontiers Comput. Sci. | 5 |
| 2025 | A Coarse-to-Fine Robotic Fabric Alignment System Integrating Visual Servoing and Admittance ControlabstractFabric alignment is essential to key production processes such as cutting, sewing, and fusing in garment manufacturing. Traditionally, this task has relied heavily on the dexterity and expertise of skilled human workers. Although automated systems have been introduced, they often lack the flexibility required for complex alignment tasks. In this paper, we present a novel robotic fabric alignment framework that fully automates the process with high precision and adaptability. First, we propose a coarse-to-fine alignment strategy, where an initial imprecise target position is roughly computed based on a basic perception module and eye-to-hand calibration. This is followed by a sliding mode control (SMC)-based visual servoing approach (in an eye-in-hand configuration) to ensure a close-up view of feedback features for the fine alignment process. Additionally, we consider system disturbances estimated by a fuzzy logic system (FLS) and combine it with the controller to further enhance the system’s robustness. Finally, we developed an advanced end-effector equipped with force/torque (F/T) sensors and air-powered needle grippers for gentle fabric manipulation using admittance control. We validate our framework through a series of experiments that demonstrate its effectiveness in fabric alignment tasks. Jiaming Qi, Liang Lu 0005, Lei Yang 0048, Yan Ding 0002, Pai Zheng, David Navarro-Alarcon, Jia Pan 0001, Peng Zhou 0018 |
IEEE Trans Autom. Sci. Eng. | 4 |
| 2024 | Joint Optimization for Quality Selection and Resource Allocation of Live Video Streaming in Internet of VehiclesabstractLive Video Streaming (LVS) services are critical in supporting real-time applications in Internet of Vehicles (IoV) by transmitting real-time generated video content from streaming server to vehicles. Due to restricted spectrum resources and high vehicle mobility, LVS suffers from notable performance degradation. Moreover, existing strategies such as buffer size control and edge caching, are designed for video-on-demand service, which is ineffective for LVS in IoV. Accordingly, we investigate the problem of LVS-IoV by synthesizing multicasting and Scalable Video Coding-based encoding with the goal of maximizing Quality of Experience (QoE), which is defined as the weighted sum of video quality, rebuffering time, and quality variation. The LVS-IoV is decoupled into three sub-problems: vehicle grouping, quality selection, and resource allocation. Firstly, we propose a K-means-based vehicle grouping method that considers geographical distribution, velocity, and dynamic channels. Secondly, we determine the quality selection of each group based on the Value Decomposition Network for maximizing overall video quality. This network utilizes global value function decomposition and centralized training to achieve fast convergence, followed by distributed execution. Lastly, we propose a sub-gradient algorithm to achieve optimal resource allocation. We build simulation model and perform extensive evaluation, which demonstrates its superiority compared to other competitive methods. Penglin Dai, Meiting Wu, Ke Li 0020, Xiao Wu 0001, Yan Ding 0002 |
IEEE Trans. Serv. Comput. | 5 |
| 2023 | Symbolic State Space Optimization for Long Horizon Mobile Manipulation PlanningabstractIn existing task and motion planning (TAMP) research, it is a common assumption that experts manually specify the state space for task-level planning. A well-developed state space enables the desirable distribution of limited computational resources between task planning and motion planning. However, developing such task-level state spaces can be non-trivial in practice. In this paper, we consider a long horizon mobile manipulation domain including repeated navigation and manipulation. We propose Symbolic State Space Optimization (S3O) for computing a set of abstracted locations and their 2D geometric groundings for generating task-motion plans in such domains. Our approach has been extensively evaluated in simulation and demonstrated on a real mobile manipulator working on clearing up dining tables. Results show the superiority of the proposed method over TAMP baselines in task completion rate and execution time. Xiaohan Zhang 0002, Yan Ding 0002, Yuqian Jiang, Yuke Zhu, Peter Stone 0001, Shiqi Zhang 0001 |
IROS | 3 |
| 2023 | Task and Motion Planning with Large Language Models for Object RearrangementabstractMulti-object rearrangement is a crucial skill for service robots, and commonsense reasoning is frequently needed in this process. However, achieving commonsense arrangements requires knowledge about objects, which is hard to transfer to robots. Large language models (LLMs) are one potential source of this knowledge, but they do not naively capture information about plausible physical arrangements of the world. We propose LLM-GROP, which uses prompting to extract commonsense knowledge about semantically valid object configurations from an LLM and instantiates them with a task and motion planner in order to generalize to varying scene geometry. LLM-GROP allows us to go from natural-language commands to human-aligned object rearrangement in varied environments. Based on human evaluations, our approach achieves the highest rating while outperforming competitive baselines in terms of success rate while maintaining comparable cumulative action costs. Finally, we demonstrate a practical implementation of LLM-GROP on a mobile manipulator in real-world scenarios. Supplementary materials are available at: https://sites.google.com/view/llm-grop Yan Ding 0002, Xiaohan Zhang 0002, Chris Paxton 0001, Shiqi Zhang 0001 |
IROS | 1 |
| 2023 | Learning to reason about contextual knowledge for planning under uncertaintyabstractSequential decision-making (SDM) methods enable AI agents to compute an action policy toward achieving long-term goals under uncertainty. Existing research has shown that contextual knowledge in declarative forms can be used for improving the performance of SDM methods. However, the contextual knowledge from people tends to be incomplete and sometimes inaccurate, which greatly limits the applicability of knowledge-based SDM methods. In this paper, we develop a novel algorithm for knowledge-based SDM, called PERIL, that learns from interaction experience to reason about contextual knowledge, as applied to urban driving scenarios. Experiments have been conducted using CARLA, a widely used autonomous driving simulator. Results demonstrate PERIL’s superiority in comparison to existing knowledge-based SDM baselines. Cheng Cui, Saeid Amiri, Yan Ding 0002, Xingyue Zhan, Shiqi Zhang 0001 |
UAI | 3 |
| 2022 | Visually Grounded Task and Motion Planning for Mobile ManipulationabstractTask and motion planning (TAMP) algorithms aim to help robots achieve task-level goals, while maintaining motion-level feasibility. This paper focuses on TAMP domains that involve robot behaviors that take extended periods of time (e.g., long-distance navigation). In this paper, we develop a visual grounding approach to help robots probabilistically evaluate action feasibility, and introduce a TAMP algorithm, called GROP, that optimizes both feasibility and efficiency. We have collected a dataset that includes 96, 000 simulated trials of a robot conducting mobile manipulation tasks, and then used the dataset to learn to ground symbolic spatial relationships for action feasibility evaluation. Compared with competitive TAMP baselines, GROP exhibited a higher task-completion rate while maintaining lower or comparable action costs. In addition to these extensive experiments in simulation, GROP is fully implemented and tested on a real robot system. Xiaohan Zhang 0002, Yan Ding 0002, Yuke Zhu, Peter Stone 0001, Shiqi Zhang 0001 |
ICRA | 3 |
| 2022 | A Force-Directed Approach to Seeking Route Recommendation in Ride-on-Demand Service Using Multi-Source Urban DataabstractThe rapidly-growing business of ride-on-demand (RoD) service such as Uber, Lyft and Didi proves the effectiveness of their new service model – using mobile apps and dynamic pricing to coordinate between drivers, passengers and the service provider, to manipulate the supply and demand, and to improve service responsiveness as well as quality. Despite its success, dynamic pricing creates a new problem for drivers: how to seek for passengers to maximize revenue under dynamic prices. Seeking route recommendation has already been studied extensively in traditional taxi service, but most studies do not consider the effects of taxis and passengers on the seeking taxi simultaneously. Further, in RoD service it is necessary to consider more factors such as dynamic prices, the status of other transportation services, etc. In this paper, we employ a force-directed approach to model, by analogy, the relationship between vacant cars and passengers as that between positive and negative charges in electrostatic field. We extract features from multi-source urban data to describe dynamic prices, the status of RoD, taxi and public transportation services, and incorporate them into our model. The model is then used in route recommendation in every intersection so that a driver in a vacant RoD car knows which road segment to take next. We conduct extensive experiments based on our multi-source urban data, including RoD service operational data, taxi GPS trajectory data and public transportation distribution data, and results not only show that our approach outperforms existing baselines, but also justify the need to incorporate multi-source urban data and dynamic prices. Suiming Guo, Chao Chen 0004, Jingyuan Wang 0001, Yan Ding 0002, Yaxiao Liu, Ke Xu 0002, Zhiwen Yu 0001, Daqing Zhang 0001 |
IEEE Trans. Mob. Comput. | 4 |
| 2021 | Task and Situation Structures for Case-Based Planning
Tavan Eftekhar, Chad Esselink, Yan Ding 0002, Shiqi Zhang 0001 |
ICCBR | 4 |
| 2020 | Task-Motion Planning for Safe and Efficient Urban DrivingabstractAutonomous vehicles need to plan at the task level to compute a sequence of symbolic actions, such as merging left and turning right, to fulfill people's service requests, where efficiency is the main concern. At the same time, the vehicles must compute continuous trajectories to perform actions at the motion level, where safety is the most important. Task-motion planning in autonomous driving faces the problem of maximizing task-level efficiency while ensuring motion-level safety. To this end, we develop algorithm Task-Motion Planning for Urban Driving (TMPUD) that, for the first time, enables the task and motion planners to communicate about the safety level of driving behaviors. TMPUD has been evaluated using a realistic urban driving simulation platform. Results suggest that TMPUD performs significantly better than competitive baselines from the literature in efficiency, while ensuring the safety of driving behaviors. Yan Ding 0002, Xiaohan Zhang 0002, Xingyue Zhan, Shiqi Zhang 0001 |
IROS | 1 |
| 2020 | TrajCompressor: An Online Map-matching-based Trajectory Compression Framework Leveraging Vehicle Heading Direction and ChangeabstractMassive and redundant vehicle trajectory data are continuously sent to the data center via vehicle-mounted GPS devices, causing a number of sustainable issues, such as storage, communication, and computation. Online trajectory compression becomes a promising way to alleviate these issues. In this paper, we present an online trajectory compression framework running under the mobile environment. The framework consists of two phases, i.e., online trajectory mapping and trajectory compression. In the phase of online trajectory mapping, we develop a light-weighted yet efficient map matcher, namely, Spatial-Directional Matching (SD-Matching), to align the noisy and sparse GPS points upon the underlying road network, which fully explores the usage of vehicle heading direction collected from the GPS trajectory data. In the phase of online trajectory compression, we propose a novel compressor based on the heading change at intersections, namely, Heading Change Compression (HCC), aiming at finding a concise and compact trajectory representation. Finally, we conduct experiments to evaluate the effectiveness and efficiency of the proposed framework using real-world datasets in the city of Beijing, China. We further deploy the system in the real world in the city of Chongqing, China. The experimental results demonstrate that: 1) the SD-Matching algorithm achieves a higher mean accuracy but consumes less time than the state-of-the-art algorithm, namely, Spatial-Temporal Matching (ST-Matching) and 2) the HCC algorithm also outperforms baselines in trading-off compression ratio and computation time. Chao Chen 0004, Yan Ding 0002, Xuefeng Xie, Shu Zhang 0003, Zhu Wang 0001, Liang Feng 0001 |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2018 | Fuel Consumption Estimation of Potential Driving Paths by Leveraging Online Route APIs
Yan Ding 0002, Chao Chen 0004, Xuefeng Xie, Zhikai Yang |
GPC | 1 |
| 2018 | An Online Trajectory Compression System Applied to Resource-Constrained GPS Devices in VehiclesabstractThe raw vehicle trajectory data gathered by GPS devices is typically large and needs to be compressed online. However, GPS devices have limited resources, and cannot afford such burdensome task. To alleviate this issue, we design an online trajectory compression system consisting of Trajectory Mapping, Trajectory Compressing and Front-End Visualizer, which is implemented in the mobile phone to migrate the computation burdens. The proposed trajectory compression method does not need extra data during compressing suitable for online applications. Experiment results demonstrate our system has excellent performances regarding effectiveness, efficiency and so on. Yan Ding 0002, Chao Chen 0004, Xuefeng Xie, Kai Liu 0001, Liang Feng 0001 |
SECON | 1 |
| 2017 | GreenPlanner: Planning personalized fuel-efficient driving routes using multi-sourced urban dataabstractGreenhouse gas emission by the increasing number of vehicles have become a significant problem in modern cities. To save energy and protect environment, recommending fuel-efficient routes to drivers becomes a promising way to alleviate this issue. To this end, in this paper, we present a novel fuel-efficient path-planning framework called GreenPlanner, which contains two phases. In the first phase, we build a personalized fuel consumption model (PFCM) for each driver, based on the individual driving behaviors and the physical features (e.g., traffic lights, stop signs, road network topology) along the routes. In the second phase, with the real-time traffic information collected via the mobile crowdsensing manner, we are able to estimate and compare the cost fuel among different routes for a given driver, and recommend him/her with the most fuel-efficient one. We evaluate the two-phase framework using the real-world datasets, consisting of road network, POI, the GPS trajectory data and the OBD-II data generated by 559 taxis in one day in the city of Beijing, China. Experimental results demonstrate that, compared to the baseline models, the proposed model achieves the best accuracy, with a mean fuel consumption error of less 7% for paths longer than 10 km. Moreover, users could save about 20% fuel consumption on average if driving along our suggested routes in our case studies. Yan Ding 0002, Chao Chen 0004, Shu Zhang 0003, Bin Guo 0001, Zhiwen Yu 0001, Yasha Wang |
PerCom | 1 |