EDBT 2026 Demo / reviewers in the wild / expert
Nieqing Cao
dblp:167/1813
· DBLP profile ↗
6ranked-venue papers
0as first author
6since 2021 · last 2025
0000-0003-3414-4603ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4 · 4 since 2021Systems, architecture and hardware · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | SKE-Layout: Spatial Knowledge Enhanced Layout Generation with LLMsabstractGenerating layouts from textual descriptions by large language models (LLMs) plays a crucial role in precise spatial reasoning-induced domains such as robotic object rearrangement and text-to-image generation. However, current methods face challenges in limited real-world examples, handling diverse layout descriptions and varying levels of granularity. To address these issues, a novel framework named Spatial Knowledge Enhanced Layout (SKE-Layout), is introduced. SKE-Layout integrates mixed spatial knowledge sources, leveraging both real and synthetic data to enhance spatial contexts. It utilizes diverse representations tailored to specific tasks and employs contrastive learning and multitask learning techniques for accurate spatial knowledge retrieval. This framework generates more accurate and fine-grained visual layouts for object rearrangement and text-to-image generation tasks, achieving improvements of 5%-30% compared to existing methods. Nieqing Cao, Yan Ding 0002, Mengying Xie, Fuqiang Gu, Chao Chen 0004 |
CVPR | 2 |
| 2025 | MoMa-Pos: An Efficient Object-Kinematic-Aware Base Placement Determination Framework for Mobile Manipulation
Beichen Shao, Nieqing Cao, Yan Ding 0002, Fuqiang Gu, Chao Chen 0004 |
ICA3PP (3) | 2 |
| 2025 | OVA-Fields: Weakly Supervised Open-Vocabulary Affordance Fields for Robot Operational Part Detection
Heng Su, Mengying Xie, Nieqing Cao, Yan Ding 0002, Beichen Shao, Xianlei Long, Fuqiang Gu, Chao Chen 0004 |
ICCV | 3 |
| 2025 | AlignBot: Aligning VLM-Powered Customized Task Planning with User Reminders Through Fine-Tuning for Household RobotsabstractThis paper presents AlignBot, a novel framework designed to optimize VLM-powered customized task planning for household robots by effectively aligning with user reminders. In domestic settings, aligning task planning with user reminders poses significant challenges due to the limited quantity, diversity, and multimodal nature of the reminders. To address these challenges, AlignBot employs a fine-tuned LLaVA-7B model, functioning as an adapter for GPT-40. This adapter model internalizes diverse forms of user reminders-such as personalized preferences, corrective guidance, and contextual assistance-into structured instruction-formatted cues that prompt GPT-40 in generating customized task plans. Additionally, AlignBot integrates a dynamic retrieval mechanism that selects task-relevant historical successes as prompts for GPT-40, further enhancing task planning accuracy. To validate the effectiveness of AlignBot, experiments are conducted in real-world household environments, which are constructed within the laboratory to replicate typical household settings. A multimodal dataset with over 1,500 entries derived from volunteer reminders is used for training and evaluation. The results demonstrate that AlignBot significantly improves customized task planning, outperforming existing LLM- and VLM-powered planners by interpreting and aligning with user reminders, achieving 86.8 % success rate compared to the vanilla GPT-40 baseline at 21.6%, reflecting a 65% improvement and over four times greater effectiveness. Supplementary materials are available at: https://yding25.com/AlignBot/ Zhaxizhuoma, Pengan Chen, Ziniu Wu, Dong Wang 0028, Peng Zhou 0018, Nieqing Cao, Yan Ding 0002, Bin Zhao 0001, Xuelong Li 0001 |
ICRA | 7 |
| 2025 | DualCLIP: Bridging 3D Geometry and Multimodal Semantics for Robotic PerceptionabstractCurrent approaches to integrating CLIP into language-driven robotics face a fundamental dilemma: While robotic implementations overlook cutting-edge 3D classification adaptations of CLIP, existing 3D-oriented CLIP methods prove inadequate for interpreting color-critical instructions prevalent in manipulation tasks. We resolve this through DualCLIP, a contrastive multimodal fusion framework that hierarchically integrates depth-aligned CLIP encoders. Our approach first aligns depth and CLIP RGB encoders using synthetic RGB-D pairs, then performs multimodal fusion via contrastive learning with language-triplet optimization. This joint training preserves 3D geometric coherence and color semantics. Evaluations demonstrate DualCLIP’s combined strength — surpassing CLIP2Point in 3D classification while showing promising improvements for CLIPORT in color-sensitive robotic manipulation. This work establishes a paradigm for translating vision-language models into 3D-aware robotic systems without compromising task-specific modality sensitivity. Yinghao Liu, Penglin Dai, Yan Ding 0002, Nieqing Cao |
IROS | 4 |
| 2025 | BestMan: a modular mobile manipulator platform for embodied AI with unified simulation-hardware APIs
Kui Yang, Nieqing Cao, Beichen Shao, Yan Ding 0002, Chao Chen 0004 |
Frontiers Comput. Sci. | 2 |