EDBT 2026 Demo / reviewers in the wild / expert
Yiran Qin
dblp:354/2276
· DBLP profile ↗
15ranked-venue papers
5as first author
15since 2021 · last 2025
0009-0008-4561-0685ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 5 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 6 since 2021Systems, architecture and hardware · 3 · 1 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | T2ISafety: Benchmark for Assessing Fairness, Toxicity, and Privacy in Image GenerationabstractText-to-image (T2I) models have rapidly advanced, enabling the generation of high-quality images from text prompts across various domains. However, these models present notable safety concerns, including the risk of generating harmful, biased, or private content. Current research on assessing T2I safety remains in its early stages. While some efforts have been made to evaluate models on specific safety dimensions, many critical risks remain unexplored. To address this gap, we introduce T2ISafety, a safety benchmark that evaluates T2I models across three key domains: toxicity, fairness, and bias. We build a detailed hierarchy of 12 tasks and 44 categories based on these three domains, and meticulously collect 70K corresponding prompts. Based on this taxonomy and prompt set, we build a large-scale T2I dataset with 68K manually annotated images and train an evaluator capable of detecting critical risks that previous work has failed to identify, including risks that even ultra-large proprietary models like GPTs cannot correctly detect. We evaluate 15 prominent diffusion models on T2ISafety and reveal several concerns including persistent issues with racial fairness, a tendency to generate toxic content, and significant variation in privacy protection across the models, even with defense methods like concept erasing. Data and evaluator are released under https://github.com/adwardlee/t2i_safety. Zhelun Shi, Xuhao Hu, Bowen Dong 0001, Yiran Qin, Xihui Liu, Lu Sheng |
CVPR | 5 |
| 2025 | ReSo: A Reward-driven Self-organizing LLM-based Multi-Agent System for Reasoning TasksabstractMulti-agent systems have emerged as a promising approach for enhancing the reasoning capabilities of large language models in complex problem-solving.However, current MAS frameworks are limited by poor flexibility and scalability, with underdeveloped optimization strategies.To address these challenges, we propose ReSo, which integrates task graph generation with a reward-driven two-stage agent selection process.The core of ReSo is the proposed Collaborative Reward Model, which can provide fine-grained reward signals for MAS cooperation for optimization.We also introduce an automated data synthesis framework for generating MAS benchmarks, without human annotations.Experimentally, ReSo matches or outperforms existing methods.ReSo achieves 33.7% and 32.3% accuracy on Math-MAS and SciBench-MAS SciBench, while other methods completely fail.The code and data are available at Reso. Hejia Geng, Xiangyuan Xue, Yiran Qin, Zhiyong Wang 0001, Zhenfei Yin, Lei Bai 0001 |
EMNLP | 5 |
| 2025 | RoboFactory: Exploring Embodied Agent Collaboration with Compositional ConstraintsabstractDesigning effective embodied multi-agent systems is critical for solving complex real-world tasks across domains. Due to the complexity of multi-agent embodied systems, existing methods fail to automatically generate safe and efficient training data for such systems. To this end, we propose the concept of compositional constraints for embodied multi-agent systems, addressing the challenges arising from collaboration among embodied agents. We design various interfaces tailored to different types of constraints, enabling seamless interaction with the physical world. Leveraging compositional constraints and specifically designed interfaces, we develop an automated data collection framework for embodied multi-agent systems and introduce the first benchmark for embodied multi-agent manipulation, RoboFactory. Based on RoboFactory benchmark, we adapt and evaluate the method of imitation learning and analyzed its performance in different difficulty agent tasks. Furthermore, we explore the architectures and training strategies for multi-agent imitation learning, aiming to build safe and efficient embodied multi-agent systems. Yiran Qin, Xiufeng Song, Zhenfei Yin, Xiaohong Liu 0001, Xihui Liu, Ruimao Zhang, Lei Bai 0001 |
ICCV | 1 |
| 2025 | GameFactorly: Creating New Games with Generative Interactive Videos
Jiwen Yu, Yiran Qin, Xintao Wang 0002, Pengfei Wan 0001, Di Zhang 0026, Xihui Liu |
ICCV | 2 |
| 2025 | High-Dynamic Radar Sequence Prediction for Weather Nowcasting Using Spatiotemporal Coherent Gaussian RepresentationabstractWeather nowcasting is an essential task that involves predicting future radar echo sequences based on current observations, offering significant benefits for disaster management, transportation, and urban planning. Current prediction methods are limited by training and storage efficiency, mainly focusing on 2D spatial predictions at specific altitudes. Meanwhile, 3D volumetric predictions at each timestamp remain largely unexplored. To address such a challenge, we introduce a comprehensive framework for 3D radar sequence prediction in weather nowcasting, using the newly proposed SpatioTemporal Coherent Gaussian Splatting (STC-GS) for dynamic radar representation and GauMamba for efficient and accurate forecasting. Specifically, rather than relying on a 4D Gaussian for dynamic scene reconstruction, STC-GS optimizes 3D scenes at each frame by employing a group of Gaussians while effectively capturing their movements across consecutive frames. It ensures consistent tracking of each Gaussian over time, making it particularly effective for prediction tasks. With the temporally correlated Gaussian groups established, we utilize them to train GauMamba, which integrates a memory mechanism into the Mamba framework. This allows the model to learn the temporal evolution of Gaussian groups while efficiently handling a large volume of Gaussian tokens. As a result, it achieves both efficiency and accuracy in forecasting a wide range of dynamic meteorological radar signals. The experimental results demonstrate that our STC-GS can efficiently represent 3D radar sequences with over $16\times$ higher spatial resolution compared with the existing 3D representation methods, while GauMamba outperforms state-of-the-art methods in forecasting a broad spectrum of high-dynamic weather conditions. Yiran Qin, Ruimao Zhang |
ICLR | 2 |
| 2025 | WorldSimBench: Towards Video Generation Models as World SimulatorsabstractRecent advancements in predictive models have demonstrated exceptional capabilities in predicting the future state of objects and scenes. However, the lack of categorization based on inherent characteristics continues to hinder the progress of predictive model development. Additionally, existing benchmarks are unable to effectively evaluate higher-capability, highly embodied predictive models from an embodied perspective. In this work, we classify the functionalities of predictive models into a hierarchy and take the first step in evaluating World Simulators by proposing a dual evaluation framework called WorldSimBench. WorldSimBench includes Explicit Perceptual Evaluation and Implicit Manipulative Evaluation, encompassing human preference assessments from the visual perspective and action-level evaluations in embodied tasks, covering three representative embodied scenarios: Open-Ended Embodied Environment, Autonomous, Driving, and Robot Manipulation. In the Explicit Perceptual Evaluation, we introduce the HF-Embodied Dataset, a video assessment dataset based on fine-grained human feedback, which we use to train a Human Preference Evaluator that aligns with human perception and explicitly assesses the visual fidelity of World Simulater. In the Implicit Manipulative Evaluation, we assess the video-action consistency of World Simulators by evaluating whether the generated situation-aware video can be accurately translated into the correct control signals in dynamic environments. Our comprehensive evaluation offers key insights that can drive further innovation in video generation models, positioning World Simulators as a pivotal advancement toward embodied artificial intelligence. Yiran Qin, Zhelun Shi, Jiwen Yu, Enshen Zhou, Zhenfei Yin, Xihui Liu, Lu Sheng, Lei Bai 0001, Ruimao Zhang |
ICML | 1 |
| 2025 | NavigateDiff: Visual Predictors are Zero-Shot Navigation AssistantsabstractNavigating unfamiliar environments presents significant challenges for household robots, requiring the ability to recognize and reason about novel decoration and layout. Existing reinforcement learning methods cannot be directly transferred to new environments, as they typically rely on extensive mapping and exploration, leading to time-consuming and inefficient. To address these challenges, we try to transfer the logical knowledge and the generalization ability of pretrained foundation models to zero-shot navigation. By integrating a large vision-language model with a diffusion network, our approach named NavigateDiff constructs a visual predictor that continuously predicts the agent's potential observations in the next step which can assist robots generate robust actions. Furthermore, to adapt the temporal property of navigation, we introduce temporal historical information to ensure that the predicted image is aligned with the navigation scene. We then carefully designed an information fusion framework that embeds the predicted future frames as guidance into goalreaching policy to solve downstream image navigation tasks. This approach enhances navigation control and generalization across both simulated and real-world environments. Through extensive experimentation, we demonstrate the robustness and versatility of our method, showcasing its potential to improve the efficiency and effectiveness of robotic navigation in diverse settings. Project Page: https://21styouth.github.io/NavigateDiff/. Yiran Qin, Yuze Hong, Benyou Wang, Ruimao Zhang |
ICRA | 1 |
| 2025 | Chain-of-Imagination for Reliable Instruction Following in Decision MakingabstractEnabling the embodied agent to imagine step-by-step the future states and sequentially approach these situation-aware states can enhance its capability to make reliable action decisions from textual instructions. In this work, we introduce a simple but effective mechanism called Chain-of-Imagination (CoI), which repeatedly employs a Multimodal Large Language Model (MLLM) equipped with diffusion model to facilitate imagining and acting upon the series of intermediate situation-aware visual sub-goals one by one, resulting in more reliable instruction-following capability. Based on the CoI mechanism, we propose an embodied agent DecisionDreamer as the low-level controller that can be adapted to different open-world scenarios. Extensive experiments demonstrate that Decision-Dreamer can achieve more reliable and accurate decision-making and significantly outperform the state-of-the-art generalist agents in the Minecraft and CALVIN sandbox simulators, regarding the instruction-following capability. For more demos, please see https://sites.google.com/view/decisiondreamer. Enshen Zhou, Yiran Qin, Zhenfei Yin, Zhelun Shi, Yuzhou Huang, Ruimao Zhang, Lu Sheng |
IROS | 2 |
| 2025 | VIKI‑R: Coordinating Embodied Multi-Agent Cooperation via Reinforcement LearningabstractCoordinating multiple embodied agents in dynamic environments remains a core challenge in artificial intelligence, requiring both perception-driven reasoning and scalable cooperation strategies. While recent works have leveraged large language models (LLMs) for multi-agent planning, a few have begun to explore vision-language models (VLMs) for visual reasoning. However, these VLM-based approaches remain limited in their support for diverse embodiment types. In this work, we introduce VIKI-Bench, the first hierarchical benchmark tailored for embodied multi-agent cooperation, featuring three structured levels: agent activation, task planning, and trajectory perception. VIKI-Bench includes diverse robot embodiments, multi-view visual observations, and structured supervision signals to evaluate reasoning grounded in visual inputs. To demonstrate the utility of VIKI-Bench, we propose VIKI-R, a two-stage framework that fine-tunes a pretrained vision-language model (VLM) using Chain-of-Thought annotated demonstrations, followed by reinforcement learning under multi-level reward signals. Our extensive experiments show that VIKI-R significantly outperforms baselines method across all task levels. Furthermore, we show that reinforcement learning enables the emergence of compositional cooperation patterns among heterogeneous agents. Together, VIKI-Bench and VIKI-R offer a unified testbed and method for advancing multi-agent, visual-driven cooperation in embodied AI systems. Xiufeng Song, Yiran Qin, Jie Yang 0009, Xiaohong Liu 0001, Philip Torr 0001, Lei Bai 0001, Zhenfei Yin |
NeurIPS | 4 |
| 2025 | GauDP: Reinventing Multi-Agent Collaboration through Gaussian-Image Synergy in Diffusion PoliciesabstractDespite significant advances in robotic policy generation, effective coordination in embodied multi-agent systems remains a fundamental challenge—particularly in scenarios where agents must balance individual perspectives with global environmental awareness.
Existing approaches often struggle to balance fine-grained local control with comprehensive scene understanding, resulting in limited scalability and compromised collaboration quality.
In this paper, we present GauDP, a novel Gaussian-image synergistic representation that facilitates scalable, perception-aware imitation learning in multi-agent collaborative systems.
Specifically, GauDP reconstructs a globally consistent 3D Gaussian field from local-view RGB images, allowing all agents to dynamically query task-relevant features from a shared scene representation.
This design facilitates both fine-grained control and globally coherent behavior without requiring additional sensing modalities.
We evaluate GauDP on the RoboFactory benchmark, which includes diverse multi-arm manipulation tasks.
Our method achieves superior performance over existing image-based methods and approaches the effectiveness of point-cloud-driven methods, while maintaining strong scalability as the number of agents increases.
Extensive ablations and visualizations further demonstrate the robustness and efficiency of our unified local-global perception framework for multi-agent embodied learning. Yiran Qin, Jiahua Ma, Zhanglin Peng, Lei Bai 0001, Ruimao Zhang |
NeurIPS | 3 |
| 2025 | Context as Memory: Scene-Consistent Interactive Long Video Generation with Memory RetrievalabstractRecent advances in interactive video generation have shown promising results, yet existing approaches struggle with scene-consistent memory capabilities in long video generation due to limited use of historical context. In this work, we propose Context-as-Memory, which utilizes historical context as memory for video generation. It includes two simple yet effective designs: (1) storing context in frame format without additional post-processing; (2) conditioning by concatenating context and frames to be predicted along the frame dimension at the input, requiring no external control modules. Furthermore, considering the enormous computational overhead of incorporating all historical context, we propose the Memory Retrieval module to select truly relevant context frames by determining FOV (Field of View) overlap between camera poses, which significantly reduces the number of candidate frames without substantial information loss. Experiments demonstrate that Context-as-Memory achieves superior memory capabilities in interactive long video generation compared to SOTAs, even generalizing effectively to open-domain scenarios not seen during training. Our project page are publicly available at https://context-as-memory.github.io/. Jiwen Yu, Jianhong Bai, Yiran Qin, Quande Liu, Xintao Wang 0002, Pengfei Wan 0001, Di Zhang 0026, Xihui Liu |
SIGGRAPH Asia | 3 |
| 2025 | Boosting 3D Object Detection via Self-Distilling Introspective Dataabstract3D object detection is a fundamental yet critical task for autonomous driving. In this paper, we investigate a novel self-distilling paradigm by proposing Self-distilling Introspective Data (SID) to boost the accuracy of 3D object detection in both LiDAR-based and LiDAR-Camera-based scenarios. The proposed SID significantly improves the applicability of the distillation approach since it does not require extra training data or complex teacher network design. Specifically, we first employ an introspective data augmentation method to enrich object-aware information in sparse point clouds through geometric or semantic injection. We then utilize this enhanced data to train a robust teacher model. In contrast to traditional distillation that relies on larger models to enhance the representations of smaller ones, the teacher model in SID shares the same architecture as the student model but exhibits exceptionally high discriminative ability. This enables the effective transfer of rich feature representations to the student model. Rooted on such a scheme, when conducting LiDAR-based detectors, SID significantly enhances the semantic representation capabilities of sparse point clouds. Additionally, in the LiDAR-Camera-based setting, SID also effectively supervises the fusion of the two modalities at the feature level, ensuring more reasonable cross-modal learning. Extensive experiments show the proposed SID improves a variety of detectors. For the LiDAR-based detector, the SID gains 2.31% mAP improvements for the hard objects in KITTI, while 1.76% NDS improvements on nuScenes. For the LiDAR-Camera-based detectors, the SID boosts the detection accuracy significantly, with 1.5% mAP promotion on KITTI and 2.15% NDS improvements on the nuScenes benchmark. Chaoqun Wang 0012, Yiran Qin, Zijian Kang, Ningning Ma, Yukai Shi, Zhen Li 0026, Ruimao Zhang |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2024 | MP5: A Multi-modal Open-ended Embodied System in Minecraft via Active PerceptionabstractIt is a long-lasting goal to design an embodied system that can solve long-horizon open-world tasks in human-like ways. However, existing approaches usually struggle with compound difficulties caused by the logic-aware decompo-sition and context-aware execution of these tasks. To this end, we introduce MP5, an open-ended multimodal em-bodied system built upon the challenging Minecraft sim-ulator, which can decompose feasible sub-objectives, de-sign sophisticated situation-aware plans, and perform em-bodied action control, with frequent communication with a goal-conditioned active perception scheme. Specifically, MP5 is developed on top of recent advances in Multimodal Large Language Models (MLLMs), and the system is mod-ulated into functional modules that can be scheduled and collaborated to ultimately solve pre-defined context- and process-dependent tasks. Extensive experiments prove that MP5 can achieve a 22% success rate on difficult process-dependent tasks and a 91 % success rate on tasks that heav-ily depend on the context. Moreover, MP5 exhibits a re-markable ability to address many open-ended tasks that are entirely novel. Please see the project page at https: //iranqin. github.io/MP5. github.io/. Yiran Qin, Enshen Zhou, Qichang Liu, Zhenfei Yin, Lu Sheng, Ruimao Zhang, Yu Qiao 0001 |
CVPR | 1 |
| 2024 | Toward Accurate Camera-based 3D Object Detection via Cascade Depth Estimation and CalibrationabstractRecent camera-based 3D object detection is limited by the precision of transforming from image to 3D feature spaces, as well as the accuracy of object localization within the 3D space. This paper aims to address such a fundamental problem of camera-based 3D object detection: How to effectively learn depth information for accurate feature lifting and object localization. Different from previous methods which directly predict depth distributions by using a supervised estimation model, we propose a cascade framework consisting of two depth-aware learning paradigms. First, a depth estimation (DE) scheme leverages relative depth information to realize the effective feature lifting from 2D to 3D spaces. Furthermore, a depth calibration (DC) scheme introduces depth reconstruction to further adjust the 3D object localization perturbation along the depth axis. In practice, the DE is explicitly realized by using both the absolute and relative depth optimization loss to promote the precision of depth prediction, while the capability of DC is implicitly embedded into the detection Transformer through a depth denoising mechanism in the training phase. The entire model training is accomplished through an end-to-end manner. We propose a baseline detector and evaluate the effectiveness of our proposal with +2.2%/+2.7% NDS/mAP improvements on NuScenes benchmark, and gain a comparable performance with 55.9%/45.7% NDS/mAP. Furthermore, we conduct extensive experiments to demonstrate its generality based on various detectors with about +2% NDS improvements. Chaoqun Wang 0012, Yiran Qin, Zijian Kang, Ningning Ma, Ruimao Zhang |
ICRA | 2 |
| 2023 | SupFusion: Supervised LiDAR-Camera Fusion for 3D Object DetectionabstractLiDAR-Camera fusion-based 3D detection is a critical task for automatic driving. In recent years, many LiDAR-Camera fusion approaches sprung up and gained promising performances compared with single-modal detectors, but always lack carefully designed and effective supervision for the fusion process. In this paper, we propose a novel training strategy called SupFusion, which provides an auxiliary feature level supervision for effective LiDAR-Camera fusion and significantly boosts detection performance. Our strategy involves a data enhancement method named Polar Sampling, which densifies sparse objects and trains an assistant model to generate high-quality features as the supervision. These features are then used to train the LiDAR-Camera fusion model, where the fusion feature is optimized to simulate the generated high-quality features. Furthermore, we propose a simple yet effective deep fusion module, which contiguously gains superior performance compared with previous fusion methods with SupFusion strategy. In such a manner, our proposal shares the following advantages. Firstly, SupFusion introduces auxiliary feature-level supervision which could boost LiDAR-Camera detection performance without introducing extra inference costs. Secondly, the proposed deep fusion could continuously improve the detector’s abilities. Our proposed SupFusion and deep fusion module is plug-and-play, we make extensive experiments to demon-strate its effectiveness. Specifically, we gain around 2% 3D mAP improvements on KITTI benchmark based on multiple LiDAR-Camera 3D detectors. Our code is available at https://github.com/IranQin/SupFusion. Yiran Qin, Chaoqun Wang 0012, Zijian Kang, Ningning Ma, Zhen Li 0026, Ruimao Zhang |
ICCV | 1 |