EDBT 2026 Demo / reviewers in the wild / expert
Guyue Zhou
dblp:133/4199
· DBLP profile ↗
57ranked-venue papers
2as first author
53since 2021 · last 2025
0000-0002-3894-9858ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 45 · 2 first-author · 43 since 2021Systems, architecture and hardware · 27 · 2 first-author · 24 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 13 since 2021Human-computer interaction and ubiquitous computing · 5 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | H2O+: An Improved Framework for Hybrid Offline-and-Online RL with Dynamics GapsabstractSolving real-world complex tasks using reinforcement learning (RL) without high-fidelity simulation environments or large amounts of offline data can be quite challenging. Online RL agents trained in imperfect simulation environments can suffer from severe sim-to-real issues. Offline RL approaches although bypass the need for simulators, often pose demanding requirements on the size and quality of the offline datasets. The recently emerged hybrid offline-and-online RL provides an attractive framework that enables joint use of limited offline data and imperfect simulator for transferable policy learning. In this paper, we develop a new algorithm, called$\mathrm{H} 2 \mathrm{O}+$, which offers great flexibility to bridge various choices of offline and online learning methods, while also accounting for dynamics gaps between the real and simulation environments. Through extensive simulation and real-world robotics experiments, we demonstrate superior performance and flexibility of$\mathbf{H 2 O}+$over advanced cross-domain online and offline RL algorithms. Tianying Ji, Bingqi Liu, Haocheng Zhao, Jianying Zheng, Guyue Zhou, Jianming Hu, Xianyuan Zhan |
ICRA | 8 |
| 2025 | OpenBench: A New Benchmark and Baseline for Semantic Navigation in Smart LogisticsabstractThe increasing demand for efficient last-mile delivery in smart logistics underscores the role of autonomous robots in enhancing operational efficiency and reducing costs. Traditional navigation methods, which depend on highprecision maps, are resource-intensive, while learning-based approaches often struggle with generalization in real-world scenarios. To address these challenges, this work proposes the Openstreetmap-enhanced oPen-air sEmantic Navigation (OPEN) system that combines foundation models with classic algorithms for scalable outdoor navigation. The system uses off-the-shelf OpenStreetMap (OSM) for flexible map representation, thereby eliminating the need for extensive pre-mapping efforts. It also employs Large Language Models (LLMs) to comprehend delivery instructions and Vision-Language Models (VLMs) for global localization, map updates, and house number recognition. To compensate the limitations of existing benchmarks that are inadequate for assessing last-mile delivery, this work introduces a new benchmark specifically designed for outdoor navigation in residential areas, reflecting the real-world challenges faced by autonomous delivery systems. Extensive experiments in simulated and real-world environments demonstrate the proposed system's efficacy in enhancing navigation efficiency and reliability. To facilitate further research, our code and benchmark are publicly available11https://ei-nav.github.io/OpenBench/. Dongjie Huo, Zehui Xu, Yongliang Shi, Yimin Yan, Yan Qiao 0004, Guyue Zhou |
ICRA | 9 |
| 2025 | DISCOVERSE: Efficient Robot Simulation in Complex High-Fidelity EnvironmentsabstractWe present Discoverse, the first unified, modular, open-source 3DGS-based simulation framework for Real2Sim2Real robot learning. It features a holistic Real2Sim pipeline that synthesizes hyper-realistic geometry and appearance of complex real-world scenarios, paving the way for analyzing and bridging the Sim2Real gap. Powered by Gaussian Splatting and MuJoCo, Discoverse enables massively parallel simulation of multiple sensor modalities and accurate physics, with inclusive supports for existing 3D assets, robot models, and ROS plugins, empowering large-scale robot learning and complex robotic benchmarks. Through extensive experiments on imitation learning, Dis coverse demonstrates state-of-the-art zero-shot Sim2Real transfer performance compared to existing simulators. For code and demos: https://air-discoverse.github.io/. Yufei Jia, Junzhe Wu, Yupei Zeng, Haonan Lin, Haizhou Ge, Weibin Gu, Kairui Ding, Zike Yan, Yunjie Cheng, Chuxuan Li, Wei Sui, Guanzhong Tian, Ruqi Huang, Guyue Zhou |
IROS | 20 |
| 2025 | UniLegs: Universal Multi-Legged Robot Control through Morphology-Agnostic Policy DistillationabstractDeveloping controllers that generalize across diverse robot morphologies remains a significant challenge in legged locomotion. Traditional approaches either create specialized controllers for each morphology or compromise performance for generality. This paper introduces a two-stage teacher-student framework that bridges this gap through policy distillation. First, we train specialized teacher policies optimized for individual morphologies, capturing the unique optimal control strategies for each robot design. Then, we distill this specialized expertise into a single Transformer-based student policy capable of controlling robots with varying leg configurations. Our experiments across five distinct legged morphologies demonstrate that our approach preserves morphology-specific optimal behaviors, with the Transformer architecture achieving 94.47% of teacher performance on training morphologies and 72.64% on unseen robot designs. Comparative analysis reveals that Transformer-based architectures consistently outperform MLP baselines by leveraging attention mechanisms to effectively model joint relationships across different kinematic structures. We validate our approach through successful deployment on a physical quadruped robot, demonstrating the practical viability of our morphology-agnostic control framework. This work presents a scalable solution for developing universal legged robot controllers that maintain near-optimal performance while generalizing across diverse morphologies. Weijie Xi, Zhanxiang Cao, Chenlin Ming, Jianying Zheng, Guyue Zhou |
IROS | 5 |
| 2025 | Self-Aligning Depth-Regularized Radiance Fields for Asynchronous RGB-D SequencesabstractIt has been shown that learning radiance fields with depth rendering and depth supervision can effectively promote the quality and convergence of view synthesis. However, this paradigm requires input RGB-D sequences to be synchronized. In the UAV city modeling scenario, there exists asynchrony between RGB images and depth images due to the different frequencies of the solid-state LiDAR and RGB sensors. To synthesize high-quality views in such a scenario, we propose a novel time-pose function, which is an implicit network that maps timestamps to SE(3) elements. To train this function, we also design a joint optimization scheme to jointly learn the large-scale depth-regularized radiance fields and the time-pose function. Furthermore, we propose a large synthetic dataset with diverse controlled mismatches and ground truth to evaluate this new problem setting systematically. The proposed approach has been evaluated on both datasets and in a real drone. To evaluate the impact of view density, each algorithm was test on three different trajectories with different view densities. Compared to state-of-the-art baseline methods, the proposed approach reduces reconstruction error by 35.26% in city modeling scenarios. Our code is available at github.com/saythe17/AsyncNeRF. Andong Yang, Yuantao Chen, Runyi Yang, Zhenxin Zhu, Hao Zhao 0002, Guyue Zhou |
WACV | 8 |
| 2025 | OPEN: Lightweight Map-Based Semantic Navigation for GPS-Free Last-Mile DeliveryabstractThe growing demand for efficient last-mile delivery highlights the need for autonomous robots to improve operational efficiency and reduce costs. Traditional navigation methods rely on high-precision maps, which are expensive to produce and maintain, while learning-based approaches often struggle to adapt to diverse real-world environments. To address these challenges, this paper presents OpenStreetMap-enhanced oPen-air sEmantic Navigation (OPEN), a system that combines foundation models with classic navigation algorithms to enable scalable outdoor navigation. By leveraging off-the-shelf OpenStreetMap (OSM), OPEN eliminates the need for extensive pre-mapping and provides a lightweight, readily available map representation. The system uses Large Language Models (LLMs) to interpret delivery instructions and Vision Language Models (VLMs) for global localization without relying on GPS, real-time map updates, and entrance recognition, ensuring robust navigation in complex environments. To further enhance adaptability, OPEN incorporates a local replanning method that dynamically adjusts waypoints in response to environmental changes and OSM inaccuracies. Since existing benchmarks do not adequately reflect the challenges of last-mile delivery, this work introduces a new benchmark designed for residential navigation. Experiments conducted in both simulated and real-world settings demonstrate that OPEN improves navigation accuracy, efficiency, and reliability, outperforming existing semantic navigation methods. To facilitate further research, the code and benchmark are publicly available. Dongjie Huo, Yongliang Shi, Yan Qiao 0004, Guyue Zhou |
IEEE Trans Autom. Sci. Eng. | 6 |
| 2025 | Masked PaCONet: Self-Supervised Part-Aware Implicit Shape Reconstruction Scalability, Flexibility, Multi-scale and Semantic ConsistencyabstractLocalized neural implicit representation methods have recently been proven effective for shape reconstruction. However, while some recent neural implicit representation-based approaches have investigated part awareness, there is still room for improvement in leveraging the rich geometry information contained in parts, which is crucial for accurate reconstruction. This study aims to enhance the accuracy of shape reconstruction by incorporatingpart awareness. This principle faces a fundamental technical challenge: manually defining parts across various categories is ambiguous and expensive. To address it, we propose a new self-supervised learning paradigm that automatically discovers meaningful parts. Our proposed paradigm has several prominent advantages as compared with the prior arts: (1) It allows masked part modeling thatscaleswell with available data; (2) It is aflexibleformulation that allows a variable number of parts; (3) It allows the fusion ofmulti-scale(global-level and part-level) features at an arbitrarily given coordinate; (4) Thesemantic consistencyof learned parts leads to transferable features. Extensive experiments validate our approach, named Masked PaCONet, showcasing its superiority in qualitative and quantitative results on public benchmarks, even under challenging settings. Codes and models will be released. Tianyu Liu 0008, Hao Zhao 0002, Bohuan Xue, Guyue Zhou, Ming Liu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | Locate N' Rotate: Two-Stage Openable Part Detection with Foundation Model Priors
Siqi Li 0009, Xiaoxue Chen, Haoyu Cheng, Guyue Zhou, Hao Zhao 0002, Guanzhong Tian |
ACCV (7) | 4 |
| 2024 | Time-conditioned Illumination for Inverse Rendering of Outdoor Scenes
Xiaoxue Chen, Hao Zhao 0002, Guyue Zhou, Ya-Qin Zhang |
BMVC | 3 |
| 2024 | "It Must Be Gesturing Towards Me": Gesture-Based Interaction between Autonomous Vehicles and PedestriansabstractInteracting with pedestrians understandably and efficiently is one of the toughest challenges faced by autonomous vehicles (AVs) due to the limitations of current algorithms and external human-machine interfaces (eHMIs). In this paper, we design eHMIs based on gestures inspired by the most popular method of interaction between pedestrians and human drivers. Eight common gestures were selected to convey AVs’ yielding or non-yielding intentions at uncontrolled crosswalks from previous literature. Through a VR experiment (N1 = 31) and a following online survey (N2 = 394), we discovered significant differences in the usability of gesture-based eHMIs compared to current eHMIs. Good gesture-based eHMIs increase the efficiency of pedestrian-AV interaction while ensuring safety. Poor gestures, however, cause misinterpretation. The underlying reasons were explored: ambiguity regarding the recipient of the signal and whether the gestures are precise, polite, and familiar to pedestrians. Based on this empirical evidence, we discuss potential opportunities and provide valuable insights into developing comprehensible gesture-based eHMIs in the future to support better interaction between AVs and other road users. Xiang Chang, Zihe Chen, Xiaoyan Dong, Tingmin Yan, Haolin Cai, Zherui Zhou, Guyue Zhou, Jiangtao Gong |
CHI | 8 |
| 2024 | Understanding Human-AI Collaboration in Music Therapy Through Co-Design with TherapistsabstractThe rapid development of musical AI technologies has expanded the creative potential of various musical activities, ranging from music style transformation to music generation. However, little research has investigated how musical AIs can support music therapists, who urgently need new technology support. This study used a mixed method, including semi-structured interviews and a participatory design approach. By collaborating with music therapists, we explored design opportunities for musical AIs in music therapy. We presented the co-design outcomes involving the integration of musical AIs into a music therapy process, which was developed from a theoretical framework rooted in emotion-focused therapy. After that, we concluded the benefits and concerns surrounding music AIs from the perspective of music therapists. Based on our findings, we discussed the opportunities and design implications for applying musical AIs to music therapy. Our work offers valuable insights for developing human-AI collaborative music systems in therapy involving complex procedures and specific requirements. Guyue Zhou, Yucheng Jin 0001, Jiangtao Gong |
CHI | 3 |
| 2024 | Structured-NeRF: Hierarchical Scene Graph with Neural Representation
Zhide Zhong, Jiakai Cao, Songen Gu, Sirui Xie, Liyi Luo, Hao Zhao 0002, Guyue Zhou, Haoang Li, Zike Yan |
ECCV (35) | 7 |
| 2024 | Block-Map-Based Localization in Large-Scale EnvironmentabstractAccurate localization is an essential technology for the flexible navigation of robots in large-scale environments. Both SLAM-based and map-based localization will increase the computing load due to the increase in map size, which will affect downstream tasks such as robot navigation and services. To this end, we propose a localization system based on Block Maps (BMs) to reduce the computational load caused by maintaining large-scale maps. Firstly, we introduce a method for generating block maps and the corresponding switching strategies, ensuring that the robot can estimate the state in large-scale environments by loading local map information. Secondly, global localization according to Branch-and-Bound Search (BBS) in the 3D map is introduced to provide the initial pose. Finally, a graph-based optimization method is adopted with a dynamic sliding window that determines what factors are being marginalized whether a robot is exposed to a BM or switching to another one, which maintains the accuracy and efficiency of pose tracking. Comparison experiments are performed on publicly available large-scale datasets. Results show that the proposed method can track the robot pose even though the map scale reaches more than 6 kilometers, while efficient and accurate localization is still guaranteed on NCLT [6] and M2DGR [35]. Codes and data will be publicly available on https://github.com/YixFeng/blocklocalization. Yixiao Feng, Yongliang Shi, Yunlong Feng, Hao Zhao 0002, Guyue Zhou |
ICRA | 7 |
| 2024 | Camera Relocalization in Shadow-free Neural Radiance FieldsabstractCamera relocalization is a crucial problem in computer vision and robotics. Recent advancements in neural radiance fields (NeRFs) have shown promise in synthesizing photo-realistic images. Several works have utilized NeRFs for refining camera poses, but they do not account for lighting changes that can affect scene appearance and shadow regions, causing a degraded pose optimization process. In this paper, we propose a two-staged pipeline that normalizes images with varying lighting and shadow conditions to improve camera relocalization. We implement our scene representation upon a hash-encoded NeRF which significantly boosts up the pose optimization process. To account for the noisy image gradient computing problem in grid-based NeRFs, we further propose a re-devised truncated dynamic low-pass filter (TDLF) and a numerical gradient averaging technique to smoothen the process. Experimental results on several datasets with varying lighting conditions demonstrate that our method achieves state-of-the-art results in camera relocalization under varying lighting conditions. Code and data will be made publicly available. Shiyao Xu, Caiyun Liu 0004, Yuantao Chen, Zhenxin Zhu, Zike Yan, Yongliang Shi, Hao Zhao 0002, Guyue Zhou |
ICRA | 8 |
| 2024 | A Comprehensive Survey of Cross-Domain Policy Transfer for Embodied Agents
Jianming Hu, Guyue Zhou, Xianyuan Zhan |
IJCAI | 3 |
| 2024 | PreAfford: Universal Affordance-Based Pre-Grasping for Diverse Objects and EnvironmentsabstractRobotic manipulation with two-finger grippers is challenged by objects lacking distinct graspable features. Traditional pre-grasping methods, which typically involve repositioning objects or utilizing external aids like table edges, are limited in their adaptability across different object categories and environments. To overcome these limitations, we introduce PreAfford, a novel pre-grasping planning framework incorporating a point-level affordance representation and a relay training approach. Our method significantly improves adaptability, allowing effective manipulation across a wide range of environments and object types. When evaluated on the ShapeNet-v2 dataset, PreAfford not only enhances grasping success rates by 69% but also demonstrates its practicality through successful real-world experiments. These improvements highlight PreAfford’s potential to redefine standards for robotic handling of complex manipulation tasks in diverse settings. Kairui Ding, Boyuan Chen 0009, Ruihai Wu, Zongzheng Zhang, Huan-ang Gao, Siqi Li 0009, Guyue Zhou, Yixin Zhu 0001, Hao Dong 0003, Hao Zhao 0002 |
IROS | 8 |
| 2024 | SurrealDriver: Designing LLM-powered Generative Driver Agent Framework based on Human Drivers' Driving-thinking DataabstractLeveraging advanced reasoning capabilities and extensive world knowledge of large language models (LLMs) to construct generative agents for solving complex real-world problems is a major trend. However, LLMs inherently lack embodiment as humans, resulting in suboptimal performance in many embodied decision-making tasks. In this paper, we introduce a framework for building human-like generative driving agents using post-driving self-report driving-thinking data from human drivers as both demonstration and feedback. To capture high-quality, natural language data from drivers, we conducted urban driving experiments, recording drivers’ verbalized thoughts under various conditions to serve as chain-of-thought prompts and demonstration examples for the LLM-Agent. The framework’s effectiveness was evaluated through simulations and human assessments. Results indicate that incorporating expert demonstration data significantly reduced collision rates by 81.04% and increased human likeness by 50% compared to a baseline LLM-based agent. Our study provides insights into using natural language-based human demonstration data for embodied tasks. The driving-thinking dataset is available at https://github.com/AIR-DISCOVER/Driving-Thinking-Dataset. Zhijie Yi, Xiaoxi Shen, Huiling Peng, Xiaoan Liu, Jingli Qin, Jintao Xie, Peizhong Gao, Guyue Zhou, Jiangtao Gong |
IROS | 11 |
| 2024 | Active Neural Mapping at ScaleabstractWe introduce a NeRF-based active mapping system that enables efficient and robust exploration of large-scale indoor environments. The key to our approach is the extraction of a generalized Voronoi graph (GVG) from the continually updated neural map, leading to the synergistic integration of scene geometry, appearance, topology, and uncertainty. Anchoring uncertain areas induced by the neural map to the vertices of GVG allows the exploration to undergo adaptive granularity along a safe path that traverses unknown areas efficiently. Harnessing a modern hybrid NeRF representation, the proposed system achieves competitive results in terms of reconstruction accuracy, coverage completeness, and exploration efficiency even when scaling up to large indoor environments. Extensive results at different scales validate the efficacy of the proposed system. Zijia Kuang, Zike Yan, Hao Zhao 0001, Guyue Zhou, Hongbin Zha |
IROS | 4 |
| 2024 | Arm-Constrained Curriculum Learning for Loco-Manipulation of a Wheel-Legged RobotabstractIncorporating a robotic manipulator into a wheellegged robot enhances its agility and expands its potential for practical applications. However, the presence of potential instability and uncertainties presents additional challenges for control objectives. In this paper, we introduce an arm-constrained curriculum learning architecture to tackle the issues introduced by adding the manipulator. Firstly, we develop an arm-constrained reinforcement learning algorithm to ensure safety and reliability in control performance after equipping the manipulator. Additionally, to address discrepancies in reward settings between the arm and the base, we propose a reward-aware curriculum learning method. The policy is first trained in Isaac gym and transferred to the physical robot to complete grasping tasks, including the door-opening task, fan-twitching task and the relay-baton-picking and following task. The results demonstrate that our proposed approach effectively controls the arm-equipped wheel-legged robot to master grasping abilities including the dynamic grasping skills, allowing it to chase and catch a moving object while in motion. Please refer to our website (https://acodedog.github.io/wheel-legged-loco-manipulation/) for the code and supplemental videos. Yufei Jia, Haizhou Zhao, Jinni Zhou, Jun Ma 0008, Guyue Zhou |
IROS | 9 |
| 2024 | Blending Distributed NeRFs with Tri-stage Robust Pose OptimizationabstractDue to the limited model capacity, leveraging distributed Neural Radiance Fields (NeRFs) for modeling extensive urban environments has become a necessity. However, current distributed NeRF registration approaches encounter aliasing artifacts, arising from discrepancies in rendering resolutions and suboptimal pose precision. These factors collectively deteriorate the fidelity of pose estimation within NeRF frameworks, resulting in occlusion artifacts during the NeRF blending stage. In this paper, we present a distributed NeRF system with tri-stage pose optimization. In the first stage, precise poses of images are achieved by bundle adjusting Mip-NeRF 360 with a coarse-to-fine strategy. In the second stage, we incorporate the inverting Mip-NeRF 360, coupled with the truncated dynamic low-pass filter, to enable the achievement of robust and precise poses, termed Frame2Model optimization. On top of this, we obtain a coarse transformation between NeRFs in different coordinate systems. In the third stage, we fine-tune the transformation between NeRFs by Model2Model pose optimization. After obtaining precise transformation parameters, we proceed to implement NeRF blending, showcasing superior performance metrics in both real-world and simulation scenarios. Codes and data will be publicly available at https://github.com/boilcy/Distributed-NeRF. Baijun Ye, Caiyun Liu 0004, Xiaoyu Ye, Yuantao Chen, Yuhai Wang, Zike Yan, Yongliang Shi, Hao Zhao 0002, Guyue Zhou |
IROS | 9 |
| 2024 | Closed-Loop Visuomotor Control with Generative Expectation for Robotic ManipulationabstractDespite significant progress in robotics and embodied AI in recent years, deploying robots for long-horizon tasks remains a great challenge. Majority of prior arts adhere to an open-loop philosophy and lack real-time feedback, leading to error accumulation and undesirable robustness. A handful of approaches have endeavored to establish feedback mechanisms leveraging pixel-level differences or pre-trained visual representations, yet their efficacy and adaptability have been found to be constrained. Inspired by classic closed-loop control systems, we propose CLOVER, a closed-loop visuomotor control framework that incorporates feedback mechanisms to improve adaptive robotic control. CLOVER consists of a text-conditioned video diffusion model for generating visual plans as reference inputs, a measurable embedding space for accurate error quantification, and a feedback-driven controller that refines actions from feedback and initiates replans as needed. Our framework exhibits notable advancement in real-world robotic tasks and achieves state-of-the-art on CALVIN benchmark, improving by 8% over previous open-loop counterparts. Code and checkpoints are maintained at https://github.com/OpenDriveLab/CLOVER. Qingwen Bu, Li Chen 0008, Yanchao Yang 0001, Guyue Zhou, Junchi Yan, Ping Luo 0002, Heming Cui, Yi Ma 0001, Hongyang Li 0001 |
NeurIPS | 5 |
| 2024 | More Than Routing: Joint GPS and Route Modeling for Refine Trajectory Representation LearningabstractTrajectory representation learning plays a pivotal role in supporting various downstream tasks, such as travel time estimation, trajectory classification and Top-k similar trajectory search. Traditional methods in order to filter the noise in GPS trajectories tend to focus on routing-based methods to simplify the trajectories. However, these approaches ignore the motion details contained in the GPS data, limiting the representation capability of trajectory representation learning. To fill this gap, we propose a novel representation learning framework that is Jointly G PS and Route Modeling based on self-supervised technology, namely JGRM. We consider GPS trajectory and route trajectory as the two modals of a single movement observation and fuse information through inter-modal information interaction. Specifically, we develop two encoders, each tailored to capture representations of GPS trajectories and route trajectories respectively. The representations from these two modalities are fed into a shared transformer for inter-modal information interaction. Eventually, we design three self-supervised tasks to train the model. We validate the effectiveness of the proposed method on two real-world datasets through extensive experiments. The experimental results show that JGRM significantly outperforms existing methods in both road segment representation and trajectory representation tasks. Our source code is available at Github https://github.com/mamazi0131/JGRM. Zheyan Tu, Xinhai Chen 0002, Yan Zhang 0122, Deguo Xia, Guyue Zhou, Yu Zheng 0004, Jiangtao Gong |
WWW | 6 |
| 2024 | ECT: Fine-grained edge detection with learned cause tokens
Shaocong Xu, Xiaoxue Chen, Yuhang Zheng 0004, Guyue Zhou, Yurong Chen 0001, Hongbin Zha, Hao Zhao 0002 |
Image Vis. Comput. | 4 |
| 2024 | City-scale continual neural semantic mapping with three-layer sampling and panoptic representation
Yongliang Shi, Runyi Yang, Zirui Wu, Pengfei Li 0007, Caiyun Liu 0004, Hao Zhao 0002, Guyue Zhou |
Knowl. Based Syst. | 7 |
| 2024 | Adaptive Surface Normal Constraint for Geometric Estimation From Monocular ImagesabstractWe introduce a novel approach to learn geometries such as depth and surface normal from images while incorporating geometric context. The difficulty of reliably capturing geometric context in existing methods impedes their ability to accurately enforce the consistency between the different geometric properties, thereby leading to a bottleneck of geometric estimation quality. We therefore propose the Adaptive Surface Normal (ASN) constraint, a simple yet efficient method. Our approach extracts geometric context that encodes the geometric variations present in the input image and correlates depth estimation with geometric constraints. By dynamically determining reliable local geometry from randomly sampled candidates, we establish a surface normal constraint, where the validity of these candidates is evaluated using the geometric context. Furthermore, our normal estimation leverages the geometric context to prioritize regions that exhibit significant geometric variations, which makes the predicted normals accurately capture intricate and detailed geometric information. Through the integration of geometric context, our method unifies depth and surface normal estimations within a cohesive framework, which enables the generation of high-quality 3D geometry from images. We validate the superiority of our approach over state-of-the-art methods through extensive evaluations and comparisons on diverse indoor and outdoor datasets, showcasing its efficiency and robustness. Xiaoxiao Long, Yuhang Zheng 0004, Yupeng Zheng, Beiwen Tian, Cheng Lin 0001, Lingjie Liu, Hao Zhao 0002, Guyue Zhou, Wenping Wang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 8 |
| 2023 | Work with AI and Work for AI: Autonomous Vehicle Safety Drivers' Lived ExperiencesabstractThe development of Autonomous Vehicle (AV) has created a novel job, the safety driver, recruited from experienced drivers to supervise and operate AV in numerous driving missions. Safety drivers usually work with non-perfect AV in high-risk real-world traffic environments for road testing tasks. However, this group of workers is under-explored in the HCI community. To fill this gap, we conducted semi-structured interviews with 26 safety drivers. Our results present how safety drivers cope with defective algorithms and shape and calibrate their perceptions while working with AV. We found that, as front-line workers, safety drivers are forced to take risks accumulated from the AV industry upstream and are also confronting restricted self-development in working for AV development. We contribute the first empirical evidence of the lived experience of safety drivers, the first passengers in the development of AV, and also the grassroots workers for AV, which can shed light on future human-AI interaction research. Mengdi Chu, Keyu Zong, Xin Shu 0009, Jiangtao Gong, Zhicong Lu, Kaimin Guo, Xinyi Dai, Guyue Zhou |
CHI | 8 |
| 2023 | MR.Brick: Designing A Remote Mixed-reality Educational Game System for Promoting Children's Social & Collaborative SkillsabstractChildren are one of the groups most influenced by COVID-19-related social distancing, and a lack of contact with peers can limit their opportunities to develop social and collaborative skills. However, remote socialization and collaboration as an alternative approach is still a great challenge for children. This paper presents MR.Brick, a Mixed Reality (MR) educational game system that helps children adapt to remote collaboration. A controlled experimental study involving 24 children aged six to ten was conducted to compare MR.Brick with the traditional video game by measuring their social and collaborative skills and analyzing their multi-modal playing behaviours. The results showed that MR.Brick was more conducive to children’s remote collaboration experience than the traditional video game. Given the lack of training systems designed for children to collaborate remotely, this study may inspire interaction design and educational research in related fields. Yudan Wu, Shanhe You, Zixuan Guo 0003, Guyue Zhou, Jiangtao Gong |
CHI | 5 |
| 2023 | "I am the follower, also the boss": Exploring Different Levels of Autonomy and Machine Forms of Guiding Robots for the Visually ImpairedabstractGuiding robots, in the form of canes or cars, have recently been explored to assist blind and low vision (BLV) people. Such robots can provide full or partial autonomy when guiding. However, the pros and cons of different forms and autonomy for guiding robots remain unknown. We sought to fill this gap. We designed autonomy-switchable guiding robotic cane and car. We conducted a controlled lab-study (N=12) and a field study (N=9) on BLV. Results showed that full autonomy received better walking performance and subjective ratings in the controlled study, whereas participants used more partial autonomy in the natural environment as demanding more control. Besides, the car robot has demonstrated abilities to provide a higher sense of safety and navigation efficiency compared with the cane robot. Our findings offered empirical evidence about how the BLV community perceived different machine forms and autonomy, which can inform the design of assistive robots. Yan Zhang 0122, Haole Guo, Qihe Chen, Mingming Fan 0001, Guyue Zhou, Jiangtao Gong |
CHI | 8 |
| 2023 | DPF: Learning Dense Prediction Fields with Weak SupervisionabstractNowadays, many visual scene understanding problems are addressed by dense prediction networks. But pixel-wise dense annotations are very expensive (e.g., for scene parsing) or impossible (e.g., for intrinsic image decomposition), motivating us to leverage cheap point-level weak supervision. However, existing pointly-supervised methods still use the same architecture designed for full supervision. In stark contrast to them, we propose a new paradigm that makes predictions for point coordinate queries, as inspired by the recent success of implicit representations, like distance or radiance fields. As such, the method is named as dense prediction fields (DPFs). DPFs generate expressive intermediate features for continuous sub-pixel locations, thus allowing outputs of an arbitrary resolution. DPFs are naturally compatible with point-level supervision. We showcase the effectiveness of DPFs using two substantially different tasks: high-level semantic parsing and low-level intrinsic image decomposition. In these two cases, supervision comes in the form of single-point semantic category and two-point relative reflectance, respectively. As benchmarked by three large-scale public datasets PASCALContext, ADE20K and IIW, DPFs set new state-of-the-art performance on all of them with significant margins. Code can be accessed at https://github.com/cxx226/DPF. Xiaoxue Chen, Yuhang Zheng 0004, Yupeng Zheng, Hao Zhao 0002, Guyue Zhou, Ya-Qin Zhang |
CVPR | 6 |
| 2023 | Delving into Shape-aware Zero-shot Semantic SegmentationabstractThanks to the impressive progress of large-scale vision-language pretraining, recent recognition models can classify arbitrary objects in a zero-shot and open-set manner, with a surprisingly high accuracy. However, translating this success to semantic segmentation is not trivial, because this dense prediction task requires not only accurate semantic understanding but also fine shape delineation and existing vision-language models are trained with image-level language descriptions. To bridge this gap, we pursue shape-aware zero-shot semantic segmentation in this study. Inspired by classical spectral methods in the image segmentation literature, we propose to leverage the eigen vectors of Laplacian matrices constructed with self-supervised pixel-wise features to promote shape-awareness. Despite that this simple and effective technique does not make use of the masks of seen classes at all, we demonstrate that it out-performs a state-of-the-art shape-aware formulation that aligns ground truth and predicted edges during training. We also delve into the performance gains achieved on different datasets using different backbones and draw several interesting and conclusive observations: the benefits of promoting shape-awareness highly relates to mask compactness and language embedding locality. Finally, our method sets new state-of-the-art performance for zero-shot semantic segmentation on both Pascal and COCO, with significant margins. Code and models will be accessed at SAZS. Beiwen Tian, Kehua Sheng, Bo Zhang 0106, Hao Zhao 0002, Guyue Zhou |
CVPR | 8 |
| 2023 | DQS3D: Densely-matched Quantization-aware Semi-supervised 3D DetectionabstractIn this paper, we study the problem of semi-supervised 3D object detection, which is of great importance considering the high annotation cost for cluttered 3D indoor scenes. We resort to the robust and principled framework of self-teaching, which has triggered notable progress for semi-supervised learning recently. While this paradigm is natural for image-level or pixel-level prediction, adapting it to the detection problem is challenged by the issue of proposal matching. Prior methods are based upon two-stage pipelines, matching heuristically selected proposals generated in the first stage and resulting in spatially sparse training signals. In contrast, we propose the first semi-supervised 3D detection algorithm that works in the single-stage manner and allows spatially dense training signals. A fundamental issue of this new design is the quantization error caused by point-to-voxel discretization, which inevitably leads to misalignment between two transformed views in the voxel domain. To this end, we derive and implement closed-form rules that compensate this misalignment on-the-fly. Our results are significant, e.g., promoting Scan-Net [email protected] from 35.2% to 48.5% using 20% annotation. Codes and data are publicly available1. Huan-ang Gao, Beiwen Tian, Pengfei Li 0007, Hao Zhao 0002, Guyue Zhou |
ICCV | 5 |
| 2023 | INT2: Interactive Trajectory Prediction at IntersectionsabstractMotion forecasting is an important component in autonomous driving systems. One of the most challenging problems in motion forecasting is interactive trajectory prediction, whose goal is to jointly forecasts the future trajectories of interacting agents. To this end, we present a large-scale interactive trajectory prediction dataset named INT2 for INTeractive trajectory prediction at INTersections. INT2 includes 612,000 scenes, each lasting 1 minute, containing up to 10,200 hours of data. The agent trajectories are auto-labeled by a high-performance offline temporal detection and fusion algorithm, whose quality is further inspected by human judges. Vectorized semantic maps and traffic light information are also included in INT2. Additionally, the dataset poses an interesting domain mismatch challenge. For each intersection, we treat rush-hour and non-rush-hour segments as different domains. We benchmark the best open-sourced interactive trajectory prediction method on INT2 and Waymo Open Motion, under in-domain and cross-domain settings. The dataset, code and models are publicly available at https://github.com/AIRDISCOVER/INT2. Zhijie Yan, Pengfei Li 0007, Zheng Fu, Shaocong Xu, Yongliang Shi, Xiaoxue Chen, Yuhang Zheng 0004, Yang Li 0178, Tianyu Liu 0008, Chuxuan Li, Nairui Luo, Zuoxu Wang, Yifeng Shi, Zhengxiao Han, Jirui Yuan, Jiangtao Gong, Guyue Zhou, Hang Zhao 0021, Hao Zhao 0002 |
ICCV | 20 |
| 2023 | 3D Implicit Transporter for Temporally Consistent Keypoint DiscoveryabstractKeypoint-based representation has proven advantageous in various visual and robotic tasks. However, the existing 2D and 3D methods for detecting keypoints mainly rely on geometric consistency to achieve spatial alignment, neglecting temporal consistency. To address this issue, the Transporter method was introduced for 2D data, which reconstructs the target frame from the source frame to incorporate both spatial and temporal information. However, the direct application of the Transporter to 3D point clouds is infeasible due to their structural differences from 2D images. Thus, we propose the first 3D version of the Transporter, which leverages hybrid 3D representation, cross attention, and implicit reconstruction. We apply this new learning system on 3D articulated objects and non-rigid animals (humans and rodents) and show that learned keypoints are spatio-temporally consistent. Additionally, we propose a closed-loop control strategy that utilizes the learned keypoints for 3D object manipulation and demonstrate its superior performance. Codes are available at https://github.com/zhongcl-thu/3D-Implicit-Transporter. Chengliang Zhong, Yuhang Zheng 0004, Yupeng Zheng, Hao Zhao 0002, Li Yi 0001, Xiaodong Mu, Ling Wang 0001, Pengfei Li 0007, Guyue Zhou, Chao Yang 0026, Jian Zhao 0006 |
ICCV | 9 |
| 2023 | Understanding Embodied Reference with Touch-Line Transformer
Yang Li 0178, Xiaoxue Chen, Hao Zhao 0002, Jiangtao Gong, Guyue Zhou, Federico Rossano, Yixin Zhu 0001 |
ICLR | 5 |
| 2023 | From Semi-supervised to Omni-supervised Room Layout Estimation Using Point CloudsabstractRoom layout estimation is a long-existing robotic vision task that benefits both environment sensing and motion planning. However, layout estimation using point clouds (PCs) still suffers from data scarcity due to annotation difficulty. As such, we address the semi-supervised setting of this task based upon the idea of model exponential moving averaging. But adapting this scheme to the state-of-the-art (SOTA) solution for PC-based layout estimation is not straightforward. To this end, we define a quad set matching strategy and several consistency losses based upon metrics tailored for layout quads. Besides, we propose a new online pseudo-label harvesting algorithm that decomposes the distribution of a hybrid distance measure between quads and PC into two components. This technique does not need manual threshold selection and intuitively encourages quads to align with reliable layout points. Surprisingly, this framework also works for the fully-supervised setting, achieving a new SOTA on the ScanNet benchmark. Last but not least, we also push the semi-supervised setting to the realistic omni-supervised setting, demonstrating significantly promoted performance on a newly annotated ARKitScenes testing set. Our codes, data and models are made publicly available**Code: https://github.com/AIR-DISCOVER/Omni-PQ. Huan-ang Gao, Beiwen Tian, Pengfei Li 0007, Xiaoxue Chen, Hao Zhao 0002, Guyue Zhou, Yurong Chen 0001, Hongbin Zha |
ICRA | 6 |
| 2023 | ADAPT: Action-aware Driving Caption TransformerabstractEnd-to-end autonomous driving has great potential in the transportation industry. However, the lack of transparency and interpretability of the automatic decision-making process hinders its industrial adoption in practice. There have been some early attempts to use attention maps or cost volume for better model explainability which is difficult for ordinary passengers to understand. To bridge the gap, we propose an end-to-end transformer-based architecture, ADAPT (Action-aware Driving cAPtion Transformer), which provides user-friendly natural language narrations and reasoning for each decision making step of autonomous vehicular control and action. ADAPT jointly trains both the driving caption task and the vehicular control prediction task, through a shared video representation. Experiments on BDD-X (Berkeley DeepDrive eXplanation) dataset demonstrate state-of-the-art performance of the ADAPT framework on both automatic metrics and human evaluation. To illustrate the feasibility of the proposed framework in real-world applications, we build a novel deployable system that takes raw car videos as input and outputs the action narrations and reasoning in real time. The code, models and data are available at https://github.com/jxbbb/ADAPT. Bu Jin, Yupeng Zheng, Pengfei Li 0007, Hao Zhao 0002, Yuhang Zheng 0004, Guyue Zhou |
ICRA | 8 |
| 2023 | LODE: Locally Conditioned Eikonal Implicit Scene Completion from Sparse LiDARabstractScene completion refers to obtaining dense scene representation from an incomplete perception of complex 3D scenes. This helps robots detect multi-scale obstacles and analyse object occlusions in scenarios such as autonomous driving. Recent advances show that implicit representation learning can be leveraged for continuous scene completion and achieved through physical constraints like Eikonal equations. However, former Eikonal completion methods only demonstrate results on watertight meshes at a scale of tens of meshes. None of them are successfully done for non-watertight LiDAR point clouds of open large scenes at a scale of thousands of scenes. In this paper, we propose a novel Eikonal formulation that conditions the implicit representation on localized shape priors which function as dense boundary value constraints, and demonstrate it works on SemanticKITTI and SemanticPOSS. It can also be extended to semantic Eikonal scene completion with only small modifications to the network architecture. With extensive quantitative and qualitative results, we demonstrate the benefits and drawbacks of existing Eikonal methods, which naturally leads to the new locally conditioned formulation. Notably, we improve IoU from 31.7% to 51.2% on SemanticKITTI and from 40.5% to 48.7% on SemanticPOSS. We extensively ablate our methods and demonstrate that the proposed formulation is robust to a wide spectrum of implementation hyper-parameters. Codes and models are publicly available at https://github.com/AIR-DISCOVER/LODE Pengfei Li 0007, Ruowen Zhao, Yongliang Shi, Hao Zhao 0002, Jirui Yuan, Guyue Zhou, Ya-Qin Zhang |
ICRA | 6 |
| 2023 | Planning Assembly Sequence with Graph TransformerabstractAssembly Sequence Planning (ASP) is the essential process for modern manufacturing, proven to be NP-complete thus its effective and efficient solution has been a challenge for researchers in the field. In this paper, we present a graph-transformer based framework for the ASP problem which is trained and demonstrated on a self-collected ASP database. The ASP database contains a self-collected set of LEGO models. The LEGO model is abstracted to a heterogeneous graph structure after a thorough analysis of the original structure and feature extraction. The ground truth assembly sequence is first generated by brute-force search and then adjusted manually to be in line with human rational habits. Based on this self-collected ASP dataset, we propose a heterogeneous graph-transformer framework to learn the latent rules for assembly planning. We evaluated the proposed framework in a series of experiments. The results show that the similarity of the predicted and ground truth sequences can reach 0.44, a medium correlation measured by Kendall's τ. Meanwhile, we compared the different effects of node features and edge features and generated a feasible and reasonable assembly sequence as a benchmark for further research. Our dataset and code are available on: htps://github.com/AIR-DISCOVER/ICRA_ASP. Lin Ma 0002, Jiangtao Gong, Hao Chen 0062, Hao Zhao 0002, Wenbing Huang 0001, Guyue Zhou |
ICRA | 7 |
| 2023 | Unsupervised Road Anomaly Detection with Language AnchorsabstractRoad anomaly detection is critical to safe autonomous driving, because current road scene understanding models are usually trained in a closed-set manner and fail to identify unknown objects. What's worse, it is difficult, if not impossible, to collect a large-scale dataset with anomaly annotations. So this paper studies unsupervised anomaly detection which finds out anomaly regions using scene parsing logits solely. While former methods depend on the weights learned from the closed training set as anchors for logit generation, we resort to language anchors that are learned from enormous paired vision and language data. Thanks to rich open-set semantic information contained in these language anchors, our method performs better than former unsupervised counterparts while maintaining the advantage of training without accessing any out-of-distribution data. We delve into this new paradigm and identify the superiority of using pair-wise binary logits, which we credit to a better understanding of the negation language anchor. Last but not least, we find that the former top-1 selection of semantic labels for uncertainty measurement is problematic in many cases and a new blended standardization strategy brings clear improvements to our solution. We report state-of-the-art performance on FS LostAndFound, LostAndFound and RoadAnomaly datasets among comparable methods. The codes are publicly available at https://github.com/TB5z035/URAD-LA.git Beiwen Tian, Mingdao Liu, Huan-ang Gao, Pengfei Li 0007, Hao Zhao 0002, Guyue Zhou |
ICRA | 6 |
| 2023 | Enable Natural Tactile Interaction for Robot Dog based on Large-format Distributed Flexible Pressure SensorsabstractTouch is an important channel for human-robot interaction, while it is challenging for robots to recognize human touch accurately and make appropriate responses. In this paper, we design and implement a set of large-format distributed flexible pressure sensors on a robot dog to enable natural human-robot tactile interaction. Through a heuristic study, we sorted out 81 tactile gestures commonly used when humans interact with real dogs and 44 dog reactions. A gesture classification algorithm based on ResNet is proposed to recognize these 81 human gestures, and the classification accuracy reaches 98.7%. In addition, an action prediction algorithm based on Transformer is proposed to predict dog actions from human gestures, reaching a 1-gram BLEU score of 0.87. Finally, we compare the tactile interaction with the voice interaction during a freedom human-robot-dog interactive playing study. The results show that tactile interaction plays a more significant role in alleviating user anxiety, stimulating user excitement and improving the acceptability of robot dogs. Lishuang Zhan, Yancheng Cao, Qitai Chen, Haole Guo, Jiasi Gao, Yiyue Luo, Shihui Guo, Guyue Zhou, Jiangtao Gong |
ICRA | 8 |
| 2023 | Annotating Covert Hazardous Driving Scenarios Online: Utilizing Drivers' Electroencephalography (EEG) SignalsabstractAs autonomous driving systems prevail, it is becoming increasingly critical that the systems learn from databases containing fine-grained driving scenarios. Most databases currently available are human-annotated; they are expensive, time-consuming, and subject to behavioral biases. In this paper, we provide initial evidence supporting a novel technique utilizing drivers' electroencephalography (EEG) signals to implicitly label hazardous driving scenarios while passively viewing recordings of real-road driving, thus sparing the need for manual annotation and avoiding human annotators' behavioral biases during explicit report. We conducted an EEG experiment using real-life and animated recordings of driving scenarios and asked participants to report danger explicitly whenever necessary. Behavioral results showed the participants tended to report danger only when overt hazards (e.g., a vehicle or a pedestrian appearing unexpectedly from behind an occlusion) were in view. By contrast, their EEG signals were enhanced at the sight of both an overt hazard and a covert hazard (e.g., an occlusion signalling possible appearance of a vehicle or a pedestrian from behind). Thus, EEG signals were more sensitive to driving hazards than explicit reports. Further, the Time-Series AI (TSAI, [1]) successfully classified EEG signals corresponding to overt and covert hazards. We discuss future steps necessary to materialize the technique in real life. Chen Zheng 0005, Muxiao Zi, Mengdi Chu, Yan Zhang 0122, Jirui Yuan, Guyue Zhou, Jiangtao Gong |
ICRA | 7 |
| 2023 | STEPS: Joint Self-supervised Nighttime Image Enhancement and Depth EstimationabstractSelf-supervised depth estimation draws a lot of attention recently as it can promote the 3D sensing capa-bilities of self-driving vehicles. However, it intrinsically relies upon the photometric consistency assumption, which hardly holds during nighttime. Although various supervised night-time image enhancement methods have been proposed, their generalization performance in challenging driving scenarios is not satisfactory. To this end, we propose the first method that jointly learns a nighttime image enhancer and a depth estimator, without using ground truth for either task. Our method tightly entangles two self-supervised tasks using a newly proposed uncertain pixel masking strategy. This strategy originates from the observation that nighttime images not only suffer from underexposed regions but also from overexposed regions. By fitting a bridge-shaped curve to the illumination map distribution, both regions are suppressed and two tasks are bridged naturally. We benchmark the method on two established datasets: nuScenes and RobotCar and demonstrate state-of-the-art performance on both of them. Detailed ablations also reveal the mechanism of our proposal. Last but not least, to mitigate the problem of sparse ground truth of existing datasets, we provide a new photo-realistically enhanced nighttime dataset based upon CARLA. It brings meaningful new challenges to the community. Codes, data, and models are available at https://github.com/ucaszyp/STEPS. Yupeng Zheng, Chengliang Zhong, Pengfei Li 0007, Huan-ang Gao, Yuhang Zheng 0004, Bu Jin, Ling Wang 0001, Hao Zhao 0002, Guyue Zhou, Dongbin Zhao |
ICRA | 9 |
| 2023 | LATITUDE: Robotic Global Localization with Truncated Dynamic Low-pass Filter in City-scale NeRFabstractNeural Radiance Fields (NeRFs) have made great success in representing complex 3D scenes with high-resolution details and efficient memory. Nevertheless, current NeRF - based pose estimators have no initial pose prediction and are prone to local optima during optimization. In this paper, we present LATITUDE: Global Localization with Truncated Dynamic Low-pass Filter, which introduces a two-stage localization mechanism in city-scale NeRF. In place recognition stage, we train a regressor through images generated from trained NeRFs, which provides an initial value for global localization. In pose optimization stage, we minimize the residual between the observed image and rendered image by directly optimizing the pose on the tangent plane. To avoid falling into local optimum, we introduce a Truncated Dynamic Low-pass Filter (TDLF) for coarse-to-fine pose registration. We evaluate our method on both synthetic and real-world data and show its potential applications for high-precision navigation in large-scale city scenes. Codes and dataset will be publicly available at https://github.com/jike5/LATITUDE. Zhenxin Zhu, Yuantao Chen, Zirui Wu, Yongliang Shi, Chuxuan Li, Pengfei Li 0007, Hao Zhao 0002, Guyue Zhou |
ICRA | 9 |
| 2023 | Real is Better than Perfect: Sim-to-Real Robotic System in Secondary School EducationabstractSimulation systems of robots can facilitate the prediction, development, and debugging of robotic systems. However, they seldom applied in robotics education for primary and secondary school students. In this paper, we present a sim-to-real robotic system that enables students to optimize their algorithms in a simulated environment and validate them in a remote physical laboratory with data logs and remote cameras. Moreover, the system employs an automated submit-test-reset subsystem that minimizes the need for human intervention and provides 24/7 testing support. Experimental data from a trial with 28 students in remote areas show that the sim-to-real robotic experimental environment has comparable learning outcomes to a pure real robot environment and is significantly better than a pure simulation environment. Given the results, we validate that our system can substantially reduce the costs of teaching equipment and space while maintaining high-quality robotics education. Jiasi Gao, Haole Guo, Zhanxiang Cao, Guyue Zhou |
IROS | 5 |
| 2023 | Can Quadruped Guide Robots be Used as Guide Dogs?abstractQuadruped robots have the potential to guide blind and low vision (BLV) people due to their highly flexible locomotion and emotional value provided by their bionic forms. However, the development of quadruped guide robots rarely involves BLV users' participatory designs and evaluations. In this paper, we conducted two empirical experiments both in indoor controlled and outdoor field scenarios, exploring the benefits and drawbacks of quadruped guide robots. The results show that the nowadays commercial quadruped robots exposed significant disadvantages in usability and trust compared with wheeled robots. It is concluded that the moving gait and walking noise of quadruped robots would limit the guiding effectiveness to a certain extent, and the empathetic effect of its bionic form for BLV users could not be fully reflected. Based on the findings of wheeled robots and quadruped robots' advantages, we discuss the design implications for the future guide robot design for BLV users. This paper reports the first empirical experiment about quadruped guide robots with BLV users and preliminary explores their potential improvement space in substituting guide dogs, which can inspire the further specialized design of quadruped guide robots. Qihe Chen, Yan Zhang 0122, Tingmin Yan, Guyue Zhou, Jiangtao Gong |
IROS | 7 |
| 2023 | PAD: A Dataset and Benchmark for Pose-agnostic Anomaly DetectionabstractObject anomaly detection is an important problem in the field of machine vision and has seen remarkable progress recently. However, two significant challenges hinder its research and application. First, existing datasets lack comprehensive visual information from various pose angles. They usually have an unrealistic assumption that the anomaly-free training dataset is pose-aligned, and the testing samples have the same pose as the training data. However, in practice, anomaly may exist in any regions on a object, the training and query samples may have different poses, calling for the study on pose-agnostic anomaly detection. Second, the absence of a consensus on experimental protocols for pose-agnostic anomaly detection leads to unfair comparisons of different methods, hindering the research on pose-agnostic anomaly detection. To address these issues, we develop Multi-pose Anomaly Detection (MAD) dataset and Pose-agnostic Anomaly Detection (PAD) benchmark, which takes the first step to address the pose-agnostic anomaly detection problem. Specifically, we build MAD using 20 complex-shaped LEGO toys including 4K views with various poses, and high-quality and diverse 3D anomalies in both simulated and real environments. Additionally, we propose a novel method OmniposeAD, trained using MAD, specifically designed for pose-agnostic anomaly detection. Through comprehensive evaluations, we demonstrate the relevance of our dataset and method. Furthermore, we provide an open-source benchmark library, including dataset and baseline methods that cover 8 anomaly detection paradigms, to facilitate future research and application in this domain. Code, data, and models are publicly available at https://github.com/EricLee0224/PAD. Weize Li 0001, Lihan Jiang, Guoliang Wang 0002, Guyue Zhou, Shanghang Zhang, Hao Zhao 0002 |
NeurIPS | 5 |
| 2022 | Cerberus Transformer: Joint Semantic, Affordance and Attribute ParsingabstractMulti-task indoor scene understanding is widely considered as an intriguing formulation, as the affinity of different tasks may lead to improved performance. In this paper, we tackle the new problem of Joint semantic, affordance and attribute parsing. However, successfully resolving it requires a model to capture long-range dependency, learn from weakly aligned data and properly balance sub-tasks during training. To this end, we propose an attention-based architecture named Cerberus and a tailored training framework. Our method effectively addresses aforementioned challenges and achieves state-of-the-art performance on all three tasks. Moreover, an in-depth analysis shows concept affinity consistent with human cognition, which inspires us to explore the possibility of weakly supervised learning. Surprisingly, Cerberus achieves strong results using only 0.1%–1% annotation. Visualizations further confirm that this success is credited to common attention maps across tasks. Code and models can be accessed at https://github.com/OPEN-AIR-SUN/Cerberus. Xiaoxue Chen, Tianyu Liu 0008, Hao Zhao 0001, Guyue Zhou, Ya-Qin Zhang |
CVPR | 4 |
| 2022 | Brick Yourself within 3 MinutesabstractThis paper presents an intelligent machine which can automatically convert the captured portrait into a physical gadget made up of LEGO bricks. On the contrary to synthesising a 2D image or a virtual 3D object, generating physical 3D assembly object needs to take physical properties and assembly process into consideration, leading to more challenges. To generate brick models for arbitrary portraits, we formulate the transformation between the attribute space (extracted from 2D images) and the brick model space as a constraint integer programming problem which can be solved with a heuristic search method. Furthermore, as the bricks are physically scattered, we propose an algorithm to generate corresponding assembly instructions for customized figure-featured-bricks to facilitate users' assembly. Meanwhile, we deploy the proposed algorithms on an automatic machine which integrates a camera, a printer, a laptop, and a brick operation unit. Finally, the generated brick models and assembly instructions are evaluated by a large number of users. It is worth noting that the whole system works as an intelligent vending machine, producing a 150-brick-model within 3 minutes. Guyue Zhou, Liyi Luo, Haole Guo, Hao Zhao 0002 |
ICRA | 1 |
| 2022 | Learning with Yourself: a Tangible Twin Robot System to Promote STEM EducationabstractThis paper presents a customized programmable robotic system, TanTwin (Tangible Twin), designed to promote STEM education for K-12 children. Firstly, TanTwin is implemented based on a wheel-robot with standard LEGO bricks. With several deep neural networks, a child can convert a captured portrait of himself/herself into standard LEGO bricks, therefore he/she can build a tangible twin robot of him-selflherself automatically. Besides, to adapt to the customized appearance, the corresponding visual element and content of the robotic system were also changed by a rule-based adaption algorithm. To demonstrate the effectiveness of TanTwin and to investigate whether tangible twin robots could contribute to children's learning, we conducted a controlled experimental study to compare learning with a TanTwin and with a standard robot system through measuring students' cognitive learning outcomes. The pre-/post- knowledge test results indicated that learning with a tangible twin robot leads to significantly better learning outcomes. Given the results, we validate our system and customization technology can promote STEM education. Jiasi Gao, Jiangtao Gong, Guyue Zhou, Haole Guo, Tong Qi |
IROS | 3 |
| 2022 | TOIST: Task Oriented Instance Segmentation Transformer with Noun-Pronoun DistillationabstractCurrent referring expression comprehension algorithms can effectively detect or segment objects indicated by nouns, but how to understand verb reference is still under-explored. As such, we study the challenging problem of task oriented detection, which aims to find objects that best afford an action indicated by verbs like sit comfortably on. Towards a finer localization that better serves downstream applications like robot interaction, we extend the problem into task oriented instance segmentation. A unique requirement of this task is to select preferred candidates among possible alternatives. Thus we resort to the transformer architecture which naturally models pair-wise query relationships with attention, leading to the TOIST method. In order to leverage pre-trained noun referring expression comprehension models and the fact that we can access privileged noun ground truth during training, a novel noun-pronoun distillation framework is proposed. Noun prototypes are generated in an unsupervised manner and contextual pronoun features are trained to select prototypes. As such, the network remains noun-agnostic during inference. We evaluate TOIST on the large-scale task oriented dataset COCO-Tasks and achieve +10.7% higher $\rm{mAP^{box}}$ than the best-reported results. The proposed noun-pronoun distillation can boost $\rm{mAP^{box}}$ and $\rm{mAP^{mask}}$ by +2.6% and +3.6%. Codes and models are publicly available. Pengfei Li 0007, Beiwen Tian, Yongliang Shi, Xiaoxue Chen, Hao Zhao 0002, Guyue Zhou, Ya-Qin Zhang |
NeurIPS | 6 |
| 2022 | When to Trust Your Simulator: Dynamics-Aware Hybrid Offline-and-Online Reinforcement LearningabstractLearning effective reinforcement learning (RL) policies to solve real-world complex tasks can be quite challenging without a high-fidelity simulation environment. In most cases, we are only given imperfect simulators with simplified dynamics, which inevitably lead to severe sim-to-real gaps in RL policy learning. The recently emerged field of offline RL provides another possibility to learn policies directly from pre-collected historical data. However, to achieve reasonable performance, existing offline RL algorithms need impractically large offline data with sufficient state-action space coverage for training. This brings up a new question: is it possible to combine learning from limited real data in offline RL and unrestricted exploration through imperfect simulators in online RL to address the drawbacks of both approaches? In this study, we propose the Dynamics-Aware Hybrid Offline-and-Online Reinforcement Learning (H2O) framework to provide an affirmative answer to this question. H2O introduces a dynamics-aware policy evaluation scheme, which adaptively penalizes the Q function learning on simulated state-action pairs with large dynamics gaps, while also simultaneously allowing learning from a fixed real-world dataset. Through extensive simulation and real-world tasks, as well as theoretical analysis, we demonstrate the superior performance of H2O against other cross-domain online and offline RL algorithms. H2O provides a brand new hybrid offline-and-online RL paradigm, which can potentially shed light on future RL algorithm design for solving practical real-world tasks. Yiwen Qiu, Guyue Zhou, Jianming Hu, Xianyuan Zhan |
NeurIPS | 5 |
| 2022 | SNAKE: Shape-aware Neural 3D Keypoint FieldabstractDetecting 3D keypoints from point clouds is important for shape reconstruction, while this work investigates the dual question: can shape reconstruction benefit 3D keypoint detection? Existing methods either seek salient features according to statistics of different orders or learn to predict keypoints that are invariant to transformation. Nevertheless, the idea of incorporating shape reconstruction into 3D keypoint detection is under-explored. We argue that this is restricted by former problem formulations. To this end, a novel unsupervised paradigm named SNAKE is proposed, which is short for shape-aware neural 3D keypoint field. Similar to recent coordinate-based radiance or distance field, our network takes 3D coordinates as inputs and predicts implicit shape indicators and keypoint saliency simultaneously, thus naturally entangling 3D keypoint detection and shape reconstruction. We achieve superior performance on various public benchmarks, including standalone object datasets ModelNet40, KeypointNet, SMPL meshes and scene-level datasets 3DMatch and Redwood. Intrinsic shape awareness brings several advantages as follows. (1) SNAKE generates 3D keypoints consistent with human semantic annotation, even without such supervision. (2) SNAKE outperforms counterparts in terms of repeatability, especially when the input point clouds are down-sampled. (3) the generated keypoints allow accurate geometric registration, notably in a zero-shot setting. Codes and models are available at https://github.com/zhongcl-thu/SNAKE. Chengliang Zhong, Peixing You, Xiaoxue Chen, Hao Zhao 0002, Fuchun Sun 0001, Guyue Zhou, Xiaodong Mu, Chuang Gan 0001, Wenbing Huang 0001 |
NeurIPS | 6 |
| 2022 | Distance-Aware Occlusion Detection With Focused AttentionabstractFor humans, understanding the relationships between objects using visual signals is intuitive. For artificial intelligence, however, this task remains challenging. Researchers have made significant progress studying semantic relationship detection, such as human-object interaction detection and visual relationship detection. We take the study of visual relationships a step further from semantic to geometric. In specific, we predict relative occlusion and relative distance relationships. However, detecting these relationships from a single image is challenging. Enforcing focused attention to task-specific regions plays a critical role in successfully detecting these relationships. In this work, (1) we propose a novel three-decoder architecture as the infrastructure for focused attention; 2) we use the generalized intersection box prediction task to effectively guide our model to focus on occlusion-specific regions; 3) our model achieves a new state-of-the-art performance on distance-aware relationship detection. Specifically, our model increases the distance F1-score from 33.8% to 38.6% and boosts the occlusion F1-score from 34.4% to 41.2%. Our code and data will be publicly available. Yang Li 0178, Yucheng Tu, Xiaoxue Chen, Hao Zhao 0002, Guyue Zhou |
IEEE Trans. Image Process. | 5 |
| 2018 | FlyCap: Markerless Motion Capture Using Multiple Autonomous Flying CamerasabstractAiming at automatic, convenient and non-instrusive motion capture, this paper presents a new generation markerless motion capture technique, the FlyCap system, to capture surface motions of moving characters using multiple autonomous flying cameras (autonomous unmanned aerial vehicles(UAVs) each integrated with an RGBD video camera). During data capture, three cooperative flying cameras automatically track and follow the moving target who performs large-scale motions in a wide space. We propose a novel non-rigid surface registration method to track and fuse the depth of the three flying cameras for surface motion tracking of the moving target, and simultaneously calculate the pose of each flying camera. We leverage the using of visual-odometry information provided by the UAV platform, and formulate the surface tracking problem in a non-linear objective function that can be linearized and effectively minimized through a Gaussian-Newton method. Quantitative and qualitative experimental results demonstrate the plausible surface and motion reconstruction results. Lan Xu 0003, Yebin Liu, Guyue Zhou, Qionghai Dai, Lu Fang 0001 |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2017 | Beyond SIFT using binary features in Loop Closure DetectionabstractIn this paper a binary feature based Loop Closure Detection (LCD) method is proposed, which for the first time achieves higher precision-recall (PR) performance compared with state-of-the-art SIFT feature based approaches. The proposed system originates from our previous work Multi-Index hashing for Loop closure Detection (MILD), which employs Multi-Index Hashing (MIH) [1] for Approximate Nearest Neighbor (ANN) search of binary features. As the accuracy of MILD is limited by repeating textures and inaccurate image similarity measurement, burstiness handling is introduced to solve this problem and achieves considerable accuracy improvement. Additionally, a comprehensive theoretical analysis on MIH used in MILD is conducted to further explore the potentials of hashing methods for ANN search of binary features from probabilistic perspective. This analysis provides more freedom on best parameter choosing in MIH for different application scenarios. Experiments on popular public datasets show that the proposed approach achieved the highest accuracy compared with state-of-the-art while running at 30Hz for databases containing thousands of images. Guyue Zhou, Lan Xu 0003, Lu Fang 0001 |
IROS | 2 |
| 2014 | On-board inertial-assisted visual odometer on an embedded systemabstractIn this paper, we propose a novel inertial-assisted visual odometry system intended for low-cost micro aerial vehicles (MAVs). The system sensor assembly consists of two downward-facing cameras and an inertial measurement unit (IMU) with three-axis accelerometers/gyroscopes. Real-time implementation of the system is enabled by a low-cost embedded system via two important features: firstly, simple pixel-level algorithms are integrated in a low-end FPGA and accelerated via pipeline and combinational logic techniques; secondly, a fast yaw-and-translation estimation algorithm works well with a novel outlier rejection scheme based on probabilistic predetermined operations rather than hypothesis testing iterations. We illustrate the performance of our system by hovering a MAV in a GPS-denied environment. Its feasibility and robustness is also illustrated in complex outdoor environments. Guyue Zhou, Jiaxin Ye, Zexiang Li 0001 |
ICRA | 1 |
| 2013 | Personal photo album compression and managementabstractThe advance in multimedia technologies have resulted in an explosive growth of pictures in personal computers and in cloud. Typically many pictures taken in the same occasion are similar. The cost to store and transmit them can be very significant. Thus it is important to find an efficient method to store these pictures. This paper proposed a compression scheme for similar images. Our approach is to arrange all the similar images into tree structure then apply video coding technique along each branch. To maximize the inter-image correlation between adjacent photos, we consider the minimum spanning tree (MST) subjecting to a maximum depth limit to ensure fast access to all images. This structure is encoded by the latest video coding technique High Efficiency Video Coding (HEVC), which is reported to has advantage in high definition video/image compression. It also supports deleting, adding and modifying images. Experiments show that the proposed method saved 75% space comparing to JPEG format. Ruobing Zou, Oscar C. Au, Guyue Zhou, Wei Dai 0002, Wei Hu 0003, Pengfei Wan 0001 |
ISCAS | 3 |