EDBT 2026 Demo / reviewers in the wild / expert
Jia Pan 0001
dblp:97/896-1
· DBLP profile ↗
105ranked-venue papers
13as first author
63since 2021 · last 2026
0000-0001-9003-2054ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 75 · 10 first-author · 39 since 2021Systems, architecture and hardware · 47 · 5 first-author · 21 since 2021Graphics, computer vision, multimedia, augmented reality and games · 23 · 3 first-author · 17 since 2021Applied, interdisciplinary, general and emerging computing · 16 · 1 first-author · 12 since 2021Human-computer interaction and ubiquitous computing · 4 · 3 since 2021Computer networks · 3 · 3 since 2021Databases, data management, data science and information retrieval · 2 · 2 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | NeuPAN: Direct Point Robot Navigation with End-to-End Model-Based Learning (Abstract Reprint)abstractNavigating a nonholonomic robot in a cluttered, unknown environment requires accurate perception and precise motion control for real-time collision avoidance. This article presents neural proximal alternating-minimization network (NeuPAN): a real-time, highly accurate, map-free, easy-to-deploy, and environment-invariant robot motion planner. Leveraging a tightly coupled perception-to-control framework, NeuPAN has two key innovations compared to existing approaches: first, it directly maps raw point cloud data to a latent distance feature space for collision-free motion generation, avoiding error propagation from the perception to control pipeline; second, it is interpretable from an end-to-end model-based learning perspective. The crux of NeuPAN is solving an end-to-end mathematical model with numerous point-level constraints using a plug-and-play proximal alternating-minimization network, incorporating neurons in the loop. This allows NeuPAN to generate real-time, physically interpretable motions. It seamlessly integrates data and knowledge engines, and its network parameters can be fine-tuned via back propagation. We evaluate NeuPAN on a ground mobile robot, a wheel-legged robot, and an autonomous vehicle, in extensive simulated and real-world environments. Results demonstrate that NeuPAN outperforms existing baselines in terms of accuracy, efficiency, robustness, and generalization capabilities across various environments, including the cluttered sandbox, office, corridor, and parking lot. We show that NeuPAN works well in unknown and unstructured environments with arbitrarily shaped objects, transforming impassable paths into passable ones. Ruihua Han, Shuai Wang 0004, Zeqing Zhang, Shijie Lin, Cheng-Zhong Xu 0001, Yonina C. Eldar, Qi Hao 0003, Jia Pan 0001 |
AAAI | 11 |
| 2026 | BiKC+: Bimanual Hierarchical Imitation With Keypose-Conditioned Coordination-Aware Consistency PoliciesabstractRobots are essential in industrial manufacturing due to their reliability and efficiency. They excel in performing simple and repetitive unimanual tasks but still face challenges with bimanual manipulation. This difficulty arises from the complexities of coordinating dual arms and handling multi-stage processes. Recent integration of generative models into imitation learning (IL) has made progress in tackling specific challenges. However, few approaches explicitly consider the multi-stage nature of bimanual tasks while also emphasizing the importance of inference speed. In multi-stage tasks, failures or delays at any stage can cascade over time, impacting the success and efficiency of subsequent sub-stages and ultimately hindering overall task performance. In this paper, we propose a novel keypose-conditioned coordination-aware consistency policy tailored for bimanual manipulation. Our framework instantiates hierarchical imitation learning with a high-level keypose predictor and a low-level trajectory generator. The predicted keyposes serve as sub-goals for trajectory generation, indicating targets for individual sub-stages. The trajectory generator is formulated as a consistency model, generating action sequences based on historical observations and predicted keyposes in a single inference step. In particular, we devise an innovative approach for identifying bimanual keyposes, considering both robot-centric action features and task-centric operation styles. Simulation and real-world experiments illustrate that our approach significantly outperforms baseline methods in terms of success rates and operational efficiency. Hang Xu 0001, Dongjie Yu, Jia Pan 0001 |
IEEE Trans Autom. Sci. Eng. | 5 |
| 2026 | Internal State Estimation in Crowds via Active Information GatheringabstractAccurately estimating human internal states, such as personality traits or behavioral patterns, is critical for enhancing the effectiveness of human–robot interaction, particularly in multi-agent settings. These insights are key in applications ranging from social navigation to autism diagnosis. However, prior methods are limited by scalability and passive observation, making real-time estimation in complex, multi-human settings difficult. In this work, we propose a practical method for active human personality estimation in crowds, with a focus on applications related to Autism Spectrum Disorder (ASD). Our method combines a personality-conditioned behavior model, based on the Eysenck 3-Factor theory, with an active robot information-gathering policy that triggers human behaviors through a receding-horizon planner. The robot’s belief about human personality is then updated via Bayesian inference. We demonstrate the effectiveness of our approach through proof-of-concept studies in simulation, user studies with typical adults, and preliminary experiments involving participants with ASD. Our results show that our method can scale to tens of humans and reduce personality estimation error by 29.2% and uncertainty by 79.9% in simulation compared to the passive baseline. User studies with typical adults confirm the method’s ability to generalize across complex personality distributions. Additionally, we explore its application in autism-related scenarios, demonstrating that the method can identify the difference between neurotypical and autistic behavior. The results suggest that our framework could serve as a foundation for future ASD-specific applications. Xuebo Ji, Zherong Pan, Xifeng Gao, Lei Yang 0048, Xinxin Du, Kaiyun Li, Yong-Jin Liu 0001, Wenping Wang 0001, Changhe Tu, Jia Pan 0001 |
ACM Trans. Hum. Robot Interact. | 10 |
| 2026 | EROAM: Event-Based Camera Rotational Odometry and Mapping in Real Time
Wanli Xing 0002, Shijie Lin, Linhan Yang, Zeqing Zhang, Yanjun Du, Maolin Lei, Yipeng Pan, Chen Wang 0123, Jia Pan 0001 |
IEEE Trans. Robotics | 9 |
| 2026 | DecoRec: Decomposed 3D Scene Reconstruction From Single-View Images via Object-Level DiffusionabstractIn this paper, we introduce DecoRec, a novel system designed to elevate single-view 2D images to a decomposed 3D scene mesh. Current methods for single-view scene reconstruction typically rely on object retrieval or the regression of coarse 3D voxels or surfaces, leading to inaccuracies in capturing the appearance and geometry of the input image. The lack of high-quality large-scale scene-level datasets further complicates direct 3D scene generation from single-view images. To achieve high-quality 3D scene generation from a single-view image, DecoRec takes advantage of recent diffusion-based single-view object reconstruction methods to reconstruct individual objects separately. Subsequently, a refinement pipeline is proposed to effectively merge these reconstructed objects, enhancing appearance and geometry through a differentiable rendering technique and diffusion-guided refinement. Our results demonstrate that DecoRec facilitates high-quality single-view scene reconstruction in both geometry and novel synthesis, offering significant benefits for downstream applications like room interior design. Yuhan Ping, Yuan Liu 0025, Xiaoxiao Long, Peng Wang 0099, Junhui Hou, Jianyi Zheng, Jia Pan 0001, Xin Li 0003, Cheng Lin 0001 |
IEEE Trans. Vis. Comput. Graph. | 7 |
| 2025 | Optimizing Efficiency of Mixed Traffic Through Reinforcement Learning: A Topology-Independent Approach and BenchmarkabstractThis paper presents a mixed traffic control policy designed to optimize traffic efficiency across diverse road topologies, addressing issues of congestion prevalent in urban environments. A model-free reinforcement learning (RL) approach is developed to manage large-scale traffic flow, using data collected by autonomous vehicles to influence human-driven vehicles. A real-world mixed traffic control benchmark is also released, which includes 444 scenarios from 20 countries, representing a wide geographic distribution and covering a variety of scenarios and road topologies. This benchmark serves as a foundation for future research, providing a realistic simulation environment for the development of effective policies. Comprehensive experiments demonstrate the effectiveness and adaptability of the proposed method, achieving better performance than existing traffic control methods in both intersection and roundabout scenarios. To the best of our knowledge, this is the first project to introduce a real-world complex scenarios mixed traffic control benchmark. Videos and code of our work are available at https://sites.google.com/berkeley.edu/mixedtrafficplus/home Chuyang Xiao, Dawei Wang 0006, Xinzheng Tang, Jia Pan 0001, Yuexin Ma |
ICRA | 4 |
| 2025 | Magnetometer-Calibrated Hybrid Transformer for Robust Inertial Tracking in RoboticsabstractInertial tracking is vital for autonomous robots and has gained popularity with the ubiquity of low-cost Inertial Measurement Units (IMUs) and deep learning-powered tracking algorithms. Existing works, however, have not fully utilized IMU measurements, particularly magnetometers, nor maximized the potential of deep learning to achieve the desired accuracy. To bridge the gap, we introduce NeurIT, which employs a Time-Frequency Block-recurrent Transformer (TF-BRT) at its core, combining RNN and Transformer to learn both time-frequency representative features. To fully utilize IMU information, we strategically employ differentiation of body-frame magnetometers for orientation calibration in a sensor fusion manner. Experiments conducted in diverse environments show that NeurIT maintains a mere 1 -meter tracking error over a 300 - meter distance, surpassing state-of-the-art baselines by 48.21 % on unseen data. NeurIT also performs comparably to the visual-inertial approach (Tango Phone) in vision-favored conditions and surpasses it in plain environments. We share the code and data to promote further research: https://github.com/aiot-lab/NeurIT. Xinzhe Zheng 0001, Sijie Ji, Yipeng Pan, Kaiwen Zhang 0016, Jia Pan 0001, Chenshu Wu |
ICRA | 5 |
| 2025 | EventSync: Joint Recovery of Temporal Offsets and Relative Orientations for Wide-Baseline Event CamerasabstractEvent-Based cameras offer significant advantages due to their high temporal resolution and low power consumption. However, when deploying multiple such cameras, a critical challenge emerges: each camera operates on an independent time system, resulting in temporal misalignment that severely degrades performance in multi-event camera applications. Traditional hardware-based synchronization methods face significant limitations in compatibility and are impractical for wide-baseline configurations. We introduce EventSync, a software-based algorithm that achieves millisecond-level synchronization by exploiting the motion of objects in the cameras’ shared field of view, while simultaneously estimating the relative orientation between cameras. Our approach eliminates the need for physical connections, making it particularly valuable for wide-baseline deployments. Through comprehensive evaluation in both simulated environments and real-world indoor/outdoor scenarios, we demonstrate robust synchronization accuracy and precise extrinsic calibration across varying camera configurations, significantly outperforming existing methods. Code: https://github.com/wlxing1901/event-sync Wanli Xing 0002, Shijie Lin, Guangze Zheng 0001, Linhan Yang, Yanjun Du, Jia Pan 0001 |
IROS | 6 |
| 2025 | Lattice Boltzmann Model for Learning Real-World Pixel DynamicityabstractThis work proposes the Lattice Boltzmann Model (LBM) to learn real-world pixel dynamicity for visual tracking.
LBM decomposes visual representations into dynamic pixel lattices and solves pixel motion states through collision-streaming processes.
Specifically, the high-dimensional distribution of the target pixels is acquired through a multilayer predict-update network to estimate the pixel positions and visibility. The predict stage formulates lattice collisions among the spatial neighborhood of target pixels and develops lattice streaming within the temporal visual context. The update stage rectifies the pixel distributions with online visual representations. Compared with existing methods, LBM demonstrates practical applicability in an online and real-time manner, which can efficiently adapt to real-world visual tracking tasks. Comprehensive evaluations of real-world point tracking benchmarks such as TAP-Vid and RoboTAP validate LBM's efficiency. A general evaluation of large-scale open-world object tracking benchmarks such as TAO, BFT, and OVT-B further demonstrates LBM's real-world practicality. Guangze Zheng 0001, Shijie Lin, Haobo Zuo, Si Si, Ming-Shan Wang, Changhong Fu 0001, Jia Pan 0001 |
NeurIPS | 7 |
| 2025 | Research on Enhanced Gait Phase Segmentation Based on Multimodal Spatiotemporal Information FusionabstractGait phase segmentation, pivotal for understanding lower limb motion, finds applications in diverse fields like medicine and sports. While existing method often struggle with accuracy and adaptability in real-world settings, this study presents a novel methodology employing particle filters for precise lower limb motion capture (MoCap) utilizing inertial sensors, which can be used in more everyday environments and in a wider range of applications over a longer period of time. The innovative approach adeptly tracks walking movements, labeling six gait phases via skeleton reconstruction facilitated by the MoCap algorithm. Subsequently, we propose a neural network architecture amalgamating temporal convolutional network (TCN), graph convolutional network (GCN), and long short-term memory (LSTM). This architecture integrates raw data from inertial sensors with joint angles derived from reconstructed motion, achieving accurate segmentation of the six gait phases. Experimental validation compares the MoCap algorithm against an optical motion capture system, and the neural network’s performance against state-of-the-art methods. Results demonstrate our method’s superior accuracy of 96.94%, highlighting its efficacy in addressing gait phase segmentation challenges and propelling advancements in gait analysis. Hao Zhang 0170, Xiaofeng Liu 0006, Jie Li 0009, Jia Pan 0001, Chu Kiong Loo, Angelo Cangelosi |
IEEE Internet Things J. | 4 |
| 2025 | Talking Face Generation With Lip and Identity PriorsabstractABSTRACT Speech‐driven talking face video generation has attracted growing interest in recent research. While person‐specific approaches yield high‐fidelity results, they require extensive training data from each individual speaker. In contrast, general‐purpose methods often struggle with accurate lip synchronization, identity preservation, and natural facial movements. To address these limitations, we propose a novel architecture that combines an alignment model with a rendering model. The rendering model synthesizes identity‐consistent lip movements by leveraging facial landmarks derived from speech, a partially occluded target face, multi‐reference lip features, and the input audio. Concurrently, the alignment model estimates optical flow using the occluded face and a static reference image, enabling precise alignment of facial poses and lip shapes. This collaborative design enhances the rendering process, resulting in more realistic and identity‐preserving outputs. Extensive experiments demonstrate that our method significantly improves lip synchronization and identity retention, establishing a new benchmark in talking face video generation. Frederick W. B. Li, Gary K. L. Tam, Bailin Yang, Fangzhe Nan, Jia Pan 0001 |
Comput. Animat. Virtual Worlds | 6 |
| 2025 | A Coarse-to-Fine Robotic Fabric Alignment System Integrating Visual Servoing and Admittance ControlabstractFabric alignment is essential to key production processes such as cutting, sewing, and fusing in garment manufacturing. Traditionally, this task has relied heavily on the dexterity and expertise of skilled human workers. Although automated systems have been introduced, they often lack the flexibility required for complex alignment tasks. In this paper, we present a novel robotic fabric alignment framework that fully automates the process with high precision and adaptability. First, we propose a coarse-to-fine alignment strategy, where an initial imprecise target position is roughly computed based on a basic perception module and eye-to-hand calibration. This is followed by a sliding mode control (SMC)-based visual servoing approach (in an eye-in-hand configuration) to ensure a close-up view of feedback features for the fine alignment process. Additionally, we consider system disturbances estimated by a fuzzy logic system (FLS) and combine it with the controller to further enhance the system’s robustness. Finally, we developed an advanced end-effector equipped with force/torque (F/T) sensors and air-powered needle grippers for gentle fabric manipulation using admittance control. We validate our framework through a series of experiments that demonstrate its effectiveness in fabric alignment tasks. Jiaming Qi, Liang Lu 0005, Lei Yang 0048, Yan Ding 0002, Pai Zheng, David Navarro-Alarcon, Jia Pan 0001, Peng Zhou 0018 |
IEEE Trans Autom. Sci. Eng. | 7 |
| 2025 | Swarm Robotic Flocking With Aggregation Ability PrivacyabstractWe address the challenge of achieving flocking behavior in swarm robotic systems without compromising the privacy of individual robots’ aggregation capabilities. Traditional flocking algorithms are susceptible to privacy breaches, as adversaries can deduce the identity and aggregation abilities of robots by observing their movements. We introduce a novel control mechanism for privacy-preserving flocking, leveraging the Laplace mechanism within the framework of differential privacy. Our method mitigates privacy breaches by introducing a controlled level of noise, thus obscuring sensitive information. We explore the trade-off between privacy and utility by varying the differential privacy parameter$\epsilon$. Our quantitative analysis reveals that$\epsilon \leq 0.13$represents a lower threshold where private information is almost completely protected, whereas$\epsilon \geq 0.85$marks an upper threshold where private information cannot be protected at all. Empirical results validate that our approach effectively maintains privacy of the robots’ aggregation abilities throughout the flocking process.Note to Practitioners—This paper was motivated by the problem of preserving privacy of individual robots in a swarm robotic system. Existing approaches to address this issue generally consider that accomplishing complex tasks requiring explicit information sharing between robots, while explicit communication in public channel carries the risk of information leakage. It is not always like this in real adversarial environments, and this assumption restricts the investigation of privacy in autonomous systems. This paper suggests that an individual robot can use its sensors onboard to perceive states of other neighbors in a distributed way without explicit communication. Despite avoiding information leakage during explicit information sharing between robots, the configuration of swarm can still reveal sensitive information about the ability of each robot. In this paper, we propose a privacy-preserving approach for flocking control using the Laplace mechanism based on the concept of differential privacy. The solution prevents an adversary with full knowledge of the swarm’s configuration from learning the sensitive information of individual robots, thus ensuring the security of swarm robots in terms of sensitive information during ongoing missions. Shuai Zhang 0024, Yunke Huang, Weizi Li, Jia Pan 0001 |
IEEE Trans Autom. Sci. Eng. | 4 |
| 2025 | NeuPAN: Direct Point Robot Navigation With End-to-End Model-Based LearningabstractNavigating a nonholonomic robot in a cluttered, unknown environment requires accurate perception and precise motion control for real-time collision avoidance. This article presents neural proximal alternating-minimization network (NeuPAN): a real-time, highly accurate, map-free, easy-to-deploy, and environment-invariant robot motion planner. Leveraging a tightly coupled perception-to-control framework, NeuPAN has two key innovations compared to existing approaches: first, it directly maps raw point cloud data to a latent distance feature space for collision-free motion generation, avoiding error propagation from the perception to control pipeline; second, it is interpretable from an end-to-end model-based learning perspective. The crux of NeuPAN is solving an end-to-end mathematical model with numerous point-level constraints using a plug-and-play proximal alternating-minimization network, incorporating neurons in the loop. This allows NeuPAN to generate real-time, physically interpretable motions. It seamlessly integrates data and knowledge engines, and its network parameters can be fine-tuned via backpropagation. We evaluate NeuPAN on a ground mobile robot, a wheel-legged robot, and an autonomous vehicle, in extensive simulated and real-world environments. Results demonstrate that NeuPAN outperforms existing baselines in terms of accuracy, efficiency, robustness, and generalization capabilities across various environments, including the cluttered sandbox, office, corridor, and parking lot. We show that NeuPAN works well in unknown and unstructured environments with arbitrarily shaped objects, transforming impassable paths into passable ones. Ruihua Han, Shuai Wang 0004, Zeqing Zhang, Shijie Lin, Cheng-Zhong Xu 0001, Yonina C. Eldar, Qi Hao 0003, Jia Pan 0001 |
IEEE Trans. Robotics | 11 |
| 2025 | Autonomous Tomato Harvesting With Top-Down Fusion Network for Limited DataabstractUsing robots for tomato truss harvesting represents a promising approach to agricultural production. However, incomplete acquisition of perception information and clumsy operations often result in low harvest success rates or crop damage. To address this issue, we designed a new method for tomato truss perception, an autonomous harvesting method, and a novel circular rotary cutting end-effector. The robot performs object detection and keypoint detection on tomato trusses using the proposed Top-down Fusion Network, making decisions on suitable targets for harvesting based on phenotyping and pose estimation. The designed end-effector moves gradually from the bottom up to wrap around the tomato truss, cutting the peduncle to complete the harvest. Experiments conducted in real-world scenarios for robotic perception and autonomous harvesting of tomato trusses show that the proposed method increases accuracy by up to 11.42% and 22.29% for complete and limited dataset conditions, compared to baseline models. Furthermore, we have implemented an automatic tomato harvesting system based on TDFNet, which reaches an average harvest success rate of 89.58% in the greenhouse. Xingxu Li, Yiheng Han, Nan Ma 0012, Yong-Jin Liu 0001, Jia Pan 0001, Siyi Zheng |
IEEE Trans. Robotics | 5 |
| 2025 | A Potential Field Method for Tooth Motion Planning in Orthodontic TreatmentabstractInvisible orthodontics, commonly known as clear alignment treatment, offers a more comfortable and aesthetically pleasing alternative in orthodontic care, attracting considerable attention in the dental community in recent years. It replaces conventional metal braces with a series of removable, and transparent aligners. Each aligner is crafted to facilitate a gradual adjustment of the teeth, ensuring progressive stages of dental correction. This necessitates the design for teeth motion. Here we present an automatic method and a system for generating collision-free teeth motion planning while avoiding gaps between adjacent teeth, which is unacceptable in clinical practice. To tackle this task, we formulate it as a constrained optimization problem and utilize the interior point method for its solution. We also developed an interactive system that enables dentists to easily visualize and edit the paths. Our method significantly speeds up the clear aligner planning process, creating the desired motion paths for a full set of teeth in under five minutes-a task that typically requires several hours of manual work. Our experiments and user studies confirm the effectiveness of this method in planning teeth movement, showcasing its potential to streamline orthodontic procedures. Yuexin Ma, Lei Yang 0048, Congyi Zhang 0001, Guangshun Wei, Runnan Chen, Min Gu 0003, Jia Pan 0001, Zhengbao Yang, Taku Komura, Shi-Qing Xin, Yuanfeng Zhou, Changhe Tu, Wenping Wang 0001 |
IEEE Trans. Vis. Comput. Graph. | 8 |
| 2025 | A Rule-Based Optimization Method for Tooth AlignmentabstractWhile tooth alignment is crucial for digital dentistry, especially in orthodontic treatment, existing computer-aided methods mainly focus on the 3D dental crown but overlook the entire teeth, which is essential for applications in orthodontics. Besides, clinical orthodontic rules are not fully considered in these methods, i.e., there should be no collisions and gaps between teeth, the upper jaw and lower jaw should have correct occlusion relationships, the teeth should comply with a reasonable dental arch curve, etc. To generate optimal tooth alignment results, we propose a rule-based optimization method for solving the tooth alignment problem that takes into consideration the clinical rules functionally and aesthetically. We optimize rule-driven objective functions by adjusting the 6-DoF transformations of each tooth. Besides, our optimization formulation supports customization for different clinical scenarios by specifying the various energy terms. Extensive experiments, ablation studies, and user studies have been conducted to validate the effectiveness of our method. Quantitative and qualitative comparisons demonstrate that our method generates better tooth alignments than previous methods. Yuhan Ping, Guodong Wei, Guangshun Wei, Congyi Zhang 0001, Noha A. SAID, Jia Pan 0001, Shi-Qing Xin, Yuanfeng Zhou, Changhe Tu, Min Gu 0003, Wenping Wang 0001 |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2024 | NetTrack: Tracking Highly Dynamic Objects with a NetabstractThe complex dynamicity of open-world objects presents non-negligible challenges for multi-object tracking (MOT), often manifested as severe deformations, fast motion, and occlusions. Most methods that solely depend on coarse-grained object cues, such as boxes and the overall appearance of the object, are susceptible to degradation due to distorted internal relationships of dynamic objects. To address this problem, this work proposes Net Track, an efficient, generic, and affordable tracking framework to introduce fine-grained learning that is robust to dynamicity. Specifically, N etTrack constructs a dynamicity-aware association with a fine-grained Net, leveraging point-level visual cues. Correspondingly, a fine-grained sampler and matching method have been incorporated. Furthermore, NetTrack learns object-text correspondence for fine-grained localization. To evaluate MOT in extremely dynamic open-world scenarios, a bird flock tracking (BFT) dataset is constructed, which exhibits high dynamicity with diverse species and open-world scenarios. Comprehensive evaluation on BFT validates the effectiveness of fine-grained learning on object dynamicity, and thorough transfer experiments on challenging open-world benchmarks, i.e., TAO, TAO-OW, AnimalTrack, and GMOT-40, validate the strong generalization ability of NetTrack even without finetuning. Guangze Zheng 0001, Shijie Lin, Haobo Zuo, Changhong Fu 0001, Jia Pan 0001 |
CVPR | 5 |
| 2024 | LASIL: Learner-Aware Supervised Imitation Learning For Long-Term Microscopic Traffic SimulationabstractMicroscopic traffic simulation plays a crucial role in transportation engineering by providing insights into in-dividual vehicle behavior and overall traffic flow. How-ever, creating a realistic simulator that accurately repli-cates human driving behaviors in various traffic conditions presents significant challenges. Traditional simulators relying on heuristic models often fail to deliver accurate simulations due to the complexity of real-world traffic environments. Due to the covariate shift issue, existing imitation learning-based simulators often fail to generate stable long-term simulations. In this paper, we propose a novel approach called learner-aware supervised imitation learning to address the covariate shift problem in multi-agent imi-tation learning. By leveraging a variational autoencoder simultaneously modeling the expert and learner state distribution, our approach augments expert states such that the augmented state is aware of learner state distribution. Our method, applied to urban traffic simulation, demon-strates significant improvements over existing state-of-the-art baselines in both short-term microscopic and long-term macroscopic realism when evaluated on the real-world dataset pNEUMA. Zhenwei Miao, Weizi Li, Dayang Hao, Jia Pan 0001 |
CVPR | 7 |
| 2024 | Efficient Semantic Segmentation for Compressed VideoabstractRobots, constrained by limited onboard computing resources, often encounter situations wherein high-resolution and high-bit-rate videos captured by their cameras necessitate compression before further analysis. In this paper, we propose a novel video semantic segmentation paradigm for compressed video. Specifically, our framework draws the inspiration from the principle of Wavelet Transform, and thus we design the network structure, WTDecomNet, approximating the decomposition of high-resolution image into its low-resolution counterpart and axial details. The aim is to well preserve the image content through decomposition and maintain model efficiency by obtaining semantics from low-resolution image. To facilitate this purpose, we propose an efficient axial subband approximation module for extracting axial details and a lightweight temporal alignment module for associating keyframes and non-keyframes of compressed video. Through comprehensive experiments, we show that our model can achieve the state-of-the-art performance on public benchmarks. Especially on CamVid, comparing to baseline, our proposed model reduces the computational overhead by ∼70% while improving mIoU by ∼4%. Qi Li 0038, Jia Pan 0001, Wenxi Liu |
ICRA | 4 |
| 2024 | Mixed Traffic Control and Coordination from PixelsabstractTraffic congestion is a persistent problem in our society. Previous methods for traffic control have proven futile in alleviating current congestion levels leading researchers to explore ideas with robot vehicles given the increased emergence of vehicles with different levels of autonomy on our roads. This gives rise to mixed traffic control, where robot vehicles regulate human-driven vehicles through reinforcement learning (RL). However, most existing studies use precise observations that require domain expertise and hand engineering for each road network’s observation space. Additionally, precise observations use global information, such as environment outflow, and local information, i.e., vehicle positions and velocities. Obtaining this information requires updating existing road infrastructure with vast sensor environments and communication to potentially unwilling human drivers. We consider image observations, a modality that has not been extensively explored for mixed traffic control via RL, as the alternative: 1) images do not require a complete re-imagination of the observation space from environment to environment; 2) images are ubiquitous through satellite imagery, in-car camera systems, and traffic monitoring systems; and 3) images only require communication to equipment. In this work, we show robot vehicles using image observations can achieve competitive performance to using precise information on environments, including ring, figure eight, intersection, merge, and bottleneck. In certain scenarios, our approach even outperforms using precision observations, e.g., up to 8% increase in average vehicle velocity in the merge environment, despite only using local traffic information as opposed to global traffic information. Michael Villarreal, Bibek Poudel, Jia Pan 0001, Weizi Li |
ICRA | 3 |
| 2024 | Progressive Representation Learning for Real-Time UAV TrackingabstractVisual object tracking has significantly promoted autonomous applications for unmanned aerial vehicles (UAVs). However, learning robust object representations for UAV tracking is especially challenging in complex dynamic environments, when confronted with aspect ratio change and occlusion. These challenges severely alter the original information of the object. To handle the above issues, this work proposes a novel progressive representation learning framework for UAV tracking, i.e., PRL-Track. Specifically, PRL-Track is divided into coarse representation learning and fine representation learning. For coarse representation learning, two innovative regulators, which rely on appearance and semantic information, are designed to mitigate appearance interference and capture semantic information. Furthermore, for fine representation learning, a new hierarchical modeling generator is developed to intertwine coarse object representations. Exhaustive experiments demonstrate that the proposed PRL-Track delivers exceptional performance on three authoritative UAV tracking benchmarks. Real-world tests indicate that the proposed PRL-Track realizes superior tracking performance with 42.6 frames per second on the typical UAV platform equipped with an edge smart camera. The code, model, and demo videos are available at https://github.com/vision4robotics/PRL-Track. Changhong Fu 0001, Xiang Lei, Haobo Zuo, Liangliang Yao, Guangze Zheng 0001, Jia Pan 0001 |
IROS | 6 |
| 2024 | Prompt-Driven Temporal Domain Adaptation for Nighttime UAV TrackingabstractNighttime UAV tracking under low-illuminated scenarios has achieved great progress by domain adaptation (DA). However, previous DA training-based works are deficient in narrowing the discrepancy of temporal contexts for UAV trackers. To address the issue, this work proposes a prompt-driven temporal domain adaptation training framework to fully utilize temporal contexts for challenging nighttime UAV tracking, i.e., TDA. Specifically, the proposed framework aligns the distribution of temporal contexts from daytime and nighttime domains by training the temporal feature generator against the discriminator. The temporal-consistent discriminator progressively extracts shared domain-specific features to generate coherent domain discrimination results in the time series. Additionally, to obtain high-quality training samples, a prompt-driven object miner is employed to precisely locate objects in unannotated nighttime videos. Moreover, a new benchmark for long-term nighttime UAV tracking is constructed. Exhaustive evaluations on both public and self-constructed nighttime benchmarks demonstrate the remarkable performance of the tracker trained in TDA framework, i.e., TDA-Track. Real-world tests at nighttime also show its practicality. The code and demo videos are available at https://github.com/vision4robotics/TDA-Track. Changhong Fu 0001, Yiheng Wang 0001, Liangliang Yao, Guangze Zheng 0001, Haobo Zuo, Jia Pan 0001 |
IROS | 6 |
| 2024 | DaDiff: Domain-aware Diffusion Model for Nighttime UAV TrackingabstractDomain adaptation is an inspiring solution to the misalignment issue of day/night image features for nighttime UAV tracking. However, the one-step adaptation paradigm is inadequate in addressing the prevalent difficulties posed by low-resolution (LR) objects when viewed from the UAVs at night, owing to the blurry edge contour and limited detail information. Moreover, these approaches struggle to perceive LR objects disturbed by nighttime noise. To address these challenges, this work proposes a novel progressive alignment paradigm, named domain-aware diffusion model (DaDiff), aligning nighttime LR object features to the daytime by virtue of progressive and stable generations. The proposed DaDiff includes an alignment encoder to enhance the detail information of nighttime LR objects, a tracking-oriented layer designed to achieve close collaboration with tracking tasks, and a successive distribution discriminator presented to distinguish different feature distributions at each diffusion timestep successively. Furthermore, an elaborate nighttime UAV tracking benchmark is constructed for LR objects, namely NUT-LR, consisting of 100 annotated sequences. Exhaustive experiments have demonstrated the robustness and feature alignment ability of the proposed DaDiff. The source code and video demo are available at https://github.com/vision4robotics/DaDiff. Haobo Zuo, Changhong Fu 0001, Guangze Zheng 0001, Liangliang Yao, Kunhan Lu, Jia Pan 0001 |
IROS | 6 |
| 2024 | Evolution Strategy and Controlled Residual Convolutional Neural Networks for ADC Calibration in the Absence of Ground TruthabstractCalibrating ADCs in the absence of ground truth presents a significant challenge for high-precision applications. This paper addresses this issue by introducing a novel two-step approach that combines evolutionary strategy and deep learning techniques. First, we employ covariance matrix adaptation evolution strategy to obtain ground truth signal samples with optimal SFDR values. This serves as a robust foundation for the subsequent calibration process. Second, we propose a new calibration neural network architecture called controlled residual convolutional neural networks. This architecture introduces a controlled residual branch within the network, allowing for more effective learning and calibration. The controlled residual branch is designed to adaptively adjust the network’s focus between the main and residual paths, thereby enhancing its calibration capabilities. Experimental results underscore the efficacy of our proposed method. Specifically, we observed a 29.01dB improvement in SFDR, representing the maximum enhancement relative to previous methods. These results validate the effectiveness of our approach in achieving high-precision ADC calibration without the need for the information of ground truth signals, thereby making it feasible for background calibration. Jia Pan 0001, Xizhu Peng |
ISCAS | 4 |
| 2024 | Large-Scale Mixed Traffic Control Using Dynamic Vehicle Routing and Privacy-Preserving CrowdsourcingabstractControlling and coordinating urban traffic flow through robot vehicles is emerging as a novel transportation paradigm for the future. While this approach garners growing attention from researchers and practitioners, effectively managing and coordinating large-scale mixed traffic remains a challenge. We introduce an effective framework for large-scale mixed traffic control via privacy-preserving crowdsourcing and dynamic vehicle routing. Our framework consists of three modules: a privacy-protecting crowdsensing method, a graph propagation-based traffic forecasting method, and a privacy-preserving route selection mechanism. We evaluate our framework using a real-world road network. The results show that our framework accurately forecasts traffic flow, efficiently mitigates network-wide RV shortage issue, and coordinates large-scale mixed traffic. Compared to other baseline methods, our framework not only reduces the RV shortage issue up to 69.4% but also reduces the average waiting time of all vehicles in the network up to 27%. Dawei Wang 0006, Weizi Li, Jia Pan 0001 |
IEEE Internet Things J. | 3 |
| 2024 | Monocular BEV Perception of Road Scenes via Front-to-Top View ProjectionabstractHD map reconstruction is crucial for autonomous driving. LiDAR-based methods are limited due to expensive sensors and time-consuming computation. Camera-based methods usually need to perform road segmentation and view transformation separately, which often causes distortion and missing content. To push the limits of the technology, we present a novel framework that reconstructs a local map formed by road layout and vehicle occupancy in the bird's-eye view given a front-view monocular image only. We propose a front-to-top view projection (FTVP) module, which takes the constraint of cycle consistency between views into account and makes full use of their correlation to strengthen the view transformation and scene understanding. In addition, we apply multi-scale FTVP modules to propagate the rich spatial information of low-level features to mitigate spatial deviation of the predicted object location. Experiments on public benchmarks show that our method achieves various tasks on road layout estimation, vehicle occupancy estimation, and multi-class semantic estimation, at a performance level comparable to the state-of-the-arts, while maintaining superior efficiency. Wenxi Liu, Qi Li 0038, Weixiang Yang, Yuanlong Yu 0001, Yuexin Ma, Shengfeng He, Jia Pan 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 8 |
| 2024 | Prototype learning based generic multiple object tracking via point-to-box supervision
Wenxi Liu, Qi Li 0038, Yinhua She, Yuanlong Yu 0001, Jia Pan 0001, Jason Gu |
Pattern Recognit. | 6 |
| 2024 | Learning Autonomous Viewpoint Adjustment from Human Demonstrations for TelemanipulationabstractTeleoperation systems find many applications from earlier search-and-rescue to more recent daily tasks. It is widely acknowledged that using external sensors can decouple the view of the remote scene from the motion of the robot arm during manipulation, facilitating the control task. However, this design requires the coordination of multiple operators or may exhaust a single operator as s/he needs to control both the manipulator arm and the external sensors. To address this challenge, our work introduces a viewpoint prediction model, the first data-driven approach that autonomously adjusts the viewpoint of a dynamic camera to assist in telemanipulation tasks. This model is parameterized by a deep neural network and trained on a set of human demonstrations. We propose a contrastive learning scheme that leverages viewpoints in a camera trajectory as contrastive data for network training. We demonstrated the effectiveness of the proposed viewpoint prediction model by integrating it into a real-world robotic system for telemanipulation. User studies reveal that our model outperforms several camera control methods in terms of control experience and reduces the perceived task load compared to manual camera control. As an assistive module of a telemanipulation system, our method significantly reduces task completion time for users who choose to adopt its recommendation. Ruixing Jia, Lei Yang 0048, Ying Cao 0001, Calvin K. L. Or, Wenping Wang 0001, Jia Pan 0001 |
ACM Trans. Hum. Robot Interact. | 6 |
| 2024 | Neuromorphic Synergy for Video BinarizationabstractBimodal objects, such as the checkerboard pattern used in camera calibration, markers for object tracking, and text on road signs, to name a few, are prevalent in our daily lives and serve as a visual form to embed information that can be easily recognized by vision systems. While binarization from intensity images is crucial for extracting the embedded information in the bimodal objects, few previous works consider the task of binarization of blurry images due to the relative motion between the vision sensor and the environment. The blurry images can result in a loss in the binarization quality and thus degrade the downstream applications where the vision system is in motion. Recently, neuromorphic cameras offer new capabilities for alleviating motion blur, but it is non-trivial to first deblur and then binarize the images in a real-time manner. In this work, we propose an event-based binary reconstruction method that leverages the prior knowledge of the bimodal target's properties to perform inference independently in both event space and image space and merge the results from both domains to generate a sharp binary image. We also develop an efficient integration method to propagate this binary image to high frame rate binary video. Finally, we develop a novel method to naturally fuse events and images for unsupervised threshold identification. The proposed method is evaluated in publicly available and our collected data sequence, and shows the proposed method can outperform the SOTA methods to generate high frame rate binary video in real-time on CPU-only devices. Shijie Lin, Xiang Zhang 0022, Lei Yang 0048, Lei Yu 0006, Wenping Wang 0001, Jia Pan 0001 |
IEEE Trans. Image Process. | 8 |
| 2024 | A Distributed Outmost Push Approach for Multirobot HerdingabstractThis article presents a distributed control strategy for herding groups of evaders towards a predefined goal region using a team of robotic herders. In herding problems, evaders tend to move away from each other to increase their coverage regions. This makes it challenging to develop control solutions since the wandering evaders need to be collected while driving the herd. To address this, we propose the distributed outmost push strategy, where each robotic herder pushes the evader that is farthest from the goal region. The intuition behind this strategy is that robotic herders should focus on evaders that are further from the goal region as they are more likely to be missed during the herding process. The outmost evaders are selected from the local field of view, and the robotic herders make decisions in a decentralized manner. We also analyze the convergence of the designed dynamics and the minimum sensing range required for herders. The proposal's effectiveness and generality are validated through numerical simulations and real robotic experiments. Shuai Zhang 0024, Xiaokang Lei, Mengyuan Duan, Xingguang Peng, Jia Pan 0001 |
IEEE Trans. Robotics | 5 |
| 2024 | Heterogeneous Targets Trapping With Swarm Robots by Using Adaptive Density-Based InteractionabstractHomogeneous swarm robots are of significant research interest due to their robustness, flexibility, and scalability in completing complex tasks across various applications. This paper focuses on trapping heterogeneous targets using swarm robots, with emphasis on their different strengths. These targets consist of weak, strong, and group-moving individuals, where stronger targets exhibit larger body size, higher physical strength, and stronger resistance ability. Our goal is to develop an adaptive controller that enables swarm robots to self-organize and distribute themselves to trap targets, while adjusting encirclement thickness and robot group size based on target strength. We leverage local implicit information generated from density-based interaction to improve inter-robot and robot-target interactions through adaptive allocation and transformation mechanisms. The feasibility of our approach is validated through numerical simulations and experiments involving up to 50 physical robots and one human-controlled transformer. Shuai Zhang 0024, Xiaokang Lei, Xingguang Peng, Jia Pan 0001 |
IEEE Trans. Robotics | 4 |
| 2023 | mCLIP: Multilingual CLIP via Cross-lingual TransferabstractGuanhua Chen, Lu Hou, Yun Chen, Wenliang Dai, Lifeng Shang, Xin Jiang, Qun Liu, Jia Pan, Wenping Wang. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Guanhua Chen 0001, Lu Hou 0002, Yun Chen 0007, Wenliang Dai, Lifeng Shang, Xin Jiang 0002, Qun Liu 0001, Jia Pan 0001, Wenping Wang 0001 |
ACL (1) | 8 |
| 2023 | Hierarchical Temporal Transformer for 3D Hand Pose Estimation and Action Recognition from Egocentric RGB VideosabstractUnderstanding dynamic hand motions and actions from egocentric RGB videos is a fundamental yet challenging task due to self-occlusion and ambiguity. To address occlusion and ambiguity, we develop a transformer-based framework to exploit temporal information for robust estimation. Noticing the different temporal granularity of and the semantic correlation between hand pose estimation and action recognition, we build a network hierarchy with two cascaded transformer encoders, where the first one exploits the short-term temporal cue for hand pose estimation, and the latter aggregates per-frame pose and object information over a longer time span to recognize the action. Our approach achieves competitive results on two first-person hand action benchmarks, namely FPHA and H2O. Extensive ablation studies verify our design choices. Yilin Wen 0001, Hao Pan 0001, Lei Yang 0048, Jia Pan 0001, Taku Komura, Wenping Wang 0001 |
CVPR | 4 |
| 2023 | Evaluating Explanation Methods for Vision-and-Language NavigationabstractThe ability to navigate robots with natural language instructions in an unknown environment is a crucial step for achieving embodied artificial intelligence (AI). With the improving performance of deep neural models proposed in the field of vision-and-language navigation (VLN), it is equally interesting to know what information the models utilize for their decision-making in the navigation tasks. To understand the inner workings of deep neural models, various explanation methods have been developed for promoting explainable AI (XAI). But they are mostly applied to deep neural models for image or text classification tasks and little work has been done in explaining deep neural models for VLN tasks. In this paper, we address these problems by building quantitative benchmarks to evaluate explanation methods for VLN models in terms of faithfulness. We propose a new erasure-based evaluation pipeline to measure the step-wise textual explanation in the sequential decision-making setting. We evaluate several explanation methods for two representative VLN models on two popular VLN datasets and reveal valuable findings through our experiments. Guanqi Chen, Lei Yang 0048, Guanhua Chen 0001, Jia Pan 0001 |
ECAI | 4 |
| 2023 | Fast Event-based Double Integral for Real-time RoboticsabstractMotion deblurring is a critical ill-posed problem that is important in many vision-based robotics applications. The recently proposed event-based double integral (EDI) provides a theoretical framework for solving the deblurring prob-lem with the event camera and generating clear images at high frame-rate. However, the original EDI is mainly designed for offline computation and does not support real-time requirement in many robotics applications. In this paper, we propose the fast EDI, an efficient implementation of EDI that can achieve real-time online computation on single-core CPU devices, which is common for physical robotic platforms used in practice. In experiments, our method can handle event rates at as high as 13 million event per second in a wide variety of challenging lighting conditions. We demonstrate the benefit on multiple downstream real-time applications, including localization, vi-sual tag detection, and feature matching. Shijie Lin, Yingqiang Zhang, Dongyue Huang, Jia Pan 0001 |
ICRA | 6 |
| 2023 | Vision-based Six-Dimensional Peg-in-Hole for Practical Connector InsertionabstractWe study six-dimensional (6D) perceptive peg-in-hole problem for practical connector insertion task in this paper. To enable the manipulator system to handle different types of pegs in complex environment, we develop a perceptive robotic assembly system that utilizes an in-hand RGB-D camera for peg-in-hole with multiple types of pegs. The proposed framework addresses the critical hole detection and pose estimation problem through combining the learning-based detection with model-based pose estimation strategies. By exploiting the structure of the peg-in-hole task, we consider a rectangle-shape based characterization for modeling the candidate socket. Such a characterization allows us to design simple learning-based methods to detect and estimate the 6D pose of the target socket that balances between processing speed and accuracy. To validate our method, we test the performance of the proposed perceptive peg-in-hole solution using a KUKA iiwa7 robotic arm to accomplish the socket insertion task with two types of practical sockets (RJ45/HDMI). Without the need of additional search, our method achieves an acceptable success rate in the connector insertion tasks. The results confirm the reliability of our method and show that our method is suitable for real world application. Kun Zhang 0017, Chen Wang 0123, Hua Chen 0007, Jia Pan 0001, Michael Yu Wang, Wei Zhang 0013 |
ICRA | 4 |
| 2023 | Polymer-Based Self-Calibrated Optical Fiber Tactile SensorabstractHuman skin can accurately sense the self-decoupled normal and shear forces when in contact with objects of different sizes. Although there exist many soft and conformable tactile sensors on robotic applications able to decouple the normal force and shear forces, the impact of the size of object in contact on the force calibration model has been commonly ignored. Here, using the principle that contact force can be derived from the light power loss in the soft optical fiber core, we present a soft tactile sensor that decouples normal and shear forces and calibrates the measurement results based on the object size, by designing a two-layered weaved polymer-based optical fiber anisotropic structure embedded in a soft elastomer. Based on the anisotropic response of optical fibers, we developed a linear calibration algorithm to simultaneously measure the size of the contact object and the decoupled normal and shear forces calibrated the object size. By calibrating the sensor at the robotic arm tip, we show that robots can reconstruct the force vector at an average accuracy of 0.15N for normal forces, 0.17N for shear forces in X-axis, and 0.18N for shear forces in Y-axis, within the sensing range of 0-2N in all directions, and the average accuracy of object size measurement of 0.4mm, within the test indenter diameter range of 5-12mm. Youcan Yan, Zeqing Zhang, Lei Yang 0048, Jia Pan 0001 |
IROS | 5 |
| 2023 | Tight Collision Probability for UAV Motion Planning in Uncertain EnvironmentabstractOperating unmanned aerial vehicles (UAVs) in complex environments that feature dynamic obstacles and external disturbances poses significant challenges, primarily due to the inherent uncertainty in such scenarios. Additionally, inaccurate robot localization and modeling errors further exacerbate these challenges. Recent research on UAV motion planning in static environments has been unable to cope with the rapidly changing surroundings, resulting in trajectories that may not be feasible. Moreover, previous approaches that have addressed dynamic obstacles or external disturbances in isolation are insufficient to handle the complexities of such environments. This paper proposes a reliable motion planning framework for UAVs, integrating various uncertainties into a chance constraint that characterizes the uncertainty in a probabilistic manner. The chance constraint provides a probabilistic safety certificate by calculating the collision probability between the robot's Gaussian-distributed forward reachable set and states of obstacles. To reduce the conservatism of the planned trajectory, we propose a tight upper bound of the collision probability and evaluate it both exactly and approximately. The approximated solution is used to generate motion primitives as a reference trajectory, while the exact solution is leveraged to iteratively optimize the trajectory for better results. Our method is thoroughly tested in simulation and real-world experiments, verifying its reliability and effectiveness in uncertain environments. Fu Zhang 0002, Fei Gao 0011, Jia Pan 0001 |
IROS | 4 |
| 2023 | POMDP-Guided Active Force-Based Search for Robotic InsertionabstractIn robotic insertion tasks where the uncertainty exceeds the allowable tolerance, a good search strategy is essential for successful insertion and significantly influences efficiency. The commonly used blind search method is time-consuming and does not exploit the rich contact information. In this paper, we propose a novel search strategy that actively utilizes the information contained in the contact configuration and shows high efficiency. In particular, we formulate this problem as a Partially Observable Markov Decision Process (POMDP) with carefully designed primitives based on an in-depth analysis of the contact configuration's static stability. From the formulated POMDP, we can derive a novel search strategy. Thanks to its simplicity, this search strategy can be incorporated into a Finite-State-Machine (FSM) controller. The behaviors of the FSM controller are realized through a low-level Cartesian Impedance Controller. Our method is based purely on the robot's proprioceptive sensing and does not need visual or tactile sensors. To evaluate the effectiveness of our proposed strategy and control framework, we conduct extensive comparison experiments in simulation, where we compare our method with the baseline approach. The results demonstrate that our proposed method achieves a higher success rate with a shorter search time and search trajectory length compared to the baseline method. Additionally, we show that our method is robust to various initial displacement errors. Chen Wang 0123, Haoxiang Luo, Kun Zhang 0017, Hua Chen 0007, Jia Pan 0001, Wei Zhang 0013 |
IROS | 5 |
| 2023 | ChatHRC: Personalized Human-Robot Collaboration using Fuzzy Reinforcement Learning with Natural Language RewardsabstractCollaboration between humans and robots can be challenging because robots may have difficulty understanding a specific person’s intentions, particularly in complicated tasks such as co-manipulation and assembly in computer, communication, and consumer electronics (3C) manufacturing. These tasks require different weights on accuracy and speed for various fabrication steps, making traditional physical interaction inadequate. In this paper, we introduce a fuzzy reinforcement learning-based admittance controller that can infer humans’ intentions not only through physical interaction but also through natural language. During training, the natural language is encoded into a reward term to help the robot reach the human-intended convergence point, allowing us to develop a “personalized” policy. During testing, the language serves as a tool to help the robot understand and obey humans’ intentions when physical interaction alone is insufficient. For example, if the user finds it difficult to push the robot and needs it to move faster, they can say “it’s really slow,” while a request for high-accuracy operation can be conveyed through “the damping is too small.” With this algorithm, the robot can comprehend the intentions and act accordingly in such situations. Further results and videos can be found at: https://sites.google.com/view/hri-nlp. Weifeng Lu, Yu Zheng 0001, Jia Pan 0001 |
RO-MAN | 4 |
| 2023 | Coorp: Satisfying Low-Latency and High-Throughput Requirements of Wireless Network for Coordinated Robotic LearningabstractIn coordinated robotic learning, multiple robots share the same wireless channel for communication, and bring together latency-sensitive (LS) network flows for control and bandwidth-hungry (BH) flows for distributed learning. Unfortunately, existing wireless network supporting systems cannot coordinate these two network flows to meet their own requirements: 1) prioritized contention systems (e.g., EDCA) prevent LS messages from timely acquiring the wireless channel because multiple wireless network interface cards (WNICs) with BH messages are contending for the channel 2) global planning systems (e.g., SchedWiFi) have to reserve a notable time window in the shared channel for each LS flow, suffering from severe bandwidth degradation (up to 42%). We present the coordinated preemption method to meet both requirements for LS flows and BH flows. Globally (among multiple robots), coordinated preemption eliminates unnecessary contention of BH flows by making them transmit in a round-robin manner, such that LS flows have the highest chance to win the contention against BH flows, without sacrificing overall bandwidth from the perspective of coordinated robotic learning applications. Locally (within the same robot), coordinated preemption in real time predicts the periodic transmission of LS flows from the upper application and conservatively limits packets of BH flows buffered in the WNIC only before LS packets arriving, reducing the bandwidth devoted to preemption. COORP, our implementation of coordinated preemption, reduced the violation of latency requirements from 53.9% (EDCA) to 8.8% (comparable to SchedWiFi). Regarding learning quality, COORP achieved a comparable (at times the same) learning reward with EDCA, which grew up to 76% faster than SchedWiFi. Shengliang Deng, Xiuxian Guan, Zekai Sun, Shixiong Zhao, Tianxiang Shen, Xusheng Chen, Tianyang Duan, Jia Pan 0001, Libo Zhang 0001, Heming Cui |
IEEE Internet Things J. | 9 |
| 2023 | Monocular Camera-Based Complex Obstacle Avoidance via Efficient Deep Reinforcement LearningabstractDeep reinforcement learning has achieved great success in laser-based collision avoidance works because the laser can sense accurate depth information without too much redundant data, which can maintain the robustness of the algorithm when it is migrated from the simulation environment to the real world. However, high-cost laser devices are not only difficult to deploy for a large scale of robots but also demonstrate unsatisfactory robustness towards the complex obstacles, including irregular obstacles, e.g., tables, chairs, and shelves, as well as complex ground and special materials. In this paper, we propose a novel monocular camera-based complex obstacle avoidance framework. Particularly, we innovatively transform the captured RGB images to pseudo-laser measurements for efficient deep reinforcement learning. Compared to the traditional laser measurement captured at a certain height that only contains one-dimensional distance information away from the neighboring obstacles, our proposed pseudo-laser measurement fuses the depth and semantic information of the captured RGB image, which makes our method effective for complex obstacles. We also design a feature extraction guidance module to weight the input pseudo-laser measurement, and the agent has more reasonable attention for the current state, which is conducive to improving the accuracy and efficiency of the obstacle avoidance policy. Besides, we adaptively add the synthesized noise to the laser measurement during the training stage to decrease the sim-to-real gap and increase the robustness of our model in the real environment. Finally, the experimental results show that our framework achieves state-of-the-art performance in several virtual and real-world scenarios. Jianchuan Ding, Lingping Gao, Wenxi Liu, Haiyin Piao, Jia Pan 0001, Zhenjun Du, Xin Yang 0011 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2023 | Scale-Aware Siamese Object Tracking for Vision-Based UAM ApproachingabstractIn many industrial applications of unmanned aerial manipulator (UAM), visual approaching the object is crucial to subsequent manipulating. In comparison with the widely-studied manipulating, the key to efficient vision-based UAM approaching, i.e., UAM object tracking, is still limited. Since traditional model-based UAM tracking is costly and cannot track arbitrary objects, an intuitive solution is to introduce state-of-the-art model-free Siamese trackers from the visual tracking field. Although Siamese tracking is most suitable for the onboard embedded processors, severe object scale variation in UAM tracking brings formidable challenges. To address these problems, this work proposes a novel model-free scale-aware Siamese tracker (SiamSA). Specifically, a scale attention network is proposed to emphasize scale awareness in feature processing. A scale-aware anchor proposal network is designed to achieve anchor proposing. Besides, two novel UAM tracking benchmarks are first recorded. Comprehensive experiments on benchmarks validate the effectiveness of SiamSA. Furthermore, real-world tests also confirm practicality for industrial UAM approaching tasks with high efficiency and robustness. Guangze Zheng 0001, Changhong Fu 0001, Junjie Ye 0004, Bowen Li 0007, Geng Lu, Jia Pan 0001 |
IEEE Trans. Ind. Informatics | 6 |
| 2023 | Grasping Living Objects With Adversarial Behaviors Using Inverse Reinforcement LearningabstractLiving objects are difficult to grasp since they can actively elude capture by adopting adversarial behaviors that are extremely hard to model or predict. In this case, an inappropriately strong contact force may hurt the struggling living objects and a grasping algorithm that can minimize the contact force whenever possible is required. To solve this challenging task, in this article, we present a reinforcement-learning (RL)-based algorithm with two stages: the pregrasp stage and the in-hand stage. In the pregrasp stage, the robot focuses on the living object's adversarial behavior and approaches it in a reliable manner. In particular, we use inverse RL to encode the living object's adversarial behavior into a reward function. The negative value of the learned reward function is then used to train a high-quality grasping policy that can compete with the living object's adversarial behavior with the RL framework. In the in-hand stage, we use RL to train a grasp policy such that the dexterous hand can grab the living object with the minimal force. A set of dense rewards are also specifically designed to encourage the robot to grasp and hold the living object persistently. To further improve the grasp performance, we explicitly take into account the structure of the dexterous robot hand by treating the hand as a graph and adopting graph convolutional network to formulate the grasping policy. We conduct a set of experiments to demonstrate the performance of our proposed method, in which the robot can grasp living objects with the success rate of 90% and 95% in the pregrasp and in-hand stages, respectively. The contact force applied by the robotic hand to the living object is dramatically reduced in comparison with the baseline grasping policy. Yu Zheng 0001, Jia Pan 0001 |
IEEE Trans. Robotics | 3 |
| 2023 | Sampling-Based Planning for Retrieving Near-Cylindrical Objects in Cluttered Scenes Using Hierarchical GraphsabstractWe present an incremental sampling-based task and motion planner for retrieving near-cylindrical objects, like bottle, in cluttered scenes, which computes a plan for removing obstacles to generate a collision-free motion of a robot to retrieve the target object. Our proposed planner uses a two-level hierarchy, including the first-level roadmap for the target object motion and the second-level retrieval graph for the entire robot motion, to aid in deciding the order and trajectory of object removal. We use an incremental expansion strategy to update the roadmap and retrieval graph from the collisions between the target object, the robot, and the obstacles, in order to optimize the object removal sequence. The performance of our method is highlighted in several benchmark scenes, including a fixed robotic arm in a cluttered scene with known obstacle locations and a scene, where locations of some objects or even the target object are unknown due to occlusions. Our method can also efficiently solve the high-dimensional planning problem of object retrieval using a mobile manipulator and be combined with the symbolic planner to plan complex multistep tasks. We deploy our method to a physical robot and integrate it with nonprehensile actions to improve operational efficiency. Compared to the state-of-the-art approaches, our method reduces task and motion planning time up to 24.6$\%$with a higher success rate, and still provides a near-optimal plan. Hao Tian 0003, Chaoyang Song 0001, Changbo Wang, Xinyu Zhang 0002, Jia Pan 0001 |
IEEE Trans. Robotics | 5 |
| 2022 | Towards Making the Most of Cross-Lingual Transfer for Zero-Shot Neural Machine TranslationabstractGuanhua Chen, Shuming Ma, Yun Chen, Dongdong Zhang, Jia Pan, Wenping Wang, Furu Wei. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Guanhua Chen 0001, Shuming Ma, Yun Chen 0007, Dongdong Zhang 0001, Jia Pan 0001, Wenping Wang 0001, Furu Wei |
ACL (1) | 5 |
| 2022 | End-to-End Trajectory Distribution Prediction Based on Occupancy Grid MapsabstractIn this paper, we aim to forecast a future trajectory distribution of a moving agent in the real world, given the social scene images and historical trajectories. Yet, it is a challenging task because the ground-truth distribution is unknown and unobservable, while only one of its samples can be applied for supervising model learning, which is prone to bias. Most recent works focus on predicting diverse trajectories in order to cover all modes of the real distribution, but they may despise the precision and thus give too much credit to unrealistic predictions. To address the issue, we learn the distribution with symmetric cross-entropy using occupancy grid maps as an explicit and scene-compliant approximation to the ground-truth distribution, which can effectively penalize unlikely predictions. In specific, we present an inverse reinforcement learning based multi-modal trajectory distribution forecasting framework that learns to plan by an approximate value iteration network in an end-to-end manner. Besides, based on the predicted distribution, we generate a small set of representative trajectories through a differentiable Transformer-based network, whose attention mechanism helps to model the relations of trajectories. In experiments, our method achieves state-of-the-art performance on the Stanford Drone Dataset and Intersection Drone Dataset. Wenxi Liu, Jia Pan 0001 |
CVPR | 3 |
| 2022 | Autofocus for Event CamerasabstractFocus control (FC) is crucial for cameras to capture sharp images in challenging real-world scenarios. The autofocus (AF) facilitates the FC by automatically adjusting the focus settings. However, due to the lack of effective AF methods for the recently introduced event cameras, their FC still relies on naive AF like manual focus adjustments, leading to poor adaptation in challenging real-world conditions. In particular, the inherent differences between event and frame data in terms of sensing modality, noise, temporal resolutions, etc., bring many challenges in designing an effective AF method for event cameras. To address these challenges, we develop a novel event-based autofocus framework consisting of an event-specific focus measure called event rate (ER) and a robust search strategy called event-based golden search (EGS). To verify the performance of our method, we have collected an event-based autofocus dataset (EAD) containing well-synchronized frames, events, and focal positions in a wide variety of challenging scenes with severe lighting and motion conditions. The experiments on this dataset and additional real-world scenarios demonstrated the superiority of our method over state-of-the-art approaches in terms of efficiency and accuracy. Shijie Lin, Yinqiang Zhang, Lei Yu 0006, Jia Pan 0001 |
CVPR | 6 |
| 2022 | High-resolution Face Swapping via Latent Semantics DisentanglementabstractWe present a novel high-resolution face swapping method using the inherent prior knowledge of a pre-trained GAN model. Although previous research can leverage generative priors to produce high-resolution results, their quality can suffer from the entangled semantics of the latent space. We explicitly disentangle the latent semantics by utilizing the progressive nature of the generator, deriving structure at-tributes from the shallow layers and appearance attributes from the deeper ones. Identity and pose information within the structure attributes are further separated by introducing a landmark-driven structure transfer latent direction. The disentangled latent code produces rich generative features that incorporate feature blending to produce a plausible swapping result. We further extend our method to video face swapping by enforcing two spatio-temporal constraints on the latent space and the image space. Extensive experiments demonstrate that the proposed method outperforms state-of-the-art image/video face swapping methods in terms of hallucination quality and consistency. Code can be found at: https://github.com/cnnlstm/FSLSD_HiRes. Yangyang Xu 0003, Bailin Deng, Junle Wang, Yanqing Jing, Jia Pan 0001, Shengfeng He |
CVPR | 5 |
| 2022 | Faithful Extreme Rescaling via Generative Prior Reciprocated Invertible RepresentationsabstractThis paper presents a Generative prior ReciprocAted Invertible rescaling Network (GRAIN) for generating faithful high-resolution (HR) images from low-resolution (LR) invertible images with an extreme upscaling factor (64×). Previous researches have leveraged the prior knowledge of a pretrained GAN model to generate high-quality upscaling results. However, they fail to produce pixel-accurate results due to the highly ambiguous extreme mapping process. We remedy this problem by introducing a reciprocated invertible image rescaling process, in which high-resolution information can be delicately embedded into an invertible low-resolution image and generative prior for a faithful HR reconstruction. In particular, the invertible LR features not only carry significant HR semantics, but also are trained to predict scale-specific latent codes, yielding a preferable utilization of generative features. On the other hand, the enhanced generative prior is re-injected to the rescaling process, compensating the lost details of the invertible rescaling. Our reciprocal mechanism perfectly integrates the advantages of invertible encoding and generative prior, leading to the first feasible extreme rescaling solution. Extensive experiments demonstrate superior performance against state-of-the-art upscaling methods. Code is available at https://github.com/cszzx/GRAIN. Zhixuan Zhong, Liangyu Chai, Yang Zhou 0038, Bailin Deng, Jia Pan 0001, Shengfeng He |
CVPR | 5 |
| 2022 | ModLaNets: Learning Generalisable Dynamics via Modularity and Physical Inductive BiasabstractDeep learning models are able to approximate one specific dynamical system but struggle at learning generalisable dynamics, where dynamical systems obey the same laws of physics but contain different numbers of elements (e.g., double- and triple-pendulum systems). To relieve this issue, we proposed the Modular Lagrangian Network (ModLaNet), a structural neural network framework with modularity and physical inductive bias. This framework models the energy of each element using modularity and then construct the target dynamical system via Lagrangian mechanics. Modularity is beneficial for reusing trained networks and reducing the scale of networks and datasets. As a result, our framework can learn from the dynamics of simpler systems and extend to more complex ones, which is not feasible using other relevant physics-informed neural networks. We examine our framework for modelling double-pendulum or three-body systems with small training datasets, where our models achieve the best data efficiency and accuracy performance compared with counterparts. We also reorganise our models as extensions to model multi-pendulum and multi-body systems, demonstrating the intriguing reusable feature of our framework. Yupu Lu, Shijie Lin, Guanqi Chen, Jia Pan 0001 |
ICML | 4 |
| 2022 | DynamicFilter: an Online Dynamic Objects Removal Framework for Highly Dynamic EnvironmentsabstractEmergence of massive dynamic objects will diversify spatial structures when robots navigate in urban environments. Therefore, the online removal of dynamic objects is critical. In this paper, we introduce a novel online removal framework for highly dynamic urban environments. The framework consists of the scan-to-map front-end and the map-to-map back-end modules. Both the front- and back-ends deeply integrate the visibility-based approach and map-based approach. The experiments validate the framework in highly dynamic simulation scenarios and real-world dataset. Tingxiang Fan, Bowen Shen, Hua Chen 0007, Wei Zhang 0013, Jia Pan 0001 |
ICRA | 5 |
| 2022 | Siamese Object Tracking for Vision-Based UAM Approaching with Pairwise Scale-Channel AttentionabstractAlthough the manipulating of the unmanned aerial manipulator (UAM) has been widely studied, vision-based UAM approaching, which is crucial to the subsequent manipulating, generally lacks effective design. The key to the visual UAM approaching lies in object tracking, while current UAM tracking typically relies on costly model-based methods. Besides, UAM approaching often confronts more severe object scale variation issues, which makes it inappro-priate to directly employ state-of-the-art model-free Siamese-based methods from the object tracking field. To address the above problems, this work proposes a novel Siamese network with pairwise scale-channel attention (SiamSA) for vision-based UAM approaching. Specifically, SiamSA consists of a pairwise scale-channel attention network (PSAN) and a scale-aware anchor proposal network (SA-APN). PSAN acquires valuable scale information for feature processing, while SA-APN mainly attaches scale awareness to anchor proposing. Moreover, a new tracking benchmark for UAM approaching, namely UAMT100, is recorded with 35K frames on a flying UAM platform for evaluation. Exhaustive experiments on the benchmarks and real-world tests validate the efficiency and practicality of SiamSA with a promising speed. Both the code and UAMT100 benchmark are now available at https://github.com/vision4robotics/SiamSA. Guangze Zheng 0001, Changhong Fu 0001, Junjie Ye 0004, Bowen Li 0007, Geng Lu, Jia Pan 0001 |
IROS | 6 |
| 2022 | Visual-tactile Sensing for Real-time Liquid Volume Estimation in GraspingabstractWe propose a deep visuo-tactile model for real-time estimation of the liquid inside a deformable container in a proprioceptive way. We fuse two sensory modalities, i.e., the raw visual inputs from the RGB camera and the tactile cues from our specific tactile sensor without any extra sensor calibrations. The robotic system is well controlled and adjusted based on the estimation model in real time. The main contributions and novelties of our work are listed as follows: 1) Explore a proprioceptive way for liquid volume estimation by developing an end-to-end predictive model with multi-modal convolutional networks, which achieve a high precision with an error of ~ 2 ml in the experimental validation. 2) Propose a multi-task learning architecture which comprehensively considers the losses from both classification and regression tasks, and comparatively evaluate the performance of each variant on the collected data and actual robotic platform. 3) Utilize the proprioceptive robotic system to accurately serve and control the requested volume of liquid, which is continuously flowing into a deformable container in real time. 4) Adaptively adjust the grasping plan to achieve more stable grasping and manipulation according to the real-time liquid volume prediction. Ruixing Jia, Lei Yang 0048, Youcan Yan, Zheng Wang 0002, Jia Pan 0001, Wenping Wang 0001 |
IROS | 6 |
| 2022 | Optimization-Based Online Flow Fields Estimation for AUVs Navigation
Hao Xu 0026, Yupu Lu, Jia Pan 0001 |
ISRR | 3 |
| 2021 | Projecting Your View Attentively: Monocular Road Scene Layout Estimation via Cross-View TransformationabstractHD map reconstruction is crucial for autonomous driving. LiDAR-based methods are limited due to the deployed expensive sensors and time-consuming computation. Camera-based methods usually need to separately perform road segmentation and view transformation, which often causes distortion and the absence of content. To push the limits of the technology, we present a novel framework that enables reconstructing a local map formed by road layout and vehicle occupancy in the bird’s-eye view given a front-view monocular image only. In particular, we propose a cross-view transformation module, which takes the constraint of cycle consistency between views into account and makes full use of their correlation to strengthen the view transformation and scene understanding. Considering the relationship between vehicles and roads, we also design a context-aware discriminator to further refine the results. Experiments on public benchmarks show that our method achieves the state-of-the-art performance in the tasks of road layout estimation and vehicle occupancy estimation. Especially for the latter task, our model outperforms all competitors by a large margin. Furthermore, our model runs at 35 FPS on a single GPU, which is efficient and applicable for real-time panorama HD map reconstruction. Weixiang Yang, Qi Li 0038, Wenxi Liu, Yuanlong Yu 0001, Yuexin Ma, Shengfeng He, Jia Pan 0001 |
CVPR | 7 |
| 2021 | Zero-Shot Cross-Lingual Transfer of Neural Machine Translation with Multilingual Pretrained EncodersabstractPrevious work mainly focuses on improving cross-lingual transfer for NLU tasks with a multilingual pretrained encoder (MPE), or improving the performance on supervised machine translation with BERT.However, it is under-explored that whether the MPE can help to facilitate the cross-lingual transferability of NMT model.In this paper, we focus on a zero-shot cross-lingual transfer task in NMT.In this task, the NMT model is trained with parallel dataset of only one language pair and an off-the-shelf MPE, then it is directly tested on zero-shot language pairs.We propose SixT, a simple yet effective model for this task.SixT leverages the MPE with a two-stage training schedule and gets further improvement with a position disentangled encoder and a capacity-enhanced decoder.Using this method, SixT significantly outperforms mBART, a pretrained multilingual encoderdecoder model explicitly designed for NMT, with an average improvement of 7.1 BLEU on zero-shot any-to-English test sets across 14 source languages.Furthermore, with much less training computation cost and training data, our model achieves better performance on 15 any-to-English test sets than CRISS and m2m-100, two strong multilingual NMT baselines. Guanhua Chen 0001, Shuming Ma, Yun Chen 0007, Li Dong 0004, Dongdong Zhang 0001, Jia Pan 0001, Wenping Wang 0001, Furu Wei |
EMNLP (1) | 6 |
| 2021 | An Overconstrained Robotic Leg with Coaxial Quasi-direct Drives for Omni-directional Ground MobilityabstractPlanar mechanisms dominate modern designs of legged robots with remote actuator placement for robust agility in ground mobility. This paper presents a novel design of robotic leg modules using the Bennett linkage, driven by two coaxially arranged quasi-direct actuators capable of omnidirectional ground locomotion. The Bennett linkage belongs to a family of overconstrained linkages with three-dimensional spatial motion and unparalleled joint axes. We present the first work regarding the design, modeling, and optimization of the Bennett leg module, enabling lateral locomotion, like the crabs, that was not capable with robotic legs designed with common planar mechanisms. We further explored the concept of overconstrained robots, which is a class of advanced robots based on the design reconfiguration of the Bennett leg modules, serving as a potential direction for future research. Shihao Feng, Yuping Gu, Weijie Guo 0003, Yuqin Guo, Fang Wan 0002, Jia Pan 0001, Chaoyang Song 0001 |
ICRA | 6 |
| 2021 | Efficient SE(3) Reachability Map Generation via Interplanar Integration of Intra-planar ConvolutionsabstractConvolution has been used for fast computation of reachability maps, but it has high computational costs when performing SE(3) convolution operations for general joint arrangements in industrial robots and 3D workspace. Its application is also limited to planar robots, 2D workspace, or robots with special spatial arrangements for joints. In this paper, we find that the SE(3) convolution can be decomposed into a set of SE(2) convolutions, which significantly reduces the computational complexity when computing the reachability map of high-DOF robotic manipulators in the 3D workspace. We also leverage GPU parallel computing and Fast Fourier transform to further accelerate the computation procedure. We demonstrate the time efficiency and quality of our approach using a set of numerical experiments for constructing reachability maps and also present a multi-robot plant phenotyping system that uses the computed reachability map for efficient viewpoint selection and path planning. Yiheng Han, Jia Pan 0001, Mengfei Xia, Long Zeng 0001, Yong-Jin Liu 0001 |
ICRA | 2 |
| 2021 | Encirclement Guaranteed Cooperative Pursuit with Robust Model Predictive ControlabstractThis paper studies a novel encirclement guaranteed cooperative pursuit problem involving N pursuers and a single evader in an unbounded two-dimensional game domain. Throughout the game, the pursuers are required to maintain encirclement of the evader, i.e., the evader should always stay inside the convex hull generated by all the pursuers, in addition to achieving the classical capture condition. To tackle this challenging cooperative pursuit problem, a robust model predictive control (RMPC) based formulation framework is first introduced, which simultaneously accounts for the encirclement and capture requirements under the assumption that the evader’s action is unavailable to all pursuers. Despite the reformulation, the resulting RMPC problem involves a bilinear constraint due to the encirclement requirement. To further handle such a bilinear constraint, a novel encirclement guaranteed partitioning scheme is devised that simplifies the original bilinear RMPC problem to a number of linear tube MPC (TMPC) problems solvable in a decentralized manner. Simulation experiments demonstrate the effectiveness of the proposed solution framework. Furthermore, comparisons with existing approaches show that the explicit consideration of the encirclement condition significantly improves the chance of successful capture of the evader in various scenarios. Chen Wang 0123, Hua Chen 0007, Jia Pan 0001, Wei Zhang 0013 |
IROS | 3 |
| 2021 | A Computational Framework for Robot Hand Design via Reinforcement LearningabstractRobot hand is essential for a fully functional robot and designing a good robot hand is a sophisticated job that challenges the designer’s knowledge and experience. This paper presents a computational framework for automatic optimal robot hand design based on reinforcement learning (RL), which considers desired grasping tasks, grasp control strategies, and performance quality measures altogether. The RL-based framework intends to grow finger joints with different types and link lengths at different positions from null. Then, the reward function for such a growing action is defined in terms of quality indexes of the generated robot hand to perform desired grasping tasks under expected control strategies. To demonstrate the effectiveness of this framework, in this paper we set the desired task to simply grasping objects of three primitive shapes (i.e., box, cylinder, and sphere) with predefined hand positions and strategies to close fingers to achieve grasps for each object. The force closure condition, quantitative stability indexes, and energy consumption of grasps as well as some penalty terms are used to assemble the reward function. Through simulation and practical prototype experiments, we show that capable robot hands can be automatically generated by the proposed framework. Potential factors that affect the output of the framework and deserve further exploration are also discussed. Zhong Zhang 0015, Yu Zheng 0001, Lezhang Liu, Xuan Zhao 0006, Xiong Li 0001, Jia Pan 0001 |
IROS | 7 |
| 2021 | Intermittent Contextual Learning for Keyfilter-Aware UAV Object Tracking Using Deep Convolutional FeatureabstractVisual tracking, one of the most favorable multimedia applications, has been widely used in unmanned aerial vehicle (UAV) for civil infrastructure monitoring, aerial cinematography, autonomous navigation, etc. Most existing trackers utilize deep convolutional feature to enhance tracking robustness in scenarios of various appearance variation. However, they commonly neglect speed which is crucial for UAV with restricted calculation resources. In this work, a novel correlation filter-based keyfilter-aware tracker with a new intermittent context learning strategy is proposed to efficiently and effectively alleviate the problems of background clutter, deficient description, occlusion, illumination change, etc. Specifically, context information is utilized to empower the filter higher discriminating ability through response repression of the omnidirectional context patches. Furthermore, keyfilter is produced from the periodically selected keyframe. The latest produced keyfilter is used to restrain the current filter's corrupted changes. Most importantly, context learning of correlation filter is implemented intermittently to fully increase the tracking efficiency. This intermittent learning strategy can ensure every filter maintain context awareness owing to the restriction of keyfilter, periodically enhancing the context awareness. Substantial experiments on three challenging UAV benchmarks totally with 213 image sequences have shown that our tracker surpasses the state-of-the-art results, and exhibits a remarkable generality in short-term and long-term UAV tracking tasks as well as a variety of challenging attributes. Yiming Li 0003, Changhong Fu 0001, Ziyuan Huang 0003, Yinqiang Zhang, Jia Pan 0001 |
IEEE Trans. Multim. | 5 |
| 2020 | Learning Resilient Behaviors for Navigation Under UncertaintyabstractDeep reinforcement learning has great potential to acquire complex, adaptive behaviors for autonomous agents automatically. However, the underlying neural network polices have not been widely deployed in real-world applications, especially in these safety-critical tasks (e.g., autonomous driving). One of the reasons is that the learned policy cannot perform flexible and resilient behaviors as traditional methods to adapt to diverse environments. In this paper, we consider the problem that a mobile robot learns adaptive and resilient behaviors for navigating in unseen uncertain environments while avoiding collisions. We present a novel approach for uncertainty-aware navigation by introducing an uncertainty-aware predictor to model the environmental uncertainty, and we propose a novel uncertainty-aware navigation network to learn resilient behaviors in the prior unknown environments. To train the proposed uncertainty-aware network more stably and efficiently, we present the temperature decay training paradigm, which balances exploration and exploitation during the training process. Our experimental evaluation demonstrates that our approach can learn resilient behaviors in diverse environments and generate adaptive trajectories according to environmental uncertainties. Tingxiang Fan, Pinxin Long, Wenxi Liu, Jia Pan 0001, Ruigang Yang, Dinesh Manocha |
ICRA | 4 |
| 2020 | Real-Time UAV Path Planning for Autonomous Urban Scene ReconstructionabstractUnmanned aerial vehicles (UAVs) are frequently used for large-scale scene mapping and reconstruction. However, in most cases, drones are operated manually, which should be more effective and intelligent. In this article, we present a method of real-time UAV path planning for autonomous urban scene reconstruction. Considering the obstacles and time costs, we utilize the top view to generate the initial path. Then we estimate the building heights and take close-up pictures that reveal building details through a SLAM framework. To predict the coverage of the scene, we propose a novel method which combines information on reconstructed point clouds and possible coverage areas. The experimental results reveal that the reconstruction quality of our method is good enough. Our method is also more time-saving than the state-of-the-arts. Qi Kuang, Jia Pan 0001 |
ICRA | 3 |
| 2020 | Keyfilter-Aware Real-Time UAV Object TrackingabstractCorrelation filter-based tracking has been widely applied in unmanned aerial vehicle (UAV) with high efficiency. However, it has two imperfections, i.e., boundary effect and filter corruption. Several methods enlarging the search area can mitigate boundary effect, yet introducing undesired background distraction. Existing frame-by-frame context learning strategies for repressing background distraction nevertheless lower the tracking speed. Inspired by keyframe-based simultaneous localization and mapping, keyfilter is proposed in visual tracking for the first time, in order to handle the above issues efficiently and effectively. Keyfilters generated by periodically selected keyframes learn the context intermittently and are used to restrain the learning of filters, so that 1) context awareness can be transmitted to all the filters via keyfilter restriction, and 2) filter corruption can be repressed. Compared to the state-of-the-art results, our tracker performs better on two challenging benchmarks, with enough speed for UAV real-time applications. Yiming Li 0003, Changhong Fu 0001, Ziyuan Huang 0003, Yinqiang Zhang, Jia Pan 0001 |
ICRA | 5 |
| 2020 | An Actor-Critic Approach for Legible Robot Motion PlannerabstractIn human-robot collaboration, it is crucial for the robot to make its intentions clear and predictable to the human partners. Inspired by the mutual learning and adaptation of human partners, we suggest an actor-critic approach for a legible robot motion planner. This approach includes two neural networks and a legibility evaluator: 1) A policy network based on deep reinforcement learning (DRL); 2) A Recurrent Neural Networks (RNNs) based sequence to sequence (Seq2Seq) model as a motion predictor; 3) A legibility evaluator that maps motion to legible reward. Through a series of human-subject experiments, we demonstrate that with a simple handicraft function and no real-human data, our method lead to improved collaborative performance against a baseline method and a non-prediction method. Xuan Zhao 0006, Tingxiang Fan, Dawei Wang 0006, Tao Han 0008, Jia Pan 0001 |
ICRA | 6 |
| 2020 | Augmented Memory for Correlation Filters in Real-Time UAV TrackingabstractThe outstanding computational efficiency of discriminative correlation filter (DCF) fades away with various complicated improvements. Previous appearances are also gradually forgotten due to the exponential decay of historical views in traditional appearance updating scheme of DCF framework, reducing the model's robustness. In this work, a novel tracker based on DCF framework is proposed to augment memory of previously appeared views while running at real-time speed. Several historical views and the current view are simultaneously introduced in training to allow the tracker to adapt to new appearances as well as memorize previous ones. A novel rapid compressed context learning is proposed to increase the discriminative ability of the filter efficiently. Substantial experiments on UAVDT and UAV123 datasets have validated that the proposed tracker performs competitively against other 26 top DCF and deep-based trackers with over 40 FPS on CPU. Yiming Li 0003, Changhong Fu 0001, Fangqiang Ding, Ziyuan Huang 0003, Jia Pan 0001 |
IROS | 5 |
| 2020 | Configuration Space Decomposition for Learning-based Collision Checking in High-DOF RobotsabstractMotion planning for robots of high degrees-of-freedom (DOFs) is an important problem in robotics with sampling-based methods in configuration space $\mathcal{C}$ as one popular solution. Recently, machine learning methods have been introduced into sampling-based motion planning methods, which train a classifier to distinguish collision free subspace from in-collision subspace in $\mathcal{C}$. In this paper, we propose a novel configuration space decomposition method and show two nice properties resulted from this decomposition. Using these two properties, we build a composite classifier that works compatibly with previous machine learning methods by using them as the elementary classifiers. Experimental results are presented, showing that our composite classifier outperforms state-of-the-art single-classifier methods by a large margin. A real application of motion planning in a multi-robot system in plant phenotyping using three UR5 robotic arms is also presented. Yiheng Han, Wang Zhao 0001, Jia Pan 0001, Yong-Jin Liu 0001 |
IROS | 3 |
| 2020 | DeepMNavigate: Deep Reinforced Multi-Robot Navigation Unifying Local & Global Collision AvoidanceabstractWe present a novel algorithm (DeepMNavigate) for global multi-agent navigation in dense scenarios using deep reinforcement learning (DRL). Our approach uses local and global information for each robot from motion information maps. We use a three-layer CNN that takes these maps as input to generate a suitable action to drive each robot to its goal position. Our approach is general, learns an optimal policy using a multi-scenario, multi-state training algorithm, and can directly handle raw sensor measurements for local observations. We demonstrate the performance on dense, complex benchmarks with narrow passages and environments with tens of agents. We highlight the algorithm’s benefits over prior learning methods and geometric decentralized algorithms in complex scenarios. Qingyang Tan, Tingxiang Fan, Jia Pan 0001, Dinesh Manocha |
IROS | 3 |
| 2020 | A Flexible Dual-Core Optical Waveguide Sensor for Simultaneous and Continuous Measurement of Contact Force and PositionabstractHaving the merits of chemical inertness and immunity to electromagnetic interference, light weight, small size, and softness, optical waveguides have attracted much attention in making tactile sensors recently. This paper presents a new design of waveguide using two layers of cores, one of which has an uniform width and the other has an incremental width. It is deduced and verified that the contact force can be derived from the light power loss in the uniform-width core, while the contact position can be derived from the light power loss in the other core together with the estimated force. By this dual-core design, a single waveguide can simultaneously and continuously measure the contact force and position along it, which makes it very suited for integration on some thin long robotic parts, such as robotic fingers. A hardware experiment has been conducted to demonstrate its effectiveness on a two-finger gripper in an assembly task. The dual-core waveguide achieves 2 mm spatial resolution and 0.1 N sensitivity. Zhong Zhang 0015, Yu Zheng 0001, Jia Pan 0001, Xiong Li 0001, Zhengyou Zhang |
IROS | 3 |
| 2019 | Context-Aware Spatio-Recurrent Curvilinear Structure SegmentationabstractCurvilinear structures are frequently observed in various images in different forms, such as blood vessels or neuronal boundaries in biomedical images. In this paper, we propose a novel curvilinear structure segmentation approach using context-aware spatio-recurrent networks. Instead of directly segmenting the whole image or densely segmenting fixed-sized local patches, our method recurrently samples patches with varied scales from the target image with learned policy and processes them locally, which is similar to the behavior of changing retinal fixations in the human visual system and it is beneficial for capturing the multi-scale or hierarchical modality of the complex curvilinear structures. In specific, the policy of choosing local patches is attentively learned based on the contextual information of the image and the historical sampling experience. In this way, with more patches sampled and refined, the segmentation of the whole image can be progressively improved. To validate our approach, comparison experiments on different types of image data are conducted and the sampling procedures for exemplar images are illustrated. We demonstrate that our method achieves the state-of-the-art performance in public datasets. Feigege Wang, Wenxi Liu, Yuanlong Yu 0001, Shengfeng He, Jia Pan 0001 |
CVPR | 6 |
| 2019 | Visualizing the Invisible: Occluded Vehicle Segmentation and RecoveryabstractIn this paper, we propose a novel iterative multi-task framework to complete the segmentation mask of an occluded vehicle and recover the appearance of its invisible parts. In particular, firstly, to improve the quality of the segmentation completion, we present two coupled discriminators that introduce an auxiliary 3D model pool for sampling authentic silhouettes as adversarial samples. In addition, we propose a two-path structure with a shared network to enhance the appearance recovery capability. By iteratively performing the segmentation completion and the appearance recovery, the results will be progressively refined. To evaluate our method, we present a dataset, Occluded Vehicle dataset, containing synthetic and real-world occluded vehicle images. Based on this dataset, we conduct comparison experiments and demonstrate that our model outperforms the state-of-the-arts in both tasks of recovering segmentation mask and appearance for occluded vehicles. Moreover, we also demonstrate that our appearance recovery approach can benefit the occluded vehicle tracking in real-world videos. Xiaosheng Yan, Yuanlong Yu 0001, Feigege Wang, Wenxi Liu, Shengfeng He, Jia Pan 0001 |
ICCV | 6 |
| 2019 | Compact Reachability Map for Excavator Motion PlanningabstractIn this paper, we propose a novel compact reachability map representation for excavator motion planning. The constructed reachability map can concisely encode the bucket’s reachable pose and the translation capability limited by excavator’s kinematic structure. By explicitly exploiting the property that the basic excavation motion lies on the excavation plane determined by excavator links, we further reduce the construction of the map from 3D Euclidean space to 2D excavation plane. We show the pre-computed reachability map can be used to develop new excavator motion planning approach. By indexing on the pre-computed reachability map, we can efficiently compute the feasible full-bucket trajectory for single step excavation operation. We highlight the results of the reachability map construction and demonstrate the simulation results of motion planning using a commercial dynamic simulator. Yajue Yang, Liangjun Zhang, Xinjing Cheng, Jia Pan 0001, Ruigang Yang |
IROS | 4 |
| 2018 | Manipulating Highly Deformable Materials Using a Visual Feedback DictionaryabstractThe complex physical properties of highly deformable materials such as clothes pose significant challenges for autonomous robotic manipulation systems. We present a novel visual feedback dictionary-based method for manipulating deformable objects towards a desired configuration. Our approach is based on visual servoing and we use an efficient technique to extract key features from the RGB sensor stream in the form of a histogram of deformable model features. These histogram features serve as high-level representations of the state of the deformable material. Next, we collect manipulation data and use a visual feedback dictionary that maps the velocity in the high-dimensional feature space to the velocity of the robotic end-effectors for manipulation. We have evaluated our approach on a set of complex manipulation tasks and human-robot manipulation tasks on different cloth pieces with varying material characteristics. Biao Jia, Jia Pan 0001, Dinesh Manocha |
ICRA | 3 |
| 2018 | Towards Optimally Decentralized Multi-Robot Collision Avoidance via Deep Reinforcement LearningabstractDeveloping a safe and efficient collision avoidance policy for multiple robots is challenging in the decentralized scenarios where each robot generates its paths without observing other robots' states and intents. While other distributed multi-robot collision avoidance systems exist, they often require extracting agent-level features to plan a local collision-free action, which can be computationally prohibitive and not robust. More importantly, in practice the performance of these methods are much lower than their centralized counterparts. We present a decentralized sensor-level collision avoidance policy for multi-robot systems, which directly maps raw sensor measurements to an agent's steering commands in terms of movement velocity. As a first step toward reducing the performance gap between decentralized and centralized methods, we present a multi-scenario multi-stage training framework to learn an optimal policy. The policy is trained over a large number of robots on rich, complex environments simultaneously using a policy gradient based reinforcement learning algorithm. We validate the learned sensor-level collision avoidance policy in a variety of simulated scenarios with thorough performance evaluations and show that the final learned policy is able to find time efficient, collision-free paths for a large-scale robot system. We also demonstrate that the learned policy can be well generalized to new scenarios that do not appear in the entire training period, including navigating a heterogeneous group of robots and a large-scale scenario with 100 robots. Videos are available at https://sites.google.com/view/drlmaca. Pinxin Long, Tingxiang Fan, Xinyi Liao, Wenxi Liu, Hao Zhang 0170, Jia Pan 0001 |
ICRA | 6 |
| 2018 | Considering Human Behavior in Motion Planning for Smooth Human-Robot Collaboration in Close ProximityabstractIt is well-known that a deep understanding of coworkers' behavior and preference is important for collaboration effectiveness. In this work, we present a method to accomplish smooth human-robot collaboration in close proximity by taking into account the human's behavior while planning the robot's trajectory. In particular, we first use an occupancy map to summarize human's movement preference over time, and such prior information is then considered in an optimization-based motion planner via two cost items: 1) avoidance of the workspace previously occupied by human, to eliminate the interruption and to increase the task success rate; 2) tendency to keep a safe distance between the human and the robot to improve the safety. In the experiments, we compare the collaboration performance among planners using different combinations of human-aware cost items, including the avoidance factor, both the avoidance and safe distance factor, and a baseline where no human-related factors are considered. The trajectories generated are tested in both simulated and real-world environments, and the results show that our method can significantly increase the collaborative task success rates and is also human-friendly. Xuan Zhao 0006, Jia Pan 0001 |
RO-MAN | 2 |
| 2017 | Efficient multi-agent global navigation using interpolating bridgesabstractWe present a novel approach for collision-free global navigation for continuous-time multi-agent systems with general linear dynamics. Our approach is general and can be used to perform collision-free navigation in 2D and 3D workspaces with narrow passages and crowded regions. As part of pre-computation, we compute multiple bridges in the narrow or tight regions in the workspace using kinodynamic RRT algorithms. Our bridge has certain geometric properties that enable us to calculate a collision-free trajectory for each agent using simple interpolation at runtime. Moreover, we combine interpolated bridge trajectories with local multi-agent navigation algorithms to compute global collision-free paths for each agent. The overall approach combines the performance benefits of coupled multi-agent algorithms with the precomputed trajectories of the bridges to handle challenging scenarios. In practice, our approach can perform global navigation for tens to hundreds of agents on a single CPU core in 2D and 3D workspaces. Liang He 0008, Jia Pan 0001, Dinesh Manocha |
ICRA | 2 |
| 2017 | Parallel Motion Planning Using Poisson-Disk SamplingabstractWe present a rapidly exploring-random-tree-based parallel motion planning algorithm that uses the maximal Poisson-disk sampling scheme. Our approach exploits the free-disk property of the maximal Poisson-disk samples to generate nodes and perform tree expansion. Furthermore, we use an adaptive scheme to generate more samples in challenging regions of the configuration space. The Poisson-disk sampling results in improved parallel performance and we highlight the performance benefits on multicore central processing units as well as manycore graphics processing units on different benchmarks. Chonhyon Park, Jia Pan 0001, Dinesh Manocha |
IEEE Trans. Robotics | 2 |
| 2016 | Analyzing the utility of a support pin in sequential robotic manipulationabstractPick-and-place regrasp is an important manipulation skill for a robot. It helps a robot accomplish tasks that cannot be achieved within a single grasp, due to constraints such as kinematics or collisions between the robot and the environment. Previous work on pick-and-place regrasp only leveraged flat surfaces for intermediate placements, and thus is limited in the capability to reorient an object. In this paper, we extend the reorientation capability of a pick-and-place regrasp by adding a vertical pin on the working surface and using it as the intermediate location for regrasping. In particular, our method automatically computes the stable placements of an object leaning against a vertical pin, finds several force-closure grasps, generates a graph of regrasp actions, and searches for the regrasp sequence. To compare the regrasping performance with and without using pins, we evaluate the success rate and the length of regrasp sequences while performing tasks on various models. Experiments on reorientation and assembly tasks validate the benefit of using support pins for regrasping. Weiwei Wan, Jia Pan 0001, Kensuke Harada |
ICRA | 3 |
| 2016 | Proxemic group behaviors using reciprocal multi-agent navigationabstractWe present a decentralized algorithm for group-based coherent and reciprocal multi-agent navigation. In addition to generating collision-free trajectories for each agent, our approach is able to simulate macroscopic group movements and proxemic behaviors that result in coherent navigation. Our approach is general, makes no assumptions about the size or shape of the group, and can generate smooth trajectories for the agents. Furthermore, it can dynamically adapt to obstacles or the behavior of other agents. The additional overhead of generating proxemic group behaviors is relatively small and our approach can simulate hundreds of agents in real-time. We highlight its benefits on different benchmarks. Liang He 0008, Jia Pan 0001, Wenping Wang 0001, Dinesh Manocha |
ICRA | 2 |
| 2016 | Rope caging and graspingabstractWe present a novel method for caging grasps in this paper by stretching ropes on the surface of a 3D object. Both topology and shape of a model to be grasped has been considered in our approach. Our algorithm can guarantee generating local minimal rings on every topological branches of a given model with the help of a Reeb graph. Cages and grasps can then be computed from these rings, and physical experimental tests have been conducted to verify the robustness of our approach. Tsz-Ho Kwok, Weiwei Wan, Jia Pan 0001, Charlie C. L. Wang, Jianjun Yuan 0003, Kensuke Harada, Yong Chen 0017 |
ICRA | 3 |
| 2016 | An empirical comparison among the effect of different supports in sequential robotic manipulationabstractPick-and-place regrasp extends the manipulation capability of a robot by using a sequence of regrasps to accomplish tasks that are not possible using a single grasp due to constraints such as kinematics or collisions between the robot and the environment. Previous work on pick-and-place only leveraged static passive devices for intermediate placements, and thus is limited in the flexibility and robustness to reorient an object. In this paper, we extend the reorientation capability of a pick-and-place regrasp by adding an actively actuated gripper fixed in the working cell, and using it as the intermediate location for regrasping. In particular, our method automatically computes the stable placements of an object being hold in the gripper support, finds a rich set of force-closure grasps, performs k-means based grasp clustering, generates a graph of regrasp actions, and searches for the optimal regrasp sequence. To compare the regrasping performance with typical passive supports, we evaluate the success rate while performing tasks on various models. Experiments on reorientation tasks validate the benefit of using an actively actuated gripper for regrasp placement. Weiwei Wan, Jia Pan 0001, Kensuke Harada |
IROS | 3 |
| 2016 | Efficient global penetration depth computation for articulated models
Hao Tian 0003, Xinyu Zhang 0002, Changbo Wang, Jia Pan 0001, Dinesh Manocha |
Comput. Aided Des. | 4 |
| 2015 | Multi-contour initial pose estimation for 3D registrationabstractReliable manipulation of everyday household objects is essential to the success of service robots. In order to accurately manipulate these objects, robots need to know objects' full 6-DOF pose, which is challenging due to sensor noise, clutters and occlusions. In this paper, we present a new approach for effectively guessing the object pose given an observation of just a small patch of the object, by leveraging the fact that many household objects can only keep stable on a planar surface under a small set of poses. In particular, for each stable pose of an object, we slice the object with horizontal planes and extract multiple cross-section contours. The pose estimation is then reduced to find a stable pose whose contour matches best with that of the sensor data, and this can be solved efficiently by convolution. Experiments on the manipulation tasks in the DARPA Robotics Challenge validate our approach. In addition, we also investigate our method's performance on object recognition tasks raising in the challenge. Ernest Cheung, Jia Pan 0001 |
IROS | 3 |
| 2015 | Leveraging appearance priors in non-rigid registration, with application to manipulation of deformable objectsabstractManipulation of deformable objects is a widely applicable but challenging task in robotics. One promising nonparametric approach for this problem is trajectory transfer, in which a non-rigid registration is computed between the starting scene of the demonstration and the scene at test time. This registration is extrapolated to find a function from ℝ3to ℝ3, which is then used to warp the demonstrated robot trajectory to generate a proposed trajectory to execute in the test scene. In prior work [1] [2], only depth information from the scenes has been used to compute this warp function. This approach ignores appearance information, but there are situations in which using both shape and appearance information is necessary for finding high quality non-rigid warp functions. In this paper, we describe an approach to learn relevant appearance information about deformable objects using deep learning, and use this additional information to improve the quality of non-rigid registration between demonstration and test scenes. Our method better registers areas of interest on deformable objects that are crucial for manipulation, such as rope crossings and towel corners and edges. We experimentally validate our approach in both simulation and in the real world, and show that the utilization of appearance information leads to a significant improvement in both selecting the best matching demonstration scene for a given test scene, and finding a high quality non-rigid registration between those two scenes. Sandy H. Huang, Jia Pan 0001, George Mulcaire, Pieter Abbeel |
IROS | 2 |
| 2015 | Planning Curvature and Torsion Constrained Ribbons in 3D With Application to Intracavitary BrachytherapyabstractWe present an approach for planning ensembles of channels, ribbons, within 3D printed implants for facilitating radiation therapy treatment of cancer. The ribbons are traced out by sweeping a constant width rigid body (cuboid) along spatial curves. We propose a method for planning multiple disjoint and mutually collision-free ribbons of finite thickness along curvature and torsion constrained curves in 3D space. This is equivalent to planning motions for the cross section of the ribbon along a spatial curve such that the cross section is oriented along the unit binormal to the curve defined according to the Frenet-Serret frame. We propose a two-stage planning approach. In the first stage, a customized sampling-based planner uses rapidly exploring random trees (RRTs) to generate feasible curvature and torsion constrained ribbons. In the second stage, the curvature and torsion along each ribbon is locally optimized using sequential quadratic programming (SQP). We use this approach to design curved radiation delivery channels inside custom 3D printed implants that allow temporary insertion of a high-dose radioactive source that is threaded through the channels using a wire and allowed to dwell for specified times to expose cancerous tumors for intracavitary brachytherapy treatment. Constraints on the curvature and torsion are required for 3D printing (to allow flushing of sacrificial material) and for smooth insertion of radioactive sources. In simulation experiments, this approach achieves an improvement of 46% in tumor coverage compared with a greedy approach that generates channels sequentially. Sachin Patil, Jia Pan 0001, Pieter Abbeel, Kenneth Y. Goldberg |
IEEE Trans Autom. Sci. Eng. | 2 |
| 2014 | Predicting initialization effectiveness for trajectory optimizationabstractTrajectory optimization is a method for solving motion planning problems by formulating them as non-convex constrained optimization problems. The optimization process, however, can get stuck in local optima that are in collision. As a consequence, these methods typically require multiple initializations. This poses the problem of deciding which initializations to use when given a limited computational budget. In this paper we propose a machine learning approach to predict whether a collision-free solution will be found from a given initialization. We present a set of trajectory features that encode the obstacle distribution locally around a robot. These features are designed for generalization across different tasks. Our experiments on various planning benchmarks demonstrate the performance of our approach. Jia Pan 0001, Pieter Abbeel |
ICRA | 1 |
| 2014 | Poisson-RRTabstractWe present an RRT-based motion planning algorithm that uses the maximal Poisson-disk sampling scheme. Our approach exploits the free-disk property of the maximal Poisson-disk samples to generate nodes and perform tree expansion. Furthermore, we use an adaptive scheme to generate more samples in challenging regions of the configuration space. Our approach can be easily parallelized on multi-core CPUs and many-core GPUs. We highlight the performance of our algorithm on different benchmarks. Chonhyon Park, Jia Pan 0001, Dinesh Manocha |
ICRA | 2 |
| 2014 | Motion planning under uncertainty for on-road autonomous drivingabstractWe present a motion planning framework for autonomous on-road driving considering both the uncertainty caused by an autonomous vehicle and other traffic participants. The future motion of traffic participants is predicted using a local planner, and the uncertainty along the predicted trajectory is computed based on Gaussian propagation. For the autonomous vehicle, the uncertainty from localization and control is estimated based on a Linear-Quadratic Gaussian (LQG) framework. Compared with other safety assessment methods, our framework allows the planner to avoid unsafe situations more efficiently, thanks to the direct uncertainty information feedback to the planner. We also demonstrate our planner's ability to generate safer trajectories compared to planning only with a LQG framework. Wenda Xu, Jia Pan 0001, Junqing Wei, John M. Dolan |
ICRA | 2 |
| 2014 | Planning Curvature and Torsion Constrained Ribbons in 3D with Application to Intracavitary Brachytherapy
Sachin Patil, Jia Pan 0001, Pieter Abbeel, Kenneth Y. Goldberg |
WAFR | 2 |
| 2013 | Real-time collision detection and distance computation on point cloud sensor dataabstractMost prior techniques for proximity computations are designed for synthetic models and assume exact geometric representations. However, real robots construct representations of the environment using their sensors, and the generated representations are more cluttered and less precise than synthetic models. Furthermore, this sensor data is updated at high frequency. In this paper, we present new collision- and distance-query algorithms, which can efficiently handle large amounts of point cloud sensor data received at real-time rates. We present two novel techniques to accelerate the computation of broad-phase data structures: 1) we present a progressive technique that incrementally computes a high-quality dynamic AABB tree for fast culling, and 2) we directly use an octree representation of the point cloud data as a proximity data structure. We assign a probability value to each leaf node of the tree, and the algorithm computes the nodes corresponding to high collision probability. In practice, our new approaches can be an order of magnitude faster than previous methods. We demonstrate the performance of the new methods on both synthetic data and on sensor data collected using a Kinect™ for motion planning for a mobile manipulator robot. Jia Pan 0001, Ioan Alexandru Sucan, Sachin Chitta, Dinesh Manocha |
ICRA | 1 |
| 2013 | Real-time optimization-based planning in dynamic environments using GPUsabstractWe present a novel algorithm to compute collision-free trajectories in dynamic environments. Our approach is general and does not require a priori knowledge about the obstacles or their motion. We use a replanning framework that interleaves optimization-based planning with execution. Furthermore, we describe a parallel formulation that exploits a high number of cores on commodity graphics processors (GPUs) to compute a high-quality path in a given time interval. We derive bounds on how parallelization can improve the responsiveness of the planner and the quality of the trajectory. Chonhyon Park, Jia Pan 0001, Dinesh Manocha |
ICRA | 2 |
| 2013 | Efficient penetration depth approximation using active learningabstractWe present a new method for efficiently approximating the global penetration depth between two rigid objects using machine learning techniques. Our approach consists of two phases: offline learning and performing run-time queries. In the learning phase, we precompute an approximation of the contact space of a pair of intersecting objects from a set of samples in the configuration space. We use active and incremental learning algorithms to accelerate the precomputation and improve the accuracy. During the run-time phase, our algorithm performs a nearest-neighbor query based on translational or rotational distance metrics. The run-time query has a small overhead and computes an approximation to global penetration depth in a few milliseconds. We use our algorithm for collision response computations in Box2D or Bullet game physics engines and complex 3D models and observe more than an order of magnitude improvement over prior PD computation techniques. Jia Pan 0001, Xinyu Zhang 0002, Dinesh Manocha |
ACM Trans. Graph. | 1 |
| 2012 | Bi-level Locality Sensitive Hashing for k-Nearest Neighbor ComputationabstractWe present a new Bi-level LSH algorithm to perform approximate k-nearest neighbor search in high dimensional spaces. Our formulation is based on a two-level scheme. In the first level, we use a RP-tree that divides the dataset into sub-groups with bounded aspect ratios and is used to distinguish well-separated clusters. During the second level, we compute a single LSH hash table for each sub-group along with a hierarchical structure based on space-filling curves. Given a query, we first determine the sub-group that it belongs to and perform k-nearest neighbor search within the suitable buckets in the LSH hash table corresponding to the sub-group. Our algorithm also maps well to current GPU architectures and can improve the quality of approximate KNN queries as compared to prior LSH-based algorithms. We highlight its performance on two large, high-dimensional image datasets. Given a runtime budget, Bi-level LSH can provide better accuracy in terms of recall or error ration. Moreover, our formulation reduces the variation in runtime cost or the quality of results. Jia Pan 0001, Dinesh Manocha |
ICDE | 1 |
| 2012 | FCL: A general purpose library for collision and proximity queriesabstractWe present a new collision and proximity library that integrates several techniques for fast and accurate collision checking and proximity computation. Our library is based on hierarchical representations and designed to perform multiple proximity queries on different model representations. The set of queries includes discrete collision detection, continuous collision detection, separation distance computation and penetration depth estimation. The input models may correspond to triangulated rigid or deformable models and articulated models. Moreover, FCL can perform probabilistic collision checking between noisy point clouds that are captured using cameras or LIDAR sensors. The main benefit of FCL lies in the fact that it provides a unified interface that can be used by various applications. Furthermore, its flexible architecture makes it easier to implement new algorithms within this framework. The runtime performance of the library is comparable to state of the art collision and proximity algorithms. We demonstrate its performance on synthetic datasets as well as motion planning and grasping computations performed using a two-armed mobile manipulation robot. Jia Pan 0001, Sachin Chitta, Dinesh Manocha |
ICRA | 1 |
| 2012 | Real-Time Optimization-Based Planning in Dynamic Environments Using GPUsabstractWe present a novel algorithm to compute collision-free trajectories in dynamic environments. Our approach is general and makes no assumption about the obstacles or their motion. We use a replanning framework that interleaves optimization-based planning with execution. Furthermore, we describe a parallel formulation that exploits high number of cores on commodity graphics processors (GPUs) to compute a high-quality path in a given time interval. Overall, we show that search in configuration spaces can be significantly accelerated by using GPU parallelism. Chonhyon Park, Jia Pan 0001, Dinesh Manocha |
SOCS | 2 |
| 2012 | Faster Sample-Based Motion Planning Using Instance-Based Learning
Jia Pan 0001, Sachin Chitta, Dinesh Manocha |
WAFR | 1 |
| 2011 | Fast GPU-based locality sensitive hashing for k-nearest neighbor computationabstractWe present an efficient GPU-based parallel LSH algorithm to perform approximate k-nearest neighbor computation in high-dimensional spaces. We use the Bi-level LSH algorithm, which can compute k-nearest neighbors with higher accuracy and is amenable to parallelization. During the first level, we use the parallel RP-tree algorithm to partition datasets into several groups so that items similar to each other are clustered together. The second level involves computing the Bi-Level LSH code for each item and constructing a hierarchical hash table. The hash table is based on parallel cuckoo hashing and Morton curves. In the query step, we use GPU-based work queues to accelerate short-list search, which is one of the main bottlenecks in LSH-based algorithms. We demonstrate the results on large image datasets with 200,000 images which are represented as 512 dimensional vectors. In practice, our GPU implementation can obtain more than 40X acceleration over a single-core CPU-based LSH implementation. Jia Pan 0001, Dinesh Manocha |
GIS | 1 |
| 2011 | Probabilistic Collision Detection Between Noisy Point Clouds Using Robust Classification
Jia Pan 0001, Sachin Chitta, Dinesh Manocha |
ISRR | 1 |
| 2010 | g-Planner: Real-time Motion Planning and Global Navigation using GPUsabstractWe present novel randomized algorithms for solving global motion planning problems that exploit the computational capabilities of many-core GPUs. Our approach uses thread and data parallelism to achieve high performance for all components of sample-based algorithms, including random sampling, nearest neighbor computation, local planning, collision queries and graph search. The approach can efficiently solve both the multi-query and single-query versions of the problem and obtain considerable speedups over prior CPU-based algorithms. We demonstrate the efficiency of our algorithms by applying them to a number of 6DOF planning benchmarks in 3D environments. Overall, this is the first algorithm that can perform real-time motion planning and global navigation using commodity hardware. Jia Pan 0001, Christian Lauterbach, Dinesh Manocha |
AAAI | 1 |
| 2010 | Retraction-based RRT planner for articulated modelsabstractWe present a new retraction algorithm for high DOF articulated models and use our algorithm to improve the performance of RRT planners in narrow passages. The retraction step is formulated as a constrained optimization problem and performs iterative refinement on the boundary of C-Obstacle space. We also combine the retraction algorithm with decomposition planners to handle very high DOF articulated models. The performance of our approach is analyzed using Voronoi diagrams and we show that our retraction algorithm provides a good approximation to the ideal RRT-extension in constrained environments. We have implemented our algorithm and tested its performance on robots with more than 40 DOFs in complex environments. In practice, we observe significant performance (2-80X) improvement over prior RRT planners on challenging scenarios with narrow passages. Jia Pan 0001, Liangjun Zhang, Dinesh Manocha |
ICRA | 1 |
| 2010 | Efficient nearest-neighbor computation for GPU-based motion planningabstractWe present a novel k-nearest neighbor search algorithm (KNNS) for proximity computation in motion planning algorithm that exploits the computational capabilities of many-core GPUs. Our approach uses locality sensitive hashing and cuckoo hashing to construct an efficient KNNS algorithm that has linear space and time complexity and exploits the multiple cores and data parallelism effectively. In practice, we see magnitude improvement in speed and scalability over prior GPU-based KNNS algorithm. On some benchmarks, our KNNS algorithm improves the performance of overall planner by 20-40 times for CPU-based planner and up to 2 times for GPU-based planner. Jia Pan 0001, Christian Lauterbach, Dinesh Manocha |
IROS | 1 |
| 2010 | GPU-Based Parallel Collision Detection for Real-Time Motion Planning
Jia Pan 0001, Dinesh Manocha |
WAFR | 1 |
| 2010 | A hybrid approach for simulating human motion in constrained environmentsabstractAbstract We present a new algorithm to generate plausible motions for high‐DOF human‐like articulated figures in constrained environments with multiple obstacles. Our approach is general and makes no assumptions about the articulated model or the environment. The algorithm combines hierarchical model decomposition with sample‐based planning to efficiently compute a collision‐free path in tight spaces. Furthermore, we use path perturbation and replanning techniques to satisfy the kinematic and dynamic constraints on the motion. In order to generate realistic human‐like motion, we present a new motion blending algorithm that refines the path computed by the planner with motion capture data to compute a smooth and plausible trajectory. We demonstrate the results of generating motion corresponding to placing or lifting object, walking, and bending for a 38‐DOF articulated model. Copyright © 2010 John Wiley & Sons, Ltd. Jia Pan 0001, Liangjun Zhang, Ming C. Lin, Dinesh Manocha |
Comput. Animat. Virtual Worlds | 1 |