EDBT 2026 Demo / reviewers in the wild / expert
Zhenglong Sun 0001
dblp:02/3837
· DBLP profile ↗
25ranked-venue papers
1as first author
18since 2021 · last 2025
0000-0002-8135-1659ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 19 · 1 first-author · 12 since 2021Systems, architecture and hardware · 15 · 1 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | An Open-Source Snake Hole-Digging Inspired Safety-Critical Insertion Planning and Replanning Framework for Continuum RobotsabstractContinuum robots, with their slender, flexible structures, are increasingly utilized for navigating confined spaces via follow-the-leader (FTL) motion to minimize trauma. However, traditional FTL algorithms fail to adjust the entire robot configuration, including the tail, and struggle with navigation in dynamic environments requiring real-time re-planning. We propose a novel FTL motion planning framework inspired by snake hole-digging, enabling optimal shape configuration during insertion. Our approach outperforms the baseline FTL by 75.36% in three environments (circular, confined circular, maze) with six tests. The framework integrates control barrier functions (CBF) and quadratic programming (QP) for real-time obstacle avoidance, with control Lyapunov functions (CLF) ensuring minimal deviation. The target reaching error decreases by 48.43% when using CLF-CBF-QP compared to CBF-QP in four dynamic circular environments with obstacles from different directions. This work, with open-source code, provides a robust solution for continuum robot FTL motion planning in dynamic environments. Guanglin Ji, Zhenglong Sun 0001 |
IROS | 2 |
| 2025 | Towards Fully Autonomous Robotic Ultrasound-guided Biopsy for Superficial OrgansabstractUltrasound-guided therapeutic procedures rely heavily on operator skill, leading to variability and high training costs. The shortage of trained ultra-sonographers further exacerbates the issue, increasing workloads and associated health risks. Robotic technology has the potential to effectively tackle these issues, yet there has been limited research on fully autonomous robotic ultrasound-guided biopsy systems based on the entire workflow. To address this challenge, this paper presents an autonomous robotic operative framework for superficial organ biopsy. The system integrates real-time slice-to-volume registration and navigation, along with a needle insertion mechanism, following operational protocols to autonomously perform the entire biopsy procedure. The feasibility, robustness, and generalizability of the system are validated through experimental studies. Xiaoqiang Ji 0001, Zhenglong Sun 0001 |
IROS | 5 |
| 2025 | High-dynamic Tactile Sensing for Tactile Servo Manipulation: Let Robots Swing a HammerabstractHigh-dynamic tactile sensing and tactile servo control present challenges in robustness and real-time performance. This paper proposes a closed-loop tactile servo control strategy for robotic nail hammering, by allowing controlled hammer slide within a rigid robotic 2-finger gripper. The proposed approach detects tactile information of continuous sliding and sliding-induced vibrations in real time and modulates gripping force. The control encourages rotational sliding to enhance impact and reduce recoil while restricting parallel slippage to maintain grip stability. To achieve real-time processing and effective sliding feature extraction, we employ Short-Time Fourier Transform (STFT) and a dual-stream Physics-Informed Machine Learning (PIML) model, processing tactile data at 1 kHz with an average latency of 1.04 ms. Experimental results show that, compared to conventional methods, controlling hammer slippage reduces arm joint recoil by 64.26% (223.30 N → 79.81 N) while increasing hammer impact force by 179.97% (28.06 N → 78.56 N). The method adapts to hammers with varying mass distributions, significantly improving impact resilience and manipulation performance in high-dynamic interactions. These advancements pave the way for more dexterous and robust robotic systems with embodied intelligence. Yingtian Xu, Zhenglong Sun 0001, Ziya Wang |
IROS | 2 |
| 2025 | Toward high-efficiency, low-resource, and explainable neuropeptide prediction with MSKDNPabstractNeuropeptides are essential signaling molecules produced in the nervous system that regulate diverse physiological processes and are closely implicated in the pathogenesis of neurodegenerative and neuropsychiatric disorders. Investigating neuropeptides contributes to a better understanding of their regulatory mechanisms and offers new insights into therapeutic strategies for related diseases. Therefore, accurate identification of neuropeptides is crucial for advancing biomedical research and drug development. Due to the high cost of experimental validation, various artificial intelligence methods have been developed for rapid neuropeptide identification. However, existing approaches often suffer from high computational resource consumption, slow processing speed, and poor deploy ability. Moreover, a user-friendly web server for practical application is still lacking. To this end, we propose MSKDNP, a neuropeptide prediction model based on a multi-stage knowledge distillation framework. With only 1.2% of the parameters, MSKDNP attains performance comparable to a fully fine-tuned protein language model while achieving state-of-the-art results in neuropeptide recognition. Moreover, MSKDNP provides favorable interpretability, facilitating biological understanding. A freely accessible web server is available at https://awi.cuhk.edu.cn/∼biosequence/MSKDNP/index.php. Peilin Xie, Jiahui Guan, Yulan Liu, Zhang Cheng, Xuxin He, Zhenglong Sun 0001, Tzong-Yi Lee, Lantian Yao, Ying-Chih Chiang |
Briefings Bioinform. | 9 |
| 2025 | Teaching Masked Autoencoder With Strong AugmentationsabstractMasked autoencoder (MAE) has been regarded as a capable self-supervised learner for various downstream tasks. Nevertheless, the model still lacks high-level discriminability, which results in poor linear probing performance. In view of the fact that strong augmentation plays an essential role in contrastive learning, can we capitalize on strong augmentation in MAE? The difficulty originates from the pixel uncertainty caused by strong augmentation that may affect the reconstruction, and thus, directly introducing strong augmentation into MAE often hurts the performance. In this article, we delve into the potential of strong augmented views to enhance MAE while maintaining MAE's advantages. To this end, we propose a simple yet effective masked Siamese autoencoder (MSA) model, which consists of a student branch and a teacher branch. The student branch derives MAE's advanced architecture, and the teacher branch treats the unmasked strong view as an exemplary teacher to impose high-level discrimination onto the student branch. We demonstrate that our MSA can improve the model's spatial perception capability and, therefore, globally favors interimage discrimination. Empirical evidence shows that the model pretrained by MSA provides superior performances across different downstream tasks. Notably, linear probing performance on frozen features extracted from MSA leads to 6.1% gains over MAE on ImageNet-1k. Fine-tuning (FT) the network on VQAv2 task finally achieves 67.4% accuracy, outperforming 1.6% of the supervised method DeiT and 1.2% of MAE. Codes and models are available at https://github.com/KimSoybean/MSA. Rui Zhu 0014, Yalong Bai, Ting Yao 0003, Jingen Liu, Zhenglong Sun 0001, Tao Mei 0001, Chang Wen Chen |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2025 | Dynamic Hysteresis Compensation for Tendon-Sheath Mechanism in Flexible Surgical Robots Without Distal PerceptionabstractThe accurate position transmission of tendon-sheath mechanisms (TSMs) is challenging but of significance to the flexible robot for minimally invasive surgery (MIS). The challenges are mainly attributed to 1) the tendon-elongation and its caused hysteresis that depend on the route configuration of the TSM and could result in misaligned position transmission, 2) realistic surgical scenarios requiring the TSM with arbitrary and even time-varying route configurations, and 3) the absence of distal sensory feedback due to strict spatial constraints. Existing works are always devoted to tackling 1) yet evade 2) and 3). Here a route-related tendon-elongation model is formulated to resolve 1), and in response to 2), a route-sensing optical fiber is used. Obeying 3), a feedforward hysteresis compensator is then developed to align the distal position of the tendon with the desired position. Our final contribution gives an application-oriented remedy for the foregoing methodologies. Applying our compensator on the challenging position transmission tasks subject to 2) and 3), the positional accuracy can be still maintained at around 97.50%; guided by the provided remedy, the surgical end-effector achieves sub-millimeter tip position accuracy. Extensive tests demonstrate that the pending concerns yet of great practical importance in existing related works are well resolved. Guanglin Ji, Minyi Sun, Yin Xiao, Huaiyuan Rao, Zhenglong Sun 0001 |
IEEE Trans. Robotics | 6 |
| 2025 | Design and Analysis of a Closed-Loop Emotion Regulation System Based on Multimodal Affective Computing and Emotional Markov ChainabstractIn our daily lives, emotions are extremely important. However, predicting and regulating emotion is still a critical problem to be solved in the research of the human-computer interaction (HCI). In this study, we explore this problem using a regulation method. A novel multimodal affective computing algorithm is proposed and implemented in an emotion regulation system. The selection of music stimuli and the modeling of dynamic emotions are done using emotional Markov chains. This regulation system can monitor the user’s emotion and play music, selected by the regulation policy until the user can maintain the desired emotion. Our system was verified by two experiments. In the first experiment, by predicting the participants’ affective states, we tested the precision of our multimodal affective computing system. In the second experiment, we tested the regulation algorithm embedded in a closed-loop regulation system by comparing it with playing music without feedback. The results suggest that participants can regulate and maintain the desired affective state by using the emotion regulation system. Xingchao Wang, Chen-Zhong Li, Zhenglong Sun 0001, Yangsheng Xu |
IEEE Trans. Syst. Man Cybern. Syst. | 3 |
| 2024 | SD-DiT: Unleashing the Power of Self-Supervised Discrimination in Diffusion Transformer*abstractDiffusion Transformer (DiT) has emerged as the new trend of generative diffusion models on image generation. In view of extremely slow convergence in typical DiT, recent breakthroughs have been driven by mask strategy that significantly improves the training efficiency of DiT with additional intra-image contextual learning. Despite this progress, mask strategy still suffers from two inherent limitations: (a) training-inference discrepancy and (b) fuzzy relations between mask reconstruction & generative diffusion process, resulting in sub-optimal training of DiT. In this work, we address these limitations by novelly unleashing the self-supervised discrimination knowledge to boost DiT training. Technically, we frame our DiT in a teacher-student manner. The teacher-student discriminative pairs are built on the diffusion noises along the same Probability Flow Ordinary Differential Equation (PF-ODE). Instead of applying mask reconstruction loss over both DiT encoder and decoder, we decouple DiT encoder and decoder to separately tackle discriminative and generative objectives. In particular, by encoding discriminative pairs with student and teacher DiT encoders, a new discriminative loss is designed to encourage the inter-image alignment in the selfsupervised embedding space. After that, student samples are fed into student DiT decoder to perform the typical generative diffusion task. Extensive experiments are conducted on ImageNet dataset, and our method achieves a competitive balance between training cost and generative capacity. Rui Zhu 0014, Yingwei Pan, Yehao Li, Ting Yao 0003, Zhenglong Sun 0001, Tao Mei 0001, Chang Wen Chen |
CVPR | 5 |
| 2024 | Prompt-Based Learning for Unpaired Image CaptioningabstractUnpaired Image Captioning (UIC) has been developed to learn image descriptions from unaligned vision-language sample pairs. Existing works usually tackle this task using adversarial learning and visual concept reward based on reinforcement learning. However, these existing works were only able to learn limited cross-domain information in vision and language domains, which restrains the captioning performance of UIC. Inspired by the success of Vision-Language Pre-Trained Models (VL-PTMs) in this research, we attempt to infer the cross-domain cue information about a given image from the large VL-PTMs for the UIC task. This research is also motivated by recent successes of prompt learning in many downstream multi-modal tasks, including image-text retrieval and vision question answering. In this work, a semantic prompt is introduced and aggregated with visual features for more accurate caption prediction under the adversarial learning framework. In addition, a metric prompt is designed to select high-quality pseudo image-caption samples obtained from the basic captioning model and refine the model in an iterative manner. Extensive experiments on the COCO and Flickr30 K datasets validate the promising captioning ability of the proposed model. We expect that the proposed prompt-based UIC model will stimulate a new line of research for the VL-PTMs based captioning. Peipei Zhu, Xiao Wang 0014, Lin Zhu 0012, Zhenglong Sun 0001, Wei-Shi Zheng 0001, Yaowei Wang 0001, Chang Wen Chen |
IEEE Trans. Multim. | 4 |
| 2023 | Unpaired Image Captioning by Image-Level Weakly-Supervised Visual Concept RecognitionabstractThe goal of unpaired image captioning (UIC) is to describe images without using image-caption pairs in the training phase. Although challenging, we expect the task can be accomplished by leveraging images aligned with visual concepts. Most existing studies use off-the-shelf algorithms to obtain the visual concepts because the Bounding Box (BBox) labels or relationship-triplet labels used for training are expensive to acquire. To avoid exhaustive annotations, we propose a novel approach to achieve cost-effective UIC. Specifically, we adopt image-level labels to optimize the UIC model in a weakly-supervised manner. For each image, we assume that only the image-level labels are available without specific locations and numbers. The image-level labels are utilized to train a weakly-supervised object recognition model to extract object information (e.g., instance), and the extracted instances are adopted to infer the relationships among different objects using an enhanced graph neural network (GNN). The proposed approach achieves comparable or even better performance compared with previous methods without expensive annotations. Furthermore, we design an unrecognized object (UnO) loss to improve the alignment of the inferred object and relationship information with the images. It can effectively alleviate the issue encountered by existing UIC models when generating sentences with nonexistent objects. To the best of our knowledge, this is the first attempt to address the problem of Weakly-Supervised visual concept recognition for UIC (WS-UIC) based only on image-level labels. Extensive experiments demonstrate that the proposed method achieves inspiring results on the COCO dataset while significantly reducing the labeling cost. Peipei Zhu, Xiao Wang 0014, Yong Luo 0002, Zhenglong Sun 0001, Wei-Shi Zheng 0001, Yaowei Wang 0001, Chang Wen Chen |
IEEE Trans. Multim. | 4 |
| 2022 | Augmented Pointing Gesture Estimation for Human-Robot InteractionabstractWith recent advancements in CV (computer vision) and AI (Artificial Intelligence) technologies, pointing gesture is becoming an emerging trend for human-robot interaction. Its intuitive and deictic nature makes it an ideal way for giving commands, especially referring spatial information to the robots. In this paper, we propose an augmented pointing gesture estimation method to enable richer and programmable instructions to be given to the robots. We propose five pointing gestures and demonstrate the idea using a collaborative robot with a multi-finger robotic gripper. Experiments are designed and conducted to test the pointing accuracy in space and in gesture estimation. The results show that our proposed method can achieve a mean drift of 8.3 cm and an estimation accuracy of 94.08%. Zhixian Hu, Yingtian Xu, Waner Lin, Ziya Wang, Zhenglong Sun 0001 |
ICRA | 5 |
| 2022 | Fixed and Sliding FBG Sensors-Based Triaxial Tip Force Sensing for Cable-Driven Continuum RobotsabstractTip force sensing for cable-driven continuum robots are vital to provide the force information for safe and reliable human-robot interaction. However, traditional triaxial force sensors usually have a complicated structure occupying its inner lumen, without enough space for additional instrumental tools. To solve this, this paper proposes a fixed and sliding fiber Bragg grating (FBG) sensors-based triaxial force sensing method for cable-driven continuum robots. The fixed FBG sensors are attached to the circumferential surface of continuum robot at the tip and base, and the sliding optical fibers with FBG sensors are located in the actuation channels as the sensing integrated pulling cables. This configuration guarantees a compact structure and large inner lumen. Two five-degreed-of-freedom (5-DOF) electromagnetic (EM) and a 6-DOF EM sensors are assembled to the tip and the base of the robot respectively, which can obtain the pose of the tip with respect to the base. The tip force in three directions can be decoupled using the information of the Bragg wavelength changes and EM sensors. Results show that the mean errors of force sensing along x-direction, y-direction, and z-direction are 4.1%, 4.7%, and 9.8%, respectively. The proposed sensing method does not rely on the elasticity of continuum robot, enabling its wide applicability for other cable-driven pseudo-continuum robots. Zecai Lin, Huanghua Liu, Xiaojie Ai, Weidong Chen 0001, Anzhu Gao, Zhenglong Sun 0001, Guang-Zhong Yang, Huan Jia |
ICRA | 6 |
| 2022 | Recurrent neural networks as kinematics estimator and controller for redundant manipulators subject to physical constraints
Ning Tan 0003, Peng Yu 0003, Shen Liao, Zhenglong Sun 0001 |
Neural Networks | 4 |
| 2021 | Improving Contrastive Learning by Visualizing Feature TransformationabstractContrastive learning, which aims at minimizing the distance between positive pairs while maximizing that of negative ones, has been widely and successfully applied in unsupervised feature learning, where the design of positive and negative (pos/neg) pairs is one of its keys. In this paper, we attempt to devise a feature-level data manipulation, differing from data augmentation, to enhance the generic contrastive self-supervised learning. To this end, we first design a visualization scheme for pos/neg score1distribution, which enables us to analyze, interpret and understand the learning process. To our knowledge, this is the first attempt of its kind. More importantly, leveraging this tool, we gain some significant observations, which inspire our novel Feature Transformation proposals including the extrapolation of positives. This operation creates harder positives to boost the learning because hard positives enable the model to be more view-invariant. Besides, we propose the interpolation among negatives, which provides diversified negatives and makes the model more discriminative. It is the first attempt to deal with both challenges simultaneously. Experiment results show that our proposed Feature Transformation can improve at least 6.0% accuracy on ImageNet-100 over MoCo baseline, and about 2.0% accuracy on ImageNet-1K over the MoCoV2 baseline. Transferring to the downstream tasks successfully demonstrate our model is less task-bias. Visualization tools and codes: https://github.com/DTennant/CL-Visualizing-Feature-Transformation. Rui Zhu 0014, Bingchen Zhao, Jingen Liu, Zhenglong Sun 0001, Chang Wen Chen |
ICCV | 4 |
| 2021 | Long-Range Hand Gesture Recognition via Attention-based SSD NetworkabstractHand gesture recognition plays an essential role in the human-robot interaction (HRI) field. Most previous research only studies hand gesture recognition in a short distance, which cannot be applied for interaction with mobile robots like unmanned aerial vehicles (UAVs) at a longer and safer distance. Therefore, we investigate the challenging long-range hand gesture recognition problem for the interaction between humans and UAVs. To this end, we propose a novel attention-based single shot multibox detector (SSD) model that incorporates both spatial and channel attention for hand gesture recognition. We notably extend the recognition distance from 1 meter to 7 meters through the proposed model without sacrificing speed. Besides, we present a long-range hand gesture (LRHG) dataset collected by the USB camera mounted on mobile robots. The hand gestures are collected at discrete distance levels from 1 meter to 7 meters, where most of the hand gestures are small and at low resolution. Experiments with the self-built LRHG dataset show our methods reach the surprising performance-boosting over the state-of-the-art method like the SSD network on both short-range (1 meter) and long-range (up to 7 meters) hand gesture recognition tasks. Liguang Zhou, Chenping Du, Zhenglong Sun 0001, Tin Lun Lam, Yangsheng Xu |
ICRA | 3 |
| 2021 | Design of an SSVEP-based BCI Stimuli System for Attention-based Robot Navigation in Robotic TelepresenceabstractBrain-computer interface (BCI)-based robotic telepresence provides an opportunity for people with disabilities to control robots remotely without any actual physical movement. However, traditional BCI systems usually require the user to select the navigation direction from visual stimuli in a fixed background, which makes it difficult to control the robot in a dynamic environment during the locomotion. In this paper, a novel SSVEP-based BCI stimuli system is proposed for robotic telepresence. The novel system utilized the live video streamed from the robot onboard camera as the input. By altering and flickering the detected objects in the scene with different frequencies predefined based on their relative positions on the screen, the robot can be navigated based on the user’s attention in a dynamic manner. In order to better differentiate multiple objects (more than the number of frequencies predefined), the task-related component analysis (TRCA) model was trained with a priori offline experimental data to select the front objects with priority. Experiments were conducted to validate the proposed system. Using the system, four human subjects are able to control a humanoid robot to navigate through multiple objects to reach the desired goal. The success rate reaches 87.5% in average. Xingchao Wang, Xiaopeng Huang, Liguang Zhou, Zhenglong Sun 0001, Yangsheng Xu |
IROS | 5 |
| 2021 | BORM: Bayesian Object Relation Model for Indoor Scene RecognitionabstractScene recognition is a fundamental task in robotic perception. For human beings, scene recognition is reasonable because they have abundant object knowledge of the real world. The idea of transferring prior object knowledge from humans to scene recognition is significant but still less exploited. In this paper, we propose to utilize meaningful object representations for indoor scene representation. First, we utilize an improved object model (IOM) as a baseline that enriches the object knowledge by introducing a scene parsing algorithm pretrained on the ADE20K dataset with rich object categories related to the indoor scene. To analyze the object co-occurrences and pairwise object relations, we formulate the IOM from a Bayesian perspective as the Bayesian object relation model (BORM). Meanwhile, we incorporate the proposed BORM with the PlacesCNN model as the combined Bayesian object relation model (CBORM) for scene recognition and significantly outperforms the state-of-the-art methods on the reduced Places365 dataset, and SUN RGB-D dataset without retraining, showing the excellent generalization ability of the proposed method. Code can be found at https://github.com/FreeformRobotics/BORM. Liguang Zhou, Jun Cen, Xingchao Wang, Zhenglong Sun 0001, Tin Lun Lam, Yangsheng Xu |
IROS | 4 |
| 2021 | Trajectory Tracking of Soft Continuum Robots with Unknown Models Based on Varying Parameter Recurrent Neural NetworksabstractBio-inspired robots, e.g., soft continuum robots, have broad application prospects due to their structural dexterity and interaction safety. But these features also bring great challenges to the precise control of soft continuum robots. In this work, we investigate how to achieve the kinematic control of soft continuum robots without knowing model parameters of the robots. To this end, a model-free scheme based on varying-parameter recurrent neural networks (VP-RNN) is proposed. The scheme involves two components, one of which solves the inverse kinematics problem based on a VP-RNN model, and the other employs another VP-RNN model to estimate the pseudo-inverse of Jacobian matrix of continuum robots. Finally, the feasibility and robustness of the proposed control strategy are validated by simulations, including comparisons with other methods and case study with jammed actuation. Ning Tan 0003, Peng Yu 0003, Fenglei Ni, Zhenglong Sun 0001 |
SMC | 4 |
| 2020 | CCRobot-III: a Split-type Wire-driven Cable Climbing Robot for Cable-stayed Bridge Inspection*abstractThis paper presents a novel Cable Climbing Robot CCRobot-III, which is the third version designed for bridge cable inspection tasks, aiming at surpassing previous versions in terms of climbing speed and payload capacity. Benefiting from Split-type Wire-driven design, CCRobot-III can climb along a 90-110mm diameter bridge cable in inchworm-like gait at a speed of up to 12m/min, and carrying more than 40kg payload at the same time. CCRobot-III consists of a climbing precursor and a main-body frame. The two parts are connected and driven by steel wires. The climbing precursor, acting as a mobile anchor, moves quickly on a bridge cable. The mainbody frame, acting as a mobile winch, carries payload and pulls itself to a certain position with steel wires. Both parts have one or two pairs of palm-based gripper, which is the key component for providing strong adhesion to support the robot climbing. Experimental results have shown that CCRobotIII possesses outstanding climbing performance, high payload capacity, and good adaptability to complex conditions of cable surface. Moreover, it has potential engineering applications on the cable-stayed bridge for fieldwork. Ning Ding 0003, Zhenliang Zheng, Junlin Song, Zhenglong Sun 0001, Tin Lun Lam, Huihuan Qian |
ICRA | 4 |
| 2020 | A Novel Solar Tracker Driven by Waves: From Idea to ImplementationabstractTraditional solar trackers often adopt motors to automatically adjust the attitude of the solar panels towards the sun for maximum power efficiency. In this paper, a novel design of solar tracker for the ocean environment is introduced. Utilizing the fluctuations due to the waves, electromagnetic brakes are utilized instead of motors to adjust the attitude of the solar panels. Compared with the traditional solar trackers, the proposed one is simpler in hardware while the harvesting efficiency is similar. The desired attitude is calculated out of the local location and time. Then based on the dynamic model of the system, the angular acceleration of the solar panels is estimated and a control algorithm is proposed to decide the release and lock states of the brakes. In such a manner, the adjustment of the attitude of the solar panels can be achieved by using two brakes only. Experiments are conducted to validate the acceleration estimator and the dynamic model. At last, the feasibility of the proposed solar tracker is tested on the real water surface. The results show that the system is able to adjust 40° in two dimensions within 28 seconds. Hengli Liu, Chongfeng Liu, Zhenglong Sun 0001, Tin Lun Lam, Huihuan Qian |
ICRA | 4 |
| 2020 | DiPE: Deeper into Photometric Errors for Unsupervised Learning of Depth and Ego-motion from Monocular VideosabstractUnsupervised learning of depth and ego-motion from unlabelled monocular videos has recently drawn great attention, which avoids the use of expensive ground truth in the supervised one. It achieves this by using the photometric errors between the target view and the synthesized views from its adjacent source views as the loss. Despite significant progress, the learning still suffers from occlusion and scene dynamics. This paper shows that carefully manipulating photometric errors can tackle these difficulties better. The primary improvement is achieved by a statistical technique that can mask out the invisible or nonstationary pixels in the photometric error map and thus prevents misleading the networks. With this outlier masking approach, the depth of objects moving in the opposite direction to the camera can be estimated more accurately. To the best of our knowledge, such scenarios have not been seriously considered in the previous works, even though they pose a higher risk in applications like autonomous driving. We also propose an efficient weighted multi-scale scheme to reduce the artifacts in the predicted depth maps. Extensive experiments on the KITTI dataset show the effectiveness of the proposed approaches. The overall system achieves state-of-the-art performance on both depth and ego-motion estimation. Hualie Jiang, Laiyan Ding, Zhenglong Sun 0001, Rui Huang 0001 |
IROS | 3 |
| 2020 | OceanVoy: A Hybrid Energy Planning System for Autonomous SailboatabstractTowards long range and high endurance sailing, energy is of utmost importance. Moreover, benefiting from the dominance of the sailboat itself, it is energy-saving and environment-friendly. Thus, the sailboat with energy planning problem is meaningful. However, until now, the sailboat energy optimization problem has rarely been considered. In this paper, we focus on the energy consumption optimization of an autonomous sailboat. It has been formulated as a Nonlinear Programming problem (NLP). We deal with it with a hybrid control scheme, in which pseudo-spectral (PS) optimal control method is used in heading control, and a model-free framework guided by Extreme Seeking Control (ESC) is used in sail control. The optimal path is generated with the optimal input motor torques in time series. As a result, both simulation and experiments have validated motion planning and energy planning performance. Notably, about 7% of energy is saved on average. Our proposed method can make sailboats sailing longer and sustainable. Qinbo Sun, Weimin Qi, Hengli Liu, Zhenglong Sun 0001, Tin Lun Lam, Huihuan Qian |
IROS | 4 |
| 2020 | A Two-stage Automatic Latching System for The USVs Charging in Disturbed BerthabstractAutomatic latching for charging in a disturbed environment for Unmanned Surface Vehicle (USVs) is always a challenging problem. In this paper, we propose a two-stage automatic latching system for USVs charging in berth. In Stage I, a vision-guided algorithm is developed to calculate an optimal latching position for charging. In Stage II, a novel latching mechanism is designed to compensate the movement misalignments from the water disturbance. A set of experiments have been conducted in real-world environments. The results show the latching success rate has been improved from 40% to 73.3% in the best cases with our proposed system. Furthermore, the vision-guided algorithm provides a methodology to optimize the design radius of the latching mechanism with respect to different disturbance levels accordingly. Outdoor experiments have validated the efficiency of our proposed automatic latching system. The proposed system improves the autonomy intelligence of the USVs and provides great benefits for practical applications. Chongfeng Liu, Hengli Liu, Zhenglong Sun 0001, Tin Lun Lam, Huihuan Qian |
IROS | 5 |
| 2019 | Joint Torque Estimation toward Dynamic and Compliant Control for Gear-Driven Torque Sensorless Quadruped RobotabstractThis paper investigates dynamic and compliant control based on joint output torque estimation for electrically actuated quadruped robots with large-reduction-ratio harmonic gear. Compared with position control, force control exhibits better performance of dynamics and compliance for the robot's interactions with complex environments. However, force control without direct feedbacks from torque sensors may come with poor tracking performance of joint compliance when the robot equipped with gears of high reduction. To solve this problem, we propose a new method to estimate joint torque from motor current and rotation velocity detected on each joint, using a more precise friction model of the harmonic gear. We also introduce a pre-stance phase to the whole cycle of leg alternating swing/stance based on hybrid force and position control to dynamically absorb feet impacts on the ground. Our controller performance is validated by standing experiment and walking experiment. Bingchen Jin, Caiming Sun, Aidong Zhang 0002, Ning Ding 0003, Ganyu Deng, Zuwen Zhu, Zhenglong Sun 0001 |
IROS | 8 |
| 2009 | Development and preliminary data of novel integrated optical micro-force sensing tools for retinal microsurgeryabstractThis paper reports the development of novel micro-force sensing tools for retinal microsurgery. Retinal microsurgery requires extremely delicate manipulation of retinal tissue, and tool-to-tissue interaction forces are frequently below human perceptual thresholds. Further, the interaction between the tool shaft and sclera makes accurate sensing of forces exerted on the retina very difficult with previously developed force sensing schemes, in which the sensor is located outside the eye. In the work reported here, we incorporate 160 µm Fiber Bragg Grating (FBG) strain sensors into the tool shaft to sense forces distal to the sclera. The sensor is applicable both with robotically manipulated and freehand tools. Preliminary results with a 1 degree-of-freedom (DOF) sensor have demonstrated 0.25 mN resolution, and work is underway to develop 2 and 3 DOF tools. The design and analysis of the force sensing tool is presented with preliminary testing data and some initial experiments using the tool with both freehand and robotic manipulation. Zhenglong Sun 0001, Marcin Balicki, Jin U. Kang, James Handa, Russell H. Taylor, Iulian Iordachita |
ICRA | 1 |