EDBT 2026 Demo / reviewers in the wild / expert
Kai Huang 0001
dblp:86/489-1
· DBLP profile ↗
162ranked-venue papers
11as first author
95since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 81 · 9 first-author · 36 since 2021Artificial intelligence and machine learning · 72 · 1 first-author · 49 since 2021Applied, interdisciplinary, general and emerging computing · 26 · 1 first-author · 20 since 2021Graphics, computer vision, multimedia, augmented reality and games · 21 · 20 since 2021Software engineering, systems software and programming languages · 10 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 4 · 3 since 2021Computer networks · 1Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | An Environment-Aware Verification Framework for LLM-Generated Robot Control ProgramsabstractLarge language models (LLMs) are increasingly used in robotics to translate natural language instructions into executable control programs via task-specific prompts. However, existing approaches often lack correctness guarantees for LLM-generated programs, leading to compilation errors and runtime failures. While some methods consider a verification mechanism, they typically assume complete prior knowledge of the environment, making them unsuitable for complex environments where such knowledge is unavailable. This paper introduces VeBot, an environment-aware verification framework designed to ensure the correctness of robot control programs generated by LLMs. Specifically, VeBot introduces: (i) an LLM-friendly robot control language (RCL) that facilitates the program generation by abstracting away the complex Python code details, (ii) a compiler that translates LLM-generated RCL programs into a control flow graph (CFG) while verifying the lexical, syntactic, and semantic correctness, and (iii) a runtime verification mechanism that checks the CFG and compiles the verified segments into executable Python code, avoiding collisions or planning failures during execution. We illustrate the VeBot framework using a household scenario, and the evaluation shows that it consistently outperforms existing methods across a range of LLMs and tasks, achieving high success rates even with lightweight LLMs. ZhanShang Nie, Xuanming Liu, Kai Huang 0001, Shuai Zhao 0004 |
DATE | 7 |
| 2026 | Respiratory Motion Compensation Based on Mid-axis Plane for Dynamic Human Point Cloud Inpainting
Shengtao Li, Jiadun Wang, Daosong Hu, Kai Huang 0001 |
ICIC (17) | 4 |
| 2026 | A Passive-LoRa Tag Chip Achieving 78m Battery-Free Bidirectional Communication with Standard LoRa Devices
Qijing Xiao, Weixiao Wang, Guanjie Gu, Changgui Yang, Hanli Liu, Kai Huang 0001, Bo Zhao 0003 |
ISCAS | 7 |
| 2026 | A Photovoltaic Energy-Harvesting Chip Featuring Self-Adaptive-Monitoring-Time MPPT
Chenyang Tao, Changgui Yang, Kai Huang 0001, Bo Zhao 0003 |
ISCAS | 5 |
| 2026 | K-STAR: Knowledge-Guided Submap Tracking and Recovery for Monocular SLAM in Degraded Visual Conditions
Boyang Li 0009, Shuai Zhao 0004, Kai Huang 0001 |
KSEM (6) | 5 |
| 2026 | A frequency-guided denoising framework based on convolutional transformer for electrocardiogram signals
Mingyue Cui, Yewei Gan, Jiepeng Chen, Yanchong Xie, Daosong Hu, Yuning Cui 0001, Kai Huang 0001 |
Eng. Appl. Artif. Intell. | 8 |
| 2026 | From Edge to Edge: A Flow-Inspired Scheduling Planner for Multi-Robot SystemsabstractTrajectory planning is crucial in multi-robot systems, particularly in environments with numerous obstacles. While extensive research has been conducted in this field, the challenge of coordinating multiple robots to flow collectively from one side of the map to the other—such as in crossing missions through obstacle-rich spaces—has received limited attention. This paper focuses on this directional traversal scenario by introducing a real-time scheduling scheme that enables multi-robot systems to move from edge to edge, emulating the smooth and efficient flow of water. Inspired by network flow optimization, our scheme decomposes the environment into a flow-based network structure, enabling the efficient allocation of robots to paths based on real-time congestion levels. The proposed scheduling planner operates on top of existing collision avoidance algorithms, aiming to minimize overall traversal time by balancing detours and waiting times. Simulation results demonstrate the effectiveness of the proposed scheme in achieving fast and coordinated traversal. Furthermore, real-world flight tests with ten drones validate its practical feasibility. This work contributes a flow-inspired, real-time scheduling planner tailored for directional multi-robot traversal in complex, obstacle-rich environments. Code: https://github.com/chengji253/FlowPlanner. Mingyue Cui, Boyang Li 0009, Tianjiang Hu, Kai Huang 0001 |
IEEE Trans Autom. Sci. Eng. | 6 |
| 2026 | A Self-Attention-Based LiDAR Point Cloud Compression Framework in Autonomous Driving EnvironmentsabstractLight detection and ranging (LiDAR) sensors are crucial for autonomous vehicles to accurately perceive the surrounding environment. However, the sparsity and irregularity of large-scale LiDAR point clouds (LPCs) bring challenges for storage and transmission. Meanwhile, existing works usually adopt insufficient context and bring intolerable computation complexity, especially for high-precision LPC reconstruction. To address these problems, we propose a novel self-attention-based framework for LPC compression and reconstruction in autonomous driving environments. Specifically, our approach employs a robust backbone for octree-based feature extraction, which can be pretrained and easily extended to various tasks, thereby reducing the need for extensive task-specific architectural modifications. The backbone constructs node sequences of octree by nonoverlapping context windows and shares the result of a multihead self-attention (MSA) operation among them. Considering the similarity in features among sibling nodes, we design a locally enhanced module for exploiting sibling features and a positional encoding generator for enhancing the translation invariance of the octree node sequence. During postprocessing, we further propose an offset prediction model to reduce coordinate distortions caused by voxelization. Experimental results indicate that compared to the benchmark geometry-based point cloud compression (GPCC), our approach achieves gains of up to 54.4% for geometry and 6.8% for intensity, while compared to the attention-based baseline, we achieve up to 99% reduction in coding time. We believe that our approach effectively mines the spatial geometric features in LPCs and has low coupling for specific tasks, which will boost the related applications from algorithm optimization to industrial products. Mingyue Cui, Junhua Long, Mingjian Feng, Juncheng Tao, Yuyang Zhong, Yehua Ling, Daosong Hu, Kai Huang 0001 |
IEEE Trans. Ind. Informatics | 8 |
| 2026 | LNet: Lightweight Network for Driver Attention Estimation via Scene and Gaze ConsistencyabstractIn resource-constrained vehicle systems, establishing consistency between multi-view scenes and driver gaze remains challenging. Prior methods mainly focus on cross-source data fusion, estimating gaze or attention maps through unidirectional implicit links between scene and facial features. Although bidirectional projection can correct misalignment between predictions and ground truth, the high resolution of scene images and complex semantic extraction incur heavy computational loads. To address these issues, we propose a lightweight driver-attention estimation framework that leverages geometric consistency between scene and gaze to guide feature extraction bidirectionally, thereby strengthening representation. Specifically, we first introduce a lightweight feature extraction module that captures global and local information in parallel through dual asymmetric branches to efficiently extract facial and scene features. An information cross fusion module is then designed to promote interaction between the scene and gaze streams. The multi-branch architecture extracts gaze and geometric cues at multiple scales, reducing the computational redundancy caused by mixed features when modeling geometric consistency across both views. Experiments on a large public dataset show that incorporating scene information introduces no significant computational overhead and yields a better trade-off between accuracy and efficiency. Moreover, leveraging bidirectional projection and the temporal continuity of gaze, we preliminarily explore the framework's potential for predicting attention trends. Daosong Hu, Mingyue Cui, Kai Huang 0001 |
IEEE Trans. Image Process. | 4 |
| 2026 | Pushing Physical Limits and Uncovering Motion Templates of Spine-Based Quadruped Locomotion via Reinforcement LearningabstractFlexible spines are critical to the remarkable agility and speed of animals. Translating this biological advantage to quadruped robots presents a significant control challenge, particularly in coordinating the spine and limbs for maximal velocity. In this work, we utilize reinforcement learning (RL) to develop high-speed locomotion for a bioinspired mouse robot with a lateral flexible spine. The resulting controller achieves motor performance that demonstrably surpasses non-spined and model-based methods. More importantly, our analysis reveals the principles behind this performance: the emergence of two distinct motion templates. For high-speed walking, the robot learns a “whip-like” spinal oscillation to increase leg swing frequency, while for agile turning, it adopts a dynamic “bend-and-straighten” pattern. These findings demonstrate the capability of RL to not only generate high-performance controllers but also to produce emergent strategies that, upon analysis, reveal underlying principles of high-speed, spine-driven locomotion. Zhenshan Bing, Yulong Xiao, Yuhong Huang, Long Cheng 0007, Biao Hu 0001, Gang Chen 0023, Yang Gao 0001, Fuchun Sun 0001, Kai Huang 0001, Alois C. Knoll |
IEEE Trans. Robotics | 10 |
| 2025 | CLAPS: A CLIP-Unified Auto-Prompt Segmentation for Multi-Modal Retinal ImagingabstractRecent advancements in foundation models, such as the Segment Anything Model (SAM), have significantly impacted medical image segmentation, especially in retinal imaging, where precise segmentation is vital for diagnosis. Despite this progress, current methods face critical challenges: 1) modality ambiguity in textual disease descriptions, 2) a continued reliance on manual prompting for SAM-based workflows, and 3) a lack of a unified framework, with most methods being modalityand task-specific. To overcome these hurdles, we propose CLIP-unified Auto-Prompt Segmentation (CLAPS), a novel method for unified segmentation across diverse tasks and modalities in retinal imaging. Our approach begins by pre-training a CLIP-based image encoder on a large, multi-modal retinal dataset to handle data scarcity and distribution imbalance. We then leverage GroundingDINO to automatically generate spatial bounding box prompts by detecting local lesions. To unify tasks and resolve ambiguity, we use text prompts enhanced with a unique “modality signature” for each imaging modality. Ultimately, these automated textual and spatial prompts guide SAM to execute precise segmentation, creating a fully automated and unified pipeline. Extensive experiments on 12 diverse datasets across 11 critical segmentation categories show that CLAPS achieves performance on par with specialized expert models while surpassing existing benchmarks across most metrics, demonstrating its broad generalizability as a foundation model. Yinzheng Zhao, Junjie Yang 0001, Xiangtong Yao, Quanmin Liang, Shahrooz Faghih Roohi, Kai Huang 0001, Nassir Navab, M. Ali Nasseri |
BIBM | 7 |
| 2025 | UOPSL: Unpaired OCT Predilection Sites Learning for Fundus Image Diagnosis AugmentationabstractSignificant advancements in AI-driven multimodal medical image diagnosis have led to substantial improvements in ophthalmic disease identification in recent years. However, acquiring paired multimodal ophthalmic images remains prohibitively expensive. While fundus photography is simple and cost-effective, the limited availability of OCT data and inherent modality imbalance hinder further progress. Conventional approaches that rely solely on fundus or textual features often fail to capture fine-grained spatial information, as each imaging modality provides distinct cues about lesion predilection sites. In this study, we propose a novel unpaired multimodal framework UOPSL that utilizes extensive OCT-derived spatial priors to dynamically identify predilection sites, enhancing fundus imagebased disease recognition. Our approach bridges unpaired fundus and OCTs via extended disease text descriptions. Initially, we employ contrastive learning on a large corpus of unpaired OCT and fundus images while simultaneously learning the predilection sites matrix in the OCT latent space. Through extensive optimization, this matrix captures lesion localization patterns within the OCT feature space. During the fine-tuning or inference phase of the downstream classification task based solely on fundus images, where paired OCT data is unavailable, we eliminate OCT input and utilize the predilection sites matrix to assist in fundus image classification learning. Extensive experiments conducted on 9 diverse datasets across 28 critical categories demonstrate that our framework outperforms existing benchmarks. Yinzheng Zhao, Junjie Yang 0001, Xiangtong Yao, Quanmin Liang, Daniel Zapp, Kai Huang 0001, Nassir Navab, M. Ali Nasseri |
BIBM | 7 |
| 2025 | FIFA: Fine-grained Inter-frame Attention for Driver's Video Gaze EstimationabstractGaze direction serves as a pivotal indicator for assessing the level of driver attention. While image-based gaze estimation has been extensively researched, there has been a recent shift towards capturing gaze direction from video sequences. This approach encounters notable challenges, including the comprehension of the dynamic pupil evolution across frames and the extraction of head pose information from a relatively static background. To surmount these challenges, we introduce a dual-stream deep learning framework that explicitly models the displacement changes of the pupil through a fine-grained inter-frame attention mechanism and generates weights to adjust gaze embeddings. This technique transforms the face into a set of distinct patches and employs cross-attention to ascertain the correlation between pixel displacements in various patches and adjacent frames, thereby tracking spatial dynamics within the sequence. Our method is validated using two publicly available driver gaze datasets, and the results indicate that it achieves state-of-the-art performance or is on par with the best outcomes while reducing the parameters. Daosong Hu, Mingyue Cui, Kai Huang 0001 |
CVPR | 3 |
| 2025 | FSHNet: Fully Sparse Hybrid Network for 3D Object DetectionabstractFully sparse 3D detectors have recently gained significant attention due to their efficiency in long-range detection. However, sparse 3D detectors extract features only from non-empty voxels, which impairs long-range interactions and causes the center feature missing. The former weakens the feature extraction capability, while the latter hinders network optimization. To address these challenges, we introduce the Fully Sparse Hybrid Network (FSHNet). FSHNet incorporates a proposed SlotFormer block to enhance the long-range feature extraction capability of existing sparse encoders. The SlotFormer divides sparse voxels using a slot partition approach, which, compared to traditional window partition, provides a larger receptive field. Additionally, we propose a dynamic sparse label assignment strategy to deeply optimize the network by providing more high-quality positive samples. To further enhance performance, we introduce a sparse upsampling module to refine downsampled voxels, preserving fine-grained details crucial for detecting small objects. Extensive experiments on the Waymo, nuScenes, and Argoverse2 benchmarks demonstrate the effectiveness of FSHNet. The code is available at https://github.com/Say2L/FSHNet. Shuai Liu 0009, Mingyue Cui, Boyang Li 0009, Quanmin Liang, Tinghe Hong, Yunxiao Shan, Kai Huang 0001 |
CVPR | 7 |
| 2025 | HAS-GPU: Efficient Hybrid Auto-scaling with Fine-Grained GPU Allocation for SLO-Aware Serverless Inferences
Jianfeng Gu 0001, Puxuan Wang, Isaac David Núñez Araya, Kai Huang 0001, Michael Gerndt |
Euro-Par (1) | 4 |
| 2025 | Region Expansion: Optimization of Patch-Fetching Method for Point Cloud Denoising
Shengtao Li, Jiadun Wang, Daosong Hu, Kai Huang 0001 |
ICANN (2) | 5 |
| 2025 | Efficient Event Camera Data Pretraining with Adaptive Prompt Fusion
Quanmin Liang, Shuai Liu 0009, Xinzi Cao, Jinyi Lu, Feidiao Yang, Wei Zhang 0161, Kai Huang 0001, Yonghong Tian 0001 |
ICCV | 8 |
| 2025 | Planar KNN for Multi-camera Interference Mitigation of Point Cloud
Shengtao Li, Jiadun Wang, Daosong Hu, Kai Huang 0001 |
ICIC (1) | 4 |
| 2025 | Stable Tracking of Eye Gaze Direction During Ophthalmic SurgeryabstractOphthalmic surgical robots offer superior stability and precision by reducing the natural hand tremors of human surgeons, enabling delicate operations in confined surgical spaces. Despite the advancements in developing vision- and force-based control methods for surgical robots, preoperative navigation remains heavily reliant on manual operation, limiting the consistency and increasing the uncertainty. Existing eye gaze estimation techniques in the surgery, whether traditional or deep learning-based, face challenges including dependence on additional sensors, occlusion issues in surgical environments, and the requirement for facial detection. To address these limitations, this study proposes an innovative eye localization and tracking method that combines machine learning with traditional algorithms, eliminating the requirements of landmarks and maintaining stable iris detection and gaze estimation under varying lighting and shadow conditions. Extensive real-world experiment results show that our proposed method has an average estimation error of 0.58 degrees for eye orientation estimation and 2.08-degree average control error for the robotic arm's movement based on the calculated orientation. Tinghe Hong, Shenlin Cai, Boyang Li 0009, Kai Huang 0001 |
ICRA | 4 |
| 2025 | Autonomous Continuous Capsulorhexis Based on a Force-Vision-Guided Robot SystemabstractCapsulorhexis is challenging in cataract surgery, since the size, centering, and circularity of the capsule are important. Those indicators are closely related to the subsequent step of phacoemulsification and the postoperative position of the intraocular lens. It takes 3-5 years for a resident to practice, while the occurrence of deficient capsulorhexis is still inevitable. This paper proposes a robotic system to automate Continuous Curvilinear Capsulorhexis (CCC) in cataract surgery. A typical ophthalmic microscope system and a triaxial force sensor are utilized to guide the robot system with a force-vision method. The constraint of a Remote Center of Motion (RCM) is designed to perform the surgery route. The experimental results on exvivo porcine eyes show our autonomous method can achieve a satisfactory 6 mm capsule. With an average centering deviation below 7.6 % and circularity of 0.993, the consistency of the capsulorhexis is comparable to a surgeon-made one. Hongli Liang, M. Ali Nasseri, Haotian Lin 0001, Kai Huang 0001 |
ICRA | 5 |
| 2025 | Intraoperative Trocar-Based Eyeball Rotation Estimation Using Only 2D Microscope ImagesabstractIn ophthalmic surgery, surgeons or robots manipulate a light probe and an instrument around two separated trocars following sclerotomy to achieve orbital control for eyeball pose adjustment and subsequent surgical tasks referring to microscope frames. However, current methods face significant challenges in directly extracting the eyeball pose from real-time microscope frames due to the limited microscope perspective and the darkened operating room (OR). This paper decomposes eyeball rotations only along the x and y axes. Then, a method of calculating eyeball poses using eyeball geometry and microscopic trocar positions is presented. This method is tested by simulation and a phantom system with current [2.0, 2.8] degree error, providing assistant intraoperative eyeball status in the dark OR with extended method discussions. Junjie Yang 0001, Satoshi Inagaki, Daniel Zapp, Mathias Maier, Peter C. Issa, Kai Huang 0001, Nassir Navab, M. Ali Nasseri |
ICRA | 7 |
| 2025 | Learning-Based Quadruped Robot Framework for Locomotion on Dynamic Rigid PlatformsabstractTypical robot controllers assume firm ground, limiting their effectiveness in controlling robots on dynamic platforms such as trucks or ships. To address this limitation, we propose a reinforcement learning framework for robot locomotion on dynamic rigid platforms and a simulation in which 6-DoF dynamic platforms emulating ship oscillation. The framework enables a reinforcement learning model to estimate platform motion during robot locomotion control. In the simulation, our framework significantly reduces the quadruped robot’s fall rate and trajectory deviation compared to baseline controllers. Experiments on a real robot show that our framework enabled a quadruped robot to adapt to platform motions, including those that threw the robot into the air, while baseline models struggled in this case. Thus, our framework can advance the deployment of robots in real-world marine and vehicular applications. Kai Huang 0001, Heming Feng, Tianjiang Hu |
IROS | 1 |
| 2025 | EANS: Reducing Energy Consumption for UAV with an Environmental Adaptive Navigation StrategyabstractUnmanned Aerial Vehicles (Uavs) are limited by the onboard energy. Refinement of the navigation strategy directly affects both the flight velocity and the trajectory based on the adjustment of key parameters in the Uavs pipeline, thus reducing energy consumption. However, existing techniques tend to adopt static and conservative strategies in dynamic scenarios, leading to inefficient energy reduction. Dynamically adjusting the navigation strategy requires overcoming the challenges including the task pipeline interdependencies, the environmental-strategy correlations, and the selecting parameters. To solve the aforementioned problems, this paper proposes a method to dynamically adjust the navigation strategy of the Uavs by analyzing its dynamic characteristics and the temporal characteristics of the autonomous navigation pipeline, thereby reducing Uavs energy consumption in response to environmental changes. We compare our method with the baseline through hardware-in-the-loop (HIL) simulation and real-world experiments, showing our method 3.2X and 2.6X improvements in mission time, 2.4X and 1.6X improvements in energy, respectively. Boyang Li 0009, Long Chen 0005, Kai Huang 0001 |
IROS | 5 |
| 2025 | A Two-Stage Method for Specular Highlight Detection and Removal in Medical Images
Zefeng Li, Mingyue Cui, Daosong Hu, Jin Gong, Jingchong Weng, Lele Tian, Kai Huang 0001 |
MICCAI (10) | 9 |
| 2025 | ESOD: Event-Based Small Object DetectionabstractEvent-based object detection plays a crucial role in scenarios involving high-speed motion, extreme lighting conditions, and high-frequency detection. However, existing methods fail to address the challenges posed by small objects, including discriminative feature deficiency, the loss of critical information, and the inherent sparsity of event data. Moreover, the lack of benchmark datasets has significantly hindered progress in this field. To tackle these issues, we propose the Fully Deformable Detection Network (FDDNet), a lightweight framework that dynamically adapts to extract key features. First, we introduce a Long-Term Deformable Temporal Receptive Module (LDTR), which aligns critical features across consecutive event streams and leverages a State Space Model for long-range temporal modeling, enhancing the detection of high-speed small objects. Second, to address the sparsity of event data and the concentration of key features along object edges, we design a Sparse Feature Aggregation Block (SFAB) within the backbone and a coarse-to-fine deformable detection head, enabling hierarchical feature refinement from local to global, and improving the detection quality of sparse targets. Finally, to mitigate the lack of event-based small object datasets, we develop a high-quality, annotation-free data acquisition method and collect a real-world benchmark dataset for validation. Extensive experiments demonstrate that our approach achieves state-of-the-art (SOTA) performance on event-based small object detection tasks, with a mAP of 37.4% (+2.4%) on our benchmark and runs at 88 FPS, showcasing both accuracy and real-time capability. Our code and Supplement are available at https://github.com/Lqm26/ESOD. Quanmin Liang, Jinyi Lu, Shuai Liu 0009, Yinzheng Zhao, Wei Zhang 0161, Kai Huang 0001, Yonghong Tian 0001 |
ACM Multimedia | 8 |
| 2025 | GaussianFusion: Gaussian-Based Multi-Sensor Fusion for End-to-End Autonomous DrivingabstractMulti-sensor fusion is crucial for improving the performance and robustness of end-to-end autonomous driving systems. Existing methods predominantly adopt either attention-based flatten fusion or bird’s eye view fusion through geometric transformations. However, these approaches often suffer from limited interpretability or dense computational overhead. In this paper, we introduce GaussianFusion, a Gaussian-based multi-sensor fusion framework for end-to-end autonomous driving. Our method employs intuitive and compact Gaussian representations as intermediate carriers to aggregate information from diverse sensors. Specifically, we initialize a set of 2D Gaussians uniformly across the driving scene, where each Gaussian is parameterized by physical attributes and equipped with explicit and implicit features. These Gaussians are progressively refined by integrating multi-modal features. The explicit features capture rich semantic and spatial information about the traffic scene, while the implicit features provide complementary cues beneficial for trajectory planning. To fully exploit rich spatial and semantic information in Gaussians, we design a cascade planning head that iteratively refines trajectory predictions through interactions with Gaussians. Extensive experiments on the NAVSIM and Bench2Drive benchmarks demonstrate the effectiveness and robustness of the proposed GaussianFusion framework. The source code is included in the supplementary material and will be released publicly. Shuai Liu 0009, Quanmin Liang, Zefeng Li, Boyang Li 0009, Kai Huang 0001 |
NeurIPS | 5 |
| 2025 | TeamFed: Teamwork Principles-Inspired Federated Learning for 3D Object Detection
Siheng Ren, Boyang Li 0009, Shuai Liu 0009, Jiahui Liao, Mingyue Cui, Kai Huang 0001 |
PRCV (11) | 6 |
| 2025 | GHR-2D: Gaze and head redirection via disentanglement and diffusion for gaze estimation
Daosong Hu, Mingyue Cui, Kai Huang 0001 |
Eng. Appl. Artif. Intell. | 3 |
| 2025 | A variable-gain fixed-time convergent neurodynamic network for time-variant quadratic programming under unknown noises
Biao Song, Tinghe Hong, Weibing Li, Gang Chen 0023, Yongping Pan 0001, Kai Huang 0001 |
Neurocomputing | 6 |
| 2025 | FedPillarNet: Unifying personalized and global features for federated 3D LiDAR object detection
Boyang Li 0009, Siheng Ren, Shuai Zhao 0004, Mingyue Cui, Kai Huang 0001 |
J. Syst. Archit. | 5 |
| 2025 | Energy Efficient Scheduling for Position Reconfiguration of Swarm DronesabstractEnhancing the energy efficiency of drones, particularly in extending the flight lifetime, has emerged as a crucial area. Position reconfiguration has been explored as a mechanism to achieve this goal for swarm drones. Building on this concept, we investigate how position reconfiguration can be applied within urban wind environments to further extend the lifetime of drone swarms. Despite its potential, efficiently implementing position reconfiguration remains challenging. To address it, we propose an efficient position reconfiguration scheme that reduces the energy consumption imbalance of the swarm and prolongs the lifetime. The scheme includes: (1) a MIP (mixed integer programming)-based optimization method. (2) an approximation algorithm that runs in pseudo-polynomial time and without the need for an optimization solver. The scheme provides a complete position reconfiguration solution that determines (i) the number of position reconfiguration; (ii) when to perform reconfiguration; (iii) who to change positions. Simulation and experimental results demonstrate the effectiveness of our scheme. Note to Practitioners—In urban environments, the significant variation in wind speeds leads to an energy imbalance among swarm drones performing tasks. This paper addresses the practical issue of extending the lifetime of drones in such environments by optimizing position reconfiguration. Specifically, drones operating in high wind speed areas require more energy to maintain hovering, resulting in faster battery depletion. By allowing drones with more remaining energy to exchange positions with those experiencing higher energy consumption, the overall energy usage can be balanced, thus extending the mission duration. We propose an energy-efficient scheduling scheme to determine when and which drones should reconfigure their positions. The scheme strikes a balance between the benefits of reconfiguration and the associated energy costs, preventing unnecessary movement that could waste energy while ensuring drones do not deplete their batteries prematurely. This solution is particularly suited for drone swarms operating in urban environments. Future research could further explore the integration of this scheme into real-time drone fleet management systems. Mingxin Wei, Shuai Zhao 0004, Hui Cheng 0002, Kai Huang 0001 |
IEEE Trans Autom. Sci. Eng. | 5 |
| 2025 | GAEM: Graph-Driven Attention-Based Entropy Model for LiDAR Point Cloud CompressionabstractHigh-quality LiDAR point cloud (LPC) coding is essential for efficiently transmitting and storing the vast amounts of data required for accurate 3D environmental representation. The Octree-based entropy coding framework has emerged as the predominant method, however, previous study usually overly relies on large-scale attention-based context prediction to encode Octree nodes, overlooking the inherent correlational properties of this structure. In this paper, we propose a novel Graph-driven Attention-based Entropy Model (GAEM), which adopts partitioned graph attention mechanisms to uncover contextual dependencies among neighboring nodes. Different from the Cartesian coordinate-based coding mode with higher redundancy, GAEM uses the multi-level spherical Octree to organize point clouds, improving the quality of LPC reconstruction. GAEM combines graph convolution for node feature embedding and grouped-graph attention for exploiting dependency among contexts, which preserves performance in low-computation using localized nodes. Besides, to further increase the receptive field, we design a high-resolution cross-attention module introducing sibling nodes. Experimental results show that our method achieves state-of-the-art performance on the LiDAR benchmark SemanticKITTI and MPEG-specified dataset Ford, compared to all baselines. Compared to the benchmark GPCC, our method achieves gains of up to 53.9% and 53.6% on SemanticKITTI and Ford while compared to the sibling-introduced methods, we achieve up to 42.3% and 44.7% savings in encoding/decoding time. In particular, our GAEM allows for extension to downstream tasks (i.e.,vehicle detection and semantic segmentation), further demonstrating the practicality of the method. Mingyue Cui, Yuyang Zhong, Mingjian Feng, Junhua Long, Yehua Ling, Kai Huang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2025 | DAPCC: Diverse Attention-Based Entropy Model for Dynamic LiDAR Point Cloud CompressionabstractLiDAR point cloud (LPC) compression is an indispensable component for 3D vision tasks, especially for dynamic point clouds. However, the existing methods based on traditional spatial-temporal attention are immature, causing little improvement in inter-frame feature extraction. In this paper, we propose Diverse Attention-based Point Cloud Compression (DAPCC), an LPC compression entropy model combining aggregation embedding modules for temporal point matching and spatial-temporal attention blocks for dynamic Octree node encoding, which can effectively utilize the change information of dynamic point clouds. Specifically, we first introduce aggregation embedding to match the Octree sequences from two sweeps to establish temporal correlation. To effectively capture the feature details, we further design local and global combined attention for the spatial-temporal information of point clouds which can focus on the whole context. Finally, we organize a symmetric MLP module capable of strengthening vital features. We conduct experiments of static and dynamic compression on both indoor/outdoor point cloud benchmark datasets (i.e., ScanNet, SemanticKITTI, and MPEG Common Test Conditions (CTC) Category 3 datasets) and downstream applications (i.e., vehicle detection and semantic segmentation). Compared with the previous state-of-the-art methods, our method achieves up to 14.7% bpp and 45% decoding time savings and adapts to the downstream tasks with almost no impact on performance. Mingyue Cui, Yuyang Zhong, Mingjian Feng, Yehua Ling, Junhua Long, Jinhong Xia, Kai Huang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2025 | UnMoDE: Uncertainty Modeling for Driver Gaze Estimation via Feature DisentanglementabstractGaze estimation can be used for assessing the attention level of drivers. Current works predominantly focus on enhancing model accuracy, often overlooking the influence of input sample and label uncertainty. In this paper, we propose a framework for uncertainty modeling in driver gaze estimation via feature disentanglement, referred to as UnMoDE. Our approach begins by extracting facial information into distinct feature spaces using an asymmetric dual-branch encoder to obtain gaze features. Subsequently, a multi-layer perceptron (MLP) is employed to project gaze features and labels into an embedding space, representing them as Gaussian distributions. The uncertainty is described using a covariance matrix. Random sampling is applied to derive samples from the gaze embedding distribution to estimate the most probable embedding representation. This estimated representation is then used to regress the gaze direction and is projected back into the gaze feature space, along with identity information, to facilitate facial reconstruction. Extensive experimental evaluations demonstrate that UnMoDE significantly outperforms baseline and state-of-the-art methods on the latest benchmark datasets collected for drivers, particularly in reducing the number of samples with significant errors. Daosong Hu, Mingyue Cui, Kai Huang 0001 |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2024 | Event-Based Image Enhancement Under High Dynamic Range Scenarios
Jingchong Weng, Boyang Li 0009, Kai Huang 0001 |
ACCV (6) | 3 |
| 2024 | Fault-tolerant DAG Scheduling with Runtime Reconfiguration on Multicore Real-Time SystemsabstractFault tolerance and real-time performance are two essential goals for directed acyclic graph (DAG) scheduling. However, the redundant tasks to compensate for faults can significantly prolong the completion time of a DAG, i.e., the makespan. In addition, the unpredictable runtime failure status, i.e., a fault may or may not occur during task execution, results in a huge difference between the offline schedule and the actual execution. Existing list scheduling methods can support fault-tolerant execution with all redundant tasks taken into account before runtime. However, such methods cannot effectively reduce the actual makespan as conditional execution of redundant tasks is not considered during scheduling. To address the above issues, this paper proposes a fault-tolerant DAG scheduling method with a runtime reconfiguration facility to optimize the actual makespan. First, the fault-aware makespan is optimized by a fine-grained offline scheduling method considering the worst-case scenario and the runtime flexibility. Then, the runtime reconfiguration mechanism safely moves the influential nodes ahead using the additional time interval to minimize the actual makespan. The experimental results indicate that the proposed method outperforms the state-of-the-art (SOTA) methods in terms of schedulability and actual makespan. Yuanhai Zhang, Shuai Zhao 0004, Gang Chen 0023, Kai Huang 0001 |
ASAP | 4 |
| 2024 | 4D-CAT: Synthesis of 4D Coronary Artery Trees from Systole and DiastoleabstractThe three-dimensional vascular model reconstructed from CT images is widely used in medical diagnosis. At different phases, the beating of the heart can cause deformation of vessels, resulting in different vascular imaging states and false positive diagnostic results. The 4D model can simulate a complete cardiac cycle. Due to the dose limitation of contrast agent injection in patients, it is valuable to synthesize a 4D coronary artery trees through finite phases imaging. In this paper, we propose a method for generating a 4D coronary artery trees, which maps the systole to the diastole through deformation field prediction, interpolates on the timeline, and the motion trajectory of points are obtained. Specifically, the centerline is used to represent vessels and to infer deformation fields using cube-based sorting and neural networks. Adjacent vessel points are aggregated and interpolated based on the deformation field of the centerline point to obtain displacement vectors of different phases. Finally, the proposed method is validated through experiments to achieve the registration of non-rigid vascular points and the generation of 4D coronary trees. Daosong Hu, Ruomeng Wang, Mingyue Cui, Kai Huang 0001 |
BIBM | 6 |
| 2024 | A Tremor Evaluation Method for Retinal Injection Micro-needle by Microscopic VideosabstractVarious types of surgical robots have gradually occupied more positions in modern medical care due to safety and stability reasons. Retinal surgical robots are still in the development stage, so there is still a lack of some surgical indicators to evaluate the safety and stability of surgical robots. In order to evaluate the performance of robot-assisted surgery more accurately, this paper proposes a method to evaluate the tremor of tools based on video collected by a microscope in surgery. This method is mainly dedicated to evaluating the tremor level of tools in clinical retinal surgery, in order to verify the pulling of the wound during surgery. The method is mainly divided into three parts: tool semantic segmentation, key point extraction, tip-end reasoning and tremor evaluation. Firstly, the method extracts the outline label of the tool through semantic segmentation. Then, some key points on the outline, which could reflect the posture, could be obtained by skeleton extraction, edge straight line fitting. After that, these points could be used to assess the tilt angle of the needle. Also, according to the angle and the actual length of the needle tip, the position of the needle tip could be assessed in the image, and finally the algorithm could evaluate the tremor level based on the change in position. This method could evaluate the tremor of the tool in clinical surgery, which can help to evaluate the safety and stability of the retinal injection surgery. Yanlin Li 0006, Rihui Song, Andi Xu, Haotian Lin 0001, Kai Huang 0001 |
BIBM | 5 |
| 2024 | Extrapolating Prospective Glaucoma Fundus Images through Diffusion in Irregular Longitudinal SequencesabstractThe utilization of longitudinal datasets for glaucoma progression prediction offers a compelling approach to support early therapeutic interventions. Predominant methodologies in this domain have primarily focused on the direct prediction of glaucoma stage labels from longitudinal datasets. However, such methods may not adequately encapsulate the nuanced developmental trajectory of the disease. To enhance the diagnostic acumen of medical practitioners, we propose a novel diffusion-based model to predict prospective images by extrapolating from existing longitudinal fundus images of patients. The methodology delineated in this study distinctively leverages sequences of images as inputs. Subsequently, a time-aligned mask is employed to select a specific year for image generation. During the training phase, the time-aligned mask resolves the issue of irregular temporal intervals in longitudinal image sequence sampling. Additionally, we utilize a strategy of randomly masking a frame in the sequence to establish the ground truth. This methodology aids the network in continuously acquiring knowledge regarding the internal relationships among the sequences throughout the learning phase. Moreover, the introduction of textual labels is instrumental in categorizing images generated within the sequence. The empirical findings from the conducted experiments indicate that our proposed model not only effectively generates longitudinal data but also significantly improves the precision of downstream classification tasks. Junjie Yang 0001, Shahrooz Faghih Roohi, Yinzheng Zhao, Daniel Zapp, Kai Huang 0001, Nassir Navab, M. Ali Nasseri |
BIBM | 6 |
| 2024 | KLDD: Kalman Filter based Linear Deformable Diffusion Model in Retinal Image SegmentationabstractAI-based vascular segmentation is becoming increasingly common in enhancing the screening and treatment of ophthalmic diseases. Deep learning structures based on U-Net have achieved relatively good performance in vascular segmentation. However, small blood vessels and capillaries tend to be lost during segmentation when passed through the traditional U-Net downsampling module. To address this gap, this paper proposes a novel Kalman filter based Linear Deformable Diffusion (KLDD) model for retinal vessel segmentation. Our model employs a diffusion process that iteratively refines the segmentation, leveraging the flexible receptive fields of deformable convolutions in feature extraction modules to adapt to the detailed tubular vascular structures. More specifically, we first employ a feature extractor with linear deformable convolution to capture vascular structure information form the input images. To better optimize the coordinate positions of deformable convolution, we employ the Kalman filter to enhance the perception of vascular structures in linear deformable convolution. Subsequently, the features of the vascular structures extracted are utilized as a conditioning element within a diffusion model by the Cross-Attention Aggregation module (CAAM) and the Channel-wise Soft Attention module (CSAM). These aggregations are designed to enhance the diffusion model’s capability to generate vascular structures. Experiments are evaluated on retinal fundus image datasets (DRIVE, CHASE DB1) as well as the 3mm and 6mm of the OCTA-500 dataset, and the results show that the diffusion model proposed in this paper outperforms other methods. Yinzheng Zhao, Junjie Yang 0001, Kai Huang 0001, Nassir Navab, M. Ali Nasseri |
BIBM | 4 |
| 2024 | Bilateral Event Mining and Complementary for Event Stream Super-ResolutionabstractEvent Stream Super-Resolution (ESR) aims to address the challenge of insufficient spatial resolution in event streams, which holds great significance for the application of event cameras in complex scenarios. Previous works for ESR often process positive and negative events in a mixed paradigm. This paradigm limits their ability to effectively model the unique characteristics of each event and mutually refine each other by considering their correlations. In this paper, we propose a bilateral event mining and complementary network (BMCNet) to fully leverage the potential of each event and capture the shared information to complement each other simultaneously. Specifically, we resort to a two-stream network to accomplish comprehensive mining of each type of events individually. To facilitate the exchange of information between two streams, we propose a bilateral information exchange (BIE) module. This module is layer-wisely embedded between two streams, enabling the effective propagation of hierarchical global information while alleviating the impact of invalid information brought by inherent characteristics of events. The experimental results demonstrate that our approach outperforms the previous state-of-the-art methods in ESR, achieving performance improvements of over 11% on both real and synthetic datasets. Moreover, our method significantly enhances the performance of event-based downstream tasks such as object recognition and video reconstruction. Our code is available at https://github.com/Lqm26/BMCNet-ESR. Zhilin Huang, Quanmin Liang, Yijie Yu 0001, Chujun Qin, Xiawu Zheng, Kai Huang 0001, Zikun Zhou, Wenming Yang |
CVPR | 6 |
| 2024 | A Du-Octree based Cross-Attention Model for LiDAR Geometry CompressionabstractPoint cloud compression is an essential technology for efficient storage and transmission of 3D data. Previous methods usually use hierarchical tree data structures for encoding the spatial sparseness of point clouds. However, the node context within the tree is not fully discovered since the feature space among nodes varies significantly. To address this problem, we innovatively represent the LiDAR points in a two-octree structure instead of using traditional single-octree coding, and then design the cross-attention model to capture the hierarchical features between different octrees, of which each octree incorporates a transformer-based deep entropy model and an arithmetic encoder. Besides, we introduce the untied cross-aware position encoding with principal component analysis and different projection matrices, which enhances the correlations over two octrees’ attention feature embeddings. Experimental results show that our method outperforms the previous state-of-the-art works, achieving up to 8.2% Bpp savings on point cloud benchmark datasets with different lasers. Mingyue Cui, Mingjian Feng, Junhua Long, Daosong Hu, Shuai Zhao 0004, Kai Huang 0001 |
ICRA | 6 |
| 2024 | Optimizing Dynamic Balance in a Rat Robot via the Lateral Flexion of a Soft Actuated SpineabstractBalancing oneself using the spine is a physiological alignment of the body posture in the most efficient manner by the muscular forces for mammals. For this reason, we can see many disabled quadruped animals can still stand or walk even with three limbs. This paper investigates the optimization of dynamic balance during trot gait based on the spatial relationship between the center of mass (CoM) and support area influenced by spinal flexion. During trotting, the robot balance is significantly influenced by the distance of the CoM to the support area formed by diagonal footholds. In this context, lateral spinal flexion, which is able to modify the position of footholds, holds promise for optimizing balance during trotting. This paper explores this phenomenon using a rat robot equipped with a soft actuated spine. Based on the lateral flexion of the spine, we establish a kinematic model to quantify the impact of spinal flexion on robot balance during trot gait. Subsequently, we develop an optimized controller for spinal flexion, designed to enhance balance without altering the leg locomotion. The effectiveness of our proposed controller is evaluated through extensive simulations and physical experiments conducted on a rat robot. Compared to both a non-spine based trot gait controller and a trot gait controller with lateral spinal flexion, our proposed optimized controller effectively improves the dynamic balance of the robot and retains the desired locomotion during trotting. Yuhong Huang, Zhenshan Bing, Zitao Zhang, Genghang Zhuang, Kai Huang 0001, Alois C. Knoll |
ICRA | 5 |
| 2024 | Shadow-Based 3D Pose Estimation of Intraocular Instrument Using Only 2D ImagesabstractIn ophthalmic surgeries, such as vitreoretinal operations, surgeons rely on imaging systems, primarily microscopes, for real-time instrument monitoring and motion planning. However, novice surgeons struggle to extract 3D instrument positions from 2D microscope frames, necessitating extensive trial-and-error experience with the background that additional imaging modalities such as iOCT remain inaccessible in most operating rooms. Targeting intraocular assessment within the current surgical setup, this paper presents an imagebased pose estimation method to obtain real-time instrument tip positions in a standard 12mm-radius spherical eyeball model, which links floating instruments with on-the-retinal objects based on the intraocular shadowing principle. We validate this estimation method in a Unity simulator and verify its depth estimation capability using a specially designed eyeball phantom. Both simulator and phantom experiments demonstrate an average needle-tip estimation error within [1.0, 2.0] mm using only 2D microscope frames. Junjie Yang 0001, Mathias Maier, Kai Huang 0001, Nassir Navab, M. Ali Nasseri |
ICRA | 4 |
| 2024 | Efficient Event Stream Super-Resolution with Recursive Multi-Branch Fusion
Quanmin Liang, Zhilin Huang, Xiawu Zheng, Feidiao Yang, Jun Peng 0007, Kai Huang 0001, Yonghong Tian 0001 |
IJCAI | 6 |
| 2024 | DCDet: Dynamic Cross-based 3D Object Detector
Shuai Liu 0009, Boyang Li 0009, Zhiyu Fang, Kai Huang 0001 |
IJCAI | 4 |
| 2024 | An Efficient Position Reconfiguration Approach for Maximizing Lifetime of Fixed-wing Swarm DronesabstractWith the development and application of swarm drones, some researchers have tried to replicating the migration patterns of geese in drones swarm formation to extend their lifetime. However, the problem of performing appropriate position reconfiguration based on the battery energy still remains an unsolved issue. This paper proposes an efficient position reconfiguration approach that reduces the energy consumption imbalance of the swarm and prolongs the lifetime. The approach includes: (1) a two-step MIP (mixed-integer programming)-based optimization method. (2) a two-step heuristic algorithm that can run in pseudo-polynomial time and without the need for an optimization solver. The approach provides a complete position reconfiguration solution that determines (i) the number of position reconfiguration; (ii) which drones need to exchange positions in every position reconfiguration; (iii) the length of time to maintain each position before next reconfiguration. Finally, the approach is compared with other three methods in experiments which demonstrate the effectiveness of it. Mingyue Cui, Yunxiao Shan, Shuai Zhao 0004, Kai Huang 0001 |
IROS | 6 |
| 2024 | An Online Rcm Adjusting System for Robot-Assisted Retinal SurgeriesabstractIn robot-assisted retinal surgery, a Remote Center of Motion (Rcm) allows the surgical instrument to rotate around a distal fixed point without any lateral translations. The Rcm point should be perfectly aligned inside the trocar. Otherwise, unexpected tool translations at the expected remote center will enlarge the force applied to the trocar and consequently result in post-operative complications. Due to the narrow size of the trocar and the lack of real-time detection equipment, the Rcm point is hard to be perfectly located inside the trocar. Even if the Rcm is perfectly aligned, the movement of the tissue around the eyeball could make it inappropriate again. In this paper, inspired by the control strategy of surgeons, an online Rcm adjusting strategy is proposed. Instead of only using one fixed Rcm point, to restrict the force between the surgical tool and the trocar, the proposed strategy adjusts the position of the Rcm point during the motion. The results show our approach significantly reduces the force between the robot end-effector and surgical port by 64.2%. In addition, the results also demonstrate that our approach complies the Rcm trajectories without deforming or spoiling the working space, which is significantly important for obeying surgeon’s instructions in practice. Ting Wang 0028, Huanqi Ni, Yanlin Li 0006, Ruoxi Chen, M. Ali Nasseri, Haotian Lin 0001, Kai Huang 0001 |
IROS | 8 |
| 2024 | Shadow Maintenance for Automatic Light-Probe Control in Ophthalmic Surgeries Using Only 2D informationabstractIn ophthalmic surgeries, the light probe is responsible for providing safe intraocular illumination and ensuring the visibility of the instrument and its shadow as the only available reference for qualitative depth estimation and landing point prediction in fundus microscopic images. To achieve sustainable shadow-based estimation during surgeries, we propose controlling the light probe automatically to limit the shadow position around the instrument tip using only 2D information from the microscope. We also integrate an intensity balancing sub-module to guarantee the normal intensity distribution and the safe depth of light-tip placement. Without motor-based pose coordination between the light probe and the instrument, experiments analyze the performance of our image-based shadow maintenance with only image information under the constraints of RCM and discuss the working volume and segmentation limitations during simulation and real-robot tests. Junjie Yang 0001, Satoshi Inagaki, Daniel Zapp, Mathias Maier, Kai Huang 0001, Nassir Navab, M. Ali Nasseri |
IROS | 6 |
| 2024 | Intraocular Reflection Modeling and Avoidance Planning in Image-Guided Ophthalmic SurgeriesabstractIntuitive enhancement of surgical precision in robotic retinal surgery highly depends on the stable acquisition of intraocular imaging data. Such acquisition requires segmenting intraocular components, especially instrument-tip positions, to achieve state estimation and subsequent navigation and motion control. However, intraocular light reflections and glares significantly impact instrument segmentation, state estimation, and subsequent visual servoing in retinal surgery. At the same time, light reflections are among the sources of information for intraoperative navigation. In this work, we propose a method for modeling and optimizing light reflections using microscopy as the standard surgical imaging modality. Beyond optimization, our approach seamlessly integrates the optimized reflection with path planning, strategically circumventing reflection areas and ensuring uninterrupted visibility of instrument tips throughout the surgical procedure. Experiments demonstrate the methodology’s efficacy in avoiding glare affections during eye surgeries. Junjie Yang 0001, Yinzheng Zhao, Daniel Zapp, Mathias Maier, Kai Huang 0001, Nassir Navab, M. Ali Nasseri |
IROS | 6 |
| 2024 | FFAM: Feature Factorization Activation Map for Explanation of 3D DetectorsabstractLiDAR-based 3D object detection has made impressive progress recently, yet most existing models are black-box, lacking interpretability. Previous explanation approaches primarily focus on analyzing image-based models and are not readily applicable to LiDAR-based 3D detectors. In this paper, we propose a feature factorization activation map (FFAM) to generate high-quality visual explanations for 3D detectors. FFAM employs non-negative matrix factorization to generate concept activation maps and subsequently aggregates these maps to obtain a global visual explanation. To achieve object-specific visual explanations, we refine the global visual explanation using the feature gradient of a target object. Additionally, we introduce a voxel upsampling strategy to align the scale between the activation map and input point cloud. We qualitatively and quantitatively analyze FFAM with multiple detectors on several datasets. Experimental results validate the high-quality visual explanations produced by FFAM. The code is available at \url{https://anonymous.4open.science/r/FFAM-B9AF}. Shuai Liu 0009, Boyang Li 0009, Zhiyu Fang, Mingyue Cui, Kai Huang 0001 |
NeurIPS | 5 |
| 2024 | FedDeSnowNet: Federated De-snowing Network for LiDAR Point Clouds
Zhiyu Fang, Boyang Li 0009, Jiahui Liao, Siheng Ren, Kai Huang 0001 |
NPC (2) | 5 |
| 2024 | Robotic Control of Endoscope Assistance in Skull Base Surgery Based on Adaptive RCM Point
Tinghe Hong, Boyang Li 0009, Weibing Li, Kai Huang 0001 |
PRICAI (5) | 4 |
| 2024 | NIV-SSD: Neighbor IoU-voting single-stage object detector from point cloud
Shuai Liu 0009, Di Wang 0011, Quan Wang 0006, Kai Huang 0001 |
Neurocomputing | 4 |
| 2024 | Timing-accurate scheduling and allocation for parallel I/O operations in real-time systems
Yuanhai Zhang, Shuai Zhao 0004, Gang Chen 0023, Haoyu Luo, Kai Huang 0001 |
J. Syst. Archit. | 5 |
| 2024 | Context-Based Meta-Reinforcement Learning With Bayesian Nonparametric ModelsabstractDeep reinforcement learning agents usually need to collect a large number of interactions to solve a single task. In contrast, meta-reinforcement learning (meta-RL) aims to quickly adapt to new tasks using a small amount of experience by leveraging the knowledge from training on a set of similar tasks. State-of-the-art context-based meta-RL algorithms use the context to encode the task information and train a policy conditioned on the inferred latent task encoding. However, most recent works are limited to parametric tasks, where a handful of variables control the full variation in the task distribution, and also failed to work in non-stationary environments due to the few-shot adaptation setting. To address those limitations, we propose MEta-reinforcement Learning with Task Self-discovery (MELTS), which adaptively learns qualitatively different nonparametric tasks and adapts to new tasks in a zero-shot manner. We introduce a novel deep clustering framework (DPMM-VAE) based on an infinite mixture of Gaussians, which combines the Dirichlet process mixture model (DPMM) and the variational autoencoder (VAE), to simultaneously learn task representations and cluster the tasks in a self-adaptive way. Integrating DPMM-VAE into MELTS enables it to adaptively discover the multi-modal structure of the nonparametric task distribution, which previous methods using isotropic Gaussian random variables cannot model. In addition, we propose a zero-shot adaptation mechanism and a recurrence-based context encoding strategy to improve the data efficiency and make our algorithm applicable in non-stationary environments. On various continuous control tasks with both parametric and nonparametric variations, our algorithm produces a more structured and self-adaptive task latent space and also achieves superior sample efficiency and asymptotic performance compared with state-of-the-art meta-RL algorithms. Zhenshan Bing, Yuqi Yun, Kai Huang 0001, Alois C. Knoll |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2024 | Towards Human-Compatible Autonomous Car: A Study of Non-Verbal Turing Test in Automated Driving With Affective Transition ModellingabstractAutonomous cars are indispensable when humans go further down the hands-free route. Although existing literature highlights that the acceptance of the autonomous car will increase if it drives in a human-like manner, sparse research offers the naturalistic experience from a passenger's seat perspective to examine the humanness of current autonomous cars. The present study tested whether the AI driver could create a human-like ride experience for passengers based on 69 participants' feedback in a real-road scenario. We designed a ride experience-based version of the non-verbal Turing test for automated driving. Participants rode in autonomous cars (driven by either human or AI drivers) as a passenger and judged whether the driver was human or AI. The AI driver failed to pass our test because passengers detected the AI driver above chance. In contrast, when the human driver drove the car, the passengers' judgement was around chance. We further investigated how human passengers ascribe humanness in our test. Based on Lewin's field theory, we advanced a computational model combining signal detection theory with pre-trained language models to predict passengers' humanness rating behaviour. We employed affective transition between pre-study baseline emotions and corresponding post-stage emotions as the signal strength of our model. Results showed that the passengers' ascription of humanness would increase with the greater affective transition. Our study suggested an important role of affective transition in passengers' ascription of humanness, which might become a future direction for autonomous driving. Zhaoning Li, Qiaoli Jiang, Zhengming Wu, Anqi Liu 0001, Haiyan Wu, Miner Huang, Kai Huang 0001, Yixuan Ku |
IEEE Trans. Affect. Comput. | 7 |
| 2024 | Semi-Supervised Multitask Learning Using Gaze Focus for Gaze EstimationabstractGaze estimation can be applied in various scenarios, seeking to comprehend human visual attention through camera images. Contemporary research predominantly employs deep learning to directly output gaze from facial or ocular images. however, most methods concentrate solely on estimating gaze direction, overlooking gaze point. We propose two multitask learning frameworks for estimating gaze point and gaze direction, with the objective of achieving unsupervised learning of gaze point and supervised gaze estimation via gaze intersection. Two attention layers are proposed to guide the generation of facial features, addressing the challenge posed by unlabeled gaze point. The focus attention layer employs the eyes to guide facial features, connecting both features and utilizing similarity to enhance eye information. Another approach utilizes only the full face image, employing self-attention to enhance pertinent information. Four loss functions are employed to constrain networks in 2D and 3D spaces. The combination of eye position constraints and attention layers ensures the accuracy of gaze point prediction. Gaze intersection can be used to obtain gaze depth, thereby solving the problem of depth-overlapping. The advantages of the proposed method in gaze tracking are verified through comprehensive experiments. Daosong Hu, Kai Huang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | Model-Free Synchronous Motion Generation of Multiple Heterogeneous Continuum RobotsabstractHeterogeneous continuum robots (HCRs) with different structures have been designed for different purposes, whereas the coordination of multiple HCRs has received little attention. On one hand, multiple HCRs coordination brings the possibility of performing complicated tasks. On the other hand, the structural diversity of HCRs poses great difficulties to their modeling and control. This article proposes a model-free scheme for the synchronous motion control of multiple HCRs. The control problem is formulated as two convex optimization problems in a model-free closed-loop framework, including control quantity estimation and Jacobian matrix estimation. The proposed approach aims at addressing the synchronous motion problem of multiple HCRs in a decentralized way. The design of model-free feedback control guarantees the high adaptability of the proposed method for a wide range of HCRs. Simulation studies are performed to verify the effectiveness and adaptability of the proposed scheme for multiple HCRs. Comparative studies verify that the tracking error synthesized by the proposed method is about two-thirds lower than that of the existing method while the computational cost is similar, which reveals the merit of the proposed method in terms of accuracy. Finally, the feasibility and effectiveness of the proposed method are also verified by hardware-in-loop simulations and physical experiments on the synchronous motion control of a cable-driven continuum robot and a concentric-tube robot. Peng Yu 0003, Ning Tan 0003, Yuyang Wu, Binbin Qiu, Kai Huang 0001 |
IEEE Trans. Ind. Informatics | 5 |
| 2024 | MENet: Multi-Modal Mapping Enhancement Network for 3D Object Detection in Autonomous DrivingabstractTo achieve more accurate perception performance, LiDAR and camera are gradually chosen to improve 3D object detection simultaneously. However, it is still a non-trivial task to build an effective fusion mechanism, and this is hindering the development of multi-modal based method. Especially, the mapping relationship construction between two modalities is far from fully explored. Canonical cross-modal mapping suffers from failure when the calibration matrix is incorrect, and it also greatly wastes the amount and density of RGB image information. This paper aims to extend the traditional one-to-one alignment relationship between LiDAR and camera. For all projected point clouds, we enhance their cross-modal mapping relationship through aggregating color-texture related feature and shape-contour related feature. Further, a mapping pyramid is proposed to leverage the semantic representation of the image feature at different stages. Based on the above mapping enhancement strategies, our method increases the engagement rate of image. Finally, we design a fusion module based on an attention mechanism to improve the point cloud feature with the auxiliary image feature. Extensive experiments on the KITTI dataset and SUN-RGBD dataset show that our model achieves satisfactory 3D object detection, especially for categories with sparse point clouds compared with other multi-modal fusion networks. Moyun Liu, Youping Chen, Jingming Xie, Yang Zhang 0053, Zhenshan Bing, Genghang Zhuang, Kai Huang 0001, Joey Tianyi Zhou |
IEEE Trans. Intell. Transp. Syst. | 9 |
| 2024 | Meta-Reinforcement Learning in Nonstationary and Nonparametric EnvironmentsabstractRecent state-of-the-art artificial agents lack the ability to adapt rapidly to new tasks, as they are trained exclusively for specific objectives and require massive amounts of interaction to learn new skills. Meta-reinforcement learning (meta-RL) addresses this challenge by leveraging knowledge learned from training tasks to perform well in previously unseen tasks. However, current meta-RL approaches limit themselves to narrow parametric and stationary task distributions, ignoring qualitative differences and nonstationary changes between tasks that occur in the real world. In this article, we introduce a Task-Inference-based meta-RL algorithm using explicitly parameterized Gaussian variational autoencoders (VAEs) and gated Recurrent units (TIGR), designed for nonparametric and nonstationary environments. We employ a generative model involving a VAE to capture the multimodality of the tasks. We decouple the policy training from the task-inference learning and efficiently train the inference mechanism on the basis of an unsupervised reconstruction objective. We establish a zero-shot adaptation procedure to enable the agent to adapt to nonstationary task changes. We provide a benchmark with qualitatively distinct tasks based on the half-cheetah environment and demonstrate the superior performance of TIGR compared with state-of-the-art meta-RL approaches in terms of sample efficiency (three to ten times faster), asymptotic performance, and applicability in nonparametric and nonstationary environments with zero-shot adaptation. Videos can be viewed at https://videoviewsite.wixsite.com/tigr. Zhenshan Bing, Lukas Knak, Long Cheng 0007, Fabrice O. Morin, Kai Huang 0001, Alois C. Knoll |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2023 | OctFormer: Efficient Octree-Based Transformer for Point Cloud Compression with Local EnhancementabstractPoint cloud compression with a higher compression ratio and tiny loss is essential for efficient data transportation. However, previous methods that depend on 3D convolution or frequent multi-head self-attention operations bring huge computations. To address this problem, we propose an octree-based Transformer compression method called OctFormer, which does not rely on the occupancy information of sibling nodes. Our method uses non-overlapped context windows to construct octree node sequences and share the result of a multi-head self-attention operation among a sequence of nodes. Besides, we introduce a locally-enhance module for exploiting the sibling features and a positional encoding generator for enhancing the translation invariance of the octree node sequence. Compared to the previous state-of-the-art works, our method obtains up to 17% Bpp savings compared to the voxel-context-based baseline and saves an overall 99% coding time compared to the attention-based baseline. Mingyue Cui, Junhua Long, Mingjian Feng, Boyang Li 0009, Kai Huang 0001 |
AAAI | 5 |
| 2023 | Meta-Reinforcement Learning Based on Self-Supervised Task Representation LearningabstractMeta-reinforcement learning enables artificial agents to learn from related training tasks and adapt to new tasks efficiently with minimal interaction data. However, most existing research is still limited to narrow task distributions that are parametric and stationary, and does not consider out-of-distribution tasks during the evaluation, thus, restricting its application. In this paper, we propose MoSS, a context-based Meta-reinforcement learning algorithm based on Self-Supervised task representation learning to address this challenge. We extend meta-RL to broad non-parametric task distributions which have never been explored before, and also achieve state-of-the-art results in non-stationary and out-of-distribution tasks. Specifically, MoSS consists of a task inference module and a policy module. We utilize the Gaussian mixture model for task representation to imitate the parametric and non-parametric task variations. Additionally, our online adaptation strategy enables the agent to react at the first sight of a task change, thus being applicable in non-stationary tasks. MoSS also exhibits strong generalization robustness in out-of-distributions tasks which benefits from the reliable and robust task representation. The policy is built on top of an off-policy RL algorithm and the entire network is trained completely off-policy to ensure high sample efficiency. On MuJoCo and Meta-World benchmarks, MoSS outperforms prior works in terms of asymptotic performance, sample efficiency (3-50x faster), adaptation efficiency, and generalization robustness on broad and diverse task distributions. Mingyang Wang 0003, Zhenshan Bing, Xiangtong Yao, Shuai Wang 0007, Kai Huang 0001, Hang Su 0001, Chenguang Yang 0001, Alois C. Knoll |
AAAI | 5 |
| 2023 | Accelerated Optimization for Simulation of Brain Spiking Neural Network on GPGPUs
Fangzhou Zhang, Mingyue Cui, Jiakang Zhang, Yehua Ling, Kai Huang 0001 |
ICA3PP (6) | 6 |
| 2023 | Dense Depth Estimation for Surgical Endoscope Robot with Multi-Baseline Depth Map FusionabstractDense depth estimation in endoscopic images can provide surgeons with important information for performing accurate minimally invasive surgeries. However, it is difficult to estimate the absolute depth of the scene based on monocular endoscope. Depth values in endoscopic images change drastically during the operation, which make it hard to estimate them with a fixed baseline. In this paper, we propose a depth estimation scheme with multiple baselines. The monocular endoscope is moved horizontally by a robotic endoscope holder to generate stereo images. A pixel-level depth map fusion algorithm is designed to combine depth values estimated with different baselines. Experimental results show that the proposed method improves the accuracy of depth estimation and the visual quality of depth maps. Zhidong Tan, Rihui Song, Kai Huang 0001 |
ICIP | 3 |
| 2023 | Selective Frequency Network for Image Restoration
Yuning Cui 0001, Zhenshan Bing, Wenqi Ren, Xinwei Gao, Xiaochun Cao, Kai Huang 0001, Alois C. Knoll |
ICLR | 7 |
| 2023 | GFNet: Gaze Focus Network using Attention for Gaze EstimationabstractGaze estimation can be applied to human visual attention understanding. The current methods mainly obtain gaze mapping from facial or eye images, and most of them only focus on gaze point or gaze direction estimation. In this paper, we propose a multitask gaze focus network for gaze point and gaze direction estimation. Focus attention layer is used to guide the generation of facial features. By connecting eye and face features, feature similarity is used to get attention weights, and make it tend to eyes position. We propose four loss functions to constrain the network in 2D and 3D spaces. The combination of eye position constraint and focus attention layer ensures the accuracy of gaze point estimation. Gaze focus is used to obtain gaze depth. Through comprehensive experiments, the advantages of proposed method in gaze tracking are verified. In addition, the application prospect of proposed method in depth-overlapping is proved. Daosong Hu, Kai Huang 0001 |
ICME | 2 |
| 2023 | Real-Time Instance Segmentation and Tip Detection for Neuroendoscopic Surgical Instruments
Rihui Song, Silu Guo, Yehua Ling, Jin Gong, Kai Huang 0001 |
ICONIP (10) | 6 |
| 2023 | Meta-Reinforcement Learning via Language InstructionsabstractAlthough deep reinforcement learning has recently been very successful at learning complex behaviors, it requires a tremendous amount of data to learn a task. One of the fundamental reasons causing this limitation lies in the nature of the trial-and-error learning paradigm of reinforcement learning, where the agent communicates with the environment and pro-gresses in the learning only relying on the reward signal. This is implicit and rather insufficient to learn a task well. On the con-trary, humans are usually taught new skills via natural language instructions. Utilizing language instructions for robotic motion control to improve the adaptability is a recently emerged topic and challenging. In this paper, we present a meta-RL algorithm that addresses the challenge of learning skills with language instructions in multiple manipulation tasks. On the one hand, our algorithm utilizes the language instructions to shape its in-terpretation of the task, on the other hand, it still learns to solve task in a trial-and-error process. We evaluate our algorithm on the robotic manipulation benchmark (Meta-World) and it significantly outperforms state-of-the-art methods in terms of training and testing task success rates. Codes are available at https://tumi6robot.wixsite.com/million. Zhenshan Bing, Alexander W. Koch, Xiangtong Yao, Kai Huang 0001, Alois C. Knoll |
ICRA | 4 |
| 2023 | Smooth Stride Length Change of Rat Robot with a Compliant Actuated Spine Based on CPG ControllerabstractThe aim of this research is to investigate the relationship between spinal flexion and quadruped locomotion in a rat robot equipped with a compliant spine, controlled by a central pattern generator (CPG). The study reveals that spinal flexion can enhance limb stride length, but it may also cause significant and unexpected motion disturbances during stride length variations. To address this issue, this paper proposes a CPG model driven by spinal flexion and a novel oscillator that incorporates a circular limit cycle and accounts for the anticipated stride length transition process. This approach effectively matches the torque change with the dynamics of stride length changes, leading to lower energy consumption. Extensive simulations are conducted to evaluate the efficacy of the proposed oscillator and compare it with the original kinetic model and other CPG models. The results demonstrate that the designed CPG model with the proposed oscillator yields smoother gait transitions during stride length variations and reduces energy consumption. Yuhong Huang, Zhenshan Bing, Zitao Zhang, Kai Huang 0001, Fabrice O. Morin, Alois C. Knoll |
IROS | 4 |
| 2023 | Learning from Symmetry: Meta-Reinforcement Learning with Symmetrical Behaviors and Language InstructionsabstractMeta-reinforcement learning (meta-RL) is a promising approach that enables the agent to learn new tasks quickly. However, most meta-RL algorithms show poor generalization in multi-task scenarios due to the insufficient task information provided only by rewards. Language-conditioned meta-RL improves the generalization capability by matching language instructions with the agent's behaviors. While both behaviors and language instructions have symmetry, which can speed up human learning of new knowledge. Thus, combining symmetry and language instructions into meta-RL can help improve the algorithm's generalization and learning efficiency. We propose a dual-MDP meta-reinforcement learning method that enables learning new tasks efficiently with symmetrical behav-iors and language instructions. We evaluate our method in mul-tiple challenging manipulation tasks, and experimental results show that our method can greatly improve the generalization and learning efficiency of meta-reinforcement learning. Videos are available at https://tumi6robot.wixsite.com/symmetry/. Xiangtong Yao, Zhenshan Bing, Genghang Zhuang, Kejia Chen 0005, Kai Huang 0001, Alois C. Knoll |
IROS | 6 |
| 2023 | An Energy-Efficient Lane-Keeping System Using 3D LiDAR Based on Spiking Neural NetworkabstractLane keeping, as a fundamental functionality of autonomous navigation, remains a challenging task for autonomous robots and vehicles. Recently, spiking neural networks (SNNs) have gained attention and research interest due to their biological plausibility and application potential on neuromorphic processors. SNNs have also been successfully deployed on robots to solve autonomous navigation problems. However, lane keeping with a LiDAR sensor is still an open problem for SNNs. In this work, we propose an end-to-end approach based on an SNN to solve the lane-keeping problem using a 3D LiDAR sensor. For the first time, we explore the capability of the proposed SNN controller to perceive the LiDAR input and exploit the features to perform reward-based feedback learning. To ensure the effectiveness of the controller, the proposed method is deployed and evaluated on two high-fidelity simulators. The experimental results demonstrate the high applicability and performance in different scenarios. Furthermore, experiments show that the SNN is capable of performing lane keeping in a simulated urban environment with only 18 control neurons and 32 synapse connections, producing on average only a 17cm deviation from lane center, which is 4.3 % of the lane width. Genghang Zhuang, Zhenshan Bing, Xiangtong Yao, Yuhong Huang, Kai Huang 0001, Alois C. Knoll |
IROS | 6 |
| 2023 | Label-Preserving Data Augmentation in Latent Space for Diabetic Retinopathy Recognition
Junjie Yang 0001, Shahrooz Faghih Roohi, Kai Huang 0001, Mathias Maier, Nassir Navab, M. Ali Nasseri |
MICCAI (3) | 4 |
| 2023 | Event-Diffusion: Event-Based Image Reconstruction and Restoration with Diffusion ModelsabstractEvent cameras offer the advantages of low latency, high temporal resolution and HDR compared to conventional cameras. Due to the asynchronous and sparse nature of events, many existing algorithms cannot be directly applied, necessitating the reconstruction of intensity frames. However, existing reconstruction methods often result in artifacts and edge blurring due to noise and event accumulation. In this paper, we argue that the key to event-based image reconstruction is to enhance the edge information of objects and restore the artifacts in the reconstructed images. To explain, edge information is one of the most important features in the event stream, providing information on the shape and contour of objects. Considering the extraordinary capabilities of Denoising Diffusion Probabilistic Models (DDPMs) in image generation, reconstruction, and restoration, we propose a new framework which incorporate it into the reconstruction pipeline to obtain high-quality results which effectively remove artifacts and blur in reconstructed images. Specifically, we first extract edge information from the event stream using the proposed event-based denoising method. It employs the contrast maximization framework to remove noise from the event stream and extract clear object edge information. And then, the edge information is further adopted to our diffusion model, which is used to enhance the edges of objects in the reconstructed images, thus improving the restoration effect. Experimental results show that our method achieves significant improvements in the mean squared error (MSE), the structural similarity (SSIM), and the perceptual similarity (LPIPS) metrics, with average improvements of 40%, 15%, and 25%, respectively, compared to previous state-of-the-art models, and has good generalization performance. Quanmin Liang, Xiawu Zheng, Kai Huang 0001, Yan Zhang 0109, Jie Chen 0001, Yonghong Tian 0001 |
ACM Multimedia | 3 |
| 2023 | Dense Depth Estimation for Monocular Endoscope Robot with an Adaptive BaselineabstractDepth information is useful to surgeons and surgical assistance systems. However, it is a challenging task to estimate the depth of various surgical scenes based on a monocular endoscope. We propose a depth estimation approach for a monocular endoscope with a stereo matching algorithm. The monocular endoscope is moved horizontally by a robotic endoscope holder to simulate a stereo vision system. The main challenge is how to obtain a proper baseline for better depth information generation as the depth range of a surgical scene is unknown beforehand. We design a baseline evaluation and selection algorithm to search for suitable baselines for surgical scenes with different depth ranges. Experimental results show that our approach improves the average accuracy of different depth scenarios by 10.8% when the error range is 2mm. Rihui Song, Zhidong Tan, Hongli Liang, Yehua Ling, Gang Chen 0023, Kai Huang 0001, Jin Gong |
SMC | 6 |
| 2023 | An efficient real-time accelerator for high-accuracy DNN-based optical flow estimation in FPGA
Yuanxing Yan, Yehua Ling, Kai Huang 0001, Gang Chen 0023 |
J. Syst. Archit. | 3 |
| 2023 | FTSC: Fault-tolerant scheduling and control co-design for distributed real-time system
Yuanhai Zhang, Zijin Xu, Nan Guan, Shuai Zhao 0004, Gang Chen 0023, Kai Huang 0001 |
J. Syst. Archit. | 7 |
| 2023 | Meta-Reinforcement Learning in Non-Stationary and Dynamic EnvironmentsabstractIn recent years, the subject of deep reinforcement learning (DRL) has developed very rapidly, and is now applied in various fields, such as decision making and control tasks. However, artificial agents trained with RL algorithms require great amounts of training data, unlike humans that are able to learn new skills from very few examples. The concept of meta-reinforcement learning (meta-RL) has been recently proposed to enable agents to learn similar but new skills from a small amount of experience by leveraging a set of tasks with a shared structure. Due to the task representation learning strategy with few-shot adaptation, most recent work is limited to narrow task distributions and stationary environments, where tasks do not change within episodes. In this work, we address those limitations and introduce a training strategy that is applicable to non-stationary environments, as well as a task representation based on Gaussian mixture models to model clustered task distributions. We evaluate our method on several continuous robotic control benchmarks. Compared with state-of-the-art literature that is only applicable to stationary environments with few-shot adaption, our algorithm first achieves competitive asymptotic performance and superior sample efficiency in stationary environments with zero-shot adaption. Second, our algorithm learns to perform successfully in non-stationary settings as well as a continual learning setting, while learning well-structured task representations. Last, our algorithm learns basic distinct behaviors and well-structured task representations in task distributions with multiple qualitatively distinct tasks. Zhenshan Bing, David Lerch, Kai Huang 0001, Alois C. Knoll |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2023 | Toward Intelligent Sensing: Optimizing Lidar Beam Distribution for Autonomous DrivingabstractLiDAR (Light Detection And Ranging) sensors have been widely used in autonomous vehicles as the main sensors. According to the specification details of the widely used 3D LiDAR products in the market, the distribution of vertical beam channels is set according to a uniform angular resolution, which is not ideally efficient for specific autonomous tasks. In this paper, we propose a novel approach to find the optimized angular distribution of the vertical beam channels for different application scenarios and installation configurations. The experimental results in a study case suggest that concerning the vehicle detection task, the optimized LiDARs perform almost two times better than the ones with the same number of channels in terms of the detection range, and have perception performances close to the LiDARs with double channels in the long distance. Genghang Zhuang, Zhenshan Bing, Xiangtong Yao, Yuhong Huang, Kai Huang 0001, Alois C. Knoll |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2023 | Robotic Manipulation in Dynamic Scenarios via Bounding-Box-Based Hindsight Goal GenerationabstractBy relabeling past experience with heuristic or curriculum goals, state-of-the-art reinforcement learning (RL) algorithms such as hindsight experience replay (HER), hindsight goal generation (HGG), and graph-based HGG (G-HGG) have been able to solve challenging robotic manipulation tasks in multigoal settings with sparse rewards. HGG outperforms HER in challenging tasks in which goals are difficult to explore by learning from a curriculum, in which intermediate goals are selected based on the Euclidean distance to target goals. G-HGG enhances HGG by selecting intermediate goals from a precomputed graph representation of the environment, which enables its applicability in an environment with stationary obstacles. However, G-HGG is not applicable to manipulation tasks with dynamic obstacles, since its graph representation is only valid in static scenarios and fails to provide any correct information to guide the exploration. In this article, we propose bounding-box-based HGG (Bbox-HGG), an extension of G-HGG selecting hindsight goals with the help of image observations of the environment, which makes it applicable to tasks with dynamic obstacles. We evaluate Bbox-HGG on four challenging manipulation tasks, where significant enhancements in both sample efficiency and overall success rate are shown over state-of-the-art algorithms. The videos can be viewed at https://videoviewsite.wixsite.com/bbhgg. Zhenshan Bing, Erick Álvarez, Long Cheng 0007, Fabrice O. Morin, Rui Li 0077, Xiaojie Su, Kai Huang 0001, Alois C. Knoll |
IEEE Trans. Neural Networks Learn. Syst. | 7 |
| 2022 | FlowAcc: Real-Time High-Accuracy DNN-based Optical Flow Accelerator in FPGAabstractRecently, accelerator architectures have been designed to use deep neural networks (DNNs) to accelerate computer vision tasks, possessing the advantages of both accuracy and speed. Optical flow accelerator is however not among these architectures that DNNs have been successfully deployed. Existing hardware accelerators for optical flow estimation are all designed for classic methods and generally perform poorly in estimated accuracy. In this paper, we present FlowAcc, a dedicated hardware accelerator for DNN-based optical flow estimation, adopting a pipelined hardware design for real-time processing of image streams. We design an efficient multiplexing binary neural network (BNN) architecture for pyramidal feature extraction to significantly reduce the hardware cost and make it independent of the pyramid level number. Furthermore, efficient hamming distance calculation and competent flow regularization are utilized for hierarchical optical flow estimation to greatly improve the system efficiency. Comprehensive experimental results demonstrate that FlowAcc achieves state-of-the-art estimation accuracy and real-time performance on the Middlebury dataset when compared with the existing optical flow accelerators. Yehua Ling, Yuanxing Yan, Kai Huang 0001, Gang Chen 0023 |
DATE | 3 |
| 2022 | Ultra-Flow: An Ultra-fast and High-quality Optical Flow Accelerator with Deep Feature Matching on FPGAabstractDense and accurate optical flow estimation is an important requirement for dynamic scene perception in autonomous systems. However, most of the existing FPGA accelerators are based on classic methods, which cannot deal with large displacements of moving objects in ultra-fast scenes. In this paper, we present Ultra-Flow, an ultra-fast pipelined architecture for efficient optical flow estimation and refinement. Ultra-Flow utilizes binary neural networks to generate the robust feature map, on which hierarchical matching is directly performed. Therefore, multiple usages of neural networks at hierarchical levels can be avoided to achieve hardware efficiency in Ultra-Flow. Optimizations, including local flow regularization and enhanced matching, are further used to improve the throughput and refine the optical flow to obtain higher accuracy. Evaluation results show that, compared to state-of-the-art FPGA accelerators, Ultra-Flow achieves leading accuracy in the Middlebury sequences at ultra-fast processing speed up to 687.92 frames/s for 640 × 480 pixel images. Yehua Ling, Yuanxing Yan, Kai Huang 0001, Gang Chen 0023 |
FPL | 3 |
| 2022 | De-snowing LiDAR Point Clouds With Intensity and Spatial-Temporal FeaturesabstractPoint clouds from 3D light detection and ranging (LiDAR) are widely used. Noise caused by falling snow reduces the availability of point clouds. Due to the sparseness of LiDAR point clouds and the fact that the snow point clouds are easily affected by multi factors such as wind or snowfall conditions, it is difficult to accurately remove the snow while preserving the details of the point clouds. To solve the problem, this paper presents a de-snowing approach combining the intensity and spatial-temporal features. An intensity-based filter firstly removes the snow. Then a repairing method restores the non-snow points based on the spatial-temporal features. Experimental results demonstrate that our approach outperforms existing work in the literature and performs the least damage to the point clouds in different snowfall scenarios. Boyang Li 0009, Jieling Li, Gang Chen 0023, Hejun Wu, Kai Huang 0001 |
ICRA | 5 |
| 2022 | ColibriDoc: an Eye-in-Hand Autonomous Trocar Docking SystemabstractRetinal surgery is a complex medical procedure that requires exceptional expertise and dexterity. For this purpose, several robotic platforms are currently under development to enable or improve the outcome of microsurgical tasks. Since the control of such robots is often designed for navigation inside the eye in proximity to the retina, successful trocar docking and insertion of the instrument into the eye represents an additional cognitive effort, and is therefore one of the open challenges in robotic retinal surgery. For this purpose, we present a platform for autonomous trocar docking that combines computer vision and a robotic setup. Inspired by the Cuban Colibri (hummingbird) aligning its beak to a flower using only vision, we mount a camera onto the endeffector of a robotic system. By estimating the position and pose of the trocar, the robot is able to autonomously align and navigate the instrument towards the Trocar Entry Point (TEP) and finally perform the insertion. Our experiments show that the proposed method is able to accurately estimate the position and pose of the trocar and achieve repeatable autonomous docking. The aim of this work is to reduce the complexity of the robotic setup prior to the surgical task and therefore, increase the intuitiveness of the system integration into clinical workflow. Shervin Dehghani, Michael Sommersperger, Junjie Yang 0001, Mehrdad Salehi, Benjamin Busam, Kai Huang 0001, Peter Gehlbach, Iulian Iordachita, Nassir Navab, M. Ali Nasseri |
ICRA | 6 |
| 2022 | Enhanced Quadruped Locomotion of a Rat Robot Based on the Lateral Flexion of a Soft Actuated SpineabstractIn nature, the movement of quadrupeds is completed under the combined action of the spine and the legs. Inspired by this, this paper explores the effect of a lateral flexing spine on the locomotion of a rat robot. Benefiting from the regular lateral flexion of a soft actuated spine, the rat robot exhibits enhance step length of its hind legs and increased translational velocity by coordinating the opposite movements of the left and right sides. Furthermore, this paper introduces a mathematical model of the effect of the flexible spine on the robot velocity. Finally, extensive experiments are conducted in simulations and on the physical rat robot. Compared with the locomotion without a flexing spine, the simulation results show that the velocity of the robot can be increased up to 218.29%, which is in line with the theoretical results from the proposed mathematical model. Limited by the gap between simulation and the real world, the experiment results of the physical rat robot show a slight performance than the theoretical results. But the physical rat robot can still enhance its translational velocity with the help of a lateral flexing spine. Yuhong Huang, Zhenshan Bing, Florian Walter, Alex Rohregger, Zitao Zhang, Kai Huang 0001, Fabrice O. Morin, Alois C. Knoll |
IROS | 6 |
| 2022 | A Biologically-Inspired Simultaneous Localization and Mapping System Based on LiDAR SensorabstractSimultaneous localization and mapping (SLAM) is one of the essential techniques and functionalities used by robots to perform autonomous navigation tasks. Inspired by the rodent hippocampus, this paper presents a biologically inspired SLAM system based on a LiDAR sensor using a hippocampal model to build a cognitive map and estimate the robot pose in indoor environments. Based on the biologically inspired models mimicking boundary cells, place cells, and head direction cells, the SLAM system using LiDAR point cloud data is capable of leveraging the self-motion cues from the LiDAR odometry and the boundary cues from the LiDAR boundary cells to build a cognitive map and estimate the robot pose. Experiment results show that with the LiDAR boundary cells the proposed SLAM system greatly outperforms the camera-based brain-inspired method in both simulation and indoor environments, and is competitive with the conventional LiDAR-based SLAM methods. Genghang Zhuang, Zhenshan Bing, Yuhong Huang, Kai Huang 0001, Alois C. Knoll |
IROS | 4 |
| 2022 | A Biologically-Inspired Global Localization System for Mobile Robots Using LiDAR SensorabstractLocalization in the environment is an essential navigational capability for animals and indoor robotic vehicles. In indoor environments, it is still challenging to perfectly solve the global localization problem using probabilistic methods. However, animals are able to instinctively localize themselves with much less effort. Therefore, an intriguing and promising approach is to seek biological inspiration from animals. In this paper, we present a biologically-inspired global localization system using a LiDAR sensor that utilizes a hippocampal model and a landmark-based relocalization approach. The experiment results show that the proposed method is competitive with Monte Carlo Localization, and the results demonstrate the high accuracy, applicability, and reliability of the proposed biologically-inspired localization system in various localization scenarios. Genghang Zhuang, Carlo Cagnetta, Zhenshan Bing, Hu Cao, Kai Huang 0001, Alois C. Knoll |
IV | 6 |
| 2022 | Lite-Stereo: A Resource-Efficient Hardware Accelerator for Real-Time High-Quality Stereo Estimation Using Binary Neural NetworkabstractStereo estimation plays a key role in many autonomous systems, such as robotics and self-driving cars. Recent work on StereoEngine, an FPGA-based accelerator for deep neural network (DNN)-based stereo estimation, has been demonstrated as a promising solution to achieve both real-time and high accuracy performance for depth sensing. However, this solution still suffers from over-utilizing the hardware resource of FPGAs. In this article, we present Lite-Stereo, a resource-efficient DNN-based stereo vision accelerator to improve the hardware efficiency for StereoEngine running on a resource-constrained FPGA. To achieve this, we design a set of optimized hardware architectures for resource-demanding bottleneck modules. In order to balance the gap between the processing speed and resource efficiency, the process elements in binary neural network modules are shared within and across modules. In addition, we provide reusing strategies on path aggregation and neighbor calculation to improve the resource efficiency of the semi-global matching module. Evaluation results demonstrate that Lite-Stereo reduces the hardware cost of ALUTs and RAM bits by 60% and 29%, respectively, without compromising the accuracy and energy efficiency compared with StereoEngine. Yehua Ling, Haitao Meng, Kai Huang 0001, Gang Chen 0023 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2022 | Complex Robotic Manipulation via Graph-Based Hindsight Goal GenerationabstractReinforcement learning algorithms, such as hindsight experience replay (HER) and hindsight goal generation (HGG), have been able to solve challenging robotic manipulation tasks in multigoal settings with sparse rewards. HER achieves its training success through hindsight replays of past experience with heuristic goals but underperforms in challenging tasks in which goals are difficult to explore. HGG enhances HER by selecting intermediate goals that are easy to achieve in the short term and promising to lead to target goals in the long term. This guided exploration makes HGG applicable to tasks in which target goals are far away from the object's initial position. However, the vanilla HGG is not applicable to manipulation tasks with obstacles because the Euclidean metric used for HGG is not an accurate distance metric in such an environment. Although, with the guidance of a handcrafted distance grid, grid-based HGG can solve manipulation tasks with obstacles, a more feasible method that can solve such tasks automatically is still in demand. In this article, we propose graph-based hindsight goal generation (G-HGG), an extension of HGG selecting hindsight goals based on shortest distances in an obstacle-avoiding graph, which is a discrete representation of the environment. We evaluated G-HGG on four challenging manipulation tasks with obstacles, where significant enhancements in both sample efficiency and overall success rate are shown over HGG and HER. Videos can be viewed at https://videoviewsite.wixsite.com/ghgg. Zhenshan Bing, Matthias Brucker, Fabrice O. Morin, Rui Li 0077, Xiaojie Su, Kai Huang 0001, Alois C. Knoll |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2022 | Toward Cognitive Navigation: Design and Implementation of a Biologically Inspired Head Direction Cell NetworkabstractAs a vital cognitive function of animals, the navigation skill is first built on the accurate perception of the directional heading in the environment. Head direction cells (HDCs), found in the limbic system of animals, are proven to play an important role in identifying the directional heading allocentrically in the horizontal plane, independent of the animal's location and the ambient conditions of the environment. However, practical HDC models that can be implemented in robotic applications are rarely investigated, especially those that are biologically plausible and yet applicable to the real world. In this article, we propose a computational HDC network that is consistent with several neurophysiological findings concerning biological HDCs and then implement it in robotic navigation tasks. The HDC network keeps a representation of the directional heading only relying on the angular velocity as an input. We examine the proposed HDC model in extensive simulations and real-world experiments and demonstrate its excellent performance in terms of accuracy and real-time capability. Zhenshan Bing, Amir E. I. Sewisy, Genghang Zhuang, Florian Walter, Fabrice O. Morin, Kai Huang 0001, Alois C. Knoll |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2021 | Understanding the Property of Long Term Memory for the LSTM with Attention MechanismabstractRecent trends of incorporating LSTM network with different attention mechanisms in time series forecasting have led researchers to consider the attention module as an essential component. While existing studies revealed the effectiveness of attention mechanism with some visualization experiments, the underlying rationale behind their outstanding performance on learning long-term dependencies remains hitherto obscure. In this paper, we aim to elaborate on this fundamental question by conducting a thorough investigation of the memory property for LSTM network with attention mechanism. We present a theoretical analysis of LSTM integrated with attention mechanism, and demonstrate that it is capable of generating an adaptive decay rate which dynamically controls the memory decay according to the obtained attention score. In particular, our theory shows that attention mechanism brings significantly slower decays than the exponential decay rate of a standard LSTM. Experimental results on four real-world time series datasets demonstrate the superiority of the attention mechanism for maintaining long-term memory when compared to the state-of-the-art methods, and further corroborate our theoretical analysis. Wendong Zheng, Putian Zhao, Kai Huang 0001, Gang Chen 0023 |
CIKM | 3 |
| 2021 | Peak temperature analysis and optimization for pipelined hard real-time systems
Long Cheng 0007, Kai Huang 0001, Liang Mi, Gang Chen 0023, Alois C. Knoll |
Inf. Sci. | 2 |
| 2021 | An efficient GPU-accelerated inference engine for binary neural network on mobile phones
Shengyu He, Haitao Meng, Zhaoheng Zhou, Kai Huang 0001, Gang Chen 0023 |
J. Syst. Archit. | 5 |
| 2021 | Efficient runtime slack management for EDF-VD-based mixed-criticality scheduling
Junjie Yang 0001, Guangyi Xu 0004, Gang Chen 0023, Nan Guan, Kai Huang 0001 |
J. Syst. Archit. | 5 |
| 2021 | SiamFPN: A Deep Learning Method for Accurate and Real-Time Maritime Ship TrackingabstractVisual object tracking plays an essential role in various maritime applications. However, most of the existing tracking methods belong to generative models, which only focus on the features of the object and require the target has significant visual saliency for accurate tracking. While the visual saliency is available in most of the common tracking conditions, these methods may fail when facing challenging situations. In this paper, a deep learning based tracking method is proposed to track maritime ships, namely, SiamFPN. In SiamFPN, a modified Siamese Network is combined with multi-RPNs to build a tracking pipeline. Concretely, A ResNet-50 with an FPN structure is used as the CNN of the detection subnetwork of Siamese, and a template subnetwork is parallel to the detection. In order to strengthen the discriminative ability, three RPNs are deployed to process the output of Siamese Network. Moreover, a historical impacts based proposal selection method is developed for selecting correct target areas. Finally, a dataset is collected for training and testing SiamFPN and validating our excellent performance over the other four recent SOTA trackers. Based on the experimental results, we achieved 74 % on average accuracy with real-time speed. Yunxiao Shan, Xiaomei Zhou, Shanghua Liu, Kai Huang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2020 | PhoneBit: Efficient GPU-Accelerated Binary Neural Network Inference Engine for Mobile PhonesabstractOver the last years, a great success of deep neural networks (DNNs) has been witnessed in computer vision and other fields. However, performance and power constraints make it still challenging to deploy DNNs on mobile devices due to their high computational complexity. Binary neural networks (BNNs) have been demonstrated as a promising solution to achieve this goal by using bit-wise operations to replace most arithmetic operations. Currently, existing GPU-accelerated implementations of BNNs are only tailored for desktop platforms. Due to architecture differences, mere porting of such implementations to mobile devices yields suboptimal performance or is impossible in some cases. In this paper, we propose PhoneBit, a GPU-accelerated BNN inference engine for Android-based mobile devices that fully exploits the computing power of BNNs on mobile GPUs. PhoneBit provides a set of operator-level optimizations including locality-friendly data layout, bit packing with vectorization and layers integration for efficient binary convolution. We also provide a detailed implementation and parallelization optimization for PhoneBit to optimally utilize the memory bandwidth and computing power of mobile GPUs. We evaluate PhoneBit with AlexNet, YOLOv2 Tiny and VGG16 with their binary version. Our experiment results show that PhoneBit can achieve significant speedup and energy efficiency compared with state-of-the-art frameworks for mobile devices. Gang Chen 0023, Shengyu He, Haitao Meng, Kai Huang 0001 |
DATE | 4 |
| 2020 | Offline Practising and Runtime Training Framework for Autonomous Motion Control of Snake RobotsabstractThis paper proposes an offline and runtime combined framework for the autonomous motion of snake robots. With the dynamic feedback of its state during runtime, the robot utilizes the linear regression to update its control parameters for better performance and thus adaptively reacts to the environment. To reduce interference from infeasible samples and improve efficiency, the data set for runtime training is chosen from one in several clusters categorized from samples collected in offline practice. Moreover, only the most sensitive control parameter is updated at one iteration for better robustness and efficiency. The effectiveness and efficiency of our approach are evaluated by a set of case studies of pole climbing. Experimental results demonstrate that with the proposed framework, the snake robot can adapt its locomotion gait to poles with different unknown diameters. Long Cheng 0007, Zhiyong Jian, Yuhong Huang, Kai Huang 0001 |
ICRA | 6 |
| 2020 | 3D Object Detection and Tracking Based on Streaming DataabstractRecent approaches for 3D object detection have made tremendous progresses due to the development of deep learning. However, previous researches are mostly based on individual frames, leading to limited exploitation of information between frames. In this paper, we attempt to leverage the temporal information in streaming data and explore 3D streaming based object detection as well as tracking. Toward this goal, we set up a dual-way network for 3D object detection based on keyframes, and then propagate predictions to non-key frames through a motion based interpolation algorithm guided by temporal information. Our framework is not only shown to have significant improvements on object detection compared with frame-by-frame paradigm, but also proven to produce competitive results on KITTI Object Tracking Benchmark, with 76.68% in MOTA and 81.65% in MOTP respectively. Xusen Guo, Jianfeng Gu 0001, Silu Guo, Zixiao Xu, Chengzhang Yang, Shanghua Liu, Long Cheng 0007, Kai Huang 0001 |
ICRA | 8 |
| 2020 | Microscope-Guided Autonomous Clear Corneal IncisionabstractClear Corneal Incision, a challenging step in cataract surgery, and important to the overall quality of the surgery. New surgeons usually spend one full year trying to perfect their incision, but even after such rigorous training deficient incisions can still occur. This paper proposes an autonomous robotic system for this self-sealing incision. A conventional ophthalmic microscope system with a monocular camera is utilized to capture the surgical scene, ascertain the robot's position, and estimate depth information. Kinematics with a remote centre of motion (RCM) is designed for a multi-axes robot to perform the incision route. The experimental results on ex-vivo porcine eyes show the autonomous Clear Corneal Incision has a stricter three-plane structure than a surgeon-made incision, which is closer to the ideal incision. Sean J. Bergunder, Duoru Lin, Shengzhi Lin, M. Ali Nasseri, Mingchuan Zhou, Haotian Lin 0001, Kai Huang 0001 |
ICRA | 9 |
| 2020 | Target Tracking Control of a Wheel-less Snake Robot Based on a Supervised Multi-layered SNNabstractThe snake-like robot without wheels is a bio-inspired robot whose high degree of freedom results in a challenge in autonomous locomotion control. The use of a Spiking Neural Network (SNN) which is a biologically plausible artificial neural network can help to achieve the autonomous locomotion behavior of snake robots in an energy-efficient manner. Approaches that use an SNN without hidden layers have been applied in the single-target tracking task. However, due to the complexity of the 3D gaits on a wheel-less snake robot and the imprecision of the pose control while in motion, they have some fluctuation that adversely affects their performances. In this work, we design two multi-layered SNNs with different topology for a wheel-less snake robot to track a certain moving object. The visual signals obtained from a Dynamic Vision Sensor (DVS) are fed into the SNN to drive the locomotion controller. Furthermore, the Reward-modulated Spike-Timing-Dependent Plasticity (R-STDP) learning rule is utilized to train the SNN end-to-end. Compared to the SNN without hidden layers, the proposed multi-layered SNN with a separated hidden layer shows its advantage in terms of robustness. Zhuangyi Jiang, Richard Otto, Zhenshan Bing, Kai Huang 0001, Alois C. Knoll |
IROS | 4 |
| 2020 | Offloading Autonomous Driving Services via Edge ComputingabstractA key challenge for autonomous driving is to process a massive amount of sensor data and make safe and reliable decisions in real time. However, autonomous vehicles often have insufficient onboard resources to provide the required computation capacity. To address this problem, this article advocates a novel approach to offload computation-intensive autonomous driving services to roadside units and cloud for swift executions. Our approach combines an integer linear programming (ILP) formulation for offline optimization of the scheduling strategy and a fast heuristics algorithm for online adaptation. We verify our technique with both synthetic task graphs and real-world deployment. The experimental results show that our approach can improve system performance effectively. Mingyue Cui, Shipeng Zhong, Boyang Li 0009, Xu Chen 0004, Kai Huang 0001 |
IEEE Internet Things J. | 5 |
| 2020 | Fault-tolerant real-time tasks scheduling with dynamic fault handling
Gang Chen 0023, Nan Guan, Kai Huang 0001, Wang Yi 0001 |
J. Syst. Archit. | 3 |
| 2020 | Energy-efficient and damage-recovery slithering gait design for a snake-like robot based on reinforcement learning and inverse reinforcement learning
Zhenshan Bing, Christian Lemke, Long Cheng 0007, Kai Huang 0001, Alois C. Knoll |
Neural Networks | 4 |
| 2020 | Indirect and direct training of spiking neural networks for end-to-end control of a lane-keeping vehicle
Zhenshan Bing, Claus Meschede, Guang Chen 0001, Alois C. Knoll, Kai Huang 0001 |
Neural Networks | 5 |
| 2020 | StereoEngine: An FPGA-Based Accelerator for Real-Time High-Quality Stereo Estimation With Binary Neural NetworkabstractStereo estimation is essential to many applications such as mobile autonomous robots, most of which ask for real-time response, high energy, and storage efficiency. Deep neural networks (DNNs) have shown to yield significant gains in improving accuracy. However, these DNN-based algorithms are challenging to be deployed on energy and resource-constrained devices due to the high computational complexities of DNNs. In this article, we present StereoEngine, a fully pipelined end-to-end stereo vision accelerator that computes accurate dense depth in a real-time and energy-efficient manner. An efficient stereo algorithm is developed and optimized for a high-quality hardware-friendly implementation, that leverages binary neural network (BNN) to learn discriminative binary descriptors to improve the disparity. The design of StereoEngine is a standalone DNN-based stereo vision system where all processing procedures are implemented on a hardware platform. The effectiveness of StereoEngine is evaluated by comprehensive experiments. Compared with software-based implementations on the highend and embedded Nvidia GPUs, StereoEngine achieves up to 3×, 13×, and 50× speedups, as well as up to 211×, 58×, and 73× energy efficiency improvement, respectively. Furthermore, StereoEngine achieves leading accuracy when compared to state-of-the-art hardware implementations on the challenging KITTI dataset. Gang Chen 0023, Yehua Ling, Haitao Meng, Shengyu He, Kai Huang 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 7 |
| 2020 | GPU-Accelerated Real-Time Stereo Estimation With Binary Neural NetworkabstractDepth estimation from stereo images is essential to many applications such as robotics and autonomous vehicles, most of which ask for the real-time response, high energy and storage efficiency. Recent work has shown deep neural networks (DNN) perform extremely well for stereo estimation. However, these state-of-the-art DNN based algorithms are challenging to be deployed into real-world applications due to the high computational complexities of DNNs. Most of them are too slow for real-time inference and require several seconds of GPU computation to process image frames. In this article, we address the problem of fast stereo estimation and propose an efficient and light-weighted stereo matching system, called StereoBit, to produce a disparity map in a real-time manner while achieving close to state-of-the-art accuracy. To achieve this goal, we propose a binary neural network to generate weighted Hamming distance for an efficient similarity join in stereo estimation. In addition, we propose a novel approximation approach to derive StereoBit network directly from the well-trained network with the cosine similarity. Our approximation strategies enable a significant speedup while maintaining almost the same accuracy compared to the network with the cosine similarity. Furthermore, we present an optimization framework for fully exploiting the computing power of StereoBit. The framework provides a significant speedup of stereo estimation routines, and at the same time, reduces the memory usage for storing parameters. The effectiveness of StereoBit is evaluated by comprehensive experiments. StereoBit can achieve 60 frames per second on an NVIDIA TITAN Xp GPU on KITTI 2012 benchmark while achieving 3-pixel non-occluded stereo error 3.56 percent. Gang Chen 0023, Haitao Meng, Yucheng Liang, Kai Huang 0001 |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2020 | Efficient Performance Estimation and Work-Group Size Pruning for OpenCL Kernels on GPUsabstractGraphic Processing Units (GPUs) play a vital role in state-of-the-art high-performance scientific computing realm and research work towards its performance analysis is crucial but nontrivial. Extant GPU performance models are far from practical use, while fine-grained GPU simulation requires a considerably large time cost. Moreover, massive amounts of designs with various program inputs and parameter settings pose a challenge for efficient performance estimation and tuning of parallel GPU applications. To this end, this article presents a hybrid framework for the efficient performance estimation and work-group size pruning of OpenCL workloads on GPUs. The framework contains a static module used to extract the kernel execution trace from the high-level source code and a dynamical module used to mimic the kernel execution flow to estimate the runtime performance. For the design space pruning, an extra analysis is performed to filter out the redundant work-group sizes with duplicated execution traces and inferior pipelines. The proposed framework does not require any program runs to estimate the performance and find the optimal or near-optimal designs. Experiments on four Commercial Off-The-Shelf (COTS) Nvidia GPUs show that the framework can predict the runtime performance with an average error of 17.04 percent and reduce the program design space by an average of 78.47 percent. Xiebing Wang, Xuehai Qian, Alois C. Knoll, Kai Huang 0001 |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2019 | A Hybrid Framework for Fast and Accurate GPU Performance Estimation through Source-Level Analysis and Trace-Based SimulationabstractThis paper proposes a hybrid framework for fast and accurate performance estimation of OpenCL kernels running on GPUs. The kernel execution flow is statically analyzed and thereupon the execution trace is generated via a loop-based bidirectional branch search. Then the trace is dynamically simulated to perform a dummy execution of the kernel to obtain the estimated time. The framework does not rely on profiling or measurement results which are used in conventional performance estimation techniques. Moreover, the lightweight trace-based simulation consumes much less time than a fine-grained GPU simulator. Our framework can accurately grasp the variation trend of the execution time in the design space and robustly predict the performance of the kernels across two generations of recent Nvidia GPU architectures. Experiments on four Commercial Off-The-Shelf (COTS) GPUs show that our framework can predict the runtime performance with average Mean Absolute Percentage Error (MAPE) of 17.04% and time consumption of a few seconds. We also demonstrate the practicability of our framework with a real-world application. Xiebing Wang, Kai Huang 0001, Alois C. Knoll, Xuehai Qian |
HPCA | 2 |
| 2019 | End to End Learning of a Multi-Layered Snn Based on R-Stdp for a Target Tracking Snake-Like RobotabstractThis paper introduces an end-to-end learning approach based on Reward-modulated Spike-Timing-Dependent Plasticity (R-STDP) for a multi-layered spiking neural network (SNN). As a case study, a snake-like robot is used as an agent to perform target tracking tasks on the basis of our proposed approach. Since the key of R-STDP is to use rewards to modulate synapse strengthens, we first propose a general way to propagate the reward back through a multi-layered SNN. Upon the proposed approach, we build up an SNN controller that drives a snake-like robot for performing target tracking tasks. We demonstrate the practicability and advantage of our approach in terms of lateral tracking accuracy by comparing it to other state-of-the-art learning algorithms for SNNs based on R-STDP. Zhenshan Bing, Zhuangyi Jiang, Long Cheng 0007, Caixia Cai, Kai Huang 0001, Alois C. Knoll |
ICRA | 5 |
| 2019 | Mixed Frame-/Event-Driven Fast Pedestrian DetectionabstractPedestrian detection has attracted enormous research attention in the field of Intelligent Transportation System (ITS) due to that pedestrians are the most vulnerable traffic participants. So far, almost all pedestrian detection solutions are based on the conventional frame-based camera. However, they cannot perform very well in scenarios with bad light condition and high-speed motion. In this work, a Dynamic and Active Pixel Sensor (DAVIS), whose two channels concurrently output conventional gray-scale frames and asynchronous low-latency temporal contrast events of light intensity, was first used to detect pedestrians in a traffic monitoring scenario. Data from two camera channels were fed into Convolutional Neural Networks (CNNs) including three YOLOv3 models and three YOLO-tiny models to gather bounding boxes of pedestrians with respective confidence map. Furthermore, a confidence map fusion method combining the CNN-based detection results from both DAVIS channels was proposed to obtain higher accuracy. The experiments were conducted on a custom dataset collected on TUM campus. Benefiting from the high speed, low latency and wide dynamic range of the event channel, our method achieved higher frame rate and lower latency than those only using a conventional camera. Additionally, it reached higher average precision by using the fusion approach. Zhuangyi Jiang, Kai Huang 0001, Walter Stechele, Guang Chen 0001, Zhenshan Bing, Alois C. Knoll |
ICRA | 3 |
| 2019 | Needle Localization for Robot-assisted Subretinal Injection based on Deep Learning
Mingchuan Zhou, Xijia Wang, Jakob Weiss, Abouzar Eslami, Kai Huang 0001, Mathias Maier, Chris P. Lohmann, Nassir Navab, Alois C. Knoll, M. Ali Nasseri |
ICRA | 5 |
| 2019 | Energy-Efficient Slithering Gait Exploration for a Snake-Like Robot Based on Reinforcement LearningabstractSimilar to their counterparts in nature, the flexible bodies of snake-like robots enhance their movement capability and adaptability in diverse environments. However, this flexibility corresponds to a complex control task involving highly redundant degrees of freedom, where traditional model-based methods usually fail to propel the robots energy-efficiently. In this work, we present a novel approach for designing an energy-efficient slithering gait for a snake-like robot using a model-free reinforcement learning (RL) algorithm. Specifically, we present an RL-based controller for generating locomotion gaits at a wide range of velocities, which is trained using the proximal policy optimization (PPO) algorithm. Meanwhile, a traditional parameterized gait controller is presented and the parameter sets are optimized using the grid search and Bayesian optimization algorithms for the purposes of reasonable comparisons. Based on the analysis of the simulation results, we demonstrate that this RL-based controller exhibits very natural and adaptive movements, which are also substantially more energy-efficient than the gaits generated by the parameterized controller. Videos are shown at https://videoviewsite.wixsite.com/rlsnake . Zhenshan Bing, Christian Lemke, Zhuangyi Jiang, Kai Huang 0001, Alois C. Knoll |
IJCAI | 4 |
| 2019 | LiDAR Based Navigable Region Detection for Unmanned Surface VehiclesabstractDetection of the navigable regions for the unmanned surface vehicles (USVs) sailing on the narrow rivers is very important. Existing detection methods mostly depend on the cameras, which is sensitive to environments and cannot provide reliable navigable regions for sailing. In this paper, we propose a scheme to process 3D LiDAR data to achieve an accurate and robust navigable regions detection. We conduct field experiments in a narrow river in different scenarios to prove the performance of the proposed scheme, which reaches on average 93.8% precision and 92.7% recall. Xiangtong Yao, Yunxiao Shan, Jieling Li, Donghui Ma, Kai Huang 0001 |
IROS | 5 |
| 2019 | FFOB: efficient online mode-switch procrastination in mixed-criticality systems
Biao Hu 0001, Lothar Thiele, Pengcheng Huang 0001, Kai Huang 0001, Christoph Griesbeck, Alois C. Knoll |
Real Time Syst. | 4 |
| 2019 | High-Speed Scene Flow on Embedded Commercial Off-the-Shelf SystemsabstractScene flow is an essential part of a stereo-based perception system for autonomous driving and mobile robotics. As in most of these platforms, the computing resource is limited but the computing requirement is high, embedded and parallelized algorithms are of vital importance for real-time tasks. This paper develops a cross-platform embedded scene flow algorithm by using an OpenCL (Open Computing Language) programming. Meanwhile, we propose a method to achieve a good performance by using a novel coarse-grained software pipeline for the embedded stream application. Experimental results show that the proposed algorithm can boost the average processing speed to 50 fps for different commercial off-the-shelf (COTS) hardware, including desktop graphics processing units (GPUs), field-programmable gate arrays (FPGAs), and mobile phone platforms. For certain GPUs, the peak frame rates can also reach 1000 fps. By comparing the efficiency among the serial platform, we illustrate that with the help of OpenCL programming, COTS platforms can provide enough computing resources for the stereo-based perception algorithm. Long Chen 0005, Mingyue Cui, Feihu Zhang, Biao Hu 0001, Kai Huang 0001 |
IEEE Trans. Ind. Informatics | 5 |
| 2018 | Implementing and Parallelizing Real-time Lane Detection on Heterogeneous PlatformsabstractLane detection is a cardinal functionality in state-of-the-art Advanced Driver Assistant Systems (ADAS). However, it is still not straightforward to fulfill the real-time performance demand of processing High Definition (HD) images with high robustness and scalability. To address this problem, we propose an improved lane detection algorithm based on top-view image transformation and two-stage RANdom SAmple Consensus (RANSAC) model fitting. By virtue of off-line affine homography matrix adaption to bound an adaptive Region Of Interest (ROI) for subsequent on-line Warp Perspective Mapping (WPM) transformation, the algorithm can analyze arbitrary on-road videos and generate adaptive ROI without priori knowledge about camera parameter. To ensure the scalability, we present a comprehensive parallel design of the application in a heterogeneous system consisting of multi-core CPU, GPU and FPGA. We show in detail how the potentially parallel task loads are implemented and optimized so that they can be mapped to the most suitable processor so as to achieve optimal performance. Experimental results reveal that our improved algorithm can robustly process the video streams with a higher accuracy. Moreover, the heterogeneous executions are capable of processing HD 1920×1080 images with runtime performance of 81.6 fps and 47.9 fps, respectively, on an AMD FirePro W7100 GPU and a Terasic Arria 10 FPGA. Xiebing Wang, Christopher Kiwus, Canhao Wu, Biao Hu 0001, Kai Huang 0001, Alois C. Knoll |
ASAP | 5 |
| 2018 | Scheduling and shaping of complex task activations for mixed-criticality systemsabstractIn this paper, we present a new schedulability test that can cope with complex activation patterns for mixed-criticality systems. Under this analysis we proceed to present a shaping approach that can adaptively make use of system slack to improve the quality of service to less critical tasks. Compared with the state-of-the-art scheduling analysis, our scheduling analysis is more effective in handling the case that activation events can be backlogged and task deadlines can be arbitrary; and the shaping approach furthermore reduces the dropped jobs of less critical tasks without jeopardizing the guarantee to critical tasks. Extensive simulations and real-life deployment in Raspberry Pi 3 board confirm the effectiveness of our proposed schedulability test and shaping approach. Biao Hu 0001, Kai Huang 0001 |
ASP-DAC | 2 |
| 2018 | Resource-Aware Design for Reliable Autonomous Applications with Multiple Periods
Rongjie Yan, Yiqi Lv, Junjie Yang 0001, Kai Huang 0001 |
FM | 6 |
| 2018 | Design Verification and Validation for Reliable Safety-Critical Autonomous Control SystemsabstractProviding guarantees on the system behavior is mandatory for safety-critical autonomous vehicles. Among these guarantees, proving the fulfillment of real-time constraints and reliability requirements on the system is a key issue, as their violation could result in unexpected and unsafe behaviors. The violation may come from the complicated interaction between software and hardware modules, or transient hardware faults. AUTOSAR, the most popular industrial standard in the automotive domain, provides an open standardized architecture for software development, where an application can be deployed on multiple electronic control units (ECUs). We present a verification and validation method for the design of such safety-critical autonomous control systems that could tolerate transient faults. The embedded implementation of an AUTOSAR model is transformed into a three-layer system model in timed automata, so that system behavior can be evaluated and checked with hard real-time constraints and the implementing architecture. We demonstrate the feasibility of the method with a simplified controller developed for the autonomous vehicles. Rongjie Yan, Junjie Yang 0001, Kai Huang 0001 |
ICECCS | 4 |
| 2018 | A Learning Based Recovery for Damaged Snake-Like Robots
Zhuoqun Guan, Zhiyong Jian, Long Cheng 0007, Kai Huang 0001 |
ICONIP (7) | 6 |
| 2018 | End to End Learning of Spiking Neural Network Based on R-STDP for a Lane Keeping VehicleabstractLearning-based methods have demonstrated clear advantages in controlling robot tasks, such as the information fusion abilities, strong robustness, and high accuracy. Meanwhile, the on-board systems of robots have limited computation and energy resources, which are contradictory with state-of-the-art learning approaches. They are either too lightweight to solve complex problems or too heavyweight to be used for mobile applications. On the other hand, training spiking neural networks (SNNs) with biological plausibility has great potentials of performing fast computation and energy efficiency. However, the lack of effective learning rules for SNNs impedes their wide usage in mobile robot applications. This paper addresses the problem by introducing an end to end learning approach of spiking neural networks for a lane keeping vehicle. We consider the reward-modulated spike-timing-dependent-plasticity (R-STDP) as a promising solution in training SNNs, since it combines the advantages of both reinforcement learning and the well-known STDP. We test our approach in three scenarios that a Pioneer robot is controlled to keep lanes based on an SNN. Specifically, the lane information is encoded by the event data from a neuromorphic vision sensor. The SNN is constructed using R-STDP synapses in an all-to-all fashion. We demonstrate the advantages of our approach in terms of the lateral localization accuracy by comparing with other state-of-the-art learning algorithms based on SNNs. Zhenshan Bing, Claus Meschede, Kai Huang 0001, Guang Chen 0001, Florian Röhrbein, Mahmoud Akl, Alois C. Knoll |
ICRA | 3 |
| 2018 | Precision Needle Tip Localization Using Optical Coherence Tomography Images for Subretinal InjectionabstractSubretinal injection is a delicate and complex microsurgery, which requires surgeons to inject the therapeutic substance in a pre-operatively defined and intra-operatively updated subretinal target area. Due to the lack of subretinal visual feedback, it is hard to sense the insertion depth during the procedure, thus affecting the results of surgical outcome and hindering the widespread use of this treatment. This paper presents a novel approach to estimate the 3D position of the needle under the retina using the information from microscope-integrated Intraoperative Optical Coherence Tomography (iOCT). We evaluated our approach on both tissue phantom and ex-vivo porcine eyes. Evaluation results show that the average error in distance measurement is 4.7 μm (maximum of 16.5 μm). We furthermore, verified the feasibility of the proposed method to track the insertion depth of needle in robot-assisted subretinal injection. Mingchuan Zhou, Kai Huang 0001, Abouzar Eslami, Hessam Roodaki, Daniel Zapp, Mathias Maier, Chris P. Lohmann, Alois C. Knoll, M. Ali Nasseri |
ICRA | 2 |
| 2018 | Planecell: Representing Structural Space with Plane ElementsabstractReconstruction based on the stereo camera has received considerable attention recently, but two particular challenges still remain. The first concerns the need to present and compress data in an effective way, and the second is to maintain as much of the available information as possible while ensuring sufficient accuracy. To overcome these issues, we propose a new 3D representation method, namely, planecell, that extracts planarity from the depth-assisted image segmentation and then directly projects these depth planes into the 3D world. The proposed method demonstrates its advancement especially dealing with large-scale structural environment, such as autonomous driving scene. The reconstruction result of our method achieves equal accuracy compared to dense point clouds and compresses the output file 200 times. To further obtain global surfaces, an energy function formulated from Conditional Random Field that generalizes the planar relationships is maximized. We evaluate our method with reconstruction baselines on the KITTI outdoor scene dataset, and the results indicate the superiorities compared to other 3D space representation methods in accuracy, memory requirements and the scope of applications. Lei Fan 0005, Long Chen 0005, Kai Huang 0001, Dongpu Cao |
Intelligent Vehicles Symposium | 3 |
| 2018 | A Reliable Road Segmentation and Edge Extraction for Sparse 3D Lidar DataabstractPrecise segmentation of road areas using cheap Lidar is a tough and critical task due to data sparsity problem. With sparse point clouds, reliable perception of environment is difficult due to the lack of available information and loss of object features. This paper presents a new approach to use sparse 3D Lidar data for road segmentation by fusing multiple frames of point cloud. With registration of multiple frames into a same coordinate system, reliable data can be provided for later ground segmentation and edge extraction. The accuracies of extensive experiments on three kinds of roads demonstrate that the proposed approach obtains high precision and reliability. Jianfeng Gu 0001, Yuehui Wang, Long Chen 0005, Zhe Xuanyuan, Kai Huang 0001 |
Intelligent Vehicles Symposium | 6 |
| 2018 | Online Cooperative 3D Mapping for Autonomous DrivingabstractAutonomous driving requires 3D representations of the environments as high definition maps. In many cases, it is not efficient for a single vehicle to map the entire large environment. Therefore, a group of vehicles could cooperate to build maps. In this paper, we propose an approach for cooperative 3D mapping by multiple vehicles working simultaneously as a team. Each vehicle uses 3D LIDAR sensor and local mapping algorithms to build local map and the global map can be obtained by merging all the local maps in an consistent manner. The challenges in cooperative mapping lie in both accuracy and efficiency. We show that our cooperative mapping approach can save mapping time as well as reduce the accumulated error often suffered by single vehicle mapping algorithms. Meanwhile, real world experiments results indicate that our mapping algorithm can be implemented online with minimum burden imposed on communication channel and computation resources on each vehicle. Zhe Xuanyuan, Boyang Li 0009, Long Chen 0005, Kai Huang 0001 |
Intelligent Vehicles Symposium | 5 |
| 2018 | LiSeg: Lightweight Road-object Semantic Segmentation In 3D LiDAR Scans For Autonomous DrivingabstractLiDAR based perception module plays an important role in autonomous driving. However, the present CNN models are designed for image processing but not LiDAR point clouds. The performances of such models are limited by the great memory consumption and heavy computation cost. In this work, a lightweight CNN model, Liseg, is proposed to perform real-time road-object semantic segmentation on LiDAR point cloud scans for autonomous driving. The model size of Liseg is several times smaller than others, while achieving high accuracy. Wenquan Zhang, Chancheng Zhou, Junjie Yang 0001, Kai Huang 0001 |
Intelligent Vehicles Symposium | 4 |
| 2018 | Utilization-Based Scheduling of Flexible Mixed-Criticality Real-Time TasksabstractMixed-criticality models are an emerging paradigm for the design of real-time systems because of their significantly improved resource efficiency. However, formal mixed-criticality models have traditionally been characterized by two impractical assumptions: once any high-criticality task overruns, all low-criticality tasks are suspended and all other high-criticality tasks are assumed to exhibit high-criticality behaviors at the same time. In this paper, we propose a more realistic mixed-criticality model, called the flexible mixed-criticality (FMC) model, in which these two issues are addressed in a combined manner. In this new model, only the overrun task itself is assumed to exhibit high-criticality behavior, while other high-criticality tasks remain in the same mode as before. The guaranteed service levels of low-criticality tasks are gracefully degraded with the overruns of high-criticality tasks. We derive a utilization-based technique to analyze the schedulability of this new mixed-criticality model under EDF-VD scheduling. During run time, the proposed test condition serves an important criterion for dynamic service level tuning, by means of which the maximum available execution budget for low-criticality tasks can be directly determined with minimal overhead while guaranteeing mixed-criticality schedulability. Experiments demonstrate the effectiveness of the FMC scheme compared with state-of-the-art techniques. Gang Chen 0023, Nan Guan, Di Liu 0002, Qingqiang He, Kai Huang 0001, Todor P. Stefanov, Wang Yi 0001 |
IEEE Trans. Computers | 5 |
| 2018 | Adas on Cots with OpenCL: A Case Study with Lane DetectionabstractThe concept of autonomous cars is driving a boost for car electronics and the size of automotive electronics market is foreseen to double by 2025. How to benefit from this boost is an interesting question. This article presents a case study to test the feasibility of using OpenCL as the programming language and Cots components as the underlying computing platforms for Adas development. For representative Adas applications, a scalable lane detection is developed that can tune the trade-off between detection accuracy and speed. Our OpenCL implementation is tested on 14 video streams from different data-sets with different road scenarios on 5 Cots platforms. We demonstrate that the Cots platforms can provide more than sufficient computing power for the lane detection in the meanwhile our OpenCL implementation can exploit the massive parallelism provided by the Cots platforms. Kai Huang 0001, Biao Hu 0001, Long Chen 0005, Alois C. Knoll, Zhihua Wang 0001 |
IEEE Trans. Computers | 1 |
| 2017 | Online workload monitoring with the feedback of actual execution time for real-time systemsabstractGuaranteeing the system workload within design bounds is a basic requirement for a real-time system. Design-time bounds are usually based on worst-case activation patterns and worst-case execution time. While using the worst-case assumptions for online monitoring can guarantee the system safety, it also introduces unexplored slacks due to tasks consuming less than their worst-case execution times. In this paper, we introduce a monitoring scheme with the feedback of actual execution time for real-time systems. By using this runtime feedback instead of offline assumptions, this monitoring scheme can accept events that are considered as violations offline, and thereby improve the system utilization. In the experiments of both MATLAB simulation and MicroC/OS-II running in a softcore processor implemented on an FPGA, different probability distributions of actual execution time are used in analyzing how much the benefit can be gained from the feedback scheme. Biao Hu 0001, Kai Huang 0001, Gang Chen 0023, Long Cheng 0007, Alois C. Knoll |
DATE | 2 |
| 2017 | Slope angle estimation based on multi-sensor fusion for a snake-like robotabstractIn this paper, we report on a body state and ground profile estimator for a snake-like robot executing a rolling gait to travel from flat ground to a slope. With the help of the estimator, the snake-like robot can adaptively adjust the body shape and locomotion speed by changing the gait parameters for the purpose of tackling a steep slope. Specifically, we propose a repeating sequence of continuous time dynamical models to fuse kinematic encoder data with on-board Inertial Measurement Unit (IMU) measurements based on extended Kalman filter (EKF). All the sensors are mounted inside each module of the snake-like robot, which measure the joint position, the three-axis acceleration, and the three-axis angular velocity. Further, the robot changes its moving pattern under our policy, judging by the estimated angle of the ground profile. We implement this estimation procedure off-line, using data extracted from repeated runs of the snake-like robot by simulation and evaluate its performance compared to the ground truth. Zhenshan Bing, Long Cheng 0007, Alois C. Knoll, Anyang Zhong, Kai Huang 0001, Feihu Zhang |
FUSION | 5 |
| 2017 | Exploring FPGA-GPU Heterogeneous Architecture for ADAS: Towards Performance and Energy
Xiebing Wang, Kai Huang 0001, Alois C. Knoll |
ICA3PP | 3 |
| 2017 | Event-Based Target Tracking Control for a Snake Robot Using a Dynamic Vision Sensor
Zhuangyi Jiang, Zhenshan Bing, Kai Huang 0001, Guang Chen 0001, Long Cheng 0007, Alois C. Knoll |
ICONIP (6) | 3 |
| 2017 | CPG-based control of smooth transition for body shape and locomotion speed of a snake-like robotabstractIn this paper, a lightweight central pattern generator(CPG) model is designed for a snake-like robot, to achieve smooth transition of body shape and locomotion speed. First, based on the convergence behavior of the gradient system, a lightweight CPG model with fast computing time is designed and compared with other widely adopted CPG models. Then, the body shape and locomotion speed transitions in rolling gait are simulated based on the proposed CPG model. Compared with the sinusoid-based method, a smooth transition process can be achieved, without generating undesired movement or abnormal torque. Finally, extensive prototype experiments are conducted to demonstrate that the CPG-based control can effectively ensure smooth transition process and avoid abnormal torque, when the body shape and locomotion speed are changed. Zhenshan Bing, Long Cheng 0007, Kai Huang 0001, Mingchuan Zhou, Alois C. Knoll |
ICRA | 3 |
| 2017 | RGB-T SLAM: A flexible SLAM framework by combining appearance and thermal informationabstractVisual SLAM in low illumination scenes remains a considerably challenging task since the available amount of appearance information frequently stays insufficient. To tackle with this problem, we propose a novel SLAM framework by using both appearance information and thermal information, which possesses illumination-free recognizable contents, in a flexible manner. The key idea is to continuously update a RGB-T map, which contains both RGB and thermal map points to implement location and mapping. More specifically, in our SLAM system, we detect features in both RGB and thermal images and combine them together to update the RGB-T map and implement simultaneous location and mapping. Both quantitative and qualitative results demonstrate the effectiveness of our framework, especially under low illumination environments. Long Chen 0005, Libo Sun 0002, Lei Fan 0005, Kai Huang 0001, Zhe Xuanyuan |
ICRA | 5 |
| 2017 | Towards autonomous locomotion: Slithering gait design of a snake-like robot for target observation and trackingabstractIn this paper, a biologically inspired 3D slithering gait for a snake-like robot is designed and implemented for the purpose of target tracking. First, by balancing the forward speed and the stability of the robot, a straight slithering gait is modelled, under which the robot can march straight, fast, and stably. Then, for the purpose of steering, the straight slithering gait is modified into a biased slithering gait. The relationship between turning radius and gait parameters is analyzed by the resistive force theory. With the head composition algorithm, we investigate the orientation problem of the snake robot's head module to obtain stable visual information during the locomotion process. Finally, with the guidance of the vision sensor mounted in the head module, target tracking simulations and prototype experiments are conducted to demonstrate the practicality and effectiveness of the slithering gait in autonomous locomotion scenarios. Zhenshan Bing, Long Cheng 0007, Kai Huang 0001, Zhuangyi Jiang, Guang Chen 0001, Florian Röhrbein, Alois C. Knoll |
IROS | 3 |
| 2017 | Real-time scene flow on COTS embedded systems by coarse-grained software pipelineabstractScene flow is a key function of stereo-based environment perception system for mobile robotics and autonomous vehicle. Due to the heavy computing requirement and the limited computing resource, parallelized and embedded algorithms become quite important for the application of the mobile robotics. This paper develops a cross-platform embedded scene flow algorithm by using a coarse-grained software pipeline and OpenCL programming language. Our OpenCL algorithm is tested on 10 video streams from different datasets with different scenarios on different commercial-off-the-shelf (COTS) hardware. The average frame rates for the 10 videos can reach about 50 fps on both GPU and mobile device. The peak frame rates for certain videos on GPU can reach almost 450 fps. We also demonstrate that the COTS platform can provide sufficient computing power for stereo-based perception algorithm potentially by using OpenCL programming. Long Chen 0005, Mingyue Cui, Kai Huang 0001, Zhe Xuanyuan |
Intelligent Vehicles Symposium | 3 |
| 2017 | Moving-Object Detection From Consecutive Stereo Pairs Using Slanted Plane SmoothingabstractDetecting moving objects is of great importance for autonomous unmanned vehicle systems, and a challenging task especially in complex dynamic environments. This paper proposes a novel approach for the detection of moving objects and the estimation of their motion states using consecutive stereo image pairs on mobile platforms. First, we use a variant of the semi-global matching algorithm to compute initial disparity maps. Second, assisted by the initial disparities, boundaries in the image segmentation produced by simple linear iterative clustering are classified into coplanar, hinge, and occlusion. Moving points are obtained during ego-motion estimation by a modified random sample consensus) algorithm without resorting to time-consuming dense optical flow. Finally, the moving objects are extracted by merging superpixels according to the boundary types and their movements. The proposed method is accelerated on the GPU at 20 frames per second. The data which we use for testing and benchmarking is released, thus completing similar data sets. It includes 812 image pairs and 924 moving objects with ground truth for better algorithms evaluation. Experimental results demonstrate that the proposed method achieves competitive results in terms of moving-object detection and their motion state estimation in challenging urban scenarios. Long Chen 0005, Lei Fan 0005, Guodong Xie, Kai Huang 0001, Andreas Nüchter |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2016 | Minimizing peak temperature for pipelined hard real-time systems
Long Cheng 0007, Kai Huang 0001, Gang Chen 0023, Biao Hu 0001, Alois C. Knoll |
DATE | 2 |
| 2016 | A scalable lane detection algorithm on COTSs with OpenCL
Kai Huang 0001, Biao Hu 0001, Jan Botsch, Nikhil Madduri, Alois C. Knoll |
DATE | 1 |
| 2016 | On-the-fly fast overrun budgeting for mixed-criticality systemsabstractIn mixed-criticality scheduling, the widely assumed mode-switch scheme assumes that both high- and low-criticality tasks are schedulable when no tasks overrun (normal mode) and all high-criticality tasks are schedulable even when they overrun (critical mode, where low-criticality tasks are abandoned/degraded). However, this scheme triggers a mode-switch immediately after any task overruns, which can be abrupt and pessimistic. In this paper, we tackle dual-criticality systems scheduled by earliest-deadline-first, and propose light-weight mode-switch schemes that are effective in keeping the system "away" from the critical mode. Our main idea is to perform overrun budgeting for all tasks as a whole, by monitoring task executions and updating a common overrun budget. This way, the overrun budget is shared among all tasks, and adaptively replenished leveraging run-time information; consequently, mode-switch can be postponed as much as possible. Experimental results demonstrate that the proposed mode-switch schemes outperform existing solutions to a large extent, in reducing the abandoned jobs and mode-switch frequencies, as well as in increasing the time ratio that all tasks are scheduled in the system. Biao Hu 0001, Kai Huang 0001, Pengcheng Huang 0001, Lothar Thiele, Alois C. Knoll |
EMSOFT | 2 |
| 2016 | Evaluation and Improvements of Runtime Monitoring Methods for Real-Time Event StreamsabstractRuntime monitoring is of great importance as a safeguard to guarantee the correctness of system runtime behaviors. Two state-of-the-art methods, dynamic counters and l -repetitive function, were recently developed to tackle the runtime monitoring for real-time systems. While both are reported to be efficient in monitoring arbitrary events, the monitoring performance between them has not yet been evaluated. This article evaluates both methods in depth, to identify their strengths and weaknesses. New methods are proposed to efficiently monitor the many-to-one connections that are abstracted as AND and OR components on multiple inputs. Representative scenarios are used as our case studies to quantitatively demonstrate the evaluations. Both methods are implemented in hardware F pga . The timing overhead and resource usages of implementing the two methods are evaluated. Biao Hu 0001, Kai Huang 0001, Gang Chen 0023, Long Cheng 0007, Alois C. Knoll |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2016 | Adaptive Workload Management in Mixed-Criticality Systems
Biao Hu 0001, Kai Huang 0001, Gang Chen 0023, Long Cheng 0007, Alois C. Knoll |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2015 | Evaluation of runtime monitoring methods for real-time event streamsabstractRuntime monitoring is of great importance as a safe guard to guarantee the correctness of system runtime behaviors. Two new methods, i.e., dynamic counters and l-repetitive function, are recently developed to tackle the runtime monitoring for hard real-time systems. This paper investigates in depth these two newly developed runtime monitoring methods, trying to evaluate and identify their strengths and weaknesses. Representative scenarios are used as our case studies to quantitatively demonstrate our comparisons. We also provide FPGA implementations and resource usages of both methods. Biao Hu 0001, Kai Huang 0001, Gang Chen 0023, Alois C. Knoll |
ASP-DAC | 2 |
| 2015 | Adaptive runtime shaping for mixed-criticality systemsabstractThis paper investigates runtime shaping for mixed-criticality systems to increase the system QoS. Unlike the previous work in the literature that enforces an offline workload bound, an adaptively shaping approach is proposed where the incoming workload of the low-critical tasks is regulated by the actual demand of the high-critical tasks. This actual demand is adaptively updated using the historical arrival information of the high-critical tasks and thus can maximize the runtime QoS of low-critical tasks. To reduce the online overheads of computing the workload demand, a lightweight scheme with the complexity of O(n log(m)) is developed. Experiments are also provided to demonstrate the effectiveness and efficiency of our approach. Biao Hu 0001, Kai Huang 0001, Gang Chen 0023, Long Cheng 0007, Alois C. Knoll |
EMSOFT | 2 |
| 2015 | Profiling and annotation combined method for multimedia application specific MPSoC performance estimationabstractAccurate and fast performance estimation is necessary to drive design space exploration and thus support important design decisions. Current techniques are either time consuming or not accurate enough. In this paper, we solve these problems by presenting a hybrid method for multimedia multiprocessor system-on-chip (MPSoC) performance estimation. A general coverage analysis tool GNU gcov is employed to profile the execution statistics during the native simulation. To tackle the complexity and keep the analysis and simulation manageable, the orthogonalization of communication and computation parts is adopted. The estimation result of the computation part is annotated to a transaction accurate model for further analysis, by which a gradual refinement of MPSoC performance estimation is supported. The implementation and its experimental results prove the feasibility and efficiency of the proposed method. Kai Huang 0002, Siwen Xiu, Dandan Zheng 0001, Min Yu 0006, De Ma, Kai Huang 0001, Gang Chen 0023, Xiaolang Yan |
Frontiers Inf. Technol. Electron. Eng. | 7 |
| 2015 | Applying Pay-Burst-Only-Once Principle for Periodic Power Management in Hard Real-Time Pipelined Multiprocessor SystemsabstractPipelined computing is a promising paradigm for embedded system design. Designing a power management policy to reduce the power consumption of a pipelined system with nondeterministic workload is, however, nontrivial. In this article, we study the problem of energy minimization for coarse-grained pipelined systems under hard real-time constraints and propose new approaches based on an inverse use of the pay-burst-only-once principle. We formulate the problem by means of the resource demands of individual pipeline stages and propose two new approaches, a quadratic programming-based approach and fast heuristic, to solve the problem. In the quadratic programming approach, the problem is transformed into a standard quadratic programming with box constraint and then solved by a standard quadratic programming solver. Observing the problem is NP-hard, the fast heuristic is designed to solve the problem more efficiently. Our approach is scalable with respect to the numbers of pipeline stages. Simulation results using real-life applications are presented to demonstrate the effectiveness of our methods. Gang Chen 0023, Kai Huang 0001, Christian Buckl, Alois C. Knoll |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2014 | Timing anomalies in multi-core architectures due to the interference on the shared resourcesabstractTiming anomalies in single-core processors have been theoretically explained and well understood phenomenon. This paper presents new timing anomalies which occur in multi-core architectures due to the interference on the shared resources. We derive formulation to capture these anomalies and provide practical evidences using real applications from the M̈alardalen WCET benchmark suit executing on NIOS II multi-core architecture on an Altera FPGA. Hardik Shah, Kai Huang 0001, Alois C. Knoll |
ASP-DAC | 2 |
| 2014 | Abstract: Shared L2 Cache Management in Multicore Real-Time SystemabstractIn multicore system, shared cache interference has been recognized as one of the major factors that degrade the average performance as well as predictability of system. How to manage the shared cache in order to optimize the system performance while guaranteeing the system predictability is still an open issue. State-of-the-art techniques on this topic use page coloring to partition the shared cache at OS level. In this paper, we present a shared cache management scheme for multicore system. This shared cache management scheme supports way-based cache partitioning at hardware level, building task-level time-triggered reconfigurable-cache multicore system. We evaluated the proposed scheme w.r.t. different numbers of cores and cache modules and prototyped the constructed MPSoCs on FPGA. Gang Chen 0023, Biao Hu 0001, Kai Huang 0001, Alois C. Knoll, Di Liu 0002 |
FCCM | 3 |
| 2014 | Adaptive dynamic power management for hard real-time pipelined Multiprocessor SystemsabstractEnergy efficiency is a critical design concern for embedded systems. Dynamic power management (DPM) schemes in Multiprocessor System on Chips (MPSoCs) has been wildly used to explore the idleness of processors and dynamically reduce the energy consumption by putting idle processors to low-power states. In this paper, we explore how to effectively apply dynamic power management in adaptive manner to reduce leakage power consumption for coarse-grained pipelined systems under hard real-time requirements. At each adaptive point, a system transformation is proposed to model the pipeline system with unfinished events as multi-stream system. By using extended pay-burst-only-once principle, the service curves for corresponding stream can be computed as a constraint for a minimal resource demand and energy minimization problem can be formulated with respect to the resource demands at each adaptive point. One light-weight heuristic, called balance workload scheme (BWS), is proposed in this paper to solve the minimization problem. Simulation results using real-life applications are presented to demonstrate the effectiveness of our approach. Gang Chen 0023, Kai Huang 0001, Alois C. Knoll |
RTCSA | 2 |
| 2014 | Energy optimization for real-time multiprocessor system-on-chip with optimal DVFS and DPM combinationabstractEnergy optimization is a critical design concern for embedded systems. Combining D VFS +D PM is considered as one preferable technique to reduce energy consumption. There have been optimal D VFS +D PM algorithms for periodic independent tasks running on uniprocessor in the literature. Optimal combination of D VFS and D PM for periodic dependent tasks on multicore systems is however not yet reported. The challenge of this problem is that the idle intervals of cores are not easy to model. In this article, a novel technique is proposed to directly model the idle intervals of individual cores such that both D VFS and D PM can be optimized at the same time. Based on this technique, the energy optimization problem is formulated by means of mixed integrated linear programming. We also present techniques to prune the exploration space of the formulation. Experimental results using real-world benchmarks demonstrate the effectiveness of our approach compared to existing approaches. Gang Chen 0023, Kai Huang 0001, Alois C. Knoll |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2013 | Cache partitioning and scheduling for energy optimization of real-time MPSoCsabstractCache partitioning is a promising technique to reduce energy consumption of the cache subsystem for MPSoCs. Currently, most existing techniques focus primarily on static partition on core level. In this paper, we present a task-level approach and show that it outperforms core-level strategies. By taking the interference patterns of individual tasks into account, our approach generates optimal task-level cache partition schemes as well as feasible schedules at compilation time by means of a mixed integer linear programming formulation. We also present techniques to prune the exploration space of our formulation. Experimental results using real-world benchmarks demonstrate that our approach achieves 33% energy savings on average compared to core-based cache partition approaches. Gang Chen 0023, Kai Huang 0001, Alois C. Knoll |
ASAP | 2 |
| 2013 | Energy optimization with worst-case deadline guarantee for pipelined multiprocessor systemsabstractPipelined computing is a promising paradigm for embedded system design. Designing the scheduling policy for a pipelined system is however more involved. In this paper, we study the problem of the energy minimization for coarse-grained pipelined systems under hard real-time constraints and propose a method based on an inverse use of the pay-burst-only-once principle. We formulate the problem by means of the resource demands of individual pipeline stages and solve it by quadratic programming. Our approach is scalable w.r.t the number of the pipeline stages. Simulation results using real-life applications as well as commercialized processors are presented to demonstrate the effectiveness of our method. Gang Chen 0023, Kai Huang 0001, Christian Buckl, Alois C. Knoll |
DATE | 2 |
| 2013 | Effective Online Power Management with Adaptive Interplay of DVS and DPM for Embedded Real-Time SystemabstractEffective power management is an important design concern for modern embedded systems. In this paper, we present an effective framework to integrate both DVS and DPM to optimize the overall energy consumption. We propose an online algorithm to determine the optimal operating frequency and mode transition of a processor based on the runtime workload. Our algorithm runs in O(n) time, where n is the number of the events stored in the system buffer. A feasibility analysis is also presented, which serves as a criteria for setting the system buffer as well as runtime schedulabilty check. The evaluations with specifications of two commercial processors show that our algorithm is more energy-efficient compared to existing schemes in the literature. Gang Chen 0023, Kai Huang 0001, Christian Buckl, Alois C. Knoll |
DSD | 2 |
| 2012 | Conforming the runtime inputs for hard real-time embedded systemsabstractTiming is an important concern when designing an embedded system. While lots of researches on hard real-time systems focus on design-time analysis, monitoring the corresponding runtime behaviors are seldom investigated. In this paper, we investigate the conformity problem for runtime inputs of a hard real-time system. We adopt the widely used arrival curve model which captures the worst/best-cases event arrivals in the time interval domain and propose an algorithm to on-the-fly evaluate the conformity of the system input w.r.t. given arrival curves. The developed algorithm is lightweight in terms of both computation and memory overheads, which is particularly suitable for resource-constrained embedded systems. We also provide proofs and an Fpga implementation to demonstrate the effectiveness of our approach. Kai Huang 0001, Gang Chen 0023, Christian Buckl, Alois C. Knoll |
DAC | 1 |
| 2012 | Towards fault-tolerant embedded systems with imperfect fault detectionabstractMany state-of-the-art approaches on fault-tolerant system design make the simplifying assumption that all faults are detected within a certain time interval. However, based on a detailed experimental analysis, we observe that perfect fault detection is not only an impractical assumption but even if implementable also a suboptimal design decision. This paper presents an approach that takes imperfect fault detection into account. Novel analysis and optimization techniques are developed, which distinguish detectable and undetectable faults in the overall workflow. Besides synthesizing the task schedules, our approach also decides which of the available fault detectors is selected for each task instance. Experimental results show that our approach finds solutions with several orders of magnitude higher reliability than current approaches. Kai Huang 0001, Andreas Raabe, Christian Buckl, Alois C. Knoll |
DAC | 2 |
| 2012 | Embedding formal performance analysis into the design cycle of MPSoCs for real-time streaming applicationsabstractModern real-time streaming applications are increasingly implemented on multiprocessor systems-on-chip (MPSoC). The implementation, as well as the verification of real-time applications executing on MPSoCs, are difficult tasks, however. A major challenge is the performance analysis of MPSoCs, which is required for early design space exploration and final system verification. Simulation-based methods are not well-suited for this purpose, due to long runtimes and non-exhaustive corner-case coverage. To overcome these limitations, formal performance analysis methods that provide guarantees for meeting real-time constraints have been developed. Embedding formal performance analysis into the MPSoC design cycle requires the generation of a faithful analysis model and its calibration with the system-specific parameters. In this article, a design flow that automates these steps is presented. In particular, we integrate modular performance analysis (MPA) into the distributed operation layer (DOL) MPSoC programming environment. The result is an MPSoC software design flow that allows for automatically generating the system implementation, together with an analysis model for system verification. Kai Huang 0001, Wolfgang Haid, Iuliana Bacivarov, Lothar Thiele |
ACM Trans. Embed. Comput. Syst. | 1 |
| 2011 | Energy-Efficient Scheduling Algorithms for Periodic Power Management for Real-Time Event StreamsabstractAs modern VLSI technology is scaling to the deep sub-micron domain, embedded systems face a power-efficiency problem, i.e., static power consumption caused by the leakage current. This paper explores how to use dynamic power management to reduce static power consumption while guaranteeing hard real-time properties. To tackle event arrivals with non-deterministic patterns, the arrival curve model is adopted to describe event arrivals in the interval domain. To reduce runtime overhead, periodic power management is investigated, which turns on and off a system with a fixed period. To reduce the timing complexity of computing such a periodic scheme, two algorithms, which are based on a linear-segmented representation of the arrival curve model, are proposed to trade the complexity with accuracy for energy reduction. We also present simulation results to demonstrate the effectiveness of our algorithms. Kai Huang 0001, Jian-Jia Chen, Lothar Thiele |
RTCSA (1) | 1 |
| 2011 | Applying real-time interface and calculus for dynamic power management in hard real-time systems
Kai Huang 0001, Luca Santinelli, Jian-Jia Chen, Lothar Thiele, Giorgio C. Buttazzo |
Real Time Syst. | 1 |
| 2010 | Adaptive power management for real-time event streamsabstractDynamic power management has become essential for battery-driven embedded systems. This paper explores how to efficiently and effectively reduce the energy consumption of a device (system) for serving multiple event streams. Considering two different preemptive scheduling, i.e., earliest deadline first and fixed priority, we propose new method to adaptively control the power mode of the device according to historical arrivals of events. Our method can not only tackle arbitrary event arrivals but also provide hard real-time guarantees with respect to both timing and backlog constraints. Simulation results are presented as well to demonstrate the effectiveness of our approach. Kai Huang 0001, Luca Santinelli, Jian-Jia Chen, Lothar Thiele, Giorgio C. Buttazzo |
ASP-DAC | 1 |
| 2009 | Adaptive Dynamic Power Management for Hard Real-Time SystemsabstractPower dissipation has constrained the performance boosting of modern computer systems in the past decade. Dynamic power management has been widely applied to change the system (or device) state dynamically to reduce the power consumption. This paper explores how to effectively reduce the energy consumption to handle event streams with hard real-time guarantees. We adopt Real-Time Calculus to describe the event arrival and resource service by arrival curves and service curves in the interval domain, respectively. We develop online algorithms to adaptively control the power mode of the device, postponing the processing of arrival events as late as possible. Profited from the worst-case interval-based abstraction, our algorithms can on one hand tackle arbitrary event arrivals (even with burstiness) and on the other hand guarantee hard real-time requirements in terms of both timing and backlog constraints. We also present simulation results to demonstrate the effectiveness of our algorithms. Kai Huang 0001, Luca Santinelli, Jian-Jia Chen, Lothar Thiele, Giorgio C. Buttazzo |
RTSS | 1 |
| 2007 | Windowed FIFOs for FPGA-based Multiprocessor SystemsabstractFPGA-based multiprocessor systems are viable solutions for stream-based embedded applications. They provide a software abstraction which enables coarse-grained parallel deployment on an FPGA chip. A widely used model for such a deployment is the class of Kahn process networks despite their limitation to pure FIFO communications. In this paper, a new mechanism denoted as windowed FIFO is introduced, extending the functionality for data transfer. The new concept allows non-destructive read, reordering, and skipping of data within a communication channel. We present the behavior, the software interface and the hardware design of this mechanism. We introduce our abstraction of WFIFO process network which is suitable for systematic and automated synthesis while still inheriting the nice property of Kahn process networks, i.e. being determinate. Also, we present illuminating examples to demonstrate the practicality of the outlined approach. Kai Huang 0001, D. Grunert, Lothar Thiele |
ASAP | 1 |
| 2007 | Performance analysis of multimedia applications using correlated streamsabstractIn modern embedded systems, data streams are often partitioned into separate sub-streams which are processed on parallel hardware components. To analyze the performance of these systems with high accuracy, correlations between event streams must be taken into account. No methods are known so far that are able to model such a scenario with the desired accuracy. In this paper, we present a new approach to analyze correlations and we embed this analysis method into a well-established modular performance analysis framework. The presented approach enables system-level performance analysis of complete systems by taking into account stream correlations and blocking-read semantics. Experimental results on a hardware-software prototyping system are provided that show the accuracy of the analysis in a practical application Kai Huang 0001, Lothar Thiele |
DATE | 1 |