EDBT 2026 Demo / reviewers in the wild / expert
He Kong 0001
dblp:137/3221-1
· DBLP profile ↗
23ranked-venue papers
0as first author
23since 2021 · last 2026
0000-0002-1382-4186ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 11 · 11 since 2021Artificial intelligence and machine learning · 9 · 9 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Intention-Aware Diffusion Model for Pedestrian Trajectory PredictionabstractPredicting pedestrian motion trajectories is critical for the path planning and motion control of autonomous vehicles. Recent diffusion-based models have shown promising results in capturing the inherent stochasticity of pedestrian behavior for trajectory prediction. However, the absence of explicit semantic modelling of pedestrian intent in many diffusion-based methods may result in misinterpreted behaviors and reduced prediction accuracy. To address the above challenges, we propose a diffusion-based pedestrian trajectory prediction framework that incorporates both short-term and long-term motion intentions. Short-term intent is modelled using a residual polar representation, which decouples direction and magnitude to capture fine-grained local motion patterns. Long-term intent is estimated through a learnable, token-based endpoint predictor that generates multiple candidate goals with associated probabilities, enabling multimodal and context-aware intention modelling. Furthermore, we enhance the diffusion process by incorporating adaptive guidance and a residual noise predictor that dynamically refines denoising accuracy. The proposed framework is evaluated on the widely used ETH, UCY, NBA, and SDD benchmarks, demonstrating competitive results against state-of-the-art methods. Yu Liu 0163, Xiao Ren, Youfu Li 0001, He Kong 0001 |
AAAI | 5 |
| 2026 | Stabilization of Fully Actuated Nonlinear Systems: Inverse Optimal Control Design With Stability Margins
Weizhen Liu, Guangren Duan 0001, Menghua Zhang, Mehdi Golestani, He Kong 0001 |
IEEE Trans Autom. Sci. Eng. | 6 |
| 2026 | Prescribed-Time Tracking of Uncertain Nonlinear Systems With Unknown Control CoefficientsabstractIn this paper, the problem of prescribed-time tracking control with unified prescribed performance is studied for multi-input multi-output (MIMO) nonlinear systems with mismatched nonvanishing disturbances, actuator faults, and time-varying control coefficients whose sign and magnitude are both unknown. On the one hand, a novel prescribed-time stability criterion using Nussbaum functions is proposed to deal with the issues raised by the presence of mismatched nonvanishing disturbances, actuator faults, and time-varying control coefficients. This criterion is of independent interest and can be used beyond the control problem addressed in this paper. On the other hand, based on the proposed stability criterion, a prescribed-time tracking control framework is developed so that the tracking error converges to zero within a prescribed time, in the presence of the aforementioned complicating factors. Compared with existing asymptotic stability results for uncertain MIMO nonlinear systems subject to unknown control coefficients, the proposed framework guarantees that the tracking error remains within the unified prescribed performance boundary, which is uniform with respect to different initial tracking errors, thereby eliminating the need for controller redesign and stability reanalysis. The proposed control method is verified via an electromechanical system and a robot manipulator system in numerical simulation. Guangtai Tian, Wuquan Li, Mehdi Golestani, Mingming Shi, Guangren Duan 0001, He Kong 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 6 |
| 2026 | Occlusion-Aware Diffusion Model for Pedestrian Intention PredictionabstractPredicting pedestrian crossing intentions is crucial for the navigation of mobile robots and intelligent vehicles. Although recent deep learning-based models have shown significant success in forecasting intentions, few consider incomplete observation under occlusion scenarios. To tackle this challenge, we propose an Occlusion-Aware Diffusion Model (ODM) that reconstructs occluded motion patterns and leverages them to guide future intention prediction. During the denoising stage, we introduce an occlusion-aware diffusion transformer architecture to estimate noise features associated with occluded patterns, thereby enhancing the model’s ability to capture contextual relationships in occluded semantic scenarios. Furthermore, an occlusion mask-guided reverse process is introduced to effectively utilize observation information, reducing the accumulation of prediction errors and enhancing the accuracy of reconstructed motion features. The performance of the proposed method under various occlusion scenarios is comprehensively evaluated and compared with existing methods on popular benchmarks, namely PIE and JAAD. Extensive experimental results demonstrate that the proposed method achieves more robust performance than existing methods in the literature. To benefit the community, we open-source our code athttps://github.com/AISLAB-sustech/ODM Yu Liu 0163, Zedong Yang, Youfu Li 0001, He Kong 0001 |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2025 | Improved Extrinsic Calibration of Acoustic Cameras via Batch OptimizationabstractAcoustic cameras have found many applications in practice. Accurate and reliable extrinsic calibration of the microphone array and visual sensors within acoustic cameras is crucial for fusing visual and auditory measurements. Existing calibration methods either require prior knowledge of the microphone array geometry or rely on grid search which suffers from slow iteration speed or poor convergence. To overcome these limitations, in this paper, we propose an automatic calibration technique using a calibration board with both visual and acoustic markers to identify each microphone position in the camera frame. We formulate the extrinsic calibration problem (between microphones and the visual sensor) as a nonlinear least squares problem and employ a batch optimization strategy to solve the associated problem. Extensive numerical simulations and real-world experiments show that the proposed method improves both the accuracy and robustness of extrinsic parameter calibration for acoustic cameras, in comparison to existing methods. To benefit the community, we open-source all the codes and data at https://github.com/AISLAB-sustech/AcousticCamera. He Kong 0001 |
ICASSP | 4 |
| 2025 | Calibration of Multiple Asynchronous Microphone Arrays using Hybrid TDOAabstractAccurate calibration of acoustic sensing systems made of multiple asynchronous microphone arrays is essential for satisfactory performance in sound sour ce localization and tracking. State-of-the-art calibration methods for this type of system rely on the time difference of arrival and direction of arrival measurements among the microphone arrays (denoted as TDOA-M and DOA, respectively). In this paper, to enhance calibration accuracy, we propose to incorporate the time difference of arrival measurements between adjacent sound events (TDOA-S) with respect to the microphone arrays. More specifically, we propose a two-stage calibration approach, including an initial value estimation (IVE) procedure and the final joint optimization step. The IVE stage first initializes all parameters except for microphone array orientations, using hybrid TDOA (i.e., TDOA-M and TDOA-S), odometer data from a moving robot carrying a speaker, and DOA. Subsequently, microphone orientations are estimated through the iterative closest point method. The final joint optimization step estimates multiple microphone array locations, orientations, time offsets, clock drift rates, and sound source locations simultaneously. Both simulation and experiment results show that for scenarios with low or moderate TDOA noise levels, our approach outperforms existing methods in terms of accuracy. All code and data are available at https://github.com/AISLAB-sustech/Hybrid-TDOA-Multi-Calib. Wenda Pan, He Kong 0001 |
ICASSP | 4 |
| 2025 | AuralNet: Hierarchical Attention-based 3D Binaural Localization of Overlapping Speakers
Linya Fu, Yu Liu 0163, Zedong Yang, Youfu Li 0001, He Kong 0001 |
INTERSPEECH | 7 |
| 2025 | SAGENet: Binaural Echo-Based 3D Depth Estimation with Sparse Angular Queries and Refined Geometric CuesabstractIn this paper, we propose SAGENet that utilizes only binaural echoes (i.e., for scenarios when vision perception seriously degrades) for scene depth estimation. Unlike previous methods that implicitly learn spatial features from echoes, which may cause shape and scale drift, SAGENet explicitly extracts spatial cues, effectively enhancing depth estimation accuracy. First, we leverage signal processing to generate coarse 2D geometric cues, which contain scene scale and shape information, as additional input for the 3D depth estimation network. This approach aids the network in better reconstructing depth information from the scene. Given the substantial noise in the 2D geometric cues, we design a geometric cue consistency denoising loss function to help the network accurately interpret the scale and shape information embedded in the features. Second, we initialize learnable queries with angular spectrum peaks and fuse them with audio features via self-attention to guide the network to focus on the first few reflections echo dominant feature, while effectively suppressing interference from reverberation. Finally, Our experimental results on the Replica and real-world BatVision datasets show that the proposed method outperforms the existing binaural echo-based methods (including BatVision) by more than 5% and 10% in absolute relative error, respectively. To benefit the community, we open-source the code at https://github.com/zjuersdsd/SAGENet.git. Guangyao Liu, Weimeng Cui, Yuzhang Xi, Liu Yang 0021, Peixuan Hu, He Kong 0001, Zhi Wang 0003 |
IROS | 6 |
| 2025 | Observability-driven Assignment of Heterogeneous Sensors for Multi-Target TrackingabstractThis paper addresses the challenge of assigning heterogeneous sensors (i.e., robots with varying sensing capabilities) for multi-target tracking. We classify robots into two categories: (1) sufficient sensing robots, equipped with range and bearing sensors, capable of independently tracking targets, and (2) limited sensing robots, which are equipped with only range or bearing sensors and need to at least form a pair to collaboratively track a target. Our objective is to optimize tracking quality by minimizing uncertainty in target state estimation through efficient robot-to-target assignment. By leveraging matroid theory, we propose a greedy assignment algorithm that dynamically allocates robots to targets to maximize tracking quality. The algorithm guarantees constant-factor approximation bounds of 1/3 for arbitrary tracking quality functions and 1/2 for submodular functions, while maintaining polynomial-time complexity. Extensive simulations demonstrate the algorithm’s effectiveness in accurately estimating and tracking targets over extended periods. Furthermore, numerical results confirm that the algorithm’s performance is close to that of the optimal assignment, highlighting its robustness and practical applicability. Seyed Ali Rakhshan, Mehdi Golestani, He Kong 0001 |
IROS | 3 |
| 2025 | Single-Microphone-Based Sound Source Localization for Mobile Robots in Reverberant EnvironmentsabstractAccurately estimating sound source positions is crucial for robot audition. However, existing sound source localization methods typically rely on a microphone array with at least two spatially preconfigured microphones. This requirement hinders the applicability of microphone-based robot audition systems and technologies. To alleviate these challenges, we propose an online sound source localization method that uses a single microphone mounted on a mobile robot in reverberant environments. Specifically, we develop a lightweight neural network model with only 43k parameters to perform real-time distance estimation by extracting temporal information from reverberant signals. The estimated distances are then processed using an extended Kalman filter to achieve online sound source localization. To the best of our knowledge, this is the first work to achieve online sound source localization using a single microphone on a moving robot, a gap that we aim to fill in this work. Extensive experiments demonstrate the effectiveness and merits of our approach. To benefit the broader research community, we have open-sourced our code at https://github.com/JiangWAV/single-mic-SSL. Runwu Shi, Benjamin Yen 0001, He Kong 0001, Kazuhiro Nakadai |
IROS | 4 |
| 2025 | A Novel Feasibility Condition-Free Approach for Achieving Desired Precision and Unified Performance Within Prescribed TimeabstractThis paper proposes a low-complexity tracking control framework for uncertain nonlinear systems in strict feedback and normal forms, respectively. By leveraging a smooth scaling function, these control schemes ensure unified prescribed performance for the output tracking error of strict feedback nonlinear systems and the full-state tracking errors of normal form nonlinear systems. The notion of unified prescribed performance allows for different performance behaviors via performance functions, which can be either constant or time-varying with arbitrarily large initial values. The main contribution is achieving unified prescribed performance for full-state tracking errors without imposing feasibility conditions, a limitation of existing approaches. To eliminate these strict conditions, we introduce a uniform transformation independent of initial conditions. Additionally, the proposed control schemes are low-complexity since they do not require adaptive mechanisms or function approximation to deal with uncertainties and disturbances. The effectiveness of these frameworks is demonstrated through comparative analysis. Mehdi Golestani, Yongduan Song 0001, Tao Liu 0011, Xiang Xu 0003, Guangren Duan 0001, He Kong 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 6 |
| 2025 | A Novel Control Approach Accommodating Dynamic Process and Steady-State AccuracyabstractThis paper proposes an adaptive tracking control framework for nonlinear systems with unmodeled dynamics, ensuring both practical prescribed-time convergence and prescribed performance for full-state errors. Existing methods often depend on unbounded gains, focus only on output tracking error, or rely on initial conditions, restricting their practical applicability. To overcome these issues, we propose a novel adaptive control framework that constrains full-state errors independent of initial conditions and drives them to a prescribed region within a predefined time. This is achieved by using a bounded, continuously differentiable, prescribed-time gain. An adaptive mechanism with a dissipating term is designed to handle unmodeled dynamics and guarantee zero tracking error even under nonvanishing disturbances. Moreover, a smooth scaling function is introduced to enforce desired transient and steady-state performance while reducing large initial control effort. Numerical simulations demonstrate the superiority of the proposed method compared to existing approaches. Mehdi Golestani, Guangtai Tian, Yongduan Song 0001, Guangren Duan 0001, He Kong 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 5 |
| 2025 | Prescribed-Time Control of Nonlinear Systems With Global Prescribed Performance for State ErrorsabstractThis paper studies the prescribed-time tracking control problem for nonlinear systems with unknown time-varying parameters, mismatched nonvanishing uncertainties, unknown control coefficients, and potential actuator faults. The proposed control strategy employs a prescribed-time adjustment function to guarantee that state errors converge to zero within a specified time, despite the presence of nonvanishing mismatched uncertainties. The proposed controller avoids the need to use adaptive mechanisms and is therefore simple to implement. Moreover, the proposed control strategy does not require the control coefficient bounds to be known. Based on a prescribed-time scaling function and a barrier function, prescribed performance for state errors is guaranteed, which is uniform with respect to initial conditions, eliminating the need for an offline optimization algorithm to determine the controller gains. The simulation results demonstrate the effectiveness of the proposed control framework. Guangtai Tian, Mehdi Golestani, James Lam, Guangren Duan 0001, He Kong 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 5 |
| 2024 | Information-Aware Joint Calibration of Microphone Array and Sound Source LocalizationabstractAccurate calibration of microphone arrays is essential for various applications. However, existing graph SLAM-based methods, which introduce additional poselandmark constraints to enhance calibration accuracy, face challenges related to data redundancy and increased computational costs. In this paper, we propose a novel approach that combines the graph SLAM framework with an information-theoretic data selection strategy for efficient joint microphone array calibration and sound source localization. Our method evaluates and selects the most informative data segments by leveraging mutual information, effectively filtering out redundant measurements. This approach results in a streamlined and efficient calibration dataset that facilitates high estimation accuracy while significantly reducing the computational burden. Extensive simulations and real-world experiments validate that our method outperforms existing full batch data processing methods, demonstrating its potential for realtime applications. All the codes and datasets are publicly available at https://github.com/AISLAB-sustech/IA-Microphone-Calibration. Haowen Deng, Linya Fu, He Kong 0001 |
IPIN | 5 |
| 2024 | I-ASM: Iterative Acoustic Scene Mapping for Enhanced Robot Auditory Perception in Complex Indoor EnvironmentsabstractThis paper addresses the challenge of acoustic scene mapping (ASM) in complex indoor environments with multiple sound sources. Unlike existing methods that rely on prior data association or SLAM frameworks, we propose a novel particle filter-based iterative framework, termed I-ASM, for ASM using a mobile robot equipped with a microphone array and LiDAR. I-ASM harnesses an innovative "implicit association" to align sound sources with Direction of Arrival (DoA) observations without requiring explicit pairing, thereby streamlining the mapping process. Given inputs including an occupancy map, DoA estimates from various robot positions, and corresponding robot pose data, I-ASM performs multi-source mapping through an iterative cycle of "Filtering-Clustering-Implicit Associating". The proposed framework has been tested in real-world scenarios with up to 10 concurrent sound sources, demonstrating its robustness against missing and false DoA estimates while achieving high-quality ASM results. To benefit the community, we open-source all the codes and data at https://github.com/AISLAB-sustech/Acoustic-Scene-Mapping Linya Fu, Yuanzheng He, Xu Qiao, He Kong 0001 |
IROS | 5 |
| 2024 | Asynchronous Microphone Array Calibration using Hybrid TDOA InformationabstractAsynchronous microphone array calibration is a prerequisite for many audition robot applications. A popular solution to the above calibration problem is the batch form of Simultaneous Localisation and Mapping (SLAM), using the time difference of arrival measurements between two microphones (TDOA-M), and the robot (which serves as a moving sound source during calibration) odometry information. In this paper, we introduce a new form of measurement for microphone array calibration, i.e. the time difference of arrival between adjacent sound events (TDOA-S) with respect to the microphone channels. We propose to use TDOA-S and TDOA-M, called hybrid TDOA, together with odometry measurements for bath SLAM-based calibration of asynchronous microphone arrays. Extensive simulation and real-world experiments show that our method is more independent of microphone number, less sensitive to initial values (when using off-the-shelf algorithms such as Gauss-Newton iterations), and has better calibration accuracy and robustness under various TDOA noises. Simulation results also demonstrate that our method has a lower Cramér-Rao lower bound (CRLB) for microphone parameters. To benefit the community, we open-source our code and data at https://github.com/AISLAB-sustech/Hybrid-TDOA-Calib. He Kong 0001 |
IROS | 3 |
| 2024 | Robust Adaptive Control of High-Order Fully-Actuated Systems: Command Filtered Backstepping With Concurrent LearningabstractThis paper investigates the problem of tracking control for high-order strict-feedback systems (HOSFSs) with both parametric uncertainties and nonlinear function uncertainties. Based on the high-order fully-actuated (HOFA) system approach, a direct high-order robust adaptive command filtered backstepping (HORACFB) design is proposed. We adopt the concurrent learning (CL) technique to identify the unknown parameters through the examination of linear independence within the recorded data. To do so, we have introduced a novel parametric model for constructing the parameter update law under the presence of both parametric uncertainties and nonlinear function uncertainties. This is achieved by introducing new filtering variables to avoid utilizing high-order derivative information of system states required by the existing CL technique. The proposed framework can guarantee both reference tracking and unknown parameter estimation convergence (to their true values) with arbitrary accuracy that can be tuned by the designer. The closed-loop system proves to be uniformly ultimately bounded. Last but not the least, the proposed control framework avoids converting the high-order systems into first-order ones, thereby reducing unnecessary backstepping steps. Weizhen Liu, Guangren Duan 0001, He Kong 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 4 |
| 2024 | Control of Uncertain High-Order Fully Actuated Strict-Feedback Systems: A Backstepping Approach With High-Gain Observer-Based Derivative ApproximationabstractIn this article, a high-gain observer (HGO)-based differentiator is proposed to approximate the derivatives of the virtual control to ease the "explosion of complexity" problem in backstepping for the second-order strict-feedback systems (SOSFSs) and high-order strict-feedback systems (HOSFSs). Unlike the existing high-order command-filtered backstepping or extended dynamic surface control that needs to tune a large number of parameters, this proposed high-order HGO-based backstepping (HOHGOB) scheme can improve the derivative approximation performance by only tuning a single parameter, i.e., the observer gain. A further advantage of the proposed HOHGOB scheme, in addition to its simplicity and ease of implementation, is that the estimation error of the derivatives shrinks to zero as the observer gain grows to infinity. We have also rigorously established that the states of the closed-loop system achieve uniform ultimate boundedness under the designed high-order backstepping controller. Additionally, the output tracking error can be made arbitrarily small by the designer. The efficacy of the proposed scheme is numerically validated through a benchmark application to a single-link robot arm. Weizhen Liu, Guangren Duan 0001, Mehdi Golestani, He Kong 0001 |
IEEE Trans. Cybern. | 5 |
| 2024 | Adaptive Tracking Control for Underactuated Double Pendulum Overhead Cranes With Variable Cable LengthabstractAlthough the literature on control of overhead crane systems is extensive and relatively mature, there is still a need to develop strategies that can simultaneously handle factors such as the double pendulum effect, variable cable length, input saturation, input dead zones, and external disturbances. This article is concerned with adaptive tracking control for underactuated overhead cranes in the presence of the above-mentioned challenging effects. The proposed controller is composed of the following two components. First, a tracking signal vector that effectively reduces system swing magnitudes is constructed to improve the transient performance and guarantee smooth operation of the system. Second, an adaptive law is designed to estimate and compensate for the overall effects of the friction, the external disturbances, and certain nonlinearities. The system stability has been proved rigorously via the Lyapunov method and Barbalat's lemma. Extensions to the cases with input saturation and dead zones have also been discussed. Extensive numerical simulations have been conducted to verify the performance and robustness of the proposed controller, in comparison to some existing methods. Fuxing Yao, Ai-Guo Wu 0001, Mehdi Golestani, Derong Liu 0001, Guangren Duan 0001, He Kong 0001 |
IEEE Trans. Cybern. | 6 |
| 2024 | SLAM-Based Joint Calibration of Multiple Asynchronous Microphone Arrays and Sound Source LocalizationabstractRobot audition systems with multiple microphone arrays have many applications in practice. However, the accurate calibration of multiple microphone arrays remains challenging because there are many unknown parameters to be identified, including the relative transforms (i.e., orientation and translation) and asynchronous factors (i.e., initial time offset and sampling clock difference) between microphone arrays. To tackle these challenges, in this article, we adopt batch simultaneous localization and mapping (SLAM) for joint calibration of multiple asynchronous microphone arrays and sound source localization. Using the Fisher information matrix (FIM) approach, we first conduct the observability analysis (i.e., parameter identifiability) of the abovementioned calibration problem and establish necessary/sufficient conditions under which the FIM and the Jacobian matrix have full column rank, which implies the identifiability of the unknown parameters. We also discover several scenarios where the unknown parameters are not uniquely identifiable. Subsequently, we propose an effective framework to initialize the unknown parameters, which is used as the initial guess in batch SLAM for multiple microphone array calibration, aiming to further enhance optimization accuracy and convergence. Extensive numerical simulations and real experiments have been conducted to verify the performance of the proposed method. The experimental results show that the proposed pipeline achieves higher accuracy with fast convergence in comparison to methods that use the noise-corrupted ground truth of the unknown parameters as the initial guess in the optimization and other existing frameworks. Yuanzheng He, Daobilige Su, Katsutoshi Itoyama, Kazuhiro Nakadai, Junfeng Wu 0001, Shoudong Huang, Youfu Li 0001, He Kong 0001 |
IEEE Trans. Robotics | 9 |
| 2023 | SLAM-Based Joint Calibration of Differential RSS Sensor Array and Source LocalizationabstractSensor arrays generating differential received signal strength (DRSS) measurements have found many applications in robotics. However, accurate calibration of these sensor arrays remains a challenge. Most existing methods are impractical in that they assume to know signal source positions or certain parameters (i.e., path loss exponent), and try to estimate the others. In this paper, we adopt graph simultaneous localization and mapping (SLAM) as a general framework for jointly estimating the source positions and parameters of the DRSS sensor array. Our contributions are twofold. On the one hand, by using a Fisher information matrix approach, we conduct a systematic observability analysis of the corresponding SLAM setup for the calibration problem. On the other hand, we propose an effective procedure to select the initial value which is fed to Levenberg-Marquardt iterations for further improving optimization accuracy and convergence. Extensive simulation and hardware experiments show that the proposed method renders high-quality calibration results. All the codes and data are publicly available at https://github.com/SUSTech2022/DRSS-sensor-array-calibration. Linya Fu, Xu Qiao, Shoudong Huang, Guoqiang Mao, Zhiyun Lin, Youfu Li 0001, He Kong 0001 |
IECON | 7 |
| 2022 | One-Shot Learning-Based Animal Video SegmentationabstractDeep learning-based video segmentation methods can offer a good performance after being trained on the large-scale pixel labeled datasets. However, a pixel-wise manual labeling of animal images is challenging and time consuming due to irregular contours and motion blur. To achieve desirable tradeoffs between the accuracy and speed, a novel one-shot learning-based approach is proposed in this article to segment animal video with only one labeled frame. The proposed approach consists of the following three main modules: guidance frame selection utilizes “BubbleNet” to choose one frame for manual labeling, which can leverage the fine-tuning effects of the only labeled frame; Xception-based fully convolutional network localizes dense prediction using depthwise separable convolutions based on one single labeled frame; and postprocessing is used to remove outliers and sharpen object contours, which consists of two submodules—test time augmentation and conditional random field. Extensive experiments have been conducted on the DAVIS 2016 animal dataset. Our proposed video segmentation approach achieved mean intersection-over-union score of 89.5% on the DAVIS 2016 animal dataset with less run time, and outperformed the state-of-art methods (OSVOS and OSMN). The proposed one-shot learning-based approach achieves real-time and automatic segmentation of animals with only one labeled video frame. This can be potentially used further as a baseline for intelligent perception-based monitoring of animals and other domain-specific applications.11The source code, datasets, and pre-trained weights for this work are publicly [Online]. Available:https://github.com/tengfeixue-victor/One-Shot-Animal-Video-Segmentation. Tengfei Xue, Yongliang Qiao, He Kong 0001, Daobilige Su, Shirui Pan, Khalid Rafique, Salah Sukkarieh |
IEEE Trans. Ind. Informatics | 3 |
| 2021 | Necessary and Sufficient Conditions for Observability of SLAM-Based TDOA Sensor Array Calibration and Source LocalizationabstractSensor array-based systems, which adopt time difference of arrival (TDOA) measurements among the sensors, have found many robotic applications. However, for existing frameworks and systems to be useful, the sensor array needs to be calibrated accurately. Of particular interest in this article are microphone array-based robot audition systems. In our recent work, by using a moving sound source, and the graph-based formulation of simultaneous localization and mapping (SLAM), we have proposed a framework for joint sound source localization and calibration of microphone array geometrical information, together with the estimation of microphone time offset and clock difference/drift rates. However, a thorough study on the identifiability question, termed observability analysis here, in the SLAM framework for microphone array calibration and sound source localization, is still lacking in the literature. In this article, we will fill the abovementioned gap via a Fisher information matrix approach. Motivated by the equivalence between the full column rankness of the Fisher information matrix and the Jacobian matrix, we leverage the structure of the latter associated with the SLAM formulation, and present necessary and sufficient conditions guaranteeing its full column rankness, which lead to parameter identifiability. We have thoroughly discussed the 3-D case with asynchronous (with both time offset and clock drifts, or with only one of them) and synchronous microphone array, respectively. These conditions are closely related to the motion varieties of the sound source and the microphone array configuration, and have intuitive and physical interpretations. Based on the established conditions, we have also discovered some particular cases where observability is impossible. Connections with calibration of other sensors will also be discussed, amongst others. To our best knowledge, this is the first systematic work on observability analysis of SLAM-based microphone array calibration and sound source localization. The tools and concepts used in this article are also applicable to other TDOA sensing modalities such as ultrawide band (UWB) sensors. Daobilige Su, He Kong 0001, Salah Sukkarieh, Shoudong Huang |
IEEE Trans. Robotics | 2 |