EDBT 2026 Demo / reviewers in the wild / expert
Youfu Li 0001
dblp:32/238 · also You Fu Li 0001, You-Fu Li 0001
· DBLP profile ↗
162ranked-venue papers
5as first author
51since 2021 · last 2026
0000-0002-5227-1326ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 97 · 2 first-author · 26 since 2021Systems, architecture and hardware · 53 · 2 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 43 · 22 since 2021Applied, interdisciplinary, general and emerging computing · 24 · 3 first-author · 8 since 2021Human-computer interaction and ubiquitous computing · 5Computer networks · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Intention-Aware Diffusion Model for Pedestrian Trajectory PredictionabstractPredicting pedestrian motion trajectories is critical for the path planning and motion control of autonomous vehicles. Recent diffusion-based models have shown promising results in capturing the inherent stochasticity of pedestrian behavior for trajectory prediction. However, the absence of explicit semantic modelling of pedestrian intent in many diffusion-based methods may result in misinterpreted behaviors and reduced prediction accuracy. To address the above challenges, we propose a diffusion-based pedestrian trajectory prediction framework that incorporates both short-term and long-term motion intentions. Short-term intent is modelled using a residual polar representation, which decouples direction and magnitude to capture fine-grained local motion patterns. Long-term intent is estimated through a learnable, token-based endpoint predictor that generates multiple candidate goals with associated probabilities, enabling multimodal and context-aware intention modelling. Furthermore, we enhance the diffusion process by incorporating adaptive guidance and a residual noise predictor that dynamically refines denoising accuracy. The proposed framework is evaluated on the widely used ETH, UCY, NBA, and SDD benchmarks, demonstrating competitive results against state-of-the-art methods. Yu Liu 0163, Xiao Ren, Youfu Li 0001, He Kong 0001 |
AAAI | 4 |
| 2026 | Exploiting structure-semantic consistency for photorealistic SLAM with 3D Gaussian splatting
Qianang Zhou, Hai Liu 0004, Youfu Li 0001, Junlin Xiong |
Neurocomputing | 4 |
| 2026 | SC2R: similarity cues-aware evolutionary relationship mining for fine-grained bird image classification
Hai Liu 0004, Feifei Li 0003, Zhiyi Du, Tingting Liu 0006, Zhaoli Zhang, Youfu Li 0001 |
Pattern Recognit. | 8 |
| 2026 | PhyTrans: Learning Phylogenetic Relationships for FBIC via Hierarchical Taxonomy RepresentationabstractHow to accurately identify endangered bird species in complex natural environments has become an important research topic jointly concerned by the computer vision and biological conservation communities. However, they remain limited in systematically modeling cross-species semantic similarity and effectively exploiting structural stability under pose variations, making robust discrimination in highly similar species scenarios difficult. To address these challenges, we propose PhyTrans, a phylogeny-driven fine-grained bird recognition framework that achieves unified representation learning by jointly modeling inter-species phylogenetic relationships and intra-image skeletal invariance across different poses. Specifically, a phylogenetic token construction (PTC) module is designed to leverage hierarchical taxonomic information, ranging from class to species, and embed phylogenetic relationships into a hyperbolic space, which preserves hierarchical semantic distances while explicitly modeling appearance similarity induced by evolutionary relatedness. Building upon this, phylogenetic representations and intra-image skeletal structural cues are further integrated within a unified Transformer architecture through the proposed phylogenetic relationship mining (PRM) module, enabling collaborative modeling of cross-species similarity and structural invariance. Extensive experiments on the CUB-200-2011 and NABirds datasets demonstrate that PhyTrans outperforms state-of-the-art approaches, validating the critical role of phylogenetic relationships in advancing ecological visual recognition. Hai Liu 0004, Tingting Liu 0006, Dazhen Shen, Zhaoli Zhang, Youfu Li 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2026 | MHPE: Learning Morphology Relationships for Robust Head Pose Estimation With Facial Rotation RepresentationabstractAlthough accurate head pose estimation is critical for natural human-computer interaction, it remains challenging due to occlusion, extreme poses, illumination conditions, and data ambiguity issues. To address these challenges, a novel morphology aware Transformer framework (MHPE) is proposed, which can learn morphological relationships during facial rotation. The methodology is based on two key findings: cross-region geometric dependencies and angle-specific morphodynamic representations. The proposed framework incorporates two key components: adversarial feature generation, which generates robust rotation representations by adaptive multi-scale feature interaction; and morphology relationship inference, which establishes long-range dependencies between facial features through a cross-modal attention mechanism that incorporates morphological priors. Extensive evaluations on three demanding benchmarks (BIWI, AFLW2000, and 300W-LP) demonstrate state-of-the-art performance, particularly in demanding scenarios. The Python implementation will be available on request to facilitate reproducibility. Tingting Liu 0006, Jianping Ju, Zhixiong Song, Shijia Qian, Ning Rao, Hai Liu 0004, Youfu Li 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2026 | HPCTrans: Heterogeneous Plumage Cues-Aware Texton Correlation Representation for FBIC via TransformersabstractFine-grained bird image classification (FBIC) for distinguishing bird subspecies is challenging because of several issues, including a camouflaged appearance, body occlusion, and an arbitrary bird posture. To address these challenges, we propose a novel heterogeneous plumage cues-aware texton correlation representation for FBIC, which leverages texton correlation in various functional plumage regions for effective learning. Two key findings are revealed: 1) texton structural discrepancies of heterogeneous plumage; and 2) abstract region information for specific birds. On this basis, this model introduces texton coherence extraction module (TCEM) and abstract representation selection (ARS). Specifically, considering bird characteristics, TCEM is introduced to exploit the spatial statistical properties of local textons in heterogeneous plumage. To the best of our knowledge, this study is the first to introduce heterogeneous plumage cues for mining texton correlation relationship representations in FBIC tasks. In addition, a Multiscale Information Cross-Attention Transformer (MICAformer) is proposed for better modeling texton correlation representation. The experimental results on the CUB-200-2011 dataset and NABirds show the effectiveness of the proposed HPCTrans model over the state-of-the-art methods. Hai Liu 0004, Shuang Zeng, Liqian Deng, Tingting Liu 0006, Xionghua Liu, Zhaoli Zhang, Youfu Li 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2026 | ResFlow: Fine-Tuning Residual Optical Flow for Event-Based High Temporal Resolution Motion Estimation
Qianang Zhou, Junhui Hou, Yongjian Deng, Youfu Li 0001, Junlin Xiong |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2026 | Texture Affinity Cue-Aware Relationship Representation via Transformers for Facial Expression Recognition in Affective RobotsabstractAutomatic facial expression recognition (FER) from facial videos is a key component in enabling machines to understand human emotional states, which is crucial for affective robots designed to be interactive companions and applied in smart healthcare. However, FER is susceptible to challenges such as occlusion, arbitrary orientations, and illumination, making it difficult to implement precise FER models in robots. To address these issues, we propose a texture affinity cues-aware relationship representation method (FTATrans), which learns to associate facial texture with facial expressions in videos. The research reveals two key findings: 1) interaction of facial textures, and 2) texture affinity effects. On this basis, FTATrans mainly consists of two key networks: semantic-information feature generation (SFG) and texture-affinity relationship mining (TAR). In particular, the semantic relationships between different facial regions can be learned through SFG. TAR is used to capture the texture affinity relationship and integrate them with the overall facial expression information. Additionally, a loss function focused on expression-specific texture variations is proposed to guide the model in learning discriminative expression information. Experiments conducted on five video-based FER datasets demonstrate that the FTATrans model achieves state-of-the-art performance. Hai Liu 0004, Feifei Li 0003, Tingting Liu 0006, Zhaoli Zhang, Naixue Xiong, Youfu Li 0001 |
IEEE Trans. Ind. Informatics | 7 |
| 2026 | Occlusion-Aware Diffusion Model for Pedestrian Intention PredictionabstractPredicting pedestrian crossing intentions is crucial for the navigation of mobile robots and intelligent vehicles. Although recent deep learning-based models have shown significant success in forecasting intentions, few consider incomplete observation under occlusion scenarios. To tackle this challenge, we propose an Occlusion-Aware Diffusion Model (ODM) that reconstructs occluded motion patterns and leverages them to guide future intention prediction. During the denoising stage, we introduce an occlusion-aware diffusion transformer architecture to estimate noise features associated with occluded patterns, thereby enhancing the model’s ability to capture contextual relationships in occluded semantic scenarios. Furthermore, an occlusion mask-guided reverse process is introduced to effectively utilize observation information, reducing the accumulation of prediction errors and enhancing the accuracy of reconstructed motion features. The performance of the proposed method under various occlusion scenarios is comprehensively evaluated and compared with existing methods on popular benchmarks, namely PIE and JAAD. Extensive experimental results demonstrate that the proposed method achieves more robust performance than existing methods in the literature. To benefit the community, we open-source our code athttps://github.com/AISLAB-sustech/ODM Yu Liu 0163, Zedong Yang, Youfu Li 0001, He Kong 0001 |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2026 | SkeFormer: Skeletal Cues-Aware Bone Point Relationship Learning for Efficient FBIC via TransformersabstractHow to identify endangered bird species in complex outdoor environments has attracted significant attention in the fields of computer vision and machine learning. Previous studies on fine-grained bird image classification (FBIC) face numerous challenges, such as environmental occlusions and arbitrary postures, which limit the accuracy and robustness of existing methods. To address these challenges and enable more reliable bird species identification in extreme outdoor conditions, we propose a novel skeletal cues-aware bone point relationship learning for efficient FBIC via Transformers (SkeFormer). To the best of our knowledge, this is the first time skeletal relationships have been introduced to the FBIC task. Our model introduces three key modules: the skeletal relationship mining (SRM) module, the multilevel feature generation (MFG) module, and the key feature selection (KFS) module. Specifically, in SRM, the model mines the skeletal relationships among different bird species. In MFG, multiscale information is aggregated by connecting features across multiple layers. The KFS module selects key immutable regions of birds based on the learned skeletal relationships. Extensive experiments on two benchmark datasets, CUB-200-2011 and NABirds, show that SkeFormer outperforms existing state-ofthe- art models. The code for SkeFormer will be publicly available. Hai Liu 0004, Tingting Liu 0006, Zhaoli Zhang, Youfu Li 0001 |
IEEE Trans. Multim. | 7 |
| 2026 | Spatially-Guided Temporal Aggregation for Robust Event-RGB Optical Flow EstimationabstractCurrent optical flow methods exploit the stable appearance of frame (or RGB) data to establish robust correspondences across time. Event cameras, on the other hand, provide high-temporal-resolution motion cues and excel in challenging scenarios. These complementary characteristics underscore the potential of integrating frame and event data for optical flow estimation. However, most cross-modal approaches fail to fully utilize the complementary advantages, relying instead on simply stacking information. This study introduces a novel approach that uses a spatially dense modality to guide the aggregation of the temporally dense event modality, achieving effective cross-modal fusion. Specifically, we propose an event-enhanced frame representation that preserves the rich texture of frames and the basic structure of events. We use the enhanced representation as the guiding modality and employ events to capture temporally dense motion information. The robust motion features derived from the guiding modality direct the aggregation of motion information from events. To further enhance fusion, we propose a transformer-based module that complements sparse event motion features with spatially rich frame information and enhances global information propagation. Additionally, a mix-fusion encoder is designed to extract comprehensive spatiotemporal contextual features from both modalities. Extensive experiments on the MVSEC and DSEC-Flow datasets demonstrate the effectiveness of our framework. Leveraging the complementary strengths of frames and events, our method achieves leading performance on the DSEC-Flow dataset. Compared to the event-only model, frame guidance improves accuracy by 10%. Furthermore, it outperforms the state-of-the-art fusion-based method with a 4% accuracy gain and a 45% reduction in inference time. The code is publicly available athttps://github.com/ZhouQianang/STFlow. Qianang Zhou, Junhui Hou, Yongjian Deng, Youfu Li 0001, Junlin Xiong |
IEEE Trans. Multim. | 5 |
| 2026 | HomLLM: Exploiting Semantic Homology Relationship for Fine-Grained Bird Image Classification via Large Language ModelsabstractHow to recognize endangered bird species in complex outdoor environments has attracted considerable attention in the fields of computer vision and machine learning. However, fine-grained bird image classification (FBIC) is susceptible to problems such as arbitrary postures, interclass discriminability, and occlusions. We propose a novel semantic homology relationship representation learning for fine-grained bird classification with large language models, namely HomLLM, to address these challenges in FBIC effectively. Our proposed model aims to learn homology relationship representations adaptively by identifying invariant structural correspondences between visual features and semantic descriptions, using limited bird data and base class labels. Our approach yields two key findings: 1) invariant homology in key regions of birds that maintain structural consistency across different postures and 2) homological relationship that establish essential taxonomic markers among similar bird classes. Based on these insights, we propose two new modules of the model: the semantic homology generation (SHG) module and homology relationship mining (HRM) module. Specifically, in SHG, bird features are described at multiple granularities through a large language model (LLM) to establish semantic homology. In HRM, feature adaptation is performed separately for textual and visual information, and cross-modal homological interaction is performed hierarchically. In addition, we propose a hierarchical homology interaction scheme to integrate multilevel homological features while preserving structural consistency. Experiments on the commonly used bird datasets CUB-200-2011 and NABirds demonstrate that HomLLM exhibits better performance than state-of-the-art (SOTA) methods. Hai Liu 0004, Tingting Liu 0006, Lin Chen 0033, Zhaoli Zhang, Youfu Li 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 7 |
| 2026 | EvSAM: Segment Anything Model with Event-based AssistanceabstractThe general-purpose Segment Anything Model (SAM) is limited by the inherent constraints of RGB sensors, which render it inadequate for challenging real-world scenarios such as adverse lighting conditions and rapid motion. In contrast, event cameras, a novel type of bio-inspired visual sensor, offer distinct imaging advantages, including high temporal resolution and a high dynamic range. The event streams generated by these cameras provide spatiotemporal dynamic cues that are often absent in conventional image frames. To overcome the limitations of RGB-based models, we propose SAM with Event-based Assistance (EvSAM) , a novel RGB-event multi-modal semantic segmentation framework. EvSAM leverages the strong generalization capabilities of SAM while incorporating the complementary characteristics of event data to enhance scene comprehension, particularly under adverse conditions. To address the challenges of fusing two modals (image and event) with large data format discrepancy, we introduce two core components: the Multi-spatiotemporal-scale Patch Alignment Block (MS 2 PAB) and the Event-based Feature Injector (EFInj) for SAM. Specifically, the MS \({}^{2}\) PAB captures spatiotemporal semantic coherence from the event stream and transforms it into a frame-based complementary representation using a multi-spatiotemporal alignment strategy. The EFInj introduces a dynamic event feature update mechanism, wherein the fused features at a given layer guide the adaptive generation of deeper event representations. This process facilitates the integration of RGB spatial semantics with event-based motion cues. Owing to these core designs, EvSAM demonstrates superior performance on event-based semantic segmentation datasets, thereby fully validating its distinct advantages in handling extreme visual scenarios. Furthermore, we extend our model to the task of depth estimation, which further demonstrates its strong generalization ability and scalability for various downstream applications. Yuhan Liu 0021, Hao Chen 0034, Ding Ding 0002, Zhen Yang 0004, Youfu Li 0001, Yongjian Deng |
ACM Trans. Multim. Comput. Commun. Appl. | 7 |
| 2025 | Know Where You Are From: Event-Based Segmentation via Spatio-Temporal PropagationabstractEvent cameras have gained attention in segmentation due to their higher temporal resolution and dynamic range compared to traditional cameras. However, they struggle with issues like lack of color perception and triggering only at motion edges, making it hard to distinguish objects with similar contours or segment spatially continuous objects. Our work aims to address these often overlooked issues. Based on the assumption that various objects exhibit different motion patterns, we believe that embedding the historical motion states of objects into segmented scenes can effectively address these challenges. Inspired by this, we propose the ESS framework ``Know Where You Are From" (KWYAF), which incorporates past motion cues through spatio-temporal propagation embedding. This framework features two core components: the Sequential Motion Encoding Module (SME) and the Event-Based Reliable Region Selection Mechanism (ER²SM). SMEs construct prior motion features through spatio-temporal correlation modeling for boosting final segmentation, while ER²SM adapts to identify high-confidence regions, embedding motion more precisely through local window masks and reliable region selection. A large number of experiments have demonstrated the effectiveness of our proposed framework in terms of both quantity and quality. Gengyu Lyu, Hao Chen 0034, Bochen Xie, Zhen Yang 0004, Youfu Li 0001, Yongjian Deng |
AAAI | 6 |
| 2025 | In-Plane Manipulation of Soft Micro-Fiber with Ultrasonic Transducer Array and MicroscopeabstractNoncontact manipulation of soft micro-fibers has great potential in advanced manufacturing, materials science, and biomedical engineering. However, current noncontact manipulation techniques primarily focus on objects with regular shapes, e.g., solid particles, cells, or droplets, with fewer solutions available for manipulating flexible and elongated structures. In this paper, an automated ultrasonic manipulation system is introduced for in-plane soft micro-fiber manipulation, which mainly consists of an ultrasonic transducer array and a microscope. A real-time trap generation algorithm is designed to manipulate the micro-fibers by the visual feedback from microscope. An adequate theoretical analysis is also provided for explanation of the deformation behavior of micro-fiber under external forces. The system is capable of precise in-plane positioning and motion trajectory planning to micro-fiber end, and in-plane morphological reshaping to the micro-fiber. Experiments validated the effectiveness of the proposed system for the in-plane manipulation of soft micro-fibers. Finally, the system was showcased by the practical application of material property characterization. Jieyun Zou, Siyuan An, Jiaqi Li 0029, Yalin Shi, Youfu Li 0001, Song Liu 0003 |
ICRA | 6 |
| 2025 | AuralNet: Hierarchical Attention-based 3D Binaural Localization of Overlapping Speakers
Linya Fu, Yu Liu 0163, Zedong Yang, Youfu Li 0001, He Kong 0001 |
INTERSPEECH | 6 |
| 2025 | A Variable Stiffness Supernumerary Robotic Limb with Pneumatic-Tendon Coupled Actuation *abstractSupernumerary robotic limbs (SRLs) can assist humans in achieving efficient and comfortable work in daily life or industrial assembly scenarios, requiring SRLs to switch between rigidity and flexibility to perform compliant movements while also providing stable support for humans to reduce fatigue from prolonged standing, existing SRLs struggle to achieve this transition. In this study, a variable stiffness supernumerary robotic limb (VSSRL) is implemented, capable of adjusting its position and stiffness through pneumatic-tendon coupled actuation. The position of the VSSRL is accurately modulated by tendons, while its stiffness is controlled by pneumatic-tendon coupled actuation, tendons significantly increase the overall stiffness of the VSSRL, and the fiber-reinforced actuators (FRAs) can dynamically adjust its stiffness in response to changes in dynamic loads. Furthermore, a kinematic model of the VSSRL and a stiffness model under the coupling of FRAs and tendons are developed. Then, the trajectory and stiffness of the VSSRL in task execution are assigned based on human motion, and a multi-objective control system for both position and stiffness of the VSSRL is designed based on reinforcement learning (RL) algorithm, achieving collaborative control of position and stiffness for the VSSRL. The accuracy of the control system is validated through experiments, which demonstrate that the load capacity of the VSSRL is significantly enhanced under the action of tendons and FRAs, and that the VSSRL is able to provide various modes of assistance for daily life activities. Mengcheng Zhao, Juanxia Zhou, Kaizhen Huang, Aihong Ji, Xuyan Hou, Guoli Song, Youfu Li 0001 |
IROS | 10 |
| 2025 | EPA: Boosting Event-based Video Frame Interpolation with Perceptually Aligned LearningabstractEvent cameras, with their capacity to provide high temporal resolution information between frames, are increasingly utilized for video frame interpolation (VFI) in challenging scenarios characterized by high-speed motion and significant occlusion. However, prevalent issues of blur and distortion within the keyframes and ground truth data used for training and inference in these demanding conditions are frequently overlooked. This oversight impedes the perceptual realism and multi-scene generalization capabilities of existing event-based VFI (E-VFI) methods when generating interpolated frames. Motivated by the observation that semantic-perceptual discrepancies between degraded and pristine images are considerably smaller than their image-level differences, we introduce EPA. This novel E-VFI framework diverges from approaches reliant on direct image-level supervision by constructing multilevel, degradation-insensitive semantic perceptual supervisory signals to enhance the perceptual realism and multi-scene generalization of the model's predictions. Specifically, EPA operates in two phases: it first employs a DINO-based perceptual extractor, a customized style adapter, and a reconstruction generator to derive multi-layered, degradation-insensitive semantic-perceptual features ($\mathcal{S}$). Second, a novel Bidirectional Event-Guided Alignment (BEGA) module utilizes deformable convolutions to align perceptual features from keyframes to ground truth with inter-frame temporal guidance extracted from event signals. By decoupling the learning process from direct image-level supervision, EPA enhances model robustness against degraded keyframes and unreliable ground truth information. Extensive experiments demonstrate that this approach yields interpolated frames more consistent with human perceptual preferences. *The code will be released upon acceptance.* Yuhan Liu 0021, Linghui Fu, Zhen Yang 0004, Hao Chen 0034, Youfu Li 0001, Yongjian Deng |
NeurIPS | 5 |
| 2025 | Event-based video interpolation via complementary motion information
Yuhan Liu 0021, Linghui Fu, Hao Chen 0034, Zhen Yang 0004, Youfu Li 0001, Yongjian Deng |
Eng. Appl. Artif. Intell. | 5 |
| 2025 | Pixel-Level Semantics Boosted Fine-Grained Bird Image Classification
Yongjian Deng, Bochen Xie, Hai Liu 0004, Youfu Li 0001, Zhen Yang 0004 |
Eng. Appl. Artif. Intell. | 6 |
| 2025 | Neuromorphic event-based recognition boosted by motion-aware learning
Yuhan Liu 0021, Yongjian Deng, Bochen Xie, Hai Liu 0004, Zhen Yang 0004, Youfu Li 0001 |
Neurocomputing | 6 |
| 2025 | Vision-Based Closed-Loop Control With Spatiotemporal Multiplexing Strategy for Noncontact Trapping of Multiple Micro-ParticlesabstractNoncontact trapping of micro objects has great application potential in fields like material science and biomedical engineering due to its label-freeness and biocompatibility. In this paper, an automated acoustic micro-particle trapping system implemented with phased transducer array (PTA) is prototyped. The system is incorporated with a stereo vision to provide visual feedback benefited from localization of the invisible acoustic field through hydrophone scanning. Binocular vision calibration and stereo matching are realized using image Jacobian matrix. An efficient phase modulation algorithm is proposed for the calculation of desired PTA phase profile in real-time and a spatiotemporal multiplexing control strategy is adopted to dynamically generate multiple trappings. Experimental results well demonstrated that the stable trapping of multiple particles can be robustly realized by the system, leading to the improvements of robotic noncontact manipulation with invisible acoustic end-effector. Note to Practitioners—This paper is motivated by the problem that previous classic acoustic trapping was achieved as a physical phenomenon that particles within the trapping zone would be automatically trapped and thus required people to place the particle into the invisible trapping zone, which is neither precision nor efficient. Such problem is a crucial factor that limits acoustic tweezer to be further readily usable in bioengineering, surface manufacturing, and quantitative micromechanical characterization. In this work, automated acoustic trapping is presented in the context of robotics, as grasping task in conventional industrial robots, that can generate the acoustic trap exactly in the location where particles are detected (by microscopic vision, or micro-CT or acoustic imaging, etc.). This paper proposes a full pipeline to automatically trap multiple particles using ultrasonic transducer array and binocular microscopic vision. The experiments verified the ability of proposed method in simultaneously trapping three micro particles with opposite acoustic properties. Such trapping method is the foundational technology for further acoustic manipulation such as arraying and sorting, which will be the tasks in our future work. Jiaqi Li 0029, Chengxi Zhong, Teng Li 0017, Zhenhuan Sun, Youfu Li 0001, Hu Su, Song Liu 0003 |
IEEE Trans Autom. Sci. Eng. | 6 |
| 2025 | Mirror Adaptive Impedance Control of Multi-Mode Soft Exoskeleton With Reinforcement LearningabstractSoft exoskeleton robots (exosuits) have exhibited promising potentials in walking assistance with comfortable wearing experience. In this paper, a twisted string actuator (TSA) is developed and equipped with the exosuit to provide powerful driving force and variable assistance intensity for hemiplegic patients, which provides human-domain and robot-domain training modes for subjects with different movement capabilities. Since the human-exosuit coupling dynamics is difficult to be modeled due to the soft structure of the exosuit and incomplete knowledge of the wearer’s performance, accurate control and efficient assistance cannot be guaranteed in current exosuits. By taking advantage of the motion characteristic of hemiplegic patients, a mirror adaptive impedance control is proposed, where the robotic actuation is modulated based on the motion and physiological reference of the healthy limb (HL) as well as the performance of the impaired limb (IL). A linear quadratic regulation (LQR) is formulated to minimize the bilateral trajectory tracking errors and human effort, and the adaptation between the human-domain and robot-domain modes can be realized. A reinforcement learning (RL) algorithm is designed to solve the given LQR problem to optimize the impedance parameters with little information of the human or robot model. The proposed robotic system is validated through experiments to perform its effectiveness and superiority. Note to Practitioners—To assist walking for hemiplegic patients, it is crucial to provide comfortable and compliant driving force that can minimize the patients’ voluntary effort. The development of soft exoskeleton is able to realize comfortable wearing experience, and the proposed TSA can output compliant driving force with different training modes and assistance intensities. To address the problem in accurately building the human-exosuit coupled model, this work achieves the adaptive control strategy for patients with various movement capabilities in two steps. Firstly, a mirror adaptive impedance controller is proposed to make the patient’s HL tightly follow the IL’s motion for high safety guarantee performance and training autonomy. Secondly, a reinforcement learning-based LQR framework is constructed to minimize the patient’s voluntary effort by optimizing the prescribed impedance model parameters, which can significantly facilitate assistance efficiency for different patients. The experiments demonstrate that the proposed robotic system can obtain appropriate training modes and efficient walking assistance for human subjects. In the future study, it will be investigated how to assist patients in walking on different terrains with highly stable and adaptive actuation, which will accelerate the application of the proposed exosuit into activities of daily living. Kaizhen Huang, Mengcheng Zhao, Aihong Ji, Youfu Li 0001 |
IEEE Trans Autom. Sci. Eng. | 6 |
| 2025 | TransIFC: Invariant Cues-Aware Feature Concentration Learning for Efficient Fine-Grained Bird Image ClassificationabstractFine-grained bird image classification (FBIC) is not only meaningful for endangered bird observation and protection but also a prevalent task for image classification in multimedia processing and computer vision. However, FBIC suffers from several challenges, such as bird molting, complex background, and arbitrary bird posture. To effectively tackle these challenges, we present a novel invariant cues-aware feature concentration Transformer (TransIFC), which learns invariant and core information in bird images. To this end, two novel modules are proposed to leverage the characteristics of bird images, namely, the hierarchy stage feature aggregation (HSFA) module and the feature in feature abstraction (FFA) module. The HSFA module aggregates the multiscale information of bird images by concatenating multilayer features. The FFA module extracts the invariant cues of birds through feature selection based on discrimination scores. Transformer is employed as the backbone to reveal the long-dependent semantic relationships in bird images. Moreover, abundant visualizations are provided to prove the interpretability of the HSFA and FFA modules in TransIFC. Comprehensive experiments demonstrate that TransIFC can achieve state-of-the-art performance on the CUB-200-2011 dataset (91.0%) and the NABirds dataset (90.9%). Finally, extended experiments have been conducted on the Stanford Cars dataset to suggest the potential of generalizing our method on other fine-grained visual classification tasks. Hai Liu 0004, Cheng Zhang 0020, Yongjian Deng, Bochen Xie, Tingting Liu 0006, Youfu Li 0001 |
IEEE Trans. Multim. | 6 |
| 2024 | A Dynamic GCN with Cross-Representation Distillation for Event-Based LearningabstractRecent advances in event-based research prioritize sparsity and temporal precision. Approaches learning sparse point-based representations through graph CNNs (GCN) become more popular. Yet, these graph techniques hold lower performance than their frame-based counterpart due to two issues: (i) Biased graph structures that don't properly incorporate varied attributes (such as semantics, and spatial and temporal signals) for each vertex, resulting in inaccurate graph representations. (ii) A shortage of robust pretrained models. Here we solve the first problem by proposing a new event-based GCN (EDGCN), with a dynamic aggregation module to integrate all attributes of vertices adaptively. To address the second problem, we introduce a novel learning framework called cross-representation distillation (CRD), which leverages the dense representation of events as a cross-representation auxiliary to provide additional supervision and prior knowledge for the event graph. This frame-to-graph distillation allows us to benefit from the large-scale priors provided by CNNs while still retaining the advantages of graph-based models. Extensive experiments show our model and learning framework are effective and generalize well across multiple vision tasks. Yongjian Deng, Hao Chen 0034, Youfu Li 0001 |
AAAI | 3 |
| 2024 | SAM-Event-Adapter: Adapting Segment Anything Model for Event-RGB Semantic SegmentationabstractSemantic segmentation, a fundamental visual task ubiquitously employed in sectors ranging from transportation and robotics to healthcare, has always captivated the research community. In the wake of rapid advancements in large model research, the foundation model for semantic segmentation tasks, termed the Segment Anything Model (SAM), has been introduced. This model substantially addresses the dilemma of poor generalizability of previous segmentation models and the disadvantage in requiring to retrain the whole model on variant datasets. Nonetheless, segmentation models developed on SAM remain constrained by the inherent limitations of RGB sensors, particularly in scenarios characterized by complex lighting conditions and high-speed motion. Motivated by these observations, a natural recourse is to adapt SAM to additional visual modalities without compromising its robust generalizability. To achieve this, we introduce a lightweight SAM-Event-Adapter (SE-Adapter) module, which incorporates event camera data into a cross-modal learning architecture based on SAM, with only limited tunable parameters incremental. Capitalizing on the high dynamic range and temporal resolution afforded by event cameras, our proposed multi-modal Event-RGB learning architecture effectively augments the performance of semantic segmentation tasks. In addition, we propose a novel paradigm for representing event data in a patch format compatible with transformer-based models, employing multi-spatiotemporal scale encoding to efficiently extract motion and semantic correlations from event representations. Exhaustive empirical evaluations conducted on the DSEC-Semantic and DDD17 datasets provide validation of the effectiveness and rationality of our proposed approach. Yongjian Deng, Yuhan Liu 0021, Hao Chen 0034, Youfu Li 0001, Zhen Yang 0004 |
ICRA | 5 |
| 2024 | Real-Time Particle Cluster Manipulation with Holographic Acoustic End-Effector under MicroscopeabstractNon-contact particle cluster manipulation holds significant promise in the realms of advanced manufacturing, chemistry, and pharmacy. However, achieving precise and dynamic control over the spatial kinematics of particle clusters remains a significant challenge, necessitating real-time and accurately programmable robotic end-effector. To this end, we develop an innovative non-contact, precise particle cluster manipulation system with ultrasonic phased array transducer (PAT) under microscope. This system combines a physics-based deep learning algorithm for real-time calculation of phase-only holograms (POHs), supporting PAT to dynamically form acoustic fields, namely holographic acoustic end-effector (HAE). Leveraging the dynamically and accurately generated HAEs by our system, kinematics control of particle clusters including aggregation, rotation, and translation is yielded. The extensive experiments well demonstrated the effectiveness of proposed system for particle cluster manipulation. Siyuan An, Chengxi Zhong, Haojian Lu, Jiaqi Li 0029, Youfu Li 0001, Song Liu 0003 |
IROS | 7 |
| 2024 | NanoNeRF: Robot-assisted Nanoscale 360° reconstruction with neural radiance field under scanning electron microscopeabstractThe pursuit of 3D reconstruction from 2D images for nanomanipulation under scanning electron microscopy stands as a critical research endeavor. Previous methods either necessitates additional lighting which is difficult in standard SEM devices or relies on feature matching with low resolution and precision, further constraining reconstruction performance. In this paper, we propose a novel robot-assisted nanoscale 360° reconstruction approach, which simplifies SEM setups and maximizes the utilization of robot motion and feedback. By harnessing a nanorobotic system, we capture 360°multi-view images automatically with precise mapping information and camera postures. Sequentially, neural radiance field reconstruct the pixel-wise structure and synthesizing images from diverse perspectives. Experimental results using two real datasets demonstrates our approach’s efficacy, achieving PSNR of 28.1 and SSIM of 0.93 for nanotube reconstruction, and PSNR of 32.8 and SSIM of 0.98 for AFM cantilever reconstruction. These results validate the reliability and robustness of our proposed robot-assisted reconstruction method. Haojian Lu, Jiaqi Li 0029, Youfu Li 0001, Hu Su, Song Liu 0003 |
IROS | 6 |
| 2024 | Human-Robot Interaction Control for Multi-Mode Exosuit with Reinforcement LearningabstractSoft exoskeleton robots have promising potential in walking assistance with comfortable wearing experience. In this study, an exosuit equipped with a twisted string actuator (TSA) is developed to provide powerful driving force and diverse operating modes for hemiplegic patients in daily life. It is challenging to establish the human-robot coupling dynamic model due to the soft structure of the exosuit and tight coupling, precise control and effective assistance are difficult to guaranteed in current exosuits. Considering the impedance characteristics of human-robot interaction, an adaptive impedance control method based on reinforcement learning (RL) is proposed, where human motion intention is utilized to optimize impedance parameters and adjust the robot's operating mode. A nonlinear disturbance observer is proposed to compensate for the effects of model estimation errors, joint friction, and external disturbances. Experimental verification demonstrates the effectiveness and superiority of the robotic system. Kaizhen Huang, Mengcheng Zhao, Aihong Ji, Guoli Song, Youfu Li 0001 |
IROS | 7 |
| 2024 | Binary Amplitude-Only Hologram Generation for Acoustic End-Effector Design by Physics-based deep learningabstractAcoustic holography has emerged as a cutting-edge technique for constructing a micro-robot acoustic end-effector for non-contact manipulation. As one of typical implementations of acoustic holography, Binary Amplitude-Only Hologram (BAOH) featured with a simple structure provides an efficient alternative for modulating acoustic fields that support micro-robotic manipulation. In the present study, we propose a deep learning based BAOH generation method for constructing precise and high-resolution end-effector based on acoustic field. Specifically, we model the BAOH generation problem into an optimization framework. The framework combines an acoustic wave propagation model with the deep neural network, in favor of bypassing the laborious collection of labeled data and facilitating the model to learn the inverse mapping. Additionally, to address the issues of gradient invalidation and information loss caused by binarization, the framework uses an adaptive binarization layer consisting of differentiable binarization and adaptive threshold automatically learned during training, which facilitates to realize end-to-end optimization and increase the non-linear capacity of the model. The simulation experiments show that the proposed method is capable to predict BAOH that supports precise, robust, versatile and real-time construction of acoustic end-effector, enjoying broad prospects in various applications related to micro-robotic manipulation. Qing Liu 0025, Hu Su, Jiaqi Li 0029, Youfu Li 0001, Song Liu 0003 |
IROS | 4 |
| 2024 | Design and Control of a Soft Supernumerary Robotic Limb Based on Fiber-Reinforced ActuatorabstractSupernumerary robotic limbs (SRLs) provide additional wearable limbs to enhance the user’s physical abilities. Most SRLs employ rigid structures, resulting in uncomfortable wearing experience and insufficient flexible manipulation. As a new type of SRL, soft SRLs offer operational flexibility, lightweight structure, and wearing safety, compensating for the shortcomings of rigid SRLs. However, due to the complex actuation mechanisms, soft SRLs pose challenges in multiple deformations and accurate controlling. In this paper, a soft SRL actuated by fiber-reinforced actuators (FRAs) is proposed. A kinematic model is established to capture the posture of the SRL. A control system is proposed to adjust the SRL posture precisely by configuration of the FRAs. Finally, the accuracy of the proposed control strategy is verified through experiments, and the SRL prototype exhibits flexibility and adaptability to various scenarios, effectively assisting users in accomplishing complex tasks. Yonghua Lu, Mengcheng Zhao, Kaizhen Huang, Bai Chen 0002, Xuyan Hou, Youfu Li 0001 |
IROS | 8 |
| 2024 | A temporal densely connected recurrent network for event-based human pose estimation
Zhanpeng Shao, Wuzhen Wang, Jianyu Yang 0002, Youfu Li 0001 |
Pattern Recognit. | 6 |
| 2024 | Visual Servo Control for Workspace Navigation of Nanorobot End-Effector Inside SEMabstractDespite of the promising advancements in the last decade, the majority of scanning electron microscope (SEM)-based nanomanipulation tasks remain manually performed. Even in automated tasks, human intervention is required, at least during the task preparatory stage, where both the object of interest and robot end-effector are adequately enclosed in the field of view (FOV) of SEM at moderate magnification. This paper proposes a fully automated visual servo control method for workspace navigation of nanorobot end-effector that actively maintains the end-effector position offset from the center of FOV during passive working scene zooming and translation operations. We also propose using the scaling image Jacobian matrix theory to adaptively establish the hand-eye relationship of the nanorobotic system at uncalibrated magnifications, without the need for hardware regulation. This proposed method for workspace navigation is applicable to almost all commercial nanomanipulation systems. To the authors’ best knowledge, there has been no dedicated research on this problem in the literature. Experiments show that the proposed method significantly improves workspace navigation efficiency by about two-thirds, even for a skilled operator.Note to Practitioners—Workspace navigation refers to the process of moving the FOV to visually enclose the object or feature of interest at a moderate magnification, which is prerequisite for initializing either an automated or manual nanomanipulation task. This process involves consecutive zooming of the SEM in and out, as well as translating the sample stage, which can cause the end-effector to move out of the FOV. Therefore, it is necessary to move the robot end-effector with visual assistance along workspace navigation to avoid unexpected collisions. Typically, skilled operators perform this process manually in almost all nanomanipulation tasks, which is tedious and time-consuming. Therefore, developing an automated visual servo control method for workspace navigation of end-effector will significantly contribute to the nanomanipulation community. Teng Li 0017, Zhenhuan Sun, Youfu Li 0001, Song Liu 0003 |
IEEE Trans Autom. Sci. Eng. | 5 |
| 2024 | Event Voxel Set Transformer for Spatiotemporal Representation Learning on Event StreamsabstractEvent cameras are neuromorphic vision sensors that record a scene as sparse and asynchronous event streams. Most event-based methods project events into dense frames and process them using conventional vision models, resulting in high computational complexity. A recent trend is to develop point-based networks that achieve efficient event processing by learning sparse representations. However, existing works may lack robust local information aggregators and effective feature interaction operations, thus limiting their modeling capabilities. To this end, we propose an attention-aware model named Event Voxel Set Transformer (EVSTr) for efficient spatiotemporal representation learning on event streams. It first converts the event stream into voxel sets and then hierarchically aggregates voxel features to obtain robust representations. The core of EVSTr is an event voxel transformer encoder that consists of two well-designed components, including the Multi-Scale Neighbor Embedding Layer (MNEL) for local information aggregation and the Voxel Self-Attention Layer (VSAL) for global feature interaction. Enabling the network to incorporate a long-range temporal structure, we introduce a segment modeling strategy (S2TM) to learn motion patterns from a sequence of segmented voxel sets. The proposed model is evaluated on two recognition tasks, including object classification and action recognition. To provide a convincing model evaluation, we present a new event-based action recognition dataset (NeuroHAR) recorded in challenging scenarios. Comprehensive experiments show that EVSTr achieves state-of-the-art performance while maintaining low model complexity. Bochen Xie, Yongjian Deng, Zhanpeng Shao, Qingsong Xu 0002, Youfu Li 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | MMATrans: Muscle Movement Aware Representation Learning for Facial Expression Recognition via TransformersabstractHow to automatically recognize facial expression has caused concerns in industrial human–robot interaction. However, facial expression recognition (FER) is susceptible to problems, such as occlusion, arbitrary orientations, and illumination. To effectively address these challenges in FER, we present a novel facial muscle movement aware representation learning that can learn the semantic relationships of facial muscle movements in facial expression images. Two key findings are revealed: 1) muscle movements from different facial regions often show semantic relationships; and 2) not all facial muscle regions have equal contributions for different facial expressions. On this basis, this model presents two novel modules, namely, discriminative feature generation (DFG) and muscle relationship mining (MRM). Specifically, in DFG, the memory of our model for mislabeling decreases. In MRM, muscle–motion interaction among diverse facial regions is learned through visual transformers (MMATrans). Experiments on three in-the-wild FER datasets (RAF-DB, FERPlus, and AffectNet) show that our MMATrans yields better performance compared with state-of-the-art methods. Hai Liu 0004, Qiyun Zhou, Cheng Zhang 0020, Junyan Zhu, Tingting Liu 0006, Zhaoli Zhang, Youfu Li 0001 |
IEEE Trans. Ind. Informatics | 7 |
| 2024 | EHPE: Skeleton Cues-Based Gaussian Coordinate Encoding for Efficient Human Pose EstimationabstractHuman pose estimation (HPE) has many wide applications such as multimedia processing, behavior understanding and human-computer interaction. Most previous studies have encountered many constraints, such as restricted scenarios and RGB inputs. To mitigate constraints to estimating the human poses in general scenarios, we present an efficient human pose estimation model (i.e., EHPE) with joint direction cues and Gaussian coordinate encoding. Specifically, we propose an anisotropic Gaussian coordinate coding method to describe the skeleton direction cues among adjacent keypoints. To the best of our knowledge, this is the first time that the skeleton direction cues is introduced to the heatmap encoding in HPE task. Then, a multi-loss function is proposed to constrain the output to prevent the overfitting. The Kullback-Leibler divergence is introduced to measure the predication label and its ground truth one. The performance of EHPE is evaluated on two HPE datasets: MS COCO and MPII. Experimental results demonstrate that EHPE can obtain robust results, and it significantly outperforms existing state-of-the-art HPE methods. Lastly, we extend the experiments on infrared images captured by our research group. The experiments achieved the impressive results regardless of insufficient color and texture information. Hai Liu 0004, Tingting Liu 0006, Zhaoli Zhang, Youfu Li 0001 |
IEEE Trans. Multim. | 5 |
| 2024 | EISNet: A Multi-Modal Fusion Network for Semantic Segmentation With Events and ImagesabstractBio-inspired event cameras record a scene as sparse and asynchronous “events” by detecting per-pixel brightness changes. Such cameras show great potential in challenging scene understanding tasks, benefiting from the imaging advantages of high dynamic range and high temporal resolution. Considering the complementarity between event and standard cameras, we propose a multi-modal fusion network (EISNet) to improve the semantic segmentation performance. The key challenges of this topic lie in (i) how to encode event data to represent accurate scene information and (ii) how to fuse multi-modal complementary features by considering the characteristics of two modalities. To solve the first challenge, we propose an Activity-Aware Event Integration Module (AEIM) to convert event data into frame-based representations with high-confidence details via scene activity modeling. To tackle the second challenge, we introduce the Modality Recalibration and Fusion Module (MRFM) to recalibrate modal-specific representations and then aggregate multi-modal features at multiple stages. MRFM learns to generate modal-oriented masks to guide the merging of complementary features, achieving adaptive fusion. Based on these two core designs, our proposed EISNet adopts an encoder-decoder transformer architecture for accurate semantic segmentation using events and images. Experimental results show that our model outperforms state-of-the-art methods by a large margin on event-based semantic segmentation datasets. The code is publicly available athttps://github.com/bochenxie/EISNet. Bochen Xie, Yongjian Deng, Zhanpeng Shao, Youfu Li 0001 |
IEEE Trans. Multim. | 4 |
| 2024 | SLAM-Based Joint Calibration of Multiple Asynchronous Microphone Arrays and Sound Source LocalizationabstractRobot audition systems with multiple microphone arrays have many applications in practice. However, the accurate calibration of multiple microphone arrays remains challenging because there are many unknown parameters to be identified, including the relative transforms (i.e., orientation and translation) and asynchronous factors (i.e., initial time offset and sampling clock difference) between microphone arrays. To tackle these challenges, in this article, we adopt batch simultaneous localization and mapping (SLAM) for joint calibration of multiple asynchronous microphone arrays and sound source localization. Using the Fisher information matrix (FIM) approach, we first conduct the observability analysis (i.e., parameter identifiability) of the abovementioned calibration problem and establish necessary/sufficient conditions under which the FIM and the Jacobian matrix have full column rank, which implies the identifiability of the unknown parameters. We also discover several scenarios where the unknown parameters are not uniquely identifiable. Subsequently, we propose an effective framework to initialize the unknown parameters, which is used as the initial guess in batch SLAM for multiple microphone array calibration, aiming to further enhance optimization accuracy and convergence. Extensive numerical simulations and real experiments have been conducted to verify the performance of the proposed method. The experimental results show that the proposed pipeline achieves higher accuracy with fast convergence in comparison to methods that use the noise-corrupted ground truth of the unknown parameters as the initial guess in the optimization and other existing frameworks. Yuanzheng He, Daobilige Su, Katsutoshi Itoyama, Kazuhiro Nakadai, Junfeng Wu 0001, Shoudong Huang, Youfu Li 0001, He Kong 0001 |
IEEE Trans. Robotics | 8 |
| 2023 | TokenHPE: Learning Orientation Tokens for Efficient Head Pose Estimation via TransformersabstractHead pose estimation (HPE) has been widely used in the fields of human machine interaction, self-driving, and attention estimation. However, existing methods cannot deal with extreme head pose randomness and serious occlusions. To address these challenges, we identify three cues from head images, namely, neighborhood similarities, significant facial changes, and critical minority relationships. To leverage the observed findings, we propose a novel critical minority relationship-aware method based on the Transformer architecture in which the facial part relationships can be learned. Specifically, we design several orientation tokens to explicitly encode the basic orientation regions. Meanwhile, a novel token guide multiloss function is designed to guide the orientation tokens as they learn the desired regional similarities and relationships. We evaluate the proposed method on three challenging benchmark HPE datasets. Experiments show that our method achieves better performance compared with state-of-the-art methods. Our code is publicly available at https://github.com/zc2023/TokenHPE. Cheng Zhang 0020, Hai Liu 0004, Yongjian Deng, Bochen Xie, Youfu Li 0001 |
CVPR | 5 |
| 2023 | SLAM-Based Joint Calibration of Differential RSS Sensor Array and Source LocalizationabstractSensor arrays generating differential received signal strength (DRSS) measurements have found many applications in robotics. However, accurate calibration of these sensor arrays remains a challenge. Most existing methods are impractical in that they assume to know signal source positions or certain parameters (i.e., path loss exponent), and try to estimate the others. In this paper, we adopt graph simultaneous localization and mapping (SLAM) as a general framework for jointly estimating the source positions and parameters of the DRSS sensor array. Our contributions are twofold. On the one hand, by using a Fisher information matrix approach, we conduct a systematic observability analysis of the corresponding SLAM setup for the calibration problem. On the other hand, we propose an effective procedure to select the initial value which is fed to Levenberg-Marquardt iterations for further improving optimization accuracy and convergence. Extensive simulation and hardware experiments show that the proposed method renders high-quality calibration results. All the codes and data are publicly available at https://github.com/SUSTech2022/DRSS-sensor-array-calibration. Linya Fu, Xu Qiao, Shoudong Huang, Guoqiang Mao, Zhiyun Lin, Youfu Li 0001, He Kong 0001 |
IECON | 6 |
| 2023 | Orientation Cues-Aware Facial Relationship Representation for Head Pose Estimation via TransformerabstractHead pose estimation (HPE) is an indispensable upstream task in the fields of human-machine interaction, self-driving, and attention detection. However, practical head pose applications suffer from several challenges, such as severe occlusion, low illumination, and extreme orientations. To address these challenges, we identify three cues from head images, namely, critical minority relationships, neighborhood orientation relationships, and significant facial changes. On the basis of the three cues, two key insights on head poses are revealed: 1) intra-orientation relationship and 2) cross-orientation relationship. To leverage two key insights above, a novel relationship-driven method is proposed based on the Transformer architecture, in which facial and orientation relationships can be learned. Specifically, we design several orientation tokens to explicitly encode basic orientation regions. Besides, a novel token guide multi-loss function is accordingly designed to guide the orientation tokens as they learn the desired regional similarities and relationships. Experimental results on three challenging benchmark HPE datasets show that our proposed TokenHPE achieves state-of-the-art performance. Moreover, qualitative visualizations are provided to verify the effectiveness of the token-learning methodology. Hai Liu 0004, Cheng Zhang 0020, Yongjian Deng, Tingting Liu 0006, Zhaoli Zhang, Youfu Li 0001 |
IEEE Trans. Image Process. | 6 |
| 2022 | A Voxel Graph CNN for Object Classification with Event CamerasabstractEvent cameras attract researchers' attention due to their low power consumption, high dynamic range, and extremely high temporal resolution. Learning models on event-based object classification have recently achieved massive success by accumulating sparse events into dense frames to apply traditional 2D learning methods. Yet, these approaches necessitate heavy-weight models and are with high computational complexity due to the redundant information introduced by the sparse-to-dense conversion, limiting the potential of event cameras on real-life applications. This study aims to address the core problem of balancing accuracy and model complexity for event-based classification models. To this end, we introduce a novel graph representation for event data to exploit their sparsity better and customize a lightweight voxel graph convolutional neural network (EV-VGCNN) for event-based classification. Specifically, (1) using voxel-wise vertices rather than previous point-wise inputs to explicitly exploit regional 2D semantics of event streams while keeping the sparsity; (2) proposing a multi-scale feature relational layer (MFRL) to extract spatial and motion cues from each vertex discriminatively concerning its distances to neighbors. Comprehensive experiments show that our model can advance state-of-the-art classification accuracy with extremely low model complexity (merely 0.84M parameters). Yongjian Deng, Hao Chen 0034, Hai Liu 0004, Youfu Li 0001 |
CVPR | 4 |
| 2022 | Multi-stream feature refinement network for human object interaction detection
Zhanpeng Shao, Zhongyan Hu, Jianyu Yang 0002, Youfu Li 0001 |
J. Vis. Commun. Image Represent. | 4 |
| 2022 | MVF-Net: A Multi-View Fusion Network for Event-Based Object ClassificationabstractEvent-based object recognition has drawn increasing attention for event cameras’ distinguished advantages of low power consumption and high dynamic range. For this new modality, previous works based on customizing low-level descriptors are vulnerable to noise and with limited generalizability. Although recent works turn to design various deep neural networks to extract event features, they either suffer from data insufficiency to fully train the event-based model or fail to encode spatial and temporal cues simultaneously with their single view network. In this work, we address these limitations by proposing a multi-view attention-aware network, in which an event stream is projected to multi-view 2D maps to utilize well-trained 2D models and explore spatio-temporal complements. Besides, the attention mechanism is used to boost the complements in different streams for better joint inference. Comprehensive experiments show the large superiority of our model over state-of-the-art methods as well as the efficacy of our multi-view fusion framework for event data. Yongjian Deng, Hao Chen 0034, Youfu Li 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2022 | ARHPE: Asymmetric Relation-Aware Representation Learning for Head Pose Estimation in Industrial Human-Computer InteractionabstractHead pose estimation (HPE) has wide industrial applications, such as online education, human–robot interaction, and automatic manufacturing. In this article, we address two key problems in HPE based on label learning and asymmetric relation cues: 1) how to bridge the gap between the better prediction performance of networks and incorrectly label pose images in the HPE datasets and 2) how to take full advantage of the adjacent poses information around the centered pose image. We reconstruct all the incorrect labels as a two-dimensional Lorentz distribution to tackle the first problem. Instead of directly adopting the angle values ashardlabels, we assign part of the probability values (softlabels) to adjacent labels for learning discriminative feature representations. To address the second problem, we reveal the asymmetric relation nature of HPE datasets. The yaw direction and pitch direction are assigned different weights by introducing the half at half-maximum of the Lorentz distribution. Compared with the traditional end-to-end frameworks, the proposed one can leverage the asymmetric relation cues for predicting the head pose angle in the incorrect label scenarios. Extensive experiments on two public datasets and our infrared dataset demonstrate that the proposed ARHPE network significantly outperforms other state-of-the-art approaches. Hai Liu 0004, Tingting Liu 0006, Zhaoli Zhang, Arun Kumar Sangaiah, Youfu Li 0001 |
IEEE Trans. Ind. Informatics | 6 |
| 2021 | Simultaneous Precision Assembly of Multiple Objects through Coordinated Micro-robot ManipulationabstractSimultaneous assembly of multiple objects is a key technology to form solid connections among objects to get compact structures in precision assembly and micro-assembly. Dramatically different from traditional assembly of two objects, the interaction among multiple objects is more complicated on analysis and control. During simultaneous assembly of multiple objects, there are multiple mutually effected contact surfaces, and multiple force sensors are needed to perceive the interaction status. In this paper, a coordinated micro-robot manipulation strategy is proposed for simultaneous assembly problem, which is based on microscopic vision and force information. Taking simultaneous assembly of three objects as an instance, the proposed method is well articulated, including calibration of assembly system, force analysis for each contacting surface, and insertion control strategy for assembly process. The proposed method is applicable also to case with more objects. Experiment results demonstrate effectiveness of the proposed method. Song Liu 0003, Yuyu Jia, Youfu Li 0001, Yao Guo 0002, Haojian Lu |
ICRA | 3 |
| 2021 | A Multi-Level Network for Human Pose EstimationabstractAlthough multi-person human pose estimation has made great progress in recent years, the challenges such as various scales of persons, occluded keypoints, and crowded backgrounds in complex scenes are still remained to be solved. In this paper, we propose a novel multi-level pose estimation network (MLPE) to learn multi-level features that can preserve both the strong semantic clues and spatial resolution for keypoint prediction and location. More specifically, a multi-level prediction network with a feature enhancement strategy is first proposed to learn multi-level features to achieve a good trade-off between the global context information and spatial resolution. We then build a high-resolution fine network to restore high spatial resolution information based on transposed convolutions to accurately locate the keypoints. We have conducted extensive experiments on the challenging MS COCO dataset, which has proved the effectiveness of our proposed method. Code†and the experimental results are publicly online available for further research. Zhanpeng Shao, Youfu Li 0001, Jianyu Yang 0002, Xiaolong Zhou 0001 |
ICRA | 3 |
| 2021 | CNN-Based RGB-D Salient Object Detection: Learn, Select, and Fuse
Hao Chen 0034, Youfu Li 0001, Yongjian Deng, Guosheng Lin |
Int. J. Comput. Vis. | 2 |
| 2021 | Anisotropic angle distribution learning for head pose estimation and attention understanding in human-computer interaction
Hai Liu 0004, Hanwen Nie, Zhaoli Zhang, Youfu Li 0001 |
Neurocomputing | 4 |
| 2021 | Learning Representations From Skeletal Self-Similarities for Cross-View Action RecognitionabstractExisting research attention in vision-based action recognition is generally paid on recognizing actions from the same views seen in the training data. One of the big challenges in action recognition lies in the large variations of action representations as actions are captured from totally different viewpoints. This paper addresses this problem by learning view-invariant representations from skeletal self-similarities of varying scales with a very light multi-stream neural network (MSNN). As human skeletons have been proved to be an effective feature modality used for action recognition and are easy to obtain, we first create a view-invariant action description by formulating skeletal self-similarities at each frame as an image (SSI), which can show a high structural stability under view changes. Accordingly, a MSNN is designed based on 3D CNN and LSTM units to learn representations from SSIs of multiple scales, where the scheme of multiple scales provides our method with a good robustness to view changes. In addition, we integrate the computation of SSIs into the MSNN by wrapping it as a custom learnable layer thanks to its simplicity, instead of normalizing and transforming skeletons using a hand-crafted preprocessing. Extensive experimental evaluations on three challenging cross-view datasets demonstrate the effectiveness of our proposed method, which achieves superior performance to the state-of-the-art algorithms on cross-view recognition. The source code of this work will be released shortly to facilitate future studies in this field. Zhanpeng Shao, Youfu Li 0001, Hong Zhang 0013 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2021 | Learning From Images: A Distillation Learning Framework for Event CamerasabstractEvent cameras have recently drawn massive attention in the computer vision community because of their low power consumption and high response speed. These cameras produce sparse and non-uniform spatiotemporal representations of a scene. These characteristics of representations make it difficult for event-based models to extract discriminative cues (such as textures and geometric relationships). Consequently, event-based methods usually perform poorly compared to their conventional image counterparts. Considering that traditional images and event signals share considerable visual information, this paper aims to improve the feature extraction ability of event-based models by using knowledge distilled from the image domain to additionally provide explicit feature-level supervision for the learning of event data. Specifically, we propose a simple yet effective distillation learning framework, including multi-level customized knowledge distillation constraints. Our framework can significantly boost the feature extraction process for event data and is applicable to various downstream tasks. We evaluate our framework on high-level and low-level tasks, i.e., object classification and optical flow prediction. Experimental results show that our framework can effectively improve the performance of event-based models on both tasks by a large margin. Furthermore, we present a 10K dataset (CEP-DVS) for event-based object classification. This dataset consists of samples recorded under random motion trajectories that can better evaluate the motion robustness of the event-based model and is compatible with multi-modality vision tasks. Yongjian Deng, Hao Chen 0034, Youfu Li 0001 |
IEEE Trans. Image Process. | 4 |
| 2020 | 3D Gaze Estimation for Head-Mounted Devices based on Visual SaliencyabstractCompared with the maturity of 2D gaze tracking technology, 3D gaze tracking has gradually become a research hotspot in recent years. The head-mounted gaze tracker has shown great potential for gaze estimation in 3D space due to its appealing flexibility and portability. The general challenge for 3D gaze tracking algorithms is that calibration is necessary before the usage, and calibration targets cannot be easily applied in some situations or might be blocked by moving human and objects. Besides, the accuracy on depth direction has always come to be a crucial problem. Regarding the issues mentioned above, a 3D gaze estimation with auto-calibration method is proposed in this study. We use an RGBD camera as the scene camera to acquire the accurate 3D structure of the environment. The automatic calibration is achieved by uniting gaze vectors with saliency maps of the scene which aligned depth information. Finally, we determine the 3D gaze point through a point cloud generated from the RGBD camera. The experiment result demonstrates that our proposed method achieves 4.34° of average angle error in the field from 0.5m to 3m and the average depth error is 23.22mm, which is sufficient for 3D gaze estimation in the real scene. Meng Liu 0021, Youfu Li 0001, Hai Liu 0004 |
IROS | 2 |
| 2020 | Infrared head pose estimation with multi-scales feature fusion on the IRHP database for human attention recognition
Hai Liu 0004, Xiang Wang 0024, Wei Zhang 0139, Zhaoli Zhang, Youfu Li 0001 |
Neurocomputing | 5 |
| 2020 | Infrared facial expression recognition via Gaussian-based label distribution learning in the dark illumination environment for human emotion detection
Zhaoli Zhang, Chenghang Lai, Hai Liu 0004, Youfu Li 0001 |
Neurocomputing | 4 |
| 2020 | A Novel Dual-Probe-Based Micrograsping System Allowing Dexterous 3-D Orientation AdjustmentabstractThis article proposes a two-finger-based micrograsping system with high compliant borosilicate 3.3 glass probes and the corresponding sensing and control algorithms, which enables the orientation manipulation of microparts in three-dimensional (3-D) space. Compared with the existing research, the novelty of this article relies on three aspects: 1) the end-effector of the microgripper is designed to be with high compliance so that the squeeze force exerted on microparts can be more accurately regulated and the proposed microgripper is capable of manipulating fragile microparts; 2) the micrograsping system is endowed the capability to fully manipulate microparts' orientation without recurring to auxiliary probes or rotary stages; and 3) the vibration characteristic of the grasping arm is investigated as cantilever beam for gasping stability analysis and squeeze force maintaining. In specific, taking spherical microparts with dimensions in the range from tens to hundreds of micrometers as target, the grasping system configuration, and the contact model between the probe and the microparts are first presented. Afterward, kinematics-based motion control strategy for position adjustment and orientation manipulation of microparts is clarified. Then, squeeze force regulation strategy is proposed, including adhesive force evaluation, vision-based squeeze force estimation, and the micropart releasing method. Finally, the vibration characteristic of the grasping arm is investigated as cantilever beam for grasping stability analysis and the squeeze force maintaining. The reliability and availability of the proposed micrograsping system is validated by well-designed experiments. Song Liu 0003, Youfu Li 0001 |
IEEE Trans Autom. Sci. Eng. | 2 |
| 2020 | Sensing and Control for Simultaneous Precision Peg-in-Hole Assembly of Multiple ObjectsabstractThe problem of simultaneous precision assembly of multiple objects is quite practical one to form compact physical structures and functionalities in mechatronics and advanced robotics. The core research aspects facing the problem are the contact status perception between each two object and the motion planning of each separate object. These two aspects mutually affect each other and cannot be discussed separately. In this paper, we first strategically discuss the possible approaches to solve the simultaneous assembly problem and analyze their advantages and drawbacks. Then, a probabilistic control method is developed based on the incomplete perceived information of the assembly process, which can achieve the highest assembly efficiency from the strategic perspective. Specifically, by fully utilizing the mechanical properties of materials in micrometer scale, the interaction between objects is first characterized as stochastic state-transition process. Second, adopting the simultaneous feeding strategy instead of serial feeding, the current contact status between each two object is determined based on the state-transition equation as a probability distribution along a hyperline. Finally, the motion planning technique is designed taking all possible radial forces on every contact surface into consideration. The experimental results demonstrate the effectiveness of the proposed method. Song Liu 0003, Youfu Li 0001, Dengpeng Xing |
IEEE Trans Autom. Sci. Eng. | 2 |
| 2020 | Discriminative Cross-Modal Transfer Learning and Densely Cross-Level Feedback Fusion for RGB-D Salient Object DetectionabstractThis article addresses two key issues in RGB-D salient object detection based on the convolutional neural network (CNN). 1) How to bridge the gap between the "data-hungry" nature of CNNs and the insufficient labeled training data in the depth modality? 2) How to take full advantages of the complementary information among two modalities. To solve the first problem, we model the depth-induced saliency detection as a CNN-based cross-modal transfer learning problem. Instead of directly adopting the RGB CNN as initialization, we additionally train a modality classification network (MCNet) to encourage discriminative modality-specific representations in minimizing the modality classification loss. To solve the second problem, we propose a densely cross-level feedback topology, in which the cross-modal complements are combined in each level and then densely fed back to all shallower layers for sufficient cross-level interactions. Compared to traditional two-stream frameworks, the proposed one can better explore, select, and fuse cross-modal cross-level complements. Experiments show the significant and consistent improvements of the proposed CNN framework over other state-of-the-art methods. Hao Chen 0034, Youfu Li 0001, Dan Su 0001 |
IEEE Trans. Cybern. | 2 |
| 2020 | A High-Precision Automatic Wire Wrapping Approach Based on Microscopic Vision and Force InformationabstractThis paper proposes a wire wrapping approach, which enables high-precision, fully automatic, quality controllable, and visually monitored wire wrapping by actively rotating the rod and coordinately translating the wire based on microscopic vision and force information. Viewing the wire as a one-dimensional object, the proposed paper contributes to both the precision manipulation field and the engineering utilities to fabricate precision helical structures used in many fields. The basic technical contents involved are the active rotation of the rod and the coordinated translation of the manipulator, both of which are designed to keep the relative spatial relationship and the interactive force between the rod and the wire. Extensive experiments were conducted to validate the effectiveness of the proposed method to achieve high-precision wire wrapping. Experimental results show that with discretized active rotation increment about the rod axis of 6°, the local tensile deformation of the wire can be controlled within ±2 μm error range, while the average local helical angle can be controlled within ±0.3° error range with standard deviation less than 1.5°. Song Liu 0003, Youfu Li 0001 |
IEEE Trans. Ind. Informatics | 2 |
| 2020 | Flexible FTIR Spectral Imaging Enhancement for Industrial Robot Infrared Vision SensingabstractInfrared (IR) spectral imaging sensing is a powerful visual technique for industrial material recognition in robot vision systems. However, the imaging sensing data have issues of random noise and band overlap. Resolution enhancement is usually the first step in the preprocessing procedure of industrial robot vision sensing. In this article, we develop a resolution-enhancement algorithm with total variation (TV) constraints for the degraded Fourier transform IR (FTIR) spectrum due to overlap and noise degradation in the robot vision sensing. The kernel function is calculated using the spectrometer imaging systems and Fourier optical theory. The proposed model not only can remove noises effectively but also can estimate the kernel function because of the adaptive TV as constraint regularization. This model is examined by a set of simulated FTIR spectra with the Poisson noises and a series of real FTIR spectra. The proposed model is compared with the other state-of-the-art methods in terms of performance. Experimental results demonstrate that the proposed approach can split the overlap band effectively while the spectral structure details are retained satisfactorily. The enhanced high-resolution imaging spectrum data can raise the robot vision sensing accuracy in industrial intelligent systems. Tingting Liu 0006, Hai Liu 0004, Youfu Li 0001, Zengzhao Chen, Zhaoli Zhang, Sannyuya Liu |
IEEE Trans. Ind. Informatics | 3 |
| 2020 | Cross-Validated Locally Polynomial Modeling for 2-D/3-D Gaze Tracking With Head-Worn DevicesabstractIn the context of wearable gaze tracking techniques, the problems of two-dimensional (2-D) and three-dimensional (3-D) gaze estimation can be viewed as inferring 2-D epipolar lines and 3-D visual axes from eye monitoring cameras. To this end, in this article, a simple local polynomial model is proposed to back-project a pupil center onto its corresponding visual axis. Based on this approximation, a homographylike relation is derived in a local manner, and via the Leave-One-Out cross-validation criterion, training gaze samples at one certain depth is leveraged to partition entire input space into multiple overlapping subregions. Then, the gaze data at another depth are utilized to recover the epipolar point, i.e., the image eyeball center. Thus, given a pupil image, the corresponding epipolar line can be determined by the resolved homographylike mapping and the epipolar point. By using the same partition structure, 3-D gaze prediction model can be inferred by solving a nonlinear optimization problem, which aims to minimize the angular disparities between training visual directions and prediction ones. Meanwhile, it is necessary to form a good starting point and suitable constraints for the optimization problem. Otherwise, it may end up with trivial solutions, i.e., faraway eye positions. To facilitate the practical implementation of our proposed method, we also analyze how the spatial distribution of calibration points impacts the model learning accuracy. The experiment results justify the effectiveness of our proposed gaze estimation method for both the normal vision and eyewear users. Dan Su 0001, Youfu Li 0001, Hao Chen 0034 |
IEEE Trans. Ind. Informatics | 2 |
| 2020 | RGBD Salient Object Detection via Disentangled Cross-Modal FusionabstractDepth is beneficial for salient object detection (SOD) for its additional saliency cues. Existing RGBD SOD methods focus on tailoring complicated cross-modal fusion topologies, which although achieve encouraging performance, are with a high risk of over-fitting and ambiguous in studying cross-modal complementarity. Different from these conventional approaches combining cross-modal features entirely without differentiating, we concentrate our attention on decoupling the diverse cross-modal complements to simplify the fusion process and enhance the fusion sufficiency. We argue that if cross-modal heterogeneous representations can be disentangled explicitly, the cross-modal fusion process can hold less uncertainty, while enjoying better adaptability. To this end, we design a disentangled cross-modal fusion network to expose structural and content representations from both modalities by cross-modal reconstruction. For different scenes, the disentangled representations allow the fusion module to easily identify, and incorporate desired complements for informative multi-modal fusion. Extensive experiments show the effectiveness of our designs and a large outperformance over state-of-the-art methods. Hao Chen 0034, Yongjian Deng, Youfu Li 0001, Tzu-Yi Hung, Guosheng Lin |
IEEE Trans. Image Process. | 3 |
| 2019 | DISR: Deep Infrared Spectral Restoration Algorithm for Robot Sensing and Intelligent Visual Tracking SystemsabstractInfrared imaging spectrometer (IRIS) often suffers from overlapped bands and random noises, which limit the precision of subsequent processing in robot vision sensing. To address this problem, we propose a novel Gabor transform-based infrared spectrum restoration method by successfully exploring the intrinsic structure of the clean IR spectrum from the degraded one. At first, a total variation (TV) regularized Gabor coefficients adjustment descriptor is designed and incorporated into the spectrum restoration model. Then, the proposed model is inferred via an efficient optimization approach based on split Bregman iteration method. Comprehensive experiments illustrate the significant and consistent improvements of the developed model over state-of-the-art approaches. The restored high-resolution spectrum can be utilized for detecting the different materials in the robot visual tracking systems. Hai Liu 0004, Youfu Li 0001, Dan Su 0001, Zhaoli Zhang, Sannyuya Liu, Tingting Liu 0006 |
IROS | 2 |
| 2019 | Region-wise Polynomial Regression for 3D Mobile Gaze EstimationabstractIn the context of mobile gaze tracking techniques, a 3D gaze point can be calculated as the middle point between two 3D visual axes. To infer gaze directions and eyeball positions, a nonlinear optimization problem is typically formulated to minimize the angular disparities between the training gaze directions and prediction ones. Nonetheless, the experimental results reported by some previous works show that this kind of approaches are very likely to yield large prediction errors hence considered less useful for human-machine interactions. In this study, we aim to address this widespread issue in three aspects. At first, instead of using a global regression model, a simple local polynomial model is proposed to back-project a pupil center onto its corresponding visual axis. Based on the Leave-One-Out cross-validation criterion, the partition structure is automatically learned in the process of resolving a homography-like relationship. Secondly, a good starting point for nonlinear-optimization is obtained by the image eyeball center, which can be estimated by systematic parallax errors. Meanwhile, it is necessary to add the suitable constraints for 3D eye positions. Otherwise, the optimization may end up with trivial solutions, i.e., faraway eye positions. Thirdly, we explore a strategy for designing the spatial distribution of calibration points in a principled manner. The experiment results demonstrate that an encouraging gaze estimation accuracy can be achieved by our proposed framework for both the normal vision and eyewear users. Dan Su 0001, Youfu Li 0001, Hao Chen 0034 |
IROS | 2 |
| 2019 | Multi-modal fusion network with multi-scale multi-path and cross-modal interactions for RGB-D salient object detection
Hao Chen 0034, Youfu Li 0001, Dan Su 0001 |
Pattern Recognit. | 2 |
| 2019 | A Hierarchical Model for Human Action Recognition From Body-PartsabstractAs increasing attention is paid to human action recognition from skeleton data, this paper focuses on such tasks by proposing a hierarchical model to discover the structure information of body-parts involved in actions for better analysis of human actions in the skeleton data. Considering human actions as simultaneous motions of body-parts of the human skeleton, we propose a hierarchical model to simultaneously apply discriminative body-parts selection at a same scale and group coupling of bundles of body-parts at different scales, while we decompose the human skeleton into a hierarchy of body-parts of varying scales. To represent such hierarchy of body-parts, we accordingly build a hierarchical rotation and relative velocity (HRRV) descriptor. The hierarchical representations encoded by Fisher vectors of the HRRV descriptors are properly formulated into the hierarchical model via the proposed mixed norm, to apply the sparse selection of body-parts and regularize the structure of such hierarchy of body-parts. The extensive evaluations on three challenging datasets demonstrate the effectiveness of our proposed approach, which achieves superior performance compared to the state-of-the-art algorithms on datasets with various sizes, showing it is more widely applicable than existing approaches. Zhanpeng Shao, Youfu Li 0001, Yao Guo 0002, Xiaolong Zhou 0001, Shengyong Chen |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2019 | Toward Precise Gaze Estimation for Mobile Head-Mounted Gaze Tracking SystemsabstractThe gaze estimation in the mobile scenario often suffers from the extrapolation and parallax errors. In this paper, we propose a novel calibration framework to achieve the precise gaze estimation for head-mounted gaze trackers. Our proposed framework consists of two steps to learn a point-to-point and a point-to-line relations, respectively. The aim of step I is to infer the relation between pupil centers and spatially constrained points of regard. By adopting the “CalibMe” gaze data acquisition method, a sparse Gaussian Process using pseudo-inputs is used to capture the smooth residual field unmodeled by the polynomial function. Meanwhile, a distraction detection criterion is introduced to identify the moment when user's attention is taken away from the calibration point thereby removing outliers. By combining with the point-to-point relation inferred in step I, the observed parallax errors are leveraged in step II to obtain a point-to-line relation, i.e., each pupil center will correspond to an epipolar line. Thus, the real image gaze point projected from different depths is predicted as the intersection of two epipolar lines inferred from binocular data. The simulation and experimental results show the effectiveness of our proposed calibration framework for head-mounted gaze trackers. Dan Su 0001, Youfu Li 0001, Hao Chen 0034 |
IEEE Trans. Ind. Informatics | 2 |
| 2019 | Three-Stream Attention-Aware Network for RGB-D Salient Object DetectionabstractPrevious RGB-D fusion systems based on convolutional neural networks (CNNs) typically employ a two-stream architecture, in which RGB and depth inputs are learnt independently. The multi-modal fusion stage is typically performed by concatenating the deep features from each stream in the inference process. The traditional two-stream architecture might experience insufficient multi-modal fusion due to two following limitations: (1) The cross-modal complementarity is rarely studied in the bottom-up path, wherein we believe the crossmodal complements can be combined to learn new discriminative features to enlarge the RGB-D representation community; (2) The cross-modal channels are typically combined by undifferentiated concatenation, which appears ambiguous to select cross-modal complementary features. In this work, we address these two limitations by proposing a novel three-stream attention-aware multi-modal fusion network. In the proposed architecture, a cross-modal distillation stream, accompanying the RGB-specific and depth-specific streams, is introduced to extract new RGB-D features in each level in the bottom-up path. Furthermore, the channel-wise attention mechanism is innovatively introduced to the cross-modal cross-level fusion problem to adaptively select complementary feature maps from each modality in each level. Extensive experiments report the effectiveness of the proposed architecture and the significant improvement over state-of-theart RGB-D salient object detection methods. Hao Chen 0034, Youfu Li 0001 |
IEEE Trans. Image Process. | 2 |
| 2018 | Progressively Complementarity-Aware Fusion Network for RGB-D Salient Object DetectionabstractHow to incorporate cross-modal complementarity sufficiently is the cornerstone question for RGB-D salient object detection. Previous works mainly address this issue by simply concatenating multi-modal features or combining unimodal predictions. In this paper, we answer this question from two perspectives: (1) We argue that if the complementary part can be modelled more explicitly, the cross-modal complement is likely to be better captured. To this end, we design a novel complementarity-aware fusion (CA-Fuse) module when adopting the Convolutional Neural Network (CNN). By introducing cross-modal residual functions and complementarity-aware supervisions in each CA-Fuse module, the problem of learning complementary information from the paired modality is explicitly posed as asymptotically approximating the residual function. (2) Exploring the complement across all the levels. By cascading the CA-Fuse module and adding level-wise supervision from deep to shallow densely, the cross-level complement can be selected and combined progressively. The proposed RGB-D fusion network disambiguates both cross-modal and cross-level fusion processes and enables more sufficient fusion results. The experiments on public datasets show the effectiveness of the proposed CA-Fuse module and the RGB-D salient object detection network. Hao Chen 0034, Youfu Li 0001 |
CVPR | 2 |
| 2018 | A Hierarchical Model for Action Recognition Based on Body PartsabstractAs increasing attention is paid on human action recognition from skeleton data, this paper focuses on such tasks by proposing a hierarchical model to discover the structure information of body-parts involved in human actions. Considering human actions as simultaneous motions of different body-parts of the human skeleton, we propose a hierarchical model to simultaneously apply discriminative body-parts selection at a same scale and group coupling of bundles of body-parts at different scales, while we decompose the human skeleton into a hierarchy of body-parts of varying scales. To represent such hierarchy of body-parts, we accordingly build a hierarchical RRV (Rotation and Relative Velocity) descriptors. The hierarchical representations encoded by Fisher vectors of the hierarchical RRV descriptors are properly formulated into the hierarchical model via the proposed hierarchical mixed norm, to apply sparse selection of body-parts and regularize the structure of such hierarchy of body-parts. The extensive evaluations on three challenging datasets demonstrate the effectiveness of our proposed approach, which achieves superior performance compared to state-of-the-art results on different sizes of datasets, showing it is more widely applicable than existing approaches. Zhanpeng Shao, Youfu Li 0001, Yao Guo 0002, Jianyu Yang 0002, Zhenhua Wang 0003 |
ICRA | 2 |
| 2018 | Attention-Aware Cross-Modal Cross-Level Fusion Network for RGB-D Salient Object DetectionabstractConvolutional neural networks have achieved wide success in RGB saliency detection. Recently, the advent of RGB-D sensors such as Kinect provide additional geometric saliency cues. However, the key challenge for RGB-D salient object detection that how to fuse RGB and depth information sufficiently is still under-studied. Traditional works mainly follow the two-stream architecture and combine RGB and depth features/decisions in an early or late point. The multi-modal fusion stage is performed by directly concatenating the features from two modalities without selection. In this work, we address this question by proposing a novel network with a distinguished insight: A selection module is significantly helpful for more informative and sufficient cross-modal cross-level combination. To this end, we introduce a top-down RGB-D fusion network which integrates an attention-aware cross-modal cross-level fusion block in each level to select discriminative features from each level and each modality. Extensive experiments on public datasets show that the proposed network is able to solve the key problems in RGB-D fusion and achieves state-of-the-art performance on RGB-D salient object detection. Hao Chen 0034, Youfu Li 0001, Dan Su 0001 |
IROS | 2 |
| 2018 | Indoor Mapping and Localization for Pedestrians using Opportunistic Sensing with SmartphonesabstractIndoor localization for pedestrians has gained increasing popularity among the rich body of literature for the last decade. In this paper, a low-cost indoor mapping and localization solution is proposed using the opportunistic signals from ambient indoor environments with a smartphone. It is composed of GraphSLAM-based offline mapping and Bayesian filtering-based online localization using generated signal maps. The GraphSLAM front-end is constructed by motion constraints from pedestrian dead-reckoning (PDR), loop-closure constraints identified by magnetic sequence matching with WiFi signal similarity validation, and observation constraints from opportunistic magnetic headings after error rejection. Globally consistent trajectories are created by graph optimization, after which signal maps (e.g., WiFi, magnetic fields, lights) are generated by Gaussian Processes Regression (GPR) for later localization. We propose to use the pseudo-wall constraints from the GPR variance map of magnetic fields and the lights measurements as observations for particle filtering. The proposed method is evaluated on several datasets collected from both the in-compass office buildings and outside public areas. Real-time localization is demonstrated on a smartphone in an office building covering 2000 square meters with the 50- and 90-percentile accuracies being 2.30 m and 3.41 m, respectively. Lujia Wang 0001, Youfu Li 0001, Ming Liu 0001 |
IROS | 3 |
| 2018 | Plugo: A Scalable Visible Light Communication System Towards Low-Cost Indoor LocalizationabstractIndoor localization is critical to many location-aware applications, however, a low-cost solution with guaranteed accuracies has not yet come. Visible Light Communication (VLC-) based localization techniques are very promising to fill this gap. In this paper, we propose Plugo, a novel VLC system with random multiple access towards low-cost indoor localization. Compared to conventional RF-based approaches that rely on dedicated wireless access points as location beacons, the proposed system has the potential to deliver better accuracies with reduced cost. Specifically, we build a handful of compact VLC-compatible LED bulbs out of low-cost offthe-shelf components (around $10 total cost for each assembly) and recover VLC signals using a cheap photodiode receiver. The basic framed slotted Additive Links On-line Hawaii Area (ALOHA) is exploited to achieve random multiple access over the shared optical medium. We show its effectiveness in beacon broadcasting by experiments, and further, demonstrate a preliminary localization result with sound accuracy by using fingerprinting-based methods in a customized testbed. Lujia Wang 0001, Youfu Li 0001, Ming Liu 0001 |
IROS | 3 |
| 2018 | Visual adaptive tracking for monocular omnidirectional camera
Yazhe Tang, Zhi Gao 0005, Feng Lin 0003, Youfu Li 0001, Fei Wen 0004 |
J. Vis. Commun. Image Represent. | 4 |
| 2018 | DSRF: A flexible trajectory descriptor for articulated human action recognition
Yao Guo 0002, Youfu Li 0001, Zhanpeng Shao |
Pattern Recognit. | 2 |
| 2018 | RRV: A Spatiotemporal Descriptor for Rigid Body Motion RecognitionabstractThe motion behaviors of a rigid body can be characterized by a six degrees of freedom motion trajectory, which contains the 3-D position vectors of a reference point on the rigid body and 3-D rotations of this rigid body over time. This paper devises a rotation and relative velocity (RRV) descriptor by exploring the local translational and rotational invariants of rigid body motion trajectories, which is insensitive to noise, invariant to rigid transformation and scale. The RRV descriptor is then applied to characterize motions of a human body skeleton modeled as articulated interconnections of multiple rigid bodies. To show the descriptive ability of our RRV descriptor, we explore its potentials and applications in different rigid body motion recognition tasks. The experimental results on benchmark datasets demonstrate that our RRV descriptor learning discriminative motion patterns can achieve superior results for various recognition tasks. Yao Guo 0002, Youfu Li 0001, Zhanpeng Shao |
IEEE Trans. Cybern. | 2 |
| 2017 | Development of precise mobile gaze tracking system based on online sparse Gaussian process regression and smooth-pursuit identificationabstractIn this paper, we aim to address two challenges in the implementation of mobile gaze tracking systems, i.e., the parallax error and the inflexible calibration procedure. Our proposed method mainly involves two steps and all the calibration process can be completed without needs to receive user's commands. At first, instead of fixating at calibration points successively, users are required to fixate at one calibration point while smoothly and persistently varying the head position. In this case, the eye movements will compensate for head movements due to the activation of the smooth pursuit system. Then the PCA analysis is applied for distinguishing the smooth pursuits from other kinds of eye movements to get reliable training data. Meanwhile, a online sparse Gaussian Process using the FITC approximation is proposed to model the relationship between image pupil centers and gaze points. The next step aims to compensate the parallax error by recovering the epipolar geometry of gaze tracking systems and eyeballs. Users are asked to fixate at the points with different distances and detected parallax errors can be applied to estimate the epipolar geometry of mobile gaze trackers by solving a nonlinear optimization problem. Thus the real image gaze point with different depths can be straightforwardly estimated as the intersection point of two epipolar lines derived from binocular data. The simulation and experimental results demonstrate the effectiveness of our proposed method. Dan Su 0001, Youfu Li 0001 |
ICRA | 2 |
| 2017 | Two-eye model-based gaze estimation from a Kinect sensorabstractIn this paper, we present an effective and accurate gaze estimation method based on two-eye model of a subject with the tolerance of free head movement from a Kinect sensor. To accurately and efficiently determine the point of gaze, i) we employ two-eye model to improve the estimation accuracy; ii) we propose an improved convolution-based means of gradients method to localize the iris center in 3D space; iii) we present a new personal calibration method that only needs one calibration point. The method approximates the visual axis as a line from the iris center to the gaze point to determine the eyeball centers and the Kappa angles. The final point of gaze can be calculated by using the calibrated personal eye parameters. We experimentally evaluate the proposed gaze estimation method on eleven subjects. Experimental results demonstrate that our gaze estimation method has an average estimation accuracy around 1.99°, which outperforms many leading methods in the state-of-the-art. Xiaolong Zhou 0001, Haibin Cai, Youfu Li 0001, Honghai Liu 0001 |
ICRA | 3 |
| 2017 | RGB-D Saliency Detection by Multi-stream Late Fusion Network
Hao Chen 0034, Youfu Li 0001, Dan Su 0001 |
ICVS | 2 |
| 2017 | M3Net: Multi-scale multi-path multi-modal fusion network and example application to RGB-D salient object detectionabstractFusing RGB and depth data is compelling in boosting performance for various robotic and computer vision tasks. Typically, the streams of RGB and depth information are merged into a single fusion point in an early or late stage to generate combined features or decisions. The single fusion point also means single fusion path, which is congested and inflexible to fuse all the information from different modalities. As a result, the fusion process is brute-force and consequently insufficient. To address this problem, we propose a multi-scale multi-path multi-modal fusion network (M3Net), in which the fusion path is scattered to diversify the contributions of each modality from global and local perspectives. Specially, the CNN streams of each modality are fused with a global understanding path and meanwhile a local capturing path. By filtering and regulating information flow in a multi-path way, the M3Net is equipped with more adaptive and flexible fusion mechanism, thus easing the gradient-based learning process, improving the directness and transparency of the fusion process and simultaneously facilitating the fusion process with multi-scale perspectives. Comprehensive experiments demonstrate the significant and consistent improvements of the proposed approach over state-of-the-art methods. Hao Chen 0034, Youfu Li 0001, Dan Su 0001 |
IROS | 2 |
| 2017 | MSM-HOG: A flexible trajectory descriptor for rigid body motion recognitionabstractThis paper proposes a flexible descriptor for representing 6-D rigid body motion trajectories, which not only shows strong invariances and descriptive ability but also achieves satisfactory results in both recognition accuracy and efficiency. 6-D rigid body motion trajectories are first transformed into the Multi-layer Self-similarity Matrices (MSM) representation. The MSM is the combination of the square similarity matrices in three layers, which captures both local and global spatiotemporal features of the trajectories. Next, the Histogram of Oriented Gradients (HOG) features extracted from the MSM representation are concatenated as the final MSM-HOG trajectory descriptor. Then we train the Support Vector Machine (SVM) classifier with the linear kernel for multicalss motion recognition. Finally, rigid body motion recognition experiments on two public datasets are conducted to verify the effectiveness and efficiency of the proposed method. Yao Guo 0002, Youfu Li 0001, Zhanpeng Shao |
IROS | 2 |
| 2017 | On Multiscale Self-Similarities Description for Effective Three-Dimensional/Six-Dimensional Motion Trajectory RecognitionabstractMotion trajectories provide compact informative clues in characterizing motion behaviors of human bodies, robots, and moving objects. This paper devises an invariant and unified descriptor for three-dimensional/six-dimensional (3-D/6-D) motion trajectories recognition by exploring the latent motion patterns in the multiscale self-similarity matrices (MSM) within a motion trajectory and its components. The MSM approach transforms a motion trajectory in Euclidean space into a set of similarity matrices and exhibits strong invariances, in which each matrix can be regarded as a grayscale image. Next, the histograms of oriented gradients features extracted from the MSM representation are concatenated as the final trajectory descriptor. In addition, an improved kernel MSM is raised by calculating the pairwise kernel distances. Finally, extensive 3-D/6-D motion trajectory recognition experiments on three public datasets with a linear support vector machine classifier are conducted to verify the effectiveness and efficiency of the proposed approach. Yao Guo 0002, Youfu Li 0001, Zhanpeng Shao |
IEEE Trans. Ind. Informatics | 2 |
| 2016 | Invariant multi-scale descriptor for shape representation, matching and retrieval
Jianyu Yang 0002, Hongxing Wang 0001, Junsong Yuan 0001, Youfu Li 0001, Jianyang Liu |
Comput. Vis. Image Underst. | 4 |
| 2016 | Parsing 3D motion trajectory for gesture recognition
Jianyu Yang 0002, Junsong Yuan 0001, Youfu Li 0001 |
J. Vis. Commun. Image Represent. | 3 |
| 2016 | Parameterized Distortion-Invariant Feature for Robust Tracking in Omnidirectional VisionabstractCentral catadioptric omnidirectional images exhibit serious nonlinear distortions due to the involved quadratic mirrors. Therefore, features based on the conventional pin-hole model are hard to achieve satisfactory performances when directly applied to the distorted omnidirectional images. This paper analyzes the catadioptric geometry to facilitate modeling the nonlinear distortions of omnidirectional images. Different to the conventional imaging model, the prior information is considered in catadioptric system. A parameterized neighborhood mapping model is proposed to efficiently calculate the neighborhood of an object based on its measurable radial distance in the image plane. On the basis of the parameterized nonlinear model, a distortion-invariant fragment-based joint-feature mixture model of Gaussian is presented for human target tracking in omnidirectional vision. Under the framework of Gaussian Mixture Model, the problem of feature matching is converted into the feature clustering. The joint probability distribution of a joint-feature class is modeled by a mixture of Gaussian. A weight contribution mechanism is designed to flexibly weight the fragments contribution based on their responses, which leads to a robust tracking even under serious partial occlusion. Finally, experiments validate the advantage of the proposed algorithm over other conventional approaches. Catadioptric omnidirectional cameras have been widely used in robotics and surveillance fields for visual sensing due to its big field-of-view. However, conventional visual models use large-scale statistical sampling for feature extraction in catadioptric sensor, which may consume lot of computational cost. For practical applications, a parameterized model that can accurately and efficiently formulate distortion of catadioptric image is desirable. Integrating of the priori of system, a parameterized neighborhood model is presented to directly extract distorted image content in image, which can significantly improve the efficiency of algorithm. To robustly handle challenging occlusion in the distorted image, a flexible fragment-based joint-feature framework is presented for robust non-rigid human target tracking. Compared with the conventional tracking methods applied to catadioptric vision, the proposed tracking approaches leads to much better performance from the perspective of efficiency and robustness. Yazhe Tang, Youfu Li 0001, Shuzhi Sam Ge, Jun Luo 0006, Hongliang Ren 0001 |
IEEE Trans Autom. Sci. Eng. | 2 |
| 2016 | On Integral Invariants for Effective 3-D Motion Trajectory Matching and RecognitionabstractMotion trajectories tracked from the motions of human, robots, and moving objects can provide an important clue for motion analysis, classification, and recognition. This paper defines some new integral invariants for a 3-D motion trajectory. Based on two typical kernel functions, we design two integral invariants, the distance and area integral invariants. The area integral invariants are estimated based on the blurred segment of noisy discrete curve to avoid the computation of high-order derivatives. Such integral invariants for a motion trajectory enjoy some desirable properties, such as computational locality, uniqueness of representation, and noise insensitivity. Moreover, our formulation allows the analysis of motion trajectories at a range of scales by varying the scale of kernel function. The features of motion trajectories can thus be perceived at multiscale levels in a coarse-to-fine manner. Finally, we define a distance function to measure the trajectory similarity to find similar trajectories. Through the experiments, we examine the robustness and effectiveness of the proposed integral invariants and find that they can capture the motion cues in trajectory matching and sign recognition satisfactorily. Zhanpeng Shao, Youfu Li 0001 |
IEEE Trans. Cybern. | 2 |
| 2015 | Distortion invariant joint-feature for visual tracking in catadioptric omnidirectional visionabstractCentral catadioptric omnidirectional images exhibit serious nonlinear distortions due to quadratic mirrors involved. Conventional visual features developed based on the perspective model are hard to achieve a satisfactory performance when directly applied to the distorted omnidirectional image. This paper presents a parameterized neighborhood model to efficiently calculate the adaptive neighborhood of an object based on the measurable radial distance in image plane. On the basis of the parameterized neighborhood model, a distortion invariant joint-feature framework implemented with contour-color fragment mixture model of Gaussian is proposed for visual tracking in catadioptric omnidirectional camera system. Under the framework of Gaussian Mixture Model, the problem of feature matching is converted into feature clustering. A weight contribution mechanism is presented to flexibly weight the fragments based on their responses, which makes the system robustly guided by limited visible fragments even when serious partial occlusion happens. The experiments validate the performance of the proposed algorithm. Yazhe Tang, Youfu Li 0001, Shuzhi Sam Ge, Jun Luo 0006, Hongliang Ren 0001 |
ICRA | 2 |
| 2015 | Flexible Trajectory Indexing for 3D Motion RecognitionabstractMotion trajectory analysis is important for human motion recognition and human computer interaction. In this paper, we propose a flexible 3D trajectory indexing method for complex 3D motion recognition. Based on both point level and primitive-level descriptors, trajectories are represented in the sub-primitive level, the level between the point level and primitive level. Primitives are flexibly segmented into sub-primitives in various scales, and the sub-primitives retain more detailed information than primitives. The detailed level of sub-primitives can be adjusted by controlling segmentation scales according to motion complexities. The proposed approach is suitable for spatial motion trajectory, which is view-invariant in 3D space. A cluster model is also proposed to represent motion classes and motion recognition performed based on maximum a posteriori (MAP) criterion. The experiments on benchmark datasets validate the effectiveness of the proposed approach. Jianyu Yang 0002, Junsong Yuan 0001, Youfu Li 0001 |
WACV | 3 |
| 2015 | Integral invariants for space motion trajectory matching and recognition
Zhanpeng Shao, Youfu Li 0001 |
Pattern Recognit. | 2 |
| 2015 | Learning Local Appearances With Sparse Representation for Robust and Fast Visual TrackingabstractIn this paper, we present a novel appearance model using sparse representation and online dictionary learning techniques for visual tracking. In our approach, the visual appearance is represented by sparse representation, and the online dictionary learning strategy is used to adapt the appearance variations during tracking. We unify the sparse representation and online dictionary learning by defining a sparsity consistency constraint that facilitates the generative and discriminative capabilities of the appearance model. An elastic-net constraint is enforced during the dictionary learning stage to capture the characteristics of the local appearances that are insensitive to partial occlusions. Hence, the target appearance is effectively recovered from the corruptions using the sparse coefficients with respect to the learned sparse bases containing local appearances. In the proposed method, the dictionary is undercomplete and can thus be efficiently implemented for tracking. Moreover, we employ a median absolute deviation based robust similarity metric to eliminate the outliers and evaluate the likelihood between the observations and the model. Finally, we integrate the proposed appearance model with the particle filter framework to form a robust visual tracking algorithm. Experiments on benchmark video sequences show that the proposed appearance model outperforms the other state-of-the-art approaches in tracking performance. Tianxiang Bai, Youfu Li 0001, Xiaolong Zhou 0001 |
IEEE Trans. Cybern. | 2 |
| 2015 | An Exponential-Rayleigh Model for RSS-Based Device-Free Localization and TrackingabstractA common technical difficulty in device-free localization and tracking (DFLT) with a wireless sensor network is that the change of the received signal strength (RSS) of the link often becomes more unpredictable due to the multipath interferences. This challenge can lead to unsatisfactory or even unstable DFLT performance. This work focuses on developing a new RSS model, called Exponential-Rayleigh (ER) model, for addressing this challenge. Based on data from our extensive experiments, we first develop the ER model of the received signal strength. This model consists of two parts: the large-scale exponential attenuation part and the small-scale Rayleigh enhancement part. The new consideration on using the Rayleigh model is to depict the target-induced multipath components. We then explore the use of the ER model with a particle filter in the context of multi-target localization and tracking. Finally, we experimentally demonstrate that our ER model outperforms the existing models. The experimental results highlight the advantages of using the Rayleigh model in mitigating the multipath interferences thus improving the DFLT performance. Yao Guo 0002, Kaide Huang, Nanyong Jiang, Xuemei Guo, Youfu Li 0001, Guoli Wang 0001 |
IEEE Trans. Mob. Comput. | 5 |
| 2014 | Online visual object tracking with supervised sparse representation and learningabstractIn this paper, an online visual object tracking algorithm based on the discriminative sparse representation framework with supervised learning is proposed. Different from the generative sparse representation based tracking algorithms, the proposed method casts the tracking problem into a binary classification task. A linear classifier is embedded into the sparse representation model by incorporating the classification error into the objective function to achieve discriminative classification. The dictionary and the classifier are jointly trained using the online dictionary learning algorithm, thus allow the model can adapt the dynamic variations of target appearance and background environment. The target locations are updated based on the classification score and the greedy search motion model. The proposed method is evaluated using four benchmark datasets and is compared with three state-of-the-art tracking algorithms. The results show that the discriminative sparse representation facilitates the tracking performance. Tianxiang Bai, Youfu Li 0001, Zhanpeng Shao |
ICARCV | 2 |
| 2014 | Radial distortion invariants and lens evaluation under a single-optical-axis omnidirectional camera
Yihong Wu 0002, Zhanyi Hu, Youfu Li 0001 |
Comput. Vis. Image Underst. | 3 |
| 2014 | Entropy distribution and coverage rate-based birth intensity estimation in GM-PHD filter for multi-target visual tracking
Xiaolong Zhou 0001, Youfu Li 0001, Bingwei He |
Signal Process. | 2 |
| 2014 | Robust Visual Tracking Using Flexible Structured Sparse RepresentationabstractIn this work, we propose a robust and flexible appearance model based on the structured sparse representation framework. In our method, we model the complex nonlinear appearance manifold and the occlusion as a sparse linear combination of structured union of subspaces in a basis library, which consists of multiple incremental learned target subspaces and a partitioned occlusion template set. In order to enhance the discriminative power of the model, a number of clustered background subspaces are also added into the basis library and updated during tracking. With the Block Orthogonal Matching Pursuit (BOMP) algorithm, we show that the new flexible structured sparse representation based appearance model facilitates the tracking performance compared with the prototype structured sparse representation model and other state of the art tracking algorithms. Tianxiang Bai, Youfu Li 0001 |
IEEE Trans. Ind. Informatics | 2 |
| 2014 | GM-PHD-Based Multi-Target Visual Tracking Using Entropy Distribution and Game TheoryabstractTracking multiple moving targets in a video is a challenge because of several factors, including noisy video data, varying number of targets, and mutual occlusion problems. The Gaussian mixture probability hypothesis density (GM-PHD) filter, which aims to recursively propagate the intensity associated with the multi-target posterior density, can overcome the difficulty caused by the data association. This paper develops a multi-target visual tracking system that combines the GM-PHD filter with object detection. First, a new birth intensity estimation algorithm based on entropy distribution and coverage rate is proposed to automatically and accurately track the newborn targets in a noisy video. Then, a robust game-theoretical mutual occlusion handling algorithm with an improved spatial color appearance model is proposed to effectively track the targets in mutual occlusion. The spatial color appearance model is improved by incorporating interferences of other targets within the occlusion region. Finally, the experiments conducted on publicly available videos demonstrate the good performance of the proposed visual tracking system. Xiaolong Zhou 0001, Youfu Li 0001, Bingwei He, Tianxiang Bai |
IEEE Trans. Ind. Informatics | 2 |
| 2013 | Learning appearance manifolds with structured sparse representation for robust visual trackingabstractThis paper presents a novel algorithm for robust visual object tracking based on the structured sparse representation framework. Conventional structured sparse representation based tracker models the nonlinear appearance manifold with a single subspace that is difficult to handle significant pose and illumination changes. Different from the afore-mentioned method, the proposed algorithm approximates the nonlinear appearance manifold by multiple low dimensional subspaces computed by an incremental learning scheme based on the merging and insert strategy. In order to enhance the discriminative power of the model, a number of clustered background subspaces are also added into the basis library and updated during tracking. With the Block Orthogonal Matching Pursuit (BOMP) algorithm, we show that the complex nonlinear appearance manifold can effectively represent by a sparse linear combination of structured union of subspaces. Experiments on benchmark video sequences show that the new structured sparse representation model improves the robustness of tracking. Tianxiang Bai, Youfu Li 0001, Zhanpeng Shao |
ICRA | 2 |
| 2013 | A new descriptor for multiple 3D motion trajectories recognitionabstractMotion trajectory gives a meaningful and informative clue in characterizing the motions of human, robots or moving objects. Hence, the descriptor for motion trajectory plays an importance role in motion recognition for many robotic tasks. However, an effective and compact descriptor for multiple 3D motion trajectories under complicated situation is lacking. In this paper, we propose a novel invariant descriptor for multiple motion trajectories based on the kinematic relation among multiple moving parts. There are two kinds of kinematic relation among multiple trajectories: articulated and independent trajectories. Spherical coordinate system is introduced to get a uniform and compact representation for both kinds of trajectories, where the relative trajectory concept are firstly defined based on orientation and distance changes in favor of acquiring relative movement features of each child trajectory with respect to the root trajectory. Then, by incorporating both the differential invariants of root trajectory and orientation, distance variations of each relative trajectory respectively, the new descriptor is constructed. Finally, effectiveness and robustness of our proposed new descriptor for multiple trajectories under complex circumstance are validated by the conducted two experiments for sign language and human action recognition. Zhanpeng Shao, Youfu Li 0001 |
ICRA | 2 |
| 2013 | Multi-target visual tracking with game theory-based mutual occlusion handlingabstractTracking multiple moving targets in video is still a challenge because of mutual occlusion problem. This paper presents a Gaussian mixture probability hypothesis density-based visual tracking system with game theory-based mutual occlusion handling. First, a two-step occlusion reasoning algorithm is proposed to determine the occlusion region. Then, the spatial constraint-based appearance model with other interacting targets¶ interferences is modeled. Finally, an n-person, non-zero-sum, non-cooperative game is constructed to handle the mutual occlusion problem. The individual measurements within the occlusion region are regarded as the players in the constructed game competing for the maximum utilities by using the certain strategies. The Nash Equilibrium of the game is the optimal estimation of the locations of the players within the occlusion region. Experiments conducted on publicly available videos demonstrate the good performance of the proposed occlusion handling algorithm. Xiaolong Zhou 0001, Youfu Li 0001, Bingwei He, Tianxiang Bai |
IROS | 2 |
| 2013 | Adaptive weighted learning for linear regression problems via Kullback-Leibler divergence
Zhizheng Liang, Youfu Li 0001, Shixiong Xia |
Pattern Recognit. | 2 |
| 2013 | Game-theoretical occlusion handling for multi-target visual tracking
Xiaolong Zhou 0001, Youfu Li 0001, Bingwei He |
Pattern Recognit. | 2 |
| 2013 | Feature extraction based on Lp-norm generalized principal component analysis
Zhizheng Liang, Shixiong Xia, Yong Zhou 0003, Lei Zhang 0029, Youfu Li 0001 |
Pattern Recognit. Lett. | 5 |
| 2013 | Finding Optimal Focusing Distance and Edge Blur Distribution for Weakly Calibrated 3-D Visionabstract3-D Vision is now a common sensing method frequently used in industrial applications. With the convenience of an uncalibrated system, 3-D reconstruction by a self-calibration technique is possible, but always incomplete or unreliable. This paper presents a novel method to analyze the blur distribution in an image and find the optimal focusing distance so that additional constraints can be used to generate absolute measurement of the models. With the assumption of a Gaussian distribution model of the point spread function, this paper applies two theorems to efficiently compute the defocusing extent on stripe edges. Because the blurring diameter implies the distance from the sensor to the surface, we can upgrade the 3-D map obtained from self-calibration with the known scaling factor. Through theoretical and experimental analysis, we find that not only the technology is feasible, but also both the accuracy and the efficiency are satisfactory. Shengyong Chen, Youfu Li 0001 |
IEEE Trans. Ind. Informatics | 2 |
| 2012 | Robust and fast visual tracking using constrained sparse coding and dictionary learningabstractWe present a novel appearance model using sparse coding with online sparse dictionary learning techniques for robust visual tracking. In the proposed appearance model, the target appearance is modeled via online sparse dictionary learning technique with an “elastic-net constraint”. This scheme allows us to capture the characteristics of the target local appearance, and promotes the robustness against partial occlusions during tracking. Additionally, we unify the sparse coding and online dictionary learning by defining a “sparsity consistency constraint” that facilitates the generative and discriminative capabilities of the appearance model. Moreover, we propose a robust similarity metric that can eliminate the outliers from the corrupted observations. We then integrate the proposed appearance model with the particle filter framework to form a robust visual tracking algorithm. Experiments on publicly available benchmark video sequences demonstrate that the proposed appearance model improves the tracking performance compared with other state-of-the-art approaches. Tianxiang Bai, Youfu Li 0001, Xiaolong Zhou 0001 |
IROS | 2 |
| 2012 | Birth intensity online estimation in GM-PHD filter for multi-target visual trackingabstractMulti-target tracking in video is a challenge due to noisy video data, varying number of targets, and the data association problems. In this paper, a multi-target visual tracking system that incorporates object detection with the Gaussian mixture PHD filter is developed. The main contribution of this paper is to propose a new birth intensity online estimation method that based on the entropy distribution and the coverage rate. First, the birth intensity is initialized by using the previously obtained targets' states and measurements. The measurements are obtained by object detection and classified into the birth measurements and the survival measurements. Then it is updated according to the currently obtained birth measurements. In the update stage, the instability of the entropy distribution is applied to remove components like noises within the birth intensity which are irrelevant with the currently obtained birth measurements. And the coverage rate between each birth intensity component and corresponding birth measurement is computed to further eliminate the noises. Finally, experiments are implemented to show the performance of the proposed visual tracking system, especially to show the good performance for tracking the newborn targets. Xiaolong Zhou 0001, Youfu Li 0001, Bingwei He, Tianxiang Bai, Yazhe Tang |
IROS | 2 |
| 2012 | Robust visual tracking with structured sparse representation appearance model
Tianxiang Bai, Youfu Li 0001 |
Pattern Recognit. | 2 |
| 2012 | A Hierarchical Model Incorporating Segmented Regions and Pixel Descriptors for Video Background SubtractionabstractBackground subtraction is important for detecting moving objects in videos. Currently, there are many approaches to performing background subtraction. However, they usually neglect the fact that the background images consist of different objects whose conditions may change frequently. In this paper, a novel hierarchical background model is proposed based on segmented background images. It first segments the background images into several regions by the mean-shift algorithm. Then, a hierarchical model, which consists of the region models and pixel models, is created. The region model is a kind of approximate Gaussian mixture model extracted from the histogram of a specific region. The pixel model is based on the cooccurrence of image variations described by histograms of oriented gradients of pixels in each region. Benefiting from the background segmentation, the region models and pixel models corresponding to different regions can be set to different parameters. The pixel descriptors are calculated only from neighboring pixels belonging to the same object. The experimental results are carried out with a video database to demonstrate the effectiveness, which is applied to both static and dynamic scenes by comparing it with some well-known background subtraction methods. Shengyong Chen, Jianhua Zhang 0002, Youfu Li 0001, Jianwei Zhang 0001 |
IEEE Trans. Ind. Informatics | 3 |
| 2012 | Overall Well-Focused Catadioptric Image Acquisition With Multifocal Images: A Model-Based MethodabstractWhen a catadioptric imaging system suffers from limited depth of field, a single image can not capture all the objects with clear focus. To solve this problem, a set of multi-focal images can be used to extend the depth of field by fusing the best focused image regions to an overall well focused composite image. In this paper, we proposed a novel model based method that avoids the computational cost problem of previous extended-depth-of-field algorithms. Based on the special optical geometry properties of catadioptric systems, the proposed model describes the shapes of the best focused image regions in multi-focal images by a series of neighboring concentric annuluses. Then we proposed a method to estimate the model parameters. Based on this model, an overall well focused image is obtained by combining the best focused regions with a fast and reliable online operation. Experiments on catadioptric images of a variety of different scenes and camera settings verify the validity of the model and the robust performance of the proposed method. Youfu Li 0001 |
IEEE Trans. Image Process. | 2 |
| 2011 | Structured sparse representation appearance model for robust visual trackingabstractWe propose a robust visual tracker based on structured sparse representation appearance model. The appearance of tracking target is modeled as a sparse linear combination of Eigen templates plus a sparse error due to occlusions. We address the structured sparse representation that preferably matches the practical visual tracking problem by taking the contiguous spatial distribution of occlusion into account. The sparsity is achieved by Block Orthogonal Matching Pursuit (BOMP) for solving structured sparse representation problem more efficiently. The model update scheme, based on incremental Singular Value Decomposition (SVD), guarantees the Eigen templates that are able to capture the variations of target appearance online. Then the approximation error is adopted to build a probabilistic observation model that integrates with a stochastic affine motion model to form a particle filter framework for visual tracking. Thanks to the block structure of sparse representation and BOMP, our proposed tracker demonstrates superiority on both efficiency and robustness improvement in comparison experiments with publicly available benchmark video sequences. Tianxiang Bai, Youfu Li 0001, Yazhe Tang |
ICRA | 2 |
| 2011 | Microscopic photometric stereo: A dense microstructure 3D measurement methodabstractWe propose a novel Microscopic Photometric Stereo (MPS) method based on optical microscope and uncalibrated photometric stereo (UPS) for dense and refined microstructure 3D measurement. The UPS method, which does not require a priori knowledge of the light-source direction or the light-source intensity, is employed to recover surface normals and albedos from the captured micro-images. For resolving the inherent generalized bas-relief (GBR) ambiguity of the UPS, we present a GBR disambiguation method based on a framework of entropy minimization, and extend it using a graph cuts energy minimization to decrease the influence of noise and further refine the recovered surface normal. The proposed MPS method has been tested on synthetic as well as real images and very encouraging results have been obtained. Youfu Li 0001 |
ICRA | 2 |
| 2011 | An analytical solution to optimal focal distance in catadioptric imaging systemsabstractCatadioptric imaging systems are important in many computer vision and robotics applications. This work addresses the issue of optimally setting the focal distance of a lens based camera in a catadioptric imaging system in order to acquire a best focused image. To this end, understanding the spatial distribution of virtual feature formed by mirror reflection is important. It is known that the virtual features of infinite range of scene depth are limited to a finite depth extent, which is named a caustic volume. In this work we further find that for a variety of quadric mirror based catadioptric systems, when the objects are located at a certain distance to the system, the corresponding virtual features can be considered to be located on the caustic volume boundary. We verify this property with real catadioptric images. Based on this property, an analytical solution is derived for the optimal focal distance setting, which can only be calculated by software simulation or numerical approaches in previous work. This solution is compared with a numerical solution of previous method and is also verified by a simulation of the optical process. Youfu Li 0001 |
ICRA | 2 |
| 2011 | Generic radial distortion calibration of a novel single camera based panoramic stereoscopic systemabstractThis work presents a novel panoramic stereoscopic system consisting of a fisheye lens camera and a hyperbolic mirror with co-axis installation. From the overlapping field of view captured through the fisheye lens and the reflection of the mirror, position of an object point in 3D Euclidean space can be reconstructed once the system geometry is calibrated. To deal with the non-single viewpoint issue in the catadioptric image, a generic radial distortion model is used to describe the imaging process with a series of viewing cones. The parameters of the viewing cones are estimated using a homography based method with observations of an LCD panel at a few unknown positions. Following this, a closed form solution for 3D reconstruction is used with a non-linear optimization to obtain an optimal calibration. A prototype of the proposed design is constructed. Quantitative experiments are conducted to evaluate the calibration result in terms of 3D reconstruction precision. With the calibration result, we also present potential robotic applications of the proposed system such as 3D environment reconstruction in a 360 degree horizontal field of view. Youfu Li 0001, Dong Sun 0001, Beiwei Zhang 0001 |
ICRA | 2 |
| 2011 | Invariant trajectory indexing for real time 3D motion recognitionabstractMotion trajectory description plays an important role in motion recognition for many robotic tasks. The previous methods are too high in computational cost for real time recognition. In this paper, we propose a novel indexing method for motion trajectories as a high level invariant descriptor. Trajectories are segmented into basic segments and represented in segment level rather than point level. We index the trajectories by their segment sequences to recognize them. The computational cost is significantly decreased and the experimental results show that this method is effective and it can be used for real time motion recognition. The accuracy is also preserved for long term trajectories. Youfu Li 0001 |
IROS | 2 |
| 2011 | Blockwise projection matrix versus blockwise data on undersampled problems: Analysis, comparison and applications
Zhizheng Liang, Shixiong Xia, Yong Zhou 0003, Youfu Li 0001 |
Pattern Recognit. | 4 |
| 2010 | A regularization framework for robust dimensionality reduction with applications to image reconstruction and feature extraction
Zhizheng Liang, Youfu Li 0001 |
Pattern Recognit. | 2 |
| 2010 | Motion trajectory reproduction from generalized signature description
Shandong Wu, Youfu Li 0001 |
Pattern Recognit. | 2 |
| 2010 | Projected gradient method for kernel discriminant nonnegative matrix factorization and the applications
Zhizheng Liang, Youfu Li 0001, Tuo Zhao |
Signal Process. | 2 |
| 2009 | A Model Based Method for Overall Well Focused Catadioptric Image Acquisition with Multi-focal Images
Youfu Li 0001, Yihong Wu 0002 |
CAIP | 2 |
| 2009 | A Novel View Planning Method for Automatic Reconstruction of Unknown 3-D Objects Based on the Limit Visual SurfaceabstractAutomatic reconstruction of unknown 3-D objects has been of great importance in the areas of machine vision, object recognition, and automatic modeling. In this paper, a new planning approach of generating 3-D models automatically is proposed. The new algorithm incorporates the limit visual surfaces of unknown model which are obtained according to both of the known object boundary knowledge and the visual region of the vision system and selects the suitability of viewpoints as the next best view on scanning coverage. The limit visual surfaces are used to predict the maximal information of unknown model and then the visibility criterion of next viewpoint is determined. And the position which can obtain the maximal visual surface area is defined as the next best view position. The reconstruction result of real model with proposed method show the efficiency in practical implementation. Xiaolong Zhou 0001, Bingwei He, Youfu Li 0001 |
ICIG | 3 |
| 2009 | Step function based turning maneuvers in biomimetic robotic fishabstractThis paper presents a new turning maneuver generation method for a multilink biomimetic robotic fish, in which smooth step functions are introduced to dynamically trigger directed offsets in active and asymmetric swimming. With the proposed method, three basic turning modes can be unified into a general framework by choosing appropriate step-function combinations and dynamic bias. Furthermore, this method can be employed to maneuver the robotic fish agilely in the path planning, which promises more flexibility and steadiness in potential applications to bio-inspired autonomous underwater vehicles. Junzhi Yu 0001, Ming Wang 0001, Min Tan 0001, Youfu Li 0001 |
ICRA | 4 |
| 2009 | Probabilistic Cluster Signature for Modeling Motion ClassesabstractIn this paper, a novel 3-D motion trajectory signature is introduced to serve as an effective description to the raw trajectory. More importantly, based on the trajectory signature, a probabilistic model-based cluster signature is further developed for modeling a motion class. The cluster signature is a mixture model-based motion description that is useful for motion class perception, recognition and to benefit a generalized robot task representation. The signature modeling process is supported by integrating the EM and IPRA algorithms. The conducted experiments verified the cluster signature's effectiveness. Shandong Wu, Youfu Li 0001, Jianwei Zhang 0001 |
IROS | 2 |
| 2009 | Incremental support vector machine learning in the primal and applications
Zhizheng Liang, Youfu Li 0001 |
Neurocomputing | 2 |
| 2009 | Flexible signature descriptions for adaptive motion trajectory representation, perception and recognition
Shandong Wu, Youfu Li 0001 |
Pattern Recognit. | 2 |
| 2009 | Dynamic View Planning by Effective Particles for Three-Dimensional TrackingabstractIn this paper, we propose a new approach to dynamically manage the viewpoint of a vision system for optimal 3-D tracking using particle techniques. We adopt the effective sample size in the proposed particle filter as a criterion for evaluating tracking performance and employ it to guide the view-planning process for finding the best viewpoint configuration. In our approach, the vision system is designed and configured to achieve the largest number of effective particles, which minimizes tracking error by revealing the system to a better swarm of importance samples and interpreting posterior states in a better way. Superiorities of our method are shown by comparison with the resampling particle filter and other view-planning methods. Youfu Li 0001 |
IEEE Trans. Syst. Man Cybern. Part B | 2 |
| 2008 | A hierarchical motion trajectory signature descriptorabstractMotion trajectory is a compact clue for motion characterization. However, it is normally used directly in its raw data form in most work and effective trajectory description is lacking. In this paper, we propose a novel hierarchical motion trajectory signature descriptor, which can not only fully capture motion features for detailed perception, but also can be used for probabilistic fast recognition. The hierarchy enables the signature to exhibit high functional adaptability meeting different application requirements. At the first-level, differential invariants are employed to describe trajectory features and a nonlinear signature warping method is developed to perceive and recognize trajectories. The second-level signature is the condensation of the first-level signature by applying PCA based dimension optimization. It behaves more efficiently in recognition based on the Gaussian Mixture modeling and Bayesian classifier. The conducted experiments verified the signature’s effectiveness. Shandong Wu, Youfu Li 0001, Jianwei Zhang 0001 |
ICRA | 2 |
| 2008 | The detection of multiple moving objects using fast level set methodabstractA novel method for the detection of multiple moving objects is proposed in this paper. In order to get the detailed information of the objects, fast level set method which is only based on the evolution of single link list is mainly used to detect the boundaries of moving objects. The whole process consists of two main procedures: the coarse detection and the fine localization. During the coarse detection procedure, the velocity function is defined according to the modified Otsu method which is more effective to eliminate the split phenomena of the whole motion area and get the consecutive boundaries. As to the fine localization, improved region competition is applied to obtain the smooth and exact contours. The proposed method has been tested on several different video sequences, and the efficiency of the method has been verified. Youfu Li 0001 |
IJCNN | 3 |
| 2008 | Invariant signature description and trajectory reproduction for robot Learning by DemonstrationabstractIn most reported works about robot learning by demonstration (LbD), the demonstration is normally limited to simple gestures or grasp actions. In this paper, motion trajectory oriented LbD is studied in which free form 3-D motion trajectory is extracted to characterize certain human demonstrations. We propose to build effective description to motion trajectories to be learned by a robot instead of learning the raw trajectory data. A novel signature descriptor is formulated which serves as a generic and invariant description for motion trajectories. More importantly, a trajectory reproduction algorithm based on the learned signature is investigated to enable a robot to repeat/follow the reproduced trajectory instance. Experiments are reported to show the signature description and the reproduction algorithm for further application to the LbD. Shandong Wu, Youfu Li 0001, Jianwei Zhang 0001 |
IROS | 2 |
| 2008 | Detecting and Handling Unreliable Points for Camera Parameter Estimation
Yihong Wu 0002, Youfu Li 0001, Zhanyi Hu |
Int. J. Comput. Vis. | 2 |
| 2008 | A general recursive linear method and unique solution pattern design for the perspective-n-point problem
De Xu, Youfu Li 0001, Min Tan 0001 |
Image Vis. Comput. | 2 |
| 2008 | A note on two-dimensional linear discriminant analysis
Zhizheng Liang, Youfu Li 0001 |
Pattern Recognit. Lett. | 2 |
| 2008 | Vision Processing for Realtime 3-D Data Acquisition Based on Coded Structured LightabstractStructured light vision systems have been successfully used for accurate measurement of 3-D surfaces in computer vision. However, their applications are mainly limited to scanning stationary objects so far since tens of images have to be captured for recovering one 3-D scene. This paper presents an idea for real-time acquisition of 3-D surface data by a specially coded vision system. To achieve 3-D measurement for a dynamic scene, the data acquisition must be performed with only a single image. A principle of uniquely color-encoded pattern projection is proposed to design a color matrix for improving the reconstruction efficiency. The matrix is produced by a special code sequence and a number of state transitions. A color projector is controlled by a computer to generate the desired color patterns in the scene. The unique indexing of the light codes is crucial here for color projection since it is essential that each light grid be uniquely identified by incorporating local neighborhoods so that 3-D reconstruction can be performed with only local analysis of a single image. A scheme is presented to describe such a vision processing method for fast 3-D data acquisition. Practical experimental performance is provided to analyze the efficiency of the proposed methods. Shengyong Chen, Youfu Li 0001, Jianwei Zhang 0001 |
IEEE Trans. Image Process. | 2 |
| 2008 | A New Active Visual System for Humanoid RobotsabstractIn this paper, a new active visual system is developed, which is based on bionic vision and is insensitive to the property of the cameras. The system consists of a mechanical platform and two cameras. The mechanical platform has two degrees of freedom of motion in pitch and yaw, which is equivalent to the neck of a humanoid robot. The cameras are mounted on the platform. The directions of the optical axes of the two cameras can be simultaneously adjusted in opposite directions. With these motions, the object's images can be located at the centers of the image planes of the two cameras. The object's position is determined with the geometry information of the visual system. A more general model for active visual positioning using two cameras without a neck is also investigated. The position of an object can be computed via the active motions. The presented model is less sensitive to the intrinsic parameters of cameras, which promises more flexibility in many applications such as visual tracking with changeable focusing. Experimental results verify the effectiveness of the proposed methods. De Xu, Youfu Li 0001, Min Tan 0001 |
IEEE Trans. Syst. Man Cybern. Part B | 2 |
| 2007 | Image Segmentation Algorithm using Watershed Transform and Level Set MethodabstractA novel method for image segmentation is proposed in this paper, which combines the watershed transform and region-based level set method. The watershed transform is first used to presegment the image so as to get the initial partition of it. Some useful information of the primitive regions and boundaries can be obtained. The region-based level set method is then applied for extracting the boundaries of objects on the basis of the presegmentation. The consumed time does not depend on the size of the image but the number of presegmented regions because only label level set function is updated instead of the level set function for each pixel. Therefore, the proposed method is computationally efficient. Moreover, the algorithm can localize the boundary of the regions exactly due to the edges obtained by the watersheds. The efficiency and accuracy of the algorithm is demonstrated by the experiments on the MR brain images. Miaomiao Liu 0001, Youfu Li 0001 |
ICASSP (1) | 3 |
| 2007 | Realtime Structured Light Vision with the Principle of Unique Color CodesabstractTo date, several successful structured light vision systems for accurate 3D measurement in machine vision have been set up. However, these are usually limited to scanning stationary objects or static environments since tens of images have to be captured for recovering one 3D scene, which results in the industry largely avoiding this technology. This paper presents a method of grid-pattern design based on the principles of uniquely color-encoded structured light, to improve the reconstruction efficiency for real-time processing. For a live scene, the 3D measurement is desired to only capture a single image. To realize this, an important problem for the color-encoded projection is the unique indexing of the color codes in the image. It is essential that each light grid be uniquely identified by incorporating the local neighborhoods in the pattern so that 3D reconstruction can be performed with only local analysis of a single image. This paper describes such a method in the design of the special grid patterns and its corresponding 3D reconstruction method for fast vision perception. Shengyong Chen, Youfu Li 0001, Jianwei Zhang 0001 |
ICRA | 2 |
| 2007 | Active Illumination for Robot VisionabstractA vision sensor is the robot's eye to perceive its environment, but the perception performance can be significantly affected by illumination conditions. This paper presents strategies of adaptive illumination control for robot vision to achieve the best scene interpretation. It investigates how to obtain the most comfortable illumination conditions for a vision sensor. In a "comfort" condition the image reflects the natural properties of the concerned object. "Discomfort" may occur if some scene information is lost. Strategies are proposed to optimize the pose and optical parameters of the luminaire and the sensor, with emphasis on controlling the intensity and avoiding glare. Shengyong Chen, Jianwei Zhang 0001, Houxiang Zhang, Wanliang Wang, Youfu Li 0001 |
ICRA | 5 |
| 2007 | A Study on the Relative Pose Problem in an Active Vision System with varying Focal LengthsabstractIn this work, the relative pose problem is addressed in our structured light system. Assuming that there is an arbitrary planar structure in the scene, we suggest a method for estimating the rotation matrix and translation vector between the camera and the projector. In this system, the varying focal lengths of the camera are allowed and can be obtained without any further assumptions. Finally, we give some experimental results to validate this method. Youfu Li 0001 |
ICRA | 2 |
| 2007 | Self-recalibration of a structured light system via plane-based homography
Youfu Li 0001 |
Pattern Recognit. | 2 |
| 2007 | Support Vector Networks in Adaptive Friction CompensationabstractThis paper presents our research on how support vector regression (SVR) and parametric adaptive learning, which are normally used independently, can be exploited together to benefit adaptive neural control. In the context of friction compensation for servo-motion control systems, we present the notion of support vector networks which play an essential role in combining SVR and adaptive neural network (NN) in cooperation for friction estimation. The analysis shows that the proposed support vector network contributes not only to the performance improvement but also to the practical usefulness in adaptive friction compensation. Experimental results are reported to demonstrate the effectiveness of the proposed approach. Guoli Wang 0001, Youfu Li 0001, D. X. Bi |
IEEE Trans. Neural Networks | 2 |
| 2006 | Support Vector Network Enhanced Adaptive Friction CompensationabstractThis paper explores the notation of support vector networks, a new paradigm of combining support vector regression (SVR) parametrization with adaptive neural mechanism, in friction compensation for servo-motion systems. The contribution of this work is twofold. The first is to develop an enhanced adaptive friction compensator via SVR parametrization; the second is to present an analysis that shows the evidences of the performance improvement and practical usefulness enhancement due to SVR parametrization. The experimental study was conducted to validate the proposed method Guoli Wang 0001, Youfu Li 0001, D. X. Bi |
ICRA | 2 |
| 2006 | A Focal Cue for Metric Measurement of 3D SurfacesabstractThis paper finds a method for computing the best-focused location from an image and using it as a dimensional cue for acquisition of a 3D scene surface. In some situations in 3D vision, an object cannot be reconstructed into a 3D model with metric dimensions. Rather, it can only be reconstructed into a 3D structure up to a similarity transformation. To upgrade the 3D model from a similarity transformation to a Euclidean transformation, we propose a method based on the best-focused locations. By analyzing the blur distribution in an image, this method finds the best-focused locations from an image, which provides an additional cue for upgrading the reconstructed 3D structure. Hence, we can obtain not only the object's shape, but also the dimensions and sizes of surface features Shengyong Chen, Youfu Li 0001, Jianwei Zhang 0001 |
IROS | 2 |
| 2006 | Easy Calibration for Para-catadioptric-like CameraabstractFor omnidirectional cameras, most of the previous calibration methods from lines use conic fitting. This paper presents a calibration method for para-catadioptric-like cameras from lines without conic fitting under a single view. We establish equations on the five camera intrinsic parameters. These equations are linear for the focal lengths and skew factor once the principal point is known. The principal point can be approximated well by the center of the imaged mirror contour in practice or can be accurately estimated by quadric equations. After obtaining the principal point, we propose an algorithm to calibrate the focal lengths and skew factor. The algorithm needs neither prior structure knowledge nor conic fitting and is linear, which make it easy to implement. Other omnidirectional cameras can also use this presented work if high accuracy is not required. Experiments demonstrate the efficiency of the proposed algorithm. Yihong Wu 0002, Youfu Li 0001, Zhanyi Hu |
IROS | 2 |
| 2006 | A Visual Positioning Method Based on Relative Orientation Detection for Mobile RobotsabstractIn this paper, a new visual positioning method based on corresponding points at two adjacent views is developed for mobile robots. A camera is mounted on a wheeled mobile robot with nonholonomic constraints. The camera whose intrinsic parameters are well calibrated can rotate around an axis perpendicular to the ground plane. The relative orientation and scaled position offsets of the mobile robot are computed from the corresponding points in the common part of the views despite their unknown positions in Cartesian space. The relative orientation is used in a simple visual dead reckoning method to modify the odometry information. Then, based on the relative orientation and modified odometry information, equations are given to determine the position and orientation of the mobile robot. Experiments are performed to verify the effectiveness of the proposed methods De Xu, Youfu Li 0001, Min Tan 0001 |
IROS | 2 |
| 2006 | Deep compression of remotely rendered viewsabstractThree-dimensional (3-D) models are information-rich and provide compelling visualization effects. However downloading and viewing 3-D scenes over the network may be excessive. In addition low-end devices typically have insufficient power and/or memory to render the scene interactively in real-time. Alternatively,3-D image warping, an image-based-rendering technique that renders a two-dimensional(2-D) depth view to form new views intended from different viewpoints and/or orientations, may be employed on a limited device. In a networked 3-D environment,the warped views may be further compensated by the graphically rendered views and transmitted to clients at times. Depth views can be considered as a compact model of 3-D scenes enabling the remote rendering of complex 3-D environment on relatively low-end devices. The major overhead of the 3-D image warping environment is the transmission of the depth views of the initial and subsequent references. This paper addresses the issue by presenting an effective remote rendering environment based on the deep compression of depth views utilizing the context statistics structure present in depth views. The warped image quality is also explored by reducing the resolution of the depth map. It is shown that proposed deep compression of the remote rendered view significantly outperforms the JPEG2000 and enables the realtime rendering of remote 3-D scene while the degradation of warped image quality is visually imperceptible for the benchmark scenes. Paul Bao, Douglas Gourlay, Youfu Li 0001 |
IEEE Trans. Multim. | 3 |
| 2006 | New Pose-Detection Method for Self-Calibrated Cameras Based on Parallel Lines and Its Application in Visual Control SystemabstractIn this paper, a new method is proposed to detect the pose of an object with two cameras. First, the intrinsic parameters of the cameras are self-calibrated with two pairs of parallel lines that are orthogonal. Then, the poses of the cameras relative to the parallel lines are deduced, and the rotational transformation between the two cameras is calculated. With the intrinsic parameters and the relative pose of the two cameras, a method is proposed to obtain the poses of a line, plane, and rigid object. Furthermore, a new visual-control method is developed using a pose detection rather than a three-dimensional reconstruction. Experiments are conducted to verify the effectiveness of the proposed method. De Xu, Youfu Li 0001, Min Tan 0001 |
IEEE Trans. Syst. Man Cybern. Part B | 2 |
| 2005 | Dynamic Recalibration of an Active Vision System via Generic HomographyabstractActive vision systems are widely used for 3D reconstruction. However, they normally have to be calibrated before vision tasks. In this paper, we present a novel method for dynamic recalibration for a structured light vision system. Given a single image in the scene, a closed-form solution for the pose of the camera is derived by enforcing the constraints via its generic Homography. Then 3D reconstruction can be performed immediately even when the system is displaced or its configuration is changed. Experiments have been carried out and satisfactory results are obtained by evaluating dimensional measurements and back-projection errors. Youfu Li 0001 |
ICRA | 2 |
| 2005 | Dynamic calibration of a structured light system via planar motionabstractActive vision is widely used in robotics. However, it should be calibrated carefully before a vision task. In this paper, we propose a closed-form solution for dynamic calibration and 3D reconstruction using structured light system. In this system, we assume that the light planes cast from the DLP projector are calibrated and kept constant while the principal point of the camera is known. When the camera undergoes a planar motion, its effective focal lengths (measured in width and height of the pixels respectively) and motion parameters can be calibrated dynamically. Then the 3D reconstruction can be effectively carried out with only a single image of the concerned scene. This feature is significant for many dynamic environments. In some practical applications, such as robot navigation along ground plane, planar motion is sufficient and our method provides an effective solution. Beiwei Zhang 0001, Youfu Li 0001 |
IROS | 2 |
| 2005 | Information entropy-based viewpoint planning for 3-D object reconstructionabstractIn this paper, we present an information entropy-based viewpoint-planning approach for reconstruction of freeform surfaces of three-dimensional objects. To achieve the reconstruction, the object is first sliced into a series of cross section curves, with each curve to be reconstructed by a closed B-spline curve. In the framework of Bayesian statistics, we propose an improved Bayesian information criterion (BIC) for determining the B-spline model complexity. Then, we analyze the uncertainty of the model using entropy as the measurement. Based on this analysis, we predict the information gain for each cross section curve for the next measurement. After predicting the information gain of each curve, we obtain the information change for all the B-spline models. This information gain is then mapped into the view space. The viewpoint that contains maximal information gain about the object is selected as the next best view. Experimental results show successful implementation of our view planning method for digitization and reconstruction of freeform objects. Youfu Li 0001, Z. G. Liu |
IEEE Trans. Robotics | 1 |
| 2005 | Vision sensor planning for 3-D model acquisitionabstractA novel method is proposed in this paper for automatic acquisition of three-dimensional (3-D) models of unknown objects by an active vision system, in which the vision sensor is to be moved from one viewpoint to the next around the target to obtain its complete model. In each step, sensing parameters are determined automatically for incrementally building the 3-D target models. The method is developed by analyzing the target's trend surface, which is the regional feature of a surface for describing the global tendency of change. While previous approaches to trend analysis are usually focused on generating polynomial equations for interpreting regression surfaces in three dimensions, this paper proposes a new mathematical model for predicting the unknown area of the object surface. A uniform surface model is established by analyzing the surface curvatures. Furthermore, a criterion is defined to determine the exploration direction, and an algorithm is developed for determining the parameters of the next view. Implementation of the method is carried out to validate the proposed method. Shengyong Chen, Youfu Li 0001 |
IEEE Trans. Syst. Man Cybern. Part B | 2 |
| 2004 | Active Viewpoint Planning for Model ConstructionabstractThis paper presents a novel method of viewpoint planning for incrementally building the models of unknown objects or environments by an active vision system. The proposed method is based on the model of trend surface, which is the regional feature of a surface for describing the global tendency of change. A new mathematical model is developed for predicting the unknown area of the object surface. A unique surface model is established by analyzing the surface curvature. Furthermore, a criterion is defined to determine the exploration direction. The algorithm is developed for determining the next view pose, which satisfies the placement constraints such as resolution, focus, and field of view. Finally, implementation of the method is carried out to verify the proposed method. Shengyong Chen, Youfu Li 0001 |
ICRA | 2 |
| 2004 | Uncalibrated Euclidean 3-D reconstruction using an active vision systemabstractUncalibrated reconstruction of a scene is desired in many practical applications of computer vision. However, using a single camera with unconstrained motion and unknown parameters, a true Euclidean three-dimensional (3-D) model of the scene cannot be reconstructed. In this paper, we present a method for true Euclidean 3-D reconstruction using an active vision system consisting of a pattern projector and a camera. When the intrinsic and extrinsic parameters of the camera are changed during the reconstruction, they can be self-calibrated and the real 3-D model of the scene can then be reconstructed. The parameters of the projector are precalibrated and are kept constant during the reconstruction process. This allows the configuration of the vision system to be varied during a reconstruction task, which increases its self-adaptability to the environment or scene structure in which it is to work. Youfu Li 0001, R. S. Lu |
IEEE Trans. Robotics Autom. | 1 |
| 2004 | Integrated sensing and filter design for a single-link flexible manipulatorabstractThis paper addresses the problem of using different sensors in filter design that can simultaneously satisfy multiple specifications. A novel approach is taken in the design paradigm that integrates the sensing strategy with the filter design, which improves the filtering performance. An application to the estimation of the endpoint vibration rate of a single-link flexible manipulator is presented with experimental verifications. Guoli Wang 0001, Youfu Li 0001 |
IEEE Trans. Robotics | 2 |
| 2003 | Context Modeling based Depth Image Compression for Distributed Virtual EnvironmentabstractDepth images, comprising pixel intensities and depth map, are viewed as a compact model of 3D scenes and used in 3D image warping for the distributed or collaborative virtual environment to enable the distributed rendering of complex 3D scenes at relatively low cost. The major overhead of the model is the transmission of the depth images of the initial reference and the subsequent new views that may be up to a few Mbs in size depending on the screen resolution. This paper presents efficient compression techniques specifically designed for the depth images. Additionally we experiment with the warped image quality by reducing the resolution of the depth map. We show that significant compression can be achieved at a cost of a very modest impairment in the perceptual quality of the warped image. Paul Bao, Douglas Gourlay, Youfu Li 0001 |
CW | 3 |
| 2003 | Dynamically reconfigurable visual sensing for 3D perceptionabstractIn many applications, a vision sensor often needs to move from one place to another and change its configuration for perception of different object features. A dynamic reconfigurable vision sensor is useful in such a case to gaze at the features. This paper introduces this concept and investigates the issues in self-recalibrating a 6-DOF structured light system under changing sensing configuration. The relative pose between the projector and camera of the system is calibrated by taking a single view of the scene, so that the 3D measurements and reconstruction can be performed immediately when and if the configuration of the system is changed. Experiments were carried out to demonstrate the implementation of the proposed method. Shengyong Chen, Youfu Li 0001 |
ICRA | 2 |
| 2003 | Automatic recalibration of an active structured light vision systemabstractA structured light vision system using pattern projection is useful for robust reconstruction of three-dimensional objects. One of the major tasks in using such a system is the calibration of the sensing system. This paper presents a new method by which a two-degree-of-freedom structured light system can be automatically recalibrated, if and when the relative pose between the camera and the projector is changed. A distinct advantage of this method is that neither an accurately designed calibration device nor the prior knowledge of the motion of the camera or the scene is required. Several important cues for self-recalibration are explored. The sensitivity analysis shows that high accuracy in-depth value can be achieved with this calibration method. Some experimental results are presented to demonstrate the calibration technique. Youfu Li 0001, Shengyong Chen |
IEEE Trans. Robotics Autom. | 1 |
| 2002 | Optimum viewpoint planning for model-based robot visionabstractIn some model-based vision tasks, such as automatic inspection of industrial parts, a set of viewpoints must be planned for sampling all features of interest around the object. This paper presents the techniques of deciding the optimal viewpoint distribution and a shortest path through these viewpoints, which are achieved by the genetic algorithm and Christofides algorithm respectively. Shengyong Chen, Youfu Li 0001 |
IEEE Congress on Evolutionary Computation | 2 |
| 2002 | Self Recalibration of a Structured Light Vision System from a Single ViewabstractStructured-light system is widely used for reconstructing 3D objects in machine vision. One of the major tasks in establishing such a system is the laborious and tedious calibration of the sensors. This paper presents a new method which dynamically calibrates the system automatically, if and when the relative pose between the camera and the projector is changed. A distinct advantage of this method is that neither the design of a calibration pattern/device nor the pre-knowledge of the movement of camera or scene is required. Several important cues for self-recalibration, including geometrical cue and focus cue, are explored in this paper Finally, some experimental observations are presented to illustrate the implementation of this new method. Shengyong Chen, Youfu Li 0001 |
ICRA | 2 |
| 2002 | A Method of Automatic Sensor Placement for Robot Vision in Inspection TasksabstractThis paper presents an automatic sensor placement technique for robot vision in inspection tasks. In such vision systems, a sensor often needs to be moved from one pose to another around the object to sample all features of interest. Multiple 3D images are taken from different vantage points. The technique involves deciding the optimal sensor placements and a shortest path through these viewpoints for automatic generation of an inspection plan. A viewpoint is expressed by N parameters and a topology of viewpoints is achieved by genetic algorithm. The inspection plan is evaluated using a min-max criterion and the shortest path is determined by Christofides algorithm. In addition, a computation example is presented to illustrate the techniques and algorithms. Shengyong Chen, Youfu Li 0001 |
ICRA | 2 |
| 2001 | Modifying the shape of NURBS surfaces with geometric constraints
Shi-Min Hu 0001, Youfu Li 0001 |
Comput. Aided Des. | 2 |
| 2000 | General constrained deformations based on generalized metaballs
Xiaogang Jin 0001, Youfu Li 0001, Qunsheng Peng 0001 |
Comput. Graph. | 2 |
| 1998 | General Constrained Deformations based on Generalized MetaballsabstractSpace deformation is an important tool in computer animation and shape design. We propose a new local deformation model based on generalized metaballs. The user specifies a series of constraints, which can be made up of points, lines, surfaces and volumes, their effective radii and maximum displacements; the deformation model creates a generalized metaball for each constraint. Each generalized metaball is associated with a potential function centered on the constraint, the potential function drops from 1 on the constraint to 0 on the effective radius. This deformation model operates on the local space and is independent of the underlining representation of the object to be deformed. The deformation can be finely controlled by adjusting the parameters of the generalized metaballs. We also present some extensions and the extended deformation model to include scale and rotation constraints. Experiments show that this deformation model is efficient and intuitive. It can deal with various constraints, which is difficult for traditional deformation model. Xiaogang Jin 0001, Youfu Li 0001, Qunsheng Peng 0001 |
PG | 2 |
| 1995 | Robotic Grasping of Complex Objects without Full Geometrical Knowledge of the ShapeabstractThis paper describes a control method for searching suitable gripping points on the outline of generic shapes. It is aimed at robotic grasping tasks where no full geometrical knowledge of the shape is assumed. We first describe the method for a range of 2D generic shapes and then instantiate the method for pick and place operations on known shapes. The algorithms are run on the shape as it appears on the computer screen directly from a vision system. Virtual robotic fingers and associated sensors are then configured and positioned on the screen with three variables being simultaneously controlled. When the control systems reach a steady state, information about virtual finger position and orientation are then used to drive the manipulator. M. A. Rodrigues, Youfu Li 0001, Mark H. Lee, Jem J. Rowland, C. King |
ICRA | 2 |
| 1994 | A Visually Guided Robot System for Food Handling ApplicationsabstractThis paper presents the initial development of a visually guided robotic system for handling food products that are often presented in an unstructured manner. In order to handle the communications between the host and robot controller more effectively, we developed multi-level commands within the host programming environment. Vision guided grasping is described in the context of vector manipulations. The system is able to calculate the grasp vector according to both online visual information and off-line data based on the food product and gripper type in consideration. A case study is presented where the system handles trapezoidal shaped fish products.> Youfu Li 0001, Mark H. Lee, M. A. Rodrigues, Jem J. Rowland |
ICRA | 1 |
| 1994 | Supervisory robotic control under vision guidanceabstractAdvanced robotic applications often require a robot to work as a part of an integrated system. Supervisory control of the robot in such cases readily permits incorporation of external sensing systems such as vision. To handle communication between a host and robot controller more effectively, we have developed multilevel commands within the host programming environment. An example application is presented which is a visually guided robot food handling task. The work cell comprises a robot manipulator, a vision system, and a host controller whereby the supervisory control of the robot is achieved.> Youfu Li 0001, Mark H. Lee, M. A. Rodrigues, Jem J. Rowland |
IROS | 1 |