EDBT 2026 Demo / reviewers in the wild / expert
Ye Liu 0005
dblp:96/2615-5
· DBLP profile ↗
35ranked-venue papers
10as first author
20since 2021 · last 2026
0000-0002-2686-3002ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 27 · 7 first-author · 16 since 2021Artificial intelligence and machine learning · 11 · 3 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Perceive More with Less: LiDAR Point Cloud Compression at Just Recognizable Distortion for 3D Scene UnderstandingabstractExisting LiDAR point cloud (LPC) data coding methods primarily focus on balancing compression efficiency and reconstruction quality according to the human vision system (HVS). However, these methods rarely consider the requirements of downstream scene understanding tasks from the perspective of the machine vision system (MVS). To address this challenge, we explore the maximum degree of LPC compression that has negligible impact on perception accuracy, called LPC-based just recognizable compression distortion (lpcJRCD). Specifically, we introduce a novel point-wise quantization approach for constructing a MVS-based LiDAR dataset and present a new lpcJRCD-guided intelligent compression framework tailored for MVS applications. To enhance MVS-based LPC compression efficiency, we develop a dual-feature interaction (DFI) module that fuses point and voxel features. Additionally, we propose a mask-based loss function to ensure accurate point-wise quality level prediction. Experimental results demonstrate the effectiveness of our proposed model in reducing the average bit rate by up to 94.98% while preserving perception accuracy in autonomous vehicles. Miaohui Wang, Runnan Huang, Taojun Liu, Shuyuan Lin, Ye Liu 0005, Yun Song |
AAAI | 5 |
| 2026 | The Last Byte: Learning Just Enough for Machine-Oriented Image CompressionabstractJust recognizable distortion (JRD) has been introduced for image compression for machines, aiming to quantify the maximum coding distortion that can be tolerated by a specific perception model, thereby defining the upper bound of machine vision redundancy (MVR). However, existing JRD-based redundancy estimation methods face three key challenges: limited dataset annotation accuracy, low prediction efficiency, and insufficient perception accuracy, all of which hinder their practical deployment. To address these limitations, we propose a new MVR-Net, a frame-wise efficient JRD prediction method that generates the optimal encoding quantization map in a single inference pass. Furthermore, we refine the annotation standard for JRD datasets based on experimental insights, enhancing the precision of recognizable redundancy measurement. Compared to stateof-the-art methods, MVR-Net achieves a superior balance between bitrate reduction and perception accuracy in JRD-guided compression, while offering up to a 40,000× speed improvement, demonstrating its practicality and efficiency for real-world applications. Wuyuan Xie, Zhenming Li, Ye Liu 0005, Yun Song, Miaohui Wang |
AAAI | 3 |
| 2026 | DPTracker: Dynamic prompter for RGB-D tracking
Junzhe Zhao, Jintao Su, Ye Liu 0005, Jun Liu 0036, Miaohui Wang |
Pattern Recognit. Lett. | 3 |
| 2026 | APNet: Accurate Prompting Network With Modality Guidance and Structural Awareness for RGB-D Semantic SegmentationabstractParameter-efficient fine-tuning (PEFT) is promising for RGB-D semantic segmentation, as lightweight prompters enable frozen pre-trained RGB backbones to leverage massive RGB pretraining knowledge without full fine-tuning on limited paired RGB-D data. However, existing PEFT methods have two critical limitations: static modal fusion ignores the dynamic reliability of RGB and depth across scenes, leading to suboptimal performance in complex environments; conventional prompts lack structural awareness, causing the loss of edge and texture details essential for dense prediction. To solve these problems, we propose the Accurate Prompting Network (APNet) for precise prompt injection in frozen backbones with two core modules. A Modality Effectiveness Guider (MEG) conducts input-level modal reliability assessment and dynamically generates scene-adaptive modality weights by capturing scene characteristics (e.g., illumination, texture richness). A Structural Awareness Prompter (SAP) injects directional structural priors into prompts via multi-directional gating convolution, endowing prompts with explicit edge and texture information to match semantic segmentation demands. MEG and SAP collaboratively form a precise prompting mechanism that realizes dynamic modal contribution allocation and structural detail preservation, facilitating efficient and accurate cross-modal knowledge transfer to the frozen backbone. Extensive experiments on NYUDv2 and SUN RGB-D show that APNet achieves state-of-the-art mIoU of 59.6% and 52.6% with only 6.2M trainable parameters, realizing a superior trade-off between segmentation accuracy and parameter efficiency. Junzhe Zhao, Jintao Su, Jun Liu 0036, Miaohui Wang, Ye Liu 0005 |
IEEE Signal Process. Lett. | 5 |
| 2025 | mmFAS: Multimodal Face Anti-Spoofing Using Multi-Level Alignment and Switch-Attention FusionabstractThe increasing number of presentation attacks on reliable face matching has raised concerns and garnered attention towards face anti-spoofing (FAS). However, existing methods for FAS modeling commonly fuse multiple visual modalities (e.g., RGB, Depth, and Infrared) in a straightforward manner, disregarding latent feature gaps that can hinder representation learning. To address this challenge, we propose a novel multimodal FAS framework (mmFAS) that focuses on explicit alignment and fusion of latent features across different modalities. Specifically, we develop a multimodal alignment module to alleviate the latent feature gap by using instance-level contrastive learning and class-level matching simultaneously. Further, we explore a new switch-attention based fusion module to automatically aggregate complementary information and control model complexity. To evaluate the anti-spoofing performance more effectively, we adopt a challenging yet meaningful cross-database protocol involving four benchmark multimodal FAS datasets to simulate realworld scenarios. Extensive experimental results demonstrate the effectiveness of mmFAS in improving the accuracy of FAS systems, outperforming 10 representative methods. Geng Chen 0006, Wuyuan Xie, Di Lin 0002, Ye Liu 0005, Miaohui Wang |
AAAI | 4 |
| 2025 | Adaptive Skeleton Prompt Tuning for Cross-Dataset 3D Human Pose EstimationabstractInconsistency of distributions in human actions and camera viewpoints can lead to significant deviations when the pre-trained 3D pose estimators are tested on cross-datasets. In practical applications, the estimators usually follow the standard full fine-tuning paradigm on the target dataset, which requires updating and saving a complete set of training parameters for different tasks, resulting in a large waste of resources and distorting pre-trained features. Taking inspiration from the widely used prompt learning in NLP, we explore the parameter-efficient fine-tuning solution of 3D pose estimators for the first time and propose the Adaptive Skeleton Prompt Tuning (ASP-Tuning) method, which freezes the backbone of the pre-trained model and generates a series of pose generic promptings as well as adaptive promptings specific to the input skeleton features to learn distribution transformation. Extensive experiments on multiple estimator backbones and datasets show that our method is superior to other fine-tuning methods and achieves state-of-the-art performance. Haolun Li 0001, Fuchen Zheng, Ye Liu 0005, Jian Xiong 0005, Haidong Hu, Hao Gao 0005 |
ICASSP | 3 |
| 2025 | Diffusion Models are Good Unsupervised Class-agnostic Shape Part SegmentatorsabstractShape part segmentation is a critical task in computer graphics and robotics. However, traditional supervised methods rely heavily on large amounts of labeled data, which poses significant challenges in many real-world scenarios where such data is often scarce or difficult to obtain. To address this issue, we propose an unsupervised, class-agnostic part segmentation method called Point Diffusion Segmentation (PDS). Our research demonstrates that unconditional point cloud diffusion models can capture abstract object concepts within their sub-attention layers. By extracting preliminary point cloud features from these attention maps, PDS generates efficacious segmentation results. This method fully leverages unlabeled data and proves to be highly applicable in various downstream tasks, including zero-shot part segmentation. Without resorting to any labeled data, PDS improves the zero-shot part segmentation performance of PointClipV2 by 3.1% on the ShapeNet Part dataset, setting a new state-of-the-art baseline and demonstrating significant potential of PDS. Zhongbin Jiang, Tianhao Shi, Hao Gao 0005, Jun Liu 0036, Ye Liu 0005 |
ICASSP | 5 |
| 2025 | MMPX: Multi-modal Mamba Prompter to Large Vision Foundation Model for RGB-X Semantic SegmentationabstractMulti-modal semantic segmentation leverages multiple types of input data to perform pixel-level classification of images, enhancing the accuracy and robustness of segmentation tasks. Mainstream methods use small-scale models which have limited generalization ability. Training large-scale multi-modal models requires massive multi-modal data, which is difficult to obtain. Thus, fine-tuning large-scale vision foundation models (LVFMs) trained with abundant RGB data for multi-modal segmentation is a more practical solution. Existing methods only tap into the potential of LVFMs in the RGB modality which ignores their potential in non-RGB modalities. In this paper, we propose an innovative universal prompting framework, MMPX. Specifically, the effective multi-modal Mamba fuser (MMF) explores the potential of LVFMs in the integrated representation of RGB+X. On the other hand, we introduce multi-modal Mamba prompters (MMPs) to fine-tune large-scale foundation models. This prompter takes the integrated RGB+X representation as input and dynamically adjusts the model parameters using Mamba, eliminating redundant information while retaining key features, thus achieving efficient prompt generation. The proposed method achieves SOTA performance on five multi-modal benchmarks, including RGB+Depth, RGB+Thermal, RGB+Event, which fully validate the effectiveness and generalization ability of the approach. The code and results are available at: https://github.com/CauchyCat/MMPX. Ye Liu 0005, Hao Gao 0005, Jun Liu 0036 |
ICME | 2 |
| 2025 | Frequency Decoupled Masked Auto-Encoder for Self-Supervised Skeleton-Based Action RecognitionabstractIn 3D skeleton-based action recognition, the limited availability of supervised data has driven interest in self-supervised learning methods. The reconstruction paradigm using masked auto-encoder (MAE) is an effective and mainstream self-supervised learning approach. However, recent studies indicate that MAE models tend to focus on features within a certain frequency range, which may result in the loss of important information. To address this issue, we propose a frequency decoupled MAE. Specifically, by incorporating a scale-specific frequency feature reconstruction module, we delve into leveraging frequency information as a direct and explicit target for reconstruction, which augments the MAE's capability to discern and accurately reproduce diverse frequency attributes within the data. Moreover, in order to address the issue of unstable gradient updates caused by more complex optimization objectives with frequency reconstruction, we introduce a dual-path network combined with an exponential moving average (EMA) parameter updating strategy to guide the model in stabilizing the training process. We have conducted extensive experiments which have demonstrated the effectiveness of the proposed method. Ye Liu 0005, Tianhao Shi, Mingliang Zhai, Jun Liu 0036 |
IEEE Signal Process. Lett. | 1 |
| 2025 | CPAL: Cross-Prompting Adapter With LoRAs for RGB+X Semantic SegmentationabstractAs sensor technology evolves, RGB+X systems combine traditional RGB cameras with another type of auxiliary sensor, which enhances perception capabilities and provides richer information for important tasks such as semantic segmentation. However, acquiring massive RGB+X data is difficult due to the need for specific acquisition equipment. Therefore, traditional RGB+X segmentation methods often perform pretraining on relatively abundant RGB data. However, these methods lack corresponding mechanisms to fully exploit the pretrained model, and the scope of the pretraining RGB dataset remains limited. Recent works have employed prompt learning to tap into the potential of pretrained foundation models, but these methods adopt a unidirectional prompting approach i.e., using X or RGB+X modality to prompt pretrained foundation models in RGB modality, neglecting the potential in non-RGB modalities. In this paper, we are dedicated to developing the potential of pretrained foundation models in both RGB and non-RGB modalities simultaneously, which is non-trivial due to the semantic gap between modalities. Specifically, we present the CPAL (Cross-prompting Adapter with LoRAs), a framework that features a novel bi-directional adapter to simultaneously fully exploit the complementarity and bridging the semantic gap between modalities. Additionally, CPAL introduces low-rank adaption (LoRA) to fine-tune the foundation model of each modal. With the support of these elements, we have successfully unleashed the potential of RGB foundation models in both RGB and non-RGB modalities simultaneously. Our method achieves state-of-the-art (SOTA) performance on five multi-modal benchmarks, including RGB+Depth, RGB+Thermal, RGB+Event, and a multi-modal video object segmentation benchmark, as well as four multi-modal salient object detection benchmarks. The code and results are available at:https://github.com/abelny56/CPAL. Ye Liu 0005, Miaohui Wang, Jun Liu 0036 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2025 | Coupled Noise Suppression and Feature Enhancement Network for Skeleton-Based Action RecognitionabstractIn recent years, remarkable progress has been made in skeleton-based action recognition. However, there is a significant amount of noise in skeleton data, which is simply overlooked by most existing methods. Some methods have designed specialized mechanisms to handle noise, but these mechanisms are either based on prior knowledge or require additional supervision information. To overcome these problems, we propose in this article a fully implicit solution, which embeds a soft-thresholding-based denoising module into existing networks, which can automatically learn to remove noise without any prior knowledge or additional supervision information. In addition, by relaxing the nonnegative constraint, the module gains the ability to adaptively enhance key features. Based on this, we further propose a two-staged method for coupled noise suppression and feature enhancement. The proposed method achieves state-of-the-art performance on public datasets. Moreover, on noise polluted datasets, the proposed method demonstrates significant performance advantages over existing methods. Ye Liu 0005, Tianyong Wu, Tianhao Shi, Miaohui Wang, Hao Gao 0005, Jun Liu 0036 |
IEEE Trans. Ind. Informatics | 1 |
| 2025 | suLPCC: A Novel LiDAR Point Cloud Compression Framework for Scene Understanding TasksabstractLight detection and ranging (LiDAR) point cloud compression (LPCC) plays an important role in managing the storage, transmission, and perception of the rapidly expanding volume of LiDAR point cloud (LPC) data. However, there has been a noticeable lack of comprehensive investigation into LPCC methods specifically designed for environmental perception and understanding. To address this gap, we propose a new LPCC framework aimed at meeting the unique requirements of various scene understanding tasks, enhancing the adaptability of LPCCs in real-world scenarios. Specifically, we divide the input LPCs into an object and a scene component through a distinction module, design a new point completion-based method to encode object LPCs, and develop novel structure-aware intracoding and motion-optimized intercoding schemes to compress scene LPCs. Experimental results on three benchmark datasets demonstrate the effectiveness of our proposed method on the localization, mapping, and detection tasks. We believe that the findings presented in this article will contribute to a deeper understanding of LPCCs as well as promote further development of LiDAR sensor-based systems. Miaohui Wang, Runnan Huang, Ye Liu 0005, Yanshan Li, Wuyuan Xie |
IEEE Trans. Ind. Informatics | 3 |
| 2025 | Co-MOT: Exploring the Collaborative Relations in Traffic Flow for 3D Multi-Object TrackingabstractIn real-world scenes, vehicles and pedestrians on the road often exhibit consistency in their overall motion, forming the traffic flow we observe. Exploring this global collective motion consistency to aid in 3D multi-object tracking (MOT) tasks is an under-investigated issue in existing research. Recently, Graph Neural Networks (GNN) have been introduced to model interactions between targets in 3D tracking problems, achieving remarkable performance. However, existing GNN based methods usually employ neighborhood-based approaches to construct graphs which are unable to fully exploit collective relations in traffic flow. In this paper, we propose a GNN based 3D MOT method which effectively utilizes the collective motion consistency in traffic flow. Collective motion is modeled with a densely connected intra-flow graph within the collective group, allowing information to flow quickly. To build the intra-flow graph, we propose an effective collinearity condition to distinguish potential collective groups from the detected objects. For reasoning on the graph, we propose a progressive serial message-passing solver which enables the network to learn complex group movement relationships based on a thorough understanding of simple neighborhood relations. Our proposed method achieves state-of-the-art performance on public datasets: NuScenes and KITTI tracking benchmark. We have conducted extensive experiments to evaluate the comprehensive performance of our proposed method which demonstrates the effectiveness of the proposed method. Ye Liu 0005, Xingdi Liu, Zhongbin Jiang, Jun Liu 0036 |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2024 | Deep learning based object detection from multi-modal sensors: an overview
Ye Liu 0005, Shiyang Meng, Hongzhang Wang, Jun Liu 0036 |
Multim. Tools Appl. | 1 |
| 2024 | Learning graph-based representations for scene flow estimation
Mingliang Zhai, Hao Gao 0005, Ye Liu 0005, Jianhui Nie, Kang Ni |
Multim. Tools Appl. | 3 |
| 2023 | Just Noticeable Visual Redundancy Forecasting: A Deep Multimodal-Driven ApproachabstractJust noticeable difference (JND) refers to the maximum visual change that human eyes cannot perceive, and it has a wide range of applications in multimedia systems. However, most existing JND approaches only focus on a single modality, and rarely consider the complementary effects of multimodal information. In this article, we investigate the JND modeling from an end-to-end homologous multimodal perspective, namely hmJND-Net. Specifically, we explore three important visually sensitive modalities, including saliency, depth, and segmentation. To better utilize homologous multimodal information, we establish an effective fusion method via summation enhancement and subtractive offset, and align homologous multimodal features based on a self-attention driven encoder-decoder paradigm. Extensive experimental results on eight different benchmark datasets validate the superiority of our hmJND-Net over eight representative methods. Wuyuan Xie, Shukang Wang, Sukun Tian, Lirong Huang, Ye Liu 0005, Miaohui Wang |
AAAI | 5 |
| 2022 | SingleMatch: a point cloud coarse registration method with single match point and deep-learning describer
Jianhui Nie, Hao Gao 0005, Ye Liu 0005, Haotian Lu 0006 |
Multim. Tools Appl. | 4 |
| 2022 | Localizing and tracking dense crowd of microbes by joint association and detection refinement
Ye Liu 0005, Shuohong Wang, Jianhui Nie, Hao Gao 0005 |
Vis. Comput. | 1 |
| 2021 | Enhancement of ridge-valley features in point cloud based on position and normal guidance
Jianhui Nie, Zhaochen Zhang, Ye Liu 0005, Hao Gao 0005, Feng Xu 0005, Wenkai Shi |
Comput. Graph. | 3 |
| 2021 | Bas-relief generation from point clouds based on normal space compression with real-time adjustment on CPU
Jianhui Nie, Wenkai Shi, Ye Liu 0005, Hao Gao 0005, Feng Xu 0005, Zhaochen Zhang |
Graph. Model. | 3 |
| 2020 | An artificial bee algorithm with a leading group and its application into image registration
Haidong Hu, Chi-Man Pun, Ye Liu 0005, Xiangjing Lai, Hao Gao 0005 |
Multim. Tools Appl. | 3 |
| 2020 | High-quality-guided artificial bee colony algorithm for designing loudspeaker
Hao Gao 0005, Haolun Li 0001, Ye Liu 0005, Huimin Lu 0001, Hyoungseop Kim, Chi-Man Pun |
Neural Comput. Appl. | 3 |
| 2019 | Context-Aware Three-Dimensional Mean-Shift With Occlusion Handling for Robust Object Tracking in RGB-D VideosabstractDepth cameras have recently become popular and many vision problems can be better solved with depth information. But, how to integrate depth information into a visual tracker to overcome the challenges such as occlusion and background distraction is still underinvestigated in current literature on visual tracking. In this paper, we investigate a 3-D extension of a classical mean-shift tracker whose greedy gradient ascend strategy is generally considered as unreliable in conventional 2-D tracking. However, through careful study of the physical property of 3-D point clouds, we reveal that objects which may appear to be adjacent on a 2-D image will form distinctive modes in the 3-D probability distribution approximated by kernel density estimation, and finding the nearest mode using 3-D mean-shift can always work in tracking. Based on the understanding of 3-D mean-shift, we propose two important mechanisms to further boost the tracker's robustness: one is to enable the tracker to be aware of potential distractions and make corresponding adjustments to the appearance model; and the other is to enable the tracker to detect and recover from tracking failures caused by total occlusion. The proposed method is both effective and computationally efficient. On a conventional personal computer, it runs at more than 60 FPS without graphical processing unit acceleration. Ye Liu 0005, Xiaoyuan Jing, Jianhui Nie, Hao Gao 0005, Jun Liu 0036, Guoping Jiang |
IEEE Trans. Multim. | 1 |
| 2018 | Physical blob detector and Multi-Channel Color Shape Descriptor for human detection
Guyue Zhang, Jun Liu 0036, Ye Liu 0005, Luchao Tian, Yan Qiu Chen |
J. Vis. Commun. Image Represent. | 3 |
| 2018 | Early diagnosis of cirrhosis via automatic location and geometric description of liver capsule
Shuohong Wang, Xiang Liu 0003, Ye Liu 0005, Yan Qiu Chen |
Vis. Comput. | 4 |
| 2017 | 3D tracking swimming fish school with learned kinematic model using LSTM networkabstractThis paper proposes a reliable 3D fish tracking method using a novel master-slave camera setup. Instead of conventional dynamic models that rely on prior knowledge about target kinematics, the proposed method learns the kinematic model with a Long Short-Term Memory (LSTM) network. On this basis, the 3D state of fish at each moment is predicted by LSTM network. We propose to use an innovative master-view-tracking-first strategy. The fish are first tracked in the master view. Cross-view association is then established utilizing motion continuity and epipolar constraint cues. Experiments on data sets of different fish densities show that the proposed method is effective and outperforms the state-of-the-art methods. Shuohong Wang, Xiang Liu 0003, Zhiming Qian, Ye Liu 0005, Yan Qiu Chen |
ICASSP | 5 |
| 2017 | Tracking the 3D position and orientation of flying swarms with learned kinematic pattern using LSTM networkabstractAccurately and reliably tracking the 3D position and orientation of individuals in large flying swarms is valuable not only for scientific researches but also practical applications. However, large quantity, frequent occlusions, similar appearance, tiny body size and abrupt motion make it remain an open problem. The 3D flying swarm tracking method proposed in this paper tracks both position and orientation of each individual in the swarm using the particle filter framework. Particles are scattered more pertinently by the dynamic model based on the learned kinematic pattern of a single target with a Long Short-Term Memory (LSTM) network. In addition, the observation model combines the Weighted Occupancy Ratio (WOR) and Temporal Appearance Coherency (TAC) cues in each view to improve the accuracy and robustness of the reconstructed body orientation. Experiments on both simulation and real-world data sets demonstrate the effectiveness and superiority of the proposed method. Shuohong Wang, Hai-Feng Su, Xi En Cheng, Ye Liu 0005, Aike Quo, Yan Qiu Chen |
ICME | 4 |
| 2016 | 3D tracking swimming fish school using a master view tracking first strategyabstract3D motion data of fish school is more valuable than 2D data for behavior and other researches. This paper proposes to use a master view tracking first strategy based on a novel master-slave camera setup. On this basis, fish are firstly tracked in master view in 2D after being extracted via an eye-focused Gaussian Mixture Model (E-GMM) detector. Then 3D trajectories are reconstructed by associating 2D tracking results in master view and detection results in slave views after fish in slave views are localized using an eye-focused Gabor (E-Gabor) detector. Experiments on data sets with different fish densities demonstrate that the proposed method outperforms two state-of-the-art methods in terms of 5 evaluation metrics. Shuohong Wang, Xiang Liu 0003, Ye Liu 0005, Yan Qiu Chen |
BIBM | 4 |
| 2016 | Robust Real-Time Human Perception with Depth CameraabstractPerception of the presence and position of human is crucial for many kinds of Artificial Intelligence (AI) applications. In this paper, we have developed a novel two-staged method for realtime human detection in depth image. The first stage is to quickly scan through the image to detect possible head-top locations in order to ensure all the candidate locations are included. The second stage is to use a novel head-shoulder descriptor (HSD) which jointly encodes the One-hot Depth Difference information and local geometric characteristics of human upper body to filter the detections so as to keep the genuine human locations and discard false positives. The results show that our approach using only depth data is superior to other methods using color and depth images on four datasets. In addition, our method performs well under weak illumination conditions or even total darkness. Moreover, our system is also able to run in real-time on conventional PC without GPU acceleration. Guyue Zhang, Luchao Tian, Ye Liu 0005, Jun Liu 0036, Xiang An Liu, Yang Liu 0003, Yan Qiu Chen |
ECAI | 3 |
| 2016 | Automatic 3D tracking system for large swarm of moving objects
Ye Liu 0005, Shuohong Wang, Yan Qiu Chen |
Pattern Recognit. | 1 |
| 2015 | An ultra-fast human detection method for color-depth camera
Jun Liu 0036, Guyue Zhang, Ye Liu 0005, Luchao Tian, Yan Qiu Chen |
J. Vis. Commun. Image Represent. | 3 |
| 2015 | Detecting and tracking people in real time with RGB-D camera
Jun Liu 0036, Ye Liu 0005, Guyue Zhang, Peiru Zhu, Yan Qiu Chen |
Pattern Recognit. Lett. | 2 |
| 2013 | Real-time human detection and tracking in complex environments using single RGBD cameraabstractThis paper presents a new approach to real-time human detection and tracking in cluttered and dynamic environments by integration of RGB and depth data. We introduce the notion of Point Ensemble Image, which fully encodes both RGB and depth information from a virtual plan-view perspective, and we reveal that human detection and tracking in 3D space can be performed very effectively based on this new representation. Our human detector is able to take advantage of depth data by effectively locate physically plausible candidates as a first step, and then both depth and color information is made full use of in a supervised learning manner at the second stage. 3D trajectories of humans are finally generated by data association in which joint statistics of color and height are computed and compared. Experimental results show that the system is able to work satisfactorily in complex real-world situations. Jun Liu 0036, Ye Liu 0005, Ying Cui 0003, Yan Qiu Chen |
ICIP | 2 |
| 2012 | Automatic Tracking of a Large Number of Moving Targets in 3D
Ye Liu 0005, Yan Qiu Chen |
ECCV (4) | 1 |
| 2012 | 3D tracking of deformable surface by propagating feature correspondences
Ye Liu 0005, Yan Qiu Chen |
ICPR | 1 |