Yinghao Cai

dblp:28/2227 · DBLP profile ↗
← Back
36ranked-venue papers
9as first author
16since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 27 · 6 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 16 · 7 first-author · 1 since 2021Systems, architecture and hardware · 10 · 7 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021
YearPublicationVenuePosition
2026 HOCOpt: Hand-Object Contact Optimization to Improve Pose Estimation in Physical Interactions
abstract
Reconstructing hand–object physical interaction through visual sensing is crucial to understanding human intentions and guaranteeing the safety of human–robot collaboration in industrial applications. Due to heavy occlusion and cluttered backgrounds, existing methods generate inaccurate pose estimation, leading to unrealistic physical interactions between hands and objects. In this article, we present a novel contact-driven pose optimization framework called hand–object contact optimization (HOCOpt) to achieve accurate hand–object pose estimation. The HOCOpt includes two parts: contact estimation and pose optimization. In the contact estimation, we propose a contact region estimation network (CREN) to predict the potential contact across the hand–object meshes with inaccurate poses. A novel contact entropy weight and an auxiliary network are introduced to the training process of CREN to accelerate the model learning and improve the prediction accuracy. For pose optimization, a two-stage hand–object pose optimization method is utilized to refine inaccurate poses by considering both contact distribution and contact stability. During optimization, an orientation-aware differentiable contact model is introduced to account for hand deformation and contact forces to achieve accurate contact modeling. Extensive experiments on ContactPose, HO3D, and DexYCB datasets show that our approach outperforms the existing baselines. Besides, experiments on physical interaction tasks for human–robot collaboration are conducted to demonstrate the practical significance of HOCOpt in industrial scenarios.
Xiaoge Cao, Tao Lu 0006, Wenhao Yu 0011, Yinghao Cai, Shuo Wang 0001
IEEE Trans. Ind. Informatics5
2025 VSLCG-U: A UNet-Based Model with Mamba Gated Connections for Dinosaur Footprint Segmentation
Yinghao Cai, Shaoning Zeng, Jianhang Zhou
ICONIP (2)1
2025 NeuGrasp: Generalizable Neural Surface Reconstruction with Background Priors for Material-Agnostic Object Grasp Detection
abstract
Robotic grasping in scenes with transparent and specular objects presents great challenges for methods relying on accurate depth information. In this paper, we introduce NeuGrasp, a neural surface reconstruction method that leverages background priors for material-agnostic grasp detection. NeuGrasp integrates transformers and global prior volumes to aggregate multi-view features with spatial encoding, enabling robust surface reconstruction in narrow and sparse viewing conditions. By focusing on foreground objects through residual feature enhancement and refining spatial perception with an occupancy-prior volume, NeuGrasp excels in handling objects with transparent and specular surfaces. Extensive experiments in both simulated and real-world scenarios show that NeuGrasp outperforms state-of-the-art methods in grasping while maintaining comparable reconstruction quality. More details are available at https://neugrasp.github.io/.
Qingyu Fan, Yinghao Cai, Wenzhe He, Tao Lu 0006, Shuo Wang 0001
ICRA2
2025 MISCGrasp: Leveraging Multiple Integrated Scales and Contrastive Learning for Enhanced Volumetric Grasping
abstract
Robotic grasping faces challenges in adapting to objects with varying shapes and sizes. In this paper, we introduce MISCGrasp, a volumetric grasping method that integrates multi-scale feature extraction with contrastive feature enhancement for self-adaptive grasping. We propose a query-based interaction between high-level and low-level features through the Insight Transformer, while the Empower Transformer selectively attends to the highest-level features, which synergistically strikes a balance between focusing on fine geometric details and overall geometric structures. Furthermore, MISCGrasp utilizes multi-scale contrastive learning to exploit similarities among positive grasp samples, ensuring consistency across multi-scale features. Extensive experiments in both simulated and real-world environments demonstrate that MISCGrasp outperforms baseline and variant methods in tabletop decluttering tasks. More details are available at https://miscgrasp.github.io/.
Qingyu Fan, Yinghao Cai, Chunting Jiao, Tao Lu 0006, Shuo Wang 0001
IROS2
2025 SENIOR: Efficient Query Selection and Preference-Guided Exploration in Preference-based Reinforcement Learning
abstract
Preference-based Reinforcement Learning (PbRL) methods provide a solution to avoid reward engineering by learning reward models based on human preferences. However, poor feedback- and sample- efficiency still remain the problems that hinder the application of PbRL. In this paper, we present a novel efficient query selection and preference-guided exploration method, called SENIOR, which could select the meaningful and easy-to-comparison behavior segment pairs to improve human feedback-efficiency and accelerate policy learning with the designed preference-guided intrinsic rewards. Our key idea is twofold: (1) We designed a Motion-Distinction-based Selection scheme (MDS). It selects segment pairs with apparent motion and different directions through kernel density estimation of states, which is more task-related and easy for human preference labeling; (2) We proposed a novel preference-guided exploration method (PGE). It encourages the exploration towards the states with high preference and low visits and continuously guides the agent achieving the valuable samples. The synergy between the two mechanisms could significantly accelerate the progress of reward and policy learning. Our experiments show that SENIOR outperforms other five existing methods in both human feedback-efficiency and policy convergence speed on six complex robot manipulation tasks from simulation and four real-worlds. Videos can be found on our project website: https://2025senior.github.io/
Hexian Ni, Tao Lu 0006, Haoyuan Hu, Yinghao Cai, Shuo Wang 0001
IROS4
2025 Learn-Gen-Plan: Bridging the Gap Between Vision Language Models and Real-World Long-Horizon Dexterous Manipulations
abstract
Long-horizon dexterous tasks have been a long-standing problem in robotic manipulation. Previous studies have developed task and motion planning, imitation learning, and reinforcement learning methods for long-horizon manipulations. However, these methods are hard to achieve efficient planning for new tasks. Empowered with the Vision Language Model (VLM), recent studies significantly improve the generalization of robot systems. However, these works are only verified in simple pick-and-place tasks due to limited skills. To this end, we propose the Learn-Gen-Plan (LGP), which combines the VLM and learning-based primitives to endow robots with the ability to efficiently plan and complete various long-horizon dexterous tasks. LGP contains two key phases: skill generation and task planning. In skill generation, the Skill Generator is proposed to utilize the learned key primitives and hand-crafted trivial primitives to generate adaptive robot skills. In task planning, the Multimodal Planner generates the robot plan based on image observation, generated skills, and text prompts. We set up a series of dexterous tasks (e.g., cable routing, peg-in-hole assembly) in a real-world lighting circuit wiring scenario to evaluate LGP. The experimental results show that LGP efficiently generates robot plans with learned skills, controlling the robot to complete various multi-step cable wiring tasks.
Peng Hao 0003, Shaowei Cui, Junhang Wei, Tao Lu 0006, Yinghao Cai, Shuo Wang 0001
IEEE Trans Autom. Sci. Eng.5
2024 Exploring Consistency in Graph Representations: from Graph Kernels to Graph Neural Networks
abstract
Graph Neural Networks (GNNs) have emerged as a dominant approach in graph representation learning, yet they often struggle to capture consistent similarity relationships among graphs. To capture similarity relationships, while graph kernel methods like the Weisfeiler-Lehman subtree (WL-subtree) and Weisfeiler-Lehman optimal assignment (WLOA) perform effectively, they are heavily reliant on predefined kernels and lack sufficient non-linearities. Our work aims to bridge the gap between neural network methods and kernel approaches by enabling GNNs to consistently capture relational structures in their learned representations. Given the analogy between the message-passing process of GNNs and WL algorithms, we thoroughly compare and analyze the properties of WL-subtree and WLOA kernels. We find that the similarities captured by WLOA at different iterations are asymptotically consistent, ensuring that similar graphs remain similar in subsequent iterations, thereby leading to superior performance over the WL-subtree kernel. Inspired by these findings, we conjecture that the consistency in the similarities of graph representations across GNN layers is crucial in capturing relational structures and enhancing graph classification performance. Thus, we propose a loss to enforce the similarity of graph representations to be consistent across different layers. Our empirical analysis verifies our conjecture and shows that our proposed consistency loss can significantly enhance graph classification performance across several GNN backbones on various datasets.
Xuyuan Liu, Yinghao Cai, Qihui Yang, Yujun Yan
NeurIPS2
2022 Meta-Imitation Learning by Watching Video Demonstrations
Tao Lu 0006, Xiaoge Cao, Yinghao Cai, Shuo Wang 0001
ICLR4
2022 Joint Self-Supervised Monocular Depth Estimation and SLAM
abstract
Classical monocular Simultaneous Localization and Mapping (SLAM) and convolutional neural networks (CNNs) based monocular depth estimation represent two different methods towards reconstructing the 3D geometry of the scene. In this paper, we leverage SLAM and depth estimation for their respective advantages to further improve the performance of both tasks. For SLAM, running pseudo RGBD-SLAM with CNN-predicted depths improves the accuracy of visual odometry and mapping compared with the monocular SLAM baseline. For depth estimation, we use 3D scene structures from geometric SLAM to refine the pre-trained monocular depth estimation network to update the model which did not reach the optimum due to the photometric inconsistency. Moreover, the proposed method incorporates an optional Sparse Auxiliary Network [1] into the original depth estimation network, from which the sparse depth features are dynamically combined with RGB features for predicting the depth map. Experimental results on KITTI and TUM RGB-D datasets show that our method achieves state-of-the-art performances on both depth prediction and pose estimation tasks.
Xiaoxia Xing, Yinghao Cai, Tao Lu 0006, Dayong Wen
ICPR2
2022 Learning-based Six-axis Force/Torque Estimation Using GelStereo Fingertip Visuotactile Sensing
abstract
Visuotactile sensors have recently attracted much attention in robot communities due to the benefit of high spatial resolution sensing. However, force/torque estimation by visuotactile sensors remains a challenging problem. In this paper, we propose a learning-based six-axis force/torque estimation network using GelStereo visuotactile sensor, which can provide two-dimensional (2D) and three-dimensional (3D) displacements of markers embedded in the sensor surface. The convolutional neural networks are employed to extract multi-modal tactile deformation features; and a novel contact positional encoding method is proposed to eliminate the influence of translation invariance in convolutional operators. The well-trained model achieves the best RMSE of 0.290 N in force and 0.0084 Nm in torque. Furthermore, the proposed force/torque estimation network is integrated with a force-feedback policy for adaptive grasping tasks. The experimental results demonstrate the effectiveness of the proposed method and its potential application in robotic grasping and manipulation tasks.
Chaofan Zhang, Shaowei Cui, Yinghao Cai, Jingyi Hu, Rui Wang 0031, Shuo Wang 0001
IROS3
2022 VGPN: 6-DoF Grasp Pose Detection Network Based on Hough Voting
abstract
In this paper, we propose a novel Voting based Grasp Pose Network (VGPN) to detect 6-DoF grasps in cluttered scenes. The motivation of this paper is that local object geometry can provide useful clues about where the object can be grasped. Generated by the sampled seed points from raw point cloud, the votes allow seed points in different object regions to contribute to locations where the object can be grasped. Geometric features from various local regions are aggregated to generate grasps in a more confident and dense space, which enables grasp prediction utilizing more global context features. The search space of grasp pose detection is also greatly reduced. Experimental results on both simulation and real-world environments show that our proposed method outperforms state-of-the-art approaches in terms of both success rate and coverage of the ground truth grasps. The objects can be grasped with fewer attempts which is critical in real-world applications.
Yinghao Cai, Tao Lu 0006, Shuo Wang 0001
IROS2
2022 Manipulation skill learning on multi-step complex task based on explicit and implicit curriculum learning
Naijun Liu, Tao Lu 0006, Yinghao Cai, Rui Wang 0031, Shuo Wang 0001
Sci. China Inf. Sci.3
2021 Hierarchical Learning from Demonstrations for Long-Horizon Tasks
abstract
Although reinforcement learning (RL) has achieved great success in robotic manipulation skills learning, it is still challenging for long-horizon tasks. Combining RL with demonstrations is an effective solution. In this paper, we propose a novel hierarchical learning from demonstrations method for long-horizon tasks, which leverages (i) object-centered segmentation of demonstrations to automatically segment the teaching trajectories into episodes. (ii) a bi-level hierarchical imitation learning method with a parallel training mechanism to train the two-level policies simultaneously. Experimental results on three challenging long-horizon tasks with sparse rewards show that our proposed method significantly outperforms state-of-art approaches in terms of both sample-efficiency and success rate. Moreover, our method is the only one which achieves satisfactory performance in tasks of multi-object stack and multi-object push&stack.
Boyao Li, Tao Lu 0006, Yinghao Cai, Shuo Wang 0001
ICRA4
2021 DIMSAN: Fast Exploration with the Synergy between Density-based Intrinsic Motivation and Self-adaptive Action Noise
abstract
Exploration in environments with sparse rewards remains a challenging problem in Deep Reinforcement Learning (DRL). For the off-policy method, it usually needs a large number of training samples. With the growing dimensions of state and action space, this method becomes more and more sample-inefficient. In this paper, we propose a novel fast exploration method for off-policy reinforcement learning, called Density-based Intrinsic Motivation and Self-adaptive Action Noise (DIMSAN). Our main contribution is twofold: (1) We propose a Density-based Intrinsic Motivation (DIM) method. It introduces a new intrinsic-reward generation mechanism based on samples’ density estimation during experience replay and encourages the agent to seek novel and unfamiliar states. (2) We propose a Self-adaptive Action Noise (SAN) to deal with the exploration-exploitation tradeoffs, which could automatically change the exploration step through adding adaptive action space noise. The synergy between DIM and SAN could guide the agent to search the state and action space with high efficiency. We evaluate our method on the benchmark manipulation tasks and the designed challenging ones. Empirical results show that our method outperforms the existing methods in terms of convergence speed and sample efficiency, especially in challenging tasks.
Boyao Li, Tao Lu 0006, Yinghao Cai, Shuo Wang 0001
ICRA5
2021 3DTDesc: learning local features using 2D and 3D cues
Xiaoxia Xing, Yinghao Cai, Tao Lu 0006, Dayong Wen
Mach. Vis. Appl.2
2021 Correction to: 3DTDesc: learning local features using 2D and 3D cues
Xiaoxia Xing, Yinghao Cai, Tao Lu 0006, Dayong Wen
Mach. Vis. Appl.2
2020 Dynamic Guided Network for Monocular Depth Estimation
abstract
Self-attention and encoder-decoder have been widely used in the deep neural network for monocular depth estimation. The self-attention mechanism is capable of capturing long-range dependencies by computing the representation of each image position by a weighted sum of the features at all positions, while the encoder-decoder can capture detailed structural information by gradually recovering spatial information. In this work, we combine the advantages of both methods. Specifically, our proposed model, DGNet, extends EMANet [1] by adding an effective decoder module to progressively refine the coarse depth map. In the decoder stage, we design a dynamic guided upsampling module that employs dynamically generated kernel conditioned on low-level features to guide the upsampling of the coarse depth map. Experimental results demonstrate that our method obtains higher accuracy and generates visually pleasant depth maps.
Xiaoxia Xing, Yinghao Cai, Tao Lu 0006, Dayong Wen
ICPR2
2020 ACDER: Augmented Curiosity-Driven Experience Replay
abstract
Exploration in environments with sparse feed-back remains a challenging research problem in reinforcement learning (RL). When the RL agent explores the environment randomly, it results in low exploration efficiency, especially in robotic manipulation tasks with high dimensional continuous state and action space. In this paper, we propose a novel method, called Augmented Curiosity-Driven Experience Replay (ACDER), which leverages (i) a new goal-oriented curiosity-driven exploration to encourage the agent to pursue novel and task-relevant states more purposefully and (ii) the dynamic initial states selection as an automatic exploratory curriculum to further improve the sample-efficiency. Our approach complements Hindsight Experience Replay (HER) by introducing a new way to pursue valuable states. Experiments conducted on four challenging robotic manipulation tasks with binary rewards, including Reach, Push, Pick&Place and Multi-step Push. The empirical results show that our proposed method significantly outperforms existing methods in the first three basic tasks and also achieves satisfactory performance in multi-step robotic task learning.
Boyao Li, Tao Lu 0006, Yinghao Cai, Shuo Wang 0001
ICRA5
2019 Localizing Discriminative Visual Landmarks for Place Recognition
abstract
We address the problem of visual place recognition with perceptual changes. The fundamental problem of visual place recognition is generating robust image representations which are not only insensitive to environmental changes but also distinguishable to different places. Taking advantage of the feature extraction ability of Convolutional Neural Networks (CNNs), we further investigate how to localize discriminative visual landmarks that positively contribute to the similarity measurement, such as buildings and vegetations. In particular, a Landmark Localization Network (LLN) is designed to indicate which regions of an image are used for discrimination. Detailed experiments are conducted on open source datasets with varied appearance and viewpoint changes. The proposed approach achieves superior performance against state-of-the-art methods.
Zhe Xin, Yinghao Cai, Tao Lu 0006, Xiaoxia Xing, Shaojun Cai, Jixiang Zhang 0001
ICRA2
2019 Self-modeling Tracking Control of Crawler Fire Fighting Robot Based on Causal Network*
abstract
In this paper, a self-modeling method based on a causal network is proposed for the tracking control of the Crawler Fire Fighting Robot (CFFR). The method mainly consists of two parts, one is a motion model, based on data driving, learning to establish the correspondence between control signal sequence and vehicle motion, estimating the motion state of the next moment from historical data, eliminating complex CFFR modeling. The other is the tracking network. Based on the simulation data of the motion model, the relationship between the target trajectory and the current control command is learned, which simplifies the design and cumbersome tuning of the complex controller. The effectiveness of the proposed method is verified in both simulated and real-world environments. Qualitative and quantitative experimental results verify the accuracy of the tracking.
Wenkai Chang, Caiyun Yang, Tao Lu 0006, Yinghao Cai, Shuo Wang 0001
IROS5
2018 3DTNet: Learning Local Features Using 2D and 3D Cues
abstract
We present an approach to learn 3D local descriptor by combining both 2D texture and 3D geometric information, which can be used to register partial 3D data for a variety of vision applications. Unlike previous approaches which simply concatenate features learned from multiple sources into one feature descriptor, we learn 2D and 3D feature representations jointly. We design a network, 3DTNet with an architecture particularly designed for learning robust local feature representation leveraging both texture and geometric information. Two types of information are interacted with each other which results in more robust and stable feature representation. Finally, feature representations of multi-scale neighborhoods are aggregated to further improve the performance of feature matching. Extensive experimental results show that our method outperforms state-of-art 2D or 3D descriptors in terms of both accuracy and efficiency.
Xiaoxia Xing, Yinghao Cai, Tao Lu 0006, Shaojun Cai, Dayong Wen
3DV2
2018 Visual Localization in Changing Environments using Place Recognition Techniques
abstract
This paper proposes a visual localization system combining Convolutional Neural Networks (CNNs) and sparse point features to estimate the 6-DOF pose of the robot. The challenges of visual localization across time lie in that the same place captured across time appears dramatically different due to different illumination and weather conditions, viewpoint variations and dynamic objects. In this paper, a novel CNN-based place recognition approach is proposed, which requires no time-consuming feature generation process and no task-specific training. Moreover, we demonstrate that the rich semantic context information obtained from place recognition can greatly improve the subsequent feature matching process for pose estimation. The semantic constraint performs much better than traditional Bag-of-Words based methods for establishing correspondences between the query image and the map. To evaluate the robustness of the algorithm, the proposed system is integrated into ORB-SLAM2 and verified on the data collected over various illumination and weather conditions. Extensive experimental results show that even with weak ORB descriptors, the proposed system can significantly improve the success rate of localization under severe appearance changes.
Zhe Xin, Yinghao Cai, Shaojun Cai, Jixiang Zhang 0001
ICPR2
2018 Probabilistic Voting for Sequence Based Visual Place Recognition
abstract
Visual place recognition is the task of recognizing the query image in a set of dataset images. It is a challenging problem in computer vision due to frequent and unpredictable environmental changes. In this paper, a novel approach is proposed for visual place recognition. We consider the problem of visual place recognition as a probabilistic voting problem on coherent image sequences. According to the co-visibility relationship of images in the dataset, each query can be represented by a categorical variable. Therefore, the whole sequence is the distribution of several independent and non-identical categorical variables. Introducing the probabilistic framework not only removes the need for heuristic parameters but also recognizes location efficiently and effectively. Two widely used datasets are used to evaluate the performance of the proposed method. The probabilistic voting algorithm achieves superior performance compared with state-of-the-art methods and satisfies the real-time requirement.
Zhe Xin, Yinghao Cai, Jixiang Zhang 0001
ICPR2
2016 Optimization of radial distortion self-calibration for structure from motion from uncalibrated UAV images
abstract
Structure from motion (SfM) and self-calibration from images of unknown radial distortions could fail under some critical configurations and produce distorted reconstruction results. In this paper, we propose an effective approach to optimize the estimation of radial distortion coefficient by taking full advantage of GPS information, which allows for more accurate SfM results. A feedback function is designed as the metric to indicate the magnitude of the distortion error. Heuristic search strategies are applied to search for the optimal distortion coefficient. Extensive experimental results show that our approach can effectively reduce the distorted deformation error and improve the estimation accuracy of the distortion coefficient.
Yonglu Li 0002, Yinghao Cai, Dayong Wen
ICPR2
2016 Persistent people tracking and face capture using a PTZ camera
Yinghao Cai, Gérard G. Medioni
Mach. Vis. Appl.1
2015 The multi-strand graph for a PTZ tracker
abstract
High-resolution images can be used to resolve matching ambiguities between trajectory fragments (tracklets), which is one of the main challenges in multiple target tracking. A PTZ camera, which can pan, tilt and zoom, is a powerful and efficient tool that offers both close-up views and wide area coverage on demand. The wide-area view makes it possible to track many targets while the close-up view allows individuals to be identified from high-resolution images of their faces. A central component of a PTZ tracking system is a scheduling algorithm that determines which target to zoom in on. In this paper we study this scheduling problem from a theoretical perspective, where the high resolution images are also used for tracklet matching. We propose a novel data structure, the Multi-Strand Tracking Graph (MSG), which represents the set of tracklets computed by a tracker and the possible associations between them. The MSG allows efficient scheduling as well as resolving - directly or by elimination - matching ambiguities between tracklets. The main feature of the MSG is the auxiliary data saved in each vertex, which allows efficient computation while avoiding time-consuming graph traversal. Synthetic data simulations are used to evaluate our scheduling algorithm and to demonstrate its superiority over a naïve one.
Shachaf Melman, Yael Moses, Gérard G. Medioni, Yinghao Cai
AVSS4
2015 Learning predictable binary codes for face indexing
Ran He 0001, Yinghao Cai, Tieniu Tan, Larry Davis 0001
Pattern Recognit.2
2014 Exploring context information for inter-camera multiple target tracking
abstract
In this paper, we present a new solution to inter-camera multiple target tracking with non-overlapping fields of view. The identities of people are maintained when they are moving from one camera to another. Instead of matching snapshots of people across cameras, we mainly explore what kind of context information from videos can be used for inter-camera tracking. We introduce two kinds of context information, spatio-temporal context and relative appearance context in this paper. The spatio-temporal context indicates a way of collecting samples for discriminative appearance learning where target-specific appearance models are learned to distinguish different people from each other. The relative appearance context models inter-object appearance similarities for people walking in proximity. The relative appearance model helps disambiguate individual appearance matching across cameras. We show improved performance with context information for inter-camera tracking. Our method achieves promising results in two crowded scenes compared with state-of-art methods.
Yinghao Cai, Gérard G. Medioni
WACV1
2013 Towards a practical PTZ face detection and tracking system
abstract
We address the problem of automatic face detection and tracking in uncontrolled scenarios using a pan-tilt-zoom (PTZ) network camera, which could prove most helpful in forensic applications. The detected faces are associated with the corresponding people and trajectories. The dynamic nature of real-world scenarios and real-time restrictions complicate our task. Different from previous work which use a mixture of wide angle cameras and PTZ cameras, we explore the limits to what can be expected from a single PTZ camera. The system first detects and tracks pedestrians in zoomed-out mode, then selects, using a scheduler, a person to zoom in to. After zoom in, we come back to wide area mode, and solve the person-to-person, face-to-person and face-to-face data association problems. Extensive experiments in challenging indoor and outdoor uncontrolled conditions demonstrate the effectiveness of the proposed system.
Yinghao Cai, Gérard G. Medioni, Thang Ba Dinh
WACV1
2010 Recovering the Topology of Multiple Cameras by Finding Continuous Paths in a Trellis
abstract
In this paper, we propose an unsupervised method for recovering the topology of multiple cameras with non-overlapping fields of view. The nodes in the topology graph are defined as entry/exit zones in each camera while the connectivity between nodes is inferred through finding continuous paths in a trellis where appearance information and temporal information of moving objects are encoded. Unlike previous methods which assume a single mode transition distribution between nodes, our method is capable of dealing with multi-modal transition situations when both cars and pedestrians are in the scene. Results on simulated and real-life datasets demonstrate the effectiveness of the proposed method.
Yinghao Cai, Kaiqi Huang, Tieniu Tan, Matti Pietikäinen
ICPR1
2010 Matching Groups of People by Covariance Descriptor
abstract
In this paper, we present a new solution to the problem of matching groups of people across multiple non-overlapping cameras. Similar to the problem of matching individuals across cameras, matching groups of people also faces challenges such as variations of illumination conditions, poses and camera parameters. Moreover, people often swap their positions while walking in a group. In this paper, we propose to use covariance descriptor in appearance matching of group images. Covariance descriptor is shown to be a discriminative descriptor which captures both appearance and statistical properties of image regions. Furthermore, it presents a natural way of combining multiple heterogeneous features together with a relatively low dimensionality. Experimental results on two different datasets demonstrate the effectiveness of the proposed method.
Yinghao Cai, Valtteri Takala, Matti Pietikäinen
ICPR1
2010 Boosting Clusters of Samples for Sequence Matching in Camera Networks
abstract
This study introduces a novel classification algorithm for learning and matching sequences in view independent object tracking. The proposed learning method uses adaptive boosting and classification trees on a wide collection (shape, pose, color, texture, etc.) of image features that constitute a model for tracked objects. The temporal dimension is taken into account by using k-mean clusters of sequence samples. Most of the utilized object descriptors have a temporal quality also. We argue that with a proper boosting approach and decent number of reasonably descriptive image features it is feasible to do view-independent sequence matching in sparse camera networks. The experiments on real-life surveillance data support this statement.
Valtteri Takala, Yinghao Cai, Matti Pietikäinen
ICPR2
2008 Matching tracking sequences across widely separated cameras
abstract
In this paper, we present a new solution to the problem of matching tracking sequences across different cameras. Unlike snapshot-based appearance matching which matches objects by a single image, we focus on sequence matching to alleviate the uncertainties brought by segmentation errors and partial occlusions. By incorporating multiple snapshots of the same object, the influence of the variation is alleviated. At the training stage, given the sequence of a queried person under one camera, the appearance model is formulated by concatenating feature vectors with the majority of votes over the sequence. At the testing stage, Bayesian inference is incorporated into the identification framework to accumulate the temporal information in the sequence. Experimental results demonstrate the effectiveness of the proposed method.
Yinghao Cai, Kaiqi Huang, Tieniu Tan
ICIP1
2008 Human appearance matching across multiple non-overlapping cameras
abstract
In this paper, we present a new solution to the problem of appearance matching across multiple non-overlapping cameras. Objects of interest, pedestrians are represented by a set of region signatures centered at points sampled from edges. The problem of frame-to-frame appearance matching is formulated as finding corresponding points in two images as minimization of a cost function over the space of correspondence. The correspondence problem is solved under integer optimization framework where the cost function is determined by similarity of region signatures as well as geometric constraints between points. Experimental results demonstrate the effectiveness of the proposed method.
Yinghao Cai, Kaiqi Huang, Tieniu Tan
ICPR1
2007 Continuously Tracking Objects Across Multiple Widely Separated Cameras
Yinghao Cai, Wei Chen 0012, Kaiqi Huang, Tieniu Tan
ACCV (1)1
2007 Real-Time Moving Object Classification with Automatic Scene Division
abstract
We address the problem of moving object classification. Our aim is to classify moving objects of traffic scene videos into pedestrians, bicycles and vehicles. Instead of supervised learning and manual labeling of large training samples, our classifiers are initialized and refined online automatically. With efficient features extracted and organized, the approach can be real-time and achieve high classification accuracy. Once the view or scene changes detected, the algorithm can automatically refine the classifiers and adapt them to new environments. Experimental results demonstrate the effectiveness and robustness of the proposed approach.
Zhaoxiang Zhang 0001, Yinghao Cai, Kaiqi Huang, Tieniu Tan
ICIP (5)2