EDBT 2026 Demo / reviewers in the wild / expert
Jinshi Cui
dblp:21/2565
· DBLP profile ↗
46ranked-venue papers
7as first author
11since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 36 · 6 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 21 · 2 first-author · 8 since 2021Systems, architecture and hardware · 10 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Databases, data management, data science and information retrieval · 2Human-computer interaction and ubiquitous computing · 2 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Robust Chronic Stress Detection on Consumer-Grade Wearable Data: A Two-Stage Learning Framework
An Dai, Jinshi Cui |
HealthCom | 2 |
| 2025 | Gaze4ASD: A Novel Dataset and Visual Saliency Map-Based Method for Autism ScreeningabstractAutism spectrum disorder (ASD) affects social communication and behavior, highlighting the need for accessible and reliable screening methods. While eye gaze-based approaches can potentially address the limitations of traditional methods, they face challenges such as small sample sizes and subtle gaze differences between ASD and typically developing (TD) children in existing datasets. To overcome these barriers, we developed Gaze4ASD1, an eye gaze dataset comprising eye gaze data from 133 TD and 33 ASD children using 30 high-discrimination image stimuli. Leveraging this dataset, we proposed a visual saliency map-based method that transforms eye gaze data into saliency maps, incorporating novel feature extraction and fusion techniques based on ASD/TD gaze patterns, achieving 88.5% accuracy. Psychological paradigms further validated its ability to capture ASD-specific gaze patterns, establishing a scalable foundation for ASD screening. Yizhang Yang, Jinshi Cui, Junshi Lu |
ICME | 2 |
| 2025 | Multitask Stress Detection on Low-frequency Physiological SignalsabstractStress is an inevitable aspect of modern life. Real-time stress detection methods can assist individuals in regulating their mental states and enhancing psychological well-being. This study utilized data from over 180 participants, who were divided into biological stressor, social stressor, and control groups. Physiological indicators included respiratory sinus arrhythmia (RSA), skin conductance level (SCL), and cortisol levels. Zdevelop new algorithms on low-frequency wearable physiological data. Considering the problem of high noise in this type of data, we first proposed and verified the effectiveness of the algorithm on high-frequency physiological instrument data, and then transplanted the effective algorithm to low-frequency wearable data. Ultimately, this paper provides a comprehensive dataset covering various types of stressors and demonstrates encouraging results in stress detection, achieving a precision of 83.3% on watch data. Yiming Rong, Jinshi Cui, Guosong Gao |
IJCNN | 2 |
| 2024 | Image-Feature Weak-to-Strong Consistency: An Enhanced Paradigm for Semi-supervised Learning
Zhiyu Wu, Jinshi Cui |
ECCV (15) | 2 |
| 2024 | AllMatch: Exploiting All Unlabeled Data for Semi-Supervised Learning
Zhiyu Wu, Jinshi Cui |
IJCAI | 2 |
| 2023 | BERT-ERC: Fine-Tuning BERT Is Enough for Emotion Recognition in ConversationabstractPrevious works on emotion recognition in conversation (ERC) follow a two-step paradigm, which can be summarized as first producing context-independent features via fine-tuning pretrained language models (PLMs) and then analyzing contextual information and dialogue structure information among the extracted features. However, we discover that this paradigm has several limitations. Accordingly, we propose a novel paradigm, i.e., exploring contextual information and dialogue structure information in the fine-tuning step, and adapting the PLM to the ERC task in terms of input text, classification structure, and training strategy. Furthermore, we develop our model BERT-ERC according to the proposed paradigm, which improves ERC performance in three aspects, namely suggestive text, fine-grained classification module, and two-stage training. Compared to existing methods, BERT-ERC achieves substantial improvement on four datasets, indicating its effectiveness and generalization capability. Besides, we also set up the limited resources scenario and the online prediction scenario to approximate real-world scenarios. Extensive experiments demonstrate that the proposed paradigm significantly outperforms the previous one and can be adapted to various scenes. Xiangyu Qin, Zhiyu Wu, Yanran Li, Jian Luan 0001, Bin Wang 0004, Li Wang 0114, Jinshi Cui |
AAAI | 8 |
| 2023 | LA-Net: Landmark-Aware Learning for Reliable Facial Expression Recognition under Label NoiseabstractFacial expression recognition (FER) remains a challenging task due to the ambiguity of expressions. The derived noisy labels significantly harm the performance in real-world scenarios. To address this issue, we present a new FER model named Landmark-Aware Net (LA-Net), which leverages facial landmarks to mitigate the impact of label noise from two perspectives. Firstly, LA-Net uses landmark information to suppress the uncertainty in expression space and constructs the label distribution of each sample by neighborhood aggregation, which in turn improves the quality of training supervision. Secondly, the model incorporates landmark information into expression representations using the devised expression-landmark contrastive loss. The enhanced expression feature extractor can be less susceptible to label noise. Our method can be integrated with any deep neural network for better training supervision without introducing extra inference costs. We conduct extensive experiments on both in-the-wild datasets and synthetic noisy datasets and demonstrate that LA-Net achieves state-of-the-art performance. Zhiyu Wu, Jinshi Cui |
ICCV | 2 |
| 2022 | Psychology-Inspired Interaction Process Analysis based on Time SeriesabstractInteraction portrays the process of exchanging emotions and states between individuals in a specific temporal and spatial context. Interactions are layered on a temporal scale, including social signals at an instant, intrinsic structures of a process, and enduring relationships between the participants. Existing studies of affective computing focus on the identification of social signals at specific moments, which are precise but limited to the microscopic level; while psychological studies focus on the process and relationship analysis, which are semantic but qualitative. In this paper, we focused on interaction process analysis. By combining psychological knowledge and time series analysis techniques, we proposed quantitative methods for features of personal status, dyad correlation, synchrony coefficients, and ‘lead-follow’ structures. We conducted experiments on the mother-infant interaction dataset. We first compared the differences in interaction patterns between social signals. We found that gaze direction in the free-play phase, emotion synchrony in the recov-ery phase, and engagement in both phases, play an important role in differentiating attachment styles. We then explored the differences in interaction patterns. We found the ‘infant dropping while mother rising’ asynchrony coefficient and the ‘infant leads, mother follows’ structure related to secure attachment styles. Finally, we successfully predicted infant attachment formed one year later based on machine learning methods, confirming the effectiveness of our feature extraction method. Jiaheng Han, Honggai Li, Jinshi Cui, Qili Lan |
ICPR | 3 |
| 2021 | Dense Relation Distillation With Context-Aware Aggregation for Few-Shot Object DetectionabstractConventional deep learning based methods for object detection require a large amount of bounding box annotations for training, which is expensive to obtain such high quality annotated data. Few-shot object detection, which learns to adapt to novel classes with only a few annotated examples, is very challenging since the fine-grained feature of novel object can be easily overlooked with only a few data available. In this work, aiming to fully exploit features of annotated novel object and capture fine-grained features of query object, we propose Dense Relation Distillation with Context-aware Aggregation (DCNet) to tackle the few-shot detection problem. Built on the meta-learning based framework, Dense Relation Distillation module targets at fully exploiting support features, where support features and query feature are densely matched, covering all spatial locations in a feed-forward fashion. The abundant usage of the guidance information endows model the capability to handle common challenges such as appearance changes and occlusions. Moreover, to better capture scale-aware features, Context-aware Aggregation module adaptively harnesses features from different scales for a more comprehensive feature representation. Extensive experiments illustrate that our proposed approach achieves state-of-the-art results on PASCAL VOC and MS COCO datasets. Code will be made available at https://github.com/hzhupku/DCNet. Hanzhe Hu, Shuai Bai, Aoxue Li, Jinshi Cui, Liwei Wang 0001 |
CVPR | 4 |
| 2021 | Region-aware Contrastive Learning for Semantic SegmentationabstractRecent works have made great success in semantic segmentation by exploiting contextual information in a local or global manner within individual image and supervising the model with pixel-wise cross entropy loss. However, from the holistic view of the whole dataset, semantic relations not only exist inside one single image, but also prevail in the whole training data, which makes solely considering intra-image correlations insufficient. Inspired by recent progress in unsupervised contrastive learning, we propose the region-aware contrastive learning (RegionContrast) for semantic segmentation in the supervised manner. In order to enhance the similarity of semantically similar pixels while keeping the discrimination from others, we employ contrastive learning to realize this objective. With the help of memory bank, we explore to store all the representative features into the memory. Without loss of generality, to efficiently incorporate all training data into the memory bank while avoiding taking too much computation resource, we propose to construct region centers to represent features from different categories for every image. Hence, the proposed region-aware contrastive learning is performed in a region level for all the training data, which saves much more memory than methods exploring the pixel-level relations. The proposed RegionContrast brings little computation cost during training and requires no extra overhead for testing. Extensive experiments demonstrate that our method achieves state-of-the-art performance on three benchmark datasets including Cityscapes, ADE20K and COCO Stuff. Hanzhe Hu, Jinshi Cui, Liwei Wang 0001 |
ICCV | 2 |
| 2021 | Semi-Supervised Semantic Segmentation via Adaptive Equalization LearningabstractDue to the limited and even imbalanced data, semi-supervised semantic segmentation tends to have poor performance on some certain categories, e.g., tailed categories in Cityscapes dataset which exhibits a long-tailed label distribution. Existing approaches almost all neglect this problem, and treat categories equally. Some popular approaches such as consistency regularization or pseudo-labeling may even harm the learning of under-performing categories, that the predictions or pseudo labels of these categories could be too inaccurate to guide the learning on the unlabeled data. In this paper, we look into this problem, and propose a novel framework for semi-supervised semantic segmentation, named adaptive equalization learning (AEL). AEL adaptively balances the training of well and badly performed categories, with a confidence bank to dynamically track category-wise performance during training. The confidence bank is leveraged as an indicator to tilt training towards under-performing categories, instantiated in three strategies: 1) adaptive Copy-Paste and CutMix data augmentation approaches which give more chance for under-performing categories to be copied or cut; 2) an adaptive data sampling approach to encourage pixels from under-performing category to be sampled; 3) a simple yet effective re-weighting method to alleviate the training noise raised by pseudo-labeling. Experimentally, AEL outperforms the state-of-the-art methods by a large margin on the Cityscapes and Pascal VOC benchmarks under various data partition protocols. Code is available at https://github.com/hzhupku/SemiSeg-AEL. Hanzhe Hu, Fangyun Wei, Han Hu 0001, Qiwei Ye, Jinshi Cui, Liwei Wang 0001 |
NeurIPS | 5 |
| 2020 | Boundary-aware Graph Convolution for Semantic SegmentationabstractRecent works have made great progress in semantic segmentation by exploiting contextual information in a local or global manner with dilated convolutions, pyramid pooling or self-attention mechanism. However, few works have focused on harvesting boundary information to improve the segmentation performance. In order to enhance the feature similarity within the object and keep discrimination from other objects, we propose a boundary-aware graph convolution (BGC) module to propagate features within the object. The graph reasoning is performed among pixels of the same object apart from the boundary pixels. Based on the proposed BGC module, we further introduce the Boundary-aware Graph Convolution N et-work(BGCN et), which consists of two main components including a basic segmentation network and the BGC module, forming a coarse-to-fine paradigm. Specifically, the BGC module takes the coarse segmentation feature map as node features and boundary prediction to guide graph construction. After graph convolution, the reasoned feature and the input feature are fused together to get the refined feature, producing the refined segmentation result. We conduct extensive experiments on three popular semantic segmentation benchmarks including Cityscapes, PASCAL VOC 2012 and COCO Stuff, and achieve state-of-the-art performance on all three benchmarks. Hanzhe Hu, Jinshi Cui, Hongbin Zha |
ICPR | 2 |
| 2019 | Pose-Aware Face Alignment based on CNN and 3DMM
Songjiang Li, Honggai Li, Jinshi Cui, Hongbin Zha |
BMVC | 3 |
| 2018 | Recognition of Infants' Gaze Behaviors and EmotionsabstractThis paper proposes a system for recognition of infants' gaze behaviors and emotions from videos. In the current work, researchers believed that the information of eye region is crucial for gaze behavior recognition, and emotion recognition mostly depends on the appearance of face. However, because the differentiation of all parts of infant's body has not finished yet, we found that infants always express their intentions and emotions using their whole body, especially moving their heads. Therefore, we incorporate the head pose information as features into the gaze behavior recognition, and we extract the gaze features to improve the recognition of infants' emotions. In addition, we combine several Deep Neural Networks, which can not only capture the details of the images very well, but also make full use of the temporal features. In order to recognize infants' gaze behaviors, we design a feature-extraction convolutional neural network which can obtain the features of infants' gaze direction, then we feed these features with head pose into the next gaze behavior recurrent neural network. Moreover, we combine the features of facial express and gaze behavior to characterize infants' emotions, and expand this system with an emotion recurrent neural network. In the end, we achieve the recognition accuracy 98.31% and 94.71% respectively on our data set. Bikun Yang, Jinshi Cui, Yuqiang Tong, Hongbin Zha |
ICPR | 2 |
| 2017 | Improving Children's Gaze Prediction via Separate Facial Areas and Attention Shift CueabstractTo predict and assess visual attention, saliencybased visual attention modeling is a popular approach. However, state-of-the-art models are developed for adults, in which children are not considered. Additionally, these models consider neither social cues like face, nor attention learning cues. The face is a vital part of visual attention. Psychological studies reveal that sub-facial areas are different in visual attention. Some models highlight faces in social scenes, but sub-facial areas are not taken into account. Attention learning reveals internal processing of visual attention. By learning how the cognitive system deals with visual stimuli, it is possible to predict visual attention behavior. In this paper, we propose a multilevel visual attention model to predict fixations of children when watching a talking face. Based on traditional saliency maps, the proposed model includes both separate facial areas and attention shift cue. An eye-tracking experiment is conducted to evaluate the model. Results show that the proposed model significantly outperforms conventional models in talking face scenes. Songjiang Li, Wen Cui, Jinshi Cui, Hongbin Zha |
FG | 3 |
| 2017 | Specialized gaze estimation for children by convolutional neural network and domain adaptationabstractChildren's social gaze behavior modeling and evaluation has obtained increasing attentions in various research areas. In psychology research, eye gaze behavior is very important to developmental disorders diagnosis and assessment. In robotics area, gaze interaction between children and robots also draws more and more attention. However, there exists no specific gaze estimator for children in social interaction context. Current approaches usually use models trained with adults' data to estimate children's gaze. Since gaze behaviors and eye appearances of children are different from those of adults, the current approaches, especially those with free-calibration assumptions which are utilized in usual human-robot interaction systems, will result in big errors. Note that children data is difficult to collect and label, so directly learning from children data is hard to achieve. We propose a new system to solve this problem, which combines a CNN feature extractor trained from adult data and a domain adaptation unit using geodesic flow kernel to adapt the source domain (adults) classifier to the target domain (children). Our system performs well in children's gaze estimation. Wen Cui, Jinshi Cui, Hongbin Zha |
ICIP | 2 |
| 2017 | Ego-centric traffic behavior understanding through multi-level vehicle trajectory analysisabstractThis study proposes a multi-level trajectory analysis method for modeling traffic behavior from an ego-centric view, where on-road vehicle trajectories are collected based on the authors' previous studies of an on-board system consisting of multiple 2D lidar sensors. From an input set of trajectories, a set of hot regions (topics) that trajectory points most frequently present are first discovered using a sticky HDP-HMM; then, the major paths of the trajectories' transitions across different hot regions are extracted by recursively mining frequent subsequences of topics; and finally, paths are modeled using a hierarchical hidden Markov model (HHMM), where the intra-path dynamics is represented using an HMM, in which each state corresponds to a hot region, while the inter-path transition is assumed to be Markovian. The model could be used for behavior prediction, i.e. whenever a vehicle is detected in a scene, predicting which route it will probably follow and how its trajectory will probably develop over time, which is essential to interpreting the potential risks for longer time horizons. Experiments are conducted using a large set of vehicle trajectories collected from motorways in Beijing, and promising results are presented. Donghao Xu, Huijing Zhao, Jinshi Cui, Hongbin Zha, Franck Guillemard, Stéphane Géronimi, François Aioun |
ICRA | 4 |
| 2017 | Object Tracking via Temporal Consistency Dictionary LearningabstractSparse representation-based methods have been successfully applied to visual tracking. However, complex and inefficient optimization limits their deployment in practical tracking scenarios. In this paper, we propose a temporal consistency dictionary learning tracking algorithm to enable efficient dictionary learning and tracking executive. First, we present an objective function which introduces the fixed dictionary and variance dictionary to reconstruct the object's appearance. In particular, the proposed method takes the temporal consistency into account by adding a regularization term into the objective function to constrain the difference of object appearance at adjacent frames. Then the optimization problem is solved in an iteration way. Moreover, the proposed method can encode the object's local structural information, and the local patches from the same candidate altogether for a global appearance representation. Second, we develop an effective observation likelihood function based on the proposed model. It takes the influence of patches with large reconstruction errors into consideration, thereby, alleviating the drifting of the object. Finally, we present an appearance updating strategy to adapt to the object's appearance variations by the online dictionary learning. Experimental evaluations on the TB50 and TB100 datasets show that the proposed tracking method outperforms sparse representation related visual tracking as well as other state-of-the-art tracking methods. Xu Cheng 0003, Yifeng Zhang 0001, Jinshi Cui, Lin Zhou 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 3 |
| 2014 | Joint estimation of head pose and visual focus of attentionabstractHead pose is an important indicator of a person's visual focus of attention (VFoA). A traditional way to recognize VFoA is to consider accurate head pose or gaze estimations. However, these estimations usually degrade drastically in middle or low resolution video data. In this paper, a joint estimation of head pose and VFoA is proposed to address this issue; both head pose and VFoA are iteratively refined until convergence. This approach is evaluated in a specific scenario involving children around a table playing together with toys. Datasets are acquired and annotated by psychologists in Peking university. The experimental results demonstrate the usefulness of the join estimation process to recognize visual focus of attention in middle resolution video sequences. Yingning Huang, Dingrui Duan, Jinshi Cui, Franck Davoine, Hongbin Zha |
ICIP | 3 |
| 2014 | Calibration method for multiple 2D LIDARs systemabstractMany robotic and mobile mapping systems have been developed using multiple 2D LIDARs (briefly multi-LIDAR system) to sense environment. In such systems, extrinsic calibration of all LIDARs is essential for making collaborative use of the data from different sensors. This research aims at developing a calibration method for multi-LIDAR systems at the general scene, such as an outdoor place or an underground parking-lot, without modification to environment by putting calibration targets. In this paper, the calibration method is proposed by aligning the 3D data of different LIDARs. They are concerned at two-levels: 1) reference calibration, i.e. finding the transformation from a reference LIDAR to the platform frame; 2) multi-LIDAR calibration, i.e. finding the LIDARs' relative geometries by referring to the reference one. The method is examined in calibrating the multiple 2D LIDARs on an intelligent vehicle platform POSS-V, where the data collected through a driving in an underground parking-lot are registered to find sensors' geometry. Calibration accuracy is examined by comparing with a CAD model of the scene, which was measured by using a total station. Mengwen He, Huijing Zhao, Jinshi Cui, Hongbin Zha |
ICRA | 3 |
| 2014 | Monocular visual localization using road structural featuresabstractPrecise localization is an essential issue for autonomous driving applications, where GPS-based systems are challenged to meet requirements such as lane-level accuracy. This paper introduces a new visual-based localization approach in dynamic traffic environments, focusing on and exploiting properties of structured roads like straight roads or intersections. Such environments show several line segments on lane markings, curbs, poles, building edges, etc., which demonstrate the road's longitude, latitude and vertical directions. Based on this observation, we define a Road Structural Feature (RSF) as sets of segments along three perpendicular axes together with feature points. At each video frame, the proper road structure (or multiple road structures in case of an intersection) is predicted based on the geometric information given by a 2D map. The RSF is then detected from line segments and points extracted from the image, and used to estimate the pose of the vehicle. Experiments are conducted using video streams collected on major roads in downtown Beijing, which are structured and with intense dynamic traffic. GPS/IMU data have been collected and synchronized with the video streams as a reference in validation. The results show good performance compared with that of a more traditional visual odometry method. Future work will be addressed on using visual approach to improve GPS localization accuracy. Huijing Zhao, Franck Davoine, Jinshi Cui, Hongbin Zha |
Intelligent Vehicles Symposium | 4 |
| 2013 | Pairwise LIDAR calibration using multi-type 3D geometric features in natural sceneabstractIt has become a well-known technology that 3D measurement of a large environment could be achieved by using a number of 2D LIDARs on a mobile platform. In such a system, calibration is essential for making collaborative use of different LIDAR data, while existing methods usually require modifications to the environments, such as putting calibration targets, or rely on special facilities, which is labor intensive and put many restrictions to potential applications. This research aims at developing a calibration method for multiple 2D LIDAR sensing systems, which could be conducted in a general outdoor environment using the features of a nature scene. Special focus is cast on solving the noisy sensing in a complex environment and the occlusions caused by largely different sensor viewpoints. A multi-type geometric feature based calibration algorithm is proposed, which extracts the features such as points, lines, planes and quadrics from the 3D points of each LIDAR sensing. Transformation parameters from each sensor to the frame on a moving platform is estimated by matching the multi-type features. Experiments are conducted using the data sets of an intelligent vehicle platform (POSS-V) through a driving in the campus of Peking University. Results of calibrating two LIDAR sensors with largely different viewpoints are presented, and the accuracy and robustness concerning noisy feature extractions are examined intensively. Mengwen He, Huijing Zhao, Franck Davoine, Jinshi Cui, Hongbin Zha |
IROS | 4 |
| 2013 | Laser-based tracking of multiple interacting pedestrians via on-line learning
Xuan Song 0001, Jinshi Cui, Huijing Zhao, Hongbin Zha, Ryosuke Shibasaki |
Neurocomputing | 2 |
| 2013 | A fully online and unsupervised system for large and high-density area surveillance: Tracking, semantic scene learning and abnormality detectionabstractFor reasons of public security, an intelligent surveillance system that can cover a large, crowded public area has become an urgent need. In this article, we propose a novel laser-based system that can simultaneously perform tracking, semantic scene learning, and abnormality detection in a fully online and unsupervised way. Furthermore, these three tasks cooperate with each other in one framework to improve their respective performances. The proposed system has the following key advantages over previous ones: (1) It can cover quite a large area (more than 60×35m), and simultaneously perform robust tracking, semantic scene learning, and abnormality detection in a high-density situation. (2) The overall system can vary with time, incrementally learn the structure of the scene, and perform fully online abnormal activity detection and tracking. This feature makes our system suitable for real-time applications. (3) The surveillance tasks are carried out in a fully unsupervised manner, so that there is no need for manual labeling and the construction of huge training datasets. We successfully apply the proposed system to the JR subway station in Tokyo, and demonstrate that it can cover an area of 60×35m, robustly track more than 150 targets at the same time, and simultaneously perform online semantic scene learning and abnormality detection with no human intervention. Xuan Song 0001, Xiaowei Shao, Quanshi Zhang, Ryosuke Shibasaki, Huijing Zhao, Jinshi Cui, Hongbin Zha |
ACM Trans. Intell. Syst. Technol. | 6 |
| 2013 | An online system for multiple interacting targets tracking: Fusion of laser and vision, tracking and learningabstractMultitarget tracking becomes significantly more challenging when the targets are in close proximity or frequently interact with each other. This article presents a promising online system to deal with these problems. The novelty of this system is that laser and vision are integrated with tracking and online learning to complement each other in one framework: when the targets do not interact with each other, the laser-based independent trackers are employed and the visual information is extracted simultaneously to train some classifiers online for “possible interacting targets”. When the targets are in close proximity, the classifiers learned online are used alongside visual information to assist in tracking. Therefore, this mode of cooperation not only deals with various tough problems encountered in tracking, but also ensures that the entire process can be completely online and automatic. Experimental results demonstrate that laser and vision fully display their respective advantages in our system, and it is easy for us to obtain a good trade-off between tracking accuracy and the time-cost factor. Xuan Song 0001, Huijing Zhao, Jinshi Cui, Xiaowei Shao, Ryosuke Shibasaki, Hongbin Zha |
ACM Trans. Intell. Syst. Technol. | 3 |
| 2013 | Tracking Generic Human Motion via Fusion of Low- and High-Dimensional ApproachesabstractTracking generic human motion is highly challenging due to its high-dimensional state space and the various motion types involved. In order to deal with these challenges, a fusion formulation which integrates low- and high-dimensional tracking approaches into one framework is proposed. The low-dimensional approach successfully overcomes the high-dimensional problem of tracking the motions with available training data by learning motion models, but it only works with specific motion types. On the other hand, although the high-dimensional approach may recover the motions without learned models by sampling directly in the pose space, it lacks robustness and efficiency. Within the framework, the two parallel approaches, low- and high-dimensional, are fused via a probabilistic approach at each time step. This probabilistic fusion approach ensures that the overall performance of the system is improved by concentrating on the respective advantages of the two approaches and resolving their weak points. The experimental results, after qualitative and quantitative comparisons, demonstrate the effectiveness of the proposed approach in tracking generic human motion. Jinshi Cui, Ye Liu 0002, Yuandong Xu, Huijing Zhao, Hongbin Zha |
IEEE Trans. Syst. Man Cybern. Syst. | 1 |
| 2012 | Fusion of low-and high-dimensional approaches by trackers sampling for generic human motion tracking
Ye Liu 0002, Jinshi Cui, Huijing Zhao, Hongbin Zha |
ICPR | 2 |
| 2012 | Omni-directional detection and tracking of on-road vehicles using multiple horizontal laser scannersabstractThis research aims at generating an omnidirectional perception at the host vehicle's surroundings, extracting accurate and continuous motion trajectories of the nearby vehicles using low cost laser scanners. A system of detecting and tracking on-road vehicles using multiple laser scanners is developed, where focuses are cast on solving data association of simultaneous measurements from multiple sensors at different viewpoints, and state estimation in case of partial observations in dense dynamic situations. Experimental results in freeways in Beijing are presented, system efficiency is demonstrated, where motion trajectories describing driving behaviors such as overtaking, lane changing and other interactions between driving objects are captured. In addition, the accuracy in vehicle detection and tracking is examined using a reference vehicle with a ground truth GPS. Huijing Zhao, Chao Wang 0060, Franck Davoine, Jinshi Cui, Hongbin Zha |
Intelligent Vehicles Symposium | 5 |
| 2012 | Detection and Tracking of Moving Objects at Intersections Using a Network of Laser ScannersabstractIn our previous work, we reported a system that monitors an intersection using a network of horizontal laser scanners. This paper focuses on an algorithm for moving-object detection and tracking, given a sequence of distributed laser scan data of an intersection. The goal is to detect each moving object that enters the intersection; estimate state parameters such as size; and track its location, speed, and direction while it passes through the intersection. This work is unique, to the best of the authors' knowledge, in that the data is novel, which provides new possibilities but with great challenges; the algorithm is the first proposal that uses such data in detecting and tracking all moving objects that pass through a large crowded intersection with focus on achieving robustness to partial observations, some of which result from occlusions, and on performing correct data associations in crowded situations. Promising results are demonstrated using experimental data from real intersections, whereby, for 1063 objects moving through an intersection over 20 min, 988 are perfectly tracked from entrance to exit with an excellent tracking ratio of 92.9%. System advantages, limitations, and future work are discussed. Huijing Zhao, Jie Sha, Yipu Zhao, Junqiang Xi, Jinshi Cui, Hongbin Zha, Ryosuke Shibasaki |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2011 | Tracking Generic Human Motion via Fusion of Low- and High-Dimensional Approaches
Yuandong Xu, Jinshi Cui, Huijing Zhao, Hongbin Zha |
BMVC | 2 |
| 2011 | A novel laser-based system: Fully online detection of abnormal activity via an unsupervised methodabstractAbnormal activity detection plays a crucial role in surveillance applications, and such system has become an urgent need for public security. In this paper, we propose a novel laser-based system, which can perform the online detection of abnormal activity with an unsupervised way. The proposed system has the following key features that make it advantageous over previous ones: (1) It can cover quite a large and crowded area, such as subway station, public square, intersection and etc. (2) The overall system can vary with time period, incrementally learn the behavior pattern of pedestrians and perform the fully online detection of abnormal activity. This feature makes our system be quite suitable for the real-time applications. (3) The abnormal activity detection is carried out with a fully unsupervised way, there is no need for manual labelling and constructing the huge training datasets. We successfully applied the proposed system into the JR subway station of Tokyo, which can cover a 60×35m area, track more 150 targets at the same time and simultaneously perform the robust detection of abnormal activity with no human intervention. Xuan Song 0001, Xiaowei Shao, Ryosuke Shibasaki, Huijing Zhao, Jinshi Cui, Hongbin Zha |
ICRA | 5 |
| 2010 | An online approach: Learning-Semantic-Scene-by-Tracking and Tracking-by-Learning-Semantic-SceneabstractLearning the knowledge of scene structure and tracking a large number of targets are both active topics of computer vision in recent years, which plays a crucial role in surveillance, activity analysis, object classification and etc. In this paper, we propose a novel system which simultaneously performs the Learning-Semantic-Scene and Tracking, and makes them supplement each other in one framework. The trajectories obtained by the tracking are utilized to continually learn and update the scene knowledge via an online un-supervised learning. On the other hand, the learned knowledge of scene in turn is utilized to supervise and improve the tracking results. Therefore, this “adaptive learning-tracking loop” can not only perform the robust tracking in high density crowd scene, dynamically update the knowledge of scene structure and output semantic words, but also ensures that the entire process is completely automatic and online. We successfully applied the proposed system into the JR subway station of Tokyo, which can dynamically obtain the semantic scene structure and robustly track more than 150 targets at the same time. Xuan Song 0001, Xiaowei Shao, Huijing Zhao, Jinshi Cui, Ryosuke Shibasaki, Hongbin Zha |
CVPR | 4 |
| 2010 | Fusion of laser and vision for multiple targets tracking via on-line learningabstractMulti-target tracking becomes significantly more challenging when the targets are in close proximity or frequently interact with each other. This paper presents a promising tracking system to deal with these problems. The novelty of this system is that laser and vision, tracking and learning are integrated and can complement each other in one framework: when the targets do not interact with each other, the laser-based independent trackers are employed and the visual information is extracted simultaneously to train some classifiers for the “possible interacting targets”. When the targets are in close proximity, the learned classifiers and visual information are used to assist in tracking. Therefore, this mode of co-operation between them not only deals with various tough problems encountered in the tracking, but also ensures that the entire process can be completely on-line and automatic. Experimental results demonstrated that laser and vision fully display their respective advantages in our system, and it is easy for us to obtain a perfect trade-off between tracking accuracy and time-cost. Xuan Song 0001, Huijing Zhao, Jinshi Cui, Xiaowei Shao, Ryosuke Shibasaki, Hongbin Zha |
ICRA | 3 |
| 2009 | Moving object classification using horizontal laser scan dataabstractMotivated by two potential applications, i.e. enhancing driving safety and traffic data collection, a system has been developed using a single-layer horizontal laser scanner as the major sensor for both localization and perception of the surroundings in a large dynamic urban environment. This research focuses on a classification method, that given a stream of laser measurements, classify the moving object into either a person, a group of people, a bicycle or a car. In this research, a number of features are defined after examining the property of data appearance. A classification method is proposed after examining the likelihood measures between each pair of feature and class. Experimental results are presented, demonstrating that the algorithm has efficiency with respect to both driving safety and traffic data collection in highly dynamic environment. Huijing Zhao, Quanshi Zhang, Masaki Chiba, Ryosuke Shibasaki, Jinshi Cui, Hongbin Zha |
ICRA | 5 |
| 2009 | Combining Laser-Scanning Data and Images for Target Tracking and Scene Modeling
Hongbin Zha, Huijing Zhao, Jinshi Cui, Xuan Song 0001, Xianghua Ying |
ISRR | 3 |
| 2009 | A Laser-Scanner-Based Approach Toward Driving Safety and Traffic Data CollectionabstractThis work is motivated by the following two potential applications: 1) enhancing driving safety and 2) collecting traffic data in a large dynamic urban environment. A laser-scanner-based approach is proposed. The problem is formulated as a simultaneous localization and mapping (SLAM) with object tracking and classification, where the focus is on managing a mixture of data from both dynamic and static objects in a highly dynamic environment. A trajectory-oriented closure is also proposed using the sporadically available global positioning system (GPS) measurements in urban areas to assist for global accuracy, particularly when the vehicle makes a noncyclical measurement in a large outdoor environment. Experiments are conducted using the data that were collected along a course near 4.5 km in a highly dynamic environment. Possibilities of the approaches toward the two potential applications are demonstrated, and avenues for future works are discussed. Huijing Zhao, Masaki Chiba, Ryosuke Shibasaki, Xiaowei Shao, Jinshi Cui, Hongbin Zha |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2008 | Probabilistic Detection-based Particle Filter for Multi-target TrackingabstractIn this paper, we present a Probabilistic Detection-based Particle Filter (PD-PF) for tracking a variable number of interacting targets. When the objects do not interact with each other, our method performs like the deterministic detection-base methods. When the objects are in close proximity, the interactions and occlusions are modelled by a mixed proposal constructed by probabilistic detections and information from dynamic models. Specially, prior of detection-reliability minimizes the influence of non-detection or false alarm in the tracking. Moreover, we run independent PD-PF for each target, such that particles are sampled in a small state space, thus our method not only obtains a better approximation of posterior than joint particle filter or independent particle filter when interactions occur, but also has an acceptable computational complexity. Different evaluations demonstrate the validity and efficiency of the proposed method. 1 Xuan Song 0001, Jinshi Cui, Hongbin Zha, Huijing Zhao |
BMVC | 2 |
| 2008 | Vision-Based Multiple Interacting Targets Tracking via On-Line Supervised Learning
Xuan Song 0001, Jinshi Cui, Hongbin Zha, Huijing Zhao |
ECCV (3) | 2 |
| 2008 | Tracking interacting targets with laser scanner via on-line supervised learningabstractSuccessful multi-target tracking requires locating the targets and labeling their identities. For the laser based tracking system, the latter becomes significantly more challenging when the targets frequently interact with each other. This paper presents a novel on-line supervised learning based method for tracking interacting targets with laser scanner. When the targets do not interact with each other, we collect samples and train a classifier for each target. When the targets are in close proximity, we use these classifiers to assist in tracking. Different evaluations demonstrate that this method has a better tracking performance than previous methods when interactions occur, and can maintain correct tracking under various complex tracking situations. Xuan Song 0001, Jinshi Cui, Xulei Wang, Huijing Zhao, Hongbin Zha |
ICRA | 2 |
| 2008 | SLAM in a dynamic large outdoor environment using a laser scannerabstractIn this research, we propose a method of SLAM in a dynamic large outdoor environment using a laser scanner. Focus are cast on solving two major problems: 1) achieving global accuracy especially in non-cyclical environment, 2) tackling a mixture of data from both dynamic and static objects. Algorithms are developed, where GPS data and control inputs are used to diagnose pose error and guide to achieve a global accuracy; Classification of laser points and objects are conducted not in an independent module but across the processing in a framework of SLAM with moving object detection and tracking. Experiments are conducted using the data from two test-bed vehicles, and performance of the algorithms are demonstrated. Huijing Zhao, Masaki Chiba, Ryosuke Shibasaki, Xiaowei Shao, Jinshi Cui, Hongbin Zha |
ICRA | 5 |
| 2008 | Multi-modal tracking of people using laser scanners and video camera
Jinshi Cui, Hongbin Zha, Huijing Zhao, Ryosuke Shibasaki |
Image Vis. Comput. | 1 |
| 2007 | Synchronized Ego-Motion Recovery of Two Face-to-Face Cameras
Jinshi Cui, Yasushi Yagi, Hongbin Zha, Yasuhiro Mukaigawa, Kazuaki Kondo |
ACCV (1) | 1 |
| 2007 | Laser-based detection and tracking of multiple people in crowds
Jinshi Cui, Hongbin Zha, Huijing Zhao, Ryosuke Shibasaki |
Comput. Vis. Image Underst. | 1 |
| 2006 | Fusion of Detection and Matching Based Approaches for Laser Based Multiple People TrackingabstractMost of visual tracking algorithms have been achieved by matching-based searching strategies or detection-based data association algorithms. In this paper, our objective is to analysis laser scan image sequences to track multiple people in a crowded environment. Due to the poor features provided by laser scan images, neither of the above two approaches can achieves good tracking. To address the problem, we propose a novel multiple-target tracking algorithm fusing both detection and matching based strategies. First, target to detected measurement data association is incorporated to the joint state proposal, to form a mixture proposal that combines information from the dynamic model and the detected measurements. And then, we utilize a MCMC sampling step to obtain a more efficient multi-target filter. Our approach has been applied to the real laser scan image data. Evaluations show that the proposed method is a robust and effective multi-target tracking algorithm. Jinshi Cui, Huijing Zhao, Ryosuke Shibasaki |
CVPR (1) | 1 |
| 2006 | Laser-based Interacting People Tracking Using Multi-level ObservationsabstractLaser based people tracking systems have been developed for mobile robotics and intelligent surveillance areas. Existing systems rely on simple laser point clustering methods to extract object locations. However, when dealing with multiple interacting people, laser points of different persons are often interlaced and undistinguishable due to measurement noise and they can not provide reliable features. It causes current systems quite fragile and unreliable. In this paper, we try to explore potentials from multi-level observations including weakly detected features, stably extracted features and foreground points. For inference, detection incorporated joint particle filter is used. And stably extracted features are utilized to properly estimate parameters of dynamic model for each target. In real experiments, we obtain raw data from multiple registered laser scanners, which measure two legs for each people. Evaluations with real data show that the proposed method is more robust and effective than existing approaches Jinshi Cui, Hongbin Zha, Huijing Zhao, Ryosuke Shibasaki |
IROS | 1 |
| 2005 | Tracking multiple people using laser and visionabstractWe present a novel system that aims at reliably detecting and tracking multiple people in an open area. Multiple single-row laser scanners and one video camera are utilized. Feet trajectory tracking based on registration of distance information from multiple laser scanners and visual body region tracking based on color histogram are combined in a Bayesian formulation. Results from tests in a real environment are reported to demonstrate that the system can detect and track multiple people simultaneously with reliable and real-time performance. Jinshi Cui, Hongbin Zha, Huijing Zhao, Ryosuke Shibasaki |
IROS | 1 |