Huijun Di

dblp:88/1144 · DBLP profile ↗
← Back
51ranked-venue papers
5as first author
14since 2021 · last 2026
0000-0002-6432-2127ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 32 · 3 first-author · 10 since 2021Artificial intelligence and machine learning · 23 · 3 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 2Systems, architecture and hardware · 1 · 1 since 2021Security and privacy · 1
YearPublicationVenuePosition
2026 RadarMP: Motion Perception for 4D mmWave Radar in Autonomous Driving
abstract
Accurate 3D scene motion perception significantly enhances the safety and reliability of an autonomous driving system. Benefiting from its all-weather operational capability and unique perceptual properties, 4D mmWave radar has emerged as an essential component in advanced autonomous driving. However, sparse and noisy radar points often lead to imprecise motion perception, leaving autonomous vehicles with limited sensing capabilities when optical sensors degrade under adverse weather conditions. In this paper, we propose RadarMP, a novel method for precise 3D scene motion perception using low-level radar echo signals from two consecutive frames. Unlike existing methods that separate radar target detection and motion estimation, RadarMP jointly models both tasks in a unified architecture, enabling consistent radar point cloud generation and pointwise 3D scene flow prediction. Tailored to radar characteristics, we design specialized self-supervised loss functions guided by Doppler shifts and echo intensity, effectively supervising spatial and motion consistency without explicit annotations. Extensive experiments on the public dataset demonstrate that RadarMP achieves reliable motion perception across diverse weather and illumination conditions, outperforming radar-based decoupled motion perception pipelines and enhancing perception capabilities for full-scenario autonomous driving systems.
Ruiqi Cheng, Huijun Di, Wei Liang 0008
AAAI2
2026 SFGFusion: Surface fitting guided 3D object detection with 4D radar and camera fusion
Xiaozhi Li, Huijun Di, Wei Liang 0008
Pattern Recognit.2
2025 FloNa: Floor Plan Guided Embodied Visual Navigation
abstract
Humans naturally rely on floor plans to navigate in unfamiliar environments, as they are readily available, reliable, and provide rich geometrical guidance. However, existing visual navigation settings overlook this valuable prior knowledge, leading to limited efficiency and accuracy. To eliminate this gap, we introduce a novel navigation task: Floor Plan Visual Navigation (FloNa), the first attempt to incorporate floor plans into embodied visual navigation. While the floor plan offers significant advantages, two key challenges emerge: (1) handling the spatial inconsistency between the floor plan and the actual scene layout for collision-free navigation, and (2) aligning observed images with the floor plan sketch despite their distinct modalities. To address these challenges, we propose FloDiff, a novel diffusion policy framework incorporating a localization module to facilitate alignment between the current observation and the floor plan. We further collect 20k navigation episodes across 117 scenes in the iGibson simulator to support the training and evaluation. Extensive experiments demonstrate the effectiveness and efficiency of our framework in unfamiliar scenes using floor plan knowledge.
Weiqi Huang, Wei Liang 0008, Huijun Di
AAAI5
2025 Emergency Evacuation Map Guided Navigation via Topological Alignment and VLM Reasoning
Canzhi Chen, Weiqi Huang, Huijun Di, Wei Liang 0008
PRCV (6)5
2024 Visual Loop Closure Detection with Thorough Temporal and Spatial Context Exploitation
abstract
Despite advancements in visual Simultaneous Localization and Mapping (SLAM), prevailing visual Loop Closure Detection (LCD) methods primarily rely on computationally intensive image similarity comparisons, neglecting temporal-spatial context during long-term exploration. To address this issue, we propose TOSA, a novel visual LCD algorithm harnessing TempOral and SpAtial context for efficient LCD. Specifically, as the agent explores through time, our approach recurrently updates a latent feature incorporating historical information via a Long Short-Term Memory (LSTM) module. Upon receiving a query frame, TOSA seamlessly fuses the latent feature with the query feature to predict the candidates’ distribution, thus averting intensive similarity computation. Additionally, TOSA integrates a temporal-spatial convolution for candidate refinement by thoroughly exploiting the temporal consistency and spatial correlation to enhance selected candidates, further boosting the performance. Extensive experiments across four standard datasets showcase the superiority of our method over existing state-of-the-art techniques, demonstrating the effectiveness of utilizing rich temporal-spatial contexts.
Huijun Di, Wei Liang 0008
IROS3
2024 Self-Supervised Interactive Image Segmentation
abstract
Although interactive image segmentation techniques have made significant progress, supervised learning-based methods rely heavily on large-scale labeled data which is difficult to obtain in certain domains such as medicine, biology, etc. Models trained on natural images also struggle to achieve satisfactory results when directly applied to these domains. To solve this dilemma, we propose a Self-supervised Interactive Segmentation (SIS) method that achieves superior generalization performance. By clustering features from unlabeled data, we obtain classifiers that assign pseudo-labels to pixels in images. After refinement by super-pixel voting, these pseudo-labels are then used to train our segmentation network. To enable our network to better adapt to cross-domain images, we introduce correction learning and anti-forgetting regularization to conduct test-time adaptation. Our experiment results on five datasets show that our approach significantly outperforms other interactive segmentation methods across natural image datasets in the same conditions and achieves even better performance than some supervised methods when across to medical image domain. The code and models are available at https://github.com/leal0110/SIS.
Qingxuan Shi, Huijun Di, Enyi Wu
IEEE Trans. Circuits Syst. Video Technol.3
2022 Detail-Preserving Transformer for Light Field Image Super-resolution
abstract
Recently, numerous algorithms have been developed to tackle the problem of light field super-resolution (LFSR), i.e., super-resolving low-resolution light fields to gain high-resolution views. Despite delivering encouraging results, these approaches are all convolution-based, and are naturally weak in global relation modeling of sub-aperture images necessarily to characterize the inherent structure of light fields. In this paper, we put forth a novel formulation built upon Transformers, by treating LFSR as a sequence-to-sequence reconstruction task. In particular, our model regards sub-aperture images of each vertical or horizontal angular view as a sequence, and establishes long-range geometric dependencies within each sequence via a spatial-angular locally-enhanced self-attention layer, which maintains the locality of each sub-aperture image as well. Additionally, to better recover image details, we propose a detail-preserving Transformer (termed as DPT), by leveraging gradient maps of light field to guide the sequence learning. DPT consists of two branches, with each associated with a Transformer for learning from an original or gradient image sequence. The two branches are finally fused to obtain comprehensive feature representations for reconstruction. Evaluations are conducted on a number of light field datasets, including real-world scenes and synthetic data. The proposed method achieves superior performance comparing with other state-of-the-art schemes. Our code is publicly available at: https://github.com/BITszwang/DPT.
Shunzhou Wang, Tianfei Zhou, Yao Lu 0001, Huijun Di
AAAI4
2022 Reconstructing 3D Contour Models of General Scenes from RGB-D Sequences
Huijun Di, Lingxiao Song
MMM (2)2
2022 Cascade Scale-Aware Distillation Network for Lightweight Remote Sensing Image Super-Resolution
Haowei Ji, Huijun Di, Shunzhou Wang, Qingxuan Shi
PRCV (4)2
2022 Distillation Remote Sensing Object Counting via Multi-Scale Context Feature Aggregation
abstract
Remote sensing object counting is an important issue in remote sensing analysis. Remote sensing object counting has many challenges, such as large-scale variations and complex backgrounds. The previous counting methods have many shortboards, such as only focusing on local appearance features of target scenes and ignoring the self-supervision ability of the network itself. To remedy the above problems, in this article, we propose a novel remote sensing object counting method, which contains the adaptive multi-scale context aggregation module (AMCAM) and the self-context distillation module (SCDM). The AMCAM can model and fuse context information from different receptive fields effectively. It also keeps detailed information through multiple pixel attention (PA) modules step by step. The SCDM can improve the representation learning without adding any additional supervision information. SCDM uses feature maps from the deeper layer of the network to supervise feature maps from the earlier layer of the network. Our method has achieved good performance on the remote sensing object counting dataset, RSOC, and mainstream crowd counting datasets, such as ShanghaiTech and UCF-QNRF datasets.
Zuodong Duan, Shunzhou Wang, Huijun Di, Jiahao Deng
IEEE Trans. Geosci. Remote. Sens.3
2022 Contextual Transformation Network for Lightweight Remote-Sensing Image Super-Resolution
abstract
Current super-resolution networks typically reduce network parameters and multiadds operations by designing lightweight structures, but lightening the convolution layer is often ignored. In this work, we observe that$3 \times 3$convolutions occupy a high percentage of network parameters in most lightweight super-resolution networks. This motivates us to consider lightening super-resolution networks by replacing$3 \times 3$convolutions with lightweight convolutions, while maintaining the performance. To achieve this, we propose a lightweight convolution layer named contextual transformation layer (CTL). It can yield efficient contextual features through a context feature extraction module and enrich extracted contextual features through a context feature transformation module. Based on CTLs, we build a lightweight super-resolution network called contextual transformation network (CTN) for remote-sensing image super-resolution. Specifically, we use two CTLs to construct a contextual transformation block (CTB) for hierarchical feature learning. Interleaved with a CTB, a context enhancement module (CEM) is employed to enhance the extracted feature representations. All extracted features are processed by a contextual feature aggregation module for final remote-sensing image super-resolution. Extensive experiments are performed on a remote-sensing image super-resolution benchmark named UC Merced. Our method achieves superior results to the other state-of-the-art methods. To demonstrate the generalization ability of our CTL, we extend our CTN to two relevant tasks: natural image super-resolution and natural image denoising. Experimental results on natural image super-resolution benchmarks (i.e., Set5, Set14, B100, Urban100, and Manga109) and natural image denoising benchmarks (i.e., SIDD and DND) further prove the superiority of our method. Our code is publicly available athttps://github.com/BITszwang/CTNet.
Shunzhou Wang, Tianfei Zhou, Yao Lu 0001, Huijun Di
IEEE Trans. Geosci. Remote. Sens.4
2021 A Decomposition Model for Stereo Matching
abstract
In this paper, we present a decomposition model for stereo matching to solve the problem of excessive growth in computational cost (time and memory cost) as the resolution increases. In order to reduce the huge cost of stereo matching at the original resolution, our model only runs dense matching at a very low resolution and uses sparse matching at different higher resolutions to recover the disparity of lost details scale-by-scale. After the decomposition of stereo matching, our model iteratively fuses the sparse and dense disparity maps from adjacent scales with an occlusion-aware mask. A refinement network is also applied to improving the fusion result. Compared with high-performance methods like PSMNet and GANet, our method achieves 10−100× speed increase while obtaining comparable disparity estimation results.
Chengtang Yao, Yunde Jia, Huijun Di, Pengxiang Li 0002, Yuwei Wu 0001
CVPR3
2021 Multi-homography Estimation and Inference Driven by Contour Alignment
Yunde Jia, Huijun Di, Yuwei Wu 0001
ICIG (1)3
2021 SA-InterNet: Scale-Aware Interaction Network for Joint Crowd Counting and Localization
Xiuqi Chen, Huijun Di, Shunzhou Wang
PRCV (1)3
2020 GCVNet: Geometry Constrained Voting Network to Estimate 3D Pose for Fine-Grained Object Categories
Yaohang Han, Huijun Di, Hanfeng Zheng, Jianyong Qi, Jianwei Gong
PRCV (1)2
2020 TriSpaSurf: A Triple-View Outline Guided 3D Surface Reconstruction of Vehicles from Sparse Point Cloud
Hanfeng Zheng, Huijun Di, Yaohang Han, Jianwei Gong
PRCV (1)2
2020 SCLNet: Spatial context learning network for congested crowd counting
Shunzhou Wang, Yao Lu 0001, Tianfei Zhou, Huijun Di, Lin Zhang 0033
Neurocomputing4
2020 Face Spoofing Detection Using Relativity Representation on Riemannian Manifold
abstract
Face recognition and verification systems are susceptible to spoofing attacks using photographs, videos or masks. Most existing methods focus on spoofing detection in Euclidean space, and ignore the features' manifold structure and interrelationships, thus limiting their capabilities of discrimination and generalization. In this paper, we propose a relativity representation on Riemannian manifold for face spoofing detection. The relativity representation improves generalization capability while ensuring discriminability, at both levels of feature description and classification score. The feature-level relativity representation generalizes information by modeling interrelationships among basic features, and would not depend too much on characteristics of a particular dataset. The score-level relativity representation makes decisions relatively, not absolutely, according to interrelationships (via Riemannian metric) and competitions (via example reweighting) among data samples on Riemannian manifold. The discriminability is ensured by the high-order nature of the feature-level relativity representation as well as Riemannian reweighted discriminative learning of the score-level relativity representation. Moreover, we integrate an attack-sensitive SVM classifier in Euclidean space to improve spoofing detection. Experiments demonstrate the effectiveness of our method on both intra-dataset and cross-dataset testing.
Chengtang Yao, Yunde Jia, Huijun Di, Yuwei Wu 0001
IEEE Trans. Inf. Forensics Secur.3
2020 GAIM: Graph Attention Interaction Model for Collective Activity Recognition
abstract
Unbalanced interaction relationships at personal and group levels play a pivotal role in collective activity recognition, which has not been adaptively and jointly explored by previous approaches. In this paper, we propose a graph attention interaction model (GAIM) embedded with the graph attention block (GAB) to explicitly and adaptively infer unbalanced interaction relations at personal and group levels in a unified architecture, and further to learn the spatial and temporal evolutions of the collective activity from these interactions to predict the activity labels. We first design the spatiotemporal graphs tailored to the collective activity where the concurrent person and group nodes, respectively, represent individuals' actions and the collective activity. The graphs provide both spatial structures and semantic appearance features for the collective activity. Then, GAB performs convolution-like filters on the graphs to infer unequal and two-level interaction relations in the collective activity by implementing graph convolutional networks with a shared attention mechanism. At the personal level, the GAB learns different levels of interactions for each person node from its neighbor person nodes under the guidance from the group node. At the group level, the GAB assesses various degrees of interactions to the group node contributed by person nodes. Equipped with the GRUs network, the GAIM learns the spatial and temporal evolutions of individuals' actions as well as the collective activity from the captured interactions, and finally predicts the label of the collective activity. Experiments on four publicly available datasets and ablation studies are conducted to evaluate the performance of our GAIM, and the improved performance demonstrates the effectiveness of our model.
Yao Lu 0001, Ruizhe Yu, Huijun Di, Lin Zhang 0033, Shunzhou Wang
IEEE Trans. Multim.4
2019 SAF: Semantic Attention Fusion Mechanism for Pedestrian Detection
Ruizhe Yu, Shunzhou Wang, Yao Lu 0001, Huijun Di, Lin Zhang 0033
PRICAI (2)4
2019 Spatio-temporal attention mechanisms based model for collective activity recognition
Huijun Di, Yao Lu 0001, Lin Zhang 0033, Shunzhou Wang
Signal Process. Image Commun.2
2018 CPFG-SLAM: a Robust Simultaneous Localization and Mapping based on LIDAR in Off-Road Environment
abstract
Simultaneous localization and mapping (SLAM), as an important tool for vehicle positioning and mapping, plays an important role in the unmanned vehicle technology. This paper mainly presents a new solution to the LIDAR-based SLAM for unmanned vehicles in the off-road environment. Many methods have been proposed to solve the SLAM problems well. However, in complex environment, especially off-road environment, it is difficult to obtain stable positioning results due to the rough road and scene diversity. We propose a SLAM algorithm based on grid which combining probability and feature by Expectation-maximization (EM). The algorithm is mainly divided into three steps: data preprocessing, pose estimation, updating feature grid map. Our algorithm has strong robustness and real-time performance. We have tested our algorithm with our datasets of the multiple off-road scenes which obtained by LIDAR. Our algorithm performs pose estimation and feature map updating in parallel, which guarantees the real-time performance of the algorithm. The average processing time of each frame is about 55ms, and the average relative translation error is around 0.94%. Compared with several state-of-the-art algorithms, our algorithm has better performance in robustness and location accuracy.
Kaijin Ji, Huiyan Chen, Huijun Di, Jianwei Gong, Guangming Xiong, Jianyong Qi
Intelligent Vehicles Symposium3
2018 Real-Time 6D Lidar SLAM in Large Scale Natural Terrains for UGV
abstract
Simultaneous Localization And Mapping (SLAM) plays a more and more important role in the environment perception system of Unmanned Ground Vehicle (UGV), most SLAM technologies used to be applied indoor or in urban scenarios, we present a real-time 6D SLAM approach suitable for large scale natural terrain with the help of an Inertial Measurement Unit(IMU) and two 3D Lidars. Besides dividing the entire map into many submaps which consists of large numbers of tree structure based voxels, we use probabilistic methods to represent the possibility of one voxel being occupied/null. A Sparse Pose Adjustment (SPA) method has been used to solve 6D global pose optimization with some relative poses as pose constraints and relative motions computed from IMU data as kinetics constraints. A place recognition method integrated a method named Rotation Histogram Matching (RHM) and a Branch and Bound Search (BBS) based Iterative Closest Points (ICP) algorithm is applied to realize a real-time loop closure detection. We complete global pose optimization with the help of Ceres. Experimental results obtained from a real large scale natural environment shows an effective reduction for Lidar odometry pose accumulative error and a good performance for 3D mapping.
Zhongze Liu, Huiyan Chen, Huijun Di, Jianwei Gong, Guangming Xiong, Jianyong Qi
Intelligent Vehicles Symposium3
2018 A two-level attention-based interaction model for multi-person activity recognition
Huijun Di, Yao Lu 0001, Lin Zhang 0033, Shunzhou Wang
Neurocomputing2
2017 Improving Deep Crowd Density Estimation via Pre-classification of Density
Shunzhou Wang, Huailin Zhao, Weiren Wang, Huijun Di
ICONIP (3)4
2017 Video pose estimation with global motion cues
Qingxuan Shi, Huijun Di, Yao Lu 0001, Feng Lv, Xuedong Tian
Neurocomputing2
2017 Region-based Mixture Models for human action recognition in low-resolution videos
Huijun Di, Jian Zhang 0002, Yao Lu 0001, Feng Lv
Neurocomputing2
2017 Locality-Constrained Collaborative Model for Robust Visual Tracking
abstract
This paper presents a novel discriminative, generative, and collaborative appearance model for robust object tracking. In contrast to existing methods, we use different appearance manifolds to represent the target in the discriminative and generative appearance models and propose a novel collaborative scheme to combine these two components. In particular: 1) for the discriminative component, we develop a graph regularized discriminant analysis (GRDA) algorithm that can find a projection to more effectively distinguish the target from the background; 2) for the generative component, we introduce a simple yet effective coding method for object representation. The method involves no optimization, and thus better efficiency can be achieved; and 3) for the collaborative model, we apply GRDA again to find a subspace for discriminating the likelihood features (generated from the discriminative and generative appearance models) and use the nearest neighbor criterion to determine the final likelihood. Besides, all the components are online updated so that our tracker can deal with appearance changes effectively. The experimental results over 23 challenging image sequences demonstrate that the proposed algorithm achieves better performance compared with other state-of-the-art methods.
Tianfei Zhou, Yao Lu 0001, Huijun Di
IEEE Trans. Circuits Syst. Video Technol.3
2016 Video pose estimation via medium granularity graphical model with spatial-temporal symmetric constraint part model
abstract
We address the problem of full body human pose estimation in video. Most previous work consider body part, pose or trajectory of body part as basic unit to compose the pose sequence. In contrast, we consider tracklet of body part as the basic unit. Based on this medium granularity representation we develop a spatio-temporal graphical model to select an optimal tracklet for each part in each video segment. In our model, tracklet nodes of symmetric parts are coupled to one node to overcome the double counting problem. Through iterative spatial and temporal parsing, optimal solution is achieved in polynomial time. We apply our model on three publicly available datasets and show remarkable quantitative and qualitative improvements over the state-of-the-art approaches.
Qingxuan Shi, Huijun Di, Yao Lu 0001, Ming Qin, Xuedong Tian
ICIP2
2016 Recognizing human actions from low-resolution videos by region-based mixture models
abstract
Recognizing human action from low-resolution (LR) videos is essential for many applications including large-scale video surveillance, sports video analysis and intelligent aerial vehicles. Currently, state-of-the-art performance in action recognition is achieved by the use of dense trajectories which are extracted by optical flow algorithms. However, the optical flow algorithms are far from perfect in LR videos. In addition, the spatial and temporal layout of features is a powerful cue for action discrimination. While, most existing methods encode the layout by previously segmenting body parts which is not feasible in LR videos. Addressing the problems, we adopt the Layered Elastic Motion Tracking (LEMT) method to extract a set of long-term motion trajectories and a long-term common shape from each video sequence, where the extracted trajectories are much denser than those of sparse interest points(SIPs); then we present a hybrid feature representation to integrate both of the shape and motion features; and finally we propose a Region-based Mixture Model (RMM) to be utilized for action classification. The RMM models the spatial layout of features without any needs of body parts segmentation. Experiments are conducted on two publicly available LR human action datasets. Among which, the UT-Tower dataset is very challenging because the average height of human figures is only about 20 pixels. The proposed approach attains near-perfect accuracy on both of the datasets.
Huijun Di, Jian Zhang 0002, Yao Lu 0001, Feng Lv
ICME2
2016 Video object segmentation aggregation
abstract
We present an approach for unsupervised object segmentation in unconstrained videos. Driven by the latest progress in this field, we argue that segmentation performance can be largely improved by aggregating the results generated by state-of-the-art algorithms. Initially, objects in individual frames are estimated through a per-frame aggregation procedure using majority voting. While this can predict relatively accurate object location, the initial estimation fails to cover the parts that are wrongly labeled by more than half of the algorithms. To address this, we build a holistic appearance model using non-local appearance cues by linear regression. Then, we integrate the appearance priors and spatio-temporal information into an energy minimization framework to refine the initial estimation. We evaluate our method on challenging benchmark videos and demonstrate that it outperforms state-of-the-art algorithms.
Tianfei Zhou, Yao Lu 0001, Huijun Di, Jian Zhang 0002
ICME3
2016 A Background Basis Selection-Based Foreground Detection Method
abstract
Foreground detection plays a fundamental role in video analysis. Frames with only background information are usually beneficial for many foreground detection algorithms, especially for regression-based methods where the background is recovered from a background basis matrix. However, many regression-based methods ignore the basis selection process or select bases by simple sampling, which may limit their performance. In this paper, a regression-based foreground detection method with a novel background basis selection process is proposed. The proposed basis selection method, which includes basis matrix construction and basis matrix update processes, aims to build an effective background basis matrix which helps to boost the performance of our foreground detection method. In our algorithm, the basis matrix construction process first builds the basis matrix locally with a multiple clustering evaluation process. With the locally constructed basis matrix, a modified linear regression-based foreground detection method is proposed for separating foreground and background globally. To further increase the representativeness and the adaptiveness of the background basis matrix, a basis matrix update algorithm is designed to incrementally replace the ineffective bases with new selected ones. Extensive experiments on challenging sequences demonstrate the effectiveness and the advantages of our method.
Ming Qin, Yao Lu 0001, Huijun Di, Wei Huang 0030
IEEE Trans. Multim.3
2015 Contour Flow: Middle-Level Motion Estimation by Combining Motion Segmentation and Contour Alignment
abstract
Our goal is to estimate contour flow (the contour pairs with consistent point correspondence) from inconsistent contours extracted independently in two video frames. We formulate the contour flow estimation locally as a motion segmentation problem where motion patterns grouped from optical flow field are exploited for local correspondence measurement. To solve local ambiguities, contour flow estimation is further formulated globally as a contour alignment problem. We propose a novel two-staged strategy to obtain global consistent point correspondence under various contour transitions such as splitting, merging and branching. The goal of the first stage is to obtain possible accurate contour-to-contour alignments, and the second stage aims to make a consistent fusion of many partial alignments. Such a strategy can properly balance the accuracy and the consistency, which enables a middle-level motion representation to be constructed by just concatenating frame-by-frame contour flow estimation. Experiments prove the effectiveness of our method.
Huijun Di, Qingxuan Shi, Feng Lv, Ming Qin, Yao Lu 0001
ICCV1
2015 Adaptive Piecewise Elastic Motion Estimation
Huijun Di, Linmi Tao, Guangyou Xu
ICIC (2)1
2015 Robust Contour Tracking via Constrained Separate Tracking of Location and Shape
Huijun Di, Linmi Tao, Guangyou Xu
ICIG (3)1
2015 Human pose estimation with global motion cues
abstract
We present a novel method to estimate full-body human pose in video sequence by incorporating global motion cues. It has been demonstrated that temporal constraints can largely enhance the pose estimation. Most current approaches typically employ local motion to propagate pose detections to supplement the pose candidates. However, the local motion estimation is often inaccurate under fast movements of body parts and unhelpful when no strong detections achieved in adjacent frames. In this paper, we propose to propagate the detection in each frame using the global motion estimation. Benefiting from the strong detections, our algorithm first produces reasonable trajectory hypotheses for each body part. Then, we cast pose estimation as an optimization problem defined on these trajectories with spatial links between body parts. In the optimization process, we select body part trajectory rather than body part candidate to infer the human pose. Experimental results demonstrate significant performance improvement in comparison with the state-of-the-art methods.
Qingxuan Shi, Huijun Di, Yao Lu 0001, Feng Lv
ICIP2
2015 Background basis selection from multiple clustering on local neighborhood structure
abstract
Foreground detection with dynamic background is a challenging task in video surveillance analysis. When clean background bases are constructed, regression based foreground detection usually becomes more effective. In this paper, a novel basis selection method based on local neighborhood structure is proposed. The present method first constructs local neighborhood relationships among the basis candidates in a reconstruction manner. Then a multiple clustering strategy is designed to evaluate these basis candidates on local neighborhood structure. According to the evaluation score given by multiple clustering process, clean background bases (including dynamic background) are separated from candidates corrupted by foreground. By adding the proposed basis selection process to a modified linear regression framework, the foreground detection can be implemented in a more effective way. Experimental results on multiple videos show that the modified framework with basis selection is competitive with the state of the art.
Ming Qin, Yao Lu 0001, Huijun Di, Wei Huang 0030
ICME3
2015 Abrupt motion tracking via nearest neighbor field driven stochastic sampling
Tianfei Zhou, Yao Lu 0001, Feng Lv, Huijun Di, Qingjie Zhao, Jian Zhang 0002
Neurocomputing4
2014 Nearest neighbor field driven stochastic sampling for abrupt motion tracking
abstract
Stochastic sampling based trackers have shown good performance for abrupt motion tracking so that they have gained popularity in recent years. However, the existing methods tend to explore the whole state space uniformly with an inefficiency preliminary sampling phase. In this paper, we propose a nearest neighbor field(NNF) driven stochastic sampling framework for abrupt motion tracking in which NNF provides us promising regions the target may exist, and thus can help to explore the state space more effectively. Our approach firstly computes NNF to determine the promising regions; subsequently, we adopt Smoothing Stochastic Approximate Monte Carlo(SSAMC) sampling scheme to accurately localize the target. SSAMC is robust to handle the noises in NNF by propagating a sample's information to its neighboring regions. Finally, we refine the result with sparse representation based template matching technique. The experimental results on challenging sequences show that our tracker outperforms other related methods by better accuracy and higher robustness.
Tianfei Zhou, Yao Lu 0001, Huijun Di
ICME3
2014 Depth Super-resolution by Fusing Depth Imaging and Stereo Vision with Structural Determinant Information Inference
abstract
In this paper, we present a depth super-resolution framework by fusing depth imaging and stereo vision for high-resolution and high-accuracy depth maps. Depth cameras and stereo vision have their own limitations in some aspects, but their characteristics of range sensing are complementary. Thus, combining both approaches can produce more satisfactory results than either one. Unlike previous fusion methods, we initially taking the noisy depth observation from depth camera as prior information of scene structure. The prior information of scene structure is also utilized to infer structural determinant information, like depth discontinuity and occlusion, which is essential to improve the quality of depth map in the fusion process. In succession, the prior knowledge helps to overcome difficulties of intensity inconsistency in image observation from stereo vision component. Experimental results demonstrate effectiveness and accuracy of the proposed method.
Yucheng Wang 0003, Huijun Di, Wei Liang 0008, Jian Zhang 0002, Yunde Jia
ICPR2
2012 A Scalable Distributed Architecture for Intelligent Vision System
abstract
The complexity of intelligent computer vision systems demands novel system architectures that are capable of integrating various computer vision algorithms into a working system with high scalability. The real-time applications of human-centered computing are based on multiple cameras in current systems, which require a transparent distributed architecture. This paper presents an application-oriented service share model for the generalization of vision processing. Based on the model, a vision system architecture is presented that can readily integrate computer vision processing and make application modules share services and exchange messages transparently. The architecture provides a standard interface for loading various modules and a mechanism for modules to acquire inputs and publish processing results that can be used as inputs by others. Using this architecture, a system can load specific applications without considering the common low-layer data processing. We have implemented a prototype vision system based on the proposed architecture. The latency performance and 3-D track function were tested with the prototype system. The architecture is scalable and open, so it will be useful for supporting the development of an intelligent vision system, as well as a distributed sensor system.
Guojian Wang, Linmi Tao, Huijun Di, Xiyong Ye, Yuanchun Shi
IEEE Trans. Ind. Informatics3
2010 A Globally Optimal Approach for 3D Elastic Motion Estimation from Stereo Sequences
Qifan Wang 0001, Linmi Tao, Huijun Di
ECCV (4)3
2010 A Robust Approach for Person Localization in Multi-camera Environment
abstract
Person localization is fundamental in human centered computing, since person should be localized before being actively serviced. This paper proposed a robust approach to localize person based on the geometric constraints in multi-camera environment. The proposed algorithm has several advantages: 1) no assumption on the positions and orientations of cameras except the cameras should have certain common field of view; 2) no assumption on the visibility of particular body part (e.g., feet), except a portion of person should be observed in at least two views; 3) reliability in terms of tolerating occlusion, body posture change and inaccurate motion detection. It can also provide error control and be further extended to measure person height. The efficacy of the approach is demonstrated on challenging real-world scenarios.
Luo Sun, Huijun Di, Linmi Tao, Guangyou Xu
ICPR2
2009 Visual Focus of Attention Recognition in the Ambient Kitchen
Ligeng Dong, Huijun Di, Linmi Tao, Guangyou Xu, Patrick Olivier
ACCV (3)2
2009 A Mixture of Transformed Hidden Markov Models for Elastic Motion Estimation
abstract
Elastic motion is a nonrigid motion constrained only by some degree of smoothness and continuity. Consequently, elastic motion estimation by explicit feature matching actually contains two correlated subproblems: shape registration and motion tracking, which account for spatial smoothness and temporal continuity, respectively. If we ignore their interrelationship, solving each of them alone will be rather challenging, especially when the cluttered features are involved. To integrate them into a probabilistic model, one straightforward approach is to draw the dependence between their hidden states. With regard to their separated states, there are, however, two different explanations of motion which are still made under the individual constraint of smoothness or continuity. Each one can be error-prone, and their coupling causes error propagation. Therefore, it is highly desirable to design a probabilistic model in which a unified state is shared by the two subproblems. This paper is intended to propose such a model, i.e., a Mixture of Transformed Hidden Markov Models (MTHMM), where a unique explanation of motion is made simultaneously under the spatiotemporal constraints. As a result, the MTHMM could find a coherent global interpretation of elastic motion from local cluttered edge features, and experiments show its robustness under ambiguities, data missing, and outliers.
Huijun Di, Linmi Tao, Guangyou Xu
IEEE Trans. Pattern Anal. Mach. Intell.1
2009 Group Interaction Analysis in Dynamic Context
abstract
Computer understanding of human actions and interactions is one of the key research issues in human computing. In this regard, context plays an essential role in semantic understanding of human behavioral and social signals from sensor data. This paper put forward an event-based dynamic context model to address the problems of context awareness in the analysis of group interaction scenarios. Event-driven multilevel dynamic Bayesian network is correspondingly proposed to detect multilevel events, which underlies the context awareness mechanism. Online analysis can be achieved, which is superior over previous works. Experiments in our smart meeting room demonstrate the effectiveness of our approach.
Huijun Di, Ligeng Dong, Linmi Tao, Guangyou Xu
IEEE Trans. Syst. Man Cybern. Part B2
2008 Background modeling from a free-moving camera by Multi-Layer Homography Algorithm
abstract
This paper proposes a novel Multi-layer Homography Algorithm for background modeling from a free-moving camera. Background is composed of many planes. Different planes satisfy with different homographies which can be found by our algorithm. Each pixel except for the moving pixel definitely belongs to some plane. Rectified by the corresponding homography, each static pixel in the shared view can find its match in the previous frame. Thus, frames can be rectified to a specific viewpoint for background modeling. Experiment shows it is effective. Our approach can be used in motion detection from a free-moving camera.
Linmi Tao, Huijun Di, Naveed Iqbal Rao, Guangyou Xu
ICIP3
2008 Group Interaction Analysis in Dynamic Context
abstract
Computer understanding of human actions and interactions is one of the key research issues in human computing. In this regard, context plays an essential role in semantic understanding of human behavioral and social signals from sensor data. This paper put forward an event-based dynamic context model to address the problems of context awareness in the analysis of group interaction scenarios. Event-driven multilevel dynamic Bayesian network is correspondingly proposed to detect multilevel events, which underlies the context awareness mechanism. Online analysis can be achieved, which is superior over previous works. Experiments in our smart meeting room demonstrate the effectiveness of our approach.
Huijun Di, Ligeng Dong, Linmi Tao, Guangyou Xu
IEEE Trans. Syst. Man Cybern. Part B2
2007 Groupwise Shape Registration on Raw Edge Sequence via A Spatio-Temporal Generative Model
abstract
Groupwise shape registration of raw edge sequence is addressed. Automatically extracted edge maps are treated as noised input shape of the deformable object and their registration are considered, results can be used to build statistical shape models without laborious manual labeling process. Dealing with raw edges poses several challenges, to fight against them a novel spatio-temporal generative model is proposed which joints shape registration and trajectory tracking. Mean shape, consistent correspondences among edge sequence and associated non-rigid transformations are jointly inferred under EM framework. Our algorithm is tested on real video sequences of a dancing ballerina, talking face, and walking person. Results achieved are interesting, promising, and prove the robustness of our method. Potential applications can be found in statistical shape analysis, action recognition, object tracking, etc.
Huijun Di, Naveed Iqbal Rao, Guangyou Xu, Linmi Tao
CVPR1
2006 Estimating Illumination Parameters in Real Space with Application to Image Relighting
Linmi Tao, Guangyou Xu, Huijun Di
ACCV (1)4
2006 Refine Stereo Correspondence Using Bayesian Network and Dynamic Programming on a Color Based Minimal Span Tree
Naveed Iqbal Rao, Huijun Di, Guangyou Xu
ACIVS2