Hongbin Zha

dblp:20/5020 · DBLP profile ↗
← Back
290ranked-venue papers
18as first author
45since 2021 · last 2026
0000-0001-5860-4673ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 185 · 11 first-author · 26 since 2021Graphics, computer vision, multimedia, augmented reality and games · 156 · 7 first-author · 22 since 2021Systems, architecture and hardware · 44 · 3 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 35 · 5 first-author · 11 since 2021Human-computer interaction and ubiquitous computing · 20 · 6 first-authorDatabases, data management, data science and information retrieval · 6
YearPublicationVenuePosition
2026 RealNet: Efficient and Unsupervised Detection of AI-Generated Images via Real-Only Representation Learning
abstract
Detecting AI-generated images remains a persistent challenge, as existing detectors often struggle to generalize to forgeries produced by previously unseen generative models. This generalization gap mainly stems from entanglement with semantic content and overfitting to model-specific artifacts. Moreover, many state-of-the-art methods rely on large pre-trained backbones or computationally intensive pipelines, which limit their applicability in real-world, resource-constrained environments. We propose RealNet, a lightweight and unsupervised framework that constructs a disentangled, forgery-aware representation space using only real images. RealNet first extracts semantic-agnostic representations through a dual adversarial denoising mechanism, producing compact features with low intra-class variance. These representations are then perturbed in feature space to generate pseudo-negative samples, which are combined with the original real features to train a lightweight discriminator, enabling robust detection without any dependence on synthetic images during training. Comprehensive evaluations across GAN, diffusion, and emerging VAR-based paradigms demonstrate that RealNet achieves superior cross-model generalization and robustness. RealNet surpasses previous state-of-the-art approaches by 4.51% in accuracy and 3.93% in average precision, while maintaining significantly lower computational cost. Furthermore, we introduce a medically relevant synthetic image dataset and show RealNet remains effective under severe distribution shifts, highlighting its potential for deployment in high-stakes real-world scenarios. Together, these advantages position RealNet as a practical, scalable and socially impactful solution for robust AI-generated image detection.
Shuaibo Li, Laixin Zhang, Wei Ma 0008, Jianwei Guo 0003, Shibiao Xu, Zhijie Qiu, Hongbin Zha
AAAI7
2026 Rotation and semantic co-aware Transformer for oriented object detection in remote sensing images
Shaojun Lv, Wei Ma 0008, Hongbin Zha
Eng. Appl. Artif. Intell.4
2026 Low-light light field image enhancement based on illumination-guided implicit gradient representation
Deyang Liu, Xiaofei Zhou 0003, Ping An 0001, Caifeng Shan, Hongbin Zha
Neurocomputing6
2026 Degradation-Aware Blind Light-Field Image Quality Assessment With Linear Attention
abstract
Blind Light-Field Image Quality Assessment (LFIQA) is challenging, as degradations are microlens-dependent and spatially non-uniform, while perceptual quality relies on both spatial fidelity and angular consistency. However, many existing methods either assume globally stationary distortions or adopt global pooling or self-attention, which can be biased by locally corrupted lenslets and become computationally prohibitive when modeling long-range spatial-angular dependencies. Therefore, in this letter, we present a degradation-aware framework that first predicts a microlens reliability map to make quality inference robust to spatially non-uniform, lenslet-varying corruption. It then extracts spatial and angular features, applies a shared Receptance Weighted Key Value (RWKV) module for linear-time long-range context, and fuses them to predict perceptual quality. Experiments show a higher correlation with subjective ratings with competitive efficiency.
Youzhi Zhang 0004, Jianyu Qian, Deyang Liu, Xiaofei Zhou 0003, Hongbin Zha, Caifeng Shan
IEEE Signal Process. Lett.5
2025 Proactive Scene Decomposition and Reconstruction
abstract
Human behaviors are the major causes of scene dynamics and inherently contain rich cues regarding the dynamics. This paper formalizes a new task of proactive scene decomposition and reconstruction, an online approach that leverages human-object interactions to iteratively disassemble and reconstruct the environment. By observing these intentional interactions, we can dynamically refine the decomposition and reconstruction process, addressing inherent ambiguities in static object-level reconstruction. The proposed system effectively integrates multiple tasks in dynamic environments such as accurate camera and object pose estimation, instance decomposition, and online map updating, capitalizing on cues from human-object interactions in egocentric live streams for a flexible, progressive alternative to conventional object-level reconstruction methods. Aided by the Gaussian splatting technique, accurate and consistent dynamic scene modeling is achieved with photorealistic and efficient rendering. The efficacy is validated in multiple real-world scenarios with promising advantages.
Baicheng Li, Zike Yan, Hongbin Zha
ICCV4
2025 Training-free Fourier Phase Diffusion for Style Transfer
abstract
Diffusion models have shown significant potential for image style transfer tasks. However, achieving effective stylization while preserving content in a training-free setting remains a challenging issue due to the tightly coupled representation space and inherent randomness of the models. In this paper, we propose a Fourier phase diffusion model that addresses this challenge. Given that the Fourier phase spectrum encodes an image's edge structures, we propose modulating the intermediate diffusion samples with the Fourier phase of a content image to conditionally guide the diffusion process. This ensures content retention while fully utilizing the diffusion model's style generation capabilities. To implement this, we introduce a content phase spectrum incorporation method that aligns with the characteristics of the diffusion process, preventing interference with generative stylization. To further enhance content preservation, we integrate homomorphic semantic features extracted from the content image at each diffusion stage. Extensive experimental results demonstrate that our method outperforms state-of-the-art models in both content preservation and stylization. Code is available at https://github.com/zhang2002forwin/Fourier-Phase-Diffusion-for-Style-Transfer.
Wei Ma 0008, Hongbin Zha
IJCAI5
2025 Structured 3D gaussian splatting for novel view synthesis based on single RGB-LiDAR View
Zhiqun Zhao, Wei Ma 0008, Hongbin Zha
Appl. Intell.5
2025 Multi-Task Gradual Inference with a Single Encoder-Decoder Network for Automatic Portrait Matting
abstract
This paper presents a multi-task gradual inference model, MTGINet, for automatic portrait matting. It handles the subtasks of automatic portrait matting, namely portrait-transition-background trimap segmentation and transition region matting, with a single encoder-decoder structure. First, we enrich the highest stage of features from the encoder with portrait shape context via a shape context aggregation (SCA) module for trimap segmentation. Then, we fuse the SCA-enhanced features with detailed clues from the encoder for transition-region-aware alpha matting. The gradual inference model naturally allows sufficient interaction between the subtasks via forward computation and backwards propagation during training, and therefore achieves high accuracy while maintaining low complexity. In addition, considering the discrepancies in feature requirements across subtasks, we adapt the features from the encoders before reusing them via a feature rectification module. In addition to the MTGINet model, we have constructed a new large-scale dataset, HPM-17K, for half-body portrait matting. It consists of 16,967 images with diverse backgrounds. Comparative experiments with existing deep models on the public P3M-10K dataset and our HPM-17K dataset demonstrate that the proposed model exhibits state-of-the-art performance.
Wenbing Yang, Wei Ma 0008, Qing Mi, Hongbin Zha
Comput. Vis. Media4
2025 Learning and aggregating principal semantics for semantic edge detection in images
Lijun Dong, Wei Ma 0008, Hongbin Zha
Expert Syst. Appl.4
2025 Toward Semantically-Consistent Deformable 2D-3D Registration for 3D Craniofacial Structure Estimation From a Single-View Lateral Cephalometric Radiograph
abstract
The deep neural networks combined with the statistical shape model have enabled efficient deformable 2D-3D registration and recovery of 3D anatomical structures from a single radiograph. However, the recovered volumetric image tends to lack the volumetric fidelity of fine-grained anatomical structures and explicit consideration of cross-dimensional semantic correspondence. In this paper, we introduce a simple but effective solution for semantically-consistent deformable 2D-3D registration and detailed volumetric image recovery by inferring a voxel-wise registration field between the cone-beam computed tomography and a single lateral cephalometric radiograph (LC). The key idea is to refine the initial statistical model-based registration field with craniofacial structural details and semantic consistency from the LC. Specifically, our framework employs a self-supervised scheme to learn a voxel-level refiner of registration fields to provide fine-grained craniofacial structural details and volumetric fidelity. We also present a weakly supervised semantic consistency measure for semantic correspondence, relieving the requirements of volumetric image collections and annotations. Experiments showcase that our method achieves deformable 2D-3D registration with performance gains over state-of-the-art registration and radiograph-based volumetric reconstruction methods. The source code is available at https://github.com/Jyk-122/SC-DREG.
Yikun Jiang, Yuru Pei, Tianmin Xu, Xiaoru Yuan, Hongbin Zha
IEEE Trans. Medical Imaging5
2025 EdgeMaskFormer: Adapting Mask Transformer for Semantic Edge Detection
abstract
Semantic Edge Segmentation (SED) is crucial for intelligent agents to understand and interact with their environments, as it enables them to locate and recognize semantic boundaries. The prevailing framework in the field of SED is multi-label learning, which identifies edges and their semantics by learning to assign multiple labels that indicate the categories of the objects forming the edges. However, this framework has demonstrated limited performance when dealing with complex scenarios. In this paper, we propose a mask classification framework specifically tailored for the SED task, termed EdgeMaskFormer. Within this framework, we develop a query-based edge semantic extractor to learn semantic embeddings for edge mask classification with assistance from regional semantic supervision. Additionally, we design a context-aware hierarchical edge extractor to serve as an edge mask head, which can capture multi-scale edges of different categories under the guidance from the semantic embeddings via dynamic convolution. Furthermore, we develop matching and supervision mechanisms specifically for edge mask classification in order to reduce edge noise and address the imbalance between edge and non-edge samples. Our extensive experiments on three public datasets demonstrate that the proposed approach achieves outstanding performance in semantic edge detection, particularly on those datasets with complex scenarios.
Lijun Dong, Wei Ma 0008, Hongbin Zha
IEEE Trans. Multim.3
2024 Depth Reconstruction with Neural Signed Distance Fields in Structured Light Systems
abstract
We introduce a novel depth estimation technique for multi-frame structured light setups using neural implicit representations of 3D space. Our approach employs a neural signed distance field (SDF), trained through self-supervised differentiable rendering. Unlike passive vision, where joint estimation of radiance and geometry fields is necessary, we capitalize on known radiance fields from projected patterns in structured light systems. This enables isolated optimization of the geometry field, ensuring convergence and network efficacy with fixed device positioning. To enhance geometric fidelity, we incorporate an additional color loss based on object surfaces during training. Real-world experiments demonstrate our method’s superiority in geometric performance for few-shot scenarios, while achieving comparable results with increased pattern availability.
Rukun Qiao, Hiroshi Kawasaki, Hongbin Zha
3DV3
2024 Adaptive VIO: Deep Visual-Inertial Odometry with Online Continual Learning
abstract
Visual-inertial odometry (VIO) has demonstrated re-markable success due to its low-cost and complementary sensors. However, existing VIO methods lack the general-ization ability to adjust to different environments and sen-sor attributes. In this paper, we propose Adaptive VIO, a new monocular visual-inertial odometry that combines online continual learning with traditional nonlinear opti-mization. Adaptive VIO comprises two networks to pre-dict visual correspondence and IMU bias. Unlike end-to-end approaches that use networks to fuse the features from two modalities (camera and IMU) and predict poses directly, we combine neural networks with visual-inertial bundle adjustment in our VIO system. The optimized esti-mates will be fed back to the visual and IMU bias networks, refining the networks in a self-supervised manner. Such a learning-optimization-combined framework and feedback mechanism enable the system to perform online contin-ual learning. Experiments demonstrate that our Adaptive VIO manifests adaptive capability on EuRoC and TUM-VI datasets. The overall performance exceeds the currently known learning-based VIO methods and is comparable to the state-of-the-art optimization-based methods.
Youqi Pan, Wugen Zhou, Yingdian Cao, Hongbin Zha
CVPR4
2024 PanoRecon: Real-Time Panoptic 3D Reconstruction from Monocular Video
abstract
We introduce the Panoptic 3D Reconstruction task, a unified and holistic scene understanding task for a monocular video. And we present PanoRecon - a novel framework to address this new task, which realizes an online geometry reconstruction alone with dense semantic and instance labeling. Specifically, PanoRecon incrementally performs panoptic 3D reconstruction for each video fragment consisting of multiple consecutive key frames, from a volumetric feature representation using feed-forward neural networks. We adopt a depth-guided back-projection strategy to sparse and purify the volumetric feature representation. We further introduce a voxel clustering module to get object instances in each local fragment, and then design a tracking and fusion algorithm for the integration of instances from different fragments to ensure temporal co-herence. Such design enables our PanoRecon to yield a coherent and accurate panoptic 3D reconstruction. Exper-iments on ScanNetV2 demonstrate a very competitive geometry reconstruction result compared with state-of-the-art reconstruction methods, as well as promising 3D panoptic segmentation result with only RGB input, while being real-time. Code is available at: https://github.com/Riser6/PanoRecon.
Zike Yan, Hongbin Zha
CVPR3
2024 Learn to Memorize and to Forget: A Continual Learning Perspective of Dynamic SLAM
Baicheng Li, Zike Yan, Hanqing Jiang, Hongbin Zha
ECCV (77)5
2024 Active Neural Mapping at Scale
abstract
We introduce a NeRF-based active mapping system that enables efficient and robust exploration of large-scale indoor environments. The key to our approach is the extraction of a generalized Voronoi graph (GVG) from the continually updated neural map, leading to the synergistic integration of scene geometry, appearance, topology, and uncertainty. Anchoring uncertain areas induced by the neural map to the vertices of GVG allows the exploration to undergo adaptive granularity along a safe path that traverses unknown areas efficiently. Harnessing a modern hybrid NeRF representation, the proposed system achieves competitive results in terms of reconstruction accuracy, coverage completeness, and exploration efficiency even when scaling up to large indoor environments. Extensive results at different scales validate the efficacy of the proposed system.
Zijia Kuang, Zike Yan, Hao Zhao 0001, Guyue Zhou, Hongbin Zha
IROS5
2024 Stochastic Anomaly Simulation for Maxilla Completion from Cone-Beam Computed Tomography
Yixiao Guo, Yuru Pei, Zhi-bo Zhou, Tianmin Xu, Hongbin Zha
MICCAI (8)6
2024 ECT: Fine-grained edge detection with learned cause tokens
Shaocong Xu, Xiaoxue Chen, Yuhang Zheng 0004, Guyue Zhou, Yurong Chen 0001, Hongbin Zha, Hao Zhao 0002
Image Vis. Comput.6
2024 A Low-Cost and Scalable Framework to Build Large-Scale Localization Benchmark for Augmented Reality
abstract
Nowadays the application of AR is expanding from small or medium environments to large-scale environments, where the visual-based localization in the large-scale environments becomes a critical demand. Current visual-based localization techniques face robustness challenges in complex large-scale environments, requiring tremendous number of data with groundtruth localization for algorithm benchmarking or model training. The previous groundtruth solutions can only be used outdoors, or require high equipment/labor costs, so they cannot be scalable to large environments for both indoors and outdoors, nor can they produce large amounts of data at a feasible cost. In this work, we propose LSFB, a novel low-cost and scalable framework to build localization benchmark in large-scale indoor and outdoor environments. The key is to reconstruct an accurate HD map of the environment. For each visual-inertial sequence captured in the environment, the groundtruth poses are obtained by joint optimization taking both the HD map and visual-inertial constraints. The experiments demonstrate the obtained groundtruth poses have cm-level accuracy. We use the proposed method to collect a localization dataset by mobile phones and AR glasses in various environments with various motions, and release the dataset as the first large-scale localization benchmark for AR.
Haomin Liu, Linsheng Zhao, Weijian Xie, Mingxuan Jiang, Hongbin Zha, Hujun Bao, Guofeng Zhang 0001
IEEE Trans. Circuits Syst. Video Technol.6
2024 MOUNT: Learning 6DoF Motion Prediction Based on Uncertainty Estimation for Delayed AR Rendering
abstract
The delay of rendering on AR devices requires prediction of head motion using sensor data acquired tens of even one hundred milliseconds ago to avoid misalignment between the virtual content and the physical world, where the misalignment will lead to a sense of time latency and dizziness for users. To solve the problem, we propose a method for the 6DoF motion prediction to compensate for the time latency. Compared with traditional hand-crafted methods, our method is based on deep learning, which has better motion prediction ability to deal with complex human motion. In particular, we propose a MOtion UNcerTainty encode decode network (MOUNT) that estimates the uncertainty of input data and predicts the uncertainty of output motion to improve the prediction accuracy and smoothness. Experiments on the EuRoC and our collected dataset demonstrate that our method significantly outperforms the traditional method and greatly improves AR visual effects.
Haoran Chen 0010, Lantian Wei, Haomin Liu, Boxin Shi, Guofeng Zhang 0001, Hongbin Zha
IEEE Trans. Vis. Comput. Graph.6
2023 Active Neural Mapping
abstract
We address the problem of active mapping with a continually-learned neural scene representation, namely Active Neural Mapping. The key lies in actively finding the target space to be explored with efficient agent movement, thus minimizing the map uncertainty on-the-fly within a previously unseen environment. In this paper, we examine the weight space of the continually-learned neural field, and show empirically that the neural variability, the prediction robustness against random weight perturbation, can be directly utilized to measure the instant uncertainty of the neural map. Together with the continuous geometric information inherited in the neural map, the agent can be guided to find a traversable path to gradually gain knowledge of the environment. We present for the first time an online active mapping system with a coordinate-based implicit neural representation. Experiments in the visually-realistic Gibson and Matterport3D environment demonstrate the efficacy of the proposed method.
Zike Yan, Haoxiang Yang, Hongbin Zha
ICCV3
2023 From Semi-supervised to Omni-supervised Room Layout Estimation Using Point Clouds
abstract
Room layout estimation is a long-existing robotic vision task that benefits both environment sensing and motion planning. However, layout estimation using point clouds (PCs) still suffers from data scarcity due to annotation difficulty. As such, we address the semi-supervised setting of this task based upon the idea of model exponential moving averaging. But adapting this scheme to the state-of-the-art (SOTA) solution for PC-based layout estimation is not straightforward. To this end, we define a quad set matching strategy and several consistency losses based upon metrics tailored for layout quads. Besides, we propose a new online pseudo-label harvesting algorithm that decomposes the distribution of a hybrid distance measure between quads and PC into two components. This technique does not need manual threshold selection and intuitively encourages quads to align with reliable layout points. Surprisingly, this framework also works for the fully-supervised setting, achieving a new SOTA on the ScanNet benchmark. Last but not least, we also push the semi-supervised setting to the realistic omni-supervised setting, demonstrating significantly promoted performance on a newly annotated ARKitScenes testing set. Our codes, data and models are made publicly available**Code: https://github.com/AIR-DISCOVER/Omni-PQ.
Huan-ang Gao, Beiwen Tian, Pengfei Li 0007, Xiaoxue Chen, Hao Zhao 0002, Guyue Zhou, Yurong Chen 0001, Hongbin Zha
ICRA8
2023 Online Adaptive Disparity Estimation for Dynamic Scenes in Structured Light Systems
abstract
In recent years, deep neural networks have shown remarkable progress in dense disparity estimation from dynamic scenes in monocular structured light systems. However, their performance significantly drops when applied in unseen environments. To address this issue, self-supervised online adaptation has been proposed as a solution to bridge this performance gap. Unlike traditional fine-tuning processes, online adaptation performs test-time optimization to adapt networks to new domains. Therefore, achieving fast convergence during the adaptation process is critical for attaining satisfactory accuracy. In this paper, we propose an unsupervised loss function based on long sequential inputs. It ensures better gradient directions and faster convergence. Our loss function is designed using a multi-frame pattern flow, which comprises a set of sparse trajectories of the projected pattern along the sequence. We estimate the sparse pseudo ground truth with a confidence mask using a filter-based method, which guides the online adaptation process. Our proposed framework significantly improves the online adaptation speed and achieves superior performance on unseen data. The code is available on https://github.com/CodePointer/TIDENet.
Rukun Qiao, Hiroshi Kawasaki, Hongbin Zha
IROS3
2023 View-relation constrained global representation learning for multi-view-based 3D object recognition
Ruchang Xu, Qing Mi, Wei Ma 0008, Hongbin Zha
Appl. Intell.4
2023 DSC-MDE: Dual structural contexts for monocular depth estimation
Wubin Yan, Lijun Dong, Wei Ma 0008, Qing Mi, Hongbin Zha
Knowl. Based Syst.5
2023 Bi-Graph Reasoning for Masticatory Muscle Segmentation From Cone-Beam Computed Tomography
abstract
Automated segmentation of masticatory muscles is a challenging task considering ambiguous soft tissue attachments and image artifacts of low-radiation cone-beam computed tomography (CBCT) images. In this paper, we propose a bi-graph reasoning model (BGR) for the simultaneous detection and segmentation of multi-category masticatory muscles from CBCTs. The BGR exploits the local and long-range interdependencies of regions of interest and category-specific prior knowledge of masticatory muscles by reasoning on the category graph and the region graph. The category graph of the learnable muscle prior knowledge handles high-level dependencies of muscle categories, enhancing the feature representation with noise-agnostic category knowledge. The region graph models both local and global dependencies of the candidate muscle regions of interest. The proposed BGR accommodates the high-level dependencies and enhances the region features in the presence of entangled soft tissue and image artifacts. We evaluated the proposed approach by segmenting masticatory muscles on clinically acquired CBCTs. Extensive experimental results show that the BGR effectively segments masticatory muscles with state-of-the-art accuracy.
Yicheng Zhong, Yuru Pei, Kaichen Nie, Yungeng Zhang, Tianmin Xu, Hongbin Zha
IEEE Trans. Medical Imaging6
2022 SC-wLS: Towards Interpretable Feed-forward Camera Re-localization
Hao Zhao 0002, Shunkai Li, Yingdian Cao, Hongbin Zha
ECCV (1)5
2022 ReINView: Re-interpreting Views for Multi-view 3D Object Recognition
abstract
Multi-view-based 3D object recognition is important in robot-environment interaction. However, recent methods simply extract features from each view via convolutional neural networks (CNNs) and then fuse these features together to make predictions. These methods ignore the inherent ambiguities of each view caused due to 3D-2D projection. To address this problem, we propose a novel deep framework for multi-view-based 3D object recognition. Instead of fusing the multi-view features directly, we design a re-interpretation module (ReINView) to eliminate the ambiguities at each view. To achieve this, ReINView re-interprets view features patch by patch by using their context from nearby views, considering that local patches are generally co-visible at nearby viewpoints. Since contour shapes are essential for 3D object recognition as well, ReINView further performs view-level re-interpretation, in which we use all the views as context sources since the target contours to be re-interpreted are globally observable. The re-interpreted multi-view features can better reflect the 3D global and local structures of the object. Experiments on both ModelNet40 and ModelNet10 show that the proposed model outperforms state-of-the-art methods in 3D object recognition.
Ruchang Xu, Wei Ma 0008, Qing Mi, Hongbin Zha
IROS4
2022 Multiview Feature Aggregation for Facade Parsing
abstract
Facade image parsing is essential to the semantic understanding and 3-D reconstruction of urban scenes. Considering the occlusion and appearance ambiguity in single-view images and the easy acquisition of multiple views, in this letter, we propose a multiview enhanced deep architecture for facade parsing. The highlight of this architecture is a cross-view feature aggregation module that can learn to choose and fuse useful convolutional neural network (CNN) features from nearby views to enhance the representation of a target view. Benefitting from the multiview enhanced representation, the proposed architecture can better deal with the ambiguity and occlusion issues. Moreover, our cross-view feature aggregation module can be straightforwardly integrated into existing single-image parsing frameworks. Extensive comparison experiments and ablation studies are conducted to demonstrate the good performance of the proposed method and the validity and transportability of the cross-view feature aggregation module.
Wenguang Ma, Shibiao Xu, Wei Ma 0008, Hongbin Zha
IEEE Geosci. Remote. Sens. Lett.4
2022 Dense correspondence of deformable volumetric images via deep spectral embedding and descriptor learning
Diya Sun, Yuru Pei, Yungeng Zhang, Tianmin Xu, Tianbing Wang, Hongbin Zha
Medical Image Anal.6
2022 Face Restoration via Plug-and-Play 3D Facial Priors
abstract
State-of-the-art face restoration methods employ deep convolutional neural networks (CNNs) to learn a mapping between degraded and sharp facial patterns by exploring local appearance knowledge. However, most of these methods do not well exploit facial structures and identity information, and only deal with task-specific face restoration (e.g., face super-resolution or deblurring). In this paper, we propose cross-tasks and cross-models plug-and-play 3D facial priors to explicitly embed the network with the sharp facial structures for general face restoration tasks. Our 3D priors are the first to explore 3D morphable knowledge based on the fusion of parametric descriptions of face attributes (e.g., identity, facial expression, texture, illumination, and face pose). Furthermore, the priors can easily be incorporated into any network and are very efficient in improving the performance and accelerating the convergence speed. Firstly, a 3D face rendering branch is set up to obtain 3D priors of salient facial structures and identity knowledge. Secondly, for better exploiting this hierarchical information (i.e., intensity similarity, 3D facial structure, and identity content), a spatial attention module is designed for the image restoration problems. Extensive face restoration experiments including face super-resolution and deblurring demonstrate that the proposed 3D priors achieve superior face restoration results over the state-of-the-art algorithms.
Xiaobin Hu, Wenqi Ren, Jiaolong Yang, Xiaochun Cao, David P. Wipf, Bjoern Menze, Xin Tong 0001, Hongbin Zha
IEEE Trans. Pattern Anal. Mach. Intell.8
2022 Deep Visual Odometry With Adaptive Memory
abstract
We propose a novel deep visual odometry (VO) method that considers global information by selecting memory and refining poses. Existing learning-based methods take the VO task as a pure tracking problem via recovering camera poses from image snippets, leading to severe error accumulation. Global information is crucial for alleviating accumulated errors. However, it is challenging to effectively preserve such information for end-to-end systems. To deal with this challenge, we design an adaptive memory module, which progressively and adaptively saves the information from local to global in a neural analogue of memory, enabling our system to process long-term dependency. Benefiting from global information in the memory, previous results are further refined by an additional refining module. With the guidance of previous outputs, we adopt a spatial-temporal attention to select features for each view based on the co-visibility in feature domain. Specifically, our architecture consisting of Tracking, Remembering and Refining modules works beyond tracking. Experiments on the KITTI and TUM-RGBD datasets demonstrate that our approach outperforms state-of-the-art methods by large margins and produces competitive results against classic approaches in regular scenes. Moreover, our model achieves outstanding performance in challenging scenarios such as texture-less regions and abrupt motions, where classic algorithms tend to fail.
Xin Wang 0072, Junqiu Wang, Hongbin Zha
IEEE Trans. Pattern Anal. Mach. Intell.4
2022 All-Higher-Stages-In Adaptive Context Aggregation for Semantic Edge Detection
abstract
Convolutional Neural Networks (CNNs) can reveal local variation details and multi-scale spatial context in images via low-to-high stages of feature expression; effective fusion of these raw features is key to Semantic Edge Detection (SED). The methods available in the field generally fuse features across stages in a position-aligned mode, which cannot satisfy the requirements of diverse semantic context in categorizing different pixels. In this paper, we propose a deep framework for SED, the core of which is a new multi-stage feature fusion structure, called All-HiS-In ACA (All-Higher-Stages-In Adaptive Context Aggregation). All-HiS-In ACA can adaptively select semantic context from all higher-stages for detailed features via a cross-stage self-attention paradigm, and thus can obtain fused features with high-resolution details for edge localization and rich semantics for edge categorization. In addition, we develop a non-parametric Inter-layer Complementary Enhancement (ICE) module to supplement clues at each stage with their counterparts in adjacent stages. The ICE-enhanced multi-stage features are then fed into the All-HiS-In ACA module. We also construct an Object-level Semantic Integration (OSI) module to further refine the fused features by enforcing the consistency of the features within the same object. Extensive experiments demonstrate the superior performance of the proposed method over state-of-the-art works.
Qihan Bo, Wei Ma 0008, Yukun Lai, Hongbin Zha
IEEE Trans. Circuits Syst. Video Technol.4
2022 Progressive Feature Learning for Facade Parsing With Occlusions
abstract
Existing deep models for facade parsing often fail in classifying pixels in heavily occluded regions of facade images due to the difficulty in feature representation of these pixels. In this paper, we solve facade parsing with occlusions by progressive feature learning. To this end, we locate the regions contaminated by occlusions via Bayesian uncertainty evaluation on categorizing each pixel in these regions. Then, guided by the uncertainty, we propose an occlusion-immune facade parsing architecture in which we progressively re-express the features of pixels in each contaminated region from easy to hard. Specifically, the outside pixels, which have reliable context from visible areas, are re-expressed at early stages; the inner pixels are processed at late stages when their surroundings have been decontaminated at the earlier stages. In addition, at each stage, instead of using regular square convolution kernels, we design a context enhancement module (CEM) with directional strip kernels, which can aggregate structural context to re-express facade pixels. Extensive experiments on popular facade datasets demonstrate that the proposed method achieves state-of-the-art performance.
Wenguang Ma, Shibiao Xu, Wei Ma 0008, Xiaopeng Zhang 0001, Hongbin Zha
IEEE Trans. Image Process.5
2022 Unsupervised Domain Adaptation for Semantic Segmentation of Urban Street Scenes Reflected by Convex Mirrors
abstract
Reflective convex mirrors are often used on street corners or as passenger-side mirrors on cars to obtain scene information by reflecting blind spots in the field of view, which can provide safety for pedestrians and drivers on roads, driveways, and alleys that lack of visibility. In recent years, deep learning based scene understanding methods (e.g., semantic segmentation) have been rapidly developed. However, due to gaps in the geometric domain, models trained on normal images are not directly applicable to scenes with convex mirror reflections. In this paper, we propose a novel framework to reduce the domain gap between normal images and convex mirror reflection images. In particular, we geometrically model convex mirrors to obtain a differentiable convex mirror simulation layer, CMSL. With the help of CMSL, we perform adversarial domain adaptation on edges in the input space and semantic boundaries in the output space to reduce the geometric appearance gap between the synthetic and real images. To verify the effectiveness of our algorithm, we construct the first convex mirror reflection scene dataset CMR1K, which contains 268 images with fine annotations. Extensive experimental results show that our algorithm can significantly outperform the baseline and previous methods. For example, our method surpasses the baseline and AdvEnt by 10% and 3% in mIoU, respectively.
Yongjie Shi, Xianghua Ying, Hongbin Zha
IEEE Trans. Intell. Transp. Syst.3
2022 Deep Volumetric Descriptor Learning for Dense Correspondence of Cone-Beam Computed Tomography via Spectral Maps
Diya Sun, Yungeng Zhang, Yuru Pei, Peixin Li, Kaichen Nie, Tianmin Xu, Tianbing Wang, Hongbin Zha
IEEE Trans. Medical Imaging8
2021 Generalizing to the Open World: Deep Visual Odometry With Online Adaptation
abstract
Despite learning-based visual odometry (VO) has shown impressive results in recent years, the pretrained networks may easily collapse in unseen environments. The large domain gap between training and testing data makes them difficult to generalize to new scenes. In this paper, we propose an online adaptation framework for deep VO with the assistance of scene-agnostic geometric computations and Bayesian inference. In contrast to learning-based pose estimation, our method solves pose from optical flow and depth while the single-view depth estimation is continuously improved with new observations by online learned uncertainties. Meanwhile, an online learned photometric uncertainty is used for further depth and pose optimization by a differentiable Gauss-Newton layer. Our method enables fast adaptation of deep VO networks to unseen environments in a self-supervised manner. Extensive experiments including Cityscapes to KITTI and outdoor KITTI to indoor TUM demonstrate that our method achieves state-of-the-art generalization ability among self-supervised VO methods.
Shunkai Li, Yingdian Cao, Hongbin Zha
CVPR4
2021 Online Learning of a Probabilistic and Adaptive Scene Representation
abstract
Constructing and maintaining a consistent scene model on-the-fly is the core task for online spatial perception, interpretation, and action. In this paper, we represent the scene with a Bayesian nonparametric mixture model, seamlessly describing per-point occupancy status with a continuous probability density function. Instead of following the conventional data fusion paradigm, we address the problem of online learning the process how sequential point cloud data are generated from the scene geometry. An incremental and parallel inference is performed to update the parameter space in real-time. We experimentally show that the proposed representation achieves state-of-the-art accuracy with promising efficiency. The consistent probabilistic formulation assures a generative model that is adaptive to different sensor characteristics, and the model complexity can be dynamically adjusted on-the-fly according to different data scales.
Zike Yan, Xin Wang 0072, Hongbin Zha
CVPR3
2021 Continual Neural Mapping: Learning An Implicit Scene Representation from Sequential Observations
abstract
Recent advances have enabled a single neural network to serve as an implicit scene representation, establishing the mapping function between spatial coordinates and scene properties. In this paper, we make a further step towards continual learning of the implicit scene representation directly from sequential observations, namely Continual Neural Mapping. The proposed problem setting bridges the gap between batch-trained implicit neural representations and commonly used streaming data in robotics and vision communities. We introduce an experience replay approach to tackle an exemplary task of continual neural mapping: approximating a continuous signed distance function (SDF) from sequential depth images as a scene geometry representation. We show for the first time that a single network can represent scene geometry over time continually without catastrophic forgetting, while achieving promising trade-offs between accuracy and efficiency.
Zike Yan, Xuesong Shi, Peng Wang 0001, Hongbin Zha
ICCV6
2021 Spectral Embedding Approximation and Descriptor Learning for Craniofacial Volumetric Image Correspondence
Diya Sun, Yungeng Zhang, Yuru Pei, Tianmin Xu, Hongbin Zha
MICCAI (4)5
2021 Learning Dual Transformer Network for Diffeomorphic Registration
Yungeng Zhang, Yuru Pei, Hongbin Zha
MICCAI (4)3
2021 Robust 3D face reconstruction from single noisy depth image through semantic consistency
abstract
Abstract This paper addresses the 3D face reconstruction and semantic annotation from a single‐view noisy depth image. A deep neural network‐based coarse‐to‐fine framework is presented to take advantage of 3D morphable model (3DMM) regression and per‐vertex geometry refinement. The low‐dimensional subspace coefficients of the 3DMM initialize the global facial geometry, being prone to be over‐smooth because of the low‐pass characteristics of the shape subspace. The proposed geometry refinement subnetwork predicts per‐vertex displacements to enrich local details, which is learned from unlabelled noisy depth images based on the registration‐like loss. In order to guarantee the semantic correspondence between the resultant 3D face and the depth image, a semantic consistency constraint is introduced to adapt an annotation model learned from the synthetic data to real noisy depth images. The resultant depth annotations are required to be consistent with the label propagation from the coarse and refined parametric 3D faces. The proposed coarse‐to‐fine reconstruction scheme and the semantic consistency constraint are evaluated on the depth‐based 3D face reconstruction and semantic annotation. The series of experiments demonstrate that the proposed approach achieves the performance improvements over compared methods regarding 3D face reconstruction and depth image annotation.
Peixin Li, Yuru Pei, Yicheng Zhong, Yuke Guo, Hongbin Zha
IET Comput. Vis.5
2021 Pyramid ALKNet for Semantic Parsing of Building Facade Image
abstract
The semantic parsing of building facade images is a fundamental yet challenging task in urban scene understanding. Existing works sought to tackle this task by using facade grammars or convolutional neural networks (CNNs). The former can hardly generate parsing results coherent with real images while the latter often fails to capture relationships among facade elements. In this letter, we propose a pyramid atrous large kernel (ALK) network (ALKNet) for the semantic segmentation of facade images. The pyramid ALKNet captures long-range dependencies among building elements by using ALK modules in multiscale feature maps. It makes full use of the regular structures of facades to aggregate useful nonlocal context information and thereby is capable of dealing with challenging image regions caused by occlusions, ambiguities, and so on. Experiments on both rectified and unrectified facade data sets show that ALKNet has better performances than those of state-of-the-art methods.
Wenguang Ma, Wei Ma 0008, Shibiao Xu, Hongbin Zha
IEEE Geosci. Remote. Sens. Lett.4
2021 Reconstruction regularized low-rank subspace learning for cross-modal retrieval
Jianlong Wu, Xingxu Xie, Liqiang Nie, Zhouchen Lin, Hongbin Zha
Pattern Recognit.5
2021 Line Flow Based Simultaneous Localization and Mapping
abstract
In this article, we propose a visual simultaneous localization and mapping (SLAM) method by predicting and updating line flows that represent sequential 2-D projections of 3-D line segments. While feature-based SLAM methods have achieved excellent results, they still face problems in challenging scenes containing occlusions, blurred images, and repetitive textures. To address these problems, we leverage a line flow to encode the coherence of line segment observations of the same 3-D line along the temporal dimension, which has been neglected in prior SLAM systems. Thanks to this line flow representation, line segments in a new frame can be predicted according to their corresponding 3-D lines and their predecessors along the temporal dimension. We create, update, merge, and discard line flows on-the-fly. We model the proposed line flow based SLAM (LF-SLAM) using a Bayesian network. Extensive experimental results demonstrate that the proposed LF-SLAM method achieves state-of-the-art results due to the utilization of line flows. Specifically, LF-SLAM obtains good localization and mapping results in challenging scenes with occlusions, blurred images, and repetitive textures.
Qiuyuan Wang, Zike Yan, Junqiu Wang, Wei Ma 0008, Hongbin Zha
IEEE Trans. Robotics6
2020 FC-vSLAM: Integrating Feature Credibility in Visual SLAM
abstract
Feature-based visual SLAM (vSLAM) systems compute camera poses and scene maps by detecting and matching 2D features, mostly being points and line segments, from image sequences. These systems often suffer from unreliable detections. In this paper, we define feature credibility (FC) for both points and line segments, formulate it into vSLAMs and develop an FC-vSLAM system based on the widely used ORB-SLAM framework. Compared with existing credibility definitions, the proposed one, considering both temporal observation stability and perspective triangulation reliability, is more comprehensive. We formulate the credibility in our SLAM system to suppress the influences from unreliable features on the pose and map optimization. We also present a way to improve the line end observations by their multi-view correspondences, to improve the integrity of the 3D maps. Experiments on both the TUM and 7-Scenes datasets demonstrate that our feature credibility and the multi-view line optimization are effective; the developed FC-vSLAM system outperforms existing popular feature-based systems in both localization and mapping.
Shuai Xie, Wei Ma 0008, Qiuyuan Wang, Ruchang Xu, Hongbin Zha
3DV5
2020 Unified Graph and Low-Rank Tensor Learning for Multi-View Clustering
abstract
Multi-view clustering aims to take advantage of multiple views information to improve the performance of clustering. Many existing methods compute the affinity matrix by low-rank representation (LRR) and pairwise investigate the relationship between views. However, LRR suffers from the high computational cost in self-representation optimization. Besides, compared with pairwise views, tensor form of all views' representation is more suitable for capturing the high-order correlations among all views. Towards these two issues, in this paper, we propose the unified graph and low-rank tensor learning (UGLTL) for multi-view clustering. Specifically, on the one hand, we learn the view-specific affinity matrix based on projected graph learning. On the other hand, we reorganize the affinity matrices into tensor form and learn its intrinsic tensor based on low-rank tensor approximation. Finally, we unify these two terms together and jointly learn the optimal projection matrices, affinity matrices and intrinsic low-rank tensor. We also propose an efficient algorithm to iteratively optimize the proposed model. To evaluate the performance of the proposed method, we conduct extensive experiments on multiple benchmarks across different scenarios and sizes. Compared with the state-of-the-art approaches, our method achieves much better performance.
Jianlong Wu, Xingyu Xie, Liqiang Nie, Zhouchen Lin, Hongbin Zha
AAAI5
2020 Fully Convolutional Network for Consistent Voxel-Wise Correspondence
abstract
In this paper, we propose a fully convolutional network-based dense map from voxels to invertible pair of displacement vector fields regarding a template grid for the consistent voxel-wise correspondence. We parameterize the volumetric mapping using a convolutional network and train it in an unsupervised way by leveraging the spatial transformer to minimize the gap between the warped volumetric image and the template grid. Instead of learning the unidirectional map, we learn the nonlinear mapping functions for both forward and backward transformations. We introduce the combinational inverse constraints for the volumetric one-to-one maps, where the pairwise and triple constraints are utilized to learn the cycle-consistent correspondence maps between volumes. Experiments on both synthetic and clinically captured volumetric cone-beam CT (CBCT) images show that the proposed framework is effective and competitive against state-of-the-art deformable registration techniques.
Yungeng Zhang, Yuru Pei, Yuke Guo, Gengyu Ma, Tianmin Xu, Hongbin Zha
AAAI6
2020 An Unsupervised Approach for 3D Face Reconstruction from a Single Depth Image
Peixin Li, Yuru Pei, Yicheng Zhong, Yuke Guo, Gengyu Ma, Wenhai Wu, Hongbin Zha
CGI9
2020 Self-Supervised Deep Visual Odometry With Online Adaptation
abstract
Self-supervised VO methods have shown great success in jointly estimating camera pose and depth from videos. However, like most data-driven methods, existing VO networks suffer from a notable decrease in performance when confronted with scenes different from the training data, which makes them unsuitable for practical applications. In this paper, we propose an online meta-learning algorithm to enable VO networks to continuously adapt to new environments in a self-supervised manner. The proposed method utilizes convolutional long short-term memory (convLSTM) to aggregate rich spatial-temporal information in the past. The network is able to memorize and learn from its past experience for better estimation and fast adaptation to the current frame. When running VO in the open world, in order to deal with the changing environment, we propose an online feature alignment method by aligning feature distributions at different time. Our VO network is able to seamlessly adapt to different environments. Extensive experiments on unseen outdoor scenes, virtual to real world and outdoor to indoor environments demonstrate that our method consistently outperforms state-of-the-art self-supervised VO baselines considerably.
Shunkai Li, Xin Wang 0072, Yingdian Cao, Zike Yan, Hongbin Zha
CVPR6
2020 RDCFace: Radial Distortion Correction for Face Recognition
abstract
The effects of radial lens distortion often appear in wide-angle cameras of surveillance and safeguard systems, which may severely degrade performances of previous face recognition algorithms. Traditional methods for radial lens distortion correction usually employ line features in scenarios that are not suitable for face images. In this paper, we propose a distortion-invariant face recognition system called RDCFace, which directly and only utilize the distorted images of faces, to alleviate the effects of radial lens distortion. RDCFace is an end-to-end trainable cascade network, which can learn rectification and alignment parameters to achieve a better face recognition performance without requiring supervision of facial landmarks and distortion parameters. We design sequential spatial transformer layers to optimize the correction, alignment, and recognition modules jointly. The feasibility of our method comes from implicitly using the statistics of the layout of face features learned from the large-scale face data. Extensive experiments indicate that our method is distortion robust and gains significant improvements on LFW, YTF, CFP, and RadialFace, a real distorted face benchmark compared with state-of-the-art methods.
He Zhao 0006, Xianghua Ying, Yongjie Shi, Xin Tong 0007, Jingsi Wen, Hongbin Zha
CVPR6
2020 Face Denoising and 3D Reconstruction from A Single Depth Image
abstract
The reconstruction of 3D face shapes and expressions from a single depth image obtained by a consumer depth camera is a challenging issue considering device-specific noise, the data missing, and the lack of textual constraints. In order to relieve the computationally-intensive nonlinear optimization of traditional template-fitting-based methods, we aim to build an end-to-end regression framework between a depth image and a 3D face encoded by the identity, the expression, and the pose parameters. Concerning the lack of paired depth images and 3D faces, we utilize the unsupervised CycleGAN-based network to adapt the regression model learned from the synthetic data to the real-captured noisy depth images. Instead of separate depth image denoising and 3D face inference, we present a task-specific coupled loss for end-to-end 3D face estimation. We propose a three-tier constraint for the shape consistency in the joint embedding, the depth image, and the surface space to avoid shape distortions in the unsupervised domain adaptation network. We report promising qualitative results for the task of the face denoising and 3D face reconstruction from a single depth image.
Yicheng Zhong, Yuru Pei, Peixin Li, Yuke Guo, Gengyu Ma, Wenhai Wu, Hongbin Zha
FG9
2020 A Simple Yet Effective Pipeline For Radial Distortion Correction
abstract
Eliminating the radial lens distortion of an image is a crucial preprocessing step for many computer vision applications. This paper explores a simple yet effective pipeline for radial distortion correction. Different from existing state-of-the-art methods that design complex network structure and concatenate multi-branch features. Our model uses a single network without any additional supervision. We design two differentiable layers to synthesize and rectify distorted images efficiently. Based on these layers, an online data synthesis strategy, a sampling grid loss, and an image reprojection loss are proposed to improve the distortion correction accuracy. Compared with the state-of-the-art methods, our model achieves the best rectification quality on both the synthetic and real distorted images with dozens of times faster inference speed. The training data and codes will be released.11https://github.com/MccreeZhao/RDCPipeline
He Zhao 0006, Yongjie Shi, Xin Tong 0007, Xianghua Ying, Hongbin Zha
ICIP5
2020 Qamface: Quadratic Additive Angular Margin Loss For Face Recognition
abstract
The angular-based softmax losses and their variants achieve great success in face recognition based on deep learning. ArcFace [1] which directly maximize decision boundary in angular space is one of the most popular and effective loss function. In this paper, we analyze the inherent limitations of ArcFace, including the non-monotonic logit and gradient curve, and inappropriate trend of loss value. To address these problems, we propose a novel loss function named the Quadratic Additive Angular Margin Loss (QAMFace). It takes the value of the angle through a quadratic function rather than cosine function as the target logit. Our QAMFace is easy to implement and only adds negligible computational overhead. Experiments on several relevant benchmarks show that QAMFace performs better in convergence on feature embedding, and consistently outperforms the state-of-the-art face recognition methods. Our codes will be released soon.1
He Zhao 0006, Yongjie Shi, Xin Tong 0007, Xianghua Ying, Hongbin Zha
ICIP5
2020 Boundary-aware Graph Convolution for Semantic Segmentation
abstract
Recent works have made great progress in semantic segmentation by exploiting contextual information in a local or global manner with dilated convolutions, pyramid pooling or self-attention mechanism. However, few works have focused on harvesting boundary information to improve the segmentation performance. In order to enhance the feature similarity within the object and keep discrimination from other objects, we propose a boundary-aware graph convolution (BGC) module to propagate features within the object. The graph reasoning is performed among pixels of the same object apart from the boundary pixels. Based on the proposed BGC module, we further introduce the Boundary-aware Graph Convolution N et-work(BGCN et), which consists of two main components including a basic segmentation network and the BGC module, forming a coarse-to-fine paradigm. Specifically, the BGC module takes the coarse segmentation feature map as node features and boundary prediction to guide graph construction. After graph convolution, the reasoned feature and the input feature are fused together to get the refined feature, producing the refined segmentation result. We conduct extensive experiments on three popular semantic segmentation benchmarks including Cityscapes, PASCAL VOC 2012 and COCO Stuff, and achieve state-of-the-art performance on all three benchmarks.
Hanzhe Hu, Jinshi Cui, Hongbin Zha
ICPR3
2020 Position-aware and Symmetry Enhanced GAN for Radial Distortion Correction
abstract
This paper presents a novel method based on the generative adversarial network for radial distortion correction. Instead of generating a corrected image, our generator predicts a pixel flow map to measure the pixel offset between the distorted and corrected image. The quality of the generated pixel flow map and the warped image are judged by the discriminator. As texture far away from the image center has strong distortion, we develop an Adaptive Inverted Foveal layer which can transform the deformation to the intensity of the image to exploit this property. Rotation symmetry enhanced convolution kernels are applied to extract geometric features of different orientations explicitly. These learned features are recalibrated using the Squeeze-and-Excitation block to assign different weights for different directions. Moreover, we construct a first real-world radial distorted image dataset RD600 annotated with ground truth to evaluate our proposed method. We conduct extensive experiments to validate the effectiveness of each part of our framework. The further experiment shows our approach outperforms previous methods in both synthetic and real-world datasets quantitatively and qualitatively.
Yongjie Shi, Xin Tong 0007, Jingsi Wen, He Zhao 0006, Xianghua Ying, Hongbin Zha
ICPR6
2020 G-FAN: Graph-Based Feature Aggregation Network for Video Face Recognition
abstract
In this paper, we propose a graph-based feature aggregation network (G-FAN) for video face recognition. Compared with the still image, video face recognition exhibits great challenges due to huge intra-class variability and high interclass ambiguity. To address this problem, our G-FAN first uses a Convolutional Neural Network to extract deep features for every input face of a subject. Then, we build an affinity graph based on the relationship between facial features and apply Graph Convolutional Network to generate fine-grained quality vectors for each frame. Finally, the features among multiple frames are adaptively aggregated into a discriminative vector to represent a video face. Different from previous works that take a single image as input, our G-FAN could utilize the correlation information between image pairs and aggregate a template of face images simultaneously. The experiments on video face recognition benchmarks, including YTF, IJB-A, and IJB-C show that: (i) G-FAN automatically learns to advocate high-quality frames while repelling low-quality ones. (ii) G-FAN significantly boosts recognition accuracy and outperforms other state-of-the-art aggregation methods.
He Zhao 0006, Yongjie Shi, Xin Tong 0007, Jingsi Wen, Xianghua Ying, Hongbin Zha
ICPR6
2020 3D Orientation Estimation and Vanishing Point Extraction from Single Panoramas Using Convolutional Neural Network
abstract
3D orientation estimation is a key component of many important computer vision tasks such as autonomous navigation and 3D scene understanding. This paper presents a new CNN architecture to estimate the 3D orientation of an omnidirectional camera with respect to the world coordinate system from a single spherical panorama. To train the proposed architecture, we leverage a dataset of panoramas named VOP60K from Google Street View with labeled 3D orientation, including 50 thousand panoramas for training and 10 thousand panoramas for testing. Previous approaches usually estimate 3D orientation under pinhole cameras. However, for a panorama, due to its larger field of view, previous approaches cannot be suitable. In this paper, we propose an edge extractor layer to utilize the low-level and geometric information of panorama, an attention module to fuse different features generated by previous layers. A regression loss for two column vectors of the rotation matrix and classification loss for the position of vanishing points are added to optimize our network simultaneously. The proposed algorithm is validated on our benchmark, and experimental results clearly demonstrate that it outperforms previous methods.
Yongjie Shi, Xin Tong 0007, Jingsi Wen, He Zhao 0006, Xianghua Ying, Hongbin Zha
ICRA6
2020 Automatic Tooth Segmentation and Dense Correspondence of 3D Dental Model
Diya Sun, Yuru Pei, Peixin Li, Guangying Song, Yuke Guo, Hongbin Zha, Tianmin Xu
MICCAI (4)6
2019 Pose-Aware Face Alignment based on CNN and 3DMM
Songjiang Li, Honggai Li, Jinshi Cui, Hongbin Zha
BMVC4
2019 Beyond Tracking: Selecting Memory and Refining Poses for Deep Visual Odometry
abstract
Most previous learning-based visual odometry (VO) methods take VO as a pure tracking problem. In contrast, we present a VO framework by incorporating two additional components called Memory and Refining. The Memory component preserves global information by employing an adaptive and efficient selection strategy. The Refining component ameliorates previous results with the contexts stored in the Memory by adopting a spatial-temporal attention mechanism for feature distilling. Experiments on the KITTI and TUM-RGBD benchmark datasets demonstrate that our method outperforms state-of-the-art learning-based methods by a large margin and produces competitive results against classic monocular VO approaches. Especially, our model achieves outstanding performance in challenging scenarios such as texture-less regions and abrupt motions, where classic VO algorithms tend to fail.
Xin Wang 0072, Shunkai Li, Qiuyuan Wang, Junqiu Wang, Hongbin Zha
CVPR6
2019 Sequential Adversarial Learning for Self-Supervised Deep Visual Odometry
abstract
We propose a self-supervised learning framework for visual odometry (VO) that incorporates correlation of consecutive frames and takes advantage of adversarial learning. Previous methods tackle self-supervised VO as a local structure from motion (SfM) problem that recovers depth from single image and relative poses from image pairs by minimizing photometric loss between warped and captured images. As single-view depth estimation is an ill-posed problem, and photometric loss is incapable of discriminating distortion artifacts of warped images, the estimated depth is vague and pose is inaccurate. In contrast to previous methods, our framework learns a compact representation of frame-to-frame correlation, which is updated by incorporating sequential information. The updated representation is used for depth estimation. Besides, we tackle VO as a self-supervised image generation task and take advantage of Generative Adversarial Networks (GAN). The generator learns to estimate depth and pose to generate a warped target image. The discriminator evaluates the quality of generated image with high-level structural perception that overcomes the problem of pixel-wise loss in previous methods. Experiments on KITTI and Cityscapes datasets show that our method obtains more accurate depth with details preserved and predicted pose outperforms state-of-the-art self-supervised methods significantly.
Shunkai Li, Xin Wang 0072, Zike Yan, Hongbin Zha
ICCV5
2019 Deep Comprehensive Correlation Mining for Image Clustering
abstract
Recent developed deep unsupervised methods allow us to jointly learn representation and cluster unlabelled data. These deep clustering methods %like DAC start with mainly focus on the correlation among samples, e.g., selecting high precision pairs to gradually tune the feature representation, which neglects other useful correlations. In this paper, we propose a novel clustering framework, named deep comprehensive correlation mining~(DCCM), for exploring and taking full advantage of various kinds of correlations behind the unlabeled data from three aspects: 1) Instead of only using pair-wise information, pseudo-label supervision is proposed to investigate category information and learn discriminative features. 2) The features' robustness to image transformation of input space is fully explored, which benefits the network learning and significantly improves the performance. 3) The triplet mutual information among features is presented for clustering problem to lift the recently discovered instance-level deep mutual information to a triplet-level formation, which further helps to learn more discriminative features. Extensive experiments on several challenging datasets show that our method achieves good performance, e.g., attaining 62.3% clustering accuracy on CIFAR-10, which is 10.1% higher than the state-of-the-art results.
Jianlong Wu, Keyu Long, Fei Wang 0032, Chen Qian 0006, Cheng Li 0009, Zhouchen Lin, Hongbin Zha
ICCV7
2019 Local Supports Global: Deep Camera Relocalization With Sequence Enhancement
abstract
We propose to leverage the local information in a image sequence to support global camera relocalization. In contrast to previous methods that regress global poses from single images, we exploit the spatial-temporal consistency in sequential images to alleviate uncertainty due to visual ambiguities by incorporating a visual odometry (VO) component. Specifically, we introduce two effective steps called content-augmented pose estimation and motion-based refinement. The content-augmentation step focuses on alleviating the uncertainty of pose estimation by augmenting the observation based on the co-visibility in local maps built by the VO stream. Besides, the motion-based refinement is formulated as a pose graph, where the camera poses are further optimized by adopting relative poses provided by the VO component as additional motion constraints. Thus, the global consistency can be guaranteed. Experiments on the public indoor 7-Scenes and outdoor Oxford RobotCar benchmark datasets demonstrate that benefited from local information inherent in the sequence, our approach outperforms state-of-the-art methods, especially in some challenging cases, e.g., insufficient texture, highly repetitive textures, similar appearances, and over-exposure.
Xin Wang 0072, Zike Yan, Qiuyuan Wang, Junqiu Wang, Hongbin Zha
ICCV6
2019 Three Orthogonal Vanishing Points Estimation in Structured Scenes Using Convolutional Neural Networks
abstract
Inferring 3D geometric cues is a crucial step, whereas vanishing point plays a very important role in image understanding from a single image of structured scenes. In this paper, we construct a 330-thousand-item image database of structured scenes labeled by vanishing points, focal length and camera orientation. We grab over 300 thousand Google Street View images which cover the downtown and neighboring areas of New York, Los Angeles, Chicago and etc. The prediction error is characterized by a loss function by imposing a regularization item derived from the geometric constraint of orthogonal vanishing points and focal length. Moreover, we collect about 30 thousand indoor images using a full 360-degree panorama camera taken by ourselves in room, office, library and etc. We also using Convolutional Neural Networks to transfer learning from street view images to indoor images. Extensive experiments demonstrate that our algorithm outperforms state-of-the-art non-learned approaches.
Yongjie Shi, Danfeng Zhang, Jingsi Wen, Xin Tong 0007, He Zhao 0006, Xianghua Ying, Hongbin Zha
ICIP7
2019 Camera and LiDAR Fusion for On-road Vehicle Tracking with Reinforcement Learning
abstract
We formulate camera and LiDAR fusion tracking as a sequential decision-making process. With our deep reinforcement learning framework, we try to optimize the tracking trajectory to be as accurate, smooth, and long as possible. In contrast to traditional fusion algorithms involving complex feature and strategy design and hyperparameters tuned for different scenarios, our fusion agent can learn the confidence of each input by tracking the results from raw observation in a data-driven fashion. Given the input states of different sensors, our approach chooses one input with a higher expected cumulative reward as the observation of a Kalman filter to iteratively predict the target position. The expected cumulative reward is estimated with a convolutional neural network, trained with a modified DQN algorithm, which takes inputs from both LiDAR and a camera. Through case studies and quantitative result evaluation on our dataset from the 4th Ring Road in Beijing, our algorithm is validated to achieve more accurate and robust tracking performance.
Yongkun Fang, Huijing Zhao, Hongbin Zha, Xijun Zhao
IV3
2019 Visual Odometry with Deep Bidirectional Recurrent Neural Networks
Xin Wang 0072, Qiuyuan Wang, Junqiu Wang, Hongbin Zha
PRCV (3)5
2019 Structure-aware SLAM with planes and lines in man-made environment
Jiyuan Zhang 0001, Hongbin Zha
Pattern Recognit. Lett.3
2019 Essential Tensor Learning for Multi-View Spectral Clustering
abstract
Recently, multi-view clustering attracts much attention, which aims to take advantage of multi-view information to improve the performance of clustering. However, most recent work mainly focuses on the self-representation-based subspace clustering, which is of high computation complexity. In this paper, we focus on the Markov chain-based spectral clustering method and propose a novel essential tensor learning method to explore the high-order correlations for multi-view representation. We first construct a tensor based on multi-view transition probability matrices of the Markov chain. By incorporating the idea from the robust principle component analysis, tensor singular value decomposition (t-SVD)-based tensor nuclear norm is imposed to preserve the low-rank property of the essential tensor, which can well capture the principle information from multiple views. We also employ the tensor rotation operator for this task to better investigate the relationship among views as well as reduce the computation complexity. The proposed method can be efficiently optimized by the alternating direction method of multipliers (ADMM). Extensive experiments on seven real-world datasets corresponding to five different applications show that our method achieves superior performance over other state-of-the-art methods.
Jianlong Wu, Zhouchen Lin, Hongbin Zha
IEEE Trans. Image Process.3
2019 On-Road Vehicle Tracking Using Part-Based Particle Filter
abstract
In this paper, we propose a part-based particle filter for on-road vehicle tracking. The proposed model combines a part-based strategy with a particle filter. By introducing a hidden state representing the center position of the vehicle, particles corresponding to vehicle parts sharing the same motion can be collectively updated in an efficient manner. By using a pre-trained appearance and geometric model, the tracker can distinguish parts with rich information from invalid parts to make more precise predictions. Meanwhile, some prior knowledge about the motion patterns of vehicles in a well-structured on-road environment is learned and can be used to infer measurement and motion models to improve tracking performance and efficiency. Experiments were conducted using the real data collected in Beijing to examine the performance of the method in different situations in terms of both its advantages and challenges. The collected Beijing highway dataset for on-road vehicle tracking will be made publicly available. We compare our method with the state-of-the-art approaches. The results demonstrate that the proposed algorithm is able to handle occlusion and the aspect ratio changes in the on-road vehicle tracking problem.
Yongkun Fang, Chao Wang 0060, Xijun Zhao, Huijing Zhao, Hongbin Zha
IEEE Trans. Intell. Transp. Syst.6
2019 Flow-based SLAM: From geometry computation to learning
abstract
Simultaneous localization and mapping (SLAM) has attracted considerable research interest from the robotics and computer-vision communities for >30 years. With steady and progressive efforts being made, modern SLAM systems allow robust and online applications in real-world scenes. We examined the evolution of this powerful perception tool in detail and noticed that the insights concerning incremental computation and temporal guidance are persistently retained. Herein, we denote this temporal continuity as a flow basis and present for the first time a survey that specifically focuses on the flow-based nature, ranging from geometric computation to the emerging learning techniques. We start by reviewing two essential stages for geometric computation, presenting the de facto standard pipeline and problem formulation, along with the utilization of temporal cues. The recently emerging techniques are then summarized, covering a wide range of areas, such as learning techniques, sensor fusion, and continuoustime trajectory modeling. This survey aims at arousing public attention on how robust SLAM systems benefit from a continuously observing nature, as well as the topics worthy of further investigation for better utilizing the temporal cues.
Zike Yan, Hongbin Zha
Virtual Real. Intell. Hardw.2
2018 Continuous-Time Stereo Visual Odometry Based on Dynamics Model
Xin Wang 0072, Zike Yan, Qiuyuan Wang, Hongbin Zha
ACCV (6)6
2018 Guided Feature Selection for Deep Visual Odometry
Qiuyuan Wang, Xin Wang 0072, Junqiu Wang, Hongbin Zha
ACCV (6)6
2018 Dense Correspondence of Cone-Beam Computed Tomography Images Using Oblique Clustering Forest
Diya Sun, Yuru Pei, Yuke Guo, Gengyu Ma, Tianmin Xu, Hongbin Zha
BMVC6
2018 Joint Dictionary Learning and Semantic Constrained Latent Subspace Projection for Cross-Modal Retrieval
abstract
With the increasing of multi-modal data on the internet, cross-modal retrieval has received a lot of attention in recent years. It aims to use one type of data as query and retrieve results of another type. For different modality data, how to reduce their heterogeneous property and preserve their local relationship are two main challenges. In this paper, we present a novel joint dictionary learning and semantic constrained latent subspace learning method for cross-modal retrieval~(JDSLC) to deal with above two issues. In this unified framework, samples from different modalities are encoded by their corresponding dictionaries to reduce the semantic gap. In the meantime, we learn modality-specific projection matrices to map the sparse coefficients into the shared latent subspace. Meanwhile, we impose a novel cross-modal similarity constraint to make the representations of samples that belong to same class but from different modalities as close as possible in the latent subspace. An efficient algorithm is proposed to jointly optimize the proposed model and learn the optimal dictionary, coefficients and projection matrix for each modality. Extensive experimental results on multiple benchmark datasets show that our method outperforms the state-of-the-art approaches.
Jianlong Wu, Zhouchen Lin, Hongbin Zha
CIKM3
2018 PSDF Fusion: Probabilistic Signed Distance Function for On-the-fly 3D Data Fusion and Scene Reconstruction
Qiuyuan Wang, Xin Wang 0072, Hongbin Zha
ECCV (9)4
2018 Recurrent Squeeze-and-Excitation Context Aggregation Net for Single Image Deraining
Xia Li 0005, Jianlong Wu, Zhouchen Lin, Hong Liu 0008, Hongbin Zha
ECCV (7)5
2018 Alternating Multi-bit Quantization for Recurrent Neural Networks
Jianqiang Yao, Zhouchen Lin, Wenwu Ou, Yuanbin Cao, Zhirong Wang, Hongbin Zha
ICLR (Poster)7
2018 Radial Lens Distortion Correction by Adding a Weight Layer with Inverted Foveal Models to Convolutional Neural Networks
abstract
Radial lens distortion often exists in images taken by commercial cameras, which does not satisfy the assumption of pinhole camera model. Eliminating the radial lens distortion of an image is necessary as a preprocessing step for many vision applications. Some paper has employed Convolutional Neural Networks (CNNs), to achieve radial distortion correction. They generated images with a large number of images of high variation of radial distortion, which can be well exploited by deep CNN with a high learning capacity, and reach the state-of-the-art results. In this paper, we claim that a weight layer with inverted foveal models can be added to these existing CNNs methods for radial distortion correction. In the widely used very deep Resnet-18 model, our method achieves about 20 percent decrease in the loss function with faster convergence compared to the previous methods.
Yongjie Shi, Danfeng Zhang, Jingsi Wen, Xin Tong 0007, Xianghua Ying, Hongbin Zha
ICPR6
2018 Recognition of Infants' Gaze Behaviors and Emotions
abstract
This paper proposes a system for recognition of infants' gaze behaviors and emotions from videos. In the current work, researchers believed that the information of eye region is crucial for gaze behavior recognition, and emotion recognition mostly depends on the appearance of face. However, because the differentiation of all parts of infant's body has not finished yet, we found that infants always express their intentions and emotions using their whole body, especially moving their heads. Therefore, we incorporate the head pose information as features into the gaze behavior recognition, and we extract the gaze features to improve the recognition of infants' emotions. In addition, we combine several Deep Neural Networks, which can not only capture the details of the images very well, but also make full use of the temporal features. In order to recognize infants' gaze behaviors, we design a feature-extraction convolutional neural network which can obtain the features of infants' gaze direction, then we feed these features with head pose into the next gaze behavior recurrent neural network. Moreover, we combine the features of facial express and gaze behavior to characterize infants' emotions, and expand this system with an emotion recurrent neural network. In the end, we achieve the recognition accuracy 98.31% and 94.71% respectively on our data set.
Bikun Yang, Jinshi Cui, Yuqiang Tong, Hongbin Zha
ICPR5
2018 Scalable Monocular SLAM by Fusing and Connecting Line Segments with Inverse Depth Filter
abstract
In this paper we propose a fast and robust line-based approach to monocular SLAM. It relies on a novel inverse depth representation of lines capable of tracking line segments in long image consequences. Tracked lines through frames provide crucial directional and positional knowledge for boosting localization performance, and they are more informative in charactering environments than points especially for urban outdoor and indoor scenes. The developed two-parameter inverse depth representation of lines is applicable for Kalman filter to achieve an efficient solver due to its linearity, which has lower computational cost compared to binary descriptors. This filter is also harmonious with inverse depth filter of points, both of which are incorporated under a unified minimization framework to enhance the performance of monocular SLAM. Real world monocular sequences have demonstrated that the proposed SLAM system outperforms the state-of-the-art and produces accurate results in both indoor and outdoor scenes.
Jiyuan Zhang 0001, Hongbin Zha
ICPR3
2018 An Efficient Volumetric Mesh Representation for Real-Time Scene Reconstruction Using Spatial Hashing
abstract
Mesh plays an indispensable role in dense realtime reconstruction essential in robotics. Efforts have been made to maintain flexible data structures for 3D data fusion, yet an efficient incremental framework specifically designed for online mesh storage and manipulation is missing. We propose a novel framework to compactly generate, update, and refine mesh for scene reconstruction upon a volumetric representation. Maintaining a spatial-hashed field of cubes, we distribute vertices with continuous value on discrete edges that supportO(1) vertex accessing and forbid memory redundancy. By introducing Hamming distance in mesh refinement, we further improve the mesh quality regarding the triangle type consistency with a low cost. Lock-based and lock-free operations were applied to avoid thread conflicts in GPU parallel computation. Experiments demonstrate that the mesh memory consumption is significantly reduced while the running speed is kept in the online reconstruction process.
Jieqi Shi, Weijie Tang, Xin Wang 0072, Hongbin Zha
ICRA5
2018 Consistent Correspondence of Cone-Beam CT Images Using Volume Functional Maps
Yungeng Zhang, Yuru Pei, Yuke Guo, Gengyu Ma, Tianmin Xu, Hongbin Zha
MICCAI (1)6
2018 Robust Matrix Factorization by Majorization Minimization
abstract
-norm based low rank matrix factorization in the presence of missing data and outliers remains a hot topic in computer vision. Due to non-convexity and non-smoothness, all the existing methods either lack scalability or robustness, or have no theoretical guarantee on convergence. In this paper, we apply the Majorization Minimization technique to solve this problem. At each iteration, we upper bound the original function with a strongly convex surrogate. By minimizing the surrogate and updating the iterates accordingly, the objective function has sufficient decrease, which is stronger than just being non-increasing that other methods could offer. As a consequence, without extra assumptions, we prove that any limit point of the iterates is a stationary point of the objective function. In comparison, other methods either do not have such a convergence guarantee or require extra critical assumptions. Extensive experiments on both synthetic and real data sets testify to the effectiveness of our algorithm. The speed of our method is also highly competitive.
Zhouchen Lin, Hongbin Zha
IEEE Trans. Pattern Anal. Mach. Intell.3
2018 Scene-Adaptive Off-Road Detection Using a Monocular Camera
abstract
This paper studies vision-based road detection for a robot's path following in off-road environments. We define the problem as detecting the region in front of the robot that is mechanically traversable (i.e., mechanical traversability), that is apt to be chosen by a human to drive through (i.e., human selection), and that extends for a distance to show the road's direction, shape, or even network of the intersection ahead (i.e., far-field capability). An algorithm framework is designed that contains two parts: inference and learning. In inference, the problem is formulated as a consecutive road type classification and road region segmentation to address the diversity of terrain surfaces. In model learning, the robot is first driven by a human being, with image samples on the track of the robot being collected that meet the prerequisites of both mechanical traversability and human selection. Evaluation measures are defined to examine the three requirements of mechanical traversability, human selection, and far-field capability. The performances of the above aspects are demonstrated on a data set using LiDAR, track and manual references, which will be released together with this publication.
Jilin Mei, Huijing Zhao, Hongbin Zha
IEEE Trans. Intell. Transp. Syst.4
2018 Spatially Consistent Supervoxel Correspondences of Cone-Beam Computed Tomography Images
abstract
Establishing dense correspondences of cone-beam computed tomography (CBCT) images is a crucial step for the attribute transfer and morphological variation assessment in clinical orthodontics. In this paper, a novel method, unsupervised spatially consistent clustering forest, is proposed to tackle the challenges for automatic supervoxel-wise correspondences of CBCT images. A complexity analysis of the proposed method with respect to the clustering hypotheses is provided with a data-dependent learning guarantee. The learning bound considers both the sequential tree traversals determined by questions stored in branch nodes and the clustering compactness of leaf nodes. A novel tree-pruning algorithm, guided by the learning bound, is also proposed to remove locally inconsistent leaf nodes. The resulting forest yields spatially consistent affinity estimations, thanks to the pruning penalizing trees with inconsistent leaf assignments and the combinational contextual feature channels used to learn the forest. A forest-based metric is utilized to derive the pairwise affinities and dense correspondences of CBCT images. The proposed method has been applied to the label propagation of clinically captured CBCT images. In the experiments, the method outperforms variants of both supervised and unsupervised forest-based methods and state-of-the-art label-propagation methods, achieving the mean dice similarity coefficients of 0.92, 0.89, 0.94, and 0.93 for the mandible, the maxilla, the zygoma arch, and the teeth data, respectively.
Yuru Pei, Yunai Yi, Gengyu Ma, Tae-Kyun Kim 0001, Yuke Guo, Tianmin Xu, Hongbin Zha
IEEE Trans. Medical Imaging7
2017 A Unified Convex Surrogate for the Schatten-p Norm
abstract
The Schatten-p norm (0 < p < 1) has been widely used to replace the nuclear norm for better approximating the rank function. However, existing methods are either 1) not scalable for large scale problems due to relying on singular value decomposition (SVD) in every iteration, or 2) specific to some p values, e.g., 1/2, and 2/3. In this paper, we show that for any p, p1, and p2 > 0 satisfying 1/p = 1/p1 + 1/p2, there is an equivalence between the Schatten-p norm of one matrix and the Schatten-p1 and the Schatten-p2 norms of its two factor matrices. We further extend the equivalence to multiple factor matrices and show that all the factor norms can be convex and smooth for any p > 0. In contrast, the original Schatten-p norm for 0 < p < 1 is non-convex and non-smooth. As an example we conduct experiments on matrix completion. To utilize the convexity of the factor matrix norms, we adopt the accelerated proximal alternating linearized minimization algorithm and establish its sequence convergence. Experiments on both synthetic and real datasets exhibit its superior performance over the state-of-the-art methods. Its speed is also highly competitive.
Zhouchen Lin, Hongbin Zha
AAAI3
2017 Improving Children's Gaze Prediction via Separate Facial Areas and Attention Shift Cue
abstract
To predict and assess visual attention, saliencybased visual attention modeling is a popular approach. However, state-of-the-art models are developed for adults, in which children are not considered. Additionally, these models consider neither social cues like face, nor attention learning cues. The face is a vital part of visual attention. Psychological studies reveal that sub-facial areas are different in visual attention. Some models highlight faces in social scenes, but sub-facial areas are not taken into account. Attention learning reveals internal processing of visual attention. By learning how the cognitive system deals with visual stimuli, it is possible to predict visual attention behavior. In this paper, we propose a multilevel visual attention model to predict fixations of children when watching a talking face. Based on traditional saliency maps, the proposed model includes both separate facial areas and attention shift cue. An eye-tracking experiment is conducted to evaluate the model. Results show that the proposed model significantly outperforms conventional models in talking face scenes.
Songjiang Li, Wen Cui, Jinshi Cui, Hongbin Zha
FG6
2017 Specialized gaze estimation for children by convolutional neural network and domain adaptation
abstract
Children's social gaze behavior modeling and evaluation has obtained increasing attentions in various research areas. In psychology research, eye gaze behavior is very important to developmental disorders diagnosis and assessment. In robotics area, gaze interaction between children and robots also draws more and more attention. However, there exists no specific gaze estimator for children in social interaction context. Current approaches usually use models trained with adults' data to estimate children's gaze. Since gaze behaviors and eye appearances of children are different from those of adults, the current approaches, especially those with free-calibration assumptions which are utilized in usual human-robot interaction systems, will result in big errors. Note that children data is difficult to collect and label, so directly learning from children data is hard to achieve. We propose a new system to solve this problem, which combines a CNN feature extractor trained from adult data and a domain adaptation unit using geodesic flow kernel to adapt the source domain (adults) classifier to the target domain (children). Our system performs well in children's gaze estimation.
Wen Cui, Jinshi Cui, Hongbin Zha
ICIP3
2017 Ego-centric traffic behavior understanding through multi-level vehicle trajectory analysis
abstract
This study proposes a multi-level trajectory analysis method for modeling traffic behavior from an ego-centric view, where on-road vehicle trajectories are collected based on the authors' previous studies of an on-board system consisting of multiple 2D lidar sensors. From an input set of trajectories, a set of hot regions (topics) that trajectory points most frequently present are first discovered using a sticky HDP-HMM; then, the major paths of the trajectories' transitions across different hot regions are extracted by recursively mining frequent subsequences of topics; and finally, paths are modeled using a hierarchical hidden Markov model (HHMM), where the intra-path dynamics is represented using an HMM, in which each state corresponds to a hot region, while the inter-path transition is assumed to be Markovian. The model could be used for behavior prediction, i.e. whenever a vehicle is detected in a scene, predicting which route it will probably follow and how its trajectory will probably develop over time, which is essential to interpreting the potential risks for longer time horizons. Experiments are conducted using a large set of vehicle trajectories collected from motorways in Beijing, and promising results are presented.
Donghao Xu, Huijing Zhao, Jinshi Cui, Hongbin Zha, Franck Guillemard, Stéphane Géronimi, François Aioun
ICRA5
2017 Factorization for projective and metric reconstruction via truncated nuclear norm
abstract
Structure from motion (SfM) is a crucial and widely studied problem in computer vision. Recently, the factorization framework for SfM was formulated as a low rank approximation problem: the rank of rescaled measurement matrix is always smaller than four. Since the rank function is non-convex, a common practice is to replace with its convex surrogate, i.e., the nuclear norm. However, nuclear norm sometimes gets unsatisfactory results. In this paper, we apply the recently proposed truncated nuclear norm to handle the factorization framework in a non-convex way, which heavily penalizes the singular values beyond the desired rank. We further introduce weighted ℓ1-norm to handle missing data and outliers uniformly. Based on truncated nuclear norm, we propose two factorization models for projective reconstruction and metric reconstruction, respectively. We also proposed an extremely efficient algorithm to tackle one of the optimization sub-problems. Extensive experiments on synthetic and real datasets verify the effectiveness of our method for projective and metric reconstructions. Our method achieves higher accuracy in 3D reconstruction and is more robust to missing data and outliers.
Zhouchen Lin, Tong Lin 0002, Hongbin Zha
IJCNN5
2017 On-road vehicle tracking using part-based particle filter
abstract
In this paper, we propose a part-based particle filter for on-road vehicle tracking. The proposed model takes part-based strategies into account in a particle filter. By introducing a hidden state vehicle center position, vehicle parts particles can be updated efficiently as a whole sharing same motion. With a pre-trained appearance and geometric model, tracker can distinguish parts with rich information from invalid parts to make a more precise prediction. Meanwhile some priori knowledge about the moving pattern of vehicles in well-structured on-road environment is learned, and can be used in the inference of measurement model and motion model to improve tracking performance and efficiency. Experiments were conducted with real data collected in Beijing to examine the performance in different situations on both the advantages and challenges. The Beijing highway dataset for on-road vehicle tracking will be opened to the society. We compare our method with the state-of-the-art approaches. Result demonstrate that the proposed algorithm are able to handle occlusion and aspect ratio change in on-road vehicle tracking problem.
Yongkun Fang, Chao Wang 0060, Huijing Zhao, Hongbin Zha
IROS4
2017 Mixed Metric Random Forest for Dense Correspondence of Cone-Beam Computed Tomography Images
Yuru Pei, Yunai Yi, Gengyu Ma, Yuke Guo, Tianmin Xu, Hongbin Zha
MICCAI (1)7
2017 Joint Latent Subspace Learning and Regression for Cross-Modal Retrieval
abstract
Cross-modal retrieval has received much attention in recent years. It is a commonly used method to project multi-modality data into a common subspace and then retrieve. However, nearly all existing methods directly adopt the space defined by the binary class label information without learning as the shared subspace for regression. In this paper, we first adopt the spectral regression method to learn the optimal latent space shared by data of all modalities based on the orthogonal constraints. Then we construct a graph model to project the multi-modality data into the latent space. Finally, we combine these two processes together to jointly learn the latent space and regress. We conduct extensive experiments on multiple benchmark datasets and our proposed method outperforms the state-of-the-art approaches.
Jianlong Wu, Zhouchen Lin, Hongbin Zha
SIGIR3
2017 Locality-constrained linear coding based bi-layer model for multi-view facial expression recognition
Jianlong Wu, Zhouchen Lin, Wenming Zheng, Hongbin Zha
Neurocomputing4
2017 The Shape Interaction Matrix-Based Affine Invariant Mismatch Removal for Partial-Duplicate Image Search
abstract
Mismatch removal is a key step in many computer vision problems. In this paper, we handle the mismatch removal problem by adopting shape interaction matrix (SIM). Given the homogeneous coordinates of the two corresponding point sets, we first compute the SIMs of the two point sets. Then, we detect the mismatches by picking out the most different entries between the two SIMs. Even under strong affine transformations, outliers, noises, and burstiness, our method can still work well. Actually, this paper is the first non-iterative mismatch removal method that achieves affine invariance. Extensive results on synthetic 2D points matching data sets and real image matching data sets verify the effectiveness, efficiency, and robustness of our method in removing mismatches. Moreover, when applied to partial-duplicate image search, our method reaches higher retrieval precisions with shorter time cost compared with the state-of-the-art geometric verification methods.
Zhouchen Lin, Hongbin Zha
IEEE Trans. Image Process.3
2016 Relaxed Majorization-Minimization for Non-Smooth and Non-Convex Optimization
abstract
We propose a new majorization-minimization (MM) method for non-smooth and non-convex programs, which is general enough to include the existing MM methods. Besides the local majorization condition, we only require that the difference between the directional derivatives of the objective function and its surrogate function vanishes when the number of iterations approaches infinity, which is a very weak condition. So our method can use a surrogate function that directly approximates the non-smooth objective function. In comparison, all the existing MM methods construct the surrogate function by approximating the smooth component of the objective function. We apply our relaxed MM methods to the robust matrix factorization (RMF) problem with different regularizations, where our locally majorant algorithm shows advantages over the state-of-the-art approaches for RMF. This is the first algorithm for RMF ensuring, without extra assumptions, that any limit point of the iterates is a stationary point.
Zhouchen Lin, Hongbin Zha
AAAI4
2016 Edge Enhanced Direct Visual Odometry
Xin Wang 0072, Mingcai Zhou, Renju Li, Hongbin Zha
BMVC5
2016 Camera Calibration from Periodic Motion of a Pedestrian
abstract
Camera calibration directly from image sequences of a pedestrian without using any calibration object is a really challenging task and should be well solved in computer vision, especially in visual surveillance. In this paper, we propose a novel camera calibration method based on recovering the three orthogonal vanishing points (TOVPs), just using an image sequence of a pedestrian walking in a straight line, without any assumption of scenes or motions, e.g., control points with known 3D coordinates, parallel or perpendicular lines, non-natural or pre-designed special human motions, as often necessary in previous methods. The traces of shoes of a pedestrian carry more rich and easily detectable metric information than all other body parts in the periodic motion of a pedestrian, but such information is usually overlooked by previous work. In this paper, we employ the images of the toes of the shoes on the ground plane to determine the vanishing point corresponding to the walking direction, and then utilize harmonic conjugate properties in projective geometry to recover the vanishing point corresponding to the perpendicular direction of the walking direction in the horizontal plane and the vanishing point corresponding to the vertical direction. After recovering all of the TOVPs, the intrinsic and extrinsic parameters of the camera can be determined. Experiments on various scenes and viewing angles prove the feasibility and accuracy of the proposed method.
Shiyao Huang, Xianghua Ying, Jiangpeng Rong, Zeyu Shang, Hongbin Zha
CVPR5
2016 Volumetric reconstruction of craniofacial structures from 2D lateral cephalograms by regression forest
abstract
The 3D reconstruction is an essential step to measure the craniofacial morphological changes from the historical growth database with only 2D cephalograms. In this paper, we propose a novel regression-forest-based method to estimate the volumetric intensity images from a lateral cephalogram. The regression forest can produce a prediction of the volumetric craniofacial structure as a mixture of Gaussian by weighted aggregating the distributions from trees. The dense anatomical structure can be reconstructed with no time-consuming digitally-reconstructed-radiographs (DRR) in the online testing process. The experiments demonstrate the proposed method can reconstruct volumetric intensity images from the lateral cephalograms effectively.
Yuru Pei, Fanfan Dai, Tianmin Xu, Hongbin Zha, Gengyu Ma
ICIP4
2016 Anatomical structure similarity estimation by random forest
abstract
The morphological similarity of anatomical structures is essential to the study of the species evolution. In this paper, we investigate the unsupervised shape similarity analysis by a random-forest-based metric. The dense continuous deformation fields are employed as the shape descriptors. The forest is built when given the unlabeled deformation fields, where the leaves can be seen as an optimal clustering of the data set. The salient region is defined based on the dominant feature channels determined by the forests. The pairwise shape distance is computed efficiently with just binary comparisons stored in tree branches. We have applied our method to several skeletal data sets, including the skulls, teeth, radii, and metatarsals. Our experiments demonstrate the proposed method can handle the taxonomic classification effectively.
Yuru Pei, Lei Kou, Hongbin Zha
ICIP3
2016 Multi-view common space learning for emotion recognition in the wild
abstract
It is a very challenging task to recognize emotion in the wild. Recently, combining information from various views or modalities has attracted more attention. Cross modality features and features extracted by different methods are regarded as multi-view information of the sample. In this paper, we propose a method to analyse multi-view features of emotion samples and automatically recognize the expression as part of the fourth Emotion Recognition in the Wild Challenge (EmotiW 2016). In our method, we first extract multi-view features such as BoF, CNN, LBP-TOP and audio features for each expression sample. Then we learn the corresponding projection matrices to map multi-view features into a common subspace. In the meantime, we impose l_2,1-norm penalties on projection matrices for feature selection. We apply both this method and PLSR to emotion recognition. We conduct experiments on both AFEW and HAPPEI datasets, and achieve superior performance. The best recognition accuracy of our method is 0.5531 on the AFEW dataset for video based emotion recognition in the wild. The minimum RMSE for group happiness intensity recognition is 0.9525 on HAPPEI dataset. Both of them are much better than that of the challenge baseline.
Jianlong Wu, Zhouchen Lin, Hongbin Zha
ICMI3
2016 Perceptual enhancement for stereoscopic videos based on horopter consistency
abstract
Audience discomfort, such as eye strain and dizziness, is one of the urgent issues that virtual reality and 3D movie technologies should tackle. Except for inappropriate horizontal and vertical disparity, one major problem is that people's binocular vergence and focal length in the cinema remain inconsistent from normal visual habits. Psychologists discovered the horopter and Panum's fusional area to describe zero-disparity points projected on the retinas based on accommodation-convergence consistency. In this paper, inspired by these concepts, we propose a stereoscopic effect correction system for perceptual enhancement according to fixated region and scene information. As a preprocessing step, tracking and stereo matching algorithms are implemented to prepare cues for further transformation in 3D space. Then in order to accomplish certain visual effects, we describe a geometric framework for disparity refinement and image warping based on parameter adjustment of the virtual stereoscopic rig. For evaluation, subjective experiments have been conducted to prove the effectiveness of our method. Therefore, our work provides a possibility to improve the audience experience from a formerly underexplored perspective.
Zeyu Wang 0003, Xiaohan Jin, Renju Li, Hongbin Zha, Katsushi Ikeuchi
VRST5
2016 Evidential calibration of binary SVM classifiers
Philippe Xu, Franck Davoine, Hongbin Zha, Thierry Denoeux
Int. J. Approx. Reason.3
2016 Nonlinear Dimensionality Reduction by Local Orthogonality Preserving Alignment
Liwei Wang 0001, Hongbin Zha
J. Comput. Sci. Technol.5
2016 Probabilistic Inference for Occluded and Multiview On-road Vehicle Detection
abstract
Visual-based approaches have been extensively studied for on-road vehicle detection; however, it faces great challenges as the visual appearance of a vehicle may greatly change across different viewpoints and as a partial observation sometimes happens due to occlusions from infrastructure or scene dynamics and/or a limited camera vision field. This paper presents a visual-based on-road vehicle detection algorithm for a multilane traffic scene. A probabilistic inference framework based on part models is proposed to overcome the challenges from a multiview and partial observation. Geometric models are learned for each dominant viewpoint to describe the configuration of vehicle parts and their spatial relations in probabilistic representations. Viewpoint maps are generated based on the knowledge of the road structure and driving patterns, which provide a prediction of the viewpoints of a vehicle whenever it happens at a certain location. Extensive experiments are conducted using an onboard camera on multilane motor ways in Beijing. A large-scale data set that contains more than 30 000 labeled ground truths for both fully and partially observed vehicles in different viewpoints across various traffic density scenes is developed. The data set will be opened to the society together with this publication.
Chao Wang 0060, Yongkun Fang, Huijing Zhao, Chunzhao Guo, Seiichi Mita, Hongbin Zha
IEEE Trans. Intell. Transp. Syst.6
2015 Multi-modal Brain Image Registration Based on Subset Definition and Manifold-to-Manifold Distance
abstract
Image registration is an important procedure in multi-modal brain image processing. The main challenge is the variations of intensity distributions in different image modalities. The efficient SSD based method cannot handle this kind of variations. And other approaches based on modality independent descriptors and metrics are usually time-consuming. In this article, we propose a novel similarity metric based on manifold-to-manifold distance imposed on the subset of original images. We define a subset for a compact representation of the original image. Manifold learning technique is employed to reveal the intrinsic structure of the sampled data. Instead of comparing the images in the original feature space, we use the manifold-to-manifold distance to measure the difference. By minimizing the distance between the manifolds, we iteratively obtain the optimal registration of the original image pair. Experiment results show that our approach is effective to deal with the multi-modal image registration on the BrainWeb dataset.
Yuru Pei, Hongbin Zha
ICIG (2)3
2015 Radial lens distortion correction using cascaded one-parameter division model
abstract
Radial lens distortion is the most significant lens distortion in current cameras, and many models are proposed to describe it. Fitzgibbon presented the prestigious division distortion model with just a single parameter. Based on the fact that a line in 3D space may be projected onto a curve in the image plane due to radial distortion, Alemán et al. utilized the Hough transform with three parameters to automatically correct radial lens distortion, namely, one parameter coming from the division model, and the other two arising from the corrected image line. However, in some cases, especially for wide angle lenses, the corrected results are not very satisfactory. Someone may suggest that we can use the so-called extended division model with more than one parameter. Unfortunately, the problem will become very hard to solve, since the dimensions of the Hough parameter space become higher than three. In this paper, we propose a cascaded one-parameter division model to deal with the problem. In each stage the Hough parameter space is always three-dimensional. Enormous experiments on real images illustrate the ability of our method.
Xiang Mei, Jiangpeng Rong, Xianghua Ying, Shiyao Huang, Hongbin Zha
ICIP6
2015 Ellipse-specific fitting by relaxing the 3L constraints with semidefinite programming
abstract
This paper presents a new efficient method to increase the accuracy and the robustness of ellipse fitting, by utilizing the 3L algorithm and semidefinite programming (SDP). The novelty lies on the combination of relaxed geometric distance constraints and semidefinite programming framework. Due to the relaxed 3L constraints, the proposed approach provides high robustness in the presence of noise. The accuracy of the final solution is prominently increased even if the data suffer from strong occlusions or noises. The proposed method represents significant advantages in both accuracy and robustness. Experimental results and comparisons with state-of-the-art fitting methods demonstrate the improvements in ellipse fitting.
Jiangpeng Rong, Xiang Mei, Xianghua Ying, Shiyao Huang, Hongbin Zha
ICIP6
2015 Multiple Models Fusion for Emotion Recognition in the Wild
abstract
Emotion recognition in the wild is a very challenging task. In this paper, we propose a multiple models fusion method to automatically recognize the expression in the video clip as part of the third Emotion Recognition in the Wild Challenge (EmotiW 2015). In our method, we first extract dense SIFT, LBP-TOP and audio features from each video clip. For dense SIFT features, we use the bag of features (BoF) model with two different encoding methods (locality-constrained linear coding and group saliency based coding) to further represent it. During the classification process, we use partial least square regression to calculate the regression value of each model. By learning the optimal weight of each model based on the regression value, we fuse these models together. We conduct experiments on the given validation and test datasets, and achieve superior performance. The best recognition accuracy of our fusion method is 52.50% on the test dataset, which is 13.17% higher than the challenge baseline accuracy of 39.33%.
Jianlong Wu, Zhouchen Lin, Hongbin Zha
ICMI3
2015 Visual-based on-road vehicle detection: A transnational experiment and comparison
abstract
As a key technique in ADAS (Advanced Driving Assistant System) or autonomous driving systems, visual-based on-road vehicle detection has been studied widely, while it faces still great challenges, among which are the complexity, diversity and unpredictable changes of the real-world environments. In the authors' previous work, an algorithm was developed in a probabilistic inference framework with its focus on solving the multi-view and occlusion problems at multi-lane motor way scenes. In this research, we seek to answer the questions: how efficient is the system during a long-term operation across a large area of changed conditions? To this end, a large scale experiment is conducted, where three testing data sets are developed containing the samples of more than 30,000 on Beijing's ring roads, 800 on Nagoya's fast road, and 3,000 on Nagoya's downtown streets, and the performance of visual-based vehicle detection concerning the multi-view and occlusion problems across extensive regions and at transnational environments are studied. We present our preliminary findings in this paper, which leads to a more extensive study in future work.
Chao Wang 0060, Huijing Zhao, Chunzhao Guo, Seiichi Mita, Hongbin Zha
Intelligent Vehicles Symposium5
2015 Supervised learning via Euler's Elastica models
Tong Lin 0002, Hanlin Xue, Hongbin Zha
J. Mach. Learn. Res.5
2015 Learning to Detect Anomalies in Surveillance Video
abstract
Detecting anomalies in surveillance videos, that is, finding events or objects with low probability of occurrence, is a practical and challenging research topic in computer vision community. In this paper, we put forward a novel unsupervised learning framework for anomaly detection. At feature level, we propose a Sparse Semi-nonnegative Matrix Factorization (SSMF) to learn local patterns at each pixel, and a Histogram of Nonnegative Coefficients (HNC) can be constructed as local feature which is more expressive than previously used features like Histogram of Oriented Gradients (HOG). At model level, we learn a probability model which takes the spatial and temporal contextual information into consideration. Our framework is totally unsupervised requiring no human-labeled training data. With more expressive features and more complicated model, our framework can accurately detect and localize anomalies in surveillance video. We carried out extensive experiments on several benchmark video datasets for anomaly detection, and the results demonstrate the superiority of our framework to state-of-the-art approaches, validating the effectiveness of our framework.
Tan Xiao, Chao Zhang 0001, Hongbin Zha
IEEE Signal Process. Lett.3
2014 Anomaly Detection via Local Coordinate Factorization and Spatio-Temporal Pyramid
Tan Xiao, Chao Zhang 0001, Hongbin Zha, Fangyun Wei
ACCV (5)3
2014 Imposing Differential Constraints on Radial Distortion Correction
Xianghua Ying, Xiang Mei, Ganwen Wang, Jiangpeng Rong, Hongbin Zha
ACCV (1)6
2014 Similarity-Aware Patchwork Assembly for Depth Image Super-resolution
abstract
This paper describes a patchwork assembly algorithm for depth image super-resolution. An input low resolution depth image is disassembled into parts by matching similar regions on a set of high resolution training images, and a super-resolution image is then assembled using these corresponding matched counterparts. We convert the super resolution problem into a Markov Random Field (MRF) labeling problem, and propose a unified formulation embedding (1) the consistency between the resolution enhanced image and the original input, (2) the similarity of disassembled parts with the corresponding regions on training images, (3) the depth smoothness in local neighborhoods, (4) the additional geometric constraints from self-similar structures in the scene, and (5) the boundary coincidence between the resolution enhanced depth image and an optional aligned high resolution intensity image. Experimental results on both synthetic and real-world data demonstrate that the proposed algorithm is capable of recovering high quality depth images with X4 resolution enhancement along each coordinate direction, and that it outperforms state-of-the-arts [14] in both qualitative and quantitative evaluations.
Zhichao Lu, Rui Gan, Hongbin Zha
CVPR5
2014 Joint estimation of head pose and visual focus of attention
abstract
Head pose is an important indicator of a person's visual focus of attention (VFoA). A traditional way to recognize VFoA is to consider accurate head pose or gaze estimations. However, these estimations usually degrade drastically in middle or low resolution video data. In this paper, a joint estimation of head pose and VFoA is proposed to address this issue; both head pose and VFoA are iteratively refined until convergence. This approach is evaluated in a specific scenario involving children around a table playing together with toys. Datasets are acquired and annotated by psychologists in Peking university. The experimental results demonstrate the usefulness of the join estimation process to recognize visual focus of attention in middle resolution video sequences.
Yingning Huang, Dingrui Duan, Jinshi Cui, Franck Davoine, Hongbin Zha
ICIP6
2014 L1-norm global geometric consistency for partial-duplicate image retrieval
abstract
In all feature point based partial-duplicate image retrieval systems, false matching is a common issue. To tackle the problem, geometric contexts are widely applied to filter the inconsistent matches. This paper presents a novel method called ℓ1-norm global geometric consistency. We first form the squared distance matrices of all the matched feature points, which remain invariant under translation and rotation between partial-duplicated images. Then we find the scale difference by solving a one-variable ℓ1-norm error minimization problem, where the large sparse errors correspond to the locations of inconsistent matches. By adopting the Golden Section Search method the minimization problem can be solved efficiently. Extensive experimental results show that our method reaches higher precisions than state-of-the-art geometric verification methods in detecting inconsistent matches. Its speed is also highly competitive even when compared to local geometric consistency based methods.
Zhouchen Lin, Hongbin Zha
ICIP5
2014 Radial distortion correction from a single image of a planar calibration pattern using convex optimization
abstract
In Hartley-Kang's paper [7], they directly treated a planar calibration pattern as an image to construct an image pair together with a radial distorted image of the planar calibration pattern, and then proposed a very efficient method to determine the center of radial distortion by estimating the epipole in the radial distorted image. After determined the center of radial distortion, a least square method was utilized to recover the radial distortion function using the monotonicity constraints. In this paper, we present a convex optimization method to recover the radial distortion function using the same constraints as those required by Hartley-Kang's method, whereas our method can obtain better results of radial distortion correction. The experiments validate our approach.
Xianghua Ying, Xiang Mei, Ganwen Wang, Hongbin Zha
ICIP5
2014 Multiple View Based Building Modeling with Multi-box Grammar
abstract
This paper describes a multiple view based approach for building modeling via a novel multi-box grammar, which represents an occlusion relationship among the projections of a set of buildings sharing a common Manhattan World coordinate system. We formulate the building modeling problem as an energy minimization to combine the constraints from the multi-box grammar with (1) the semantic labeling information from appearance models, (2) the directional information w.r.t the vanishing points in each single view, and (3) the planar homography correspondence among multiple views. We further propose a two-step coarse-to-fine approach to achieve the optimal solution. First we employ super-pixels and a simplified edition of the grammar to reduce the searching space, and obtain an initial layout to accelerate the convergence speed. At the second stage, the scene model is refined to achieve pixel-level accuracy by minimizing the energy using Random Walk. Experiments on street view images demonstrate the capability of our method in reconstructing multiple buildings at different distances, and also the robustness in handling occlusion.
Ruiling Deng, Qiuliang Wang, Rui Gan, Hongbin Zha
ICPR5
2014 Sandwich Cut: An Algorithm for Temporally-Coherent Video Bilayer Segmentation
abstract
In this paper, we introduce the Sandwich Cut: an algorithm for temporally-coherent bilayer segmentation in monocular video sequences. Building upon the observation that temporally-coherent segmentation relys on the frame with its temporal neighbors, the key idea is looking ahead to a future frame for the current segmentation by constructing a three-frame graph like a sandwich. The sandwich, three-frame graph, combines the cues of previous segmentation, color, spatial and temporal color differences, where interlayer weights are estimated by temporal gradients. In each step of our sequential system the final result of the current frame and the temporary result of the future one are obtained, and the latter is used to estimate an adaptive coherence factor for the next sandwich cut. The presentation of the system is complemented with the quantitative evaluations and comparisons with the latest techniques on a public database. Moreover, a method for evaluating the temporal coherence of the results is introduced and tested. Compared with the classic optical flow based method, our system performances better with higher efficiency.
Songtao Pu, Hongbin Zha
ICPR2
2014 Low Rank Global Geometric Consistency for Partial-Duplicate Image Search
abstract
All existing feature point based partial-duplicate image retrieval systems are confronted with the false feature point matching problem. To resolve this issue, geometric contexts are widely used to verify the geometric consistency in order to remove false matches. However, most of the existing methods focus on local geometric contexts rather than global. Seeking global contexts has attracted a lot of attention in recent years. This paper introduces a novel global geometric consistency, based on the low rankness of squared distance matrices of feature points, to detect false matches. We cast the problem of detecting false matches as a problem of decomposing a squared distance matrix into a low rank matrix, which models the global geometric consistency, and a sparse matrix, which models the mismatched feature points. So we arrive at a model of Robust Principal Component Analysis. Our Low Rank Global Geometric Consistency (LRGGC) is simple yet effective and theoretically sound. Extensive experimental results show that our LRGGC is much more accurate than state of the art geometric verification methods in detecting false matches and is robust to all kinds of similarity transformation (scaling, rotation, and translation) and even slight change in 3D views. Its speed is also highly competitive even compared with local geometric consistency based methods.
Zhouchen Lin, Hongbin Zha
ICPR4
2014 The Perspective-3-Point Problem When Using a Planar Mirror
abstract
The Perspective-3-Point problem (P3P) is a classical and fundamental problem in computer vision. All possible solution sets for the P3P problem are from 1 to 4 solutions. In this paper, we propose a very simple way to reduce the ambiguity of numbers of possible solutions in P3P using a planar mirror. For three reference points, if they and their reflections in a planar mirror are both observed, we may obtain two P3P problems: One is from the three original reference points, and the other is from their reflections. A trivial procedure may be suggested: Solve for each of the two P3P problems, and then find the intersections of the two solution sets. Different from the trivial case, we propose an efficient method which employs the ratio relations of the unknowns in the two P3P problems. The ratio relations are arise from mirror reflection, and can be easily determined before solving the two P3P problems. With the ratio relations, a system of 6 equations with 3 unknowns can be determined. To solve the over-constraint problem, we utilize an efficient algorithm by finding all local minima of least-squares residual. Experiments validate our approach.
Xianghua Ying, Ganwen Wang, Xiang Mei, Hongbin Zha
ICPR5
2014 Calibration method for multiple 2D LIDARs system
abstract
Many robotic and mobile mapping systems have been developed using multiple 2D LIDARs (briefly multi-LIDAR system) to sense environment. In such systems, extrinsic calibration of all LIDARs is essential for making collaborative use of the data from different sensors. This research aims at developing a calibration method for multi-LIDAR systems at the general scene, such as an outdoor place or an underground parking-lot, without modification to environment by putting calibration targets. In this paper, the calibration method is proposed by aligning the 3D data of different LIDARs. They are concerned at two-levels: 1) reference calibration, i.e. finding the transformation from a reference LIDAR to the platform frame; 2) multi-LIDAR calibration, i.e. finding the LIDARs' relative geometries by referring to the reference one. The method is examined in calibrating the multiple 2D LIDARs on an intelligent vehicle platform POSS-V, where the data collected through a driving in an underground parking-lot are registered to find sensors' geometry. Calibration accuracy is examined by comparing with a CAD model of the scene, which was measured by using a total station.
Mengwen He, Huijing Zhao, Jinshi Cui, Hongbin Zha
ICRA4
2014 On-road vehicle detection through part model learning and probabilistic inference
abstract
Visual based approach has been studied extensively for on-road vehicle detection, while it faces great challenges, as visual appearance of a vehicle may change greatly across different viewpoints, and partial observation happens sometime due to occlusions from infrastructure or scene dynamics, and/or limited camera vision field. Inspired by the works on part-based detection, this research proposes a probabilistic framework for on-road vehicle detection, where focus is cast on vehicle pose inference on the set of part instances by addressing the issues of partial observation and varying viewpoints. To this end, geometric models describing the configuration of vehicle parts as well as their spatial relations in probabilistic representations are learned for each dominant viewpoint, and viewpoint maps are generated on each typical road structure, which provide probabilistic prediction to the viewpoints of a vehicle at each location at ego frame. Experiments have been conducted using a data set that was developed in the authors' previous work on the ring roads in Beijing. Viewpoint-discriminative part appearance models (VDPAM) and viewpoint-discriminative part-based geometric models (VDPGM) are learned on the image samples of the data set, and the road structure-based probabilistic viewpoint maps (RSPVM) are generated by taking the statistics of the Lidar-based vehicle detection results. On-road vehicle detection is examined using an on-road video stream that has been labelled with ground truth. Experimental results are presented and efficiency on detecting the partially observed vehicles on varying viewpoints is demonstrated.
Chao Wang 0060, Huijing Zhao, Chunzhao Guo, Seiichi Mita, Hongbin Zha
IROS5
2014 Monocular visual localization using road structural features
abstract
Precise localization is an essential issue for autonomous driving applications, where GPS-based systems are challenged to meet requirements such as lane-level accuracy. This paper introduces a new visual-based localization approach in dynamic traffic environments, focusing on and exploiting properties of structured roads like straight roads or intersections. Such environments show several line segments on lane markings, curbs, poles, building edges, etc., which demonstrate the road's longitude, latitude and vertical directions. Based on this observation, we define a Road Structural Feature (RSF) as sets of segments along three perpendicular axes together with feature points. At each video frame, the proper road structure (or multiple road structures in case of an intersection) is predicted based on the geometric information given by a 2D map. The RSF is then detected from line segments and points extracted from the image, and used to estimate the pose of the vehicle. Experiments are conducted using video streams collected on major roads in downtown Beijing, which are structured and with intense dynamic traffic. GPS/IMU data have been collected and synchronized with the video streams as a reference in validation. The results show good performance compared with that of a more traditional visual odometry method. Future work will be addressed on using visual approach to improve GPS localization accuracy.
Huijing Zhao, Franck Davoine, Jinshi Cui, Hongbin Zha
Intelligent Vehicles Symposium5
2014 Trinary-Projection Trees for Approximate Nearest Neighbor Search
abstract
We address the problem of approximate nearest neighbor (ANN) search for visual descriptor indexing. Most spatial partition trees, such as KD trees, VP trees, and so on, follow the hierarchical binary space partitioning framework. The key effort is to design different partition functions (hyperplane or hypersphere) to divide the points so that 1) the data points can be well grouped to support effective NN candidate location and 2) the partition functions can be quickly evaluated to support efficient NN candidate location. We design a trinary-projection direction-based partition function. The trinary-projection direction is defined as a combination of a few coordinate axes with the weights being 1 or -1. We pursue the projection direction using the widely adopted maximum variance criterion to guarantee good space partitioning and find fewer coordinate axes to guarantee efficient partition function evaluation. We present a coordinate-wise enumeration algorithm to find the principal trinary-projection direction. In addition, we provide an extension using multiple randomized trees for improved performance. We justify our approach on large-scale local patch indexing and similar image search.
Jingdong Wang 0001, Naiyan Wang, You Jia, Jian Li 0015, Hongbin Zha, Xian-Sheng Hua 0001
IEEE Trans. Pattern Anal. Mach. Intell.6
2013 Anatomical Structure Sketcher for Cephalograms by Bimodal Deep Learning
abstract
Lateral cephalogram X-ray (LCX) images are essential to provide patientspecific morphological information of anatomical structures. The automatic annotation of anatomical structures in cephalograms has been performed in the biomedical engineering for nearly twenty years. Most systems only handle a portion of salient craniofacial landmark set [1, 2, 3]. Although model-based methods can produce a full set of markers [5, 7], the pattern fitting can fail to converge in blurry images. It is challenging to annotate LCX images with high fidelity. In this work, we propose a novel cephalogram sketcher system as shown in Fig. 1 for the automatic anatomical-structure annotation, especially for the blemished images due to structure overlappings and devicespecific distortions during projection. Firstly, we introduce an hierarchical extension of a pictorial model to detect anatomical structures. Secondly, the bimodal deep Boltzmann machine (DBM) is employed to sketch the structure contours. Specifically, the contour sketcher takes advantages of the path in the DBM to extract the contour definitions from the patch textures by alternating Gibbs sampling. Given a cephalogram I, the structure definition S, and the parameters Θ = (Θq,Θr) with respect to the intraand inter-layer correlations, the posterior probability distribution according to the Bayes rule is defined as P(S|I,Θ) ∝ P(I|S,Θ)P(S|Θ), where P(S|Θ) is a shape prior distribution. P(I|S,Θ) is the image likelihood given the hierarchical architecture and the model parameters. The likelihood can be factorized as a product of likelihoods of local structures.
Yuru Pei, Hongbin Zha, Tianmin Xu
BMVC3
2013 Supervised Kernel Descriptors for Visual Recognition
abstract
In visual recognition tasks, the design of low level image feature representation is fundamental. The advent of local patch features from pixel attributes such as SIFT and LBP, has precipitated dramatic progresses. Recently, a kernel view of these features, called kernel descriptors (KDES), generalizes the feature design in an unsupervised fashion and yields impressive results. In this paper, we present a supervised framework to embed the image level label information into the design of patch level kernel descriptors, which we call supervised kernel descriptors (SKDES). Specifically, we adopt the broadly applied bag-of-words (BOW) image classification pipeline and a large margin criterion to learn the low-level patch representation, which makes the patch features much more compact and achieve better discriminative ability than KDES. With this method, we achieve competitive results over several public datasets comparing with state-of-the-art methods.
Peng Wang 0001, Jingdong Wang 0001, Weiwei Xu 0003, Hongbin Zha, Shipeng Li 0001
CVPR5
2013 Unsupervised Random Forest Manifold Alignment for Lipreading
abstract
Lip reading from visual channels remains a challenging topic considering the various speaking characteristics. In this paper, we address an efficient lip reading approach by investigating the unsupervised random forest manifold alignment (RFMA). The density random forest is employed to estimate affinity of patch trajectories in speaking facial videos. We propose novel criteria for node splitting to avoid the rank-deficiency in learning density forests. By virtue of the hierarchical structure of random forests, the trajectory affinities are measured efficiently, which are used to find embeddings of the speaking video clips by a graph-based algorithm. Lip reading is formulated as matching between manifolds of query and reference video clips. We employ the manifold alignment technique for matching, where the L∞-norm-based manifold-to-manifold distance is proposed to find the matching pairs. We apply this random forest manifold alignment technique to various video data sets captured by consumer cameras. The experiments demonstrate that lip reading can be performed effectively, and outperform state-of-the-arts.
Yuru Pei, Tae-Kyun Kim 0001, Hongbin Zha
ICCV3
2013 Streaming video object segmentation with the adaptive coherence factor
abstract
In this paper, we present a motion-adaptive algorithm for streaming object segmentation in monocular video sequences. To segment each frame, we fuse the cues from color, spatiotemporal contrast, and the bilayer labels on the previous frame in a two-frame graph. In the graph the coherence factor, the weight of the temporal smoothing term, is online estimated by the current frame, the previous frames and their segmentation results. The algorithm builds upon the observation that the amount of the object movement is approximately linear related to the summation of the temporal contrasts between adjacent frames. With the adaptive coherence factor we can improve the temporal coherence of the results when the object movement is changed. Finally each frame is segmented by binary graph cut. Experimental results show the effects of the adaptive coherence factor and validate the effectiveness of our proposed algorithm.
Songtao Pu, Hongbin Zha
ICIP2
2013 Canonicalized central absolute moment for edge-based color constancy
abstract
In the recent paper by Weijer et al. [9], the authors proposed the Grey-Edge hypothesis, which assumes that the average edge difference in the scene is achromatic. The Minkowski norm of color derivatives is used to approximate the light source color. In this paper, we point out that the Minkowski norm of these derivatives can actually be interpreted by the raw moment of the distribution of these derivatives. Furthermore, we discovered that the central moment of the distribution can also be utilized to estimate the light source color, and comparable results may be obtained.
Xianghua Ying, Lulu Hou, Yongbo Hou, Hongbin Zha
ICIP5
2013 Pairwise LIDAR calibration using multi-type 3D geometric features in natural scene
abstract
It has become a well-known technology that 3D measurement of a large environment could be achieved by using a number of 2D LIDARs on a mobile platform. In such a system, calibration is essential for making collaborative use of different LIDAR data, while existing methods usually require modifications to the environments, such as putting calibration targets, or rely on special facilities, which is labor intensive and put many restrictions to potential applications. This research aims at developing a calibration method for multiple 2D LIDAR sensing systems, which could be conducted in a general outdoor environment using the features of a nature scene. Special focus is cast on solving the noisy sensing in a complex environment and the occlusions caused by largely different sensor viewpoints. A multi-type geometric feature based calibration algorithm is proposed, which extracts the features such as points, lines, planes and quadrics from the 3D points of each LIDAR sensing. Transformation parameters from each sensor to the frame on a moving platform is estimated by matching the multi-type features. Experiments are conducted using the data sets of an intelligent vehicle platform (POSS-V) through a driving in the campus of Peking University. Results of calibrating two LIDAR sensors with largely different viewpoints are presented, and the accuracy and robustness concerning noisy feature extractions are examined intensively.
Mengwen He, Huijing Zhao, Franck Davoine, Jinshi Cui, Hongbin Zha
IROS5
2013 Lane change trajectory prediction by using recorded human driving data
abstract
Being able to predict the trajectory of a human driver's potential lane change behavior in urban high way scenario is crucial for lane change risk assessment task. A good prediction of the driver's lane change trajectory makes it possible to evaluate the risk and warn the driver beforehand. Rather than generating such a trajectory only using a mathematical model, this paper develops a lane change trajectory prediction approach based on real human driving data stored in a database. In real-time, the system generates parametric trajectories by interpolating k human lane change trajectory instances from the pre-collected database that are similar to the current driving situation. In order to build this real lane change database, a human lane change data collection vehicle platform is developed. Extensive experiments have been carried out in urban highway environments to build a significant database with more than 200 lane changes. Real results show that this approach produces lane change trajectories that are quite similar to real ones which makes it suitable to predict humanlike lane change maneuvers.
Huijing Zhao, Philippe Bonnifait, Hongbin Zha
Intelligent Vehicles Symposium4
2013 Structure-Sensitive Superpixels via Geodesic Distance
Peng Wang 0001, Rui Gan, Jingdong Wang 0001, Hongbin Zha
Int. J. Comput. Vis.5
2013 Laser-based tracking of multiple interacting pedestrians via on-line learning
Xuan Song 0001, Jinshi Cui, Huijing Zhao, Hongbin Zha, Ryosuke Shibasaki
Neurocomputing4
2013 Self-Calibration of Catadioptric Camera with Two Planar Mirrors from Silhouettes
abstract
If an object is interreflected between two planar mirrors, we may take an image containing both the object and its multiple reflections, i.e., simultaneously imaging multiple views of an object by a single pinhole camera. This paper emphasizes the problem of recovering both the intrinsic and extrinsic parameters of the camera using multiple silhouettes from one single image. View pairs among views in a single image can be divided into two kinds by the relationship between the two views in the pair: reflected by some mirror (real or virtual) and in a circular motion. Epipoles in the first kind of pairs can be easily determined from intersections of common tangent lines of silhouettes. Based on the projective properties of these epipoles, efficient methods are proposed to recover both the imaged circular points and the included angle between two mirrors. Epipoles in the second kind of pairs can be recovered simultaneously with the projection of intersection line between two mirrors by solving a simple 1D optimization problem using the consistency constraint of epipolar tangent lines. Fundamental matrices among views in a single image are all recovered. Using the estimated intrinsic and extrinsic parameters of the camera, a euclidean reconstruction can be obtained. Experiments validate the proposed approach.
Xianghua Ying, Yongbo Hou, Sheng Guan, Hongbin Zha
IEEE Trans. Pattern Anal. Mach. Intell.6
2013 A fully online and unsupervised system for large and high-density area surveillance: Tracking, semantic scene learning and abnormality detection
abstract
For reasons of public security, an intelligent surveillance system that can cover a large, crowded public area has become an urgent need. In this article, we propose a novel laser-based system that can simultaneously perform tracking, semantic scene learning, and abnormality detection in a fully online and unsupervised way. Furthermore, these three tasks cooperate with each other in one framework to improve their respective performances. The proposed system has the following key advantages over previous ones: (1) It can cover quite a large area (more than 60×35m), and simultaneously perform robust tracking, semantic scene learning, and abnormality detection in a high-density situation. (2) The overall system can vary with time, incrementally learn the structure of the scene, and perform fully online abnormal activity detection and tracking. This feature makes our system suitable for real-time applications. (3) The surveillance tasks are carried out in a fully unsupervised manner, so that there is no need for manual labeling and the construction of huge training datasets. We successfully apply the proposed system to the JR subway station in Tokyo, and demonstrate that it can cover an area of 60×35m, robustly track more than 150 targets at the same time, and simultaneously perform online semantic scene learning and abnormality detection with no human intervention.
Xuan Song 0001, Xiaowei Shao, Quanshi Zhang, Ryosuke Shibasaki, Huijing Zhao, Jinshi Cui, Hongbin Zha
ACM Trans. Intell. Syst. Technol.7
2013 An online system for multiple interacting targets tracking: Fusion of laser and vision, tracking and learning
abstract
Multitarget tracking becomes significantly more challenging when the targets are in close proximity or frequently interact with each other. This article presents a promising online system to deal with these problems. The novelty of this system is that laser and vision are integrated with tracking and online learning to complement each other in one framework: when the targets do not interact with each other, the laser-based independent trackers are employed and the visual information is extracted simultaneously to train some classifiers online for “possible interacting targets”. When the targets are in close proximity, the classifiers learned online are used alongside visual information to assist in tracking. Therefore, this mode of cooperation not only deals with various tough problems encountered in tracking, but also ensures that the entire process can be completely online and automatic. Experimental results demonstrate that laser and vision fully display their respective advantages in our system, and it is easy for us to obtain a good trade-off between tracking accuracy and the time-cost factor.
Xuan Song 0001, Huijing Zhao, Jinshi Cui, Xiaowei Shao, Ryosuke Shibasaki, Hongbin Zha
ACM Trans. Intell. Syst. Technol.6
2013 Tracking Generic Human Motion via Fusion of Low- and High-Dimensional Approaches
abstract
Tracking generic human motion is highly challenging due to its high-dimensional state space and the various motion types involved. In order to deal with these challenges, a fusion formulation which integrates low- and high-dimensional tracking approaches into one framework is proposed. The low-dimensional approach successfully overcomes the high-dimensional problem of tracking the motions with available training data by learning motion models, but it only works with specific motion types. On the other hand, although the high-dimensional approach may recover the motions without learned models by sampling directly in the pose space, it lacks robustness and efficiency. Within the framework, the two parallel approaches, low- and high-dimensional, are fused via a probabilistic approach at each time step. This probabilistic fusion approach ensures that the overall performance of the system is improved by concentrating on the respective advantages of the two approaches and resolving their weak points. The experimental results, after qualitative and quantitative comparisons, demonstrate the effectiveness of the proposed approach in tracking generic human motion.
Jinshi Cui, Ye Liu 0002, Yuandong Xu, Huijing Zhao, Hongbin Zha
IEEE Trans. Syst. Man Cybern. Syst.5
2012 A Bayesian Approach to Uncertainty-Based Depth Map Super Resolution
Rui Gan, Hongbin Zha, Long Wang 0001
ACCV (4)4
2012 Extending Fitts' law to account for the effects of movement direction on 2d pointing
abstract
Fitts' law is the most widely applied model in the field of HCI. However, this model and its existing extensions are still limited for 2D pointing task especially when the effects of movement direction (Θ) remain in the task. In this paper, we employ the concept of projection to account for the effects of target width (W) and height (H) on movement time so that we seamlessly integrate the four factors, i.e. Θ, amplitude (A), W and H, into the new extension of Fitts' law, which can uncover not only the periodicity of the asymmetrical impacts of W and H with the variation of Θ but also their interrelation. Carrying out two experiments, we verify that the vertical projection of W and the horizontal projection of H in the line of movement direction can be viewed as the determinants of movement time. Finally, we offer recommendations for 2D pointing experiments and discuss the implications for interface designs.
Xinyong Zhang, Hongbin Zha, Wenxin Feng 0001
CHI2
2012 A Game-Theoretical Approach to Image Segmentation
Rui Gan, Hongbin Zha, Long Wang 0001
CVM4
2012 Salient object detection for searched web images via global saliency
abstract
In this paper, we deal with the problem of detecting the existence and the location of salient objects for thumbnail images on which most search engines usually perform visual analysis in order to handle web-scale images. Different from previous techniques, such as sliding window-based or segmentation-based schemes for detecting salient objects, we propose to use a learning approach, random forest in our solution. Our algorithm exploits global features from multiple saliency indicators to directly predict the existence and the position of the salient object. To validate our algorithm, we constructed a large image database collected from Bing image search, that contains hundreds of thousands of manually labeled web images. The experimental results using this new database and the resized MSRA database [16] demonstrate that our algorithm outperforms previous state-of-the-art methods.
Peng Wang 0001, Jingdong Wang 0001, Jie Feng 0012, Hongbin Zha, Shipeng Li 0001
CVPR5
2012 Random-sampling-based spatial-temporal feature for consumer video concept classification
abstract
Concept classification for consumer videos is a challenging task considering the co-occurrence of a variety objects and arbitrary motions in video segments. In this paper, we present a novel video concept classification framework with random-sampling-based spatialtemporal features. Short-term random-sampled point tracks are obtained within video segments. The spatial-temporal features are extracted from these tracks. Concept codebooks are constructed using Multiple Instance Learning upon the spatial-temporal features. The SVM classifiers are trained over codebook-based histograms for an online concept detection. We performed experiments on a video database taken from YouTube. The experimental results demonstrate that the consumer videos can be efficiently assigned concept labels by our approach.
Anjun Wei, Yuru Pei, Hongbin Zha
ICIP3
2012 Total Variation and Euler's Elastica for Supervised Learning
Tong Lin 0002, Hanlin Xue, Hongbin Zha
ICML4
2012 Incoherent dictionary learning for sparse representation
Tong Lin 0002, Hongbin Zha
ICPR3
2012 Fusion of low-and high-dimensional approaches by trackers sampling for generic human motion tracking
Ye Liu 0002, Jinshi Cui, Huijing Zhao, Hongbin Zha
ICPR4
2012 Spatial consistency based selective reranking for content based object retrieval
Tan Xiao, Chao Zhang 0001, Hongbin Zha
ICPR3
2012 Direct least square fitting of ellipsoids
Xianghua Ying, Yongbo Hou, Sheng Guan, Hongbin Zha
ICPR6
2012 Laser-based intelligent surveillance and abnormality detection in extremely crowded scenarios
abstract
Abnormal activity detection plays a crucial role in surveillance applications, and a surveillance system that can perform robustly in the extremely crowded area has become an urgent need for public security. In this paper, we propose a novel laser-based system which can simultaneously perform the tracking, semantic scene learning and abnormality detection in the large and crowded environment. In our system, a novel abnormality detection model is proposed, and it considers and combines various factors that will influence human activity. Moreover, this model intensively investigate the relationship between pedestrians' social behaviors and their walking scenarios. We successfully applied the proposed system to the JR subway station of Tokyo, which can cover a 60×35m area, robustly track more than 180 targets at the same time and simultaneously perform the online semantic scene learning and abnormality detection with no human intervention.
Xuan Song 0001, Xiaowei Shao, Quanshi Zhang, Ryosuke Shibasaki, Huijing Zhao, Hongbin Zha
ICRA6
2012 A real-time motion planner with trajectory optimization for autonomous vehicles
abstract
In this paper, an efficient real-time autonomous driving motion planner with trajectory optimization is proposed. The planner first discretizes the plan space and searches for the best trajectory based on a set of cost functions. Then an iterative optimization is applied to both the path and speed of the resultant trajectory. The post-optimization is of low computational complexity and is able to converge to a higher-quality solution within a few iterations. Compared with the planner without optimization, this framework can reduce the planning time by 52% and improve the trajectory quality. The proposed motion planner is implemented and tested both in simulation and on a real autonomous vehicle in three different scenarios. Experiments show that the planner outputs high-quality trajectories and performs intelligent driving behaviors.
Wenda Xu, Junqing Wei, John M. Dolan, Huijing Zhao, Hongbin Zha
ICRA5
2012 Computing object-based saliency in urban scenes using laser sensing
abstract
It becomes a well-known technology that a low-level map of complex environment containing 3D laser points can be generated using a robot with laser scanners. Given a cloud of 3D laser points of an urban scene, this paper proposes a method for locating the objects of interest, e.g. traffic signs or road lamps, by computing object-based saliency. Our major contributions are: 1) a method for extracting simple geometric features from laser data is developed, where both range images and 3D laser points are analyzed; 2) an object is modeled as a graph used to describe the composition of geometric features; 3) a graph matching based method is developed to locate the objects of interest on laser data. Experimental results on real laser data depicting urban scenes are presented; efficiency as well as limitations of the method are discussed.
Yipu Zhao, Mengwen He, Huijing Zhao, Franck Davoine, Hongbin Zha
ICRA5
2012 A system of automated training sample generation for visual-based car detection
abstract
This paper presents a system to automatically generate car sample dataset for visual-based car detector training. The dataset contains multi-view car samples labeled with the car's pose, so that a view-discriminative training and car detection is also available. There are mainly two parts in the system: laser-based car detection and tracking generates motion trajectories of on-road cars, and then visual samples are extracted by fusing the detection and tracking results with visual-based detection. A multi-modal sensor system is developed for the omni-directional data collection on a test-bed vehicle. By processing the data of experiment conducted on the freeway of Beijing, a large number of multi-view car samples with pose information were generated. The samples' quality is evaluated by applying it in a visual car detector's training and testing procedure.
Chao Wang 0060, Huijing Zhao, Franck Davoine, Hongbin Zha
IROS4
2012 Learning lane change trajectories from on-road driving data
abstract
Lane change is one of the most principle driving behaviors on structure roads. It frequently happens in daily driving. A key issue in lane change technique is trajectory planning, where a set of trajectories describing possible vehicle motions are generated by applying a parametric function, and by uniformly sampling the end states in configuration space; the trajectories are then examined to find an optimal one for execution. However, such a trajectory set has poor efficiency due to the large sample number. Many trajectories in this set seldom happen in real human driving behaviors. In this research, lane change trajectories are collected from real driving data of different drivers. Their statistics are analyzed, through which, a simplified trajectory set is generated. Experiment results show that the trajectory set has much less number of samples but can still guarantee to cover usual lane change behaviors of human being.
Huijing Zhao, Franck Davoine, Hongbin Zha
Intelligent Vehicles Symposium4
2012 Omni-directional detection and tracking of on-road vehicles using multiple horizontal laser scanners
abstract
This research aims at generating an omnidirectional perception at the host vehicle's surroundings, extracting accurate and continuous motion trajectories of the nearby vehicles using low cost laser scanners. A system of detecting and tracking on-road vehicles using multiple laser scanners is developed, where focuses are cast on solving data association of simultaneous measurements from multiple sensors at different viewpoints, and state estimation in case of partial observations in dense dynamic situations. Experimental results in freeways in Beijing are presented, system efficiency is demonstrated, where motion trajectories describing driving behaviors such as overtaking, lane changing and other interactions between driving objects are captured. In addition, the accuracy in vehicle detection and tracking is examined using a reference vehicle with a ground truth GPS.
Huijing Zhao, Chao Wang 0060, Franck Davoine, Jinshi Cui, Hongbin Zha
Intelligent Vehicles Symposium6
2012 Contour canonical form: an efficient intrinsic embedding approach to matching non-rigid 3D objects
abstract
Evaluating the intrinsic similarities between non-rigid 3D shapes is of vital importance in content-based shape retrieval. In this paper, we present a novel intrinsic embedding technique, the contour canonical form, to express the isometry-invariant shape representation. The basic idea is to generate an unbent mapping shape for each subpart by aligning the geodesic contours. In details, we first extract the feature points on the non-rigid shape. Then, their canonical mapping positions are calculated, which are globally optimized under geodesic constraints defined on the shape surface. Guided by these positions, an embedding shape is finally obtained by adaptively rotating and translating the geodesic contours around the corresponding feature point. Compared with existing spectral embedding methods, our approach excels on both the preservation of geometric information and the computational efficiency. In the experiment, the contour canonical form is applied in retrieving non-rigid 3D shapes from the McGill articulated benchmark. The appealing results clearly demonstrate a significant performance improvement of our approach over state-of-the-art methods.
Xulei Wang, Hongbin Zha
ICMR2
2012 Unsupervised Image Matching Based on Manifold Alignment
abstract
This paper challenges the issue of automatic matching between two image sets with similar intrinsic structures and different appearances, especially when there is no prior correspondence. An unsupervised manifold alignment framework is proposed to establish correspondence between data sets by a mapping function in the mutual embedding space. We introduce a local similarity metric based on parameterized distance curves to represent the connection of one point with the rest of the manifold. A small set of valid feature pairs can be found without manual interactions by matching the distance curve of one manifold with the curve cluster of the other manifold. To avoid potential confusions in image matching, we propose an extended affine transformation to solve the nonrigid alignment in the embedding space. The comparatively tight alignments and the structure preservation can be obtained simultaneously. The point pairs with the minimum distance after alignment are viewed as the matchings. We apply manifold alignment to image set matching problems. The correspondence between image sets of different poses, illuminations, and identities can be established effectively by our approach.
Yuru Pei, Fengchun Huang, Fuhao Shi, Hongbin Zha
IEEE Trans. Pattern Anal. Mach. Intell.4
2012 A Fast Algorithm for Multidimensional Ellipsoid-Specific Fitting by Minimizing a New Defined Vector Norm of Residuals Using Semidefinite Programming
abstract
A quadratic surface in n-dimensional space is defined as the locus of zeros of a quadratic polynomial. The quadratic polynomial may be compactly written in notation by an (n+1)-vector and a real symmetric matrix of order n+1, where the vector represents homogenous coordinates of an n-D point, and the symmetric matrix is constructed from the quadratic coefficients. If an n-D quadratic surface is an n-D ellipsoid, the leading n × n principal submatrix of the symmetric matrix would be positive or opposite definite. As we know, to impose a matrix being positive or opposite definite, perhaps the best choice may be to employ semidefinite programming (SDP). From such straightforward and intuitive knowledge, in the literature until 2002, Calafiore first proposed a feasible method for multidimensional ellipsoid-specific fitting using SDP, which minimizes the 2--norm of the algebraic residual vector. However, the runtime of the method is significantly long and memory is often out when the number of fitted points is greater than several thousand. In this paper, we propose a fast and easily implemented algorithm for multidimensional ellipsoid-specific fitting by minimizing a new defined vector norm of the algebraic residual vector using SDP, which drastically decreases the size of the SDP problem while preserving accuracy. The proposed fast method can handle several million fitted points without any difficulty.
Xianghua Ying, Hongbin Zha
IEEE Trans. Pattern Anal. Mach. Intell.3
2012 Vanishing point detection using cascaded 1D Hough Transform from single images
Bo Li 0018, Xianghua Ying, Hongbin Zha
Pattern Recognit. Lett.4
2012 Detection and Tracking of Moving Objects at Intersections Using a Network of Laser Scanners
abstract
In our previous work, we reported a system that monitors an intersection using a network of horizontal laser scanners. This paper focuses on an algorithm for moving-object detection and tracking, given a sequence of distributed laser scan data of an intersection. The goal is to detect each moving object that enters the intersection; estimate state parameters such as size; and track its location, speed, and direction while it passes through the intersection. This work is unique, to the best of the authors' knowledge, in that the data is novel, which provides new possibilities but with great challenges; the algorithm is the first proposal that uses such data in detecting and tracking all moving objects that pass through a large crowded intersection with focus on achieving robustness to partial observations, some of which result from occlusions, and on performing correct data associations in crowded situations. Promising results are demonstrated using experimental data from real intersections, whereby, for 1063 objects moving through an intersection over 20 min, 988 are perfectly tracked from entrance to exit with an excellent tracking ratio of 92.9%. System advantages, limitations, and future work are discussed.
Huijing Zhao, Jie Sha, Yipu Zhao, Junqiang Xi, Jinshi Cui, Hongbin Zha, Ryosuke Shibasaki
IEEE Trans. Intell. Transp. Syst.6
2011 Tracking Generic Human Motion via Fusion of Low- and High-Dimensional Approaches
Yuandong Xu, Jinshi Cui, Huijing Zhao, Hongbin Zha
BMVC4
2011 Camera Resectioning from Image Edges with the L∞-Norm Using Linear Programming
Xianghua Ying, Lulu Hou, Hongbin Zha
BMVC5
2011 Sorted Random Projections for robust texture classification
abstract
This paper presents a simple and highly effective system for robust texture classification, based on (1) random local features, (2) a simple global Bag-of-Words (BoW) representation, and (3) Support Vector Machines (SVMs) based classification. The key contribution in this work is to apply a sorting strategy to a universal yet information-preserving random projection (RP) technique, then comparing two different texture image representations (histograms and signatures) with various kernels in the SVMs. We have tested our texture classification system on six popular and challenging texture databases for exemplar based texture classification, comparing with 12 recent state-of-the-art methods. Experimental results show that our texture classification system yields the best classification rates of which we are aware of 99.37% for CUReT, 97.16% for Brodatz, 99.30% for UMD and 99.29% for KTH-TIPS. Moreover, combining random features significantly outperforms the state-of-the-art descriptors in material categorization.
Li Liu 0002, Paul W. Fieguth, Gangyao Kuang, Hongbin Zha
ICCV4
2011 Robust consistent correspondence between 3D non-rigid shapes based on "Dual Shape-DNA"
abstract
In this paper, we propose a novel framework to construct dense and high-quality consistent correspondence between non-rigid surfaces. Our correspondence framework exploits “Dual Shape-DNA” (dual Laplace-Beltrami spectral embedding) to capture global characteristics of objects, and converts two originally different and complex shapes into two similar and simple shapes to facilitate the correspondence. Since our method avoids the computation of geodesic distances, it is robust to local topology changes. By exploiting the excellent properties of the dual domain, our dual spectral framework can robustly construct Laplace-Beltrami embeddings on highly non-regular 3D meshes. After performing initial non-rigid matching in the dual Laplace-Beltrami spectral domain, we return 3D spatial domain and apply a shape-preserving non-rigid deformation to produce the final dense consistent correspondence. We show that our framework is suitable for non-rigid consistent correspondence, and the high-quality correspondence results are achieved.
Huai-Yu Wu, Hongbin Zha
ICCV2
2011 Structure-sensitive superpixels via geodesic distance
abstract
Over-segments (i.e. superpixels) have been commonly used as supporting regions for feature vectors and primitives to reduce computational complexity in various image analysis tasks. In this paper, we describe a structuresensitive over-segmentation technique by exploiting Lloyd's algorithm with a geodesic distance. It generates smaller superpixels to achieve lower under-segmentation in structure-dense regions with high intensity or color variation, and produces larger segments to increase computational efficiency in structure-sparse regions with homogeneous appearance. We adopt geometric flows to compute the geodesic distances amongst pixels, and in the segmentation procedure, the density of over-segments is automatically adjusted according to an energy functional that embeds color homogeneity, structure density and compactness constraints. Comparative experiments with the Berkeley database show that the proposed algorithm outperforms prior arts while offering a comparable computational efficiency with fast methods, such as TurboPixels.
Peng Wang 0001, Jingdong Wang 0001, Rui Gan, Hongbin Zha
ICCV5
2011 A novel laser-based system: Fully online detection of abnormal activity via an unsupervised method
abstract
Abnormal activity detection plays a crucial role in surveillance applications, and such system has become an urgent need for public security. In this paper, we propose a novel laser-based system, which can perform the online detection of abnormal activity with an unsupervised way. The proposed system has the following key features that make it advantageous over previous ones: (1) It can cover quite a large and crowded area, such as subway station, public square, intersection and etc. (2) The overall system can vary with time period, incrementally learn the behavior pattern of pedestrians and perform the fully online detection of abnormal activity. This feature makes our system be quite suitable for the real-time applications. (3) The abnormal activity detection is carried out with a fully unsupervised way, there is no need for manual labelling and constructing the huge training datasets. We successfully applied the proposed system into the JR subway station of Tokyo, which can cover a 60×35m area, track more 150 targets at the same time and simultaneously perform the robust detection of abnormal activity with no human intervention.
Xuan Song 0001, Xiaowei Shao, Ryosuke Shibasaki, Huijing Zhao, Jinshi Cui, Hongbin Zha
ICRA6
2011 A vehicle model for micro-traffic simulation in dynamic urban scenarios
abstract
In order to improve energy efficiency of transport systems, eco-driving strategies are studied world-widely. However, most literatures on eco-driving based on traditional traffic flow models, are greatly simplified, and can not evaluate the effects on detailed driving behaviors. By referring to robot motion planning approaches, in this research a microscopic vehicle model is developed and it can represent different driving behaviors, such as aggressive or conservative driving; a collision detection algorithm is proposed that takes O(1) time to check for a trajectory's collision, enabling realtime planning; and a traffic simulation system is developed by incorporating traffic rules, so that the driving behaviors such as observing or not observing traffic rules can also be represented. Experiments are conducted on the simulation platform, and the performance of different driving behaviors on travel time, mileage, comfort and eco is studied.
Wenda Xu, Wen ZhaYao, Huijing Zhao, Hongbin Zha
ICRA4
2011 3D Line Drawing for Archaeological Illustration
Tao Luo 0003, Renju Li, Hongbin Zha
Int. J. Comput. Vis.3
2011 Partwise Cross-Parameterization via Nonregular Convex Hull Domains
abstract
In this paper, we propose a novel partwise framework for cross-parameterization between 3D mesh models. Unlike most existing methods that use regular parameterization domains, our framework uses nonregular approximation domains to build the cross-parameterization. Once the nonregular approximation domains are constructed for 3D models, different (and complex) input shapes are transformed into similar (and simple) shapes, thus facilitating the cross-parameterization process. Specifically, a novel nonregular domain, the convex hull, is adopted to build shape correspondence. We first construct convex hulls for each part of the segmented model, and then adopt our convex-hull cross-parameterization method to generate compatible meshes. Our method exploits properties of the convex hull, e.g., good approximation ability and linear convex representation for interior vertices. After building an initial cross-parameterization via convex-hull domains, we use compatible remeshing algorithms to achieve an accurate approximation of the target geometry and to ensure a complete surface matching. Experimental results show that the compatible meshes constructed are well suited for shape blending and other geometric applications.
Huai-Yu Wu, Chunhong Pan, Hongbin Zha, Qing Yang 0002, Songde Ma
IEEE Trans. Vis. Comput. Graph.3
2010 Modeling dwell-based eye pointing target acquisition
abstract
We propose a quantitative model for dwell-based eye pointing tasks. Using the concepts of information theory to analogize eye pointing, we define an index of difficulty (IDeye) for the corresponding tasks in a similar manner to the definition that Fitts made for hand pointing. According to our validations in different situations, IDeye, which takes account of the distinct characteristics of rapid saccades and involuntary eye jitters, can accurately and meaningfully describe eye pointing tasks. To the best of our knowledge, this work is the first successful attempt to model eye gaze interactions.
Xinyong Zhang, Xiangshi Ren, Hongbin Zha
CHI3
2010 Optimizing kd-trees for scalable visual descriptor indexing
abstract
In this paper, we attempt to scale up the kd-tree indexing methods for large-scale vision applications, e.g., indexing a large number of SIFT features and other types of visual descriptors. To this end, we propose an effective approach to generate near-optimal binary space partitioning and need low time cost to access the nodes in the query stage. First, we relax the coordinate-axis-alignment constraint in partition axis selection used in conventional kd-trees, and form a partition axis with the great variance by combining a few coordinate axes in a binary manner for each node, which yields better space partitioning and requires almost the same time cost to visit internal nodes during the query stage thanks to cheap projection operations. Then, we introduce a simple but very effective scheme to guarantee the partition axis of each internal node is orthogonal to or parallel with those of its ancestors, which leads to efficient distance computation between a query point and the cell associated with each node and yields fast priority search. Compared with the conventional kd-trees, our approach takes a little more tree construction time, but obtains much better nearest neighbor search performance. Experimental results on large scale local patch indexing and image search with tiny images show that our approach outperforms the state-of-the-art kd-tree based indexing methods.
You Jia, Jingdong Wang 0001, Hongbin Zha, Xian-Sheng Hua 0001
CVPR4
2010 An online approach: Learning-Semantic-Scene-by-Tracking and Tracking-by-Learning-Semantic-Scene
abstract
Learning the knowledge of scene structure and tracking a large number of targets are both active topics of computer vision in recent years, which plays a crucial role in surveillance, activity analysis, object classification and etc. In this paper, we propose a novel system which simultaneously performs the Learning-Semantic-Scene and Tracking, and makes them supplement each other in one framework. The trajectories obtained by the tracking are utilized to continually learn and update the scene knowledge via an online un-supervised learning. On the other hand, the learned knowledge of scene in turn is utilized to supervise and improve the tracking results. Therefore, this “adaptive learning-tracking loop” can not only perform the robust tracking in high density crowd scene, dynamically update the knowledge of scene structure and output semantic words, but also ensures that the entire process is completely automatic and online. We successfully applied the proposed system into the JR subway station of Tokyo, which can dynamically obtain the semantic scene structure and robustly track more than 150 targets at the same time.
Xuan Song 0001, Xiaowei Shao, Huijing Zhao, Jinshi Cui, Ryosuke Shibasaki, Hongbin Zha
CVPR6
2010 Global and local isometry-invariant descriptor for 3D shape comparison and partial matching
abstract
In this paper, based on manifold harmonics, we propose a novel framework for 3D shape similarity comparison and partial matching. First, we propose a novel symmetric mean-value representation to robustly construct high-quality manifold harmonic bases on nonuniform-sampling meshes. Then, based on the manifold harmonic bases constructed, a novel shape descriptor is presented to capture both of global and local features of 3D shape. This feature descriptor is isometry-invariant, i.e., invariant to rigid-body transformations and non-rigid bending. After characterizing 3D models with the shape features, we perform 3D retrieval with a up-to-date discriminative kernel. This kernel is a dimension-free approach to quantifying the similarity between two unordered feature-sets, thus especially suitable for our high-dimensional feature data. Experimental results show that our framework can be effectively used for both comprehensive comparison and partial matching among non-rigid 3D shapes.
Huai-Yu Wu, Hongbin Zha, Tao Luo 0003, Xulei Wang, Songde Ma
CVPR2
2010 Geometric properties of multiple reflections in catadioptric camera with two planar mirrors
abstract
A catadioptric system consisting of a pinhole camera and two planar mirrors is deeply investigated in this paper. The two mirrors combine to form a corner and face-to-face with the pinhole. Their relative pose is unknown. An object will be reflected in the mirror corner one-time or multiple-times. Using the pinhole, we may take an image containing the object and its reflections, i.e., simultaneously imaging multiple views of an object by a single camera. We discovered that each 3D point and its reflections lie on a circle. We call the point set composed of a 3D point and its reflections, a Reflection Point Group (RPG), and the circle related to a RPG is called a RPG circle. All RPG circles are parallel to one another. Furthermore, each RPG can be partitioned into two separate subgroups. Shape formed by all points in a subgroup is invariant with respect to location of the 3D point. From these geometric properties, two calibration approaches can be utilized: One is based on parallel circles; the other is using 2D homographies among invariant shapes. Experiments validate our approaches.
Xianghua Ying, Ren Ren 0002, Hongbin Zha
CVPR4
2010 Interactive viewpoint-space navigation for visual-audio exhibition of painting
abstract
In this paper, we present a system for exhibiting a Chinese landscape painting about 900 years old. There are three parts in our system: (1) we allocate a voice dubbing or background music, which is treated as a point sound source, onto the 2D painting and obtain its position in the 2D space. All of the audio data are then located in a 3D hidden space, by projecting their 2D positions to the 3D space through a projection model. (2) A two-layer directed graph structure is proposed to well organize the audio data in a 4D space (with 1D temporal and 3D spatial). (3) The exhibition is defined as an active exploration in a viewpoint space, which faces both the image and the 3D world where the sound sources reside. The 3D space and the two-layer graph structure generate a natural and meaningful stereo audio field. Meanwhile, compared to videos with guided walk through, the active exploration makes the exhibition more attractive.
Wei Ma 0008, Yang Liu 0006, Yizhou Wang 0001, Ying-Qing Xu, Hongbin Zha, Wen Gao 0001
ICME5
2010 CDP Mixture Models for Data Clustering
abstract
In Dirichlet process (DP) mixture models, the number of components is implicitly determined by the sampling parameters of Dirichlet process. However, this kind of models usually produces lots of small mixture components when modeling real-world data, especially high-dimensional data. In this paper, we propose a new class of Dirichlet process mixture models with some constrained principles, named constrained Dirichlet process (CDP) mixture models. Based on general DP mixture models, we add a resampling step to obtain latent parameters. In this way, CDP mixture models can suppress noise and generate the compact patterns of the data. Experimental results on data clustering show the remarkable performance of the CDP mixture models.
Yangfeng Ji, Tong Lin 0002, Hongbin Zha
ICPR3
2010 One-Shot Scanning Using a Color Stripe Pattern
abstract
Structured light 3D scanning has many applications such as 3D modeling, animation, motion analysis, deformation measurement and so on. Traditional structured light methods make use of a sequence of patterns to obtain the dense 3D data of objects. However, few methods have been proposed to achieve pixel wise reconstruction using one pattern only. In this paper, we proposes a one-shot scanning system based on a novel stripe pattern. This pattern uses color stripes with quadratic intensity distribution in each stripe. The color distribution is based on a De Bruijn sequence with six colors and order three. Graph cut is utilized to decode the color information and the resulting code is calculated using local intensity. Compared with traditional methods, the proposed method uses one pattern only and achieves pixel wise reconstruction. Experimental results show that our one-shot scanning system can robustly capture 3D data with high accuracy.
Renju Li, Hongbin Zha
ICPR2
2010 Single View Metrology Along Orthogonal Directions
abstract
In this paper, we describe how 3D metric measurements can be determined from a single uncalibrated image, when only minimal geometric information are available in the image. The minimal information just is orthogonal vanishing points. Given such limited information, we show that the length ratios on different orthogonal directions can be directly computed. The exciting discovery of the method seems to oppose common senses: Usually, in the calibration process, all edge-lengths of cuboid are known, in this paper, cuboid edge-lengths are unknown but its edge-lengths ratios can be recovered from image. 3D metric measurements can be directly computed from the image using our linear method.
Lulu Hou, Ren Ren 0002, Xianghua Ying, Hongbin Zha
ICPR5
2010 A dynamic subgoal path planner for unpredictable environments
abstract
Although lots of planning algorithms have focused on the planning of fixed manipulators and mobile robots in moderate dynamic environments, seldom planning algorithms can be employed to deal with mobile agents in the presence of large scenario scales and unpredictable changing obstacles. Path planning for mobile robots in unpredictable environments would be an extreme challenge since computational complexity increase dramatically with high dimensionality, unpredictability and large scales. In this paper, a novel and real-time approach is proposed to solve this problem by generating subgoals dynamically according to time and potential values. This dynamic subgoal based approach includes two procedures, the subgoal generator and the inter-subgoal or inner replanner. On the one hand, a set of high-level subgoals is generated dynamically by an improved single shot strategy that could tailor itself adaptively. On the other hand, a roadmap is built during the preprocessing phase by employing a localized Dynamic Roadmap Mapping (local-DRM) for inter-subgoal replanning. Finally these two procedures will collaborate according to the potential field criterion to ensure completeness. Our approach can not only generate paths rapidly enough to satisfy the requirements of an anytime planer but also work for large scenario scales. Experimental results on different kinds of mobile agents, in large scenario scales and in the presence of unpredictable changing obstacles show that our approach can find out a collision free path on an average of 0.11s for a single planning, indicating an anytime planner.
Hong Liu 0008, Weiwei Wan, Hongbin Zha
ICRA3
2010 Fusion of laser and vision for multiple targets tracking via on-line learning
abstract
Multi-target tracking becomes significantly more challenging when the targets are in close proximity or frequently interact with each other. This paper presents a promising tracking system to deal with these problems. The novelty of this system is that laser and vision, tracking and learning are integrated and can complement each other in one framework: when the targets do not interact with each other, the laser-based independent trackers are employed and the visual information is extracted simultaneously to train some classifiers for the “possible interacting targets”. When the targets are in close proximity, the learned classifiers and visual information are used to assist in tracking. Therefore, this mode of co-operation between them not only deals with various tough problems encountered in the tracking, but also ensures that the entire process can be completely on-line and automatic. Experimental results demonstrated that laser and vision fully display their respective advantages in our system, and it is easy for us to obtain a perfect trade-off between tracking accuracy and time-cost.
Xuan Song 0001, Huijing Zhao, Jinshi Cui, Xiaowei Shao, Ryosuke Shibasaki, Hongbin Zha
ICRA6
2010 Scene understanding in a large dynamic environment through a laser-based sensing
abstract
It became a well known technology that a map of complex environment containing low-level geometric primitives (such as laser points) can be generated using a robot with laser scanners. This research is motivated by the need of obtaining semantic knowledge of a large urban outdoor environment after the robot explores and generates a low-level sensing data set. An algorithm is developed with the data represented in a range image, while each pixel can be converted into a 3D coordinate. Using an existing segmentation method that models only geometric homogeneities, the data of a single object of complex geometry, such as people, cars, trees etc., is partitioned into different segments. Such a segmentation result will greatly restrict the capability of object recognition. This research proposes a framework of simultaneous segmentation and classification of range image, where the classification of each segment is conducted based on its geometric properties, and homogeneity of each segment is evaluated conditioned on each object class. Experiments are presented using the data of a large dynamic urban outdoor environment, and performance of the algorithm is evaluated.
Huijing Zhao, Yipu Zhao, Hongbin Zha
ICRA5
2010 Segmentation and classification of range image from an intelligent vehicle in urban environment
abstract
As the rapid development of sensing and mapping techniques, it becomes a well-known technology that a map of complex environment can be generated using a robot carrying sensors. However, most of the existing researches represent environments directly using the integration of point clouds or other low-level geometric primitives. It remains an open problem to automatically convert these low-level map representations to semantic descriptions in order to effectively support high-level decision of a robot. Based on another representation of 3D point clouds, i.e. range image, this paper proposes a framework of segmentation and classification of range image, the objective of which is to annotate class labels to the data clusters that are obtained through a graph-based segmentation. Experimental results are presented and evaluated demonstrating that the proposed algorithm has efficiency in understanding the semantic knowledge of a large dynamic urban outdoor environment.
Huijing Zhao, Yipu Zhao, Hongbin Zha
IROS5
2010 Learning Robust Similarity Measures for 3D Partial Shape Retrieval
Xulei Wang, Hua-Yan Wang, Hongbin Zha, Hong Qin 0001
Int. J. Comput. Vis.4
2010 Model Transduction for Triangle Meshes
Huai-Yu Wu, Chunhong Pan, Hongbin Zha, Songde Ma
J. Comput. Sci. Technol.3
2009 Crease Detection on Noisy Meshes via Probabilistic Scale Selection
Tao Luo 0003, Huai-Yu Wu, Hongbin Zha
ACCV (2)3
2009 Visyllable-specific facial transition motion embedding and extraction
abstract
The visual facial appearances are important to the speaking perception. The effective and reasonable extraction of the transition motions between the keyframes is desirable to the facial speech animation. In this paper, we present the visyllable-specific transition motion embedding with the temporal extension of the Laplacian eigenmaps (TLE). By imposing the temporal constraints, the TLE-based embedding preserves the possible transitions between the keyshapes inside the visyllable sequence. Given the keyframe pair, the in-between transition motions can be extracted in the latent space by the shortest path searching algorithm. Our experiments demonstrate an effective engine for embedding and extracting the transition motions specific to the visyllables.
Yuru Pei, Hongbin Zha
ICIP2
2009 Interactive modeling of 3D facial expressions with hierarchical Gaussian process latent variable models
abstract
The natural expressions play an important role in the daily communication. The efficient and intuitive facial expression editing based on the limited constraints is desirable in the facial animation. In this paper, we present an interactive 3D facial expression editing system with the hierarchical Gaussian process latent variable model (HGPLVM). The hierarchical model incorporates the joint work of the local facial features to produce the natural expressions. To deal with the holistic expression modeling from the local constraints, the inverse mapping between the low-level feature nodes and the high-level facial region nodes is established by the RBF regression model in the latent space. A propagation algorithm is introduced to predict the holistic facial configurations. The experiments demonstrate the 3D facial expressions satisfying the user constraints can be produced efficiently.
Fuhao Shi, Yuru Pei, Hongbin Zha
ICIP3
2009 3D facial expression editing based on the dynamic graph model
abstract
To model a detailed 3D expressive face based on the limited user constraints is a challenge work. In this paper, we present the facial expression editing technique based on a dynamic graph model. The probabilistic relations between facial expressions and the complex combination of local facial features, as well as the temporal behaviors of facial expressions are represented by the hierarchical dynamic Bayesian network. Given limited user-constraints on the sparse feature mesh, the system can infer the basis expression probabilities, which are used to locate the corresponding expressive mesh in the shape space spanned by the basis models. The experiments demonstrate the 3D dense facial meshes corresponding to the user-constraints can be synthesized effectively.
Yuru Pei, Hongbin Zha
ICME2
2009 Mahalanobis Distance Based Non-negative Sparse Representation for Face Recognition
abstract
Sparse representation for machine learning has been exploited in past years. Several sparse representation based classification algorithms have been developed for some applications, for example, face recognition. In this paper, we propose an improved sparse representation based classification algorithm. Firstly, for a discriminative representation, a non-negative constraint of sparse coefficient is added to sparse representation problem. Secondly, Mahalanobis distance is employed instead of Euclidean distance to measure the similarity between original data and reconstructed data. The proposed classification algorithm for face recognition has been evaluated under varying illumination and pose using standard face databases. The experimental results demonstrate that the performance of our algorithm is better than that of the up-to-date face recognition algorithm based on sparse representation.
Yangfeng Ji, Tong Lin 0002, Hongbin Zha
ICMLA3
2009 Moving object classification using horizontal laser scan data
abstract
Motivated by two potential applications, i.e. enhancing driving safety and traffic data collection, a system has been developed using a single-layer horizontal laser scanner as the major sensor for both localization and perception of the surroundings in a large dynamic urban environment. This research focuses on a classification method, that given a stream of laser measurements, classify the moving object into either a person, a group of people, a bicycle or a car. In this research, a number of features are defined after examining the property of data appearance. A classification method is proposed after examining the likelihood measures between each pair of feature and class. Experimental results are presented, demonstrating that the algorithm has efficiency with respect to both driving safety and traffic data collection in highly dynamic environment.
Huijing Zhao, Quanshi Zhang, Masaki Chiba, Ryosuke Shibasaki, Jinshi Cui, Hongbin Zha
ICRA6
2009 Combining Laser-Scanning Data and Images for Target Tracking and Scene Modeling
Hongbin Zha, Huijing Zhao, Jinshi Cui, Xuan Song 0001, Xianghua Ying
ISRR1
2009 Modeling plants with sensor data
Wei Ma 0008, Bo Xiang, Hongbin Zha, Xiaopeng Zhang 0001
Sci. China Ser. F Inf. Sci.3
2009 Robust human tracking based on multi-cue integration and mean-shift
Hong Liu 0008, Hongbin Zha, Yuexian Zou
Pattern Recognit. Lett.3
2009 A Laser-Scanner-Based Approach Toward Driving Safety and Traffic Data Collection
abstract
This work is motivated by the following two potential applications: 1) enhancing driving safety and 2) collecting traffic data in a large dynamic urban environment. A laser-scanner-based approach is proposed. The problem is formulated as a simultaneous localization and mapping (SLAM) with object tracking and classification, where the focus is on managing a mixture of data from both dynamic and static objects in a highly dynamic environment. A trajectory-oriented closure is also proposed using the sporadically available global positioning system (GPS) measurements in urban areas to assist for global accuracy, particularly when the vehicle makes a noncyclical measurement in a large outdoor environment. Experiments are conducted using the data that were collected along a course near 4.5 km in a highly dynamic environment. Possibilities of the approaches toward the two potential applications are demonstrated, and avenues for future works are discussed.
Huijing Zhao, Masaki Chiba, Ryosuke Shibasaki, Xiaowei Shao, Jinshi Cui, Hongbin Zha
IEEE Trans. Intell. Transp. Syst.6
2008 Dimension Amnesic Pyramid Match Kernel
Xulei Wang, Hongbin Zha
AAAI3
2008 Probabilistic Detection-based Particle Filter for Multi-target Tracking
abstract
In this paper, we present a Probabilistic Detection-based Particle Filter (PD-PF) for tracking a variable number of interacting targets. When the objects do not interact with each other, our method performs like the deterministic detection-base methods. When the objects are in close proximity, the interactions and occlusions are modelled by a mixed proposal constructed by probabilistic detections and information from dynamic models. Specially, prior of detection-reliability minimizes the influence of non-detection or false alarm in the tracking. Moreover, we run independent PD-PF for each target, such that particles are sampled in a small state space, thus our method not only obtains a better approximation of posterior than joint particle filter or independent particle filter when interactions occur, but also has an acceptable computational complexity. Different evaluations demonstrate the validity and efficiency of the proposed method. 1
Xuan Song 0001, Jinshi Cui, Hongbin Zha, Huijing Zhao
BMVC3
2008 Improving eye cursor's stability for eye pointing tasks
abstract
In order to improve the stability of eye cursor, we introduce three methods, force field (FF), speed reduction (SR), and warping to target center (TC) to modulate eye cursor trajectories by counteracting eye jitter, which is the main cause of destabilizing the eye cursor. We evaluate these methods using two controlled experiments. One is an attention task experiment, which indicates that both FF and SR significantly alleviate the instability of eye cursor, but TC is not as we anticipated. The other is a 2D pointing task experiment, which shows that FF and SR as well as the improved implementation of SR (iSR) indeed improve human performance in dominant dwell-based eye pointing tasks of eye-based interactions. The method iSR is especially effective to accelerate eye pointing (10.5% and 8.5%) and reduce error rate (6.1% and 2.7%) when target diameter D = 45 and 60 pixels.
Xinyong Zhang, Xiangshi Ren, Hongbin Zha
CHI3
2008 Vision-Based Multiple Interacting Targets Tracking via On-Line Supervised Learning
Xuan Song 0001, Jinshi Cui, Hongbin Zha, Huijing Zhao
ECCV (3)3
2008 Multi-scale creases detection on noisy meshes
abstract
We propose a multi-scale approach to detecting creases on noisy meshes. Given a noisy mesh as input, first we generate its discrete multi-scale representation based on an improved anisotropic diffusion method. Then, a method is presented for detecting salient vertices on a noisy mesh by exploiting curvatures estimated at different scales. Finally, creases are generated by connecting curvature extrema points detected on the mesh edges with the salient vertices as end points. During this process, the mesh features are preserved. Experimental results demonstrate that our approach works well on noisy meshes reconstructed from raw scanning data, by showing that geometrically and perceptually salient creases are well detected on noisy meshes.
Tao Luo 0003, Hongbin Zha
ICIP2
2008 Decomposition of branching volume data by tip detection
abstract
We present an approach to decomposing branching volume data into sub-branches. First, a metric is proposed for evaluating local convexities in volumetric data, and it is a criterion for global selection of tip points. Second, a multi-path growing strategy is adopted to segment the volumes based on a DFS transformation starting from the tips. Experiments show that this approach is capable of generating desirable components and reasonable segmentation boundaries of a volume.
Wei Ma 0008, Bo Xiang, Xiaopeng Zhang 0001, Hongbin Zha
ICIP4
2008 Creating a face model from an unknown skull based on the tissue map
abstract
The craniofacial reconstruction is employed as an initialization of the identification in forensics. In this paper we present a tissue map based craniofacial reconstruction technique. The reconstruction is formulated as the superimposition of the selected tissue map onto the novel skull. The key problem is the accurate map registration, which is implemented as a warping guided by 2D feature patterns. Given a novel skull, the feature patterns are extracted automatically under an energy minimization framework. The feature configuration on the warped tissue map is expected to resemble that on the novel skull. The target facial model is reconstructed by a smooth interpolation of the point cloud, which results from a simple addition of range images. The presented experiments demonstrate the facial outlook can be reconstructed from the tissue map feasibly and efficiently.
Yuru Pei, Hongbin Zha, Zhongbiao Yuan
ICIP2
2008 Adaptive feature-spatial representation for Mean-shift tracker
abstract
Mean-shift tracker plays an important role in computer vision applications due to its efficiency in mode seeking. By encoding the spatial information appropriately, the robustness of tracking could be greatly enhanced. However, to account for the deformation and other sources of variation of the tracking object, the spatial configuration should not be fixed apriori and it is more suitable to be adapted online. To this end, this paper presents a novel method to formulate an adaptive feature-spatial representation (FSR) for mean-shift tracking. By encoding blocking features of the tracking object with a set of adaptively weighted and spatially distributed tunable kernels, the object variations, like deformations and partial occlusions, can be handled appropriately. Extensive experiments under various conditions clearly demonstrate the obvious advantage of our approach compared to the classical mean-shift trackers.
Hong Liu 0008, Hongbin Zha
ICIP4
2008 Dirichlet component analysis: feature extraction for compositional data
abstract
We consider feature extraction (dimensionality reduction) for compositional data, where the data vectors are constrained to be positive and constant-sum. In real-world problems, the data components (variables) usually have complicated "correlations" while their total number is huge. Such scenario demands feature extraction. That is, we shall de-correlate the components and reduce their dimensionality. Traditional techniques such as the Principle Component Analysis (PCA) are not suitable for these problems due to unique statistical properties and the need to satisfy the constraints in compositional data. This paper presents a novel approach to feature extraction for compositional data. Our method first identifies a family of dimensionality reduction projections that preserve all relevant constraints, and then finds the optimal projection that maximizes the estimated Dirichlet precision on projected data. It reduces the compositional data to a given lower dimensionality while the components in the lower-dimensional space are de-correlated as much as possible. We develop theoretical foundation of our approach, and validate its effectiveness on some synthetic and real-world datasets.
Hua-Yan Wang, Qiang Yang 0001, Hong Qin 0001, Hongbin Zha
ICML4
2008 Adaptive p-posterior mixture-model kernels for multiple instance learning
abstract
In multiple instance learning (MIL), how the instances determine the bag-labels is an essential issue, both algorithmically and intrinsically. In this paper, we show that the mechanism of how the instances determine the bag-labels is different for different application domains, and does not necessarily obey the traditional assumptions of MIL. We therefore propose an adaptive framework for MIL that adapts to different application domains by learning the domain-specific mechanisms merely from labeled bags. Our approach is especially attractive when we are encountered with novel application domains, for which the mechanisms may be different and unknown. Specifically, we exploit mixture models to represent the composition of each bag and an adaptable kernel function to represent the relationship between the bags. We validate on synthetic MIL datasets that the kernel function automatically adapts to different mechanisms of how the instances determine the bag-labels. We also compare our approach with state-of-the-art MIL techniques on real-world benchmark datasets.
Hua-Yan Wang, Qiang Yang 0001, Hongbin Zha
ICML3
2008 Convenient reconstruction of natural plants by images
abstract
Convenient reconstruction of natural plants is a difficult task because of their intrinsic complex geometry. In this paper, we propose a convenient image-based approach to modeling and rendering large-leaf plants by view synthesis based on an approximate geometric model. The model we adopt consists of a set of billboard clusters, obtained by automatically decomposing a volume based on the flat property of leaves and assigning billboards corresponding to all input image viewpoints for each cluster. The volume is recovered from a set of images. In the final rendering procedure, a view-dependent billboard in each cluster is generated by interpolating already constructed ones. The whole process can be finished in half an hour and requires no user intervention.
Wei Ma 0008, Hongbin Zha
ICPR2
2008 Image-based plant modeling by knowing leaves from their apexes
abstract
In the paper, we present a novel approach to modeling plants from images by detecting apex features. First, an effective algorithm is proposed to extract apex features in volumetric data recovered from the images. It provides position and pose information for assigning 3D generic leaves. Then, the 3D leaf shapes are determined by an optimization based on the volume. Finally, Branches are modeled by using a particle flow approach. The proposed method is simply with limited manual intervention and has the obvious benefit of knowing a leaf by its visible apex part.
Wei Ma 0008, Hongbin Zha, Xiaopeng Zhang 0001, Bo Xiang
ICPR2
2008 Facial feature estimation from the local structural diversity of skulls
abstract
In forensics, the craniofacial reconstruction is employed as an initialization of the identification from skulls. It is a challenging work to develop such a system due to the ambiguity in the relationship between the shape of the skull and the face. In this paper, we present a facial feature estimation method based on the local structural diversity of skulls. A mapping system between the skull structural measurements and the facial feature shapes is established via a RBF regression model. The PCA subspaces are established for the local facial features and the skull structures. Moreover, we investigate the attribute vector of the facial feature polyhedron and the distance graph of the skull structure as the shape descriptors. The experiments demonstrate the feature outlooks can be estimated feasibly and efficiently.
Yuru Pei, Hongbin Zha, Zhongbiao Yuan
ICPR2
2008 Efficient detection of projected concentric circles using four intersection points on a secant line
abstract
Concentric circles are often used for calibration. Based on the geometric properties of concentric circles, we proved that only four intersection edge points on one secant line of the two images of the concentric circles (ICC) are sufficient to determine the whole parameters of the ICC, when the ratio of the radii of two concentric circles is given. Experimental results validate the proposed approach.
Xianghua Ying, Hongbin Zha
ICPR2
2008 Tracking interacting targets with laser scanner via on-line supervised learning
abstract
Successful multi-target tracking requires locating the targets and labeling their identities. For the laser based tracking system, the latter becomes significantly more challenging when the targets frequently interact with each other. This paper presents a novel on-line supervised learning based method for tracking interacting targets with laser scanner. When the targets do not interact with each other, we collect samples and train a classifier for each target. When the targets are in close proximity, we use these classifiers to assist in tracking. Different evaluations demonstrate that this method has a better tracking performance than previous methods when interactions occur, and can maintain correct tracking under various complex tracking situations.
Xuan Song 0001, Jinshi Cui, Xulei Wang, Huijing Zhao, Hongbin Zha
ICRA5
2008 SLAM in a dynamic large outdoor environment using a laser scanner
abstract
In this research, we propose a method of SLAM in a dynamic large outdoor environment using a laser scanner. Focus are cast on solving two major problems: 1) achieving global accuracy especially in non-cyclical environment, 2) tackling a mixture of data from both dynamic and static objects. Algorithms are developed, where GPS data and control inputs are used to diagnose pose error and guide to achieve a global accuracy; Classification of laser points and objects are conducted not in an independent module but across the processing in a framework of SLAM with moving object detection and tracking. Experiments are conducted using the data from two test-bed vehicles, and performance of the algorithms are demonstrated.
Huijing Zhao, Masaki Chiba, Ryosuke Shibasaki, Xiaowei Shao, Jinshi Cui, Hongbin Zha
ICRA6
2008 Predictive model for path planning by using k-near dynamic bridge builder and Inner Parzen Window
abstract
Robotic path planning in changing environments with difficult regions is an extremely challenge. Since the structure of configuration space (C-space) will change when obstacles move in workspace (W-space), the planner should have the capacity of building approximate structure of C-space, while avoiding intense computational complexity. Further, difficult regions will also change their positions, which requires the planner should be able to identify them fast and increase the free nodes inside them efficiently. This paper presents a novel approach for path planning in changing environments using predictive model, which is inspired by the idea of active learning. With the help of W-C nodes mapping, this predictive model is built to capture the approximate structure of C-space, while avoiding intense computational complexity. This model include two steps: K-near Dynamic Bridge Builder (K-near DBB) is proposed to identify difficult passages in the space first, and then Inner Parzen Window is adopted to sample points in these difficult regions without invoking any collision checker. Experiments are carried out with two 6-DOF manipulators, and our approach can find a path with high time efficiency and low error rate, even if the environment is complex.
Hong Liu 0008, Weiwei Wan, Hongbin Zha
IROS4
2008 The Craniofacial Reconstruction from the Local Structural Diversity of Skulls
abstract
Abstract The craniofacial reconstruction is employed as an initialization of the identification from skulls in forensics. In this paper, we present a two‐level craniofacial reconstruction framework based on the local structural diversity of the skulls. On the low level, the holistic reconstruction is formulated as the superimposition of the selected tissue map on the novel skull. The crux is the accurate map registration, which is implemented as a warping guided by the 2D feature curve patterns. The curve pattern extraction under an energy minimization framework is proposed for the automatic feature labeling on the skull depth map. The feature configuration on the warped tissue map is expected to resemble that on the novel skull. In order to make the reconstructed faces personalized, on the high level, the local facial features are estimated from the skull measurements via a RBF model. The RBF model is learnt from a dataset of the skull and the face feature pairs extracted from the head volume data. The experiments demonstrate the facial outlooks can be reconstructed feasibly and efficiently.
Yuru Pei, Hongbin Zha, Zhongbiao Yuan
Comput. Graph. Forum2
2008 Identical Projective Geometric Properties of Central Catadioptric Line Images and Sphere Images with Applications to Calibration
Xianghua Ying, Hongbin Zha
Int. J. Comput. Vis.2
2008 Multi-modal tracking of people using laser scanners and video camera
Jinshi Cui, Hongbin Zha, Huijing Zhao, Ryosuke Shibasaki
Image Vis. Comput.2
2008 Riemannian Manifold Learning
abstract
Recently, manifold learning has been widely exploited in pattern recognition, data analysis, and machine learning. This paper presents a novel framework, called Riemannian manifold learning (RML), based on the assumption that the input high-dimensional data lie on an intrinsically low-dimensional Riemannian manifold. The main idea is to formulate the dimensionality reduction problem as a classical problem in Riemannian geometry, i.e., how to construct coordinate charts for a given Riemannian manifold? We implement the Riemannian normal coordinate chart, which has been the most widely used in Riemannian geometry, for a set of unorganized data points. First, two input parameters (the neighborhood size k and the intrinsic dimension d) are estimated based on an efficient simplicial reconstruction of the underlying manifold. Then, the normal coordinates are computed to map the input high-dimensional data into a low-dimensional space. Experiments on synthetic data as well as real world images demonstrate that our algorithm can learn intrinsic geometric structures of the data, preserve radial geodesic distances, and yield regular embeddings.
Tong Lin 0002, Hongbin Zha
IEEE Trans. Pattern Anal. Mach. Intell.2
2008 Skew detection for complex document images using robust borderlines in both text and non-text regions
Hong Liu 0008, Hongbin Zha
Pattern Recognit. Lett.3
2007 Synchronized Ego-Motion Recovery of Two Face-to-Face Cameras
Jinshi Cui, Yasushi Yagi, Hongbin Zha, Yasuhiro Mukaigawa, Kazuaki Kondo
ACCV (1)3
2007 Camera Calibration Using Principal-Axes Aligned Conics
Xianghua Ying, Hongbin Zha
ACCV (1)2
2007 Fuzzy Decision Method for Motion Deadlock Resolving in Robot Soccer Games
Hong Liu 0008, Hongbin Zha
ICIC (1)3
2007 Collaborative Mean Shift Tracking Based on Multi-Cue Integration and Auxiliary Objects
abstract
Colour-based mean shift is an effective and fast algorithm for tracking colour blobs. However, it is vulnerable to full occlusion and target out of range for a few frames. This paper proposes a tracking method based on multi-cue integration and auxiliary objects to deal with these problems. A colour-location-prediction integration mean shift method is proposed to track each auxiliary object. Motivated by the idea of tuning weight of each cue according to their performances, these three cues are integrated adaptively according to their quality functions. Moreover, auxiliary objects get effective relative information with targets automatically, and update the information ceaselessly. When the target disappears, auxiliary objects will export useful information to estimate the location of the target. Experiments show that this method can adapt the weight of multi-cue efficiently, reinitialize the targets after long time disappearance, and increase the robustness of tracking in various conditions.
Hong Liu 0008, Hongbin Zha
ICIP (3)4
2007 An Efficient Method for the Detection of Projected Concentric Circles
abstract
Concentric circles are often used as calibration features since they possess good geometric properties. This paper presents an efficient method for the detection of projected concentric circles in the image plane while considering their special geometric properties. The proposed method is capable of detecting partially visible concentric circles. Experimental results demonstrate the validity of the proposed approach.
Xianghua Ying, Hongbin Zha
ICIP (6)2
2007 Dirichlet aggregation: unsupervised learning towards an optimal metric for proportional data
abstract
Proportional data (normalized histograms) have been frequently occurring in various areas, and they could be mathematically abstracted as points residing in a geometric simplex. A proper distance metric on this simplex is of importance in many applications including classification and information retrieval. In this paper, we develop a novel framework to learn an optimal metric on the simplex. Major features of our approach include: 1) its flexibility to handle correlations among bins/dimensions; 2) widespread applicability without being limited to ad hoc backgrounds; and 3) a "real" global solution in contrast to existing traditional local approaches. The technical essence of our approach is to fit a parametric distribution to the observed empirical data in the simplex. The distribution is parameterized by affinities between simplex vertices, which is learned via maximizing likelihood of observed data. Then, these affinities induce a metric on the simplex, defined as the earth mover's distance equipped with ground distances derived from simplex vertex affinities.
Hua-Yan Wang, Hongbin Zha, Hong Qin 0001
ICML2
2007 A dynamic bridge builder to identify difficult regions for path planning in changing environments
abstract
This paper presents an efficient path planner to identify difficult regions for path planning in changing environments, in which obstacles can move randomly. The difficult regions consist of narrow passages and the boundaries of obstacles in robot Configuration Space (C-space). These regions exert significantly negative influence on finding a valid path in static environments. The problem becomes more complicated in changing environments, because that the regions will change their positions when obstacles move. Besides, it is necessary to identify difficult regions in real time since obstacles may move frequently. To identify difficult regions fast when they change their positions, a dynamic bridge builder is proposed based on a W-C nodes mapping and a Bridge planner method. The W-C nodes mapping is used not only to conserve the validity of nodes in C-space, but also to provide the information about where a "bridge" should be built, i.e. the positions of narrow passages, and where the boundaries of obstacles are. Furthermore, a hierarchy sampling strategy is employed to boost the density of nodes in difficult regions efficiently. In the query phase, a Lazy-edges evaluation method is adopted to validate the edges in a found path. Simulated experiments for a dual-manipulator system show that our method is efficient for path planning in changing environments.
Hong Liu 0008, Xuezhi Deng, Hongbin Zha
IROS4
2007 Camera calibration from a circle and a coplanar point at infinity with applications to sports scenes analyses
abstract
Circles, such as concentric circles, arbitrary coplanar or parallel circles, are often employed for camera calibration. An interesting observation in football or basketball scenes is that a circle called the center circle existing in the midfield. By taking into account the midfield line which passes through the center circle’ center, a novel camera calibration method using the images of the midfield is proposed in this paper. A point at infinity can be determined from the center circle and the midfield line using the pole-polar relationship. The image of a point at infinity is called a vanishing point. The image of a circle is called a circle image in this paper. From one circle image and one coplanar vanishing point, a cubic constraint on the IAC can be obtained. The degenerate cases are discussed, and some applications are given in this paper.
Xianghua Ying, Hongbin Zha
IROS2
2007 Automatic seal image retrieval method by using shape features of Chinese characters
abstract
In many eastern countries, a large number of seal images need to be identified every day. The system compares an input seal image with its reference seal and validates the authenticity of it. However, the reference seal is usually found manually. As each reference seal has quite a long corresponding ID, inputting ID to get the needed seal one by one costs quite a lot of time and this manual stage has become the bottleneck of automatic seal identification. To make the seal verification system more automatic, a new retrieval method based on Chinese characters' shape features will be introduced in this paper. Firstly, the main characters region is obtained by transforming the round seal image to a rectangle one and choosing the main part for each seal. Secondly, every single character is segmented using correlation method, and the number of characters in a seal can be got. Thirdly, four horizontal and four vertical features are extracted for each character in a seal and an eigenvector called position code is defined. Finally, the retrieval strategy is mainly based on transform from position code to a weight to decide which two seal images have the most similarity. Experiments on a database of 1000 testing seal images provide the retrieval ability for the proposed approach. The correct rate goes to 95.3%.
Hong Liu 0008, Hongbin Zha
SMC4
2007 Laser-based detection and tracking of multiple people in crowds
Jinshi Cui, Hongbin Zha, Huijing Zhao, Ryosuke Shibasaki
Comput. Vis. Image Underst.2
2007 Stylized synthesis of facial speech motions
abstract
Abstract Stylized synthesis of facial speech motions is central to facial animation. Most synthesis algorithms put emphasis on the reasonable concatenation of captured motion segments. The dynamic modeling of speech units, e.g. visemes and visyllables (the visual appearance of a syllable), has not drawn much attention. In this paper, we address the fundamental issues regarding the stylized dynamic modeling of visyllables. The decomposable generalized model is learnt for the stylized motion synthesis. The visyllable modeling includes two parts: (1) A dynamic model for each kind of visyllable that is learnt based on a Gaussian Process Dynamical Model; (2) A multilinear model based unified mapping between the high dimensional observation space and low dimensional latent space. The dynamic visyllable model embeds the high dimensional motion data, and constructs the dynamic mapping in the latent space simultaneously. To generalize the visyllable model from several instances, the mapping coefficient matrices are assembled to a tensor, which is decomposed into independent modes, e.g. identity and uttering styles. Therefore, with the linear combination of components in each mode, the novel stylized motions can be synthesized. Copyright © 2007 John Wiley & Sons, Ltd.
Yuru Pei, Hongbin Zha
Comput. Animat. Virtual Worlds2
2007 Transferring of Speech Movements from Video to 3D Face Space
abstract
We present a novel method for transferring speech animation recorded in low quality videos to high resolution 3D face models. The basic idea is to synthesize the animated faces by an interpolation based on a small set of 3D key face shapes which span a 3D face space. The 3D key shapes are extracted by an unsupervised learning process in 2D video space to form a set of 2D visemes which are then mapped to the 3D face space. The learning process consists of two main phases: 1) Isomap-based nonlinear dimensionality reduction to embed the video speech movements into a low-dimensional manifold and 2) K-means clustering in the low-dimensional space to extract 2D key viseme frames. Our main contribution is that we use the Isomap-based learning method to extract intrinsic geometry of the speech video space and thus to make it possible to define the 3D key viseme shapes. To do so, we need only to capture a limited number of 3D key face models by using a general 3D scanner. Moreover, we also develop a skull movement recovery method based on simple anatomical structures to enhance 3D realism in local mouth movements. Experimental results show that our method can achieve realistic 3D animation effects with a small number of 3D key face models.
Yuru Pei, Hongbin Zha
IEEE Trans. Vis. Comput. Graph.2
2007 Editorial
Hongbin Zha, Bingfeng Zhou
Vis. Comput.1
2006 Vision Based Speech Animation Transferring with Underlying Anatomical Structure
Yuru Pei, Hongbin Zha
ACCV (1)2
2006 Fisheye Lenses Calibration Using Straight-Line Spherical Perspective Projection Constraint
Xianghua Ying, Zhanyi Hu, Hongbin Zha
ACCV (2)3
2006 Interpreting Sphere Images Using the Double-Contact Theorem
Xianghua Ying, Hongbin Zha
ACCV (1)2
2006 Shape Topics: A Compact Representation and New Algorithms for 3D Partial Shape Retrieval
abstract
This paper develops an efficient new method for 3D partial shape retrieval. First, a Monte Carlo sampling strategy is employed to extract local shape signatures from each 3D model. After vector quantization, these features are represented by using a bag-of-words model. The main contributions of this paper are threefold as follows: 1) a partial shape dissimilarity measure is proposed to rank shapes according to their distances to the input query, without using any timeconsuming alignment procedure; 2) by applying the probabilistic text analysis technique, a highly compact representation "Shape Topics" and accompanying algorithms are developed for efficient 3D partial shape retrieval, the mapping from "Shape Topics" to "object categories" is established using multi-class SVMs; and 3) a method for evaluating the performance of partial shape retrieval is proposed and tested. To our best knowledge, very few existing methods are able to perform well online partial shape retrieval for large 3D shape repositories. Our experimental results are expected to validate the efficacy and effectiveness of our novel approach.
Hongbin Zha, Hong Qin 0001
CVPR (2)2
2006 Riemannian Manifold Learning for Nonlinear Dimensionality Reduction
Tony Lin 0001, Hongbin Zha, Sang Uk Lee
ECCV (1)2
2006 Laser-based Interacting People Tracking Using Multi-level Observations
abstract
Laser based people tracking systems have been developed for mobile robotics and intelligent surveillance areas. Existing systems rely on simple laser point clustering methods to extract object locations. However, when dealing with multiple interacting people, laser points of different persons are often interlaced and undistinguishable due to measurement noise and they can not provide reliable features. It causes current systems quite fragile and unreliable. In this paper, we try to explore potentials from multi-level observations including weakly detected features, stably extracted features and foreground points. For inference, detection incorporated joint particle filter is used. And stably extracted features are utilized to properly estimate parameters of dynamic model for each target. In real experiments, we obtain raw data from multiple registered laser scanners, which measure two legs for each people. Evaluations with real data show that the proposed method is more robust and effective than existing approaches
Jinshi Cui, Hongbin Zha, Huijing Zhao, Ryosuke Shibasaki
IROS2
2006 A Path Planner in Changing Environments by Using W-C Nodes Mapping Coupled with Lazy Edges Evaluation
abstract
This paper presents a path planner based on PRM framework for robots operating in changing environments, in which obstacles can move randomly and robots may change their original shapes, e.g. a robot manipulator grasps an object. W-C nodes mapping coupled with lazy edges evaluation is used to ensure a generated path containing only valid nodes and edges when constructed probabilistic roadmap becomes partially invalid in changing environments. Our method combines merits of DRM and Lazy PRM methods. W-C nodes mapping, which is constructed in pre-processing phase by mapping every basic cell in workspace to nodes of roadmap, is preserved to indicate invalid nodes of roadmap fast whenever obstacles move. W-C edges mapping, which is another mapping relationship of DRM method, is skipped since it is much more complicated and time-consuming for construction. The simplified mapping with smaller size can be recomputed or modified fast when robots change their original shapes. Instead, lazy edges evaluation is used to ensure all edges valid along a found path. Simulated experiments for a dual-manipulator system show that our method is efficient for path planning in changing environments
Hong Liu 0008, Xuezhi Deng, Hongbin Zha
IROS3
2006 Registration of Multiple Laser Scans Based on 3D Contour Features
abstract
When 3D laser scanner captures range data of real scenes, one of most important problems is how to align all range data into a common coordinate system. In this paper, we propose an algorithm of registration of multiple range data from real scenes using 3D contour features. Firstly, 3D contour features are extracted using self-adaptive curve fitting, and a searching structure of octree is built from the 3D contour features. Secondly, using mahalanobis distance, the leaf nodes are matched between two scans to compute original transform matrix, and then transform matrix is refined step by step through ICP until a best transform matrix is obtained. Lastly, a new global registration strategy is given based on the nearby principle. The experiments of multiple range data registration from indoor scenes, outdoor scenes and ancient buildings are done, and the results show the proposed algorithm is robust.
Shaoxing Hu, Hongbin Zha, Aiwu Zhang
IV2
2006 Robust Mean Shift Tracking Based on Multi-Cue Integration
abstract
Color-based mean shift has been addressed as an effective and fast algorithm for tracking color blobs. This deterministic searching method suffers from low saturation color object, color clutter in backgrounds and complete occlusion for several frames. This paper proposes a direct motion-color integration method to solve the low saturation color problem and the color background clutter problem. Based on the direct cue integration, an occlusion handler that is able to deal with long term full occlusion is proposed to solve the complete occlusion problem as well. Moreover, motivated by the idea of tuning weight of each cue according to its performance, a method of adaptive multi-cue integration based mean shift is proposed. Weights of each cue are adjusted according to a quality function, which is used to evaluate the performance of each cue in the adaptive integration scheme. Extensive experiments show that this method can adapt the weight of individual cue efficiently, and increase the robustness of tracking in various conditions.
Hong Liu 0008, Hongbin Zha
SMC3
2006 The Generalized Shape Distributions for Shape Matching and Analysis
abstract
This paper presents a novel 3D shape descriptor "the generalized shape distributions" for effective shape matching and analysis, by taking advantage of both local and global shape signatures. We start this process by generating spin images on meshes. These local shape descriptors are then quantized via k-means clustering. The key contribution of this paper is to represent a global 3D shape as the spatial configuration of a set of specific local shapes. We achieve this goal by computing the distributions of the Euclidean distance of pairs of local shape clusters. Because of the spatial, sparse distribution of local shapes defined over a 3D model, an indexing data structure is adopted to reduce the space complexity of the proposed shape descriptor. The technical merits of our new approach are at least two-fold: (1) it is robust to non-trivial shape occlusions and deformations, since there are statistically a large number of chances that some local shape signatures and their spatial layouts are unchanged and users can easily identify those unchanged parts; (2) it is more discriminative than a simple collection of local shape signatures, since the spatial layouts of a global shape are explicitly computed. Our preliminary experiments have shown the effectiveness of this new approach for shape comparison and analysis
Hongbin Zha, Hong Qin 0001
SMI2
2006 Geometric Interpretations of the Relation between the Image of the Absolute Conic and Sphere Images
abstract
A spherical object has been introduced into camera calibration for several years through utilizing the properties of an image conic, which is the projection of the occluding contour of a sphere in the perspective image. However, in literature, only an algebraic interpretation was presented for the relation between the image of the absolute conic and sphere images. In this paper, we propose two geometric interpretations of this relation and two novel camera calibration methods using sphere images are derived from these geometric interpretations.
Xianghua Ying, Hongbin Zha
IEEE Trans. Pattern Anal. Mach. Intell.2
2006 Coarse-to-Fine Vision-Based Localization by Indexing Scale-Invariant Features
abstract
This paper presents a novel coarse-to-fine global localization approach inspired by object recognition and text retrieval techniques. Harris-Laplace interest points characterized by scale-invariant transformation feature descriptors are used as natural landmarks. They are indexed into two databases: a location vector space model (LVSM) and a location database. The localization process consists of two stages: coarse localization and fine localization. Coarse localization from the LVSM is fast, but not accurate enough, whereas localization from the location database using a voting algorithm is relatively slow, but more accurate. The integration of coarse and fine stages makes fast and reliable localization possible. If necessary, the localization result can be verified by epipolar geometry between the representative view in the database and the view to be localized. In addition, the localization system recovers the position of the camera by essential matrix decomposition. The localization system has been tested in indoor and outdoor environments. The results show that our approach is efficient and reliable.
Junqiu Wang, Hongbin Zha, Roberto Cipolla
IEEE Trans. Syst. Man Cybern. Part B2
2005 Linear Approaches to Camera Calibration from Sphere Images or Active Intrinsic Calibration Using Vanishing Points
abstract
Spherical objects and vanish points are often used for camera calibration. An occluding contour of a sphere is projected to a conic in the perspective image, and using a moving active camera, the trajectory of a vanishing point in the perspective images is also a conic when the camera is rotated about a fixed 3D axis whereas the translation of the camera is arbitrary. In fact, the problems of camera calibration using conics from spheres or vanishing points can be described by same mathematic representations. Two linear approaches to the problems are proposed in this paper: one based on the geometric interpretation of the relation between image conics and the image of the absolute conic, and the other using the special structure of the problems in algebra. Only three such conics are needed for the two linear approaches, and the minimum number for previous nonlinear optimization methods is also three. All five intrinsic parameters are recovered linearly without making assumptions, such as, zero-skew or unitary aspect ratio which are often used in previous methods. The two linear algorithms have been tested in extensive experiments with respect to noise sensitivity and also made comparisons with recent calibration techniques.
Xianghua Ying, Hongbin Zha
ICCV2
2005 Document Image Retrieval Based on Density Distribution Feature and Key Block Feature
abstract
Document image retrieval is an important part of many document image processing systems such as paperless office systems, digital libraries and so on. Its task is to help users find out the most similar document images from a document image database. For developing a system of document image retrieval among different resolutions, different formats document images with hybrid characters of multiple languages, a new retrieval method based on document image density distribution features and key block features is proposed in this paper. Firstly, the density distribution and key block features of a document image are defined and extracted based on documents' print-core. Secondly, the candidate document images are attained based on the density distribution features. Thirdly, to improve reliability of the retrieval results, a confirmation procedure using key block features is applied to those candidates. Experimental results on a large scale document image database, which contains 10385 document images, show that the proposed method is efficient and robust to retrieve different kinds of document images in real time.
Hong Liu 0008, Suoqian Feng, Hongbin Zha
ICDAR3
2005 Combining interest points and edges for content-based image retrieval
abstract
This paper presents a novel approach using combined features to retrieve images containing specific objects, scenes or buildings. The content of an image is characterized by two kinds of features: Harris-Laplace interest points described by the SIFT descriptor and edges described by the edge color histogram. Edges and corners contain the maximal amount of information necessary for image retrieval. The feature detection in this work is an integrated process: edges are detected directly based on the Harris function; Harris interest points are detected at several scales and Harris-Laplace interest points are found using the Laplace function. The combination of edges and interest points brings efficient feature detection and high recognition ratio to the image retrieval system. Experimental results show this system has good performance.
Junqiu Wang, Hongbin Zha, Roberto Cipolla
ICIP (3)2
2005 Camera pose determination from a single view of parallel lines
abstract
In this paper, we present a method for finding the closed form solutions to the problem of determining the pose of a camera with respect to a given set of parallel lines in 3D space from a single view, while it can not be solved by previous methods for the perspective-n-line (PnL) problem. The main idea of our method is that, firstly the distances from the optical center of the camera to these parallel lines are determined, and then the pose parameters are recovered using the obtained distances. The problem of finding these optical-center-to-line distances in fact is the degenerated perspective-n-point (PnP) problem, and we proved that there are at most two solutions for the degenerated P3P problem. An application of our method is to distinguish crosswalks and staircases aiding for the partially sighted. The method also provides a different way to investigate the problem of shape from texture.
Xianghua Ying, Hongbin Zha
ICIP (3)2
2005 Vision-based Global Localization Using a Visual Vocabulary
abstract
This paper presents a novel coarse-to-fine global localization approach that is inspired by object recognition and text retrieval techniques. Harris-Laplace interest points characterized by SIFT descriptors are used as natural landmarks. These descriptors are indexed into two databases: an inverted index and a location database. The inverted index is built based on a visual vocabulary learned from the feature descriptors. In the location database, each location is directly represented by a set of scale invariant descriptors. The localization process consists of two stages: coarse localization and fine localization. Coarse localization from the inverted index is fast but not accurate enough; whereas localization from the location database using voting algorithm is relatively slow but more accurate. The combination of coarse and fine stages makes fast and reliable localization possible. In addition, if necessary, the localization result can be verified by epipolar geometry between the representative view in database and the view to be localized. Experimental results show that our approach is efficient and reliable.
Junqiu Wang, Roberto Cipolla, Hongbin Zha
ICRA3
2005 Tracking multiple people using laser and vision
abstract
We present a novel system that aims at reliably detecting and tracking multiple people in an open area. Multiple single-row laser scanners and one video camera are utilized. Feet trajectory tracking based on registration of distance information from multiple laser scanners and visual body region tracking based on color histogram are combined in a Bayesian formulation. Results from tests in a real environment are reported to demonstrate that the system can detect and track multiple people simultaneously with reliable and real-time performance.
Jinshi Cui, Hongbin Zha, Huijing Zhao, Ryosuke Shibasaki
IROS2
2005 A planning method for safe interaction between human arms and robot manipulators
abstract
This paper presents a planning method based on mapping moving obstacles into C-space for safe interaction between human arms and robot manipulators. In pre-processing phase, a hybrid distance metric is defined to select neighboring sampled nodes in C-space to construct a roadmap. Then, two kinds of mapping are constructed to determine invalid and dangerous edges in the roadmap for each basic cell decomposed in workspace. For updating the roadmap when an obstacle is moving, basic cells covering the obstacle's surfaces are mapped into the roadmap by using new positions of the surfaces points sampled on the obstacle. In query phase, in order to predict and avoid coming collisions and reach the goal efficiently, an interaction strategy with six kinds of planning actions of searching, updating, walking, waiting, dodging and pausing are designed. Simulated experiments show that the proposed method is efficient for safe interaction between two working robot manipulators and two randomly moving human arms.
Hong Liu 0008, Xuezhi Deng, Hongbin Zha
IROS3
2005 Modeling facial expression space for recognition
abstract
In this paper, we present a method of modeling facial expression space for facial expression recognition by fuzzy integral. In traditional expression recognition methods using shape features, there are problems in describing both the uncertainty in facial expression classification and the relationship between facial features and facial expressions. Using facial expression space model, those problems can be solved easily. Firstly, we use values of fuzzy integral in different facial expression spaces to describe the uncertainty of facial expression. Secondly, by the fuzzy measure automatically constructed in each facial expression space, we deal with different effects of facial features for facial expression classification. Experiments show this method has a good ability of describing the uncertainty of facial expression and acquires good results of classification.
Hong Liu 0008, Hongbin Zha
IROS3
2005 Simultaneously calibrating catadioptric camera and detecting line features using Hough transform
abstract
A line in space is projected to a conic in a central catadioptric image, and such a conic is called a line image. This paper proposes a novel approach to calibrating catadioptric camera and detecting line images simultaneously by using Hough transform. Previous approaches to catadioptric cameras calibration employ the traditional conic detecting or fitting methods for line images, and then use these recovered conies to estimate the intrinsic parameters based on some properties of line images. However, the type of a line image can be line, circle, ellipse, hyperbola or parabola, and in general only a small arc of the conic is visible in the image, which brings novel challenges for conic detection and fitting where traditional conic detecting and fitting methods may fail. As we know, the accuracy of the estimated intrinsic parameters highly depends on the accuracy of the extracted conies. The main contribution of this work is we show that all line images from catadioptric cameras with the same intrinsic parameters must belong to a family of conies with only two degree-of-freedom, and such a family is called a line image family. Therefore, we present a novel special Hough transform for line image detection which ensures that all detected conies must belong to a line image family related to certain intrinsic parameters. For all possible values of the unknown intrinsic parameters, the line image special Hough transform are performed. The one with the highest confidence is chosen as the estimated values for these unknown intrinsic parameters, and the corresponding results of line image detection are chosen as the estimated values for line images. In order to make the searching process more efficient, the hierarchical approaches are employed in this paper. The validity of our proposed approach is illustrated by experiments.
Xianghua Ying, Hongbin Zha
IROS2
2005 Real 3D Digital Method for Large-Scale Cultural Heritage Sites
abstract
Digital preservation of cultural heritage sites has become a global problem. The paper proposes a real 3D digital method for culture heritage sites using 3D laser scanners and CCD cameras. Firstly, we preprocess the laser scans for noise removal and hole filling. Next step is using an improved ICP algorithm we present step-by-step registration to align all range scans into a common coordinate system. And then, we proposed a filtering of 3D data compression and use a volumetric-based algorithm for the construction of a coherent 3D mesh that encloses all range scans. Finally, through texture mapping, we obtain real 3D and real texture models. The example of the construction of the 3D model of buildings and grottos are presented.
Shaoxing Hu, Hongbin Zha, Aiwu Zhang
IV2
2005 Interactive Visual Retrieval System for Large Scale 3D Models Database
abstract
This paper focuses on the key algorithms and techniques for developing an interactive visual retrieval system for large scale 3D databases, and a novel 3D model retrieval and visualization engine, 3DMIRACLES, has been developed, which integrates effective algorithms and techniques for both shape-based retrieval of 3D models and real-time visualization of the retrieval results in realistic 3D interactive mode. In the retrieval system, interactive visualization for the retrieval user interface and 3D shape retrieval computation are the two most important functional modules. For interactive visualization, a novel 3D viewer has been developed, which implements hybrid rendering method to make much simplification and shortcut processing of 3D rendering computation for achieving high speed and efficient visualization of large scale database; for retrieval computation, new algorithms for 3D shape feature extraction and similarity matching have been developed and implemented.
Weibin Liu, Yusuke Uehara, Daiki Masumoto, Jiantao Pu, Hongbin Zha
MMM7
2005 Omni-directional vision based human motion detection for autonomous mobile robots
abstract
This paper presents a novel human motion detection system using an omni-camera on a platform of autonomous mobile robot. Compared with other related works, the system using an omni-directional camera provides a larger FOV (field of view) for motion detection, and makes a better effect using temporal differencing method based on compensation for ego-motion. Furthermore, the method, respectively combined with color information and feature point's on human body, is employed to determine human motion contours, which consequently improves the performance of the temporal differencing method. Experimental results show that the system provides an efficient way to track human motion in an indoor environment.
Hong Liu 0008, Hongbin Zha
SMC3
2004 Interactive Rendering with LOD Control and Occlusion Culling Based on Polygon Hierarchies
abstract
This work presents a new method of combining dynamic control of LOD and conservative occlusion culling based on a new hierarchical data structure of polygons. Our method is effective for rendering a large amount of data in complex environments.
Tokuo Tsuji, Hongbin Zha, Ryo Kurazume, Tsutomu Hasegawa
Computer Graphics International2
2004 3D model based head pose tracking by using weighted depth and brightness constraints
abstract
This paper proposes a robust method of tracking human head poses from a sequence of monocular images. First we estimate the head pose parameters in the first frame by an affine correspondence based method developed in our lab. Then both the linear brightness and depth constraint equations derived from the small interframe rigid motion assumption are used to implement the fast tracking of the head poses. We also take advantage of geometry information of the features on the face surface to weight the brightness and depth constraint equations to get more accurate results. Finally, in order to diminish the effects of gradual illumination changes and occlusions, we estimate the reliability of the features frame by frame and dynamically update the reliable feature set. Experiments show the proposed method can robustly track the head poses especially for the types of motions which make obvious depth variation.
Guoyuan Liang, Hongbin Zha, Hong Liu 0008
ICIG2
2004 Tissue map based craniofacial reconstruction and facial deformation using RBF network
abstract
In this paper we present a novel craniofacial reconstruction method employing statistical tissue thickness information. The tissue thickness data gotten from CT images are represented as 2D tissue maps. The input (target) skull model is parameterized onto a 2D planar map and the landmarks are utilized to train a RBFN (radial basis function network), which realizes warping of planar maps between the target tissue and the generic tissue. The generic tissue is aligned onto the target skull by applying the trained network onto it, and thus the target facial map can be obtained by a simple addition of the warped generic maps. Finally, we interactively deform the model based on a RBFN to make the facial meshes more personalized, and map the texture from orthogonal photos onto the reconstructed model to improve rendering effects. Experiment results show that the proposed approach is helpful in improving the recognition ability in forensic applications.
Yuru Pei, Hongbin Zha, Zhongbiao Yuan
ICIG2
2004 Efficient View-dependent LOD Control for Large 3D Unclosed Mesh Models of Environments
abstract
The rapid progress of 3D modeling techniques have brought more and more applications of 3D range data in robotics. However, the 3D data we acquire with range finders mounted on mobile robots are usually too large to process in real-time, and holes and boundaries are unavoidable on the surface it represents. One method to resolve this problem is to perform mesh simplification and level of detail (LOD) control on these models. A new method based on progressive meshes (PMs) is proposed in this paper, which adopts different simplification and remeshing strategies on different boundary conditions, and therefore well preserves geometrical features on the boundaries through the simplification. The detail records are organized in a forest-shaped data structure at the same time. Hence, when a scene is reconstructed using the multiresolution models, the local details can be rapidly added by selective refinement. Furthermore, continuous view-dependent distributions of LOD are also supported to provide efficient scene representation and rendering.
Jie Feng 0012, Hongbin Zha
ICRA2
2004 A Robust Method for Shape-Based 3D Model Retrieval
abstract
Proliferation of 3D models necessitates developing efficient methods for indexing or retrieving the models in a large database. Many previous methods for this purpose defined functions on concentric spheres as approximation of 3D geometry for spherical harmonic transform (SHT). In this paper, we point out that this is not robust as the surface of a model may shift between different shells under perturbation, and multi-layer of surfaces may exist in one shell, making the function definition ambiguous. To solve these problems, we propose a method to characterize 3D shape using delta functions. Then, spherical functions are defined by sampling in the frequency domain of the delta functions for SHT. By doing so, our method can support retrieval with controllable acuity, which benefits wider range of applications and facilitates customization to different users. Experiments have shown that our method is more robust than previous approaches.
Jiantao Pu, Guyu Xin, Hongbin Zha, Weibin Liu, Yusuke Uehara
PG4
2003 Mobile robot localization with an incomplete map in non-stationary environments
abstract
One of the fundamental problems of the mobile robots is self-localization, i.e., to estimate the self-position by comparing sensor data and a map. In non-stationary environments, a robot should avoid to use changed objects as landmarks in the localization. However, in most previous localization methods, it is assumed that there is no change, or changes are easily identified by sensing. In this paper, we propose a self-localization method that is robust against changes in environments. The method identifies changes from noisy and ambiguous sensor data. Since an object with a random shape may be added at a random position, it generates and utilizes multiple hypotheses about the changes. A number of simulation experiments have been performed in various environments, to demonstrate the effectiveness of the method.
Kanji Tanaka 0001, Tsutomu Hasegawa, Hongbin Zha, Eiji Kondo, Nobuhiro Okada
ICRA3
2003 Viewpoint planning in map updating task for improving utility of a map
abstract
Many mobile robots use maps for path planning. However, due to changes of obstacle configurations in the environment, the utility of the map is often deteriorated. In this paper, we address a problem of map updating task, where a robot automatically explores the environment to update a map. This task is different from well known map making task in that the robot can estimate global structure of the environment by using the map of a past environment. This enables a map updating robot effective viewpoint planning. In this paper, we propose an efficient method for motion planning in the task, in which utility of the map for its users can be taken into account.
Kanji Tanaka 0001, Hongbin Zha, Tsutomu Hasegawa
IROS2
2003 Special issue on 3-D image analysis and modeling
Hongbin Zha, Hideo Saito 0001, Vittorio Murino, Andrea Fusiello
IEEE Trans. Syst. Man Cybern. Part B1
2001 Reconstruction of surfaces from medical slices using a multi-scale strategy
abstract
In this paper, we propose a method of reconstructing triangular surfaces from given medical slices. The method tries to solve the problem by using contours interactively extracted from the slices. After that the problem changes to connecting the contours into a triangular mesh. For the purpose, the algorithm has to deal with two difficult problems: contour correspondence and branching. The correspondence problem is solved here by exploiting information both of the extracted contours and the original slices. Then, we triangulate the contours by a piecewise-linear interpolation scheme with ability to handle degenerate portions in the mesh. The most significant advantage in integrating the grey-level slice information is that we can match the contours more correctly to find reliable contour correspondence for modeling complex anatomical structures.
Jinguo He, Hongbin Zha, Qingyun Shi
SMC2
2001 A new subdivision method for modeling 3D objects with significant discontinuities
abstract
The authors present a new subdivision method for modeling 3D objects with significant discontinuities. We employ a differential operator to detect significant feature curves, and then use a combined interpolatory subdivision algorithm to refine the initial control mesh. The limit surface of the combined subdivision scheme is C/sup 1/ smooth, except near the detected feature curves where it is C/sup 0/ continuous.
Hongbin Zha, Xingzhou Yu
SMC1
2000 3-D Shapes Modeling which has Hierarchical Structure Based on B-Spline Surfaces with Non-Uniform Knots
abstract
A method for 3D shapes modeling which has hierarchical structure is proposed. For representing 3D surface shapes, 4th order non-uniform rational B-spline functions with controllable knots are employed as the surface model. Consequently, a multiresolution representation technique based on multiresolution wavelet transform can be implemented with the corresponding B-wavelets. The surface models at each resolution level are obtained by performing a decomposition algorithm, after the surface model at the highest level is estimated. In order to estimate the surface model as accurately as possible, a regularization problem is solved by an iterative algorithm. Through several experiments using real data of range images, the effectiveness of the proposed method was confirmed.
Makoto Maeda, Kousuke Kumamaru, Hongbin Zha, Katsuhiro Inoue
ICPR3
2000 Dynamic Control of Mesh LODs (Levels of Detail) by Using a Multiresolution Mesh Data Structure
abstract
We present a new method for controlling levels of detail (LODs) of triangulated mesh objects by a gaze-driven approach. The method relies upon a hierarchical mesh representation, called a merge-and-split tree of vertices, in which objects are described at nearly continuous LODs. The basic idea here is to adjust the mesh shape quality adaptively under a geometrical definition of gaze points and fovea regions. Not only does it generate a circular fovea region with smooth transition of LODs, but it also makes it possible to update the whole mesh quickly for a moving gaze.
Hongbin Zha, Yoshinobu Makimoto, Tsutomu Hasegawa
ICPR1
2000 Topology-adaptive modeling of objects by using variable-size ball marching
abstract
The paper proposes a novel method of modeling multiobject scenes or objects of arbitrary topology by a marching method based on mathematical morphology operations with a variable-size structuring element. We assume that 3D data on the whole object surface has been acquired. A narrow-band stopping region is built by merging the neighborhood of every data point. A single close-surface surrounding the data point set is given as an initial shape. The surface shrinks and splits smoothly by repeated erosions with a ball structuring element, whose size increases continuously as the curvature of the surface increases. Every surface point stops moving if it lies in the stopping region. After this surface fitting to the stopping region, the region is narrowed, and the fitting-and-narrowing cycle is repeated until the model is consistent with the object surface on the basis of a topological check. After the above rough estimation, for model refinement, we extract a quadrangular mesh from the resultant surface, and then perform precise fitting to the 3D data by using an energy minimisation.
Kenji Hara, Hongbin Zha, Tsutomu Hasegawa
SMC2
2000 Registration of range images with different scanning resolutions
abstract
The paper presents a novel method of registering two range images that are taken at different 3D viewpoints. Distinct from classic methods, it allows an accurate estimation of a scaling parameter of the images in addition to the general motion parameters. It is made possible owing to the following efforts: 1) extended signature images (ESIs) are utilized to establish precise correspondence between image objects even if they are different in scale; 2) fine image alignment is achieved by incorporating scale estimates into a modified iterative closest point (ICP) algorithm. Results of experiments demonstrate the method is of practical use even for images acquired by real digitizers.
Hongbin Zha, Makoto Ikuta, Tsutomu Hasegawa
SMC1
2000 Fast multiresolution modeling of 3D objects using mesh-based wavelet analysis
abstract
We propose a method of generating multiresolution triangulated meshes from range data clouds for modeling 3D objects. Such meshes can be obtained by applying wavelet analysis techniques directly to a fine mesh if it is equivalent in topology to a tetrahedron subdivided by a recursive 4-to-1 splitting algorithm. However, a mesh of this special characteristic is difficult to construct by a traditional meshing method, and hence another time-consuming remeshing process has to be used. To solve the problem, we try to get a topologically proper coarse mesh automatically from any given dataset. Then, finer meshes are produced by applying the 4-to-1 subdivision algorithm recursively, resulting in multiresolution meshes in an efficient manner.
Hongbin Zha, Tomoo Mitsutomi, Tsutomu Hasegawa
SMC1
1999 Topology-Adaptive Modeling of Objects by a Level Set Method with Multi-Level Stopping Conditions
abstract
Level set methods were proposed mainly by mathematicians for extracting bounding contours of multiple regions or shapes with holes. In the paper, we propose a new method of constructing a surface/solid model of a 3-D object of arbitrary topology on the basis of the approach. Given sensor data covering the whole object surface, the method begins with an initial approximation of the object by evolving a closed-surface into a model typologically equivalent to the real object. The refined approximation is then performed by deforming it at an optimally tuned speed. By numerically solving the level set equations for tracking the evolving surfaces, it becomes possible to handle 3-D objects of arbitrary topology while maintaining high accuracy in the shape fitting.
Kenji Hara, Hongbin Zha, Tsutomu Hasegawa
ICIP (4)2
1999 Observation planning for map updating tasks by predicting changes in environments
abstract
For robots working in a dynamic environment, there is a great need for a global environment map that should be updated timely. The paper presents an observation planning method for the map updating purpose. The task assigned to the robot is to detect changes of environment at planned observation positions and update the map into the newest version. After the observation at a current position, the robot gets quantitative estimations on the changes that may occur at possible sequences of future observations. Then, it evaluates the sequences for the purpose of map updating and selects the most favored next position as the planned one. Efforts are made in searching the optimal position by a dynamic programming algorithm with constraints from previous searching results.
Kanji Tanaka 0001, Hongbin Zha, Tsutomu Hasegawa
IROS2
1999 Repositioning planning of autonomous arms for an unstable pushing task
abstract
Distributed robot systems using autonomous manipulators (arms) are widely employed in complex tasks such as object pushing, handling, tumbling, or combinations of them. In such cases, the problem of arm reconfiguration is encountered if the object has to be moved into another orientation. In the paper, we describe an object-pushing robot system in which emphasis is placed on efficient planning on the changing of pushing positions, which we refer to as arm repositioning. A new algorithm is proposed for the repositioning mainly based on evaluation of the object stability and load distribution. The applicability of the algorithm is illustrated by results of experiments using simulated or real arm systems.
Hongbin Zha, Hiroki Nagahama, Tsutomu Hasegawa
IROS1
1998 A Recursive Fitting-and-Splitting Algorithm for 3-D Object Modeling Based on Superquadrics
Hongbin Zha, Tsuyoshi Hoshide, Tsutomu Hasegawa
ACCV (1)1
1998 Next Best Viewpoint (NBV) Planning for Active Object Modeling Based on a Learning-by-Showing Approach
Hongbin Zha, Ken'ichi Morooka, Tsutomu Hasegawa
ACCV (2)1
1998 Regularization-based 3D object modeling from multiple range images
abstract
Several methods have been proposed for building 3D object models from multiple range images mainly on the basis of the deformable models and pseudo-dynamics. However, the patches are usually placed crossing over the edge regions, and the surface orientation discontinuities cannot be properly preserved as a result. This paper proposes a new 3D object modeling method that generates refined triangular meshes while preserving the surface discontinuities. We formulate the process into a problem of minimizing an energy functional by using the idea of line processes and Markov random fields. Moreover, we show that near-optimal solutions can be obtained by utilizing the continuation method.
Kenji Hara, Hongbin Zha, Tsutomu Hasegawa
ICPR2
1998 Surface recovery by using regularization theory and its application to multiresolution analysis
abstract
In this paper, a surface recovery method using multiresolution wavelet transform is proposed. For representing 3D surface shapes, 4th order B-spline functions with uniform knots are introduced as scaling functions of spline wavelets. In order to estimate the surface function, a regularization problem is solved by an iterative algorithm. The estimated surface function can, be decomposed into an approximate surface function at the lowest resolution and the corresponding wavelet components. Consequently, by reducing the noise components which the wavelet components include, the surface recovery method can give an accurate estimation of the surface function. Through several experiments, both the robustness to noises and the edge-preserving property in recovering the surface have been confirmed.
Makoto Maeda, Kousuke Kumamaru, Katsuhiro Inoue, Hongbin Zha
ICPR4
1998 Next best viewpoint (NBV) planning for active object modeling based on a learning-by-showing approach
abstract
The paper presents a method of creating a complete model of a curved object from a sequence of range images acquired by a fixed range finder. To accomplish the modeling fast and accurately in an optimal manner, we propose a new online viewpoint planning algorithm to choose the next best viewpoint (NBV) based on the already obtained partial model. The NBV is determined by evaluating factors such as possibility of merging new data, local shape changes, registration accuracy and control point distribution.
Ken'ichi Morooka, Hongbin Zha, Tsutomu Hasegawa
ICPR2
1998 A recursive fitting-and-splitting algorithm for 3-D object modeling using superquadrics
abstract
The paper proposes a new approach to 3D object modeling by integrating superquadric-fitting and segmentation into a top-down, recursive algorithm. Given sensor data, which are a set of multiview range data covering the whole object surface, the method begins with an initial approximation of the object by fitting a single superquadric. The fitting residuals are then examined to pick up data points either in deep concave regions or too far away from the approximating surface. A dividing plane is extracted from the points to partition the original data set into two disjoint subsets, which are further treated respectively with the same fitting-and-splitting scheme. This process is repeated until the whole data are decomposed into primitive superquadrics all within some preset error tolerance.
Hongbin Zha, Tsuyoshi Hoshide, Tsutomu Hasegawa
ICPR1
1998 Multi-resolution surface description of 3D objects by shape-adaptive triangular meshes
abstract
Automatic 3D object modeling is a problem to be solved for many vision applications. In this paper, we propose a method of constructing a B-rep model of a 3D object from multi-view range images by utilizing a dynamic balloon scheme. The model is represented with triangular meshes, and by dynamic subdivision of the triangles, the meshes can approximate the local shapes adaptively. Moreover, it is also possible to derive a hierarchical mesh representation, which is controlled by a sequence of prescribed error tolerances. Results of experiments using real range images are presented to show the practical feasibility of the method.
Hongbin Zha, Soshiro Tahira, Tsutomu Hasegawa
ICPR1
1998 3-dimensional object model construction from range images taken by a range finder on a mobile robot
abstract
To construct a 3D model of an object in a computer, a range finder is used to obtain range images of the object. The range finder takes range images from many view points to eliminate the influence of occlusions. A registration algorithm is, then, used to merge the range images. When the range finder is mounted on a mobile robot and measures while moving, range images taken by it will be distorted. We extended the registration algorithm to remove the deformation. The extended algorithm can remove it and is used to construct a 3D model even for a case where the speed parameters of the robot are not well known.
Nobuhiro Okada, Hongbin Zha, Tadashi Nagata, Eiji Kondo, Ken'ichi Morooka
IROS2
1998 Planning observation positions for a mobile robot to update incomplete maps of indoor environments
abstract
For mobile robots working in a dynamic environment, there is a great need for a global environment map that should be updated timely. In the paper, we propose methods for the map updating purpose by using a specialized mobile robot. The map updating system is composed of two basic processing components: local change detection and global observation planning. Principal strategies for the observation position planning are developed making good use of knowledge both from an outmoded initial map and from up-to-date sensory information. The paper gives a detailed description of the new method: moving around object clusters. Simulation results show this method is efficient even for detecting large changes in a complex environment.
Kanji Tanaka 0001, Hongbin Zha, Tsutomu Hasegawa
IROS2
1998 Surface reconstruction in overlapping range images for generating close-surface 3-D object models
abstract
The paper addresses the problem of building close-surface (CS) models of 3-D objects by integrating multi-view range images of the object. For the purpose, we have to use overlapping regions between the images to register them into a common view-independent coordinate system. In this paper, we propose a method of removing redundancy in the registered multi-view data due to the image overlap. It is assumed that each image has been triangulated into a connected mesh after the image acquisition. We firstly remove triangle edges in the overlapping regions, and then reconstruct the regions by forcing both border lines to advance toward each other in an optimal retriangulation procedure.
Hongbin Zha, Y. Malkimoto, Tsutomu Hasegawa
SMC1
1998 Distributed multi-arm systems for complex pushing tasks
abstract
Distributed robot systems using autonomous manipulators are widely employed in complex tasks such as object pushing, handling, tumbling, or combinations of them. In such cases, the problem of arm reconfiguration is encountered if the object has to be moved into another orientation. In the paper, we describe an object-pushing robot system in which emphasis is placed on efficient planning on the changing of pushing positions, which we call arm repositioning. A new algorithm is proposed for the repositioning mainly based on an evaluation of the object stability and load distribution. The applicability of the algorithm is illustrated by results of experiments using simulated or real arm systems.
Hongbin Zha, Hiroki Nagahama, Tsutomu Hasegawa
SMC1
1998 Recognizing 3-D objects by using a Hopfield-style optimization algorithm for matching patch-based descriptions
Hongbin Zha, Hideki Nanamegi, Tadashi Nagata
Pattern Recognit.1
1997 Detecting changes in a dynamic environment for updating its maps by using a mobile robot
abstract
For mobile robots working in a dynamic environment, there is a great need for a global environment map that is updated frequently. In the paper, we propose a new method by which a mobile robot with an original environment map automatically detects changes in the environment and modifies the map into a new version. The changes are identified by performing robot localization and change detection alternatively at some planned observation positions. To deal with robot position errors caused by dead reckoning, available information on the original map and a local map recording already detected changes is utilized as much as possible. Simulation results show this method is efficient even for detecting large changes in a complex environment.
Hongbin Zha, Kanji Tanaka 0001, Tsutomu Hasegawa
IROS1
1996 3-D object recognition from range images by using a model-based Hopfield-style matching algorithm
abstract
A new method is proposed for recognizing 3-D objects by using a Hopfield-style optimization algorithm based on matching between patch-based image and model descriptions. To obtain the image descriptions, range images are employed to extract reliable high-level patch features. In the optimization process the objective function is a Liapunov function that is minimized in a Hopfield net with its interconnections encoding the imposed geometrical constraints on the model descriptions. At first, the paper describes a pre-processing method for deriving the necessary image description. It then presents the structure of the used Hopfield net which is able to recognize multiple objects all at once.
Hongbin Zha, Hideki Nanamegi, Tadashi Nagata
ICPR1
1995 Cooperative manipulations based on genetic algorithms using contact information
abstract
This paper presents a new method for cooperative manipulations based on genetic algorithms using contact information. In order to apply genetic algorithms, each robot arm generates new information about object postures using its own contact information without using of any vision systems. The object postures are derived by calculating normal vectors of object surfaces. The system proposed on the basis of genetic algorithms determines forces which robot arms should apply to the objects by solving an optimization problem utilizing the generated and obtained information. Owing to genetic algorithms, the system has not only the adaptation to the change of the working environment, but also can determine the forces for robot arms without physical inconsistency.
Tadashi Nagata, Kosuke Konishi, Hongbin Zha
IROS (2)3
1992 Quantifying saliency of feature points on 3-D curved surfaces from range images
abstract
A method for quantifying saliency of 3D feature points from range images of model surfaces is proposed. The saliency is characterized by the discriminating power of the feature points for distinguishing the model objects they belong to from others. It is measured by local shape similarity, and the similarity coefficients of surface points are calculated by parameterizing local surface representations into the local orientation coordinate systems. An algorithm for performing the Hough transform weighted by the similarity coefficients is developed to optimally determine the saliency coefficients of the feature points. It is shown that the method is especially useful for dealing with the occlusion problem that must be solved when constructing a flexible robot vision system.>
Hongbin Zha, Tadashi Nagata, Kousuke Kumamaru
ICRA1
1991 Recognizing and locating a known object from multiple images
abstract
An approach to recognizing and locating a partially visible object from multiple images for a pile of parts is proposed. The image-to-model correspondence is established through an examination of the consistency between the surface patches extracted from the images and the patches described in an object model. To obtain the scene description, which is composed of parameterized surface patches, the scene's needle map is derived by a modified photometric stereo method. The needle map is then segmented into primitive surfaces, and the resultant patches are identified and described by using a modified Hough transformation based on the Gaussian spherical maps of the patches. To interpret the scene description, the object model is built with a frame structure to describe geometric features of the model patches. By a search process guided by some heuristic planning strategies based on the fuzzy-set concept, sets of patches considered to be possibly on the surface of the model object are extracted from the scene description. The candidate sets are tested and the locations and orientations of the corresponding object instances are computed. Results of two experiments are reported.>
Tadashi Nagata, Hongbin Zha
IEEE Trans. Robotics Autom.2
1988 Determining orientation, location and size of primitive surfaces by a modified hough transformation technique
Tadashi Nagata, Hongbin Zha
Pattern Recognit.2