EDBT 2026 Demo / reviewers in the wild / expert
Zhanyi Hu
dblp:35/6476
· DBLP profile ↗
133ranked-venue papers
18as first author
21since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 87 · 10 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 59 · 7 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 23 · 6 first-author · 6 since 2021Systems, architecture and hardware · 4 · 2 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FILTER: A Framework for Defending Against Backdoor Attacks in Vertical Federated LearningabstractVertical Federated Learning (VFL) is a distributed machine learning paradigm in which participants train models with vertically partitioned data. Many previous studies have identified backdoor vulnerabilities in VFL systems. However, limited effort has been devoted to developing defenses against such attacks. Unlike centralized machine learning or horizontal FL, VFL poses new challenges for defending against backdoor attacks, particularly because the central server lacks control over the entire model. In this paper, we first explore defenses against backdoor attacks in VFL when the attacker possesses sufficient knowledge of the label information. Specifically, we propose FILTER, a framework for defending against backdoor attacks in VFL to ensure the integrity of VFL systems during training in the presence of malicious participants. To address backdoor risks in VFL, it incorporates two novel filters: an embedding-based filter and a loss-based filter, which effectively identify and remove poisoned samples in later stages of training. Through extensive experiments on five benchmark datasets against four state-of-the-art backdoor attacks, we demonstrate that FILTER significantly reduces the success rate of attacks while maintaining accuracy on clean data close to that of the models trained without such defenses. Zhanyi Hu, Cen Chen 0001, Yanhao Wang 0001 |
AAAI | 1 |
| 2026 | Active learning-based structure parsing of ancient Chinese architectures: a benchmark dataset and baseline
Wei Wang 0347, Yixing Wang, Qiulei Dong, Zhanyi Hu |
Expert Syst. Appl. | 6 |
| 2026 | AdaArc: Adaptive Weighted Bidirectional Angular Margin Loss for Single-Stage Large-Scale Visual Place Recognition
Wei Gao 0014, Zhanyi Hu |
Mach. Learn. | 3 |
| 2025 | Bad-PFL: Exploiting Backdoor Attacks against Personalized Federated LearningabstractData heterogeneity and backdoor attacks rank among the most significant challenges facing federated learning (FL). For data heterogeneity, personalized federated learning (PFL) enables each client to maintain a private personalized model to cater to client-specific knowledge. Meanwhile, vanilla FL has proven vulnerable to backdoor attacks. However, recent advancements in PFL community have demonstrated a potential immunity against such attacks. This paper explores this intersection further, revealing that existing federated backdoor attacks fail in PFL because backdoors about manually designed triggers struggle to survive in personalized models. To tackle this, we degisn Bad-PFL, which employs features from natural data as our trigger. As long as the model is trained on natural data, it inevitably embeds the backdoor associated with our trigger, ensuring its longevity in personalized models. Moreover, our trigger undergoes mutual reinforcement training with the model, further solidifying the backdoor's durability and enhancing attack effectiveness. The large-scale experiments across three benchmark datasets demonstrate the superior performance of Bad-PFL against various PFL methods, even when equipped with state-of-the-art defense mechanisms. Mingyuan Fan 0003, Zhanyi Hu, Fuyi Wang, Cen Chen 0001 |
ICLR | 2 |
| 2024 | Efficient Region-Based 3-D Urban Building Reconstruction From TomoSAR ImagesabstractThe tomographic synthetic aperture radar (TomoSAR) technique has been gaining attention because it can retrieve the 3-D structures of urban buildings by synthesizing apertures along the elevation direction. However, most existing TomoSAR methods in literature calculate elevations pixel by pixel and overlook the correlation between elevations, leading to low accuracy and efficiency. To solve these problems, this study introduces an efficient region-based 3-D urban building reconstruction method that incorporates different geometric primitives (i.e., points, planes, and models). Specifically, the proposed method, under the constraints constructed by different geometric primitives, follows three steps to reconstruct the box-like models of urban buildings: 1) it detects double-bounce regions and reconstructs box-like models based on plane sweeping and region growing; 2) it reconstructs box-like models based on multiplane fitting and optimization for nondouble-bounce regions; and 3) it regularizes box-like models based on building layout priors (e.g., collinearity and proximity). The experimental results on two datasets show that the proposed method can efficiently produce reliable results and outperforms several existing methods both qualitatively and quantitatively. Wei Wang 0347, Liankun Yu, Qiulei Dong, Zhanyi Hu |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2023 | ASPPR: active single-image piecewise planar 3D reconstruction based on geometric priors
Wei Wang 0347, Qiulei Dong, Zhanyi Hu |
Sci. China Inf. Sci. | 3 |
| 2023 | Interactive piecewise planar building reconstruction from a single image based on geometric priors
Wei Wang 0347, Qiulei Dong, Zhanyi Hu |
Expert Syst. Appl. | 3 |
| 2022 | Superpoint-guided Semi-supervised Semantic Segmentation of 3D Point Cloudsabstract3D point cloud semantic segmentation is a challenging topic in the computer vision field. Most of the existing methods in literature require a large amount of fully labeled training data, but it is extremely time-consuming to obtain these training data by manually labeling massive point clouds. Addressing this problem, we propose a superpoint-guided semi-supervised segmentation network for 3D point clouds, which jointly utilizes a small portion of labeled scene point clouds and a large number of unlabeled point clouds for network training. The proposed network is iteratively updated with its predicted pseudo labels, where a superpoint generation module is introduced for extracting superpoints from 3D point clouds, and a pseudo-label optimization module is explored for automatically assigning pseudo labels to the unlabeled points under the constraint of the extracted superpoints. Additionally, there are some 3D points without pseudo-label supervision. We propose an edge prediction module to constrain features of edge points. A superpoint feature aggregation module and a superpoint feature consistency loss function are introduced to smooth superpoint features. Extensive experimental results on two 3D public datasets demonstrate that our method can achieve better performance than several state-of-the-art point cloud segmentation networks and several popular semi-supervised segmentation methods with few labeled scenes. Shuang Deng, Qiulei Dong, Bo Liu 0035, Zhanyi Hu |
ICRA | 4 |
| 2022 | Robust observer design and Nash-game-driven fuzzy optimization for uncertain dynamical systems
Zhanyi Hu, Jin Huang 0002, Zeyu Yang 0002 |
Fuzzy Sets Syst. | 1 |
| 2022 | Design and experimental validation of event-triggered multi-vehicle cooperation in conflicting scenariosabstractPlatoon control is widely studied for coordinating connected and automated vehicles (CAVs) on highways due to its potential for improving traffic throughput and road safety. Inspired by platoon control, the cooperation of multiple CAVs in conflicting scenarios can be greatly simplified by virtual platooning. Vehicle-to-vehicle communication is an essential ingredient in virtual platoon systems. Massive data transmission with limited communication resources incurs inevitable imperfections such as transmission delay and dropped packets. As a result, unnecessary transmission needs to be avoided to establish a reliable wireless network. To this end, an event-triggered robust control method is developed to reduce the use of communication resources while ensuring the stability of the virtual platoon system with time-varying uncertainty. The uniform boundedness, uniform ultimate boundedness, and string stability of the closed-loop system are analytically proved. As for the triggering condition, the uncertainty of the boundary information is considered, so that the threshold can be estimated more reasonably. Simulation and experimental results verify that the proposed method can greatly reduce data transmission while creating multi-vehicle cooperation. The threshold affects the tracking ability and communication burden, and hence an optimization framework for choosing the threshold is worth exploring in future research. Zhanyi Hu, Yingjun Qiao, Jin Huang 0002, Yi-fan Jia 0001 |
Frontiers Inf. Technol. Electron. Eng. | 1 |
| 2022 | Gated Feature Aggregation for Height Estimation From Single Aerial ImagesabstractHeight estimation from single images, strictly speaking, is an ill-posed problem. However, recently, it is shown that it is both possible and feasible to learn a mapping from image statistics to height information. In spite of recent efforts in this field, how to learn fine-shape preserving features, such as object boundaries and contours, is still an open issue. In this work, we propose a progressive learning network to estimate height information from single aerial images in a coarse-to-fine manner. In particular, a gated feature aggregation module is introduced to effectively combine low-level and high-level features. The proposed method is validated on three public datasets, including the Vaihingen dataset, the Potsdam dataset, and the DFC2019 dataset. Both quantitative and qualitative experimental results demonstrate that the proposed method can achieve more accurate height estimation from single aerial images, especially with better object boundary and contour preserving capability, than four related height estimation methods. Qiulei Dong, Zhanyi Hu |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2022 | Active Learning Based 3D Semantic Labeling From Images and Videosabstract3D semantic segmentation is one of the most fundamental problems for 3D scene understanding and has attracted much attention in the field of computer vision. In this paper, we propose an active learning based 3D semantic labeling method for large-scale 3D mesh model generated from images or videos. Taking as input a 3D mesh model reconstructed from the image based 3D modeling system, coupled with the calibrated images, our method outputs a fine 3D semantic mesh model in which each facet is assigned a semantic label. There are three major steps in our framework: 2D semantic segmentation, 2D-3D semantic fusion, and batch image selection. A limited annotation image set is first used to fine-tune a pre-trained semantic segmentation network for obtaining the pixel-wise semantic probability maps. Then all these maps are back-projected into 3D space and fused on the 3D mesh model using Markov Random Field optimization, thus yield a preliminary 3D semantic mesh model and a heat model showing each facet’s confidence. This 3D semantic model is used as a reliable supervisor to select the parts that are not well segmented for manual annotation to boost the performance of the 2D semantic segmentation network, as well as the 3D mesh labeling, in the next iteration. This Training-Fusion-Selection process continues until the label assignment of the 3D mesh model becomes steady. By this means, we significantly reduce the amount for annotation but not the labeling quality of 3D semantic models. Extensive experiments demonstrate the effectiveness and generalization ability of our method on a wide variety of datasets. Mengqi Rong, Hainan Cui, Zhanyi Hu, Hanqing Jiang, Hongmin Liu 0001, Shuhan Shen |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2022 | Robust Camera Translation Estimation via Rank EnforcementabstractCamera translation averaging, aiming to recover the global camera locations from a given set of camera translation directions, is a challenging problem for Structure from Motion (SfM) in the field of computer vision, largely due to the fact that the given relative translation directions from a set of noisy essential matrices are generally of low accuracy. To tackle this problem, we first reveal a novel but a simple property of the camera translation matrix consisting of all the pairwise camera translations among an arbitrary set of cameras that the rank of this translation matrix is always smaller or equal to 4. Then, by explicitly enforcing this rank property, a novel translation estimation method for computing global camera locations is proposed, called TERE. Moreover, to further improve the performances of the explored TERE in the two aspects of accuracy and speed, an iterative batch-based translation estimation method is proposed, called B-TERE, where a small-scale batch of cameras is selected without replacement from the given set of cameras according to a simple camera selection strategy at each iterative step, and the locations of the selected cameras are estimated by the proposed TERE accordingly. Extensive experimental results on various datasets demonstrate that our proposed methods could achieve better performances in comparison to several state-of-the-art methods. Qiulei Dong, Xiang Gao 0009, Hainan Cui, Zhanyi Hu |
IEEE Trans. Cybern. | 4 |
| 2022 | Pursuing 3-D Scene Structures With Optical Satellite Images From Affine Reconstruction to Euclidean ReconstructionabstractHow to use multiple optical satellite images to recover the 3D scene structure is a challenging and important problem in the remote sensing field. Most existing methods in literature have been explored based on the classical RPC (Rational Polynomial Coefficients) camera model which requires at least 39 GCPs (ground control points), however, it is non-trivial to obtain such a large number of GCPs in many real scenes. Addressing this problem, we propose a hierarchical reconstruction framework based on multiple optical satellite images, which needs only 4 GCPs to fully-automated reconstruct the 3D scene structure. The proposed framework is independent from the RPC model and composed of a dense affine reconstruction stage and a followed affine-to-Euclidean upgrading stage: At the dense affine reconstruction stage, a dense affine reconstruction approach is explored for pursuing the 3D affine scene structure without any GCP from input satellite images. Then at the affine-to-Euclidean upgrading stage, the obtained 3D affine structure is upgraded to a Euclidean one with 4 GCPs. Experimental results on two public datasets demonstrate that the proposed method significantly outperforms several state-of-the-art methods in most cases. Pinhe Wang, Limin Shi, Bao Chen, Zhanyi Hu, Jianzhong Qiao, Qiulei Dong |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | SAR-to-Optical Image Translation With Hierarchical Latent FeaturesabstractDue to the all-weather and all-time imaging capability of Synthetic Aperture Radar (SAR), SAR remote sensing analysis has attracted much attention recently. However, compared with optical images, SAR images are more difficult to be interpreted. If a SAR image could be translated into its corresponding optical image, then the generated optical image would be helpful for assisting the interpretation. Addressing this issue, we investigate how to translate SAR images to optical ones in this work, and propose a parallel generative adversarial model for SAR-to-optical image translation, called Parallel-GAN, consisting of a backbone image translation sub-network and an adjoint optical image reconstruction sub-network. Under the proposed model, the backbone image translation sub-network is designed to translate SAR images to optical ones, and simultaneously some of its intermediate layers are required to output similar latent features to those from the corresponding layers of the adjoint image reconstruction sub-network. Thanks to the imposed hierarchical latent optical features, the proposed Parallel-GAN could achieve the SAR-to-optical image translation effectively. Extensive experimental results on three public datasets demonstrate that the proposed method outperforms ten state-of-the-art methods for SAR-to-optical image translation. Haixia Wang 0002, Zhanyi Hu, Qiulei Dong |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2021 | Hardness Sampling for Self-Training Based Transductive Zero-Shot LearningabstractTransductive zero-shot learning (T-ZSL) which could alleviate the domain shift problem in existing ZSL works, has received much attention recently. However, an open problem in T-ZSL: how to effectively make use of unseen-class samples for training, still remains. Addressing this problem, we first empirically analyze the roles of unseen-class samples with different degrees of hardness in the training process based on the uneven prediction phenomenon found in many ZSL methods, resulting in three observations. Then, we propose two hardness sampling approaches for selecting a subset of diverse and hard samples from a given unseen-class dataset according to these observations. The first one identifies the samples based on the class-level frequency of the model predictions while the second enhances the former by normalizing the class frequency via an approximate class prior estimated by an explored prior estimation algorithm. Finally, we design a new Self-Training framework with Hardness Sampling for T-ZSL, called STHS, where an arbitrary inductive ZSL method could be seamlessly embedded and it is iteratively trained with unseen-class samples selected by the hardness sampling approach. We introduce two typical ZSL methods into the STHS framework and extensive experiments demonstrate that the derived T-ZSL methods outperform many state-of-the-art methods on three public benchmarks. Besides, we note that the unseen-class dataset is separately used for training in some existing transductive generalized ZSL (T-GZSL) methods, which is not strict for a GZSL task. Hence, we suggest a more strict T-GZSL data setting and establish a competitive baseline on this setting by introducing the proposed STHS framework to T-GZSL. Bo Liu 0035, Qiulei Dong, Zhanyi Hu |
CVPR | 3 |
| 2021 | Rotation Transformation Network: Learning View-Invariant Point Cloud For Classification And SegmentationabstractMany recent works show that a spatial manipulation module could boost the performances of deep neural networks (DNNs) for 3D point cloud analysis. In this paper, we aim to provide an insight into spatial manipulation modules. Firstly, we find that the smaller the rotational degree of freedom (RDF) of objects is, the more easily these objects are handled by these DNNs. Then, we investigate the effect of the popular T-Net module and find that it could not reduce the RDF of objects. Motivated by the above two issues, we propose a rotation transformation network for point cloud analysis, called RTN, which could reduce the RDF of input 3D objects to 0. The RTN could be seamlessly inserted into many existing DNNs for point cloud analysis. Extensive experimental results on 3D point cloud classification and segmentation tasks demonstrate that the proposed RTN could improve the performances of several state-of-the-art methods significantly. Shuang Deng, Bo Liu 0035, Qiulei Dong, Zhanyi Hu |
ICME | 4 |
| 2021 | Optimization-Based Visual-Inertial SLAM Tightly Coupled with Raw GNSS MeasurementsabstractUnlike loose coupling approaches and the EKF-based approaches in the literature, we propose an optimization-based visual-inertial SLAM tightly coupled with raw Global Navigation Satellite System (GNSS) measurements, a first attempt of this kind in the literature to our knowledge. More specifically, reprojection error, IMU pre-integration error and raw GNSS measurement error are jointly minimized within a sliding window, in which the asynchronism between images and raw GNSS measurements is accounted for. In addition, issues such as marginalization, noisy measurements removal, as well as tackling vulnerable situations are also addressed. Experimental results on public dataset in complex urban scenes show that our proposed approach outperforms state-of-the-art visual-inertial SLAM, GNSS single point positioning, as well as a loose coupling approach, including scenes mainly containing low-rise buildings and those containing urban canyons. Jinxu Liu, Wei Gao 0014, Zhanyi Hu |
ICRA | 3 |
| 2021 | Semantic-diversity transfer network for generalized zero-shot learning via inner disagreement based OOD detector
Bo Liu 0035, Qiulei Dong, Zhanyi Hu |
Knowl. Based Syst. | 3 |
| 2021 | Urban Scene LOD Vectorized Modeling From Photogrammetry MeshesabstractUrban scene modeling is a challenging task for the photogrammetry and computer vision community due to its large scale, structural complexity, and topological delicacy. This paper presents an efficient multistep modeling framework for large-scale urban scenes from aerial images. It takes aerial images and a textured 3D mesh model generated by an image-based modeling system as the input and outputs compact polygon models with semantics at different levels of detail (LODs). Based on the key observation that urban buildings usually have piecewise planar rooftops and vertical walls, we propose a segment-based modeling method, which consists of three major stages: scene segmentation, roof contour extraction, and building modeling. By combining the deep neural network predictions with geometric constraints of the 3D mesh, the scene is first segmented into three classes. Then, for each building mesh, the 2D line segments are detected and used to slice the ground into polygon cells, followed by assigning each cell a roof plane via a MRF optimization. Finally, the LOD model is obtained by extruding cells to their corresponding planes. Compared with direct modeling in 3D space, we transform the mesh into a uniform 2D image grid representation and most of the modeling work is performed in 2D space, which has the advantages of low computational complexity and high robustness. In addition, our method doesn't require any global prior, such as the Manhattan or Atlanta world assumption, making it flexible to model scenes with different characteristics and complexity. Experiments on both single buildings and large-scale urban scenes demonstrate that by combining 2D photometric with 3D geometric information, the proposed algorithm is robust and efficient in urban scene LOD vectorized modeling compared with the state-of-the-art approaches. Jiali Han, Lingjie Zhu, Xiang Gao 0009, Zhanyi Hu, Liyang Zhou, Hongmin Liu 0001, Shuhan Shen |
IEEE Trans. Image Process. | 4 |
| 2021 | An Iterative Co-Training Transductive Framework for Zero Shot LearningabstractIn zero-shot learning (ZSL) community, it is generally recognized that transductive learning performs better than inductive one as the unseen-class samples are also used in its training stage. How to generate pseudo labels for unseen-class samples and how to use such usually noisy pseudo labels are two critical issues in transductive learning. In this work, we introduce an iterative co-training framework which contains two different base ZSL models and an exchanging module. At each iteration, the two different ZSL models are co-trained to separately predict pseudo labels for the unseen-class samples, and the exchanging module exchanges the predicted pseudo labels, then the exchanged pseudo-labeled samples are added into the training sets for the next iteration. By such, our framework can gradually boost the ZSL performance by fully exploiting the potential complementarity of the two models' classification capabilities. In addition, our co-training framework is also applied to the generalized ZSL (GZSL), in which a semantic-guided OOD detector is proposed to pick out the most likely unseen-class samples before class-level classification to alleviate the bias problem in GZSL. Extensive experiments on three benchmarks show that our proposed methods could significantly outperform about 31 state-of-the-art ones. Bo Liu 0035, Lihua Hu, Qiulei Dong, Zhanyi Hu |
IEEE Trans. Image Process. | 4 |
| 2020 | Zero-Shot Learning from Adversarial Feature Residual to Compact Visual FeatureabstractRecently, many zero-shot learning (ZSL) methods focused on learning discriminative object features in an embedding feature space, however, the distributions of the unseen-class features learned by these methods are prone to be partly overlapped, resulting in inaccurate object recognition. Addressing this problem, we propose a novel adversarial network to synthesize compact semantic visual features for ZSL, consisting of a residual generator, a prototype predictor, and a discriminator. The residual generator is to generate the visual feature residual, which is integrated with a visual prototype predicted via the prototype predictor for synthesizing the visual feature. The discriminator is to distinguish the synthetic visual features from the real ones extracted from an existing categorization CNN. Since the generated residuals are generally numerically much smaller than the distances among all the prototypes, the distributions of the unseen-class features synthesized by the proposed network are less overlapped. In addition, considering that the visual features from categorization CNNs are generally inconsistent with their semantic features, a simple feature selection strategy is introduced for extracting more compact semantic visual features. Extensive experimental results on six benchmark datasets demonstrate that our method could achieve a significantly better performance than existing state-of-the-art methods by ∼1.2-13.2% in most cases. Bo Liu 0035, Qiulei Dong, Zhanyi Hu |
AAAI | 3 |
| 2020 | 3D Semantic Labeling of Photogrammetry Meshes Based on Active LearningabstractAs different urban scenes are similar but still not completely consistent, coupled with the complexity of labeling directly in 3D, high-level understanding of 3D scenes has always been a tricky problem. In this paper, we propose a procedural approach for 3D semantic expression of urban scenes based on active learning. We first start with a small labeled image set to fine-tune a semantic segmentation network and then project its probability map onto a 3D mesh model for fusion, finally outputs a 3D semantic mesh model in which each facet has a semantic label and a heat model showing each facet's confidence. Our key observation is that our algorithm is iterative, in each iteration, we use the output semantic model as a supervision to select several valuable images for annotation to co-participate in the fine-tuning for overall improvement. In this way, we reduce the workload of labeling but not the quality of 3D semantic model. Using urban areas from two different cities, we show the potential of our method and demonstrate its effectiveness. Mengqi Rong, Shuhan Shen, Zhanyi Hu |
ICPR | 3 |
| 2020 | Effective two-view line segment reconstruction based on structure priors
Wei Wang 0347, Hainan Cui, Wei Gao 0014, Zhanyi Hu |
Sci. China Inf. Sci. | 4 |
| 2020 | Effective piecewise planar modeling based on sparse 3D points and convolutional neural network
Wei Wang 0347, Wei Gao 0014, Zhanyi Hu |
Neurocomputing | 3 |
| 2020 | Complete Scene Reconstruction by Merging Images and Laser ScansabstractImage based modeling and laser scanning are two commonly used approaches in large-scale architectural scene reconstruction nowadays. In order to generate a complete scene reconstruction, an effective way is to completely cover the scene using ground and aerial images, supplemented by laser scanning on certain regions with low textures and complicated structures. Thus, the key issue is to accurately calibrate cameras and register laser scans in a unified framework. To this end, we proposed a three-step pipeline for complete scene reconstruction by merging images and laser scans. First, images are captured around the architecture in a multiview and multiscale way and are feed into a structure-from-motion (SfM) pipeline to generate SfM points. Then, based on the SfM result, the laser scanning locations are automatically planned by considering textural richness, structural complexity of the scene and spatial layout of the laser scans. Finally, the images and laser scans are accurately merged in a coarse-to-fine manner. Experimental evaluations on two ancient Chinese architecture datasets demonstrate the effectiveness of our proposed complete scene reconstruction pipeline. Xiang Gao 0009, Shuhan Shen, Lingjie Zhu, Tianxin Shi, Zhiheng Wang 0001, Zhanyi Hu |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2019 | Active Semantic Labeling of Street View Point CloudsabstractSemantic 3D models have shown their importance in many fields such as autonomous driving. However, it remains a tough task to assign semantic labels to various scenes. In this paper, we propose an Active Learning based method for semantic labeling of street view point clouds with a small amount of annotated data samples. The proposed method takes a point cloud and registrated images as the input, and yields a point cloud with semantic labels. We iteratively fine-tunes a network with the ever-enlarging training set to exploit the semantic information of the scene, and fuse the semantic labels in 3D space. To deal with the imbalanced data in street view scenes, a label biased criterion for query selection is proposed to help select images to efficiently improve the performance of the network and the quality of the semantic model. Experimental result shows that the proposed method demands limited human labor and works well in assigning semantic labels to the imbalanced scenes like street view scenes. Shuhan Shen, Zhanyi Hu |
ICME | 3 |
| 2019 | Visual-Inertial Odometry Tightly Coupled with Wheel Encoder Adopting Robust Initialization and Online Extrinsic CalibrationabstractCombining camera, IMU and wheel encoder is a wise choice for car positioning because of the low cost and complementarity of the sensors. We propose a novel extended visual-inertial odometry algorithm tightly fusing data from the above three sensors. Firstly we propose an IMU-odometer pre-integration approach utilizing complete IMU measurements and wheel encoder readings, to make scale estimation more accurate in subsequent 4-degrees of freedom (DoF) optimization. Secondly we develop an original initialization module where encoder readings are fully utilized to refine gravity direction and provide an initial value for camera pose in real scale. Thirdly, we design a computationally efficient online extrinsic calibration method by fixing the linearization point for the rotational component of IMU-odometer extrinsic parameters, which is deployed depending on the convergence of accelerometer bias. Experimental results prove the robustness of our initialization module and the accuracy of the whole trajectory, as well as the improvement brought about by online extrinsic calibration. Our program can also run on an Nvidia Jetson TX2 module in real time. Jinxu Liu, Wei Gao 0014, Zhanyi Hu |
IROS | 3 |
| 2019 | Effectively modeling piecewise planar urban scenes based on structure priors and CNN
Wei Wang 0347, Wei Gao 0014, Zhanyi Hu |
Sci. China Inf. Sci. | 3 |
| 2019 | Ground and aerial meta-data integration for localization and reconstruction: A review
Xiang Gao 0009, Shuhan Shen, Zhanyi Hu, Zhiheng Wang 0001 |
Pattern Recognit. Lett. | 3 |
| 2018 | Fine-Level Semantic Labeling of Large-Scale 3D Model by Active LearningabstractSemantic labeling of 3D models has been a challenging task in recent years. Due to the various categories and shapes of 3D objects in different scenes, it is hard to develop a versatile method suitable for most scenes. In this paper, we propose an Active Learning based method to tackle the problem. The proposed method takes a 3D mesh model generated from images using SfM and MVS, as well as the calibrated images, as the input, and outputs a semantic mesh model in which each facet takes a fine-level semantic label. Starting with a small annotated image set, we progressively fine-tune a Convolutional Neural Network (CNN) with the ever-enlarging annotated image set for image semantic segmentation. In each iteration, by back-projecting the pixel labels to the 3D model and fusing them in 3D space, a semantic 3D model is generated. The semantic 3D model functions as a supervisor to select a batch of worthy images for annotation to boost the performance of the CNN in next iteration. This process iterates until the label assignment of the 3D model becomes steady. By making full use of the 3D geometric information, the proposed method could significantly reduce the annotation cost without losing the labeling quality of 3D models. Experimental results of fine-level labeling on two large-scale ancient Chinese architectures demonstrate the effectiveness of the proposed method. Shuhan Shen, Zhanyi Hu |
3DV | 3 |
| 2018 | Large Scale Urban Scene Modeling from MVS Meshes
Lingjie Zhu, Shuhan Shen, Xiang Gao 0009, Zhanyi Hu |
ECCV (11) | 4 |
| 2018 | High-Quality and Memory-Efficient Volumetric Integration of Depth Maps Using Plane PriorsabstractVolumetric integration method is widely used to fuse depth maps in dense 3D reconstruction systems. High memory footprint is one of its main disadvantages. We introduce a method to de-noise depth maps and save memory usage during volumetric integration of depth maps with the use of plane priors. We develop a new planar region detection method with the use of depth gradients and then de-noise the planar region of depth maps. During volumetric integration we allocate the voxels and integrate depth maps with the use of plane priors as well. Extensive experiments show that our method saves approximately 30% memory footprint and has higher reconstruction quality compared with some of the current state-of-the-art systems. These characteristics enable our method to be used for 3D scanning on mobile devices which have limited memory resources. Yangdong Liu, Wei Gao 0014, Zhanyi Hu |
ICPR | 3 |
| 2018 | Online Temporal Calibration of Camera and IMU using Nonlinear OptimizationabstractIn this paper, we aim to calibrate the time delay of timestamps of cameras and IMU measurements provided by Android smart phones or other low-cost devices whose camera and IMU are not temporally aligned. The time delay is estimated online in an iterative way through nonlinear optimization in sliding windows. We add new terms that are relative to time delay to the pre-integration results of IMU measurements instead of feature observations in order to improve the precision of temporal calibration. The experimental results indicate that our calibration result is closer to the real value compared with the state-of-the-art system and that our method appears to converge faster. By using our temporal calibration, the visual inertial odometry algorithm is less likely to suffer from fast turning or sudden stop. Jinxu Liu, Zhanyi Hu, Wei Gao 0014 |
ICPR | 2 |
| 2018 | Learning stratified 3D reconstruction
Qiulei Dong, Mao Shu, Hainan Cui, Huarong Xu, Zhanyi Hu |
Sci. China Inf. Sci. | 5 |
| 2018 | Modern physiognomy: an investigation on predicting personality traits and intelligence from the human face
Rizhen Qin, Wei Gao 0014, Huarong Xu, Zhanyi Hu |
Sci. China Inf. Sci. | 4 |
| 2018 | Statistics of Visual Responses to Image Object Stimuli from Primate AIT Neurons to DNN NeuronsabstractUnder the goal-driven paradigm, Yamins et al. ( 2014 ; Yamins & DiCarlo, 2016 ) have shown that by optimizing only the final eight-way categorization performance of a four-layer hierarchical network, not only can its top output layer quantitatively predict IT neuron responses but its penultimate layer can also automatically predict V4 neuron responses. Currently, deep neural networks (DNNs) in the field of computer vision have reached image object categorization performance comparable to that of human beings on ImageNet, a data set that contains 1.3 million training images of 1000 categories. We explore whether the DNN neurons (units in DNNs) possess image object representational statistics similar to monkey IT neurons, particularly when the network becomes deeper and the number of image categories becomes larger, using VGG19, a typical and widely used deep network of 19 layers in the computer vision field. Following Lehky, Kiani, Esteky, and Tanaka ( 2011 , 2014 ), where the response statistics of 674 IT neurons to 806 image stimuli are analyzed using three measures (kurtosis, Pareto tail index, and intrinsic dimensionality), we investigate the three issues in this letter using the same three measures: (1) the similarities and differences of the neural response statistics between VGG19 and primate IT cortex, (2) the variation trends of the response statistics of VGG19 neurons at different layers from low to high, and (3) the variation trends of the response statistics of VGG19 neurons when the numbers of stimuli and neurons increase. We find that the response statistics on both single-neuron selectivity and population sparseness of VGG19 neurons are fundamentally different from those of IT neurons in most cases; by increasing the number of neurons in different layers and the number of stimuli, the response statistics of neurons at different layers from low to high do not substantially change; and the estimated intrinsic dimensionality values at the low convolutional layers of VGG19 are considerably larger than the value of approximately 100 reported for IT neurons in Lehky et al. ( 2014 ), whereas those at the high fully connected layers are close to or lower than 100. To the best of our knowledge, this work is the first attempt to analyze the response statistics of DNN neurons with respect to primate IT neurons in image object representation. Qiulei Dong, Zhanyi Hu |
Neural Comput. | 3 |
| 2018 | Accurate and efficient ground-to-aerial model alignment
Xiang Gao 0009, Lihua Hu, Hainan Cui, Shuhan Shen, Zhanyi Hu |
Pattern Recognit. | 5 |
| 2018 | Learning Depth From Single Images With Deep Neural Network Embedding Focal LengthabstractLearning depth from a single image, as an important issue in scene understanding, has attracted a lot of attention in the past decade. The accuracy of the depth estimation has been improved from conditional Markov random fields, non-parametric methods, to deep convolutional neural networks most recently. However, there exist inherent ambiguities in recovering 3D from a single 2D image. In this paper, we first prove the ambiguity between the focal length and monocular depth learning, and verify the result using experiments, showing that the focal length has a great influence on accurate depth recovery. In order to learn monocular depth by embedding the focal length, we propose a method to generate synthetic varying-focal-length dataset from fixed-focal-length datasets, and a simple and effective method is implemented to fill the holes in the newly generated images. For the sake of accurate depth recovery, we propose a novel deep neural network to infer depth through effectively fusing the middle-level information on the fixed-focal-length dataset, which outperforms the state-of-the-art methods built on pretrained VGG. Furthermore, the newly generated varying-focallength dataset is taken as input to the proposed network in both learning and inference phases. Extensive experiments on the fixed- and varying-focal-length datasets demonstrate that the learned monocular depth with embedded focal length is significantly improved compared to that without embedding the focal length information. Lei He 0004, Guanghui Wang 0001, Zhanyi Hu |
IEEE Trans. Image Process. | 3 |
| 2017 | Batched Incremental Structure-from-MotionabstractThe incremental Structure-from-Motion (SfM) technique has advanced in both robustness and accuracy, but the efficiency and scalability remain its key challenges. In this paper, we propose a novel batched incremental SfM technique to tackle these problems in a unified framework, where two iteration loops are contained. The inner loop is a tracks triangulation loop, where a novel tracks selection method is proposed to find a compact subset of tracks for the bundle adjustment (BA). The outer loop is a camera registration loop, where a batch of cameras are simultaneously added to alleviate the drifting risk and reduce the running times of BA. By the tracks selection and batched camera registration, we find these two iteration loops converge fast. Extensive experiments demonstrate that our new SfM system performs similarly or better than many of the state-of-the-art SfM systems in terms of camera calibration accuracy, while is more efficient, robust and scalable for large-scale scene reconstruction. Hainan Cui, Shuhan Shen, Xiang Gao 0009, Zhanyi Hu |
3DV | 4 |
| 2017 | Variational Building Modeling from Urban MVS MeshesabstractIn this paper, we introduce a method for building LOD (levels of detail) modeling from urban multi-view stereo (MVS) meshes. Using city MVS meshes as input, our algorithm proceeds in three main steps: segmentation, contour extraction and modeling. With the prior knowledge and span constraint, we first segment the scene with an adapted variational measure to discover the underlying structures. The next contour extraction step projects the vertical structures onto the ground as line segments and extract the facade contours from them with a Markov random field. In the last modeling step, the contours are used to label the roof sections out and extruded to generate models of LODs with semantics. Experiments on complex and noisy urban meshes show that our approach could generate compact and accurate building models when compared with stateof- art methods. Lingjie Zhu, Shuhan Shen, Lihua Hu, Zhanyi Hu |
3DV | 4 |
| 2017 | HSfM: Hybrid Structure-from-MotionabstractStructure-from-Motion (SfM) methods can be broadly categorized as incremental or global according to their ways to estimate initial camera poses. While incremental system has advanced in robustness and accuracy, the efficiency remains its key challenge. To solve this problem, global reconstruction system simultaneously estimates all camera poses from the epipolar geometry graph, but it is usually sensitive to outliers. In this work, we propose a new hybrid SfM method to tackle the issues of efficiency, accuracy and robustness in a unified framework. More specifically, we propose an adaptive community-based rotation averaging method first to estimate camera rotations in a global manner. Then, based on these estimated camera rotations, camera centers are computed in an incremental way. Extensive experiments show that our hybrid method performs similarly or better than many of the state-of-the-art global SfM approaches, in terms of computational efficiency, while achieves similar reconstruction accuracy and robustness with two other state-of-the-art incremental SfM approaches. Hainan Cui, Xiang Gao 0009, Shuhan Shen, Zhanyi Hu |
CVPR | 4 |
| 2017 | Robust 3D Indoor Map Building via RGB-D SLAM with Adaptive IMU Fusion on Robot
Xinrui Meng, Wei Gao 0014, Zhanyi Hu |
ICIG (1) | 3 |
| 2017 | CSFM: Community-based structure from motionabstractStructure-from-Motion approaches could be broadly divided into two classes: incremental and global. While incremental manner is robust to outliers, it suffers from error accumulation and heavy computation load. The global manner has the advantage of simultaneously estimating all camera poses, but it is usually sensitive to epipolar geometry outliers. In this paper, we propose an adaptive community-based SfM (CSfM) method which takes both robustness and efficiency into consideration. First, the epipolar geometry graph is partitioned into separate communities. Then, the reconstruction problem is solved for each community in parallel. Finally, the reconstruction results are merged by a novel global similarity averaging method, which solves three convex L1 optimization problems. Experimental results show that our method performs better than many of the state-of-the-art global SfM approaches in terms of computational efficiency, while achieves similar or better reconstruction accuracy and robustness than many of the state-of-the-art incremental SfM approaches. Hainan Cui, Shuhan Shen, Xiang Gao 0009, Zhanyi Hu |
ICIP | 4 |
| 2017 | Accurate mesh-based alignment for ground and aerial multi-view stereo modelsabstractWe propose a method for accurate alignment of ground and aerial multi-view stereo (MVS) models. We achieve this goal by reconstructing the surface meshes from MVS point clouds generated by aerial and ground images respectively, and then iteratively removing the gap between them. The key issue is how to establish reliable correspondences between two meshes. To address this issue, we introduce a new set called the skeleton facet set (SFS) to represent the locally smooth part on the mesh, and then compute the transformation matrix by comparing the depths of the facets in SFS between aerial and ground models. Experimental results show that the proposed method is able to yield accurate alignment results and is robust to noise as well. Shuhan Shen, Xiang Gao 0009, Zhanyi Hu |
ICIP | 4 |
| 2017 | Global fusion of generalized camera model for efficient large-scale structure from motion
Hainan Cui, Shuhan Shen, Zhanyi Hu |
Sci. China Inf. Sci. | 3 |
| 2017 | Energy-based multi-view piecewise planar stereo
Wei Wang 0347, Lihua Hu, Zhanyi Hu |
Sci. China Inf. Sci. | 3 |
| 2017 | Tracks selection for robust, efficient and scalable large-scale structure from motion
Hainan Cui, Shuhan Shen, Zhanyi Hu |
Pattern Recognit. | 3 |
| 2017 | Two-Stream Deep Correlation Network for Frontal Face RecoveryabstractPose and textural variations are two dominant factors to affect the performance of face recognition. It is widely believed that generating the corresponding frontal face from a face image of an arbitrary pose is an effective step toward improving the recognition performance. In the literature, however, the frontal face is generally recovered by only exploring textural characteristic. In this letter, we propose a two-stream deep correlation network, which incorporates both geometric and textural features for frontal face recovery. Given a face image under an arbitrary pose as input, geometric and textural characteristics are first extracted from two separate streams. The extracted characteristics are then fused through the proposed multiplicative patch correlation layer. These two steps are integrated into one network for end-to-end training and prediction, which is demonstrated effective compared with state-of-the-art methods on the benchmark datasets. Ting Zhang 0006, Qiulei Dong, Ming Tang 0001, Zhanyi Hu |
IEEE Signal Process. Lett. | 4 |
| 2017 | Dynamic Graph Cuts in ParallelabstractThis paper aims at bridging the two important trends in efficient graph cuts in the literature, the one is to decompose a graph into several smaller subgraphs to take the advantage of parallel computation, the other is to reuse the solution of the max-flow problem on a residual graph to boost the efficiency on another similar graph. Our proposed parallel dynamic graph cuts algorithm takes the advantages of both, and is extremely efficient for certain dynamically changing MRF models in computer vision. The performance of our proposed algorithm is validated on two typical dynamic graph cuts problems: the foreground-background segmentation in video, where similar graph cuts problems need to be solved in sequential and GrabCut, where graph cuts are used iteratively. Miao Yu 0005, Shuhan Shen, Zhanyi Hu |
IEEE Trans. Image Process. | 3 |
| 2016 | Pursuing face identity from view-specific representation to view-invariant representationabstractHow to learn view-invariant facial representations is an important task for view-invariant face recognition. The recent work [1] discovered that the brain of the macaque monkey has a face-processing network, where some neurons are view-specific. Motivated by this discovery, this paper proposes a deep convolutional learning model for face recognition, which explicitly enforces this view-specific mechanism for learning view-invariant facial representations. The proposed model consists of two concatenated modules: the first one is a convolutional neural network (CNN) for learning the corresponding viewing pose to the input face image; the second one consists of multiple CNNs, each of which learns the corresponding frontal image of an image under a specific viewing pose. This method is of low computational cost, and it can be well trained with a relatively small number of samples. The experimental results on the MultiPIE dataset demonstrate the effectiveness of our proposed convolutional model in contrast to three state-of-the-art works. Ting Zhang 0006, Qiulei Dong, Zhanyi Hu |
ICIP | 3 |
| 2016 | Robust global translation averaging with feature tracksabstractHow to average translations is the single most difficult task in global structure-from-motion (SfM) to fully tap its potentials in terms of reconstruction efficiency and accuracy since usually only noisy translation directions can be factored out from essential matrices due to the inevitable matching outliers. To tackle this problem, this work proposes a two-step strategy. Firstly, a “2-point method” is introduced to refine the epipolar geometry by which a more accurate track set is generated. Then, translation lengths are computed by solving a convex L1 optimization according to the adjacent triangles induced by the selected tracks and translations. Extensive experiments show that our method performs similarly or better than the state-of-art SfM approaches in terms of the reconstruction accuracy, completeness and efficiency. Hainan Cui, Shuhan Shen, Zhanyi Hu |
ICPR | 3 |
| 2016 | Automatic building extraction from oblique aerial imagesabstractIn this paper we propose an automatic urban building extraction method for oblique aerial images. Five steps are included in this method: point cloud generation, grid partition, feature extraction, building detection and building reconstruction. Taking advantages of recent progress in large-scale Structure from Motion (SfM) and Multiple View Stereo (MVS), dense point cloud is generated first. Then, we project the point cloud into a regularly spaced grid in XY plan, and convert the building extraction problem into an image segmentation problem. By combining the strength of the geometric attribute and spectral attribute, three complementary features are extracted and a MRF based graph model along with an energy function is created. Points belonging to buildings are recognized by minimizing this function, and prismatic 3D building models are reconstructed accordingly. Shuhan Shen, Zhanyi Hu |
ICPR | 3 |
| 2016 | Triangulation and metric of lines based on geometric error
Fuchao Wu, Ming Zhang 0031, Guanghui Wang 0001, Zhanyi Hu |
Comput. Vis. Image Underst. | 4 |
| 2016 | Dynamic Parallel and Distributed Graph CutsabstractGraph cuts are widely used in computer vision. To speed up the optimization process and improve the scalability for large graphs, Strandmark and Kahl introduced a splitting method to split a graph into multiple subgraphs for parallel computation in both shared and distributed memory models. However, this parallel algorithm (the parallel BK-algorithm) does not have a polynomial bound on the number of iterations and is found to be non-convergent in some cases due to the possible multiple optimal solutions of its sub-problems. To remedy this non-convergence problem, in this paper, we first introduce a merging method capable of merging any number of those adjacent sub-graphs that can hardly reach agreement on their overlapping regions in the parallel BK-algorithm. Based on the pseudo-boolean representations of graph cuts, our merging method is shown to be effectively reused all the computed flows in these sub-graphs. Through both splitting and merging, we further propose a dynamic parallel and distributed graph cuts algorithm with guaranteed convergence to the globally optimal solutions within a predefined number of iterations. In essence, this paper provides a general framework to allow more sophisticated splitting and merging strategies to be employed to further boost performance. Our dynamic parallel algorithm is validated with extensive experimental results. Miao Yu 0005, Shuhan Shen, Zhanyi Hu |
IEEE Trans. Image Process. | 3 |
| 2015 | Efficient Large-Scale Structure From Motion by Fusing Auxiliary Imaging InformationabstractOne of the potentially effective means for large-scale 3D scene reconstruction is to reconstruct the scene in a global manner, rather than incrementally, by fully exploiting available auxiliary information on the imaging condition, such as camera location by Global Positioning System (GPS), orientation by inertial measurement unit (or compass), focal length from EXIF, and so on. However, such auxiliary information, though informative and valuable, is usually too noisy to be directly usable. In this paper, we present an approach by taking advantage of such noisy auxiliary information to improve structure from motion solving. More specifically, we introduce two effective iterative global optimization algorithms initiated with such noisy auxiliary information. One is a robust rotation averaging algorithm to deal with contaminated epipolar graph, the other is a robust scene reconstruction algorithm to deal with noisy GPS data for camera centers initialization. We found that by exclusively focusing on the estimated inliers at the current iteration, the optimization process initialized by such noisy auxiliary information could converge well and efficiently. Our proposed method is evaluated on real images captured by unmanned aerial vehicle, StreetView car, and conventional digital cameras. Extensive experimental results show that our method performs similarly or better than many of the state-of-art reconstruction approaches, in terms of reconstruction accuracy and completeness, but is more efficient and scalable for large-scale image data sets. Hainan Cui, Shuhan Shen, Wei Gao 0014, Zhanyi Hu |
IEEE Trans. Image Process. | 4 |
| 2015 | An efficient approach for 2D to 3D video conversion based on structure from motion
Wei Liu 0023, Yihong Wu 0002, Fusheng Guo, Zhanyi Hu |
Vis. Comput. | 4 |
| 2014 | Fusion of Auxiliary Imaging Information for Robust, Scalable and Fast 3D Reconstruction
Hainan Cui, Shuhan Shen, Wei Gao 0014, Zhanyi Hu |
ACCV (1) | 4 |
| 2014 | Radial distortion invariants and lens evaluation under a single-optical-axis omnidirectional camera
Yihong Wu 0002, Zhanyi Hu, Youfu Li 0001 |
Comput. Vis. Image Underst. | 2 |
| 2014 | How to Select Good Neighboring Images in Depth-Map Merging Based 3D ModelingabstractDepth-map merging based 3D modeling is an effective approach for reconstructing large-scale scenes from multiple images. In addition to generate high quality depth maps at each image, how to select suitable neighboring images for each image is also an important step in the reconstruction pipeline, unfortunately to which little attention has been paid in the literature until now. This paper is intended to tackle this issue for large scale scene reconstruction where many unordered images are captured and used with substantial varying scale and view-angle changes. We formulate the neighboring image selection as a combinatorial optimization problem and use the quantum-inspired evolutionary algorithm to seek its optimal solution. Experimental results on the ground truth data set show that our approach can significantly improve the quality of the depth-maps as well as final 3D reconstruction results with high computational efficiency. Shuhan Shen, Zhanyi Hu |
IEEE Trans. Image Process. | 2 |
| 2013 | Self-Calibration Under the Cayley Framework
Fuchao Wu, Zhanyi Hu |
Int. J. Comput. Vis. | 3 |
| 2012 | Automatic real-time SLAM relocalization based on a hierarchical bipartite graph model
Qiulei Dong, Zhaopeng Gu, Zhanyi Hu |
Sci. China Inf. Sci. | 3 |
| 2012 | Hybrid Parallel Bundle Adjustment for 3D Scene Reconstruction with Massive Points
Wei Gao 0014, Zhanyi Hu |
J. Comput. Sci. Technol. | 3 |
| 2012 | Rotationally Invariant Descriptors Using Intensity Order PoolingabstractThis paper proposes a novel method for interest region description which pools local features based on their intensity orders in multiple support regions. Pooling by intensity orders is not only invariant to rotation and monotonic intensity changes, but also encodes ordinal information into a descriptor. Two kinds of local features are used in this paper, one based on gradients and the other on intensities; hence, two descriptors are obtained: the Multisupport Region Order-Based Gradient Histogram (MROGH) and the Multisupport Region Rotation and Intensity Monotonic Invariant Descriptor (MRRID). Thanks to the intensity order pooling scheme, the two descriptors are rotation invariant without estimating a reference orientation, which appears to be a major error source for most of the existing methods, such as Scale Invariant Feature Transform (SIFT), SURF, and DAISY. Promising experimental results on image matching and object recognition demonstrate the effectiveness of the proposed descriptors compared to state-of-the-art descriptors. Bin Fan 0001, Fuchao Wu, Zhanyi Hu |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2012 | Robust line matching through line-point invariants
Bin Fan 0001, Fuchao Wu, Zhanyi Hu |
Pattern Recognit. | 3 |
| 2011 | Aggregating gradient distributions into intensity orders: A novel local image descriptorabstractA novel local image descriptor is proposed in this paper, which combines intensity orders and gradient distributions in multiple support regions. The novelty lies in three aspects: 1) The gradient is calculated in a rotation invariant way in a given support region; 2) The rotation invariant gradients are adaptively pooled spatially based on intensity orders in order to encode spatial information; 3) Multiple support regions are used for constructing descriptor which further improves its discriminative ability. Therefore, the proposed descriptor encodes not only gradient information but also information about relative relationship of intensities as well as spatial information. In addition, it is truly rotation invariant in theory without the need of computing a dominant orientation which is a major error source of most existing methods, such as SIFT. Results on the standard Oxford dataset and 3D objects have shown a significant improvement over the state-of-the-art methods under various image transformations. Bin Fan 0001, Fuchao Wu, Zhanyi Hu |
CVPR | 3 |
| 2011 | Efficient Suboptimal Solutions to the Optimal Triangulation
Fuchao Wu, Zhanyi Hu |
Int. J. Comput. Vis. | 3 |
| 2011 | Towards reliable matching of images containing repetitive patterns
Bin Fan 0001, Fuchao Wu, Zhanyi Hu |
Pattern Recognit. Lett. | 3 |
| 2010 | Line matching leveraged by point correspondencesabstractA novel method for line matching is proposed. The basic idea is to use tentative point correspondences, which can be easily obtained by keypoint matching methods, to significantly improve line matching performance, even when the point correspondences are severely contaminated by outliers. When matching a pair of image lines, a group of corresponding points that may be coplanar with these lines in 3D space is firstly obtained from all corresponding image points in the local neighborhoods of these lines. Then given such a group of corresponding points, the similarity between this pair of lines is calculated based on an affine invariant from one line and two points. The similarity is defined on the basis of median statistic in order to handle the problem of inevitable incorrect correspondences in the group of point correspondences. Furthermore, the relationship of rotation between the reference and query images is estimated from all corresponding points to filter out those pairs of lines which are obviously impossible to be matches, hence speeding up the matching process as well as further improving its robustness. Extensive experiments on real images demonstrate the good performance of the proposed method as well as its superiority to the state-of-the-art methods. Bin Fan 0001, Fuchao Wu, Zhanyi Hu |
CVPR | 3 |
| 2010 | Stereo matching with adaptive support-weight correlation and Graph CutsabstractConstructing a reliable data term and occlusion handling are two important issues for energy model based stereo method. In this paper, we at first use a 2-step adaptive support-weight correlation approach to get a reliable correlation volume. Then a pixel classification is proposed which classifies pixels into three classes: occluded, unstable and stable. For each pixel, according its class, a confidence weight is assigned. After that a new energy model is then constructed by integrating the correlation volume and the confidence weight. Finally, through minimizing this energy using Graph cuts, a better disparity map is obtained. Experimental results on the Middlebury data set show that our proposed method has the similar good performance with the top rank Graph Cuts based algorithms listed on the Middlebury homepage. Limin Shi, Fusheng Guo, Wei Gao 0014, Zhanyi Hu |
SMC | 4 |
| 2010 | Rejecting Mismatches by Correspondence Function
Xiangru Li 0001, Zhanyi Hu |
Int. J. Comput. Vis. | 2 |
| 2010 | Degeneracy from Twisted Cubic Under Two Views
Yihong Wu 0002, Zhanyi Hu |
J. Comput. Sci. Technol. | 3 |
| 2010 | Modeling Stereopsis via Markov Random FieldabstractMarkov random field (MRF) and belief propagation have given birth to stereo vision algorithms with top performance. This article explores their biological plausibility. First, an MRF model guided by physiological and psychophysical facts was designed. Typically an MRF-based stereo vision algorithm employs a likelihood function that reflects the local similarity of two regions and a potential function that models the continuity constraint. In our model, the likelihood function is constructed on the basis of the disparity energy model because complex cells are considered as front-end disparity encoders in the visual pathway. Our likelihood function is also relevant to several psychological findings. The potential function in our model is constrained by the psychological finding that the strength of the cooperative interaction minimizing relative disparity decreases as the separation between stimuli increases. Our model is tested on three kinds of stereo images. In simulations on images with repetitive patterns, we demonstrate that our model could account for the human depth percepts that were previously explained by the second-order mechanism. In simulations on random dot stereograms and natural scene images, we demonstrate that false matches introduced by the disparity energy model can be reliably removed using our model. A comparison with the coarse-to-fine model shows that our model is able to compute the absolute disparity of small objects with larger relative disparity. We also relate our model to several physiological findings. The hypothesized neurons of the model are selective for absolute disparity and have facilitative extra receptive field. There are plenty of such neurons in the visual cortex. In conclusion, we think that stereopsis can be implemented by neural networks resembling MRF. Yansheng Ming, Zhanyi Hu |
Neural Comput. | 2 |
| 2009 | Twisted Cubic: Degeneracy Degree and Relationship with General Degeneracy
Yihong Wu 0002, Zhanyi Hu |
ACCV (2) | 3 |
| 2009 | Cayley Transformation and Numerical Stability of Calibration Equation
Fuchao Wu, Zhanyi Hu |
Int. J. Comput. Vis. | 3 |
| 2009 | MSLD: A robust descriptor for line matching
Fuchao Wu, Zhanyi Hu |
Pattern Recognit. | 3 |
| 2009 | Modeling neuronal response to disparity gradient
Lianqing Yu, Zhanyi Hu |
Soft Comput. | 2 |
| 2009 | Pointwise Motion Image (PMI): A Novel Motion Representation and Its Applications to Abnormality Detection and Behavior RecognitionabstractIn this paper, we propose a novel motion representation and apply it to abnormality detection and behavior recognition. At first, pointwise correspondences for the foreground in two consecutive video frames are established by performing a salient-region-based pointwise matching algorithm. Then, based on the established pointwise correspondences, a pointwise motion image (PMI) for each frame is built up to represent the motion status of the foreground. The PMI is more suitable for video analysis as it encapsulates a variety of motion information such as pointwise motion speed, pointwise motion orientation, pointwise motion duration, as well as the global shape of the foreground. In addition, it represents all of these pieces of information by a color image in the HSV space, by which many popular techniques in the image processing field can be straightforwardly adopted. By combining the PMI and AdaBoost, a method for abnormality detection and behavior recognition is proposed. The proposed method is shown to possess a high discriminative ability and is capable of dealing with local motion, global motion, and similar motions with different speeds. Experiments including a comparison with two existing methods demonstrate the effectiveness of the proposed representation in abnormality detection and behavior recognition. Qiulei Dong, Yihong Wu 0002, Zhanyi Hu |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2008 | Detecting and Handling Unreliable Points for Camera Parameter Estimation
Yihong Wu 0002, Youfu Li 0001, Zhanyi Hu |
Int. J. Comput. Vis. | 3 |
| 2008 | A new linear algorithm for calibrating central catadioptric cameras
Fuchao Wu, Fuqing Duan, Zhanyi Hu, Yihong Wu 0002 |
Pattern Recognit. | 3 |
| 2008 | Pose determination and plane measurement using a trapezium
Fuqing Duan, Fuchao Wu, Zhanyi Hu |
Pattern Recognit. Lett. | 3 |
| 2008 | A new normalized method on line-based homography estimation
Xiaoming Deng 0001, Zhanyi Hu |
Pattern Recognit. Lett. | 3 |
| 2007 | MAPACo-Training: A Novel Online Learning Algorithm of Behavior Models
Heping Li, Zhanyi Hu, Yihong Wu 0002, Fuchao Wu |
ACCV (1) | 2 |
| 2007 | Multi-Camera Calibration with One-Dimensional Object under General MotionsabstractIt is well known that in order to calibrate a single camera with a one-dimensional (1D) calibration object, the object must undertake some constrained motions, in other words, it is impossible to calibrate a single camera if the object motion is of general one. For a multi-camera setup, i.e., when the number of camera is more than one, can the cameras be calibrated by a 1D object under general motions? In this work, we prove that all cameras can indeed be calibrated and a calibration algorithm is also proposed and experimentally tested. In contrast to other multi-camera calibration method, no one calibrated "base" camera is needed. In addition, we show that for such multi-camera cases, the minimum condition of calibration and critical motions are similar to those of calibrating a single camera with 1D calibration object. Liang Wang 0021, Fuchao Wu, Zhanyi Hu |
ICCV | 3 |
| 2007 | A note on the convergence of the mean shift
Xiangru Li 0001, Zhanyi Hu, Fuchao Wu |
Pattern Recognit. | 2 |
| 2007 | FOE estimation: Can image measurement errors be totally "corrected" by the geometric method?
Fuchao Wu, Liang Wang 0021, Zhanyi Hu |
Pattern Recognit. | 3 |
| 2007 | Structure and motion of nonrigid object under perspective projection
Guanghui Wang 0001, Hung-Tat Tsui, Zhanyi Hu |
Pattern Recognit. Lett. | 3 |
| 2006 | Gesture Recognition Using Quadratic Curves
Qiulei Dong, Yihong Wu 0002, Zhanyi Hu |
ACCV (1) | 3 |
| 2006 | Detecting Critical Configuration of Six Points
Yihong Wu 0002, Zhanyi Hu |
ACCV (2) | 2 |
| 2006 | Fisheye Lenses Calibration Using Straight-Line Spherical Perspective Projection Constraint
Xianghua Ying, Zhanyi Hu, Hongbin Zha |
ACCV (2) | 2 |
| 2006 | An Affine Invariant of Parallelograms and Its Application to Camera Calibration and 3D Reconstruction
Fuchao Wu, Fuqing Duan, Zhanyi Hu |
ECCV (2) | 3 |
| 2006 | Easy Calibration for Para-catadioptric-like CameraabstractFor omnidirectional cameras, most of the previous calibration methods from lines use conic fitting. This paper presents a calibration method for para-catadioptric-like cameras from lines without conic fitting under a single view. We establish equations on the five camera intrinsic parameters. These equations are linear for the focal lengths and skew factor once the principal point is known. The principal point can be approximated well by the center of the imaged mirror contour in practice or can be accurately estimated by quadric equations. After obtaining the principal point, we propose an algorithm to calibrate the focal lengths and skew factor. The algorithm needs neither prior structure knowledge nor conic fitting and is linear, which make it easy to implement. Other omnidirectional cameras can also use this presented work if high accuracy is not required. Experiments demonstrate the efficiency of the proposed algorithm. Yihong Wu 0002, Youfu Li 0001, Zhanyi Hu |
IROS | 3 |
| 2006 | A robust method to recognize critical configuration for camera calibration
Yihong Wu 0002, Zhanyi Hu |
Image Vis. Comput. | 2 |
| 2006 | Coplanar circles, quasi-affine invariance and calibration
Yihong Wu 0002, Xinju Li, Fuchao Wu, Zhanyi Hu |
Image Vis. Comput. | 4 |
| 2006 | Euclidean reconstruction of a circular truncated cone only from its uncalibrated contours
Yihong Wu 0002, Guanghui Wang 0001, Fuchao Wu, Zhanyi Hu |
Image Vis. Comput. | 4 |
| 2006 | The Number of Independent Kruppa Constraints from N Images
Zhanyi Hu, Yihong Wu 0002, Fuchao Wu, Songde Ma |
J. Comput. Sci. Technol. | 1 |
| 2006 | The LLE and a linear mapping
Fuchao Wu, Zhanyi Hu |
Pattern Recognit. | 2 |
| 2005 | Geometric Invariants and Applications under Catadioptric Camera ModelabstractThis paper presents geometric invariants of points and their applications under central catadioptric camera model. Although the image has severe distortions under the model, we establish some accurate projective geometric invariants of scene points and their image points. These invariants, being functions of principal point, are useful, from which a method for calibrating the camera principal point and a method for recovering planar scene structures are proposed. The main advantage of using these in variants for plane reconstruction is that neither camera motion nor the intrinsic parameters, except for the principal point, is needed. The theoretical correctness of the established invariants and robustness of the proposed methods are demonstrated by experiments. In addition, our results are found to be applicable to some more general camera models other than the catadioptric one Yihong Wu 0002, Zhanyi Hu |
ICCV | 2 |
| 2005 | 8-Point Algorithm Revisited: Factorized 8-Point AlgorithmabstractIn this paper, a novel algorithm for the fundamental matrix estimation, called factorized 8-point algorithm, is presented. The factorized 8-point algorithm is composed of three steps: (1) The measurement matrix in the traditional 8-point algorithm is decomposed into two factor matrices; (2) By introducing some auxiliary variables, a new linear minimization problem is formed, where every element of its associated measurement matrix is simply either a measurement datum or a constant; (3) The fundamental matrix is determined by solving this minimization problem by a least squares method. Like the traditional 8-point algorithm and Hartley's normalized 8-point algorithm, the factorized 8-point algorithm is also completely linear. But unlike the normalized 8-point algorithm, the factorized 8-point algorithm does not need any pre-normalization step. Since every element of the measurement matrix in the factorized 8-point algorithm is a measurement datum or a constant, no amplification of measurement error is involved; the factorized 8-point algorithm can boost effectively the robustness of the estimation. Large numbers of experiments show that the factorized 8-point algorithm consistently outperforms the traditional 8-point algorithm. In addition, although the factorized 8-point algorithm is specially designed for fundamental matrix estimation, its basic principle can be generalized to other estimation problems in computer vision, such as camera projection matrix estimation, homography estimation, focus of expansion estimation, and trifocal tensor estimation. Fuchao Wu, Zhanyi Hu, Fuqing Duan |
ICCV | 2 |
| 2005 | Single view metrology from scene constraints
Guanghui Wang 0001, Zhanyi Hu, Fuchao Wu, Hung-Tat Tsui |
Image Vis. Comput. | 2 |
| 2005 | Camera calibration and 3D reconstruction from a single view based on scene constraints
Guanghui Wang 0001, Hung-Tat Tsui, Zhanyi Hu, Fuchao Wu |
Image Vis. Comput. | 3 |
| 2005 | A General Sufficient Condition of Four Positive Solutions of the P3P Problem
Zhanyi Hu |
J. Comput. Sci. Technol. | 2 |
| 2005 | Camera calibration with moving one-dimensional objects
Fuchao Wu, Zhanyi Hu, Haijiang Zhu |
Pattern Recognit. | 2 |
| 2005 | Reconstruction of structured scenes from two uncalibrated images
Guanghui Wang 0001, Hung-Tat Tsui, Zhanyi Hu |
Pattern Recognit. Lett. | 3 |
| 2005 | A new constraint on the imaged absolute conic from aspect ratio and its application
Yihong Wu 0002, Zhanyi Hu |
Pattern Recognit. Lett. | 2 |
| 2004 | Camera Calibration from the Quasi-affine Invariance of Two Parallel Circles
Yihong Wu 0002, Haijiang Zhu, Zhanyi Hu, Fuchao Wu |
ECCV (1) | 3 |
| 2004 | Can We Consider Central Catadioptric Cameras and Fisheye Cameras within a Unified Imaging Model
Xianghua Ying, Zhanyi Hu |
ECCV (1) | 2 |
| 2004 | Single View Based Measurement on Space Planes
Guanghui Wang 0001, Zhanyi Hu, Fuchao Wu |
J. Comput. Sci. Technol. | 2 |
| 2004 | Catadioptric Camera Calibration Using Geometric InvariantsabstractCentral catadioptric cameras are imaging devices that use mirrors to enhance the field of view while preserving a single effective viewpoint. In this paper, we propose a novel method for the calibration of central catadioptric cameras using geometric invariants. Lines and spheres in space are all projected into conics in the catadioptric image plane. We prove that the projection of a line can provide three invariants whereas the projection of a sphere can only provide two. From these invariants, constraint equations for the intrinsic parameters of catadioptric camera are derived. Therefore, there are two kinds of variants of this novel method. The first one uses projections of lines and the second one uses projections of spheres. In general, the projections of two lines or three spheres are sufficient to achieve catadioptric camera calibration. One important conclusion in this paper is that the method based on projections of spheres is more robust and has higher accuracy than that based on projections of lines. The performances of our method are demonstrated by both the results of simulations and experiments with real images. Xianghua Ying, Zhanyi Hu |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2003 | A Linear Trinocular Rectification Method for Accurate Stereoscopic MatchingabstractIn this paper we propose and study a simple trinocular rectification method in which stratification to projective and affine components gives the rectifying homographies in a closed form. The class of trinocular rectifications which has 6 DOF is parametrized by an independent set of parameters with a geometric meaning. This offers the possibility to minimize rectification distortion in a natural way. It is shown experimentally on real data that our algorithm performs the rectification task correctly. As shown on groundtruth data using Confidently Stable Matching, trinocular matching significantly improves disparity map density and mismatch error, both depending on texture strength. Matching results on real complex scenes are reported. 1 Huaifeng Zhang, Jan Cech, Radim Sára, Fuchao Wu, Zhanyi Hu |
BMVC | 5 |
| 2003 | Catadioptric Camera Calibration Using Geometric InvariantsabstractCentral catadioptric cameras are imaging devices that use mirrors to enhance the field of view while preserving a single effective viewpoint. In this paper, we propose a novel method for the calibration of central catadioptric cameras using geometric invariants. Lines in space are projected into conics in the catadioptric image plane as well as spheres in space. We proved that the projection of a line can provide three invariants whereas the projection of a sphere can provide two. From these invariants, constraint equations for the intrinsic parameters of catadioptric camera are derived. Therefore, there are two variants of this novel method. The first one uses the projections of lines and the second one uses the projections of spheres. In general, the projections of two lines or three spheres are sufficient to achieve the catadioptric camera calibration. One important observation in this paper is that the method based on the projections of spheres is more robust and has higher accuracy than that using the projections of lines. The performances of our method are demonstrated by the results of simulations and experiments with real images. Xianghua Ying, Zhanyi Hu |
ICCV | 2 |
| 2003 | The Invariant Representations of a Quadric Cone and a Twisted CubicabstractUp to now, the shortest invariant representation of a quadric has 138 summands and there has been no invariant representation of a twisted cubic in 3D projective space, which limit to some extent the applications of invariants in 3D space. We give a very short invariant representation of a quadric cone, a special quadric, which has only two summands similar to the invariant representation of a planar conic, and give a short invariant representation of a twisted cubic. Then, a completely linear algorithm for generating the parametric equations of a twisted cubic is provided also. Finally, we exemplify some applications of our proposed invariant representations in the fields of computer vision and automated geometric theorem proving. Yihong Wu 0002, Zhanyi Hu |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2003 | A new easy camera calibration technique based on circular points
Xiaoqiao Meng, Zhanyi Hu |
Pattern Recognit. | 2 |
| 2003 | The impossibility of affine reconstruction from perspective image pairs obtained by a translating camera with varying parameters
Zhanyi Hu, Fuchao Wu, Guanghui Wang 0001 |
Pattern Recognit. Lett. | 1 |
| 2002 | A Note on the Number of Solutions of the Noncoplanar P4P ProblemabstractIn the literature, the PnP problem is indistinguishably defined as either to determine the distances of the control points from the camera's optical center or to determine the transformation matrices from the object-centered frame to the camera-centered frame. We show that these two definitions are generally not equivalent. In particular, we prove that, if the four control points are not coplanar, the upper bound of the P4P problem under the distance-based definition is 5 and also attainable, whereas the upper bound of the P4P problem under the transformation-based definition is only 4. Finally, we study the conditions under which at least two, three, four, and five different positive solutions exist in the distance based noncoplanar P4P problem. Zhanyi Hu, Fuchao Wu |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2001 | Robot Self-Location by Line Correspondences
Zhanyi Hu, Hung-Tat Tsui |
J. Comput. Sci. Technol. | 1 |
| 2000 | A New Easy Camera Calibration Technique Based on Circular PointsabstractInspired by Zhang's work on flexible calibration technique, a new easy technique for calibrating a camera based on circular points is proposed. The proposed technique only requires the camera to observe a newly designed planar calibration pattern (referred to as the model plane hereinafter) which includes a circle and a pencil of lines passing through the circle's center, at a few (at least three) different unknown orientations, then all the five intrinsic parameters can be determined linearly. The main advantage of our new technique is that it needs to know neither any metric measurement on the model plane, nor the correspondences between points on the model plane and image ones, hence the whole calibration process becomes extremely simple. The proposed technique is particularly useful for those people who are not familiar with computer vision. Experiments with simulated data as well as with real images show that our new technique is robust and accurate. (C) 2002 Pattern Recognition Society. Published by Elsevier Science Ltd. All rights reserved. Xiaoqiao Meng, Zhanyi Hu |
BMVC | 3 |
| 2000 | A New Automatic Quasar Recognition Technique Based on PCA and the Hough TransformabstractThe main purpose of quasar recognition is to determine the observed quasar spectrum's redshift value. In the past, the template of the quasar rest frame was mainly constructed based on astronomers' inference. Due to the inaccuracy of such a template, it is hard to determine the redshift value by matching the observed quasar spectrum with the template directly. This paper's main contributions are two-fold. Firstly, the template in our paper is constructed by the principal component analysis (PCA) method from some selected spectra with known readshift values, hence the resulting template is more realistic. Secondly, a 2D Hough transform, rather than a 1D Hough transforms is used. In our 2D Hough transform, in addition to the redshift parameter, a new parameter, named scale parameter, is also introduced to enhance the discriminability. The experiments show that our proposed technique is workable and the correct recognition rate can reach about as high as 90%. Ling-yun Huang, Zhanyi Hu, Fengmei Sun |
ICPR | 2 |
| 2000 | A Novel Method for Camera Planar Motion Detection and Robust Estimation of the 1D Trifocal TensorabstractA camera moving in a plane can often simplify a computer vision job. Camera self-calibration and robot self-location are good examples. We focus on the problem of camera planar motion and its application to the camera self-calibration method of Faugeras et al. (1998). We have made three new contributions to the camera planar motion detection. First, we prove that the trifocal lines in different views of the same planar motion must have the same line representation in the 2D retinal plane. This conclusion greatly simplifies the planar motion detection problem. Second, we distinguish the usage of three different cases of planar motion: ordinary planar motion, co-linear planar motion and rotation planar motion. Third, we propose the robust planar motion detection method and the method of estimation of trifocal lines in the uniform framework under the above three configurations. We have also purposed a method for eliminating the 2D image points whose 1D projection points are inaccurate and cause significant errors on the estimation of the 1D trifocal tensor. Experiments with our new techniques using simulated data and real images had obtained very good results, which are better than those reported in the above article. Le Lu 0001, Hung-Tat Tsui, Zhanyi Hu |
ICPR | 3 |
| 2000 | Planar Conic Based Camera CalibrationabstractInspired by the technique proposed by Zhang (1998), we proposed a camera calibration technique, which only requires observing three or more planar concentric conics at a few (at least two) different orientations. All computations involved are linear matrix manipulations. Compared with the classical techniques where an expensive calibration pattern is commonly used, our technique is easy to implement and more flexible. Using conics also simplifies the problem of correspondence. Both computer simulation and real data are used to test the proposed technique. Changjiang Yang, Fengmei Sun, Zhanyi Hu |
ICPR | 3 |
| 1999 | An inherent probabilistic aspect of the Hough transform
Zhanyi Hu, Changjiang Yang, Songde Ma |
J. Comput. Sci. Technol. | 1 |
| 1998 | Direct triangle extraction by a randomized Hough techniqueabstractThe macro features, such as triangle, quadrilateral, polygon, play a very important role in many computer vision applications such as matching, visual inspection, object tracking. However, effective ways to extract such macro features from images are still not available in the literature up to now. Worse, there are few reports on the matter. The paper proposes a randomized Hough technique to directly extract general triangles (i.e., with unknown size and orientation) from images. Extensive simulations as well as experiments with real images show that the results are satisfactory. Zhanyi Hu, Hung-Tat Tsui |
ICPR | 1 |
| 1998 | In defense of the Hough transformabstractThe Hough transform has been a widely used technique for geometric primitive extraction. However, recently, a new family of techniques based on optimization, such as the genetic algorithm, the tabu search algorithm, the algorithm based on random samples of minimum subset, claimed their superiority over the Hough transform. In this paper, based on a reasonable criterion, namely the expected number of random samples of minimum subset for a single successful primitive extraction, the performance of the two families of technique is compared. We show that the Hough transform generally outperforms optimization based techniques. In particular, based on a large number of simulations and experiments with real images, we show that with a comparable performance, the randomized Hough transform (RHT), a representative of Hough techniques, is about twice as fast as the random sample consensus (RANSAC), a representative of optimization based techniques, in both line extraction and circle extraction. Zhanyi Hu, Hung-Tat Tsui |
ICPR | 1 |
| 1998 | An intrinsic parameters self-calibration technique for active vision systemabstractThis paper presents a new camera intrinsic parameters self-calibration technique for ordinary active vision system. By controlling a pan-tilt-translation camera platform to do a sequence of specially designed motions (called a camera motion configuration here), we rigorously proved that the camera intrinsic parameters can be determined linearly under such two configurations: (1)regulating the camera’s orientation by 3 tilts, at each camera’s orientation, controlling the camera to translate twice along 2 orthogonal directions; (2)regulating the camera’s orientation by 1 pan and 2 tilts, at each camera’s orientation, controlling the camera to translate twice along 2 orthogonal directions. Furthermore, based on extensive simulations of stability analysis, it is shown that the configuration 2 is robust, whereas the configuration 1 is numerically unstable and sensitive to noise. Experiments with real data were carried out and the calibration results have been verified by a stereo vision experiment. A comparison with other camera calibration approaches is also reported here. 1. Changjiang Yang, Zhanyi Hu |
ICPR | 2 |
| 1998 | A new definition of the Hough transform
Zhanyi Hu, Songde Ma |
J. Comput. Sci. Technol. | 1 |
| 1997 | Performance prediction of the hough transform
Zhanyi Hu, Songde Ma |
J. Comput. Sci. Technol. | 1 |
| 1996 | Towards a new framework of the Hough transformabstractThis paper's main contributions are three-fold. Firstly, it is shown that the two existing template matching-like definitions of the Hough transform proposed by Princen, Illingworth and Kittler (1992) and by Bergen and Shvaytser (1991) are inadequate. The principal reason behind this is that the common implicit assumption of these two definitions, that every feature point within the template associated with a given accumulator cell E/sub 0/ in Hough space votes equally to E/sub 0/, is not reasonable. Secondly, an inherent probabilistic aspect of the Hough transform embedded in the transformation process from the image space to the parameter space is clarified. It is concluded that when the Hough transform is used to detect a pattern, an appropriate curve (surface, if the number of the parameters to be detected is more than 2) density function, which depends on the parameterization of the pattern, must be implicitly or explicitly provided to eliminate the uncertainties resulting from such a probabilistic aspect. Thirdly, a new framework of the Hough transform is proposed which mainly consists of two parts, namely parameterization and associated curve (surface) density function. Zhanyi Hu, Songde Ma |
ICIP (3) | 1 |
| 1996 | Uniform line parameterization
Zhanyi Hu, Songde Ma |
Pattern Recognit. Lett. | 1 |
| 1995 | Active Vision Based Stereo Vision
Zhanyi Hu, Songde Ma |
ACCV | 2 |
| 1995 | The three conditions of a good line parameterization
Zhanyi Hu, Songde Ma |
Pattern Recognit. Lett. | 1 |
| 1993 | Parameter probability density analysis for the Hough transform
Zhanyi Hu, Jacques Destiné |
Signal Process. | 1 |
| 1992 | Performance comparison of line parametrizationsabstractThe performance of the Hough transform depends on parametrization schemes as well as mapping procedures. Based on criteria such as normalized-variance and risk-ratio the authors find that normal parametrization is the best among the four frequently mentioned ones in literature, if uniform noise is present in the image space.> Zhanyi Hu, Jacques Destiné |
ICPR (3) | 1 |
| 1992 | Parameter density analysis in Hough spaceabstractPresents a general method to determine the parameter density distribution f/sup 1->M/( rho , theta ) (or the f/sup M->1/( rho , theta )) of a curve when the image points on the curve are transformed under the 1->M (the M->1) mapping. This method converts a relatively complicated problem of determining the density functions into an easier one of calculating the intersection points between the curve and the line rho =xcos theta +ysin theta and the derivatives of the curve at these points. Moreover, the authors have found the analytical result of the parameter point spread function of an infinitely long line, a problem that has been qualitatively discussed by Brown (1983).> Zhanyi Hu, Jacques Destiné |
ICPR (3) | 1 |