Xianghua Ying

dblp:00/131 · DBLP profile ↗
← Back
67ranked-venue papers
22as first author
29since 2021 · last 2026
0000-0002-9785-0727ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 51 · 18 first-author · 22 since 2021Graphics, computer vision, multimedia, augmented reality and games · 42 · 15 first-author · 14 since 2021Systems, architecture and hardware · 4 · 2 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 MVFormer: Multi-View Point Cloud Transformer for 3D Mechanical Component Recognition
Ruibin Wang, Xianghua Ying, Bowei Xing
Int. J. Comput. Vis.2
2025 Normal-NeRF: Ambiguity-Robust Normal Estimation for Highly Reflective Scenes
abstract
Neural Radiance Fields (NeRF) often struggle with reconstructing and rendering highly reflective scenes. Recent advancements have developed various reflection-aware appearance models to enhance NeRF's capability to render specular reflections. However, the robust reconstruction of highly reflective scenes is still hindered by the inherent shape ambiguity on specular surfaces. Existing methods typically rely on additional geometry priors to regularize the shape prediction, but this can lead to oversmoothed geometry in complex scenes. Observing the critical role of surface normals in parameterizing reflections, we introduce a transmittance-gradient-based normal estimation technique that remains robust even under ambiguous shape conditions. Furthermore, we propose a dual activated densities module that effectively bridges the gap between smooth surface normals and sharp object boundaries. Combined with a reflection-aware appearance model, our proposed method achieves robust reconstruction and high-fidelity rendering of scenes featuring both highly specular reflections and intricate geometric structures. Extensive experiments demonstrate that our method outperforms existing state-of-the-art methods on various datasets.
Ji Shi 0003, Xianghua Ying, Ruohao Guo, Bowei Xing, Wenzhen Yue
AAAI2
2025 Audio-Visual Instance Segmentation
abstract
In this paper, we propose a new multi-modal task, termed audio-visual instance segmentation (AVIS), which aims to simultaneously identify, segment and track individual sounding object instances in audible videos. To facilitate this research, we introduce a high-quality benchmark named AVISeg, containing over 90K instance masks from 26 semantic categories in 926 long videos. Additionally, we propose a strong baseline model for this task. Our model first localizes sound source within each frame, and condenses object-specific contexts into concise tokens. Then it builds long-range audio-visual dependencies between these tokens using window-based attention, and tracks sounding objects among the entire video sequences. Extensive experiments reveal that our method performs best on AVISeg, surpassing the existing methods from related tasks. We further conduct the evaluation on several multi-modal large models. Unfortunately, they exhibits subpar performance on instance-level sound source localization and temporal perception. We expect that AVIS will inspire the community towards a more comprehensive multi-modal understanding. Dataset and code is available at https://github.com/ruohaoguo/avis.
Ruohao Guo, Xianghua Ying, Yaru Chen 0003, Dantong Niu, Guangyao Li 0001, Liao Qu, Yanyu Qi, Jinxing Zhou, Bowei Xing, Wenzhen Yue, Ji Shi 0003, Qixun Wang 0002, Peiliang Zhang, Buwen Liang
CVPR2
2025 Can In-context Learning Really Generalize to Out-of-distribution Tasks?
abstract
In this work, we explore the mechanism of in-context learning (ICL) on out-of-distribution (OOD) tasks that were not encountered during training. To achieve this, we conduct synthetic experiments where the objective is to learn OOD mathematical functions through ICL using a GPT-2 model. We reveal that Transformers may struggle to learn OOD task functions through ICL. Specifically, ICL performance resembles implementing a function within the pretraining hypothesis space and optimizing it with gradient descent based on the in-context examples. Additionally, we investigate ICL's well-documented ability to learn unseen abstract labels in context. We demonstrate that such ability only manifests in the scenarios without distributional shifts and, therefore, may not serve as evidence of new-task-learning ability. Furthermore, we assess ICL's performance on OOD tasks when the model is pretrained on multiple tasks. Both empirical and theoretical analyses demonstrate the existence of the \textbf{low-test-error preference} of ICL, where it tends to implement the pretraining function that yields low test error in the testing context. We validate this through numerical experiments. This new theoretical result, combined with our empirical findings, elucidates the mechanism of ICL in addressing OOD tasks.
Qixun Wang 0002, Yifei Wang 0001, Xianghua Ying, Yisen Wang 0001
ICLR3
2025 FreEformer: Frequency Enhanced Transformer for Multivariate Time Series Forecasting
abstract
This paper presents FreEformer, a simple yet effective model that leverages a Frequency Enhanced Transformer for multivariate time series forecasting. Our work is based on the assumption that the frequency spectrum provides a global perspective on the composition of series across various frequencies and is highly suitable for robust representation learning. Specifically, we first convert time series into the complex frequency domain using the Discrete Fourier Transform (DFT). The Transformer architecture is then applied to the frequency spectra to capture cross-variate dependencies, with the real and imaginary parts processed independently. However, we observe that the vanilla attention matrix exhibits a low-rank characteristic, thus limiting representation diversity. To address this, we enhance the vanilla attention mechanism by introducing an additional learnable matrix to the original attention matrix, followed by row-wise L1 normalization. Theoretical analysis demonstrates that this enhanced attention mechanism improves both feature diversity and gradient flow. Extensive experiments demonstrate that FreEformer consistently outperforms state-of-the-art models on eighteen real-world benchmarks covering electricity, traffic, weather, healthcare and finance. Notably, the enhanced attention mechanism also consistently improves the performance of state-of-the-art Transformer-based forecasters. Code is available at https://anonymous.4open.science/r/FreEformer.
Wenzhen Yue, Xianghua Ying, Bowei Xing, Ruohao Guo, Ji Shi 0003
IJCAI3
2025 OLinear: A Linear Model for Time Series Forecasting in Orthogonally Transformed Domain
abstract
This paper presents $\mathbf{OLinear}$, a $\mathbf{linear}$-based multivariate time series forecasting model that operates in an $\mathbf{o}$rthogonally transformed domain. Recent forecasting models typically adopt the temporal forecast (TF) paradigm, which directly encode and decode time series in the time domain. However, the entangled step-wise dependencies in series data can hinder the performance of TF. To address this, some forecasters conduct encoding and decoding in the transformed domain using fixed, dataset-independent bases (e.g., sine and cosine signals in the Fourier transform). In contrast, we propose $\mathbf{OrthoTrans}$, a data-adaptive transformation based on an orthogonal matrix that diagonalizes the series' temporal Pearson correlation matrix. This approach enables more effective encoding and decoding in the decorrelated feature domain and can serve as a plug-in module to enhance existing forecasters. To enhance the representation learning for multivariate time series, we introduce a customized linear layer, $\mathbf{NormLin}$, which employs a normalized weight matrix to capture multivariate dependencies. Empirically, the NormLin module shows a surprising performance advantage over multi-head self-attention, while requiring nearly half the FLOPs. Extensive experiments on 24 benchmarks and 140 forecasting tasks demonstrate that OLinear consistently achieves state-of-the-art performance with high efficiency. Notably, as a plug-in replacement for self-attention, the NormLin module consistently enhances Transformer-based forecasters. The code and datasets are available at https://github.com/jackyue1994/OLinear.
Wenzhen Yue, Hao Wang 0179, Haoxuan Li 0001, Xianghua Ying, Ruohao Guo, Bowei Xing, Ji Shi 0003
NeurIPS5
2025 SalienTR: A closer look at multi-modal transformer for RGB-T salient object detection
Ruohao Guo, Wenzhen Yue, Liao Qu, Yanyu Qi, Dantong Niu, Xianghua Ying
Expert Syst. Appl.6
2025 Improving Bird's Eye View based 3D object detection via Learnable Perspective View
Ruibin Wang, Bowei Xing, Xianghua Ying
Expert Syst. Appl.3
2025 FPSMix: data augmentation strategy for point cloud classification
Taiyan Chen, Xianghua Ying
Frontiers Comput. Sci.2
2025 Multi-modal Prompt Alignment with Fine-grained LLM Knowledge for Unsupervised Domain Adaptation
Bowei Xing, Xianghua Ying, Ruibin Wang, Ruohao Guo
Int. J. Comput. Vis.2
2024 VPDETR: End-to-End Vanishing Point DEtection TRansformers
abstract
In the field of vanishing point detection, previous works commonly relied on extracting and clustering straight lines or classifying candidate points as vanishing points. This paper proposes a novel end-to-end framework, called VPDETR (Vanishing Point DEtection TRansformer), that views vanishing point detection as a set prediction problem, applicable to both Manhattan and non-Manhattan world datasets. By using the positional embedding of anchor points as queries in Transformer decoders and dynamically updating them layer by layer, our method is able to directly input images and output their vanishing points without the need for explicit straight line extraction and candidate points sampling. Additionally, we introduce an orthogonal loss and a cross-prediction loss to improve accuracy on the Manhattan world datasets. Experimental results demonstrate that VPDETR achieves competitive performance compared to state-of-the-art methods, without requiring post-processing.
Taiyan Chen, Xianghua Ying, Jinfa Yang, Ruibin Wang, Ruohao Guo, Bowei Xing, Ji Shi 0003
AAAI2
2024 Hierarchical Unsupervised Relation Distillation for Source Free Domain Adaptation
Bowei Xing, Xianghua Ying, Ruibin Wang, Ruohao Guo, Ji Shi 0003, Wenzhen Yue
ECCV (50)2
2024 Masked Local-Global Representation Learning for 3D Point Cloud Domain Adaptation
abstract
Point cloud is a popular and widely used geometric representation, which has attracted significant attention in 3D vision. However, the geometric variability of point cloud representations across different datasets can cause domain discrepancies, which hinder knowledge transfer and model generalization, resulting in degraded performance in target domain. In this paper, we present a novel approach to improve point cloud domain adaptation by employing masked representation learning in a self-supervised manner. Specifically, our method combines masked feature prediction and masked sample consistency to encode both local structure and global semantic information for learning invariant point cloud representation across domains. Moreover, to learn domain-specific representation and transfer knowledge from source to target, we propose prototype-calibrated self-training. By exploiting class-wise prototypes in the shared feature space, the soft pseudo labels can be adaptively denoised, which benefits the decision boundary learning in target domain. We conduct experiments on PointDA-10 and PointSegDA for 3D point cloud shape classification and semantic segmentation, respectively. The results demonstrate the effectiveness of our method and show that we can achieve the new state-of-the-art performance on point cloud domain adaptation.
Bowei Xing, Xianghua Ying, Ruibin Wang
ICRA2
2024 Sub-Adjacent Transformer: Improving Time Series Anomaly Detection with Reconstruction Error from Sub-Adjacent Neighborhoods
Wenzhen Yue, Xianghua Ying, Ruohao Guo, Ji Shi 0003, Bowei Xing, Yuqing Zhu 0001, Taiyan Chen
IJCAI2
2024 Instance-Level Panoramic Audio-Visual Saliency Detection and Ranking
abstract
Panoramic audio-visual saliency detection is to segment the most attention-attractive regions in 360° panoramic videos with sound. To meticulously delineate the detected salient regions and effectively model human attention shift, we extend this task to more fine-grained instance scenarios: identifying salient object instances and inferring their saliency ranks. In this paper, we propose the first instance-level framework that can simultaneously be applied to segmentation and ranking of multiple salient objects in panoramic videos. Specifically, it consists of a distortion-aware pixel decoder to overcome panoramic distortions, a sequential audio-visual fusion module to integrate audio-visual information, and a spatio-temporal object decoder to separate individual instances and predict their saliency scores. Moreover, owing to the absence of such annotations, we create the ground-truth saliency ranks for the PAVS10K benchmark. Extensive experiments demonstrate that our model is capable of achieving state-of-the-art performance on the PAVS10K for both saliency detection and ranking tasks. The code is available at https://github.com/ruohaoguo/pavsodr.
Ruohao Guo, Dantong Niu, Liao Qu, Yanyu Qi, Ji Shi 0003, Wenzhen Yue, Bowei Xing, Taiyan Chen, Xianghua Ying
ACM Multimedia9
2024 Open-Vocabulary Audio-Visual Semantic Segmentation
abstract
Audio-visual semantic segmentation (AVSS) aims to segment and classify sounding objects in videos with acoustic cues. However, most approaches operate on the close-set assumption and only identify pre-defined categories from training data, lacking the generalization ability to detect novel categories in practical applications. In this paper, we introduce a new task: open-vocabulary audio-visual semantic segmentation, extending AVSS task to open-world scenarios beyond the annotated label space. This is a more challenging task that requires recognizing all categories, even those that have never been seen nor heard during training. Moreover, we propose the first open-vocabulary AVSS framework, OV-AVSS, which mainly consists of two parts: 1) a universal sound source localization module to perform audio-visual fusion and locate all potential sounding objects and 2) an open-vocabulary classification module to predict categories with the help of the prior knowledge from large-scale pre-trained vision-language models. To properly evaluate the open-vocabulary AVSS, we split zero-shot training and testing subsets based on the AVSBench-semantic benchmark, namely AVSBench-OV. Extensive experiments demonstrate the strong segmentation and zero-shot generalization ability of our model on all categories. On the AVSBench-OV dataset, OV-AVSS achieves 55.43% mIoU on base categories and 29.14% mIoU on novel categories, exceeding the state-of-the-art zero-shot method by 41.88%/20.61% and open-vocabulary method by 10.2%/11.6%. The code is available at https://github.com/ruohaoguo/ovavss.
Ruohao Guo, Liao Qu, Dantong Niu, Yanyu Qi, Wenzhen Yue, Ji Shi 0003, Bowei Xing, Xianghua Ying
ACM Multimedia8
2024 Dissecting the Failure of Invariant Learning on Graphs
abstract
Enhancing node-level Out-Of-Distribution (OOD) generalization on graphs remains a crucial area. In this paper, we develop a Structural Causal Model (SCM) to theoretically dissect the performance of two prominent invariant learning methods--Invariant Risk Minimization (IRM) and Variance-Risk Extrapolation (VREx)--in node-level OOD settings. Our analysis reveals a critical limitation: these methods may struggle to identify invariant features due to the complexities introduced by the message-passing mechanism, which can obscure causal features within a range of neighboring samples. To address this, we propose Cross-environment Intra-class Alignment (CIA), which explicitly eliminates spurious features by aligning representations within the same class, bypassing the need for explicit knowledge of underlying causal patterns. To adapt CIA to node-level OOD scenarios where environment labels are hard to obtain, we further propose CIA-LRA (Localized Reweighting Alignment) that leverages the distribution of neighboring labels to selectively align node representations, effectively distinguishing and preserving invariant features while removing spurious ones, all without relying on environment labels. We theoretically prove CIA-LRA's effectiveness by deriving an OOD generalization error bound based on PAC-Bayesian analysis. Experiments on graph OOD benchmarks validate the superiority of CIA and CIA-LRA, marking a significant advancement in node-level OOD generalization.
Qixun Wang 0002, Yifei Wang 0001, Yisen Wang 0001, Xianghua Ying
NeurIPS4
2024 Tensor decompositions for temporal knowledge graph completion with time perspective
Jinfa Yang, Xianghua Ying, Yongjie Shi, Bowei Xing
Expert Syst. Appl.2
2024 Rectifying self-training with neighborhood consistency and proximity for source-free domain adaptation
Bowei Xing, Xianghua Ying, Ruibin Wang
Neurocomputing2
2024 UniTR: A Unified TRansformer-Based Framework for Co-Object and Multi-Modal Saliency Detection
abstract
Recent years have witnessed a growing interest in co-object segmentation and multi-modal salient object detection. Many efforts are devoted to segmenting co-existed objects among a group of images or detecting salient objects from different modalities. Albeit the appreciable performance achieved on respective benchmarks, each of these methods is limited to a specific task and cannot be generalized to other tasks. In this paper, we develop aUnifiedTRansformer-based framework, namelyUniTR, aiming at tackling the above tasks individually with a unified architecture. Specifically, a transformer module (CoFormer) is introduced to learn the consistency of relevant objects or complementarity from different modalities. To generate high-quality segmentation maps, we adopt a dual-stream decoding paradigm that allows the extracted consistent or complementary information to better guide mask prediction. Moreover, a feature fusion module (ZoomFormer) is designed to enhance backbone features and capture multi-granularity and multi-semantic information. Extensive experiments show that our UniTR performs well on17benchmarks, and surpasses existing state-of-the-art approaches.
Ruohao Guo, Xianghua Ying, Yanyu Qi, Liao Qu
IEEE Trans. Multim.2
2024 Exploiting Temporal Correlations for 3D Human Pose Estimation
abstract
Exploiting the rich temporal information in human pose sequences to facilitate 3D pose estimation has garnered particular attention. While various learning architectures have been designed for temporal exploiting, these architectures are usually trained via the 3D pose loss independently imposed on every single frame, without explicit temporal signals introduced for supervision. This inevitably increases the difficulty of temporal exploiting, since the network must reason about the meaningful temporal information based on the non-temporal single-frame supervision first. Only then, the network can utilize this information to guide sequence modeling. Recently, some work introduce temporal smoothness as an explicit supervision signal, which makes the network more straightforwardly reaches the temporal information from the supervision signal, thus improving the temporal exploiting. However, the temporal smoothness only roughly measures the short-term temporal properties between adjacent frame pairs. In this work, we propose to generalize the supervision of temporal smoothness to temporal correlations, letting the network precisely consider more comprehensive temporal properties in sequences. We contribute two novel correlation-based loss functions, which adopt different strategies to respectively regularize the encoder and decoder sides of the network for temporal exploiting. Besides, we design a pre-training scheme to ensure a general convergence of existing pose estimators under our correlation losses. Experiments on three benchmarks demonstrate that our method can be compatible with different networks, improving their temporal exploiting ability to output more accurate and robust pose estimations.
Ruibin Wang, Xianghua Ying, Bowei Xing
IEEE Trans. Multim.2
2024 Improving static and temporal knowledge graph embedding using affine transformations of entities
abstract
To find a suitable embedding for a knowledge graph (KG) remains a big challenge nowadays. By measuring the distance or plausibility of triples and quadruples in static and temporal knowledge graphs, many reliable knowledge graph embedding (KGE) models are proposed. However, these classical models may not be able to represent and infer various relation patterns well, such as TransE cannot represent symmetric relations, DistMult cannot represent inverse relations, RotatE cannot represent multiple relations, etc.. In this paper, we improve the ability of these models to represent various relation patterns by introducing the affine transformation framework. Specifically, we first utilize a set of affine transformations related to each relation or timestamp to operate on entity vectors, and then these transformed vectors can be applied not only to static KGE models, but also to temporal KGE models. The main advantage of using affine transformations is their good geometry properties with interpretability. Our experimental results demonstrate that the proposed intuitive design with affine transformations provides a statistically significant increase in performance with adding a few extra processing steps and keeping the same number of embedding parameters. Taking TransE as an example, we employ the scale transformation (the special case of an affine transformation). Surprisingly, it even outperforms RotatE to some extent on various datasets. We also introduce affine transformations into RotatE, Distmult, ComplEx, TTransE and TComplEx respectively, and experiments demonstrate that affine transformations consistently and significantly improve the performance of state-of-the-art KGE models on both static and temporal knowledge graph benchmarks.
Jinfa Yang, Xianghua Ying, Yongjie Shi, Ruibin Wang
J. Web Semant.2
2023 ECO-3D: Equivariant Contrastive Learning for Pre-training on Perturbed 3D Point Cloud
abstract
In this work, we investigate contrastive learning on perturbed point clouds and find that the contrasting process may widen the domain gap caused by random perturbations, making the pre-trained network fail to generalize on testing data. To this end, we propose the Equivariant COntrastive framework which closes the domain gap before contrasting, further introduces the equivariance property, and enables pre-training networks under more perturbation types to obtain meaningful features. Specifically, to close the domain gap, a pre-trained VAE is adopted to convert perturbed point clouds into less perturbed point embedding of similar domains and separated perturbation embedding. The contrastive pairs can then be generated by mixing the point embedding with different perturbation embedding. Moreover, to pursue the equivariance property, a Vector Quantizer is adopted during VAE training, discretizing the perturbation embedding into one-hot tokens which indicate the perturbation labels. By correctly predicting the perturbation labels from the perturbed point cloud, the property of equivariance can be encouraged in the learned features. Experiments on synthesized and real-world perturbed datasets show that ECO-3D outperforms most existing pre-training strategies under various downstream tasks, achieving SOTA performance for lots of perturbations.
Ruibin Wang, Xianghua Ying, Bowei Xing, Jinfa Yang
AAAI2
2023 Cross-Modal Contrastive Learning for Domain Adaptation in 3D Semantic Segmentation
abstract
Domain adaptation for 3D point cloud has attracted a lot of interest since it can avoid the time-consuming labeling process of 3D data to some extent. A recent work named xMUDA leveraged multi-modal data to domain adaptation task of 3D semantic segmentation by mimicking the predictions between 2D and 3D modalities, and outperformed the previous single modality methods only using point clouds. Based on it, in this paper, we propose a novel cross-modal contrastive learning scheme to further improve the adaptation effects. By employing constraints from the correspondences between 2D pixel features and 3D point features, our method not only facilitates interaction between the two different modalities, but also boosts feature representations in both labeled source domain and unlabeled target domain. Meanwhile, to sufficiently utilize 2D context information for domain adaptation through cross-modal learning, we introduce a neighborhood feature aggregation module to enhance pixel features. The module employs neighborhood attention to aggregate nearby pixels in the 2D image, which relieves the mismatching between the two different modalities, arising from projecting relative sparse point cloud to dense image pixels. We evaluate our method on three unsupervised domain adaptation scenarios, including country-to-country, day-to-night, and dataset-to-dataset. Experimental results show that our approach outperforms existing methods, which demonstrates the effectiveness of the proposed method.
Bowei Xing, Xianghua Ying, Ruibin Wang, Jinfa Yang, Taiyan Chen
AAAI2
2023 Improving point cloud classification and segmentation via parametric veronese mapping
abstract
Deep learning based 3D point cloud classification and segmentation has achieved remarkable success. Existing methods are usually implemented in the original space with 3D coordinates as inputs. However, we find that point networks taking only information of first-order coordinates hardly learn geometric features of higher order, such as point cloud normals or poses. In this study, we propose to map the input point clouds into a non-linear space to facilitate networks learning and leveraging high-order features. Firstly, we design the Parametric Veronese Mapping (PVM) function which automatically learns to map point clouds into a non-linear space. As a result, the mapped point clouds are enriched with high-order elements and maintain the basic point set properties as in the original 3D space. We can then exploit existing networks to learn high-order features from mapped point clouds. Secondly, we contribute a two-stage transformation learning module that modifies the previous one-stage module to better leverage high-order features for aligning point clouds in the projective space. Finally, an interaction module is designed to learn more discriminative features by aggregating information from both the original and projective space. Extensive experiments demonstrate that our method successfully improves the ability of most existing networks to learn high-order features and thus contributing to more accurate classification and segmentation. Moreover, the resulting models show stronger robustness to affine transformations and real-world perturbations.
Ruibin Wang, Xianghua Ying, Bowei Xing, Xin Tong 0007, Taiyan Chen, Jinfa Yang, Yongjie Shi
Pattern Recognit.2
2022 Learning Hierarchy-Aware Quaternion Knowledge Graph Embeddings with Representing Relations as 3D Rotations
abstract
Knowledge graph embedding aims to represent entities and relations as low-dimensional vectors, which is an effective way for predicting missing links. It is crucial for knowledge graph embedding models to model and infer various relation patterns, such as symmetry/antisymmetry. However, many existing approaches fail to model semantic hierarchies, which are common in the real world. We propose a new model called HRQE, which represents entities as pure quaternions. The relational embedding consists of two parts: (a) Using unit quaternions to represent the rotation part in 3D space, where the head entities are rotated by the corresponding relations through Hamilton product. (b) Using scale parameters to constrain the modulus of entities to make them have hierarchical distributions. To the best of our knowledge, HRQE is the first model that can encode symmetry/antisymmetry, inversion, composition, multiple relation patterns and learn semantic hierarchies simultaneously. Experimental results demonstrate the effectiveness of HRQE against some of the SOTA methods on four well-established knowledge graph completion benchmarks.
Jinfa Yang, Xianghua Ying, Yongjie Shi, Xin Tong 0007, Ruibin Wang, Taiyan Chen, Bowei Xing
COLING2
2022 Transformer Based Line Segment Classifier with Image Context for Real-Time Vanishing Point Detection in Manhattan World
abstract
Previous works on vanishing point detection usually use geometric prior for line segment clustering. We find that image context can also contribute to accurate line classification. Based on this observation, we propose to classify line segments into three groups according to three unknown-but-sought vanishing points with Manhattan world assumption, using both geometric information and image context in this work. To achieve this goal, we propose a novel Transformer based Line segment Classifier (TLC) that can group line segments in images and estimate the corresponding vanishing points. In TLC, we design a line segment descriptor to represent line segments using their positions, directions and local image contexts. Transformer based feature fusion module is used to capture global features from all line segments, which is proved to improve the classification performance significantly in our experiments. By using a network to score line segments for outlier rejection, vanishing points can be got by Singular Value Decomposition (SVD) from the classified lines. The proposed method runs at 25 fps on one NVIDIA 2080Ti card for vanishing point detection. Experimental results on synthetic and real-world datasets demonstrate that our method is superior to other state-of-the-art methods on the balance between accuracy and efficiency, while keeping stronger generalization capability when trained and evaluated on different datasets.
Xin Tong 0007, Xianghua Ying, Yongjie Shi, Ruibin Wang, Jinfa Yang
CVPR2
2022 Unsupervised Domain Adaptation for Semantic Segmentation of Urban Street Scenes Reflected by Convex Mirrors
abstract
Reflective convex mirrors are often used on street corners or as passenger-side mirrors on cars to obtain scene information by reflecting blind spots in the field of view, which can provide safety for pedestrians and drivers on roads, driveways, and alleys that lack of visibility. In recent years, deep learning based scene understanding methods (e.g., semantic segmentation) have been rapidly developed. However, due to gaps in the geometric domain, models trained on normal images are not directly applicable to scenes with convex mirror reflections. In this paper, we propose a novel framework to reduce the domain gap between normal images and convex mirror reflection images. In particular, we geometrically model convex mirrors to obtain a differentiable convex mirror simulation layer, CMSL. With the help of CMSL, we perform adversarial domain adaptation on edges in the input space and semantic boundaries in the output space to reduce the geometric appearance gap between the synthetic and real images. To verify the effectiveness of our algorithm, we construct the first convex mirror reflection scene dataset CMR1K, which contains 268 images with fine annotations. Extensive experimental results show that our algorithm can significantly outperform the baseline and previous methods. For example, our method surpasses the baseline and AdvEnt by 10% and 3% in mIoU, respectively.
Yongjie Shi, Xianghua Ying, Hongbin Zha
IEEE Trans. Intell. Transp. Syst.2
2021 Towards Cross-View Consistency in Semantic Segmentation While Varying View Direction
abstract
Several images are taken for the same scene with many view directions. Given a pixel in any one image of them, its correspondences may appear in the other images. However, by using existing semantic segmentation methods, we find that the pixel and its correspondences do not always have the same inferred label as expected. Fortunately, from the knowledge of multiple view geometry, if we keep the position of a camera unchanged, and only vary its orientation, there is a homography transformation to describe the relationship of corresponding pixels in such images. Based on this fact, we propose to generate images which are the same as real images of the scene taken in certain novel view directions for training and evaluation. We also introduce gradient guided deformable convolution to alleviate the inconsistency, by learning dynamic proper receptive field from feature gradients. Furthermore, a novel consistency loss is presented to enforce feature consistency. Compared with previous approaches, the proposed method gets significant improvement in both cross-view consistency and semantic segmentation performance on images with abundant view directions, while keeping comparable or better performance on the existing datasets.
Xin Tong 0007, Xianghua Ying, Yongjie Shi, He Zhao 0006, Ruibin Wang
IJCAI2
2020 RDCFace: Radial Distortion Correction for Face Recognition
abstract
The effects of radial lens distortion often appear in wide-angle cameras of surveillance and safeguard systems, which may severely degrade performances of previous face recognition algorithms. Traditional methods for radial lens distortion correction usually employ line features in scenarios that are not suitable for face images. In this paper, we propose a distortion-invariant face recognition system called RDCFace, which directly and only utilize the distorted images of faces, to alleviate the effects of radial lens distortion. RDCFace is an end-to-end trainable cascade network, which can learn rectification and alignment parameters to achieve a better face recognition performance without requiring supervision of facial landmarks and distortion parameters. We design sequential spatial transformer layers to optimize the correction, alignment, and recognition modules jointly. The feasibility of our method comes from implicitly using the statistics of the layout of face features learned from the large-scale face data. Extensive experiments indicate that our method is distortion robust and gains significant improvements on LFW, YTF, CFP, and RadialFace, a real distorted face benchmark compared with state-of-the-art methods.
He Zhao 0006, Xianghua Ying, Yongjie Shi, Xin Tong 0007, Jingsi Wen, Hongbin Zha
CVPR2
2020 A Simple Yet Effective Pipeline For Radial Distortion Correction
abstract
Eliminating the radial lens distortion of an image is a crucial preprocessing step for many computer vision applications. This paper explores a simple yet effective pipeline for radial distortion correction. Different from existing state-of-the-art methods that design complex network structure and concatenate multi-branch features. Our model uses a single network without any additional supervision. We design two differentiable layers to synthesize and rectify distorted images efficiently. Based on these layers, an online data synthesis strategy, a sampling grid loss, and an image reprojection loss are proposed to improve the distortion correction accuracy. Compared with the state-of-the-art methods, our model achieves the best rectification quality on both the synthetic and real distorted images with dozens of times faster inference speed. The training data and codes will be released.11https://github.com/MccreeZhao/RDCPipeline
He Zhao 0006, Yongjie Shi, Xin Tong 0007, Xianghua Ying, Hongbin Zha
ICIP4
2020 Qamface: Quadratic Additive Angular Margin Loss For Face Recognition
abstract
The angular-based softmax losses and their variants achieve great success in face recognition based on deep learning. ArcFace [1] which directly maximize decision boundary in angular space is one of the most popular and effective loss function. In this paper, we analyze the inherent limitations of ArcFace, including the non-monotonic logit and gradient curve, and inappropriate trend of loss value. To address these problems, we propose a novel loss function named the Quadratic Additive Angular Margin Loss (QAMFace). It takes the value of the angle through a quadratic function rather than cosine function as the target logit. Our QAMFace is easy to implement and only adds negligible computational overhead. Experiments on several relevant benchmarks show that QAMFace performs better in convergence on feature embedding, and consistently outperforms the state-of-the-art face recognition methods. Our codes will be released soon.1
He Zhao 0006, Yongjie Shi, Xin Tong 0007, Xianghua Ying, Hongbin Zha
ICIP4
2020 Position-aware and Symmetry Enhanced GAN for Radial Distortion Correction
abstract
This paper presents a novel method based on the generative adversarial network for radial distortion correction. Instead of generating a corrected image, our generator predicts a pixel flow map to measure the pixel offset between the distorted and corrected image. The quality of the generated pixel flow map and the warped image are judged by the discriminator. As texture far away from the image center has strong distortion, we develop an Adaptive Inverted Foveal layer which can transform the deformation to the intensity of the image to exploit this property. Rotation symmetry enhanced convolution kernels are applied to extract geometric features of different orientations explicitly. These learned features are recalibrated using the Squeeze-and-Excitation block to assign different weights for different directions. Moreover, we construct a first real-world radial distorted image dataset RD600 annotated with ground truth to evaluate our proposed method. We conduct extensive experiments to validate the effectiveness of each part of our framework. The further experiment shows our approach outperforms previous methods in both synthetic and real-world datasets quantitatively and qualitatively.
Yongjie Shi, Xin Tong 0007, Jingsi Wen, He Zhao 0006, Xianghua Ying, Hongbin Zha
ICPR5
2020 G-FAN: Graph-Based Feature Aggregation Network for Video Face Recognition
abstract
In this paper, we propose a graph-based feature aggregation network (G-FAN) for video face recognition. Compared with the still image, video face recognition exhibits great challenges due to huge intra-class variability and high interclass ambiguity. To address this problem, our G-FAN first uses a Convolutional Neural Network to extract deep features for every input face of a subject. Then, we build an affinity graph based on the relationship between facial features and apply Graph Convolutional Network to generate fine-grained quality vectors for each frame. Finally, the features among multiple frames are adaptively aggregated into a discriminative vector to represent a video face. Different from previous works that take a single image as input, our G-FAN could utilize the correlation information between image pairs and aggregate a template of face images simultaneously. The experiments on video face recognition benchmarks, including YTF, IJB-A, and IJB-C show that: (i) G-FAN automatically learns to advocate high-quality frames while repelling low-quality ones. (ii) G-FAN significantly boosts recognition accuracy and outperforms other state-of-the-art aggregation methods.
He Zhao 0006, Yongjie Shi, Xin Tong 0007, Jingsi Wen, Xianghua Ying, Hongbin Zha
ICPR5
2020 3D Orientation Estimation and Vanishing Point Extraction from Single Panoramas Using Convolutional Neural Network
abstract
3D orientation estimation is a key component of many important computer vision tasks such as autonomous navigation and 3D scene understanding. This paper presents a new CNN architecture to estimate the 3D orientation of an omnidirectional camera with respect to the world coordinate system from a single spherical panorama. To train the proposed architecture, we leverage a dataset of panoramas named VOP60K from Google Street View with labeled 3D orientation, including 50 thousand panoramas for training and 10 thousand panoramas for testing. Previous approaches usually estimate 3D orientation under pinhole cameras. However, for a panorama, due to its larger field of view, previous approaches cannot be suitable. In this paper, we propose an edge extractor layer to utilize the low-level and geometric information of panorama, an attention module to fuse different features generated by previous layers. A regression loss for two column vectors of the rotation matrix and classification loss for the position of vanishing points are added to optimize our network simultaneously. The proposed algorithm is validated on our benchmark, and experimental results clearly demonstrate that it outperforms previous methods.
Yongjie Shi, Xin Tong 0007, Jingsi Wen, He Zhao 0006, Xianghua Ying, Hongbin Zha
ICRA5
2019 Three Orthogonal Vanishing Points Estimation in Structured Scenes Using Convolutional Neural Networks
abstract
Inferring 3D geometric cues is a crucial step, whereas vanishing point plays a very important role in image understanding from a single image of structured scenes. In this paper, we construct a 330-thousand-item image database of structured scenes labeled by vanishing points, focal length and camera orientation. We grab over 300 thousand Google Street View images which cover the downtown and neighboring areas of New York, Los Angeles, Chicago and etc. The prediction error is characterized by a loss function by imposing a regularization item derived from the geometric constraint of orthogonal vanishing points and focal length. Moreover, we collect about 30 thousand indoor images using a full 360-degree panorama camera taken by ourselves in room, office, library and etc. We also using Convolutional Neural Networks to transfer learning from street view images to indoor images. Extensive experiments demonstrate that our algorithm outperforms state-of-the-art non-learned approaches.
Yongjie Shi, Danfeng Zhang, Jingsi Wen, Xin Tong 0007, He Zhao 0006, Xianghua Ying, Hongbin Zha
ICIP6
2018 Radial Lens Distortion Correction by Adding a Weight Layer with Inverted Foveal Models to Convolutional Neural Networks
abstract
Radial lens distortion often exists in images taken by commercial cameras, which does not satisfy the assumption of pinhole camera model. Eliminating the radial lens distortion of an image is necessary as a preprocessing step for many vision applications. Some paper has employed Convolutional Neural Networks (CNNs), to achieve radial distortion correction. They generated images with a large number of images of high variation of radial distortion, which can be well exploited by deep CNN with a high learning capacity, and reach the state-of-the-art results. In this paper, we claim that a weight layer with inverted foveal models can be added to these existing CNNs methods for radial distortion correction. In the widely used very deep Resnet-18 model, our method achieves about 20 percent decrease in the loss function with faster convergence compared to the previous methods.
Yongjie Shi, Danfeng Zhang, Jingsi Wen, Xin Tong 0007, Xianghua Ying, Hongbin Zha
ICPR5
2016 Radial Lens Distortion Correction Using Convolutional Neural Networks Trained with Synthesized Images
Jiangpeng Rong, Shiyao Huang, Zeyu Shang, Xianghua Ying
ACCV (3)4
2016 Camera Calibration from Periodic Motion of a Pedestrian
abstract
Camera calibration directly from image sequences of a pedestrian without using any calibration object is a really challenging task and should be well solved in computer vision, especially in visual surveillance. In this paper, we propose a novel camera calibration method based on recovering the three orthogonal vanishing points (TOVPs), just using an image sequence of a pedestrian walking in a straight line, without any assumption of scenes or motions, e.g., control points with known 3D coordinates, parallel or perpendicular lines, non-natural or pre-designed special human motions, as often necessary in previous methods. The traces of shoes of a pedestrian carry more rich and easily detectable metric information than all other body parts in the periodic motion of a pedestrian, but such information is usually overlooked by previous work. In this paper, we employ the images of the toes of the shoes on the ground plane to determine the vanishing point corresponding to the walking direction, and then utilize harmonic conjugate properties in projective geometry to recover the vanishing point corresponding to the perpendicular direction of the walking direction in the horizontal plane and the vanishing point corresponding to the vertical direction. After recovering all of the TOVPs, the intrinsic and extrinsic parameters of the camera can be determined. Experiments on various scenes and viewing angles prove the feasibility and accuracy of the proposed method.
Shiyao Huang, Xianghua Ying, Jiangpeng Rong, Zeyu Shang, Hongbin Zha
CVPR2
2015 Radial lens distortion correction using cascaded one-parameter division model
abstract
Radial lens distortion is the most significant lens distortion in current cameras, and many models are proposed to describe it. Fitzgibbon presented the prestigious division distortion model with just a single parameter. Based on the fact that a line in 3D space may be projected onto a curve in the image plane due to radial distortion, Alemán et al. utilized the Hough transform with three parameters to automatically correct radial lens distortion, namely, one parameter coming from the division model, and the other two arising from the corrected image line. However, in some cases, especially for wide angle lenses, the corrected results are not very satisfactory. Someone may suggest that we can use the so-called extended division model with more than one parameter. Unfortunately, the problem will become very hard to solve, since the dimensions of the Hough parameter space become higher than three. In this paper, we propose a cascaded one-parameter division model to deal with the problem. In each stage the Hough parameter space is always three-dimensional. Enormous experiments on real images illustrate the ability of our method.
Xiang Mei, Jiangpeng Rong, Xianghua Ying, Shiyao Huang, Hongbin Zha
ICIP4
2015 Ellipse-specific fitting by relaxing the 3L constraints with semidefinite programming
abstract
This paper presents a new efficient method to increase the accuracy and the robustness of ellipse fitting, by utilizing the 3L algorithm and semidefinite programming (SDP). The novelty lies on the combination of relaxed geometric distance constraints and semidefinite programming framework. Due to the relaxed 3L constraints, the proposed approach provides high robustness in the presence of noise. The accuracy of the final solution is prominently increased even if the data suffer from strong occlusions or noises. The proposed method represents significant advantages in both accuracy and robustness. Experimental results and comparisons with state-of-the-art fitting methods demonstrate the improvements in ellipse fitting.
Jiangpeng Rong, Xiang Mei, Xianghua Ying, Shiyao Huang, Hongbin Zha
ICIP4
2014 Imposing Differential Constraints on Radial Distortion Correction
Xianghua Ying, Xiang Mei, Ganwen Wang, Jiangpeng Rong, Hongbin Zha
ACCV (1)1
2014 Radial distortion correction from a single image of a planar calibration pattern using convex optimization
abstract
In Hartley-Kang's paper [7], they directly treated a planar calibration pattern as an image to construct an image pair together with a radial distorted image of the planar calibration pattern, and then proposed a very efficient method to determine the center of radial distortion by estimating the epipole in the radial distorted image. After determined the center of radial distortion, a least square method was utilized to recover the radial distortion function using the monotonicity constraints. In this paper, we present a convex optimization method to recover the radial distortion function using the same constraints as those required by Hartley-Kang's method, whereas our method can obtain better results of radial distortion correction. The experiments validate our approach.
Xianghua Ying, Xiang Mei, Ganwen Wang, Hongbin Zha
ICIP1
2014 The Perspective-3-Point Problem When Using a Planar Mirror
abstract
The Perspective-3-Point problem (P3P) is a classical and fundamental problem in computer vision. All possible solution sets for the P3P problem are from 1 to 4 solutions. In this paper, we propose a very simple way to reduce the ambiguity of numbers of possible solutions in P3P using a planar mirror. For three reference points, if they and their reflections in a planar mirror are both observed, we may obtain two P3P problems: One is from the three original reference points, and the other is from their reflections. A trivial procedure may be suggested: Solve for each of the two P3P problems, and then find the intersections of the two solution sets. Different from the trivial case, we propose an efficient method which employs the ratio relations of the unknowns in the two P3P problems. The ratio relations are arise from mirror reflection, and can be easily determined before solving the two P3P problems. With the ratio relations, a system of 6 equations with 3 unknowns can be determined. To solve the over-constraint problem, we utilize an efficient algorithm by finding all local minima of least-squares residual. Experiments validate our approach.
Xianghua Ying, Ganwen Wang, Xiang Mei, Hongbin Zha
ICPR1
2013 Canonicalized central absolute moment for edge-based color constancy
abstract
In the recent paper by Weijer et al. [9], the authors proposed the Grey-Edge hypothesis, which assumes that the average edge difference in the scene is achromatic. The Minkowski norm of color derivatives is used to approximate the light source color. In this paper, we point out that the Minkowski norm of these derivatives can actually be interpreted by the raw moment of the distribution of these derivatives. Furthermore, we discovered that the central moment of the distribution can also be utilized to estimate the light source color, and comparable results may be obtained.
Xianghua Ying, Lulu Hou, Yongbo Hou, Hongbin Zha
ICIP1
2013 Self-Calibration of Catadioptric Camera with Two Planar Mirrors from Silhouettes
abstract
If an object is interreflected between two planar mirrors, we may take an image containing both the object and its multiple reflections, i.e., simultaneously imaging multiple views of an object by a single pinhole camera. This paper emphasizes the problem of recovering both the intrinsic and extrinsic parameters of the camera using multiple silhouettes from one single image. View pairs among views in a single image can be divided into two kinds by the relationship between the two views in the pair: reflected by some mirror (real or virtual) and in a circular motion. Epipoles in the first kind of pairs can be easily determined from intersections of common tangent lines of silhouettes. Based on the projective properties of these epipoles, efficient methods are proposed to recover both the imaged circular points and the included angle between two mirrors. Epipoles in the second kind of pairs can be recovered simultaneously with the projection of intersection line between two mirrors by solving a simple 1D optimization problem using the consistency constraint of epipolar tangent lines. Fundamental matrices among views in a single image are all recovered. Using the estimated intrinsic and extrinsic parameters of the camera, a euclidean reconstruction can be obtained. Experiments validate the proposed approach.
Xianghua Ying, Yongbo Hou, Sheng Guan, Hongbin Zha
IEEE Trans. Pattern Anal. Mach. Intell.1
2012 Direct least square fitting of ellipsoids
Xianghua Ying, Yongbo Hou, Sheng Guan, Hongbin Zha
ICPR1
2012 A Fast Algorithm for Multidimensional Ellipsoid-Specific Fitting by Minimizing a New Defined Vector Norm of Residuals Using Semidefinite Programming
abstract
A quadratic surface in n-dimensional space is defined as the locus of zeros of a quadratic polynomial. The quadratic polynomial may be compactly written in notation by an (n+1)-vector and a real symmetric matrix of order n+1, where the vector represents homogenous coordinates of an n-D point, and the symmetric matrix is constructed from the quadratic coefficients. If an n-D quadratic surface is an n-D ellipsoid, the leading n × n principal submatrix of the symmetric matrix would be positive or opposite definite. As we know, to impose a matrix being positive or opposite definite, perhaps the best choice may be to employ semidefinite programming (SDP). From such straightforward and intuitive knowledge, in the literature until 2002, Calafiore first proposed a feasible method for multidimensional ellipsoid-specific fitting using SDP, which minimizes the 2--norm of the algebraic residual vector. However, the runtime of the method is significantly long and memory is often out when the number of fitted points is greater than several thousand. In this paper, we propose a fast and easily implemented algorithm for multidimensional ellipsoid-specific fitting by minimizing a new defined vector norm of the algebraic residual vector using SDP, which drastically decreases the size of the SDP problem while preserving accuracy. The proposed fast method can handle several million fitted points without any difficulty.
Xianghua Ying, Hongbin Zha
IEEE Trans. Pattern Anal. Mach. Intell.1
2012 Vanishing point detection using cascaded 1D Hough Transform from single images
Bo Li 0018, Xianghua Ying, Hongbin Zha
Pattern Recognit. Lett.3
2011 Camera Resectioning from Image Edges with the L∞-Norm Using Linear Programming
Xianghua Ying, Lulu Hou, Hongbin Zha
BMVC2
2010 Geometric properties of multiple reflections in catadioptric camera with two planar mirrors
abstract
A catadioptric system consisting of a pinhole camera and two planar mirrors is deeply investigated in this paper. The two mirrors combine to form a corner and face-to-face with the pinhole. Their relative pose is unknown. An object will be reflected in the mirror corner one-time or multiple-times. Using the pinhole, we may take an image containing the object and its reflections, i.e., simultaneously imaging multiple views of an object by a single camera. We discovered that each 3D point and its reflections lie on a circle. We call the point set composed of a 3D point and its reflections, a Reflection Point Group (RPG), and the circle related to a RPG is called a RPG circle. All RPG circles are parallel to one another. Furthermore, each RPG can be partitioned into two separate subgroups. Shape formed by all points in a subgroup is invariant with respect to location of the 3D point. From these geometric properties, two calibration approaches can be utilized: One is based on parallel circles; the other is using 2D homographies among invariant shapes. Experiments validate our approaches.
Xianghua Ying, Ren Ren 0002, Hongbin Zha
CVPR1
2010 Single View Metrology Along Orthogonal Directions
abstract
In this paper, we describe how 3D metric measurements can be determined from a single uncalibrated image, when only minimal geometric information are available in the image. The minimal information just is orthogonal vanishing points. Given such limited information, we show that the length ratios on different orthogonal directions can be directly computed. The exciting discovery of the method seems to oppose common senses: Usually, in the calibration process, all edge-lengths of cuboid are known, in this paper, cuboid edge-lengths are unknown but its edge-lengths ratios can be recovered from image. 3D metric measurements can be directly computed from the image using our linear method.
Lulu Hou, Ren Ren 0002, Xianghua Ying, Hongbin Zha
ICPR4
2009 Combining Laser-Scanning Data and Images for Target Tracking and Scene Modeling
Hongbin Zha, Huijing Zhao, Jinshi Cui, Xuan Song 0001, Xianghua Ying
ISRR5
2008 Efficient detection of projected concentric circles using four intersection points on a secant line
abstract
Concentric circles are often used for calibration. Based on the geometric properties of concentric circles, we proved that only four intersection edge points on one secant line of the two images of the concentric circles (ICC) are sufficient to determine the whole parameters of the ICC, when the ratio of the radii of two concentric circles is given. Experimental results validate the proposed approach.
Xianghua Ying, Hongbin Zha
ICPR1
2008 Identical Projective Geometric Properties of Central Catadioptric Line Images and Sphere Images with Applications to Calibration
Xianghua Ying, Hongbin Zha
Int. J. Comput. Vis.1
2007 Camera Calibration Using Principal-Axes Aligned Conics
Xianghua Ying, Hongbin Zha
ACCV (1)1
2007 An Efficient Method for the Detection of Projected Concentric Circles
abstract
Concentric circles are often used as calibration features since they possess good geometric properties. This paper presents an efficient method for the detection of projected concentric circles in the image plane while considering their special geometric properties. The proposed method is capable of detecting partially visible concentric circles. Experimental results demonstrate the validity of the proposed approach.
Xianghua Ying, Hongbin Zha
ICIP (6)1
2007 Camera calibration from a circle and a coplanar point at infinity with applications to sports scenes analyses
abstract
Circles, such as concentric circles, arbitrary coplanar or parallel circles, are often employed for camera calibration. An interesting observation in football or basketball scenes is that a circle called the center circle existing in the midfield. By taking into account the midfield line which passes through the center circle’ center, a novel camera calibration method using the images of the midfield is proposed in this paper. A point at infinity can be determined from the center circle and the midfield line using the pole-polar relationship. The image of a point at infinity is called a vanishing point. The image of a circle is called a circle image in this paper. From one circle image and one coplanar vanishing point, a cubic constraint on the IAC can be obtained. The degenerate cases are discussed, and some applications are given in this paper.
Xianghua Ying, Hongbin Zha
IROS1
2006 Fisheye Lenses Calibration Using Straight-Line Spherical Perspective Projection Constraint
Xianghua Ying, Zhanyi Hu, Hongbin Zha
ACCV (2)1
2006 Interpreting Sphere Images Using the Double-Contact Theorem
Xianghua Ying, Hongbin Zha
ACCV (1)1
2006 Geometric Interpretations of the Relation between the Image of the Absolute Conic and Sphere Images
abstract
A spherical object has been introduced into camera calibration for several years through utilizing the properties of an image conic, which is the projection of the occluding contour of a sphere in the perspective image. However, in literature, only an algebraic interpretation was presented for the relation between the image of the absolute conic and sphere images. In this paper, we propose two geometric interpretations of this relation and two novel camera calibration methods using sphere images are derived from these geometric interpretations.
Xianghua Ying, Hongbin Zha
IEEE Trans. Pattern Anal. Mach. Intell.1
2005 Linear Approaches to Camera Calibration from Sphere Images or Active Intrinsic Calibration Using Vanishing Points
abstract
Spherical objects and vanish points are often used for camera calibration. An occluding contour of a sphere is projected to a conic in the perspective image, and using a moving active camera, the trajectory of a vanishing point in the perspective images is also a conic when the camera is rotated about a fixed 3D axis whereas the translation of the camera is arbitrary. In fact, the problems of camera calibration using conics from spheres or vanishing points can be described by same mathematic representations. Two linear approaches to the problems are proposed in this paper: one based on the geometric interpretation of the relation between image conics and the image of the absolute conic, and the other using the special structure of the problems in algebra. Only three such conics are needed for the two linear approaches, and the minimum number for previous nonlinear optimization methods is also three. All five intrinsic parameters are recovered linearly without making assumptions, such as, zero-skew or unitary aspect ratio which are often used in previous methods. The two linear algorithms have been tested in extensive experiments with respect to noise sensitivity and also made comparisons with recent calibration techniques.
Xianghua Ying, Hongbin Zha
ICCV1
2005 Camera pose determination from a single view of parallel lines
abstract
In this paper, we present a method for finding the closed form solutions to the problem of determining the pose of a camera with respect to a given set of parallel lines in 3D space from a single view, while it can not be solved by previous methods for the perspective-n-line (PnL) problem. The main idea of our method is that, firstly the distances from the optical center of the camera to these parallel lines are determined, and then the pose parameters are recovered using the obtained distances. The problem of finding these optical-center-to-line distances in fact is the degenerated perspective-n-point (PnP) problem, and we proved that there are at most two solutions for the degenerated P3P problem. An application of our method is to distinguish crosswalks and staircases aiding for the partially sighted. The method also provides a different way to investigate the problem of shape from texture.
Xianghua Ying, Hongbin Zha
ICIP (3)1
2005 Simultaneously calibrating catadioptric camera and detecting line features using Hough transform
abstract
A line in space is projected to a conic in a central catadioptric image, and such a conic is called a line image. This paper proposes a novel approach to calibrating catadioptric camera and detecting line images simultaneously by using Hough transform. Previous approaches to catadioptric cameras calibration employ the traditional conic detecting or fitting methods for line images, and then use these recovered conies to estimate the intrinsic parameters based on some properties of line images. However, the type of a line image can be line, circle, ellipse, hyperbola or parabola, and in general only a small arc of the conic is visible in the image, which brings novel challenges for conic detection and fitting where traditional conic detecting and fitting methods may fail. As we know, the accuracy of the estimated intrinsic parameters highly depends on the accuracy of the extracted conies. The main contribution of this work is we show that all line images from catadioptric cameras with the same intrinsic parameters must belong to a family of conies with only two degree-of-freedom, and such a family is called a line image family. Therefore, we present a novel special Hough transform for line image detection which ensures that all detected conies must belong to a line image family related to certain intrinsic parameters. For all possible values of the unknown intrinsic parameters, the line image special Hough transform are performed. The one with the highest confidence is chosen as the estimated values for these unknown intrinsic parameters, and the corresponding results of line image detection are chosen as the estimated values for line images. In order to make the searching process more efficient, the hierarchical approaches are employed in this paper. The validity of our proposed approach is illustrated by experiments.
Xianghua Ying, Hongbin Zha
IROS1
2004 Can We Consider Central Catadioptric Cameras and Fisheye Cameras within a Unified Imaging Model
Xianghua Ying, Zhanyi Hu
ECCV (1)1
2004 Catadioptric Camera Calibration Using Geometric Invariants
abstract
Central catadioptric cameras are imaging devices that use mirrors to enhance the field of view while preserving a single effective viewpoint. In this paper, we propose a novel method for the calibration of central catadioptric cameras using geometric invariants. Lines and spheres in space are all projected into conics in the catadioptric image plane. We prove that the projection of a line can provide three invariants whereas the projection of a sphere can only provide two. From these invariants, constraint equations for the intrinsic parameters of catadioptric camera are derived. Therefore, there are two kinds of variants of this novel method. The first one uses projections of lines and the second one uses projections of spheres. In general, the projections of two lines or three spheres are sufficient to achieve catadioptric camera calibration. One important conclusion in this paper is that the method based on projections of spheres is more robust and has higher accuracy than that based on projections of lines. The performances of our method are demonstrated by both the results of simulations and experiments with real images.
Xianghua Ying, Zhanyi Hu
IEEE Trans. Pattern Anal. Mach. Intell.1
2003 Catadioptric Camera Calibration Using Geometric Invariants
abstract
Central catadioptric cameras are imaging devices that use mirrors to enhance the field of view while preserving a single effective viewpoint. In this paper, we propose a novel method for the calibration of central catadioptric cameras using geometric invariants. Lines in space are projected into conics in the catadioptric image plane as well as spheres in space. We proved that the projection of a line can provide three invariants whereas the projection of a sphere can provide two. From these invariants, constraint equations for the intrinsic parameters of catadioptric camera are derived. Therefore, there are two variants of this novel method. The first one uses the projections of lines and the second one uses the projections of spheres. In general, the projections of two lines or three spheres are sufficient to achieve the catadioptric camera calibration. One important observation in this paper is that the method based on the projections of spheres is more robust and has higher accuracy than that using the projections of lines. The performances of our method are demonstrated by the results of simulations and experiments with real images.
Xianghua Ying, Zhanyi Hu
ICCV1