Xinyu Huang 0001

dblp:91/2102-1 · DBLP profile ↗
← Back
30ranked-venue papers
5as first author
15since 2021 · last 2025
0000-0002-5786-3101ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 25 · 4 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 20 · 4 first-author · 10 since 2021Systems, architecture and hardware · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2Databases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2025 Depth Any Camera: Zero-Shot Metric Depth Estimation from Any Camera
abstract
While recent depth foundation models exhibit strong zero-shot generalization, achieving accurate metric depth across diverse camera types—particularly those with large fields of view (FoV) such as fisheye and 360-degree cameras—remains a significant challenge. This paper presents Depth Any Camera (DAC), a powerful zero-shot metric depth estimation framework that extends a perspective-trained model to effectively handle cameras with varying FoVs. The framework is designed to ensure that all existing 3D data can be leveraged, regardless of the specific camera types used in new applications. Remarkably, DAC is trained exclusively on perspective images but generalizes seamlessly to fisheye and 360-degree cameras without the need for specialized training data. DAC employs Equi-Rectangular Projection (ERP) as a unified image representation, enabling consistent processing of images with diverse FoVs. Its core components include pitch-aware Image-to-ERP conversion with efficient online augmentation to simulate distorted ERP patches from undistorted inputs, FoV alignment operations to enable effective training across a wide range of FoVs, and multi-resolution data augmentation to further address resolution disparities between training and testing. DAC achieves state-of-the-art zero-shot metric depth estimation, improving δ1accuracy by up to 50% on multiple fisheye and 360-degree datasets compared to prior metric depth foundation models, demonstrating robust generalization across camera types.
Yuliang Guo, Sparsh Garg, S. Mahdi H. Miangoleh, Xinyu Huang 0001, Liu Ren 0001
CVPR4
2025 Online Language Splatting
abstract
To enable AI agents to interact seamlessly with both humans and 3D environments, they must not only perceive the 3D world accurately but also align human language with 3D spatial representations. While prior work has made significant progress by integrating language features into geometrically detailed 3D scene representations using 3D Gaussian Splatting (GS), these approaches rely on computationally intensive offline preprocessing of language features for each input image, limiting adaptability to new environments. In this work, we introduce Online Language Splatting, the first framework to achieve online, near real-time, open-vocabulary language mapping within a 3DGS-SLAM system without requiring pre-generated language features. The key challenge lies in efficiently fusing high-dimensional language features into 3D representations while balancing the computation speed, memory usage, rendering quality and open-vocabulary capability. To this end, we innovatively design: (1) a high-resolution CLIP embedding module capable of generating detailed language feature maps in 18ms per frame, (2) a two-stage online auto-encoder that compresses 768-dimensional CLIP features to 15 dimensions while preserving open-vocabulary capabilities, and (3) a color-language disentangled optimization approach to improve rendering quality. Experimental results show that our online method not only surpasses the state-of-the-art offline methods in accuracy but also achieves more than 40x efficiency boost, demonstrating the potential for dynamic and interactive AI applications.
Saimouli Katragadda, Cho-Ying Wu, Yuliang Guo, Xinyu Huang 0001, Guoquan Huang 0001, Liu Ren 0001
ICCV4
2025 CHARM3R: Towards Unseen Camera Height Robust Monocular 3D Detector
Abhinav Kumar 0004, Yuliang Guo, Xinyu Huang 0001, Liu Ren 0001, Xiaoming Liu 0002
ICCV4
2025 SMART: Advancing Scalable Map Priors for Driving Topology Reasoning
abstract
Topology reasoning is crucial for autonomous driving as it enables comprehensive understanding of connec-tivity and relationships between lanes and traffic elements. While recent approaches have shown success in perceiving driving topology using vehicle-mounted sensors, their scalability is hindered by the reliance on training data captured by consistent sensor configurations. We identify that the key factor in scalable lane perception and topology reasoning is the elimination of this sensor-dependent feature. To address this, we propose SMART, a scalable solution that leverages easily available standard-definition (SD) and satellite maps to learn a map prior model, supervised by large-scale geo-referenced high-definition (HD) maps independent of sensor settings. Attributed to scaled training, SMART alone achieves superior offline lane topology understanding using only SD and satellite inputs. Extensive experiments further demonstrate that SMART can be seamlessly integrated into any online topology reasoning methods, yielding significant improvements of up to 28% on the OpenLane-V2 benchmark. Project page: https://jay-ye.github.io/smart.
Junjie Ye 0007, David Paz, Hengyuan Zhang 0001, Yuliang Guo, Xinyu Huang 0001, Henrik I. Christensen, Yue Wang 0041, Liu Ren 0001
ICRA5
2025 MapGS: Generalizable Pretraining and Data Augmentation for Online Mapping via Novel View Synthesis
abstract
Online mapping reduces the reliance of au-tonomous vehicles on high-definition (HD) maps, significantly enhancing scalability. However, recent advancements often overlook cross-sensor configuration generalization, leading to performance degradation when models are deployed on vehicles with different camera intrinsics and extrinsics. With the rapid evolution of novel view synthesis methods, we investigate the extent to which these techniques can be leveraged to address the sensor configuration generalization challenge. We propose a novel framework leveraging Gaussian splatting to reconstruct scenes and render camera images in target sensor configurations. The target config sensor data, along with labels mapped to the target config, are used to train online mapping models. Our proposed framework on the nuScenes and Ar-goverse 2 datasets demonstrates a performance improvement of 18 % through effective dataset augmentation, achieves faster convergence and efficient training, and exceeds state-of-the-art performance when using only 25 % of the original training data. This enables data reuse and reduces the need for laborious data labeling. Project page at https://henryzhangzhy.github.io/mapgs.
Hengyuan Zhang 0001, David Paz, Yuliang Guo, Xinyu Huang 0001, Henrik I. Christensen, Liu Ren 0001
IV4
2024 SeaBird: Segmentation in Bird's View with Dice Loss Improves Monocular 3D Detection of Large Objects
abstract
Monocular 3D detectors achieve remarkable performance on cars and smaller objects. However, their performance drops on larger objects, leading to fatal accidents. Some attribute the failures to training data scarcity or the receptive field requirements of large objects. In this paper, we highlight this understudied problem of generalization to large objects. We find that modern frontal detectors struggle to generalize to large objects even on nearly balanced datasets. We argue that the cause of failure is the sensitivity of depth regression losses to noise of larger objects. To bridge this gap, we comprehensively investigate regression and dice losses, examining their robustness under varying error levels and object sizes. We mathematically prove that the dice loss leads to superior noise-robustness and model convergence for large objects compared to regression losses for a simplified case. Leveraging our theoretical insights, we propose SeaBird (Segmentation in Bird's View) as the first step towards generalizing to large objects. SeaBird effectively integrates BEV segmentation on foreground objects for 3D detection, with the segmentation head trained with the dice loss. SeaBird achieves SoTA results on the KITTI-360 leaderboard and improves existing detectors on the nuScenes leaderboard, particularly for large objects.
Abhinav Kumar 0004, Yuliang Guo, Xinyu Huang 0001, Liu Ren 0001, Xiaoming Liu 0002
CVPR3
2024 Behind the Veil: Enhanced Indoor 3D Scene Reconstruction with Occluded Surfaces Completion
abstract
In this paper, we present a novel indoor 3D reconstruction method with occluded surface completion, given a sequence of depth readings. Prior state-of-the-art (SOTA) methods only focus on the reconstruction of the visible areas in a scene, neglecting the invisible areas due to the occlusions, e.g., the contact surface between furniture, occluded wall and floor. Our method tackles the task of completing the occluded scene surfaces, resulting in a complete 3D scene mesh. The core idea of our method is learning 3D geometry prior from various complete scenes to infer the occluded geometry of an unseen scene from solely depth measurements. We design a coarse-fine hierarchical octree representation coupled with a dual-decoder architecture, i.e., Geo-decoder and 3D Inpainter, which jointly reconstructs the complete 3D scene geometry. The Geo-decoder with detailed representation at fine levels is optimized online for each scene to reconstruct visible surfaces. The 3D Inpainter with abstract representation at coarse levels is trained offline using various scenes to complete occluded surfaces. As a result, while the Geo-decoder is specialized for an individual scene, the 3D Inpainter can be generally applied across different scenes. We evaluate the proposed method on the 3D Completed Room Scene (3D-CRS) and iTHOR datasets, significantly outperforming the SOTA methods by a gain of 16.8% and 24.2% in terms of the completeness of 3D reconstruction. 3D-CRS dataset including a complete 3D mesh of each scene is provided on project webpage11https://github.com/BoschRHI3NA/3D-CRS-dataset.
Su Sun, Cheng Zhao 0002, Yuliang Guo, Ruoyu Wang 0012, Xinyu Huang 0001, Victor Y. Chen, Liu Ren 0001
CVPR5
2024 SUP-NeRF: A Streamlined Unification of Pose Estimation and NeRF for Monocular 3D Object Reconstruction
Yuliang Guo, Abhinav Kumar 0004, Cheng Zhao 0002, Ruoyu Wang 0012, Xinyu Huang 0001, Liu Ren 0001
ECCV (69)5
2024 TCLC-GS: Tightly Coupled LiDAR-Camera Gaussian Splatting for Autonomous Driving: Supplementary Materials
Cheng Zhao 0002, Su Sun, Ruoyu Wang 0012, Yuliang Guo, Jun-Jun Wan, Xinyu Huang 0001, Victor Y. Chen, Liu Ren 0001
ECCV (63)7
2024 Enhancing Online Road Network Perception and Reasoning with Standard Definition Maps
abstract
Autonomous driving for urban and highway driving applications often requires High Definition (HD) maps to generate a navigation plan. Nevertheless, various challenges arise when generating and maintaining HD maps at scale. While recent online mapping methods have started to emerge, their performance especially for longer ranges is limited by heavy occlusion in dynamic environments. With these considerations in mind, our work focuses on leveraging lightweight and scalable priors–Standard Definition (SD) maps–in the development of online vectorized HD map representations. We first examine the integration of prototypical rasterized SD map representations into various online mapping architectures. Furthermore, to identify lightweight strategies, we extend the OpenLane-V2 dataset with OpenStreetMaps and evaluate the benefits of graphical SD map representations. A key finding from designing SD map integration components is that SD map encoders are model agnostic and can be quickly adapted to new architectures that utilize bird’s eye view (BEV) encoders. Our results show that making use of SD maps as priors for the online mapping task can significantly speed up convergence and boost the performance of the online centerline perception task by 30% (mAP). Furthermore, we show that the introduction of the SD maps leads to a reduction of the number of parameters in the perception and reasoning task by leveraging SD map graphs while improving the overall performance. Project Page: https://henryzhangzhy.github.io/sdhdmap/.
Hengyuan Zhang 0001, David Paz, Yuliang Guo, Arun Das 0007, Xinyu Huang 0001, Karsten Haug, Henrik I. Christensen, Liu Ren 0001
IROS5
2023 3D Copy-Paste: Physically Plausible Object Insertion for Monocular 3D Detection
abstract
A major challenge in monocular 3D object detection is the limited diversity and quantity of objects in real datasets. While augmenting real scenes with virtual objects holds promise to improve both the diversity and quantity of the objects, it remains elusive due to the lack of an effective 3D object insertion method in complex real captured scenes. In this work, we study augmenting complex real indoor scenes with virtual objects for monocular 3D object detection. The main challenge is to automatically identify plausible physical properties for virtual assets (e.g., locations, appearances, sizes, etc.) in cluttered real scenes. To address this challenge, we propose a physically plausible indoor 3D object insertion approach to automatically copy virtual objects and paste them into real scenes. The resulting objects in scenes have 3D bounding boxes with plausible physical locations and appearances. In particular, our method first identifies physically feasible locations and poses for the inserted objects to prevent collisions with the existing room layout. Subsequently, it estimates spatially-varying illumination for the insertion location, enabling the immersive blending of the virtual objects into the original scene with plausible appearances and cast shadows. We show that our augmentation method significantly improves existing monocular 3D object models and achieves state-of-the-art performance. For the first time, we demonstrate that a physically plausible 3D object insertion, serving as a generative data augmentation technique, can lead to significant improvements for discriminative downstream tasks such as monocular 3D object detection. Project website: https://gyhandy.github.io/3D-Copy-Paste/.
Yunhao Ge, Hong-Xing Yu, Cheng Zhao 0002, Yuliang Guo, Xinyu Huang 0001, Liu Ren 0001, Laurent Itti, Jiajun Wu 0001
NeurIPS5
2022 Blind Removal of Facial Foreign Shadows
Yaojie Liu, Andrew Z. Hou, Xinyu Huang 0001, Liu Ren 0001, Xiaoming Liu 0002
BMVC3
2022 OmniFusion: 360 Monocular Depth Estimation via Geometry-Aware Fusion
abstract
A well-known challenge in applying deep-learning methods to omnidirectional images is spherical distortion. In dense regression tasks such as depth estimation, where structural details are required, using a vanilla CNN layer on the distorted 360 image results in undesired information loss. In this paper, we propose a 360 monocular depth estimation pipeline, OmniFusion, to tackle the spherical distortion issue. Our pipeline transforms a 360 image into less-distorted perspective patches (i.e. tangent images) to obtain patch-wise predictions via CNN, and then merge the patch-wise results for final output. To handle the discrepancy between patch-wise predictions which is a major issue affecting the merging quality, we propose a new framework with the following key components. First, we propose a geometry-aware feature fusion mechanism that combines 3D geometric features with 2D image features to compensate for the patch-wise discrepancy. Second, we employ the self-attention-based transformer architecture to conduct a global aggregation of patch-wise information, which further improves the consistency. Last, we introduce an iterative depth refinement mechanism, to further refine the estimated depth based on the more accurate geometric features. Experiments show that our method greatly mitigates the distortion issue, and achieves state-of-the-art performances on several 360 monocular depth estimation benchmark datasets. Our code is available at https://github.com/yuyanli0831/OmniFusion.
Yuliang Guo, Zhixin Yan, Xinyu Huang 0001, Ye Duan, Liu Ren 0001
CVPR4
2022 Symmetry and Uncertainty-Aware Object SLAM for 6DoF Object Pose Estimation
abstract
We propose a keypoint-based object-level SLAM framework that can provide globally consistent 6DoF pose estimates for symmetric and asymmetric objects alike. To the best of our knowledge, our system is among the first to utilize the camera pose information from SLAM to provide prior knowledge for tracking keypoints on symmetric objects - ensuring that new measurements are consistent with the current 3D scene. Moreover, our semantic key-point network is trained to predict the Gaussian covariance for the keypoints that captures the true error of the prediction, and thus is not only useful as a weight for the residuals in the system's optimization problems, but also as a means to detect harmful statistical outliers without choosing a manual threshold. Experiments show that our method provides competitive performance to the state of the art in 6DoF object pose estimation, and at a real-time speed. Our code, pre-trained models, and keypoint labels are available https://github.com/rpng/suo_slam.
Nathaniel W. Merrill, Yuliang Guo, Xingxing Zuo 0001, Xinyu Huang 0001, Stefan Leutenegger, Liu Ren 0001, Guoquan Huang 0001
CVPR4
2022 Part-Level Car Parsing and Reconstruction in Single Street View Images
abstract
Part information has been proven to be resistant to occlusions and viewpoint changes, which are main difficulties in car parsing and reconstruction. However, in the absence of datasets and approaches incorporating car parts, there are limited works that benefit from it. In this paper, we propose the first part-aware approach for joint part-level car parsing and reconstruction in single street view images. Without labor-intensive part annotations on real images, our approach simultaneously estimates pose, shape, and semantic parts of cars. There are two contributions in this paper. First, our network introduces dense part information to facilitate pose and shape estimation, which is further optimized with a novel 3D loss. To obtain part information in real images, a class-consistent method is introduced to implicitly transfer part knowledge from synthesized images. Second, we construct the first high-quality dataset containing 348 car models with physical dimensions and part annotations. Given these models, 60K synthesized images with randomized configurations are generated. Experimental results demonstrate that part knowledge can be effectively transferred with our class-consistent method, which significantly improves part segmentation performance on real street views. By fusing dense part information, our pose and shape estimation results achieve the state-of-the-art performance on the ApolloCar3D and outperform previous approaches by large margins in terms of both A3DP-Abs and A3DP-Rel.
Qichuan Geng, Hong Zhang 0009, Feixiang Lu, Xinyu Huang 0001, Sen Wang 0003, Zhong Zhou, Ruigang Yang
IEEE Trans. Pattern Anal. Mach. Intell.4
2020 The ApolloScape Open Dataset for Autonomous Driving and Its Application
abstract
Autonomous driving has attracted tremendous attention especially in the past few years. The key techniques for a self-driving car include solving tasks like 3D map construction, self-localization, parsing the driving road and understanding objects, which enable vehicles to reason and act. However, large scale data set for training and system evaluation is still a bottleneck for developing robust perception models. In this paper, we present the ApolloScape dataset [1] and its applications for autonomous driving. Compared with existing public datasets from real scenes, e.g., KITTI [2] or Cityscapes [3] , ApolloScape contains much large and richer labelling including holistic semantic dense point cloud for each site, stereo, per-pixel semantic labelling, lanemark labelling, instance segmentation, 3D car instance, high accurate location for every frame in various driving videos from multiple sites, cities and daytimes. For each task, it contains at lease 15x larger amount of images than SOTA datasets. To label such a complete dataset, we develop various tools and algorithms specified for each task to accelerate the labelling process, such as joint 3D-2D segment labeling, active labelling in videos etc. Depend on ApolloScape, we are able to develop algorithms jointly consider the learning and inference of multiple tasks. In this paper, we provide a sensor fusion scheme integrating camera videos, consumer-grade motion sensors (GPS/IMU), and a 3D semantic map in order to achieve robust self-localization and semantic segmentation for autonomous driving. We show that practically, sensor fusion and joint learning of multiple tasks are beneficial to achieve a more robust and accurate system. We expect our dataset and proposed relevant algorithms can support and motivate researchers for further development of multi-sensor fusion and multi-task learning in the field of computer vision.
Xinyu Huang 0001, Peng Wang 0001, Xinjing Cheng, Dingfu Zhou, Qichuan Geng, Ruigang Yang
IEEE Trans. Pattern Anal. Mach. Intell.1
2019 Mask-off: Synthesizing Face Images in the Presence of Head-mounted Displays
abstract
Wearable VR/AR devices provide users with fully immersive experience in a virtual environment, enabling possibilities to reshape the forms of entertainment and telepresence. While the body language is a crucial element in effective communication, wearing a head-mounted display (HMD) could severely hinder the eye contact and block facial expressions. We present a novel headset removal technique that enables high-quality occlusion-free communication in virtual environment. In particular, our solution synthesizes photoreal faces in the occluded region with faithful reconstruction of facial expressions and eye movements. Towards this goal, we develop a novel capture setup that consists of two near-infrared (NIR) cameras inside the HMD for eye capturing and one external RGB camera for recording visible face regions. To enable realistic face synthesis with consistent illuminations, we propose a data-driven approach to fuse the narrow-field-of-view NIR images with the RGB image captured from the external camera. In addition, to generate pho-torealistic eyes, a dedicated algorithm is proposed to colorize the NIR eye images and further rectify the color distortion caused by the non-linear mapping of IR light sensitivity. Experimental results demonstrate that our framework is capable to synthesize high-fidelity unoccluded facial images with accurate tracking of head motion, facial expression and eye movement.
Qingguo Xu, Weikai Chen 0001, Jun Xing, Xinyu Huang 0001, Ruigang Yang
VR6
2017 Multiscale spatially regularised correlation filters for visual tracking
abstract
Recently, discriminative correlation filter based trackers have achieved extremely successful results in many competitions and benchmarks. These methods utilise a periodic assumption of the training samples to efficiently learn a classifier. However, this assumption will produce unwanted boundary effects which severely degrade the tracking performance. Correlation filters with limited boundaries and spatially regularised discriminative correlation filters were proposed to reduce boundary effects. However, their methods use the fixed scale mask or pre‐designed weights function, respectively, which are unsuitable for large‐scale variation. In this study, the authors proposed multiscale spatially regularised correlation filters (MSRCF) for visual tracking. The authors’ augmented objective can reduce the boundary effect even in large‐scale variation, leading to more discriminative model. The proposed multiscale regularisation matrix makes MSRCF fast convergence. The authors’ online tracking algorithm performs favourably against state‐of‐the‐art trackers on OTB‐2013 and OTB‐2015 Benchmark in terms of efficiency, accuracy and robustness.
Xiaodong Gu 0006, Xinyu Huang 0001, Alade O. Tokuta
IET Comput. Vis.2
2015 Fusion of color images and LiDAR data for lane classification
abstract
Lane classification is a fundamental problem for autonomous driving and map-aided localization. Many existing algorithms rely on special designed 1D or 2D filters to extract features of lane markings from either color images or LiDAR data. However, these handcrafted features could not be robust under various driving and lighting conditions.
Xiaodong Gu 0006, Andi Zang, Xinyu Huang 0001, Alade O. Tokuta
SIGSPATIAL/GIS3
2015 Counting and Classification of Highway Vehicles by Regression Analysis
abstract
In this paper, we describe a novel algorithm that counts and classifies highway vehicles based on regression analysis. This algorithm requires no explicit segmentation or tracking of individual vehicles, which is usually an important part of many existing algorithms. Therefore, this algorithm is particularly useful when there are severe occlusions or vehicle resolution is low, in which extracted features are highly unreliable. There are mainly two contributions in our proposed algorithm. First, a warping method is developed to detect the foreground segments that contain unclassified vehicles. The common used modeling and tracking (e.g., Kalman filtering) of individual vehicles are not required. In order to reduce vehicle distortion caused by the foreshortening effect, a nonuniform mesh grid and a projective transformation are estimated and applied during the warping process. Second, we extract a set of low-level features for each foreground segment and develop a cascaded regression approach to count and classify vehicles directly, which has not been used in the area of intelligent transportation systems. Three different regressors are designed and evaluated. Experiments show that our regression-based algorithm is accurate and robust for poor quality videos, from which many existing algorithms could fail to extract reliable features.
Mingpei Liang, Xinyu Huang 0001, Chung-Hao Chen, Alade O. Tokuta
IEEE Trans. Intell. Transp. Syst.2
2014 Video face beautification
abstract
This paper presents a novel system framework of face beautification. Unlike prior works that deal with single images, the proposed beautification framework is designed for an input video and it is able to improve both the appearance and the shape of a face. Our system adopts a state-of-the-art algorithm to synthesize and track 3D face models using blendshapes. The personalized 3D model can be edited to satisfy personal preference. This interactive process is needed only once per subject. Based on the tracking result and the modified face model, we present an algorithm to beautify the face video efficiently and consistently. Furthermore we develop a variant of content preserving warping to reduce warping distortions along the face boundary. Finally we adopt real time bilateral filtering to remove wrinkles, freckles, and unwanted blemishes. This framework is evaluated on a set of videos. The experiments demonstrate that our framework can generate consistent and pleasant results over video frames while the original expressions and features are persevered naturally.
Xinyu Huang 0001, Jizhou Gao, Alade O. Tokuta, Cha Zhang, Ruigang Yang
ICME2
2014 A Performance Comparison between Circular and Spline-Based Methods for Iris Segmentation
abstract
Iris segmentation is an important module of iris recognition that can substantially affect recognition performance. Since iris and pupil boundaries usually are not exactly circular, spline-based methods have been used to model irregular iris and pupil boundaries recently. However, in most existing methods, many other factors or modules in the iris recognition pipeline are evaluated together and their mixed effects are assumed to be negligible. More importantly, the splines that model irregularity of the boundaries could not be enough to model the internal nonlinear deformations of an iris pattern (e.g., caused by iris dilation). As a result, it remains unclear whether spline-based methods can provide significant improvements. In this paper, we conduct a complete performance comparison between circular and spline-based methods. There are mainly two contributions. Firstly, for the purpose of comparison, we propose a spline estimator that is robust to outliers caused by eyelashes, eyelids, highlights, and shadows. Secondly, we analyze the relation between iris matching distances and segmentation results by using circular and spline-based methods. Based on our experiments, we found that, even with the proposed robust spline estimator, the improvement of recognition performance is still limited (around 6%). Therefore, in case that less robust spline estimators are used due to the real-time requirement in practical systems, the actual recognition improvement by using splines could be far below the expectation.
Changpeng Ti, Xinyu Huang 0001, Alade O. Tokuta, Ruigang Yang
ICPR3
2013 An experimental study of pupil constriction for liveness detection
abstract
As iris recognition systems have been deployed in many security areas, liveness detection that can distinguish between real iris patterns and fake ones becomes an important module. Most existing algorithms focus on the appearance difference between real and fake iris (for example, printed patterns, cosmetic contact lenses etc.) which is a very difficult problem. Instead of studying image properties of fake irises, we show that pupil constriction, the fundamental characteristic of real and live irises, can be very robust for liveness detection. In this experimental study, we first build an iris acquisition system that can acquire two eye images under two different illumination conditions in a less intrusive environment. Second, in order to model the process of pupil constriction, we propose a feature descriptor that consists of similarity measurement between iris patches and ratio of iris and pupil diameters. Third, the performance of liveness prediction is evaluated based on the training of a Support Vector Machine (SVM) classifier. The high success prediction rate shows that the classifier is effective without knowing any prior knowledge of fake irises.
Xinyu Huang 0001, Changpeng Ti, Qi-zhen Hou, Alade O. Tokuta, Ruigang Yang
WACV1
2013 Measurement of mirror surfaces using specular reflection and analytical computation
Zhen Zhou Wang, Xinyu Huang 0001, Ruigang Yang
Mach. Vis. Appl.2
2011 Interreflection removal for photometric stereo by using spectrum-dependent albedo
abstract
We present a novel method that can separate m-bounced light and remove the interreflections in a photometric stereo setup. Under the assumption of a uniformly colored lambertian surface, the intensity of a point in the scene is the sum of 1-bounced light through m-bounced light rays. Ruled by the law of diffuse reflection, whenever a light ray is bounced by the surface, its intensity will be attenuated by the factor of albedo ρ. This implies that the measured intensity value can be written as a polynomial function of ρ, and the intensity contribution of the m-bounced light rays are expressed by the term of ρm. Therefore, when we change the surface albedo, the intensity of the m-bounced light is changed to the order of m. This non-linearity gives us the possibility to separate the m-bounced light. In practice, we illuminate the scene with different light colors to effectively simulate different surface albedos since albedo is spectrum dependent. Once the m-bounced light rays are separated, we can perform the photometric stereo algorithm on the 1-bounced light (direct lighting) images to produce the 3D shape without the impact of interreflections. Experiments have shown that we get significantly improved scene reconstruction with a minimum of two color images.
Miao Liao, Xinyu Huang 0001, Ruigang Yang
CVPR2
2009 Manifold Estimation in View-Based Feature Space for Face Synthesis across Poses
Xinyu Huang 0001, Jizhou Gao, Sen-Ching S. Cheung, Ruigang Yang
ACCV (1)1
2009 Image deblurring for less intrusive iris capture
abstract
For most iris capturing scenarios, captured iris images could easily blur when the user is out of the depth of field (DOF) of the camera, or when he or she is moving. The common solution is to let the user try the capturing process again as the quality of these blurred iris images is not good enough for recognition. In this paper, we propose a novel iris deblurring algorithm that can be used to improve the robustness and nonintrusiveness for iris capture. Unlike other iris deblurring algorithms, the key feature of our algorithm is that we use the domain knowledge inherent in iris images and iris capture settings to improve the performance, which could be in the form of iris image statistics, characteristics of pupils or highlights, or even depth information from the iris capturing system itself. Our experiments on both synthetic and real data demonstrate that our deblurring algorithm can significantly restore blurred iris patterns and therefore improve the robustness of iris capture.
Xinyu Huang 0001, Liu Ren 0001, Ruigang Yang
CVPR1
2008 Illumination and Person-Insensitive Head Pose Estimation Using Distance Metric Learning
Xianwang Wang, Xinyu Huang 0001, Jizhou Gao, Ruigang Yang
ECCV (2)2
2008 Toward the Light Field Display: Autostereoscopic Rendering via a Cluster of Projectors
abstract
Ultimately, a display device should be capable of reproducing the visual effects observed in reality. In this paper we introduce an autostereoscopic display that uses a scalable array of digital light projectors and a projection screen augmented with microlenses to simulate a light field for a given three-dimensional scene. Physical objects emit or reflect light in all directions to create a light field that can be approximated by the light field display. The display can simultaneously provide many viewers from different viewpoints a stereoscopic effect without head tracking or special viewing glasses. This work focuses on two important technical problems related to the light field display; calibration and rendering. We present a solution to automatically calibrate the light field display using a camera and introduce two efficient algorithms to render the special multi-view images by exploiting their spatial coherence. The effectiveness of our approach is demonstrated with a four-projector prototype that can display dynamic imagery with full parallax.
Ruigang Yang, Xinyu Huang 0001, Sifang Li, Christopher O. Jaynes
IEEE Trans. Vis. Comput. Graph.2
2007 Calibrating Pan-Tilt Cameras with Telephoto Lenses
Xinyu Huang 0001, Jizhou Gao, Ruigang Yang
ACCV (1)1