Zihan Zhou 0001

dblp:00/6525-1 · DBLP profile ↗
← Back
41ranked-venue papers
7as first author
12since 2021 · last 2026
0000-0002-1697-2168ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 33 · 6 first-author · 8 since 2021Artificial intelligence and machine learning · 28 · 4 first-author · 8 since 2021Systems, architecture and hardware · 3 · 2 since 2021Databases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 SPATIALGEN: Layout-Guided 3D Indoor Scene Generation
Chuan Fang, Heng Li 0009, Yixun Liang, Jia Zheng 0002, Yongsen Mao, Yuan Liu 0025, Rui Tang 0015, Zihan Zhou 0001, Ping Tan 0002
3DV8
2025 From 2D CAD Drawings to 3D Parametric Models: A Vision-Language Approach
abstract
In this paper, we present CAD2Program, a new method for reconstructing 3D parametric models from 2D CAD drawings. Our proposed method is inspired by recent successes in vision-language models (VLMs), and departs from traditional methods which rely on task-specific data representations and/or algorithms. Specifically, on the input side, we simply treat the 2D CAD drawing as a raster image, regardless of its original format, and encode the image with a standard ViT model. We show that such an encoding scheme achieves competitive performance against existing methods that operate on vector-graphics inputs, while imposing substantially fewer restrictions on the 2D drawings. On the output side, our method auto-regressively predicts a general-purpose language describing 3D parametric models in text form. Compared to other sequence modeling methods for CAD which use domain-specific sequence representations with fixed-size slots, our text-based representation is more flexible, and can be easily extended to arbitrary geometric entities and semantic or functional properties. Experimental results on a large-scale dataset of cabinet models demonstrate the effectiveness of our method.
Xilin Wang, Jia Zheng 0002, Yuanchao Hu, Hao Zhu 0004, Zihan Zhou 0001
AAAI6
2025 SpatialLM: Training Large Language Models for Structured Indoor Modeling
abstract
SpatialLM is a large language model designed to process 3D point cloud data and generate structured 3D scene understanding outputs. These outputs include architectural elements like walls, doors, windows, and oriented object boxes with their semantic categories. Unlike previous methods which exploit task-specific network designs, our model adheres to the standard multimodal LLM architecture and is fine-tuned directly from open-source LLMs. To train SpatialLM, we collect a large-scale, high-quality synthetic dataset consisting of the point clouds of 12,328 indoor scenes (54,778 rooms) with ground-truth 3D annotations, and conduct a careful study on various modeling and training decisions. On public benchmarks, our model gives state-of-the-art performance in layout estimation and competitive results in 3D object detection. With that, we show a feasible path for enhancing the spatial understanding capabilities of modern LLMs for applications in augmented reality, embodied robotics, and more.
Yongsen Mao, Junhao Zhong, Chuan Fang, Jia Zheng 0002, Rui Tang 0015, Hao Zhu 0004, Ping Tan 0002, Zihan Zhou 0001
NeurIPS8
2025 Computer-Aided Layout Generation for Building Design: A Review
abstract
Generating realistic building layouts for automatic building design has been studied in both computer vision and architectural domains. Traditional approaches in the latter, which are based on optimization techniques or heuristic design guidelines, can synthesize desirable layouts, but usually require post-processing and involve human interaction in the design pipeline, making them costly and time-consuming. The advent of deep generative models has significantly improved the fidelity and diversity of the generated architecture layouts, reducing the workload of designers and making the process much more efficient. This paper presents a comprehensive review of three major research topics in architectural layout design and generation: floorplan layout generation, scene layout synthesis, and generation of various other formats of building layouts. For each topic, we overview the leading paradigms, categorized either by research domains (architecture or machine learning) or by user input conditions or constraints. We then introduce commonly-adopted benchmark datasets used to verify the effectiveness of the methods, as well as corresponding evaluation metrics. Finally, we identify the well-solved problems and limitations of existing approaches, and then propose promising directions for future research. This survey has an associated project which aims to maintain the resources, at https://github.com/jcliu0428/awesome-building-layout-generation.
Yuan Xue 0002, Haomiao Ni, Rui Yu 0002, Zihan Zhou 0001, Sharon X. Huang
Comput. Vis. Media5
2024 NeRF-Enhanced Outpainting for Faithful Field-of-View Extrapolation
abstract
In various applications, such as robotic navigation and remote visual assistance, expanding the field of view (FOV) of the camera proves beneficial for enhancing environmental perception. Unlike image outpainting techniques aimed solely at generating aesthetically pleasing visuals, these applications demand an extended view that faithfully represents the scene. To achieve this, we formulate a new problem of faithful FOV extrapolation that utilizes a set of pre-captured images as prior knowledge of the scene. To address this problem, we present a simple yet effective solution called NeRF-Enhanced Outpainting (NEO) that uses extended-FOV images generated through NeRF to train a scene-specific image outpainting model. To assess the performance of NEO, we conduct comprehensive evaluations on three photorealistic datasets and one real-world dataset. Extensive experiments on the benchmark datasets showcase the robustness and potential of our method in addressing this challenge. We believe our work lays a strong foundation for future exploration within the research community.
Rui Yu 0002, Zihan Zhou 0001, Sharon X. Huang
ICRA3
2023 PlankAssembly: Robust 3D Reconstruction from Three Orthographic Views with Learnt Shape Programs
abstract
In this paper, we develop a new method to automatically convert 2D line drawings from three orthographic views into 3D CAD models. Existing methods for this problem reconstruct 3D models by back-projecting the 2D observations into 3D space while maintaining explicit correspondence between the input and output. Such methods are sensitive to errors and noises in the input, thus often fail in practice where the input drawings created by human designers are imperfect. To overcome this difficulty, we leverage the attention mechanism in a Transformer-based sequence generation model to learn flexible mappings between the input and output. Further, we design shape programs which are suitable for generating the objects of interest to boost the reconstruction accuracy and facilitate CAD modeling applications. Experiments on a new benchmark dataset show that our method significantly outperforms existing ones when the inputs are noisy or incomplete.
Jia Zheng 0002, Zixin Zhang 0002, Xiaojun Yuan 0002, Jian Yin 0001, Zihan Zhou 0001
ICCV6
2023 Two-stage Content-Aware Layout Generation for Poster Designs
abstract
Automatic layout generation models can generate numerous design layouts in a few seconds, which significantly reduces the amount of repetitive work for designers. However, most of these models consider the layout generation task as arranging layout elements with different attributes on a blank canvas, thus struggle to handle the case when an image is used as the layout background. Additionally, existing layout generation models often fail to incorporate explicit aesthetic principles such as alignment and non-overlap, and neglect implicit aesthetic principles which are hard to model. To address these issues, this paper proposes a two-stage content-aware layout generation framework for poster layout generation. Our framework consists of an aesthetics-conditioned layout generation module and a layout ranking module. The diffusion model based layout generation module utilizes an aesthetics-guided layout denoising process to sample layout proposals that meet explicit aesthetic constraints. The Auto-Encoder based layout ranking module then measures the distance between those proposals and real designs to determine the layout that best meets implicit aesthetic principles. Quantitative and qualitative experiments demonstrate that our method outperforms state-of-the-art content-aware layout generation models.
Shang Chai, Liansheng Zhuang, Fengying Yan, Zihan Zhou 0001
ACM Multimedia4
2022 Neural Face Identification in a 2D Wireframe Projection of a Manifold Object
abstract
In computer-aided design (CAD) systems, 2D line drawings are commonly used to illustrate 3D object designs. To reconstruct the 3D models depicted by a single 2D line drawing, an important key is finding the edge loops in the line drawing which correspond to the actual faces of the 3D object. In this paper, we approach the classical problem of face identification from a novel data-driven point of view. We cast it as a sequence generation problem: starting from an arbitrary edge, we adopt a variant of the popular Transformer model to predict the edges associated with the same face in a natural order. This allows us to avoid searching the space of all possible edge loops with various handcrafted rules and heuristics as most existing methods do, deal with challenging cases such as curved surfaces and nested edge loops, and leverage additional cues such as face types. We further discuss how possibly imperfect predictions can be used for 3D object reconstruction. The project page is at https://manycore-research.github.io/faceformer.
Jia Zheng 0002, Zihan Zhou 0001
CVPR3
2022 Deep Depth from Focus with Differential Focus Volume
abstract
Depth-from-focus (DFF) is a technique that infers depth using the focus change of a camera. In this work, we propose a convolutional neural network (CNN) to find the best-focused pixels in a focal stack and infer depth from the focus estimation. The key innovation of the network is the novel deep differential focus volume (DFV). By computing the first-order derivative with the stacked features over different focal distances, DFV is able to capture both the focus and context information for focus analysis. Besides, we also introduce a probability regression mechanism for focus estimation to handle sparsely sampled focal stacks and provide uncertainty estimation to the final prediction. Comprehensive experiments demonstrate that the proposed model achieves state-of-the-art performance on multiple datasets with good generalizability and fast speed.
Fengting Yang, Sharon X. Huang, Zihan Zhou 0001
CVPR3
2022 End-to-End Graph-Constrained Vectorized Floorplan Generation with Panoptic Refinement
Yuan Xue 0002, José Pinto Duarte, Krishnendra Shekhawat, Zihan Zhou 0001, Sharon X. Huang
ECCV (15)5
2022 Iterative Design and Prototyping of Computer Vision Mediated Remote Sighted Assistance
abstract
Remote sighted assistance (RSA) is an emerging navigational aid for people with visual impairments (PVI). Using scenario-based design to illustrate our ideas, we developed a prototype showcasing potential applications for computer vision to support RSA interactions. We reviewed the prototype demonstrating real-world navigation scenarios with an RSA expert, and then iteratively refined the prototype based on feedback. We reviewed the refined prototype with 12 RSA professionals to evaluate the desirability and feasibility of the prototyped computer vision concepts. The RSA expert and professionals were engaged by, and reacted insightfully and constructively to the proposed design ideas. We discuss what we learned about key resources, goals, and challenges of the RSA prosthetic practice through our iterative prototype review, as well as implications for the design of RSA systems and the integration of computer vision technologies into RSA.
Jingyi Xie 0001, Madison Reddie, Sooyeon Lee, Syed Masum Billah, Zihan Zhou 0001, Chun-Hua Tsai, John M. Carroll 0001
ACM Trans. Comput. Hum. Interact.5
2021 Towards Robust Human Trajectory Prediction in Raw Videos
abstract
Human trajectory prediction has received increased attention lately due to its importance in applications such as autonomous vehicles and indoor robots. However, most existing methods make predictions based on human-labeled trajectories and ignore the errors and noises in detection and tracking. In this paper, we study the problem of human trajectory forecasting in raw videos, and show that the prediction accuracy can be severely affected by various types of tracking errors. Accordingly, we propose a simple yet effective strategy to correct the tracking failures by enforcing prediction consistency over time. The proposed "re-tracking" algorithm can be applied to any existing tracking and prediction pipelines. Experiments on public benchmark datasets demonstrate that the proposed method can improve both tracking and prediction performance in challenging real-world scenarios. The code and data are available at https://git.io/retracking-prediction.
Rui Yu 0002, Zihan Zhou 0001
IROS2
2020 Superpixel Segmentation With Fully Convolutional Networks
abstract
In computer vision, superpixels have been widely used as an effective way to reduce the number of image primitives for subsequent processing. But only a few attempts have been made to incorporate them into deep neural networks. One main reason is that the standard convolution operation is defined on regular grids and becomes inefficient when applied to superpixels. Inspired by an initialization strategy commonly adopted by traditional superpixel algorithms, we present a novel method that employs a simple fully convolutional network to predict superpixels on a regular image grid. Experimental results on benchmark datasets show that our method achieves state-of-the-art superpixel segmentation performance while running at about 50fps. Based on the predicted superpixels, we further develop a downsampling/upsampling scheme for deep networks with the goal of generating high-resolution outputs for dense prediction tasks. Specifically, we modify a popular network architecture for stereo matching to simultaneously predict superpixels and disparities. We show that improved disparity estimation accuracy can be obtained on public datasets.
Fengting Yang, Hailin Jin, Zihan Zhou 0001
CVPR4
2020 Neural Wireframe Renderer: Learning Wireframe to Image Translations
Yuan Xue 0002, Zihan Zhou 0001, Sharon X. Huang
ECCV (26)2
2020 Structured3D: A Large Photo-Realistic Dataset for Structured 3D Modeling
Jia Zheng 0002, Jing Li 0117, Rui Tang 0015, Shenghua Gao, Zihan Zhou 0001
ECCV (9)6
2020 Data-driven Distributed State Estimation and Behavior Modeling in Sensor Networks
abstract
Nowadays, the prevalence of sensor networks has enabled tracking of the states of dynamic objects for a wide spectrum of applications from autonomous driving to environmental monitoring and urban planning. However, tracking realworld objects often faces two key challenges: First, due to the limitation of individual sensors, state estimation needs to be solved in a collaborative and distributed manner. Second, the objects' movement behavior model is unknown, and needs to be learned using sensor observations. In this work, for the first time, we formally formulate the problem of simultaneous state estimation and behavior learning in a sensor network. We then propose a simple yet effective solution to this new problem by extending the Gaussian process-based Bayes filters (GPBayesFilters) to an online, distributed setting. The effectiveness of the proposed method is evaluated on tracking objects with unknown movement behaviors using both synthetic data and data collected from a multi-robot platform.
Rui Yu 0002, Zhenyuan Yuan, Zihan Zhou 0001
IROS4
2019 Single-Image Piece-Wise Planar 3D Reconstruction via Associative Embedding
abstract
Single-image piece-wise planar 3D reconstruction aims to simultaneously segment plane instances and recover 3D plane parameters from an image. Most recent approaches leverage convolutional neural networks (CNNs) and achieve promising results. However, these methods are limited to detecting a fixed number of planes with certain learned order. To tackle this problem, we propose a novel two-stage method based on associative embedding, inspired by its recent success in instance segmentation. In the first stage, we train a CNN to map each pixel to an embedding space where pixels from the same plane instance have similar embeddings. Then, the plane instances are obtained by grouping the embedding vectors in planar regions via an efficient mean shift clustering algorithm. In the second stage, we estimate the parameter for each plane instance by considering both pixel-level and instance-level consistencies. With the proposed method, we are able to detect an arbitrary number of planes. Extensive experiments on public datasets validate the effectiveness and efficiency of our method. Furthermore, our method runs at 30 fps at the testing time, thus could facilitate many real-time applications such as visual SLAM and human-robot interaction. Code is available at https://github.com/svip-lab/PlanarReconstruction.
Zehao Yu 0002, Jia Zheng 0002, Dongze Lian, Zihan Zhou 0001, Shenghua Gao
CVPR4
2018 Learning to Parse Wireframes in Images of Man-Made Environments
abstract
In this paper, we propose a learning-based approach to the task of automatically extracting a "wireframe" representation for images of cluttered man-made environments. The wireframe (see Fig. 1) contains all salient straight lines and their junctions of the scene that encode efficiently and accurately large-scale geometry and object shapes. To this end, we have built a very large new dataset of over 5,000 images with wireframes thoroughly labelled by humans. We have proposed two convolutional neural networks that are suitable for extracting junctions and lines with large spatial support, respectively. The networks trained on our dataset have achieved significantly better performance than state-of-the-art methods for junction detection and line segment detection, respectively. We have conducted extensive experiments to evaluate quantitatively and qualitatively the wireframes obtained by our method, and have convincingly shown that effectively and efficiently parsing wireframes for images of man-made environments is a feasible goal within reach. Such wireframes could benefit many important visual tasks such as feature correspondence, 3D reconstruction, vision-based mapping, localization, and navigation. The data and source code are available at https://github.com/huangkuns/wireframe.
Kun Huang 0001, Zihan Zhou 0001, Tianjiao Ding, Shenghua Gao, Yi Ma 0001
CVPR3
2018 Recovering 3D Planes from a Single Image via Convolutional Neural Networks
Fengting Yang, Zihan Zhou 0001
ECCV (10)2
2018 Discovering Triangles in Portraits for Supporting Photographic Creation
abstract
Incorporating the concept of triangles in photos is an effective composition technique used by professional photographers for making pictures more interesting or dynamic. Information on the locations of the embedded triangles is valuable for comparing the composition of portrait photos which can be further leveraged by a retrieval system or used by the photographers. This paper presents a system to automatically detect embedded triangles in portrait photos. The problem is challenging because the triangles used in portraits are often not clearly defined by straight lines. The system first extracts a set of filtered line segments as candidate triangle sides and then utilizes a modified random sample consensus algorithm to fit triangles onto the set of line segments. We propose two metrics Continuity Ratio and Total Ratio to evaluate the fitted triangles; those with high fitting scores are taken as detected triangles. Experimental results have demonstrated high accuracy in locating preeminent triangles in portraits without dependence on the camera or lens parameters. To demonstrate the benefits of our method to digital photography we have developed two novel applications that aim to help users compose high-quality photos. In the first application we develop a human position and pose recommendation system by retrieving and presenting compositionally similar photos taken by competent photographers. The second application is a novel sketch-based triangle retrieval system which searches for photos containing a specific triangular configuration. User studies have been conducted to validate the effectiveness of these approaches.
Siqiong He, Zihan Zhou 0001, Farshid Farhat, James Z. Wang 0001
IEEE Trans. Multim.2
2017 Multi-scale FCN with Cascaded Instance Aware Segmentation for Arbitrary Oriented Word Spotting in the Wild
abstract
Scene text detection has attracted great attention these years. Text potentially exist in a wide variety of images or videos and play an important role in understanding the scene. In this paper, we present a novel text detection algorithm which is composed of two cascaded steps: (1) a multi-scale fully convolutional neural network (FCN) is proposed to extract text block regions, (2) a novel instance (word or line) aware segmentation is designed to further remove false positives and obtain word instances. The proposed algorithm can accurately localize word or text line in arbitrary orientations, including curved text lines which cannot be handled in a lot of other frameworks. Our algorithm achieved state-of-the-art performance in ICDAR 2013 (IC13), ICDAR 2015 (IC15) and CUTE80 and Street View Text (SVT) benchmark datasets.
Dafang He, Xiao Yang 0004, Chen Liang 0001, Zihan Zhou 0001, Alexander Ororbia, Daniel Kifer, C. Lee Giles
CVPR4
2017 Improving Offline Handwritten Chinese Character Recognition by Iterative Refinement
abstract
We present an iterative refinement module that can be applied to the output feature maps of any existing convolutional neural networks in order to further improve classification accuracy. The proposed module, implemented by an attention-based recurrent neural network, can iteratively use its previous predictions to update attention and thereafter refine current predictions. In this way, the model is able to focus on a sub-region of input images to distinguish visually similar characters (see Figure 1 for an example). We evaluate its effectiveness on handwritten Chinese character recognition (HCCR) task and observe significant performance gain. HCCR task is challenging due to large number of classes and small differences between certain characters. To overcome these difficulties, we further propose a novel convolutional architecture that utilizes both low-level visual cues and high-level structural information. Together with the proposed iterative refinement module, our approach achieves an accuracy of 97.37%, outperforming previous methods that use raw images as input on ICDAR-2013 dataset [1].
Xiao Yang 0004, Dafang He, Zihan Zhou 0001, Daniel Kifer, C. Lee Giles
ICDAR3
2017 Learning to Read Irregular Text with Attention Mechanisms
abstract
We present a robust end-to-end neural-based model to attentively recognize text in natural images. Particularly, we focus on accurately identifying irregular (perspectively distorted or curved) text, which has not been well addressed in the previous literature. Previous research on text reading often works with regular (horizontal and frontal) text and does not adequately generalize to processing text with perspective distortion or curving effects. Our work proposes to overcome this difficulty by introducing two learning components: (1) an auxiliary dense character detection task that helps to learn text specific visual patterns, (2) an alignment loss that provides guidance to the training of an attention model. We show with experiments that these two components are crucial for achieving fast convergence and high classification accuracy for irregular text recognition. Our model outperforms previous work on two irregular-text datasets: SVT-Perspective and CUTE80, and is also highly-competitive on several regular-text datasets containing primarily horizontal and frontal text.
Xiao Yang 0004, Dafang He, Zihan Zhou 0001, Daniel Kifer, C. Lee Giles
IJCAI3
2017 Label Information Guided Graph Construction for Semi-Supervised Learning
abstract
In the literature, most existing graph-based semi-supervised learning methods only use the label information of observed samples in the label propagation stage, while ignoring such valuable information when learning the graph. In this paper, we argue that it is beneficial to consider the label information in the graph learning stage. Specifically, by enforcing the weight of edges between labeled samples of different classes to be zero, we explicitly incorporate the label information into the state-of-the-art graph learning methods, such as the low-rank representation (LRR), and propose a novel semi-supervised graph learning method called semi-supervised low-rank representation. This results in a convex optimization problem with linear constraints, which can be solved by the linearized alternating direction method. Though we take LRR as an example, our proposed method is in fact very general and can be applied to any self-representation graph learning methods. Experiment results on both synthetic and real data sets demonstrate that the proposed graph learning method can better capture the global geometric structure of the data, and therefore is more effective for semi-supervised learning tasks.
Liansheng Zhuang, Zihan Zhou 0001, Shenghua Gao, Jingwen Yin, Zhouchen Lin, Yi Ma 0001
IEEE Trans. Image Process.2
2017 Detecting Dominant Vanishing Points in Natural Scenes with Application to Composition-Sensitive Image Retrieval
abstract
Linear perspective is widely used in landscape photography to create the impression of depth on a 2D photo. Automated understanding of linear perspective in landscape photography has several real-world applications, including aesthetics assessment, image retrieval, and on-site feedback for photo composition, yet adequate automated understanding has been elusive. We address this problem by detecting the dominant vanishing point and the associated line structures in a photo. However, natural landscape scenes pose great technical challenges because often the number of strong edges converging to the dominant vanishing point is inadequate. To overcome this difficulty, we propose a novel vanishing point detection method that exploits global structures in the scene via contour detection. We show that our method significantly outperforms state-of-the-art methods on a public ground truth landscape image dataset that we have created. Based on the detection results, we further demonstrate how our approach to linear perspective understanding provides on-site guidance to amateur photographers on their work through a novel viewpoint-specific image retrieval system.
Zihan Zhou 0001, Farshid Farhat, James Z. Wang 0001
IEEE Trans. Multim.1
2016 Robust Plane-Based Calibration of Multiple Non-Overlapping Cameras
abstract
The availability of commodity multi-camera systems such as Google Jump, Jaunt, and Lytro Immerge have brought new demand for reliable and efficient extrinsic camera calibration. State-of-the-art solutions generally require that adjacent, if not all, cameras observe a common area or employ known scene structures. In this paper, we present a novel multi-camera calibration technique that eliminates such requirements. Our approach extends the single-pair hand-eye calibration used in robotics to multi-camera systems. Specifically, we make use of (possibly unknown) planar structures in the scene and combine plane-based structure from motion, camera pose estimation, and task-specific bundle adjustment for extrinsic calibration. Experiments on several multi-camera setups demonstrate that our scheme is highly accurate, robust, and efficient.
Zihan Zhou 0001, Ziran Xing, Yanbing Dong, Yi Ma 0001, Jingyi Yu 0001
3DV2
2016 Aggregating Local Context for Accurate Scene Text Detection
Dafang He, Xiao Yang 0004, Wenyi Huang, Zihan Zhou 0001, Daniel Kifer, C. Lee Giles
ACCV (5)4
2016 Detecting Arbitrary Oriented Text in the Wild with a Visual Attention Model
abstract
Text embedded in images provides important semantic information about a scene and its content. Detecting text in an unconstrained environment is a challenging task because of the many fonts, sizes, backgrounds, and alignments of the characters. We present a novel attention model for detecting arbitrary oriented and curved scene text. Inspired by the attention mechanisms in the human visual system, our model utilizes a spatial glimpse network to processes the attended area and deploys a recurrent neural network that aggregates the information over time to determine the attention movement. Combining this with an off-the-shelf region proposal method, the model achieves the state-of-the-art performance on the highly cited ICDAR2013 dataset, and the MSRA-TD500 dataset which contains arbitrary oriented text.
Wenyi Huang, Dafang He, Xiao Yang 0004, Zihan Zhou 0001, Daniel Kifer, C. Lee Giles
ACM Multimedia4
2015 Modeling Perspective Effects in Photographic Composition
abstract
Automatic understanding of photo composition is a valuable technology in multiple areas including digital photography, multimedia advertising, entertainment, and image retrieval. In this paper, we propose a method to model geometrically the compositional effects of linear perspective. Comparing with existing methods which have focused on basic rules of design such as simplicity, visual balance, golden ratio, and the rule of thirds, our new quantitative model is more comprehensive whenever perspective is relevant. We first develop a new hierarchical segmentation algorithm that integrates classic photometric cues with a new geometric cue inspired by perspective geometry. We then show how these cues can be used directly to detect the dominant vanishing point in an image without extracting any line segments, a technique with implications for multimedia applications beyond this work. Finally, we demonstrate an interesting application of the proposed method for providing on-site composition feedback through an image retrieval system.
Zihan Zhou 0001, Siqiong He, Jia Li 0001, James Z. Wang 0001
ACM Multimedia1
2013 Plane-Based Content Preserving Warps for Video Stabilization
abstract
Recently, a new image deformation technique called content-preserving warping (CPW) has been successfully employed to produce the state-of-the-art video stabilization results in many challenging cases. The key insight of CPW is that the true image deformation due to viewpoint change can be well approximated by a carefully constructed warp using a set of sparsely constructed 3D points only. However, since CPW solely relies on the tracked feature points to guide the warping, it works poorly in large texture less regions, such as ground and building interiors. To overcome this limitation, in this paper we present a hybrid approach for novel view synthesis, observing that the texture less regions often correspond to large planar surfaces in the scene. Particularly, given a jittery video, we first segment each frame into piecewise planar regions as well as regions labeled as non-planar using Markov random fields. Then, a new warp is computed by estimating a single homography for regions belong to the same plane, while inheriting results from CPW in the non-planar regions. We demonstrate how the segmentation information can be efficiently obtained and seamlessly integrated into the stabilization framework. Experimental results on a variety of real video sequences verify the effectiveness of our method.
Zihan Zhou 0001, Hailin Jin, Yi Ma 0001
CVPR1
2013 Single-Sample Face Recognition with Image Corruption and Misalignment via Sparse Illumination Transfer
abstract
Single-sample face recognition is one of the most challenging problems in face recognition. We propose a novel face recognition algorithm to address this problem based on a sparse representation based classification (SRC) framework. The new algorithm is robust to image misalignment and pixel corruption, and is able to reduce required training images to one sample per class. To compensate the missing illumination information typically provided by multiple training images, a sparse illumination transfer (SIT) technique is introduced. The SIT algorithms seek additional illumination examples of face images from one or more additional subject classes, and form an illumination dictionary. By enforcing a sparse representation of the query image, the method can recover and transfer the pose and illumination information from the alignment stage to the recognition stage. Our extensive experiments have demonstrated that the new algorithms significantly outperform the existing algorithms in the single-sample regime and with less restrictions. In particular, the face alignment accuracy is comparable to that of the well-known Deformable SRC algorithm using multiple training images, and the face recognition accuracy exceeds those of the SRC and Extended SRC algorithms using hand labeled alignment initialization.
Liansheng Zhuang, Allen Y. Yang, Zihan Zhou 0001, S. Shankar Sastry, Yi Ma 0001
CVPR3
2013 Fast 퓁 1 -Minimization Algorithms for Robust Face Recognition
abstract
l1-minimization refers to finding the minimum l1-norm solution to an underdetermined linear system [Formula: see text]. Under certain conditions as described in compressive sensing theory, the minimum l1-norm solution is also the sparsest solution. In this paper, we study the speed and scalability of its algorithms. In particular, we focus on the numerical implementation of a sparsity-based classification framework in robust face recognition, where sparse representation is sought to recover human identities from high-dimensional facial images that may be corrupted by illumination, facial disguise, and pose variation. Although the underlying numerical problem is a linear program, traditional algorithms are known to suffer poor scalability for large-scale applications. We investigate a new solution based on a classical convex optimization framework, known as augmented Lagrangian methods. We conduct extensive experiments to validate and compare its performance against several popular l1-minimization solvers, including interior-point method, Homotopy, FISTA, SESOP-PCD, approximate message passing, and TFOCS. To aid peer evaluation, the code for all the algorithms has been made publicly available.
Allen Y. Yang, Zihan Zhou 0001, A. G. Balasubramanian, S. Shankar Sastry, Yi Ma 0001
IEEE Trans. Image Process.2
2012 Robust plane-based structure from motion
abstract
We introduce a new approach to structure and motion recovery directly from one or more large planes in the scene. When such a plane exists, we demonstrate how to automatically detect and track it robustly and consistently over a long video sequence, and how to efficiently self-calibrate the camera using the homographies induced by this plane. We build a complete structure from motion system which does not use any additional off-the-plane information about the scene, and show its advantage over conventional systems in handling two important issues which often occur in real world videos, namely, the plane degeneracy and the dynamic foreground problems. Experimental results on a variety of real video sequences verify the effectiveness and efficiency of our system.
Zihan Zhou 0001, Hailin Jin, Yi Ma 0001
CVPR1
2012 Toward a Practical Face Recognition System: Robust Alignment and Illumination by Sparse Representation
abstract
Many classic and contemporary face recognition algorithms work well on public data sets, but degrade sharply when they are used in a real recognition system. This is mostly due to the difficulty of simultaneously handling variations in illumination, image misalignment, and occlusion in the test image. We consider a scenario where the training images are well controlled and test images are only loosely controlled. We propose a conceptually simple face recognition system that achieves a high degree of robustness and stability to illumination variation, image misalignment, and partial occlusion. The system uses tools from sparse representation to align a test face image to a set of frontal training images. The region of attraction of our alignment algorithm is computed empirically for public face data sets such as Multi-PIE. We demonstrate how to capture a set of training images with enough illumination variation that they span test images taken under uncontrolled illumination. In order to evaluate how our algorithms work under practical testing conditions, we have implemented a complete face recognition system, including a projector-based training acquisition system. Our system can efficiently and effectively recognize faces under a variety of realistic conditions, using only frontal images under the proposed illuminations as training.
Andrew Wagner, John Wright 0001, Arvind Ganesh, Zihan Zhou 0001, Hossein Mobahi, Yi Ma 0001
IEEE Trans. Pattern Anal. Mach. Intell.4
2010 Towards a robust face recognition system using compressive sensing
abstract
An application of compressive sensing (CS) theory in imagebased robust face recognition is considered. Most contemporary face recognition systems suffer from limited abilities to handle image nuisances such as illumination, facial disguise, and pose misalignment. Motivated by CS, the problem has been recently cast in a sparse representation framework: The sparsest linear combination of a query image is sought using all prior training images as an overcomplete dictionary, and the dominant sparse coefficients reveal the identity of the query image. The ability to perform dense error correction directly in the image space also provides an intriguing solution to compensate pixel corruption and improve the recognition accuracy exceeding most existing solutions. Furthermore, a local iterative process can be applied to solve for an image transformation applied to the face region when the query image is misaligned. Finally, we discuss the state of the art in fast ℓ1-minimization to improve the speed of the robust face recognition system. The paper also provides useful guidelines to practitioners working in similar fields, such as acoustic/speech recognition. Index Terms: face recognition, compressive sensing, ℓ1minimization 1.
Allen Y. Yang, Zihan Zhou 0001, Yi Ma 0001, S. Shankar Sastry
INTERSPEECH2
2010 Stable Principal Component Pursuit
abstract
In this paper, we study the problem of recovering a low-rank matrix (the principal components) from a high-dimensional data matrix despite both small entry-wise noise and gross sparse errors. Recently, it has been shown that a convex program, named Principal Component Pursuit (PCP), can recover the low-rank matrix when the data matrix is corrupted by gross sparse errors. We further prove that the solution to a related convex program (a relaxed PCP) gives an estimate of the low-rank matrix that is simultaneously stable to small entry-wise noise and robust to gross sparse errors. More precisely, our result shows that the proposed convex program recovers the low-rank matrix even though a positive fraction of its entries are arbitrarily corrupted, with an error bound proportional to the noise level. We present simulation results to support our result and demonstrate that the new convex program accurately recovers the principal components (the low-rank matrix) under quite broad conditions. To our knowledge, this is the first result that shows the classical Principal Component Analysis (PCA), optimal for small i.i.d. noise, can be made robust to gross sparse errors; or the first that shows the newly proposed PCP can be made stable to small entry-wise perturbations.
Zihan Zhou 0001, Xiaodong Li 0005, John Wright 0001, Emmanuel J. Candès, Yi Ma 0001
ISIT1
2009 Towards a practical face recognition system: Robust registration and illumination by sparse representation
abstract
Most contemporary face recognition algorithms work well under laboratory conditions but degrade when tested in less-controlled environments. This is mostly due to the difficulty of simultaneously handling variations in illumination, alignment, pose, and occlusion. In this paper, we propose a simple and practical face recognition system that achieves a high degree of robustness and stability to all these variations. We demonstrate how to use tools from sparse representation to align a test face image with a set of frontal training images in the presence of significant registration error and occlusion. We thoroughly characterize the region of attraction for our alignment algorithm on public face datasets such as Multi-PIE. We further study how to obtain a sufficient set of training illuminations for linearly interpolating practical lighting conditions. We have implemented a complete face recognition system, including a projector-based training acquisition system, in order to evaluate how our algorithms work under practical testing conditions. We show that our system can efficiently and effectively recognize faces under a variety of realistic conditions, using only frontal images under the proposed illuminations as training.
Andrew Wagner, John Wright 0001, Arvind Ganesh, Zihan Zhou 0001, Yi Ma 0001
CVPR4
2009 Separation of a subspace-sparse signal: Algorithms and conditions
abstract
In this paper, we show how two classical sparse recovery algorithms, Orthogonal Matching Pursuit and Basis Pursuit, can be naturally extended to recover block-sparse solutions for subspace-sparse signals. A subspace-sparse signal is sparse with respect to a set of subspaces, instead of atoms. By generalizing the notion of mutual incoherence to the set of subspaces, we show that all classical sufficient conditions remain exactly the same for these algorithms to work for subspace-sparse signals, in both noiseless and noisy cases. The sufficient conditions provided are easy to verify for large systems. We conduct simulations to compare the performance of the proposed algorithms.
Arvind Ganesh, Zihan Zhou 0001, Yi Ma 0001
ICASSP2
2009 Face recognition with contiguous occlusion using markov random fields
abstract
Partially occluded faces are common in many applications of face recognition. While algorithms based on sparse representation have demonstrated promising results, they achieve their best performance on occlusions that are not spatially correlated (i.e. random pixel corruption). We show that such sparsity-based algorithms can be significantly improved by harnessing prior knowledge about the pixel error distribution. We show how a Markov Random Field model for spatial continuity of the occlusion can be integrated into the computation of a sparse representation of the test image with respect to the training images. Our algorithm efficiently and reliably identifies the corrupted regions and excludes them from the sparse representation. Extensive experiments on both laboratory and real-world datasets show that our algorithm tolerates much larger fractions and varieties of occlusion than current state-of-the-art algorithms.
Zihan Zhou 0001, Andrew Wagner, Hossein Mobahi, John Wright 0001, Yi Ma 0001
ICCV1
2008 Demo: Robust face recognition via sparse representation
abstract
This work builds on the method of [6] to create a prototype access control system, capable of handling variations in illumination and expression, as well as significant occlusion or disguise. Our demonstration will allow participants to interact with the algorithm, gaining a better understanding strengths and limitations of sparse representation as a tool for robust recognition.
John Wright 0001, Arvind Ganesh, Zihan Zhou 0001, Andrew Wagner, Yi Ma 0001
FG3
2008 Nearest-Subspace Patch Matching for face recognition under varying pose and illumination
abstract
We consider the problem of recognizing human faces despite variations in both pose and illumination, using only frontal training images. We propose a very simple algorithm, called nearest-subspace patch matching, which combines a local translational model for deformation due to pose with a linear subspace model for lighting variations. This algorithm gives surprisingly competitive performance for moderate variations in both pose and illumination, a domain that encompasses most face recognition applications, such as access control. The results also provide a baseline for justifying the use of more complicated face models or more advanced learning methods to handle more extreme situations. Extensive experiments on publicly available databases verify the efficacy of the proposed method and clarify its operating range.
Zihan Zhou 0001, Arvind Ganesh, John Wright 0001, Shen-Fu Tsai, Yi Ma 0001
FG1