VLDB 2026 Research / reviewers in the wild / expert
Sing Bing Kang
dblp:k/SingBingKang
· DBLP profile ↗
150ranked-venue papers
33as first author
11since 2021 · last 2025
0000-0003-2016-2915ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 107 · 17 first-author · 10 since 2021Artificial intelligence and machine learning · 102 · 24 first-author · 7 since 2021Systems, architecture and hardware · 5 · 5 first-authorHuman-computer interaction and ubiquitous computing · 5 · 3 first-authorApplied, interdisciplinary, general and emerging computing · 3 · 3 first-authorDatabases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | TFM2: Training-Free Mask Matching for Open-Vocabulary Semantic SegmentationabstractThe potential of Open-Vocabulary Semantic Segmentation (OVSS) in few-shot scenarios is not fully explored due to the complexity of extending few-shot concepts to seman-tic segmentation tasks. To address this challenge, we propose Training-Free Mask Matching (TFM2), an efficient, mask-based adapter method that enhances OVSS models for the few-shot open vocabulary semantic segmentation task. TFM2 is a key-value cache that explicitly designed for image masks. We introduce three modules to construct and refine the mask cache, subsequently enhancing the OVSS mask classification performance. Comprehensive experiments demonstrate that TFM2 improves the performance of state-of-the-art OVSS methods by a margin of 1% to 5% across different settings. Moreover, TFM2 is not limited to any specific methods or backbones. This work underscores the importance and potential of few-shot data in OVSS and presents a significant step toward leveraging this potential Yaoxin Zhuo, Zachary Bessinger, Lichen Wang, Naji Khosravan, Baoxin Li, Sing Bing Kang |
WACV | 6 |
| 2024 | iBARLE: imBalance-Aware Room Layout EstimationabstractRoom layout estimation predicts layouts from a single panorama. It requires datasets with large-scale and diverse room shapes to well train the models. However, there are significant imbalances in real-world datasets including the dimensions of layout complexity, camera locations, and variation in scene appearance. These issues considerably influence the model training performance. In this work, we propose imBalance-Aware Room Layout Estimation (iBARLE) framework to address these issues. iBARLE consists of: (1) Appearance Variation Generation (AVG) module, which promotes visual appearance domain generalization, (2) Complex Structure Mix-up (CSMix) module, which enhances generalizability w.r.t. room structure, and (3) a gradient-based layout objective function, which allows more effective accounting for occlusions in complex layouts. All modules are jointly trained and help each other to achieve the best performance. Experiments and ablation studies based on ZInD [6] dataset illustrate that iBARLE has state-of-the-art performance compared with other layout estimation baselines. Taotao Jing, Lichen Wang, Naji Khosravan, Zhiqiang Wan, Zachary Bessinger, Zhengming Ding, Sing Bing Kang |
WACV | 7 |
| 2024 | Polarimetric Helmholtz StereopsisabstractHelmholtz stereopsis (HS) exploits the reciprocity principle of light propagation (i.e., the Helmholtz reciprocity) for 3D reconstruction of surfaces with arbitrary reflectance. In this paper, we present the polarimetric Helmholtz stereopsis (polar-HS), which extends the classical HS by considering the polarization state of light in the reciprocal paths. With the additional phase information from polarization, polar-HS requires only one reciprocal image pair. We derive the reciprocity relationship of Mueller matrix and formulate new reciprocity constraint that takes polarization state into account. We also utilize polarimetric constraints and extend them to the case of perspective projection. For the recovery of surface depths and normals, we incorporate reciprocity constraint with diffuse/specular polarimetric constraints in a unified optimization framework. For depth estimation, we further propose to utilize the consistency of diffuse angle of polarization. For normal estimation, we develop a normal refinement strategy based on degree of linear polarization. Using a hardware prototype, we show that our approach produces high-quality 3D reconstruction for different types of surfaces, ranging from diffuse to highly specular. Yuqi Ding, Yu Ji 0001, Mingyuan Zhou, Sing Bing Kang, Jinwei Ye |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2022 | PSMNet: Position-aware Stereo Merging Network for Room Layout EstimationabstractIn this paper, we propose a new deep learning-based method for estimating room layout given a pair of 360° panoramas. Our system, called Position-aware Stereo Merging Network or PSMNet, is an end-to-end joint layout-pose estimator. PSMNet consists of a Stereo Pano Pose (SP2) transformer and a novel Cross-Perspective Projection (CP2) layer. The stereo-view SP2 transformer is used to implicitly infer correspondences between views, and can handle noisy poses. The pose-aware CP2layer is designed to render features from the adjacent view to the anchor (reference) view, in order to perform view fusion and estimate the visible layout. Our experiments and analysis validate our method, which significantly outperforms the state-of-the-art layout estimators, especially for large and complex room spaces. Haiyan Wang 0019, Will Hutchcroft, Yuguang Li, Zhiqiang Wan, Ivaylo Boyadzhiev, Yingli Tian, Sing Bing Kang |
CVPR | 7 |
| 2022 | LASER: LAtent SpacE Rendering for 2D Visual LocalizationabstractWe present LASER, an image-based Monte Carlo Localization (MCL) framework for 2D floor maps. LASER introduces the concept of latent space rendering, where 2D pose hypotheses on the floor map are directly rendered into a geometrically-structured latent space by aggregating viewing ray features. Through a tightly coupled rendering codebook scheme, the viewing ray features are dynamically determined at rendering-time based on their geometries (i.e. length, incident-angle), endowing our representation with view-dependent fine-grain variability. Our codebook scheme effectively disentangles feature encoding from rendering, allowing the latent space rendering to run at speeds above 10KHz. Moreover, through metric learning, our geometrically-structured latent space is common to both pose hypotheses and query images with arbitrary field of views. As a result, LASER achieves state-of-the-art performance on large-scale indoor localization datasets (i. e. ZInD [5] and Structured3D [38]) for both panorama and perspective image queries, while significantly outperforming existing learning-based methods in speed. Zhixiang Min, Naji Khosravan, Zachary Bessinger, Manjunath Narayana, Sing Bing Kang, Enrique Dunn, Ivaylo Boyadzhiev |
CVPR | 5 |
| 2022 | CoVisPose: Co-visibility Pose Transformer for Wide-Baseline Relative Pose Estimation in 360$^\circ $ Indoor Panoramas
Will Hutchcroft, Yuguang Li, Ivaylo Boyadzhiev, Zhiqiang Wan, Haiyan Wang 0019, Sing Bing Kang |
ECCV (32) | 6 |
| 2022 | SALVe: Semantic Alignment Verification for Floorplan Reconstruction from Sparse Panoramas
John Lambert, Yuguang Li, Ivaylo Boyadzhiev, Lambert Wixson, Manjunath Narayana, Will Hutchcroft, James Hays, Frank Dellaert, Sing Bing Kang |
ECCV (31) | 9 |
| 2022 | Generating Topological Structure of Floorplans from Room AttributesabstractAnalysis of indoor spaces requires topological information. In this paper, we propose to extract topological information from room attributes using what we call Iterative and adaptive graph Topology Learning (ITL). ITL progressively predicts multiple relations between rooms; at each iteration, it improves node embeddings, which in turn facilitates the generation of a better topological graph structure. This notion of iterative improvement of node embeddings and topological graph structure is in the same spirit as [5]. However, while [5] computes the adjacency matrix based on node similarity, we learn the graph metric using a relational decoder to extract room correlations. Experiments using a new challenging indoor dataset validate our proposed method. Qualitative and quantitative evaluation for layout topology prediction and floorplan generation applications also demonstrate the effectiveness of ITL. Yu Yin 0001, Will Hutchcroft, Naji Khosravan, Ivaylo Boyadzhiev, Yun Fu 0001, Sing Bing Kang |
ICMR | 6 |
| 2022 | Semantically supervised appearance decomposition for virtual staging from a single panoramaabstractWe describe a novel approach to decompose a single panorama of an empty indoor environment into four appearance components: specular, direct sunlight, diffuse and diffuse ambient without direct sunlight. Our system is weakly supervised by automatically generated semantic maps (with floor, wall, ceiling, lamp, window and door labels) that have shown success on perspective views and are trained for panoramas using transfer learning without any further annotations. A GAN-based approach supervised by coarse information obtained from the semantic map extracts specular reflection and direct sunlight regions on the floor and walls. These lighting effects are removed via a similar GAN-based approach and a semantic-aware inpainting step. The appearance decomposition enables multiple applications including sun direction estimation, virtual furniture insertion, floor material replacement, and sun direction change, providing an effective tool for virtual home staging. We demonstrate the effectiveness of our approach on a large and recently released dataset of panoramas of empty homes. Tiancheng Zhi, Bowei Chen 0004, Ivaylo Boyadzhiev, Sing Bing Kang, Martial Hebert, Srinivasa G. Narasimhan |
ACM Trans. Graph. | 4 |
| 2021 | Zillow Indoor Dataset: Annotated Floor Plans With 360deg Panoramas and 3D Room LayoutsabstractWe present Zillow Indoor Dataset (ZInD): A large indoor dataset with 71,474 panoramas from 1,524 real unfurnished homes. ZInD provides annotations of 3D room layouts, 2D and 3D floor plans, panorama location in the floor plan, and locations of windows and doors. The ground truth construction took over 1,500 hours of annotation work. To the best of our knowledge, ZInD is the largest real dataset with layout annotations. A unique property is the room layout data, which follows a real world distribution (cuboid, more general Manhattan, and non-Manhattan layouts) as opposed to the mostly cuboid or Manhattan layouts in current publicly available datasets. Also, the scale and annotations provided are valuable for effective research related to room layout and floor plan analysis. To demonstrate ZInD’s benefits, we benchmark on room layout estimation from single panoramas and multi-view registration. Steve Cruz, Will Hutchcroft, Yuguang Li, Naji Khosravan, Ivaylo Boyadzhiev, Sing Bing Kang |
CVPR | 6 |
| 2021 | Polarimetric Helmholtz StereopsisabstractHelmholtz stereopsis (HS) exploits the reciprocity principle of light propagation (i.e., the Helmholtz reciprocity) for 3D reconstruction of surfaces with arbitrary reflectance. In this paper, we present the polarimetric Helmholtz stereopsis (polar-HS), which extends the classical HS by considering the polarization state of light in the reciprocal paths. With the additional phase information from polarization, polar-HS requires only one reciprocal image pair. We formulate new reciprocity and diffuse/specular polarimetric constraints to recover surface depths and normals using an optimization framework. Using a hardware prototype, we show that our approach produces high-quality 3D reconstruction for different types of surfaces, ranging from diffuse to highly specular. Yuqi Ding, Yu Ji 0001, Mingyuan Zhou, Sing Bing Kang, Jinwei Ye |
ICCV | 4 |
| 2020 | 3D Face Reconstruction using Color Photometric Stereo with Uncalibrated Near Point LightsabstractWe present a new color photometric stereo (CPS) method that recovers high quality, detailed 3D face geometry in a single shot. Our system uses three uncalibrated near point lights of different colors and a single camera. For robust self-calibration of the light sources, we use 3D morphable model (3DMM) [1] and semantic segmentation of facial parts. For reconstruction, we address the inherent spectral ambiguity in color photometric stereo by incorporating albedo consensus, albedo similarity, and proxy prior into a unified framework. In this way, we jointly exploit multiple cues to resolve under-determinedness, without the need for spatial constancy of albedo. Experiments show that our new approach produces state-of-the-art results from single image with high-fidelity geometry that includes details such as wrinkles. Yu Ji 0001, Mingyuan Zhou, Sing Bing Kang, Jingyi Yu 0001 |
ICCP | 4 |
| 2019 | Revealing Scenes by Inverting Structure From Motion ReconstructionsabstractMany 3D vision systems localize cameras within a scene using 3D point clouds. Such point clouds are often obtained using structure from motion (SfM), after which the images are discarded to preserve privacy. In this paper, we show, for the first time, that such point clouds retain enough information to reveal scene appearance and compromise privacy. We present a privacy attack that reconstructs color images of the scene from the point cloud. Our method is based on a cascaded U-Net that takes as input, a 2D multichannel image of the points rendered from a specific viewpoint containing point depth and optionally color and SIFT descriptors and outputs a color image of the scene from that viewpoint. Unlike previous feature inversion methods, we deal with highly sparse and irregular 2D point distributions and inputs where many point attributes are missing, namely keypoint orientation and scale, the descriptor image source and the 3D point visibility. We evaluate our attack algorithm on public datasets and analyze the significance of the point cloud attributes. Finally, we show that novel views can also be generated thereby enabling compelling virtual tours of the underlying scene. Francesco Pittaluga, Sanjeev J. Koppal, Sing Bing Kang, Sudipta N. Sinha |
CVPR | 3 |
| 2019 | Privacy Preserving Image-Based LocalizationabstractImage-based localization is a core component of many augmented/mixed reality (AR/MR) and autonomous robotic systems. Current localization systems rely on the persistent storage of 3D point clouds of the scene to enable camera pose estimation, but such data reveals potentially sensitive scene information. This gives rise to significant privacy risks, especially as for many applications 3D mapping is a background process that the user might not be fully aware of. We pose the following question: How can we avoid disclosing confidential information about the captured 3D scene, and yet allow reliable camera pose estimation? This paper proposes the first solution to what we call privacy preserving image-based localization. The key idea of our approach is to lift the map representation from a 3D point cloud to a 3D line cloud. This novel representation obfuscates the underlying scene geometry while providing sufficient geometric constraints to enable robust and accurate 6-DOF camera pose estimation. Extensive experiments on several datasets and localization scenarios underline the high practical relevance of our proposed approach. Pablo Speciale, Johannes L. Schönberger, Sing Bing Kang, Sudipta N. Sinha, Marc Pollefeys |
CVPR | 3 |
| 2019 | Content Aware Image Pre-CompensationabstractThe goal of image pre-compensation is to process an image such that after being convolved with a known kernel, will appear close to the sharp reference image. In a practical setting, the pre-compensated image has significantly higher dynamic range than the latent image. As a result, some form of tone mapping is needed. In this paper, we show how global tone mapping functions affect contrast and ringing in image pre-compensation. We further enhance contrast and reduce ringing by considering the visual saliency. Specifically, we prioritize contrast preservation in salient regions while tolerating more blurriness elsewhere. For quantitative analysis, we design new metrics to measure the contrast of an image with ringing. Specifically, we set out to find its "equivalent ringing-free" image that matches its intensity histogram and uses its contrast as the measure. We illustrate our approach on projector defocus compensation and visual acuity enhancement. Compared with the state-of-the-art, our approach significantly improves the contrast. We also perform user studies to demonstrate that our method can effectively improve the viewing experience for users with impaired vision. Jinwei Ye, Yu Ji 0001, Mingyuan Zhou, Sing Bing Kang, Jingyi Yu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2019 | Hyperspectral Light Field Stereo MatchingabstractIn this paper, we describe how scene depth can be extracted using a hyperspectral light field capture (H-LF) system. Our H-LF system consists of a 5 ×6 array of cameras, with each camera sampling a different narrow band in the visible spectrum. There are two parts to extracting scene depth. The first part is our novel cross-spectral pairwise matching technique, which involves a new spectral-invariant feature descriptor and its companion matching metric we call bidirectional weighted normalized cross correlation (BWNCC). The second part, namely, H-LF stereo matching, uses a combination of spectral-dependent correspondence and defocus cues. These two new cost terms are integrated into a Markov Random Field (MRF) for disparity estimation. Experiments on synthetic and real H-LF data show that our approach can produce high-quality disparity maps. We also show that these results can be used to produce the complete plenoptic cube in addition to synthesizing all-focus and defocused color images under different sensor spectral responses. Kang Zhu, Yujia Xue, Qiang Fu 0002, Sing Bing Kang, Xilin Chen 0001, Jingyi Yu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2019 | Personalized Exposure Control Using Adaptive Metering and Reinforcement LearningabstractWe propose a reinforcement learning approach for real-time exposure control of a mobile camera that is personalizable. Our approach is based on Markov Decision Process (MDP). In the camera viewfinder or live preview mode, given the current frame, our system predicts the change in exposure so as to optimize the trade-off among image quality, fast convergence, and minimal temporal oscillation. We model the exposure prediction function as a fully convolutional neural network that can be trained through Gaussian policy gradient in an end-to-end fashion. As a result, our system can associate scene semantics with exposure values; it can also be extended to personalize the exposure adjustments for a user and device. We improve the learning performance by incorporating an adaptive metering module that links semantics with exposure. This adaptive metering module generalizes the conventional spot or matrix metering techniques. We validate our system using the MIT FiveK [1] and our own datasets captured using iPhone 7 and Google Pixel. Experimental results show that our system exhibits stable real-time behavior while improving visual quality compared to what is achieved through native camera control. Huan Yang 0005, Baoyuan Wang, Noranart Vesdapunt, Minyi Guo, Sing Bing Kang |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2018 | Automatic 3D Indoor Scene Modeling From Single PanoramaabstractWe describe a system that automatically extracts 3D geometry of an indoor scene from a single 2D panorama. Our system recovers the spatial layout by finding the floor, walls, and ceiling; it also recovers shapes of typical indoor objects such as furniture. Using sampled perspective sub-views, we extract geometric cues (lines, vanishing points, orientation map, and surface normals) and semantic cues (saliency and object detection information). These cues are used for ground plane estimation and occlusion reasoning. The global spatial layout is inferred through a constraint graph on line segments and planar superpixels. The recovered layout is then used to guide shape estimation of the remaining objects using their normal information. Experiments on synthetic and real datasets show that our approach is state-of-the-art in both accuracy and efficiency. Our system can handle cluttered scenes with complex geometry that are challenging to existing techniques. Ruiyang Liu, Sing Bing Kang, Jingyi Yu 0001 |
CVPR | 4 |
| 2018 | Semantic-Driven Generation of Hyperlapse from 360 Degree VideoabstractWe present a system for converting a fully panoramic (360 degree) video into a normal field-of-view (NFOV) hyperlapse for an optimal viewing experience. Our system exploits visual saliency and semantics to non-uniformly sample in space and time for generating hyperlapses. In addition, users can optionally choose objects of interest for customizing the hyperlapses. We first stabilize an input 360 degree video by smoothing the rotation between adjacent frames and then compute regions of interest and saliency scores. An initial hyperlapse is generated by optimizing the saliency and motion smoothness followed by the saliency-aware frame selection. We further smooth the result using an efficient 2D video stabilization approach that adaptively selects the motion model to generate the final hyperlapse. We validate the design of our system by showing results for a variety of scenes and comparing against the state-of-the-art method through a large-scale user study. Wei-Sheng Lai, Yujia Huang, Neel Joshi, Chris Buehler, Ming-Hsuan Yang 0001, Sing Bing Kang |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2017 | Personalized Cinemagraphs Using Semantic Understanding and Collaborative LearningabstractCinemagraphs are a compelling way to convey dynamic aspects of a scene. In these media, dynamic and still elements are juxtaposed to create an artistic and narrative experience. Creating a high-quality, aesthetically pleasing cinemagraph requires isolating objects in a semantically meaningful way and then selecting good start times and looping periods for those objects to minimize visual artifacts (such a tearing). To achieve this, we present a new technique that uses object recognition and semantic segmentation as part of an optimization method to automatically create cinemagraphs from videos that are both visually appealing and semantically meaningful. Given a scene with multiple objects, there are many cinemagraphs one could create. Our method evaluates these multiple candidates and presents the best one, as determined by a model trained to predict human preferences in a collaborative way. We demonstrate the effectiveness of our approach with multiple results and a user study. Tae-Hyun Oh, Kyungdon Joo, Neel Joshi, Baoyuan Wang, In-So Kweon, Sing Bing Kang |
ICCV | 6 |
| 2017 | Visual attribute transfer through deep image analogyabstractWe propose a new technique for visual attribute transfer across images that may have very different appearance but have perceptually similar semantic structure. By visual attribute transfer, we mean transfer of visual information (such as color, tone, texture, and style) from one image to another. For example, one image could be that of a painting or a sketch while the other is a photo of a real scene, and both depict the same type of scene. Our technique finds semantically-meaningful dense correspondences between two input images. To accomplish this, it adapts the notion of "image analogy" [Hertzmann et al. 2001] with features extracted from a Deep Convolutional Neutral Network for matching; we call our technique deep image analogy. A coarse-to-fine strategy is used to compute the nearest-neighbor field for generating the results. We validate the effectiveness of our proposed method in a variety of cases, including style/texture transfer, color/style swap, sketch/painting to photo, and time lapse. Jing Liao 0001, Lu Yuan 0001, Gang Hua 0001, Sing Bing Kang |
ACM Trans. Graph. | 5 |
| 2016 | Holoportation: Virtual 3D Teleportation in Real-timeabstractWe present an end-to-end system for augmented and virtual reality telepresence, called Holoportation. Our system demonstrates high-quality, real-time 3D reconstructions of an entire space, including people, furniture and objects, using a set of new depth cameras. These 3D models can also be transmitted in real-time to remote users. This allows users wearing virtual or augmented reality displays to see, hear and interact with remote participants in 3D, almost as if they were present in the same physical space. From an audio-visual perspective, communicating and interacting with remote users edges closer to face-to-face communication. This paper describes the Holoportation technical system in full, its key interactive capabilities, the application scenarios it enables, and an initial qualitative study of using this new communication medium. Sergio Orts, Christoph Rhemann, Sean Ryan Fanello, Wayne Chang, Adarsh Kowdle, Yury Degtyarev, David Kim 0002, Philip Davidson, Sameh Khamis, Mingsong Dou, Vladimir Tankovich, Charles T. Loop, Qin Cai, Philip A. Chou, Sarah Mennicken, Julien P. C. Valentin, Vivek Pradeep, Shenlong Wang, Sing Bing Kang, Pushmeet Kohli, Yuliya Lutchyn, Cem Keskin, Shahram Izadi |
UIST | 19 |
| 2016 | Temporally coherent completion of dynamic videoabstractWe present an automatic video completion algorithm that synthesizes missing regions in videos in a temporally coherent fashion. Our algorithm can handle dynamic scenes captured using a moving camera. State-of-the-art approaches have difficulties handling such videos because viewpoint changes cause image-space motion vectors in the missing and known regions to be inconsistent. We address this problem by jointly estimating optical flow and color in the missing regions. Using pixel-wise forward/backward flow fields enables us to synthesize temporally coherent colors. We formulate the problem as a non-parametric patch-based optimization. We demonstrate our technique on numerous challenging videos. Jia-Bin Huang 0001, Sing Bing Kang, Narendra Ahuja, Johannes Kopf 0001 |
ACM Trans. Graph. | 2 |
| 2016 | Enhancing Light Fields through Ray-Space StitchingabstractLight fields (LFs) have been shown to enable photorealistic visualization of complex scenes. In practice, however, an LF tends to have a relatively small angular range or spatial resolution, which limits the scope of virtual navigation. In this paper, we show how seamless virtual navigation can be enhanced by stitching multiple LFs. Our technique consists of two key components: LF registration and LF stitching. To register LFs, we use what we call the ray-space motion matrix (RSMM) to establish pairwise ray-ray correspondences. Using Plücker coordinates, we show that the RSMM is a 5 ×6 matrix, which reduces to a 5 ×5 matrix under pure translation and/or in-plane rotation. The final LF stitching is done using multi-resolution, high-dimensional graph-cut in order to account for possible scene motion, imperfect RSMM estimation, and/or undersampling. We show how our technique allows us to create LFs with various enhanced features: extended horizontal and/or vertical field-of-view, larger synthetic aperture and defocus blur, and larger parallax. Xinqing Guo, Sing Bing Kang, Haiting Lin, Jingyi Yu 0001 |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2015 | A light transport model for mitigating multipath interference in Time-of-flight sensorsabstractContinuous-wave Time-of-flight (TOF) range imaging has become a commercially viable technology with many applications in computer vision and graphics. However, the depth images obtained from TOF cameras contain scene dependent errors due to multipath interference (MPI). Specifically, MPI occurs when multiple optical reflections return to a single spatial location on the imaging sensor. Many prior approaches to rectifying MPI rely on sparsity in optical reflections, which is an extreme simplification. In this paper, we correct MPI by combining the standard measurements from a TOF camera with information from direct and global light transport. We report results on both simulated experiments and physical experiments (using the Kinect sensor). Our results, evaluated against ground truth, demonstrate a quantitative improvement in depth accuracy. Nikhil Naik 0003, Achuta Kadambi, Christoph Rhemann, Shahram Izadi, Ramesh Raskar, Sing Bing Kang |
CVPR | 6 |
| 2015 | Ambient occlusion via compressive visibility estimationabstractThere has been emerging interest on recovering traditionally challenging intrinsic scene properties. In this paper, we present a novel computational imaging solution for recovering the ambient occlusion (AO) map of an object. AO measures how much light from all different directions can reach a surface point without being blocked by self-occlusions. Previous approaches either require obtaining highly accurate surface geometry or acquiring a large number of images. We adopt a compressive sensing framework that captures the object under strategically coded lighting directions. We show that this incident illumination field exhibits some unique properties suitable for AO recovery: every ray's contribution to the visibility function is binary while their distribution for AO measurement is sparse. This enables a sparsity-prior based solution for iteratively recovering the surface normal, the surface albedo, and the visibility function from a small number of images. To physically implement the scheme, we construct an encodable directional light source using the light field probe. Experiments on synthetic and real scenes show that our approach is both reliable and accurate with significantly reduced size of input. Wei Yang 0034, Yu Ji 0001, Haiting Lin, Sing Bing Kang, Jingyi Yu 0001 |
CVPR | 5 |
| 2015 | Depth Recovery from Light Field Using Focal Stack SymmetryabstractWe describe a technique to recover depth from a light field (LF) using two proposed features of the LF focal stack. One feature is the property that non-occluding pixels exhibit symmetry along the focal depth dimension centered at the in-focus slice. The other is a data consistency measure based on analysis-by-synthesis, i.e., the difference between the synthesized focal stack given the hypothesized depth map and that from the LF. These terms are used in an iterative optimization framework to extract scene depth. Experimental results on real Lytro and Raytrix data demonstrate that our technique outperforms state-of-the-art solutions and is significantly more robust to noise and under-sampling. Haiting Lin, Can Chen 0004, Sing Bing Kang, Jingyi Yu 0001 |
ICCV | 3 |
| 2015 | Resolving Scale Ambiguity via XSlit Aspect Ratio AnalysisabstractIn perspective cameras, images of a frontal-parallel 3D object preserve its aspect ratio invariant to its depth. Such an invariance is useful in photography but is unique to perspective projection. In this paper, we show that alternative non-perspective cameras such as the crossed-slit or XSlit cameras exhibit a different depth-dependent aspect ratio (DDAR) property that can be used to 3D recovery. We first conduct a comprehensive analysis to characterize DDAR, infer object depth from its AR, and model recoverable depth range, sensitivity, and error. We show that repeated shape patterns in real Manhattan World scenes can be used for 3D reconstruction using a single XSlit image. We also extend our analysis to model slopes of lines. Specifically, parallel 3D lines exhibit depth-dependent slopes (DDS) on their images which can also be used to infer their depths. We validate our analyses using real XSlit cameras, XSlit panoramas, and catadioptric mirrors. Experiments show that DDAR and DDS provide important depth cues and enable effective single-image scene reconstruction. Wei Yang 0034, Haiting Lin, Sing Bing Kang, Jingyi Yu 0001 |
ICCV | 3 |
| 2015 | Change-Based Image Cropping with Exclusion and Compositional Features
Jianzhou Yan, Stephen Lin 0001, Sing Bing Kang, Xiaoou Tang |
Int. J. Comput. Vis. | 3 |
| 2015 | Fast Edge-Aware Denoising by Approximated Patch Geodesic PathsabstractPatch-based denoising, while effective, requires expensive pairwise patch comparisons. We present a novel fast patch-based denoising technique based on patch geodesic paths (PatchGPs). PatchGPs treat image patches as nodes and patch differences as edge weights for computing the shortest (geodesic) paths. The distance defined by the PatchGP can then be used as a similarity metric for image denoising. We first show that, for natural images, PatchGPs can be approximated by minimum hop paths (MHPs) that correspond to Euclidean line paths connecting two patch nodes. The denoising kernel is constructed using patches along discretized MHP search directions. We apply a weight propagation scheme to robustly and efficiently compute the path distance for each MHP. Our technique handles noise at multiple scales by analyzing the noise distribution (through wavelet decomposition) at each scale. Experiments show that our approach maintains the high quality of patch-based denoising but is a few orders of magnitude faster. We also demonstrate how PatchGP can be used for fast Bayer pattern (raw) denoising and image detail enhancement. Sing Bing Kang, Jie Yang 0002, Jingyi Yu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2014 | Light Field Stereo Matching Using Bilateral Statistics of Surface CamerasabstractIn this paper, we introduce a bilateral consistency metric on the surface camera (SCam) [26] for light field stereo matching to handle significant occlusions. The concept of SCam is used to model angular radiance distribution with respect to a 3D point. Our bilateral consistency metric is used to indicate the probability of occlusions by analyzing the SCams. We further show how to distinguish between on-surface and free space, textured and non-textured, and Lambertian and specular through bilateral SCam analysis. To speed up the matching process, we apply the edge preserving guided filter [14] on the consistency-disparity curves. Experimental results show that our technique outperforms both the state-of-the-art and the recent light field stereo matching methods, especially near occlusion boundaries. Can Chen 0004, Haiting Lin, Sing Bing Kang, Jingyi Yu 0001 |
CVPR | 4 |
| 2014 | Image Pre-compensation: Balancing Contrast and RingingabstractThe goal of image pre-compensation is to process an image such that after being convolved with a known kernel, will appear close to the sharp reference image. In a practical setting, the pre-compensated image has significantly higher dynamic range than the latent image. As a result, some form of tone mapping is needed. In this paper, we show how global tone mapping functions affect contrast and ringing in image pre-compensation. In particular, we show that linear tone mapping eliminates ringing but incurs severe contrast loss, while non-linear tone mapping functions such as Gamma curves slightly enhances contrast but introduces ringing. To enable quantitative analysis, we design new metrics to measure the contrast of an image with ringing. Specifically, we set out to find its "equivalent ringing-free" image that matches its intensity histogram and uses its contrast as the measure. We illustrate our approach on projector defocus compensation and visual acuity enhancement. Compared with the state-of-the-art, our approach significantly improves the contrast. We believe our technique is the first to analytically trade-off between contrast and ringing. Yu Ji 0001, Jinwei Ye, Sing Bing Kang, Jingyi Yu 0001 |
CVPR | 3 |
| 2014 | A Learning-to-Rank Approach for Image Color EnhancementabstractWe present a machine-learned ranking approach for automatically enhancing the color of a photograph. Unlike previous techniques that train on pairs of images before and after adjustment by a human user, our method takes into account the intermediate steps taken in the enhancement process, which provide detailed information on the person's color preferences. To make use of this data, we formulate the color enhancement task as a learning-to-rank problem in which ordered pairs of images are used for training, and then various color enhancements of a novel input image can be evaluated from their corresponding rank values. From the parallels between the decision tree structures we use for ranking and the decisions made by a human during the editing process, we posit that breaking a full enhancement sequence into individual steps can facilitate training. Our experiments show that this approach compares well to existing methods for automatic color enhancement. Jianzhou Yan, Stephen Lin 0001, Sing Bing Kang, Xiaoou Tang |
CVPR | 3 |
| 2014 | Time-Mapping Using Space-Time SaliencyabstractWe describe a new approach for generating regular-speed, low-frame-rate (LFR) video from a high-frame-rate (HFR) input while preserving the important moments in the original. We call this time-mapping, a time-based analogy to high dynamic range to low dynamic range spatial tone-mapping. Our approach makes these contributions: (1) a robust space-time saliency method for evaluating visual importance, (2) a re-timing technique to temporally resample based on frame importance, and (3) temporal filters to enhance the rendering of salient motion. Results of our space-time saliency method on a benchmark dataset show it is state-of-the-art. In addition, the benefits of our approach to HFR-to-LFR time-mapping over more direct methods are demonstrated in a user study. Sing Bing Kang, Michael F. Cohen |
CVPR | 2 |
| 2014 | Collaborative Personalization of Image Enhancement
Ashish Kapoor, Juan C. Caicedo, Dani Lischinski, Sing Bing Kang |
Int. J. Comput. Vis. | 4 |
| 2014 | Depth Transfer: Depth Extraction from Video Using Non-Parametric SamplingabstractWe describe a technique that automatically generates plausible depth maps from videos using non-parametric depth sampling. We demonstrate our technique in cases where past methods fail (non-translating cameras and dynamic scenes). Our technique is applicable to single images as well as videos. For videos, we use local motion cues to improve the inferred depth maps, while optical flow is used to ensure temporal depth consistency. For training and evaluation, we use a Kinect-based system to collect a large data set containing stereoscopic videos with known depths. We show that our depth estimation technique outperforms the state-of-the-art on benchmark databases. Our technique can be used to automatically convert a monoscopic video into stereo for 3D visualization, and we demonstrate this through a variety of visually pleasing results for indoor and outdoor scenes, including results from the feature film Charade. Kevin Karsch, Ce Liu 0001, Sing Bing Kang |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2014 | Image completion using planar structure guidanceabstractWe propose a method for automatically guiding patch-based image completion using mid-level structural cues. Our method first estimates planar projection parameters, softly segments the known region into planes, and discovers translational regularity within these planes. This information is then converted into soft constraints for the low-level completion algorithm by defining prior probabilities for patch offsets and transformations. Our method handles multiple planes, and in the absence of any detected planes falls back to a baseline fronto-parallel image completion algorithm. We validate our technique through extensive comparisons with state-of-the-art algorithms on a variety of scenes. Jia-Bin Huang 0001, Sing Bing Kang, Narendra Ahuja, Johannes Kopf 0001 |
ACM Trans. Graph. | 2 |
| 2014 | Learning to be a depth camera for close-range human capture and interactionabstractWe present a machine learning technique for estimating absolute, per-pixel depth using any conventional monocular 2D camera, with minor hardware modifications. Our approach targets close-range human capture and interaction where dense 3D estimation of hands and faces is desired. We use hybrid classification-regression forests to learn how to map from near infrared intensity images to absolute , metric depth in real-time. We demonstrate a variety of human-computer interaction and capture scenarios. Experiments show an accuracy that outperforms a conventional light fall-off baseline, and is comparable to high-quality consumer depth cameras, but with a dramatically reduced cost, power consumption, and form-factor. Sean Ryan Fanello, Cem Keskin, Shahram Izadi, Pushmeet Kohli, David Kim 0002, David Sweeney, Antonio Criminisi, Jamie Shotton, Sing Bing Kang, Tim Paek |
ACM Trans. Graph. | 9 |
| 2013 | Fast Patch-Based Denoising Using Approximated Patch Geodesic PathsabstractPatch-based methods such as Non-Local Means (NLM) and BM3D have become the de facto gold standard for image denoising. The core of these approaches is to use similar patches within the image as cues for denoising. The operation usually requires expensive pair-wise patch comparisons. In this paper, we present a novel fast patch-based denoising technique based on Patch Geodesic Paths (PatchGP). PatchGPs treat image patches as nodes and patch differences as edge weights for computing the shortest (geodesic) paths. The path lengths can then be used as weights of the smoothing/denoising kernel. We first show that, for natural images, PatchGPs can be effectively approximated by minimum hop paths (MHPs) that generally correspond to Euclidean line paths connecting two patch nodes. To construct the denoising kernel, we further discretize the MHP search directions and use only patches along the search directions. Along each MHP, we apply a weight propagation scheme to robustly and efficiently compute the path distance. To handle noise at multiple scales, we conduct wavelet image decomposition and apply PatchGP scheme at each scale. Comprehensive experiments show that our approach achieves comparable quality as the state-of-the-art methods such as NLM and BM3D but is a few orders of magnitude faster. Sing Bing Kang, Jie Yang 0002, Jingyi Yu 0001 |
CVPR | 2 |
| 2013 | Recovering Stereo Pairs from AnaglyphsabstractAn anaglyph is a single image created by selecting complementary colors from a stereo color pair, the user can perceive depth by viewing it through color-filtered glasses. We propose a technique to reconstruct the original color stereo pair given such an anaglyph. We modified SIFT-Flow and use it to initially match the different color channels across the two views. Our technique then iteratively refines the matches, selects the good matches (which defines the "anchor" colors), and propagates the anchor colors. We use a diffusion-based technique for the color propagation, and added a step to suppress unwanted colors. Results on a variety of inputs demonstrate the robustness of our technique. We also extended our method to anaglyph videos by using optic flow between time frames. Armand Joulin, Sing Bing Kang |
CVPR | 2 |
| 2013 | Learning the Change for Automatic Image CroppingabstractImage cropping is a common operation used to improve the visual quality of photographs. In this paper, we present an automatic cropping technique that accounts for the two primary considerations of people when they crop: removal of distracting content, and enhancement of overall composition. Our approach utilizes a large training set consisting of photos before and after cropping by expert photographers to learn how to evaluate these two factors in a crop. In contrast to the many methods that exist for general assessment of image quality, ours specifically examines differences between the original and cropped photo in solving for the crop parameters. To this end, several novel image features are proposed to model the changes in image content and composition when a crop is applied. Our experiments demonstrate improvements of our method over recent cropping algorithms on a broad range of images. Jianzhou Yan, Stephen Lin 0001, Sing Bing Kang, Xiaoou Tang |
CVPR | 3 |
| 2013 | Post-processing approach for radiometric self-calibration of videoabstractWe present a novel data-driven technique for radiometric self-calibration of video from an unknown camera. Our approach self-calibrates radiometric variations in video, and is applied as a post-process; there is no need to access the camera, and in particular it is applicable to internet videos. This technique builds on empirical evidence that in video the camera response function (CRF) should be regarded time variant, as it changes with scene content and exposure, instead of relying on a single camera response function. We show that a time-varying mixture of responses produces better accuracy and consistently reduces the error in mapping intensity to irradiance when compared to a single response model. Furthermore, our mixture model counteracts the effects of possible nonlinear exposure-dependent intensity perturbations and white-balance changes caused by proprietary camera firmware. We further show how radiometrically calibrated video improves the performance of other video analysis algorithms, enabling a video segmentation algorithm to be invariant to exposure and gain variations over the sequence. We validate our data-driven technique on videos from a variety of cameras and demonstrate the generality of our approach by applying it to internet video. Matthias Grundmann 0002, Chris McClanahan, Sing Bing Kang, Irfan A. Essa |
ICCP | 3 |
| 2013 | Transformation guided image completionabstractIn this paper, we describe a new interactive image completion system that allows users to easily specify various forms of mid-level structures in the image. Our system supports the specification of four basic symmetric types: reflection, translation, rotation, and glide. The user inputs are automatically converted into guidance maps that encode possible candidate shifts and, indirectly, local transformations of rotation and scale. These guidance maps are used in conjunction with a color matching cost for image completion. We show that our system is capable of handling a variety of challenging examples. Jia-Bin Huang 0001, Johannes Kopf 0001, Narendra Ahuja, Sing Bing Kang |
ICCP | 4 |
| 2013 | Single-Image Vignetting Correction from Gradient Distribution SymmetriesabstractWe present novel techniques for single-image vignetting correction based on symmetries of two forms of image gradients: semicircular tangential gradients (SCTG) and radial gradients (RG). For a given image pixel, an SCTG is an image gradient along the tangential direction of a circle centered at the presumed optical center and passing through the pixel. An RG is an image gradient along the radial direction with respect to the optical center. We observe that the symmetry properties of SCTG and RG distributions are closely related to the vignetting in the image. Based on these symmetry properties, we develop an automatic optical center estimation algorithm by minimizing the asymmetry of SCTG distributions, and also present two methods for vignetting estimation based on minimizing the asymmetry of RG distributions. In comparison to prior approaches to single-image vignetting correction, our methods do not rely on image segmentation and they produce more accurate results. Experiments show our techniques to work well for a wide range of images while achieving a speed-up of 3-5 times compared to a state-of-the-art method. Yuanjie Zheng, Stephen Lin 0001, Sing Bing Kang, Rui Xiao 0001, James C. Gee, Chandra Kambhamettu |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2012 | Context-Based Automatic Local Image Enhancement
Sung Ju Hwang, Ashish Kapoor, Sing Bing Kang |
ECCV (1) | 3 |
| 2012 | Depth Extraction from Video Using Non-parametric Sampling
Kevin Karsch, Ce Liu 0001, Sing Bing Kang |
ECCV (5) | 3 |
| 2012 | Improving sub-pixel correspondence through upsampling
Li Xu 0001, Jiaya Jia, Sing Bing Kang |
Comput. Vis. Image Underst. | 3 |
| 2012 | Image Restoration by Matching Gradient DistributionsabstractThe restoration of a blurry or noisy image is commonly performed with a MAP estimator, which maximizes a posterior probability to reconstruct a clean image from a degraded image. A MAP estimator, when used with a sparse gradient image prior, reconstructs piecewise smooth images and typically removes textures that are important for visual realism. We present an alternative deconvolution method called iterative distribution reweighting (IDR) which imposes a global constraint on gradients so that a reconstructed image should have a gradient distribution similar to a reference distribution. In natural images, a reference distribution not only varies from one image to another, but also within an image depending on texture. We estimate a reference distribution directly from an input image for each texture segment. Our algorithm is able to restore rich mid-frequency textures. A large-scale user study supports the conclusion that our algorithm improves the visual realism of reconstructed images compared to those of MAP estimators. Taeg Sang Cho, C. Lawrence Zitnick, Neel Joshi, Sing Bing Kang, Richard Szeliski, William T. Freeman |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2012 | State of the JournalabstractT year 2012 will mark the end of my term as Editor-in-Chief of the IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI). I believe that we have made substantial progress on one of the core challenges that TPAMI faces, namely, the continued growth of machine learning. As I have mentioned, the IEEE does not have a journal whose focus is modern machine learning methods, such as SVMs. Yet many papers in this area are submitted to TPAMI. One factor is that machine learning falls within TPAMI’s scope statement, but perhaps a more important reason is the journal’s excellence in computer vision, an area where machine learning is having a substantial and increasing impact. There is no possibility for TPAMI to ignore this area and continue to thrive, so the journal out of necessity must rise to the challenge of becoming a leading publication in machine learning. My predecessor David Kriegman saw this development clearly, and responded by appointing Zoubin Gharamani, a famous machine learning expert, as an Associate Editor in Chief (AEIC). Zoubin served his complete 4-year term with distinction, and has now moved up to the TPAMI Advisory Board. Over the last few years the number of machine learning submissions has continued to grow substantially, and we have clearly needed additional help. Machine learning is an area where TPAMI faces some distinct challenges. TPAMI has not published a body of truly fundamental papers in machine learning that is comparable to our accomplishments in computer vision, biometrics, or other areas that are closer to the journal’s traditional strengths. As a result, many of the submissions we receive have fallen short of TPAMI’s high standards. This has posed a diffi cult problem because it is challenging to attract top-notch researchers in machine learning as reviewers or AEs when most of the papers they handle must be rejected, and many would, in all honesty, never be submitted to a major machine learning journal. My primary focus throughout my term as EIC has been to address this situation, and I am pleased to report signifi cant progress. As you know, Max Welling joined us as an AEIC. I am happy to announce that Neil Lawrence has also agreed to serve as an AEIC. Neil has served with distinction as an AE for TPAMI, is on the board of JMLR, and will be program chair for AISTATS. (I note with amusement that Max Welling also held these three roles, which suggests a simple automatic classifi er to detect TPAMI AEICs in machine learning!) Neil has published two books in machine learning, and is primarily interested in probabilistic models. With both Max and Neil on board as AEICs we now have suffi cient manpower to address our main challenges. We have raised the bar for machine learning papers to be sent out for review by rejecting papers early on that would have eventually been rejected anyway. Hopefully, the effects of this will be clearly felt by everyone involved in the reviewing process, and the authors of high-quality submissions will benefi t from the increased availability of reviewing resources. Coupled with this effort, Max and Neil are developing a number of high quality special issues on important topics in machine learning (see the call for papers on page 207 of this issue for the fi rst such initiative). The fi eld of machine learning has a major advantage in its commitment to Open Access, which is an issue that the IEEE (along with most publishers) is struggling with. The top journal in machine learning (JMLR) is Open Access, while perhaps the best conference (NIPS) is making its proceedings available in arXiv. This has enormous benefi ts to the machine learning community. I personally believe that TPAMI will, over time, end up moving to an Open Access model, and I will hazard a guess that this will be one of the main challenges that the next EIC will face. On the operational side, the reviewing process on the whole is fairly timely, although exceptions do occur for a variety of reasons, and I want to yet again apologize to the authors whose papers get stalled in the process for one reason or another. To provide some numbers, there were 999 submissions in 2010 (I must confess I was really hoping for one more to come in at the very end). We are on track for a similar number in 2011, with 795 received as I write. The acceptance rate for 2010 submissions so far is 14 percent, though it is important to realize this does not imply 86 percent have been rejected, since a number of such papers are still undergoing revisions. The typical time from submission to final decision is about six months, which is unchanged from last year. Approximately 30 percent of submissions are rejected without review; while this is unpleasant for the authors, it saves them time from having their paper rejected at the end of the full review process and lets them quickly revise their papers for submission to a more appropriate journal. I am happy to report that the issue with the print queue is now under control, and papers now typically appear in print approximately 5.5 months after the fi nal material is uploaded. Short papers are generally published even faster, and authors are urged to consider this option. Of course, papers continue to be published online quite quickly after acceptance. On the topic of online publication, TPAMI is now available in the IEEE Computer Society’s new OnlinePlus format, at a signifi cant discount to the print subscription price. Over time the number of subscribers to the printed journal is falling, and readers who wish to see this format continue should be sure to sign up for print subscriptions. Sing Bing Kang, Jiri Matas, Max Welling, Ramin Zabih |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2012 | Editor's Note
Ramin Zabih, Sing Bing Kang, Neil D. Lawrence, Jiri Matas, Max Welling |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2012 | Editor's Note
Ramin Zabih, Sing Bing Kang, Neil D. Lawrence, Jiri Matas, Max Welling |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2012 | Quality prediction for image completionabstractWe present a data-driven method to predict the quality of an image completion method. Our method is based on the state-of-the-art non-parametric framework of Wexleret al. [2007]. It uses automatically derived search space constraints for patch source regions, which lead to improved texture synthesis and semantically more plausible results. These constraints also facilitate performance prediction by allowing us to correlate output quality against features of possible regions used for synthesis. We use our algorithm to first crop and then complete stitched panoramas. Our predictive ability is used to find an optimal crop shapebeforethe completion is computed, potentially saving significant amounts of computation. Our optimized crop includes as much of the original panorama as possible while avoiding regions that can be less successfully filled in. Our predictor can also be applied for hole filling in the interior of images. In addition to extensive comparative results, we ran several user studies validating our predictive feature, good relative quality of our results against those of other state-of-the-art algorithms, and our automatic cropping algorithm. Johannes Kopf 0001, Wolf Kienzle, Steven Mark Drucker, Sing Bing Kang |
ACM Trans. Graph. | 4 |
| 2012 | Video Snapshots: Creating High-Quality Images from Video ClipsabstractWe describe a unified framework for generating a single high-quality still image ("snapshot") from a short video clip. Our system allows the user to specify the desired operations for creating the output image, such as super resolution, noise and blur reduction, and selection of best focus. It also provides a visual summary of activity in the video by incorporating saliency-based objectives in the snapshot formation process. We show examples on a number of different video clips to illustrate the utility and flexibility of our system. Kalyan Sunkavalli, Neel Joshi, Sing Bing Kang, Michael F. Cohen, Hanspeter Pfister |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2011 | Collaborative personalization of image enhancementabstractWhile most existing enhancement tools for photographs have universal auto-enhancement functionality, recent research shows that users can have personalized preferences. In this paper, we explore whether such personalized preferences in image enhancement tend to cluster and whether users can be grouped according to such preferences. To this end, we analyze a comprehensive data set of image enhancements collected from 336 users via Amazon Mechanical Turk. We find that such clusters do exist and can be used to derive methods to learn statistical preference models from a group of users. We also present a probabilistic framework that exploits the ideas behind collaborative filtering to automatically enhance novel images for new users. Experiments show that inferring clusters in image enhancement preferences results in better prediction of image enhancement preferences and outperforms generic auto-correction tools. Juan C. Caicedo, Ashish Kapoor, Sing Bing Kang |
CVPR | 3 |
| 2011 | Noise suppression in low-light images through joint denoising and demosaicingabstractWe address the effects of noise in low-light images in this paper. Color images are usually captured by a sensor with a color filter array (CFA). This requires a demosaicing process to generate a full color image. The captured images typically have low signal-to-noise ratio, and the demosaicing step further corrupts the image, which we show to be the leading cause of visually objectionable random noise patterns (splotches). To avoid this problem, we propose a combined framework of denoising and demosaicing, where we use information about the image inferred in the denoising step to perform demosaicing. Our experiments show that such a framework results in sharper low-light images that are devoid of splotches and other noise artifacts. Priyam Chatterjee, Neel Joshi, Sing Bing Kang, Yasuyuki Matsushita |
CVPR | 3 |
| 2011 | Guest Editors' Introduction to the Special Section on Award-Winning Papers from the IEEE Conference on Computer Vision and Pattern Recognition 2009 (CVPR 2009)abstract. Kaiming He, Jian Sun, and Xiaoou Tang, “Single Image Haze Removal Using Dark Channel Prior,” winner of the Best Paper Award (sponsored by Microsoft). . Anat Levin, Yair Weiss, Fredo Durand, and Bill Freeman, “Understanding and evaluating blind deconvolution algorithms,” winner of the Best Paper-Honorable Mention Award (sponsored by Honeywell). . Ce Liu, Jenny Yuen, and Antonio Torralba, “Nonparametric Scene Parsing: Label Transfer via Dense Scene Alignment,” winner of the Best Student Paper Award (sponsored by MERL). . Olivier Duchenne, Francis Bach, In So Kweon, and Jean Ponce, “A Tensor-Based Algorithm for HighOrder Graph Matching,” winner of the Best Student Paper-Honorable Mention Award (sponsored by Hewlett-Packard). CVPR 2009 received 1,464 complete submissions by the 20 November 2008 deadline. We worked with 46 Area Chairs, who are well-respected members of the computer vision community, and 749 reviewers to select 61 papers as Orals and 322 papers as Posters. Twenty of these accepted papers were recommended to the Awards Committee for consideration. The Awards Committee consisted of five senior members of the vision community, three of whom were Area Chairs. The committee members had no conflicts with the candidate papers. The four award papers were selected after three phases of reviewing. These award papers were presented in the only singletrack session of the main conference. The authors of the award-winning papers and honorable mentions were invited to submit an extended version of their paper as a journal submission to TPAMI. The papers were reviewed by expert reviewers in the field, following the usual TPAMI procedure. All of the papers were accepted, in most cases after relatively minor revisions. CVPR 2009 also presented the Longuet-Higgins Prize for fundamental contributions in computer vision that have withstood the test of time. The awards committee that selected the CVPR 2009 awards was also asked to assist in the selection of this award. The prize (sponsored by IBM) was awarded to the following two papers that were published in CVPR 1999: Irfan A. Essa, Sing Bing Kang, Marc Pollefeys |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2011 | Editor's Note
Ramin Zabih, Zoubin Ghahramani, Sing Bing Kang, Jiri Matas |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2011 | Editorial
Ramin Zabih, Zoubin Ghahramani, Sing Bing Kang, Jiri Matas |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2011 | Editor's Note
Ramin Zabih, Sing Bing Kang, Jiri Matas, Max Welling |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2011 | Editor's Note
Ramin Zabih, Sing Bing Kang, Jiri Matas, Max Welling |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2011 | Matting and compositing of transparent and refractive objectsabstractThis article introduces a new approach for matting and compositing transparent and refractive objects in photographs. The key to our work is an image-based matting model, termed the Attenuation-Refraction Matte (ARM), that encodes plausible refractive properties of a transparent object along with its observed specularities and transmissive properties. We show that an object's ARM can be extracted directly from a photograph using simple user markup. Once extracted, the ARM is used to paste the object onto a new background with a variety of effects, including compound compositing, Fresnel effect, scene depth, and even caustic shadows. User studies find our results favorable to those obtained with Photoshop as well as perceptually valid in most cases. Our approach allows photo editing of transparent and refractive objects in a manner that produces realistic effects previously only possible via 3D models or environment matting. Sai-Kit Yeung, Chi-Keung Tang, Michael S. Brown, Sing Bing Kang |
ACM Trans. Graph. | 4 |
| 2010 | Image-Based and Sketch-Based Modeling of Plants and Trees
Sing Bing Kang |
ACCV (1) | 1 |
| 2010 | Removing rolling shutter wobbleabstractWe present an algorithm to remove wobble artifacts from a video captured with a rolling shutter camera undergoing large accelerations or jitter. We show how estimating the rapid motion of the camera can be posed as a temporal super-resolution problem. The low-frequency measurements are the motions of pixels from one frame to the next. These measurements are modeled as temporal integrals of the underlying high-frequency jitter of the camera. The estimated high-frequency motion of the camera is then used to re-render the sequence as though all the pixels in each frame were imaged at the same time. We also present an auto-calibration algorithm that can estimate the time between the capture of subsequent rows in the camera. Simon Baker, Eric P. Bennett, Sing Bing Kang, Richard Szeliski |
CVPR | 3 |
| 2010 | A content-aware image priorabstractIn image restoration tasks, a heavy-tailed gradient distribution of natural images has been extensively exploited as an image prior. Most image restoration algorithms impose a sparse gradient prior on the whole image, reconstructing an image with piecewise smooth characteristics. While the sparse gradient prior removes ringing and noise artifacts, it also tends to remove mid-frequency textures, degrading the visual quality. We can attribute such degradations to imposing an incorrect image prior. The gradient profile in fractal-like textures, such as trees, is close to a Gaussian distribution, and small gradients from such regions are severely penalized by the sparse gradient prior. To address this issue, we introduce an image restoration algorithm that adapts the image prior to the underlying texture. We adapt the prior to both low-level local structures as well as mid-level textural characteristics. Improvements in visual quality is demonstrated on deconvolution and denoising tasks. Taeg Sang Cho, Neel Joshi, C. Lawrence Zitnick, Sing Bing Kang, Richard Szeliski, William T. Freeman |
CVPR | 4 |
| 2010 | Personalization of image enhancementabstractWe address the problem of incorporating user preference in automatic image enhancement. Unlike generic tools for automatically enhancing images, we seek to develop methods that can first observe user preferences on a training set, and then learn a model of these preferences to personalize enhancement of unseen images. The challenge of designing such system lies at intersection of computer vision, learning, and usability; we use techniques such as active sensor selection and distance metric learning in order to solve the problem. The experimental evaluation based on user studies indicates that different users do have different preferences in image enhancement, which suggests that personalization can further help improve the subjective quality of generic image enhancements. Sing Bing Kang, Ashish Kapoor, Dani Lischinski |
CVPR | 1 |
| 2010 | Generating sharp panoramas from motion-blurred videosabstractIn this paper, we show how to generate a sharp panorama from a set of motion-blurred video frames. Our technique is based on joint global motion estimation and multi-frame deblurring. It also automatically computes the duty cycle of the video, namely the percentage of time between frames that is actually exposure time. The duty cycle is necessary for allowing the blur kernels to be accurately extracted and then removed. We demonstrate our technique on a number of videos. Yunpeng Li 0002, Sing Bing Kang, Neel Joshi, Steven M. Seitz, Daniel P. Huttenlocher |
CVPR | 2 |
| 2010 | Using optical defocus to denoiseabstractEffective reduction of noise is generally difficult because of the possible tight coupling of noise with high-frequency image structure. The problem is worse under low-light conditions. In this paper, we propose slightly optically defocusing the image in order to loosen this noise-image structure coupling. This allows us to more effectively reduce noise and subsequently restore the small defocus. We analytically show how this is possible, and demonstrate our technique on a number of examples that include low-light images. Qi Shan, Jiaya Jia, Sing Bing Kang, Zenglu Qin |
CVPR | 3 |
| 2010 | Image deblurring using inertial measurement sensorsabstractWe present a deblurring algorithm that uses a hardware attachment coupled with a natural image prior to deblur images from consumer cameras. Our approach uses a combination of inexpensive gyroscopes and accelerometers in an energy optimization framework to estimate a blur function from the camera's acceleration and angular velocity during an exposure. We solve for the camera motion at a high sampling rate during an exposure and infer the latent image using a joint optimization. Our method is completely automatic, handles per-pixel, spatially-varying blur, and out-performs the current leading image-based methods. Our experiments show that it handles large kernels -- up to at least 100 pixels, with a typical size of 30 pixels. We also present a method to perform "ground-truth" measurements of camera motion blur. We use this method to validate our hardware and deconvolution approach. To the best of our knowledge, this is the first work that uses 6 DOF inertial sensors for dense, per-pixel spatially-varying image deblurring and the first work to gather dense ground-truth measurements for camera-shake blur. Neel Joshi, Sing Bing Kang, C. Lawrence Zitnick, Richard Szeliski |
ACM Trans. Graph. | 2 |
| 2009 | Human video texturesabstractThis paper describes a data-driven approach for generating photorealistic animations of human motion. Each animation sequence follows a user-choreographed path and plays continuously by seamlessly transitioning between different segments of the captured data. To produce these animations, we capitalize on the complementary characteristics of motion capture data and video. We customize our capture system to record motion capture data that are synchronized with our video source. Candidate transition points in video clips are identified using a new similarity metric based on 3-D marker trajectories and their 2-D projections into video. Once the transitions have been identified, a video-based motion graph is constructed. We further exploit hybrid motion and video data to ensure that the transitions are seamless when generating animations. Motion capture marker projections serve as control points for segmentation of layers and nonrigid transformation of regions. This allows warping and blending to generate seamless in-between frames for animation. We show a series of choreographed animations of walks and martial arts scenes as validation of our approach. Matthew Flagg, Atsushi Nakazawa, Qiushuang Zhang, Sing Bing Kang, Young Kee Ryu, Irfan A. Essa, James M. Rehg |
SI3D | 4 |
| 2009 | Texture SplicingabstractAbstract We propose a new texture editing operation called texture splicing. For this operation, we regard a texture as having repetitive elements (textons) seamlessly distributed in a particular pattern. Taking two textures as input, texture splicing generates a new texture by selecting the texton appearance from one texture and distribution from the other. Texture splicing involves self‐similarity search to extract the distribution, distribution warping, context‐dependent warping, and finally, texture refinement to preserve overall appearance. We show a variety of results to illustrate this operation. Yiming Liu 0001, Jiaping Wang, Su Xue, Xin Tong 0001, Sing Bing Kang, Baining Guo |
Comput. Graph. Forum | 5 |
| 2009 | Single-Image Vignetting CorrectionabstractIn this paper, we propose a method for robustly determining the vignetting function given only a single image. Our method is designed to handle both textured and untextured regions in order to maximize the use of available information. To extract vignetting information from an image, we present adaptations of segmentation techniques that locate image regions with reliable data for vignetting estimation. Within each image region, our method capitalizes on the frequency characteristics and physical properties of vignetting to distinguish it from other sources of intensity variation. Rejection of outlier pixels is applied to improve the robustness of vignetting estimation. Comprehensive experiments demonstrate the effectiveness of this technique on a broad range of images with both simulated and natural vignetting effects. Causes of failures using the proposed algorithm are also analyzed. Yuanjie Zheng, Stephen Lin 0001, Chandra Kambhamettu, Jingyi Yu 0001, Sing Bing Kang |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2009 | An MRF-Based DeInterlacing Algorithm With Exemplar-Based RefinementabstractIn this paper, we propose an MRF-based deinterlacing algorithm that combines the benefits of rule-based algorithms such as motion-adaptation, edge-directed interpolation, and motion compensation, with those of an MRF formulation. MRF-based interpolation and enhancement algorithms are typically formulated as an optimization over pixel intensities or colors, which can make them relatively slow. In comparison, our MRF-based deinterlacing algorithm uses interpolation functions as labels.We use seven interpolants (three spatial, three temporal, and one for motion compensation). The core dynamic programming algorithm is, therefore, sped up greatly over the direct use of intensity as labels. We also show how an exemplar-based learning algorithm can be used to refine the output of our MRF-based algorithm. The training set can be augmented with exemplars from static regions of the same video, as a form of "self-learning." Shengyang Dai, Simon Baker, Sing Bing Kang |
IEEE Trans. Image Process. | 3 |
| 2009 | Matte-Based Restoration of Vintage VideoabstractIn this paper, we describe a new method for restoring digitized vintage video with film wear artifacts. Such artifacts result in partially or completely missing information. To maximize use of observed data, we cast the problem as that of recovering mattes of artifacts. More specifically, we extract the distributions of artifact color and its fractional (alpha) contribution to the frame. To account for spatial color discontinuity and pixel occlusion or disocclusion, we introduce the alpha-modulated bilateral filter. The problem is solved as a 3-D spatio-temporal conditional random field (CRF) with artifact color and (discretized) alpha as states. Inference is done through belief propagation. Results verify the effectiveness of our method. Furthermore, we can produce a synthetically generated vintage footage using extracted artifact information from actual vintage video. Jiading Gai, Sing Bing Kang |
IEEE Trans. Image Process. | 2 |
| 2008 | Single-image vignetting correction using radial gradient symmetryabstractIn this paper, we present a novel single-image vignetting method based on the symmetric distribution of the radial gradient (RG). The radial gradient is the image gradient along the radial direction with respect to the image center. We show that the RG distribution for natural images without vignetting is generally symmetric. However, this distribution is skewed by vignetting. We develop two variants of this technique, both of which remove vignetting by minimizing asymmetry of the RG distribution. Compared with prior approaches to single-image vignetting correction, our method does not require segmentation and the results are generally better. Experiments show our technique works for a wide range of images and it achieves a speed-up of 4’5 times compared with a state-of-the-art method. Yuanjie Zheng, Jingyi Yu 0001, Sing Bing Kang, Stephen Lin 0001, Chandra Kambhamettu |
CVPR | 3 |
| 2008 | Automatic Estimation and Removal of Noise from a Single ImageabstractImage denoising algorithms often assume an additive white Gaussian noise (AWGN) process that is independent of the actual RGB values. Such approaches are not fully automatic and cannot effectively remove color noise produced by todays CCD digital camera. In this paper, we propose a unified framework for two tasks: automatic estimation and removal of color noise from a single image using piecewise smooth image models. We introduce the noise level function (NLF), which is a continuous function describing the noise level as a function of image brightness. We then estimate an upper bound of the real noise level function by fitting a lower envelope to the standard deviations of per-segment image variances. For denoising, the chrominance of color noise is significantly removed by projecting pixel values onto a line fit to the RGB values in each segment. Then, a Gaussian conditional random field (GCRF) is constructed to obtain the underlying clean image from the noisy input. Extensive experiments are conducted to test the proposed algorithm, which is shown to outperform state-of-the-art denoising algorithms. Ce Liu 0001, Richard Szeliski, Sing Bing Kang, C. Lawrence Zitnick, William T. Freeman |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2008 | Sketching reality: Realistic interpretation of architectural designsabstractIn this article, we introduce sketching reality , the process of converting a freehand sketch into a realistic-looking model. We apply this concept to architectural designs. As the sketch is being drawn, our system periodically interprets its 2.5D-geometry by identifying new junctions, edges, and faces, and then analyzing the extracted topology. The user can add detailed geometry and textures through sketches as well. This is possible through the use of databases that match partial sketches to models of detailed geometry and textures. The final product is a realistic texture-mapped 2.5D-model of the building. We show a variety of buildings that have been created using this system. Xuejin Chen, Sing Bing Kang, Ying-Qing Xu, Julie Dorsey, Harry Shum |
ACM Trans. Graph. | 2 |
| 2008 | Sketch-based tree modeling using Markov random fieldabstractIn this paper, we describe a new system for converting a user's freehand sketch of a tree into a full 3D model that is both complex and realistic-looking. Our system does this by probabilistic optimization based on parameters obtained from a database of tree models. The best matching model is selected by comparing its 2D projections with the sketch. Branch interaction is modeled by a Markov random field, subject to the constraint of 3D projection to sketch. Our system then uses the notion of self-similarity to add new branches before finally populating all branches with leaves of the user's choice. We show a variety of natural-looking tree models generated from freehand sketches with only a few strokes. Xuejin Chen, Boris Neubert, Ying-Qing Xu, Oliver Deussen, Sing Bing Kang |
ACM Trans. Graph. | 5 |
| 2007 | Automatic Removal of Chromatic Aberration from a Single ImageabstractMany high resolution images exhibit chromatic aberration (CA), where the color channels appear shifted. Unfortunately, merely compensating for these shifts is sometimes inadequate, because the intensities are modified by other effects such as spatially-varying defocus and (surprisingly) in-camera sharpening. In this paper, we start from the basic principles of image formation to characterize CA, and show how its effects can be substantially reduced. We also show results of CA correction on a number of high-resolution images taken with different cameras. Sing Bing Kang |
CVPR | 1 |
| 2007 | Inferring Temporal Order of Images From 3D StructureabstractIn this paper, we describe a technique to temporally sort a collection of photos that span many years. By reasoning about persistence of visible structures, we show how this sorting task can be formulated as a constraint satisfaction problem (CSP). Casting this problem as a CSP allows us to efficiently find a suitable ordering of the images despite the large size of the solution space (factorial in the number of images) and the presence of occlusions. We present experimental results for photographs of a city acquired over a one hundred year period. Grant Schindler, Frank Dellaert, Sing Bing Kang |
CVPR | 3 |
| 2007 | Flash Cut: Foreground Extraction with Flash and No-flash Image PairsabstractIn this paper, we propose a novel approach for foreground layer extraction using flash/no-flash image pairs, which we call flash cut. Flash cut is based on the simple observation that only the foreground is significantly brightened by the flash and the background appearance change is very small, if the background is distant. Changes due to flash, motion, and color information are fused in an MRF framework to produce high quality segmentation results. Flash cut handles some amount of camera shake, and foreground motion, which makes it practical for anyone with a flash-equipped camera to use. We validate our approach on a variety of indoor and outdoor examples. Jian Sun 0009, Jian Sun 0001, Sing Bing Kang, Zongben Xu, Xiaoou Tang, Harry Shum |
CVPR | 3 |
| 2007 | Layered Depth PanoramasabstractRepresentations for interactive photorealistic visualization of scenes range from compact 2D panoramas to data-intensive 4D light fields. In this paper, we propose a technique for creating a layered representation from a sparse set of images taken with a hand-held camera. This representation, which we call a layered depth panorama (LDP), allows the user to experience 3D by off-axis panning. It combines the compelling experience of panoramas with limited 3D navigation. Our choice of representation is motivated by ease of capture and compactness. We formulate the problem of constructing the LDP as the recovery of color and geometry in a multi-perspective cylindrical disparity space. We leverage a graph cut approach to sequentially determine the disparity and color of each layer using multi-view stereo. Geometry visible through the cracks at depth discontinuities in a frontmost layer is determined and assigned to layers behind the frontmost layer. All layers are then used to render novel panoramic views with parallax. We demonstrate our approach on a variety of complex outdoor and indoor scenes. Ke Colin Zheng, Sing Bing Kang, Michael F. Cohen, Richard Szeliski |
CVPR | 2 |
| 2007 | Using Photographs to Enhance Videos of a Static Scene
Pravin Bhat, C. Lawrence Zitnick, Noah Snavely, Aseem Agarwala, Maneesh Agrawala, Michael F. Cohen, Brian Curless, Sing Bing Kang |
Rendering Techniques | 8 |
| 2007 | Stereo for Image-Based Rendering using Image Over-Segmentation
C. Lawrence Zitnick, Sing Bing Kang |
Int. J. Comput. Vis. | 2 |
| 2007 | Parameter-Free Radial Distortion Correction with Center of Distortion EstimationabstractWe propose a method of simultaneously calibrating the radial distortion function of a camera and the other internal calibration parameters. The method relies on the use of a planar (or, alternatively, nonplanar) calibration grid which is captured in several images. In this way, the determination of the radial distortion is an easy add-on to the popular calibration method proposed by Zhang [24]. The method is entirely noniterative and, hence, is extremely rapid and immune to the problem of local minima. Our method determines the radial distortion in a parameter-free way, not relying on any particular radial distortion model. This makes it applicable to a large range of cameras from narrow-angle to fish-eye lenses. The method also computes the center of radial distortion, which, we argue, is important in obtaining optimal results. Experiments show that this point may be significantly displaced from the center of the image or the principal point of the camera. Richard I. Hartley, Sing Bing Kang |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2007 | Image-based tree modelingabstractIn this paper, we propose an approach for generating 3D models of natural-looking trees from images that has the additional benefit of requiring little user intervention. While our approach is primarily image-based, we do not model each leaf directly from images due to the large leaf count, small image footprint, and widespread occlusions. Instead, we populate the tree with leaf replicas from segmented source images to reconstruct the overall tree shape. In addition, we use the shape patterns of visible branches to predict those of obscured branches. We demonstrate our approach on a variety of trees. Ping Tan 0002, Jingdong Wang 0001, Sing Bing Kang, Long Quan |
ACM Trans. Graph. | 4 |
| 2007 | High Resolution Animated Scenes from StillsabstractCurrent techniques for generating animated scenes involve either videos (whose resolution is limited) or a single image (which requires a significant amount of user interaction). In this paper, we describe a system that allows the user to quickly and easily produce a compelling-looking animation from a small collection of high resolution stills. Our system has two unique features. First, it applies an automatic partial temporal order recovery algorithm to the stills in order to approximate the original scene dynamics. The output sequence is subsequently extracted using a second-order Markov Chain model. Second, a region with large motion variation can be automatically decomposed into semiautonomous regions such that their temporal orderings are softly constrained. This is to ensure motion smoothness throughout the original region. The final animation is obtained by frame interpolation and feathering. Our system also provides a simple-to-use interface to help the user to fine-tune the motion of the animated scene. Using our system, an animated scene can be generated in minutes. We show results for a variety of scenes. Zhouchen Lin, Lifeng Wang 0001, Yunbo Wang, Sing Bing Kang, Tian Fang |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2006 | Noise Estimation from a Single ImageabstractIn order to work well, many computer vision algorithms require that their parameters be adjusted according to the image noise level, making it an important quantity to estimate. We show how to estimate an upper bound on the noise level from a single image based on a piecewise smooth image prior model and measured CCD camera response functions. We also learn the space of noise level functions how noise level changes with respect to brightness and use Bayesian MAP inference to infer the noise level function from a single image. We illustrate the utility of this noise estimation for two algorithms: edge detection and featurepreserving smoothing through bilateral filtering. For a variety of different noise levels, we obtain good results for both these algorithms with no user-specified inputs. Ce Liu 0001, William T. Freeman, Richard Szeliski, Sing Bing Kang |
CVPR (1) | 4 |
| 2006 | Video Completion by Motion Field TransferabstractExisting methods for video completion typically rely on periodic color transitions, layer extraction, or temporally local motion. However, periodicity may be imperceptible or absent, layer extraction is difficult, and temporally local motion cannot handle large holes. This paper presents a new approach for video completion using motion field transfer to avoid such problems. Unlike prior methods, we fill in missing video parts by sampling spatio-temporal patches of local motion instead of directly sampling color. Once the local motion field has been computed within the missing parts of the video, color can then be propagated to produce a seamless hole-free video. We have validated our method on many videos spanning a variety of scenes. We can also use the same approach to perform frame interpolation using motion fields from different videos. Takaaki Shiratori, Yasuyuki Matsushita, Xiaoou Tang, Sing Bing Kang |
CVPR (1) | 4 |
| 2006 | Reconstructing Occluded Surfaces Using Synthetic Apertures: Stereo, Focus and Robust MeasuresabstractMost algorithms for 3D reconstruction from images use cost functions based on SSD, which assume that the surfaces being reconstructed are visible to all cameras. This makes it difficult to reconstruct objects which are partially occluded. Recently, researchers working with large camera arrays have shown it is possible to "see through" occlusions using a technique called synthetic aperture focusing. This suggests that we can design alternative cost functions that are robust to occlusions using synthetic apertures. Our paper explores this design space. We compare classical shape from stereo with shape from synthetic aperture focus. We also describe two variants of multi-view stereo based on color medians and entropy that increase robustness to occlusions. We present an experimental comparison of these cost functions on complex light fields, measuring their accuracy against the amount of occlusion. Vaibhav Vaish, Marc Levoy, Richard Szeliski, C. Lawrence Zitnick, Sing Bing Kang |
CVPR (2) | 5 |
| 2006 | Single-Image Vignetting CorrectionabstractIn this paper, we propose a method for determining the vignetting function given only a single image. Our method is designed to handle both textured and untextured regions in order to maximize the use of available information. To extract vignetting information from an image, we present adaptations of segmentation techniques that locate image regions with reliable data for vignetting estimation. Within each image region, our method capitalizes on frequency characteristics and physical properties of vignetting to distinguish it from other sources of intensity variation. The vignetting data acquired from regions are weighted according to a presented reliability measure to promote robustness in estimation. Comprehensive experiments demonstrate the effectiveness of this technique on a broad range of images. Yuanjie Zheng, Stephen Lin 0001, Sing Bing Kang |
CVPR (1) | 3 |
| 2006 | Video and Image Bayesian Demosaicing with a Two Color Image Prior
Eric P. Bennett, Matthew Uyttendaele, C. Lawrence Zitnick, Richard Szeliski, Sing Bing Kang |
ECCV (1) | 5 |
| 2006 | Boundary matting for view synthesis
Samuel W. Hasinoff, Sing Bing Kang, Richard Szeliski |
Comput. Vis. Image Underst. | 2 |
| 2006 | Stereo Matching with Linear Superposition of LayersabstractIn this paper, we address stereo matching in the presence of a class of non-Lambertian effects, where image formation can be modeled as the additive superposition of layers at different depths. The presence of such effects makes it impossible for traditional stereo vision algorithms to recover depths using direct color matching-based methods. We develop several techniques to estimate both depths and colors of the component layers. Depth hypotheses are enumerated in pairs, one from each layer, in a nested plane sweep. For each pair of depth hypotheses, matching is accomplished using spatial-temporal differencing. We then use graph cut optimization to solve for the depths of both layers. This is followed by an iterative color update algorithm which we proved to be convergent. Our algorithm recovers depth and color estimates for both synthetic and real image sequences. Yanghai Tsin, Sing Bing Kang, Richard Szeliski |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2006 | Image-based plant modelingabstractIn this paper, we propose a semi-automatic technique for modeling plants directly from images. Our image-based approach has the distinct advantage that the resulting model inherits the realistic shape and complexity of a real plant. We designed our modeling system to be interactive, automating the process of shape recovery while relying on the user to provide simple hints on segmentation. Segmentation is performed in both image and 3D spaces, allowing the user to easily visualize its effect immediately. Using the segmented image and 3D data, the geometry of each leaf is then automatically recovered from the multiple views by fitting a deformable leaf model. Our system also allows the user to easily reconstruct branches in a similar manner. We show realistic reconstructions of a variety of plants, and demonstrate examples of plant editing. Long Quan, Ping Tan 0002, Lu Yuan 0001, Jingdong Wang 0001, Sing Bing Kang |
ACM Trans. Graph. | 6 |
| 2006 | Flash mattingabstractIn this paper, we propose a novel approach to extract mattes using a pair of flash/no-flash images. Our approach, which we call flash matting , was inspired by the simple observation that the most noticeable difference between the flash and no-flash images is the foreground object if the background scene is sufficiently distant. We apply a new matting algorithm called joint Bayesian flash matting to robustly recover the matte from flash/no-flash images, even for scenes in which the foreground and the background are similar or the background is complex. Experimental results involving a variety of complex indoors and outdoors scenes show that it is easy to extract high-quality mattes using an off-the-shelf, flash-equipped camera. We also describe extensions to flash matting for handling more general scenes. Jian Sun 0001, Yin Li 0003, Sing Bing Kang, Harry Shum |
ACM Trans. Graph. | 3 |
| 2006 | Animating Chinese paintings through stroke-based decompositionabstractThis article proposes a technique to animate a Chinese style painting given its image. We first extract descriptions of the brush strokes that hypothetically produced it. The key to the extraction process is the use of a brush stroke library, which is obtained by digitizing single brush strokes drawn by an experienced artist. The steps in our extraction technique are first to segment the input image, then to find the best set of brush strokes that fit the regions, and, finally, to refine these strokes to account for local appearance. We model a single brush stroke using its skeleton and contour, and we characterize texture variation within each stroke by sampling perpendicularly along its skeleton. Once these brush descriptions have been obtained, the painting can be animated at the brush stroke level. In this article, we focus on Chinese paintings with relatively sparse strokes. The animation is produced using a graphical application we developed. We present several animations of real paintings using our technique. Songhua Xu, Ying-Qing Xu, Sing Bing Kang, David Salesin, Yunhe Pan, Harry Shum |
ACM Trans. Graph. | 3 |
| 2005 | Symmetric Stereo Matching for Occlusion HandlingabstractIn this paper, we propose a symmetric stereo model to handle occlusion in dense two-frame stereo. Our occlusion reasoning is directly based on the visibility constraint that is more general than both ordering and uniqueness constraints used in previous work. The visibility constraint requires occlusion in one image and disparity in the other to be consistent. We embed the visibility constraint within an energy minimization framework, resulting in a symmetric stereo model that treats left and right images equally. An iterative optimization algorithm is used to approximate the minimum of the energy using belief propagation. Our stereo model can also incorporate segmentation as a soft constraint. Experimental results on the Middlebury stereo images show that our algorithm is state-of-the-art. Jian Sun 0001, Yin Li 0003, Sing Bing Kang |
CVPR (2) | 3 |
| 2005 | Parameter-Free Radial Distortion Correction with Centre of Distortion EstimationabstractWe propose a method of simultaneously calibrating the radial distortion function of a camera along with the other internal calibration parameters. The method relies on the use of a planar (or alternatively nonplanar) calibration grid, which is captured in several images. In this way, the determination of the radial distortion is an easy add-on to the popular calibration method proposed by Zhang [1999]. The method is entirely noniterative, and hence is extremely rapid and immune from the problem of local minima. Our method determines the radial distortion in a parameter-free way, not relying on any particular radial distortion model. This makes it applicable to a large range of cameras from narrow-angle to fish-eye lenses. The method also computes the centre of radial distortion, which we argue is important in obtaining optimal results. Experiments show that this point may be significantly displaced from the centre of the image, or the principal point of the camera. Richard I. Hartley, Sing Bing Kang |
ICCV | 2 |
| 2005 | Separating Reflections in Human Iris Images for Illumination EstimationabstractA method is presented for separating corneal reflections in an image of human irises to estimate illumination from the surrounding scene. Previous techniques for reflection separation have demonstrated success in only limited cases, such as for uniform colored lighting and simple object textures, so they are not applicable to irises which exhibit intricate textures and complicated reflections of the environment. To make this problem feasible, we present a method that capitalizes on physical characteristics of human irises to obtain an illumination estimate that encompasses the prominent light contributors in the scene. Results of this algorithm are presented for eyes of different colors, including light colored eyes for which reflection separation is necessary to determine a valid illumination estimate. Huiqiong Wang, Stephen Lin 0001, Xiaopei Liu, Sing Bing Kang |
ICCV | 4 |
| 2005 | Consistent Segmentation for Optical Flow EstimationabstractIn this paper, we propose a method for jointly computing optical flow and segmenting video while accounting for mixed pixels (matting). Our method is based on statistical modeling of an image pair using constraints on appearance and motion. Segments are viewed as overlapping regions with fractional (/spl alpha/) contributions. Bidirectional motion is estimated based on spatial coherence and similarity of segment colors. Our model is extended to video by chaining the pairwise models to produce a joint probability distribution to be maximized. To make the problem more tractable, we factorize the posterior distribution and iteratively minimize its parts. We demonstrate our method on frame interpolation. C. Lawrence Zitnick, Nebojsa Jojic, Sing Bing Kang |
ICCV | 3 |
| 2005 | Extracting layers and analyzing their specular properties using epipolar-plane-image analysis
Antonio Criminisi, Sing Bing Kang, Rahul Swaminathan, Richard Szeliski, P. Anandan 0001 |
Comput. Vis. Image Underst. | 2 |
| 2004 | Estimating Intrinsic Images from Image Sequences with Biased Illumination
Yasuyuki Matsushita, Stephen Lin 0001, Sing Bing Kang, Harry Shum |
ECCV (2) | 3 |
| 2004 | Extracting View-Dependent Depth Maps from a Collection of Images
Sing Bing Kang, Richard Szeliski |
Int. J. Comput. Vis. | 1 |
| 2004 | Error Analysis of Pure Rotation-Based Self-CalibrationabstractSelf-calibration using pure rotation is a well-known technique and has been shown to be a reliable means for recovering intrinsic camera parameters. However, in practice, it is virtually impossible to ensure that the camera motion for this type of self-calibration is a pure rotation. In this paper, we present an error analysis of recovered intrinsic camera parameters due to the presence of translation. We derived closed-form error expressions for a single pair of images with nondegenerate motion; for multiple rotations for which there are no closed-form solutions, analysis was done through repeated experiments. Among others, we show that translation-independent solutions do exist under certain practical conditions. Our analysis can be used to help choose the least error-prone approach (if multiple approaches exist) for a given set of conditions. Sing Bing Kang, Harry Shum, Guangyou Xu |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2004 | High-quality video view interpolation using a layered representationabstractThe ability to interactively control viewpoint while watching a video is an exciting application of image-based rendering. The goal of our work is to render dynamic scenes with interactive viewpoint control using a relatively small number of video cameras. In this paper, we show how high-quality video-based rendering of dynamic scenes can be accomplished using multiple synchronized video streams combined with novel image-based modeling and rendering algorithms. Once these video streams have been processed, we can synthesize any intermediate view between cameras at any time, with the potential for space-time manipulation.In our approach, we first use a novel color segmentation-based stereo algorithm to generate high-quality photoconsistent correspondences across all camera views. Mattes for areas near depth discontinuities are then automatically extracted to reduce artifacts during view synthesis. Finally, a novel temporal two-layer compressed representation that handles matting is developed for rendering at interactive rates. C. Lawrence Zitnick, Sing Bing Kang, Matthew Uyttendaele, Simon A. J. Winder, Richard Szeliski |
ACM Trans. Graph. | 2 |
| 2003 | Directional Histogram Model for Three-Dimensional Shape SimilarityabstractIn this paper, we propose a novel shape representation we call directional histogram model (DHM). It captures the shape variation of an object and is invariant to scaling and rigid transforms. The DHM is computed by first extracting a directional distribution of thickness histogram signatures, which are translation invariant. We show how the extraction of the thickness histogram distribution can be accelerated using conventional graphics hardware. Orientation invariance is achieved by computing the spherical harmonic transform of this distribution. Extensive experiments show that the DHM is capable of high discrimination power and is robust to noise. Xinguo Liu, Robin Sun, Sing Bing Kang, Harry Shum |
CVPR (1) | 3 |
| 2003 | Stereo Matching with Reflections and TranslucencyabstractIn this paper, we address the stereo matching problem in the presence of reflections and translucency, where image formation can be modeled as the additive superposition of layers at different depth. The presence of such effects violates the Lambertian assumption underlying traditional stereo vision algorithms, making it impossible to recover component depths using direct color matching based methods. We develop several techniques to estimate both depths and colors of the component layers. Depth hypotheses are enumerated in pairs, one from each layer, in a nested plane sweep. For each pair of depth hypotheses, we compute a component-color-independent matching error per pixel, using a spatial-temporal differencing technique. We then use graph cut optimization to solve for the depths of both layers. This is followed by an iterative color update algorithm whose convergence is proven in our paper. We show convincing results of depth and color estimates for both synthetic and real image sequences. Yanghai Tsin, Sing Bing Kang, Richard Szeliski |
CVPR (1) | 2 |
| 2003 | Large environment rendering using plenoptic primitivesabstractOne of the most difficult tasks in computer graphics is to enable virtual walkthroughs in very large and complicated environments that are photorealistic, seamless, and in real time. Current image-based rendering techniques, while capable of photorealism and interactive speeds, have failed in practice to extend to visualizations of such environments. We demonstrate an approach that defines a virtual walkthrough experience using plenoptic primitives (PPs). A PP can be any type of local visual experience: 360/spl deg/ static panorama, panoramic video (PV), lumigraph/light field representation, or concentric mosaics (CMs). By combining them judiciously, user experience can be authored with significantly reduced effort while maintaining high-quality user experience. We illustrate our technique on synthetic and real environments using PVs and CMs and show how the problem of achieving smooth transitions among PVs and CMs can be solved by using position-dependent local geometries. Sing Bing Kang, Minsheng Wu, Yin Li 0003, Harry Shum |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2003 | Survey of image-based representations and compression techniquesabstractWe survey the techniques for image-based rendering (IBR) and for compressing image-based representations. Unlike traditional three-dimensional (3-D) computer graphics, in which 3-D geometry of the scene is known, IBR techniques render novel views directly from input images. IBR techniques can be classified into three categories according to how much geometric information is used: rendering without geometry, rendering with implicit geometry (i.e., correspondence), and rendering with explicit geometry (either with approximate or accurate geometry). We discuss the characteristics of these categories and their representative techniques. IBR techniques demonstrate a surprising diverse range in their extent of use of images and geometry in representing 3-D scenes. We explore the issues in trading off the use of images and geometry by revisiting plenoptic-sampling analysis and the notions of view dependency and geometric proxies. Finally, we highlight compression techniques specifically designed for image-based representations. Such compression techniques are important in making IBR techniques practical. Harry Shum, Sing Bing Kang, S. C. Chan 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2003 | High dynamic range videoabstractTypical video footage captured using an off-the-shelf camcorder suffers from limited dynamic range. This paper describes our approach to generate high dynamic range (HDR) video from an image sequence of a dynamic scene captured while rapidly varying the exposure of each frame. Our approach consists of three parts: automatic exposure control during capture, HDR stitching across neighboring frames, and tonemapping for viewing. HDR stitching requires accurately registering neighboring frames and choosing appropriate pixels for computing the radiance map. We show examples for a variety of dynamic scenes. We also show how we can compensate for scene and camera movement when creating an HDR still from a series of bracketed still photographs. Sing Bing Kang, Matthew Uyttendaele, Simon A. J. Winder, Richard Szeliski |
ACM Trans. Graph. | 1 |
| 2002 | Diffuse-Specular Separation and Depth Recovery from Image Sequences
Stephen Lin 0001, Yuanzhen Li, Sing Bing Kang, Xin Tong 0001, Harry Shum |
ECCV (3) | 3 |
| 2002 | On the Motion and Appearance of Specularities in Image Sequences
Rahul Swaminathan, Sing Bing Kang, Richard Szeliski, Antonio Criminisi, Shree K. Nayar |
ECCV (1) | 2 |
| 2002 | Single-Image Reflectance Estimation for Relighting by Iterative Soft GroupingabstractReflectance values for image-based relighting are often estimated from grouped pixels with similar reflectance, but such groupings are difficult to compute with certainty for sparse image data. To address this problem, we propose an iterative method that aggregates BRDF data in a single image with known geometry and lighting by soft grouping, where pixels contribute to one another's estimate according to their degree of reflectance similarity. Estimation of specular reflectance is further improved by albedo-independent soft grouping of pixels based on shape continuity. With recovered reflectances, we demonstrate realistic relighting for synthetic and real scenes, including surfaces with spatially-varying reflectance. Yuanzhen Li, Stephen Lin 0001, Sing Bing Kang, Hanqing Lu, Harry Shum |
PG | 3 |
| 2002 | Lighting Interpolation by Shadow Morphing Using Intrinsic LumigraphsabstractDensely-sampled image representations such as the light field or lumigraph have been effective in enabling photorealistic image synthesis. Unfortunately, lighting interpolation with such representations has not been shown to be possible without the use of accurate 3D geometry and surface reflectance properties. In this paper we propose an approach to image-based lighting interpolation that is based on estimates of geometry and shading from relatively few images. We decompose captured light fields at different lighting conditions into intrinsic images (reflectance and illumination images), and estimate view-dependent scene geometries using multi-view stereo. We call the resulting representation an intrinsic lumigraph. In the same way that the lumigraph uses geometry to permit more accurate view interpolation, the intrinsic lumigraph uses both geometry and intrinsic images to allow high-quality interpolation at different views and lighting conditions. Joint use of geometry and intrinsic images is effective in the computation of shadow masks for shadow prediction at new lighting conditions. We illustrate our approach with images of real scenes. Yasuyuki Matsushita, Sing Bing Kang, Stephen Lin 0001, Harry Shum, Xin Tong 0001 |
PG | 2 |
| 2002 | Appearance-Based Structure from Motion Using Linear Classes of 3-D Models
Sing Bing Kang |
Int. J. Comput. Vis. | 1 |
| 2001 | Handling Occlusions in Dense Multi-view StereoabstractWhile stereo matching was originally formulated as the recovery of 3D shape from a pair of images, it is now generally recognized that using more than two images can dramatically improve the quality of the reconstruction. Unfortunately, as more images are added, the prevalence of semi-occluded regions (pixels visible in some but not all images) also increases. We propose some novel techniques to deal with this problem. Our first idea is to use a combination of shiftable windows and a dynamically selected subset of the neighboring images to do the matches. Our second idea is to explicitly label occluded pixels within a global energy minimization framework, and to reason about visibility within this framework so that only truly visible pixels are matched. Experimental results show a dramatic improvement using the first idea over conventional multibaseline stereo, especially when used in conjunction with a global energy minimization technique. These results also show that explicit occlusion labeling and visibility reasoning do help, but not significantly, if the spatial and temporal selection is applied first. Sing Bing Kang, Richard Szeliski, Jinxiang Chai |
CVPR (1) | 1 |
| 2001 | Optimal Texture Map Reconstruction from Multiple ViewsabstractThe recovery of 3D models from multiple reference images involves not only the extraction of 3D shape, but also of texture. Assuming that all surfaces are Lambertian, the resulting final texture is typically computed as a linear combination of reference textures. This is, however, not the optimal means for reconstructing textures, since this does not model the anisotropy in the texture projection. Furthermore, the spatial image sampling may be quite variable within a fore-shortened surface. This also has important implications for computer vision techniques that involve analysis by synthesis and the image-based rendering (IBR) technique of view-dependent texture mapping (VDTM). Starting with sampling theory, we show how weights should be spatially distributed for optimal texture construction. The local weights take into consideration the effects of anisotropy and variable spatial image sampling. We also present experimental results to verify our analysis. Lifeng Wang 0001, Sing Bing Kang, Richard Szeliski, Harry Shum |
CVPR (1) | 2 |
| 2001 | Error Analysis of Pure Rotation-Based Self-Calibration
Leslie Wang, Sing Bing Kang, Harry Shum, Guangyou Xu |
ICCV | 2 |
| 2000 | Catadioptric Self-CalibrationabstractWe have assembled a standalone, movable system that can capture long sequences of omnidirectional images (up to 1,500 images at 6.7 Hz and a resolution of 1140/spl times/090). The goal of this system is to reconstruct complex large environments, such as an entire floor of a building, from the captured images only. In this paper, we address the important issue of how to calibrate such a system. Our method uses images of the environment to calibrate the camera, without the use of any special calibration pattern, knowledge of camera motion, or knowledge of scene geometry. It uses the consistency of pairwise tracked point features across a sequence based on the characteristics of catadioptric imaging. We also show how the projection equation for this catadioptric camera can be formulated to be equivalent to that of a typical rectilinear perspective camera with just a simple transformation. Sing Bing Kang |
CVPR | 1 |
| 2000 | Visual Tunnel Analysis for Visibility Prediction and Camera PlanningabstractA sequence of images taken along a camera trajectory captures a subset of scene appearance. If visibility space is the space that encapsulates the appearance of the scene at every conceivable pose and viewing angle, then the act of acquiring the image sequence constitutes "carving a volume in visibility space." We call such a volume a visual tunnel. The analysis of the visual tunnel allows us to do the following: predict the range of virtual camera poses in which the images can be reconstructed totally using the captured rays, predict which parts of the image can be generated for a given virtual camera pose, and plan camera paths for scene visualization at desired locations. We describe our visual tunnel concept and provide illustrative examples in 2D and 3D. Sing Bing Kang, Peter-Pike J. Sloan, Steven M. Seitz |
CVPR | 1 |
| 2000 | Can We Calibrate a Camera Using an Image of a Flat, Textureless Lambertian Surface?
Sing Bing Kang, Richard Weiss 0001 |
ECCV (2) | 1 |
| 2000 | The Geometry-Image Representation Tradeoff for RenderingabstractIt is generally recognized that 3-D models are compact representations for rendering. While pure image-based rendering techniques are capable of producing highly photorealistic outputs, the size of the input "model" is usually very large. The important issues in trading off geometry versus images include compactness of representation, photorealism of reconstructed views, and speed of rendering. We describe our past work in modeling and rendering, and articulate lessons learnt. We then delineate our vision of an ideal rendering system. Sing Bing Kang, Richard Szeliski, P. Anandan 0001 |
ICIP | 1 |
| 2000 | Video Editing Using Figure Tracking and Image-Based RenderingabstractWe describe a new approach to video editing based on the semi-automatic segmentation of video into multiple layers and the composition of layers using image-based rendering. Using figure tracking and background motion estimation, we can segment a moving figure and reconstruct the background. Using geometrically-correct pixel reprojection, layers can be composited on the basis of the geometry of the underlying scene and the position of a virtual camera. We have implemented a prototype editing system called SpliceWorld. James M. Rehg, Sing Bing Kang, Tat-Jen Cham |
ICIP | 2 |
| 2000 | Combined Spline- and Block-Based Motion Estimation for Video CodingabstractWe propose a technique to estimate motion for video coding. Our technique combines spline-based registration and block matching motion estimation. It replaces the initial compute-intensive gross block matching search with spline-based registration, which is more efficient in recovering full-image motion fields. This results in a smooth motion field representative of the true motion in the scene, which can be more efficiently encoded and be guaranteed of high visual quality as well. Furthermore, the implementation is computationally cost-effective. Experimental results on well-known test image sequences show that the method results in an increase in coding efficiency along with a reduction in computational cost. Frédéric Dufaux, Sing Bing Kang |
ICPR | 2 |
| 2000 | Review of image-based rendering techniques
Harry Shum, Sing Bing Kang |
VCIP | 2 |
| 1999 | Multi-Layered Image-Based Rendering
Sing Bing Kang, Huong Quynh Dinh |
Graphics Interface | 1 |
| 1999 | Characterization of Errors in Compositing Panoramic Images
Sing Bing Kang, Richard Weiss 0001 |
Comput. Vis. Image Underst. | 1 |
| 1999 | Registration and integration of textured 3D data
Andrew E. Johnson 0002, Sing Bing Kang |
Image Vis. Comput. | 2 |
| 1998 | Virtual Navigation of Complex Scenes using Clusters of Cylindrical Panoramic Images
Sing Bing Kang, Pavan K. Desikan |
Graphics Interface | 1 |
| 1998 | Hands-free navigation in VR environments by tracking the head
Sing Bing Kang |
Int. J. Hum. Comput. Stud. | 1 |
| 1997 | Characterization of errors in compositing panoramic imagesabstractIn this paper we describe the effect of errors in the intrinsic camera parameters on reconstructed panoramic images. A panoramic image is created by first capturing a sequence of images while rotating the camera about a vertical axis a full 360/spl deg/. The subsequent steps are projecting the original rectilinear images onto cylindrical surfaces and compositing them to form the panoramic image. Our analysis has led to a technique that allows simultaneous recovery of the camera focal length and properly composited panoramic images. The correct focal length can be determined by iterating the processes of projecting the original rectilinear images onto cylindrical surfaces given an estimate of the focal length and compositing the resulting images to yield an increasingly better estimate of the focal length. This paper shows that the convergence towards the correct focal length is exponential. Sing Bing Kang, Richard Weiss 0001 |
CVPR | 1 |
| 1997 | A Parallel Feature Tracker for Extended Image Sequences
Sing Bing Kang, Richard Szeliski, Harry Shum |
Comput. Vis. Image Underst. | 1 |
| 1997 | 3-D Scene Data Recovery Using Omnidirectional Multibaseline Stereo
Sing Bing Kang, Richard Szeliski |
Int. J. Comput. Vis. | 1 |
| 1997 | Shape Ambiguities in Structure From MotionabstractThis paper examines the fundamental ambiguities and uncertainties inherent in recovering structure from motion. By examining the eigenvectors associated with null or small eigenvalues of the Hessian matrix, we can quantify the exact nature of these ambiguities and predict how they affect the accuracy of the reconstructed shape. Our results for orthographic cameras show that the bas-relief ambiguity is significant even with many images, unless a large amount of rotation is present. Similar results for perspective cameras suggest that three or more frames and a large amount of rotation are required for metrically accurate reconstruction. Richard Szeliski, Sing Bing Kang |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 1997 | Toward automatic robot instruction from perception-mapping human grasps to manipulator graspsabstractOur approach of programming a robot is by direct human demonstration. The system observes a human performing the task, recognizes the human grasp, and maps it onto the manipulator. This paper describes how an observed human grasp can be mapped to that of a given general-purpose manipulator for task replication. Planning the manipulator grasp based upon the observed human grasp is done at two levels: the functional and physical levels. Initially, at the functional level, grasp mapping is achieved at the virtual finger level; the virtual finger is a group of fingers acting against an object surface in a similar manner. Subsequently, at the physical level, the geometric properties of the object and manipulator are considered in fine-tuning the manipulator grasp. Our work concentrates on power or enveloping grasps and the fingertip precision grasps. We conclude by showing an example of an entire programming cycle from human demonstration to robot execution. Sing Bing Kang, Katsushi Ikeuchi |
IEEE Trans. Robotics Autom. | 1 |
| 1996 | 3-D Scene Data Recovery using Omnidirectional Multibaseline StereoabstractA traditional approach to extracting geometric information from a large scene is to compute multiple 3-D depth maps from stereo pairs or direct range finders, and then to merge the 3-D data. However, the resulting merged depth maps may be subject to merging errors if the relative poses between depth maps are not known exactly. In addition, the 3-D data may also have to be resampled before merging, which adds additional complexity and potential sources of errors. This paper provides a means of directly extracting 3-D data covering a very wide field of view, thus by-passing the need for numerous depth map merging. In our work, cylindrical images are first composited from sequences of images taken while the camera is rotated 360/spl deg/ about a vertical axis. By taking such image panoramas at different camera locations, we can recover 3-D data of the scene using a set of simple techniques: feature tracking, an 8-point structure from motion algorithm, and multibaseline stereo. We also investigate the effect of median filtering on the recovered 3-D point distributions, and show the results of our approach applied to both synthetic and real scenes. Sing Bing Kang, Richard Szeliski |
CVPR | 1 |
| 1996 | Shape Ambiguities in Structure from Motion
Richard Szeliski, Sing Bing Kang |
ECCV (1) | 2 |
| 1995 | A Robot System that Observes and Replicates Grasping TasksabstractTo alleviate the problem of overwhelming complexity in grasp synthesis and path planning associated with robot task planning, we adopt the approach of teaching the robot by demonstrating in front of it. The system has four components: the observation system, the grasping task recognition module, the task translator and the robot system. The observation system comprises an active multibaseline stereo system and a dataglove. The data stream recorded is then used to track object motion; this paper illustrates how complimentary sensory data can be used for this purpose. The data stream is also interpreted by the grasping task recognition module, which produces higher levels of abstraction to describe both the motion and actions taken in the task. The resulting information are provided to the task translator which creates commands for the robot system to replicate the observed task. In this paper we describe how these components work with special emphasis on the observation system. The robot system that we use to perform the grasping tasks comprises the PUMA 560 arm and the Utah/MIT hand.> Sing Bing Kang, Katsushi Ikeuchi |
ICCV | 1 |
| 1995 | A Multibaseline Stereo System with Active Illumination and Real-Time Image AcquisitionabstractWe describe our implementation of a parallel depth recovery scheme for a four-camera multibaseline stereo in a convergent configuration. Our system is capable of image capture at video rate. This is critical in applications that require three-dimensional tracking. We obtain dense stereo depth data by projecting a light pattern of frequency modulated sinusoidally varying intensity onto the scene, thus increasing the local discriminability at each pixel and facilitating matches. In addition, we make most of the camera view areas by converging them at a volume of interest. Results show that we are able to extract stereo depth data that are, on the average, less than 1 mm in error at distances between 1.5 to 3.5 m away from the cameras.> Sing Bing Kang, Jon A. Webb, C. Lawrence Zitnick, Takeo Kanade |
ICCV | 1 |
| 1995 | Toward automatic robot instruction from perception-temporal segmentation of tasks from human hand motionabstractThis paper describes work on the temporal segmentation of grasping task sequences based on human hand motion. The segmentation process results in the identification of motion breakpoints separating the different constituent phases of the grasping task. A grasping task is composed of three basic phases: pregrasp phase, static grasp phase, and manipulation phase. We show that by analyzing the fingertip polygon area (which is an indication of the hand preshape) and the speed of hand movement (which is an indication of the hand transportation), we can divide a task into meaningful action segments such as approach object (which corresponds to the pregrasp phase), grasp object, manipulate object, place object, and depart (a special case of the pregrasp phase which signals the termination of the task). We introduce a measure called the volume sweep rate, which is the product of the fingertip polygon area and the hand speed. The profile of this measure is also used in the determination of the task breakpoints.> Sing Bing Kang, Katsushi Ikeuchi |
IEEE Trans. Robotics Autom. | 1 |
| 1994 | Determination of Motion Breakpoints in a Task Sequence from Human Hand MotionabstractThis paper describes the authors' work on the temporal segmentation of grasping task sequences based on human hand motion. The segmentation process results in the identification of motion breakpoints separating the different constituent phases of the grasping task. A grasping task is composed of three basic phases: pregrasp phase, static grasp phase, and manipulation phase. The authors show that by analyzing the fingertip polygon (preshape) area and the speed of hand movement, they can divide a task into meaningful action segments such as approach object, grasp object, manipulate object, place object, and depart. The authors introduce a measure called the volume sweep rate, which is the product of the fingertip polygon area and the hand speed. The profile of this measure is also used in the determination of the task breakpoints. The temporal task segmentation process is important as it serves as a preprocessing step to the characterization of the task phases. Once the breakpoints have been identified, further analyses such as grasp recognition and object motion extraction can then be carried out.> Sing Bing Kang, Katsushi Ikeuchi |
ICRA | 1 |
| 1994 | Grasp Recoguition and Manipulative Motion Characterization from Human Hand Motion SequencesabstractWe are developing a system capable of observing a human performing a task and understanding the task well enough to replicate it. This approach is called Assembly Plan from Observation. In order to replicate the observed task, we have to analyze the entire sequence. This can be done by first segmenting the task sequence into its constituent pre-grasp, grasp, and manipulation phases. This paper describes the different analyses that can be done subsequent to the temporal segmentation. These include human grasp recognition, extraction of object motion, and the spatiofrequency (spectrogram) analysis of the manipulation phase.> Sing Bing Kang, Katsushi Ikeuchi |
ICRA | 1 |
| 1994 | Robot task programming by human demonstration: mapping human grasps to manipulator graspsabstractTo alleviate the problem of overwhelming complexity in grasp synthesis and path planning associated with robot task planning, we adopt the approach of teaching the robot by demonstrating in front of it. A system with this programming technique is able to temporally segment a task into separate and meaningful parts for further individual analysis and recognize the human grasp employed in the task. With such derived information, this system would then map the human grasp to that of the given manipulator plan its trajectory, and proceed to execute the task. This paper describes how grasp mapping can be accomplished in our system. The mapping process essentially comprises three steps. The first step is local functional mapping, in which grasps of functionally equivalent fingers are established. This is followed by gross physical mapping which produces a kinematically feasible manipulator grasp. Finally, by carrying out local grasp adjustment using some task-related criterion, we arrive at a locally optimal manipulator grasp. We describe these steps in detail in this paper and show results of example grasp mappings.> Sing Bing Kang, Katsushi Ikeuchi |
IROS | 1 |
| 1994 | Recovering 3D Shape and Motion from Image Streams Using Nonlinear Least Squares
Richard Szeliski, Sing Bing Kang |
J. Vis. Commun. Image Represent. | 2 |
| 1993 | Recovering 3D shape and motion from image streams using nonlinear least squaresabstractA shape and motion estimation algorithm based on nonlinear least squares applied to the tracks of features through time is presented. While the authors' approach requires iteration, it quickly converges to the desired solution, even in the absence of a priori knowledge about the shape or motion. Important features of the algorithm include its ability to handle partial point tracks and true perspective, its ability to use line segment matches and point matches simultaneously, and its use of an object-centered representation for faster and more accurate structure and motion recovery.> Richard Szeliski, Sing Bing Kang |
CVPR | 2 |
| 1993 | A grasp abstraction hierarchy for recognition of grasping tasks from observationabstractThis work focuses on the abstraction hierarchy for a grasp which has been recognized from low-level hand-object interaction data. Previous work done on grasp classification and recognition is discussed. The proposed abstraction hierarchy is presented with illustrations, implementation issues, as well as experimental results. Issues pertaining to the conceptual analysis of the other aspects of recognizing grasping tasks are also presented. The authors report on the current status of the project and future work. Sing Bing Kang, Katsushi Ikeuchi |
IROS | 1 |
| 1993 | The Complex EGI: A New Representation for 3-D Pose DeterminationabstractThe complex extended Gaussian image (CEGI), a 3D object representation that can be used to determine the pose of an object, is described. In this representation, the weight associated with each outward surface normal is a complex weight. The normal distance of the surface from the predefined origin is encoded as the phase of the weight, whereas the magnitude of the weight is the visible area of the surface. This approach decouples the orientation and translation determination into two distinct least-squares problems. The justification for using such a scheme is twofold: it not only allows the pose of the object to be extracted, but it also distinguishes a convex object from a nonconvex object having the same EGI representation. The CEGI scheme has the advantage of not requiring explicit spatial object-model surface correspondence in determining object orientation and translation. Experiments involving synthetic data of two polyhedral and two smooth objects are presented to illustrate the feasibility of this method.> Sing Bing Kang, Katsushi Ikeuchi |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 1993 | Toward automatic robot instruction from perception-recognizing a grasp from observationabstractDeals with the programming of robots to perform grasping tasks. To do this, the assembly plan from observation (APO) paradigm is adopted, where the key idea is to enable a system to observe a human performing a grasping task, understand it, and perform the task with minimal human intervention. A grasping task is composed of three phases: pregrasp phase, static grasp phase, and manipulation phase. The first step in recognizing a grasping task is identifying the grasp itself. The proposed strategy of identifying the grasp is to map the low-level hand configuration to increasingly more abstract grasp descriptions. To achieve the mapping, a grasp representation is introduced, called the contact web, which is composed of a pattern of effective contact points between the hand and the object. A grasp taxonomy based on the contact web is also proposed as a tool to systematically identify a grasp. The grasp can be described at higher conceptual levels using a certain mapping function that results in an index called the grasp cohesive index. This index can be used to identify the grasp. Results from grasping experiments show that it is possible to distinguish between various types of grasps using the proposed contact web, grasp taxonomy and grasp cohesive index.> Sing Bing Kang, Katsushi Ikeuchi |
IEEE Trans. Robotics Autom. | 1 |
| 1992 | Grasp Recognition Using The Contact WebabstractWe propose an approach to teach robots to ger- form grasping tasks. This approach is based on the Assembly Plan from Observation (APO) paradigm, where the key idea is to enable a system to observe a human performing a grasping task, understand it, and perform the task with minimal human intervention. A grasping task is composed of three phases: pre-grasp phase, static grasp phase, and manipulation phase. The first step in recognizing a grasping task is to identify the grasp itself (within the static grasp phase). We propose to identify the grasp by means of a grasp representation called the contact web which is composed of a pattern of effective contact points between the hand and the object. We also propose a grasp taxonomy based on the contact web to systematically identify a grasp. Results from grasping experiments show that it is possi- ble to distinguish between various types of grasps using the proposed contact web and grasp taxonomy. Sing Bing Kang, Katsushi Ikeuchi |
IROS | 1 |
| 1991 | Determining 3-D object pose using the complex extended Gaussian imageabstractA method based on the extended Gaussian image (EGI) which can be used to determine the pose of a 3-D object is presented. In this scheme, the weight associated with each outward surface normal is a complex weight. The normal distance of the surface from the predefined origin is encoded as the phase of the weight, while the magnitude of the weight is the visible area of the surface. This approach decouples the orientation and translation determination into two distinct least-squares problems. Experiments involving synthetic data of two polyhedral and two smooth objects as well as real range data of the same smooth objects indicate the feasibility of this method.> Sing Bing Kang, Katsushi Ikeuchi |
CVPR | 1 |