Kristin J. Dana

dblp:36/4636 · DBLP profile ↗
← Back
60ranked-venue papers
7as first author
15since 2021 · last 2025
0000-0002-2356-6950ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 33 · 6 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 33 · 5 first-author · 8 since 2021Computer networks · 9 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021
YearPublicationVenuePosition
2025 A Landmark-Aware Visual Navigation Dataset for Map Representation Learning
abstract
Map representations learned by expert demonstrations have shown promising research value. However, the field of visual navigation still faces challenges due to the lack of real-world human-navigation datasets that can support efficient, supervised, representation learning of environments. We present a Landmark-Aware Visual Navigation (LAVN) dataset to allow for supervised learning of human-centric exploration policies and map building. We collect RGBD observation and human point-click pairs as a human annotator explores virtual and real-world environments with the goal of full coverage exploration of the space. The human annotators also provide distinct landmark examples along each trajectory, which we intuit will simplify the task of map or graph building and localization. These human point-clicks serve as direct supervision for waypoint prediction when learning to explore in environments. Our dataset covers a wide spectrum of scenes, including rooms in indoor environments, as well as walkways outdoors. We releaseour dataset with detailed documentation at https://huggingface.co/datasets/visnavdataset/lavn (DOI: l0.57967/hf/2386) and a plan for long-term preservation.
Faith M. Johnson, Kristin J. Dana, Bryan Bo Cao, Shubham Jain 0003, Ashwin Ashok
HRI2
2025 Agtech Framework for Cranberry-Ripening Analysis Using Vision Foundation Models
abstract
Agricultural domains are being transformed by recent advances in AI and computer vision that support quantitative visual evaluation. Using aerial and ground imaging over a time series, we develop a framework for characterizing the ripening process of cranberry crops, a crucial component for precision agriculture tasks such as comparing crop breeds (high-throughput phenotyping) and detecting disease. Using drone imaging, we capture images from 20 waypoints across multiple bogs, and using ground-based imaging (hand-held camera), we image same bog patch using fixed fiducial markers. Both imaging methods are repeated to gather a multi-week time series spanning the entire growing season. Aerial imaging provides multiple samples to compute a distribution of albedo values. Ground imaging enables tracking of individual berries for a detailed view of berry appearance changes. Using vision transformers (ViT) for feature detection after segmentation, we extract a high dimensional feature descriptor of berry appearance. Interpretability of appearance is critical for plant biologists and cranberry growers to support crop breeding decisions (e.g. comparison of berry varieties from breeding programs). For interpretability, we create a 2D manifold of cranberry appearance by using a UMAP dimensionality reduction on ViT features. This projection enables quantification of ripening paths and a useful metric of ripening rate. We demonstrate the comparison of four cranberry varieties based on our ripening assessments. This work is the first of its kind and has future impact for cranberries and for other crops including wine grapes, olives, blueberries, and maize. Aerial and ground datasets are made publicly available.
Faith M. Johnson, Ryan Meegan, Jack Lowry, Peter Oudemans, Kristin J. Dana
WACV5
2025 Editorial: Introduction to the Special Section on Best of CVPR'2022
Kristin J. Dana, Gang Hua 0001, Stefan Roth 0001, Dimitris Samaras
IEEE Trans. Pattern Anal. Mach. Intell.1
2024 WATCH: Wide-Area Terrestrial Change Hypercube
abstract
Monitoring Earth activity using data collected from multiple satellite imaging platforms in a unified way is a significant challenge, especially with large variability in image resolution, spectral bands, and revisit rates. Further, the availability of sensor data varies across time as new platforms are launched. In this work, we introduce an adaptable framework and network architecture capable of predicting on subsets of the available platforms, bands, or temporal ranges it was trained on. Our system, called WATCH, is highly general and can be applied to a variety of geospatial tasks. In this work, we analyze the performance of WATCH using the recent IARPA SMART public dataset and metrics. We focus primarily on the problem of broad area search for heavy construction sites. Experiments validate the robustness of WATCH during inference to limited sensor availability, as well the the ability to alter inference-time spatial or temporal sampling. WATCH is open source and available for use on this or other remote sensing problems. Code and model weights are available at: https://gitlab.kitware.com/computer-vision/geowatch
Connor Greenwell, Jon Crall, Matthew Purri, Kristin J. Dana, Nathan Jacobs, Armin Hadzic, Scott Workman, Matthew J. Leotta
WACV4
2023 Learning a Pedestrian Social Behavior Dictionary
Faith M. Johnson, Kristin J. Dana
BMVC2
2023 Self-Supervised Object Detection from Egocentric Videos
abstract
Understanding the visual world from human perspectives has been a long-standing challenge in computer vision. Egocentric videos exhibit high scene complexity and irregular motion flows compared to typical video understanding tasks. With the egocentric domain in mind, we address the problem of self-supervised, class-agnostic object detection, aiming to locate all objects in a given view, without any annotations or pre-trained weights. Our method, self-supervised object detection from egocentric videos (DEVI), generalizes appearance-based methods to learn features end-to-end that are category-specific and invariant to viewing angle and illumination. Our approach leverages natural human behavior in egocentric perception to sample diverse views of objects for our multi-view and scale-regression losses, and our cluster residual module learns multi-category patches for complex scene understanding. DEVI results in gains up to 4.11% AP50, 0.11% AR1, 1.32% AR10, and 5.03% AR100on recent egocentric datasets, while significantly reducing model complexity. We also demonstrate competitive performance on out-of-domain datasets without additional training or fine-tuning.
Peri Akiva, Jing Huang 0020, Kevin J. Liang, Rama Kovvuri, Matt Feiszli, Kristin J. Dana, Tal Hassner
ICCV7
2023 Single Stage Weakly Supervised Semantic Segmentation of Complex Scenes
abstract
The costly process of obtaining semantic segmentation labels has driven research towards weakly supervised semantic segmentation (WSSS) methods, using only image-level, point, or box labels. Such annotations introduce limitations and challenges that results in overly-tuned methods specialized in specific domains or scene types. The over reliance of image-level based methods on generation of high quality class activation maps (CAMs) results in limited applicable dataset complexity range, mostly focusing on object centric scenes. Additionally, the lack of dense annotations requires methods to increase network complexity to obtain additional semantic information, often done through multiple stages of training and refinement. Here, we present a single-stage approach generalizable to a wide range of dataset complexities, that is trainable from scratch, without any dependency on pre-trained backbones, classification, or separate refinement tasks. We utilize point annotations to generate reliable, on-the-fly pseudo-masks through refined and spatially filtered features. We are to demonstrate SOTA performance on benchmark datasets (PascalVOC 2012), as well as significantly outperform other SOTA WSSS methods on recent real-world datasets (CRAID, CityPersons, IAD, ADE20K, CityScapes) with up to 28.1% and 22.6% performance boosts compared to our single-stage and multi-stage baselines respectively.
Peri Akiva, Kristin J. Dana
WACV2
2022 Finch: A Remote Sensing Dataset Management Framework
abstract
This paper describes an open-source dataset framework de-signed for machine learning remote sensing datasets called Finch after the finches Charles Darwin studied while devel-oping the theory of evolution. The Finch Framework enables datasets to be created from annotations-to-images and to ef-ficiently modify an existing dataset. Commonly, datasets are released once and never updated to become more challenging or to correct mistakes. Our proposed framework encourages datasets to evolve over time and provides a version control system for consistent comparisons between experiments. We utilize other open-source tools and unique attributes of re-mote sensing data to characterize entire datasets with a few lightweight configuration files. Generated datasets are ac-companied with a dataset manifest file, providing a consistent and powerful format for referencing all data. The standard dataset representation reduces the time required to integrate datasets into machine learning pipelines.
Matthew Purri, Kristin J. Dana
IGARSS2
2022 Vi-Fi: Associating Moving Subjects across Vision and Wireless Sensors
abstract
In this paper, we present Vi-Fi, a multi-modal system that leverages a user's smartphone WiFi Fine Timing Measurements (FTM) and inertial measurement unit (IMU) sensor data to associate the user detected on a camera footage with their corresponding smartphone identifier (e.g. WiFi MAC address). Our approach uses a recurrent multi-modal deep neural network that exploits FTM and IMU measurements along with distance between user and camera (depth information) to learn affinity matrices. As a baseline method for comparison, we also present a traditional non deep learning approach that uses bipartite graph matching. To facilitate evaluation, we collected a multi-modal dataset that comprises camera videos with depth information (RGB-D), WiFi FTM and IMU measurements for multiple participants at diverse real-world settings. Using association accuracy as the key metric for evaluating the fidelity of Vi-Fi in associating human users on camera feed with their phone IDs, we show that Vi-Fi achieves between 81% (real-time) to 91% (offline) association accuracy.
Hansi Liu, Abrar Alali, Mohamed Ibrahim Ahmed 0001, Bryan Bo Cao, Nicholas Meegan, Marco Gruteser, Shubham Jain 0003, Kristin J. Dana, Ashwin Ashok, Bin Cheng 0002, Hongsheng Lu
IPSN9
2022 ViTag: Online WiFi Fine Time Measurements Aided Vision-Motion Identity Association in Multi-person Environments
abstract
In this paper, we present ViTag to associate user identities across multimodal data, particularly those obtained from cameras and smartphones. ViTag associates a sequence of vision tracker generated bounding boxes with Inertial Mea-surement Unit (IMU) data and Wi-Fi Fine Time Measurements (FTM) from smartphones. We formulate the problem as association by sequence to sequence (seq2seq) translation. In this two-step process, our system first performs cross-modal translation using a multimodal LSTM encoder-decoder network (X-Translator) that translates one modality to another, e.g. recon-structing IMU and FTM readings purely from camera bounding boxes. Second, an association module finds identity matches between camera and phone domains, where the translated modality is then matched with the observed data from the same modality. In contrast to existing works, our proposed approach can associate identities in multi-person scenarios where all users may be performing the same activity. Extensive experiments in real-world indoor and outdoor environments demonstrate that online association on camera and phone data (IMU and FTM) achieves an average Identity Precision Accuracy (IDP) of 88.39% on a 1 to 3 seconds window, outperforming the state-of-the-art Vi-Fi (82.93%). Further study on modalities within the phone domain shows the FTM can improve association performance by 12.56% on average. Finally, results from our sensitivity experiments demonstrate the robustness of ViTag under different noise and environment variations.
Bryan Bo Cao, Abrar Alali, Hansi Liu, Nicholas Meegan, Marco Gruteser, Kristin J. Dana, Ashwin Ashok, Shubham Jain 0003
SECON6
2022 Differential Viewpoints for Ground Terrain Material Recognition
abstract
Computational surface modeling that underlies material recognition has transitioned from reflectance modeling using in-lab controlled radiometric measurements to image-based representations based on internet-mined single-view images captured in the scene. We take a middle-ground approach for material recognition that takes advantage of both rich radiometric cues and flexible image capture. A key concept is differential angular imaging, where small angular variations in image capture enables angular-gradient features for an enhanced appearance representation that improves recognition. We build a large-scale material database, Ground Terrain in Outdoor Scenes (GTOS) database, to support ground terrain recognition for applications such as autonomous driving and robot navigation. The database consists of over 30,000 images covering 40 classes of outdoor ground terrain under varying weather and lighting conditions. We develop a novel approach for material recognition called texture-encoded angular network (TEAN) that combines deep encoding pooling of RGB information and differential angular images for angular-gradient features to fully leverage this large dataset. With this novel network architecture, we extract characteristics of materials encoded in the angular and spatial gradients of their appearance. Our results show that TEAN achieves recognition performance that surpasses single view performance and standard (non-differential/large-angle sampling) multiview performance.
Jia Xue, Hang Zhang 0005, Ko Nishino, Kristin J. Dana
IEEE Trans. Pattern Anal. Mach. Intell.4
2021 Shape From Sky: Polarimetric Normal Recovery Under the Sky
abstract
The sky exhibits a unique spatial polarization pattern by scattering the unpolarized sun light. Just like insects use this unique angular pattern to navigate, we use it to map pixels to directions on the sky. That is, we show that the unique polarization pattern encoded in the polarimetric appearance of an object captured under the sky can be decoded to reveal the surface normal at each pixel. We derive a polarimetric reflection model of a diffuse plus mirror surface lit by the sun and a clear sky. This model is used to recover the per-pixel surface normal of an object from a single polarimetric image or from multiple polarimetric images captured under the sky at different times of the day. We experimentally evaluate the accuracy of our shape-from-sky method on a number of real objects of different surface compositions. The results clearly show that this passive approach to fine-geometry recovery that fully leverages the unique illumination made by nature is a viable option for 3D sensing. With the advent of quad-Bayer polarization chips, we believe the implications of our method span a wide range of domains.
Tomoki Ichikawa, Matthew Purri, Ryo Kawahara, Shohei Nobuhara, Kristin J. Dana, Ko Nishino
CVPR5
2021 Lost and Found!: associating target persons in camera surveillance footage with smartphone identifiers
abstract
We demonstrate an application of finding target persons on a surveillance video. Each visually detected participant is tagged with a smartphone ID and the target person with the query ID is highlighted. This work is motivated by the fact that establishing associations between subjects observed in camera images and messages transmitted from their wireless devices can enable fast and reliable tagging. This is particularly helpful when target pedestrians need to be found on public surveillance footage, without the reliance on facial recognition. The underlying system uses a multi-modal approach that leverages WiFi Fine Timing Measurements (FTM) and inertial sensor (IMU) data to associate each visually detected individual with a corresponding smartphone identifier. These smartphone measurements are combined strategically with RGB-D information from the camera, to learn affinity matrices using a multi-modal deep learning network.
Hansi Liu, Abrar Alali, Mohamed Ibrahim Ahmed 0001, Marco Gruteser, Shubham Jain 0003, Kristin J. Dana, Ashwin Ashok, Bin Cheng 0002, Hongsheng Lu
MobiSys7
2021 H2O-Net: Self-Supervised Flood Segmentation via Adversarial Domain Adaptation and Label Refinement
abstract
Accurate flood detection in near real time via high resolution, high latency satellite imagery is essential to prevent loss of lives by providing quick and actionable information. Instruments and sensors useful for flood detection are only available in low resolution, low latency satellites with region re-visit periods of up to 16 days, making flood alerting systems that use such satellites unreliable. This work presents H2O-Network, a self-supervised deep learning method to segment floods from satellites and aerial imagery by bridging domain gap between low and high latency satellite and coarse-to-fine label refinement. H2O-Net learns to synthesize signals highly correlative with water presence as a domain adaptation step for semantic segmentation in high resolution satellite imagery. Our work also proposes a self-supervision mechanism, which does not require any hand annotation, used during training to generate high quality ground truth data. We demonstrate that H2O-Net outperforms the state-of-the-art semantic segmentation methods on satellite imagery by 10% and 12% pixel accuracy and mIoU respectively for the task of flood segmentation. We emphasize the generalizability of our model by transferring model weights trained on satellite imagery to drone imagery, a highly different sensor and domain.
Peri Akiva, Matthew Purri, Kristin J. Dana, Beth Tellman, Tyler Anderson
WACV3
2021 AI on the Bog: Monitoring and Evaluating Cranberry Crop Risk
abstract
Machine vision for precision agriculture has attracted considerable research interest in recent years. The goal of this paper is to develop an end-to-end cranberry health monitoring system to enable and support real time cranberry over-heating assessment to facilitate informed decisions that may sustain the economic viability of the farm. Toward this goal, we propose two main deep learning-based modules for: 1) cranberry fruit segmentation to delineate the exact fruit regions in the cranberry field image that are exposed to sun, 2) prediction of cloud coverage conditions and sun irradiance to estimate the inner temperature of exposed cranberries. We develop drone-based field data and ground-based sky data collection systems to collect video imagery at multiple time points for use in crop health analysis. Extensive evaluation on the data set shows that it is possible to predict exposed fruit's inner temperature with high accuracy (0.02% MAPE). The sun irradiance prediction error was found to be 8.41-20.36% MAPE in the 5-20 minutes time horizon. With 62.54% mIoU for segmentation and 13.46 MAE for counting accuracies in exposed fruit identification, this system is capable of giving informed feedback to growers to take precautionary action (e.g., irrigation) in identified crop field regions with higher risk of sunburn in the near future. Though this novel system is applied for cranberry health monitoring, it represents a pioneering step forward for efficient farming and is useful in precision agriculture beyond the problem of cranberry overheating.
Peri Akiva, Benjamin Planche, Kristin J. Dana, Peter Oudemans, Michael Mars
WACV4
2020 Teaching Cameras to Feel: Estimating Tactile Physical Properties of Surfaces from Images
Matthew Purri, Kristin J. Dana
ECCV (27)2
2020 Angular Luminance for Material Segmentation
abstract
Moving cameras provide multiple intensity measurements per pixel, yet often semantic segmentation, material recognition, and object recognition do not utilize this information. With basic alignment over several frames of a moving camera sequence, a distribution of intensities over multiple angles is obtained. It is well known from prior work that luminance histograms and the statistics of natural images provide a strong material recognition cue. We utilize per-pixel angular luminance distributions as a key feature in discriminating the material of the surface. The angle-space sampling in a multiview satellite image sequence is an unstructured sampling of the underlying reflectance function of the material. For real-world materials there is significant intra-class variation that can be managed by building a angular luminance network (AngLNet). This network combines angular reflectance cues from multiple images with spatial cues as input to fully convolutional networks for material segmentation. We demonstrate the increased performance of AngLNet over prior state-of-the-art in material segmentation from satellite imagery.
Jia Xue, Matthew Purri, Kristin J. Dana
IGARSS3
2019 Light Field Messaging With Deep Photographic Steganography
abstract
We develop Light Field Messaging (LFM), a process of embedding, transmitting, and receiving hidden information in video that is displayed on a screen and captured by a handheld camera. The goal of the system is to minimize perceived visual artifacts of the message embedding, while simultaneously maximizing the accuracy of message recovery on the camera side. LFM requires photographic steganography for embedding messages that can be displayed and camera-captured. Unlike digital steganography, the embedding requirements are significantly more challenging due to the combined effect of the screen's radiometric emittance function, the camera's sensitivity function, and the camera-display relative geometry. We devise and train a network to jointly learn a deep embedding and recovery algorithm that requires no multi-frame synchronization. A key novel component is the camera display transfer function (CDTF) to model the camera-display pipeline. To learn this CDTF we introduce a dataset (Camera-Display 1M) of 1,000,000 camera-captured images collected from 25 camera-display pairs. The result of this work is a high-performance real-time LFM system using consumer-grade displays and smartphone cameras.
Eric Wengrowski, Kristin J. Dana
CVPR2
2019 Photo-Realistic Facial Texture Transfer
abstract
Style transfer methods have achieved significant success in recent years with the use of convolutional neural networks. However, many of these methods concentrate on artistic style transfer with few constraints on the output image appearance. We address the challenging problem of transferring face texture from a style face image to a content face image in a photorealistic manner without changing the identity of the original content image. Our framework for face texture transfer (FaceTex) augments the prior work of MRF-CNN with a novel facial semantic regularization that incorporates a face prior regularization smoothly suppressing the changes around facial meso-structures (e.g eyes, nose and mouth) and a facial structure loss function which implicitly preserves the facial structure so that face texture can be transferred without changing the original identity. We demonstrate results on face images and compare our approach with recent state-of-the-art methods. Our results demonstrate superior texture transfer because of the ability to maintain the identity of the original face image.
Hang Zhang 0005, Kristin J. Dana
WACV3
2018 Context Encoding for Semantic Segmentation
abstract
Recent work has made significant progress in improving spatial resolution for pixelwise labeling with Fully Convolutional Network (FCN) framework by employing Dilated/Atrous convolution, utilizing multi-scale features and refining boundaries. In this paper, we explore the impact of global contextual information in semantic segmentation by introducing the Context Encoding Module, which captures the semantic context of scenes and selectively highlights class-dependent featuremaps. The proposed Context Encoding Module significantly improves semantic segmentation results with only marginal extra computation cost over FCN. Our approach has achieved new state-of-the-art results 51.7% mIoU on PASCAL-Context, 85.9% mIoU on PASCAL VOC 2012. Our single model achieves a final score of 0.5567 on ADE20K test set, which surpasses the winning entry of COCO-Place Challenge 2017. In addition, we also explore how the Context Encoding Module can improve the feature representation of relatively shallow networks for the image classification on CIFAR-10 dataset. Our 14 layer network has achieved an error rate of 3.45%, which is comparable with state-of-the-art approaches with over 10Ã- more layers. The source code for the complete system are publicly available1.
Hang Zhang 0005, Kristin J. Dana, Jianping Shi, Xiaogang Wang 0001, Ambrish Tyagi, Amit Agrawal 0002
CVPR2
2018 Deep Texture Manifold for Ground Terrain Recognition
abstract
We present a texture network called Deep Encoding Pooling Network (DEP) for the task of ground terrain recognition. Recognition of ground terrain is an important task in establishing robot or vehicular control parameters, as well as for localization within an outdoor environment. The architecture of DEP integrates orderless texture details and local spatial information and the performance of DEP surpasses state-of-the-art methods for this task. The GTOS database (comprised of over 30,000 images of 40 classes of ground terrain in outdoor scenes) enables supervised recognition. For evaluation under realistic conditions, we use test images that are not from the existing GTOS dataset, but are instead from hand-held mobile phone videos of similar terrain. This new evaluation dataset, GTOS-mobile, consists of 81 videos of 31 classes of ground terrain such as grass, gravel, asphalt and sand. The resultant network shows excellent performance not only for GTOS-mobile, but also for more general databases (MINC and DTD). Leveraging the discriminant features learned from this network, we build a new texture manifold called DEP-manifold. We learn a parametric distribution in feature space in a fully supervised manner, which gives the distance relationship among classes and provides a means to implicitly represent ambiguous class boundaries. The source code and database are publicly available1.
Jia Xue, Hang Zhang 0005, Kristin J. Dana
CVPR3
2017 Differential Angular Imaging for Material Recognition
abstract
Material recognition for real-world outdoor surfaces has become increasingly important for computer vision to support its operation in the wild. Computational surface modeling that underlies material recognition has transitioned from reflectance modeling using in-lab controlled radiometric measurements to image-based representations based on internet-mined images of materials captured in the scene. We propose to take a middle-ground approach for material recognition that takes advantage of both rich radiometric cues and flexible image capture. We realize this by developing a framework for differential angular imaging, where small angular variations in image capture provide an enhanced appearance representation and significant recognition improvement. We build a large-scale material database, Ground Terrain in Outdoor Scenes (GTOS) database, geared towards real use for autonomous agents. The database consists of over 30,000 images covering 40 classes of outdoor ground terrain under varying weather and lighting conditions. We develop a novel approach for material recognition called a Differential Angular Imaging Network (DAIN) to fully leverage this large dataset. With this novel network architecture, we extract characteristics of materials encoded in the angular and spatial gradients of their appearance. Our results show that DAIN achieves recognition performance that surpasses single view or coarsely quantized multiview images. These results demonstrate the effectiveness of differential angular imaging as a means for flexible, in-place material recognition.
Jia Xue, Hang Zhang 0005, Kristin J. Dana, Ko Nishino
CVPR3
2017 Deep TEN: Texture Encoding Network
abstract
We propose a Deep Texture Encoding Network (Deep-TEN) with a novel Encoding Layer integrated on top of convolutional layers, which ports the entire dictionary learning and encoding pipeline into a single model. Current methods build from distinct components, using standard encoders with separate off-the-shelf features such as SIFT descriptors or pre-trained CNN features for material recognition. Our new approach provides an end-to-end learning framework, where the inherent visual vocabularies are learned directly from the loss function. The features, dictionaries, encoding representation and the classifier are all learned simultaneously. The representation is orderless and therefore is particularly useful for material and texture recognition. The Encoding Layer generalizes robust residual encoders such as VLAD and Fisher Vectors, and has the property of discarding domain specific information which makes the learned convolutional features easier to transfer. Additionally, joint training using multiple datasets of varied sizes and class labels is supported resulting in increased recognition performance. The experimental results show superior performance as compared to state-of-the-art methods using gold-standard databases such as MINC-2500, Flickr Material Database, KTH-TIPS-2b, and two recent databases 4D-Light-Field-Material and GTOS. The source code for the complete system are publicly available1.
Hang Zhang 0005, Jia Xue, Kristin J. Dana
CVPR3
2017 Reading between the pixels: Photographic steganography for camera display messaging
abstract
We exploit human color metamers to send light-modulated messages decipherable by cameras, but camouflaged to human vision. These time-varying messages are concealed in ordinary images and videos. Unlike previous methods which rely on visually obtrusive intensity modulation, embedding with color reduces visible artifacts. The mismatch in human and camera spectral sensitivity creates a unique opportunity for hidden messaging. Each color pixel in an electronic display image is modified by shifting the base color along a particular color gradient. The challenge is to find the set of color gradients that maximizes camera response and minimizes human response. Our approach does not require a priori measurement of these sensitivity curves. We learn an ellipsoidal partitioning of the 6-dimensional space of base colors and color gradients. This partitioning creates metamer sets defined by the base color of each display pixel and the corresponding color gradient for message encoding. We sample from the learned metamer sets to find optimal color steps for arbitrary base colors. Ordinary displays and cameras are used, so there is no need for high speed cameras or displays. Our primary contribution is a method to map pixels in an arbitrary image to metamer pairs for steganographic camera-display messaging.
Eric Wengrowski, Kristin J. Dana, Marco Gruteser, Narayan B. Mandayam
ICCP2
2016 Friction from Reflectance: Deep Reflectance Codes for Predicting Physical Surface Properties from One-Shot In-Field Reflectance
Hang Zhang 0005, Kristin J. Dana, Ko Nishino
ECCV (4)2
2016 Hybrid deep learning for Reflectance Confocal Microscopy skin images
abstract
Reflectance Confocal Microscopy (RCM) is used for evaluation of human skin disorders and the effects of skin treatments by imaging the skin layers at different depths. Traditionally, clinical experts manually categorize the images captured into different skin layers. This time-consuming labeling task impedes the convenient analysis of skin image datasets. In recent automated image recognition tasks, deep learning with convolutional neural nets (CNN) has achieved remarkable results. However in many clinical settings, training data is often limited and insufficient for CNN training. For recognition of RCM skin images, we demonstrate that a CNN trained on a moderate size dataset leads to low accuracy. We introduce a hybrid deep learning approach which uses traditional texton-based feature vectors as input to train a deep neural network. This hybrid method uses fixed filters in the input layer instead of tuned filters, yet superior performance is achieved. Our dataset consists of 1500 images from 15 RCM stacks belonging to six different categories of skin layers. We show that our hybrid deep learning approach performs with a test accuracy of 82% compared with 51% for CNN. We also compare the results with additional proposed methods for RCM image recognition and show improved accuracy.
Kristin J. Dana, Gabriela Oana Cula, M. Catherine Mack
ICPR2
2016 High-rate flicker-free screen-camera communication with spatially adaptive embedding
abstract
Embedded screen-camera communication techniques encode information in screen imagery that can be decoded with a camera receiver yet remains unobtrusive to the human observer. These techniques have applications in tagging content on screens similar to QR-code tagging for other objects. This paper characterizes the design space for flicker-free embedded screen-camera communication. In particular, we identify an orthogonal dimension to prior work: spatial content-adaptive encoding, and observe that it is essential to combine multiple dimensions to achieve both high capacity and minimal flicker. From these insights, we develop content-adaptive encoding techniques that exploit visual features such as edges and texture to unobtrusively communicate information. These can then be layered over existing techniques to further boost the capacity. Our experimental results show that there is potential to achieve an average goodput of about 22 kbps, significantly outperforming existing work while remaining flicker-free.
Viet Nguyen, Yaqin Tang, Ashwin Ashok, Marco Gruteser, Kristin J. Dana, Eric Wengrowski, Narayan B. Mandayam
INFOCOM5
2016 Optimal radiometric calibration for camera-display communication
abstract
We present a novel method for communicating between a moving camera and an electronic display by embedding and recovering hidden, dynamic information within an image. A small intensity pattern is added to alternate frames of a time-varying display. A handheld camera pointed at the display can receive not only the display image, but also an underlying message. Differencing the camera-captured alternate frames leaves the small intensity pattern, but results in errors due to photometric effects that depend on camera pose. Detecting and robustly decoding the message requires careful photometric modeling for message recovery. The key innovation of our approach is an algorithm that performs simultaneous radiometric calibration and message recovery in one convex optimization problem. By modeling the photometry of the system using a camera-display transfer function (CDTF), we derive an optimal online radiometric calibration (OORC) for robust computational messaging as demonstrated with nine different commercial cameras and displays. The online radiometric calibration algorithms described in this paper significantly reduces message recovery errors, especially for low intensity messages and oblique camera angles.
Eric Wengrowski, Wenjia Yuan, Kristin J. Dana, Ashwin Ashok, Marco Gruteser, Narayan B. Mandayam
WACV3
2016 Modified balanced iterative reducing and clustering using hierarchies (m-BIRCH) for visual clustering
Siddharth K. Madan, Kristin J. Dana
Pattern Anal. Appl.2
2016 Automated Crack Detection on Concrete Bridges
abstract
Detection of cracks on bridge decks is a vital task for maintaining the structural health and reliability of concrete bridges. Robotic imaging can be used to obtain bridge surface image sets for automated on-site analysis. We present a novel automated crack detection algorithm, the STRUM (spatially tuned robust multifeature) classifier, and demonstrate results on real bridge data using a state-of-the-art robotic bridge scanning system. By using machine learning classification, we eliminate the need for manually tuning threshold parameters. The algorithm uses robust curve fitting to spatially localize potential crack regions even in the presence of noise. Multiple visual features that are spatially tuned to these regions are computed. Feature computation includes examining the scale-space of the local feature in order to represent the information and the unknown salient scale of the crack. The classification results are obtained with real bridge data from hundreds of crack regions over two bridges. This comprehensive analysis shows a peak STRUM classifier performance of 95% compared with 69% accuracy from a more typical image-based approach. In order to create a composite global view of a large bridge span, an image sequence from the robot is aligned computationally to create a continuous mosaic. A crack density map for the bridge mosaic provides a computational description as well as a global view of the spatial patterns of bridge deck cracking. The bridges surveyed for data collection and testing include Long-Term Bridge Performance program's (LTBP) pilot project bridges at Haymarket, VA, USA, and Sacramento, CA, USA.
Prateek Prasanna, Kristin J. Dana, Nenad Gucunski, Basily Basily, Hung Manh La, Ronny Salim Lim, Hooman Parvardeh
IEEE Trans Autom. Sci. Eng.2
2016 Automated GPR Rebar Analysis for Robotic Bridge Deck Evaluation
abstract
Ground penetrating radar (GPR) is used to evaluate deterioration of reinforced concrete bridge decks based on measuring signal attenuation from embedded rebar. The existing methods for obtaining deterioration maps from GPR data often require manual interaction and offsite processing. In this paper, a novel algorithm is presented for automated rebar detection and analysis. We test the process with comprehensive measurements obtained using a novel state-of-the-art robotic bridge inspection system equipped with GPR sensors. The algorithm achieves robust performance by integrating machine learning classification using image-based gradient features and robust curve fitting of the rebar hyperbolic signature. The approach avoids edge detection, thresholding, and template matching that require manual tuning and are known to perform poorly in the presence of noise and outliers. The detected hyperbolic signatures of rebars within the bridge deck are used to generate deterioration maps of the bridge deck. The results of the rebar region detector are compared quantitatively with several methods of image-based classification and a significant performance advantage is demonstrated. High rates of accuracy are reported on real data that includes thousands of individual hyperbolic rebar signatures from three real bridge decks.
Kristin J. Dana, Francisco A. Romero, Nenad Gucunski
IEEE Trans. Cybern.2
2016 What Am I Looking At? Low-Power Radio-Optical Beacons for In-View Recognition on Smart-Glass
abstract
Applications on wearable personal imaging devices, or Smart-glasses as they are called, can largely benefit from accurate and energy-efficient recognition of objects that are within the user's view. Existing solutions such as optical or computer vision approaches are too energy intensive, while low-power active radio tags suffer from imprecise orientation estimates. To address this challenge, this paper presents the design, implementation, and evaluation of a radio-optical hybrid system where a radio-optical transmitter, or tag, whose radio-optical beacons are used for accurate relative orientation tracking of tagged objects by a wearable radio-optical receiver. A low-power radio link that conveys identity is used to reduce the battery drain by synchronizing the radio-optical transmitter and receiver so that extremely short optical (infrared) pulses are sufficient for orientation (angle and distance) estimation. Through extensive experiments with our prototype we show that our system can achieve orientation estimates with 1-to-2 degree accuracy and within 40 cm ranging error, with a maximum range of 9 m in typical indoor use cases. With a tag and receiver battery power consumption of 81 μW and 90 mW, respectively, our radio-optical tags and receiver are at least 1.5x energy efficient than prior works in this space.
Ashwin Ashok, Chenren Xu, Tam Vu 0001, Marco Gruteser, Richard E. Howard, Yanyong Zhang, Narayan B. Mandayam, Wenjia Yuan, Kristin J. Dana
IEEE Trans. Mob. Comput.9
2015 Reflectance hashing for material recognition
abstract
We introduce a novel method for using reflectance to identify materials. Reflectance offers a unique signature of the material but is challenging to measure and use for recognizing materials due to its high-dimensionality. In this work, one-shot reflectance of a material surface which we refer to as a reflectance disk is capturing using a unique optical camera. The pixel coordinates of these reflectance disks correspond to the surface viewing angles. The reflectance has class-specific stucture and angular gradients computed in this reflectance space reveal the material class. These reflectance disks encode discriminative information for efficient and accurate material recognition. We introduce a framework called reflectance hashing that models the reflectance disks with dictionary learning and binary hashing. We demonstrate the effectiveness of reflectance hashing for material recognition with a number of real-world materials.
Hang Zhang 0005, Kristin J. Dana, Ko Nishino
CVPR2
2015 Low-Power Radio-Optical Beacons for In-View Recognition
abstract
Object recognition on wearable devices using computer vision is too energy intensive and challenging when objects are similar looking, while low-power active radio frequency identification (RFID) systems suffer from imprecise orientation (angle and distance) estimates. To address this challenge, this paper presents a novel radio-optical based recognition system where a radio-optical transmitter, or tag, that emits a beacon whose infra-red (IR) signal strength is used for accurate relative orientation tracking of tagged objects at a wearable radio-optical receiver. A low-power radio link that conveys identity is used to reduce the battery drain by synchronizing the radio- optical transmitter and receiver so that extremely short optical pulses are sufficient for precise orientation estimation. Through extensive experiments with our prototype we show that our system can achieve orientation estimates with 1-2° accuracy and within 40cm ranging error, with a maximum range of 9m in typical indoor use cases. With a tag battery power consumption of 86μW, the radio-optical tags show potential to achieve about half a decade lifetimes.
Ashwin Ashok, Chenren Xu, Tam Vu 0001, Marco Gruteser, Richard E. Howard, Yanyong Zhang, Narayan B. Mandayam, Wenjia Yuan, Kristin J. Dana
VTC Fall9
2014 Phase messaging method for time-of-flight cameras
abstract
Ubiquitous light emitting devices and low-cost commercial digital cameras facilitate optical wireless communication system such as visual MIMO where handheld cameras communicate with electronic displays. While intensity-based optical communications are more prevalent in camera-display messaging, we present a novel method that uses modulated light phase for messaging and time-of-flight (ToF) cameras for receivers. With intensity-based methods, light signals can be degraded by reflections and ambient illumination. By comparison, communication using ToF cameras is more robust against challenging lighting conditions. Additionally, the concept of phase messaging can be combined with intensity messaging for a significant data rate advantage. In this work, we design and construct a phase messaging array (PMA), which is the first of its kind, to communicate to a ToF depth camera by manipulating the phase of the depth camera's infrared light signal. The array enables message variation spatially using a plane of infrared light emitting diodes and temporally by varying the induced phase shift. In this manner, the phase messaging array acts as the transmitter by electronically controlling the light signal phase. The ToF camera acts as the receiver by observing and recording a time-varying depth. We show a complete implementation of a 3×3 prototype array with custom hardware and demonstrating average bit accuracy as high as 97.8%. The prototype data rate with this approach is 1 Kbps that can be extended to approximately 10 Mbps.
Wenjia Yuan, Richard E. Howard, Kristin J. Dana, Ramesh Raskar, Ashwin Ashok, Marco Gruteser, Narayan B. Mandayam
ICCP3
2014 Capacity of pervasive camera based communication under perspective distortions
abstract
Cameras are ubiquitous and increasingly being used not just for capturing images but also for communicating information. For example, the pervasive QR codes can be viewed as communicating a short code to camera-equipped sensors and recent research has explored using screen-to-camera communications for larger data transfers. Such communications could be particularly attractive in pervasive camera based applications, where such camera communications can reuse the existing camera hardware and also leverage from the large pixel array structure for high data-rate communication. While several prototypes have been constructed, the fundamental capacity limits of this novel communication channel in all but the simplest scenarios remains unknown. The visual medium differs from RF in that the information capacity of this channel largely depends on the perspective distortions while multipath becomes negligible. In this paper, we create a model of this communication system to allow predicting the capacity based on receiver perspective (distance and angle to the transmitter). We calibrate and validate this model through lab experiments wherein information is transmitted from a screen and received with a tablet camera. Our capacity estimates indicate that tens of Mbps is possible using a smartphone camera even when the short code on the screen images onto only 15% of the camera frame. Our estimates also indicate that there is room for at least 2.5x improvement in throughput of existing screen - camera communication prototypes.
Ashwin Ashok, Shubham Jain 0003, Marco Gruteser, Narayan B. Mandayam, Wenjia Yuan, Kristin J. Dana
PerCom6
2013 BiFocus: using radio-optical beacons for an augmented reality search application
abstract
Augmented Reality (AR) applications benefit from accurate detection of the objects that are within a person's view. Typically, it is not only desirable to identify what is currently within view, but also to navigate the users view to the item of interest - for example, finding a misplaced object. In this paper we demonstrate a low-power hybrid radio-optical beaconing system, where objects of interest are tagged with battery-powered RFID-like tags equipped with infrared light emitting diodes (LED) that emit periodic infrared beacons. These beacons are used for accurately estimating the angle and distance from the object to the receiver so as to locate it. The beacons are synchronized using the radio link that is also used to convey the object's unique ID.
Ashwin Ashok, Chenren Xu, Tam Vu 0001, Marco Gruteser, Richard E. Howard, Yanyong Zhang, Narayan B. Mandayam, Wenjia Yuan, Kristin J. Dana
MobiSys9
2012 Dynamic and invisible messaging for visual MIMO
abstract
The growing ubiquity of cameras in hand-held devices and the prevalence of electronic displays in signage creates a novel framework for wireless communications. Traditionally, the term MIMO is used for multiple-input multiple-output where the multiple-input component is a set of radio transmitters and the multiple-output component is a set of radio receivers. We employ the concept of visual MIMO where pixels are transmitters and cameras are receivers. In this manner, the techniques of computer vision can be combined with principles from wireless communications to create an optical line-of-sight communications channel. Two major challenges are addressed: (1) The message for transmission must be embedded in the observed display so that the message is hidden from the observer and the electronic display can simultaneously be used for its originally intended purpose (e.g. signage, advertisements, maps); (2) Photometric and geometric distortions during the imaging process corrupt the information channel between the transmitter display and the receiver camera. These distortions must be modeled and removed. In this paper, we present a real-time messaging paradigm and its implementation in an operational visual MIMO optical systems. As part of the system, we develop a novel algorithm for photographic message extraction which includes automatic display detection, message embedding and message retrieval. Experiments show that the system achieves an average accuracy of 94.6% at the bitrate of 6222.2 bps.
Wenjia Yuan, Kristin J. Dana, Ashwin Ashok, Marco Gruteser, Narayan B. Mandayam
WACV2
2011 Demo: visual MIMO based LED - camera communication applied to automobile safety
abstract
The inherent limitations in RF spectrum availability and susceptibility to interference make it difficult to meet the reliability required for automotive safety applications. To address this challenge, this work explores an alternative communication system called Visual MIMO that uses light emitting arrays as transmitters and cameras as receivers. Visual MIMO applied to vehicular communication proposes to reuse existing LED rear and headlights as transmitters and existing cameras (e.g. those used for parking assistance, rear-view cameras) as receivers. In this work we show a proof of concept based demonstration of the Visual MIMO system consisting of an LED transmitter array and a high-speed camera.
Michael Varga, Ashwin Ashok, Marco Gruteser, Narayan B. Mandayam, Wenjia Yuan, Kristin J. Dana
MobiSys6
2011 Rate adaptation in visual MIMO
abstract
We propose a rate adaptation scheme for visual MIMO camera-based communications, wherein parallel data transmissions from light emitting arrays are received by multiple receive elements of a CCD/CMOS camera image sensor. Unlike RF MIMO, multipath fading is negligible in the visual MIMO channel. Instead, the channel is largely dependent on receiver perspective (distance and angle) and visibility issues (partial line-of-sight availability and occlusions). This allows for slower adaptation but requires the adaptation algorithm to choose among a more complex set of modes. In this paper, we define a set of operating modes for visual MIMO transmitters and propose a rate adaptation scheme to switch between these modes. Our Visual MIMO Rate Adaptation (VMRA) is a packet based rate adaptation protocol that bases its rate selection decisions on the packet error rate feedback. Using trace-based simulation results for a vehicle-to-vehicle communication scenario, we illustrate how our VMRA algorithms can adapt over distance as well as visibility variations in an optical link and achieve a higher average throughput.
Ashwin Ashok, Marco Gruteser, Narayan B. Mandayam, Taekyoung Kwon 0002, Wenjia Yuan, Michael Varga, Kristin J. Dana
SECON7
2010 Challenge: mobile optical networks through visual MIMO
abstract
Mobile optical communications has so far largely been limited to short ranges of about ten meters, since the highly directional nature of optical transmissions would require costly mechanical steering mechanisms. Advances in CCD and CMOS imaging technology along with the advent of visible and infrared (IR) light sources such as (light emitting diode) LED arrays presents an exciting and challenging concept which we call as visual-MIMO (multiple-input multiple-output) where optical transmissions by multiple transmitter elements are received by an array of photodiode elements (e.g. pixels in a camera). Visual-MIMO opens a new vista of research challenges in PHY, MAC and Network layer research and this paper brings together the networking, communications and computer vision fields to discuss the feasibility of this as well as the underlying opportunities and challenges. Example applications range from household/factory robotic to tactical to vehicular networks as well pervasive computing, where RF communications can be interference-limited and prone to eavesdropping and security lapses while the less observable nature of highly directional optical transmissions can be beneficial. The impact of the characteristics of such technologies on the medium access and network layers has so far received little consideration. Example characteristics are a strong reliance on computer vision algorithms for tracking, a form of interference cancellation that allows successfully receiving packets from multiple transmitters simultaneously, and the absence of fast fading but a high susceptibility to outages due to line-of-sight interruptions. These characteristics lead to significant challenges and opportunities for mobile networking research
Ashwin Ashok, Marco Gruteser, Narayan B. Mandayam, Jayant Silva, Michael Varga, Kristin J. Dana
MobiCom6
2007 Surface detail in computer models
Kristin J. Dana, Gabriela Oana Cula, Jing Wang 0008
Image Vis. Comput.1
2007 Polarization Multiplexing and Demultiplexing for Appearance-Based Modeling
abstract
Polarization has been used in numerous prior studies for separating diffuse and specular reflectance components, but in this work we show that it also can be used to separate surface reflectance contributions from individual light sources. Our approach is called polarization multiplexing and it has a significant impact in appearance modeling where the image as a function of illumination direction is needed. Multiple unknown light sources can illuminate the scene simultaneously, and the individual contributions to the overall surface reflectance are estimated. Polarization multiplexing relies on the relationship between the light source direction and the intensity modulation. Inverting this transformation enables the individual intensity contributions to be estimated. In addition to polarization multiplexing, we show that phase histograms from the intensity modulations can be used to estimate scene properties including the number of light sources.
Gabriela Oana Cula, Kristin J. Dana, Dinesh K. Pai
IEEE Trans. Pattern Anal. Mach. Intell.2
2006 Relief Texture from Specularities
abstract
In vision and graphics, advanced object models require not only 3D shape, but also surface detail. While several scanning devices exist to capture the global shape of an object, few methods concentrate on capturing the fine-scale detail. Fine-scale surface geometry (relief texture), such as surface markings, roughness, and imprints, is essential in highly realistic rendering and accurate prediction. We present a novel approach for measuring the relief texture of specular or partially specular surfaces using a specialized imaging device with a concave parabolic mirror to view multiple angles in a single image. Laser scanning typically fails for specular surfaces because of light scattering, but our method is explicitly designed for specular surfaces. Also, the spatial resolution of the measured geometry is significantly higher than standard methods, so very small surface details are captured. Furthermore, spatially varying reflectance is measured simultaneously, i.e., both texture color and texture shape are retrieved.
Jing Wang 0008, Kristin J. Dana
IEEE Trans. Pattern Anal. Mach. Intell.2
2005 Polarization Multiplexing for Bidirectional Imaging
abstract
Our goal is to incorporate polarization in appearance-based modeling in an efficient and meaningful way. Polarization has been used in numerous prior studies for separating diffuse and specular reflectance components, but in this work we show that it also can be used to separate surface reflectance contributions from individual light sources. Our approach is called polarization multiplexing and it has significant impact in appearance modeling and bidirectional imaging where the image as a function of illumination direction is needed. Multiple unknown light sources can illuminate the scene simultaneously, and the individual contributions to the overall surface reflectance can be estimated. To develop the method of polarization multiplexing, we use a relationship between light source direction and intensity modulation. Inverting this transformation enables the individual intensity contributions to be estimated. In addition to polarization multiplexing, we show that phase histograms from the intensity modulations can be used to estimate scene properties including the number of light sources.
Gabriela Oana Cula, Kristin J. Dana, Dinesh K. Pai
CVPR (2)2
2005 Skin Texture Modeling
Gabriela Oana Cula, Kristin J. Dana, Frank P. Murphy, Babar K. Rao
Int. J. Comput. Vis.2
2004 Hybrid Textons: Modeling Surfaces with Reflectance and Geometry
Jing Wang 0008, Kristin J. Dana
CVPR (1)2
2004 3D Texture Recognition Using Bidirectional Feature Histograms
Gabriela Oana Cula, Kristin J. Dana
Int. J. Comput. Vis.2
2004 Clustering and blending for texture synthesis
Jasvinder Singh, Kristin J. Dana
Pattern Recognit. Lett.2
2003 A Novel Approach For Texture Shape Recovery
abstract
In vision and graphics, there is a sustained interest in capturing accurate 3D shape with various scanning devices. However, the resulting geometric representation is only part of the story. Surface texture of real objects is also an important component of the representation and fine-scale surface geometry such as surface markings, roughness, and imprints, are essential in highly realistic rendering and accurate prediction. We present a novel approach for measuring the fine-scale surface shape of specular surfaces using a curved mirror to view multiple angles in a single image. A distinguishing aspect of our method is that it is designed for specular surfaces, unlike many methods (e.g. laser scanning) which cannot handle highly specular objects. Also, the spatial resolution is very high so that it can resolve very small surface details that are beyond the resolution of standard devices. Furthermore, our approach incorporates the simultaneous use of a bidirectional texture measurement method, so that spatially varying bidirectional reflectance is measured at the same time as surface shape.
Jing Wang 0008, Kristin J. Dana
ICCV2
2001 Compact Representation of Bidirectional Texture Functions
abstract
A bidirectional texture function (BTF) describes image texture as it varies with viewing and illumination direction. Many real world surfaces such as skin, fur, gravel, etc. exhibit fine-scale geometric surface detail. Accordingly, variations in appearance with viewing and illumination direction may be quite complex due to local foreshortening, masking and shadowing. Representations of surface texture that support robust recognition must account for these effects. We construct a representation which captures the underlying statistical distribution of features in the image texture as well as the variations in this distribution with viewing and illumination direction. The representation combines clustering to learn characteristic image features and principle components analysis to reduce the space of feature histograms. This representation is based on a core image set as determined by a quantitative evaluation of importance of individual images in the overall representation. The result is a compact representation and a recognition method where a single novel image of unknown viewing and illumination direction can be classified efficiently. The CUReT (Columbia-Utrecht reflectance and texture) database is used as a test set for evaluation of these methods.
Gabriela Oana Cula, Kristin J. Dana
CVPR (1)2
2001 BRDF/BTF Measurement Device
abstract
Capturing surface appearance is important for a large number of applications. Appearance of real world surfaces is difficult to model as it varies with the direction of illumination as well as the direction from which it is viewed. Consequently, measurements of the BRDF (bidirectional reflectance distribution function) have been important. In addition, many applications require measuring how the entire surface reflects light, i.e. spatially varying BRDF measurements are important as well. For compactness we refer to a spatially varying BRDF as a BTF (bidirectional texture function). Measurements of BRDF and/or BTF typically require significant resources in time and equipment. In this work, a device for BRDF/BTF measurement is presented that is compact, economical and convenient. The device uses the approach of curved mirrors to remove the need for hemispherical positioning of the camera and illumination source. Instead, simple planar translations of optical components are used to vary the illumination direction and to scan the surface. Furthermore, the measurement process is fast because the device enables simultaneous measurements of multiple viewing directions.
Kristin J. Dana
ICCV1
1999 Correlation Model for 3D Texture
abstract
While an exact definition of texture is somewhat elusive, texture can be qualitatively described as a distribution of color, albedo or local normal on a surface. In the literature, the word texture is often used to describe a color or albedo variation on a smooth surface. We refer to such texture as 2D texture. In real world scenes, texture is often due to surface height variations and can be termed 3D texture. Because of local foreshortening and masking, oblique views of 3D texture are not simple transformations of the frontal view. Consequently, texture representations such as the correlation function or power spectrum are also affected by local foreshortening and masking. This work presents a correlation model for a particular class of 3D textures. The model characterizes the spatial relationship among neighboring pixels in an image of 3D texture and the change of this spatial relationship with viewing direction.
Kristin J. Dana, Shree K. Nayar
ICCV1
1999 Texture histograms as a function of irradiation and viewing direction
Bram van Ginneken, Jan J. Koenderink, Kristin J. Dana
Int. J. Comput. Vis.3
1999 Bidirectional Reflection Distribution Function of Thoroughly Pitted Surfaces
Jan J. Koenderink, Andrea J. van Doorn, Kristin J. Dana, Shree K. Nayar
Int. J. Comput. Vis.3
1999 Reflectance and Texture of Real-World Surfaces
abstract
In this work, we investigate the visual appearance of real-world surfaces and the dependence of appearance on the geometry of imaging conditions. We discuss a new texture representation called the BTF (bidirectional texture function) which captures the variation in texture with illumination and viewing direction. We present a BTF database with image textures from over 60 different samples, each observed with over 200 different combinations of viewing and illumination directions. We describe the methods involved in collecting the database as well as the importqance and uniqueness of this database for computer graphics. A related quantity to the BTF is the familiar BRDF (bidirectional reflectance distribution function). The measurement methods involved in the BTF database are conducive to simultaneous measurement of the BRDF. Accordingly, we also present a BRDF database with reflectance measurements for over 60 different samples, each observed with over 200 different combinations of viewing and illumination directions. Both of these unique databases are publicly available and have important implications for computer graphics.
Kristin J. Dana, Bram van Ginneken, Shree K. Nayar, Jan J. Koenderink
ACM Trans. Graph.1
1998 Histogram Model for 3D Textures
abstract
Image texture can arise not only from surface albedo variations (2D texture) but also from surface height variations (3D texture). Since the appearance of 3D texture depends on the illumination and viewing direction in a complicated manner, such image texture can be called a bidirectional texture function. A fundamental representation of image texture is the histogram of pixel intensities. Since the histogram of 3D texture also depends on the illumination and viewing directions in a complex fashion, we refer to it as a bidirectional histogram. In this work, we present a concise analytical model for the bidirectional histogram of Lambertian, isotropic, randomly rough surfaces, which are common in real-world scenes. We demonstrate the accuracy of the histogram model by fitting to several samples from the Columbia-Utrecht texture database. The parameters obtained from the model fits are roughness measures which can be used in texture recognition schemes. In addition, the model has potential application in estimating illumination direction in scenes where surfaces of known tilt and roughness are visible. We demonstrate the usefulness of our model by employing it in a novel 3D texture synthesis procedure.
Kristin J. Dana, Shree K. Nayar
CVPR1
1997 Reflectance and Texture of Real-World Surfaces Authors
abstract
In this work, we investigate the visual appearance of real-world surfaces and the dependence of appearance on imaging conditions. We present a BRDF (bidirectional reflectance distribution function) database with reflectance measurements for over 60 different samples, each observed with over 200 different combinations of viewing and source directions. We fit the BRDF measurements to two recent models to obtain a BRDF parameter database. These BRDF parameters can be directly used for both image analysis and image synthesis. Finally, we present a BTF (bidirectional texture function) database with image textures from over 60 different samples, each observed with over 200 different combinations of viewing and source directions. Each of these unique databases has important implications for a variety of vision algorithms and each is made publicly available.
Kristin J. Dana, Shree K. Nayar, Bram van Ginneken, Jan J. Koenderink
CVPR1
1994 Real-time scene stabilization and mosaic construction
abstract
We describe a real-time system designed to construct a stable view of a scene through aligning images of an incoming video stream and dynamically constructing an image mosaic. This system uses a video processing unit developed by the David Sarnoff Research Center called the Vision Front End (VFE-100) for the pyramid-based image processing tasks required to implement this process. This paper includes a description of the multiresolution coarse-to-fine image registration strategy, the techniques used for mosaic construction, the implementation of this process on the VFE-100 system, and experimental results showing image mosaics constructed with the VFE-100.>
Michael W. Hansen, P. Anandan 0001, Kristin J. Dana, Gooitzen S. van der Wal, Peter J. Burt
WACV3
1994 Frameless registration of MR and CT 3D volumetric data sets
abstract
In this paper we present techniques for frameless registration of 3D Magnetic Resonance (MR) and Computed Tomography (CT) volumetric data of the head and spine. We present techniques for estimating a 3D affine or rigid transform which can be used to resample the CT (or MR) data to align with the MR (or CT) data. Our technique transforms the MR and CT data sets with spatial filters so they can be directly matched. The matching is done by a direct optimization technique using a gradient based descent approach and a coarse-to-fine control strategy over a 4D pyramid. We present results on registering the head and spine data by matching 3D edges and results on registering cranial ventricle data by matching images filtered by a Laplacian of a Gaussian.>
Rakesh Kumar 0001, Kristin J. Dana, P. Anandan 0001, Neil E. Okamoto, James R. Bergen, Paul F. Hemler, Thilaka S. Sumanaweera, Petra A. van den Elsen, John R. Adler Jr.
WACV2