EDBT 2026 Demo / reviewers in the wild / expert
Scott McCloskey
dblp:30/44
· DBLP profile ↗
40ranked-venue papers
17as first author
10since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 36 · 15 first-author · 9 since 2021Artificial intelligence and machine learning · 25 · 11 first-author · 6 since 2021Security and privacy · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | EBS-EKF: Accurate and High Frequency Event-based Star TrackingabstractEvent-Based sensors (EBS) are a promising new technology for star tracking due to their low latency and power efficiency, but prior work has thus far been evaluated exclusively in simulation with simplified signal models. We propose a novel algorithm for event-based star tracking, grounded in an analysis of the EBS circuit and an extended Kalman filter (EKF). We quantitatively evaluate our method using real night sky data, comparing its results with those from a space-ready active-pixel sensor (APS) star tracker. We demonstrate that our method is an order-of-magnitude more accurate than existing methods due to improved signal modeling and state estimation, while providing more frequent updates and greater motion tolerance than conventional APS trackers. We provide all code*and the first dataset of events synchronized with APS solutions. Albert W. Reed, Connor Hashemi, Dennis Melamed, Nitesh Menon, Keigo Hirakawa, Scott McCloskey |
CVPR | 6 |
| 2025 | Re-identifying People in Video via Learned Temporal Attention and Multi-modal Foundation ModelsabstractBiometric recognition from security camera video is a challenging problem when the individuals change clothes or when they are partly occluded. Others have recently demonstrated that CLIP's visual encoder performs well in this domain, but existing methods fail to make use of the model's text encoder or temporal information available in video. In this paper, we present VCLIP, a method for person identification in videos captured in challenging poses and with changes to a person's clothing. Harnessing the power of pre-trained vision-language models, we Jointly train a temporal fusion network while fine-tuning the visual encoder. To leverage the cross-modal embedding space, we use learned biometric pedestrian attribute features to further enhance our model's person re-identification (Re-ID) ability. We demonstrate significant performance improvements via experiments with the MEVID and CCVID datasets, particularly in the more challenging clothes-changing conditions. In support of this and future methods that use textual attributes for Re-ID with multimodal models, we release a dataset of annotated pedestrian attributes for the popular MEVID dataset [4]. Cole Hill, Florence Yellin, Krishna Regmi, Dawei Du, Scott McCloskey |
WACV | 5 |
| 2025 | MetaVIn: Meteorological and Visual Integration for Atmospheric Turbulence Strength EstimationabstractLong-range image understanding is a challenging task for computer vision due to the presence of atmospheric turbulence. Turbulence can degrade image quality (blur and geometric distortion) due to the medium's spatio-temporal varying index of refraction bending light rays. The strength of atmospheric turbulence is quantified by the refractive index structure parameter$C_n^2$, and estimating it is important both as an indicator of image degradation and is useful for downstream tasks including video restoration and estimating true shape and range/depth. However, traditional methods for estimating$C_{n}^{2}$involve expensive and complex optical equipment, limiting their practicality. In this paper, we propose MetaVIn: a Meteorological and Visual Integration system to predict atmospheric turbulence strength. Our method leverages image quality metrics to capture sharpness and blur, combined with meteorological information within a Kolmogorov Arnold Network (KAN). We demonstrate that this approach provides a more accurate and generalizable estimation of$C_{n}^{2}$, outperforming previous state-of-the-art methods in both blind image quality assessment and passive video-based turbulence strength estimation on a large dataset of 35,364 image samples with accompanying ground truth scintillometer measurements for$C_n^2$. Our method enables better prediction and mitigation of atmospheric image degradation while being useful in applications such as shape and range estimation, enhancing the practical utility of our approach. Ripon K. Saha, Scott McCloskey, Suren Jayasuriya |
WACV | 2 |
| 2024 | High-Resolution Image Enumeration for Low-Resolution Face RecognitionabstractImage enhancement is a long-studied problem that is understood to be fundamentally ill-posed, meaning that there are many high-quality images that are consistent with any given low-quality observation. When image enhancement is applied to biometric samples, such as sharpening or super-resolving a face image, this ill-posedness is often ignored and a single high-quality sample is estimated without ensuring that it's consistent with the observation. In this work, we describe a method to enumerate multiple high-quality samples from a single input, all of which are consistent with the low-quality input, and use this to estimate confidence in a face-based search result. This method quantifies the sample- and gallery-conditioned uncertainty by enumerating multiple high-quality images from a single low-quality sample, ensuring that each is consistent with the input. We demonstrate this by first showing that, even with modern deep-face features, low-resolution face recognition is still ill-posed even when applied to frontal face images. We enumerate multiple high-resolution (HR) images that are consistent with a low-resolution (LR) face sample, represent these with modern deep features, and demonstrate a subject-specific ill-posedness in recognition. Can Chen 0004, Scott McCloskey |
FG | 2 |
| 2024 | Concurrent Band Selection and Traversability Estimation from Long-Wave Hyperspectral Imagery in Off-Road SettingsabstractAutonomous navigation has become increasingly popular in recent years; However, most existing methods focus on on-road navigation and utilize active sensors, such as LiDAR. This paper instead focuses on autonomous off-road navigation using traversability estimation from passive sensors, specifically long-wave (LW) hyperspectral imagery (HSI). We present a method for selecting a subset of hyperspectral bands that are most useful for traversability estimation by designing a band selection module that designs a minimal sensor that measures sparsely-sampled spectral bands while jointly training a semantic segmentation network for traversability estimation. The effectiveness of our method is demonstrated using our dataset of LW HSI from diverse off-road scenes including forest, desert, snow, ponds, and open fields. Our dataset includes imagery collected both during the daytime and nighttime during various weather conditions, including challenging scenes with a wide range of obstacles. Using our method, we learn a small subset (2%) of all the HSI bands that can achieve competitive or better traversability estimation accuracy to that achieved when utilizing all hyperspectral bands. Using only 5 bands, our method is able to achieve a mean class accuracy that is only 1.3% less than that achieved using full 256-band HSI and only 0.1% less than that achieved using 250-band HSI, demonstrating the success of our method. Florence Yellin, Scott McCloskey, Cole Hill, Brian Clipp |
WACV | 2 |
| 2023 | DOERS: Distant Observation Enhancement and Recognition SystemabstractIn order to recognize people across long distances and from elevated viewpoints, biometric systems must handle the challenges of imaging through atmospheric turbulence and non-frontal presentations, in addition to the traditional A-PIE challenges of aging, pose, illumination, and expression. While individual biometric modalities such as facial appearance, gait, and whole body appearance each have a role to play, no single modality can address all of these challenges. This paper describes a novel multi-modal biometric recognition system that addresses the challenges of atmospheric turbulence, occlusions, and elevated viewpoints by combining these modalities. We demonstrate our system on both $R G B$ video-based identity verification and both open and closed-world search. Dawei Du, Cole Hill, Gabriel Bertocco, Maurício Pamplona Segundo, Wes Robbins, Brandon RichardWebster, Roderic Collins, Sudeep Sarkar, Terrance E. Boult, Scott McCloskey |
IJCB | 10 |
| 2023 | AG-ReID 2023: Aerial-Ground Person Re-identification Challenge ResultsabstractPerson re-identification (Re-ID) on aerial-ground platforms has emerged as an intriguing topic within computer vision, presenting a plethora of unique challenges. Highflying altitudes of aerial cameras make persons appear differently in terms of viewpoints, poses, and resolution compared to the images of the same person viewed from ground cameras. Despite its potential, few algorithms have been developed for person re-identification on aerial-ground data, mainly due to the absence of comprehensive datasets. In response, we have collected a large-scale dataset and organized the Aerial-Ground person Re-IDentification Challenge (AG-ReID2023) to foster advancements in the field. The dataset comprises 100,502 images with 1,615 unique identities, including 51,530 training images featuring 807 identities. The test set is divided into two subsets: Aerial to Ground (808 ids, 4,348 query images, 19,259 gallery images) and Ground to Aerial (808 ids, 4,151 query images, 21,214 gallery images). In addition, we manually annotate individuals with their matching IDs across cameras and provide 15 soft attribute labels. The AG-ReID2023 Challenge in conjunction with the 7thIEEE International Joint Conference on Biometrics (IJCB) has garnered interest from numerous institutes, resulting in the submission of five distinct algorithms. We provide an in-depth examination of the evaluation outcomes and present our findings from the contest. For additional details, kindly refer to the official website1.1https://agreid23.github.io. Kien Nguyen Thanh, Clinton Fookes, Sridha Sridharan, Feng Liu 0037, Xiaoming Liu 0002, Arun Ross, Dana Michalski, Debayan Deb, Mahak Kothari, Manisha Saini, Dawei Du, Scott McCloskey, Gabriel Bertocco, Fernanda A. Andaló, Terrance E. Boult, Anderson Rocha 0001, Haidong Zhu, Zhaoheng Zheng, Ramakant Nevatia, Zaigham A. Randhawa, Sinan Sabri, Gianfranco Doretto |
IJCB | 13 |
| 2023 | MEVID: Multi-view Extended Videos with Identities for Video Person Re-IdentificationabstractIn this paper, we present the Multi-view Extended Videos with Identities (MEVID) dataset for large-scale, video person re-identification (ReID) in the wild. To our knowledge, MEVID represents the most-varied video person ReID dataset, spanning an extensive indoor and outdoor environment across nine unique dates in a 73-day window, various camera viewpoints, and entity clothing changes. Specifically, we label the identities of 158 unique people wearing 598 outfits taken from 8, 092 tracklets, average length of about 590 frames, seen in 33 camera views from the very-large-scale MEVA person activities dataset. While other datasets have more unique identities, MEVID emphasizes a richer set of information about each individual, such as: 4 outfits/identity vs. 2 outfits/identity in CCVID, 33 viewpoints across 17 locations vs. 6 in 5 simulated locations for MTA, and 10 million frames vs. 3 million for LS-VID. Being based on the MEVA video dataset, we also inherit data that is intentionally demographically balanced to the continental United States. To accelerate the annotation process, we developed a semi-automatic annotation framework and GUI that combines state-of-the-art real-time models for object detection, pose estimation, person ReID, and multi-object tracking. We evaluate several state-of-the-art methods on MEVID challenge problems and comprehensively quantify their robustness in terms of changes of outfit, scale, and background location. Our quantitative analysis on the realistic, unique aspects of MEVID shows that there are significant remaining challenges in video person ReID and indicates important directions for future research. Daniel Davila, Dawei Du, Bryon Lewis, Christopher Funk, Joseph VanPelt, Roderic Collins, Kellie Corona, Matt S. Brown, Scott McCloskey, Anthony Hoogs, Brian Clipp |
WACV | 9 |
| 2022 | Resolution Transfer for Object Detection from Satellite ImageryabstractSmallsat constellations are an increasingly common source of global-scale overhead imagery that are refreshed with a higher frequency than traditional satellites. The smaller size and lower cost of smallsats enable frequent revisits, but result in images with lower resolution and quality than the high resolution (HR) images from traditional satellites. In order to benefit from the increased temporal frequency provided by smallsat constellations, new approaches are needed to automatically detect objects in their imagery. We present a super resolution (SR) approach that incorporates domain adaptation (DA) to enable object detection from low resolution (LR) images without the need for paired training data or annotations in the LR domain. Our Resolution Transfer approach addresses the resolution and quality loss for smallsats, as demonstrated via airplane detection. Florence Yellin, Michael Albright, Scott McCloskey |
ICPR | 4 |
| 2021 | Bridging the Gap Between Computational Photography and Visual RecognitionabstractWhat is the current state-of-the-art for image restoration and enhancement applied to degraded images acquired under less than ideal circumstances? Can the application of such algorithms as a pre-processing step improve image interpretability for manual analysis or automatic visual recognition to classify scene content? While there have been important advances in the area of computational photography to restore or enhance the visual quality of an image, the capabilities of such techniques have not always translated in a useful way to visual recognition tasks. Consequently, there is a pressing need for the development of algorithms that are designed for the joint problem of improving visual appearance and recognition, which will be an enabling factor for the deployment of visual recognition tools in many real-world scenarios. To address this, we introduce the UG$^2$dataset as a large-scale benchmark composed of video imagery captured under challenging conditions, and two enhancement tasks designed to test algorithmic impact on visual quality and automatic object recognition. Furthermore, we propose a set of metrics to evaluate the joint improvement of such tasks as well as individual algorithmic advances, including a novel psychophysics-based evaluation regime for human assessment and a realistic set of quantitative measures for object recognition performance. We introduce six new algorithms for image restoration or enhancement, which were created as part of the IARPA sponsored UG$^2$Challenge workshop held at CVPR 2018. Under the proposed evaluation regime, we present an in-depth analysis of these algorithms and a host of deep learning-based and classic baseline approaches. From the observed results, it is evident that we are in the early days of building a bridge between computational photography and visual recognition, leaving many opportunities for innovation in this area. Rosaura G. VidalMata, Sreya Banerjee, Brandon RichardWebster, Michael Albright, Pedro Davalos, Scott McCloskey, Ben Miller, Asong Tambo, Sushobhan Ghosh, Sudarshan Nagesh, Ye Yuan 0012, Yueyu Hu, Wenhan Yang, Xiaoshuai Zhang, Jiaying Liu 0001, Zhangyang Wang, Hwann-Tzong Chen, Tzu-Wei Huang, Wen-Chi Chin, Yi-Chun Li, Mahmoud Lababidi, Charles Otto, Walter J. Scheirer |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2019 | Jittered Exposures for Light Field Super-Resolution
Nianyi Li, Scott McCloskey, Jingyi Yu 0001 |
ICIP | 2 |
| 2019 | Detecting GAN-Generated Imagery Using Saturation CuesabstractImage forensics is an increasingly relevant problem, as it can potentially address online disinformation campaigns and mitigate problematic aspects of social media. Of particular interest, given its recent successes, is the detection of imagery produced by Generative Adversarial Networks (GANs), e.g. `deepfakes'. Leveraging large training sets and extensive computing resources, recent GANs can be trained to generate synthetic imagery which is (in some ways) indistinguishable from real imagery. We analyze the structure of the generating network of a popular GAN implementation [1], and show that the network's treatment of exposure is markedly different from a real camera. We further show that this cue can be used to distinguish GAN-generated imagery from camera imagery, including effective discrimination between GAN imagery and real camera images used to train the GAN. Scott McCloskey, Michael Albright |
ICIP | 1 |
| 2019 | Analyzing Modern Camera Response FunctionsabstractCamera Response Functions (CRFs) map the irradiance incident at a sensor pixel to an intensity value in the corresponding image pixel. The nonlinearity of CRFs impact physics-based and low-level computer vision methods like de-blurring, photometric stereo, etc. In addition, CRFs have been used for forensics to identify regions of an image spliced in from a different camera. Despite its importance, the process of radiometrically calibrating a camera's CRF is significantly harder and less standardized than geometric calibration. Competing methods use different mathematical models of the CRF, some of which are derived from an outdated dataset. We present a new dataset of 178 CRFs from modern digital cameras, derived from 1565 camera review images available online, and use it to answer a series of questions about CRFs. Which mathematical models are best for CRF estimation? How have they changed over time? And how unique are CRFs from camera to camera? Can Chen 0004, Scott McCloskey, Jingyi Yu 0001 |
WACV | 2 |
| 2019 | Low-and Semantic-Level Cues for Forensic Splice DetectionabstractImage forensics is increasingly of interest, due to the deluge of online photo-sharing platforms and high-quality image editing software. While advances in computer vision enable these editing tools, they also provide a means by which to detect such tampering. Recent work in large-scale image phylogeny, for instance, aims to infer relationships between multiple images which may have contributed to a manipulation. A key task within this scope is detecting splicing between a pair of images, which we address with a combination of low-and semantic-level cues in order to provide fast detection with the false alarm rate demanded by large image collections. We show that, while deep learning approaches contribute to these ends, traditional feature-based detection still forms the basis of a useful detector, providing scale invariance without incurring the complexity associated with deep networks. Our method out-performs the state-of-the-art approach based on deep learning, measured on a challenge dataset for splice detection, and can also be used for detection of copy-move manipulations. Asong Tambo, Michael Albright, Scott McCloskey |
WACV | 3 |
| 2018 | Focus Manipulation Detection via Photometric Histogram AnalysisabstractWith the rise of misinformation spread via social media channels, enabled by the increasing automation and realism of image manipulation tools, image forensics is an increasingly relevant problem. Classic image forensic methods leverage low-level cues such as metadata, sensor noise fingerprints, and others that are easily fooled when the image is re-encoded upon upload to facebook, etc. This necessitates the use of higher-level physical and semantic cues that, once hard to estimate reliably in the wild, have become more effective due to the increasing power of computer vision. In particular, we detect manipulations introduced by artificial blurring of the image, which creates inconsistent photometric relationships between image intensity and various cues. We achieve 98% accuracy on the most challenging cases in a new dataset of blur manipulations, where the blur is geometrically correct and consistent with the scene's physical arrangement. Such manipulations are now easily generated, for instance, by smartphone cameras having hardware to measure depth, e.g. 'Portrait Mode' of the iPhone7Plus. We also demonstrate good performance on a challenge dataset evaluating a wider range of manipulations in imagery representing 'in the wild' conditions. Can Chen 0004, Scott McCloskey, Jingyi Yu 0001 |
CVPR | 2 |
| 2017 | Image Splicing Detection via Camera Response Function AnalysisabstractRecent advances on image manipulation techniques have made image forgery detection increasingly more challenging. An important component in such tools is to fake motion and/or defocus blurs through boundary splicing and copy-move operators, to emulate wide aperture and slow shutter effects. In this paper, we present a new technique based on the analysis of the camera response functions (CRF) for efficient and robust splicing and copy-move forgery detection and localization. We first analyze how non-linear CRFs affect edges in terms of the intensity-gradient bivariable histograms. We show distinguishable shape differences on real vs. forged blurs near edges after a splicing operation. Based on our analysis, we introduce a deep-learning framework to detect and localize forged edges. In particular, we show the problem can be transformed to a handwriting recognition problem an resolved by using a convolutional neural network. We generate a large dataset of forged images produced by splicing followed by retouching and comprehensive experiments show our proposed method outperforms the state-of-the-art techniques in accuracy and robustness. Can Chen 0004, Scott McCloskey, Jingyi Yu 0001 |
CVPR | 2 |
| 2017 | Temporally Coded Illumination for Rolling Shutter Motion De-blurringabstractRolling shutter image sensors are increasingly used over global shutter sensors due to their lower cost and reduced size. As is well known in the computer vision community, one drawback of rolling shutter sensors is that they introduce geometric distortions, such as skew or wobble, when either the sensor or objects in the scene move. This problem has received a great deal of attention, and robust solutions are now widely available. Less well known is the fact that, when used in conjunction with active illumination (i.e., a flash), rolling shutter images are often motion blurred due to timing issues between the sensor and illuminator. We address this in the context of barcode scanning, where blur significantly limits the motion tolerance of rolling shutter-based devices. Building on past work which uses temporally-coded illumination patterns to improve the invertibility of motion blur, we modulate the illuminator in a particular temporal pattern in order to ensure that the sharp image can be recovered despite spatially-varying blur. Experimentally, we modify a commercial, off-the-shelf scanner to demonstrate an ability to correctly decode a barcode moving faster than the stated motion tolerance. Scott McCloskey, Sharath Venkatesha |
WACV | 1 |
| 2016 | Fast, high dynamic range light field processing for pattern recognitionabstractWe present a light field processing method to quickly produce an image for pattern recognition. Unlike processing for aesthetic purposes, our objective is not to produce the best-looking image, but to produce a recognizable image as fast as possible. By leveraging the recognition algorithm's dynamic range and robustness to optical defocus, we develop carefully-chosen tradeoffs to ensure recognition at a much lower level of computational complexity. Capitalizing on the algorithm's dynamic range yields large speedups by minimizing the number of light field views used in refocusing. Robustness to optical defocus allows us to quantize the refocus parameter and minimize the number of interpolations. The resulting joint optimization is performed via dynamic programming to choose the set of views which, when combined, produce a recognizable refocused image in the least possible computing time. We demonstrate the improved recognition dynamic range of barcode scanning using a Lytro camera, and dramatic reductions in computational complexity on a low-power embedded processor. Scott McCloskey, Ben Miller |
ICCP | 1 |
| 2015 | A Low-Noise Fluttering Shutter Camera Handling Accelerated MotionabstractWe address the problem of motion blur removal using a computational camera with a fluttering shutter. While there are several prototype flutter shutter cameras, and many scenarios in which motion blur is problematic, there are few real-world uses of flutter shutter cameras due to two important limitations. The first is that the shutter mechanisms used to date - primarily Liquid Crystal Display (LCD) elements or electronic shutters - increase noise due to reduced light efficiency or multiple readouts, respectively. Secondly, the class of motions to which the flutter shutter is applicable has been limited to linear, constant velocity motion. We address the first limitation by developing a prototype flutter shutter camera with a reflective element providing high light efficiency and a single-read imaging system. In addition to improved noise performance, this method of exposure modulation imposes fewer limitations on the shutter sequence, allowing us to extend the flutter shutter technique to cases with constant (non-zero) acceleration. We demonstrate both the noise reduction and improved reconstructions in the case of an accelerating camera. Scott McCloskey, Sharath Venkatesha, Kelly Muldoon, Ryan Eckman |
WACV | 1 |
| 2014 | Latent Domains Modeling for Visual Domain AdaptationabstractTo improve robustness to significant mismatches between source domain and target domain - arising from changes such as illumination, pose and image quality - domain adaptation is increasingly popular in computer vision. But most of methods assume that the source data is from single domain, or that multi-domain datasets provide the domain label for training instances. In practice, most datasets are mixtures of multiple latent domains, and difficult to manually provide the domain label of each data point. In this paper, we propose a model that automatically discovers latent domains in visual datasets. We first assume the visual images are sampled from multiple manifolds, each of which represents different domain, and which are represented by different subspaces. Using the neighborhood structure estimated from images belonging to the same category, we approximate the local linear invariant subspace for each image based on its local structure, eliminating the category-specific elements of the feature. Based on the effectiveness of this representation, we then propose a squared-loss mutual information based clustering model with category distribution prior in each domain to infer the domain assignment for images. In experiment, we test our approach on two common image datasets, the results show that our method outperforms the existing state-of-the-art methods, and also show the superiority of multiple latent domain discovery. Caiming Xiong, Scott McCloskey, Shao-Hang Hsieh, Jason J. Corso |
AAAI | 2 |
| 2014 | Improved Motion Invariant Deblurring through Motion Estimation
Scott McCloskey |
ECCV (4) | 1 |
| 2014 | Motion aware motion invarianceabstractWe address motion de-blurring using a computational camera that captures an image while the stabilizing optical element moves in a modified Canon IS lens. Our work builds on that of Levin et al. [11], who introduce parabolic motion as a means of achieving invariance to unknown subject velocity in an a priori known direction. While the previous work addresses a specific scenario - exact knowledge of motion orientation and a uniform, symmetric prior on its magnitude - we generalize this to address scenarios where the motion of objects in the scene or the camera itself are known to various extents. We describe a motion invariant camera based on an off-the-shelf lens, and show how its motion and position sensors can be used to inform both the image capture and de-blurring. We demonstrate that our changes to motion invariance improve the quality of captured images in the case of both object and camera motion. Scott McCloskey, Kelly Muldoon, Sharath Venkatesha |
ICCP | 1 |
| 2014 | Masking Light Fields to Remove Partial OcclusionabstractWe address partial occlusion due to objects close to a micro lens-based light field camera. Partial occlusion degrades the quality of the image, and may obscure important details of the background scene. In order to remove the effects of partial occlusion post-capture, previous methods with traditional cameras have required the photographer to capture multiple, precisely registered, images under different settings. The use of a light field camera eliminates this requirement, as the camera simultaneously captures multiple views of the scene, making it possible to remove partial occlusion from a single image captured by a hand-held camera. Relative to past approaches for light field completion, we show significantly better performance for the small viewpoint changes inherent to a handheld light field camera, and avoid the need for time-domain data for occlusion estimation. Scott McCloskey |
ICPR | 1 |
| 2014 | Recognition of 3D package shapes for single camera metrologyabstractMany applications of 3D object measurement have become commercially viable due to the recent availability of low-cost range cameras such as the Microsoft Kinect. We address the application of measuring an object's dimensions for the purpose of billing in shipping transactions, where high accuracy is required for certification. In particular, we address cases where an object's pose reduces the accuracy with which we can estimate dimensions from a single camera. Because the class of object shapes is limited in the shipping domain, we perform a closed-world recognition in order to determine a shape model which can account for missing parts, and/or to induce the user to reposition the object for higher accuracy. Our experiments demonstrate that the addition of this recognition step significantly improves system accuracy. Ryan Lloyd, Scott McCloskey |
WACV | 2 |
| 2014 | Multimedia event detection with multimodal feature fusion and temporal concept localization
Sangmin Oh, Scott McCloskey, Ilseo Kim, Arash Vahdat, Kevin J. Cannons, Hossein Hajimirsadeghi, Greg Mori, A. G. Amitha Perera, Megha Pandey, Jason J. Corso |
Mach. Vis. Appl. | 2 |
| 2012 | Local Expert Forest of Score Fusion for Video Event Classification
Jingchen Liu, Scott McCloskey, Yanxi Liu 0001 |
ECCV (5) | 2 |
| 2012 | Training data recycling for multi-level learning
Jingchen Liu, Scott McCloskey, Yanxi Liu 0001 |
ICPR | 2 |
| 2012 | Activity detection in the wild using video metadata
Scott McCloskey, Pedro Davalos |
ICPR | 1 |
| 2012 | Design and Estimation of Coded Exposure Point Spread FunctionsabstractWe address the problem of motion deblurring using coded exposure. This approach allows for accurate estimation of a sharp latent image via well-posed deconvolution and avoids lost image content that cannot be recovered from images acquired with a traditional shutter. Previous work in this area has used either manual user input or alpha matting approaches to estimate the coded exposure Point Spread Function (PSF) from the captured image. In order to automate deblurring and to avoid the limitations of matting approaches, we propose a Fourier-domain statistical approach to coded exposure PSF estimation that allows us to estimate the latent image in cases of constant velocity, constant acceleration, and harmonic motion. We further demonstrate that previously used criteria to choose a coded exposure PSF do not produce one with optimal reconstruction error, and that an additional 30 percent reduction in Root Mean Squared Error (RMSE) of the latent image estimate can be achieved by incorporating natural image statistics. Scott McCloskey, Yuanyuan Ding, Jingyi Yu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2011 | Motion invariance and custom blur from lens motionabstractWe demonstrate that image stabilizing hardware included in many camera lenses can be used to implement motion invariance and custom blur effects. Motion invariance is intended to capture images where objects within a range of velocities appear defocused with the same point spread function, obviating the need for blur estimation in advance of de-blurring. We show that the necessary parabolic motion can be implemented with stabilizing lens motion, but that the range of velocities to which capture is invariant decreases with increasing exposure time. We also show that, when that range is expanded through increased lens displacement, lens motion becomes less repeatable. In addition to motion invariance, we demonstrate that stabilizing lens motion can be used to design custom defocus kernels for aesthetic purposes, and can replace lens accessories. Scott McCloskey, Kelly Muldoon, Sharath Venkatesha |
ICCP | 1 |
| 2011 | Temporally coded flash illumination for motion deblurringabstractWe use temporally sequenced flash illumination to capture coded exposure images of fast-moving objects in low light environments. These coded flash images allow for accurate estimation of blur-free latent images in the presence of object motion. By distributing flashes over a window of time, we lessen eye safety concerns associated with powerful all-at-once flashes. We show how our flash-based coded exposure system has better robustness to increasing object velocity than shutter-based exposure coding, thereby obviating the need for pre-exposure velocity estimation. We also show that the quality of the estimated sharp image is robust to varying levels of ambient illumination. This and other benefits of our coded flash system are demonstrated with real images acquired using prototype hardware. Scott McCloskey |
ICCV | 1 |
| 2011 | 2D Barcode localization and motion deblurring using a flutter shutter cameraabstractWe describe a system for localizing and deblurring motion-blurred 2D barcodes. Previous work on barcode detection and deblurring has mainly focused on 1D barcodes, and has employed traditional image acquisition which is not robust to motion blur. Our solution is based on coded exposure imaging which, as we show, enables well-posed de-convolution and decoding over a wider range of velocities. To serve this solution, we developed a simple and effective approach for 2D barcode localization under motion blur, a metric for evaluating the quality of the deblurred 2D barcodes, and an approach for motion direction estimation in coded exposure images. We tested our system on real camera images of three popular 2D barcode symbologies: Data Matrix, PDF417 and Aztec Code. Scott McCloskey |
WACV | 2 |
| 2011 | Removal of Partial Occlusion from Single ImagesabstractThis paper examines large partial occlusions in an image which occur near depth discontinuities when the foreground object is severely out of focus. We model these partial occlusions using matting, with the alpha value determined by the convolution of the blur kernel with a pinhole projection of the occluder. The main contribution is a method for removing the image contribution of the foreground occluder in regions of partial occlusion, which improves the visibility of the background scene. The method consists of three steps. First, the region of complete occlusion is estimated using a curve evolution method. Second, the alpha value at each pixel in the partly occluded region is estimated. Third, the intensity contribution of the foreground occluder is removed in regions of partial occlusion. Experiments demonstrate the method's ability to remove the effects of partial occlusion in single images with minimal user input. Scott McCloskey, Michael S. Langer, Kaleem Siddiqi |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2010 | Analysis of Motion Blur with a Flutter Shutter Camera for Non-linear Motion
Yuanyuan Ding, Scott McCloskey, Jingyi Yu 0001 |
ECCV (1) | 2 |
| 2010 | Velocity-Dependent Shutter Sequences for Motion Deblurring
Scott McCloskey |
ECCV (6) | 1 |
| 2010 | Removing Partial Occlusion from Blurred Thin OccludersabstractWe present a method to remove partial occlusion that arises from out-of-focus thin foreground occluders such as wires, branches, or a fence. Such partial occlusion causes the irradiance at a pixel to be a weighted sum of the radiances of a blurred foreground occluder and that of the background. The result is that the background component has lower contrast than it would if seen without the occluder. In order to remove the contribution of the foreground in such regions, we characterize the position and size of the occluder in a narrow aperture image. In subsequent images with wider apertures, we use this characterization to remove the contribution of the foreground, thereby restoring contrast in the background. We demonstrate our method on real camera images without assuming that the background is static. Scott McCloskey, Michael S. Langer, Kaleem Siddiqi |
ICPR | 1 |
| 2009 | Planar orientation from blur gradients in a single imageabstractWe present a focus-based method to recover the orientation of a textured planar surface patch from a single image. The method exploits the relationship between the orientation of equifocal (i.e. uniformly-blurred) contours in the image and the plane's tilt and slant angles. Compared to previous methods that determine planar orientation, we make fewer assumptions about the texture and remove the restriction that images must be acquired through a pinhole aperture. Our method estimates slant and tilt of an image patch in a single image, as compared to depth from defocus methods that require two or more input images. Experiments are performed using a large set of test images. Scott McCloskey, Michael S. Langer |
CVPR | 1 |
| 2009 | Incremental Multiple Kernel Learning for object recognitionabstractA good training dataset, representative of the test images expected in a given application, is critical for ensuring good performance of a visual categorization system. Obtaining task specific datasets of visual categories is, however, far more tedious than obtaining a generic dataset of the same classes. We propose an Incremental Multiple Kernel Learning (IMKL) approach to object recognition that initializes on a generic training database and then tunes itself to the classification task at hand. Our system simultaneously updates the training dataset as well as the weights used to combine multiple information sources. We demonstrate our system on a vehicle classification problem in a video stream overlooking a traffic intersection. Our system updates itself with images of vehicles in poses more commonly observed in the scene, as well as with image patches of the background, leading to an increase in performance. A considerable change in the kernel combination weights is observed as the system gathers scene specific training data over time. The system is also seen to adapt itself to the illumination change in the scene as day transitions to night. Aniruddha Kembhavi, Behjat Siddiquie, Roland Miezianko, Scott McCloskey, Larry Davis 0001 |
ICCV | 4 |
| 2007 | Automated Removal of Partial Occlusion Blur
Scott McCloskey, Michael S. Langer, Kaleem Siddiqi |
ACCV (1) | 1 |
| 2007 | Evolving Measurement Regions for Depth from Defocus
Scott McCloskey, Michael S. Langer, Kaleem Siddiqi |
ACCV (2) | 1 |