Richard P. Wildes

dblp:12/5222 · also Richard Wildes, Rick Wildes · DBLP profile ↗
← Back
64ranked-venue papers
11as first author
10since 2021 · last 2025
0000-0003-3433-1329ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 58 · 8 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 44 · 7 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author
YearPublicationVenuePosition
2025 Quantifying and Learning Static vs. Dynamic Information in Deep Spatiotemporal Networks
abstract
There is limited understanding of the information captured by deep spatiotemporal models in their intermediate representations. For example, while evidence suggests that action recognition algorithms are heavily influenced by visual appearance in single frames, no quantitative methodology exists for evaluating such static bias in the latent representation compared to bias toward dynamics. We tackle this challenge by proposing an approach for quantifying the static and dynamic biases of any spatiotemporal model, and apply our approach to three tasks, action recognition, automatic video object segmentation (AVOS) and video instance segmentation (VIS). Our key findings are: (i) Most examined models are biased toward static information. (ii) Some datasets that are assumed to be biased toward dynamics are actually biased toward static information. (iii) Individual channels in an architecture can be biased toward static, dynamic or jointly encode a combination static and dynamic information. (iv) Most models converge to their culminating biases in the first half of training. We then explore how these biases affect performance on dynamically biased datasets. For action recognition, we propose StaticDropout, a semantically guided dropout that debiases a model from static information toward dynamics. For AVOS, we design a better combination of fusion and cross connection layers compared with previous architectures.
Matthew Kowal, Mennatullah Siam, Md. Amirul Islam, Neil D. B. Bruce, Richard P. Wildes, Konstantinos G. Derpanis
IEEE Trans. Pattern Anal. Mach. Intell.5
2024 Selective, Interpretable and Motion Consistent Privacy Attribute Obfuscation for Action Recognition
abstract
Concerns for the privacy of individuals captured in public imagery have led to privacy-preserving action recognition. Existing approaches often suffer from issues arising through obfuscation being applied globally and a lack of interpretability. Global obfuscation hides privacy sensitive regions, but also contextual regions important for action recognition. Lack of interpretability erodes trust in these new technologies. We highlight the limitations of current paradigms and propose a solution: Human selected privacy templates that yield interpretability by design, an ob-fuscation scheme that selectively hides attributes and also induces temporal consistency, which is important in action recognition. Our approach is architecture agnostic and directly modifies input imagery, while existing approaches generally require architecture training. Our approach offers more flexibility, as no training is required, and outperforms alternatives on three widely used datasets.
Filip Ilic, He Zhao 0004, Thomas Pock, Richard P. Wildes
CVPR4
2024 Visual Concept Connectome (VCC): Open World Concept Discovery and Their Interlayer Connections in Deep Models
abstract
Understanding what deep network models capture in their learned representations is a fundamental challenge in computer vision. We present a new methodology to understanding such vision models, the Visual Concept Con-nectome (VCC), which discovers human interpretable concepts and their interlayer connections in a fully unsuper-vised manner. Our approach simultaneously reveals fine-grained concepts at a layer, connection weightings across all layers and is amendable to global analysis of network structure (e.g. branching pattern of hierarchical concept assemblies). Previous work yielded ways to extract inter-pretable concepts from single layers and examine their im-pact on classification, but did not afford multilayer concept analysis across an entire network architecture. Quantitative and qualitative empirical results show the effectiveness of VCCs in the domain of image classification. Also, we lever-age VCCs for the application of failure mode debugging to reveal where mistakes arise in deep networks.
Matthew Kowal, Richard P. Wildes, Konstantinos G. Derpanis
CVPR2
2023 StepFormer: Self-Supervised Step Discovery and Localization in Instructional Videos
abstract
Instructional videos are an important resource to learn procedural tasks from human demonstrations. However, the instruction steps in such videos are typically short and sparse, with most of the video being irrelevant to the procedure. This motivates the need to temporally localize the instruction steps in such videos, i.e. the task called key-step localization. Traditional methods for key-step localization require video-level human annotations and thus do not scale to large datasets. In this work, we tackle the problem with no human supervision and introduce StepFormer, a self-supervised model that discovers and localizes instruction steps in a video. StepFormer is a transformer decoder that attends to the video with learnable queries, and produces a sequence of slots capturing the key-steps in the video. We train our system on a large dataset of instructional videos, using their automatically-generated subtitles as the only source of supervision. In particular, we supervise our system with a sequence of text narrations using an order-aware loss function that filters out irrelevant phrases. We show that our model outperforms all previous unsupervised and weakly-supervised approaches on step detection and localization by a large margin on three challenging benchmarks. Moreover, our model demonstrates an emergent property to solve zero-shot multi-step localization and outperforms all relevant baselines at this task.
Nikita Dvornik, Isma Hadji, Konstantinos G. Derpanis, Richard P. Wildes, Allan Douglas Jepson
CVPR5
2023 MED-VT: Multiscale Encoder-Decoder Video Transformer with Application to Object Segmentation
abstract
Multiscale video transformers have been explored in a wide variety of vision tasks. To date, however, the multiscale processing has been confined to the encoder or decoder alone. We present a unified multiscale encoder-decoder transformer that is focused on dense prediction tasks in videos. Multiscale representation at both encoder and decoder yields key benefits of implicit extraction of spatiotemporal features (i.e. without reliance on input optical flow) as well as temporal consistency at encoding and coarse-to-fine detection for high-level (e.g. object) semantics to guide precise localization at decoding. Moreover, we propose a transductive learning scheme through many-to-many label propagation to provide temporally consistent predictions. We showcase our Multiscale Encoder-Decoder Video Transformer (MED-VT) on Automatic Video Object Segmentation (AVOS) and actor/action segmentation, where we outperform state-of-the-art approaches on multiple benchmarks using only raw images, without using optical flow.
Rezaul Karim, He Zhao 0004, Richard P. Wildes, Mennatullah Siam
CVPR3
2022 P3IV: Probabilistic Procedure Planning from Instructional Videos with Weak Supervision
abstract
In this paper, we study the problem of procedure planning in instructional videos. Here, an agent must produce a plausible sequence of actions that can transform the environment from a given start to a desired goal state. When learning procedure planning from instructional videos, most recent work leverages intermediate visual observations as supervision, which requires expensive annotation efforts to localize precisely all the instructional steps in training videos. In contrast, we remove the need for expensive temporal video annotations and propose a weakly supervised approach by learning from natural language instructions. Our model is based on a transformer equipped with a memory module, which maps the start and goal observations to a sequence of plausible actions. Furthermore, we augment our model with a probabilistic generative module to capture the uncertainty inherent to procedure planning, an aspect largely overlooked by previous work. We evaluate our model on three datasets and show our weakly-supervised approach outperforms previous fully supervised state-of-the-art models on multiple metrics.
He Zhao 0004, Isma Hadji, Nikita Dvornik, Konstantinos G. Derpanis, Richard P. Wildes, Allan Douglas Jepson
CVPR5
2022 A Deeper Dive Into What Deep Spatiotemporal Networks Encode: Quantifying Static vs. Dynamic Information
abstract
Deep spatiotemporal models are used in a variety of computer vision tasks, such as action recognition and video object segmentation. Currently, there is a limited understanding of what information is captured by these models in their intermediate representations. For example, while it has been observed that action recognition algorithms are heavily influenced by visual appearance in single static frames, there is no quantitative methodology for evaluating such static bias in the latent representation compared to bias toward dynamic information (e.g. motion). We tackle this challenge by proposing a novel approach for quantifying the static and dynamic biases of any spatiotemporal model. To show the efficacy of our approach, we analyse two widely studied tasks, action recognition and video object segmentation. Our key findings are threefold: (i) Most examined spatiotemporal models are biased toward static information; although, certain two-stream architectures with cross-connections show a better balance between the static and dynamic information captured. (ii) Some datasets that are commonly assumed to be biased toward dynamics are actually biased toward static information. (iii) Individual units (channels) in an architecture can be biased toward static, dynamic or a combination of the two.11Project page and code
Matthew Kowal, Mennatullah Siam, Md. Amirul Islam, Neil D. B. Bruce, Richard P. Wildes, Konstantinos G. Derpanis
CVPR5
2022 Is Appearance Free Action Recognition Possible?
Filip Ilic, Thomas Pock, Richard P. Wildes
ECCV (4)3
2022 Sports Video Analysis on Large-Scale Data
Dekun Wu, He Zhao 0004, Xingce Bao, Richard P. Wildes
ECCV (37)4
2021 Where are you heading? Dynamic Trajectory Prediction with Expert Goal Examples
abstract
Goal-conditioned approaches recently have been found very useful to human trajectory prediction, when adequate goal estimates are provided. Yet, goal inference is difficult in itself and often incurs extra learning effort. We propose to predict pedestrian trajectories via the guidance of goal expertise, which can be obtained with modest expense through a novel goal-search mechanism on already seen training examples. There are three key contributions in our study. First, we devise a framework that exploits nearest examples for high-quality goal position inquiry. This approach naturally considers multi-modality, physical constraints, compatibility with existing methods and is nonparametric; it therefore does not require additional learning effort typical in goal inference. Second, we present an end-to-end trajectory predictor that can efficiently associate goal retrievals to past motion information and dynamically infer possible future trajectories. Third, with these two novel techniques in hand, we conduct a series of experiments on two broadly explored datasets (SDD and ETH/UCY) and show that our approach surpasses previous state-of-the-art performance by notable margins and reduces the need for additional parameters. Code can be found at our project page1.
He Zhao 0004, Richard P. Wildes
ICCV2
2020 On Diverse Asynchronous Activity Anticipation
He Zhao 0004, Richard P. Wildes
ECCV (29)2
2020 Graph Neural Net Using Analytical Graph Filters and Topology Optimization for Image Denoising
abstract
While convolutional neural nets (CNNs) have achieved remarkable performance for a wide range of inverse imaging applications, the filter coefficients are computed in a purely data-driven manner and are not explainable. Inspired by an analytically derived CNN by Hadji et al., in this paper we construct a new layered graph neural net (GNN) using GraphBio as our graph filter. Unlike convolutional filters in previous GNNs, our employed GraphBio is analytically defined and requires no training, and we optimize the end-to-end system only via learning of appropriate graph topology at each layer. In signal filtering terms, it means that our linear graph filter at each layer is always intrepretable as low-pass with known biorthogonal conditions, while the graph spectrum itself is optimized via data training. As an example application, we show that our analytical GNN achieves image denoising performance comparable to a state-of-the-art CNN-based scheme when the training and testing data share the same statistics, and when they differ, our analytical GNN outperforms it by more than 1dB in PSNR.
Weng-Tai Su, Gene Cheung, Richard P. Wildes, Chia-Wen Lin
ICASSP3
2020 Deep Insights into Convolutional Networks for Video Recognition
abstract
Abstract As the success of deep models has led to their deployment in all areas of computer vision, it is increasingly important to understand how these representations work and what they are capturing. In this paper, we shed light on deep spatiotemporal representations by visualizing the internal representation of models that have been trained to recognize actions in video. We visualize multiple two-stream architectures to show that local detectors for appearance and motion objects arise to form distributed representations for recognizing human actions. Key observations include the following. First, cross-stream fusion enables the learning of true spatiotemporal features rather than simply separate appearance and motion features. Second, the networks can learn local representations that are highly class specific, but also generic representations that can serve a range of classes. Third, throughout the hierarchy of the network, features become more abstract and show increasing invariance to aspects of the data that are unimportant to desired distinctions (e.g. motion patterns across various speeds). Fourth, visualizations can be used not only to shed light on learned representations, but also to reveal idiosyncrasies of training data and to explain failure cases of the system.
Christoph Feichtenhofer, Axel Pinz, Richard P. Wildes, Andrew Zisserman
Int. J. Comput. Vis.3
2019 Spatiotemporal Feature Residual Propagation for Action Prediction
abstract
Recognizing actions from limited preliminary video observations has seen considerable recent progress. Typically, however, such progress has been had without explicitly modeling fine-grained motion evolution as a potentially valuable information source. In this study, we address this task by investigating how action patterns evolve over time in a spatial feature space. There are three key components to our system. First, we work with intermediate-layer ConvNet features, which allow for abstraction from raw data, while retaining spatial layout, which is sacrificed in approaches that rely on vectorized global representations. Second, instead of propagating features per se, we propagate their residuals across time, which allows for a compact representation that reduces redundancy while retaining essential information about evolution over time. Third, we employ a Kalman filter to combat error build-up and unify across prediction start times. Extensive experimental results on the JHMDB21, UCF101 and BIT datasets show that our approach leads to a new state-of-the-art in action prediction.
He Zhao 0004, Richard P. Wildes
ICCV2
2018 What Have We Learned From Deep Representations for Action Recognition?
abstract
As the success of deep models has led to their deployment in all areas of computer vision, it is increasingly important to understand how these representations work and what they are capturing. In this paper, we shed light on deep spatiotemporal representations by visualizing what two-stream models have learned in order to recognize actions in video. We show that local detectors for appearance and motion objects arise to form distributed representations for recognizing human actions. Key observations include the following. First, cross-stream fusion enables the learning of true spatiotemporal features rather than simply separate appearance and motion features. Second, the networks can learn local representations that are highly class specific, but also generic representations that can serve a range of classes. Third, throughout the hierarchy of the network, features become more Abstract and show increasing invariance to aspects of the data that are unimportant to desired distinctions (e.g. motion patterns across various speeds). Fourth, visualizations can be used not only to shed light on learned representations, but also to reveal idiosyncracies of training data and to explain failure cases of the system. This document is best viewed offline where figures play on click.
Christoph Feichtenhofer, Axel Pinz, Richard P. Wildes, Andrew Zisserman
CVPR3
2018 A New Large Scale Dynamic Texture Dataset with Application to ConvNet Understanding
Isma Hadji, Richard P. Wildes
ECCV (14)2
2017 Temporal Residual Networks for Dynamic Scene Recognition
abstract
This paper combines three contributions to establish a new state-of-the-art in dynamic scene recognition. First, we present a novel ConvNet architecture based on temporal residual units that is fully convolutional in spacetime. Our model augments spatial ResNets with convolutions across time to hierarchically add temporal residuals as the depth of the network increases. Second, existing approaches to video-based recognition are categorized and a baseline of seven previously top performing algorithms is selected for comparative evaluation on dynamic scenes. Third, we introduce a new and challenging video database of dynamic scenes that more than doubles the size of those previously available. This dataset is explicitly split into two subsets of equal size that contain videos with and without camera motion to allow for systematic study of how this variable interacts with the defining dynamics of the scene per se. Our evaluations verify the particular strengths and weaknesses of the baseline algorithms with respect to various scene classes and camera motion parameters. Finally, our temporal ResNet boosts recognition performance and establishes a new state-of-the-art on dynamic scene recognition, as well as on the complementary task of action recognition.
Christoph Feichtenhofer, Axel Pinz, Richard P. Wildes
CVPR3
2017 Spatiotemporal Multiplier Networks for Video Action Recognition
abstract
This paper presents a general ConvNet architecture for video action recognition based on multiplicative interactions of spacetime features. Our model combines the appearance and motion pathways of a two-stream architecture by motion gating and is trained end-to-end. We theoretically motivate multiplicative gating functions for residual networks and empirically study their effect on classification accuracy. To capture long-term dependencies we inject identity mapping kernels for learning temporal relationships. Our architecture is fully convolutional in spacetime and able to evaluate a video in a single forward pass. Empirical investigation reveals that our model produces state-of-the-art results on two standard action recognition datasets.
Christoph Feichtenhofer, Axel Pinz, Richard P. Wildes
CVPR3
2017 A Spatiotemporal Oriented Energy Network for Dynamic Texture Recognition
abstract
This paper presents a novel hierarchical spatiotemporal orientation representation for spacetime image analysis. It is designed to combine the benefits of the multilayer architecture of ConvNets and a more controlled approach to spacetime analysis. A distinguishing aspect of the approach is that unlike most contemporary convolutional networks no learning is involved; rather, all design decisions are specified analytically with theoretical motivations. This approach makes it possible to understand what information is being extracted at each stage and layer of processing as well as to minimize heuristic choices in design. Another key aspect of the network is its recurrent nature, whereby the output of each layer of processing feeds back to the input. To keep the network size manageable across layers, a novel cross-channel feature pooling is proposed. The multilayer architecture that results systematically reveals hierarchical image structure in terms of multiscale, multiorientation properties of visual spacetime. To illustrate its utility, the network has been applied to the task of dynamic texture recognition. Empirical evaluation on multiple standard datasets shows that it sets a new state-of-the-art.
Isma Hadji, Richard P. Wildes
ICCV2
2016 Spatiotemporal Residual Networks for Video Action Recognition
abstract
Two-stream Convolutional Networks (ConvNets) have shown strong performance for human action recognition in videos. Recently, Residual Networks (ResNets) have arisen as a new technique to train extremely deep architectures. In this paper, we introduce spatiotemporal ResNets as a combination of these two approaches. Our novel architecture generalizes ResNets for the spatiotemporal domain by introducing residual connections in two ways. First, we inject residual connections between the appearance and motion pathways of a two-stream architecture to allow spatiotemporal interaction between the two streams. Second, we transform pretrained image ConvNets into spatiotemporal networks by equipping these with learnable convolutional filters that are initialized as temporal residual connections and operate on adjacent feature maps in time. This approach slowly increases the spatiotemporal receptive field as the depth of the model increases and naturally integrates image ConvNet design principles. The whole model is trained end-to-end to allow hierarchical learning of complex spatiotemporal features. We evaluate our novel spatiotemporal ResNet using two widely used action recognition benchmarks where it exceeds the previous state-of-the-art.
Christoph Feichtenhofer, Axel Pinz, Richard P. Wildes
NIPS3
2016 Dynamic Scene Recognition with Complementary Spatiotemporal Features
abstract
This paper presents Dynamically Pooled Complementary Features (DPCF), a unified approach to dynamic scene recognition that analyzes a short video clip in terms of its spatial, temporal and color properties. The complementarity of these properties is preserved through all main steps of processing, including primitive feature extraction, coding and pooling. In the feature extraction step, spatial orientations capture static appearance, spatiotemporal oriented energies capture image dynamics and color statistics capture chromatic information. Subsequently, primitive features are encoded into a mid-level representation that has been learned for the task of dynamic scene recognition. Finally, a novel dynamic spacetime pyramid is introduced. This dynamic pooling approach can handle both global as well as local motion by adapting to the temporal structure, as guided by pooling energies. The resulting system provides online recognition of dynamic scenes that is thoroughly evaluated on the two current benchmark datasets and yields best results to date on both datasets. In-depth analysis reveals the benefits of explicitly modeling feature complementarity in combination with the dynamic spacetime pyramid, indicating that this unified approach should be well-suited to many areas of video analysis.
Christoph Feichtenhofer, Axel Pinz, Richard P. Wildes
IEEE Trans. Pattern Anal. Mach. Intell.3
2015 Dynamically encoded actions based on spacetime saliency
abstract
Human actions typically occur over a well localized extent in both space and time. Similarly, as typically captured in video, human actions have small spatiotemporal support in image space. This paper capitalizes on these observations by weighting feature pooling for action recognition over those areas within a video where actions are most likely to occur. To enable this operation, we define a novel measure of spacetime saliency. The measure relies on two observations regarding foreground motion of human actors: They typically exhibit motion that contrasts with that of their surrounding region and they are spatially compact. By using the resulting definition of saliency during feature pooling we show that action recognition performance achieves state-of-the-art levels on three widely considered action recognition datasets. Our saliency weighted pooling can be applied to essentially any locally defined features and encodings thereof. Additionally, we demonstrate that inclusion of locally aggregated spatiotemporal energy features, which efficiently result as a by-product of the saliency computation, further boosts performance over reliance on standard action recognition features alone.
Christoph Feichtenhofer, Axel Pinz, Richard P. Wildes
CVPR3
2014 Bags of Spacetime Energies for Dynamic Scene Recognition
abstract
This paper presents a unified bag of visual word (BoW) framework for dynamic scene recognition. The approach builds on primitive features that uniformly capture spatial and temporal orientation structure of the imagery (e.g., video), as extracted via application of a bank of spatiotemporally oriented filters. Various feature encoding techniques are investigated to abstract the primitives to an intermediate representation that is best suited to dynamic scene representation. Further, a novel approach to adaptive pooling of the encoded features is presented that captures spatial layout of the scene even while being robust to situations where camera motion and scene dynamics are confounded. The resulting overall approach has been evaluated on two standard, publically available dynamic scene datasets. The results show that in comparison to a representative set of alternatives, the proposed approach outperforms the previous state-of-the-art in classification accuracy by 10%.
Christoph Feichtenhofer, Axel Pinz, Richard P. Wildes
CVPR3
2014 The Applicability of Spatiotemporal Oriented Energy Features to Region Tracking
abstract
This paper proposes the novel application of an uncommonly rich feature representation to the domain of visual tracking. The proposed representation for tracking models both the spatial structure and dynamics of a target in a unified fashion, while simultaneously offering robustness to illumination variations. Specifically, the proposed feature is derived from spatiotemporal energy measurements that are computed by filtering in 3D, (x, y, t), image spacetime. These spatiotemporal energy measurements capture the underlying local spacetime orientation structure of the target across multiple scales. The breadth of applicability of these features within the field of visual tracking is demonstrated by their instantiation within three disparate tracking paradigms that are representative of the various basic types of region trackers in the field. Instantiation within these three tracking paradigms requires that the raw oriented energy measurements be post-processed using different methodologies that range from histogram accumulation to the identity transform. Qualitative and quantitative empirical evaluation on a challenging suite of videos demonstrates the strength and applicability of the proposed representation to tracking, as it outperforms other commonly-used features across all tracking paradigms. Moreover, it is shown that overall high tracking accuracy can be obtained with this proposed representation, as spatiotemporal oriented energy instantiations are shown to outperform several recent, state-of-the-art trackers.
Kevin J. Cannons, Richard P. Wildes
IEEE Trans. Pattern Anal. Mach. Intell.2
2014 Spacetime Stereo and 3D Flow via Binocular Spatiotemporal Orientation Analysis
abstract
This paper presents a novel approach to recovering estimates of 3D structure and motion of a dynamic scene from a sequence of binocular stereo images. The approach is based on matching spatiotemporal orientation distributions between left and right temporal image streams, which encapsulates both local spatial and temporal structure for disparity estimation. By capturing spatial and temporal structure in this unified fashion, both sources of information combine to yield disparity estimates that are naturally temporal coherent, while helping to resolve matches that might be ambiguous when either source is considered alone. Further, by allowing subsets of the orientation measurements to support different disparity estimates, an approach to recovering multilayer disparity from spacetime stereo is realized. Similarly, the matched distributions allow for direct recovery of dense, robust estimates of 3D scene flow. The approach has been implemented with real-time performance on commodity GPUs using OpenCL. Empirical evaluation shows that the proposed approach yields qualitatively and quantitatively superior estimates in comparison to various alternative approaches, including the ability to provide accurate multilayer estimates in the presence of (semi)transparent and specular surfaces.
Mikhail Sizintsev, Richard P. Wildes
IEEE Trans. Pattern Anal. Mach. Intell.2
2013 Spacetime Forests with Complementary Features for Dynamic Scene Recognition
abstract
This paper presents spacetime forests defined over complementary spatial and temporal features for recognition of naturally occurring dynamic scenes. The approach improves on the previous state-of-the-art in both classification and execution rates. A particular improvement is with increased robustness to camera motion, where previous approaches have experienced difficulty. There are three key novelties in the approach. First, a novel spacetime descriptor is employed that exploits the complementary nature of spatial and temporal information, as inspired by previous research on the role of orientation features in scene classification. Second, a forest-based classifier is used to learn a multi-class representation of the feature distributions. Third, the video is processed in temporal slices with scale matched preferentially to scene dynamics over camera motion. Slicing allows for temporal alignment to be handled as latent information in the classifier and for efficient, incremental processing. The integrated approach is evaluated empirically on two publically available datasets to document its outstanding performance.
Christoph Feichtenhofer, Axel Pinz, Richard P. Wildes
BMVC3
2013 Egomotion Estimation Using Binocular Spatiotemporal Oriented Energy
abstract
Camera egomotion estimation is concerned with the recovery of a camera's motion (e.g., instantaneous translation and rotation) as it moves through its environment.It has been demonstrated to be of both theoretical and practical interest.This paper documents a novel algorithm for egomotion estimation based on binocularly matched spatiotemporal oriented energy distributions.Basing the estimation on oriented energy measurements makes it possible to recover egomotion without the need to establish temporal correspondences or convert disparity into 3D world coordinates.The resulting algorithm has been realized in software and evaluated quantitatively on a novel laboratory dataset with groundtruth as well as qualitatively on both indoor and outdoor real-world datasets.Performance is evaluated relative to comparable alternative algorithms and shown to exhibit best overall performance.
Richard P. Wildes
BMVC2
2013 Action Spotting and Recognition Based on a Spatiotemporal Orientation Analysis
abstract
This paper provides a unified framework for the interrelated topics of action spotting, the spatiotemporal detection and localization of human actions in video, and action recognition, the classification of a given video into one of several predefined categories. A novel compact local descriptor of video dynamics in the context of action spotting and recognition is introduced based on visual spacetime oriented energy measurements. This descriptor is efficiently computed directly from raw image intensity data and thereby forgoes the problems typically associated with flow-based features. Importantly, the descriptor allows for the comparison of the underlying dynamics of two spacetime video segments irrespective of spatial appearance, such as differences induced by clothing, and with robustness to clutter. An associated similarity measure is introduced that admits efficient exhaustive search for an action template, derived from a single exemplar video, across candidate video sequences. The general approach presented for action spotting and recognition is amenable to efficient implementation, which is deemed critical for many important applications. For action spotting, details of a real-time GPU-based instantiation of the proposed approach are provided. Empirical evaluation of both action spotting and action recognition on challenging datasets suggests the efficacy of the proposed approach, with state-of-the-art performance documented on standard datasets.
Konstantinos G. Derpanis, Mikhail Sizintsev, Kevin J. Cannons, Richard P. Wildes
IEEE Trans. Pattern Anal. Mach. Intell.4
2012 Spatiotemporal Salience via Centre-Surround Comparison of Visual Spacetime Orientations
Andrei Zaharescu, Richard P. Wildes
ACCV (3)2
2012 Dynamic scene understanding: The role of orientation features in space and time in scene classification
abstract
Natural scene classification is a fundamental challenge in computer vision. By far, the majority of studies have limited their scope to scenes from single image stills and thereby ignore potentially informative temporal cues. The current paper is concerned with determining the degree of performance gain in considering short videos for recognizing natural scenes. Towards this end, the impact of multiscale orientation measurements on scene classification is systematically investigated, as related to: (i) spatial appearance, (ii) temporal dynamics and (iii) joint spatial appearance and dynamics. These measurements in visual space, x-y, and spacetime, x-y-t, are recovered by a bank of spatiotemporal oriented energy filters. In addition, a new data set is introduced that contains 420 image sequences spanning fourteen scene categories, with temporal scene information due to objects and surfaces decoupled from camera-induced ones. This data set is used to evaluate classification performance of the various orientation-related representations, as well as state-of-the-art alternatives. It is shown that a notable performance increase is realized by spatiotemporal approaches in comparison to purely spatial or purely temporal methods.
Konstantinos G. Derpanis, Matthieu Lecce, Kostas Daniilidis, Richard P. Wildes
CVPR4
2012 Spacetime Texture Representation and Recognition Based on a Spatiotemporal Orientation Analysis
abstract
This paper is concerned with the representation and recognition of the observed dynamics (i.e., excluding purely spatial appearance cues) of spacetime texture based on a spatiotemporal orientation analysis. The term "spacetime texture" is taken to refer to patterns in visual spacetime, (x,y,t), that primarily are characterized by the aggregate dynamic properties of elements or local measurements accumulated over a region of spatiotemporal support, rather than in terms of the dynamics of individual constituents. Examples include image sequences of natural processes that exhibit stochastic dynamics (e.g., fire, water, and windblown vegetation) as well as images of simpler dynamics when analyzed in terms of aggregate region properties (e.g., uniform motion of elements in imagery, such as pedestrians and vehicular traffic). Spacetime texture representation and recognition is important as it provides an early means of capturing the structure of an ensuing image stream in a meaningful fashion. Toward such ends, a novel approach to spacetime texture representation and an associated recognition method are described based on distributions (histograms) of spacetime orientation structure. Empirical evaluation on both standard and original image data sets shows the promise of the approach, including significant improvement over alternative state-of-the-art approaches in recognizing the same pattern from different viewpoints.
Konstantinos G. Derpanis, Richard P. Wildes
IEEE Trans. Pattern Anal. Mach. Intell.2
2012 Spatiotemporal Stereo and Scene Flow via Stequel Matching
abstract
This paper is concerned with the recovery of temporally coherent estimates of 3D structure and motion of a dynamic scene from a sequence of binocular stereo images. A novel approach is presented based on matching of spatiotemporal quadric elements (stequels) between views, as this primitive encapsulates both spatial and temporal image structure for 3D estimation. Match constraints are developed for bringing stequels into correspondence across binocular views. With correspondence established, temporally coherent disparity estimates are obtained without explicit motion recovery. Further, the matched stequels also will be shown to support direct recovery of scene flow estimates. Extensive algorithmic evaluation with ground truth data incorporated in both local and global correspondence paradigms shows the considerable benefit of using stequels as a matching primitive and its advantages in comparison to alternative methods of enforcing temporal coherence in disparity estimation. Additional experiments document the usefulness of stequel matching for 3D scene flow estimation.
Mikhail Sizintsev, Richard P. Wildes
IEEE Trans. Pattern Anal. Mach. Intell.2
2011 Spatiotemporal oriented energies for spacetime stereo
abstract
This paper presents a novel approach to recovering temporally coherent estimates of 3D structure of a dynamic scene from a sequence of binocular stereo images. The approach is based on matching spatiotemporal orientation distributions between left and right temporal image streams, which encapsulates both local spatial and temporal structure for disparity estimation. By capturing spatial and temporal structure in this unified fashion, both sources of information combine to yield disparity estimates that are naturally temporal coherent, while helping to resolve matches that might be ambiguous when either source is considered alone. Further, by allowing subsets of the orientation measurements to support different disparity estimates, an approach to recovering multilayer disparity from spacetime stereo is realized. The approach has been implemented with real-time performance on commodity GPUs. Empirical evaluation shows that the approach yields qualitatively and quantitatively superior disparity estimates in comparison to various alternative approaches, including the ability to provide accurate multilayer estimates in the presence of (semi)transparent and specular surfaces.
Mikhail Sizintsev, Richard P. Wildes
ICCV2
2011 Classification of traffic video based on a spatiotemporal orientation analysis
abstract
This paper describes a system for classifying traffic congestion videos based on their observed visual dynamics. Central to the proposed system is treating traffic flow identification as an instance of dynamic texture classification. More specifically, a recent discriminative model of dynamic textures is adapted for the special case of traffic flows. This approach avoids the need for segmentation, tracking and motion estimation that typify extant approaches. Classification is based on matching distributions (or histograms) of spacetime orientation structure. Empirical evaluation on a publicly available data set shows high classification performance and robustness to typical environmental conditions (e.g., variable lighting).
Konstantinos G. Derpanis, Richard P. Wildes
WACV2
2010 Efficient action spotting based on a spacetime oriented structure representation
abstract
This paper addresses action spotting, the spatiotemporal detection and localization of human actions in video. A novel compact local descriptor of video dynamics in the context of action spotting is introduced based on visual spacetime oriented energy measurements. This descriptor is efficiently computed directly from raw image intensity data and thereby forgoes the problems typically associated with flow-based features. An important aspect of the descriptor is that it allows for the comparison of the underlying dynamics of two spacetime video segments irrespective of spatial appearance, such as differences induced by clothing, and with robustness to clutter. An associated similarity measure is introduced that admits efficient exhaustive search for an action template across candidate video sequences. Empirical evaluation of the approach on a set of challenging natural videos suggests its efficacy.
Konstantinos G. Derpanis, Mikhail Sizintsev, Kevin J. Cannons, Richard P. Wildes
CVPR4
2010 Dynamic texture recognition based on distributions of spacetime oriented structure
abstract
This paper addresses the challenge of recognizing dynamic textures based on their observed visual dynamics. Typically, the term dynamic texture is used with reference to image sequences of various natural processes that exhibit stochastic dynamics (e.g., smoke, water and windblown vegetation); although, it applies equally well to images of simpler dynamics when analyzed in terms of aggregate region properties (e.g., uniform motion of elements in traffic video). In this paper, a novel approach to dynamic texture representation and an associated recognition method are proposed. The approach pursued here recognizes dynamic textures based on matching distributions (histograms) of spacetime orientation structure. Empirical evaluation on a standard database with controls to remove the effects of identical viewpoint demonstrates that the proposed approach achieves superior performance over alternative state-of-the-art methods.
Konstantinos G. Derpanis, Richard P. Wildes
CVPR2
2010 Visual Tracking Using a Pixelwise Spatiotemporal Oriented Energy Representation
Kevin J. Cannons, Jacob M. Gryn, Richard P. Wildes
ECCV (4)3
2010 Anomalous Behaviour Detection Using Spatiotemporal Oriented Energies, Subset Inclusion Histogram Comparison and Event-Driven Processing
Andrei Zaharescu, Richard P. Wildes
ECCV (1)2
2010 Coarse-to-fine stereo vision with accurate 3D boundaries
Mikhail Sizintsev, Richard P. Wildes
Image Vis. Comput.2
2010 The Structure of Multiplicative Motions in Natural Imagery
abstract
A theoretical investigation of the frequency structure of multiplicative image motion signals is presented, e.g., as associated with translucency phenomena. Previous work has claimed that the multiplicative composition of visual signals generally results in the annihilation of oriented structure in the spectral domain. As a result, research has focused on multiplicative signals in highly specialized scenarios where highly structured spectral signatures are prevalent, or introduced a nonlinearity to transform the multiplicative image signal to an additive one. In contrast, in this paper, it is shown that oriented structure is present in multiplicative cases when natural domain constraints are taken into account. This analysis suggests that the various instances of naturally occurring multiple motion structures can be treated in a unified manner. As an example application of the developed theory, a multiple motion estimator previously proposed for translation, additive transparency, and occlusion is adapted to multiplicative image motions. This estimator is shown to yield superior performance over the alternative practice of introducing a nonlinear preprocessing step.
Konstantinos G. Derpanis, Richard P. Wildes
IEEE Trans. Pattern Anal. Mach. Intell.2
2009 Detecting Spatiotemporal Structure Boundaries: Beyond Motion Discontinuities
Konstantinos G. Derpanis, Richard P. Wildes
ACCV (2)2
2009 Early spatiotemporal grouping with a distributed oriented energy representation
abstract
Spatiotemporal data is associated with vast amounts of raw samples. Given the limited computational resources typically available, an initial organization of this data supporting semantically meaningful lines of inquiry would facilitate efficient processing. In this paper, a new representation for grouping raw image data into a set of coherent spacetime regions is proposed. Unique in this proposal is that coherency is related to a richer description of local spacetime structure than generally considered. In particular, the representation describes the presence of particular oriented spacetime structures in a distributed manner. A key advantage of this representation is its ability to signal the presence of multiple oriented structures at a given spacetime location. More generally, the abstraction allows for the description and grouping of motion and non-motion-related patterns in a uniform manner. Empirical evaluation of the grouping method on synthetic and challenging natural imagery suggests its efficacy.
Konstantinos G. Derpanis, Richard P. Wildes
CVPR2
2009 Spatiotemporal stereo via spatiotemporal quadric element (stequel) matching
abstract
Spatiotemporal stereo is concerned with the recovery of the 3D structure of a dynamic scene from a temporal sequence of multiview images. This paper presents a novel method for computing temporally coherent disparity maps from a sequence of binocular images through an integrated consideration of image spacetime structure and without explicit recovery of motion. The approach is based on matching spatiotemporal quadric elements (stequels) between views, as it is shown that this matching primitive provides a natural way to encapsulate both local spatial and temporal structure for disparity estimation. Empirical evaluation with laboratory based imagery with ground truth and more typical natural imagery shows that the approach provides considerable benefit in comparison to alternative methods for enforcing temporal coherence in disparity estimation.
Mikhail Sizintsev, Richard P. Wildes
CVPR2
2009 Detecting motion patterns via direction maps with application to surveillance
Jacob M. Gryn, Richard P. Wildes, John K. Tsotsos
Comput. Vis. Image Underst.2
2008 Definition and recovery of kinematic features for recognition of American sign language movements
Konstantinos G. Derpanis, Richard P. Wildes, John K. Tsotsos
Image Vis. Comput.2
2007 Spatiotemporal Oriented Energy Features for Visual Tracking
Kevin J. Cannons, Richard P. Wildes
ACCV (1)2
2006 Efficient Stereo with Accurate 3-D Boundaries
abstract
This paper presents methods for recovering accurate binocular disparity es-timates in the vicinity of 3D surface discontinuities. Of particular concern are methods that impact coarse-to-fine, block matching as it forms the ba-sis of the fastest and resource efficient disparity estimation procedures. Two advances are put forth. First, a novel approach to coarse-to-fine processing is presented that adapts match window support across scale to ameliorate corruption of disparity estimates near 3D boundaries. Second, a novel for-mulation of half-occlusion cues within the coarse-to-fine, block matching framework is described to inhibit false matches that can arise in regions near occlusions. Empirical results show that incorporation of these advances in coarse-to-fine, block matching reduces disparity errors by more than a factor of two, while performing little extra computation. 1
Mikhail Sizintsev, Richard P. Wildes
BMVC2
2004 Hand Gesture Recognition within a Linguistics-Based Framework
Konstantinos G. Derpanis, Richard P. Wildes, John K. Tsotsos
ECCV (1)2
2004 A stereo confidence metric using single view imagery with comparison to five alternative approaches
Geoffrey Egnal, Max Mintz, Richard P. Wildes
Image Vis. Comput.3
2002 Detecting Binocular Half-Occlusions: Empirical Comparisons of Five Approaches
abstract
Binocular half-occlusion points are those that are visible in one of the two views provided by a binocular imaging system. Due to their importance in binocular matching as well as, subsequent interpretation tasks, a number of approaches have been developed for dealing with such points. In the current paper, we consider five methods that explicitly detect half-occlusions and report on a more uniform comparison than has previously been performed. Taking a disparity image and its associated match goodness image as input, we generate images that show the half-occluded points in the underlying scene. We quantitatively and qualitatively compare these methods under a variety of conditions.
Geoffrey Egnal, Richard P. Wildes
IEEE Trans. Pattern Anal. Mach. Intell.2
2001 Video to Reference Image Alignment in the Presence of Sparse Features and Appearance Change
abstract
A robust, multi-frame, progressive refinement framework for registering narrow field of view video to reference imagery is presented. A major strength of the approach is its effectiveness in the presence of dissimilar video and reference image appearance. Normalized oriented energy image pyramids are employed to enable alignment of images with global visual dissimilarities, yet local feature commonality. Local matching is then applied coarse-to-fine, along four dimensions: spatial frequency, local support, search range, and model order (a robust parametric model fit is used to reject outliers at each iteration). Globally optimal multi-frame alignment is obtained with respect to several constraints: frame-to-reference local matches, recovered frame-to-frame motion, and optional a priori estimates of sensor pose. The framework is described in detail and applied to two examples: aerial video to geographic reference image alignment (georegistration) and retinal slit lamp video to fundus image alignment.
David J. Hirvonen, Bogdan Matei, Richard P. Wildes, Steven C. Hsu
CVPR (2)3
2001 Video Georegistration: Algorithm and Quantitative Evaluation
abstract
An algorithm is presented for video georegistration, with a particular concern for aerial video, i.e., video captured from an airborne platform. The algorithm's input is a video stream with telemetry (camera model specification sufficient to define an initial estimate of the view) and geodetically calibrated reference imagery (coaligned digital orthoimage and elevation map). The output is a spatial registration of the video to the reference so that it inherits the available geodetic coordinates. The video is processed in a continuous fashion to yield a corresponding stream of georegistered results. Quantitative results of evaluating the developed approach with real world aerial video also are presented. The results suggest that the developed approach may provide valuable input to the analysis and interpretation of aerial video.
Richard P. Wildes, David J. Hirvonen, Steven C. Hsu, Rakesh Kumar 0001, W. Brian Lehman, Bogdan Matei, Wenyi Zhao
ICCV1
2001 Aerial video surveillance and exploitation
abstract
There is growing interest in performing aerial surveillance using video cameras. Compared to traditional framing cameras, video cameras provide the capability to observe ongoing activity within a scene and to automatically control the camera to track the activity. However, the high data rates and relatively small field of view of video cameras present new technical challenges that must be overcome before such cameras can be widely used. In this paper, we present a framework and details of the key components for real-time, automatic exploitation of aerial video for surveillance applications. The framework involves separating an aerial video into the natural components corresponding to the scene. Three major components of the scene are the static background geometry, moving objects, and appearance of the static and dynamic components of the scene. In order to delineate videos into these scene components, we have developed real time, image-processing techniques for 2-D/3-D frame-to-frame alignment, change detection, camera control, and tracking of independently moving objects in cluttered scenes. The geo-location of video and tracked objects is estimated by registration of the video to controlled reference imagery, elevation maps, and site models. Finally static, dynamic and reprojected mosaics may be constructed for compression, enhanced visualization, and mapping applications.
Rakesh Kumar 0001, Harpreet Sawhney, Supun Samarasekera, Steven C. Hsu, Yanlin Guo, Keith J. Hanna, Art Pope, Richard P. Wildes, David J. Hirvonen, Michael W. Hansen, Peter Burt
Proc. IEEE9
2000 Detecting Binocular Half-Occlusions: Empirical Comparisons of Four Approaches
abstract
Binocular half-occlusion points are those that are visible in one of the two views provided by a binocular imaging system. Due to their importance in binocular matching as well as subsequent interpretation tasks, a number of approaches have been developed for dealing with such points. In the current paper, we consider four methods that explicitly detect half-occlusions and report on a more uniform comparison than has previously been performed. Taking a disparity image and its associated match goodness image as input, we generate images that show the half-occluded points in the underlying scene. We quantitatively and qualitatively compare these methods under a variety of conditions.
Geoffrey Egnal, Richard P. Wildes
CVPR2
2000 Qualitative Spatiotemporal Analysis Using an Oriented Energy Representation
Richard P. Wildes, James R. Bergen
ECCV (2)1
2000 Recovering Estimates of Fluid Flow from Image Sequence Data
Richard P. Wildes, Michael J. Amabile, Ann-Marie Lanzillotto, Tzong-Shyng Leu
Comput. Vis. Image Underst.1
1998 A Measure of Motion Salience for Surveillance Applications
Richard P. Wildes
ICIP (3)1
1997 Physically based fluid flow recovery from image sequences
abstract
This paper presents an approach to measuring fluid flow from image sequences. The approach centers around a motion recovery algorithm that is based on principles from fluid mechanics: The algorithm is constrained so that recovered flows observe conservation of mass as well as physically motivated boundary conditions. Results are presented from application of the algorithm to transmittance imagery of fluid flows, where the fluids contained a contrast medium. In these experiments, the algorithm recovered accurate and precise estimates of the flow. The significance of this work is two fold. First, from a theoretical point of view it is shown how information derived from the physical behavior of fluids can be used to motivate a flow recovery algorithm. Second, from an applications point of view the developed algorithm can be used to augment the tools that are available for the measurement of fluid dynamics; other imaged flows that observe compatible constraints might benefit in a similar fashion.
Richard P. Wildes, Michael J. Amabile, Ann-Marie Lanzillotto, Tzong-Shyng Leu
CVPR1
1997 Iris recognition: an emerging biometric technology
abstract
This paper examines automated iris recognition as a biometrically based technology for personal identification and verification. The motivation for this endeavor stems from the observation that the human iris provides a particularly interesting structure on which to base a technology for noninvasive biometric assessment. In particular the biomedical literature suggests that irises are as distinct as fingerprints or patterns of retinal blood vessels. Further, since the iris is an overt body, its appearance is amenable to remote examination with the aid of a machine vision system. The body of this paper details issues in the design and operation of such systems. For the sake of illustration, extant systems are described in some amount of detail.
Richard P. Wildes
Proc. IEEE1
1996 A machine-vision system for iris recognition
Richard P. Wildes, Jane C. Asmuth, Gilbert L. Green, Steven C. Hsu, Raymond J. Kolczynski, James R. Matey, Sterling E. McBride
Mach. Vis. Appl.1
1994 Singularities of the visual motion field: 3D rotation or 3D translation
abstract
This paper addresses the manner in which the structure and motion of the 3D world are manifest in the stable singularities of its imaged visual motion field. By focusing on stable structures the approach enjoys a degree of robustness and invariance to the particulars of recovering visual motion. The specific results to be described address situations where an optical sensor is undergoing pure rotational or pure translational motion through its environment. For the case of pure rotation it is shown that the singularities provide information about the axes, senses and relative magnitudes of rotation. For the case of pure translation it is shown that the singularities provide information about the shape of viewed surfaces as well as information about the translation itself. Finally, empirical results demonstrate the feasibility of applying the analysis to natural images.
Richard P. Wildes
ICPR (1)1
1994 A system for automated iris recognition
abstract
This paper describes a prototype system for personnel verification based on automated iris recognition. The motivation for this endeavour stems from the observation that the human iris provides a particularly interesting structure on which to base a technology for noninvasive biometric measurement. In particular, it is known in the biomedical community that irises are as distinct as fingerprints or patterns of retinal blood vessels. Further, since the iris is an overt body its appearance is amenable to remote examination with the aid of a computer vision system. The body of this paper details the design and operation of such a system. Also presented are the results of an empirical study where the system exhibits flawless performance in the evaluation of 520 iris images.>
Richard P. Wildes, Jane C. Asmuth, Gilbert L. Green, Stephen C. Hsu, Raymond J. Kolczynski, James R. Matey, Sterling E. McBride
WACV1
1993 On the Qualitative Structure of Temporally Evolving Visual Motion Fields
Richard P. Wildes
AAAI1
1991 Direct Recovery of Three-Dimensional Scene Geometry From Binocular Stereo Disparity
abstract
An analysis of disparity is presented. It makes explicit the geometric relations between a stereo disparity field and a differentially project scene. These results show how it is possible to recover three-dimensional surface geometry through first-order (i.e., distance and orientation of a surface relative to an observer) and binocular viewing parameters in a direct fashion from stereo disparity. As applications of the analysis, algorithms have been developed for recovering three-dimensional surface orientation and discontinuities from stereo disparity. The results of applying these algorithms to natural image binocular stereo disparity information are presented.>
Richard P. Wildes
IEEE Trans. Pattern Anal. Mach. Intell.1