EDBT 2026 Demo / reviewers in the wild / expert
Axel Pinz
dblp:00/5595
· DBLP profile ↗
63ranked-venue papers
7as first author
0since 2021 · last 2020
0000-0001-7914-619XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 48 · 5 first-authorGraphics, computer vision, multimedia, augmented reality and games · 44 · 5 first-authorApplied, interdisciplinary, general and emerging computing · 3 · 1 first-authorSystems, architecture and hardware · 2Databases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
17 papers |
Video understanding and tracking · 56% Trustworthy machine learning · 17% Deep learning architectures and training · 13% | |
| Computer graphics and multimedia
4 papers |
Image and video processing · 42% Multimedia analysis and retrieval · 40% Computational photography and imaging · 15% |
Topics — the 30 heaviest of 35, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Video understanding and tracking
action recognition |
1.8 | 7 | 2018 | What Have We Learned From Deep Representations for Action Recognition? · CVPR 2018 Spatiotemporal Multiplier Networks for Video Action Recognition · CVPR 2017 Temporal Residual Networks for Dynamic Scene Recognition · CVPR 2017 |
Machine learning › Trustworthy machine learning
interpretability |
0.8 | 2 | 2020 | Deep Insights into Convolutional Networks for Video Recognition · Int. J. Comput. Vis. 2020 What Have We Learned From Deep Representations for Action Recognition? · CVPR 2018 |
Computer vision › Video understanding and tracking › dynamic scene analysis › video scene understanding
dynamic scene recognition |
0.5 | 2 | 2017 | Temporal Residual Networks for Dynamic Scene Recognition · CVPR 2017 Bags of Spacetime Energies for Dynamic Scene Recognition · CVPR 2014 |
Machine learning › Trustworthy machine learning › interpretability › visual explanation
deep network visualization |
0.4 | 1 | 2020 | Deep Insights into Convolutional Networks for Video Recognition · Int. J. Comput. Vis. 2020 |
Computer vision › Video understanding and tracking
video classification |
0.4 | 1 | 2020 | Deep Insights into Convolutional Networks for Video Recognition · Int. J. Comput. Vis. 2020 |
Computer vision › Video understanding and tracking › multi-object tracking
object detection and tracking |
0.3 | 1 | 2017 | Detect to Track and Track to Detect · ICCV 2017 |
Computer vision › Video understanding and tracking › spatio-temporal modeling
spatiotemporal CNN |
0.3 | 1 | 2017 | Temporal Residual Networks for Dynamic Scene Recognition · CVPR 2017 |
Computer vision › Video understanding and tracking › video object detection
spatio-temporal object detection |
0.3 | 1 | 2017 | Detect to Track and Track to Detect · ICCV 2017 |
Machine learning › Deep learning architectures and training
two-stream network |
0.3 | 1 | 2017 | Spatiotemporal Multiplier Networks for Video Action Recognition · CVPR 2017 |
Computer vision › Image recognition and object detection
object detection |
0.3 | 4 | 2008 | Learning an Alphabet of Shape and Appearance for Multi-Class Object Detection · Int. J. Comput. Vis. 2008 A Boundary-Fragment-Model for Object Detection · ECCV (2) 2006 Incremental learning of object detectors using a visual shape alphabet · CVPR (1) 2006 |
Machine learning › Deep learning architectures and training › convolutional neural network
convolutional neural network architecture |
0.2 | 1 | 2016 | Spatiotemporal Residual Networks for Video Action Recognition · NIPS 2016 |
Machine learning › Deep learning architectures and training › skip connections
residual connection |
0.2 | 1 | 2016 | Spatiotemporal Residual Networks for Video Action Recognition · NIPS 2016 |
Computer vision › Video understanding and tracking › spatio-temporal modeling
spatiotemporal fusion |
0.2 | 1 | 2016 | Convolutional Two-Stream Network Fusion for Video Action Recognition · CVPR 2016 |
Multimedia analysis and retrieval
video analysis |
0.2 | 1 | 2016 | Dynamic Scene Recognition with Complementary Spatiotemporal Features · IEEE Trans. Pattern Anal. Mach. Intell. 2016 |
Machine learning › Representation and self-supervised learning › visual representation › image representation
bag of visual words |
0.2 | 1 | 2014 | Bags of Spacetime Energies for Dynamic Scene Recognition · CVPR 2014 |
Computer vision › Video understanding and tracking › spatio-temporal modeling
spatio-temporal feature encoding |
0.2 | 1 | 2014 | Bags of Spacetime Energies for Dynamic Scene Recognition · CVPR 2014 |
Computer vision › Image recognition and object detection › object detection
multi-class object detection |
0.1 | 2 | 2008 | Learning an Alphabet of Shape and Appearance for Multi-Class Object Detection · Int. J. Comput. Vis. 2008 Incremental learning of object detectors using a visual shape alphabet · CVPR (1) 2006 |
Machine learning › Deep learning architectures and training
convolutional neural network |
0.1 | 1 | 2020 | Deep Insights into Convolutional Networks for Video Recognition · Int. J. Comput. Vis. 2020 |
Computational photography and imaging › camera geometry › camera motion estimation
relative pose estimation |
0.1 | 1 | 2007 | Influence of numerical conditioning on the accuracy of relative orientation · CVPR 2007 |
Computer vision › 3D vision
camera pose estimation |
0.1 | 1 | 2006 | Robust Pose Estimation from a Planar Target · IEEE Trans. Pattern Anal. Mach. Intell. 2006 |
Computer vision › Image recognition and object detection
object recognition |
0.1 | 1 | 2006 | Generic Object Recognition with Boosting · IEEE Trans. Pattern Anal. Mach. Intell. 2006 |
Computer vision › Image recognition and object detection › object recognition
weakly supervised object recognition |
0.1 | 1 | 2006 | Generic Object Recognition with Boosting · IEEE Trans. Pattern Anal. Mach. Intell. 2006 |
Computer vision › Image recognition and object detection › image classification
object classification |
0.0 | 1 | 2008 | Learning an Alphabet of Shape and Appearance for Multi-Class Object Detection · Int. J. Comput. Vis. 2008 |
Computer vision › 3D vision › local feature descriptor
ordinal measures |
0.0 | 1 | 1999 | The Discriminatory Power of Ordinal Measures - Towards a New Coefficient · CVPR 1999 |
Computer vision › 3D vision › stereo vision
shape from stereo |
0.0 | 1 | 1999 | The Discriminatory Power of Ordinal Measures - Towards a New Coefficient · CVPR 1999 |
Computer vision › 3D vision › stereo vision
stereo matching |
0.0 | 1 | 1999 | The Discriminatory Power of Ordinal Measures - Towards a New Coefficient · CVPR 1999 |
Computer vision › 3D vision
stereo vision |
0.0 | 1 | 1999 | The Discriminatory Power of Ordinal Measures - Towards a New Coefficient · CVPR 1999 |
Computational photography and imaging
camera calibration |
0.0 | 1 | 2007 | Influence of numerical conditioning on the accuracy of relative orientation · CVPR 2007 |
Machine learning › Learning paradigms
incremental learning |
0.0 | 1 | 2006 | Incremental learning of object detectors using a visual shape alphabet · CVPR (1) 2006 |
Geometric modeling and processing
shape representation |
0.0 | 1 | 2006 | A Boundary-Fragment-Model for Object Detection · ECCV (2) 2006 |
Methods — techniques the papers use, named apart from their topics
two-stream network · 0.4feature visualization · 0.4two-stream network visualization · 0.3cross-stream fusion analysis · 0.3temporal residual units · 0.3multiplicative interactions · 0.3motion gating · 0.3identity mapping kernels · 0.3fully convolutional spacetime network · 0.3convolutional neural network · 0.3spatiotemporal oriented energy · 0.2feature coding and pooling · 0.2dynamic spacetime pyramid · 0.2numerical conditioning · 0.1five-point algorithm · 0.1eight-point algorithm · 0.1ordinal correlation coefficient · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2020 | Representing Objects in Video as Space-Time Volumes by Combining Top-Down and Bottom-Up ProcessesabstractAs top-down based approaches of object recognition from video are getting more powerful, a structured way to combine them with bottom-up grouping processes becomes feasible. When done right, the resulting representation is able to describe objects and their decomposition into parts at appropriate spatio-temporal scales. We propose a method that uses a modern object detector to focus on salient structures in video, and a dense optical flow estimator to supplement feature extraction. From these structures we extract space-time volumes of interest (STVIs) by smoothing in spatio-temporal Gaussian Scale Space that guides bottom-up grouping. The resulting novel representation enables us to analyze and visualize the decomposition of an object into meaningful parts while preserving temporal object continuity. Our experimental validation is twofold. First, we achieve competitive results on a common video object segmentation benchmark. Second, we extend this benchmark with high quality object part annotations, DAVIS Parts1, on which we establish a strong baseline by showing that our method yields spatio-temporally meaningful object parts. Our new representation will support applications that require high-level space-time reasoning at the parts level. Filip Ilic, Axel Pinz |
WACV | 2 |
| 2020 | Synthesizing human-like sketches from natural images using a conditional convolutional decoderabstractHumans are able to precisely communicate diverse concepts by employing sketches, a highly reduced and abstract shape based representation of visual content. We propose, for the first time, a fully convolutional end-to-end architecture that is able to synthesize human-like sketches of objects in natural images with potentially cluttered background. To enable an architecture to learn this highly abstract mapping, we employ the following key components: (1) a fully convolutional encoder-decoder structure, (2) a perceptual similarity loss function operating in an abstract feature space and (3) conditioning of the decoder on the label of the object that shall be sketched. Given the combination of these architectural concepts, we can train our structure in an end-to-end supervised fashion on a collection of sketch-image pairs. The generated sketches of our architecture can be classified with 85.6% Top-5 accuracy and we verify their visual quality via a user study. We find that deep features as a perceptual similarity metric enable image translation with large domain gaps and our findings further show that convolutional neural networks trained on image classification tasks implicitly learn to encode shape information. Moritz Kampelmühler, Axel Pinz |
WACV | 2 |
| 2020 | Deep Insights into Convolutional Networks for Video RecognitionabstractAbstract As the success of deep models has led to their deployment in all areas of computer vision, it is increasingly important to understand how these representations work and what they are capturing. In this paper, we shed light on deep spatiotemporal representations by visualizing the internal representation of models that have been trained to recognize actions in video. We visualize multiple two-stream architectures to show that local detectors for appearance and motion objects arise to form distributed representations for recognizing human actions. Key observations include the following. First, cross-stream fusion enables the learning of true spatiotemporal features rather than simply separate appearance and motion features. Second, the networks can learn local representations that are highly class specific, but also generic representations that can serve a range of classes. Third, throughout the hierarchy of the network, features become more abstract and show increasing invariance to aspects of the data that are unimportant to desired distinctions (e.g. motion patterns across various speeds). Fourth, visualizations can be used not only to shed light on learned representations, but also to reveal idiosyncrasies of training data and to explain failure cases of the system. Christoph Feichtenhofer, Axel Pinz, Richard P. Wildes, Andrew Zisserman |
Int. J. Comput. Vis. | 2 |
| 2018 | What Have We Learned From Deep Representations for Action Recognition?abstractAs the success of deep models has led to their deployment in all areas of computer vision, it is increasingly important to understand how these representations work and what they are capturing. In this paper, we shed light on deep spatiotemporal representations by visualizing what two-stream models have learned in order to recognize actions in video. We show that local detectors for appearance and motion objects arise to form distributed representations for recognizing human actions. Key observations include the following. First, cross-stream fusion enables the learning of true spatiotemporal features rather than simply separate appearance and motion features. Second, the networks can learn local representations that are highly class specific, but also generic representations that can serve a range of classes. Third, throughout the hierarchy of the network, features become more Abstract and show increasing invariance to aspects of the data that are unimportant to desired distinctions (e.g. motion patterns across various speeds). Fourth, visualizations can be used not only to shed light on learned representations, but also to reveal idiosyncracies of training data and to explain failure cases of the system. This document is best viewed offline where figures play on click. Christoph Feichtenhofer, Axel Pinz, Richard P. Wildes, Andrew Zisserman |
CVPR | 2 |
| 2017 | Temporal Residual Networks for Dynamic Scene RecognitionabstractThis paper combines three contributions to establish a new state-of-the-art in dynamic scene recognition. First, we present a novel ConvNet architecture based on temporal residual units that is fully convolutional in spacetime. Our model augments spatial ResNets with convolutions across time to hierarchically add temporal residuals as the depth of the network increases. Second, existing approaches to video-based recognition are categorized and a baseline of seven previously top performing algorithms is selected for comparative evaluation on dynamic scenes. Third, we introduce a new and challenging video database of dynamic scenes that more than doubles the size of those previously available. This dataset is explicitly split into two subsets of equal size that contain videos with and without camera motion to allow for systematic study of how this variable interacts with the defining dynamics of the scene per se. Our evaluations verify the particular strengths and weaknesses of the baseline algorithms with respect to various scene classes and camera motion parameters. Finally, our temporal ResNet boosts recognition performance and establishes a new state-of-the-art on dynamic scene recognition, as well as on the complementary task of action recognition. Christoph Feichtenhofer, Axel Pinz, Richard P. Wildes |
CVPR | 2 |
| 2017 | Spatiotemporal Multiplier Networks for Video Action RecognitionabstractThis paper presents a general ConvNet architecture for video action recognition based on multiplicative interactions of spacetime features. Our model combines the appearance and motion pathways of a two-stream architecture by motion gating and is trained end-to-end. We theoretically motivate multiplicative gating functions for residual networks and empirically study their effect on classification accuracy. To capture long-term dependencies we inject identity mapping kernels for learning temporal relationships. Our architecture is fully convolutional in spacetime and able to evaluate a video in a single forward pass. Empirical investigation reveals that our model produces state-of-the-art results on two standard action recognition datasets. Christoph Feichtenhofer, Axel Pinz, Richard P. Wildes |
CVPR | 2 |
| 2017 | Detect to Track and Track to DetectabstractRecent approaches for high accuracy detection and tracking of object categories in video consist of complex multistage solutions that become more cumbersome each year. In this paper we propose a ConvNet architecture that jointly performs detection and tracking, solving the task in a simple and effective way. Our contributions are threefold: (i) we set up a ConvNet architecture for simultaneous detection and tracking, using a multi-task objective for frame-based object detection and across-frame track regression; (ii) we introduce correlation features that represent object co-occurrences across time to aid the ConvNet during tracking; and (iii) we link the frame level detections based on our across-frame tracklets to produce high accuracy detections at the video level. Our ConvNet architecture for spatiotemporal object detection is evaluated on the large-scale ImageNet VID dataset where it achieves state-of-the-art results. Our approach provides better single model performance than the winning method of the last ImageNet challenge while being conceptually much simpler. Finally, we show that by increasing the temporal stride we can dramatically increase the tracker speed. Christoph Feichtenhofer, Axel Pinz, Andrew Zisserman |
ICCV | 2 |
| 2016 | Convolutional Two-Stream Network Fusion for Video Action RecognitionabstractRecent applications of Convolutional Neural Networks (ConvNets) for human action recognition in videos have proposed different solutions for incorporating the appearance and motion information. We study a number of ways of fusing ConvNet towers both spatially and temporally in order to best take advantage of this spatio-temporal information. We make the following findings: (i) that rather than fusing at the softmax layer, a spatial and temporal network can be fused at a convolution layer without loss of performance, but with a substantial saving in parameters, (ii) that it is better to fuse such networks spatially at the last convolutional layer than earlier, and that additionally fusing at the class prediction layer can boost accuracy, finally (iii) that pooling of abstract convolutional features over spatiotemporal neighbourhoods further boosts performance. Based on these studies we propose a new ConvNet architecture for spatiotemporal fusion of video snippets, and evaluate its performance on standard benchmarks where this architecture achieves state-of-the-art results. Christoph Feichtenhofer, Axel Pinz, Andrew Zisserman |
CVPR | 2 |
| 2016 | Spatiotemporal Residual Networks for Video Action RecognitionabstractTwo-stream Convolutional Networks (ConvNets) have shown strong performance for human action recognition in videos. Recently, Residual Networks (ResNets) have arisen as a new technique to train extremely deep architectures. In this paper, we introduce spatiotemporal ResNets as a combination of these two approaches. Our novel architecture generalizes ResNets for the spatiotemporal domain by introducing residual connections in two ways. First, we inject residual connections between the appearance and motion pathways of a two-stream architecture to allow spatiotemporal interaction between the two streams. Second, we transform pretrained image ConvNets into spatiotemporal networks by equipping these with learnable convolutional filters that are initialized as temporal residual connections and operate on adjacent feature maps in time. This approach slowly increases the spatiotemporal receptive field as the depth of the model increases and naturally integrates image ConvNet design principles. The whole model is trained end-to-end to allow hierarchical learning of complex spatiotemporal features. We evaluate our novel spatiotemporal ResNet using two widely used action recognition benchmarks where it exceeds the previous state-of-the-art. Christoph Feichtenhofer, Axel Pinz, Richard P. Wildes |
NIPS | 2 |
| 2016 | Dynamic Scene Recognition with Complementary Spatiotemporal FeaturesabstractThis paper presents Dynamically Pooled Complementary Features (DPCF), a unified approach to dynamic scene recognition that analyzes a short video clip in terms of its spatial, temporal and color properties. The complementarity of these properties is preserved through all main steps of processing, including primitive feature extraction, coding and pooling. In the feature extraction step, spatial orientations capture static appearance, spatiotemporal oriented energies capture image dynamics and color statistics capture chromatic information. Subsequently, primitive features are encoded into a mid-level representation that has been learned for the task of dynamic scene recognition. Finally, a novel dynamic spacetime pyramid is introduced. This dynamic pooling approach can handle both global as well as local motion by adapting to the temporal structure, as guided by pooling energies. The resulting system provides online recognition of dynamic scenes that is thoroughly evaluated on the two current benchmark datasets and yields best results to date on both datasets. In-depth analysis reveals the benefits of explicitly modeling feature complementarity in combination with the dynamic spacetime pyramid, indicating that this unified approach should be well-suited to many areas of video analysis. Christoph Feichtenhofer, Axel Pinz, Richard P. Wildes |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2015 | Cultural Heritage Acquisition: Geometry-Based Radiometry in the WildabstractMeasuring radiometric surface properties "in the wild" is challenging, because the appearance of a surface strongly depends on ambient illumination, especially when direct sunlight or cast shadows need to be handled, and when the scene must not be altered by shrouding devices or by reference targets. We present a purely photogram metric measurement method that integrates 3D surface geometry, camera poses, known artificial illumination, and calibration of the measurement device. For geometry measurement, we extend Structure-from-Motion with constrained bundle adjustment post processing to obtain geo-referenced Euclidean reconstructions. For radiometry measurement, we perform frame differencing of images with and without intense artificial illumination to eliminate the influence of ambient illumination, and we calculate radiometric surface properties based on known illumination configuration and radiometric calibration. Calculating the relations between surface normal, camera pose, and incident light leads to the final result of dense 3D point clouds that represent radiometric surface properties. In our experimental validation, we apply this method to the field of cultural heritage acquisition of prehistoric rock art and achieve excellent results compared to ground truth. Further, the measurement method itself is quite general and will be applicable to many research areas where both, precise surface geometry and radiometry is required. Thomas Höll, Axel Pinz |
3DV | 2 |
| 2015 | Dynamically encoded actions based on spacetime saliencyabstractHuman actions typically occur over a well localized extent in both space and time. Similarly, as typically captured in video, human actions have small spatiotemporal support in image space. This paper capitalizes on these observations by weighting feature pooling for action recognition over those areas within a video where actions are most likely to occur. To enable this operation, we define a novel measure of spacetime saliency. The measure relies on two observations regarding foreground motion of human actors: They typically exhibit motion that contrasts with that of their surrounding region and they are spatially compact. By using the resulting definition of saliency during feature pooling we show that action recognition performance achieves state-of-the-art levels on three widely considered action recognition datasets. Our saliency weighted pooling can be applied to essentially any locally defined features and encodings thereof. Additionally, we demonstrate that inclusion of locally aggregated spatiotemporal energy features, which efficiently result as a by-product of the saliency computation, further boosts performance over reliance on standard action recognition features alone. Christoph Feichtenhofer, Axel Pinz, Richard P. Wildes |
CVPR | 2 |
| 2014 | Bags of Spacetime Energies for Dynamic Scene RecognitionabstractThis paper presents a unified bag of visual word (BoW) framework for dynamic scene recognition. The approach builds on primitive features that uniformly capture spatial and temporal orientation structure of the imagery (e.g., video), as extracted via application of a bank of spatiotemporally oriented filters. Various feature encoding techniques are investigated to abstract the primitives to an intermediate representation that is best suited to dynamic scene representation. Further, a novel approach to adaptive pooling of the encoded features is presented that captures spatial layout of the scene even while being robust to situations where camera motion and scene dynamics are confounded. The resulting overall approach has been evaluated on two standard, publically available dynamic scene datasets. The results show that in comparison to a representative set of alternatives, the proposed approach outperforms the previous state-of-the-art in classification accuracy by 10%. Christoph Feichtenhofer, Axel Pinz, Richard P. Wildes |
CVPR | 2 |
| 2014 | Exploiting temporal and spatial constraints in traffic sign detection from a moving vehicle
Sinisa Segvic, Karla Brkic, Zoran Kalafatic, Axel Pinz |
Mach. Vis. Appl. | 4 |
| 2013 | Spacetime Forests with Complementary Features for Dynamic Scene RecognitionabstractThis paper presents spacetime forests defined over complementary spatial and temporal features for recognition of naturally occurring dynamic scenes. The approach improves on the previous state-of-the-art in both classification and execution rates. A particular improvement is with increased robustness to camera motion, where previous approaches have experienced difficulty. There are three key novelties in the approach. First, a novel spacetime descriptor is employed that exploits the complementary nature of spatial and temporal information, as inspired by previous research on the role of orientation features in scene classification. Second, a forest-based classifier is used to learn a multi-class representation of the feature distributions. Third, the video is processed in temporal slices with scale matched preferentially to scene dynamics over camera motion. Slicing allows for temporal alignment to be handled as latent information in the classifier and for efficient, incremental processing. The integrated approach is evaluated empirically on two publically available datasets to document its outstanding performance. Christoph Feichtenhofer, Axel Pinz, Richard P. Wildes |
BMVC | 2 |
| 2009 | A level set framework using a new incremental, robust Active Shape Model for object segmentation and tracking
Michael Fussenegger, Peter M. Roth, Horst Bischof, Rachid Deriche, Axel Pinz |
Image Vis. Comput. | 5 |
| 2008 | Globally Optimal O(n) Solution to the PnP Problem for General Camera ModelsabstractWe present a novel and fast algorithm to solve the Perspective-n-Point problem. The PnP problem- estimating the pose of a calibrated camera based on measurements and known 3D scene, is recasted as a minimization problem of the Object Space Cost. Instead of limiting the algorithm to perspective cameras, we use a formulation for general camera models. The minimization problem, together with a quaternion based representation of the rotation, is transferred into a semi definite positive program (SDP). This transfer is done in O(n) time and leads to an SDP of constant size. The solution of the SDP is a global minimizer of the PnP problem, which can be estimated in less than 0.15 seconds for 100 points. 1 Gerald Schweighofer, Axel Pinz |
BMVC | 2 |
| 2008 | Online/Realtime Structure and Motion for General Camera ModelsabstractThis paper presents a novel algorithm for online structure and motion estimation. The algorithm works for general camera models and minimizes object space error, it does not rely on gradient-based optimization, and it is provably globally convergent. In comparison to previous work, which reports cubic complexity in the number of frames, our major contribution is a significant reduction of complexity. The new algorithm requires constant time per frame and can thus be used in online applications. Experimental results show high reconstruction accuracy with respect to simulated ground truth data. We also present two applications in artificial marker reconstruction and handheld augmented reality. Gerald Schweighofer, Sinisa Segvic, Axel Pinz |
WACV | 3 |
| 2008 | Learning an Alphabet of Shape and Appearance for Multi-Class Object DetectionabstractWe present a novel algorithmic approach to object categorization and detection that can learn category specific detectors, using Boosting, from a visual alphabet of shape and appearance. The alphabet itself is learnt incrementally during this process. The resulting representation consists of a set of category-specific descriptors—basic shape features are represented by boundary-fragments, and appearance is represented by patches—where each descriptor in combination with centroid vectors for possible object centroids (geometry) forms an alphabet entry. Our experimental results highlight several qualities of this novel representation. First, we demonstrate the power of purely shape-based representation with excellent categorization and detection results using a Boundary-Fragment-Model (BFM), and investigate the capabilities of such a model to handle changes in scale and viewpoint, as well as intra- and inter-class variability. Second, we show that incremental learning of a BFM for many categories leads to a sub-linear growth of visual alphabet entries by sharing of shape features, while this generalization over categories at the same time often improves categorization performance (over independently learning the categories). Finally, the combination of basic shape and appearance (boundary-fragments and patches) features can further improve results. Certain feature types are preferred by certain categories, and for some categories we achieve the lowest error rates that have been reported so far. Andreas Opelt, Axel Pinz, Andrew Zisserman |
Int. J. Comput. Vis. | 2 |
| 2007 | Influence of numerical conditioning on the accuracy of relative orientationabstractWe study the influence of numerical conditioning on the accuracy of two closed-form solutions to the overconstrained relative orientation problem. We consider the well known eight-point algorithm and the recent five-point algorithm, and evaluate changes in their performance due to Hartley's normalization and Muehlich's equilibration. The need for numerical conditioning is introduced by explaining the known occurence of the bias of the eight-point algorithm towards the forward motion. Then it is shown how conditioning can be used to improve the results of the recent five-point algorithm. This is not straightforward since the conditioning disturbs the calibration of the input data. The conditioning therefore needs to be reverted before enforcing the internal cubic constraints of the essential matrix. The obtained improvements are less dramatic than in the case of the eight-point algorithm, for which we offer a plausible explanation. The theoretical claims are backed up with extensive experimentation on noisy artificial datasets, under a variety of geometric and imaging parameters. Sinisa Segvic, Gerald Schweighofer, Axel Pinz |
CVPR | 3 |
| 2007 | An augmented reality human-computer interface for object localization in a cognitive vision system
Hannes Siegl, Marc Hanheide, Sebastian Wrede 0001, Axel Pinz |
Image Vis. Comput. | 4 |
| 2006 | Multiregion Level Set Tracking with Transformation Invariant Shape Priors
Michael Fussenegger, Rachid Deriche, Axel Pinz |
ACCV (1) | 3 |
| 2006 | A Multiphase Level Set Based Segmentation Framework with Pose Invariant Shape Priors
Michael Fussenegger, Rachid Deriche, Axel Pinz |
ACCV (2) | 3 |
| 2006 | Fusing Shape and Appearance Information for Object Category DetectionabstractWe present methods for recognizing object categories which are able to combine various feature types (e.g. image patches and edge boundaries). Our objective is to detect object instances in an image, as opposed to the easier task of image categorization. To this end, we investigate two algorithms for learning and detecting object categories which both benefit from combining features. The first uses a naive combination method for detectors each employing only one type of feature, the second learns the best features (from a pool of patches and boundaries). In experiments we achieve comparable results to the state of the art over a number of datasets, and for some categories we even achieve the lowest errors that have been reported so far. The results also show that certain object categories prefer certain feature types (e.g. boundary fragments for airplanes). Andreas Opelt, Andrew Zisserman, Axel Pinz |
BMVC | 3 |
| 2006 | Fast and Globally Convergent Structure and Motion Estimation for General Camera ModelsabstractThis paper presents a novel algorithm to solve the Structure and Motion problem. The novelty is in the use of a general camera model, which does not constrain the algorithm to a specific camera, and the use of the Object Space Error for General Camera Models as cost function. We show that, using this cost function, the structure and the translation part of the motion can be estimated from the rotation part of the motion in closed form. So only the rotation part of the motion needs to be optimized to estimate the minimum of the total cost function. This results in an iterative algorithm which has a theoretical speedup factor of 8 compared to the bundle adjustment method. We also prove the global convergence of the presented algorithm. 1 Gerald Schweighofer, Axel Pinz |
BMVC | 2 |
| 2006 | Incremental learning of object detectors using a visual shape alphabetabstractWe address the problem of multiclass object detection. Our aims are to enable models for new categories to benefit from the detectors built previously for other categories, and for the complexity of the multiclass system to grow sublinearly with the number of categories. To this end we introduce a visual alphabet representation which can be learnt incrementally, and explicitly shares boundary fragments (contours) and spatial configurations (relation to centroid) across object categories. We develop a learning algorithm with the following novel contributions: (i) AdaBoost is adapted to learn jointly, based on shape features; (ii) a new learning schedule enables incremental additions of new categories; and (iii) the algorithm learns to detect objects (instead of categorizing images). Furthermore, we show that category similarities can be predicted from the alphabet. We obtain excellent experimental results on a variety of complex categories over several visual aspects. We show that the sharing of shape features not only reduces the number of features required per category, but also often improves recognition performance, as compared to individual detectors which are trained on a per-class basis. Andreas Opelt, Axel Pinz, Andrew Zisserman |
CVPR (1) | 2 |
| 2006 | A Boundary-Fragment-Model for Object Detection
Andreas Opelt, Axel Pinz, Andrew Zisserman |
ECCV (2) | 2 |
| 2006 | Generic Object Recognition with BoostingabstractThis paper explores the power and the limitations of weakly supervised categorization. We present a complete framework that starts with the extraction of various local regions of either discontinuity or homogeneity. A variety of local descriptors can be applied to form a set of feature vectors for each local region. Boosting is used to learn a subset of such feature vectors (weak hypotheses) and to combine them into one final hypothesis for each visual category. This combination of individual extractors and descriptors leads to recognition rates that are superior to other approaches which use only one specific extractor/descriptor setting. To explore the limitation of our system, we had to set up new, highly complex image databases that show the objects of interest at varying scales and poses, in cluttered background, and under considerable occlusion. We obtain classification results up to 81 percent ROC-equal error rate on the most complex of our databases. Our approach outperforms all comparable solutions on common databases. Andreas Opelt, Axel Pinz, Michael Fussenegger, Peter Auer |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2006 | Robust Pose Estimation from a Planar TargetabstractIn theory, the pose of a calibrated camera can be uniquely determined from a minimum of four coplanar but noncollinear points. In practice, there are many applications of camera pose tracking from planar targets and there is also a number of recent pose estimation algorithms which perform this task in real-time, but all of these algorithms suffer from pose ambiguities. This paper investigates the pose ambiguity for planar targets viewed by a perspective camera. We show that pose ambiguities--two distinct local minima of the according error function--exist even for cases with wide angle lenses and close range targets. We give a comprehensive interpretation of the two minima and derive an analytical solution that locates the second minimum. Based on this solution, we develop a new algorithm for unique and robust pose estimation from a planar target. In the experimental evaluation, this algorithm outperforms four state-of-the-art pose estimation algorithms. Gerald Schweighofer, Axel Pinz |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2004 | Weak Hypotheses and Boosting for Generic Object Detection and Recognition
Andreas Opelt, Michael Fussenegger, Axel Pinz, Peter Auer |
ECCV (2) | 3 |
| 2004 | A New High Speed CMOS Camera for Real-time Tracking ApplicationsabstractThere are many potential applications for very high-speed vision sensors in robotics. Maybe tracking is the most obvious one. The main idea of this paper is to combine existing standard technology (CMOS imaging sensors, FPGA "glue logic", and USB 2.0 interface) to use direct pixel access capabilities for real-time tracking. We designed and built such a prototype camera system, including FPGA programmed functionality for Fixed Pattern Noise (FPN) calibration, subsampling and direct subwindow access. The performance of this new camera is on one hand limited by the maximum pixel clock of the sensor, on the other hand by the USB 2.0 microframe timing and bandwith constraints. We achieve "frame rates" of up to 2.5 kHz for small subwindows which can be randomly and individually addressed for each update cycle. Besides all given technical details and specifications of the system, we show a demonstration application of a high-speed blob tracking system which verifies the usability of our new camera for highly demanding tracking applications. Ulrich Muehlmann, Miguel Ribo, Peter Lang, Axel Pinz |
ICRA | 4 |
| 2003 | Real-Time Camera Pose in a Room
Manmohan Krishna Chandraker, Christoph Stock, Axel Pinz |
ICVS | 3 |
| 2001 | Synthesizing realistic facial animations using energy minimization for model-based coding
Lijun Yin 0001, Anup Basu, Stefan Bernögger, Axel Pinz |
Pattern Recognit. | 4 |
| 2001 | Automated Melanoma RecognitionabstractA system for the computerized analysis of images obtained from ELM has been developed to enhance the early recognition of malignant melanoma. As an initial step, the binary mask of the skin lesion is determined by several basic segmentation algorithms together with a fusion strategy. A set of features containing shape and radiometric features as well as local and global parameters is calculated to describe the malignancy of a lesion. Significant features are then selected from this set by application of statistical feature subset selection methods. The final kNN classification delivers a sensitivity of 87% with a specificity of 92%. Harald Ganster, Axel Pinz, Reinhard Röhrer, Ernst Wildling, Michael Binder, Harald Kittler |
IEEE Trans. Medical Imaging | 2 |
| 2000 | Fuzzy Relative Positions for Qualitative Pose EstimationabstractSimilarity of object part embedding (description of how each part is surrounded by its neighbors) can be used to qualitatively infer similarity of images. We present an approach using fuzzy relative positions to qualitatively describe and reason with object part embedding. In the framework of a qualitative active object recognition experiment, we demonstrate the usefulness of the approach in order to recover the pose of the object under inspection with regard to the camera viewing direction. Jean-Philippe Andreu, Axel Pinz |
ICPR | 2 |
| 2000 | 3D Model Based Pose Determination in Real-Time: Strategies, Convergence, AccuracabstractIn this paper, a new real-time model based pose determination system based on the Gauss-Newton method is described. Various model fitting strategies, which use gradient search, are discussed and compared. A novel strategy, the adaptive perpendicular search, is proposed for both matching ambiguity avoidance and convergence enforcement. The system is shown to be highly effective in practice, while keeping the computational cost low. A sufficient initial pose estimate as an input for the Gauss-Newton based methods is obtained from the parametric eigenspace. With certain restrictions, the overall pose determination system performs in real-time, as required in industrial applications. Thomas Auer, Gernot Bachler, Stefan Scherer, Axel Pinz |
ICPR | 5 |
| 2000 | Learning Temporal Context in Active Object Recognition Using Bayesian AnalysisabstractActive object recognition is a successful strategy to reduce the uncertainty of single view recognition, by planning sequences of views, actively obtaining these views, and integrating multiple recognition results. Understanding recognition as a sequential decision problem challenges the visual agent to select discriminative information sources. The presented system emphasizes the importance of temporal context in disambiguating initial object hypotheses, provides the corresponding theory for Bayesian fusion processes, and demonstrates its performance to be superior to alternative view planning schemes. Instance based learning proposed to estimate the control function enables then real-time processing with improved performance characteristics. Lucas Paletta, Manfred Prantl, Axel Pinz |
ICPR | 3 |
| 2000 | Appearance-based active object recognition
Hermann Borotschnig, Lucas Paletta, Manfred Prantl, Axel Pinz |
Image Vis. Comput. | 4 |
| 2000 | An automatic assessment scheme for steel quality inspection
Klaus Wiltschi, Axel Pinz, Tony Lindeberg |
Mach. Vis. Appl. | 2 |
| 1999 | A Vision Driven Automatic Assembly Unit
Gernot Bachler, Reinhard Röhrer, Stefan Scherer, Axel Pinz |
CAIP | 5 |
| 1999 | Real-Time Optical Edge and Corner Tracking at Subpixel Accuracy
Stefan Brantner, Thomas Auer, Axel Pinz |
CAIP | 3 |
| 1999 | Subpixel Stereo Matching by Robust Estimation of Local Distortion Using Gabor Filters
Peter Werth, Stefan Scherer, Axel Pinz |
CAIP | 3 |
| 1999 | The Discriminatory Power of Ordinal Measures - Towards a New CoefficientabstractPerspective distortion, occlusion and specular reflection are challenging problems in shape-from-stereo. In this paper we review one recently published area-based stereo matching algorithm (Bhat and Nayar, 1998) designed to be robust in these cases. Although the algorithm is an important contribution to stereo-matching, we show that its coefficient has a low discriminatory power, which leads to a significant number of multiple best matches. In order to cope with this drawback we introduce a new normalized ordinal correlation coefficient. Experiments showing the behavior of the proposed coefficient are performed on various datasets including real data with ground truth. The new coefficient reduces the occurrence of multiple best matches to almost zero per cent. It also shows a more robust and equally accurate behavior. These benefits are achieved at almost no additional computational costs. Stefan Scherer, Axel Pinz, Peter Werth |
CVPR | 2 |
| 1999 | The integration of optical and magnetic tracking for multi-user augmented reality
Thomas Auer, Stefan Brantner, Axel Pinz |
EGVE | 3 |
| 1999 | Visual object detection for autonomous sewer robotsabstractThe goal of the proposed detection system is to identify objects, e.g. inlets, in sewage pipes. A camera attached to an autonomous sewer robot provides images that are interpreted by an attention driven recognition module. Local appearances in the input image are represented in an environment specific description subspace extracted by principal component analysis. The object class posterior interpretation in terms of a radial basis function network constitutes an attention filter constraining further processing on receptive fields filter resolutions. Multiresolution decision fusion is the framework used to combine detection confidences to enhance robustness in the global classification. The vision system is evaluated in various experiments where it proves successful with respect to the local classification rate, to the generalization behavior in recognizing similar objects, and to detection that requires a minimum of positive falses. Lucas Paletta, Erich Rome, Axel Pinz |
IROS | 3 |
| 1999 | The integration of optical and magnetic tracking for multi-user augmented reality
Thomas Auer, Axel Pinz |
Comput. Graph. | 2 |
| 1998 | Active Object Recognition in Parametric EigenspaceabstractWe present an efficient method within an active vision framework for recognizing objects which are ambiguous from certain viewpoints. The system is allowed to reposition the camera to capture additional views and, therefore, to resolve the classification result obtained from a single view. The approach uses an appearance based object representation, namely the parametric eigenspace, and augments it by probability distributions. This captures possible variations in the input images due to errors in the pre-processing chain or the imaging system. Furthermore, the use of probability distributions gives us a gauge to view planning. View planning is shown to be of great use in reducing the number of images to be captured when compared to a random strategy. 1 Introduction Most computer vision systems found in the literature perform object recognition on the basis of the information gathered from a single image. Typically, a set of features is extracted and matched against object ... Hermann Borotschnig, Lucas Paletta, Manfred Prantl, Axel Pinz |
BMVC | 4 |
| 1998 | Eye tracking and animation for MPEG-4 codingabstractAccurate localization and tracking of facial features are crucial for developing high quality model-based coding (MPEG-4) systems. For teleconferencing applications at very low bit rates, it is necessary to track eye and lip movements accurately over time. These movements can be coded and transmitted to a remote site, where animation techniques can be used to synthesize facial movements on a model of a face. In this paper we describe and simple heuristics which are effective in improving the results of well-known facial feature detection and tracking algorithms. Animation models are also presented, along with experimental results to demonstrate the system being developed. We focus our discussion only on the detection, tracking and modeling of eye movements. Stefan Bernögger, Lijun Yin 0001, Anup Basu, Axel Pinz |
ICPR | 4 |
| 1998 | Qualitative spatial reasoning to infer the camera position in generic object recognitionabstractQualitative spatial reasoning and qualitative representation of space is required in many applications of computer vision. We present a new approach using fuzzy spatial relations to qualitatively estimate the current camera position in an active object recognition experiment. Starting from a single image of an object and a corresponding mapping to a 3D CAD-prototype, the visibility and the occlusion of parts of the prototype are used to infer possible viewing directions on a view sphere (i.e. initial object pose estimation). This representation can be used for several tasks, e.g. refining the current viewpoint estimation, obtaining a new view and verifying the current object hypothesis. Experiments demonstrate initial view point estimations for simple CAD prototypes. This new method is applicable to generic object recognition and to other areas of qualitative vision. Axel Pinz, Jean-Philippe Andreu |
ICPR | 1 |
| 1998 | Feature selection in melanoma recognitionabstractMelanoma, one of the most aggressive types of cancer, can be healed, if recognized in early stages. In order to automate the early recognition of skin cancer; a system that analyses digital epiluminescence microscopic images is used. After segmentation, 33 features representing shape and radiometric properties are calculated. In the paper the quality of the features is evaluated by applying several feature selection methods. The results show that with each selection method the feature set can be reduced to dimension four with nearly no loss of information. Results with classification rates of up to 75% are achieved and relations between selected features and medical criteria are observed. Reinhard Röhrer, Harald Ganster, Axel Pinz, Michael Binder |
ICPR | 3 |
| 1998 | Robust adaptive window matching by homogeneity constraint and integration of descriptionsabstractThe crucial element of existing binocular stereo algorithms is to establish point to point correspondences in the two images. In order to overcome the problem of distortions, a promising approach is to locally adapt the correlation window size. In this paper we derive two major simplifications of an existing adaptive window approach. These simplifications are justified for images of objects with constant reflection properties. For this purpose the homogeneity constraint is introduced. In order to increase the robustness of the simplified algorithm the integration with a local description matching approach is proposed. The new algorithm is extended by subpixel approximation. Experiments are performed on real data and compared with manual measurements. A numerical analysis of the results demonstrates the achieved accuracy anal robustness. Stefan Scherer, Wilfried Andexer, Axel Pinz |
ICPR | 3 |
| 1998 | Mapping the Human Brain RetinaabstractThe new therapeutic method of scotoma-based photocoagulation (SBP) developed at the Vienna Eye Clinic for diagnosis and treatment of age-related macular degeneration requires retinal maps from scanning laser ophthalmoscope images. This paper describes in detail all necessary image analysis steps for map generation. A prototype software system for fully automatic map generation has been implemented and tested on a representative dataset selected from a clinical study with 50 patients. The map required for the SBP treatment can be reliably extracted in all cases. Thus, algorithms presented in this paper should be directly applicable in daily clinical routine without major modifications. Axel Pinz, Stefan Bernögger |
IEEE Trans. Medical Imaging | 1 |
| 1997 | Classification of carbide distributions using scale selection and directional distributionsabstractWe present an automatic method for the classification of steel quality based on scale-space operations. The carbide distribution of microscopic specimen images is assessed by classifying according to so-called 'degree' and 'type' of the specimen. 'Degree' is represented by features extracted with automatic scale selection, and 'type' information is computed from second-moment descriptors. In combination with a morphological verification scheme, this pattern classifier shares large similarities with current manual techniques. Compared to previous work, the new classification scheme has several advantages: the significant scale of the carbide agglomeration is calculated explicitly, and the method is less sensitive to the variance of spatial connectivity than a morphological approach. Klaus Wiltschi, Tony Lindeberg, Axel Pinz |
ICIP (3) | 3 |
| 1996 | Active fusion using Bayesian networks applied to multi-temporal remote sensing imageryabstractImage processing applications and especially those in the area of remote sensing are often characterized by a high degree of complexity. We introduce a general framework, called 'active fusion', that actively selects and combines information from multiple sources in order to obtain a reliable result at reasonable costs. A sample implementation of parts of the framework is given using Bayesian networks and decision theoretic techniques for the task of agricultural field classification. This experiment shows a significant reduction in the number of information sources required for a reliable decision. Manfred Prantl, Harald Ganster, Axel Pinz |
ICPR | 3 |
| 1996 | Topological investigations of object modelsabstractThe purpose of this paper is the development of a theoretical basis for the investigation of a 3D model's topological structure which is currently a more or less neglected aspect in the areas of object reconstruction and vision. The knowledge of this structure gives rise to detailed statements on the correspondence between a real world object and a digital model which cannot be made using other methods. We describe a generalization of the snake concept to arbitrary triangulated surfaces. Our aim is the calculation of so-called minimal length loops which may be used for the determination and localization of tunnels. Minimal length loops can be used in practice for applications like the assessment of object reconstruction results, autonomous reconstruction and motion planing. Peter Uray, Axel Pinz |
ICPR | 2 |
| 1996 | Active fusion - A new method applied to remote sensing image interpretation
Axel Pinz, Manfred Prantl, Harald Ganster, Hermann Borotschnig |
Pattern Recognit. Lett. | 1 |
| 1995 | Affine Matching of Intermediate Symbolic Representations
Axel Pinz, Manfred Prantl, Harald Ganster |
CAIP | 1 |
| 1993 | Layout and analysis: Finding text, titles, and photos in digital images of newspaper pagesabstractAn important step in the analysis of printed documents is the segmentation and classification of blocks into categories such as photographs, titles, paragraphs, etc. The authors present an approach to enhance and combine two commonly used methods, a merging bottom-up approach and a cutting top down approach, to segment pages of a newspaper. The implementation of a layout analysis system as a preprocessing module for a commercial product is described.> Johann Wieser, Axel Pinz |
ICDAR | 2 |
| 1992 | Neural Network "Surgery": Transplantation of Hidden Units
Axel Pinz, Horst Bischof |
ECAI | 1 |
| 1992 | Visualization methods for neural networksabstractThe interpretation of neural network behavior is of particular interest in neural network research. Visualization methods provide the necessary means to simultaneously analyze the huge amount of information hidden in the network. The authors propose a framework for visualization methods suited for feed forward neural networks. The basic idea is to use the spatial information available outside the network to arrange the data to be visualized (weights, activations of units) in the spatial domain of the display. Several examples which illustrate the proposed framework are presented.> Horst Bischof, Axel Pinz, Walter G. Kropatsch |
ICPR (2) | 2 |
| 1992 | Information fusion in image understandingabstractComputer vision and image understanding processes are not very robust; small changes in exposure parameters or in internal parameters of algorithms can lead to significantly different results. A combination (fusion) of these results is profitable. The authors introduce an extended fusion concept dealing with different sources of information at external (world, scene, image) and internal (image description, scene description) levels and define the process of fusion. Each level requires its own procedure of quality measure and information fusion in order to yield a combination of components from several sources. Related work in the field is reviewed. Examples from the authors' own work cover remote sensing (improvement of classification results by fusion at the image level), medical image processing of ocular fundus images (automatic control point selection by fusion at the image description level) and the interpretation of Billard scenes (object identification by fusion at the scene description level).> Axel Pinz, Renate Bartl |
ICPR (1) | 1 |
| 1992 | Multispectral classification of Landsat-images using neural networksabstractThe authors report the application of three-layer back-propagation networks for classification of Landsat TM data on a pixel-by-pixel basis. The results are compared to Gaussian maximum likelihood classification. First, it is shown that the neural network is able to perform better than the maximum likelihood classifier. Secondly, in an extension of the basic network architecture it is shown that textural information can be integrated into the neural network classifier without the explicit definition of a texture measure. Finally, the use of neural networks for postclassification smoothing is examined.> Horst Bischof, Werner Schneider, Axel Pinz |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 1990 | Constructing a neural network for the interpretation of the species of trees in aerial photographsabstractA neural network with a three-layer feedforward architecture was used to interpret the species of trees in aerial photographs. Weight-visualization (WV) diagrams were developed to interpret the weights and the behavior of the hidden units easily. Several networks that were trained on different parameter settings were combined to construct a better-performing network using the WV diagrams as a tool.> Axel Pinz, Horst Bischof |
ICPR (1) | 1 |