VLDB 2026 Research / reviewers in the wild / expert
Arnold W. M. Smeulders
dblp:15/5400
· DBLP profile ↗
174ranked-venue papers
8as first author
6since 2021 · last 2022
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 118 · 4 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 101 · 3 first-author · 5 since 2021Databases, data management, data science and information retrieval · 12 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 5 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | Contrasting Quadratic Assignments for Set-Based Representation Learning
Artem Moskalev, Ivan Sosnovik, Arnold W. M. Smeulders |
ECCV (27) | 4 |
| 2022 | LieGG: Studying Learned Lie Group GeneratorsabstractSymmetries built into a neural network have appeared to be very beneficial for a wide range of tasks as it saves the data to learn them. We depart from the position that when symmetries are not built into a model a priori, it is advantageous for robust networks to learn symmetries directly from the data to fit a task function. In this paper, we present a method to extract symmetries learned by a neural network and to evaluate the degree to which a network is invariant to them. With our method, we are able to explicitly retrieve learned invariances in a form of the generators of corresponding Lie-groups without prior knowledge of symmetries in the data. We use the proposed method to study how symmetrical properties depend on a neural network's parameterization and configuration. We found that the ability of a network to learn symmetries generalizes over a range of architectures. However, the quality of learned symmetries depends on the depth and the number of parameters. Artem Moskalev, Anna Sepliarskaia, Ivan Sosnovik, Arnold W. M. Smeulders |
NeurIPS | 4 |
| 2021 | Human-object Interaction Detection without Alignment Supervision
Mert Kilickaya, Arnold W. M. Smeulders |
BMVC | 2 |
| 2021 | DISCO: accurate Discrete Scale Convolutions
Ivan Sosnovik, Artem Moskalev, Arnold W. M. Smeulders |
BMVC | 3 |
| 2021 | Structured Visual Search via Composition-aware LearningabstractThis paper studies visual search using structured queries. The structure is in the form of a 2D composition that encodes the position and the category of the objects. The transformation of the position and the category of the objects leads to a continuous-valued relationship between visual compositions, which carries highly beneficial information, although not leveraged by previous techniques. To that end, in this work, our goal is to leverage these continuous relationships by using the notion of symmetry in equivariance. Our model output is trained to change symmetrically with respect to the input transformations, leading to a sensitive feature space. Doing so leads to a highly efficient search technique, as our approach learns from fewer data using a smaller feature space. Experiments on two large-scale benchmarks of MS-COCO and HICO-DET demonstrates that our approach leads to a considerable gain in the performance against competing techniques. Mert Kilickaya, Arnold W. M. Smeulders |
WACV | 2 |
| 2021 | Scale Equivariance Improves Siamese TrackingabstractSiamese trackers turn tracking into similarity estimation between a template and the candidate regions in the frame. Mathematically, one of the key ingredients of success of the similarity function is translation equivariance. Non-translation-equivariant architectures induce a positional bias during training, so the location of the target will be hard to recover from the feature space. In real life scenarios, objects undergo various transformations other than translation, such as rotation or scaling. Unless the model has an internal mechanism to handle them, the similarity may degrade. In this paper, we focus on scaling and we aim to equip the Siamese network with additional built-in scale equivariance to capture the natural variations of the target a priori. We develop the theory for scale-equivariant Siamese trackers, and provide a simple recipe for how to make a wide range of existing trackers scale-equivariant. We present SE-SiamFC, a scale-equivariant variant of SiamFC built according to the recipe. We conduct experiments on OTB and VOT benchmarks and on the synthetically generated T-MNIST and S-MNIST datasets. We demonstrate that a built-in additional scale equivariance is useful for visual object tracking. Ivan Sosnovik, Artem Moskalev, Arnold W. M. Smeulders |
WACV | 3 |
| 2020 | Cloth in the Wind: A Case Study of Physical Measurement Through SimulationabstractFor many of the physical phenomena around us, we have developed sophisticated models explaining their behavior. Nevertheless, measuring physical properties from visual observations is challenging due to the high number of causally underlying physical parameters -- including material properties and external forces. In this paper, we propose to measure latent physical properties for cloth in the wind without ever having seen a real example before. Our solution is an iterative refinement procedure with simulation at its core. The algorithm gradually updates the physical model parameters by running a simulation of the observed phenomenon and comparing the current simulation to a real-world observation. The correspondence is measured using an embedding function that maps physically similar examples to nearby points. We consider a case study of cloth in the wind, with curling flags as our leading example - a seemingly simple phenomena but physically highly involved. Based on the physics of cloth and its visual manifestation, we propose an instantiation of the embedding function. For this mapping, modeled as a deep network, we introduce a spectral layer that decomposes a video volume into its temporal spectral power and corresponding frequencies. Our experiments demonstrate that the proposed method compares favorably to prior work on the task of measuring cloth material properties and external wind force from a real-world video. Tom F. H. Runia, Kirill Gavrilyuk, Cees Snoek, Arnold W. M. Smeulders |
CVPR | 4 |
| 2020 | Scale-Equivariant Steerable Networks
Ivan Sosnovik, Michal Szmaja, Arnold W. M. Smeulders |
ICLR | 3 |
| 2020 | Model Decay in Long-Term TrackingabstractTo account for appearance variations, tracking models need to be updated during the course of inference. However, updating the tracker model with adverse bounding box predictions adds an unavoidable bias term to the learning. This bias term, which we refer to as model decay, offsets the learning and causes tracking drift. While its adverse affect might not be visible in short-term tracking, accumulation of this bias over a long-term can eventually lead to a permanent loss of the target. In this paper, we look at the problem of model bias from a mathematical perspective. Further, we briefly examine the effect of various sources of tracking error on model decay, using a correlation filter (ECO) and a Siamese (SINT) tracker. Based on observations and insights, we propose simple additions that help to reduce model decay in long-term tracking. The proposed tracker is evaluated on four long-term and one short-term tracking benchmarks, demonstrating superior accuracy and robustness, even on 30 minute long videos. Efstratios Gavves, Ran Tao 0004, Arnold W. M. Smeulders |
ICPR | 4 |
| 2020 | Tackling Occlusion in Siamese Tracking with Structured DropoutsabstractOcclusion is one of the most difficult challenges in object tracking to model. This is because unlike other challenges, where data augmentation can be of help, occlusion is hard to simulate as the occluding object can be anything in any shape. In this paper, we propose a simple solution to simulate the effects of occlusion in the latent space. Specifically, we present structured dropout to mimick the change in latent codes under occlusion. We present three forms of dropout (channel dropout, segment dropout and slice dropout) with the various forms of occlusion in mind. To demonstrate its effectiveness, the dropouts are incorporated into two modern Siamese trackers (SiamFC and SiamRPN++). The outputs from multiple dropouts are combined using an encoder network to obtain the final prediction. Experiments on several tracking benchmarks show the benefits of structured dropouts, while due to their simplicity requiring only small changes to the existing tracker models. Efstratios Gavves, Arnold W. M. Smeulders |
ICPR | 3 |
| 2020 | Self-Selective Context for Interaction RecognitionabstractHuman-object interaction recognition aims for identifying the relationship between a human subject and an object. Researchers incorporate global scene context into the early layers of deep Convolutional Neural Networks as a solution. They report a significant increase in the performance since generally interactions are correlated with the scene (i.e., riding bicycle on the city street). However, this approach leads to the following problems. It increases the network size in the early layers, therefore not efficient. It leads to noisy filter responses when the scene is irrelevant, therefore not accurate. It only leverages scene context whereas human-object interactions offer a multitude of contexts, therefore incomplete. To circumvent these issues, in this work, we propose Self-Selective Context (SSC). SSC operates on the joint appearance of human-objects and context to bring the most discriminative context(s) into play for recognition. We devise novel contextual features that model the locality of human-object interactions and show that SSC can seamlessly integrate with the State-of-the-art interaction recognition models. Our experiments show that SSC leads to an important increase in interaction recognition performance, while using much fewer parameters. Mert Kilickaya, Noureldien Hussein, Efstratios Gavves, Arnold W. M. Smeulders |
ICPR | 4 |
| 2020 | Explaining with Counter Visual Attributes and ExamplesabstractIn this paper, we aim to explain the decisions of neural networks by utilizing multimodal information. That is counter-intuitive attributes and counter visual examples which appear when perturbed samples are introduced. Different from previous work on interpreting decisions using saliency maps, text, or visual patches we propose to use attributes and counter-attributes, and examples and counter-examples as part of the visual explanations. When humans explain visual decisions they tend to do so by providing attributes and examples. Hence, inspired by the way of human explanations in this paper we provide attribute-based and example-based explanations. Moreover, humans also tend to explain their visual decisions by adding counter-attributes and counter-examples to explain what isnot seen. We introduce directed perturbations in the examples to observe which attribute values change when classifying the examples into the counter classes. This delivers intuitive counter-attributes and counter-examples. Our experiments with both coarse and fine-grained datasets show that attributes provide discriminating and human-understandable intuitive and counter-intuitive explanations. Sadaf Gulshad, Arnold W. M. Smeulders |
ICMR | 2 |
| 2019 | Timeception for Complex Action RecognitionabstractThis paper focuses on the temporal aspect for recognizing human activities in videos; an important visual cue that has long been undervalued. We revisit the conventional definition of activity and restrict it to Complex Action: a set of one-actions with a weak temporal pattern that serves a specific purpose. Related works use spatiotemporal 3D convolutions with fixed kernel size, too rigid to capture the varieties in temporal extents of complex actions, and too short for long-range temporal modeling. In contrast, we use multi-scale temporal convolutions, and we reduce the complexity of 3D convolutions. The outcome is Timeception convolution layers, which reasons about minute-long temporal patterns, a factor of 8 longer than best related works. As a result, Timeception achieves impressive accuracy in recognizing the human activities of Charades, Breakfast Actions and MultiTHUMOS. Further, we demonstrate that Timeception learns long-range temporal dependencies and tolerate temporal extents of complex actions. Noureldien Hussein, Efstratios Gavves, Arnold W. M. Smeulders |
CVPR | 3 |
| 2019 | Repetition EstimationabstractVisual repetition is ubiquitous in our world. It appears in human activity (sports, cooking), animal behavior (a bee’s waggle dance), natural phenomena (leaves in the wind) and in urban environments (flashing lights). Estimating visual repetition from realistic video is challenging as periodic motion is rarely perfectly static and stationary . To better deal with realistic video, we elevate the static and stationary assumptions often made by existing work. Our spatiotemporal filtering approach, established on the theory of periodic motion, effectively handles a wide variety of appearances and requires no learning. Starting from motion in 3D we derive three periodic motion types by decomposition of the motion field into its fundamental components. In addition, three temporal motion continuities emerge from the field’s temporal dynamics. For the 2D perception of 3D motion we consider the viewpoint relative to the motion; what follows are 18 cases of recurrent motion perception. To estimate repetition under all circumstances, our theory implies constructing a mixture of differential motion maps: \(\mathbf {F}\) , \({\varvec{\nabla }}\mathbf {F}\) , \({\varvec{\nabla }}{\varvec{\cdot }} \mathbf {F}\) and \({\varvec{\nabla }}{\varvec{\times }} \mathbf {F}\) . We temporally convolve the motion maps with wavelet filters to estimate repetitive dynamics. Our method is able to spatially segment repetitive motion directly from the temporal filter responses densely computed over the motion maps. For experimental verification of our claims, we use our novel dataset for repetition estimation, better-reflecting reality with non-static and non-stationary repetitive motion. On the task of repetition counting, we obtain favorable results compared to a deep learning alternative. Tom F. H. Runia, Cees Snoek, Arnold W. M. Smeulders |
Int. J. Comput. Vis. | 3 |
| 2018 | Estimating small differences in car-pose from orbits
Berkay Kicanaoglu, Ran Tao 0004, Arnold W. M. Smeulders |
BMVC | 3 |
| 2018 | Real-World Repetition Estimation by Div, Grad and CurlabstractWe consider the problem of estimating repetition in video, such as performing push-ups, cutting a melon or playing violin. Existing work shows good results under the assumption of static and stationary periodicity. As realistic video is rarely perfectly static and stationary, the often preferred Fourier-based measurements is inapt. Instead, we adopt the wavelet transform to better handle non-static and non-stationary video dynamics. From the flow field and its differentials, we derive three fundamental motion types and three motion continuities of intrinsic periodicity in 3D. On top of this, the 2D perception of 3D periodicity considers two extreme viewpoints. What follows are 18 fundamental cases of recurrent perception in 2D. In practice, to deal with the variety of repetitive appearance, our theory implies measuring time-varying flow Ftand its differentials ΔFt, Δ·Ftand Δ×Ftover segmented foreground motion. For experiments, we introduce the new QUVA Repetition dataset, reflecting reality by including non-static and non-stationary videos. On the task of counting repetitions in video, we obtain favorable results compared to a deep learning alternative. Tom F. H. Runia, Cees Snoek, Arnold W. M. Smeulders |
CVPR | 3 |
| 2018 | Long-Term Tracking in the Wild: A Benchmark
Jack Valmadre, Luca Bertinetto, João F. Henriques, Ran Tao 0004, Andrea Vedaldi, Arnold W. M. Smeulders, Philip Torr 0001, Efstratios Gavves |
ECCV (3) | 6 |
| 2018 | i-RevNet: Deep Invertible Networks
Jörn-Henrik Jacobsen, Arnold W. M. Smeulders, Edouard Oyallon |
ICLR (Poster) | 2 |
| 2018 | Asymmetric kernel in Gaussian Processes for learning target variance
Silvia L. Pintea, Jan C. van Gemert, Arnold W. M. Smeulders |
Pattern Recognit. Lett. | 3 |
| 2017 | Dynamic Steerable Blocks in Deep Residual Networks
Jörn-Henrik Jacobsen, Bert De Brabandere, Arnold W. M. Smeulders |
BMVC | 3 |
| 2017 | Unified Embedding and Metric Learning for Zero-Exemplar Event DetectionabstractEvent detection in unconstrained videos is conceived as a content-based video retrieval with two modalities: textual and visual. Given a text describing a novel event, the goal is to rank related videos accordingly. This task is zero-exemplar, no video examples are given to the novel event. Related works train a bank of concept detectors on external data sources. These detectors predict confidence scores for test videos, which are ranked and retrieved accordingly. In contrast, we learn a joint space in which the visual and textual representations are embedded. The space casts a novel event as a probability of pre-defined events. Also, it learns to measure the distance between an event and its related videos. Our model is trained end-to-end on publicly available EventNet. When applied to TRECVID Multimedia Event Detection dataset, it outperforms the state-of-the-art by a considerable margin. Noureldien Hussein, Efstratios Gavves, Arnold W. M. Smeulders |
CVPR | 3 |
| 2017 | Tracking by Natural Language SpecificationabstractThis paper strives to track a target object in a video. Rather than specifying the target in the first frame of a video by a bounding box, we propose to track the object based on a natural language specification of the target, which provides a more natural human-machine interaction as well as a means to improve tracking results. We define three variants of tracking by language specification: one relying on lingual target specification only, one relying on visual target specification based on language, and one leveraging their joint capacity. To show the potential of tracking by natural language specification we extend two popular tracking datasets with lingual descriptions and report experiments. Finally, we also sketch new tracking scenarios in surveillance and other live video streams that become feasible with a lingual specification of the target. Ran Tao 0004, Efstratios Gavves, Cees Snoek, Arnold W. M. Smeulders |
CVPR | 5 |
| 2017 | Searching for A ThingabstractFor humans, one picture usually suffices to identify an object of search. I am looking for this little girl, have you seen her? or Do you have such another one? are two ways to specify a target even to someone who has never seen the object of search before. Searching from one example in digital multimedia retrieval is a hard problem. From the one example one needs to derive an accurate estimate of all accidental variations in the target picture as well as the structural variation of the target in all other potential pictures. From the one example one needs to derive an accurate estimate of all accidental variations the target instance might have. Arnold W. M. Smeulders, Ran Tao 0004 |
ICMR | 1 |
| 2017 | ACM SIGMM Award for Outstanding Technical Contributions to Multimedia Computing, Communications and ApplicationsabstractThe 2017 winner of the prestigious ACM Special Interest Group on Multimedia (SIGMM) award for Outstanding Technical Contributions to Multimedia Computing, Communications and Applications is Prof. Dr. Arnold Smeulders. The award is given in recognition of his outstanding and pioneering contributions to defining and bridging the semantic gap in content-based image retrieval. During the early years of his scientific career Dr. Smeulders studied the invariant fundamentals of lines, shapes, textures and colors. It resulted in several PAMI papers (IEEE Transactions on Pattern Analysis and Machine Intelligence) that are still being cited today. Besides the written record, Arnold always had the drive to showcase academic results in real-world systems. In 1989 he introduced the Diagnostic Encyclopedia Workstation: a system containing 3,000 images from pathology combined with what we would now call a handcrafted ontology. Already then he showed his great ability to generalize results as in 1991 he launched one of the world's first image search engines that combined automatic indexing, interactive retrieval, and evaluation. During this period he was also instrumental in building our community, organizing the first conferences, and defining the semantic gap as the fundamental problem of image retrieval. The end of this era culminated in what became the most cited paper of our discipline: Content-based image retrieval at the end of the early years. .... Arnold W. M. Smeulders |
ACM Multimedia | 1 |
| 2017 | Segmentation models diversity for object proposals
Marco Manfredi, Costantino Grana, Rita Cucchiara, Arnold W. M. Smeulders |
Comput. Vis. Image Underst. | 4 |
| 2017 | Point Light Source Position Estimation From RGB-D Images by Learning Surface AttributesabstractLight source position (LSP) estimation is a difficult yet an important problem in computer vision. A common approach for estimating the LSP assumes Lambert's law. However, in real-world scenes, Lambert's law does not hold for all different types of surfaces. Instead of assuming all that surfaces follow Lambert's law, our approach classifies image surface segments based on their photometric and geometric surface attributes (i.e. glossy, matte, curved, and so on) and assigns weights to image surface segments based on their suitability for LSP estimation. In addition, we propose the use of the estimated camera pose to globally constrain LSP for RGB-D video sequences. Experiments on Boom and a newly collected RGB-D video data sets show that the state-of-the-art methods are outperformed by the proposed method. The results demonstrate that weighting image surface segments based on their attributes outperform the state-of-the-art methods in which the image surface segments are considered to equally contribute. In particular, by using the proposed surface weighting, the angular error for LSP estimation is reduced from 12.6° to 8.2° and 24.6° to 4.8° for Boom and RGB-D video data sets, respectively. Moreover, using the camera pose to globally constrain LSP provides higher accuracy (4.8°) compared with using single frames (8.5°). Sezer Karaoglu, Yang Liu 0009, Theo Gevers, Arnold W. M. Smeulders |
IEEE Trans. Image Process. | 4 |
| 2017 | Words Matter: Scene Text for Image Classification and RetrievalabstractText in natural images typically adds meaning to an object or scene. In particular, text specifies which business places serve drinks (e.g., cafe, teahouse) or food (e.g., restaurant, pizzeria), and what kind of service is provided (e.g., massage, repair). The mere presence of text, its words, and meaning are closely related to the semantics of the object or scene. This paper exploits textual contents in images for fine-grained business place classification and logo retrieval. There are four main contributions. First, we show that the textual cues extracted by the proposed method are effective for the two tasks. Combining the proposed textual and visual cues outperforms visual only classification and retrieval by a large margin. Second, to extract the textual cues, a generic and fully unsupervised word box proposal method is introduced. The method reaches state-of-the-art word detection recall with a limited number of proposals. Third, contrary to what is widely acknowledged in text detection literature, we demonstrate that high recall in word detection is more important than high f-score at least for both tasks considered in this work. Last, this paper provides a large annotated text detection dataset with 10 K images and 27 601 word boxes. Sezer Karaoglu, Ran Tao 0004, Theo Gevers, Arnold W. M. Smeulders |
IEEE Trans. Multim. | 4 |
| 2016 | Structured Receptive Fields in CNNsabstractLearning powerful feature representations with CNNs is hard when training data are limited. Pre-training is one way to overcome this, but it requires large datasets sufficiently similar to the target domain. Another option is to design priors into the model, which can range from tuned hyperparameters to fully engineered representations like Scattering Networks. We combine these ideas into structured receptive field networks, a model which has a fixed filter basis and yet retains the flexibility of CNNs. This flexibility is achieved by expressing receptive fields in CNNs as a weighted sum over a fixed basis which is similar in spirit to Scattering Networks. The key difference is that we learn arbitrary effective filter sets from the basis rather than modeling the filters. This approach explicitly connects classical multiscale image analysis with general CNNs. With structured receptive field networks, we improve considerably over unstructured CNNs for small and medium dataset scenarios as well as over Scattering for large datasets. We validate our findings on ILSVRC2012, Cifar-10, Cifar-100 and MNIST. As a realistic small dataset example, we show state-of-the-art classification results on popular 3D MRI brain-disease datasets where pre-training is difficult due to a lack of large public datasets in a similar domain. Jörn-Henrik Jacobsen, Jan C. van Gemert, Zhongyu Lou, Arnold W. M. Smeulders |
CVPR | 4 |
| 2016 | Siamese Instance Search for TrackingabstractIn this paper we present a tracker, which is radically different from state-of-the-art trackers: we apply no model updating, no occlusion detection, no combination of trackers, no geometric matching, and still deliver state-of-the-art tracking performance, as demonstrated on the popular online tracking benchmark (OTB) and six very challenging YouTube videos. The presented tracker simply matches the initial patch of the target in the first frame with candidates in a new frame and returns the most similar patch by a learned matching function. The strength of the matching function comes from being extensively trained generically, i.e., without any data of the target, using a Siamese deep neural network, which we design for tracking. Once learned, the matching function is used as is, without any adapting, to track previously unseen targets. It turns out that the learned matching function is so powerful that a simple tracker built upon it, coined Siamese INstance search Tracker, SINT, which only uses the original observation of the target from the first frame, suffices to reach state-of-the-art performance. Further, we show the proposed tracker even allows for target re-identification after the target was absent for a complete video shot. Ran Tao 0004, Efstratios Gavves, Arnold W. M. Smeulders |
CVPR | 3 |
| 2016 | Featureless: Bypassing feature extraction in action categorizationabstractThis method introduces an efficient manner of learning action categories without the need of feature estimation. The approach starts from low-level values, in a similar style to the successful CNN methods. However, rather than extracting general image features, we learn to predict specific video representations from raw video data. The benefit of such an approach is that at the same computational expense it can predict 2D video representations as well as 3D ones, based on motion. The proposed model relies on discriminative Wald-boost, which we enhance to a multiclass formulation for the purpose of learning video representations. The suitability of the proposed approach as well as its time efficiency are tested on the UCF11 action recognition dataset. Silvia L. Pintea, Pascal Mettes, Jan C. van Gemert, Arnold W. M. Smeulders |
ICIP | 4 |
| 2016 | Large scale Gaussian Process for overlap-based object proposal scoring
Silvia L. Pintea, Sezer Karaoglu, Jan C. van Gemert, Arnold W. M. Smeulders |
Comput. Vis. Image Underst. | 4 |
| 2015 | Attributes and categories for generic instance search from one exampleabstractThis paper aims for generic instance search from one example where the instance can be an arbitrary 3D object like shoes, not just near-planar and one-sided instances like buildings and logos. Firstly, we evaluate state-of-the-art instance search methods on this problem. We observe that what works for buildings loses its generality on shoes. Secondly, we propose to use automatically learned category-specific attributes to address the large appearance variations present in generic instance search. On the problem of searching among instances from the same category as the query, the category-specific attributes outperform existing approaches by a large margin. On a shoe dataset containing 6624 shoe images recorded from all viewing angles, we improve the performance from 36.73 to 56.56 using category-specific attributes. Thirdly, we extend our methods to search objects without restricting to the specifically known category. We show the combination of category-level information and the category-specific attributes is superior to combining category-level information with low-level features such as Fisher vector. Ran Tao 0004, Arnold W. M. Smeulders, Shih-Fu Chang |
CVPR | 2 |
| 2015 | Local Alignments for Fine-Grained Categorization
Efstratios Gavves, Basura Fernando, Cees Snoek, Arnold W. M. Smeulders, Tinne Tuytelaars |
Int. J. Comput. Vis. | 4 |
| 2014 | Fisher and VLAD with FLAIRabstractA major computational bottleneck in many current algorithms is the evaluation of arbitrary boxes. Dense local analysis and powerful bag-of-word encodings, such as Fisher vectors and VLAD, lead to improved accuracy at the expense of increased computation time. Where a simplification in the representation is tempting, we exploit novel representations while maintaining accuracy. We start from state-of-the-art, fast selective search, but our method will apply to any initial box-partitioning. By representing the picture as sparse integral images, one per codeword, we achieve a Fast Local Area Independent Representation. FLAIR allows for very fast evaluation of any box encoding and still enables spatial pooling. In FLAIR we achieve exact VLAD's difference coding, even with L2 and power-norms. Finally, by multiple codeword assignments, we achieve exact and approximate Fisher vectors with FLAIR. The results are a 18x speedup, which enables us to set a new state-of-the-art on the challenging 2010 PASCAL VOC objects and the fine-grained categorization of the CUB-2011 200 bird species. Plus, we rank number one in the official ImageNet 2013 detection challenge. Koen E. A. van de Sande, Cees Snoek, Arnold W. M. Smeulders |
CVPR | 3 |
| 2014 | Locality in Generic Instance Search from One ExampleabstractThis paper aims for generic instance search from a single example. Where the state-of-the-art relies on global image representation for the search, we proceed by including locality at all steps of the method. As the first novelty, we consider many boxes per database image as candidate targets to search locally in the picture using an efficient point-indexed representation. The same representation allows, as the second novelty, the application of very large vocabularies in the powerful Fisher vector and VLAD to search locally in the feature space. As the third novelty we propose an exponential similarity function to further emphasize locality in the feature space. Locality is advantageous in instance search as it will rest on the matching unique details. We demonstrate a substantial increase in generic instance search performance from one example on three standard datasets with buildings, logos, and scenes from 0.443 to 0.620 in mAP. Ran Tao 0004, Efstratios Gavves, Cees Snoek, Arnold W. M. Smeulders |
CVPR | 4 |
| 2014 | Déjà Vu: - Motion Prediction in Static Images
Silvia L. Pintea, Jan C. van Gemert, Arnold W. M. Smeulders |
ECCV (3) | 3 |
| 2014 | Visual dictionaries in the Brain: Comparing HMAX and BOWabstractThe human visual system is thought to use features of intermediate complexity for scene representation. How the brain computationally represents intermediate features is, however, still unclear. Here we tested and compared two widely used computational models — the biologically plausible HMAX model and Bag of Words (BoW) model from computer vision against human brain activity. These computational models use visual dictionaries, candidate features of intermediate complexity, to represent visual scenes, and the models have been proven effective in automatic object and scene recognition. We analyzed where in the brain and to what extent human fMRI responses to natural scenes can be accounted for by the HMAX and BoW representations. Voxel-wise application of a distance-based variation partitioning method reveals that HMAX explains significant brain activity in early visual regions and also in higher regions such as LO, TO while the BoW primarily explains brain acitvity in the early visual area. Notably, both HMAX and BoW explain the most brain activity in higher areas such as V4 and TO. These results suggest that visual dictionaries might provide a suitable computation for the representation of intermediate features in the brain. Kandan Ramakrishnan, Iris I. A. Groen, H. Steven Scholte, Arnold W. M. Smeulders, Sennay Ghebreab |
ICME | 4 |
| 2014 | Evaluating multimedia features and fusion for example-based event detectionabstractMultimedia event detection (MED) is a challenging problem because of the heterogeneous content and variable quality found in large collections of Internet videos. To study the value of multimedia features and fusion for representing and learning events from a set of example video clips, we created SESAME, a system for video SEarch with Speed and Accuracy for Multimedia Events. SESAME includes multiple bag-of-words event classifiers based on single data types: low-level visual, motion, and audio features; high-level semantic visual concepts; and automatic speech recognition. Event detection performance was evaluated for each event classifier. The performance of low-level visual and motion features was improved by the use of difference coding. The accuracy of the visual concepts was nearly as strong as that of the low-level visual features. Experiments with a number of fusion methods for combining the event detection scores from these classifiers revealed that simple fusion methods, such as arithmetic mean, perform as well as or better than other, more complex fusion methods. SESAME’s performance in the 2012 TRECVID MED evaluation was one of the best reported. Gregory K. Myers, Ramesh Nallapati, Julien van Hout, Stephanie Pancoast, Ramakant Nevatia, Chen Sun 0002, AmirHossein Habibian, Dennis C. Koelma, Koen E. A. van de Sande, Arnold W. M. Smeulders, Cees Snoek |
Mach. Vis. Appl. | 10 |
| 2014 | Visual Tracking: An Experimental SurveyabstractThere is a large variety of trackers, which have been proposed in the literature during the last two decades with some mixed success. Object tracking in realistic scenarios is a difficult problem, therefore, it remains a most active area of research in computer vision. A good tracker should perform well in a large number of videos involving illumination changes, occlusion, clutter, camera motion, low contrast, specularities, and at least six more aspects. However, the performance of proposed trackers have been evaluated typically on less than ten videos, or on the special purpose datasets. In this paper, we aim to evaluate trackers systematically and experimentally on 315 video fragments covering above aspects. We selected a set of nineteen trackers to include a wide variety of algorithms often cited in literature, supplemented with trackers appearing in 2010 and 2011 for which the code was publicly available. We demonstrate that trackers can be evaluated objectively by survival curves, Kaplan Meier statistics, and Grubs testing. We find that in the evaluation practice the F-score is as effective as the object tracking accuracy (OTA) score. The analysis under a large variety of circumstances provides objective insight into the strengths and weaknesses of trackers. Arnold W. M. Smeulders, Dung Manh Chu, Rita Cucchiara, Simone Calderara, Afshin Dehghan, Mubarak Shah |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2013 | Fine-Grained Categorization by AlignmentsabstractThe aim of this paper is fine-grained categorization without human interaction. Different from prior work, which relies on detectors for specific object parts, we propose to localize distinctive details by roughly aligning the objects using just the overall shape, since implicit to fine-grained categorization is the existence of a super-class shape shared among all classes. The alignments are then used to transfer part annotations from training images to test images (supervised alignment), or to blindly yet consistently segment the object in a number of regions (unsupervised alignment). We furthermore argue that in the distinction of fine grained sub-categories, classification-oriented encodings like Fisher vectors are better suited for describing localized information than popular matching oriented features like HOG. We evaluate the method on the CU-2011 Birds and Stanford Dogs fine-grained datasets, outperforming the state-of-the-art. Efstratios Gavves, Basura Fernando, Cees Snoek, Arnold W. M. Smeulders, Tinne Tuytelaars |
ICCV | 4 |
| 2013 | Codemaps - Segment, Classify and Search Objects LocallyabstractIn this paper we aim for segmentation and classification of objects. We propose codemaps that are a joint formulation of the classification score and the local neighborhood it belongs to in the image. We obtain the codemap by reordering the encoding, pooling and classification steps over lattice elements. Other than existing linear decompositions who emphasize only the efficiency benefits for localized search, we make three novel contributions. As a preliminary, we provide a theoretical generalization of the sufficient mathematical conditions under which image encodings and classification becomes locally decomposable. As first novelty we introduce l2 normalization for arbitrarily shaped image regions, which is fast enough for semantic segmentation using our Fisher codemaps. Second, using the same lattice across images, we propose kernel pooling which embeds nonlinearities into codemaps for object classification by explicit or approximate feature mappings. Results demonstrate that l2 normalized Fisher codemaps improve the state-of-the-art in semantic segmentation for PASCAL VOC. For object classification the addition of nonlinearities brings us on par with the state-of-the-art, but is 3x faster. Because of the codemaps' inherent efficiency, we can reach significant speed-ups for localized search as well. We exploit the efficiency gain for our third novelty: object seg- ment retrieval using a single query image only. Efstratios Gavves, Koen E. A. van de Sande, Cees Snoek, Arnold W. M. Smeulders |
ICCV | 5 |
| 2013 | Searching Things in Large Sets of Images
Arnold W. M. Smeulders |
SOFSEM | 1 |
| 2013 | Selective Search for Object Recognition
Jasper R. R. Uijlings, Koen E. A. van de Sande, Theo Gevers, Arnold W. M. Smeulders |
Int. J. Comput. Vis. | 4 |
| 2013 | Bootstrapping Visual Categorization With Relevant NegativesabstractLearning classifiers for many visual concepts are important for image categorization and retrieval. As a classifier tends to misclassify negative examples which are visually similar to positive ones, inclusion of such misclassified and thus relevant negatives should be stressed during learning. User-tagged images are abundant online, but which images are the relevant negatives remains unclear. Sampling negatives at random is the de facto standard in the literature. In this paper, we go beyond random sampling by proposing Negative Bootstrap. Given a visual concept and a few positive examples, the new algorithm iteratively finds relevant negatives. Per iteration, we learn from a small proportion of many user-tagged images, yielding an ensemble of meta classifiers. For efficient classification, we introduce Model Compression such that the classification time is independent of the ensemble size. Compared with the state of the art, we obtain relative gains of 14% and 18% on two present-day benchmarks in terms of mean average precision. For concept search in one million images, model compression reduces the search time from over 20 h to approximately 6 min. The effectiveness and efficiency, without the need of manually labeling any negatives, make negative bootstrap appealing for learning better visual concept classifiers. Xirong Li 0001, Cees Snoek, Marcel Worring, Dennis C. Koelma, Arnold W. M. Smeulders |
IEEE Trans. Multim. | 5 |
| 2012 | Convex reduction of high-dimensional kernels for visual classificationabstractLimiting factors of fast and effective classifiers for large sets of images are their dependence on the number of images analyzed and the dimensionality of the image representation. Considering the growing number of images as a given, we aim to reduce the image feature dimensionality in this paper. We propose reduced linear kernels that use only a portion of the dimensions to reconstruct a linear kernel. We formulate the search for these dimensions as a convex optimization problem, which can be solved efficiently. Different from existing kernel reduction methods, our reduced kernels are faster and maintain the accuracy benefits from non-linear embedding methods that mimic non-linear SVMs. We show these properties on both the Scenes and PASCAL VOC 2007 datasets. In addition, we demonstrate how our reduced kernels allow to compress Fisher vector for use with non-linear embeddings, leading to high accuracy. What is more, without using any labeled examples the selected and weighed kernel dimensions appear to correspond to visually meaningful patches in the images. Efstratios Gavves, Cees Snoek, Arnold W. M. Smeulders |
CVPR | 3 |
| 2012 | Fusing concept detection and geo context for visual searchabstractGiven the proliferation of geo-tagged images, the question of how to exploit geo tags and the underlying geo context for visual search is emerging. Based on the observation that the importance of geo context varies over concepts, we propose a concept-based image search engine which fuses visual concept detection and geo context in a concept-dependent manner. Compared to individual content-based and geo-based concept detectors and their uniform combination, concept-dependent fusion shows improvements. Moreover, since the proposed search engine is trained on social-tagged images alone without the need of human interaction, it is flexible to cope with many concepts. Search experiments on 101 popular visual concepts justify the viability of the proposed solution. In particular, for 79 out of the 101 concepts, the learned weights yield improvements over the uniform weights, with a relative gain of at least 5% in terms of average precision. Xirong Li 0001, Cees Snoek, Marcel Worring, Arnold W. M. Smeulders |
ICMR | 4 |
| 2012 | All vehicles are cars: subclass preferences in container conceptsabstractThis paper investigates the natural bias humans display when labeling images with a container label like vehicle or carnivore. Using three container concepts as subtree root nodes, and all available concepts between these roots and the images from the ImageNet Large Scale Visual Recognition Challenge (ILSVRC) dataset, we analyze the differences between the images labeled at these varying levels of abstraction and the union of their constituting leaf nodes. We find that for many container concepts, a strong preference for one or a few different constituting leaf nodes occurs. These results indicate that care is needed when using hierarchical knowledge in image classification: if the aim is to classify vehicles the way humans do, then cars and buses may be the only correct results. Daan T. J. Vreeswijk, Cees Snoek, Koen E. A. van de Sande, Arnold W. M. Smeulders |
ICMR | 4 |
| 2012 | Visual synonyms for landmark image retrieval
Efstratios Gavves, Cees Snoek, Arnold W. M. Smeulders |
Comput. Vis. Image Underst. | 3 |
| 2012 | The Visual Extent of an Object - Suppose We Know the Object LocationsabstractThe visual extent of an object reaches beyond the object itself. This is a long standing fact in psychology and is reflected in image retrieval techniques which aggregate statistics from the whole image in order to identify the object within. However, it is unclear to what degree and how the visual extent of an object affects classification performance. In this paper we investigate the visual extent of an object on the Pascal VOC dataset using a Bag-of-Words implementation with (colour) SIFT descriptors. Our analysis is performed from two angles. (a) Not knowing the object location, we determine where in the image the support for object classification resides. We call this the normal situation. (b) Assuming that the object location is known, we evaluate the relative potential of the object and its surround, and of the object border and object interior. We call this the ideal situation. Our most important discoveries are: (i) Surroundings can adequately distinguish between groups of classes: furniture, animals, and land-vehicles. For distinguishing categories within one group the surroundings become a source of confusion. (ii) The physically rigid plane, bike, bus, car, and train classes are recognised by interior boundaries and shape, not by texture. The non-rigid animals dog, cat, cow, and sheep are recognised primarily by texture, i.e. fur, as their projected shape varies greatly. (iii) We confirm an early observation from human psychology (Biederman in Perceptual Organization, pp. 213–263, 1981 ): in the ideal situation with known object locations, recognition is no longer improved by considering surroundings. In contrast, in the normal situation with unknown object locations, the surroundings significantly contribute to the recognition of most classes. Jasper R. R. Uijlings, Arnold W. M. Smeulders, Remko J. H. Scha |
Int. J. Comput. Vis. | 2 |
| 2012 | Content-Based Analysis Improves Audiovisual Archive RetrievalabstractContent-based video retrieval is maturing to the point where it can be used in real-world retrieval practices. One such practice is the audiovisual archive, whose users increasingly require fine-grained access to broadcast television content. In this paper, we take into account the information needs and retrieval data already present in the audiovisual archive, and demonstrate that retrieval performance can be significantly improved when content-based methods are applied to search. To the best of our knowledge, this is the first time that the practice of an audiovisual archive has been taken into account for quantitative retrieval evaluation. To arrive at our main result, we propose an evaluation methodology tailored to the specific needs and circumstances of the audiovisual archive, which are typically missed by existing evaluation initiatives. We utilize logged searches, content purchases, session information, and simulators to create realistic query sets and relevance judgments. To reflect the retrieval practice of both the archive and the video retrieval community as closely as possible, our experiments with three video search engines incorporate archive-created catalog entries as well as state-of-the-art multimedia content analysis results. A detailed query-level analysis indicates that individual content-based retrieval methods such as transcript-based retrieval and concept-based retrieval yield approximately equal performance gains. When combined, we find that content-based video retrieval incorporated into the archive's practice results in significant performance increases for shot retrieval and for retrieving entire television programs. The time has come for audiovisual archives to start accommodating content-based video retrieval methods into their daily practice. Bouke Huurnink, Cees Snoek, Maarten de Rijke, Arnold W. M. Smeulders |
IEEE Trans. Multim. | 4 |
| 2012 | Harvesting Social Images for Bi-Concept SearchabstractSearching for the co-occurrence of two visual concepts in unlabeled images is an important step towards answering complex user queries. Traditional visual search methods use combinations of the confidence scores of individual concept detectors to tackle such queries. In this paper we introduce the notion of bi-concepts, a new concept-based retrieval method that is directly learned from social-tagged images. As the number of potential bi-concepts is gigantic, manually collecting training examples is infeasible. Instead, we propose a multimedia framework to collect de-noised positive as well as informative negative training examples from the social web, to learn bi-concept detectors from these examples, and to apply them in a search engine for retrieving bi-concepts in unlabeled images. We study the behavior of our bi-concept search engine using 1.2 M social-tagged images as a data source. Our experiments indicate that harvesting examples for bi-concepts differs from traditional single-concept methods, yet the examples can be collected with high accuracy using a multi-modal approach. We find that directly learning bi-concepts is better than oracle linear fusion of single-concept detectors, with a relative improvement of 100%. This study reveals the potential of learning high-order semantics from social images, for free, suggesting promising new lines of research. Xirong Li 0001, Cees Snoek, Marcel Worring, Arnold W. M. Smeulders |
IEEE Trans. Multim. | 4 |
| 2011 | Segmentation as selective search for object recognitionabstractFor object recognition, the current state-of-the-art is based on exhaustive search. However, to enable the use of more expensive features and classifiers and thereby progress beyond the state-of-the-art, a selective search strategy is needed. Therefore, we adapt segmentation as a selective search by reconsidering segmentation: We propose to generate many approximate locations over few and precise object delineations because (1) an object whose location is never generated can not be recognised and (2) appearance and immediate nearby context are most effective for object recognition. Our method is class-independent and is shown to cover 96.7% of all objects in the Pascal VOC 2007 test set using only 1,536 locations per image. Our selective search enables the use of the more expensive bag-of-words method which we use to substantially improve the state-of-the-art by up to 8.5% for 8 out of 20 classes on the Pascal VOC 2010 detection challenge. Koen E. A. van de Sande, Jasper R. R. Uijlings, Theo Gevers, Arnold W. M. Smeulders |
ICCV | 4 |
| 2011 | Social negative bootstrapping for visual categorizationabstractTo learn classifiers for many visual categories, obtaining labeled training examples in an efficient way is crucial. Since a classifier tends to misclassify negative examples which are visually similar to positive examples, inclusion of such informative negatives should be stressed in the learning process. However, they are unlikely to be hit by random sampling, the de facto standard in literature. In this paper, we go beyond random sampling by introducing a novel social negative bootstrapping approach. Given a visual category and a few positive examples, the proposed approach adaptively and iteratively harvests informative negatives from a large amount of social-tagged images. To label negative examples without human interaction, we design an effective virtual labeling procedure based on simple tag reasoning. Virtual labeling, in combination with adaptive sampling, enables us to select the most misclassified negatives as the informative samples. Learning from the positive set and the informative negative sets results in visual classifiers with higher accuracy. Experiments on two present-day image benchmarks employing 650K virtually labeled negative examples show the viability of the proposed approach. On a popular visual categorization benchmark our precision at 20 increases by 34%, compared to baselines trained on randomly sampled negatives. We achieve more accurate visual categorization without the need of manually labeling any negatives. Xirong Li 0001, Cees Snoek, Marcel Worring, Arnold W. M. Smeulders |
ICMR | 4 |
| 2011 | Instant Bag-of-Words served on a laptopabstractThis demo showcases our realtime implementation of concept classification using the Bag-of-Words method embedded within MediaTable, our interactive categorization tool for large multimedia collections. MediaTable allows the users to open images from disk or download these directly from the internet. Each image is then processed using the Bag-of-Words method, which computes classification scores for 20 distinct concepts classes on the fly. These are then seamlessly displayed in the interface. Jasper R. R. Uijlings, Ork de Rooij, Daan Odijk, Arnold W. M. Smeulders, Marcel Worring |
ICMR | 4 |
| 2011 | Personalizing automated image annotation using cross-entropyabstractAnnotating the increasing amounts of user-contributed images in a personalized manner is in great demand. However, this demand is largely ignored by the mainstream of automated image annotation research. In this paper we aim for personalizing automated image annotation by jointly exploiting personalized tag statistics and content-based image annotation. We propose a cross-entropy based learning algorithm which personalizes a generic annotation model by learning from a user's multimedia tagging history. Using cross-entropy-minimization based Monte Carlo sampling, the proposed algorithm optimizes the personalization process in terms of a performance measurement which can be flexibly chosen. Automatic image annotation experiments with 5,315 realistic users in the social web show that the proposed method compares favorably to a generic image annotation method and a method using personalized tag statistics only. For 4,442 users the performance improves, where for 1,088 users the absolute performance gain is at least 0.05 in terms of average precision. The results show the value of the proposed method. Xirong Li 0001, Efstratios Gavves, Cees Snoek, Marcel Worring, Arnold W. M. Smeulders |
ACM Multimedia | 5 |
| 2011 | Internet video searchabstractIn this tutorial, we focus on the challenges in internet video search, present methods how to achieve state-of-the-art performance while maintaining efficient execution, and indicate how to obtain improvements in the near future. Moreover, we give an overview of the latest developments and future trends in the field on the basis of the TRECVID competition - the leading competition for video search engines run by NIST - where we have achieved consistent top performance over the past years, including the 2008, 2009 and 2010 editions. Cees Snoek, Arnold W. M. Smeulders |
ACM Multimedia | 2 |
| 2011 | Text and image subject classifiers: dense works betterabstractWe investigate the feasibility of training visual concept detectors for such abstract subject categories as biology and history with the aim of employing these for full-text to image linking. We show that using dense sampling methods can lead to image classifiers that perform well enough for interactive search. Echoing this dense sampling in the image domain, we also show that using term frequencies as text features outperforms using a topic abstraction method. Finally, we use these monomodal classifiers for the task of linking texts to images, improving more than 50% over the state-of-the-art, thereby showing that dense is better. Daan T. J. Vreeswijk, Bouke Huurnink, Arnold W. M. Smeulders |
ACM Multimedia | 3 |
| 2011 | Identifying distributed and overlapping clusters of hemodynamic synchrony in fMRI data setsabstractNatural sensory stimuli elicit complex brain responses that manifest in fMRI as widely distributed and overlapping clusters of hemodynamic responses. We propose a statistical signal processing method for finding synchronous hemodynamic activity that directly or transiently reflects information about the experimental condition. When applied to fMRI data, the method searches for voxels with activation patterns exhibiting high coherence and simultaneously high variance across brain scans. The crux of the method is functional principal component analysis (fPCA) of activation patterns stored in a two-dimensional data matrix, with rows and columns representing voxels and scans, respectively. Without external information, fPCA is performed directly on this data matrix. Otherwise, the data matrix is first transformed to highlight a specific source of variation, enabling fully or partially supervised fPCA with a single parameter determining the degree of supervision. We evaluated our method on a public benchmark of fMRI scans of subjects viewing natural movies. Our method turns out to be very suitable for flexibly uncovering distributed and overlapping hemodynamic patterns that distinguish well between experimental conditions or cognitive states. Sennay Ghebreab, Arnold W. M. Smeulders |
Pattern Anal. Appl. | 2 |
| 2010 | Thirteen Hard Cases in Visual TrackingabstractVisual tracking is a fundamental task in computer vision. However there has been no systematic way of analyzing visual trackers so far. In this paper we propose a method that can help researchers determine strengths and weaknesses of any visual tracker. To this end, we consider visual tracking as an isolated problem and decompose it into fundamental and independent subproblems. Each subproblem is designed to associate with a different tracking circumstance. By evaluating a visual tracker onto a specific subproblem, we can determine how good it is with respect to that dimension. In total we come up with thirteen subproblems in our decomposition. We demonstrate the use of our proposed method by analyzing working conditions of two state-of-the-art trackers. Dung Manh Chu, Arnold W. M. Smeulders |
AVSS | 2 |
| 2010 | Video search engines: advances in multimedia retrieval, part ii [1]abstractIn this tutorial, we focus on the challenges in video search, present methods how to achieve state-of-the-art performance, and indicate how to obtain improvements in the near future. Moreover, we give an overview of the latest developments and future trends in the field on the basis of the TRECVID competition - the leading competition for video search engines run by NIST - where we have achieved consistent top performance over the years, including the 2008 and 2009 editions. Cees Snoek, Arnold W. M. Smeulders |
ACM Multimedia | 2 |
| 2010 | Comparing compact codebooks for visual categorization
Jan C. van Gemert, Cees Snoek, Cor J. Veenman, Arnold W. M. Smeulders, Jan-Mark Geusebroek |
Comput. Vis. Image Underst. | 4 |
| 2010 | Visual Word AmbiguityabstractThis paper studies automatic image classification by modeling soft assignment in the popular codebook model. The codebook model describes an image as a bag of discrete visual words selected from a vocabulary, where the frequency distributions of visual words in an image allow classification. One inherent component of the codebook model is the assignment of discrete visual words to continuous image features. Despite the clear mismatch of this hard assignment with the nature of continuous features, the approach has been successfully applied for some years. In this paper, we investigate four types of soft assignment of visual words to image features. We demonstrate that explicitly modeling visual word assignment ambiguity improves classification performance compared to the hard assignment of the traditional codebook model. The traditional codebook model is compared against our method for five well-known data sets: 15 natural scenes, Caltech-101, Caltech-256, and Pascal VOC 2007/2008. We demonstrate that large codebook vocabulary sizes completely deteriorate the performance of the traditional model, whereas the proposed model performs consistently. Moreover, we show that our method profits in high-dimensional feature spaces and reaps higher benefits when increasing the number of image categories. Jan C. van Gemert, Cor J. Veenman, Arnold W. M. Smeulders, Jan-Mark Geusebroek |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2010 | Stages as Models of Scene GeometryabstractReconstruction of 3D scene geometry is an important element for scene understanding, autonomous vehicle and robot navigation, image retrieval, and 3D television. We propose accounting for the inherent structure of the visual world when trying to solve the scene reconstruction problem. Consequently, we identify geometric scene categorization as the first step toward robust and efficient depth estimation from single images. We introduce 15 typical 3D scene geometries called stages, each with a unique depth profile, which roughly correspond to a large majority of broadcast video frames. Stage information serves as a first approximation of global depth, narrowing down the search space in depth estimation and object localization. We propose different sets of low-level features for depth estimation, and perform stage classification on two diverse data sets of television broadcasts. Classification results demonstrate that stages can often be efficiently learned from low-dimensional image representations. Vladimir Nedovic, Arnold W. M. Smeulders, André Redert, Jan-Mark Geusebroek |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2010 | Real-Time Visual Concept ClassificationabstractAs datasets grow increasingly large in content-based image and video retrieval, computational efficiency of concept classification is important. This paper reviews techniques to accelerate concept classification, where we show the trade-off between computational efficiency and accuracy. As a basis, we use the Bag-of-Words algorithm that in the 2008 benchmarks of TRECVID and PASCAL lead to the best performance scores. We divide the evaluation in three steps: 1) Descriptor Extraction, where we evaluate SIFT, SURF, DAISY, and Semantic Textons. 2) Visual Word Assignment, where we compare a k-means visual vocabulary with a Random Forest and evaluate subsampling, dimension reduction with PCA, and division strategies of the Spatial Pyramid. 3) Classification, where we evaluate the χ2, RBF, and Fast Histogram Intersection kernel for the SVM. Apart from the evaluation, we accelerate the calculation of densely sampled SIFT and SURF, accelerate nearest neighbor assignment, and improve accuracy of the Histogram Intersection kernel. We conclude by discussing whether further acceleration of the Bag-of-Words pipeline is possible. Our results lead to a 7-fold speed increase without accuracy loss, and a 70-fold speed increase with 3% accuracy loss. The latter system does classification in real-time, which opens up new applications for automatic concept classification. For example, this system permits five standard desktop PCs to automatically tag for 20 classes all images that are currently uploaded to Flickr. Jasper R. R. Uijlings, Arnold W. M. Smeulders, Remko J. H. Scha |
IEEE Trans. Multim. | 2 |
| 2009 | What is the spatial extent of an object?abstractThis paper discusses the question: Can we improve the recognition of objects by using their spatial context? We start from Bag-of-Words models and use the Pascal 2007 dataset. We use the rough object bounding boxes that come with this dataset to investigate the fundamental gain context can bring. Our main contributions are: (I) The result of Zhang et al. in CVPR07 that context is superfluous derived from the Pascal 2005 data set of 4 classes does not generalize to this dataset. For our larger and more realistic dataset context is important indeed. (II) Using the rough bounding box to limit or extend the scope of an object during both training and testing, we find that the spatial extent of an object is determined by its category: (a) well-defined, rigid objects have the object itself as the preferred spatial extent. (b) Non-rigid objects have an unbounded spatial extent : all spatial extents produce equally good results. (c) Objects primarily categorised based on their function have the whole image as their spatial extent. Finally, (III) using the rough bounding box to treat object and context separately, we find that the upper bound of improvement is 26% (12% absolute) in terms of mean average precision, and this bound is likely to be higher if the localisation is done using segmentation. It is concluded that object localisation, if done sufficiently precise, helps considerably in the recognition of objects for the Pascal 2007 dataset. Jasper R. R. Uijlings, Arnold W. M. Smeulders, Remko J. H. Scha |
CVPR | 2 |
| 2009 | Periodic event detection and recognition in videoabstractPeriodicity attracts special attention in human cognition. Hence it is important to consider that in automatic analysis of motion events. This paper presents a method for representing periodic events with which events can be compared irrespective of their duration. The effectiveness of such a representation is verified with event classification. E. P. Vivek, Erik Pogalin, Arnold W. M. Smeulders |
ICASSP | 3 |
| 2009 | A Biologically Plausible Model for Rapid Natural Scene IdentificationabstractContrast statistics of the majority of natural images conform to a Weibull distribution. This property of natural images may facilitate efficient and very rapid extraction of a scenes visual gist. Here we investigate whether a neural response model based on the Weibull contrast distribution captures visual information that humans use to rapidly identify natural scenes. In a learning phase, we measure EEG activity of 32 subjects viewing brief flashes of 800 natural scenes. From these neural measurements and the contrast statistics of the natural image stimuli, we derive an across subject Weibull response model. We use this model to predict the responses to a large set of new scenes and estimate which scene the subject viewed by finding the best match between the model predictions and the observed EEG responses. In almost 90 percent of the cases our model accurately predicts the observed scene. Moreover, in most failed cases, the scene mistaken for the observed scene is visually similar to the observed scene itself. These results suggest that Weibull contrast statistics of natural images contain a considerable amount of scene gist information to warrant rapid identification of natural images. Sennay Ghebreab, H. Steven Scholte, Victor A. F. Lamme, Arnold W. M. Smeulders |
NIPS | 4 |
| 2008 | Visual quasi-periodicityabstractPeriodicity is at the core of the recognition of many actions. This paper takes the following steps to detect and measure periodicity. 1) We establish a conceptual framework of classifying periodicity in 10 essential cases, the most important of which are flashing (of a traffic light), pulsing (of an anemone), swinging (of wings), spinning (of a swimmer), turning (of a conductor), shuttling (of a brush), drifting (of an escalator) and thrusting (of a kangaroo). 2) We present an algorithm to detect all cases by the one and the same algorithm. It tracks the object independent of the objectpsilas appearance, then performs probabilistic PCA and spectral analysis followed by detection and frequency measurement. The method shows good performance with fixed parameters for examples of all above cases assembled from the Internet. 3) Application of the method, completely unaltered, to a random half hour of CNN news has led to an 80% score. Erik Pogalin, Arnold W. M. Smeulders, Andrew H. C. Thean |
CVPR | 2 |
| 2008 | Kernel Codebooks for Scene Categorization
Jan C. van Gemert, Jan-Mark Geusebroek, Cor J. Veenman, Arnold W. M. Smeulders |
ECCV (3) | 4 |
| 2008 | Analyzing video concept detectors visuallyabstractIn this demonstration we showcase an interactive analysis tool for researchers working on concept-based video retrieval. By visualizing intermediate concept detection analysis stages, the tool aids in understanding the success and failure of video concept detection methods. We demonstrate the tool on the domain of pop concert video. Cees Snoek, Richard van Balen, Dennis C. Koelma, Arnold W. M. Smeulders, Marcel Worring |
ICME | 4 |
| 2008 | Quadratic boosting
Thang V. Pham, Arnold W. M. Smeulders |
Pattern Recognit. | 2 |
| 2007 | Predictive Modeling of fMRI Brain States Using Functional Canonical Correlation Analysis
Sennay Ghebreab, Arnold W. M. Smeulders, Pieter W. Adriaans |
AIME | 2 |
| 2007 | Recent Advances and Challenges of Semantic Image/Video SearchabstractWe present an overview of recent advances and major challenges in image and video search, with a specific focus on large-scale semantic concept detection and indexing. Such semantic indexing paradigm has been driven by the increasing availability of the large resources of corpora, novel labeling approaches, innovative image features, and machine learning techniques for visual content recognition. We will discus key approaches, recent results, and novel applications in text-to-concept semantic search and multi-modal retrieval models. Open issues and major opportunities are also presented. Shih-Fu Chang, Wei-Ying Ma, Arnold W. M. Smeulders |
ICASSP (4) | 3 |
| 2007 | The Mediamill Semantic Video Search EngineabstractIn this paper we present the methods underlying the MediaMill semantic video search engine. The basis for the engine is a semantic indexing process which is currently based on a lexicon of 491 concept detectors. To support the user in navigating the collection, the system defines a visual similarity space, a semantic similarity space, a semantic thread space, and browsers to explore them. We compare the different browsers and their utility within the TRECVID benchmark. In 2005, we obtained a top-3 result for 19 out of 24 search topics. In 2006 for 14 out of 24. Marcel Worring, Cees Snoek, Ork de Rooij, Giang P. Nguyen, Arnold W. M. Smeulders |
ICASSP (4) | 5 |
| 2007 | Depth Information by Stage ClassificationabstractRecently, methods for estimating 3D scene geometry or absolute scene depth information from 2D image content have been proposed. However, general applicability of these methods in depth estimation may not be realizable, as inconsistencies may be introduced due to a large variety of possible pictorial content. We identify scene categorization as the first step towards efficient and robust depth estimation from single images. To that end, we describe a limited number of typical 3D scene geometries, called stages, each having a unique depth pattern and thus providing a specific context for stage objects. This type of scene information narrows down the possibilities with respect to individual objects' locations, scales and identities. We show how these stage types can be efficiently learned and how they can lead to robust extraction of depth information. Our results indicate that stages without much variation and object clutter can be detected robustly, with up to 60% success rate. Vladimir Nedovic, Arnold W. M. Smeulders, André Redert, Jan-Mark Geusebroek |
ICCV | 2 |
| 2007 | The Role of Visual Content and Style for Concert Video IndexingabstractThis paper contributes to the automatic indexing of concert video. In contrast to traditional methods, which rely primarily on audio information for summarization applications, we explore how a visual-only concept detection approach could be employed. We investigate how our recent method for news video indexing -which takes into account the role of content and style -generalizes to the concert domain. We analyze concert video on three levels of visual abstraction, namely: content, style, and their fusion. Experiments with 12 concept detectors, on 45 hours of visually challenging concert video, show that the automatically learned best approach is concept-dependent. Moreover, these results suggest that the visual modality provides ample opportunity for more effective indexing and retrieval of concert video when used in addition to the auditory modality. Cees Snoek, Marcel Worring, Arnold W. M. Smeulders, Bauke Freiburg |
ICME | 3 |
| 2007 | The Distribution Family of Similarity DistancesabstractAssessing similarity between features is a key step in object recognition and scene categorization tasks. We argue that knowledge on the distribution of distances generated by similarity functions is crucial in deciding whether features are similar or not. Intuitively one would expect that similarities between features could arise from any distribution. In this paper, we will derive the contrary, and report the theoretical result that $L_p$-norms --a class of commonly applied distance metrics-- from one feature vector to other vectors are Weibull-distributed if the feature values are correlated and non-identically distributed. Besides these assumptions being realistic for images, we experimentally show them to hold for various popular feature extraction algorithms, for a diverse range of images. This fundamental insight opens new directions in the assessment of feature similarity, with projected improvements in object and scene recognition algorithms. Erratum: The authors of paper have declared that they have become convinced that the reasoning in the reference is too simple as a proof of their claims. As a consequence, they withdraw their theorems. Gertjan J. Burghouts, Arnold W. M. Smeulders, Jan-Mark Geusebroek |
NIPS | 2 |
| 2007 | Predicting Brain States from fMRI Data: Incremental Functional Principal Component RegressionabstractWe propose a method for reconstruction of human brain states directly from functional neuroimaging data. The method extends the traditional multivariate regression analysis of discretized fMRI data to the domain of stochastic functional measurements, facilitating evaluation of brain responses to naturalistic stimuli and boosting the power of functional imaging. The method searches for sets of voxel timecourses that optimize a multivariate functional linear model in terms of Rsquare-statistic. Population based incremental learning is used to search for spatially distributed voxel clusters, taking into account the variation in Haemodynamic lag across brain areas and among subjects by voxel-wise non-linear registration of stimuli to fMRI data. The method captures spatially distributed brain responses to naturalistic stimuli without attempting to localize function. Application of the method for prediction of naturalistic stimuli from new and unknown fMRI data shows that the approach is capable of identifying distributed clusters of brain locations that are highly predictive of a specific stimuli. Sennay Ghebreab, Arnold W. M. Smeulders, Pieter W. Adriaans |
NIPS | 2 |
| 2007 | Spatio-Temporal Context for Robust Multitarget TrackingabstractIn multitarget tracking, the main challenge is to maintain the correct identity of targets even under occlusions or when differences between the targets are small. The paper proposes a new approach to this problem by incorporating the context information. The context of a target in an image sequence has two components: the spatial context including the local background and nearby targets and the temporal context including all appearances of the targets that have been seen previously. The paper considers both aspects. We propose a new model for multitarget tracking based on the classification of each target against its spatial context. The tracker searches a region similar to the target while avoiding nearby targets. The temporal context is included by integrating the entire history of target appearance based on probabilistic principal component analysis (PPCA). We have developed a new incremental scheme that can learn the full set of PPCA parameters accurately online. The experiments show robust tracking performance under the condition of severe clutter, occlusions, and pose changes. Hieu Tat Nguyen, Arnold W. M. Smeulders |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2007 | Interactive Search by Direct Manipulation of Dissimilarity SpaceabstractIn this paper, we argue to learn dissimilarity for interactive search in content based image retrieval. In literature, dissimilarity is often learned via the feature space by feature selection, feature weighting or by adjusting the parameters of a function of the features. Other than existing techniques, we use feedback to adjust the dissimilarity space independent of feature space. This has the great advantage that it manipulates dissimilarity directly. To create a dissimilarity space, we use the method proposed by Pekalska and Duin, selecting a set of images called prototypes and computing distances to those prototypes for all images in the collection. After the user gives feedback, we apply active learning with a one-class support vector machine to decide the movement of images such that relevant images stay close together while irrelevant ones are pushed away (the work of Guo ). The dissimilarity space is then adjusted accordingly. Results on a Corel dataset of 10000 images and a TrecVid collection of 43907 keyframes show that our proposed approach is not only intuitive, it also significantly improves the retrieval performance. Giang P. Nguyen, Marcel Worring, Arnold W. M. Smeulders |
IEEE Trans. Multim. | 3 |
| 2007 | A Learned Lexicon-Driven Paradigm for Interactive Video RetrievalabstractEffective video retrieval is the result of interplay between interactive query selection, advanced visualization of results, and a goal-oriented human user. Traditional interactive video retrieval approaches emphasize paradigms, such as query-by-keyword and query-by-example, to aid the user in the search for relevant footage. However, recent results in automatic indexing indicate that query-by-concept is becoming a viable resource for interactive retrieval also. We propose in this paper a new video retrieval paradigm. The core of the paradigm is formed by first detecting a large lexicon of semantic concepts. From there, we combine query-by-concept, query-by-example, query-by-keyword, and user interaction into the MediaMill semantic video search engine. To measure the impact of increasing lexicon size on interactive video retrieval performance, we performed two experiments against the 2004 and 2005 NIST TRECVID benchmarks, using lexicons containing 32 and 101 concepts, respectively. The results suggest that from all factors that play a role in interactive retrieval, a large lexicon of semantic concepts matters most. Indeed, by exploiting large lexicons, many video search questions are solvable without using query-by-keyword and query-by-example. In addition, we show that the lexicon-driven search engine outperforms all state-of-the-art video retrieval systems in both TRECVID 2004 and 2005 Cees Snoek, Marcel Worring, Dennis C. Koelma, Arnold W. M. Smeulders |
IEEE Trans. Multim. | 4 |
| 2006 | Robust multi-target tracking using spatio-temporal contextabstractIn multi-target tracking, the maintaining of the correct identity of targets is challenging. In the presented tracking method, accurate target identification is achieved by incorporating the appearance information of the spatial and temporal context of each target. The spatial context of a target involves local background and nearby targets. The first contribution of the paper is to provide a new discriminative model for multi-target tracking with the embedded classification of each target against its context. As a result, the tracker not only searches for the image region similar to the target but also avoids latching on nearby targets or on a background region. The temporal context of a target includes its appearances seen during tracking in the past. The past appearances are used to train a probabilistic PCA that is used as the measurement model of the target at the present. As the second contribution, we develop a new incremental scheme for probabilistic PCA. It can update accurately the full set of parameters including a noise parameter still ignored in related literature. The experiments show robust tracking performance under the condition of severe clutter, occlusions and pose changes. Hieu Tat Nguyen, Arnold W. M. Smeulders |
CVPR (1) | 3 |
| 2006 | The Semantic Pathfinder for Generic News Video IndexingabstractThis paper presents the semantic pathfinder architecture for generic indexing of video archives. The pathfinder automatically extracts semantic concepts from video based on the exploration of different paths through three consecutive analysis steps, closely linked to the video production process, namely: content analysis, style analysis, and context analysis. The virtue of the semantic pathfinder is its learned ability to find a best path of analysis steps on a per-concept basis. To show the generality of this indexing approach we develop detectors for a lexicon of 32 concepts and we evaluate the semantic pathfinder against the 2004 NIST TRECVID video retrieval benchmark, using a news archive of 64 hours. Top ranking performance indicates the merit of the semantic pathfinder Cees Snoek, Marcel Worring, Jan-Mark Geusebroek, Dennis C. Koelma, Frank J. Seinstra, Arnold W. M. Smeulders |
ICME | 6 |
| 2006 | The influence of cross-validation on video classification performanceabstractDigital video is sequential in nature. When video data is used in a semantic concept classification task, the episodes are usually summarized with shots. The shots are annotated as containing, or not containing, a certain concept resulting in a labeled dataset. These labeled shots can subsequently be used by supervised learning methods (classifiers) where they are trained to predict the absence or presence of the concept in unseen shots and episodes. The performance of such automatic classification systems is usually estimated with cross-validation. By taking random samples from the dataset for training and testing as such, part of the shots from an episode are in the training set and another part from the same episode is in the test set. Accordingly, data dependence between training and test set is introduced, resulting in too optimistic performance estimates. In this paper, we experimentally show this bias, and propose how this bias can be prevented using episode-constrained crossvalidation. Moreover, we show that a 17% higher classifier performance can be achieved by using episode constrained cross-validation for classifier parameter tuning. Jan C. van Gemert, Cees Snoek, Cor J. Veenman, Arnold W. M. Smeulders |
ACM Multimedia | 4 |
| 2006 | The challenge problem for automated detection of 101 semantic concepts in multimediaabstractWe introduce the challenge problem for generic video indexing to gain insight in intermediate steps that affect performance of multimedia analysis methods, while at the same time fostering repeatability of experiments. To arrive at a challenge problem, we provide a general scheme for the systematic examination of automated concept detection methods, by decomposing the generic video indexing problem into 2 unimodal analysis experiments, 2 multimodal analysis experiments, and 1 combined analysis experiment. For each experiment, we evaluate generic video indexing performance on 85 hours of international broadcast news data, from the TRECVID 2005/2006 benchmark, using a lexicon of 101 semantic concepts. By establishing a minimum performance on each experiment, the challenge problem allows for component-based optimization of the generic indexing issue, while simultaneously offering other researchers a reference for comparison during indexing methodology development. To stimulate further investigations in intermediate analysis steps that inuence video indexing performance, the challenge offers to the research community a manually annotated concept lexicon, pre-computed low-level multimedia features, trained classifier models, and five experiments together with baseline performance, which are all available at http://www.mediamill.nl/challenge/. Cees Snoek, Marcel Worring, Jan C. van Gemert, Jan-Mark Geusebroek, Arnold W. M. Smeulders |
ACM Multimedia | 5 |
| 2006 | Robust Tracking Using Foreground-Background Texture Discrimination
Hieu Tat Nguyen, Arnold W. M. Smeulders |
Int. J. Comput. Vis. | 2 |
| 2006 | Sparse Representation for Coarse and Fine Object RecognitionabstractThis paper offers a sparse, multiscale representation of objects. It captures the object appearance by selection from a very large dictionary of Gaussian differential basis functions. The learning procedure results from the matching pursuit algorithm, while the recognition is based on polynomial approximation to the bases, turning image matching into a problem of polynomial evaluation. The method is suited for coarse recognition between objects and, by adding more bases, also for fine recognition of the object pose. The advantages over the common representation using PCA include storing sampled points for recognition is not required, adding new objects to an existing data set is trivial because retraining other object models is not needed, and significantly in the important case where one has to scan an image over multiple locations in search for an object, the new representation is readily available as opposed to PCA projection at each location. The experimental result on the COIL-100 data set demonstrates high recognition accuracy with real-time performance. Thang V. Pham, Arnold W. M. Smeulders |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2006 | The Semantic Pathfinder: Using an Authoring Metaphor for Generic Multimedia IndexingabstractThis paper presents the semantic pathfinder architecture for generic indexing of multimedia archives. The semantic pathfinder extracts semantic concepts from video by exploring different paths through three consecutive analysis steps, which we derive from the observation that produced video is the result of an authoring-driven process. We exploit this authoring metaphor for machine-driven understanding. The pathfinder starts with the content analysis step. In this analysis step, we follow a data-driven approach of indexing semantics. The style analysis step is the second analysis step. Here, we tackle the indexing problem by viewing a video from the perspective of production. Finally, in the context analysis step, we view semantics in context. The virtue of the semantic pathfinder is its ability to learn the best path of analysis steps on a per-concept basis. To show the generality of this novel indexing approach, we develop detectors for a lexicon of 32 concepts and we evaluate the semantic pathfinder against the 2004 NIST TRECVID video retrieval benchmark, using a news archive of 64 hours. Top ranking performance in the semantic concept detection task indicates the merit of the semantic pathfinder for generic indexing of multimedia archives. Cees Snoek, Marcel Worring, Jan-Mark Geusebroek, Dennis C. Koelma, Frank J. Seinstra, Arnold W. M. Smeulders |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2006 | Learning spatial relations in object recognition
Thang V. Pham, Arnold W. M. Smeulders |
Pattern Recognit. Lett. | 2 |
| 2006 | Robust photometric invariant features from the color tensorabstractLuminance-based features are widely used as low-level input for computer vision applications, even when color data is available. The extension of feature detection to the color domain prevents information loss due to isoluminance and allows us to exploit the photometric information. To fully exploit the extra information in the color data, the vector nature of color data has to be taken into account and a sound framework is needed to combine feature and photometric invariance theory. In this paper, we focus on the structure tensor, or color tensor, which adequately handles the vector nature of color images. Further, we combine the features based on the color tensor with photometric invariant derivatives to arrive at photometric invariant features. We circumvent the drawback of unstable photometric invariants by deriving an uncertainty measure to accompany the photometric invariant derivatives. The uncertainty is incorporated in the color tensor, hereby allowing the computation of robust photometric invariant features. The combination of the photometric invariance theory and tensor-based features allows for detection of a variety of features such as photometric invariant edges, corners, optical flow, and curvature. The proposed features are tested for noise characteristics and robustness to photometric changes. Experiments show that the proposed features are robust to scene incidental events and that the proposed uncertainty measure improves the applicability of full invariants. Joost van de Weijer 0001, Theo Gevers, Arnold W. M. Smeulders |
IEEE Trans. Image Process. | 3 |
| 2005 | Early versus late fusion in semantic video analysisabstractSemantic analysis of multimodal video aims to index segments of interest at a conceptual level. In reaching this goal, it requires an analysis of several information streams. At some point in the analysis these streams need to be fused. In this paper, we consider two classes of fusion schemes, namely early fusion and late fusion. The former fuses modalities in feature space, the latter fuses modalities in semantic space. We show by experiment on 184 hours of broadcast video data and for 20 semantic concepts, that late fusion tends to give slightly better performance for most concepts. However, for those concepts where early fusion performs better the difference is more significant. Cees Snoek, Marcel Worring, Arnold W. M. Smeulders |
ACM Multimedia | 3 |
| 2005 | Object recognition with uncertain geometry and uncertain part detection
Thang V. Pham, Arnold W. M. Smeulders |
Comput. Vis. Image Underst. | 2 |
| 2005 | The Amsterdam Library of Object Images
Jan-Mark Geusebroek, Gertjan J. Burghouts, Arnold W. M. Smeulders |
Int. J. Comput. Vis. | 3 |
| 2005 | A Six-Stimulus Theory for Stochastic Texture
Jan-Mark Geusebroek, Arnold W. M. Smeulders |
Int. J. Comput. Vis. | 2 |
| 2005 | The UvA color document dataset
Leon Todoran, Marcel Worring, Arnold W. M. Smeulders |
Int. J. Document Anal. Recognit. | 3 |
| 2005 | Color texture measurement and segmentation
Minh Anh Hoang, Jan-Mark Geusebroek, Arnold W. M. Smeulders |
Signal Process. | 3 |
| 2004 | An Information-Based Measure for Grouping Quality
Erik A. Engbers, Michael Lindenbaum, Arnold W. M. Smeulders |
ECCV (3) | 3 |
| 2004 | Tracking Aspects of the Foreground against the Background
Hieu Tat Nguyen, Arnold W. M. Smeulders |
ECCV (2) | 2 |
| 2004 | Active learning using pre-clusteringabstractThe paper is concerned with two-class active learning. While the common approach for collecting data in active learning is to select samples close to the classification boundary, better performance can be achieved by taking into account the prior data distribution. The main contribution of the paper is a formal framework that incorporates clustering into active learning. The algorithm first constructs a classifier on the set of the cluster representatives, and then propagates the classification decision to the other samples via a local noise model. The proposed model allows to select the most representative samples as well as to avoid repeatedly labeling samples in the same cluster. During the active learning process, the clustering is adjusted using the coarse-to-fine strategy in order to balance between the advantage of large clusters and the accuracy of the data representation. The results of experiments in image databases show a better performance of our algorithm compared to the current methods. Hieu Tat Nguyen, Arnold W. M. Smeulders |
ICML | 2 |
| 2004 | Guest Editorial
Arnold W. M. Smeulders, Thomas S. Huang, Theo Gevers |
Int. J. Comput. Vis. | 1 |
| 2004 | Thick 2D relations for document understanding
Marco Aiello 0001, Arnold W. M. Smeulders |
Inf. Sci. | 2 |
| 2004 | Fast Occluded Object Tracking by a Robust Appearance FilterabstractWe propose a new method for object tracking in image sequences using template matching. To update the template, appearance features are smoothed temporally by robust Kalman filters, one to each pixel. The resistance of the resulting template to partial occlusions enables the accurate detection and handling of more severe occlusions. Abrupt changes of lighting conditions can also be handled, especially when photometric invariant color features are used. The method has only a few parameters and is computationally fast enough to track objects in real time. Hieu Tat Nguyen, Arnold W. M. Smeulders |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2004 | Population-based incremental interactive concept learning for image retrieval by stochastic string segmentationsabstractWe propose a method for concept-based medical image retrieval that is a superset of existing semantic-based image retrieval methods. We conceive of a concept as an incremental and interactive formalization of the user's conception of an object in an image. The premise is that such a concept is closely related to a user's specific preferences and subjectivity and, thus, allows to deal with the complexity and content-dependency of medical image content. We describe an object in terms of multiple continuous boundary features and represent an object concept by the stochastic characteristics of an object population. A population-based incrementally learning technique, in combination with relevance feedback, is then used for concept customization. The user determines the speed and direction of concept customization using a single parameter that defines the degree of exploration and exploitation of the search space. Images are retrieved from a database in a limited number of steps based upon the customized concept. To demonstrate our method we have performed concept-based image retrieval on a database of 292 digitized X-ray images of cervical vertebrae with a variety of abnormalities. The results show that our method produces precise and accurate results when doing a direct search. In an open-ended search our method efficiently and effectively explores the search space. Sennay Ghebreab, Conrade C. Jaffe, Arnold W. M. Smeulders |
IEEE Trans. Medical Imaging | 3 |
| 2003 | Fragmentation in the Vision of ScenesabstractNatural images are highly structured in their spatial configuration. Where one would expect a different spatial distribution for every image, as each image has a different spatial layout, we show that the spatial statistics of recorded images can be explained by a single process of sequential fragmentation. The observation by a resolution limited sensory system turns out to have a profound influence on the observed statistics of natural images. The power-law and normal distribution represent the extreme cases of sequential fragmentation. Between these two extremes, spatial detail statistics deform from power-law to normal through the Weibull type distribution as receptive field size increases relative to image detail size. Jan-Mark Geusebroek, Arnold W. M. Smeulders |
ICCV | 2 |
| 2003 | Components and systems for interactive video indexingabstractThe process of video indexing determines the quality of video retrieval. We present a modularization of indexing systems in which dependencies of components are made explicit. We stress the impact of human interaction in the architectural scheme, as the semantic gap between automatic abstractions and semantic indices requires human intervention. We discuss the components for efficient indexing, interaction and visualization in detail. Jeroen Vendrig, Marcel Worring, Arnold W. M. Smeulders |
ICME | 3 |
| 2003 | Video retrieval and summarization
Nicu Sebe, Michael S. Lew, Arnold W. M. Smeulders |
Comput. Vis. Image Underst. | 3 |
| 2003 | Design Considerations for Generic Grouping in VisionabstractGrouping in vision can be seen as the process that organizes image entities into higher-level structures. Despite its importance, there is little consistency in the statement of the grouping problem in literature. In addition, most grouping algorithms in vision are inspired on a specific technique, rather than being based on desired characteristics, making it cumbersome to compare the behavior of various methods. We discuss six precisely formulated considerations for the design of generic grouping algorithms in vision: proper definition, invariance, multiple interpretations, multiple solutions, simplicity and robustness. We observe none of the existing algorithms for grouping in vision meet all the considerations. We present a simple algorithm as an extension of a classical algorithm, where the extension is based on taking the considerations into account. The algorithm is applied to three examples: grouping point sets, grouping poly-lines, and grouping flow-field vectors. The complexity of the greedy algorithm is O(nO/sub G/), where O/sub G/ is the complexity of the grouping measure. Erik A. Engbers, Arnold W. M. Smeulders |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2003 | Strings: Variational Deformable Models of Multivariate Continuous Boundary FeaturesabstractWe propose a new image segmentation technique called strings. A string is a variational deformable model that is learned from a collection of example objects rather than built from a priori analytical or geometrical knowledge. As opposed to existing approaches, an object boundary is represented by a one-dimensional multivariate curve in functional space, a feature function, rather than by a point in vector space. In the learning phase, feature functions are defined by extraction of multiple shape and image features along continuous object boundaries in a given learning set. The feature functions are aligned, then subjected to functional principal components analysis and functional principal regression to summarize the feature space and to model its content, respectively. Also, a Mahalanobis distance model is constructed for evaluation of boundaries in terms of their feature functions, taking into account the natural variations seen in the learning set. In the segmentation phase, an object boundary in a new image is searched for with help of a curve. The curve gives rise to a feature function, a string, that is weighted by the regression model and evaluated by the Mahalanobis model. The curve is deformed in an iterative procedure to produce feature functions with minimal Mahalanobis distance. Strings have been compared with active shape models on 145 vertebra images, showing that strings produce better results when initialized close to the target boundary, and comparable results otherwise. Sennay Ghebreab, Arnold W. M. Smeulders |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2003 | Color constancy from physical principles
Jan-Mark Geusebroek, Rein van den Boomgaard, Arnold W. M. Smeulders, Theo Gevers |
Pattern Recognit. Lett. | 3 |
| 2003 | Fast anisotropic Gauss filteringabstractWe derive the decomposition of the anisotropic Gaussian in a one-dimensional (1-D) Gauss filter in the x-direction followed by a 1-D filter in a nonorthogonal direction phi. So also the anisotropic Gaussian can be decomposed by dimension. This appears to be extremely efficient from a computing perspective. An implementation scheme for normal convolution and for recursive filtering is proposed. Also directed derivative filters are demonstrated. For the recursive implementation, filtering an 512 x 512 image is performed within 40 msec on a current state of the art PC, gaining over 3 times in performance for a typical filter, independent of the standard deviations and orientation of the filter. Accuracy of the filters is still reasonable when compared to truncation error or recursive approximation error. The anisotropic Gaussian filtering method allows fast calculation of edge and ridge maps, with high spatial and angular accuracy. For tracking applications, the normal anisotropic convolution scheme is more advantageous, with applications in the detection of dashed lines in engineering drawings. The recursive implementation is more attractive in feature detection applications, for instance in affine invariant edge and ridge detection in computer vision. The proposed computational filtering method enables the practical applicability of orientation scale-space analysis. Jan-Mark Geusebroek, Arnold W. M. Smeulders, Joost van de Weijer 0001 |
IEEE Trans. Image Process. | 2 |
| 2002 | Data GroundTruth, Complexity, and Evaluation Measures for Color Document AnalysisabstractPublications on color document image analysis present results on small, non-publicly available datasets.We propose in this paper a well defined and groundtruthed color dataset existing of over 1000 pages, with associated tools for evaluation. The color data groundtruthing and evaluation tools are based on a well defined document model, complexity measures to assess the inherent dificulty of analyzing a page, and well founded evaluation measures. Together they form a suitable basis for evaluating diverse applications in color document analysis. These keywords were added by machine and not by the authors. This process is experimental and the keywords may be updated as the learning algorithm improves. Leon Todoran, Marcel Worring, Arnold W. M. Smeulders |
Document Analysis Systems | 3 |
| 2002 | Fast Anisotropic Gauss Filtering
Jan-Mark Geusebroek, Arnold W. M. Smeulders, Joost van de Weijer 0001 |
ECCV (1) | 2 |
| 2002 | Template tracking using color invariant pixel featuresabstractIn our method for tracking objects, appearance features are smoothed by robust and adaptive Kalman filters, one to each pixel, making the method robust against occlusions. While existing methods use only intensity to model the object appearance, we concentrate on multivalue features. Specifically, one option is to use photometric invariant color features, making the method robust to illumination effects such as shadow and object geometry. The method is able to track objects in real time. Hieu Tat Nguyen, Arnold W. M. Smeulders |
ICIP (1) | 2 |
| 2002 | Interactive Indexing and Retrieval of Multimedia Content
Marcel Worring, Andrew D. Bagdanov, Jan C. van Gemert, Jan-Mark Geusebroek, Hoang Minh, Guus Schreiber, Cees Snoek, Jeroen Vendrig, Jan Wielemaker, Arnold W. M. Smeulders |
SOFSEM | 10 |
| 2002 | Necklaces: Inhomogeneous and Point-Enhanced Deformable Models
Sennay Ghebreab, Arnold W. M. Smeulders, Pia R. Pfluger |
Comput. Vis. Image Underst. | 2 |
| 2002 | Face detection by aggregated Bayesian network classifiers
Thang V. Pham, Marcel Worring, Arnold W. M. Smeulders |
Pattern Recognit. Lett. | 3 |
| 2002 | Tracking nonparameterized object contours in videoabstractWe propose a new method for contour tracking in video. The inverted distance transform of the edge map is used as an edge indicator function for contour detection. Using the concept of topographical distance, the watershed segmentation can be formulated as a minimization. This new viewpoint gives a way to combine the results of the watershed algorithm on different surfaces. In particular, our algorithm determines the contour as a combination of the current edge map and the contour, predicted from the tracking result in the previous frame. We also show that the problem of background clutter can be relaxed by taking the object motion into account. The compensation with object motion allows to detect and remove spurious edges in background. The experimental results confirm the expected advantages of the proposed method over the existing approaches. Hieu Tat Nguyen, Marcel Worring, Rein van den Boomgaard, Arnold W. M. Smeulders |
IEEE Trans. Image Process. | 4 |
| 2001 | Color Constant Ratio Gradients for Image Segmentation and Similarity of Texture ObjectsabstractWe aim for content-based image retrieval of texture objects in natural scenes under varying illumination and viewing conditions. To achieve this, image retrieval is based on matching feature distributions derived from color invariant gradients. To cope with object cluttering, region-based texture segmentation is applied on the target images prior to the actual image retrieval process. The retrieval scheme is empirically verified on color images taken from texture objects under different lighting, conditions. Theo Gevers, Arnold W. M. Smeulders |
CVPR (1) | 2 |
| 2001 | Invariant representation in image processingabstractThe paper discusses the role of invariance in image processing, specifically the desire to discriminate against unwanted variations in the scene while maintaining the power to tell the difference between object-intrinsic characteristics and scene-accidental conditions. It provides an analysis and references of what are directly observables in a general scene. Arnold W. M. Smeulders, Jan-Mark Geusebroek, Theo Gevers |
ICIP (3) | 1 |
| 2001 | Model Based Interactive Story Unit SegmentationabstractLogical Story Unit segmentation in general domains, such as movies and television series, requires interaction between experts and automatic tools. We present an interaction model that allows users, who are non-experts in the field of video processing, to segment videos into Logical Story Units by tuning segmentation model parameters rather than manually adjust results of automatic methods. Suitable features are determined interactively by visualizing their values in the context of the original video data. Jeroen Vendrig, Marcel Worring, Arnold W. M. Smeulders |
ICME | 3 |
| 2001 | A Minimum Cost Approach for Segmenting Networks of Lines
Jan-Mark Geusebroek, Arnold W. M. Smeulders, Hugo Geerts |
Int. J. Comput. Vis. | 2 |
| 2001 | Scale Dependency of Image Derivatives for Feature Measurement in Curvilinear Structures
Geert J. Streekstra, Rein van den Boomgaard, Arnold W. M. Smeulders |
Int. J. Comput. Vis. | 3 |
| 2001 | Interaction in the segmentation of medical images: A survey
Sílvia Delgado Olabarriaga, Arnold W. M. Smeulders |
Medical Image Anal. | 2 |
| 2001 | Filter Image Browsing: Interactive Image Retrieval by Using Database Overviews
Jeroen Vendrig, Marcel Worring, Arnold W. M. Smeulders |
Multim. Tools Appl. | 3 |
| 2001 | Color InvarianceabstractThis paper presents the measurement of colored object reflectance, under different, general assumptions regarding the imaging conditions. We exploit the Gaussian scale-space paradigm for color images to define a framework for the robust measurement of object reflectance from color images. Object reflectance is derived from a physical reflectance model based on the Kubelka-Munk theory for colorant layers. Illumination and geometrical invariant properties are derived from the reflectance model. Invariance and discriminative power of the color invariants is experimentally investigated, showing the invariants to be successful in discounting shadow, illumination, highlights, and noise. Extensive experiments show the different invariants to be highly discriminative, while maintaining invariance properties. The presented framework for color measurement is well-founded in the physics of color as well as in measurement science. Hence, the proposed invariants are considered more adequate for the measurement of invariant color features than existing methods. Jan-Mark Geusebroek, Rein van den Boomgaard, Arnold W. M. Smeulders, Hugo Geerts |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2000 | Image Retrieval and Segmentation based on Color InvariantsabstractWe will demonstrate our CVPR2000 paper "Measurement of Color Invariants" for the cases of image retrieval based on query by example and for color image segmentation. Both are of importance in content based access of image and video data. We demonstrate the usefulness of the proposed color invariants in image retrieval by example systems. We show that an image retrieval query should include the type of invariance expected in the result. We demonstrate such queries by using the "ImageSurf" retrieval system. Segmentation of images based on the proposed color invariants is demonstrated by the "PicToVision" system. The system provides image processing functionality through the world wide web, and is publicly accessible at www.science. uva.nl/research/isis.pictovision.html. Jan-Mark Geusebroek, Dennis C. Koelma, Arnold W. M. Smeulders, Theo Gevers |
CVPR | 3 |
| 2000 | Measurement of Color InvariantsabstractThis paper presents the measurement of object reflectance from color images. We exploit the Gaussian scale-space paradigm to define framework for the robust measurement of object reflectance from color images. Illumination and geometrical invariant properties are derived from a physical reflectance model based on the Kubelka-Munk theory. Imaging conditions are assumed to be white illumination and matte, dull object or general object, respectively. Invariance is denoted by +, whereas sensitivity to the imaging condition is indicated by -. Invariance, discriminative power and localization accuracy of the color invariants is extensively investigated, showing the invariants to be successful in discounting shadow, illumination intensity, highlights, and noise. Experiments show the different invariants to be highly discriminative while maintaining invariance properties. The presented framework for color measurement is well-founded in physics as well as measurement science. The framework is thoroughly evaluated experimentally. Hence is considered more adequate than existing methods for the measurement of invariant color features. Jan-Mark Geusebroek, Arnold W. M. Smeulders, Rein van den Boomgaard |
CVPR | 2 |
| 2000 | Color and Scale: The Spatial Structure of Color Images
Jan-Mark Geusebroek, Rein van den Boomgaard, Arnold W. M. Smeulders, Anuj Dev |
ECCV (1) | 3 |
| 2000 | Scale Dependent Differential Geometry for the Measurement of Center Line and Diameter in 3D Curvilinear Structures
Geert J. Streekstra, Rein van den Boomgaard, Arnold W. M. Smeulders |
ECCV (1) | 3 |
| 2000 | Content-Based Image Retrieval at the End of the Early YearsabstractPresents a review of 200 references in content-based image retrieval. The paper starts with discussing the working conditions of content-based retrieval: patterns of use, types of pictures, the role of semantics, and the sensory gap. Subsequent sections discuss computational steps for image retrieval systems. Step one of the review is image processing for retrieval sorted by color, texture, and local geometry. Features for retrieval are discussed next, sorted by: accumulative and global features, salient points, object and shape features, signs, and structural combinations thereof. Similarity of pictures and objects in pictures is reviewed for each of the feature types, in close connection to the types and means of feedback the user of the systems is capable of giving by interaction. We briefly discuss aspects of system engineering: databases, system architecture, and evaluation. In the concluding section, we present our view on: the driving force of the field, the heritage from computer vision, the influence on computer vision, the role of similarity and of interaction, the need for databases, the problem of evaluation, and the role of the semantic gap. Arnold W. M. Smeulders, Marcel Worring, Simone Santini, Amarnath Gupta, Ramesh Jain 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2000 | PicToSeek: combining color and shape invariant features for image retrievalabstractWe aim at combining color and shape invariants for indexing and retrieving images. To this end, color models are proposed independent of the object geometry, object pose, and illumination. From these color models, color invariant edges are derived from which shape invariant features are computed. Computational methods are described to combine the color and shape invariants into a unified high-dimensional invariant feature set for discriminatory object retrieval. Experiments have been conducted on a database consisting of 500 images taken from multicolored man-made objects in real world scenes. From the theoretical and experimental results it is concluded that object retrieval based on composite color and shape invariant features provides excellent retrieval accuracy. Object retrieval based on color invariants provides very high retrieval accuracy whereas object retrieval based entirely on shape invariants yields poor discriminative power. Furthermore, the image retrieval scheme is highly robust to partial occlusion, object clutter and a change in the object's pose. Finally, the image retrieval scheme is integrated into the PicToSeek system on-line at http://www.wins.uva.nl/research/isis/PicToSeek/ for searching images on the World Wide Web. Theo Gevers, Arnold W. M. Smeulders |
IEEE Trans. Image Process. | 2 |
| 2000 | Spectral Volume RenderingabstractVolume renderers for interactive analysis must be sufficiently versatile to render a broad range of volume images: unsegmented "raw" images as recorded by a 3D scanner, labeled segmented images, multimodality images, or any combination of these. The usual strategy is to assign to each voxel a three component RGB color and an opacity value /spl alpha/. This so-called RGB/spl alpha/ approach offers the possibility of distinguishing volume objects by color. However, these colors are connected to the objects themselves, thereby bypassing the idea that in reality the color of an object is also determined by the light source and light detectors c.q. human eyes. The physically realistic approach presented, models light interacting with the materials inside a voxel causing spectral changes in the light. The radiated spectrum falls upon a set of RGB detectors. The spectral approach is investigated to see whether it could enhance the visualization of volume data and interactive tools. For that purpose, a material is split into an absorbing part (the medium) and a scattering part (small particles). The medium is considered to be either achromatic or chromatic, while the particles are considered to scatter the light achromatically, elastically, or inelastically. Inelastic scattering particles combined with an achromatic absorbing medium offer additional visual features: objects are made visible through the surface structure of a surrounding volume object and volume and surface structures can be made visible at the same time. With one or two materials the method is faster than the RGB/spl alpha/ approach, with three materials the performance is equal. The spectral approach can be considered as an extension of the RGB/spl alpha/ approach with greater visual flexibility and a better balance between quality and speed. Herke Jan Noordmans, Hans T. M. van der Voort, Arnold W. M. Smeulders |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 1999 | Grammatical Inference of Dashed Lines
Arnold Jonk, Rein van den Boomgaard, Arnold W. M. Smeulders |
Comput. Vis. Image Underst. | 3 |
| 1999 | Content based internet access to paper documents
Marcel Worring, Arnold W. M. Smeulders |
Int. J. Document Anal. Recognit. | 2 |
| 1999 | Content-based image retrieval by viewpoint-invariant color indexing
Theo Gevers, Arnold W. M. Smeulders |
Image Vis. Comput. | 2 |
| 1999 | Color-based object recognition
Theo Gevers, Arnold W. M. Smeulders |
Pattern Recognit. | 2 |
| 1998 | Color Invariant SnakesabstractSnakes provide high-level information in the form of continuity constraints and minimum energy constraints related to the contour shape and image features. These image features are usually based on intensity edges. However, intensity edges may appear in the scene without a material/color transition to support it. As a consequence, when using intensity edges as image features, the image segmentation results obtained by snakes may be negatively affected by the imaging-process (e.g. shadows, shading and highlights). In this paper, we aim at using color invariant gradient information to guide the deformation process to obtain snake boundaries which correspond to material boundaries in images discounting the disturbing influences of surface orientation, illumination, shadows and highlights. Experiments conducted on various color images show that the proposed color invariant snake successfully find material contours discounting other "accidental" edges types (e.g. shadows, shading and highlight transitions). Comparison with intensity-based snakes shows that the intensity-based snake is dramatically outperformed by the presented color invariant snake. 1 Theo Gevers, Sennay Ghebreab, Arnold W. M. Smeulders |
BMVC | 3 |
| 1998 | Photometric Invariant Region DetectionabstractIn this paper, we concentrate on determining homogeneously colored regions invariant to surface orientation change, illumination, shadows and highlights. To this end, the influence of various well-known color models (e.g. , , , , , , , and ) are examined, in theory, for the dichromatic reflection model and, in practice, for two distinct region-based segmentation methods: the k-means clustering technique and the split&merge algorithm. Experiments are conducted on color images taken from colored objects in real-world scenes. On the basis of the theoretical and experimental results it is concluded that , , , , and all detect regions invariant to a change in surface orientation, viewpoint of the camera, and illumination intensity. Furthermore, and also detect regions independent of highlights. , , , , ,a nd provide segmentation results which are all sensitive to surface orientation and illumination intensity as well as color models incorporating brightness into their systems: in , in ,a nd in . Theo Gevers, Arnold W. M. Smeulders, Harro M. G. Stokman |
BMVC | 2 |
| 1998 | Image Indexing using Composite Color and Shape Invariant FeaturesabstractNew sets of color models are proposed for object recognition invariant to a change in view point, object geometry and illumination. Further, computational methods are presented to combine color and shape invariants to produce a high-dimensional invariant feature set for discriminatory object recognition. Experiments on a database of 500 images show that object recognition based on composite color and shape invariant features provides excellent recognition accuracy. Furthermore, object recognition based on color invariants provides very high recognition accuracy whereas object recognition based entirely on shape invariants yields very poor discriminative power. The image database and the performance of the recognition scheme can be experienced within PicToSeek: on-line as part of the ZOMAX system at: http://www.wins.uva.nl/research/isis/zomax/. Theo Gevers, Arnold W. M. Smeulders |
ICCV | 2 |
| 1998 | Detection and Characterization of Isolated and Overlapping Spots
Herke Jan Noordmans, Arnold W. M. Smeulders |
Comput. Vis. Image Underst. | 2 |
| 1998 | High accuracy tracking of 2D/3D curved line-structures by consecutive cross-section matching
Herke Jan Noordmans, Arnold W. M. Smeulders |
Pattern Recognit. Lett. | 2 |
| 1997 | Combining Region Splitting and Edge Detection through Guided Delaunay Image SubdivisionabstractIn this paper, an adaptive split-and-merge segmentation method is proposed. The splitting phase of the algorithm employs the incremental Delaunay triangulation competent of forming grid edges of arbitrary orientation, and position. The tessellation grid, defined by the Delaunay triangulation, is adjusted to the semantics of the image data by combining similarity and difference information among pixels. Experimental results on synthetic images show that the method is robust to different object edge orientations, partially weak object edges and very noisy homogeneous regions. Experiments on a real image indicate that the method yields good segmentation results even when there is a quadratic sloping of intensities particularly suited for segmenting natural scenes of man-made objects. Theo Gevers, Arnold W. M. Smeulders |
CVPR | 2 |
| 1997 | From Linear to Non-Linear Reading: A Case Study to Provide Internet Access to Paper DocumentsabstractThe authors consider the construction of hypertext from scanned paper material. They consider textual links as well as text-figure and text-figure label links. This process of hypertext creation encompasses a number of document structures and methods for going from one document representation to the other. These are described and exemplified by a case study of turning a manual into an electronic hypertext book. They further discuss the Netscape based interface to the system. The hypermanual considered in the paper is accessible through their Web page. Marcel Worring, Arnold W. M. Smeulders |
ICDAR | 2 |
| 1997 | Fast volume render techniques for interactive analysis
Herke Jan Noordmans, Arnold W. M. Smeulders, Hans T. M. van der Voort |
Vis. Comput. | 2 |
| 1996 | BESSI: an experimentation system for vision module evaluationabstractIn past years, the complexity of computer vision software systems has grown considerably. As in other software development areas, there is a need for accurate descriptions of system behavior in practical situations. One element is software documentation, another more important one is the performance for various data inputs. These descriptions are essential for maintainability and successful application of software in the long term. We present an experimentation system for the evaluation of computer vision modules. The system is capable of collecting data, executing experiments and analysing the generated data. In the paper the architecture of BESSI is discussed and the system is illustrated by an example. George A. Den Boer, Arnold W. M. Smeulders |
ICPR | 2 |
| 1996 | Color-metric pattern-card matching for viewpoint invariant image retrievalabstractIn this paper, viewpoint independent image retrieval by color-metric pattern-card matching is presented. First, a photometric color invariant is proposed measuring, a local color property of a pixel and its neighboring pixels while discounting the disturbing influences of shading, shadows and highlights. Color-metric pattern-cards are constructed on the basis of the photometric color invariant indicating whether a particular discrete photometric color invariant value is present in an image. To express similarity between color-metric pattern-cards, similarity functions are proposed and evaluated on a database of 500 images taken from 2-D and 3-D colored man-made objects in real world 3-D scenes. The experimental results show that high image retrieval accuracy is achieved by two distinct similarity functions depending on the presence of object clutter in the scene. Furthermore, image retrieval by color-metric pattern-card matching is to a large degree robust to partial occlusion and a change in viewing position. Good run-time performance of the pattern-card matching process is achieved allowing for fast image retrieval by example image. Theo Gevers, Arnold W. M. Smeulders |
ICPR | 2 |
| 1996 | Parameterized Feasible Boundaries in Gradient Vector FieldsabstractSegmentation of (noisy) images containing a complex ensemble of objects is difficult to achieve on the basis of local image information only. It is advantageous to attack the problem of object boundary extraction by a model-based segmentation procedure. Segmentation is achieved by tuning the parameters of the geometrical model in such a way that the boundary template locates and describes the object in the image in an optimal way. The optimality of the solution is based on an objective function taking into account image information as well as the shape of the template. Objective functions in literature are mainly based on the gradient magnitude and a measure describing the smoothness of the template. In this contribution, we propose a new image objective function based on directional gradient information derived from Gaussian smoothed derivatives of the image data. The proposed method is designed to accurately locate an object boundary even in the case of a conflicting object positioned close to the object of interest. We further introduce a new smoothness objective to ensure the physical feasibility of the contour. The method is evaluated on artificial data. Results on real medical images show that the method is very effective in accurately locating object boundaries in very complex images. Marcel Worring, Arnold W. M. Smeulders, Lawrence H. Staib, James S. Duncan |
Comput. Vis. Image Underst. | 2 |
| 1995 | An axiomatic approach to clustering line-segmentsabstractIn this paper we consider the problem of clustering line-segments into new ones. The clustering-hierarchy gives an answer to the question what original line segments are combined into larger ones. Such a clustering is defined as a hierarchical ordering of a set of line-segments. Criteria on a clustering-method are presented. The difference between edges and lines in relation to scale-invariant clustering is demonstrated. Existing approaches are evaluated using the presented criteria. It is shown that these approaches do not meet desirable criteria such as scale-invariance. A new method is described that adheres the formulated criteria. Finally an experiment is presented that illustrates the usefulness of the new method. Arnold Jonk, Arnold W. M. Smeulders |
ICDAR | 2 |
| 1995 | Evaluation of an interactive tool for handwritten form descriptionabstractA highly time-consuming activity in many areas of commerce and business is the manual entry into computer of data handwritten on forms. All forms in widespread use contain discrete fields where specific information can be entered. Automatic recognition of these forms could be achieved using existing state-of-the-art OCR algorithms for numerals, alphabetic characters, cursive words, signatures and mark sensing if they could be rapidly configured along with any inter-relationships and dependencies for different forms. This paper describes an initial implementation of an interactive graphical tool to allow the handwritten fields of a form and their inter-relationships to be described and defined for automatic linking with appropriate OCR algorithms. Results indicate that the main requirement is for the operator to have a full understanding of the handwritten form and an ability to describe its contents. Marcel Worring, Rein van den Boomgaard, Arnold W. M. Smeulders |
ICDAR | 3 |
| 1995 | Digitized Circular Arcs: Characterization and Parameter EstimationabstractThe digitization of a circular arc causes an inherent loss of geometrical information. Arcs with slightly different local curvature or position may lead to exactly the same digital pattern. In this paper the authors give a characterization of all centers and radii of circular arcs yielding the same digitization pattern. The radius of the arcs varies over the set. However, only one curvature or radius estimate can be assigned to the digital pattern. The authors derive an optimal estimator and give expressions for the bound on the precision of estimation. This bound due to digitization is the deterministic equivalent of the Cramer/Rao bound known from parameter estimation theory. Consider the estimation of the local curvature and local radius of a smooth object. Typically such parameters are estimated by moving a window along the digital boundary. Methods in literature show a poor precision in estimating curvature values, relative errors of over 40% are often found. From the definition of curvature it follows that locally the curve can be considered a circular arc and hence the method presented in this paper can be applied to the pattern in the window giving estimates with optimal precision and a measure for the remaining error. On the practical side the authors present examples of the residual error due to the discrete grid. The estimation of the radius or curvature of a circular arc at random position with an estimation window containing 10 points (coded with nine Freemancodes) has a relative deviation exceeding 2%. For a full disk the deviation is below 1% when the radius r exceeds four grid units. The presented method is particularly useful for problems where some prior knowledge on the distribution of radii is known and where there is a noise-free sampling.> Marcel Worring, Arnold W. M. Smeulders |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 1994 | ScilImage: an environment for collaborative use and development of image processing softwareabstractThis paper presents a brief overview of a portable multi-layered interactive environment for software development and use, and a class based infrastucture for image processing. The environment consists of a library handler, a C interpreter, a command expander, a menu and dialog generator, and a visual programming interface. It forms a proper vehicle for development and exchange of (image processing) software across multiple research institutions. Richard van Balen, Arnold W. M. Smeulders |
ICPR (1) | 2 |
| 1994 | Location estimation of cylinders from a 2-D imageabstractThis paper describes an approach to estimate the location from a cylinder considering both diffuse and specular models of reflection. Given the proper orientation of the cylinder, the Marquardt-Levenberg method from computational mathematics is used to estimate the parameters in the models based on several intensity profiles perpendicular to the cylinders orientation. While the nonphysical axis of the cylinder is used to determine its position, radius lighting direction and other reflectance parameters are also estimated. Experimental results for both synthesised and real images show that the method is robust with respect to noise. Onno Wink, Arnold W. M. Smeulders, Dennis C. Koelma |
ICPR (1) | 2 |
| 1994 | Discrete circular arcsabstractIn this contribution we consider in all detail the effect of digitization on circular arcs. Given a specific discrete circular arc we find the set of all continuous arcs which by digitization would result in this pattern. From this characterization we provide optimal estimates of the radius (or curvature) of the original arc. This estimator achieves the ultimate precision one can reach in estimation which we call the geometric minimum variance bound (GMVB). Marcel Worring, Arnold W. M. Smeulders |
ICPR (1) | 2 |
| 1994 | The Morphological Structure of Images: The Differential Equations of Morphological Scale-SpaceabstractWe introduce a class of nonlinear differential equations that are solved using morphological operations. The erosion and dilation act as morphological propagators propagating the initial condition into the "scale-space", much like the Gaussian convolution is the propagator for the linear diffusion equation. The analysis starts in the set domain, resulting in the description of erosions and dilations in terms of contour propagation. We show that the structuring elements to be used must have the property that at each point of the contour there is a well-defined and unique normal vector. Then given the normal at a point of the dilated contour we can find the corresponding point (point-of-contact) on the original contour. In some situations we can even link the normal of the dilated contour with the normal in the point-of-contact of the original contour. The results of the set domain are then generalized to grey value images. The role of the normal is replaced with the function gradient. The same analysis also holds for the erosion. Using a family of increasingly larger structuring functions we are then able to link infinitesimal changes in grey value with the gradient in the image.> Rein van den Boomgaard, Arnold W. M. Smeulders |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 1994 | A visual programming interface for an image processing environment
Dennis C. Koelma, Arnold W. M. Smeulders |
Pattern Recognit. Lett. | 2 |
| 1994 | Measurement of 3D-line shaped objects
Marcel Worring, Pia R. Pfluger, Arnold W. M. Smeulders, Adriaan B. Houtsmuller |
Pattern Recognit. Lett. | 3 |
| 1993 | An Approach to Image Retrieval for Image Databases
Theo Gevers, Arnold W. M. Smeulders |
DEXA | 2 |
| 1992 | The morphological structure of imagesabstractThe authors investigate the use of mathematical morphology to construct scale-spaces. These scale-spaces are based on differential equations, which are solved by morphological operators, describing the evolution of images in scale-space.> Rein van den Boomgaard, Arnold W. M. Smeulders |
ICPR (3) | 2 |
| 1992 | Σnigma: an image retrieval systemabstractPresents a system which retrieves images on the basis of automatically generated indexes (i.e. semantic image representations, obtained by automatic image analysis, indicating the content of the images). The system consists of two parts: an off-line indexing part and an on-line image retrieval part. The indexing component is used to automatically generate semantic representations of images so that the image retrieval component can use this information to enable image retrieval. The man-machine communication of the image retrieval component is based on an iconical graphical query language to accomplish geographical query specification for image access. Experiments have been carried out on three different sets of images from the following domains: MRI images of the chest, electronic schemas and topographic maps. The experiments show encouraging results especially for domains which have a high degree of formality in their pictorial expression, such as electronic schemas and topographic maps, and to a less extent for domains having a weak degree of formality such as MRI images of the chest.> Theo Gevers, Arnold W. M. Smeulders |
ICPR (2) | 2 |
| 1992 | Quantitative 3-D texture analysis of interphase cell nucleiabstractIn order to investigate the spatio-temporal structure of chromatin in interphase nuclei the authors present two 3-D texture parameters based on the grey-weighted distance transform that quantify the accessibility and the homogeneity of a nucleus. Results of experiments on computer generated textures show that these texture parameters are shape independent.> Karel C. Strasters, Arnold W. M. Smeulders, Hans T. M. van der Voort, Ian T. Young, Nanne Nanninga |
ICPR (3) | 2 |
| 1992 | Multi-scale analysis of discrete point setsabstractPresents the shape of a sparse point set S in R/sup 2/. A crucial step in finding the shape of a sparse point set is the definition of its boundary. This boundary is a graph indicating a relation among the elements of S. No well defined definition of such a boundary is found in literature. For continuous point sets this problem does not exist as the boundary has a unique definition. The authors pose general criteria a boundary definition should satisfy and show that the alpha -graph satisfies those criteria. The boundary is a function of the scale parameter alpha . The authors further show that the alpha -graph has a strong relation with mathematical morphology. As an application the use of the alpha -graph in the multi-scale recognition of industrial objects is shown.> Marcel Worring, Arnold W. M. Smeulders |
ICPR (1) | 2 |
| 1992 | The accuracy and precision of curvature estimation methodsabstractDeals with the estimation of curvature from digital image data, especially the selection of a curvature estimation procedure based on its accuracy and precision. The authors establish that almost all curvature estimation techniques from literature suffer from a severe directional inaccuracy and/or poor precision (errors depend on the method, orientation and scale ranging from 1% to more than 200%). A practical solution to the curvature estimation problem is presented.> Marcel Worring, Arnold W. M. Smeulders |
ICPR (3) | 2 |
| 1992 | Optimization of length measurements for isotropic distance transformations in three dimension
A. L. D. Beckers, Arnold W. M. Smeulders |
CVGIP Image Underst. | 2 |
| 1990 | The probability of a random straight line in two and three dimensions
A. L. D. Beckers, Arnold W. M. Smeulders |
Pattern Recognit. Lett. | 2 |
| 1990 | SCILAIM: A multi-level interactive image processing environment
Ton K. ten Kate, Richard van Balen, Arnold W. M. Smeulders, Frans C. A. Groen, George A. Den Boer |
Pattern Recognit. Lett. | 3 |
| 1989 | A comment on "a note on 'distance transformations in digital images'"
A. L. D. Beckers, Arnold W. M. Smeulders |
Comput. Vis. Graph. Image Process. | 2 |
| 1989 | Human chromosome classification based on local band descriptors
Frans C. A. Groen, Ton K. ten Kate, Arnold W. M. Smeulders, Ian T. Young |
Pattern Recognit. Lett. | 3 |
| 1988 | "Image segmentation and uncertainty" by R. Wilson and M. Spann
Arnold W. M. Smeulders |
Pattern Recognit. Lett. | 1 |
| 1987 | Length estimators for digitized contours
Leo Dorst, Arnold W. M. Smeulders |
Comput. Vis. Graph. Image Process. | 2 |
| 1986 | Best Linear Unbiased Estimators for Properties of Digitized Straight LinesabstractThis paper considers the problem of measuring properties of digitized straight lines from the viewpoint of measurement methodology. The measurement and estimation process is described in detail, revealing the importance of a step called ``characterization'' which was not recognized explicitly before. Using this new concept, BLUE (Best Linear Unbiased) estimators are found. These are calculated for various properties of digitized straight lines, and are briefly compared to previous work. Leo Dorst, Arnold W. M. Smeulders |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 1986 | Correction to "Best Linear Unbiased Estimators for Properties of Digitized Straight Lines"
Leo Dorst, Arnold W. M. Smeulders |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 1984 | Discrete Representation of Straight LinesabstractIf a continuous straight line segment is digitized on a regular grid, obviously a loss of information occurs. As a result, the discrete representation obtained (e.g., a chaincode string) can be coded more conveniently than the continuous line segment, but measurements of properties (such as line length) performed on the representation have an intrinsic inaccuracy due to the digitization process. In this paper, two fundamental properties of the quantization of straight line segments are treated. 1) It is proved that every ``straight'' chaincode string can be represented by a set of four unique integer parameters. Definitions of these parameters are given. 2) A mathematical expression is derived for the set of all continuous line segments which could have generated a given chaincode string. The relation with the chord property is briefly discussed. Leo Dorst, Arnold W. M. Smeulders |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 1982 | Vector code probability and metrication error in the representation of straight lines of finite length
Albert M. Vossepoel, Arnold W. M. Smeulders |
Comput. Graph. Image Process. | 2 |
| 1982 | Vector code probability and metrication error in the representation of straight lines of finite length
Albert M. Vossepoel, Arnold W. M. Smeulders |
Comput. Graph. Image Process. | 2 |