Theo Gevers

dblp:12/6600 · DBLP profile ↗
← Back
175ranked-venue papers
26as first author
39since 2021 · last 2026
0000-0002-1190-5492ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 126 · 19 first-author · 23 since 2021Artificial intelligence and machine learning · 111 · 18 first-author · 30 since 2021Applied, interdisciplinary, general and emerging computing · 2Databases, data management, data science and information retrieval · 1 · 1 first-author
YearPublicationVenuePosition
2026 Edge-Centric Relational Reasoning for 3D Scene Graph Prediction
abstract
3D scene graph prediction aims to abstract complex 3D environments into structured graphs consisting of objects and their pairwise relationships. Existing approaches typically adopt object-centric graph neural networks, where relation edge features are iteratively updated by aggregating messages from connected object nodes. However, this design inherently restricts relation representations to pairwise object context, making it difficult to capture high-order relational dependencies that are essential for accurate relation prediction. To address this limitation, we propose a Link-guided Edge-centric relational reasoning framework with Object-aware fusion, namely LEO, which enables progressive reasoning from relation-level context to object-level understanding. Specifically, LEO first predicts potential links between object pairs to suppress irrelevant edges, and then transforms the original scene graph into a line graph where each relation is treated as a node. A line graph neural network is applied to perform edge-centric relational reasoning to capture inter-relation context. The enriched relation features are subsequently integrated into the original object-centric graph to enhance object-level reasoning and improve relation prediction. Our framework is model-agnostic and can be integrated with any existing object-centric method. Experiments on the 3DSSG dataset with two competitive baselines show consistent improvements, highlighting the effectiveness of our edge-to-object reasoning paradigm.
Yanni Ma, Hao Liu 0061, Yulan Guo, Theo Gevers, Martin R. Oswald
AAAI4
2025 A Large-Scale Dataset of Gaussian Splats and Their Self-Supervised Pretraining
abstract
3D Gaussian Splatting (3DGS) has become the de facto method of 3D representation in many vision tasks. This calls for the 3D understanding directly in this representation space. To facilitate the research in this direction, we first build a large-scale dataset of 3DGS using the commonly used ShapeNet and ModelNet datasets. Our dataset ShapeSplat consists of 65K objects from 87 unique categories, whose labels are in accordance with the respective datasets. The creation of this dataset utilized the computing equivalent of 2 GPU years on a TITAN XP GPU. We utilize our dataset for unsupervised pretraining and supervised finetuning for classification and segmentation tasks. To this end, we introduce Gaussian-MAE, which highlights the unique benefits of representation learning from Gaussian parameters. Through exhaustive experiments, we provide several valuable insights. In particular, we show that (1) the distribution of the optimized GS centroids significantly differs from the uniformly sampled point cloud (used for initialization) counterpart; (2) this change in distribution results in degradation in classification but improvement in segmentation tasks when using only the centroids; (3) to leverage additional Gaussian parameters, we propose Gaussian feature grouping in a normalized feature space, along with splats pooling layer, offering a tailored solution to effectively group and embed similar Gaussians, which leads to notable improvement in finetuning tasks. Our dataset and model are publicly available at ShapeSplat.
Yue Li 0036, Bin Ren 0005, Nicu Sebe, Ender Konukoglu, Theo Gevers, Luc Van Gool, Danda Pani Paudel
3DV6
2025 3D-AVS: LiDAR-based 3D Auto-Vocabulary Segmentation
abstract
Open-vocabulary segmentation methods offer promising capabilities in detecting unseen object categories, but the category must be aware and needs to be provided by a human, either via a text prompt or pre-labeled datasets, thus limiting their scalability. We propose 3D-AVS, a method for Auto-Vocabulary Segmentation of 3D point clouds for which the vocabulary is unknown and auto-generated for each input at runtime, thus eliminating the human in the loop and typically providing a substantially larger vocabulary for richer annotations. 3D-AVS first recognizes semantic entities from image or point cloud data and then segments all points with the automatically generated vocabulary. Our method incorporates both image-based and point-based recognition, enhancing robustness under challenging lighting conditions where geometric information from Li-DAR is especially valuable. Our point-based recognition features a Sparse Masked Attention Pooling (SMAP) module to enrich the diversity of recognized objects. To address the challenges of evaluating unknown vocabularies and avoid annotation biases from label synonyms, hierarchies, or semantic overlaps, we introduce the annotation-free Text-Point Semantic Similarity (TPSS) metric for assessing generated vocabulary quality. Our evaluations on nuScenes and ScanNet200 demonstrate 3D-AVS’s ability to generate semantic classes with accurate point-wise segmentations.
Weijie Wei 0001, Osman Ülger, Fatemeh Karimi Nejadasl, Theo Gevers, Martin R. Oswald
CVPR4
2025 LumiNet: Latent Intrinsics Meets Diffusion Models for Indoor Scene Relighting
abstract
We introduce LumiNet, a novel architecture that leverages generative models and latent intrinsic representations for transferring lighting from one image to another. Given a source image and a target lighting image, LumiNet generates a relit version of the source scene that captures the target’s lighting. Our approach makes two key contributions: a data curation strategy from the StyleGAN-based relighting model for our training, and a modified diffusion-based Con-trolNet that processes both latent intrinsic properties from the source image and latent extrinsic properties from the target image. We further improve lighting transfer through a learned adaptor that injects the target’s latent extrinsic properties via cross-attention and light-weight fine-tuning.Unlike traditional ControlNet, which generates images with conditional maps from a single scene, LumiNet processes latent representations from two different images -preserving geometry and albedo from the source while transferring lighting characteristics from the target. Experiments demonstrate that our method successfully transfers complex lighting phenomena including specular highlights and indirect illumination across scenes with varying spatial layouts and materials, outperforming existing approaches on challenging indoor scenes using only images as input.
Xiaoyan Xing, Konrad Groh, Sezer Karaoglu, Theo Gevers, Anand Bhattad
CVPR4
2025 MAGiC-SLAM: Multi-Agent Gaussian Globally Consistent SLAM
abstract
Simultaneous localization and mapping (SLAM) systems with novel view synthesis capabilities are widely used in computer vision, with applications in augmented reality, robotics, and autonomous driving. However, existing approaches are limited to single-agent operation. Recent work has addressed this problem using a distributed neural scene representation. Unfortunately, existing methods are slow, cannot accurately render real-world data, are restricted to two agents, and have limited tracking accuracy. In contrast, we propose a rigidly deformable 3D Gaussian-based scene representation that dramatically speeds up the system. However, improving tracking accuracy and reconstructing a globally consistent map from multiple agents remains challenging due to trajectory drift and discrepancies across agents’ observations. Therefore, we propose new tracking and map-merging mechanisms and integrate loop closure in the Gaussian-based SLAM pipeline. We evaluate MAGiC-SLAM on synthetic and real-world datasets and find it more accurate and faster than the state of the art.
Vladimir Yugay, Theo Gevers, Martin R. Oswald
CVPR2
2025 SDFit: 3D Object Pose and Shape by Fitting a Morphable SDF to a Single Image
abstract
Recovering 3D object pose and shape from a single image is a challenging and ill-posed problem. This is due to strong (self-)occlusions, depth ambiguities, the vast intra- and inter-class shape variance, and the lack of 3D ground truth for natural images. Existing deep-network methods are trained on synthetic datasets to predict 3D shapes, so they often struggle generalizing to real-world images. Moreover, they lack an explicit feedback loop for refining noisy estimates, and primarily focus on geometry without directly considering pixel alignment. To tackle these limitations, we develop a novel render-and-compare optimization framework, called SDFit. This has three key innovations: First, it uses a learned category-specific and morphable signed-distance-function (mSDF) model, and fits this to an image by iteratively refining both 3D pose and shape. The mSDF robustifies inference by constraining the search on the manifold of valid shapes, while allowing for arbitrary shape topologies. Second, SDFit retrieves an initial 3D shape that likely matches the image, by exploiting foundational models for efficient look-up into 3D shape databases. Third, SDFit initializes pose by establishing rich 2D-3D correspondences between the image and the mSDF through foundational features. We evaluate SDFit on three image datasets, i.e., Pix3D, Pascal3D+, and COMIC. SDFit performs on par with SotA feed-forward networks for unoccluded images and common poses, but is uniquely robust to occlusions and uncommon poses. Moreover, it requires no retraining for unseen images. Thus, SDFit contributes new insights for generalizing in the wild. Code is available at https://anticdimi.github.io/sdfit.
Dimitrije Antic, Georgios Paschalidis, Shashank Tripathi, Theo Gevers, Saikumar Dwivedi, Dimitrios Tzionas
ICCV4
2025 SceneSplat: Gaussian Splatting-Based Scene Understanding with Vision-Language Pretraining
abstract
Recognizing arbitrary or previously unseen categories is essential for comprehensive real-world 3D scene understanding. Currently, all existing methods rely on 2D or textual modalities during training or together at inference. This highlights the clear absence of a model capable of processing 3D data alone for learning semantics end-to-end, along with the necessary data to train such a model. Meanwhile, 3D Gaussian Splatting (3DGS) has emerged as the de facto standard for 3D scene representation across various vision tasks. However, effectively integrating semantic reasoning into 3DGS in a generalizable manner remains an open challenge. To address these limitations, we introduce SceneSplat, to our knowledge the first large-scale 3D indoor scene understanding approach that operates natively on 3DGS. Furthermore, we propose a self-supervised learning scheme that unlocks rich 3D feature learning from unlabeled scenes. To power the proposed methods, we introduce SceneSplat-7K, the first large-scale 3DGS dataset for indoor scenes, comprising 7916 scenes derived from seven established datasets, such as ScanNet and Matterport3D. Generating SceneSplat-7K required computational resources equivalent to 150 GPU days on an L4 GPU, enabling standardized benchmarking for 3DGS-based reasoning for indoor scenes. Our exhaustive experiments on SceneSplat-7K demonstrate the significant benefit of the proposed method over the established baselines.
Yue Li 0036, Runyi Yang, Huapeng Li, Mengjiao Ma, Bin Ren 0005, Nikola Popovic 0001, Nicu Sebe, Ender Konukoglu, Theo Gevers, Luc Van Gool, Martin R. Oswald, Danda Pani Paudel
ICCV10
2025 SceneSplat++: A Large Dataset and Comprehensive Benchmark for Language Gaussian Splatting
abstract
3D Gaussian Splatting (3DGS) serves as a highly performant and efficient encoding of scene geometry, appearance, and semantics. Moreover, grounding language in 3D scenes has proven to be an effective strategy for 3D scene understanding. Current Language Gaussian Splatting line of work fall into three main groups: (i) per-scene optimization-based, (ii) per-scene optimization-free, and (iii) generalizable approach. However, most of them are evaluated only on rendered 2D views of a handful of scenes and viewpoints close to the training views, limiting ability and insight into holistic 3D understanding. To address this gap, we propose the first large-scale benchmark that systematically assesses these three groups of methods directly in 3D space, evaluating on 1060 scenes across three indoor datasets and one outdoor dataset. Benchmark results demonstrate a clear advantage of the generalizable paradigm, particularly in relaxing the scene-specific limitation, enabling fast feed-forward inference on novel scenes, and achieving superior segmentation performance. We further introduce SceneSplat-49K -- a carefully curated 3DGS dataset comprising of around 49K diverse indoor and outdoor scenes trained from multiple sources, with which we demonstrate generalizable approach could harness strong data priors. Our codes, benchmark, and datasets are available.
Mengjiao Ma, Yue Li 0036, Jiahuan Cheng, Runyi Yang, Bin Ren 0005, Nikola Popovic 0001, Mingqiang Wei, Nicu Sebe, Ender Konukoglu, Luc Van Gool, Theo Gevers, Martin R. Oswald, Danda Pani Paudel
NeurIPS12
2025 Training-free diffusion for controlling illumination conditions in images
Xiaoyan Xing, Vincent Tao Hu, Jan Hendrik Metzen, Konrad Groh, Sezer Karaoglu, Theo Gevers
Comput. Vis. Image Underst.6
2025 Exploring dynamic plane representations for neural scene reconstruction
Ruihong Yin, Yunlu Chen, Sezer Karaoglu, Theo Gevers
Pattern Recognit.4
2025 3D human pose estimation and action recognition using fisheye cameras: A survey and benchmark
abstract
3D human pose estimation based on visual information aims to predict 3D poses of humans in images or videos. The aim of human action recognition is to classify what kind of actions people do. Both topics are widely studied in the field of computer vision. Existing methods mainly focus on 3D human pose estimation and human action recognition using images/videos recorded by perspective cameras. In contrast to perspective cameras, fisheye cameras use wide-angle lenses capturing wider field-of-views (FOV). Fisheye cameras are used in many applications such as surveillance and autonomous driving. In this paper, a survey is given on monocular 3D human pose estimation and action recognition. A new benchmark dataset is proposed using a fisheye camera to quantitatively compare and analyze existing methods.
Shaodi You, Sezer Karaoglu, Theo Gevers
Pattern Recognit.4
2024 Learning Generalized Segmentation for Foggy-Scenes by Bi-directional Wavelet Guidance
abstract
Learning scene semantics that can be well generalized to foggy conditions is important for safety-crucial applications such as autonomous driving. Existing methods need both annotated clear images and foggy images to train a curriculum domain adaptation model. Unfortunately, these methods can only generalize to the target foggy domain that has seen in the training stage, but the foggy domains vary a lot in both urban-scene styles and fog styles. In this paper, we propose to learn scene segmentation well generalized to foggy-scenes under the domain generalization setting, which does not involve any foggy images in the training stage and can generalize to any arbitrary unseen foggy scenes. We argue that an ideal segmentation model that can be well generalized to foggy-scenes need to simultaneously enhance the content, de-correlate the urban-scene style and de-correlate the fog style. As the content (e.g., scene semantic) rests more in low-frequency features while the style of urban-scene and fog rests more in high-frequency features, we propose a novel bi-directional wavelet guidance (BWG) mechanism to realize the above three objectives in a divide-and-conquer manner. With the aid of Haar wavelet transformation, the low frequency component is concentrated on the content enhancement self-attention, while the high frequency component is shifted to the style and fog self-attention for de-correlation purpose. It is integrated into existing mask-level Transformer segmentation pipelines in a learnable fashion. Large-scale experiments are conducted on four foggy-scene segmentation datasets under a variety of interesting settings. The proposed method significantly outperforms existing directly-supervised, curriculum domain adaptation and domain generalization segmentation methods. Source code is available at https://github.com/BiQiWHU/BWG.
Qi Bi, Shaodi You, Theo Gevers
AAAI3
2024 Learning Content-Enhanced Mask Transformer for Domain Generalized Urban-Scene Segmentation
abstract
Domain-generalized urban-scene semantic segmentation (USSS) aims to learn generalized semantic predictions across diverse urban-scene styles. Unlike generic domain gap challenges, USSS is unique in that the semantic categories are often similar in different urban scenes, while the styles can vary significantly due to changes in urban landscapes, weather conditions, lighting, and other factors. Existing approaches typically rely on convolutional neural networks (CNNs) to learn the content of urban scenes. In this paper, we propose a Content-enhanced Mask TransFormer (CMFormer) for domain-generalized USSS. The main idea is to enhance the focus of the fundamental component, the mask attention mechanism, in Transformer segmentation models on content information. We have observed through empirical analysis that a mask representation effectively captures pixel segments, albeit with reduced robustness to style variations. Conversely, its lower-resolution counterpart exhibits greater ability to accommodate style variations, while being less proficient in representing pixel segments. To harness the synergistic attributes of these two approaches, we introduce a novel content-enhanced mask attention mechanism. It learns mask queries from both the image feature and its down-sampled counterpart, aiming to simultaneously encapsulate the content and address stylistic variations. These features are fused into a Transformer decoder and integrated into a multi-resolution content-enhanced mask attention learning scheme. Extensive experiments conducted on various domain-generalized urban-scene segmentation datasets demonstrate that the proposed CMFormer significantly outperforms existing CNN-based methods by up to 14.0% mIoU and the contemporary HGFormer by up to 1.7% mIoU. The source code is publicly available at https://github.com/BiQiWHU/CMFormer.
Qi Bi, Shaodi You, Theo Gevers
AAAI3
2024 SceneTeller: Language-to-3D Scene Generation
Basak Melis Öcal, Maxim Tatarchenko, Sezer Karaoglu, Theo Gevers
ECCV (85)4
2024 T-MAE : Temporal Masked Autoencoders for Point Cloud Representation Learning
Weijie Wei 0001, Fatemeh Karimi Nejadasl, Theo Gevers, Martin R. Oswald
ECCV (11)3
2024 Ray-Distance Volume Rendering for Neural Scene Reconstruction
Ruihong Yin, Yunlu Chen, Sezer Karaoglu, Theo Gevers
ECCV (14)4
2024 FewViewGS: Gaussian Splatting with Few View Matching and Multi-stage Training
abstract
The field of novel view synthesis from images has seen rapid advancements with the introduction of Neural Radiance Fields (NeRF) and more recently with 3D Gaussian Splatting. Gaussian Splatting became widely adopted due to its efficiency and ability to render novel views accurately. While Gaussian Splatting performs well when a sufficient amount of training images are available, its unstructured explicit representation tends to overfit in scenarios with sparse input images, resulting in poor rendering performance. To address this, we present a 3D Gaussian-based novel view synthesis method using sparse input images that can accurately render the scene from the viewpoints not covered by the training images. We propose a multi-stage training scheme with matching-based consistency constraints imposed on the novel views without relying on pre-trained depth estimation or diffusion models. This is achieved by using the matches of the available training images to supervise the generation of the novel views sampled between the training frames with color, geometry, and semantic losses. In addition, we introduce a locality preserving regularization for 3D Gaussians which removes rendering artifacts by preserving the local color structure of the scene. Evaluation on synthetic and real-world datasets demonstrates competitive or superior performance of our method in few-shot novel view synthesis compared to existing state-of-the-art methods.
Ruihong Yin, Vladimir Yugay, Yue Li 0036, Sezer Karaoglu, Theo Gevers
NeurIPS5
2024 Image semantic segmentation of indoor scenes: A survey
abstract
This survey provides a comprehensive evaluation of various deep learning-based segmentation architectures. It covers a wide range of models, from traditional ones like FCN and PSPNet to more modern approaches like SegFormer and FAN. In addition to assessing the methods in terms of segmentation accuracy, we propose to also evaluate the methods in terms of temporal consistency and corruption vulnerability. Most of the existing surveys on semantic segmentation focus on outdoor datasets. In contrast, this survey focuses on indoor scenarios to enhance the applicability of segmentation methods in this specific domain. Furthermore, our evaluation consists of a performance analysis of the methods in prevalent real-world segmentation scenarios that pose particular challenges. These complex situations involve scenes impacted by diverse forms of noise, blur corruptions, camera movements, optical aberrations, among other factors. By jointly exploring the segmentation accuracy, temporal consistency, and corruption vulnerability in challenging real-world situations, our survey offers insights that go beyond existing surveys, facilitating the understanding and development of better image segmentation methods for indoor scenes.
Ronny Velastegui, Maxim Tatarchenko, Sezer Karaoglu, Theo Gevers
Comput. Vis. Image Underst.4
2024 Kinship similarity for open sets
Wei Wang 0469, Shaodi You, Sezer Karaoglu, Theo Gevers
Pattern Recognit.4
2023 Geometry-guided Feature Learning and Fusion for Indoor Scene Reconstruction
abstract
In addition to color and textural information, geometry provides important cues for 3D scene reconstruction. However, current reconstruction methods only include geometry at the feature level thus not fully exploiting the geometric information.In contrast, this paper proposes a novel geometry integration mechanism for 3D scene reconstruction. Our approach incorporates 3D geometry at three levels, i.e. feature learning, feature fusion, and network supervision. First, geometry-guided feature learning encodes geometric priors to contain view-dependent information. Second, a geometry-guided adaptive feature fusion is introduced which utilizes the geometric priors as a guidance to adaptively generate weights for multiple views. Third, at the supervision level, taking the consistency between 2D and 3D normals into account, a consistent 3D normal loss is designed to add local constraints.Large-scale experiments are conducted on the ScanNet dataset, showing that volumetric methods with our geometry integration mechanism outperform state-of-the-art methods quantitatively as well as qualitatively. Volumetric methods with ours also show good generalization on the 7-Scenes and TUM RGB-D datasets.
Ruihong Yin, Sezer Karaoglu, Theo Gevers
ICCV3
2023 Learning rotation equivalent scene representation from instance-level semantics: A novel top-down perspective
abstract
This paper focuses on rotation variant scene recognition. Different from existing rotation invariant recognition approaches which learn from either rotated images or rotated convolutional filters in a bottom-up manner, a new top-down perspective by learning is explored from instance-level semantic representation. The goal is to eliminate the convolutional feature differences in bottom-up feature propagation caused by the rotation sensitive nature of convolution operation. Our rotation equivalent convolutional neural network (RE-CNN) scheme consists of three components. Firstly, our key instance selection module highlights the instances strongly related to the scene scheme regardless of their orientation. Secondly, our key instance aggregation module builds a scene representation invariant to the position change of each instance caused by rotation. Finally, our semantic fusion module allows the framework to be organized as a whole and implements rotation regularization. Notably, our RE-CNN scheme can be adapted to existing CNNs in a plug-in-and-play manner. Extensive experiments on rotation variant scene recognition benchmarks from four domains demonstrate the state-of-the-art performance and generalization capability of the proposed RE-CNN.
Qi Bi, Shaodi You, Wei Ji 0011, Theo Gevers
Comput. Vis. Image Underst.4
2023 A survey on kinship verification
abstract
In this survey, kinship verification is defined as the automatic process of verifying whether two or more persons are blood relatives (kin) by analyzing images of their faces. Kinship verification is an important research field in computer vision with many applications such as finding missing persons, family album organization, and online image search. Although substantial progress has been made in kinship verification in the past decade, there are still challenges such as intrinsic (face i.e., differences in facial appearance) and extrinsic (acquisition i.e., varying imaging conditions) problems. And there is still a demand for more diverse datasets. Therefore, this paper provides a survey on kinship verification methods and datasets. The survey starts with the definition of kinship verification and its corresponding intrinsic and extrinsic challenges. Then, an overview of kinship verification methods and datasets is given. Finally, a new multi-modal dataset (Nemo-Kinship Dataset) is proposed as a benchmark dataset addressing large inter-subject age variations consisting of 4216 videos of 248 persons from 85 families. The newly collected dataset is used to systematically test and analyze state-of-the-art methods.
Wei Wang 0469, Shaodi You, Sezer Karaoglu, Theo Gevers
Neurocomputing4
2023 Geometric Back-Propagation in Morphological Neural Networks
abstract
This paper provides a definition of back-propagation through geometric correspondences for morphological neural networks. In addition, dilation layers are shown to learn probe geometry by erosion of layer inputs and outputs. A proof-of-principle is provided, in which predictions and convergence of morphological networks significantly outperform convolutional networks.
Rick Groenendijk, Leo Dorst, Theo Gevers
IEEE Trans. Pattern Anal. Mach. Intell.3
2023 Interactive Learning of Intrinsic and Extrinsic Properties for All-Day Semantic Segmentation
abstract
Scene appearance changes drastically throughout the day. Existing semantic segmentation methods mainly focus on well-lit daytime scenarios and are not well designed to cope with such great appearance changes. Naively using domain adaption does not solve this problem because it usually learns a fixed mapping between the source and target domain and thus have limited generalization capability on all-day scenarios (i. e., from dawn to night). In this paper, in contrast to existing methods, we tackle this challenge from the perspective of image formulation itself, where the image appearance is determined by both intrinsic (e. g., semantic category, structure) and extrinsic (e. g., lighting) properties. To this end, we propose a novel intrinsic-extrinsic interactive learning strategy. The key idea is to interact between intrinsic and extrinsic representations during the learning process under spatial-wise guidance. In this way, the intrinsic representation becomes more stable and, at the same time, the extrinsic representation gets better at depicting the changes. Consequently, the refined image representation is more robust to generate pixel-wise predictions for all-day scenarios. To achieve this, we propose an All-in-One Segmentation Network (AO-SegNet) in an end-to-end manner. Large scale experiments are conducted on three real datasets (Mapillary, BDD100K and ACDC) and our proposed synthetic All-day CityScapes dataset. The proposed AO-SegNet shows a significant performance gain against the state-of-the-art under a variety of CNN and ViT backbones on all the datasets.
Qi Bi, Shaodi You, Theo Gevers
IEEE Trans. Image Process.3
2022 Distortion-aware Depth Estimation with Gradient Priors from Panoramas of Indoor Scenes
abstract
Compared to 2D perspective images, panoramic images capture a larger field-of-view (FOV). Depth estimation from panoramas is an important task for 3D scene understanding and has made significant progress with the development of CNNs. However, existing CNN-based methods still suffer from the Equirectangular Projection (ERP) problem to deal with panoramic distortions (e.g. same receptive fields near the equator and the two poles) and have difficulty generating accurate depth boundaries. In contrast to existing CNN-based methods, in this paper, a novel Transformer-based method is proposed which is able to cope with panoramic distortions and to generate accurate depth boundaries. A Distortion-aware Transformer is designed using a yaw-invariant cycle shift and a distortion-guided partitioning. The aim is to alleviate the distortion effect by enlarging the receptive fields in both horizontal and vertical directions. Then, a Gradient Transformer is proposed to enhance the features around the boundaries. Gradient information is adopted as a boundary prior. Large-scale experimental results show an improvement compared to state-of-the-art methods. Our method also shows strong generalization capabilities. Finally, our method is extended to panorama semantic segmentation.
Ruihong Yin, Sezer Karaoglu, Theo Gevers
3DV3
2022 Pose Guided Human Motion Transfer by Exploiting 2D and 3D Information
abstract
Human motion transfer aims to animate the pose of a human in a source image driven by the poses of a human in a target video. To warp (transfer) human poses, most of the existing methods are based on optical flow or affine transformations as an intermediate representation followed by a generator module to perform the motion transfer. Existing methods perform well in terms of reconstruction quality. However, the quality of the human pose transfer has received less attention although it is an important part of the motion transfer process. Therefore, in this paper, we propose a method focusing on both the reconstruction quality as well as pose consistency. In contrast to existing methods, performing warping procedures in 2D- or 3D-space, we introduce a strategy to combine the warped features in both 2D- and 3D-space to alleviate the self-occlusion problem. In this way, our method benefits from 2D (robustness) and 3D (steering) information to guide the generation process. To reduce the pose error caused by inaccurate 3D estimation, a method is proposed to maintain semantic consistency between predictions and target images at arm and leg regions. Experiments conducted on large scale datasets show that the proposed method outperforms existing methods. Ablation studies clarify the benefits of using feature fusion and semantic consistency.
Shaodi You, Sezer Karaoglu, Theo Gevers
3DV4
2022 MorphPool: Efficient Non-linear Pooling & Unpooling in CNNs
Rick Groenendijk, Leo Dorst, Theo Gevers
BMVC3
2022 PIE-Net: Photometric Invariant Edge Guided Network for Intrinsic Image Decomposition
abstract
Intrinsic image decomposition is the process of recovering the image formation components (reflectance and shading) from an image. Previous methods employ either explicit priors to constrain the problem or implicit constraints as formulated by their losses (deep learning). These methods can be negatively influenced by strong illumination conditions causing shading-reflectance leakages. Therefore, in this paper, an end-to-end edge-driven hybrid CNN approach is proposed for intrinsic image decomposition. Edges correspond to illumination invariant gradients. To handle hard negative illumination transitions, a hierarchical approach is taken including global and local refinement layers. We make use of attention layers to further strengthen the learning process. An extensive ablation study and large scale experiments are conducted showing that it is beneficial for edge-driven hybrid IID networks to make use of illumination invariant descriptors and that separating global and local cues helps in improving the performance of the network. Finally, it is shown that the proposed method obtains state of the art performance and is able to generalise well to real world images. The project page with pretrained models, finetuned models and network code can be found at https://ivi.fnwi.uva.nl/cv/pienet/.
Partha Das, Sezer Karaoglu, Theo Gevers
CVPR3
2022 Intrinsic image decomposition using physics-based cues and CNNs
abstract
Intrinsic image decomposition is the decomposition of an image into its reflectance and shading components. The intrinsic image decomposition problem is inherently ill-posed, since there can be multiple solutions to compute the intrinsic components forming the same image. In this paper, we explore the use of physics-based priors. We also propose a new architecture that separates the learning components in a stacked manner. We explore various ways of integrating such priors into a deep learning system. Our method is trained and tested on a large synthetic garden dataset to assess its performance. It is evaluated and compared to state-of-the-art methods using two standard intrinsic datasets. Finally, the pre-trained network is tested on real world images to show the generalisation capabilities of the network.
Partha Das, Sezer Karaoglu, Theo Gevers
Comput. Vis. Image Underst.3
2022 Multi-person 3D pose estimation from a single image captured by a fisheye camera
abstract
Multi-person 3D pose estimation with absolute depths for a fisheye camera is a challenging task but with valuable applications in daily life, especially for video surveillance. However, to the best of our knowledge, such problem has not been explored so far, leaving a gap in practical applications. In this work, we first propose a method for multi-person 3D pose estimation from a single image taken by a fisheye camera. Our method consists of two branches to estimate absolute 3D human poses: (1) a 2D-to-3D lifting module to predict root-relative 3D human poses (HPoseNet); (2) a root regression module to estimate absolute root locations in the camera coordinate (HRootNet). Finally, we propose a fisheye re-projection module without using ground-truth camera parameters to connect two branches, alleviating the impact of image distortions on 3D pose estimation and further regularizing prediction absolute 3D poses. Experimental results demonstrate that our method achieves the state-of-the-art performance on two public multi-person 3D pose datasets with synthetic fisheye images and our newly collected dataset with real fisheye images. The code and new dataset will be made publicly available.
Shaodi You, Sezer Karaoglu, Theo Gevers
Comput. Vis. Image Underst.4
2022 Self-Supervised Face Image Manipulation by Conditioning GAN on Face Decomposition
abstract
We present a novel architecture for manipulating facial expressions, head poses, and lighting conditions from a single monocular image. Recent methods based on Generative Adversarial Networks show promising results in expression manipulation. However, the variation is either defined by a limited number of classes or not well suitable for explicit manipulation of different attributes such as pose and lighting conditions. Besides, state-of-the-art methods are mostly focused on frontal faces. Therefore, in this paper, a new Generative Adversarial Network architecture is proposed by explicitly conditioning on the appearance image space which is the product of direct manipulation of facial expressions, light and pose conditions of the face model in 3D space. In addition, the method only requires video sequences for training. Therefore, it is self-supervised. Unlike other face manipulation methods, the proposed method does not require target specific training. Large scale experiments show that our method outperforms state-of-the-art methods for different scenarios.
Minh Ngô, Sezer Karaoglu, Theo Gevers
IEEE Trans. Multim.3
2021 Multi-Loss Weighting with Coefficient of Variations
abstract
Many interesting tasks in machine learning and computer vision are learned by optimising an objective function defined as a weighted linear combination of multiple losses. The final performance is sensitive to choosing the correct (relative) weights for these losses. Finding a good set of weights is often done by adopting them into the set of hyper- parameters, which are set using an extensive grid search. This is computationally expensive. In this paper, we propose a weighting scheme based on the coefficient of variations and set the weights based on properties observed while training the model1. The proposed method incorporates a measure of uncertainty to balance the losses, and as a result the loss weights evolve during training without requiring another (learning based) optimisation. In contrast to many loss weighting methods in literature, we focus on single-task multi-loss problems, such as monocular depth estimation and semantic segmentation, and show that multi-task approaches for loss weighting do not work on those single-tasks. The validity of the approach is shown empirically for depth estimation and semantic segmentation on multiple datasets.
Rick Groenendijk, Sezer Karaoglu, Theo Gevers, Thomas Mensink
WACV3
2021 EDEN: Multimodal Synthetic Dataset of Enclosed GarDEN Scenes
abstract
Multimodal large-scale datasets for outdoor scenes are mostly designed for urban driving problems. The scenes are highly structured and semantically different from scenarios seen in nature-centered scenes such as gardens or parks. To promote machine learning methods for nature-oriented applications, such as agriculture and gardening, we propose the multimodal synthetic dataset for Enclosed garDEN scenes (EDEN). The dataset features more than 300K images captured from more than 100 garden models. Each image is annotated with various low/high-level vision modalities, including semantic segmentation, depth, surface normals, intrinsic colors, and optical flow. Experimental results on the state-of-the-art methods for semantic segmentation and monocular depth prediction, two important tasks in computer vision, show positive impact of pre-training deep networks on our dataset for unstructured natural scenes. The dataset and related materials will be available at https://lhoangan.github.io/eden.
Hoang-An Le, Thomas Mensink, Partha Das, Sezer Karaoglu, Theo Gevers
WACV5
2021 Identity Unbiased Deception Detection by 2D-to-3D Face Reconstruction
abstract
Deception is a common phenomenon in society, both in our private and professional lives. However, humans are notoriously bad at accurate deception detection. Based on the literature, human accuracy of distinguishing between lies and truthful statements is 54% on average, in other words, it is slightly better than a random guess. While people do not much care about this issue, in high-stakes situations such as interrogations for series crimes and for evaluating the testimonies in court cases, accurate deception detection methods are highly desirable. To achieve a reliable, covert, and non-invasive deception detection, we propose a novel method that disentangles facial expression and head pose related features using 2D-to-3D face reconstruction technique from a video sequence and uses them to learn characteristics of deceptive behavior. We evaluate the proposed method on the Real-Life Trial (RLT) dataset that contains high-stakes deceits recorded in courtrooms. Our results show that the proposed method (with an accuracy of 68%) improves the state of the art. Besides, a new dataset has been collected, for the first time, for low-stake deceit detection. In addition, we compare high-stake deceit detection methods on the newly collected low-stake deceits.
Minh Ngô, Wei Wang 0469, Burak Mandira, Sezer Karaoglu, Henri Bouma, Hamdi Dibeklioglu, Theo Gevers
WACV7
2021 Automatic Calibration of the Fisheye Camera for Egocentric 3D Human Pose Estimation from a Single Image
abstract
We propose a method for egocentric 3D human pose estimation from a single image captured by a fisheye camera. The problem of estimating the egocentric 3D pose for a fisheye camera is that images may be subject to strong image distortions (e.g. 2D poses on the image plane that pass through the line of sight of the fisheye lens).Therefore, in this paper, we approach this problem by an automatic calibration module. Given a single image, our network first estimates 3D joint locations of a human in camera coordinates. To alleviate the impact of image distortions on 3D human pose estimation, we then use the automatic calibration to further regularize the 3D predictions. Experimental results demonstrate that the proposed method achieves state-of-the-art performance.
Shaodi You, Theo Gevers
WACV3
2021 Physics-based shading reconstruction for intrinsic image decomposition
abstract
We investigate the use of photometric invariance and deep learning to compute intrinsic images (albedo and shading). We propose albedo and shading gradient descriptors which are derived from physics-based models. Using the descriptors, albedo transitions are masked out and an initial sparse shading map is calculated directly from the corresponding RGB image gradients in a learning-free unsupervised manner. Then, an optimization method is proposed to reconstruct the full dense shading map. Finally, we integrate the generated shading map into a novel deep learning framework to refine it and also to predict corresponding albedo image to achieve intrinsic image decomposition. By doing so, we are the first to directly address the texture and intensity ambiguity problems of the shading estimations. Large scale experiments show that our approach steered by physics-based invariant descriptors achieve superior results on MIT Intrinsics, NIR-RGB Intrinsics, Multi-Illuminant Intrinsic Images, Spectral Intrinsic Images, As Realistic As Possible, and competitive results on Intrinsic Images in the Wild datasets while achieving state-of-the-art shading estimations.
Anil S. Baslamisli, Yang Liu 0009, Sezer Karaoglu, Theo Gevers
Comput. Vis. Image Underst.4
2021 Pose invariant age estimation of face images in the wild
Wei Wang 0469, Sezer Karaoglu, Wei Zeng 0016, Theo Gevers
Comput. Vis. Image Underst.5
2021 Automatic generation of dense non-rigid optical flow
abstract
There hardly exists any large-scale datasets with dense optical flow of non-rigid motion from real-world imagery as of today. The reason lies mainly in the required setup to derive ground truth optical flows: a series of images with known camera poses along its trajectory, and an accurate 3D model from a textured scene. Human annotation is not only too tedious for large databases, it can simply hardly contribute to accurate optical flow. To circumvent the need for manual annotation, we propose a framework to automatically generate optical flow from real-world videos. The method extracts and matches objects from video frames to compute initial constraints, and applies a deformation over the objects of interest to obtain dense optical flow fields. We propose several ways to augment the optical flow variations. Extensive experimental results show that training on our automatically generated optical flow outperforms methods that are trained on rigid synthetic data using FlowNet-S, LiteFlowNet, PWC-Net, and RAFT. Datasets and implementation of our optical flow generation framework are released at https://github.com/lhoangan/arap_flow.
Hoang-An Le, Tushar Nimbhorkar, Thomas Mensink, Anil S. Baslamisli, Sezer Karaoglu, Theo Gevers
Comput. Vis. Image Underst.6
2021 ShadingNet: Image Intrinsics by Fine-Grained Shading Decomposition
abstract
Abstract In general, intrinsic image decomposition algorithms interpret shading as one unified component including all photometric effects. As shading transitions are generally smoother than reflectance (albedo) changes, these methods may fail in distinguishing strong photometric effects from reflectance variations. Therefore, in this paper, we propose to decompose the shading component into direct (illumination) and indirect shading (ambient light and shadows) subcomponents. The aim is to distinguish strong photometric effects from reflectance variations. An end-to-end deep convolutional neural network (ShadingNet) is proposed that operates in a fine-to-coarse manner with a specialized fusion and refinement unit exploiting the fine-grained shading model. It is designed to learn specific reflectance cues separated from specific photometric effects to analyze the disentanglement capability. A large-scale dataset of scene-level synthetic images of outdoor natural environments is provided with fine-grained intrinsic image ground-truths. Large scale experiments show that our approach using fine-grained shading decompositions outperforms state-of-the-art algorithms utilizing unified shading on NED, MPI Sintel, GTA V, IIW, MIT Intrinsic Images, 3DRMS and SRD datasets.
Anil S. Baslamisli, Partha Das, Hoang-An Le, Sezer Karaoglu, Theo Gevers
Int. J. Comput. Vis.5
2020 MMD Based Discriminative Learning for Face Forgery Detection
Theo Gevers
ACCV (5)2
2020 Unified Application of Style Transfer for Face Swapping and Reenactment
Minh Ngô, Christian aan de Wiel, Sezer Karaoglu, Theo Gevers
ACCV (5)4
2020 Novel View Synthesis from Single Images via Point Cloud Transformation
Hoang-An Le, Thomas Mensink, Partha Das, Theo Gevers
BMVC4
2020 Pano2Scene: 3D Indoor Semantic Scene Reconstruction from a Single Indoor Panorama Image
Wei Zeng 0016, Sezer Karaoglu, Theo Gevers
BMVC3
2020 Kinship Identification Through Joint Learning Using Kinship Verification Ensembles
Wei Wang 0469, Shaodi You, Theo Gevers
ECCV (22)3
2020 Joint 3D Layout and Depth Prediction from a Single Indoor Panorama Image
Wei Zeng 0016, Sezer Karaoglu, Theo Gevers
ECCV (16)3
2020 Object features and face detection performance: Analyses with 3D-rendered synthetic data
abstract
This paper is to provide an overview of how object features from images influence face detection performance, and how to select synthetic faces to address specific features. To this end, we investigate the effects of occlusion, scale, viewpoint, background, and noise by using a novel synthetic image generator based on 3DU Face Dataset. To examine the effects of different features, we selected three detectors (Faster RCNN, HR, SSH) as representative of various face detection methodologies. Comparing different configurations of synthetic data on face detection systems, it showed that our synthetic dataset could complement face detectors to become more robust against features in the real world. Our analysis also demonstrated that a variety of data augmentation is necessary to address nuanced differences in performance.
Sezer Karaoglu, Hoang-An Le, Theo Gevers
ICPR4
2020 Orthographic Projection Linear Regression for Single Image 3D Human Pose Estimation
abstract
3D human pose estimation from a single 2D image in the wild is an important computer vision task but yet extremely challenging. Unlike images taken from indoor and well constrained environments, 2D outdoor images in the wild are extremely complex because of varying imaging conditions. Furthermore, 2D images usually do not have corresponding 3D pose ground truth making a supervised approach ill-constrained. Therefore, in this paper, we propose to associate the 3D human pose, the 2D human pose projection and the 2D image appearance through a new orthographic projection based linear regression module. Unlike existing reprojection based approaches, our orthographic projection and regression do not suffer from small angle problems, which usually lead to overfitting in the depth dimension. Hence, we propose a deep neural network which adopts the 2D pose, 3D pose regression and orthographic projection linear regression module. The proposed method shows state-of-the-art performance on the Human3.6M dataset and generalizes well to in-the-wild images.
Shaodi You, Theo Gevers
ICPR3
2020 On the benefit of adversarial training for monocular depth estimation
abstract
In this paper we address the benefit of adding adversarial training to the task of monocular depth estimation. A model can be trained in a self-supervised setting on stereo pairs of images, where depth (disparities) are an intermediate result in a right-to-left image reconstruction pipeline. For the quality of the image reconstruction and disparity prediction, a combination of different losses is used, including L1 image reconstruction losses and left–right disparity smoothness. These are local pixel-wise losses, while depth prediction requires global consistency. Therefore, we extend the self-supervised network to become a Generative Adversarial Network (GAN), by including a discriminator which should tell apart reconstructed (fake) images from real images. We evaluate Vanilla GANs, LSGANs and Wasserstein GANs in combination with different pixel-wise reconstruction losses. Based on extensive experimental evaluation, we conclude that adversarial training is beneficial if and only if the reconstruction loss is not too constrained. Even though adversarial training seems promising because it promotes global consistency, non-adversarial training outperforms (or is on par with) any method trained with a GAN when a constrained reconstruction loss is used in combination with batch normalisation. Based on the insights of our experimental evaluation we obtain state-of-the art monocular depth estimation results by using batch normalisation and different output scales.
Rick Groenendijk, Sezer Karaoglu, Theo Gevers, Thomas Mensink
Comput. Vis. Image Underst.3
2020 Spatial-temporal dual-actor CNN for human interaction prediction in video
Mahlagha Afrasiabi, Hassan Khotanlou, Theo Gevers
Multim. Tools Appl.3
2020 Automatic Estimation of Taste Liking Through Facial Expression Dynamics
abstract
The level of taste liking is an important measure for a number of applications such as the prediction of long-term consumer acceptance for different food and beverage products. Based on the fact that facial expressions are spontaneous, instant and heterogeneous sources of information, this paper aims to automatically estimate the level of taste liking through facial expression videos. Instead of using handcrafted features, the proposed approach deep learns the regional expression dynamics, and encodes them to a Fisher vector for video representation. Regional Fisher vectors are then concatenated, and classified by linear SVM classifiers. The aim is to reveal the hidden patterns of taste-elicited responses by exploiting expression dynamics such as the speed and acceleration of facial movements. To this end, we have collected the first large-scale beverage tasting database in the literature. The database has 2,970 videos of taste-induced facial expressions collected from 495 subjects. Our large-scale experiments on this database show that the proposed approach achieves an accuracy of 70.37 percent for distinguishing between three levels of taste-liking. Furthermore, we assess the human performance recruiting 45 participants, and show that humans are significantly less reliable for estimating taste appreciation from facial expressions in comparison to the proposed method.
Hamdi Dibeklioglu, Theo Gevers
IEEE Trans. Affect. Comput.2
2018 Three for one and one for three: Flow, Segmentation, and Surface Normals
Hoang-An Le, Anil S. Baslamisli, Thomas Mensink, Theo Gevers
BMVC4
2018 CNN Based Learning Using Reflection and Retinex Models for Intrinsic Image Decomposition
abstract
Most of the traditional work on intrinsic image decomposition rely on deriving priors about scene characteristics. On the other hand, recent research use deep learning models as in-and-out black box and do not consider the well-established, traditional image formation process as the basis of their intrinsic learning process. As a consequence, although current deep learning approaches show superior performance when considering quantitative benchmark results, traditional approaches are still dominant in achieving high qualitative results. In this paper, the aim is to exploit the best of the two worlds. A method is proposed that (1) is empowered by deep learning capabilities, (2) considers a physics-based reflection model to steer the learning process, and (3) exploits the traditional approach to obtain intrinsic images by exploiting reflectance and shading gradient information. The proposed model is fast to compute and allows for the integration of all intrinsic components. To train the new model, an object centered large-scale datasets with intrinsic ground-truth images are created. The evaluation results demonstrate that the new model outperforms existing methods. Visual inspection shows that the image formation loss function augments color reproduction and the use of gradient information produces sharper edges. Datasets, models and higher resolution images are available at https://ivi.fnwi.uva.nl/cv/retinet.
Anil S. Baslamisli, Hoang-An Le, Theo Gevers
CVPR3
2018 Joint Learning of Intrinsic Images and Semantic Segmentation
Anil S. Baslamisli, Thomas T. Groenestege, Partha Das, Hoang-An Le, Sezer Karaoglu, Theo Gevers
ECCV (6)6
2018 Expression-Invariant Age Estimation Using Structured Learning
abstract
In this paper, we investigate and exploit the influence of facial expressions on automatic age estimation. Different from existing approaches, our method jointly learns the age and expression by introducing a new graphical model with a latent layer between the age/expression labels and the features. This layer aims to learn the relationship between the age and expression and captures the face changes which induce the aging and expression appearance, and thus obtaining expression-invariant age estimation. Conducted on three age-expression datasets (FACES , Lifespan and NEMO ), our experiments illustrate the improvement in performance when the age is jointly learnt with expression in comparison to expression-independent age estimation. The age estimation error is reduced by 14.43, 37.75 and 9.30 percent for the FACES, Lifespan and NEMO datasets respectively. The results obtained by our graphical model, without prior-knowledge of the expressions of the tested faces, are better than the best reported ones for all datasets. The flexibility of the proposed model to include more cues is explored by incorporating gender together with age and expression. The results show performance improvements for all cues.
Zhongyu Lou, Fares Alnajar, José M. Álvarez 0004, Ninghang Hu, Theo Gevers
IEEE Trans. Pattern Anal. Mach. Intell.5
2017 Auto-Calibrated Gaze Estimation Using Human Gaze Patterns
abstract
We present a novel method to auto-calibrate gaze estimators based on gaze patterns obtained from other viewers. Our method is based on the observation that the gaze patterns of humans are indicative of where a new viewer will look at. When a new viewer is looking at a stimulus, we first estimate a topology of gaze points (initial gaze points). Next, these points are transformed so that they match the gaze patterns of other humans to find the correct gaze points. In a flexible uncalibrated setup with a web camera and no chin rest, the proposed method is tested on ten subjects and ten images. The method estimates the gaze points after looking at a stimulus for a few seconds with an average error below $$4.5^{\circ }$$ . Although the reported performance is lower than what could be achieved with dedicated hardware or calibrated setup, the proposed method still provides sufficient accuracy to trace the viewer attention. This is promising considering the fact that auto-calibration is done in a flexible setup , without the use of a chin rest, and based only on a few seconds of gaze initialization data. To the best of our knowledge, this is the first work to use human gaze patterns in order to auto-calibrate gaze estimators.
Fares Alnajar, Theo Gevers, Roberto Valenti, Sennay Ghebreab
Int. J. Comput. Vis.2
2017 Point Light Source Position Estimation From RGB-D Images by Learning Surface Attributes
abstract
Light source position (LSP) estimation is a difficult yet an important problem in computer vision. A common approach for estimating the LSP assumes Lambert's law. However, in real-world scenes, Lambert's law does not hold for all different types of surfaces. Instead of assuming all that surfaces follow Lambert's law, our approach classifies image surface segments based on their photometric and geometric surface attributes (i.e. glossy, matte, curved, and so on) and assigns weights to image surface segments based on their suitability for LSP estimation. In addition, we propose the use of the estimated camera pose to globally constrain LSP for RGB-D video sequences. Experiments on Boom and a newly collected RGB-D video data sets show that the state-of-the-art methods are outperformed by the proposed method. The results demonstrate that weighting image surface segments based on their attributes outperform the state-of-the-art methods in which the image surface segments are considered to equally contribute. In particular, by using the proposed surface weighting, the angular error for LSP estimation is reduced from 12.6° to 8.2° and 24.6° to 4.8° for Boom and RGB-D video data sets, respectively. Moreover, using the camera pose to globally constrain LSP provides higher accuracy (4.8°) compared with using single frames (8.5°).
Sezer Karaoglu, Yang Liu 0009, Theo Gevers, Arnold W. M. Smeulders
IEEE Trans. Image Process.3
2017 Con-Text: Text Detection for Fine-Grained Object Classification
abstract
This paper focuses on fine-grained object classification using recognized scene text in natural images. While the state-of-the-art relies on visual cues only, this paper is the first work which proposes to combine textual and visual cues. Another novelty is the textual cue extraction. Unlike the state-of-the-art text detection methods, we focus more on the background instead of text regions. Once text regions are detected, they are further processed by two methods to perform text recognition, i.e., ABBYY commercial OCR engine and a state-of-the-art character recognition algorithm. Then, to perform textual cue encoding, bi- and trigrams are formed between the recognized characters by considering the proposed spatial pairwise constraints. Finally, extracted visual and textual cues are combined for fine-grained classification. The proposed method is validated on four publicly available data sets: ICDAR03, ICDAR13, Con-Text, and Flickr-logo. We improve the state-of-the-art end-to-end character recognition by a large margin of 15% on ICDAR03. We show that textual cues are useful in addition to visual cues for fine-grained classification. We show that textual cues are also useful for logo retrieval. Adding textual cues outperforms visual- and textual-only in fine-grained classification (70.7% to 60.3%) and logo retrieval (57.4% to 54.8%).
Sezer Karaoglu, Ran Tao 0004, Jan C. van Gemert, Theo Gevers
IEEE Trans. Image Process.4
2017 Words Matter: Scene Text for Image Classification and Retrieval
abstract
Text in natural images typically adds meaning to an object or scene. In particular, text specifies which business places serve drinks (e.g., cafe, teahouse) or food (e.g., restaurant, pizzeria), and what kind of service is provided (e.g., massage, repair). The mere presence of text, its words, and meaning are closely related to the semantics of the object or scene. This paper exploits textual contents in images for fine-grained business place classification and logo retrieval. There are four main contributions. First, we show that the textual cues extracted by the proposed method are effective for the two tasks. Combining the proposed textual and visual cues outperforms visual only classification and retrieval by a large margin. Second, to extract the textual cues, a generic and fully unsupervised word box proposal method is introduced. The method reaches state-of-the-art word detection recall with a limited number of proposals. Third, contrary to what is widely acknowledged in text detection literature, we demonstrate that high recall in word detection is more important than high f-score at least for both tasks considered in this work. Last, this paper provides a large annotated text detection dataset with 10 K images and 27 601 word boxes.
Sezer Karaoglu, Ran Tao 0004, Theo Gevers, Arnold W. M. Smeulders
IEEE Trans. Multim.3
2016 Detect2Rank: Combining Object Detectors Using Learning to Rank
abstract
Object detection is an important research area in the field of computer vision. Many detection algorithms have been proposed. However, each object detector relies on specific assumptions of the object appearance and imaging conditions. As a consequence, no algorithm can be considered universal. With the large variety of object detectors, the subsequent question is how to select and combine them. In this paper, we propose a framework to learn how to combine object detectors. The proposed method uses (single) detectors like Deformable Part Models, Color Names and Ensemble of Exemplar-SVMs, and exploits their correlation by high-level contextual features to yield a combined detection list. Experiments on the PASCAL VOC07 and VOC10 data sets show that the proposed method significantly outperforms single object detectors, DPM (8.4%), CN (6.8%) and EES (17.0%) on VOC07 and DPM (6.5%), CN (5.5%) and EES (16.2%) on VOC10. We show with an experiment that there are no constraints on the type of the detector. The proposed method outperforms (2.4%) the state-of-the-art object detector (RCNN) on VOC07 when Regions with Convolutional Neural Network is combined with other detectors used in this paper.
Sezer Karaoglu, Yang Liu 0009, Theo Gevers
IEEE Trans. Image Process.3
2015 Color Constancy by Deep Learning
abstract
Computational color constancy aims to estimate the color of the light source.The performance of many vision tasks, such as object detection and scene understanding, may benefit from color constancy by using the corrected object colors.Since traditional color constancy methods are based on specific assumptions, none of those methods can be used as a universal predictor.Further, shallow learning schemes are used for trainingbased color constancy, possibly suffering from limited learning capabilities.In this paper, we propose a new framework using Deep Neural Networks (DNNs) to obtain accurate light source estimation.We reformulate color constancy as a DNN-based regression approach to estimate the color of the light source.The model is trained using datasets of more than a million images.Experiments show that the proposed algorithm outperforms the state-of-the-art by 9%.Especially in cross dataset validation, our approach reduces the median angular error by 35%.Our algorithm operates at more than 100 fps during testing.
Zhongyu Lou, Theo Gevers, Ninghang Hu, Marcel P. Lucassen
BMVC2
2015 Age estimation under changes in image quality: An experimental study
abstract
In this paper, we investigate the influence of image quality on the performance of aging features. Age estimation systems used or designed a number of aging features to capture the aging cues from the face such as skin texture and wrinkles. These aging cues are sensitive to small changes in the imaging conditions which suggests considering the imaging quality when extracting such information. Although interesting performances are reported on various datasets, the effect of image quality has not been addressed. We introduce a scheme to explore the influence of image quality on the performance of appearance aging features. A number of datasets are experimented on where artifacts resulted from different types of noise are considered. Finally, we propose a method to automatically apply the most suitable features based on the quality of the image. The results show that better or comparable performance is obtained when automatically applying different features, based on image quality, in comparison to a single (best) feature type.
Fares Alnajar, Theo Gevers, Sezer Karaoglu
ICIP2
2015 Per-patch metric learning for robust image matching
abstract
We propose a patch-specific metric learning method to improve matching performance of local descriptors. Existing methodologies typically focus on invariance, by completely considering, or completely disregarding all variations. We propose a metric learning method that is robust to only a range of variations. The ability to choose the level of robustness allows us to fine-tune the trade-off between invariance and discriminative power. We learn a distance metric for each patch independently by sampling from a set of relevant image transformations. These transformations give a-priori knowledge about the behavior of the query patch under the applied transformation in feature space. We learn the robust metric by either fully generating only the relevant range of transformations, or by a novel direct metric. The matching between query patch and data is performed with this new metric. Results on the ALOI dataset show that the proposed method improves performance of SIFT by 6.22% for geometric and 4.43% for photometric transformations.
Sezer Karaoglu, Ivo Everts, Jan C. van Gemert, Theo Gevers
ICIP4
2015 Color constancy by combining low-mid-high level image cues
Yang Liu 0009, Theo Gevers
Comput. Vis. Image Underst.2
2015 SuperPixel based mid-level image description for image recognition
H. Emrah Tasli, Ronan Sicre, Theo Gevers
J. Vis. Commun. Image Represent.3
2015 Combining Facial Dynamics With Appearance for Age Estimation
abstract
Estimating the age of a human from the captured images of his/her face is a challenging problem. In general, the existing approaches to this problem use appearance features only. In this paper, we show that in addition to appearance information, facial dynamics can be leveraged in age estimation. We propose a method to extract and use dynamic features for age estimation, using a person's smile. Our approach is tested on a large, gender-balanced database with 400 subjects, with an age range between 8 and 76. In addition, we introduce a new database on posed disgust expressions with 324 subjects in the same age range, and evaluate the reliability of the proposed approach when used with another expression. State-of-the-art appearance-based age estimation methods from the literature are implemented as baseline. We demonstrate that for each of these methods, the addition of the proposed dynamic features results in statistically significant improvement. We further propose a novel hierarchical age estimation architecture based on adaptive age grouping. We test our approach extensively, including an exploration of spontaneous versus posed smile dynamics, and gender-specific age estimation. We show that using spontaneity information reduces the mean absolute error by up to 21%, advancing the state of the art for facial age estimation.
Hamdi Dibeklioglu, Fares Alnajar, Albert Ali Salah, Theo Gevers
IEEE Trans. Image Process.4
2015 Estimation of Sunlight Direction Using 3D Object Models
abstract
The direction of sunlight is an important informative cue in a number of applications in image processing, such as augmented reality and object recognition. In general, existing methods to estimate the direction of the sunlight rely on different image features (e.g., sky, texture, shadows, and shading). These features can be considered as weak informative cues as no single feature can reliably estimate the sunlight direction. Moreover, existing methods may require that the camera parameters are known limiting their applicability. In this paper, we present a new method to estimate the sunlight direction from a single (outdoor) image by inferring casts shadows through object modeling and recognition. First, objects (e.g., cars or persons) are first (automatically) recognized in images by exemplar-SVMs. Instead of training the Support Vector Machine (SVMs) using natural images (limited variation in viewpoints), we propose to train on 2D object samples generated from 3D object models. Then, the recognized objects are used as sundial cues (probes) to estimate the sunlight direction by inferring the corresponding shadows generated by 3D object models considering different illumination directions. We demonstrate the effectiveness of our approach on synthetic and real images. Experiments show that our method estimates the azimuth angle accurately within a quadrant (smaller than 45°) and compute the zenith angle with mean angular error of 23°.
Yang Liu 0009, Theo Gevers
IEEE Trans. Image Process.2
2015 Extracting 3D Layout From a Single Image Using Global Image Structures
abstract
Extracting the pixel-level 3D layout from a single image is important for different applications, such as object localization, image, and video categorization. Traditionally, the 3D layout is derived by solving a pixel-level classification problem. However, the image-level 3D structure can be very beneficial for extracting pixel-level 3D layout since it implies the way how pixels in the image are organized. In this paper, we propose an approach that first predicts the global image structure, and then we use the global structure for fine-grained pixel-level 3D layout extraction. In particular, image features are extracted based on multiple layout templates. We then learn a discriminative model for classifying the global layout at the image-level. Using latent variables, we implicitly model the sublevel semantics of the image, which enrich the expressiveness of our model. After the image-level structure is obtained, it is used as the prior knowledge to infer pixel-wise 3D layout. Experiments show that the results of our model outperform the state-of-the-art methods by 11.7% for 3D structure classification. Moreover, we show that employing the 3D structure prior information yields accurate 3D scene layout segmentation.
Zhongyu Lou, Theo Gevers, Ninghang Hu
IEEE Trans. Image Process.2
2015 Recognition of Genuine Smiles
abstract
Automatic distinction between genuine (spontaneous) and posed expressions is important for visual analysis of social signals. In this paper, we describe an informative set of features for the analysis of face dynamics, and propose a completely automatic system to distinguish between genuine and posed enjoyment smiles. Our system incorporates facial landmarking and tracking, through which features are extracted to describe the dynamics of eyelid, cheek, and lip corner movements. By fusing features over different regions, as well as over different temporal phases of a smile, we obtain a very accurate smile classifier. We systematically investigate age and gender effects, and establish that age-specific classification significantly improves the results, even when the age is automatically estimated. We evaluate our system on the 400-subject UvA-NEMO database we have recently collected, as well as on three other smile databases from the literature . Through an extensive experimental evaluation, we show that our system improves the state of the art in smile classification and provides useful insights in smile psychophysics.
Hamdi Dibeklioglu, Albert Ali Salah, Theo Gevers
IEEE Trans. Multim.3
2014 Expression-Invariant Age Estimation
Fares Alnajar, Zhongyu Lou, José M. Álvarez 0004, Theo Gevers
BMVC4
2014 DENSE sampling of features for image retrieval
abstract
This paper focuses on the image retrieval task. We propose the use of dense feature points computed on several color channels to improve the retrieval system. To validate our approach, an evaluation of various SIFT extraction strategies is performed. Detected SIFT are compared with dense SIFT. Dense color descriptors: C-SIFT and T-SIFT are then utilized. A comparison between standard and rotation invariant features is further achieved. Finally, several encoding strategies are studied: Bag of Visual Words (BOW), Fisher vectors, and vector of locally aggregated descriptors (VLAD). The presented approaches are evaluated on several datasets and we show a large improvement over the baseline.
Ronan Sicre, Theo Gevers
ICIP2
2014 Geometry-constrained spatial pyramid adaptation for image classification
abstract
This paper proposes a geometry-constrained spatial pyramid adaptation approach for the image classification task. Scene geometry is used as an input parameter for generating the spatial pyramid definitions. The resulting region adaptation is performed in accordance with the predefined geometric guidelines and underlying image characteristics. Using an approximate global geometric correspondence, exploits the idea that images of the same category share a spatial similarity. This assumption is evaluated and justified in an object classification framework, in which generated region segments are used as an enhancement to the widely utilized “spatial pyramid” method. Fixed region pyramids are replaced by the proposed locally coherent geometrically consistent region segments. Performance of the proposed method on object classification framework is evaluated on the 20 class Pascal VOC 2007 dataset. The proposed method shows consistent increase in the mean average precision (MAP) score for different experimental scenarios.
H. Emrah Tasli, Ronan Sicre, Theo Gevers, A. Aydin Alatan
ICIP3
2014 SuperPixel Based Angular Differences as a Mid-level Image Descriptor
abstract
This paper focuses on the object recognition task and aims at improving the accuracy with an emphasis on the feature extraction step. Feature extraction is widely used in image classification as an initial step in the pipeline. In this paper, we propose a method to explore the conventional feature extraction techniques from the perspective that mid-level information could be incorporated in order to obtain a superior scene description. We hypothesize that the commonly used pixel based low-level descriptions are useful but can be improved with the introduction of mid-level region information. Hence, we investigate super pixel based image representation to acquire such mid-level information in order to improve the classification accuracy. Detailed experimental evaluations on classification and retrieval tasks are performed in order to validate the proposed hypothesis. A consistent increase is observed in the mean average precision (MAP) score for different experimental scenarios and image categories.
Ronan Sicre, H. Emrah Tasli, Theo Gevers
ICPR3
2014 Color Constancy Using 3D Scene Geometry Derived From a Single Image
abstract
The aim of color constancy is to remove the effect of the color of the light source. As color constancy is inherently an ill-posed problem, most of the existing color constancy algorithms are based on specific imaging assumptions (e.g., gray-world and white patch assumption). In this paper, 3D geometry models are used to determine which color constancy method to use for the different geometrical regions (depth/layer) found in images. The aim is to classify images into stages (rough 3D geometry models). According to stage models, images are divided into stage regions using hard and soft segmentation. After that, the best color constancy methods are selected for each geometry depth. To this end, we propose a method to combine color constancy algorithms by investigating the relation between depth, local image statistics, and color constancy. Image statistics are then exploited per depth to select the proper color constancy method. Our approach opens the possibility to estimate multiple illuminations by distinguishing nearby light source from distant illuminations. Experiments on state-of-the-art data sets show that the proposed algorithm outperforms state-of-the-art single color constancy algorithms with an improvement of almost 50% of median angular error. When using a perfect classifier (i.e, all of the test images are correctly classified into stages); the performance of the proposed method achieves an improvement of 52% of the median angular error compared with the best-performing single color constancy algorithm.
Noha M. Elfiky, Theo Gevers, Arjan Gijsenij, Jordi Gonzàlez 0001
IEEE Trans. Image Process.2
2014 Evaluation of Color Spatio-Temporal Interest Points for Human Action Recognition
abstract
This paper considers the recognition of realistic human actions in videos based on spatio-temporal interest points (STIPs). Existing STIP-based action recognition approaches operate on intensity representations of the image data. Because of this, these approaches are sensitive to disturbing photometric phenomena, such as shadows and highlights. In addition, valuable information is neglected by discarding chromaticity from the photometric representation. These issues are addressed by color STIPs. Color STIPs are multichannel reformulations of STIP detectors and descriptors, for which we consider a number of chromatic and invariant representations derived from the opponent color space. Color STIPs are shown to outperform their intensity-based counterparts on the challenging UCF sports, UCF11 and UCF50 action recognition benchmarks by more than 5% on average, where most of the gain is due to the multichannel descriptors. In addition, the results show that color STIPs are currently the single best low-level feature choice for STIP-based approaches to human action recognition.
Ivo Everts, Jan C. van Gemert, Theo Gevers
IEEE Trans. Image Process.3
2014 Robustifying Descriptor Instability Using Fisher Vectors
abstract
Many computer vision applications, including image classification, matching, and retrieval use global image representations, such as the Fisher vector, to encode a set of local image patches. To describe these patches, many local descriptors have been designed to be robust against lighting changes and noise. However, local image descriptors are unstable when the underlying image signal is low. Such low-signal patches are sensitive to small image perturbations, which might come e.g., from camera noise or lighting effects. In this paper, we first quantify the relation between the signal strength of a patch and the instability of that patch, and second, we extend the standard Fisher vector framework to explicitly take the descriptor instabilities into account. In comparison to common approaches to dealing with descriptor instabilities, our results show that modeling local descriptor instability is beneficial for object matching, image retrieval, and classification.
Ivo Everts, Jan C. van Gemert, Thomas Mensink, Theo Gevers
IEEE Trans. Image Process.4
2014 Combining Priors, Appearance, and Context for Road Detection
abstract
Detecting the free road surface ahead of a moving vehicle is an important research topic in different areas of computer vision, such as autonomous driving or car collision warning. Current vision-based road detection methods are usually based solely on low-level features. Furthermore, they generally assume structured roads, road homogeneity, and uniform lighting conditions, constraining their applicability in real-world scenarios. In this paper, road priors and contextual information are introduced for road detection. First, we propose an algorithm to estimate road priors online using geographical information, providing relevant initial information about the road location. Then, contextual cues, including horizon lines, vanishing points, lane markings, 3-D scene layout, and road geometry, are used in addition to low-level cues derived from the appearance of roads. Finally, a generative model is used to combine these cues and priors, leading to a road detection method that is, to a large degree, robust to varying imaging conditions, road types, and scenarios.
José M. Álvarez 0004, Antonio M. López 0001, Theo Gevers, Felipe Lumbreras
IEEE Trans. Intell. Transp. Syst.3
2014 Image Alignment by Piecewise Planar Region Matching
abstract
Robust image registration is a challenging problem, especially when dealing with severe changes in illumination and viewpoint. Previous methods assume a global geometric model (e.g., homography) and, hence, are only able to align images under predefined constraints (e.g., planar scenes and parallax-free camera motion). However, these constraints may not hold for natural scenes and uncontrolled imaging conditions. Therefore, this paper proposes a novel method which approximates image regions with planes by incorporating piecewise local geometric models. The approximated planar regions are obtained by exploiting a hierarchical figure-ground segmentation method. Each such planar region assumes an affine transformation. To achieve the alignment of the planar regions, an energy function is defined which employs intensity, a key-point descriptor, and geometric information under a global constraint. By re-segmenting and re-merging planar regions iteratively in an energy minimization framework, the method is able to align images even under significant changes in illumination and viewpoint. Experiments on two datasets show that the proposed method outperforms state-of-the-art, especially in the case of large appearance variations and it is, therefore, applicable to web-images (i.e., unconstrained setting) which are taken from the same scene with different viewpoints.
Zhongyu Lou, Theo Gevers
IEEE Trans. Multim.2
2014 Extracting Primary Objects by Video Co-Segmentation
abstract
Video object segmentation is a challenging problem. Without human annotation or other prior information, it is hard to select a meaningful primary object from a single video, so extracting the primary object across videos is a more promising approach. However, existing algorithms consider the problem as foreground/background segmentation. Therefore, we propose an algorithm that learns the model of the primary object by representing the frames/videos as a graphical model. The probabilistic graphical model is built across a set of videos based on an object proposal algorithm. Our approach considers appearance, spatial, and temporal consistency of the primary objects. A new dataset is created to evaluate the proposed method and to compare it to the state-of-the-art on video object co-segmentation. The experiments show that our method obtains state-of-the-art results, outperforming other algorithms by 1.5% (pixel accuracy) on the MOViCS dataset and 9.6% (pixel accuracy) on the new dataset.
Zhongyu Lou, Theo Gevers
IEEE Trans. Multim.2
2013 Evaluation of Color STIPs for Human Action Recognition
abstract
This paper is concerned with recognizing realistic human actions in videos based on spatio-temporal interest points (STIPs). Existing STIP-based action recognition approaches operate on intensity representations of the image data. Because of this, these approaches are sensitive to disturbing photometric phenomena such as highlights and shadows. Moreover, valuable information is neglected by discarding chromaticity from the photometric representation. These issues are addressed by Color STIPs. Color STIPs are multi-channel reformulations of existing intensity-based STIP detectors and descriptors, for which we consider a number of chromatic representations derived from the opponent color space. This enhanced modeling of appearance improves the quality of subsequent STIP detection and description. Color STIPs are shown to substantially outperform their intensity-based counterparts on the challenging UCF~sports, UCF11 and UCF50 action recognition benchmarks. Moreover, the results show that color STIPs are currently the single best low-level feature choice for STIP-based approaches to human action recognition.
Ivo Everts, Jan C. van Gemert, Theo Gevers
CVPR3
2013 Calibration-Free Gaze Estimation Using Human Gaze Patterns
abstract
We present a novel method to auto-calibrate gaze estimators based on gaze patterns obtained from other viewers. Our method is based on the observation that the gaze patterns of humans are indicative of where a new viewer will look at. When a new viewer is looking at a stimulus, we first estimate a topology of gaze points (initial gaze points). Next, these points are transformed so that they match the gaze patterns of other humans to find the correct gaze points. In a flexible uncalibrated setup with a web camera and no chin rest, the proposed method was tested on ten subjects and ten images. The method estimates the gaze points after looking at a stimulus for a few seconds with an average accuracy of 4:3°. Although the reported performance is lower than what could be achieved with dedicated hardware or calibrated setup, the proposed method still provides a sufficient accuracy to trace the viewer attention. This is promising considering the fact that auto-calibration is done in a flexible setup, without the use of a chin rest, and based only on a few seconds of gaze initialization data. To the best of our knowledge, this is the first work to use human gaze patterns in order to auto-calibrate gaze estimators.
Fares Alnajar, Theo Gevers, Roberto Valenti, Sennay Ghebreab
ICCV2
2013 Like Father, Like Son: Facial Expression Dynamics for Kinship Verification
abstract
Kinship verification from facial appearance is a difficult problem. This paper explores the possibility of employing facial expression dynamics in this problem. By using features that describe facial dynamics and spatio-temporal appearance over smile expressions, we show that it is possible to improve the state of the art in this problem, and verify that it is indeed possible to recognize kinship by resemblance of facial expressions. The proposed method is tested on different kin relationships. On the average, 72.89% verification accuracy is achieved on spontaneous smiles.
Hamdi Dibeklioglu, Albert Ali Salah, Theo Gevers
ICCV3
2013 Super pixel extraction via convexity induced boundary adaptation
abstract
This study presents an efficient super-pixel extraction algorithm with major contributions to the state-of-the-art in terms of accuracy and computational complexity. Segmentation accuracy is improved through convexity constrained geodesic distance utilization; while computational efficiency is achieved by replacing complete region processing with boundary adaptation idea. Starting from the uniformly distributed rectangular equal-sized super-pixels, region boundaries are adapted to intensity edges iteratively by assigning boundary pixels to the most similar neighboring super-pixels. At each iteration, super-pixel regions are updated and hence progressively converging to compact pixel groups. Experimental results with state-of-the-art comparisons, validate the performance of the proposed technique in terms of both accuracy and speed.
H. Emrah Tasli, Cevahir Çigla, Theo Gevers, A. Aydin Alatan
ICME3
2013 Con-text: text detection using background connectivity for fine-grained object classification
abstract
This paper focuses on fine-grained classification by detecting photographed text in images. We introduce a text detection method that does not try to detect all possible foreground text regions but instead aims to reconstruct the scene background to eliminate non-text regions. Object cues such as color, contrast, and objectiveness are used in corporation with a random forest classifier to detect background pixels in the scene. Results on two publicly available datasets ICDAR03 and a fine-grained Building subcategories of ImageNet shows the effectiveness of the proposed method.
Sezer Karaoglu, Jan C. van Gemert, Theo Gevers
ACM Multimedia3
2013 Spot the differences: from a photograph burst to the single best picture
abstract
With the rise of the digital camera, people nowadays typically take several near-identical photos of the same scene to maximize the chances of a good shot. This paper proposes a user-friendly tool for exploring a personal photo gallery for selecting or even creating the best shot of a scene between its multiple alternatives. This functionality is realized through a graphical user interface where the best viewpoint can be selected from a generated panorama of the scene. Once the viewpoint is selected, the user is able to go explore possible alternatives coming from the other images. Using this tool, one can explore a photo gallery efficiently. Moreover, additional compositions from other images are also possible. With such additional compositions, one can go from a burst of photographs to the single best one. Even funny compositions of images, where you can duplicate a person in the same image, are possible with our proposed tool.
H. Emrah Tasli, Jan C. van Gemert, Theo Gevers
ACM Multimedia3
2013 Selective Search for Object Recognition
Jasper R. R. Uijlings, Koen E. A. van de Sande, Theo Gevers, Arnold W. M. Smeulders
Int. J. Comput. Vis.3
2013 Joint Attention by Gaze Interpolation and Saliency
abstract
Joint attention, which is the ability of coordination of a common point of reference with the communicating party, emerges as a key factor in various interaction scenarios. This paper presents an image-based method for establishing joint attention between an experimenter and a robot. The precise analysis of the experimenter's eye region requires stability and high-resolution image acquisition, which is not always available. We investigate regression-based interpolation of the gaze direction from the head pose of the experimenter, which is easier to track. Gaussian process regression and neural networks are contrasted to interpolate the gaze direction. Then, we combine gaze interpolation with image-based saliency to improve the target point estimates and test three different saliency schemes. We demonstrate the proposed method on a human-robot interaction scenario. Cross-subject evaluations, as well as experiments under adverse conditions (such as dimmed or artificial illumination or motion blur), show that our method generalizes well and achieves rapid gaze estimation for establishing joint attention.
Zeynep Yücel, Albert Ali Salah, Çetin Meriçli, Tekin Meriçli, Roberto Valenti, Theo Gevers
IEEE Trans. Cybern.6
2013 Road Geometry Classification by Adaptive Shape Models
abstract
Vision-based road detection is important for different applications in transportation, such as autonomous driving, vehicle collision warning, and pedestrian crossing detection. Common approaches to road detection are based on low-level road appearance (e.g., color or texture) and neglect of the scene geometry and context. Hence, using only low-level features makes these algorithms highly depend on structured roads, road homogeneity, and lighting conditions. Therefore, the aim of this paper is to classify road geometries for road detection through the analysis of scene composition and temporal coherence. Road geometry classification is proposed by building corresponding models from training images containing prototypical road geometries. We propose adaptive shape models where spatial pyramids are steered by the inherent spatial structure of road images. To reduce the influence of lighting variations, invariant features are used. Large-scale experiments show that the proposed road geometry classifier yields a high recognition rate of 73.57% ± 13.1, clearly outperforming other state-of-the-art methods. Including road shape information improves road detection results over existing appearance-based methods. Finally, it is shown that invariant features and temporal information provide robustness against disturbing imaging conditions.
José M. Álvarez 0004, Theo Gevers, Ferran Diego, Antonio M. López 0001
IEEE Trans. Intell. Transp. Syst.2
2012 Improving HOG with Image Segmentation: Application to Human Detection
Yainuvis Socarrás Salas, David Vázquez 0001, Antonio M. López 0001, David Gerónimo Gómez, Theo Gevers
ACIVS5
2012 Road Scene Segmentation from a Single Image
José M. Álvarez 0004, Theo Gevers, Yann LeCun, Antonio M. López 0001
ECCV (7)2
2012 Are You Really Smiling at Me? Spontaneous versus Posed Enjoyment Smiles
Hamdi Dibeklioglu, Albert Ali Salah, Theo Gevers
ECCV (3)3
2012 Per-patch Descriptor Selection Using Surface and Scene Properties
Ivo Everts, Jan C. van Gemert, Theo Gevers
ECCV (6)3
2012 Edge classification using photo-geometric features
Josep M. Gonfaus, Theo Gevers, Arjan Gijsenij, F. Xavier Roca, Jordi Gonzàlez 0001
ICPR2
2012 A smile can reveal your age: enabling facial dynamics in age estimation
abstract
Estimation of a person's age from the facial image has many applications, ranging from biometrics and access control to cosmetics and entertainment. Many image-based methods have been proposed for this problem. In this paper, we propose a method for the use of dynamic features in age estimation, and show that 1) the temporal dynamics of facial features can be used to improve image-based age estimation; 2) considered alone, static image-based features are more accurate than dynamic features. We have collected and annotated an extensive database of face videos from 400 subjects with an age range between 8 and 76, which allows us to extensively analyze the relevant aspects of the problem. The proposed system, which fuses facial appearance and expression dynamics, performs with a mean absolute error of 4.81 (4.87) years. This represents a significant improvement of accuracy in comparison to the sole use of appearance-based features.
Hamdi Dibeklioglu, Theo Gevers, Albert Ali Salah, Roberto Valenti
ACM Multimedia2
2012 What Are You Looking at? - Improving Visual Gaze Estimation by Saliency
abstract
In this paper we present a novel mechanism to obtain enhanced gaze estimation for subjects looking at a scene or an image. The system makes use of prior knowledge about the scene (e.g. an image on a computer screen), to define a probability map of the scene the subject is gazing at, in order to find the most probable location. The proposed system helps in correcting the fixations which are erroneously estimated by the gaze estimation device by employing a saliency framework to adjust the resulting gaze point vector. The system is tested on three scenarios: using eye tracking data, enhancing a low accuracy webcam based eye tracker, and using a head pose tracker. The correlation between the subjects in the commercial eye tracking data is improved by an average of 13.91%. The correlation on the low accuracy eye gaze tracker is improved by 59.85%, and for the head pose tracker we obtain an improvement of 10.23%. These results show the potential of the system as a way to enhance and self-calibrate different visual gaze estimation systems.
Roberto Valenti, Nicu Sebe, Theo Gevers
Int. J. Comput. Vis.3
2012 Learning-based encoding with soft assignment for age estimation under unconstrained imaging conditions
Fares Alnajar, Caifeng Shan, Theo Gevers, Jan-Mark Geusebroek
Image Vis. Comput.3
2012 Improving Color Constancy by Photometric Edge Weighting
abstract
Edge-based color constancy methods make use of image derivatives to estimate the illuminant. However, different edge types exist in real-world images, such as material, shadow, and highlight edges. These different edge types may have a distinctive influence on the performance of the illuminant estimation. Therefore, in this paper, an extensive analysis is provided of different edge types on the performance of edge-based color constancy methods. First, an edge-based taxonomy is presented classifying edge types based on their photometric properties (e.g., material, shadow-geometry, and highlights). Then, a performance evaluation of edge-based color constancy is provided using these different edge types. From this performance evaluation, it is derived that specular and shadow edge types are more valuable than material edges for the estimation of the illuminant. To this end, the (iterative) weighted Gray-Edge algorithm is proposed in which these edge types are more emphasized for the estimation of the illuminant. Images that are recorded under controlled circumstances demonstrate that the proposed iterative weighted Gray-Edge algorithm based on highlights reduces the median angular error with approximately 25 percent. In an uncontrolled environment, improvements in angular error up to 11 percent are obtained with respect to regular edge-based color constancy.
Arjan Gijsenij, Theo Gevers, Joost van de Weijer 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2012 Accurate Eye Center Location through Invariant Isocentric Patterns
abstract
Locating the center of the eyes allows for valuable information to be captured and used in a wide range of applications. Accurate eye center location can be determined using commercial eye-gaze trackers, but additional constraints and expensive hardware make these existing solutions unattractive and impossible to use on standard (i.e., visible wavelength), low-resolution images of eyes. Systems based solely on appearance are proposed in the literature, but their accuracy does not allow us to accurately locate and distinguish eye centers movements in these low-resolution settings. Our aim is to bridge this gap by locating the center of the eye within the area of the pupil on low-resolution images taken from a webcam or a similar device. The proposed method makes use of isophote properties to gain invariance to linear lighting changes (contrast and brightness), to achieve in-plane rotational invariance, and to keep low-computational costs. To further gain scale invariance, the approach is applied to a scale space pyramid. In this paper, we extensively test our approach for its robustness to changes in illumination, head pose, scale, occlusion, and eye rotation. We demonstrate that our system can achieve a significant improvement in accuracy over state-of-the-art techniques for eye center location in standard low-resolution imagery.
Roberto Valenti, Theo Gevers
IEEE Trans. Pattern Anal. Mach. Intell.2
2012 A Statistical Method for 2-D Facial Landmarking
abstract
Many facial-analysis approaches rely on robust and accurate automatic facial landmarking to correctly function. In this paper, we describe a statistical method for automatic facial-landmark localization. Our landmarking relies on a parsimonious mixture model of Gabor wavelet features, computed in coarse-to-fine fashion and complemented with a shape prior. We assess the accuracy and the robustness of the proposed approach in extensive cross-database conditions conducted on four face data sets (Face Recognition Grand Challenge, Cohn-Kanade, Bosphorus, and BioID). Our method has 99.33% accuracy on the Bosphorus database and 97.62% accuracy on the BioID database on the average, which improves the state of the art. We show that the method is not significantly affected by low-resolution images, small rotations, facial expressions, and natural occlusions such as beard and mustache. We further test the goodness of the landmarks in a facial expression recognition application and report landmarking-induced improvement over baseline on two separate databases for video-based expression recognition (Cohn-Kanade and BU-4DFE).
Hamdi Dibeklioglu, Albert Ali Salah, Theo Gevers
IEEE Trans. Image Process.3
2012 Color Constancy for Multiple Light Sources
abstract
Color constancy algorithms are generally based on the simplifying assumption that the spectral distribution of a light source is uniform across scenes. However, in reality, this assumption is often violated due to the presence of multiple light sources. In this paper, we will address more realistic scenarios where the uniform light-source assumption is too restrictive. First, a methodology is proposed to extend existing algorithms by applying color constancy locally to image patches, rather than globally to the entire image. After local (patch-based) illuminant estimation, these estimates are combined into more robust estimations, and a local correction is applied based on a modified diagonal model. Quantitative and qualitative experiments on spectral and real images show that the proposed methodology reduces the influence of two light sources simultaneously present in one scene. If the chromatic difference between these two illuminants is more than 1°, the proposed framework outperforms algorithms based on the uniform light-source assumption (with error-reduction up to approximately 30%). Otherwise, when the chromatic difference is less than 1° and the scene can be considered to contain one (approximately) uniform light source, the performance of the proposed method framework is similar to global color constancy methods.
Arjan Gijsenij, Theo Gevers
IEEE Trans. Image Process.3
2012 Sparse Color Interest Points for Image Retrieval and Object Categorization
abstract
Interest point detection is an important research area in the field of image processing and computer vision. In particular, image retrieval and object categorization heavily rely on interest point detection from which local image descriptors are computed for image matching. In general, interest points are based on luminance, and color has been largely ignored. However, the use of color increases the distinctiveness of interest points. The use of color may therefore provide selective search reducing the total number of interest points used for image matching. This paper proposes color interest points for sparse image representation. To reduce the sensitivity to varying imaging conditions, light-invariant interest points are introduced. Color statistics based on occurrence probability lead to color boosted points, which are obtained through saliency-based feature selection. Furthermore, a principal component analysis-based scale selection method is proposed, which gives a robust scale estimation per interest point. From large-scale experiments, it is shown that the proposed color interest point detector has higher repeatability than a luminance-based one. Furthermore, in the context of image retrieval, a reduced and predictable number of color features show an increase in performance compared to state-of-the-art interest points. Finally, in the context of object recognition, for the Pascal VOC 2007 challenge, our method gives comparable performance to state-of-the-art methods using only a small fraction of the features, reducing the computing time considerably.
Julian Stöttinger, Allan Hanbury, Nicu Sebe, Theo Gevers
IEEE Trans. Image Process.4
2012 Combining Head Pose and Eye Location Information for Gaze Estimation
abstract
Head pose and eye location for gaze estimation have been separately studied in numerous works in the literature. Previous research shows that satisfactory accuracy in head pose and eye location estimation can be achieved in constrained settings. However, in the presence of nonfrontal faces, eye locators are not adequate to accurately locate the center of the eyes. On the other hand, head pose estimation techniques are able to deal with these conditions; hence, they may be suited to enhance the accuracy of eye localization. Therefore, in this paper, a hybrid scheme is proposed to combine head pose and eye location information to obtain enhanced gaze estimation. To this end, the transformation matrix obtained from the head pose is used to normalize the eye regions, and in turn, the transformation matrix generated by the found eye location is used to correct the pose estimation procedure. The scheme is designed to enhance the accuracy of eye location estimations, particularly in low-resolution videos, to extend the operative range of the eye locators, and to improve the accuracy of the head pose tracker. These enhanced estimations are then combined to obtain a novel visual gaze estimation system, which uses both eye location and head information to refine the gaze estimates. From the experimental results, it can be derived that the proposed unified scheme improves the accuracy of eye estimations by 16% to 23%. Furthermore, it considerably extends its operating range by more than 15° by overcoming the problems introduced by extreme head poses. Moreover, the accuracy of the head pose tracker is improved by 12% to 24%. Finally, the experimentation on the proposed combined gaze estimation system shows that it is accurate (with a mean error between 2° and 5°) and that it can be used in cases where classic approaches would fail without imposing restraints on the position of the head.
Roberto Valenti, Nicu Sebe, Theo Gevers
IEEE Trans. Image Process.3
2011 Segmentation as selective search for object recognition
abstract
For object recognition, the current state-of-the-art is based on exhaustive search. However, to enable the use of more expensive features and classifiers and thereby progress beyond the state-of-the-art, a selective search strategy is needed. Therefore, we adapt segmentation as a selective search by reconsidering segmentation: We propose to generate many approximate locations over few and precise object delineations because (1) an object whose location is never generated can not be recognised and (2) appearance and immediate nearby context are most effective for object recognition. Our method is class-independent and is shown to cover 96.7% of all objects in the Pascal VOC 2007 test set using only 1,536 locations per image. Our selective search enables the use of the more expensive bag-of-words method which we use to substantially improve the state-of-the-art by up to 8.5% for 8 out of 20 classes on the Pascal VOC 2010 detection challenge.
Koen E. A. van de Sande, Jasper R. R. Uijlings, Theo Gevers, Arnold W. M. Smeulders
ICCV3
2011 Color Constancy Using Natural Image Statistics and Scene Semantics
abstract
Existing color constancy methods are all based on specific assumptions such as the spatial and spectral characteristics of images. As a consequence, no algorithm can be considered as universal. However, with the large variety of available methods, the question is how to select the method that performs best for a specific image. To achieve selection and combining of color constancy algorithms, in this paper natural image statistics are used to identify the most important characteristics of color images. Then, based on these image characteristics, the proper color constancy algorithm (or best combination of algorithms) is selected for a specific image. To capture the image characteristics, the Weibull parameterization (e.g., grain size and contrast) is used. It is shown that the Weibull parameterization is related to the image attributes to which the used color constancy methods are sensitive. An MoG-classifier is used to learn the correlation and weighting between the Weibull-parameters and the image attributes (number of edges, amount of texture, and SNR). The output of the classifier is the selection of the best performing color constancy method for a certain image. Experimental results show a large improvement over state-of-the-art single algorithms. On a data set consisting of more than 11,000 images, an increase in color constancy performance up to 20 percent (median angular error) can be obtained compared to the best-performing single algorithm. Further, it is shown that for certain scene categories, one specific color constancy algorithm can be used instead of the classifier considering several algorithms.
Arjan Gijsenij, Theo Gevers
IEEE Trans. Pattern Anal. Mach. Intell.2
2011 Computational Color Constancy: Survey and Experiments
abstract
Computational color constancy is a fundamental prerequisite for many computer vision applications. This paper presents a survey of many recent developments and state-of-the-art methods. Several criteria are proposed that are used to assess the approaches. A taxonomy of existing algorithms is proposed and methods are separated in three groups: static methods, gamut-based methods, and learning-based methods. Further, the experimental setup is discussed including an overview of publicly available datasets. Finally, various freely available methods, of which some are considered to be state of the art, are evaluated on two datasets.
Arjan Gijsenij, Theo Gevers, Joost van de Weijer 0001
IEEE Trans. Image Process.2
2011 Empowering Visual Categorization With the GPU
abstract
Visual categorization is important to manage large collections of digital images and video, where textual metadata is often incomplete or simply unavailable. The bag-of-words model has become the most powerful method for visual categorization of images and video. Despite its high accuracy, a severe drawback of this model is its high computational cost. As the trend to increase computational power in newer CPU and GPU architectures is to increase their level of parallelism, exploiting this parallelism becomes an important direction to handle the computational cost of the bag-of-words approach. When optimizing a system based on the bag-of-words approach, the goal is to minimize the time it takes to process batches of images. this paper, we analyze the bag-of-words model for visual categorization in terms of computational cost and identify two major bottlenecks: the quantization step and the classification step. We address these two bottlenecks by proposing two efficient algorithms for quantization and classification by exploiting the GPU hardware and the CUDA parallel programming model. The algorithms are designed to (1) keep categorization accuracy intact, (2) decompose the problem, and (3) give the same numerical results. In the experiments on large scale datasets, it is shown that, by using a parallel implementation on the Geforce GTX260 GPU, classifying unseen images is 4.8 times faster than a quad-core CPU version on the Core i7 920, while giving the exact same numerical results. In addition, we show how the algorithms can be generalized to other applications, such as text retrieval and video retrieval. Moreover, when the obtained speedup is used to process extra video frames in a video retrieval benchmark, the accuracy of visual categorization is improved by 29%.
Koen E. A. van de Sande, Theo Gevers, Cees Snoek
IEEE Trans. Multim.2
2010 3D Scene priors for road detection
abstract
Vision-based road detection is important in different areas of computer vision such as autonomous driving, car collision warning and pedestrian crossing detection. However, current vision-based road detection methods are usually based on low-level features and they assume structured roads, road homogeneity, and uniform lighting conditions. Therefore, in this paper, contextual 3D information is used in addition to low-level cues. Low-level photometric invariant cues are derived from the appearance of roads. Contextual cues used include horizon lines, vanishing points, 3D scene layout and 3D road stages. Moreover, temporal road cues are included. All these cues are sensitive to different imaging conditions and hence are considered as weak cues. Therefore, they are combined to improve the overall performance of the algorithm. To this end, the low-level, contextual and temporal cues are combined in a Bayesian framework to classify road sequences. Large scale experiments on road sequences show that the road detection method is robust to varying imaging conditions, road types, and scenarios (tunnels, urban and highway). Further, using the combined cues outperforms all other individual cues. Finally, the proposed method provides highest road detection accuracy when compared to state-of-the-art methods.
José M. Álvarez 0004, Theo Gevers, Antonio M. López 0001
CVPR2
2010 Visual Gaze Estimation by Joint Head and Eye Information
abstract
In this paper, we present an unconstrained visual gaze estimation system. The proposed method extracts the visual field of view of a person looking at a target scene in order to estimate the approximate location of interest (visual gaze). The novelty of the system is the joint use of head pose and eye location information to fine tune the visual gaze estimated by the head pose only, so that the system can be used in multiple scenarios. The improvements obtained by the proposed approach are validated using the Boston University head pose dataset, on which the standard deviation of the joint visual gaze estimation improved by 61:06% horizontally and 52:23% vertically with respect to the gaze estimation obtained by the head pose only. A user study shows the potential of the proposed system.
Roberto Valenti, Adel Lablack, Nicu Sebe, Chaabane Djeraba, Theo Gevers
ICPR5
2010 The Impact of Color on Bag-of-Words Based Object Recognition
abstract
In recent years several works have aimed at exploiting color information in order to improve the bag-of-words based image representation. There are two stages in which color information can be applied in the bag-of-words framework. Firstly, feature detection can be improved by choosing highly informative color-based regions. Secondly, feature description, typically focusing on shape, can be improved with a color description of the local patches. Although both approaches have been shown to improve results the combined merits have not yet been analyzed. Therefore, in this paper we investigate the combined contribution of color to both the feature detection and extraction stages. Experiments performed on two challenging data sets, namely Flower and Pascal VOC 2009; clearly demonstrate that incorporating color in both feature detection and extraction significantly improves the overall performance.
David Augusto Rojas Vigo, Fahad Shahbaz Khan, Joost van de Weijer 0001, Theo Gevers
ICPR4
2010 Geographic information for vision-based road detection
abstract
Road detection is a vital task for the development of autonomous vehicles. The knowledge of the free road surface ahead of the target vehicle can be used for autonomous driving, road departure warning, as well as to support advanced driver assistance systems like vehicle or pedestrian detection. Using vision to detect the road has several advantages in front of other sensors: richness of features, easy integration, low cost or low power consumption. Common vision-based road detection approaches use low-level features (such as color or texture) as visual cues to group pixels exhibiting similar properties. However, it is difficult to foresee a perfect clustering algorithm since roads are in outdoor scenarios being imaged from a mobile platform. In this paper, we propose a novel high-level approach to vision-based road detection based on geographical information. The key idea of the algorithm is exploiting geographical information to provide a rough detection of the road. Then, this segmentation is refined at low-level using color information to provide the final result. The results presented show the validity of our approach.
José M. Álvarez 0004, Felipe Lumbreras, Theo Gevers, Antonio M. López 0001
Intelligent Vehicles Symposium3
2010 Eyes do not lie: spontaneous versus posed smiles
abstract
Automatic detection of spontaneous versus posed facial expressions received a lot of attention in recent years. However, almost all published work in this area use complex facial features or multiple modalities, such as head pose and body movements with facial features. Besides, the results of these studies are not given on public databases. In this paper, we focus on eyelid movements to classify spontaneous versus posed smiles and propose distance-based and angular features for eyelid movements. We assess the reliability of these features with continuous HMM, k-NN and naive Bayes classifiers on two different public datasets. Experimentation shows that our system provides classification rates up to 91 per cent for posed smiles and up to 80 per cent for spontaneous smiles by using only eyelid movements. We additionally compare the discrimination power of movement features from different facial regions for the same task.
Hamdi Dibeklioglu, Roberto Valenti, Albert Ali Salah, Theo Gevers
ACM Multimedia4
2010 Learning Photometric Invariance for Object Detection
José M. Álvarez 0004, Theo Gevers, Antonio M. López 0001
Int. J. Comput. Vis.2
2010 Generalized Gamut Mapping using Image Derivative Structures for Color Constancy
abstract
The gamut mapping algorithm is one of the most promising methods to achieve computational color constancy. However, so far, gamut mapping algorithms are restricted to the use of pixel values to estimate the illuminant. Therefore, in this paper, gamut mapping is extended to incorporate the statistical nature of images. It is analytically shown that the proposed gamut mapping framework is able to include any linear filter output. The main focus is on the local n -jet describing the derivative structure of an image. It is shown that derivatives have the advantage over pixel values to be invariant to disturbing effects (i.e. deviations of the diagonal model) such as saturated colors and diffuse light. Further, as the n -jet based gamut mapping has the ability to use more information than pixel values alone, the combination of these algorithms are more stable than the regular gamut mapping algorithm. Different methods of combining are proposed. Based on theoretical and experimental results conducted on large scale data sets of hyperspectral, laboratory and real-world scenes, it can be derived that (1) in case of deviations of the diagonal model, the derivative-based approach outperforms the pixel-based gamut mapping, (2) state-of-the-art algorithms are outperformed by the n -jet based gamut mapping, (3) the combination of the different n -jet based gamut mappings provide more stable solutions, and (4) the fusion strategy based on the intersection of feasible sets provides better color constancy results than the union of the feasible sets.
Arjan Gijsenij, Theo Gevers, Joost van de Weijer 0001
Int. J. Comput. Vis.2
2010 Evaluating Color Descriptors for Object and Scene Recognition
abstract
Image category recognition is important to access visual information on the level of objects and scene types. So far, intensity-based descriptors have been widely used for feature extraction at salient points. To increase illumination invariance and discriminative power, color descriptors have been proposed. Because many different descriptors exist, a structured overview is required of color invariant descriptors in the context of image category recognition. Therefore, this paper studies the invariance properties and the distinctiveness of color descriptors (software to compute the color descriptors from this paper is available from http://www.colordescriptors.com) in a structured way. The analytical invariance properties of color descriptors are explored, using a taxonomy based on invariance properties with respect to photometric transformations, and tested experimentally using a data set with known illumination conditions. In addition, the distinctiveness of color descriptors is assessed experimentally using two benchmarks, one from the image domain and one from the video domain. From the theoretical and experimental results, it can be derived that invariance to light intensity changes and light color changes affects category recognition. The results further reveal that, for light intensity shifts, the usefulness of invariance is category-specific. Overall, when choosing a single descriptor and no prior knowledge about the data set and object and scene categories is available, the OpponentSIFT is recommended. Furthermore, a combined set of color descriptors outperforms intensity-based SIFT and improves category recognition by 8 percent on the PASCAL VOC 2007 and by 7 percent on the Mediamill Challenge.
Koen E. A. van de Sande, Theo Gevers, Cees Snoek
IEEE Trans. Pattern Anal. Mach. Intell.2
2009 Learning photometric invariance from diversified color model ensembles
abstract
Color is a powerful visual cue for many computer vision applications such as image segmentation and object recognition. However, most of the existing color models depend on the imaging conditions affecting negatively the performance of the task at hand. Often, a reflection model (e.g., Lambertian or dichromatic reflectance) is used to derive color invariant models. However, those reflection models might be too restricted to model real-world scenes in which different reflectance mechanisms may hold simultaneously. Therefore, in this paper, we aim to derive color invariance by learning from color models to obtain diversified color invariant ensembles. First, a photometrical orthogonal and non-redundant color model set is taken on input composed of both color variants and invariants. Then, the proposed method combines and weights these color models to arrive at a diversified color ensemble yielding a proper balance between invariance (repeatability) and discriminative power (distinctiveness). To achieve this, the fusion method uses a multi-view approach to minimize the estimation error. In this way, the method is robust to data uncertainty and produces properly diversified color invariant ensembles. Experiments are conducted on three different image datasets to validate the method. From the theoretical and experimental results, it is concluded that the method is robust against severe variations in imaging conditions. The method is not restricted to a certain reflection model or parameter tuning. Further, the method outperforms state-of- the-art detection techniques in the field of object, skin and road recognition.
José M. Álvarez 0004, Theo Gevers, Antonio M. López 0001
CVPR2
2009 Physics-based edge evaluation for improved color constancy
abstract
Edge-based color constancy makes use of image derivatives to estimate the illuminant. However, different edge types exist in real-world images such as shadow, geometry, material and highlight edges. These different edge types may have a distinctive influence on the performance of the illuminant estimation.
Arjan Gijsenij, Theo Gevers, Joost van de Weijer 0001
CVPR2
2009 Robustifying eye center localization by head pose cues
abstract
Head pose and eye location estimation are two closely related issues which refer to similar application areas. In recent years, these problems have been studied individually in numerous works in the literature. Previous research shows that cylindrical head models and isophote based schemes provide satisfactory precision in head pose and eye location estimation, respectively. However, the eye locator is not adequate to accurately locate eye in the presence of extreme head poses. Therefore, head pose cues may be suited to enhance the accuracy of eye localization in the presence of severe head poses. In this paper, a hybrid scheme is proposed in which the transformation matrix obtained from the head pose is used to normalize the eye regions and, in turn the transformation matrix generated by the found eye location is used to correct the pose estimation procedure. The scheme is designed to (1) enhance the accuracy of eye location estimations in low resolution videos, (2) to extend the operating range of the eye locator and (3) to improve the accuracy and re-initialization capabilities of the pose tracker. From the experimental results it can be derived that the proposed unified scheme improves the accuracy of eye estimations by 16% to 23%. Further, it considerably extends its operating range by more than 15°, by overcoming the problems introduced by extreme head poses. Finally, the accuracy of the head pose tracker is improved by 12% to 24%.
Roberto Valenti, Zeynep Yücel, Theo Gevers
CVPR3
2009 Color constancy using 3D scene geometry
abstract
The aim of color constancy is to remove the effect of the color of the light source. As color constancy is inherently an ill-posed problem, most of the existing color constancy algorithms are based on specific imaging assumptions such as the grey-world and white patch assumptions.
Arjan Gijsenij, Theo Gevers, Vladimir Nedovic, De Xu, Jan-Mark Geusebroek
ICCV3
2009 Image saliency by isocentric curvedness and color
abstract
In this paper we propose a novel computational method to infer visual saliency in images. The method is based on the idea that salient objects should have local characteristics that are different than the rest of the scene, being edges, color or shape. By using a novel operator, these characteristics are combined to infer global information. The obtained information is used as a weighting for the output of a segmentation algorithm so that the salient object in the scene can easily be distinguished from the background. The proposed approach is fast and it does not require any learning. The experimentation shows that the system can enhance interesting objects in images and it is able to correctly locate the same object annotated by humans with an F-measure of 85.61% when the object size is known, and 79.19% when the object size is unknown, improving the state of the art performance on a public dataset.
Roberto Valenti, Nicu Sebe, Theo Gevers
ICCV3
2009 Vision-based road detection using road models
abstract
Vision-based road detection is very challenging since the road is in an outdoor scenario imaged from a mobile platform. In this paper, a new top-down road detection algorithm is proposed. The method is based on scene (road) classification which provides the probability that an image contains certain type of road geometry (straight, left/right curve, etc.). During the training of the classifier a road probability map is also learned for each road geometry. Then, the proper pixel-based method is selected and fused to provide an improved road detection approach. From experiments it is concluded that the proposed method outperforms state-of-the-art algorithms in a frame by frame context.
José M. Álvarez 0004, Theo Gevers, Antonio M. López 0001
ICIP2
2009 Shadow edge detection using geometric and photometric features
abstract
The detection of shadow and shading edges is a first step towards reducing the imaging effects that are caused by interactions of the light source with surfaces that are in the scene. As most of the algorithms for shadow edge detection use photometric information, geometric information have been ignored so far. In this paper, the aim is to include geometric features for more robust shadow edge detection. First, thousands of patches are annotated as either containing a shadow edge or not. Then, geometric features of these patches are analyzed and it is shown that the combination of photometric and geometric features improves the classification of shadow edges with respect to using either one of these features with 14%. These results demonstrate the added value of geometric features, in addition to photometric features, for the detection of shadow edges.
Arjan Gijsenij, Theo Gevers
ICIP2
2009 Color constancy using stage classification
abstract
The aim of color constancy is to remove the effect of the color of the light source. Since color constancy is inherently an ill-posed problem, different assumptions have been proposed. Because existing color constancy algorithms are based on specific assumptions, none of them can be considered as universal. Therefore, how to select a proper algorithm for a given imaging configuration is an important question.
Arjan Gijsenij, Theo Gevers, Koen E. A. van de Sande, Jan-Mark Geusebroek, De Xu
ICIP3
2009 Isocentric color saliency in images
abstract
In this paper we propose a novel computational method to infer visual saliency in images. The computational method is based on the idea that salient objects should have local characteristics that are different than the rest of the scene, being edges, color or shape, and that these characteristics can be combined to infer global information. The proposed approach is fast, does not require any learning and the experimentation shows that it can enhance interesting objects in images, improving the state of the art performance on a public dataset.
Roberto Valenti, Nicu Sebe, Theo Gevers
ICIP3
2008 Evaluation of color descriptors for object and scene recognition
abstract
Image category recognition is important to access visual information on the level of objects and scene types. So far, intensity-based descriptors have been widely used. To increase illumination invariance and discriminative power, color descriptors have been proposed only recently. As many descriptors exist, a structured overview of color invariant descriptors in the context of image category recognition is required.
Koen E. A. van de Sande, Theo Gevers, Cees Snoek
CVPR2
2008 Accurate eye center location and tracking using isophote curvature
abstract
The ubiquitous application of eye tracking is precluded by the requirement of dedicated and expensive hardware, such as infrared high definition cameras. Therefore, systems based solely on appearance (i.e. not involving active infrared illumination) are being proposed in literature. However, although these systems are able to successfully locate eyes, their accuracy is significantly lower than commercial eye tracking devices. Our aim is to perform very accurate eye center location and tracking, using a simple Web cam. By means of a novel relevance mechanism, the proposed method makes use of isophote properties to gain invariance to linear lighting changes (contrast and brightness), to achieve rotational invariance and to keep low computational costs. In this paper we test our approach for accurate eye location and robustness to changes in illumination and pose, using the BioIDand the Yale Face B databases, respectively. We demonstrate that our system can achieve a considerable improvement in accuracy over state of the art techniques.
Roberto Valenti, Theo Gevers
CVPR2
2008 A Perceptual Comparison of Distance Measures for Color Constancy Algorithms
Arjan Gijsenij, Theo Gevers, Marcel P. Lucassen
ECCV (1)2
2007 Color Constancy using Natural Image Statistics
abstract
Although many color constancy methods exist, they are all based on specific assumptions such as the set of possible light sources, or the spatial and spectral characteristics of images. As a consequence, no algorithm can be considered as universal. However, with the large variety of available methods, the question is how to select the method that induces equivalent classes for different image characteristics. Furthermore, the subsequent question is how to combine the different algorithms in a proper way. To achieve selection and combining of color constancy algorithms, in this paper, natural image statistics are used to identify the most important characteristics of color images. Then, based on these image characteristics, the proper color constancy algorithm (or best combination of algorithms) is selected for a specific image. To capture the image characteristics, the Weibull parameterization (e.g. texture and contrast) is used. Experiments show that, on a large data set of 11,000 images, our approach outperforms current state-of-the-art single algorithms, as well as simple alternatives for combining several algorithms.
Arjan Gijsenij, Theo Gevers
CVPR2
2007 Color Constancy using Image Regions
abstract
Color constancy is important for various applications such as image segmentation, object recognition and image retrieval where object color features are extracted invariant to the illumination conditions. Different color constancy methods have been proposed. These methods, in general, compute color constancy based on all image colors. However, not all pixels contain relevant information for color constancy. Eventually, biased pixel values may decrease the performance of color constancy methods. To this end, in this paper, we propose a method based on low-level image features using subsets of pixels. Hence, instead of using the entire pixel set for estimating the illuminant, only relevant pixels in the image are used. Therefore, prior segmentation is performed to learn for different image categories (e.g. open country, street, indoor) which pixel set (i.e. image parts) is most appropriate for a reliable estimation. Based on large scale experiments on real-world scenes, it can be derived that for certain categories, like open country and street, the estimation is far more accurate using image parts than when using the entire image.
Arjan Gijsenij, Theo Gevers
ICIP (3)2
2007 Do Colour Interest Points Improve Image Retrieval?
abstract
In image retrieval scenarios, many methods use interest point detection at an early stage to find regions in which descriptors are calculated. Finding salient locations in image data is crucial for these tasks. Observing that most current methods use only the luminance information of the images, we investigate the use of colour information in interest point detection. A way to use multi-channel information in the Harris corner detector is explored and different colour spaces are evaluated. To determine the characteristic scale of an interest point, a new colour scale selection method is presented. We show that using colour information and boosting salient colours results in improved performance in retrieval tasks.
Julian Stöttinger, Allan Hanbury, Nicu Sebe, Theo Gevers
ICIP (1)4
2007 Authentic facial expression analysis
Nicu Sebe, Michael S. Lew, Yafei Sun, Ira Cohen, Theo Gevers, Thomas S. Huang
Image Vis. Comput.5
2007 Selection and Fusion of Color Models for Image Feature Detection
abstract
The choice of a color model is of great importance for many computer vision algorithms (e.g., feature detection, object recognition, and tracking) as the chosen color model induces the equivalence classes to the actual algorithms. As there are many color models available, the inherent difficulty is how to automatically select a single color model or, alternatively, a weighted subset of color models producing the best result for a particular task. The subsequent hurdle is how to obtain a proper fusion scheme for the algorithms so that the results are combined in an optimal setting. To achieve proper color model selection and fusion of feature detection algorithms, in this paper, we propose a method that exploits nonperfect correlation between color models or feature detection algorithms derived from the principles of diversification. As a consequence, a proper balance is obtained between repeatability and distinctiveness. The result is a weighting scheme which yields maximal feature discrimination. The method is verified experimentally for three different image feature detectors. The experimental results show that the fusion method provides feature detection results having a higher discriminative power than the standard weighting scheme. Further, it is experimentally shown that the color model selection scheme provides a proper balance between color invariance (repeatability) and discriminative power (distinctiveness).
Harro M. G. Stokman, Theo Gevers
IEEE Trans. Pattern Anal. Mach. Intell.2
2007 Edge-Based Color Constancy
abstract
Color constancy is the ability to measure colors of objects independent of the color of the light source. A well-known color constancy method is based on the gray-world assumption which assumes that the average reflectance of surfaces in the world is achromatic. In this paper, we propose a new hypothesis for color constancy namely the gray-edge hypothesis, which assumes that the average edge difference in a scene is achromatic. Based on this hypothesis, we propose an algorithm for color constancy. Contrary to existing color constancy algorithms, which are computed from the zero-order structure of images, our method is based on the derivative structure of images. Furthermore, we propose a framework which unifies a variety of known (gray-world, max-RGB, Minkowski norm) and the newly proposed gray-edge and higher order gray-edge algorithms. The quality of the various instantiations of the framework is tested and compared to the state-of-the-art color constancy methods on two large data sets of images recording objects under a large number of different light sources. The experiments show that the proposed color constancy algorithms obtain comparable results as the state-of-the-art color constancy methods with the merit of being computationally more efficient.
Joost van de Weijer 0001, Theo Gevers, Arjan Gijsenij
IEEE Trans. Image Process.2
2007 A Spatially Constrained Generative Model and an EM Algorithm for Image Segmentation
abstract
In this paper, we present a novel spatially constrained generative model and an expectation-maximization (EM) algorithm for model-based image segmentation. The generative model assumes that the unobserved class labels of neighboring pixels in the image are generated by prior distributions with similar parameters, where similarity is defined by entropic quantities relating to the neighboring priors. In order to estimate model parameters from observations, we derive a spatially constrained EM algorithm that iteratively maximizes a lower bound on the data log-likelihood, where the penalty term is data-dependent. Our algorithm is very easy to implement and is similar to the standard EM algorithm for Gaussian mixtures with the main difference that the labels posteriors are "smoothed" over pixels between each E- and M-step by a standard image filter. Experiments on synthetic and real images show that our algorithm achieves competitive segmentation results compared to other Markov-based methods, and is in general faster.
Aristeidis Diplaros, Nikos Vlassis, Theo Gevers
IEEE Trans. Neural Networks3
2006 Boosting Color Saliency in Image Feature Detection
abstract
The aim of salient feature detection is to find distinctive local events in images. Salient features are generally determined from the local differential structure of images. They focus on the shape-saliency of the local neighborhood. The majority of these detectors are luminance-based, which has the disadvantage that the distinctiveness of the local color information is completely ignored in determining salient image features. To fully exploit the possibilities of salient point detection in color images, color distinctiveness should be taken into account in addition to shape distinctiveness. In this paper, color distinctiveness is explicitly incorporated into the design of saliency detection. The algorithm, called color saliency boosting, is based on an analysis of the statistics of color image derivatives. Color saliency boosting is designed as a generic method easily adaptable to existing feature detectors. Results show that substantial improvements in information content are acquired by targeting color salient features.
Joost van de Weijer 0001, Theo Gevers, Andrew D. Bagdanov
IEEE Trans. Pattern Anal. Mach. Intell.2
2006 Combining color and shape information for illumination-viewpoint invariant object recognition
abstract
In this paper, we propose a new scheme that merges color- and shape-invariant information for object recognition. To obtain robustness against photometric changes, color-invariant derivatives are computed first. Color invariance is an important aspect of any object recognition scheme, as color changes considerably with the variation in illumination, object pose, and camera viewpoint. These color invariant derivatives are then used to obtain similarity invariant shape descriptors. Shape invariance is equally important as, under a change in camera viewpoint and object pose, the shape of a rigid object undergoes a perspective projection on the image plane. Then, the color and shape invariants are combined in a multidimensional color-shape context which is subsequently used as an index. As the indexing scheme makes use of a color-shape invariant context, it provides a high-discriminative information cue robust against varying imaging conditions. The matching function of the color-shape context allows for fast recognition, even in the presence of object occlusion and cluttering. From the experimental results, it is shown that the method recognizes rigid objects with high accuracy in 3-D complex scenes and is robust against changing illumination, camera viewpoint, object pose, and noise.
Aristeidis Diplaros, Theo Gevers, Ioannis Patras
IEEE Trans. Image Process.2
2006 Robust photometric invariant features from the color tensor
abstract
Luminance-based features are widely used as low-level input for computer vision applications, even when color data is available. The extension of feature detection to the color domain prevents information loss due to isoluminance and allows us to exploit the photometric information. To fully exploit the extra information in the color data, the vector nature of color data has to be taken into account and a sound framework is needed to combine feature and photometric invariance theory. In this paper, we focus on the structure tensor, or color tensor, which adequately handles the vector nature of color images. Further, we combine the features based on the color tensor with photometric invariant derivatives to arrive at photometric invariant features. We circumvent the drawback of unstable photometric invariants by deriving an uncertainty measure to accompany the photometric invariant derivatives. The uncertainty is incorporated in the color tensor, hereby allowing the computation of robust photometric invariant features. The combination of the photometric invariance theory and tensor-based features allows for detection of a variety of features such as photometric invariant edges, corners, optical flow, and curvature. The proposed features are tested for noise characteristics and robustness to photometric changes. Experiments show that the proposed features are robust to scene incidental events and that the proposed uncertainty measure improves the applicability of full invariants.
Joost van de Weijer 0001, Theo Gevers, Arnold W. M. Smeulders
IEEE Trans. Image Process.2
2005 Selection and Fusion of Color Models for Feature Detection
abstract
The choice of a color space is of great importance for many computer vision algorithms (e.g. edge detection and object recognition). It induces the equivalence classes to the actual algorithms. However, the problem is how to automatically select the color space that produces the best result for a particular task. The subsequent difficulty then is how to obtain a proper weighting scheme for the algorithms so that the results are combined in an optimal setting. To achieve proper color space selection and fusion of feature detectors, in this paper, we propose a method that exploits non-perfect correlation between the color models derived from the principles of diversification. As a consequence, the weighting scheme yields maximal color discrimination. The method is verified experimentally for two different feature detectors. The experimental results show that the model provides feature detection results having a discriminative power of 30 percent higher than the standard weighting scheme.
Harro M. G. Stokman, Theo Gevers
CVPR (1)2
2005 Boosting Saliency in Color Image Features
abstract
The aim of salient point detection is to find distinctive events in images. Salient features are generally determined from the local differential structure of images. They focus on the shape saliency of the local neighborhood. The majority of these detectors is luminance based which has the disadvantage that the distinctiveness of the local color information is completely ignored. To fully exploit the possibilities of color image salient point detection, color distinctiveness should be taken into account next to shape distinctiveness. In this paper color distinctiveness is explicitly incorporated into the design of saliency detection. The algorithm, called color saliency boosting, is based on an analysis of the statistics of color image derivatives. Isosalient color derivatives can be closely approximated by ellipsoidal surfaces in color derivative space. Based on this remarkable statistical finding, isosalient derivatives are transformed by color boosting to have equal impact on the saliency. Color saliency boosting is designed as a generic method easily adaptable to existing feature detectors. Results show that substantial improvements in information content are acquired by targeting color salient features. Further, the generality of the method is illustrated by applying color boosting to multiple existing saliency methods.
Joost van de Weijer 0001, Theo Gevers
CVPR (1)2
2005 Color feature detection and classification by learning
abstract
In this paper, we aim at the classification of the physical nature of local image structures in color images on the basis of geometrical and photometrical information. To this end, a framework is proposed to combine the local differential structure (i.e. geometrical information such as edges, corners, T-junctions etc) and color (i.e. photometrical information such as shadows, shading, illumination, highlights) in a multi-dimensional feature space. This framework is used to yield a proper classifier to classify salient image structures on the basis of their physical nature. The proposed framework is empirically verified on a set of images. From the theoretical and experimental results it is concluded that the proposed classification scheme successfully classifies local image structures robust to image translation and rotation, illumination intensity variations, and noise.
Theo Gevers, Simon Voortman, Frank Aldershoff
ICIP (2)1
2005 Color constancy based on the Grey-edge hypothesis
abstract
A well-known color constancy method is based on the Grey-World assumption i.e. the average reflectance of surfaces in the world is achromatic. In this article we propose a new hypothesis for color constancy, namely the Grey-Edge hypothesis assuming that the average edge difference in a scene is achromatic. Based on this hypothesis, we propose an algorithm for color constancy. Recently, the Grey-World hypothesis and the max-RGB method were shown to be two instantiations of a Minkowski norm based color constancy method. Similarly we also propose a more general version of the Grey-Edge hypothesis which assumes that the Minkowsky norm of derivatives of the reflectance of surfaces is achromatic. The algorithms are tested on a large data set of images under different illuminants, and the results show that the new method outperforms the Grey-World assumption and the max-RGB method. Results are comparable to more elaborate algorithms, however at lower computational costs.
Joost van de Weijer 0001, Theo Gevers
ICIP (2)2
2005 Learning probabilistic classifiers for human-computer interaction applications
Nicu Sebe, Ira Cohen, Fábio G. Cozman, Theo Gevers, Thomas S. Huang
Multim. Syst.4
2005 Edge and Corner Detection by Photometric Quasi-Invariants
abstract
Feature detection is used in many computer vision applications such as image segmentation, object recognition, and image retrieval. For these applications, robustness with respect to shadows, shading, and specularities is desired. Features based on derivatives of photometric invariants, which we will call full invariants, provide the desired robustness. However, because computation of photometric invariants involves nonlinear transformations, these features are Instable and, therefore, impractical for many applications. We propose a new class of derivatives which we refer to as quasi-invariants. These quasi-invariants are derivatives which share with full photometric invariants the property that they are insensitive for certain photometric edges, such as shadows or specular edges, but without the inherent instabilities of full photometric invariants. Experiments show that the quasi-invariant derivatives are less sensitive to noise and introduce less edge displacement than full invariant derivatives. Moreover, quasi-invariants significantly outperform the full invariant derivatives in terms of discriminative power.
Joost van de Weijer 0001, Theo Gevers, Jan-Mark Geusebroek
IEEE Trans. Pattern Anal. Mach. Intell.2
2004 Color invariant density estimation for image segmentation and object tracking
abstract
In this paper, we formulate a novel density estimation scheme derived from color invariants for image segmentation and object tracking. The advantage of color invariants is that they are robust against varying illumination. However, color invariants are ill-defined when the intensity or saturation is low. Therefore, to achieve robust density estimation, computational methods are presented to estimate the amount of sensor noise through these color invariant images. The obtained uncertainty is subsequently used as a weighting term in the density estimation process to achieve robust image segmentation and object tracking. Experiments are conducted on image sequences recorded from complex 3D scenes. From the experimental results it is shown that the proposed method successfully segments and finds objects robust against illumination and noisy data.
Theo Gevers, Frank Aldershoff
ICIP1
2004 Robust optical flow from photometric invariants
Joost van de Weijer 0001, Theo Gevers
ICIP2
2004 Color for Image Indexing and Retrieval
Theo Gevers, Graham D. Finlayson, Raimondo Schettini
Comput. Vis. Image Underst.1
2004 Guest Editorial
Arnold W. M. Smeulders, Thomas S. Huang, Theo Gevers
Int. J. Comput. Vis.3
2004 Robust Histogram Construction from Color Invariants for Object Recognition
Theo Gevers, Harro M. G. Stokman
IEEE Trans. Pattern Anal. Mach. Intell.1
2004 Robust segmentation and tracking of colored objects in video
abstract
Segmenting and tracking of objects in video is of great importance for video-based encoding, surveillance, and retrieval. However, the inherent difficulty of object segmentation and tracking is to distinguish changes in the displacement of objects from disturbing effects such as noise and illumination changes. Therefore, in this paper, we formulate a color-based deformable model which is robust against noisy data and changing illumination. Computational methods are presented to measure color constant gradients. Further, a model is given to estimate the amount of sensor noise through these color constant gradients. The obtained uncertainty is subsequently used as a weighting term in the deformation process. Experiments are conducted on image sequences recorded from three-dimensional scenes. From the experimental results, it is shown that the proposed color constant deformable method successfully finds object contours robust against illumination, and noisy, but homogeneous regions.
Theo Gevers
IEEE Trans. Circuits Syst. Video Technol.1
2003 Reflectance-based Classification of Color Edges
abstract
We aim at using color information to classify the physical nature of edges in video. To achieve physics-based edge classification, we first propose a novel approach to color edge detection by automatic noise-adaptive thresholding derived from sensor noise analysis. Then, we present a taxonomy on color edge types. As a result, a parameter-free edge classifier is obtained by labeling color transitions into one of the following types: (1) shadow-geometry, (2) highlight edges, (3) material edges. The proposed method is empirically verified on images showing complex real world scenes.
Theo Gevers
ICCV1
2003 Color Edge Detection by Photometric Quasi-Invariants
abstract
Photometric invariance is used in many computer vision applications. The advantage of photometric invariance is the robustness against shadows, shading and illumination conditions. However, the drawbacks of photometric invariance are the loss of discriminative power and the inherent instabilities caused by the nonlinear transformations to compute the invariants. In this paper, we propose a new class of derivatives which we refer to as photometric quasi-invariants. These quasi-invariants share with full invariants the nice property that they are robust against photometric edges, such as shadows or specular edges. Further, these quasi-invariants do not have the inherent instabilities of full photometric invariants. We will apply these quasi-invariant derivatives in the context of photometric invariant edge detection and classification. Experiments show that the quasi-invariant derivatives are stable and they significantly outperform the full invariant derivatives in discriminative power.
Joost van de Weijer 0001, Theo Gevers, Jan-Mark Geusebroek
ICCV2
2003 Classifying multimedia documents by merging textual and pictorial information
abstract
In this paper, we study computational models and techniques to merge textual and image features to classify multimedia documents into semantically meaningful groups. A vector-based framework is used to index documents on the basis of textual, pictorial and composite (textual-pictorial) information. The scheme makes use of weighted document terms and color invariant image features to obtain a high-dimensional image descriptor in vector form to be used as an index. Based on supervised learning, a classifier is used to organize the multimedia documents. Due to space limitations, in this paper, we focus on the application of classifying/finding pictures of people on the Internet. Performance evaluations are reported on the accuracy of merging textual and pictorial information for classification.
Theo Gevers, Frank Aldershoff
ICIP (3)1
2003 Robust Photometric Invariant Region Detection in Multispectral Images
Theo Gevers, Harro M. G. Stokman
Int. J. Comput. Vis.1
2003 Color constancy from physical principles
Jan-Mark Geusebroek, Rein van den Boomgaard, Arnold W. M. Smeulders, Theo Gevers
Pattern Recognit. Lett.4
2003 Classifying color edges in video into shadow-geometry, highlight, or material transitions
abstract
We aim at using color information to classify the physical nature of edges in video. To achieve physics-based edge classification, we first propose a novel approach to color edge detection by automatic noise-adaptive thresholding derived from sensor noise analysis. Then, we present a taxonomy on color edge types. As a result, a parameter-free edge classifier is obtained labeling color transitions into one of the following types: 1) shadow-geometry, 2) highlight edges, and 3) material edges. The proposed method is empirically verified on images showing complex real world scenes.
Theo Gevers, Harro M. G. Stokman
IEEE Trans. Multim.1
2002 Adaptive Image Segmentation by Combining Photometric Invariant Region and Edge Information
abstract
An adaptive image segmentation scheme is proposed employing the Delaunay triangulation for image splitting. The tessellation grid of the Delaunay triangulation is adapted to the semantics of the image data by combining region and edge information. To achieve robustness against imaging conditions (e.g. shading, shadows, illumination and highlights), photometric invariant similarity measures and edge computation are proposed. Experimental results on synthetic and real images show that the segmentation method is robust to edge orientation, partially weak object boundaries and noisy-but-homogeneous regions. Furthermore, the method is robust, to a large degree, to varying imaging conditions.
Theo Gevers
IEEE Trans. Pattern Anal. Mach. Intell.1
2002 Image segmentation and similarity of color-texture objects
abstract
We aim for content-based image retrieval of textured objects in natural scenes under varying illumination and viewing conditions. To achieve this, image retrieval is based on matching feature distributions derived from color invariant gradients. To cope with object cluttering, region-based texture segmentation is applied on the target images prior to the actual image retrieval process. The retrieval scheme is empirically verified on color images taken from textured objects under different lighting conditions.
Theo Gevers
IEEE Trans. Multim.1
2001 Color Constant Ratio Gradients for Image Segmentation and Similarity of Texture Objects
abstract
We aim for content-based image retrieval of texture objects in natural scenes under varying illumination and viewing conditions. To achieve this, image retrieval is based on matching feature distributions derived from color invariant gradients. To cope with object cluttering, region-based texture segmentation is applied on the target images prior to the actual image retrieval process. The retrieval scheme is empirically verified on color images taken from texture objects under different lighting, conditions.
Theo Gevers, Arnold W. M. Smeulders
CVPR (1)1
2001 Robust Histogram Construction from Color Invariants
Theo Gevers
ICCV1
2001 Invariant representation in image processing
abstract
The paper discusses the role of invariance in image processing, specifically the desire to discriminate against unwanted variations in the scene while maintaining the power to tell the difference between object-intrinsic characteristics and scene-accidental conditions. It provides an analysis and references of what are directly observables in a general scene.
Arnold W. M. Smeulders, Jan-Mark Geusebroek, Theo Gevers
ICIP (3)3
2001 Color mode filtering
abstract
In this paper mode filtering of color images is explored. An existing framework based on local histograms is extended to multi-channel images. Within this framework three color mode operations are proposed; 1. Global mode operation for edge sharpening, noise reduction and small object removal, 2. Constrained mode operation for white noise filtering while preserving detail, and 3. Uncertain data mode filtering to incorporate prior knowledge about the certainty of the measurements into the mode computation. Results obtained for a variety of images indicate the feasibility of color mode filtering.
Joost van de Weijer 0001, Theo Gevers
ICIP (1)2
2000 Colour Constancy from Hyper-Spectral Data
abstract
This paper aims for color constant identification of object colors through the analysis of spectral color data. New computational color models are proposed which are not only invariant to illumination variations (color constancy) but also robust to a change in viewpoint and object geometry (color invariance). Color constancy and invariance is achieved by spectral imaging using a white reference, and based on color ratio’s (without a white reference). From the theoretical and experimental results it is concluded that the proposed computational methods for color constancy and invariance are highly robust to a change in SPD of the light source as well as a change in the pose of the object.
Theo Gevers, Harro M. G. Stokman, Joost van de Weijer 0001
BMVC1
2000 Image Retrieval and Segmentation based on Color Invariants
abstract
We will demonstrate our CVPR2000 paper "Measurement of Color Invariants" for the cases of image retrieval based on query by example and for color image segmentation. Both are of importance in content based access of image and video data. We demonstrate the usefulness of the proposed color invariants in image retrieval by example systems. We show that an image retrieval query should include the type of invariance expected in the result. We demonstrate such queries by using the "ImageSurf" retrieval system. Segmentation of images based on the proposed color invariants is demonstrated by the "PicToVision" system. The system provides image processing functionality through the world wide web, and is publicly accessible at www.science. uva.nl/research/isis.pictovision.html.
Jan-Mark Geusebroek, Dennis C. Koelma, Arnold W. M. Smeulders, Theo Gevers
CVPR4
2000 Classifying Color Transitions into Shadow-Geometry, Illumination, Highlight or Material Edges
abstract
We aim at using color information to classify the physical nature of a color edge: that is whether the transition is due to shadows, abrupt surface orientation changes, illumination, highlights or material changes. To achieve a physics-based edge classification, we propose a taxonomy of color invariant edges. The taxonomy is based upon the sensitivity of the various color edges with respect to different imaging dependencies i.e. shadows, object shape, shading (i.e. illumination intensity changes), highlights and material characteristics. From this taxonomy, the edge classifier is derived labeling color transitions into the following types: (1) shadow, geometry or shading edges, (2) highlight edges, (3) material edges. Experiments conducted with the edge classification technique on color and hyperspectral images show that the proposed method successfully discriminates the different edge types.
Theo Gevers, Harro M. G. Stokman
ICIP1
2000 Color Measurement by Imaging Spectrometry
Harro M. G. Stokman, Theo Gevers, Jan J. Koenderink
Comput. Vis. Image Underst.2
2000 PicToSeek: combining color and shape invariant features for image retrieval
abstract
We aim at combining color and shape invariants for indexing and retrieving images. To this end, color models are proposed independent of the object geometry, object pose, and illumination. From these color models, color invariant edges are derived from which shape invariant features are computed. Computational methods are described to combine the color and shape invariants into a unified high-dimensional invariant feature set for discriminatory object retrieval. Experiments have been conducted on a database consisting of 500 images taken from multicolored man-made objects in real world scenes. From the theoretical and experimental results it is concluded that object retrieval based on composite color and shape invariant features provides excellent retrieval accuracy. Object retrieval based on color invariants provides very high retrieval accuracy whereas object retrieval based entirely on shape invariants yields poor discriminative power. Furthermore, the image retrieval scheme is highly robust to partial occlusion, object clutter and a change in the object's pose. Finally, the image retrieval scheme is integrated into the PicToSeek system on-line at http://www.wins.uva.nl/research/isis/PicToSeek/ for searching images on the World Wide Web.
Theo Gevers, Arnold W. M. Smeulders
IEEE Trans. Image Process.1
1999 Detection and Classification of Hyper-Spectral Edges
abstract
Intensity-based edge detectors cannot distinguish whether an edge is caused by material changes, shadows, surface orientation changes or by highlights. Therefore, our aim is to classify the physical cause of an edge using hyperspectra obtained by a spectrograph. Methods are presented to detect edges in hyperspectral images. In theory, the effect of varying imaging conditions is analyzed for "raw" hyper-spectra, for normalized hyper-spectra, and for hue computed from hyper-spectra. From this analysis, an edge classifier is derived which distinguishes hyper-spectral edges into the following types: (1) a shadow or geometry edge, (2) a highlight edge, (3) a material edge. 1 Introduction Edge information from an image can be used to measure or recognize objects in images. Edges correspond to significant changes in the image, ideally at the boundary between two different regions. However, false edges are often detected, and (parts of) important edges are missing. Thus, after edge d...
Harro M. G. Stokman, Theo Gevers
BMVC2
1999 Content-based image retrieval by viewpoint-invariant color indexing
Theo Gevers, Arnold W. M. Smeulders
Image Vis. Comput.1
1999 Color-based object recognition
Theo Gevers, Arnold W. M. Smeulders
Pattern Recognit.1
1998 Color Invariant Snakes
abstract
Snakes provide high-level information in the form of continuity constraints and minimum energy constraints related to the contour shape and image features. These image features are usually based on intensity edges. However, intensity edges may appear in the scene without a material/color transition to support it. As a consequence, when using intensity edges as image features, the image segmentation results obtained by snakes may be negatively affected by the imaging-process (e.g. shadows, shading and highlights). In this paper, we aim at using color invariant gradient information to guide the deformation process to obtain snake boundaries which correspond to material boundaries in images discounting the disturbing influences of surface orientation, illumination, shadows and highlights. Experiments conducted on various color images show that the proposed color invariant snake successfully find material contours discounting other "accidental" edges types (e.g. shadows, shading and highlight transitions). Comparison with intensity-based snakes shows that the intensity-based snake is dramatically outperformed by the presented color invariant snake. 1
Theo Gevers, Sennay Ghebreab, Arnold W. M. Smeulders
BMVC1
1998 Photometric Invariant Region Detection
abstract
In this paper, we concentrate on determining homogeneously colored regions invariant to surface orientation change, illumination, shadows and highlights. To this end, the influence of various well-known color models (e.g. , , , , , , , and ) are examined, in theory, for the dichromatic reflection model and, in practice, for two distinct region-based segmentation methods: the k-means clustering technique and the split&merge algorithm. Experiments are conducted on color images taken from colored objects in real-world scenes. On the basis of the theoretical and experimental results it is concluded that , , , , and all detect regions invariant to a change in surface orientation, viewpoint of the camera, and illumination intensity. Furthermore, and also detect regions independent of highlights. , , , , ,a nd provide segmentation results which are all sensitive to surface orientation and illumination intensity as well as color models incorporating brightness into their systems: in , in ,a nd in .
Theo Gevers, Arnold W. M. Smeulders, Harro M. G. Stokman
BMVC1
1998 Image Indexing using Composite Color and Shape Invariant Features
abstract
New sets of color models are proposed for object recognition invariant to a change in view point, object geometry and illumination. Further, computational methods are presented to combine color and shape invariants to produce a high-dimensional invariant feature set for discriminatory object recognition. Experiments on a database of 500 images show that object recognition based on composite color and shape invariant features provides excellent recognition accuracy. Furthermore, object recognition based on color invariants provides very high recognition accuracy whereas object recognition based entirely on shape invariants yields very poor discriminative power. The image database and the performance of the recognition scheme can be experienced within PicToSeek: on-line as part of the ZOMAX system at: http://www.wins.uva.nl/research/isis/zomax/.
Theo Gevers, Arnold W. M. Smeulders
ICCV1
1997 Combining Region Splitting and Edge Detection through Guided Delaunay Image Subdivision
abstract
In this paper, an adaptive split-and-merge segmentation method is proposed. The splitting phase of the algorithm employs the incremental Delaunay triangulation competent of forming grid edges of arbitrary orientation, and position. The tessellation grid, defined by the Delaunay triangulation, is adjusted to the semantics of the image data by combining similarity and difference information among pixels. Experimental results on synthetic images show that the method is robust to different object edge orientations, partially weak object edges and very noisy homogeneous regions. Experiments on a real image indicate that the method yields good segmentation results even when there is a quadratic sloping of intensities particularly suited for segmenting natural scenes of man-made objects.
Theo Gevers, Arnold W. M. Smeulders
CVPR1
1996 Color-metric pattern-card matching for viewpoint invariant image retrieval
abstract
In this paper, viewpoint independent image retrieval by color-metric pattern-card matching is presented. First, a photometric color invariant is proposed measuring, a local color property of a pixel and its neighboring pixels while discounting the disturbing influences of shading, shadows and highlights. Color-metric pattern-cards are constructed on the basis of the photometric color invariant indicating whether a particular discrete photometric color invariant value is present in an image. To express similarity between color-metric pattern-cards, similarity functions are proposed and evaluated on a database of 500 images taken from 2-D and 3-D colored man-made objects in real world 3-D scenes. The experimental results show that high image retrieval accuracy is achieved by two distinct similarity functions depending on the presence of object clutter in the scene. Furthermore, image retrieval by color-metric pattern-card matching is to a large degree robust to partial occlusion and a change in viewing position. Good run-time performance of the pattern-card matching process is achieved allowing for fast image retrieval by example image.
Theo Gevers, Arnold W. M. Smeulders
ICPR1
1994 Image segmentation by directed region subdivision
abstract
In this paper, an image segmentation method based on directed image region partitioning is proposed. The method consists of two separate stages: a splitting phase followed by a merging phase. The splitting phase starts with an initial coarse triangulation and employs the incremental Delaunay triangulation as a directed image region splitting technique. The triangulation process is accomplished by adding points as vertices one by one into the triangulation. A top-down point selection strategy is proposed for selecting these points in the image domain of grey-value and color images. The merging phase coalesces the oversegmentation, generated by the splitting phase, into homogeneous image regions. Because images might be negatively affected by changes in intensity due to shading or surface orientation change, the authors propose homogeneity criteria which are robust to intensity changes caused by these phenomena for both grey-value and color images. Performance of the image segmentation method has been evaluated by experiments on test images.
Theo Gevers, V. K. Kajcovski
ICPR (1)1
1993 An Approach to Image Retrieval for Image Databases
Theo Gevers, Arnold W. M. Smeulders
DEXA1
1992 Σnigma: an image retrieval system
abstract
Presents a system which retrieves images on the basis of automatically generated indexes (i.e. semantic image representations, obtained by automatic image analysis, indicating the content of the images). The system consists of two parts: an off-line indexing part and an on-line image retrieval part. The indexing component is used to automatically generate semantic representations of images so that the image retrieval component can use this information to enable image retrieval. The man-machine communication of the image retrieval component is based on an iconical graphical query language to accomplish geographical query specification for image access. Experiments have been carried out on three different sets of images from the following domains: MRI images of the chest, electronic schemas and topographic maps. The experiments show encouraging results especially for domains which have a high degree of formality in their pictorial expression, such as electronic schemas and topographic maps, and to a less extent for domains having a weak degree of formality such as MRI images of the chest.>
Theo Gevers, Arnold W. M. Smeulders
ICPR (2)1