Sezer Karaoglu

dblp:75/9557 · DBLP profile ↗
← Back
38ranked-venue papers
6as first author
24since 2021 · last 2025
0000-0001-9073-9420ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 25 · 18 since 2021Graphics, computer vision, multimedia, augmented reality and games · 23 · 6 first-author · 11 since 2021
YearPublicationVenuePosition
2025 LumiNet: Latent Intrinsics Meets Diffusion Models for Indoor Scene Relighting
abstract
We introduce LumiNet, a novel architecture that leverages generative models and latent intrinsic representations for transferring lighting from one image to another. Given a source image and a target lighting image, LumiNet generates a relit version of the source scene that captures the target’s lighting. Our approach makes two key contributions: a data curation strategy from the StyleGAN-based relighting model for our training, and a modified diffusion-based Con-trolNet that processes both latent intrinsic properties from the source image and latent extrinsic properties from the target image. We further improve lighting transfer through a learned adaptor that injects the target’s latent extrinsic properties via cross-attention and light-weight fine-tuning.Unlike traditional ControlNet, which generates images with conditional maps from a single scene, LumiNet processes latent representations from two different images -preserving geometry and albedo from the source while transferring lighting characteristics from the target. Experiments demonstrate that our method successfully transfers complex lighting phenomena including specular highlights and indirect illumination across scenes with varying spatial layouts and materials, outperforming existing approaches on challenging indoor scenes using only images as input.
Xiaoyan Xing, Konrad Groh, Sezer Karaoglu, Theo Gevers, Anand Bhattad
CVPR3
2025 Training-free diffusion for controlling illumination conditions in images
Xiaoyan Xing, Vincent Tao Hu, Jan Hendrik Metzen, Konrad Groh, Sezer Karaoglu, Theo Gevers
Comput. Vis. Image Underst.5
2025 Exploring dynamic plane representations for neural scene reconstruction
Ruihong Yin, Yunlu Chen, Sezer Karaoglu, Theo Gevers
Pattern Recognit.3
2025 3D human pose estimation and action recognition using fisheye cameras: A survey and benchmark
abstract
3D human pose estimation based on visual information aims to predict 3D poses of humans in images or videos. The aim of human action recognition is to classify what kind of actions people do. Both topics are widely studied in the field of computer vision. Existing methods mainly focus on 3D human pose estimation and human action recognition using images/videos recorded by perspective cameras. In contrast to perspective cameras, fisheye cameras use wide-angle lenses capturing wider field-of-views (FOV). Fisheye cameras are used in many applications such as surveillance and autonomous driving. In this paper, a survey is given on monocular 3D human pose estimation and action recognition. A new benchmark dataset is proposed using a fisheye camera to quantitatively compare and analyze existing methods.
Shaodi You, Sezer Karaoglu, Theo Gevers
Pattern Recognit.3
2024 SceneTeller: Language-to-3D Scene Generation
Basak Melis Öcal, Maxim Tatarchenko, Sezer Karaoglu, Theo Gevers
ECCV (85)3
2024 Ray-Distance Volume Rendering for Neural Scene Reconstruction
Ruihong Yin, Yunlu Chen, Sezer Karaoglu, Theo Gevers
ECCV (14)3
2024 FewViewGS: Gaussian Splatting with Few View Matching and Multi-stage Training
abstract
The field of novel view synthesis from images has seen rapid advancements with the introduction of Neural Radiance Fields (NeRF) and more recently with 3D Gaussian Splatting. Gaussian Splatting became widely adopted due to its efficiency and ability to render novel views accurately. While Gaussian Splatting performs well when a sufficient amount of training images are available, its unstructured explicit representation tends to overfit in scenarios with sparse input images, resulting in poor rendering performance. To address this, we present a 3D Gaussian-based novel view synthesis method using sparse input images that can accurately render the scene from the viewpoints not covered by the training images. We propose a multi-stage training scheme with matching-based consistency constraints imposed on the novel views without relying on pre-trained depth estimation or diffusion models. This is achieved by using the matches of the available training images to supervise the generation of the novel views sampled between the training frames with color, geometry, and semantic losses. In addition, we introduce a locality preserving regularization for 3D Gaussians which removes rendering artifacts by preserving the local color structure of the scene. Evaluation on synthetic and real-world datasets demonstrates competitive or superior performance of our method in few-shot novel view synthesis compared to existing state-of-the-art methods.
Ruihong Yin, Vladimir Yugay, Yue Li 0036, Sezer Karaoglu, Theo Gevers
NeurIPS4
2024 Image semantic segmentation of indoor scenes: A survey
abstract
This survey provides a comprehensive evaluation of various deep learning-based segmentation architectures. It covers a wide range of models, from traditional ones like FCN and PSPNet to more modern approaches like SegFormer and FAN. In addition to assessing the methods in terms of segmentation accuracy, we propose to also evaluate the methods in terms of temporal consistency and corruption vulnerability. Most of the existing surveys on semantic segmentation focus on outdoor datasets. In contrast, this survey focuses on indoor scenarios to enhance the applicability of segmentation methods in this specific domain. Furthermore, our evaluation consists of a performance analysis of the methods in prevalent real-world segmentation scenarios that pose particular challenges. These complex situations involve scenes impacted by diverse forms of noise, blur corruptions, camera movements, optical aberrations, among other factors. By jointly exploring the segmentation accuracy, temporal consistency, and corruption vulnerability in challenging real-world situations, our survey offers insights that go beyond existing surveys, facilitating the understanding and development of better image segmentation methods for indoor scenes.
Ronny Velastegui, Maxim Tatarchenko, Sezer Karaoglu, Theo Gevers
Comput. Vis. Image Underst.3
2024 Kinship similarity for open sets
Wei Wang 0469, Shaodi You, Sezer Karaoglu, Theo Gevers
Pattern Recognit.3
2023 Geometry-guided Feature Learning and Fusion for Indoor Scene Reconstruction
abstract
In addition to color and textural information, geometry provides important cues for 3D scene reconstruction. However, current reconstruction methods only include geometry at the feature level thus not fully exploiting the geometric information.In contrast, this paper proposes a novel geometry integration mechanism for 3D scene reconstruction. Our approach incorporates 3D geometry at three levels, i.e. feature learning, feature fusion, and network supervision. First, geometry-guided feature learning encodes geometric priors to contain view-dependent information. Second, a geometry-guided adaptive feature fusion is introduced which utilizes the geometric priors as a guidance to adaptively generate weights for multiple views. Third, at the supervision level, taking the consistency between 2D and 3D normals into account, a consistent 3D normal loss is designed to add local constraints.Large-scale experiments are conducted on the ScanNet dataset, showing that volumetric methods with our geometry integration mechanism outperform state-of-the-art methods quantitatively as well as qualitatively. Volumetric methods with ours also show good generalization on the 7-Scenes and TUM RGB-D datasets.
Ruihong Yin, Sezer Karaoglu, Theo Gevers
ICCV2
2023 A survey on kinship verification
abstract
In this survey, kinship verification is defined as the automatic process of verifying whether two or more persons are blood relatives (kin) by analyzing images of their faces. Kinship verification is an important research field in computer vision with many applications such as finding missing persons, family album organization, and online image search. Although substantial progress has been made in kinship verification in the past decade, there are still challenges such as intrinsic (face i.e., differences in facial appearance) and extrinsic (acquisition i.e., varying imaging conditions) problems. And there is still a demand for more diverse datasets. Therefore, this paper provides a survey on kinship verification methods and datasets. The survey starts with the definition of kinship verification and its corresponding intrinsic and extrinsic challenges. Then, an overview of kinship verification methods and datasets is given. Finally, a new multi-modal dataset (Nemo-Kinship Dataset) is proposed as a benchmark dataset addressing large inter-subject age variations consisting of 4216 videos of 248 persons from 85 families. The newly collected dataset is used to systematically test and analyze state-of-the-art methods.
Wei Wang 0469, Shaodi You, Sezer Karaoglu, Theo Gevers
Neurocomputing3
2022 Distortion-aware Depth Estimation with Gradient Priors from Panoramas of Indoor Scenes
abstract
Compared to 2D perspective images, panoramic images capture a larger field-of-view (FOV). Depth estimation from panoramas is an important task for 3D scene understanding and has made significant progress with the development of CNNs. However, existing CNN-based methods still suffer from the Equirectangular Projection (ERP) problem to deal with panoramic distortions (e.g. same receptive fields near the equator and the two poles) and have difficulty generating accurate depth boundaries. In contrast to existing CNN-based methods, in this paper, a novel Transformer-based method is proposed which is able to cope with panoramic distortions and to generate accurate depth boundaries. A Distortion-aware Transformer is designed using a yaw-invariant cycle shift and a distortion-guided partitioning. The aim is to alleviate the distortion effect by enlarging the receptive fields in both horizontal and vertical directions. Then, a Gradient Transformer is proposed to enhance the features around the boundaries. Gradient information is adopted as a boundary prior. Large-scale experimental results show an improvement compared to state-of-the-art methods. Our method also shows strong generalization capabilities. Finally, our method is extended to panorama semantic segmentation.
Ruihong Yin, Sezer Karaoglu, Theo Gevers
3DV2
2022 Pose Guided Human Motion Transfer by Exploiting 2D and 3D Information
abstract
Human motion transfer aims to animate the pose of a human in a source image driven by the poses of a human in a target video. To warp (transfer) human poses, most of the existing methods are based on optical flow or affine transformations as an intermediate representation followed by a generator module to perform the motion transfer. Existing methods perform well in terms of reconstruction quality. However, the quality of the human pose transfer has received less attention although it is an important part of the motion transfer process. Therefore, in this paper, we propose a method focusing on both the reconstruction quality as well as pose consistency. In contrast to existing methods, performing warping procedures in 2D- or 3D-space, we introduce a strategy to combine the warped features in both 2D- and 3D-space to alleviate the self-occlusion problem. In this way, our method benefits from 2D (robustness) and 3D (steering) information to guide the generation process. To reduce the pose error caused by inaccurate 3D estimation, a method is proposed to maintain semantic consistency between predictions and target images at arm and leg regions. Experiments conducted on large scale datasets show that the proposed method outperforms existing methods. Ablation studies clarify the benefits of using feature fusion and semantic consistency.
Shaodi You, Sezer Karaoglu, Theo Gevers
3DV3
2022 PIE-Net: Photometric Invariant Edge Guided Network for Intrinsic Image Decomposition
abstract
Intrinsic image decomposition is the process of recovering the image formation components (reflectance and shading) from an image. Previous methods employ either explicit priors to constrain the problem or implicit constraints as formulated by their losses (deep learning). These methods can be negatively influenced by strong illumination conditions causing shading-reflectance leakages. Therefore, in this paper, an end-to-end edge-driven hybrid CNN approach is proposed for intrinsic image decomposition. Edges correspond to illumination invariant gradients. To handle hard negative illumination transitions, a hierarchical approach is taken including global and local refinement layers. We make use of attention layers to further strengthen the learning process. An extensive ablation study and large scale experiments are conducted showing that it is beneficial for edge-driven hybrid IID networks to make use of illumination invariant descriptors and that separating global and local cues helps in improving the performance of the network. Finally, it is shown that the proposed method obtains state of the art performance and is able to generalise well to real world images. The project page with pretrained models, finetuned models and network code can be found at https://ivi.fnwi.uva.nl/cv/pienet/.
Partha Das, Sezer Karaoglu, Theo Gevers
CVPR2
2022 Intrinsic image decomposition using physics-based cues and CNNs
abstract
Intrinsic image decomposition is the decomposition of an image into its reflectance and shading components. The intrinsic image decomposition problem is inherently ill-posed, since there can be multiple solutions to compute the intrinsic components forming the same image. In this paper, we explore the use of physics-based priors. We also propose a new architecture that separates the learning components in a stacked manner. We explore various ways of integrating such priors into a deep learning system. Our method is trained and tested on a large synthetic garden dataset to assess its performance. It is evaluated and compared to state-of-the-art methods using two standard intrinsic datasets. Finally, the pre-trained network is tested on real world images to show the generalisation capabilities of the network.
Partha Das, Sezer Karaoglu, Theo Gevers
Comput. Vis. Image Underst.2
2022 Multi-person 3D pose estimation from a single image captured by a fisheye camera
abstract
Multi-person 3D pose estimation with absolute depths for a fisheye camera is a challenging task but with valuable applications in daily life, especially for video surveillance. However, to the best of our knowledge, such problem has not been explored so far, leaving a gap in practical applications. In this work, we first propose a method for multi-person 3D pose estimation from a single image taken by a fisheye camera. Our method consists of two branches to estimate absolute 3D human poses: (1) a 2D-to-3D lifting module to predict root-relative 3D human poses (HPoseNet); (2) a root regression module to estimate absolute root locations in the camera coordinate (HRootNet). Finally, we propose a fisheye re-projection module without using ground-truth camera parameters to connect two branches, alleviating the impact of image distortions on 3D pose estimation and further regularizing prediction absolute 3D poses. Experimental results demonstrate that our method achieves the state-of-the-art performance on two public multi-person 3D pose datasets with synthetic fisheye images and our newly collected dataset with real fisheye images. The code and new dataset will be made publicly available.
Shaodi You, Sezer Karaoglu, Theo Gevers
Comput. Vis. Image Underst.3
2022 Self-Supervised Face Image Manipulation by Conditioning GAN on Face Decomposition
abstract
We present a novel architecture for manipulating facial expressions, head poses, and lighting conditions from a single monocular image. Recent methods based on Generative Adversarial Networks show promising results in expression manipulation. However, the variation is either defined by a limited number of classes or not well suitable for explicit manipulation of different attributes such as pose and lighting conditions. Besides, state-of-the-art methods are mostly focused on frontal faces. Therefore, in this paper, a new Generative Adversarial Network architecture is proposed by explicitly conditioning on the appearance image space which is the product of direct manipulation of facial expressions, light and pose conditions of the face model in 3D space. In addition, the method only requires video sequences for training. Therefore, it is self-supervised. Unlike other face manipulation methods, the proposed method does not require target specific training. Large scale experiments show that our method outperforms state-of-the-art methods for different scenarios.
Minh Ngô, Sezer Karaoglu, Theo Gevers
IEEE Trans. Multim.2
2021 Multi-Loss Weighting with Coefficient of Variations
abstract
Many interesting tasks in machine learning and computer vision are learned by optimising an objective function defined as a weighted linear combination of multiple losses. The final performance is sensitive to choosing the correct (relative) weights for these losses. Finding a good set of weights is often done by adopting them into the set of hyper- parameters, which are set using an extensive grid search. This is computationally expensive. In this paper, we propose a weighting scheme based on the coefficient of variations and set the weights based on properties observed while training the model1. The proposed method incorporates a measure of uncertainty to balance the losses, and as a result the loss weights evolve during training without requiring another (learning based) optimisation. In contrast to many loss weighting methods in literature, we focus on single-task multi-loss problems, such as monocular depth estimation and semantic segmentation, and show that multi-task approaches for loss weighting do not work on those single-tasks. The validity of the approach is shown empirically for depth estimation and semantic segmentation on multiple datasets.
Rick Groenendijk, Sezer Karaoglu, Theo Gevers, Thomas Mensink
WACV2
2021 EDEN: Multimodal Synthetic Dataset of Enclosed GarDEN Scenes
abstract
Multimodal large-scale datasets for outdoor scenes are mostly designed for urban driving problems. The scenes are highly structured and semantically different from scenarios seen in nature-centered scenes such as gardens or parks. To promote machine learning methods for nature-oriented applications, such as agriculture and gardening, we propose the multimodal synthetic dataset for Enclosed garDEN scenes (EDEN). The dataset features more than 300K images captured from more than 100 garden models. Each image is annotated with various low/high-level vision modalities, including semantic segmentation, depth, surface normals, intrinsic colors, and optical flow. Experimental results on the state-of-the-art methods for semantic segmentation and monocular depth prediction, two important tasks in computer vision, show positive impact of pre-training deep networks on our dataset for unstructured natural scenes. The dataset and related materials will be available at https://lhoangan.github.io/eden.
Hoang-An Le, Thomas Mensink, Partha Das, Sezer Karaoglu, Theo Gevers
WACV4
2021 Identity Unbiased Deception Detection by 2D-to-3D Face Reconstruction
abstract
Deception is a common phenomenon in society, both in our private and professional lives. However, humans are notoriously bad at accurate deception detection. Based on the literature, human accuracy of distinguishing between lies and truthful statements is 54% on average, in other words, it is slightly better than a random guess. While people do not much care about this issue, in high-stakes situations such as interrogations for series crimes and for evaluating the testimonies in court cases, accurate deception detection methods are highly desirable. To achieve a reliable, covert, and non-invasive deception detection, we propose a novel method that disentangles facial expression and head pose related features using 2D-to-3D face reconstruction technique from a video sequence and uses them to learn characteristics of deceptive behavior. We evaluate the proposed method on the Real-Life Trial (RLT) dataset that contains high-stakes deceits recorded in courtrooms. Our results show that the proposed method (with an accuracy of 68%) improves the state of the art. Besides, a new dataset has been collected, for the first time, for low-stake deceit detection. In addition, we compare high-stake deceit detection methods on the newly collected low-stake deceits.
Minh Ngô, Wei Wang 0469, Burak Mandira, Sezer Karaoglu, Henri Bouma, Hamdi Dibeklioglu, Theo Gevers
WACV4
2021 Physics-based shading reconstruction for intrinsic image decomposition
abstract
We investigate the use of photometric invariance and deep learning to compute intrinsic images (albedo and shading). We propose albedo and shading gradient descriptors which are derived from physics-based models. Using the descriptors, albedo transitions are masked out and an initial sparse shading map is calculated directly from the corresponding RGB image gradients in a learning-free unsupervised manner. Then, an optimization method is proposed to reconstruct the full dense shading map. Finally, we integrate the generated shading map into a novel deep learning framework to refine it and also to predict corresponding albedo image to achieve intrinsic image decomposition. By doing so, we are the first to directly address the texture and intensity ambiguity problems of the shading estimations. Large scale experiments show that our approach steered by physics-based invariant descriptors achieve superior results on MIT Intrinsics, NIR-RGB Intrinsics, Multi-Illuminant Intrinsic Images, Spectral Intrinsic Images, As Realistic As Possible, and competitive results on Intrinsic Images in the Wild datasets while achieving state-of-the-art shading estimations.
Anil S. Baslamisli, Yang Liu 0009, Sezer Karaoglu, Theo Gevers
Comput. Vis. Image Underst.3
2021 Pose invariant age estimation of face images in the wild
Wei Wang 0469, Sezer Karaoglu, Wei Zeng 0016, Theo Gevers
Comput. Vis. Image Underst.3
2021 Automatic generation of dense non-rigid optical flow
abstract
There hardly exists any large-scale datasets with dense optical flow of non-rigid motion from real-world imagery as of today. The reason lies mainly in the required setup to derive ground truth optical flows: a series of images with known camera poses along its trajectory, and an accurate 3D model from a textured scene. Human annotation is not only too tedious for large databases, it can simply hardly contribute to accurate optical flow. To circumvent the need for manual annotation, we propose a framework to automatically generate optical flow from real-world videos. The method extracts and matches objects from video frames to compute initial constraints, and applies a deformation over the objects of interest to obtain dense optical flow fields. We propose several ways to augment the optical flow variations. Extensive experimental results show that training on our automatically generated optical flow outperforms methods that are trained on rigid synthetic data using FlowNet-S, LiteFlowNet, PWC-Net, and RAFT. Datasets and implementation of our optical flow generation framework are released at https://github.com/lhoangan/arap_flow.
Hoang-An Le, Tushar Nimbhorkar, Thomas Mensink, Anil S. Baslamisli, Sezer Karaoglu, Theo Gevers
Comput. Vis. Image Underst.5
2021 ShadingNet: Image Intrinsics by Fine-Grained Shading Decomposition
abstract
Abstract In general, intrinsic image decomposition algorithms interpret shading as one unified component including all photometric effects. As shading transitions are generally smoother than reflectance (albedo) changes, these methods may fail in distinguishing strong photometric effects from reflectance variations. Therefore, in this paper, we propose to decompose the shading component into direct (illumination) and indirect shading (ambient light and shadows) subcomponents. The aim is to distinguish strong photometric effects from reflectance variations. An end-to-end deep convolutional neural network (ShadingNet) is proposed that operates in a fine-to-coarse manner with a specialized fusion and refinement unit exploiting the fine-grained shading model. It is designed to learn specific reflectance cues separated from specific photometric effects to analyze the disentanglement capability. A large-scale dataset of scene-level synthetic images of outdoor natural environments is provided with fine-grained intrinsic image ground-truths. Large scale experiments show that our approach using fine-grained shading decompositions outperforms state-of-the-art algorithms utilizing unified shading on NED, MPI Sintel, GTA V, IIW, MIT Intrinsic Images, 3DRMS and SRD datasets.
Anil S. Baslamisli, Partha Das, Hoang-An Le, Sezer Karaoglu, Theo Gevers
Int. J. Comput. Vis.4
2020 Unified Application of Style Transfer for Face Swapping and Reenactment
Minh Ngô, Christian aan de Wiel, Sezer Karaoglu, Theo Gevers
ACCV (5)3
2020 Pano2Scene: 3D Indoor Semantic Scene Reconstruction from a Single Indoor Panorama Image
Wei Zeng 0016, Sezer Karaoglu, Theo Gevers
BMVC2
2020 Joint 3D Layout and Depth Prediction from a Single Indoor Panorama Image
Wei Zeng 0016, Sezer Karaoglu, Theo Gevers
ECCV (16)2
2020 Object features and face detection performance: Analyses with 3D-rendered synthetic data
abstract
This paper is to provide an overview of how object features from images influence face detection performance, and how to select synthetic faces to address specific features. To this end, we investigate the effects of occlusion, scale, viewpoint, background, and noise by using a novel synthetic image generator based on 3DU Face Dataset. To examine the effects of different features, we selected three detectors (Faster RCNN, HR, SSH) as representative of various face detection methodologies. Comparing different configurations of synthetic data on face detection systems, it showed that our synthetic dataset could complement face detectors to become more robust against features in the real world. Our analysis also demonstrated that a variety of data augmentation is necessary to address nuanced differences in performance.
Sezer Karaoglu, Hoang-An Le, Theo Gevers
ICPR2
2020 On the benefit of adversarial training for monocular depth estimation
abstract
In this paper we address the benefit of adding adversarial training to the task of monocular depth estimation. A model can be trained in a self-supervised setting on stereo pairs of images, where depth (disparities) are an intermediate result in a right-to-left image reconstruction pipeline. For the quality of the image reconstruction and disparity prediction, a combination of different losses is used, including L1 image reconstruction losses and left–right disparity smoothness. These are local pixel-wise losses, while depth prediction requires global consistency. Therefore, we extend the self-supervised network to become a Generative Adversarial Network (GAN), by including a discriminator which should tell apart reconstructed (fake) images from real images. We evaluate Vanilla GANs, LSGANs and Wasserstein GANs in combination with different pixel-wise reconstruction losses. Based on extensive experimental evaluation, we conclude that adversarial training is beneficial if and only if the reconstruction loss is not too constrained. Even though adversarial training seems promising because it promotes global consistency, non-adversarial training outperforms (or is on par with) any method trained with a GAN when a constrained reconstruction loss is used in combination with batch normalisation. Based on the insights of our experimental evaluation we obtain state-of-the art monocular depth estimation results by using batch normalisation and different output scales.
Rick Groenendijk, Sezer Karaoglu, Theo Gevers, Thomas Mensink
Comput. Vis. Image Underst.2
2018 Joint Learning of Intrinsic Images and Semantic Segmentation
Anil S. Baslamisli, Thomas T. Groenestege, Partha Das, Hoang-An Le, Sezer Karaoglu, Theo Gevers
ECCV (6)5
2017 Point Light Source Position Estimation From RGB-D Images by Learning Surface Attributes
abstract
Light source position (LSP) estimation is a difficult yet an important problem in computer vision. A common approach for estimating the LSP assumes Lambert's law. However, in real-world scenes, Lambert's law does not hold for all different types of surfaces. Instead of assuming all that surfaces follow Lambert's law, our approach classifies image surface segments based on their photometric and geometric surface attributes (i.e. glossy, matte, curved, and so on) and assigns weights to image surface segments based on their suitability for LSP estimation. In addition, we propose the use of the estimated camera pose to globally constrain LSP for RGB-D video sequences. Experiments on Boom and a newly collected RGB-D video data sets show that the state-of-the-art methods are outperformed by the proposed method. The results demonstrate that weighting image surface segments based on their attributes outperform the state-of-the-art methods in which the image surface segments are considered to equally contribute. In particular, by using the proposed surface weighting, the angular error for LSP estimation is reduced from 12.6° to 8.2° and 24.6° to 4.8° for Boom and RGB-D video data sets, respectively. Moreover, using the camera pose to globally constrain LSP provides higher accuracy (4.8°) compared with using single frames (8.5°).
Sezer Karaoglu, Yang Liu 0009, Theo Gevers, Arnold W. M. Smeulders
IEEE Trans. Image Process.1
2017 Con-Text: Text Detection for Fine-Grained Object Classification
abstract
This paper focuses on fine-grained object classification using recognized scene text in natural images. While the state-of-the-art relies on visual cues only, this paper is the first work which proposes to combine textual and visual cues. Another novelty is the textual cue extraction. Unlike the state-of-the-art text detection methods, we focus more on the background instead of text regions. Once text regions are detected, they are further processed by two methods to perform text recognition, i.e., ABBYY commercial OCR engine and a state-of-the-art character recognition algorithm. Then, to perform textual cue encoding, bi- and trigrams are formed between the recognized characters by considering the proposed spatial pairwise constraints. Finally, extracted visual and textual cues are combined for fine-grained classification. The proposed method is validated on four publicly available data sets: ICDAR03, ICDAR13, Con-Text, and Flickr-logo. We improve the state-of-the-art end-to-end character recognition by a large margin of 15% on ICDAR03. We show that textual cues are useful in addition to visual cues for fine-grained classification. We show that textual cues are also useful for logo retrieval. Adding textual cues outperforms visual- and textual-only in fine-grained classification (70.7% to 60.3%) and logo retrieval (57.4% to 54.8%).
Sezer Karaoglu, Ran Tao 0004, Jan C. van Gemert, Theo Gevers
IEEE Trans. Image Process.1
2017 Words Matter: Scene Text for Image Classification and Retrieval
abstract
Text in natural images typically adds meaning to an object or scene. In particular, text specifies which business places serve drinks (e.g., cafe, teahouse) or food (e.g., restaurant, pizzeria), and what kind of service is provided (e.g., massage, repair). The mere presence of text, its words, and meaning are closely related to the semantics of the object or scene. This paper exploits textual contents in images for fine-grained business place classification and logo retrieval. There are four main contributions. First, we show that the textual cues extracted by the proposed method are effective for the two tasks. Combining the proposed textual and visual cues outperforms visual only classification and retrieval by a large margin. Second, to extract the textual cues, a generic and fully unsupervised word box proposal method is introduced. The method reaches state-of-the-art word detection recall with a limited number of proposals. Third, contrary to what is widely acknowledged in text detection literature, we demonstrate that high recall in word detection is more important than high f-score at least for both tasks considered in this work. Last, this paper provides a large annotated text detection dataset with 10 K images and 27 601 word boxes.
Sezer Karaoglu, Ran Tao 0004, Theo Gevers, Arnold W. M. Smeulders
IEEE Trans. Multim.1
2016 Large scale Gaussian Process for overlap-based object proposal scoring
Silvia L. Pintea, Sezer Karaoglu, Jan C. van Gemert, Arnold W. M. Smeulders
Comput. Vis. Image Underst.2
2016 Detect2Rank: Combining Object Detectors Using Learning to Rank
abstract
Object detection is an important research area in the field of computer vision. Many detection algorithms have been proposed. However, each object detector relies on specific assumptions of the object appearance and imaging conditions. As a consequence, no algorithm can be considered universal. With the large variety of object detectors, the subsequent question is how to select and combine them. In this paper, we propose a framework to learn how to combine object detectors. The proposed method uses (single) detectors like Deformable Part Models, Color Names and Ensemble of Exemplar-SVMs, and exploits their correlation by high-level contextual features to yield a combined detection list. Experiments on the PASCAL VOC07 and VOC10 data sets show that the proposed method significantly outperforms single object detectors, DPM (8.4%), CN (6.8%) and EES (17.0%) on VOC07 and DPM (6.5%), CN (5.5%) and EES (16.2%) on VOC10. We show with an experiment that there are no constraints on the type of the detector. The proposed method outperforms (2.4%) the state-of-the-art object detector (RCNN) on VOC07 when Regions with Convolutional Neural Network is combined with other detectors used in this paper.
Sezer Karaoglu, Yang Liu 0009, Theo Gevers
IEEE Trans. Image Process.1
2015 Age estimation under changes in image quality: An experimental study
abstract
In this paper, we investigate the influence of image quality on the performance of aging features. Age estimation systems used or designed a number of aging features to capture the aging cues from the face such as skin texture and wrinkles. These aging cues are sensitive to small changes in the imaging conditions which suggests considering the imaging quality when extracting such information. Although interesting performances are reported on various datasets, the effect of image quality has not been addressed. We introduce a scheme to explore the influence of image quality on the performance of appearance aging features. A number of datasets are experimented on where artifacts resulted from different types of noise are considered. Finally, we propose a method to automatically apply the most suitable features based on the quality of the image. The results show that better or comparable performance is obtained when automatically applying different features, based on image quality, in comparison to a single (best) feature type.
Fares Alnajar, Theo Gevers, Sezer Karaoglu
ICIP3
2015 Per-patch metric learning for robust image matching
abstract
We propose a patch-specific metric learning method to improve matching performance of local descriptors. Existing methodologies typically focus on invariance, by completely considering, or completely disregarding all variations. We propose a metric learning method that is robust to only a range of variations. The ability to choose the level of robustness allows us to fine-tune the trade-off between invariance and discriminative power. We learn a distance metric for each patch independently by sampling from a set of relevant image transformations. These transformations give a-priori knowledge about the behavior of the query patch under the applied transformation in feature space. We learn the robust metric by either fully generating only the relevant range of transformations, or by a novel direct metric. The matching between query patch and data is performed with this new metric. Results on the ALOI dataset show that the proposed method improves performance of SIFT by 6.22% for geometric and 4.43% for photometric transformations.
Sezer Karaoglu, Ivo Everts, Jan C. van Gemert, Theo Gevers
ICIP1
2013 Con-text: text detection using background connectivity for fine-grained object classification
abstract
This paper focuses on fine-grained classification by detecting photographed text in images. We introduce a text detection method that does not try to detect all possible foreground text regions but instead aims to reconstruct the scene background to eliminate non-text regions. Object cues such as color, contrast, and objectiveness are used in corporation with a random forest classifier to detect background pixels in the scene. Results on two publicly available datasets ICDAR03 and a fine-grained Building subcategories of ImageNet shows the effectiveness of the proposed method.
Sezer Karaoglu, Jan C. van Gemert, Theo Gevers
ACM Multimedia1