Emre Akbas

dblp:78/1103 · DBLP profile ↗
← Back
33ranked-venue papers
6as first author
17since 2021 · last 2026
0000-0002-3760-6722ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 27 · 5 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 20 · 3 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 Q-Former Autoencoder: A Modern Framework for Medical Anomaly Detection
abstract
Anomaly detection in medical images is an important yet challenging task due to the diversity of possible anomalies and the practical impossibility of collecting comprehensively annotated data sets. In this work, we tackle unsupervised medical anomaly detection proposing a modernized autoencoder-based framework, the Q-Former Autoencoder, that leverages state-of-the-art pretrained vision foundation models, such as DINO, DINOv2 and Masked Autoencoder. Instead of training encoders from scratch, we directly utilize frozen vision foundation models as feature extractors, enabling rich, multi-stage, high-level representations without domain-specific fine-tuning. We propose the usage of the Q-Former architecture as the bottleneck, which enables the control of the length of the reconstruction sequence, while efficiently aggregating multi-scale features. Additionally, we incorporate a perceptual loss computed using features from a pretrained Masked Autoencoder, guiding the reconstruction towards semantically meaningful structures. Our framework is evaluated on four diverse medical anomaly detection benchmarks, achieving state-of-the-art results on BraTS2021, RESC, and RSNA. Our results highlight the potential of vision foundation model encoders, pretrained on natural images, to generalize effectively to medical image analysis tasks without further fine-tuning. We release the code and models at https://github.com/emirhanbayar/QFAE.
Francesco Dalmonte, Emirhan Bayar, Emre Akbas, Mariana-Iuliana Georgescu
WACV3
2026 Representation recycling for streaming video analysis
Can Ufuk Ertenli, Ramazan Gokberk Cinbis, Emre Akbas
Neurocomputing3
2026 Intrinsic dimensionality as a model-free measure of class imbalance
Çagri Eser, Zeynep Sonat Baltaci, Emre Akbas, Sinan Kalkan
Neurocomputing3
2025 WAIT: Feature warping for animation to illustration video translation using GANs
Samet Hicsonmez, Nermin Samet, Fidan Samet, Oguz Bakir, Emre Akbas, Pinar Duygulu
Neurocomputing5
2024 RankED: Addressing Imbalance and Uncertainty in Edge Detection Using Ranking-based Losses
abstract
Detecting edges in images suffers from the problems of (PI) heavy imbalance between positive and negative classes as well as (P2) label uncertainty owing to disagreement be-tween different annotators. Existing solutions address PI using class-balanced cross-entropy loss and dice loss and P2 by only predicting edges agreed upon by most annota-tors. In this paper, we propose RankED, a unified ranking-based approach that addresses both the imbalance problem (P 1) and the uncertainty problem (P2). RankED tackles these two problems with two components: One component which ranks positive pixels over negative pixels, and the second which promotes high confidence edge pixels to have more label certainty. We show that RankED outperforms previous studies and sets a new state-of-the-art on NYUD-v2, BSDS500 and Multi-cue datasets. Code is available at https://ranked-cvpr24.github.io.
Bedrettin Çetinkaya, Sinan Kalkan, Emre Akbas
CVPR3
2024 Bucketed Ranking-Based Losses for Efficient Training of Object Detectors
Feyza Yavuz, Baris Can Cam, Adnan Harun Dogan, Kemal Oksuz, Emre Akbas, Sinan Kalkan
ECCV (59)5
2024 SegIns: A simple extension to instance discrimination task for better localization learning
Melih Baydar, Emre Akbas
J. Vis. Commun. Image Represent.2
2023 Correlation Loss: Enforcing Correlation between Classification and Localization
abstract
Object detectors are conventionally trained by a weighted sum of classification and localization losses. Recent studies (e.g., predicting IoU with an auxiliary head, Generalized Focal Loss, Rank & Sort Loss) have shown that forcing these two loss terms to interact with each other in non-conventional ways creates a useful inductive bias and improves performance. Inspired by these works, we focus on the correlation between classification and localization and make two main contributions: (i) We provide an analysis about the effects of correlation between classification and localization tasks in object detectors. We identify why correlation affects the performance of various NMS-based and NMS-free detectors, and we devise measures to evaluate the effect of correlation and use them to analyze common detectors. (ii) Motivated by our observations, e.g., that NMS-free detectors can also benefit from correlation, we propose Correlation Loss, a novel plug-in loss function that improves the performance of various object detectors by directly optimizing correlation coefficients: E.g., Correlation Loss on Sparse R-CNN, an NMS-free method, yields 1.6 AP gain on COCO and 1.8 AP gain on Cityscapes dataset. Our best model on Sparse R-CNN reaches 51.0 AP without test-time augmentation on COCO test-dev, reaching state-of-the-art. Code is available at: https://github.com/fehmikahraman/CorrLoss.
Fehmi Kahraman, Kemal Oksuz, Sinan Kalkan, Emre Akbas
AAAI4
2023 HoughNet: Integrating Near and Long-Range Evidence for Visual Detection
abstract
This paper presents HoughNet, a one-stage, anchor-free, voting-based, bottom-up object detection method. Inspired by the Generalized Hough Transform, HoughNet determines the presence of an object at a certain location by the sum of the votes cast on that location. Votes are collected from both near and long-distance locations based on a log-polar vote field. Thanks to this voting mechanism, HoughNet is able to integrate both near and long-range, class-conditional evidence for visual recognition, thereby generalizing and enhancing current object detection methodology, which typically relies on only local evidence. On the COCO dataset, HoughNet's best model achieves 46.4$AP$(and 65.1$AP_{50}$), performing on par with the state-of-the-art in bottom-up object detection and outperforming most major one-stage and two-stage methods. We further validate the effectiveness of our proposal in other visual detection tasks, namely, video object detection, instance segmentation, 3D object detection and keypoint detection for human pose estimation, and an additional “labels to photo” image generation task, where the integration of our voting module consistently improves performance in all cases. Code is available athttps://github.com/nerminsamet/houghnet.
Nermin Samet, Samet Hicsonmez, Emre Akbas
IEEE Trans. Pattern Anal. Mach. Intell.3
2022 Streaming Multiscale Deep Equilibrium Models
Can Ufuk Ertenli, Emre Akbas, Ramazan Gokberk Cinbis
ECCV (11)2
2022 Does depth estimation help object detection?
Bedrettin Çetinkaya, Sinan Kalkan, Emre Akbas
Image Vis. Comput.3
2022 One Metric to Measure Them All: Localisation Recall Precision (LRP) for Evaluating Visual Detection Tasks
abstract
Despite being widely used as a performance measure for visual detection tasks, Average Precision (AP) is limited in (i) reflecting localisation quality, (ii) interpretability and (iii) robustness to the design choices regarding its computation, and its applicability to outputs without confidence scores. Panoptic Quality (PQ), a measure proposed for evaluating panoptic segmentation (Kirillov et al., 2019), does not suffer from these limitations but is limited to panoptic segmentation. In this paper, we propose Localisation Recall Precision (LRP) Error as the average matching error of a visual detector computed based on both its localisation and classification qualities for a given confidence score threshold. LRP Error, initially proposed only for object detection by Oksuz et al. (2018), does not suffer from the aforementioned limitations and is applicable to all visual detection tasks. We also introduce Optimal LRP (oLRP) Error as the minimum LRP Error obtained over confidence scores to evaluate visual detectors and obtain optimal thresholds for deployment. We provide a detailed comparative analysis of LRP Error with AP and PQ, and use nearly 100 state-of-the-art visual detectors from seven visual detection tasks (i.e. object detection, keypoint detection, instance segmentation, panoptic segmentation, visual relationship detection, zero-shot detection and generalised zero-shot detection) using ten datasets to empirically show that LRP Error provides richer and more discriminative information than its counterparts. Code available at: https://github.com/kemaloksuz/LRP-Error.
Kemal Oksuz, Baris Can Cam, Sinan Kalkan, Emre Akbas
IEEE Trans. Pattern Anal. Mach. Intell.4
2021 Mask-aware IoU for Anchor Assignment in Real-time Instance Segmentation
Kemal Oksuz, Baris Can Cam, Fehmi Kahraman, Zeynep Sonat Baltaci, Sinan Kalkan, Emre Akbas
BMVC6
2021 Rank & Sort Loss for Object Detection and Instance Segmentation
abstract
We propose Rank & Sort (RS) Loss, a ranking-based loss function to train deep object detection and instance segmentation methods (i.e. visual detectors). RS Loss supervises the classifier, a sub-network of these methods, to rank each positive above all negatives as well as to sort positives among themselves with respect to (wrt.) their localisation qualities (e.g. Intersection-over-Union - IoU). To tackle the non-differentiable nature of ranking and sorting, we reformulate the incorporation of error-driven update with back-propagation as Identity Update, which enables us to model our novel sorting error among positives. With RS Loss, we significantly simplify training: (i) Thanks to our sorting objective, the positives are prioritized by the classifier without an additional auxiliary head (e.g. for centerness, IoU, mask-IoU), (ii) due to its ranking-based nature, RS Loss is robust to class imbalance, and thus, no sampling heuristic is required, and (iii) we address the multi-task nature of visual detectors using tuning-free task-balancing coefficients. Using RS Loss, we train seven diverse visual detectors only by tuning the learning rate, and show that it consistently outperforms baselines: e.g. our RS Loss improves (i) Faster R-CNN by ∼ 3 box AP and aLRP Loss (ranking-based baseline) by ∼ 2 box AP on COCO dataset, (ii) Mask R-CNN with repeat factor sampling (RFS) by 3.5 mask AP (∼ 7 AP for rare classes) on LVIS dataset; and also outperforms all counterparts. Code is available at: https://github.com/kemaloksuz/RankSortLoss.
Kemal Oksuz, Baris Can Cam, Emre Akbas, Sinan Kalkan
ICCV3
2021 Adversarial Segmentation Loss For Sketch Colorization
abstract
We introduce a new method for generating color images from sketches or edge maps. Current methods either require some form of additional user-guidance or are limited to the “paired” translation approach. We argue that segmentation information could provide valuable guidance for sketch colorization. To this end, we propose to leverage semantic image segmentation, as provided by a general purpose panoptic segmentation network, to create an additional adversarial loss function. Our loss function can be integrated to any baseline GAN model. Our method is not limited to datasets that contain segmentation labels, and it can be trained for “unpaired” translation tasks. We show the effectiveness of our method on four different datasets spanning scene level indoor, outdoor, and children book illustration images using qualitative, quantitative and user study analysis. Our model improves its baseline up to 35 points on the FID metric. Our code and pretrained models can be found at https://github.com/giddyyupp/AdvSegLoss.
Samet Hicsonmez, Nermin Samet, Emre Akbas, Pinar Duygulu
ICIP3
2021 HPRNet: Hierarchical point regression for whole-body human pose estimation
Nermin Samet, Emre Akbas
Image Vis. Comput.2
2021 Imbalance Problems in Object Detection: A Review
abstract
In this paper, we present a comprehensive review of the imbalance problems in object detection. To analyze the problems in a systematic manner, we introduce a problem-based taxonomy. Following this taxonomy, we discuss each problem in depth and present a unifying yet critical perspective on the solutions in the literature. In addition, we identify major open issues regarding the existing imbalance problems as well as imbalance problems that have not been discussed before. Moreover, in order to keep our review up to date, we provide an accompanying webpage which catalogs papers addressing imbalance problems, according to our problem-based taxonomy. Researchers can track newer studies on this webpage available at: https://github.com/kemaloksuz/ObjectDetectionImbalance.
Kemal Oksuz, Baris Can Cam, Sinan Kalkan, Emre Akbas
IEEE Trans. Pattern Anal. Mach. Intell.4
2020 Reducing Label Noise in Anchor-Free Object Detection
Nermin Samet, Samet Hicsonmez, Emre Akbas
BMVC3
2020 HoughNet: Integrating Near and Long-Range Evidence for Bottom-Up Object Detection
Nermin Samet, Samet Hicsonmez, Emre Akbas
ECCV (25)3
2020 A Ranking-based, Balanced Loss Function Unifying Classification and Localisation in Object Detection
abstract
We propose average Localisation-Recall-Precision (aLRP), a unified, bounded, balanced and ranking-based loss function for both classification and localisation tasks in object detection. aLRP extends the Localisation-Recall-Precision (LRP) performance metric (Oksuz et al., 2018) inspired from how Average Precision (AP) Loss extends precision to a ranking-based loss function for classification (Chen et al., 2020). aLRP has the following distinct advantages: (i) aLRP is the first ranking-based loss function for both classification and localisation tasks. (ii) Thanks to using ranking for both tasks, aLRP naturally enforces high-quality localisation for high-precision classification. (iii) aLRP provides provable balance between positives and negatives. (iv) Compared to on average ~6 hyperparameters in the loss functions of state-of-the-art detectors, aLRP Loss has only one hyperparameter, which we did not tune in practice. On the COCO dataset, aLRP Loss improves its ranking-based predecessor, AP Loss, up to around 5 AP points, achieves 48.9 AP without test time augmentation and outperforms all one-stage detectors. Code available at: https://github.com/kemaloksuz/aLRPLoss .
Kemal Oksuz, Baris Can Cam, Emre Akbas, Sinan Kalkan
NeurIPS3
2020 Generating Positive Bounding Boxes for Balanced Training of Object Detectors
abstract
Two-stage deep object detectors generate a set of regions-of-interest RoIs in the first stage, then, in the second stage, identify objects among the proposed RoIs that sufficiently overlap with a ground truth (GT) box. The second stage is known to suffer from a bias towards RoIs that have low intersection-over-union (IoU) with the associated GT boxes. To address this issue, we first propose a sampling method to generate bounding boxes (BB) that overlap with a given reference box more than a given IoU threshold. Then, we use this BB generation method to develop a positive RoI (pRoI) generator that, for the second stage, produces RoIs following any desired spatial or IoU distribution. We show that our pRoI generator is able to simulate other sampling methods for positive examples such as hard example mining and prime sampling. Using our generator as an analysis tool, we show that (i) IoU imbalance has an adverse effect on performance, (ii) hard positive example mining improves the performance only for certain input IoU distributions, and (iii) the imbalance among the foreground classes has an adverse effect on performance and that it can be alleviated at the batch level. Finally, we train Faster R-CNN using our pRoI generator and, compared to conventional training, obtain better or on-par performance for low IoUs and significant improvements when trained for higher IoUs for Pascal VOC and MS COCO datasets. The code is available at: https://github.com/kemaloksuz/BoundingBoxGenerator.
Kemal Oksuz, Baris Can Cam, Emre Akbas, Sinan Kalkan
WACV3
2020 Low-level multiscale image segmentation and a benchmark for its evaluation
Emre Akbas, Narendra Ahuja
Comput. Vis. Image Underst.1
2020 GANILLA: Generative adversarial networks for image to illustration translation
Samet Hicsonmez, Nermin Samet, Emre Akbas, Pinar Duygulu
Image Vis. Comput.3
2019 Self-Supervised Learning of 3D Human Pose Using Multi-View Geometry
abstract
Training accurate 3D human pose estimators requires large amount of 3D ground-truth data which is costly to collect. Various weakly or self supervised pose estimation methods have been proposed due to lack of 3D data. Nevertheless, these methods, in addition to 2D ground-truth poses, require either additional supervision in various forms (e.g. unpaired 3D ground truth data, a small subset of labels) or the camera parameters in multiview settings. To address these problems, we present EpipolarPose, a self-supervised learning method for 3D human pose estimation, which does not need any 3D ground-truth data or camera extrinsics. During training, EpipolarPose estimates 2D poses from multi-view images, and then, utilizes epipolar geometry to obtain a 3D pose and camera geometry which are subsequently used to train a 3D pose estimator. We demonstrate the effectiveness of our approach on standard benchmark datasets (i.e. Human3.6M and MPI-INF-3DHP) where we set the new state-of-the-art among weakly/self-supervised methods. Furthermore, we propose a new performance measure Pose Structure Score (PSS) which is a scale invariant, structure aware measure to evaluate the structural plausibility of a pose with respect to its ground truth. Code and pretrained models are available at https://github.com/mkocabas/EpipolarPose.
Muhammed Kocabas, Salih Karagoz, Emre Akbas
CVPR3
2018 MultiPoseNet: Fast Multi-Person Pose Estimation Using Pose Residual Network
Muhammed Kocabas, Salih Karagoz, Emre Akbas
ECCV (11)3
2018 Localization Recall Precision (LRP): A New Performance Metric for Object Detection
Kemal Oksuz, Baris Can Cam, Emre Akbas, Sinan Kalkan
ECCV (7)3
2017 Object detection through search with a foveated visual system
abstract
Humans and many other species sense visual information with varying spatial resolution across the visual field (foveated vision) and deploy eye movements to actively sample regions of interests in scenes. The advantage of such varying resolution architecture is a reduced computational, hence metabolic cost. But what are the performance costs of such processing strategy relative to a scheme that processes the visual field at high spatial resolution? Here we first focus on visual search and combine object detectors from computer vision with a recent model of peripheral pooling regions found at the V1 layer of the human visual system. We develop a foveated object detector that processes the entire scene with varying resolution, uses retino-specific object detection classifiers to guide eye movements, aligns its fovea with regions of interest in the input image and integrates observations across multiple fixations. We compared the foveated object detector against a non-foveated version of the same object detector which processes the entire image at homogeneous high spatial resolution. We evaluated the accuracy of the foveated and non-foveated object detectors identifying 20 different objects classes in scenes from a standard computer vision data set (the PASCAL VOC 2007 dataset). We show that the foveated object detector can approximate the performance of the object detector with homogeneous high spatial resolution processing while bringing significant computational cost savings. Additionally, we assessed the impact of foveation on the computation of bottom-up saliency. An implementation of a simple foveated bottom-up saliency model with eye movements showed agreement in the selection of top salient regions of scenes with those selected by a non-foveated high resolution saliency model. Together, our results might help explain the evolution of foveated visual systems with eye movements as a solution that preserves perceptual performance in visual search while resulting in computational and metabolic savings to the brain.
Emre Akbas, Miguel P. Eckstein
PLoS Comput. Biol.1
2014 Low-Level Hierarchical Multiscale Segmentation Statistics of Natural Images
abstract
This paper is aimed at obtaining the statistics as a probabilistic model pertaining to the geometric, topological and photometric structure of natural images. The image structure is represented by its segmentation graph derived from the low-level hierarchical multiscale image segmentation. We first estimate the statistics of a number of segmentation graph properties from a large number of images. Our estimates confirm some findings reported in the past work, as well as provide some new ones. We then obtain a Markov random field based model of the segmentation graph which subsumes the observed statistics. To demonstrate the value of the model and the statistics, we show how its use as a prior impacts three applications: image classification, semantic image segmentation and object detection.
Emre Akbas, Narendra Ahuja
IEEE Trans. Pattern Anal. Mach. Intell.1
2010 Pedestrian Recognition with a Learned Metric
Mert Dikmen, Emre Akbas, Thomas S. Huang, Narendra Ahuja
ACCV (4)2
2010 Low-Level Image Segmentation Based Scene Classification
abstract
This paper is aimed at evaluating the semantic information content of multiscale, low-level image segmentation. As a method of doing this, we use selected features of segmentation for semantic classification of real images. To estimate the relative measure of the information content of our features, we compare the results of classifications we obtain using them with those obtained by others using the commonly used patch/grid based features. To classify an image using segmentation based features, we model the image in terms of a probability density function, a Gaussian mixture model (GMM) to be specific, of its region features. This GMM is fit to the image by adapting a universal GMM which is estimated so it fits all images. Adaptation is done using a maximum-aposteriori criterion. We use kernelized versions of Bhattacharyya distance to measure the similarity between two GMMs and support vector machines to perform classification. We outperform previously reported results on a publicly available scene classification dataset. These results suggest further experimentation in evaluating the promise of low level segmentation in image classification.
Emre Akbas, Narendra Ahuja
ICPR1
2009 From Ramp Discontinuities to Segmentation Tree
Emre Akbas, Narendra Ahuja
ACCV (1)1
2007 Automatic Image Annotation by Ensemble of Visual Descriptors
abstract
Automatic image annotation systems available in the literature concatenate color, texture and/or shape features in a single feature vector to learn a set of high level semantic categories using a single learning machine. This approach is quite naive to map the visual features to high level semantic information concerning the categories. Concatenation of many features with different visual properties and wide dynamical ranges may result in curse of dimensionality and redundancy problems. Additionally, it usually requires normalization which may cause an undesirable distortion in the feature space. An elegant way of reducing the effects of these problems is to design a dedicated feature space for each image category, depending on its content, and learn a range of visual properties of the whole image from a variety of feature sets. For this purpose, a two-layer ensemble learning system, called Supervised Annotation by Descriptor Ensemble (SADE), is proposed. SADE, initially, extracts a variety of low-level visual descriptors from the image. Each descriptor is, then, fed to a separate learning machine in the first layer. Finally, the meta-layer classifier is trained on the output of the first layer classifiers and the images are annotated by using the decision of the meta-layer classifier. This approach not only avoids normalization, but also reduces the effects of dimensional curse and redundancy. The proposed system outperforms a state-of-the-art automatic image annotation system, in an equivalent experimental setup.
Emre Akbas, Fatos T. Yarman-Vural
CVPR1
2006 A Hierarchical Classification System Based on Adaptive Resonance Theory
abstract
In this study, we propose a hierarchical classification system, which emulates the eye-brain channel in two hierarchical layers. In the first layer, a set of classifiers are trained by using low level, low dimensional features. In the second layer, the recognition results of the first layer are fed to the Fuzzy ARTMAP (FAM) classifier which implements the Adaptive Resonance Theory. Experiments indicate that the hierarchical approach proposed in this paper, increases the classification performances compared to the available methods.
Mutlu Uysal, Emre Akbas, Fatos T. Yarman-Vural
ICIP2