EDBT 2026 Demo / reviewers in the wild / expert
Emre Akbas
dblp:78/1103
· DBLP profile ↗
33ranked-venue papers
6as first author
17since 2021 · last 2026
0000-0002-3760-6722ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 27 · 5 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 20 · 3 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Q-Former Autoencoder: A Modern Framework for Medical Anomaly DetectionabstractAnomaly detection in medical images is an important yet challenging task due to the diversity of possible anomalies and the practical impossibility of collecting comprehensively annotated data sets. In this work, we tackle unsupervised medical anomaly detection proposing a modernized autoencoder-based framework, the Q-Former Autoencoder, that leverages state-of-the-art pretrained vision foundation models, such as DINO, DINOv2 and Masked Autoencoder. Instead of training encoders from scratch, we directly utilize frozen vision foundation models as feature extractors, enabling rich, multi-stage, high-level representations without domain-specific fine-tuning. We propose the usage of the Q-Former architecture as the bottleneck, which enables the control of the length of the reconstruction sequence, while efficiently aggregating multi-scale features. Additionally, we incorporate a perceptual loss computed using features from a pretrained Masked Autoencoder, guiding the reconstruction towards semantically meaningful structures. Our framework is evaluated on four diverse medical anomaly detection benchmarks, achieving state-of-the-art results on BraTS2021, RESC, and RSNA. Our results highlight the potential of vision foundation model encoders, pretrained on natural images, to generalize effectively to medical image analysis tasks without further fine-tuning. We release the code and models at https://github.com/emirhanbayar/QFAE. Francesco Dalmonte, Emirhan Bayar, Emre Akbas, Mariana-Iuliana Georgescu |
WACV | 3 |
| 2026 | Representation recycling for streaming video analysis
Can Ufuk Ertenli, Ramazan Gokberk Cinbis, Emre Akbas |
Neurocomputing | 3 |
| 2026 | Intrinsic dimensionality as a model-free measure of class imbalance
Çagri Eser, Zeynep Sonat Baltaci, Emre Akbas, Sinan Kalkan |
Neurocomputing | 3 |
| 2025 | WAIT: Feature warping for animation to illustration video translation using GANs
Samet Hicsonmez, Nermin Samet, Fidan Samet, Oguz Bakir, Emre Akbas, Pinar Duygulu |
Neurocomputing | 5 |
| 2024 | RankED: Addressing Imbalance and Uncertainty in Edge Detection Using Ranking-based LossesabstractDetecting edges in images suffers from the problems of (PI) heavy imbalance between positive and negative classes as well as (P2) label uncertainty owing to disagreement be-tween different annotators. Existing solutions address PI using class-balanced cross-entropy loss and dice loss and P2 by only predicting edges agreed upon by most annota-tors. In this paper, we propose RankED, a unified ranking-based approach that addresses both the imbalance problem (P 1) and the uncertainty problem (P2). RankED tackles these two problems with two components: One component which ranks positive pixels over negative pixels, and the second which promotes high confidence edge pixels to have more label certainty. We show that RankED outperforms previous studies and sets a new state-of-the-art on NYUD-v2, BSDS500 and Multi-cue datasets. Code is available at https://ranked-cvpr24.github.io. Bedrettin Çetinkaya, Sinan Kalkan, Emre Akbas |
CVPR | 3 |
| 2024 | Bucketed Ranking-Based Losses for Efficient Training of Object Detectors
Feyza Yavuz, Baris Can Cam, Adnan Harun Dogan, Kemal Oksuz, Emre Akbas, Sinan Kalkan |
ECCV (59) | 5 |
| 2024 | SegIns: A simple extension to instance discrimination task for better localization learning
Melih Baydar, Emre Akbas |
J. Vis. Commun. Image Represent. | 2 |
| 2023 | Correlation Loss: Enforcing Correlation between Classification and LocalizationabstractObject detectors are conventionally trained by a weighted sum of classification and localization losses. Recent studies (e.g., predicting IoU with an auxiliary head, Generalized Focal Loss, Rank & Sort Loss) have shown that forcing these two loss terms to interact with each other in non-conventional ways creates a useful inductive bias and improves performance. Inspired by these works, we focus on the correlation between classification and localization and make two main contributions: (i) We provide an analysis about the effects of correlation between classification and localization tasks in object detectors. We identify why correlation affects the performance of various NMS-based and NMS-free detectors, and we devise measures to evaluate the effect of correlation and use them to analyze common detectors. (ii) Motivated by our observations, e.g., that NMS-free detectors can also benefit from correlation, we propose Correlation Loss, a novel plug-in loss function that improves the performance of various object detectors by directly optimizing correlation coefficients: E.g., Correlation Loss on Sparse R-CNN, an NMS-free method, yields 1.6 AP gain on COCO and 1.8 AP gain on Cityscapes dataset. Our best model on Sparse R-CNN reaches 51.0 AP without test-time augmentation on COCO test-dev, reaching state-of-the-art. Code is available at: https://github.com/fehmikahraman/CorrLoss. Fehmi Kahraman, Kemal Oksuz, Sinan Kalkan, Emre Akbas |
AAAI | 4 |
| 2023 | HoughNet: Integrating Near and Long-Range Evidence for Visual DetectionabstractThis paper presents HoughNet, a one-stage, anchor-free, voting-based, bottom-up object detection method. Inspired by the Generalized Hough Transform, HoughNet determines the presence of an object at a certain location by the sum of the votes cast on that location. Votes are collected from both near and long-distance locations based on a log-polar vote field. Thanks to this voting mechanism, HoughNet is able to integrate both near and long-range, class-conditional evidence for visual recognition, thereby generalizing and enhancing current object detection methodology, which typically relies on only local evidence. On the COCO dataset, HoughNet's best model achieves 46.4$AP$(and 65.1$AP_{50}$), performing on par with the state-of-the-art in bottom-up object detection and outperforming most major one-stage and two-stage methods. We further validate the effectiveness of our proposal in other visual detection tasks, namely, video object detection, instance segmentation, 3D object detection and keypoint detection for human pose estimation, and an additional “labels to photo” image generation task, where the integration of our voting module consistently improves performance in all cases. Code is available athttps://github.com/nerminsamet/houghnet. Nermin Samet, Samet Hicsonmez, Emre Akbas |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2022 | Streaming Multiscale Deep Equilibrium Models
Can Ufuk Ertenli, Emre Akbas, Ramazan Gokberk Cinbis |
ECCV (11) | 2 |
| 2022 | Does depth estimation help object detection?
Bedrettin Çetinkaya, Sinan Kalkan, Emre Akbas |
Image Vis. Comput. | 3 |
| 2022 | One Metric to Measure Them All: Localisation Recall Precision (LRP) for Evaluating Visual Detection TasksabstractDespite being widely used as a performance measure for visual detection tasks, Average Precision (AP) is limited in (i) reflecting localisation quality, (ii) interpretability and (iii) robustness to the design choices regarding its computation, and its applicability to outputs without confidence scores. Panoptic Quality (PQ), a measure proposed for evaluating panoptic segmentation (Kirillov et al., 2019), does not suffer from these limitations but is limited to panoptic segmentation. In this paper, we propose Localisation Recall Precision (LRP) Error as the average matching error of a visual detector computed based on both its localisation and classification qualities for a given confidence score threshold. LRP Error, initially proposed only for object detection by Oksuz et al. (2018), does not suffer from the aforementioned limitations and is applicable to all visual detection tasks. We also introduce Optimal LRP (oLRP) Error as the minimum LRP Error obtained over confidence scores to evaluate visual detectors and obtain optimal thresholds for deployment. We provide a detailed comparative analysis of LRP Error with AP and PQ, and use nearly 100 state-of-the-art visual detectors from seven visual detection tasks (i.e. object detection, keypoint detection, instance segmentation, panoptic segmentation, visual relationship detection, zero-shot detection and generalised zero-shot detection) using ten datasets to empirically show that LRP Error provides richer and more discriminative information than its counterparts. Code available at: https://github.com/kemaloksuz/LRP-Error. Kemal Oksuz, Baris Can Cam, Sinan Kalkan, Emre Akbas |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2021 | Mask-aware IoU for Anchor Assignment in Real-time Instance Segmentation
Kemal Oksuz, Baris Can Cam, Fehmi Kahraman, Zeynep Sonat Baltaci, Sinan Kalkan, Emre Akbas |
BMVC | 6 |
| 2021 | Rank & Sort Loss for Object Detection and Instance SegmentationabstractWe propose Rank & Sort (RS) Loss, a ranking-based loss function to train deep object detection and instance segmentation methods (i.e. visual detectors). RS Loss supervises the classifier, a sub-network of these methods, to rank each positive above all negatives as well as to sort positives among themselves with respect to (wrt.) their localisation qualities (e.g. Intersection-over-Union - IoU). To tackle the non-differentiable nature of ranking and sorting, we reformulate the incorporation of error-driven update with back-propagation as Identity Update, which enables us to model our novel sorting error among positives. With RS Loss, we significantly simplify training: (i) Thanks to our sorting objective, the positives are prioritized by the classifier without an additional auxiliary head (e.g. for centerness, IoU, mask-IoU), (ii) due to its ranking-based nature, RS Loss is robust to class imbalance, and thus, no sampling heuristic is required, and (iii) we address the multi-task nature of visual detectors using tuning-free task-balancing coefficients. Using RS Loss, we train seven diverse visual detectors only by tuning the learning rate, and show that it consistently outperforms baselines: e.g. our RS Loss improves (i) Faster R-CNN by ∼ 3 box AP and aLRP Loss (ranking-based baseline) by ∼ 2 box AP on COCO dataset, (ii) Mask R-CNN with repeat factor sampling (RFS) by 3.5 mask AP (∼ 7 AP for rare classes) on LVIS dataset; and also outperforms all counterparts. Code is available at: https://github.com/kemaloksuz/RankSortLoss. Kemal Oksuz, Baris Can Cam, Emre Akbas, Sinan Kalkan |
ICCV | 3 |
| 2021 | Adversarial Segmentation Loss For Sketch ColorizationabstractWe introduce a new method for generating color images from sketches or edge maps. Current methods either require some form of additional user-guidance or are limited to the “paired” translation approach. We argue that segmentation information could provide valuable guidance for sketch colorization. To this end, we propose to leverage semantic image segmentation, as provided by a general purpose panoptic segmentation network, to create an additional adversarial loss function. Our loss function can be integrated to any baseline GAN model. Our method is not limited to datasets that contain segmentation labels, and it can be trained for “unpaired” translation tasks. We show the effectiveness of our method on four different datasets spanning scene level indoor, outdoor, and children book illustration images using qualitative, quantitative and user study analysis. Our model improves its baseline up to 35 points on the FID metric. Our code and pretrained models can be found at https://github.com/giddyyupp/AdvSegLoss. Samet Hicsonmez, Nermin Samet, Emre Akbas, Pinar Duygulu |
ICIP | 3 |
| 2021 | HPRNet: Hierarchical point regression for whole-body human pose estimation
Nermin Samet, Emre Akbas |
Image Vis. Comput. | 2 |
| 2021 | Imbalance Problems in Object Detection: A ReviewabstractIn this paper, we present a comprehensive review of the imbalance problems in object detection. To analyze the problems in a systematic manner, we introduce a problem-based taxonomy. Following this taxonomy, we discuss each problem in depth and present a unifying yet critical perspective on the solutions in the literature. In addition, we identify major open issues regarding the existing imbalance problems as well as imbalance problems that have not been discussed before. Moreover, in order to keep our review up to date, we provide an accompanying webpage which catalogs papers addressing imbalance problems, according to our problem-based taxonomy. Researchers can track newer studies on this webpage available at: https://github.com/kemaloksuz/ObjectDetectionImbalance. Kemal Oksuz, Baris Can Cam, Sinan Kalkan, Emre Akbas |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2020 | Reducing Label Noise in Anchor-Free Object Detection
Nermin Samet, Samet Hicsonmez, Emre Akbas |
BMVC | 3 |
| 2020 | HoughNet: Integrating Near and Long-Range Evidence for Bottom-Up Object Detection
Nermin Samet, Samet Hicsonmez, Emre Akbas |
ECCV (25) | 3 |
| 2020 | A Ranking-based, Balanced Loss Function Unifying Classification and Localisation in Object DetectionabstractWe propose average Localisation-Recall-Precision (aLRP), a unified, bounded, balanced and ranking-based loss function for both classification and localisation tasks in object detection. aLRP extends the Localisation-Recall-Precision (LRP) performance metric (Oksuz et al., 2018) inspired from how Average Precision (AP) Loss extends precision to a ranking-based loss function for classification (Chen et al., 2020). aLRP has the following distinct advantages: (i) aLRP is the first ranking-based loss function for both classification and localisation tasks. (ii) Thanks to using ranking for both tasks, aLRP naturally enforces high-quality localisation for high-precision classification. (iii) aLRP provides provable balance between positives and negatives. (iv) Compared to on average ~6 hyperparameters in the loss functions of state-of-the-art detectors, aLRP Loss has only one hyperparameter, which we did not tune in practice. On the COCO dataset, aLRP Loss improves its ranking-based predecessor, AP Loss, up to around 5 AP points, achieves 48.9 AP without test time augmentation and outperforms all one-stage detectors. Code available at: https://github.com/kemaloksuz/aLRPLoss . Kemal Oksuz, Baris Can Cam, Emre Akbas, Sinan Kalkan |
NeurIPS | 3 |
| 2020 | Generating Positive Bounding Boxes for Balanced Training of Object DetectorsabstractTwo-stage deep object detectors generate a set of regions-of-interest RoIs in the first stage, then, in the second stage, identify objects among the proposed RoIs that sufficiently overlap with a ground truth (GT) box. The second stage is known to suffer from a bias towards RoIs that have low intersection-over-union (IoU) with the associated GT boxes. To address this issue, we first propose a sampling method to generate bounding boxes (BB) that overlap with a given reference box more than a given IoU threshold. Then, we use this BB generation method to develop a positive RoI (pRoI) generator that, for the second stage, produces RoIs following any desired spatial or IoU distribution. We show that our pRoI generator is able to simulate other sampling methods for positive examples such as hard example mining and prime sampling. Using our generator as an analysis tool, we show that (i) IoU imbalance has an adverse effect on performance, (ii) hard positive example mining improves the performance only for certain input IoU distributions, and (iii) the imbalance among the foreground classes has an adverse effect on performance and that it can be alleviated at the batch level. Finally, we train Faster R-CNN using our pRoI generator and, compared to conventional training, obtain better or on-par performance for low IoUs and significant improvements when trained for higher IoUs for Pascal VOC and MS COCO datasets. The code is available at: https://github.com/kemaloksuz/BoundingBoxGenerator. Kemal Oksuz, Baris Can Cam, Emre Akbas, Sinan Kalkan |
WACV | 3 |
| 2020 | Low-level multiscale image segmentation and a benchmark for its evaluation
Emre Akbas, Narendra Ahuja |
Comput. Vis. Image Underst. | 1 |
| 2020 | GANILLA: Generative adversarial networks for image to illustration translation
Samet Hicsonmez, Nermin Samet, Emre Akbas, Pinar Duygulu |
Image Vis. Comput. | 3 |
| 2019 | Self-Supervised Learning of 3D Human Pose Using Multi-View GeometryabstractTraining accurate 3D human pose estimators requires large amount of 3D ground-truth data which is costly to collect. Various weakly or self supervised pose estimation methods have been proposed due to lack of 3D data. Nevertheless, these methods, in addition to 2D ground-truth poses, require either additional supervision in various forms (e.g. unpaired 3D ground truth data, a small subset of labels) or the camera parameters in multiview settings. To address these problems, we present EpipolarPose, a self-supervised learning method for 3D human pose estimation, which does not need any 3D ground-truth data or camera extrinsics. During training, EpipolarPose estimates 2D poses from multi-view images, and then, utilizes epipolar geometry to obtain a 3D pose and camera geometry which are subsequently used to train a 3D pose estimator. We demonstrate the effectiveness of our approach on standard benchmark datasets (i.e. Human3.6M and MPI-INF-3DHP) where we set the new state-of-the-art among weakly/self-supervised methods. Furthermore, we propose a new performance measure Pose Structure Score (PSS) which is a scale invariant, structure aware measure to evaluate the structural plausibility of a pose with respect to its ground truth. Code and pretrained models are available at https://github.com/mkocabas/EpipolarPose. Muhammed Kocabas, Salih Karagoz, Emre Akbas |
CVPR | 3 |
| 2018 | MultiPoseNet: Fast Multi-Person Pose Estimation Using Pose Residual Network
Muhammed Kocabas, Salih Karagoz, Emre Akbas |
ECCV (11) | 3 |
| 2018 | Localization Recall Precision (LRP): A New Performance Metric for Object Detection
Kemal Oksuz, Baris Can Cam, Emre Akbas, Sinan Kalkan |
ECCV (7) | 3 |
| 2017 | Object detection through search with a foveated visual systemabstractHumans and many other species sense visual information with varying spatial resolution across the visual field (foveated vision) and deploy eye movements to actively sample regions of interests in scenes. The advantage of such varying resolution architecture is a reduced computational, hence metabolic cost. But what are the performance costs of such processing strategy relative to a scheme that processes the visual field at high spatial resolution? Here we first focus on visual search and combine object detectors from computer vision with a recent model of peripheral pooling regions found at the V1 layer of the human visual system. We develop a foveated object detector that processes the entire scene with varying resolution, uses retino-specific object detection classifiers to guide eye movements, aligns its fovea with regions of interest in the input image and integrates observations across multiple fixations. We compared the foveated object detector against a non-foveated version of the same object detector which processes the entire image at homogeneous high spatial resolution. We evaluated the accuracy of the foveated and non-foveated object detectors identifying 20 different objects classes in scenes from a standard computer vision data set (the PASCAL VOC 2007 dataset). We show that the foveated object detector can approximate the performance of the object detector with homogeneous high spatial resolution processing while bringing significant computational cost savings. Additionally, we assessed the impact of foveation on the computation of bottom-up saliency. An implementation of a simple foveated bottom-up saliency model with eye movements showed agreement in the selection of top salient regions of scenes with those selected by a non-foveated high resolution saliency model. Together, our results might help explain the evolution of foveated visual systems with eye movements as a solution that preserves perceptual performance in visual search while resulting in computational and metabolic savings to the brain. Emre Akbas, Miguel P. Eckstein |
PLoS Comput. Biol. | 1 |
| 2014 | Low-Level Hierarchical Multiscale Segmentation Statistics of Natural ImagesabstractThis paper is aimed at obtaining the statistics as a probabilistic model pertaining to the geometric, topological and photometric structure of natural images. The image structure is represented by its segmentation graph derived from the low-level hierarchical multiscale image segmentation. We first estimate the statistics of a number of segmentation graph properties from a large number of images. Our estimates confirm some findings reported in the past work, as well as provide some new ones. We then obtain a Markov random field based model of the segmentation graph which subsumes the observed statistics. To demonstrate the value of the model and the statistics, we show how its use as a prior impacts three applications: image classification, semantic image segmentation and object detection. Emre Akbas, Narendra Ahuja |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2010 | Pedestrian Recognition with a Learned Metric
Mert Dikmen, Emre Akbas, Thomas S. Huang, Narendra Ahuja |
ACCV (4) | 2 |
| 2010 | Low-Level Image Segmentation Based Scene ClassificationabstractThis paper is aimed at evaluating the semantic information content of multiscale, low-level image segmentation. As a method of doing this, we use selected features of segmentation for semantic classification of real images. To estimate the relative measure of the information content of our features, we compare the results of classifications we obtain using them with those obtained by others using the commonly used patch/grid based features. To classify an image using segmentation based features, we model the image in terms of a probability density function, a Gaussian mixture model (GMM) to be specific, of its region features. This GMM is fit to the image by adapting a universal GMM which is estimated so it fits all images. Adaptation is done using a maximum-aposteriori criterion. We use kernelized versions of Bhattacharyya distance to measure the similarity between two GMMs and support vector machines to perform classification. We outperform previously reported results on a publicly available scene classification dataset. These results suggest further experimentation in evaluating the promise of low level segmentation in image classification. Emre Akbas, Narendra Ahuja |
ICPR | 1 |
| 2009 | From Ramp Discontinuities to Segmentation Tree
Emre Akbas, Narendra Ahuja |
ACCV (1) | 1 |
| 2007 | Automatic Image Annotation by Ensemble of Visual DescriptorsabstractAutomatic image annotation systems available in the literature concatenate color, texture and/or shape features in a single feature vector to learn a set of high level semantic categories using a single learning machine. This approach is quite naive to map the visual features to high level semantic information concerning the categories. Concatenation of many features with different visual properties and wide dynamical ranges may result in curse of dimensionality and redundancy problems. Additionally, it usually requires normalization which may cause an undesirable distortion in the feature space. An elegant way of reducing the effects of these problems is to design a dedicated feature space for each image category, depending on its content, and learn a range of visual properties of the whole image from a variety of feature sets. For this purpose, a two-layer ensemble learning system, called Supervised Annotation by Descriptor Ensemble (SADE), is proposed. SADE, initially, extracts a variety of low-level visual descriptors from the image. Each descriptor is, then, fed to a separate learning machine in the first layer. Finally, the meta-layer classifier is trained on the output of the first layer classifiers and the images are annotated by using the decision of the meta-layer classifier. This approach not only avoids normalization, but also reduces the effects of dimensional curse and redundancy. The proposed system outperforms a state-of-the-art automatic image annotation system, in an equivalent experimental setup. Emre Akbas, Fatos T. Yarman-Vural |
CVPR | 1 |
| 2006 | A Hierarchical Classification System Based on Adaptive Resonance TheoryabstractIn this study, we propose a hierarchical classification system, which emulates the eye-brain channel in two hierarchical layers. In the first layer, a set of classifiers are trained by using low level, low dimensional features. In the second layer, the recognition results of the first layer are fed to the Fuzzy ARTMAP (FAM) classifier which implements the Adaptive Resonance Theory. Experiments indicate that the hierarchical approach proposed in this paper, increases the classification performances compared to the available methods. Mutlu Uysal, Emre Akbas, Fatos T. Yarman-Vural |
ICIP | 2 |