EDBT 2026 Demo / reviewers in the wild / expert
Sinan Kalkan
dblp:62/2714
· DBLP profile ↗
54ranked-venue papers
1as first author
31since 2021 · last 2026
0000-0003-0915-5917ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 40 · 1 first-author · 23 since 2021Graphics, computer vision, multimedia, augmented reality and games · 29 · 1 first-author · 17 since 2021Systems, architecture and hardware · 6 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Investigating Bias and Fairness in Appearance-based Gaze EstimationabstractWhile appearance-based gaze estimation has achieved significant improvements in accuracy and domain adaptation, the fairness of these systems across different demographic groups remains largely unexplored. To date, there is no comprehensive benchmark quantifying algorithmic bias in gaze estimation. This paper presents the first extensive evaluation of fairness in appearance-based gaze estimation, focusing on ethnicity and gender attributes. We establish a fairness baseline by analyzing state-of-the-art models using standard fairness metrics, revealing significant performance disparities. Furthermore, we evaluate the effectiveness of existing bias mitigation strategies when applied to the gaze domain and show that their fairness contributions are limited. We summarize key insights and open issues. Overall, our work calls for research into developing robust, equitable gaze estimators. To support future research and reproducibility, we publicly release our annotations, code, and trained models at: github.com/ akgulburak/gaze-estimation-fairness. Burak Akgül, Erol Sahin, Sinan Kalkan |
FG | 3 |
| 2026 | Gaze4HRI: Zero-shot Benchmarking Gaze Estimation Neural-Networks for Human-Robot InteractionabstractWhile zero-shot appearance-based 3D gaze estimation offers significant cost-efficiency by directly mapping RGB images to gaze vectors, its reliability in Human-Robot Interaction (HRI) settings remains uncertain. Existing benchmarks frequently overlook fundamental HRI conditions, such as dynamic camera viewpoints and moving targets in video. Furthermore, current cross-dataset evaluations often suffer from a complexity gap, where methods trained on diverse datasets are tested on significantly smaller and less varied sets, failing to assess true robustness. To bridge these gaps, we introduce Gaze4HRI, a large-scale dataset (50+ subjects, 3,000+ videos, 600,000+ frames) designed to evaluate state-of-the-art performance against critical HRI variables: illumination, headgaze conflict, as well as the motion of camera and gaze target in video. Our benchmark reveals that all evaluated methods fail in at least one condition, identifying steeply-downward gaze as a universal failure point. Notably, PureGaze trained on the ETH-X-Gaze dataset uniquely maintains resilience across all other conditions. These results challenge the recent focus in the literature on complex spatial-temporal modeling and Transformer-based architectures. Instead, our findings suggest that extensive data diversity, as exemplified by the ETH-X-Gaze dataset, serves as the primary driver of zero-shot robustness in unconstrained environments, while resilienceenhancing frameworks, such as PureGaze's self-adversarial loss for gaze feature purification, provide a substantial further improvement. Ultimately, this study establishes a rigorous benchmark that provides practical guidelines for practitioners as well as reshaping future research. The dataset and codes are available at https://gazeforhri.github.io. Berk Sezer, Ali Görkem Küçük, Erol Sahin, Sinan Kalkan |
FG | 4 |
| 2026 | PDV: Prompt Directional Vectors for Zero-shot Composed Image RetrievalabstractZero-shot Composed Image Retrieval (ZS-CIR) enables image search using a reference image and a text prompt without requiring specialized text-image composition networks trained on large-scale paired data. However, current ZS-CIR approaches suffer from three critical limitations in their reliance on composed text embeddings: static query embedding representations, insufficient utilization of image embeddings, and suboptimal performance when fusing text and image embeddings. To address these challenges, we introduce the Prompt Directional Vector (PDV), a simple yet effective training-free enhancement that captures semantic modifications induced by user prompts. PDV enables three key improvements: (1) Dynamic composed text embeddings where prompt adjustments are controllable via a scaling factor, (2) composed image embeddings through semantic transfer from text prompts to image features, and (3) weighted fusion of composed text and image embeddings that enhances retrieval by balancing visual and semantic similarity. Our approach serves as a plug-and-play enhancement for existing ZS-CIR methods with minimal computational overhead. Extensive experiments across multiple benchmarks demonstrate that PDV consistently improves retrieval performance when integrated with state-of-the-art ZS-CIR approaches, particularly for methods that generate accurate compositional embeddings. The code will be released upon publication. Osman Tursun, Sinan Kalkan, Simon Denman, Clinton Fookes |
WACV | 2 |
| 2026 | Intrinsic dimensionality as a model-free measure of class imbalance
Çagri Eser, Zeynep Sonat Baltaci, Emre Akbas, Sinan Kalkan |
Neurocomputing | 4 |
| 2026 | ms-mamba: Multi-scale mamba for time-series forecasting
Yusuf Meric Karadag, Ismail Talaz, Ipek Gursel Dino, Sinan Kalkan |
Neurocomputing | 4 |
| 2025 | CorDis: A Novel Correlation-Based Disentanglement MeasureabstractDisentangled representation learning aims to decompose images into meaningful independent factors of variation. However, measuring the extent of disentanglement remains a challenge. Available measures, such as β-VAE, MIG, SAP score or Explicitness score rely on classifier assumptions, sampling schemes, or mutual information estimators, introducing biases and dependencies to the model. In this paper, we present a novel correlation-based measure, CorDis, which mitigates these dependencies while preserving robust and interpretable insights into disentanglement. We systematically compare CorDis with existing measures. Experimental results demonstrate that CorDis provides a more principled and assumption-light approach to measuring the amount of disentanglement, contributing to the development of universal and consistent benchmarks. Hazal Mogultay Ozcan, Sinan Kalkan, Fatos T. Yarman-Vural |
ICIP | 2 |
| 2025 | Transfer learning and parameter-efficient fine-tuning for heating energy consumption prediction using urban building energy models (UBEM)
Ilkim Canli, Yusuf Meric Karadag, Sevval Ucar, Ismail Talaz, Fatma Ece Gursoy, Ipek Gursel Dino, Sinan Kalkan |
Adv. Eng. Informatics | 7 |
| 2025 | FanNet: A mesh convolution operator for learning dense maps
Günes Sucu, Sinan Kalkan, Yusuf Sahillioglu |
Comput. Graph. | 2 |
| 2025 | Multimodal multimedia information retrieval through the integration of fuzzy clustering, OWA-based fusion, and Siamese neural networks
Saeid Sattari, Sinan Kalkan, Adnan Yazici |
Fuzzy Sets Syst. | 2 |
| 2025 | Generalized variational autoencoders for learning disentangled representation
Hazal Mogultay Ozcan, Sinan Kalkan, Fatos T. Yarman-Vural |
Neurocomputing | 2 |
| 2025 | L-VAE: variational auto-encoder with learnable beta for disentangled representation
Hazal Mogultay Ozcan, Sinan Kalkan, Fatos T. Yarman-Vural |
Mach. Vis. Appl. | 2 |
| 2024 | RankED: Addressing Imbalance and Uncertainty in Edge Detection Using Ranking-based LossesabstractDetecting edges in images suffers from the problems of (PI) heavy imbalance between positive and negative classes as well as (P2) label uncertainty owing to disagreement be-tween different annotators. Existing solutions address PI using class-balanced cross-entropy loss and dice loss and P2 by only predicting edges agreed upon by most annota-tors. In this paper, we propose RankED, a unified ranking-based approach that addresses both the imbalance problem (P 1) and the uncertainty problem (P2). RankED tackles these two problems with two components: One component which ranks positive pixels over negative pixels, and the second which promotes high confidence edge pixels to have more label certainty. We show that RankED outperforms previous studies and sets a new state-of-the-art on NYUD-v2, BSDS500 and Multi-cue datasets. Code is available at https://ranked-cvpr24.github.io. Bedrettin Çetinkaya, Sinan Kalkan, Emre Akbas |
CVPR | 2 |
| 2024 | Bucketed Ranking-Based Losses for Efficient Training of Object Detectors
Feyza Yavuz, Baris Can Cam, Adnan Harun Dogan, Kemal Oksuz, Emre Akbas, Sinan Kalkan |
ECCV (59) | 6 |
| 2024 | FairReFuse: Referee-Guided Fusion for Multi-Modal Causal Fairness in Depression Detection
Jiaee Cheong, Sinan Kalkan, Hatice Gunes |
IJCAI | 2 |
| 2024 | BaSeNet: A Learning-based Mobile Manipulator Base Pose Sequence Planning for Pickup TasksabstractIn many applications, a mobile manipulator robot is required to grasp a set of objects distributed in space. This may not be feasible from a single base pose and the robot must plan the sequence of base poses for grasping all objects, minimizing the total navigation and grasping time. This is a Combinatorial Optimization problem that can be solved using exact methods, which provide optimal solutions but are computationally expensive, or approximate methods, which offer computationally efficient but sub-optimal solutions. Recent studies have shown that learning-based methods can solve Combinatorial Optimization problems, providing near-optimal and computationally efficient solutions.In this work, we present BaSeNet - a learning-based approach to plan the sequence of base poses for the robot to grasp all the objects in the scene. We propose a Reinforcement Learning based solution that learns the base poses for grasping individual objects and the sequence in which the objects should be grasped to minimize the total navigation and grasping costs using Layered Learning. As the problem has a varying number of states and actions, we represent states and actions as a graph and use Graph Neural Networks for learning. We show that the proposed method can produce comparable solutions to exact and approximate methods with significantly less computation time. The code and Reinforcement Learning environments will be made available on the project webpage*. Lakshadeep Naik, Sinan Kalkan, Sune Lundø Sørensen, Mikkel Baun Kjærgaard, Norbert Krüger |
IROS | 2 |
| 2024 | Uncertainty as a Fairness MeasureabstractUnfair predictions of machine learning (ML) models impede their broad acceptance in real-world settings. Tackling this arduous challenge first necessitates defining what it means for an ML model to be fair. This has been addressed by the ML community with various measures of fairness that depend on the prediction outcomes of the ML models, either at the group-level or the individual-level. These fairness measures are limited in that they utilize point predictions, neglecting their variances, or uncertainties, making them susceptible to noise, missingness and shifts in data. In this paper, we first show that a ML model may appear to be fair with existing point-based fairness measures but biased against a demographic group in terms of prediction uncertainties. Then, we introduce new fairness measures based on different types of uncertainties, namely, aleatoric uncertainty and epistemic uncertainty. We demonstrate on many datasets that (i) our uncertaintybased measures are complementary to existing measures of fairness, and (ii) they provide more insights about the underlying issues leading to bias. Selim Kuzucu, Jiaee Cheong, Hatice Gunes, Sinan Kalkan |
J. Artif. Intell. Res. | 4 |
| 2023 | Correlation Loss: Enforcing Correlation between Classification and LocalizationabstractObject detectors are conventionally trained by a weighted sum of classification and localization losses. Recent studies (e.g., predicting IoU with an auxiliary head, Generalized Focal Loss, Rank & Sort Loss) have shown that forcing these two loss terms to interact with each other in non-conventional ways creates a useful inductive bias and improves performance. Inspired by these works, we focus on the correlation between classification and localization and make two main contributions: (i) We provide an analysis about the effects of correlation between classification and localization tasks in object detectors. We identify why correlation affects the performance of various NMS-based and NMS-free detectors, and we devise measures to evaluate the effect of correlation and use them to analyze common detectors. (ii) Motivated by our observations, e.g., that NMS-free detectors can also benefit from correlation, we propose Correlation Loss, a novel plug-in loss function that improves the performance of various object detectors by directly optimizing correlation coefficients: E.g., Correlation Loss on Sparse R-CNN, an NMS-free method, yields 1.6 AP gain on COCO and 1.8 AP gain on Cityscapes dataset. Our best model on Sparse R-CNN reaches 51.0 AP without test-time augmentation on COCO test-dev, reaching state-of-the-art. Code is available at: https://github.com/fehmikahraman/CorrLoss. Fehmi Kahraman, Kemal Oksuz, Sinan Kalkan, Emre Akbas |
AAAI | 3 |
| 2023 | Towards Gender Fairness for Mental Health PredictionabstractMental health is becoming an increasingly prominent health challenge. Despite a plethora of studies analysing and mitigating bias for a variety of tasks such as face recognition and credit scoring, research on machine learning (ML) fairness for mental health has been sparse to date. In this work, we focus on gender bias in mental health and make the following contributions. First, we examine whether bias exists in existing mental health datasets and algorithms. Our experiments were conducted using Depresjon, Psykose and D-Vlog. We identify that both data and algorithmic bias exist. Second, we analyse strategies that can be deployed at the pre-processing, in-processing and post-processing stages to mitigate for bias and evaluate their effectiveness. Third, we investigate factors that impact the efficacy of existing bias mitigation strategies and outline recommendations to achieve greater gender fairness for mental health. Upon obtaining counter-intuitive results on D-Vlog dataset, we undertake further experiments and analyses, and provide practical suggestions to avoid hampering bias mitigation efforts in ML for mental health. Jiaee Cheong, Selim Kuzucu, Sinan Kalkan, Hatice Gunes |
IJCAI | 3 |
| 2023 | TMO-Det: Deep tone-mapping optimized with and for object detection
Ismail Hakki Kocdemir, Alper Koz, Ahmet Oguz Akyüz, Alan Chalmers, A. Aydin Alatan, Sinan Kalkan |
Pattern Recognit. Lett. | 6 |
| 2022 | Segment Augmentation and Differentiable Ranking for Logo RetrievalabstractLogo retrieval is a challenging problem since the definition of similarity is more subjective than image retrieval, and the set of known similarities is very scarce. In this paper, to tackle this challenge, we propose a simple but effective segment-based augmentation strategy to introduce artificially similar logos for training deep networks for logo retrieval. In this novel augmentation strategy, we first find segments in a logo and apply transformations such as rotation, scaling, and color change, on the segments, unlike the conventional strategies that perform augmentation at the image level. Moreover, we evaluate suitability of using ranking-based losses (namely Smooth-AP) for learning similarity for logo retrieval. On the METU and the LLD datasets, we show that (i) our segment-based augmentation strategy improves retrieval performance compared to the baseline model or image-level augmentation strategies, and (ii) Smooth-AP indeed performs better than conventional losses for logo retrieval. Feyza Yavuz, Sinan Kalkan |
ICPR | 2 |
| 2022 | AssembleRL: Learning to Assemble Furniture from Their Point CloudsabstractThe rise of simulation environments has enabled learning-based approaches for assembly planning, which is otherwise a labor-intensive and daunting task. Assembling furniture is especially interesting since furniture are intricate and pose challenges for learning-based approaches. Surprisingly, humans can solve furniture assembly mostly given a 2D snapshot of the assembled product. Although recent years have witnessed promising learning-based approaches for furniture assembly, they assume the availability of correct connection labels for each assembly step, which are expensive to obtain in practice. In this paper, we alleviate this assumption and aim to solve furniture assembly with as little human expertise and supervision as possible. To be specific, we assume the availability of the assembled point cloud, and comparing the point cloud of the current assembly and the point cloud of the target product, obtain a novel reward signal based on two measures: Incorrectness and incompleteness. We show that our novel reward signal can train a deep network to successfully assemble different types of furniture. Code and networks available here: https://github.com/METU-KALFA/AssembleRL. Özgür Aslan, Burak Bolat, Batuhan Bal, Tugba Tümer, Erol Sahin, Sinan Kalkan |
IROS | 6 |
| 2022 | Object Detection for Autonomous Driving: High-Dynamic Range vs. Low-Dynamic Range ImagesabstractAn important problem in autonomous driving is to perceive objects even under challenging illumination conditions. Despite this problem, existing solutions use low-dynamic range (LDR) images for object detection for autonomous driving. In this paper, we provide a novel analysis on whether high-dynamic range (HDR) images can provide better performance for object detection for autonomous driving. To this end, we choose a seminal deep object detector and systematically evaluate its performance when trained with (i) LDR images, (ii) HDR images, and (iii) tone-mapped LDR images for scenes with different illuminations. We show that a detector with HDR images pre-processed with normalization and gamma correction can only marginally perform better than a detector with LDR or tone-mapped LDR images. Our analysis of this unexpected finding reveals that a detector with HDR images requires significantly more samples as the space of HDR images is significantly larger than that of LDR images. Ismail Hakki Kocdemir, Ahmet Oguz Akyüz, Alper Koz, Alan Chalmers, A. Aydin Alatan, Sinan Kalkan |
MMSP | 6 |
| 2022 | Vision-based estimation of the number of occupants using video cameras
Ipek Gursel Dino, M. Esat Kalfaoglu, Orcun Koral Iseri, Bilge Erdogan, Sinan Kalkan, A. Aydin Alatan |
Adv. Eng. Informatics | 5 |
| 2022 | Does depth estimation help object detection?
Bedrettin Çetinkaya, Sinan Kalkan, Emre Akbas |
Image Vis. Comput. | 2 |
| 2022 | Hand-crafted versus learned representations for audio event detection
Selver Ezgi Küçükbay, Adnan Yazici, Sinan Kalkan |
Multim. Tools Appl. | 3 |
| 2022 | One Metric to Measure Them All: Localisation Recall Precision (LRP) for Evaluating Visual Detection TasksabstractDespite being widely used as a performance measure for visual detection tasks, Average Precision (AP) is limited in (i) reflecting localisation quality, (ii) interpretability and (iii) robustness to the design choices regarding its computation, and its applicability to outputs without confidence scores. Panoptic Quality (PQ), a measure proposed for evaluating panoptic segmentation (Kirillov et al., 2019), does not suffer from these limitations but is limited to panoptic segmentation. In this paper, we propose Localisation Recall Precision (LRP) Error as the average matching error of a visual detector computed based on both its localisation and classification qualities for a given confidence score threshold. LRP Error, initially proposed only for object detection by Oksuz et al. (2018), does not suffer from the aforementioned limitations and is applicable to all visual detection tasks. We also introduce Optimal LRP (oLRP) Error as the minimum LRP Error obtained over confidence scores to evaluate visual detectors and obtain optimal thresholds for deployment. We provide a detailed comparative analysis of LRP Error with AP and PQ, and use nearly 100 state-of-the-art visual detectors from seven visual detection tasks (i.e. object detection, keypoint detection, instance segmentation, panoptic segmentation, visual relationship detection, zero-shot detection and generalised zero-shot detection) using ten datasets to empirically show that LRP Error provides richer and more discriminative information than its counterparts. Code available at: https://github.com/kemaloksuz/LRP-Error. Kemal Oksuz, Baris Can Cam, Sinan Kalkan, Emre Akbas |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2021 | Mask-aware IoU for Anchor Assignment in Real-time Instance Segmentation
Kemal Oksuz, Baris Can Cam, Fehmi Kahraman, Zeynep Sonat Baltaci, Sinan Kalkan, Emre Akbas |
BMVC | 5 |
| 2021 | AULA-Caps: Lifecycle-Aware Capsule Networks for Spatio-Temporal Analysis of Facial ActionsabstractMost state-of-the-art approaches for Facial Action Unit (AU) detection rely on evaluating static frames, encoding a snapshot of heightened facial activity. In real-world interactions, however, facial expressions are more subtle and evolve over time requiring AU detection models to learn spatial as well as temporal information. In this work, we focus on both spatial and spatio-temporal features encoding the temporal evolution of facial AU activation. We propose the Action Unit Lifecycle-Aware Capsule Network (AULA-Caps) for AU detection using both frame and sequence-level features. While, at the frame-level, the capsule layers of AULA-Caps learn spatial feature primitives to determine AU activations, at the sequence-level, it learns temporal dependencies between contiguous frames by focusing on relevant spatio-temporal segments in the sequence. The learnt feature capsules are routed together such that the model learns to selectively focus on spatial or spatio-temporal information depending upon the AU lifecycle. The proposed model is evaluated on popular benchmarks, namely BP4D and GFT datasets, obtaining state-of-the-art results for both. Nikhil Churamani, Sinan Kalkan, Hatice Gunes |
FG | 2 |
| 2021 | Rank & Sort Loss for Object Detection and Instance SegmentationabstractWe propose Rank & Sort (RS) Loss, a ranking-based loss function to train deep object detection and instance segmentation methods (i.e. visual detectors). RS Loss supervises the classifier, a sub-network of these methods, to rank each positive above all negatives as well as to sort positives among themselves with respect to (wrt.) their localisation qualities (e.g. Intersection-over-Union - IoU). To tackle the non-differentiable nature of ranking and sorting, we reformulate the incorporation of error-driven update with back-propagation as Identity Update, which enables us to model our novel sorting error among positives. With RS Loss, we significantly simplify training: (i) Thanks to our sorting objective, the positives are prioritized by the classifier without an additional auxiliary head (e.g. for centerness, IoU, mask-IoU), (ii) due to its ranking-based nature, RS Loss is robust to class imbalance, and thus, no sampling heuristic is required, and (iii) we address the multi-task nature of visual detectors using tuning-free task-balancing coefficients. Using RS Loss, we train seven diverse visual detectors only by tuning the learning rate, and show that it consistently outperforms baselines: e.g. our RS Loss improves (i) Faster R-CNN by ∼ 3 box AP and aLRP Loss (ranking-based baseline) by ∼ 2 box AP on COCO dataset, (ii) Mask R-CNN with repeat factor sampling (RFS) by 3.5 mask AP (∼ 7 AP for rare classes) on LVIS dataset; and also outperforms all counterparts. Code is available at: https://github.com/kemaloksuz/RankSortLoss. Kemal Oksuz, Baris Can Cam, Emre Akbas, Sinan Kalkan |
ICCV | 4 |
| 2021 | HDR Image Construction from Trifocal Multiexposure ImagesabstractWith the progress of autonomous vehicles, the sensing of the environment in more detail with higher dynamic ranges has become more important to classify surrounding objects and obstacles. While stereo HDR images for this purpose can provide advantages compared to the conventional LDR images, they suffer from limited dynamic ranges and spike-like noises due to the inaccuracies in disparity estimation. In this paper, we formulate the HDR image construction problem from trifocal multi-exposure images and develop a method which improves disparity estimation for better HDR image construction. Given the symmetric geometry of the trifocal setup, the proposed method uses the equivalence of disparities from middle to left and middle to right images to determine the reliable regions. The HDR radiance for the pixels in these reliable regions are estimated by using the weighted average of the warped images in different exposures and the middle image, whereas the radiance values outside the reliable regions are estimated by using only the middle image. The experiments with different exposure combinations for left, middle and right images reveal better performances of the proposed method compared to the stereo HDR imaging. It is also observed that the improvements are more apparent for larger disparities between the cameras. Alper Koz, Baris Demirkiliç, Yunus Bilge Kurt, Ahmet Oguz Akyüz, Sinan Kalkan, A. Aydin Alatan, Alan Chalmers |
MMSP | 5 |
| 2021 | Imbalance Problems in Object Detection: A ReviewabstractIn this paper, we present a comprehensive review of the imbalance problems in object detection. To analyze the problems in a systematic manner, we introduce a problem-based taxonomy. Following this taxonomy, we discuss each problem in depth and present a unifying yet critical perspective on the solutions in the literature. In addition, we identify major open issues regarding the existing imbalance problems as well as imbalance problems that have not been discussed before. Moreover, in order to keep our review up to date, we provide an accompanying webpage which catalogs papers addressing imbalance problems, according to our problem-based taxonomy. Researchers can track newer studies on this webpage available at: https://github.com/kemaloksuz/ObjectDetectionImbalance. Kemal Oksuz, Baris Can Cam, Sinan Kalkan, Emre Akbas |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2020 | Transformer-Encoder Detector Module: Using Context to Improve Robustness to Adversarial Attacks on Object DetectionabstractDeep neural network approaches have demonstrated high performance in object recognition (CNN) and detection (Faster-RCNN) tasks, but experiments have shown that such architectures are vulnerable to adversarial attacks (FFF, UAP): low amplitude perturbations, barely perceptible by the human eye, can lead to a drastic reduction in labelling performance. This article proposes a new context module, called Transformer-Encoder Detector Module, that can be applied to an object detector to (i) improve the labelling of object instances; and (ii) improve the detector's robustness to adversarial attacks. The proposed model achieves higher mAP, F1 scores and AUC average score of up to 13% compared to the baseline Faster-RCNN detector, and an mAP score 8 points higher on images subjected to FFF or UAP attacks due to the inclusion of both contextual and visual features extracted from scene and encoded into the model. The result demonstrates that a simple ad-hoc context module can improve the reliability of object detectors significantly. Faisal Alamri, Sinan Kalkan, Nicolas Pugeault |
ICPR | 2 |
| 2020 | A Ranking-based, Balanced Loss Function Unifying Classification and Localisation in Object DetectionabstractWe propose average Localisation-Recall-Precision (aLRP), a unified, bounded, balanced and ranking-based loss function for both classification and localisation tasks in object detection. aLRP extends the Localisation-Recall-Precision (LRP) performance metric (Oksuz et al., 2018) inspired from how Average Precision (AP) Loss extends precision to a ranking-based loss function for classification (Chen et al., 2020). aLRP has the following distinct advantages: (i) aLRP is the first ranking-based loss function for both classification and localisation tasks. (ii) Thanks to using ranking for both tasks, aLRP naturally enforces high-quality localisation for high-precision classification. (iii) aLRP provides provable balance between positives and negatives. (iv) Compared to on average ~6 hyperparameters in the loss functions of state-of-the-art detectors, aLRP Loss has only one hyperparameter, which we did not tune in practice. On the COCO dataset, aLRP Loss improves its ranking-based predecessor, AP Loss, up to around 5 AP points, achieves 48.9 AP without test time augmentation and outperforms all one-stage detectors. Code available at: https://github.com/kemaloksuz/aLRPLoss . Kemal Oksuz, Baris Can Cam, Emre Akbas, Sinan Kalkan |
NeurIPS | 4 |
| 2020 | Continual Learning for Affective Robotics: Why, What and How?abstractCreating and sustaining closed-loop dynamic and social interactions with humans require robots to continually adapt towards their users' behaviours, their affective states and moods while keeping them engaged in the task they are performing. Analysing, understanding and appropriately responding to human nonverbal behaviour and affective states are the central objectives of affective robotics research. Conventional machine learning approaches do not scale well to the dynamic nature of such real-world interactions as they require samples from stationary data distributions. The real-world is not stationary, it changes continuously. In such contexts, the training data and learning objectives may also change rapidly. Continual Learning (CL), by design, is able to address this very problem by learning incrementally. In this paper, we argue that CL is an essential paradigm for creating fully adaptive affective robots (why). To support this argument, we first provide an introduction to CL approaches and what they can offer for various dynamic (interactive) situations (what). We then formulate guidelines for the affective robotics community on how to utilise CL for perception and behaviour learning with adaptation (how). For each case, we reformulate the problem as a CL problem and outline a corresponding CL-based solution. We conclude the paper by highlighting the potential challenges to be faced and by providing specific recommendations on how to utilise CL for affective robotics. Nikhil Churamani, Sinan Kalkan, Hatice Gunes |
RO-MAN | 2 |
| 2020 | Generating Positive Bounding Boxes for Balanced Training of Object DetectorsabstractTwo-stage deep object detectors generate a set of regions-of-interest RoIs in the first stage, then, in the second stage, identify objects among the proposed RoIs that sufficiently overlap with a ground truth (GT) box. The second stage is known to suffer from a bias towards RoIs that have low intersection-over-union (IoU) with the associated GT boxes. To address this issue, we first propose a sampling method to generate bounding boxes (BB) that overlap with a given reference box more than a given IoU threshold. Then, we use this BB generation method to develop a positive RoI (pRoI) generator that, for the second stage, produces RoIs following any desired spatial or IoU distribution. We show that our pRoI generator is able to simulate other sampling methods for positive examples such as hard example mining and prime sampling. Using our generator as an analysis tool, we show that (i) IoU imbalance has an adverse effect on performance, (ii) hard positive example mining improves the performance only for certain input IoU distributions, and (iii) the imbalance among the foreground classes has an adverse effect on performance and that it can be alleviated at the batch level. Finally, we train Faster R-CNN using our pRoI generator and, compared to conventional training, obtain better or on-par performance for low IoUs and significant improvements when trained for higher IoUs for Pascal VOC and MS COCO datasets. The code is available at: https://github.com/kemaloksuz/BoundingBoxGenerator. Kemal Oksuz, Baris Can Cam, Emre Akbas, Sinan Kalkan |
WACV | 4 |
| 2019 | Searching for Ambiguous Objects in Videos using Relational Referring Expressions
Hazan Anayurt, Sezai Artun Ozyegin, Ulfet Cetin, Utku Aktas, Sinan Kalkan |
BMVC | 5 |
| 2019 | Learning to Generate Unambiguous Spatial Referring Expressions for Real-World EnvironmentsabstractReferring to objects in a natural and unambiguous manner is crucial for effective human-robot interaction. Previous research on learning-based referring expressions has focused primarily on comprehension tasks, while generating referring expressions is still mostly limited to rule-based methods. In this work, we propose a two-stage approach that relies on deep learning for estimating spatial relations to describe an object naturally and unambiguously with a referring expression. We compare our method to the state of the art algorithm in ambiguous environments (e.g., environments that include very similar objects with similar relationships). We show that our method generates referring expressions that people find to be more accurate (~30% better) and would prefer to use (~32% more often). Fethiye Irmak Dogan, Sinan Kalkan, Iolanda Leite |
IROS | 2 |
| 2019 | Deep 3D semantic scene extrapolation
Ali Abbasi 0006, Sinan Kalkan, Yusuf Sahillioglu |
Vis. Comput. | 2 |
| 2018 | Localization Recall Precision (LRP): A New Performance Metric for Object Detection
Kemal Oksuz, Baris Can Cam, Emre Akbas, Sinan Kalkan |
ECCV (7) | 4 |
| 2018 | What is (Missing or Wrong) in the Scene? A Hybrid Deep Boltzmann Machine for Contextualized Scene ModelingabstractScene models allow robots to reason about what is in the scene, what else should be in it, and what should not be in it. In this paper, we propose a hybrid Boltzmann Machine (BM) for scene modeling where relations between objects are integrated. To be able to do that, we extend BM to include tri-way edges between visible (object) nodes and make the network to share the relations across different objects. We evaluate our method against several baseline models (Deep Boltzmann Machines, and Restricted Boltzmann Machines) on a scene classification dataset, and show that it performs better in several scene reasoning tasks. Ilker Bozcan, Yagmur Oymak, Idil Zeynep Alemdar, Sinan Kalkan |
ICRA | 4 |
| 2018 | A Deep Incremental Boltzmann Machine for Modeling Context in RobotsabstractContext is an essential capability for robots that are to be as adaptive as possible in challenging environments. Although there are many context modeling efforts, they assume a fixed structure and number of contexts. In this paper, we propose an incremental deep model that extends Restricted Boltzmann Machines. Our model gets one scene at a time, and gradually extends the contextual model when necessary, either by adding a new context or a new context layer to form a hierarchy. We show on a scene classification benchmark that our method converges to a good estimate of the contexts of the scenes, and performs better or on-par on several tasks compared to other incremental models or non-incremental models. Fethiye Irmak Dogan, Hande Çelikkanat, Sinan Kalkan |
ICRA | 3 |
| 2018 | CINet: A Learning Based Approach to Incremental Context Modeling in RobotsabstractThere have been several attempts at modeling context in robots. However, either these attempts assume a fixed number of contexts or use a rule-based approach to determine when to increment the number of contexts. In this paper, we pose the task of when to increment as a learning problem, which we solve using a Recurrent Neural Network. We show that the network successfully (with 98% testing accuracy) learns to predict when to increment, and demonstrate, in a scene modeling problem (where the correct number of contexts is not known), that the robot increments the number of contexts in an expected manner (i.e., the entropy of the system is reduced). We also present how the incremental model can be used for various scene reasoning tasks. Fethiye Irmak Dogan, Ilker Bozcan, Mehmet Celik, Sinan Kalkan |
IROS | 4 |
| 2017 | Using deep networks for drone detectionabstractDrone detection is the problem of finding the smallest rectangle that encloses the drone(s) in a video sequence. In this study, we propose a solution using an end-to-end object detection model based on convolutional neural networks. To solve the scarce data problem for training the network, we propose an algorithm for creating an extensive artificial dataset by combining background-subtracted real images. With this approach, we can achieve precision and recall values both of which are high at the same time. Cemal Aker, Sinan Kalkan |
AVSS | 2 |
| 2017 | Drone-vs-Bird detection challenge at IEEE AVSS2017abstractSmall drones are a rising threat due to their possible misuse for illegal activities, in particular smuggling and terrorism. The project SafeShore, funded by the European Commission under the Horizon 2020 program, has launched the “drone-vs-bird detection challenge” to address one of the many technical issues arising in this context. The goal is to detect a drone appearing at some point in a video where birds may be also present: the algorithm should raise an alarm and provide a position estimate only when a drone is present, while not issuing alarms on birds. This paper reports on the challenge proposal, evaluation, and results. Angelo Coluccia, Marian Ghenescu, Tomas Piatrik, Geert De Cubber, Arne Schumann, Lars Wilko Sommer, Johannes Klatte, Tobias Schuchert, Jürgen Beyerer, Mohammad Farhadi, Ruhallah Amandi, Cemal Aker, Sinan Kalkan, Nabin Sharma, Sultan Daud Khan, Khan Makkah, Michael Blumenstein |
AVSS | 13 |
| 2015 | An iterative adaptive multi-modal stereo-vision method using mutual information
Mustafa Yaman, Sinan Kalkan |
J. Vis. Commun. Image Represent. | 2 |
| 2013 | Learning Social Affordances and Using Them for Planning
Kadir Firat Uyanik, Yigit Çaliskan, Asil Kaan Bozcuoglu, Onur Yürüten, Sinan Kalkan, Erol Sahin |
CogSci | 5 |
| 2013 | Deep Hierarchies in the Primate Visual Cortex: What Can We Learn for Computer Vision?abstractComputational modeling of the primate visual system yields insights of potential relevance to some of the challenges that computer vision is facing, such as object recognition and categorization, motion detection and activity recognition, or vision-based navigation and manipulation. This paper reviews some functional principles and structures that are generally thought to underlie the primate visual cortex, and attempts to extract biological principles that could further advance computer vision research. Organized for a computer vision audience, we present functional principles of the processing hierarchies present in the primate visual system considering recent discoveries in neurophysiology. The hierarchical processing in the primate visual system is characterized by a sequence of different levels of processing (on the order of 10) that constitute a deep hierarchy in contrast to the flat vision architectures predominantly used in today's mainstream computer vision. We hope that the functional description of the deep hierarchies realized in the primate visual system provides valuable insights for the design of computer vision algorithms, fostering increasingly productive interaction between biological and computer vision research. Norbert Krüger, Peter Janssen, Sinan Kalkan, Markus Lappe, Ales Leonardis, Justus H. Piater, Antonio Jose Rodríguez-Sánchez, Laurenz Wiskott |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2012 | Human and robotics hands grasping dangerabstractBehavioural and neuroscience studies have shown that observing objects activates affordances, evoking motor responses. The aim of the present study is twofold. First, we intend to investigate whether children are sensitive to the distinction between neutral/graspable (affordances) and dangerous objects. Second, we aim to verify whether children's responses are modulated also by the agent who is interacting with the objects (human hand contrasted with robot hand, and male hand contrasted to female hand). We conducted an experiment on school-age children using a priming paradigm: a prime given by a hand or a control object was followed by graspable or dangerous objects. Children were required to categorize them into artefacts or natural objects by pressing two keys on a keyboard. Our results clearly showed that children are able to distinguish between neutral and dangerous objects: the latter produced an interference effect. In addition, we demonstrated that children are sensitive to the difference between actions performed by biological and non-biological agents: responses were faster when the prime was a grasping hand of a human compared to control stimuli. Results were interpreted in terms of gradient of vulnerability (female hand induced the most inhibition, while robot hand induced the least one) and of motor resonance (resonance is higher when the similarity between the hand prime and the participant's hand is higher). Filomena Anelli, Roberto Nicoletti, Sinan Kalkan, Erol Sahin, Anna M. Borghi |
IJCNN | 3 |
| 2012 | Disparity disambiguation by fusion of signal- and symbolic-level information
Jarno Ralli, Javier Díaz 0001, Sinan Kalkan, Norbert Krüger, Eduardo Ros Vidal |
Mach. Vis. Appl. | 3 |
| 2010 | Learning Affordances for Categorizing Objects and Their PropertiesabstractIn this paper, we demonstrate that simple interactions with objects in the environment leads to a manifestation of the perceptual properties of objects. This is achieved by deriving a condensed representation of the effects of actions (called effect prototypes in the paper), and investigating the relevance between perceptual features extracted from the objects and the actions that can be applied to them. With this at hand, we show that the agent can categorize (i.e., partition) its raw sensory perceptual feature vector, extracted from the environment, which is an important step for development of concepts and language. Moreover, after learning how to predict the effect prototypes of objects, the agent can categorize objects based on the predicted effects of actions that can be applied on them. Nilgün Dag, Ilkay Atil, Sinan Kalkan, Erol Sahin |
ICPR | 3 |
| 2010 | Using multi-modal 3D contours and their relations for vision and robotics
Emre Baseski, Nicolas Pugeault, Sinan Kalkan, Leon Bodenhagen, Justus H. Piater, Norbert Krüger |
J. Vis. Commun. Image Represent. | 3 |
| 2009 | Continuous dimensionality characterization of image structures
Michael Felsberg, Sinan Kalkan, Norbert Krüger |
Image Vis. Comput. | 2 |
| 2007 | A Scene Representation Based on Multi-Modal 2D and 3D FeaturesabstractVisually extracted 2D and 3D information have their own advantages and disadvantages that complement each other. Therefore, it is important to be able to switch between the different dimensions according to the requirements of the problem and use them together to combine the reliability of 2D information with the richness of 3D information. In this article, we use 2D and 3D information in a feature-based vision system and demonstrate their complementary properties on different applications (namely: depth prediction, scene interpretation, grasping from vision and object learning). Emre Baseski, Nicolas Pugeault, Sinan Kalkan, Dirk Kraft, Florentin Wörgötter, Norbert Krüger |
ICCV | 3 |
| 2006 | Statistical Analysis of Local 3D Structure in 2D ImagesabstractFor the analysis of images, a deeper understanding of their intrinsic structure is required. This has been obtained for 2D images by means of statistical analysis [15, 18]. Here, we analyze the relation between local image structures (i.e., homogeneous, edge-like, corner-like or texturelike structures) and the underlying local 3D structure, represented in terms of continuous surfaces and different kinds of 3D discontinuities, using 3D range data with the true color information. We find that homogeneous image patches correspond to continuous surfaces, and discontinuities are mainly formed by edge-like or corner-like structures. The results are discussed with regard to existing and potential computer vision applications and the assumptions made by these applications. Sinan Kalkan, Florentin Wörgötter, Norbert Krüger |
CVPR (1) | 1 |