EDBT 2026 Demo / reviewers in the wild / expert
Savas Özkan
dblp:146/2328
· DBLP profile ↗
20ranked-venue papers
14as first author
10since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 15 · 11 first-author · 9 since 2021Artificial intelligence and machine learning · 5 · 3 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 3 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Mem-MLP: Real-Time 3D Human Motion Generation from Sparse InputsabstractRealistic and smooth full-body tracking is crucial for immersive AR/VR applications. Existing systems primarily track head and hands via Head Mounted Devices (HMDs) and controllers, making the 3D full-body reconstruction incomplete. One potential approach is to generate the full-body motions from sparse inputs collected from limited sensors using a Neural Network (NN) model. In this paper, we propose a novel method based on a multi-layer perceptron (MLP) backbone that is enhanced with residual connections and a novel NN-component called Memory-Block. In particular, Memory-Block represents missing sensor data with trainable code-vectors, which are combined with the sparse signals from previous time instances to improve the temporal consistency. Furthermore, we formulate our solution as a multi-task learning problem, allowing our MLP-backbone to learn robust representations that boost accuracy. Our experiments show that our method outperforms state-of-the-art baselines by substantially reducing prediction errors. Moreover, it achieves 72 FPS on mobile HMDs that ultimately improves the accuracy-running time tradeoff. Sinan Mutlu, Georgios-Fotios Angelis, Savas Özkan, Paul Wisbey, Anastasios Drosou, Mete Ozay |
WACV | 3 |
| 2026 | Guided Model Merging for Hybrid Data Learning: Leveraging Centralized Data to Refine Decentralized ModelsabstractCurrent network training paradigms primarily focus on either centralized or decentralized data regimes. However, in practice, data availability often exhibits a hybrid nature, where both regimes coexist. This hybrid setting presents new opportunities for model training, as the two regimes offer complementary trade-offs: decentralized data is abundant but subject to heterogeneity and communication constraints, while centralized data—though limited in volume and potentially unrepresentative—enables better curation and high-throughput access. Despite its potential, effectively combining these paradigms remains challenging, and few frameworks are tailored to hybrid data regimes. To address this, we propose a novel framework that constructs a model atlas from decentralized models and leverages centralized data to refine a global model within this structured space. The refined model is then used to reinitialize the decentralized models. Our method synergizes federated learning (to exploit decentralized data) and model merging (to utilize centralized data), enabling effective training under hybrid data availability. Theoretically, we show that our approach achieves faster convergence than methods relying solely on decentralized data, due to variance reduction in the merging process. Extensive experiments demonstrate that our framework consistently outperforms purely centralized, purely decentralized, and existing hybrid-adaptable methods. Notably, our method remains robust even when the centralized and decentralized data domains differ or when decentralized data contains noise, significantly broadening its applicability. Junyi Zhu 0002, Ruicong Yao, Taha Ceritli, Savas Özkan, Matthew B. Blaschko, Eunchung Noh, Jeongwon Min, Cho Jung Min, Mete Ozay |
WACV | 4 |
| 2025 | Accurate Scene Text Recognition with Efficient Model Scaling and Cloze Self-DistillationabstractScaling architectures have been proven effective for improving Scene Text Recognition (STR), but the individual contribution of vision encoder and text decoder scaling remain under-explored. In this work, we present an in-depth empirical analysis and demonstrate that, contrary to previous observations, scaling the decoder yields significant performance gains, always exceeding those achieved by encoder scaling alone. We also identify label noise as a key challenge in STR, particularly in real-world data, which can limit the effectiveness of STR models. To address this, we propose Cloze Self-Distillation (CSD), a method that mitigates label noise by distilling a student model from context-aware soft predictions and pseudolabels generated by a teacher model. Additionally, we enhance the decoder architecture by introducing differential cross-attention for STR. Our methodology achieves state-of-the-art performance on 10 out of 11 benchmarks using only real data, while significantly reducing the parameter size and computational costs. Andrea Maracani, Savas Özkan, Sijun Cho, Hyowon Kim, Eunchung Noh, Jeongwon Min, Cho Jung Min, Dookun Park, Mete Ozay |
CVPR | 2 |
| 2025 | A Study of Improving The Privacy-Utility Trade-off of Task-specific Models with Learnable PrivacyabstractIn recent years, machine learning (ML) models have been integrated into various applications and products to improve user experience. However, this approach raises significant concerns about the protection of private user data utilized for training the models. One limitation of vanilla privacy methods is that they can improve the robustness of the models against privacy attacks at the cost of accuracy while performing ML tasks. We propose a framework for implementing privacy models that learn privacy budgets to improve the trade-off between privacy of user data, task models, and their utility (task accuracy). The experimental results show that our framework dramatically enhances the task accuracy of ML models in image classification tasks while providing better privacy protection compared to the state-of-the-art methods. Savas Özkan, Taha Ceritli, Jeongwon Min, Eunchung Noh, Jung Min Cho, Dookun Park, Mete Ozay |
ICASSP | 1 |
| 2025 | Hyper-Refinement for Low-Rank AdaptationabstractParameter-efficient fine-tuning (PEFT) is utilized to adapt large pre-trained machine learning (ML) models to new tasks using a small number of trainable parameters. In particular, Low-Rank Adaptation (LoRA) is one of the prominent PEFT methods. To this end, we introduce a novel method that exploits data to improve the accuracy of models fine-tuned by LoRA. Our method implements a hypernetwork that generates refinement parameters using data to update the low-rank parameters of LoRA. The experimental results validate that our method improves the accuracy of large language models (LLMs) on language understanding tasks by 3% on average compared to LoRA and its variants. Savas Özkan, Taha Ceritli, Jeongwon Min, Eunchung Noh, Jung Min Cho, Dookun Park, Mete Ozay |
ICASSP | 1 |
| 2025 | Efficient and Accurate Scene Text Recognition with Cascaded-TransformersabstractIn recent years, vision transformers with text decoder have demonstrated remarkable performance on Scene Text Recognition (STR) due to their ability to capture long-range dependencies and contextual relationships with high learning capacity. However, the computational and memory demands of these models are significant, limiting their deployment in resource-constrained applications. To address this challenge, we propose an efficient and accurate STR system. Specifically, we focus on improving the efficiency of encoder models by introducing a cascaded-transformers structure. This structure progressively reduces the vision token size during the encoding step, effectively eliminating redundant tokens and reducing computational cost. Our experimental results confirm that our STR system achieves comparable performance to state-of-the-art baselines while substantially decreasing computational requirements. In particular, for large-models, the accuracy remains same, 92.77 → 92.68, while computational complexity is almost halved with our structure. Savas Özkan, Andrea Maracani, Mete Ozay, Hyowon Kim, Sijun Cho, Eunchung Noh, Jeongwon Min, Jung Min Cho |
MMSys | 1 |
| 2024 | Texture and Normal Map Estimation for 3D Face Reconstructionabstract3D Morphable Models (3DMMs) can effectively capture human face shape and texture by exploiting face statistics computed on 3D scanned faces. However, the capacity of these models imposes limitations on shape details and texture expressiveness. Consequently, they provide low-detailed texture and normal map estimations. We propose a novel framework designed to estimate high-quality texture and normal maps in the UV space of a single input image. This framework comprises encoder, normalization and decoder models. To this end, we aim to learn rich and robust representations that remain unaffected under changing pose, expression and illumination by enhancing estimation details. Experimental results demonstrate that our framework outperforms the baseline significantly in terms of realism and identity preservation. Savas Özkan, Mete Ozay, Tom Robinson |
ICASSP | 1 |
| 2024 | Towards Better Control Of Latent Spaces For Face EditingabstractGenerative models can synthesize diverse and photo-realistic images that have demonstrated remarkable success in computer vision. Notably, Generative Adversarial Networks (GANs) trained for faces (i.e., StyleGAN2) can be considered as a powerful image generation pipeline. However, the entangled space of features learned by GANs restricts precise control of modifying the content of generated images. This paper introduces a framework designed to enhance the control for face editing in image generation by disentanglement of feature spaces of GANs (a.k.a. GAN spaces). Our framework aims to enable better control for modification of face concepts such as pose, expression, and illumination. For this purpose, the framework first learns multiple latent spaces for parameterization of face concepts learned by 3D Morphable face models. Then, it employs a Identity-Conditioned Attention Mechanism (ICAM) to decouple face identity representations from the parameterized face concepts. Moreover, we adapt variational autoencoders to model the hierarchical structure of GAN features by incorporating transformer networks for end-to-end optimization of model parameters. Our results show that our ICAM achieves state-of-the-art identity preservation and editing precision accuracy on benchmark datasets with improved training time and memory usage. Savas Özkan, Mete Ozay |
ICIP | 1 |
| 2023 | Conceptual and Hierarchical Latent Space Decomposition for Face EditingabstractGenerative Adversarial Networks (GANs) can produce photo-realistic results using an unconditional image-generation pipeline. However, the images generated by GANs (e.g., StyleGAN) are entangled in feature spaces, which makes it difficult to interpret and control the contents of images. In this paper, we present an encoder-decoder model that decomposes the entangled GAN space into a conceptual and hierarchical latent space in a self-supervised manner. The outputs of 3D morphable face models are leveraged to independently control image synthesis parameters like pose, expression, and illumination. For this purpose, a novel latent space decomposition pipeline is introduced using transformer networks and generative models. Later, this new space is used to optimize a transformer-based GAN space controller for face editing. In this work, a StyleGAN2 model for faces is utilized. Since our method manipulates only GAN features, the photo-realism of Style-GAN2 is fully preserved. The results demonstrate that our method qualitatively and quantitatively outperforms baselines in terms of identity preservation and editing precision. Savas Özkan, Mete Ozay, Tom Robinson |
ICCV | 1 |
| 2021 | CHAOS Challenge - combined (CT-MR) healthy abdominal organ segmentation
A. Emre Kavur, Naciye Sinem Gezer, Mustafa Baris, Sinem Aslan, Pierre-Henri Conze, Vladimir Groza, Duc Duy Pham, Soumick Chatterjee, Philipp Ernst, Savas Özkan, Bora Baydar, Dmitry A. Lachinov, Shuo Han 0001, Josef Pauli, Fabian Isensee, Matthias Perkonigg, Rachana Sathish, Ronnie Rajan, Debdoot Sheet, Gurbandurdy Dovletov, Oliver Speck, Andreas Nürnberger, Klaus H. Maier-Hein, Gozde Bozdagi Akar, Gozde Unal, Oguz Dicle, M. Alper Selver |
Medical Image Anal. | 10 |
| 2020 | Exploiting Local Indexing and Deep Feature Confidence Scores for Fast Image-to-Video SearchabstractThe cost-effective visual representation and fast query-by-example search are two challenging goals that should be maintained for web-scale visual retrieval tasks on moderate hardware. This paper introduces a fast and robust method that ensures both of these goals by obtaining state-of-the-art performance for an image-to-video search scenario. Hence, we present critical enhancements to well-known indexing and visual representation techniques by promoting faster, better and moderate retrieval performance. We also boost the superiority of our method for some visual challenges by exploiting individual decisions of local and global descriptors at query time. For instance, local content descriptors represent copied/duplicated scenes with large geometric deformations such as scale, orientation and affine transformation. In contrast, the use of global content descriptors is more practical for near-duplicate and semantic searches. Experiments are conducted on a large-scale Stanford I2V dataset. The experimental results show that our method is useful in terms of complexity and query processing time for large-scale visual retrieval scenarios, even if local and global representations are used together. The proposed method is superior and achieves state-of-the-art performance based on the mean average precision (MAP) score of this dataset. Lastly, we report additional MAP scores after updating the ground annotations unveiled by retrieval results of the proposed method, and it shows that the actual performance. Savas Özkan, Gozde Bozdagi Akar |
ICPR | 1 |
| 2019 | EndNet: Sparse AutoEncoder Network for Endmember Extraction and Hyperspectral UnmixingabstractData acquired from multichannel sensors are a highly valuable asset to interpret the environment for a variety of remote sensing applications. However, low spatial resolution is a critical limitation for previous sensors, and the constituent materials of a scene can be mixed in different fractions due to their spatial interactions. Spectral unmixing is a technique that allows us to obtain the material spectral signatures and their fractions from hyperspectral data. In this paper, we propose a novel endmember extraction and hyperspectral unmixing scheme, so-called EndNet, that is based on a two-staged autoencoder network. This well-known structure is completely enhanced and restructured by introducing additional layers and a projection metric [i.e., spectral angle distance (SAD) instead of inner product] to achieve an optimum solution. Moreover, we present a novel loss function that is composed of a Kullback-Leibler divergence term with SAD similarity and additional penalty terms to improve the sparsity of the estimates. These modifications enable us to set the common properties of endmembers, such as nonlinearity and sparsity for autoencoder networks. Finally, due to the stochastic-gradient-based approach, the method is scalable for large-scale data and it can be accelerated on graphical processing units. To demonstrate the superiority of our proposed method, we conduct extensive experiments on several well-known data sets. The results confirm that the proposed method considerably improves the performance compared to the state-of-the-art techniques in the literature. Savas Özkan, Berk Kaya, Gozde Bozdagi Akar |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2018 | Kinshipgan: Synthesizing of Kinship Faces from Family Photos by Regularizing a Deep Face NetworkabstractIn this paper, we propose a kinship generator network that can synthesize a possible child face by analyzing his/her parent's photo. For this purpose, we focus on to handle the scarcity of kinship datasets throughout the paper by proposing novel solutions in particular. To extract robust features, we integrate a pre-trained face model to the kinship face generator. Moreover, the generator network is regularized with an additional face dataset and adversarial loss to decrease the overfitting of the limited samples. Lastly, we adapt cycle-domain transformation to attain a more stable results. Experiments are conducted on Families in the Wild (FIW) dataset. The experimental results show that the contributions presented in the paper provide important performance improvements compared to the baseline architecture and our proposed method yields promising perceptual results. Savas Özkan, Akin Orkan |
ICIP | 1 |
| 2018 | Deep Spectral Convolution Network for Hyperspectral UnmixingabstractIn this paper, we propose a novel hyperspectral unmixing technique based on deep spectral convolution networks (DSCN). Particularly, three important contributions are presented throughout this paper. First, fully-connected linear operation is replaced with spectral convolutions to extract local spectral characteristics from hyperspectral signatures with a deeper network architecture. Second, instead of batch normalization, we propose a spectral normalization layer which improves the selectivity of filters by normalizing their spectral responses. Third, we introduce two fusion configurations that produce ideal abundance maps by using the abstract representations computed from previous layers. In experiments, we use two real datasets to evaluate the performance of our method with other baseline techniques. The experimental results validate that the proposed method outperforms baselines based on Root Mean Square Error (RMSE). Savas Özkan, Gozde Bozdagi Akar |
ICIP | 1 |
| 2018 | Cloud Detection from RGB Color Remote Sensing Images with Deep Pyramid NetworksabstractCloud detection from remotely observed data is a critical pre-processing step for various remote sensing applications. In particular, this problem becomes even harder for RGB color images, since there is no distinct spectral pattern for clouds, which is directly separable from the Earth surface. In this paper, we adapt a deep pyramid network (DPN) to tackle this problem. For this purpose, the network is enhanced with a pre-trained parameter model at the encoder layer. Moreover, the method is able to obtain accurate pixel-level segmentation and classification results from a set of noisy labeled RGB color images. In order to demonstrate the superiority of the method, we collect and label data with the corresponding cloud/non-cloudy masks acquired from low-orbit Gokturk-2 and RASAT satellites. The experimental results validates that the proposed method outperforms several baselines even for hard cases (e.g. snowy mountains) that are perceptually difficult to distinguish by human eyes. Savas Özkan, Mehmet Efendioglu, Caner Demirpolat |
IGARSS | 1 |
| 2018 | Large-scale image retrieval using transductive support vector machines
Hakan Çevikalp, Merve Elmas, Savas Özkan |
Comput. Vis. Image Underst. | 3 |
| 2014 | Enhanced spatio-temporal video copy detection by combining trajectory and spatial consistencyabstractThe recent improvements on internet technologies and video coding techniques cause an increase in copyright infringements especially for video. Frequently, image-based approaches appear as an essential solution due to the fact that joint usage of quantization-based indexing and weak geometric consistency stages give a capability to compare duplicate videos quickly. However, exploiting purely spatial content ignores the temporal variation of video. In this work, we propose a system that combines the state-of-the-art quantization-based indexing scheme with a novel trajectory-based geometric consistency on spatio-temporal features. This combination improves duplicate video matching task significantly. Briefly, spatial mean and variance of the trajectories are incorporated to establish a weak geometric consistency among pair of frames. To show the success of the proposed method, content-based video copy detection field is selected and TRECVID 2009 dataset is utilized. The experimental results show that constituting trajectory-based consistency on corresponding feature pairs outperforms the performances of merely utilizing spatiotemporal signature and visual signature with enhanced weak geometric consistency. Savas Özkan, Ersin Esen, Gozde Bozdagi Akar |
ICIP | 1 |
| 2014 | Visual Group Binary Signature for Video Copy DetectionabstractNeed for automatic video copy detection is increased with the recent technical developments in the internet technologies and video recording. Even though image-based techniques with bag-of-word kind of representations are accepted as the best solution because of robustness and speed, they discard the convenient geometric relation which exists among interest points. In this work, we propose a novel geometric relation which computes a binary signature leveraging existence and non-existence of interest points in the neighborhood area. The experimental results on TRECVID 2009 content-based video copy detection dataset show that combination of our method with recently proposed quantization-based indexing and weak geometric consistency schemes outperforms classical representations. Savas Özkan, Ersin Esen, Gozde Bozdagi Akar |
ICPR | 1 |
| 2014 | Performance Analysis of State-of-the-Art Representation Methods for Geographical Image Retrieval and CategorizationabstractThis letter studies the performance of various image representation schemes used for image search problems for the purpose of geographic image retrieval from satellite imagery. We compare the most widely adopted method of the bag-of-words (BoW) approach with the more recently introduced vector of locally aggregated descriptors (VLAD) and its more compact binary version product quantized VLAD (VLAD-PQ). We show with the experiments on a publicly available 21-class land-use/land-cover data set that the VLAD-based representation outperforms BoW at the cost of increased query time, but the more compact VLAD-PQ representation achieves very similar performance as VLAD without the increased time requirement. Savas Özkan, Tayfun Ates, Engin Tola, Medeni Soysal, Ersin Esen |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2014 | Multimodal concept detection in broadcast media: KavTan
Medeni Soysal, K. Berker Logoglu, Mashar Tekin, Ersin Esen, Ahmet Saracoglu, Banu Oskay Acar, Ezgi C. Ozan, Tugrul K. Ates, Hakan Sevimli, Ayça Müge Sevinç, Ilkay Atil, Savas Özkan, Mehmet Ali Arabaci, Seda Tankiz, Talha Karadeniz 0002, Duygu Oskay Önür, Sezin Selçuk, A. Aydin Alatan, Tolga Çiloglu |
Multim. Tools Appl. | 12 |