Levent Karacan

dblp:138/3443 · DBLP profile ↗
← Back
11ranked-venue papers
7as first author
7since 2021 · last 2026
0000-0003-2764-5258ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 9 · 7 first-author · 5 since 2021Artificial intelligence and machine learning · 6 · 2 first-author · 5 since 2021
YearPublicationVenuePosition
2026 FISformer: Replacing Self-Attention With a Fuzzy Inference System in Transformer Models for Time Series Forecasting
abstract
Transformers have achieved remarkable progress in time series forecasting, yet their reliance on deterministic dot-product attention limits their capacity to model uncertainty and nonlinear dependencies across multivariate temporal dimensions. To address this limitation, we propose FISformer, a Fuzzy Inference System-Driven Transformer that replaces conventional attention with a FIS Interaction mechanism. In this framework, each query-key pair undergoes a fuzzy inference process for every feature dimension, where learnable membership functions and rule-based reasoning estimate token-wise relational strengths. These FIS-derived interaction weights capture uncertainty and provide interpretable, continuous mappings between tokens. A softmax operation is applied along the token axis to normalize these weights, which are then combined with the corresponding value features through element-wise multiplication to yield the final context-enhanced token representations. This design fuses the interpretability and uncertainty modeling of fuzzy logic with the representational power of Transformers. Extensive experiments on multiple benchmark datasets demonstrate that FISformer achieves superior forecasting accuracy, noise robustness, and interpretability compared to state-of-the-art Transformer variants, establishing fuzzy inference as an effective alternative to conventional attention mechanisms.
Bülent Haznedar, Levent Karacan
IEEE Trans. Fuzzy Syst.2
2025 One-Step Specular Highlight Removal with Adapted Diffusion Models
Mahir Atmis, Levent Karacan, Mehmet Sarigul
ICCV2
2025 Full-Frame Video Stabilization via Spatiotemporal Transformers
abstract
Traditional video stabilization methods use a warping operation to smooth the camera path but result in missing regions in the video frames. To solve this issue, full-frame video stabilization techniques attempt to fill in the unidentified boundary regions, but their effectiveness is limited. In this work, we propose a full-frame video stabilization method using spatiotemporal transformers to fill the missing boundary regions after the warping operation. For training, we adopt a self-supervised strategy and improve it by incorporating temporal information. The proposed approach allows the utilization of redundant video information spatially and temporally while filling in missing regions. Experimental results show that our approach achieves superior results on popular video stabilization datasets. The code, pre-trained model, and video results are available at https://github.com/leventkaracan/VidStabFormer.
Levent Karacan, Mehmet Sarigul
Comput. Vis. Media1
2023 VidStyleODE: Disentangled Video Editing via StyleGAN and NeuralODEs
abstract
We propose VidStyleODE, a spatiotemporally continuous disentangled video representation based upon StyleGAN and Neural-ODEs. Effective traversal of the latent space learned by Generative Adversarial Networks (GANs) has been the basis for recent breakthroughs in image editing. However, the applicability of such advancements to the video domain has been hindered by the difficulty of representing and controlling videos in the latent space of GANs. In particular, videos are composed of content (i.e., appearance) and complex motion components that require a special mechanism to disentangle and control. To achieve this, VidStyleODE encodes the video content in a pre-trained StyleGAN ${\mathcal{W}_ + }$ space and benefits from a latent ODE component to summarize the spatiotemporal dynamics of the input video. Our novel continuous video generation process then combines the two to generate high-quality and temporally consistent videos with varying frame rates. We show that our proposed method enables a variety of applications on real videos: text-guided appearance manipulation, motion manipulation, image animation, and video interpolation and extrapolation. Project website: https://cyberiada.github.io/VidStyleODE
Moayed Haji Ali, Andrew Bond, Levent Karacan, Tolga Birdal, Erkut Erdem, Duygu Ceylan, Aykut Erdem
ICCV3
2023 Region contrastive camera localization
Mehmet Sarigul, Levent Karacan
Pattern Recognit. Lett.2
2023 Multi-image transformer for multi-focus image fusion
Levent Karacan
Signal Process. Image Commun.1
2022 Disentangling Content and Motion for Text-Based Neural Video Manipulation
Levent Karacan, Tolga Kerimoglu, Ismail Inan, Tolga Birdal, Erkut Erdem, Aykut Erdem
BMVC1
2020 Manipulating Attributes of Natural Scenes via Hallucination
abstract
In this study, we explore building a two-stage framework for enabling users to directly manipulate high-level attributes of a natural scene. The key to our approach is a deep generative network that can hallucinate images of a scene as if they were taken in a different season (e.g., during winter), weather condition (e.g., on a cloudy day), or at a different time of the day (e.g., at sunset). Once the scene is hallucinated with the given attributes, the corresponding look is then transferred to the input image while preserving the semantic details intact, giving a photo-realistic manipulation result. As the proposed framework hallucinates what the scene will look like, it does not require any reference style image as commonly utilized in most of the appearance or style transfer approaches. Moreover, it allows to simultaneously manipulate a given scene according to a diverse set of transient attributes within a single model, eliminating the need of training multiple networks per each translation task. Our comprehensive set of qualitative and quantitative results demonstrates the effectiveness of our approach against the competing methods.
Levent Karacan, Zeynep Akata, Aykut Erdem, Erkut Erdem
ACM Trans. Graph.1
2017 Alpha Matting With KL-Divergence-Based Sparse Sampling
abstract
In this paper, we present a new sampling-based alpha matting approach for the accurate estimation of foreground and background layers of an image. Previous sampling-based methods typically rely on certain heuristics in collecting representative samples from known regions, and thus their performance deteriorates if the underlying assumptions are not satisfied. To alleviate this, we take an entirely new approach and formulate sampling as a sparse subset selection problem where we propose to pick a small set of candidate samples that best explains the unknown pixels. Moreover, we describe a new dissimilarity measure for comparing two samples which is based on KL-divergence between the distributions of features extracted in the vicinity of the samples. The proposed framework is general and could be easily extended to video matting by additionally taking temporal information into account in the sampling process. Evaluation on standard benchmark data sets for image and video matting demonstrates that our approach provides more accurate results compared with the state-of-the-art methods.
Levent Karacan, Aykut Erdem, Erkut Erdem
IEEE Trans. Image Process.1
2015 Image Matting with KL-Divergence Based Sparse Sampling
abstract
Previous sampling-based image matting methods typically rely on certain heuristics in collecting representative samples from known regions, and thus their performance deteriorates if the underlying assumptions are not satisfied. To alleviate this, in this paper we take an entirely new approach and formulate sampling as a sparse subset selection problem where we propose to pick a small set of candidate samples that best explains the unknown pixels. Moreover, we describe a new distance measure for comparing two samples which is based on KL-divergence between the distributions of features extracted in the vicinity of the samples. Using a standard benchmark dataset for image matting, we demonstrate that our approach provides more accurate results compared with the state-of-the-art methods.
Levent Karacan, Aykut Erdem, Erkut Erdem
ICCV1
2013 Structure-preserving image smoothing via region covariances
abstract
Recent years have witnessed the emergence of new image smoothing techniques which have provided new insights and raised new questions about the nature of this well-studied problem. Specifically, these models separate a given image into its structure and texture layers by utilizing non-gradient based definitions for edges or special measures that distinguish edges from oscillations. In this study, we propose an alternative yet simple image smoothing approach which depends on covariance matrices of simple image features, aka the region covariances. The use of second order statistics as a patch descriptor allows us to implicitly capture local structure and texture information and makes our approach particularly effective for structure extraction from texture. Our experimental results have shown that the proposed approach leads to better image decompositions as compared to the state-of-the-art methods and preserves prominent edges and shading well. Moreover, we also demonstrate the applicability of our approach on some image editing and manipulation tasks such as image abstraction, texture and detail enhancement, image composition, inverse halftoning and seam carving.
Levent Karacan, Erkut Erdem, Aykut Erdem
ACM Trans. Graph.1