EDBT 2026 Demo / reviewers in the wild / expert
Erkut Erdem
dblp:79/6569
· DBLP profile ↗
56ranked-venue papers
5as first author
30since 2021 · last 2026
0000-0002-6744-8614ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 33 · 2 first-author · 17 since 2021Artificial intelligence and machine learning · 29 · 3 first-author · 16 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Systems, architecture and hardware · 1Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Spherical Vision Transformers for Audio-Visual Saliency Prediction in 360$^{\circ }$∘ VideosabstractOmnidirectional videos (ODVs) are redefining viewer experiences in virtual reality (VR) by offering an unprecedented full field-of-view (FOV). This study extends the domain of saliency prediction to 360$^\circ$∘ environments, addressing the complexities of spherical distortion and the integration of spatial audio. Contextually, ODVs have transformed user experience by adding a spatial audio dimension that aligns sound direction with the viewer's perspective in spherical scenes. Motivated by the lack of comprehensive datasets for 360$^\circ$∘ audio-visual saliency prediction, our study curates YT360-EyeTracking, a new dataset of 81 ODVs, each observed under varying audio-visual conditions. Our goal is to explore how to utilize audio-visual cues to effectively predict visual saliency in 360$^\circ$∘ videos. Towards this aim, we propose two novel saliency prediction models: SalViT360, a vision-transformer-based framework for ODVs equipped with spherical geometry-aware spatio-temporal attention layers, and SalViT360-AV, which further incorporates transformer adapters conditioned on audio input. Our results on a number of benchmark datasets, including our YT360-EyeTracking, demonstrate that SalViT360 and SalViT360-AV significantly outperform existing methods in predicting viewer attention in 360$^\circ$∘ scenes. Interpreting these results, we suggest that integrating spatial audio cues in the model architecture is crucial for accurate saliency prediction in omnidirectional videos. Mert Cokelek, Halit Ozsoy, Nevrez Imamoglu, Cagri Ozcinar, Inci Ayhan, Erkut Erdem, Aykut Erdem |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2025 | GaussianVideo: Efficient Video Representation via Hierarchical Gaussian SplattingabstractEfficient neural representations for dynamic video scenes are critical for applications ranging from video compression to interactive simulations. Yet, existing methods often face challenges related to high memory usage, lengthy training times, and temporal consistency. To address these issues, we introduce a novel neural video representation that combines 3D Gaussian splatting with continuous camera motion modeling. By leveraging Neural ODEs, our approach learns smooth camera trajectories while maintaining an explicit 3D scene representation through Gaussians. Additionally, we introduce a spatiotemporal hierarchical learning strategy, progressively refining spatial and temporal features to enhance reconstruction quality and accelerate convergence. This memory-efficient approach achieves high-quality rendering at impressive speeds. Experimental results show that our hierarchical learning, combined with robust camera motion modeling, captures complex dynamic scenes with strong temporal consistency, achieving state-of-the-art performance across diverse video datasets in both high- and low-motion scenarios. Andrew Bond, Jui-Hsien Wang, Long Mai, Erkut Erdem, Aykut Erdem |
ICCV | 4 |
| 2024 | ViLMA: A Zero-Shot Benchmark for Linguistic and Temporal Grounding in Video-Language ModelsabstractWith the ever-increasing popularity of pretrained Video-Language Models (VidLMs), there is a pressing need to develop robust evaluation methodologies that delve deeper into their visio-linguistic capabilities. To address this challenge, we present ViLMA (Video Language Model Assessment), a task-agnostic benchmark that places the assessment of fine-grained capabilities of these models on a firm footing. Task-based evaluations, while valuable, fail to capture the complexities and specific temporal aspects of moving images that VidLMs need to process. Through carefully curated counterfactuals, ViLMA offers a controlled evaluation suite that sheds light on the true potential of these models, as well as their performance gaps compared to human-level understanding. ViLMA also includes proficiency tests, which assess basic capabilities deemed essential to solving the main counterfactual tests. We show that current VidLMs’ grounding abilities are no better than those of vision-language models which use static images. This is especially striking once the performance on proficiency tests is factored in. Our benchmark serves as a catalyst for future research on VidLMs, helping to highlight areas that still need to be explored. Ilker Kesen, Andrea Pedrotti, Mustafa Dogan, Michele Cafagna, Emre Can Acikgoz, Letitia Parcalabescu, Iacer Calixto, Anette Frank, Albert Gatt, Aykut Erdem, Erkut Erdem |
ICLR | 11 |
| 2024 | Self-Supervised Calibration of the Denoising Networks for HSIabstractTypically, neural networks are trained using supervised learning (SL) and evaluated on unseen data. This type of training relies on a substantial amount of data, including clean images. However, in the case of hyperspectral images (HSIs), acquiring a large number of images along with clean versions can be challenging and expensive. This study proposes a two-stage learning strategy to train the model for HSI data with previously unseen noise patterns. The first stage involves supervised learning to train the model on noisy and clean data pairs. The second stage incorporates self-supervised calibration using only noisy data to adapt the model to specific noise patterns. For the latter, to estimate the middle spectral band, we leverage the information from its neighboring band as a target. To ensure the network learns meaningful relationships rather than merely copying the input, we strategically create a blind spot by excluding the target band from the input data. Therefore, our self-supervised learning technique is named as Blind Band Self-Supervised (BBSS) Learning. Our approach has been shown to improve the accuracy of the model for noisy HSIs, even when the network did not previously encounter the specific noise patterns in SL. Orhan Torun, Seniha Esen Yüksel, Erkut Erdem, Aykut Erdem |
IGARSS | 3 |
| 2024 | Sequential Compositional Generalization in Multimodal ModelsabstractSemih Yagcioglu, Osman Batur İnce, Aykut Erdem, Erkut Erdem, Desmond Elliott, Deniz Yuret. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Semih Yagcioglu, Osman Batur Ince, Aykut Erdem, Erkut Erdem, Desmond Elliott, Deniz Yuret |
NAACL-HLT | 4 |
| 2024 | CLIPAway: Harmonizing focused embeddings for removing objects via diffusion modelsabstractAdvanced image editing techniques, particularly inpainting, are essential for seamlessly removing unwanted elements while preserving visual integrity. Traditional GAN-based methods have achieved notable success, but recent advancements in diffusion models have produced superior results due to their training on large-scale datasets, enabling the generation of remarkably realistic inpainted images.
Despite their strengths, diffusion models often struggle with object removal tasks without explicit guidance, leading to unintended hallucinations of the removed object. To address this issue, we introduce CLIPAway, a novel approach leveraging CLIP embeddings to focus on background regions while excluding foreground elements. CLIPAway enhances inpainting accuracy and quality by identifying embeddings that prioritize the background, thus achieving seamless object removal. Unlike other methods that rely on specialized training datasets or costly manual annotations, CLIPAway provides a flexible, plug-and-play solution compatible with various diffusion-based inpainting techniques. Yigit Ekin, Ahmet Burak Yildirim, Erdem Eren Caglar, Aykut Erdem, Erkut Erdem, Aysegul Dundar |
NeurIPS | 5 |
| 2024 | HyperGAN-CLIP: A Unified Framework for Domain Adaptation, Image Synthesis and Manipulation
Abdul Basit Anees, Ahmet Canberk Baykal, Muhammed Burak Kizil, Duygu Ceylan, Erkut Erdem, Aykut Erdem |
SIGGRAPH Asia | 5 |
| 2024 | Omnidirectional image quality assessment with local-global vision transformers
Nafiseh Jabbari Tofighi, Mohamed Hedi Elfkir, Nevrez Imamoglu, Cagri Ozcinar, Aykut Erdem, Erkut Erdem |
Image Vis. Comput. | 6 |
| 2024 | Hyperspectral image denoising via self-modulating convolutional neural networks
Orhan Torun, Seniha Esen Yüksel, Erkut Erdem, Nevrez Imamoglu, Aykut Erdem |
Signal Process. | 3 |
| 2024 | HyperE2VID: Improving Event-Based Video Reconstruction via HypernetworksabstractEvent-based cameras are becoming increasingly popular for their ability to capture high-speed motion with low latency and high dynamic range. However, generating videos from events remains challenging due to the highly sparse and varying nature of event data. To address this, in this study, we propose HyperE2VID, a dynamic neural network architecture for event-based video reconstruction. Our approach uses hypernetworks to generate per-pixel adaptive filters guided by a context fusion module that combines information from event voxel grids and previously reconstructed intensity images. We also employ a curriculum learning strategy to train the network more robustly. Our comprehensive experimental evaluations across various benchmark datasets reveal that HyperE2VID not only surpasses current state-of-the-art methods in terms of reconstruction quality but also achieves this with fewer parameters, reduced computational requirements, and accelerated inference times. Burak Ercan, Onur Eker, Canberk Saglam, Aykut Erdem, Erkut Erdem |
IEEE Trans. Image Process. | 5 |
| 2023 | Spherical Vision Transformer for 360° Video Saliency Prediction
Mert Cokelek, Nevrez Imamoglu, Cagri Ozcinar, Erkut Erdem, Aykut Erdem |
BMVC | 4 |
| 2023 | ST360IQ: No-Reference Omnidirectional Image Quality Assessment With Spherical Vision TransformersabstractOmnidirectional images, aka 360° images, can deliver immersive and interactive visual experiences. As their popularity has increased dramatically in recent years, evaluating the quality of 360° images has become a problem of interest since it provides insights for capturing, transmitting, and consuming this new media. However, directly adapting quality assessment methods proposed for standard natural images for omnidirectional data poses certain challenges. These models need to deal with very high-resolution data and implicit distortions due to the spherical form of the images. In this study, we present a method for no-reference 360° image quality assessment. Our proposed ST360IQ model extracts tangent viewports from the salient parts of the input omnidirectional image and employs a vision-transformers based module processing saliency selective patches/tokens that estimates a quality score from each viewport. Then, it aggregates these scores to give a final quality score. Our experiments on two benchmark datasets, namely OIQA and CVIQ datasets, demonstrate that as compared to the state-of-the-art, our approach predicts the quality of an omnidirectional image correlated with the human-perceived image quality. The code has been available on https://github.com/Nafiseh-Tofighi/ST360IQ Nafiseh Jabbari Tofighi, Mohamed Hedi Elfkir, Nevrez Imamoglu, Cagri Ozcinar, Erkut Erdem, Aykut Erdem |
ICASSP | 5 |
| 2023 | VidStyleODE: Disentangled Video Editing via StyleGAN and NeuralODEsabstractWe propose VidStyleODE, a spatiotemporally continuous disentangled video representation based upon StyleGAN and Neural-ODEs. Effective traversal of the latent space learned by Generative Adversarial Networks (GANs) has been the basis for recent breakthroughs in image editing. However, the applicability of such advancements to the video domain has been hindered by the difficulty of representing and controlling videos in the latent space of GANs. In particular, videos are composed of content (i.e., appearance) and complex motion components that require a special mechanism to disentangle and control. To achieve this, VidStyleODE encodes the video content in a pre-trained StyleGAN ${\mathcal{W}_ + }$ space and benefits from a latent ODE component to summarize the spatiotemporal dynamics of the input video. Our novel continuous video generation process then combines the two to generate high-quality and temporally consistent videos with varying frame rates. We show that our proposed method enables a variety of applications on real videos: text-guided appearance manipulation, motion manipulation, image animation, and video interpolation and extrapolation. Project website: https://cyberiada.github.io/VidStyleODE Moayed Haji Ali, Andrew Bond, Levent Karacan, Tolga Birdal, Erkut Erdem, Duygu Ceylan, Aykut Erdem |
ICCV | 5 |
| 2023 | CLIP-guided StyleGAN Inversion for Text-driven Real Image EditingabstractResearchers have recently begun exploring the use of StyleGAN-based models for real image editing. One particularly interesting application is using natural language descriptions to guide the editing process. Existing approaches for editing images using language either resort to instance-level latent code optimization or map predefined text prompts to some editing directions in the latent space. However, these approaches have inherent limitations. The former is not very efficient, while the latter often struggles to effectively handle multi-attribute changes. To address these weaknesses, we present CLIPInverter, a new text-driven image editing approach that is able to efficiently and reliably perform multi-attribute changes. The core of our method is the use of novel, lightweight text-conditioned adapter layers integrated into pretrained GAN-inversion networks. We demonstrate that by conditioning the initial inversion step on the Contrastive Language-Image Pre-training (CLIP) embedding of the target description, we are able to obtain more successful edit directions. Additionally, we use a CLIP-guided refinement step to make corrections in the resulting residual latent codes, which further improves the alignment with the text prompt. Our method outperforms competing approaches in terms of manipulation accuracy and photo-realism on various domains including human faces, cats, and birds, as shown by our qualitative and quantitative results. Ahmet Canberk Baykal, Abdul Basit Anees, Duygu Ceylan, Erkut Erdem, Aykut Erdem, Deniz Yuret |
ACM Trans. Graph. | 4 |
| 2022 | Disentangling Content and Motion for Text-Based Neural Video Manipulation
Levent Karacan, Tolga Kerimoglu, Ismail Inan, Tolga Birdal, Erkut Erdem, Aykut Erdem |
BMVC | 5 |
| 2022 | How scene attributes and sound influence visual exploration of omnidirectional panoramic scenes
Halit Ozsoy, Mert Cokelek, Inci Ayhan, Erkut Erdem, Aykut Erdem |
CogSci | 4 |
| 2022 | Perception-Distortion Trade-Off in the SR Space Spanned by Flow ModelsabstractFlow-based generative super-resolution (SR) models learn to produce a diverse set of feasible SR solutions, called the SR space. Diversity of SR solutions increases with the temperature (τ) of latent variables, which introduces random variations of texture among sample solutions, resulting in visual artifacts and low fidelity. In this paper, we present a simple but effective image ensembling/fusion approach to obtain a single SR image eliminating random artifacts and improving fidelity without significantly compromising perceptual quality. We achieve this by benefiting from a diverse set of feasible photorealistic solutions in the SR space spanned by flow models. We propose different image ensembling and fusion strategies which offer multiple paths to move sample solutions in the SR space to more desired destinations in the perception-distortion plane in a controllable manner depending on the fidelity vs. perceptual quality requirements of the task at hand. Experimental results demonstrate that our image ensembling/fusion strategy achieves more promising perception-distortion trade-off compared to sample SR images produced by flow models and adversarially trained models in terms of both quantitative metrics and visual quality. Cansu Korkmaz, A. Murat Tekalp, Zafer Dogan, Erkut Erdem, Aykut Erdem |
ICIP | 4 |
| 2022 | Neural Natural Language Generation: A Survey on Multilinguality, Multimodality, Controllability and LearningabstractDeveloping artificial learning systems that can understand and generate natural language has been one of the long-standing goals of artificial intelligence. Recent decades have witnessed an impressive progress on both of these problems, giving rise to a new family of approaches. Especially, the advances in deep learning over the past couple of years have led to neural approaches to natural language generation (NLG). These methods combine generative language learning techniques with neural-networks based frameworks. With a wide range of applications in natural language processing, neural NLG (NNLG) is a new and fast growing field of research. In this state-of-the-art report, we investigate the recent developments and applications of NNLG in its full extent from a multidimensional view, covering critical perspectives such as multimodality, multilinguality, controllability and learning strategies. We summarize the fundamental building blocks of NNLG approaches from these aspects and provide detailed reviews of commonly used preprocessing steps and basic neural architectures. This report also focuses on the seminal applications of these NNLG models such as machine translation, description generation, automatic speech recognition, abstractive summarization, text simplification, question answering and generation, and dialogue generation. Finally, we conclude with a thorough discussion of the described frameworks by pointing out some open research directions. Erkut Erdem, Menekse Kuyu, Semih Yagcioglu, Anette Frank, Letitia Parcalabescu, Barbara Plank, Andrii Babii, Oleksii Turuta, Aykut Erdem, Iacer Calixto, Elena Lloret, Elena Apostol, Ciprian-Octavian Truica, Branislava Sandrih, Sanda Martincic-Ipsic, Gábor Berend, Albert Gatt, Grazina Korvel |
J. Artif. Intell. Res. | 1 |
| 2022 | Leveraging semantic saliency maps for query-specific video summarization
Kemal Cizmeciler, Erkut Erdem, Aykut Erdem |
Multim. Tools Appl. | 2 |
| 2021 | Cross-lingual Visual Pre-training for Multimodal Machine TranslationabstractOzan Caglayan, Menekse Kuyu, Mustafa Sercan Amac, Pranava Madhyastha, Erkut Erdem, Aykut Erdem, Lucia Specia. Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. 2021. Ozan Caglayan, Menekse Kuyu, Mustafa Sercan Amac, Pranava Swaroop Madhyastha, Erkut Erdem, Aykut Erdem, Lucia Specia |
EACL | 5 |
| 2021 | SLAMP: Stochastic Latent Appearance and Motion PredictionabstractMotion is an important cue for video prediction and often utilized by separating video content into static and dynamic components. Most of the previous work utilizing motion is deterministic but there are stochastic methods that can model the inherent uncertainty of the future. Existing stochastic models either do not reason about motion explicitly or make limiting assumptions about the static part. In this paper, we reason about appearance and motion in the video stochastically by predicting the future based on the motion history. Explicit reasoning about motion without history already reaches the performance of current stochastic models. The motion history further improves the results by allowing to predict consistent dynamics several frames into the future. Our model performs comparably to the state-of-the-art models on the generic video prediction datasets, however, significantly outperforms them on two challenging real-world autonomous driving datasets with complex motion and dynamic background. Adil Kaan Akan, Erkut Erdem, Aykut Erdem, Fatma Güney |
ICCV | 2 |
| 2021 | NOVA: Rendering Virtual Worlds with Humans for Computer Vision TasksabstractAbstract Today, the cutting edge of computer vision research greatly depends on the availability of large datasets, which are critical for effectively training and testing new methods. Manually annotating visual data, however, is not only a labor‐intensive process but also prone to errors. In this study, we present NOVA, a versatile framework to create realistic‐looking 3D rendered worlds containing procedurally generated humans with rich pixel‐level ground truth annotations. NOVA can simulate various environmental factors such as weather conditions or different times of day, and bring an exceptionally diverse set of humans to life, each having a distinct body shape, gender and age. To demonstrate NOVA's capabilities, we generate two synthetic datasets for person tracking. The first one includes 108 sequences, each with different levels of difficulty like tracking in crowded scenes or at nighttime and aims for testing the limits of current state‐of‐the‐art trackers. A second dataset of 97 sequences with normal weather conditions is used to show how our synthetic sequences can be utilized to train and boost the performance of deep‐learning based trackers. Our results indicate that the synthetic data generated by NOVA represents a good proxy of the real‐world and can be exploited for computer vision tasks. Abdulrahman Kerim, Cem Aslan, Ufuk Celikcan, Erkut Erdem, Aykut Erdem |
Comput. Graph. Forum | 4 |
| 2021 | From Noon to Sunset: Interactive Rendering, Relighting, and Recolouring of Landscape Photographs by Modifying Solar PositionabstractAbstract Image editing is a commonly studied problem in computer graphics. Despite the presence of many advanced editing tools, there is no satisfactory solution to controllably update the position of the sun using a single image. This problem is made complicated by the presence of clouds, complex landscapes, and the atmospheric effects that must be accounted for. In this paper, we tackle this problem starting with only a single photograph. With the user clicking on the initial position of the sun, our algorithm performs several estimation and segmentation processes for finding the horizon, scene depth, clouds, and the sky line. After this initial process, the user can make both fine‐ and large‐scale changes on the position of the sun: it can be set beneath the mountains or moved behind the clouds practically turning a midday photograph into a sunset (or vice versa). We leverage a precomputed atmospheric scattering algorithm to make all of these changes not only realistic but also in real‐time. We demonstrate our results using both clear and cloudy skies, showing how to add, remove, and relight clouds, all the while allowing for advanced effects such as scattering, shadows, light shafts, and lens flares. Murat Türe, Mustafa Ege Çiklabakkal, Aykut Erdem, Erkut Erdem, Pinar Satilmis, Ahmet Oguz Akyüz |
Comput. Graph. Forum | 4 |
| 2021 | Using synthetic data for person tracking under adverse weather conditions
Abdulrahman Kerim, Ufuk Celikcan, Erkut Erdem, Aykut Erdem |
Image Vis. Comput. | 3 |
| 2021 | mustGAN: multi-stream Generative Adversarial Networks for MR Image Synthesis
Mahmut Yurt, Salman Ul Hassan Dar, Aykut Erdem, Erkut Erdem, Kader Karli Oguz, Tolga Çukur |
Medical Image Anal. | 4 |
| 2021 | MSVD-Turkish: a comprehensive multimodal video dataset for integrated vision and language research in Turkish
Begüm Çitamak Erdinç, Ozan Caglayan, Menekse Kuyu, Erkut Erdem, Aykut Erdem, Pranava Swaroop Madhyastha, Lucia Specia |
Mach. Transl. | 4 |
| 2021 | Leveraging auxiliary image descriptions for dense video captioning
Emre Boran, Aykut Erdem, Nazli Ikizler-Cinbis, Erkut Erdem, Pranava Swaroop Madhyastha, Lucia Specia |
Pattern Recognit. Lett. | 4 |
| 2021 | Generating visual story graphs with application to photo album summarization
Bora Celikkale, Goksu Erdogan, Aykut Erdem, Erkut Erdem |
Signal Process. Image Commun. | 4 |
| 2021 | Synthetic18K: Learning better representations for person re-ID and attribute recognition from 1.4 million synthetic images
Onur Can Uner, Cem Aslan, Burak Ercan, Tayfun Ates, Ufuk Celikcan, Aykut Erdem, Erkut Erdem |
Signal Process. Image Commun. | 7 |
| 2021 | Burst Photography for Learning to Enhance Extremely Dark ImagesabstractCapturing images under extremely low-light conditions poses significant challenges for the standard camera pipeline. Images become too dark and too noisy, which makes traditional enhancement techniques almost impossible to apply. Recently, learning-based approaches have shown very promising results for this task since they have substantially more expressive capabilities to allow for improved quality. Motivated by these studies, in this paper, we aim to leverage burst photography to boost the performance and obtain much sharper and more accurate RGB images from extremely dark raw images. The backbone of our proposed framework is a novel coarse-to-fine network architecture that generates high-quality outputs progressively. The coarse network predicts a low-resolution, denoised raw image, which is then fed to the fine network to recover fine-scale details and realistic textures. To further reduce the noise level and improve the color accuracy, we extend this network to a permutation invariant structure so that it takes a burst of low-light images as input and merges information from multiple images at the feature-level. Our experiments demonstrate that our approach leads to perceptually more pleasing results than the state-of-the-art methods by producing more detailed and considerably higher quality images. Ahmet Serdar Karadeniz, Erkut Erdem, Aykut Erdem |
IEEE Trans. Image Process. | 2 |
| 2020 | Belief Regulated Dual Propagation Nets for Learning Action Effects on Groups of Articulated Objects
Ahmet Ercan Tekden, Aykut Erdem, Erkut Erdem, Mert Imre, M. Yunus Seker, Emre Ugur |
ICRA | 3 |
| 2020 | Hedging static saliency models to predict dynamic saliency
Yasin Kavak, Erkut Erdem, Aykut Erdem |
Signal Process. Image Commun. | 2 |
| 2020 | Manipulating Attributes of Natural Scenes via HallucinationabstractIn this study, we explore building a two-stage framework for enabling users to directly manipulate high-level attributes of a natural scene. The key to our approach is a deep generative network that can hallucinate images of a scene as if they were taken in a different season (e.g., during winter), weather condition (e.g., on a cloudy day), or at a different time of the day (e.g., at sunset). Once the scene is hallucinated with the given attributes, the corresponding look is then transferred to the input image while preserving the semantic details intact, giving a photo-realistic manipulation result. As the proposed framework hallucinates what the scene will look like, it does not require any reference style image as commonly utilized in most of the appearance or style transfer approaches. Moreover, it allows to simultaneously manipulate a given scene according to a diverse set of transient attributes within a single model, eliminating the need of training multiple networks per each translation task. Our comprehensive set of qualitative and quantitative results demonstrates the effectiveness of our approach against the competing methods. Levent Karacan, Zeynep Akata, Aykut Erdem, Erkut Erdem |
ACM Trans. Graph. | 4 |
| 2019 | Procedural Reasoning Networks for Understanding Multimodal ProceduresabstractThis paper addresses the problem of comprehending procedural commonsense knowledge. This is a challenging task as it requires identifying key entities, keeping track of their state changes, and understanding temporal and causal relations. Contrary to most of the previous work, in this study, we do not rely on strong inductive bias and explore the question of how multimodality can be exploited to provide a complementary semantic signal. Towards this end, we introduce a new entity-aware neural comprehension model augmented with external relational memory units. Our model learns to dynamically update entity states in relation to each other while reading the text instructions. Our experimental analysis on the visual reasoning tasks in the recently proposed RecipeQA dataset reveals that our approach improves the accuracy of the previously reported models by a large margin. Moreover, we find that our model learns effective dynamic representations of entities even though we do not use any supervision at the level of entity states. Mustafa Sercan Amac, Semih Yagcioglu, Aykut Erdem, Erkut Erdem |
CoNLL | 4 |
| 2018 | RecipeQA: A Challenge Dataset for Multimodal Comprehension of Cooking RecipesabstractUnderstanding and reasoning about cooking recipes is a fruitful research direction towards enabling machines to interpret procedural text.In this work, we introduce RecipeQA, a dataset for multimodal comprehension of cooking recipes.It comprises of approximately 20K instructional recipes with multiple modalities such as titles, descriptions and aligned set of images.With over 36K automatically generated question-answer pairs, we design a set of comprehension and reasoning tasks that require joint understanding of images and text, capturing the temporal flow of events and making sense of procedural knowledge.Our preliminary results indicate that RecipeQA will serve as a challenging test bed and an ideal benchmark for evaluating machine comprehension systems.The data and leaderboard are available at http://hucvl.github.io/recipeqa. Semih Yagcioglu, Aykut Erdem, Erkut Erdem, Nazli Ikizler-Cinbis |
EMNLP | 3 |
| 2018 | Spatio-Temporal Saliency Networks for Dynamic Saliency PredictionabstractComputational saliency models for still images have gained significant popularity in recent years. Saliency prediction from videos, on the other hand, has received relatively little interest from the community. Motivated by this, in this paper, we study the use of deep learning for dynamic saliency prediction and propose the so-called spatio-temporal saliency networks. The key to our models is the architecture of two-stream networks where we investigate different fusion mechanisms to integrate spatial and temporal information. We evaluate our models on the dynamic images and eye movements and University of Central Florida-Sports datasets and present highly competitive results against the existing state-of-the-art models. We also carry out some experiments on a number of still images from the MIT300 dataset by exploiting the optical flow maps predicted from these images. Our results show that considering inherent motion information in this way can be helpful for static saliency estimation. Çagdas Bak, Aysun Kocak, Erkut Erdem, Aykut Erdem |
IEEE Trans. Multim. | 3 |
| 2017 | Re-evaluating Automatic Metrics for Image CaptioningabstractMert Kilickaya, Aykut Erdem, Nazli Ikizler-Cinbis, Erkut Erdem. Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics: Volume 1, Long Papers. 2017. Mert Kilickaya, Aykut Erdem, Nazli Ikizler-Cinbis, Erkut Erdem |
EACL (1) | 4 |
| 2017 | Automatic Description Generation from Images: A Survey of Models, Datasets, and Evaluation Measures (Extended Abstract)abstractAutomatic image description generation is a challenging problem that has recently received a large amount of interest from the computer vision and natural language processing communities. In this survey, we classify the known approaches based on how they conceptualise this problem and provide a review of existing models, highlighting their advantages and disadvantages. Moreover, we give an overview of the benchmark image-text datasets and the evaluation measures that have been developed to assess the quality of machine-generated descriptions. Finally we explore future directions in the area of automatic image description. Raffaella Bernardi, Ruken Cakici, Desmond Elliott, Aykut Erdem, Erkut Erdem, Nazli Ikizler-Cinbis, Frank Keller, Adrian Muscat, Barbara Plank |
IJCAI | 5 |
| 2017 | Data-driven image captioning via salient region discoveryabstractIn the past few years, automatically generating descriptions for images has attracted a lot of attention in computer vision and natural language processing research. Among the existing approaches, data‐driven methods have been proven to be highly effective. These methods compare the given image against a large set of training images to determine a set of relevant images, then generate a description using the associated captions. In this study, the authors propose to integrate an object‐based semantic image representation into a deep features‐based retrieval framework to select the relevant images. Moreover, they present a novel phrase selection paradigm and a sentence generation model which depends on a joint analysis of salient regions in the input and retrieved images within a clustering framework. The authors demonstrate the effectiveness of their proposed approach on Flickr8K and Flickr30K benchmark datasets and show that their model gives highly competitive results compared with the state‐of‐the‐art models. Mert Kilickaya, Burak Kerim Akkus, Ruken Cakici, Aykut Erdem, Erkut Erdem, Nazli Ikizler-Cinbis |
IET Comput. Vis. | 5 |
| 2017 | A comparative study for feature integration strategies in dynamic saliency estimation
Yasin Kavak, Erkut Erdem, Aykut Erdem |
Signal Process. Image Commun. | 2 |
| 2017 | Alpha Matting With KL-Divergence-Based Sparse SamplingabstractIn this paper, we present a new sampling-based alpha matting approach for the accurate estimation of foreground and background layers of an image. Previous sampling-based methods typically rely on certain heuristics in collecting representative samples from known regions, and thus their performance deteriorates if the underlying assumptions are not satisfied. To alleviate this, we take an entirely new approach and formulate sampling as a sparse subset selection problem where we propose to pick a small set of candidate samples that best explains the unknown pixels. Moreover, we describe a new dissimilarity measure for comparing two samples which is based on KL-divergence between the distributions of features extracted in the vicinity of the samples. The proposed framework is general and could be easily extended to video matting by additionally taking temporal information into account in the sampling process. Evaluation on standard benchmark data sets for image and video matting demonstrates that our approach provides more accurate results compared with the state-of-the-art methods. Levent Karacan, Aykut Erdem, Erkut Erdem |
IEEE Trans. Image Process. | 3 |
| 2016 | An Objective Deghosting Quality Metric for HDR ImagesabstractAbstract Reconstructing high dynamic range (HDR) images of a complex scene involving moving objects and dynamic backgrounds is prone to artifacts. A large number of methods have been proposed that attempt to alleviate these artifacts, known as HDR deghosting algorithms. Currently, the quality of these algorithms are judged by subjective evaluations, which are tedious to conduct and get quickly outdated as new algorithms are proposed on a rapid basis. In this paper, we propose an objective metric which aims to simplify this process. Our metric takes a stack of input exposures and the deghosting result and produces a set of artifact maps for different types of artifacts. These artifact maps can be combined to yield a single quality score. We performed a subjective experiment involving 52 subjects and 16 different scenes to validate the agreement of our quality scores with subjective judgements and observed a concordance of almost 80%. Our metric also enables a novel application that we call as hybrid deghosting, in which the output of different deghosting algorithms are combined to obtain a superior deghosting result. Okan Tarhan Tursun, Ahmet Oguz Akyüz, Aykut Erdem, Erkut Erdem |
Comput. Graph. Forum | 4 |
| 2016 | Automatic Description Generation from Images: A Survey of Models, Datasets, and Evaluation MeasuresabstractAutomatic description generation from natural images is a challenging problem that has recently received a large amount of interest from the computer vision and natural language processing communities. In this survey, we classify the existing approaches based on how they conceptualize this problem, viz., models that cast description as either generation problem or as a retrieval problem over a visual or multimodal representational space. We provide a detailed review of existing models, highlighting their advantages and disadvantages. Moreover, we give an overview of the benchmark image datasets and the evaluation measures that have been developed to assess the quality of machine-generated image descriptions. Finally we extrapolate future directions in the area of automatic image description generation. Raffaella Bernardi, Ruken Cakici, Desmond Elliott, Aykut Erdem, Erkut Erdem, Nazli Ikizler-Cinbis, Frank Keller, Adrian Muscat, Barbara Plank |
J. Artif. Intell. Res. | 5 |
| 2016 | Deformable part-based tracking by coupled global and local correlation filters
Osman Akin, Erkut Erdem, Aykut Erdem, Krystian Mikolajczyk |
J. Vis. Commun. Image Represent. | 2 |
| 2015 | Image Matting with KL-Divergence Based Sparse SamplingabstractPrevious sampling-based image matting methods typically rely on certain heuristics in collecting representative samples from known regions, and thus their performance deteriorates if the underlying assumptions are not satisfied. To alleviate this, in this paper we take an entirely new approach and formulate sampling as a sparse subset selection problem where we propose to pick a small set of candidate samples that best explains the unknown pixels. Moreover, we describe a new distance measure for comparing two samples which is based on KL-divergence between the distributions of features extracted in the vicinity of the samples. Using a standard benchmark dataset for image matting, we demonstrate that our approach provides more accurate results compared with the state-of-the-art methods. Levent Karacan, Aykut Erdem, Erkut Erdem |
ICCV | 3 |
| 2015 | City Scale Image Geolocalization via Dense Scene AlignmentabstractPredicting where a photo was taken is quite important and yet a challenging task for computer vision algorithms. Our motivation is to solve this difficult problem in a city scale setting by employing a data-driven approach. In order to pursue this goal, we developed a fast and robust scene matching method that follows a coarse-to-fine strategy. In particular, we combine scene retrieval via global features and dense scene alignment and use a large set of geo-tagged images of downtown San Francisco in our evaluation. The experimental results show that the proposed approach, despite its simplicity, is surprisingly effective and achieves comparable results with the state-of-the-art. Semih Yagcioglu, Erkut Erdem, Aykut Erdem |
WACV | 2 |
| 2015 | The State of the Art in HDR Deghosting: A Survey and EvaluationabstractAbstract Obtaining a high quality high dynamic range (HDR) image in the presence of camera and object movement has been a long‐standing challenge. Many methods, known as HDR deghosting algorithms, have been developed over the past ten years to undertake this challenge. Each of these algorithms approaches the deghosting problem from a different perspective, providing solutions with different degrees of complexity, solutions that range from rudimentary heuristics to advanced computer vision techniques. The proposed solutions generally differ in two ways: (1) how to detect ghost regions and (2) what to do to eliminate ghosts. Some algorithms choose to completely discard moving objects giving rise to HDR images which only contain the static regions. Some other algorithms try to find the best image to use for each dynamic region. Yet others try to register moving objects from different images in the spirit of maximizing dynamic range in dynamic regions. Furthermore, each algorithm may introduce different types of artifacts as they aim to eliminate ghosts. These artifacts may come in the form of noise, broken objects, under‐ and over‐exposed regions, and residual ghosting. Given the high volume of studies conducted in this field over the recent years, a comprehensive survey of the state of the art is required. Thus, the first goal of this paper is to provide this survey. Secondly, the large number of algorithms brings about the need to classify them. Thus the second goal of this paper is to propose a taxonomy of deghosting algorithms which can be used to group existing and future algorithms into meaningful classes. Thirdly, the existence of a large number of algorithms brings about the need to evaluate their effectiveness, as each new algorithm claims to outperform its precedents. Therefore, the last goal of this paper is to share the results of a subjective experiment which aims to evaluate various state‐of‐the‐art deghosting algorithms. Okan Tarhan Tursun, Ahmet Oguz Akyüz, Aykut Erdem, Erkut Erdem |
Comput. Graph. Forum | 4 |
| 2015 | Predicting memorability of images using attention-driven spatial pooling and image semantics
Bora Celikkale, Aykut Erdem, Erkut Erdem |
Image Vis. Comput. | 3 |
| 2014 | Top down saliency estimation via superpixel-based discriminative dictionaries
Aysun Kocak, Kemal Cizmeciler, Aykut Erdem, Erkut Erdem |
BMVC | 4 |
| 2013 | Structure-preserving image smoothing via region covariancesabstractRecent years have witnessed the emergence of new image smoothing techniques which have provided new insights and raised new questions about the nature of this well-studied problem. Specifically, these models separate a given image into its structure and texture layers by utilizing non-gradient based definitions for edges or special measures that distinguish edges from oscillations. In this study, we propose an alternative yet simple image smoothing approach which depends on covariance matrices of simple image features, aka the region covariances. The use of second order statistics as a patch descriptor allows us to implicitly capture local structure and texture information and makes our approach particularly effective for structure extraction from texture. Our experimental results have shown that the proposed approach leads to better image decompositions as compared to the state-of-the-art methods and preserves prominent edges and shading well. Moreover, we also demonstrate the applicability of our approach on some image editing and manipulation tasks such as image abstraction, texture and detail enhancement, image composition, inverse halftoning and seam carving. Levent Karacan, Erkut Erdem, Aykut Erdem |
ACM Trans. Graph. | 2 |
| 2012 | Fragments based tracking with adaptive cue integration
Erkut Erdem, Séverine Dubuisson, Isabelle Bloch |
Comput. Vis. Image Underst. | 1 |
| 2012 | Visual tracking by fusing multiple cues with context-sensitive reliabilities
Erkut Erdem, Séverine Dubuisson, Isabelle Bloch |
Pattern Recognit. | 1 |
| 2009 | Segmentation using the edge strength function as a shape prior within a local deformation modelabstractThis paper presents a new image segmentation framework which employs a shape prior in the form of an edge strength function to introduce a higher-level influence on the segmentation process. We formulate segmentation as the minimization of three coupled functionals, respectively, defining three processes: prior-guided segmentation, shape feature extraction and local deformation estimation. Particularly, the shape feature extraction process is in charge of estimating an edge strength function from the evolving object region. The local deformation estimation process uses this function to determine a meaningful correspondence between a given prior and the evolving object region, and the deformation map estimated in return supervises the segmentation by enforcing the evolving object boundary towards the prior shape. Erkut Erdem, Sibel Tari, Luminita A. Vese |
ICIP | 1 |
| 2008 | Disconnected Skeleton: Shape at Its Absolute ScaleabstractWe present a new skeletal representation along with a matching framework to address the deformable shape recognition problem. The disconnectedness arises as a result of excessive regularization that we use to describe a shape at an attainably coarse scale. Our motivation is to rely on stable properties the shape instead of inaccurately measured secondary details. The new representation does not suffer from the common instability problems of the traditional connected skeletons, and the matching process gives quite successful results on a diverse database of 2D shapes. An important difference of our approach from the conventional use of skeleton is that we replace the local coordinate frame with a global Euclidean frame supported by additional mechanisms to handle articulations and local boundary deformations. As a result, we can produce descriptions that are sensitive to any combination of changes in scale, position, orientation and articulation, as well as invariant ones. Cagri Aslan, Aykut Erdem, Erkut Erdem, Sibel Tari |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2003 | Computer vision based unistroke keyboard system and mouse for the handicappedabstractIn this paper, a unistroke keyboard based on computer vision is described for the handicapped. The keyboard can be made of paper or fabric containing an image of a keyboard, which has an upside down U-shape. It can even be displayed on a computer screen. Each character is represented by a non-overlapping rectangular region on the keyboard image and the user enters a character by illuminating a character region with a laser pointer. The keyboard image is monitored by a camera and illuminated key locations are recognized. During the text entry process the user neither have to turn the laser light off nor raise the laser light from the keyboard. A disabled person who has difficulty using his/her hands may attach the laser pointer to an eyeglass and easily enter text by moving his/her head to point the laser beam on a character location. In addition, a mouse-like device can be developed based on the same principle. The user can move the cursor by moving the laser light on the computer screen which is monitored by a camera. Erkut Erdem, Aykut Erdem, Volkan Atalay, A. Enis Çetin |
ICME | 1 |
| 2002 | Computer vision based mouseabstractWe describe a computer vision based mouse, which can control and command the cursor of a computer or a computerized system using a camera. In order to move the cursor on the computer screen the user simply moves the mouse shaped passive device placed on a surface within the viewing area of the camera. The video generated by the camera is analyzed using computer vision techniques and the computer moves the cursor according to mouse movements. The computer vision based mouse has regions corresponding to buttons for clicking. To click a button the user simply covers one of these regions with his/her finger. Aykut Erdem, Erkut Erdem, Yasemin Yardimci, Volkan Atalay, A. Enis Çetin |
ICASSP | 2 |