Akin Caliskan

dblp:189/3916 · DBLP profile ↗
← Back
9ranked-venue papers
3as first author
7since 2021 · last 2026
0000-0003-2918-5603ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 2 first-author · 7 since 2021Artificial intelligence and machine learning · 6 · 2 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 Polyglot: Multilingual Style Preserving Speech-Driven Facial Animation
Federico Nocentini, Kwanggyoon Seo, Qingju Liu, Claudio Ferrari, Stefano Berretti, David Ferman, Hyeongwoo Kim, Pablo Garrido 0001, Akin Caliskan
FG9
2025 MVL-Net: Pairwise Learning for Multi-View Multiple People Labelling
abstract
In the multi-view domain, it is challenging to correctly label multiple people across viewpoints because of occlusions, visual ambiguities, appearance variation, etc. Deep learning, although having witnessed remarkable success in computer vision tasks, still remains underexplored for the multi-view labelling task, due to the lack of labelled multi-view datasets. In this paper, we propose a novel end-to-end deep neural network named Multi-View Labelling network (MVL-net) that addresses this issue. To overcome the dataset shortage, a large-scale multi-view dataset is generated by combining 3D human models and panoramic backgrounds, along with human poses and realistic rendering. In the proposed MVL-net, we first incorporate Transformer blocks to capture the non-local information for multi-view feature extraction. A matching net is then introduced to achieve multiple people labelling, by predicting matching confidence scores for pairwise instances from two views, thus addressing the problem of the unknown number of people when labelling across views. An additional geometry feature obtained from the epipolar geometry is integrated to leverage multi-view cues during training. To the best of our knowledge, the MVL-net is the first work using deep learning to train a multi-view labelling network. Comprehensive experiments on both synthetic and real-world datasets demonstrate the effectiveness of the proposed method, which outperforms the existing state-of-the-art approaches.
Yue Zhang 0082, Akin Caliskan, Mai Xu, Adrian Hilton 0001, Jean-Yves Guillemaut
IEEE Trans. Multim.2
2024 PAV: Personalized Head Avatar from Unstructured Video Collection
Akin Caliskan, Berkay Kicanaoglu, Hyeongwoo Kim
ECCV (41)1
2023 RANA: Relightable Articulated Neural Avatars
abstract
We propose RANA, a relightable and articulated neural avatar for the synthesis of humans under arbitrary viewpoints, body poses, and lighting. We only require a short video clip of the person to create the avatar and assume no knowledge about the lighting environment. We present a novel framework to model humans while disentangling their geometry, texture, and lighting environment from monocular RGB videos. To simplify this otherwise ill-posed task we first estimate the coarse geometry and texture of the person via SMPL+D model fitting and then learn an articulated neural representation for higher quality image synthesis. RANA first generates the normal and albedo maps of the person in any given target body pose and then uses spherical harmonics lighting to generate the shaded image in the target lighting environment. We also propose to pre-train RANA using synthetic images and demonstrate that it leads to better disentanglement between geometry and texture while also improving robustness to novel body poses. Finally, we also present a new photo-realistic synthetic dataset, Relighting Human, to quantitatively evaluate the performance of the proposed approach.
Umar Iqbal 0001, Akin Caliskan, Koki Nagano, Sameh Khamis, Pavlo Molchanov 0001, Jan Kautz
ICCV2
2022 LiP-Flow: Learning Inference-Time Priors for Codec Avatars via Normalizing Flows in Latent Space
Emre Aksan, Shugao Ma, Akin Caliskan, Stanislav Pidhorskyi, Alexander Richard, Shih-En Wei, Jason M. Saragih, Otmar Hilliges
ECCV (26)3
2021 Multi-Person Implicit Reconstruction From a Single Image
abstract
We present a new end-to-end learning framework to obtain detailed and spatially coherent reconstructions of multiple people from a single image. Existing multi-person methods suffer from two main drawbacks: they are often model-based and therefore cannot capture accurate 3D models of people with loose clothing and hair; or they require manual intervention to resolve occlusions or interactions. Our method addresses both limitations by introducing the first end-to-end learning approach to perform model-free implicit reconstruction for realistic 3D capture of multiple clothed people in arbitrary poses (with occlusions) from a single image. Our network simultaneously estimates the 3D geometry of each person and their 6DOF spatial locations, to obtain a coherent multi-human reconstruction. In addition, we introduce a new synthetic dataset that depicts images with a varying number of inter-occluded humans and a variety of clothing and hair styles. We demonstrate robust, high-resolution reconstructions on images of multiple humans with complex occlusions, loose clothing and a large variety of poses and scenes. Our quantitative evaluation on both synthetic and real world datasets demonstrates state-of-the-art performance with significant improvements in the accuracy and completeness of the reconstructions over competing approaches.
Armin Mustafa, Akin Caliskan, Lourdes Agapito, Adrian Hilton 0001
CVPR2
2021 A Novel Multi-View Labelling Network Based on Pairwise Learning
abstract
Correct labelling of multiple people from different viewpoints in complex scenes is a challenging task due to occlusions, visual ambiguities, as well as variations in appearance and illumination. In recent years, deep learning approaches have proved very successful at improving the performance of a wide range of recognition and labelling tasks such as person re-identification and video tracking. However, to date, applications to multi-view tasks have proved more challenging due to the lack of suitably labelled multi-view datasets, which are difficult to collect and annotate. The contributions of this paper are two-fold. First, a synthetic dataset is generated by combining 3D human models and panoramas along with human poses and appearance detail rendering to overcome the shortage of real dataset for multi-view labelling. Second, a novel framework named Multi-View Labelling network (MVL-net) is introduced to leverage the new dataset and unify the multi-view multiple people detection, segmentation and labelling tasks in complex scenes. To the best of our knowledge, this is the first work using deep learning to train a multi-view labelling network. Experiments conducted on both synthetic and real datasets demonstrate that the proposed method outperforms the existing state-of-the-art approaches.
Yue Zhang 0082, Akin Caliskan, Adrian Hilton 0001, Jean-Yves Guillemaut
ICIP2
2020 Multi-view Consistency Loss for Improved Single-Image 3D Reconstruction of Clothed People
Akin Caliskan, Armin Mustafa, Evren Imre, Adrian Hilton 0001
ACCV (1)1
2016 Superpixel based hyperspectral target detection
abstract
Using the spectral signature of a target by means of matching the signature with the pixels of an acquired hyperspectral image has been proven as an effective way of classifying hyperspectral pixels in most of the proposed methods in hyperspectral image analysis. A disadvantage of these methods is however to use only the spectral characteristics of pixels for detection while ignoring the spatial relations between the neighbouring pixels. In this paper, we propose a hyperspectral target detection method which uses also the spatial neigboorhood information as well as the spectral characteristics of hyperspectral pixels. To this end, we first utilize superpixelization method [1] to describe the neigborhood relation between the hyperspectral pixels, which has been previously developed and proved to be better compared to a pioneer state-of-the-art superpixel algorithm, SLIC [2]. Second, we investigate the best representatives for superpixels among different alternatives, such as centroids, medoid and mean, and modify the well-known hyperspectral target detection algorithm using orthogonal subspace projection, DTDCA [3], appropriately for superpixels. The improvements of the proposed approach over DTDCA in terms of the detection and false detection rates are verified on real hyperspectral images taken from wheat and corn fields with a VNIR camera.
Akin Caliskan, Emrecan Bati, Alper Koz, A. Aydin Alatan
IGARSS1