EDBT 2026 Demo / reviewers in the wild / expert
Riza Alp Güler
dblp:157/3659
· DBLP profile ↗
16ranked-venue papers
5as first author
5since 2021 · last 2024
0000-0001-5149-5195ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 14 · 4 first-author · 4 since 2021Artificial intelligence and machine learning · 13 · 4 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
12 papers |
3D vision · 48% Generative modeling · 23% Face, body and person analysis · 16% | |
| Computer graphics and multimedia
4 papers |
Visual content generation and editing · 51% Geometric modeling and processing · 34% Computer animation and physical simulation · 15% |
Topics — the 29 heaviest of 34, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Generative modeling
diffusion model |
1.4 | 2 | 2024 | SPAD: Spatially Aware Multi-View Diffusers · CVPR 2024 Repurposing Diffusion Inpainters for Novel View Synthesis · SIGGRAPH Asia 2023 |
Computer vision › 3D vision
human mesh recovery |
1.2 | 2 | 2024 | MeshPose: Unifying DensePose and 3D Body Mesh reconstruction · CVPR 2024 BLSM: A Bone-Level Skinned Model of the Human Mesh · ECCV (5) 2020 |
Computer vision › 3D vision
3d human reconstruction |
1.1 | 2 | 2024 | MeshPose: Unifying DensePose and 3D Body Mesh reconstruction · CVPR 2024 HoloPose: Holistic 3D Human Reconstruction In-The-Wild · CVPR 2019 |
Computer vision › 3D vision
novel view synthesis |
0.9 | 2 | 2024 | Repurposing Diffusion Inpainters for Novel View Synthesis · SIGGRAPH Asia 2023 SPAD: Spatially Aware Multi-View Diffusers · CVPR 2024 |
Computer vision › Face, body and person analysis › human pose estimation › 3d pose estimation
dense pose estimation |
0.8 | 3 | 2019 | Slim DensePose: Thrifty Learning From Sparse Annotations and Motion Cues · CVPR 2019 DensePose: Dense Human Pose Estimation in the Wild · CVPR 2018 HoloPose: Holistic 3D Human Reconstruction In-The-Wild · CVPR 2019 |
Machine learning › Generative modeling › diffusion model › 3d-aware diffusion
multi-view diffusion |
0.8 | 1 | 2024 | SPAD: Spatially Aware Multi-View Diffusers · CVPR 2024 |
Visual content generation and editing
3d content generation |
0.8 | 1 | 2024 | SPAD: Spatially Aware Multi-View Diffusers · CVPR 2024 |
Visual content generation and editing › image generation
multi-view image generation |
0.8 | 1 | 2024 | SPAD: Spatially Aware Multi-View Diffusers · CVPR 2024 |
Machine learning › Generative modeling › diffusion model › image restoration
diffusion-based inpainting |
0.7 | 1 | 2023 | Repurposing Diffusion Inpainters for Novel View Synthesis · SIGGRAPH Asia 2023 |
Computer vision › 3D vision › novel view synthesis
single-image novel view synthesis |
0.7 | 1 | 2023 | Repurposing Diffusion Inpainters for Novel View Synthesis · SIGGRAPH Asia 2023 |
Computer animation and physical simulation
character animation |
0.7 | 1 | 2023 | Invertible Neural Skinning · CVPR 2023 |
Geometric modeling and processing › shape modeling › human body modeling
clothed human modeling |
0.7 | 1 | 2023 | Invertible Neural Skinning · CVPR 2023 |
Geometric modeling and processing › shape modeling
human body modeling |
0.7 | 1 | 2023 | Invertible Neural Skinning · CVPR 2023 |
Computer vision › 3D vision › correspondence estimation
dense correspondence |
0.6 | 2 | 2018 | DensePose: Dense Human Pose Estimation in the Wild · CVPR 2018 DenseReg: Fully Convolutional Dense Shape Regression In-the-Wild · CVPR 2017 |
Computer vision › 3D vision › pose estimation
3d hand pose estimation |
0.4 | 1 | 2020 | Weakly-Supervised Mesh-Convolutional Hand Reconstruction in the Wild · CVPR 2020 |
Computer vision › 3D vision › 3d human reconstruction
hand mesh reconstruction |
0.4 | 1 | 2020 | Weakly-Supervised Mesh-Convolutional Hand Reconstruction in the Wild · CVPR 2020 |
Computer vision › 3D vision
human body modeling |
0.4 | 1 | 2020 | BLSM: A Bone-Level Skinned Model of the Human Mesh · ECCV (5) 2020 |
Computer vision › Face, body and person analysis
human pose estimation |
0.4 | 1 | 2019 | HoloPose: Holistic 3D Human Reconstruction In-The-Wild · CVPR 2019 |
Computer vision › 3D vision › 3d human reconstruction
single-view human reconstruction |
0.4 | 1 | 2019 | HoloPose: Holistic 3D Human Reconstruction In-The-Wild · CVPR 2019 |
Machine learning › Deep learning architectures and training
autoencoder |
0.3 | 1 | 2018 | Deforming Autoencoders: Unsupervised Disentangling of Shape and Appearance · ECCV (10) 2018 |
Machine learning › Representation and self-supervised learning › representation learning
disentangled representation learning |
0.3 | 1 | 2018 | Deforming Autoencoders: Unsupervised Disentangling of Shape and Appearance · ECCV (10) 2018 |
Machine learning › Generative modeling › image translation
human pose transfer |
0.3 | 1 | 2018 | Dense Pose Transfer · ECCV (3) 2018 |
Machine learning › Representation and self-supervised learning › representation learning › disentangled representation learning
shape and appearance disentanglement |
0.3 | 1 | 2018 | Deforming Autoencoders: Unsupervised Disentangling of Shape and Appearance · ECCV (10) 2018 |
Visual content generation and editing
image generation |
0.3 | 1 | 2018 | Dense Pose Transfer · ECCV (3) 2018 |
Visual content generation and editing › image generation › person image generation
pose-guided person image synthesis |
0.3 | 1 | 2018 | Dense Pose Transfer · ECCV (3) 2018 |
Computer vision › Face, body and person analysis
face alignment |
0.3 | 1 | 2017 | DenseReg: Fully Convolutional Dense Shape Regression In-the-Wild · CVPR 2017 |
Machine learning › Deep learning architectures and training › feedforward neural network
invertible neural network |
0.2 | 1 | 2023 | Invertible Neural Skinning · CVPR 2023 |
Machine learning › Learning paradigms
weakly supervised learning |
0.1 | 1 | 2020 | Weakly-Supervised Mesh-Convolutional Hand Reconstruction in the Wild · CVPR 2020 |
Geometric modeling and processing
mesh deformation |
0.1 | 1 | 2020 | BLSM: A Bone-Level Skinned Model of the Human Mesh · ECCV (5) 2020 |
Methods — techniques the papers use, named apart from their topics
linear blend skinning · 2.2plücker coordinates · 1.5epipolar geometry · 1.5cross-view attention · 1.5invertible neural network · 1.3differentiable rendering · 1.3weak supervision · 1.2end-to-end training · 0.8monocular depth estimation · 0.7epipolar masking · 0.7skinning · 0.4dense pose transfer · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | SPAD: Spatially Aware Multi-View DiffusersabstractWe present SPAD, a novel approach for creating con-sistent multi-view images from text prompts or single images. To enable multi-view generation, we repurpose a pre-trained 2D diffusion model by extending its self-attention layers with cross-view interactions, and fine-tune it on a high quality subset of Objaverse. We find that a naive extension of the self-attention proposed in prior work (e.g., MV-Dream) leads to content copying between views. Therefore, we explicitly constrain the cross-view attention based on epipolar geometry. To further enhance 3D consistency, we utilize Plücker coordinates derived from camera rays and inject them as positional encoding. This enables SPAD to reason over spatial proximity in 3D well. Compared to concurrent works that can only generate views at fixed azimuth and elevation (e.g., MVDream, SyncDreamer), SPAD offers full camera control and achieves state-of-the-art results in novel view synthesis on unseen objects from the Objaverse and Google Scanned Objects datasets. Finally, we demon-strate that text-to-3D generation using SPAD prevents the multi-face Janus issue. Yash Kant, Aliaksandr Siarohin, Ziyi Wu 0002, Michael Vasilkovsky, Guocheng Qian, Jian Ren 0005, Riza Alp Güler, Bernard Ghanem, Sergey Tulyakov, Igor Gilitschenski |
CVPR | 7 |
| 2024 | MeshPose: Unifying DensePose and 3D Body Mesh reconstructionabstractDensePose provides a pixel-accurate association of images with 3D mesh coordinates, but does not provide a 3D mesh, while Human Mesh Reconstruction (HMR) systems have high 2D reprojection error, as measured by DensePose localization metrics. In this work we introduce MeshPose to jointly tackle DensePose and HMR. For this we first introduce new losses that allow us to use weak DensePose supervision to accurately localize in 2D a subset of the mesh vertices (‘VertexPose’). We then lift these vertices to 3D, yielding a low-poly body mesh (‘MeshPose’). Our system is trained in an end -to-end manner and is the first HMR method to attain competitive DensePose accuracy, while also being lightweight and amenable to efficient inference, making it suitable for real-time AR applications. Eric-Tuan Le, Antonis Kakolyris, Petros Koutras, Himmy Tam, Efstratios Skordos, George Papandreou, Riza Alp Güler, Iasonas Kokkinos |
CVPR | 7 |
| 2023 | Invertible Neural SkinningabstractBuilding animatable and editable models of clothed humans from raw 3D scans and poses is a challenging problem. Existing reposing methods suffer from the limited expressiveness of Linear Blend Skinning (LBS), require costly mesh extraction to generate each new pose, and typically do not preserve surface correspondences across different poses. In this work, we introduce Invertible Neural Skinning (INS) to address these shortcomings. To maintain correspondences, we propose a Pose-conditioned Invertible Network (PIN) architecture, which extends the LBS process by learning additional pose-varying deformations. Next, we combine PIN with a differentiable LBS module to build an expressive and end-to-end Invertible Neural Skinning (INS) pipeline. We demonstrate the strong performance of our method by outperforming the state-of-the-art reposing techniques on clothed humans and preserving surface correspondences, while being an order of magnitude faster. We also perform an ablation study, which shows the usefulness of our pose-conditioning formulation, and our qualitative results display that INS can rectify artefacts introduced by LBS well. Yash Kant, Aliaksandr Siarohin, Riza Alp Güler, Menglei Chai, Jian Ren 0005, Sergey Tulyakov, Igor Gilitschenski |
CVPR | 3 |
| 2023 | Repurposing Diffusion Inpainters for Novel View SynthesisabstractIn this paper, we present a method for generating consistent novel views from a single source image. Our approach focuses on maximizing the reuse of visible pixels from the source image. To achieve this, we use a monocular depth estimator that transfers visible pixels from the source view to the target view. Starting from a pre-trained 2D inpainting diffusion model, we train our method on the large-scale Objaverse dataset to learn 3D object priors. While training we use a novel masking mechanism based on epipolar lines to further improve the quality of our approach. This allows our framework to perform zero-shot novel view synthesis on a variety of objects. We evaluate the zero-shot abilities of our framework on three challenging datasets: Google Scanned Objects, Ray Traced Multiview, and Common Objects in 3D. Yash Kant, Aliaksandr Siarohin, Michael Vasilkovsky, Riza Alp Güler, Jian Ren 0005, Sergey Tulyakov, Igor Gilitschenski |
SIGGRAPH Asia | 4 |
| 2022 | Context-Self Contrastive Pretraining for Crop Type Semantic SegmentationabstractIn this paper, we propose a fully supervised pre-training scheme based on contrastive learning particularly tailored to dense classification tasks. The proposedContext-Self Contrastive Loss(CSCL) learns an embedding space that makes semantic boundaries pop-up by use of a similarity metric between every location in a training sample and its local context. For crop type semantic segmentation fromSatellite Image Time Series(SITS) we find performance at parcel boundaries to be a critical bottleneck and explain how CSCL tackles the underlying cause of that problem, improving the state-of-the-art performance in this task. Additionally, using images from theSentinel-2(S2) satellite missions we compile the largest, to our knowledge, SITS dataset densely annotated by crop type and parcel identities, which we make publicly available together with the data generation pipeline. Using that data we find CSCL, even with minimal pre-training, to improve all respective baselines and present a process for semantic segmentation at greater resolution than that of the input images for obtaining crop classes at a more granular level. The code and instructions to download the data can be found in https://github.com/michaeltrs/DeepSatModels. Michail Tarasiou, Riza Alp Güler, Stefanos Zafeiriou |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2020 | Weakly-Supervised Mesh-Convolutional Hand Reconstruction in the WildabstractWe introduce a simple and effective network architecture for monocular 3D hand pose estimation consisting of an image encoder followed by a mesh convolutional decoder that is trained through a direct 3D hand mesh reconstruction loss. We train our network by gathering a large-scale dataset of hand action in YouTube videos and use it as a source of weak supervision. Our weakly-supervised mesh convolutions-based system largely outperforms state-of-the-art methods, even halving the errors on the in the wild benchmark. The dataset and additional resources are available at https://arielai.com/mesh_hands. Dominik Kulon, Riza Alp Güler, Iasonas Kokkinos, Michael M. Bronstein, Stefanos Zafeiriou |
CVPR | 2 |
| 2020 | BLSM: A Bone-Level Skinned Model of the Human Mesh
Haoyang Wang 0002, Riza Alp Güler, Iasonas Kokkinos, George Papandreou, Stefanos Zafeiriou |
ECCV (5) | 2 |
| 2019 | Single Image 3D Hand Reconstruction with Mesh Convolutions
Dominik Kulon, Haoyang Wang 0002, Riza Alp Güler, Michael M. Bronstein, Stefanos Zafeiriou |
BMVC | 3 |
| 2019 | HoloPose: Holistic 3D Human Reconstruction In-The-WildabstractWe introduce HoloPose, a method for holistic monocular 3D human body reconstruction. We first introduce a part-based model for 3D model parameter regression that allows our method to operate in-the-wild, gracefully handling severe occlusions and large pose variation. We further train a multi-task network comprising 2D, 3D and Dense Pose estimation to drive the 3D reconstruction task. For this we introduce an iterative refinement method that aligns the model-based 3D estimates of 2D/3D joint positions and DensePose with their image-based counterparts delivered by CNNs, achieving both model-based, global consistency and high spatial accuracy thanks to the bottom-up CNN processing. We validate our contributions on challenging benchmarks, showing that our method allows us to get both accurate joint and 3D surface estimates while operating at more than 10fps in-the-wild. More information about our approach, including videos and demos is available at http://arielai.com/holopose. Riza Alp Güler, Iasonas Kokkinos |
CVPR | 1 |
| 2019 | Slim DensePose: Thrifty Learning From Sparse Annotations and Motion CuesabstractDensePose supersedes traditional landmark detectors by densely mapping image pixels to body surface coordinates. This power, however, comes at a greatly increased annotation cost, as supervising the model requires to manually label hundreds of points per pose instance. In this work, we thus seek methods to significantly slim down the DensePose annotations, proposing more efficient data collection strategies. In particular, we demonstrate that if annotations are collected in video frames, their efficacy can be multiplied for free by using motion cues. To explore this idea, we introduce DensePose-Track, a dataset of videos where selected frames are annotated in the traditional DensePose manner. Then, building on geometric properties of the DensePose mapping, we use the video dynamic to propagate ground-truth annotations in time as well as to learn from Siamese equivariance constraints. Having performed exhaustive empirical evaluation of various data annotation and learning strategies, we demonstrate that doing so can deliver significantly improved pose estimation results over strong baselines. However, despite what is suggested by some recent works, we show that merely synthesizing motion patterns by applying geometric transformations to isolated frames is significantly less effective, and that motion cues help much more when they are extracted from videos. Natalia Neverova, James Thewlis, Riza Alp Güler, Iasonas Kokkinos, Andrea Vedaldi |
CVPR | 3 |
| 2018 | DensePose: Dense Human Pose Estimation in the WildabstractIn this work we establish dense correspondences between an RGB image and a surface-based representation of the human body, a task we refer to as dense human pose estimation. We gather dense correspondences for 50K persons appearing in the COCO dataset by introducing an efficient annotation pipeline. We then use our dataset to train CNN-based systems that deliver dense correspondence 'in the wild', namely in the presence of background, occlusions and scale variations. We improve our training set's effectiveness by training an inpainting network that can fill in missing ground truth values and report improvements with respect to the best results that would be achievable in the past. We experiment with fully-convolutional networks and region-based models and observe a superiority of the latter. We further improve accuracy through cascading, obtaining a system that delivers highly-accurate results at multiple frames per second on a single gpu. Supplementary materials, data, code, and videos are provided on the project page http://densepose.org. Riza Alp Güler, Natalia Neverova, Iasonas Kokkinos |
CVPR | 1 |
| 2018 | Dense Pose Transfer
Natalia Neverova, Riza Alp Güler, Iasonas Kokkinos |
ECCV (3) | 2 |
| 2018 | Deforming Autoencoders: Unsupervised Disentangling of Shape and Appearance
Zhixin Shu, Mihir Sahasrabudhe, Riza Alp Güler, Dimitris Samaras, Nikos Paragios, Iasonas Kokkinos |
ECCV (10) | 3 |
| 2017 | DenseReg: Fully Convolutional Dense Shape Regression In-the-WildabstractIn this paper we propose to learn a mapping from image pixels into a dense template grid through a fully convolutional network. We formulate this task as a regression problem and train our network by leveraging upon manually annotated facial landmarks "in-the-wild". We use such landmarks to establish a dense correspondence field between a three-dimensional object template and the input image, which then serves as the ground-truth for training our regression system. We show that we can combine ideas from semantic segmentation with regression networks, yielding a highly-accurate quantized regression architecture. Our system, called DenseReg, allows us to estimate dense image-to-template correspondences in a fully convolutional manner. As such our network can provide useful correspondence information as a stand-alone system, while when used as an initialization for Statistical Deformable Models we obtain landmark localization results that largely outperform the current state-of-the-art on the challenging 300W benchmark. We thoroughly evaluate our method on a host of facial analysis tasks, and demonstrate its use for other correspondence estimation tasks, such as the human body and the human ear. DenseReg code is made available at http://alpguler.com/DenseReg.html along with supplementary materials. Riza Alp Güler, George Trigeorgis, Epameinondas Antonakos, Patrick Snape, Stefanos Zafeiriou, Iasonas Kokkinos |
CVPR | 1 |
| 2016 | Landmarks inside the shape: Shape matching using image descriptors
Riza Alp Güler, Sibel Tari, Gozde Unal |
Pattern Recognit. | 1 |
| 2014 | Screened Poisson Hyperfields for Shape CodingabstractWe present a novel perspective on shape characterization using the screened Poisson equation. We discuss that the effect of the screening parameter is a change of measure of the underlying metric space. Screening also indicates a conditioned random walker biased by the choice of measure. A continuum of shape fields is created by varying the screening parameter or, equivalently, the bias of the random walker. In addition to creating a regional encoding of the diffusion with a different bias, we further break down the influence of boundary interactions by considering a number of independent random walks, each emanating from a certain boundary point, whose superposition yields the screened Poisson field. Probing the screened Poisson equation from these two complementary perspectives leads to a high-dimensional hyperfield: a rich characterization of the shape that encodes global, local, interior, and boundary interactions. To extract particular shape information as needed in a compact way from the hyperfield, we apply various decompositions either to unveil parts of a shape or parts of a boundary or to create consistent mappings. The latter technique involves lower-dimensional embeddings, which we call screened Poisson encoding maps (SPEM). The expressive power of the SPEM is demonstrated via illustrative experiments as well as a quantitative shape retrieval experiment over a public benchmark database on which the SPEM method shows a high-ranking performance among the existing state-of-the-art shape retrieval methods. Riza Alp Güler, Sibel Tari, Gozde Unal |
SIAM J. Imaging Sci. | 1 |