Stavros Tsogkas

dblp:119/1444 · DBLP profile ↗
← Back
17ranked-venue papers
3as first author
7since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 2 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 3 first-author · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
13 papers
3D vision · 40% Robot manipulation · 22% Segmentation and scene understanding · 15%
Computer graphics and multimedia
4 papers
Geometric modeling and processing · 84% Image and video processing · 16%

Topics — the 27 heaviest of 30, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › 3D vision
3d shape representation
1.222025
Probabilistic Directed Distance Fields for Ray-Based Shape Representations · IEEE Trans. Pattern Anal. Mach. Intell. 2025
Geometric Disentanglement for Generative Latent Shape Models · ICCV 2019
Robotics › Robot manipulation
grasping
1.222023
Fast-Grasp'D: Dexterous Multi-finger Grasp Generation Through Differentiable Simulation · ICRA 2023
Grasp'D: Differentiable Contact-Rich Grasp Synthesis for Multi-Fingered Hands · ECCV (6) 2022
Computer vision › 3D vision › 3d reconstruction
single-view 3d reconstruction
1.022022
Representing 3D Shapes with Probabilistic Directed Distance Fields · CVPR 2022
Few-Shot Single-View 3-D Object Reconstruction with Compositional Priors · ECCV (25) 2020
Computer vision › Segmentation and scene understanding
object skeleton detection
0.922021
DeepFlux for Skeleton Detection in the Wild · Int. J. Comput. Vis. 2021
DeepFlux for Skeletons in the Wild · CVPR 2019
Robotics › Robot manipulation › grasping
multifingered grasping
0.712023
Fast-Grasp'D: Dexterous Multi-finger Grasp Generation Through Differentiable Simulation · ICRA 2023
Geometric modeling and processing
shape modeling
0.712023
Disentangling Geometric Deformation Spaces in Generative Latent Shape Models · Int. J. Comput. Vis. 2023
Geometric modeling and processing › shape modeling › implicit modeling
implicit shape representation
0.612022
Representing 3D Shapes with Probabilistic Directed Distance Fields · CVPR 2022
Geometric modeling and processing
shape representation
0.612022
Representing 3D Shapes with Probabilistic Directed Distance Fields · CVPR 2022
Computer vision › 3D vision › 3d shape analysis › symmetry analysis
symmetry detection
0.522019
DeepFlux for Skeletons in the Wild · CVPR 2019
Learning-Based Symmetry Detection in Natural Images · ECCV (7) 2012
Computer vision › 3D vision
3d reconstruction
0.412020
Few-Shot Single-View 3-D Object Reconstruction with Compositional Priors · ECCV (25) 2020
Image and video processing
image segmentation
0.412020
Appearance Shock Grammar for Fast Medial Axis Extraction From Real Images · CVPR 2020
Geometric modeling and processing › skeletonization
medial axis transform
0.412020
Appearance Shock Grammar for Fast Medial Axis Extraction From Real Images · CVPR 2020
Geometric modeling and processing › shape descriptor
shock graph
0.412020
Appearance Shock Grammar for Fast Medial Axis Extraction From Real Images · CVPR 2020
Machine learning › Representation and self-supervised learning › representation learning › disentangled representation learning › disentanglement
latent space disentanglement
0.412019
Geometric Disentanglement for Generative Latent Shape Models · ICCV 2019
Computer vision › 3D vision › 3d shape analysis
shape decomposition
0.312017
AMAT: Medial Axis Transform for Natural Images · ICCV 2017
Computer vision › 3D vision › neural rendering
differentiable rendering
0.312025
Probabilistic Directed Distance Fields for Ray-Based Shape Representations · IEEE Trans. Pattern Anal. Mach. Intell. 2025
Computer vision › Image recognition and object detection
attribute recognition
0.212014
Understanding Objects in Detail with Fine-Grained Attributes · CVPR 2014
Computer vision › Image recognition and object detection › object detection › part-based object detection
deformable part model
0.212014
Segmentation-Aware Deformable Part Models · CVPR 2014
Computer vision › Image recognition and object detection › image classification
fine-grained image classification
0.212014
Understanding Objects in Detail with Fine-Grained Attributes · CVPR 2014
Computer vision › Image recognition and object detection
object detection
0.212014
Segmentation-Aware Deformable Part Models · CVPR 2014
Computer vision › Image recognition and object detection › object detection
part-based object detection
0.212014
Understanding Objects in Detail with Fine-Grained Attributes · CVPR 2014
Computer vision › Segmentation and scene understanding › image segmentation › region-based segmentation
superpixel segmentation
0.212014
Segmentation-Aware Deformable Part Models · CVPR 2014
Machine learning › Generative modeling
3d generative model
0.212022
Representing 3D Shapes with Probabilistic Directed Distance Fields · CVPR 2022
Robotics › Robot manipulation › grasping
multifingered hand
0.212022
Grasp'D: Differentiable Contact-Rich Grasp Synthesis for Multi-Fingered Hands · ECCV (6) 2022
Machine learning › Generative modeling
variational autoencoder
0.112019
Geometric Disentanglement for Generative Latent Shape Models · ICCV 2019
Image and video processing
image reconstruction
0.112017
AMAT: Medial Axis Transform for Natural Images · ICCV 2017
Computer vision › Segmentation and scene understanding
part segmentation
0.112014
Understanding Objects in Detail with Fine-Grained Attributes · CVPR 2014

Methods — techniques the papers use, named apart from their topics

probabilistic modeling · 2.0generative latent shape models · 1.3differentiable simulation · 1.2differentiable rendering · 1.1implicit representation · 0.9gradient-based optimization · 0.7contact modeling · 0.6deep flux · 0.5shock grammar · 0.4few-shot learning · 0.4compositional priors · 0.4appearance-based criteria · 0.4weighted geometric set cover · 0.3clustering · 0.3
YearPublicationVenuePosition
2025 Probabilistic Directed Distance Fields for Ray-Based Shape Representations
abstract
In modern computer vision, the optimal representation of 3D shape remains task-dependent. One fundamental operation applied to such representations is differentiable rendering, which enables learning-based inverse graphics approaches. Standard explicit representations are often easily rendered, but can suffer from limited geometric fidelity, among other issues. On the other hand, implicit representations generally preserve greater fidelity, but suffer from difficulties with rendering, limiting scalability. In this work, we devise Directed Distance Fields (DDFs), which map a ray or oriented point (position and direction) to surface visibility and depth. This enables efficient differentiable rendering, obtaining depth with a single forward pass per pixel, as well as higher-order geometry with only additional backward passes. Using probabilistic DDFs (PDDFs), we can model the inherent discontinuities in the underlying field. We then apply DDFs to single-shape fitting, generative modelling, and 3D reconstruction, showcasing strong performance with simple architectural components via the versatility of our representation. Finally, since the dimensionality of DDFs permits view-dependent geometric artifacts, we conduct a theoretical investigation of the constraints necessary for view consistency. We find a small set of field properties that are sufficient to guarantee a DDF is consistent, without knowing which shape the field is expressing.
Tristan Aumentado-Armstrong, Stavros Tsogkas, Sven J. Dickinson, Allan Douglas Jepson
IEEE Trans. Pattern Anal. Mach. Intell.2
2023 Fast-Grasp'D: Dexterous Multi-finger Grasp Generation Through Differentiable Simulation
abstract
Multi-finger grasping relies on high quality training data, which is hard to obtain: human data is hard to transfer and synthetic data relies on simplifying assumptions that reduce grasp quality. By making grasp simulation differentiable, and contact dynamics amenable to gradient-based optimization, we accelerate the search for high-quality grasps with fewer limiting assumptions. We present Grasp'D-1M: a large-scale dataset for multi-finger robotic grasping, synthesized with Fast-Grasp'D, a novel differentiable grasping simulator. Grasp'D-1M contains one million training examples for three robotic hands (three, four and five-fingered), each with multimodal visual inputs (RGB+depth+segmentation, available in mono and stereo). Grasp synthesis with Fast-Grasp'D is 10x faster than GraspIt! [1] and 20x faster than the prior Grasp'D differentiable simulator [2]. Generated grasps are more stable and contact-rich than GraspIt! grasps, regardless of the distance threshold used for contact generation. We validate the usefulness of our dataset by retraining an existing vision-based grasping pipeline [3] on Grasp'D-1M, and showing a dramatic increase in model performance, predicting grasps with 30% more contact, a 33% higher epsilon metric, and 35% lower simulated displacement. Additional details at fast-graspd.github.io.
Dylan Turpin, Tao Zhong 0003, Shutong Zhang, Guanglei Zhu, Eric Heiden, Miles Macklin, Stavros Tsogkas, Sven J. Dickinson, Animesh Garg
ICRA7
2023 Efficient Flow-Guided Multi-frame De-fencing
abstract
Taking photographs "in-the-wild" is often hindered by fence obstructions that stand between the camera user and the scene of interest, and which are hard or impossible to avoid. De-fencing is the algorithmic process of automatically removing such obstructions from images, revealing the invisible parts of the scene. While this problem can be formulated as a combination of fence segmentation and image inpainting, this often leads to implausible hallucinations of the occluded regions. Existing multi-frame approaches rely on propagating information to a selected keyframe from its temporal neighbors, but they are often inefficient and struggle with alignment of severely obstructed images. In this work we draw inspiration from the video completion literature, and develop a simplified framework for multi-frame de-fencing that computes high quality flow maps directly from obstructed frames, and uses them to accurately align frames. Our primary focus is efficiency and practicality in a real world setting: the input to our algorithm is a short image burst (5 frames) – a data modality commonly available in modern smartphones– and the output is a single reconstructed keyframe, with the fence removed. Our approach leverages simple yet effective CNN modules, trained on carefully generated synthetic data, and outperforms more complicated alternatives real bursts, both quantitatively and qualitatively, while running real-time.
Stavros Tsogkas, Fengjia Zhang, Allan Douglas Jepson, Alex Levinshtein
WACV1
2023 Disentangling Geometric Deformation Spaces in Generative Latent Shape Models
Tristan Aumentado-Armstrong, Stavros Tsogkas, Sven J. Dickinson, Allan Douglas Jepson
Int. J. Comput. Vis.2
2022 Representing 3D Shapes with Probabilistic Directed Distance Fields
abstract
Differentiable rendering is an essential operation in modern vision, allowing inverse graphics approaches to 3D understanding to be utilized in modern machine learning frameworks. Explicit shape representations (voxels, point clouds, or meshes), while relatively easily rendered, often suffer from limited geometric fidelity or topological con-straints. On the other hand, implicit representations (occu-pancy, distance, or radiance fields) preserve greater fidelity, but suffer from complex or inefficient rendering processes, limiting scalability. In this work, we endeavour to address both shortcomings with a novel shape representation that allows fast differentiable rendering within an implicit ar-chitecture. Building on implicit distance representations, we define Directed Distance Fields (DDFs), which map an oriented point (position and direction) to surface visibility and depth. Such a field can render a depth map with a single forward pass per pixel, enable differential surface geometry extraction (e.g., surface normals and curvatures) via network derivatives, be easily composed, and permit extraction of classical unsigned distance fields. Using probabilistic DDFs (PDDFs), we show how to model inherent discontinuities in the underlying field. Finally, we apply our method to fitting single shapes, unpaired 3D-aware generative image modelling, and single-image 3D reconstruction tasks, showcasing strong performance with simple architectural components via the versatility of our representation.
Tristan Aumentado-Armstrong, Stavros Tsogkas, Sven J. Dickinson, Allan Douglas Jepson
CVPR2
2022 Grasp'D: Differentiable Contact-Rich Grasp Synthesis for Multi-Fingered Hands
Dylan Turpin, Liquan Wang, Eric Heiden, Yun-Chun Chen, Miles Macklin, Stavros Tsogkas, Sven J. Dickinson, Animesh Garg
ECCV (6)6
2021 DeepFlux for Skeleton Detection in the Wild
Yongchao Xu, Yukang Wang, Stavros Tsogkas, Jianqiang Wan, Xiang Bai, Sven J. Dickinson, Kaleem Siddiqi
Int. J. Comput. Vis.3
2020 Cycle-Consistent Generative Rendering for 2D-3D Modality Translation
abstract
For humans, visual understanding is inherently generative: given a 3D shape, we can postulate how it would look in the world; given a 2D image, we can infer the 3D structure that likely gave rise to it. We can thus translate between the 2D visual and 3D structural modalities of a given object. In the context of computer vision, this corresponds to a learnable module that serves two purposes: (i) generate a realistic rendering of a 3D object (shape-to-image translation) and (ii) infer a realistic 3D shape from an image (image-to-shape translation). In this paper, we learn such a module while being conscious of the difficulties in obtaining large paired 2D-3D datasets. By leveraging generative domain translation methods, we are able to define a learning algorithm that requires only weak supervision, with unpaired data. The resulting model is not only able to perform 3D shape, pose, and texture inference from 2D images, but can also generate novel textured 3D shapes and renders, similar to a graphics pipeline. More specifically, our method (i) infers an explicit 3D mesh representation, (ii) utilizes example shapes to regularize inference, (iii) requires only an image mask (no keypoints or camera extrinsics), and (iv) has generative capabilities. While prior work explores subsets of these properties, their combination is novel. We demonstrate the utility of our learned representation, as well as its performance on image generation and unpaired 3D shape inference tasks.
Tristan Aumentado-Armstrong, Alex Levinshtein, Stavros Tsogkas, Konstantinos G. Derpanis, Allan Douglas Jepson
3DV3
2020 Appearance Shock Grammar for Fast Medial Axis Extraction From Real Images
abstract
We combine ideas from shock graph theory with more recent appearance-based methods for medial axis extraction from complex natural scenes, improving upon the present best unsupervised method, in terms of efficiency and performance. We make the following specific contributions: i) we extend the shock graph representation to the domain of real images, by generalizing the shock type definitions using local, appearance-based criteria; ii) we then use the rules of a Shock Grammar to guide our search for medial points, drastically reducing run time when compared to other methods, which exhaustively consider all points in the input image; iii) we remove the need for typical post-processing steps including thinning, non-maximum suppression, and grouping, by adhering to the Shock Grammar rules while deriving the medial axis solution; iv) finally, we raise some fundamental concerns with the evaluation scheme used in previous work and propose a more appropriate alternative for assessing the performance of medial axis extraction from scenes. Our experiments on the BMAX500 and SK-LARGE datasets demonstrate the effectiveness of our approach. We outperform the present state-of-the-art, excelling particularly in the high-precision regime, while running an order of magnitude faster and requiring no post-processing.
Charles-Olivier Dufresne Camaro, Morteza Rezanejad, Stavros Tsogkas, Kaleem Siddiqi, Sven J. Dickinson
CVPR3
2020 Few-Shot Single-View 3-D Object Reconstruction with Compositional Priors
Mateusz Michalkiewicz, Sarah Parisot, Stavros Tsogkas, Mahsa Baktash, Anders P. Eriksson, Eugene Belilovsky
ECCV (25)3
2019 DeepFlux for Skeletons in the Wild
abstract
Computing object skeletons in natural images is challenging, owing to large variations in object appearance and scale, and the complexity of handling background clutter. Many recent methods frame object skeleton detection as a binary pixel classification problem, which is similar in spirit to learning-based edge detection, as well as to semantic segmentation methods. In the present article, we depart from this strategy by training a CNN to predict a two-dimensional vector field, which maps each scene point to a candidate skeleton pixel, in the spirit of flux-based skeletonization algorithms. This ``image context flux'' representation has two major advantages over previous approaches. First, it explicitly encodes the relative position of skeletal pixels to semantically meaningful entities, such as the image points in their spatial context, and hence also the implied object boundaries. Second, since the skeleton detection context is a region-based vector field, it is better able to cope with object parts of large width. We evaluate the proposed method on three benchmark datasets for skeleton detection and two for symmetry detection, achieving consistently superior performance over state-of-the-art methods.
Yukang Wang, Yongchao Xu, Stavros Tsogkas, Xiang Bai, Sven J. Dickinson, Kaleem Siddiqi
CVPR3
2019 Geometric Disentanglement for Generative Latent Shape Models
abstract
Representing 3D shapes is a fundamental problem in artificial intelligence, which has numerous applications within computer vision and graphics. One avenue that has recently begun to be explored is the use of latent representations of generative models. However, it remains an open problem to learn a generative model of shapes that is interpretable and easily manipulated, particularly in the absence of supervised labels. In this paper, we propose an unsupervised approach to partitioning the latent space of a variational autoencoder for 3D point clouds in a natural way, using only geometric information, that builds upon prior work utilizing generative adversarial models of point sets. Our method makes use of tools from spectral geometry to separate intrinsic and extrinsic shape information, and then considers several hierarchical disentanglement penalties for dividing the latent space in this manner. We also propose a novel disentanglement penalty that penalizes the predicted change in the latent representation of the output,with respect to the latent variables of the initial shape. We show that the resulting latent representation exhibits intuitive and interpretable behaviour, enabling tasks such as pose transfer that cannot easily be performed by models with an entangled representation.
Tristan Aumentado-Armstrong, Stavros Tsogkas, Allan Douglas Jepson, Sven J. Dickinson
ICCV2
2017 AMAT: Medial Axis Transform for Natural Images
abstract
We introduce Appearance-MAT (AMAT), a generalization of the medial axis transform for natural images, that is framed as a weighted geometric set cover problem. We make the following contributions: i) we extend previous medial point detection methods for color images, by associating each medial point with a local scale; ii) inspired by the invertibility property of the binary MAT, we also associate each medial point with a local encoding that allows us to invert the AMAT, reconstructing the input image; iii) we describe a clustering scheme that takes advantage of the additional scale and appearance information to group individual points into medial branches, providing a shape decomposition of the underlying image regions. In our experiments, we show state-of-the-art performance in medial point detection on Berkeley Medial AXes (BMAX500), a new dataset of medial axes based on the BSDS500 database, and good generalization on the SK506 and WH-SYMMAX datasets. We also measure the quality of reconstructed images from BMAX500, obtained by inverting their computed AMAT. Our approach delivers significantly better reconstruction quality w.r.t. to three baselines, using just 10% of the image pixels. Our code and annotations are available at https://github.com/tsogkas/amat.
Stavros Tsogkas, Sven J. Dickinson
ICCV1
2016 Prior-Based Coregistration and Cosegmentation
Mahsa Shakeri, Enzo Ferrante, Stavros Tsogkas, Sarah Lippé, Samuel Kadoury, Iasonas Kokkinos, Nikos Paragios
MICCAI (2)3
2014 Segmentation-Aware Deformable Part Models
abstract
In this work we propose a technique to combine bottom-up segmentation, coming in the form of SLIC superpixels, with sliding window detectors, such as Deformable Part Models (DPMs). The merit of our approach lies in "cleaning up" the low-level HOG features by exploiting the spatial support of SLIC superpixels, this can be understood as using segmentation to split the feature variation into object-specific and background changes. Rather than committing to a single segmentation we use a large pool of SLIC superpixels and combine them in a scale-, position- and object-dependent manner to build soft segmentation masks. The segmentation masks can be computed fast enough to repeat this process over every candidate window, during training and detection, for both the root and part filters of DPMs. We use these masks to construct enhanced, background-invariant features to train DPMs. We test our approach on the PASCAL VOC 2007, outperforming the standard DPM in 17 out of 20 classes, yielding an average increase of 1.7% AP. Additionally, we demonstrate the robustness of this approach, extending it to dense SIFT descriptors for large displacement optical flow.
Eduard Trulls, Stavros Tsogkas, Iasonas Kokkinos, Alberto Sanfeliu, Francesc Moreno-Noguer
CVPR2
2014 Understanding Objects in Detail with Fine-Grained Attributes
abstract
We study the problem of understanding objects in detail, intended as recognizing a wide array of fine-grained object attributes. To this end, we introduce a dataset of 7, 413 airplanes annotated in detail with parts and their attributes, leveraging images donated by airplane spotters and crowd-sourcing both the design and collection of the detailed annotations. We provide a number of insights that should help researchers interested in designing fine-grained datasets for other basic level categories. We show that the collected data can be used to study the relation between part detection and attribute prediction by diagnosing the performance of classifiers that pool information from different parts of an object. We note that the prediction of certain attributes can benefit substantially from accurate part detection. We also show that, differently from previous results in object detection, employing a large number of part templates can improve detection accuracy at the expenses of detection speed. We finally propose a coarse-to-fine approach to speed up detection through a hierarchical cascade algorithm.
Andrea Vedaldi, Siddharth Mahendran, Stavros Tsogkas, Subhransu Maji, Ross B. Girshick, Juho Kannala, Esa Rahtu, Iasonas Kokkinos, Matthew B. Blaschko, David J. Weiss, Ben Taskar, Karen Simonyan, Naomi Saphra, Sammy Mohamed
CVPR3
2012 Learning-Based Symmetry Detection in Natural Images
Stavros Tsogkas, Iasonas Kokkinos
ECCV (7)1