Nikolaos Gkanatsios

dblp:225/5677 · DBLP profile ↗
← Back
8ranked-venue papers
4as first author
5since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 4 since 2021Artificial intelligence and machine learning · 6 · 2 first-author · 5 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
3D vision · 22% Motion planning and robot control · 15% Segmentation and scene understanding · 13%

Topics — the 14 heaviest of 16, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › 3D vision › 3d scene understanding
3d instance segmentation
0.812024
ODIN: A Single Model for 2D and 3D Segmentation · CVPR 2024
Robotics › Motion planning and robot control › motion planning › learning-based motion planning
diffusion-based trajectory planning
0.812024
Diffusion-ES: Gradient-Free Planning with Diffusion for Autonomous and Instruction-Guided Driving · CVPR 2024
Machine learning › Generative modeling
diffusion model
0.812024
Diffusion-ES: Gradient-Free Planning with Diffusion for Autonomous and Instruction-Guided Driving · CVPR 2024
Machine learning › Optimization for machine learning › evolutionary computation
genetic algorithms
0.812024
Diffusion-ES: Gradient-Free Planning with Diffusion for Autonomous and Instruction-Guided Driving · CVPR 2024
Robotics › Autonomous driving
planning for self-driving vehicles
0.812024
Diffusion-ES: Gradient-Free Planning with Diffusion for Autonomous and Instruction-Guided Driving · CVPR 2024
Robotics › Motion planning and robot control
trajectory optimization
0.812024
Diffusion-ES: Gradient-Free Planning with Diffusion for Autonomous and Instruction-Guided Driving · CVPR 2024
Computer vision › 3D vision › 3d scene understanding
3d scene parsing
0.712023
Analogy-Forming Transformers for Few-Shot 3D Parsing · ICLR 2023
Knowledge, reasoning and agents › Knowledge representation and reasoning
analogical reasoning
0.712023
Analogy-Forming Transformers for Few-Shot 3D Parsing · ICLR 2023
Machine learning › Deep learning architectures and training
transformer
0.712023
Analogy-Forming Transformers for Few-Shot 3D Parsing · ICLR 2023
Computer vision › 3D vision
point cloud analysis
0.612022
Bottom Up Top Down Detection Transformers for Language Grounding in Images and Point Clouds · ECCV (36) 2022
Computer vision › Vision and language
visual grounding
0.612022
Bottom Up Top Down Detection Transformers for Language Grounding in Images and Point Clouds · ECCV (36) 2022
Computer vision › Segmentation and scene understanding
scene graph generation
0.512021
Grounding Consistency: Distilling Spatial Common Sense for Precise Visual Relationship Detection · ICCV 2021
Computer vision › Vision and language
visual relationship detection
0.512021
Grounding Consistency: Distilling Spatial Common Sense for Precise Visual Relationship Detection · ICCV 2021
Computer vision › 3D vision
3d scene understanding
0.212024
ODIN: A Single Model for 2D and 3D Segmentation · CVPR 2024

Methods — techniques the papers use, named apart from their topics

transformer · 1.4positional encoding · 0.8gradient-free optimization · 0.8evolutionary search · 0.8diffusion model · 0.8cross-view fusion · 0.8LLM prompting · 0.8few-shot learning · 0.7analogy · 0.7detection transformer · 0.6
YearPublicationVenuePosition
2024 ODIN: A Single Model for 2D and 3D Segmentation
abstract
State-of-the-art models on contemporary 3D segmentation benchmarks like ScanNet consume and label dataset- provided 3D point clouds, obtained through post processing of sensed multiview RGB-D images. They are typically trained in-domain, forego large-scale 2D pre-training and outperform alternatives that featurize the posed RGB- D multiview images instead. The gap in performance between methods that consume posed images versus post- processed 3D point clouds has fueled the belief that 2D and 3D perception require distinct model architectures. In this paper, we challenge this view and propose ODIN (Omni-Dimensional INstance segmentation), a model that can segment and label both 2D RGB images and 3D point clouds, using a transformer architecture that alternates between 2D within-view and 3D cross-view information fusion. Our model differentiates 2D and 3D feature oper-ations through the positional encodings of the tokens in-volved, which capture pixel coordinates for 2D patch tokens and 3D coordinates for 3D feature tokens. ODIN achieves state-of-the-art performance on ScanNet200, Matterport3D and AI2THOR 3D instance segmentation benchmarks, and competitive performance on ScanNet, S3DIS and COCO. It outperforms all previous works by a wide margin when the sensed 3D point cloud is used in place of the point cloud sampled from 3D mesh. When used as the 3D perception engine in an instructable embodied agent architecture, it sets a new state-of-the-art on the TEACh action-from-dialogue benchmark. Our code and checkpoints can be found at the project website
Pushkal Katara, Nikolaos Gkanatsios, Adam W. Harley, Gabriel Sarch, Kriti Aggarwal, Vishrav Chaudhary, Katerina Fragkiadaki
CVPR3
2024 Diffusion-ES: Gradient-Free Planning with Diffusion for Autonomous and Instruction-Guided Driving
abstract
Diffusion models excel at modeling complex and multi-modal trajectory distributions for decision-making and control. Reward-gradient guided denoising has been recently proposed to generate trajectories that maximize both a differentiable reward function and the likelihood under the data distribution captured by a diffusion model. Reward-gradient guided denoising requires a differentiable reward function fitted to both clean and noised samples, limiting its applicability as a general trajectory optimizer. In this paper, we propose Diffusion-ES, a method that combines gradient-free optimization with trajectory denoising to optimize black-box non-differentiable objectives while staying in the data manifold. Diffusion-ES samples trajectories during evolutionary search from a diffusion model and scores them using a black-box reward function. It mutates high-scoring trajecto-ries using a truncated diffusion process that applies a small number of noising and denoising steps, allowing for much more efficient exploration of the solution space. We show that Diffusion-Ex achieves state-of-the-art performance on nuPlan, an established closed-loop planning benchmark for autonomous driving. Diffusion-ES outperforms existing sampling-based planners, reactive deterministic or diffusion-based policies, and reward-gradient guidance. Additionally, we show that unlike prior guidance methods, our method can optimize non-differentiable language-shaped reward functions generated by few-shot LLM prompting. When guided by a human teacher that issues instructions to follow, our method can generate novel, highly complex behaviors, such as aggressive lane weaving, which are not present in the training data. This allows us to solve the hardest nuPlan scenarios which are beyond the capabilities of existing tra-jectory optimization methods and driving policies.11Project page: diffusion-es.github.io
Brian Yang, Huangyuan Su, Nikolaos Gkanatsios, Tsung-Wei Ke, Jeff G. Schneider, Katerina Fragkiadaki
CVPR3
2023 Analogy-Forming Transformers for Few-Shot 3D Parsing
Nikolaos Gkanatsios, Mayank Singh 0016, Zhaoyuan Fang, Shubham Tulsiani, Katerina Fragkiadaki
ICLR1
2022 Bottom Up Top Down Detection Transformers for Language Grounding in Images and Point Clouds
Nikolaos Gkanatsios, Ishita Mediratta, Katerina Fragkiadaki
ECCV (36)2
2021 Grounding Consistency: Distilling Spatial Common Sense for Precise Visual Relationship Detection
abstract
Scene Graph Generators (SGGs) are models that, given an image, build a directed graph where each edge represents a predicted subject predicate object triplet. Most SGGs silently exploit datasets' bias on relationships' context, i.e. its subject and object, to improve recall and neglect spatial and visual evidence, e.g. having seen a glut of data for person wearing shirt, they are overconfident that every person is wearing every shirt. Such imprecise predictions are mainly ascribed to the lack of negative examples for most relationships, which obstructs models from meaningfully learning predicates, even those that have ample positive examples. We first present an indepth investigation of the context bias issue to showcase that all examined state-of-the-art SGGs share the above vulnerabilities. In response, we propose a semi-supervised scheme that forces predicted triplets to be grounded consistently back to the image, in a closed-loop manner. The developed spatial common sense can be then distilled to a student SGG and substantially enhance its spatial reasoning ability. This Grounding Consistency Distillation (GCD) approach is model-agnostic and benefits from the superfluous unlabeled samples to retain the valuable context information and avert memorization of annotations. Furthermore, we demonstrate that current metrics disregard unlabeled samples, rendering themselves incapable of reflecting context bias, then we mine and incorporate during evaluation hard-negatives to reformulate precision as a reliable metric. Extensive experimental comparisons exhibit large quantitative - up to 70% relative precision boost on VG200 dataset - and qualitative improvements to prove the significance of our GCD method and our metrics towards refocusing graph generation as a core aspect of scene understanding. Code available at https://github.com/deeplab-ai/grounding-consistent-vrd.
Markos Diomataris, Nikolaos Gkanatsios, Vassilis Pitsikalis, Petros Maragos
ICCV2
2020 From Saturation to Zero-Shot Visual Relationship Detection Using Local Context
Nikolaos Gkanatsios, Vassilis Pitsikalis, Petros Maragos
BMVC1
2019 Deeply Supervised Multimodal Attentional Translation Embeddings for Visual Relationship Detection
abstract
Detecting visual relationships, i.e.triplets, has been a challenging Scene Understanding task approached in the past via linguistic priors or spatial information in a single feature branch. We introduce a new deeply supervised two-branch architecture, the Multimodal Attentional Translation Embeddings, where the visual features of each branch are driven by a multimodal attentional mechanism that exploits spatio-linguistic similarities in a low-dimensional space. We present a variety of experiments comparing against all related approaches in the literature, as well as by re-implementing and fine-tuning several of them. Results on the commonly employed VRD dataset [1] show that the proposed method clearly outperforms all others, while we also justify our claims both quantitatively and qualitatively.
Nikolaos Gkanatsios, Vassilis Pitsikalis, Petros Koutras, Athanasia Zlatintsi, Petros Maragos
ICIP1
2019 Showcasing Deeply Supervised Multimodal Attentional Translation Embeddings: a Demo for Visual Relationship Detection
abstract
We address the task of Visual Relationship Detection, i.e. the detection oftriplets in an image, introducing Multimodal Attentional Translation Embeddings (ICIP 2019 paper, id 3642). Motivated by the need of visualization and interpretation of the results, as well as the lack of other tools for online predictions on this task, we design and implement the first architecture for live inference of visual relationships on video streams and wild images, including research and engineering extensions, ablation models and a lightweight CPU-version. The code is available at https://bitbucket.org/deeplabai/vrd.
Nikolaos Gkanatsios, Vassilis Pitsikalis, Petros Koutras, Athanasia Zlatintsi, Petros Maragos
ICIP1