Andrew Melnik

dblp:218/5571 · DBLP profile ↗
← Back
6ranked-venue papers
1as first author
5since 2021 · last 2026
0000-0002-7252-9267ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Robot manipulation · 60% Generative modeling · 35% 3D vision · 6%
Computer graphics and multimedia
1 paper
Visual content generation and editing · 100%

Topics — the 8 heaviest of 9, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Robotics › Robot manipulation
grasping
1.012026
GRIM: Task-Oriented Grasping with Conditioning on Generative Examples · AAAI 2026
Robotics › Robot manipulation › grasping › grasp planning
task-oriented grasping
1.012026
GRIM: Task-Oriented Grasping with Conditioning on Generative Examples · AAAI 2026
Machine learning › Generative modeling › face synthesis
GAN-based face generation
0.812024
Face Generation and Editing With StyleGAN: A Survey · IEEE Trans. Pattern Anal. Mach. Intell. 2024
Machine learning › Generative modeling › generative adversarial network
StyleGAN
0.812024
Face Generation and Editing With StyleGAN: A Survey · IEEE Trans. Pattern Anal. Mach. Intell. 2024
Visual content generation and editing
face editing
0.812024
Face Generation and Editing With StyleGAN: A Survey · IEEE Trans. Pattern Anal. Mach. Intell. 2024
Visual content generation and editing › image editing
GAN inversion
0.812024
Face Generation and Editing With StyleGAN: A Survey · IEEE Trans. Pattern Anal. Mach. Intell. 2024
Computer vision › 3D vision
object alignment
0.312026
GRIM: Task-Oriented Grasping with Conditioning on Generative Examples · AAAI 2026
Machine learning › Generative modeling › image manipulation
deepfake
0.212024
Face Generation and Editing With StyleGAN: A Survey · IEEE Trans. Pattern Anal. Mach. Intell. 2024

Methods — techniques the papers use, named apart from their topics

pg-GAN · 1.5StyleGAN3 · 1.5StyleGAN · 1.5principal component analysis · 1.0geometric cues · 1.0DINO features · 1.0
YearPublicationVenuePosition
2026 GRIM: Task-Oriented Grasping with Conditioning on Generative Examples
abstract
Task-Oriented Grasping (TOG) presents a significant challenge, requiring a nuanced understanding of task semantics, object affordances, and the functional constraints dictating how an object should be grasped for a specific task. To address these challenges, we introduce GRIM (Grasp Re-alignment via Iterative Matching), a novel training-free framework for task-oriented grasping. Initially, a coarse alignment strategy is developed using a combination of geometric cues and principal component analysis (PCA)-reduced DINO features for similarity scoring. Subsequently, the full grasp pose associated with the retrieved memory instance is transferred to the aligned scene object and further refined against a set of task-agnostic, geometrically stable grasps generated for the scene object, prioritizing task compatibility. In contrast to existing learning-based methods, GRIM demonstrates strong generalization capabilities, achieving robust performance with only a small number of conditioning examples.
Shailesh, Nayan Kumar, Priya Shukla, Andrew Melnik, Michael Beetz, Gora Chand Nandi
AAAI5
2024 Zero-Shot Imitation Policy Via Search In Demonstration Dataset
abstract
Behavioral cloning uses a dataset of demonstrations to learn a policy. To overcome computationally expensive training procedures and address the policy adaptation problem, we propose to use latent spaces of pre-trained foundation models to index a demonstration dataset, instantly access similar relevant experiences, and copy behavior from these situations. Actions from a selected similar situation can be performed by the agent until representations of the agent’s current situation and the selected experience diverge in the latent space. Thus, we formulate our control problem as a dynamic search problem over a dataset of experts’ demonstrations. We test our approach on BASALT MineRL-dataset in the latent representation of a Video Pre-Training model. We compare our model to state-of-the-art, Imitation Learning-based Minecraft agents. Our approach can effectively recover meaningful demonstrations and show human-like behavior of an agent in the Minecraft environment in a wide variety of scenarios. Experimental results reveal that performance of our search-based approach clearly wins in terms of accuracy and perceptual evaluation over learning-based models.
Federico Malato, Florian Leopold, Andrew Melnik, Ville Hautamäki
ICASSP3
2024 Face Generation and Editing With StyleGAN: A Survey
abstract
Our goal with this survey is to provide an overview of the state of the art deep learning methods for face generation and editing using StyleGAN. The survey covers the evolution of StyleGAN, from PGGAN to StyleGAN3, and explores relevant topics such as suitable metrics for training, different latent representations, GAN inversion to latent spaces of StyleGAN, face image editing, cross-domain face stylization, face restoration, and even Deepfake applications. We aim to provide an entry point into the field for readers that have basic knowledge about the field of deep learning and are looking for an accessible introduction and overview.
Andrew Melnik, Maksim Miasayedzenkau, Dzianis Makarovets, Dzianis Pirshtuk, Eren Akbulut, Dennis Holzmann, Tarek Renusch, Gustav Reichert, Helge J. Ritter
IEEE Trans. Pattern Anal. Mach. Intell.1
2022 A Graph-based U-Net Model for Predicting Traffic in unseen Cities
abstract
Accurate traffic prediction is a key ingredient to enable traffic management like rerouting cars to reduce road congestion or regulating traffic via dynamic speed limits to maintain a steady flow. One way to represent traffic data is as temporally changing heatmaps visualizing attributes of traffic, such as speed and volume. In recent approaches, U-Net models have shown state of the art performance on traffic forecasting from such heatmaps. We propose to combine the U-Net architecture with graph layers which improves spatial generalization to unseen road networks compared to a Vanilla U-Net. In particular, we specialize existing graph operations to be sensitive to geographical topology and generalize pooling and upsampling operations to be applicable to graphs.
Luca Hermes, Barbara Hammer, Andrew Melnik, Riza Velioglu, Markus Vieth 0001, Malte Schilling
IJCNN3
2021 Decentralized control and local information for robust and adaptive decentralized Deep Reinforcement Learning
abstract
Decentralization is a central characteristic of biological motor control that allows for fast responses relying on local sensory information. In contrast, the current trend of Deep Reinforcement Learning (DRL) based approaches to motor control follows a centralized paradigm using a single, holistic controller that has to untangle the whole input information space. This motivates to ask whether decentralization as seen in biological control architectures might also be beneficial for embodied sensori-motor control systems when using DRL. To answer this question, we provide an analysis and comparison of eight control architectures for adaptive locomotion that were derived for a four-legged agent, but with their degree of decentralization varying systematically between the extremes of fully centralized and fully decentralized. Our comparison shows that learning speed is significantly enhanced in distributed architectures-while still reaching the same high performance level of centralized architectures-due to smaller search spaces and local costs providing more focused information for learning. Second, we find an increased robustness of the learning process in the decentralized cases-it is less demanding to hyperparameter selection and less prone to becoming trapped in poor local minima. Finally, when examining generalization to uneven terrains-not used during training-we find best performance for an intermediate architecture that is decentralized, but integrates only local information from both neighboring legs. Together, these findings demonstrate beneficial effects of distributing control into decentralized units and relying on local information. This appears as a promising approach towards more robust DRL and better generalization towards adaptive behavior.
Malte Schilling, Andrew Melnik, Frank W. Ohl, Helge J. Ritter, Barbara Hammer
Neural Networks2
2019 Jointly Trained Variational Autoencoder for Multi-Modal Sensor Fusion
Timo Korthals, Marc Hesse, Jürgen Leitner, Andrew Melnik, Ulrich Rückert 0001
FUSION4