Zoltán Ádám Milacski

dblp:161/2866 · also Zoltan A. Milacski · DBLP profile ↗
← Back
12ranked-venue papers
5as first author
7since 2021 · last 2025
0000-0002-3135-2936ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 4 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Custom Condition Generation for Zero-Shot Human-Scene Interactions Synthesis
abstract
Existing methods for creating human interactions within scenes show promise for common interactions, but often fail with less frequent ones. To overcome this, we introduce a new approach that creates tailored conditions for generating these interactions without previously seen examples. This method leverages the strengths of both large language models (LLMs) and vision-language models (VLMs). Unlike the GenZI, the current state-of-the-art approach, which struggles with rare interactions due to its reliance on VLM inpainting, our method follows a three-step process: first, we generate a preliminary human posture using VLMs and then estimate this posture in three dimensions. Next, we refine the conditions to fit the specific scene and interaction by analyzing the inputs with both LLMs and VLMs. Finally, we fine-tune the placement, orientation, and posture of the human figure using specific optimization techniques. Our experimental results show that this method performs well across a wide range of interactions, including those that are less common.
Ryosuke Kawamura, Zoltán Ádám Milacski, Fernando De la Torre, László A. Jeni, Koichiro Niinuma
FG2
2025 OASIS: Object-guided Attention for Text-conditional Diffusion Synthesis of Human Interaction Sequences
abstract
Analyzing and synthesizing human-object interaction is crucial for advancing intelligent systems that engage with the physical environment. However, simultaneous tracking of human and object data presents inherent challenges, resulting in limitations in dataset scale, diversity, and annotation quality within this domain, thereby hindering the generalization ability of trained models. This study introduces OASIS, a novel framework that extends pretrained text-conditional human motion diffusion models to address the complex task of fullbody 3D hand-object interaction generation. Specifically, we freeze the parameters of the pretrained motion diffusion model, while incorporating additional object-guided attention layers, which we train to adapt the human motion latents to match the input object motion sequence and the text. Our method can be understood as a ControlNet [38] for interaction. Through extensive experimentation, we demonstrate the effectiveness and robustness of our framework in generating realistic handobject interactions from textual descriptions. Our method surpasses the state-of-the-art performance in FID and accuracy interaction fidelity metrics compared to the prior best method IMoS [10], with improvements of 0.08 in FID and $2 \%$ in accuracy for body motion synthesis, and 0.15 in FID and $10 \%$ in accuracy for hand motion synthesis.
Chih-Chun Yang 0005, Tianhui Cai, Zoltán Ádám Milacski, Aayush Prakash, Shingo Takagi 0001, Daeil Kim, Fernando De la Torre
FG3
2025 GHOST: Grounded Human Motion Generation with Open Vocabulary Scene-and-Text Contexts
abstract
The connection between our 3D surroundings and the descriptive language that characterizes them would be well-suited for localizing and generating human motion in context but for one problem. The complexity introduced by multiple modalities makes capturing this connection challenging with a fixed set of descriptors. Specifically, closed vocabulary scene encoders, which require learning text-scene associations from scratch, have been favored in the literature, often resulting in inaccurate motion grounding. In this paper, we propose a method that integrates an open vocabulary scene encoder into the architecture, establishing a robust connection between text and scene. Our two-step approach starts with pretraining the scene encoder through knowledge distillation from an existing open vocabulary semantic image segmentation model, ensuring a shared text-scene feature space. Subsequently, the scene encoder is fine-tuned for conditional motion generation, incorporating two novel regularization losses that regress the category and size of the goal object. Our methodology achieves up to a 30% reduction in the goal object distance metric compared to the prior state-of-the-art baseline model on the HUMANISE dataset. This improvement is demonstrated through evaluations conducted using three implementations of our framework, a perceptual study, and an open vocabulary experiment. Additionally, our method is designed to accommodate future 2D open vocabulary segmentation methods for distillation in a plug-and-play manner.
Zoltán Ádám Milacski, Koichiro Niinuma, Ryosuke Kawamura, Fernando De la Torre, László A. Jeni
WACV1
2024 CoGS: Controllable Gaussian Splatting
abstract
Capturing and re-animating the 3D structure of artic-ulated objects present significant barriers. On one hand, methods requiring extensively calibrated multi-view setups are prohibitively complex and resource-intensive, limiting their practical applicability. On the other hand, while single-camera Neural Radiance Fields (NeRFs) offer a more streamlined approach, they have excessive training and rendering costs. 3D Gaussian Splatting would be a suitable alternative but for two reasons. Firstly, existing methods for 3D dynamic Gaussians require synchronized multi- view cameras, and secondly, the lack of controllability in dynamic scenarios. We present CoGS, a methodfor Controllable Gaussian Splatting, that enables the direct ma-nipulation of scene elements, offering real-time control of dynamic scenes without the prerequisite of pre-computing control signals. We evaluated CoGS using both synthetic and real-world datasets that include dynamic objects that differ in degree of difficulty. In our evaluations, CoGS con-sistently outperformed existing dynamic and controllable neural representations in terms of visual fidelity.
Joel Julin, Zoltán Ádám Milacski, Koichiro Niinuma, László A. Jeni
CVPR3
2024 MotionGPT: Human Motion Synthesis with Improved Diversity and Realism via GPT-3 Prompting
abstract
There are numerous applications for human motion synthesis, including animation, gaming, robotics, or sports science. In recent years, human motion generation from natural language has emerged as a promising alternative to costly and labor-intensive data collection methods relying on motion capture or wearable sensors (e.g., suits). Despite this, generating human motion from textual descriptions remains a challenging and intricate task, primarily due to the scarcity of large-scale supervised datasets capable of capturing the full diversity of human activity.This study proposes a new approach, called MotionGPT, to address the limitations of previous text-based human motion generation methods by utilizing the extensive semantic information available in large language models (LLMs). We first pretrain a doubly text-conditional motion diffusion model on both coarse ("high-level") and detailed ("low-level") ground truth text data. Then during inference, we improve motion diversity and alignment with the training set, by zero-shot prompting GPT-3 for additional "low-level" details. Our method achieves new state-of-the-art quantitative results in terms of Fréchet Inception Distance (FID) and motion diversity metrics, and improves all considered metrics. Furthermore, it has strong qualitative performance, producing natural results. Code is available at https://github.com/humansensinglab/MotionGPT
José Ribeiro-Gomes, Tianhui Cai, Zoltán Ádám Milacski, Aayush Prakash, Shingo Takagi 0001, Amaury Aubel, Daeil Kim, Alexandre Bernardino, Fernando De la Torre
WACV3
2023 DyLiN: Making Light Field Networks Dynamic
abstract
Light Field Networks, the re-formulations of radiance fields to oriented rays, are magnitudes faster than their coordinate network counterparts, and provide higher fidelity with respect to representing 3D structures from 2D observations. They would be well suited for generic scene representation and manipulation, but suffer from one problem: they are limited to holistic and static scenes. In this paper, we propose the Dynamic Light Field Network (DyLiN) method that can handle non-rigid deformations, including topological changes. We learn a deformation field from input rays to canonical rays, and lift them into a higher dimensional space to handle discontinuities. We further introduce CoDyLiN, which augments DyLiN with controllable attribute inputs. We train both models via knowledge distillation from pretrained dynamic radiance fields. We evaluated DyLiN using both synthetic and real world datasets that include various non-rigid deformations. DyLiN qualitatively outperformed and quantitatively matched state-of-the-art methods in terms of visual fidelity, while being 25 – 71× computationally faster. We also tested CoDyLiN on attribute annotated data and it surpassed its teacher model. Project page: https://dylin2023.github.io.
Joel Julin, Zoltán Ádám Milacski, Koichiro Niinuma, László A. Jeni
CVPR3
2021 MADGAN: unsupervised medical anomaly detection GAN using multiple adjacent brain MRI slice reconstruction
abstract
BACKGROUND: Unsupervised learning can discover various unseen abnormalities, relying on large-scale unannotated medical images of healthy subjects. Towards this, unsupervised methods reconstruct a 2D/3D single medical image to detect outliers either in the learned feature space or from high reconstruction loss. However, without considering continuity between multiple adjacent slices, they cannot directly discriminate diseases composed of the accumulation of subtle anatomical anomalies, such as Alzheimer's disease (AD). Moreover, no study has shown how unsupervised anomaly detection is associated with either disease stages, various (i.e., more than two types of) diseases, or multi-sequence magnetic resonance imaging (MRI) scans. RESULTS: We propose unsupervised medical anomaly detection generative adversarial network (MADGAN), a novel two-step method using GAN-based multiple adjacent brain MRI slice reconstruction to detect brain anomalies at different stages on multi-sequence structural MRI: (Reconstruction) Wasserstein loss with Gradient Penalty + 100 [Formula: see text] loss-trained on 3 healthy brain axial MRI slices to reconstruct the next 3 ones-reconstructs unseen healthy/abnormal scans; (Diagnosis) Average [Formula: see text] loss per scan discriminates them, comparing the ground truth/reconstructed slices. For training, we use two different datasets composed of 1133 healthy T1-weighted (T1) and 135 healthy contrast-enhanced T1 (T1c) brain MRI scans for detecting AD and brain metastases/various diseases, respectively. Our self-attention MADGAN can detect AD on T1 scans at a very early stage, mild cognitive impairment (MCI), with area under the curve (AUC) 0.727, and AD at a late stage with AUC 0.894, while detecting brain metastases on T1c scans with AUC 0.921. CONCLUSIONS: Similar to physicians' way of performing a diagnosis, using massive healthy training data, our first multiple MRI slice reconstruction approach, MADGAN, can reliably predict the next 3 slices from the previous 3 ones only for unseen healthy images. As the first unsupervised various disease diagnosis, MADGAN can reliably detect the accumulation of subtle anatomical anomalies and hyper-intense enhancing lesions, such as (especially late-stage) AD and brain metastases on multi-sequence MRI scans.
Leonardo Rundo, Kohei Murao, Tomoyuki Noguchi, Yuki Shimahara, Zoltán Ádám Milacski, Saori Koshino, Evis Sala, Hideki Nakayama, Shin'ichi Satoh 0001
BMC Bioinform.6
2020 VideoOneNet: Bidirectional Convolutional Recurrent OneNet with Trainable Data Steps for Video Processing
abstract
Deep Neural Networks (DNNs) achieve the state-of-the-art results on a wide range of image processing tasks, however, the majority of such solutions are problem-specific, like most AI algorithms. The One Network to Solve Them All (OneNet) procedure has been suggested to resolve this issue by exploiting a DNN as the proximal operator in Alternating Direction Method of Multipliers (ADMM) solvers for various imaging problems. In this work, we make two contributions, both facilitating end-to-end learning using backpropagation. First, we generalize OneNet to videos by augmenting its convolutional prior network with bidirectional recurrent connections; second, we extend the fixed fully connected linear ADMM data step with another trainable bidirectional convolutional recurrent network. In our computational experiments on the Rotated MNIST, Scanned CIFAR-10 and UCF-101 data sets, the proposed modifications improve performance by a large margin compared to end-to-end convolutional OneNet and 3D Wavelet sparsity on several video processing problems: pixelwise inpainting-denoising, blockwise inpainting, scattered inpainting, super resolution, compressive sensing, deblurring, frame interpolation, frame prediction and colorization. Our two contributions are complementary, and using them together yields the best results.
Zoltán Ádám Milacski, Barnabás Póczos, András Lörincz
ICML1
2020 Multi Object Tracking for Similar Instances: A Hybrid Architecture
Áron Fóthi, Kinga Bettina Faragó, László Kopácsi, Zoltán Ádám Milacski, Viktor Varga, András Lörincz
ICONIP (1)4
2019 Differentiable Unrolled Alternating Direction Method of Multipliers for OneNet
Zoltán Ádám Milacski, Barnabás Póczos, András Lörincz
BMVC1
2019 Group k-Sparse Temporal Convolutional Neural Networks: Unsupervised Pretraining for Video Classification
abstract
In this paper we propose Group k-Sparse Temporal Convolutional Neural Networks for unsupervised pretraining using video data. Our work is the first to consider the recurrent extension of structured sparsity, thus enhancing representational power and explainability. We show that our architecture is able to outperform several state-of-the-art baselines on Rotated MNIST, Scanned CIFAR-10, COIL-100 and NEC Animal pretraining benchmarks for video classification using limited labeled data.
Zoltán Ádám Milacski, Barnabás Póczos, András Lörincz
IJCNN1
2015 Robust Detection of Anomalies via Sparse Methods
Zoltán Ádám Milacski, Marvin Ludersdorfer, András Lörincz, Patrick van der Smagt
ICONIP (3)1