Julian Tanke

dblp:217/2077 · DBLP profile ↗
← Back
9ranked-venue papers
3as first author
7since 2021 · last 2025
0009-0004-7670-6808ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 2 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2025 TalkCuts: A Large-Scale Dataset for Multi-Shot Human Speech Video Generation
abstract
In this work, we present TalkCuts, a large-scale dataset designed to facilitate the study of multi-shot human speech video generation. Unlike existing datasets that focus on single-shot, static viewpoints, TalkCuts offers 164k clips totaling over 500 hours of high-quality 1080P human speech videos with diverse camera shots, including close-up, half-body, and full-body views. The dataset includes detailed textual descriptions, 2D keypoints and 3D SMPL-X motion annotations, covering over 10k identities, enabling multimodal learning and evaluation. As a first attempt to showcase the value of the dataset, we present Orator, an LLM-guided multi-modal generation framework as a simple baseline, where the language model functions as a multi-faceted director, orchestrating detailed specifications for camera transitions, speaker gesticulations, and vocal modulation. This architecture enables the synthesis of coherent long-form videos through our integrated multi-modal video generation module. Extensive experiments in both pose-guided and audio-driven settings show that training on TalkCuts significantly enhances the cinematographic coherence and visual appeal of generated multi-shot speech videos. We believe TalkCuts provides a strong foundation for future work in controllable, multi-shot speech video generation and broader multimodal learning.
Jiaben Chen, Ailing Zeng, Xueyang Yu, Siyuan Cen, Julian Tanke, Koichi Saito, Yuki Mitsufuji, Chuang Gan 0001
NeurIPS7
2023 Social Diffusion: Long-term Multiple Human Motion Anticipation
abstract
We propose Social Diffusion, a novel method for short-term and long-term forecasting of the motion of multiple persons as well as their social interactions. Jointly forecasting motions for multiple persons involved in social activities is inherently a challenging problem due to the interdependencies between individuals. In this work, we leverage a diffusion model conditioned on motion histories and causal temporal convolutional networks to forecast individually and contextually plausible motions for all participants. The contextual plausibility is achieved via an order-invariant aggregation function. As a second contribution, we design a new evaluation protocol that measures the plausibility of social interactions which we evaluate on the Haggling dataset, which features a challenging social activity where people are actively taking turns to talk and switching their attention. We evaluate our approach on four datasets for multi-person forecasting where our approach outperforms the state-of-the-art in terms of motion realism and contextual plausibility.
Julian Tanke, Linguang Zhang, Amy Zhao, Chengcheng Tang, Yujun Cai, Lezi Wang, Po-Chen Wu, Juergen Gall, Cem Keskin
ICCV1
2023 Humans in Kitchens: A Dataset for Multi-Person Human Motion Forecasting with Scene Context
abstract
Forecasting human motion of multiple persons is very challenging. It requires to model the interactions between humans and the interactions with objects and the environment. For example, a person might want to make a coffee, but if the coffee machine is already occupied the person will haveto wait. These complex relations between scene geometry and persons ariseconstantly in our daily lives, and models that wish to accurately forecasthuman behavior will have to take them into consideration. To facilitate research in this direction, we propose Humans in Kitchens, alarge-scale multi-person human motion dataset with annotated 3D human poses, scene geometry and activities per person and frame.Our dataset consists of over 7.3h recorded data of up to 16 persons at the same time in four kitchen scenes, with more than 4M annotated human poses, represented by a parametric 3D body model. In addition, dynamic scene geometry and objects like chair or cupboard are annotated per frame. As first benchmarks, we propose two protocols for short-term and long-term human motion forecasting.
Julian Tanke, Oh-Hun Kwon, Felix B. Mueller, Andreas Doering, Juergen Gall
NeurIPS1
2023 ElliPose: Stereoscopic 3D Human Pose Estimation by Fitting Ellipsoids
abstract
One of the most relevant tasks for augmented and virtual reality applications is the interaction of virtual objects with real humans which requires accurate 3D human pose predictions. Obtaining accurate 3D human poses requires careful camera calibration which is difficult for nontechnical personal or in a pop-up scenario. Recent markerless motion capture approaches require accurate camera calibration at least for the final triangulation step. Instead, we solve this problem by presenting ElliPose, Stereoscopic 3D Human Pose Estimation by Fitting Ellipsoids, where we jointly estimate the 3D human as well as the camera pose. We exploit the fact that bones do not change in length over the course of a sequence and thus their relative trajectories have to lie on the surface of a sphere which we can utilize to iteratively correct the camera and 3D pose estimation. As another use-case we demonstrate that our approach can be used as replacement for ground-truth 3D poses to train monocular 3D pose estimators. We show that our method produces competitive results even when comparing with state-of-the-art methods that use more cameras or ground-truth camera extrinsics.
Christian Grund, Julian Tanke, Juergen Gall
WACV2
2022 TAVA: Template-free Animatable Volumetric Actors
Ruilong Li, Julian Tanke, Minh Vo, Michael Zollhöfer, Juergen Gall, Angjoo Kanazawa, Christoph Lassner
ECCV (32)2
2021 Intention-based Long-Term Human Motion Anticipation
abstract
Recently, a few works have been proposed to model the uncertainty of the future human motion. These works do not forecast a single sequence but multiple sequences for the same observation. While these works focused on increasing the diversity, this work focuses on keeping a high quality of the forecast sequences even for very long time horizons of up to 30 seconds. In order to achieve this goal, we propose to forecast the intention of the person ahead of time. This has the advantage that the generated human motion remains goal oriented and that the motion transitions between two actions are smooth and highly realistic. We furthermore propose a new quality score for evaluation that correlates better with human perception than other metrics. The results and a user study show that our approach forecasts multiple sequences that are more plausible compared to the state-of-the-art.
Julian Tanke, Chintan Zaveri, Juergen Gall
3DV1
2021 ANR: Articulated Neural Rendering for Virtual Avatars
abstract
The combination of traditional rendering with neural networks in Deferred Neural Rendering (DNR) [38] provides a compelling balance between computational complexity and realism of the resulting images. Using skinned meshes for rendering articulating objects is a natural extension for the DNR framework and would open it up to a plethora of applications. However, in this case the neural shading step must account for deformations that are possibly not captured in the mesh, as well as alignment in-accuracies and dynamics—which can confound the DNR pipeline. We present Articulated Neural Rendering (ANR), a novel framework based on DNR which explicitly addresses its limitations for virtual human avatars. We show the superiority of ANR not only with respect to DNR but also with methods specialized for avatar creation and animation. In two user studies, we observe a clear preference for our avatar model and we demonstrate state-of-the-art performance on quantitative evaluation metrics. Perceptually, we observe better temporal stability, level of detail and plausibility. More results are available at our project page: https://anr-avatars.github.io.
Amit Raj, Julian Tanke, James Hays, Minh Vo, Carsten Stoll, Christoph Lassner
CVPR2
2020 Recursive Bayesian Filtering for Multiple Human Pose Tracking from Multiple Cameras
Oh-Hun Kwon, Julian Tanke, Juergen Gall
ACCV (2)2
2018 An adaptive peer-sampling protocol for building networks of browsers
Brice Nédelec, Julian Tanke, Davide Frey, Pascal Molli, Achour Mostéfaoui
World Wide Web2