Thorsten Gernoth

dblp:23/6414 · DBLP profile ↗
← Back
4ranked-venue papers
1as first author
2since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer graphics and multimedia
1 paper
Visual content generation and editing · 100%
Artificial intelligence
2 papers
3D vision · 85% Generative modeling · 15%

Topics — the 7 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Visual content generation and editing › video generation › controllable video generation
camera-controlled video generation
0.912025
Cavia: Camera-controllable Multi-view Video Diffusion with View-Integrated Attention · ICML 2025
Visual content generation and editing › video generation
multi-view video generation
0.912025
Cavia: Camera-controllable Multi-view Video Diffusion with View-Integrated Attention · ICML 2025
Visual content generation and editing
video generation
0.912025
Cavia: Camera-controllable Multi-view Video Diffusion with View-Integrated Attention · ICML 2025
Computer vision › 3D vision › point cloud registration
partial point cloud registration
0.512021
DeepPRO: Deep Partial Point Cloud Registration of Objects · ICCV 2021
Computer vision › 3D vision
point cloud registration
0.512021
DeepPRO: Deep Partial Point Cloud Registration of Objects · ICCV 2021
Computer vision › 3D vision › pose estimation
rigid body pose estimation
0.512021
DeepPRO: Deep Partial Point Cloud Registration of Objects · ICCV 2021
Machine learning › Generative modeling
diffusion model
0.312025
Cavia: Camera-controllable Multi-view Video Diffusion with View-Integrated Attention · ICML 2025

Methods — techniques the papers use, named apart from their topics

view-integrated attention · 1.7diffusion model · 1.7dense correspondence · 0.5deep neural network · 0.5
YearPublicationVenuePosition
2025 Cavia: Camera-controllable Multi-view Video Diffusion with View-Integrated Attention
abstract
In recent years there have been remarkable breakthroughs in image-to-video generation. However, the 3D consistency and camera controllability of generated frames have remained unsolved. Recent studies have attempted to incorporate camera control into the generation process, but their results are often limited to simple trajectories or lack the ability to generate consistent videos from multiple distinct camera paths for the same scene. To address these limitations, we introduce Cavia, a novel framework for camera-controllable, multi-view video generation, capable of converting an input image into multiple spatiotemporally consistent videos. Our framework extends the spatial and temporal attention modules into view-integrated attention modules, improving both viewpoint and temporal consistency. This flexible design allows for joint training with diverse curated data sources, including scene-level static videos, object-level synthetic multi-view dynamic videos, and real-world monocular dynamic videos. To the best of our knowledge, Cavia is the first framework that enables users to generate multiple videos of the same scene with precise control over camera motion, while simultaneously preserving object motion. Extensive experiments demonstrate that Cavia surpasses state-of-the-art methods in terms of geometric consistency and perceptual quality.
Dejia Xu, Yifan Jiang 0001, Liangchen Song, Thorsten Gernoth, Liangliang Cao, Zhangyang Wang, Hao Tang 0001
ICML5
2021 DeepPRO: Deep Partial Point Cloud Registration of Objects
abstract
We consider the problem of online and real-time registration of partial point clouds obtained from an unseen real-world rigid object without knowing its 3D model. The point cloud is partial as it is obtained by a depth sensor capturing only the visible part of the object from a certain viewpoint. It introduces two main challenges: 1) two partial point clouds do not fully overlap and 2) keypoints tend to be less reliable when the visible part of the object does not have salient local structures. To address these issues, we propose DeepPRO, a keypoint-free and an end-to-end trainable deep neural network. Its core idea is inspired by how humans align two point clouds: we can imagine how two point clouds will look like after the registration based on their shape. To realize the idea, DeepPRO has inputs of two partial point clouds and directly predicts the point-wise location of the aligned point cloud. By preserving the ordering of points during the prediction, we enjoy dense correspondences between input and predicted point clouds when inferring rigid transform parameters. We conduct extensive experiments on the real-world Linemod and synthetic ModelNet40 datasets. In addition, we collect and evaluate on the PRO1k dataset, a large-scale version of Linemod meant to test generalization to real-world scans. Results show that DeepPRO achieves the best accuracy against thirteen strong baseline methods, e.g., 2.2mm ADD on the Linemod dataset, while running 50 fps on mobile devices.
Onur C. Hamsici, Steven Feng, Prachee Sharma, Thorsten Gernoth
ICCV5
2010 Face recognition under pose variations using shape-adapted texture features
abstract
The complexity in face recognition emerges from the variability of the appearance of a human face. While the identity is preserved, the appearance of a face may change due to factors such as illumination, pose or facial expression. To recognize a person independent of pose, we want to separate shape from texture information. We concentrate on the texture part in this work. We first fit an active appearance model to a given facial image. The shape information is used to transform the face into a shape-free representation. We decompose the transformed face into local regions and extract texture features from these not necessarily rectangular regions using a shape-adapted discrete cosine transform. The texture features we use for face recognition are independent of pose and shape of the face. We show that these features contain sufficient discriminative information to recognize persons across changes in pose.
Thorsten Gernoth, André Gooßen, Rolf-Rainer Grigat
ICIP1
2008 Local binary patterns for lip motion analysis
abstract
Lip motion analysis can enhance a security system based on face recognition significantly. Examining speaker dependent lip movements and visually determining the spoken content can prevent attacks using photographs or pre-recorded video sequences. Our system operates under active near infrared illumination. We investigate the use of a new type of features, namely local binary patterns, to model lip motions with hidden Markov models. We evaluate the classification accuracy with the TUNIR database, which we made available to the public for the future comparison of competing approaches.
Ralph Kricke, Thorsten Gernoth, Rolf-Rainer Grigat
ICIP2