Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Daniel Koceja

dblp:405/1358 · DBLP profile ↗
← Back
1ranked-venue papers
0as first author
1since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Generative modeling · 70% Deep learning architectures and training · 30%

Topics — the 4 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Generative modeling › video generation
long video generation
0.912025
One-Minute Video Generation with Test-Time Training · CVPR 2025
Machine learning › Deep learning architectures and training
sequence modeling
0.912025
One-Minute Video Generation with Test-Time Training · CVPR 2025
Machine learning › Generative modeling
video generation
0.912025
One-Minute Video Generation with Test-Time Training · CVPR 2025
Machine learning › Generative modeling › video generation
text-to-video generation
0.312025
One-Minute Video Generation with Test-Time Training · CVPR 2025

Methods — techniques the papers use, named apart from their topics

transformer · 0.9test-time training · 0.9mamba layers · 0.9
YearPublicationVenuePosition
2025 One-Minute Video Generation with Test-Time Training
abstract
Transformers today still struggle to generate one-minute videos because self-attention layers are inefficient for long context. Alternatives such as Mamba layers struggle to produce coherent scenes because their hidden states are small and less expressive. We experiment with Test-Time Training (TTT) layers, whose hidden states themselves can be neural networks, therefore larger and more expressive. Adding TTT layers into a pre-trained Transformer enables it to generate one-minute videos from text storyboards. We curate a dataset based on Tom and Jerry cartoons as a proof-of-concept benchmark. Compared to baselines such as Mamba 2, Gated DeltaNet, and sliding-window attention layers, TTT layers generate much more coherent videos that tell complete stories, leading by 34 Elo points in a human evaluation of 100 videos per method. Although promising, our results are still limited in physical realism, and the efficiency of our implementation can be further improved.Sample videos, code and annotations are available at: https://test-time-training.github.io/video-dit
Karan Dalal, Daniel Koceja, Yue Zhao 0006, Shihao Han, Ka Chun Cheung, Jan Kautz, Yejin Choi 0001, Yu Sun 0020, Xiaolong Wang 0004
CVPR2