VLDB 2026 Research / reviewers in the wild / expert
Daniel Koceja
dblp:405/1358
· DBLP profile ↗
1ranked-venue papers
0as first author
1since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
1 paper |
Generative modeling · 70% Deep learning architectures and training · 30% |
Topics — the 4 heaviest of 4, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Generative modeling › video generation
long video generation |
0.9 | 1 | 2025 | One-Minute Video Generation with Test-Time Training · CVPR 2025 |
Machine learning › Deep learning architectures and training
sequence modeling |
0.9 | 1 | 2025 | One-Minute Video Generation with Test-Time Training · CVPR 2025 |
Machine learning › Generative modeling
video generation |
0.9 | 1 | 2025 | One-Minute Video Generation with Test-Time Training · CVPR 2025 |
Machine learning › Generative modeling › video generation
text-to-video generation |
0.3 | 1 | 2025 | One-Minute Video Generation with Test-Time Training · CVPR 2025 |
Methods — techniques the papers use, named apart from their topics
transformer · 0.9test-time training · 0.9mamba layers · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | One-Minute Video Generation with Test-Time TrainingabstractTransformers today still struggle to generate one-minute videos because self-attention layers are inefficient for long context. Alternatives such as Mamba layers struggle to produce coherent scenes because their hidden states are small and less expressive. We experiment with Test-Time Training (TTT) layers, whose hidden states themselves can be neural networks, therefore larger and more expressive. Adding TTT layers into a pre-trained Transformer enables it to generate one-minute videos from text storyboards. We curate a dataset based on Tom and Jerry cartoons as a proof-of-concept benchmark. Compared to baselines such as Mamba 2, Gated DeltaNet, and sliding-window attention layers, TTT layers generate much more coherent videos that tell complete stories, leading by 34 Elo points in a human evaluation of 100 videos per method. Although promising, our results are still limited in physical realism, and the efficiency of our implementation can be further improved.Sample videos, code and annotations are available at: https://test-time-training.github.io/video-dit Karan Dalal, Daniel Koceja, Yue Zhao 0006, Shihao Han, Ka Chun Cheung, Jan Kautz, Yejin Choi 0001, Yu Sun 0020, Xiaolong Wang 0004 |
CVPR | 2 |