VLDB 2026 Research / reviewers in the wild / expert
Karan Dalal
dblp:359/3103
· DBLP profile ↗
2ranked-venue papers
1as first author
2since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
Generative modeling · 43% Deep learning architectures and training · 38% Transfer learning and domain adaptation · 19% |
Topics — the 5 heaviest of 5, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Deep learning architectures and training
sequence modeling |
1.7 | 2 | 2025 | Learning to (Learn at Test Time): RNNs with Expressive Hidden States · ICML 2025 One-Minute Video Generation with Test-Time Training · CVPR 2025 |
Machine learning › Generative modeling › video generation
long video generation |
0.9 | 1 | 2025 | One-Minute Video Generation with Test-Time Training · CVPR 2025 |
Machine learning › Transfer learning and domain adaptation › test-time adaptation
test-time training |
0.9 | 1 | 2025 | Learning to (Learn at Test Time): RNNs with Expressive Hidden States · ICML 2025 |
Machine learning › Generative modeling
video generation |
0.9 | 1 | 2025 | One-Minute Video Generation with Test-Time Training · CVPR 2025 |
Machine learning › Generative modeling › video generation
text-to-video generation |
0.3 | 1 | 2025 | One-Minute Video Generation with Test-Time Training · CVPR 2025 |
Methods — techniques the papers use, named apart from their topics
transformer · 0.9test-time training · 0.9self-supervised learning · 0.9mamba layers · 0.9linear attention · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | One-Minute Video Generation with Test-Time TrainingabstractTransformers today still struggle to generate one-minute videos because self-attention layers are inefficient for long context. Alternatives such as Mamba layers struggle to produce coherent scenes because their hidden states are small and less expressive. We experiment with Test-Time Training (TTT) layers, whose hidden states themselves can be neural networks, therefore larger and more expressive. Adding TTT layers into a pre-trained Transformer enables it to generate one-minute videos from text storyboards. We curate a dataset based on Tom and Jerry cartoons as a proof-of-concept benchmark. Compared to baselines such as Mamba 2, Gated DeltaNet, and sliding-window attention layers, TTT layers generate much more coherent videos that tell complete stories, leading by 34 Elo points in a human evaluation of 100 videos per method. Although promising, our results are still limited in physical realism, and the efficiency of our implementation can be further improved.Sample videos, code and annotations are available at: https://test-time-training.github.io/video-dit Karan Dalal, Daniel Koceja, Yue Zhao 0006, Shihao Han, Ka Chun Cheung, Jan Kautz, Yejin Choi 0001, Yu Sun 0020, Xiaolong Wang 0004 |
CVPR | 1 |
| 2025 | Learning to (Learn at Test Time): RNNs with Expressive Hidden StatesabstractSelf-attention performs well in long context but has quadratic complexity. Existing RNN layers have linear complexity, but their performance in long context is limited by the expressive power of their hidden states. We present a practical framework for instantiating sequence modeling layers with linear complexity and expressive hidden states. The key idea is to make the hidden state a machine learning model itself, and the update rule a step of self-supervised learning. Since the hidden state is updated by training even on test sequences, our layers are called Test-Time Training (TTT) layers. We consider two instantiations: TTT-Linear and TTT-MLP, whose hidden state is a linear model and a two-layer MLP respectively. We evaluate our instantiations at the scale of 125M to 1.3B parameters, comparing with a strong Transformer and Mamba, a modern RNN. Similar to Transformer, TTT-Linear and TTT-MLP can keep reducing perplexity by conditioning on more tokens, while Mamba cannot after 16k context. TTT-MLP still faces challenges in memory I/O, but shows larger potential in long context, pointing to a promising direction for future research. Yu Sun 0020, Karan Dalal, Arjun Vikram, Genghan Zhang, Yann Dubois, Xinlei Chen, Xiaolong Wang 0004, Oluwasanmi Koyejo, Tatsunori B. Hashimoto, Carlos Guestrin |
ICML | 3 |