VLDB 2026 Research / reviewers in the wild / expert
Tieyuan Chen
dblp:347/1572
· DBLP profile ↗
6ranked-venue papers
4as first author
6since 2021 · last 2026
0009-0005-7939-7139ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 3 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
4 papers |
Video understanding and tracking · 40% Knowledge representation and reasoning · 20% 3D vision · 17% |
Topics — the 10 heaviest of 11, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Video understanding and tracking › deep video understanding › video reasoning
video causal reasoning |
1.8 | 2 | 2026 | MECD+: Unlocking Event-Level Causal Graph Discovery for Video Reasoning · IEEE Trans. Pattern Anal. Mach. Intell. 2026 MECD: Unlocking Multi-Event Causal Discovery in Video Reasoning · NeurIPS 2024 |
Computer vision › Video understanding and tracking
action recognition |
1.0 | 1 | 2026 | Few-Shot Action Recognition via Intra- and Inter-Video Information Maximization · IEEE Trans. Pattern Anal. Mach. Intell. 2026 |
Knowledge, reasoning and agents › Knowledge representation and reasoning › causal reasoning
causal graph discovery |
1.0 | 1 | 2026 | MECD+: Unlocking Event-Level Causal Graph Discovery for Video Reasoning · IEEE Trans. Pattern Anal. Mach. Intell. 2026 |
Knowledge, reasoning and agents › Knowledge representation and reasoning
causal reasoning |
1.0 | 1 | 2026 | MECD+: Unlocking Event-Level Causal Graph Discovery for Video Reasoning · IEEE Trans. Pattern Anal. Mach. Intell. 2026 |
Computer vision › Video understanding and tracking › action recognition
few-shot action recognition |
1.0 | 1 | 2026 | Few-Shot Action Recognition via Intra- and Inter-Video Information Maximization · IEEE Trans. Pattern Anal. Mach. Intell. 2026 |
Computer vision › 3D vision › correspondence estimation
dense correspondence |
0.9 | 1 | 2025 | Unleashing Diffusion Transformers for Visual Correspondence by Modulating Massive Activations · NeurIPS 2025 |
Computer vision › 3D vision › correspondence estimation
image correspondence |
0.9 | 1 | 2025 | Unleashing Diffusion Transformers for Visual Correspondence by Modulating Massive Activations · NeurIPS 2025 |
Machine learning › Probabilistic and Bayesian machine learning
causal inference |
0.8 | 1 | 2024 | MECD: Unlocking Multi-Event Causal Discovery in Video Reasoning · NeurIPS 2024 |
Machine learning › Probabilistic and Bayesian machine learning › causal inference › causal discovery
granger causality |
0.8 | 1 | 2024 | MECD: Unlocking Multi-Event Causal Discovery in Video Reasoning · NeurIPS 2024 |
Computer vision › Video understanding and tracking
video question answering |
0.3 | 1 | 2026 | MECD+: Unlocking Event-Level Causal Graph Discovery for Video Reasoning · IEEE Trans. Pattern Anal. Mach. Intell. 2026 |
Methods — techniques the papers use, named apart from their topics
mask-based event prediction · 1.8front-door adjustment · 1.8counterfactual inference · 1.8mutual information maximization · 1.0granger causality · 1.0adaptive spatial-temporal sampling · 1.0channel discard strategy · 0.9adaptive layer normalization · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Looking Beyond Visible Cues: Implicit Video Question Answering via Dual-Clue Reasoning
Tieyuan Chen, Huabin Liu 0001, Yi Wang 0033, Chaofan Gan, Mingxi Lv, Ziran Qin, Li Shen 0008, Junhui Hou, Weiyao Lin |
Int. J. Comput. Vis. | 1 |
| 2026 | MECD+: Unlocking Event-Level Causal Graph Discovery for Video ReasoningabstractVideo causal reasoning aims to achieve a high-level understanding of videos from a causal perspective. However, it exhibits limitations in its scope, primarily executed in a question-answering paradigm and focusing on brief video segments containing isolated events and basic causal relations, lacking comprehensive and structured causality analysis for videos with multiple interconnected events. To fill this gap, we introduce a new task and dataset, Multi-Event Causal Discovery (MECD). It aims to uncover the causal relations between events distributed chronologically across long videos. Given visual segments and textual descriptions of events, MECD identifies the causal associations between these events to derive a comprehensive and structured event-level video causal graph explaining why and how the result event occurred. To address the challenges of MECD, we devise a novel framework inspired by the Granger Causality method, incorporating an efficient mask-based event prediction model to perform an Event Granger Test. It estimates causality by comparing the predicted result event when premise events are masked versus unmasked. Furthermore, we integrate causal inference techniques such as front-door adjustment and counterfactual inference to mitigate challenges in MECD like causality confounding and illusory causality. Additionally, context chain reasoning is introduced to conduct more robust and generalized reasoning. Experiments validate the effectiveness of our framework in reasoning complete causal relations, outperforming GPT-4o and VideoChat2 by 5.77% and 2.70%, respectively. Further experiments demonstrate that causal relation graphs can also contribute to downstream video understanding tasks such as video question answering and video event prediction. Tieyuan Chen, Huabin Liu 0001, Yi Wang 0033, Yihang Chen 0002, Tianyao He, Chaofan Gan, Huanyu He, Weiyao Lin |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2026 | Few-Shot Action Recognition via Intra- and Inter-Video Information MaximizationabstractCurrent few-shot action recognition involves two primary sources of information for classification: (1) intra-video information, determined by frame content within a single video clip, and (2) inter-video information, measured by relationships (e.g., feature similarity) among videos. However, existing methods inadequately exploit these two information sources. In terms of intra-video information, current sampling operations for input videos may omit critical action information, reducing the utilization efficiency of video data. For the inter-video information, the action misalignment among videos makes it challenging to calculate precise relationships. Moreover, how to jointly consider both inter- and intra-video information remains under-explored for few-shot action recognition. To this end, we propose a novel framework, Video Information Maximization (VIM), for few-shot video action recognition. VIM is equipped with an adaptive spatial-temporal video sampler and a spatial-temporal action alignment model to maximize intra- and inter-video information, respectively. The video sampler adaptively selects important frames and amplifies critical spatial regions for each input video based on the task at hand. This preserves and emphasizes informative parts of video clips while eliminating interference at the data level. The alignment model performs temporal and spatial action alignment sequentially at the feature level, leading to more precise measurements of inter-video similarity. Finally, based on the mutual information measurement, we introduce a new training objective into few-shot learning, which provides explicit guidance in jointly maximizing intra- and inter-video information in our VIM. Extensive experimental results on public datasets for few-shot action recognition demonstrate the effectiveness of our framework. Huabin Liu 0001, Tieyuan Chen, Yuxi Li 0009, Shuyuan Li, John See, Weiyao Lin |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2025 | Unleashing Diffusion Transformers for Visual Correspondence by Modulating Massive ActivationsabstractPre-trained stable diffusion models (SD) have shown great advances in visual correspondence.
In this paper, we investigate the capabilities of Diffusion Transformers (DiTs) for accurate dense correspondence. Distinct from SD, DiTs exhibit a critical phenomenon in which very few feature activations exhibit significantly larger values than others, known as massive activations, leading to uninformative representations and significant performance degradation for DiTs.
The massive activations consistently concentrate at very few fixed dimensions across all image patch tokens, holding little local information.
We analyze these dimension-concentrated massive activations and uncover that their concentration is inherently linked to the Adaptive Layer Normalization (AdaLN) in DiTs.
Building on these findings, we propose the Diffusion Transformer Feature (DiTF), a training-free AdaLN-based framework that extracts semantically discriminative features from DiTs.
Specifically, DiTF leverages AdaLN to adaptively localize and normalize massive activations through channel-wise modulation.
Furthermore, a channel discard strategy is introduced to mitigate the adverse effects of massive activations.
Experimental results demonstrate that our DiTF outperforms both DINO and SD-based models and establishes a new state-of-the-art performance for DiTs in different visual correspondence tasks (e.g., with +9.4\% on Spair-71k and +4.4\% on AP-10K-C.S.). Chaofan Gan, Yuanpeng Tu, Tieyuan Chen, Yuxi Li 0009, Mehrtash Harandi, Weiyao Lin |
NeurIPS | 4 |
| 2025 | CSTA: Spatial-Temporal Causal Adaptive Learning for Exemplar-Free Video Class-Incremental LearningabstractContinual learning aims to acquire new knowledge while retaining past information. Class-incremental learning (CIL) presents a challenging scenario where classes are introduced sequentially. For video data, the task becomes more complex than image data because it requires learning and preserving both spatial appearance and temporal action involvement. To address this challenge, we propose a novel exemplar-free framework that equips separate spatiotemporal adapters to learn new class patterns, accommodating the incremental information representation requirements unique to each class. While separate adapters are proven to mitigate forgetting and fit unique requirements, naively applying them hinders the intrinsic connection between spatial and temporal information increments, affecting the efficiency of representing newly learned class information. Motivated by this, we introduce two key innovations from a causal perspective. First, a causal distillation module is devised to maintain the relation between spatial-temporal knowledge for a more efficient representation. Second, a causal compensation mechanism is proposed to reduce the conflicts during increment and memorization between different types of information. Extensive experiments conducted on benchmark datasets demonstrate that our framework can achieve new state-of-the-art results, surpassing current example-based methods by 4.2% in accuracy on average. The codes are accessible in https://github.com/tychen-SJTU/CSTA. Tieyuan Chen, Huabin Liu 0001, Chern Hong Lim, John See, Xing Gao 0005, Junhui Hou, Weiyao Lin |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2024 | MECD: Unlocking Multi-Event Causal Discovery in Video ReasoningabstractVideo causal reasoning aims to achieve a high-level understanding of video content from a causal perspective. However, current video reasoning tasks are limited in scope, primarily executed in a question-answering paradigm and focusing on short videos containing only a single event and simple causal relationships, lacking comprehensive and structured causality analysis for videos with multiple events. To fill this gap, we introduce a new task and dataset, Multi-Event Causal Discovery (MECD). It aims to uncover the causal relationships between events distributed chronologically across long videos. Given visual segments and textual descriptions of events, MECD requires identifying the causal associations between these events to derive a comprehensive, structured event-level video causal diagram explaining why and how the final result event occurred. To address MECD, we devise a novel framework inspired by the Granger Causality method, using an efficient mask-based event prediction model to perform an Event Granger Test, which estimates causality by comparing the predicted result event when premise events are masked versus unmasked. Furthermore, we integrate causal inference techniques such as front-door adjustment and counterfactual inference to address challenges in MECD like causality confounding and illusory causality. Experiments validate the effectiveness of our framework in providing causal relationships in multi-event videos, outperforming GPT-4o and VideoLLaVA by 5.7% and 4.1%, respectively. Tieyuan Chen, Huabin Liu 0001, Tianyao He, Yihang Chen 0002, Chaofan Gan, Yang Zhang 0002, Weiyao Lin |
NeurIPS | 1 |