Chaofan Gan

dblp:383/8768 · DBLP profile ↗
← Back
5ranked-venue papers
2as first author
5since 2021 · last 2026
0009-0001-5297-2202ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Video understanding and tracking · 23% Knowledge representation and reasoning · 22% 3D vision · 20%
Databases, data mining, and information retrieval
1 paper
Information retrieval · 100%

Topics — the 11 heaviest of 12, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Video understanding and tracking › deep video understanding › video reasoning
video causal reasoning
1.822026
MECD+: Unlocking Event-Level Causal Graph Discovery for Video Reasoning · IEEE Trans. Pattern Anal. Mach. Intell. 2026
MECD: Unlocking Multi-Event Causal Discovery in Video Reasoning · NeurIPS 2024
Knowledge, reasoning and agents › Knowledge representation and reasoning › causal reasoning
causal graph discovery
1.012026
MECD+: Unlocking Event-Level Causal Graph Discovery for Video Reasoning · IEEE Trans. Pattern Anal. Mach. Intell. 2026
Knowledge, reasoning and agents › Knowledge representation and reasoning
causal reasoning
1.012026
MECD+: Unlocking Event-Level Causal Graph Discovery for Video Reasoning · IEEE Trans. Pattern Anal. Mach. Intell. 2026
Computer vision › 3D vision › correspondence estimation
dense correspondence
0.912025
Unleashing Diffusion Transformers for Visual Correspondence by Modulating Massive Activations · NeurIPS 2025
Computer vision › 3D vision › correspondence estimation
image correspondence
0.912025
Unleashing Diffusion Transformers for Visual Correspondence by Modulating Massive Activations · NeurIPS 2025
Machine learning › Probabilistic and Bayesian machine learning
causal inference
0.812024
MECD: Unlocking Multi-Event Causal Discovery in Video Reasoning · NeurIPS 2024
Machine learning › Probabilistic and Bayesian machine learning › causal inference › causal discovery
granger causality
0.812024
MECD: Unlocking Multi-Event Causal Discovery in Video Reasoning · NeurIPS 2024
Machine learning › Trustworthy machine learning › robustness
learning with noisy labels
0.812024
DAC: 2D-3D Retrieval with Noisy Labels via Divide-and-Conquer Alignment and Correction · ACM Multimedia 2024
Information retrieval › cross-modal retrieval
2d-3d retrieval
0.812024
DAC: 2D-3D Retrieval with Noisy Labels via Divide-and-Conquer Alignment and Correction · ACM Multimedia 2024
Information retrieval
cross-modal retrieval
0.812024
DAC: 2D-3D Retrieval with Noisy Labels via Divide-and-Conquer Alignment and Correction · ACM Multimedia 2024
Computer vision › Video understanding and tracking
video question answering
0.312026
MECD+: Unlocking Event-Level Causal Graph Discovery for Video Reasoning · IEEE Trans. Pattern Anal. Mach. Intell. 2026

Methods — techniques the papers use, named apart from their topics

mask-based event prediction · 1.8front-door adjustment · 1.8counterfactual inference · 1.8self-correction · 1.5divide-and-conquer · 1.5adaptive alignment · 1.5granger causality · 1.0channel discard strategy · 0.9adaptive layer normalization · 0.9
YearPublicationVenuePosition
2026 Looking Beyond Visible Cues: Implicit Video Question Answering via Dual-Clue Reasoning
Tieyuan Chen, Huabin Liu 0001, Yi Wang 0033, Chaofan Gan, Mingxi Lv, Ziran Qin, Li Shen 0008, Junhui Hou, Weiyao Lin
Int. J. Comput. Vis.4
2026 MECD+: Unlocking Event-Level Causal Graph Discovery for Video Reasoning
abstract
Video causal reasoning aims to achieve a high-level understanding of videos from a causal perspective. However, it exhibits limitations in its scope, primarily executed in a question-answering paradigm and focusing on brief video segments containing isolated events and basic causal relations, lacking comprehensive and structured causality analysis for videos with multiple interconnected events. To fill this gap, we introduce a new task and dataset, Multi-Event Causal Discovery (MECD). It aims to uncover the causal relations between events distributed chronologically across long videos. Given visual segments and textual descriptions of events, MECD identifies the causal associations between these events to derive a comprehensive and structured event-level video causal graph explaining why and how the result event occurred. To address the challenges of MECD, we devise a novel framework inspired by the Granger Causality method, incorporating an efficient mask-based event prediction model to perform an Event Granger Test. It estimates causality by comparing the predicted result event when premise events are masked versus unmasked. Furthermore, we integrate causal inference techniques such as front-door adjustment and counterfactual inference to mitigate challenges in MECD like causality confounding and illusory causality. Additionally, context chain reasoning is introduced to conduct more robust and generalized reasoning. Experiments validate the effectiveness of our framework in reasoning complete causal relations, outperforming GPT-4o and VideoChat2 by 5.77% and 2.70%, respectively. Further experiments demonstrate that causal relation graphs can also contribute to downstream video understanding tasks such as video question answering and video event prediction.
Tieyuan Chen, Huabin Liu 0001, Yi Wang 0033, Yihang Chen 0002, Tianyao He, Chaofan Gan, Huanyu He, Weiyao Lin
IEEE Trans. Pattern Anal. Mach. Intell.6
2025 Unleashing Diffusion Transformers for Visual Correspondence by Modulating Massive Activations
abstract
Pre-trained stable diffusion models (SD) have shown great advances in visual correspondence. In this paper, we investigate the capabilities of Diffusion Transformers (DiTs) for accurate dense correspondence. Distinct from SD, DiTs exhibit a critical phenomenon in which very few feature activations exhibit significantly larger values than others, known as massive activations, leading to uninformative representations and significant performance degradation for DiTs. The massive activations consistently concentrate at very few fixed dimensions across all image patch tokens, holding little local information. We analyze these dimension-concentrated massive activations and uncover that their concentration is inherently linked to the Adaptive Layer Normalization (AdaLN) in DiTs. Building on these findings, we propose the Diffusion Transformer Feature (DiTF), a training-free AdaLN-based framework that extracts semantically discriminative features from DiTs. Specifically, DiTF leverages AdaLN to adaptively localize and normalize massive activations through channel-wise modulation. Furthermore, a channel discard strategy is introduced to mitigate the adverse effects of massive activations. Experimental results demonstrate that our DiTF outperforms both DINO and SD-based models and establishes a new state-of-the-art performance for DiTs in different visual correspondence tasks (e.g., with +9.4\% on Spair-71k and +4.4\% on AP-10K-C.S.).
Chaofan Gan, Yuanpeng Tu, Tieyuan Chen, Yuxi Li 0009, Mehrtash Harandi, Weiyao Lin
NeurIPS1
2024 DAC: 2D-3D Retrieval with Noisy Labels via Divide-and-Conquer Alignment and Correction
abstract
With the recent burst of 2D and 3D data, cross-modal retrieval has attracted increasing attention recently. However, manual labeling by non-experts will inevitably introduce corrupted annotations given ambiguous 2D/3D content. Though previous works have addressed this issue by designing a naive division strategy with hand-crafted thresholds, their performance generally exhibits great sensitivity to the threshold value. Besides, they fail to fully utilize the valuable supervisory signals within each divided subset. To tackle this problem, we propose a Divide-and-conquer 2D-3D cross-modal Alignment and Correction framework (DAC), which comprises Multimodal Dynamic Division (MDD) and Adaptive Alignment and Correction (AAC). Specifically, the former performs accurate sample division by adaptive credibility modeling for each sample based on the compensation information within multimodal loss distribution. Then in AAC, samples in distinct subsets are exploited with different alignment strategies to fully enhance the semantic compactness and meanwhile alleviate over-fitting to noisy labels, where a self-correction strategy is introduced to improve the quality of representation. Moreover. To evaluate the effectiveness in real-world scenarios, we introduce a challenging noisy benchmark, namely Objaverse-N200, which comprises 200k-level samples annotated with 1156 realistic noisy labels. Extensive experiments on both traditional and the newly proposed benchmarks demonstrate the generality and superiority of our DAC, where DAC outperforms state-of-the-art models by a large margin. (i.e., with +5.9% gain on ModelNet40 and +5.8% on Objaverse-N200).
Chaofan Gan, Yuanpeng Tu, Yuxi Li 0009, Weiyao Lin
ACM Multimedia1
2024 MECD: Unlocking Multi-Event Causal Discovery in Video Reasoning
abstract
Video causal reasoning aims to achieve a high-level understanding of video content from a causal perspective. However, current video reasoning tasks are limited in scope, primarily executed in a question-answering paradigm and focusing on short videos containing only a single event and simple causal relationships, lacking comprehensive and structured causality analysis for videos with multiple events. To fill this gap, we introduce a new task and dataset, Multi-Event Causal Discovery (MECD). It aims to uncover the causal relationships between events distributed chronologically across long videos. Given visual segments and textual descriptions of events, MECD requires identifying the causal associations between these events to derive a comprehensive, structured event-level video causal diagram explaining why and how the final result event occurred. To address MECD, we devise a novel framework inspired by the Granger Causality method, using an efficient mask-based event prediction model to perform an Event Granger Test, which estimates causality by comparing the predicted result event when premise events are masked versus unmasked. Furthermore, we integrate causal inference techniques such as front-door adjustment and counterfactual inference to address challenges in MECD like causality confounding and illusory causality. Experiments validate the effectiveness of our framework in providing causal relationships in multi-event videos, outperforming GPT-4o and VideoLLaVA by 5.7% and 4.1%, respectively.
Tieyuan Chen, Huabin Liu 0001, Tianyao He, Yihang Chen 0002, Chaofan Gan, Yang Zhang 0002, Weiyao Lin
NeurIPS5