EDBT 2026 Demo / reviewers in the wild / expert
Donglin Di
dblp:242/4517
· DBLP profile ↗
4ranked-venue papers in the field
0as first author
3since 2021 · last 2025
0000-0002-2270-3378ORCID · verified
Domains — venue-derived; a paper can count in several
Other / Interdisciplinary · 3Information Retrieval & Web Search · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | GAOT: Generating Articulated Objects Through Text-Guided Diffusion ModelsabstractArticulated object generation has seen increasing advancements, yet existing models often lack the ability to be conditioned on text prompts. To address the significant gap between textual descriptions and 3D articulated object representations, we propose GAOT, a three-phase framework that generates articulated objects from text prompts, leveraging diffusion models and hypergraph learning in a three-step process. Lei Fan 0007, Donglin Di, Shaohui Liu |
MMAsia | 3 |
| 2025 | UniCP: A Unified Caching and Pruning Framework for Efficient Video GenerationabstractDiffusion Transformers (DiT) excel in video generation but encounter significant computational challenges due to the quadratic complexity of attention. Notably, attention differences between adjacent diffusion steps follow a U-shaped pattern. Current methods leverage this property by caching attention blocks; however, they still struggle with sudden error spikes and large discrepancies. To address these issues, we propose UniCP—a unified caching and pruning framework for efficient video generation. UniCP optimizes both temporal and spatial dimensions through: Error-Aware Dynamic Cache Window (EDCW): Dynamically adjusts cache window sizes for different blocks at various timesteps to adapt to abrupt error changes. PCA-based Slicing (PCAS) and Dynamic Weight Shift (DWS): PCAS prunes redundant attention components, while DWS integrates caching and pruning by enabling dynamic switching between pruned and cached outputs. By adjusting cache windows and pruning redundant components, UniCP enhances computational efficiency and maintains video detail fidelity. Experimental results show that UniCP outperforms existing methods, delivering superior performance and efficiency. Wenzhang Sun, Qirui Hou, Donglin Di, Yongjia Ma, Jianxun Cui |
MMAsia | 3 |
| 2025 | MoCA: Identity-Preserving Text-to-Video Generation via Mixture of Cross AttentionabstractAchieving ID-preserving text-to-video (T2V) generation remains challenging despite recent advances in diffusion-based models. Existing approaches often fail to capture fine-grained facial dynamics or maintain temporal identity coherence. To address these limitations, we propose MoCA, a novel Video Diffusion Model built on a Diffusion Transformer (DiT) backbone, incorporating a Mixture of Cross-Attention mechanism inspired by the Mixture-of-Experts paradigm. Our framework improves inter-frame identity consistency by embedding MoCA layers into each DiT block, where Hierarchical Temporal Pooling captures identity features over varying timescales, and Temporal-Aware Cross-Attention Experts dynamically model spatiotemporal relationships. We further incorporate a Latent Video Perceptual Loss to enhance identity coherence and fine-grained details across video frames. To train this model, we collect CelebIPVid, a dataset of 10,000 high-resolution videos from 1,000 diverse individuals, promoting cross-ethnic generalization. Extensive experiments on CelebIPVid show that MoCA outperforms existing T2V methods by over 5% across facial similarity. Qi Xie 0009, Yongjia Ma, Donglin Di, Xuehao Gao, Xun Yang 0001 |
MMAsia | 3 |
| 2019 | Annotating Objects and Relations in User-Generated VideosabstractUnderstanding the objects and relations between them is indispensable to fine-grained video content analysis, which is widely studied in recent research works in multimedia and computer vision. However, existing works are limited to evaluating with either small datasets or indirect metrics, such as the performance over images. The underlying reason is that the construction of a large-scale video dataset with dense annotation is tricky and costly. In this paper, we address several main issues in annotating objects and relations in user-generated videos, and propose an annotation pipeline that can be executed at a modest cost. As a result, we present a new dataset, named VidOR, consisting of 10k videos (84 hours) together with dense annotations that localize 80 categories of objects and 50 categories of predicates in each video. We have made the training and validation set public and extendable for more tasks to facilitate future research on video object and relation recognition. Xindi Shang, Donglin Di, Junbin Xiao, Xun Yang 0001, Tat-Seng Chua |
ICMR | 2 |