EDBT 2026 Demo / reviewers in the wild / expert
Yongjia Ma
dblp:366/7630
· DBLP profile ↗
11ranked-venue papers
2as first author
11since 2021 · last 2026
0000-0002-0070-4134ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 8 · 2 first-author · 8 since 2021Artificial intelligence and machine learning · 5 · 5 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Computer networks · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DiTalker: A unified DiT-based framework for high-quality and style-controllable portrait animation
Yongjia Ma, Lei Fan 0007, Donglin Di, Tonghua Su |
Comput. Vis. Image Underst. | 2 |
| 2026 | Tuning-Free Long Video Generation via Global-Local Collaborative DiffusionabstractCreating high-fidelity, coherent long videos is a sought-after aspiration. While recent video diffusion models have shown promising potential, they still grapple with spatiotemporal inconsistencies and high computational resource demands. We propose Global-Local Collaborative Diffusion (GLC-Diffusion), a tuning-free method for long video generation. It models the long video denoising process by establishing denoising trajectories through Global-Local Collaborative Denoising (GLCD) to ensure overall content consistency and temporal coherence between frames. Additionally, we introduce a Noise Reinitialization strategy which combines local noise shuffling with frequency fusion to improve global content consistency and visual diversity. Further, we propose a Video Motion Consistency Refinement (VMCR) module that computes the gradient of pixel-wise and frequency-wise losses to enhance visual consistency and temporal smoothness. Extensive experiments, including quantitative and qualitative evaluations on videos of varying lengths (e.g., 3× and 6× longer), demonstrate that our method effectively integrates with existing video diffusion models, producing coherent, high-fidelity long videos superior to previous approaches. Yongjia Ma, Junlin Chen, Donglin Di, Qi Xie 0009, Lei Fan 0007, Wei Chen 0089, Na Zhao 0004, Xun Yang 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2025 | DH-FaceVid-1K: A Large-Scale High-Quality Dataset for Face Video Generation
Donglin Di, Wenzhang Sun, Yongjia Ma, Hao Li 0030, Wei Chen 0089, Lei Fan 0007, Tonghua Su, Xun Yang 0001 |
ICCV | 4 |
| 2025 | QR-LoRA: Efficient and Disentangled Fine-Tuning via QR Decomposition for Customized GenerationabstractExisting text-to-image models often rely on parameter fine-tuning techniques such as Low-Rank Adaptation (LoRA) to customize visual attributes. However, when combining multiple LoRA models for content-style fusion tasks, unstructured modifications of weight matrices often lead to undesired feature entanglement between content and style attributes. We propose QR-LoRA, a novel fine-tuning framework leveraging QR decomposition for structured parameter updates that effectively separate visual attributes. Our key insight is that the orthogonal Q matrix naturally minimizes interference between different visual features, while the upper triangular R matrix efficiently encodes attribute-specific transformations. Our approach fixes both Q and R matrices while only training an additional task-specific $ΔR$ matrix. This structured design reduces trainable parameters to half of conventional LoRA methods and supports effective merging of multiple adaptations without cross-contamination due to the strong disentanglement properties between $ΔR$ matrices. Experiments demonstrate that QR-LoRA achieves superior disentanglement in content-style fusion tasks, establishing a new paradigm for parameter-efficient, disentangled fine-tuning in generative models. The project page is available at: https://luna-ai-lab.github.io/QR-LoRA/. Yongjia Ma, Donglin Di, Jianxun Cui, Hao Li 0030, Wei Chen 0089, Xun Yang 0001, Wangmeng Zuo |
ICCV | 2 |
| 2025 | Multi-scale Feature Field with Anti-brightness-sensitivity Postprocessing for Few-shot Neural Panoptic SegmentationabstractNeural scene segmentation, which achieves 3D scene reconstruction and segmentation via 2D posed inputs, has garnered substantial attention. However, previous methods have relied on dense-view manual labels, while some improved few-shot or zero-shot models based on semantic-aware features or generated labels still encounter two challenges: (i) Lack of comprehensive perceptual ability, which limits their performance especially for complex scene understanding. (ii) Misleading by brightness-sensitive priors, which affects the segmentation particularly for specular highlight regions. Therefore, in this work, we propose Multi-scale Distilled & Anti-brightness-sensitivity Postprocessed Neural Feature Field (MDAP-NFF), a method for few-shot neural panoptic segmentation. To enhance perceptual ability, multi-scale feature field is designed for distilling multi-scale semantic-aware, e.g., DINO, features. For the robustness improvement against brightness-sensitivity in distilled feature field, we propose the anti-brightness-sensitivity postprocessing module, incorporating our designed 3Donv-VM technique. By training the feature field via 3D distillation and further postprocessing the features under supervision of sparse-view labels, our method achieves reliable 3D-consistent panoptic segmentation. Experimental results show that our model can perform high-quality panoptic segmentation under sparse-view supervision. Code and more results will be available at https://David-Dou.github.io/MDAP-NFF Bin Dou, Yongjia Ma, Zejian Yuan |
ICMR | 2 |
| 2025 | UniCP: A Unified Caching and Pruning Framework for Efficient Video GenerationabstractDiffusion Transformers (DiT) excel in video generation but encounter significant computational challenges due to the quadratic complexity of attention. Notably, attention differences between adjacent diffusion steps follow a U-shaped pattern. Current methods leverage this property by caching attention blocks; however, they still struggle with sudden error spikes and large discrepancies. To address these issues, we propose UniCP—a unified caching and pruning framework for efficient video generation. UniCP optimizes both temporal and spatial dimensions through: Error-Aware Dynamic Cache Window (EDCW): Dynamically adjusts cache window sizes for different blocks at various timesteps to adapt to abrupt error changes. PCA-based Slicing (PCAS) and Dynamic Weight Shift (DWS): PCAS prunes redundant attention components, while DWS integrates caching and pruning by enabling dynamic switching between pruned and cached outputs. By adjusting cache windows and pruning redundant components, UniCP enhances computational efficiency and maintains video detail fidelity. Experimental results show that UniCP outperforms existing methods, delivering superior performance and efficiency. Wenzhang Sun, Qirui Hou, Donglin Di, Yongjia Ma, Jianxun Cui |
MMAsia | 5 |
| 2025 | MoCA: Identity-Preserving Text-to-Video Generation via Mixture of Cross AttentionabstractAchieving ID-preserving text-to-video (T2V) generation remains challenging despite recent advances in diffusion-based models. Existing approaches often fail to capture fine-grained facial dynamics or maintain temporal identity coherence. To address these limitations, we propose MoCA, a novel Video Diffusion Model built on a Diffusion Transformer (DiT) backbone, incorporating a Mixture of Cross-Attention mechanism inspired by the Mixture-of-Experts paradigm. Our framework improves inter-frame identity consistency by embedding MoCA layers into each DiT block, where Hierarchical Temporal Pooling captures identity features over varying timescales, and Temporal-Aware Cross-Attention Experts dynamically model spatiotemporal relationships. We further incorporate a Latent Video Perceptual Loss to enhance identity coherence and fine-grained details across video frames. To train this model, we collect CelebIPVid, a dataset of 10,000 high-resolution videos from 1,000 diverse individuals, promoting cross-ethnic generalization. Extensive experiments on CelebIPVid show that MoCA outperforms existing T2V methods by over 5% across facial similarity. Qi Xie 0009, Yongjia Ma, Donglin Di, Xuehao Gao, Xun Yang 0001 |
MMAsia | 2 |
| 2025 | EverybodyDance: Bipartite Graph-Based Identity Correspondence for Multi-Character AnimationabstractConsistent pose‐driven character animation has achieved remarkable progress in single‐character scenarios. However, extending these advances to multi‐character settings is non‐trivial, especially when position swap is involved. Beyond mere scaling, the core challenge lies in enforcing correct Identity Correspondence (IC) between characters in reference and generated frames. To address this, we introduce EverybodyDance, a systematic solution targeting IC correctness in multi-character animation. EverybodyDance is built around the **Identity Matching Graph (IMG)**, which models characters in the generated and reference frames as two node sets in a weighted complete bipartite graph. Edge weights, computed via our proposed Mask–Query Attention (MQA), quantify the affinity between each pair of characters. Our key insight is to formalize IC correctness as a graph structural metric and to optimize it during training. We also propose a series of targeted strategies tailored for multi-character animation, including identity-embedded guidance, a multi-scale matching strategy, and pre-classified sampling, which work synergistically. Finally, to evaluate IC performance, we curate the **Identity Correspondence Evaluation** benchmark, dedicated to multi‐character IC correctness. Extensive experiments demonstrate that EverybodyDance substantially outperforms state‐of‐the‐art baselines in both IC and visual fidelity. Haotian Ling, Zequn Chen, Qiuying Chen, Donglin Di, Yongjia Ma, Hao Li 0030, Zhulin Tao, Xun Yang 0001 |
NeurIPS | 5 |
| 2025 | TrAME: Trajectory-Anchored Multi-View Editing for Text-Guided 3D Gaussian ManipulationabstractDespite significant strides in the field of 3D scene editing, current methods encounter substantial challenge, particularly in preserving 3D consistency during the multi-view editing process. To tackle this challenge, we propose a progressive 3D editing strategy that ensures multi-view consistency via a Trajectory-Anchored Scheme (TAS) with a dual-branch editing mechanism. Specifically, TAS facilitates a tightly coupled iterative process between 2D view editing and 3D updating, preventing error accumulation yielded from the text-to-image process. Additionally, we explore the connection between optimization-based methods and reconstruction-based methods, offering a unified perspective for selecting superior design choices, supporting the rationale behind the designed TAS. We further present a tuning-free View-Consistent Attention Control (VCAC) module that leverages cross-view semantic and geometric reference from the source branch to yield aligned views from the target branch during the editing of 2D views. To validate the effectiveness of our method, we analyze 2D examples to demonstrate the improved consistency with the VCAC module. Extensive quantitative and qualitative results in text-guided 3D scene editing clearly indicate that our method can achieve superior editing quality compared with state-of-the-art 3D scene editing methods. Our project site is athttps://fkcptlst.github.io/TrAME/ Chaofan Luo, Donglin Di, Xun Yang 0001, Yongjia Ma, Zhou Xue, Wei Chen 0089, Xiaofei Gou, Yebin Liu |
IEEE Trans. Multim. | 4 |
| 2024 | RD-NERF: Neural Robust Distilled Feature Fields for Sparse-View Scene SegmentationabstractWe propose Neural Robust Distilled Feature Fields (RD-NeRF) for achieving robust 3D semantic feature distillation and 3D consistent scene segmentation with sparse-view labels. Specifically, we introduce a two-stage pipeline. In the distillation stage, we employ the pre-trained image feature extractor, DINO-ViT, as the teacher network. RD-NeRF distills semantic knowledge into 3D space and utilizes the Vector-Matrix (VM) tensor decomposition method to represent semantic field with volumetric rendering. For the training process, we utilize the distance-wise and angle-wise distillation loss. This enables the student network to capture high-level semantics, enhance scene reconstruction and segmentation performance, and improve robustness and effectiveness in distillation. In the segmentation stage, hash features and distilled semantic features are inputs for the segmentation MLP, which is supervised by the sparse-view labels. The experimental results demonstrate that our model performs well in 3D-consistent scene segmentation under sparse-view supervision. Yongjia Ma, Bin Dou, Zejian Yuan |
ICASSP | 1 |
| 2024 | Learning Segmented 3D Gaussians via Efficient Feature Unprojection for Zero-Shot Neural Scene Segmentation
Bin Dou, Yongjia Ma, Zejian Yuan, Nanning Zheng 0001 |
ICONIP (8) | 4 |