VLDB 2026 Research / reviewers in the wild / expert
Tengjiao Sun
dblp:402/7645
· DBLP profile ↗
2ranked-venue papers
0as first author
2since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
1 paper |
Generative modeling · 50% Representation and self-supervised learning · 50% | |
| Computer graphics and multimedia
1 paper |
Computer animation and physical simulation · 100% |
Topics — the 4 heaviest of 4, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Generative modeling
autoregressive model |
1.0 | 1 | 2026 | MOGO: Residual Quantized Hierarchical Causal Transformer for Real-Time and Infinite-Length 3D Human Motion Generation · AAAI 2026 |
Machine learning › Representation and self-supervised learning
vector quantization |
1.0 | 1 | 2026 | MOGO: Residual Quantized Hierarchical Causal Transformer for Real-Time and Infinite-Length 3D Human Motion Generation · AAAI 2026 |
Computer animation and physical simulation › motion synthesis
human motion synthesis |
1.0 | 1 | 2026 | MOGO: Residual Quantized Hierarchical Causal Transformer for Real-Time and Infinite-Length 3D Human Motion Generation · AAAI 2026 |
Computer animation and physical simulation › motion synthesis › human motion synthesis
text-to-motion generation |
1.0 | 1 | 2026 | MOGO: Residual Quantized Hierarchical Causal Transformer for Real-Time and Infinite-Length 3D Human Motion Generation · AAAI 2026 |
Methods — techniques the papers use, named apart from their topics
residual vector quantization · 2.0chain-of-thought · 2.0causal transformer · 2.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MOGO: Residual Quantized Hierarchical Causal Transformer for Real-Time and Infinite-Length 3D Human Motion GenerationabstractRecent advances in transformer-based text-to-motion generation have significantly improved motion quality. However, achieving both real-time performance and long-horizon scalability remains an open challenge. In this paper, we present MOGO (Motion Generation with One-pass), a novel autoregressive framework for efficient and scalable 3D human motion generation. MOGO consists of two key components. First, we introduce MoSA-VQ, a motion scale-adaptive residual vector quantization module that hierarchically discretizes motion sequences through learnable scaling parameters, enabling dynamic allocation of representation capacity and producing compact yet expressive multi-level representations. Second, we design the RQHC-Transformer, a residual quantized hierarchical causal transformer that decodes motion tokens in a single forward pass. Each transformer block aligns with one quantization level, allowing hierarchical abstraction and temporally coherent generation with strong semantic flow. Compared to diffusion- and LLM-based approaches, MOGO achieves lower inference latency while preserving high motion fidelity. Moreover, its hierarchical latent design enables seamless and controllable infinite-length motion generation, with stable transitions and the ability to adaptively incorporate updated control signals at arbitrary points in time. To further enhance generalization and interpretability, we introduce Textual Condition Alignment (TCA), which leverages large language models with Chain-of-Thought reasoning to bridge the gap between real-world prompts and training data. TCA not only improves zero-shot performance on unseen datasets but also enriches motion comprehension for in-distribution prompts through explicit intent decomposition. Extensive experiments on HumanML3D, KIT-ML, and the unseen CMP dataset demonstrate that MOGO outperforms prior methods in generation quality, inference efficiency, and temporal scalability. Tengjiao Sun, Pengcheng Fang, Xiaohao Cai, Hansung Kim 0001 |
AAAI | 2 |
| 2025 | UniTMGE: Uniform Text-Motion Generation and Editing Model via DiffusionabstractCurrent methods have shown promising results in applying diffusion models to motion generation given text input. However, these methods are limited to unimodal inputs and outputs, restricted to motion generation alone, and lacking multimodal control capabilities. To address these issues, we introduce UniTMGE, a text-motion multimodal generation and editing framework based on diffusion. UniTMGE overcomes single-modality limitations, enabling exceptional performance and strong generalization across multiple tasks like text-driven motion generation, motion captioning, motion completion, and multimodal motion editing. UniTMGE comprises three components: UTMV for mapping text and motion into a shared latent space using contrastive learning, a controllable diffusion model customized for the UTMV space, and MCRE for unifying multimodal conditions into CLIP representations, enabling precise multimodal control and flexible motion editing through simple linear operations. We conducted both closed-world experiments and open-world experiments using the Motion-X dataset with detailed text descriptions, with results demonstrating our model's effectiveness and generalizability across multiple tasks. Yangfan He, Tengjiao Sun |
WACV | 3 |