EDBT 2026 Demo / reviewers in the wild / expert
Yunyang Ge
dblp:389/2589
· DBLP profile ↗
4ranked-venue papers
0as first author
4since 2021 · last 2026
0000-0002-9525-9079ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
4 papers |
Generative modeling · 94% 3D vision · 6% | |
| Computer graphics and multimedia
4 papers |
Visual content generation and editing · 63% Virtual and augmented reality · 30% Image and video processing · 6% |
Topics — the 10 heaviest of 12, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Generative modeling
diffusion model |
1.7 | 2 | 2025 | Identity-Preserving Text-to-Video Generation by Frequency Decomposition · CVPR 2025 RoomPainter: View-Integrated Diffusion for Consistent Indoor Scene Texturing · CVPR 2025 |
Virtual and augmented reality › immersive video
360-degree video |
1.0 | 1 | 2026 | 360Explorer: Exploring 4D Controllable World in Panoramic Videos · AAAI 2026 |
Visual content generation and editing › video generation
controllable video generation |
1.0 | 1 | 2026 | 360Explorer: Exploring 4D Controllable World in Panoramic Videos · AAAI 2026 |
Machine learning › Generative modeling › diffusion model
diffusion transformer |
0.9 | 1 | 2025 | Identity-Preserving Text-to-Video Generation by Frequency Decomposition · CVPR 2025 |
Machine learning › Generative modeling › conditional generative model › controlled generation
identity-preserving generation |
0.9 | 1 | 2025 | Identity-Preserving Text-to-Video Generation by Frequency Decomposition · CVPR 2025 |
Machine learning › Generative modeling › video generation
text-to-video generation |
0.9 | 1 | 2025 | Identity-Preserving Text-to-Video Generation by Frequency Decomposition · CVPR 2025 |
Visual content generation and editing
texture synthesis |
0.9 | 1 | 2025 | RoomPainter: View-Integrated Diffusion for Consistent Indoor Scene Texturing · CVPR 2025 |
Machine learning › Generative modeling › diffusion model
text-to-image generation |
0.8 | 1 | 2024 | Prompt2Poster: Automatically Artistic Chinese Poster Creation from Prompt Only · ACM Multimedia 2024 |
Computer vision › 3D vision › 3d scene modeling › scene representation
dynamic scene representation |
0.3 | 1 | 2026 | 360Explorer: Exploring 4D Controllable World in Panoramic Videos · AAAI 2026 |
Image and video processing › image transform
frequency decomposition |
0.3 | 1 | 2025 | Identity-Preserving Text-to-Video Generation by Frequency Decomposition · CVPR 2025 |
Methods — techniques the papers use, named apart from their topics
reverse warping · 2.04d scene representation · 2.0zero-shot diffusion adaptation · 1.7tuning-free pipeline · 1.7neighbor-integrated attention · 1.7frequency-aware heuristic control · 1.7attention-guided multi-view integrated sampling · 1.7large language model · 1.5controllable layout generation · 1.5
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | 360Explorer: Exploring 4D Controllable World in Panoramic VideosabstractWe present 360Explorer, a novel approach for generating 4D controllable panoramic videos conditioned on user-provided 3D instructions for exploring and manipulating dynamic worlds. Compared to existing perspective-based methods struggle to address spatial consistency during camera rotation in place, we introduce the panoramic view in controllable video generation models to inherently maintain the view recall consistency. By introducing dynamic point clouds as the 4D scene representations, 360Explorer unifies the modeling of camera transformations and object movements as incomplete renders to describe precise control instructions in 3D worlds. To tackle the data limitation in acquiring multi-viewpoint panoramic videos, we further propose a reverse warping strategy to construct the training dataset on easily accessible monocular panoramic videos. Extensive experiments demonstrate that 360Explorer achieves superior performance in creating 4D controllable panoramic videos with camera transformation and object movements aligned with diverse provided instructions. Xinhua Cheng, Haiyang Zhou, Wangbo Yu, Tanghui Jia, Bin Lin 0014, Yunyang Ge, Li Yuan 0007 |
AAAI | 6 |
| 2025 | RoomPainter: View-Integrated Diffusion for Consistent Indoor Scene TexturingabstractIndoor scene texture synthesis has garnered significant interest due to its important potential applications in virtual reality, digital media and creative arts. Existing diffusion-model-based researches either rely on per-view inpainting techniques, which are plagued by severe cross-view inconsistencies and conspicuous seams, or adopt optimization-based approaches that involve substantial computational overhead. In this work, we present RoomPainter, a frame-work that seamlessly integrates efficiency and consistency to achieve high-fidelity texturing of indoor scenes. The core of RoomPainter features a zero-shot technique that effectively adapts a 2D diffusion model for 3D-consistent texture synthesis, along with a two-stage generation strategy that ensures both global and local consistency. Specifically, we introduce Attention-Guided Multi-View Integrated Sampling (MVIS) combined with a neighbor-integrated attention mechanism for zero-shot texture map generation. Using the MVIS, we firstly generate texture map for the entire room to ensure global consistency, then adopt its variant, namely Attention-Guided Multi-View Integrated Repaint Sampling (MVRS) to repaint individual instances within the room, thereby further enhancing local consistency and addressing the occlusion problem. Experiments demonstrate that RoomPainter achieves superior performance for indoor scene texture synthesis in visual quality, global consistency and generation efficiency. Zhipeng Huang 0001, Wangbo Yu, Xinhua Cheng, ChengShu Zhao, Yunyang Ge, Mingyi Guo, Li Yuan 0007, Yonghong Tian 0001 |
CVPR | 5 |
| 2025 | Identity-Preserving Text-to-Video Generation by Frequency DecompositionabstractIdentity-Preserving text-to-video (IPT2V) generation aims to create high-fidelity videos with consistent human identity. It is an important task in video generation but remains an open problem for generative models. This paper pushes the technical frontier of IPT2V in two directions that have not been resolved in the literature: (1) A tuning-free pipeline without tedious case-by-case finetuning, and (2) A frequency-aware heuristic identity-preserving Diffusion Transformer (DiT)-based control scheme. To achieve these goals, we propose ConsisID, a tuning-free DiT-based controllable IPT2V model to keep human-identity consistent in the generated video. Inspired by prior findings in frequency analysis of vision/diffusion transformers, it employs identity-control signals base on frequency domain, since facial features can be decomposed into low-frequency global features (e.g., profile, proportions) and high-frequency intrinsic features (e.g., identity markers that remain unaffected by pose changes). Extensive experiments demonstrate that our frequency-aware heuristic scheme provides an optimal control solution for DiT-based models, making strides towards more effective IPT2V. Shenghai Yuan 0002, Jinfa Huang, Xianyi He, Yunyang Ge, Yujun Shi, Liuhan Chen, Jiebo Luo 0001, Li Yuan 0007 |
CVPR | 4 |
| 2024 | Prompt2Poster: Automatically Artistic Chinese Poster Creation from Prompt OnlyabstractAs a critical component in graphic design, artistic posters are widely applied in the advertising and entertainment industry, thus the automatic poster creation from user-provided prompts has become increasingly desired recently. Although existing Text2Image methods create impressive images aligned with given prompts, they fail to generate ideal artistic posters, especially with Chinese texts. To create desired artistic Chinese posters including an aligned background, reasonable layouts, and stylized graphical texts from given prompts only, we propose an automatic poster creation framework, named Prompt2Poster. Our framework utilizes the capacity of the powerful Large Language Model (LLM) to extract user intention from provided prompts and generate the aligned background. Although only taking a user prompt as the input, linguistic, visual, and geometrical information is fully utilized in the framework, bringing the ability to fit different distributions. To achieve the use of multi-modal information in the framework, two carefully designed modules, Controllable Layout Generator (CLG) and Graphical Text Generator (GTG) are proposed, leading to accurate and pleasurable visual results. Comprehensive experiments demonstrate that our Prompt2Poster achieves superior performance, especially in text quality and visual harmony. Yunyang Ge, Liuhan Chen, Haiyang Zhou, Qian Wang 0062, Xinhua Cheng, Li Yuan 0007 |
ACM Multimedia | 2 |