EDBT 2026 Demo / reviewers in the wild / expert
Haoning Wu 0002
dblp:264/5802-2
· DBLP profile ↗
10ranked-venue papers
3as first author
10since 2021 · last 2026
0009-0001-8717-338XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 9 · 3 first-author · 9 since 2021Artificial intelligence and machine learning · 6 · 2 first-author · 6 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SceneGen: Single-Image 3D Scene Generation in One Feedforward Passabstract3D content generation has recently attracted significant research interest, driven by its critical applications in VR/AR and embodied AI. In this work, we tackle the challenging task of synthesizing multiple 3D assets within a single scene image. Concretely, our contributions are fourfold: (i) we present SceneGen, a novel framework that takes a scene image and corresponding object masks as input, simultaneously producing multiple 3D assets with geometry and texture. Notably, SceneGen operates with no need for extra optimization or asset retrieval; (ii) we introduce a novel feature aggregation module that integrates local and global scene information from visual and geometric encoders within the feature extraction module. Coupled with a position head, this enables the generation of 3D assets and their relative spatial positions in a single feedforward pass; (iii) we demonstrate SceneGen's direct extensibility to multi-image input scenarios. Despite being trained solely on single-image inputs, our architecture yields improved generation performance when multiple images are provided; and (iv) extensive quantitative and qualitative evaluations confirm the efficiency and robustness of our approach. We believe this paradigm offers a novel solution for high-quality 3D content generation, potentially advancing its practical applications in downstream tasks. The code and model will be publicly available at: https://mengmouxu.github.io/SceneGen. “Everything you can imagine is real.” —Pablo Picasso Yanxu Meng, Haoning Wu 0002, Ya Zhang 0002, Weidi Xie |
3DV | 2 |
| 2025 | Towards Universal Soccer Video UnderstandingabstractAs a globally celebrated sport, soccer has attracted widespread interest from fans all over the world. This paper aims to develop a comprehensive multi-modal framework for soccer video understanding. Specifically, we make the following contributions in this paper: (i) we introduce SoccerReplay-1988, the largest multi-modal soccer dataset to date, featuring videos and detailed annotations from 1,988 complete matches, with an automated annotation pipeline; (ii) we present an advanced soccer-specific visual encoder, MatchVision, which leverages spatiotemporal information across soccer videos and excels in various downstream tasks; (iii) we conduct extensive experiments and ablation studies on event classification, commentary generation, and multi-view foul recognition. MatchVision demonstrates state-of-the-art performance on all of them, substantially outperforming existing models, which highlights the superiority of our proposed data and model. We believe that this work will offer a standard paradigm for sports understanding research.“Football is one of the world’s best means of communication. It is impartial, apolitical, and universal.”—— Franz Beckenbauer (1945 - 2024) Jiayuan Rao, Haoning Wu 0002, Ya Zhang 0002, Yanfeng Wang 0001, Weidi Xie |
CVPR | 2 |
| 2025 | MRGen: Segmentation Data Engine for Underrepresented MRI Modalities
Haoning Wu 0002, Ya Zhang 0002, Yanfeng Wang 0001, Weidi Xie |
ICCV | 1 |
| 2025 | Multi-Agent System for Comprehensive Soccer Understanding
Jiayuan Rao, Haoning Wu 0002, Ya Zhang 0002, Yanfeng Wang 0001, Weidi Xie |
ACM Multimedia | 3 |
| 2025 | MegaFusion: Extend Diffusion Models towards Higher-resolution Image Generation without Further TuningabstractDiffusion models have emerged as frontrunners in text-to-image generation, but their fixed image resolution during training often leads to challenges in high-resolution image generation, such as semantic deviations and object replication. This paper introduces MegaFusion, a novel approach that extends existing diffusion-based text-to-image models towards efficient higher-resolution generation without additional fine-tuning or adaptation. Specifically, we employ an innovative truncate and relay strategy to bridge the denoising processes across different resolutions, allowing for high-resolution image generation in a coarse-to-fine manner. Moreover, by integrating dilated convolutions and noise re-scheduling, we further adapt the model's priors for higher resolution. The versatility and efficacy of MegaFusion make it universally applicable to both latent-space and pixel-space diffusion models, along with other derivative models. Extensive experiments confirm that MegaFusion significantly boosts the capability of existing models to pro-duce images of megapixels and various aspect ratios, while only requiring about 40% of the original computational cost. Code is available at https://haoningwu3639.github.io/MegaFusion/. Haoning Wu 0002, Shaocheng Shen, Qiang Hu 0003, Xiaoyun Zhang 0001, Ya Zhang 0002, Yanfeng Wang 0001 |
WACV | 1 |
| 2024 | Intelligent Grimm - Open-ended Visual Storytelling via Latent Diffusion ModelsabstractGenerative models have recently exhibited exceptional capabilities in text-to-image generation, but still struggle to generate image sequences coherently. In this work, we focus on a novel, yet challenging task of generating a co-herent image sequence based on a given storyline, denoted as open-ended visual storytelling. We make the following three contributions: (i) to fulfill the task of visual sto-rytelling, we propose a learning-based auto-regressive im-age generation model, termed as Story Gen, with a novel vision-language context module, that enables to generate the current frame by conditioning on the corresponding text prompt and preceding image-caption pairs; (ii) to ad-dress the data shortage of visual storytelling, we collect paired image-text sequences by sourcing from online videos and open-source E-books, establishing processing pipeline for constructing a large-scale dataset with diverse characters, storylines, and artistic styles, named StorySalon; (iii) Quantitative experiments and human evaluations have vali-dated the superiority of our StoryGen, where we show it can generalize to unseen characters without any optimization, and generate image sequences with coherent content and consistent character. Code, dataset, and models are avail-able at https://haoningwu3639.github.io/StoryGen_Webpage/. “Mirror mirror on the wall, who's the fairest of them all?” -Grimms' Fairy Tales Chang Liu 0079, Haoning Wu 0002, Xiaoyun Zhang 0001, Yanfeng Wang 0001, Weidi Xie |
CVPR | 2 |
| 2024 | MatchTime: Towards Automatic Soccer Game Commentary GenerationabstractSoccer is a globally popular sport with a vast audience, in this paper, we consider constructing an automatic soccer game commentary model to improve the audiences’ viewing experience. In general, we make the following contributions: First, observing the prevalent video-text misalignment in existing datasets, we manually annotate timestamps for 49 matches, establishing a more robust benchmark for soccer game commentary generation, termed as SN-Caption-test-align; Second, we propose a multi-modal temporal alignment pipeline to automatically correct and filter the existing dataset at scale, creating a higher-quality soccer game commentary dataset for training, denoted as MatchTime; Third, based on our curated dataset, we train an automatic commentary generation model, named MatchVoice. Extensive experiments and ablation studies have demonstrated the effectiveness of our alignment pipeline, and training model on the curated datasets achieves state-of-the-art performance for commentary generation, showcasing that better alignment can lead to significant performance improvements in downstream tasks. Jiayuan Rao, Haoning Wu 0002, Chang Liu 0079, Yanfeng Wang 0001, Weidi Xie |
EMNLP | 2 |
| 2023 | Boost Video Frame Interpolation via Motion Adaptation
Haoning Wu 0002, Xiaoyun Zhang 0001, Weidi Xie, Ya Zhang 0002, Yanfeng Wang 0001 |
BMVC | 1 |
| 2023 | NeRF-SDP: Efficient Generalizable Neural Radiance Field with Scene Depth PerceptionabstractIn recent years, neural radiance fields have exhibited impressive performance in novel view synthesis. However, exploiting complex network structures to achieve generalizable NeRF usually results in inefficient rendering. Existing methods for accelerating rendering directly employ simpler inference networks or fewer sampling points, leading to unsatisfactory synthesis quality. To address the challenge of balancing rendering speed and quality in generalizable NeRF, we propose a novel framework, NeRF-SDP, which achieves both efficiency and high fidelity by introducing scene depth perception. We incorporate more scene information into the radiance field by using our proposed geometry feature extraction and depth-encoded ray transformer to improve the model’s inference capabilities with sparse points. With the aid of scene depth perception, NeRF-SDP can better understand the scene’s structure, thus better reconstructing the objects’ edges with significantly fewer artifacts. Experimental results demonstrate that NeRF-SDP achieves comparable synthesis quality to state-of-the-art methods while significantly improving rendering efficiency. Furthermore, ablation studies confirm that the depth-encoded ray transformer enhances the model’s robustness to varying numbers of sampling points. Qiuwen Wang, Shuai Guo 0002, Haoning Wu 0002, Rong Xie 0004, Li Song 0001, Wenjun Zhang 0001 |
MMAsia | 3 |
| 2022 | LAR-SR: A Local Autoregressive Model for Image Super-ResolutionabstractPrevious super-resolution (SR) approaches often formulate SR as a regression problem and pixel wise restoration, which leads to a blurry and unreal SR output. Recent works combine adversarial loss with pixel-wise loss to train a GAN-based model or introduce normalizing flows into SR problems to generate more realistic images. As another powerful generative approach, autoregressive (AR) model has not been noticed in low level tasks due to its limitation. Based on the fact that given the structural in-formation, the textural details in the natural images are locally related without long term dependency, in this paper we propose a novel autoregressive model-based SR approach, namely LAR-SR, which can efficiently generate realistic SR images using a novel local autoregressive (LAR) module. The proposed LAR module can sample all the patches of textural components in parallel, which greatly reduces the time consumption. In addition to high time efficiency, it is also able to leverage contextual information of pixels and can be optimized with a consistent loss. Experimental results on the widely-used datasets show that the proposed LAR-SR approach achieves superior performance on the vi-sual quality and quantitative metrics compared with other generative models such as GAN, Flow, and is competitive with the mixture generative model. Baisong Guo, Xiaoyun Zhang 0001, Haoning Wu 0002, Yu Wang 0027, Ya Zhang 0002, Yanfeng Wang 0001 |
CVPR | 3 |