Wen Wang 0015

dblp:29/4680-15 · DBLP profile ↗
← Back
16ranked-venue papers
3as first author
16since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 2 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 6 since 2021Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 Beyond Hard Masks: Progressive Token Evolution for Diffusion Language Models
abstract
Linhao Zhong, Linyu Wu, Bozhen Fang, Tianjian Feng, Chenchen Jing, Wen Wang, Jiaheng Zhang, Hao Chen, Chunhua Shen. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Linhao Zhong 0001, Linyu Wu, Bozhen Fang, Tianjian Feng, Chenchen Jing, Wen Wang 0015, Jiaheng Zhang, Hao Chen 0041, Chunhua Shen
ACL (1)6
2026 Efficient Self-Evaluation for Diffusion Language Models via Sequence Regeneration
abstract
Linhao Zhong, Linyu Wu, Wen Wang, Yuling Xi, Chenchen Jing, Jiaheng Zhang, Hao Chen, Chunhua Shen. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Linhao Zhong 0001, Linyu Wu, Wen Wang 0015, Yuling Xi, Chenchen Jing, Jiaheng Zhang, Hao Chen 0041, Chunhua Shen
ACL (1)3
2026 FreerCustom: Training-Free Multi-Concept Customization for Image and Video Generation
Canyu Zhao, Ganggui Ding, Wen Wang 0015, Zhen Yang 0009, Zide Liu, Hao Chen 0041, Chunhua Shen
Int. J. Comput. Vis.3
2025 MovieBench: A Hierarchical Movie Level Dataset for Long Video Generation
abstract
Recent advancements in video generation models, like Stable Video Diffusion, show promising results, but primarily focus on short, single-scene videos. These models struggle with generating long videos that involve multiple scenes, coherent narratives, and consistent characters. Furthermore, there is no publicly available dataset tailored for the analysis, evaluation, and training of long video generation models. In this paper, we present MovieBench: A Hierarchical Movie-Level Dataset for Long Video Generation, which addresses these challenges by providing unique contributions: (1) movie-length videos featuring rich, coherent storylines and multi-scene narratives, (2) consistency of character appearance and audio across scenes, and (3) hierarchical data structure contains high-level movie information and detailed shot-level descriptions. Experiments demonstrate that MovieBench brings some new insights and challenges, such as maintaining character ID consistency across multiple scenes for various characters. The dataset will be public and continuously maintained, aiming to advance the field of long video generation. Data can be found at: MovieBench.
Weijia Wu 0001, Xi Xia, Haoen Feng, Wen Wang 0015, Qinghong Lin, Chunhua Shen, Zheng Shou 0001
CVPR6
2025 Framer: Interactive Frame Interpolation
abstract
We propose Framer for interactive frame interpolation, which targets producing smoothly transitioning frames between two images as per user creativity. Concretely, besides taking the start and end frames as inputs, our approach supports customizing the transition process by tailoring the trajectory of some selected keypoints. Such a design enjoys two clear benefits. First, incorporating human interaction mitigates the issue arising from numerous possibilities of transforming one image to another, and in turn enables finer control of local motions. Second, as the most basic form of interaction, keypoints help establish the correspondence across frames, enhancing the model to handle challenging cases (e.g., objects on the start and end frames are of different shapes and styles). It is noteworthy that our system also offers an "autopilot" mode, where we introduce a module to estimate the keypoints and refine the trajectory automatically, to simplify the usage in practice. Extensive experimental results demonstrate the appealing performance of Framer on various applications, such as image morphing, time-lapse video generation, cartoon interpolation, etc. The code, model, and interface are publicly accessible at https://github.com/aim-uofa/Framer.
Wen Wang 0015, Qiuyu Wang, Kecheng Zheng, Hao Ouyang, Zhekai Chen, Biao Gong, Hao Chen 0041, Yujun Shen, Chunhua Shen
ICLR1
2025 MovieDreamer: Hierarchical Generation for Coherent Long Visual Sequences
abstract
Recent advancements in video generation have primarily leveraged diffusion models for short-duration content. However, these approaches often fall short in modeling complex narratives and maintaining character consistency over extended periods, which is essential for long-form video production like movies. We propose MovieDreamer, a novel hierarchical framework that integrates the strengths of autoregressive models with diffusion-based rendering to pioneer long-duration video generation with intricate plot progressions and high visual fidelity. Our approach utilizes autoregressive models for global narrative coherence, predicting sequences of visual tokens that are subsequently transformed into high-quality video frames through diffusion rendering. This method is akin to traditional movie production processes, where complex stories are factorized down into manageable scene capturing. Further, we employ a multimodal script that enriches scene descriptions with detailed character information and visual style, enhancing continuity and character identity across scenes. We present extensive experiments across various movie genres, demonstrating that our approach not only achieves superior visual and narrative quality but also effectively extends the duration of generated content significantly beyond current capabilities.
Canyu Zhao, Wen Wang 0015, Fan Wang 0019, Hao Chen 0041, Bo Zhang 0025, Chunhua Shen
ICLR3
2025 Efficient Multiple-Precision Floating-Point Multiply-Add Architecture for Deep Learning Applications
abstract
To fully exploit the potential of parallel computing while minimizing hardware costs, floating-point multiply-add (FMA) units that support multiple precisions are widely used in deep learning. However, achieving a balance between accuracy, performance, and hardware overhead remains a significant challenge. This paper presents a unified multiple-precision floating-point FMA architecture that supports four precision formats: single-precision (SP), Bfloat16 (BF16), TensorFloat32+ (TF32+), and INT8. The proposed architecture allows for the parallel execution of nine BF16 FMA operations, two TF32+ FMA operations, one SP FMA operation, or nine INT8 multiply-add (MA) operations. Through careful data format selection and an optimized architectural design, the architecture achieves multiplier utilization rates of 100%, 88.9%, 100%, and 100% for the four precision modes, respectively, with all multipliers operating at full bit width. Compared to state-of-the-art multiple-precision FMA designs, this architecture delivers over nine times the BF16 throughput while increasing the area by only 49%. The flexible data format configuration makes the proposed architecture suitable for a wide range of deep learning applications.
Songtai Liang, Bingjie Xia, Wen Wang 0015, Peng Liu 0016
ISCAS3
2025 Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration
abstract
Long-horizon video-audio reasoning and fine-grained pixel understanding impose conflicting requirements on omnimodal models: dense temporal coverage demands many low-resolution frames, whereas precise grounding calls for high-resolution inputs. We tackle this trade-off with a two-system architecture: a Global Reasoning System selects informative keyframes and rewrites the task at low spatial cost, while a Detail Understanding System performs pixel-level grounding on the selected high-resolution snippets. Because "optimal" keyframe selection and reformulation are ambiguous and hard to supervise, we formulate them as a reinforcement-learning (RL) problem and present Omni-R1, an end-to-end RL framework built on Group Relative Policy Optimization. Omni-R1 trains the Global Reasoning System through hierarchical rewards obtained via online collaboration with the Detail Understanding System, requiring only one epoch of RL on small task splits. Experiments on two challenging benchmarks, Referring Audio-Visual Segmentation (RefAVS) and Reasoning Video Object Segmentation (REVOS), show that Omni-R1 not only surpasses strong supervised baselines but also outperforms specialized state-of-the-art models, while substantially improving out-of-domain generalization and mitigating multimodal hallucination. Our results demonstrate the first successful application of RL to large-scale omnimodal reasoning and highlight a scalable path toward universally foundation models.
Muzhi Zhu, Zongze Du, Canyu Zhao, Wen Wang 0015, Hao Chen 0041, Chunhua Shen
NeurIPS7
2025 AutoStory: Generating Diverse Storytelling Images with Minimal Human Efforts
Wen Wang 0015, Canyu Zhao, Hao Chen 0041, Zhekai Chen, Kecheng Zheng, Chunhua Shen
Int. J. Comput. Vis.1
2024 FreeCustom: Tuning-Free Customized Image Generation for Multi-Concept Composition
abstract
Benefiting from large-scale pre-trained text-to-image (T2I) generative models, impressive progress has been achieved in customized image generation, which aims to generate user-specified concepts. Existing approaches have extensively focused on single-concept customization and still encounter challenges when it comes to complex scenarios that involve combining multiple concepts. These approaches often require retraining/fine-tuning using a few images, leading to time-consuming training processes and impeding their swift implementation. Furthermore, the reliance on multiple images to represent a singular concept increases the difficulty of customization. To this end, we propose FreeCustom, a novel tuning-free method to generate customized images of multi-concept composition based on reference concepts, using only one image per concept as input. Specifically, we introduce a new multi-reference self-attention (MRSA) mechanism and a weighted mask strategy that enables the generated image to access and focus more on the reference concepts. In addition, MRSA leverages our key finding that input concepts are better preserved when providing images with context interactions. Experiments show that our method's produced images are consistent with the given concepts and better aligned with the input text. Our method outperforms or performs on par with other training-based methods in terms of multi-concept composition and single-concept customization, but is simpler. Codes can be found here.
Ganggui Ding, Canyu Zhao, Wen Wang 0015, Zhen Yang 0009, Zide Liu, Hao Chen 0041, Chunhua Shen
CVPR3
2024 FreeCompose: Generic Zero-Shot Image Composition with Diffusion Prior
Zhekai Chen, Wen Wang 0015, Zhen Yang 0009, Zeqing Yuan, Hao Chen 0041, Chunhua Shen
ECCV (17)2
2024 Object-Aware Inversion and Reassembly for Image Editing
abstract
Diffusion-based image editing methods have achieved remarkable advances in text-driven image editing. The editing task aims to convert an input image with the original text prompt into the desired image that is well-aligned with the target text prompt. By comparing the original and target prompts, we can obtain numerous editing pairs, each comprising an object and its corresponding editing target. To allow editability while maintaining fidelity to the input image, existing editing methods typically involve a fixed number of inversion steps that project the whole input image to its noisier latent representation, followed by a denoising process guided by the target prompt. However, we find that the optimal number of inversion steps for achieving ideal editing results varies significantly among different editing pairs, owing to varying editing difficulties. Therefore, the current literature, which relies on a fixed number of inversion steps, produces sub-optimal generation quality, especially when handling multiple editing pairs in a natural image. To this end, we propose a new image editing paradigm, dubbed Object-aware Inversion and Reassembly (OIR), to enable object-level fine-grained editing. Specifically, we design a new search metric, which determines the optimal inversion steps for each editing pair, by jointly considering the editability of the target and the fidelity of the non-editing region. We use our search metric to find the optimal inversion step for each editing pair when editing an image. We then edit these editing pairs separately to avoid \concept. Subsequently, we propose an additional reassembly step to seamlessly integrate the respective editing results and the non-editing region to obtain the final edited image. To systematically evaluate the effectiveness of our method, we collect two datasets called OIRBench for benchmarking single- and multi-object editing, respectively. Experiments demonstrate that our method achieves superior performance in editing object shapes, colors, materials, categories, \textit{etc.}, especially in multi-object editing scenarios. The project page can be found in https://aim-uofa.github.io/OIR-Diffusion/.
Zhen Yang 0009, Ganggui Ding, Wen Wang 0015, Hao Chen 0041, Bohan Zhuang, Chunhua Shen
ICLR3
2024 Mantissa-Aware Floating-Point Eight-Term Fused Dot Product Unit
abstract
Floating-point dot product is widely used in various applications. A conventional discrete construction of multipliers and adders often leads to accumulated errors and lower speed. This paper presents a floating-point eight-term fused dot product unit with mantissa-aware hardware design to attack these problems. In the proposed design, multiplication and addition of numbers are operated in a fused manner with exception controller. A pre-shift and post-shift combined scheme for mantissa alignment is utilized to eliminate the latency between significand multiplication and mantissa alignment. A one-path mantissa compress and addition structure is employed that can effectively reduce the footprint. An impact of internal mantissa datapath width on calculation accuracy is analyzed. Compared to the discrete method, the proposed design delivers a significant reduction up to 45.7%, 26.9%, and 33.0%, in terms of latency, area, and power, respectively.
Wen Wang 0015, Bingjie Xia, Peng Liu 0016
ISCAS1
2023 EVA: Exploring the Limits of Masked Visual Representation Learning at Scale
abstract
We launch EVA, a vision-centric foundation model to Explore the limits of Visual representation at scAle using only publicly accessible data. EVA is a vanilla ViT pre-trained to reconstruct the masked out image-text aligned vision features conditioned on visible image patches. Via this pretext task, we can efficiently scale up EVA to one billion parameters, and sets new records on a broad range of representative vision downstream tasks, such as image recognition, video action recognition, object detection, instance segmentation and semantic segmentation without heavy supervised training. Moreover, we observe quantitative changes in scaling EVA result in qualitative changes in transfer learning performance that are not present in other models. For instance, EVA takes a great leap in the challenging large vocabulary instance segmentation task: our model achieves almost the same state-of-the-art performance on LVIS dataset with over a thousand categories and COCO dataset with only eighty categories. Beyond a pure vision encoder, EVA can also serve as a vision-centric, multi-modal pivot to connect images and text. We find initializing the vision tower of a giant CLIP from EVA can greatly stabilize the training and outperform the training from scratch counterpart with much fewer samples and less compute, providing a new direction for scaling up and accelerating the costly training of multi-modal foundation models.
Wen Wang 0015, Binhui Xie, Ledell Wu, Xinggang Wang, Tiejun Huang 0001, Yue Cao 0001
CVPR2
2023 Images Speak in Images: A Generalist Painter for In-Context Visual Learning
abstract
In-context learning, as a new paradigm in NLP, allows the model to rapidly adapt to various tasks with only a handful of prompts and examples. But in computer vision, the difficulties for in-context learning lie in that tasks vary significantly in the output representations, thus it is unclear how to define the general-purpose task prompts that the vision model can understand and transfer to out-of-domain tasks. In this work, we present Painter, a generalist model which addresses these obstacles with an “image”-centric solution, that is, to redefine the output of core vision tasks as images, and specify task prompts as also images. With this idea, our training process is extremely simple, which performs standard masked image modeling on the stitch of input and output image pairs. This makes the model capable of performing tasks conditioned on visible image patches. Thus, during inference, we can adopt a pair of input and output images from the same task as the input condition, to indicate which task to perform. Without bells and whistles, our generalist Painter can achieve competitive performance compared to well-established task-specific models, on seven representative vision tasks ranging from high-level visual understanding to low-level image processing. In addition, Painter significantly outperforms recent generalist models on several challenging tasks.
Wen Wang 0015, Yue Cao 0001, Chunhua Shen, Tiejun Huang 0001
CVPR2
2023 SegGPT: Towards Segmenting Everything In Context
abstract
We present SegGPT, a generalist model for segmenting everything in context. We unify various segmentation tasks into a generalist in-context learning framework that accommodates different kinds of segmentation data by transforming them into the same format of images. The training of SegGPT is formulated as an in-context coloring problem with random color mapping for each data sample. The objective is to accomplish diverse tasks according to the context, rather than relying on specific colors. After training, SegGPT can perform arbitrary segmentation tasks in images or videos via in-context inference, such as object instance, stuff, part, contour, and text. SegGPT is evaluated on a broad range of tasks, including few-shot semantic segmentation, video object segmentation, semantic segmentation, and panoptic segmentation. Our results show strong capabilities in segmenting in-domain and out-of-domain targets, either qualitatively or quantitatively.
Yue Cao 0001, Wen Wang 0015, Chunhua Shen, Tiejun Huang 0001
ICCV4