EDBT 2026 Demo / reviewers in the wild / expert
Shoufa Chen
dblp:187/4654
· DBLP profile ↗
18ranked-venue papers
8as first author
17since 2021 · last 2026
0000-0002-6126-2595ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 16 · 7 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 5 first-author · 8 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FlashVideo: Flowing Fidelity to Detail for Efficient High-Resolution Video GenerationabstractDiT models have achieved great success in text-to-video generation, leveraging their scalability in model capacity and data scale. High content and motion fidelity aligned with text prompts, however, often require large model parameters and a substantial number of function evaluations (NFEs). Realistic and visually appealing details are typically reflected in high-resolution outputs, further amplifying computational demands—especially for single-stage DiT models. To address these challenges, we propose a novel two-stage framework, FlashVideo, which strategically allocates model capacity and NFEs across stages to balance generation fidelity and quality. In the first stage, prompt fidelity is prioritized through a low-resolution generation process utilizing large parameters and sufficient NFEs to enhance computational efficiency. The second stage achieves a nearly straight ODE trajectory between low and high resolutions via flow matching, effectively generating fine details and fixing artifacts with minimal NFEs. To ensure a seamless connection between the two independently trained stages during inference, we carefully design degradation strategies during the second-stage training. Quantitative and visual results demonstrate that FlashVideo achieves state-of-the-art high-resolution video generation with superior computational efficiency. Additionally, the two-stage design enables users to preview the initial output and accordingly adjust the prompt before committing to full-resolution generation, thereby significantly reducing computational costs and wait times as well as enhancing commercial viability. Shoufa Chen, Chongjian Ge, Peize Sun, Yi Jiang 0009, Zehuan Yuan, Bingyue Peng, Ping Luo 0002 |
AAAI | 3 |
| 2025 | Goku: Flow Based Video Generative Foundation ModelsabstractThis paper introduces Goku, a state-of-the-art family of joint image-and-video generation models leveraging rectified flow Transformers to achieve industry-leading performance. We detail the foundational elements enabling high-quality visual generation, including the data curation pipeline, model architecture design, flow formulation, and advanced infrastructure for efficient and robust large-scale training. The Goku models demonstrate superior performance in both qualitative and quantitative evaluations, setting new benchmarks across major tasks. Specifically, Goku achieves 0.76 on GenEval and 83.65 on DPG-Bench for text-to-image generation, and 84.85 on VBench for text-to-video tasks. We believe that this work provides valuable insights and practical advancements for the research community in developing joint image-and-video generation models. Shoufa Chen, Chongjian Ge, Fengda Zhu, Hao Yang 0044, Hongxiang Hao, Zhichao Lai, Yifei Hu, Ting-Che Lin, Yanghua Peng, Peize Sun, Ping Luo 0002, Yi Jiang 0009, Zehuan Yuan, Bingyue Peng |
CVPR | 1 |
| 2025 | Prompt-A-Video: Prompt your Video Diffusion Model via Preference-Aligned LLMabstractText-to-video models have made remarkable advancements through optimization on high-quality text-video pairs, where the textual prompts play a pivotal role in determining quality of output videos. However, achieving the desired output often entails multiple revisions and iterative inference to refine user-provided prompts. Current automatic methods for refining prompts encounter challenges such as Modality-Inconsistency, Cost-Discrepancy, and Model-Unaware when applied to text-to-video diffusion models. To address these problem, we introduce an LLM-based prompt adaptation framework, termed as Prompt-A-Video, which excels in crafting Video-Centric, Labor-Free and Preference-Aligned prompts tailored to specific video diffusion model. Our approach involves a meticulously crafted two-stage optimization and alignment system. Initially, we conduct a reward-guided prompt evolution pipeline to automatically create optimal prompts pool and leverage them for supervised fine-tuning (SFT) of the LLM. Then multi-dimensional rewards are employed to generate pairwise data for the SFT model, followed by the direct preference optimization (DPO) algorithm to further facilitate preference alignment. Through extensive experimentation and comparative analyses, we validate the effectiveness of Prompt-A-Video across diverse generation models, highlighting its potential to push the boundaries of video generation. Yatai Ji, Jie Wu 0001, Shoufa Chen, Chongjian Ge, Peize Sun, Wenqi Shao, Xuefeng Xiao 0001, Ping Luo 0002 |
ICCV | 5 |
| 2025 | ControlAR: Controllable Image Generation with Autoregressive ModelsabstractAutoregressive (AR) models have reformulated image generation as next-token prediction, demonstrating remarkable potential and emerging as strong competitors to diffusion models. However, control-to-image generation, akin to ControlNet, remains largely unexplored within AR models. Although a natural approach, inspired by advancements in Large Language Models, is to tokenize control images into tokens and prefill them into the autoregressive model before decoding image tokens, it still falls short in generation quality compared to ControlNet and suffers from inefficiency. To this end, we introduce ControlAR, an efficient and effective framework for integrating spatial controls into autoregressive image generation models. Firstly, we explore control encoding for AR models and propose a lightweight control encoder to transform spatial inputs (e.g., canny edges or depth maps) into control tokens. Then ControlAR exploits the conditional decoding method to generate the next image token conditioned on the per-token fusion between control and image tokens, similar to positional encodings. Compared to prefilling tokens, using conditional decoding significantly strengthens the control capability of AR models but also maintains the model efficiency. Furthermore, the proposed ControlAR surprisingly empowers AR models with arbitrary-resolution image generation via conditional decoding and specific controls. Extensive experiments can demonstrate the controllability of the proposed ControlAR for the autoregressive control-to-image generation across diverse inputs, including edges, depths, and segmentation masks. Furthermore, both quantitative and qualitative results indicate that ControlAR surpasses previous state-of-the-art
controllable diffusion models, e.g., ControlNet++. Zongming Li, Tianheng Cheng, Shoufa Chen, Peize Sun, Haocheng Shen, Longjin Ran, Xiaoxin Chen 0001, Wenyu Liu 0001, Xinggang Wang |
ICLR | 3 |
| 2025 | WorldWeaver: Generating Long-Horizon Video Worlds via Rich PerceptionabstractGenerative video modeling has made significant strides, yet ensuring structural and temporal consistency over long sequences remains a challenge. Current methods predominantly rely on RGB signals, leading to accumulated errors in object structure and motion over extended durations. To address these issues, we introduce WorldWeaver, a robust framework for long video generation that jointly models RGB frames and perceptual conditions within a unified long-horizon modeling scheme. Our training framework offers three key advantages. First, by jointly predicting perceptual conditions and color information from a unified representation, it significantly enhances temporal consistency and motion dynamics. Second, by leveraging depth cues, which we observe to be more resistant to drift than RGB, we construct a memory bank that preserves clearer contextual information, improving quality in long-horizon video generation. Third, we employ segmented noise scheduling for training prediction groups, which further mitigates drift and reduces computational cost. Extensive experiments on both diffusion and rectified flow-based models demonstrate the effectiveness of WorldWeaver in reducing temporal drift and improving the fidelity of generated videos. Xueqing Deng, Shoufa Chen, Angtian Wang, Qiushan Guo, Mingfei Han 0003, Zeyue Xue, Mengzhao Chen, Ping Luo 0002 |
NeurIPS | 3 |
| 2024 | GenTron: Diffusion Transformers for Image and Video GenerationabstractIn this study, we explore Transformer-based diffusion models for image and video generation. Despite the dominance of Transformer architectures in various fields due to their flexibility and scalability, the visual generative domain primarily utilizes CNN-based U-Net architectures, particularly in diffusion-based models. We introduce GenTron, a family of Generative models employing Transformer-based diffusion, to address this gap. Our initial step was to adapt Diffusion Transformers (DiTs) from class to text conditioning, a process involving thorough empirical exploration of the conditioning mechanism. We then scale GenTron from approximately 900M to over 3B parameters, observing improvements in visual quality. Furthermore, we extend GenTron to text-to-video generation, incorporating novel motion-free guidance to enhance video quality. In human evaluations against SDXL, GenTron achieves a 51.1% win rate in visual quality (with a 19.8% draw rate), and a 42.3% win rate in text alignment (with a 42.9% draw rate). GenTron notably performs well in T2I-CompBench, highlighting its compositional generation ability. We hope GenTron could provide meaningful insights and serve as a valuable reference for future research. Please refer to the website11https://www.shoufachen.com/gentron_website/ and the arXiv version for the most up-to-date results: https://arxiv.org/abs/2312.04557. Shoufa Chen, Mengmeng Xu 0006, Jiawei Ren 0001, Yuren Cong, Sen He 0001, Yanping Xie, Animesh Sinha, Ping Luo 0002, Tao Xiang 0002, Juan-Manuel Pérez-Rúa |
CVPR | 1 |
| 2024 | FLATTEN: optical FLow-guided ATTENtion for consistent text-to-video editingabstractText-to-video editing aims to edit the visual appearance of a source video conditional on textual prompts.
A major challenge in this task is to ensure that all frames in the edited video are visually consistent.
Most recent works apply advanced text-to-image diffusion models to this task by inflating 2D spatial attention in the U-Net into spatio-temporal attention.
Although temporal context can be added through spatio-temporal attention, it may introduce some irrelevant information for each patch and therefore cause inconsistency in the edited video.
In this paper, for the first time, we introduce optical flow into the attention module in diffusion model's U-Net to address the inconsistency issue for text-to-video editing.
Our method, FLATTEN, enforces the patches on the same flow path across different frames to attend to each other in the attention module, thus improving the visual consistency in the edited videos.
Additionally, our method is training-free and can be seamlessly integrated into any diffusion based text-to-video editing methods and improve their visual consistency.
Experiment results on existing text-to-video editing benchmarks show that our proposed method achieves the new state-of-the-art performance. In particular, our method excels in maintaining the visual consistency in the edited videos. Yuren Cong, Mengmeng Xu 0006, Christian Simon, Shoufa Chen, Jiawei Ren 0001, Yanping Xie, Juan-Manuel Pérez-Rúa, Bodo Rosenhahn, Tao Xiang 0002, Sen He 0001 |
ICLR | 4 |
| 2024 | RoboCodeX: Multimodal Code Generation for Robotic Behavior SynthesisabstractRobotic behavior synthesis, the problem of understanding multimodal inputs and generating precise physical control for robots, is an important part of Embodied AI. Despite successes in applying multimodal large language models for high-level understanding, it remains challenging to translate these conceptual understandings into detailed robotic actions while achieving generalization across various scenarios. In this paper, we propose a tree-structured multimodal code generation framework for generalized robotic behavior synthesis, termed RoboCodeX. RoboCodeX decomposes high-level human instructions into multiple object-centric manipulation units consisting of physical preferences such as affordance and safety constraints, and applies code generation to introduce generalization ability across various robotics platforms. To further enhance the capability to map conceptual and perceptual understanding into control commands, a specialized multimodal reasoning dataset is collected for pre-training and an iterative self-updating methodology is introduced for supervised fine-tuning. Extensive experiments demonstrate that RoboCodeX achieves state-of-the-art performance in both simulators and real robots on four different kinds of manipulation tasks and one embodied navigation task. Yao Mu 0001, Shoufa Chen, Qiaojun Yu, Chongjian Ge, Runjian Chen, Zhixuan Liang, Mengkang Hu, Chaofan Tao, Peize Sun, Haibao Yu, Chao Yang 0026, Wenqi Shao, Wenhai Wang, Jifeng Dai, Yu Qiao 0001, Mingyu Ding, Ping Luo 0002 |
ICML | 4 |
| 2023 | DiffusionDet: Diffusion Model for Object DetectionabstractWe propose DiffusionDet, a new framework that formulates object detection as a denoising diffusion process from noisy boxes to object boxes. During the training stage, object boxes diffuse from ground-truth boxes to random distribution, and the model learns to reverse this noising process. In inference, the model refines a set of randomly generated boxes to the output results in a progressive way. Our work possesses an appealing property of flexibility, which enables the dynamic number of boxes and iterative evaluation. The extensive experiments on the standard benchmarks show that DiffusionDet achieves favorable performance compared to previous well-established detectors. For example, DiffusionDet achieves 5.3 AP and 4.8 AP gains when evaluated with more boxes and iteration steps, under a zero-shot transfer setting from COCO to CrowdHuman. Our code is available at https://github.com/ShoufaChen/DiffusionDet. Shoufa Chen, Peize Sun, Yibing Song, Ping Luo 0002 |
ICCV | 1 |
| 2023 | Going Denser with Open-Vocabulary Part SegmentationabstractObject detection has been expanded from a limited number of categories to open vocabulary. Moving forward, a complete intelligent vision system requires understanding more fine-grained object descriptions, object parts. In this paper, we propose a detector with the ability to predict both open-vocabulary objects and their part segmentation. This ability comes from two designs. First, we train the detector on the joint of part-level, object-level and image-level data to build the multi-granularity alignment between language and image. Second, we parse the novel object into its parts by its dense semantic correspondence with the base object. These two designs enable the detector to largely benefit from various data sources and foundation models. In open-vocabulary part segmentation experiments, our method outperforms the baseline by 3.3∼7.3 mAP in cross-dataset generalization on PartImageNet, and improves the baseline by 7.3 novel AP50in cross-category generalization on Pascal Part. Finally, we train a detector that generalizes to a wide range of part segmentation datasets while achieving better performance than dataset-specific training. Peize Sun, Shoufa Chen, Chenchen Zhu, Fanyi Xiao, Ping Luo 0002, Saining Xie, Zhicheng Yan 0001 |
ICCV | 2 |
| 2023 | Soft Neighbors are Positive Supporters in Contrastive Visual Representation Learning
Chongjian Ge, Jiangliu Wang, Zhan Tong, Shoufa Chen, Yibing Song, Ping Luo 0002 |
ICLR | 4 |
| 2023 | Image parallel block compressive sensing scheme using DFT measurement matrixabstractAbstract Compressive sensing (CS)-based image coding has been widely studied in the field of image processing. However, the CS-based image encoder has a significant gap in image reconstruction performance compared with the conventional image compression methods. In order to improve the reconstruction quality of CS-based image encoder, we proposed an image parallel block compressive sensing (BCS) coding scheme, which is based on discrete Cosine transform (DCT) sparse basis matrix and partial discrete Fourier transform (DFT) measurement matrix. In the proposed parallel BCS scheme, each column of an image block is sampled by the same DFT measurement matrix. Due to the complex property of DFT measurement matrix, the compressed image data is complex. Then, the real part and imaginary part of the resulting BCS data are quantized and transformed into two bit streams, respectively. At the reconstruction stage, the resulting two bit streams are transformed back into two real signals using inverse quantization operation. The resulting two real signals are combined into one complex signal, which is served as the input data of the CS reconstructed algorithm. The theoretical analysis based on minimum Frobenius norm method demonstrates that the proposed DFT measurement matrix outperforms the other conventional measurement matrices. The simulation results show that the reconstructed performance of the proposed DFT measurement matrix is better than that of the other conventional measurement matrices for the proposed parallel BCS. Specifically, we analyzed the impact of quantization on the reconstruction performance of CS. The experiment results show that the effect of the quantization on reconstruction performance in BCS framework can nearly be ignored. Zhongpeng Wang, Yannan Jiang, Shoufa Chen |
Multim. Tools Appl. | 3 |
| 2023 | CycleMLP: A MLP-Like Architecture for Dense Visual PredictionsabstractThis article presents a simple yet effective multilayer perceptron (MLP) architecture, namely CycleMLP, which is a versatile neural backbone network capable of solving various tasks of dense visual predictions such as object detection, segmentation, and human pose estimation. Compared to recent advanced MLP architectures such as MLP-Mixer (Tolstikhin et al. 2021), ResMLP (Touvron et al. 2021), and gMLP (Liu et al. 2021), whose architectures are sensitive to image size and are infeasible in dense prediction tasks, CycleMLP has two appealing advantages: 1) CycleMLP can cope with various spatial sizes of images; 2) CycleMLP achieves linear computational complexity with respect to the image size by using local windows. In contrast, previous MLPs have$O(N^{2})$computational complexity due to their full connections in space. 3) The relationship between convolution, multi-head self-attention in Transformer, and CycleMLP are discussed through an intuitive theoretical analysis. We build a family of models that can surpass state-of-the-art MLP and Transformer models e.g., Swin Transformer (Liu et al. 2021), while using fewer parameters and FLOPs. CycleMLP expands the MLP-like models’ applicability, making them versatile backbone networks that achieve competitive results on dense prediction tasks For example, CycleMLP-Tiny outperforms Swin-Tiny by 1.3% mIoU on ADE20 K dataset with fewer FLOPs. Moreover, CycleMLP also shows excellent zero-shot robustness on ImageNet-C dataset. Shoufa Chen, Enze Xie, Chongjian Ge, Runjian Chen, Ding Liang, Ping Luo 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2022 | CycleMLP: A MLP-like Architecture for Dense Prediction
Shoufa Chen, Enze Xie, Chongjian Ge, Runjian Chen, Ding Liang, Ping Luo 0002 |
ICLR | 1 |
| 2022 | CtrlFormer: Learning Transferable State Representation for Visual Control via TransformerabstractTransformer has achieved great successes in learning vision and language representation, which is general across various downstream tasks. In visual control, learning transferable state representation that can transfer between different control tasks is important to reduce the training sample size. However, porting Transformer to sample-efficient visual control remains a challenging and unsolved problem. To this end, we propose a novel Control Transformer (CtrlFormer), possessing many appealing benefits that prior arts do not have. Firstly, CtrlFormer jointly learns self-attention mechanisms between visual tokens and policy tokens among different control tasks, where multitask representation can be learned and transferred without catastrophic forgetting. Secondly, we carefully design a contrastive reinforcement learning paradigm to train CtrlFormer, enabling it to achieve high sample efficiency, which is important in control problems. For example, in the DMControl benchmark, unlike recent advanced methods that failed by producing a zero score in the “Cartpole” task after transfer learning with 100k samples, CtrlFormer can achieve a state-of-the-art score with only 100k samples while maintaining the performance of previous tasks. The code and models are released in our project homepage. Yao Mu 0001, Shoufa Chen, Mingyu Ding, Jianyu Chen 0002, Runjian Chen, Ping Luo 0002 |
ICML | 2 |
| 2022 | AdaptFormer: Adapting Vision Transformers for Scalable Visual RecognitionabstractPretraining Vision Transformers (ViTs) has achieved great success in visual recognition. A following scenario is to adapt a ViT to various image and video recognition tasks. The adaptation is challenging because of heavy computation and memory storage. Each model needs an independent and complete finetuning process to adapt to different tasks, which limits its transferability to different visual domains.To address this challenge, we propose an effective adaptation approach for Transformer, namely AdaptFormer, which can adapt the pre-trained ViTs into many different image and video tasks efficiently.It possesses several benefits more appealing than prior arts.Firstly, AdaptFormer introduces lightweight modules that only add less than 2% extra parameters to a ViT, while it is able to increase the ViT's transferability without updating its original pre-trained parameters, significantly outperforming the existing 100\% fully fine-tuned models on action recognition benchmarks.Secondly, it can be plug-and-play in different Transformers and scalable to many visual tasks.Thirdly, extensive experiments on five image and video datasets show that AdaptFormer largely improves ViTs in the target domains. For example, when updating just 1.5% extra parameters, it achieves about 10% and 19% relative improvement compared to the fully fine-tuned models on Something-Something~v2 and HMDB51, respectively. Code is available at https://github.com/ShoufaChen/AdaptFormer. Shoufa Chen, Chongjian Ge, Zhan Tong, Jiangliu Wang, Yibing Song, Jue Wang 0001, Ping Luo 0002 |
NeurIPS | 1 |
| 2021 | Watch Only Once: An End-to-End Video Action Detection FrameworkabstractWe propose an end-to-end pipeline, named Watch Once Only (WOO), for video action detection. Current methods either decouple video action detection task into separated stages of actor localization and action classification or train two separated models within one stage. In contrast, our approach solves the actor localization and action classification simultaneously in a unified network. The whole pipeline is significantly simplified by unifying the backbone network and eliminating many hand-crafted components. WOO takes a unified video backbone to simultaneously extract features for actor location and action classification. In addition, we introduce spatial-temporal action embeddings into our framework and design a spatial-temporal fusion module to obtain more discriminative features with richer information, which further boosts the action classification performance. Extensive experiments on AVA and JHMDB datasets show that WOO achieves state-of-the-art performance, while still reduces up to 16.7% GFLOPs compared with existing methods. We hope our work can inspire re-thinking the convention of action detection and serve as a solid baseline for end-to-end action detection. Code is available at https://github.com/ShoufaChen/WOO. Shoufa Chen, Peize Sun, Enze Xie, Chongjian Ge, Jiannan Wu, Ping Luo 0002 |
ICCV | 1 |
| 2019 | Single-Shot Detector with Multiple Inference PathsabstractIn this paper, we investigate the problem of resource-constrained object detection using deep learning, which is a challenging problem in real-world applications. To address this problem, we propose a single-shot detector with multiple inference paths based on a multi-scale DenseNet. Experiments are carried out on the PASCAL VOC and COCO datasets, and results show that, with significant computation reduction, our detection network obtains comparable performance corresponding to the single-shot detectors, such as YOLO and SSD. Shoufa Chen, Xinggang Wang |
ICIP | 1 |