VLDB 2026 Research / reviewers in the wild / expert
Shuchen Weng
dblp:220/4303
· DBLP profile ↗
26ranked-venue papers
11as first author
23since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 25 · 10 first-author · 23 since 2021Graphics, computer vision, multimedia, augmented reality and games · 15 · 8 first-author · 12 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ReContraster: Making Your Posters Stand Out with Regional ContrastabstractEffective poster design requires rapidly capturing attention and clearly conveying messages.Inspired by the "contrast effects" principle, we propose ReContraster, the first training-free model to leverage regional contrast to make posters stand out.By emulating the cognitive behaviors of a poster designer, ReContraster introduces the compositional multi-agent system to identify elements, organize layout, and evaluate generated poster candidates.To further ensure harmonious transitions across region boundaries, ReContraster integrates the hybrid denoising strategy during the diffusion process.We additionally contribute a new benchmark dataset for comprehensive evaluation.Seven quantitative metrics and four user studies confirm its superiority over relevant state-of-the-art methods, producing visually striking and aesthetically appealing posters. Peixuan Zhang, Zijian Jia, Ziqi Cai, Shuchen Weng, Si Li 0001, Boxin Shi |
ACL (1) | 4 |
| 2026 | L-VOCAL: Language-based Video Colorization with Audio Alignment
Shuchen Weng, Huan Ouyang, Yuchen Hong, Lihan Lin, Si Li 0001, Boxin Shi |
Int. J. Comput. Vis. | 2 |
| 2026 | Affective Image Editing: Shaping Emotional Factors via Text Descriptions
Peixuan Zhang, Shuchen Weng, Chengxuan Zhu, Binghao Tang, Zijian Jia, Si Li 0001, Boxin Shi |
Int. J. Comput. Vis. | 2 |
| 2026 | L-C4: Language-based video colorization for creative and consistent color
Shuchen Weng, Huan Ouyang, Lihan Lin, Yu Li 0003, Si Li 0001, Boxin Shi |
Neurocomputing | 2 |
| 2026 | Toward Deeper Emotional Reflection: Crafting Affective Image Filters With Generative PriorsabstractSocial media platforms enable users to express emotions by posting text with accompanying images. In this paper, we propose the Affective Image Filter (AIF) task, which aims to reflect visually-abstract emotionsfrom text into visually-concrete images, thereby creating emotionally compelling results. We first introduce the AIF dataset and the formulation of the AIF models. Then, we present AIF-B as an initial attempt based on a multi-modal transformer architecture. After that, we propose AIF-D as an extension of AIF-B towards deeper emotional reflection, effectively leveraging generative priors from pre-trained large-scale diffusion models. Quantitative and qualitative experiments demonstrate that AIF models achieve superior performance for both content consistency and emotional fidelity compared to state-of-the-art methods. Extensive user study experiments demonstrate that AIF models are significantly more effective at evoking specific emotions. Based on the presented results, we comprehensively discuss the value and potential of AIF models. Peixuan Zhang, Shuchen Weng, Jiajun Tang 0001, Si Li 0001, Boxin Shi |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2025 | PhyS-EdiT: Physics-aware Semantic Image Editing with Text DescriptionabstractAchieving joint control over material properties, lighting, and high-level semantics in images is essential for applications in digital media, advertising, and interactive design. Existing methods often isolate these properties, lacking a cohesive approach to manipulating them simultaneously. We introduce PhyS-EdiT, a novel diffusion-based model that enables precise control over four critical material properties: roughness, metallicity, albedo, and transparency while integrating lighting and semantic adjustments within a single framework. To facilitate this disentangled control, we present PR-TIPS, a large and diverse synthetic dataset designed to improve the disentanglement of material and lighting effects. PhyS-EdiT incorporates a dual-network architecture and robust training strategies to balance low-level physical realism with high-level semantic coherence, supporting localized and continuous property adjustments. Extensive experiments demonstrate the superiority of PhyS-EdiT in editing both synthetic and real-world images, achieving state-of-the-art performance on material, lighting, and semantic editing tasks. Ziqi Cai, Shuchen Weng, Boxin Shi |
CVPR | 2 |
| 2025 | VIRES: Video Instance Repainting via Sketch and Text Guided GenerationabstractWe introduce VIRES, a video instance repainting method with sketch and text guidance, enabling video instance repainting, replacement, generation, and removal. Existing approaches struggle with temporal consistency and accurate alignment with the provided sketch sequence. VIRES leverages the generative priors of text-to-video models to maintain temporal consistency and produce visually pleasing results. We propose the Sequential ControlNet with the standardized self-scaling, which effectively extracts structure layouts and adaptively captures high-contrast sketch details. We further augment the diffusion transformer backbone with the sketch attention to interpret and inject fine-grained sketch semantics. A sketch-aware encoder ensures that repainted results are aligned with the provided sketch sequence. Additionally, we contribute the VIRESET, a dataset with detailed annotations tailored for training and evaluating video instance editing methods. Experimental results demonstrate the effectiveness of VIRES, which outperforms state-of-the-art methods in visual quality, temporal consistency, condition alignment, and human ratings. The code, dataset and pretrained models are available at: https://hjzheng.net/projects/VIRES. Shuchen Weng, Haojie Zheng, Peixuan Zhang, Yuchen Hong, Si Li 0001, Boxin Shi |
CVPR | 1 |
| 2025 | Audio-Sync Video Generation with Multi-Stream Temporal ControlabstractAudio is inherently temporal and closely synchronized with the visual world, making it a naturally aligned and expressive control signal for controllable video generation (e.g., movies).
Beyond control, directly translating audio into video is essential for understanding and visualizing rich audio narratives (e.g., Podcasts or historical recordings).
However, existing approaches fall short in generating high-quality videos with precise audio-visual synchronization, especially across diverse and complex audio types.
In this work, we introduce MTV, a versatile framework for audio-sync video generation. MTV explicitly separates audios into speech, effects, and music tracks, enabling disentangled control over lip motion, event timing, and visual mood, respectively—resulting in fine-grained and semantically aligned video generation.
To support the framework, we additionally present DEMIX, a dataset comprising high-quality cinematic videos and demixed audio tracks. DEMIX is structured into five overlapped subsets, enabling scalable multi-stage training for diverse generation scenarios.
Extensive experiments demonstrate that MTV achieves state-of-the-art performance across six standard metrics spanning video quality, text-video consistency, and audio-video alignment. Shuchen Weng, Haojie Zheng, Si Li 0001, Boxin Shi |
NeurIPS | 1 |
| 2025 | PanoWan: Lifting Diffusion Video Generation Models to 360° with Latitude/Longitude-aware Mechanisms
Shuchen Weng, Jingqi Liu, Chengxuan Zhu, Minggui Teng, Zijian Jia, Boxin Shi |
NeurIPS | 2 |
| 2025 | OpenCIR: Conditional Image Repainting With Open Condition MixtureabstractIn this paper, we introduce OpenCIR, a fully-functional Conditional Image Repainting (CIR) model designed for local image editing. Given an image and a combination of conditions related to geometry, texture, and color, CIR models are required to repaint instances and seamlessly composite them with the original images. Previous CIR models suffer from limited object categories, restricted condition modalities, and demanded geometry precision. In contrast, leveraging the generative priors from pre-trained models, OpenCIR could repaint open object categories. Equipped with redesigned condition injection modules and the condition extension strategy, OpenCIR is able to understand open condition modalities. Adopting the contour refinement strategy, OpenCIR allows users to specify instances with open geometry precision. In addition, we contribute the Open-CIR dataset, which includes detailed annotations, tailored for the comprehensive training and evaluation of the OpenCIR model. Extensive experiments demonstrate that OpenCIR outperforms relevant state-of-the-art methods, achieving superior visual quality, and more favorable results by human evaluators. Shuchen Weng, Xiaocheng Gong, Haojie Zheng, Si Li 0001, Boxin Shi |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2024 | Colorizing Monochromatic Radiance FieldsabstractThough Neural Radiance Fields (NeRF) can produce colorful 3D representations of the world by using a set of 2D images, such ability becomes non-existent when only monochromatic images are provided. Since color is necessary in representing the world, reproducing color from monochromatic radiance fields becomes crucial. To achieve this goal, instead of manipulating the monochromatic radiance fields directly, we consider it as a representation-prediction task in the Lab color space. By first constructing the luminance and density representation using monochromatic images, our prediction stage can recreate color representation on the basis of an image colorization module. We then reproduce a colorful implicit model through the representation of luminance, density, and color. Extensive experiments have been conducted to validate the effectiveness of our approaches. Our project page: https://liquidammonia.github.io/color-nerf. Yean Cheng, Renjie Wan, Shuchen Weng, Chengxuan Zhu, Yakun Chang, Boxin Shi |
AAAI | 3 |
| 2024 | Language-guided Image Reflection SeparationabstractThis paper studies the problem of language-guided re-flection separation, which aims at addressing the ill-posed reflection separation problem by introducing language de-scriptions to provide layer content. We propose a unified framework to solve this problem, which leverages the cross-attention mechanism with contrastive learning strategies to construct the correspondence between language descriptions and image layers. A gated network design and a ran-domized training strategy are employed to tackle the rec-ognizable layer ambiguity. The effectiveness of the pro-posed method is validated by the significant performance advantage over existing reflection separation methods on both quantitative and qualitative comparisons. Haofeng Zhong, Yuchen Hong, Shuchen Weng, Jinxiu Liang, Boxin Shi |
CVPR | 3 |
| 2024 | L-DiffER: Single Image Reflection Removal with Language-Based Diffusion Model
Yuchen Hong, Haofeng Zhong, Shuchen Weng, Jinxiu Liang, Boxin Shi |
ECCV (20) | 3 |
| 2024 | Conditional Image RepaintingabstractA number of advanced image editing technologies have demonstrated impressive performance in synthesizing visually pleasing results in accordance with user instructions. In this paper, we further extend the practicalities of image editing technology by proposing the conditional image repainting (CIR) task, which requires the model to synthesize realistic visual content based on multiple cross-modality conditions provided by the user. We first define condition inputs and formulate two-phased CIR models as the baseline. After that, we further design unified CIR models with novel condition fusion modules to improve the performance. For allowing users to express their intent more freely, our CIR models support both attributes and language to represent colors of repainted visual content. We demonstrate the effectiveness of CIR models by collecting and processing four datasets. Finally, we present a number of practical application scenarios of CIR models to demonstrate its usability. Shuchen Weng, Boxin Shi |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2023 | L-CoIns: Language-based Colorization With Instance AwarenessabstractLanguage-based colorization produces plausible colors consistent with the language description provided by the user. Recent studies introduce additional annotation to prevent color-object coupling and mismatch issues, but they still have difficulty in distinguishing instances corresponding to the same object words. In this paper, we propose a transformer-based framework to automatically aggregate similar image patches and achieve instance awareness without any additional knowledge. By applying our presented luminance augmentation and counter-color loss to break down the statistical correlation between luminance and color words, our model is driven to synthesize colors with better descriptive consistency. We further collect a dataset to provide distinctive visual characteristics and detailed language descriptions for multiple instances in the same image. Extensive experiments demonstrate our advantages of synthesizing visually pleasing and description-consistent results of instance-aware colorization. Shuchen Weng, Peixuan Zhang, Yu Li 0003, Si Li 0001, Boxin Shi |
CVPR | 2 |
| 2023 | Affective Image Filter: Reflecting Emotions from Text to ImagesabstractUnderstanding the emotions in text and presenting them visually is a very challenging problem that requires a deep understanding of natural language and high-quality image synthesis simultaneously. In this work, we propose Affective Image Filter (AIF), a novel model that is able to understand the visually-abstract emotions from the text and reflect them to visually-concrete images with appropriate colors and textures. We build our model based on the multi-modal transformer architecture, which unifies both images and texts into tokens and encodes the emotional prior knowledge. Various loss functions are proposed to understand complex emotions and produce appropriate visualization. In addition, we collect and contribute a new dataset with abundant aesthetic images and emotional texts for training and evaluating the AIF model. We carefully design four quantitative metrics and conduct a user study to comprehensively evaluate the performance, which demonstrates our AIF model outperforms state-of-the-art methods and could evoke specific emotional responses from human observers. Shuchen Weng, Peixuan Zhang, Si Li 0001, Boxin Shi |
ICCV | 1 |
| 2023 | L-CAD: Language-based Colorization with Any-level Descriptions using Diffusion PriorsabstractLanguage-based colorization produces plausible and visually pleasing colors under the guidance of user-friendly natural language descriptions. Previous methods implicitly assume that users provide comprehensive color descriptions for most of the objects in the image, which leads to suboptimal performance. In this paper, we propose a unified model to perform language-based colorization with any-level descriptions. We leverage the pretrained cross-modality generative model for its robust language understanding and rich color priors to handle the inherent ambiguity of any-level descriptions. We further design modules to align with input conditions to preserve local spatial structures and prevent the ghosting effect. With the proposed novel sampling strategy, our model achieves instance-aware colorization in diverse and complex scenarios. Extensive experimental results demonstrate our advantages of effectively handling any-level descriptions and outperforming both language-based and automatic colorization methods. The code and pretrained models
are available at: https://github.com/changzheng123/L-CAD. Shuchen Weng, Peixuan Zhang, Yu Li 0003, Si Li 0001, Boxin Shi |
NeurIPS | 2 |
| 2023 | LuminAIRe: Illumination-Aware Conditional Image Repainting for Lighting-Realistic GenerationabstractWe present the ilLumination-Aware conditional Image Repainting (LuminAIRe) task to address the unrealistic lighting effects in recent conditional image repainting (CIR) methods. The environment lighting and 3D geometry conditions are explicitly estimated from given background images and parsing masks using a parametric lighting representation and learning-based priors. These 3D conditions are then converted into illumination images through the proposed physically-based illumination rendering and illumination attention module. With the injection of illumination images, physically-correct lighting information is fed into the lighting-realistic generation process and repainted images with harmonized lighting effects in both foreground and background regions can be acquired, whose superiority over the results of state-of-the-art methods is confirmed through extensive experiments. For facilitating and validating the LuminAIRe task, a new dataset Car-LuminAIRe with lighting annotations and rich appearance variants is collected. Jiajun Tang 0001, Haofeng Zhong, Shuchen Weng, Boxin Shi |
NeurIPS | 3 |
| 2022 | L-CoDe: Language-Based Colorization Using Color-Object Decoupled ConditionsabstractColorizing a grayscale image is inherently an ill-posed problem with multi-modal uncertainty. Language-based colorization offers a natural way of interaction to reduce such uncertainty via a user-provided caption. However, the color-object coupling and mismatch issues make the mapping from word to color difficult. In this paper, we propose L-CoDe, a Language-based Colorization network using color-object Decoupled conditions. A predictor for object-color corresponding matrix (OCCM) and a novel attention transfer module (ATM) are introduced to solve the color-object coupling problem. To deal with color-object mismatch that results in incorrect color-object correspondence, we adopt a soft-gated injection module (SIM). We further present a new dataset containing annotated color-object pairs to provide supervisory signals for resolving the coupling problem. Experimental results show that our approach outperforms state-of-the-art methods conditioned on captions. Shuchen Weng, Jiajun Tang 0001, Si Li 0001, Boxin Shi |
AAAI | 1 |
| 2022 | UniCoRN: A Unified Conditional Image Repainting NetworkabstractConditional image repainting (CIR) is an advanced image editing task, which requires the model to generate visual content in user-specified regions conditioned on multiple cross-modality constraints, and composite the visual content with the provided background seamlessly. Existing methods based on two-phase architecture design assume dependency between phases and cause color-image incongruity. To solve these problems, we propose a novel Unified Conditional image Repainting Network (UniCoRN). We break the two-phase assumption in the CIR task by constructing the interaction and dependency relationship between background and other conditions. We further introduce the hierarchical structure into cross-modality similarity model to capture feature patterns at different levels and bridge the gap between visual content and color condition. A new Landscape-CIR dataset is collected and annotated to expand the application scenarios of the CIR task. Experiments show that UniCoRN achieves higher synthetic quality, better condition consistency, and more realistic compositing effect. Jimeng Sun 0002, Shuchen Weng, Si Li 0001, Boxin Shi |
CVPR | 2 |
| 2022 | L-CoDer: Language-Based Colorization with Color-Object Decoupling Transformer
Shuchen Weng, Yu Li 0003, Si Li 0001, Boxin Shi |
ECCV (18) | 2 |
| 2022 | CT2: Colorization Transformer via Color Tokens
Shuchen Weng, Jimeng Sun 0002, Yu Li 0003, Si Li 0001, Boxin Shi |
ECCV (7) | 1 |
| 2022 | Instance Contour Adjustment via Structure-Driven CNN
Shuchen Weng, Ming-Ching Chang, Boxin Shi |
ECCV (7) | 1 |
| 2020 | MISC: Multi-Condition Injection and Spatially-Adaptive Compositing for Conditional Person Image SynthesisabstractIn this paper, we explore synthesizing person images with multiple conditions for various backgrounds. To this end, we propose a framework named ``MISC" for conditional image generation and image compositing. For conditional image generation, we improve the existing condition injection mechanisms by leveraging the inter-condition correlations. For the image compositing, we theoretically prove the weaknesses of the cutting-edge methods, and make it more robust by removing the spatially-invariance constraint, and enabling the bounding mechanism and the spatial adaptability. We show the effectiveness of our method on the Video Instance-level Parsing dataset, and demonstrate the robustness through controllability tests. Shuchen Weng, Wenbo Li 0001, Dawei Li 0006, Hongxia Jin, Boxin Shi |
CVPR | 1 |
| 2020 | Conditional Image Repainting via Semantic Bridge and Piecewise Value Function
Shuchen Weng, Wenbo Li 0001, Dawei Li 0006, Hongxia Jin, Boxin Shi |
ECCV (9) | 1 |
| 2019 | Dual-stream CNN for Structured Time Series ClassificationabstractThe structured time series (STS) classification problem requires the modeling of interweaved spatiotemporal dependency. Most previous methods model these two dependencies independently. Due to the complexity of the STS data, we argue that a desirable method should be a holistic framework that is adaptive and flexible. This motivates us to design a deep neural network with such merits. Inspired by the dual-stream hypothesis in neural science, we propose a novel dual-stream framework for modeling the interweaved spatiotemporal dependency, and develop a convolutional neural network within this framework that aims to achieve high adaptability and flexibility in STS configurations of sequential order and dependency range. Our model is highly modularized and scalable, making it easy to be adapted to specific tasks. The effectiveness of our model is demonstrated through experiments on benchmark datasets for skeleton based activity recognition. Shuchen Weng, Wenbo Li 0001, Yi Zhang 0070, Siwei Lyu |
ICASSP | 1 |