EDBT 2026 Demo / reviewers in the wild / expert
Junjie Shen 0008
dblp:349/4492
· DBLP profile ↗
6ranked-venue papers
0as first author
6since 2021 · last 2026
0009-0008-6983-5213ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
Generative modeling · 82% Language models and text generation · 18% | |
| Computer graphics and multimedia
3 papers |
Visual content generation and editing · 100% |
Topics — the 9 heaviest of 10, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Generative modeling
diffusion model |
1.0 | 1 | 2026 | MoFu: Scale-Aware Modulation and Fourier Fusion for Multi-Subject Video Generation · AAAI 2026 |
Machine learning › Generative modeling › diffusion model
video diffusion model |
1.0 | 1 | 2026 | MoFu: Scale-Aware Modulation and Fourier Fusion for Multi-Subject Video Generation · AAAI 2026 |
Visual content generation and editing
video generation |
1.0 | 1 | 2026 | MoFu: Scale-Aware Modulation and Fourier Fusion for Multi-Subject Video Generation · AAAI 2026 |
Machine learning › Generative modeling › multimodal generation
multimodal generative model |
0.9 | 1 | 2025 | CTR-Driven Advertising Image Generation with Multimodal Large Language Models · WWW 2025 |
Natural language and speech › Language models and text generation › large language model fine-tuning
multimodal large language model fine-tuning |
0.9 | 1 | 2025 | CTR-Driven Advertising Image Generation with Multimodal Large Language Models · WWW 2025 |
Visual content generation and editing
layout generation |
0.9 | 1 | 2025 | Uni-Layout: Integrating Human Feedback in Unified Layout Generation and Evaluation · ACM Multimedia 2025 |
Machine learning › Generative modeling
image generation |
0.8 | 1 | 2024 | Towards Reliable Advertising Image Generation Using Human Feedback · ECCV (20) 2024 |
Machine learning › Generative modeling › video generation
text-to-video generation |
0.3 | 1 | 2026 | MoFu: Scale-Aware Modulation and Fourier Fusion for Multi-Subject Video Generation · AAAI 2026 |
Human-AI interaction
human feedback |
0.2 | 1 | 2024 | Towards Reliable Advertising Image Generation Using Human Feedback · ECCV (20) 2024 |
Methods — techniques the papers use, named apart from their topics
human feedback · 2.4scale-aware modulation · 2.0fast fourier transform · 2.0LLM-guided feature modulation · 2.0reinforcement learning from reward model · 1.7preference optimization · 1.7multimodal pretraining · 1.7reinforcement learning from human feedback · 1.5layout evaluation · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MoFu: Scale-Aware Modulation and Fourier Fusion for Multi-Subject Video GenerationabstractMulti-subject video generation aims to synthesize videos from textual prompts and multiple reference images, ensuring that each subject preserves natural scale and visual fidelity. However, current methods face two challenges: scale inconsistency, where variations in subject size lead to unnatural generation, and permutation sensitivity, where the order of reference inputs causes subject distortion. In this paper, we propose MoFu, a unified framework that tackles both challenges. For scale inconsistency, we introduce Scale-Aware Modulation (SMO), an LLM-guided module that extracts implicit scale cues from the prompt and modulates features to ensure consistent subject sizes. To address permutation sensitivity, we present a simple yet effective Fourier Fusion strategy that processes the frequency information of reference features via the Fast Fourier Transform to produce a unified representation. Besides, we design a Scale-Permutation Stability Loss to jointly encourage scale-consistent and permutation-invariant generation. To further evaluate these challenges, we establish a dedicated benchmark with controlled variations in subject scale and reference permutation. Extensive experiments demonstrate that MoFu significantly outperforms existing methods in preserving natural scale, subject fidelity, and overall visual quality. Run Ling, Ke Cao 0001, Ao Ma 0005, Runze He, Changwei Wang 0001, Rongtao Xu, Yihua Shao, Zhanjie Zhang, Guibing Guo, Jingjing Lv, Junjie Shen 0008, Ching Law, Xingwei Wang 0001 |
AAAI | 16 |
| 2025 | Generate E-commerce Product Background by Integrating Category Commonality and Personalized StyleabstractThe state-of-the-art methods for e-commerce product background generation suffer from the inefficiency of designing product-wise prompts when scaling up the production, as well as the ineffectiveness of describing fine-grained styles when customizing personalized backgrounds for some specific brands. To address these obstacles, we integrate the category commonality and personalized style into diffusion models. Concretely, we propose a Category-Wise Generator to enable large-scale background generation with only one model for the first time. A unique identifier in the prompt is assigned to each category, whose attention is located on the background by a mask-guided cross attention layer to learn the category-wise style. Furthermore, for products with specific and fine-grained requirements in layout, elements, etc, a Personality-Wise Generator is devised to learn such personalized style directly from a reference image to resolve textual ambiguities, and is trained in a self-supervised manner for more efficient training data usage. To advance research in this field, the first large-scale e-commerce product background generation dataset BG60k is constructed, which covers more than 60k product images from over 2k categories. Experiments demonstrate that our method could generate high-quality backgrounds for different categories, and maintain the personalized background style of reference images. BG60k will be available at https://github.com/Whileherham/BG60k. Haohan Wang, Jingjing Lv, Junjie Shen 0008, Zhangang Lin, Jingping Shao |
ICASSP | 6 |
| 2025 | Uni-Layout: Integrating Human Feedback in Unified Layout Generation and Evaluation
Shuo Lu, Yanyin Chen, Fengheng Li, Jingjing Lv, Junjie Shen 0008, Ching Law, Jian Liang 0001 |
ACM Multimedia | 8 |
| 2025 | CTR-Driven Advertising Image Generation with Multimodal Large Language ModelsabstractIn web data, advertising images are crucial for capturing user attention and improving advertising effectiveness. Most existing methods generate background for products primarily focus on the aesthetic quality, which may fail to achieve satisfactory online performance. To address this limitation, we explore the use of Multimodal Large Language Models (MLLMs) for generating advertising images by optimizing for Click-Through Rate (CTR) as the primary objective. Firstly, we build targeted pre-training tasks, and leverage a large-scale e-commerce multimodal dataset to equip MLLMs with initial capabilities for advertising image generation tasks. To further improve the CTR of generated images, we propose a novel reward model to fine-tune pre-trained MLLMs through Reinforcement Learning (RL), which can jointly utilize multimodal features and accurately reflect user click preferences. Meanwhile, a product-centric preference optimization strategy is developed to ensure that the generated background content aligns with the product characteristics after fine-tuning, enhancing the overall relevance and effectiveness of the advertising images. Extensive experiments have demonstrated that our method achieves state-of-the-art performance in both online and offline metrics. Our code and pre-trained models are publicly available at: https://github.com/Chenguoz/CAIG. Xingye Chen, Zhenbang Du, Yanyin Chen, Haohan Wang, Linkai Liu 0002, Jinyuan Zhao, Jingjing Lv, Junjie Shen 0008, Zhangang Lin, Jingping Shao, Yuanjie Shao, Xinge You, Changxin Gao, Nong Sang |
WWW | 13 |
| 2024 | Towards Reliable Advertising Image Generation Using Human Feedback
Zhenbang Du, Haohan Wang, Jingsen Wang, Jingjing Lv, Xin Zhu 0008, Junsheng Jin, Junjie Shen 0008, Zhangang Lin, Jingping Shao |
ECCV (20) | 11 |
| 2023 | Relation-Aware Diffusion Model for Controllable Poster Layout GenerationabstractPoster layout is a crucial aspect of poster design. Prior methods primarily focus on the correlation between visual content and graphic elements. However, a pleasant layout should also consider the relationship between visual and textual contents and the relationship between elements. In this study, we introduce a relation-aware diffusion model for poster layout generation that incorporates these two relationships in the generation process. Firstly, we devise a visual-textual relation-aware module that aligns the visual and textual representations across modalities, thereby enhancing the layout's efficacy in conveying textual information. Subsequently, we propose a geometry relation-aware module that learns the geometry relationship between elements by comprehensively considering contextual information. Additionally, the proposed method can generate diverse layouts based on user constraints. To advance research in this field, we have constructed a poster layout dataset named CGL-Dataset V2. Our proposed method outperforms state-of-the-art methods on CGL-Dataset V2. The data and code will be available at https://github.com/liuan0803/RADM. Fengheng Li, Honghe Zhu, Jingjing Lv, Xin Zhu 0008, Junjie Shen 0008, Zhangang Lin, Jingping Shao |
CIKM | 9 |