Jingbo Wang 0001

dblp:14/9860 · also Jingbo B. Wang · DBLP profile ↗
← Back
5ranked-venue papers
0as first author
5since 2021 · last 2025
0000-0001-7544-0084ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Segmentation and scene understanding · 41% Generative modeling · 32% Video understanding and tracking · 14%
Computer graphics and multimedia
2 papers
Visual content generation and editing · 100%
Theoretical computer science
1 paper
Quantum computing and quantum information · 91% Computational complexity · 9%

Topics — the 14 heaviest of 16, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Visual content generation and editing › image generation
personalized image generation
1.722025
DiffSensei: Bridging Multi-Modal LLMs and Diffusion Models for Customized Manga Generation · CVPR 2025
DreamRelation: Bridging Customization and Relation Generation · CVPR 2025
Machine learning › Generative modeling
diffusion model
0.912025
DreamRelation: Bridging Customization and Relation Generation · CVPR 2025
Computer vision › Segmentation and scene understanding
panoptic segmentation
0.912025
RMP-SAM: Towards Real-Time Multi-Purpose Segment Anything · ICLR 2025
Machine learning › Generative modeling › diffusion model › text-to-image generation
text-to-image diffusion model
0.912025
DreamRelation: Bridging Customization and Relation Generation · CVPR 2025
Computer vision › Video understanding and tracking
video instance segmentation
0.912025
RMP-SAM: Towards Real-Time Multi-Purpose Segment Anything · ICLR 2025
Visual content generation and editing
image generation
0.912025
DiffSensei: Bridging Multi-Modal LLMs and Diffusion Models for Customized Manga Generation · CVPR 2025
Visual content generation and editing › multimodal content generation
story visualization
0.912025
DiffSensei: Bridging Multi-Modal LLMs and Diffusion Models for Customized Manga Generation · CVPR 2025
Quantum computing and quantum information › quantum state tomography
classical shadow
0.912025
Learning the Complexity of Weakly Noisy Quantum States · ICLR 2025
Quantum computing and quantum information
quantum learning
0.912025
Learning the Complexity of Weakly Noisy Quantum States · ICLR 2025
Quantum computing and quantum information › quantum complexity theory
quantum state complexity
0.912025
Learning the Complexity of Weakly Noisy Quantum States · ICLR 2025
Computer vision › Segmentation and scene understanding
interactive segmentation
0.312025
RMP-SAM: Towards Real-Time Multi-Purpose Segment Anything · ICLR 2025
Computer vision › Vision and language › vision-language model
multimodal large language model
0.312025
DiffSensei: Bridging Multi-Modal LLMs and Diffusion Models for Customized Manga Generation · CVPR 2025
Machine learning › Generative modeling › diffusion model
text-to-image generation
0.312025
DreamRelation: Bridging Customization and Relation Generation · CVPR 2025
Computational complexity › learning theory
sample complexity
0.312025
Learning the Complexity of Weakly Noisy Quantum States · ICLR 2025

Methods — techniques the papers use, named apart from their topics

keypoint matching loss · 1.7identity-relation disentanglement · 1.7diffusion model · 1.7transformer · 1.4prompt-driven decoding · 0.9multimodal LLMs · 0.9multimodal LLM · 0.9masked cross-attention · 0.9masked cross attention · 0.9learning algorithms · 0.9dynamic convolution · 0.9classical shadow · 0.9
YearPublicationVenuePosition
2025 DreamRelation: Bridging Customization and Relation Generation
abstract
Customized image generation is essential for creating personalized content based on user prompts, allowing large-scale text-to-image diffusion models to more effectively meet individual needs. However, existing models often neglect the relationships between customized objects in generated images. In contrast, this work addresses this gap by focusing on relation-aware customized image generation, which seeks to preserve the identities from image prompts while maintaining the relationship specified in text prompts. Specifically, we introduce DreamRelation, a framework that disentangles identity and relation learning using a carefully curated dataset. Our training data consists of relation-specific images, independent object images containing identity information, and text prompts to guide relation generation. Then, we propose two key modules to tackle the two main challenges—generating accurate and natural relationships, especially when significant pose adjustments are required, and avoiding object confusion in cases of overlap. First, we introduce a keypoint matching loss that effectively guides the model in adjusting object poses closely tied to their relationships. Second, we incorporate local features of the image prompts to better distinguish between objects, preventing confusion in overlapping cases. Extensive results on our proposed benchmarks demonstrate the superiority of DreamRelation in generating precise relations while preserving object identities across a diverse set of objects and relationships.
Lu Qi 0001, Jianzong Wu, Jinbin Bai, Jingbo Wang 0001, Yunhai Tong, Xiangtai Li
CVPR5
2025 DiffSensei: Bridging Multi-Modal LLMs and Diffusion Models for Customized Manga Generation
abstract
Story visualization, the task of creating visual narratives from textual descriptions, has seen progress with text-to-image generation models. However, these models often lack effective control over character appearances and interactions, particularly in multi-character scenes. To address these limitations, we propose a new task: customized manga generation and introduce DiffSensei, an innovative framework specifically designed for generating manga with dynamic multi-character control. DiffSensei integrates a diffusion-based image generator with a multimodal large language model (MLLM) that acts as a text-compatible identity adapter. Our approach employs masked cross attention to seamlessly incorporate character features, enabling precise layout control without direct pixel transfer. Additionally, the MLLM-based adapter adjusts character features to align with panel-specific text cues, allowing flexible adjustments in character expressions, poses, and actions. We also introduce MangaZero, a large-scale dataset tailored to this task, containing 43,264 manga pages and 427,147 annotated panels, supporting the visualization of varied character interactions and movements across sequential frames. Extensive experiments demonstrate that DiffSensei outperforms existing models, marking a significant advancement in manga generation by enabling text- adaptable character customization. The code, model, and dataset are open-sourced to the community.1
Jianzong Wu, Jingbo Wang 0001, Yanhong Zeng, Xiangtai Li, Yunhai Tong
CVPR3
2025 Learning the Complexity of Weakly Noisy Quantum States
abstract
Quantifying the complexity of quantum states is a longstanding key problem in various subfields of science, ranging from quantum computing to the black-hole theory. The lower bound on quantum pure state complexity has been shown to grow linearly with system size [J. Haferkamp et al., 2022, *Nat. Phys.*]. However, extending this result to noisy circuit environments, which better reflect real quantum devices, remains an open challenge. In this paper, we explore the complexity of weakly noisy quantum states via the quantum learning method. We present an efficient learning algorithm, that leverages the classical shadow representation of target quantum states, to predict the circuit complexity of weakly noisy quantum states. Our algorithm is proved to be optimal in terms of sample complexity accompanied with polynomial classical processing time. Our result builds a bridge between the learning algorithm and quantum state complexity, meanwhile highlighting the power of learning algorithm in characterizing intrinsic properties of quantum states.
Bujiao Wu, Yanqi Song, Xiao Yuan 0002, Jingbo Wang 0001
ICLR5
2025 RMP-SAM: Towards Real-Time Multi-Purpose Segment Anything
abstract
Recent segmentation methods, which adopt large-scale data training and transformer architecture, aim to create one foundation model that can perform multiple tasks. However, most of these methods rely on heavy encoder and decoder frameworks, hindering their performance in real-time scenarios. To explore real-time segmentation, recent advancements primarily focus on semantic segmentation within specific environments, such as autonomous driving. However, they often overlook the generalization ability of these models across diverse scenarios. Therefore, to fill this gap, this work explores a novel real-time segmentation setting called real-time multi-purpose segmentation. It contains three fundamental sub-tasks: interactive segmentation, panoptic segmentation, and video instance segmentation. Unlike previous methods, which use a specific design for each task, we aim to use only a single end-to-end model to accomplish all these tasks in real-time. To meet real-time requirements and balance multi-task learning, we present a novel dynamic convolution-based method, Real-Time Multi-Purpose SAM (RMP-SAM). It contains an efficient encoder and an efficient decoupled adapter to perform prompt-driven decoding. Moreover, we further explore different training strategies and one new adapter design to boost co-training performance further. We benchmark several strong baselines by extending existing works to support our multi-purpose segmentation. Extensive experiments demonstrate that RMP-SAM is effective and generalizes well on proposed benchmarks and other specific semantic tasks. Our implementation of RMP-SAM achieves the optimal balance between accuracy and speed for these tasks. The code is released at \url{https://github.com/xushilin1/RAP-SAM}
Shilin Xu 0001, Haobo Yuan, Lu Qi 0001, Jingbo Wang 0001, Kai Chen 0026, Yunhai Tong, Bernard Ghanem, Xiangtai Li, Ming-Hsuan Yang 0001
ICLR5
2022 Fashionformer: A Simple, Effective and Unified Baseline for Human Fashion Segmentation and Recognition
Shilin Xu 0001, Xiangtai Li, Jingbo Wang 0001, Yunhai Tong, Dacheng Tao
ECCV (37)3