Yibo Wang 0039

dblp:25/4764-39 · DBLP profile ↗
← Back
7ranked-venue papers
2as first author
7since 2021 · last 2025
0009-0007-1774-0039ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 2 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Generative modeling · 21% Language models and text generation · 18% Image recognition and object detection · 16%
Computer graphics and multimedia
1 paper
Visual content generation and editing · 100%
Network and information security
1 paper
Security and privacy of machine learning · 100%

Topics — the 19 heaviest of 20, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Generative modeling
diffusion model
1.522024
GeoDiffusion: Text-Prompted Geometric Control for Object Detection Data Generation · ICLR 2024
DetDiffusion: Synergizing Generative and Perceptive Models for Enhanced Data Generation and Perception · CVPR 2024
Computer vision › Image recognition and object detection
object detection
1.522024
GeoDiffusion: Text-Prompted Geometric Control for Object Detection Data Generation · ICLR 2024
DetDiffusion: Synergizing Generative and Perceptive Models for Enhanced Data Generation and Perception · CVPR 2024
Machine learning › Efficient and distributed learning
adaptive computation
0.912025
Ada-R1: Hybrid-CoT via Bi-Level Adaptive Reasoning Optimization · NeurIPS 2025
Natural language and speech › Language models and text generation › large language model reasoning
adaptive reasoning
0.912025
Ada-R1: Hybrid-CoT via Bi-Level Adaptive Reasoning Optimization · NeurIPS 2025
Natural language and speech › Language models and text generation
chain-of-thought reasoning
0.912025
Ada-R1: Hybrid-CoT via Bi-Level Adaptive Reasoning Optimization · NeurIPS 2025
Machine learning › Reinforcement learning › policy optimization
group relative policy optimization
0.912025
R1-ShareVL: Incentivizing Reasoning Capabilities of Multimodal Large Language Models via Share-GRPO · NeurIPS 2025
Machine learning › Efficient and distributed learning
inference efficiency
0.912025
Ada-R1: Hybrid-CoT via Bi-Level Adaptive Reasoning Optimization · NeurIPS 2025
Natural language and speech › Language models and text generation
large language model safety
0.912025
Panacea: Mitigating Harmful Fine-tuning for Large Language Models via Post-fine-tuning Perturbation · NeurIPS 2025
Computer vision › Vision and language › vision-language model
multimodal large language model
0.912025
R1-ShareVL: Incentivizing Reasoning Capabilities of Multimodal Large Language Models via Share-GRPO · NeurIPS 2025
Computer vision › Vision and language
multimodal reasoning
0.912025
R1-ShareVL: Incentivizing Reasoning Capabilities of Multimodal Large Language Models via Share-GRPO · NeurIPS 2025
Machine learning › Reinforcement learning
policy optimization
0.912025
R1-ShareVL: Incentivizing Reasoning Capabilities of Multimodal Large Language Models via Share-GRPO · NeurIPS 2025
Machine learning › Trustworthy machine learning › AI safety › safety alignment
safety alignment preservation
0.912025
Panacea: Mitigating Harmful Fine-tuning for Large Language Models via Post-fine-tuning Perturbation · NeurIPS 2025
Computer vision › Image recognition and object detection › object detection
data augmentation for detection
0.812024
DetDiffusion: Synergizing Generative and Perceptive Models for Enhanced Data Generation and Perception · CVPR 2024
Machine learning › Generative modeling › diffusion model › controllable generation
layout-conditioned image generation
0.812024
DetDiffusion: Synergizing Generative and Perceptive Models for Enhanced Data Generation and Perception · CVPR 2024
Machine learning › Generative modeling › diffusion model
text-to-image generation
0.812024
GeoDiffusion: Text-Prompted Geometric Control for Object Detection Data Generation · ICLR 2024
Visual content generation and editing
style transfer
0.812024
StyleCrafter: Taming Artistic Video Diffusion with Reference-Augmented Adapter Learning · ACM Trans. Graph. 2024
Visual content generation and editing › video-to-video translation
stylized video generation
0.812024
StyleCrafter: Taming Artistic Video Diffusion with Reference-Augmented Adapter Learning · ACM Trans. Graph. 2024
Visual content generation and editing
video generation
0.812024
StyleCrafter: Taming Artistic Video Diffusion with Reference-Augmented Adapter Learning · ACM Trans. Graph. 2024
Robotics › Autonomous driving
perception
0.212024
GeoDiffusion: Text-Prompted Geometric Control for Object Detection Data Generation · ICLR 2024

Methods — techniques the papers use, named apart from their topics

random perturbation · 1.7adaptive perturbation · 1.7diffusion model · 1.5reinforcement learning · 0.9preference optimization · 0.9model merging · 0.9data transformation · 0.9advantage estimation · 0.9perception-aware loss · 0.8perception-aware attribute · 0.8data augmentation · 0.8adapter learning · 0.8
YearPublicationVenuePosition
2025 Ada-R1: Hybrid-CoT via Bi-Level Adaptive Reasoning Optimization
abstract
Recently, long-thought reasoning models achieve strong performance on complex reasoning tasks, but often incur substantial inference overhead, making efficiency a critical concern. Our empirical analysis reveals that the benefit of using Long-CoT varies across problems: while some problems require elaborate reasoning, others show no improvement—or even degraded accuracy. This motivates adaptive reasoning strategies that tailor reasoning depth to the input. However, prior work primarily reduces redundancy within long reasoning paths, limiting exploration of more efficient strategies beyond the Long-CoT paradigm. To address this, we propose a novel two-stage framework for adaptive and efficient reasoning. First, we construct a hybrid reasoning model by merging long and short CoT models to enable diverse reasoning styles. Second, we apply bi-level preference training to guide the model to select suitable reasoning styles (group-level), and prefer concise and correct reasoning within each style group (instance-level). Experiments demonstrate that our method significantly reduces inference costs compared to other baseline approaches, while maintaining performance. Notably, on five mathematical datasets, the average length of reasoning is reduced by more than 50\%, highlighting the potential of adaptive strategies to optimize reasoning efficiency in large language models.
Haotian Luo, Haiying He, Yibo Wang 0039, Jinluan Yang, Naiqiang Tan, Xiaochun Cao, Dacheng Tao, Li Shen 0008
NeurIPS3
2025 Panacea: Mitigating Harmful Fine-tuning for Large Language Models via Post-fine-tuning Perturbation
abstract
Harmful fine-tuning attack introduces significant security risks to the fine-tuning services. Main-stream defenses aim to vaccinate the model such that the later harmful fine-tuning attack is less effective. However, our evaluation results show that such defenses are fragile-- with a few fine-tuning steps, the model still can learn the harmful knowledge. To this end, we do further experiment and find that an embarrassingly simple solution-- adding purely random perturbations to the fine-tuned model, can recover the model from harmful behaviors, though it leads to a degradation in the model’s fine-tuning performance. To address the degradation of fine-tuning performance, we further propose \methodname, which optimizes an adaptive perturbation that will be applied to the model after fine-tuning. \methodname maintains model's safety alignment performance without compromising downstream fine-tuning performance. Comprehensive experiments are conducted on different harmful ratios, fine-tuning tasks and mainstream LLMs, where the average harmful scores are reduced by up-to 21.2%, while maintaining fine-tuning performance. As a by-product, we analyze the adaptive perturbation and show that different layers in various LLMs have distinct safety coefficients. Source code available at https://github.com/w-yibo/Panacea.
Yibo Wang 0039, Tiansheng Huang, Li Shen 0008, Huanjin Yao, Haotian Luo, Naiqiang Tan, Jiaxing Huang 0001, Dacheng Tao
NeurIPS1
2025 Mulberry: Empowering MLLM with o1-like Reasoning and Reflection via Collective Monte Carlo Tree Search
abstract
In this work, we aim to develop an MLLM that understands and solves questions by learning to create each intermediate step of the reasoning involved till the final answer. To this end, we propose Collective Monte Carlo Tree Search (CoMCTS), a new learning-to-reason method for MLLMs, which introduces the concept of collective learning into ``tree search'' for effective and efficient reasoning-path searching and learning. The core idea of CoMCTS is to leverage collective knowledge from multiple models to collaboratively conjecture, search and identify effective reasoning paths toward correct answers via four iterative operations including Expansion, Simulation and Error Positioning, Backpropagation, and Selection. Using CoMCTS, we construct Mulberry-260k, a multimodal dataset with a tree of rich, explicit and well-defined reasoning nodes for each question. With Mulberry-260k, we perform collective SFT to train our model, Mulberry, a series of MLLMs with o1-like step-by-step Reasoning and Reflection capabilities. Extensive experiments demonstrate the superiority of our proposed methods on various benchmarks. Code is available at https://github.com/HJYao00/Mulberry.
Huanjin Yao, Jiaxing Huang 0001, Jingyi Zhang 0005, Yibo Wang 0039, Shunyu Liu 0001, YuXin Song 0001, Haocheng Feng, Li Shen 0008, Dacheng Tao
NeurIPS5
2025 R1-ShareVL: Incentivizing Reasoning Capabilities of Multimodal Large Language Models via Share-GRPO
abstract
In this work, we aim to incentivize the reasoning ability of Multimodal Large Language Models (MLLMs) via reinforcement learning (RL) and develop an effective approach that mitigates the sparse reward and advantage vanishing issues during RL. To this end, we propose Share-GRPO, a novel RL approach that tackle these issues by exploring and sharing diverse reasoning trajectories over expanded question space. Specifically, Share-GRPO first expands the question space for a given question via data transformation techniques, and then encourages MLLM to effectively explore diverse reasoning trajectories over the expanded question space and shares the discovered reasoning trajectories across the expanded questions during RL. In addition, Share-GRPO also shares reward information during advantage computation, which estimates solution advantages hierarchically across and within question variants, allowing more accurate estimation of relative advantages and improving the stability of policy training. Extensive evaluations over 6 widely-used reasoning benchmarks showcase the superior performance of our method. Code is available at https://github.com/HJYao00/R1-ShareVL.
Huanjin Yao, Qixiang Yin, Jingyi Zhang 0005, Min Yang 0007, Yibo Wang 0039, Li Shen 0008, Minghui Qiu, Dacheng Tao, Jiaxing Huang 0001
NeurIPS5
2024 DetDiffusion: Synergizing Generative and Perceptive Models for Enhanced Data Generation and Perception
abstract
Current perceptive models heavily depend on resource-intensive datasets, prompting the need for innovative solutions. Leveraging recent advances in diffusion models, synthetic data, by constructing image inputs from various annotations, proves beneficial for downstream tasks. While prior methods have separately addressed generative and perceptive models, DetDiffusion, for the first time, harmonizes both, tackling the challenges in generating effective data for perceptive models. To enhance image generation with perceptive models, we introduce perception-aware loss (P.A. loss) through segmentation, improving both quality and controllability. To boost the performance of specific perceptive models, our method customizes data augmentation by extracting and utilizing perception-aware attribute (P.A. Attr) during generation. Experimental results from the object detection task highlight DetDiffusion's superior performance, establishing a new state-of-the-art in layout-guided generation. Furthermore, image syntheses from DetDiffusion can effectively augment training data, significantly enhancing downstream detection performance.
Yibo Wang 0039, Ruiyuan Gao 0001, Kai Chen 0023, Kaiqiang Zhou, Yingjie Cai, Lanqing Hong, Zhenguo Li, Lihui Jiang, Dit-Yan Yeung, Qiang Xu 0001, Kai Zhang 0008
CVPR1
2024 GeoDiffusion: Text-Prompted Geometric Control for Object Detection Data Generation
abstract
Diffusion models have attracted significant attention due to the remarkable ability to create content and generate data for tasks like image classification. However, the usage of diffusion models to generate the high-quality object detection data remains an underexplored area, where not only image-level perceptual quality but also geometric conditions such as bounding boxes and camera views are essential. Previous studies have utilized either copy-paste synthesis or layout-to-image (L2I) generation with specifically designed modules to encode the semantic layouts. In this paper, we propose the GeoDiffusion, a simple framework that can flexibly translate various geometric conditions into text prompts and empower pre-trained text-to-image (T2I) diffusion models for high-quality detection data generation. Unlike previous L2I methods, our GeoDiffusion is able to encode not only the bounding boxes but also extra geometric conditions such as camera views in self-driving scenes. Extensive experiments demonstrate GeoDiffusion outperforms previous L2I methods while maintaining 4x training time faster. To the best of our knowledge, this is the first work to adopt diffusion models for layout-to-image generation with geometric conditions and demonstrate that L2I-generated images can be beneficial for improving the performance of object detectors.
Kai Chen 0023, Enze Xie, Yibo Wang 0039, Lanqing Hong, Zhenguo Li, Dit-Yan Yeung
ICLR4
2024 StyleCrafter: Taming Artistic Video Diffusion with Reference-Augmented Adapter Learning
abstract
Text-to-video (T2V) models have shown remarkable capabilities in generating diverse videos. However, they struggle to produce user-desired artistic videos due to (i) text's inherent clumsiness in expressing specific styles and (ii) the generally degraded style fidelity. To address these challenges, we introduce StyleCrafter, a generic method that enhances pretrained T2V models with a style control adapter, allowing video generation in any style by feeding a reference image. Considering the scarcity of artistic video data, we propose to first train a style control adapter using style-rich image datasets, then transfer the learned stylization ability to video generation through a tailor-made finetuning paradigm. To promote content-style disentanglement, we employ carefully designed data augmentation strategies to enhance decoupled learning. Additionally, we propose a scale-adaptive fusion module to balance the influences of text-based content features and image-based style features, which helps generalization across various text and style combinations. StyleCrafter efficiently generates high-quality stylized videos that align with the content of the texts and resemble the style of the reference images. Experiments demonstrate that our approach is more flexible and efficient than existing competitors. Project page: https://gongyeliu.github.io/StyleCrafter.github.io/
Gongye Liu, Menghan Xia, Yong Zhang 0034, Haoxin Chen, Jinbo Xing, Yibo Wang 0039, Xintao Wang 0002, Ying Shan, Yujiu Yang 0001
ACM Trans. Graph.6