Jianshu Zeng

dblp:392/5051 · DBLP profile ↗
← Back
3ranked-venue papers
1as first author
3since 2021 · last 2025
0009-0005-5326-4442ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Vision and language · 19% Language models and text generation · 19% Reinforcement learning · 19%
Computer graphics and multimedia
1 paper
Visual content generation and editing · 100%

Topics — the 8 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation
alignment
0.912025
MM-RLHF: The Next Step Forward in Multimodal LLM Alignment · ICML 2025
Machine learning › Generative modeling
diffusion model
0.912025
ROSE: Remove Objects with Side Effects in Videos · NeurIPS 2025
Computer vision › Vision and language › vision-language model
multimodal large language model
0.912025
MM-RLHF: The Next Step Forward in Multimodal LLM Alignment · ICML 2025
Machine learning › Reinforcement learning
reinforcement learning from human feedback
0.912025
MM-RLHF: The Next Step Forward in Multimodal LLM Alignment · ICML 2025
Computer vision › Video understanding and tracking › video reconstruction
video inpainting
0.912025
ROSE: Remove Objects with Side Effects in Videos · NeurIPS 2025
Visual content generation and editing
video editing
0.912025
ROSE: Remove Objects with Side Effects in Videos · NeurIPS 2025
Visual content generation and editing › video editing
video object removal
0.912025
ROSE: Remove Objects with Side Effects in Videos · NeurIPS 2025
Machine learning › Trustworthy machine learning
interpretability
0.312025
MM-RLHF: The Next Step Forward in Multimodal LLM Alignment · ICML 2025

Methods — techniques the papers use, named apart from their topics

synthetic data generation · 1.7diffusion transformer · 1.73d rendering · 1.7reward modeling · 0.9dynamic reward scaling · 0.9
YearPublicationVenuePosition
2025 MM-RLHF: The Next Step Forward in Multimodal LLM Alignment
abstract
Existing efforts to align multimodal large language models (MLLMs) with human preferences have only achieved progress in narrow areas, such as hallucination reduction, but remain limited in practical applicability and generalizability. To this end, we introduce **MM-RLHF**, a dataset containing **120k** fine-grained, human-annotated preference comparison pairs. This dataset represents a substantial advancement over existing resources, offering superior size, diversity, annotation granularity, and quality. Leveraging this dataset, we propose several key innovations to improve both the quality of reward models and the efficiency of alignment algorithms. Notably, we introduce the **Critique-Based Reward Model**, which generates critiques of model outputs before assigning scores, offering enhanced interpretability and more informative feedback compared to traditional scalar reward mechanisms. Additionally, we propose **Dynamic Reward Scaling**, a method that adjusts the loss weight of each sample according to the reward signal, thereby optimizing the use of high-quality comparison pairs. Our approach is rigorously evaluated across **10** distinct dimensions, encompassing **27** benchmarks, with results demonstrating significant and consistent improvements in model performance (Figure.1).
Yifan Zhang 0004, Haochen Tian 0001, Chaoyou Fu, Peiyan Li 0001, Jianshu Zeng, Wulin Xie, Yang Shi 0009, Huanyu Zhang 0002, Junkang Wu, Xue Wang 0010, Yibo Hu 0001, Tingting Gao, Zhang Zhang 0001, Fan Yang 0094, Di Zhang 0026, Liang Wang 0001, Rong Jin 0001
ICML6
2025 ROSE: Remove Objects with Side Effects in Videos
abstract
Video object removal has achieved advanced performance due to the recent success of video generative models. However, when addressing the side effects of objects, \textit{e.g.,} their shadows and reflections, existing works struggle to eliminate these effects for the scarcity of paired video data as supervision. This paper presents \method, termed \textbf{R}emove \textbf{O}bjects with \textbf{S}ide \textbf{E}ffects, a framework that systematically studies the object's effects on environment, which can be categorized into five common cases: shadows, reflections, light, translucency and mirror. Given the challenges of curating paired videos exhibiting the aforementioned effects, we leverage a 3D rendering engine for synthetic data generation. We carefully construct a fully-automatic pipeline for data preparation, which simulates a large-scale paired dataset with diverse scenes, objects, shooting angles, and camera trajectories. ROSE is implemented as an video inpainting model built on diffusion transformer. To localize all object-correlated areas, the entire video is fed into the model for reference-based erasing. Moreover, additional supervision is introduced to explicitly predict the areas affected by side effects, which can be revealed through the differential mask between the paired videos. To fully investigate the model performance on various side effect removal, we presents a new benchmark, dubbed ROSE-Bench, incorporating both common scenarios and the five special side effects for comprehensive evaluation. Experimental results demonstrate that \method achieves superior performance compared to existing video object erasing models and generalizes well to real-world video scenarios.
Chenxuan Miao, Yutong Feng, Jianshu Zeng, Zixiang Gao, Hantang Liu, Yunfeng Yan, Donglian Qi, Xi Chen 0119, Hengshuang Zhao
NeurIPS3
2024 Distill the Knowledge of Multimodal Large Language Model into Text-to-Image Vehicle Re-identification
Jianshu Zeng
ICPR (18)1