EDBT 2026 Demo / reviewers in the wild / expert
Jinsong Li 0001
dblp:29/3923-1
· DBLP profile ↗
5ranked-venue papers
0as first author
5since 2021 · last 2026
—ORCID · unresolved
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
4 papers |
Vision and language · 85% Trustworthy machine learning · 9% Reinforcement learning · 4% | |
| Computer graphics and multimedia
1 paper |
Visual content generation and editing · 50% Image and video processing · 50% |
Topics — the 14 heaviest of 15, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Vision and language
visual grounding |
1.0 | 1 | 2026 | Game Ground Bench: Probing the Limits of LVLMs in Complex Semantic Grounding Across Game Universes · AAAI 2026 |
Image and video processing › video processing
temporal consistency |
0.9 | 1 | 2025 | Light-a-Video: Training-Free Video Relighting via Progressive Light Fusion · ICCV 2025 |
Visual content generation and editing › visual effects
video relighting |
0.9 | 1 | 2025 | Light-a-Video: Training-Free Video Relighting via Progressive Light Fusion · ICCV 2025 |
Machine learning › Trustworthy machine learning
data leakage |
0.8 | 1 | 2024 | Are We on the Right Way for Evaluating Large Vision-Language Models? · NeurIPS 2024 |
Computer vision › Vision and language
image captioning |
0.8 | 1 | 2024 | ShareGPT4V: Improving Large Multi-modal Models with Better Captions · ECCV (17) 2024 |
Computer vision › Vision and language
multimodal benchmark |
0.8 | 1 | 2024 | Are We on the Right Way for Evaluating Large Vision-Language Models? · NeurIPS 2024 |
Computer vision › Vision and language › multimodal evaluation
multimodal capability evaluation |
0.8 | 1 | 2024 | Are We on the Right Way for Evaluating Large Vision-Language Models? · NeurIPS 2024 |
Computer vision › Vision and language › vision-language model
multimodal large language model |
0.8 | 1 | 2024 | ShareGPT4V: Improving Large Multi-modal Models with Better Captions · ECCV (17) 2024 |
Computer vision › Vision and language
video captioning |
0.8 | 1 | 2024 | ShareGPT4Video: Improving Video Understanding and Generation with Better Captions · NeurIPS 2024 |
Computer vision › Vision and language
video-language model |
0.8 | 1 | 2024 | ShareGPT4Video: Improving Video Understanding and Generation with Better Captions · NeurIPS 2024 |
Computer vision › Vision and language
vision-language model |
0.8 | 1 | 2024 | Are We on the Right Way for Evaluating Large Vision-Language Models? · NeurIPS 2024 |
Computer vision › Vision and language › vision-language model
vision-language model evaluation |
0.8 | 1 | 2024 | Are We on the Right Way for Evaluating Large Vision-Language Models? · NeurIPS 2024 |
Machine learning › Reinforcement learning
policy optimization |
0.3 | 1 | 2026 | Game Ground Bench: Probing the Limits of LVLMs in Complex Semantic Grounding Across Game Universes · AAAI 2026 |
Machine learning › Generative modeling › video generation
text-to-video generation |
0.2 | 1 | 2024 | ShareGPT4Video: Improving Video Understanding and Generation with Better Captions · NeurIPS 2024 |
Methods — techniques the papers use, named apart from their topics
reinforcement learning · 1.0fine-tuning · 1.0diffusion model · 0.9attention mechanism · 0.9multimodal large language model · 0.8differential video captioning · 0.8data leakage metrics · 0.8benchmark construction · 0.8GPT-4V annotation · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Game Ground Bench: Probing the Limits of LVLMs in Complex Semantic Grounding Across Game UniversesabstractLarge Vision-Language Models (LVLMs) have demonstrated remarkable capabilities, yet their ability to ground language in complex, interactive environments such as video games remains a critical frontier. Existing benchmarks are inadequate for this purpose: real-world datasets like RefCOCO introduce a domain gap; GUI-centric benchmarks lack the complexity of modern game interfaces; and existing game-specific benchmarks are often too simplistic or narrow, failing to assess fine-grained, generalizable grounding capabilities. To address this issue, we propose GGBench — a large-scale, cross-genre benchmark designed to probe the grounding capabilities of LVLMs in diverse gaming scenarios. GGBench features unprecedented genre diversity, encompassing 10 categories including card games, first-person shooters, and role-playing games, with a total of 1335 test images. It focuses on tasks that require connecting natural language instructions to specific in-game objects and UI elements. Experimental results show existing models perform poorly on GGBench, with weak grounding abilities, especially in complex game scenarios. Due to limited data scale, fine-tuning them for gaming scenarios is also challenging. To address this, we propose Game-R1, a novel training method centered on the Grounded Reinforcement Policy Optimization (GRPO) algorithm. GRPO maximizes limited interaction data utility and enables robust few-shot generalization across games. Extensive experiments show Game-R1 significantly outperforms existing LVLMs on GGBench, validating our approach. GGBench provides a solid and comprehensive evaluation platform for subsequent research on agents in gaming environments, which strongly promotes development in this field. Zhangyang Qi, Jinsong Li 0001, Hongjian Wu, Jiaqi Wang 0003, Hengshuang Zhao |
AAAI | 2 |
| 2025 | Light-a-Video: Training-Free Video Relighting via Progressive Light FusionabstractRecent advancements in image relighting models, driven by large-scale datasets and pre-trained diffusion models, have enabled the imposition of consistent lighting. However, video relighting still lags, primarily due to the excessive training costs and the scarcity of diverse, high-quality video relighting datasets. A simple application of image relighting models on a frame-by-frame basis leads to several issues: lighting source inconsistency and relighted appearance inconsistency, resulting in flickers in the generated videos. In this work, we propose Light-A-Video, a training-free approach to achieve temporally smooth video relighting. Adapted from image relighting models, Light-A-Video introduces two key techniques to enhance lighting consistency. First, we design a Consistent Light Attention (CLA) module, which enhances cross-frame interactions within the self-attention layers of the image relight model to stabilize the generation of the background lighting source. Second, leveraging the physical principle of light transport independence, we apply linear blending between the source video's appearance and the relighted appearance, using a Progressive Light Fusion (PLF) strategy to ensure smooth temporal transitions in illumination. Experiments show that Light-A-Video improves the temporal consistency of relighted video while maintaining the relighted image quality, ensuring coherent lighting transitions across frames. Project page: https://bujiazi.github.io/light-a-video.github.io/. Jiazi Bu, Pengyang Ling, Pan Zhang 0001, Qidong Huang, Jinsong Li 0001, Xiaoyi Dong, Yuhang Zang, Yuhang Cao, Anyi Rao, Jiaqi Wang 0003, Li Niu 0002 |
ICCV | 7 |
| 2024 | ShareGPT4V: Improving Large Multi-modal Models with Better Captions
Lin Chen 0026, Jinsong Li 0001, Xiaoyi Dong, Pan Zhang 0001, Conghui He, Jiaqi Wang 0003, Feng Zhao 0004, Dahua Lin |
ECCV (17) | 2 |
| 2024 | ShareGPT4Video: Improving Video Understanding and Generation with Better CaptionsabstractWe present the ShareGPT4Video series, aiming to facilitate the video understanding of large video-language models (LVLMs) and the video generation of text-to-video models (T2VMs) via dense and precise captions. The series comprises: 1) ShareGPT4Video, 40K GPT4V annotated dense captions of videos with various lengths and sources, developed through carefully designed data filtering and annotating strategy. 2) ShareCaptioner-Video, an efficient and capable captioning model for arbitrary videos, with 4.8M high-quality aesthetic videos annotated by it. 3) ShareGPT4Video-8B, a simple yet superb LVLM that reached SOTA performance on three advancing video benchmarks. To achieve this, taking aside the non-scalable costly human annotators, we find using GPT4V to caption video with a naive multi-frame or frame-concatenation input strategy leads to less detailed and sometimes temporal-confused results. We argue the challenge of designing a high-quality video captioning strategy lies in three aspects: 1) Inter-frame precise temporal change understanding. 2) Intra-frame detailed content description. 3) Frame-number scalability for arbitrary-length videos. To this end, we meticulously designed a differential video captioning strategy, which is stable, scalable, and efficient for generating captions for videos with arbitrary resolution, aspect ratios, and length. Based on it, we construct ShareGPT4Video, which contains 40K high-quality videos spanning a wide range of categories, and the resulting captions encompass rich world knowledge, object attributes, camera movements, and crucially, detailed and precise temporal descriptions of events. Based on ShareGPT4Video, we further develop ShareCaptioner-Video, a superior captioner capable of efficiently generating high-quality captions for arbitrary videos. We annotated 4.8M aesthetically appealing videos by it and verified their effectiveness on a 10-second text2video generation task. For video understanding, we verified the effectiveness of ShareGPT4Video on several current LVLM architectures and presented our superb new LVLM ShareGPT4Video-8B. All the models, strategies, and annotations will be open-sourced and we hope this project can serve as a pivotal resource for advancing both the LVLMs and T2VMs community. Lin Chen 0016, Xilin Wei, Jinsong Li 0001, Xiaoyi Dong, Pan Zhang 0001, Yuhang Zang, Haodong Duan, Lin Bin, Zhenyu Tang 0004, Li Yuan 0007, Yu Qiao 0001, Dahua Lin, Feng Zhao 0004, Jiaqi Wang 0003 |
NeurIPS | 3 |
| 2024 | Are We on the Right Way for Evaluating Large Vision-Language Models?abstractLarge vision-language models (LVLMs) have recently achieved rapid progress, sparking numerous studies to evaluate their multi-modal capabilities. However, we dig into current evaluation works and identify two primary issues: 1) Visual content is unnecessary for many samples. The answers can be directly inferred from the questions and options, or the world knowledge embedded in LLMs. This phenomenon is prevalent across current benchmarks. For instance, GeminiPro achieves 42.7% on the MMMU benchmark without any visual input, and outperforms the random choice baseline across six benchmarks near 24% on average. 2) Unintentional data leakage exists in LLM and LVLM training. LLM and LVLM could still answer some visual-necessary questions without visual content, indicating the memorizing of these samples within large-scale training data. For example, Sphinx-X-MoE gets 43.6% on MMMU without accessing images, surpassing its LLM backbone with 17.9%. Both problems lead to misjudgments of actual multi-modal gains and potentially misguide the study of LVLM. To this end, we present MMStar, an elite vision-indispensable multi-modal benchmark comprising 1,500 samples meticulously selected by humans. MMStar benchmarks 6 core capabilities and 18 detailed axes, aiming to evaluate LVLMs' multi-modal capacities with carefully balanced and purified samples. These samples are first roughly selected from current benchmarks with an automated pipeline, human review is then involved to ensure each curated sample exhibits visual dependency, minimal data leakage, and requires advanced multi-modal capabilities. Moreover, two metrics are developed to measure data leakage and actual performance gain in multi-modal training. We evaluate 16 leading LVLMs on MMStar to assess their multi-modal capabilities, and on 7 benchmarks with the proposed metrics to investigate their data leakage and actual multi-modal gain. Jinsong Li 0001, Xiaoyi Dong, Pan Zhang 0001, Yuhang Zang, Haodong Duan, Jiaqi Wang 0003, Yu Qiao 0001, Dahua Lin, Feng Zhao 0004 |
NeurIPS | 2 |