VLDB 2026 Research / reviewers in the wild / expert
Shitian Zhao
dblp:364/2271
· DBLP profile ↗
9ranked-venue papers
1as first author
9since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 1 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
5 papers |
Vision and language · 29% Language models and text generation · 27% Generative modeling · 22% | |
| Computer graphics and multimedia
2 papers |
Visual content generation and editing · 94% Image and video processing · 6% |
Topics — the 17 heaviest of 18, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Vision and language › vision-language model
multimodal large language model |
1.6 | 2 | 2025 | To Think or Not To Think: A Study of Thinking in Rule-Based Visual Reinforcement Fine-Tuning · NeurIPS 2025 SPHINX-X: Scaling Data and Parameters for a Family of Multi-modal Large Language Models · ICML 2024 |
Machine learning › Generative modeling › diffusion model
text-to-image generation |
1.0 | 1 | 2026 | Lumina-mGPT: Flexible Photorealistic Autoregressive Text-to-Image Generation · Int. J. Comput. Vis. 2026 |
Machine learning › Generative modeling
image generation |
0.9 | 1 | 2025 | Fontanimate: High Quality Few-Shot Font Generation Via Animating Font Transfer Process · ICCV 2025 |
Machine learning › Reinforcement learning › reinforcement learning for NLP
reinforcement fine-tuning |
0.9 | 1 | 2025 | To Think or Not To Think: A Study of Thinking in Rule-Based Visual Reinforcement Fine-Tuning · NeurIPS 2025 |
Visual content generation and editing › image generation
controllable image generation |
0.9 | 1 | 2025 | PixWizard: Versatile Image-to-Image Visual Assistant with Open-Language Instructions · ICLR 2025 |
Visual content generation and editing › font generation
few-shot font generation |
0.9 | 1 | 2025 | Fontanimate: High Quality Few-Shot Font Generation Via Animating Font Transfer Process · ICCV 2025 |
Visual content generation and editing
font generation |
0.9 | 1 | 2025 | Fontanimate: High Quality Few-Shot Font Generation Via Animating Font Transfer Process · ICCV 2025 |
Visual content generation and editing
image editing |
0.9 | 1 | 2025 | PixWizard: Versatile Image-to-Image Visual Assistant with Open-Language Instructions · ICLR 2025 |
Natural language and speech › Language models and text generation
context generation |
0.8 | 1 | 2024 | Causal-CoG: A Causal-Effect Look at Context Generation for Boosting Multi-Modal Language Models · CVPR 2024 |
Natural language and speech › Language models and text generation
large language model training |
0.8 | 1 | 2024 | SPHINX-X: Scaling Data and Parameters for a Family of Multi-modal Large Language Models · ICML 2024 |
Natural language and speech › Language models and text generation
prompting |
0.8 | 1 | 2024 | Causal-CoG: A Causal-Effect Look at Context Generation for Boosting Multi-Modal Language Models · CVPR 2024 |
Computer vision › Vision and language
visual question answering |
0.8 | 1 | 2024 | Causal-CoG: A Causal-Effect Look at Context Generation for Boosting Multi-Modal Language Models · CVPR 2024 |
Computer vision › Image recognition and object detection
image classification |
0.3 | 1 | 2025 | To Think or Not To Think: A Study of Thinking in Rule-Based Visual Reinforcement Fine-Tuning · NeurIPS 2025 |
Image and video processing
image restoration |
0.3 | 1 | 2025 | PixWizard: Versatile Image-to-Image Visual Assistant with Open-Language Instructions · ICLR 2025 |
Machine learning › Trustworthy machine learning › interpretability
causal analysis |
0.2 | 1 | 2024 | Causal-CoG: A Causal-Effect Look at Context Generation for Boosting Multi-Modal Language Models · CVPR 2024 |
Machine learning › Trustworthy machine learning
interpretability |
0.2 | 1 | 2024 | Causal-CoG: A Causal-Effect Look at Context Generation for Boosting Multi-Modal Language Models · CVPR 2024 |
Computer vision › Image recognition and object detection
visual recognition |
0.2 | 1 | 2024 | SPHINX-X: Scaling Data and Parameters for a Family of Multi-modal Large Language Models · ICML 2024 |
Methods — techniques the papers use, named apart from their topics
font transfer · 1.7animation · 1.7instruction tuning · 1.6verifiable reward · 0.9reinforcement fine-tuning · 0.9diffusion transformer · 0.9chain-of-thought · 0.9prompting · 0.8multimodal language model · 0.8multi-stage training · 0.8causal inference · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Lumina-mGPT: Flexible Photorealistic Autoregressive Text-to-Image Generation
Yi Xin 0003, Shitian Zhao, Le Zhuo, Weifeng Lin, Xinyue Li 0001, Guangtao Zhai, Xiaohong Liu 0001, Hongsheng Li 0001, Yu Qiao 0001, Peng Gao 0007 |
Int. J. Comput. Vis. | 3 |
| 2026 | Med-R1: Reinforcement Learning for Generalizable Medical Reasoning in Vision-Language ModelsabstractVision-language models (VLMs) have achieved impressive progress in natural image reasoning, yet their potential in medical imaging remains underexplored. Medical vision-language tasks demand precise understanding and clinically coherent answers, which are difficult to achieve due to complexity of medical data and the scarcity of high-quality expert annotations. These challenges limit the effectiveness of conventional supervised fine-tuning (SFT) and Chain-of-Thought (CoT) strategies that work well in general domains. To address these challenges, we propose Med-R1, a reinforcement learning (RL)-enhanced VLM designed to improve generalization and reliability in medical reasoning. Med-R1 adopts Group Relative Policy Optimization (GRPO) to encourage reward-guided learning beyond static annotations. We comprehensively evaluate Med-R1 across eight distinct medical imaging modalities. Med-R1 achieves a 29.94% improvement in average accuracy over its base model Qwen2-VL-2B, and even outperforms Qwen2-VL-72B-a model with $36\times $ more parameters. To assess cross-task generalization, we further evaluate Med-R1 on five question types. Med-R1 outperforms Qwen2-VL-2B by 32.06% in question-type generalization, also surpassing Qwen2-VL-72B. We further explore the thinking process in Med-R1, a crucial component of Deepseek-R1. Our results show that omitting intermediate rationales (No-Thinking Med-R1) not only improves cross-domain generalization with less training, but also challenges the common assumption that more reasoning always helps. Nevertheless, we also find that the Think-After Med-R1 variant further improves performance while maintaining interpretability. These findings suggest that, in medical VQA, the mere presence of explicit reasoning does not guarantee better performance. Instead, performance depends on the quality of the reasoning and the position where the reasoning is generated. Yuxiang Lai, Jike Zhong, Shitian Zhao, Konstantinos Psounis, Xiaofeng Yang 0005 |
IEEE Trans. Medical Imaging | 4 |
| 2025 | Fontanimate: High Quality Few-Shot Font Generation Via Animating Font Transfer Process
Kainan Yan, Shitian Zhao, Jie Wen 0001, Junjun He, Peng Gao 0007 |
ICCV | 4 |
| 2025 | PixWizard: Versatile Image-to-Image Visual Assistant with Open-Language InstructionsabstractThis paper presents a versatile image-to-image visual assistant, PixWizard, designed for image generation, manipulation, and translation based on free-from language instructions. To this end, we tackle a variety of vision tasks into a unified image-text-to-image generation framework and curate an Omni Pixel-to-Pixel Instruction-Tuning Dataset. By constructing detailed instruction templates in natural language, we comprehensively include a large set of diverse vision tasks such as text-to-image generation, image restoration, image grounding, dense image prediction, image editing, controllable generation, inpainting/outpainting, and more. Furthermore, we adopt Diffusion Transformers (DiT) as our foundation model and extend its capabilities with a flexible any resolution mechanism, enabling the model to dynamically process images based on the aspect ratio of the input, closely aligning with human perceptual processes. The model also incorporates structure-aware and semantic-aware guidance to facilitate effective fusion of information from the input image. Our experiments demonstrate that PixWizard not only shows impressive generative and understanding abilities for images with diverse resolutions but also exhibits generalization capabilities with unseen tasks and human instructions. Weifeng Lin, Renrui Zhang, Le Zhuo, Shitian Zhao, Siyuan Huang 0004, Junlin Xie, Peng Gao 0007, Hongsheng Li 0001 |
ICLR | 5 |
| 2025 | Sekai: A Video Dataset towards World ExplorationabstractVideo generation techniques have made remarkable progress, promising to be the foundation of interactive world exploration.However, existing video generation datasets are not well-suited for world exploration training as they suffer from some limitations: limited locations, short duration, static scenes, and a lack of annotations about exploration and the world.In this paper, we introduce Sekai (meaning "world" in Japanese), a high-quality first-person view worldwide video dataset with rich annotations for world exploration. It consists of over 5,000 hours of walking or drone view (FPV and UVA) videos from over 100 countries and regions across 750 cities. We develop an efficient and effective toolbox to collect, pre-process and annotate videos with location, scene, weather, crowd density, captions, and camera trajectories.Comprehensive analyses and experiments demonstrate the dataset’s scale, diversity, annotation quality, and effectiveness for training video generation models.We believe Sekai will benefit the area of video generation and world exploration, and motivate valuable applications. Zhen Li 0026, Chuanhao Li 0001, Xiaofeng Mao, Shaoheng Lin, Ming Li 0010, Shitian Zhao, Zhaopan Xu, Xinyue Li 0001, Yukang Feng, Zizhen Li, Fanrui Zhang, Jiaxin Ai, Yuwei Wu 0001, Tong He 0001, Yunde Jia, Kaipeng Zhang |
NeurIPS | 6 |
| 2025 | To Think or Not To Think: A Study of Thinking in Rule-Based Visual Reinforcement Fine-TuningabstractThis paper investigates the role of explicit thinking process in rule-based reinforcement fine-tuning (RFT) for multi-modal large language models (MLLMs). We first extend \textit{Thinking-RFT} to image classification task, using verifiable rewards for fine-tuning~(FT). Experiments show {Thinking-RFT} significantly outperforms supervised FT and yields a cross-dataset generalization effect. We then rethink and question whether explicit thinking in RFT is always necessary and beneficial. Challenging the convention that explicit thinking is crucial for the success of RFT, we introduce \textit{No-Thinking-RFT}, exploring RFT without thinking by introducing a simple equality accuracy reward. We evaluate No-Thinking-RFT on six diverse tasks across different model sizes and types. Experiment results reveal four key findings: \textbf{(1).} Visual perception tasks do not require thinking during RFT, as No-Thinking-RFT consistently outperforms or matches Thinking-RFT across model sizes and types. \textbf{(2).} Models with limited capabilities struggle to generate high-quality CoT for RFT, making Thinking-RFT less effective than No-Thinking-RFT. \textbf{(3).} There are inconsistencies between the answers in the thinking tags and answer tags for some responses of Thinking-RFT, which show lower average accuracy than the overall accuracy. \textbf{(4).} The performance gain of No-Thinking-RFT mainly stems from improved learning during no thinking FT and the avoidance of inference overthinking, as evidenced by the partial gains from appending empty thinking tags at inference time of Thinking-RFT. We hypothesize that explicit thinking before verifiable answers may hinder reward convergence and reduce performance in certain scenarios. To test this, we propose \textit{Think-After-Answer}, which places thinking after the answer to mitigate this effect for experimental verification. Lastly, we conduct a pilot study to explore whether MLLMs can learn when to think during RFT, introducing an \textit{Adaptive-Thinking} method. Experiments show that model converges to either thinking or not depending on model capability, achieving comparable or better performance than both Thinking and No-Thinking-RFT. Our findings suggest MLLMs can adaptively decide to think or not based on their capabilities and task complexity, offering insights into the thinking process in RFT. Jike Zhong, Shitian Zhao, Yuxiang Lai, Haoquan Zhang, Wang Zhu 0001, Kaipeng Zhang |
NeurIPS | 3 |
| 2025 | Debiasing Medical Knowledge for Prompting Universal Model in CT Image SegmentationabstractWith the assistance of large language models, which offer universal medical prior knowledge via text prompts, state-of-the-art Universal Models (UM) have demonstrated considerable potential in the field of medical image segmentation. Semantically detailed text prompts, on the one hand, indicate comprehensive knowledge; on the other hand, they bring biases that may not be applicable to specific cases involving heterogeneous organs or rare cancers. To this end, we propose a Debiased Universal Model (DUM) to consider instance-level context information and remove knowledge biases in text prompts from the causal perspective. We are the first to discover and mitigate the bias introduced by universal knowledge. Specifically, we propose to extract organ-level text prompts via language models and instance-level context prompts from the visual features of each image. We aim to highlight more on factual instance-level information and mitigate organ-level's knowledge bias. This process can be derived and theoretically supported by a causal graph, and instantiated by designing a standard UM (SUM) and a biased UM. The debiased output is finally obtained by subtracting the likelihood distribution output by biased UM from that of the SUM. Experiments on three large-scale multi-center external datasets and MSD internal tumor datasets show that our method enhances the model's generalization ability in handling diverse medical scenarios and reducing the potential biases, even with an improvement of 4.16% compared with popular universal model on the AbdomenAtlas dataset, showing the strong generalizability. The code is publicly available at https://github.com/DeepMed-Lab-ECNU/DUM. Boxiang Yun, Shitian Zhao, Qingli Li, Alex Chichung Kot, Yan Wang 0033 |
IEEE Trans. Medical Imaging | 2 |
| 2024 | Causal-CoG: A Causal-Effect Look at Context Generation for Boosting Multi-Modal Language ModelsabstractWhile Multi-modal Language Models (MLMs) demonstrate impressive multimodal ability, they still struggle on providing factual and precise responses for tasks like visual question answering (VQA). In this paper, we address this challenge from the perspective of contextual information. We propose Causal Context Generation, Causal-CoG, which is a prompting strategy that engages contextual information to enhance precise VQA during inference. Specifically, we prompt MLMs to generate contexts, i.e, text description of an image, and engage the generated contexts for question answering. Moreover, we investigate the ad-vantage of contexts on VQA from a causality perspective, introducing causality filtering to select samples for which contextual information is helpful. To show the effectiveness of Causal-CoG, we run extensive experiments on 10 multimodal benchmarks and show consistent improvements, e.g., +6.30% on POPE, +13.69% on Vizwiz and +6.43% on VQAv2 compared to direct decoding, surpassing existing methods. We hope Casual-CoG inspires explorations of context knowledge in multimodal models, and serves as a plug-and-play strategy for MLM decoding.11Code is released zhaoshitian/Causal-CoG Shitian Zhao, Zhuowan Li, Yadong Lu, Alan L. Yuille, Yan Wang 0033 |
CVPR | 1 |
| 2024 | SPHINX-X: Scaling Data and Parameters for a Family of Multi-modal Large Language ModelsabstractWe propose SPHINX-X, an extensive Multi-modality Large Language Model (MLLM) series developed upon SPHINX. To improve the architecture and training efficiency, we modify the SPHINX framework by removing redundant visual encoders, bypassing fully-padded sub-images with skip tokens, and simplifying multi-stage training into a one-stage all-in-one paradigm. To fully unleash the potential of MLLMs, we assemble a comprehensive multi-domain and multi-modal dataset covering publicly available resources in language, vision, and vision-language tasks. We further enrich this collection with our curated OCR intensive and Set-of-Mark datasets, extending the diversity and generality. By training over different base LLMs including TinyLlama-1.1B, InternLM2-7B, LLaMA2-13B, and Mixtral-8$\times$7B, we obtain a spectrum of MLLMs that vary in parameter size and multilingual capabilities. Comprehensive benchmarking reveals a strong correlation between the multi-modal performance with the data and parameter scales. Code and models are released at https://github.com/Alpha-VLLM/LLaMA2-Accessory. Renrui Zhang, Longtian Qiu, Siyuan Huang 0004, Weifeng Lin, Shitian Zhao, Shijie Geng, Kaipeng Zhang, Wenqi Shao, Conghui He, Junjun He, Hao Shao, Pan Lu, Yu Qiao 0001, Hongsheng Li 0001, Peng Gao 0007 |
ICML | 6 |