VLDB 2026 Research / reviewers in the wild / expert
Runqi Qiao
dblp:379/9974
· DBLP profile ↗
6ranked-venue papers
3as first author
6since 2021 · last 2025
0009-0000-4508-201XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 3 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
6 papers |
Vision and language · 50% Language models and text generation · 34% Question answering and dialogue systems · 8% | |
| Interdisciplinary, comprehensive, and emerging computing
2 papers |
Computing education · 79% Computational social science and digital humanities · 21% | |
| Computer graphics and multimedia
1 paper |
Visual content generation and editing · 100% |
Topics — the 12 heaviest of 13, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Vision and language
visual reasoning |
1.7 | 2 | 2025 | We-Math: Does Your Large Multimodal Model Achieve Human-like Mathematical Reasoning? · ACL (1) 2025 V-Oracle: Making Progressive Reasoning in Deciphering Oracle Bones for You and Me · ACL (1) 2025 |
Computer vision › Vision and language
multimodal reasoning |
1.6 | 2 | 2025 | V-Oracle: Making Progressive Reasoning in Deciphering Oracle Bones for You and Me · ACL (1) 2025 Making Visual Sense of Oracle Bones for You and Me · CVPR 2024 |
Natural language and speech › Language models and text generation
instruction following |
0.9 | 1 | 2025 | Toward Verifiable Instruction-Following Alignment for Retrieval Augmented Generation · AAAI 2025 |
Natural language and speech › Question answering and dialogue systems
knowledge-intensive question answering |
0.9 | 1 | 2025 | Toward Verifiable Instruction-Following Alignment for Retrieval Augmented Generation · AAAI 2025 |
Natural language and speech › Language models and text generation
large language model evaluation |
0.9 | 1 | 2025 | CS-Bench: A Comprehensive Benchmark for Large Language Models towards Computer Science Mastery · ICLR 2025 |
Natural language and speech › Language models and text generation
mathematical reasoning |
0.9 | 1 | 2025 | We-Math: Does Your Large Multimodal Model Achieve Human-like Mathematical Reasoning? · ACL (1) 2025 |
Computer vision › Vision and language › vision-language model
multimodal large language model |
0.9 | 1 | 2025 | OCR-Critic: Aligning Multimodal Large Language Models' Perception through Critical Feedback · ACM Multimedia 2025 |
Computer vision › Image recognition and object detection › text recognition
optical character recognition |
0.9 | 1 | 2025 | OCR-Critic: Aligning Multimodal Large Language Models' Perception through Critical Feedback · ACM Multimedia 2025 |
Natural language and speech › Language models and text generation
retrieval-augmented generation |
0.9 | 1 | 2025 | Toward Verifiable Instruction-Following Alignment for Retrieval Augmented Generation · AAAI 2025 |
Computer vision › Vision and language
visual question answering |
0.9 | 1 | 2025 | V-Oracle: Making Progressive Reasoning in Deciphering Oracle Bones for You and Me · ACL (1) 2025 |
Visual content generation and editing › image generation
conditional image synthesis |
0.8 | 1 | 2024 | Making Visual Sense of Oracle Bones for You and Me · CVPR 2024 |
Computational social science and digital humanities › cultural heritage
cultural heritage computing |
0.2 | 1 | 2024 | Making Visual Sense of Oracle Bones for You and Me · CVPR 2024 |
Methods — techniques the papers use, named apart from their topics
foundation model prompting · 2.3multilingual evaluation · 1.7benchmark construction · 1.7large multimodal model · 1.7text-to-image generation · 1.5image denoising · 1.5visual instruction tuning · 0.9synthetic data generation · 0.9oracle alignment tuning · 0.9instruction tuning · 0.9dynamic alignment strategy · 0.9data augmentation · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Toward Verifiable Instruction-Following Alignment for Retrieval Augmented GenerationabstractFollowing natural instructions is crucial for the effective application of Retrieval-Augmented Generation (RAG) systems. Despite recent advancements in Large Language Models (LLMs), research on assessing and improving instruction-following (IF) alignment within the RAG domain remains limited. To address this issue, we propose VIF-RAG, an automated, scalable, and verifiable synthetic pipeline for instruction-following alignment in RAG systems. We start by manually crafting a minimal set of atomic instructions (100k) through automated processes. To further bridge the gap in instruction-following auto-evaluation for RAG systems, we introduce FollowRAG Benchmark, which includes approximately 3K test samples, covering 22 categories of general instruction constraints and four knowledge-intensive QA datasets. Due to its robust pipeline design, FollowRAG can seamlessly integrate with different RAG benchmarks. Using FollowRAG and eight widely-used IF and foundational abilities benchmarks for LLMs, we demonstrate that VIF-RAG markedly enhances LLM performance across a broad range of general instruction constraints while effectively leveraging its capabilities in RAG scenarios. Further analysis offers practical insights for achieving IF alignment in RAG systems. Guanting Dong 0001, Xiaoshuai Song, Yutao Zhu 0001, Runqi Qiao, Zhicheng Dou, Ji-Rong Wen |
AAAI | 4 |
| 2025 | V-Oracle: Making Progressive Reasoning in Deciphering Oracle Bones for You and MeabstractOracle Bone Script (OBS) is a vital treasure of human civilization, rich in insights from ancient societies. However, the evolution of written language over millennia complicates its decipherment. In this paper, we propose V-Oracle, an innovative framework that utilizes Large Multi-modal Models (LMMs) for interpreting OBS. V-Oracle applies principles of pictographic character formation and frames the task as a visual question-answering (VQA) problem, establishing a multi-step reasoning chain. It proposes a multi-dimensional data augmentation for synthesizing high-quality OBS samples, and also implements a multi-phase oracle alignment tuning to improve LMMs’ visual reasoning capabilities. Moreover, to bridge the evaluation gap in the OBS field, we further introduce Oracle-Bench, a comprehensive benchmark that emphasizes process-oriented assessment and incorporates both standard and out-of-distribution setups for realistic evaluation. Extensive experimental results can demonstrate the effectiveness of our method in providing quantitative analyses and superior deciphering capability. Runqi Qiao, Qiuna Tan, Guanting Dong 0001, MinhuiWu MinhuiWu, Jiapeng Wang 0005, Zhuoma Gongque, Yadong Xue, Zhimin Bao, Lan Yang 0014, Chen Li 0031, Honggang Zhang 0002 |
ACL (1) | 1 |
| 2025 | We-Math: Does Your Large Multimodal Model Achieve Human-like Mathematical Reasoning?abstractRunqi Qiao, Qiuna Tan, Guanting Dong, MinhuiWu MinhuiWu, Chong Sun, Xiaoshuai Song, Jiapeng Wang, Zhuoma GongQue, Shanglin Lei, YiFan Zhang, Zhe Wei, Miaoxuan Zhang, Runfeng Qiao, Xiao Zong, Yida Xu, Peiqing Yang, Zhimin Bao, Muxi Diao, Chen Li, Honggang Zhang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Runqi Qiao, Qiuna Tan, Guanting Dong 0001, Minhui Wu, Xiaoshuai Song, Jiapeng Wang 0005, Zhuoma Gongque, Shanglin Lei, Miaoxuan Zhang, Runfeng Qiao, Xiao Zong, Peiqing Yang 0003, Zhimin Bao, Muxi Diao, Chen Li 0031, Honggang Zhang 0002 |
ACL (1) | 1 |
| 2025 | CS-Bench: A Comprehensive Benchmark for Large Language Models towards Computer Science MasteryabstractLarge language models (LLMs) have demonstrated significant potential in advancing various fields of research and society. However, the current community of LLMs overly focuses on benchmarks for analyzing specific foundational skills (e.g. mathematics and code generation), neglecting an all-round evaluation of the computer science field. To bridge this gap, we introduce CS-Bench, the first multilingual (English, Chinese, French, German) benchmark dedicated to evaluating the performance of LLMs in computer science. CS-Bench comprises approximately 10K meticulously curated test samples, covering 26 subfields across 4 key areas of computer science, encompassing various task forms and divisions of knowledge and reasoning. Utilizing CS-Bench, we conduct a comprehensive evaluation of over 30 mainstream LLMs, revealing the relationship between CS performance and model scales. We also quantitatively analyze the reasons for failures in existing LLMs and highlight directions for improvements, including knowledge supplementation and CS-specific reasoning. Further cross-capability experiments show a high correlation between LLMs' capabilities in computer science and their abilities in mathematics and coding. Moreover, expert LLMs specialized in mathematics and coding also demonstrate strong performances in several CS subfields. Looking ahead, we envision CS-Bench serving as a cornerstone for LLM applications in the CS field and paving new avenues in assessing LLMs' diverse reasoning capabilities. Our project homepage is available at https://csbench.github.io/. Xiaoshuai Song, Muxi Diao, Guanting Dong 0001, Yujia Fu, Runqi Qiao, Zhexu Wang, Dayuan Fu, Huangxuan Wu, Weihao Zeng 0003, Yejie Wang, Zhuoma Gongque, Jianing Yu 0001, Qiuna Tan, Weiran Xu |
ICLR | 6 |
| 2025 | OCR-Critic: Aligning Multimodal Large Language Models' Perception through Critical FeedbackabstractRecent advancements in Large Multimodal Models have demonstrated impressive performance in various tasks. However, their capabilities in error detection and resolution for Optical Character Recognition (OCR) remain underexplored. To address this gap, we construct the first visual instruction tuning dataset specifically for detailed OCR error analysis. Building on this foundation, we develop a universal, plug-and-play OCR-Critic model that incorporates three novel dynamic alignment strategies. These strategies systematically mitigate LMMs' weaknesses in OCR tasks by providing coarse-to-fine error feedback. To comprehensively evaluate these capabilities, we introduce OCR-ERROR, a benchmark designed to assess LMMs' ability to detect and categorize OCR errors, covering two task types, diverse error categories, and 2,400 rigorously validated samples. Experimental results show that OCR-Critic effectively identifies fine-grained OCR errors across multiple domains. With the integration of our dynamic alignment strategies, the LMM further achieves substantial performance gains on four prominent benchmarks, demonstrating both versatility and effectiveness. Qiuna Tan, Runqi Qiao, Guanting Dong 0001, Minhui Wu, Jiapeng Wang 0005, Miaoxuan Zhang, Chen Li 0031, Honggang Zhang 0002 |
ACM Multimedia | 2 |
| 2024 | Making Visual Sense of Oracle Bones for You and MeabstractVisual perception evolves over time. This is particularly the case of oracle bone scripts, where visual glyphs seem intuitive to people from distant past prove difficult to be understood in contemporary eyes. While semantic correspon-dence of an oracle can be found via a dictionary lookup, this proves to be not enough for public viewers to connect the dots, i.e., why does this oracle mean that? Common solution relies on a laborious curation process to collect visual guide for each oracle (Fig. 1), which hinges on the case-by-case effort and taste of curators. This paper delves into one natural follow-up question: can AI take over? Begin with a comprehensive human study, we show par-ticipants could indeed make better sense of an oracle glyph subjected to a proper visual guide and its efficacy can be approximated via a novel metric termed TransOV (Trans-ferable Oracle Visuals). We then define a new conditional visual generation task based on an oracle glyph and its se-mantic meaning and importantly approach it by circumventing any form of model training in the presence of fatal lack of oracle data. At its heart is to leverage foundation model like GPT-4V to reason about the visual cues hidden inside an oracle and take advantage of an existing text-to-image model for final visual guide generation. Extensive empirical evidence shows our AI-enabled visual guides achieve signif-icantly comparable TransOV performance compared with those collected under manual efforts. Finally, we demon-strate the versatility of our system under a more complex setting, where it is required to work alongside with an AI image denoiser to cope with raw oracle scan image inputs (cf processed clean oracle glyphs). Code is available at https://github.com/RQ-Lab/OBS-Visual. Runqi Qiao, Lan Yang 0014, Kaiyue Pang, Honggang Zhang 0002 |
CVPR | 1 |