VLDB 2026 Research / reviewers in the wild / expert
Ruilin Luo
dblp:10/8501
· DBLP profile ↗
6ranked-venue papers
3as first author
6since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 3 first-author · 6 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
4 papers |
Vision and language · 28% Language models and text generation · 27% Reinforcement learning · 27% | |
| Databases, data mining, and information retrieval
1 paper |
Data models and query languages · 100% |
Topics — the 12 heaviest of 12, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation
mathematical reasoning |
1.7 | 2 | 2025 | Unlocking Multimodal Mathematical Reasoning via Process Reward Model · NeurIPS 2025 Critical Tokens Matter: Token-Level Contrastive Estimation Enhances LLM's Reasoning Capability · ICML 2025 |
Computer vision › Vision and language › visual question answering
medical visual question answering |
1.0 | 1 | 2026 | Act as you think: Reinforcing Consistent Reasoning in Medical Visual Question Answering · ACL (1) 2026 |
Natural language and speech › Question answering and dialogue systems
reasoning consistency |
1.0 | 1 | 2026 | Act as you think: Reinforcing Consistent Reasoning in Medical Visual Question Answering · ACL (1) 2026 |
Machine learning › Reinforcement learning
reinforcement learning for reasoning |
1.0 | 1 | 2026 | Act as you think: Reinforcing Consistent Reasoning in Medical Visual Question Answering · ACL (1) 2026 |
Computer vision › Vision and language
visual question answering |
1.0 | 1 | 2026 | Act as you think: Reinforcing Consistent Reasoning in Medical Visual Question Answering · ACL (1) 2026 |
Machine learning › Representation and self-supervised learning
contrastive estimation |
0.9 | 1 | 2025 | Critical Tokens Matter: Token-Level Contrastive Estimation Enhances LLM's Reasoning Capability · ICML 2025 |
Computer vision › Vision and language › multimodal reasoning
multimodal mathematical reasoning |
0.9 | 1 | 2025 | Unlocking Multimodal Mathematical Reasoning via Process Reward Model · NeurIPS 2025 |
Machine learning › Reinforcement learning › reinforcement learning from human feedback
process reward model |
0.9 | 1 | 2025 | Unlocking Multimodal Mathematical Reasoning via Process Reward Model · NeurIPS 2025 |
Machine learning › Reinforcement learning
reinforcement learning from process rewards |
0.9 | 1 | 2025 | Unlocking Multimodal Mathematical Reasoning via Process Reward Model · NeurIPS 2025 |
Natural language and speech › Language models and text generation
large language model reasoning |
0.8 | 1 | 2024 | PTD-SQL: Partitioning and Targeted Drilling with LLMs in Text-to-SQL · EMNLP 2024 |
Data models and query languages › natural language interface › natural language interface to database
text-to-SQL |
0.8 | 1 | 2024 | PTD-SQL: Partitioning and Targeted Drilling with LLMs in Text-to-SQL · EMNLP 2024 |
Natural language and speech › Language models and text generation › chain-of-thought reasoning
multimodal chain-of-thought |
0.3 | 1 | 2025 | Unlocking Multimodal Mathematical Reasoning via Process Reward Model · NeurIPS 2025 |
Methods — techniques the papers use, named apart from their topics
partitioning · 1.5in-context learning · 1.5reinforcement learning · 1.0rollout sampling · 0.9process reward model · 0.9group relative policy optimization · 0.9direct preference optimization · 0.9contrastive estimation · 0.9chain-of-thought · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Act as you think: Reinforcing Consistent Reasoning in Medical Visual Question AnsweringabstractSongtao Jiang, Yuan Wang, Ruizhe Chen, Yan Zhang, Ruilin Luo, Bohan Lei, Yeying Jin, Sibo Song, ZhiBo Yang, Jimeng Sun, Jian Wu, Zuozhu Liu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Songtao Jiang, Ruizhe Chen, Yan Zhang 0004, Ruilin Luo, Bohan Lei, Yeying Jin, Sibo Song, Jimeng Sun 0001, Jian Wu 0001, Zuozhu Liu |
ACL (1) | 5 |
| 2025 | Critical Tokens Matter: Token-Level Contrastive Estimation Enhances LLM's Reasoning CapabilityabstractMathematical reasoning tasks pose significant challenges for large language models (LLMs) because they require precise logical deduction and sequence analysis. In this work, we introduce the concept of critical tokens – elements within reasoning trajectories that significantly influence incorrect outcomes. We present a novel framework for identifying these tokens through rollout sampling and demonstrate their substantial divergence from traditional error tokens. Through extensive experiments on datasets such as GSM8K and MATH500, we show that identifying and replacing critical tokens significantly improves model accuracy. We propose an efficient methodology for pinpointing these tokens in large-scale datasets using contrastive estimation and extend this framework to enhance model training processes with direct preference optimization (DPO). Experimental results on GSM8K and MATH500 benchmarks with the widely used models Llama-3 (8B and 70B) and Deepseek-math (7B) demonstrate the effectiveness of the proposed approach, cDPO. Our results underscore the potential of leveraging critical tokens to reduce errors in reasoning tasks, advancing the development of AI systems capable of robust logical deduction. Zicheng Lin, Qiuzhi Liu, Xing Wang 0007, Ruilin Luo, Chufan Shi, Siheng Li, Yujiu Yang 0001, Zhaopeng Tu |
ICML | 6 |
| 2025 | Unlocking Multimodal Mathematical Reasoning via Process Reward ModelabstractProcess Reward Models (PRMs) have shown promise in enhancing the mathematical reasoning capabilities of Large Language Models (LLMs) through Test-Time Scaling (TTS). However, their integration into multimodal reasoning remains largely unexplored. In this work, we take the first step toward unlocking the potential of PRMs in multimodal mathematical reasoning. We identify three key challenges: (i) the scarcity of high-quality reasoning data constrains the capabilities of foundation Multimodal Large Language Models (MLLMs), which imposes further limitations on the upper bounds of TTS and reinforcement learning (RL); (ii) a lack of automated methods for process labeling within multimodal contexts persists; (iii) the employment of process rewards in unimodal RL faces issues like reward hacking, which may extend to multimodal scenarios. To address these issues, we introduce URSA, a three-stage Unfolding multimodal pRocess-Supervision Aided training framework. We first construct MMathCoT-1M, a high-quality large-scale multimodal Chain-of-Thought (CoT) reasoning dataset, to build a stronger math reasoning foundation MLLM, URSA-8B. Subsequently, we go through an automatic process to synthesize process supervision data, which emphasizes both logical correctness and perceptual consistency. We introduce DualMath-1.1M to facilitate the training of URSA-8B-RM. Finally, we propose Process-Supervised Group-Relative-Policy-Optimization (PS-GRPO), pioneering a multimodal PRM-aided online RL method that outperforms vanilla GRPO. With PS-GRPO application, URSA-8B-PS-GRPO outperforms Gemma3-12B and GPT-4o by 8.4% and 2.7% on average across 6 benchmarks. Ruilin Luo, Zhuofan Zheng, Xinzhe Ni, Zicheng Lin, Songtao Jiang, Yiyao Yu, Chufan Shi, Ruihang Chu, Yujiu Yang 0001 |
NeurIPS | 1 |
| 2024 | Prior Relational Schema Assists Effective Contrastive Learning for Inductive Knowledge Graph CompletionabstractKnowledge Graph Completion (KGC) is a task aimed at uncovering the inherent relationships among known knowledge triplets in a Knowledge Graph (KG) and subsequently predicting missing links. Presently, there is a rising interest in inductive knowledge graph completion, where missing links may pertain to previously unobserved entities. Previous inductive KGC methods mainly rely on descriptive information of entities to improve the representation of unseen entities, neglecting to provide effective prior knowledge for relation modeling. To tackle this challenge, we capture prior schema-level interactions related to relations by leveraging entity type information, thereby furnishing effective prior constraints when reasoning with newly introduced entities. Moreover, We employ normal in-batch negatives and introduce schema-guided negatives to bolster the efficiency of normal contrastive representation learning. Experimental results demonstrate that our approach consistently achieves state-of-the-art performance on various established metrics across multiple benchmark datasets for link prediction. Notably, our method achieves a 20.5% relative increase in Hits@1 on the HumanWiki-Ind dataset. Ruilin Luo, Jiayi Li 0002, Jianghangfan Zhang, Jing Xiao 0006, Yujiu Yang 0001 |
LREC/COLING | 1 |
| 2024 | PTD-SQL: Partitioning and Targeted Drilling with LLMs in Text-to-SQLabstractLarge Language Models (LLMs) have emerged as powerful tools for Text-to-SQL tasks, exhibiting remarkable reasoning capabilities.Different from tasks such as math word problems and commonsense reasoning, SQL solutions have a relatively fixed pattern.This facilitates the investigation of whether LLMs can benefit from categorical thinking, mirroring how humans acquire knowledge through inductive reasoning based on comparable examples.In this study, we propose that employing query group partitioning allows LLMs to focus on learning the thought processes specific to a single problem type, consequently enhancing their reasoning abilities across diverse difficulty levels and problem categories.Our experiments reveal that multiple advanced LLMs, when equipped with PTD-SQL, can either surpass or match previous state-of-theart (SOTA) methods on the Spider and BIRD datasets.Intriguingly, models with varying initial performances have exhibited significant improvements, mainly at the boundary of their capabilities after targeted drilling, suggesting a parallel with human progress.Code is available at https://github.com/lrlbbzl/PTD-SQL. Ruilin Luo, Binghuai Lin, Zicheng Lin, Yujiu Yang 0001 |
EMNLP | 1 |
| 2024 | Prior Bilinear-Based Models for Knowledge Graph Completion
Jiayi Li 0002, Ruilin Luo, Jing Xiao 0006, Yujiu Yang 0001 |
ECML/PKDD (3) | 2 |