VLDB 2026 Research / reviewers in the wild / expert
Runyang You
dblp:402/4585
· DBLP profile ↗
3ranked-venue papers
2as first author
3since 2021 · last 2026
0009-0008-6018-1129ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 2 · 2 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Network and information security
1 paper |
Security and privacy of machine learning · 100% | |
| Artificial intelligence
3 papers |
Language models and text generation · 47% Reinforcement learning · 41% Vision and language · 12% | |
| Databases, data mining, and information retrieval
1 paper |
Recommender systems · 100% |
Topics — the 7 heaviest of 7, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation › large language model reasoning › inference-time reasoning
latent reasoning |
1.0 | 1 | 2026 | Parallel Test-Time Scaling for Latent Reasoning Models · ACL (1) 2026 |
Machine learning › Reinforcement learning
reinforcement learning for recommendation |
0.9 | 1 | 2025 | R$^2$ec: Towards Large Recommender Models with Reasoning · NeurIPS 2025 |
Recommender systems › explainable recommendation
reasoning-based recommendation |
0.9 | 1 | 2025 | R$^2$ec: Towards Large Recommender Models with Reasoning · NeurIPS 2025 |
Security and privacy of machine learning › large language model safety
jailbreak defense |
0.9 | 1 | 2025 | Towards Harmless Multimodal Assistants with Blind Preference Optimization · ACM Multimedia 2025 |
Security and privacy of machine learning › large language model safety
multimodal large language model safety |
0.9 | 1 | 2025 | Towards Harmless Multimodal Assistants with Blind Preference Optimization · ACM Multimedia 2025 |
Security and privacy of machine learning › large language model alignment
safety alignment |
0.9 | 1 | 2025 | Towards Harmless Multimodal Assistants with Blind Preference Optimization · ACM Multimedia 2025 |
Computer vision › Vision and language › vision-language model
multimodal large language model |
0.3 | 1 | 2025 | Towards Harmless Multimodal Assistants with Blind Preference Optimization · ACM Multimedia 2025 |
Methods — techniques the papers use, named apart from their topics
reinforcement learning · 1.7fused reward · 1.7dual-head architecture · 1.7direct preference optimization · 1.7blind preference optimization · 1.7monte carlo dropout · 1.0latent reward model · 1.0contrastive learning · 1.0additive gaussian noise · 1.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Parallel Test-Time Scaling for Latent Reasoning ModelsabstractParallel test-time scaling (TTS) is a pivotal approach for enhancing large language models (LLMs), typically by sampling multiple tokenbased chains-of-thought in parallel and aggregating outcomes through voting or search.Recent advances in latent reasoning, where intermediate reasoning unfolds in continuous vector spaces, offer a more efficient alternative to explicit Chain-of-Thought, yet whether such latent models can similarly benefit from parallel TTS remains open, mainly due to the absence of sampling mechanisms in continuous space, and the lack of probabilistic signals for advanced trajectory aggregation.This work enables parallel TTS for latent reasoning models by addressing the above issues.For sampling, we introduce two uncertainty-inspired stochastic strategies: Monte Carlo Dropout and Additive Gaussian Noise.For aggregation, we design a Latent Reward Model (LatentRM) trained with step-wise contrastive objective to score and guide latent reasoning.Extensive experiments and visualization analyses show that both sampling strategies scale effectively with compute and exhibit distinct exploration dynamics, while LatentRM enables effective trajectory selection.Together, our explorations open a new direction for scalable inference in continuous spaces. Runyang You, Yongqi Li 0001, Meng Liu 0006, Wenjie Wang 0007, Liqiang Nie, Wenjie Li 0002 |
ACL (1) | 1 |
| 2025 | Towards Harmless Multimodal Assistants with Blind Preference OptimizationabstractMultimodal Large Language Models (MLLMs) have demonstrated impressive capabilities in multimodal understanding, reasoning, and interaction. Given the extensive applications of MLLMs, the associated safety issues have become increasingly critical. Due to the effectiveness of preference optimization in aligning MLLMs with human preferences, there is an urgent need for safety-related preference data for MLLMs. To address this, we construct the MMSafe-PO preference dataset towards harmless multimodal assistants, featuring multimodal instructions, the conversational format, and ranked paired responses from human feedback. We also identify two insightful observations: modality co-defense and modality cheating, which illustrate that MLLMs possess a certain level of inherent defense while still presenting unique safety challenges. Based on these observations, we propose the Blind Preference Optimization (BPO) approach. Comprehensive experiments on three benchmarks show that BPO effectively enhances the safety capabilities of MLLMs. Notably, BPO significantly improves the safety rate of the base MLLM by 45.0%, outperforming the DPO approach. Additionally, applying BPO to the MMSafe-PO dataset greatly reduces the base MLLM's unsafe rate on other safety benchmarks (14.5% on MM-SafetyBench and 82.9% on HarmEval), demonstrating the effectiveness and robustness of both the dataset and the approach. Yongqi Li 0001, Lu Yang 0008, Jian Wang 0054, Runyang You, Wenjie Li 0002, Liqiang Nie |
ACM Multimedia | 4 |
| 2025 | R$^2$ec: Towards Large Recommender Models with ReasoningabstractLarge recommender models have extended LLMs as powerful recommenders via encoding or item generation, and recent breakthroughs in LLM reasoning synchronously motivate the exploration of reasoning in recommendation.
In this work, we propose R$^2$ec, a unified large recommender model with intrinsic reasoning capability.
R$^2$ec introduces a dual-head architecture that supports both reasoning chain generation and efficient item prediction in a single model, significantly reducing inference latency. To overcome the lack of annotated reasoning data, we design RecPO, a reinforcement learning framework that optimizes reasoning and recommendation jointly with a novel fused reward mechanism.
Extensive experiments on three datasets demonstrate that R$^2$ec outperforms traditional, LLM-based, and reasoning-augmented recommender baselines, while further analyses validate its competitive efficiency among conventional LLM-based recommender baselines
and strong adaptability to diverse recommendation scenarios. Code and checkpoints available at https://github.com/YRYangang/RRec. Runyang You, Yongqi Li 0001, Xinyu Lin 0001, Xin Zhang 0097, Wenjie Wang 0007, Wenjie Li 0002, Liqiang Nie |
NeurIPS | 1 |