Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Runyang You

dblp:402/4585 · DBLP profile ↗
← Back
3ranked-venue papers
2as first author
3since 2021 · last 2026
0009-0008-6018-1129ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 2 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Network and information security
1 paper
Security and privacy of machine learning · 100%
Artificial intelligence
3 papers
Language models and text generation · 47% Reinforcement learning · 41% Vision and language · 12%
Databases, data mining, and information retrieval
1 paper
Recommender systems · 100%

Topics — the 7 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation › large language model reasoning › inference-time reasoning
latent reasoning
1.012026
Parallel Test-Time Scaling for Latent Reasoning Models · ACL (1) 2026
Machine learning › Reinforcement learning
reinforcement learning for recommendation
0.912025
R$^2$ec: Towards Large Recommender Models with Reasoning · NeurIPS 2025
Recommender systems › explainable recommendation
reasoning-based recommendation
0.912025
R$^2$ec: Towards Large Recommender Models with Reasoning · NeurIPS 2025
Security and privacy of machine learning › large language model safety
jailbreak defense
0.912025
Towards Harmless Multimodal Assistants with Blind Preference Optimization · ACM Multimedia 2025
Security and privacy of machine learning › large language model safety
multimodal large language model safety
0.912025
Towards Harmless Multimodal Assistants with Blind Preference Optimization · ACM Multimedia 2025
Security and privacy of machine learning › large language model alignment
safety alignment
0.912025
Towards Harmless Multimodal Assistants with Blind Preference Optimization · ACM Multimedia 2025
Computer vision › Vision and language › vision-language model
multimodal large language model
0.312025
Towards Harmless Multimodal Assistants with Blind Preference Optimization · ACM Multimedia 2025

Methods — techniques the papers use, named apart from their topics

reinforcement learning · 1.7fused reward · 1.7dual-head architecture · 1.7direct preference optimization · 1.7blind preference optimization · 1.7monte carlo dropout · 1.0latent reward model · 1.0contrastive learning · 1.0additive gaussian noise · 1.0
YearPublicationVenuePosition
2026 Parallel Test-Time Scaling for Latent Reasoning Models
abstract
Parallel test-time scaling (TTS) is a pivotal approach for enhancing large language models (LLMs), typically by sampling multiple tokenbased chains-of-thought in parallel and aggregating outcomes through voting or search.Recent advances in latent reasoning, where intermediate reasoning unfolds in continuous vector spaces, offer a more efficient alternative to explicit Chain-of-Thought, yet whether such latent models can similarly benefit from parallel TTS remains open, mainly due to the absence of sampling mechanisms in continuous space, and the lack of probabilistic signals for advanced trajectory aggregation.This work enables parallel TTS for latent reasoning models by addressing the above issues.For sampling, we introduce two uncertainty-inspired stochastic strategies: Monte Carlo Dropout and Additive Gaussian Noise.For aggregation, we design a Latent Reward Model (LatentRM) trained with step-wise contrastive objective to score and guide latent reasoning.Extensive experiments and visualization analyses show that both sampling strategies scale effectively with compute and exhibit distinct exploration dynamics, while LatentRM enables effective trajectory selection.Together, our explorations open a new direction for scalable inference in continuous spaces.
Runyang You, Yongqi Li 0001, Meng Liu 0006, Wenjie Wang 0007, Liqiang Nie, Wenjie Li 0002
ACL (1)1
2025 Towards Harmless Multimodal Assistants with Blind Preference Optimization
abstract
Multimodal Large Language Models (MLLMs) have demonstrated impressive capabilities in multimodal understanding, reasoning, and interaction. Given the extensive applications of MLLMs, the associated safety issues have become increasingly critical. Due to the effectiveness of preference optimization in aligning MLLMs with human preferences, there is an urgent need for safety-related preference data for MLLMs. To address this, we construct the MMSafe-PO preference dataset towards harmless multimodal assistants, featuring multimodal instructions, the conversational format, and ranked paired responses from human feedback. We also identify two insightful observations: modality co-defense and modality cheating, which illustrate that MLLMs possess a certain level of inherent defense while still presenting unique safety challenges. Based on these observations, we propose the Blind Preference Optimization (BPO) approach. Comprehensive experiments on three benchmarks show that BPO effectively enhances the safety capabilities of MLLMs. Notably, BPO significantly improves the safety rate of the base MLLM by 45.0%, outperforming the DPO approach. Additionally, applying BPO to the MMSafe-PO dataset greatly reduces the base MLLM's unsafe rate on other safety benchmarks (14.5% on MM-SafetyBench and 82.9% on HarmEval), demonstrating the effectiveness and robustness of both the dataset and the approach.
Yongqi Li 0001, Lu Yang 0008, Jian Wang 0054, Runyang You, Wenjie Li 0002, Liqiang Nie
ACM Multimedia4
2025 R$^2$ec: Towards Large Recommender Models with Reasoning
abstract
Large recommender models have extended LLMs as powerful recommenders via encoding or item generation, and recent breakthroughs in LLM reasoning synchronously motivate the exploration of reasoning in recommendation. In this work, we propose R$^2$ec, a unified large recommender model with intrinsic reasoning capability. R$^2$ec introduces a dual-head architecture that supports both reasoning chain generation and efficient item prediction in a single model, significantly reducing inference latency. To overcome the lack of annotated reasoning data, we design RecPO, a reinforcement learning framework that optimizes reasoning and recommendation jointly with a novel fused reward mechanism. Extensive experiments on three datasets demonstrate that R$^2$ec outperforms traditional, LLM-based, and reasoning-augmented recommender baselines, while further analyses validate its competitive efficiency among conventional LLM-based recommender baselines and strong adaptability to diverse recommendation scenarios. Code and checkpoints available at https://github.com/YRYangang/RRec.
Runyang You, Yongqi Li 0001, Xinyu Lin 0001, Xin Zhang 0097, Wenjie Wang 0007, Wenjie Li 0002, Liqiang Nie
NeurIPS1