Victor Rühle

dblp:277/8100 · DBLP profile ↗
← Back
8ranked-venue papers
0as first author
7since 2021 · last 2026
0000-0002-8957-7628ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 5 since 2021Systems, architecture and hardware · 1 · 1 since 2021Security and privacy · 1Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Language models and text generation · 65% Efficient and distributed learning · 35%
Network and information security
2 papers
Privacy and data protection · 71% Security and privacy of machine learning · 29%
Databases, data mining, and information retrieval
1 paper
Information retrieval · 100%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Cloud and datacenter computing · 100%

Topics — the 10 heaviest of 12, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Information retrieval
retrieval-augmented generation
1.012026
Cost-Aware Retrieval-Augmentation Reasoning Models with Adaptive Retrieval Depth · WWW 2026
Natural language and speech › Language models and text generation › large language model evaluation › capability evaluation
long-context evaluation
0.912025
Minerva: A Programmable Memory Test Benchmark for Language Models · ICML 2025
Natural language and speech › Language models and text generation
model routing
0.912025
BEST-Route: Adaptive LLM Routing with Test-Time Optimal Compute · ICML 2025
Natural language and speech › Language models and text generation › large language model inference
hybrid inference
0.812024
Hybrid LLM: Cost-Efficient and Quality-Aware Query Routing · ICLR 2024
Machine learning › Efficient and distributed learning › inference serving
large language model serving
0.812024
Hybrid LLM: Cost-Efficient and Quality-Aware Query Routing · ICLR 2024
Privacy and data protection
differential privacy
0.712023
Bayesian Estimation of Differential Privacy · ICML 2023
Cloud and datacenter computing
resource management
0.712023
Snape: Reliable and Low-Cost Computing with Mixture of Spot and On-Demand VMs · ASPLOS (3) 2023
Security and privacy of machine learning
membership inference
0.622023
Analyzing Information Leakage of Updates to Natural Language Models · CCS 2020
Bayesian Estimation of Differential Privacy · ICML 2023
Privacy and data protection › information leakage
training data leakage
0.412020
Analyzing Information Leakage of Updates to Natural Language Models · CCS 2020
Natural language and speech › Language models and text generation
large language model inference
0.312025
BEST-Route: Adaptive LLM Routing with Test-Time Optimal Compute · ICML 2025

Methods — techniques the papers use, named apart from their topics

retrieval augmentation · 1.0test-time compute scaling · 0.9response sampling · 0.9programmatic benchmark generation · 0.9difficulty prediction · 0.8differentially private SGD · 0.7constrained reinforcement learning · 0.7bayesian posterior estimation · 0.7differential score · 0.4differential rank · 0.4
YearPublicationVenuePosition
2026 Cost-Aware Retrieval-Augmentation Reasoning Models with Adaptive Retrieval Depth
Helia Hashemi, Victor Rühle, Saravan Rajmohan
WWW2
2025 BEST-Route: Adaptive LLM Routing with Test-Time Optimal Compute
abstract
Large language models (LLMs) are powerful tools but are often expensive to deploy at scale. LLM query routing mitigates this by dynamically assigning queries to models of varying cost and quality to obtain a desired tradeoff. Prior query routing approaches generate only one response from the selected model and a single response from a small (inexpensive) model was often not good enough to beat a response from a large (expensive) model due to which they end up overusing the large model and missing out on potential cost savings. However, it is well known that for small models, generating multiple responses and selecting the best can enhance quality while remaining cheaper than a single large-model response. We leverage this idea to propose BEST-Route, a novel routing framework that chooses a model and the number of responses to sample from it based on query difficulty and the quality thresholds. Experiments on real-world datasets demonstrate that our method reduces costs by up to 60% with less than 1% performance drop.
Dujian Ding, Ankur Mallick, Chi Wang 0001, Daniel Madrigal 0001, Mirian Hipolito Garcia, Menglin Xia, Laks V. S. Lakshmanan, Qingyun Wu, Victor Rühle
ICML10
2025 Minerva: A Programmable Memory Test Benchmark for Language Models
abstract
How effectively can LLM-based AI assistants utilize their memory (context) to perform various tasks? Traditional data benchmarks, which are often manually crafted, suffer from several limitations: they are static, susceptible to overfitting, difficult to interpret, and lack actionable insights–failing to pinpoint the specific capabilities a model lacks when it does not pass a test. In this paper, we present a framework for automatically generating a comprehensive set of tests to evaluate models’ abilities to use their memory effectively. Our framework extends the range of capability tests beyond the commonly explored (passkey, key-value, needle in the haystack) search, a dominant focus in the literature. Specifically, we evaluate models on atomic tasks such as searching, recalling, editing, matching, comparing information in context memory, performing basic operations when inputs are structured into distinct blocks, and maintaining state while operating on memory, simulating real-world data. Additionally, we design composite tests to investigate the models’ ability to perform more complex, integrated tasks. Our benchmark enables an interpretable, detailed assessment of memory capabilities of LLMs.
Menglin Xia, Victor Rühle, Saravan Rajmohan, Reza Shokri
ICML2
2024 Hybrid LLM: Cost-Efficient and Quality-Aware Query Routing
abstract
Large language models (LLMs) excel in most NLP tasks but also require expensive cloud servers for deployment due to their size, while smaller models that can be deployed on lower cost (e.g., edge) devices, tend to lag behind in terms of response quality. Therefore in this work we propose a hybrid inference approach which combines their respective strengths to save cost and maintain quality. Our approach uses a router that assigns queries to the small or large model based on the predicted query difficulty and the desired quality level. The desired quality level can be tuned dynamically at test time to seamlessly trade quality for cost as per the scenario requirements. In experiments our approach allows us to make up to 40% fewer calls to the large model, with no drop in response quality.
Dujian Ding, Ankur Mallick, Chi Wang 0001, Robert Sim, Subhabrata Mukherjee, Victor Rühle, Laks V. S. Lakshmanan, Ahmed Awadallah 0001
ICLR6
2023 Snape: Reliable and Low-Cost Computing with Mixture of Spot and On-Demand VMs
abstract
Cloud providers often have resources that are not being fully utilized, and they may offer them at a lower cost to make up for the reduced availability of these resources. However, customers may be hesitant to use such offerings (such as spot VMs) as making trade-offs between cost and resource availability is not always straightforward. In this work, we propose Snape (Spot On-demand Perfect Mixture), an intelligent framework to optimize the cost and resource availability by dynamically mixing on-demand VMs with spot VMs. Through a detailed characterization based on real production traces, we verify that the eviction of spot VMs is predictable to some extent. Snape also leverages constrained reinforcement learning to adjust the mixture policy online. Experiments across different configurations show that Snape achieves 44% savings compared to using only on-demand VMs while maintaining 99.96% availability, which is 2.77% higher than using only spot VMs.
Fangkai Yang, Lu Wang 0029, Zhenyu Xu 0003, Liqun Li, Bo Qiao 0001, Camille Couturier, Chetan Bansal, Soumya Ram, Si Qin, Íñigo Goiri, Eli Cortez, Terry Yang, Victor Rühle, Saravan Rajmohan, Qingwei Lin, Dongmei Zhang 0001
ASPLOS (3)15
2023 Bayesian Estimation of Differential Privacy
abstract
Algorithms such as Differentially Private SGD enable training machine learning models with formal privacy guarantees. However, because these guarantees hold with respect to unrealistic adversaries, the protection afforded against practical attacks is typically much better. An emerging strand of work empirically estimates the protection afforded by differentially private training as a confidence interval for the privacy budget $\hat{\varepsilon}$ spent with respect to specific threat models. Existing approaches derive confidence intervals for $\hat{\varepsilon}$ from confidence intervals for false positive and false negative rates of membership inference attacks, which requires training an impractically large number of models to get intervals that can be acted upon. We propose a novel, more efficient Bayesian approach that brings privacy estimates within the reach of practitioners. Our approach reduces sample size by computing a posterior for $\hat{\varepsilon}$ (not just a confidence interval) from the joint posterior of the false positive and false negative rates of membership inference attacks. We implement an end-to-end system for privacy estimation that integrates our approach and state-of-the-art membership inference attacks, and evaluate it on text and vision classification tasks. For the same number of samples, we see a reduction in interval width of up to 40% compared to prior work.
Santiago Zanella-Béguelin, Lukas Wutschitz, Shruti Tople, Ahmed Salem 0001, Victor Rühle, Andrew Paverd, Mohammad Naseri, Boris Köpf
ICML5
2021 Privacy Regularization: Joint Privacy-Utility Optimization in LanguageModels
abstract
Fatemehsadat Mireshghallah, Huseyin Inan, Marcello Hasegawa, Victor Rühle, Taylor Berg-Kirkpatrick, Robert Sim. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Niloofar Mireshghallah, Huseyin A. Inan, Marcello Hasegawa, Victor Rühle, Taylor Berg-Kirkpatrick, Robert Sim
NAACL-HLT4
2020 Analyzing Information Leakage of Updates to Natural Language Models
abstract
To continuously improve quality and reflect changes in data, machine learning applications have to regularly retrain and update their core models. We show that a differential analysis of language model snapshots before and after an update can reveal a surprising amount of detailed information about changes in the training data. We propose two new metrics---differential score and differential rank---for analyzing the leakage due to updates of natural language models. We perform leakage analysis using these metrics across models trained on several different datasets using different methods and configurations. We discuss the privacy implications of our findings, propose mitigation strategies and evaluate their effect.
Santiago Zanella-Béguelin, Lukas Wutschitz, Shruti Tople, Victor Rühle, Andrew Paverd, Olga Ohrimenko, Boris Köpf, Marc Brockschmidt
CCS4