VLDB 2026 Research / reviewers in the wild / expert
Rui Liu 0002
dblp:42/469-2
· DBLP profile ↗
15ranked-venue papers
8as first author
6since 2021 · last 2024
0000-0001-9061-787XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 6 · 4 first-author · 4 since 2021Systems, architecture and hardware · 3 · 1 first-author · 2 since 2021Security and privacy · 2Software engineering, systems software and programming languages · 2 · 2 first-authorComputer networks · 1Human-computer interaction and ubiquitous computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Towards Resource-adaptive Query Execution in Cloud Native Databases
Rui Liu 0002, Jun Hyuk Chang, Riki Otaki, Zhe Heng Eng, Aaron J. Elmore, Michael J. Franklin, Sanjay Krishnan |
CIDR | 1 |
| 2024 | Riveter: Adaptive Query Suspension and Resumption Framework for Cloud Native DatabasesabstractIn modern cloud environments, ephemeral resources with intermittent availability and fluctuating monetary costs are becoming common. This dynamic nature presents a new challenge when deploying cloud-native databases: adaptive query execution, which can suspend queries when the resources are scarce or costs unexpectedly soar, and then resume them when the resources become available or cost-effective. Addressing this challenge requires the design and implementation of query suspension and resumption with a mechanism that can adaptively determine when, if, and how to suspend queries. In this paper, we propose Riveter, a query suspension and resumption framework that can adaptively pause ongoing queries using various strategies, including (1) a redo strategy that terminates queries and subsequently re-runs them, (2) a pipeline-level strategy that suspends a query once one of its pipelines has completed to reduce the storage requirements for intermediate data, (3) and a process-level strategy that enables the suspension of query execution processes at any given moment but generates a substantial volume of intermediate data for query resumption. We also devise a cost model to estimate query latency using various strategies and an algorithm to select the one that causes minimum latency. To demonstrate the effectiveness of Riveter, we conduct evaluations based on the TPC-H benchmark to investigate intermediate data persistence, strategy selection, and cost model-based estimation. Our results not only present the difference among the strategies of Riveter in terms of the size of persisted intermediate data and the time of triggering the suspension but also confirm the adaptive and efficient query suspension and resumption delivered by Riveter. Rui Liu 0002, Aaron J. Elmore, Michael J. Franklin, Sanjay Krishnan |
ICDE | 1 |
| 2023 | Rotary: A Resource Arbitration Framework for Progressive Iterative AnalyticsabstractIncreasingly modern computing applications employ progressive iterative analytics, as best exemplified by two prevalent cases, approximate query processing (AQP) and deep learning training (DLT). In comparison to classic computing applications that only return the results after processing all the input data, progressive iterative analytics keep providing approximate or partial results to users by performing computations on a subset of the entire dataset until either the users are satisfied with the results, or the predefined completion criteria are achieved. Typically, progressive iterative analytic jobs have various completion criteria, produce diminishing returns, and process data at different rates, which necessitates a novel resource arbitration that can continuously prioritize the progressive iterative analytic jobs and determine if/when to reallocate and preempt the resources. We propose and design a resource arbitration framework, Rotary, and implement two resource arbitration systems, Rotary-AQP and Rotary-DLT, for approximate query processing and deep learning training. We build a TPC-H based AQP workload and a survey-based DLT workload to evaluate the two systems, respectively. The evaluation results demonstrate that Rotary-AQP and Rotary-DLT outperform the state-of-the-art systems and confirm the generality and practicality of the proposed resource arbitration framework. Rui Liu 0002, Aaron J. Elmore, Michael J. Franklin, Sanjay Krishnan |
ICDE | 1 |
| 2023 | Optimizing Data Pipelines for Machine Learning in Feature StoresabstractData pipelines (i.e., converting raw data to features) are critical for machine learning (ML) models, yet their development and management is time-consuming. Feature stores have recently emerged as a new "DBMS-for-ML" with the premise of enabling data scientists and engineers to define and manage their data pipelines. While current feature stores fulfill their promise from a functionality perspective, they are resource-hungry---with ample opportunities for implementing database-style optimizations to enhance their performance. In this paper, we propose a novel set of optimizations specifically targeted for point-in-time join, which is a critical operation in data pipelines. We implement these optimizations on top of Feathr: a widely-used feature store, and evaluate them on use cases from both the TPCx-AI benchmark and real-world online retail scenarios. Our thorough experimental analysis shows that our optimizations can accelerate data pipelines by up to 3× over state-of-the-art baselines. Rui Liu 0002, Kwanghyun Park 0001, Fotis Psallidas, Xiaoyong Zhu, Jinghui Mo, Rathijit Sen, Matteo Interlandi, Konstantinos Karanasos, Yuanyuan Tian 0001, Jesús Camacho-Rodríguez |
Proc. VLDB Endow. | 1 |
| 2022 | Bring orders into uncertainty: enabling efficient uncertain graph processing via novel path sampling on multi-accelerator systemsabstractUncertain or probabilistic graphs have been ubiquitously used to represent noisy, incomplete, and inaccurate linked data in many emerging big-data mining and analytics applications. It is impractical to solve uncertain graph problems exactly as it requires to evaluate an exponential number of certain instances (or "possible worlds") generated from an uncertain graph. Previously, several CPU-based techniques were proposed to use sampling for uncertain graph processing. However, we observe that (1) they suffer from low computation efficiency and large memory overhead due to unnecessary edge sampling at runtime; (2) they cannot leverage the massive parallelism provided by modern general-purpose accelerators; and (3) there lacks a general programming framework for high-performance uncertain graph processing. To tackle these challenges, we propose a novel runtime path sampling method, which is able to identify and eliminate unnecessary edge sampling via incremental path identification and filtering, resulting in significant reduction in computation and data movement. Centered around this idea, we introduce a general uncertain graph processing framework for multi-GPU systems, named BPGraph1. BPGraph provides general support for users to design and optimize a wide-range of uncertain graph algorithms and applications without concerning about the underlying complexity. Extensive evaluation on a variety of real-world uncertain graph applications demonstrates an average speedup of 26X (up to 43X) and better scalability from BPGraph over the state-of-the-art frameworks. Heng Zhang 0005, Lingda Li, Hang Liu 0001, Donglin Zhuang, Rui Liu 0002, Chengying Huan, Dingwen Tao, Yongchao Liu 0004, Charles He, Shuaiwen Song |
ICS | 5 |
| 2021 | An efficient uncertain graph processing framework for heterogeneous architecturesabstractUncertain or probabilistic graphs have been ubiquitously used in many emerging applications. Previously CPU based techniques were proposed to use sampling but suffer from (1) low computation efficiency and large memory overhead, (2) low degree of parallelism, and (3) nonexistent general framework to effectively support programming uncertain graph applications. To tackle these challenges, we propose a general uncertain graph processing framework for multi-GPU systems, named BPGraph. Integrated with our highly-efficient path sampling method, BPGraph can support a wide range of uncertain graph algorithms' development and optimization. Extensive evaluation demonstrates a significant performance improvement from BPGraph over the state-of-the-art uncertain graph sampling techniques. Heng Zhang 0005, Lingda Li, Donglin Zhuang, Rui Liu 0002, Dingwen Tao, Shuaiwen Song |
PPoPP | 4 |
| 2019 | Participant Incentive Mechanism Toward Quality-Oriented Sensing: Understanding and ApplicationabstractThe ubiquity of ever-more-capable mobile devices, especially smartphones, brings forth participatory sensing to collect and interpret information. It can achieve unprecedented quantity of data. However, it is arduous to guarantee quality of data because everyone can contribute data without scrutinization. It is an important issue in quality-oriented participatory sensing. Our idea to address this issue is motivating participants to contribute accurate data for improving data quality directly. In this article, we propose a reputation-based incentive mechanism, RIM, to realize the idea. More specifically, we identify the participants who collect the accurate data and regard them as the reputable ones. Then, the reputable participants are granted a higher chance to obtain rewards so that other people will try to follow such users and become reputable as well. Namely, RIM can encourage and steer users to collect accurate data in the long term. We analyze our incentive mechanism by formalization and premise implications. For a feasibility study of participatory sensing and verification of the implications, we implement and deploy a participatory sensing application focusing on monitoring environmental noise in a specific location as a case study and conduct a simulation based on the case study to further evaluate the proposed incentive mechanism. The results from the case study and the simulation present that RIM can remarkably increase the quality of collected data in participatory sensing while corroborating our theoretical implications. Ruiyun Yu, Jiannong Cao 0001, Rui Liu 0002, Wenyu Gao, Xingwei Wang 0001, Junbin Liang |
ACM Trans. Sens. Networks | 3 |
| 2019 | Understanding Mobile Users' Privacy Expectations: A Recommendation-Based Method Through CrowdsourcingabstractPrivacy is a pivotal issue of mobile apps because there is a plethora of personal and sensitive information in smartphones. Many mechanisms and tools are proposed to detect and mitigate privacy leaks. However, they rarely consider users' preferences and expectations. Users hold various expectation towards different mobile apps. For example, users may allow a social app to access their photos rather than a game app because it goes beyond users' expectation to access personal photos. Therefore, we believe it is practical and beneficial to understand users' privacy expectations on various mobile apps and help them mitigate privacy risks introduced by smartphones. To achieve this objective, we propose and implement PriWe, a system based on crowdsourcing driven by users who contribute privacy permission settings of the apps installed on their smartphones. PriWe leverages the crowdsourced permission settings to understand users' privacy expectations and provides app specific recommendations to mitigate information leakage. We deployed PriWe in the real world for evaluation. According to the feedback of 78 users who evaluated our system and 422 participants who completed our survey, PriWe is able to make proper recommendations which can match participants' privacy expectations and are mostly accepted by users, thereby help them to mitigate privacy disclosure in smartphones. Rui Liu 0002, Junbin Liang, Jiannong Cao 0001, Kehuan Zhang, Wenyu Gao, Lei Yang 0024, Ruiyun Yu |
IEEE Trans. Serv. Comput. | 1 |
| 2018 | Privacy-based recommendation mechanism in mobile participatory sensing systems using crowdsourced users' preferences
Rui Liu 0002, Junbin Liang, Wenyu Gao, Ruiyun Yu |
Future Gener. Comput. Syst. | 1 |
| 2018 | Accessing mobile user's privacy based on IME personalization: Understanding and practical attacksabstractInput Method Editor (IME) is an indispensable component on current smartphones. With its assistance, the number of key presses is reduced, and non-Latin characters could be inputted. Furthermore, modern IMEs integrate several personalized features like reordering suggestion lists and predicting the next words based on user’s input history. Such optimization improves the user experience but turns the IME dictionary into a pool of user privacy. Previous works have discussed the privacy risks coming from malicious IMEs. Indeed, they could cause security and privacy issues if installed by common users, but their impact is limited as the majority of IMEs are well-behaved. However, whether legitimate IMEs are bullet-proof is not answered before. In this paper, we make the first attempt to study the security implications of IME personalization and the back-end infrastructure on Android devices. In the end, we identify a critical vulnerability lying under the Android KeyEvent processing framework, which can be exploited to launch cross-app KeyEvent injection (CAKI) attack and bypass the app-isolation mechanism. By abusing such design flaw, an adversary can harvest entries from the personalized user dictionary of IME through an ostensibly innocuous app only asking for common permissions. Our evaluation over a broad spectrum of Android OSes, devices, and IMEs suggests such issue should be fixed immediately. All Android versions we examined (from very old 2.3.4 to the latest 6.0.1) and most IME apps we surveyed (11 out of 18) are vulnerable. User’s private information, like contact names, location, etc., can be easily exfiltrated. Up to hundreds of millions of mobile users are under this threat. To mitigate this security issue, we propose a practical defense mechanism which augments the existing KeyEvent processing framework without forcing any change to IME apps. Wenrui Diao, Rui Liu 0002, Zhe Zhou 0001, Zhou Li 0001, Kehuan Zhang |
J. Comput. Secur. | 2 |
| 2018 | Untangling Blockchain: A Data Processing View of Blockchain SystemsabstractBlockchain technologies are gaining massive momentum in the last few years. Blockchains are distributed ledgers that enable parties who do not fully trust each other to maintain a set of global states. The parties agree on the existence, values, and histories of the states. As the technology landscape is expanding rapidly, it is both important and challenging to have a firm grasp of what the core technologies have to offer, especially with respect to their data processing capabilities. In this paper, we first survey the state of the art, focusing on private blockchains (in which parties are authenticated). We analyze both in-production and research systems in four dimensions: distributed ledger, cryptography, consensus protocol, and smart contract. We then present BLOCKBENCH, a benchmarking framework for understanding performance of private blockchains against data processing workloads. We conduct a comprehensive evaluation of three major blockchain systems based on BLOCKBENCH, namely Ethereum, Parity, and Hyperledger Fabric. The results demonstrate several trade-offs in the design space, as well as big performance gaps between blockchain and database systems. Drawing from design principles of database systems, we discuss several research directions for bringing blockchain performance closer to the realm of databases. Tien Tuan Anh Dinh, Rui Liu 0002, Meihui Zhang 0001, Gang Chen 0001, Beng Chin Ooi, Ji Wang 0006 |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2018 | When Privacy Meets Usability: Unobtrusive Privacy Permission Recommendation System for Mobile Apps Based on CrowdsourcingabstractPeople nowadays almost want everything at their fingertips, from business to entertainment, and meanwhile they do not want to leak their sensitive data. Strong information protection can be a competitive advantage, but preserving privacy is a real challenge when people use the mobile apps in the smartphone. If they are too lax with privacy preserving, important or sensitive information could be lost. If they are too tight with privacy, making users jump through endless hoops to access the data they need to get their work done, productivity can nosedive. Thus, striking a balance between privacy and usability in mobile applications can be difficult. Leveraging the privacy permission settings in mobile operating systems, our basic idea to address this issue is to provide proper recommendations about the settings so that the users can preserve their sensitive information and maintain the usability of apps. In this paper, we propose an unobtrusive recommendation system to implement this idea, which can crowdsource users' privacy permission settings and generate the recommendations for them accordingly. Besides, our system allows users to provide feedback to revise the recommendations for getting better performance and adapting different scenarios. For the evaluation, we collected users' preferences from 382 participants on Amazon Technical Turks and released our system to users in the real world for 10 days. According to the study, our system can make appropriate recommendations which can meet participants' privacy expectation and mobile apps' usability. Rui Liu 0002, Jiannong Cao 0001, Kehuan Zhang, Wenyu Gao, Junbin Liang, Lei Yang 0024 |
IEEE Trans. Serv. Comput. | 1 |
| 2017 | BLOCKBENCH: A Framework for Analyzing Private BlockchainsabstractBlockchain technologies are taking the world by storm. Public blockchains, such as Bitcoin and Ethereum, enable secure peer-to-peer applications like crypto-currency or smart contracts. Their security and performance are well studied. This paper concerns recent private blockchain systems designed with stronger security (trust) assumption and performance requirement. These systems target and aim to disrupt applications which have so far been implemented on top of database systems, for example banking, finance and trading applications. Multiple platforms for private blockchains are being actively developed and fine tuned. However, there is a clear lack of a systematic framework with which different systems can be analyzed and compared against each other. Such a framework can be used to assess blockchains' viability as another distributed data processing platform, while helping developers to identify bottlenecks and accordingly improve their platforms. Tien Tuan Anh Dinh, Ji Wang 0006, Gang Chen 0001, Rui Liu 0002, Beng Chin Ooi, Kian-Lee Tan |
SIGMOD Conference | 4 |
| 2017 | Vulnerable GPU Memory Management: Towards Recovering Raw Data from GPUabstractAbstract According to previous reports, information could be leaked from GPU memory; however, the security implications of such a threat were mostly over-looked, because only limited information could be indirectly extracted through side-channel attacks. In this paper, we propose a novel algorithm for recovering raw data directly from the GPU memory residues of many popular applications such as Google Chrome and Adobe PDF reader. Our algorithm enables harvesting highly sensitive information including credit card numbers and email contents from GPU memory residues. Evaluation results also indicate that nearly all GPU-accelerated applications are vulnerable to such attacks, and adversaries can launch attacks without requiring any special privileges both on traditional multi-user operating systems, and emerging cloud computing scenarios. Zhe Zhou 0001, Wenrui Diao, Zhou Li 0001, Kehuan Zhang, Rui Liu 0002 |
Proc. Priv. Enhancing Technol. | 6 |
| 2016 | PriMe: Human-centric privacy measurement based on user preferences towards data sharing in mobile participatory sensing systemsabstractMobile participatory sensing systems allow people with mobile devices to collect, interpret, and share data from their respective environments. One of the main obstacles for long-term participation in such systems is the users' privacy concerns. Due to the nature of these systems, users have to agree to provide some personalized information. Typically, however, people are reluctant to share any information, as it may be sensitive. This is especially the case if the content of the data in question is not completely transparent. In order to increase users' willingness to participate in such systems, we should help users identify which data they can share without violating their personal privacy policies. However, the perception of how sensitive a piece of information is may differ from user to user. In this paper, we propose the human-centric privacy measurement method PriMe, which quantifies privacy risks based on user preferences towards data sharing in participatory sensing systems. Further, we implemented and deployed PriMe in the real world as a user study for evaluation. The study shows that PriMe provides accurate ratings that fit users' individual perceptions of privacy, and is accepted by users as a trustworthy tool. Rui Liu 0002, Jiannong Cao 0001, Sebastian VanSyckel, Wenyu Gao |
PerCom | 1 |