Ruochen Jiang

dblp:186/8193 · DBLP profile ↗
← Back
4ranked-venue papers
2as first author
3since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 4 · 2 first-author · 3 since 2021
YearPublicationVenuePosition
2025 ArrayMorph: Optimizing Hyperslab Queries on the Cloud for Machine Learning Pipelines
abstract
Cloud storage services such as Amazon S3, Azure Blob Storage, and Google Cloud Storage are widely used to store raw data for machine learning applications. When the data is later processed, the analysis predominantly focuses on regions of interest (such as a small bounding box in a larger image) and discards uninteresting regions. Machine learning applications can significantly accelerate their I/O if they push this data filtering step to the cloud. Prior work has proposed different methods to partially read array (tensor) objects, such as chunking, reading a contiguous byte range, and evaluating a lambda function. No method is optimal; estimating the total time and cost of a data retrieval requires an understanding of the data serialization order, the chunk size and platform-specific properties. This paper introduces ArrayMorph, a cloud-based array data storage system that automatically determines which is the best method to use to retrieve regions of interest from data on the cloud. ArrayMorph formulates data accesses as hyperslab queries, and optimizes them using a multi-phase cost-based approach. ArrayMorph seamlessly integrates with Python/PyTorch-based ML applications, and is experimentally shown to transfer up to 9.8X less data than existing systems. This makes ML applications run up to 1.7X faster and 9X cheaper than prior solutions.
Ruochen Jiang, Spyros Blanas
Proc. VLDB Endow.1
2022 Evaluation of Sustainable Economic and Environmental Development Evidence From OECD Countries
abstract
This study analyzes the economic and environmental performance of OECD countries over 2000–2019. A by-production approach is applied and the efficiency score is decomposed into its economic and environmental components. Unlike previous studies, we apply a refined model that allows for the correct modeling of by-production technology. The refined model can provide clear economic illustrations for balancing economic growth and environmental protection. The results indicate that environmental inefficiency is higher than the potential economic improvement. The environmental efficiency of OECD countries is improving, while economic performance is worsening over time. Therefore, instead of highly polluting energy, clean energy should be used to build a low-carbon economy. Worldwide, carbon-emitting countries and developed countries should shoulder their responsibilities to reduce carbon emissions and provide emission reduction funds for developing countries, while simultaneously sharing the green production technologies needed to reduce emissions.
Jingyu Zhou, Xingyu Xu 0004, Ruochen Jiang, Jinyang Cai
J. Glob. Inf. Manag.3
2021 Jigsaw: A Data Storage and Query Processing Engine for Irregular Table Partitioning
abstract
The physical data layout significantly impacts performance when database systems access cold data. In addition to the traditional row store and column store designs, recent research proposes to partition tables hierarchically, starting from either horizontal or vertical partitions and then determining the best partitioning strategy on the other dimension independently for each partition. All these partitioning strategies naturally produce rectangular partitions. Coarse-grained rectangular partitioning reads unnecessary data when a table cannot be partitioned along one dimension for all queries. Fine-grained rectangular partitioning produces many small partitions which negatively impacts I/O performance and possibly introduces a high tuple reconstruction overhead.
Donghe Kang, Ruochen Jiang, Spyros Blanas
SIGMOD Conference2
2020 Towards Extracting Highlights From Recorded Live Videos: An Implicit Crowdsourcing Approach
abstract
Live streaming platforms need to store a lot of recorded live videos on a daily basis. An important problem is how to automatically extract highlights (i.e., attractive short video clips) from these massive, long recorded live videos. However, algorithmic approaches are either domain-specific, which require experts to spend a long time to design, or resource-intensive, which require a lot of training data and/or computing resources. In this paper, we propose LIGHTOR, a novel implicit crowd-sourcing approach to overcome these limitations. The key insight is to collect users' natural interactions with a live streaming platform, and then leverage them to detect highlights. We recruit hundreds of users from Amazon Mechanical Turk, and evaluate the performance of LIGHTOR using two popular games in Twitch. The results show that LIGHTOR can achieve high extraction precision with a small set of training data and low computing resources.
Ruochen Jiang, Changbo Qu, Jiannan Wang 0001, Chi Wang 0001, Yudian Zheng
ICDE1