Yinhao Hong

dblp:325/2497 · DBLP profile ↗
← Back
4ranked-venue papers
1as first author
4since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 3 · 1 first-author · 3 since 2021
YearPublicationVenuePosition
2025 A Hybrid Approach to Integrating Deterministic and Non-deterministic Concurrency Control in Database Systems
abstract
Deterministic and non-deterministic concurrency control algorithms have shown respective advantages under diverse workloads. Thus, a natural idea is to blend them together. However, because deterministic algorithms work with stringent assumptions, e.g., batched execution and non-interactive transactions, they hardly work together with non-deterministic algorithms. To address this issue, we propose HDCC, a hybrid approach that adaptively employs Calvin and OCC, which have distinct concurrency control and logging schemes, in the same database system. To ensure serializability and recovery correctness, we introduce lock-sharing, global validation, and two-log-interleaving mechanisms. Additionally, we introduce a rule-based assignment mechanism to dynamically select Calvin or OCC based on workload characteristics. Experimental results using TPC-C and YCSB benchmarks demonstrate that HDCC surpasses existing hybrid approaches by up to 3.1×.
Yinhao Hong, Hongyao Zhao, Wei Lu 0015, Xiaoyong Du 0001, Yuxing Chen 0003, Anqun Pan, Lixiong Zheng
Proc. VLDB Endow.1
2025 An Efficient Two-Round Distributed Transaction Processing Approach over Heterogeneous Networks
Hongyao Zhao, Wei Lu 0015, Zhanhao Zhao, Yinhao Hong, Quanqing Xu, Jinliang Xiao, Fusheng Han, Chuanhui Yang, Xiaoyong Du 0001
VLDB J.4
2022 Taming the Big Data Monster: Managing Petabytes of Data with Multi-Model Databases
abstract
With the development of big data technology, the amount of business data that Internet companies need to handle has reached the petabyte level, which poses great pressure on the system processing capacity. For example, the peak order volume of Alibaba's Global Shopping Festival in 2020 reached 583,000 orders per second. Even worse, multi-model data are involved in real business. The inability to perform high-throughput, lowlatency transaction processing can result in a poor user experience that can lead to serious financial losses due to customer churn. Although numerous optimizations have been proposed, they can fail in the face of petabytes of data, or be significantly less effective. In this paper, we propose a novel and practical multi-model big data system that can manage petabytes of data. Particularly, we show three special designs for processing the petabytes of data. First, we perform partition to reduce the amount of unnecessary data to be scanned. Second, we adaptively adopt row storage mode for big tables that are frequently updated and column storage mode for tables that are frequently queried to improve the system efficiency. Third, we conduct compression to accelerate IO access speed. We analyze Alibaba's two real PB-level business scenarios, Double 11 and Zhixingtong, and generate workloads and benchmark accordingly to verify our system. Experiments show that our system can efficiently manage petabyte-scale data in real scenarios, providing high-performance querying of terabyte-scale datasets, and be suitable for various workloads.
Feng Zhang 0007, Yinhao Hong, Yunpeng Chai, Wei Lu 0015, Hong Chen 0001, Xiaoyong Du 0001, Le Mi, Xilin Tang, Yanliang Zhou, Peng Zhang 0077, Fengyi Chen
SBAC-PAD3
2022 An Adaptive Elastic Multi-model Big Data Analysis and Information Extraction System
abstract
Abstract With the diverse applications to industry and domain-specific context, multi-source information extraction on semi-structured and unstructured data, as well as across data models, is becoming more common. However, multi-model information extraction often requires the deployment of multiple data model management, storage, and analysis subsystems on the cloud, many subsystems are not high-resource utilization at the same time, and the resource waste phenomenon is often serious. Therefore, an adaptive scalable multi-model big data analysis and information extraction system is designed and implemented in this paper, which can support data maintenance and cross-model query of relational, graph, document, key and other data models, and can provide efficient cross-model information extraction. On this basis, we can achieve the system resource allocation on demand and fast scaling mechanism, according to the real-time requirements of multi-model big data analysis, and dynamic adjustment of each subsystem resource allocation. Therefore, our solution not only guarantees multi-model query and information extraction performance and quality of service, but also significantly reduces the total consumption of system resources and cost.
Qiang Yin 0004, Sheng Du, Jianquan Leng, Yinhao Hong, Feng Zhang 0007, Yunpeng Chai, Xiao Zhang 0014, Xiaonan Zhao, Wei Lu 0015
Data Sci. Eng.6