Yanpeng Hu

dblp:13/10805 · DBLP profile ↗
← Back
15ranked-venue papers
4as first author
13since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 3 first-author · 5 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Computer networks · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 Pome: Parallelizing I/Os and Computations for Efficient LSM-tree-based Data Storage
abstract
CPU computations and I/O operations are fundamental to data storage systems. Storage systems conduct computations with their user threads, such as sorting data for orderliness. They handle I/Os through system calls (syscalls) including file write, read, and fsync, which the OS’s kernel threads perform with storage devices.
Yanpeng Hu, Chundong Wang 0001
HPDC1
2025 Bridging a Shortcut to Exempt the Software Tax Charged on Logging I/Os
abstract
Databases widely adopt the technique of logging for durability and consistency, while it introduces severe performance overhead. Efforts to mitigate the logging overhead include optimizing the fsync syscall and using preallocated log files with fdatasync to persist necessary data only. Additionally, hardware advancements like power loss protection (PLP) in modern solid-state drives (SSDs) have reduced the time cost for each I/O operation at the hardware level. Yet, logging I/Os still bottleneck many databases like OceanBase due to the traversal through multiple software layers that jointly impose a significant software tax. In this paper, we find that preal-location establishes a log file's stable structure, minimizing the likelihood of space reallocations and permission changes. Leveraging this, we propose Éxitos, which reorganizes the in-memory offset-to-block mapping per log file for quick lookups and creates a direct I/O path from database to SSD, bypassing most of the software layers. Implemented with eBPF on an NVMe SSD, Éxitos improves the performance of OceanBase by up to 2.3× with write-intensive workloads.
Yanpeng Hu, Yunxin Yang, Chundong Wang 0001
HotStorage1
2025 Extending Applications Safely and Efficiently
Yusheng Zheng, Yiwei Yang 0002, Yanpeng Hu, Xiaozheng Lai, Dan Williams 0001, Andi Quinn
OSDI4
2025 Transmission scheduling for remote state estimation with quantized communication against tampering attacks
Mengqi Li 0003, Yanpeng Hu, Jin Guo 0003
Ad Hoc Networks2
2025 Optimal Consensus Control Strategy for Multi-Agent Systems Under Cyber Attacks via a Stackelberg Game Approach
abstract
Based on the Stackelberg game theory, this paper studies the optimal consensus control problem of the multi-agent system (MAS) under multiple network attacks. Multiple attackers interfere with the information interaction between agents by intervening in the transmission channel of the multi-agent wireless network, thereby changing the communication topology of the MAS. Under the limited energy of the system and attackers, the energy allocation of transmission channels is used as the performance index, and the importance of the transmission channel to the MAS is used as an auxiliary indicator to establish a Stackelberg game model between the MAS and attackers under incomplete information. By solving the Stackelberg equilibrium solution of the game model, the optimal attack energy allocation scheme and the optimal energy allocation scheme of the system are obtained, respectively. Then, based on the equilibrium solution, the Hamilton-Jacobi-Bellman equation with coupling terms for the MAS under attacks is given. In view of the broad learning system theory, online algorithms that approximate the optimal performance function and the optimal controller are designed respectively, and sufficient conditions are obtained for the proposed optimal controller to ensure the stability and consistency of the MAS. Finally, the effectiveness of the method is verified by a numerical simulation.
Yanpeng Hu, Yinghui Wang 0010, Ruizhe Jia, Jin Guo 0003
IEEE Trans Autom. Sci. Eng.2
2024 Adaptive Compensation of Range Space-Time Varying Dispersion for Hypersonic Vehicle Coated by Plasma Sheath
abstract
When the hypersonic vehicle travel through the atmosphere, its surface will be covered by the plasma sheath, which will affect the radar imaging seriously, such as the inverse synthetic aperture radar (ISAR) and synthetic aperture radar (SAR). The dispersion effect of the space-time varying plasma sheath brings different phase errors to the range signals, resulting in the complex defocusing of the radar image. In this article, the error model of the range signal with the space-time varying dispersion effect is established, and the quadratic phase error (QPE) is focused on analysis. Then, an adaptive compensation of range space-time varying dispersion based on autofocus is proposed for the hypersonic vehicles coated by plasma sheath. Simulating with the typical parameters, the affected range signal is well focused after compensation. Meanwhile, the imaging quality is quantitatively evaluated with the promotion of the resolution, the peak side lobe ratio (PSLR) and the integral side lobe ratio (ISLR) from 0.45 m to 0.13 m, −0.32 dB to −13.26 dB and 4.55 dB to −10.03 dB, which proves the effectiveness of the proposed compensation method.
Yanpeng Hu, Wei Guo 0025, Fangfang Shen, Peng Xiao 0001
IGARSS1
2024 Summarizing Charts of Financial Document via Context-Aware Multi-Modeling
abstract
In the field of financial analysis, investment research analysts depend on a detailed understanding of complex financial documents to guide their decision-making process. Charts, while providing visual insights into data, present challenges in summarization. To address this issue, we present a novel approach that leverages contextual awareness, both in terms of textual semantics and visual perception. Our method begins with object detection technology to accurately locate and identify charts. Subsequently, a pre-trained language model is employed for vectorizing text and chart captions, enabling effective correlation between charts and their textual descriptions. Utilizing a large language model and strategic prompt engineering, we generate concise yet informative chart summaries, and incorporate visual saliency to assign scores, quantifying the importance of each chart for more effective data interpretation. Our study, supported by dedicated datasets, validates efficiency and accuracy improvements in financial analysis, expediting well-informed investment decisions.
Xiaoyue Huang, Yaxuan Zheng, Xiping Wang, Yanpeng Hu, Changbo Wang, Chenhui Li 0001
IJCNN4
2024 Tampering Attack Detection for Remote State Estimation With Quantized Observations
abstract
Aiming at the security problem of cyber physical system (CPS), this paper designs and evaluates a tampering attack detection scheme for remote state estimation with quantized observations. When the transmitted information is quantized and the attacker randomly tampers it with a certain probability, we design a detection algorithm with a given error threshold based on the received data. Moreover, we proposes detection performance indexes, leakage detection rate and false detection rate, for the above algorithm, which are used to measure the algorithm performance to discriminate the tampered data, and further analyzes the impact of the attack strategy and the error threshold on performance indexes. Finally, simulations are designed to verify the theoretical results.
Mengqi Li 0003, Yanpeng Hu, Jin Guo 0003
IEEE Signal Process. Lett.2
2023 Exploring Architectural Implications to Boost Performance for in-NVM B+-Tree
abstract
Computer architecture keeps evolving to support the byte-addressable non-volatile memory (NVM). Researchers have tailored the prevalent B+-tree with NVM, crafting a history of utilizing architectural supports to gain both high performance and crash consistency. The latest architecture-level changes for NVM, e.g., the eADR, motivate us to further explore architectural implications in the design and implementation of in-NVM B+-tree. Our quantitative study finds that eADR makes the cache misses impact increasingly on an in-NVM B+-tree's performance. We hence propose Conan for the conflict-aware node allocation based on theoretical justifications. Conan decomposes the virtual addresses of B+-tree nodes regarding a VIPT cache and intentionally places them into different cache sets. Experiments show that Conan evidently reduces cache conflicts and boosts the performance of state-of-the-art in-NVM B+-tree.
Yanpeng Hu, Qisheng Jiang 0001, Chundong Wang 0001
ASP-DAC1
2023 Asynchronous and Adaptive Checkpoint for WAL-based Data Storage Systems
abstract
Write-ahead logging (WAL) is widely utilized to ensure data’s integrity for data storage systems. Modified data is firstly written to a WAL file. Then data is persistently flushed to original home location for in-place update. These two steps are referred to as commit and checkpoint. In this paper, we take SQLite in the WAL mode to study the impact of checkpoint. Once 1,000 pages accumulate in the WAL file, SQLite checkpoints them to the database file with fsync. Such a periodical checkpoint fashion causes substantial spikes to the user-facing latency of inserting or updating data over time. Also, the fixed checkpoint frequency of every 1,000 pages does not consider the runtime write/read access pattern. We propose an algorithm named Walack. Walack conducts fsync asynchronously for each checkpoint. By observing write and read requests, it online adjusts the checkpoint frequency. These two strategies jointly enable Walack to gain both high performance and space efficiency. Experiments show that Walack reduces the user-facing tail latency by up to 92.3% for write requests, with both average write and read performances retained.
Yanpeng Hu, Chundong Wang 0001
ICPADS2
2023 Estimating Market Value of Companies Based on Finance Statement through Data Fusion
abstract
The evaluation of a company's value can serve as a guide for investors to assess the company and make informed investment decisions. However, conventional valuation techniques are not applicable to Initial Public Offering (IPO) companies in China, mainly due to the absence of historical market performance. In contrast, a company's finance statement provides a periodic overview of the company's operational and production activities, which is linked to its market performance. Traditional methods often rely on the selection of a limited number of financial indicators from the finance statement and the application of regression analysis. These approaches fail to fully exploit the comprehensive data available in the finance statement. This study proposes a comprehensive method that leverages all relevant information contained in the finance statement, including industry interconnections, financial indices, and additional insights obtained from the report. The structured data is analyzed through tree models, while the interrelationships between different companies are modeled through graph neural networks. Our approach offers a multi-perspective evaluation of IPO companies. The results of our experiments demonstrate that our method can effectively utilize the valuable information in finance statements and improve outcomes.
Shiqi Jiang 0001, Yaxuan Zheng, Wenli Xiong, Yanpeng Hu, Changbo Wang, Chenhui Li 0001
IJCNN6
2022 NobLSM: an LSM-tree with non-blocking writes for SSDs
abstract
Solid-state drives (SSDs) are gaining popularity. Meanwhile, key-value stores built on log-structured merge-tree (LSM-tree) are widely deployed for data management. LSM-tree frequently calls syncs to persist newly-generated files for crash consistency. The blocking syncs are costly for performance. We revisit the necessity of syncs for LSM-tree. We find that Ext4 journaling embraces asynchronous commits to implicitly persist files. Hence, we design NobLSM that makes LSM-tree and Ext4 cooperate to substitute most syncs with non-blocking asynchronous commits, without losing consistency. Experiments show that NobLSM significantly outperforms state-of-the-art LSM-trees with higher throughput on an ordinary SSD.
Haoran Dang, Chongnan Ye, Yanpeng Hu, Chundong Wang 0001
DAC3
2021 Industry Chain Graph Building Based on Text Semantic Association Mining
abstract
The current volume of data in the field of securities investment is increasing dramatically. Simultaneously, the linkage of data from multiple parties makes investment reasoning decisions more challenging than ever. In response to this problem, the financial field's knowledge graph can improve the efficiency, depth, and breadth of financial practitioners' information analysis. Some existing financial knowledge graphs analyze the shareholding relationship between companies. Still, because they are limited to observing data from the company's perspective, users without professional industry background cannot quickly find the industry factors of stock market changes. This paper proposes a financial knowledge graph from the industry chain's perspective. This paper builds upstream and downstream relationships between industries through Transformer-based bidirectional encoder to mine potential industry chain associations from text data and completes the long industry chain of the stock market. This paper also builds a visualization system to display and explore the connection between listed companies and industries. Users can inspect the industry chain's composition and each company's revenue status and stock market conditions in the industry chain. The experiment shows that when the market price fluctuation is detected, the stock price fluctuation can be traced back to its origin in the knowledge graph.
Jipeng Li, Yujing Sun 0003, Chenhui Li 0001, Yanpeng Hu, Changbo Wang
IJCNN4
2019 Visual Analysis of Retailing Store Location Selection
abstract
An appropriate location of the retailing store is vital for achieving business success. However, a huge amount of complex information needs to be considered in location selection, such as customer flow, the business environment, and current business performance. Unlike traditional location recommendation method on account of statistical sampling, we establish a model of business-district attractiveness based on customer flow and provide a method of data-driven visual comparisons. In addition, we build an interactive visual analysis system with a user-friendly interface for an interactive visual query about complex business and environment information. Our system can help users select retailing store location, support interactive visual queries, and display rich information to facilitate manager in decision-making.
Kelin Li, Yi-Na Li, Yanpeng Hu, Changbo Wang
VINCI4
2011 Dacoop: Accelerating Data-Iterative Applications on Map/Reduce Cluster
abstract
Map/reduce is a popular parallel processing framework for massive-scale data-intensive computing. The data-iterative application is composed of a serials of map/reduce jobs and need to repeatedly process some data files among these jobs. The existing implementation of map/reduce framework focus on perform data processing in a single pass with one map/reduce job and do not directly support the data-iterative applications, particularly in term of the explicit specification of the repeatedly processed data among jobs. In this paper, we propose an extended version of Hadoop map/reduce framework called Dacoop. Dacoop extends Map/Reduce programming interface to specify the repeatedly processed data, introduces the shared memory-based data cache mechanism to cache the data since its first access, and adopts the caching-aware task scheduling so that the cached data can be shared among the map/reduce jobs of data-iterative applications. We evaluate Dacoop on two typical data-iterative applications: k-means clustering and the domain rule reasoning in sementic web, with real and synthetic datasets. Experimental results show that the data-iterative applications can gain better performance on Dacoop than that on Hadoop. The turnaround time of a data-iterative application can be reduced by the maximum of 15.1%.
Guangrui Li 0004, Lei Wang 0004, Yanpeng Hu
PDCAT4