VLDB 2026 Research / reviewers in the wild / expert
Siyu He
dblp:163/8359
· DBLP profile ↗
11ranked-venue papers
2as first author
6since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Computer networks · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
2 papers |
Storage systems · 86% High-performance computing · 14% | |
| Artificial intelligence
1 paper |
Deep learning architectures and training · 100% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Computational science and engineering · 100% |
Topics — the 5 heaviest of 6, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Storage systems
computational storage |
0.7 | 1 | 2023 | λ-IO: A Unified IO Stack for Computational Storage · FAST 2023 |
Storage systems › i/o architecture › i/o subsystem
i/o stack |
0.7 | 1 | 2023 | λ-IO: A Unified IO Stack for Computational Storage · FAST 2023 |
Storage systems › computational storage
storage offload |
0.7 | 1 | 2023 | λ-IO: A Unified IO Stack for Computational Storage · FAST 2023 |
High-performance computing
large-scale training |
0.3 | 1 | 2018 | CosmoFlow: using deep learning to learn the universe at scale · SC 2018 |
Computational science and engineering
cosmology |
0.1 | 1 | 2018 | CosmoFlow: using deep learning to learn the universe at scale · SC 2018 |
Methods — techniques the papers use, named apart from their topics
deep learning · 1.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FxFEDS-ASLM-ELCC : Filtered-x fast Euclidean direction search algorithm based on ASLM with enhanced low-cost center clustering
Xiuwen Yan, Lu Lu 0005, Siyu He, Tao Yu 0004, Wenxing Yang |
Signal Process. | 3 |
| 2025 | Grouping Strategy Based Secure Cross-User Deduplication in Cloud StorageabstractRandom chunks generation attack is a more sophisticated form of side channel attack in the process of cross-user client side deduplication. By constructing a deduplication request consists of one or more target chunks together with a number of randomly generated non-duplicate ones, the situation that every one of the target chunks is duplicate cannot be differentiated from that none or part of them is duplicate, which hinders the introduction of obfuscation. Once such an attack is launched in a hybrid mode, the risk is even more significant since a proportion of target chunks is replaced with public ones, which makes the existence privacy of remaining target chunks easier to be exposed. In order to deal with these challenges, we propose a novel grouping strategy based secure cross-user deduplication scheme for cloud storage, which takes the lead to achieve security against the hybrid attack mentioned above. Specifically, we propose a secure grouping obfuscation strategy and an aggregated response generation mechanism, allowing the cloud service provider to randomly distribute query tags of chunks in a certain request into several groups, and generate an obfuscated response for each one of them according to an adaptive response table before returning an aggregated result. The security analysis and experimental results show that the proposed scheme is able to achieve resistance against the aforementioned attacks with controllable overhead introduced no matter how many target chunks are involved. Siyu He, Haixin Chen, Zhangwei Cao |
ICC | 2 |
| 2025 | Splicing Strategy based Chunk Level Deduplication Scheme for Cloud StorageabstractRandom chunks generation attack is a more sophisticated form of side-channel attack, where a deduplication request containing one or more target chunks and several randomly generated non-duplicate chunks is able to be constructed. Once the number of required chunks or linear combinations in response is equal to that of the non-duplicate chunks, the existence privacy of target chunks is immediately exposed. Once such an attack is launched in a hybrid mode, which means some of the target chunks are replaced with public ones with known existence, the probability of privacy exposure further increases. In order to address these challenges, we propose a novel splicing strategy based chunk level deduplication scheme, which effectively protects the existence privacy of target chunks. Specifically, with the help of short hash functions, we introduce a secure splicing strategy for tags that divides query tags into multiple ternary tag groups in an unpredictable manner from the perspective of the attacker, which is the basis for chunk splicing. Furthermore, a response table is constructed accordingly to require uploading of the corresponding number of linear combinations based on the number of non-duplicate tag groups, which is able to introduce obfuscation in a lightweight way. Security analysis and experimental results demonstrate that the proposed scheme can effectively resist the aforementioned attacks with controllable overhead, regardless of the number of target chunks in request. Haixin Chen, Siyu He, Yamin Yang, Luchao Jin |
TrustCom | 3 |
| 2025 | A Slim Prompt-Averaged Consistency prompt learning for vision-language model
Siyu He, Sheng-Sheng Wang 0001, Sifan Long 0001 |
Knowl. Based Syst. | 1 |
| 2024 | PTMGS: A Cost-Optimal LLM Training Tasks Migration MethodabstractComputing power infrastructure has a significant impact on the training speed and cost during the training process of Large Language Models (LLMs). Due to differences in GPU hardware, network architecture, and training frameworks, the cost and unit price of computing power fluctuate greatly, and even within the same cloud service provider, there are huge differences in prices for different clusters. The feature of regularly saving checkpoints during the training process of large language models makes it possible to dynamically migrate tasks between GPU clusters during the training process. Based on this background, this article proposes a set of greedy strategy heuristic algorithms with the goal of minimizing training costs, which choose the best time and target for task migration, thereby achieving overall optimization of training costs. Through multiple rounds of experimental verification on real task samples and simulation data, the results show that the proposed greedy strategy based preemptive algorithm can reduce the overall training cost by more than 10% . This study provides new solutions and methods for efficient computing power supply for large language model training, which can help reduce training costs for computing infrastructure operators and customers. Gangyi Luo, Siyu He, Genning Zhang, Zhuzhong Qian |
HPCC | 4 |
| 2023 | λ-IO: A Unified IO Stack for Computational Storage
Zhe Yang 0012, Youyou Lu, Xiaojian Liao, Youmin Chen, Siyu He, Jiwu Shu |
FAST | 6 |
| 2018 | A computational method for detecting the associations between multiple loci and phenotypes
Zhongmeng Zhao, Jiali Huang, Mingzhe Xu, Ruoyu Liu, Siyu He, Xuanping Zhang |
BIBM | 5 |
| 2018 | Correcting genomic deletion calls with complex boundaries from next generation sequencing data
Zhongmeng Zhao, Zewen Tian, Yu Geng 0001, Siyu He, Xuanping Zhang |
BIBM | 4 |
| 2018 | CosmoFlow: using deep learning to learn the universe at scale
Amrita Mathuriya, Deborah Bard, Peter Mendygral, Lawrence Meadows, James Arnemann, Siyu He, Tuomas Kärnä, Diana Moise, Simon J. Pennycook, Kristyn J. Maschhoff, Jason Sewall, Nalini Kumar, Shirley Ho, Michael F. Ringenburg, Prabhat, Victor W. Lee |
SC | 7 |
| 2017 | A crowdsourcing method for correcting sequencing errors for the third-generation sequencing dataabstractThe third generation sequencing data exposes great advantage on read length, which extremely benefits the genomic analyses. However, the third generation sequencing data implies error models different from the ones that the second generation data brings. It is suggested to correct sequencing errors, which could significantly reduce false positives in downstream analyses. Existing error correction approaches often suffer accuracy loss when the hybrid reads present diversity or the coverage varies. In this paper, we propose a novel method based on crowdsourcing strategy, which is implemented as CLTC. CLTC is also a hybrid correction algorithm, which consists of four steps. The second generation reads are first collected and mapped to the third generation reads. Then, the base difficult level is defined to describe the diversities on a base among a group of 2nd-generation reads covered it. The capability is evaluated for each 2nd-generation read, which considers the base difficult levels across the read, the consistency among overlapped reads and the mapping quality between the 2nd- and 3rd-generation reads. A heuristic algorithm is designed for the calculation of capabilities. An expectation-maximization algorithm is finally used to compute the corrected result for each base-pair. We test CLTC on different datasets and compare to the existing approaches. The results demonstrate that CLTC is able to achieve higher accuracy and performs faster than the existing ones. Yu Geng 0001, Zhongmeng Zhao, Zhaofang Du, Siyu He, Xuanping Zhang |
BIBM | 6 |
| 2014 | Comparison of three implementations of digital average current control for DC-DC convertersabstractProposed in this paper are three different implementations for digital average current mode control for DC-DC converters operating in the continuous conduction mode. The advantages and disadvantages of each implementation are described. Design procedures for the both the voltage and current loops are described. Using a boost converter prototype, the dynamic performance of all three implementations has been evaluated and is presented here. Siyu He, R. Mark Nelms |
IECON | 1 |