Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Siyu He

dblp:163/8359 · DBLP profile ↗
← Back
11ranked-venue papers
2as first author
6since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Computer networks · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
Storage systems · 86% High-performance computing · 14%
Artificial intelligence
1 paper
Deep learning architectures and training · 100%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Computational science and engineering · 100%

Topics — the 5 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Storage systems
computational storage
0.712023
λ-IO: A Unified IO Stack for Computational Storage · FAST 2023
Storage systems › i/o architecture › i/o subsystem
i/o stack
0.712023
λ-IO: A Unified IO Stack for Computational Storage · FAST 2023
Storage systems › computational storage
storage offload
0.712023
λ-IO: A Unified IO Stack for Computational Storage · FAST 2023
High-performance computing
large-scale training
0.312018
CosmoFlow: using deep learning to learn the universe at scale · SC 2018
Computational science and engineering
cosmology
0.112018
CosmoFlow: using deep learning to learn the universe at scale · SC 2018

Methods — techniques the papers use, named apart from their topics

deep learning · 1.0
YearPublicationVenuePosition
2026 FxFEDS-ASLM-ELCC : Filtered-x fast Euclidean direction search algorithm based on ASLM with enhanced low-cost center clustering
Xiuwen Yan, Lu Lu 0005, Siyu He, Tao Yu 0004, Wenxing Yang
Signal Process.3
2025 Grouping Strategy Based Secure Cross-User Deduplication in Cloud Storage
abstract
Random chunks generation attack is a more sophisticated form of side channel attack in the process of cross-user client side deduplication. By constructing a deduplication request consists of one or more target chunks together with a number of randomly generated non-duplicate ones, the situation that every one of the target chunks is duplicate cannot be differentiated from that none or part of them is duplicate, which hinders the introduction of obfuscation. Once such an attack is launched in a hybrid mode, the risk is even more significant since a proportion of target chunks is replaced with public ones, which makes the existence privacy of remaining target chunks easier to be exposed. In order to deal with these challenges, we propose a novel grouping strategy based secure cross-user deduplication scheme for cloud storage, which takes the lead to achieve security against the hybrid attack mentioned above. Specifically, we propose a secure grouping obfuscation strategy and an aggregated response generation mechanism, allowing the cloud service provider to randomly distribute query tags of chunks in a certain request into several groups, and generate an obfuscated response for each one of them according to an adaptive response table before returning an aggregated result. The security analysis and experimental results show that the proposed scheme is able to achieve resistance against the aforementioned attacks with controllable overhead introduced no matter how many target chunks are involved.
Siyu He, Haixin Chen, Zhangwei Cao
ICC2
2025 Splicing Strategy based Chunk Level Deduplication Scheme for Cloud Storage
abstract
Random chunks generation attack is a more sophisticated form of side-channel attack, where a deduplication request containing one or more target chunks and several randomly generated non-duplicate chunks is able to be constructed. Once the number of required chunks or linear combinations in response is equal to that of the non-duplicate chunks, the existence privacy of target chunks is immediately exposed. Once such an attack is launched in a hybrid mode, which means some of the target chunks are replaced with public ones with known existence, the probability of privacy exposure further increases. In order to address these challenges, we propose a novel splicing strategy based chunk level deduplication scheme, which effectively protects the existence privacy of target chunks. Specifically, with the help of short hash functions, we introduce a secure splicing strategy for tags that divides query tags into multiple ternary tag groups in an unpredictable manner from the perspective of the attacker, which is the basis for chunk splicing. Furthermore, a response table is constructed accordingly to require uploading of the corresponding number of linear combinations based on the number of non-duplicate tag groups, which is able to introduce obfuscation in a lightweight way. Security analysis and experimental results demonstrate that the proposed scheme can effectively resist the aforementioned attacks with controllable overhead, regardless of the number of target chunks in request.
Haixin Chen, Siyu He, Yamin Yang, Luchao Jin
TrustCom3
2025 A Slim Prompt-Averaged Consistency prompt learning for vision-language model
Siyu He, Sheng-Sheng Wang 0001, Sifan Long 0001
Knowl. Based Syst.1
2024 PTMGS: A Cost-Optimal LLM Training Tasks Migration Method
abstract
Computing power infrastructure has a significant impact on the training speed and cost during the training process of Large Language Models (LLMs). Due to differences in GPU hardware, network architecture, and training frameworks, the cost and unit price of computing power fluctuate greatly, and even within the same cloud service provider, there are huge differences in prices for different clusters. The feature of regularly saving checkpoints during the training process of large language models makes it possible to dynamically migrate tasks between GPU clusters during the training process. Based on this background, this article proposes a set of greedy strategy heuristic algorithms with the goal of minimizing training costs, which choose the best time and target for task migration, thereby achieving overall optimization of training costs. Through multiple rounds of experimental verification on real task samples and simulation data, the results show that the proposed greedy strategy based preemptive algorithm can reduce the overall training cost by more than 10% . This study provides new solutions and methods for efficient computing power supply for large language model training, which can help reduce training costs for computing infrastructure operators and customers.
Gangyi Luo, Siyu He, Genning Zhang, Zhuzhong Qian
HPCC4
2023 λ-IO: A Unified IO Stack for Computational Storage
Zhe Yang 0012, Youyou Lu, Xiaojian Liao, Youmin Chen, Siyu He, Jiwu Shu
FAST6
2018 A computational method for detecting the associations between multiple loci and phenotypes
Zhongmeng Zhao, Jiali Huang, Mingzhe Xu, Ruoyu Liu, Siyu He, Xuanping Zhang
BIBM5
2018 Correcting genomic deletion calls with complex boundaries from next generation sequencing data
Zhongmeng Zhao, Zewen Tian, Yu Geng 0001, Siyu He, Xuanping Zhang
BIBM4
2018 CosmoFlow: using deep learning to learn the universe at scale
Amrita Mathuriya, Deborah Bard, Peter Mendygral, Lawrence Meadows, James Arnemann, Siyu He, Tuomas Kärnä, Diana Moise, Simon J. Pennycook, Kristyn J. Maschhoff, Jason Sewall, Nalini Kumar, Shirley Ho, Michael F. Ringenburg, Prabhat, Victor W. Lee
SC7
2017 A crowdsourcing method for correcting sequencing errors for the third-generation sequencing data
abstract
The third generation sequencing data exposes great advantage on read length, which extremely benefits the genomic analyses. However, the third generation sequencing data implies error models different from the ones that the second generation data brings. It is suggested to correct sequencing errors, which could significantly reduce false positives in downstream analyses. Existing error correction approaches often suffer accuracy loss when the hybrid reads present diversity or the coverage varies. In this paper, we propose a novel method based on crowdsourcing strategy, which is implemented as CLTC. CLTC is also a hybrid correction algorithm, which consists of four steps. The second generation reads are first collected and mapped to the third generation reads. Then, the base difficult level is defined to describe the diversities on a base among a group of 2nd-generation reads covered it. The capability is evaluated for each 2nd-generation read, which considers the base difficult levels across the read, the consistency among overlapped reads and the mapping quality between the 2nd- and 3rd-generation reads. A heuristic algorithm is designed for the calculation of capabilities. An expectation-maximization algorithm is finally used to compute the corrected result for each base-pair. We test CLTC on different datasets and compare to the existing approaches. The results demonstrate that CLTC is able to achieve higher accuracy and performs faster than the existing ones.
Yu Geng 0001, Zhongmeng Zhao, Zhaofang Du, Siyu He, Xuanping Zhang
BIBM6
2014 Comparison of three implementations of digital average current control for DC-DC converters
abstract
Proposed in this paper are three different implementations for digital average current mode control for DC-DC converters operating in the continuous conduction mode. The advantages and disadvantages of each implementation are described. Design procedures for the both the voltage and current loops are described. Using a boost converter prototype, the dynamic performance of all three implementations has been evaluated and is presented here.
Siyu He, R. Mark Nelms
IECON1