Sijie Shen

dblp:202/7006 · DBLP profile ↗
← Back
14ranked-venue papers
5as first author
11since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 2 first-author · 5 since 2021Software engineering, systems software and programming languages · 3 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 GoCache: Accelerating Out-Of-Core Graph Queries with Pattern-Driven Caching
Lixiao Cui, Luofan Chen, Chongzhuo Yang, Xiaojian Luo, Sijie Shen, Wenyuan Yu, Jingren Zhou 0001, Cheng Li 0001
ICDE7
2026 A Blockchain-Based Fine-Grained Reputation-Enhanced Consensus Mechanism for Secure Health Data Trading
abstract
With the deep integration of wearable devices and information technology, health data collection has become more efficient and accurate. This has brought about significant changes to health management and public health. However, due to the high sensitivity and privacy of health data, data sharing faces major challenges. Most existing studies focus on encryption algorithms and access control to ensure data security. They often ignore the credibility of data providers, which affects data quality and reduces user participation. To address these issues, this article proposes a blockchain-based and reputation-enhanced health data trading model. A fine-grained reputation value calculation method based on the Beta distribution is introduced to objectively evaluate the behavior of data providers. Based on this, a consensus mechanism linked to reputation value is designed to improve consensus efficiency and avoid centralization of node selection. Furthermore, this article uses evolutionary game theory to analyze the reward and punishment mechanism in the trading process. It explores the dynamic balance between platform cost and user willingness to share. Experimental results show that the model ensures secure health data sharing, while effectively improving data usability, system fairness, and user participation.
Sijie Shen, Taochun Wang, Fulong Chen 0002, Dong Xie 0005, Chuanxin Zhao
IEEE Trans. Comput. Soc. Syst.1
2025 Moko: Marrying Python with Big Data Systems
abstract
Python stands as the preferred language for data science, thanks to its user-friendly syntax and a robust ecosystem that effortlessly accommodates a variety of data types and workloads, such as relational/tabular data, tensors, and graphs. While Python thrives in smaller data settings, it struggles to scale in distributed big data environments. MOKO is an IR-based execution framework designed to extend Python's reach into the distributed big data domain by generating code that can utilize existing systems such as Spark, Dask, Torch, and GRAPE. Moko preserves Python's key features---interoperability, ease of use, and support for multi-model data types and workloads---while enabling efficient execution in a distributed setting. Our evaluation indicates that MOKO can accelerate Python applications by up to 11× across diverse systems, diminish data alignment overhead by 28×, and outperform hand-optimized solutions by 2.5×.
Tao He 0013, Sijie Shen, Lei Wang 0004, Wenyuan Yu, Jingren Zhou 0001
EuroSys3
2025 KVCache Cache in the Wild: Characterizing and Optimizing KVCache Cache at a Large Cloud Provider
Jinbo Han, Xingda Wei, Sijie Shen, Dingyan Zhang, Chenguang Fang, Rong Chen 0001, Wenyuan Yu, Haibo Chen 0001
USENIX ATC4
2024 Traceable Health Data Sharing Based-on Blockchain
Taochun Wang, Sijie Shen, Qingshan Wu, Fulong Chen 0002, Chuanxin Zhao, Shuhan Wan
WASA (1)2
2024 LSMGraph: A High-Performance Dynamic Graph Storage System with Multi-Level CSR
abstract
The growing volume of graph data may exhaust the main memory. It is crucial to design a disk-based graph storage system to ingest updates and analyze graphs efficiently. However, existing dynamic graph storage systems suffer from read or write amplification and face the challenge of optimizing both read and write performance simultaneously. To address this challenge, we propose LSMGraph, a novel dynamic graph storage system that combines the write-friendly LSM-tree and the read-friendly CSR. It leverages the multi-level structure of LSM-trees to optimize write performance while utilizing the compact CSR structures embedded in the LSM-trees to boost read performance. LSMGraph uses a new memory structure, MemGraph, to efficiently cache graph updates and uses a multi-level index to speed up reads within the multi-level structure. Furthermore, LSMGraph incorporates a vertex-grained version control mechanism to mitigate the impact of LSM-tree compaction on read performance and ensure the correctness of concurrent read and write operations. Our evaluation shows that LSMGraph significantly outperforms state-of-the-art (graph) storage systems on both graph update and graph analytical workloads.
Song Yu 0004, Shufeng Gong 0001, Sijie Shen, Yanfeng Zhang 0001, Wenyuan Yu, Pengxi Liu, Hongfu Li, Xiaojian Luo, Ge Yu 0001, Jingren Zhou 0001
Proc. ACM Manag. Data4
2023 GLogS: Interactive Graph Pattern Matching Query At Large Scale
Longbin Lai, Zhibin Wang 0002, Sijie Shen, Bingqing Lyu, Wenyuan Yu, Zhengping Qian, Chen Tian 0001, Sheng Zhong 0002, Yeh-Ching Chung, Jingren Zhou 0001
USENIX ATC6
2023 Bridging the Gap between Relational OLTP and Graph-based OLAP
Sijie Shen, Zihang Yao, Lei Wang 0004, Longbin Lai, Li Su 0005, Rong Chen 0001, Wenyuan Yu, Haibo Chen 0001, Binyu Zang, Jingren Zhou 0001
USENIX ATC1
2022 Incorporating domain knowledge through task augmentation for front-end JavaScript code generation
abstract
Code generation aims to generate a code snippet automatically from natural language descriptions. Generally, the mainstream code generation methods rely on a large amount of paired training data, including both the natural language description and the code. However, in some domain-specific scenarios, building such a large paired corpus for code generation is difficult because there is no directly available pairing data, and a lot of effort is required to manually write the code descriptions to construct a high-quality training dataset. Due to the limited training data, the generation model cannot be well trained and is likely to be overfitting, making the model's performance unsatisfactory for real-world use. To this end, in this paper, we propose a task augmentation method that incorporates domain knowledge into code generation models through auxiliary tasks and a Subtoken-TranX model by extending the original TranX model to support subtoken-level code generation. To verify our proposed approach, we collect a real-world code generation dataset and conduct experiments on it. Our experimental results demonstrate that the subtoken-level TranX model outperforms the original TranX model and the Transformer model on our dataset, and the exact match accuracy of Subtoken-TranX improves significantly by 12.75% with the help of our task augmentation method. The model performance on several code categories has satisfied the requirements for application in industrial systems. Our proposed approach has been adopted by Alibaba's BizCook platform. To the best of our knowledge, this is the first domain code generation system adopted in industrial development environments.
Sijie Shen, Yihong Dong, Qizhi Guo, Yankun Zhen, Ge Li 0001
ESEC/SIGSOFT FSE1
2022 DrTM+B: Replication-Driven Live Reconfiguration for Fast and General Distributed Transaction Processing
abstract
Recent in-memory database systems leverage advanced hardware features like RDMA to provide transaction processing at millions of transactions per second. Distributed transaction processing systems can scale to even higher rates, especially for partitionable workloads. Unfortunately, it is challenging to sustain such high rates during live reconfiguration of partitions. In this article, we observe that state-of-the-art approaches would cause notable performance disruption under fast transaction processing. To this end, this article presents DrTM+B, a live reconfiguration approach that seamlessly repartitions data with little performance disruption to running transactions. DrTM+B uses a pre-copy-based mechanism to avoid excessive data transfer by leveraging common properties in recent transactional systems. DrTM+B's reconfiguration plans reduce data movement by preferring existing data replicas, while copying data from multiple replicas asynchronously and in parallel. It further reuses the log forwarding mechanism in primary-backup replication to seamlessly track and forward dirty database tuples and avoids iterative copying costs. To commit a reconfiguration plan in a transactional-safe way, DrTM+B designs a cooperative commit protocol for synchronization of data and state among replicas. To boost the performance during data migration, DrTM+B combines the pre-copy and post-copy schemes to propose a hybrid copy scheme. The live reconfiguration approach can also coexist with fault-tolerance mechanisms of primary-backup replication to provide high availability. Evaluation on a working system based on DrTM+R with 3-way replication using typical OLTP workloads like TPC-C and SmallBank shows that DrTM+B incurs only very small performance degradation during live reconfiguration and provides high availability. Both the reconfiguration time and the downtime are also minimal.
Sijie Shen, Xingda Wei, Rong Chen 0001, Haibo Chen 0001, Binyu Zang
IEEE Trans. Parallel Distributed Syst.1
2021 Retrofitting High Availability Mechanism to Tame Hybrid Transaction/Analytical Processing
Sijie Shen, Rong Chen 0001, Haibo Chen 0001, Binyu Zang
OSDI1
2020 LSTM-based argument recommendation for non-API methods
Guangjie Li, Hui Liu 0003, Ge Li 0001, Sijie Shen, Hanlin Tang 0001
Sci. China Inf. Sci.4
2020 Modular Tree Network for Source Code Representation Learning
abstract
Learning representation for source code is a foundation of many program analysis tasks. In recent years, neural networks have already shown success in this area, but most existing models did not make full use of the unique structural information of programs. Although abstract syntax tree (AST)-based neural models can handle the tree structure in the source code, they cannot capture the richness of different types of substructure in programs. In this article, we propose a modular tree network that dynamically composes different neural network units into tree structures based on the input AST. Different from previous tree-structural neural network models, a modular tree network can capture the semantic differences between types of AST substructures. We evaluate our model on two tasks: program classification and code clone detection. Our model achieves the best performance compared with state-of-the-art approaches in both tasks, showing the advantage of leveraging more elaborate structure information of the source code.
Wenhan Wang, Ge Li 0001, Sijie Shen, Xin Xia 0001, Zhi Jin 0001
ACM Trans. Softw. Eng. Methodol.3
2017 Replication-driven Live Reconfiguration for Fast Distributed Transaction Processing
Xingda Wei, Sijie Shen, Rong Chen 0001, Haibo Chen 0001
USENIX ATC2