EDBT 2026 Demo / reviewers in the wild / expert
Shangyu Wu
dblp:241/1185
· DBLP profile ↗
13ranked-venue papers
4as first author
11since 2021 · last 2026
0000-0002-1961-143XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ColdCode: Cold Data Encoding for Enhanced Reliability and Lifetime in 3D NAND FlashabstractCold storage, which stores rarely accessed data, dominates modern data centers but is poorly served by NAND flash's data randomization. As a common practice today by flash vendors, data randomization is applied in NAND flash chips to avoid extreme data patterns that generate the worst-case raw bit error rate (RBER). However, this paper demonstrates that data randomization rules out the opportunity to explore data patterns with very low RBERs, through a comprehensive analysis on data randomization in 3D NAND flash chips (across 8 models). Motivated by this, we propose ColdCode, a novel data coding framework to replace the conventional randomizer in 3D high-density flash for cold data storage. Using a tag that indicates coldness information passed from the file system to solid-state drive (SSD) controllers, the controller encodes cold data to enhance reliability and extend the lifetime. ColdCode employs two coding techniques: skewed coding and reversed Huffman coding, applied based on the data entropy. These techniques effectively reduce the RBER of encoded data compared to conventional randomization. Experimental results on real high-density 3D flash chips show that, under the same error conditions, the skewed coding and reversed Huffman coding reduce the average RBER by 42% and 30%, respectively. Consequently, the lifetime of the flash chips is prolonged by factors of 2.97× and 1.7×, respectively, compared to data randomization. Qiao Li 0001, Shangyu Wu, Yufei Cui, Jie Zhang 0048, Chun Jason Xue |
EuroSys | 2 |
| 2025 | GeneQuery: Generalized Gene Expression Prediction from Histology Images via Image-Gene QAabstractGene expression profiling provides profound insights into molecular mechanisms, but its time-consuming and costly nature often presents significant challenges. Recent advancements have utilized histological images to predict spatially resolved gene expression profiles. The gene prediction problem has two main characteristics, i.e., spatial heterogeneity and gene interdependency. Existing works only focus on addressing spatial heterogeneity, ignoring the importance of relationships between genes. To address the above limitation, this paper presents GeneQuery, which aims to solve this gene expression prediction task in a question-answering (QA) manner for better generality and flexibility. Specifically, GeneQuery takes gene meta-information as queries and whole-slide images as contexts and then predicts the queried gene expression values. GeneQuery learns to dynamically fuse spatial image features with gene semantics, and uses an attention mechanism to explicitly capture the spatial information and a shared regressor to implicitly capture the gene relationship. A variant of GeneQuery can use the attention mechanism to capture gene relationships for more complex tissue cases. This QA-based reformulation also grants the model the ability to generalize, enabling the prediction of unseen gene expression without requiring model retraining. Comprehensive experiments on spatial transcriptomics datasets show that the proposed GeneQuery outperforms existing state-of-the-art methods on known and unseen genes. More results also demonstrate that GeneQuery can help analyze the tissue structure. Linjing Liu, Yufei Cui, Shangyu Wu, Xue (Steve) Liu, Antoni B. Chan, Chun Jason Xue |
BIBM | 4 |
| 2025 | MLWQ: Efficient Small Language Model Deployment via Multi-Level Weight QuantizationabstractSmall language models (SLMs) are gaining attention for their lower computational and memory needs while maintaining strong performance.However, efficiently deploying SLMs on resource-constrained devices remains a significant challenge.Post-training quantization (PTQ) is a widely used compression technique that reduces memory usage and inference computation, yet existing methods face challenges in inefficient bit-width allocation and insufficient fine-grained quantization adjustments, leading to suboptimal performance, particularly at lower bit-widths.To address these challenges, we propose multi-level weight quantization (MLWQ), which facilitates the efficient deployment of SLMs.Our method enables more effective bit-width allocation by jointly considering inter-layer loss and intralayer salience.Furthermore, we propose a finegrained partitioning of intra-layer salience to support the tweaking of quantization parameters within each group.Experimental results indicate that MLWQ achieves competitive performance compared to state-of-the-art methods, providing an effective approach for the efficient deployment of SLMs while maintaining model accuracy. Chun Hu, Shangyu Wu, Chun Jason Xue, Qing'an Li |
EMNLP | 3 |
| 2025 | EvoP: Robust LLM Inference via Evolutionary Pruning
Shangyu Wu, Hongchao Du, Tei-Wei Kuo, Nan Guan, Chun Jason Xue |
NLPCC (1) | 1 |
| 2024 | CHESS: Optimizing LLM Inference via Channel-Wise Thresholding and Selective SparsificationabstractDeploying large language models (LLMs) on edge devices presents significant challenges due to the substantial computational overhead and memory requirements.Activation sparsification can mitigate these resource challenges by reducing the number of activated neurons during inference.Existing methods typically employ thresholding-based sparsification based on the statistics of activation tensors.However, they do not model the impact of activation sparsification on performance, resulting in suboptimal performance degradation.To address the limitations, this paper reformulates the activation sparsification problem to explicitly capture the relationship between activation sparsity and model performance.Then, this paper proposes CHESS , a general activation sparsification approach via CHannel-wise thrEsholding and Selective Sparsification.First, channel-wise thresholding assigns a unique threshold to each activation channel in the feed-forward network (FFN) layers.Then, selective sparsification involves applying thresholding-based activation sparsification to specific layers within the attention modules.Finally, we detail the implementation of sparse kernels to accelerate LLM inference.Experimental results demonstrate that the proposed CHESS achieves lower performance degradation over eight downstream tasks while activating fewer parameters than existing methods, thus speeding up the LLM inference by up to 1.27x. Shangyu Wu, Weidong Wen, Chun Jason Xue, Qing'an Li |
EMNLP | 2 |
| 2024 | LeaderKV: Improving Read Performance of KV Stores via Learned Index and Decoupled KV TableabstractLog-structured merge-tree (LSM-tree) is a storage architecture widely used in key-value (KV) stores. To enhance the read efficiency of LSM-tree, recent works utilize the learned index to learn the mapping between keys and locations. However, in existing learned-index-aided KV stores, inefficient design of the learned index and disk access significantly impact the read performance. How to design a learned KV store to improve index efficiency and minimize disk access remains a critical problem. This paper presents LeaderKV, a read-optimized LSM-tree-based KV store. LeaderKV employs decoupled KV tables (DK-Table) and efficient learned indexes for data retrieval. DKTables are storage files in Leader Kvbecause they avoid reading irrelevant data in collaboration with learned indexes during queries. A learned index called Leader is proposed to accelerate data retrieval within DKTable. Leader is composed of precise models and approximate models. A redirect mechanism is designed to reduce the cost of mispredictions in Leader. We integrate DKTable and Leader into LeaderKV and demonstrate its effectiveness using a variety of datasets and workloads. Experimental results show that LeaderKV significantly improves the read performance compared to representative schemes. Yi Wang 0003, Jianan Yuan, Shangyu Wu, Jiaxian Chen, Chenlin Ma, Jianbin Qin |
ICDE | 3 |
| 2024 | ReFusion: Improving Natural Language Understanding with Computation-Efficient Retrieval Representation FusionabstractRetrieval-based augmentations (RA) incorporating knowledge from an external database into language models have greatly succeeded in various knowledge-intensive (KI) tasks. However, integrating retrievals in non-knowledge-intensive (NKI) tasks is still challenging.
Existing works focus on concatenating retrievals with inputs to improve model performance. Unfortunately, the use of retrieval concatenation-based augmentations causes an increase in the input length, substantially raising the computational demands of attention mechanisms.
This paper proposes a new paradigm of RA named \textbf{ReFusion}, a computation-efficient \textbf{Re}trieval representation \textbf{Fusion} with bi-level optimization. Unlike previous works, ReFusion directly fuses the retrieval representations into the hidden states of models.
Specifically, ReFusion leverages an adaptive retrieval integrator to seek the optimal combination of the proposed ranking schemes across different model layers. Experimental results demonstrate that the proposed ReFusion can achieve superior and robust performance in various NKI tasks. Shangyu Wu, Yufei Cui, Xue (Steve) Liu, Buzhou Tang, Tei-Wei Kuo, Chun Jason Xue |
ICLR | 1 |
| 2023 | Tidal-Tree-Mem: Toward Read-Intensive Key-Value Stores With Tidal Structure Based on LSM-TreeabstractThe log-structured merge-tree (LSM-tree)-based key-value store has been widely adopted by many large-scale data storage applications for its excellent write performance. However, such write performance gains mainly come from scarifying read performance due to the leveled and log-structured intrinsic characteristics of the LSM-tree. Therefore, the critical challenge of the existing LSM-tree is how to improve the read efficiency by reducing read amplification. This article for the first time proposes Tidal-tree-Mem, a novel data structure where data flow inside the LSM-tree-like Tidal waves. First, a floating strategy is proposed to allow frequently accessed files at the bottom of the LSM-tree to move to higher positions, reducing read amplification. Second, a stretching strategy is proposed to vary the shape of the LSM-tree to adapt to workloads with different characteristics. To evaluate the performance of Tidal-tree-Mem, we conduct a series of experiments using standard benchmarks from YCSB. The experimental results show that Tidal-tree-Mem can effectively reduce read amplification and the overall latency by over 71.94% and 47.34%, respectively, compared with representative schemes. Chenlin Ma, Shangyu Wu, Yi Wang 0003, Rui Mao 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2022 | Work-in-Progress: Lark: A Learned Secondary Index Toward LSM-tree for Resource-Constrained Embedded Storage SystemsabstractLSM-tree-based key-value stores are popular in embedded storage systems. With the growing demands of data analysis, the secondary index is created to support non-primary-key lookups. However, the lookup efficiency and space consumption of secondary index remain for further optimization. Inspired by the learned index, this paper presents Lark, a learned secondary index toward LSM-tree for resource-constrained embedded storage systems. Lark employs machine learning to speed up the non-primary-key queries and compress secondary indexes. Our preliminary evaluations show that, in comparison with traditional secondary index schemes, Lark achieves better lookup performance with less space consumption. Jianan Yuan, Shangyu Wu, Yiquan Lin, Chenlin Ma, Rui Mao 0001, Yi Wang 0003 |
CODES+ISSS | 3 |
| 2022 | NFL: Robust Learned Index via Distribution TransformationabstractRecent works on learned index open a new direction for the indexing field. The key insight of the learned index is to approximate the mapping between keys and positions with piece-wise linear functions. Such methods require partitioning key space for a better approximation. Although lots of heuristics are proposed to improve the approximation quality, the bottleneck is that the segmentation overheads could hinder the overall performance. This paper tackles the approximation problem by applying a distribution transformation to the keys before constructing the learned index. A two-stage Normalizing-Flow-based Learned index framework (NFL) is proposed, which first transforms the original complex key distribution into a near-uniform distribution, then builds a learned index leveraging the transformed keys. For effective distribution transformation, we propose a Numerical Normalizing Flow (Numerical NF). Based on the characteristics of the transformed keys, we propose a robust After-Flow Learned Index (AFLI). To validate the performance, comprehensive evaluations are conducted on both synthetic and real-world workloads, which shows that the proposed NFL produces the highest throughput and the lowest tail latency compared to the state-of-the-art learned indexes. Shangyu Wu, Yufei Cui, Jinghuan Yu, Xuan Sun 0003, Tei-Wei Kuo, Chun Jason Xue |
Proc. VLDB Endow. | 1 |
| 2022 | Bits-Ensemble: Toward Light-Weight Robust Deep Ensemble by Bits-SharingabstractRobustness and uncertainty estimation is crucial to the safety of deep neural networks (DNNs) deployed on the edge. The deep ensemble model, composed of a set of individual DNNs (namely members), has strong performance in accuracy, uncertainty estimation, and robustness to out-of-distribution data and adversarial attacks. However, the storage and memory consumption increases linearly with the number of members within an ensemble. Previous works focus on selecting better members, layer-wise low-rank approximation of ensemble parameters, and designing partial ensemble model for reducing the ensemble size, thus lowering storage and memory consumption. In this work, we pay attention to the quantization of the ensemble, which serves as the last mile of network deployment. We propose a differentiable and parallelizable bit sharing scheme that allows the members to share the less significant bits of parameters, without hurting the performance, leaving alone the more significant bits. The intuition is that, numerically, more significant bits (e.g., the bit for the sign) are more useful in distinguishing a member from other members. For real deployment of the bit-sharing scheme, we further propose an efficient encoding-decoding scheme with minimal storage overhead. The experimental results show that, BitsEnsemble reduces the storage size of ensemble for over$22\times $, with only$0.36\times $increase in training latency, and no sacrifice of inference latency. The code is available inhttps://github.com/ralphc1212/bitsensemble. Yufei Cui, Shangyu Wu, Qiao Li 0001, Antoni B. Chan, Tei-Wei Kuo, Chun Jason Xue |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2020 | Towards Read-Intensive Key-Value Stores with Tidal Structure Based on LSM-TreeabstractKey-value store has played a critical role in many large-scale data storage applications. The log-structured merge-tree (LSM-tree) based key-value store achieves excellent performance on write-intensive workloads which is mainly benefited from the mechanism of converting a batch of random writes into sequential writes. However, LSM-tree doesn't improve a lot in read-intensive workloads which takes a higher latency. The main reason lies in the hierarchical search mechanism in LSM-tree structure. The key challenge is how to propose new strategies based on the existing LSM-tree structure to improve read efficiency and reduce read amplifications.This paper proposes Tidal-tree, a novel data structure where data flows inside LSM-tree like Tidal waves. Tidal-tree targets at improving read efficiency in read-intensive workloads. Tidal-tree allows frequently accessed files at the bottom of LSM-tree to move to higher positions, thereby reducing read latency. Tidal-tree also makes LSM-tree into a variable shape to cater for different characteristic workloads. To evaluate the performance of Tidal-tree, we conduct a series of experiments using standard benchmarks from YCSB. The experimental results show that Tidal-tree can significantly improve read efficiency and reduce read amplifications compared with representative schemes. Yi Wang 0003, Shangyu Wu, Rui Mao 0001 |
ASP-DAC | 2 |
| 2019 | Towards Cross-Platform Inference on Edge Devices with Emerging Neuromorphic ArchitectureabstractDeep convolutional neural networks have become the mainstream solution for many artificial intelligence applications. However, they are still rarely deployed on mobile or edge devices due to the cost of a substantial amount of data movement among limited resources. The emerging processing-inmemory neuromorphic architecture offers a promising direction to accelerate the inference process. The key issue becomes how to effectively allocate the processing of inference between computing and storage resources on an edge device.This paper presents Mobile-I, a resource allocation scheme to accelerate the Inference process on Mobile or edge devices. Mobile-I targets at the emerging 3D neuromorphic architecture to reduce the processing latency among computing resources and fully utilize the limited on-chip storage resources. We formulate the target problem as a resource allocation problem and use a software-based solution to offer the cross-platform deployment across multiple mobile or edge devices. We conduct a set of experiments using realistic workloads that are generated from Intel Movidius neural compute stick. Experimental results show that Mobile-I can effectively reduce the processing latency and improve the utilization of computing resources with negligible overhead in comparison with representative schemes. Shangyu Wu, Yi Wang 0003, Amelie Chi Zhou, Rui Mao 0001, Zili Shao, Tao Li 0006 |
DATE | 1 |