VLDB 2026 Research / reviewers in the wild / expert
Yong Sheng
dblp:35/3373
· DBLP profile ↗
9ranked-venue papers
4as first author
2since 2021 · last 2024
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 3Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 first-authorComputer networks · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Storage systems · 100% | |
| Computer networks
1 paper |
Wireless sensing and localization · 77% Wireless networking · 23% |
Topics — the 7 heaviest of 7, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Storage systems › key-value storage
compaction |
0.8 | 1 | 2024 | LavaStore: ByteDance's Purpose-built, High-performance, Cost-effective Local Storage Engine for Cloud Services · Proc. VLDB Endow. 2024 |
Storage systems
file systems |
0.8 | 1 | 2024 | LavaStore: ByteDance's Purpose-built, High-performance, Cost-effective Local Storage Engine for Cloud Services · Proc. VLDB Endow. 2024 |
Storage systems
key-value storage |
0.8 | 1 | 2024 | LavaStore: ByteDance's Purpose-built, High-performance, Cost-effective Local Storage Engine for Cloud Services · Proc. VLDB Endow. 2024 |
Storage systems › key-value storage
LSM-tree |
0.8 | 1 | 2024 | LavaStore: ByteDance's Purpose-built, High-performance, Cost-effective Local Storage Engine for Cloud Services · Proc. VLDB Endow. 2024 |
Wireless sensing and localization
received signal strength |
0.1 | 1 | 2008 | Detecting 802.11 MAC Layer Spoofing Using Received Signal Strength · INFOCOM 2008 |
Network security
wireless network security |
0.1 | 1 | 2008 | Detecting 802.11 MAC Layer Spoofing Using Received Signal Strength · INFOCOM 2008 |
Wireless networking › WLAN
IEEE 802.11 |
0.0 | 1 | 2008 | Detecting 802.11 MAC Layer Spoofing Using Received Signal Strength · INFOCOM 2008 |
Methods — techniques the papers use, named apart from their topics
write-ahead logging · 0.8KV separation · 0.8gaussian mixture model · 0.2RSS profiling · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | A multimodal data sensing and feature learning-based self-adaptive hybrid approach for machining quality prediction
Yong Sheng, Yingfeng Zhang, Ming Luo 0005, Yifan Pang, Qinan Wang |
Adv. Eng. Informatics | 1 |
| 2024 | LavaStore: ByteDance's Purpose-built, High-performance, Cost-effective Local Storage Engine for Cloud ServicesabstractPersistent key-value (KV) stores are widely used by cloud services at ByteDance as local storage engines, and RocksDB used to be the de facto implementation since it can be tailored to a variety of workloads and requirements. In this paper, we provide key insights into local storage engine usage at ByteDance, explain why the combination of highly write-intensive workloads and stringent requirements on cost efficiency and point lookup tail latency may pose challenges to a general-purpose local storage engine such as RocksDB, and present the design and implementation of LavaStore , a high-performance cost-effective local storage engine purpose-built to address these challenges. LavaStore achieves its design goals by selectively customizing a few components of a RocksDB-based, general-purpose local storage engine, including a distinct KV separation design that decouples garbage collection from compaction, a specialized engine type for the commonly recurring Write-Ahead-Logging workload, and a customized user-space append-only filesystem. LavaStore has been deployed to production with hundreds of thousands of running instances, storing more than 100 PB of data and serving billions of requests per second, bringing significant performance improvements and cost reductions to customers over their original local storage engines. For example, a ByteDance proprietary distributed OLTP database service has experienced a reduction in average write and read latency by 61% and 16%, respectively, and a ByteDance proprietary caching service has gained an 87% increase in write throughput with no more than 6% space overhead. Jiaxin Ou, Sheng Qiu, Yizheng Jiao, Qizhong Mao, Zhengyu Yang 0012, Yang Liu 0442, Jianyang Hu, Jinrui Liu, Yong Sheng, Cao Lixun, Hongde Li, Lei Zhang 0213, Jianjun Chen 0001 |
Proc. VLDB Endow. | 14 |
| 2014 | Formal Verification of Fault-Tolerant and Recovery Mechanisms for Safe Node Sequence ProtocolabstractFault-tolerance has huge impact on embedded safety-critical systems. As technology that assists to the development of such improvement, Safe Node Sequence Protocol (SNSP) is designed to make part of such impact. In this paper, we present a mechanism for fault-tolerance and recovery based on the Safe Node Sequence Protocol (SNSP) to strengthen the system robustness, from which the correctness of a fault-tolerant prototype system is analyzed and verified. In order to verify the correctness of more than thirty failure modes, we have partitioned the complete protocol state machine into several subsystems, followed to the injection of corresponding fault classes into dedicated independent models. Experiments demonstrate that this method effectively reduces the size of overall state space, and verification results indicate that the protocol is able to recover from the fault model in a fault-tolerant system and continue to operate as errors occur. Rui Zhou 0005, Rong Min, Chanjuan Li, Yong Sheng, Qingguo Zhou, Kuanching Li |
AINA | 5 |
| 2013 | A Server Model for Reliable Communication on Cell/B.EabstractIn most cases of safety-related systems, the network is an indispensable part. At this point, the system reliability is as important as the system communication quality. With the emergence of multi-core architectures, the first generation usually aims to provide reliable and deterministic computing resources. Therefore, with the boost requirement of reliability and throughput that cannot be satisfied by general single-core processors, the deployment of safety-related systems is transferred and processed multi-core environments. In this paper, we propose Reliable Communication Server on SPU (RCSoS), which is a server model for reliable communication utilizing SPU (Synergistic Processor Unit) in Cell/B.E (Cell Broadband Engine). It simulates SPU as a communication server and guarantees the reliability and determinacy by the isolation mode of SPU and contract model. We have implemented RCSoS in PlayStation 3, which dynamically adjust parameters, and inform applications on contract violations. Experiments show the performance of this model. Rui Zhou 0005, Huaming Chen, Yong Sheng, Qingguo Zhou, Kuanching Li |
ICPP | 4 |
| 2013 | XtratuM/PPC: a hypervisor for partitioned system on PowerPC processors
Rui Zhou 0005, Qingguo Zhou, Yong Sheng, Kuanching Li |
J. Supercomput. | 3 |
| 2013 | Erratum to: XtratuM/PPC: a hypervisor for partitioned system on PowerPC processors
Rui Zhou 0005, Qingguo Zhou, Yong Sheng, Kuanching Li |
J. Supercomput. | 3 |
| 2008 | Detecting 802.11 MAC Layer Spoofing Using Received Signal StrengthabstractMAC addresses can be easily spoofed in 802.11 wireless LANs. An adversary can exploit this vulnerability to launch a large number of attacks. For example, an attacker may masquerade as a legitimate access point to disrupt network services or to advertise false services, tricking nearby wireless stations. On the other hand, the received signal strength (RSS) is a measurement that is hard to forge arbitrarily and it is highly correlated to the transmitter's location. Assuming the attacker and the victim are separated by a reasonable distance, RSS can be used to differentiate them to detect MAC spoofing, as recently proposed by several researchers. By analyzing the RSS pattern of typical 802.11 transmitters in a 3-floor building covered by 20 air monitors, we observed that the RSS readings followed a mixture of multiple Gaussian distributions. We discovered that this phenomenon was mainly due to antenna diversity, a widely-adopted technique to improve the stability and robustness of wireless connectivity. This observation renders existing approaches ineffective because they assume a single RSS source. We propose an approach based on Gaussian mixture models, building RSS profiles for spoofing detection. Experiments on the same testbed show that our method is robust against antenna diversity and significantly outperforms existing approaches. At a 3% false positive rate, we detect 73.4%, 89.6% and 97.8% of attacks using the three proposed algorithms, based on local statistics of a single AM, combining local results from AMs, and global multi-AM detection, respectively. Yong Sheng, Keren Tan, David Kotz, Andrew T. Campbell |
INFOCOM | 1 |
| 2005 | Distance measures for nonparametric weak process modelsabstractNonparametric versions of hidden Markov models, what we call weak models, are robust for process detection and easy to construct, as the assumption of knowing precise probabilities in HMMs is weakened to {0,1}-values of reachabilities. Weak models are shown to be equivalent to DFAs/ NFAs. The concept of minimal unifilar weak model (/spl mu/-WM) is introduced. The spectral radius of the transition matrix of /spl mu/-WM determines the growth rate of acceptable observation sequences. An absolute weak model distance is defined for model clustering purpose, while a relative distance is a measure of how fast the performance of detection gets improved as more observations arrive. Convergence of the distance measures is proved. Yong Sheng, George Cybenko |
SMC | 1 |
| 2005 | A parallel decision tree-based method for user authentication based on keystroke patternsabstractWe propose a Monte Carlo approach to attain sufficient training data, a splitting method to improve effectiveness, and a system composed of parallel decision trees (DTs) to authenticate users based on keystroke patterns. For each user, approximately 19 times as much simulated data was generated to complement the 387 vectors of raw data. The training set, including raw and simulated data, is split into four subsets. For each subset, wavelet transforms are performed to obtain a total of eight training subsets for each user. Eight DTs are thus trained using the eight subsets. A parallel DT is constructed for each user, which contains all eight DTs with a criterion for its output that it authenticates the user if at least three DTs do so; otherwise it rejects the user. Training and testing data were collected from 43 users who typed the exact same string of length 37 nine consecutive times to provide data for training purposes. The users typed the same string at various times over a period from November through December 2002 to provide test data. The average false reject rate was 9.62% and the average false accept rate was 0.88%. Yong Sheng, Vir V. Phoha, S. M. Rovnyak |
IEEE Trans. Syst. Man Cybern. Part B | 1 |