Tao Cai 0003

dblp:90/443-3 · DBLP profile ↗
← Back
25ranked-venue papers
13as first author
20since 2021 · last 2026
0000-0003-1423-2710ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 2 first-author · 9 since 2021Systems, architecture and hardware · 8 · 7 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 4 first-author · 5 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 first-author · 2 since 2021Computer networks · 1
YearPublicationVenuePosition
2026 Point-Patch Transformer for Multivariate Time Series Forecasting
Wenyi Xiao, DeJiao Niu, Tao Cai 0003, Yikang Deng, Liushan Zhang, Qiujing Huang
DASFAA (6)3
2026 A Cache Friendly LSM Tree Based on Extendible Hash
abstract
ABSTRACT The LSM‐tree based on extendible hashing has been adopted in key‐value storage systems due to its high write throughput, good scalability, and balanced read‐write performance. However, it faces several challenges when accessing SSTables, including low cache efficiency, uneven data distribution across hash buckets, and frequent directory expansions. To address these issues, this paper proposes a cache‐friendly LSM‐tree based on extendible hashing. To solve the problem that similar keys in SSTables are not stored adjacently in SSTables, a key‐to‐value mapping strategy based on Locality‐Sensitive Hashing (LSH) is employed. Second, to address the uneven data distribution across hash buckets in extendible hashing, a logically uniform extendible hashing scheme is designed, along with a novel HTable structure to replace the traditional SSTables in LSM‐tree. In addition, an HTable indexing strategy based on the LSH‐Cuckoo filter is proposed to accurately locate the target HTable. Based on the Intel Optane DC Persistent Memory driver, a prototype of a cache‐friendly key‐value storage system named DLMS was implemented on non‐volatile memory (NVM), and evaluated using the YCSB benchmark. Experimental results show that, compared to the LSM‐tree‐based storage system RocksDB, DLMS achieves an average improvement of 9.8% in read throughput and 11.9% in write throughput, while reducing insertion latency by 7.8%.
Tao Cai 0003, Qiujing Huang, Jianfei Dai, DeJiao Niu, Yikang Deng
Concurr. Comput. Pract. Exp.1
2026 A stable self-organized spatial pooling algorithm for hierarchical temporal memory
DeJiao Niu, Tao Cai 0003
Expert Syst. Appl.3
2026 DSAC-Hash: Distribution-Similarity-Aware Cross-modal Hashing
Mutaz Ibrahim Mohammed Ahmed Ibrahim, DeJiao Niu, Tao Cai 0003
Image Vis. Comput.3
2025 MMSEH: Metadata Management System Based on New Extendible Hash
abstract
Metadata management systems are essential for locating and managing files. However, traditional strategies based on B-trees or hash structures are primarily designed for block-based interfaces used in HDDs and SSDs, making them less effective in leveraging the byte-addressability and high-speed characteristics of emerging non-volatile memory (NVM) devices. In this paper, we propose a new metadata management framework that incorporates a custom hash function to convert file access paths into hash values, thereby improving search efficiency while preserving file hierarchy. We also design a new extendible hashing scheme, FBHASH, which distinguishes between file and directory metadata, extends hash directories and data buckets, and enhances overall metadata access efficiency. We implement a prototype system, HANVFS, by embedding this design directly into the PMEM driver. Using three benchmark tools—Filebench, Fio, and YCSB—we evaluate its performance. The results show that HANVFS can improve read/write throughput by up to 30.6% and read/write bandwidth by up to 24.7% compared to NOVA and Ext4 loaded on the PMEM driver.
Tao Cai 0003, Yikang Deng, DeJiao Niu, Qiangqiang Ni
HPCC1
2025 ZQuery-LSM-Tree for ZNS SSD
Tao Cai 0003, Danping Zou, DeJiao Niu, Zihao Yinyi, Qiujing Huang, Yikang Deng
ICA3PP (5)1
2025 SWD-HTM: A Novel Hierarchical Temporal Memory Model Integrating Optimal Transport and Sparse Autoencoder
DeJiao Niu, Tao Cai 0003
PRICAI3
2024 A LLC-Friendly LSM-tree
abstract
The LSM-tree is important for the efficiency of Key-value store. There is large capacity Last-Level Cache in modern CPU. However, it is crucial to effectively utilize LLC in LSM-tree, which is currently an urgent issue that needs to be addressed. In this paper, we propose a LLC-friendly LSM-tree named LLC-LSM-tree. The structure of it is given to distribute the KV pairs among LLC, memory and storage devices. Meanwhile, the LLC-oriented KV pair packet strategy is presented to optimize the LLC slices and combines Direct Data I/O (DDIO) to improve the write throughput. Additionally, FinTable is designed as a replacement for the existing SSTable in the LSM-tree to reduce the size of the index and adapt to the limited capacity of the LLC. Moreover, we propose a prefetching strategy for FinTable to enhance query efficiency by the characteristics of different operation. We implement a prototype of a Key-value store with LLC-LSM-tree called CFMS based on Intel open source Optane DC Persistent Memory drivers and DDIO source code. The prototype is tested and analyzed using YCSB and db_bench. The results show that compared to other LSM-tree-based Key-value storage systems like RocksDB, WiscKey, and ListDB, CFMS can achieve an average improvement of 19.37% in read throughput and 8.43% in write throughput, while reducing write amplification by 27.91%.
Tao Cai 0003, Jianfei Dai, DeJiao Niu, Qiangqiang Ni
ASAP1
2024 One Process Spatiotemporal Learning of Transformers via Vcls Token for Multivariate Time Series Forecasting
Tao Cai 0003, Haixiang Wu, DeJiao Niu, Xuewen Xia, Jingzehua Xu
ICANN (6)1
2024 A Hierarchical Multi-scale Cortical Learning Algorithm for Time Series Forecasting
DeJiao Niu, Tao Cai 0003, Xuewen Xia
ICIC (4)3
2024 DNZ-LSM-Tree for Hybrid Storage Systems
abstract
The Log-Structured Merge tree (LSM-tree) can transform random writes into sequential writes, adapting to the high sequential write performance of external storage devices. This has led to its widespread application in various Key-Value (KV) storage systems. However, the LSM-tree has issues such as write amplification and periodic sharp declines in performance. Non-Volatile Memory (NVM) and Zoned Namespace (ZNS) Solid State Drive (SSD) are emerging storage devices that differ significantly from traditional SSDs. Directly applying LSM-tree to NVM and ZNS SSD is challenging due to their unique advantages and characteristics. This paper proposes the DNZ-LSM-Tree, tailored for hybrid external storage constructed using NVM and ZNS SSD. Initially, we present the structure of the DNZ-LSM-Tree, which reconstructs the LSM-tree by utilizing the distinct features of memory, NVM, and ZNS SSD. Subsequently, to address the cascading compaction issue that significantly affects the performance of LSM-tree, we construct a Compaction Cache in NVM and design a layered distribution strategy for Sorted String Tables (SSTables) with cascading compactions. In addition, the zone allocation strategy based on key overlap ratio estimation and the wear leveling strategy for DNZ-LSM-Tree are designed to manage the zones of ZNS SSD. Finally, a prototype of KV storage system based on hybrid storage devices called HNZMS is implemented by employing DNZ-LSM-Tree and tested by YCSB. The results indicate that compared to the storage system called ListDB based on LSM-tree, HNZMS can increase the write throughput by 36.2% and reduce the write amplification by 51.9%.
Tao Cai 0003, DeJiao Niu, Qiangqiang Ni, Zihao Yinyi, Danping Zou
ISPA1
2024 A Hierarchical Heterogeneous IoT Time Series Data Index for NVM
abstract
The index plays an important role in the performance of IoT time series data storage systems. However, the current index designed for HDD or SSD can not adapt to the characteristics of IoT time series data and effectively leverage the performance advantages of NVM. A new hierarchical heterogeneous index is designed for IoT time series storage system based on NVM. The structure is given first and the group consists of several data blocks used to manage IoT time series data. The ordered construction strategy is designed for skip list by the sustained generation of IoT time series data and creating the index for each IoT time series data block group. Meanwhile, a compression and reconstruction strategy for skip list is given to effectively utilize NVM. Then, a TS- Radix tree is presented to index IoT time series data block groups by the temporal characteristics of IoT time series data. Finally, a prototype of it is implemented and YCSB-TS is used to evaluate. The results show that this new index can effectively improve the throughput of random and range queries by up to 262.4%, surpassing the performance of InfluxDB, OpenTSDB, and KairosDB.
Tao Cai 0003, Tianle Lei, DeJiao Niu, Jianfei Dai, Qiangqiang Ni
SMC1
2024 CSTformer: Cross Spatio-Temporal Transformer for Multivariate Time Series Forecasting
abstract
Transformer-based models have shown remarkable success in Multivariate Time Series Forecasting (MTSF). Previous methods apply Attention mechanisms to capture temporal and spatial (variable) dependencies separately. However, they struggle to model intricate local spatiotemporal correlations. To address this limitation and enhance performance, we propose CSTformer, a novel Transformer-based model that enables capturing Cross Spatio- Temporal (CST) dependency for MTSF. In CSTformer, through a Variate Compete Linear Attention (VCLA) mechanism, each variable efficiently achieves specific CST features, which compete against the backdrop of the local multivariate. Additionally, we develop a Mixture of Latent (MoL) module to provide adaptive predictions for variables with varying degrees of CST dependencies. Our experimental results on nine benchmarks indicate that, compared with the state-of-the-art method, CSTformer yields a 2.7% relative improvement.
Tao Cai 0003, Haixiang Wu, DeJiao Niu
SMC1
2023 A New Spatial Pooler Algorithm Based on Heterogeneous Hash Group
abstract
Hierarchical Temporal Memory (HTM) is an emerging neural network model that simulates the working principles of human brain neocortex. HTM forms sparse distributed representation (SDR) and learns temporal relations of sequential data through two core learning algorithms, spatial pooler (SP) and temporal memory (TM). However, the conventional SP algorithm is time consuming as all columns need to be iteratively trained and a large amount of computations is required to generate the SDR. Moreover, the stability of SDR is hard to guarantee. To improve efficiency and construct stable representation, we propose a new SP algorithm based on heterogeneous hash group (SP_HHG), where three different hash functions are designed for fast data representation. The former two functions originate from the locality sensitive hash and use different sensitive fields, and the third function is essentially a random hash. The representations generated by the combined hash values satisfy the principles of SDR in HTM. Extensive experiments are carried out on three sequential datasets and the results show that SP_HHG can effectively improve the efficiency of SP and produce more stable SDR which lead to the enhanced performance of HTM.
DeJiao Niu, Tao Cai 0003
IJCNN3
2023 A new hierarchical temporal memory based on recurrent learning unit
abstract
Hierarchical temporal memory is an emerging machine learning technology that aims to model the structural and algorithmic properties of the neocortex. It is particularly suitable for learning and predicting sequential data. However, when dealing with long time series or complex sequences, the accuracy is relatively lower than desired. In this paper, a novel hierarchical temporal memory based on recurrent learning unit is proposed, where a feedback mechanism is involved into the model. The original cell is extended with a recurrent unit to capture long temporal dependencies of synaptic connections between neurons. The temporal pooler algorithm is then modified to adapt to the recurrent learning unit, and the supervised gradient information is combined with the Hebbian synaptogenesis learning rule in speeding up the training. The prototype of the proposed hierarchical temporal memory is implemented and extensive experiments are carried out on two public datasets under various settings. Experimental results show that the proposed model obtains an accuracy increase by up to 32% and a perplexity drop by up to 14% on sequence prediction and text generation tasks, respectively, which indicates the hierarchical temporal memory with recurrent feedback outperforms the original model on sequence learning.
DeJiao Niu, Tianquan Liu, Tao Cai 0003
J. Exp. Theor. Artif. Intell.4
2023 NEHASH: high-concurrency extendible hashing for non-volatile memory
abstract
Extendible hashing is an effective way to manage increasingly large file system metadata, but it suffers from low concurrency and lack of optimization for non-volatile memory (NVM). In this paper, a multilevel hash directory based on lazy expansion is designed to improve the concurrency and efficiency of extendible hashing, and a hash bucket management algorithm based on groups is presented to improve the efficiency of hash key management by reducing the size of the hash bucket, thereby improving the performance of extendible hashing. Meanwhile, a hierarchical storage strategy of extendible hashing for NVM is given to take advantage of dynamic random access memory (DRAM) and NVM. Furthermore, on the basis of the device driver for Intel Optane DC Persistent Memory, the prototype of high-concurrency extendible hashing named NEHASH is implemented. Yahoo cloud serving benchmark (YCSB) is used to test and compare with CCEH, level hashing, and cuckoo hashing. The results show that NEHASH can improve read throughput by up to 16.5% and write throughput by 19.3%.
Tao Cai 0003, DeJiao Niu, Yueming Ma, Tianle Lei, Jianfei Dai
Frontiers Inf. Technol. Electron. Eng.1
2022 A New IoT Storage System Based on Raw NVM
abstract
NVM storage devices have the advantages of high read-write speed, non-volatile and large capacity, which provides support for efficient storage and management of a large number of time series data collected by IoT devices. However, how to study a new IoT storage system for time series data according to the advantages of NVM storage devices is an important issue that needs to be solved. This paper first analyses the characteristics of accessing to IoT time series data and then based on the characteristics of NVM storage devices, a multi-granularity auto-converting structure for time series data is designed. It not only reflects the timeliness of accessing to IoT time series data, but also avoids additional data replication, which improves the storage efficiency of IoT time series data. A timeliness-based heterogeneous query strategy is designed to improve the query efficiency according to the IoT time series data storage structure and accessing characteristics. The prototype of a new IoT storage system based on raw NVM named NBTSMS is implemented based on the Intel open-source NVM storage device driver PMEM. InfluxDB, OpenTSDB, and TimescaleDB are used for evaluation with YCSB-TS. Results show that NSTSMS can improve write throughput by 137%, random query throughput by 153.7%, scan throughput by 189.4%, and mixed-operations throughput by 55%.
Tao Cai 0003, Yueming Ma, DeJiao Niu, Tianle Lei, Jianfei Dai
IEEE Big Data1
2022 A Stable Spatial Pooling Algorithm for Hierarchical Temporal Memory Network
abstract
Hierarchical Temporal Memory (HTM) is an emerging neural network technology which is inspired by the structure and the working principles of human neocortex. HTM uses a set of active columns to represent the input and inhibits frequently activated columns by a boosting mechanism. However, the boosting mechanism may cause an instable representation and thus affects the HTM performance. To address this issue, we propose a stable spatial pooling algorithm for HTM, where the column loadability is introduced, and a new column activation strategy based on the loadability is presented. We implement the proposed algorithm and carry out various experiments on both the synthetic and real-world sequential datasets. The results show that our HTM can generate more stable internal representation for the input data. The prediction accuracy of the proposed HTM outperforms that of the conventional HTM and the popular LSTM networks.
DeJiao Niu, Tao Cai 0003
ICTAI3
2021 The Novel Efficient Transformer for NLP
Benjamin Mensa-Bonsu, Tao Cai 0003, Tresor Y. Koffi, DeJiao Niu
KSEM2
2021 A Two Tier Hybrid Metadata Management Mechanism for NVM Storage System
Tao Cai 0003, Fuli Chen, DeJiao Niu, Yueming Ma
NPC1
2018 The Classified and Active Caching Strategy for Iterative Application in Spark
abstract
The Resilient Distributed Dataset (RDD) cache is an important way to improve the efficiency of application on Spark. There are some specific characteristics of iterative application. We design the classified and active caching strategy for iterative application. It can reduce the time delay of RDD creation of data set for iterative application and avoid creating or recomputing it again during the iterative application. At the same time, the RDD of parameter set should be renewed and cached after every iteration to reduce the time overhead of the subsequent iteration. The prototype is implemented based on Spark and one iterative application is used to test and compared with Spark. The results show that the classified and active caching strategy can reduce the 12%-45% time overhead of iterative application.
DeJiao Niu, Tao Cai 0003, ZhiPeng Chen
ICCCN3
2018 ALSTM: Adaptive LSTM for Durative Sequential Data
abstract
Long short-term memory (LSTM) network is an effective model architecture for deep learning approaches to sequence modeling tasks. However, the current LSTMs can't use the property of sequential data when dealing with the sequence components, which last for a certain period of time. This may make the model unable to benefit from the inherent characteristics of time series and result in poor performance as well as lower efficiency. In this paper, we present a novel adaptive LSTM for durative sequential data which exploits the temporal continuance of the input data in designing a new LSTM unit. By adding a new mask gate and maintaining span, the cell's memory update is not only determined by the input data but also affected by its duration. An adaptive memory update method is proposed according to the change of the sequence input at each time step. This breaks the limitation that the cells calculate the cell state and hidden output for each input always in a unified manner, making the model more suitable for processing the sequences with continuous data. The experimental results on various sequence training tasks show that under the same iteration epochs, the proposed method can achieve higher accuracy, but need relatively less training time compared with the standard LSTM architecture.
DeJiao Niu, Zheng Xia, Tao Cai 0003, Tianquan Liu, Yongzhao Zhan 0001
ICTAI4
2018 Improving Write Performance and Extending Endurance of Object-Based NAND Flash Devices
abstract
Write amplification is a major cause of performance and endurance degradations in NAND flash-based storage systems. In an object-based NAND flash device (ONFD), two causes of write amplification are onode partial update and cascading update. Here, onode is a type of small-sized object metadata, and multiple onodes are stored in one NAND flash page. Updating one onode invokes partial page update (i.e., onode partial update), incurring unnecessary migration of the un-updated data. Cascading update denotes updating object metadata in a cascading manner due to object data update or migration. Although there are only several bytes that need to be updated in the object metadata, one or more pages have to be re-written accordingly. In this work, we propose a system design to alleviate the write amplification issue in the object-based NAND flash device. The proposed design includes (1) a multi-level garbage collection technique to minimize unnecessary data migration incurred by onode partial update and (2) a B+ table tree, Semantics-Aware Flexible (SAF) data layout, and selective cache design to reduce the write operations associated with cascading update. To guarantee system consistency, we also propose a power failure handling technique. Experiment results show that our proposed design can achieve up to 20% write reduction compared to the best states of the art.
Jie Guo 0002, Chuhan Min, Tao Cai 0003, Yiran Chen 0001
ACM Trans. Embed. Comput. Syst.3
2016 NVMCFS: Complex File System for Hybrid NVM
abstract
Due to the price limitation and the number of DIMM slot, Byte and Block addressable NVM devices should coexist in the massive Storage Class Memory(SCM). But they have many differences such as interface, access granularity, I/O performance and storage capacity. Therefore, the existing main memory and file system management algorithms cannot be applied in it directly. In this paper, we present a complex file system named NVMCFS for Hybrid NVM. The head-tail layout and space management based on two layer radix-tree is provided to unify logic space between two type NVM devices. The complex file structures, dynamic file data distributed strategy, buffer for an individual file and asymmetric call in strategy are used to speed up the access response and improve I/O performance. The hybrid consistent mechanism is given and it can reduce the performance loss NVMCFS. Finally, the prototype of NVMCFS is implemented and evaluated by various benchmark. Compared to Ext2 and Ext4 on PMBD, NVMCFS improves sequential read speed 4.4x and 5x, sequential write speed 2.8x and 1.9x, IOPS 45% and 62%, and has the similar I/O performance with PMFS. At the same time, NVMCFS reduces the total overhead of consistency by 50%~92% compared to Ext4.
Tao Cai 0003, DeJiao Niu, Yeqing Zhu
ICPADS1
2005 Research and implementation of graphic service grid system
abstract
In this paper, we present a special model G-GRID for graphic service grid system. The main function of each layer in the model is described in detail, the service description and system architecture are defined as well. Then we propose the key algorithms of services discovery, management and schedule of the grid system. Some implementation interfaces of the graphic service grid system are provided in the end. Then prove the validity of graphic service grid system.
Tao Cai 0003, Shiguang Ju, Xiangmei Song, DeJiao Niu
CSCWD (1)1