VLDB 2026 Research / reviewers in the wild / expert
Chengxiang Si
dblp:62/5048
· DBLP profile ↗
32ranked-venue papers
2as first author
17since 2021 · last 2026
0000-0003-2646-6100ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 10 · 2 first-authorGraphics, computer vision, multimedia, augmented reality and games · 7 · 7 since 2021Artificial intelligence and machine learning · 6 · 4 since 2021Computer networks · 5 · 4 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Security and privacy · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Software engineering, systems software and programming languages · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | You can run but you can never hide: A multi-module collaborative detection framework based on network traffic
Chengxiang Si, Zhou Zhou 0007, Zhenyu Cheng 0001, Peishuai Sun |
Comput. Networks | 2 |
| 2025 | MS-NHHO: A Swarm Intelligence Optimization Algorithm Incorporating Cognitive Science for Malicious Traffic Detection
Zhou Zhou 0007, Chengxiang Si, Qingyun Liu 0001 |
CogSci | 3 |
| 2025 | APTSniffer: Detecting APT Attack Traffic Using Retrieval-Augmented Large Language ModelsabstractAdvanced Persistent Threats (APT) differ from traditional attacks by using more complex and covert strategies for long-term assaults, posing a severe threat to organizational and national security. Due to problems like the shortage of APT traffic data and encrypted traffic obfuscation, existing methods cannot accurately identify APT traffic with just a few traffic samples. To overcome the above limitation, we propose a novel encrypted APT traffic detection model, APTSniffer, which combines large language models (LLM) and retrieval-augmented technology. APTSniffer utilizes the few-shot inference and generalization abilities of large language models by converting raw traffic data into natural language inference examples understandable by the LLM. Experimental results show that, compared to other baseline models, APTSniffer exhibits SOTA performance. It achieves F1 scores above 97% on three APT datasets, making it practically applicable for APT traffic detection tasks. Chengxiang Si, Zhou Zhou 0007, Chenxu Wang 0006, Peishuai Sun, Qingyun Liu 0001 |
ICASSP | 2 |
| 2025 | PORTIA: A Multi-Granularity APT Detection Model Based on Provenance GraphsabstractAdvanced Persistent Threats (APTs) have become a major cybersecurity threat due to their stealthy attack methods and long latency periods. Traditional signature-based detection struggles to detect novel attacks, and while unsupervised methods using Graph Neural Networks (GNNs) can model system behavior, they face challenges in handling large-scale provenance graphs and accurately incorporating system operational states for detection.This paper presents PORTIA, a multi-granularity APT detection model based on graph representation learning. PORTIA constructs provenance graphs by integrating temporal information from audit logs and uses a graph mask autoencoder to model normal system behavior, detecting anomalies through embedding shifts. As an unsupervised model, PORTIA can swiftly identify anomalous system states without relying on attack signatures, achieving fine-grained detection by incorporating system operational states. Evaluations on multiple datasets show that PORTIA detects both standard and APT attacks with high precision, outperforming existing detection systems. Haoqiang Wang, Zhou Zhou 0007, Chengxiang Si, Qingyun Liu 0001 |
IJCNN | 6 |
| 2025 | MOLE: Provenance Graph Generation Framework Based on LLM PromptingabstractIn the increasingly complex landscape of cyber-attacks, logs have become a critical source of data for detecting system threats. Currently, most log-based detection systems rely on converting audit logs into provenance graphs during the process of attack investigation. However, this construction process is still heavily dependent on manually written code with regular expressions tailored to each specific log type. In this paper, we propose MOLE, a provenance graph generation framework based on prompting with large language models (LLMs). Unlike traditional approaches, MOLE does not rely on prior knowledge and is adaptable to diverse types of log data. The framework automatically generates provenance graph extraction templates through instruction generation and parses logs locally to produce the final provenance graph.MOLE leverages the log patterns and structures learned by LLMs from large-scale data during training. As a result, tasks that previously required several days of manual coding to generate a provenance graph can now be completed in just a few minutes. Furthermore, when processing 50 million log entries, the entire provenance graph generation process consumed only 20k tokens. Haoqiang Wang, Zhou Zhou 0007, Chengxiang Si, Qingyun Liu 0001 |
IJCNN | 5 |
| 2025 | FlowMiner: A Powerful Model Based on Flow Correlation Mining for Encrypted Traffic Classification
Chengxiang Si, Zhenyu Cheng 0001, Chenxu Wang 0006, Jiang Xie 0004, Peishuai Sun, Qingyun Liu 0001 |
INFOCOM | 2 |
| 2025 | MTDIR: A Malicious Traffic Detection Model Based on the Image Retrieval PerspectiveabstractTraffic typically reflects network behavior, enabling the detection of network attacks through malicious traffic analysis. Existing methods often suffer from feature redundancy during extraction, reducing model efficiency. Additionally, these methods struggle to capture long-range dependencies, impacting feature learning. A single-perspective approach also limits the ability to learn universal patterns from multiple features, restricting the model's applicability. To address these shortcomings, this paper proposes a malicious traffic detection model based on image retrieval (MTDIR), which leverages depthwise separable convolution for multi-level feature retrieval, reducing computational overhead and redundancy. By learning features across multiple dimensions, MTDIR enhances detection performance. Ablation experiments confirm each module's effectiveness. MTDIR outperforms the control group in detection accuracy while maintaining low time overhead and considerable generalization and robustness, making it highly applicable. Zhou Zhou 0007, Chengxiang Si |
ICMR | 4 |
| 2025 | Two Heads are Better than One: A Network Attack Detection Model Based on Multimodal and Multimedia RetrievalabstractTraffic, as a carrier of network behavior, can be used to detect attacks through malicious traffic detection. However, the existing methods have limited effectiveness in detecting covert attacks and are insufficiently resistant to interference in complex environments. In addition, the redundant information in the feature extraction process will reduce the model's efficiency. Therefore, this paper proposes a network attack detection model based on multimodal and multimedia retrieval (NADMR), which utilizes two modalities, namely, image and time series, to mine the key features of the traffic from a multi-dimensional perspective. Secondly, lightweight spatial attention and split channel attention are designed to extract discriminative features of image modality and analyze feature information of time series modality from multiple domains. Meanwhile, ghost convolution is introduced to improve efficiency. The ablation experiments verify the effectiveness of each technique, and the comparison experiments show that NADMR performs better in detection performance and efficiency. Its good generalization ability and robustness, as well as lower complexity, make it suitable for more scenarios. Zhou Zhou 0007, Chengxiang Si |
ICMR | 3 |
| 2025 | Zero in on the Target: A Composite Robust Model for Retrieving Information in Traffic Data to Discover Network AttacksabstractWith the popularization of the Internet and the diversification of attack methods, web security has become an important part of information security. As a carrier of network behavior, traffic can reveal attack behaviors in the Web environment through malicious traffic detection. Since images can fully express spatial features and local associations, it is feasible to visualize traffic as images and retrieve key feature information to detect malicious traffic. However, existing methods are prone to redundancy during feature extraction. Secondly, a single perspective makes it difficult to learn patterns with universality from diverse feature information. In addition, the selection of segmentation thresholds in the preprocessing is closely related to the model's information retrieval effect. The commonly adopted preset thresholds are difficult to cope with the changes in traffic data, limiting the applicability of the existing methods. Therefore, this paper proposes a Multi-Module-Based Composite Robust Model for Network Attack Detection (MCNAD). The model adopts depthwise separable convolution (DSC) to reduce redundant information, proposes a multi-scale feature learning module to enhance the model characterization ability, and proposes a gray level co-occurrence matrix segmentation algorithm with adaptive threshold (GLCM-AT) to optimize data preprocessing. The results show that MCNAD improves detection performance with better detection efficiency, generalization ability, and robustness, demonstrating its wide applicability in multiple scenarios. Chengxiang Si, Zhenyu Cheng 0001 |
ACM Multimedia | 2 |
| 2025 | AdvTG: An Adversarial Traffic Generation Framework to Deceive DL-Based Malicious Traffic Detection ModelsabstractDeep learning-based (DL-based) malicious traffic detection models are effective but vulnerable to adversarial attacks. Existing adversarial attacks have shown promising results when targeting traffic detection models based on statistics and sequence features. However, these attacks are less effective against models that rely on payload analysis. The main reason is the difficulty in generating semantic, compliant, and functional payloads, which limits their practical application. Peishuai Sun, Xiao-chun Yun, Chengxiang Si, Jiang Xie 0004 |
WWW | 5 |
| 2025 | Sample analysis and multi-label classification for malicious sample datasets
Jiang Xie 0004, Xiao-chun Yun, Chengxiang Si |
Comput. Networks | 4 |
| 2024 | MLMTD: A Multi-Layer Malicious Traffic Detection Model Based on Multi-Branch Octave Convolution and Attention MechanismabstractMalicious traffic detection is important for the safe operation of cyberspace. Existing methods are difficult to extract discriminative features, leading to the detection rate bottleneck. In addition, the performance is significantly degraded in sample imbalanced scenarios, with poor generalization ability and insufficient scalability. Therefore, this paper proposes a multi-layer malicious traffic detection model based on multi-branch octave convolution and attention mechanism (MLMTD), which adopts multiple modules to extract more diverse feature information and learn important features more adequately in multiple dimensions. Ablation experiments validate the effectiveness of each module. The results in several experimental scenarios show that MLMTD can achieve better results with fewer features and attain superior robustness and generalization ability compared to the control models. Chengxiang Si, Zhenyu Cheng 0001 |
ICASSP | 2 |
| 2024 | A Targeted Adversarial Attack Method for Multi-Classification Malicious Traffic DetectionabstractLeveraging deep learning to detect malicious network traffic is a crucial technology in network management and network security. However, deep learning security has raised concerns among scholars. In this work, we explore executing targeted adversarial attacks for multi-classification malicious traffic detection with limited interactions. Specifically, we constrain the number of interactions with detection and employ a hop-skip-jump attack (HSJA) to generate a small number of adversarial samples. These adversarial samples are then heuristically used to train a generative adversarial network (GAN) to generate a substantial quantity of adversarial samples. Experiments demonstrate that our method is more adversarial and displays a certain degree of generalization compared with other methods. Peishuai Sun, Chengxiang Si, Zhenyu Cheng 0001, Qingyun Liu 0001 |
ICASSP | 2 |
| 2024 | MTDM-MS: A Malicious Traffic Detection Model Based on Multi-Category SignalsabstractThe demand for malicious traffic detection continues to rise along with the development of the Internet. Existing methods perform with flaws in complex feature extraction processes and interference by obfuscation techniques and other means. In addition, the performance is unstable in new scenarios, and the generalization ability is not good enough. Therefore, this paper proposes a malicious traffic detection model based on multi-category signals (MTDM-MS), which adopts multiple modules to extract the features of text sequence signals and image signals of the traffic, respectively, to realize the interaction of various feature information and improve the model's characterization ability. Ablation experiments verify the effectiveness of each module. Experimental results in several datasets show that MTDM-MS possesses considerable detection performance and generalization ability with a 2.2% to 8.4% improvement in macro-F1 compared with the control models. Chengxiang Si, Zhenyu Cheng 0001 |
ICME | 2 |
| 2024 | Blockchain-Based Covert Communication: A Detection Attack and Efficient ImprovementabstractCovert channels in blockchain networks achieve undetectable and reliable communication, while transactions incorporating secret data are perpetually stored on the chain, thereby leaving the secret data continuously susceptible to extraction. MTMM (IEEE Transactions on Computers 2023) is a state-of-the-art blockchain-based covert channel. It utilizes Bitcoin network traffic that will not be recorded on the chain to embed data, thus mitigating the above issues. However, we identify a distinctive pattern in MTMM, based on which we propose a comparison attack to accurately detect MTMM traffic. To defend against the attack, we present an improvement named ORIM, which exploits the permutation of transaction hashes within inventory messages to transmit secret data. ORIM leverages a pseudo-random function to obscure the transaction hashes involved in the permutation to ensure unobservability. The obfuscated values, rather than the original transaction hashes, are utilized to encode the confidential data. Furthermore, we introduce a variable-length encoding scheme predicated on complete binary trees. This scheme considerably amplifies the bandwidth and facilitates efficient encoding and decoding of secret data. Experimental results indicate that ORIM maintains unobservability and that ORIM’s bandwidth is approximately$3.7\times $of MTMM. Zhuo Chen 0001, Liehuang Zhu, Peng Jiang 0007, Zijian Zhang 0001, Chengxiang Si |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2022 | A Secure and Anonymous Communicate Scheme over the Internet of ThingsabstractAnonymous exchange of data has a strong demand in many scenarios. With the development of IoT and wireless networks, plenty of smart devices are interconnected through wireless technologies such as 5G and Wi-Fi, making it possible to use them for information exchanging. The authors find a P2P network model for secure and anonymous communication, which is a typical Crowds system and the operating mechanism meets the characteristics of limited-resources of IoT devices. Based on this network model, the authors design a lightweight communication scheme for the remote-control system in this work, using two kinds ofVirtual-Spaces to achieve the purpose of identity announced and data exchanged. The authors implemented a prototype system of the scheme and tested it over theFreenet, proving that the scheme can effectively resist the impact of flow analysis on the anonymity of communication while ensuring communication data security. By analyzing the scheme’s performance, the author believes that the scheme is practical and is suitable for scenarios that are not time-sensitive but require high anonymity. Qindong Sun, Chengxiang Si, Yanyue Xu, Shancang Li, Prosanta Gope |
ACM Trans. Sens. Networks | 3 |
| 2021 | DBUL: A User Identity Linkage Method across Social Networks Based on Spatiotemporal DataabstractWith the increasing availability of spatiotemporal data, user identity linkage across social networks based on spatiotemporal data has attracted more and more attention. The existing methods have some problems, such as trajectory processing is not suitable for sparse data, grid based processing leads to information loss and anomaly. To solve the above problems, we propose a DBSCAN clustering based method DBUL to solve the problem of user identity linkage based on spatiotemporal data. According to the sparsity, heterogeneity and imbalance of spatiotemporal data in social networks, this method can represent the user identity as the form of cluster centers, and link user identities by calculating the similarity between cluster center representations. We compare this method with several state-of-the-art user identity linkage methods based on spatiotemporal data on real datasets, and the results show that this method outperforms the baseline methods in terms of effectiveness and efficiency. Chengxiang Si |
ICTAI | 3 |
| 2020 | Parallel Belief Propagation Optimized by Coloring on GPUs
Junteng Hou, Chengxiang Si, Guangjun Wu |
ICA3PP (1) | 2 |
| 2020 | MG-Hybrid: A Strongly Connected Components Detection Algorithm using Multiple GPUsabstractDetection of strongly connected component (SCC) on the GPU has become a fundamental operation to accelerate graph computing. Existing SCC detection methods on multiple GPUs introduce massive unnecessary data transformation between multiple GPUs. In this paper, we propose a novel distributed SCC detection approach using multiple GPUs plus CPU. Our approach includes three key ideas: (1) segmentation and labeling over large-scale datasets; (2) collecting and merging the segmented SCCs; and (3) running tasks assignment over multiples GPUs and CPU. We implement our approach under a hybrid distributed architecture with multiple GPUs plus CPU. Our approach can achieve device-level optimization and can be compatible with the state-of-the-art algorithms. We conduct extensive theoretical and experimental analysis to demonstrate efficiency and accuracy of our approach. The experimental results expose that our approach can achieves 11.2×, 1.2×, 1.2× speedup for SCC detection using NVIDIA K80 compared with Tarjan's, FB-Trim, and FB-Hybrid algorithms respectively. Junteng Hou, Guangjun Wu, Bingnan Ma, Chengxiang Si, Siyu Jia |
ISCAS | 5 |
| 2020 | A lightweight and aggregated system for indoor/outdoor detection using smart devices
Zheng Qin 0003, Houbing Song, Chengxiang Si, Renwei Zhang |
Future Gener. Comput. Syst. | 4 |
| 2020 | DAF: An adaptive computing framework for multimedia data streams analysisabstractWe consider the problem of efficiently online computing/filtering or analysis multimedia streams. In this scenario, we register a large scale of continuous analysis queries to filter pornographic stream items. Each query is a conjunction of filters. For instance, the query “does this image contain a people basking in the beach?” can be resolved by applying the conjunction of water, people, sand, sea filters successively on the stream item. However, the online evaluation of multimedia filters is indeed very expensive, fortunately there usually exist multiple filters shared among a lot of queries. In other words, each filter may occur in multiple queries. An open problem in such a filtering scenario is how to order the filters in an optimal sequence to achieve significant performance. Existing methods are based on a greedy strategy which orders the filters according to three factors (selectivity, popularity, cost). Although all these methods achieve good results, there are still some problems that haven’t addressed yet. First, the selectivity factor is set empirically, which can not adaptively adjust with multimedia stream. Second, the proportion relationships among the three factors (selectivity, cost, popularity) were not considerably explored. Under these observations,in this paper, we propose a Dynamic-Analytic hierarchy process Framework (DAF) which use a time-based compositional forecasting method, which is based on the idea of exponential smoothing, to deal with the factors’ proportion relationships dynamics. Experiments on both synthetic and real lift multimedia streams demonstrate that our proposed framework (DAF) provides much great adaptability in modeling the factors proportion relationships changing over multimedia stream environment. Jun Li 0076, Yanzhao Liu, Chengxiang Si |
Intell. Data Anal. | 5 |
| 2020 | Passive browser identification with multi-scale Convolutional Neural Networks
Saeid Samizade, Chao Shen 0001, Chengxiang Si, Xiaohong Guan |
Neurocomputing | 3 |
| 2020 | Identifying vulnerabilities of SSL/TLS certificate verification in Android apps with static and dynamic analysis
Guangquan Xu, Weixuan Mao, Chengxiang Si, Witold Pedrycz, Wei Wang 0012 |
J. Syst. Softw. | 5 |
| 2019 | Deanonymizing Tor in a Stealthy WayabstractThe Tor network is one of the largest anonymity networks and has attracted millions of users worldwide. Since deployed in October 2002, many forms of attacks against Tor had been proposed, aiming to deanonymize the network. The protocol-level attacks, which can deanonymize users by manipulating a cell, are effective to compromise the anonymity of Tor by controlling the entry node and the exit node in a circuit. However, due to the absence of stealthiness, it may be noticed by the victim since the connection will be released. To address this issue, we present two types of attacks. The Type-I is based on protocol-level strategy but attempts to keep the connections alive by fixing damaged cells in the network, thereby making it much stealthier in some cases and it is quite difficult to be defended. In the Type-II attack, the entry node sends a signal to the accomplice exit node via a special type of outbound cells to prevent the connection from being closed, thus the attack can keep stealthy in more general situations. We also propose some countermeasures to keep the network away from the Type- II attack. An evaluation of the two types of attacks has been performed. And the results showed that both the two types of attacks are effective and pose serious threats to the Tor network. Jianjun Lin, Zhenhao Wu, Chengxiang Si |
IPCCC | 4 |
| 2013 | ADS-B Data Authentication Based on AH ProtocolabstractWith the evolution of traditional civil aviation into "e-enabled" aviation, automatic dependent surveillance-broadcast (ADS-B) system plays an important role to replace radar to become the cornerstone of the next generation air traffic management. However, ADS-B system is a broadcast-type data link and ADS-B signals are unauthenticated, thus inserting a false aircraft into the ADS-B system is easy. In this paper, to filter spoofed targets, we present ADS-B data authentication scheme based on AH protocol. Security analysis demonstrates that the proposed scheme can achieve integrity of ADS-B messages, authenticity of data origin sources and resistance against replay attacks. Rui-dong Chen, Chengxiang Si, Haomiao Yang, Xiaosong Zhang 0001 |
DASC | 2 |
| 2010 | An Effective, Low-Overhead, Improved Replacement Algorithm for Mail Service Applications in Storage SystemabstractThis paper analyzed the performance characteristics of classic 2Q algorithm when it was performed on mail-service workloads, and proposes an improved algorithm, called 2Q*. The simulation results show that 2Q* algorithm can outperform the other replacement algorithms, including the classic 2Q algorithm, for all the cache sizes and various mail-service workloads. To verify the simulation results in real system, we integrated the algorithm into FlexiCache, a partitioned buffer cache system, and joined it with a popular adaptive sequential pre-fetch policy properly. The experiment results verify the effectiveness of 2Q* algorithm for mail service applications. By joint with the pre-fetch policy, the performance is further improved. Moreover, its runtime overhead is also fairly low. Chengxiang Si, Xiaoxuan Meng, Yuanfei Chen, Lu Xu 0001 |
ISPA | 1 |
| 2009 | A Flexible Two-Layer Buffer Caching Scheme for Shared Storage CacheabstractThis paper presents a flexible two-layer buffer caching scheme to improve the performance of storage cache which is used to serve multiple concurrently accessing applications with diverse access patterns. To achieve this, the proposed scheme dynamically partitions the cache among applications. At the first layer, it uses a configurable global cache allocation policy to make adaptive cache allocations in response to the evolving access patterns of competing applications. In contrast to traditional global replacement approach, the allocation policy adopted in our scheme utilizes applicationpsilas marginal utility but not cache demand for allocation, so as to minimize the total number of cache misses. At the second laver, our scheme tries to maximize the utilization of allocated cache blocks by applying each application with an appropriate local replacement algorithm based on its access pattern. We have implemented our scheme in Linux kernel 2.6.18 as a pseudo device driver and measured its performance using various real-life workloads. The experiment results show that compared with Linux page cache, the proposed scheme can reduce the overall response time by up to 2 times with an average of 50% and reduce the overall disk load by up to 31% with an average of 21%. Xiaoxuan Meng, Chengxiang Si, Wenwu Na, Haroon-Ur-Rashid Khan, Lu Xu 0001 |
HPCC | 2 |
| 2009 | Volume Based Metadata Isolation in Blue Whale Cluster File SystemabstractThe mainly traditional File Systems are constructed on single device where the metadata and data access interfere with each other which will lead to performance degradations. In this paper we propose a volume based mechanism in BWFS to separate metadata from data into different devices by isolations both in store location and IO path. Test result has verified the effectiveness of metadata isolation. By separating metadata into high-performance storage, the application IO bandwidth is promoted by 2.6~2.8 times, meanwhile, the OPS is improved by 2~5 times. Jingliang Zhang, Chengxiang Si, Yajun Jia, Jiangang Zhang, Lu Xu 0001 |
HPCC | 2 |
| 2009 | A Replacement Algorithm Designed for the Web Search Engine and Its Application in Storage CacheabstractWith popularity of different kind of search engines on WWW, it requires the backend storage system to provide better physical I/O performance to speedup the query service perceived by end users. However, existing general purpose designed replacement algorithm canpsilat performs well for the web search applications. This paper first studies the access pattern of various real-life web search workload and then propose a new replacement algorithm RED-LRU based on the observed access properties. The simulation results shows that our proposed algorithm uniformly outperform the other replacement algorithms for all the workloads and cache size. To validate the simulation results, we integrate RED-LRU algorithm into a real storage cache DPCache. The experiment results in real system confirm the effectiveness of our proposed algorithm in improving the caching performance for web search application. Moreover, the runtime overhead of RED-LRU is also fairly low in practice. Xiaoxuan Meng, Chengxiang Si, Jiangang Zhang, Lu Xu 0001 |
ISPA | 2 |
| 2009 | P-Cache: Providing Prioritized Caching Service for Storage SystemabstractP-Cache to provide prioritized caching service for storage server which is used to serve multiple concurrently accessing applications with diverse access patterns and unequal importance. Given the replacement algorithm and the application access patterns, the end performance of each individual application in a shared cache is actually determined by its allocated cache resource. So, P-Cache adopts a dynamic partitioning approach to explicitly divide cache resource among applications and utilizes a global cache allocation policy to make adaptive cache allocations to guarantee the preset relative caching priority among competing applications. We have implemented P-Cache in Linux kernel 2.6.18 as a pseudo device driver and measured its performance using synthetic benchmark and real-life workloads. The experiment results show that the prioritized caching service provided by P-Cache can not only be used to support application priority but can also be utilized to improve the overall storage system performance. Its runtime overhead is also smaller compared with Linux page cache. Xiaoxuan Meng, Chengxiang Si, Wenwu Na, Lu Xu 0001 |
ISPA | 2 |
| 2009 | Enhancing the Scalability of Blue Whale Cluster File System in Video-editing EnvironmentabstractScalability, which indicates the capacity of system to support the maximal number of clients, is very important to cluster file system. But in practical applications, the bandwidth that storage devices provide doesnpsilat linearly grow with the increasing number of clients which leads to bad QoS and restricts the scalability of cluster file system. This paper analyses the bottleneck of scalability of Blue Whale Cluster File System (BWFS) in video-editing application, and present a solution that clients employ the spare disk resource to cache data from remote server. By careful chosen of cache strategy and replacement algorithm, clients can access remote data from local disk, which significantly reduce random access to server and efficiently reduce the workload of storage servers. Therefore, the scalability of the whole system is significantly enhanced. We also implement a prototype to validate the solution and the result shows that it can enhance the scalability of the Blue Whale System by 80% at most. Chengxiang Si, Xiaoxuan Meng, Junwei Zhang 0003, Lu Xu 0001 |
NAS | 1 |
| 2008 | A Novel Network RAID Architecture with Out-of-Band Virtualization and Redundant ManagementabstractThe paper presents a novel network RAID storage system based on the out-of-band virtualization architecture and the backend centralized redundant management. The application servers can inquire the mapping information of virtual disk from the out-of-band virtualization server and directly access the storage nodes. The read request can fetch the data from special storage node, while the write request is not only stored into the storage node, and also mirrored into the redundant server by the storage node. The redundant server can cache updated data in local disks with log-structured mode, and calculate the parity of RAID5 in the background process when the system is idle. It relieves the bottle problem of I/O performance and low reliability danger of single controller in the front-end centralized management system. And the layout of RAID1/RAID5 on the data block has acquired the trade-off among performance, reliability and cost. Wenwu Na, Xiaoxuan Meng, Chengxiang Si, Jian Ke, Qingzhong Bu, Lu Xu 0001 |
ICPADS | 3 |