VLDB 2026 Research / reviewers in the wild / expert
Xiaoshe Dong
dblp:50/2755
· DBLP profile ↗
74ranked-venue papers
1as first author
24since 2021 · last 2026
0000-0002-9003-2625ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 44 · 21 since 2021Applied, interdisciplinary, general and emerging computing · 11 · 1 since 2021Human-computer interaction and ubiquitous computing · 5 · 1 first-authorArtificial intelligence and machine learning · 4 · 1 since 2021Software engineering, systems software and programming languages · 3Computer networks · 2Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Portable GPU Kernel Performance Modeling Method Based on LLVM IR Dynamic Feature Prediction
Qiang Wang 0062, Hao Zheng 0004, Yiru Liu, Chaojun Deng, Ziheng Wang 0002, Xiaoshe Dong |
CF | 8 |
| 2026 | GPU Kernel Optimization Beyond Full Builds: An LLM Framework with Minimal Executable Programs
Ruifan Chu, Anbang Wang, Xiuxiu Bai, Xiaoshe Dong |
PAKDD (4) | 5 |
| 2026 | DP-SWAP: Fast Swapping Strategy Based on Dynamic Programming
Weiduo Chen, Xiaoshe Dong, Qiang Wang 0062 |
Future Gener. Comput. Syst. | 2 |
| 2026 | Mapo: Performance model driven GPU memory access code optimization
Xiaoshe Dong, Junkai Cao, Ruifan Chu, Ziheng Wang 0002, Qiang Wang 0062, Xiuxiu Bai |
Future Gener. Comput. Syst. | 2 |
| 2026 | BLG-Tuning: Benchmark-Based Low-Cost General-Purpose I/O Modeling and TuningabstractI/O performance has become a major bottleneck for many data-intensive applications. Each layer of the parallel I/O stack provides parameters that can optimize I/O performance, but determining the optimal performance parameters based on the operating configuration is a challenge. Previous work has required separate performance models for different programs for tuning, which is very costly in term of measurement data. We propose BLG-Tuning: a B enchmark-based L ow-cost G eneral-purpose I/O Modeling and Tuning. BLG-Tuning maps application I/O loads to benchmark parameters and uses the benchmark-trained performance model to achieve I/O performance prediction and thus avoid the additional computing and communication overhead for measurement. For applications, BLG-Tuning collects the application characteristics to calibrate the performance model and improve prediction accuracy. Experience shows that BLG-Tuning predicts the I/O time of MADbench2, Flash-IO, S3D-IO, BT-IO, and LAMMPS with MAPE of 22.2%, 18.2%, 29.3%, 21.5%, and 34.8%, respectively. After tuning, the five applications obtain I/O speedup from 5.6× to 27.3×. Ziheng Wang 0002, Yuchao Wu, Xiaoshe Dong |
ACM Trans. Archit. Code Optim. | 6 |
| 2025 | ATP: Achieving Throughput Peak for DNN Training via Smart GPU Memory ManagementabstractDue to the limited GPU memory, the performance of large DNNs training is constrained by the unscalable batch size. Existing studies partially address the issue of GPU memory limit through tensor recomputation and swapping, but overlook the exploration of optimal performance. In response, we propose ATP, a recomputation and swapping based GPU memory management framework that aims to maximize training performance by breaking GPU memory constraints. ATP utilizes a throughput model and we propose to evaluate the theoretical peak performance achievable by DNN training on GPU, and provide the optimum memory size required for recomputation and swapping. We optimize the mechanisms for GPU memory pool and CUDA stream control, employ an optimization method to search for specific tensors requiring recomputation and swapping, thereby bringing the actual DNN training performance on ATP closer to theoretical values. Evaluations with different types of large DNN models indicate that ATP achieve throughput improvements ranging from 1.14∼ 1.49×, while support model training exceeding the GPU memory limit by up to 9.2×. Weiduo Chen, Xiaoshe Dong, Fan Zhang 0139, Bowen Li 0009, Yufei Wang 0008, Qiang Wang 0062 |
ACM Trans. Archit. Code Optim. | 2 |
| 2025 | CUSPX: Efficient GPU Implementations of Post-Quantum Signature SPHINCS+abstractQuantum computers pose a serious threat to existing cryptographic systems. While Post-Quantum Cryptography (PQC) offers resilience against quantum attacks, its performance limitations often hinder widespread adoption. Among the three National Institute of Standards and Technology (NIST)-selected general-purpose PQC schemes, SPHINCS${}^{+}$is particularly susceptible to these limitations. We introduce CUSPX (CUDASPHINCS${}^{+}$), the first large-scale parallel implementation of SPHINCS${}^{+}$capable of running across 10,000 cores. CUSPX leverages a novel three-level parallelism framework, applying it toalgorithmic parallelism,data parallelism, andhybrid parallelism. Notably, CUSPX introduces parallel Merkle tree construction algorithms for arbitrary parallel scales and several load-balancing solutions, further enhancing performance. By treating tasks parallelism as the top level of parallelism, CUSPX provides a four-level parallel scheme that can run with any number of tasks. Evaluated on a single GeForce RTX 3090 using the SPHINCS${}^{+}$-SHA-256-128s-simple parameter set, CUSPX achieves a single task's signature generation latency of 0.67 ms, demonstrating a 5,105$\times$speedup over a single-thread version and an 18.50$\times$speedup over the previous fastest implementation. Ziheng Wang 0002, Xiaoshe Dong, Heng Chen 0002, Yan Kang 0005, Qiang Wang 0062 |
IEEE Trans. Computers | 2 |
| 2024 | pommDNN: Performance optimal GPU memory management for deep neural network training
Weiduo Chen, Xiaoshe Dong, Xinhang Chen, Song Liu 0007, Qin Xia, Qiang Wang 0062 |
Future Gener. Comput. Syst. | 2 |
| 2024 | Flimm: Foreground traffic aware data migration manager for distributed storage system
Bowen Li 0009, Xiaoshe Dong, Jue Mi, Yufei Wang 0008, Weiduo Chen |
Future Gener. Comput. Syst. | 2 |
| 2024 | An Example of Parallel Merkle Tree Traversal: Post-Quantum Leighton-Micali Signature on the GPUabstractThe hash-based signature (HBS) is the most conservative and time-consuming among many post-quantum cryptography (PQC) algorithms. Two HBSs, LMS and XMSS, are the only PQC algorithms standardised by the National Institute of Standards and Technology (NIST) now. Existing HBSs are designed based on serial Merkle tree traversal, which is not conducive to taking full advantage of the computing power of parallel architectures such as CPUs and GPUs. We propose a parallel Merkle tree traversal (PMTT), which is tested by implementing LMS on the GPU. This is the first work accelerating LMS on the GPU, which performs well even with over 10,000 cores. Considering different scenarios of algorithmic parallelism and data parallelism, we implement corresponding variants for PMTT. The design of PMTT for algorithmic parallelism mainly considers the execution efficiency of a single task, while that for data parallelism starts with the full utilisation of GPU performance. In addition, we are the first to design a CPU-GPU collaborative processing solution for traversal algorithms to reduce the communication overhead between CPU and GPU. For algorithmic parallelism, our implementation is still 4.48× faster than the ideal time of the state-of-the-art traversal algorithm. For data parallelism, when the number of cores increases from 1 to 8,192, the parallel efficiency is 78.39%. In comparison, our LMS implementation outperforms most existing LMS and XMSS implementations. Ziheng Wang 0002, Xiaoshe Dong, Yan Kang 0005, Heng Chen 0002, Qiang Wang 0062 |
ACM Trans. Archit. Code Optim. | 2 |
| 2024 | Parallel implementations of post-quantum leighton-Micali signature on multiple nodes
Yan Kang 0005, Xiaoshe Dong, Ziheng Wang 0002, Heng Chen 0002, Qiang Wang 0062 |
J. Supercomput. | 2 |
| 2023 | Simplified High Level Parallelism Expression on Heterogeneous Systems through Data Partition Pattern DescriptionabstractAbstract With the development of heterogeneous systems, the demand for high-level programming methods that ease heterogeneous programming and produce portable applications has become more urgent. This paper proposes DACL, the data associated computing language. DACL introduces data partition patterns to achieve architecture-independent parallelism expression. Meanwhile, DACL provides simplified language extensions, as well as programming features such as serialization of the computing process, parameterization of data attributes and modularity, thus reducing the difficulty of heterogeneous programming and improving programming productivity. The operational semantics show that DACL enables different levels of parallelism degree calculation and retains data access patterns, reserving optimization potential. To support cross-platform execution, the currently implemented source-to-source compilers employ OpenMP and OpenCL as the backend. We reconstructed multiple benchmarks selected from the Parboil and Rodinia benchmark suits with DACL and conducted a comparison test on CPU, GPU and MIC platforms. The code size of each rebuilt benchmark is roughly equivalent to that of the serial code, which is only 13%–64% of the benchmark OpenCL code. With the support of the compilation system, the reconstructed code can execute on different processors without modification, yielding a competitive or better performance to that of the manually written benchmark code. Shusen Wu, Xiaoshe Dong, Heng Chen 0002, Qiang Wang 0062, Zhengdong Zhu |
Comput. J. | 2 |
| 2023 | Parallel SHA-256 on SW26010 many-core processor for hashing of multiple messages
Ziheng Wang 0002, Xiaoshe Dong, Yan Kang 0005, Heng Chen 0002 |
J. Supercomput. | 2 |
| 2023 | Efficient GPU Implementations of Post-Quantum Signature XMSSabstractThe National Institute of Standards and Technology (NIST) approved XMSS as part of the post-quantum cryptography (PQC) development effort in 2018. XMSS is currently one of only two standardized PQC algorithms, but its performance limits its use. For example, the fastest record for some standardized parameters still takes more than a minute to generate a keypair. In this article, we present the first GPU implementation for XMSS and its variant XMSS$^{\mathsf {MT}}$. The high parallelism of GPUs is especially effective for reducing latency in key generation and improving throughput for signing and verifying. In order to meet various application scenarios, we provide three parallel XMSS schemes:algorithmic parallelism,multi-keypair data parallelism, andsingle-keypair data parallelism. For these schemes, we design custom parallel strategies that use more than 10,000 cores for all parameters provided by NIST. In addition, we analyze the availability of most previous serial optimizations and explore numerous techniques to fully exploit GPU performance. Our evaluations are made with the XMSSMT-SHA2_20/2_256 parameter set on a GeForce RTX 3090. The result shows the key generation latency is 3.20 ms, a speedup of 21,899× compared to the GPU ported version, which is also 54× speedup faster than the fastest work (174 ms). When 16384 tasks are executed, the throughput (task/s) for signing/verifying in the single-key and multi-key cases is 311,424/415,100 and 145,100/419,887, respectively. Compared to the throughput for signing/verifying (1695/4000) of the fastest work, we obtain a speedup of 184×/104× and 86×/105× in single-key and multi-key cases, respectively. Ziheng Wang 0002, Xiaoshe Dong, Heng Chen 0002, Yan Kang 0005 |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2022 | Status, challenges and trends of data-intensive supercomputing
Jia Wei 0002, Pei Ren, Yujia Lei, Yuqi Qu, Qiyu Jiang, Xiaoshe Dong, Weiguo Wu, Qiang Wang 0062, Xingjun Zhang |
CCF Trans. High Perform. Comput. | 8 |
| 2022 | A high-applicability heterogeneous cloud data centers resource management algorithm based on trusted virtual machine migration
Bin Liang 0005, Xiaoshe Dong, Yufei Wang 0008, Xingjun Zhang |
Expert Syst. Appl. | 2 |
| 2022 | LogSC: Model-based one-sided communication performance estimation
Ziheng Wang 0002, Heng Chen 0002, Xiaoshe Dong, Weilin Cai, Xingjun Zhang |
Future Gener. Comput. Syst. | 3 |
| 2022 | C-Lop: Accurate contention-based modeling of MPI concurrent communication
Ziheng Wang 0002, Heng Chen 0002, Weiling Cai, Xiaoshe Dong, Xingjun Zhang |
Parallel Comput. | 4 |
| 2022 | Optimizing Small-Sample Disk Fault Detection Based on LSTM-GAN ModelabstractIn recent years, researches on disk fault detection based on SMART data combined with different machine learning algorithms have been proven to be effective. However, these methods require a large amount of data. In the early stages of the establishment of a data center or the deployment of new storage devices, the amount of reliability data for disks is relatively limited, and the amount of failed disk data is even less, resulting in the unsatisfactory detection performances of machine learning algorithms. To solve the above problems, we propose a novel small sample disk fault detection (SSDFD) 1 optimizing method based on Generative Adversarial Networks (GANs). Combined with the characteristics of hard disk reliability data, the generator of the original GAN is improved based on Long Short-Term Memory (LSTM), making it suitable for the generation of failed disk data. To alleviate the problem of data imbalance and expand the failed disk dataset with reduced amounts of original data, the proposed model is trained through adversarial training, which focuses on the generation of failed disk data. Experimental results on real HDD datasets show that SSDFD can generate enough virtual failed disk data to enable the machine learning algorithm to detect disk faults with increased accuracy under the condition of a few original failed disk data. Furthermore, the model trained with 300 original failed disk data has a significant effect on improving the accuracy of HDD fault detection. The optimal amount of generated virtual data are, 20–30 times that of the original data. Yufei Wang 0008, Xiaoshe Dong, Weiduo Chen, Xingjun Zhang |
ACM Trans. Archit. Code Optim. | 2 |
| 2022 | SunwayURANS: 3D full-annulus URANS simulations of transonic axial compressors on Sunway TaihuLight
Heng Chen 0002, Ziheng Wang 0002, Xiaoshe Dong, Xingjun Zhang |
J. Supercomput. | 5 |
| 2022 | Extending τ-Lop to model MPI blocking primitives on shared memory
Ziheng Wang 0002, Heng Chen 0002, Xiaoshe Dong, Weilin Cai, Yan Kang 0005, Xingjun Zhang |
J. Supercomput. | 3 |
| 2021 | Performance evaluation of convolutional neural network on Tianhe-3 prototype
Weiduo Chen, Xiaoshe Dong, Heng Chen 0002, Qiang Wang 0062, Xingda Yu, Xingjun Zhang |
J. Supercomput. | 2 |
| 2021 | OKCM: improving parallel task scheduling in high-performance computing systems using online learning
Xingjun Zhang, Zeyu Ji, Xiaoshe Dong, Chenglong Hu |
J. Supercomput. | 5 |
| 2021 | A multi-dimensional double descending maximum padding priority algorithm for cloud data centers
Bin Liang 0005, Xiaoshe Dong, Yufei Wang 0008 |
J. Supercomput. | 2 |
| 2020 | Memory-aware resource management algorithm for low-energy cloud data centers
Bin Liang 0005, Xiaoshe Dong, Yufei Wang 0008, Xingjun Zhang |
Future Gener. Comput. Syst. | 2 |
| 2020 | H2Pregel : A partition-based hybrid hierarchical graph computation approach
Xiaoshe Dong, Heng Chen 0002, Xingjun Zhang |
Future Gener. Comput. Syst. | 2 |
| 2020 | Structured mesh-oriented framework design and optimization for a coarse-grained parallel CFD solver based on hybrid MPI/OpenMP programming
Xiaoshe Dong, Nianjun Zou, Weiguo Wu, Xingjun Zhang |
J. Supercomput. | 2 |
| 2020 | A low-power task scheduling algorithm for heterogeneous cloud computing
Bin Liang 0005, Xiaoshe Dong, Yufei Wang 0008, Xingjun Zhang |
J. Supercomput. | 2 |
| 2020 | NADE: nodes performance awareness and accurate distance evaluation for degraded read in heterogeneous distributed erasure code-based storage
Xingjun Zhang, Xiaoshe Dong |
J. Supercomput. | 5 |
| 2020 | Fine-grained scheduling in multi-resource clusters
Mosong Zhou, Xiaoshe Dong, Heng Chen 0002, Xingjun Zhang |
J. Supercomput. | 2 |
| 2019 | NoT: a high-level no-threading parallel programming method for heterogeneous systems
Shusen Wu, Xiaoshe Dong, Xingjun Zhang, Zhengdong Zhu |
J. Supercomput. | 2 |
| 2018 | A Runtime Available Resource Capacity Evaluation Model Based on the Concept of Similar TasksabstractA mismatch between resource supply and demand in cloud computing leads to inefficient utilization of resources or performance degradation. Therefore, this paper establishes a runtime model to evaluate the available capacity of computing resources on the basis of similar tasks. This model takes advantage of a characteristic of cloud workload; that is, similar tasks in cloud computing have a similar execution logic. The model evaluates the available resource capacity according to task similarity, thus avoiding any impact on the resource consumption of existing benchmarks. We apply the model to propose a resource capacity evaluation method called Caipan, which considers numerous factors according to resource type. This method obtains accurate results in a timely manner at little cost. We use the results of Caipan to develop some algorithms that aim to match resource supply and demand, and improve cloud platform performance. We test the Caipan method and the Caipan-based algorithms in both dedicated and real-world cloud environments. The test results show that the Caipan method obtains the available resource capacity both accurately and in a timely manner, and effectively supports the optimization of both algorithms and platforms. Moreover, algorithms based on Caipan reduce the mismatch between resource supply and demand, and significantly improve cloud platform performance. Mosong Zhou, Xiaoshe Dong, Heng Chen 0002, Xingjun Zhang |
Comput. J. | 2 |
| 2018 | IncPregel: an incremental graph parallel computation model
Xiaoshe Dong, Heng Chen 0002, Yinfeng Wang |
Frontiers Comput. Sci. | 2 |
| 2018 | Power and discrete rate adaptation in BER constrained wireless powered communication networksabstractOptimal system throughput is a crucial design issue in wireless powered communication networks. Unlike related literature, this study considers the problem of maximising throughput (MTP) for a practical scenario, in which each user node (UN) selects its own rate from a discrete rate set and controls its transmission power according to its own channel state. The MTP with a bit‐error‐rate constraint is investigated based on two general uplink access methods. The formulation size is exponentially large with respect to the problem input. A novel greedy algorithm based on the column generation method (GA‐CGM) is proposed to solve the problem efficiently. The GA‐CGM decomposes the problem into a master problem and a subproblem. The master problem is solved by the simplex method. The subproblem is solved by a greedy algorithm due to its non‐linearity programming model. The max–min throughput problem (MMTP) is also considered, because of the ‘doubly near‐far’ phenomenon which leads to the unfair throughput among different UNs. Experimental results demonstrate that the proposed solutions are close to the optimal solutions for MTP and MMTP problems. They also show increasing the transmission power of a base station or decreasing the path‐loss exponent improves the throughput performance. Ming Lei 0003, Xingjun Zhang, Bocheng Yu, Xiaoshe Dong |
IET Commun. | 4 |
| 2018 | Stochastic geometry modeling and energy efficiency analysis of millimeter wave cellular networks
Song Cen, Xingjun Zhang, Ming Lei 0003, Scott Fowler, Xiaoshe Dong |
Wirel. Networks | 5 |
| 2017 | Small files storing and computing optimization in Hadoop parallel renderingabstractSummary Hadoop framework has been widely used in animation industries to build a large scale, high performance parallel rendering system. However, Hadoop Distributed File System (HDFS) and the MapReduce programming model are designed to manage large files and suffer performance penalty while rendering and storing small files in a rendering system. Therefore, a method that merges small files based on two intelligent algorithms is proposed to solve the problem. The method uses Particle Swarm Optimization (PSO) to select the optimal merge values for multiple sets of scenes and then uses Support Vector Machine (SVM) to generate a general SVM model which can be used to get the optimal merge value for any scene, by mainly considering the rendering time, memory limitation and other indicators. Then, the method takes advantage of frame‐to‐frame coherence to merge files in the same scene in an interval‐based way with the optimal merge value. Finally, the proposed method is compared with the naive method under three different render scenes. Experimental results show that the proposed method significantly reduces the number of small files and render tasks, and improves the storage efficiency and computing efficiency. Copyright © 2016 John Wiley & Sons, Ltd. Yizhi Zhang, Zhengdong Zhu, Honglin Cui, Xiaoshe Dong, Heng Chen 0002 |
Concurr. Comput. Pract. Exp. | 4 |
| 2017 | Studying shadow page cache to improve isolated drivers' performanceabstractSummary Using virtualization technology, users can reuse various existing operating systems and drivers to customize their application environments. However, faults that exist in the reused drivers threaten the users' customized environments. Chariot has been proposed to solve this problem. It isolates a driver by monitoring its real‐time performance and examining the correctness of its write operations. To capture the write operations, we found that it needs to keep the write permissions of shadow pages as read‐only, which may cause many page faults in the virtual machine monitor and has adversely affect the driver's performance. To address this, we propose a new algorithm that uses the shadow page cache to remit the page faults, based on the principle of delaying the closing of the write permissions of frequently accessed shadow pages. Experimental results show that using these shadow page caches, we can greatly the performance of isolated drivers without significantly impacting Chariot's isolation efficiency. Hao Zheng 0004, Zhengdong Zhu, Xiaoshe Dong, Baoke Chen, Chengzhe Liu |
Concurr. Comput. Pract. Exp. | 3 |
| 2016 | The Optimization of Memory Access Congestion for MapReduce Applications on Manycore SystemsabstractThe prevalence of heterogeneous manycore processors has been shown as a promising alternative to web-based data parallel MapReduce applications. The differences of manycore from multicore raise new challenges to designing and implementing efficient and scalable MapReduce applications, such as memory access congestion. This paper argues that it is more efficient using multiple small groups of cores than using a single group with all cores to address these challenges. We propose a group-based MapReduce implementation, called Grouped-MapReduce (GMR). It uses grouping to align tasks so that the tasks in each group are congestion-free. Further, it uses multiplexing to control the implementation of groups so that multiple groups can efficiently share the limited memory bandwidth without causing serious memory access congestion. We have implemented a prototype of GMR based on Phoenix++, an already highly optimized MapReduce runtime. Experiments on six benchmarks show that GMR implements and scales well on manycore systems and obtains an impressive improvement over Phoenix++ from 1.04x to 1.77x without artificially tuning the existing application code. Endong Wang, Xiaoshe Dong, Zhengdong Zhu |
Comput. J. | 3 |
| 2016 | Improving the Reliability of the Operating System Inside a VMabstractVirtualization technology can provide reusability and strong isolation between different virtual machines (VMs). However, there is no effective isolation mechanism inside a VM to solve an operating system's reliability problems, including driver faults. This paper describes Chariot, an architecture that provides effective and transparent driver isolation inside the VM, achieves fine-grained driver isolation and retains the reusability advantage of virtualization technology. First, Chariot transparently monitors an isolated driver with monitoring wrappers, and establishes an access control table (ACT) in a timely manner that records the driver write permissions. Secondly, Chariot protects the shadow page table of the VM (where the driver resides) in due time to capture its write operations. Next, the ACT examines the correctness of the write operations. Finally, if an illegal write operation is detected, Chariot recovers the faulty driver and prevents the spread of driver faults in the VM. The experimental results show that Chariot effectively isolates more than 90% of injected faults (with performance losses of |$<$|20% in most benchmarks) and effectively improves the reliability of the VM. In addition, Chariot can be easily extended to isolate new drivers and ported to other versions of OSs in the virtualization environment. Hao Zheng 0004, Xiaoshe Dong, Zhengdong Zhu, Baoke Chen, Xiuxiu Bai, Xingjun Zhang, Endong Wang |
Comput. J. | 2 |
| 2016 | TextGen: a realistic text data content generation method for modern storage system benchmarksabstractModern storage systems incorporate data compressors to improve their performance and capacity. As a result, data content can significantly influence the result of a storage system benchmark. Because real-world proprietary datasets are too large to be copied onto a test storage system, and most data cannot be shared due to privacy issues, a benchmark needs to generate data synthetically. To ensure that the result is accurate, it is necessary to generate data content based on the characterization of real-world data properties that influence the storage system performance during the execution of a benchmark. The existing approach, called SDGen, cannot guarantee that the benchmark result is accurate in storage systems that have built-in word-based compressors. The reason is that SDGen characterizes the properties that influence compression performance only at the byte level, and no properties are characterized at the word level. To address this problem, we present TextGen, a realistic text data content generation method for modern storage system benchmarks. TextGen builds the word corpus by segmenting real-world text datasets, and creates a word-frequency distribution by counting each word in the corpus. To improve data generation performance, the word-frequency distribution is fitted to a lognormal distribution by maximum likelihood estimation. The Monte Carlo approach is used to generate synthetic data. The running time of TextGen generation depends only on the expected data size, which means that the time complexity of TextGen is O ( n ). To evaluate TextGen, four real-world datasets were used to perform an experiment. The experimental results show that, compared with SDGen, the compression performance and compression ratio of the datasets generated by TextGen deviate less from real-world datasets when end-tagged dense code, a representative of word-based compressors, is evaluated. Xiaoshe Dong, Xingjun Zhang, Yinfeng Wang, Tao Ju 0002, Guofu Feng |
Frontiers Inf. Technol. Electron. Eng. | 2 |
| 2015 | Thread Count Prediction Model: Dynamically Adjusting Threads for Heterogeneous Many-Core SystemsabstractDetermining an appropriate thread count for a multithread application running on a heterogeneous many-core system is crucial for improving computing performance and reducing energy consumption. This paper investigates the interrelation between thread count and computing performance of applications, and designs a prediction model of the optimum thread count on the basis of Amdahl's law combined with regression analysis theory to improve computing performance and reduce energy consumption. The prediction model can estimate the optimum tread count relying on the program running behaviors and the architecture characteristics of heterogeneous many-core system. Using the estimated optimum thread count, the number of the active hardware threads and processing cores on the many-core processor is dynamically adjusted in the process of thread mapping to improve the energy efficiency of entire heterogeneous many-core system. The experimental results show that, using this paper proposed thread count prediction model, on an average, the computing performance is improved by 48.6%, energy consumption is reduced by 59%, and additional overhead introduced is 2.03% compared with that of the traditional thread mapping for the PARSEC benchmark programs run on an Intel MIC heterogeneous many-core system. Tao Ju 0002, Weiguo Wu, Heng Chen 0002, Zhengdong Zhu, Xiaoshe Dong |
ICPADS | 5 |
| 2015 | The Performance Survey of in Memory DatabaseabstractTo satisfy the ever-increasing performance demand of Big Data and critical applications the data management needs to offer the flexible schema, high availability, light weight replica, high volume and high scalability features so as to facilitate the transaction. The in memory database (IMDB) eliminates the I/O bottleneck by storing data in main memory. We give a deeper analysis of current main-stream IMDB systems performance which focuses on the data structure, architecture, volume, concurrency, availability and scalability. The V3 performance model is proposed to evaluate the Velocity, Volume and Varity of the 19 IMDB systems, in order to highlight the candidates with realtime transaction and high volume processing capacity coordinately. Test results clearly demonstrate that NewSQL is better at dealing with high-frequency trading models. To fully utilize the advantages of the multi-core and many-core processors capability improvements, a three-level optimization design strategy, which includes the memory-access level, the kernel-speedup level and the data-partition level also be proposed using the hardware parallelism for achieving task-level and data-level parallelism of IMDB programs, guarantees the IMDB could accelerate the real-time transaction in an efficient way. We believe that IMDB should become a compulsive option for enterprise users. Yinfeng Wang, Guiquan Zhong, Lin Kun, Huang Kai, Fuliang Guo, Chengzhe Liu, Xiaoshe Dong |
ICPADS | 8 |
| 2015 | Research on Algorithms to Capture Drivers' Write OperationsabstractFull virtualization technology is highly reusable. Using this property, various types and versions of existing operating systems and drivers can be reused in a virtual machine to customize users’ application environments. However, these environments are threatened by drivers’ write operation faults, which are caused by bugs in reused drivers. Chariot is a reliability architecture that has been developed to solve this problem. This architecture captures a driver's write operations by maintaining the write permissions of shadow pages as read-only to examine their correctness. Nevertheless, this capture method produces many page faults in the virtual machine monitor and has an adverse impact on the performance of isolated drivers. To reduce performance losses, this paper examines two algorithms that cache recently used shadow pages using different structures to avoid frequent page faults. The experimental results show that the performance of isolated drivers can be greatly improved using these shadow page caches without significantly impacting the isolation efficiency of Chariot. Hao Zheng 0004, Xiaoshe Dong, Zhengdong Zhu, Baoke Chen, Yizhi Zhang, Xingjun Zhang |
Comput. J. | 2 |
| 2015 | Edge Propagation KD-Trees: Computing Approximate Nearest Neighbor FieldsabstractPropagation-assisted kd-tree is a state-of-the-art method for computing approximate nearest neighbor (ANN) fields. In this method, each query patch needs descending search in the kd-tree and propagation search in the nearby patches. We observed that the query patches in the edge region need descending search, while other query patches only need propagation search. This can be an opportunity to save plenty of search time. In this letter, we propose edge propagation kd-trees to quickly compute ANNs. Our method can distinguish between edge patches and propagation patches in choosing the proper search. Experiments on public data set VidPairs show that our search method is 2-3 times faster than the propagation-assisted kd-tree search method at nearly the same accuracy. Xiuxiu Bai, Xiaoshe Dong, Yuanqi Su |
IEEE Signal Process. Lett. | 2 |
| 2015 | A scalability prediction approach for multi-threaded applications on manycore processors
Xiuxiu Bai, Endong Wang, Xiaoshe Dong, Xingjun Zhang |
J. Supercomput. | 3 |
| 2014 | Improving MapReduce Performance in a Heterogeneous Cloud: A Measurement StudyabstractHybrid clouds, geo-distributed cloud and continuous upgrades of computing, storage and networking resources in the cloud have driven datacenters evolving towards heterogeneous clusters. Unfortunately, most of MapReduce implementations are designed for homogeneous computing environments and perform poorly in heterogeneous clusters. Although a fair of research efforts have dedicated to improve MapReduce performance, there still lacks of in-depth understanding of the key factors that affect the performance of MapReduce jobs in heterogeneous clusters. In this paper, we present an extensive experimental study on two categories of factors: system configuration and task scheduling. Our measurement study shows that an in-depth understanding of these factors is critical for improving MapReduce performance in a heterogeneous environment. We conclude with five key findings: (1) Early shuffle, though effective for reducing the latency of MapReduce jobs, can impact the performance of map tasks and reduce tasks differently when running on different types of nodes. (2) Two phases in map tasks have different sensitive to input block size and the ratio of sort phase with different block size is different for different type of nodes. (3) Scheduling map or reduce tasks dynamically with node capacity and workload awareness can further enhance the job performance and improve resource consumption efficiency. (4) Although random scheduling of reduce tasks works well in homogeneous clusters, it can significantly degrade the performance in heterogeneous clusters when shuffled data size is large. (5) Phase-aware progress rate estimation and speculation strategy can provide substantial performance gain over the state of art speculation scheduler. Ling Liu 0001, Qi Zhang 0009, Xiaoshe Dong |
IEEE CLOUD | 4 |
| 2014 | Thread Mapping and Parallel Optimization for MIC Heterogeneous Parallel Systems
Tao Ju 0002, Zhengdong Zhu, Yinfeng Wang, Xiaoshe Dong |
ICA3PP (2) | 5 |
| 2014 | An Availability Approached Task Scheduling Algorithm in Heterogeneous Fault-Tolerant SystemabstractIn heterogeneous fault-tolerant system, especially high performance computer system, the issue of providing system with high availability assurance for real-time applications which have availability requirements has been widespread concerned. While, few research concentrates on combining real-time application availability requirement with scheduling algorithm. In this paper, an availability approached task scheduling algorithm is proposed. On the basis of heterogeneous fault-tolerant system scheduler and scheduling algorithm designation, we can improve the system availability without increasing additional hardware costs, and shorten task average response time, in addition, schedule task with high efficiency and reliability. Experiment results show that, such availability approached task scheduling algorithm has a system performance advantage over the traditional system task scheduling algorithms, it achieves the goal of balancing availability and task response time in heterogeneous fault-tolerant system, thus improve the system availability. Xiaoshe Dong, Xingjun Zhang, Yinfeng Wang |
NAS | 2 |
| 2014 | A run-time optimization approach for reducing data movements using locality-aware searching
Endong Wang, Xingjun Zhang, Tao Ju 0002, Xiaoshe Dong |
J. Supercomput. | 6 |
| 2013 | A Novel Method for Computing Encoding Delay and Bandwidth on Network Coding NodeabstractComputing the encoding delay and bandwidth on the nodes is important for the deployment and implementation of network coding. However, the existing researches merely use the methods of experimental measurement or rough estimation to investigate it. In order to accurately compute the delay and bandwidth on the nodes with network coding implemented, this paper proposes a novel method to quantify encoding delay and encoding bandwidth. Firstly, network coding is applied to design a new network. And then, the encoding procedure is systematically investigated on the source node, the intermediate node, and also the decoding procedure is systematically investigated at the destination node. And the method, which can be used to compute the node encoding delay and encoding bandwidth, is presented, Thus the method of computing total delay of network with network coding is derived. Finally, the experiments validate the correctness of this method. Yuxing Wu, Xingjun Zhang, Song Cen, Xiaoshe Dong |
NAS | 5 |
| 2013 | Chariot: A High Compatible Architecture to Improve Virtual Machine ReliabilityabstractCurrently, the virtualization technologies can integrate multiple operating systems into a high-performance server to maximize the utilization of the server's resources. This server can serves more users. However, the driver faults in virtual machine still seriously affect the reliability of the virtual machine, and even affect the reliability of the entire server. This paper presents Chariot, a high compatible architecture to improve virtual machine reliability. If the driver is loaded by the Chariot's isolation loading mechanism, its memory usage will be timely monitored by Chariot, and its access control table will be established. Through setting the corresponding shadow page table of the whole kernel space of the virtual machine, Chariot captures the write operations of the isolated driver. Combing the access control table, Chariot can determine the correctness of these writing operations. Chariot has an effective errors isolation capability, and is easy to develop. Also Chariot has an excellent compatibility and needs not to modify the drivers and the operation system in the virtual machine. Experimental results show that Chariot can effectively isolate the driver faults, and improve the reliability of operation system in the virtual machine environments. Hao Zheng 0004, Xiaoshe Dong, Endong Wang, Baoke Chen, Weifeng Gong, Xingjun Zhang |
NAS | 2 |
| 2013 | An Undirected Graph Traversal Based Grouping Prediction Method for Data De-duplicationabstractThe data capacity of the de-duplication system, which is limited by the memory, is difficult to carry out a large-scale expansion. To solve this problem, the paper proposes a hash table grouping prediction method based on undirected graph traversal. This method exploits the indexing table replacement, which is similar to the virtual memory cache replacement, to expand the storage capacity of the data de-duplication system without increasing the system memory. The hit rate of the grouping prediction and system performance are improved by grouping index entries based on undirected graph traversal. Experimental results show that, based on the cache prefetching and the hash table grouping, the memory consuming takes up 10% of the index table size while the capacity equally rise to 10 times of original. The method can make the index table cache hit rate increased to 87.6%, comparing 47% without group in dataset 1 of our experiment, make the performance acceptable. Xingjun Zhang, Guofeng Zhu, Yueguang Zhu, Xiaoshe Dong |
SNPD | 5 |
| 2013 | Improving Virtual Machine Reliability with Driver Fault IsolationabstractWith the development of virtualization technologies, the server resources are maximized by integrating multiple operating systems into a high-performance server. So the server is possible to provide services to more users simultaneously. However, the driver fault still impacts the reliability of the operating system in the virtual machine, and impacts the continuity and stability of services. This paper proposes a architecture to improve the reliability of the virtual machine environments. By monitoring the driver's memory usage, the architecture creates the authorization table. Through setting the corresponding shadow page table in the virtual machine manager of the whole kernel space of the virtual machine, the architecture captures the write operation of the virtual machine. Combing with the authorization table, the correctness of the writing operations can be determined. Our architecture needn't to modify the drivers and is easy to develop. Experimental results show that the architecture can effectively isolate the driver faults, and improve the reliability of the virtual machine environments. Hao Zheng 0004, Xiaoshe Dong, Endong Wang, Baoke Chen, Xingjun Zhang |
SNPD | 2 |
| 2013 | A dual process redundancy approach to transient fault tolerance for ccNUMA architecture
Xingjun Zhang, Endong Wang, Feilong Tang 0001, Meishun Yang, Hengyi Wei, Xiaoshe Dong |
Neurocomputing | 6 |
| 2013 | Performance Analysis of Network I/O Workloads in Virtualized Data CentersabstractServer consolidation and application consolidation through virtualization are key performance optimizations in cloud-based service delivery industry. In this paper, we argue that it is important for both cloud consumers and cloud providers to understand the various factors that may have significant impact on the performance of applications running in a virtualized cloud. This paper presents an extensive performance study of network I/O workloads in a virtualized cloud environment. We first show that current implementation of virtual machine monitor (VMM) does not provide sufficient performance isolation to guarantee the effectiveness of resource sharing across multiple virtual machine instances (VMs) running on a single physical host machine, especially when applications running on neighboring VMs are competing for computing and communication resources. Then we study a set of representative workloads in cloud-based data centers, which compete for either CPU or network I/O resources, and present the detailed analysis on different factors that can impact the throughput performance and resource sharing effectiveness. For example, we analyze the cost and the benefit of running idle VM instances on a physical host where some applications are hosted concurrently. We also present an in-depth discussion on the performance impact of colocating applications that compete for either CPU or network I/O resources. Finally, we analyze the impact of different CPU resource scheduling strategies and different workload rates on the performance of applications running on different VMs hosted by the same physical machine. Yiduo Mei, Ling Liu 0001, Xing Pu, Sankaran Sivathanu, Xiaoshe Dong |
IEEE Trans. Serv. Comput. | 5 |
| 2012 | Powered Grid Scheduling by Ant Algorithm
Xiaoshe Dong |
ICIC (1) | 2 |
| 2012 | A performance and energy optimization mechanism for cooperation-oriented multiple server clusters
Zhenghua Xue, Xiaoshe Dong, Leijun Hu |
Future Gener. Comput. Syst. | 2 |
| 2010 | mPlogP: A Parallel Computation Model for Heterogeneous Multi-core ComputerabstractDue to the heterogeneity and the multigrain parallelism of the heterogeneous multi-core computer, communication and memory access show hierarchical characteristics ignored by other models. In this paper, a new model named mPlogP, is presented on the basis of the PlogP model, in which communication and memory access is abstracted by considering these new characteristics of the heterogeneous multi-core computer. It uses memory access to model the behavior of computation, estimates the execution time of every part of applications and guides the optimization of effective parallel programs. Finally this proposed model is validated by experiments that it can precisely evaluate the execution of parallel applications under the heterogeneous multi-core computer. Xingjun Zhang, Jinghua Feng, Xiaoshe Dong |
CCGRID | 4 |
| 2008 | A Distributed Trust Management Based on Authorizing Negotiation in Open and Dynamic EnvironmentsabstractTrust has been recognized as an important factor for information security in open and dynamic environments, such as Internet applications, p2p systems etc. On the basis of analyzing existing trust management systems, this paper proposes a Distributed Trust Management based on Authorizing Negotiation (DTMAN). DTMAN presents a number of innovative features. First, it can authorize strangers by using authorizing negotiation, so it is very suitable for open and dynamic environments. Second, a high efficient algorithm for compliance checking is developed to support DTMAN, whose time complexity and space complexity are both O(n) (where n is the cardinality of the set of authorization credentials). The experimental result shows that the algorithm of DTMAN is more efficient than others. Shangyuan Guan, Xiaoshe Dong, Weiguo Wu, Yiduo Mei, Guofu Feng |
AINA | 2 |
| 2008 | FORT: A decentralized automated trust negotiation framework for gridsabstractTrust has been recognized as an important factor for grid security. This paper proposes a decentralized automated trust negotiation framework, FORT, to establish trust relationship between service providers and service requesters in grids. FORT presents many innovative features. First, FORT is decentralized, so it scales well and is well-suited for large-scale grids. Second, FORT refines its policy language with attribute constraint, so it can provide the support for effective protection of sensitive information of the two negotiation parties and flexible limitation of delegation range. Last, we employ multithreaded technology to speed up negotiation. This paper depicts the implementation of FORT and designs experiments to evaluate its performance. Experimental results show that FORT can effectively protect sensitive services at the cost of little performance of systems and is scalable. Shangyuan Guan, Xiaoshe Dong, Yiduo Mei, Xingjun Zhang |
CSCWD | 2 |
| 2008 | EntityTrust: Feedback credibility-based global reputation mechanism in cooperative computing systemabstractTrust and reputation are important decision-making factors in cooperative computing systems. It is a fundamental but challenging task for reputation systems to estimate entity's reputation accurately and efficiently in distributed environment. We propose EntityTrust, a global reputation mechanism in cooperative computing systems. EntityTrust introduces a new direct feedback metrics to reflect the dynamic feature of trust. Besides that, an entity's feedback credibility can affect its reputation in a direct way. Experimental results show that EntityTrust can improve the accuracy and efficiency of evaluation for global reputation. Moreover, it can effectively combat malicious behaviors presented in the cooperative communities. Yiduo Mei, Xiaoshe Dong, Zhenhua Tian, Shangyuan Guan, Heng Chen 0002 |
CSCWD | 2 |
| 2008 | A Profile-based Memory Access Optimizing Technology on CBE ArchitectureabstractIn the paper, we investigate the memory access technology on cell broadband engine architecture (CBEA), and develop a profiling infrastructure for memory management on the architecture. By registering the dynamic memory allocation and providing details of trace of memory access, the infrastructure provides the data partition information automatically which alleviates the burdens of programmer and provides a safety guarantee for aggressive data prefetch for computing task. On the other hand, the profile information is useful for analyzing the patterns of memory access and helpful for further performance optimization. Experimental results show that applications implemented based on our SDK library not only support aggressive memory access method without the requirement of external data partition information, but also could be optimized aggressively under the guideline of the profile information provided by the proposed SDK library. Guofu Feng, Xiaoshe Dong, Jinghua Feng, Xuhao Wang |
HPCC | 2 |
| 2008 | CASTTE: A Trust Management for Securing the GridabstractIt has been a fundamental but challenging problem to gain assurance of the trustworthiness of service providers or requesters and ensure their interests. We present the formal definition of trust management, and then propose a trust management, CASTTE, to secure sensitive services and requesters in grids. CASTTE verifies access trust by using trust negotiation so as to protect sensitive services, and protects sensitive information of the two negotiators effectively by using a negotiation strategy based on protection tree. Furthermore, we utilize trust force to specify provision trust and apply trust force to service selection. This paper implements CASTTE and designs experiments to evaluate its performance. The experimental results show that it can not only protect sensitive services at the cost of little performance of systems, but also identify good services from bad ones effectively. Shangyuan Guan, Xiaoshe Dong, Yiduo Mei, Zhao Wang 0001, Zhengdong Zhu |
HPCC | 2 |
| 2008 | An Authorization Mechanism Based on Privilege Negotiation Policy in GridabstractWith the dynamic change of users and resources in different secure domains of Grid, the overall consistency of privileges defined would be broken. This would compromise Grid system and waste system overhead on dealing with the increasing grid jobs with invalid privileges. To address the problem, this paper proposes an authorization mechanism based on privilege negotiation policy. This mechanism can detect timely the change of privileges, negotiate automatically and resume quickly the overall consistency of privileges between different secure domains. The test result of the mechanism implementation shows that it shortens greatly the period of resuming the overall consistency of privileges between different secure domains when the consistency was broken. This reduces the number of grid jobs with invalid privileges. Thereby, it avoids wasting more system overhead of dealing with the increasing grid jobs with invalid privileges and improves system performance. Runlian Zhang, Xiaonian Wu, Xiaoshe Dong, Shangyuan Guan |
HPCC | 3 |
| 2008 | A mechanism of automated monitoring deployment in grid environmentabstractConstructing monitoring systems beforehand cannot satisfy the requirement of dynamically organizing resources. This paper describes a script-based approach to automate the deployment of grid monitoring service components. An automated deployment monitoring mechanism uses the information of dynamical Yinfeng Wang, Xiaoshe Dong |
QSHINE | 3 |
| 2007 | SDRD: A Novel Approach to Resource Discovery in Grid Environments
Yiduo Mei, Xiaoshe Dong, Weiguo Wu, Shangyuan Guan, Junyang Li 0004 |
APPT | 2 |
| 2007 | Rapid and Automated Deployment of Monitoring Services in Grid EnvironmentsabstractThe monitoring service is a crucial component in the service oriented grid infrastructure. Objectives and requirements for the grid monitoring system have been summarized in this paper. To cope with the dynamic and large-scale nature of the grid, a scalable distributed monitoring system is proposed, which can support easy and rapid deployment of monitoring services in the grid environments. The BitTorrent protocol is adopted to facilitate the distribution and automated deployment of the monitoring system. The deployment of monitoring service for the newly joined node and the update of the deployed components can be initiated by the nodes within the system to reduce the administrator's participation. The maintenance cost and human errors might happen in the deployment process can be reduced. With the help of the peer-to-peer networks, our proposed system can automatically adapt to failures in network connections or nodes. Service capacity of our proposed system is given out. Simulation results show that our proposed system supports efficient and rapid deployment to a large scale grid and provides a robust platform to efficiently monitor the grid resources. Yiduo Mei, Xiaoshe Dong, Junyang Li 0004, Xu Jing, Zhenghua Xue |
APSCC | 2 |
| 2007 | An Energy-Efficient Management Mechanism for Large-Scale Server ClustersabstractWith the increase of the computing demand, high performance server clusters are becoming one of the most important computing infrastructures. The current clusters are designed to meet peak load with all the computing resources keeping running. However, this static reservation with full computing resources can not adapt to the time-varying computing requirement, and may incur low resource utilization and needless power consumption when the cluster system is underloaded. In this paper, we present an extensible architecture of cluster management system. This architecture promises a good extensibility by integrating job scheduler and resource manager in loose couple. Concentrating on the power saving of large-scale clusters, we describe the power model of servers, and based on the presented management system architecture, we propose a novel resource management way, adaptive pool based resource management (APRM) method, for adaptive provision of computing resources in accordance with the time- varying workload demand. APRM enables a cost- effective operating by providing dynamic computing capacity with automatic resource control. We validated APRM on the energy efficiency and quality of service (QoS) by simulation measurement, and the results showed that APRM yields significant power saving with little impact on QoS. Zhenghua Xue, Xiaoshe Dong, Shengqun Fan, Yiduo Mei |
APSCC | 2 |
| 2007 | Application of uncertainty reasoning theory to satellite fault detection and diagnosisabstractReasoning theories are divided into certainty reasoning theories and uncertainty reasoning theories. Now, only certainty reasoning theories are used to detect and diagnose satellite faults. However, in practice, it is difficult to detect and diagnose some faults of the satellite automatically only by use of certainty reasoning theories. The reason is that detection and diagnosis of these faults require a rational reasoning and a fault-tolerant capability. Fortunately, uncertainty reasoning theories can meet these requirements. It is attracting attention of many experts in the space field all over the world that uncertainty reasoning theories are applied to detect and diagnose satellite faults. Uncertainty reasoning theories include several kinds of theories, such as Inclusion Degree Theory, Rough Set Theory, Evidence Reasoning Theory, Probabilistic Reasoning Theory, Fuzzy Reasoning Theory, and so on. Inclusion Degree Theory, Rough Set Theory and Evidence Reasoning Theory are three advanced ones. Based on these three theories respectively, this paper introduces three new methods to detect and diagnose satellite faults. It is shown that the methods, suitable for detecting and diagnosing satellite faults, especially uncertainty faults, can remedy the defects of the current methods. Tianshe Yang, Zheng Xi, Xiaoshe Dong, YongXuan Huang |
SMC | 4 |
| 2007 | Research on methods to simulate spacecraft systemsabstractSystem simulation plays very important role in spacecraft engineering and is an essential part of spacecraft technology. Three simulation methods used to test and verify performances and functions of spacecraft and five simulation methods to establish a universal simulation system for spacecraft ground TT&C system are proposed. A simulation system for spacecraft, especially for control system of the spacecraft, is analyzed. Tianshe Yang, Zheng Xi, Xiaoshe Dong, YongXuan Huang |
SMC | 4 |
| 2006 | GHIDS: Defending Computational Grids against Misusing of Shared ResourcesabstractDetecting intrusions at host level is vital to protecting shared resources in grid, but traditional host-based intrusion detecting system (HIDS) is not suitable for grid environment. Grid-specific attacks are different from traditional ones, and traditional HIDS can not recognize a grid user and always with high performance overhead. This paper proposes a grid-specific host-based intrusion detection system (GHIDS) which employs bottleneck verification approach to detect intrusions with low false alarm rate and high detection rate. Working within operating system kernel and performing bottleneck verification by integer comparison, GHIDS achieves high efficiency and accuracy. Security reports generated by GHIDS are indexed not only by local user ID, but also by grid user ID. That is more useful for analyzing grid user behaviors globally by both host administrators and high level grid-based IDS Guofu Feng, Xiaoshe Dong, Weizhe Liu, Junyang Li 0004 |
APSCC | 2 |
| 2006 | AOCMS: An Adaptive and Scalable Monitoring System for Large-Scale ClustersabstractIn this paper, we present the design and implementation of AOCMS, an adaptive, scalable and efficient monitoring system for a large-scale cluster. We describe an adaptive architecture of AOCMS in detail, and focus on the discussion about some techniques as to enhancing the adaptation, scalability and efficiency of AOCMS. These techniques include: a solution to monitor a heterogeneous cluster; a universal applet-servlet communicating controller responsible for communication between the clients and the Web server; adaptive pools providing threads or connections to the database for the monitoring tasks on demand; and an AOP-based alarm decoupling the alarming logic from the monitoring logic. Moreover, we measured the performance of AOCMS. The results show that AOCMS runs with low overheads and responds to clients quickly Zhenghua Xue, Xiaoshe Dong, Weiguo Wu |
APSCC | 2 |
| 2006 | Key Techniques of Software Sharing for on Demand Service-Oriented Computing
Xiaoshe Dong, Yinfeng Wang, Fang Zheng 0002, Zhongsheng Qin, Guofu Feng |
GPC | 1 |
| 2003 | Design and Implementation of Heartbeat in Multi-Machine EnvironmentabstractHeartbeat mechanism is widely used in the high availability field of monitor network service and server nodes. In this paper, we design a distributed heartbeat mechanism, which consists of one Master node and multiple Standby nodes, all running the same process of heartbeat but in different states. We experimentally demonstrate that the mechanism improves the scalability over that of the dual computer nodes and supports nice fail-over. Moreover, the process of the heartbeat only needs to maintain few states, which dramatically reduces the overheads of the system and simplifies the management of the nodes. When there are only two nodes in the network, the mechanism will degenerate to dual computer heartbeat. In current version, the mechanism only supports the switch-based architecture; it is easy to make it supporting other architectures such as Ring, etc. Zonghao Hou, Yongxiang Huang, Shouqi Zheng, Xiaoshe Dong, Bingkang Wang |
AINA | 4 |