EDBT 2026 Demo / reviewers in the wild / expert
Ruibo Wang
dblp:32/7079
· DBLP profile ↗
52ranked-venue papers
9as first author
39since 2021 · last 2027
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 22 · 3 first-author · 18 since 2021Artificial intelligence and machine learning · 11 · 2 first-author · 6 since 2021Computer networks · 8 · 3 first-author · 8 since 2021Databases, data management, data science and information retrieval · 7 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | Revisiting free space fragmentation: A new garbage collection scheme for F2FSabstractFlash Friendly File System (F2FS) is a log-structured file system (LFS) optimized for Flash memory characteristics and is widely deployed on mobile devices, embedded systems, and some Linux platforms that use NAND Flash. File fragmentation and free space fragmentation both affect the performance of F2FS. In this work, we investigate the performance impact of fragmentation through energy consumption characterization. Our measurements indicate that energy consumption increases with the number of file and free space fragments, based on experiments evaluating F2FS while serving I/O requests across diverse workload scenarios. While considerable efforts have been devoted to mitigating file fragmentation, comparatively less attention has been paid to understanding and optimizing free space fragmentation, which predominantly arises from the distribution of invalid blocks. We observe that reclaiming invalid blocks via background garbage collection (GC) incurs over 100 mJ of energy per invocation, yet yields only marginal reductions in free space fragmentation. This motivates us to improve GC effectiveness in reducing free space fragmentation. We propose the free space frag mentation-aware (FragGC) and file system perf ormance-aware GC (PerfGC) scheme. We seek to both reduce GCs and enhance the efficiency of each GC operation. We reassess the definition of a free space fragment through empirical analysis and introduce the free space fragmentation factor as a lightweight metric to quantify the degree of free space fragmentation at the segment level. FragGC optimizes victim segment selection and valid block migration based on this metric. PerfGC adjusts GC frequency according to the impact of free space fragmentation on file system performance. Experimental results on a real platform demonstrate that FragGC and PerfGC reduce GC count compared to traditional F2FS and its latest GC optimization, ATGC. FragGC reduces the time to replay traces by 24.4% to 40.9% for large-scale applications. Ruibo Wang, Yong Dong, Weizhao Lin |
Future Gener. Comput. Syst. | 2 |
| 2026 | DeloopSGNN: Revisiting Spectral GNNs Through the Lens of Spatial AggregationabstractGraph Neural Networks (GNNs) have been studied from two primary perspectives: spectral, which employs global graph signal filtering and is theoretically more expressive, and spatial, which builds on local neighborhood aggregation and generalizes well across diverse graph structures. While spectral GNNs are expected to perform better in theory, they often underperform in practice compared to spatial models. To better understand this gap, we introduce a novel theoretical framework for converting spectral GNNs into the spatial domain, allowing for more intuitive analysis. This transformation reveals that signal looping and repeated high-order aggregation are major causes of over-smoothing in spectral GNNs. By addressing these issues in the spatial domain and converting the model back to the spectral domain, we propose DeloopSGNN, a spectral GNN with improved expressive capacity. Experiments on benchmark datasets show that DeloopSGNN achieves consistently strong performance in terms of accuracy and adversarial robustness, demonstrating that spectral GNNs can benefit significantly from careful architectural design grounded in our proposed framework. Duanyu Li, Huijun Wu 0001, Kai Lu 0001, Zhenwei Wu, Yong Dong, Ruibo Wang |
AAAI | 8 |
| 2026 | Chain-of-Thought Prompting for Frame Identification with Large Language Models
Xuefei Cao, Yan Xue, Ruibo Wang |
KSEM (6) | 4 |
| 2026 | PallasGNN: Curriculum-Based Pattern Mining for Robust GNNs
Kaiwen Xia, Huijun Wu 0001, Ruibo Wang, Zhenwei Wu, Yong Dong |
PAKDD (3) | 4 |
| 2026 | Performance Limits and Probabilistic Shaping in PPM Optical Links With Multi-Pixel SPDsabstractThis paper investigates the performance limits of pulse position modulation (PPM) optical communication systems using multi-pixel single-photon detectors (SPDs) under the constraint of detector dead time. A detailed analytical framework is developed by modeling the temporal detection behavior of SPDs as a Markov process, capturing the effects of dead time across consecutive PPM symbols. Closed-form expressions are derived for slot-wise detection probabilities, symbol transition matrices, and symbol error rate (SER) under maximum-likelihood detection. To enhance spectral and energy efficiency, a probabilistic shaping scheme is introduced and optimized using a truncated Blahut-Arimoto algorithm, allowing the transmitter to adapt its symbol distribution to the nonlinear detection characteristics of the SPD array. The proposed model is validated through extensive Monte Carlo simulations, demonstrating excellent agreement with theoretical predictions. Results show that probabilistic shaping significantly improves communication sensitivity, reducing the required number of photons per bit by up to 25% in photon-starved or high-noise regimes. Ziyuan Shi, Shunyuan Shang, Ruibo Wang, Mohamed-Slim Alouini |
IEEE Trans. Commun. | 3 |
| 2026 | mtGEMM: An Efficient GEMM Library for Modern Multi-Core DSPsabstractThe General Matrix Multiplication (GEMM) is a crucial subprogram in high-performance computing (HPC). With the increasing importance of power and energy consumption, modern Digital Signal Processors (DSPs) are being integrated into general-purpose HPC systems. However, due to architecture disparities, traditional optimizations for CPUs and GPUs are not easily applicable to modern DSPs. This paper shares our experience of optimizing the GEMM operation using a CPU-DSP platform as a case study. Our work employs a set of strategies to improve the performance and scalability of GEMM. These strategies focus on developing micro-kernels based on heterogeneous on-chip memory, addressing the memory access bottleneck in multi-core parallelism, and facilitating efficient transpose-GEMM. These approaches, collectively referred to as an efficient and practical library (a.k.a.mtGEMM), maximize computational capabilities and bandwidth utilization of multi-core DSPs, while achieving high performance for variously-shaped GEMMs. Our experimental results demonstrate thatmtGEMMcan attain between 92% and 96% of the hardware peak, with the multi-core scalability being almost linear. Jianbin Fang, Kainan Yu, Peng Zhang 0061, Dezun Dong, Xinxin Qi, Xingyu Hou, Ruibo Wang, Kai Lu 0001 |
IEEE Trans. Parallel Distributed Syst. | 7 |
| 2026 | PHIDE: A Parallel Hybrid Direct-Iterative Eigensolver for Hermitian Eigenvalue ProblemsabstractIn this paper, we propose a Parallel Hybrid Direct-Iterative Eigensolver for Hermitian Eigenvalue Problems without tridiagonalization, denoted byPHIDE, which combines direct and iterative methods.PHIDEfirst reduces a Hermitian matrix to banded form, then applies a spectrum slicing algorithm to the banded matrix, and finally computes the eigenvectors of the original matrix via backtransformation. Compared with conventional direct eigensolvers,PHIDEavoids tridiagonalization, which involves many memory-bound operations. InPHIDE, the banded eigenvalue problem is solved using the contour integral method implemented in FEAST, which may yield slightly lower accuracy than tridiagonalization-based approaches. For sequences of correlated Hermitian eigenvalue problems arising in density functional theory (DFT),PHIDEachieves an average speedup of$1.22\times$over the state-of-the-art direct solver in ELPA when using 1024 processes. Numerical experiments are conducted on dense Hermitian matrices from real applications as well as large sparse matrices from the SuiteSparse and ELSES collections. Shengguo Li, Xinzhe Wu, José E. Román, Ziyang Yuan, Ruibo Wang, Xuguang Chen |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2026 | Fully Decentralized Data Distribution for Large-Scale HPC SystemsabstractFor many years, in the HPC data distribution scenario, as the scale of the HPC system continues to increase, manufacturers have to increase the number of data providers to improve the IO parallelism to match the data demanders. In large-scale, especially exascale HPC systems, this mode of decoupling the demander and provider presents significant scalability limitations and incurs substantial costs. In our view, only a distribution model in which the demander also acts as the provider can fundamentally cope with changes in scale and have the best scalability, which is called all-to-all data distribution mode in this paper. We design and implement the BitTorrent protocol on computing networks in HPC systems and propose FD3, a fully decentralized data distribution method. We design the Requested-to-Validated Table (RVT) and the Highest ranking and Longest consecutive piece segment First (HLF) policy based on the features of the HPC networking environment to improve the performance of FD3. In addition, we design a torrent-tree to accelerate the distribution of seed file data and the aggregation of distribution state, and release the tracker load with neighborhood local-generation algorithm. Experimental results show that FD3 can scale smoothly to 11k+ computing nodes, and its performance is much better than that of the parallel file system. Compared with the original BitTorrent, the performance is improved by 8-15 times. FD3 highlights the considerable potential of the all-to-all model in HPC data distribution scenarios. Furthermore, the work of this paper can further stimulate the exploration of future distributed parallel file systems and provide a foundation and inspiration for the design of data access patterns for Exscale HPC systems. Ruibo Wang, Mingtian Shao, Huijun Wu 0001, Yiqin Dai, Kai Lu 0001 |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2026 | Coverage and Rate Analysis of Follower-Based LEO Satellite Networks: A Stochastic Geometry ApproachabstractTo mitigate inter-satellite interference and payload limits in LEO mega-constellations, satellite clusters, groups of small cooperative satellites have been proposed to improve performance and reduce interference. The typical configuration divides the cluster into a leader satellite with full processing and control capabilities and multiple simpler follower satellites that assist with coverage and throughput. These clusters enhance coverage and throughput, prompting interest in their performance gains and optimal deployment. Given that the spherical stochastic geometry (SG) model has been proven effective for modeling such structures, we establish a performance evaluation framework based on the SG approach for the leader-follower satellite architecture, enabling an assessment of communication performance under different deployment configurations quantitatively. We derive analytical expressions for the outage probability and average data rate to evaluate the communication performance of the satellite system, along with low-complexity approximations. Numerical results demonstrate the performance advantages of the leader-follower architecture over a single leader satellite and explore optimal deployment configurations for the follower satellites. Juanjuan Ru, Ruibo Wang, Mohamed-Slim Alouini |
IEEE Trans. Wirel. Commun. | 2 |
| 2025 | MergeFS: Optimizing Node-Local Burst Buffers for Complex HPC WorkflowsabstractHigh-performance computing (HPC) applications are increasingly transitioning from traditional numerical simulations to an intelligent fusion paradigm integrating AI algorithms and big data analytics, exemplified by initiatives such as AI4Science. This evolution, coupled with rising problem complexity, results in workflows composed of interdependent subtasks. Existing HPC storage solutions, particularly burst buffer systems, have yet to adequately address the unique challenges posed by such workflows, including efficient cross-task data sharing and namespace fusion, leading to suboptimal resource utilization and performance bottlenecks in complex dependency scenarios. In this paper, we present MergeFS, a lightweight, workflow-aware burst buffer file system that incorporates a treestructured workflow registry for precise dependency management alongside a dynamic multi-namespace mechanism enabling rapid and isolated data access. MergeFS effectively integrates workflow management, namespace control, and data view fusion. Experimental evaluations demonstrate that MergeFS significantly outperforms current workflow-centric burst buffer optimizations in runtime performance with low management overhead. Zhaohao Zhong, Huijun Wu 0001, Yong Dong, Zhenwei Wu, Ruibo Wang |
ICPADS | 8 |
| 2025 | From Islands to Archipelago: Towards Collaborative and Adaptive Burst Buffer for HPC SystemsabstractModern supercomputers increasingly use node-local storage as burst buffers (BB) to address I/O bottlenecks.However, current BBs do not naturally support workflows, a common workload in HPC consisting of many interconnected subtasks.Workflow I/O can be divided into three types: intratask I/O, inter-task I/O, and stage-in/out I/O.While BBs can accelerate intra-task I/O, they often overlook the other two.Inter-task I/O relies on migrating data through the Parallel File System (PFS), which can slow down overall performance.Although allocating more resources to create larger BBs could help, it increases costs.Additionally, temporary BBs lack permanent storage, requiring data migration between the PFS and BB for stage-in and stage-out I/O.This process often involves multiple data copies and reduces I/O efficiency.Even for intra-task I/O, unbalanced data distribution can cause bottlenecks on heavily loaded nodes.To improve workflow acceleration in BB systems, it is important to address the needs of all the above-mentioned three Mingtian Shao, Ruibo Wang, Kai Lu 0001, Yiqin Dai, Huijun Wu 0001 |
ICS | 2 |
| 2025 | ASC-Hook: Efficient System Call Interception for ARMabstractSystem call interception is essential for tools that modify or monitor application behavior. However, current system call interception solutions on ARM platforms still face challenges related to performance and completeness. This paper introduces ASC-Hook, an efficient and comprehensive binary rewriting framework specifically designed for intercepting system calls on ARM architectures. ASC-Hook tackles two critical challenges: the misalignment of the target address caused by directly replacing the SVC instruction with BR x8, and the return to the original control flow after system call interception. To achieve this, we propose a hybrid replacement strategy combined with a customized trampoline mechanism. Additionally, multiple completeness strategies tailored for system call interception are implemented to guarantee thorough coverage. Experimental evaluations demonstrate that ASC-Hook reduces overhead to as low as 1/29 of existing solutions, while incurring an average performance loss of only 3.8% in system call-intensive applications. Ruibo Wang, Gen Zhang |
LCTES | 5 |
| 2025 | Performance Analysis of Single Photon Detector-Based High-Speed Communication SystemsabstractSingle-photon detector (SPD)-based communication systems, and in particular those employing superconducting nanowire single-photon detectors (SNSPDs), are playing an irreplaceable role in scenarios such as quantum communication and deep-space communication. Due to the high cost of experimental equipment, establishing a theoretical framework to analyze the performance of SPD-based systems has become an effective and low-cost solution. However, due to the unique detection mechanism of SPDs and the decisive impact of dead time, there is currently no analytical framework suitable for evaluating the performance of high-speed communication systems with SPDs. To fill this gap, we propose an analytical framework tailored to SPD-based pulse-position modulation (PPM) systems, based on a Markov detection model, and use this framework to evaluate the system’s symbol error rate and achievable symbol rate. The framework demonstrates significant advantages over simulations and experiments, particularly in its ability to predict theoretical performance limits. Based on this framework, we reveal unique characteristics of the SPD-based PPM system, such as channel asymmetry and detection dependence. In addition, optimization guidelines for three representative system configurations are provided. Ziyuan Shi, Ruibo Wang, Mohamed-Slim Alouini |
IEEE Trans. Commun. | 2 |
| 2025 | MIST: Towards MPI Instant Startup and Termination on Tianhe HPC SystemsabstractAs the size of MPI programs grows with expanding HPC resources and parallelism demands, the overhead of MPI startup and termination escalates due to the inclusion of less scalable global operations. Global operations involving extensive cross-machine communication and synchronization are crucial for ensuring semantic correctness. The current focus is on optimizing and accelerating these global operations rather than removing them, as the latter involves systematic changes to the system software stack and may impact program semantics. Given this background, we propose a systematic solution named MIST to safely eliminate global operations in MPI startup and termination. Through optimizing the generation of communication addresses, designing reliable communication protocols, and exploiting the resource release mechanism, MIST eliminates all global operations to achieve MPI instant startup and termination while ensuring correct program execution. Experiments on Tianhe-2A supercomputer demonstrate that MIST can reduce theMPI_Init()time by 32.5-77.6% and theMPI_Finalize()time by 28.9-85.0%. Yiqin Dai, Ruibo Wang, Yong Dong, Juan Chen 0001, Huijun Wu 0001, Mingtian Shao, Kai Lu 0001 |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2024 | YFLM: An Improved Levenberg-Marquardt Algorithm for Global Bundle Adjustment
Jie Liu 0002, Ruibo Wang |
CGI (2) | 5 |
| 2024 | Fully Decentralized Data Distribution for Exascale-HPC: End of the Provider-Demander Matching PuzzleabstractFor many years, in the HPC data distribution scenario, as the scale of the HPC system continues to increase, manufacturers have to increase the number of data providers to improve the IO parallelism to match the data demanders. In the era of Exascale Computing, this mode of decoupling the demander and provider has limited scalability and huge costs. In our view, only a distribution model in which the demander also acts as the provider can fundamentally cope with changes in scale and have the best scalability, which is called all-to-all data distribution mode in this paper. We design and implement the BitTorrent protocol on computing networks in HPC systems and propose FD3, a fully decentralized data distribution method. We design the Requested-to-Validated Table (RVT) and the Nearest and Longest consecutive piece Segment First (NLSF) policy based on the features of the HPC networking environment to improve the performance of FD3. Experimental results show that FD3 can scale smoothly to 11k+ computing nodes, and its performance is much better than that of the parallel file system. Compared with the original BitTorrent, the performance is improved by 7–11 times. FD3 shows the great potential of the all-to-all model in HPC data distribution scenarios. At the same time, the work of this paper can further stimulate the exploration of future distributed parallel file systems and provide a foundation and inspiration for the design of data access patterns for Exscale HPC systems. Mingtian Shao, Ruibo Wang, Huijun Wu 0001, Yiqin Dai, Kai Lu 0001 |
CLUSTER | 3 |
| 2024 | Joint Robust Secure Beamforming Designs for ISAC-Enabled LEO Satellite SystemsabstractThe security of low-earth orbit (LEO) satellite communication systems faces challenges due to the high-speed movement characteristic. In this paper, we study the secure transmission for an ISAC-enabled LEO satellite system by considering the sensing function to enhance security. With the assistance of the sensing capability, more precise angle information of potential targets/eavesdroppers (Eve) can be estimated, thereby strengthening the system's performance. However, it is impossible to avoid the errors. To tackle this problem, we consider the channel uncertainty model and maximize the sum secrecy rate by jointly designing the secure transmit beamforming and radar receive filters. Furthermore, the non-convex problem is transformed into a series of convex optimization sub-problems, and the required optimization parameters are obtained through iterative calculation. Simulation results verify the advantages of the proposed scheme in achieving secure transmission of the LEO satellite system. Ruibo Wang, Bodong Shang, Mohamed-Slim Alouini |
ICC | 2 |
| 2024 | Optimizing General Matrix Multiplications on Modern Multi-core DSPsabstractGeneral Matrix Multiplication (GEMM) is a key subprogram in high-performance computing (HPC) and deep learning workloads. With the rising significance of power and energy consumption in HPC systems, accelerators based on Digital Signal Processors (DSPs) have been integrated into general-purpose HPC systems. Due to the architecture disparities, the GEMM optimization techniques used on conventional multi-core CPUs and GPGPUs are not always applicable to DSPs. This paper shares our experience in optimizing GEMM on multi-core GPDSPs, using a CPU-DSP processor as a case study. Our approach employs a range of techniques to optimize performance for DSP architectures. These include data partitioning, three-level pipelining, dedicated micro-kernel design, and improved vector reduction. These optimizations maximize the overlap between computation and communication while fully exploiting the capabilities of floating-point arithmetic units to achieve high performance. Our experimental results demonstrate that the performance attained by our optimization is up to 96% of the theoretical peak performance of the hardware. Kainan Yu, Xinxin Qi, Peng Zhang 0061, Jianbin Fang, Dezun Dong, Ruibo Wang, Tao Tang 0001, Chun Huang 0006, Yonggang Che, Zheng Wang 0001 |
IPDPS | 6 |
| 2024 | Towards Highly Compatible I/O-Aware Workflow Scheduling on HPC SystemsabstractScientific workflows on High-Performance Computing (HPC) consist of multiple data processing and computing tasks with dependencies. Efficiently scheduling computing resources and multi-tier storage across workflow tasks is crucial for optimizing performance. Existing solutions often fall short in achieving the co-scheduling of computing and 1/O resources and lack compatibility with HPC system software. In this paper, we introduce a performance model for scheduling workflows on HPC systems to enhance the understanding of workflow scheduling and facilitate the testing of the scheduling algorithm. Additionally, we propose THman, an open-source scientific workflow scheduler featuring our heuristic scheduling algorithm, Highest Contribution First (HCF). THman achieves online co-scheduling of computing and I/O resources for workflows and is designed to work with traditional HPC batch schedulers for high compatibility. We evaluate THman using simulated workloads and real-world workflow applications. Experimental results show that THman reduces workflow makespan by up to 30.9% compared to alternative methods. Yiqin Dai, Ruibo Wang, Yong Dong, Kai Lu 0001 |
SC | 2 |
| 2024 | Enhancing Physical-Layer Security in LEO Satellite-Enabled IoT Network CommunicationsabstractThe extensive deployment of low earth orbit (LEO) satellites introduces significant security challenges for communication security issues in Internet of Things (IoT) networks. With the rising number of satellites potentially acting as eavesdroppers, integrating physical-layer security (PLS) into satellite communications has become increasingly critical. However, these studies are facing challenges, such as dealing with dynamic topology difficulties, limitations in interference analysis, and the high complexity of performance evaluation. To address these challenges, for the first time, we investigate the PLS strategies in satellite communications using the stochastic geometry (SG) analytical framework. We consider the uplink communication scenario in an LEO-enabled IoT network, where the multitier satellites from different operators, respectively, serve as the legitimate receivers and eavesdroppers. In this scenario, we derive low-complexity analytical expressions for the security performance metrics, namely availability probability, successful communication probability, and secure communication probability. By introducing the power allocation parameters, we incorporate the artificial noise (AN) technique, which is an important PLS strategy, into this the analytical framework, and evaluate the gains it brings to secure transmission. In addition to the AN technique, we also analyse the impact of constellation configuration, physical-layer parameters, and network layer parameters on the aforementioned metrics. Anna Talgat, Ruibo Wang, Mustafa A. Kishk, Mohamed-Slim Alouini |
IEEE Internet Things J. | 2 |
| 2024 | Ultra Reliable Low Latency Routing in LEO Satellite Constellations: A Stochastic Geometry ApproachabstractIn recent years, LEO satellite constellations have become envisioned as a core component of the next-generation wireless communication networks. The successive establishment of mega satellite constellations has triggered further demands for satellite communication advanced features: high reliability and low latency. In this article, we first establish a multi-objective optimization problem that simultaneously maximizes reliability and minimizes latency, then we solve it by two methods. According to the optimal solution, ideal upper bounds for reliability and latency performance of LEO satellite routing can be derived. Next, we design an algorithm for relay satellite subset selection, which can approach the ideal upper bounds in terms of performance. Furthermore, we derive analytical expressions for satellite availability, coverage probability, and latency under the stochastic geometry (SG) framework, and the accuracy is verified by Monte Carlo simulation. In the numerical results, we study the routing performance of three existing mega constellations and the impact of different constellation parameter configurations on performance. By comparing with existing routing strategies, we demonstrate the advantages of our proposed routing strategy and extend the scope of our research. Ruibo Wang, Mustafa A. Kishk, Mohamed-Slim Alouini |
IEEE J. Sel. Areas Commun. | 1 |
| 2024 | Towards adaptive graph neural networks via solving prior-data conflictsabstractGraph neural networks (GNNs) have achieved remarkable performance in a variety of graph-related tasks. Recent evidence in the GNN community shows that such good performance can be attributed to the homophily prior; i.e., connected nodes tend to have similar features and labels. However, in heterophilic settings where the features of connected nodes may vary significantly, GNN models exhibit notable performance deterioration. In this work, we formulate this problem as prior-data conflict and propose a model called the mixture-prior graph neural network (MPGNN). First, to address the mismatch of homophily prior on heterophilic graphs, we introduce the non-informative prior, which makes no assumptions about the relationship between connected nodes and learns such relationship from the data. Second, to avoid performance degradation on homophilic graphs, we implement a soft switch to balance the effects of homophily prior and non-informative prior by learnable weights. We evaluate the performance of MPGNN on both synthetic and real-world graphs. Results show that MPGNN can effectively capture the relationship between connected nodes, while the soft switch helps select a suitable prior according to the graph characteristics. With these two designs, MPGNN outperforms state-of-the-art methods on heterophilic graphs without sacrificing performance on homophilic graphs. Xugang Wu, Huijun Wu 0001, Ruibo Wang, Xu Zhou 0004, Kai Lu 0001 |
Frontiers Inf. Technol. Electron. Eng. | 3 |
| 2024 | SNCL: a supernode OpenCL implementation for hybrid computing arrays
Tao Tang 0001, Kai Lu 0001, Lin Peng 0001, Yingbo Cui 0001, Jianbin Fang, Chun Huang 0006, Ruibo Wang, Canqun Yang, Yifei Guo |
J. Supercomput. | 7 |
| 2024 | Faster and Scalable MPI Applications LaunchingabstractDistributed parallel MPI applications are the dominant workload in many high-performance computing systems. While optimizing MPI application execution is a well-studied field, little work has considered optimizing the initial MPI application launching phase, which incurs extensive cross-machine communications and synchronization. The overhead of MPI application launching can be expensive, accounting for more than million core hours per 10K nodes annually on the production Tianhe-2A supercomputer, which will increase as the number of parallel machines used grows. Therefore, it is critical to optimize the MPI application launching process. This paper presents a novel approach to optimizing the MPI application launch. Our approach adopts a location-aware address generation rule to eliminate the need for address exchange and a topology-aware global communication scheme to optimize cross-machine synchronization. We then design a new application launch procedure to support the proposed optimizations to further reduce the pressure of the shared I/O system. Our techniques have been deployed to production in the Tianhe-2A supercomputer and the Next Generation Tianhe Supercomputer. Experimental results show that our approach scales well and outperforms alternative schemes, reducing the MPI application launching time by over 29% with 320K MPI processes. Yong Dong, Yiqin Dai, Kai Lu 0001, Ruibo Wang, Juan Chen 0001, Mingtian Shao, Zheng Wang 0001 |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2024 | Reliability Analysis of Multi-Hop Routing in Multi-Tier LEO Satellite NetworksabstractThis article studies the reliability of multi-hop routing in a multi-tier hybrid satellite-terrestrial relay network (HSTRN). We evaluate the reliability of multi-hop routing by introducing interruption probability, which is the probability that no relay device (ground gateway or satellite) is available during a hop. The single-hop interruption probability is derived and extended to the multi-hop interruption probability using a stochastic geometry-based approach. Since the interruption probability in HSTRN highly depends on the priority of selecting communication devices at different tiers, we propose three priority strategies: (i) stationary optimal priority strategy, (ii) single-hop interruption probability inspired strategy, and (iii) density inspired strategy. Among them, the interruption probability under the stationary optimal priority strategy can approach the ideal lower bound. However, when analyzing an HSTRN with a large number of tiers, the stationary optimal priority strategy is computationally expensive. The single-hop interruption probability inspired strategy is expected to be a low-complexity but less reliable alternative to the stationary optimal priority strategy. In numerical results, we study the complementarity between terrestrial devices and satellites. Furthermore, analytical results for reliability are also applicable to the analysis of satellite availability, coverage probability, and ultra-reliable and low latency communications (URLLC) rate. Finally, we extend our original routing strategy into a multi-flow one with dynamic priority strategy. Ruibo Wang, Mustafa A. Kishk, Mohamed-Slim Alouini |
IEEE Trans. Wirel. Commun. | 1 |
| 2023 | We Need to Talk About Reproducibility in NLP Model ComparisonabstractNLPers frequently face reproducibility crisis in a comparison of various models of a realworld NLP task.Many studies have empirically showed that the standard splits tend to produce low reproducible and unreliable conclusions, and they attempted to improve the splits by using more random repetitions.However, the improvement on the reproducibility in a comparison of NLP models is limited attributed to a lack of investigation on the relationship between the reproducibility and the estimator induced by a splitting strategy.In this paper, we formulate the reproducibility in a model comparison into a probabilistic function with regard to a conclusion.Furthermore, we theoretically illustrate that the reproducibility is qualitatively dominated by the signal-tonoise ratio (SNR) of a model performance estimator obtained on a corpus splitting strategy.Specifically, a higher value of the SNR of an estimator probably indicates a better reproducibility.On the basis of the theoretical motivations, we develop a novel mixture estimator of the performance of an NLP model with a regularized corpus splitting strategy based on a blocked 3 × 2 cross-validation.We conduct numerical experiments on multiple NLP tasks to show that the proposed estimator achieves a high SNR, and it substantially increases the reproducibility.Therefore, we recommend the NLP practitioners to use the proposed method to compare NLP models instead of the methods based on the widely-used standard splits and the random splits with multiple repetitions. Yan Xue, Xuefei Cao, Xingli Yang, Yu Wang 0092, Ruibo Wang, Jihong Li |
EMNLP | 5 |
| 2023 | An Improved Cross-Validated Adversarial Validation Method
Zhengjiang Liu, Yan Xue, Ruibo Wang, Xuefei Cao, Jihong Li |
KSEM (1) | 4 |
| 2023 | Leveraging Free Labels to Power up Heterophilic Graph Learning in Weakly-Supervised Settings: An Empirical Study
Xugang Wu, Huijun Wu 0001, Ruibo Wang, Duanyu Li, Xu Zhou 0004, Kai Lu 0001 |
ECML/PKDD (3) | 3 |
| 2023 | MT-office: parallel password recovery program for office on domestic heterogeneous multi-core processor
Yongtao Luo, Bo Yang 0023, Jie Liu 0002, Ruibo Wang, Jinmin Wen, Tiaojie Xiao, Xuguang Chen, Chunye Gong |
CCF Trans. High Perform. Comput. | 4 |
| 2023 | Programming bare-metal accelerators with heterogeneous threading models: a case study of Matrix-3000abstractAs the hardware industry moves toward using specialized heterogeneous many-core processors to avoid the effects of the power wall, software developers are finding it hard to deal with the complexity of these systems. In this paper, we share our experience of developing a programming model and its supporting compiler and libraries for Matrix-3000, which is designed for next-generation exascale supercomputers but has a complex memory hierarchy and processor organization. To assist its software development, we have developed a software stack from scratch that includes a low-level programming interface and a high-level OpenCL compiler. Our low-level programming model offers native programming support for using the bare-metal accelerators of Matrix-3000, while the high-level model allows programmers to use the OpenCL programming standard. We detail our design choices and highlight the lessons learned from developing system software to enable the programming of bare-metal accelerators. Our programming models have been deployed in the production environment of an exascale prototype system. Jianbin Fang, Peng Zhang 0061, Chun Huang 0006, Tao Tang 0001, Kai Lu 0001, Ruibo Wang, Zheng Wang 0001 |
Frontiers Inf. Technol. Electron. Eng. | 6 |
| 2023 | Ensemble Feature Selection With Block-Regularized m × 2 Cross-ValidationabstractEnsemble feature selection (EFS) has attracted significant interest in the literature due to its great potential in reducing the discovery rate of noise features and stabilizing the feature selection results. In view of the superior performance of block-regularized m × 2 cross-validation on generalization performance and algorithm comparison, a novel EFS technology based on block-regularized m × 2 cross-validation is proposed in this study. Contrary to the traditional ensemble learning with a binomial distribution, the distribution of feature selection frequency in the proposed technique is approximated by a beta distribution more accurately. Furthermore, theoretical analysis of the proposed technique shows that it yields a higher selection probability for important features, lower selected risk for noise features, more true positives, and fewer false positives. Finally, the above conclusions are verified by the simulated and real data experiments. Xingli Yang, Yu Wang 0045, Ruibo Wang, Jihong Li |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2023 | Resident Population Density-Inspired Deployment of K-Tier Aerial Cellular NetworkabstractUsing Unmanned Aerial Vehicles (UAVs) to enhance network coverage has proven a variety of benefits compared to terrestrial counterparts. One of the commonly used mathematical tools to model the locations of the UAVs is stochastic geometry (SG). However, in the existing studies, both users and UAVs are often modeled as homogeneous point processes. In this paper, we consider an inhomogeneous Poisson point process (PPP)-based model for the locations of the users that captures the degradation in the density of active users as we move away from the town center. In addition, we propose the deployment of aerial vehicles following the same inhomogeneity of the users to maximize the performance. In addition, a multi-tier network model is also considered to make better use of the rich space resources. Then, the analytical expressions of the coverage probability for a typical user and the total coverage probability are derived. Finally, we optimize the coverage probability with limitations of the total number of UAVs and the minimum local coverage probability. Finally we give the optimal UAV distribution parameters when the maximum overall coverage probability is reached. Ruibo Wang, Mustafa A. Kishk, Mohamed-Slim Alouini |
IEEE Trans. Wirel. Commun. | 1 |
| 2022 | XTree: Traversal-Based Partitioning for Extreme-Scale Graph Processing on SupercomputersabstractGraph algorithms, such as Breadth First Search (BFS), Single Source Shortest Path (SSSP), PageRank (PR), and Connected Components (CC), are increasingly important in big data processing and analytics. As graph scales (numbers of vertices and edges) have increased from billions to trillions, Supercomputers have huge numbers (up to hundreds of thousands) of computing nodes (CNs) that can provide ultra-high aggregate computing power and memory capacity, thus being particularly suitable for processing extreme-scale graphs with trillions of vertices and edges. However, existing cluster-based graph-parallel systems perform poorly when deployed on supercomputers, since their partitioning methods overlook the hierarchical nature of supercomputer networks and incur prohibitive communication storm. This paper presents XTree, an efficient traversal-based partitioning method for minimizing communication overhead of graph processing on supercomputers. We observe that supercomputers' huge numbers of CNs are usually organized into hierarchical communication domains, which can be modeled as a domain tree where communication in lower-level domains is significantly faster than that in higher-level ones. Therefore, the key idea of XTree's partitioning is to exploit hierarchical locality by viewing the graph as a BFS tree and leveraging the topology knowledge to map the graph's BFS tree onto the domain tree, We evaluate the effectiveness of XTree by running various graph algorithms, on both real-world big graphs and synthetic trillion-scale graphs. XTree substantially reduces communication overhead and achieves orders of magnitude speedup against the Graph500 reference implementations with the state-of-the-art 2D-decomposition partitioning. Xinbiao Gan, Yiming Zhang 0003, Ruigeng Zeng, Jie Liu 0002, Ruibo Wang, Li Chen 0008, Kai Lu 0001 |
ICDE | 5 |
| 2022 | The Fast and Scalable MPI Application Launch of the Tianhe HPC systemabstractFast and scalable MPI application launch helps achieve exascale performance and is becoming a common goal in high-performance computing. However, the traditional launch technique suffers from scalability deficiencies in the global information exchange and the global barrier operation. This drawback makes it challenging to launch MPI applications quickly in large-scale systems. In this paper, we propose a fast and scalable application launch technique and details its associated hardware and software support. The optimized launch technique includes a locality-aware static address generation rule for eliminating the need for address exchange and a topology-aware global communication scheme for improving global communication efficiency. We also propose an optimized application launch sequence for supporting the above launch technique. We implement and evaluate the proposed launch technique on the Tianhe-2A supercomputer and the Tianhe Exascale Prototype Upgrade System. Experimental results show that our technique can reduce the launch time by 26.1% when launching an application with 256K processes. Yiqin Dai, Yong Dong, Kai Lu 0001, Ruibo Wang, Mingtian Shao, Juan Chen 0001 |
IPDPS | 5 |
| 2022 | Towards Scalable Resource Management for SupercomputersabstractToday's supercomputers offer massive computation resources to execute a large number of user jobs. Effectively managing such large-scale hardware parallelism and workloads is essential for supercomputers. However, existing HPC resource management (RM) systems fail to capitalize on the hardware parallelism by following a centralized design used decades ago. They give poor scalability and inefficient performance on today's supercomputers, which will worsen in exascale computing. We present ESlurm, a better RM for supercomputers. As a departure from existing HPC RMs, ESlurm implements a distributed communication structure. It employs a new communication tree strategy and uses job runtime estimation to improve communications and job scheduling efficiency. ESlurm is deployed into production in a real supercomputer. We evaluate ESlurm on up to 20K nodes. Compared to state-of-the-art RM solutions, ESlurm exhibits better scalability, significantly reducing the resource usage of master nodes and improving data transfer and job scheduling efficiency by a large margin. Yiqin Dai, Yong Dong, Kai Lu 0001, Ruibo Wang, Wei Zhang 0027, Juan Chen 0001, Mingtian Shao, Zheng Wang 0001 |
SC | 4 |
| 2022 | vGraph: Memory-Efficient Multicore Graph Processing for Traversal-Centric AlgorithmsabstractTo lower the monetary/energy cost, single-machine multicore graph processing is gaining increasing attention for a wide range of traversal-centric graph algorithms such as BFS, SSSP, CC, and PageRank, of which the processing is relatively simple and the topology data (vertices and edges) dominates the memory footprint. This paper presents$v$Graph, a NUMA-aware, memory-efficient multicore graph processing system for traversal-centric algorithms.$v$Graph proposes an ultralight NUMA-aware graph preprocessing scheme which eliminates almost all complex preprocessing steps and pipelines per-NUMA graph loading and compressing, to effectively reduce inter-NUMA memory accesses while keeping both preprocessing cost and peak memory footprint low. We further optimize$v$Graph with effective HPC techniques including prefetching and work-stealing. Evaluation on a 384GB-memory, four-NUMA machine shows that compared to the state-of-the-art NUMA-aware/-unaware systems,$v$Graph can process much larger real-world and synthetic graphs with various traversal-centric algorithms, achieving significantly higher memory efficiency and lower processing time. Menghan Jia, Yiming Zhang 0003, Xinbiao Gan, Dongsheng Li 0001, Erci Xu, Ruibo Wang, Kai Lu 0001 |
SC | 6 |
| 2022 | MT-3000: a heterogeneous multi-zone processor for HPC
Kai Lu 0001, Yang Guo 0003, Chun Huang 0006, Sheng Liu 0001, Ruibo Wang, Jianbin Fang, Tao Tang 0001, Zhaoyun Chen, Biwei Liu, Zhong Liu 0003, Yuanwu Lei, Haiyan Sun |
CCF Trans. High Perform. Comput. | 6 |
| 2022 | TEES: topology-aware execution environment service for fast and agile application deployment in HPCabstractHigh-performance computing (HPC) systems are about to reach a new height: exascale. Application deployment is becoming an increasingly prominent problem. Container technology solves the problems of encapsulation and migration of applications and their execution environment. However, the container image is too large, and deploying the image to a large number of compute nodes is time-consuming. Although the peer-to-peer (P2P) approach brings higher transmission efficiency, it introduces larger network load. All of these issues lead to high startup latency of the application. To solve these problems, we propose the topology-aware execution environment service (TEES) for fast and agile application deployment on HPC systems. TEES creates a more lightweight execution environment for users, and uses a more efficient topology-aware P2P approach to reduce deployment time. Combined with a split-step transport and launch-in-advance mechanism, TEES reduces application startup latency. In the Tianhe HPC system, TEES realizes the deployment and startup of a typical application on 17 560 compute nodes within 3 s. Compared to container-based application deployment, the speed is increased by 12-fold, and the network load is reduced by 85%. Mingtian Shao, Kai Lu 0001, Wanqing Chi, Ruibo Wang, Yiqin Dai |
Frontiers Inf. Technol. Electron. Eng. | 4 |
| 2022 | TianheGraph: Customizing Graph Search for Graph500 on Tianhe SupercomputerabstractAs the era of exascale supercomputing is coming, it is vital for next-generation supercomputers to find appropriate applications with high social and economic benefit. In recent years, it has been widely accepted that extremely-large graph computation is a promising killer application for supercomputing. Although Tianhe series supercomputers are leading in the world-wide competition of supercomputing (ranked No. 1 in the Top500 list for six times), previously they had been inefficient in graph computation according to the Graph500 list. This is mainly because the previous graph processing system cannot leverage the advanced hardware features of Tianhe supercomputers. To address the problem, in this paper we present our integrated optimizations for improving the graph computation performance on our next-generation Tianhe supercomputing system, mainly including sorting with buffering for heavy vertices, vectorized searching with SVE (Scalable Vector Extension) on matrix2000+ CPUs, and group communication on the proprietary interconnection network. Performance evaluation on a subset of the Tianhe supercomputer (with 512 nodes and 196,608 cores) shows that our customized graph processing system effectively improves the graph search performance and achieves the BFS performance of 2131.98 GTEPS. Xinbiao Gan, Yiming Zhang 0003, Ruibo Wang, Tiaojie Xiao, Ruigeng Zeng, Jie Liu 0002, Kai Lu 0001 |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2020 | A survey on optimizations towards best-effort hardware transactional memory
Zhenwei Wu, Kai Lu 0001, Ruibo Wang |
CCF Trans. High Perform. Comput. | 3 |
| 2020 | Tuning Parameter Selection Based on Blocked 3˟ 2 Cross-Validation for High-Dimensional Linear Regression Model
Xingli Yang, Yu Wang 0092, Ruibo Wang, Jihong Li |
Neural Process. Lett. | 3 |
| 2019 | Bayes Test of Precision, Recall, and F1 Measure for Comparison of Two Natural Language Processing ModelsabstractDirect comparison on point estimation of the precision (P), recall (R), and F1 measure of two natural language processing (NLP) models on a common test corpus is unreasonable and results in less replicable conclusions due to a lack of a statistical test. However, the existing t-tests in cross-validation (CV) for model comparison are inappropriate because the distributions of P, R, F1 are skewed and an interval estimation of P, R, and F1 based on a t-test may exceed [0,1]. In this study, we propose to use a block-regularized 3×2 CV (3×2 BCV) in model comparison because it could regularize the difference in certain frequency distributions over linguistic units between training and validation sets and yield stable estimators of P, R, and F1. On the basis of the 3×2 BCV, we calibrate the posterior distributions of P, R, and F1 and derive an accurate interval estimation of P, R, and F1. Furthermore, we formulate the comparison into a hypothesis testing problem and propose a novel Bayes test. The test could directly compute the probabilities of the hypotheses on the basis of the posterior distributions and provide more informative decisions than the existing significance t-tests. Three experiments with regard to NLP chunking tasks are conducted, and the results illustrate the validity of the Bayes test. Ruibo Wang, Jihong Li |
ACL (1) | 1 |
| 2019 | A Comprehensive Assessment of Modis-Derived Instantaneous Net Surface Shortwave Radiation using the in-Situ Fluxnet DatabaseabstractNet Surface Shortwave Radiation (NSSR) is a key component of the surface radiation budget, which controls the energy, water exchanges, and many physical processes. The primary purpose of this study is to build a concise and feasible model of estimating NSSR with data of Moderate Resolution Imaging Spectroradiometer (MODIS) onboard the Terra satellite. Random Forest (RF) machine learning method was applied to building the model with the FULXNET in situ observations because of its powerful ability in nonlinear fitting. Total 17 variables are considered in RF model for retrieving NSSR, and thousands of parameters combinations were carried to obtained optimal parameters of the proposed model. The Bias, RMSE, and R2of estimated instantaneous NSSR in 95 selected sites during 2014 whole year are -0.085 W m-2, 28.274 W m-2and 0.989, respectively. Consequently, retrieval of instantaneous NSSR with RF method would be believed to be an efficient method in the future by considering its concise process and great accuracy. Wangmin Ying, Ruibo Wang, Lu Niu, Hua Wu 0001 |
IGARSS | 2 |
| 2019 | Block-regularized repeated learning-testing for estimating generalization error
Ruibo Wang, Jihong Li, Xingli Yang |
Inf. Sci. | 1 |
| 2019 | Calibrating GloVe model on the principle of Zipf's law
Xuefei Cao, Jihong Li, Ruibo Wang, Yu Wang 0092, Qian Niu, Junfeng Shi |
Pattern Recognit. Lett. | 3 |
| 2017 | Block-Regularized m × 2 Cross-Validated Estimator of the Generalization ErrorabstractA cross-validation method based on [Formula: see text] replications of two-fold cross validation is called an [Formula: see text] cross validation. An [Formula: see text] cross validation is used in estimating the generalization error and comparing of algorithms' performance in machine learning. However, the variance of the estimator of the generalization error in [Formula: see text] cross validation is easily affected by random partitions. Poor data partitioning may cause a large fluctuation in the number of overlapping samples between any two training (test) sets in [Formula: see text] cross validation. This fluctuation results in a large variance in the [Formula: see text] cross-validated estimator. The influence of the random partitions on variance becomes serious as [Formula: see text] increases. Thus, in this study, the partitions with a restricted number of overlapping samples between any two training (test) sets are defined as a block-regularized partition set. The corresponding cross validation is called block-regularized [Formula: see text] cross validation ([Formula: see text] BCV). It can effectively reduce the influence of random partitions. We prove that the variance of the [Formula: see text] BCV estimator of the generalization error is smaller than the variance of [Formula: see text] cross-validated estimator and reaches the minimum in a special situation. An analytical expression of the variance can also be derived in this special situation. This conclusion is validated through simulation experiments. Furthermore, a practical construction method of [Formula: see text] BCV by a two-level orthogonal array is provided. Finally, a conservative estimator is proposed for the variance of estimator of the generalization error. Ruibo Wang, Yu Wang 0092, Jihong Li, Xingli Yang |
Neural Comput. | 1 |
| 2015 | Confidence Interval for F1 Measure of Algorithm Performance Based on Blocked 3× 2 Cross-ValidationabstractIn studies on the application of machine learning such as Information Retrieval (IR), the focus is typically on the estimation of the F1measure of algorithm performance. Approximate symmetrical confidence intervals constructed by the F1value based on cross-validated L distribution are commonly used in the literature. However, theoretical analysis on the distribution of F1values shows that such distribution is actually non-symmetrical. Thus, simply using symmetrical distribution to approximate non-symmetrical distribution may be inappropriate and may result in a low degree of confidence and long interval length for the confidence interval. In the present study, a non-symmetrical confidence interval of the F1measure based on Beta prime distribution is constructed by using the F1value computed based on the average confusion matrix of a blocked 3 x 2 cross-validation. Experimental results show that in most cases, our method has high degrees of confidence. With an acceptable degree of confidence, our method has a shorter interval length than the approximate symmetrical confidence intervals based on the blocked 3 x 2 and 5 x 2 cross-validated L distributions. The approximate symmetrical confidence interval based on the 10-fold cross-validated L distribution has the shortest interval length of the four confidence intervals but with low degrees of confidence in all cases. Taking these two factors into consideration, our method is recommended. Yu Wang 0092, Jihong Li, Ruibo Wang, Xingli Yang |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2014 | Iaso: an autonomous fault-tolerant management system for supercomputers
Kai Lu 0001, Gen Li 0002, Ruibo Wang, Wanqing Chi, Yongpeng Liu, Hong-Wei Tang, Yinghui Gao |
Frontiers Comput. Sci. | 4 |
| 2014 | Blocked 3×2 Cross-Validated t-Test for Comparing Supervised Classification Learning AlgorithmsabstractIn the research of machine learning algorithms for classification tasks, the comparison of the performances of algorithms is extremely important, and a statistical test of significance for generalization error is often used to perform it in the machine learning literature. In view of the randomness of partitions in cross-validation, a new blocked 3×2 cross-validation is proposed to estimate generalization error in this letter. We then conduct an analysis of variance of the blocked 3×2 cross-validated estimator. A relatively conservative variance estimator that considers the correlation between any two two-fold cross-validations, and was previously neglected in 5×2 cross-validated t and F-tests is put forward. A corresponding test using this variance estimator is presented to compare the performances of algorithms. Simulated results show that the performance of our test is comparable with that of 5×2 cross-validated tests but with less computation complexity. Yu Wang 0092, Ruibo Wang, Huichen Jia, Jihong Li |
Neural Comput. | 2 |
| 2010 | Brief announcement: NUMA-aware transactional memoryabstractTransactional Memory (TM) research has focused on multi-core processors; limited research has been aimed at the clusters, leaving the area of NUMA (Non-Uniform Memory Access) system unexplored. The NUMA system's memory is physically distributed which brings the different access latency between local and remote memory. The existing TM design is not NUMA-aware which makes significant performance degradation on NUMA system. We introduce the latency-based conflict detection process and the forecasting-based conflict preventing method. The NUMA-aware strategies provide a good practical TM performance on NUMA system. Kai Lu 0001, Ruibo Wang, Xicheng Lu |
PODC | 2 |
| 2009 | Two-phase conflict detection for transactional memory on clustersabstractTransactional memory (TM) research has focused on multi-core processors; limited research has been aimed at the clusters. The intention of deploying TM on clusters is using more processors to solve big problems with this convenient technique. But the performance of the existing cluster's TM is poor because of the expensive remote access. The conflict detection, which is the most frequent operation of TM, is highly depending on the remote memory access. The remote memory access is usually 10 to 100 times slower than the local one in a cluster. We introduce the two-phase conflict detection strategy. By dividing the conflict detection process into two levels, the hierarchical strategy provides a good practical performance. Ruibo Wang, Kai Lu 0001, Xicheng Lu |
CLUSTER | 1 |
| 2009 | Investigating transactional memory performance on ccNUMA machinesabstractMost Software Transactional Memory (STM) research has focused on multi-core processors and small SMP machines; limited research has been aimed at the clusters, leaving the area of big SMP machines unexplored. Big SMP machine usually use Non-Uniform Memory Access (NUMA) to unburden the overloading between CPUs and the memory. In this paper, we evaluate several STM implementations on big SMP machine with cache coherent NUMA (ccNUMA) architecture. We found the remote memory access latency is the key factor influencing the STM performance. We also analyze the different design choices of STM. Finally, we conclude a specific design choice to achieve high performance in this domain. Ruibo Wang, Kai Lu 0001, Xicheng Lu |
HPDC | 1 |