EDBT 2026 Demo / reviewers in the wild / expert
Xing Cai
dblp:58/788
· DBLP profile ↗
41ranked-venue papers
6as first author
18since 2021 · last 2026
0000-0003-3706-4414ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 24 · 1 first-author · 6 since 2021Artificial intelligence and machine learning · 10 · 2 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Theory of computation · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | An efficient dominance decomposition-based deep graph evolutionary algorithm for the expensive multi-objective optimization
Xing Cai, Tong Zhang 0021, Zhen Cui 0001 |
Expert Syst. Appl. | 1 |
| 2026 | CAC: An asynchronous non-blocking consistency model with bounded staleness for distributed machine learningabstractRelaxed consistency models have been reported to significantly improve the performance of training machine learning (ML) models compared to strong consistency models. However, the existing relaxed consistency models for distributed ML either force fast workers to wait for stragglers (e.g., stale-synchronous models) or have no upper bound on data staleness (e.g., asynchronous models), thus negatively affecting the quality and performance of training ML models. We propose a new asynchronous non-blocking consistency model with bounded staleness for distributed ML that overcomes the drawbacks of blocking and unbounded staleness in previous relaxed consistency models. The new model, named Chained Asynchronous Consistency (CAC), guarantees an upper bound on data staleness in asynchronous computation without forcing fast workers to wait for slow workers. We theoretically prove that the Stochastic Gradient Descent (SGD) algorithm under CAC converges and the upper bound on the convergence expectation is independent of the number of workers. Based on the CAC model, we develop a new staleness-aware cacheable distributed object (CAC-object) for distributed ML where shared parameters are distributed among workers in a peer-to-peer manner. This approach avoids intermediate centralized storage, such as parameter servers, while maintaining simple one-sided communication (e.g., get and put ). The CAC-object allows remote parameters to be cached and reused locally following a consistency model (e.g., CAC, stale-synchronous or asynchronous model). To demonstrate the applicability of the CAC-object in asynchronous distributed ML, we introduce a new asynchronous distributed matrix completion algorithm (CAC-MF) using the CAC-object. We develop the CAC-object and CAC-MF using UPC++, a Partitioned Global Address Space (PGAS) library for high-performance computing (HPC), and evaluate them in different execution scenarios (e.g., with and without stragglers and crash failures, different minibatch sizes) on HPC clusters. Our experimental results show that the CAC model scales with increasing workers, tolerates stragglers and crash failures, and achieves better convergence than stale-synchronous and asynchronous models. Particularly, in the case of halted stragglers, CAC’s Root Mean Square Error (RMSE) is up to 25 times less than the baseline models’ RMSE for the Netflix dataset, and 170 times less for the MovieLens dataset. Phuong Hoai Ha, Xing Cai, Tan Nguyen 0001 |
Future Gener. Comput. Syst. | 2 |
| 2026 | Minimum Description Length-Driven Fragment Mining for Pretraining Molecule Property Prediction ModelabstractMolecular fragments play a crucial role in molecular property prediction. However, most existing deep learning approaches rely heavily on expert-defined substructural patterns, limiting their ability to identify novel or latent fragments. This constraint reduces the generalizability and applicability of molecular fragments in molecular representation learning. In this study, we propose the Molecular Multi-view Pre-training model with Adaptive Fragment Mining (MMP-AFM), a unified framework that facilitates the seamless integration of molecular structural information. MMP-AFM formulates fragment discovery as a combinatorial optimization problem, using description length as the objective to enable adaptive extraction of molecular fragments and dynamic construction of a fragment library. Additionally, we design a molecular multi-view self-supervised pretraining framework that aligns features from the fragment, global, and data augmentation views, ensuring a comprehensive integration of molecular substructural information. Finally, the MMP-AFM is applied to both molecular classification and regression tasks. Experimental results demonstrate that MMP-AFM consistently outperforms existing methods across multiple tasks, highlighting its broad applicability and efficiency. Xing Cai, Tong Zhang 0021, Yide Qiu, Baotong Su, Zhen Cui 0001 |
IEEE Trans. Comput. Biol. Bioinform. | 1 |
| 2025 | One for All: Universal Topological Primitive Transfer for Graph Structure LearningabstractThe non-Euclidean geometry inherent in graph structures fundamentally impedes cross-graph knowledge transfer. Drawing inspiration from texture transfer in computer vision, we pioneer topological primitives as transferable semantic units for graph structural knowledge. To address three critical barriers - the absence of specialized benchmarks, aligned semantic representations, and systematic transfer methodologies - we present G²SN-Transfer, a unified framework comprising: (i) TopoGraph-Mapping that transforms non-Euclidean graphs into transferable sequences via topological primitive distribution dictionaries; (ii) G²SN, a dual-stream architecture learning text-topology aligned representations through contrastive alignment; and (iii) AdaCross-Transfer, a data-adaptive knowledge transfer mechanism leveraging cross-attention for both full-parameter and parameter-frozen scenarios. Particularly, G²SN is a dual-stream sequence network driven by ordinary differential equations, and our theoretical analysis establishes the convergence guarantee of G²SN. We construct STA-18, the first large-scale benchmark with aligned topological primitive-text pairs across 18 diverse graph datasets. Comprehensive evaluations demonstrate that G²SN achieves state-of-the-art performance on four structural learning tasks (average 3.2\% F1-score improvement), while our transfer method yields consistent enhancements across 13 downstream tasks (5.2\% average gains) including 10 large-scale graph datasets. The datasets and code are available at https://anonymous.4open.science/r/UGSKT-C10E/. Yide Qiu, Tong Zhang 0021, Xing Cai, Zhen Cui 0001 |
NeurIPS | 3 |
| 2025 | UniHG: A Large-scale Universal Heterogeneous Graph Dataset and Benchmark for Representation Learning and Cross-Domain TransferringabstractIrregular data in the real world are usually organized as heterogeneous graphs consisting of multiple types of nodes and edges. However, current heterogeneous graph research confronts three fundamental challenges: i) Benchmark Deficiency, ii) Semantic Disalignment, and iii) Propagation Degradation. In this paper, we construct a large-scale, universal, and joint multi-domain heterogeneous graph dataset named UniHG to facilitate heterogeneous graph representation learning and cross-domain knowledge mining. Overall, UniHG contains 77.31 million nodes and 564 million directed edges with thousands of labels and attributes, which is currently the largest universal heterogeneous graph dataset available to the best of our knowledge. To perform effective learning and provide comprehensively benchmarks on UniHG , two key measures are taken, including i) the semantic alignment strategy for multi-attribute entities, which projects the feature description of multi-attribute nodes and edges into a common embedding space to facilitate information aggregation; ii) proposing the novel Heterogeneous Graph Decoupling (HGD) framework with a specifically designed Anisotropy Feature Propagation (AFP) module for learning effective multi-hop anisotropic propagation kernels. These two strategies enable efficient information propagation among a tremendous number of multi-attribute entities and meanwhile mine multi-attribute association adaptively through the multi-hop aggregation in large-scale heterogeneous graphs. Comprehensive benchmark results demonstrate that our model significantly outperforms existing methods with an accuracy improvement of 28.93\%. And the UniHG can facilitate downstream tasks, achieving an NDCG@20 improvement rate of 11.48\% and 11.71\%. The UniHG dataset and benchmark codes have been released at https://anonymous.4open.science/r/UniHG-AA78. Yide Qiu, Tong Zhang 0021, Shaoxiang Ling, Xing Cai, Ziqi Gu, Zhen Cui 0001 |
NeurIPS | 4 |
| 2025 | CPU- and GPU-initiated Communication Strategies for Conjugate Gradient Methods on Large GPU ClustersabstractThe Conjugate Gradient (CG) method is a key building block in numerous applications, yet its low computational intensity and sensitivity to communication overhead make it difficult to scale efficiently on multi-GPU systems. In light of recent advances in multi-GPU communication technologies, we revisit CG parallelization for large-scale GPU clusters. James D. Trotter, Sinan Ekmekçibasi, Dogan Sagbili, Johannes Langguth, Xing Cai, Didem Unat |
SC | 5 |
| 2025 | Fragment-Driven Progressive Alternating Diffusion for De Novo Molecular DesignabstractHigh reliability and creativity remain key goals for AI-driven de novo molecule design. In this work, we propose a fragment-driven progressive alternating diffusion (FDPAD) framework in a coarse-to-fine generation mode. By modeling molecules as fragment-structured graphs, FDPAD entails a progressive discrete diffusion process by randomly walking some sequences of fragment-structured units (FSU), thereby mitigating combinatorial complexities and facilitating the synthesis of intricate macroscopic structures. To delve deeper internal structures of FSU, we design two distinct diffusion processes: the conditioned fragment diffusion (CFD) and the inter-fragment bond diffusion (IBD). In CFD, a string-based diffusion probability model is proposed to enrich the diversity of fragments, leveraging the partially-generated molecule as condition. And in IBD, a graph-based diffusion model upon bond-related atom graph is proposed to boost the prediction of intricate chemical bond connections among molecular fragments. Through the interleaving of CFD and IBD processes, our model outperforms state-of-the-art algorithms in de novo molecular generation, particularly in generating novel and unique molecules. Xing Cai, Tong Zhang 0021, Yide Qiu, Zhen Cui 0001 |
IEEE Trans. Comput. Biol. Bioinform. | 1 |
| 2024 | MOQEA/D: Multi-Objective QEA With Decomposition Mechanism and Excellent Global Search and Its ApplicationabstractIn this paper, a large-scale multi-objective gate assignment model is constructed by considering the flight international and domestic attributes, task type, airline affiliation, and aircraft type. Then a multi-objective quantum-inspired evolutionary algorithm based on decomposition mechanism, namely MOQEA/D is developed to solve the constructed model effectively. Specifically, a new decomposition mechanism is designed to decompose the multi-objective GAP into several single-objective sub-GAPs. Each quantum bit string solves a single-objective sub-GAP independently. And a new optimal crossover strategy is proposed to limit the randomness of observation operations and maximize the preservation of excellent genes to further improve the optimization performance. Finally, the multi-objective knapsack problem and the multi-objective GAP are selected to verify the effectiveness of the MOQEA/D. The experiment results demonstrate that the MOQEA/D can effectively solve large-scale multi-objective knapsack problem and obtain ideal gate assignment results. It takes on very significance and application value in solving complex optimization problems. Wu Deng 0001, Xing Cai, Daqing Wu, Huiling Chen 0001, Xiaojuan Ran, Xiangbing Zhou, Huimin Zhao 0002 |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2023 | Dynamic hybrid mechanism-based differential evolution algorithm and its application
Xing Cai, Xiangbing Zhou, Huiling Chen 0001, Yuangang Li 0001, Wuquan Deng, Wu Deng 0001 |
Expert Syst. Appl. | 2 |
| 2023 | Multi-strategy competitive-cooperative co-evolutionary algorithm and its application
Xiangbing Zhou, Xing Cai, Wu Deng 0001 |
Inf. Sci. | 2 |
| 2023 | Targeting performance and user-friendliness: GPU-accelerated finite element computation with automated code generation in FEniCSabstractThis paper studies the use of automated code generation to provide user-friendly GPU acceleration for solving partial differential equations (PDEs) with finite element methods. By extending the FEniCS framework and its automated compiler, we have achieved that a high-level description of finite element computations written in the Unified Form Language is auto-translated to parallelised CUDA C++ code. The auto-generated code provides GPU offloading for the finite element assembly of linear equation systems which are then solved by a GPU-supported linear algebra backend. Specifically, we explore several auto-generated optimisations of the resulting CUDA C++ code. Numerical experiments show that GPU-based linear system assembly for a typical PDE with first-order elements can benefit from using a lookup table to avoid repeatedly carrying out numerous binary searches, and that further performance gains can be obtained by assembling a sparse matrix row by row. More importantly, the extended FEniCS compiler is able to seamlessly couple the assembly and solution phases for GPU acceleration, so that all unnecessary CPU–GPU data transfers are eliminated. Detailed experiments are used to quantify the negative impact of these data transfers, which can entirely destroy the potential of GPU acceleration if the assembly and solution phases are offloaded to GPU separately. Finally, a complete, auto-generated GPU-based PDE solver for a nonlinear solid mechanics application is used to demonstrate a substantial speedup over running on dual-socket multi-core CPUs, including GPU acceleration of algebraic multigrid as the preconditioner. James D. Trotter, Johannes Langguth, Xing Cai |
Parallel Comput. | 3 |
| 2023 | Detailed Modeling of Heterogeneous and Contention-Constrained Point-to-Point MPI CommunicationabstractThe network topology of modern parallel computing systems is inherently heterogeneous, with a variety of latency and bandwidth values. Moreover, contention for the bandwidth can exist on different levels when many processes communicate with each other. Many-pair, point-to-point MP communication is thus characterized by heterogeneity and contention, even on a cluster of homogeneous multicore CPU nodes. To get a detailed understanding of the individual communication cost per MPI process, we propose a new modeling methodology that incorporates both heterogeneity and contention. First, we improve the standard max-rate model to better quantify the actually achievable bandwidth depending on the number of MPI processes in competition. Then, we make a further extension that more detailedly models the bandwidth contention when the competing MPI processes have different numbers of neighbors, with also non-uniform message sizes. Thereafter, we include more flexibility by considering interactions between intra-socket and inter-socket messaging. Through a series of experiments done on different processor architectures, we show that the new heterogeneous and contention-constrained performance models can adequately explain the individual communication cost associated with each MPI process. The largest test of realistic point-to-point MPI communication involves 8,192 processes and in total 2,744,632 simultaneous messages over 64 dual-socket AMD Epyc Rome compute nodes connected by InfiniBand, for which the overall prediction accuracy achieved is 8 4%. Andreas Thune, Sven-Arne Reinemo, Tor Skeie, Xing Cai |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2022 | DKNAS: A Practical Deep Keypoint Extraction Framework Based on Neural Architecture SearchabstractKeypoint extraction including both keypoint detection and description is a fundamental step in a wide range of geometric multimedia applications. In recent years, many learning-based approaches for keypoint extraction emerge and achieve promising results. However, they usually design network architectures empirically and lack of considerations about the comprehensive performance, which leads to limited applications. In this paper, we propose a practical framework based on Neural Architecture Search (NAS) technology, DKNAS, which can search architectures automatically and maintain efficiency and effectiveness, simultaneously. To the best of our knowledge, the proposed framework is the first NAS framework for keypoint extraction. The evaluation on HPatches dataset shows that our method achieves state-of-the-art results in the metrics of repeatability, localization error, homography accuracy and matching scores. Besides, our model is applied to a traditional Simultaneous Localization and Mapping (SLAM) system, ORB-SLAM2, to replace the handcrafted keypoints. Experimental results demonstrate that the system adopting our model outperforms ORB-SLAM2 and some other deep keypoints enhanced systems. Xing Cai, Ge Li 0002, Thomas H. Li |
ICRA | 2 |
| 2022 | On Memory Traffic and Optimisations for Low-order Finite Element Assembly Algorithms on Multi-core CPUsabstractMotivated by the wish to understand the achievable performance of finite element assembly on unstructured computational meshes, we dissect the standard cellwise assembly algorithm into four kernels, two of which are dominated by irregular memory traffic. Several optimisation schemes are studied together with associated lower and upper bounds on the estimated memory traffic volume. Apart from properly reordering the mesh entities, the two most significant optimisations include adopting a lookup table in adding element matrices or vectors to their global counterparts, and using a row-wise assembly algorithm for multi-threaded parallelisation. Rigorous benchmarking shows that, due to the various optimisations, the actual volumes of memory traffic are in many cases very close to the estimated lower bounds. These results confirm the effectiveness of the optimisations, while also providing a recipe for developing efficient software for finite element assembly. James D. Trotter, Xing Cai, Simon W. Funke |
ACM Trans. Math. Softw. | 2 |
| 2021 | iPUG for Multiple Graphcore IPUs: Optimizing Performance and Scalability of Parallel Breadth-First SearchabstractParallel graph algorithms have become one of the principal applications of high-performance computing besides numerical simulations and machine learning workloads. However, due to their highly unstructured nature, graph algorithms remain extremely challenging for most parallel systems, with large gaps between observed performance and theoretical limits. Further-more, most mainstream architectures rely heavily on single instruction multiple data (SIMD) processing for high floating-point rates, which is not beneficial for graph processing which instead requires high memory bandwidth, low memory latency, and efficient processing of unstructured data. On the other hand, we are currently observing an explosion of new hardware architectures, many of which are adapted to specific purposes and diverge from traditional designs. A notable example is the Graphcore Intelligence Processing Unit (IPU), which is developed to meet the needs of upcoming machine intelligence applications. Its design eschews the traditional cache hierarchy, relying on SRAM as its main memory instead. The result is an extremely high-bandwidth, low-latency memory at the cost of capacity. In addition, the IPU consists of a large number of independent cores, allowing for true multiple instruction multiple data (MIMD) processing. Together, these features suggest that such a processor is well suited for graph processing. We test the limits of graph processing on multiple IPUs by implementing a low-level, high-performance code for breadth-first search (BFS), following the specifications of Graph500, the most widely used benchmark for parallel graph processing. Despite the simplicity of the BFS algorithm, implementing efficient parallel codes for it has proven to be a challenging task in the past. We show that our implementation scales well on a system with 8 IPUs and attains roughly twice the performance of an equal number of NVIDIA V100 GPUs using state-of-the-art CUDA code. Luk Burchard, Xing Cai, Johannes Langguth |
HiPC | 2 |
| 2021 | An improved quantum-inspired cooperative co-evolution algorithm with muli-strategy and its application
Xing Cai, Huimin Zhao 0002, Shifan Shang, Yongquan Zhou, Wu Deng 0001, Wuquan Deng |
Expert Syst. Appl. | 1 |
| 2021 | Quantum differential evolution with cooperative coevolution framework and hybrid mutation strategy for large scale optimization
Wu Deng 0001, Shifan Shang, Xing Cai, Huimin Zhao 0002, Yongquan Zhou, Wuquan Deng |
Knowl. Based Syst. | 3 |
| 2021 | An improved differential evolution algorithm and its application in optimization problem
Wu Deng 0001, Shifan Shang, Xing Cai, Huimin Zhao 0002 |
Soft Comput. | 3 |
| 2020 | Towards Loss Balance and Consistent Model in Self-supervised Monocular Depth EstimationabstractRecently, self-supervised methods based on Convolutional Neural Networks (CNN) have achieved remarkable success in monocular depth estimation. To obtain higher quality depth maps, some of these approaches leverage traditional schemes to compute rough depth maps as proxy labels and adopt classic regression loss functions to minimize the differences between network-predicted depth maps and proxy labels. However, at proxy labels with large depth values, these methods suffer from a loss imbalance problem. To address this limitation and further improve the network performance, this article offers three key contributions. Firstly, a novel regression loss function is proposed, which can alleviate the loss imbalance problem and better handle rough proxy labels. Secondly, a dynamic mask is designed to accelerate network convergence. Thirdly, an innovative consistency loss is introduced, which can produce a more accurate and consistent model by maintaining consistency between the produced depth maps of each input image and its mirror. The effectiveness of our contributions is demonstrated by a series of ablation studies. Extensive experiments on KITTI dataset reveal that our approach achieves state-of-the-art results. Lanqing Zhang, Xing Cai, Keyao Li, Ge Li 0002, Thomas H. Li |
ICTAI | 3 |
| 2020 | VONAS: Network Design in Visual Odometry using Neural Architecture SearchabstractThe end-to-end VO (visual odometry) is a complicated task with the property of highly temporal dependency, but the design of its deep networks lacks thorough investigation. Meanwhile, NAS (Neural architecture search) has been widely searched and applied in many computer vision fields due to its advantage in automatic network design. However, most of the existing NAS frameworks only consider single image tasks such as image classification, lacking the consideration of the video (multi-frames) tasks such as VO. Therefore, this paper explores the network design for the VO task and proposes a more general single path based one-shot NAS, named VONAS, which can model sequential information for video-related tasks. Extensive experiments prove that the network architecture is significant for the (un)supervised VO. The models obtained by VONAS are lightweight and achieve SOTA performance with good generalization. Xing Cai, Lanqing Zhang, Ge Li 0002, Thomas H. Li |
ACM Multimedia | 1 |
| 2020 | Cache simulation for irregular memory traffic on multi-core CPUs: Case study on performance models for sparse matrix-vector multiplication
James D. Trotter, Johannes Langguth, Xing Cai |
J. Parallel Distributed Comput. | 3 |
| 2019 | PDNet: Prior-Model Guided Depth-Enhanced Network for Salient Object DetectionabstractFully convolutional neural networks (FCNs) have shown outstanding performance in many computer vision tasks including salient object detection. However, there still remains two issues needed to be addressed in deep learning based saliency detection. One is the lack of tremendous amount of annotated data to train a network. The other is the lack of robustness for extracting salient objects in images containing complex scenes. In this paper, we present a new architecture-PDNet, a robust prior-model guided depth-enhanced network for RGB-D salient object detection. In contrast to existing works, in which RGB-D values of image pixels are fed directly to a network, the proposed architecture is composed of a master network for processing RGB values, and a sub-network making full use of depth cues and incorporate depth-based features into the master network. To overcome the limited size of the labeled RGB-D dataset for training, we employ a large conventional RGB dataset to pre-train the master network, which proves to contribute largely to the final accuracy. Extensive evaluations over five benchmark datasets demonstrate that our proposed method performs favorably against the state-of-the-art approaches. Chunbiao Zhu, Xing Cai, Kan Huang, Thomas H. Li, Ge Li 0002 |
ICME | 2 |
| 2018 | SingleGAN: Image-to-Image Translation by a Single-Generator Network Using Multiple Generative Adversarial Learning
Xiaoming Yu, Xing Cai, Zhenqiang Ying, Thomas H. Li, Ge Li 0002 |
ACCV (5) | 2 |
| 2018 | Memory Bandwidth Contention: Communication vs Computation Tradeoffs in Supercomputers with Multicore ArchitecturesabstractWe study the problem of contention for memory bandwidth between computation and communication in supercomputers that feature multicore CPUs. The problem arises when communication and computation are overlapped and both operations compete for the same memory bandwidth. This contention is most visible at the limits of scalability, when communication and computation take similar amounts of time and thus must be taken into account in order to reach maximum scalability in memory bandwidth bound applications. Typical examples of codes affected by the memory bandwidth contention problem are sparse matrix-vector computations, graph algorithms, and many machine learning problems, as they typically exhibit a high demand for both memory bandwidth and inter-node communication, while performing a relatively low number of arithmetic operations. The problem is even more relevant in truly heterogeneous computations where CPUs and accelerators are used in concert. In that case it can lead to mispredictions of expected performance and consequently to suboptimal load balancing between CPU and accelerator, which in turn can lead to idling of powerful accelerators and thus to a large decrease in performance. We propose a simple benchmark in order to quantify the loss of performance due to memory bandwidth contention. Based on that, we derive a theoretical model to determine the impact of the phenomenon on parallel memory-bound applications. We test the model on scientific computations, discuss the practical relevance of the problem and suggest possible techniques to remedy it. Johannes Langguth, Xing Cai, Mohammed Sourouri |
ICPADS | 2 |
| 2016 | The EMC2 Project on Embedded Microcontrollers: Technical Progress after Two YearsabstractSince April 2014 the Artemis/ECSEL project EMC2 is running and provides significant results. EMC2 stands for "Embedded Multi-Core Systems for Mixed Criticality Applications in Dynamic and Changeable Real-Time Environments". In this paper we report recent progress on technical work in the different workpackages and use cases. We highlight progress in the research on system architecture, design methodology, platform and operating systems, and in qualification and certification. Application cases in the fields of automotive, avionics, health care, and industry are presented exploiting the technical results achieved. Werner Weber, Alfred Hoess, Jan van Deventer, Frank Oppenheimer, Rolf Ernst, Adam Kostrzewa, Philippe Dore, Thierry Goubier, Haris Isakovic, Norbert Druml, Egon Wuchner, Daniel Schneider 0001, Erwin Schoitsch, Eric Armengaud, Thomas Soderqvist, Massimo Traversone, Sascha Uhrig, Juan-Carlos Perez-Cortes, Sergio Sáez, Juha Kuusela, Mark van Helvoort, Xing Cai, Bjørn Nordmoen, Geir Yngve Paulsen, Hans Petter Dahle, Michael Geissel, Jürgen Salecker, Peter Tummeltshammer |
DSD | 22 |
| 2016 | Enabling Tissue-Scale Cardiac Simulations Using Heterogeneous Computing on Tianhe-2abstractWe develop a simulator for 3D tissue of the human cardiac ventricle with a physiologically realistic cell model and deploy it on the supercomputer Tianhe-2. In order to attain the full performance of the heterogeneous CPU-Xeon Phi design, we use carefully optimized codes for both devices and combine them to obtain suitable load balancing. Using a large number of nodes, we are able to perform tissue-scale simulations of the electrical activity and calcium handling in millions of cells, at a level of detail that tracks the states of trillions of ryanodine receptors. We can thus simulate arrythmogenic spiral waves and other complex arrhythmogenic patterns which arise from calcium handling deficiencies in human cardiac ventricle tissue. Due to extensive code tuning and parallelization via OpenMP, MPI, and SCIF/COI, large scale simulations of 10 heartbeats can be performed in a matter of hours. Test results indicate excellent scalability, thus paving the way for detailed whole-heart simulations in future generations of leadership class supercomputers. Johannes Langguth, Qiang Lan, Namit Gaur, Xing Cai, Mei Wen, Chunyuan Zhang |
ICPADS | 4 |
| 2015 | Towards Detailed Tissue-Scale 3D Simulations of Electrical Activity and Calcium Handling in the Human Cardiac Ventricle
Qiang Lan, Namit Gaur, Johannes Langguth, Xing Cai |
ICA3PP (3) | 4 |
| 2015 | Communication-hiding programming for clusters with multi-coprocessor nodesabstractSummary Future exascale systems are expected to adopt compute nodes that incorporate many accelerators. To shed some light on the upcoming software challenge, this paper investigates the particular topic of programming clusters that have multiple Xeon Phi coprocessors in each compute node. A new offload approach is considered for intra‐node communication, which combines Intel's APIs of coprocessor offload infrastructure (COI) and symmetric communication interface (SCIF) for achieving low latency. While the conventional pragma‐based offload approach allows simpler programming, the COI‐SCIF approach has three advantages in (1) lower overhead associated with launching offloaded code, (2) higher data transfer bandwidths, and (3) more advanced asynchrony between computation and data movement. The low‐level COI‐SCIF approach is also shown to have benefits over the MPI‐OpenMP counterpart, which belongs to the symmetric usage mode. Moreover, a hybird programming strategy based on COI‐SCIF is presented for joining the computational force of all CPUs and coprocessors, while realizing communication hiding. All the programming approaches are tested by a real‐world 3D application, for which the COI‐SCIF‐based approach shows a performance advantage on Tianhe‐2. Copyright © 2015 John Wiley & Sons, Ltd. Xinnan Dong, Mei Wen, Jun Chai, Xing Cai, Mandan Zhao, Chunyuan Zhang |
Concurr. Comput. Pract. Exp. | 4 |
| 2015 | Parallel performance modeling of irregular applications in cell-centered finite volume methods over unstructured tetrahedral meshes
Johannes Langguth, Nan Wu 0003, Jun Chai, Xing Cai |
J. Parallel Distributed Comput. | 4 |
| 2015 | An analytical GPU performance model for 3D stencil computations from the angle of data traffic
Huayou Su, Xing Cai, Mei Wen, Chunyuan Zhang |
J. Supercomput. | 2 |
| 2014 | Automated Transformation of GPU-Specific OpenCL Kernels Targeting Performance Portability on Multi-Core/Many-Core CPUs
Dafei Huang, Mei Wen, Changqing Xun, Dong Chen 0015, Xing Cai, Yuran Qiao, Nan Wu 0003, Chunyuan Zhang |
Euro-Par | 5 |
| 2014 | Utilizing Multiple Xeon Phi Coprocessors on One Compute Node
Xinnan Dong, Jun Chai, Mei Wen, Nan Wu 0003, Xing Cai, Chunyuan Zhang, Zhaoyun Chen |
ICA3PP (2) | 6 |
| 2014 | Heterogeneous CPU-GPU computing for the finite volume method on 3D unstructured meshesabstractA recent trend in modern high-performance computing environments is the introduction of accelerators such as GPU and Xeon Phi, i.e. specialized computing devices that are optimized for highly parallel applications and coexist with CPUs. In regular compute-intensive applications with predictable data access patterns, these devices often outperform traditional CPUs by far and thus relegate them to pure control functions instead of computations. For irregular applications however, the gap in relative performance can be much smaller, and sometimes even reversed. Thus, maximizing overall performance in such systems requires that full use of all available computational resources is made. In this paper we study the attainable performance of the cell-centered finite volume method on 3D unstructured tetrahedral meshes using heterogeneous systems consisting of CPUs and multiple GPUs. Finite volume methods are widely used numerical strategies for solving partial differential equations. The advantages of using finite volumes include built-in support for conservation laws and suitability for unstructured meshes. Our focus lies in demonstrating how a workload distribution that maximizes overall performance can be derived from the actual performance attained by the different computing devices in the heterogeneous environment. We also highlight the dual role of partitioning software in reordering and partitioning the input mesh, thus giving rise to a new combined approach to partitioning. Johannes Langguth, Xing Cai |
ICPADS | 2 |
| 2014 | Effective multi-GPU communication using multiple CUDA streams and threadsabstractIn the context of multiple GPUs that share the same PCIe bus, we propose a new communication scheme that leads to a more effective overlap of communication and computation. Multiple CUDA streams and OpenMP threads are adopted so that data can simultaneously be sent and received. A representative 3D stencil example is used to demonstrate the effectiveness of our scheme. We compare the performance of our new scheme with an MPI-based state-of-the-art scheme. Results show that our approach outperforms the state-of-the-art scheme, being up to 1.85× faster. However, our performance results also indicate that the current underlying PCIe bus architecture needs improvements to handle the future scenario of many GPUs per node. Mohammed Sourouri, Tor Gillberg, Scott B. Baden, Xing Cai |
ICPADS | 4 |
| 2013 | On the GPU-CPU Performance Portability of OpenCL for 3D Stencil ComputationsabstractAlthough OpenCL programming provides full code portability between different hardware platforms, performance portability can be far from satisfactory. In this work, we use a set of representative 3D stencil computations to study OpenCL's performance portability between GPUs and CPUs. For each stencil computation, we have devised different implementations of the computational kernel function, all being 100% code-portable between the two architectures. The most straightforward and compact implementation gives satisfactory CPU performance but performs poorly on GPUs, because such an implementation hampers effective use of the GPU hardware. By injecting code complexity into the involved loop nests, we can create kernel functions that still have full code portability but with increased performance portability. It is found that spatial data blocking and register reuse can be beneficial for performance on both GPUs and CPUs, whereas use of OpenCL's local memory (and subsequent temporal blocking) may only have positive effects on GPUs. Huayou Su, Nan Wu 0003, Mei Wen, Chunyuan Zhang, Xing Cai |
ICPADS | 5 |
| 2013 | Resource-efficient utilization of CPU/GPU-based heterogeneous supercomputers for Bayesian phylogenetic inference
Jun Chai, Huayou Su, Mei Wen, Xing Cai, Nan Wu 0003, Chunyuan Zhang |
J. Supercomput. | 4 |
| 2012 | Using 1000+ GPUs and 10000+ CPUs for Sedimentary Basin SimulationsabstractIn cutting-edge CPU/GPU hybrid clusters, such as Tianhe-1A, the aggregate CPU computing capability may amount to up to 1/3 of the aggregate GPU computing capability. It thus goes without saying that the CPUs and GPUs should jointly carry out the computational work. However, to effectively and simultaneously use both the hardware components requires great care when developing the parallel implementations. The challenges include (1) finding a balanced division of the workload between the CPU and GPU sides, and (2) hiding various overheads by overlapping computations with CPU-GPU data transfers and/or MPI communications. We study these issues in the context of real-world sedimentary basin simulations. Numerical experiments show that an appropriately devised CPU-GPU hybrid implementation is able to handle a global mesh resolution of 131,072*131,072, and a double-precision rate of 62 TFlops is achieved by using 1024 GPUs and 12288 CPU cores on Tianhe-1A. Such an extreme computing capability will be of great importance for carrying out high-resolution and continental-scale stratigraphic simulations in future. Mei Wen, Huayou Su, Wenjie Wei, Nan Wu 0003, Xing Cai, Chunyuan Zhang |
CLUSTER | 5 |
| 2011 | Mint: realizing CUDA performance in 3D stencil methods with annotated CabstractWe present Mint, a programming model that enables the non-expert to enjoy the performance benefits of hand coded CUDA without becoming entangled in the details. Mint targets stencil methods, which are an important class of scientific applications. We have implemented the Mint programming model with a source-to-source translator that generates optimized CUDA C from traditional C source. The translator relies on annotations to guide translation at a high level. The set of pragmas is small, and the model is compact and simple. Yet, Mint is able to deliver performance competitive with painstakingly hand-optimized CUDA. We show that, for a set of widely used stencil kernels, Mint realized 80% of the performance obtained from aggressively optimized CUDA on the 200 series NVIDIA GPUs. Our optimizations target three dimensional kernels, which present a daunting array of optimizations. Didem Unat, Xing Cai, Scott B. Baden |
ICS | 2 |
| 2010 | Numerical Analysis of a Dual-Sediment Transport Model Applied to Lake Okeechobee, FloridaabstractWe have studied two numerical strategies for solving a coupled system of dictinct nonlinear equations governing sediment transport in Lake Okeechobee. Using high-resolution bathymetry data of Lake Okeechobee, Florida, we study the numerical properties of the two strategies, from 1 core to 512 cores. The fully-explicit scheme is straightforward to implement and requires a relatively small amount of computation per time step. However, this simple numerical strategy has to use small time steps to ensure stability. These small time steps may render the explicit solver impractical for long-term and high-resolution basin simulations. As a comparison, we have implemented a semi-implicit scheme, where the two partial differential equations at each time step are solved implicitly in sequence. Numerical experiments show that this semi-implicit scheme is numerically stable even for very large time steps. Using a multicore-based cluster, we have carried out parallel simulations of sediment transport along a river chanel and into Lake Okeechobee. Stuart R. Clark, Wenjie Wei, Xing Cai |
ISPDC | 3 |
| 2004 | Using the parallel algebraic recursive multilevel solver in modern physical applications
Masha Sosonkina, Yousef Saad, Xing Cai |
Future Gener. Comput. Syst. | 3 |
| 2000 | Parallel Simulation of 3D Nonlinear Acoustic Fields on a Linux-clusterabstractSimulating the propagation of 3D ultrasonic waves in a nonlinear medium is a demanding task. It requires the solution of time-dependent and nonlinear partial differential equations (PDEs). For such nonlinear PDEs, we have to use an implicit numerical method that is very CPU-intensive. Parallel simulation is therefore essential for studying those ultrasonic waves with satisfactory accuracy. We present the parallelization of an ultrasonic wave simulator and report some numerical experiments that have been run on a 24-node Linux cluster. Our CPU-measurements indicate that Linux clusters can deliver satisfactory parallel computing power for the numerical solution of PDEs. For simulating 3D ultrasonic waves in particular; we have found that our Linux cluster; which consists of 48 Pentium-III 500 MHz processors inter-connected with a 100 Mbit/s Ethernet network, is fully comparable with an SGI Origin 2000 machine. Xing Cai, Åsmund Ødegård |
CLUSTER | 1 |