EDBT 2026 Demo / reviewers in the wild / expert
Xuebin Chi
dblp:90/2174
· DBLP profile ↗
42ranked-venue papers
1as first author
21since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 27 · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6Artificial intelligence and machine learning · 4 · 4 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | swKokkos: An Athread Backend for Enhanced Kokkos with the Sunway Heterogeneous Architecture
Junlin Wei, Jinrong Jiang, Chen Li 0068, Yehong Zhang, Yue Yu 0001, Lian Zhao, Zhenjia Li, Feng Zhang 0048, Yidi Bai, Maoxue Yu, Hailong Liu 0007, Xuebin Chi |
EuroSys | 16 |
| 2026 | TAC: Cache-Based System for Accelerating Billion-Scale GNN Training on Multi-GPU PlatformabstractGraph neural networks (GNNs) have been proven to have increasingly widespread applications in the real world. In the mainstream mini-batch training mode, multiple cache-based GNN training acceleration systems have been proposed because of the possibility of selecting the same vertex multiple times during the sampling process. However, on ultra-large scale graphs, especially those exhibiting power-law characteristics, these systems are difficult to fully utilize the distribution characteristics of cached data, which limits training performance. To this end, we propose TAC, a GNN training acceleration system that fully exploits the distribution characteristics of cached data to optimize both data transmission and computational efficiency. Specifically, we have designed a data affinity optimization algorithm that significantly enhances the locality of cache access. Secondly, an adaptive sparse matrix operator for sparsity perception is proposed, which dynamically selects the optimal computing mode based on the location of data. Finally, we have constructed a fine-grained training pipeline that maximizes system parallelism by hiding the sampling and computation. The experimental results show that TAC significantly outperforms existing state-of-the-art cache acceleration systems on multiple benchmark datasets, demonstrating higher training efficiency. Jue Wang 0013, Xingguo Shi, Junyu Gu, Peng Di, Sian Li, Chunbao Zhou, Lian Zhao, Yangang Wang 0002, Xuebin Chi |
PPoPP | 13 |
| 2026 | Review and analysis of performance prediction methods and tools for heterogeneous parallel programsabstractAbstract With the increasing number of computationally intensive applications, heterogeneous systems have become an important solution for improving computing performance. In order to effectively develop and optimize parallel programs running on these systems, performance prediction has become an indispensable part. This article aims to comprehensively review the methods and tools for predicting parallel program performance in heterogeneous systems, analyze the characteristics of existing technologies, explore their development trends, and provide valuable references and guidance for researchers and developers. This article adopts a systematic review method, first sorting out the research process of parallel program performance prediction in heterogeneous systems, and then classifying and summarizing the current mainstream performance prediction methods, including analysis model-based prediction, simulation-based prediction, and machine learning based prediction. This article also summarizes the tools and platforms used to predict parallel program performance in heterogeneous systems. Through review, it was found that various performance prediction methods and tools have their own advantages in feature richness, availability, and accuracy, but they have all improved the efficiency and accuracy of parallel program performance prediction to a certain extent. The review of this article indicates that despite various methods and tools available for performance prediction, there are still many challenges and unresolved issues. Future research should further explore more accurate, efficient and intelligent prediction methods to better support the development and optimization of parallel programs in heterogeneous systems. Beibei Gu, Lian Zhao, Chen Li 0068, Xuebin Chi |
CCF Trans. High Perform. Comput. | 4 |
| 2026 | HIP-DFPT: Scalable Optimization of Irregular Workloads in Quantum Perturbation on GPU Clusters
Meng Wan, Jue Wang 0013, Shunde Li, Honghui Shang, He Bai 0005, Peng Shi 0006, Yuchen Pang, Ying Liu 0055, Jinrong Jiang, Yangang Wang 0002, Xuebin Chi |
IEEE Trans. Parallel Distributed Syst. | 12 |
| 2025 | ParGNN: A Scalable Graph Neural Network Training Framework on multi-GPUsabstractFull-batch Graph Neural Network (GNN) training is indispensable for interdisciplinary applications. Although fullbatch training has advantages in convergence accuracy and speed, it still faces challenges such as severe load imbalance and high communication traffic overhead. In order to address these challenges, we propose ParGNN, an efficient full-batch training system for GNNs, which adopts a profiler-guided adaptive load balancing method along with graph over-partition to alleviate load imbalance. Based on the over-partition results, we present a subgraph pipeline algorithm to overlap communication and computation while maintaining the accuracy of GNN training. Extensive experiments demonstrate that ParGNN can not only obtain the highest accuracy but also reach the preset accuracy in the shortest time. In the end-to-end experiments performed on the four datasets, ParGNN outperforms the two state-of-theart full-batch GNN systems, PipeGCN and DGL, achieving the highest speedup of $2.7 \times$ and $21.8 \times$ times respectively. Junyu Gu, Shunde Li, Rongqiang Cao, Jue Wang 0013, Shigang Li 0002, Chunbao Zhou, Yangang Wang 0002, Xuebin Chi |
DAC | 11 |
| 2025 | Acc-SpMM: Accelerating General-purpose Sparse Matrix-Matrix Multiplication with GPU Tensor CoresabstractGeneral-purpose Sparse Matrix-Matrix Multiplication (SpMM) is a fundamental kernel in scientific computing and deep learning. The emergence of new matrix computation units such as Tensor Cores (TCs) brings more opportunities for SpMM acceleration. However, in order to fully unleash the power of hardware performance, systematic optimization is required. In this paper, we propose Acc-SpMM, a high-performance SpMM library on TCs, with multiple optimizations, including data-affinity-based reordering, memory efficient compressed format, high-throughput pipeline, and adaptive sparsity-aware load balancing. In contrast to the state-of-the-art SpMM kernels on various NVIDIA GPU architectures with a diverse range of benchmark matrices, Acc-SpMM achieves significant performance improvements, on average 2.52x (up to 5.11x) speedup on RTX 4090, on average 1.91x (up to 4.68x) speedup on A800, and on average 1.58x (up to 3.60x) speedup on H100 over cuSPARSE. Haisha Zhao, San Li, Chunbao Zhou, Jue Wang 0013, Zhikuang Xin, Shunde Li, Yangang Wang 0002, Xuebin Chi |
PPoPP | 13 |
| 2025 | MISA-AKMC : Achieve Kinetic Monte Carlo Simulation of 20 Quadrillion Atoms on GPU ClustersabstractThe Atomic Kinetic Monte Carlo (AKMC) method provides insights into the macroscopic behavior of materials through atomistic-level simulations and finds broad applications in materials science innovation. Improving simulation scale and performance remains a consistent focus in the development of parallel AKMC software. We port the AKMC software to GPU clusters. To alleviate the memory pressure in large-scale complex system simulations, we redesign the data layout and propose the Lattice Data Compression and Vacancy Data Decompression algorithms. Additionally, We propose a multi-level pipeline scheme combined with an on-demand communication forwarding and merging strategy to reduce data transfer and communication overhead. Compared to state-of-the-art KMC software, MISA-AKMC achieves a 10.41-fold improvement in computational throughput and a 52.07-fold expansion in simulation scale. We implement the first true micrometer-scale AKMC simulation involving 20 quadrillion atoms on GPU clusters. MISA-AKMC achieves 96.03% parallel efficiency in weak scaling and 85.29% in strong scaling on 16,000 GPUs. Shunde Li, Ningming Nie, Jue Wang 0013, He Bai 0005, Genshen Chu, Xinfu He, Yangang Wang 0002, Changjun Hu, Xuebin Chi |
SC | 11 |
| 2025 | Kilometer-Scale AI-Powered and Performance-Portable Earth System Model (AP3ESM) to Achieve Year-Scale Simulation Speed on Heterogeneous SupercomputersabstractKilometer-scale Earth system models (ESMs) necessitate exascale supercomputers to facilitate realistic simulations of weather phenomena and climate variability over a time span ranging from days to decades. We present AP3ESM, an ultra‑high‑resolution, AI‑Powered, Performance‑Portable ESM coupling atmosphere, land surface, ocean, and sea ice components. By leveraging the performance portability features of Kokkos and OpenMP, the AP3ESM operates efficiently on two heterogeneous systems while incurring minimal development overhead. Advanced optimization techniques, such as adaptive parallel algorithms, AI-enhanced physical parameterizations, and mixed-precision computations, have been implemented to further boost the computational efficiency. Breaking the 1-km resolution barrier, AP3ESM delivers 0.85 and 1.98 simulated-years-per-day (SYPD) for the standalone atmosphere and ocean components on 34.1 million Sunway cores and 16085 GPUs, respectively; the holistic AP3ESM achieves 0.54 SYPD on 37.2 million Sunway cores. Notably, the forecast experiment successfully captures Super Typhoon Doksuri in 2023 and its associated extreme rainfall across China. Maoxue Yu, Yuhu Chen, Jiaying Song, Xiaohui Duan, Junwei Wei, Jiangfeng Yu, Hailong Liu 0007, Jinrong Jiang, Yi Zhang 0127, Pengfei Lin 0004, Weipeng Zheng, Jingwei Xie, Jiakang Zhang, Zilu Liu, Xiaoyu Jin, Jilin Wei, Qixin Chang, Qingxia Lin, Yanzhi Zhou, Wei Xue 0003, Haohuan Fu, Yue Yu 0001, Xuebin Chi, Lixin Wu |
SC | 30 |
| 2025 | A parallel algorithm for an Ocean General Circulation Model based on a unified dynamics framework
Xuebin Chi, Jinrong Jiang, Run Guo, Lian Zhao, Chen Li 0068, Yidi Bai, Junlin Wei, Guangqing Zhou |
CCF Trans. High Perform. Comput. | 2 |
| 2024 | Generative Evolution Attacks Portfolio SelectionabstractIt is agreed that portfolio selection is of great importance for the financial market. Numerous outstanding exact and heuristic algorithms have been proposed in the past decades. However, their development always demands meticulous human ideas and could be time-consuming. Moreover, most of them tend to suffer from performance degradation when exposed to new portfolio selection models and different investment environments. Learning-enabled approaches have recently yielded impressive results, but these methods still grapple with challenges in model design and training. In this paper, we explore the mutual facilitation of large language models (LLMs) and huristic approaches in portfolio selection, and propose a novel LLM-based multi-objective evolutionary algorithm (MOEA) named IlmPC-NSGA-II. In this algorithm, the LLM with carefully-designed well-structured prompts serves as a straightforward yet effective engine for generating new solutions, non-dominated sorting and crowding distance calculation are adopted to enable the LLM and the evolutionary process to mutually guide toward the optimal region of the solution space. Experimental results on various scales of constrained multi-objective portfolio selection models and four benchmark problems demonstrate that our proposed approach can achieve a more competitive performance compared to widely-used MOEAs and the LLM-only method. Chen Li 0068, Jinrong Jiang, Lian Zhao, Yidi Bai, Zhonghua Lu, Xuebin Chi |
CEC | 7 |
| 2024 | Large-scale Phase-Field Simulations for Solid-Solid Phase Transformations involving Elastic EnergyabstractPhase-field models have been used extensively in studying microstructure evolution in alloys and have the superiority of comprehending, predicting, and optimizing microstructure-sensitive macroscopic material properties. The elastic strain energy is a vital factor in modeling crystal structure formation in solid-solid phase transformations. Conventionally, it is computed in the reciprocal space according to the famous Khachaturyan-Shatalov theory. In large-scale simulations, the full-space Fourier transform becomes extremely time-consuming. Yaqian Gao, Jian Zhang 0070, Huang Ye, Xuebin Chi |
ICPP | 4 |
| 2024 | POSTER: ParGNN: Efficient Training for Large-Scale Graph Neural Network on GPU ClustersabstractFull-batch graph neural network (GNN) training is essential for interdisciplinary applications. Large-scale graph data is usually divided into subgraphs and distributed across multiple compute units to train GNN. The state-of-the-art load balancing method based on direct graph partition is too rough to effectively achieve true load balancing on GPU clusters. We propose ParGNN, which employs a profiler-guided load balance workflow in conjunction with graph repartition to alleviate load imbalance and minimize communication traffic. Experiments have verified that ParGNN has the capability to scale to larger clusters. Shunde Li, Junyu Gu, Jue Wang 0013, Tiechui Yao, Yumeng Shi, Shigang Li 0002, Weiting Xi, Shushen Li, Chunbao Zhou, Yangang Wang 0002, Xuebin Chi |
PPoPP | 12 |
| 2024 | A Performance-Portable Kilometer-Scale Global Ocean Model on ORISE and New Sunway Heterogeneous SupercomputersabstractOcean general circulation models (OGCMs) are indispensable for studying the multi-scale oceanic processes and climate change. High-resolution ocean simulations require immense computational power and thus become a challenge in climate science. We present LICOMK++, a performance-portable OGCM using Kokkos, to facilitate global kilometer-scale ocean simulations. The breakthroughs include: (1) we enhance cuttingedge Kokkos with the Sunway architecture, enabling LICOMK++ to become the first performance-portable OGCM on diversified architectures, i.e., Sunway processors, CUDA/HIP-based GPUs, and ARM CPUs. (2) LICOMK++ overcomes the one simulated-years-per-day (SYPD) performance challenge for global realistic OGCM at $1-\mathrm{km}$ resolution. It records $\mathbf{1. 0 5}$ and 1.70 SYPD with a parallel efficiency of 54.8% and 55.6% scaling on almost the entire new Sunway supercomputer and two-thirds of the ORISE supercomputer. (3) LICOMK++ is the first global 1-km-resolution realistic OGCM to generate scientific results. It successfully reproduces mesoscale and submesoscale structures that have considerable climate effects. Junlin Wei, Jiangfeng Yu, Jinrong Jiang, Hailong Liu 0007, Pengfei Lin 0004, Maoxue Yu, Lian Zhao, Weipeng Zheng, Jingwei Xie, Yanzhi Zhou, Tao Zhang 0096, Feng Zhang 0048, Yehong Zhang, Yue Yu 0001, Yidi Bai, Chen Li 0068, Zipeng Yu, Xuebin Chi |
SC | 24 |
| 2024 | Accelerating LASG/IAP climate system ocean model version 3 for performance portability using Kokkos
Junlin Wei, Pengfei Lin 0004, Jinrong Jiang, Hailong Liu 0007, Lian Zhao, Yehong Zhang, Feng Zhang 0048, Youyun Li, Yue Yu 0001, Xuebin Chi |
Future Gener. Comput. Syst. | 13 |
| 2023 | A Sparse Matrix Optimization Method for Graph Neural Networks Training
Tiechui Yao, Jue Wang 0013, Junyu Gu, Yumeng Shi, Yangang Wang 0002, Xuebin Chi |
KSEM (1) | 8 |
| 2023 | ANT-MOC: Scalable Neutral Particle Transport Using 3D Method of Characteristics on Multi-GPU SystemsabstractThe Method Of Characteristic (MOC) to solve the Neutron Transport Equation (NTE) is the core of full-core simulation for reactors. High resolution is enabled by discretizing the NTE through massive tracks to traverse the 3D reactor geometry. However, the 3D full-core simulation is prohibitively expensive because of the high memory consumption and the severe load imbalance. To deal with these challenges, we develop ANT-MOC1. Specifically, we build a performance model for memory footprint, computation and communication, based on which a track management strategy is proposed to overcome the resolution bottlenecks caused by limited GPU memory. Furthermore, we implement a novel multi-level load mapping strategy to ensure load balancing among nodes, GPUs, and CUs. ANT-MOC enables a 3D full-core reactor simulation with 100 billion tracks on 16,000 GPUs, with 70.69% and 89.38% parallel efficiency for strong scalability and weak scalability, respectively. Shunde Li, Zongguo Wang, Lingkun Bu, Jue Wang 0013, Zhikuang Xin, Shigang Li 0002, Yangang Wang 0002, Yangde Feng, Peng Shi 0006, Xuebin Chi |
SC | 11 |
| 2023 | Heterogeneous acceleration algorithms for shallow cumulus convection scheme over GPU clusters
Fei Li 0042, Jinrong Jiang, He Zhang 0005, Xuebin Chi |
Future Gener. Comput. Syst. | 6 |
| 2023 | LICOM3-CUDA: a GPU version of LASG/IAP climate system ocean model version 3 based on CUDA
Junlin Wei, Jinrong Jiang, Hailong Liu 0007, Feng Zhang 0048, Pengfei Lin 0004, Yongqiang Yu, Xuebin Chi, Lian Zhao, Mengrong Ding, Zipeng Yu, Weipeng Zheng |
J. Supercomput. | 8 |
| 2022 | A Multi-level Attention-Based LSTM Network for Ultra-short-term Solar Power Forecast Using Meteorological Knowledge
Tiechui Yao, Jue Wang 0013, Haizhou Cao, Yangang Wang 0002, Xuebin Chi |
KSEM (2) | 7 |
| 2022 | Sparse Reconstruction Method for Flow Fields Based on Mode Decomposition Autoencoder
Jiyan Qiu, Wu Yuan 0002, Jian Zhang 0070, Xuebin Chi |
PRICAI (1) | 5 |
| 2022 | VenusAI: An artificial intelligence platform for scientific discovery on supercomputers
Tiechui Yao, Jue Wang 0013, Meng Wan, Zhikuang Xin, Yangang Wang 0002, Rongqiang Cao, Shigang Li 0002, Xuebin Chi |
J. Syst. Archit. | 8 |
| 2018 | An efficient parallel algorithm for the coupling of global climate models and regional climate models on a large-scale multi-core cluster
Jinrong Jiang, Junqiang Zhang, Juanxiong He, He Zhang 0005, Xuebin Chi, Tianxiang Yue |
J. Supercomput. | 6 |
| 2016 | Using Supercomputer to Speed up Neural Network TrainingabstractRecent works in deep learning have shown that large models can dramatically improve performance. In this paper, we accelerated the deep network training using many GPUs. We have developed a framework based on Caffe called Caffe-HPC that can utilize computing clusters with multiple GPUs to train large models. Caffe[6] provides multimedia scientists and practitioners with a clean and modifiable framework for state-of-the-art deep learning algorithms and a collection of reference models. And Caffe-HPC retains all the features of the original Caffe, the model trained on original Caffe can be continue to trained on Caffe-HPC. It provides a convenient solution for people who are using Caffe and want to speed up the training. Using an Asynchronous Stochastic Gradient Descent optimizer, We made a good acceleration on training a CNN model on ILSVRC[5] 2012 dataset. And we have compared the convergence of different SGD algorithms. We believe our work will makes it possible to train larger networks on larger training sets in a reasonable amount of time. Jinrong Jiang, Xuebin Chi |
ICPADS | 3 |
| 2016 | Extreme-scale phase field simulations of coarsening dynamics on the sunway taihulight supercomputerabstractMany important properties of materials such as strength, ductility, hardness and conductivity are determined by the microstructures of the material. During the formation of these microstructures, grain coarsening plays an important role. The Cahn-Hilliard equation has been applied extensively to simulate the coarsening kinetics of a two-phase microstructure. It is well accepted that the limited capabilities in conducting large scale, long time simulations constitute bottlenecks in predicting microstructure evolution based on the phase field approach. We present here a scalable time integration algorithm with large stepsizes and its efficient implementation on the Sunway TaihuLight supercomputer. The highly nonlinear and severely stiff Cahn-Hilliard equations with degenerate mobility for microstructure evolution are solved at extreme scale, demonstrating that the latest advent of high performance computing platform and the new advances in algorithm design are now offering us the possibility to simulate the coarsening dynamics accurately at unprecedented spatial and time scales. Jian Zhang 0070, Chunbao Zhou, Yangang Wang 0002, Lili Ju, Qiang Du 0001, Xuebin Chi, Dexun Chen |
SC | 6 |
| 2015 | A quantitative index for measuring the development of supercomputingabstractSummary Current assessments of supercomputing (high‐performance computing) primarily focus on system performance. Quantitative methods to measure the impact of supercomputing in a broad context have not been well developed. In this paper, the basic meaning of supercomputing development is analyzed. An evaluation index system for assessing the development status of supercomputing is constructed innovatively, and the SuperComputing Development Index (SCDI) is proposed to measure supercomputing development status. SCDI is a composite index combining various indicators into one benchmark measure that monitors and compares supercomputing development in the past years. This appears to be the first attempt to quantitatively measure the supercomputing ecosystem. As an example, the SCDI of the Chinese Academy of Sciences is obtained, which is based on the data collected from 130 research groups about and covers the period from 2006 to 2012. The results have demonstrated that the proposed evaluation index system is objectively reasonable. The constructed SCDI provides a scientific method to quantitatively evaluate the development status of supercomputing for institutions or organizations. Yonghong Hu, Xuebin Chi, Debbie Chen, David K. Kahaner, David A. Yuen |
Concurr. Comput. Pract. Exp. | 2 |
| 2014 | Visual Detection of Anomalies in DNS Query Log DataabstractDNS (Domain Name System) is an essential component of the functionality of the Internet, which converts domain names to the IP addresses. The security of DNS is related to the whole Internet. DNS query log file provide the insights of the DNS security. In this paper we propose an interactive visual analysis system for the DNS log files to intuitively detect the anomalies in DNS query logs. With a theme river based ranking visualization linked with Heat-Dial-map and tree map, user could easy identify anomalies and then further analyze regional and temporal features to help the administrators figure out the reason. Moreover, the features of DNS queries in time and region could also be analysis with this system. Guihua Shan, Yang Wang 0121, Maojin Xie, Haopu Lv, Xuebin Chi |
PacificVis | 5 |
| 2012 | Interference microscopy volume illustration for biomedical dataabstractIn this paper, we propose a novel volume illustration technique inspired by interference microscopy, which has been successfully used in biological, medical and material science over decades. Our approach simulates the optical phenomenon in interference microscopy that accounts light interference over transparent specimens, in order to generate contrast enhanced and illustrative volume visualization results. Specifically, we propose PCVR (Phase- Contrast Volume Rendering) and DICVR (Differential Interference Contrast Volume Rendering) corresponding to Phase-Contrast microscopy and Differential Interference Contrast (DIC) microscopy respectively. Without complex transfer function design, our proposed method can enhance the image contrast and structure details according to the subtle change of Optical Path Differences (OPD), and illustrate the thickness change and occluded structures with interferometry metaphors. In addition, we also develop a user interface to enable slicing specimen sections in volume data. Focus+ context lens are also included in the system for convenient data navigation and exploration. As the proposed methods are based upon widely applied microscopy techniques, they are intuitive for domain experts to explore and analyze the volume data with the proposed methods. The feedbacks from domain users suggest our proposed techniques are useful volume visualization approaches complimentary to the traditional ones. Hanqi Guo 0001, Xiaoru Yuan, Guihua Shan, Xuebin Chi |
PacificVis | 5 |
| 2012 | Parallel FDTD Simulation of Photonic Crystals and Thin-Film Solar CellsabstractFinite difference time domain (FDTD) method is a robust and accurate algorithm which is widely used in computational electromagnetic field and the simulation of optical phenomenon. In this paper, parallel FDTD based on overlapped domain decomposition is used to simulate the band gap of photonic crystals and the quantum efficiency of thin film solar cells. The light-trapping effect is also analyzed by parallel FDTD, it's very important to improve light absorption. Numerical result demonstrates that the accuracy and the speedup of parallel FDTD are very high for large scale problem. Xuebin Chi, Yangde Feng, Yonghua Zhao |
PDCAT | 2 |
| 2012 | Automating Transfer Function Design with Valley Cell-Based Clustering of 2D Density PlotsabstractAbstract Two‐dimensional transfer functions are an effective and well‐accepted tool in volume classification. The design of them mostly depends on the user's experience and thus remains a challenge. Therefore, we present an approach in this paper to automate the transfer function design based on 2D density plots. By exploiting their smoothness, we adopted the Morse theory to automatically decompose the feature space into a set of valley cells. We design a simplification process based on cell separability to eliminate cells which are mainly caused by noise in the original volume data. Boundary persistence is first introduced to measure the separability between adjacent cells and to suitably merge them. Afterward, a reasonable classification result is achieved where each cell represents a potential feature in the volume data. This classification procedure is automatic and facilitates an arbitrary number and shape of features in the feature space. The opacity of each feature is determined by its persistence and size. To further incorporate the user's prior knowledge, a hierarchical feature representation is created by successive merging of the cells. With this representation, the user is allowed to merge or split features of interest and set opacity and color freely. Experiments on various volumetric data sets demonstrate the effectiveness and usefulness of our approach in transfer function generation. Yunhai Wang, Jian Zhang 0070, Dirk J. Lehmann, Holger Theisel, Xuebin Chi |
Comput. Graph. Forum | 5 |
| 2011 | Large scale plane wave pseudopotential density functional theory calculations on GPU clustersabstractIn this work, we present our implementation of the density functional theory (DFT) plane wave pseudopotential (PWP) calculations on GPU clusters. This GPU version is developed based on a CPU DFT-PWP code: PEtot, which can calculate up to a thousand atoms on thousands of processors. Our test indicates that the GPU version can have a ~10 times speed-up over the CPU version. A detail analysis of the speed-up and the scaling on the number of CPU/GPU computing units (up to 256) are presented. The success of our speed-up relies on the adoption a hybrid reciprocal-space and band-index parallelization scheme. As far as we know, this is the first GPU DFT-PWP code scalable to large number of CPU/GPU computing units. We also outlined the future work, and what is needed to further increase the computational speed by another factor of 10. Long Wang 0008, Weile Jia, Weiguo Gao, Xuebin Chi, Lin-Wang Wang |
SC | 5 |
| 2011 | Efficient opacity specification based on feature visibilities in direct volume renderingabstractAbstract Due to 3D occlusion, the specification of proper opacities in direct volume rendering is a time‐consuming and unintuitive process. The visibility histograms introduced by Correa and Ma reflect the effect of occlusion by measuring the influence of each sample in the histogram to the rendered image. However, the visibility is defined on individual samples, while volume exploration focuses on conveying the spatial relationships between features. Moreover, the high computational cost and large memory requirement limits its application in multi‐dimensional transfer function design. In this paper, we extend visibility histograms to feature visibility, which measures the contribution of each feature in the rendered image. Compared to visibility histograms, it has two distinctive advantages for opacity specification. First, the user can directly specify the visibilities for features and the opacities are automatically generated using an optimization algorithm. Second, its calculation requires only one rendering pass with no additional memory requirement. This feature visibility based opacity specification is fast and compatible with all types of transfer function design. Furthermore, we introduce a two‐step volume exploration scheme, in which an automatic optimization is first performed to provide a clear illustration of the spatial relationship and then the user adjusts the visibilities directly to achieve the desired feature enhancement. The effectiveness of this scheme is demonstrated by experimental results on several volumetric datasets. Yunhai Wang, Jian Zhang 0070, Wei Chen 0001, Huai Zhang, Xuebin Chi |
Comput. Graph. Forum | 5 |
| 2011 | Efficient Volume Exploration Using the Gaussian Mixture ModelabstractThe multidimensional transfer function is a flexible and effective tool for exploring volume data. However, designing an appropriate transfer function is a trial-and-error process and remains a challenge. In this paper, we propose a novel volume exploration scheme that explores volumetric structures in the feature space by modeling the space using the Gaussian mixture model (GMM). Our new approach has three distinctive advantages. First, an initial feature separation can be automatically achieved through GMM estimation. Second, the calculated Gaussians can be directly mapped to a set of elliptical transfer functions (ETFs), facilitating a fast pre-integrated volume rendering process. Third, an inexperienced user can flexibly manipulate the ETFs with the assistance of a suite of simple widgets, and discover potential features with several interactions. We further extend the GMM-based exploration scheme to time-varying data sets using an incremental GMM estimation algorithm. The algorithm estimates the GMM for one time step by using itself and the GMM generated from its previous steps. Sequentially applying the incremental algorithm to all time steps in a selected time interval yields a preliminary classification for each time step. In addition, the computed ETFs can be freely adjusted. The adjustments are then automatically propagated to other time steps. In this way, coherent user-guided exploration of a given time interval is achieved. Our GPU implementation demonstrates interactive performance and good scalability. The effectiveness of our approach is verified on several data sets. Yunhai Wang, Wei Chen 0001, Jian Zhang 0070, Tingxin Dong, Guihua Shan, Xuebin Chi |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2010 | Volume exploration using ellipsoidal Gaussian transfer functionsabstractThis paper presents an interactive transfer function design tool based on ellipsoidal Gaussian transfer functions (ETFs). Our approach explores volumetric features in the statistical space by modeling the space using the Gaussian mixture model (GMM) with a small number of Gaussians to maximize the likelihood of feature separation. Instant visual feedback is possible by mapping these Gaussians to ETFs and analytically integrating these ETFs in the context of the pre-integrated volume rendering process. A suite of intuitive control widgets is designed to offer automatic transfer function generation and flexible manipulations, allowing an inexperienced user to easily explore undiscovered features with several simple interactions. Our GPU implementation demonstrates interactive performance and plausible scalability which compare favorably with existing solutions. The effectiveness of our approach has been verified on several datasets. Yunhai Wang, Wei Chen 0001, Guihua Shan, Tingxin Dong, Xuebin Chi |
PacificVis | 5 |
| 2009 | A Task-Based Fault-Tolerance Mechanism to Hierarchical Master/Worker with Divisible TasksabstractThe master/worker API of the ProActive middleware provides with an easy way to use framework for parallelizing embarrassingly parallel applications. However, the traditional master/worker model faces great challenges as the development of the scalability of the distributed computing. A single-layer hierarchical master/worker has been implemented as a solution to the scalability issues of the MW API. In the new framework, the mainmaster only communicates with some submasters, and each submaster manages a set of workers. A ldquobully election algorithmrdquo and an ldquoobject discovery mechanismrdquo are implemented to solve the fault-tolerance problems of the submasters. An automatic load-balancing mechanism is implemented for the hierarchical master/worker to solve divisible tasks. Moreover, an optimization has been done to make the fault-tolerance mechanism more efficient. Zhihui Dai, Fabien Viale, Xuebin Chi, Denis Caromel, Zhonghua Lu |
HPCC | 3 |
| 2009 | A Parallel Refined Block Arnoldi Algorithm for Large Unsymmetric MatricesabstractThis paper proposed a parallel refined block Arnoldi method for computing a few eigenvalues with largest or smallest real parts. The method accelerated by Chebyshev iteration is also investigated. We report some numerical results and compare the parallel refined block methods with single vector counterparts. The results show that the proposed method is more efficient than single vector counterparts. Xuebin Chi, Jinrong Jiang, Jun Liu 0059, Zhonghua Lu |
HPCC | 2 |
| 2009 | USGPA: A User-Centric and Secure Grid Portal Architecture for High-Performance ComputingabstractA grid portal is one of the most important ways to access grid systems. The disadvantages existing in current grid portals are that they are designed for specific applications, difficult to add new applications. Besides, the security issues are not considered fully in these portals. In this paper, we proposed a user-centric and secure grid portal architecture-USGPA based on portlet. In USGPA, a security portlet model is proposed in which security issues for sensitive data and the portal server are solved. Based on the model, the application portlet model is proposed with which new applications can be encapsulated into portal quickly. Then all portlets for high-performance computing-HPC are designed based on these two models. With these portlets, users can not only manage jobs via portal but also can custom the portal by selecting and managing the applications they required. Last but important, a prototype is implemented to evaluate the scalability, security and efficiency of the architecture. Rongqiang Cao, Xuebin Chi, Zongyan Cao, Zhihui Dai, Haili Xiao |
ISPA | 2 |
| 2008 | Parallelization and Acceleration Scheme of Multilevel Fast Multipole MethodabstractThe iterative methods such as BiCGStab for solving electromagnetic field integer equations have a complexity of O(N2), which can be reduced to O(N logN) by multilevel fast multipole method (MLFMM). For large scale problems, MLFMM should be parallelized, and the iterative convergence can be accelerated by preconditioners such as incomplete inverse triangular factorization preconditioner. The interpolation based on spherical harmonic transform at each level of MLFMMpsilas octree can be further accelerated by FFT. Based on this acceleration scheme tested on distributed cluster, the results show this algorithm is feasible. Yangde Feng, Xuebin Chi |
PDCAT | 3 |
| 2007 | Rule-Based Collaborative Volume Visualization
Yunhai Wang, Xiaoru Yuan, Guihua Shan, Xuebin Chi |
CDVE | 4 |
| 2007 | An Implementation of Parallel Eigenvalue Computation Using Dual-Level Hybrid Parallelism
Yonghua Zhao, Xuebin Chi |
ICA3PP | 2 |
| 2006 | Analysis of the Bioinformatics Grid Technique Applications in China
Ang Guo, Zhonghua Lu, Yongwei Wu 0001, Xuebin Chi |
CCGRID | 6 |
| 1998 | Parallel implementation of linear algebra problems on Dawning-1000
Xuebin Chi |
J. Comput. Sci. Technol. | 1 |
| 1997 | Parallel algorithm design on some distributed systems
Jiachang Sun, Xuebin Chi, Jianwen Cao 0001 |
J. Comput. Sci. Technol. | 2 |