Xiaohui Duan

dblp:85/5001 · DBLP profile ↗
← Back
57ranked-venue papers
9as first author
37since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 35 · 9 first-author · 25 since 2021Applied, interdisciplinary, general and emerging computing · 13 · 10 since 2021Computer networks · 5Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 since 2021Artificial intelligence and machine learning · 3 · 2 since 2021
YearPublicationVenuePosition
2026 Repurposing the Cross-Segment Space on Sunway SW26010Pro for MC-Balanced Bigshare Execution
Qixin Chang, Lifeng Yan, Hailong Liu 0007, Xiaohui Duan
Euro-Par (1)5
2026 HierCut: Enabling 16-bit Format Mixed Precision for Molecular Dynamics through Hierarchical Cutoff
abstract
Mixed-precision methods offer the potential to achieve better performance while maintaining accuracy comparable to that of high-precision formats. However, the adoption of mixed precision—particularly with 16-bit formats—in scientific computing remains limited due to precision truncation.
Lin Gan 0001, Xiaohui Duan, Zhengrui Li, Jiayu Fu, Guangzhao Li, Guangwen Yang 0002
PPoPP3
2026 Multi-sequence parotid gland lesion segmentation via expert text-guided segment anything model
Zhongyuan Wu, Chuan-Xian Ren, Xiaohua Ban, Jianning Xiao, Xiaohui Duan
Expert Syst. Appl.6
2026 Accurate full segmentation of organs-at-risk in head and neck cancer based on multimodal point cloud fusion
Pengfei Xu 0007, Jie Wang 0150, Xianyi Liu, Jinping Liu 0003, Jinxiu Li, Xiaohui Duan
Medical Image Anal.7
2026 SWGOMP: Extending OpenMP for Efficient Offloading on Sunway Heterogeneous Architecture
Qixin Chang, Xiaohui Duan, Huihai An, Yi Zhang 0127, Haohuan Fu, Bin Yang 0043, Yilun Han, Dongqiang Huang, Xiting Ju, Haopeng Huang, Wei Xue 0003, Lin Gan 0008, Maoxue Yu, Jian Li 0069, Zhao Jing, Hailong Liu 0007, Lixin Wu, Ren Hu
IEEE Trans. Parallel Distributed Syst.2
2026 Accelerating Molecular Dynamics Simulations on ARM Multi-Core Processors
abstract
LAMMPS is a widely used molecular dynamics (MD) software package in materials science, computational chemistry, and biophysics, supporting parallel computing from a single CPU core to large supercomputers. The Kunpeng processor features both high memory bandwidth and core density and is therefore an interesting candidate for accelerating compute-intensive workloads. In this paper, we target the Kunpeng multi-core architecture and focus on optimizing LAMMPS for modern ARM-based platforms by using the Lennard-Jones (L-J) and Tersoff potentials as representative case studies. We investigate both common and specific optimization challenges, and present a comprehensive performance analysis addressing four key aspects: neighbor list algorithm design, force computation optimization, efficient vectorization, and multi-thread parallelization. Experimental results show that the optimized potentials achieve speedups of approximately$2 \times$and$5 \times$, reaching$4.55 \times$and$7.04\times$the performance of the original Intel version for L-J and Tersoff, respectively. Both potentials outperform Intel's acceleration library, with a peak performance up to$2.9\times$-$3.5\times$. In terms of parallel efficiency, we evaluate scalability both within a single CPU (small-scale) and across multiple nodes (large-scale). Strong and weak scaling tests within a single CPU show that when the expansion factor is 32 times, parallel efficiency remains above$90\%$. Large-scale weak scaling across multiple nodes achieves up to$86\%$efficiency when the expansion factor is 32. Using 32 nodes (18,432 processes), our implementation enables billion-atom simulations with L-J and Tersoff potentials. This work achieves breakthrough performance and provides critical support for large-scale molecular dynamics in engineering applications.
Huihai An, Zhihua Sa, Ping Gao 0005, Xiaohui Duan, Bertil Schmidt, Yizhen Chen, Lin Gan 0001, Guangwen Yang 0002
IEEE Trans. Parallel Distributed Syst.5
2026 Exploiting the Performance Potential of Extreme-Scale Earthquake Simulation: Achieving 86.7 PFLOPS With Over 39 Million Cores
abstract
Leveraging the latest Sunway supercomputer, we developed a fully optimized earthquake simulation model that accurately captures topographic effects for realistic seismic analysis. Optimizing for the SW26010Pro architecture with DMA/RMA communication mechanisms, data compression schemes, and vectorization, we achieved a speedup exceeding 160×. Our pipeline-based computation and communication overlapping scheme, combined with performance prediction models further minimized computational costs. These optimizations enabled the largest-scale curvilinear grid finite-difference method (CGFDM) earthquake simulations to date, covering 197 trillion grid points and achieving 86.7 PFLOPS on 39 million cores with a weak scaling efficiency of 97.9%. These advancements enabled the successful simulation of the 2008 Wenchuan earthquake, providing high-resolution seismic insights and robust assessments for regional hazard mitigation and disaster preparedness.
Lin Gan 0008, Wubing Wan, Zekun Yin, Zhong He, Ping Gao 0005, Xiaohui Duan, Wei Xue 0003, Haohuan Fu, Guangwen Yang 0002
IEEE Trans. Parallel Distributed Syst.9
2025 RabbitTClust2: Fast, Scalable, and Versatile Clustering for Massive Genomic Datasets
abstract
Clustering is a fundamental method for extracting meaningful information from large-scale genomic datasets. As sequencing technologies advance, efficient and scalable clustering tools have become increasingly important. Despite its outstanding efficiency in large-scale genome clustering tasks, RabbitTClust still faces certain limitations. On the one hand, as data volumes continue to grow, there remains room for further optimization of its computational performance. On the other hand, RabbitTClust is not well suited for frequent incremental data updates or fast clustering across multiple thresholds. To address these limitations, we introduce RabbitTClust2, a highly efficient and versatile tool designed for clustering large-scale genomic sequences. RabbitTClust2 integrates an efficient sketching algorithm, a pruningand inverted-index-based minimum spanning tree construction method, and strategies for reusing intermediate results. With these advancements, RabbitTClust2 is able to cluster the latest RefSeq bacterial dataset (195 k genomes, 820 GB in FASTA) within 5 minutes. Compared to previous versions, RabbitTClust2 achieves a$2.4 \times$to$4.5 \times$speedup while maintaining comparable clustering accuracy, with a 21 % reduction in memory consumption. On a distributed multi-node platform, RabbitTClust2 is capable of clustering 2.6 million genomes in approximately one hour. Furthermore, RabbitTClust2 offers significant versatility by supporting efficient incremental clustering and rapid multithreshold analysis. RabbitTClust2 utilizes incremental clustering to integrate 1,000 new sequences into a pre-clustered dataset of 194,000 genomes within 1 minute, a process that results in a$22.7 \times$speedup over RabbitTClust. In addition, we used RabbitTClust2 to generate a series of clustering results with Mash distance thresholds ranging from 0.01 to 0.2 (a total of 20 values) within 7 minutes on the RefSeq bacterial dataset. The results showed that when the clustering threshold approached 0.1, the cluster compositions changed significantly, suggesting that 0.1 may represent a critical threshold for genuslevel classification in bacteria. RabbitTClust2 is available at https://github.com/RabbitBio/RabbitTClust.
Xiaoming Xu 0004, Zekun Yin, Lifeng Yan, Yijie Gao, Xiaohui Duan, Bertil Schmidt
BIBM7
2025 SWBWA: A Highly Efficient NGS Aligner on the New Sunway Architecture
Lifeng Yan, Zekun Yin, Qixin Chang, Zhisong Wang, Xiaohui Duan, Bertil Schmidt
Euro-Par (3)6
2025 HARM3-Fusion: Hierarchical Attentional Representation Learning of Multi-modal, Multi-temporal, and Multi-sequence Fusion for Pathological Complete Response Prediction of Head and Neck Squamous Cell Carcinoma
Jianye Wang, Zhiying Gong, Lingjie Yang, Yimeng Fan, Xiaohui Duan, Weibing Zhao
MICCAI (10)9
2025 Cervical-RG: Automated Cervical Cancer Report Generation from 3D Multi-sequence MRI via CoT-Guided Hierarchical Experts
Yimeng Fan, Zhaoyi Zhan, Zheng Xing 0001, Xiaohui Duan, Weibing Zhao
MICCAI (5)11
2025 An AI-Enhanced 1km-Resolution Seamless Global Weather and Climate Model to Achieve Year-Scale Simulation Speed using 34 Million Cores
abstract
Global Storm Resolving Models (GSRMs) is crucial for understanding extreme weather events under the climate change background. In this study, we optimize Global-Regional Integrated Forecast System (GRIST), which is a unified weather-climate modeling system designed for research and operation, for the next-generation Sunway supercomputer, incorporating AI-enhanced physics suite, OpenMP-based parallelization, and mixed-precision optimizations to enhance both efficiency and performance portability, as well as the unified modeling capability. Our experiments successfully capture significant events during the "23.7" extreme rainfall over northern China influenced by super Typhoon Doksuri, at 1km resolution. Notably, our work scales to 34 million cores, enabling simulation speeds at 491 SDPD (3km) and 181 SDPD (1km).
Xiaohui Duan, Yi Zhang 0127, Haohuan Fu, Bin Yang 0043, Yilun Han, Dongqiang Huang, Huihai An, Xiting Ju, Haopeng Huang, Wei Xue 0003, Jianye Hou, Maoxue Yu, Jian Li 0069, Zhao Jing, Hailong Liu 0007, Lixin Wu
PPoPP1
2025 Trillion Ligands per Day: Performance-Portable Virtual Screening via Compound Database Optimization and Multi-Target Docking
abstract
Structure-based virtual screening confronts a grand challenge in scaling to trillion-ligand libraries for drug discovery. We present SWDOCKP2, a performance-portable virtual screening framework achieving 1.9 trillion ligand-receptor pairs daily across eight targets on the Sunway OceanLight supercomputer with 39-million cores — 10× faster than prior state-of-the-art. Key innovations combine (1) a ligand database optimizer with conformational sorting and merging, (2) multi-receptor grid alignment enabling parallel target screening and SIMD-accelerated trilinear interpolation, and (3) a Sunway architecture emulator for cross-platform efficiency. These advancements bridge computational scalability with novel drug discovery demands, offering a blueprint for next-generation supercomputing in structure-based drug design. Additionally, SWDOCKP2 will generate an unprecedented dataset of predicted protein-ligand interactions, creating a transformative resource for machine learning applications. By addressing experimental data scarcity, this dataset empowers accurate ligand prediction, generative chemistry, and AI-driven drug discovery.
Xiaohui Duan, Gaowei Chen, Yizhen Chen, Qixin Chang, Qiancheng Xia, Zekun Yin, Lin Gan 0001, Yibing Shan, Guangwen Yang 0002, Niu Huang
SC1
2025 T2-RELION: Task Parallelism, Tensor Core Accelerated RELION for Cryo-EM 3D Reconstruction
abstract
Cryo-electron microscopy (cryo-EM) is a key technique for structural biology, but its computational efficiency, particularly during 3D reconstruction, remains a bottleneck. We introduce T2-RELION, a highly optimized version of RELION for cryo-EM 3D reconstruction on CPU-GPU platforms. RELION is a widely used open-source package in the cryo-EM community. We identify and resolve key inefficiencies in RELION’s parallelization strategy and memory management by proposing task parallelism and a three-phase GPU memory management strategy. Furthermore, we leverage Tensor Cores to accelerate the hot-spot kernel for difference calculation, employing an advanced pipelining strategy to hide latency and enable thread-block-level data reuse. On a quad-A100 GPU machine, performance evaluations demonstrate that T2-RELION outperforms RELION 4.0. For the hot-spot kernel, our optimizations achieve 1.90-23.7 times speedup. For the whole application using CNG and Trpv1 datasets, we observe 3.86 times and 2.68 times speedups, respectively.
Jiayu Fu, Jingle Xu, Lin Gan 0001, Tianqi Mao 0003, Zirong Shen, Xiaohui Duan, Wei Xue 0003, Guangwen Yang 0002
SC8
2025 Kilometer-Scale AI-Powered and Performance-Portable Earth System Model (AP3ESM) to Achieve Year-Scale Simulation Speed on Heterogeneous Supercomputers
abstract
Kilometer-scale Earth system models (ESMs) necessitate exascale supercomputers to facilitate realistic simulations of weather phenomena and climate variability over a time span ranging from days to decades. We present AP3ESM, an ultra‑high‑resolution, AI‑Powered, Performance‑Portable ESM coupling atmosphere, land surface, ocean, and sea ice components. By leveraging the performance portability features of Kokkos and OpenMP, the AP3ESM operates efficiently on two heterogeneous systems while incurring minimal development overhead. Advanced optimization techniques, such as adaptive parallel algorithms, AI-enhanced physical parameterizations, and mixed-precision computations, have been implemented to further boost the computational efficiency. Breaking the 1-km resolution barrier, AP3ESM delivers 0.85 and 1.98 simulated-years-per-day (SYPD) for the standalone atmosphere and ocean components on 34.1 million Sunway cores and 16085 GPUs, respectively; the holistic AP3ESM achieves 0.54 SYPD on 37.2 million Sunway cores. Notably, the forecast experiment successfully captures Super Typhoon Doksuri in 2023 and its associated extreme rainfall across China.
Maoxue Yu, Yuhu Chen, Jiaying Song, Xiaohui Duan, Junwei Wei, Jiangfeng Yu, Hailong Liu 0007, Jinrong Jiang, Yi Zhang 0127, Pengfei Lin 0004, Weipeng Zheng, Jingwei Xie, Jiakang Zhang, Zilu Liu, Xiaoyu Jin, Jilin Wei, Qixin Chang, Qingxia Lin, Yanzhi Zhou, Wei Xue 0003, Haohuan Fu, Yue Yu 0001, Xuebin Chi, Lixin Wu
SC7
2025 RabbitSketch: a high-performance sketching library for genome analysis
abstract
SUMMARY: We present RabbitSketch, a highly optimized library of sketching algorithms such as MinHash, OrderMinHash, and HyperLogLog that can exploit the power of modern multi-core CPUs. It provides significant speedups compared to existing implementations, ranging from 2.30× to 49.55×, as well as flexible and easy-to-use interfaces for both Python and C++. As a result, the similarity analysis of 455GB genomic data can be completed in only 5 minutes using RabbitSketch with merely 20 lines of Python code. As a case study, we enhanced RabbitTClust by integrating RabbitSketch's Kssd algorithm, resulting in a 1.54× speedup with no loss in accuracy. AVAILABILITY AND IMPLEMENTATION: RabbitSketch is available at https://github.com/RabbitBio/RabbitSketch with an archived version at Zenodo: https://doi.org/10.5281/zenodo.14903962. Detailed API documentation is available at https://rabbitsketch.readthedocs.io/en/latest.
Zekun Yin, Xiaoming Xu 0004, Lifeng Yan, Fangjin Zhu, Xiaohui Duan, Bertil Schmidt
Bioinform.6
2025 SWQC: Efficient sequencing data quality control on the next-generation sunway platform
Lifeng Yan, Zekun Yin, Fangjin Zhu, Xiaohui Duan, Bertil Schmidt
Future Gener. Comput. Syst.5
2025 RabbitTrim: An Efficient and Versatile Trimmer on Multi-Core Platforms
abstract
Trimming is an essential step in sequencing data processing. However, many existing trimming tools, such as Trimmomatic and Ktrim, are limited by suboptimal implementations and fail to fully leverage the computational power of modern multi-core platforms. To address this, we introduce RabbitTrim, a highly optimized and versatile trimming tool that fully supports the functionalities of Trimmomatic and Ktrim. RabbitTrim's performance is enhanced through efficient I/O strategies, parallel (de)compression engines, block-based memory pools, bitwise operations, and vectorization techniques. Compared to Trimmomatic, RabbitTrim (in trimmomatic mode) achieves speedups ranging from 1.8x to 6.0x for plain FASTQ files and 3.7x to 14.0x for gzip-compressed FASTQ files on a 48-core Intel server. Similarly, compared to Ktrim, RabbitTrim (in ktrim mode) achieves speedups ranging from 1.5x to 2.5x for plain FASTQ files and 2.7x to 5.6x for gzip-compressed FASTQ files on the same server. Moreover, RabbitTrim is able to process 101 GB gzip-compressed sequencing data in only 5 minutes while Trimmomatic requires at least 21 minutes.
Zekun Yin, Lifeng Yan, Fangjin Zhu, Xin Li 0137, Xiaohui Duan, Bertil Schmidt
IEEE Trans. Comput. Biol. Bioinform.8
2025 RabbitBAM: Accelerating BAM File Manipulation on Multi-Core Platforms
abstract
With the continuous advancement of sequencing technology, the scale of biological data has rapidly increased. BAM format, widely used for storing aligned sequence data, is very popular due to its ease of use and good compression ratio. However, existing BAM-format file I/O libraries often fail to fully leverage the computational power of modern multi-core platforms, resulting in low CPU utilization. To address this, we introduce RabbitBAM, a fast BAM-format file I/O library. RabbitBAM employs pre-parsing and parallel parsing techniques to eliminate parsing bottlenecks and improve parallel efficiency. Additionally, we optimize multi-threaded data handling through the use of dedicated lock-free queues and memory pools. RabbitBAM achieves 2.1-3.3x speedups on next-generation sequencing data and 1-2.2x speedups on third-generation sequencing data compared to state-of-the-art SAMtools (HTSlib). We also present two case studies (BAM file quality control and sorting) using RabbitBAM, demonstrating 1.4-2.4x speedups compared to other implementations.
Lifeng Yan, Zhan Zhao, Zekun Yin, Fangjin Zhu, Xiaohui Duan, Bertil Schmidt
IEEE Trans. Comput. Biol. Bioinform.7
2024 A Single-Stage Multi-Style License Plate Recognition Method Based on Attention
Longbin Wu, Yufei Xie, Haowei Lee, Xiaohui Duan
ACML5
2024 A4-Unet: Deformable Multi-Scale Attention Network for Brain Tumor Segmentation
abstract
Brain tumor segmentation models have aided diagnosis in recent years. However, they face MRI complexity and variability challenges, including irregular shapes and unclear boundaries, leading to noise, misclassification, and incomplete segmentation, thereby limiting accuracy. To address these issues, we adhere to an outstanding Convolutional Neural Networks (CNNs) design paradigm and propose a novel network named A4-Unet. In A4-Unet, Deformable Large Kernel Attention (DLKA) is incorporated in the encoder, allowing for improved capture of multi-scale tumors. Swin Spatial Pyramid Pooling (SSPP) with cross-channel attention is employed in a bottleneck further to study long-distance dependencies within images and channel relationships. To enhance accuracy, a Combined Attention Module (CAM) with Discrete Cosine Transform (DCT) orthogonality for channel weighting and convolutional element-wise multiplication is introduced for spatial weighting in the decoder. Attention gates (AG) are added in the skip connection to highlight the foreground while suppressing irrelevant background information. The proposed network is evaluated on three authoritative MRI brain tumor benchmarks and a proprietary dataset, and it achieves a 94.4% Dice score on the BraTS 2020 dataset, thereby establishing multiple new state-of-the-art benchmarks. The code is available here: https://github.com/WendyWAAAAANG/A4-Unet.
Ruoxin Wang, Haiming Du, Yuxuan Cheng, Lingjie Yang, Xiaohui Duan, Yunfang Yu, Yu Zhou 0027, Donald Donglong Chen
BIBM7
2024 Enabling High-Performance Physical Based Rendering on New Sunway Supercomputer
abstract
Physical based rendering is widely applied in diverse fields requiring realistic scene visualization. This paper outlines our efforts in implementing a high-performance and highly scalable physical based rendering framework on the next-generation Sunway supercomputer based on PBRT. To effectively tailor the rendering application to the hardware attributes of the state-of-the-art architecture, we primarily carried out four approaches, 1) a memory management strategy, 2) a solution for runtime polymorphism, 3) measures to mitigate instruction cache misses, and 4) a two level load balancing strategy. Our design can achieve at most 41.95x speedup relative to baseline implementation. By using 32 Sunway processors, we can achieve at most 27.97x speedup relative to a 40-core 5218R CPU and at most 6.90x speedup relative to a RTX 3090 GPU. Nearly linear scalabilities are obtained when scaling up to 2,048 Sunway processors.
Lin Gan 0001, Shengye Xiang, Xiaohui Duan, Guangwen Yang 0002
IPDPS5
2024 RabbitTrim: Highly Optimized Trimming of Illumina Sequencing Data on Multi-core Platforms
Zekun Yin, Lifeng Yan, Fangjin Zhu, Xiaohui Duan, Xin Li 0137, Bertil Schmidt
ISBRA (2)6
2024 RabbitSAlign: Accelerating Short-Read Alignment for CPU-GPU Heterogeneous Platforms
Lifeng Yan, Zekun Yin, Fangjin Zhu, Xiaohui Duan, Bertil Schmidt
ISBRA (2)7
2024 O2ath: an OpenMP offloading toolkit for the sunway heterogeneous manycore platform
Lifeng Yan, Qixin Chang, Haitian Lu, Chenlin Li, Quanjie He, Xiaohui Duan, Zekun Yin, Wei Xue 0003, Haohuan Fu, Lin Gan 0001, Guangwen Yang 0002
CCF Trans. High Perform. Comput.8
2024 Acceleration of Multi-Body Molecular Dynamics With Customized Parallel Dataflow
abstract
FPGAs are drawing increasing attention in resolving molecular dynamics (MD) problems, and have already been applied in problems such as two-body potentials, force fields composed of these potentials, etc. Competitive performance is obtained compared with traditional counterparts such as CPUs and GPUs. However, as far as we know, FPGA solutions for more complex and real-world MD problems, such as multi-body potentials, are seldom to be seen. This work explores the prospects of state-of-the-art FPGAs in accelerating multi-body potential. An FPGA-based accelerator with customized parallel dataflow that features multi-body potential computation, motion update, and internode communication is designed. Major contributions include: (1) parallelization applied at different levels of the accelerator; (2) an optimized dataflow mixing atom-level pipeline and cell-level pipeline to achieve high throughput; (3) a mixed-precision method using different precision at different stages of simulations; and (4) a communication-efficient method for internode communication. Experiments show that, our single-node accelerator is over 2.7× faster than an 8-core CPU design, performing 20.501 ns/day on a 55,296-atom system for theTersoffsimulation. Regarding power efficiency, our accelerator is 28.9× higher than I7-11700 and 4.8× higher than RTX 3090 when running the same test case.
Quan Deng 0001, Qiang Liu 0011, Xiaohui Duan, Lin Gan 0008, Jinzhe Yang, Wenlai Zhao, Zhenxiang Zhang, Guiming Wu, Wayne Luk, Haohuan Fu, Guangwen Yang 0002
IEEE Trans. Parallel Distributed Syst.4
2023 Enabling Real World Scale Structural Superlubricity All-Atom Simulation on the Next-Generation Sunway Supercomputer
abstract
Molecular dynamics (MD) simulation can provide an affordable way for inspecting microscopic phenomena, which is a powerful complement to real-world experiments. But the spatial scale of MD simulations is usually magnitudes smaller than experiment systems. In this paper, we present our work, redesigning the widely used inter-layer potential in structural superlubricity. By carrying out a specialized neighbor list for inter-layer potential computation, the total memory access amount is reduced significantly. Besides, a simple but efficient vectorization strategy is implemented based on the new neighbor list. In the extreme case, our work can scale to 38 million cores to achieve a sustainable performance of 61 PFLOPS, enabling a simulation of a superlubricity system of 32 μm2 with 7.2 billion atoms at 4.75 ns/day, which is 11,834 times of reported largest scale simulation in superlubricity systems in contact area and almost ten times faster in time-to-solution. Furthermore, we have done a simulation at 9 μm2 which results in consistency with real-world experiments and verified some theoretical predictions in the mesoscopic scale.
Xiaohui Duan, Ping Gao 0005, Ming Ma 0012, Lin Gan 0001, Xin Liu 0081, Haohuan Fu, Wei Xue 0003, Dexun Chen, Guangwen Yang 0002
SC1
2023 69.7-PFlops Extreme Scale Earthquake Simulation with Crossing Multi-faults and Topography on Sunway
abstract
A high-scalable and fully optimized earthquake model is presented based on the latest Sunway supercomputer. Contributions include: 1) the curvilinear grid finite-difference method (CGFDM) and flexible model applying perfectly matched layer (PML) and enabling more accurate and realistic terrain descriptions; 2) a hybrid and non-uniform domain decomposition scheme that efficiently maps the model across different levels of the computing system; and 3) sophisticated optimizations that largely alleviate or even eliminate bottlenecks in memory, communication, etc., obtaining a speedup of over 140×. Combining all innovations, the design fully exploits the hardware potential of all aspects and enables us to perform the largest CGFDM-based earthquake simulation ever reported (69.7 PFlops using over 39 million cores). Based on our design, the Turkey earthquakes (February 6, 2023), and the Ridgecrest earthquake (July 4, 2019), are successfully simulated with a maximum resolution of 12-m. Precise hazard evaluations for the hazardous reduction of earthquake-stricken areas are also conducted.
Wubing Wan, Lin Gan 0001, Zekun Yin, Haodong Tian, Mengyuan Hua, Shengye Xiang, Zhongqiu He, Ping Gao 0005, Xiaohui Duan, Wei Xue 0003, Haohuan Fu, Guangwen Yang 0002, Yaojian Chen, Xin Liu 0081, Wei Zhang 0321
SC14
2023 Bio-ESMD: A Data Centric Implementation for Large-Scale Biological System Simulation on Sunway TaihuLight Supercomputer
abstract
Molecular dynamics (MD) simulations of biological systems are playing an increasingly important role in the research of pathogens and drugs. Most MD methods for biological simulations rely on the listed bonds which interact among specific groups of atoms identified by atom tags (unique atom tags regardless the storage location). However, efficient mapping of tags to atom locations is often challenging on modern many-core processors because data locality can not always be guaranteed for large-scale systems. In this paper, we present Bio-ESMD, a new MD implementation supporting listed bonds. Bio-ESMD is designed and developed based on our previously designed ESMD framework for many-core processors. In Bio-ESMD, we have introduced a data-centric approach for refactoring MD algorithms by reorganizing the cell list data structure to adopt bond lists with guaranteed data locality. Our implementation achieves speedups of over two compared to SW_GROMACS on Sunway TaihuLight. Furthermore, Bio-ESMD can simulate a system of 308.8 million atoms at 1.33 ns/day or 14.44 million atoms at 17.28 ns/day with linear weak scaling efficiency.
Xiaohui Duan, Junben Weng, Bertil Schmidt, Lin Gan 0001, Haohuan Fu, Wei Xue 0003, Guangwen Yang 0002
IEEE Trans. Parallel Distributed Syst.1
2023 Redesign and Accelerate the AIREBO Bond-Order Potential on the New Sunway Supercomputer
abstract
Molecular dynamics (MD) is one of the most crucial computer simulation methods for understanding real-world processes at the atomic level. Reactive potentials based on the bond order concept have the ability to model dynamic bond breaking and formation with close to quantum mechanical (QM) precision without actually requiring expensive QM calculations. In this article, we focus on the adaptive intermolecular reactive empirical bond-order (AIREBO) potential in LAMMPS for the simulation of carbon and hydrocarbon systems on the new Sunway supercomputer. To achieve scalable performance, we propose a parallel two-level building scheme and periodic buffering strategy for the tailored data design to explore data locality and data reuse. Furthermore, we design two optimized nearest-neighbor access algorithms: the redistribution of accumulated coefficients algorithm and the double-end search connectivity algorithm. Finally, we implement parallel force computation with an AoS data layout and hardware/software co-cache. In addition, we have designed a low-overhead atomic operation-based load balancing method and vectorization. The overall performance of AIREBO achieves a speedup of nearly$20\times$on a single core group (CG), and more than$5\times$and$4\times$over an Intel Xeon E5 2680 v3 core and an Intel Xeon Gold 6138 core, respectively. Compared with the Intel accelerator package in LAMMPS, our performance further achieves$3.0\times$of an Intel Xeon E5 2680 v3 core and is better than that of an Intel Xeon Gold 6138 core. We complete the validation of the results in no more than 20.5 hours on a single node with 2,000,000 running steps (i.e., 1 ns). Our experiments show that the simulation of 2,139,095,040 atoms on 798,720 ((1MPE+64CPEs) × 12,288 processes) cores exhibits a parallel efficiency of 88% under weak scaling.
Ping Gao 0005, Xiaohui Duan, Bertil Schmidt, Wubing Wan, Jiaxu Guo, Wusheng Zhang, Lin Gan 0008, Haohuan Fu, Wei Xue 0003, Guangwen Yang 0002
IEEE Trans. Parallel Distributed Syst.2
2022 FMapper: Scalable read mapper based on succinct hash index on SunWay TaihuLight
Xiaohui Duan, André Müller, Robin Kobus, Bertil Schmidt
J. Parallel Distributed Comput.2
2022 Enabling Large-Scale Simulation of CAM on the Sunway TaihuLight Supercomputer
abstract
The Community Atmosphere Model (CAM) has been ported, redesigned, and scaled to the full system of the Sunway TaihuLight, and provides peta-scale climate modeling performance. Based on a novel domain decomposition method, we have fully optimized the complete model code by using both OpenACC refactoring and more aggressive and finer-grained Athread approaches. The Athread approach enables us to achieve exceptional memory control and usage, efficient vectorization, and sophisticated utilization of the thread-level communication mechanism. We have also further refined the load-balance behaviors towards ultra-large-scale numerical simulation. By combining all these novelties, we achieved a simulation speed of 7.2 and 25.6 simulation-year-per-day (SYPD) for global 25-km and 100-km resolution, respectively (1.2- to 2.2-fold improvements over previous efforts), and a sustainable double-precision performance of 3.3 PFlops for a 750-m global simulation when using 10075000 cores.
Xiaohui Duan, Lin Gan 0001, Wubing Wan, Yuhu Chen, Jinzhe Yang, Wei Xue 0003, Haohuan Fu, Guangwen Yang 0002
IEEE Trans. Computers2
2022 Optimization of Reactive Force Field Simulation: Refactor, Parallelization, and Vectorization for Interactions
abstract
Molecular dynamics (MD) simulations are playing an increasingly important role in many areas ranging from chemical materials to biological molecules. With the continuing development of MD models, the potentials are getting larger and more complex. In this article, we focus on the reactive force field (ReaxFF) potential from LAMMPS to optimize the computation of interactions. We present our efforts on refactoring for neighbor list building, bond order computation, as well as valence angles and torsion angles computation. After redesigning these kernels, we develop a vectorized implementation for non-bonded interactions, which is nearly 100 × faster than the management processing element (MPE) on the Sunway TaihuLight supercomputer. Furthermore, we have implemented the three-body-list free torsion angles computation, and propose a line-locked software cache method to eliminate write conflicts in the torsion angle and valence angle interactions resulting in an order-of-magnitude speedup on a single Sunway TaihuLight node. In addition, we achieve a speedup of up to 3.5 compared to the KOKKOS package on an Intel Xeon Gold 6148 core. When executed on 1,024 processes, our implementation enables the simulation of 21,233,664 atoms on 66,560 cores with a performance of 0.032 ns/day and a weak scaling efficiency of 95.71 percent.
Ping Gao 0005, Xiaohui Duan, Bertil Schmidt, Wusheng Zhang, Lin Gan 0001, Haohuan Fu, Wei Xue 0003, Guangwen Yang 0002
IEEE Trans. Parallel Distributed Syst.2
2022 Scaling Poisson Solvers on Many Cores via MMEwald
abstract
The Poisson solver for the calculation of the electrostatic potential is an essential primitive in quantum mechanics calculations. In this article, we adopt the Ewald method and propose a highly-optimized and scalable framework for Poisson solver, MMEwald, on the new generation Sunway supercomputer, capable of utilizing the collection of 390-core accelerators it uses. The MMEwald is based on a grid adapted cut-plane approach to partition the points into batches and distribute the batch to the processors. Furthermore, we propose a set of architecture-specific optimizations to efficiently utilize the memory bandwidth and computation capacity of the supercomputer. Experimental results demonstrate the efficiency of the MMEwald in providing strong and weak scaling performance.
Mingchuan Wu, Yangjun Wu, Honghui Shang, Ying Liu 0055, Huimin Cui, Xiaohui Duan, Yunquan Zhang, Xiaobing Feng 0002
IEEE Trans. Parallel Distributed Syst.7
2022 Redesigning and Optimizing UCSF DOCK3.7 on Sunway TaihuLight
abstract
Molecular docking is the process of posing, scoring, and ranking small molecules at the binding sites of proteins to prioritize compounds for experimental testing. It is a widely-used computational method in the drug discovery process. However, it is a highly time-consuming procedure since a receptor may need to find favorable ligand orientations in billions of ligands. UCSF DOCK3.7 is one of the most widely used molecular docking applications. In this paper, we port and optimize UCSF DOCK3.7 on the Sunway TaihuLight supercomputer. To avoid the impact of load imbalance, we employ a producer-consumer strategy that can overlap I/O and computation in order to achieve high performance. Furthermore, we present a new binary file format to replace the mol2db2 file format for ligand storage and adopt xzip rather than gzip to compress ligand files. We show that our file format can reduce I/O time significantly while xzip saves significant storage. For the routines which determine the orientation of a ligand relative to the receptor, we present an improved algorithm to discard geometrically similar orientations. Furthermore, we fuse loops and compress memory usage to store data in fast Local Device Memory (LDM) in order to score ligand orientations with high efficiency. In addition, we propose a number of architecture-specific optimizations. Asynchronous data transfer and vectorization of computation are implemented to take full advantage of the SW26010 processor. Our experiments show that a speedup of 167 can be achieved by using the proposed strategies. Compared to a core of an Intel(R) Core(TM) i9-10900K CPU, our approach achieves speedups of 15 on a SW26010 core group. Furthermore, our implementation achieves strong scalability to hundreds of thousands of heterogeneous cores on the next-generation Sunway supercomputer.
Jinxiao Zhang, Xiaohui Duan, Xiaobo Wan, Niu Huang, Bertil Schmidt, Guangwen Yang 0002
IEEE Trans. Parallel Distributed Syst.3
2021 LMFF: efficient and scalable layered materials force field on heterogeneous many-core processors
abstract
LAMMPS is one of the most popular Molecular Dynamic (MD) packages and is widely used in the field of physics, chemistry and materials simulation. Layered Materials Force Field (LMFF) is our expansion of the LAMMPS potential function based on the Tersoff potential and inter-layer potential (ILP) in LAMMPS. LMFF is designed to study layered materials such as graphene and boron hexanitride. It is universal and does not depend on any platform. We have also carried out a series of optimizations on LMFF and the optimization work is carried out on the new generation of Sunway supercomputer, called SWLMFF. Experiments show that our implementation is efficient, scalable and portable. When generic LMFF is ported to Intel Xeon Gold 6278C, 2X performance improvement is achieved. For the optimized SWLMFF, the overall performance improvement is nearly 200--330X compared to the original ILP and Tersoff potentials. And SWLMFF has good parallel efficiency of 95%-100% under weak scaling with 2.7 million atoms on a single process. The maximum atomic system simulated by SWLMFF is close to 231 atoms. And nanosecond simulations in one day can be realized.
Ping Gao 0005, Xiaohui Duan, Jiaxu Guo, Zhenya Song, Li-Zhen Cui 0001, Xiangxu Meng, Xin Liu 0081, Wusheng Zhang, Ming Ma 0012, Dexun Chen, Haohuan Fu, Wei Xue 0003, Guangwen Yang 0002
SC2
2021 Extreme-scale ab initio quantum raman spectra simulations on the leadership HPC system in China
abstract
Raman spectroscopy provides chemical and compositional information that can serve as a structural fingerprint for various materials. Therefore, simulations of Raman spectra, including both quantum perturbation analyses and ground-state calculations, are of significant interest. However, highly accurate full quantum mechanical (QM) simulations of Raman spectra have previously been confined to small systems. For large systems such as biological materials, full QM simulations have an extremely high computational cost and remain challenging. In this work, robust new algorithms and advanced implementations on many-core architectures are employed to enable fast, accurate, and massively parallel full ab initio simulations of the Raman spectra of realistic biological systems containing up to 3006 atoms, with excellent strong and weak scaling. Up to a performance of 468.5 PFLOP/s in double-precision and 813.7 PLOPS/s in mixed-half precision is achieved on the new-generation Sunway high-performance computing system, suggesting the potential for new applications of the QM approach to biological systems.
Honghui Shang, Yunquan Zhang, You Fu, Yingxiang Gao, Yangjun Wu, Xiaohui Duan, Rongfen Lin, Xin Liu 0081, Ying Liu 0055, Dexun Chen
SC8
2020 SWMapper: Scalable Read Mapper on SunWay TaihuLight
abstract
With the rapid development of next-generation sequencing (NGS) technologies, high throughput sequencing platforms continuously produce large amounts of short read DNA data at low cost. Read mapping is a performance-critical task, being one of the first stages required for many different types of NGS analysis pipelines. We present SWMapper — a scalable and efficient read mapper for the Sunway TaihuLight supercomputer. A number of optimization techniques are proposed to achieve high performance on its heterogeneous architecture which are centered around a memory-efficient succinct hash index data structure including seed filtration, duplicate removal, dynamic scheduling, asynchronous data transfer, and overlapping I/O and computation. Furthermore, a vectorized version of the banded Myers algorithm for pairwise alignment with 256-bit vector registers is presented to fully exploit the computational power of the SW26010 processor. Our performance evaluation shows that SWMapper using all 4 compute groups of a single Sunway TaihuLight node outperforms S-Aligner on the same hardware by a factor of 6.2. In addition, compared the state-of-the-art CPU-based mappers RazerS3, BitMapper2, and Hobbes3 running on a 4-core Xeon W-2123v3 CPU, SWMapper achieves speedups of 26.5, 7.8, and 2.6, respectively. Our optimizations achieve an aggregated speedup of 11 compared to the naïve implementation on one compute group of an SW26010 processor as well as a strong scaling efficiency of 74% on 128 compute groups.
Xiaohui Duan, Xiangxu Meng, Xin Li 0137, Bertil Schmidt
ICPP2
2020 Neighbor-list-free molecular dynamics on sunway TaihuLight supercomputer
abstract
Molecular dynamics (MD) simulations are playing an increasingly important role in many research areas. Pair-wise potentials are widely used in MD simulations of bio-molecules, polymers, and nano-scale materials. Due to a low compute-to-memory-access ratio, their calculation is often bounded by memory transfer speeds. Sunway TaihuLight is one of the fastest supercomputers featuring a custom SW26010 many-core processor. Since the SW26010 has some critical limitations regarding main memory bandwidth and scratchpad memory size, it is considered as a good platform to investigate the optimization of pair-wise potentials especially in terms of data reusage. MD algorithms often use a neighbor-list data structure to reduce the computational workload. In this paper, we show that a cell-list-based approach is more suitable for the SW26010 processor. We apply a number of novel optimization methods including self-adaptable replica-summation for conflict-free parallelization, parameter profiles for flexible vectorization, and particle-cell cutoff checking filters for reducing the computational workload. We also established an open source standalone framework featuring the techniques above, ESMD1, which is at least 50% faster than the latest existing LAMMPS port on a single TaihuLight node. Furthermore, EMSD achieves a weak scaling efficiency of 88% on 4,096 nodes.
Xiaohui Duan, Ping Gao 0005, Tingjian Zhang, Hongsong Meng, Bertil Schmidt, Haohuan Fu, Lin Gan 0001, Wei Xue 0003, Guangwen Yang 0002
PPoPP1
2020 Cell-list based molecular dynamics on many-core processors: a case study on sunway TaihuLight supercomputer
abstract
Molecular dynamics (MD) simulations are playing an increasingly important role in several research areas. The most frequently used potentials in MD simulations are pair-wise potentials. Due to the memory wall, computing pair-wise potentials on many-core processors are usually memory bounded. In this paper, we take the SW26010 processor as an exemplary platform to explore the possibility to break the memory bottleneck by improving data reusage via cell-list-based methods. We use cell-lists instead of neighbor-lists in the potential computation, and apply a number of novel optimization methods. Theses methods include: an adaptive replica arrangement strategy, a parameter profile data structure, and a particle-cell cutoff checking filter. An incremental cell-list building method is also realized to accelerate the construction of cell-lists. Furthermore, we have established an open source standalone framework, ESMD, featuring the techniques above. Experiments show that ESMD is 50~170% faster than previous ports on a single node, and can scale to 1,024 nodes with a weak scalibility of 95%.
Xiaohui Duan, Ping Gao 0005, Tingjian Zhang, Hongsong Meng, Bertil Schmidt, Haohuan Fu, Lin Gan 0001, Wei Xue 0003, Guangwen Yang 0002
SC1
2020 Tuning a general purpose software cache library for TaihuLight's SW26010 processor
Xiaohui Duan, Haohuan Fu, Lin Gan 0001, Wei Xue 0003, Guangwen Yang 0002
CCF Trans. High Perform. Comput.1
2020 Millimeter-Scale and Billion-Atom Reactive Force Field Simulation on Sunway Taihulight
abstract
Large-scale molecular dynamics (MD) simulations on supercomputers play an increasingly important role in many research areas. With the capability of simulating charge equilibration (QEq), bonds and so on, Reactive force field (ReaxFF) enables the precise simulation of chemical reactions. Compared to the first principle molecular dynamics (FPMD), ReaxFF has far lower requirements on computational resources so that it can achieve higher efficiencies for large-scale simulations. In this article, we present our efforts on scaling ReaxFF on the Sunway TaihuLight Supercomputer (TaihuLight). We have carefully redesigned the force analysis and neighbor list building steps. By applying fine-grained optimizations we gain better single process performance. For the many-body interactions, we propose an isolated computation and update strategy and implement inverse trigonometric functions. For QEq, we implement a pipelined conjugate gradient (CG) approach to achieving better scalability. Furthermore, we reorganize the data layout and implement the update operation based on data locality in ReaxFF. Our experiments show that this approach can simulate chemical reactions with 1,358,954,496 atoms using 4,259,840 cores with a performance of 0.015 ns/day. To our best knowledge, this is the first realization of chemical reaction simulation with a millimeter-scale force field.
Ping Gao 0005, Xiaohui Duan, Tingjian Zhang, Bertil Schmidt, Wusheng Zhang, Lin Gan 0001, Wei Xue 0003, Haohuan Fu, Guangwen Yang 0002
IEEE Trans. Parallel Distributed Syst.2
2019 Million-Core-Scalable Simulation of the Elastic Migration Algorithm on Sunway TaihuLight Supercomputer
abstract
Migration algorithm is one of the most essential methods in seismic application to image the underground geology, and to help scientists and researchers in geophysics exploration better understand the earth system. However, due to the desire in migration algorithm for covering lager region and acquiring better resolution, many tough challenges have to be tackled for current state-of-the-art computing systems. This work optimized and scaled the elastic migration algorithm onto the Sunway TaihuLight supercomputer, one of the most powerful systems of the world. Targeting at the major process, the reverse time migration (RTM) algorithm, a set of algorithmic, process-level, and thread-level optimizations is proposed, to significantly improve the performance (up to 163× speedup in time-to-solution) on Sunway CPU. Our design is successfully scaled to over two million cores (2,662,400 cores in total) on the Sunway TaihuLight supercomputer, with nearly ideal weak-scaling efficiency. The largest run is able to achieve a sustainable performance of processing over 859 billion cells per second.
Lin Gan 0001, Jingheng Xu, Xin Wang 0233, Sihai Wu, Xiaohui Duan, Haohuan Fu, Guangwen Yang 0002
CCGRID5
2019 SW_GROMACS: accelerate GROMACS on Sunway TaihuLight
abstract
GROMACS is one of the most popular Molecular Dynamic (MD) applications and is widely used in the field of chemical and bimolecular system study. Similar to other MD applications, it needs long run-time for large-scale simulations. Therefore, many high performance platforms have been employed to accelerate it, such as Knights Landing (KNL), Cell Processor, Graphics Processing Unit (GPU) and so on. As the third fastest supercomputer in the world, Sunway TaihuLight contains 40960 SW26010 processors and SW26010 is a typical many-core processor. To make full use of the superior computation ability of TaihuLight, we port GROMACS to SW26010 with following new strategies: (1) a new deferred update strategy; (2) a new update mark strategy; (3) a full pipeline acceleration. Furthermore, we redesign GROMACS to enable all possible vectorization. Experiments show that our implementation achieves better performance than both Intel KNL and Nvidia P100 GPU when using appropriate number of SW26010 processors for a fair comparison.
Tingjian Zhang, Ping Gao 0005, Mingshan Shao, Jinxiao Zhang, Xiaohui Duan, Lin Gan 0001, Haohuan Fu, Wei Xue 0003, Guangwen Yang 0002
SC8
2018 A Robotic Auto-Focus System based on Deep Reinforcement Learning
abstract
Considering its advantages in dealing with high-dimensional visual input and learning control policies in discrete domain, Deep Q Network (DQN) could be an alternative method of traditional auto-focus means in the future. In this paper, based on Deep Reinforcement Learning, we propose an end-to-end approach that can learn auto-focus policies from visual input and finish at a clear spot automatically. We demonstrate that our method - discretizing the action space with coarse to fine steps and applying DQN is not only a solution to auto-focus but also a general approach towards vision-based control problems. Separate phases of training in virtual and real environments are applied to obtain an effective model. Virtual experiments, which are carried out after the virtual training phase, indicates that our method could achieve 100% accuracy on a certain view with different focus range. Further training on real robots could eliminate the deviation between the simulator and real scenario, leading to reliable performances in real applications.
Xiaofan Yu 0001, Runze Yu 0001, Jingsong Yang, Xiaohui Duan
ICARCV4
2018 Redesigning LAMMPS for peta-scale and hundred-billion-atom simulation on Sunway TaihuLight
Xiaohui Duan, Ping Gao 0005, Tingjian Zhang, Wusheng Zhang, Wei Xue 0003, Haohuan Fu, Lin Gan 0001, Dexun Chen, Xiangxu Meng, Guangwen Yang 0002
SC1
2017 S-Aligner: Ultrascalable Read Mapping on Sunway Taihu Light
abstract
The availability and amount of sequenced genomes have been rapidly growing in recent years because of the adoption of next-generation sequencing (NGS) technologies that enable high-throughput short-read generation at highly competitive cost. Since this trend is expected to continue in the foreseeable future, the design and implementation of efficient and scalable NGS bioinformatics algorithms are important to research and industrial applications. In this paper, we introduce S-Aligner–a highly scalable read mapper designed for the Sunway Taihu Light supercomputer and its fourth-generationShenWei many-core architecture (SW26010). S-Aligner employs a combination of optimization techniques to overcome both the memory-bound and the compute-bound bottlenecks in the read mapping algorithm. In order to make full use of the compute power of Sunway Taihu Light, our design employs three levels of parallelism: (1) internode parallelism using MPI based on a task-grid pattern, (2) intranode parallelism using multithreading and asynchronous data transfer to fully utilize all 260 cores of the SW26010 many-core processor, and (3) vectorization to exploit the available 256-bit SIMD vector registers. Moreover, we have employed asynchronous access patterns and data-sharing strategies during file I/O to overcome bandwidth limitations of the network file system. Our performance evaluation demonstrates that S-Aligner scales almost linearly with approximately 95% efficiency for up to 13,312 nodes (concurrently harnessing more than 3 millioncompute cores). Furthermore, our implementation on a single node outperforms the established RazerS3 mapper running on a platform with eight Intel Xeon E7-8860v3 CPUs while achieving highly competitive alignment accuracy.
Xiaohui Duan, Yuandong Chan, Christian Hundt 0002, Bertil Schmidt, Pavan Balaji
CLUSTER1
2017 Redesigning CAM-SE for peta-scale climate modeling performance and ultra-high resolution on Sunway TaihuLight
abstract
The Community Atmosphere Model (CAM) is ported, redesigned, and scaled to the full system of the Sunway TaihuLight, and provides peta-scale climate modeling performance. We refactored and optimized the complete code using OpenACC directives at the first stage. A more aggressive and finer-grained redesign is then applied on the CAM, to achieve finer memory control and usage, more efficient vectorization and compute and communication overlapping. We further improve the CAM performance of a 260-core Sunway processor to the range of 28 to 184 Intel CPU cores, and achieve a sustainable double-precision performance of 3.3 PFlops for a 750 m global simulation when using 10,075,000 cores. CAM on Sunway achieves the simulation speed of 3.4 and 21.5 simulation-year-per-day (SYPD) for global 25-km and 100-km resolution respectively; and enables us to perform, to our knowledge, the first simulation of the complete lifecycle of hurricane Katrina, and achieve close-to-observation simulation results for both track and intensity.
Haohuan Fu, Junfeng Liao, Nan Ding 0006, Xiaohui Duan, Lin Gan 0001, Yishuang Liang, Jinzhe Yang, Lanning Wang, Guangwen Yang 0002
SC4
2014 Interactive crowdsourcing to spontaneous reporting of Adverse Drug Reactions
abstract
Adverse Drug Reactions (ADRs) has become a worldwide problem that draws the attention of people from all racial and ethnic groups. The number of deaths caused by ADRs has greatly increased and led to many drug withdrawals in the last decades. Recent research findings indicate that most ADRs can be effectively prevented to some extent by using computer-aided information technologies. Though many spontaneous reporting systems (SRSs) have been built to enhance the pharma-covigilance, the ADRs data is still very sparse because the large amount of reports obtained from consumers contains insufficient hints to identify a possible causal relationship between an adverse event and drug. Based on this motivation, we developed Adverse-Tracking, a spontaneous reporting system of ADRs via crowd-sourcing. Our proposed system interacts with consumers through a Q&A interface and collects the ADR reports. The decision tree support vector machine (DTSVM) based on the genetic algorithm is used in our system to automate the Q&A procedure. We carried out experiments to evaluate the performance at Peking University First Hospital. As demonstrated by the results, our system is an efficient tool to track and discover adverse events in the consumers' reports of ADRs, which facilitates the detection of “signal”.
Yining Huang, Chengdong Liu, Lingchao Meng, Yunchuang Sun, Kaigui Bian, Anpeng Huang, Xiaohui Duan, Bingli Jiao
ICC9
2014 WE-CARE: An Intelligent Mobile Telecardiology System to Enable mHealth Applications
abstract
Recently, cardiovascular disease (CVD) has become one of the leading death causes worldwide, and it contributes to 41% of all deaths each year in China. This disease incurs a cost of more than 400 billion US dollars in China on the healthcare expenditures and lost productivity during the past ten years. It has been shown that the CVD can be effectively prevented by an interdisciplinary approach that leverages the technology development in both IT and electrocardiogram (ECG) fields. In this paper, we present WE-CARE , an intelligent telecardiology system using mobile 7-lead ECG devices. Because of its improved mobility result from wearable and mobile ECG devices, the WE-CARE system has a wider variety of applications than existing resting ECG systems that reside in hospitals. Meanwhile, it meets the requirement of dynamic ECG systems for mobile users in terms of the detection accuracy and latency. We carried out clinical trials by deploying the WE-CARE systems at Peking University Hospital. The clinical results clearly showed that our solution achieves a high detection rate of over 95% against common types of anomalies in ECG, while it only incurs a small detection latency around one second, both of which meet the criteria of real-time medical diagnosis. As demonstrated by the clinical results, the WE-CARE system is a useful and efficient mHealth (mobile health) tool for the cardiovascular disease diagnosis and treatment in medical platforms.
Anpeng Huang, Kaigui Bian, Xiaohui Duan, Min Chen 0003, Hongqiao Gao, Yingrui Zhang, Bingli Jiao, Linzhen Xie
IEEE J. Biomed. Health Informatics4
2013 WE-CARE: A wearable efficient telecardiology system using mobile 7-lead ECG devices
abstract
Cardiovascular disease (CVD), principally heart disease and stroke, has become the worldwide leading killer for people of all racial and ethnic groups. In China, this disease contributes 41% to all deaths each year, and it costs the country more than 400 billion US dollars in last decade, including health expenditures and lost productivity, which continues to grow as the population size increases. Recent research findings indicate that CVD can be effectively prevented to some extent by an interdisciplinary approach that combines ICT (Information and Communication Technology) and healthcare applications. Based on this motivation, we developed WE-CARE, a Wearable Efficient teleCARdiology systEm using mobile 7-lead ECG devices. Our WE-CARE system surpasses existing resting ECG systems that reside in hospitals owing to its improved mobility brought by wearable and mobile ECG devices. Meanwhile, it outperforms those conventional dynamic 1-lead or 3-lead ECG systems for mobile users in terms of the appropriate ECG information required for clinical proposes. We carried out clinical trials by deploying the WE-CARE systems at People Hospital of Peking University. The results showed that our solution achieves a high detection rate over 95% against common types of anomalies in ECG signals, while its risk alert delay is limited around one second, both of which match the criteria of real-time healthcare monitoring very well. As demonstrated by the results, the WE-CARE system is a useful and efficient tool for the cardiovascular disease prevention and daily monitoring.
Kaigui Bian, Anpeng Huang, Xiaohui Duan, Hongqiao Gao, Bingli Jiao, Linzhen Xie
ICC4
2013 To enable stable medical image and video transmission in mobile healthcare services: A Best-fit Carrier Dial-up (BCD) algorithm for GBR-oriented applications in LTE-A networks
abstract
To provide healthcare services in mobile networks, GBR (Guaranteed Bit Rate)-oriented mechanism is expected to offer stable transmission for vital life and health information. Unfortunately, a kind of GBR-based applications is aggressive to exhaust limited bandwidths in a mobile network. To facilitate such a GBR-driven application, radio bandwidth scheduling should be agile and flexible. Motivated by this trend, we propose a Best-fit Carrier Dial-up (BCD) algorithm, which can dynamically schedule demanded bandwidth for GBR-oriented healthcare applications. In fact, this proposal is developed on the basis of Carrier Aggregation (CA), which is the typical new feature in LTE-A (Long Term Evolution-Advanced) networks. Compared with the CA technology, our proposal has two favorable properties: (1) the BCD algorithm uses the LTE-A own system information to select the best-fit carrier in order to avoid extra signaling overhead. Thus it is suitable for healthcare services which are time-sensitive; (2) the BCD algorithm can dial-up bandwidth on demand at a small-scale granularity. Most valuably, this proposal is backward compatible with CA standards seamlessly. In this study, medical image and consulting video are chosen for simulation experiments. Results verify that our proposal can finely tailor radio bandwidth allocation for the time-sensitive and stable healthcare services at a required GBR level.
Yingrui Zhang, Anpeng Huang, Daoxian Wang, Xiaohui Duan, Bingli Jiao, Linzhen Xie
ICC4
2012 BER analysis for MMSE-FDE-based interleaved SC-FDMA systems over Nakagami-m fading channels
abstract
In this paper, we present an analytical study of the bit error rate (BER) for interleaved single-carrier frequency-division multiple access (SC-FDMA) systems over independent but not necessarily identically distributed (i.n.i.d.) Nakagami-m fading channels with fading parameters {m} being integers when minimum mean-square error frequency-domain equalization (MMSE-FDE) is applied. Under the assumption of independent fading characteristics among channel frequency responses (CFRs) at the allocated subcarriers for a specific user, accurate numerical BER computation for square M-ary quadrature amplitude modulation (M-QAM) is developed by exploiting the statistics of the equalized noise. More importantly, the BER derivation is based on the real distribution of the CFRs without applying the widely used approximation of the CFRs in previous literature, resulting in a more accurate BER analysis. Monte-Carlo simulations are conducted to validate the analysis.
Miaowen Wen, Xiang Cheng 0001, Zhongshan Zhang, Xiaohui Duan, Bingli Jiao
GLOBECOM4
2012 A 3R dataflow engine for restoring electrophysiological signals in telemedicine cloud platforms
abstract
Today, IT technology plays a key role in promoting health services which provide medical information over a digital information system. Unfortunately, the digitized data may be not understood by medical professionals due to the limitation of conventional medical signal recognition. To deal with this challenge, we design a 3R (Retiming, Regeneration, Reshaping) dataflow engine that can restore electrophysiological data into its original medical patterns. We carried out clinical trials on our cloud platform in PKU's People Hospital. Clinical results proved that our solution produces high playback accuracy and matches with medical diagnosis criteria very well. With the proof of our clinical tests, this solution can be a very useful tool for clinical treatment and diagnosis in medical platforms.
Yingrui Zhang, Anpeng Huang, Daoxian Wang, Xiaohui Duan, Bingli Jiao, Linzhen Xie
Healthcom4
2011 Studies on the coherence bandwidth of high-frequency vehicular communication system with antenna array
abstract
Abstract High‐frequency (HF) vehicular communication system works with the frequency‐selective fading channel, thereby the coherence bandwidth is limited due to the multipath time dispersive effect, which imposes the difficulty for achieving the high‐data rate transmission. The present paper investigates the antenna array technique for its ability in reducing the multipath delay spread of the ionosphere channel and, thus, increasing possibly the bandwidth of the HF system. To enable fast moving speed of the vehicular receiver, we propose to use single antenna at the receiver and the adaptive array at the transmitter. The adaptive beamforming is driven based on the protocol design and reciprocity principle of channel. In order to evaluate the performance of array system, a statistical HF double‐directional channel model, which includes angular information at both the transmitter and receiver, is presented. The explicit expressions are obtained for calculating the coherence bandwidth. The numerical results show the efficiency of this approach for mid‐latitude mid‐range channel. Copyright © 2009 John Wiley & Sons, Ltd.
Xiaohui Duan, Bingli Jiao
Wirel. Commun. Mob. Comput.2
2010 Sparse Representation-Based Face Recognition for One Training Image per Person
Xueping Chang, Zhonglong Zheng, Xiaohui Duan, Chenmao Xie
ICIC (1)3
2004 A Semi-Fragile Watermark Scheme For Image Authentication
abstract
In this paper, a semifragile watermarking scheme for image authentication is proposed, which addresses the issue of protecting images from illegal manipulations and modifications. The scheme extracts a signature (watermark) from the original image and inserts this signature back into the image, avoiding additional signature files. The error correction coding (ECC) is used to encode the signatures that are extracted from the image. To increase the security of this scheme, user's private key is employed for encryption and decryption of the watermark during watermark extraction and insertion procedures. Experimental result shows that if there is no change in the obtained image, the watermark will be correctly extracted, thus will pass through the authentication system. This scheme is tolerant of lossy compression such as JPEG, but malicious changes of the image will result in the breach of the watermark detection. In addition, this scheme can detect the exact locations, which are illegal modified blocks.
Xiaohui Duan, Daoxian Wang
MMM2