VLDB 2026 Research / reviewers in the wild / expert
Haohuan Fu
dblp:71/3657
· DBLP profile ↗
156ranked-venue papers
17as first author
69since 2021 · last 2026
0000-0002-6982-2235ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 97 · 13 first-author · 24 since 2021Applied, interdisciplinary, general and emerging computing · 35 · 1 first-author · 29 since 2021Artificial intelligence and machine learning · 16 · 13 since 2021Computer networks · 5 · 2 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 5 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Practical and Efficient x86-64 Emulation on RISC-VabstractAs RISC-V increasingly gains traction across various computing domains, the need to run x86-64 applications on RISC-V platforms becomes critical. Xiongchuan Tan, Yang Liu 0005, Sebastien Chevalier, Yangyu Chen 0002, Haohuan Fu |
EuroSys | 6 |
| 2026 | Continual test-time adaptation for object detection with adaptive monitoring and randomized restoration
Shilei Cao 0005, Juepeng Zheng, Baoquan Zhao, Runmin Dong, Haohuan Fu |
Expert Syst. Appl. | 8 |
| 2026 | GALA: A GlobAl-LocAl cluster approach for Multi-Source Active Domain Adaptation
Juepeng Zheng, Yibin Wen, Peifeng Zhang, Zurong Mai, Qingmei Li, Haohuan Fu |
Pattern Recognit. | 7 |
| 2026 | Evidential Graph Contrastive Alignment for Source-Free Blending-Target Domain AdaptationabstractIn this article, we first tackle a more realistic domain adaptation (DA) setting: source-free blending-target DA (SF-BTDA), where we cannot access to source-domain data while facing mixed multiple target domains without any domain labels in prior. Compared to existing DA scenarios, SF-BTDA generally faces the coexistence of different label shifts in different targets, along with noisy target pseudolabels generated from the source model. In this article, we propose a new method called evidential graph contrastive alignment (EGCA) to decouple the blending-target domain and alleviate the effect of noisy target pseudolabels. First, to improve the quality of pseudo target labels, we propose a calibrated evidential learning (CEL) module to iteratively improve both the accuracy and certainty of the resulting model and adaptively generate high-quality pseudo target labels. Second, we design a graph contrastive learning with the domain distance matrix and confidence-uncertainty criterion, to minimize the distribution gap of samples of the same class in the blending-target domain, which alleviates the coexistence of different label shifts in blended targets. We conduct a new benchmark based on three standard DA datasets, and EGCA outperforms other methods with considerable gains and achieves comparable results compared with those that have domain labels or source data in prior. Juepeng Zheng, Yibin Wen, Jinxiao Zhang, Runmin Dong, Haohuan Fu |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2026 | SWGOMP: Extending OpenMP for Efficient Offloading on Sunway Heterogeneous Architecture
Qixin Chang, Xiaohui Duan, Huihai An, Yi Zhang 0127, Haohuan Fu, Bin Yang 0043, Yilun Han, Dongqiang Huang, Xiting Ju, Haopeng Huang, Wei Xue 0003, Lin Gan 0008, Maoxue Yu, Jian Li 0069, Zhao Jing, Hailong Liu 0007, Lixin Wu, Ren Hu |
IEEE Trans. Parallel Distributed Syst. | 6 |
| 2026 | Exploiting the Performance Potential of Extreme-Scale Earthquake Simulation: Achieving 86.7 PFLOPS With Over 39 Million CoresabstractLeveraging the latest Sunway supercomputer, we developed a fully optimized earthquake simulation model that accurately captures topographic effects for realistic seismic analysis. Optimizing for the SW26010Pro architecture with DMA/RMA communication mechanisms, data compression schemes, and vectorization, we achieved a speedup exceeding 160×. Our pipeline-based computation and communication overlapping scheme, combined with performance prediction models further minimized computational costs. These optimizations enabled the largest-scale curvilinear grid finite-difference method (CGFDM) earthquake simulations to date, covering 197 trillion grid points and achieving 86.7 PFLOPS on 39 million cores with a weak scaling efficiency of 97.9%. These advancements enabled the successful simulation of the 2008 Wenchuan earthquake, providing high-resolution seismic insights and robust assessments for regional hazard mitigation and disaster preparedness. Lin Gan 0008, Wubing Wan, Zekun Yin, Zhong He, Ping Gao 0005, Xiaohui Duan, Wei Xue 0003, Haohuan Fu, Guangwen Yang 0002 |
IEEE Trans. Parallel Distributed Syst. | 12 |
| 2025 | SeasonBench-EA: A Multi-Source Benchmark for Seasonal Prediction and Numerical Model Post-Processing in East AsiaabstractSeasonal-scale climate prediction plays a critical role in supporting agricultural planning, disaster prevention, and long-term decision making. In particular, reliable forecasts issued 1-6 months in advance are essential for early warning of flood and drought risks associated with precipitation during the East Asian summer monsoon season. However, while the use of machine learning techniques has advanced rapidly in weather and subseasonal-to-seasonal forecasting, partly driven by the availability of benchmark datasets, their application to seasonal-scale prediction remains limited. Existing seasonal prediction primarily relies on ensemble forecasts from numerical models, which, while physically grounded, are subject to biases and uncertainties at long lead times. Motivated by these challenges, we propose SeasonBench-EA, a benchmark dataset for seasonal prediction in East Asia region. It features multi-resolution, multi-source data with both regional and global coverage, integrating ERA5 reanalysis data and ensemble forecasts from multiple leading forecast centers. Beyond key atmospheric fields, the dataset also includes boundary-related variables, such as ocean state, soil and solar radiation, that are essential for capturing seasonal-scale atmospheric variability. Two tasks are defined and evaluated: 1) machine learning-based seasonal prediction using ERA5 reanalysis, and 2) post-processing of seasonal forecasts from numerical model ensembles. A suite of deterministic and probabilistic metrics is provided for tasks evaluation, along with a hindcast assessment focused on precipitation during the East Asian summer monsoon, aligned with model evaluation protocols used in operations. By offering a unified data and evaluation framework, SeasonBench-EA aims to promote the development and application of data-driven methods for seasonal prediction, a challenging yet highly impactful task with board implications for society and public well-being. Our benchmark is available at https://github.com/SauryChen/SeasonBench-EA. Mengxuan Chen, Ziheng Zou, Jinxiao Zhang, Runmin Dong, Juepeng Zheng, Haohuan Fu |
NeurIPS | 8 |
| 2025 | DiffLiG: Diffusion-enhanced Liquid Graph with Attention Propagation for Grid-to-Station Precipitation CorrectionabstractModern precipitation forecasting systems, including reanalysis datasets, numerical models, and AI-based approaches, typically produce coarse-resolution gridded outputs. The process of converting these outputs to station-level predictions often introduces substantial spatial biases relative to station-level observations, especially in complex terrains or under extreme conditions. These biases stem from two core challenges: (i) $\textbf{station-level heterogeneity}$, with site-specific temporal and spatial dynamics; and (ii) $\textbf{oversmoothing}$, which blurs fine-scale variability in graph-based models. To address these issues, we propose $\textbf{DiffLiG}$ ($\underline{Diff}$usion-enhanced $\underline{Li}$quid $\underline{G}$raph with Attention Propagation), a graph neural network designed for precise spatial correction from gridded forecasts to station observations. DiffLiG integrates a GeoLiquidNet that adapts temporal encoding via site-aware OU dynamics, a graph neural network with a dynamic edge modulator that learns spatially adaptive connectivity, and a Probabilistic Diffusion Selector that generates and refines ensemble forecasts to mitigate oversmoothing. Experiments across multiple datasets show that DiffLiG consistently outperforms other methods, delivering more accurate and robust corrections across diverse geographic and climatic settings. Moreover, it achieves notable gains on other key meteorological variables, underscoring its generalizability and practical utility. Mengxuan Chen, Haohuan Fu, Juepeng Zheng |
NeurIPS | 7 |
| 2025 | Can Large Multimodal Models Understand Agricultural Scenes? Benchmarking with AgroMindabstractLarge Multimodal Models (LMMs) has demonstrated capabilities across various domains, but comprehensive benchmarks for agricultural remote sensing (RS) remain scarce. Existing benchmarks designed for agricultural RS scenarios exhibit notable limitations, primarily in terms of insufficient scene diversity in the dataset and oversimplified task design. To bridge this gap, we introduce AgroMind, a comprehensive agricultural remote sensing benchmark covering four task dimensions: spatial perception, object understanding, scene understanding, and scene reasoning, with a total of 13 task types, ranging from crop identification and health monitoring to environmental analysis. We curate a high-quality evaluation set by integrating nine public datasets and one private global parcel dataset, containing 28,482 QA pairs and 20,850 images. The pipeline begins with multi-source data pre-processing, including collection, format standardization, and annotation refinement. We then generate a diverse set of agriculturally relevant questions through the systematic definition of tasks. Finally, we employ LMMs for inference, generating responses, and performing detailed examinations. We evaluated 20 open-source LMMs and 4 closed-source models on AgroMind. Experiments reveal significant performance gaps, particularly in spatial reasoning and fine-grained recognition, it is notable that human performance lags behind several leading LMMs. By establishing a standardized evaluation framework for agricultural RS, AgroMind reveals the limitations of LMMs in domain knowledge and highlights critical challenges for future work. Data and code can be accessed at https://rssysu.github.io/AgroMind/. Qingmei Li, Zurong Mai, Shuohong Lou, Henglian Huang, Jiarui Zhang 0008, Yibin Wen, Haohuan Fu, Jianxi Huang, Juepeng Zheng |
NeurIPS | 11 |
| 2025 | SPFL: Sequential updates with Parallel aggregation for Enhanced Federated Learning under Category and Domain ShiftsabstractFederated learning (FL) has recently emerged as the primary approach to overcoming data silos,
enabling collaborative model training without sharing sensitive or proprietary data.
Parallel federated learning (PFL) aggregates models trained independently on each client’s local data, which can lead to suboptimal convergence due to limited data exposure.
In contrast, Sequential Federated Learning (SFL) allows models to traverse client datasets sequentially, enhancing data utilization.
However, SFL effectiveness is limited in real-world non-IID scenarios characterized by category shift (inconsistent class distributions) and domain shift (distribution discrepancies).
These shifts cause two critical issues: update order sensitivity, where model performance varies significantly with the sequence of client updates, and catastrophic forgetting, where the model forgets previously learned features when trained on new client data.
We propose SPFL, a novel updating method that can be integrated into existing FL methods, integrating sequential updates with parallel aggregation to enhance data utilization and ease update order sensitivity. At the same time, we give the convergence analysis of SPFL under strong convex, general convex, and non-convex conditions, proving that this update scheme is significantly better than PFL and SFL.
Additionally, we introduce the Global-Local Alignment Module to mitigate catastrophic forgetting by aligning the predictions of the global model with those of the local and previous models during training.
Our extensive experiments demonstrate that integrating SPFL into existing PFL methods significantly improves performance under category and domain shifts. Haoyuan Liang, Shilei Cao 0005, Zhiyu Ye, Haohuan Fu, Juepeng Zheng |
NeurIPS | 5 |
| 2025 | GTPBD: A Fine-Grained Global Terraced Parcel and Boundary DatasetabstractAgricultural parcels serve as basic units for conducting agricultural practices and applications, which is vital for land ownership registration, food security assessment, soil erosion monitoring, etc. However, existing agriculture parcel extraction studies only focus on mid-resolution mapping or regular plain farmlands while lacking representation of complex terraced terrains due to the demands of precision agriculture. In this paper, we introduce a more fine-grained terraced parcel dataset named GTPBD (Global Terraced Parcel and Boundary Dataset), which is the first fine-grained dataset covering major worldwide terraced regions with more than 200,000 complex terraced parcels with manually annotation. GTPBD comprises 47,537 high-resolution images with three-level labels, including pixel-level boundary labels, mask labels, and parcel labels. It covers seven major geographic zones in China and transcontinental climatic regions around the world. Compared to the existing datasets, the GTPBD dataset brings considerable challenges due to the: (1) terrain diversity; (2) complex and irregular parcel objects; and (3) multiple domain styles. Our proposed GTPBD dataset is suitable for four different tasks, including semantic segmentation, edge detection, terraced parcel extraction and unsupervised domain adaptation (UDA) tasks. Accordingly, we benchmark the GTPBD dataset on eight semantic segmentation methods, four edge extraction methods, three parcel extraction methods and five UDA methods, along with a multi-dimensional evaluation framework integrating pixel-level and object-level metrics. GTPBD fills a critical gap in terraced remote sensing research, providing a basic infrastructure for fine-grained agricultural terrain analysis and cross-scenario knowledge transfer. The code and data are available at https://github.com/Z-ZW-WXQ/GTPBD/. Yibin Wen, Shuai Yuan 0005, Haohuan Fu, Jianxi Huang, Juepeng Zheng |
NeurIPS | 5 |
| 2025 | An AI-Enhanced 1km-Resolution Seamless Global Weather and Climate Model to Achieve Year-Scale Simulation Speed using 34 Million CoresabstractGlobal Storm Resolving Models (GSRMs) is crucial for understanding extreme weather events under the climate change background. In this study, we optimize Global-Regional Integrated Forecast System (GRIST), which is a unified weather-climate modeling system designed for research and operation, for the next-generation Sunway supercomputer, incorporating AI-enhanced physics suite, OpenMP-based parallelization, and mixed-precision optimizations to enhance both efficiency and performance portability, as well as the unified modeling capability. Our experiments successfully capture significant events during the "23.7" extreme rainfall over northern China influenced by super Typhoon Doksuri, at 1km resolution. Notably, our work scales to 34 million cores, enabling simulation speeds at 491 SDPD (3km) and 181 SDPD (1km). Xiaohui Duan, Yi Zhang 0127, Haohuan Fu, Bin Yang 0043, Yilun Han, Dongqiang Huang, Huihai An, Xiting Ju, Haopeng Huang, Wei Xue 0003, Jianye Hou, Maoxue Yu, Jian Li 0069, Zhao Jing, Hailong Liu 0007, Lixin Wu |
PPoPP | 4 |
| 2025 | Kilometer-Scale AI-Powered and Performance-Portable Earth System Model (AP3ESM) to Achieve Year-Scale Simulation Speed on Heterogeneous SupercomputersabstractKilometer-scale Earth system models (ESMs) necessitate exascale supercomputers to facilitate realistic simulations of weather phenomena and climate variability over a time span ranging from days to decades. We present AP3ESM, an ultra‑high‑resolution, AI‑Powered, Performance‑Portable ESM coupling atmosphere, land surface, ocean, and sea ice components. By leveraging the performance portability features of Kokkos and OpenMP, the AP3ESM operates efficiently on two heterogeneous systems while incurring minimal development overhead. Advanced optimization techniques, such as adaptive parallel algorithms, AI-enhanced physical parameterizations, and mixed-precision computations, have been implemented to further boost the computational efficiency. Breaking the 1-km resolution barrier, AP3ESM delivers 0.85 and 1.98 simulated-years-per-day (SYPD) for the standalone atmosphere and ocean components on 34.1 million Sunway cores and 16085 GPUs, respectively; the holistic AP3ESM achieves 0.54 SYPD on 37.2 million Sunway cores. Notably, the forecast experiment successfully captures Super Typhoon Doksuri in 2023 and its associated extreme rainfall across China. Maoxue Yu, Yuhu Chen, Jiaying Song, Xiaohui Duan, Junwei Wei, Jiangfeng Yu, Hailong Liu 0007, Jinrong Jiang, Yi Zhang 0127, Pengfei Lin 0004, Weipeng Zheng, Jingwei Xie, Jiakang Zhang, Zilu Liu, Xiaoyu Jin, Jilin Wei, Qixin Chang, Qingxia Lin, Yanzhi Zhou, Wei Xue 0003, Haohuan Fu, Yue Yu 0001, Xuebin Chi, Lixin Wu |
SC | 28 |
| 2025 | Behaviour-diverse automatic penetration testing: a coverage-based deep reinforcement learning approach
Yizhou Yang, Longde Chen, Lanning Wang, Haohuan Fu, Xin Liu 0081, Zuoning Chen |
Frontiers Comput. Sci. | 5 |
| 2025 | BAN: A Universal Paradigm for Cross-Scene Classification Under Noisy Annotations From RGB and Hyperspectral Remote Sensing ImagesabstractWhile domain adaptation (DA) methods have made significant strides in remote sensing community, most current works assume that the source domain labels are accurate. However, limited emphasis has been placed on the scenario where source data are mislabeled with noisy annotations, which is more common in real applications and referred to as noisy DA (NDA). This article formulates remote sensing cross-scene classification on NDA scenarios and proposes a novel network called bilateral adaptation network (BAN), which consists of two parts: 1) forward learning (FL), which utilizes a model learning from the noisy source domain and transfers knowledge to target domain; and 2) backward learning (BL), which utilizes a dual model to acquire knowledge from the target domain and transfer it to source domain. We conduct two parts alternately and adopt a symmetrical Kullback-Leibler (KL) loss to align predictions of the model and its dual model in the same domain. This interactive strategy is able to explore bilateral relationships between domains, implicitly reducing label noise in the source domain. In addition, BAN could serve as a universal paradigm to not only improve the existing NDA methods but also enhance recent DA approaches. Comprehensive evaluations on three publicly available RGB-band remote sensing datasets and two hyperspectral datasets validate the superior effectiveness of our proposed BAN. BAN improves the average accuracy by 6.70%–15.70% on RGB datasets and overall accuracy (OA) by 1.36%–3.14% on hyperspectral datasets with flip-20% noise compared to other state-of-the-art DA and NDA approaches. Promising results indicate the potential of our approach in tackling more general and practical problems with noisy source domain. Wentang Chen, Yibin Wen, Juepeng Zheng, Jianxi Huang, Haohuan Fu |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2025 | Boosting Universal Domain Adaptation in Remote Sensing With Dual-Classifiers Consistency Discrimination and Cross-Domain Feature MixupabstractIn the field of remote sensing image classification, domain adaptation (DA) methods have been extensively utilized to overcome the challenges posed by data discrepancies between source and target domains that arise from varying imaging conditions, sensor differences, or geographical variations. Stemming from the existence of unseen classes in both the source and target domains, universal DA (UniDA) poses the greatest challenge that demands innovative solutions. Existing UniDA methods often overlook intra-domain variations within the target domain and face difficulties in distinguishing between similar known and unknown classes, which significantly hinder cross-domain transfer. To overcome these challenges, we propose a dual-classifier network tailored for cross-domain classification of remote sensing images, namedDCmix. DCmix introduces a dual-classifiers network that utilizes both closed-set and open-set classifiers to improve the accuracy of identifying unknown sample classes. To our knowledge, this is the first attempt to introduce dual classifiers into the UniDA remote sensing image classification task. We further enhance the feature generalization capability of the target domain based on sample neighborhood relations, resulting in a more adaptable and robust feature representation. A cross-domain feature mixup scheme is also designed based on the consistency discrimination of the dual classifiers, achieving smoother decision boundaries and simpler hidden layer representations. Extensive experiments conducted on four hyperspectral image datasets and three RGB datasets prove that the introduced approach attains state-of-the-art performance in remote sensing image classification under the UniDA scenario. Qingmei Li, Juepeng Zheng, Jianxi Huang, Haohuan Fu |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2025 | An Efficient Point Spread Function Inversion Method for Image-Domain One-Way Wave-Equation Least-Squares MigrationabstractImage-domain least-squares migration (LSM) has demonstrated promising potential in enhancing the spatial resolution of migration images effectively and efficiently. However, existing image-domain approaches are mostly based on a local-stationary assumption, which estimates a local-stationary deblurring filter to process the corresponding subsection of the migration image. The deblurring precision is not fine enough. A point spread function (PSF) deconvolution method has been proposed to improve the resolution of migration images on a point-wise bias. Nevertheless, the computational and storage costs, particularly during the PSF process, remain significant. To achieve high-resolution imaging with reduced costs, we propose a PSF inversion method for image-domain one-way wave-equation (OWE) LSM. Leveraging a deep-learning optimizer, we achieve a rapid convergence for inverting the PSF in the spatial domain. In addition, we introduce a rescaled loss function for the stabilization and acceleration of the PSF inversion process. The rescaled loss function also makes it possible to obtain decent deblurring results during the early stages of iterations. Through some synthetic and field dataset experiments, it can be determined that our proposed PSF inversion method can produce high-resolution images with reduced migration artifacts and balanced amplitude. Additionally, our proposed method boasts noniterative characteristics, high parallelizability, freedom from regularization, and reduced storage and computational overhead, rendering it efficient and well-suited for practical applications. Cewen Liu, Mengyao Sun 0002, Nanxun Dai, Mingjie Guo, Haohuan Fu |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2025 | Accelerating Half-Precision Seismic Simulation on Neural Processing UnitabstractDue to the superiority of handling irregular regions of interest, the curvilinear grid finite difference method (CGFDM) has become wildly used in seismic simulation for earthquake hazard evaluation and understanding of earthquake physics. This paper proposes a novel approach that optimizes a CGFDM solver on the Ascend, a cutting-edge Neural Processing(NPU) Unit using half-precision storage and mixed-precision arithmetic. The approach increases the data throughput and computing efficiency, enabling more effective seismic modeling. Furthermore, we propose an efficient matrix unit enabled 3D difference algorithm that employs matrix unit on NPU to accelerate the computation. By fully exploiting the capability of matrix unit and wide SIMD lane, our solver on Ascend achieves a speedup of 4.19 × over the performance of parallel solver on two AMD CPUs and has successfully simulated real-world Wenchuan earthquake. For the best of our knowledge, we are the first to conduct seismic simulations on NPU. Wubing Wan, Lin Gan 0008, Ping Gao 0005, Haohuan Fu, Wei Xue 0003, Guangwen Yang 0002 |
IEEE Trans. Parallel Distributed Syst. | 9 |
| 2024 | Building Bridges Across Spatial and Temporal Resolutions: Reference-Based Super-Resolution via Change Priors and Conditional Diffusion ModelabstractReference-based super-resolution (RefSR) has the potential to build bridges across spatial and temporal resolutions of remote sensing images. However, existing RefSR methods are limited by the faithfulness of content reconstruction and the effectiveness of texture transfer in large scaling factors. Conditional diffusion models have opened up new opportunities for generating realistic high-resolution images, but effectively utilizing reference images within these models remains an area for further exploration. Furthermore, content fidelity is difficult to guarantee in areas without relevant reference information. To solve these issues, we propose a change-aware diffusion model named Ref-Diff for RefSR, using the land cover change priors to guide the denoising process explicitly. Specifically, we inject the priors into the denoising model to improve the utilization of reference information in unchanged areas and regulate the reconstruction of semantically relevant content in changed areas. With this powerful guidance, we decouple the semantics-guided denoising and reference texture-guided denoising processes to improve the model performance. Extensive experiments demonstrate the superior effectiveness and robustness of the proposed method compared with state-of-the-art RefSR methods in both quantitative and qualitative evaluations. The code and data are available at https://github.com/dongrunmin/RefDiff. Runmin Dong, Shuai Yuan 0005, Mengxuan Chen, Jinxiao Zhang, Lixian Zhang 0002, Juepeng Zheng, Haohuan Fu |
CVPR | 9 |
| 2024 | Universal Domain Adaptation for Hyperspectral Image ClassificationabstractAlthough enormous Domain Adaptation (DA) approaches have been proposed for cross-scene hyperspectral image (HSI) classification, most of them strongly rely on prior knowledge of the relationship between the label sets of source and target domains (including closed-set, partial and open-set DA), which significantly limits their applications. In a real-world application scenario, we often transfer knowledge between domains without any constraints on the label sets, which is called Universal Domain Adaptation (UniDA). In this paper, we propose HyUniDA, which is the first attempt to address UniDA scenario from HSIs. HyUniDA contains two major parts: the Shared Semantic Pairing (SSP) and Domain Similarity Score (DSS). The SSP identifies pairs of clusters that have coincident semantic features as the common classes. By examining the consistency level of samples across source and target domains, DSS can estimate the number of target clusters and generate distinct clusters without prior knowledge. We evaluate our proposed method on two transfer learning tasks for four typical HSI datasets, it turns out that our proposed method yields 6.41%∼34.71% improvements compared to other state-of-the-art DA methods. Qingmei Li, Yibin Wen, Juepeng Zheng, Haohuan Fu |
IGARSS | 5 |
| 2024 | Receptive Convolution Boosts Large-Scale Multi-Class Change DetectionabstractChange detection in remote sensing is crucial for land use/land cover (LULC) change awareness. However, large-scale change detection suffers from limited contextual understanding of spatial intricacies of large changes in existing CNN-based methods. Overlapped receptive fields lead to weight sharing across feature sliders, contributing to limited detection ability on large dense multi-class changes. To address this problem, this paper presents a receptive convolution operation for large-scale multi-class change detection from high-resolution remote sensing images. Different from existing CNN-based networks, our architecture involves receptive convolutions with a large kernel size to guarantee focus on different receptive field features. Experiments on SECOND datasets show that the proposed method achieves better performance than previous counterparts. Furthermore, a large-scale LULC change detection is conducted to demonstrate the ability in large-scale applications. Shuai Yuan 0005, Lixian Zhang 0002, Haohuan Fu, Peng Gong 0002 |
IGARSS | 4 |
| 2024 | Unveiling Annual Dynamics in Large-Scale Road Networks Through a Connectivity-Aware Approach Utilizing Sentinel-2 Multi-Spectral ImageryabstractEfficient and timely assessment of road network dynamic changes is crucial for the comprehensive evaluation of urban development, transportation accessibility, and environmental impacts. While existing methods mainly focus on optimizing performance with very-high-resolution remote sensing images on public road datasets, their practical applicability remains to unlock when confronted with large-scale real-world applications utilizing multi-spectral remote sensing images. The limitations manifest in unsatisfactory model generalization and fragmented segmentation, reducing the effectiveness of road extraction outcomes. To address these challenges, this study introduces a novel connectivity-aware approach tailored to address road extraction challenges in real-world scenarios. Leveraging Sentinel-2 multi-spectral imagery, this study conducts a 6-year road change mapping over an expansive 10,097 square kilometers in Xi’an, China. Experimental and evaluation results underscore the efficacy of the proposed methodology for widespread applications in urban planning and environmental management, offering a robust solution for practical and efficient road extraction in diverse and extensive large-scale urban investigation. Lixian Zhang 0002, Kangrui Du, Shuai Yuan 0005, Runmin Dong, Juepeng Zheng, Haohuan Fu |
IGARSS | 6 |
| 2024 | Single Domain Generalization For Scene Classification Using Style-Oriented Data AugmentationabstractDomain generalization (DG), which tries to improve the performance of models trained with known domains but applied to unknown domains, is an important step towards practical solutions in real-world scenarios. In this paper, we tackle a much more difficult scenario called single domain generation in scene classification problem, where only one source domain is available during training. Existing DG methods usually focus on extracting invariant features from different known domains and often suffer from overfitting issues. Therefore, to tackle the above challenge, we propose a randomly-stylized data augmentation method, which enables randomized style perturbation of the training data, to alleviate the overfitting problem and to improve the robustness of the resulting model. On a multidomain scene classification benchmark, our method achieves an accuracy improvement of 0.4%-2.5% compared to other DG methods. Yi Zhao 0024, Guancong Lin, Juepeng Zheng, Yang You 0001, Haohuan Fu |
IGARSS | 5 |
| 2024 | DeepLight: Reconstructing High-Resolution Observations of Nighttime Light With Multi-Modal Remote Sensing Data
Lixian Zhang 0002, Runmin Dong, Shuai Yuan 0005, Jinxiao Zhang, Mengxuan Chen, Juepeng Zheng, Haohuan Fu |
IJCAI | 7 |
| 2024 | Circular Reconfigurable Parallel Processor for Edge Computing : Industrial Product ✶abstractGraphics Processing Units (GPUs) have emerged as the predominant hardware platforms for massively parallel computing. However, their inherent von-Neumann architecture still suffers performance inefficiency stemming from the sequential instruction execution and frequent data transfer overheads within the memory system. These intrinsic architectural flaws lead to heavy overhead on the latency, area, and energy efficiency, rendering GPUs suboptimal for edge computing applications. To tackle these challenges, this paper introduces a novel circular Reconfigurable Parallel Processor (RPP) to enable massively parallel applications in edge computing with high efficiency. RPP features a novel circular array of reconfigurable compute engines, enabling efficient streaming dataflow processing. In contrast to traditional Coarse Grained Reconfigurable Architecture (CGRA), the circular network topology of RPP is formed by linear switch networks with an innovative gasket memory, which reduces complicated network routing overheads while allowing versatile datapath mapping and optimized data reuse. A dedicated hierarchical memory system is proposed to support different memory access patterns and address mapping strategies, enabling flexible data access with high memory efficiency. Several hardware optimizations are further introduced to improve hardware utilization and performance such as concurrent kernel execution, register split&refill and heterogeneous scalar&vector computing. To fully utilize the hardware capability of RPP, we develop an end-to-end software stack consisting of a compiler, runtime environment, and different RPP libraries. This software stack is designed to be compatible with the GPGPU computing paradigm, enhancing its potential for broader adoption. Fabricated in a 14nm process, RPP occupies an area of 119 mm2and operates at a maximum power of 15W with a 1GHz clock frequency. From the runtime measurement of various workloads, RPP achieves up to 27.5 × higher energy efficiency than Nvidia edge GPUs in deep learning inference and up to 14062 × lower latency than AMD Ryzen 5 CPU in linear algebra operations. Jianbin Zhu, Toshio Nagata, Ryan Braidwood, Haohuan Fu, Juepeng Zheng, Wayne Luk, Hongxiang Fan |
ISCA | 7 |
| 2024 | Spatial-Temporal Context Model for Remote Sensing Imagery CompressionabstractWith the increasing spatial and temporal resolutions of obtained remote sensing (RS) images, effective compression becomes critical for storage, transmission, and large-scale in-memory processing. Although image compression methods achieve a series of breakthroughs for daily images, a straightforward application of these methods to RS domain underutilizes the properties of the RS images, such as content duplication, homogeneity, and temporal redundancy. This paper proposes a Spatial-Temporal Context model (STCM) for RS image compression, jointly leveraging context from a broader spatial scope and across different temporal images. Specifically, we propose a stacked diagonal masked module to expand the contextual reference scope, which is stackable and maintains its parallel capability. Furthermore, we propose spatial-temporal contextual adaptive coding to enable the entropy estimation to reference context across different temporal RS images at the same geographic location. Experiments show that our method outperforms previous state-of-the-art compression methods on rate-distortion (RD) performance. For downstream tasks validation, our method reduces the bitrate by 52 times for single temporal images in the scene classification task while maintaining accuracy. Jinxiao Zhang, Runmin Dong, Juepeng Zheng, Mengxuan Chen, Lixian Zhang 0002, Yi Zhao 0024, Haohuan Fu |
ACM Multimedia | 7 |
| 2024 | FUSU: A Multi-temporal-source Land Use Change Segmentation Dataset for Fine-grained Urban Semantic UnderstandingabstractFine urban change segmentation using multi-temporal remote sensing images is essential for understanding human-environment interactions in urban areas. Although there have been advances in high-quality land cover datasets that reveal the physical features of urban landscapes, the lack of fine-grained land use datasets hinders a deeper understanding of how human activities are distributed across landscapes and the impact of these activities on the environment, thus constraining proper technique development. To address this, we introduce FUSU, the first fine-grained land use change segmentation dataset for Fine-grained Urban Semantic Understanding. FUSU features the most detailed land use classification system to date, with 17 classes and 30 billion pixels of annotations. It includes bi-temporal high-resolution satellite images with 0.2-0.5 m ground sample distance and monthly optical and radar satellite time series, covering 847 km^2 across five urban areas in the southern and northern of China with different geographical features. The fine-grained land use pixel-wise annotations and high spatial-temporal resolution data provide a robust foundation for developing proper deep learning models to provide contextual insights on human activities and urbanization. To fully leverage FUSU, we propose a unified time-series architecture for both change detection and segmentation. We benchmark FUSU on various methods for several tasks. Dataset and code are available at: https://github.com/yuanshuai0914/FUSU. Shuai Yuan 0005, Guancong Lin, Lixian Zhang 0002, Runmin Dong, Jinxiao Zhang, Juepeng Zheng, Jie Wang 0036, Haohuan Fu |
NeurIPS | 9 |
| 2024 | O2ath: an OpenMP offloading toolkit for the sunway heterogeneous manycore platform
Lifeng Yan, Qixin Chang, Haitian Lu, Chenlin Li, Quanjie He, Xiaohui Duan, Zekun Yin, Wei Xue 0003, Haohuan Fu, Lin Gan 0001, Guangwen Yang 0002 |
CCF Trans. High Perform. Comput. | 13 |
| 2024 | C³DA: A Universal Domain Adaptation Method for Scene Classification From Remote Sensing ImageryabstractVarious remote sensing applications have widely used domain adaptation (DA) methods. Since it does not need to add human interpretation in the target domain, it can be used in cross-region, multi-temporal, and multi-sensor application scenarios. In order to further optimize the design of the loss function and better address the challenges of DA in remote sensing, in this paper, we propose a new universal DA method named C3DA for scene recognition of remote sensing images. It has a comprehensive C3criterion for recognizing the "unknown" classes by innovatively fusing confidence, consistency, and certainty of samples to make our network training more efficient. We evaluate the performance of our proposed method based on six transfer tasks on three remote sensing datasets. The evaluation results show that our proposed method achieves an average H-score of 58.44%, significantly higher than other SOTA universal DA methods with an average improvement of 2.32~29.43%. Compared to the baseline ResNet-50, it achieves up to 19.92% improvement, demonstrating that the proposed method outperforms in the universal DA scenario. In the future, we also plan to expand the application of this method to more scenarios. Jiaxu Guo, Yushan Lai, Jinxiao Zhang, Juepeng Zheng, Haohuan Fu, Lin Gan 0008, Liang Hu 0001, Gaochao Xu, Xilong Che |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2024 | Relational Part-Aware Learning for Complex Composite Object Detection in High-Resolution Remote Sensing ImagesabstractIn high-resolution remote sensing images (RSIs), complex composite object detection (e.g., coal-fired power plant detection and harbor detection) is challenging due to multiple discrete parts with variable layouts leading to complex weak inter-relationship and blurred boundaries, instead of a clearly defined single object. To address this issue, this article proposes an end-to-end framework, i.e., relational part-aware network (REPAN), to explore the semantic correlation and extract discriminative features among multiple parts. Specifically, we first design a part region proposal network (P-RPN) to locate discriminative yet subtle regions. With butterfly units (BFUs) embedded, feature-scale confusion problems stemming from aliasing effects can be largely alleviated. Second, a feature relation Transformer (FRT) plumbs the depths of the spatial relationships by part-and-global joint learning, exploring correlations between various parts to enhance significant part representation. Finally, a contextual detector (CD) classifies and detects parts and the whole composite object through multirelation-aware features, where part information guides to locate the whole object. We collect three remote sensing object detection datasets with four categories to evaluate our method. Consistently surpassing the performance of state-of-the-art methods, the results of extensive experiments underscore the effectiveness and superiority of our proposed method. Shuai Yuan 0005, Lixian Zhang 0002, Runmin Dong, Juepeng Zheng, Haohuan Fu, Peng Gong 0002 |
IEEE Trans. Cybern. | 6 |
| 2024 | Weakly Supervised 3-D Building Reconstruction From Monocular Remote Sensing Imagesabstract3D building reconstruction from monocular remote sensing imagery is an important research problem that has been extensively studied for several decades. Although monocular remote sensing imagery is a more economic data source compared with the LiDAR data and multi-view imagery, its limited information results in great challenges and restricts the performance of existing monocular reconstruction methods. Moreover, the expensive cost and the limited quantity of 3D annotations also restrict the application scenes of existing methods, which are mostly based on fully-supervised learning. In our previous work, we have proposed MTBR-Net, a monocular building reconstruction method that consists of a fully-supervised multi-task network and a post-processing module for optimizing the reconstruction results. In this work, we further propose WS-MTBR-Net, a weakly-supervised building reconstruction network that uses fewer 3D annotations and achieves better performance in an end-to-end manner. Specifically, our WS-MTBR-Net fully leverages the relation between different components of a 3D building instance and the property of off-nadir images to improve the footprint segmentation boundary, based on six modified tasks and a new network structure with an improved feature warping module to support weakly-supervised learning. We also design a new training strategy via a hybrid loss function that enables utilizing the training samples with different annotation levels, i.e., complete 3D annotations, 2D footprint annotations, and image-level angle annotations. Results on BONAI Shanghai and Xi’an test datasets demonstrate that our method achieves competitive performance when using 50% fewer 3D-annotated samples, and improves the footprint segmentation F1-score by around 4% compared with current state-of-the-art. Zhenghao Hu, Lingxuan Meng, Jinwang Wang, Juepeng Zheng, Runmin Dong, Conghui He, Gui-Song Xia, Haohuan Fu, Dahua Lin |
IEEE Trans. Geosci. Remote. Sens. | 9 |
| 2024 | HyUniDA: Breaking Label Set Constraints for Universal Domain Adaptation in Cross-Scene Hyperspectral Image ClassificationabstractAlthough enormous Domain Adaptation (DA) approaches have been proposed for cross-scene hyperspectral image (HSI) classification, majority DA methods strongly depend on much prior knowledge of the association among the label sets of source and target domains (encompassing closed set, partial and open set DA), thereby significantly hindering their applications. Realistic application scenarios often require knowledge transfer between domains without restrictions on the label space, which is called Universal Domain Adaptation (UniDA). In this paper, we propose HyUniDA, which is the first attempt to address UniDA scenario from HSIs. HyUniDA contains two major parts: the Shared Semantic Pairing (SSP) and Domain Similarity Score (DSS). We group both source and target domains to form discriminative clusters. The SSP identifies pairs of clusters that have coincident semantic features as the common classes. By examining the consistency level of samples across source and target domains, DSS can estimate the quantity of target clusters and generate distinct clusters without prior knowledge. Meanwhile, we apply the contrastive domain discrepancy to alleviate the offset of samples distribution, with a representative regularizer to assist distinguish target domain clusters. We evaluate our proposed method on three transfer learning tasks for six typical HSI datasets, it turns out that our proposed method yields 3.83%~37.57% improvements compared to other state-of-the-art DA methods. Qingmei Li, Yibin Wen, Juepeng Zheng, Haohuan Fu |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2024 | Acceleration of Multi-Body Molecular Dynamics With Customized Parallel DataflowabstractFPGAs are drawing increasing attention in resolving molecular dynamics (MD) problems, and have already been applied in problems such as two-body potentials, force fields composed of these potentials, etc. Competitive performance is obtained compared with traditional counterparts such as CPUs and GPUs. However, as far as we know, FPGA solutions for more complex and real-world MD problems, such as multi-body potentials, are seldom to be seen. This work explores the prospects of state-of-the-art FPGAs in accelerating multi-body potential. An FPGA-based accelerator with customized parallel dataflow that features multi-body potential computation, motion update, and internode communication is designed. Major contributions include: (1) parallelization applied at different levels of the accelerator; (2) an optimized dataflow mixing atom-level pipeline and cell-level pipeline to achieve high throughput; (3) a mixed-precision method using different precision at different stages of simulations; and (4) a communication-efficient method for internode communication. Experiments show that, our single-node accelerator is over 2.7× faster than an 8-core CPU design, performing 20.501 ns/day on a 55,296-atom system for theTersoffsimulation. Regarding power efficiency, our accelerator is 28.9× higher than I7-11700 and 4.8× higher than RTX 3090 when running the same test case. Quan Deng 0001, Qiang Liu 0011, Xiaohui Duan, Lin Gan 0008, Jinzhe Yang, Wenlai Zhao, Zhenxiang Zhang, Guiming Wu, Wayne Luk, Haohuan Fu, Guangwen Yang 0002 |
IEEE Trans. Parallel Distributed Syst. | 11 |
| 2024 | SunwayLB: Enabling Extreme-Scale Lattice Boltzmann Method Based Computing Fluid Dynamics Simulations on Advanced Heterogeneous SupercomputersabstractThe Lattice Boltzmann Method (LBM) is a class of Computational Fluid Dynamics methods which models the fluid as fictive particles. In this paper, we report our work on SunwayLB, which enables LBM based solutions aiming for industrial applications using advanced heterogeneous systems such as the Sunway supercomputers. We propose several techniques to boost the simulation speed and improve the scalability of SunwayLB, including a customized multi-level domain decomposition and data sharing scheme, a carefully orchestrated strategy to fuse kernels with different performance constraints for a more balanced workload, and optimization strategies for assembly code. Based on these optimization schemes, we manage to scale SunwayLB on three advanced supercomputers: Sunway TaihuLight, the new Sunway Supercomputer and a GPU cluster. On Sunway TaihuLight, our largest simulation involves up to 5.6 trillion lattice cells, achieving 11,245 billion cell updates per second (GLUPS), 77% memory bandwidth utilization and a sustained performance of 4.7 PFlops. We further improve the memory bandwidth utilization and computational efficiency using the unique features of a new generation of Sunway supercomputer. On the new Sunway Supercomputer, the largest simulation contains over 4.2 trillion lattice cells, resulting in 6,583 GLUPS, 81% memory bandwidth utilization and a sustained performance of 2.76 PFlops. To evaluate the portability of our code, we also adapt our code to a GPU cluster with tailored optimization techniques, resulting in 191x speedup and 83.8% memory bandwidth utilization. We demonstrate a series of computational experiments for extreme-large scale fluid flow, as examples of real-world applications, to check the validity and performance of our work. The results show that our implementation is competent to be a highly scalable and efficient solution for large-scale CFD problems on heterogeneous systems. Xuesen Chu, Xiaojing Lv, Hongsong Meng, Haohuan Fu, Guangwen Yang 0002 |
IEEE Trans. Parallel Distributed Syst. | 7 |
| 2023 | SetTron: Towards Better Generalisation in Penetration Testing with Reinforcement LearningabstractIntelligent penetration testing (pen-testing), utilising Deep Reinforcement Learning (DRL) has gained attention due to its potential for improving testing efficiency and cost-effectiveness in evaluating network system security. Nonetheless, current approaches which rely on simplistic neural network architectures suffer limitations in transferability and their ability to generalise to new tasks, thus impeding their practical application. This paper aims to address these issues by formalising the pen-testing decision process as a Host-Centric Markov decision process (HC- MDP), as well as establishing a structural representation of the relationships among the hosts within a network system. Further, we propose a flexible policy architecture, the “SetTron”, that leverages this structural representation to augment architectural inductive bias in a DRL agent and then practically evaluate our approach on pen-testing simulator platforms. The findings show SetTron to demonstrate superior performance, in terms of learning efficiency and policy convergence, compared to state-of-the-art methods and baselines with shorter penetration sequences and enhanced rewards. Besides, SetTron exhibits remarkable zero-shot generalisation capabilities, enabling perfect transfer to new tasks with randomly placed target hosts, achieving a 100 % success rate, and outperforming baselines by a factor of 6 when comparing normalised scores. Yizhou Yang, Mengxuan Chen, Haohuan Fu, Xin Liu 0081 |
GLOBECOM | 3 |
| 2023 | Large-Scale Land Cover Mapping with Fine-Grained Classes via Class-Aware Semi-Supervised Semantic SegmentationabstractSemi-supervised learning has attracted increasing attention in the large-scale land cover mapping task. However, existing methods overlook the potential to alleviate the class imbalance problem by selecting a suitable set of unlabeled data. Besides, in class-imbalanced scenarios, existing pseudo-labeling methods mostly only pick confident samples, failing to exploit the hard samples during training. To tackle these issues, we propose a unified Class-Aware Semi-Supervised Semantic Segmentation framework. The proposed framework consists of three key components. To construct a better semi-supervised learning dataset, we propose a class-aware unlabeled data selection method that is more balanced towards the minority classes. Based on the built dataset with improved class balance, we propose a Class-Balanced Cross Entropy loss, jointly considering the annotation bias and the class bias to re-weight the loss in both sample and class levels to alleviate the class imbalance problem. Moreover, we propose the Class Center Contrast method to jointly utilize the labeled and unlabeled data. Specifically, we decompose the feature embedding space using the ground truth and pseudo-labels, and employ the embedding centers for hard and easy samples of each class per image in the contrast loss to exploit the hard samples during training. Compared with state-of-the-art class-balanced pseudo-labeling methods, the proposed method improves the mean accuracy and mIoU by 4.28% and 1.70%, respectively, on the large-scale Sentinel-2 dataset with 24 land cover classes. Runmin Dong, Lichao Mou, Mengxuan Chen, Xin-Yi Tong 0003, Shuai Yuan 0005, Lixian Zhang 0002, Juepeng Zheng, Xiao Xiang Zhu 0001, Haohuan Fu |
ICCV | 10 |
| 2023 | Accelerating Large-Scale CFD Simulations with Lattice Boltzmann Method on a 40-Million-Core Sunway SupercomputerabstractThe Lattice Boltzmann Method (LBM) has gained widespread popularity due to its applicability in fluid dynamics, chemical engineering, material science, and other domains. In this work, we present an optimized implementation of the LBM, with a specific focus on achieving superior performance and scalability on advanced heterogeneous systems such as the new Sunway supercomputer. To accomplish this, we employ several techniques, including kernel fusion to enhance temporal and spatial locality, a customized multi-level domain decomposition and data sharing scheme, and pipelining strategies that are tailored to the SW26010-Pro processor. As a result of these optimizations, we have successfully scaled our code to a total of 39,000,000 CPU cores. Our largest simulation, which encompassed over 42 trillion lattice cells, achieved an impressive 67,018 billion lattice cell updates per second (GLUPS), with 82.9% memory bandwidth utilization, and a sustained performance of 28 PFlops. In order to assess the portability of our implementation, we also adapted our code to run on a GPU cluster, utilizing a range of tailored optimization techniques. Our results demonstrated a 191x speedup, along with 83.8% memory bandwidth utilization. Our proposed approach marks a significant milestone in the field of LBM implementations, as it demonstrates unprecedented scalability by effectively utilizing over 39,000,000 cores while maintaining exceptional parallel efficiency and computational performance. This achievement establishes our method as a compelling solution for addressing large-scale computational fluid dynamics challenges on heterogeneous systems. Xuesen Chu, Xiaojing Lv, Haohuan Fu, Guangwen Yang 0002 |
ICPP | 5 |
| 2023 | Fusing Time-Inconsistent Sentinel-2 Images and High-Resolution Remote Sensing ImagesabstractOne of superiorities of satellite imagery is quickly monitoring the atmosphere, land, and ocean at a large scale [1], [2]. However, due to the trade-off between spatial coverage and resolution, satellite images with a broader range usually have a lower spatial resolution [3]. Owing to the advancement of satellite remote sensing, multispectral images can now get a high resolution such as 10 m, benefiting various applications including land cover mapping, water resource management, and agricultural monitoring. Furthermore, on account of high revisit frequency (every 5 days), relatively high spatial resolution (up to 10 m), and free access, Sentinel-2 imagery has received widespread attention and become an important resource for earth system studies [4]. Runmin Dong, Haohuan Fu |
IGARSS | 2 |
| 2023 | CO-Detector: Towards Complex Object Detection with Cross-Part Feature Learning in Remote SensingabstractObject detection in remote sensing imagery builds the essential foundation of aerial and satellite image understanding, being an important role in many common real-world tasks and attracting world-wide attention. In recent years, despite the great progress of common object detection in remote sensing and the proven success of deep learning in this field, yet complex object detection which consists of multiple objects with variable layouts in remote sensing (e.g., coal-fired power plant, airport, sewage treatment plant, etc.) is still challenging for complex composite spatial relationship, non-rigid boundaries, and complicated surrounding textures. These challenges necessitate developing specific complex object detection methods to learn inter-relationship and distinctive and discriminative features in complex objects. To address this problem, in this paper, we propose a method, i.e., CO-Detector, in an end-to-end manner, to achieve various complex composite object detection in remote sensing images with high accuracy and efficiency. The effectiveness of CO-Detector is built on three main parts: (a) First, as surrounding contexts are normally complicated and similar to complex objects, we propose a Tandem Attention Network (TAN), including a channel enhanced network and a spatial enhanced network, with a K-global max/average pooling, to restrain noise disturbance and highlight complex object features and boundaries. (b) Second, we design a Part Region Proposal Network (P-RPN) to learn the interrelationship between parts in one object, generating part proposals and locating discriminative and distinctive object parts finely. (c) Third, to detect the whole complex object as well as the parts, we propose a Part Detection Network (PDN) to detect the individual parts, and detect the whole object through multi-level fused features. We train our CO-Detector model with three selected categories (i.e., coal-fired power plant, airport, oil storage tank) in three datasets, and conduct comparative experiments to evaluate and verify the performance. The comprehensive experiment results show that our CO-Detector achieves a mAP of 80.23%, outperforming 4.17%-17.83% against other cutting-edge deep learning-based detection methods. The experiment results indicate our CO-Detector has promising performance and potential in various complex object detection in highresolution remote sensing images, pending to be utilized in real large-scale applications. Shuai Yuan 0005, Juepeng Zheng, Jierui Liu, Haohuan Fu, Ray C. C. Cheung |
IGARSS | 5 |
| 2023 | Achieving 10m China Land Cover Mapping within Three Minutes Using a New Sunway SupercomputerabstractLand Cover Mapping (LCM) is an important task to detect and understand the change of the earth surface. However, most LCM methods adopt supervised classifiers, and suffer from a lack of labels at a large scale. In this paper, we propose Fast-LCM, a scalable and weakly-supervised LCM method on a new Sunway supercomputer to achieve large-scale land cover mapping, requiring no manual annotations. Fast-LCM includes two major parts: (1) a distance-guided k-means module that combines textural, spectral, and temporal features, and (2) an automatic voting-based merging strategy to give each cluster a real meaning of classification system. Through careful parallelization, our Fast-LCM method scales to over 38 million cores, and provides a sustained performance for the task of China LCM. We produce a 10m resolution land cover map of China within only 3 minutes, including 1.2 minutes for IO and only 55 seconds to finish the computation. Fast-LCM achieves an accuracy of 68.85% (25-class), with 3.64% to 7.05% higher than best existing products. Juepeng Zheng, Yi Zhao 0024, Jinxiao Zhang, Wenzhao Wu, Shuai Yuan 0005, Haohuan Fu |
IGARSS | 6 |
| 2023 | SW-LCM: A Scalable and Weakly-supervised Land Cover Mapping Method on a New Sunway SupercomputerabstractHigh-resolution land cover mapping (LCM) is an important application for studying and understanding the change of the earth surface. While deep learning (DL) methods demonstrate great potential in analyzing satellite images, they largely depend on massive high-quality labels. This paper proposes SW-LCM, a Scalable and Weakly-supervised two-stage Land Cover Mapping method on a new Sunway Supercomputer. Our method consists of a k-means clustering module as a first stage, and an iterative deep learning module as a second stage. With the k-means module providing a good enough starting point (taking inaccurate results as noisy labels), the deep learning module improves the classification results in an iterative way, without any labelling efforts required for processing large scenarios. To achieve efficiency for country-level land cover mapping, we design a customized data partition scheme and an on-the-fly assembly for k-means. Through careful parallelization and optimization, our k-means module scales to 98,304 computing nodes (over 38 million cores), and provides a sustained performance of 437.56 PFLOPS, in a real LCM task of the entire region of China; the iterative updating part scales to 24,576 nodes, with a performance of 11 PFLOPS. We produce a 10-m resolution land cover map of China, with an accuracy of 83.5% (10-class) or 73.2% (25-class), 7% to 8% higher than best existing products, paving ways for finer land surveys to support sustainability-related applications. Yi Zhao 0024, Juepeng Zheng, Haohuan Fu, Wenzhao Wu, Mengxuan Chen, Jinxiao Zhang, Lixian Zhang 0002, Runmin Dong, Zhenrong Du, Xin Liu 0081, Shaoqing Zhang, Le Yu 0001 |
IPDPS | 3 |
| 2023 | Lifetime-Based Optimization for Simulating Quantum Circuits on a New Sunway SupercomputerabstractHigh-performance classical simulator for quantum circuits, in particular the tensor network contraction algorithm, has become an important tool for the validation of noisy quantum computing. In order to address the memory limitations, the slicing technique is used to reduce the tensor dimensions, but it could also lead to additional computation overhead that greatly slows down the overall performance. This paper proposes novel lifetime-based methods to reduce the slicing overhead and improve the computing efficiency, including, an interpretation method to deal with slicing overhead, an inplace slicing strategy to find the smallest slicing set and an adaptive tensor network contraction path refiner customized for Sunway architecture. Experiments show that in most cases the slicing overhead with our inplace slicing strategy would be less than the Cotengra, which is the most used graph path optimization software at present. Finally, the resulting simulation time is reduced to 96.1s for the Sycamore quantum processor RQC, with a sustainable single-precision performance of 308.6Pflops using over 41M cores to generate 1M correlated samples, which is more than 5 times performance improvement compared to 60.4 Pflops in 2021 Gordon Bell Prize work. Yaojian Chen, Xinmin Shi, Jiawei Song, Xin Liu 0081, Lin Gan 0001, Chu Guo, Haohuan Fu, Dexun Chen, Guangwen Yang 0002 |
PPoPP | 8 |
| 2023 | Enabling Real World Scale Structural Superlubricity All-Atom Simulation on the Next-Generation Sunway SupercomputerabstractMolecular dynamics (MD) simulation can provide an affordable way for inspecting microscopic phenomena, which is a powerful complement to real-world experiments. But the spatial scale of MD simulations is usually magnitudes smaller than experiment systems. In this paper, we present our work, redesigning the widely used inter-layer potential in structural superlubricity. By carrying out a specialized neighbor list for inter-layer potential computation, the total memory access amount is reduced significantly. Besides, a simple but efficient vectorization strategy is implemented based on the new neighbor list. In the extreme case, our work can scale to 38 million cores to achieve a sustainable performance of 61 PFLOPS, enabling a simulation of a superlubricity system of 32 μm2 with 7.2 billion atoms at 4.75 ns/day, which is 11,834 times of reported largest scale simulation in superlubricity systems in contact area and almost ten times faster in time-to-solution. Furthermore, we have done a simulation at 9 μm2 which results in consistency with real-world experiments and verified some theoretical predictions in the mesoscopic scale. Xiaohui Duan, Ping Gao 0005, Ming Ma 0012, Lin Gan 0001, Xin Liu 0081, Haohuan Fu, Wei Xue 0003, Dexun Chen, Guangwen Yang 0002 |
SC | 7 |
| 2023 | 69.7-PFlops Extreme Scale Earthquake Simulation with Crossing Multi-faults and Topography on SunwayabstractA high-scalable and fully optimized earthquake model is presented based on the latest Sunway supercomputer. Contributions include: 1) the curvilinear grid finite-difference method (CGFDM) and flexible model applying perfectly matched layer (PML) and enabling more accurate and realistic terrain descriptions; 2) a hybrid and non-uniform domain decomposition scheme that efficiently maps the model across different levels of the computing system; and 3) sophisticated optimizations that largely alleviate or even eliminate bottlenecks in memory, communication, etc., obtaining a speedup of over 140×. Combining all innovations, the design fully exploits the hardware potential of all aspects and enables us to perform the largest CGFDM-based earthquake simulation ever reported (69.7 PFlops using over 39 million cores). Based on our design, the Turkey earthquakes (February 6, 2023), and the Ridgecrest earthquake (July 4, 2019), are successfully simulated with a maximum resolution of 12-m. Precise hazard evaluations for the hazardous reduction of earthquake-stricken areas are also conducted. Wubing Wan, Lin Gan 0001, Zekun Yin, Haodong Tian, Mengyuan Hua, Shengye Xiang, Zhongqiu He, Ping Gao 0005, Xiaohui Duan, Wei Xue 0003, Haohuan Fu, Guangwen Yang 0002, Yaojian Chen, Xin Liu 0081, Wei Zhang 0321 |
SC | 17 |
| 2023 | GEO-WMS: an improved approach to geoscientific workflow management system on HPC
Jiaxu Guo, Yidan Xu, Haohuan Fu, Wei Xue 0003, Lin Gan 0008, Mengxuan Tan, Tingye Wu, Yutong Shen, Xianwei Wu, Liang Hu 0001, Xilong Che |
CCF Trans. High Perform. Comput. | 3 |
| 2023 | Partial Domain Adaptation for Scene Classification From Remote Sensing ImageryabstractAlthough domain adaptation approaches have been proposed to tackle cross-regional, multitemporal, and multisensor remote sensing applications since they do not require any human interpretation in the target domain, most current works assume identical label space across the source and the target domains. However, in real-world applications, we often transfer knowledge from a large-scale dataset with rich annotations to a small-scale target dataset with scarcity of labels. In most cases, the label space of the source domain is usually large enough to subsume that of the target domain, which is termed partial domain adaptation. In this article, we propose a new partial domain adaptation algorithm for remote sensing scene classification and our proposed method contains three major parts. First, we employ a progressive auxiliary domain module to alleviate the negative transfer effect caused by outlier classes. Second, we adopt an improved domain adversarial neural network (DANN) with multiweights to better encourage domain confusion. Last but not least, we design an attentive complement entropy regularization to improve the prediction confidence for samples and avoid untransferable samples (such as the samples belonging to outlier classes in the source domain) being mistakenly classified. We collect three common remote sensing datasets to evaluate our proposed method. Our method achieves an average accuracy of 79.36%, which considerably outperforms other state-of-the-art partial domain adaptation methods with an average accuracy improvement of 1.90%–12.45% and attaining a 13.67% gain compared to the straightforward deep learning model (ResNet-50). The experiment results indicate that our approach shows promising prospects for solving more general and practical domain adaptation problems where the label space of the source domain subsumes that of the target domain. Juepeng Zheng, Yi Zhao 0024, Wenzhao Wu, Mengxuan Chen, Haohuan Fu |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2023 | Bio-ESMD: A Data Centric Implementation for Large-Scale Biological System Simulation on Sunway TaihuLight SupercomputerabstractMolecular dynamics (MD) simulations of biological systems are playing an increasingly important role in the research of pathogens and drugs. Most MD methods for biological simulations rely on the listed bonds which interact among specific groups of atoms identified by atom tags (unique atom tags regardless the storage location). However, efficient mapping of tags to atom locations is often challenging on modern many-core processors because data locality can not always be guaranteed for large-scale systems. In this paper, we present Bio-ESMD, a new MD implementation supporting listed bonds. Bio-ESMD is designed and developed based on our previously designed ESMD framework for many-core processors. In Bio-ESMD, we have introduced a data-centric approach for refactoring MD algorithms by reorganizing the cell list data structure to adopt bond lists with guaranteed data locality. Our implementation achieves speedups of over two compared to SW_GROMACS on Sunway TaihuLight. Furthermore, Bio-ESMD can simulate a system of 308.8 million atoms at 1.33 ns/day or 14.44 million atoms at 17.28 ns/day with linear weak scaling efficiency. Xiaohui Duan, Junben Weng, Bertil Schmidt, Lin Gan 0001, Haohuan Fu, Wei Xue 0003, Guangwen Yang 0002 |
IEEE Trans. Parallel Distributed Syst. | 7 |
| 2023 | Redesign and Accelerate the AIREBO Bond-Order Potential on the New Sunway SupercomputerabstractMolecular dynamics (MD) is one of the most crucial computer simulation methods for understanding real-world processes at the atomic level. Reactive potentials based on the bond order concept have the ability to model dynamic bond breaking and formation with close to quantum mechanical (QM) precision without actually requiring expensive QM calculations. In this article, we focus on the adaptive intermolecular reactive empirical bond-order (AIREBO) potential in LAMMPS for the simulation of carbon and hydrocarbon systems on the new Sunway supercomputer. To achieve scalable performance, we propose a parallel two-level building scheme and periodic buffering strategy for the tailored data design to explore data locality and data reuse. Furthermore, we design two optimized nearest-neighbor access algorithms: the redistribution of accumulated coefficients algorithm and the double-end search connectivity algorithm. Finally, we implement parallel force computation with an AoS data layout and hardware/software co-cache. In addition, we have designed a low-overhead atomic operation-based load balancing method and vectorization. The overall performance of AIREBO achieves a speedup of nearly$20\times$on a single core group (CG), and more than$5\times$and$4\times$over an Intel Xeon E5 2680 v3 core and an Intel Xeon Gold 6138 core, respectively. Compared with the Intel accelerator package in LAMMPS, our performance further achieves$3.0\times$of an Intel Xeon E5 2680 v3 core and is better than that of an Intel Xeon Gold 6138 core. We complete the validation of the results in no more than 20.5 hours on a single node with 2,000,000 running steps (i.e., 1 ns). Our experiments show that the simulation of 2,139,095,040 atoms on 798,720 ((1MPE+64CPEs) × 12,288 processes) cores exhibits a parallel efficiency of 88% under weak scaling. Ping Gao 0005, Xiaohui Duan, Bertil Schmidt, Wubing Wan, Jiaxu Guo, Wusheng Zhang, Lin Gan 0008, Haohuan Fu, Wei Xue 0003, Guangwen Yang 0002 |
IEEE Trans. Parallel Distributed Syst. | 8 |
| 2022 | Melting Glacier: A 37-Year (1984-2020) High-Resolution Glacier-Cover Record of MT. KilimanjaroabstractCommonly recognized as an important symbol of the tropics and global warming, the glacier loss on Mt. Kilimanjaro has received worldwide attention for decades. In this paper, we propose a high-resolution glacier-cover (GC) record of Mt. Kilimanjaro over the period from 1984 to 2020, using a novel deep learning-based semantic segmentation method and Google Earth images, as well as digital elevation model (DEM) and ERA5-Land (ERA5) for snowline and temperature variations analysis. Our method achieves an accuracy of 94.37%, which proves the model's capability to record the GC areas precisely. The results show that (1) the GC area dramatically decreases from 19.2 km2to 3.6 km2during 37 years, which decreases about 4% and 2% per year from 1984 to 2000 and from 2000 to 2020 respectively, (2) the snowline altitude rises from$4,651 m$to$5,088 m$by about$437 m$, and (3) the average$5,000 m$air temperature on Mt. Kilimanjaro increases from −2.1 °C to −1.1 °C by about 1 °C. This study indicates that there will be no GC within a few decades if the current loss continues. Shuai Yuan 0005, Juepeng Zheng, Lixian Zhang 0002, Runmin Dong, Yile Xing, Yuhan She, Haohuan Fu, Ray C. C. Cheung |
IGARSS | 7 |
| 2022 | Srbuildingseg-E2: An Integrated Model for End-to-End Higher-Resolution Building ExtractionabstractAutomatic and accurate extraction of buildings from remote sensing images plays a vital role in many applications. However, existing approaches for building extraction generally apply high-resolution remote sensing images as input to attain high-resolution extraction results, which is time consuming and limited due to its spatiotemporal accessibility and cost. To address this challenge, in this paper, we propose an end-to-end approach, i.e., SRBuildingSeg-E2, to achieve higher-resolution building extraction from relatively low resolution remote sensing images. By integrating super resolution and semantic segmentation techniques, our proposed approach can attain high-resolution representations using low-resolution input. The quantitative assessment results reveal its promising performance in higher-resolution building extraction. Lixian Zhang 0002, Runmin Dong, Shuai Yuan 0005, Haohuan Fu |
IGARSS | 4 |
| 2022 | A Parallel Approach for Oil Palm Tree Detection on a SW26010 Many-Core ProcessorabstractCounting and detecting oil palm trees from high-resolution remotely sensed images is a significant work for improving economy of several countries such as Malaysia, Indonesia, etc. However, rare attention have been paid on accelerating tree crown detection algorithms on high performance platforms. In this paper, we design a parallel approach for oil palm tree detection on a SW26010 many-core processor, which is used in a world-leading supercomputer, Sunway TaihuLight. Our parallel framework contains three steps, local maximum filtering, oil palm tree crown center reassignment and oil palm tree crown center merging. Experimental results indicates that our parallel approach of oil palm tree detection obtains the speedup of 32.30 times and 1.74 times for a QuickBird image with a size of$12,188\times 12,576$pixels compared with the well-optimized software implementation of the original algorithm on an Intel 12-core CPU and FPGAs. Juepeng Zheng, Wenzhao Wu, Yi Zhao 0024, Shuai Yuan 0005, Runmin Dong, Lixian Zhang 0002, Haohuan Fu |
IGARSS | 7 |
| 2022 | A fully-customized dataflow engine for 3D earthquake simulation with a complex topography
Bingwei Chen, Haohuan Fu, Wayne Luk, Guangwen Yang 0002 |
Sci. China Inf. Sci. | 2 |
| 2022 | Multisource-Domain Generalization-Based Oil Palm Tree Detection Using Very-High-Resolution (VHR) Satellite ImagesabstractProviding accurate and timely oil palm information on a large scale is essential for both economic development and ecological significance. However, owing to different sensors, photograph acquisition conditions, and environmental heterogeneity, the large volume and the variety of the data make it extremely challenging for large-scale and cross-regional oil palm tree detection. It is computationally expensive to train a model from images covering large heterogeneous regions and all environmental conditions for continuously accumulated multisource remote sensing data. In this letter, we propose a new multisource domain generalization (DG) method, Maximum Mean Discrepancy Deep Reconstruction Classification Network (MMD-DRCN). It learns representations from multiple source domains and obtains inspiring performance in an unknown and “unseen” target domain. Besides classification loss, our MMD-DRCN distills more representative features through reconstruction loss and aligns multisource latent features by MMD loss, both of which effectively enhance the capacity of generalization. MMD-DRCN achieves an average F1-score of 82.70% in all transfer tasks, attaining a 5.83% gain compared to Baseline (a straightforward convolutional neural network (CNN) model). Experimental results demonstrate DG poses a promising potential for large-scale and cross-regional oil palm tree detection without any information of the target domain. Juepeng Zheng, Wenzhao Wu, Shuai Yuan 0005, Haohuan Fu, Le Yu 0001 |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2022 | Enabling Large-Scale Simulation of CAM on the Sunway TaihuLight SupercomputerabstractThe Community Atmosphere Model (CAM) has been ported, redesigned, and scaled to the full system of the Sunway TaihuLight, and provides peta-scale climate modeling performance. Based on a novel domain decomposition method, we have fully optimized the complete model code by using both OpenACC refactoring and more aggressive and finer-grained Athread approaches. The Athread approach enables us to achieve exceptional memory control and usage, efficient vectorization, and sophisticated utilization of the thread-level communication mechanism. We have also further refined the load-balance behaviors towards ultra-large-scale numerical simulation. By combining all these novelties, we achieved a simulation speed of 7.2 and 25.6 simulation-year-per-day (SYPD) for global 25-km and 100-km resolution, respectively (1.2- to 2.2-fold improvements over previous efforts), and a sustainable double-precision performance of 3.3 PFlops for a 750-m global simulation when using 10075000 cores. Xiaohui Duan, Lin Gan 0001, Wubing Wan, Yuhu Chen, Jinzhe Yang, Wei Xue 0003, Haohuan Fu, Guangwen Yang 0002 |
IEEE Trans. Computers | 10 |
| 2022 | High-Resolution Land Cover Mapping Through Learning With Noise CorrectionabstractHigh-resolution land cover mapping over large areas is a challenging task due to the lack of high-quality labels. A potential solution is to leverage the existing knowledge contained in the freely available lower-resolution land cover products. However, the relatively low resolution and low accuracy of the products lead to numerous inaccurate labels, which harms the performance of the neural network. This article addresses the challenge by jointly optimizing the network parameters and correcting the noisy labels with a novel online noise correction approach and a synergistic noise correction loss. By incorporating the information entropy as a measurement to determine the probable correct labels, the proposed noise correction approach learns to make effective correction of the noisy labels during training and eventually boosts the performance with a training set containing less noisy labels. Experimental results show that the proposed method can effectively correct the noisy labels and reduce their negative impact on network training. By employing the proposed method, we produce a refined high-resolution (3-m) land cover map from a lower-resolution (10-m) product in China and improve the accuracy from 74.96% (10-m) to 81.32% (3-m). Such an approach that can effectively learn from noisy data sets leads to many potential opportunities for using and magnifying existing knowledge and results. Runmin Dong, Weizhen Fang, Haohuan Fu, Lin Gan 0001, Jie Wang 0036, Peng Gong 0002 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | RRSGAN: Reference-Based Super-Resolution for Remote Sensing ImageabstractRemote sensing image super-resolution (SR) plays an important role by supplementing the lack of original high-resolution (HR) images in the study scenarios of large spatial areas or long time series. However, due to the lack of imagery information in low-resolution (LR) images, single-image super-resolution (SISR) is an inherently ill-posed problem. Especially, it is difficult to reconstruct the fine textures of HR images at large upscaling factors (e.g., four times). In this work, based on Google Earth HR images, we explore the potential of the reference-based super-resolution (RefSR) method on remote sensing images, utilizing rich texture information from HR reference (Ref) images to reconstruct the details in LR images. This method can use existing HR images to help reconstruct the LR images of long time series or a specific time. We build a reference-based remote sensing SR data set (RRSSRD). Furthermore, by adopting the generative adversarial network (GAN), we propose a novel end-to-end reference-based remote sensing GAN (RRSGAN) for SR. RRSGAN can extract the Ref features and align them to the LR features. Eventually, the texture information in the Ref features can be transferred to the reconstructed HR images. In contrast to the existing RefSR methods, we propose a gradient-assisted feature alignment method that adopts the deformable convolutions to align the Ref and LR features and a relevance attention module (RAM) to improve the robustness of the model in different scenarios (e.g., land cover changes and cloud coverage). The experimental results demonstrate that RRSGAN is robust and outperforms the state-of-the-art SISR and RefSR methods in both quantitative evaluation and visual results, which indicates the great potential of the RefSR method for remote sensing tasks. Our code and data are available athttps://github.com/dongrunmin/RRSGAN. Runmin Dong, Lixian Zhang 0002, Haohuan Fu |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | Deep Learning-Based P- and S-Wave Separation for Multicomponent Vertical Seismic ProfilingabstractVertical seismic profiling (VSP) helps to derive high-resolution images around the instrumented borehole and is a cost-effective technique for CO2storage monitoring. In routine VSP data processing, P- and S-wave separation is a crucial step to extract independent single-mode waves for accurate imaging and interpretation. Conventional wave mode separation involves tedious, subjective, and non-reproducible manual interventions, especially when dealing with complex geology. To better automate the process, we propose a data-driven deep learning-based P- and S-wave separation method. Our method adapts a fully convolutional neural network that simultaneously extracts P- and S-potential data from multicomponent VSP measurements. To reduce the enormous computational cost in wave simulation while constructing training datasets with sufficient kinematic and dynamic variations, we introduce virtual wellbores where synthetic VSP data sampling wide variations in seismic kinematics and dynamics are recorded using only a dozen elastic wave simulations on a single velocity model. We qualify the separation results both directly in data space and in image space after reverse time migration (RTM). Generalization tests on various synthetic models and their corresponding RTM images demonstrate that the proposed strategy provides sufficient sampling of the high-dimensional data space and essentially ensures successful applications of the trained neural network to similar yet different geological scenarios. Yanwen Wei, Yunyue Elita Li, Jingjing Zong, Jizhong Yang, Haohuan Fu, Mengyao Sun 0002 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2022 | A Two-Stage Adaptation Network (TSAN) for Remote Sensing Scene Classification in Single-Source-Mixed-Multiple-Target Domain Adaptation (S²M²T DA) ScenariosabstractOver the past decade, domain adaptation (DA) algorithms have been proposed to address domain gap problems as they do not need any interpretation in the target domain. However, most existing efforts focus on scenarios with only one source domain and one target domain. In this article, we explore the scenario with one source domain and mixed multiple target domains for remote sensing applications and propose a new algorithm, named the two-stage adaptation network (TSAN). First, we utilize the adversarial learning approach to confuse the classifier to discriminate between the source domain and the whole mixed-multiple-target domain. Second, we adopt self-supervised learning to divide the mixed-multiple-target domain with automated generation of “pseudo”-domain labels, which guides our network to learn intrinsic features of multiple target domains. Finally, these two steps are combined as an iterative procedure. We integrate a test dataset that includes five remote sensing datasets and ten classes. Our method achieves an average accuracy of 63.25% and 73.68% with two typical backbones, considerably outperforming other DA methods with an average accuracy improvement of 4.84%–20.19% and 9.06%–17.04%, respectively. Furthermore, we identify the negative transfer effect in existing mainstream DA methods in remote sensing image classification with multiple different domains. Juepeng Zheng, Wenzhao Wu, Shuai Yuan 0005, Yi Zhao 0024, Lixian Zhang 0002, Runmin Dong, Haohuan Fu |
IEEE Trans. Geosci. Remote. Sens. | 8 |
| 2022 | Optimization of Reactive Force Field Simulation: Refactor, Parallelization, and Vectorization for InteractionsabstractMolecular dynamics (MD) simulations are playing an increasingly important role in many areas ranging from chemical materials to biological molecules. With the continuing development of MD models, the potentials are getting larger and more complex. In this article, we focus on the reactive force field (ReaxFF) potential from LAMMPS to optimize the computation of interactions. We present our efforts on refactoring for neighbor list building, bond order computation, as well as valence angles and torsion angles computation. After redesigning these kernels, we develop a vectorized implementation for non-bonded interactions, which is nearly 100 × faster than the management processing element (MPE) on the Sunway TaihuLight supercomputer. Furthermore, we have implemented the three-body-list free torsion angles computation, and propose a line-locked software cache method to eliminate write conflicts in the torsion angle and valence angle interactions resulting in an order-of-magnitude speedup on a single Sunway TaihuLight node. In addition, we achieve a speedup of up to 3.5 compared to the KOKKOS package on an Intel Xeon Gold 6148 core. When executed on 1,024 processes, our implementation enables the simulation of 21,233,664 atoms on 66,560 cores with a performance of 0.032 ns/day and a weak scaling efficiency of 95.71 percent. Ping Gao 0005, Xiaohui Duan, Bertil Schmidt, Wusheng Zhang, Lin Gan 0001, Haohuan Fu, Wei Xue 0003, Guangwen Yang 0002 |
IEEE Trans. Parallel Distributed Syst. | 6 |
| 2022 | Benchmarking 50-Photon Gaussian Boson Sampling on the Sunway TaihuLightabstractBoson sampling is expected to be an important milestone that will demonstrate quantum computational advantage (or quantum supremacy). This work establishes the benchmarking of Gaussian boson sampling (GBS) with threshold detection based on the Sunway TaihuLight supercomputer. To achieve the best performance and provide a competitive scenario for future quantum computing studies, the selected simulation algorithm is fully optimized based on a set of innovative approaches, including a parallel framework with almost perfect load balance and an instruction-level optimizing scheme based on a shortest-path-based instruction scheduling. In addition, data precision is carefully processed by an integer-instruction-based and multiple-precision fixed-point implementation, including 128- and 256-bit precison mode, which can be appropriately selected based on an adaptive precision optimizing scheme. Based on these methods, a highly efficient parallel quantum sampling algorithm is designed. The largest run enables us to obtain one Torontonian function of a$100\times 100$submatrix from 50-photon GBS within 20 hours in 128-bit precision and 2 days in 256-bit precision. To our knowledge, this was the largest quantum computing simulation based on Boson Sampling by using modern supercomputers. Lin Gan 0001, Mingcheng Chen, Yaojian Chen, Haitian Lu, Chao-Yang Lu, Jian-Wei Pan, Haohuan Fu, Guangwen Yang 0002 |
IEEE Trans. Parallel Distributed Syst. | 8 |
| 2021 | Transresnet: Transferable Resnet For Domain AdaptationabstractAlthough Deep Convolutional Neural Network (DCNN) has been admittedly witnessed as an enormous success in a wide range of applications, most of them require sufficient annotations with time-consuming and labor-exhausting efforts. Existing domain adaptation (DA) approaches delve into designing an effective loss module to minimize the distribution gap between the source and target domains. However, few studies pay attention to improve the backbone or network architecture for DA issues. In this paper, we propose a new backbone for DA specially, i.e., Transferable ResNet (TransResNet). TransResNet remedies the residual block in ResNet, separating source and target input features and highlighting more transferable channels in each block. It can be easily applied to all kinds of DA methods, without adding any extra learning parameters. We conduct substantial experiments on two general DA datasets and embed TransResNet into two seminal DA methods, including DANN and CDAN. Experimental results demonstrate TransResNet improves the transferability of the architecture, indicating that it is a great substitute for ResNet as a network backbone in DA issues. Juepeng Zheng, Wenzhao Wu, Yi Zhao 0024, Haohuan Fu |
ICIP | 4 |
| 2021 | Blind Super-Resolution on Remote Sensing Images with Blur Kernel PredictionabstractSingle image super-resolution (SISR) is essential in many remote sensing applications. Most of the existing SISR methods on remote sensing images assume that the low resolution (LR) images are synthesized from high-resolution (HR) images by bicubic downscaling. However, the performance of those methods is limited in the real-world remote sensing scenario as the actual degradation is sometimes different from the assumption. Therefore, we introduce the blind super-resolution (SR) concept and propose a super-resolution method with blur kernel prediction (BKPSR). BKPSR first predicts the blur kernel code for an image and then utilizes the blur kernel code to assist the image super-resolution. Experimental results indicate that our method outperforms existing SISR methods on real-world remote sensing images. Runmin Dong, Lixian Zhang 0002, Haohuan Fu |
IGARSS | 3 |
| 2021 | Sectoral Energy-Consumption Estimation by Unmixed Nighttime Light in Shanghai, ChinaabstractEnergy consumption management shapes the route to meet greenhouse gas emission reduction goals. Although certain types of or even total energy consumption monitor have been realized using nighttime light remote sensing data, there still leaves a gap to model energy consumption in different sectors timely and broadly. Using parcel oriented temporal linear unmixing method (POTLUM), nighttime light was decomposed into light from different land uses as sources. A pilot study in Shanghai shows that in each year, statistical energy consumptions in living, secondary and tertiary sectors correlate well with unmixed nighttime light from 2014 to 2018 (R2=0.75, 0.61, 0.24, 0.99, 0.71, respectively). This study proves the capability of unmixed nighttime light to estimate sectoral energy consumption using POTLUM, which helps to better monitor energy consumption and serves GHG emission reduction. Zhehao Ren, Lixian Zhang 0002, Bin Chen 0005, Haohuan Fu, Bing Xu 0001 |
IGARSS | 4 |
| 2021 | Monitoring Daily Nighttime Light Based on Modis and Deep Learning: A Belgium Case StudyabstractSatellite-observed night-time light has been utilized as a significant indicator for human activities and its impact on environment. Up to present, existing nighttime light (NTL) data still faces challenges in detecting the short-term human-related events due to limitation of the satellite revisit time and data quality. In this paper, we propose a promising approach for monitoring daily light during night based on deep learning and Moderate Resolution Imaging Spectroradiometer (MODIS). By modelling the relationship between MODIS and observed nighttime light, our proposed approach achieves the capability for conversion from MODIS image to Luojia-I-like daily NTL images. The quantitatively assessment of our generated NTL images demonstrates its great performance in terms of both similarity and NTL pattern. Lixian Zhang 0002, Zhehao Ren, Runmin Dong, Bing Xu 0001, Haohuan Fu |
IGARSS | 5 |
| 2021 | Coconut Trees Detection on the Tenarunga Using High-Resolution Satellite Images and Deep LearningabstractThe Coconut tree is of great importance in economic values and ecological impacts for many tropical developing countries and lots of islands in the Pacific Ocean. Detecting and counting coconut is a meaningful and valuable research. In this paper, we present a coconut tree crown detection method to detect and count the coconut trees in the Tenarunga from high-resolution satellite images acquired by Google Earth. Our coconut tree detection method contains three major procedures: feature extraction, a multi-level Region Proposal Network (RPN) and a large-scale coconut tree detection workflow. We manually annotate all coconut trees for our study regions in the Tenarunga. Eventually, we achieve a higher average F1-score of 77.14% in our four test regions than pure Faster R-CNN. Experiment results demonstrate the potential for large-scale individual coconut tree detection and counting from high-resolution satellite images using deep learning. Juepeng Zheng, Wenzhao Wu, Le Yu 0001, Haohuan Fu |
IGARSS | 4 |
| 2021 | LMFF: efficient and scalable layered materials force field on heterogeneous many-core processorsabstractLAMMPS is one of the most popular Molecular Dynamic (MD) packages and is widely used in the field of physics, chemistry and materials simulation. Layered Materials Force Field (LMFF) is our expansion of the LAMMPS potential function based on the Tersoff potential and inter-layer potential (ILP) in LAMMPS. LMFF is designed to study layered materials such as graphene and boron hexanitride. It is universal and does not depend on any platform. We have also carried out a series of optimizations on LMFF and the optimization work is carried out on the new generation of Sunway supercomputer, called SWLMFF. Experiments show that our implementation is efficient, scalable and portable. When generic LMFF is ported to Intel Xeon Gold 6278C, 2X performance improvement is achieved. For the optimized SWLMFF, the overall performance improvement is nearly 200--330X compared to the original ILP and Tersoff potentials. And SWLMFF has good parallel efficiency of 95%-100% under weak scaling with 2.7 million atoms on a single process. The maximum atomic system simulated by SWLMFF is close to 231 atoms. And nanosecond simulations in one day can be realized. Ping Gao 0005, Xiaohui Duan, Jiaxu Guo, Zhenya Song, Li-Zhen Cui 0001, Xiangxu Meng, Xin Liu 0081, Wusheng Zhang, Ming Ma 0012, Dexun Chen, Haohuan Fu, Wei Xue 0003, Guangwen Yang 0002 |
SC | 13 |
| 2021 | Closing the "quantum supremacy" gap: achieving real-time simulation of a random quantum circuit using a new Sunway supercomputerabstractWe develop a high-performance tensor-based simulator for random quantum circuits(RQCs) on the new Sunway supercomputer. Our major innovations include: (1) a near-optimal slicing scheme, and a path-optimization strategy that considers both complexity and compute density; (2) a three-level parallelization scheme that scales to about 42 million cores; (3) a fused permutation and multiplication design that improves the compute efficiency for a wide range of tensor contraction scenarios; and (4) a mixed-precision scheme to further improve the performance. Our simulator effectively expands the scope of simulatable RQCs to include the 10X10(qubits)X(1+40+1)(depth) circuit, with a sustained performance of 1.2 Eflops (single-precision), or 4.4 Eflops (mixed-precision)as a new milestone for classical simulation of quantum circuits; and reduces the simulation sampling time of Google Sycamore to 304 seconds, from the previously claimed 10,000 years. Yong (Alexander) Liu, Xin (Lucy) Liu, Fang (Nancy) Li, Haohuan Fu, Yuling Yang, Jiawei Song, Pengpeng Zhao 0006, Dajia Peng, Huarong Chen, Chu Guo, Heliang Huang, Wenzhao Wu, Dexun Chen |
SC | 4 |
| 2021 | Editorial for the special issue on large-scale AI in classical HPC environment and AI for science
Wei Xue 0003, Haohuan Fu, Weile Jia, Guangming Tan |
CCF Trans. High Perform. Comput. | 2 |
| 2021 | Training a Seismogram Discriminator Based on ResNetabstractSelecting appropriate data of good quality is the primary precondition for conducting meaningful analysis in seismological research. The data set used in a study is usually selected by setting a threshold for searching the earthquake catalog to ensure data quality, such as parameters of magnitude or hypocentral distance. While the threshold approach is useful, it cannot guarantee the consistently good quality of the selected seismograms. For that reason, a manual checking process is often required for quality control, which is inefficient and fallible for large data sets. In this study, we develop an automated seismogram discriminator that is capable of selecting quality seismograms from massive events. The discriminator is created based on a machine learning technique using a residual neural network (ResNet), a well-designed architecture in computer image recognition. Three-component seismic records of an earthquake from the catalog are considered as images by the ResNet. Using the images of seismic records, the ResNet can be trained to distinguish between good and poor seismic records. This discriminatory ability is evaluated using a blind testing data set of approximately 20 000 three-component seismic records related to the events that occurred in Sichuan between January 2014 and May 2018. The results show that our seismogram discriminator achieves an accuracy of greater than 95%. Huiyu Zhu, Mengyao Sun 0002, Haohuan Fu, Nianmao Du, Jie Zhang 0064 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2020 | Unsupervised Mixed Multi-Target Domain Adaptation for Remote Sensing Images ClassificationabstractAlthough deep learning has been successfully applied in the field of remote sensing image classification, it still requires time-consuming and costly annotations. In recent years, domain adaptation has been witnessed to address this problem as they do not need any human interpreted in the target domain dataset. However, most of the existing works dedicate effort on the circumstance where there is only one source domain and only one target domain. In this paper, we firstly explore one source and multiple target domains issue for remote sensing application and build a challenging mixed multi-target dataset to contribute to the community. Our method constitutes three parts. Firstly, as we are blind for the multitarget domain, we adopt meta learning to divide the mixed multi-target dataset and insert sub-target domain loss as the part of the loss function. Secondly, we apply the adversarial learning to confuse the classifier to discriminate between the source domain images and the whole mixed multi-target domain images. Finally, the meta learning and the adversarial learning are dynamically iterative procedures and the labels for domain classification in mixed multi-target dataset will be updated for a particular iteration. Our method is well-performed in the four common remote sensing dataset (AID, NWPU-RESISC45, UC Merced and WHU-RS19), including five classes (agriculture, forest, river, residential and parking). Our method achieved an average accuracy of 81.59% and outperformed other domain adaptation method. The experiment results indicate our method is promising for large-scale, multi-regional and multi-temporal remote sensing applications. Juepeng Zheng, Wenzhao Wu, Haohuan Fu, Runmin Dong, Lixian Zhang 0002, Shuai Yuan 0005 |
IGARSS | 3 |
| 2020 | Neighbor-list-free molecular dynamics on sunway TaihuLight supercomputerabstractMolecular dynamics (MD) simulations are playing an increasingly important role in many research areas. Pair-wise potentials are widely used in MD simulations of bio-molecules, polymers, and nano-scale materials. Due to a low compute-to-memory-access ratio, their calculation is often bounded by memory transfer speeds. Sunway TaihuLight is one of the fastest supercomputers featuring a custom SW26010 many-core processor. Since the SW26010 has some critical limitations regarding main memory bandwidth and scratchpad memory size, it is considered as a good platform to investigate the optimization of pair-wise potentials especially in terms of data reusage. MD algorithms often use a neighbor-list data structure to reduce the computational workload. In this paper, we show that a cell-list-based approach is more suitable for the SW26010 processor. We apply a number of novel optimization methods including self-adaptable replica-summation for conflict-free parallelization, parameter profiles for flexible vectorization, and particle-cell cutoff checking filters for reducing the computational workload. We also established an open source standalone framework featuring the techniques above, ESMD1, which is at least 50% faster than the latest existing LAMMPS port on a single TaihuLight node. Furthermore, EMSD achieves a weak scaling efficiency of 88% on 4,096 nodes. Xiaohui Duan, Ping Gao 0005, Tingjian Zhang, Hongsong Meng, Bertil Schmidt, Haohuan Fu, Lin Gan 0001, Wei Xue 0003, Guangwen Yang 0002 |
PPoPP | 8 |
| 2020 | Cell-list based molecular dynamics on many-core processors: a case study on sunway TaihuLight supercomputerabstractMolecular dynamics (MD) simulations are playing an increasingly important role in several research areas. The most frequently used potentials in MD simulations are pair-wise potentials. Due to the memory wall, computing pair-wise potentials on many-core processors are usually memory bounded. In this paper, we take the SW26010 processor as an exemplary platform to explore the possibility to break the memory bottleneck by improving data reusage via cell-list-based methods. We use cell-lists instead of neighbor-lists in the potential computation, and apply a number of novel optimization methods. Theses methods include: an adaptive replica arrangement strategy, a parameter profile data structure, and a particle-cell cutoff checking filter. An incremental cell-list building method is also realized to accelerate the construction of cell-lists. Furthermore, we have established an open source standalone framework, ESMD, featuring the techniques above. Experiments show that ESMD is 50~170% faster than previous ports on a single node, and can scale to 1,024 nodes with a weak scalibility of 95%. Xiaohui Duan, Ping Gao 0005, Tingjian Zhang, Hongsong Meng, Bertil Schmidt, Haohuan Fu, Lin Gan 0001, Wei Xue 0003, Guangwen Yang 0002 |
SC | 8 |
| 2020 | Tuning a general purpose software cache library for TaihuLight's SW26010 processor
Xiaohui Duan, Haohuan Fu, Lin Gan 0001, Wei Xue 0003, Guangwen Yang 0002 |
CCF Trans. High Perform. Comput. | 4 |
| 2020 | Editorial for the special issue on HPC algorithms and applications
Haohuan Fu, Wei Xue 0003, Guangming Tan |
CCF Trans. High Perform. Comput. | 1 |
| 2020 | Efficient AES implementation on Sunway TaihuLight supercomputer: A systematic approach
Liandeng Li, Jiarui Fang, Jinlei Jiang, Lin Gan 0001, Weijie Zheng 0001, Haohuan Fu, Guangwen Yang 0002 |
J. Parallel Distributed Comput. | 6 |
| 2020 | Millimeter-Scale and Billion-Atom Reactive Force Field Simulation on Sunway TaihulightabstractLarge-scale molecular dynamics (MD) simulations on supercomputers play an increasingly important role in many research areas. With the capability of simulating charge equilibration (QEq), bonds and so on, Reactive force field (ReaxFF) enables the precise simulation of chemical reactions. Compared to the first principle molecular dynamics (FPMD), ReaxFF has far lower requirements on computational resources so that it can achieve higher efficiencies for large-scale simulations. In this article, we present our efforts on scaling ReaxFF on the Sunway TaihuLight Supercomputer (TaihuLight). We have carefully redesigned the force analysis and neighbor list building steps. By applying fine-grained optimizations we gain better single process performance. For the many-body interactions, we propose an isolated computation and update strategy and implement inverse trigonometric functions. For QEq, we implement a pipelined conjugate gradient (CG) approach to achieving better scalability. Furthermore, we reorganize the data layout and implement the update operation based on data locality in ReaxFF. Our experiments show that this approach can simulate chemical reactions with 1,358,954,496 atoms using 4,259,840 cores with a performance of 0.015 ns/day. To our best knowledge, this is the first realization of chemical reaction simulation with a millimeter-scale force field. Ping Gao 0005, Xiaohui Duan, Tingjian Zhang, Bertil Schmidt, Wusheng Zhang, Lin Gan 0001, Wei Xue 0003, Haohuan Fu, Guangwen Yang 0002 |
IEEE Trans. Parallel Distributed Syst. | 11 |
| 2020 | Large-Scale Automatic K-Means Clustering for Heterogeneous Many-Core SupercomputerabstractThis article presents an automatic k-means clustering solution targeting the Sunway TaihuLight supercomputer. We first introduce a multilevel parallel partition approach that not only partitions by dataflow and centroid, but also by dimension, which unlocks the potential of the hierarchical parallelism in the heterogeneous many-core processor and the system architecture of the supercomputer. The parallel design is able to process large-scale clustering problems with up to 196,608 dimensions and over 160,000 targeting centroids, while maintaining high performance and high scalability. Furthermore, we propose an automatic hyper-parameter determination process for k-means clustering, by automatically generating and executing the clustering tasks with a set of candidate hyper-parameter, and then determining the optimal hyper-parameter using a proposed evaluation method. The proposed autoclustering solution can not only achieve high performance and scalability for problems with massive high-dimensional data, but also support clustering without sufficient prior knowledge for the number of targeted clusters, which can potentially increase the scope of k-means algorithm to new application areas. Wenlai Zhao, Pan Liu 0002, Vladimir Janjic, Xiaohan Yan, Shicai Wang, Haohuan Fu, Guangwen Yang 0002, John Thomson |
IEEE Trans. Parallel Distributed Syst. | 7 |
| 2019 | Large-scale Parallel Design for Cryo-EM Structure Determination on Heterogeneous Many-core ArchitecturesabstractCryo-EM structure determination is the most important research area in structural biology. With the development of cryo-electron microscopy, the resolution has been enhanced significantly, which leads to the huge computation to reconstruct the biomolecule in recent years. In this paper, we present a large-scale parallel design for Cryo-EM structure determination on heterogeneous many-core architectures. A novel task parallel strategy is proposed to reduce the redundant computation and improve the scalability on large-scale systems. Further, We distribute the reconstruction model to each node and rearrange the data layout to achieve high parallel efficiency and reduce the memory footprint. The proposed comprehensive parallel design shows highly parallel efficiency and scalability on large-scale heterogeneous architectures, which could significantly accelerate the whole period of Cryo-EM structure determination process. Hongkun Yu 0002, Ruixin Sun, Wenlai Zhao, Haohuan Fu, Guangwen Yang 0002 |
BIBM | 6 |
| 2019 | Million-Core-Scalable Simulation of the Elastic Migration Algorithm on Sunway TaihuLight SupercomputerabstractMigration algorithm is one of the most essential methods in seismic application to image the underground geology, and to help scientists and researchers in geophysics exploration better understand the earth system. However, due to the desire in migration algorithm for covering lager region and acquiring better resolution, many tough challenges have to be tackled for current state-of-the-art computing systems. This work optimized and scaled the elastic migration algorithm onto the Sunway TaihuLight supercomputer, one of the most powerful systems of the world. Targeting at the major process, the reverse time migration (RTM) algorithm, a set of algorithmic, process-level, and thread-level optimizations is proposed, to significantly improve the performance (up to 163× speedup in time-to-solution) on Sunway CPU. Our design is successfully scaled to over two million cores (2,662,400 cores in total) on the Sunway TaihuLight supercomputer, with nearly ideal weak-scaling efficiency. The largest run is able to achieve a sustainable performance of processing over 859 billion cells per second. Lin Gan 0001, Jingheng Xu, Xin Wang 0233, Sihai Wu, Xiaohui Duan, Haohuan Fu, Guangwen Yang 0002 |
CCGRID | 7 |
| 2019 | swATOP: Automatically Optimizing Deep Learning Operators on SW26010 Many-Core ProcessorabstractAchieving an optimized mapping of Deep Learning (DL) operators to new hardware architectures is the key to building a scalable DL system. However, handcrafted optimization involves huge engineering efforts, due to the variety of DL operator implementations and complex programming skills. Targeting the innovative many-core processor SW26010 adopted by the 3rd fastest supercomputer Sunway TaihuLight, an end-to-end automated framework called swATOP is presented as a more practical solution for DL operator optimization. Arithmetic intensive DL operators are expressed into an auto-tuning-friendly form, which is based on tensorized primitives. By describing the algorithm of a DL operator using our domain specific language (DSL), swATOP is able to derive and produce an optimal implementation by separating hardware-dependent optimization and hardware-agnostic optimization. Hardware-dependent optimization is encapsulated in a set of tensorized primitives with sufficient utilization of the underlying hardware features. The hardware-agnostic optimization contains a scheduler, an intermediate representation (IR) optimizer, an auto-tuner, and a code generator. These modules cooperate to perform an automatic design space exploration, to apply a set of programming techniques, to discover a near-optimal solution, and to generate the executable code. Our experiments show that swATOP is able to bring significant performance improvement on DL operators in over 88% of cases, compared with the best-handcrafted optimization. Compared to a black-box autotuner, the tuning and code generation time can be reduced to minutes from days using swATOP. Jiarui Fang, Wenlai Zhao, Jinzhe Yang, Long Wang 0014, Lin Gan 0001, Haohuan Fu, Guangwen Yang 0002 |
ICPP | 7 |
| 2019 | Parallelizing cryo-EM 3D reconstruction on GPU cluster with a partitioned and streamed modelabstractAs a vital approach to determine the structure of biomacromolecules, high-resolution cryo-electron microscopy (cryo-EM) 3D reconstruction is extremely compute-intensive, and has gradually migrated to GPU accelerators in recent years. With certain kernels already achieving high speedup and efficiency on GPUs, the reconstruction part, which inherently requires accesses of a large 3D model in different orientations, brings tough challenges to GPU architectures and has no effective GPU-based options. To fill the above gap, in this paper, we propose Stream3D, a novel GPU-based parallel design for cryo-EM 3D reconstruction. Our major idea is to reorganize the related problem space as streams of key-value pairs, so that we can achieve both the flexibility and efficiency to compute and accumulate the contribution to the final 3D model from all different 2D image inputs. In addition, we design a hybrid communication mechanism to reduce intra-node communications and enable the solving process on a larger scale. With the addition of our GPU-based reconstruction design, we are able to improve the performance of the reconstruction part itself by 9.50 times, and the performance of the entire processing part (the reconstruction part and the other parts with mature GPU options) by 2.83 times. Moreover, Stream3D enables using the approach at a large scale, with 65.32-fold speedup when using up to 80 GPUs. Shizhen Xu, Haohuan Fu, Hongkun Yu 0002, Wenlai Zhao, Guangwen Yang 0002 |
ICS | 3 |
| 2019 | Large-Scale Oil Palm Tree Detection from High-Resolution Remote Sensing Images Using Faster-RCNNabstractOil palm is of great importance in agricultural productivity for many tropic developing countries and accordingly investigating as well as counting oil palms is a meaningful and valuable research. In this paper, we firstly apply Faster-RCNN, one of the most popular object detection algorithms, to detect tree crowns from satellite images. Although Faster-RCNN has an excellent performance in well-known datasets of general object detection, it does not have obvious advantages in oil palm tree detection in this study compared with other classical machine learning based methods. We argue two reasons accounting for the drawbacks of Faster-RCNN: (1) the size of each oil palm tree is too small (only 17 × 17 pixels on average) in 0.6m-resolution QuickBird satellite images; (2) there are lots of other similar trees around the oil palm trees that make it difficult to detect them correctly. In order to reach a satisfying accuracy, we tailored the Region Proposal Network (RPN) and proposed a simple but practical post-processing strategy based on empirical planting rules, filtering out the wrongly detected trees (False Positives) effectively. Eventually we achieved a higher average F1-score of 94.99% (using IOU based evaluation matrices) in our six study regions compared wtih state-of-the-art oil palm detection methods. In addition, we proposed a workflow of large-scale oil palm tree detection using high-resolution remotely sensed images based deep learning methods. Juepeng Zheng, Maocai Xia, Runmin Dong, Haohuan Fu, Shuai Yuan 0005 |
IGARSS | 5 |
| 2019 | SunwayLB: Enabling Extreme-Scale Lattice Boltzmann Method Based Computing Fluid Dynamics Simulations on Sunway TaihuLightabstractThe Lattice Boltzmann Method (LBM) is a relatively new class of Computational Fluid Dynamics methods. In this paper, we report our work on SunwayLB, which enables LBM based solutions aiming for industrial applications. We propose several techniques to boost the simulation speed and improve the scalability of SunwayLB, including a customized multi-level domain decomposition and data sharing scheme, a carefully orchestrated strategy to fuse kernels with different performance constraints for a more balanced workload, and optimization strategies for assembly code, which bring up to 137x speedup. Based on these optimization schemes, we manage to perform the largest direct numerical simulation which involves up to 5.6 trillion lattice cells, achieving 11,245 billion cell updates per second (GLUPS), 77% memory bandwidth utilization and a sustained performance of 4.7 PFlops. We also demonstrate a series of computational experiments for extreme-large scale fluid flow, as examples of real-world applications, to check the validity and performance of our work. The results show that SunwayLB is competent for a practical solution for industrial applications. Xuesen Chu, Xiaojing Lv, Hongsong Meng, Shupeng Shi, Wenji Han, Jingheng Xu, Haohuan Fu, Guangwen Yang 0002 |
IPDPS | 8 |
| 2019 | GPU-based 3D cryo-EM reconstruction with key-value streams: posterabstractThe 3D reconstruction of cryo-electron microscopy (cryo-EM) structural determination process is highly compute-intensive. It inherently requires accesses of a large 3D model in different and variable orientations, brings tough challenges to GPU architecture and has no effective solutions currently. To fill this gap, we propose a novel GPU-based parallel design for cryo-EM 3D reconstruction. The major idea is to reorganize the related problem space as streams of key-value pairs, so that we can achieve both the flexibility and efficiency to compute and accumulate the contribution to the final 3D model from all different 2D image inputs. In addition, we design a hybrid communication mechanism to reduce intra-node communications and enable the solving process on a larger scale. Shizhen Xu, Hongkun Yu 0002, Haohuan Fu, Guangwen Yang 0002 |
PPoPP | 4 |
| 2019 | SW_GROMACS: accelerate GROMACS on Sunway TaihuLightabstractGROMACS is one of the most popular Molecular Dynamic (MD) applications and is widely used in the field of chemical and bimolecular system study. Similar to other MD applications, it needs long run-time for large-scale simulations. Therefore, many high performance platforms have been employed to accelerate it, such as Knights Landing (KNL), Cell Processor, Graphics Processing Unit (GPU) and so on. As the third fastest supercomputer in the world, Sunway TaihuLight contains 40960 SW26010 processors and SW26010 is a typical many-core processor. To make full use of the superior computation ability of TaihuLight, we port GROMACS to SW26010 with following new strategies: (1) a new deferred update strategy; (2) a new update mark strategy; (3) a full pipeline acceleration. Furthermore, we redesign GROMACS to enable all possible vectorization. Experiments show that our implementation achieves better performance than both Intel KNL and Nvidia P100 GPU when using appropriate number of SW26010 processors for a fair comparison. Tingjian Zhang, Ping Gao 0005, Mingshan Shao, Jinxiao Zhang, Xiaohui Duan, Lin Gan 0001, Haohuan Fu, Wei Xue 0003, Guangwen Yang 0002 |
SC | 11 |
| 2019 | Extreme-scale earthquake simulations on Sunway TaihuLight
Haohuan Fu, Bingwei Chen, Wei Zhang 0321, Guangwen Yang 0002 |
CCF Trans. High Perform. Comput. | 1 |
| 2019 | An automatic performance model-based scheduling tool for coupled climate system models
Nan Ding 0006, Wei Xue 0003, Zhenya Song, Haohuan Fu, Shiming Xu |
J. Parallel Distributed Comput. | 4 |
| 2019 | RedSync: Reducing synchronization bandwidth for distributed deep learning training system
Jiarui Fang, Haohuan Fu, Guangwen Yang 0002, Cho-Jui Hsieh |
J. Parallel Distributed Comput. | 2 |
| 2019 | Performance Tuning and Analysis for Stencil-Based Applications on POWER8 ProcessorabstractThis article demonstrates an approach for combining general tuning techniques with the POWER8 hardware architecture through optimizing three representative stencil benchmarks. Two typical real-world applications, with kernels similar to those of the winning programs of the Gordon Bell Prize 2016 and 2017, are employed to illustrate algorithm modifications and a combination of hardware-oriented tuning strategies with the application algorithms. This work fills the gap between hardware capability and software performance of the POWER8 processor, and provides useful guidance for optimizing stencil-based scientific applications on POWER systems. Jingheng Xu, Haohuan Fu, Lin Gan 0001, Wayne Luk, Guangwen Yang 0002 |
ACM Trans. Archit. Code Optim. | 2 |
| 2019 | Optimizing Finite Volume Method Solvers on Nvidia GPUsabstractAs scientific applications are increasingly ported to GPUs to benefit from both the powerful computing capacity and high throughput, accelerating explicit solvers for GPU-based finite volume methods is gaining more and more attention. In this paper, based on the detailed analysis of the FVM algorithm, we present a set of novel optimization methods, including the explicit data cache mechanism, optimal global memory loading strategy, as well as the inner-thread rescheduling method, which derives a suitable mapping from the solver algorithm to the underlying GPU hardware architecture, so as to remarkably improve the solving performance of structured mesh based FVM. We demonstrate the impact of our tuning techniques on two widely-used atmospheric dynamic kernels (3-D Euler and 2-D SWE) on five kinds of mainstream GPU platforms, and make a detailed analysis of the different tuning methodologies so as to demonstrate how to select the proper tuning strategy to different applications on various GPU platforms. Specifically, 93.9x speedup is achieved for the 3D Euler solver on Nvidia V100 over one 12-core Intel E5-2697 (v2) CPU, which is a 77 percent improvement compared with the original speedup without adopting the tuning techniques presented in this work. Jingheng Xu, Guangwen Yang 0002, Haohuan Fu, Wayne Luk, Lin Gan 0001, Wei Xue 0003, Chao Yang 0002, Yong Jiang 0001, Conghui He |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2018 | swCaffe: A Parallel Framework for Accelerating Deep Learning Applications on Sunway TaihuLightabstractThis paper reports our efforts on swCaffe, a highly efficient parallel framework for accelerating deep neural networks (DNNs) training on Sunway TaihuLight, the current fastest supercomputer in the world that adopts a unique many-core heterogeneous architecture, with 40,960 SW26010 processors connected through a customized communication network.First, we point out some insightful principles to fully exploit the performance of the innovative many-core architecture.Second, we propose a set of optimization strategies for redesigning a variety of neural network layers based on Caffe.Third, we put forward a topology-aware parameter synchronization scheme to scale the synchronous Stochastic Gradient Descent (SGD) method to multiple processors efficiently.We evaluate our framework by training a variety of widely used neural networks with the ImageNet dataset.On a single node, swCaffe can achieve 23%˜119% overall performance compared with Caffe running on K40m GPU.As compared with the Caffe on CPU, swCaffe runs 3.04˜7.84xfaster on all the networks.Finally, we present the scalability of swCaffe for training of ResNet-50 and AlexNet on the scale of 1024 nodes. Liandeng Li, Jiarui Fang, Haohuan Fu, Jinlei Jiang, Wenlai Zhao, Conghui He, Xin You 0001, Guangwen Yang 0002 |
CLUSTER | 3 |
| 2018 | PLZMA: A Parallel Data Compression Method for Cloud Computing
Xin Wang 0233, Lin Gan 0001, Jingheng Xu, Jinzhe Yang, Maocai Xia, Haohuan Fu, Xiaomeng Huang, Guangwen Yang 0002 |
ICA3PP (3) | 6 |
| 2018 | A Fast Sparse Triangular Solver for Structured-grid Problems on Sunway Many-core Processor SW26010abstractThe sparse triangular solver (SpTRSV) is one of the most essential kernels in many scientific and engineering applications. Efficiently parallelizing the SpTRSV on modern many-core architectures is considerably difficult due to inherent dependency of computation and discontinuous memory accesses. Achieving high performance of SpTRSV is even more challenging for SW26010, the new-generation customized heterogeneous many-core processor equipped in the top-rank Sunway TaihuLight supercomputer. Owing to regular sparse pattern, structured-grid triangular problems show much different computing characteristics with general ones as well as new opportunities to algorithm design on many-core architectures, which ever lacks attention. In this work, we focus on how to design and implement fast SpTRSV for structured-grid problems on SW26010. A generalized algorithm framework of parallel SpTRSV is proposed for best utilization of the features and flexibilities of SW26010 many-core architecture according to the fine-grained Producer-Consumer model. Moreover, a novel parallel structured-grid SpTRSV is presented by using direct data transfers across registers of the computing elements of SW26010. Experiments on four typical structured-grid triangular problems with different problem sizes demonstrate that our SpTRSV can achieve an average momory bandwidth utilization of 79.7% according to the stream benchmark, which leads to a speedup of 17.7 over serial version on SW26010. Furthermore, experiments with real world sparse linear problems show that our proposed SpTRSV can achieve superior preconditioning performance over the Intel Xeon E5-2670 v3 CPU and Intel Xeon Phi 7210 KNL over DDR4 memory. Wei Xue 0003, Yulong Ao, Chao Yang 0002, Haohuan Fu, Lin Gan 0001, Guangwen Yang 0002 |
ICPP | 6 |
| 2018 | Simulating the Wenchuan earthquake with accurate surface topography on Sunway TaihuLight
Bingwei Chen, Haohuan Fu, Yanwen Wei, Conghui He, Wubin Wan, Lin Gan 0001, Wei Zhang 0321, Guangwen Yang 0002 |
SC | 2 |
| 2018 | Redesigning LAMMPS for peta-scale and hundred-billion-atom simulation on Sunway TaihuLight
Xiaohui Duan, Ping Gao 0005, Tingjian Zhang, Wusheng Zhang, Wei Xue 0003, Haohuan Fu, Lin Gan 0001, Dexun Chen, Xiangxu Meng, Guangwen Yang 0002 |
SC | 8 |
| 2018 | Large-scale hierarchical k-means for heterogeneous many-core supercomputers
Liandeng Li, Wenlai Zhao, Haohuan Fu, Guangwen Yang 0002, John Thomson |
SC | 4 |
| 2018 | Application software beyond exascale: challenges and possible trendsabstractWith various exascale systems in different countries planned over the next three to five years, developing application software for such unprecedented computing capabilities and parallel scaling becomes a major challenge. In this study, we start our discussion with the current 125-Pflops Sunway TaihuLight system in China and its related application challenges and solutions. Based on our current experience with Sunway TaihuLight, we provide a projection into the next decade and discuss potential challenges and possible trends we would probably observe in future high performance computing software. Guangwen Yang 0002, Haohuan Fu |
Frontiers Inf. Technol. Electron. Eng. | 2 |
| 2018 | Optimizing Convolutional Neural Networks on the Sunway TaihuLight SupercomputerabstractThe Sunway TaihuLight supercomputer is powered by SW26010, a new 260-core processor designed with on-chip fusion of heterogeneous cores. In this article, we present our work on optimizing the training process of convolutional neural networks (CNNs) on the Sunway TaihuLight supercomputer. Specifically, a highly efficient library (swDNN) and a customized Caffe framework (swCaffe) are proposed. Architecture-oriented optimization methods targeting the many-core architecture of SW26010 are introduced and are able to achieve 48× speedup for the convolution routine in swDNN and 4× speedup for the complete training process of the VGG-16 network using swCaffe, compared to the unoptimized algorithm and framework. Compared to the cuDNN library and the Caffe framework based on the NVIDIA K40m GPU, the proposed swDNN library and swCaffe framework on SW26010 have nearly half the performance of K40m in single -precision and have 3.6× and 1.8× speedup over K40m in double precision, respectively. Wenlai Zhao, Haohuan Fu, Jiarui Fang, Weijie Zheng 0001, Lin Gan 0001, Guangwen Yang 0002 |
ACM Trans. Archit. Code Optim. | 2 |
| 2017 | A Nanosecond-Level Hybrid Table Design for Financial Market Data GeneratorsabstractThis paper proposes a hybrid sorted table design for minimizing electronic trading latency, with three main contributions. First, a hierarchical sorted table with two levels, a fast cache table in reconfigurable hardware storing megabytes of data items and a master table in software storing gigabytes of data items. Second, a full set of operations, including insertion, deletion, selection and sorting, for the hybrid table with latency in a few cycles. Third, an on-demand synchronization scheme between the cache table and the master table. An implementation has been developed that targets an FPGA-based network card in the environment of the China Financial Futures Exchange (CFFEX) which sustains 1-10Gb/s bandwidth with latency of 400 to 700 nanoseconds, providing an 80- to 125-fold latency reduction compared to a fully optimized CPU-based solution, and a 2.2-fold reduction over an existing FPGA-based solution. Haohuan Fu, Conghui He, Wayne Luk, Guangwen Yang 0002 |
FCCM | 1 |
| 2017 | Accelerating Financial Market Server through Hybrid List Design (Abstract Only)
Haohuan Fu, Conghui He, Huabin Ruan, Itay Greenspon, Wayne Luk, Yongkang Zheng, Junfeng Liao, Guangwen Yang 0002 |
FPGA | 1 |
| 2017 | Exploring the potential of reconfigurable platforms for order book updateabstractThe order book update (OBU) algorithm is widely used in financial exchanges for rebuilding order books. The number of messages produced has drastically increased over time. The software solutions become more and more difficult to scale with the growing message rate and meet the requirement of low latency. This paper explores the potential of reconfigurable platforms in revolutionizing the order book architecture, and proposes a novel order book update algorithm optimized for maximal throughput and minimal latency. Our approach has three main contributions. First, we derive a fixed tick data structure for the order book that is easier to be mapped to the hardware. Second, we design a customized cache storing the top five levels of the order book to further reduce the latency. Third, we propose a hardware-friendly order book update algorithm based on the data structures we proposed. In the experiment, our FPGA-based solution can process 1.2-1.5 million messages per second with the throughput of 10Gb/s and the latency of 132-288 nanoseconds, which is 90-157 times faster than a CPU-based solution, and 5.2-6.6 times faster than an existing FPGA-based solution. Conghui He, Haohuan Fu, Wayne Luk, Guangen Yang |
FPL | 2 |
| 2017 | An FPGA-based tree crown detection approach for remote sensing imagesabstractThe on-board data processing for remote sensing images is a popular research issue in recent years. The higher data acquisition speed and the limited power budget of satellite increases the challenges of large-scale remote sensing image processing. This paper proposes a high performance tree crown detection approach for large-scale remote sensing images on FPGAs. A pipelined-friendly and resource-economic tree crown detection algorithm (PF-TCD) is designed through reconstructing and modifying the workflow of the original algorithm. Our proposed PF-TCD obtains 18.75 times speedup for a large-scale real-world remote sensing image compared with a fully-optimized software implementation of the original algorithm on an Intel 12- core CPU. Conghui He, Haohuan Fu, Wayne Luk |
FPT | 3 |
| 2017 | Deep convolutional neural network based large-scale oil palm tree detection for high-resolution remote sensing imagesabstractThis paper proposed a deep convolutional neural network (DCNN) based framework for large-scale oil palm tree detection using high-resolution remote sensing images in Malaysia. Different from the previous palm tree or tree crown detection studies, the palm trees in our study area are very crowded and their crowns often overlap. Moreover, there are various land cover types in our study area, e.g. impervious, bare land, and other vegetation, etc. The main steps of our proposed method include large-scale and multi-class sample collection, AlexNet-based DCNN training and optimization, sliding window-based label prediction, and post-processing. Compared with the manually interpreted ground truth, our proposed method achieves detection accuracies of 92%-97% in our study area, which are greatly higher than the accuracies obtained from another two detection methods used in this paper. Haohuan Fu, Le Yu 0001 |
IGARSS | 2 |
| 2017 | 26 PFLOPS Stencil Computations for Atmospheric Modeling on Sunway TaihuLightabstractStencil computation arises from a broad set of scientific and engineering applications and often plays a critical role in the performance of extreme-scale simulations. Due to the memory bound nature, it is a challenging task to opti- mize stencil computation kernels on modern supercomputers with relatively high computing throughput whilst relatively low data-moving capability. This work serves as a demon- stration on the details of the algorithms, implementations and optimizations of a real-world stencil computation in 3D nonhydrostatic atmospheric modeling on the newly announced Sunway TaihuLight supercomputer. At the algorithm level, we present a computation-communication overlapping technique to reduce the inter-process communication overhead, a locality- aware blocking method to fully exploit on-chip parallelism with enhanced data locality, and a collaborative data accessing scheme for sharing data among different threads. In addition, a variety of effective hardware specific implementation and optimization strategies on both the process- and thread-level, from the fine-grained data management to the data layout transformation, are developed to further improve the per- formance. Our experiments demonstrate that a single-process many-core speedup of as high as 170x can be achieved by using the proposed algorithm and optimization strategies. The code scales well to millions of cores in terms of strong scalability. And for the weak-scaling tests, the code can scale in a nearly ideal way to the full system scale of more than 10 million cores, sustaining 25.96 PFLOPS in double precision, which is 20% of the peak performance. Yulong Ao, Chao Yang 0002, Wei Xue 0003, Haohuan Fu, Fangfang Liu 0004, Lin Gan 0001, Wenjing Ma |
IPDPS | 5 |
| 2017 | swDNN: A Library for Accelerating Deep Learning Applications on Sunway TaihuLightabstractTo explore the potential of training complex deep neural networks (DNNs) on other commercial chips rather than GPUs, we report our work on swDNN, which is a highly-efficient library for accelerating deep learning applications on the newly announced world-leading supercomputer, Sunway TaihuLight. Targeting SW26010 processor, we derive a performance model that guides us in the process of identifying the most suitable approach for mapping the convolutional neural networks (CNNs) onto the 260 cores within the chip. By performing a systematic optimization that explores major factors, such as organization of convolution loops, blocking techniques, register data communication schemes, as well as reordering strategies for the two pipelines of instructions, we manage to achieve a double-precision performance over 1.6 Tflops for the convolution kernel, achieving 54% of the theoretical peak. Compared with Tesla K40m with cuDNNv5, swDNN results in 1.91-9.75x performance speedup in an evaluation with over 100 parameter configurations. Jiarui Fang, Haohuan Fu, Wenlai Zhao, Bingwei Chen, Weijie Zheng 0001, Guangwen Yang 0002 |
IPDPS | 2 |
| 2017 | 18.9-Pflops nonlinear earthquake simulation on Sunway TaihuLight: enabling depiction of 18-Hz and 8-meter scenariosabstractThis paper reports our large-scale nonlinear earthquake simulation software on Sunway TaihuLight. Our innovations include: (1) a customized parallelization scheme that employs the 10 million cores efficiently at both the process and the thread levels; (2) an elaborate memory scheme that integrates on-chip halo exchange through register communcation, optimized blocking configuration guided by an analytic model, and coalesced DMA access with array fusion; (3) on-the-fly compression that doubles the maximum problem size and further improves the performance by 24%. With these innovations to remove the memory constraints of Sunway TaihuLight, our software achieves over 15% of the system's peak, better than the 11.8% efficiency achieved by a similar software running on Titan, whose byte to flop ratio is 5 times better than TaihuLight. The extreme cases demonstrate a sustained performance of over 18.9 Pflops, enabling the simulation of Tangshan earthquake as an 18-Hz scenario with an 8-meter resolution. Haohuan Fu, Conghui He, Bingwei Chen, Zekun Yin, Tingjian Zhang, Wei Xue 0003, Wanwang Yin, Guangwen Yang 0002 |
SC | 1 |
| 2017 | Redesigning CAM-SE for peta-scale climate modeling performance and ultra-high resolution on Sunway TaihuLightabstractThe Community Atmosphere Model (CAM) is ported, redesigned, and scaled to the full system of the Sunway TaihuLight, and provides peta-scale climate modeling performance. We refactored and optimized the complete code using OpenACC directives at the first stage. A more aggressive and finer-grained redesign is then applied on the CAM, to achieve finer memory control and usage, more efficient vectorization and compute and communication overlapping. We further improve the CAM performance of a 260-core Sunway processor to the range of 28 to 184 Intel CPU cores, and achieve a sustainable double-precision performance of 3.3 PFlops for a 750 m global simulation when using 10,075,000 cores. CAM on Sunway achieves the simulation speed of 3.4 and 21.5 simulation-year-per-day (SYPD) for global 25-km and 100-km resolution respectively; and enables us to perform, to our knowledge, the first simulation of the complete lifecycle of hurricane Katrina, and achieve close-to-observation simulation results for both track and intensity. Haohuan Fu, Junfeng Liao, Nan Ding 0006, Xiaohui Duan, Lin Gan 0001, Yishuang Liang, Jinzhe Yang, Lanning Wang, Guangwen Yang 0002 |
SC | 1 |
| 2017 | Designing and implementing a heuristic cross-architecture combination for graph traversal
Yang You 0001, Haohuan Fu, David A. Bader, Guangwen Yang 0002 |
J. Parallel Distributed Comput. | 2 |
| 2017 | An EnKF-based scheme to optimize hyper-parameters and features for SVM classifier
Yingsheng Ji, Yushu Chen, Haohuan Fu, Guangwen Yang 0002 |
Pattern Recognit. | 3 |
| 2017 | A Fully-Pipelined Hardware Design for Gaussian Mixture ModelsabstractGaussian Mixture Models (GMMs) are widely used in many applications such as data mining, signal processing and computer vision, for probability density modeling and soft clustering. However, the parameters of a GMM need to be estimated from data by, for example, the Expectation-Maximization algorithm for Gaussian Mixture Models (EM-GMM), which is computationally demanding. This paper presents a novel design for the EM-GMM algorithm targeting reconfigurable platforms, with five main contributions. First, a pipeline-friendly EM-GMM with diagonal covariance matrices that can easily be mapped to hardware architectures. Second, a function evaluation unit for Gaussian probability density based on fixed-point arithmetic. Third, our approach is extended to support a wide range of dimensions or/and components by fitting multiple pieces of smaller dimensions onto an FPGA chip. Fourth, we derive a cost and performance model that estimates logic resources. Fifth, our dataflow design targeting the Maxeler MPCX2000 with a Stratix-5SGSD8 FPGA can run over 200 times faster than a 6-core Xeon E5645 processor, and over 39 times faster than a Pascal TITAN-X GPU. Our design provides a practical solution to applications for training and explores better parameters for GMMs with hundreds of millions of high dimensional input instances, for low-latency and high-performance applications. Conghui He, Haohuan Fu, Ce Guo 0002, Wayne Luk, Guangwen Yang 0002 |
IEEE Trans. Computers | 2 |
| 2016 | Unleashing the performance potential of CPU-GPU platforms for the 3D atmospheric Euler solverabstractAs a traditional application on various supercomputers, atmospheric modeling has long been suffering from the low performance efficiency. In this paper, we pick the 3D Euler equation solver (the most essential dynamic component for a non-hydrostatic atmospheric model) as the target application, and explore the maximum performance efficiency that can be achieved on CPU-GPU hybrid architectures. Besides presenting the suitable hybrid domain decomposition methodology and taking proper usage of tuning techniques for both the CPU and GPU parts, we further propose a novel GPU tuning technique, namely the customizable data caching mechanism with thread warp rescheduling scheme, which is specifically designed for the Euler solver. Combining all the optimizing approaches together, remarkable performance boost has been achieved on mainstream GPU architectures including Tesla Fermi C2050, K20×, K40 and K80. Especially, on the latest Tesla K80, we demonstrate a 31.64× speedup over the performance of 12-core E5-2697 CPU. In addition, based on a hybrid CPU-GPU node with two 12-core E5-2697 CPUs and two Tesla K80 GPUs, a sustained double-precision performance of 1.04 Tflops (16% of the peak) is achieved, which is remarkably higher than the efficiency of similar optimizing tasks based on heterogeneous platforms (strictly less than 10%, as demonstrated in the related work). In addition, a nearly linear weak scaling efficiency is achieved which demonstrate the effectiveness of our domain decomposition method. Haohuan Fu, Jingheng Xu, Lin Gan 0001, Chao Yang 0002, Wei Xue 0003, Wenlai Zhao, Guangwen Yang 0002 |
ASAP | 1 |
| 2016 | Performance optimization of Jacobi stencil algorithms based on POWER8 architectureabstractIn this paper we choose the widely used Jacobi stencil algorithm as our target program to evaluate the effectiveness of tuning techniques based on the latest POWER8 processor, thus to provide optimization guidelines to similar stencil based algorithms. Jingheng Xu, Haohuan Fu, Lin Gan 0001, Hongbo Peng, Guangwen Yang 0002 |
ASAP | 2 |
| 2016 | F-CNN: An FPGA-based framework for training Convolutional Neural NetworksabstractThis paper presents a novel reconfigurable framework for training Convolutional Neural Networks (CNNs). The proposed framework is based on reconfiguring a streaming datapath at runtime to cover the training cycle for the various layers in a CNN. The streaming datapath can support various parameterized modules which can be customized to produce implementations with different trade-offs in performance and resource usage. The modules follow the same input and output data layout, simplifying configuration scheduling. For different layers, instances of the modules contain different computation kernels in parallel, which can be customized with different layer configurations and data precision. The associated models on performance, resource and bandwidth can be used in deriving parameters for the datapath to guide the analysis of design trade-offs to meet application requirements or platform constraints. They enable estimation of the implementation specifications given different layer configurations, to maximize performance under the constraints on bandwidth and hardware resources. Experimental results indicate that the proposed module design targeting Maxeler technology can achieve a performance of 62.06 GFLOPS for 32-bit floating-point arithmetic, outperforming existing accelerators. Further evaluation based on training LeNet-5 shows that the proposed framework achieves about 4 times faster than CPU implementation of Caffe and about 7.5 times more energy efficient than the GPU implementation of Caffe. Wenlai Zhao, Haohuan Fu, Wayne Luk, Yuchun Ma, Guangwen Yang 0002 |
ASAP | 2 |
| 2016 | Graph-Oriented Code Transformation Approach for Register-Limited Stencils on GPUsabstractStencil kernels play an important role in many scientific and engineering disciplines. With the development of numerical algorithms and the increasing requirements of accuracy, register-limited stencils containing massive variables and operations are widely used. However, these register-limited stencils consume vast resources when executing on GPUs. The excessive use of registers reduces the number of active threads dramatically, and consequently leads to a serious performance decline. To improve the performance of these register-limited stencils, we propose a DDG (data-dependency-graph) oriented code transformation approach in this paper. By analyzing, deleting and transforming the original stencil program on GPUs, our graph-oriented code transformation approach explores for the best trade-off between the calculation amount and the parallelism degree, and further achieves better performance. The graph-oriented code transformation approach is evaluated using the Weighted Nearly Analytic Discrete stencil, and the experimental result shows that a speedup of 2.16X can be achieved when compared with the original fairly-optimized implementation. To the best of our knowledge, our study takes the first step towards balancing the calculation amount and parallelism degree of the extremely register-limited stencils on GPUs. Mengyao Jin, Haohuan Fu, Zihong Lv, Guangwen Yang 0002 |
CCGrid | 2 |
| 2016 | Generalized GPU Acceleration for Applications Employing Finite-Volume MethodsabstractScientific HPC applications are increasingly ported to GPUs to benefit from both the high throughput and the powerful computing capacity. Many of these applications, such as atmospheric modeling and hydraulic erosion simulation, are adopting the finite volume method (FVM) as the solver algorithm. However, the communication components inside these applications generally lead to a low flop-to-byte ratio and an inefficient utilization of GPU resources. This paper aims at optimizing FVM solver based on the structured mesh. Besides a high-level overview of the finite-volume method as well as its basic optimizations on modern GPU platforms, we further present two generalized tuning techniques including an explicit cache mechanism as well as an inner-thread rescheduling method that tries to achieve a suitable mapping between the algorithm feature and the platform architecture. To the end, we demonstrate the impact of our generalized optimization methods in two typical atmospheric dynamic kernels (Euler and SWE) based on four mainstream GPU platforms. According to the experimental results of Tesla K80, speedups of 24.4x for SWE and 31.5x for Euler could be achieved over a 12-core Intel E5-2697 CPU, which is a great promotion compared with its original speedup (18x and 15.47x) without applying these two methods. Jingheng Xu, Haohuan Fu, Lin Gan 0001, Chao Yang 0002, Wei Xue 0003, Shizhen Xu, Wenlai Zhao, Bingwei Chen, Guangwen Yang 0002 |
CCGrid | 2 |
| 2016 | Cache-Friendly Design for Complex Spatially-Variable Coefficient Stencils on Many-Core ArchitecturesabstractMany-core architectures, such as the NVIDIA graphics processing unit and Intel Xeon Phi, which are characterized by high computation resources but limited on-chip memory capacity, have been used to significantly accelerate various computationally demanding tasks. Stencil operators are naturally suitable for such architectures because of their parallel calculation patterns. However, only simple stencils with points distributed along the axes and with constant coefficients have been fully investigated. This study first provides insights into optimization strategies for stencils with complex shapes, including off-axial points and spatially variable coefficients. Through our proposed stencil-decomposition schemes, we maintain read-only coefficients in on-chip caches to avoid unvectorized memory access. To alleviate the resulting severe cache-starvation situation, a generalized cache-friendly design for many-core architecture is proposed. It can reduce cache miss times and cache space consumption. The proposed methodology significantly improves the performance of stencil operations in a real seismic imaging application and introduces a new option to write highly efficient memory-bound stencil-like loops. Jiarui Fang, Haohuan Fu, Guangwen Yang 0002 |
HiPC | 2 |
| 2016 | TADE: Tight Adaptive Differential Evolution
Weijie Zheng 0001, Haohuan Fu, Guangwen Yang 0002 |
PPSN | 2 |
| 2016 | Refactoring and optimizing the community atmosphere model (CAM) on the sunway taihulight supercomputerabstractThis paper reports our efforts on refactoring and optimizing the Community Atmosphere Model (CAM) on the Sunway TaihuLight supercomputer, which uses a many-core processor that consists of management processing elements (MPEs) and clusters of computing processing elements (CPEs). To map the large code base of CAM to the millions of cores on the Sunway system, we take OpenACC-based refactoring as the major approach, and apply source-to-source translator tools to exploit the most suitable parallelism for the CPE cluster, and to fit the intermediate variable into the limited on-chip fast buffer. For individual kernels, when comparing the original ported version using only MPEs and the refactored version using both the MPE and CPE clusters, we achieve up to 22× speedup for the compute-intensive kernels. For the 25km resolution CAM global model, we manage to scale to 24,000 MPEs, and 1,536,000 CPEs, and achieve a simulation speed of 2.81 model years per day. Haohuan Fu, Junfeng Liao, Wei Xue 0003, Lanning Wang, Dexun Chen, Long Gu, Jinxiu Xu 0001, Nan Ding 0006, Conghui He, Shizhen Xu, Yishuang Liang, Jiarui Fang, Yuanchao Xu 0001, Weijie Zheng 0001, Jingheng Xu, Zhen Zheng, Wanjing Wei, Bingwei Chen, Xiaomeng Huang, Guangwen Yang 0002 |
SC | 1 |
| 2016 | 10M-core scalable fully-implicit solver for nonhydrostatic atmospheric dynamicsabstractAn ultra-scalable fully-implicit solver is developed for stiff time-dependent problems arising from the hyperbolic conservation laws in nonhydrostatic atmospheric dynamics. In the solver, we propose a highly efficient hybrid domain-decomposed multigrid preconditioner that can greatly accelerate the convergence rate at the extreme scale. For solving the overlapped subdomain problems, a geometry-based pipelined incomplete LU factorization method is designed to further exploit the on-chip fine-grained concurrency. We perform systematic optimizations on different hardware levels to achieve best utilization of the heterogeneous computing units and substantial reduction of data movement cost. The fully-implicit solver successfully scales to the entire system of the Sunway TaihuLight supercomputer with over 10.5M heterogeneous cores, sustaining an aggregate performance of 7.95 PFLOPS in double-precision, and enables fast and accurate atmospheric simulations at the 488-m horizontal resolution (over 770 billion unknowns) with 0.07 simulated-years-per-day. This is, to our knowledge, the largest fully-implicit simulation to date. Chao Yang 0002, Wei Xue 0003, Haohuan Fu, Hongtao You, Yulong Ao, Fangfang Liu 0004, Lin Gan 0001, Lanning Wang, Guangwen Yang 0002 |
SC | 3 |
| 2016 | The Sunway TaihuLight supercomputer: system and applications
Haohuan Fu, Junfeng Liao, Jinzhe Yang, Lanning Wang, Zhenya Song, Xiaomeng Huang, Chao Yang 0002, Wei Xue 0003, Fangfang Liu 0004, Fangli Qiao, Xunqiang Yin, Chaofeng Hou, Jian Zhang 0070, Yangang Wang 0002, Chunbo Zhou, Guangwen Yang 0002 |
Sci. China Inf. Sci. | 1 |
| 2015 | Optimizing Residue Number Reverse Converters through Bitwise Arithmetic on FPGAsabstractAs a promising number representation method to provide inspiring operational performance, the Residue Number System (RNS) has been widely applied in many key applications for data pocessing. However, a highly-efficient and general-purpose reverse converter, which is the key component in an RNS system, is still less to be seen, due to the costly and complex operators that require large amounts of computing resources and a long latency to accomplish. In this paper, we are targeting at reverse converters that are highly efficient and can support general moduli sets. We first propose optimizing methods based on the bit wise arithmetic to improve the performance of general reverse converters such as CRT and New CRT. The methods are capable of replacing expensive operations such as additions and multiplications with bit wise operations. We also optimize the performance of specific reverse converter through condition reduction and pre-calculation methods. Furthermore, we develop a user controlled FPGA design generator that can produce optimized reverse converter designs for a number of different moduli sets. Compared with the existing optimized converter designs, our proposed methods can further reduce the latency and resource consumption by 54.2% to 84.6% and 65% to 88.5% respectively. Bangtian Liu, Haohuan Fu, Lin Gan 0001, Wenlai Zhao, Guangwen Yang 0002 |
FCCM | 2 |
| 2015 | Optimizing Complex Spatially-Variant Coefficient Stencils for Seismic Modeling on GPUabstractThe Explicit Time Evolution (ETE) method is an innovative Finite-Difference (FD) type method to simulate the wave propagation in acoustic media with higher spatial and temporal accuracy. However, different from FD, it is difficult to achieve an efficient GPU design because of the poor memory access patterns caused by the off-axis points and spatially-variant coefficients. In this paper, we present a set of new optimization strategies for ETE stencils according to the memory hierarchy of NVIDIA GPU. To handle the problem caused by the complexity of the stencil shapes, we design a one-to-multi updating scheme for shared memory usage. To alleviate the performance damage resulted from the poor memory access pattern of reading spatially-variant coefficients, we propose a stencil decomposition method to reduce un-coalesced global memory access. Based on the state-of-the-art GPU architecture, combining with existing spatial and temporal stencil blocking schemes, we manage to achieve 9.6x and 9.9x speedups compared with a well-tuned 12-core CPUs version for 37-point and 73-point ETE stencils, respectively. Compared with a well-tuned MIC version, the best speedups for the 2 type stencils are 3.7x and 4.7x. Our designs leads to an ETE method that is 31.2x faster than conventional CPU-FD method and make it a practical seismic imaging technology. Jiarui Fang, Haohuan Fu, Nanxun Dai, Lin Gan 0001, Guangwen Yang 0002 |
ICPADS | 2 |
| 2015 | Targeted Mutation: A Novel Mutation Strategy for Differential EvolutionabstractDifferential Evolution (DE) has been shown as an effective, efficient and robust evolutionary computing algorithm. The main force to generate promising offspring is the mutation operator. Usually, two randomly selected vectors are used to generate the differential vector, which maintains the large diversity of mutant directions and ensures the possibility to find global optima. However, strong randomness also leads to the ineffective searching and slow convergence speed. A proper degree of certainty in differential vector will help the population evolve efficiently. This paper proposes a novel mutation strategy called Targeted Mutation that takes the determined target vector as the starting point of the differential vector and maintains the randomness of the ending point, which makes a better trade-off between the certainty and randomness in the differential vector. Besides, Targeted Mutation adopts the best vector as the base vector. The extensive experiments of comparison with two popular mutation operators on 20 benchmark functions demonstrate the competitive performance of our proposed targeted mutation scheme. Our method achieves better or equivalent performance over 70% of total benchmarks against the other two methods. 17 out of 20 function results can get further improved when roughly tuning parameters on each function, showing the potential ability to get even better results. In addition, an integrated evaluation scoring scheme is designed to provide a more concrete demonstration of the overall performance of different approaches, and our method gains the highest score. Weijie Zheng 0001, Haohuan Fu, Guangwen Yang 0002 |
ICTAI | 2 |
| 2015 | Scaling Support Vector Machines on modern HPC platforms
Yang You 0001, Haohuan Fu, Shuaiwen Song, Amanda Randles, Darren J. Kerbyson, Andrés Márquez 0001, Guangwen Yang 0002, Adolfy Hoisie |
J. Parallel Distributed Comput. | 2 |
| 2015 | Ultra-Scalable CPU-MIC Acceleration of Mesoscale Atmospheric Modeling on Tianhe-2abstractIn this work an ultra-scalable algorithm is designed and optimized to accelerate a 3D compressible Euler atmospheric model on the CPU-MIC hybrid system of Tianhe-2. We first reformulate the mesocale model to avoid long-latency operations, and then employ carefully designed inter-node and intra-node domain decomposition algorithms to achieve balance utilization of different computing units. Proper communication-computation overlap and concurrent data transfer methods are utilized to reduce the cost of data movement at scale. A variety of optimization techniques on both the CPU side and the accelerator side are exploited to enhance the in-socket performance. The proposed hybrid algorithm successfully scales to 6,144 Tianhe-2 nodes with a nearly ideal weak scaling efficiency, and achieve over 8 percent of the peak performance in double precision. This ultra-scalable hybrid algorithm may be of interest to the community to accelerating atmospheric models on increasingly dominated heterogeneous supercomputers. Wei Xue 0003, Chao Yang 0002, Haohuan Fu, Yangtong Xu, Junfeng Liao, Lin Gan 0001, Yutong Lu, Rajiv Ranjan 0001, Lizhe Wang 0001 |
IEEE Trans. Computers | 3 |
| 2015 | Solving the Global Atmospheric Equations through Heterogeneous Reconfigurable PlatformsabstractOne of the most essential and challenging components in climate modeling is the atmospheric model. To solve multiphysical atmospheric equations, developers have to face extremely complex stencil kernels that are costly in terms of both computing and memory resources. This article aims to accelerate the solution of global shallow water equations (SWEs), which is one of the most essential equation sets describing atmospheric dynamics. We first design a hybrid methodology that employs both the host CPU cores and the field-programmable gate array (FPGA) accelerators to work in parallel. Through a careful adjustment of the computational domains, we achieve a balanced resource utilization and a further improvement of the overall performance. By decomposing the resource-demanding SWE kernel, we manage to map the double-precision algorithm into three FPGAs. Moreover, by using fixed-point and reduced-precision floating point arithmetic, we manage to build a fully pipelined mixed-precision design on a single FPGA, which can perform 428 floating-point and 235 fixed-point operations per cycle. The mixed-precision design with four FPGAs running together can achieve a speedup of 20 over a fully optimized design on a CPU rack with two eight-core processorsand is 8 times faster than the fully optimized Kepler GPU design. As for power efficiency, the mixed-precision design with four FPGAs is 10 times more power efficient than a Tianhe-1A supercomputer node. Lin Gan 0001, Haohuan Fu, Wayne Luk, Chao Yang 0002, Wei Xue 0003, Xiaomeng Huang, Youhui Zhang, Guangwen Yang 0002 |
ACM Trans. Reconfigurable Technol. Syst. | 2 |
| 2014 | A Fully-Pipelined FPGA Design for Tree-Reweighted Message Passing Algorithm
Wenlai Zhao, Haohuan Fu, Guangwen Yang 0002 |
FCCM | 2 |
| 2014 | A highly-efficient and green data flow engine for solving euler atmospheric equationsabstractAtmospheric modeling is an essential issue in the study of climate change. However, due to the complicated algorithmic and communication models, scientists and researchers are facing tough challenges in finding efficient solutions to solve the atmospheric equations. In this paper, we accelerate a solver for the three-dimensional Euler atmospheric equations through reconfigurable data flow engines. We first propose a hybrid design that achieves efficient resource allocation and data reuse. Furthermore, through algorithmic offsetting, fast memory table, and customizable-precision arithmetic, we map a complex Euler kernel into a single FPGA chip, which can perform 956 floating point operations per cycle. In a 1U-chassis, our CPU-DFE unit with 8 FPGA chips is 18.5 times faster and 8.3 times more power efficient than a multicore system based on two 12-core Intel E5-2697 (Ivy Bridge) CPUs, and is 6.2 times faster and 5.2 times more power efficient than a hybrid unit equipped with two 12-core Intel E5-2697 (Ivy Bridge) CPUs and three Intel Xeon Phi 5120d (MIC) cards. Lin Gan 0001, Haohuan Fu, Chao Yang 0002, Wayne Luk, Wei Xue 0003, Oskar Mencer, Xiaomeng Huang, Guangwen Yang 0002 |
FPL | 2 |
| 2014 | Patra: Parallel tree-reweighted message passing architectureabstractMaximum a posteriori probability inference algorithms for Markov Random Field are widely used in many applications, such as computer vision and machine learning. Sequential tree-reweighted message passing (TRW-S) is an inference algorithm which shows good quality in finding optimal solutions. However, the performance of TRW-S in software cannot meet the requirements of many real-time applications, due to the sequential scheme and the high memory, bandwidth and computational costs. This paper proposes Patra, a novel parallel tree-reweighted message passing architecture, which involves a fully pipelined design targeting FPGA technology. We build a hybrid CPU/FPGA system to test the performance of Patra for stereo matching. Experimental results show that Patra provides about 100 times faster than a software implementation of TRW-S, and 12 times faster than a GPU-based message passing algorithm. Compared with an existing design in four FPGAs, we can achieve 2 times speedup in a single FPGA. Moreover, Patra can work at video rate in many cases, such as a rate of 167 frame/sec for a standard stereo matching test case, which makes it promising for many real-time applications. Wenlai Zhao, Haohuan Fu, Guangwen Yang 0002, Wayne Luk |
FPL | 2 |
| 2014 | Porting the Princeton Ocean Model to GPUs
Shizhen Xu, Xiaomeng Huang, Haohuan Fu, Guangwen Yang 0002 |
ICA3PP (1) | 5 |
| 2014 | Scaling and analyzing the stencil performance on multi-core and many-core architecturesabstractStencils are among the most important and time-consuming kernels in many applications. While stencil optimization has been a well-studied topic on CPU platforms, achieving higher performance and efficiency for the evolving numerical stencils on the more recent multi-core and many-core architectures is still an important issue. In this paper, we explore a number of different stencils, ranging from a basic 7-point Jacobi stencil to more complex high-order stencils used in finer numerical simulations. By optimizing and analyzing those stencils on the latest multi-core and many-core architectures (the Intel Sandy Bridge processor, the Intel Xeon Phi coprocessor, and the NVIDIA Fermi C2070 and Kepler K20x GPUs), we investigate the algorithmic and architectural factors that determine the performance and efficiency of the resulting designs. While multi-threading, vectorization, and optimization on cache and other fast buffers are still the most important techniques that provide performance, we observe that the different memory hierarchy and the different mechanism for issuing and executing parallel instructions lead to the different performance behaviors on CPU, MIC and GPU. With vector-like processing units becoming the major provider of computing power on almost all architectures, the compiler's inability to align all the computing and memory operations would become the major bottleneck from getting a high efficiency on current and future platforms. Our specific optimization of the complex WNAD stencil on GPU provides a good example of what the compiler could do to help. Lin Gan 0001, Haohuan Fu, Wei Xue 0003, Yangtong Xu, Chao Yang 0002, Zihong Lv, Yang You 0001, Guangwen Yang 0002, Kaijian Ou |
ICPADS | 2 |
| 2014 | Enabling and Scaling a Global Shallow-Water Atmospheric Model on Tianhe-2abstractThis paper presents a hybrid algorithm for the petascale global simulation of atmospheric dynamics on Tianhe-2, the world's current top-ranked supercomputer developed by China's National University of Defense Technology (NUDT). Tianhe-2 is equipped with both Intel Xeon CPUs and Intel Xeon Phi accelerators. A key idea of the hybrid algorithm is to enable flexible domain partition between an arbitrary number of processors and accelerators, so as to achieve a balanced and efficient utilization of the entire system. We also present an asynchronous and concurrent data transfer scheme to reduce the communication overhead between CPU and accelerators. The acceleration of our global atmospheric model is conducted to improve the use of the Intel MIC architecture. For the single-node test on Tianhe-2 against two Intel Ivy Bridge CPUs (24 cores), we can achieve 2.07×, 3.18×, and 4.35× speedups when using one, two, and three Intel Xeon Phi accelerators respectively. The average performance gain from SIMD vectorization on the Intel Xeon Phi processors is around 5× (out of the 8× theoretical case). Based on successful computation-communication overlapping, large-scale tests indicate that a nearly ideal weak-scaling efficiency of 93.5% is obtained when we gradually increase the number of nodes from 6 to 8,664 (nearly 1.7 million cores). In the strong-scaling test, the parallel efficiency is about 77% when the number of nodes increases from 1,536 to 8,664 for a fixed 65,664 × 5,664 × 6 mesh with 77.6 billion unknowns. Wei Xue 0003, Chao Yang 0002, Haohuan Fu, Yangtong Xu, Lin Gan 0001, Yutong Lu, Xiaoqian Zhu |
IPDPS | 3 |
| 2014 | MIC-SVM: Designing a Highly Efficient Support Vector Machine for Advanced Modern Multi-core and Many-Core ArchitecturesabstractSupport Vector Machine (SVM) has been widely used in data-mining and Big Data applications as modern commercial databases start to attach an increasing importance to the analytic capabilities. In recent years, SVM was adapted to the field of High Performance Computing for power/performance prediction, auto-tuning, and runtime scheduling. However, even at the risk of losing prediction accuracy due to insufficient runtime information, researchers can only afford to apply offline model training to avoid significant runtime training overhead. Advanced multi- and many-core architectures offer massive parallelism with complex memory hierarchies which can make runtime training possible, but form a barrier to efficient parallel SVM design. To address the challenges above, we designed and implemented MIC-SVM, a highly efficient parallel SVM for x86 based multi-core and many-core architectures, such as the Intel Ivy Bridge CPUs and Intel Xeon Phi co-processor (MIC). We propose various novel analysis methods and optimization techniques to fully utilize the multilevel parallelism provided by these architectures and serve as general optimization methods for other machine learning tools. MIC-SVM achieves 4.4-84x and 18-47x speedups against the popular LIBSVM, on MIC and Ivy Bridge CPUs respectively, for several real-world data-mining datasets. Even compared with GPUSVM, run on a top of the line NVIDIA k20x GPU, the performance of our MIC-SVM is competitive. We also conduct a cross-platform performance comparison analysis, focusing on Ivy Bridge CPUs, MIC and GPUs, and provide insights on how to select the most suitable advanced architectures for specific algorithms and input data patterns. Yang You 0001, Shuaiwen Song, Haohuan Fu, Andrés Márquez 0001, Maryam Mehri Dehnavi, Kevin J. Barker, Kirk W. Cameron, Amanda Randles, Guangwen Yang 0002 |
IPDPS | 3 |
| 2014 | A High Performance Compression Method for Climate DataabstractClimate modeling data are usually multidimensional arrays of floating-point numbers. These arrays typically have two or three spatial dimensions and one temporal dimension, describing the evolvement of climate variables in a time span. With the advances of high performance computing, the volume of climate data is expanding exponentially, bringing tough challenges for climate data archiving and sharing. In this paper, we propose a lossless compression algorithm for the time-spatial climate floating-point arrays. Our compression algorithm can eliminate more data redundancy efficiently through adaptive prediction, XOR-differencing, and multi-way compression. In addition, static regions, which are very common in climate data, can be identified and compressed more efficiently. Moreover, to utilize the multi-cores on modern computers, we proposed a method to parallelize our compression algorithm. Evaluations demonstrate that single thread version of our compression method can achieve the best balance in compression ratios, deflating throughputs and inflating throughputs. And the parallel version can achieve 800 MB/s deflating throughputs and over 2600 MB/s inflating throughputs on a 16-core server. Songbin Liu, Xiaomeng Huang, Yufang Ni, Haohuan Fu, Guangwen Yang 0002 |
ISPA | 4 |
| 2014 | CFIO2: Overlapping Communications and I/O with Computations Using RDMA Technology
Xiaomeng Huang, Shizhen Xu, Haohuan Fu, Guangwen Yang 0002 |
NPC | 5 |
| 2013 | Understanding Data Characteristics and Access Patterns in a Cloud Storage SystemabstractUnderstanding the inherent system characteristics is crucial to the design and optimization of cloud storage system, and few studies have systematically investigated its data characteristics and access patterns. This paper presents an analysis of file system snapshot and five-month access trace of a campus cloud storage system that has been deployed on Tsinghua campus for three years. The system provides online storage and data sharing services for more than 19,000 students and 500 student groups. We report several data characteristics including file size and file type, as well as some access patterns, including read/write ratio, read-write dependency and daily traffic. We find that there are many differences between cloud storage system and traditional file systems: our cloud storage system has larger file sizes, lower read/write ratio, and smaller set of active files than those of a typical traditional file system. With a trace-driven simulation, we find that the cache efficiency can be improved by 5 times using the guidance from our observations. Songbin Liu, Xiaomeng Huang, Haohuan Fu, Guangwen Yang 0002 |
CCGRID | 3 |
| 2013 | A Scalable Barotropic Mode Solver for the Parallel Ocean Program
Xiaomeng Huang, Xiaoge Wang, Haohuan Fu, Shizhen Xu, Huabin Ruan, Wei Xue 0003, Guangwen Yang 0002 |
Euro-Par | 4 |
| 2013 | Global Atmospheric Simulation on a Reconfigurable PlatformabstractSummary form only given. As the only method to study long-term climate trend and to predict potential climate risk, climate modeling is becoming a key research topic among governments and research organizations. One of the most essential and challenging components in climate modeling is the atmospheric model. To cover high resolution in climate simulation scenarios, developers have to face the challenges from billions of mesh points and extremely complex algorithms. Shallow Water Equations (SWEs) are a set of conservation laws that perform most of the essential characteristics of the atmosphere. The study of SWEs can serve as the starting point for understanding the dynamic behavior of the global atmosphere. We choose cubed-sphere mesh as the computational mesh for its better load balance in pole regions over other meshes such as the latitude-longitude mesh. The cubed-sphere mesh is obtained by mapping a cube to the surface of the sphere. The computational domain is then the six patches, each of which is covered with N × N mesh points to be calculated. When written in local coordinates, SWEs have an identical expression on the six patches, that is ∂Q/∂t + 1/Λ ∂(ΛF1)/∂x1+ 1/Λ ∂(ΛF1)/∂z2+ S=0, (1) where (x1, x2) ∈ [-π/4, π/4] are the local coordinates, Q = (h, hu1, hu2)Tis the prognostic variable, Fi= uiQ (i = 1, 2) are the convective fluxes, S is the source term. Spatially discretized with a cell-centered finite volume method and integrated with a second-order accurate TVD Runge-Kutta method, SWE solvers are transferred to the computation of a 13-point upwind stencil that exhibits a diamond shape. To get the prognostic components (h, hu1and hu2) of the central point, its neighboring 12 points need to be accessed. The stencil kernel includes at least 434 ADD/SUB operations, 570 multiplications, 99 divisions. The high arithmetic density of the SWEs algorithm makes it difficult to implement one kernel into the resource-limited FPGA card. In this study, we first proposes a hybrid algorithm that utilizes both CPUs and FPGAs to simulate the global shallow water equations (SWEs). In each of the computational patch, most of the complicated communications happen in the two layers of the outer boundary, whose value need to be exchanged with other patches. Therefore, we decompose each of the six patches into an outer part that includes two layers of the outer boundary meshes, and an inner part that is the remaining part. We assign CPU to handle the communications and the stencil calculation of the outer part, while assign FPGA to process the inner-part stencil. In this way, FPGA and CPU will work simultaneously and the CPU time for stencil and communication can be hidden in the FPGA time for stencil. For the Virtex-6 SX475T that we use in our study, the original program in double-precision will require 299% of the on-board LUTs, 283% of the FFs and 189% of the DSPs, and cannot fit into one FPGA. In order to fit the SWE kernel into one FPGA chip, we apply two algorithmic optimizations to the original design. One is to replace certain computations by lookup tables, so as to reduce the usage of computation resources. The other one is to locate common factors in the algorithm and to remove redundant computations. These two optimizations reduce the resource usage by 20%. To further reduce the resource cost and to fit the extremely complex stencil kernel into one FPGA chip, we perform optimization in the space of customizable representations and precisions. For the variables with a relatively small range, we apply fixed-point number to replace the double-precisions. For the rest parts with a wide dynamic range, we use floating-point numbers with a mixed-precision. Through mixed-precision floating-point and fixed-point arithmetic, we build a complex upwind stencil kernel on a single FPGA. The design includes a highly-efficient pipeline that can perform hundreds of floating-point and fixed-point arithmetic operations concurrently. Compared with our previous work in [1], the solution based on one FPGA acceleration card provides 100 times speedup over a 6-core CPU, and 4 times speedup over a Tianhe-1A supercomputer node that consists of 12 CPU cores and one Fermi GPU. Lin Gan 0001, Haohuan Fu, Wayne Luk, Chao Yang 0002, Wei Xue 0003, Guangwen Yang 0002 |
FCCM | 2 |
| 2013 | An FPGA-Based Data Flow Engine for Gaussian Copula ModelabstractThe Gaussian Copula Model (GCM) plays an important role in the state-of-the-art financial analysis field for modeling the dependence of financial assets. However, the existing implementations of GCM are all computationallydemanding and time-consuming. In this paper, we propose a Dataflow Engine (DFE) design to accelerate the GCM computation. Specifically, a commonly used CPU-friendly GCM algorithm is converted into a fully-pipelined dataflow graph through four steps of optimization: recomposing the algorithm to be pipeline-friendly, removing unnecessary computation, sharing common computing results, and reducing the computing precision while maintaining the same level of accuracy for the computation results. The performance of the proposed DFE design is compared with three CPU-based implementations that are well-optimized. Experimental results show that our DFE solution not only generates fairly accurate result, but also achieves a maximum of 467x speedup over a single-thread CPU-based solution, 120x speedup over a multi-thread CPUbased solution, and 47x speedup over an MPI-based solution. Huabin Ruan, Xiaomeng Huang, Haohuan Fu, Guangwen Yang 0002, Wayne Luk, Sébastien Racanière, Oliver Pell, Wenjing Han |
FCCM | 3 |
| 2013 | Accelerating solvers for global atmospheric equations through mixed-precision data flow engineabstractOne of the most essential and challenging components in a climate system model is the atmospheric model. To solve the multi-physical atmospheric equations, developers have to face extremely complex stencil kernels. In this paper, we propose a hybrid CPU-FPGA algorithm that applies single and multiple FPGAs to compute the upwind stencil for the global shallow water equations. Through mixed-precision arithmetic, we manage to build a fully pipelined upwind stencil design on a single FPGA, which can perform 428 floating-point and 235 fixed-point operations per cycle. The CPU-FPGA algorithm using one Virtex-6 FPGA provides 100 times speedup over a 6-core CPU and 4 times speedup over a hybrid node with 12 CPU cores and a Fermi GPU card. The algorithm using four FPGAs provides 330 times speedup over a 6-core CPU; it is also 14 times faster and 9 times more power efficient than the hybrid CPU-GPU node. Lin Gan 0001, Haohuan Fu, Wayne Luk, Chao Yang 0002, Wei Xue 0003, Xiaomeng Huang, Youhui Zhang, Guangwen Yang 0002 |
FPL | 2 |
| 2013 | Optimize Multidimensional Arrays Queries with Heterogeneous Replica MethodabstractMultidimensional arrays are commonly used in scientific and engineering applications. The disk layout for the multidimensional arrays will obviously affect the performance of data querying. Homogeneous Replica method are widely used to maintain the data reliability in most of the distributed storage systems and used to improve the data locality in some parallel processing systems. In this paper, we propose a novel method, that is heterogeneous replicas, to makes better use of the replica method to optimize the performance of multidimensional arrays querying. The experimental results shows that heterogeneous replicas method can significantly reduce the overhead of disk I/O for most of the queries. With three heterogeneous replicas, the performance of random generated range queries for multidimensional datasets can be improved for 30% on the average. Xiaomeng Huang, Songbin Liu, Haohuan Fu, Qiming Fang, Guangwen Yang 0002 |
NAS | 4 |
| 2013 | A peta-scalable CPU-GPU algorithm for global atmospheric simulationsabstractDeveloping highly scalable algorithms for global atmospheric modeling is becoming increasingly important as scientists inquire to understand behaviors of the global atmosphere at extreme scales. Nowadays, heterogeneous architecture based on both processors and accelerators is becoming an important solution for large-scale computing. However, large-scale simulation of the global atmosphere brings a severe challenge to the development of highly scalable algorithms that fit well into state-of-the-art heterogeneous systems. Although successes have been made on GPU-accelerated computing in some top-level applications, studies on fully exploiting heterogeneous architectures in global atmospheric modeling are still very less to be seen, due in large part to both the computational difficulties of the mathematical models and the requirement of high accuracy for long term simulations. Chao Yang 0002, Wei Xue 0003, Haohuan Fu, Lin Gan 0001, Yangtong Xu, Yutong Lu, Jiachang Sun, Guangwen Yang 0002 |
PPoPP | 3 |
| 2012 | A fully-pipelined expectation-maximization engine for Gaussian Mixture ModelsabstractGaussian Mixture Models (GMMs) are powerful tools for probability density modeling and soft clustering. They are widely used in data mining, signal processing and computer vision. In many applications, we need to estimate the parameters of a GMM from data before working with it. This task can be handled by the Expectation-Maximization algorithm for Gaussian Mixture Models (EM-GMM), which is computationally demanding. In this paper we present our FPGA-based solution for the EM-GMM algorithm. We propose a pipeline-friendly EM-GMM algorithm, a variant of the original EM-GMM algorithm that can be converted to a fully-pipelined hardware architecture. To further improve the performance, we design a Gaussian probability density function evaluation unit that works with fixed-point arithmetic. In the experiments, our FPGA-based solution generates fairly accurate results while achieving a maximum of 517 times speedup over a CPU-based solution, and 28 times speedup over a GPU-based solution. Ce Guo 0002, Haohuan Fu, Wayne Luk |
FPT | 2 |
| 2011 | Eliminating the memory bottleneck: an FPGA-based solution for 3d reverse time migrationabstractMemory-related constraints (memory bandwidth, cache size) are nowadays the performance bottleneck of most computational applications. Especially in the scenario of multiple cores, the performance does not scale with the number of cores in many cases. In our work, we present our FPGA-based solution for the 3D Reverse Time Migration (RTM) algorithm. As the most computationally demanding imaging algorithm in current oil and gas exploration, RTM involves various computational challenges, such as a high demand for storage size and bandwidth, and a poor cache behavior. Combining optimizations from both the algorithmic and architectural perspectives, our FPGA-based solution manages to remove the memory constraints and provide a high performance that can scale well with the amount of computational resources available. Compared with an optimized CPU implementation using two quad-core Intel Nehalem CPUs, our solution achieves 4x speedup on two Virtex-5 FPGAs, and 8x speedup on two Virtex-6 FPGAs. Our projection demonstrates that the performance will continue to scale with the future increase of FPGA capacities. Haohuan Fu, Robert G. Clapp |
FPGA | 1 |
| 2010 | FPGA Designs with Optimized Logarithmic ArithmeticabstractUsing a general polynomial approximation approach, we present an arithmetic library generator for the logarithmic number system (LNS). The generator produces optimized LNS arithmetic libraries that improve significantly over previous LNS designs on area and latency. We also provide area cost estimation and bit-accurate simulation tools that facilitate comparison between LNS and floating-point designs. Haohuan Fu, Oskar Mencer, Wayne Luk |
IEEE Trans. Computers | 1 |
| 2008 | Optimizing residue arithmetic on FPGAsabstractResidue Number System (RNS), which originates from the Chinese Remainder Theorem, is regarded as a promising number representation in the domain of Digital Signal Processing (DSP). This paper describes our work on optimizing residue arithmetic units on the platform of reconfigurable devices, such as FPGAs. First, we provide improved designs for residue arithmetic units. For reverse converters from RNS to binary numbers, we propose a novel design that uses only n-bit additions. Compared to previous work, the design consumes up to 14.3% less area and provides lower latency. Second, we develop a reconfigurable RNS arithmetic library generator for the moduli set {2n−1, 2n, 2n+1}. The generator supports a wide range of RNS numbers, and enables us to perform an extensive comparison between RNS and other number representations at both the arithmetic unit level and the application level. The comparison shows that, for applications involving a large number of multiplications, the RNS designs can reduce up to 1/2 DSP48s for large bit-width settings. Haohuan Fu, Oskar Mencer, Wayne Luk |
FPT | 1 |
| 2007 | Optimizing Logarithmic Arithmetic on FPGAsabstractThis paper proposes optimizations of the methods and parameters used in both mathematical approximation and hardware design for logarithmic number system (LNS) arithmetic. First, we introduce a general polynomial approximation approach with an adaptive divide-in-halves segmentation method for evaluation of LNS arithmetic functions. Second, we develop a library generator that automatically generates optimized LNS arithmetic units with a wide bit-width range from 21 to 64 bits, to support LNS application development and design exploration. The basic arithmetic units are tested on practical FPGA boards as well as software simulation. When compared with existing LNS designs, our generated units provide in most cases 6% to 37% reduction in area and 20% to 50% reduction in latency. The key challenge for LNS remains on the application level. We show the performance of LNS versus floating-point for realistic applications: digital sine/cosine waveform generator, matrix multiplication and radiative Monte Carlo simulation. Our infrastructure for fast prototyping LNS FPGA applications allows us to efficiently study LNS number representation and its tradeoffs in speed and size when compared with floating-point designs. Haohuan Fu, Oskar Mencer, Wayne Luk |
FCCM | 1 |
| 2006 | Comparing floating-point and logarithmic number representations for reconfigurable accelerationabstractThe paper investigates floating-point and logarithmic number representations for computing with FPGAs. The key issue is to select the best number format for an application to improve performance and accuracy. Using A Stream Compiler, ASC as the hardware design and compilation tool, a convenient scheme to compare the designs of both floating-point and logarithmic numbers and select the solution with the best performance and accuracy, was developed. Its contributions are: (1) optimized function evaluations for conversions between logarithmic and floating-point numbers; (2) design and implementation of logarithmic arithmetic, with optimized segmentation and polynomial degree; (3) a practical comparison case study of Monte Carlo radiative heat transfer simulation. Compared to prior work, our design supports two to six times more LNS conversion and LNS arithmetic units on one FPGA. For Monte Carlo simulation, our designs of both number systems produce 39-80% higher throughput with either a smaller area or a higher accuracy Haohuan Fu, Oskar Mencer, Wayne Luk |
FPT | 1 |
| 2005 | An efficient admission control for IEEE 802.11 networks based on throughput analyses of (Un)saturated channelabstractThis paper presents a novel analytical model and an efficient admission control algorithm for IEEE 802.11 DCF access mechanism. In contrast to the previous approaches that only analyzed the saturated status of IEEE 802.11 networks, both saturated and unsaturated states of network are analyzed and the impacts of error-frame rate and retransmission limit are also taken into account based on an improved Markov chain model. Taking the throughput difference between saturated and unsaturated states as the residual bandwidth, an efficient admission control algorithm is designed to utilize the network resources effectively. Extensive simulation data demonstrate that the admission control algorithm is efficient and can make the effective utilization of network resources Lidong Lin, Haohuan Fu, Weijia Jia 0001 |
GLOBECOM | 2 |
| 2005 | Object-Oriented Design and Implementations of 3G-324M Protocol Stack
Weijia Jia 0001, Haohuan Fu |
ICA3PP | 2 |
| 2005 | Efficient wireless link bandwidth detection for IEEE 802.11 networksabstractIn order to provide accurate and real-time bandwidth information and enhance the QoS for bandwidth-sensitive applications in a dynamically changing wireless network, the paper proposes an efficient method for wireless bandwidth detection (WBD) using a packet probing approach on mobile nodes. The novelty and contributions of WBD are three-fold: (1) efficiency - by sending probe packets of various sizes, WBD can determine the wireless link bandwidth with light load and short time duration; (2) accuracy - using different mechanisms to filter out the random time variation, a high detection accuracy can be achieved; (3) stability - our algorithm attains stable results under different cross traffic conditions. Experimental data observed through extensive simulations shows the effectiveness and efficiency of WBD. Haohuan Fu, Lidong Lin, Weijia Jia 0001 |
ICC | 1 |
| 2005 | Next Generation Networks Architecture and Layered End-to-End QoS Control
Weijia Jia 0001, Bo Han 0001, Haohuan Fu |
ISPA | 4 |
| 2005 | Efficient Multiplexing Protocol for Low Bit Rate Multi-point Video Conferencing
Haohuan Fu, Weijia Jia 0001 |
MSN | 1 |
| 2004 | An Integration Approach of Data Mining with Web Cache Pre-fetching
Yingjie Fu, Haohuan Fu, Pui-on Au |
ISPA | 2 |
| 2004 | Efficient construction of connected dominating set in wireless ad hoc networksabstractConnected dominating set based routing is a promising approach for enhancing the routing efficiency in wireless ad hoc networks. However, finding the minimum dominating set in an arbitrary graph is a NP-hard problem. We propose a simple and efficient distributed algorithm for constructing a connected dominating set in wireless ad hoc networks with time complexity O(n) and message complexity O(nlog n). The dominating set generated with our algorithm can be more reliable and load balanced for routing as compared with some well-known algorithms. The simulation results demonstrate that our algorithm outperforms previous work in terms of the size of the resultant connected dominating set. Bo Han 0001, Haohuan Fu, Lidong Lin, Weijia Jia 0001 |
MASS | 2 |
| 2004 | Performance Evaluations of Replacement Algorithms in Hierarchical Web Caching
Haohuan Fu, Pui-on Au, Weijia Jia 0001 |
WAIM | 1 |