Wenzhao Wu

dblp:190/6935 · DBLP profile ↗
← Back
11ranked-venue papers
0as first author
10since 2021 · last 2023
0009-0005-9095-8996ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 7 · 6 since 2021Systems, architecture and hardware · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2023 Automatic Deep Learning Operator Fusion on Sunway SW26010 Many-Core Processor
abstract
Deep learning networks (DNNs) have been growing rapidly in recent years, with increasing demands on computing power. Therefore, accelerating the execution of DNN models has become a research hotspot. Operator fusion is a critical optimization strategy to enhance DNN performance in Deep Learning (DL) frameworks, such as TensorFlow, Pytorch, TVM and Halide. However, these frameworks are designed for general optimization and cannot fully harness the specific features of emerging hardware. Moreover, they primarily implement operator fusion at the operator level, missing out on many fusion opportunities and heavily relying on extensive manual optimizations for fused operators. Targeting the Sunway SW26010 Many-Core processor, the basic building block of Sunway TaihuLight supercomputer, we introduce swAutoFuser, an end-to-end automatic operator fusion and code generation framework. swAutoFuser proposes a set of low-level primitives to leverage hardware features and employs an autofuser to achieve primitive level fusion, which breaks operator boundaries and enables more fusion opportunities. In addition, swAutoFuser can automatically generate high-performance fused operator implementations based on a static cost model, significantly reducing the overhead of manually optimizing fused operators. Our experiments demonstrate that swAutoFuser can improve operator performance by 10% to 56%.
Wenxiang Zhang, Wenzhao Wu, Yanjie Zhen, Wenlai Zhao, Guangwen Yang 0002
ICPADS3
2023 Achieving 10m China Land Cover Mapping within Three Minutes Using a New Sunway Supercomputer
abstract
Land Cover Mapping (LCM) is an important task to detect and understand the change of the earth surface. However, most LCM methods adopt supervised classifiers, and suffer from a lack of labels at a large scale. In this paper, we propose Fast-LCM, a scalable and weakly-supervised LCM method on a new Sunway supercomputer to achieve large-scale land cover mapping, requiring no manual annotations. Fast-LCM includes two major parts: (1) a distance-guided k-means module that combines textural, spectral, and temporal features, and (2) an automatic voting-based merging strategy to give each cluster a real meaning of classification system. Through careful parallelization, our Fast-LCM method scales to over 38 million cores, and provides a sustained performance for the task of China LCM. We produce a 10m resolution land cover map of China within only 3 minutes, including 1.2 minutes for IO and only 55 seconds to finish the computation. Fast-LCM achieves an accuracy of 68.85% (25-class), with 3.64% to 7.05% higher than best existing products.
Juepeng Zheng, Yi Zhao 0024, Jinxiao Zhang, Wenzhao Wu, Shuai Yuan 0005, Haohuan Fu
IGARSS4
2023 SW-LCM: A Scalable and Weakly-supervised Land Cover Mapping Method on a New Sunway Supercomputer
abstract
High-resolution land cover mapping (LCM) is an important application for studying and understanding the change of the earth surface. While deep learning (DL) methods demonstrate great potential in analyzing satellite images, they largely depend on massive high-quality labels. This paper proposes SW-LCM, a Scalable and Weakly-supervised two-stage Land Cover Mapping method on a new Sunway Supercomputer. Our method consists of a k-means clustering module as a first stage, and an iterative deep learning module as a second stage. With the k-means module providing a good enough starting point (taking inaccurate results as noisy labels), the deep learning module improves the classification results in an iterative way, without any labelling efforts required for processing large scenarios. To achieve efficiency for country-level land cover mapping, we design a customized data partition scheme and an on-the-fly assembly for k-means. Through careful parallelization and optimization, our k-means module scales to 98,304 computing nodes (over 38 million cores), and provides a sustained performance of 437.56 PFLOPS, in a real LCM task of the entire region of China; the iterative updating part scales to 24,576 nodes, with a performance of 11 PFLOPS. We produce a 10-m resolution land cover map of China, with an accuracy of 83.5% (10-class) or 73.2% (25-class), 7% to 8% higher than best existing products, paving ways for finer land surveys to support sustainability-related applications.
Yi Zhao 0024, Juepeng Zheng, Haohuan Fu, Wenzhao Wu, Mengxuan Chen, Jinxiao Zhang, Lixian Zhang 0002, Runmin Dong, Zhenrong Du, Xin Liu 0081, Shaoqing Zhang, Le Yu 0001
IPDPS4
2023 Partial Domain Adaptation for Scene Classification From Remote Sensing Imagery
abstract
Although domain adaptation approaches have been proposed to tackle cross-regional, multitemporal, and multisensor remote sensing applications since they do not require any human interpretation in the target domain, most current works assume identical label space across the source and the target domains. However, in real-world applications, we often transfer knowledge from a large-scale dataset with rich annotations to a small-scale target dataset with scarcity of labels. In most cases, the label space of the source domain is usually large enough to subsume that of the target domain, which is termed partial domain adaptation. In this article, we propose a new partial domain adaptation algorithm for remote sensing scene classification and our proposed method contains three major parts. First, we employ a progressive auxiliary domain module to alleviate the negative transfer effect caused by outlier classes. Second, we adopt an improved domain adversarial neural network (DANN) with multiweights to better encourage domain confusion. Last but not least, we design an attentive complement entropy regularization to improve the prediction confidence for samples and avoid untransferable samples (such as the samples belonging to outlier classes in the source domain) being mistakenly classified. We collect three common remote sensing datasets to evaluate our proposed method. Our method achieves an average accuracy of 79.36%, which considerably outperforms other state-of-the-art partial domain adaptation methods with an average accuracy improvement of 1.90%–12.45% and attaining a 13.67% gain compared to the straightforward deep learning model (ResNet-50). The experiment results indicate that our approach shows promising prospects for solving more general and practical domain adaptation problems where the label space of the source domain subsumes that of the target domain.
Juepeng Zheng, Yi Zhao 0024, Wenzhao Wu, Mengxuan Chen, Haohuan Fu
IEEE Trans. Geosci. Remote. Sens.3
2022 A Parallel Approach for Oil Palm Tree Detection on a SW26010 Many-Core Processor
abstract
Counting and detecting oil palm trees from high-resolution remotely sensed images is a significant work for improving economy of several countries such as Malaysia, Indonesia, etc. However, rare attention have been paid on accelerating tree crown detection algorithms on high performance platforms. In this paper, we design a parallel approach for oil palm tree detection on a SW26010 many-core processor, which is used in a world-leading supercomputer, Sunway TaihuLight. Our parallel framework contains three steps, local maximum filtering, oil palm tree crown center reassignment and oil palm tree crown center merging. Experimental results indicates that our parallel approach of oil palm tree detection obtains the speedup of 32.30 times and 1.74 times for a QuickBird image with a size of$12,188\times 12,576$pixels compared with the well-optimized software implementation of the original algorithm on an Intel 12-core CPU and FPGAs.
Juepeng Zheng, Wenzhao Wu, Yi Zhao 0024, Shuai Yuan 0005, Runmin Dong, Lixian Zhang 0002, Haohuan Fu
IGARSS2
2022 Multisource-Domain Generalization-Based Oil Palm Tree Detection Using Very-High-Resolution (VHR) Satellite Images
abstract
Providing accurate and timely oil palm information on a large scale is essential for both economic development and ecological significance. However, owing to different sensors, photograph acquisition conditions, and environmental heterogeneity, the large volume and the variety of the data make it extremely challenging for large-scale and cross-regional oil palm tree detection. It is computationally expensive to train a model from images covering large heterogeneous regions and all environmental conditions for continuously accumulated multisource remote sensing data. In this letter, we propose a new multisource domain generalization (DG) method, Maximum Mean Discrepancy Deep Reconstruction Classification Network (MMD-DRCN). It learns representations from multiple source domains and obtains inspiring performance in an unknown and “unseen” target domain. Besides classification loss, our MMD-DRCN distills more representative features through reconstruction loss and aligns multisource latent features by MMD loss, both of which effectively enhance the capacity of generalization. MMD-DRCN achieves an average F1-score of 82.70% in all transfer tasks, attaining a 5.83% gain compared to Baseline (a straightforward convolutional neural network (CNN) model). Experimental results demonstrate DG poses a promising potential for large-scale and cross-regional oil palm tree detection without any information of the target domain.
Juepeng Zheng, Wenzhao Wu, Shuai Yuan 0005, Haohuan Fu, Le Yu 0001
IEEE Geosci. Remote. Sens. Lett.2
2022 A Two-Stage Adaptation Network (TSAN) for Remote Sensing Scene Classification in Single-Source-Mixed-Multiple-Target Domain Adaptation (S²M²T DA) Scenarios
abstract
Over the past decade, domain adaptation (DA) algorithms have been proposed to address domain gap problems as they do not need any interpretation in the target domain. However, most existing efforts focus on scenarios with only one source domain and one target domain. In this article, we explore the scenario with one source domain and mixed multiple target domains for remote sensing applications and propose a new algorithm, named the two-stage adaptation network (TSAN). First, we utilize the adversarial learning approach to confuse the classifier to discriminate between the source domain and the whole mixed-multiple-target domain. Second, we adopt self-supervised learning to divide the mixed-multiple-target domain with automated generation of “pseudo”-domain labels, which guides our network to learn intrinsic features of multiple target domains. Finally, these two steps are combined as an iterative procedure. We integrate a test dataset that includes five remote sensing datasets and ten classes. Our method achieves an average accuracy of 63.25% and 73.68% with two typical backbones, considerably outperforming other DA methods with an average accuracy improvement of 4.84%–20.19% and 9.06%–17.04%, respectively. Furthermore, we identify the negative transfer effect in existing mainstream DA methods in remote sensing image classification with multiple different domains.
Juepeng Zheng, Wenzhao Wu, Shuai Yuan 0005, Yi Zhao 0024, Lixian Zhang 0002, Runmin Dong, Haohuan Fu
IEEE Trans. Geosci. Remote. Sens.2
2021 Transresnet: Transferable Resnet For Domain Adaptation
abstract
Although Deep Convolutional Neural Network (DCNN) has been admittedly witnessed as an enormous success in a wide range of applications, most of them require sufficient annotations with time-consuming and labor-exhausting efforts. Existing domain adaptation (DA) approaches delve into designing an effective loss module to minimize the distribution gap between the source and target domains. However, few studies pay attention to improve the backbone or network architecture for DA issues. In this paper, we propose a new backbone for DA specially, i.e., Transferable ResNet (TransResNet). TransResNet remedies the residual block in ResNet, separating source and target input features and highlighting more transferable channels in each block. It can be easily applied to all kinds of DA methods, without adding any extra learning parameters. We conduct substantial experiments on two general DA datasets and embed TransResNet into two seminal DA methods, including DANN and CDAN. Experimental results demonstrate TransResNet improves the transferability of the architecture, indicating that it is a great substitute for ResNet as a network backbone in DA issues.
Juepeng Zheng, Wenzhao Wu, Yi Zhao 0024, Haohuan Fu
ICIP2
2021 Coconut Trees Detection on the Tenarunga Using High-Resolution Satellite Images and Deep Learning
abstract
The Coconut tree is of great importance in economic values and ecological impacts for many tropical developing countries and lots of islands in the Pacific Ocean. Detecting and counting coconut is a meaningful and valuable research. In this paper, we present a coconut tree crown detection method to detect and count the coconut trees in the Tenarunga from high-resolution satellite images acquired by Google Earth. Our coconut tree detection method contains three major procedures: feature extraction, a multi-level Region Proposal Network (RPN) and a large-scale coconut tree detection workflow. We manually annotate all coconut trees for our study regions in the Tenarunga. Eventually, we achieve a higher average F1-score of 77.14% in our four test regions than pure Faster R-CNN. Experiment results demonstrate the potential for large-scale individual coconut tree detection and counting from high-resolution satellite images using deep learning.
Juepeng Zheng, Wenzhao Wu, Le Yu 0001, Haohuan Fu
IGARSS2
2021 Closing the "quantum supremacy" gap: achieving real-time simulation of a random quantum circuit using a new Sunway supercomputer
abstract
We develop a high-performance tensor-based simulator for random quantum circuits(RQCs) on the new Sunway supercomputer. Our major innovations include: (1) a near-optimal slicing scheme, and a path-optimization strategy that considers both complexity and compute density; (2) a three-level parallelization scheme that scales to about 42 million cores; (3) a fused permutation and multiplication design that improves the compute efficiency for a wide range of tensor contraction scenarios; and (4) a mixed-precision scheme to further improve the performance. Our simulator effectively expands the scope of simulatable RQCs to include the 10X10(qubits)X(1+40+1)(depth) circuit, with a sustained performance of 1.2 Eflops (single-precision), or 4.4 Eflops (mixed-precision)as a new milestone for classical simulation of quantum circuits; and reduces the simulation sampling time of Google Sycamore to 304 seconds, from the previously claimed 10,000 years.
Yong (Alexander) Liu, Xin (Lucy) Liu, Fang (Nancy) Li, Haohuan Fu, Yuling Yang, Jiawei Song, Pengpeng Zhao 0006, Dajia Peng, Huarong Chen, Chu Guo, Heliang Huang, Wenzhao Wu, Dexun Chen
SC13
2020 Unsupervised Mixed Multi-Target Domain Adaptation for Remote Sensing Images Classification
abstract
Although deep learning has been successfully applied in the field of remote sensing image classification, it still requires time-consuming and costly annotations. In recent years, domain adaptation has been witnessed to address this problem as they do not need any human interpreted in the target domain dataset. However, most of the existing works dedicate effort on the circumstance where there is only one source domain and only one target domain. In this paper, we firstly explore one source and multiple target domains issue for remote sensing application and build a challenging mixed multi-target dataset to contribute to the community. Our method constitutes three parts. Firstly, as we are blind for the multitarget domain, we adopt meta learning to divide the mixed multi-target dataset and insert sub-target domain loss as the part of the loss function. Secondly, we apply the adversarial learning to confuse the classifier to discriminate between the source domain images and the whole mixed multi-target domain images. Finally, the meta learning and the adversarial learning are dynamically iterative procedures and the labels for domain classification in mixed multi-target dataset will be updated for a particular iteration. Our method is well-performed in the four common remote sensing dataset (AID, NWPU-RESISC45, UC Merced and WHU-RS19), including five classes (agriculture, forest, river, residential and parking). Our method achieved an average accuracy of 81.59% and outperformed other domain adaptation method. The experiment results indicate our method is promising for large-scale, multi-regional and multi-temporal remote sensing applications.
Juepeng Zheng, Wenzhao Wu, Haohuan Fu, Runmin Dong, Lixian Zhang 0002, Shuai Yuan 0005
IGARSS2