EDBT 2026 Demo / reviewers in the wild / expert
Hao Zhang 0058
dblp:55/2270-58
· DBLP profile ↗
14ranked-venue papers
4as first author
13since 2021 · last 2026
0000-0003-1863-7440ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 4 first-author · 7 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Bridging the image-text gap: Reinforced Cross-modal Abnormality Driven Transformer for automatic chest X-ray report generation
Xiu-Long Yi, You Fu, Enxu Bi, Hao Zhang 0058, Jianguo Liang, Rong Hua |
Eng. Appl. Artif. Intell. | 5 |
| 2026 | Radiology report generation via visual-semantic ambivalence-aware network and focal self-critical sequence training
Xiu-Long Yi, You Fu, Enxu Bi, Jianguo Liang, Hao Zhang 0058, Jianzhi Yu, Rong Hua |
Neural Networks | 5 |
| 2025 | BITLUME: Precision-Flexible Photonic Computing for Ultra-Fast and Energy-Efficient DNN AccelerationabstractAs deep learning expands across emerging domains, computational demands are pushing traditional electronic accelerators to their limits. Silicon photonics has emerged as a promising technology for accelerating deep learning workloads, but precision remains a challenge due to noise and non-idealities. In this paper, we present BITLUME, a novel photonic computing unit that enables multiplications beyond 8-bit precision through a precision-flexible scheme. We further propose an optimized round-truncation algorithm and data mapping strategy for BITLUME to reduce optoelectronic conversions, enhance data reuse, and maintain computational accuracy. A hybrid optoelectronic architecture integrating BITLUME is developed and validated using a prototype built with FPGA, RF, and photonic components, achieving 3.7× lower end-to-end latency than the A100 GPU in dot product. Simulations of training seven DNN models at FP32 show that BITLUME achieves up to 3.35× and 10.78× speedup, and 1.53× and 4.12× energy savings, compared to the state-of-the-art photonic accelerator and A100 GPU, respectively. Chengpeng Xia, Haibo Zhang 0001, Hao Zhang 0058, Yawen Chen 0001, Amanda S. Barnard |
ICCAD | 3 |
| 2025 | ROCKET: An RNS-based Photonic Accelerator for High-Precision and Energy-Efficient DNN TrainingabstractIn recent years, the rapid development of Deep Neural Networks (DNNs) has posed significant challenges in terms of training duration and costs. High-frequency, low-power photonic computing has emerged as a highly promising solution. However, the substantial cost of data conversion and the limitations introduced by noise in photonic devices continue to hinder the realization of high-precision and energy-efficient DNN training. To address this challenge, we propose a novel photonic accelerator, ROCKET, based on the Residue Number System (RNS). RNS is based on modular arithmetic and enables support for high-precision computation through parallel multi-path low-precision operations. First, we leverage specialized lookup tables to enable high-throughput, low-latency conversions between high-precision and low-precision numerical representations. Next, we design a low-power photonic accelerator architecture utilizing intensity modulators, which minimizes the number of computational components while maximizing data reuse. Subsequently, we propose a hybrid photonic-electronic pipelined dataflow to maximize parallelism within the photonic-electronic computation path. Finally, we develop a high-frequency (4.096 GHz) hybrid photonic-electronic prototype using FPGA, Radio Frequency (RF), and photonic components to validate the feasibility of the ROCKET. Our large-scale simulations on seven mainstream DNN models show that, compared to the A100 GPU, TPU v4, and the state-of-the-art photonic accelerator Mirage, ROCKET achieves speedups of 33×, 243×, and 198×, respectively, while saving energy by factors of 64×, 204×, and 142×. Hao Zhang 0058, Haibo Zhang 0001, Chengpeng Xia, Zhiyi Huang 0001, Yawen Chen 0001, Amanda S. Barnard |
ICS | 1 |
| 2025 | ChipAI: A scalable chiplet-based accelerator for efficient DNN inference using silicon photonicsabstractTo enhance the precision of inference, deep neural network (DNN) models have been progressively growing in scale and complexity, leading to increased latency and computational resource demands. This growth necessitates scalable architectures, such as chiplet-based accelerators, to accommodate the substantial volume of deep learning inference tasks. However, the efficiency, energy consumption, and scalability of existing accelerators are severely constrained by metallic interconnects. Photonic interconnects, on the contrary, offer a promising alternative, with their advantages of low latency, high bandwidth, high energy efficiency, and simplified communication processes. In this paper, we propose ChipAI, an accelerator designed based on photonic interconnects for accelerating DNN inference tasks. ChipAI implements an efficient hybrid optical network that supports effective inter-chiplet and intra-chiplet data sharing, thereby enhancing parallel processing capabilities. Additionally, we propose a flexible dataflow leveraging the ChipAI architecture and the characteristics of DNN models, facilitating efficient architectural mapping of DNN layers. Simulation on various DNN models demonstrates that, compared to the state-of-the-art chiplet-based DNN accelerator with photonic interconnects, ChipAI can reduce the DNN inference time and energy consumption by up to 82% and 79%, respectively. Hao Zhang 0058, Haibo Zhang 0001, Zhiyi Huang 0001, Yawen Chen 0001 |
J. Syst. Archit. | 1 |
| 2025 | LHR-RFL: Linear Hybrid-Reward-Based Reinforced Focal Learning for Automatic Radiology Report GenerationabstractRadiology report generation that aims to accurately describe medical findings for given images, is pivotal in contemporary computer-aided diagnosis. Recently, despite considerable progress, current radiology report generation models still struggled to achieve consistent quality across difficult and easy samples, which dramatically impacts their clinical value. To solve this problem, we explore the difficult samples mining in radiology report generation and propose the Linear Hybrid-Reward based Reinforced Focal Learning (LHR-RFL) to effectively guide the model to allocate more attention towards some difficult samples, thereby enhancing its overall performance in both general and intricate scenarios. In implementation, we first propose the Linear Hybrid-Reward (LHR) module to better quantify the learning difficulty, which employs a linear weighting scheme that assigns varying weights to three representative Natural Language Generation (NLG) evaluation metrics. Then, we propose the Reinforced Focal Learning (RFL) to adaptively adjust the contributions of difficult samples during training, thereby augmenting their impact on model optimization. The experimental results demonstrate that our proposed LHR-RFL improves the performance of the base model across all NLG evaluation metrics, achieving an average performance improvement of 20.9% and 13.2% on IU X-ray and MIMIC-CXR datasets, respectively. Further analysis also proves that our LHR-RFL can dramatically improve the quality of reports for difficult samples. The source code will be available at https://github.com/ SKD-HPC/LHR-RFL. Xiu-Long Yi, You Fu, Jianzhi Yu, Ruiqing Liu, Hao Zhang 0058, Rong Hua |
IEEE Trans. Medical Imaging | 5 |
| 2024 | TSGET: Two-Stage Global Enhanced Transformer for Automatic Radiology Report GenerationabstractRecently, automatic radiology report generation, which targets to generate multiple sentences that can accurately describe medical observations for given X-ray images, has gained increasing attention. Existing methods commonly employ the attention mechanism for accurate word generation. However, such attention-based methods fail to leverage useful image-level global features, thereby limiting the model's reasoning ability. To tackle this challenge, we propose two-stage global enhancement layers to facilitate the Transformer to generate more reliable reports from a global perspective. Specifically, the 1st Global Enhancement Layer (1st GEL) is designed to capture the global visual context features by establishing the relationships between image-level global features and previously generated words. The 2nd Global Enhancement Layer (2nd GEL) is devised to capture the region-global level features by building the relationships between image-level global features and region-level information. The experiments demonstrate that by integrating the aforementioned two-stage global enhancement layers into the Transformer model, our proposal achieves state-of-the-art (SOTA) performance on various Natural Language Generation (NLG) evaluation metrics. Further Clinical Efficacy (CE) evaluations also validate that our proposal is able to predict more critical information. Xiu-Long Yi, You Fu, Ruiqing Liu, Hao Zhang 0058, Rong Hua |
IEEE J. Biomed. Health Informatics | 4 |
| 2023 | SEECHIP: A Scalable and Energy-Efficient Chiplet-based GPU Architecture Using Photonic LinksabstractThe continuous increase in GPU performance benefits a wide range of high-performance computing (HPC) applications. Slower growth of transistor density and limited size of chip die are now posing significant challenges to scale GPUs. The chiplet technology provides a potential solution to surpass these limitations. However, the performance of these chiplet-based GPUs is often constrained by the metallic-based interconnects between the chiplets. Emerging technologies such as photonic interconnect can overcome the limitations of metallic interconnects, offering several superior properties, such as high bandwidth density and low energy consumption. In this paper, we propose SEECHIP: a Scalable and Energy-Efficient CHIPlet-based GPU architecture using photonic links. SEECHIP introduces a novel photonic inter-chiplet network that supports both unicast and broadcast communication, providing the same transmission bandwidth at both the sending and receiving ends. In addition, we propose a tailored hierarchical memory architecture, which is more suitable for the parallelization of general-purpose HPC applications. Simulation results using 14 benchmarks show that SEECHIP can achieve and reduction in execution time and energy consumption, respectively, as compared to other GPUs with metallic or photonic interconnects. Simulation results also show that SEECHIP has good scalability compared with the other GPUs. Hao Zhang 0058, Yawen Chen 0001, Zhiyi Huang 0001, Haibo Zhang 0001 |
ICPP | 1 |
| 2023 | ESA: An efficient sequence alignment algorithm for biological database search on Sunway TaihuLight
Hao Zhang 0058, Zhiyi Huang 0001, Yawen Chen 0001, Jianguo Liang, Xiran Gao |
Parallel Comput. | 1 |
| 2023 | Comparing the performance of multi-layer perceptron training on electrical and optical network-on-chips
Yawen Chen 0001, Zhiyi Huang 0001, Haibo Zhang 0001, Hao Zhang 0058, Chengpeng Xia |
J. Supercomput. | 5 |
| 2022 | OpenACC + Athread collaborative optimization of Silicon-Crystal application on Sunway TaihuLight
Jianguo Liang, Rong Hua, Yuxi Ye, You Fu, Hao Zhang 0058 |
Parallel Comput. | 6 |
| 2022 | A novel acceleration method for molecular dynamics of crystal silicon on GPUs using OpenACCabstractAbstract Compared with CUDA and OpenCL, OpenACC has the advantages of simple programming, openness, and good portability for GPU acceleration. An OpenMP/OpenACC implementation for molecular dynamics of silicon crystal on GPUs is proposed. First, to make effective use of vectorization and streaming, data structure conversion and data dependence elimination are designed. Second, the parallel version on the single GPU is realized by adding OpenACC guidance sentences, with very few modifications. Third, a patch block strategy is proposed to realize the parallel version on single machine multi‐GPUs using OpenMP+OpenACC, which greatly simplifies the construction of shadow area and the exchange of shadow area data. Experimental results show that 23 to 25 speedup is achieved for the single GPU at different scales over the serial program on Intel(R) Xeon(R) CPU E5‐2690 v4, and 6.37 speedup is achieved over the single GPU when the number of atoms reaches 2,097,152 on 8GPUs on single machine. Jianguo Liang, You Fu, Rong Hua, Hao Zhang 0058, Yuxi Ye |
Softw. Pract. Exp. | 4 |
| 2021 | Photonic Computing and Communication for Neural Network Accelerators
Chengpeng Xia, Yawen Chen 0001, Haibo Zhang 0001, Hao Zhang 0058, Jigang Wu |
PDCAT | 4 |
| 2020 | Accelerated molecular dynamics simulation of Silicon Crystals on TaihuLight using OpenACC
Jianguo Liang, Rong Hua, Hao Zhang 0058, You Fu |
Parallel Comput. | 3 |