VLDB 2026 Research / reviewers in the wild / expert
Huachen Zhang
dblp:58/802
· DBLP profile ↗
11ranked-venue papers
2as first author
10since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 2 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Computer networks · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Configurable Streaming Accelerator for LUT-Based Super-Resolution on FPGA
Xuzhuo Hu, Jianyang Ding, Huachen Zhang, Tianshuo Lu, ZhiLei Chai |
APPT | 4 |
| 2026 | RACP: An Efficient RISC-V Domain-Specific Processor for Arbitrary-Size Kernel CNNs
Jianyang Ding, Tianshuo Lu, Huachen Zhang, Xuzhuo Hu, ZhiLei Chai |
ISCAS | 4 |
| 2026 | An efficient RISC-V processor with customized instruction set for sparse DNN acceleration on embedded system
Jianyang Ding, Huachen Zhang, Tianshuo Lu, ZhiLei Chai |
J. Syst. Archit. | 3 |
| 2025 | OMGAN: One-to-Many Generative Adversarial Network for Diagnosing Orbital Lymphoproliferative Disorders in Incomplete Multi-Parametric MRIabstractMulti-parametric magnetic resonance imaging (mpMRI) is widely used in the diagnosis of orbital lymphoproliferative disorders (OLPDs) due to its non-invasive nature. However, in clinical practice, contrast-enhanced T1-weighted (T1C) images are often unavailable due to contraindications to gadolinium-based contrast agents, meanwhile T2-weighted (T2w) images may also be omitted for time-sensitive diagnoses, making it a challenge to generate these images from T1-weighted (T1 w) image alone for multimodal differential diagnosis. Generative adversarial network (GAN)-based models partially address the issue of missing modalities in medical image analysis; however, they often suffer from unstable generation of missing images and lack integration with subsequent diagnostic tasks. To this end, we propose a One-to-Many Generative Adversarial Network (OMGAN) for diagnosing OLPDs in incomplete mpMRI, consisting of a cross-modal generator and a self-representation module, enabling multimodal diagnosis using pre-contrast images alone within a single model. Specifically, we first design an image-modality fusion module that incorporates trigonometric function coding and mixup augmentation to effectively guide the generation from T1 w to T2w and T1 C within one model. Then, we construct a cross-modal generator with a semantic disambiguation block to synthesize the missing images. Meanwhile, we use a self-representation module with a classification-guided branch to effectively extract task-relevant image features. Finally, multimodal features are fused to accomplish the differential diagnosis of OLPDs in the downstream task. Experiments on internal datasets demonstrated that OMGAN outperforms state-of-the-art GAN-based models, with the area-under-the-curve and accuracy improving by 8.18-14.04% and 12.53-16.39%, respectively. Codes are available at https://github.com/3Iasticheart/OMGAN. Yuanxin Zhao, Fengjun Zhao, Huachen Zhang, Xuelei He, Xiaowei He 0001 |
BIBM | 3 |
| 2025 | Hybrid-SANet: Hybrid Self-attention Transformer for Efficient Image Super-Resolution
Jianyang Ding, Huachen Zhang, Nachuan Zhang, Tianshuo Lu, ZhiLei Chai |
CGI (3) | 3 |
| 2025 | EVO-QNN: Efficient Mixed-Precision Quantization Inference on RISC-V-Based Edge DeviceabstractMixed-Precision Quantized Neural Network (MPQNN) helps balance inference precision and efficiency under resource constraints, while most of them lack high-energy-efficiency hardware acceleration solutions. To address these challenges, we propose a SW/HW co-design framework termed EVO-QNN for low-energy and low-latency inference. Specifically, EVO-QNN manages bit-widths of operators and extends SIMD instructions based on customized RISC-V core. Experimental results demonstrate that our framework can achieve performance improvement ranging from 1.23x to 1.58x, with only a 1.02% and 6.61% increasing in area and power consumption for 2–8 bit convolution operators. Tianshuo Lu, Jianyang Ding, Huachen Zhang, ZhiLei Chai |
FCCM | 3 |
| 2025 | RV-ESMC: Efficient Sparse Matrix Convolution Processor based on RISC-V Custom instructions for Edge PlatformsabstractAs the demand for deep neural network (DNN) inference on edge platforms grows, deploying compute-intensive DNNs on resource-constrained devices remains challenging. This paper proposes a novel sparse convolution acceleration processor, RV-ESMC, based on RISC-V architecture, with custom instructions to enable efficient edge DNN inference. RV-ESMC provides flexibility by supporting inline assembly calls in C programming. Experimental results indicate that RV-ESMC can reduce execution time by over 70% in DNNs with convolution operations compared to conventional instruction sets. The functionality of RV-ESMC is validated on an FPGA platform and its performance is comprehensively evaluated based on a 55nm CMOS process. The results show that RV-ESMC can achieve a peak energy efficiency of 675 GOPS/W. Huachen Zhang, Jianyang Ding, Tianshuo Lu, ZhiLei Chai |
FCCM | 1 |
| 2025 | Optimizing Sparse Matrix Convolution on RISC-V Core: Custom Instructions for Embedded SystemabstractWith the increasing demand for deep neural network (DNN) inference tasks on embedded platforms, deploying compute-intensive DNNs on resource-constrained embedded platforms faces challenges. While sparsification technology offers a potential solution, its implementation on edge platforms still faces difficulties. In this article, we propose a novel sparse convolution acceleration processor based on RISC-V architecture, and design specialized custom instructions to enable efficient edge DNN inference. To this end, we mainly address three technical issues. In response to numerical characteristics of sparse convolution, the designed processor can implement a hardware-friendly architecture that transforms convolutions into sparse matrix multiplication. Additionally, it employs a column-major and element-level parallel strategy to optimize load imbalance issues present in the Gustavson algorithm, thereby enhancing sparse matrix computations. To further improve computational efficiency, our work is designed by incorporating efficient execution units that reduce instruction execution overheads while minimizing memory access frequency. Compared to traditional accelerators, our work supports custom instruction formats in the C programming language, offering superior flexibility. Extensive experimental results indicate that our work can reduce execution time over 70% when running most DNNs with convolution operations compared to conventional instruction sets. Moreover, the functionality of our work is validated on an FPGA platform, and its performance is comprehensively evaluated based on a 55 nm CMOS process. The results show that our work can achieve a peak energy efficiency of 675 GOPS/W in most network inference tasks, demonstrating exceptional computational performance and energy efficiency. Huachen Zhang, Jianyang Ding, Tianshuo Lu, ZhiLei Chai |
ACM Trans. Embed. Comput. Syst. | 1 |
| 2025 | ISRLUT: Integer-Only FHD Image Super-Resolution Based on Neural Lookup Table and Near-Memory ComputingabstractWhile Deep Neural Networks (DNNs) have achieved remarkable progress in Image Super-Resolution (SR) task, they face significant challenges for edge processing FHD images. Complex DNN operators lead to high hardware resource consumption and latency. Computational inefficiency of FPU increases energy consumption, while DDR access overhead and on-chip memory overflow further constrain real-time capabilities. To address this, we propose ISRLUT, a novel accelerator architecture focused on integer-only inference and near-memory computing. Its core contributions include: (1) Fusion of Neural LUT arithmetic with reconfigurable compute units, transforming unified LUT operators from DNN operators and enhancing hardware utilization; (2) An integer-only inference and parallel architecture, eliminating floating-point dependencies and significantly reducing energy consumption; (3) An innovative internal operator memory management scheme coupled with Tile-based Buffer Overlap and Private Cache Mechanism. We deploy ISRLUT on FPGA and ASIC platforms. Experiments demonstrate that ISRLUT achieves efficient performance: For 4 \(\times\) upscaling, it requires only 36.9 KB of storage and achieves a PSNR of 30.21 dB on Set5. Hardware implementation using a 55 nm ASIC consumes merely 0.0337 W power, delivers an energy efficiency of 7278.6 Mpixels/s/W, and achieves a real-time frame rate of 118 FPS for 4 \(\times\) FHD processing, validating its superiority in energy efficiency and hardware utilization. Tianshuo Lu, Jianyang Ding, Huachen Zhang, ZhiLei Chai |
ACM Trans. Reconfigurable Technol. Syst. | 4 |
| 2023 | Orbital Lymphoproliferative Disorder Diagnosis with Incomplete Multimodal Images based on Self-/Cross-Representation and Hypergraph EnsembleabstractOrbital lymphoproliferative disorders (OLPDs) are complex orbital mass-like lesions ranging from benign to malignant. Precise preoperative diagnosis of OLPDs holds profound importance in facilitating timely and effective patient management. Recent studies have shown that exploiting multimodal images can boost the performance in identifying different orbital lesions. However, one or several imaging modalities are sometimes missing in practical applications, which has not yet been properly addressed in existing studies. To this end, we propose a novel OLPD diagnostic method with incomplete multimodal images based on self-/cross-representation and hypergraph ensemble. Specifically, in the first stage, we develop a self-representation network to extract unimodal features and a cross-representation network to impute missing features. In the second stage, by using unimodal features as input, we construct a hypergraph for each modality to make unimodal diagnosis; while for multimodal diagnosis we conduct a multi-view grouping fusion method to reduce the semantic gap between multimodal features and fuse multiple unimodal hypergraphs as multimodal hypergraph to perform multimodal diagnosis. In the third stage, we propose an ensemble strategy that incorporates unimodal diagnosis and multimodal diagnosis to accomplish the final decision. Extensive experiments demonstrate that the proposed model outperforms the state-of-the-art approaches. Xiaoyang Xie, Huachen Zhang, Yuqing Hou, Xiaowei He 0001, Fengjun Zhao |
BIBM | 2 |
| 2007 | Performance analysis and evaluation for multi-traffic networks with priority control
Wuyi Yue, Dequan Yue, Huachen Zhang, Fengsheng Tu |
Comput. Commun. | 3 |