EDBT 2026 Demo / reviewers in the wild / expert
Zhuoya Wang
dblp:221/8834
· DBLP profile ↗
8ranked-venue papers
2as first author
8since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021Systems, architecture and hardware · 3 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SASTD: Stepwise attention style transfer network based on diffusion models
Zhuoya Wang, Yongsheng Dong 0002 |
Comput. Vis. Image Underst. | 1 |
| 2025 | 3D-Domino: Ultra-Dense High-Accuracy 3D eDRAM-ROM Compute-In-Memory Based on CAA-IGZO TFT for Edge Large-Scale Model InferenceabstractThe rapid growth in the parameter count of large language models (LLMs) in recent years has placed higher demands on the density of compute-in-memory (CiM) solutions. Read-Only memory (ROM), due to its high-density advantages, has emerged as a promising CiM cell type, offering substantial task-level energy efficiency improvements over SRAM CiM. However, traditional 2D ROM CiM approaches are limited by 2D fabrication constraints, restricting scalability for LLM deployment. To address this limitation, this work explores a novel 3D back-end-of-line (BEOL)-compatible device, the channel-all-around (CAA)-IGZO TFT. Here, we propose a 3D ROM CiM with an ultra-dense cell structure and a high-throughput computing scheme. Additionally, we introduce a hybrid 3D CiM accelerator architecture that integrates both ROM and eDRAM for unprecedented density and flexibility. Evaluation results show that the proposed 3D ROM CiM, with 16 CAA-IGZO stacked layers, achieves an ultra-high memory density of 31.19 Mb/mm2/layer, a computation density of 167.6 TOPS/mm2, and high computing accuracy with a compute SNR (CSNR) of 22.6 dB, underscoring its potential for edge large-scale model acceleration. Based on this, when deployed with a LoRA-tuned GPT-2 model, the proposed hybrid 3D eDRAM-ROM architecture shows 1.7× improvement in area efficiency compared to the eDRAM-only counterpart. Zhuoya Wang, Huazhong Yang, Narayanan Vijaykrishnan, Xueqing Li 0002 |
ISCAS | 1 |
| 2024 | Target-aware Guided equivariant Diffusion model for 3D molecule GenerationabstractIn the process of targeted drug molecule design, models that incorporate three-dimensional structures show better performance than target-free models. This is because the interactions between atoms can be explicitly modeled in three-dimensional space, thereby improving the accuracy and effectiveness of drug design. In previous studies, researchers usually used static or semi-flexible crystal structures to simplify the dynamic interactions between active sites and small molecules, only calculating limited system dynamics information. However, this simplification leads to an inadequate understanding of the dynamic characteristics of active sites, neglecting the consideration of the binding dynamics during molecule generation, which impacts the generation of high-quality 3D molecules. To this end, we proposed a 3D molecule generation method based on a target-aware guided diffusion model. This method uses target structure as a condition, introducing a diffusion model to simulate the dynamic evolution of molecular structures, and further refines and optimizes ligand through the structure of target-ligand complexes. Additionally, the continuous distribution of atomic coordinates and the discrete distribution of molecular features are introduced in the molecular latent space through the equivariant graph neural network, which facilitates the representation of the parameterized reverse generation process. Experimental results show that this method can continuously achieve better performance on multiple molecular generation benchmarks and can generate more realistic multiple 3D structure molecules with high binding affinity to protein targets. Xiaotong Hu, Zhuoya Wang |
BIBM | 4 |
| 2024 | swCUDA: Auto parallel code translation framework from CUDA to ATHREAD for new generation sunway supercomputerabstractAbstract Since specific hardware characteristics and low-level programming model are adapted to both NVIDIA GPU and new generation Sunway architecture, automatically translating mature CUDA kernels to Sunway ATHREAD kernels are realistic but challenging work. To address this issue, swCUDA, an auto parallel code translation framework is proposed. To that end, we create scale affine translation to transform CUDA thread hierarchy to Sunway index, directive based memory hierarchy and data redirection optimization to assign optimal memory usage and data stride strategy, directive based grouping-calculation-asynchronous-reduction (GCAR) algorithm to provide general solution for random access issue. swCUDA utilizes code generator ANTLR as compiler frontend to parse CUDA kernel and integrate novel algorithms in the node of abstracted syntax tree (AST) depending on directives. Automatically translation is performed on the entire Polybench suite and NBody simulation benchmark. We get an average 40x speedup compared with baseline on the Sunway architecture, average speedup of 15x compared to x86 CPU and average 27 percentage higher than NVIDIA GPU. Further, swCUDA is implemented to translate major kernels of the real world application Gromacs. The translated version achieves up to 17x speedup. Maoxue Yu, Guanghao Ma, Zhuoya Wang, Yuhu Chen, Yucheng Wang 0002, Dongning Jia |
CCF Trans. High Perform. Comput. | 3 |
| 2023 | UCAPF: A Unified Processing Platform for Large-scale Virtual ScreeningabstractWith growing data volume for large-scale virtual screening, the associated data processing and management meet challenges. We have developed UCAPF, A unified platform for large-scale virtual screening. The platform provides a parallel processing framework for large-scale virtual screening data. It also enables scheduling of heterogeneous parallel architectures and hierarchical storage of massive data. The processing framework improves data quality. On the CASF-2016 dataset, the standardized molecules processed by UCAPF showed a 9.5% to 14.62% improvement in scoring performance and a 7.9% to 34.6% improvement in ranking performance compared to the raw molecules. For massive data processing, the framework provides parallel efficiency of 81.20% for molecule standardized processing and 79.51% for docking result processing on a 72-unit Hadoop cluster. In addition, the distributed database for data management improves the ability to retrieve 10,094 molecules from seventy million docking result data by a factor of 2.40 compared to the single-node storage model. Finally, we analyze the variation of input/output (I/O) over time for different phases of virtual screening to reflect the effectiveness of the scheduling strategy and tiered storage for the heterogeneous parallel architecture. Zhuoya Wang, Hong Guang, Weidan Wang, Yujie Dong |
BIBM | 5 |
| 2023 | Molecular generation strategy and optimization based on A2C reinforcement learning in de novo drug designabstractMOTIVATION: In the field of pharmacochemistry, it is a time-consuming and expensive process for the new drug development. The existing drug design methods face a significant challenge in terms of generation efficiency and quality. RESULTS: In this paper, we proposed a novel molecular generation strategy and optimization based on A2C reinforcement learning. In molecular generation strategy, we adopted transformer-DNN to retain the scaffolds advantages, while accounting for the generated molecules' similarity and internal diversity by dynamic parameter adjustment, further improving the overall quality of molecule generation. In molecular optimization, we introduced heterogeneous parallel supercomputing for large-scale molecular docking based on message passing interface communication technology to rapidly obtain bioactive information, thereby enhancing the efficiency of drug design. Experiments show that our model can generate high-quality molecules with multi-objective properties at a high generation efficiency, with effectiveness and novelty close to 100%. Moreover, we used our method to assist shandong university school of pharmacy to find several candidate drugs molecules of anti-PEDV. AVAILABILITY AND IMPLEMENTATION: The datasets involved in this method and the source code are freely available to academic users at https://github.com/wq-sunshine/MomdTDSRL.git. Zhiqiang Wei 0002, Xiaotong Hu, Zhuoya Wang, Yujie Dong, Hao Liu 0045 |
Bioinform. | 4 |
| 2023 | Efficient Large-Scale Virtual Screening Based on Heterogeneous Many-Core Supercomputing SystemabstractWith the rapid growth of virtual drug data- bases, the need for efficient molecular docking tools for large-scale screening is also growing. We have developed Vina@QNLM 2.0, a novel molecular docking system that leverages the logical processing units and computational processing arrays of heterogeneous multicore architecture processors. Compared to Vina@QNLM, the new version optimizes the docking speed without sacrificing accuracy. This greatly improves the scoring capability for large molecules (molecular weight > 500). Simultaneously, the new system provides enhanced support for applications such as reverse target finding through an improved parallel strategy. Vina@QNLM 2.0 achieves a speedup 20 times higher than that, using logical processing units only during a single docking process. Additionally, we successfully scaled the reverse target finding a task to 122,401 kernel groups with a robust scalability of 80.01%. In practice, we completed a reverse target-seeking for nine glycan molecules with 10,094 proteins within 1 hour. Hao Liu 0045, Cunji Wang, Chengchao Liu, Zhuoya Wang, Zhiqiang Wei 0002 |
IEEE J. Biomed. Health Informatics | 5 |
| 2022 | Large-Scale Simulation of Quantum Computational Chemistry on a New Sunway SupercomputerabstractQuantum computational chemistry (QCC) is the use of quantum computers to solve problems in computational quantum chemistry. We develop a high performance variational quantum eigensolver (VQE) simulator for simulating quantum computational chemistry problems on a new Sunway supercomputer. The major innovations include: (1) a Matrix Product State (MPS) based VQE simulator to reduce the amount of memory needed and increase the simulation efficiency; (2) a combination of the Density Matrix Embedding Theory with the MPS-based VQE simulator to further extend the simulation range; (3) A three-level parallelization scheme to scale up to 20 million cores; (4) Usage of the Julia script language as the main programming language, which both makes the programming easier and enables cutting edge performance as native C or Fortran; (5) Study of real chemistry systems based on the VQE simulator, achieving nearly linearly strong and weak scaling. Our simulation demonstrates the power of VQE for large quantum chemistry systems, thus paves the way for large-scale VQE experiments on near-term quantum computers. Honghui Shang, Li Shen 0001, Zhiqian Xu 0005, Chu Guo, Jie Liu 0069, Rongfen Lin, Yuling Yang, Zhuoya Wang, Yunquan Zhang |
SC | 12 |