EDBT 2026 Demo / reviewers in the wild / expert
Feng Wang 0050
dblp:90/4225-50
· DBLP profile ↗
14ranked-venue papers
2as first author
7since 2021 · last 2026
0009-0002-4009-6783ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 9 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MirageNet: A Secure, Efficient, and Scalable DNN Protection for Edge-Computing Multimedia RetrievalabstractDeploying multimedia deep neural networks (DNNs) on edge devices reduces retrieval and inference latency, but untrusted hardware in this setting can easily leak model parameters. Existing TEE-based protections have limitations: the partial-weight obfuscation methods are vulnerable due to statistical flaws, while full-weight obfuscation struggles to balance security and efficiency. To address these, we propose a convolution decomposition obfuscation scheme (MirageNet) for protecting edge multimedia DNNs in TEE–GPU heterogeneous environments. This scheme relieves the flaw that cosine similarity remains highly consistent between pre-trained and fine-tuned models. It obfuscates via convolution-kernel element-wise operations, decoy-kernel injection, and channel/kernel permutations, maximizing GPU utilization while minimizing TEE overhead. Experiments show that the MirageNet scheme reduces attack success to a black-box level, lowers runtime overhead by 15% compared to SOTA, and preserves inference accuracy identical to the original model—meeting practical edge multimedia retrieval demands. Huadi Zheng, Yuanhang Yu, Feng Wang 0050 |
ICMR | 6 |
| 2026 | A Decoupled Analytical Model for Tile Size Selection in Affine ProgramsabstractExisting tile size selection approaches are tightly coupled with compiler transformation pipelines, often leading to inaccurate modeling of cache behavior and limited effectiveness for non-rectangular tile shapes. This article presents TileMind , a decoupled analytical model that combines compile-time and runtime information for tile size selection in affine programs. It introduces a transformation-aware pre-tiling step that enables the decoupled selector to remain consistent with compiler transformations while extracting compile-time metadata. The extracted metadata is then combined with profiled runtime characteristics to construct a richer yet tractable feasible domain, within which a nonlinear objective for tile size selection is formulated. This objective is subsequently transformed into a binary product linearization problem, with its nonlinear constraints also linearized for efficient optimization. Finally, an intra-tile optimization aligns computation with data layout to enhance data reuse within tiles. Across two multi-core Intel CPUs, TileMind achieves 1.49× (sequential) and 1.33× (parallel) mean speedups on twenty PolyBench kernels, and 2.08–3.54× speedups on three deep learning workloads over the state-of-the-art analytical model Pluto-tss . Compared with TVM’s latest autotuner MetaSchedule, TileMind delivers 1.35–1.46× mean speedups while reducing tuning overhead by 2–4 orders of magnitude. While demonstrating effectiveness on selecting tile sizes for non-rectangular tile shapes and compatibility with PPCG, Pluto, and TVM, we further provide proof-of-concept results on GPUs, illustrating the potential portability of TileMind across architectures. Shihan Yuan, Zuoyan Zhang, Guanghui Song, Junhui Peng, Feng Wang 0050, Zhuo Tang, Kenli Li 0001, Jie Zhao 0002 |
ACM Trans. Archit. Code Optim. | 5 |
| 2025 | Post-Link Outlining for Code Size ReductionabstractThis paper introduces PLOS, a novel Post-Link Outlining approach designed to enhance code Size reduction for resource-constrained environments. Built on top of a post-link optimizer BOLT, PLOS maintains a holistic view of the whole-program structure and behavior, utilizing runtime information while preserving standard build system flows. The approach includes a granular outlining algorithm that matches and replaces repeated instruction sequences within/across modules and outlined functions, along with careful stack frame management to ensure correct function call handling. By integrating profiling information, PLOS balances the trade-off between code size and execution efficiency. The evaluation using eight MiBench benchmarks on an ARM-based Phytium FCT662 core demonstrates that PLOS achieves a mean code size reduction of 10.88% (up to 43.53%) and 6.61% (up to 14.78%) compared to LLVM’s and GCC's standard optimization, respectively, 1.76% (up to 4.75%) over LLVM’s aggressive code size reduction optimizations, and 2.88% (up to 8.56%) over a link-time outliner. The experimental results also show that PLOS can achieve a favorable balance between code size reduction and performance regression. Shaobai Yuan, Jihong He, Yihui Xie, Feng Wang 0050, Jie Zhao 0002 |
CC | 4 |
| 2025 | Scaling Deep Learning Molecular Dynamics to 500M Atoms on 4096-Node ARMv8 ClustersabstractMolecular dynamics (MD) simulations are essential tools for investigating large-scale molecular systems, yet achieving high performance and scalability on CPU-based architectures remains challenging. In this study, we present a highly optimized framework based on DeepMD-kit for conducting 500 millionatom MD simulations on an ARMv8 SVE high-performance computing (HPC) system. Key optimizations include leveraging OpenMP for multi-threaded acceleration of DeepMD-kit and utilizing the ARMv8 SVE instruction set to optimize doubleprecision matrix multiplication in PyTorch. These enhancements enable single ARMv8 SVE 64-core processors to achieve 1.3x the training performance of NVIDIA V100 GPU, and two ARMv8 SVE 64-core processors to achieve 1.05x the inference performance of NVIDIA V100 GPU. Leveraging this optimized framework, we achieve large-scale MD simulations across 4,096 computing nodes. Qi Du, Feng Wang 0050, Chengkun Wu, Han Wang 0006, Yongpeng Liu, Zhaoyin Zhou, Kenli Li 0001 |
CLUSTER | 2 |
| 2025 | Parallelization Strategies for DeepMD-Kit Using OpenMP: Enhancing Efficiency in Machine Learning-Based Molecular SimulationsabstractDeepMD-kit enables deep learning-based molecular dynamics (MD) simulations that require efficient parallelization to leverage modern HPC architectures. In this work, we optimize DeepMD-kit using advanced OpenMP strategies to improve scalability and computational efficiency on an ARMv8 processor-based server. Our optimizations include data parallelism for neural network inference, force calculation acceleration, NUMAaware memory management, and synchronization reductions, leading to up to 4.1× speedup and 82% higher memory band-width efficiency compared to the baseline implementation. Strong scaling analysis demonstrates superlinear speedup at mid-range core counts, with improved workload balancing and vectorized computations. However, challenges remain at ultra-large scales due to increasing synchronization overhead. Qi Du, Feng Wang 0050, Chengkun Wu |
IEEE Trans. Computers | 2 |
| 2024 | Exploring Natural Language Processing Model Acceleration in Molecular Dynamics Simulation Using High-Performance Computing and Machine LearningabstractIn molecular dynamics simulations, the integral methods used to calculate molecular trajectories, such as the Verlet integral method, involve significant computational costs. On the premise of ensuring the accuracy of simulation, how to effectively reduce the amount of computation is always a challenging problem. This paper applies Natural Language Processing models to simulate molecular motion trajectories in molecular dynamics, integrating the MPI programming model and machine learning on a high-performance computing platform to achieve MPI parallelization during model training and inference. The results indicate that different types of neural network architectures have a significant impact on inference performance and accuracy. When using the deep spatio-temporal networks model, the mean absolute error is 0.0032. At an atomic scale of 3M, the parallel efficiency with 1024 computational nodes exceeds 90%, demonstrating excellent parallel performance. Qi Du, Feng Wang 0050, Heng Wan, Chengkun Wu |
BIBM | 2 |
| 2023 | Back to Homogeneous Computing: A Tightly-Coupled Neuromorphic Processor With Neuromorphic ISAabstractIn recent years, neuromorphic processors are widely used in many scenarios, showing extreme energy efficiency over traditional architectures. However, almost all existing neuromorphic hardware are following the heterogeneous computing methodology without Instruction Set Architecture (ISA), leading to inflexibility in programming. In this paper, we first propose a RISC-V Neuromorphic Extension (RVNE) to enable fine-grained and flexible homogeneous programming for neuromorphic algorithms while utilizing SNN sparsity from different levels of granularity and computing flows. Based on RVNE, we next implement a neuromorphic micro-architecture that is tightly coupled to the CPU pipeline to accelerate neuromorphic computing. To demonstrate the proposed homogeneous neuromorphic architecture, we implement a prototype processor called NeuroRVcore based on RISC-V ISA and an open-source RISC-V core. The evaluation results show that RVNE achieves a 2.8 × −4.3 × reduction in code density compared with the general-purpose ISAs. Compared with the state-of-the-art neuromorphic processor, the proposed homogeneous computing reduces energy consumption by 3.4%−22.5% while enabling fine-grained and flexible homogeneous programming. Lei Wang 0011, Yao Wang 0002, Junbo Tie, Feng Wang 0050, LingHui Peng, Xun Xiao, Gan Zhou, Xuhu Yu, Xia Zhao 0004, Yuhua Tang, Weixia Xu 0001 |
IEEE Trans. Parallel Distributed Syst. | 6 |
| 2015 | Implementation of an Accurate and Efficient Compensated DGEMM for 64-bit ARMv8 Multi-Core ProcessorsabstractThis paper presents an implementation of an accurate and efficient compensated Double-precision General Matrix Multiplication (DGEMM) based on OpenBLAS for 64-bit ARMv8 multi-core processors. Due to cancellation phenomena in floating point arithmetic, the results of DGEMM may not be as accurate as expected. In order to increase the accuracy of DGEMM, we compensate the error introduced by its dot product kernel (GEBP) by applying an error-free transformation to rewrite the kernel in assembly language. We optimize the computations in the inner kernel through exploiting loop unrolling, instruction scheduling and software-implemented register rotation to exploit instruction level parallelism (ILP). We also conduct a priori error analysis of the derived CompDGEMM. Our compensated DGEMM is as accurate as the existing quadruple precision GEMM using MBLAS, but is up to 6.4x faster. Our parallel implementation achieves good performance and scalability under varying thread counts across a range of matrix sizes evaluated. Hao Jiang 0001, Feng Wang 0050, Kuan Li, Canqun Yang, Kejia Zhao, Chun Huang 0006 |
ICPADS | 2 |
| 2015 | Design and Implementation of a Highly Efficient DGEMM for 64-Bit ARMv8 Multi-core ProcessorsabstractThis paper presents the design and implementation of a highly efficient Double-precision General Matrix Multiplication (DGEMM) based on Open BLAS for 64-bit ARMv8 eight-core processors. We adopt a theory-guided approach by first developing a performance model for this architecture and then using it to guide our exploration. The key enabler for a highly efficient DGEMM is a highly-optimized inner kernel GEBP developed in assembly language. We have obtained GEBP by (1) maximizing its compute-to-memory access ratios across all levels of the memory hierarchy in the ARMv8 architecture with its performance-critical block sizes being determined analytically, and (2) optimizing its computations through exploiting loop unrolling, instruction scheduling and software-implemented register rotation and taking advantage of A64 instructions to support efficient FMA operations, data transfers and prefetching. We have compared our DGEMM implemented in Open BLAS with another implemented in ATLAS (also in terms of a highly-optimized GEBP in assembly). Our implementation outperforms the one in ALTAS by improving the peak performance (efficiency) of DGEMM from 3.88 Gflops (80.9%) to 4.19 Gflops (87.2%) on one core and from 30.4 Gflops (79.2%) to 32.7 Gflops (85.3%) on eight cores. These results translate into substantial performance (efficiency) improvements by 7.79% on one core and 7.70% on eight cores. In addition, the efficiency of our implementation on one core is very close to the theoretical upper bound 91.5% obtained from micro-benchmarking. Our parallel implementation achieves good performance and scalability under varying thread counts across a range of matrix sizes evaluated. Feng Wang 0050, Hao Jiang 0001, Ke Zuo, Xing Su 0004, Jingling Xue, Canqun Yang |
ICPP | 1 |
| 2014 | OpenMC: Towards Simplifying Programming for TianHe Supercomputers
Xiangke Liao, Canqun Yang, Tao Tang 0001, Huizhan Yi, Feng Wang 0050, Jingling Xue |
J. Comput. Sci. Technol. | 5 |
| 2012 | Parallelizing SOR for GPGPUs using alternate loop tiling
Peng Di, Hui Wu 0001, Jingling Xue, Feng Wang 0050, Canqun Yang |
Parallel Comput. | 4 |
| 2011 | Optimizing Linpack Benchmark on GPU-Accelerated Petascale Supercomputer
Feng Wang 0050, Canqun Yang, Yunfei Du 0001, Juan Chen 0001, Huizhan Yi, Weixia Xu 0001 |
J. Comput. Sci. Technol. | 1 |
| 2010 | Adaptive Optimization for Petascale Heterogeneous CPU/GPU ComputingabstractIn this paper, we describe our experiment developing an implementation of the Linpack benchmark for TianHe-1, a petascale CPU/GPU supercomputer system, the largest GPU-accelerated system ever attempted before. An adaptive optimization framework is presented to balance the workload distribution across the GPUs and CPUs with the negligible runtime overhead, resulting in the better performance than the static or the training partitioning methods. The CPU-GPU communication overhead is effectively hidden by a software pipelining technique, which is particularly useful for large memory-bound applications. Combined with other traditional optimizations, the Linpack we optimized using the adaptive optimization framework achieved 196.7 GFLOPS on a single compute element of TianHe-1. This result is 70.1% of the peak compute capability and 3.3 times faster than the result using the vendor's library. On the full configuration of TianHe-1 our optimizations resulted in a Linpack performance of 0.563PFLOPS, which made TianHe-1 the 5th fastest supercomputer on the Top500 list released in November 2009. Canqun Yang, Feng Wang 0050, Yunfei Du 0001, Juan Chen 0001, Jie Liu 0002, Huizhan Yi, Kai Lu 0001 |
CLUSTER | 2 |
| 2009 | Solving 2D Nonlinear Unsteady Convection-Diffusion Equations on Heterogenous Platforms with Multiple GPUsabstractSolving complex convection-diffusion equations is very important to many practical mathematical and physical problems. After the finite difference discretization, most of the time for equations solution is spent on sparse linear equation solvers. In this paper, our goal is to solve 2D Nonlinear Unsteady Convection-Diffusion Equations by accelerating an iterative algorithm named Jacobi-preconditioned QMRCGSTAB on a heterogenous platform, which is composed of a multi-core processor and multiple GPUs. Firstly, a basic implementation and evaluation for adapting the problem to this kind of platform is given. Then, we propose two optimization methods to improve the performance: kernel merging method and matrix boundary data processing. Our experimental evaluation on an AMD Opteron(tm) quad-core processor 2380 linked to an NVIDIA Tesla S1070 platform with four GPUs delivers the peak performance of 33 GFLOPS (double precision), which is a speedup of close to a factor 32 compared to the same problem running on 4 cores of the same CPU. Canqun Yang, Zhen Ge, Juan Chen 0001, Feng Wang 0050, Yunfei Du 0001 |
ICPADS | 4 |