EDBT 2026 Demo / reviewers in the wild / expert
Yun Wu 0003
dblp:32/5387-3
· DBLP profile ↗
10ranked-venue papers
6as first author
5since 2021 · last 2025
0000-0001-9332-2858ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 3 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 3 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Special Sessions - Emerging Scope and Design Challenges for Approximate Computing: Optimizing Accuracy-PPA trade-offs and BeyondabstractThe rapid growth of AI workloads is driving interest in Approximate Computing (AxC) as a means to enable low-cost, energy-efficient inference in resource-constrained systems. By introducing controlled inaccuracies, AxC can deliver substantial gains in power, performance, and area (PPA) while leveraging the inherent error tolerance of many AI models. Achieving this potential requires adapting existing frameworks to support the design and optimization of neural networks with approximate operators. Modern AxC research extends beyond accuracy-PPA trade-offs to address reliability and security, reducing redundancy overheads and exploring the distinctive side-channel implications of approximation. Application-aware approaches, such as those for spiking neural networks, show that tailoring approximation to workload-specific error behavior can surpass generic strategies. This article examines AI-guided design methods and the interplay between efficiency, reliability, and security, highlighting how these interconnected facets can advance embedded and high-performance computing. Siva Satyendra Sahoo, Bastien Deveautour, Marcello Traiola, Chongyan Gu, Yun Wu 0003, Aditya Japa, Salim Ullah, Akash Kumar 0001 |
CASES | 5 |
| 2025 | Efficient Co-Approximate Parallel Compressive Depth Reconstruction on FPGAabstractEfficient depth image reconstruction from sparse samples is crucial for machine perception applications, such as robotics, vehicle assistance and autonomy. It demands fast processing speed with low power consumption for sensing quality and safety, as well as cost reduction for FPGA and solid state implementations, within constrained resource budgets on edge devices. A new co-approximate framework of parallel approximate compressive depth reconstruction engine on FPGA is proposed using ℓ1solvers, proximal gradient decent (PGD), with instrumented frequency and voltage scaling during the iterative optimization process. By evaluating various number of parallel approximate processing units for the depth image reconstruction engine, up to 51% further power saving is achieved, and 421× speed up of parallel processing compared to the baseline, henceforth the efficiency is elevated over 43×. Yun Wu 0003, John McAllister |
ICASSP | 1 |
| 2025 | Invited Paper: Rowhammer Mitigation by Approximate Computing: A Compressed Sensing Case StudyabstractWhile Approximate Computing (AC) trades the precision for energy efficiency with tolerable errors, its security of approximate data in digital storage is not well explored for edge devices. As one of the most effective hardware security attack methods, Rowhammer attack has shown significant threats to the digital data on dynamic random access memory (DRAM) with the vulnerability of high-frequency memory row access. This work performs the first preliminary evaluation of Rowhammer attack on real-world compressed sensing applications with approximate data. By investigating Rowhammer attack on the approximate data from compact LiDar sensor, the security impact of various precisions is presented through the fidelity of reconstructed depth image. The experiments reveal considerable mitigation of Rowhammer attack by adopting AC based sensor signal processing, where up to 2× higher PSNR of output depth image is achieved comparing to those with accurate data and computations. Yuhang Hao, Yun Wu 0003, Minmin Jiang, Máire O'Neill, Chongyan Gu |
ICCAD | 2 |
| 2025 | AxRA: Approximate Rowhammer Attack for Modern DRAM SystemsabstractApproximate computing achieves high performance or less power consumption in various fault-tolerant applications, e.g., image processing, artificial intelligence (AI), etc. However, the introduction of approximate computing brings new security vulnerabilities, which threaten the entire computing system. In this paper, a novel Rowhammer attack is proposed, which utilises the approximate data stored in DRAM memories to achieve higher attack effectiveness. Compared to Rowhammer attack to DRAM memory without approximate data, the proposed method achieves more bit-flips resulting in significant data corruption. The proposed attack is implemented and evaluated on DRAM chips with a real user case, object detection using neural network. The accuracy of detection on the baseline image is employed to verify the impact of proposed attack approach. The results show that the proposed Rowhammer attack with approximate data introduces extra 33% bit-flips on victim rows than a conventional Rowhammer attack without approximate data. It also introduces up to ∼75% accuracy reduction of MNIST neural network proportionally to the increment of attack activation number. Yuhang Hao, Yun Wu 0003, Ziying Ni, Jack Miskelly, Máire O'Neill, Chongyan Gu |
ISCAS | 2 |
| 2021 | Configurable Quasi-Optimal Sphere Decoding for Scalable MIMO CommunicationsabstractSphere Decoding (SD) enables real-time quasi-optimal symbol detection for Multiple-Input Multiple-Output (MIMO) communication systems via custom circuit accelerators. Configurable SDs allow accelerator cost to be balanced with detection accuracy for the most constrained MIMO environments, such as power-constrained Internet-of-Things (IoT) scenarios. However this high detection accuracy comes at high accelerator cost. This paper proposes a novel configurable SD which addresses this issue. A Robust Bounded Spanning with Fast Enumeration (R-BSFE) approach employs novel strategies for channel matrix pre-processing and symbol enumeration to maintain quasi-ML accuracy whilst reducing complexity by up to 74%. This enables accelerators for 802.11n on Xilinx FPGA with significantly lower cost and higher throughput. To the best of the authors' knowledge, the accelerators produced are the highest performance, lowest cost quasi-ML SD accelerators on record. Yun Wu 0003, John McAllister |
IEEE Trans. Circuits Syst. I Regul. Pap. | 1 |
| 2020 | Programmable Dataflow Accelerators: A 5G OFDM Modulation/Demodulation Case StudyabstractVia OFDM technology, FFT and Inverse FFT (IFFT) operators enable the latest 5G radio standards. In these latests standards, the behaviour of FFT and IFFT needs to be flexible, supporting sub-carrier spacings from 15kHz to 480kHz and point sizes of up to 4096 point. An FFT or IFFT accelerator for 5G can take any configuration inside this spectrum both at design time, or potentially at run-time, under the control of a system control plane. This necessitates accelerators which combine high levels of flexibility and performance. This paper describes an FFT accelerator for such a context. Specifically, a novel data-driven programmable softcore processor is presented which enables run-time variable workloads and respond to data as provided by control processors in an MPSoC operating architecture. It is the first such accelerator to enable real-time IFFT/FFT for 5G, providing up to 3.89 times greater data rate than comparable accelerators. Yun Wu 0003, Peng Wang 0091, John McAllister |
ICASSP | 1 |
| 2019 | On Modified Squared Givens Rotations for Sphere Decoder PreprocessingabstractSphere Decoding for Multiple-Input Multiple-Output (MIMO) wireless systems is a complex operation, usually demanding custom accelerators in order to support real-time performance. The cost of these accelerators is disproportionately influenced by channel matrix preprocessing, which represents a relatively small fraction of the overall computational cost of detecting an OFDM MIMO frame in standards such as 802.11n, but consumes a very large amount of hardware resource. Modified Squared Givens' Rotations has been proposed to resolve this issue and shown to dramatically reduce accelerator cost. However, there is no analysis on the record of the complexity of this algorithm, nor its detection performance. This paper shows that, despite offering modest reductions in operational complexity, MFSD-SQRD enables dramatic cost reductions by explicitly addressing the overhead of matrix permutation steps. Further, it shows that for most SNR values of practical interest, the performance of MFSD-SQRD is not appreciably diminished relative to the standard SQRD approach to preprocessing. To the best of the authors' knowledge, the proposed modified SQRD preprocessing approach is the highest performance sub-optimal preprocessing approach on record. Yun Wu 0003, John McAllister |
ICASSP | 1 |
| 2018 | Architectural Synthesis of Multi-SIMD Dataflow Accelerators for FPGAabstractField Programmable Gate Array (FPGA) boast abundant resources with which to realise high-performance accelerators for computationally demanding operations. Highly efficient accelerators may be automatically derived from Signal Flow Graph (SFG) models by using architectural synthesis techniques, but in practical design scenarios, these currently operate under two important limitations - they cannot efficiently harness the programmable datapath components which make up an increasing proportion of the computational capacity of modern FPGA and they are unable to automatically derive accelerators to meet a prescribed throughput or latency requirement. This paper addresses these limitations. SFG synthesis is enabled which derives software-programmable multicore single-instruction, multiple-data (SIMD) accelerators which, via combined offline characterisation of multicore performance and compile-time program analysis, meet prescribed throughput requirements. The effectiveness of these techniques is demonstrated on tree-search and linear algebraic accelerators for 802.11n WiFi transceivers, an application for which satisfying real-time performance requirements has, to this point, proven challenging for even manually-derived architectures. Yun Wu 0003, John McAllister |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2014 | Power modelling and capping for heterogeneous ARM/FPGA SoCsabstractLow-power processors and accelerators that were originally designed for the embedded systems market are emerging as building blocks for servers. Power capping has been actively explored as a technique to reduce the energy footprint of high-performance processors. The opportunities and limitations of power capping on the new low-power processor and accelerator ecosystem are less understood. This paper presents an efficient power capping and management infrastructure for heterogeneous SoCs based on hybrid ARM/FPGA designs. The infrastructure coordinates dynamic voltage and frequency scaling with task allocation on a customised Linux system for the Xilinx Zynq SoC. We present a compiler-assisted power model to guide voltage and frequency scaling, in conjunction with workload allocation between the ARM cores and the FPGA, under given power caps. The model achieves less than 5% estimation bias to mean power consumption. In an FFT case study, the proposed power capping schemes achieve on average 97.5% of the performance of the optimal execution and match the optimal execution in 87.5% of the cases, while always meeting power constraints. Yun Wu 0003, José L. Núñez-Yáñez, Roger F. Woods, Dimitrios S. Nikolopoulos |
FPT | 1 |
| 2013 | Soft-core stream processing on FPGA: An FFT case studyabstractThe increasing design complexity associated with modern Field Programmable Gate Array (FPGA) has prompted the emergence of 'soft'-programmable processors which attempt to replace at least part of the custom circuit design problem with a problem of programming parallel processors. Despite substantial advances in this technology, its performance and resource efficiency for computationally complex operations remains in doubt. In this paper we present the first recorded implementation of a softcore Fast-Fourier Transform (FFT) on Xilinx Virtex FPGA technology. By employing a streaming processing architecture, we show how it is possible to achieve architectures which offer 1.1 GSamples/s throughput and up to 19 times speed-up against the Xilinx Radix-2 FFT dedicated circuit with comparable cost. Peng Wang 0091, John McAllister, Yun Wu 0003 |
ICASSP | 3 |