EDBT 2026 Demo / reviewers in the wild / expert
Kristof Denolf
dblp:49/799
· DBLP profile ↗
23ranked-venue papers
3as first author
12since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 16 · 2 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-authorArtificial intelligence and machine learning · 1 · 1 since 2021Software engineering, systems software and programming languages · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | STEEL: Sparsity-Aware Fused Attention for Energy-Efficient Long-Sequence Inference on AMD's XDNA™ NPUabstractTHE growing integration of Transformer-based artificial intelligence (AI) agents into core operating system functions is a key driver in modern laptop systems-on-chip (SoCs) design. While enabling powerful capabilities, their inference incurs significant compute and data-movement overhead, making them highly energy-intensive. This energy cost is a fundamental bottleneck for embedded mobile platforms with tight power and thermal constraints [2] . The Attention prefill stage is a major contributor to inference latency and energy at long sequence lengths. Consequently, significant effort has focused on optimizing attention across commercial [3] and academic platforms [4] , spanning algorithmic advances such as FlashAttention [3] and hardware enhancements including specialized non-linear units. Neural processing units (NPUs) achieve high energy efficiency through spatial dataflow architectures and explicit data-movement programming models, which expose fine-grained control over computation and memory transfers. While extensive prior work has focused on optimizing attention for graphics processing units (GPUs), comparatively few efforts have targeted attention for NPUs. Victor J. B. Jung, Gagandeep Singh 0002, Joseph Melber, Kristof Denolf, Francesco Conti 0001, Luca Benini |
FCCM | 4 |
| 2026 | Striking the Balance: GEMM Performance Optimization Across Generations of Ryzen™ AI NPUsabstractThe high computational and memory demands of modern deep learning (DL) workloads have led to the development of specialized hardware devices from cloud to edge, such as AMD's Ryzen™ AI XDNA™ NPUs. Optimizing general matrix multiplication (GEMM) algorithms for these architectures is critical for improving DL workload performance. To this end, this paper presents a common systematic methodology to optimize GEMM workloads across the two current NPU generations, namely XDNA and XDNA2. Our implementations exploit the unique architectural features of AMD's NPUs and address key performance bottlenecks at the system level. End-to-end performance evaluation across various GEMM sizes demonstrates state-of-the-art throughput of up to 6.76 TOPS (XDNA) and 38.05 TOPS (XDNA2) for 8-bit integer (int8) precision. Similarly, for brain floating-point (bf16) precision, our GEMM implementations attain up to 3.14 TOPS (XDNA) and 14.71 TOPS (XDNA2). This work provides significant insights into key performance aspects of optimizing GEMM workloads on Ryzen AI NPUs. Endri Taka, André Rösti, Joseph Melber, Pranathi Vasireddy, Kristof Denolf, Diana Marculescu |
FPGA | 5 |
| 2026 | From Loop Nests to Silicon: Mapping AI Workloads onto AMD NPUs with MLIR-AIRabstractGeneral-purpose compilers abstract away parallelism, locality, and synchronization, limiting their effectiveness on modern spatial architectures. As modern computing architectures increasingly rely on fine-grained control over data movement, execution order, and compute placement for performance, compiler infrastructure must provide explicit mechanisms for orchestrating compute and data to fully exploit such architectures. We introduce MLIR-AIR, a novel, open source compiler stack built on MLIR that bridges the semantic gap between high-level workloads and fine-grained spatial architectures such as AMD’s NPUs. MLIR-AIR defines the AIR dialect, which provides structured representations for asynchronous and hierarchical operations across compute and memory resources. AIR primitives allow the compiler to orchestrate spatial scheduling, distribute computation across hardware regions, and overlap communication with computation without relying on ad hoc runtime coordination or manual scheduling. We demonstrate MLIR-AIR’s capabilities through two case studies: matrix multiplication and the multi-head attention block from the LLaMA 2 model. For matrix multiplication, MLIR-AIR achieves up to 78.7% compute efficiency and generates implementations with performance almost identical to state-of-the-art, hand-optimized matrix multiplication written using the lower-level, close-to-metal MLIR-AIE framework. For multi-head attention, we demonstrate that the AIR interface supports fused implementations using approximately 150 lines of code, enabling tractable expression of complex workloads with efficient mapping to spatial hardware. MLIR-AIR transforms high-level structured control flow into spatial programs that efficiently utilize the compute fabric and memory hierarchy of an NPU, leveraging asynchronous execution, tiling, and communication overlap through compiler-managed scheduling. Erwei Wang, Samuel Bayliss, Andra Bisca, Zachary Blair, Sangeeta Chowdhary, Kristof Denolf, Jeff Fifield, Brandon Freiberger, Erika Hunhoff, Phil James-Roxby, Jack Lo, Joseph Melber, Stephen Neuendorffer, Eddie Richter, André Rösti, Javier Setoain, Gagandeep Singh 0002, Endri Taka, Pranathi Vasireddy, Zhewen Yu, Niansong Zhang, Jinming Zhuang |
ACM Trans. Reconfigurable Technol. Syst. | 6 |
| 2025 | Efficiency, Expressivity, and Extensibility in a Close-to-Metal NPU Programming InterfaceabstractAccelerators such as neural processing units (NPUs) deliver an enticing balance of performance and efficiency compared to general purpose compute architectures. However, effectively leveraging accelerator capabilities is not always simple: low-level programming toolkits may require substantial developer effort while high-level programming toolkits may abstract critical optimization features. This work aims to increase efficiency of designers using IRON, a toolkit for close-to-metal NPU performance engineers. We provide an updated programmer interface to IRON containing new and refined programming constructs. The new interface includes extensible features for placement and data transformation. These contributions are evaluated in terms of 1) efficiency, with analysis showing ~ 26% average reduction in lines of code and decreases in Halstead metrics for a variety of designs; 2) expressivity, demonstrating the new interface supports the wide range of features and patterns already supported by IRON; and 3) extensibility, illustrating the new tooling for placement and tiling can be extended to accommodate common use-cases. Erika Hunhoff, Joseph Melber, Kristof Denolf, Andra Bisca, Samuel Bayliss, Stephen Neuendorffer, Jeff Fifield, Jack Lo, Pranathi Vasireddy, Phil James-Roxby, Eric Keller |
FCCM | 3 |
| 2024 | Tailor: Altering Skip Connections for Resource-Efficient InferenceabstractDeep neural networks use skip connections to improve training convergence. However, these skip connections are costly in hardware, requiring extra buffers and increasing on- and off-chip memory utilization and bandwidth requirements. In this article, we show that skip connections can be optimized for hardware when tackled with a hardware-software codesign approach. We argue that while a network’s skip connections are needed for the network to learn, they can later be removed or shortened to provide a more hardware-efficient implementation with minimal to no accuracy loss. We introduce Tailor , a codesign tool whose hardware-aware training algorithm gradually removes or shortens a fully trained network’s skip connections to lower the hardware cost. Tailor improves resource utilization by up to 34% for block random access memories (BRAMs), 13% for flip-flops (FFs), and 16% for look-up tables (LUTs) for on-chip, dataflow-style architectures. Tailor increases performance by 30% and reduces memory bandwidth by 45% for a two-dimensional processing element array architecture. Olivia Weng, Gabriel Marcano, Vladimir Loncar, Alireza Khodamoradi, G. Abarajithan, Nojan Sheybani, Andres Meza 0001, Farinaz Koushanfar, Kristof Denolf, Javier M. Duarte, Ryan Kastner |
ACM Trans. Reconfigurable Technol. Syst. | 9 |
| 2024 | CHARM 2.0: Composing Heterogeneous Accelerators for Deep Learning on Versal ACAP ArchitectureabstractDense matrix multiply (MM) serves as one of the most heavily used kernels in deep learning applications. To cope with the high computation demands of these applications, heterogeneous architectures featuring both FPGA and dedicated ASIC accelerators have emerged as promising platforms. For example, the AMD/Xilinx Versal ACAP architecture combines general-purpose CPU cores and programmable logic with AI Engine processors optimized for AI/ML. An array of 400 AI Engine processors executing at 1 GHz can provide up to 6.4 TFLOPS performance for 32-bit floating-point (FP32) data. However, machine learning models often contain both large and small MM operations. While large MM operations can be parallelized efficiently across many cores, small MM operations typically cannot. We observe that executing some small MM layers from the BERT natural language processing model on a large, monolithic MM accelerator in Versal ACAP achieved less than 5% of the theoretical peak performance. Therefore, one key question arises: How can we design accelerators to fully use the abundant computation resources under limited communication bandwidth for end-to-end applications with multiple MM layers of diverse sizes? We identify the biggest system throughput bottleneck resulting from the mismatch between the massive computation resources of one monolithic accelerator and the various MM layers of small sizes in the application. To resolve this problem, we propose the CHARM framework to compose multiple diverse MM accelerator architectures working concurrently on different layers within one application. CHARM includes analytical models that guide design space exploration to determine accelerator partitions and layer scheduling. To facilitate system designs, CHARM automatically generates code, enabling thorough onboard design verification. We deploy the CHARM framework on four different deep learning applications in FP32, INT16, and INT8 data types, including BERT, ViT, NCF, and MLP, on the AMD/Xilinx Versal ACAP VCK190 evaluation board. Our experiments show that we achieve 1.46 TFLOPS, 1.61 TFLOPS, 1.74 TFLOPS, and 2.94 TFLOPS inference throughput for BERT, ViT, NCF, and MLP in FP32 data type, respectively, which obtain 5.29 \(\times\) , 32.51 \(\times\) , 1.00 \(\times\) , and 1.00 \(\times\) throughput gains compared to one monolithic accelerator. CHARM achieves the maximum throughput of 1.91 TOPS, 1.18 TOPS, 4.06 TOPS, and 5.81 TOPS in the INT16 data type for the four applications. The maximum throughput achieved by CHARM in the INT8 data type is 3.65 TOPS, 1.28 TOPS, 10.19 TOPS, and 21.58 TOPS, respectively. We have open-sourced our tools, including detailed step-by-step guides to reproduce all the results presented in this article and to enable other users to learn and leverage CHARM framework and tools in their end-to-end systems: https://github.com/arc-research-lab/CHARM . Jinming Zhuang, Jason Lau, Hanchen Ye, Zhuoping Yang, Shixin Ji, Jack Lo, Kristof Denolf, Stephen Neuendorffer, Alex K. Jones, Jingtong Hu, Yiyu Shi 0001, Deming Chen, Jason Cong, Peipei Zhou 0001 |
ACM Trans. Reconfigurable Technol. Syst. | 7 |
| 2023 | Adapting Skip Connections for Resource-Efficient FPGA InferenceabstractDeep neural networks employ skip connections – identity functions that combine the outputs of different layers-to improve training convergence; however, these skip connections are costly to implement in hardware. In particular, for inference accelerators on resource-limited platforms, they require extra buffers, increasing not only on- and off-chip memory utilization but also memory bandwidth requirements. Thus, a network that has skip connections costs more to deploy in hardware than one that has none. We argue that, for certain classification tasks, a network's skip connections are needed for the network to learn but not necessary for inference after convergence. We thus explore removing skip connections from a fully-trained network to mitigate their hardware cost. From this investigation, we introduce a fine-tuning/retraining method that adapts a network's skip connections – by either removing or shortening them-to make them fit better in hardware with minimal to no loss in accuracy. With these changes, we decrease resource utilization by up to 34% for BRAMs, 7% for FFs, and 12% LUTs when implemented on an FPGA. Olivia Weng, Gabriel Marcano, Vladimir Loncar, Alireza Khodamoradi, Nojan Sheybani, Farinaz Koushanfar, Kristof Denolf, Javier M. Duarte, Ryan Kastner |
FPGA | 7 |
| 2023 | CHARM: Composing Heterogeneous AcceleRators for Matrix Multiply on Versal ACAP ArchitectureabstractDense matrix multiply (MM) serves as one of the most heavily used kernels in deep learning applications. To cope with the high computation demands of these applications, heterogeneous architectures featuring both FPGA and dedicated ASIC accelerators have emerged as promising platforms. For example, the AMD/Xilinx Versal ACAP architecture combines general-purpose CPU cores and programmable logic (PL) with AI Engine processors (AIE) optimized for AI/ML. An array of 400 AI Engine processors executing at 1 GHz can theoretically provide up to 6.4 TFLOPs performance for 32-bit floating-point (fp32) data. However, machine learning models often contain both large and small MM operations. While large MM operations can be parallelized efficiently across many cores, small MM operations typically cannot. In our investigation, we observe that executing some small MM layers from the BERT natural language processing model on a large, monolithic MM accelerator in Versal ACAP achieved less than 5% of the theoretical peak performance. Therefore, one key question arises: How can we design accelerators to fully use the abundant computation resources under limited communication bandwidth for end-to-end applications with multiple MM layers of diverse sizes? Jinming Zhuang, Jason Lau, Hanchen Ye, Zhuoping Yang, Yubo Du, Jack Lo, Kristof Denolf, Stephen Neuendorffer, Alex K. Jones, Jingtong Hu, Deming Chen, Jason Cong, Peipei Zhou 0001 |
FPGA | 7 |
| 2023 | SPARTA: Spatial Acceleration for Efficient and Scalable Horizontal Diffusion Weather Stencil ComputationabstractFast and accurate climate simulations and weather predictions are critical for understanding and preparing for the impact of climate change. Real-world climate and weather simulations involve the use of complex compound stencil kernels, which are composed of a combination of different stencils. Horizontal diffusion is one such important compound stencil found in many climate and weather prediction models. Its computation involves a large amount of data access and manipulation that leads to two main issues on current computing systems. First, such compound stencils have high memory bandwidth demands as they require large amounts of data access. Second, compound stencils have complex data access patterns and poor data locality, as the memory access pattern is typically irregular with low arithmetic intensity. As a result, state-of-the-art CPU and GPU implementations suffer from limited performance and high energy consumption. Recent works propose using FPGAs as an alternative to traditional CPU and GPU-based systems to accelerate weather stencil kernels. However, we observe that stencil computation cannot leverage the bit-level flexibility available on an FPGA because of its complex memory access patterns, leading to high hardware resource utilization and low peak performance. Gagandeep Singh 0002, Alireza Khodamoradi, Kristof Denolf, Jack Lo, Juan Gómez-Luna, Joseph Melber, Andra Bisca, Henk Corporaal, Onur Mutlu |
ICS | 3 |
| 2021 | S2N2: A FPGA Accelerator for Streaming Spiking Neural NetworksabstractSpiking Neural Networks (SNNs) are the next generation of Artificial Neural Networks (ANNs) that utilize an event-based representation to perform more efficient computation. Most SNN implementations have a systolic array-based architecture and, by assuming high sparsity in spikes, significantly reduce computing in their designs. This work shows this assumption does not hold for applications with signals of large temporal dimension. We develop a streaming SNN (S2N2) architecture that can support fixed-per-layer axonal and synaptic delays for its network. Our architecture is built upon FINN and thus efficiently utilizes FPGA resources. We show how radio frequency processing matches our S2N2 computational model. By not performing tick-batching, a stream of RF samples can efficiently be processed by S2N2, improving the memory utilization by more than three orders of magnitude. Alireza Khodamoradi, Kristof Denolf, Ryan Kastner |
FPGA | 2 |
| 2021 | ASLR: An Adaptive Scheduler for Learning RateabstractTraining a neural network is a complicated and time-consuming task that involves adjusting and testing different combinations of hyperparameters. One of the essential hyperparameters is the learning rate, which balances the magnitude of changes at each training step. We introduce an Adaptive Scheduler for Learning Rate (ASLR) that significantly lowers the tuning effort since it only has a single hyperparameter. ASLR produces competitive results compared to the state-of-the-art for both hand-optimized learning rate schedulers and line search methods while requiring significantly less tuning effort. Our algorithm's computational cost is trivial and can be used to train various network topologies included quantized networks. Alireza Khodamoradi, Kristof Denolf, Kees A. Vissers, Ryan Kastner |
IJCNN | 2 |
| 2021 | Benchmarking vision kernels and neural network inference accelerators on embedded platforms
Murad Qasaimeh, Kristof Denolf, Alireza Khodamoradi, Michaela Blott, Jack Lo, Lisa Halder, Kees A. Vissers, Joseph Zambreno, Phillip H. Jones |
J. Syst. Archit. | 2 |
| 2019 | Analyzing the Energy-Efficiency of Vision Kernels on Embedded CPU, GPU and FPGA PlatformsabstractThis paper presents a benchmark of the energy efficiency of a wide range of vision kernels on three commonly used hardware accelerators for embedded vision applications: ARM57 CPU, Jetson TX2 GPU and ZCU102 FPGA, using their vendor optimized vision libraries: OpenCV, VisionWorks and xfOpenCV. Our results show that the GPU achieves an energy/frame reduction ratio of 1.1-3.2× compared to CPU and FPGA for simple kernels. While for more complicated kernels, the FPGA outperforms the others with energy/frame reduction ratios of 1.2-22.3×. It is also observed that the FPGA performs increasingly better as a vision kernel's complexity grows. Murad Qasaimeh, Joseph Zambreno, Phillip H. Jones, Kristof Denolf, Jack Lo, Kees A. Vissers |
FCCM | 4 |
| 2009 | Using C-to-gates to program streaming image processing kernels efficiently on FPGAsabstractEffectively exploiting the variety of computational and storage resources available in common FPGA architectures for complex applications, such as the real-time implementation of vision algorithms, is often difficult in standard HDL design methodologies. Higher-level design tools can enable a design to more quickly explore a range of different architectures. In this paper we apply algorithmic C-to-FPGA synthesis technology in a structured design approach and demonstrate its added value on two relevant vision processing kernels: optical flow and debayering. The impact of the proposed approach on the design time, the FPGA resource consumption and the throughput is measured. Kristof Denolf, Stephen Neuendorffer, Kees A. Vissers |
FPL | 1 |
| 2009 | Rate-distortion-complexity performance analysis of the SVC decoderabstractThe scalable video coding extends the H.264/AVC video coding standard by providing temporal, spatial and quality scalability. A set of new tools, such as key picture and inter-layer prediction are introduced in SVC to improve either rate-distortion performance or error resilience. This paper evaluates the impact of these tools on the performance of an optimized SVC decoder, in terms of rate, distortion, and computational complexity. The results facilitate the decision making process of choosing suitable SVC configurations under different application scenarios. Tong Gan, Bart Masschelein, Carolina Blanch, Antoine Dejonghe 0001, Kristof Denolf |
ICME | 5 |
| 2008 | Cross-layer optimization for multi-user video streaming over IEEE 802.11E HCCA wireless networksabstractIn this paper, we propose a cross-layer optimization scheme for multi-user video streaming over the uplink of IEEE 802.11e HCCA wireless network. The objective is to minimize the energy consumption of all users, including both video encoder energy and wireless transmission energy, while delivering desired video quality for each user. In our proposed scheme, both cross-layer optimization of individual mobile terminal and inter-user resource allocation are carried out, so that the video encoder and wireless transmitter configurations can be jointly steered. Experimental results show that when compared to the transmission system without cross-layer optimization, our proposed scheme is able to reduce total energy consumption from 35% up to 80%, while satisfying the same video quality target. Tong Gan, Antoine Dejonghe 0001, Gregory Lenoir, Kristof Denolf, Gauthier Lafruit, Iole Moccagatta |
ICME | 4 |
| 2007 | Modelling Energy Consumption of an ASIC MPEG-4 Simple Profile EncoderabstractIn wireless video streaming, it is desirable to minimize the energy consumption of mobile devices while achieving target video quality. For this purpose, it is essential to model (i.e., estimate) the energy cost of the streaming system, including both video encoder and wireless transmission energy, under different system settings. In this paper, an energy consumption model is proposed for a MPEG-4 simple profile encoder implemented in ASIC. Our proposed model consists of three major steps: 1) compute the encoder energy consumption through power simulations; 2) empirically model the impact of configuration parameters through curve fitting; and 3) online updating of model coefficients. Experimental results show that our proposed model works well for different types of video, and the average estimation error is below 6%. Tong Gan, Kristof Denolf, Gauthier Lafruit, Iole Moccagatta, Antoine Dejonghe 0001, Gregory Lenoir |
ICME | 2 |
| 2005 | Memory Efficient Design of an MPEG-4 Video Encoder for FPGAsabstractThe improving resolutions of new video appliances continuously increase the throughput requirements of video codecs and complicate the challenges encountered during their cost-efficient design. We propose an FPGA implementation of a high-performance MPEG-4 video encoder. The fully dedicated video pipeline is realized using a systematic design approach and exploits the inherent functional parallelism of the compression algorithm. The effect of memory and algorithmic optimizations applied at the high-level are measured on the RTL description. The resulting MPEG-4 video encoder efficiently uses the FPGA blockRAMs, uses burst oriented accesses to external memory and supports real-time processing of 30 4CIF frames per second. Kristof Denolf, Adrian Chirila-Rus, Robert D. Turney, Paul R. Schumacher, Kees A. Vissers |
FPL | 1 |
| 2005 | A scalable, multi-stream MPEG-4 video decoder for conferencing and surveillance applicationsabstractIncreasing resolutions push the throughput requirements of video codecs and complicate the challenges encountered during their cost-efficient implementations. We propose an FPGA implementation of a high-performance MPEG-4 simple profile video decoder, capable of parsing multiple bitstreams from different encoder sources. Its video pipeline architecture exploits the inherent functional parallelism and enables multi-stream support at a limited FPGA resource cost compared to a single stream version. The design is scalable with a number of added compile-time parameters - including maximum frame size and number of input bitstreams - which can be set by the user to suit his application. Paul R. Schumacher, Kristof Denolf, Adrian Chirila-Rus, Robert D. Turney, Nick Fedele, Kees A. Vissers, Jan Bormans |
ICIP (2) | 2 |
| 2005 | Memory Centric Design of an MPEG-4 Video EncoderabstractThe cost-efficient implementation of video codecs requires a set of methodologies and decision taking at different levels in the design flow. We combine upfront algorithmic tuning with memory centric optimizations to transform the video application into a system consisting of functional blocks with localized data processing and a tailored memory hierarchy. This memory optimized functional description is the leverage for the cost-efficient mapping of the system on integrated multimedia platforms. It closely reflects the real implementation constraints and consequently allows for steering the architecture selection in a correct way. The proposed approach is demonstrated on a MPEG-4 video encoder and leads to its implementation as a pipelined system. Hardware development of the motion estimation validates that the high-level memory centric concepts are applicable and realizable at the lowest level. The motion estimation kernel supports up to 30 CIF f/s with minimized processing element requirements and data input rates. Kristof Denolf, Christophe De Vleeschouwer, Robert D. Turney, Gauthier Lafruit, Jan Bormans |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2004 | A Power Optimized Display Memory Organization for Handheld User TerminalabstractToday's handheld devices become more and more multimedia capable. One subsystem of a multimedia terminal that accounts for a considerable amount of the total power consumption is the display unit. The backlight is the major culprit there. As new display units without backlights emerge, the data transfers required to put data on the screen start using up an increasingly important part of the platform's power. We have examined a novel system view that allows for power savings by decreasing the required number of memory accesses to put a frame on the screen. A two-step optimization method for existing platforms is presented. Measurements on a multimedia application show that, on average, power savings of 72% can be obtained on the display related memory accesses. For the proposed optimizations methods to work, it is important that both hardware and software designers become aware of the impact their design-time decisions have on the final power consumption of a system. Lieven Hollevoet, Andy Dewilde, Kristof Denolf, Francky Catthoor, Filip Louagie |
DATE | 3 |
| 2002 | Algorithmic and architectural co-design of a motion-estimation engine for low-power video devicesabstractDue to the large amount of data transfers it involves, the motion estimation (ME) engine is one of the most power-consuming components of any predictive video codec. As a consequence, power-optimized video coding primarily relies on a carefully designed motion estimator. This paper first presents a block ME algorithm that meets high-quality inter-frame prediction and low computational complexity requirements. It relies on a set of rules common to all recent fast and adaptive ME algorithms, but is designed in order to allow for easy and prolific data reuse. The adjacent order of the candidate positions during the search increases the locality and maintains a near-regular data flow, which results in a decrease of the data transfers and a low control complexity. Together with the computational complexity reduction, it enables cost-efficient very large scale integration realizations. A pipelined parallel architecture is then proposed and discussed. It is generic in the sense that it is suited both to the full-pel and half-pel ME. It is efficient because it allows for close to 100% hardware utilization and a sharp decrease of the peak memory bandwidth. It is suited to low-power implementation, as it enables larger data reuse factors for the most probable stages of the adaptive algorithm, which reduces the average memory bandwidth and power consumption. Christophe De Vleeschouwer, Tord Nilsson, Kristof Denolf, Jan Bormans |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2000 | 3D computational graceful degradationabstractinfo:eu-repo/semantics/published Gauthier Lafruit, Lode Nachtergaele, Kristof Denolf, Jan Bormans |
ISCAS | 3 |