VLDB 2026 Research / reviewers in the wild / expert
Per Gunnar Kjeldsberg
dblp:04/5892
· DBLP profile ↗
24ranked-venue papers
5as first author
2since 2021 · last 2026
0000-0001-9107-116XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 22 · 5 first-author · 1 since 2021Software engineering, systems software and programming languages · 4 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
4 papers |
Energy-efficient computing · 41% Electronic design automation · 20% Memory systems · 16% |
Topics — the 7 heaviest of 7, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Energy-efficient computing
power management |
0.3 | 1 | 2017 | Towards fine-grained dynamic tuning of HPC applications on modern multi-core architectures · SC 2017 |
Electronic design automation
high-level synthesis |
0.1 | 2 | 2007 | Bit-Width Constrained Memory Hierarchy Optimization for Real-Time Video Systems · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2007 Detection of Partially Simultaneously Alive Signals in Storage Requirement Estimation for Data Intensive Applications · DAC 2001 |
High-performance computing › supercomputing
exascale systems |
0.1 | 1 | 2017 | Towards fine-grained dynamic tuning of HPC applications on modern multi-core architectures · SC 2017 |
Memory systems › memory hierarchy
memory hierarchy optimization |
0.1 | 1 | 2007 | Bit-Width Constrained Memory Hierarchy Optimization for Real-Time Video Systems · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2007 |
Embedded and real-time systems › real-time signal processing
real-time video processing |
0.1 | 1 | 2007 | Bit-Width Constrained Memory Hierarchy Optimization for Real-Time Video Systems · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2007 |
Memory systems
memory optimization |
0.0 | 1 | 2003 | Data dependency size estimation for use in memory optimization · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2003 |
Electronic design automation
system-level design |
0.0 | 1 | 2003 | Data dependency size estimation for use in memory optimization · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2003 |
Methods — techniques the papers use, named apart from their topics
integer nonlinear programming · 0.1execution ordering analysis · 0.0data dependency analysis · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Exploring the Energy Storage and Voltage Control Unit Design Space in Battery-Less IoTabstractBattery-less Internet of Things (IoT) devices are emerging because it will be environmentally and economically unsustainable to power billions of IoT devices with batteries. Battery-less devices harvest energy from their environment, but the energy supplied by many such sources varies considerably over time, and there is hence a need for buffering surplus energy. This is enabled by the Energy Storage and Voltage Control Unit (ESVU), and the state-of-the-art ESVUs are REACT and CapDYN. We observe that an ideal ESVU (i) wastes minimal energy (efficiency), (ii) enables the application System-on-Chip (SoC) quickly when energy becomes available (responsiveness), (iii) can support a large variety of SoCs (applicability), and (iv) requires few or cheap components and occupies minimal area and volume (overhead). Through detailed circuit-level simulations, we demonstrate that no existing ESVU ticks all the boxes. More specifically, we find that CapDYN and REACT can provide high efficiency and responsiveness, but they fall short in applicability and overhead, respectively, and we thus propose Coulombix to fill this gap. To understand how ESVU design affects performance at the system level, we conduct a case study in which we implement hardware prototypes of REACT and Coulombix and measure the throughput and response time of a diverse set of IoT benchmarks with a solar energy harvester across three seasons. The case study supports the insights of the simulation-based study, while also exposing interesting second-order effects, highlighting that end-to-end analysis is critical when evaluating ESVUs. Lukas Liedtke, Espen Holsen, Per Gunnar Kjeldsberg, Frank Alexander Kraemer, Magnus Jahre |
ISPASS | 3 |
| 2026 | EStacker: Explaining Battery-Less IoT System Performance with Energy StacksabstractThe number of Internet of Things (IoT) devices is increasing exponentially, and it is environmentally and economically unsustainable to power all these devices with batteries. The key alternative is energy harvesting, but battery-less IoT systems require extensive evaluation to demonstrate that they are sufficiently performant across the full range of expected operating conditions. IoT developers thus need an evaluation platform that (i) ensures that each evaluated application and configuration is exposed to exactly the same energy environment and events, and (ii) provides a detailed account of what the application spends the harvested energy on. We therefore developed the EStacker evaluation platform which (i) enables fair and repeatable evaluation, and (ii) generates energy stacks. Energy stacks break down the total energy consumption of an application across hardware components and application activities, thereby explaining what the application specifically uses energy on. We augment EStacker with the ST-SP optimization which, in our experiments, reduces evaluation time by 6.3× on average while retaining the temporal behavior of the battery-less IoT system (average throughput error of 7.7%) by proportionally scaling time and power. We demonstrate the utility of EStacker through two case studies. In the first case study, we use energy stack profiles to identify a performance problem that, once addressed, improves performance by 3.3×. The second case study focuses on ST-SP, and we use it to explore the design space required to dimension the harvester and energy storage sizes of a smart parking application in roughly one week (7.7 days). Without ST-SP, sweeping this design space would have taken well over one month (41.7 days). Lukas Liedtke, Per Gunnar Kjeldsberg, Frank Alexander Kraemer, Magnus Jahre |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2018 | Runtime Precomputation of Data-Dependent Parameters in Embedded SystemsabstractIn many modern embedded systems, the available resources (e.g., CPU clock cycles, memory, and energy) are consumed nonuniformly while the system is under exploitation. Typically, the resource requirements in the system change with different input data that the system process. These data trigger different parts of the embedded software, resulting in different operations executed that require different hardware platform resources to be used. A significant research effort has been dedicated to develop mechanisms for runtime resource management (e.g., branch prediction for pipelined processors, prefetching of data from main memory to cache, and scenario-based design methodologies). All these techniques rely on the availability of information at runtime about upcoming changes in resource requirements. In this article, we propose a method for detecting upcoming resource changes based on preliminary calculation of software variables that have the most dynamic impact on resource requirements in the system. We apply the method on a modified real-life biomedical algorithm with real input data and estimate a 40% energy reduction as compared to static DVFS scheduling. Comparing to dynamic DVFS scheduling, an 18% energy reduction is demonstrated. Elena Hammari, Per Gunnar Kjeldsberg, Francky Catthoor |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2018 | Algorithm/Architecture Co-optimisation Technique for Automatic Data Reduction of Wireless Read-Out in High-Density Electrode ArraysabstractHigh-density electrode arrays used to read out neural activity will soon surpass the limits of the amount of data that can be transferred within reasonable energy budgets. This is true for wired brain implants when the required bandwidth becomes very high, and even more so for untethered brain implants that require wireless transmission of data. We propose an energy-efficient spike data extraction solution for high-density electrode arrays, capable of reducing the data to be transferred by over 85%. We combine temporal and spatial spike data analysis with low implementation complexity, where amplitude thresholds are used to detect spikes and the spatial location of the electrodes is used to extract potentially useful sub-threshold data on neighboring electrodes. We tested our method against a state-of-the-art spike detection algorithm, with prohibitively high implementation complexity, and found that the majority of spikes are extracted reliably. We obtain further improved quality results when ignoring very small spikes below 30% of the voltage thresholds, resulting in 91% accuracy. Our approach uses digital logic and is therefore scalable with an increasing number of electrodes. Yahya H. Yassin, Francky Catthoor, Fabian Kloosterman, Jyh-Jang Sun, João Couto, Per Gunnar Kjeldsberg, Nick Van Helleputte |
ACM Trans. Embed. Comput. Syst. | 6 |
| 2017 | READEX: Linking two ends of the computing continuum to improve energy-efficiency in dynamic applicationsabstractIn both the embedded systems and High Performance Computing domains, energy-efficiency has become one of the main design criteria. Efficiently utilizing the resources provided in computing systems ranging from embedded systems to current petascale and future Exascale HPC systems will be a challenging task. Suboptimal designs can potentially cause large amounts of underutilized resources and wasted energy. In both domains, a promising potential for improving efficiency of scalable applications stems from the significant degree of dynamic behaviour, e.g., runtime alternation in application resource requirements and workloads. Manually detecting and leveraging this dynamism to improve performance and energy-efficiency is a tedious task that is commonly neglected by developers. However, using an automatic optimization approach, application dynamism can be analysed at design time and used to optimize system configurations at runtime. The European Union Horizon 2020 READEX (Runtime Exploitation of Application Dynamism for Energy-efficient eX-ascale computing) project will develop a tools-aided auto-tuning methodology inspired by the system scenario methodology used in embedded systems. Dynamic behaviour of HPC applications will be exploited to achieve improved energy-efficiency and performance. Driven by a consortium of European experts from academia, HPC resource providers, and industry, the READEX project aims at developing the first of its kind generic framework to split design time and runtime automatic tuning while targeting heterogeneous system at the Exascale level. This paper describes plans for the project as well as early results achieved during its first year. Furthermore, it is shown how project results will be brought back into the embedded systems domain. Per Gunnar Kjeldsberg, Andreas Gocht, Michael Gerndt, Lubomir Riha, Joseph Schuchart, Umbreen Sabir Mian |
DATE | 1 |
| 2017 | Towards fine-grained dynamic tuning of HPC applications on modern multi-core architecturesabstractThere is a consensus that exascale systems should operate within a power envelope of 20MW. Consequently, energy conservation is still considered as the most crucial constraint if such systems are to be realized. Mohammed Sourouri, Espen Birger Raknes, Nico Reissmann, Johannes Langguth, Daniel Hackenberg, Robert Schöne, Per Gunnar Kjeldsberg |
SC | 7 |
| 2016 | Dynamic Hardware Management of the H264/AVC Encoder Control Structure Using a Framework for System ScenariosabstractMany modern applications exhibit dynamic behavior, which can be exploited for reduced energy consumption. We employ a two-phase combined design-time/run-time methodology that identifies different run-time situations and clusters similar behaviors into system scenarios. This methodology is integrated with our framework for system scenario based designs, which dynamically tune the hardware to match the application behavior. We achieve significant energy reductions for an extracted control structure of a video codec widely used in hand-held devices today. We encode a video stream consisting of different frame sizes based on measured available wireless bandwidth. Energy consumption is measured with a modified microcontroller board from Atmel with two alternative voltage and frequency settings. While maintaining the perceptual video quality and frame rate, our method results in up to 44% energy reduction for our encoded streams, even after including an average worst-case tuning overhead of 8.5%. In reality this overhead is typically negligible since the worst-case assumes very frequent tuning, while in realistic situations it will be performed at a much lower rate. This makes the expected gain over 50%. Yahya H. Yassin, Per Gunnar Kjeldsberg, Andrew Perkis, Francky Catthoor |
DSD | 2 |
| 2016 | Integrated Exploration Methodology for Data Interleaving and Data-to-Memory Mapping on SIMD ArchitecturesabstractThis work presents a methodology for efficient exploration of data interleaving and data-to-memory mapping options for Single Instruction Multiple Data (SIMD) platform architectures. The system architecture consists of a reconfigurable clustered scratch-pad memory and a SIMD functional unit, which performs the same operation on multiple input data in parallel. The memory accesses contribute substantially to the overall energy consumption of an embedded system executing a data intensive task. The scope of this work is the reduction of the overall energy consumption by increasing the utilization of the functional units and decreasing the number of memory accesses. The presented methodology is tested using a number of benchmark applications with holes in their access scheme. Potential gains are calculated based on the energy models, both for the processing and the memory part of the system. The reduction in energy consumption after efficient interleaving and mapping of data is between 40% and 80% for the complete system and the studied benchmarks. Iasonas Filippopoulos, Namita Sharma 0001, Francky Catthoor, Per Gunnar Kjeldsberg, Preeti Ranjan Panda |
ACM Trans. Embed. Comput. Syst. | 4 |
| 2014 | Systematic Exploration of Power-Aware Scenarios for IEEE 802.11ac WLAN SystemsabstractThis work explores the power management options for a transmitting wireless system using system scenarios. We exploit the variations in the communication channel and the protocol requirements during the lifetime of a transmission, in order to optimize energy usage. Both the transmission signal power and the memory subsystem are taken into consideration. Different system scenarios and the corresponding configurations capture the different resource requirements, which change dynamically during transmission. Signal power on the antenna and active memory banks are the two main platform parameters explored in this study and sufficiently detailed system models are presented for both. The trade-off between the accuracy of the generated system scenarios and the switching cost between them is analyzed. The exploration is performed for an increasing number of system scenarios, from 1 to 14, and the reported power gains are over 95% and over 25% on the signal power and the memory subsystems respectively. Nikolaos Zompakis, Iasonas Filippopoulos, Per Gunnar Kjeldsberg, Francky Catthoor, Dimitrios Soudris |
DSD | 3 |
| 2008 | Power optimization of weighted bit-product summation tree for elementary function generatorabstractIn this paper we propose a method for lowering the power consumption in our previously proposed method for approximating elementary functions. By rearranging the interconnect ordering in the summation tree we show that it is possible to lower the power consumption in the range of 5.4% to 25.6% compared to a random ordering. The reduction tree is progressively designed and the interconnect ordering is decided based on the transition activities of the partial products. The reduction in power consumption comes with no overhead in performance or area compared to the random ordering. Saeeid Tahmasbi Oskuii, Kenny Johansson, Oscar Gustafsson, Per Gunnar Kjeldsberg |
ISCAS | 4 |
| 2007 | Fast memory footprint estimation based on maximal dependency vector calculationabstractIn data dominated applications, loop transformations have a huge impact on the lifetime of array data and therefore on memory footprint. Since a locally optimal loop transformation may have a detrimental effect somewhere else, many alternative loop transformations need to be explored. Therefore, estimation of the memory footprint is essential, and this estimation has to be fast. This paper presents a fast array based memory footprint estimation technique based on counting of iteration nodes in an iteration domain constrained by a maximal lifetime. The maximal lifetime is defined by the maximal dependency vector (MDV) of the array for a given execution ordering. We further present for the first time two approaches for calculation of the MDV: a general approach based on an ILP formulation and a novel vertexes approach when iteration domains are approximated by bounding boxes. Experiments on practical test vehicles demonstrate that the estimation based on our vertexes approach is extremely fast, on average two orders of magnitude faster than the compared approaches, while still keeping the accuracy high. This enables system-level data memory footprint exploration of many different alternative transformed program codes, within interactive time limits, and on realistic complex applications Qubo Hu, Arnout Vandecappelle, Per Gunnar Kjeldsberg, Francky Catthoor, Martin Palkovic |
DATE | 3 |
| 2007 | Probabilistic gate-level power estimation using a novel waveform set methodabstractA probabilistic power estimation technique for combinational circuits is presented. A novel set of simple waveforms is the kernel of this technique. The transition density of each circuit node is estimated. Existing methods have local glitch filtering approaches that fail to model this phenomenon correctly. Glitches originated from a node may be filtered in some, but not necessarily all, of its successor nodes. Our waveform set approach allows us to utilize a global glitch filtering technique that can model the removal ofglitches in more detail. It produces error free estimates for tree structured circuits. For other circuit, experimental results using the ISCAS'85 benchmarks show that the waveform set method generally provides significantly better estimates of the transition density compared to previous techniques. Saeeid Tahmasbi Oskuii, Per Gunnar Kjeldsberg, Einar J. Aas |
ACM Great Lakes Symposium on VLSI | 2 |
| 2007 | Transition-activity aware design of reduction-stages for parallel multipliersabstractWe propose an interconnect reorganization algorithm for reduction stages in parallel multipliers. It aims at minimizing power consumption for given static probabilities at the primary inputs. In typical signal processing applications the transition probability varies between the most and least significant bits. The same is the case for individual signals within the multiplier. Our interconnect reorganization exploits this to reduce the overall switching activity, thus reducing the multiplier's power consumption. We have developed a CAD tool that reorganizes the connections within the multiplier architecture in an optimized way. Since the applied heuristic requires power estimation, we have also developed a very fast estimator fine tuned for parallel multipliers. The CAD tool automatically generates gate-level VHDL code for the optimizedmultipliers. This code and code for unoptimized multipliers have been compared using state of the art power estimation tools. The reduction in power consumption ranges from 7% up to 23% and can be achieved without any noticeable overhead in performance and area. Saeeid Tahmasbi Oskuii, Per Gunnar Kjeldsberg, Oscar Gustafsson |
ACM Great Lakes Symposium on VLSI | 2 |
| 2007 | Bit-Width Constrained Memory Hierarchy Optimization for Real-Time Video SystemsabstractThe great variety of pixel dynamics of real-time video-processing systems (RTVPS), ranging from color, grayscale, or binary pixels, means that a careful design and specification of bit widths is required. It is obvious that the bit-width specification will affect the total memory storage requirement. However, what is not so obvious is that the bit-width specification will also affect the design of the memory hierarchy, an impact similar for both hardware and software implementations. We have developed an integer-nonlinear-program formulation for the optimization of the memory hierarchy of RTVPS. An active surveillance video camera is introduced as a test case. We demonstrate how the optimization model can reduce the on-chip memory storage by 61% compared to a nonoptimal memory hierarchy. Benny Thörnberg, Martin Palkovic, Qubo Hu, Leif Olsson, Per Gunnar Kjeldsberg, Mattias O'Nils, Francky Catthoor |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2007 | Incremental hierarchical memory size estimation for steering of loop transformationsabstractModern embedded multimedia and telecommunications systems need to store and access huge amounts of data. This becomes a critical factor for the overall energy consumption, area, and performance of the systems. Loop transformations are essential to improve the data access locality and regularity in order to optimally design or utilize a memory hierarchy. However, due to abstract high-level cost functions, current loop transformation steering techniques do not take the memory platform sufficiently into account. They usually also result in only one final transformation solution. On the other hand, the loop transformation search space for real-life applications is huge, especially if the memory platform is still not fully fixed. Use of existing loop transformation techniques will therefore typically lead to suboptimal end-products. It is critical to find all interesting loop transformation instances. This can only be achieved by performing an evaluation of the effect of later design stages at the early loop transformation stage. This article presents a fast incremental hierarchical memory-size requirement estimation technique. It estimates the influence of any given sequence of loop transformation instances on the mapping of application data onto a hierarchical memory platform. As the exact memory platform instantiation is often not yet defined at this high-level design stage, a platform-independent estimation is introduced with a Pareto curve output for each loop transformation instance. Comparison among the Pareto curves helps the designer, or a steering tool, to find all interesting loop transformation instances that might later lead to low-power data mapping for any of the many possible memory hierarchy instances. Initially, the source code is used as input for estimation. However, performing the estimation repeatedly from the source code is too slow for large search space exploration. An incremental approach, based on local updating of the previous result, is therefore used to handle sequences of different loop transformations. Experiments show that the initial approach takes a few seconds, which is two orders of magnitude faster than state-of-the-art solutions but still too costly to be performed interactively many times. The incremental approach typically takes just a few milliseconds, which is another two orders of magnitude faster than the initial approach. This huge speedup allows us for the first time to handle real-life industrial-size applications and get realistic feedback during loop transformation exploration. Qubo Hu, Per Gunnar Kjeldsberg, Arnout Vandecappelle, Martin Palkovic, Francky Catthoor |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2006 | Loop Transformation Methodologies for Array-Oriented Memory ManagementabstractThe storage requirements in data-dominant signal processing systems, whose behavior is described by arraybased, loop-organized algorithmic specifications, have an important impact on the overall energy consumption, data access latency, and chip area. Applying different loop transformations on the specification code can significantly enhance the memory management of such VLSI systems, improving all the major parameters of the design space - power, area, and performance. This paper gives a global view on existing and recently proposed memory size evaluation approaches for procedural and non-procedural specifications. Moreover, it discusses typical memory management trade-offs taken into account during the exploration of system specifications by loop transformations, that can exploit these early size evaluations. Florin Balasa, Per Gunnar Kjeldsberg, Martin Palkovic, Arnout Vandecappelle, Francky Catthoor |
ASAP | 2 |
| 2006 | Hierarchical memory size estimation for loop fusion and loop shifting in data-dominated applicationsabstractLoop fusion and loop shifting are important transformations for improving data locality to reduce the number of costly accesses to off-chip memories. Since exploring the exact platform mapping for all the loop transformation alternatives is a time consuming process, heuristics steered by improved data locality are generally used. However, pure locality estimates do not sufficiently take into account the hierarchy of the memory platform. This paper presents a fast, incremental technique for hierarchical memory size requirement estimation for loop fusion and loop shifting at the early loop transformations design stage. As the exact memory platform is often not yet defined at this stage, we propose a platform-independent approach which reports the Pareto-optimal trade-off points for scratch-pad memory size and off-chip memory accesses. The estimation comes very close to the actual platform mapping. Experiments on realistic test-vehicles confirm that. It helps the designer or a tool to find the interesting loop transformations that should then be investigated in more depth afterward. Qubo Hu, Arnout Vandecappelle, Martin Palkovic, Per Gunnar Kjeldsberg, Erik Brockmeyer, Francky Catthoor |
ASP-DAC | 4 |
| 2006 | Polyhedral space generation and memory estimation from interface and memory models of real-time video systems
Benny Thörnberg, Qubo Hu, Martin Palkovic, Mattias O'Nils, Per Gunnar Kjeldsberg |
J. Syst. Softw. | 5 |
| 2004 | Memory Requirement Optimization with Loop Fusion and Loop ShiftingabstractLoop fusion and loop shifting are well recognized loop transformations for memory requirement reduction. State-of-the-art optimizations with loop fusion and shifting are based on heuristics without any evaluation of the resulting effects during each optimization step. Thus we cannot guarantee that each step results in a reduced overall memory requirement. On the other hand, most memory requirement estimations at system level are inefficient and slow. Also the estimation is not started until the optimization is done. Having to iterate between optimization and estimation is very time consuming. In this paper, we present a storage requirement optimization method which combines the optimization and estimation processes with the goal to have continuous estimates during the optimization and hence to achieve lower memory requirements. Qubo Hu, Martin Palkovic, Per Gunnar Kjeldsberg |
DSD | 3 |
| 2004 | Storage requirement estimation for optimized design of data intensive applicationsabstractA novel storage requirement estimation methodology is presented for use in the early system design phases when the data transfer ordering is only partially fixed. At that stage, none of the existing estimation tools are adequate, as they either assume a fully specified execution order or ignore it completely. A prototype CAD tool has been developed that includes major parts of the storage requirement estimation and optimization methodology. Using representative application demonstrators, we show how our techniques and tool can effectively guide the designer to achieve a transformed specification with low storage requirement. Per Gunnar Kjeldsberg, Francky Catthoor, Einar J. Aas |
ACM Trans. Design Autom. Electr. Syst. | 1 |
| 2003 | Data dependency size estimation for use in memory optimizationabstractA novel storage requirement estimation methodology is presented for use in the early system design phases when the data transfer ordering is only partly fixed. At that stage, none of the existing estimation tools are adequate, as they either assume a fully specified execution order or ignore it completely. This paper presents an algorithm for automated estimation of strict upper and lower bounds on the individual data dependency sizes in high-level application code given a partially fixed execution ordering. In the overall estimation technique, this is followed by a detection of the maximally combined size of simultaneously alive dependencies, resulting in the overall storage requirement of the application. Using representative application demonstrators, we show how our techniques can effectively guide the designer to achieve a transformed specification with low storage requirement. Per Gunnar Kjeldsberg, Francky Catthoor, Einar J. Aas |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2001 | Detection of Partially Simultaneously Alive Signals in Storage Requirement Estimation for Data Intensive ApplicationsabstractIn this paper, we propose a novel storage requirement estimation methodology for use in the early system design phases when the data transfer ordering is only partially fixed. At that stage, none of the existing estimation tools are adequate, as they either assume a fully specified execution order or ignore it completely. Using representative application demonstrators, we show how our technique can effectively guide the designer to achieve a transformed specification with low storage requirement. Per Gunnar Kjeldsberg, Francky Catthoor, Einar J. Aas |
DAC | 1 |
| 2001 | Data and memory optimization techniques for embedded systemsabstractWe present a survey of the state-of-the-art techniques used in performing data and memory-related optimizations in embedded systems. The optimizations are targeted directly or indirectly at the memory subsystem, and impact one or more out of three important cost metrics: area, performance, and power dissipation of the resulting implementation. We first examine architecture-independent optimizations in the form of code transoformations. We next cover a broad spectrum of optimization techniques that address memory architectures at varying levels of granularity, ranging from register files to on-chip memory, data caches, and dynamic memory (DRAM). We end with memory addressing related issues. Preeti Ranjan Panda, Francky Catthoor, Nikil Dutt, Koen Danckaert, Erik Brockmeyer, Chidamber Kulkarni, Arnout Vandecappelle, Per Gunnar Kjeldsberg |
ACM Trans. Design Autom. Electr. Syst. | 8 |
| 2000 | Automated Data Dependency Size Estimation with a Partially Fixed Execution OrderingabstractFor data dominated applications, the system level design trajectory should first focus on finding a good data transfer and storage solution. Since no realization details are available at this level, estimates are needed to guide the designer. This paper presents an algorithm for automated estimation of strict upper and lower bounds on the individual data dependency sizes in high level application code given a partially fixed execution ordering. Previous work has either not taken execution ordering into account at all, resulting in large overestimates, or required a fully specified ordering which is usually not available at this high level. The usefulness of the methodology is illustrated on representative application demonstrators. Per Gunnar Kjeldsberg, Francky Catthoor, Einar J. Aas |
ICCAD | 1 |