EDBT 2026 Demo / reviewers in the wild / expert
Per Larsson-Edefors
dblp:80/3712
· DBLP profile ↗
36ranked-venue papers
4as first author
5since 2021 · last 2025
0000-0001-5779-4313ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 34 · 3 first-author · 4 since 2021Software engineering, systems software and programming languages · 4Theory of computation · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Implementation Evaluation of Fixed-Point Multipliers for Complex NumbersabstractComplex multipliers are common in signal processing and scientific applications. A direct implementation of complex multiplication involves four real-valued multiplications and two real-valued additions. But it is well known that there are alternatives to the direct implementation in which some computations are shared. The rationale for these alternatives would be that fewer multiplications are used, potentially reducing hardware resource usage. We consider three distinctly different complex multiplication schemes, implement them as fixed-point complex multipliers using a 7-nm predictive technology and quantitatively evaluate their area usage and energy efficiency. Per Larsson-Edefors, Erik Börjeson |
ARITH | 1 |
| 2025 | Low-Power Complex Multiplier Pin Assignment Based on Spatial and Temporal Signal PropertiesabstractFixed-point integer multipliers are power-intensive components that are integral to many systems in computing and digital signal processing. Operating on complex fixed-point numbers, the complex multiplier is critical to a wide range of applications including communication systems. Since channel properties drift over time, communication systems require adaptive processing blocks which have to be designed for the worst-case scenario. This raises the question of how we can take advantage of performance variations of a system to reduce power dissipation. We describe how knowledge on variations in both dynamic range (the spatial dimension) and switching frequency (the temporal dimension) can be used to assign pins of complex multipliers in order to minimize power dissipation. Using netlist synthesis based on the predictive 7-nm ASAP7 cell library, we find that, for instance, if one of two 12-bit input signals of the complex multiplier has a 2-bit reduced dynamic range and a 50% reduced switching frequency, we decrease the energy per operation by 20% by selecting the optimal pin assignment. Per Larsson-Edefors, Erik Börjeson |
ISCAS | 1 |
| 2025 | FPGA-Based Wordlength Optimization for DSPabstractFixed-point representations are commonly used in DSP designs to efficiently use hardware resources. It is, however, a challenge to determine an appropriate fractional wordlength of all signals in order to reach a good balance between accuracy and hardware cost. Extensive simulations can be used to characterize a design and perform wordlength optimization (WLO), however, this tends to be slow when the DSP design is complex. FPGA emulation, which allows data to be streamed in hardware, is significantly faster than software simulation. We introduce an FPGA-accelerated WLO framework which utilizes a new WLO algorithm based on a tree-structured Parzen estimator developed for higher convergence speed. This framework is evaluated for three DSP designs; two finite-impulse response filters and one phase recovery design. The results show that our new WLO framework reduces the DSP accuracy evaluation time by a factor of 300–500 over simulation. Jinsheng Bian, Erik Börjeson, Per Larsson-Edefors |
ISCAS | 4 |
| 2025 | Energy-Efficient Computation of TensorFloat32 Numbers on an FP32 MultiplierabstractSeveral new shorter floating-point formats have been proposed to match requirements of emerging application workloads. To simplify hardware development in the presence of an increasing number of formats, one practical design option is to use as much as possible preexisting hardware, such as standard 32-bit IEEE-754 (FP32) floating-point units, to handle emerging, less complex formats. We evaluate the case where we use an FP32 multiplier to run Nvidia TensorFloat32 data. While the FP32 multiplier area is not as small as a dedicated TensorFloat32 multiplier, we show that energy per operation scales well with the mantissa width reduction and that smart pin assignment can leverage uneven input vector switching activities to significantly decrease energy for reduced precisions. Per Larsson-Edefors |
VLSI-SoC | 1 |
| 2021 | Variable-Rate VLSI Architecture for 400-Gb/s Hard-Decision Product DecoderabstractVariable-rate transceivers, which adapt to the conditions, will be central to energy-efficient communication. However, fiber-optic communication systems with high bit-rate requirements make design of flexible transceivers challenging, since additional circuits needed to orchestrate the flexibility will increase area and degrade speed. We propose a variable-rate VLSI architecture of a forward error correction (FEC) decoder based on hard-decision product codes. Variable shortening of component codes provides a mechanism by which code rate can be varied, the number of iterations offers a knob to control the coding gain, while a key-equation solver module that can swap between error-locator polynomial coefficients provides a means to change error correction capability. Our evaluations based on 28-nm netlists show that a variable-rate decoder implementation can offer a net coding gain (NCG) range of 9.96-10.38dB at a post-FEC bit-error rate of 10-15. The decoder achieves throughputs in excess of 400Gb/s, latencies below 53ns, and energy efficiencies of 1.14pJ/bit or less. While the area of the variable-rate decoder is 31% larger than a decoder with a fixed rate, the power dissipation is a mere 5% higher. The variable error correction capability feature increases the NCG range further, to above 10.5dB, but at a significant area cost. Vikram Jain, Christoffer Fougstedt, Per Larsson-Edefors |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2019 | Hardware Considerations for Selection NetworksabstractThe selection operation is a central part of a soft-decision error-correction algorithm, which is important for high-performance communication networks. High symbol rates and power-dissipation limitations motivate hardware implementation as a comparator network. We use industry-standard tools to investigate VLSI hardware implementation of selection networks with up to 512 inputs. We find theoretical network depth and size to be poor predictors of hardware performance. In a 65-nm process, we find that our novel half-life network is competitive with and in some cases superior to Zazon-Ivry's pairwise and odd/even selection networks, for delay, area, and energy per selection operation. Kenneth Peter, Lars J. Svensson, Christoffer Fougstedt, Per Larsson-Edefors |
VLSI-SoC | 4 |
| 2016 | Redesigning a tagless access buffer to require minimal ISA changesabstractEnergy efficiency is a first-order design goal for nearly all classes of processors, but it is particularly important in mobile and embedded systems. Data caches in such systems account for a large portion of the processor's energy usage, and thus techniques to improve the energy efficiency of the cache hierarchy are likely to have high impact. Our prior work reduced data cache energy via a tagless access buffer (TAB) that sits at the top of the cache hierarchy. Strided memory references are redirected from the level-one data cache (L1D) to the smaller, more energy-efficient TAB. These references need not access the data translation lookaside buffer (DTLB), and they can avoid unnecessary transfers from lower levels of the memory hierarchy. The original TAB implementation requires changing the immediate field of load and store instructions, necessitating substantial ISA modifications. Here we present a new TAB design that requires minimal instruction set changes, gives software more explicit control over TAB resource management, and remains compatible with legacy (non-TAB) code. With a line size of 32 bytes, a four-line TAB can eliminate 31% of L1D accesses, on average. Together, the new TAB, L1D, and DTLB use 22% less energy than a TAB-less hierarchy, and the TAB system decreases execution time by 1.7%. Carlos Sanchez, Peter Gavin, Daniel Moreau, Magnus Själander, David B. Whalley, Per Larsson-Edefors, Sally A. McKee |
CASES | 6 |
| 2016 | Practical way halting by speculatively accessing halt tags
Daniel Moreau, Alen Bardizbanyan, Magnus Själander, David B. Whalley, Per Larsson-Edefors |
DATE | 5 |
| 2015 | Exploring early and late ALUs for single-issue in-order pipelinesabstractIn-order processors are key components in energy-efficient embedded systems. One important design aspect of inorder pipelines is the sequence of pipeline stages: First, the position of the execute stage, in which arithmetic logic unit (ALU) operations and branch prediction are handled, impacts the number of stall cycles that are caused by data dependencies between data memory instructions and their consuming instructions and by address generation instructions that depend on an ALU result. Second, the position of the ALU inside the pipeline impacts the branch penalty. This paper considers the question on how to best make use of ALU resources inside a single-issue in-order pipeline. We begin by analyzing which is the most efficient way of placing a single ALU in an in-order pipeline. We then go on to evaluate what is the most efficient way to make use of two ALUs, one early and one late ALU, which is a technique that has revitalized commercial in-order processors in recent years. Our architectural simulations, which are based on 20 MiBench and 7 SPEC2000 integer benchmarks and a 65-nm postlayout netlist of a complete pipeline, show that utilizing two ALUs in different stages of the pipeline gives better performance and energy efficiency than any other pipeline configuration with a single ALU. Alen Bardizbanyan, Per Larsson-Edefors |
ICCD | 2 |
| 2015 | Improving Data Access Efficiency by Using Context-Aware Loads and StoresabstractMemory operations have a significant impact on both performance and energy usage even when an access hits in the level-one data cache (L1 DC). Load instructions in particular affect performance as they frequently result in stalls since the register to be loaded is often referenced before the data is available in the pipeline. L1 DC accesses also impact energy usage as they typically require significantly more energy than a register file access. Despite their impact on performance and energy usage, L1 DC accesses on most processors are performed in a general fashion without regard to the context in which the load or store operation is performed. We describe a set of techniques where the compiler enhances load and store instructions so that they can be executed with fewer stalls and/or enable the L1 DC to be accessed in a more energy-efficient manner. We show that using these techniques can simultaneously achieve a 6% gain in performance and a 43% reduction in L1 DC energy usage. Alen Bardizbanyan, Magnus Själander, David B. Whalley, Per Larsson-Edefors |
LCTES | 4 |
| 2014 | Reducing set-associative L1 data cache energy by early load data dependence detection (ELD3)abstractFast set-associative level-one data caches (L1 DCs) access all ways in parallel during load operations for reduced access latency. This is required in order to resolve data dependencies as early as possible in the pipeline, which otherwise would suffer from stall cycles. A significant amount of energy is wasted due to this fast access, since the data can only reside in one of the ways. While it is possible to reduce L1 DC energy usage by accessing the tag and data memories sequentially, hence activating only one data way on a tag match, this approach significantly increases execution time due to an increased number of stall cycles. We propose an early load data dependency detection (ELD3) technique for in-order pipelines. This technique makes it possible to detect if a load instruction has a data dependency with a subsequent instruction. If there is no such dependency, then the tag and data accesses for the load are sequentially performed so that only the data way in which the data resides is accessed. If there is a dependency, then the tag and data arrays are accessed in parallel to avoid introducing additional stall cycles. For the MiBench benchmark suite, the ELD3technique enables about 49% of all load operations to access the L1 DC sequentially. Based on 65-nm data using commercial SRAM blocks, the proposed technique reduces L1 DC energy by 13%. Alen Bardizbanyan, Magnus Själander, David B. Whalley, Per Larsson-Edefors |
DATE | 4 |
| 2014 | Assessing scrubbing techniques for Xilinx SRAM-based FPGAs in space applicationsabstractSRAM-based FPGAs are becoming increasingly attractive for use in space applications due to their reconfigurability and signal processing capabilities, as well as their increasing speed and capacity. Traditional SRAM-based FPGAs, however, are highly sensitive to the ionizing radiation environment in space, making them prone to radiation-induced memory upsets. In this paper, we evaluate and compare scrubbing techniques for Xilinx SRAM-based FPGAs with respect to radiation-induced single event upsets. A test framework using an exchangeable payload is developed for this purpose and run on a Xilinx Virtex-5 FPGA. We show that recent SRAM-based FPGAs can constitute a cost-efficient alternative to radiation-hardened or antifuse FPGAs for non-critical space application such as satellite instruments. Fredrik Brosser, Emil Milh, Vilhelm Geijer, Per Larsson-Edefors |
FPT | 4 |
| 2013 | Improving data access efficiency by using a tagless access buffer (TAB)abstractThe need for energy efficiency continues to grow for many classes of processors, including those for which performance remains vital. Data cache is crucial for good performance, but it also represents a significant portion of the processor's energy expenditure. We describe the implementation and use of a tagless access buffer (TAB) that greatly improves data access energy efficiency while slightly improving performance. The compiler recognizes memory reference patterns within loops and allocates these references to a TAB. This combined hardware/software approach reduces energy usage by (1) replacing many level-one data cache (L1D) accesses with accesses to the smaller, more power-efficient TAB; (2) removing the need to perform tag checks or data translation lookaside buffer (DTLB) lookups for TAB accesses; and (3) reducing DTLB lookups when transferring data between the L1D and the TAB. Accesses to the TAB occur earlier in the pipeline, and data lines are prefetched from lower memory levels, which result in a small performance improvement. In addition, we can avoid many unnecessary block transfers between other memory hierarchy levels by characterizing how data in the TAB are used. With a combined size equal to that of a conventional 32-entry register file, a four-entry TAB eliminates 40% of L1D accesses and 42% of DTLB accesses, on average. This configuration reduces data-access related energy by 35% while simultane-ously decreasing execution time by 3%. Alen Bardizbanyan, Peter Gavin, David B. Whalley, Magnus Själander, Per Larsson-Edefors, Sally A. McKee, Per Stenström |
CGO | 5 |
| 2013 | Speculative tag access for reduced energy dissipation in set-associative L1 data cachesabstractDue to performance reasons, all ways in set-associative level-one (L1) data caches are accessed in parallel for load operations even though the requested data can only reside in one of the ways. Thus, a significant amount of energy is wasted when loads are performed. We propose a speculation technique that performs the tag comparison in parallel with the address calculation, leading to the access of only one way during the following cycle on successful speculations. The technique incurs no execution time penalty, has an insignificant area overhead, and does not require any customized SRAM implementation. Assuming a 16kB 4-way set-associative L1 data cache implemented in a 65-nm process technology, our evaluation based on 20 different MiBench benchmarks shows that the proposed technique on average leads to a 24% data cache energy reduction. Alen Bardizbanyan, Magnus Själander, David B. Whalley, Per Larsson-Edefors |
ICCD | 4 |
| 2013 | A SiGe 8-channel comparator for application in a synthetic aperture radiometerabstractWe present a high-speed low-power 8-channel comparator tailored for the application of sampling antenna signals in a cross-correlator system for space-borne synthetic aperture radiometer instruments. Features like clock return path, perchannel offset calibration and bias current tuning make the comparator adaptable and gives the possibility to adjust the comparator for low power consumption, while keeping performance within the requirements of the cross-correlator system. The comparator has been implemented and fabricated in a 130-nm SiGe BiCMOS process. Measurements show that the comparator can perform sampling at a rate of 4.5 GS/s with a power consumption of 48 mW/channel or 1 GS/s with a power consumption of 17 mW/channel. Erik Ryman, Stefan Back Andersson, J. Riesbeck, Slavko Dejanovic, Anders Emrich, Per Larsson-Edefors |
ISCAS | 6 |
| 2013 | Designing a practical data filter cache to improve both energy efficiency and performanceabstractConventional Data Filter Cache (DFC) designs improve processor energy efficiency, but degrade performance. Furthermore, the single-cycle line transfer suggested in prior studies adversely affects Level-1 Data Cache (L1 DC) area and energy efficiency. We propose a practical DFC that is accessed early in the pipeline and transfers a line over multiple cycles. Our DFC design improves performance and eliminates a substantial fraction of L1 DC accesses for loads, L1 DC tag checks on stores, and data translation lookaside buffer accesses for both loads and stores. Our evaluation shows that the proposed DFC can reduce the data access energy by 42.5% and improve execution time by 4.2%. Alen Bardizbanyan, Magnus Själander, David B. Whalley, Per Larsson-Edefors |
ACM Trans. Archit. Code Optim. | 4 |
| 2012 | Viterbi Accelerator for Embedded Processor DatapathsabstractWe present a novel architecture for a lightweight Viterbi accelerator that can be tightly integrated inside an embedded processor datapath. We investigate the accelerator's impact on processor performance by using the EEMBC Viterbi benchmark and the in-house Viterbi Branch Metric kernel. Our evaluation based on the EEMBC benchmark shows that an accelerated 65-nm 2.7-ns processor datapath is 20% larger but 90% more cycle efficient than a datapath lacking the Viterbi accelerator, leading to an 87% overall energy reduction and a data throughput of 3.52 Mbit/s. Muhammad Waqar Azhar, Magnus Själander, Hasan Ali, Akshay Vijayashekar, Tung Thanh Hoang, Kashan Khurshid Ansari, Per Larsson-Edefors |
ASAP | 7 |
| 2012 | Feasibility study of FPGA-based equalizer for 112-Gbit/s optical fiber receiversabstractWith ever increasing demands on spectral efficiency, complex modulation schemes are being introduced in fiber communication. However, these schemes are challenging to implement as they drastically increase the computational burden at the fiber receiver's end. We perform a feasibility study of implementing a 16-QAM112-Gbit/s decision directed equalizer on a state-of-the-art FPGA platform. An FPGA offers the reconfigurability needed to allow for modulation scheme updates, however, its clock rate is limited. For this purpose, we introduce a new phase correction technique to significantly relax the delay requirement on the critical phase-recovery feedback loop. Fredrik Toft, Niclas Rousk, Jonas Mårtensson 0002, Marco Forzati, Bengt-Erik Olsson, Per Larsson-Edefors |
ISCAS | 6 |
| 2010 | Design space exploration for an embedded processor with flexible datapath interconnectabstractThe design of an embedded processor is dependent on the application domain. Traditionally, design solutions specific to an application domain have been available in three forms: VLIW-based DSP processors, ASICs and FPGAs; each respectively offering generality of application domain, energy efficiency and flexibility. However, while matching the application domain to the resources needed, the design space becomes huge. We present FlexTools, a tool framework built around the FlexCore architecture to evaluate performance and energy efficiency for different applications. Here we demonstrate FlexTools for design space exploration with a focus on the data-routing flexibility of the FlexCore processor, in search of energy-efficient interconnect configurations that are both cycle-count and hardware efficient. Evaluation results suggest that a well-optimized instance of a 65-nm multiplier-extended FlexCore processor datapath, obtained using FlexTools, executes nine integer EEMBC benchmarks with a 15% cycle count reduction and dissipates 17% less energy than a reference MIPS datapath. Tung Thanh Hoang, Ulf Jalmbrant, Erik der Hagopian, Kasyab P. Subramaniyan, Magnus Själander, Per Larsson-Edefors |
ASAP | 6 |
| 2010 | Cyclic Redundancy Checking (CRC) Accelerator for the FlexCore ProcessorabstractA proven approach to increase performance of general-purpose processors is to add hardware accelerators. In its basic configuration, the FlexCore processor has a limited set of datapath units. But thanks to a flexible datapath interconnect and a wide control word, the FlexCore datapath is explicitly designed to support integration of special units that, on demand, can accelerate certain data-intensive applications. We present the integration of a versatile accelerator for several Cyclic Redundancy Checking (CRC) keys. Furthermore, we investigate the accelerator's impact on processor execution time and energy efficiency, using the Power Stone CRC benchmark. Our evaluation shows that the accelerated 65-nm 2.7-ns FlexCore datapath is, for example, 86% more energy and cycle efficient than a datapath lacking the CRC accelerator. Muhammad Waqar Azhar, Tung Thanh Hoang, Per Larsson-Edefors |
DSD | 3 |
| 2010 | On-chip power supply noise and its implications on timingabstractWe address two problems of assessing the influence of power- supply variations on timing analysis. We present a method to assign a supply-dependent hold margin; and we describe a method to accurately characterize logic gates for the sen- sitivity of delay on supply-voltage variations. We use a com- mercial microcontroller as a design example. Lars J. Svensson, Johnny Pihl, Daniel A. Andersson, Per Larsson-Edefors |
ACM Great Lakes Symposium on VLSI | 4 |
| 2009 | Double Throughput Multiply-Accumulate unit for FlexCore processor enhancementsabstractAs a simple five-stage General-Purpose Processor (GPP), the baseline FlexCore processor has a limited set of datapath units. By utilizing a flexible datapath interconnect and a wide control word, a FlexCore processor is explicitly designed to support integration of special units that, on demand, can accelerate certain data-intensive applications. In this paper, we propose the integration of a novel Double Throughput Multiply-Accumulate (DTMAC) unit, whose different operating modes allow for on-thefly optimization of computational precision. For the two EEMBC benchmarks considered, the FlexCore processor performance is significantly enhanced when one DTMAC accelerator is included, translating into reduced execution time and energy dissipation. In comparison to the 32-bit GPP reference, the accelerated 32-bit FlexCore processor shows a 4.37× improvement in execution time and a 3.92× reduction in energy dissipation, for a benchmark with many consecutive 16-bit MAC operations. Tung Thanh Hoang, Magnus Själander, Per Larsson-Edefors |
IPDPS | 3 |
| 2009 | Multiplication Acceleration Through Twin PrecisionabstractWe present the twin-precision technique for integer multipliers. The twin-precision technique can reduce the power dissipation by adapting a multiplier to the bitwidth of the operands being computed. The technique also enables an increased computational throughput, by allowing several narrow-width operations to be computed in parallel. We describe how to apply the twin-precision technique also to signed multiplier schemes, such as Baugh-Wooley and modified-Booth multipliers. It is shown that the twin-precision delay penalty is small (5%-10%) and that a significant reduction in power dissipation (40%-70%) can be achieved, when operating on narrow-width operands. In an application case study, we show that by extending the multiplier of a general-purpose processor with the twin-precision scheme, the execution time of a Fast Fourier Transform is reduced with 15% at a 14% reduction in datapath energy dissipation. All our evaluations are based on layout-extracted data from multipliers implemented in 130-nm and 65-nm commercial process technologies. Magnus Själander, Per Larsson-Edefors |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2007 | High-Accuracy Architecture-Level Power Estimation for Partitioned SRAM Arrays in a 65-nm CMOS BPTM ProcessabstractIn this paper, we validate our previously proposed high- level power estimation models for a 65-nm BPTM process, using a physically partitioned 2-kB 6T-SRAM array. Also, we describe a new probing methodology that allows us to accurately capture not only subthreshold leakage, but also all other significant leakage mechanisms. By combining the probing methodology and the power models, we can estimate dynamic, leakage and total power of the partitioned 2-kB memory array with a 97% accuracy of that of full circuit-level simulations of the entire array. We also discuss the effect of partitioning on SRAM array power with respect to process technology scaling: Partitioning has the effect that leakage power constitutes an increasing fraction of total memory power, emphasizing the need to accurately capture leakage power in SRAM power models. Minh Quang Do, Per Larsson-Edefors, Mindaugas Drazdziulis |
DSD | 2 |
| 2006 | Multiplier reduction tree with logarithmic logic depth and regular connectivityabstractA novel partial-product reduction circuit for use in integer multiplication is presented. The high-performance multiplier (HPM) reduction tree has the ease of layout of a simple carry-save reduction array, but is in fact a high-speed low-power Dadda-style tree having a worst-case delay which depends on the logarithm (O(log TV)) of the word length N Henrik Eriksson, Per Larsson-Edefors, Mary Sheeran, Magnus Själander, Daniel Johansson, Martin Scholin |
ISCAS | 2 |
| 2006 | Toward architecture-based test-vector generation for timing verification of fast parallel multipliersabstractFast parallel multipliers that contain logarithmic partial-product reduction trees pose a challenge to simulation-based high-accuracy timing verification, since the reduction tree has many reconvergent signal branches. However, such a multiplier architecture also offers a clue as how to attack the test-vector generation problem. The timing-critical paths are intimately associated with long carry propagation. We introduce a multiplier test-vector generation method that has the ability to exercise such long carry propagation paths. Through extensive circuit simulation and static timing analysis, we evaluate the quality of the test vectors that result from the new method. Especially for fast multipliers with a pronounced carry propagation, the timing-critical vectors manage to stimulate a path, which has a delay that comes close to the true worst case delay. We investigate the complexity and run-time for the test-vector generation, and derive timing-critical vectors up to a factor word length of 54 bits. Henrik Eriksson, Per Larsson-Edefors, Daniel Eckerbert |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2005 | Accounting for the skin effect during repeater insertionabstractSince the skin effect will increase the propagation delay in an interconnect, it will also affect how to optimally select the number and size of the buffers. Failing to include the skin effect during buffer design may result in as much as 35% extra delay compared to the optimal repeater chain. We present a new method with closed-form expressions for repeater insertion where we take into account the skin effect and also the relationship between interconnect resistance, capacitance and inductance that are determined from the geometrical parameters. We also investigate the skin-effect influence on power dissipation for an optimally designed repeater chain, and find that the increase is at most 10% of the dynamic power dissipation. Daniel A. Andersson, Lars J. Svensson, Per Larsson-Edefors |
ACM Great Lakes Symposium on VLSI | 3 |
| 2004 | An Efficient Twin-Precision MultiplierabstractWe present a twin-precision multiplier that in normal operation mode efficiently performs N-b multiplications. For applications where the demand on precision is relaxed, the multiplier can perform N/2-b multiplications while expending only a fraction of the energy of a conventional N-b multiplier. For applications with high demands on throughput, the multiplier is capable of performing two independent N/2-b multiplications in parallel. A comparison between two signed 16-b multipliers, where both perform single 8-b multiplications, shows that the twin-precision multiplier has 72% lower power dissipation and 15% higher speed than the conventional one, while only requiring 8% more transistors. Magnus Själander, Henrik Eriksson, Per Larsson-Edefors |
ICCD | 3 |
| 2003 | A deep submicron power estimation methodology adaptable to variations between power characterization and estimationabstractTraditionally, RTL power estimation techniques are characterizing a component for a fixed environment (most importantly load capacitance, activity, and operating frequency). This article presents a solution to problems originating from the ineluctably changing operating conditions such as differing load capacitance due to different applications; different activity and operating frequency as power reduction techniques are more frequently employed. Daniel Eckerbert, Per Larsson-Edefors |
ASP-DAC | 2 |
| 2003 | Full-custom vs. standard-cell design flow: an adder case studyabstractFull-custom design is considered superior to standard-cell design when a high-performance circuit is requested. The structured routing of critical wires is considered to be the most important contributor to this performance gap. However, this is only true for bitsliced designs, such as ripple-carry adders, but not for designs with inter-bitslice interconnections spanning several bitslices, such as tree adders and reduction-tree multipliers. It is found that standard-cell design techniques scale better with the data width than full-custom bitsliced layouts for designs dominated by inter-bitslice interconnections. Henrik Eriksson, Per Larsson-Edefors, Tomas Henriksson, Christer Svensson |
ASP-DAC | 2 |
| 2003 | A Mixed-Mode Delay-Locked-Loop ArchitectureabstractWe present a mixed-mode delay-locked loop (DLL) architecture intended for multiple-phase clock generation. In contrast to analog DLLs, the proposed architecture allows for clock-gating; moreover, circuit simulations indicate that its performance (in terms of maximum frequency, frequency range, and low-speed power dissipation) is superior to that of a previously-reported, purely digital DLL. Daniel Eckerbert, Lars J. Svensson, Per Larsson-Edefors |
ICCD | 3 |
| 2001 | Interconnect-Driven Short-Circuit Power ModelingabstractEarly and accurate power estimation has become very important to meet the power budget in modern electronics design. In order to achieve early figures on power consumption, functional units have been modeled as black boxes and their respective power consumption has been modeled as a function of signal activity in inputs and outputs. Interconnects, and the associated load capacitances, are accounted for by simply adding the switching power consumption of the interconnect to the estimated power consumption of the functional unit driving the interconnect. Thus, the effects of load capacitance on short-circuit power in the functional units are not considered. The purpose of the present paper is two-fold: first, we put some focus on the review the effects that load capacitance has on short-circuit power. Secondly, we present two methods for interconnect-driven estimation of short-circuit power in state-of-the-art electronics design. Daniel Eckerbert, Per Larsson-Edefors |
DSD | 2 |
| 2000 | GLMC: interconnect length estimation by growth-limited multifold clusteringabstractIn this paper, interconnection length estimation is discussed and a general, simple, fast and efficient estimation technique is proposed. In contrast to traditional average length estimation techniques, such as the one based on Rent's rule, the new technique utilizes the topological information of the actual netlist and estimates the length of each interconnection separately. The result of the estimation can be directly used to assign a reasonable R and C to each interconnect, including long and wide buses. Consequently, the new technique enhances the accuracy of power and delay estimations at higher design levels of abstraction. Atila Alvandpour, Per Larsson-Edefors, Christer Svensson |
ISCAS | 2 |
| 2000 | An interconnect-driven design of a DFT processorabstractA new interconnect-driven DFT implementation is proposed in this paper. The normal way to implement the DFT is to use the FFT algorithm since it is computationally favorable. However, the increased speed comes at the cost of increased communications which give a higher power consumption. If the DFT algorithm is directly implemented instead, each channel becomes independent of all other channels and consequently communications and hence power consumption are reduced. Other benefits of using the DFT directly are the possibility to calculate a spectrum of any length, not only a power of two, and to have an irregular frequency step between channels. A number of ad hoc processing-element (PE) and system-level solutions are also proposed to reduce the power consumption even further. Daniel Eckerbert, Henrik Eriksson, Per Larsson-Edefors, Anders Edman |
ISCAS | 3 |
| 1998 | Separation and extraction of short-circuit power consumption in digital CMOS VLSI circuitsabstractIn this paper, we present a new technique which indirectly separates and extracts the total short-circuit power consumption of digital CMOS circuits. We avoid a direct encounter with the complex behavior of the short-circuit currents. Instead, we separate the dynamic power consumption from the total power and extract the total short-circuit power. The technique is based on two facts: first, the short-circuit power consumption disappears at a V/sub dd/ close to V/sub T/ and, secondly, the total capacitance depends on supply voltage in a sufficiently weak way in standard CMOS circuits. Hence, the total effective capacitance can be estimated at a low V/sub dd/. To avoid reducing V/sub dd/ below the specified forbidden level, a polynomial is used to estimate the power versus supply voltage down to V/sub T/ based on a small voltage sweep over the allowed supply voltage levels. The result shows good accuracy for the short-circuit current ranges of interest. Atila Alvandpour, Per Larsson-Edefors, Christer Svensson |
ISLPED | 2 |
| 1996 | Technology mapping onto very-high-speed standard CMOS hardwareabstractThis paper addresses technology mapping onto very-high-speed (>500 MHz) standard CMOS hardware. Both the technology mapping concept implemented in PRIMUS 2 and the hardware, on which the tool has to rely, are presented and discussed in the paper. PRIMUS 2 maps any multioutput multilevel combinational Boolean equation onto a predefined and precharacterized cell library. The exceptionally high clock rate and, hence, the strict gate delay requirement calls for a new mapping concept; PRIMUS 2 maps the set of equations onto a pipeline comprising dynamic gate cells with very low logic complexity. The only user-defined constraint on the mapping is the clock rate, and depending on this clock rate, PRIMUS 2 maps onto an adequate set of cells with adequate constraints on the placement and routing. To be able to control the timing behavior, the gate cells have to be very regular in size and shape, the transistor sizes have to be fixed and the cell library must contain not only gate cells but also gate interconnect (wire) cells. Thus, the maximal clock rate of the resulting hardware can be accurately tuned and controlled. A 1 GHz 1.0-/spl mu/m double-metal single-poly cell library has been designed and simulated in order to demonstrate the feasibility of the PRIMUS 2 technology mapping concept. Per Larsson-Edefors |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |