Wolfgang Fichtner

dblp:f/WolfgangFichtner · DBLP profile ↗
← Back
48ranked-venue papers
0as first author
0since 2021 · last 2011
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 46Software engineering, systems software and programming languages · 2Security and privacy · 1Graphics, computer vision, multimedia, augmented reality and games · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
16 papers
High-performance computing · 41% Electronic design automation · 24% Emerging computing paradigms · 16%
Network and information security
1 paper
Cryptographic primitives and cryptanalysis · 100%
Computer graphics and multimedia
1 paper
Audio and music processing · 100%

Topics — the 30 heaviest of 40, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
High-performance computing
scientific computing systems
0.122011
Atomistic nanoelectronic device engineering with sustained performances up to 1.44 PFlop/s · SC 2011
Three-dimensional numerical semiconductor device simulation: algorithms, architectures, results · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 1991
High-performance computing › performance optimization at scale
parallel scalability
0.112011
Atomistic nanoelectronic device engineering with sustained performances up to 1.44 PFlop/s · SC 2011
High-performance computing
performance optimization at scale
0.112011
Atomistic nanoelectronic device engineering with sustained performances up to 1.44 PFlop/s · SC 2011
Emerging computing paradigms › quantum computing › quantum simulation
quantum transport simulation
0.112011
Atomistic nanoelectronic device engineering with sustained performances up to 1.44 PFlop/s · SC 2011
Integrated circuit design
low-power circuit design
0.112006
Low-power architectural trade-offs in a VLSI implementation of an adaptive hearing aid algorithm · DAC 2006
Integrated circuit design
VLSI design
0.112006
Low-power architectural trade-offs in a VLSI implementation of an adaptive hearing aid algorithm · DAC 2006
Emerging computing paradigms
quantum computer architecture
0.012011
Atomistic nanoelectronic device engineering with sustained performances up to 1.44 PFlop/s · SC 2011
Cryptographic primitives and cryptanalysis › cryptographic implementation
block cipher implementation
0.012002
2Gbit/s Hardware Realizations of RIJNDAEL and SERPENT: A Comparative Analysis · CHES 2002
Cryptographic primitives and cryptanalysis › cryptographic implementation
hardware implementation
0.012002
2Gbit/s Hardware Realizations of RIJNDAEL and SERPENT: A Comparative Analysis · CHES 2002
Cryptographic primitives and cryptanalysis
symmetric cryptography
0.012002
2Gbit/s Hardware Realizations of RIJNDAEL and SERPENT: A Comparative Analysis · CHES 2002
Hardware accelerators and domain-specific architectures
cryptographic accelerator
0.012002
2Gbit/s Hardware Realizations of RIJNDAEL and SERPENT: A Comparative Analysis · CHES 2002
Electronic design automation
physical design
0.041993
Mixed element trees: a generalization of modified octrees for the generation of meshes for the simulation of complex 3-D semiconductor device structures · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 1993
Automatic rectangle-based adaptive mesh generation without obtuse angles · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 1992
A module generator based on the PQ-tree algorithm · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 1992
Electronic design automation › technology computer-aided design
semiconductor device simulation
0.041993
Mixed element trees: a generalization of modified octrees for the generation of meshes for the simulation of complex 3-D semiconductor device structures · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 1993
Three-dimensional numerical semiconductor device simulation: algorithms, architectures, results · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 1991
An adaptive grid refinement strategy for the drift-diffusion equations · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 1991
Electronic design automation
mesh generation
0.031993
Mixed element trees: a generalization of modified octrees for the generation of meshes for the simulation of complex 3-D semiconductor device structures · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 1993
Automatic rectangle-based adaptive mesh generation without obtuse angles · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 1992
Omega-an octree-based mixed element grid allocator for the simulation of complex 3-D device structures · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 1991
Electronic design automation
technology computer-aided design
0.031993
Mixed element trees: a generalization of modified octrees for the generation of meshes for the simulation of complex 3-D semiconductor device structures · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 1993
Three-dimensional numerical semiconductor device simulation: algorithms, architectures, results · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 1991
Omega-an octree-based mixed element grid allocator for the simulation of complex 3-D device structures · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 1991
Electronic design automation › technology computer-aided design
device simulation
0.021997
Optimized terminal current calculation for Monte Carlo device simulation · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 1997
Omega-an octree-based mixed element grid allocator for the simulation of complex 3-D device structures · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 1991
Electronic design automation
circuit simulation
0.051991
An adaptive grid refinement strategy for the drift-diffusion equations · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 1991
Extracting transistor changes from device simulations by gradient fitting · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 1989
A new discretization scheme for the semiconductor current continuity equations · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 1989
Audio and music processing › hearing aids
hearing aid signal processing
0.012006
Low-power architectural trade-offs in a VLSI implementation of an adaptive hearing aid algorithm · DAC 2006
Audio and music processing
speech enhancement
0.012006
Low-power architectural trade-offs in a VLSI implementation of an adaptive hearing aid algorithm · DAC 2006
Electronic design automation › technology computer-aided design › device simulation
monte carlo device simulation
0.011997
Optimized terminal current calculation for Monte Carlo device simulation · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 1997
Integrated circuit design › analog and mixed-signal circuits
device modeling
0.031989
Extracting transistor changes from device simulations by gradient fitting · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 1989
Transient Simulation of Silicon Devices and Circuits · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 1985
A Description of MOS Internodal Capacitances for Transient Simulations · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 1982
Electronic design automation › physical design
layout synthesis
0.011992
A module generator based on the PQ-tree algorithm · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 1992
Electronic design automation › physical design
module generation
0.011992
A module generator based on the PQ-tree algorithm · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 1992
Electronic design automation › physical design
placement
0.021992
An Analytic Optimization Technique for Placement of Macro-Cells · DAC 1989
A module generator based on the PQ-tree algorithm · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 1992
Integrated circuit design › semiconductor device modeling
drift-diffusion model
0.011991
An adaptive grid refinement strategy for the drift-diffusion equations · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 1991
High-performance computing › numerical linear algebra › linear solver
iterative linear solvers
0.011991
PILS: an iterative linear solver package for ill-conditioned systems · SC 1991
High-performance computing
large-scale simulation
0.011991
Three-dimensional numerical semiconductor device simulation: algorithms, architectures, results · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 1991
High-performance computing
numerical linear algebra
0.011991
PILS: an iterative linear solver package for ill-conditioned systems · SC 1991
Electronic design automation › technology computer-aided design › device simulation
three-dimensional device simulation
0.011991
Omega-an octree-based mixed element grid allocator for the simulation of complex 3-D device structures · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 1991
Electronic design automation › physical design › placement › module placement
macro placement
0.011989
An Analytic Optimization Technique for Placement of Macro-Cells · DAC 1989

Methods — techniques the papers use, named apart from their topics

wave function approach · 0.1mixed precision · 0.1load balancing · 0.1resource sharing · 0.1gate-level simulation · 0.1serpent · 0.1rijndael · 0.1variance minimization · 0.0ramo-shockley theorem · 0.0delaunay triangulation · 0.0rectangle-based meshing · 0.0ordering and coloring · 0.0finite-element discretization · 0.0error indicator · 0.0drift-diffusion equations · 0.0coupled and noncoupled schemes · 0.0hybrid finite-element discretization · 0.0newton iteration · 0.0
YearPublicationVenuePosition
2011 Atomistic nanoelectronic device engineering with sustained performances up to 1.44 PFlop/s
abstract
We present a multi-dimensional, atomistic, quantum transport simulation approach to investigate the performances of realistic nanoscale transistors for various geometries and material systems. The central computation consists in solving the Schrödinger equation with open boundary conditions several thousand times. To do that, a Wave Function approach is used since it can be relatively easily parallelized. To further improve the computational efficiency, three additional levels of parallelization are identified, the work load is optimally balanced between the CPUs, computational interleaving is applied where possible, and a mixed precision scheme is introduced. Using two different device types, a high electron mobility and a band-to-band tunneling transistor, sustained performances up to 1.28 PFlop/s in double precision (55% of the peak performance) and 1.44 PFlop/s in mixed precision are reached on 221,400 cores on the CRAY-XT5 Jaguar at Oak Ridge National Lab.
Mathieu Luisier, Timothy B. Boykin, Gerhard Klimeck, Wolfgang Fichtner
SC4
2011 A New Built-In Defect-Based Testing Technique to Achieve Zero Defects in the Automotive Environment
Vezio Malandruccolo, Mauro Ciappa, Hubert Rothleitner, Wolfgang Fichtner
J. Electron. Test.4
2010 Novel built-in methodology for defect testing of capacitor oxide in SAR analog to digital converters for critical automotive applications
abstract
Mixed signal components are increasingly used to implement controlling loops and digital signal processors in automotive applications. Successive Approximation Register (SAR) analog-to-digital converters based on switched capacitors play a major role in this evolution. Since the defectivity of the intermetal dielectric in vertical parallel plate capacitors is a major concern of this technology, dedicated screening techniques are needed to implement “zero defects” strategies. A novel screening approach is proposed in this paper, which is based on the use of a dedicated embedded circuitry with very low area consumption. A design is presented, which includes the control logic, the high voltage generation, and the leakage detection circuitry. The concept, advantages and the circuits for the proposed built-in reliability test are described in detail and illustrated by layout and circuit simulations.
Vezio Malandruccolo, Mauro Ciappa, Wolfgang Fichtner, Hubert Rothleitner
ETS3
2009 Hardware evaluation of the stream cipher-based hash functions RadioGatún and irRUPT
abstract
In the next years, new hash function candidates will replace the old MD5 and SHA-1 standards and the current SHA-2 family. The hash algorithms RadioGatun and irRUPT are potential successors based on a stream structure, which allows the achievement of high throughputs (particularly with long input messages) with minimal area occupation. In this paper, several hardware architectures of the two above mentioned hash algorithms have been investigated. The implementation on ASIC of RadioGatun with a word length of 64 bits shows a complexity of 46 k gate equivalents (GE) and reaches 5.7 Gbps throughput with a 3.64-bit input message. The same design approaches 120 Gbps on ASIC with long input messages (63.4 Gbps on a Virtex-4 FPGA with 2.9 kSlices). On the other hand, the irRUPT core turns out to be the most compact circuit (only 5.8 kGE on ASIC, and 370 Slices on FPGA) achieving 2.4 Gbps (with long input messages) on ASIC, and 1.1 Gbps on FPGA.
Luca Henzen, Flavio Carbognani, Norbert Felber, Wolfgang Fichtner
DATE4
2009 Novel Solution for the Built-in Gate Oxide Stress Test of LDMOS in Integrated Circuits for Automotive Applications
abstract
Efficient screening procedures for the control of the gate oxide defectivity are vital to limit early failures especially in critical automotive applications. Traditional strategies based on burn-in and in-line tests are able to provide the required level of reliability but they are expensive and time consuming. This paper presents a novel approach to the gate stress test of Lateral Diffused MOS transistors based on an embedded circuitry that includes logic control, high voltage generation, and leakage current monitoring. The concept, advantages and the circuit for the proposed built-in gate stress test procedure are described in very detail and illustrated by circuit simulation.
Vezio Malandruccolo, Mauro Ciappa, Wolfgang Fichtner, Hubert Rothleitner
ETS3
2009 VLSI Implementations of the Cryptographic Hash Functions MD6 and ïrRUPT
abstract
A public competition organized by the NIST recently started, with the aim of identifying a new standard for cryptographic hashing (SHA-3). Besides a high security level, candidate algorithms should show good performance on various platforms. While an average performance on high-end processors is generally not critical, implementability and flexibility in hardware is crucial, because the new standard will be implemented in a variety of lightweight devices. This paper investigates VLSI architectures of the SHA-3 candidates MD6 and irRUPT. The fastest circuit is the 16timesparallel MD6 core, reaching 16.3 Gbps at a complexity of 69.8 k gate equivalents (GE) on ASIC and 8.4 Gbps using 4465 Slices on FPGA. However, large memory requirements preclude the application of MD6 to resource-constrained systems. The most flexible and efficient circuit turns out to be our 2-irRUPT64times2-256/8 core, which achieves a throughput of 5.0 Gbps at 12.7 kGE on ASIC and 1.7 Gbps using 613 Slices on FPGA.
Luca Henzen, Flavio Carbognani, Jean-Philippe Aumasson, Sean O'Neil, Wolfgang Fichtner
ISCAS5
2009 Hardware Platform and Implementation of a Real-time Multi-user MIMO-OFDM Testbed
abstract
This paper describes a modular hardware platform of a multi-user (MU) multiple-input multiple-output (MIMO) orthogonal frequency division multiplexing (OFDM) testbed. The hardware platform is based on multiple field programmable gate arrays (FPGAs), provides four integrated radio-frequency (RF) chains, and has capabilities for extension boards. The performance and modularity of the testbed enables real-time MU-MIMO-OFDM experiments as well as offline processing experiments. To this end, the MIMO physical (PHY) layer of Haene et al., IEEE J-SAC, 2008, has been adapted to the new hardware platform and extended with bi-directional communication facilities and a basic media access control (MAC) layer equipped with Ethernet connectivity.
Markus Wenk, Peter Luethi, Patrick Maechler, Norbert Felber, Wolfgang Fichtner, Michael Lerjen
ISCAS6
2009 Live Demonstration: Hardware Platform and Implementation of a Real-time Multi-user MIMO-OFDM Testbed
abstract
The goal of the demonstration is to show the visitor how MIMO-OFDM communication works and what gains in terms of throughput, link reliability, etc. can be achieved, visualized by different experiments.
Markus Wenk, Peter Luethi, Patrick Maechler, Norbert Felber, Wolfgang Fichtner, Michael Lerjen
ISCAS6
2008 A Parallel Sparse Linear Solver for Nearest-Neighbor Tight-Binding Problems
Mathieu Luisier, Gerhard Klimeck, Andreas Schenk, Wolfgang Fichtner, Timothy B. Boykin
Euro-Par4
2008 Hardware-efficient steering matrix computation architecture for MIMO communication systems
abstract
Beamforming (BF) improves the error rate performance of multiple-input multiple-output (MIMO) wireless communication systems by spatial separation of the transmitted data streams. Spatial separation is achieved by multiplication of the transmit vector by a steering matrix, which is obtained through the singular value decomposition (SVD) of the channel matrix. In this paper, we describe a hardware-efficient VLSI architecture for steering matrix computation using a hardware- optimized SVD algorithm. Our architecture contains a high-speed Givens rotation unit which achieves high processing throughput at low area. The resulting VLSI implementation requires 3.3 mus per steering matrix computation at an expense of 41.3 kGEs and shows a 3.5-fold hardware-efficiency gain compared to a reference SVD implementation.
Christian Senning, Christoph Studer, Peter Luethi, Wolfgang Fichtner
ISCAS4
2008 VLSI architecture for data-reduced steering matrix feedback in MIMO systems
abstract
Beamforming (BF) for multiple-input multiple-output (MIMO) wireless communications systems can improve the error rate performance by spatial separation of the transmitted data streams. BF requires to feed back steering matrices from the receiver to the transmitter. The usually large amount of feedback data asks for data reduction schemes. In this paper, we investigate the error rate performance/feedback rate trade-off associated with steering matrix data-reduction schemes and present a corresponding hardware-optimized compression/decompression architecture. Our VLSI implementation achieves up to 50% data reduction for 4times4-dimensional steering matrices without a significant decrease in terms of error rate performance at a circuit complexity of only 7 k gate equivalents.
Christoph Studer, Peter Luethi, Wolfgang Fichtner
ISCAS3
2008 Transmission Gates Combined With Level-Restoring CMOS Gates Reduce Glitches in Low-Power Low-Frequency Multipliers
abstract
Various 16-bit multiplier architectures are compared in terms of dissipated energy, propagation delay, energy-delay product (EDP), and area occupation, in view of low-power low-voltage signal processing for low-frequency applications. A novel practical approach has been set up to investigate and graphically represent the mechanisms of glitch generation and propagation. It is found that spurious activity is a major cause of energy dissipation in multipliers. Measurements point out that, because of its shorter full-adder chains, the Wallace multiplier dissipates less energy than other traditional array multipliers (8.2 mu W/MHz versus 9.6 mu W/MHz for 0.18mum CMOS technology at 0.75 V). The benefits of transistor sizing are also evaluated (Wallace including minimum-size transistors dissipates 6.2 muW/MHz). By combining transmission gates with static CMOS in a Wallace architecture, a new approach is proposed to improve the energy-efficiency further (4.7 muW/MHz), beyond recently published low-power architectures. The innovation consists in suppressing glitches via resistance-capacitance low-pass filtering, while preserving unaltered driving capabilities. The reduced number ofVdd-to-ground paths also contributes to a significant decrease of static consumption.
Flavio Carbognani, Felix Bürgin, Norbert Felber, Hubert Kaeslin, Wolfgang Fichtner
IEEE Trans. Very Large Scale Integr. Syst.5
2007 Reduced-complexity mimo detector with close-to ml error rate performance
abstract
Maximum likelihood (ML) detection provides optimum error rate performance for uncoded multiple-input multiple-output (MIMO) systems. However, circuit complexity of a straightforward implementation of ML detection is uneconomic for high-rate systems. This paper addresses the VLSI implementation trade-offs of a MIMO detection algorithm that achieves close-to ML error rate performance with reduced computational complexity. The described implementations in a 0.25 μm CMOS technology for 4-4 MIMO systems feature a simple data-path, achieve high throughput, and use small silicon area. Important contributing factors to these results are efficient enumeration strategies and the application of simplified norms and sophisticated scheduling techniques together with a new low-complexity preprocessing scheme.
C. Hess, Markus Wenk, Andreas Peter Burg, Peter Luethi, Christoph Studer, Norbert Felber, Wolfgang Fichtner
ACM Great Lakes Symposium on VLSI7
2007 Regularized Frequency Domain Equalization Algorithm and its VLSI Implementation
abstract
Approximation of Toeplitz matrices with cir-culant matrices is a well-known approach to reduce the computational complexity of linear equalizers. This paper presents a novel technique to compute linear equalizer coefficients in the frequency domain. It is shown how a regularization term can help to reduce the error caused by the frequency domain approximation. A corresponding VLSI implementation provides reference for the true silicon complexity and for the complexity increase associated with the proposed algorithm.
Andreas Peter Burg, Simon Haene, Wolfgang Fichtner, Markus Rupp
ISCAS3
2007 FFT Processor for OFDM Channel Estimation
abstract
Pilot-assisted channel estimation for communication systems employing orthogonal frequency division multiplexing modulation requires significant signal processing at the receiver if the correlation among the frequency-domain channel coefficients is to be exploited in order to improve accuracy. In this work, a conventional FFT processor is extended to support all operations required by a selected channel estimation algorithm, so that both OFDM de/modulation and channel estimation can be efficiently performed on the same hardware unit. The silicon complexity of the extended processor, which was prototyped in a real-time testbed using FPGAs, is compared to a conventional FFT processor.
Simon Haene, Andreas Peter Burg, Peter Luethi, Norbert Felber, Wolfgang Fichtner
ISCAS5
2007 VLSI Implementation of a High-Speed Iterative Sorted MMSE QR Decomposition
abstract
The QR decomposition is an important, but often underestimated prerequisite for pseudo- or non-linear detection methods such as successive interference cancellation or sphere decoding for multiple-input multiple-output (MIMO) systems. The ability of concurrent iterative sorting during the QR decomposition introduces a moderate overall latency, but provides the base for an improved layered stream decoding. This paper describes the architecture and results of the first VLSI implementation of an iterative sorted QR decomposition preprocessor for MIMO receivers. The presented architecture performs MIMO channel preprocessing using Givens rotations in order to compute the minimum mean squared error QR decomposition
Peter Luethi, Andreas Peter Burg, Simon Haene, David Perels, Norbert Felber, Wolfgang Fichtner
ISCAS6
2007 Implementation of a Low-Complexity Frame-Start Detection Algorithm for MIMO Systems
abstract
Multiple-input multiple-output (MIMO) communication systems require well-designed synchronization schemes in the receiver to meet stringent QoS requirements. In particular, OFDM modulation is very sensitive to timing synchronization errors, which cause inter-symbol interference. This paper describes a frame-start detection algorithm, which relies on received signal power increase and does not require any special properties of the transmitted signal. The performance is analyzed and then, verified through simulations in a MIMO system employing orthogonal frequency division multiplexing. Finally, a low-complexity FPGA implementation of the presented algorithm is described in detail.
David Perels, Christoph Studer, Wolfgang Fichtner
ISCAS3
2007 Wireless Implant Communications for Biomedical Monitoring Sensor Network
abstract
Galvanic coupling provides a novel data transmission between sensor units for low frequency intra-body communication with electrodes attached to the human skin. In this work, that approach has been adapted to implantable miniaturized pills. A communication system for wireless data transmission in muscle tissue is developed capable of transmitting data on four channels concurrently with a throughput of 4.8kbit/s. The main focus is the future implantability of such a miniaturized system for medical long term surveillance of patients. To achieve this goal, circuit size, low power consumption and electrical safety have to be carefully considered. The implemented frequency division multiple access (FDMA) system works in the frequency range between 100 kHz and 250 kHz. Tests have been processed with 4 transmitters units. The system architecture features the possibility of migration into a single system-on-chip.
Marc Simon Wegmueller, Martin Hediger, Felix Bürgin, Wolfgang Fichtner
ISCAS5
2006 Low-power architectural trade-offs in a VLSI implementation of an adaptive hearing aid algorithm
abstract
This paper analyzes the power-area trade-off of functionally equivalent architectural implementations of a speech enhancement algorithm for hearing aids. Gate-level simulations and measurements show that an optimum degree of resource sharing (0.60 mW in a 0.25 /spl mu/m CMOS process) is more energy-efficient than both the fully time-multiplexed (1.42 mW) and the isomorphic architecture (1.54 mW), without overly large area overhead (0.77 mm/sup 2/ against 0.43 mm/sup 2/ and 4.31 mm/sup 2/, respectively).
Felix Bürgin, Flavio Carbognani, Martin Hediger, Hektor Meier, Robert Meyer-Piening, Rafael Santschi, Hubert Kaeslin, Norbert Felber, Wolfgang Fichtner
DAC9
2006 Two-phase resonant clocking for ultra-low-power hearing aid applications
abstract
Resonant clocking holds the promise of trading speed for energy in CMOS circuits that can afford to operate at low frequency, like hearing aids. An experimental chip with 110k transistors and more than 2500 latches, has been designed, fabricated and tested. The measured energy con sumption of the design at 0.8 V is 62μW/MHz, about 7.5% less than the conventional single-edge-triggered benchmark. Closer analysis reveals that much of the energy savings brought about by resonant clocking at low supply voltages are lost when a CMOS circuit is operated at higher voltages. This is because of the crossover currents that persist for much of a clock period when a circuit is driven from sine-type clock waveform.
Flavio Carbognani, Felix Bürgin, Norbert Felber, Hubert Kaeslin, Wolfgang Fichtner
DATE5
2006 A Frame-Start Detector for a 4×4 MIMO-OFDM System
abstract
Future wireless LANs will increase the peak data rate by employing multiple antennas at both transmitter and receiver. Well designed synchronization algorithms are a prerequisite for meeting stringent QoS requirements. In particular OFDM modulation, which constitutes the basics for WLAN, is very sensitive to timing synchronization errors which incur inter-symbol interference. In this paper, a novel frame synchronization algorithm is proposed that is implemented in the FPGA of a real-time MIMO-OFDM testbed. Simulations show it to be of sufficient performance in scenarios of interest, while the hardware complexity is suitable for an FPGA implementation. Additionally, the algorithm exhibits a good resilience against narrow-band interference, which causes problems in traditional frame-start detection algorithms
David Perels, Simon Haene, Andreas Peter Burg, Peter Luethi, Norbert Felber, Wolfgang Fichtner
ICASSP (4)6
2006 Algorithm and VLSI architecture for linear MMSE detection in MIMO-OFDM systems
abstract
The paper describes an algorithm and a corresponding VLSI architecture for the implementation of linear MMSE detection in packet-based MIMO-OFDM communication systems. The advantages of the presented receiver architecture are low latency, high-throughput, and efficient resource utilization, since the hardware required for the computation of the MMSE estimators is reused for the detection. The algorithm also supports the extraction of soft information for channel decoding
Andreas Peter Burg, Simon Haene, David Perels, Peter Luethi, Norbert Felber, Wolfgang Fichtner
ISCAS6
2006 42% power savings through glitch-reducing clocking strategy in a hearing aid application
abstract
Glitches are responsible for a significant proportion of overall power dissipation in digital signal processing circuits. Activity-reduction techniques that involve an optimized clocking strategy have been applied to a front-end block in a DSP adaptive directional microphone for hearing aids. Functionally equivalent implementations, differing only in their clocking scheme, have been integrated on silicon in a 0.25 mum CMOS technology. Measurements and post-layout simulations confirm a 42% reduction over single-edge-triggered clocking with clock gating. An overall power dissipation of 20 muW (@ 1.4 V, 374 kHz) has been measured. This achievement has been made possible by combining two novel techniques: a multi-stage clock gating, and a symmetric two-phase level-sensitive clocking with glitch-aware re-distribution of data-path registers
Flavio Carbognani, Felix Bürgin, Norbert Felber, Hubert Kaeslin, Wolfgang Fichtner
ISCAS5
2006 Silicon implementation of an MMSE-based soft demapper for MIMO-BICM
abstract
The performance of systems employing bit-interleaved coded modulation (BICM) critically depends on the availability of soft information. In the multi-antenna case, the extraction of optimum bit-metrics becomes prohibitively complex, so that suboptimal solutions need to be adopted for practical implementation. Instead of considering all the spatially multiplexed streams jointly, the implementation presented in this paper computes the soft information on each data stream separately, based on the output of an MMSE equalizer
Simon Haene, Andreas Peter Burg, David Perels, Peter Luethi, Norbert Felber, Wolfgang Fichtner
ISCAS6
2006 K-best MIMO detection VLSI architectures achieving up to 424 Mbps
abstract
From an error rate performance perspective, maximum likelihood (ML) detection is the preferred detection method for multiple-input multiple-output (MIMO) communication systems. However, for high transmission rates a straight forward exhaustive search implementation suffers from prohibitive complexity. The K-best algorithm provides close-to-ML bit error rate (BER) performance, while its circuit complexity is reduced compared to an exhaustive search. In this paper, a new VLSI architecture for the implementation of the K-best algorithm is presented. Instead of the mostly sequential processing that has been applied in previous VLSI implementations of the algorithm, the presented solution takes a more parallel approach. Furthermore, the application of a simplified norm is discussed. The implementation in an ASIC achieves up to 424 Mbps throughput with an area that is almost on par with current state-of-the-art implementations
Markus Wenk, Martin Zellweger, Andreas Peter Burg, Norbert Felber, Wolfgang Fichtner
ISCAS5
2004 A 2 Gb/s balanced AES crypto-chip implementation
abstract
We present a balanced 2 Gb/s en-/decryption ASIC realization of the AES algorithm that supports all standard operation modes and key lengths. Rather than optimizing only for throughput, special care is taken to balance the more involved decryption path with that of the encryption path using a number of high-level architectural and register transfer level optimizations. The fabricated en-/decryption core requires an active area of only 3.56 mm2 (less than 120,000 gate equivalents) in a modest 0.25 µm CMOS technology.
Frank K. Gürkaynak, Andreas Peter Burg, Norbert Felber, Wolfgang Fichtner, D. Gasser, Franco Hug, Hubert Kaeslin
ACM Great Lakes Symposium on VLSI4
2002 2Gbit/s Hardware Realizations of RIJNDAEL and SERPENT: A Comparative Analysis
Adrian K. Lutz, Jürg Treichler, Frank K. Gürkaynak, Hubert Kaeslin, Gérard Basler, Antonia Albani, Stephan Reichmuth, Pieter Rommens, Stephan Oetiker, Wolfgang Fichtner
CHES10
2001 PARDISO: a high-performance serial and parallel sparse linear solver in semiconductor device simulation
Olaf Schenk, Klaus Gärtner, Wolfgang Fichtner, Andreas Stricker
Future Gener. Comput. Syst.3
1999 Functional verification of intellectual properties (IP): a simulation-based solution for an application-specific instruction-set processor
abstract
Scalability and customization properties of IP modules demand for new approaches in functional verification. We present a novel simulation-based solution for an Application-specific Instruction-set Processor (ASIP). Existing assembler code preselected by IP-configurable constraints forms the verification data base (reference stimuli). A behavioral "golden model" of the IP is used to derive expected responses suitable for any possible configuration of the final ASIP (RTL) implementation. Cycle-based verification is performed by stimulating the RTL model with the assembled reference stimuli and by comparing the outputs (actual responses) against the expected responses. Primary input stimulation is accomplished by reading back interface data prior written to a memory (model) under control of the reference stimuli. The synchronization of the configaration-dependent actual responses to the non-cycle-related expected responses is achieved by a mechanism based on "interface-specific activity scheduling", which further more reduces the number of vectors efficiently, resulting in a significant simulation speed-up.
Manfred Stadler, Hubert Kaeslin, Norbert Felber, Wolfgang Fichtner, Markus Thalmann
ITC5
1997 Optimized terminal current calculation for Monte Carlo device simulation
abstract
We present a generalized Ramo-Shockley theorem (GRST) for the calculation of time-dependent terminal currents in multidimensional charge transport calculations and simulations. While analytically equivalent to existing boundary integration methods, this new domain integration technique is less sensitive to numerical error introduced by calculations of finite precision. Most significantly, we derive entirely new optimized formulas for the ensemble Monte Carlo estimation of steady-state terminal currents from the time-independent form of our GRST, which are in general not equivalent to the time-average of the true time-dependent terminal currents. We then demonstrate, both analytically and by means of example, how our new variance-minimizing terminal current estimators may be exploited to improve estimator accuracy in comparison to existing methods.
P. Douglas Yoder, Klaus Gärtner, Ulrich Krumbein, Wolfgang Fichtner
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
1993 VINCI: Secure Test of a VLSI High-Speed Encryption System
abstract
Designers of VLSI circuits for cryptographic applications are faced with a severe dilemma: Concerns of security, requiring encapsulation of cryptographic applications, and testability, requiring access to all hardware subunits to check for their correct functionality, are basically contradictory. A further security requirement is the immediate detection of failure inside the encryption component. In this paper, an approach based on a compound system test strategy is introduced that reconciles the demands of security and testability. This system test strategy based upon built-in self-test on all levels of implementation allows off-line and concurrent checking without opening access to sensitive regions via test structures. The realization of the system test scheme is a new VLSI cipher implementation, VINCI, that fulfills all security demands for immediate failure detection and supports higher level system test strategies.>
Heinz Bonnenberg, Andreas Curiger, Norbert Felber, Hubert Kaeslin, Reto Zimmermann, Wolfgang Fichtner
ITC6
1993 Mixed element trees: a generalization of modified octrees for the generation of meshes for the simulation of complex 3-D semiconductor device structures
abstract
This paper addresses the problem of the allocation of spatial grids for complex nonplanar three-dimensional (3-D) semiconductor device structures. We have characterized the class of meshes suitable for the integration of the device equations with the usual numerical schemes as being a subclass of the class of Delaunay meshes. We propose an algorithm for the efficient generation of such admissible meshes based on the iterative refinement of coarse elements. The generated meshes permit an exact geometrical modeling of rather general domain boundaries of modern silicon devices avoiding the "obtuse angle problem" by construction.>
Nancy Hitschfeld-Kahler, Paolo Conti, Wolfgang Fichtner
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
1992 Lazy-expansion symbolic expression approximation in SYNAP
abstract
A lazy-expansion technique for generating small approximate symbolic analog circuit analysis expressions is described. Statistics for this technique as implemented in the symbolic analysis program SYNAP are presented and show a two order-of-magnitude speed improvement (on larger circuits) as compared with traditional full-expansion techniques. This technique also allows larger circuits to be analyzed. Methods used in SYNAP for eliminating pole and zero movement at the design point and for handling variables representing device mismatches are also presented.>
Steven J. Seda, Marc G. R. Degrauwe, Wolfgang Fichtner
ICCAD3
1992 A module generator based on the PQ-tree algorithm
abstract
GRAPES, a system for module generation that produces faster and denser layout than conventional CAD tools, is described. The layout style produced by GRAPES resembles standard cell layout (vertical polysilicon wires crossing horizontal diffusion stripes) or sea-of-gates macro layout. Contrary to standard cells or macros, the complexity of the individual cells is not limited by any library, cells can be stretched, feedthroughs can run across them, and the transistors can be permuted and sized individually. Cells and rows are produced at the same time in a top-down procedure. A novel algorithm based on the PQ-tree algorithm orders the individual transistor gates in each row such the several rows abut, with no routing channels between the rows. The most critical nets can usually be connected by abutment, even when they connect several cells in different rows. In a comparison with a commercial standard-cell tool, GRAPES produced layout that is about two times smaller and has up to six times shorter total wire length.>
Hans-Rudolf Heeb, Wolfgang Fichtner
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
1992 Automatic rectangle-based adaptive mesh generation without obtuse angles
abstract
Mesh generators must become more powerful to keep up with the growing sophistication of device simulators. The mesh generator MESHBUILD, which produces meshes with no obtuse angles for structures with reasonably complex geometries, is described. The overall number of mesh elements is reduced compared with similar mesh generators. MESHBUILD is suitable for use in automatic grid generation and includes an interactive graphics interface. Run times for complex examples are on the order of 1 min on a Sun-4 workstation.>
Stephan Müller 0003, Kevin Kells, Wolfgang Fichtner
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
1991 PILS: an iterative linear solver package for ill-conditioned systems
Claude Pommerell, Wolfgang Fichtner
SC2
1991 Transistor sizing for large combinational digital CMOS circuits
Lucas S. Heusler, Wolfgang Fichtner
Integr.2
1991 An adaptive grid refinement strategy for the drift-diffusion equations
abstract
A method of computing the error in the solution of the semiconductor current continuity equations as well as the error in the terminal currents is proposed. An appropriate error indicator is developed based on a divergence free upwinding (finite element) discretization. The grid used for discretization is adapted to the error in the solution by dynamically adding or removing grid points in order to improve the solution and thus the terminal currents. The examples indicate that it is sufficient only for a reverse-biased p-n junction to refine the grid according to the error in the Poisson equation. In the forward biased case, it is necessary to take into account the error in the current continuity equation in order to guarantee exact terminal currents.>
Josef F. Burgler, William M. Coughran Jr., Wolfgang Fichtner
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
1991 Omega-an octree-based mixed element grid allocator for the simulation of complex 3-D device structures
abstract
The authors discuss an automatic mesh generator, Omega , developed as a front-end for the simulation of complex three-dimensional semiconductor devices. Grids generated with Omega exhibit smooth transitions from dense to coarse grid regions and a proper description of irregular geometries such as nonuniform surfaces and interfaces. In addition to the grid itself, Omega provides the dual lattice needed for the integration of the device equations. The underlying algorithm avoids the obtuse angle problem. It is shown how this problem can be formalized in 3-D, how badly shaped elements can be detected, and how they can be avoided by construction. Examples of simulations of complex 3-D devices, with grids with several tens of thousands of mesh points, demonstrate the capabilities of Omega .>
Paolo Conti, Nancy Hitschfeld-Kahler, Wolfgang Fichtner
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
1991 Three-dimensional numerical semiconductor device simulation: algorithms, architectures, results
abstract
The authors present SECOND, a program for large-scale semiconductor device simulation with truly three-dimensional grids. Since 3-D simulations necessitate large computing resources, the choice of algorithms and their implementation become of utmost importance. The authors investigated the most commonly used numerical algorithms for the solution of the classical drift-diffusion equations. The study included coupled and noncoupled point and block schemes, direct and preconditioned iterative linear solvers, and several distinct ordering and coloring techniques. Structures with regular and irregular grids were analyzed. These algorithms were compared on a variety of machines including workstations, minisupers, and supercomputers. Results of transient simulations are presented to illustrate the approach.>
Gernot Heiser, Claude Pommerell, Jürgen Weis, Wolfgang Fichtner
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
1989 An Analytic Optimization Technique for Placement of Macro-Cells
abstract
This paper discusses the placement of R (rectangular) or L-shaped blocks, such as in macro-cell design. We present a new analytic approach that simultaneously accounts for different criteria (wire length, chip area). We characterize a solution by the associated adjacency graph [1] and derive an optimization technique with an active set method [2]. Results on a test problem are given to prove the performance and feasibility of this technique.
Alexander Herrigel, Wolfgang Fichtner
DAC2
1989 A global floorplanning technique for VLSI layout
abstract
The floorplanning of rectangular cells is discussed. A new global approach that simultaneously accounts for different design goals is presented. A key aspect of this approach is a more general slicing structure representation of the floorplan that is not restricted to a special case of rectangular dissection and a two-dimensional partitioning procedure. A new model for the prediction of the associated shape functions is presented and an analytic optimization technique for the pin allocation is described.>
Alexander Herrigel, M. Glaser, Wolfgang Fichtner
ICCD3
1989 Macrocell-level compaction with automatic jog introduction
abstract
A novel algorithm for compacting a VLSI chip on the macrocell level is presented. Compared to previous algorithms, the technique can handle larger designs, produces higher-quality output, and reduces designer intervention as much as possible. Jogs are automatically introduced in the connecting wires to achieve the needed flexibility for placing cells into optimal positions.>
Alexander Herrigel, J. Kamm, Wolfgang Fichtner
ICCD3
1989 A new discretization scheme for the semiconductor current continuity equations
abstract
A hybrid finite-element method to discretize the continuity equation in semiconductor device simulation is given. Within each element of a finite element discretization, the current is uniquely determined by nodal values of the density and the potential. The authors use the integrability condition for a system of partial differential equations to obtain the equations that determine the current within the element. They then satisfy the continuity in the current flow across interelement boundaries in a weak sense. They have found that the method works in any dimension and for (d-dimensional) simplexes as well as for quadrilaterals, bricks, prisms, and so on, although they have no proof that it will not break down in particular cases.>
Josef F. Burgler, Randolph E. Bank, Wolfgang Fichtner, R. Kent Smith
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
1989 Extracting transistor changes from device simulations by gradient fitting
abstract
The results of small-signal or transient analyses from a conventional device simulator (or measured data) can be combined with gradient-fitting techniques to produce smooth spline-based MOSFET charge models for circuit simulation. These techniques are also applicable to shape-from-shading problems. Results based on simulations of some small devices are presented. The comparative efficiencies of the small-signal and transient approach are discussed as well as the relation between the so-called small- and large-signal charges. The role of spline-based table models as against compact analytical models is considered.>
William M. Coughran Jr., Wolfgang Fichtner, Eric H. Grosse
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
1988 A symbolic analysis tool for analog circuit design automation
abstract
An analysis tool has been developed to generate symbolic design equations for analog circuits. This tool (SYNAP) works in conjunction with a symbolic mathematics program (MACSYMA) to create both exact and simplified analytic expressions needed for circuit design and forms the cornerstone of a non-fixed-topology analog circuit design system. SYNAP performs DC, AC, noise, and offset analyses for time-invariant analog circuits with one stable operating point and generates code for the new design system.>
Steven J. Seda, Marc G. R. Degrauwe, Wolfgang Fichtner
ICCAD3
1985 Transient Simulation of Silicon Devices and Circuits
abstract
In this paper, we present an overview of the physical principles and numerical methods used to solve the coupled system of nonlinear partial differential equations that model the transient behavior of silicon VLSI device structures. We also describe how the same techniques are applicable to circuit simulation. A composite linear multistep formula is introduced as the time-integration scheme. Newton-iterative methods are exploited to solve the nonlinear equations that arise at each time step. We also present a simple data structure for nonsymmetric matrices with symmetric nonzero structures that facilitates iterative or direct methods with substantial efficiency gains over other storage schemes. Several computational examples, including a CMOS latchup problem, are presented and discussed.
Randolph E. Bank, William M. Coughran Jr., Wolfgang Fichtner, Eric H. Grosse, Donald J. Rose, R. Kent Smith
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
1982 A Description of MOS Internodal Capacitances for Transient Simulations
abstract
Charge versus voltage and internodal capacitance versus voltage characteristics are calculated for a short-channel MOSFET using a unified model of the dc device behavior. Velocity saturation is an important feature in the results. The importance of charge and capacitance calculations is assessed using a high speed MOS transient simulation. Device current and gate charge are determined to be the important ingredients for accurate simulation.
Geoffrey W. Taylor, Wolfgang Fichtner, J. G. Simmons
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2