VLDB 2026 Research / reviewers in the wild / expert
Heinrich Meyr
dblp:72/1999
· DBLP profile ↗
151ranked-venue papers
8as first author
0since 2021 · last 2020
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 62Computer networks · 48 · 7 first-authorSoftware engineering, systems software and programming languages · 19Graphics, computer vision, multimedia, augmented reality and games · 15 · 1 first-authorTheory of computation · 8Applied, interdisciplinary, general and emerging computing · 4Security and privacy · 2Artificial intelligence and machine learning · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer networks
36 papers |
Physical-layer communications · 97% Network measurement and analytics · 1% Internet of things and sensor networks · 1% | |
| Computer architecture, parallel and distributed computing, and storage systems
21 papers |
Electronic design automation · 37% Performance modeling and evaluation · 28% Processor architecture and microarchitecture · 23% | |
| Theoretical computer science
5 papers |
Coding theory · 96% Information theory · 4% | |
| Artificial intelligence
1 paper |
Legged, aerial and field robots · 61% Robot navigation and mapping · 39% | |
| Software engineering, system software, and programming languages
4 papers |
Compilers and program optimization · 98% Program analysis · 2% |
Topics — the 30 heaviest of 124, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Physical-layer communications › fading channels
rayleigh fading |
0.5 | 5 | 2014 | Oversampling Increases the Pre-Log of Noncoherent Rayleigh Fading Channels · IEEE Trans. Inf. Theory 2014 On the Achievable Rate of Stationary Rayleigh Flat-Fading Channels With Gaussian Inputs · IEEE Trans. Inf. Theory 2013 On the Gain of Joint Processing of Pilot and Data Symbols in Stationary Rayleigh Fading Channels · IEEE Trans. Inf. Theory 2012 |
Robotics › Legged, aerial and field robots
aerial robots |
0.4 | 1 | 2020 | Observability Analysis of Flight State Estimation for UAVs and Experimental Validation · ICRA 2020 |
Robotics › Robot navigation and mapping › state estimation
observability analysis |
0.4 | 1 | 2020 | Observability Analysis of Flight State Estimation for UAVs and Experimental Validation · ICRA 2020 |
Robotics › Legged, aerial and field robots › aerial robots
UAV state estimation |
0.4 | 1 | 2020 | Observability Analysis of Flight State Estimation for UAVs and Experimental Validation · ICRA 2020 |
Coding theory › error-correcting codes › decoding › iterative decoding › soft-input soft-output decoding
turbo decoding |
0.4 | 2 | 2017 | On the Performance Gap Between ML and Iterative Decoding of Finite-Length Turbo-Coded BICM in MIMO Systems · IEEE Trans. Commun. 2017 A Systematic Framework for Iterative Maximum Likelihood Receiver Design · IEEE Trans. Commun. 2010 |
Physical-layer communications
channel estimation |
0.4 | 9 | 2013 | On the Gain of Joint Processing of Pilot and Data Symbols in Stationary Rayleigh Fading Channels · IEEE Trans. Inf. Theory 2012 On the Achievable Rate of Stationary Rayleigh Flat-Fading Channels With Gaussian Inputs · IEEE Trans. Inf. Theory 2013 Optimum receiver design for OFDM-based broadband transmission .II. A case study · IEEE Trans. Commun. 2001 |
Physical-layer communications › information theory › capacity analysis
channel capacity |
0.4 | 2 | 2014 | Oversampling Increases the Pre-Log of Noncoherent Rayleigh Fading Channels · IEEE Trans. Inf. Theory 2014 On the Achievable Rate of Stationary Rayleigh Flat-Fading Channels With Gaussian Inputs · IEEE Trans. Inf. Theory 2013 |
Physical-layer communications › information theory
achievable rate |
0.3 | 2 | 2013 | On the Achievable Rate of Stationary Rayleigh Flat-Fading Channels With Gaussian Inputs · IEEE Trans. Inf. Theory 2013 On the Gain of Joint Processing of Pilot and Data Symbols in Stationary Rayleigh Fading Channels · IEEE Trans. Inf. Theory 2012 |
Coding theory › error-correcting codes › coded modulation
bit-interleaved coded modulation |
0.3 | 1 | 2017 | On the Performance Gap Between ML and Iterative Decoding of Finite-Length Turbo-Coded BICM in MIMO Systems · IEEE Trans. Commun. 2017 |
Coding theory › error-correcting codes › decoding › decoding algorithms › optimal decoding
maximum-likelihood decoding |
0.3 | 1 | 2017 | On the Performance Gap Between ML and Iterative Decoding of Finite-Length Turbo-Coded BICM in MIMO Systems · IEEE Trans. Commun. 2017 |
Coding theory › channel coding
turbo codes |
0.3 | 1 | 2017 | On the Performance Gap Between ML and Iterative Decoding of Finite-Length Turbo-Coded BICM in MIMO Systems · IEEE Trans. Commun. 2017 |
Physical-layer communications
channel coding |
0.3 | 4 | 2012 | Asymptotic Coded BER Analysis for MIMO BICM-ID with Quantized Extrinsic LLR · IEEE Trans. Commun. 2012 On Complexity, Energy- and Implementation-Efficiency of Channel Decoders · IEEE Trans. Commun. 2011 High-Rate Viterbi Processor: A Systolic Array Solution · IEEE J. Sel. Areas Commun. 1990 |
Physical-layer communications
MIMO |
0.3 | 3 | 2017 | Asymptotic Coded BER Analysis for MIMO BICM-ID with Quantized Extrinsic LLR · IEEE Trans. Commun. 2012 On the Performance Gap Between ML and Iterative Decoding of Finite-Length Turbo-Coded BICM in MIMO Systems · IEEE Trans. Commun. 2017 On the Gain of Joint Processing of Pilot and Data Symbols in Stationary Rayleigh Fading Channels · IEEE Trans. Inf. Theory 2012 |
Physical-layer communications
signal processing for communications |
0.3 | 5 | 2010 | Systematic Design of Iterative ML Receivers for Flat Fading Channels · IEEE Trans. Commun. 2010 A Systematic Framework for Iterative Maximum Likelihood Receiver Design · IEEE Trans. Commun. 2010 Channel tracking for RAKE receivers in closely spaced multipath environments · IEEE J. Sel. Areas Commun. 2001 |
Coding theory › error-correcting codes › decoding
iterative decoding |
0.2 | 2 | 2010 | Systematic Design of Iterative ML Receivers for Flat Fading Channels · IEEE Trans. Commun. 2010 A Systematic Framework for Iterative Maximum Likelihood Receiver Design · IEEE Trans. Commun. 2010 |
Processor architecture and microarchitecture › instruction set architecture
application-specific instruction-set processor |
0.2 | 4 | 2004 | Heterogeneous MP-SoC: the solution to energy-efficient signal processing · DAC 2004 A novel approach for flexible and consistent ADL-driven ASIP design · DAC 2004 Instruction encoding synthesis for architecture exploration using hierarchical processor models · DAC 2003 |
Performance modeling and evaluation
simulation |
0.2 | 4 | 2008 | Multiprocessor performance estimation using hybrid simulation · DAC 2008 A universal technique for fast and flexible instruction-set architecture simulation · DAC 2002 Fast Bit-True Simulation · DAC 2001 |
Electronic design automation
design space exploration |
0.1 | 5 | 2008 | Fine-grained application source code profiling for ASIP design · DAC 2005 A universal technique for fast and flexible instruction-set architecture simulation · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2004 Multiprocessor performance estimation using hybrid simulation · DAC 2008 |
Physical-layer communications › modulation › coded modulation
bit-interleaved coded modulation |
0.1 | 1 | 2012 | Asymptotic Coded BER Analysis for MIMO BICM-ID with Quantized Extrinsic LLR · IEEE Trans. Commun. 2012 |
Physical-layer communications › channel coding › error control coding › decoding
soft-output decoding |
0.1 | 1 | 2012 | Asymptotic Coded BER Analysis for MIMO BICM-ID with Quantized Extrinsic LLR · IEEE Trans. Commun. 2012 |
Robotics › Robot navigation and mapping
state estimation |
0.1 | 1 | 2020 | Observability Analysis of Flight State Estimation for UAVs and Experimental Validation · ICRA 2020 |
Physical-layer communications › channel coding › error control coding › concatenated codes
turbo codes |
0.1 | 1 | 2011 | On Complexity, Energy- and Implementation-Efficiency of Channel Decoders · IEEE Trans. Commun. 2011 |
Electronic design automation
high-level synthesis |
0.1 | 5 | 2003 | Instruction encoding synthesis for architecture exploration using hierarchical processor models · DAC 2003 Efficient building block based RTL code generation from synchronous data flow graphs · DAC 2000 System Level Fixed-Point Design Based on an Interpolative Approach · DAC 1997 |
Physical-layer communications › receiver design › optimum receiver
maximum-likelihood receiver |
0.1 | 1 | 2010 | Systematic Design of Iterative ML Receivers for Flat Fading Channels · IEEE Trans. Commun. 2010 |
Physical-layer communications › signal detection › sequence estimation
maximum-likelihood sequence estimation |
0.1 | 1 | 2010 | A Systematic Framework for Iterative Maximum Likelihood Receiver Design · IEEE Trans. Commun. 2010 |
Coding theory › error-correcting codes › coded modulation › bit-interleaved coded modulation
bit-interleaved coded modulation with iterative decoding |
0.1 | 1 | 2010 | Systematic Design of Iterative ML Receivers for Flat Fading Channels · IEEE Trans. Commun. 2010 |
Processor architecture and microarchitecture
chip multiprocessor |
0.1 | 2 | 2008 | MAPS: an integrated framework for MPSoC application parallelization · DAC 2008 Multiprocessor performance estimation using hybrid simulation · DAC 2008 |
Processor architecture and microarchitecture
instruction set architecture |
0.1 | 3 | 2003 | Instruction encoding synthesis for architecture exploration using hierarchical processor models · DAC 2003 A universal technique for fast and flexible instruction-set architecture simulation · DAC 2002 LISA - Machine Description Language for Cycle-Accurate Models of Programmable DSP Architectures · DAC 1999 |
Electronic design automation › hardware simulation
compiled simulation |
0.1 | 3 | 2004 | A universal technique for fast and flexible instruction-set architecture simulation · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2004 A universal technique for fast and flexible instruction-set architecture simulation · DAC 2002 Compiled HW/SW Co-Simulation · DAC 1996 |
Physical-layer communications
synchronization |
0.1 | 13 | 2001 | Optimum receiver design for OFDM-based broadband transmission .II. A case study · IEEE Trans. Commun. 2001 Optimum receiver design for wireless broad-band systems using OFDM. I · IEEE Trans. Commun. 1999 On sampling rate, analog prefiltering, and sufficient statistics for digital receivers · IEEE Trans. Commun. 1994 |
Methods — techniques the papers use, named apart from their topics
simulation · 0.6performance gap analysis · 0.6singular value decomposition · 0.4extended kalman filter · 0.4information-theoretic bounds · 0.3code generation · 0.2architecture exploration · 0.2information-theoretic analysis · 0.2high-SNR asymptotics · 0.2channel prediction · 0.2moment generating function · 0.1channel estimation · 0.1asymptotic BER analysis · 0.1PDF/PMF derivation · 0.1joint maximum likelihood optimization · 0.1fixed-point iteration · 0.1critical point equations · 0.1trace-driven replay · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2020 | Observability Analysis of Flight State Estimation for UAVs and Experimental ValidationabstractUAVs require reliable, cost-efficient onboard flight state estimation that achieves high accuracy and robustness to perturbation. We analyze a multi-sensor extended Kalman filter (EKF) based on the work by Leutenegger. The EKF uses measurements from a MEMS-based inertial system, static and dynamic pressure sensors as well as GPS. As opposed to other implementations we do not use a magnetic sensor because the weak magnetic field of the earth is subject to disturbances. Observability of the state is a necessary condition for the EKF to work. In this paper, we demonstrate that the system state is observable - which is in contrast to statements in the literature - if the random nature of the air mass is taken into account. Therefore, we carry out an in-depth observability analysis based on a singular value decomposition (SVD). The numerical SVD delivers a wealth of information regarding the observable (sub)spaces. We validated the theoretical findings based on sensor data recorded in test flights on a glider. Most importantly, we demonstrate that the EKF works. It is capable of absorbing large perturbations in the wind state variable converging to the undisturbed estimates. Heinrich Meyr, Meik Dörpinghaus, Gerhard P. Fettweis |
ICRA | 2 |
| 2017 | An information theoretic analysis of sequential decision-makingabstractWe provide a novel analysis of Wald's sequential probability ratio test based on information theoretic measures for symmetric thresholds, symmetric noise, and equally likely hypotheses. This test is optimal in the sense that it yields the minimum mean decision time. To analyze the decision-making process we consider information densities, which represent the stochastic information content of the observations yielding a stochastic termination time of the test. Based on this, we show that the conditional probability to decide for hypothesis H1(or the counter-hypothesis H0) given that the test terminates at time instant k is independent of time k. An analogous property has been found for a continuous-time first passage problem with two absorbing boundaries in the contexts of non-equilibrium statistical physics and communication theory. Moreover, we study the evolution of the mutual information between the binary variable to be tested and the output of the Wald test. Notably, we show that the decision time of the Wald test contains no information on which hypothesis is true beyond the decision outcome. Meik Dörpinghaus, Édgar Roldán, Izaak Neri, Heinrich Meyr, Frank Jülicher |
ISIT | 4 |
| 2017 | On the Performance Gap Between ML and Iterative Decoding of Finite-Length Turbo-Coded BICM in MIMO SystemsabstractAs real-time applications typically require short length code words, this paper analyzes the minimum achievable code word error rate of the given finite-length code words, namely, maximum likelihood (ML) decoding error probability. For coding schemes ML decoding is too complex, the key contribution of this paper is to analytically assess the performance gap between ML decoding and a practical decoding scheme. We analyze the combination of turbo codes and bit-interleaved coded-modulation that is a spectrally efficient coding scheme adopted in 3GPP long term evolution (LTE). In single-input single-output (SISO) systems, it was shown in the literature that turbo decoding delivers near-ML decoding performance. In this paper, we extend the analysis to a multi-input multi-output (MIMO) system. In contrast to the SISO case, the turbo principle-based iterative decoding scheme is subject to appreciable performance loss compared with the ML decoding, even in good MIMO channel conditions. For this reason, we further analyze potential reasons and examine possible improvements. By means of simulation, it is shown that convergence to a non-ML code word rather than non-convergence is the key reason for the observed performance loss in good channel conditions. Dan Zhang 0003, Heinrich Meyr |
IEEE Trans. Commun. | 2 |
| 2015 | Channel-aware local search (CA-LS) for iterative MIMO detectionabstractWe propose an efficient iterative multiple-input multiple-output (MIMO) detection algorithm based on the local search. Specifically, since the MIMO channel matrix twists the lattice structure of the received symbols, the proposed channel-aware local search (CA-LS) defines its search neighborhood according to the instantaneous channel realization. Such channel-dependent neighborhood can be efficiently identified by using the sphere decoder in a set that comprises the differences between pairs of QAM vectors, which is termed as delta vectors. The delta vectors with small quadratic norms with respect to the channel matrix are identified and then used as the search directions of the CA-LS. Features like sparsity and non-uniformity of delta vectors are exploited to reduce the SD complexity. Furthermore, by reformulating the detection criterion, the log-likelihood ratio (LLR) computations and searches of the CA-LS are greatly simplified. Numerical simulations demonstrate that compared with other practical iterative MIMO detectors, e.g., the list sphere decoder, the CA-LS achieves superior performance in both error rate and complexity aspects. I-Wei Lai, Chia-han Lee, Gerd Ascheid, Heinrich Meyr, Tzi-Dar Chiueh |
PIMRC | 4 |
| 2014 | Oversampling Increases the Pre-Log of Noncoherent Rayleigh Fading ChannelsabstractWe analyze the capacity of a continuous-time, time-selective, Rayleigh block-fading channel in the high signal-to-noise ratio (SNR) regime. The fading process is assumed stationary within each block and to change independently from block to block; furthermore, its realizations are not known a priori to the transmitter and the receiver (noncoherent setting). A common approach to analyzing the capacity of this channel is to assume that the receiver performs matched filtering followed by sampling at symbol rate (symbol matched filtering). This yields a discrete-time channel in which each transmitted symbol corresponds to one output sample. Liang & Veeravalli (2004) showed that the capacity of this discrete-time channel grows logarithmically with the SNR, with a capacity pre-log equal to 1-Q/N. Here, N is the number of symbols transmitted within one fading block, and Q is the rank of the covariance matrix of the discrete-time channel gains within each fading block. In this paper, we show that symbol matched filtering is not a capacity-achieving strategy for the underlying continuous-time channel. Specifically, we analyze the capacity pre-log of the discrete-time channel obtained by oversampling the continuous-time channel output, i.e., by sampling it faster than at symbol rate. We prove that by oversampling by a factor two one gets a capacity pre-log that is at least as large as 1-1/N. Since the capacity pre-log corresponding to symbol-rate sampling is 1-Q/N, our result implies indeed that symbol matched filtering is not capacity achieving at high SNR. Meik Dörpinghaus, Günther Koliander, Giuseppe Durisi, Erwin Riegler, Heinrich Meyr |
IEEE Trans. Inf. Theory | 5 |
| 2013 | On the Achievable Rate of Stationary Rayleigh Flat-Fading Channels With Gaussian InputsabstractIn this work, a discrete-time stationary Rayleigh flat-fading channel with unknown channel state information at transmitter and receiver side is studied. The law of the channel is presumed to be known to the receiver. For independent identically distributed (i.i.d.) zero-mean proper Gaussian input distributions, the achievable rate is investigated. The main contribution of this paper is the derivation of two new upper bounds on the achievable rate with Gaussian input symbols. One of these bounds is based on the one-step channel prediction error variance but is not restricted to peak power constrained input symbols like known bounds. Moreover, it is shown that Gaussian inputs yield the same pre-log as the peak power constrained capacity. The derived bounds are compared with a known lower bound on the capacity given by Deng and Haimovich and with bounds on the peak power constrained capacity given by Sethuraman et al.. Finally, the achievable rate with i.i.d. Gaussian input symbols is compared to the achievable rate using a coherent detection in combination with a solely pilot-based channel estimation. Meik Dörpinghaus, Heinrich Meyr, Rudolf Mathar |
IEEE Trans. Inf. Theory | 2 |
| 2012 | Searching for optimal scheduling of MIMO doubly iterative receivers: An ant colony optimization-based methodabstractWhen an iterative receiver has more than one iteration loop, scheduling the iterative decoding process is an important issue. To find the optimal scheduling, extrinsic information transfer (EXIT) function was the main tool in the literature. However, it is only very efficient when the codeword is significantly long. In this paper, under the consideration of a packet-based transmission scenario, we propose modeling the scheduling search problem as the foraging problem of ant colony. And then, an ant colony optimization (ACO) algorithm, labeled as max-min ant system (MMAS), is adopted and tailored for our problem. Simulation results show MMAS outperforms the EXIT function-based method when practical length codewords are used. Dan Zhang 0003, Gaojian Wang, Gerd Ascheid, Heinrich Meyr |
GLOBECOM | 4 |
| 2012 | Asymptotic Coded BER Analysis for MIMO BICM-ID with Quantized Extrinsic LLRabstractIn this paper, we derive a closed-form expression for the probability density/mass function (PDF/PMF) and the moment generating function (MGF) of the quantized detector soft output, i.e., extrinsic log-likelihood ratio (LLR), for multiple-input multiple-output (MIMO) bit-interleaved coded modulation with iterative decoding (BICM-ID) systems. The effect of either LLR clipping or LLR clipping with rounding, often applied in practical implementations, are considered. Using the derived expression, we analyze the asymptotic coded bit error rate (BER) for MIMO BICM-ID systems with the two quantization operations under a flat Rayleigh fading channel. The error rate degradation caused by quantizing the extrinsic LLR is then interpreted as an additional signal-to-noise ratio (SNR) loss. Rather than Monte Carlo simulations, this theoretical treatment provides a more convenient alternative to determining the clipping level and optimal word-length (the number of bits needed to represent signals) for LLR in a BICM-ID implementation. Finally, several other applications of this proposed theoretical analysis are also demonstrated. I-Wei Lai, Chien-Yi Wang, Tzi-Dar Chiueh, Gerd Ascheid, Heinrich Meyr |
IEEE Trans. Commun. | 5 |
| 2012 | On the Gain of Joint Processing of Pilot and Data Symbols in Stationary Rayleigh Fading ChannelsabstractIn many typical mobile communication receivers, the channel is estimated based on pilot symbols to allow for a coherent detection and decoding in a separate processing step. Currently, much work is spent on receivers which break up this separation, e.g., by enhancing channel estimation based on reliability information on the data symbols. In this paper, we evaluate the possible gain of a joint processing of data and pilot symbols in comparison to the case of a separate processing in the context of stationary Rayleigh flat-fading channels. Therefore, we discuss the nature of the possible gain of a joint processing of pilot and data symbols. We show that the additional information that can be gained by a joint processing is captured in the temporal correlation of the channel estimation error of the solely pilot-based channel estimation, which is not retrieved by the channel decoder in case of separate processing. In addition, we derive a new lower bound on the achievable rate for joint processing of pilot and data symbols. Finally, the results are extended to multiple-input multiple-output channels. Meik Dörpinghaus, Adrian Ispas, Heinrich Meyr |
IEEE Trans. Inf. Theory | 3 |
| 2011 | Asymptotic BER Analysis for MIMO-BICM with MMSE Detection and Channel EstimationabstractIn this paper, we theoretically analyze the asymptotic coded bit error rate (BER) for multiple-input multiple-output bit-interleaved coded modulation (MIMO-BICM) with linear minimum mean-squared error (MMSE) detection and estimation for a flat Rayleigh fading channel. The numerical simulations validate the accuracy of our theoretical analysis. With such study, we can model the BER improvement of MMSE detection compared with zero-forcing detection as a signal-to-noise (SNR) gain. I-Wei Lai, Gerd Ascheid, Heinrich Meyr, Tzi-Dar Chiueh |
ICC | 3 |
| 2011 | On Complexity, Energy- and Implementation-Efficiency of Channel DecodersabstractFuture wireless communication systems require efficient and flexible baseband receivers. Meaningful efficiency metrics are key for design space exploration to quantify the algorithmic and the implementation complexity of a receiver. Most of the current established efficiency metrics are based on counting operations, thus neglecting important issues like data and storage complexity. In this paper we introduce suitable energy and area efficiency metrics which resolve the afore-mentioned disadvantages. These are decoded information bit per energy and throughput per area unit. Efficiency metrics are assessed by various implementations of turbo decoders, LDPC decoders and convolutional decoders. An exploration approach is presented, which permit an appropriate benchmarking of implementation efficiency, communications performance, and flexibility trade-offs. Two case studies demonstrate this approach and show that design space exploration should result in various efficiency evaluations rather than a single snapshot metric as done often in state-of-the-art approaches. Frank Kienle, Norbert Wehn, Heinrich Meyr |
IEEE Trans. Commun. | 3 |
| 2011 | Efficient Channel-Adaptive MIMO Detection Using Just-Acceptable Error RateabstractThis paper proposes a new concept of multiple-input multiple-output (MIMO) detection, aiming at minimizing the average computational cost. The detection methods are adapted according to the estimated channel state information to deliver just-acceptable error rate (JAER) performance. Error rate models for two popular MIMO detection algorithms are derived given the channel matrix. From these models, a channel-adaptive-MIMO (CA-MIMO) receiver with detector-switching strategies is proposed. Simulation results demonstrate that the proposed CA-MIMO detector meets the JAER criterion efficiently. Compared with the sophisticated Sphere Search (SS) MIMO detector, the average saving at moderate signal-to-noise ratio (SNR) is around 58% to 72%, depending on different modulation alphabets. At high SNR, the CA-MIMO detector almost always switch to Zero-Forcing (ZF) detection, where the complexity is several orders lower than the SS detector. I-Wei Lai, Gerd Ascheid, Heinrich Meyr, Tzi-Dar Chiueh |
IEEE Trans. Wirel. Commun. | 3 |
| 2010 | Trace-based KPN composability analysis for mapping simultaneous applications to MPSoC platformsabstractNowadays, most embedded devices need to support multiple applications running concurrently. In contrast to desktop computing, very often the set of applications is known at design time and the designer needs to assure that critical applications meet their constraints in every possible use-case. In order to do this, all possible use-cases, i.e. subset of applications running simultaneously, have to be verified thoroughly. An approach to reduce the verification effort, is to perform composability analysis which has been studied for sets of applications modeled as Synchronous Dataflow Graphs. In this paper we introduce a framework that supports a more general parallel programming model based on the Kahn Process Networks Model of Computation and integrates a complete MPSoC programming environment that includes: compiler-centric analysis, performance estimation, simulation as well as mapping and scheduling of multiple applications. In our solution, composability analysis is performed on parallel traces obtained by instrumenting the application code. A case study performed on three typical embedded applications, JPEG, GSM and MPEG-2, proved the applicability of our approach. Jerónimo Castrillón, Ricardo Velasquez, Anastasia Stulova, Weihua Sheng, Jianjiang Ceng, Rainer Leupers, Gerd Ascheid, Heinrich Meyr |
DATE | 8 |
| 2010 | The achievable rate of stationary rayleigh flat-fading channels with IID input symbolsabstractIn this work, we derive a new upper bound on the achievable rate of stationary Rayleigh flat-fading channels with i.i.d. input symbols. The novelty lies in the fact that this bound is not restricted to peak power constrained input symbols like known bounds, e.g., in [1] or [2]. Therefore, the derived upper bound can also be used to evaluate the achievable rate with i.i.d. proper Gaussian input symbols, which are capacity achieving in the coherent case. The derivation of the upper bound is based on the prediction error variance of the one-step channel predictor. Meik Dörpinghaus, Heinrich Meyr, Gerd Ascheid |
ISITA | 2 |
| 2010 | BER analysis for MIMO BICM-ID assuming finite precision of extrinsic LLRabstractIn this paper, we analyze the error rate degradation caused by the finite-precision extrinsic log-likelihood ratio (LLR), i.e., demapper output, for multiple-input multiple-output (MIMO) bit-interleaved coded modulation with iterative decoding (BICM-ID). A closed-form expression for the probability density function (pdf) of the metric difference with finite precision is derived. This pdf is well-approximated by two approaches, namely the pseudo quantization noise (PQN) model and the pdf tail truncation. Moreover, we interpret such performance degradation as an additional signal-to-noise ratio (SNR) loss. As validated by Monte Carlo simulations, this SNR loss can be directly applied at the first iteration, extending our analysis to non-iterative scenario. This work provides a convenient indicator for deciding the LLR precision in real implementations. Chien-Yi Wang, I-Wei Lai, Tzi-Dar Chiueh, Gerd Ascheid, Heinrich Meyr |
ISITA | 5 |
| 2010 | A Systematic Framework for Iterative Maximum Likelihood Receiver DesignabstractIn this paper, we link the turbo principle to unconstrained maximum likelihood (ML) sequence detection and joint ML parameter estimation. First, we demonstrate for memoryless channels with complete channel state information how the turbo decoder can be systematically derived starting from the ML sequence detection criterion. In particular, we show that a method to solve the ML sequence detection problem is to iteratively solve the corresponding critical point equations of an equivalent unconstrained estimation problem by means of fixed-point iterations. The turbo decoding algorithm is obtained by approximating the overall a posteriori probabilities. Subsequently, we show how this general approximative iterative maximum likelihood (AIML) framework can be applied to general iterative ML receiver design. We consider static memoryless channels with unknown channel parameters. The time-selective fading channels with partial channel state information is the subject of a companion paper. Lars Schmitt, Heinrich Meyr |
IEEE Trans. Commun. | 2 |
| 2010 | Systematic Design of Iterative ML Receivers for Flat Fading ChannelsabstractWe extend the methodology in to flat fading channels. The systematically derived receiver structure corresponds to an iterative solution of the joint maximum likelihood (ML) optimization problem. The heuristically derived BICM with iterative decoding is shown to be the outcome of this joint ML optimization problem. Lars Schmitt, Heinrich Meyr, Dan Zhang 0003 |
IEEE Trans. Commun. | 2 |
| 2009 | Searching in the Delta Lattice: An Efficient MIMO Detection for Iterative ReceiversabstractThis paper introduces a new framework of the multiple-input multiple-output (MIMO) detection in iterative receivers. Unlike the conventional methods processing with symbol lattice, we consider the delta symbol lattice, i.e., the difference between two arbitrary points in the symbol lattice. The inherent flexible, symmetric, and sparse properties of the delta lattice enhance the detection in both complexity and performance aspects. Consequently, we propose a delta-list MIMO (DL-MIMO) detection which separately exploits the channel information and the a priori information so that a soft-input soft-output sphere decoder is dispensable. Simulation results demonstrate this hardware-friendly DL-MIMO detection delivers nearly-optimal performance at affordable cost in a practical scenario. I-Wei Lai, Chun-Hao Liao, Ernst Martin Witte, David Kammler, Filippo Borlenghi, Konstantinos Nikitopoulos, Venkatesh Ramakrishnan, Dan Zhang 0003, Tzi-Dar Chiueh, Gerd Ascheid, Heinrich Meyr |
GLOBECOM | 11 |
| 2009 | Combining orthogonalized partial metrics: Efficient enumeration for soft-input sphere decoderabstractUsing the Schnorr-Euchner (SE) order for soft-input sphere decoders is inefficient for implementation, because it requires exhaustive calculation and sorting of partial metrics of all constellation points. Instead, low-complexity methods can be applied by separating the partial metric into channel information and a priori information and solely enumerating based on one of them. With such an orthogonalization, this paper presents an algorithm that effectively combines these two enumerations to deliver an order close to the SE one. Mathematical analyses and simulation results demonstrate that this is the first algorithm allowing for a low-complexity implementation with optimal error rate performance for any number of iterations. Chun-Hao Liao, I-Wei Lai, Konstantinos Nikitopoulos, Filippo Borlenghi, David Kammler, Ernst Martin Witte, Dan Zhang 0003, Tzi-Dar Chiueh, Gerd Ascheid, Heinrich Meyr |
PIMRC | 10 |
| 2009 | Efficient implementations from libraries: Analyzing the influence of configuration parameters on key performance propertiesabstractLibrary based waveform (WF) development approaches have the potential to address one of the key challenges in software defined radio technology, developing portable and at the same time implementation-efficient WFs. The term WF, in this paper, represents a complete wireless standard with several modes. Upon standardizing the library and the interfaces of its components, different vendors can provide Flavors, which are efficient implementations, for some/all components of a WF as a board-support-package. Flavors might provide different trade-offs with respect to key performance related properties like bit error rate, throughput, latency, area, etc. This paper analyzes such a scenario by using FFT as an example. We use 10 Flavors of FFT from 2 vendors for 3 processing elements to analyze the influence of configuration parameters like input data-width, scaling, etc. on the performance properties. Analysis has been performed by doing two tests: using a simple OFDM transceiver system, calculating root mean square error. Based on our analysis, we have identified the main parameters that influence the properties. Our investigations stress the need for modeling the differences in performance on key properties in terms of the influencing parameters. Such a model can assist in selecting a set of appropriate Flavors for implementing a WF and enable tool assisted WF-development. Furthermore, it is evident from our investigations that Flavors exhibit trade-offs in properties and pose tough constraints in several aspects. Venkatesh Ramakrishnan, Joschka zur Jacobsmühlen, I-Wei Lai, Torsten Kempf, Marc Adrat, Gerd Ascheid, Markus Antweiler, Heinrich Meyr |
PIMRC | 8 |
| 2009 | Low-Complexity Channel-Adaptive MIMO Detection with Just-Acceptable Error RateabstractThis paper proposes a new concept of multiple-input multiple-output (MIMO) detection, aiming at minimizing the average computational cost. The detection methods, according to the estimated channel state information, are adapted to deliver just acceptable error rate. Performance of two popular MIMO detectors is analyzed. Based on these analyses, a channel adaptive MIMO (CA-MIMO) receiver with a novel method- selection strategy is investigated. Simulation results demonstrate that the proposed CA-MIMO detector, operating at as low SNR as the sphere search does, achieves huge average complexity saving. I-Wei Lai, Gerd Ascheid, Heinrich Meyr, Tzi-Dar Chiueh |
VTC Spring | 3 |
| 2009 | A SIMD optimization framework for retargetable compilersabstractRetargetable C compilers are currently widely used to quickly obtain compiler support for new embedded processors and to perform early processor architecture exploration. A partially inherent problem of the retargetable compilation approach, though, is the limited code quality as compared to hand-written compilers or assembly code due to the lack of dedicated optimizations techniques. This problem can be circumvented by designing flexible, retargetable code optimization techniques that apply to a certain range of target architectures. This article focuses on target machines with SIMD instruction support, a common feature in embedded processors for multimedia applications. However, SIMD optimization is known to be a difficult task since SIMD architectures are largely nonuniform, support only a limited set of data types and impose several memory alignment constraints. Additionally, such techniques require complicated loop transformations, which are tailored to the SIMD architecture in order to exhibit the necessary amount of parallelism in the code. Thus, integrating the SIMD optimization and the required loop transformations together in a single retargeting formalism is an ambitious challenge. In this article, we present an efficient and quickly retargetable SIMD code optimization framework that is integrated into an industrial retargetable C compiler. Experimental results for different processors demonstrate that the proposed technique applies to real-life target machines and that it produces code quality improvements close to the theoretical limit. Manuel Hohenauer, Felix Engel 0001, Rainer Leupers, Gerd Ascheid, Heinrich Meyr |
ACM Trans. Archit. Code Optim. | 5 |
| 2008 | MAPS: an integrated framework for MPSoC application parallelizationabstractIn the past few years, MPSoC has become the most popular solution for embedded computing. However, the challenge of programming MPSoCs also comes as the biggest side-effect of the solution. Especially, when designers have to face the legacy C code accumulated through the years, the tool support is mostly unsatisfactory. In this paper, we propose an integrated framework, MAPS, which aims at parallelizing C applications for MPSoC platforms. It extracts coarsegrained parallelism on a novel granularity level. A set of tools have been developed for the framework. We will introduce the major components and their functionalities. Two case studies will be given, which demonstrate the use of MAPS on two different kinds of applications. In both cases the proposed framework helps the programmer to extract parallelism efficiently. Jianjiang Ceng, Jerónimo Castrillón, Weihua Sheng, Hanno Scharwächter, Rainer Leupers, Gerd Ascheid, Heinrich Meyr, Tsuyoshi Isshiki, Hiroaki Kunieda |
DAC | 7 |
| 2008 | Multiprocessor performance estimation using hybrid simulationabstractWith the growing number of programmable processing elements in today's Multi Processor System-on-Chip (MPSoC) designs, the synergy required for the development of the hardware architecture and the software running on them is also increasing. In MPSoC development environment, changes in the hardware architecture can bring in extensive re-partitioning or re-parallelization of the software architecture. Fast and accurate functional simulation and performance estimation techniques are needed to cope with this co-design problem at the early phases of MPSoC design space exploration. The current paper addresses this issue by introducing a framework which combines hybrid simulation, cache simulation and online trace-driven replay techniques to accurately predict performance of programmable elements in an MPSoC environment. The resulting simulation technique can easily cope with the continuous re-organizations of software architectures during an Instruction Set Simulator (ISS) based design process. Experimental results show that this framework can improve system simulation speed by 3-5× on average while achieving accuracy closely comparable to traditional ISSes. Kingshuk Karuri, Stefan Kraemer, Rainer Leupers, Gerd Ascheid, Heinrich Meyr |
DAC | 6 |
| 2008 | High-level Modelling and Exploration of Coarse-grained Re-configurable ArchitecturesabstractThe increasing complexity of today's multimedia and wireless applications is motivating the system designers to innovate continuously. With the challenge to keep various performance metrics in a tight balance while designing a complex system, an entire range of components are now being offered as choices for system building blocks. Coarse-Grained Re-configurable Architecture (CGRA), a strongly emerging class, is currently receiving due attention for offering excellent performance as well as flexibility post fabrication. Compared to the programmable and flexible microprocessors these architectures are shown to yield stronger performance, especially in case of regular and data-driven applications. A variety of system designs are proposed of late, with CGRA as one of the key building blocks. Most of the research initiatives taken in this area have resorted to a template-based approach, where the structure of the re-configurable architecture is partially fixed with several tunable parameters. In this paper, we present a language-driven modelling and exploration framework for CGRAs. In the domain of CGRAs, this framework attempts to bring modelling ease, genericity, early exploration and path to implementation together. The modelling formalism proposed in this paper as well as the exploration capabilities are demonstrated via experiments with several algorithmic kernels. Anupam Chattopadhyay, Harold Ishebabi, Rainer Leupers, Gerd Ascheid, Heinrich Meyr |
DATE | 6 |
| 2008 | Retargetable Code Optimization for Predicated ExecutionabstractRetargetable C compilers are key components of today's embedded processor design platforms for quickly obtaining compiler support and performing early processor architecture exploration. The inherent problem of the retargetable compilation approach, though, is the well known trade-off between the compiler's flexibility and the quality of generated code. However, it can be circumvented by designing flexible, configurable code optimization techniques applicable to a certain range of target architectures. This paper focuses on target machines with predicated execution support which is wide-spread in deeply pipelined and highly parallel embedded processors used in next generation high-end video, multimedia and wireless devices. We present an efficient and quickly retargetable code optimization technique for predicated execution that is integrated into an industrial retargetable C compiler. Experimental results for several embedded processors demonstrate that the proposed technique is applicable to real-life target machines and that it produces significant code quality improvements for control intensive applications. Manuel Hohenauer, Felix Engel 0001, Rainer Leupers, Gerd Ascheid, Heinrich Meyr, Gerrit Bette, Balpreet Singh |
DATE | 5 |
| 2008 | Asymptotic BER Analysis for MIMO-BICM with Zero-Forcing Detectors Assuming Imperfect CSIabstractIn this paper, we derive the asymptotic bit error rate (BER) for multiple-input multiple-output bit-interleaved coded modulation (MIMO-BICM) with linear zero-forcing (ZF) receivers for a temporally correlated flat Rayleigh fading channel. Pilot symbol assisted modulation (PSAM) in combination with linear minimum mean-squared error (LMMSE) channel estimation is considered. We also demonstrate that the deterioration due to imperfect channel state information (CSI) can be represented by a signal-to-noise ratio (SNR) degradation. I-Wei Lai, Susanne Godtmann, Tzi-Dar Chiueh, Gerd Ascheid, Heinrich Meyr |
ICC | 5 |
| 2008 | Optimal PSK signaling over stationary Rayleigh fading channelsabstractWe consider a stationary Rayleigh flat-fading channel with temporal correlation and a compactly supported power spectral density of the channel fading process. We assume that the channel state is unknown to both transmitter and receiver, while the law of the channel is presumed to be known to the receiver. Given a set of fixed signaling sequences, the optimum input distribution, with respect to the achievable rate, has the property of a constant Kullback-Leibler distance between the output distribution and a mixture of the output distributions conditioned on the different input sequences. Based on this, we determine the set of optimum input distributions for PSK signaling. In addition, we identify the special case of transmitting one pilot symbol to acquire a phase reference as being included in the set of optimum input distributions. We derive an integral expression for the capacity constrained to PSK signaling depending on the autocorrelation of the channel and the SNR. Evaluation of the asymptotic high SNR behavior shows a loss in the constrained capacity with respect to the case of perfect channel knowledge corresponding to at least one signaling dimension, i.e., the information transmitted by one symbol. Meik Dörpinghaus, Gerd Ascheid, Heinrich Meyr, Rudolf Mathar |
ISIT | 3 |
| 2008 | Long-Term Beamforming in Single Frequency Networks using Semidefinite RelaxationabstractWe examine the application of beamforming in a single frequency network (SFN) for the provision of multicast services. A single frequency network is characterized by the simultaneous transmission of the same signal from multiple base stations. Since each user receives the signal from several base stations, the beamforming weights for many different base stations need to be jointly optimized for the transmission in order to give optimal results. The determination of the beamforming weights that maximize the minimum SNR is a nonconvex problem. A constraint relaxation however yields a convex program, which can be solved with fast algorithms at only a slight loss of performance. Furthermore, this convex formulation of the problem gives solutions that perform significantly better than omnidirectional transmission. Markus Jordan, Martin Senst, Gerd Ascheid, Heinrich Meyr |
VTC Spring | 4 |
| 2008 | Prefabrication and postfabrication architecture exploration for partially reconfigurable VLIW processorsabstractModern application-specific instruction-set processors (ASIPs) face the daunting task of delivering high performance for a wide range of applications. For enhancing the performance, architectural features, for example, pipelining, VLIW, are often employed in ASIPs, leading to high design complexity. Integrated ASIP design environments, like template-based approaches and language-driven approaches, provide an answer to this growing design complexity. At the same time, increasing hardware design costs have motivated the processor designers to introduce high flexibility in the processor. Flexibility, in its most effective form, can be introduced to the ASIP by coupling a reconfigurable unit to the base processor. Because of its obvious benefits, several reconfigurable ASIPs (rASIPs) have been designed for years. This design paradigm gained momentum with the advent of coarse-grained FPGAs, where the lack of domain-specific performance common in general-purpose FPGAs are largely overcome by choosing application-dependent basic functional units. These rASIP designs lack a generic flow from high-level specification, resulting in intuitive design decisions and hard-to-retarget processor design tools. Although partial, template-based approaches for rASIP design is existent, a clear design methodology especially for the prefabrication architecture exploration is not present. In order to address this issue, a high-level specification and design methodology for partially reconfigurable VLIW processors is proposed in this article. To show the benefit of this approach, a commercial VLIW processor is used as the base architecture and two domains of applications are studied for potential performance gain. Anupam Chattopadhyay, Harold Ishebabi, Zoltán Endre Rákossy, Kingshuk Karuri, David Kammler, Rainer Leupers, Gerd Ascheid, Heinrich Meyr |
ACM Trans. Embed. Comput. Syst. | 9 |
| 2008 | A Design Flow for Architecture Exploration and Implementation of Partially Reconfigurable ProcessorsabstractDuring the last years, the growing application complexity, design, and mask costs have compelled embedded system designers to increasingly consider partially reconfigurable application-specific instruction set processors (rASIPs) which combine a programmable base processor with a reconfigurable fabric. Although such processors promise to deliver excellent balance between performance and flexibility, their design remains a challenging task. The key to the successful design of a rASIP is combined architecture exploration of all the three major components: the programmable core, the reconfigurable fabric, and the interfaces between these two. This work presents a design flow that supports fast architecture exploration for rASIPs. The design flow is centered around a unified description of an entire rASIP in an architecture description language (ADL). This ADL description facilitates consistent modeling and exploration of all three components of a rASIP through automatic generation of the software tools (compiler tool chain and instruction set simulator) and the RTL hardware model. The generated software tools and the RTL model can be used either for final implementation of the rASIP or can serve as a preoptimized starting point for implementation that can be hand optimized afterward. The design flow is further enhanced by a number of automatic application analysis tools, including a fine-grained application profiler, an instruction set extension (ISE) generator, and a data path mapper for coarse grained reconfigurable architectures (CGRAs). We present some case studies on embedded benchmarks to show how the design space exploration process helps to efficiently design an application domain specific rASIP. Kingshuk Karuri, Anupam Chattopadhyay, David Kammler, Ling Hao, Rainer Leupers, Heinrich Meyr, Gerd Ascheid |
IEEE Trans. Very Large Scale Integr. Syst. | 7 |
| 2008 | Tight Approximation of the Bit Error Rate for BICM(-ID) Assuming Imperfect CSIabstractIn this correspondence, we analytically derive the bit error rate (BER) for bit-interleaved coded modulation (BICM) under the assumption of a temporally correlated flat Rayleigh fading channel. We assume the channel state information (CSI) to be unknown at receiver side and consider a linear minimum mean-squared error (LMMSE) channel estimator that relies on periodically inserted pilot symbols. We give asymptotic results for both non-iterative BICM and BICM with iterative decoding (BICM-ID). Furthermore, we show that the interpretation of the channel estimation error as an SNR degradation holds true in these cases. Susanne Godtmann, I-Wei Lai, Gerd Ascheid, Tzi-Dar Chiueh, Heinrich Meyr |
IEEE Trans. Wirel. Commun. | 5 |
| 2007 | A fast and generic hybrid simulation approach using C virtual machineabstractInstruction Set Simulators (ISSes) are important tools for cross-platform software development. The simulation speed is a major concern and many approaches have been proposed to improve the performance of ISSes. A prevalent technique is compiled simulation, which translates target programs into host instructions. But orders of magnitude of speed deterioration is inevitable since the difference between target and host Instruction Set Architectures (ISAs) can be large. An alternative is to emulate the program without sticking to binary compatibility. The performance problem is solved by using native execution. However, these emulators either require a special programming language, or a given Application Programming Interface (API). Last but not least, it is not trivial to integrate an emulator into a system simulator (which provides devices, external memory, etc., that the embedded programmers do care). In this paper, we propose a fast and generic hybrid simulation approach using virtualization technique to accelerate simulation and simulator-based debugging of C programs. A novel virtual coprocessor (VCP) is introduced as a processing element which executes C functions at high speed. This approach is C89 compliant and compatible with third party libraries and platform dependent code. It is also retargetable and can be integrated with existing ISSes. Two different ISAs are supported at present: MIPS and mAgic DSP. The average execution speed of the coprocessor is about 100 million simulated instructions per second. Stefan Kraemer, Rainer Leupers, Gerd Ascheid, Heinrich Meyr |
CASES | 5 |
| 2007 | Design space exploration of partially re-configurable embedded processorsabstractIn today's embedded processors, performance and flexibility have become the two key attributes. These attributes are often conflicting. The best performance is obtained from custom designed integrated circuits. In contrast, the maximum flexibility is delivered by a general purpose processor. Among the architecture types emerged over the past years to strike an optimum balance between these two attributes, two are prominent. The first ones are field programmable gate array (FPGA)-based architectures and the second ones are application-specific instruction-set processors (ASIPs). Depending on the type of application (i.e. stream-like or control-dominated) either one of the above mentioned architecture types is able to deliver high performance or flexibility or both. Consequently, a new design approach with partial re-configurability on the application-specific processor is attracting strong research interest. We call this architecture re-configurable ASIP (rASIP). Currently, the lack of a high-level abstraction of the rASIP limits the designer from trying out various design alternatives because of long and tedious exploration cycles. To address this issue, in this paper, a high-level specification for re-configurable processors is proposed. Furthermore, a seamless design space exploration methodology using this specification is proposed Anupam Chattopadhyay, W. Ahmed, Kingshuk Karuri, David Kammler, Rainer Leupers, Gerd Ascheid, Heinrich Meyr |
DATE | 7 |
| 2007 | Interactive presentation: SoftSIMD - exploiting subword parallelism using source code transformations
Stefan Kraemer, Rainer Leupers, Gerd Ascheid, Heinrich Meyr |
DATE | 4 |
| 2007 | Joint Reduction of Peak-to-Average Power Ratio and Out-of-Band Power in OFDM SystemsabstractThe high peak-to-average power ratio is a major drawback of OFDM systems. Many PAPR reduction techniques have been proposed in the literature, among them a method that uses a subset of tones that do not carry any data, but are modulated such that the PAPR of the resulting time domain signal is minimized. Another problem of OFDM systems is the high out-of-band power caused by the sidelobes of the modulated tones. The OBP can be reduced by modulating reserved tones at the edges of the occupied spectrum so that the sidelobes of the data carriers are reduced. In this paper, we propose to consider both optimization problems jointly. This way, the amount of PAPR and OBP reduction can be significantly enhanced in comparison to a system that performs two separate optimization steps. Furthermore, the joint reduction algorithm offers more flexibility, because the relative weighting of the two optimization criteria can easily be changed, resulting in a smooth trade-off curve. Martin Senst, Markus Jordan, Meik Dörpinghaus, Gerd Ascheid, Heinrich Meyr |
GLOBECOM | 6 |
| 2007 | Increasing data-bandwidth to instruction-set extensions through register clusteringabstractThe conflicting requirements of performance and flexibility in today 's embedded system market are forcing system designers to use more and more of the so called configurable or customizable processor cores. Such processors tend to meet the demanding performance constraints by accommodating application specific instruction set extensions (ISEs) which have, naturally, become a vital component of current processor customization flows. One major bottleneck in maximizing ISE performance is the limitation on the data-bandwidth between the general purpose register (GPR) file and the ISEs. For improved performance, it is desirable to have a large data-bandwidth from the GPRs to ISEs. However, the tight area constraints of modern embedded processors often restrict the GPR I/O of ISEs to save port area of the register files. This paper presents a novel approach to increase the GPR I/O of ISEs without significantly increasing the size of the GPR files. This is achieved by applying the concept of register clustering, common in many VLIW architectures, to single-issue processors with high performance ISEs. Such clustering often causes extra register moves in compiled code. This work also presents an algorithm to minimize such register moves. The benchmark results presented in this paper show that our solution can significantly reduce the area overhead of many-port GPR files without sacrificing the performance improvements through ISEs. Kingshuk Karuri, Anupam Chattopadhyay, Manuel Hohenauer, Rainer Leupers, Gerd Ascheid, Heinrich Meyr |
ICCAD | 6 |
| 2007 | On the Influence of Pilot Symbol and Data Symbol Positioning on Turbo SynchronizationabstractWhen realizing carrier frequency offset estimation at low signal-to-noise ratios, a typical feed-forward synchronization unit solely relies on known pilot symbols. The importance of the position of these pilot symbols within the burst has been elaborated on in the literature. In this paper, we discuss the importance of the pilot symbol constellation for iterative code-aided synchronization. Furthermore, we investigate the transmission order of data symbols for a turbo coded system with code-aided synchronization. Susanne Godtmann, André Pollok, Niels Hadaschik, Gerd Ascheid, Heinrich Meyr |
VTC Spring | 5 |
| 2007 | Performance Evaluation of Opportunistic Beamforming with SINR Prediction for HSDPAabstractIn this paper, opportunistic beamforming for HSDPA is considered. Opportunistic beamforming can jointly increase the system throughput and the quality of service in multiuser communication systems with proportional-fair scheduling, since with this beamforming scheme a user experiences more fading peaks in a given period of time. However, the existence of an SINR feedback delay from the user to the base station degrades the performance in such a system rapidly due to outdated and thus mismatched SINR information. Here, the prediction of the SINR in the base station is proposed with a predictor structure that is matched to the beamforming process. Markus Jordan, Gerd Ascheid, Heinrich Meyr |
VTC Spring | 3 |
| 2007 | Prediction of Downlink SNR for Opportunistic BeamformingabstractOpportunistic beamforming (OBF) is a technique to jointly increase the throughput and quality of service (QoS) in a multiuser system with fair and channel-aware scheduling. With this technique, the channel dynamics are influenced such that fading peaks occur more often within a given period of time. Channel-aware scheduling requires reliable SNR information in the base station that is obtained via feedback information from the users. Since this feedback is not instantaneous, the problem caused by a feedback delay are exacerbated by OBF because of higher channel dynamics. In this paper, a SNR predictor is derived for OBF that is matched to the beamforming process. Markus Jordan, Lars Schmitt, Gerd Ascheid, Heinrich Meyr |
WCNC | 4 |
| 2007 | ASIP architecture exploration for efficient IPSec encryption: A case studyabstractApplication-Specific Instruction-Set Processors (ASIPs) are becoming increasingly popular in the world of customized, application-driven System-on-Chip (SoC) designs. Efficient ASIP design requires an iterative architecture exploration loop---gradual refinement of the processor architecture starting from an initial template. To accomplish this task, design automation tools are used to detect bottlenecks in embedded applications, to implement application-specific processor instructions, and to automatically generate the required software tools (such as instruction-set simulator, C-compiler, assembler, and profiler), as well as to synthesize the hardware. This paper describes an architecture exploration loop for an ASIP coprocessor that implements common encryption functionality used in symmetric block cipher algorithms for internet protocol security (IPSec). The coprocessor is accessed via shared memory and, as a consequence, our approach is easily adaptable to arbitrary main processor architectures. This paper presents the extended version of our case study that has been already published on the SCOPES conference in 2004. In both papers, a MIPS architecture is used as the main processor and Blowfish as encryption algorithm. Hanno Scharwächter, David Kammler, Andreas Wieferink, Manuel Hohenauer, Kingshuk Karuri, Jianjiang Ceng, Rainer Leupers, Gerd Ascheid, Heinrich Meyr |
ACM Trans. Embed. Comput. Syst. | 9 |
| 2006 | Automatic ADL-based operand isolation for embedded processorsabstractCutting-edge applications of future embedded systems demand highest processor performance with low power consumption to get acceptable battery-life times. Therefore, low power optimization techniques are strongly applied during the development of modern application specific instruction set processors (ASIPs). Electronic system level design tools based on architecture description languages (ADL) offer a significant reduction in design time and effort by automatically generating the software tool-suite as well as the register transfer level (RTL) description of the processor. In this paper, the automation of power optimization in ADL-based RTL generation is addressed. Operand isolation is a well-known power optimization technique applicable at all stages of processor development. With increasing design complexity several efforts have been undertaken to automate operand isolation. In pipelined datapaths, where isolating signals are often implicitly available, the traditional RTL-based approach introduces unnecessary overhead. We propose an approach which extracts high-level structural information from the ADL representation and systematically uses the available control signals. Our experiments with state-of-the-art embedded processors show a significant power reduction (improvement in power efficiency) Anupam Chattopadhyay, Benedikt Geukes, David Kammler, Ernst Martin Witte, Oliver Schliebusch, Harold Ishebabi, Rainer Leupers, Gerd Ascheid, Heinrich Meyr |
DATE | 9 |
| 2006 | A SW performance estimation framework for early system-level-design using fine-grained instrumentationabstractThe increasing demands of high-performance in embedded applications under shortening time-to-market has prompted system architects in recent time to opt for multi-processor systems-on-chip (MP-SoCs) employing several programmable devices. The programmable cores provide a high amount of flexibility and reusability, and can be optimized to the requirements of the application to deliver high-performance as well. Since application software forms the basis of such designs, the need to tune the underlying SoC architecture for extracting maximum performance from the software code has become imperative. In this paper, we propose a framework that enables software development, verification and evaluation from the very beginning of MP-SoC design cycle. Unlike traditional SoC design flows where software design starts only after the initial SoC architecture is ready, our framework allows a co-development of the hardware and the software components in a tightly coupled loop where the hardware can be refined by considering the requirements of the software in a stepwise manner. The key element of this framework is the integration of a fine-grained software instrumentation tool into a system-level-design (SLD) environment to obtain accurate software performance and memory access statistics. The accuracy of such statistics is comparable to that obtained through instruction set simulation (ISS), while the execution speed of the instrumented software is almost an order of magnitude faster than ISS. Such a combined design approach assists system architects to optimize both the hardware and the software through fast exploration cycles, and can result in far shorter design cycles and high productivity. We demonstrate the generality and the efficiency of our methodology with two case studies selected from two most prominent and computationally intensive embedded application domains. Torsten Kempf, Kingshuk Karuri, Stefan Wallentowitz, Gerd Ascheid, Rainer Leupers, Heinrich Meyr |
DATE | 6 |
| 2006 | An interprocedural code optimization technique for network processors using hardware multi-threading supportabstractSophisticated C compiler support for network processors (NPUs) is required to improve their usability and consequently, their acceptance in system design. Nonetheless, high-level code compilation always introduces overhead, regarding code size and performance compared to handwritten assembly code. This overhead result partially from high-level function calls that usually introduce memory accesses in order to save and reload registers contents. A key feature of many NPU architectures is hardware multithreading support, in the form of separate register files, for fast context switching between different application tasks. In this paper, a new NPU code optimization technique to use such HW contexts is presented that minimizes the overhead for saving and reloading register contents for function calls via the runtime stack. The feasibility and the performance gain of this technique are demonstrated for the Infineon Technologies PP32 NPU architecture and typical network application kernels Hanno Scharwächter, Manuel Hohenauer, Rainer Leupers, Gerd Ascheid, Heinrich Meyr |
DATE | 5 |
| 2006 | Design of Application Specific Processors for the Cached FFT AlgorithmabstractOrthogonal frequency division multiplexing (OFDM) is a data transmission technique which is used in wired and wireless digital communication systems. In this technique, fast Fourier transformation (FFT) and inverse FFT (IFFT) are kernel processing blocks in an OFDM system, and are used for data (de)modulation. OFDM systems are increasingly required to be flexible to accommodate different standards and operation modes, in addition to being energy-efficient. A trade-off between these two conflicting requirements can be achieved by employing application-specific instruction-set processors (ASIPs). In this paper, two ASIP design concepts for the cached FFT algorithm (CFFT) are presented. A reduction in energy dissipation of up to 25% is achieved compared to an ASIP for the widely used Cooley-Tukey FFT algorithm, which was designed by using the same design methodology and technology. Further, a modified CFFT algorithm which enables a better cache utilization is presented. This modification reduces the energy dissipation by up to 10% compared to the original CFFT implementation Oguzhan Atak, Abdullah Atalar, Erdal Arikan, Harold Ishebabi, David Kammler, Gerd Ascheid, Heinrich Meyr, Mario Nicola, Guido Masera |
ICASSP (3) | 7 |
| 2006 | Enhanced Predictive Up/Down Power Control for CDMA SystemsabstractIn this paper we derive an enhanced power control algorithm, fitting into the up/down control scheme, as it is considered in the frequency division duplex (FDD) mode of the current 3GPP standard. Analysis of the classical up/down power control scheme unveils, that with increasing velocities the power control performance degrades, as the fixed step size power control is not able to track the channel fading properly. For the uplink we derive a nonlinear control algorithm generating the up/down power control commands accounting for the future of the channel fading process. Simulations show that this algorithm in combination with perfect future channel state information can partially mitigate the drawbacks of a fixed step-size up/down power control. A prerequisite for predictive power control is the acquisition of the future channel state information. In this paper we deduce a robust and adaptive structure for the prediction of the channel fading process in the context of a power controlled code division multiple access (CDMA) system based on least mean square (LMS) adaptation. Link level simulations show a signal to noise and interference ratio (SINR) gain in terms of the block error rate, enabling a decrease of the target SINR and thus leading to an enhanced spectral efficiency. Meik Dörpinghaus, Lars Schmitt, Ingo Viering, Axel Klein, Joachim Schmid 0001, Gerd Ascheid, Heinrich Meyr |
ICC | 7 |
| 2006 | An efficient parallelization technique for high throughput FFT-ASIPsabstractFast Fourier transformation (FFT) and its inverse (IFFT) are used in orthogonal frequency division multiplexing (OFDM) systems for data (de)modulation. The transformations are the kernel tasks in an OFDM implementation, and are the most processing-intensive ones. Recent trends in the electronic consumer market require OFDM implementations to be flexible, making a trade-off between area, energy-efficiency, flexibility and timing a necessity. This has spurred the development of application-specific instruction-set processors (ASIPs) for FFT processing. Parallelization is an architectural parameter that significantly influence design goals. This paper presents an analysis of the efficiency of parallelization techniques for an FFT-ASIP. It is shown that existing techniques are inefficient for high throughput applications such as ultra wideband (UWB), because of memory bottlenecks. Therefore, an interleaved execution technique which exploits temporal parallelism is proposed. With this technique, it is possible to meet the throughput requirement of UWB (409.6 Msamples/s) with only 4 non-trivial butterfly units for an ASIP that runs at 400MHz Harold Ishebabi, Gerd Ascheid, Heinrich Meyr, Oguzhan Atak, Abdullah Atalar, Erdal Arikan |
ISCAS | 3 |
| 2006 | Flexibility and low power: a contradiction in terms?abstractBoth configurable computing paradigms as well as re-configurable computing paradigms have gained significant impact within the last few years. Both paradigms have shown to be effective when power consumption is a major design constraint even though the philosophies behind are quite different: configurable approaches aim to adapt an embedded processor to an application through, for example, an extensible instruction set plus other parameters that are determined during design time. They come in two basic flavors: a) starting with a fixed core that is extended by the system designer or,b) designing the instruction set from scratch for a specific application..Re-configurable approaches on the other side gain most of their benefits through run-time re-configuration. A high degree of parallelism is needed to overcome the physical deficiencies of re-configurable fabrics (e.g. FPGAs), though.The panel will discuss advantages and disadvantages of these paradigms with respect to low power. Peter Wintermayr, Reiner W. Hartenstein, Heinrich Meyr, Steve Leibson |
ISLPED | 3 |
| 2006 | Achievable Data Rate of Wideband OFDM With Data-Aided Channel EstimationabstractThe achievable data rate of an OFDM system using data-aided channel estimation in the high bandwidth regime is evaluated under the assumption of a frequency selective, continuously fading channel. As previous results propose, the achievable data rate depends on the LMMSE channel estimate, for which a convenient representation is introduced here. The mean square estimation error is derived from this representation, allowing for an analysis with respect to the optimum amount and distribution of the pilot symbols in the wideband regime. The results on the data rate show good accordance to previous results based on non-data aided channel prediction especially in the interesting bandwidth range Niels Hadaschik, Gerd Ascheid, Heinrich Meyr |
PIMRC | 3 |
| 2005 | Instruction Set Customization of Application Specific Processors for Network Processing: A Case StudyabstractThe growth of the Internet in the last decade has made current networking applications immensely complex. Systems running such applications need special architectural support to meet the tight constraints of power and performance. This paper presents a case study of architecture exploration and optimization of an application specific instruction set processor (ASIP) for networking applications. The case study particularly focuses on the effects of instruction set customization for applications from different layers of the protocol stack. Using a state-of-the-art VLIW processor as the starting template, and architecture description language (ADL) based architecture exploration tools, this case study suggests possible instruction set and architectural modifications that can speed-up some networking applications up to 6.8 times. Moreover, this paper also shows that there exist very few similarities between diverse networking applications. Our results suggest that, it is extremely difficult to have a common set of architectural features for efficient network protocol processing and, ASIPs with specialized instruction sets can become viable solutions for such an application domain. Mohammad Mostafizur Rahman Mozumdar, Kingshuk Karuri, Anupam Chattopadhyay, Stefan Kraemer, Hanno Scharwächter, Heinrich Meyr, Gerd Ascheid, Rainer Leupers |
ASAP | 6 |
| 2005 | A framework for automated and optimized ASIP implementation supporting multiple hardware description languagesabstractArchitecture Description Languages (ADLs) are widely used to perform design space exploration for Application Specific Instruction Set Processors (ASIPs). While the design space exploration is well supported by numerous tools providing high flexibility and quality, the methodology of automated implementation is limited to simple transformations. Assuming fixed architectural templates, information given in the ADL is directly mapped to a hardware description on Register Transfer Level (RTL). Gate-Level synthesis tools are not able to perform potential optimizations, as the computational complexity grows exponential with the size of the architecture. Information such as exclusiveness, parallelism or boolean relations are spread over multiple modules and therefore hard to determine. In this paper, we present an ASIP synthesis approach from architecture description languages, based on an Intermediate Representation (IR). The IR is the key technology to provide new language-independent high-level optimizations and to realize different hardware description language backends. The feasibility of our approach is proven in a case-study. Oliver Schliebusch, Anupam Chattopadhyay, David Kammler, Gerd Ascheid, Rainer Leupers, Heinrich Meyr, Tim Kogel |
ASP-DAC | 6 |
| 2005 | Fine-grained application source code profiling for ASIP designabstractCurrent Application Specific Instruction set Processor (ASIP) design methodologies are mostly based on iterative architecture exploration that uses Architecture Description Languages (ADLs) and retargetable software development tools. However, for improved design efficiency, additional pre-architecture exploration tools are required to help narrow-down the huge design space and making coarsegrained Instruction Set Architecture (ISA) decisions before detailed ADL modeling. Extensive application code profiling is the key in such early design stages. Based on a novel code instrumentation technology, we present a microprofiling approach that fills the current gap between source-level and instruction-level profilers and combines their advantages w.r.t. speed and accuracy. We show how the microprofiler is embedded into an advanced ASIP design flow and justify its use in a case study to design an MP3 decoder ASIP. Kingshuk Karuri, Mohammad Abdullah Al Faruque, Stefan Kraemer, Rainer Leupers, Gerd Ascheid, Heinrich Meyr |
DAC | 6 |
| 2005 | C Compiler Retargeting Based on Instruction Semantics ModelsabstractEfficient architecture exploration and design of application specific instruction-set processors (ASIPs) requires retargetable software development tools, in particular C compilers that can be quickly adapted to new architectures. A widespread approach is to model the target architecture in a dedicated architecture description language (ADL) and to generate the tools automatically from the ADL specification. For C compiler generation, however, most existing systems are limited either by the manual retargeting effort or by redundancies in the ADL models that lead to potential inconsistencies. We present a new approach to retargetable compilation, based on the LISA 2.0 ADL with instruction semantics, that minimizes redundancies while simultaneously achieving a high degree of automation. The key of our approach is to generate the mapping rules needed in the compiler's code selector from the instruction semantics information. We describe the required analysis and generation techniques, and present experimental results for several embedded processors. Jianjiang Ceng, Manuel Hohenauer, Rainer Leupers, Gerd Ascheid, Heinrich Meyr, Gunnar Braun |
DATE | 5 |
| 2005 | A Modular Simulation Framework for Spatial and Temporal Task Mapping onto Multi-Processor SoC PlatformsabstractHeterogeneous multi-processor SoC (MP-SoC) platforms bear the potential to optimize conflicting performance, flexibility and energy efficiency constraints as imposed by demanding signal processing and networking applications. However, in order to take advantage of the available processing and communication resources, an optimal mapping of the application tasks on to the platform resources is of crucial importance. We propose a SystemC-based simulation framework, which enables the quantitative evaluation of application-to-platform mappings by means of an executable performance model. The key element of our approach is a configurable event-driven virtual processing unit to capture the timing behavior of multi-processor/multi-threaded MP-SoC platforms. The framework features an XML-based declarative construction mechanism of the performance model to accelerate navigation significantly in large design spaces. The capabilities of the proposed framework in terms of design space exploration is presented by a case study of a commercially available MP-SoC platform for networking applications. Focussing on the application to architecture mapping, our introduced framework highlights the potential for optimization of an efficient design space exploration environment. Torsten Kempf, Malte Doerper, Rainer Leupers, Gerd Ascheid, Heinrich Meyr, Tim Kogel, Bart Vanthournout |
DATE | 5 |
| 2005 | Improving MIMO phase noise estimation by exploiting spatial correlationsabstractPhase locked loops (PLL) for RF carrier synthesis often employ oscillators that insert a considerable amount of time varying phase noise into the received signal. That noise must then be removed in the digital baseband receiver. This phase noise is an indivisible superposition of noise components from receiver and transmitter. Regarding systems with multiple transmit and receive antennas (MIMO) and if multiple PLL for carrier synthesis are used each of the superposed phase noise processes per transmit and receive antenna pair can be measured at the receiver. This paper provides a new scheme for high SNR scenarios that exploits spatial correlation between these overlaying phase noise processes at the receiver in order to improve estimation and compensation of the phase noise. Therefore the Wiener filter approach is applied. Niels Hadaschik, Meik Dörpinghaus, Andreas Senst, Ole Harmjanz, Uwe Käufer, Gerd Ascheid, Heinrich Meyr |
ICASSP (3) | 7 |
| 2004 | A novel approach for flexible and consistent ADL-driven ASIP designabstractArchitecture description languages (ADL) have been established to aid the design of application-specific instruction-set processors (ASIP). Their main contribution is the automatic generation of a software toolkit, including C compiler, assembler, linker, and instruction-set simulator. Hence, the challenge in the design of such ADLs is to unambiguously capture the architectural information required for the toolkit generation in a single model. This is particularly difficult for C compiler and simulator, as both require information about the instructions' semantics, however, while the C compiler needs to know what an instructions does, the simulator needs to know how. Existing ADLs solve this problem by either introducing redundancy or by limiting the language's flexibility.This paper presents a novel, mixed-level approach for ADL-based instruction-set description, which offers maximum flexibility while preventing from inconsistencies. Moreover, it enables capturing instruction- and cycle-accurate descriptions in a single model. The feasibility and design efficiency of our approach is demonstrated with a number of contemporary, real-world processor architectures. Gunnar Braun, Achim Nohl, Weihua Sheng, Jianjiang Ceng, Manuel Hohenauer, Hanno Scharwächter, Rainer Leupers, Heinrich Meyr |
DAC | 8 |
| 2004 | Heterogeneous MP-SoC: the solution to energy-efficient signal processingabstractTo meet conflicting flexibility, performance and cost constraints of demanding signal processing applications, future designs in this domain will contain an increasing number of application specific programmable units combined with complex communication and memory infrastructures. Novel architecture trends like Application Specific Instruction-set Processors (ASIPs) as well as customized buses and Network-on-Chip based communication promise enormous potential for optimization. However, state-of-the-art tooling and design practice is not in a shape to take advantage of this advances in computer architecture and silicon technology. Currently, EDA industry develops two diverging strategies to cope with the design complexity of such application specific, heterogeneous MP-SoC platforms. First, the IP-driven approach emphasizes the composition of MP-SoC platforms from configurable off-the-shelf Intellectual Property blocks. On the other hand, the design-driven approach strives to take design efficiency to the required level by use of system level design methodologies and IP generation tools. In this paper, we discuss technical and economical aspects of both strategies. Based on the analysis of recent trends in computer architecture and system level design, we envision a hand-in-hand approach of signal processing platform architectures and design metholodgy to conquer the complexity crisis in emerging MP-SoC developments. Tim Kogel, Heinrich Meyr |
DAC | 2 |
| 2004 | A Methodology and Tool Suite for C Compiler Generation from ADL Processor ModelsabstractRetargetable C compilers are key tools for efficient architecture exploration for embedded processors. In this paper we describe a novel approach to retargetable compilation based on LISA, an industrial processor modeling language for efficient ASIP design. In order to circumvent the well-known trade-off between flexibility and code quality in retargetable compilation, we propose a user-guided, semiautomatic methodology that in turn builds on a powerful existing C compiler design platform. Our approach allows to include generated C compilers into the ASIP architecture exploration loop at an early stage, thereby allowing for a more efficient design process and avoiding application/architecture mismatches. We present the corresponding methodology and tool suite and provide experimental data for two real-life embedded processors that prove the feasibility of the approach. Manuel Hohenauer, Hanno Scharwächter, Kingshuk Karuri, Oliver Wahlen, Tim Kogel, Rainer Leupers, Gerd Ascheid, Heinrich Meyr, Gunnar Braun, Hans van Someren 0001 |
DATE | 8 |
| 2004 | RTL Processor Synthesis for Architecture Exploration and ImplementationabstractArchitecture description languages are widely used to perform architecture exploration for application-driven designs, whereas the RT-level is the commonly accepted level for hardware implementation. For this reason, design parameters such as timing, area or power consumption cannot be taken into consideration accurately during design space exploration. Design automation tools currently used to bridge this gap are either limited in the flexibility provided or only generate fragments of the architecture. This paper presents a synthesis tool which preserves the full flexibility of the architecture description language LISA, while being able to generate the complete architecture on RT-level using systemC. This paper also presents two real world architecture case studies to prove the feasibility of our approach. Oliver Schliebusch, Anupam Chattopadhyay, Rainer Leupers, Gerd Ascheid, Heinrich Meyr, Mario Steinert, Gunnar Braun, Achim Nohl |
DATE | 5 |
| 2004 | A System Level Processor/Communication Co-Exploration Methodology for Multi-Processor System-on-Chip PlatformabstractCurrent and future SoC designs will contain an increasing number of heterogeneous programmable units combined with a complex communication architecture to meet flexibility, performance and cost constraints. Designing such a heterogenous MP-SoC architecture bears enormous potential for optimization, but requires a system-level design environment and methodology to evaluate architectural alternatives. This paper proposes a methodology to jointly design and optimize the processor architecture together with the on-chip communication based on the LISA Processor Design Platform in combination with systemC transaction level models. The proposed methodology advocates a successive refinement flow of the system-level models of both the processor cores and the communication architecture. This allows design decisions based on the best modeling efficiency, accuracy and simulation performance possible on the respective abstraction level. The effectiveness of our approach is demonstrated by the exemplary design of a dual-processor JPEG decoding system. Andreas Wieferink, Tim Kogel, Rainer Leupers, Gerd Ascheid, Heinrich Meyr, Gunnar Braun, Achim Nohl |
DATE | 5 |
| 2004 | Performance of initial synchronization schemes for WCDMA systems with spatio-temporal correlationsabstractTwo different power-scaled noncoherent detection schemes are analyzed and compared in terms of detection probability and mean detection time in the uplink of a wideband CDMA system, where the mobile terminal transmits a pilot signal in the form of bursts of modulated chips, which are transmitted periodically and separated by long silent intervals. Both detection schemes employ temporal noncoherent averaging but differ in the way of spatial processing. One of the schemes, which is well suited for spatially uncorrelated scenarios, employs spatial noncoherent averaging, whereas the other scheme, better suited for scenarios with a distinct spatial structure, employs fixed beamforming. The performance analysis is carried out for spatially correlated frequency-selective fading channels taking channel dynamics and initial frequency offsets into account. With the presented analysis it is possible to investigate the effects of spatial correlation on the detection performance and to examine the improvement of using multiple antenna elements at the base station depending on the respective detection scheme. Lars Schmitt, Thomas Grundler, Christoph Schreyoegg, Gerd Ascheid, Heinrich Meyr |
ICC | 5 |
| 2004 | Random beamforming in correlated MISO channels for multiuser systemsabstractWe examine the technique of random beamforming to exploit multiuser diversity in the downlink of a wireless cellular communication system with an antenna array at the basestation. In random beamforming systems, the scalar signal is multiplied by a weight vector and the resulting signal vector is transmitted by the antenna array. By varying the weight vector in time, one can increase the dynamic of the effective channel resulting in faster fading and a larger variance of the effective channel. This can be exploited by a scheduler at the basestation which selects users for transmission that momentarily have a good channel. Our work focusses on the generation of the weight vectors for a uniform linear array in the basestation. We especially consider correlation between the antenna elements, but assume that the correlation matrices for the different users are not known to the basestation. Three random beamforming approaches are compared with the simplest possible case of using equal weights, where the weight vector is a constant which serves only to normalize the radiated power. We derive the mean of the received power for the different beamforming techniques and confirm our results by Monte Carlo simulations, where we evaluate the mean power of the actually scheduled user, if a proportional fair scheduling algorithm is used. One important result is the good performance of the rotating beam approach even in a not fully correlated environment. Andreas Senst, Peter Schulz-Rittich, Ulrich Krause, Gerd Ascheid, Heinrich Meyr |
ICC | 5 |
| 2004 | ASIP Architecture Exploration for Efficient Ipsec Encryption: A Case Study
Hanno Scharwächter, David Kammler, Andreas Wieferink, Manuel Hohenauer, Kingshuk Karuri, Jianjiang Ceng, Rainer Leupers, Gerd Ascheid, Heinrich Meyr |
SCOPES | 9 |
| 2004 | A universal technique for fast and flexible instruction-set architecture simulationabstractToday, designers of next-generation embedded processors and software are increasingly faced with short product lifetimes. The resulting time-to-market constraints are contradicting the continually growing processor complexity. Nevertheless, an extensive design-space exploration and product verification is indispensable for a successful market launch. In the last decade, instruction-set simulators have become an essential development tool for the design of new programmable architectures. Consequently, the simulator performance is a key factor for the overall design efficiency. Motivated by the extremely poor performance of commonly used interpretive simulators, research work on fast compiled instruction-set simulation was started ten years ago. However, due to the restrictiveness of the compiled technique, it has not been able to push through in commercial products. In this paper, we tie up with our previous research on retargetable, compiled simulation techniques, and provide a discussion about their benefits and limitations using a particular compiled scheme, static scheduling, as an example. As a conclusion, we eventually present a novel retargetable simulation technique, which combines the performance of traditional compiled simulators with the flexibility of interpretive simulation. This technique is not limited to any class of architectures or applications and can be utilized from architecture exploration up to end-user software development. We demonstrate workflow and applicability of the so-called just-in-time cache-compiled simulation technique by means of state-of-the-art real-world architectures. Gunnar Braun, Achim Nohl, Andreas Hoffmann 0002, Oliver Schliebusch, Rainer Leupers, Heinrich Meyr |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2003 | Instruction encoding synthesis for architecture exploration using hierarchical processor modelsabstractThis paper presents a novel instruction encoding generation technique for use in architecture exploration for application specific processors. The underlying exploration methodology is based on successive processor model refinement combined with simulation and profiling. Previous approaches require the tedious manual specification of binary instruction opcodes even at very early design stages due to the need to generate profiling tools. The proposed automatic technique eliminates this bottleneck in ASIP design. It is well adapted to the hierarchical processor modeling style of contemporary architecture description languages. Experimental evaluation for several real-life processor architectures confirms the practical applicability of the presented encoding techniques. Moreover, the results indicate that very compact instruction encoding schemes are generated that compete very well with hand-optimized encodings. Achim Nohl, Volker Greive, Gunnar Braun, Andreas Hoffmann 0002, Rainer Leupers, Oliver Schliebusch, Heinrich Meyr |
DAC | 7 |
| 2003 | Processor/Memory Co-Exploration on Multiple Abstraction Levels
Gunnar Braun, Andreas Wieferink, Oliver Schliebusch, Rainer Leupers, Heinrich Meyr, Achim Nohl |
DATE | 5 |
| 2003 | Initial synchronization of W-CDMA systems using a power-scaled detector with antenna diversity in frequency-selective Rayleigh fading channelsabstractAn analytical evaluation of the performance in terms of detection probability and mean detection time of a noncoherent detector at the base station is presented for a wideband CDMA system, where the mobile terminal transmits a pilot signal in the form of bursts of modulated chips, which are transmitted periodically and separated by long silent intervals. The performance analysis is carried out for a power-scaled detector employing temporal and spatial noncoherent averaging in frequency-selective fading channels taking channel dynamics and initial frequency offsets into account. With the presented analysis, which is verified by means of simulations, it is possible, and convenient, to trade off temporal noncoherent versus coherent averaging, depending on the length of the pilot bursts and channel dynamics and to examine the improvement of using multiple diversity antennas at the base station. Lars Schmitt, Volker Simon, Thomas Grundler, Christoph Schreyoegg, Heinrich Meyr |
GLOBECOM | 5 |
| 2003 | The effect of imperfect SNR knowledge on multiantenna multiuser systems with channel aware schedulingabstractWe analyze a cellular communication system in which a basestation (BTS) or access point transmits packet data to several mobile data users by means of a TDMA scheme. All users estimate their instantaneous signal-to-noise-ratio (SNR) in each slot and feed this information back to the BTS. A scheduler in the BTS then uses this information to allocate the channel resource to the user which maximizes a certain metric. We are interested in assessing the sensitivity of the system performance in terms of spectral efficiency per cell with respect to an imperfect knowledge of the multiuser channel, expressed by the estimated SNR of all users. By assuming a block fading channel model for each user, data-aided maximum-likelihood intra-slot SNR estimation can be performed if known pilot symbols are transmitted in each slot. We derive a novel SNR estimator for the block fading channel, which takes the channel statistics into account. The new estimator clearly outperforms the AWGN ML estimator in terms of system performance. Not only is the spectral efficiency larger for the system, but the optimum spectral efficiency is also achieved with fewer pilot symbols per slot. Peter Schulz-Rittich, Andreas Senst, Thomas Bilke, Heinrich Meyr |
GLOBECOM | 4 |
| 2003 | Synchronization requirements for COFDM systems with transmit diversityabstractCOFDM in combination with transmit and receive diversity is an attractive technique to increase robustness and throughput of wireless data transmission. In this paper we examine whether the sensitivity of COFDM to synchronization errors is further aggravated by these multiinput multioutput (MIMO) scenarios. A MIMO transmission model including synchronization errors is developed. Analytical expressions for the resulting disturbances are derived and verified against simulations. An analysis of the effects on optimal MIMO decoding shows that the resulting disturbances can no longer be modelled as white Gaussian noise. Michael Speth, Heinrich Meyr |
GLOBECOM | 2 |
| 2003 | Extraction of Efficient Instruction Schedulers from Cycle-True Processor Models
Oliver Wahlen, Manuel Hohenauer, Gunnar Braun, Rainer Leupers, Gerd Ascheid, Heinrich Meyr, Xiaoning Nie |
SCOPES | 6 |
| 2002 | A universal technique for fast and flexible instruction-set architecture simulationabstractIn the last decade, instruction-set simulators have become an essential development tool for the design of new programmable architectures. Consequently, the simulator performance is a key factor for the overall design efficiency. Based on the extremely poor performance of commonly used interpretive simulators, research work on fast compiled instruction-set simulation was started ten years ago. However, due to the restrictiveness of the compiled technique, it has not been able to push through in commercial products. This paper presents a new retargetable simulation technique which combines the performance of traditional compiled simulators with the flexibility of interpretive simulation. This technique is not limited to any class of architectures or applications and can be utilized from architecture exploration up to end-user software development. The work-flow and the applicability of the so-called just-in-time cache compiled simulation (JIT-CCS) technique will be demonstrated by means of state of the art real world architectures. Achim Nohl, Gunnar Braun, Oliver Schliebusch, Rainer Leupers, Heinrich Meyr, Andreas Hoffmann 0002 |
DAC | 5 |
| 2002 | Low complexity high resolution subspace-based delay estimation for DS-CDMAabstractWe consider the problem of estimating the propagation delays of a synchronous direct-sequence code division multiple access (DS-CDMA) system operating over a multipath fading channel. In a mobile receiver, the task of delay estimation can be divided into an acquisition of all multipath delays and subsequent tracking of the individual delays e.g. with a RAKE structure. Both tasks are especially challenging in indoor scenarios which commonly exhibit low delay-spread and thus require a high resolution of the estimation algorithms. A novel delay acquisition algorithm is presented, which is able to resolve multipaths whose delay difference is below one chip duration with high probability of acquisition and low computational complexity. It is based on a decomposition of the time-averaged correlation matrix of the output of a sliding correlator into signal and noise subspaces, with subsequent MUSIC spectrum computation and maximum search. The performance of the algorithm is assessed by means of computer simulations. Gunnar Fock, Peter Schulz-Rittich, Andreas Schenke, Heinrich Meyr |
ICC | 4 |
| 2002 | Designing SoC'sabstractNo abstract available. Tobias Noll, Heinrich Meyr |
ISLPED | 2 |
| 2001 | Fast Bit-True SimulationabstractThis paper presents a design environment which enables fast simulation of fixed-point signal processing algorithms. In contrast to existing approaches which use C/C++ libraries for the emulation of generic fixed-point data types, this novel approach additionally permits a code transformation to integral data types for fast simulation of the bit-true behavior. A speedup by a factor of 20 to 400 can be achieved compared to library based simulation. Holger Keding, Martin Coors, Olaf Lüthje, Heinrich Meyr |
DAC | 4 |
| 2001 | A framework for fast hardware-software co-simulationabstractWe present a new hardware-software co-simulation framework enabling fast prototyping in system-on-chip designs. On the software side, the machine description language LISA allows the generation of bit-true models of programmable architectures on various levels-from instruction-set to phase accuracy. Based on these models, a complete tool-suite consisting of fast compiled processor simulator assembler, linker HLL-compiler as well as co-simulation interface can be generated automatically. On the hardware side, the SystemC simulation class library is employed and enhanced with our generic co-simulation interface that enables the coupling of hardware and software models specified at various levels of abstraction. Besides that, a hardware modeling strategy using abstract macro-cycle based C++ processes to increase hardware modeling efficiency and simulation speed is presented. Andreas Hoffmann 0002, Tim Kogel, Heinrich Meyr |
DATE | 3 |
| 2001 | Generating production quality software development tools using a machine description languageabstractThis paper presents a methodology to automatically generate production quality software development tools for programmable architectures using the machine description language LISA. Various architectures presenting diverse architectural originalities will be presented and the feasibility of automatically generating simulator, assembler, linker and graphical debugger frontend are discussed. The presented approach is not limited to a fixed abstraction level-case studies of the Texas Instruments C62x and C54x, the Analog Devices ADSP2101 as well as the ARM7 show the applicability of the methodology from cycle/phase to instruction accurate models. Andreas Hoffmann 0002, Achim Nohl, Stefan Pees, Gunnar Braun, Heinrich Meyr |
DATE | 5 |
| 2001 | The programmable platform: does one size fit all?
A. Lock, Raúl Camposano, Heinrich Meyr |
DATE | 3 |
| 2001 | Integer code generation for the TI TMS320C62XabstractThis paper presents a methodology which enables the generation of C62/spl times/ optimized fixed-point C-code from a floating-point description of an algorithm. The FRIDGE design environment transforms floating-point ANSI-C code with local fixed-point annotations into an internal bit-true representation. From this representation we generate C62/spl times/ optimized integer C code utilizing the code transformation techniques illustrated in this paper. A benchmark is presented comparing the efficiency of the generated code with C67/spl times/ C-code, C62/spl times/ floating-point emulation and generic integer ANSI-C code. Martin Coors, Holger Keding, Olaf Lüthje, Heinrich Meyr |
ICASSP | 4 |
| 2001 | A survey on modeling issues using the machine description language LISAabstractThis paper presents a survey on modeling issues of programmable architectures using the machine description language LISA. Various architectures presenting diverse architectural characteristics are presented and the feasibility of automatically generating simulator, assembler, linker and graphical debugger frontend discussed. The presented approach is not limited to a fixed abstraction level-case studies of the Texas Instruments C62/spl times/ and C54/spl times/, the Analog Devices ADSP2101 as well as the ARM7 show the applicability of the methodology from cycle/phase to instruction accurate models. Andreas Hoffmann 0002, Achim Nohl, Gunnar Braun, Heinrich Meyr |
ICASSP | 4 |
| 2001 | A Methodology for the Design of Application Specific Instruction Set Processors (ASIP) using the Machine Description Language LISAabstractThe development of application specific instruction set processors (ASIP) is currently the exclusive domain of the semiconductor houses and core vendors. This is due to the fact that building such an architecture is a difficult task that requires expert knowledge in different domains: application software development tools, processor hardware implementation, and system integration and verification. This paper presents a retargetable framework for ASIP design which is based on machine descriptions in the LISA language. From that, software development tools can be automatically generated including HLL C-compiler, assembler, linker, simulator and debugger frontend. Moreover, synthesizable HDL code can be derived which can then be processed by standard synthesis tools. Implementation results for a low-power ASIP for DVB-T acquisition and tracking algorithms designed with the presented methodology are given. Andreas Hoffmann 0002, Oliver Schliebusch, Achim Nohl, Gunnar Braun, Oliver Wahlen, Heinrich Meyr |
ICCAD | 6 |
| 2001 | Achievable rate of MIMO channels with data-aided channel estimationabstractThe achievable rate of a coherent coded modulation (CM) digital communication system with data-aided channel estimation and a discrete, equiprobable symbol alphabet is derived under the assumption that the system operates on a flat fading MIMO channel and uses an interleaver to combat the bursty nature of the channel. It is shown that linear minimum mean square error (LMMSE) channel estimation directly follows from the derivation, and links average mutual information to the channel dynamics. Based on the assumption that known training symbols are transmitted, the achievable rate of the system is optimized with respect to the amount of training information needed. Jens Baltersee, Gunnar Fock, Heinrich Meyr |
ITW | 3 |
| 2001 | Achievable rate of MIMO channels with data-aided channel estimation and perfect interleavingabstractThe achievable rate of a coherent coded modulation digital communication system with data-aided channel estimation and a discrete equiprobable symbol alphabet is derived under the assumption that the system operates on a flat fading multiple-input/multiple-output channel and uses a perfect interleaver to combat the bursty nature of the channel. It is shown that linear minimum mean square error channel estimation directly follows from the derivation and links average mutual information to the channel dynamics. Based on the assumption that known training symbols are transmitted, the achievable rate of the system is optimized with respect to the amount of training information needed. Jens Baltersee, Gunnar Fock, Heinrich Meyr |
IEEE J. Sel. Areas Commun. | 3 |
| 2001 | Channel tracking for RAKE receivers in closely spaced multipath environmentsabstractThis paper deals with the problem of channel tracking for RAKE receivers in propagation environments characterized by closely spaced multipath components. After outlining why conventional single-path channel tracking algorithms fail in such scenarios, several new estimation algorithms are developed that are tailored to channels with closely spaced multipaths. This is achieved by removing or minimizing self-interference caused by multipath components. Other interfering users are treated as noise. Both timing tracking and phasor tracking and their interaction are covered in this paper. The derived algorithms are benchmarked against perfect channel knowledge on one hand and conventional tracking algorithms on the other hand, both in a UMTS test scenario. In moderate scenarios, the use of these new algorithms leads to performance improvements of up to 2 dB, in terms of signal-to-noise ratio (SNR) at moderate bit error rates, and even manages to track the channel in conditions where conventional tracking algorithms fail completely. Gunnar Fock, Jens Baltersee, Peter Schulz-Rittich, Heinrich Meyr |
IEEE J. Sel. Areas Commun. | 4 |
| 2001 | A novel methodology for the design of application-specificinstruction-set processors (ASIPs) using a machine description languageabstractThe development of application-specific instruction-set processors (ASIP) is currently the exclusive domain of the semiconductor houses and core vendors. This is due to the fact that building such an architecture is a difficult task that requires expertise in different domains: application software development tools, processor hardware implementation, and system integration and verification. This paper presents a retargetable framework for ASIP design which is based on machine descriptions in the LISA language. From that, software development tools can be generated automatically including high-level language C compiler, assembler, linker, simulator, and debugger frontend. Moreover, for architecture implementation, synthesizable hardware description language code can be derived, which can then be processed by standard synthesis tools. Implementation results for a low-power ASIP for digital video broadcasting terrestrial acquisition and tracking algorithms designed with the presented methodology are given. To show the quality of the generated software development tools, they are compared in speed and functionality with commercially available tools of state-of-the-art digital signal processor and /spl mu/C architectures. Andreas Hoffmann 0002, Tim Kogel, Achim Nohl, Gunnar Braun, Oliver Schliebusch, Oliver Wahlen, Andreas Wieferink, Heinrich Meyr |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 8 |
| 2001 | An information theoretic foundation of synchronized detectionabstractThe constrained capacity of a coherent coded modulation (CM) digital communication system with data-aided channel estimation and a discrete, equiprobable symbol alphabet is derived under the assumption that the system operates on a flat fading channel and uses an interleaver to combat the bursty nature of the channel. It is shown that linear minimum mean square error channel estimation directly follows from the derivation and links average mutual information to the channel dynamics. Based on the assumption that known training symbols are transmitted, the achievable rate of the system is optimized with respect to the amount of training information needed. Furthermore, the results are compared to the additive white Gaussian noise channel, and the case when ideal channel state information is available at the receiver. Jens Baltersee, Gunnar Fock, Heinrich Meyr |
IEEE Trans. Commun. | 3 |
| 2001 | Optimum receiver design for OFDM-based broadband transmission .II. A case studyabstractThis paper details on the design of OFDM receivers. Special attention is paid to the OFDM-specific receiver functions necessary to demodulate the received signal and deliver soft information to the outer receiver for decoding. In part I of the paper, the effects of nonideal transmission conditions have been thoroughly analyzed. To show the impact of the synchronization algorithms-which are most critical in OFDM-on system performance and complexity we consider the design of a complete receiver consisting of symbol synchronization, carrier/sampling clock synchronization and channel estimation. The performance of the algorithms is analyzed and a qualitative estimate of the resulting complexity is given. This allows one to draw conclusions concerning the achievable system performance under realistic complexity assumptions. Michael Speth, Stefan A. Fechtel, Gunnar Fock, Heinrich Meyr |
IEEE Trans. Commun. | 4 |
| 2000 | Efficient building block based RTL code generation from synchronous data flow graphsabstractThis paper presents a RTL-HDL code generation from synchronous data-flow graphs which supports the building block based design of data-flow oriented ASIC systems. Here, additional interfacing and controlling hardware is generated to adapt non-matching interfacing properties. In order to reduce interface register cost, a retiming approach is taken to schedule optimum building block activation times. The code generation methodology is compared to an existing approach using different case studies. Jens Horstmannshoff, Heinrich Meyr |
DAC | 2 |
| 2000 | Retargeting of Compiled Simulators for Digital Signal Processors Using a Machine Description LanguageabstractThis paper presents a methodology to retarget the technique of compiled simulation for digital signal processors (DSPs) using the modeling language LISA. In the past, the principle of compiled simulation as means for speeding up simulators has only been implemented for specific DSP architectures. The new approach presented here discusses methods of integrating compiled simulation techniques to retargetable simulation tools. The principle and the implementation are discussed in this paper and results for the TI TMS320C6201 DSP are presented. Stefan Pees, Andreas Hoffmann 0002, Heinrich Meyr |
DATE | 3 |
| 2000 | DSP core verification using automatic test case generationabstractThe verification methodology for a TMS320C25 compatible embedded DSP core is described. The DSP core has been implemented in synthesizable VHDL and has been cosimulated with the original DSP to verify correct behavior. Automatic test case generation together with hand-crafted code has been used as a means of providing stimuli to achieve increased RTL-simulation coverage. The cosimulation environment for this verification and the process of automatic test case generation is described in detail. Experimental results in terms of simulation coverage are discussed. Finally, a classification of all identified design flaws in the implementation is given and error-prone parts of the HDL design are identified. Tilman Glökler, Stefan Bitterlich, Heinrich Meyr |
ICASSP | 3 |
| 2000 | Iterative Multiuser Detection for Bit Interleaved Coed ModulationabstractIn this paper we present an approach for iterative decoding of bit interleaved coded modulation (BICM) in multi-transmitter multi-receiver scenarios. Conventional decoding methods yield results that are still way from the theoretical optimum. As a possible approach to achieve optimum performance multiuser detection via iterative cancellation is pursued. The main problem when applying this approach to BICM are cancellation errors, leading to a saturated performance. To overcome this behavior, soft-bits taking into account the reliability of the cancellation must be generated. We develop the optimum metric for this case. Using this metric we are able to approach optimum performance with few iterations. In a further step the metric is simplified yielding a complexity that is only fractions of the optimum approach. Michael Speth, Alexander Jansen, Heinrich Meyr |
ICC (2) | 3 |
| 2000 | Retargetable compiled simulation of embedded processors using a machine description languageabstractFast processor simulators are needed for the software development of embedded processors, for HW/SW cosimulation systems, and for profiling and design of application-specific processors. Such fast simulators can be generated based on the machine description language LISA. Using this language to model processor architectures enables the generation of compiled simulators on various abstraction levels, assemblers, and compiler back ends. The article discusses the requirements of software development tools on processor models and presents the approach based on the LISA language. Furthermore, the implementation of a retargetable environment consisting of compiled simulator, debugger, and assembler is presented. Measurements for a verified, cycle-based LISA model of the TI TMS320C62× DSP show that that this approach achieves between 37× and 170× higher simulation speed compared to a commercial simulator using a standard technique and the same accuracy level. Stefan Pees, Andreas Hoffmann 0002, Heinrich Meyr |
ACM Trans. Design Autom. Electr. Syst. | 3 |
| 1999 | LISA - Machine Description Language for Cycle-Accurate Models of Programmable DSP ArchitecturesabstractThis paper presents the machine description language LISA for the generation of bitand cycle accurate models of DSP processors. Based on a behavioral operation description, the architectural details and pipeline operations of modern DSP processors can be covered. Beyond the behavioral model, LISA descriptions include other architecture-related information like the instruction set. The information provided by LISA models enables automatic generation of simulators and assemblers which are essential elements of DSP software development environments. In order to proof the applicability of our approach, a realized model of the Texas Instruments TMS320C6201 DSP is presented and derived LISA code examples are given. Stefan Pees, Andreas Hoffmann 0002, Vojin Zivojnovic, Heinrich Meyr |
DAC | 4 |
| 1999 | Optimum receiver design for wireless broad-band systems using OFDM. IabstractOrthogonal frequency-division multiplexing (OFDM) is the technique of choice in digital broad-band applications that must cope with highly dispersive transmission media at low receiver implementation cost. In this paper, we focus on the inner OFDM receiver and its functions necessary to demodulate the received signal and deliver soft information to the outer receiver for decoding. The effects of relevant nonideal transmission conditions are thoroughly analyzed: imperfect channel estimation, symbol frame offset, carrier and sampling clock frequency offset, time-selective fading, and critical analog components. Through an appropriate optimization criterion (signal-to-noise ratio loss), minimum requirements on each receiver synchronization function are systematically derived. An equivalent signal model encompassing the effects of all relevant imperfections is then formulated in a generalized framework. The paper concludes with an outline of synchronization strategies. Michael Speth, Stefan A. Fechtel, Gunnar Fock, Heinrich Meyr |
IEEE Trans. Commun. | 4 |
| 1998 | FRIDGE: A Fixed-Point Design and Simulation EnvironmentabstractDigital systems, especially those for mobile applications are sensitive to power consumption, chip size and costs. Therefore they are realized using fixed-point architectures, either dedicated HW or programmable DSPs. On the other hand, system design starts from a floating-point description. These requirements have been the motivation for FRIDGE, a design environment for the specification, evaluation and implementation of fixed-point systems. FRIDGE offers a seamless design flow from a floating-point description to a fixed point implementation. Within this paper we focus on two core capabilities of FRIDGE: (1) the concept of an interactive, automated transformation of floating-point programs written in ANSI-C into fixed-point specifications, based on an interpolative approach. The design time reductions that can be achieved make FRIDGE a key component for an efficient HW/SW-codesign; (2) a fast fixed-point simulation that performs comprehensive compile-time analyses, reducing simulation time by one order of magnitude compared to existing approaches. Holger Keding, Markus Willems, Martin Coors, Heinrich Meyr |
DATE | 4 |
| 1997 | Mapping multirate dataflow to complex RT level hardware modelsabstractThe design of digital signal processing systems typically consists of an algorithm development phase carried out at a behavioral level and the selection of an efficient hardware architecture for implementation. In order to speed up the joint optimization of algorithms and architectures, a fast path to implementation must be provided. This can be achieved efficiently by directly mapping the data flow specification of the system to an RTL target architecture by means of HDL code generation. For algorithm design, communication systems are most easily modeled using multirate data flow graphs in which no notion of time is maintained. HDL code generation introduces a cycle based timing model and maps the data flow models to RTL implementations, which are usually taken from a library. Due to the increase in ASIC design complexity, these building blocks reach a high level of functionality and have complex interfacing properties. Therefore, it becomes necessary to generate additional interfacing and controlling hardware to synthesize an operable system. In this paper, we present a new approach of mapping multirate dataflow graphs to complex RTL hardware models and derive algorithms to synthesize these high-level RTL building blocks into a complete operable system. Jens Horstmannshoff, Thorsten Grötker, Heinrich Meyr |
ASAP | 3 |
| 1997 | On core and more: a design perspective for systems-on-a-chipabstractIn this survey, key drivers in design methodology are provided that enable successful design of systems-on-a-chip for the highly competitive telecommunications market. Main components of a design environment are described that fulfill the requirements of today's system design: efficient verification by means of fast simulation, integration of intellectual property, support of HW/SW co-design by means of a generic machine description language, generation of dedicated hardware blocks for high speed applications, and the link from system level performance evaluation to implementations in hardware and software. Stefan Pees, Martin Vaupel, Vojin Zivojnovic, Heinrich Meyr |
ASAP | 4 |
| 1997 | System Level Fixed-Point Design Based on an Interpolative ApproachabstractThe design process for fixed-point implementations eitherin software or in hardware requires a bit-true specificationof the algorithm in order to analyze quantization effectson an algorithmical level, abstracting from implementationaldetails.On the other hand, system design starts froma floating-point description into a fixed-point description becomesnecessary.Within this paper we present a tool thatallows an automated, interactive transformation from floating-pointANSI-C into a bit-true specification based ona new data type fixed that is introduced as an extensionto ANSI-C.The concept is rooted in a sophisticated datadependency analysis that allows to handle control structuresas well as pointers.It is part of the fixed-point designenvironment FRIDGE which includes an advanced simulatorthat covers the extended ANSI-C syntax as well astarget specific compilers which allow to generate efficientfixed-point implementations either for HW or for SW, startingfrom the bit-true algorithm specification. Markus Willems, Volker Bürsgens, Holger Keding, Thorsten Grötker, Heinrich Meyr |
DAC | 5 |
| 1997 | Unified specification of control and data flowabstractMany signal processing systems use event driven mechanisms-typically based on finite state machines (FSMs)-to control the operation of computationally intensive (data flow) parts. The state machines in turn are often fueled by external inputs as well as by feedback from the signal processing portions of the system. Packet-based transmission systems are a good example for such a close interaction between data and control flow. For an efficient design flow it is of crucial importance to be able to model and analyze the complete functionality of the system within one single design environment. Therefore, we developed a computational model that integrates the specification of control and data flow by combining the notion of data flow graphs with event driven process activation. Thorsten Grötker, Rainer Schoenen, Heinrich Meyr |
ICASSP | 3 |
| 1997 | An upper bound of the throughput of multirate multiprocessor schedulesabstractMultirate Dataflow Graphs (MR-DFGs) are used for modelling iterative computations, allowing concurrency and arbitrary data rates at ports. This model is often used for signal processing algorithms. For static scheduling the iteration period bound represents the final barrier for the computation speed, the approximation of which is often the goal of an implementation. For the singlerate case (SR-DFG), where all rates are one, an explicit bound exists and is subject of many published papers. This work presents a bound for the multirate case, which reduces to the known bound if applied to an SR-DFG. Assumptions made are a vectorized execution and a blocked schedule that organizes multiple iterations inside one period (also called execution cycle). The influence of characteristic properties in the multirate case is emphasized and related to terms from the Petri-Nets theory. Rainer Schoenen, Vojin Zivojnovic, Heinrich Meyr |
ICASSP | 3 |
| 1997 | FRIDGE: an interactive code generation environment for HW/SW codesignabstractDigital mobile systems are sensitive to power consumption, chip size and costs. Therefore they are realized using fixed-point architectures, either dedicated hardware (HW) or fixed-point processors. On the other hand, system design starts from a floating-point description. These requirements have been the motivation for FRIDGE, a design environment for the specification, evaluation and implementation of fixed-point systems. FRIDGE offers a seamless design flow from a floating-point description to a fixed-point implementation. We focus on the FRIDGE-concept of an interactive, automated transformation of floating-point programs written in ANSI-C into fixed-point specifications, based on an interpolative approach. Since HW and software (SW) implementations of the same functionality in general require different fixed-point specifications, the design time reductions that can be achieved by using FRIDGE make it a key component for an efficient HW/SW-codesign. Markus Willems, Volker Bürsgens, Thorsten Grötker, Heinrich Meyr |
ICASSP | 4 |
| 1997 | Modulo-addressing utilization in automatic software synthesis for digital signal processorsabstractDigital Signal Processors (DSPs) have become key components for the implementation of digital signal processing systems. With DSPs moving into new application domains and the increasing complexity of modern DSP architectures, efficient programming support receives major interest. Therefore, an optimizing compiler becomes a must for future DSP-architectures. Todays DSP compilers result in significant overheads both in memory consumption and program execution time compared to hand-written assembly code. This is mainly due to an inefficient compiler support of the DSP specific architectural features, such as the modulo-addressing capability which is an enabling feature for a large class of DSP algorithms. Within this paper we analyze why existing compilers fail short in supporting the modulo-addressing mode and present a compiler concept that allows the efficient utilization of this feature. We describe how an advanced compiler optimization strategy allows a near optimum support of the modulo-addressing mode, and point out why this concept is favorable to DSP-specific language extensions. Markus Willems, Holger Keding, Vojin Zivojnovic, Heinrich Meyr |
ICASSP | 4 |
| 1997 | Authors' reply [to "Comment on cycle slips in synchronizers subject to smooth narrow-band loop noise"]abstractThe claims in the comment of Popken (see ibid., vol.45, no.1, p.19-20, 1997) are wrong. We point out that the discrepancy of the cycle slip rates obtained in our paper and that of Meyr, Popken and Mueller (see ibid., vol.COM-34, no.5, p.436-45, 1986) is not due to a misconception in our paper, but rather to a different modeling of the loop noise. In order to obtain cycle slip rates by means of Fokker-Planck techniques, Meyr et al. were forced to use an oversimplified model, yielding essentially flat loop noise. In our paper, cycle slip rates were obtained by means of level crossing techniques, while preserving the smooth narrowband nature of the loop noise. Marc Moeneclaey, Heinrich Meyr |
IEEE Trans. Commun. | 2 |
| 1996 | Compiled HW/SW Co-SimulationabstractThis paper presents a technique for simulating processors and attached hardware using the principle of compiled simulation. Unlike existing, inhouse and off-the-shelf hardware/software co-simulators, which use interpretive processor simulation, the proposed technique performs instruction decoding and simulation scheduling at compile time. The technique offers up to three orders of magnitude faster simulation. The high speed allows the user to explore algorithms and hardware/software trade-offs before any hardware implementation. In this paper, the sources of the speedup and the limitations of the technique are analyzed and the realization of the simulation compiler is presented. Vojin Zivojnovic, Heinrich Meyr |
DAC | 2 |
| 1996 | The Differential CORDIC Algorithm: Constant Scale Factor Redundant Implementation without Correcting IterationsabstractThe CORDIC algorithm is a well-known iterative method for the efficient computation of vector rotations, and trigonometric and hyperbolic functions. Basically, CORDIC performs a vector rotation which is not a perfect rotation, since the vector is also scaled by a constant factor. This scaling has to be compensated for following the CORDIC iteration. Since CORDIC implementations using conventional number systems are relatively slow, current research has focused on solutions employing redundant number systems which make a much faster implementation possible. The problem with these methods is that either the scale factor becomes variable, making additional operations necessary to compensate for the scaling, or additional iterations are necessary compared to the original algorithm. In contrast we developed transformations of the usual CORDIC algorithm which result in a constant scale factor redundant implementation without additional operations. The resulting "Differential CORDIC Algorithm" (DCORDIC) makes use of on-line (most significant digit first redundant) computation. We derive parallel architectures for the radix-2 redundant number systems and present some implementation results based on logic synthesis of VHDL descriptions produced by a DCORDIC VHDL generator. We finally prove that, due to the lack of additional operations, DCORDIC compares favorably with the previously known redundant methods in terms of latency and computational complexity. Herbert Dawid, Heinrich Meyr |
IEEE Trans. Computers | 2 |
| 1996 | A CMOS IC for Gb/s Viterbi decoding: system design and VLSI implementationabstractAt present, the Viterbi algorithm (VA) is widely used in communication systems for decoding and equalization. The achievable speed of conventional Viterbi decoders (VD's) is limited by the inherent nonlinear add-compare-select (ACS) recursion. The aim of this paper is to describe system design and VLSI implementation of a complex system of fabricated ASIC's for high speed Viterbi decoding using the "minimized method" (MM) parallelized VA. We particularly emphasize the interaction between system design, architecture and VLSI implementation as well as system partitioning issues and the resulting requirements for the system design flow. Our design objectives were 1) to achieve the same decoding performance as a conventional VD using the parallelized algorithm, 2) to achieve a speed of more than 1 Gb/s, and 3) to realize a system for this task using a single cascadable ASIC. With a minimum system configuration of four identical ASIC's produced by using 1.0 /spl mu/ CMOS technology, the design objective of a decoding speed of 1.2 Gb/s is achieved. This means, compared to previous implementations of Viterbi decoders, the speed is increased by an order of magnitude. Herbert Dawid, Gerhard P. Fettweis, Heinrich Meyr |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 1995 | Digital Receiver Design Using VHDL Generation from Data Flow GraphsabstractAbstract| This paper describes a design methodology, a library of reusable VHDL descriptions and a VHDL generation tool used in the application area of digital signal processing, particularly digital receivers for communication links.The tool and the library interact with commercial system simulation and logic synthesis tools.The support of joint optimization of algorithm and architecture as well as the concept for design reuse are explained.The algorithms for generating VHDL code according to dierent user speci cations are described.An application example is used to show the bene ts and current limitations of the proposed methodology. Peter Zepter, Thorsten Grötker, Heinrich Meyr |
DAC | 3 |
| 1995 | ADEN: an environment for digital receiver ASIC designabstractDifferent levels of abstraction are suited for algorithm design and hardware architecture development. This paper presents a tool (ADEN) that provides a link from system design to VLSI implementation. It generates synchronous timed descriptions of digital hardware from dynamic data-flow system level configurations. It allows one to make use of optimized architectures available for a broad range of communication system components. These components are kept in the extensible ComBox library which provides a means to characterize their data-flow and timing properties. The design methodology together with the tool operation and the library concept are explained. An actual design example is presented to demonstrate the effectiveness of this approach. Thorsten Grötker, Peter Zepter, Heinrich Meyr |
ICASSP | 3 |
| 1995 | DSP-based mobile and satellite receivers, from algorithm to implementation: a design course at Aachen University of TechnologyabstractProfound knowledge of the interaction between algorithms and digital signal processor (DSP) architectures is required to be able to efficiently design complex communications equipment. Whereas both algorithms and architecture find treatment in many courses individually, education focusing on design methodology for DSP implementation is found to be rare. This contribution describes a concept and its implementation of a design course for DSP-based mobile and satellite communications systems attempting to fill the described gap. To illustrate the proposed concept in more detail, examples of the fall 1994 course are given. Oliver Mauss, Matthias Pankert, Ferdinand Classen, Heinrich Meyr |
ICASSP | 4 |
| 1995 | Scheduling for optimum data memory compaction in block diagram oriented software synthesisabstractFor the design of complex digital signal processing systems, block diagram oriented synthesis of real time software for programmable target processors has become an important design aid. The synthesis approach discussed in the paper is based on multirate block diagrams with scalable synchronous dataflow (SSDF) semantics. For this class of dataflow graphs the authors present scheduling techniques for optimum data memory compaction. These techniques can be employed to map signals of a block diagram onto a minimum data memory space. In order to formalize the data memory compaction problem, they first derive appropriate implementation measures. Based on these implementation measures it can be shown that optimum data memory compaction consists of optimum scheduling as well as optimum memory allocation. For the class of single appearance (SA) block diagrams with SSDF semantics, scheduling can be reduced to an integer linear programming (ILP) problem. Due to the computational complexity of ILP, the authors also present a suboptimum scheduling selection criterion, which call be used for SA and non SA-schedulers. Sebastian Ritz, Markus Willems, Heinrich Meyr |
ICASSP | 3 |
| 1995 | Test Point Insertion for an Area Efficient BISTabstractWe present a new cost based method for the insertion of test points into sequential circuits. It is especially suited for the design of an area efficient BIST using pseudorandom patterns. Testability analysis is used to detect areas of poor testability and estimate the benefit of test points. The designer can trade this benefit against increased area. Experimental results show that small sets of test points are sufficient to reach a fault coverage greater than 99%. Claus Schotten, Heinrich Meyr |
ITC | 2 |
| 1995 | Real-time algorithms and VLSI architectures for soft output MAP convolutional decodingabstractSoft output decoding algorithms have been attracting considerable attention for concatenated or iterative decoding systems. The most prominent is the soft output Viterbi algorithm (SOVA), which can be regarded as an approximation to the optimum symbol by symbol detector, the symbol by symbol MAP algorithm. The MAP algorithm is a block detector, i.e. it requires completed reception of a terminated block of received symbols, which prevents an efficient real-time implementation. Therefore, up to now, the SOVA was considered to be the more attractive alternative. In this paper, we first introduce an algebraic formulation for the MAP. Using this formulation, a novel real-time MAP algorithm (SMAP) is derived, and it is proved that in terms of hard decoding performance, the SMAP is equivalent to the VA with path memory truncation and best state decoding. The finally presented SMAP architectures provide a competitive alternative to the known SOVA architectures for soft output decoding applications. Herbert Dawid, Heinrich Meyr |
PIMRC | 2 |
| 1995 | Viterbi decoding with dual timescale traceback processingabstractA new approach to traceback processing in Viterbi decoders is presented. The approach reduces memory requirements as compared to previous approaches by using different speeds during acquisition of the best trellis path and the subsequent decoding of a block of data. This dual timescale approach allows in-place updating of the stored information and matches the constraints of commodity semi-custom technologies, where (at the considered high dock rates) write accesses to a RAM usually require more than one clock cycle. Olaf J. Joeressen, Heinrich Meyr |
PIMRC | 2 |
| 1994 | Dynamic data flow and control flow in high level DSP code synthesisabstractThe design of today's complex digital signal processing systems, such as communication equipment, increasingly relies on sophisticated CAD tools for block diagram oriented analysis and simulation. Recent work has been concerned with the integration of such simulation tools and those for implementation, i.e. for hardware or software synthesis. Data flow oriented approaches have proven to be very well suited for both tasks due to the nature of most digital signal processing applications. The control of multiple cooperating data flow tasks (including resource management) is an important issue in most digital signal processor (DSP) based systems. This paper describes the integration of control flow into data flow oriented simulation and synthesis. A heterogeneous modeling scheme is proposed. The focus is kept on maintaining the efficiency and simplicity of the data paths while offering additional expressiveness which-as does the data flow paradigm-closely matches the way of thinking of communications system engineers. The usefulness of the novel concepts is demonstrated by a prototype implementation of a digital receiver for wireless data communication. > Matthias Pankert, Oliver Mauss, Sebastian Ritz, Heinrich Meyr |
ICASSP (2) | 4 |
| 1994 | Retiming of DSP programs for optimum vectorizationabstractVectorization of digital signal processing programs given in form of data-flow graphs (DFG) is treated. It is shown that for cyclic unit-rate graphs an inherent upper bound on the linear vectorization factor exists. Using the retiming transformation of the original graph this bound can be raised up to a transformation-independent bound which is in general not tight. The authors give a sufficient condition for efficient linear vectorization (vectorization approaching the bound) and propose a corresponding algorithm. As a side result, a useful theorem for retiming of strongly connected graphs is given.> Vojin Zivojnovic, Sebastian Ritz, Heinrich Meyr |
ICASSP (2) | 3 |
| 1994 | Is it Possible to achieve a Teraflop/s on a chip? From High Performance Algorithms to ArchitecturesabstractThe forumnists address the question of high density computations on a single chip. The surface of a chip offers an ideal medium not only to store information or to process data, but also to execute computations. The 1 Giga floating point operations per second per chip mark has been achieved, we are now moving towards the teraflop mark. How is this going to happen, what are the limitations, what are the opportunities-those are central questions.> Francky Catthoor, Ed F. Deprettere, Yu Hen Hu, Jan M. Rabaey, Heinrich Meyr, Lothar Thiele |
ISCAS | 5 |
| 1994 | High Speed FIR-Filter Architectures with Scalable Sample RatesabstractFIR (finite impulse response) filters are widely used in digital signal processing. In this paper new architectures for high speed FIR filters with programmable coefficients are presented. Special efforts are undertaken to develop a structure that is well suitable for different data rates and therefore may be used within a tool (filter generator) that generates demand driven dedicated filter structures. The presented structure leads to highly efficient designs that are useable within different environments. The basic design structure is introduced and implementation considerations are discussed. Results of synthesis runs are presented.> Martin Vaupel, Heinrich Meyr |
ISCAS | 2 |
| 1994 | Improved frame synchronization for spontaneous packet transmission over frequency-selective radio channelsabstractIn digital TDMA-based packet radio channel access schemes operating over slowly time-variant frequency-selective multipath radio channels, it is necessary to achieve fast initial acquisition of receiver synchronization. The authors are concerned with improving the frame sync performance of the ML-algorithm [Fechtel and Meyr, 1993] for feedforward joint frame sync and frequency offset estimation. Emphasis is placed upon improved "non-data-dependent" (NDD) reduced-complexity variants of the ML-estimator. This is achieved by abandoning the test for signal periodicity in favor of a different preamble structure that yields a more pronounced peak of the autocorrelation at the expense of a larger preamble length. The average frame error rate and frequency estimation performance of the new frame/frequency synchronizers is assessed via simulation. For typical examples, the new algorithms are shown to yield frame sync performance gains of several decibels so that asynchronous packet transmission in the lower SNR regions becomes feasible. Stefan A. Fechtel, Heinrich Meyr |
PIMRC | 2 |
| 1994 | Frequency synchronization algorithms for OFDM systems suitable for communication over frequency selective fading channelsabstractIn this paper, the problem of carrier synchronization of OFDM systems in the presence of a substantial frequency offset is considered. New frequency estimation algorithms for the data aided (DA) mode are presented. The resulting two stage structure is able to cope with frequency offsets in the order of multiples of the spacing between subchannels. Key features of the novel scheme-which are presented in terms of estimation error variances, the required amount of training symbols and the computational load-ensure high speed synchronization with negligible decoder performance degradation at a low implementation effort.> Ferdinand Classen, Heinrich Meyr |
VTC | 2 |
| 1994 | A digital feedforward differential detection MSK receiver for packet-based mobile radioabstractWe present a differential MSK receiver with coarse frame synchronization using a RSSI (received signal strength indicator) signal and thus allowing the use of the capture phenomenon. The capture event is detected by an ML-derived circuit. Our investigations include all synchronization units allowing a tradeoff between preamble length and implementation efficiency. We provide a performance study including the effect of quantization. The results prove the suitability of our concept for a mobile communication network.> Uwe Lambrette, Heinrich Meyr |
VTC | 2 |
| 1994 | Optimal parametric feedforward estimation of frequency-selective fading radio channelsabstractEstimation of time-varying dispersive radio channels is one of the most important tasks of receiver synchronisation. State-of-the-art adaptive estimators employing decision feedback suffer from error propagation and limited robustness against faster fading. In this paper, the optimal feedforward channel estimator using only known training symbols is systematically derived. Statistical channel information is assumed to be available. For moderately rapid fading channels (snapshot assumption), the optimal estimator is shown to be divided into the two subtasks of maximum-likelihood acquisition and subsequent Wiener filtering of the acquired quantities. If perfect training sequences resulting in optimal performance are used, the Wiener filtering operation collapses into independent filtering of the individual acquired tap estimates, and the resulting channel estimator becomes efficient and flexible. The performance of the optimal estimator is evaluated for important cases. A design example of a near-optimal HF channel estimator/receiver is discussed. Linear interpolation is used to reduce the computational complexity of receiver coefficient adjustment. Simulation results confirm the superiority and robustness of feedforward synchronization having near-optimal performance at reduced complexity.> Stefan A. Fechtel, Heinrich Meyr |
IEEE Trans. Commun. | 2 |
| 1994 | On sampling rate, analog prefiltering, and sufficient statistics for digital receiversabstractWe consider the joint sequence estimation, timing and phase recovery for linear modulation. The paper differs from the classical ones in the sense that time-discrete algorithms suitable for fully digital receivers are discussed. Sufficient conditions are given such that the signal samples represent sufficient statistics. These conditions involve signal bandwidth, sampling/symbol rate and the analog prefilter characteristics. It is shown that the sampling rate need not be an exact multiple of the symbol rate, i.e., the samples can be taken from a free-running oscillator. All subsequent signal processing operations in the receiver then operate with the clock of this free-running oscillator. Timing recovery is then performed by a time-variant linear digital interpolator and a decimator. Carrier recovery and sequence estimation are performed at an average rate of one symbol per sample. The digital matched filter for this case is derived for an arbitrary colored noise spectrum.> Heinrich Meyr, Martin Oerder, Andreas Polydoros |
IEEE Trans. Commun. | 1 |
| 1993 | Efficient scalable architectures for Viterbi decodersabstractViterbi decoders (VDs) are widely used today for the decoding of convolutional codes in forward error correction schemes. Efficient deeply pipelined VLSI architectures, the generalized cascade VD and the trellis pipeline-interleaving (TPI) VD are adaptable to a given data rate only to a limited extent. The authors propose a novel unified class of deeply pipelined architectures, the scalable parallel Viterbi decoders (SPVD) that allows for a smoother adaptation to a given data rate. Therefore, the designer is able to choose an architecture that nearly exactly fulfills the throughput demands of the application without wasting silicon area by using a badly adapted architecture. This class of SPVDs contains the GCVD, TPI, node-serial and node-parallel architectures as important subclasses. Thus, it provides a framework for a unified description of the existing architectures as well. Furthermore, architectures can be derived that allow for 100% utilization making the complicated rate synchronization superfluous or trivial.> Stefan Bitterlich, Heinrich Meyr |
ASAP | 2 |
| 1993 | Optimum vectorization of scalable synchronous dataflow graphsabstractFor the design of complex digital signal processing systems, block diagram oriented synthesis of real time software for programmable target processors has become an important design aid. The synthesis approach discussed in this paper is based on multirate block diagrams with scalable synchronous dataflow (SSDF) semantics. For this class of dataflow graphs optimum vectorization techniques are introduced. Vectorization is treated as a transformation on an SSDF graph which increases the number of samples consumed or produced per activation of a block according to a specific optimization criterion. The presented optimization criterion jointly minimizes context-switching overhead caused by an activation of a block and maximizes the degree of vector processing of the important class of "single appearance minimum activation schedules" (SAMAS). This class comprises schedules in which each block appears exactly once and is activated minimum times. First, "single appearance" implies the most compact implementation of a schedule in terms of program memory. Second, "minimum activation" implies increased throughput according to optimum vectorization and minimal context-switching.> Sebastian Ritz, Matthias Pankert, V. Zivojinovic, Heinrich Meyr |
ASAP | 4 |
| 1993 | Partitioning and Surmounting the Software-Hardware Abstraction Gap in an ASIC Design ProjectabstractThe design complexity of an ASIC as well as its hardware expenditure can be optimized if software is used as far as possible and application-specific hardware executes only the critical algorithms. Experience within a chipset design project has led to various criteria to evaluate the tradeoff between hardware and software. Partitioning under great uncertainty concerning many realization parameters is demonstrated. The gap in abstraction between specification and implementation is much greater in hardware design than in software development. The concepts applied in surmounting the gap, like abstract modeling, type of model, and hierarchy, are defined and their relations are examined.> Klaus ten Hagen, Heinrich Meyr |
ICCD | 2 |
| 1993 | Design of optimum interpolation filters for digital demodulators
Vojin Zivojnovic, Heinrich Meyr |
ISCAS | 2 |
| 1993 | High-Level Software Synthesis for the Design of Communication SystemsabstractA synthesis environment that targets software programmable architectures such as digital signal processors (DSPs) is presented. These processors are well suited for implementation of real-time signal processing systems with medium throughput requirements. Techniques that tightly couple the synthesis environment to an existing communication system simulator are also presented. This enables a seamless transition between the simulation and implementation design level of communication systems. Special focus is on optimization techniques for mapping data flow oriented block diagrams onto DSPs. The combination of different mapping and optimization strategies allows comfortable synthesis of real-time code that is highly adapted to application-specific needs imposed by constraints on memory space, sampling rate, or latency. Thus, tradeoff analysis is supported by efficient interactive or automatic exploration of the design space. All presented concepts are illustrated by the design of a phase synchronizer with automatic gain control on a floating-point DSP.> Sebastian Ritz, Matthias Pankert, Vojin Zivojnovic, Heinrich Meyr |
IEEE J. Sel. Areas Commun. | 4 |
| 1992 | High speed bit-level pipelined architectures for redundant CORDIC implementationabstractThe CORDIC algorithm is well known as an efficient method for the computation of trigonometric/hyperbolic functions and vector rotations. The achievable throughput and the latency of CORDIC processors using conventional arithmetic are determined by the carry propagation occurring in additions/subtractions, since the CORDIC iterations are directed by the signs of intermediate results. Using a redundant number system, much higher throughput is possible due to the elimination of carry propagation, but an exact sign detection can not be implemented efficiently. The authors derive transformations of the original CORDIC algorithm which result in partially fixed iteration sequences no longer dependent on intermediate signs for the CORDIC vectoring mode as well as the rotation mode. Very fast and efficient carry-save architectures using redundant absolute value computation resulting from the transformed algorithms are described. A CORDIC processor (rotation mode) is presented as an implementation example which to the best of the authors knowledge is the fastest CMOS CORDIC realization today.> Herbert Dawid, Heinrich Meyr |
ASAP | 2 |
| 1992 | High-speed VLSI architectures for soft-output Viterbi decodingabstractDuring the last few years decoding algorithms that make not only the use of soft quantized inputs but also deliver soft decision outputs have attracted considerable attention because additional coding gains are obtainable in concatenated systems. A prominent member of this class of algorithms is the soft-output viterbi algorithm. In this paper two architectures for high speed VLSI implementations of the soft-output viterbi-algorithm are proposed and area estimates are given for both architectures. The well known trade-off between computational complexity and storage requirements is played to obtain new VLSI architectures with increased implementation efficiency. Area savings in excess of 40% in comparison to straightforward solutions are reported.> Olaf J. Joeressen, Martin Vaupel, Heinrich Meyr |
ASAP | 3 |
| 1992 | High level software synthesis for signal processing systemsabstractFor the design of complex digital signal processing systems, block diagram oriented simulation has become a widely accepted standard. Current research is concerned with the coupling of heterogenous simulation engines and the transition from simulation to the implementation of digital signal processing systems. Due to the difficulty in mastering complex design spaces high level hardware and software synthesis is becoming increasingly important. The authors concentrate on the block diagram oriented software synthesis of digital signal processing systems for programmable processors, such as digital signal processors (DSP). They present the synthesis environment DESCARTES illustrating novel optimization strategies. Furthermore they discuss goal directed software synthesis, by which code is interactively or automatically generated, which can be adapted to the application specific needs imposed by constraints on memory space, sampling rate or latency.> Sebastian Ritz, Matthias Pankert, Heinrich Meyr |
ASAP | 3 |
| 1992 | A new mobile digital radio transceiver concept using low-complexity combined equalization/trellis decoding and a near-optimal receiver sync strategyabstractFor land-mobile digital radio links over time-variant frequency-selective multipath channels, power and bandwidth efficient signaling and time diversity, as provided by interleaved trellis-coded modulation, are of great importance in a fading environment. A novel digital radio transceiver concept is presented that encompasses interleaved trellis coded modulation with large time diversity factor, a combined equalization/decoding scheme, and a quasi-optimal receiver sync strategy based on purely feedforward channel estimation. The resulting fully coherent and robust receiver is of remarkably low computational complexity. The average BER performance of such a transceiver operating over flat and selective 2-ray Rayleigh as well as typical land-mobile GSM channels is assessed via simulation. The results reveal the receiver's robustness and good BER performance and also shed light on the effects of insufficient interleaving at low vehicle speeds.> Stefan A. Fechtel, Heinrich Meyr |
PIMRC | 2 |
| 1991 | A new method for phase synchronization and automatic gain control of linearly modulated signals on frequency-flat fading channelsabstractAn optimal phase synchronization and automatic gain control (AGC) scheme for coherent reception of linearly modulated signals on frequency-flat mobile fading channels is presented. The channel model and receiver performance are described. It is shown that using the technique allows the irreducible error floors (due to random FM) known from the noncoherent methods to be practically eliminated. Depending on the fastness of the fading, large power gains over the noncoherent methods are achieved. Unfavorable analog signal processing and/or the high bandwidth inefficiency of the FDM-pilot coherent methods are also avoided.> Abbas Aghamohammadi, Heinrich Meyr, Gerd Ascheid |
IEEE Trans. Commun. | 2 |
| 1990 | High-Rate Viterbi Processor: A Systolic Array SolutionabstractThe main part of the Viterbi algorithm (VA) is a nonlinear feedback loop, the ACS recursion (add-compare-select recursion), which presents a bottleneck for high-speed implementations and cannot be circumvented by standard means. Because the two operations of the loop form an algebraic structure called semiring, it is shown that the ACS recursion of the Viterbi algorithm can therefore be written as a linear vector recursion. This allows the authors to employ the powerful techniques of parallel processing and pipelining, known for conventional linear systems, to achieve high throughput rates. Since the VA can be written as a linear vector recursion, it can be implemented by systolic arrays. For the class of shuffle exchange codes to be decoded by the Viterbi algorithm hardware-efficient code-optimized arrays are presented. It is shown that carry-save arithmetic can be used for the operations of ACS recursion, allowing each word-level operation to be pipelined and carried out by an efficient bit-level systolic array.> Gerhard P. Fettweis, Heinrich Meyr |
IEEE J. Sel. Areas Commun. | 2 |
| 1990 | On the error probability of linearly modulated signals on Rayleigh frequency-flat fading channelsabstractConsideration is given to optimal detection of linearly modulated signals subject to multiplicative Rayleigh-distributed distortion and additive white Gaussian noise. For coherent detection, regenerated amplitude and phase references are employed at the receiver to compensate for amplitude and phase deviations from the correct values. A system model is formulated under the assumption of perfect symbol timing and in the absence of intersymbol interference, producing a final additive noise term, applied just before the detection, which contains the effects of the original additive and multiplicative distortions and of the errors in the phase and amplitude references. By determining the probability density function of this final noise term for arbitrary types of linear modulation, it is possible to perform exact calculations of error probabilities.> Abbas Aghamohammadi, Heinrich Meyr |
IEEE Trans. Commun. | 2 |
| 1989 | Adaptive synchronization and channel parameter estimation using an extended Kalman filterabstractUnified modeling and estimation of the MD (multiplicative distortion) in finite-alphabet digital communication systems is presented. A simple form of MD is the carrier phase exp(j theta ), which has to be estimated and compensated for in a coherent receiver. A more general case with fading must, however, allow for amplitude as well as phase variations of the MD. The authors assume a state-variable model for the MD and generally obtain a nonlinear estimation problem with additional randomly varying system parameters such as received signal power, frequency offset, and Doppler spread. An extended Kalman filter is then applied as a near-optimal solution to the adaptive MD and channel parameter estimation problem. Examples are given to show the use and some advantages of this scheme.> Abbas Aghamohammadi, Heinrich Meyr, Gerd Ascheid |
IEEE Trans. Commun. | 2 |
| 1989 | An all digital receiver architecture for bandwidth efficient transmission at high data ratesabstractUsing the maximum-likelihood approach, algorithms for detection and synchronization are derived that are well suited for VLSI implementation. Special emphasis is placed on an all-digital implementation where carrier and clock synchronization do not require a feedback of signals to the analog part, which simplifies the analog front-end design (mixing oscillator and A/D converter sampling clock run at fixed frequency). An important advantage of the proposed algorithms is that a high clock rate is not required; only two-four times the symbol rate is needed, depending on amplitude quantization. Implementation aspects, e.g. architecture, and quantization, are considered. A prototype is described which was implemented to prove the feasibility of the concept and to evaluate the performance under practical conditions.> Gerd Ascheid, Martin Oerder, Johannes Stahl 0001, Heinrich Meyr |
IEEE Trans. Commun. | 4 |
| 1989 | Parallel Viterbi algorithm implementation: breaking the ACS-bottleneckabstractThe central unit of a Viterbi decoder is a data-dependent feedback loop which performs an add-compare-select (ACS) operation. This nonlinear recursion is the only bottleneck for a high-speed parallel implementation. A linear scale solution (architecture) is presented which allows the implementation of the Viterbi algorithm (VA) despite the fact that it contains a data-dependent decision feedback loop. For a fixed processing speed it allows a linear speedup in the throughput rate by a linear increase in hardware complexity. A systolic array implementation is discussed for the add-compare-select unit of the VA. The implementation of the survivor memory is considered. The method for implementing the algorithm is based on its underlying finite state feature. Thus, it is possible to transfer this method to other types of algorithms which contain a data-dependent feedback loop and have a finite state property.> Gerhard P. Fettweis, Heinrich Meyr |
IEEE Trans. Commun. | 2 |
| 1989 | A systematic approach to carrier recovery and detection of digitally phase modulated signals of fading channelsabstractThe problem of optimal carrier recovery and detection of digitally phase modulated signals on fading channels by using a nonstructured approach is presented, i.e. no constraint is placed on the receiver structure. First, the optimal receiver is derived for digitally phase-modulated signals when transmitted over a frequency-nonselective fading channel with memory. The memory results from the fact that usually the coherence time of the channel is larger than the symbol period. Symbols adjacent in time cannot be detected independently and therefore the well-known quadratic receiver is not optimal in this case. A maximum a posteriori (MAP) detector is derived and explicitly utilizes the channel memory for carrier recovery. The derivation shows that the optimal carrier recovery is, under certain conditions, a Kalman filter. Some attractive properties of this carrier recovery unit (including the absence of hang up) are discussed. Then the error rate of several digital modulation schemes is calculated taking the performance of the filter into account. The differences in susceptibility of the modulation schemes to carrier phase jitter are specified.> Reinhold Häb-Umbach, Heinrich Meyr |
IEEE Trans. Commun. | 2 |
| 1988 | Error handling performance of a token ring LANabstractThe authors review the key error recovery mechanisms of the ISO 8802/5 token-ring protocol. An analysis and evaluation which has been supported by an experimental investigation revealed that recovery times range from values of a few milliseconds up to a number of seconds. It is concluded that an appropriate reevaluation of present timer settings adapted to specific communication requirements in manufacturing systems can lead to increased recovery performance. In the case of backbone-ring interruptions there is no automatic recovery specified at the moment. The results are pertinent to the application of local area networks in the manufacturing environment.> Jürgen Tusch, Heinrich Meyr, Erwin A. Zurfluh |
LCN | 2 |
| 1988 | Cycle slips in synchronizers subject to smooth narrow-band loop noiseabstractThe cycle slipping in synchronizers subject to smooth narrowband loop noise is discussed. Such loop noise is shown to occur in repeater chains and in ring local area networks. The average rate of cycle slips and of bursts of cycle slips is evaluated, and closely approximated by a very simple expression that clearly indicates the influence of the loop noise bandwidth and the phase-detector characteristic. A comparison to the case of white loop noise reveals that, for most phase-detector characteristics, the slip rate caused by smooth narrowband loop noise is larger by several orders of magnitude than for white loop noise.> Marc Moeneclaey, Stanislaw Starzak, Heinrich Meyr |
IEEE Trans. Commun. | 3 |
| 1988 | Digital filter and square timing recoveryabstractThe digital realization of timing recovery circuits for digital data transmission is considered. A digital algorithm is proposed that can be implemented very efficiently even at high data rates. The resulting timing jitter has been computed and verified by simulations. In contrast to other known algorithms, the one presented here allows free-running sampling oscillators and a novel planar filtering method that prevents synchronization hangups.> Martin Oerder, Heinrich Meyr |
IEEE Trans. Commun. | 2 |
| 1987 | A Simple Method for Evaluating the Probability Density Function of the Sample Number for the Optimum Sequential DetectorabstractSequential detectors are often used in PN-spread-spectrum systems to synchronize the incoming signal. In this paper it is shown that the probability density function (pdf) of the sample number for the optimum sequential detector can be computed recursively, provided that the samples from the envelope correlator are appropriately modeled as statistically independent. Furthermore, the error probabilities of this detector can be determined. Heinrich Meyr, Gerhard Polzer |
IEEE Trans. Commun. | 1 |
| 1986 | Synchronization Failures in a Chain of PLL SynchronizersabstractIn digital transmission systems employing PLL repeaters, timing jitter is produced mostly by intersymbol interference (ISI). Despite the fact that IS1 represents a narrow-band disturbance to the PLL for a long chain of N repeaters, it is responsible for cycle slips of the PLL repeater. It is demonstrated that the cycle slip rate is the key parameter determining the error performance of the system. It is shown that the accumulated jitter for a large chain becomes a narrow-band process, which can be replaced by an equivalent low-pass jitter-noise process, which lends itself to analytical treatment. Numerical results of the cycle slip rate are given, based on the solution of a time-dependent Fokker-Planck equation. The results are useful for a wide range of loop transfer functions and phase-detector characteristics. The theoretical findings are confirmed by experimental results. Heinrich Meyr, Luitjens Popken, Hans R. Müller |
IEEE Trans. Commun. | 1 |
| 1983 | Transmission Design Criteria for a Synchronous Token RingabstractThis paper discusses the transmission design criteria and limiting factors of an experimental synchronous token ring implemented at the IBM Zurich Research Laboratory. The following key aspects are addressed: 1) ring topology and wiring, 2) transmission, and 3) ring synchronization with phase-locked loops. Wiring of a ring is based on a two-level hierarchy with passive wiring concentrators placed at convenient locations in a building. Data are transmitted with differential Manchester code. Special emphasis is placed on the synchronization methods and on the parameters and tolerances which limit distance and number of stations that can be attached to a ring. An analysis of the behavior of a chain of repeaters under growing jitter is given. Also the various procedures for guaranteeing high reliability are outlined. An experimental token ring has been running since 1981, and has been tested extensively under extreme jitter and noise conditions. Heinz J. Keller, Heinrich Meyr, Hans R. Müller |
IEEE J. Sel. Areas Commun. | 2 |
| 1983 | Performance Analysis for General PN-Spread-Spectrum Acquisition TechniquesabstractThis paper presents a unified approach for computing the probability density function (pdf) of the acquisition time for pseudonoise (PN) search algorithms. This approach can be applied to arbitrary search strategies and a priori distributions of the code phase. Furthermore, it is shown that the mean and the variance of the acquisition time can be obtained directly without computing the pdf first. Heinrich Meyr, Gerhard Polzer |
IEEE Trans. Commun. | 1 |
| 1982 | Real-time estimation of moving time delayabstractA two-step algorithm for the estimation of a rapidly varying time delay between two stochastic signals is described. In the first step the maximum-likelihood (ML) estimate is computed over an observation interval small enough to consider the delay to be approximately constant. Due to the short averaging interval the probability of ambiguous peaks is greatly increased. Therefore in the second step, a non-linear, adaptive postfiltering algorithm is presented that effectively suppresses 'outliers' of the ML-estimator. Heinrich Meyr, Gerhard Spies, Jörg Bohmann |
ICASSP | 1 |
| 1982 | Cycle Slips in Phase-Locked Loops: A Tutorial SurveyabstractCycle slips in phase-locked loops are statistical, nonlinear phenomena. This makes a mathematical analysis extremely difficult. As a consequence, the results of such an analysis are not easily accessible to the practicing engineer. It is the purpose of this survey paper to present a self-contained discussion of cycle slips in phase-locked loop avoiding advanced mathematical tools. Based on the results of an extensive experimental study we explain the underlying principle of the complex interaction between nonlinearity and noise. The results are complemented by simple, approximate analysis which agrees well with the experimental findings. In addition, we present a new and complete set of diagrams on cycle slip statistics not presently available in the literature. Gerd Ascheid, Heinrich Meyr |
IEEE Trans. Commun. | 2 |
| 1980 | Phase Acquisition Statistics for Phase-Locked LoopsabstractPhase acquisition probabilities for phase-locked synchronizers are derived. Both self- and aided-acquisition techniques are investigated and compared. It is shown that so called "hang-up" can be prevented by using initial quadrant estimation to control appropriate slew voltage applied to VCO. Heinrich Meyr, Luitjens Popken |
IEEE Trans. Commun. | 1 |
| 1978 | Theory of phase tracking systems of arbitrary order: Statistics of cycle slips and probability distribution of the state vectorabstractThe time-dependent distribution of the number of cycle slips in positive and negative directions, and the correlation of their time spacings, are derived from a new statistical model of an(N + l)-order phase tracking system. The probability density of the phase error and the other system variables are shown to agree with known results. Relations for the steady state are obtained in a relatively simple form. Some limiting conditions are mentioned under which the model reduces to a computationally much simpler renewal model described earlier. Dietrich Ryter, Heinrich Meyr |
IEEE Trans. Inf. Theory | 2 |
| 1977 | Complete statistical description of the phase-error process generated by correlative tracking systemsabstractThe complete statistical description of a first-order correlative tracking system with periodic nonlinearity is shown to be embedded in a renewal process. The time-dependent probability density function of the phase error, as well as the distribution of the cycle slips, is computed. The use of the renewal process approach makes it possible for the first time to compute the distribution of the positive and negative number of cycle slips within a given time interval. This information is sufficient to determine the probability density function of the absolute phase error. William C. Lindsey, Heinrich Meyr |
IEEE Trans. Inf. Theory | 2 |
| 1976 | Delay-Lock Tracking of Stochastic SignalsabstractThe measurement and tracking of the delay between two versions of a stochastic signal by cross-correlation techniques is considered. Such techniques have broad applications, e.g., interferometry, noncontact speed and distance measurement, etc. The paper begins by discussing the functional diagram of the tracking system. From this diagram a mathematically equivalent model of the system is derived and its similarities to the well known baseband model of the phase-locked loop are discussed. Using Fokker-Planck (F-P) techniques the performance of the system, as a function of fundamental system parameters, is computed and graphically illustrated. These results are then compared with experimental results obtained by computer simulation. Heinrich Meyr |
IEEE Trans. Commun. | 1 |
| 1975 | Nonlinear Analysis of Correlative Tracking Systems Using Renewal Process TheoryabstractA new method is presented which describes the behavior of an(N + 1)th-order tacking system in which the nonlinearity is either periodic [phase-locked loop (PLL) type] or a nonperiodic [delay-locked loop (DLL) type]. The cycle slipping of such systems is modeled by means of renewal Markov processes. A fundamental relation between the probability density function (pdf) of the single process and the renewal process is derived which holds in the transient as well as in the stationary state. Based on this relation it is shown that the stationary pdf, the mean time between two cycle slips, and the average number of cycles to the right (left) can be obtained by solving a single Fokker-Planck equation of the renewal process. The method is applied to the special case of a PLL and compared with the so-called periodic-extension (PE) approach. It is shown that the pdf obtained via the renewal-process approach can be reduced to agree with the PE solution for the first-order loop in the steady state only. The reasoning and its implications are discussed. In fact, it is shown that the approach based upon renewal-process theory yields more information about the system's behavior than does the PE solution. Heinrich Meyr |
IEEE Trans. Commun. | 1 |