EDBT 2026 Demo / reviewers in the wild / expert
Rudy Lauwereins
dblp:55/5766
· DBLP profile ↗
95ranked-venue papers
8as first author
4since 2021 · last 2022
0000-0002-3861-0168ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 44 · 5 first-author · 3 since 2021Software engineering, systems software and programming languages · 20 · 4 first-authorGraphics, computer vision, multimedia, augmented reality and games · 18Computer networks · 13Artificial intelligence and machine learning · 5 · 1 since 2021Theory of computation · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | Dynamic Quantization Range Control for Analog-in-Memory Neural Networks AccelerationabstractAnalog in Memory Computing (AiMC) based neural network acceleration is a promising solution to increase the energy efficiency of deep neural networks deployment. However, the quantization requirements of these analog systems are not compatible with state-of-the-art neural network quantization techniques. Indeed, while the quantization of the weights and activations is considered by modern deep neural network quantization techniques, AiMC accelerators also impose the quantization of each Matrix Vector Multiplication (MVM) result. In most demonstrated AiMC implementations, the quantization range of MVM results is considered a fixed parameter of the accelerator. This work demonstrates that dynamic control over this quantization range is possible but also desirable for analog neural networks acceleration. An AiMC compatible quantization flow coupled with a hardware aware quantization range driving technique is introduced to fully exploit these dynamic ranges. Using CIFAR-10 and ImageNet as benchmarks, the proposed solution results in networks that are both more accurate and more robust to the inherent vulnerability of analog circuits than fixed quantization range based approaches. Nathan Laubeuf, Jonas Doevenspeck, Ioannis A. Papistas, Michele Caselli, Stefan Cosemans, Peter Vrancx, Debjyoti Bhattacharjee, Arindam Mallik, Peter Debacker, Diederik Verkest, Francky Catthoor, Rudy Lauwereins |
ACM Trans. Design Autom. Electr. Syst. | 12 |
| 2022 | Efficient Backside Power Delivery for High-Performance Computing SystemsabstractIn this work, we present a thin-profile, efficient power delivery approach, including a voltage regulator with in-package power inductor and backside power delivery network (PDN). To meet 1-$\mathrm {W}/{\mathrm {mm}}^{2}$power-density target for high-performance computing (HPC) systems, a 25-high-$Q$-factor (300 MHz), 150-$\mu \text{m}$-thick, in-molding power inductor is provided for high-efficiency point-of-load (PoL) voltage regulation. Meanwhile, a novel analytical model for backside power delivery is developed for computer-aided-design (CAD) procedure to optimize the system efficiency. For the power flowing from bumps (57-$\mu \text{m} V_{\mathrm {DD}}$-bump pitch) and backside PDN to active devices, the area resistances contributed by backside PDN and the buried power rail (BPR) are 23% and 77%, respectively, if a 10-$\mu \text{m}$-horizontal-pitch nano- through-silicon via ($n$TSV) is available. The resulting impact on power dissipation is within 1% so negligible. A higher ratio (0.5) buck converter with maintained efficiency is combined to better benefit the external interconnect. The overall power delivery efficiency$\eta \,\,=83$% can be obtained for 1-$\mathrm {W}/{\mathrm {mm}}^{2}$power-density target. The power losses contributed by an air-core inductor, power switches, and PDN/BPR/redistribution layer (RDL) are 26%, 66%, and 8%, respectively. Hesheng Lin, Geert Van der Plas, Dimitrios Velenis, Francky Catthoor, Rudy Lauwereins, Eric Beyne |
IEEE Trans. Very Large Scale Integr. Syst. | 6 |
| 2022 | 84%-Efficiency Fully Integrated Voltage Regulator for Computing Systems Enabled by 2.5-D High-Density MIM CapacitorabstractWe present a$\mu \text{m}$-thin-profile power delivery solution including a charge pump with integrated passives. Targeting 1 W/mm2or higher power density, a 2.5-D high-density metal-insulator-metal (MIM) capacitor deposited on high aspect ratio (HAR) (up to 5) oxide studs is proposed. With approximately 25-nm-thick HfAlOx dielectric, its measured capacitance density is 25.4 nF/mm2for a capacitor size ranging from 1/16 mm2to 1 mm2. This shows$3.6\times $density improvement compared with the planar MIM. Theoretically, 86 nF/[email protected] bias can be obtained if a 10-nm dielectric is deposited. Moreover, the measured leakage current density is within 65 pA/mm2at 1-V bias (negligible for a 1 W/mm2-power delivery). For a backside (BS) power delivery, this 2.5-D MIM capacitor can be realized by only three BS metal layers. This enables the low-cost and thin-profile delivery system ($\sim \!\!\mu \text{m}$thickness), and the whole power delivery efficiency including a 1/2-ratio charge pump is$\eta \,\,=84$%@1 W/mm2(>5% boost in the power efficiency). Hesheng Lin, Dimitrios Velenis, Philip Nolmans, Francky Catthoor, Rudy Lauwereins, Geert Van der Plas, Eric Beyne |
IEEE Trans. Very Large Scale Integr. Syst. | 6 |
| 2021 | Noise tolerant ternary weight deep neural networks for analog in-memory inferenceabstractAnalog in memory computing (AiMC) is a promising hardware solution to efficiently perform inference with deep neural networks (DNNs). Similar to digital DNN accelerators, AiMC systems benefit from aggressively quantized DNNs. In addition, AiMC systems also suffer from noise on activations and weights. Training strategies to condition DNNs against weight noise can increase the efficiency of AiMC systems by enabling the use of more compact but more noisy weight memory devices. In this work, we utilize noise-aware training and introduce gradual noise training and network width scaling to increase the tolerance of DNNs against weight noise. Our results show that noise-aware training and gradual noise training drastically lowers the impact of weight noise without changing the network size. By utilizing network width scaling, the weight noise tolerance is increased even more with the penalty of more network parameters. Jonas Doevenspeck, Peter Vrancx, Nathan Laubeuf, Arindam Mallik, Peter Debacker, Diederik Verkest, Rudy Lauwereins, Wim Dehaene |
IJCNN | 7 |
| 2017 | Wave pipelining for majority-based beyond-CMOS technologiesabstractThe performance of some emerging nanotechnologies benefits from wave pipelining. The design of such circuits requires new models and algorithms. Thus we show how Majority-Inverter Graphs (MIG) can be used for this purpose and we extend the related optimization algorithms. The resulting designs have increased throughput, something that has traditionally been a weak point for the majority of non-charge-based technologies. We benchmark the algorithm on MIG netlists with three different technologies, Spin Wave Devices (SWD), Quantum-dot Cellular Automata (QCA), and NanoMagnetic Logic (NML). We find that the wave pipelined version of the netlists have an improvement in throughput over power of 23×, 13×, and 5× for SWD, QCA, and NML, respectively. In terms of throughput over area ratio, the improvement is 5×, 8×, and 3×, respectively. Odysseas Zografos, A. De Meester, Eleonora Testa, Mathias Soeken, Pierre-Emmanuel Gaillardon, Giovanni De Micheli, Luca G. Amarù, Praveen Raghavan, Francky Catthoor, Rudy Lauwereins |
DATE | 10 |
| 2014 | Interfacing to living cellsabstractRecent advances in More than Moore technology enable close inspection of and even direct interfacing to living cells. This paper illustrates this through three use cases. In the first use case, the type or quality of billions of cells is quickly inspected in a fluidic medium. Secondly, the effect of potential drugs is monitored in neural cell cultures. In the third use case, neural brain activity is recorded in vivo using implantable electrodes to understand how the brain functions. Rudy Lauwereins |
DATE | 1 |
| 2014 | NBTI Aging on 32-Bit Adders in the Downscaling Planar FET Technology NodesabstractReliability of advanced deeply scaled CMOS technologies is being threatened by time-dependent degradation mechanisms such as Negative Bias Temperature Instability (NBTI) phenomenon that cause workload-dependent shifts on a transistor's threshold voltage (VTH), and performance during its lifetime. In this study, NBTI-induced performance degradation of 32-bit adders (one of the most fundamental block of a processor's arithmetic logic unit) is investigated from the points of architectural topology, technology scaling (i.e. commercial 28, 45, 65nm nodes) and workload dependency. The selected adder architectures vary from basic to complex parallel-prefix ones. A workload-dependent, NBTI aging-aware digital design flow was developed within the industry standard EDA tool chain. NBTI model is based on the extracted Capture and Emission Time (CET) maps from the actual wafer measurements. Static Timing Analysis (STA) is performed to evaluate the performance degradation at the +3σ corner. Results on adders under the NBTI aging after 3 years show a performance loss up to 16%. NBTI aging results in the replacement of the time-zero critical path by an initially non-critical path during a circuit's lifetime. The time-zero critical path can shift to a new one with a probability of 89%. Technology scaling and the choice of process technology can impact the degradation by 2×. Finally, the performance degradation can vary up to 8.2× under workload variations. Halil Kukner, Pieter Weckx, Sébastien Morrison, Praveen Raghavan, Ben Kaczer, Francky Catthoor, Liesbet Van der Perre, Rudy Lauwereins, Guido Groeseneken |
DSD | 8 |
| 2014 | Towards approaching near-optimal MIMO detection performance ONAC-programmable baseband processorabstractLattice Reduction aided softoutput MIMO detectors have been demonstrated to offer a promising gain. However, computing Log-Likelihood ratios (LLR) for near-optimal MIMO detection, still poses a significant challenge for practical implementations. In this work, we present counter-ML bit-flipping algorithm for LLR generation. The proposed LLR generation algorithm has been designed to take advantage of the previously reported list generation algorithm, Multi-Tree Selective Spanning (MTSS), by maximizing the reuse of computations. Afterwards, a C-programmable MIMO detector architecture providing both data level parallelism (DLP) and instruction level parallelism (ILP), is designed for implementation. The proposed solution supports multiple MIMO detection modes, with both hard and softoutput. Performance of the proposed solution can be tuned ranging from SIC to near-ML to near-MAP, by adjusting a single parameter. In case of 4 × 4 QAM-64, it achieves peak-throughputs of 2.43Gbps and 629Mbps in case of hard and softoutput MIMO detection, with only 66.37mW and 76.14mW respective power consumption. Ubaid Ahmad, Min Li 0001, Amir Amin, Meng Li 0012, Liesbet Van der Perre, Rudy Lauwereins, Sofie Pollin |
ICASSP | 6 |
| 2014 | The value of feedback for LTE resource allocationabstractThe LTE cellular network is designed for meeting the requirements of a broad range of applications in very dynamic conditions. This explains its great flexibility in resource allocation. In order to provide the mobile stations with the exact required services, the base station relies on feedback information reported by each mobile station. In this paper, several feedback reduction schemes are analysed and compared, for a broad range of LTE resource allocation schemes and deployment scenarios, by means of simulation and a quantitative cost model. It is concluded that feedback reduction should be adapted to the scenario, as function of users, resource allocation strategy or channel properties, and up to 139% gain in overall throughput can be obtained by doing so. Alessandro Chiumento, Claude Desset, Sofie Pollin, Liesbet Van der Perre, Rudy Lauwereins |
WCNC | 5 |
| 2014 | Exploiting transport-block constraints in LTE improves downlink performanceabstractEfficient resource allocation is necessary to provide the users with the quality of service promised in modern cellular networks, such as LTE. Traditional allocation methods make use of the smallest granularity available, the physical resource block (PRB), to assign resources to each user. The selected resources assigned to a user form a transport block (TB). However, the standard constrains the use of only one modulation and coding rate per TB. This forces the channel quality of the resources to be averaged across the TB, in order to determine the best modulation and coding scheme. In state-of-the-art systems, a non-linear heuristic is used in order to perform this averaging. Unfortunately, when bad PRBs are present next to good ones, this strategy is not optimal. We show that dropping the worst PRBs can improve the performance while remaining standard-compliant. We propose a simple algorithm that can be overlaid to any existing solution and we analyse its effect on multiple state-of-the-art schedulers. A gain is obtained both in throughput (up to 8% increase) and in power consumption (up to 23% reduction). Alessandro Chiumento, Sofie Pollin, Claude Desset, Liesbet Van der Perre, Rudy Lauwereins |
WCNC | 5 |
| 2014 | Derivative-Based Scale Invariant Image Feature Detector With Error ResilienceabstractWe present a novel scale-invariant image feature detection algorithm (D-SIFER) using a newly proposed scale-space optimal 10th-order Gaussian derivative (GDO-10) filter, which reaches the jointly optimal Heisenberg's uncertainty of its impulse response in scale and space simultaneously (i.e., we minimize the maximum of the two moments). The D-SIFER algorithm using this filter leads to an outstanding quality of image feature detection, with a factor of three quality improvement over state-of-the-art scale-invariant feature transform (SIFT) and speeded up robust features (SURF) methods that use the second-order Gaussian derivative filters. To reach low computational complexity, we also present a technique approximating the GDO-10 filters with a fixed-length implementation, which is independent of the scale. The final approximation error remains far below the noise margin, providing constant time, low cost, but nevertheless high-quality feature detection and registration capabilities. D-SIFER is validated on a real-life hyperspectral image registration application, precisely aligning up to hundreds of successive narrowband color images, despite their strong artifacts (blurring, low-light noise) typically occurring in such delicate optical system setups. Pradip Mainali, Gauthier Lafruit, Klaas Tack, Luc Van Gool, Rudy Lauwereins |
IEEE Trans. Image Process. | 5 |
| 2013 | A computationally efficient soft-output Lattice Reduction-aided Selective Spanning Sphere Decoder for wireless MIMO systemsabstractIn recent years, the algorithmic optimizations and implementations of near-optimal Multiple-Input Multiple-Output (MIMO) detectors have been an area of active research. Lattice Reduction (LR) has shown to be a promising technique to improve the performance of linear MIMO detectors. However, LR-aided linear hard-output MIMO detection is still far from optimal. Practical systems use soft-output information to exploit gains from coded systems in order to yield near-optimal performance. In this paper, the LR-aided Selective Spanning Sphere Detection algorithm is proposed as a reduced-complexity candidate list generation method for soft-output MIMO detection, specifically optimized for practical MIMO-OFDM systems. This algorithm uses efficient and scalable heuristics based on simple processor-friendly operations that significantly contribute to lowering the computational complexity of the MIMO detection problem. Results from Monte Carlo simulations reveal that LR-aided SSSD is a promising algorithm that is capable of providing near-optimal performance whilst being especially computationally efficient, in comparison to other algorithms. Hoang Duy Nguyen, Ubaid Ahmad, Min Li 0001, Liesbet Van der Perre, Rudy Lauwereins, Sofie Pollin |
PIMRC | 5 |
| 2013 | SIFER: Scale-Invariant Feature Detector with Error Resilience
Pradip Mainali, Gauthier Lafruit, Qiong Yang, Bert Geelen, Luc Van Gool, Rudy Lauwereins |
Int. J. Comput. Vis. | 6 |
| 2012 | Biomedical electronics serving as physical environmental and emotional watchdogsabstractOver forty years of happy CMOS scaling brought the room-sized super-computer for the nerds into everyone's pocket, literally connecting every-body on earth. In an economy which is based on double digit growth, the obvious next step is to connect everything on earth. Rudy Lauwereins |
DAC | 1 |
| 2012 | Impact of Duty Factor, Stress Stimuli, and Gate Drive Strength on Gate Delay Degradation with an Atomistic Trap-Based BTI ModelabstractWith deeply scaled CMOS technology, Bias Temperature Instability (BTI) has become one of the most critical degradation mechanisms impacting the device reliability. In this paper, we present the BTI evaluation of a single inverter gate covering both the PMOS and NMOS degradations in a workload dependent, atomistic trap-based, stochastic BTI model. The gate propagation delay depends on the gate intrinsic delay, the input signal characteristics, and the output load. Thus, the BTI degradation is investigated due to the impact of 1) duty factor, 2) periodic clock-based and non-periodic random input sequences, 3) gate drive strength. The inverter is chosen due to its representativity of other CMOS logic gates. The applied BTI model is stochastic, and the device parameters are orthogonally generated by distributions. Results show 3% and 27% degradation shifts on the distribution mean and worst-case. In addition, it is shown that the near-critical paths with lower drive strength cells are more susceptible to the BTI degradation than the critical paths with higher drive strength cells. Halil Kukner, Pieter Weckx, Praveen Raghavan, Ben Kaczer, Francky Catthoor, Liesbet Van der Perre, Rudy Lauwereins, Guido Groeseneken |
DSD | 7 |
| 2012 | Supplementary Proof for "Equalization Algorithms in the Frequency Domain for Continuous Phase Modulations"abstractTo enable frequency domain equalization of continuous phase modulations (CPM), a block construction that ensures cyclicity over each block without disrupting the phase continuity between them was proposed in . We formalize and prove the constraints that should be respected to enable the application of this technique to any CPM scheme. Wim Van Thillo, François Horlin, Jimmy Nsenga, Valéry Ramon, André Bourdoux, Rudy Lauwereins |
IEEE Trans. Commun. | 6 |
| 2012 | Constant Time Joint Bilateral Filtering Using Joint Integral HistogramsabstractIn this brief, we present a constant time method for the joint bilateral filtering. First, we propose an image data structure, coined as joint integral histograms (JIHs). Extending the classic integral images and the integral histograms, it represents the global information of two correlated images. In a JIH, the value at each bin indicates an integral determined by the two images. Then, the joint bilateral filtering is transformed to computation and manipulation of histograms. Utilizing the JIHs, we are capable of joint bilateral filtering in constant time. Its performance is validated in a digital photography approach using Flash-noFlash image pairs. Compared with the brute-force method, the proposed method achieves a speedup factor of 2-3 orders of magnitude while producing similar filtering results. Ke Zhang 0012, Gauthier Lafruit, Rudy Lauwereins, Luc Van Gool |
IEEE Trans. Image Process. | 3 |
| 2011 | Robust Low Complexity Corner DetectorabstractCorner feature point detection with both the high-speed and high-quality is still very demanding for many real-time computer vision applications. The Harris and Kanade-Lucas-Tomasi (KLT) are widely adopted good quality corner feature point detection algorithms due to their invariance to rotation, noise, illumination, and limited view point change. Although they are widely adopted corner feature point detectors, their applications are rather limited because of their inability to achieve real-time performance due to their high complexity. In this paper, we redesigned Harris and KLT algorithms to reduce their complexity in each stage of the algorithm: Gaussian derivative, cornerness response, and non-maximum suppression (NMS). The complexity of the Gaussian derivative and cornerness stage is reduced by using an integral image. In NMS stage, we replaced a highly complex sorting and NMS by the efficient NMS followed by sorting the result. The detected feature points are further interpolated for sub-pixel accuracy of the feature point location. Our experimental results on publicly available evaluation data-sets for the feature point detectors show that our low complexity corner detector is both very fast and similar in feature point detection quality compared to the original algorithm. We achieve a complexity reduction by a factor of 9.8 and attain 50 f/s processing speed for images of size 640$\,\times\,$480 on a commodity central processing unit with 2.53 GHz and 3 GB random access memory. Pradip Mainali, Qiong Yang, Gauthier Lafruit, Luc Van Gool, Rudy Lauwereins |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2011 | Real-Time and Accurate Stereo: A Scalable Approach With Bitwise Fast Voting on CUDAabstractThis paper proposes a real-time design for accurate stereo matching on compute unified device architecture (CUDA). We present a leading local algorithm and then accelerate it by parallel computing. High matching accuracy is achieved by cost aggregation over shape-adaptive support regions and disparity refinement using reliable initial estimates. A novel sample-and-restore scheme is proposed to make the algorithm scalable, capable of attaining several times speedup at the expense of minor accuracy degradation. The refinement and the restoration are jointly realized by a local voting method. To accelerate the voting on CUDA, a graphics processing unit (GPU)-oriented bitwise fast voting method is proposed, faster than the traditional histogram-based approach with two orders of magnitude. The whole algorithm is parallelized on CUDA at a fine granularity, efficiently exploiting the computing resources of GPUs. Our design is among the fastest stereo matching methods on GPUs. Evaluated in the Middlebury stereo benchmark, the proposed design produces the most accurate results among the real-time methods. The advantages of speed, accuracy, and desirable scalability advocate our design for practical applications such as robotics systems and multiview teleconferencing. Ke Zhang 0012, Jiangbo Lu, Qiong Yang, Gauthier Lafruit, Rudy Lauwereins, Luc Van Gool |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2010 | Lococo: low complexity corner detectorabstractThe high-speed feature detection is still very demanding for many computer vision applications. In this paper, the Harris and KLT corner detectors are redesigned to reduce the complexity of the algorithm. The complexity of Harris and KLT corner detectors are reduced by using the box kernel, the integral image and efficient non-maximum suppression, achieving complexity reduction by a factor of 8. For the image of size 1000×700, our method costs only 74ms on a commodity CPU with 2GHz and 1GB RAM. Pradip Mainali, Qiong Yang, Gauthier Lafruit, Rudy Lauwereins, Luc Van Gool |
ICASSP | 4 |
| 2010 | Maximum SINR-Based Beamforming for the MISO OFDM Interference ChannelabstractWe address the problem of co-channel interference (CCI) in wireless mesh networks (WMNs). In such networks, multiple nodes communicate concurrently using the same time/frequency resources. This is recognized as the interference channel (IFC) because the CCI is a major impairment in such a scenario. To mitigate the CCI, non-cooperative beamforming techniques can be employed. Non-cooperative beamformers are simpler to implement than their cooperative counter-parts because they require neither synchronization nor the sharing of information data between the transmit nodes. In this paper, we then propose a non-cooperative beamforming scheme to improve the performance of WMNs with frequency-selective channels. We present an iterative algorithm that maximizes the signal-to-interference-and-noise ratio criterion over all subcarriers of the orthogonal frequency division multiplexing system. The performance of this algorithm is evaluated through simulations. Yann Y. L. Lebrun, Valéry Ramon, André Bourdoux, Sofie Pollin, François Horlin, Rudy Lauwereins |
ICC | 6 |
| 2010 | Robust low complexity feature trackingabstractIn this paper, we present the Kanade-Lucas-Tomasi (KLT) tracking algorithm coupled with a varying integration window, which tracks a small subset of feature points to initialize the approximate motion model between images. For the remaining larger subset of the feature points initial tracking location is predicted by using this motion model, thus improving the tracking result. For an image of size 1000×700, the computational cost is reduced by a factor of 9.5 and tracking 500 features with our method runs in only 64 ms on a commodity 2GHz CPU with 1GB RAM. Pradip Mainali, Qiong Yang, Gauthier Lafruit, Rudy Lauwereins, Luc Van Gool |
ICIP | 4 |
| 2010 | Joint integral histograms and its application in stereo matchingabstractIn this paper, we first propose a technique, referred as joint integral histograms, for weighted filtering with O(1) computational complexity. The technique is built on the classic integral images and the recent integral histograms. In a joint integral histogram, instead of remembering bin occurrences, the value at each bin indicates an integral defined by two signals. Beyond the integral histograms, our method supports weighted filtering with a more general form, where the weight could be a function of a signal different from the signal to be filtered. Then, we present a local stereo matching approach as an instantiation of the technique. Using the joint integral histograms, we achieve a speedup factor of about two orders of magnitude. Thanks to the huge speedup, the stereo method is among the best local approaches in terms of the trade-off between matching accuracy and execution speed. Experimental results demonstrate the advantages of both the joint integral histograms technique and the stereo matching approach. Ke Zhang 0012, Gauthier Lafruit, Rudy Lauwereins, Luc Van Gool |
ICIP | 3 |
| 2010 | Modeling and exploiting spatial locality trade-offs in wavelet-based applications under varying resource requirementsabstractFuture dynamic applications will require new mapping strategies to deliver power-efficient performance. Fully static design-time mappings will not be able to optimally address the unpredictably varying application characteristics and system resource requirements. Instead, the platforms will not only need to be programmable in terms of instruction set processors, but also at least partial reconfigurability will be required, while the applications themselves will need to exploit this increased freedom at runtime to adapt to the dynamism. In this context, it is important for applications to optimally exploit the memory hierarchy under varying memory availability. This article presents an analysis of spatial locality trade-offs in wavelet-based applications, to be used in dynamic execution environments: Depending on the encountered runtime conditions, the execution switches to different memory optimized instantiations or localizations, optimally exploiting temporal and spatial locality under these conditions. This is enabled by systematic mapping guidelines, indicating how the miss-rate behavior of a localization is influenced by a specific execution condition, under which conditions a certain localization is optimal and which miss-rate gains may be obtained by switching to that localization. Bert Geelen, Vissarion Ferentinos, Francky Catthoor, Gauthier Lafruit, Diederik Verkest, Rudy Lauwereins, Thanos Stouraitis |
ACM Trans. Embed. Comput. Syst. | 6 |
| 2010 | Novel block constructions using an intrafix for CPM with frequency domain equalizationabstractTo enable frequency domain equalization (FDE) for continuous phase modulation (CPM), both cyclicity of individual symbol blocks and phase continuity between different blocks have to be guaranteed. In this letter, we present new block constructions that use a subblock of data-dependent symbols, called intrafix, to satisfy both constraints for different CPM-FDE systems: using either a cyclic prefix or a training sequence (TS), both for precoded and nonprecoded CPM. The known symbols of a TS can be used to improve synchronization and channel estimation. Precoding can be applied to a certain class of CPM schemes to halve the bit error rate. Wim Van Thillo, François Horlin, Valéry Ramon, André Bourdoux, Rudy Lauwereins |
IEEE Trans. Wirel. Commun. | 5 |
| 2009 | Joint Transmit and Receive Analog Beamforming in 60 GHz MIMO Multipath ChannelsabstractAnalog beamforming (ABF) with one scalar weight per antenna is an attractive technique for low-cost, low-power 60 GHz multi-antenna wireless communication systems. However, the design of the corresponding joint transmit and receive (Tx/Rx) ABF optimization algorithms is still challenging in the case of multipath channels due to the constraint of having only one scalar weight per antenna. In this paper, we aim at maximizing the average signal to noise ratio (SNR) at the input of the equalizer and analytically derive close-to-optimal Tx/Rx scalar weights. We show that the required channel state information (CSI) for joint Tx/Rx ABF weights computation is the inner product between all Tx/Rx channel impulse response pairs. Taking the channel length into account, a training-based estimation strategy of this CSI is proposed. Simulation results carried out in a typical 60 GHz multipath environment show that the proposed scheme outperforms the existing ABF schemes in term of BER performances. Jimmy Nsenga, Wim Van Thillo, François Horlin, Valéry Ramon, André Bourdoux, Rudy Lauwereins |
ICC | 6 |
| 2009 | A Flexible Antenna Selection Scheme for 60 GHz Multi-Antenna Systems Using Interleaved ADCsabstractWe present a new scheme for maximizing the signal-to-noise ratio (SNR) in wideband multi-antenna receivers that employ time-interleaved analog-to-digital converters (ADCs). Current wideband receivers interleave a fixed number of slower ADCs into one fast ADC and assign this latter to one antenna according to an antenna selection algorithm. However, this does not always guarantee an optimal trade-off between thermal noise and quantization noise. This results in an overall SNR that is lower than what could be obtained with the same number of ADCs, assigned in a more optimal way. Therefore, we propose to adjust the number of slower ADCs assigned to a certain fast, interleaved ADC dynamically, according to the SNR of every individual antenna. Our proposed algorithm can be implemented at the expense of a very limited hardware complexity increase. The SNR gain of our new scheme can exceed 7 dB, depending on channel conditions and ADC specifications. Wim Van Thillo, Sofie Pollin, Jimmy Nsenga, Valéry Ramon, André Bourdoux, François Horlin, Rudy Lauwereins, Ahmad Bahai |
ICC | 7 |
| 2009 | Robust stereo matching with fast Normalized Cross-Correlation over shape-adaptive regionsabstractNormalized cross-correlation (NCC) is a common matching technique to tolerate radiometric differences between stereo images. However, traditional rectangle-based NCC tends to blur the depth discontinuities. This paper proposes an efficient stereo algorithm with NCC over shape-adaptive matching regions, producing depth-discontinuity preserving disparity maps while remaining the advantage of robustness to radiometric differences. To alleviate the computational intensity, we propose an acceleration algorithm using an orthogonal integral image technique, achieving a speedup factor of 10~27. In addition, a voting scheme on reliable estimates is applied to refine the initial estimates. Experiments show that, besides the robustness, the proposed method obtains accurate disparity maps at fast speed. Our method highly ranks among the local approaches in the Middlebury stereo benchmark. Ke Zhang 0012, Jiangbo Lu, Gauthier Lafruit, Rudy Lauwereins, Luc Van Gool |
ICIP | 4 |
| 2009 | Accurate and efficient stereo matching with robust piecewise votingabstractIn this paper, we propose an efficient local stereo algorithm for accurate disparity estimation. First, we attain initial disparity estimates by iterating a cross-based cost aggregation process. Then, we propose a robust voting scheme to refine the initial estimates based on a piecewise smoothness prior, improving the quality in occluded regions and low-textured regions effectively. The refinement is guided by the segmentation result of input images. Unreliable initial estimates, which are detected using an efficient left-right consistency check, are rejected to increase the reliability of the voting results. Evaluated with the Middlebury stereo benchmark, our method is among the top performing local methods in accuracy. Compared to other local methods with similar accuracy, our method is faster by a factor of about two orders. Ke Zhang 0012, Jiangbo Lu, Gauthier Lafruit, Rudy Lauwereins, Luc Van Gool |
ICME | 4 |
| 2009 | Spatial locality exploitation for runtime reordering of JPEG2000 wavelet data layoutsabstractExploitation of spatial locality is essential for memories to increase the access bandwidth and to reduce the access-related latency and energy per word. Spatial locality exploitation of a kernel can be improved by modifying placement of data in memory, but this may be felt not only by the kernel itself, but also in other application components accessing the same data. Thus care is needed to avoid global miss-rate improvements are thwarted by miss-rate increases in other application components. This article examines application-level miss-rate increases due to handling modified Wavelet Transform data layouts by explicitly reordering at runtime, exploiting the execution order freedom within a reordering buffer when the layout of surrounding components is known. For the JPEG2000 application, taking into account the reordering costs still results in 80% net WT miss-rate gains. Bert Geelen, Vissarion Ferentinos, Francky Catthoor, Gauthier Lafruit, Diederik Verkest, Rudy Lauwereins, Thanos Stouraitis |
ACM Trans. Design Autom. Electr. Syst. | 6 |
| 2009 | Low-complexity linear frequency domain equalization for continuous phase modulationabstractIn this paper, we develop a new low-complexity linear frequency domain equalization (FDE) approach for continuous phase modulated (CPM) signals. As a CPM signal is highly correlated, calculating a linear minimum mean square error (MMSE) channel equalizer requires the inversion of a nondiagonal matrix, even in the frequency domain. In order to regain the FDE advantage of reduced computational complexity, we show that this matrix can be approximated by a block-diagonal matrix without performance loss. Moreover, our MMSE equalizer can be simplified to a low-complexity zero-forcing equalizer. The proposed techniques can be applied to any CPM scheme. To support this theory we present a new polyphase matrix model, valid for any block-based CPM system. Simulation results in a 60 GHz environment show that our reduced-complexity MMSE equalizer significantly outperforms the state of the art linear MMSE receiver for large modulation indices, while it performs only slightly worse for small ones. Wim Van Thillo, François Horlin, Jimmy Nsenga, Valéry Ramon, André Bourdoux, Rudy Lauwereins |
IEEE Trans. Wirel. Commun. | 6 |
| 2008 | Panel Session - Caution Ahead: The Road to Design and Manufacturing at 32 and 22 nmabstractStart of the above-titled section of the conference proceedings record. S. Turnoy, Peter Wintermayr, Robert C. Aitken, Rudy Lauwereins, J. Tracy Weed, V. Kiefer, J. Hartmann |
DATE | 4 |
| 2008 | Spectrum Sensing over SIMO Multi-Path Fading Channels Based on Energy DetectionabstractOne of the techniques for spectrum sensing is the energy detection of signals from primary transmitters. However, when detecting in a certain frequency band, this technique suffers from the random fading introduced by the channel. This problem can be alleviated by the use of multiple antennas introducing spatial diversity. The performance of the energy detector with multiple antennas has already been analytically analyzed for flat fading channels. In this paper, we propose analytical expressions for the performance of multi-antenna energy detector with multi-paths fading channels. These expressions will enable to compare for different channel power profiles the gains brought by the two sources of diversity: spatial (or antenna) diversity and multi-path diversity. This analytical performance is compared to simulations to show the validity of the used approximations. Santiago Rodriguez-Parera, Valéry Ramon, André Bourdoux, François Horlin, Rudy Lauwereins |
GLOBECOM | 5 |
| 2008 | Complexity Reduction of High-Performance Frequency Domain Equalization for CPMabstractContinuous phase modulations (CPM) have a perfectly constant envelope and are inherently robust against analog front-end nonidealities. Therefore, they have recently been proposed for several modern wireless standards. In this paper, we further develop an existing high-performance frequency domain equalization approach for CPM signals. First, we introduce a polyphase matrix model, valid for any block-based CPM system, which has clear advantages compared to existing models. Second, the model is used to reduce the complexity of the existing minimum mean square error channel equalizer. The proposed technique can be applied to any CPM scheme. Simulation results in a 60 GHz environment confirm that it does not cause any noticeable performance degradation. Wim Van Thillo, Jimmy Nsenga, Rudy Lauwereins, Valéry Ramon, André Bourdoux, François Horlin |
GLOBECOM | 3 |
| 2008 | Spectral regrowth analysis of band-limited offset-QPSKabstractIn this paper, we present an analytical analysis to predict the power spectral density (PSD) at the output of a nonlinear power amplifier (PA). We focus on offset quadrature phase shift keying (OQPSK) waveform band-limited by a square root raised cosine (SRRC) filter. This is one of the waveforms used in wideband code division multiple access (W-CDMA) wireless standard. We show that the PA output PSD obtained by our analytical analysis matches well the simulated PSD. Furthermore, we compare the PA output PSD of QPSK and OQPSK waveforms as a function of the SRRC filter roll-off. We conclude that for small roll-off, both QPSK and OQPSK experience almost the same level of spectral re-growth. As the roll-off increases, OQPSK becomes less sensitive to PA nonlinearity relative to QPSK. Jimmy Nsenga, Wim Van Thillo, André Bourdoux, Valéry Ramon, François Horlin, Rudy Lauwereins |
ICASSP | 6 |
| 2008 | Applying frequency domain equalization to precoded CPMabstractWe show how to apply Frequency Domain Equalization (FDE) to precoded Continuous Phase Modulation (CPM) systems. It is well known that differential precoding can be applied to the specific, popular class of CPM schemes with modulation index h = 1/2Q, where Q is any integer. This precoding halves the bit error rate (BER) compared to nonprecoded CPM without any overhead or complexity increase. We apply FDE to a block-based precoded CPM system. Therefore, we show that in addition to a cyclic prefix, two subblocks of data-dependent symbols have to be inserted in each block to cope with the memory in the CPM signal and to enable correct decoding by the receiver. We explain how to calculate these subblocks. Simulation results in a 60 GHz environment confirm that the BER is halved by precoding, and that this precoding is compatible with FDE using our new technique. Wim Van Thillo, Jimmy Nsenga, Rudy Lauwereins, André Bourdoux, Valéry Ramon, François Horlin |
ICASSP | 3 |
| 2008 | A New Symbol Block Construction for CPM with Frequency Domain EqualizationabstractWe present a new symbol block construction which yields a cyclic continuous phase modulated (CPM) signal to enable frequency domain equalization. It is known that in addition to a cyclic prefix, a subblock of data-dependent symbols has to be inserted in each block to cope with the memory in the CPM signal. We propose a new subblock, called intrafix, valid for any CPM scheme. Our intrafix is shorter than what is currently known in the literature, reducing the overhead. Moreover, it can be calculated on a per-block basis, without knowledge of previous blocks. We also prove that there are constraints on the length of both the intrafix and the total block by studying the influence of the modulation index. Simulation results in a 60 GHz environment show that our new block construction satisfies all requirements. Wim Van Thillo, Jimmy Nsenga, Rudy Lauwereins, Valéry Ramon, André Bourdoux, François Horlin |
ICC | 3 |
| 2007 | Low-Complexity Frequency Domain Equalization Receiver for Continuous Phase ModulationabstractA new approach for frequency-domain equalization of continuous phase modulated (CPM) signals is presented. In contrast with state-of-the-art receivers, we separate channel equalization on the one hand and CPM demodulation on the other. This separation enables us to calculate independently of the CPM scheme a new low-complexity zero-forcing channel equalizer. We also present a new high-performance minimum mean square error (MMSE) channel equalizer for any CPM scheme and a method to lower its complexity for a popular class of CPM schemes. Simulations show that our new MMSE equalizer significantly outperforms state-of-the-art linear receivers in a 60 GHz multipath environment. Wim Van Thillo, Jimmy Nsenga, Rudy Lauwereins, Valéry Ramon, André Bourdoux, François Horlin |
GLOBECOM | 3 |
| 2007 | The Generalized Linear Decomposition of Multilevel CPM SignalsabstractMultilevel continuous phase modulated (CPM) signals feature a perfectly constant envelope, attractive spectral properties and excellent power efficiency. However, their non-linear nature makes them less tractable and their processing more complex. Fortunately, a linear decomposition exists, allowing to apply linear signal processing techniques. This decomposition was originally developed for binary CPM schemes only, and is not suited for schemes with an integer modulation index. It was extended to multilevel CPM schemes, by decomposing the multilevel input sequence in a product of binary subsequences, and applying the binary decomposition to these subsequences. When one or more of these subsequences has an integer modulation index though, this technique fails. We present a general solution, and prove how many pulses are needed to represent the CPM signal in this particular case. The decomposition of the quaternary 3RC system with h = 1/2 is given as an example. A receiver based on this solution is presented. Wim Van Thillo, Jimmy Nsenga, François Horlin, André Bourdoux, Rudy Lauwereins |
ICASSP (3) | 5 |
| 2007 | Sensitivity to Front-End Non-Idealities of Low PAPR Modulation Schemes for Communications at 60 GHzabstractThe huge bandwidth available at 60 GHz will allow short range wireless communication to deliver bit rates over 1 Gbps. However the design of millimeter wave analog blocks is more critical than at lower frequencies, leading to a possible high non ideality of the radio front-end (FE). A suitable air interface for low cost, low power 60 GHz transceivers should thus use a modulation technique that has a high level of immunity to FE non-idealities. In this paper, we compare the sensitivity to FE non-idealities of two promising air interfaces namely offset quadrature phase shift keying (OQPSK) with frequency domain equalization (FDE) and continuous phase modulation (CPM) with time domain equalization (TDE). Our study focus on three main analog imperfections that are likely to have the biggest impact on the overall system performance: phase noise generated by the voltage control oscillator (VCO), clipping and quantization errors caused by the analog-to-digital converter (ADC) and non-linearity in the power amplifier (PA). Results show that CPM based air interface is more robust than OQPSK-based regarding the FE non-idealities. However the latter is less complex to design and offers much higher data rate than the former in the same bandwidth. Jimmy Nsenga, Wim Van Thillo, François Horlin, André Bourdoux, Rudy Lauwereins |
VTC Spring | 5 |
| 2006 | Platform independent optimisation of multi-resolution 3D content to enable universal media access
Klaas Tack, Gauthier Lafruit, Francky Catthoor, Rudy Lauwereins |
Vis. Comput. | 4 |
| 2005 | Wireless platforms: GOPS for cents and MilliWattsabstractIn recent years, data communication has overtaken voice as the main force behind the growth in wireless. With this has come a proliferation of standards ranging from wide area networks at one end of the spectrum to personal area networks on the other end. The opportunities offered by this truly ubiquitous connectivity are tremendous, and are leading to revolutionary chances in the way computer, communication, and consumer systems operate and interact.Providing the necessary flexibility to seamlessly interact with the multitude of emerging network models, as well as the muscle to support the demanding multimedia functionality in a mobile environment, presents some huge challenges to the developer of the wireless implementation platforms. The power budget of the mobile terminal is typically fixed by size considerations and operation time. Cost considerations further constrain the solution space.In response to these challenges, many solutions have been floated and experimented with ranging from multi-processor architectures, advanced DSPs, reconfigurable solutions and hardwired accelerators. While these innovations break new ground in the world of embedded architectures, many questions emerge such as efficiency, flexibility and programming model.This panel will presents a "bake-off" between a number of solutions that have emerged over the recent years. Francine Bacchini, Jan M. Rabaey, Allan Cox, Frank Lane, Rudy Lauwereins, Ulrich Ramacher, David Witt |
DAC | 5 |
| 2004 | Instruction buffering exploration for low energy VLIWs with instruction clusters
Tom Vander Aa, Murali Jayapala, Francisco Barat, Geert Deconinck, Rudy Lauwereins, Francky Catthoor, Henk Corporaal |
ASP-DAC | 5 |
| 2004 | System level design technology for realizing an ambient intelligent environment
Rudy Lauwereins |
ASP-DAC | 1 |
| 2004 | How Can System-Level Design Solve the Interconnect Technology Scaling Problem?abstractThe scaling of interconnect technology hits a red brick wall: interconnect delay and power do not follow Moore's law any more. The use of new materials like Cu and low-k alleviated the problem temporarily, but physical limits are being hit. What does this mean for system level design? The session starts with an embedded tutorial, given by an interconnect semiconductor technology expert, explaining the physics behind the interconnect problem and the degrees of freedom semiconductor technology offers system designers. Panelists will then express their thoughts and discuss with you how the interconnect problem can be solved by taking these degrees of freedom into account at the system design level. Views from industrial designers, CAD vendors, IC manufacturers and researchers will be presented. Francky Catthoor, Andrea Cuomo, Grant Martin, Patrick Groeneveld, Rudy Lauwereins, Karen Maex, Patrick van de Steeg, Ron Wilson |
DATE | 5 |
| 2004 | Design Methodology for a Tightly Coupled VLIW/Reconfigurable Matrix Architecture: A Case StudyabstractCoarse-grained reconfigurable architectures have seen growing importance recently. Design tools and methodology are essential to their success. Based on our previous work on modulo scheduling algorithms and a novel architecture with tightly coupled VLIW/reconfigurable matrix, we present a C-based design flow using an MPEG-2 decoder as a design example. The application is mapped to the architecture in less than one person-week starting from a software implementation. The kernel and overall speedup over the reference VLIW are 4.84 and 3.05 respectively. The case study shows that our methodology and architecture can deliver a competitive package in terms of design efforts and performance over other programmable architectures. Bingfeng Mei, Serge Vernalde, Diederik Verkest, Rudy Lauwereins |
DATE | 4 |
| 2004 | Design-Time Data-Access Analysis for Parallel Java Programs with Shared-Memory Communication Model
Richard Stahl, Francky Catthoor, Rudy Lauwereins, Diederik Verkest |
Euro-Par | 3 |
| 2004 | Network-on-Chip for Reconfigurable Systems: From High-Level Design Down to Implementation
T. Andrei Bartic, Dirk Desmet, Jean-Yves Mignolet, Théodore Marescaux, Diederik Verkest, Serge Vernalde, Rudy Lauwereins, Frédéric Robert |
FPL | 7 |
| 2004 | High-Level Data-Access Analysis for Characterisation of (Sub)task-Level Parallelism in JavaabstractIn the era of future embedded systems the designer is confronted with multi-processor systems both for performance and energy reasons. Exploiting (sub)task-level parallelism is becoming crucial because the instruction-level parallelism alone is insufficient. The challenge is to build compiler tools that support the exploration of the task-level parallelism in the programs. To achieve this goal, we have designed an analysis framework to evaluate the potential parallelism from sequential object-oriented programs. Parallel-performance and data-access analysis are the crucial techniques for estimation of the transformation effects. We have implemented support for platform-independent data-access analysis and profiling of Java programs, which is an extension to our earlier parallel-performance analysis framework. The toolkit comprises automated design-time analysis for performance and data-access characterisation, program instrumentation, program-profiling support and post-processing analysis. We demonstrate the usability of our approach on a number of realistic Java applications. Richard Stahl, Robert Pasko, Francky Catthoor, Rudy Lauwereins, Diederik Verkest |
HIPS | 4 |
| 2004 | Fast prototyping and refinement of complex dynamic data types in multimedia applications for consumer embedded devicesabstractPortable consumer devices are increasing their capabilities more and more and can now implement new multimedia algorithms that were reserved only for powerful workstations a few years ago. Unfortunately, the original design characteristics of such algorithms do not often allow them to be ported directly to current embedded devices. These algorithms share complex and intensive dynamic memory use and actual embedded systems cannot provide efficient general-purpose memory management as it is needed. As a result, dynamic memory optimizations are a requirement when porting these applications. Within these optimizations, the refinement of the dynamically (de)allocated abstract data type implementations in the complex multimedia applications involved is one of the most important and difficult parts for an efficient mapping of the algorithms on low-power and high-speed embedded consumer devices. We describe a high-level approach for modeling and refining complex data types using abstract derived classes in C++. This approach enables the multimedia developer to compose, evaluate and refine complex data types in a conceptually straightforward way, without a time-consuming programming effort David Atienza 0001, Marc Leeman, Francky Catthoor, Geert Deconinck, Jose Manuel Mendias, Vincenzo De Florio, Rudy Lauwereins |
ICME | 7 |
| 2004 | Run-time support for heterogeneous multitasking on reconfigurable SoCs
Théodore Marescaux, Vincent Nollet, Jean-Yves Mignolet, T. Andrei Bartic, W. Moffat, Prabhat Avasare, Paul Coene, Diederik Verkest, Serge Vernalde, Rudy Lauwereins |
Integr. | 10 |
| 2003 | Exploiting Loop-Level Parallelism on Coarse-Grained Reconfigurable Architectures Using Modulo Scheduling
Bingfeng Mei, Serge Vernalde, Diederik Verkest, Hugo De Man, Rudy Lauwereins |
DATE | 5 |
| 2003 | Infrastructure for Design and Management of Relocatable Tasks in a Heterogeneous Reconfigurable System-on-Chip
Jean-Yves Mignolet, Vincent Nollet, Paul Coene, Diederik Verkest, Serge Vernalde, Rudy Lauwereins |
DATE | 6 |
| 2003 | Panel Title: Reconfigurable Computing - Different Perspectives
Wolfgang Rosenstiel, Rudy Lauwereins, Ivo Bolsens, Chris Rowen, Yankin Tanurhan, Kees A. Vissers |
DATE | 2 |
| 2003 | Low Power Coarse-Grained Reconfigurable Instruction Set Processor
Francisco Barat, Murali Jayapala, Tom Vander Aa, Rudy Lauwereins, Geert Deconinck, Henk Corporaal |
FPL | 4 |
| 2003 | Networks on Chip as Hardware Components of an OS for Reconfigurable Systems
Théodore Marescaux, Jean-Yves Mignolet, T. Andrei Bartic, W. Moffat, Diederik Verkest, Serge Vernalde, Rudy Lauwereins |
FPL | 7 |
| 2003 | ADRES: An Architecture with Tightly Coupled VLIW Processor and Coarse-Grained Reconfigurable Matrix
Bingfeng Mei, Serge Vernalde, Diederik Verkest, Hugo De Man, Rudy Lauwereins |
FPL | 5 |
| 2003 | A framework for mapping scalable networked applications on run-time reconfigurable platformsabstractDue to the heterogeneous nature of networks and end-systems in distributed multimedia systems, multimedia applications should ideally be designed to counteract fluctuations in network bandwidth and end-system processing capacities for providing end users a certain degree of quality of service (QoS). This requirement can be satisfied with scalable applications. In addition, with the current evolution in run-time reconfigurable computing, run-time reconfigurable multimedia platforms are becoming increasingly viable. In this paper, an end-to-end delivery chain framework for mapping scalable networked multimedia applications on reconfigurable platforms is presented. The framework is demonstrated by a case study of a 3D game running on a prototype run-time reconfigurable platform. Nam Pham Ngoc 0001, Gauthier Lafruit, Jean-Yves Mignolet, Serge Vernalde, Geert Deconinck, Rudy Lauwereins |
ICME | 6 |
| 2003 | Performance Analysis for Identification of (Sub-)Task-Level Parallelism in Java
Richard Stahl, Robert Pasko, Luc Rijnders, Diederik Verkest, Serge Vernalde, Rudy Lauwereins, Francky Catthoor |
SCOPES | 6 |
| 2003 | A formant filtered physical model for wind instrumentsabstractWe report on our research concerning the calibration of physical models for sound synthesis. We combine waveguide physical modeling synthesis with formant filtering, by dividing the nonlinear description of the reed mechanism into a nonlinear part and an input-dependent linear filter. We elaborate on the calibration of the model and assess its performance by comparing it to a single-reed, cylindrical bore instrument, the clarinet. Axel Nackaerts, Bart De Moor, Rudy Lauwereins |
IEEE Trans. Speech Audio Process. | 3 |
| 2003 | Search space definition and exploration for nonuniform data reuse opportunities in data-dominant applicationsabstractEfficient exploitation of temporal locality in the memory accesses on array signals can have a very large impact on the power consumption in embedded data dominated applications. The effective use of an optimized custom memory hierarchy or a customized software controlled mapping on a predefined hierarchy is crucial for this. Only recently have effective systematic techniques to deal with this specific design step begun to appear. They are still limited in their exploration scope. In this paper we construct the design space by introducing three parameters which determine how and when copies are made between different levels in a hierarchy, and determine their impact on the total memory size, storage-related power consumption, and code complexity. Strategies are then established for an efficient exploration, such that cost-effective solutions for the memory size/power trade-off can be achieved. The effectiveness of the techniques is demonstrated for several real-life image processing algorithms. Tanja Van Achteren, Francky Catthoor, Rudy Lauwereins, Geert Deconinck |
ACM Trans. Design Autom. Electr. Syst. | 3 |
| 2002 | Data Reuse Exploration Techniques for Loop-Dominated ApplicationabstractEfficient exploitation of temporal locality in the memory accesses on array signals can have a very large impact art the power consumption in embedded data dominated applications. The effective use of an optimized custom memory hierarchy or a customized software controlled mapping on a predefined hierarchy, is crucial for this. Only recently have effective systematic techniques to deal with this specific design step begun to appear They were still limited in their exploration scope. In this paper we introduce an extended formalized methodology based on an analytical model of the data reuse of a signal. The cost parameters derived from this model define the search space to explore and allow us to exploit the maximum data reuse possible. The result is an automated design technique to find power efficient memory hierarchies and generate the corresponding optimized code. Tanja Van Achteren, Geert Deconinck, Francky Catthoor, Rudy Lauwereins |
DATE | 4 |
| 2002 | Reconfigurable SoC - What Will it Look Like?
J. Bryan Lewis, Ivo Bolsens, Rudy Lauwereins, Chris Wheddon, Bhusan Gupta, Yankin Tanurhan |
DATE | 3 |
| 2002 | Adding Hardware Support to the HotSpot Virtual Machine for Domain Specific Applications
Yajun Ha, Radovan Hipik, Serge Vernalde, Diederik Verkest, Marc Engels, Rudy Lauwereins, Hugo De Man |
FPL | 6 |
| 2002 | Creating a World of Smart Re-configurable Devices
Rudy Lauwereins |
FPL | 1 |
| 2002 | Interconnection Networks Enable Fine-Grain Dynamic Multi-tasking on FPGAs
Théodore Marescaux, T. Andrei Bartic, Diederik Verkest, Serge Vernalde, Rudy Lauwereins |
FPL | 5 |
| 2002 | DRESC: a retargetable compiler for coarse-grained reconfigurable architecturesabstractCoarse-grained reconfigurable architectures have become increasingly important in recent years. Automatic design or compiling tools are essential to their success. In this paper, we present a retargetable compiler for a family of coarse-grained reconfigurable architectures. Several key issues are addressed. Program analysis and transformation prepare dataflow for scheduling. Architecture abstraction generates an internal graph representation from a concrete architecture description. A modulo scheduling algorithm is key to exploit parallelism and achieve high performance. The experimental results show up to 28.7 instructions per cycle (IPC) over tested kernels. Bingfeng Mei, Serge Vernalde, Diederik Verkest, Hugo De Man, Rudy Lauwereins |
FPT | 5 |
| 2002 | Building a Virtual Framework for Networked Reconfigurable Hardware and Software Objects
Yajun Ha, Serge Vernalde, Patrick Schaumont, Marc Engels, Rudy Lauwereins, Hugo De Man |
J. Supercomput. | 5 |
| 2002 | Reconfigurable Instruction Set Processors from a Hardware/Software PerspectiveabstractThis paper presents the design alternatives for reconfigurable instruction set processors (RISP) from a hardware/software point of view. Reconfigurable instruction set processors are programmable processors that contain reconfigurable logic in one or more of its functional units. Hardware design of such a type of processors can be split in two main tasks: the design of the reconfigurable logic and the design of the interfacing mechanisms of this logic to the rest of the processor. Among the most important design parameters are: the granularity of the reconfigurable logic, the structure of the configuration memory, the instruction encoding format, and the type of instructions supported. On the software side, code generation tools require new techniques to cope with the reconfigurability of the processor. Aside from traditional techniques, code generation requires the creation and evaluation of new reconfigurable instructions and the selection of instructions to minimize reconfiguration time. The most important design alternative on the software side is the degree of automatization present in the code generation tools. Francisco Barat, Rudy Lauwereins, Geert Deconinck |
IEEE Trans. Software Eng. | 2 |
| 2001 | Virtual Java/FPGA interface for networked reconfigurationabstractA virtual interface between Java and FPGA for networked reconfiguration is presented. Through the Java/FPGA interface, Java applications can exploit hardware accelerators with FPGAs for both functional flexibility and performance acceleration. At the same time, the interface is platform independent. It enables the networked application developers to design their applications with only one interface in mind when considering the interfacing issues. The virtual interface is part of our work to build a platform-independent deployment framework for the networked services. In the framework, both the software and hardware components of services can be platform independently described and deployed. Yajun Ha, Geert Vanmeerbeeck, Patrick Schaumont, Serge Vernalde, Marc Engels, Rudy Lauwereins, Hugo De Man |
ASP-DAC | 6 |
| 2001 | Task concurrency management methodology summaryabstractThis paper summarizes a new methodology for the design of concurrent dynamic real-time embedded systems. An embedded system can be specified at a grey-box abstraction level in a combined MTG-CDFG model. The authors believe that task concurrency management can be implemented in four major steps. Firstly, the grey box model is built, including the necessary concurrency extraction. Then transformations are applied on the specified MTG-CDFG to increase the opportunities for concurrency exploration and cost minimization. Then static scheduling will be applied on the design time analyzable parts of the grey-box model, including processor assignment in the multiple processor context. Finally, a dynamic scheduler will schedule the dynamic and coarse-grain constructs at run time on the given platform while making trade-offs based on Pareto curves. Chun Wong, Paul Marchal, Francky Catthoor, Hugo De Man, Aggeliki S. Prayati, Nathalie Cossement, Rudy Lauwereins, Diederik Verkest |
DATE | 8 |
| 2001 | A SW/HW Interface API for Java/FPGA Co-Designed Applets
Yajun Ha, Patrick Schaumont, Serge Vernalde, Marc Engels, Rudy Lauwereins, Hugo De Man |
FCCM | 5 |
| 2001 | CRISP: A Template for Reconfigurable Instruction Set Processors
Pieter Op de Beeck, Francisco Barat, Murali Jayapala, Rudy Lauwereins |
FPL | 4 |
| 2001 | Development of a Design Framework for Platform-Independent Networked Reconfiguration of Software and Hardware
Yajun Ha, Bingfeng Mei, Patrick Schaumont, Serge Vernalde, Rudy Lauwereins, Hugo De Man |
FPL | 5 |
| 1999 | TIRAN: Flexible and Portable Fault Tolerance Solutions for Cost Effective Dependable Applications
Oliver Botti, Vincenzo De Florio, Geert Deconinck, Flavio Cassinari, Susanna Donatelli, Andrea Bobbio, Axel Klein, Holger Küfner, Rudy Lauwereins, Erwin M. Thurner, Eric Verhulst |
Euro-Par | 9 |
| 1998 | Rapid prototyping of an adaptive noise canceler using GRAPE
Luc De Coster, Rudy Lauwereins, Jean A. Peperstraete |
Signal Process. | 2 |
| 1997 | Data Memory Minimisation for Synchronous Data Flow Graphs Emulated on DSP-FPGA TargetsabstractThe paper presents an algorithm to determine the close-to-smallestpossible data buffer sizes for arbitrary synchronous dataflow (SDF) applications, such that we can guarantee the existenceof a deadlock free schedule. The presented algorithm fits inthe design flow of GRAPE, an environment for the emulation andimplementation of digital signal processing (DSP) systems onarbitrary target architectures, consisting of programmable DSPprocessors and FPGAs. Reducing the size of data buffers is ofhigh importance when the application will be mapped on FieldProgrammable Gate Arrays (FPGA), since register resources arerather scarce. Marleen Adé, Rudy Lauwereins, Jean A. Peperstraete |
DAC | 2 |
| 1997 | EFTOS: A Software Framework for More Dependable Embedded HPC Applications
Geert Deconinck, Vincenzo De Florio, Rudy Lauwereins, Theodora A. Varvarigou |
Euro-Par | 3 |
| 1997 | User-triggered checkpointing: system-independent and scalable application recoveryabstractUser-triggered checkpointing and rollback is proposed as a system-independent and flexible way to integrate backward error recovery in long-running, computation-intensive message-passing applications on large parallel multicomputers. It employs library calls to coordinate the checkpointing, allowing a non-blocking and scalable approach that requires no protocol to save a consistent state because the coordination among the processes is implicit. The explicit indication of the checkpoint contents (i.e. the items of which the state must be saved) allows one to significantly reduce the amount of checkpoint data and the overhead. In contrast to other checkpointing approaches, the implementation does not rely on system-dependent features (like saving register-values or communication status) to save the state. Instead, re-executing the first part of the application brings the system-specific items into a consistent state with the rest of the checkpoint contents that is restored from the saved checkpoint data. Geert Deconinck, Rudy Lauwereins |
ISCC | 2 |
| 1996 | Implementing DSP applications on heterogeneous targets using minimal size data buffersabstractThe paper presents an algorithm to determine the smallest possible data buffer sizes for arbitrary synchronous data flow (SDF) applications, such that we can guarantee the existence of a deadlock free schedule. The presented algorithm fits in the design flow of GRAPE, an environment for the emulation and implementation of digital signal processing (DSP) systems on arbitrary target architectures, consisting of programmable DSP processors and FPGAs. Reducing the size of data buffers is of high importance when the application will be mapped on Field Programmable Gate Arrays (FPGA), since register resources are rather scarce. Marleen Adé, Rudy Lauwereins, Jean A. Peperstraete |
RSP | 2 |
| 1996 | On the Design and Implementation of Broadcast and Global Combine Operations Using the Postal ModelabstractThere are a number of models that were proposed in recent years for message passing parallel systems. Examples are the postal model and its generalization the LogP model. In the postal model a parameter /spl lambda/ is used to model the communication latency of the message-passing system. Each node during each round can send a fixed-size message and, simultaneously, receive a message of the same size. Furthermore, a message sent out during round r will incur a latency of /spl lambda/ and will arrive at the receiving node at round r+/spl lambda/-1. Our goal in this paper is to bridge the gap between the theoretical modeling and the practical implementation. In particular, we investigate a number of practical issues related to the design and implementation of two collective communication operations, namely, the broadcast operation and the global combine operation. Those practical issues include, for example, (1) techniques for measurement of the value of /spl lambda/ on a given machine, (2) creating efficient broadcast algorithms that get the latency h and the number of nodes n as parameters and (3) creating efficient global combine algorithms for parallel machines with /spl lambda/ which is not an integer. We propose solutions that address those practical issues and present results of an experimental study of the new algorithms on the Inter Delta machine. Our main conclusion is that the postal model can help in performance prediction and tuning, for example, a properly tuned broadcast improves the known implementation by more than 20%. Jehoshua Bruck, Luc De Coster, Natalie Dewulf, C. T. Howard Ho, Rudy Lauwereins |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 1995 | Cyclo-static data flowabstractThe high sample-rates involved in many DSP-applications, require the use of static schedulers wherever possible. The construction of static schedules however is classically limited to applications that fit in the synchronous data flow model. In this paper we present cyclo-static data flow as a model to describe applications with a cyclically changing behaviour. We give both a necessary and sufficient condition for the existence of a static schedule for a cyclo-static data flow graph and show how such a schedule can be constructed. The example of a video encoder is used to illustrate the importance of cyclo-static data flow for real-life DSP-systems. Greet Bilsen, Marc Engels, Rudy Lauwereins, Jean A. Peperstraete |
ICASSP | 3 |
| 1995 | Hardware-software codesign with GRAPEabstractGRAPE-II (Graphical Rapid Prototyping Environment-II) is a hardware-software codesign environment for the real-time functional emulation of synchronous DSP systems. It allows one to specify the application's data dependency graph in a target-machine-independent way. After specifying the heterogeneous target machine's architecture, it estimates the resources needed by each application subtask. Based on these requirements, it assigns the subtasks to specific target devices at compile-time, be they processors or FPGAs, establishes routing paths and determines a static schedule. It generates a main shell for each target device and generates intra-device and inter-device communication code. After downloading the executable images on to the target machine, it allows the designer to modify end-user controls and application settings at run-time. This paper situates the tool in the application design cycle, explains GRAPE-II's design flow and shows the advantages of hardware-software codesign by evaluating the achievable sampling frequency for a small example application. Marleen Adé, Rudy Lauwereins, Jean A. Peperstraete |
RSP | 2 |
| 1995 | A Compact Fault-tolerant, Deadlock-free, Minimal Routing Algorithm for n-Dimensional Wormhole Switching Based Meshes
Johan Vounckx, Geert Deconinck, Rudy Lauwereins |
SIROCCO | 3 |
| 1995 | PDG: A process-level debugger for concurrent programs in the GRAPE parallel programming environment
Chris Caerts, Rudy Lauwereins, Jean A. Peperstraete |
Future Gener. Comput. Syst. | 2 |
| 1994 | VLSI complexity reduction by piece-wise approximation of the sigmoid function
Valeriu Beiu, Jean A. Peperstraete, Joos Vandewalle, Rudy Lauwereins |
ESANN | 4 |
| 1994 | Buffer memory requirements in DSP applicationsabstractStudies synchronous multi-rate data flow graphs to determine the minimal required buffer sizes that still guarantee the construction of a deadlock-free static schedule. We develop a rule to quickly analyze a graph's consistency. A graph is split up into single and parallel paths. Single paths are analysed, as well as the most frequent parallel paths. The results are used in the rapid prototyping environment GRAPE-II in the case where the emulation hardware contains FPGAs, or when memory is critical.> Marleen Adé, Rudy Lauwereins, Jean A. Peperstraete |
RSP | 2 |
| 1994 | Geometric parallelism and cyclo-static data flow in GRAPE-IIabstractDescribes two novel features that are supported in GRAPE-II (Graphical RApid Prototyping Environment): geometric parallelism and cyclo-static data flow. GRAPE-II is intended as a system level tool for the rapid prototyping of digital signal processing (DSP) applications on multiprocessors. GRAPE-II fully supports code generation for multi-rate and asynchronous DSP applications on heterogeneous target multiprocessors. The first feature detailed in the paper, geometric parallelism, allows the programmer to efficiently specify data parallel operations, where multiple identical functions operate on different data sets. The second feature, cyclo-static data flow, enables the specification of cyclicly changing data dependencies, while still leading to static schedules.> Rudy Lauwereins, Piet Wauters, Marleen Adé, Jean A. Peperstraete |
RSP | 1 |
| 1994 | Fault-Tolerant Compact Routing Based on Reduced Structural Information in Wormhole-Switching Based Networks
Johan Vounckx, Geert Deconinck, Rudy Lauwereins, Jean A. Peperstraete |
SIROCCO | 3 |
| 1993 | Efficient decomposition of comparison and its applications
Valeriu Beiu, Jean A. Peperstraete, Joos Vandewalle, Rudy Lauwereins |
ESANN | 4 |
| 1993 | Development of a load balancing tool for the GRAPE rapid prototyping environmentabstractIn the graphical programming environment GRAPE-II intended for rapid prototyping of DSP ASICs on a multi-processor, a load balancing tool is required to map the different jobs of the DSP application on the multi-processor. Since most of the existing load balancing algorithms perform less well when they have to handle large or complex applications, the authors develop a tool that is better suited for such problems. This tool is based on three main techniques. (1) By exploiting the hierarchy that exists in the application graph, the complexity for each of the tools can be reduced. (2) Splitting the load balancing in sub tasks leads to three smaller search spaces instead of one big space. (3) Each of the smaller search spaces on its turn is reduced by using appropriate heuristic rules. As an example of this approach the scheduling tool and the basics of the assignment tool are presented.> Greet Bilsen, Marc Engels, Rudy Lauwereins, Jean A. Peperstraete |
RSP | 3 |
| 1993 | PDG: a process-level debugger for concurrent programs in the GRAPE rapid prototyping environmentabstractThe authors describe the process-level debugger of GRAPE, an environment for the rapid prototyping of digital signal processing systems on parallel computers. This debugger allows one to debug concurrent programs that are based on communicating sequential processes. Its unique feature is that it clearly separates the identification of erroneous processes from the exact localisation of the bug on the source-level. This divide-and-conquer approach is absolutely necessary for debugging complex parallel programs in a fast and systematic way, which is a key issue for rapid prototyping. The process-level debugging approach is based on an animation of program behaviour on its hierarchical graphical representation. Graphical views are used that reflect the programmer's mental picture of the application. Hierarchy allows one to employ a top-down debugging approach in which one successively refines the search-space by zooming in on suspect processes first-time-right. They consider both hierarchy that hides algorithmic details, and hierarchy that hides implementation issues. During animation a debugging kernel implementing a record-replay mechanism guarantees reproducible program behaviour even for programs containing asynchronisities.> Chris Caerts, Rudy Lauwereins, Jean A. Peperstraete |
RSP | 2 |
| 1992 | GRAPE-II: a tool for the rapid prototyping of multi-rate asynchronous DSP applications on heterogeneous multiprocessorsabstractThe second version of the graphical programming environment (GRAPE) is described. It is intended as a tool for the rapid prototyping of digital signal processing (DSP) application-specific integrated circuits (ASICs) on a multiprocessor. GRAPE-II fully supports multirate and asynchronous DSP applications and heterogeneous target multiprocessors. The extensions that are required to the programming model and the intermediate specification language to support multirate and asynchronous operation are described. The programming model is clarified by an example. The global structure of GRAPE-II is presented. The status of the project is indicated.> Rudy Lauwereins, Marc Engels, Jean A. Peperstraete |
RSP | 1 |
| 1990 | Parallel processing enables the real-time emulation of DSP ASICsabstractThe CASE tool GRAPE (GRAphical Programming Environment) is presented which allows for easy programming, compiling, debugging and evaluating high frequency real-time DSP algorithms. Its main distinctive feature is that the tool spans the whole design process, ranging from analysis over simulation and emulation up to implementation on general purpose DSP multiprocessors or integration on an Application Specific Integrated Circuit (ASIC). The DSP multiprocessor can be the target hardware or can be used for real-time emulation or accelerated simulation of an ASIC. The paper mainly focuses on the CASE tool and on the design of a flexible multiprocessor board for the real-time emulation of DSP ASIC's.> Rudy Lauwereins, Marc Engels, Jean A. Peperstraete |
RSP | 1 |
| 1987 | An Integrated Software-Hardware Multiprocesor Project
Rudy Lauwereins, Jean A. Peperstraete |
ICPP | 1 |