Gerard J. M. Smit

dblp:s/GerardJMSmit · DBLP profile ↗
← Back
74ranked-venue papers
6as first author
0since 2021 · last 2016
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 43 · 4 first-authorComputer networks · 11 · 1 first-authorSoftware engineering, systems software and programming languages · 6Artificial intelligence and machine learning · 2Graphics, computer vision, multimedia, augmented reality and games · 2Databases, data management, data science and information retrieval · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1Theory of computation · 1Applied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer networks
1 paper
Physical-layer communications · 100%
Computer architecture, parallel and distributed computing, and storage systems
3 papers
Integrated circuit design · 54% Embedded and real-time systems · 28% Processor architecture and microarchitecture · 8%

Topics — the 11 heaviest of 12, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Physical-layer communications
signal processing for communications
0.212016
Spectrum Efficient, Localized, Orthogonal Waveforms: Closing the Gap With the Balian-Low Theorem · IEEE Trans. Commun. 2016
Physical-layer communications › modulation
waveform design
0.212016
Spectrum Efficient, Localized, Orthogonal Waveforms: Closing the Gap With the Balian-Low Theorem · IEEE Trans. Commun. 2016
Integrated circuit design
low-power circuit design
0.112012
Sabrewing: A lightweight architecture for combined floating-point and integer arithmetic · ACM Trans. Archit. Code Optim. 2012
Physical-layer communications
modulation
0.112016
Spectrum Efficient, Localized, Orthogonal Waveforms: Closing the Gap With the Balian-Low Theorem · IEEE Trans. Commun. 2016
Physical-layer communications › modulation › multicarrier modulation
OFDM
0.112016
Spectrum Efficient, Localized, Orthogonal Waveforms: Closing the Gap With the Balian-Low Theorem · IEEE Trans. Commun. 2016
Embedded and real-time systems › model-based design
dataflow modeling
0.112007
Efficient Computation of Buffer Capacities for Cyclo-Static Dataflow Graphs · DAC 2007
Embedded and real-time systems
real-time system design
0.112007
Efficient Computation of Buffer Capacities for Cyclo-Static Dataflow Graphs · DAC 2007
Processor architecture and microarchitecture › special-purpose processor
digital signal processor
0.012012
Sabrewing: A lightweight architecture for combined floating-point and integer arithmetic · ACM Trans. Archit. Code Optim. 2012
Energy-efficient computing
energy-efficient architecture
0.011999
Octopus: Embracing the Energy Efficiency of Handheld Multimedia Computers · MobiCom 1999
Interconnection networks and networks-on-chip › interconnect architecture
reconfigurable interconnect
0.011999
Octopus: Embracing the Energy Efficiency of Handheld Multimedia Computers · MobiCom 1999
Embedded and real-time systems › mobile computing
mobile computing platforms
0.011999
Octopus: Embracing the Energy Efficiency of Handheld Multimedia Computers · MobiCom 1999

Methods — techniques the papers use, named apart from their topics

hermite functions · 0.2balian-low theorem · 0.2VHDL · 0.1IEEE-754 compliance testing · 0.1polynomial-time algorithm · 0.1back-pressure analysis · 0.1ATM switching · 0.0
YearPublicationVenuePosition
2016 Spectrum Efficient, Localized, Orthogonal Waveforms: Closing the Gap With the Balian-Low Theorem
abstract
The Balian-Low theorem (BLT) states the fundamental impossibility to design waveforms for L2(ℝ), which 1) form an orthogonal set, 2) are time-frequency localized, and 3) attain a critical waveform density such that they form an orthogonal basis. This article closes the gap between existing waveform designs and the BLT. The main contribution is the design of orthogonal, time-frequency localized, spectrum efficient waveforms for hexagonal lattices. The waveform design is adaptive by a single design parameter, which tradesoff time-frequency localization with the waveform density. As the orthogonalization procedure is based on employing the minimum number of most time-frequency localized waveforms (Hermite functions) it is argued that the results may be optimal in terms of combined spectrum efficiency and time-frequency localization. An example is provided for waveforms for a hexagonal lattice, which are quasi-orthogonal, time-frequency localized, and up to 99% of the critical waveform density. Although the designed waveforms are not strictly orthogonal, their cross-correlation can be made arbitrarily small. The robustness in doubly dispersive channels and the efficiency for multiuser scenarios are discussed and compared to conventional orthogonal frequency division multiplexing (OFDM).
C. Willem Korevaar, André B. J. Kokkeler, Pieter-Tjerk de Boer, Gerard J. M. Smit
IEEE Trans. Commun.4
2015 Incremental Analysis of Cyclo-Static Synchronous Dataflow Graphs
abstract
In this article, we present a mathematical characterisation of admissible schedules of cyclo-static dataflow ( csdf ) graphs. We demonstrate how algebra ic manipulation of this characterization is related to unfolding csdf actors and how this manipulation allows csdf graphs to be transformed into mrsdf graphs that are equivalent , in the sense that they admit the same set of schedules. The presented transformation allows the rich set of existing analysis techniques for mrsdf graphs to be applied to csdf graphs and generalizes the well-known transformations from csdf and mrsdf into hsdf . Moreover, it gives rise to an incremental approach to the analysis of csdf graphs, where approximate analyses are combined with exact transformations. We show the applicability of this incremental approach by demonstrating its effectiveness on the problem of optimizing buffer sizes under a throughput constraint.
Robert de Groote, Philip K. F. Hölzenspies, Jan Kuper, Gerard J. M. Smit
ACM Trans. Embed. Comput. Syst.4
2014 Declaratively Programmable Ultra Low-Latency Audio Effects Processing on FPGA
Math Verstraelen, Jan Kuper, Gerard J. M. Smit
DAFx3
2014 Analytic Clock Frequency Selection for Global DVFS
abstract
Computers can reduce their power consumption by decreasing their speed using Dynamic Voltage and Frequency Scaling (DVFS). A form of DVFS for multicore processors is global DVFS, where the voltage and clock frequency is shared among all processor cores. Because global DVFS is efficient and cheap to implement, it is used in modern multicore processors like the IBM Power 7, ARM Cortex A9 and NVIDIA Tegra 2. This theory oriented paper discusses energy optimal DVFS algorithms for such processors. There are no known provably optimal algorithms that minimize the energy consumption of nontrivial real-time applications on a global DVFS system. Such algorithms only exist for single core systems, or for simpler application models. While many DVFS algorithms focus on tasks, this theoretical study is conceptually different and focuses on the amount of parallelism. We provide a transformation from a multicore problem to a single core problem, by using the amount of parallelism of an application. Then existing single core algorithms can be used to find the optimal solution. Furthermore, we extend an existing single core algorithm such that it takes static power into account.
Marco Gerards, Johann L. Hurink, Philip K. F. Hölzenspies, Jan Kuper, Gerard J. M. Smit
PDP5
2014 Single-rate approximations of cyclo-static synchronous dataflow graphs
abstract
Exact analysis of synchronous dataflow (sdf) graphs is often considered too costly, because of the expensive transformation of the graph into a single-rate equivalent. As an alternative, several authors have proposed approximate analyses. Existing approaches to approximation are based on the operational semantics of an sdf graph.
Robert de Groote, Philip K. F. Hölzenspies, Jan Kuper, Gerard J. M. Smit
SCOPES4
2013 A dataflow-inspired CGRA for streaming applications
abstract
The herein presented research is motivated by the need for reconfigurable, flexible computing arrays targeted at streaming applications that contain a large degree of instruction-level parallelism. Such arrays are usually referred to as coarse-grained reconfigurable arrays (CGRAs). CGRAs are composed of small, reconfigurable cores that are interconnected to form a computing grid. Here, we present a complete CGRA, consisting of an architecture and a programming language. Both the architecture and the programming language are inspired by the principles found in dataflow.
Anja Niedermeier, Jan Kuper, Gerard J. M. Smit
FPL3
2013 Peak-to-average power reduction by rotation of the time-frequency representation
abstract
Multi-carrier communication is associated with a high peak-to-average power ratio (PAPR). A new PAPR reduction method is proposed which is based on rotating the time-frequency representation of a transmit signal, prior to transmission. In general, a time-frequency rotation of a multi-carrier signal would change the signal basis, affecting the robustness of the transmit signal in fading channels. Exceptions are transmit signals constructed by (modulated) Hermite functions. The PAPR has been analyzed for transmit signals based on 64 Hermite functions. Allowing a rotation over 16 angles in time-frequency, the PAPR, which occurs with a probability of 10−3, is reduced by 3.8 dB at the cost of an increased computational complexity and a minor loss in spectral efficiency.
C. Willem Korevaar, Pieter-Tjerk de Boer, André B. J. Kokkeler, Gerard J. M. Smit
GLOBECOM4
2013 Fourier-hermite communications; where Fourier meets Hermite
abstract
A new signal set, based on the Fourier and Hermite signal bases, is introduced. It combines properties of the Fourier basis signals with the perfect time-frequency localization of the Hermite functions. The signal set is characterized by both a high spectral efficiency and good time-frequency localization. Its robustness against time-frequency shifts is assessed and compared to Hermite and Fourier basis signals. The Fourier-Hermite signal set is particularly designed for communications in spectrum-scarce environments.
C. Willem Korevaar, André B. J. Kokkeler, Pieter-Tjerk de Boer, Gerard J. M. Smit
ICASSP4
2013 A correlating receiver for ES-OFDM using multiple antennas
abstract
Extended Symbol OFDM (ES-OFDM) is applied in case of a multiple antenna receiver. The receiver architecture is based on the observation that OFDM constellation points can be determined by means of correlation. Summing correlations between multiple antennas leads to an interferometer receiver. This approach gives the freedom to choose which correlations are summed. Three antenna structures are explored: a Uniform Linear Array (ULA) and a sparse array where all correlations are summed and a sparse array where only a selection of correlations are summed. The sensitivity of the Bit Error Rate (BER) of an ES-OFDM communication link to an interfering source from different directions is studied. The ULA leads to a relatively wide BER main lobe, the range of angles around the Direction of Arrival of the ES-OFDM signal where the BER is high. Outside this range, the interfering source is suppressed to low BERs, in many cases beyond requirements. By using sparse arrays, the width of the BER main lobe can be traded against the BER levels outside the BER main lobe. This effect is shown for a sparse array where all possible correlations are summed. By summing only those correlations that lead to a uniform co-array, BER levels outside the BER main lobe are lower for an interferometer receiver compared to a traditional beamforming receiver.
André B. J. Kokkeler, Gerard J. M. Smit
ICC2
2013 Nonminimum-phase channel equalization using all-pass CMA
abstract
A nonminimum-phase channel can always be decomposed into a minimum-phase part and an all-pass part. In our approach, called all-pass CMA, the dimensionality of the CMA algorithm has been reduced to improve blind equalization of a nonminimum-phase channel's all-pass part. The dimensionality reduction has been performed by parameterizing the CMA cost function in terms of the nonminimum-phase zero location of the all-pass part to be compensated. Currently, all-pass CMA can only compensate a single nonminimum-phase zero. However, compared to CMA, it typically provides a faster and more accurate compensation of this zero.
Koen C. H. Blom, Marco Gerards, André B. J. Kokkeler, Gerard J. M. Smit
PIMRC4
2013 Selection of tests for outlier detection
abstract
Integrated circuits are tested thoroughly in order to meet the high demands on quality. As an additional step, outlier detection is used to detect potential unreliable chips such that quality can be improved further. However, it is often unclear to which tests outlier detection should be applied and how the parameters must be set, such that outliers are detected and yield loss remains limited. In this paper we introduce a mathematical framework, that given a set of target devices, can select tests for outlier detection and set the parameters for each outlier detection method. We provide results on real world data and analyze the resulting yield loss and missed targets.
Harm C. M. Bossers, Johann L. Hurink, Gerard J. M. Smit
VTS3
2013 Modular Neural Tile Architecture for Compact Embedded Hardware Spiking Neural Network
Sandeep Pande, Fearghal Morgan, Seamus Cawley, Tom M. Bruintjes, Gerard J. M. Smit, Brian McGinley, Snaider Carrillo, Jim Harkin, Liam McDaid
Neural Process. Lett.5
2013 Fixed latency on-chip interconnect for hardware spiking neural network architectures
Sandeep Pande, Fearghal Morgan, Gerard J. M. Smit, Tom M. Bruintjes, Jochem H. Rutgers, Brian McGinley, Seamus Cawley, Jim Harkin, Liam McDaid
Parallel Comput.3
2012 Evaluation of a Connectionless NoC for a Real-Time Distributed Shared Memory Many-Core System
abstract
Real-time embedded systems like smartphones tend to comprise an ever increasing number of processing cores. For scalability and the need for guaranteed performance, the use of a connection-oriented network-on-chip (NoC) is advocated. Furthermore, a distributed shared memory architecture is preferred as it simplifies software development for a multicore system. In this paper, experimental evidence is provided, showing that replacing a connection-oriented NoC by a connectionless one in a distributed shared memory system reduces the hardware costs and improves the performance. We observed that our FPGA could only support an 8-core system with a connection-oriented NoC. We exchanged the NoC with our tree-shaped, connectionless network and a ring, allowing a 32-core system in the same FPGA, mainly because of a reduced number of physical connections. Although the analytical worst-case performance slightly decreased, measurements show that the latency of latency-critical memory reads was reduced by 52% on average.
Jochem H. Rutgers, Marco Bekooij, Gerard J. M. Smit
DSD3
2012 High level structural description of streaming applications
abstract
The research has been motivated by the desire for a straightforward implementation of an application described as a dataflow graph on a multicore architecture. The design has been based on an already existing language that inherently has a notion of structure: the functional programming language Haskell. Haskell has been used both for embedding the grammar for proposed language as well as to describe streaming applications with the proposed language.
Anja Niedermeier, Jan Kuper, Gerard J. M. Smit
FPL3
2012 Synchronization and matched filtering in time-frequency using the sunflower spiral
abstract
Synchronization and matched filtering of signals in time dispersive, frequency dispersive and time-frequency dispersive channels are addressed in this paper. The ‘eigenfunctions’ of these channels form the signal sets under investigation. While using channel-eigenfunctions is a first requirement for undistorted data transmission, a second necessity is to achieve good synchronization over the domains of time and frequency. The synchronization problem in time-frequency for non-stationary signals is discussed. A spiral correlation method is proposed to achieve synchronization and matched filtering in time-frequency. Spiral correlation, using the pattern of a sunflower, is simulated and evaluated. It is argued that partial spiral correlation can lead to a significant reduction in computational complexity necessary for synchronization. Generalizations and identities based on the fractional Fourier transform are provided which omit the need for fractional delay filters.
C. Willem Korevaar, André B. J. Kokkeler, Pieter-Tjerk de Boer, Gerard J. M. Smit
GLOBECOM4
2012 Multilevel Unit Commitment in Smart Grids
Maurice G. C. Bosman, Albert Molderink, Vincent Bakker, Gerard J. M. Smit, Johann L. Hurink
ICORES4
2012 Sabrewing: A lightweight architecture for combined floating-point and integer arithmetic
abstract
In spite of the fact that floating-point arithmetic is costly in terms of silicon area, the joint design of hardware for floating-point and integer arithmetic is seldom considered. While components like multipliers and adders can potentially be shared, floating-point and integer units in contemporary processors are practically disjoint. This work presents a new architecture which tightly integrates floating-point and integer arithmetic in a single datapath. It is mainly intended for use in low-power embedded digital signal processors and therefore the following design constraints were important: limited use of pipelining for the convenience of the compiler; maintaining compatibility with existing technology; minimal area and power consumption for applicability in embedded systems. The architecture is tailored to digital signal processing by combining floating-point fused multiply-add and integer multiply-accumulate . It could be deployed in a multi-core system-on-chip designed to support applications with and without dominance of floating-point calculations. The VHDL structural description of this architecture is available for download under BSD license. Besides being configurable at design time, it has been thoroughly checked for IEEE-754 compliance by means of a floating-point test suite originating from the IBM Research Labs. A proof-of-concept has also been implemented using STMicroelectronics 65nm technology. This prototype supports 32-bit signed two's complement integers and 41-bit (8-bit exponent and 32-bit significand) floating-point numbers. Our evaluations show that over 67% energy and 19% area can be saved compared to a reference design in which floating-point and integer arithmetic are implemented separately. The area overhead caused by combining floating-point and integer is less than 5%. Implemented in ST's general-purpose CMOS technology, the design can operate at a frequency of 1.35GHz, while 667MHz can be achieved in low-power CMOS. Considering that the entire datapath is partitioned in just three pipeline stages, and the fact that the design is intended for use in the low-power domain, these frequencies are adequate. They are in fact competitive with current technology low-power floating-point units. Post-layout estimates indicate that the required area of a low-power implementation can be as small as 0.04mm 2 . Power consumption is on the order of several milliwatts. Strengthened by the fact that clock gating could reduce power consumption even further, we think that a shared floating-point and integer architecture is a good choice for signal processing in low-power embedded systems.
Tom M. Bruintjes, Karel H. G. Walters, Sabih H. Gerez, Egbert Molenkamp, Gerard J. M. Smit
ACM Trans. Archit. Code Optim.5
2011 Online Univariate Outlier Detection in Final Test: A Robust Rolling Horizon Approach
abstract
We present an online outlier detection method that is applicable to Final Test. Test limits are constructed based on previous measurements and robust statistics are used to ensure a stable start to the method. We analyze our method using real-world data. Furthermore, we identified some cases which can result in performance degradation, but most experiments show that our method is robust to outliers and able to detect them in an online setting.
Harm C. M. Bossers, Johann L. Hurink, Gerard J. M. Smit
ETS3
2011 Mixed continuous/discrete time modelling with exact time adjustments
abstract
Many systems interact with their physical environment. Design of such systems need a modelling and simulation tool which can deal with both the continuous and discrete aspects. However, most current tools are not adequately able to do so, as they implement both continuous and discrete time signals as consisting of separate values at a single global simulation clock. The consequence is that simulation, of a time delay for example, either yields inaccurate results or becomes inefficient.
Kenneth C. Rovers, Jan Kuper, Marcel D. van de Burgwal, André B. J. Kokkeler, Gerard J. M. Smit
IWCMC5
2011 On the Effects of Input Unreliability on Classification Algorithms
Ardjan Zwartjes, Majid Bahrepour, Paul J. M. Havinga, Johann L. Hurink, Gerard J. M. Smit
MobiQuitous5
2011 Exploring the Use of Two Antennas for Crosscorrelation Spectrum Sensing
abstract
Spectrum sensing is one of the key characteristics of a cognitive radio. Energy detection provides maximum flexibility by not relying on any prior knowledge, but suffers from an SNR-wall due to noise uncertainty. Crosscorrelation of the outputs of two receiver paths is a technique to reduce the noise level of the total receiver, and hence improves the SNR. The reduction of the noise is limited by correlated noise originating from shared components near the antenna. In this paper we explore the use of a separate antenna for each receiver for crosscorrelation spectrum sensing. One immediate advantage is that due to the removal of the splitter, which was necessary to interface the single antenna to two receivers, the SNR improves, significantly reducing the required measurement time. A lot of the noise correlation can be removed, leading to a lower residual noise floor. The noise at each antenna will still be partially correlated due to mutual coupling, spatial noise correlation and man-made noise. We show that some signal power can be lost in the sensing process due to partial decorrelation of the signal at the two antennas. Fortunately, this seems to be a problem only in highly mobile environments, which makes the use of two-antenna crosscorrelation spectrum sensing an interesting solution towards more reliable energy detection.
Mark S. Oude Alink, A. R. Smeenge, André B. J. Kokkeler, Eric A. M. Klumperink, Gerard J. M. Smit, Bram Nauta
VTC Fall5
2011 A Correlating Receiver for OFDM at Low SNR
abstract
By extending OFDM symbols, acceptable BER performance can be achieved at low SNRs. Two alternative differential receiver architectures are presented, a receiver based on a FX correlator (Fourier transformation before correlation) and based on an XF correlator (correlation before Fourier transformation). To reduce the complexity and hence the power consumption of both the ADC and the first digital processing stage single- or two bit quantization is used. The receiver based on the XF correlator is more suited to exploit such coarse quantization. Two basic effects are visible if coarse quantization is used. First, the BER performance is reduced due to the introduction of quantization errors. Second, beyond certain SNR levels, the BER performance does not increase due to the correlation between quantization errors. Furthermore, oversampling increases BER performance considerably. For single bit quantization with oversampling, acceptable BERs (-3) can be achieved for a limited SNR range for symbol extension factors of 32 and 64. In case of two bit quantization without oversampling, the results are comparable with single bit quantization with two times oversampling. For two bit quantization in combination with two times oversampling, acceptable BERs are achieved for symbol extension factors 8, 16, 31 and 64.
André B. J. Kokkeler, Gerard J. M. Smit
VTC Spring2
2010 Run-time spatial resource management for real-time applications on heterogeneous MPSoCs
abstract
Design-time application mapping is limited to a predefined set of applications and a static platform. Resource management at run-time is required to handle future changes in the application set, and to provide some degree of fault tolerance, due to imperfect production processes and wear of materials. This paper concerns resource allocation at run-time, allowing multiple real-time applications to run simultaneously on a heterogeneous MPSoC. Low-complexity algorithms are required, in order to respond fast enough to unpredictable execution requests. We present a decomposition of this problem into four phases. The allocation of tasks to specific locations in the platform is the main contribution of this work. Experiments on a real platform show the feasibility of this approach, with execution times in tens of milliseconds for a single allocation attempt.
Timon D. ter Braak, Philip K. F. Hölzenspies, Jan Kuper, Johann L. Hurink, Gerard J. M. Smit
DATE5
2010 Adaptive Beamforming Using the Reconfigurable MONTIUM TP
abstract
Until a decade ago, the concept of phased array beam forming was mainly implemented with mechanical or analog solutions. Today, digital hardware has become powerful enough to perform the massive number of operations required for real-time digital beam forming. While more and more applications are using beam forming to improve the communication channel utilization both in space and frequency, many dedicated digital architectures are proposed for the processing. By using a reconfigurable architecture, the same hardware platform can be reused for different applications with different processing needs. In this paper, we present a reconfigurable Multi-processor System-on-Chip based solution for phased array processing that supports advanced tracking mechanisms to continuously receive signals with a mobile receiver. An adaptive beam former for DVB-S satellite reception is presented, that uses a Constant Modulus Algorithm to track satellites. The processing of a receiver with 64 antennas and 3 beams is mapped on a reconfigurable processor named Montium TP. The total implementation of such a receiver requires about 570 clock cycles on a single Montium TP, but can also be partitioned over multiple Montium TPs to support larger phased arrays.
Marcel D. van de Burgwal, Kenneth C. Rovers, Koen C. H. Blom, André B. J. Kokkeler, Gerard J. M. Smit
DSD5
2010 An Approximate Maximum Common Subgraph Algorithm for Large Digital Circuits
abstract
This paper presents an approximate Maximum Common Sub graph (MCS) algorithm, specifically for directed, cyclic graphs representing digital circuits. Because of the application domain, the graphs have nice properties: they are very sparse, have many different labels, and most vertices have only one predecessor. The algorithm iterates over all vertices once and uses heuristics to find the MCS. It is linear in computational complexity with respect to the size of the graph. Experiments show that very large common sub graphs were found in graphs of up to 200,000 vertices within a few minutes, when a quarter or less of the graphs differ. The variation in run-time and quality of the result is low.
Jochem H. Rutgers, Pascal T. Wolkotte, Philip K. F. Hölzenspies, Jan Kuper, Gerard J. M. Smit
DSD5
2010 DVB-S Signal Tracking Techniques for Mobile Phased Arrays
abstract
A system that uses adaptive beamforming techniques for mobile DVB-S reception is proposed in this paper. The purpose is to enable DVB-S reception in moving vehicles. Phased arrays are able to electronically track the desired signal during dynamic behaviour of the vehicle the array is mounted on. The proposed system uses blind beamforming to adapt the array steering vector to changing signal (conditions and) directions. Movement of the vehicle, the phased array is mounted on, leads to modulus and phase deviations at the beamformer output. An extended version of the CMA algorithm is used to adapt the steering vector weights to compensate for those deviations. For simulation of the proposed system a model of vehicle dynamics is used to generate realistic antenna data. Simulation of the proposed system based on this antenna data shows appropriate corrections for modulus and phase deviations.
Koen C. H. Blom, Marcel D. van de Burgwal, Kenneth C. Rovers, André B. J. Kokkeler, Gerard J. M. Smit
VTC Fall5
2010 Buffer capacity computation for throughput-constrained modal task graphs
abstract
Increasingly, stream-processing applications include complex control structures to better adapt to changing conditions in their environment. This adaptivity often results in task execution rates that are dependent on the processed stream. Current approaches to compute buffer capacities that are sufficient to satisfy a throughput constraint have limited applicability in case of data-dependent task execution rates. In this article, we present a dataflow model that allows tasks to have loops with an unbounded number of iterations. For instances of this dataflow model, we present efficient checks on their validity. Furthermore, we present an efficient algorithm to compute buffer capacities that are sufficient to satisfy a throughput constraint. This allows to guarantee satisfaction of a throughput constraint over different modes of a stream processing application, such as the synchronization and synchronized modes of a digital radio receiver.
Maarten Wiggers, Marco Bekooij, Gerard J. M. Smit
ACM Trans. Embed. Comput. Syst.3
2009 Monotonicity and run-time scheduling
abstract
Modern embedded multi-processors can execute several stream-processing applications concurrently. Typically, these applications are partitioned into tasks that communicate over buffers together forming a task graph. The fact that these applications are started and stopped by the user combined with the knowledge that not all applications are necessarily completely characterised makes it attractive to use run-time scheduling. We define and characterise a class of budget schedulers that by construction bound the interference from other applications. Furthermore, we will show that the worst-case effects of these schedulers can be included in dataflow process networks. The execution of the resulting dataflow process network is shown to result in tight and conservative bounds on the end-to-end temporal behaviour of the execution of the task graph on a cycle-true simulator. Given that the inter-task synchronisation of the application allows for a dataflow model that is functionally deterministic, this enables exploration of various buffer capacities and scheduler settings at a high level of abstraction.
Maarten Wiggers, Marco Bekooij, Gerard J. M. Smit
EMSOFT3
2009 An Energy and Performance Exploration of Network-on-Chip Architectures
abstract
In this paper, we explore the designs of a circuit-switched router, a wormhole router, a quality-of-service (QoS) supporting virtual channel router and a speculative virtual channel router and accurately evaluate the energy-performance tradeoffs they offer. Power results from the designs placed and routed in a 90-nm CMOS process show that all the architectures dissipate significant idle state power. The additional energy required to route a packet through the router is then shown to be dominated by the data path. This leads to the key result that, if this trend continues, the use of more elaborate control can be justified and will not be immediately limited by the energy budget. A performance analysis also shows that dynamic resource allocation leads to the lowest network latencies, while static allocation may be used to meet QoS goals. Combining the power and performance figures then allows an energy-latency product to be calculated to judge the efficiency of each of the networks. The speculative virtual channel router was shown to have a very similar efficiency to the wormhole router, while providing a better performance, supporting its use for general purpose designs. Finally, area metrics are also presented to allow a comparison of implementation costs.
Pascal T. Wolkotte, Robert Mullins 0001, Simon W. Moore, Gerard J. M. Smit
IEEE Trans. Very Large Scale Integr. Syst.5
2008 Run-time Spatial Mapping of Streaming Applications to a Heterogeneous Multi-Processor System-on-Chip (MPSOC)
abstract
In this paper, we present an algorithm for run-time allocation of hardware resources to software applications. We define the sub-problem of run-time spatial mapping and demonstrate our concept for streaming applications on heterogeneous MPSoCs. The underlying algorithm and the methods used therein are implemented and their use is demonstrated with an illustrative example.
Philip K. F. Hölzenspies, Johann L. Hurink, Jan Kuper, Gerard J. M. Smit
DATE4
2008 Computation of Buffer Capacities for Throughput Constrained and Data Dependent Inter-Task Communication
abstract
Streaming applications are often implemented as task graphs. Currently, techniques exist to derive buffer capacities that guarantee satisfaction of a throughput constraint for task graphs in which the inter-task communication is data-independent, i.e. the amount of data produced and consumed is independent of the data values in the processed stream. This paper presents a technique to compute buffer capacities that satisfy a throughput constraint for task graphs with data dependent inter-task communication, given that the task graph is a chain. We demonstrate the applicability of the approach by computing buffer capacities for an MP 3 playback application, of which the MP 3 decoder has a variable consumption rate. We are not aware of alternative approaches to compute buffer capacities that guarantee satisfaction of the throughput constraint for this application.
Maarten Wiggers, Marco Bekooij, Gerard J. M. Smit
DATE3
2008 IRIS: A Firmware Design Methodology for SIMD Architectures
abstract
Developing code for SIMD type hardware architectures is a tedious job. This is caused by the absence of both a coherent methodological framework and a hardware independent tooling. Moreover, the inherently difficult nature of programming dedicated massively parallel embedded processors, complicates the matter. This paper describes a single framework, called IRIS, to generate code for SIMD architectures. This framework is illustrated with a concrete case "Stochastic Image Quantisation". IRIS is based on an incremental construction of executable representations, which converge to the final target implementation in a semi-automated way.
Jan W. M. Jacobs, Leroy van Engelen, Jan Kuper, Gerard J. M. Smit
DSD4
2008 An oversampled filter bank multicarrier system for Cognitive Radio
abstract
Due to small sideband power leakage, filter bank multicarrier techniques are considered as interesting alternatives to traditional OFDMs for spectrum pooling Cognitive Radio. In this paper, we propose an oversampled filter bank multicarrier system for Cognitive Radio. The increased spacing between adjacent subcarriers in the oversampled filter bank multicarrier system largely reduce the intercarrier interference, the key limitation of the OFDM based Cognitive Radio. The proposed multicarrier system is compared with OFDM for BER performance and sideband power rejection. Design tradeoffs of the major parameters of the oversampled filter bank will be discussed. We also suggest a fast implementation of the proposed filter bank modulation based on generalized DFT filter bank model, followed by a computational complexity analysis.
André B. J. Kokkeler, Gerard J. M. Smit
PIMRC3
2008 Buffer Capacity Computation for Throughput Constrained Streaming Applications with Data-Dependent Inter-Task Communication
abstract
Streaming applications are often implemented as task graphs, in which data is communicated from task to task over buffers. Currently, techniques exist to compute buffer capacities that guarantee satisfaction of the throughput constraint if the amount of data produced and consumed by the tasks is known at design-time. However, applications such as audio and video decoders have tasks that produce and consume an amount of data that depends on the decoded stream. This paper introduces a dataflow model that allows for data-dependent communication, together with an algorithm that computes buffer capacities that guarantee satisfaction of a throughput constraint. The applicability of this algorithm is demonstrated by computing buffer capacities for an H.263 video decoder.
Maarten Wiggers, Marco Bekooij, Gerard J. M. Smit
IEEE Real-Time and Embedded Technology and Applications Symposium3
2008 Communication between nested loop programs via circular buffers in an embedded multiprocessor system
abstract
Multimedia applications, executed by embedded multiprocessor systems, can in some cases be represented as task graphs, with the tasks containing nested loop programs. The nested loop programs communicate via arrays and can be executed on different processors. Typically an array can be communicated via a circular buffer with a capacity smaller than the array. For such buffers, the communicating nested loop programs have to synchronize and a sufficient buffer capacity needs to be computed. In a circular buffer we use a write and a read window to support rereading, out-of-order reading or writing, and skipping of locations. A cyclo static dataflow model is derived from the application and used to compute buffer capacities that guarantee deadlock free execution. Our case-study applies circular buffers in a Digital Audio Broadcasting channel decoder application, where the frequency deinterleaver reads according to a non-affine pseudo-random function. For this application, buffer capacities are calculated that guarantee deadlock free execution.
Tjerk Bijlsma, Marco Bekooij, Pierre G. Jansen, Gerard J. M. Smit
SCOPES4
2008 Cognitive Radio Design on an MPSoC Reconfigurable Platform
abstract
Cognitive Radio has been proposed as a promising technology for solving today’s spectrum scarcity problem by means of dynamic spectrum access. The multiprocessor system-on-chip (MPSoC) reconfigurable platform is proposed as an enabling technology for cognitive radio. In this paper, we propose a design methodology based on task transaction level interface for the design of cognitive radio baseband on an MPSoC reconfigurable platform. The reconfiguration of a novel, low-complexity fast Fourier transform for orthogonal frequency-division multiplexing based Cognitive Radio is used as a design case to show the effectiveness of the methodology for modelling the dynamic behavior of Cognitive Radio and facilitating the platform implementation.
André B. J. Kokkeler, Gerard J. M. Smit
Mob. Networks Appl.3
2008 Towards Software Defined Radios Using Coarse-Grained Reconfigurable Hardware
abstract
Mobile wireless terminals tend to become multimode wireless communication devices. Furthermore, these devices become adaptive. Heterogeneous reconfigurable hardware provides the flexibility, performance, and efficiency to enable the implementation of these devices. The implementation of a wideband code division multiple access and an orthogonal frequency division multiplexing receiver using the same coarse-grained reconfigurable MONTIUM tile processor is discussed. Besides the baseband processing part of the receiver, the same reconfigurable processor has also been used to implement Viterbi and Turbo channel decoders.
Gerard K. Rauwerda, Paul M. Heysters, Gerard J. M. Smit
IEEE Trans. Very Large Scale Integr. Syst.3
2007 Efficient Computation of Buffer Capacities for Cyclo-Static Dataflow Graphs
abstract
A key step in the design of cyclo-static real-time systems is the determination of buffer capacities. In our multi-processor system, we apply back-pressure, which means that tasks wait for space in output buffers. Consequently buffer capacities affect the throughput. This requires the derivation of buffer capacities that both result in a satisfaction of the throughput constraint, and also satisfy the constraints on the maximum buffer capacities. Existing exact solutions suffer from the computational complexity that is associated with the required conversion from a cyclo-static dataflow graph to a single-rate dataflow graph. In this paper we present an algorithm, with polynomial computational complexity, that does not require this conversion and that obtains close to minimal buffer capacities. The algorithm is applied to an MP3 play-back application that is mapped on our multi-processor system. For this application, we see that a cyclo-static dataflow model can reduce the buffer capacities by 50% compared to a multi-rate dataflow model.
Maarten Wiggers, Marco Bekooij, Gerard J. M. Smit
DAC3
2007 Cyclostationary feature detection on a tiled-SoC
abstract
In this paper, a two-step methodology is introduced to analyse the mapping of cyclostationary feature detection (CFD) onto a multi-core processing platform. In the first step, the tasks to be executed by each core are determined in a structured way using techniques known from the design of array processors. In the second step, the implementation of tasks on a processing core is analysed. Using this methodology, it is shown that calculating a 127 times 127 discrete spectral correlation function requires approximately 140 mus on a tiled system on chip (SoC) with 4 Montium cores
André B. J. Kokkeler, Gerard J. M. Smit, Thijs Krol, Jan Kuper
DATE2
2007 Implementation of a 2-D 8x8 IDCT on the Reconfigurable Montium Core
abstract
This paper describes the mapping of a two-dimensional inverse discrete cosine transform (2-D IDCT) onto a word-level reconfigurable Montium® processor. This shows that the IDCT is mapped onto the Montium tile processor (TP) with reasonable effort and presents performance numbers in terms of energy consumption, speed and silicon costs. The Montium results are compared with the IDCT implementation on three other architectures: TI DSP, ASIC and ARM.
Lodewijk T. Smit, Gerard K. Rauwerda, Albert Molderink, Pascal T. Wolkotte, Gerard J. M. Smit
FPL5
2007 An Efficient FFT For OFDM Based Cognitive Radio On A Reconfigurable Architecture
abstract
Cognitive radio is a promising technology to utilize non-used parts of the spectrum that actually are assigned to licensed services. An adaptive OFDM based cognitive radio system has the capacity to nullify individual carriers to avoid interference to the licensed user. Therefore, there could be a considerably large number of zero-valued inputs/outputs for the IFFT/FFT in the OFDM transceiver. Due to the wasted operations on zero values, the standard FFT is no longer efficient. Based on this observation, we propose to use a computationally efficient IFFT/FFT as an option for OFDM based cognitive radio. Mapping this algorithm onto a reconfigurable architecture is discussed.
André B. J. Kokkeler, Gerard J. M. Smit
ICC3
2007 Using an FPGA for Fast Bit Accurate SoC Simulation
abstract
In this paper we describe a sequential simulation method to simulate large parallel homo- and heterogeneous systems on a single FPGA. The method is applicable for parallel systems were lengthy cycle and bit accurate simulations are required. It is particularly designed for systems that do not fit completely on the simulation platform (i.e. FPGA). As a case study, we use a network-on-chip (NoC) that is simulated in SystemC and on the described FPGA simulator. This enables us to observe the NoC behavior under a large variety of traffic patterns. Compared with the SystemC simulation we achieved a factor 80-300 of speed improvement, without compromising the cycle and bit level accuracy.
Pascal T. Wolkotte, Philip K. F. Hölzenspies, Gerard J. M. Smit
IPDPS3
2007 Fast, Accurate and Detailed NoC Simulations
abstract
Network-on-chip (NoC) architectures have a wide variety of parameters that can be adapted to the designer's requirements. Fast exploration of this parameter space is only possible at a high-level and several methods have been proposed. Cycle and bit accurate simulation is necessary when the actual router's RTL description needs to be evaluated and verified. However, extensive simulation of the NoC architecture with cycle and bit accuracy is prohibitively time consuming. In this paper we describe a simulation method to simulate large parallel homogeneous and heterogeneous network-on-chips on a single FPGA. The method is especially suitable for parallel systems where lengthy cycle and bit accurate simulations are required. As a case study, we use a NoC that was modelled and simulated in SystemC. We simulate the same NoC on the described FPGA simulator. This enables us to observe the NoC behavior under a large variety of traffic patterns. Compared with the SystemC simulation we achieved a speed-up of 80-300, without compromising the cycle and bit level accuracy
Pascal T. Wolkotte, Philip K. F. Hölzenspies, Gerard J. M. Smit
NOCS3
2007 Efficient Computation of Buffer Capacities for Cyclo-Static Real-Time Systems with Back-Pressure
abstract
This paper describes a conservative approximation algorithm that derives close to minimal buffer capacities for an application described as a cyclo-static dataflow graph. The resulting buffer capacities satisfy constraints on the maximum buffer capacities and end-to-end throughput and latency constraints. Furthermore we show that the effects of run-time arbitration can be included in the response times of dataflow actors. We show that modelling an MP3 playback application as a cyclo-static dataflow graph instead of a multi-rate dataflow graph results in buffer capacities that are reduced up to 39%. Furthermore, the algorithm is applied to a real-life car-radio application, in which two independent streams are processed
Maarten Wiggers, Marco Bekooij, Pierre G. Jansen, Gerard J. M. Smit
IEEE Real-Time and Embedded Technology and Applications Symposium4
2007 Modelling run-time arbitration by latency-rate servers in dataflow graphs
abstract
In order to obtain a cost-efficient solution, tasks share resources in a Multi-Processor System-on-Chip. In our architecture, shared resources are run-time scheduled. We show how the effects of Latency-Rate servers, which is a class of run-time schedulers, can be included in a dataflow model. The resulting dataflow model, which can have an arbitrary topology, enables us to provide guarantees on the temporal behaviour of the implementation.
Maarten Wiggers, Marco Bekooij, Gerard J. M. Smit
SCOPES3
2006 An optimal architecture for a DDC
abstract
Digital down conversion (DDC) is an algorithm, used to lower the amount of samples per second by selecting a limited frequency band out of a stream of samples. A possible DDC algorithm consists of two simple cascading integrating comb (CIC) filters and a finite input response (FIR) filter preceded by a modulator that is controlled with a numeric controlled oscillator (NCO). Implementations of the algorithm have been made for five architectures, two application specific integrated circuits (ASIC), a general purpose processor (GPP), a field programmable gate array (FPGA), and the Montium tile processor (TP). All architectures are functionally capable of performing the algorithm. The differences between the architectures are their performance, flexibility and energy consumption. In this paper, we compared the energy consumption of the architectures when performing the DDC algorithm. The ASIC is the best solution if digital down conversion is constantly required. When digital down conversion is needed only parts of the time, the Altera Cyclone II is the best solution due to its smaller technology size. In the spare time, the reconfigurable architectures can be reconfigured for other tasks of today's multimedia devices
Tjerk Bijlsma, Pascal T. Wolkotte, Gerard J. M. Smit
IPDPS3
2006 A pattern selection algorithm for multi-pattern scheduling
abstract
The multi-pattern scheduling algorithm is designed to schedule a graph onto a coarse-grained reconfigurable architecture, the result of which depends highly on the used patterns. This paper presents a method to select a near-optimal set of patterns. By using these patterns, the multi-pattern scheduling will result in a better schedule in the sense that the schedule will have fewer clock cycles.
Yuanqing Guo, Cornelis Hoede, Gerard J. M. Smit
IPDPS3
2005 Throughput of Streaming Applications Running on a Multiprocessor Architecture
abstract
In this paper we study the timing behaviour of streaming applications running on a multiprocessor architecture. Dependencies are derived between the application throughput and the timing characteristics of the processors and communication. Four different processor organizations that strongly influenced the results are considered and compared.
Nikolay Kavaldjiev, Gerard J. M. Smit, Pierre G. Jansen
DSD2
2005 Energy-Efficient NoC for Best-Effort Communication
abstract
A Network-on-Chip (NoC) is an energy-efficient on-chip communication architecture for Multi-Processor System-on-Chip (MPSoC) architectures. In an earlier paper we proposed a energy-efficient reconfigurable circuit-switched NoC to reduce the energy consumption compared to a packet-switched NoC. In this paper we investigate a chordal slotted ring and a bus architecture that can be used to handle the best-effort traffic in the system and configure the circuit-switched network. Both architectures are compared on their latency behavior and power consumption. At the same clock frequency, the chordal ring has the major benefit of a lower latency and higher throughput. But the bus has a lower overall power consumption at the same frequency. However, if we tune the frequency of the network to meet the throughput requirements of control network, we see that the ring consumes less energy per transported hit.
Pascal T. Wolkotte, Gerard J. M. Smit, Jens E. Becker
FPL2
2004 An Energy-Efficient Network-on-Chip for a Heterogeneous Tiled Reconfigurable Systems-on-Chip
abstract
This paper proposes a network-on-chip architecture that offers high flexibility and performance. It is used in a system-on-chip platform for future multimedia mobile devices. The network is packet switching wormhole network with virtual-channel flow control and source routing. The initial implementation results for a network router show its feasibility and size comparable with other available solutions.
Nikolay Kavaldjiev, Gerard J. M. Smit
DSD2
2004 Implementation of a flexible RAKE receiver in heterogeneous reconfigurable hardware
abstract
Mobile wireless terminals tend to become multi-mode wireless communication devices. Furthermore, these devices need to be adaptive as they can adapt to changing environmental conditions as well as changing user demands. We foresee a heterogeneous reconfigurable system-on-chip (SoC) as the key-technology for future wireless communication systems. The SoC contains processing elements of different granularities. The implementation of a W-CDMA receiver with true dynamic reconfiguration capabilities in heterogeneous reconfigurable hardware is described.
Gerard K. Rauwerda, Gerard J. M. Smit
FPT2
2004 Run-time mapping of applications to a heterogeneous reconfigurable tiled system on chip architecture
abstract
This work evaluates an algorithm that maps a number of communicating processes to a heterogeneous tiled system on chip (SoC) architecture at run-time. The mapping algorithm minimizes the total amount of energy consumption, while still providing an adequate quality of service (QoS). A realistic example is mapped using this algorithm.
Lodewijk T. Smit, Gerard J. M. Smit, Johann L. Hurink, Hajo Broersma, Daniël Paulusma, Pascal T. Wolkotte
FPT2
2004 Implementation of a HiperLAN/2 Receiver on the Reconfigurable Montium Architecture
abstract
Summary form only given. A heterogeneous system-on-chip (SoC) architecture for mobile hand-held devices is proposed to overcome the battery bottleneck in these devices. This SoC contains processing tiles of different granularities. The Montium coarse-grain reconfigurable tile processor is presented. Also, an introduction to HiperLAN/2 baseband processing is given. The implementation of a HiperLAN/2 receiver on the Montium reconfigurable architecture is explained in detail. The hardware of this implemented receiver has been simulated and the performance figures are given. The configuration overhead for the receiver is very small, which enables dynamic reconfiguration. The required computational performance can be obtained at very low clock frequencies. The Montium coarse-grain reconfigurable architecture enables an energy and area efficient implementation of a HiperLAN/2 receiver.
Paul M. Heysters, Gerard K. Rauwerda, Gerard J. M. Smit
IPDPS3
2004 The Computational Complexity of the Minimum Weight Processor Assignment Problem
Hajo Broersma, Daniël Paulusma, Gerard J. M. Smit, Frank Vlaardingerbroek, Gerhard J. Woeginger
WG3
2004 Mapping Wireless Communication Algorithms onto a Reconfigurable Architecture
Gerard K. Rauwerda, Paul M. Heysters, Gerard J. M. Smit
J. Supercomput.3
2003 Mapping Applications to an FPFA Tile
Michèl A. J. Rosien, Yuanqing Guo, Gerard J. M. Smit, Thijs Krol
DATE3
2003 A Communication Model Based on an n-Dimensional Torus Architecture Using Deadlock-Free Wormhole Routing
abstract
Routing on a two-dimensional torus architecture by means of the wormhole routing algorithm is introduced and extended to an n-dimensional torus model. To prevent blocking deadlocks caused by this algorithm, a multiple virtual channel solution is introduced. An implementation of virtual channels is introduced that allows channels with higher labels to pre-empt 'lower' channels. This algorithm is tested with a simplified model of a HiperLAN/2 receiver. The model proves to be capable of running this application on the Chameleon architecture according to http://chameleon.ctit.utwente.nl/.
Philip K. F. Hölzenspies, Erik Schepers, Wouter Bach, Mischa Jonker, Bart Sikkes, Gerard J. M. Smit, Paul J. M. Havinga
DSD6
2003 A graph covering algorithm for a coarse grain reconfigurable system
abstract
The availability of high-level design entry tooling is crucial for the viability of any reconfigurable SoC architecture. This paper presents a graph covering algorithm. The graph covering is done in two steps: template generation and template selection. The objective of template generation step is to extract functional equivalent structures, i.e. templates, from a control data flow graph. By inspecting the graph, the algorithm generates all the possible templates and the corresponding matches. Using unique serial numbers and circle numbers, the algorithm can find all distinct templates with multiple outputs. The template selection algorithm shows how this information can be used in compilers for reconfigurable systems. The objective of the template selection algorithm is to find an efficient cover for an application graph with a minimal number of distinct templates and minimal number of matches.
Yuanqing Guo, Gerard J. M. Smit, Hajo Broersma, Paul M. Heysters
LCTES2
2003 A Flexible and Energy-Efficient Coarse-Grained Reconfigurable Architecture for Mobile Systems
Paul M. Heysters, Gerard J. M. Smit, Egbert Molenkamp
J. Supercomput.2
2002 Dynamic Reconfiguration in Mobile Systems
Gerard J. M. Smit, Paul J. M. Havinga, Lodewijk T. Smit, Paul M. Heysters, Michèl A. J. Rosien
FPL1
2002 Enhancing energy efficient TCP by partial reliability
abstract
We present a study on the effects on a mobile system's energy efficiency of enhancing, with partial reliability, our energy efficient TCP variant (E/sup 2/TCP) (see Donckers, L. et al., Proc. 2nd Asian Int. Mobile Computing Conf. - AMOC2002, p.18-28, 2002). Partial reliability is beneficial for multimedia applications and is especially attractive in wireless communication environments. It allows applications to trade a controlled amount of data loss for a higher throughput, lower delays, and lower energy consumption. In the design of E/sup 2/TCP, we provide solutions to four problem areas in TCP that prevent it from reaching high levels of energy efficiency. We have implemented a simulation model of the protocol, and the results show that E/sup 2/TCP has a significant higher energy efficiency than TCP.
Lewie Donckers, Paul J. M. Havinga, Gerard J. M. Smit, Lodewijk T. Smit
PIMRC3
2001 Energy Management for Dynamically Reconfigurable Heterogeneous Mobile Systems
abstract
Dynamically reconfigurable systems offer the potential for realising efficient systems as well as providing adaptability to changing system requirements. Such systems are suitable for future mobile multimedia systems that have limited battery resources, must handle diverse data types, and must operate in dynamic application and communication environments. We propose an approach in which reconfiguration is applied dynamically at various levels of a mobile system, whereas traditionally, reconfigurable systems mainly focus at the gate level only. The research performed in the CHAMELEON project 1aims at designing such a heterogeneous reconfigurable mobile system. The two main motivations for the system are (1) to have an energy-efficient system, while (2) achieving an adequate Quality of Service for applications.
Paul J. M. Havinga, Lodewijk T. Smit, Gerard J. M. Smit, Martinus Bos, Paul M. Heysters
IPDPS3
2001 Energy-efficient wireless networking for multimedia applications
abstract
Abstract In this paper we identify the most prominent problems of wireless multimedia networking and present several state‐of‐the‐art solutions with a focus on energy efficiency. Three key problems in networked wireless multimedia systems are: (1) the need to maintain a minimum quality of service over time‐varying channels; (2) to operate with limited energy resources; and (3) to operate in a heterogeneous environment. We identify two main principles to solve these problems. The first principle is that energy efficiency should involve all layers of the system. Second, Quality of Service is an essential mechanism for mobile multimedia systems not only to give users an adequate level of service, but also as a tool to achieve an energy‐efficient system. Owing to the dynamic wireless environment, adaptability of the system will be a key issue in achieving this. Copyright © 2001 John Wiley & Sons, Ltd.
Paul J. M. Havinga, Gerard J. M. Smit
Wirel. Commun. Mob. Comput.2
2000 Energy-Efficient Adaptive Wireless Network Design
abstract
Energy efficiency is an important issue for mobile computers since they must rely on their batteries. We present an energy-efficient highly adaptive architecture of a network interface and novel data link layer protocol for wireless networks that provides quality of service (QoS) support for diverse traffic types. Due to the dynamic nature of wireless networks, adaptations are necessary to achieve energy efficiency and an acceptable quality of service. The paper provides a review of ideas and techniques relevant to the design of an energy efficient adaptive wireless network.
Paul J. M. Havinga, Gerard J. M. Smit, Martinus Bos
ISCC2
2000 Design techniques for low-power systems
Paul J. M. Havinga, Gerard J. M. Smit
J. Syst. Archit.2
2000 Energy-efficient wireless ATM design
Paul J. M. Havinga, Gerard J. M. Smit, Martinus Bos
Mob. Networks Appl.2
1999 Octopus: Embracing the Energy Efficiency of Handheld Multimedia Computers
abstract
In the MOBY DICK project we develop and define the architecture of a new generation of mobile hand-held computers called Mobile Digital Companions. The Companions must meet several major requirements: high performance, energy efficient, a notion of Quality of Service (QoS), small size, and low design complexity. To address these requirements we need to revise the architecture of the hardware, the operating system, and applications. -The approach is based on dedicated functionality and the extensive use of energy reduction techniques at all levels of system design. The Mobile Digital Companion has an unconventional architecture that saves energy by using system decomposition at different levels of the architecture and exploits locality of reference with dedicated, optimised modules. A reconfigurable internal communication network switch called Octopus exploits locality of reference and eliminates wasteful data copies. The switch is implemented as a simplified ATM switch and provides Quality of Service guarantees and enough bandwidth for multimedia applications. We have built a testbed of the architecture, of which we will present performance and energy consumption characteristics.
Paul J. M. Havinga, Gerard J. M. Smit
MobiCom2
1995 Virtual Lines, a Deadlock-Free and Real-Time Routing Mechanism for ATM Networks
Gerard J. M. Smit, Paul J. M. Havinga, Walter H. Tibboel
Inf. Sci.1
1992 The Architecture of Rattlesnake: a Real-Time Multimedia Network
Gerard J. M. Smit, Paul J. M. Havinga
NOSSDAV1
1992 On the design of a dynamic reconfigurable network switch
Gerard J. M. Smit, Paul J. M. Havinga, Pierre G. Jansen
Microprocess. Microprogramming1
1991 On hardware for generating routes in Kautz digraphs
Gerard J. M. Smit, Paul J. M. Havinga, Pierre G. Jansen, Fokke de Boer, Egbert Molenkamp
Microprocessing and Microprogramming1
1989 Hardware support for the tumult real-time scheduler
H. C. van der Bij, Gerard J. M. Smit, Paul J. M. Havinga
Microprocessing and Microprogramming2
1988 The communication processor of TUMULT-64
Gerard J. M. Smit, Pierre G. Jansen
Microprocess. Microprogramming1