EDBT 2026 Demo / reviewers in the wild / expert
Ronald P. Luijten
dblp:03/2439
· DBLP profile ↗
13ranked-venue papers
3as first author
0since 2021 · last 2019
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 6 · 2 first-authorSystems, architecture and hardware · 4 · 1 first-authorSoftware engineering, systems software and programming languages · 1Graphics, computer vision, multimedia, augmented reality and games · 1Applied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer networks
2 papers |
Optical networks · 54% Routing and switching · 46% | |
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Interconnection networks and networks-on-chip · 81% High-performance computing · 19% |
Topics — the 7 heaviest of 7, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Routing and switching › switch architecture
switch fabric |
0.1 | 1 | 2007 | Design issues in next-generation merchant switch fabrics · IEEE/ACM Trans. Netw. 2007 |
Optical networks
optical switching |
0.1 | 1 | 2005 | Viable opto-electronic HPC interconnect fabrics · SC 2005 |
Optical networks › optical switching
optoelectronic hybrid switching |
0.1 | 1 | 2005 | Viable opto-electronic HPC interconnect fabrics · SC 2005 |
Interconnection networks and networks-on-chip › high-speed networks
supercomputer interconnect |
0.1 | 1 | 2005 | Viable opto-electronic HPC interconnect fabrics · SC 2005 |
Routing and switching
switch architecture |
0.0 | 1 | 2007 | Design issues in next-generation merchant switch fabrics · IEEE/ACM Trans. Netw. 2007 |
Interconnection networks and networks-on-chip
interconnect architecture |
0.0 | 1 | 2005 | Viable opto-electronic HPC interconnect fabrics · SC 2005 |
High-performance computing
supercomputing |
0.0 | 1 | 2005 | Viable opto-electronic HPC interconnect fabrics · SC 2005 |
Methods — techniques the papers use, named apart from their topics
semiconductor optical amplifier · 0.1multistage topology · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2019 | Coherently Attached Programmable Near-Memory Acceleration Platform and its application to Stencil ProcessingabstractApplication and technology trends are increasingly forcing computer systems to be designed for specific workloads and application domains. Although memory is one of the key components impacting the performance and power consumption of state-of-art computer systems, its operation typically cannot be adapted to workload characteristics beyond some limited controller configuration options. In this paper, we present a novel near-memory acceleration platform based on an Access Processor that enables the main memory system operation to be programmed and adapted dynamically to the accelerated workload. The platform targets both ASIC and FPGA implementations integrated within IBM POWER systems. We show how this platform can be applied to accelerate stencil processing. Jan van Lunteren, Ronald P. Luijten, Dionysios Diamantopoulos, Florian Auernhammer, Christoph Hagleitner, Lorenzo Chelini, Stefano Corda, Gagandeep Singh 0002 |
DATE | 2 |
| 2014 | Holistic power analysis of implementation alternatives for a very large scale synthesis array with phased array stationsabstractThe Square Kilometre Array (SKA) will be the largest radio telescope in the world, generating data at Pb/s rates. Real-time processing will require 1018compute operations per second and system operating costs will be dominated by energy consumption. In this paper we explore design options for the aperture array of the first SKA construction phase and provide lower bounds on their power consumption. We analyze the system's components from the antenna front-end to the central signal processor and identify the main power consumers. We compare ASIC-based and FPGA-based data processing pipelines and show that ASICs can lead to 1.6 to 4 times more power efficiency. Andreea Anghel, Rik Jongerius, Gero Dittmann, Jonas R. M. Weiss, Ronald P. Luijten |
ICASSP | 5 |
| 2012 | The Network Adapter: The Missing Link between MPI Applications and Network PerformanceabstractNetwork design aspects that influence cost and performance can be classified according to their distance from the applications, into issues concerning topology, switch technology, link technology, network adapter, and communication library. The network adapter has a privileged position to take decisions with more global information than any other component in the network. It receives feedback from the switches and requests from the communication libraries and applications. Also, compared to a network switch, an adapter has access to significantly more memory (host memory and on-chip memory) and memory bandwidth (which typically exceeds network bandwidth). The potential of the adapter to improve global network performance has not yet been fully exploited. In this work we show a series of noticeable performance improvements (of at least 10% to 15%) for medium-sized message exchanges in typical HPC communication patterns by optimizing message segmentation and packet injection policies, that can be implemented in an adapter's firmware inexpensively. We also show that implementing equivalent solutions in the switch (as opposed to the adapter) leads to only marginal performance improvements as the ones obtained by controlling the segmentation and injection policy at the adapter, while involving significantly more cost. In addition, enhancing the adapter will lead to less hardware complexity in the switches, thus reducing cost and energy consumption. Cyriel Minkenberg, Ronald P. Luijten, Ramón Beivide, Patrick Geoffray, Jesús Labarta, Mateo Valero, Stephen W. Poole |
SBAC-PAD | 3 |
| 2011 | On the optimum switch radix in fat tree networksabstractBased on a realistic, yet simple cost model, we compute the switch radix that minimizes the cost of a fat tree network to support a given number of end nodes. The cost model comprises two parameters indicating the relative cost of a crosspoint vs. a link, and the crosspoint-independent base cost of a switch. These parameters can be adapted to represent a given technology used to implement links and switches. Based on these inputs, the resulting model allows a quick evaluation of the switch radix that minimizes the overall cost of the network. We demonstrate that the optimum radix depends most strongly on the relative cost of a link, and turns out to be largely independent of the network size. Using a first-order cost bounds analysis based on current CMOS and link technology, our model indicates that the optimum switch radix for large fat trees is driven almost entirely by link cost and as a result lies in the range of hundreds of ports, rather than the tens of ports being offered today by most commercial switch products today. Cyriel Minkenberg, Ronald P. Luijten |
HPSR | 2 |
| 2011 | Pinned to the walls: impact of packaging and application properties on the memory and power walls
Phillip Stanley-Marbell, Victoria Caparrós Cabezas, Ronald P. Luijten |
ISLPED | 3 |
| 2009 | Oblivious routing schemes in extended generalized Fat Tree networksabstractA family of oblivious routing schemes for fat trees and their slimmed versions is presented in this work. First, two popular oblivious routing algorithms, which we refer to as S-mod-k and D-mod-k, are analyzed in detail. S-mod-k is the default routing algorithm given as an example in the first works formally describing fat tree networks. D-mod-k has been independently proposed and investigated by several authors, who conclude in their evaluations that it achieves better performance than a random or adaptive routing approach. First, we identify the reasons why these algorithms perform well. Using this insight we extend these algorithms, originally intended for full bisection networks, to slimmed networks. Based on the lessons learned we propose a new generalized family of algorithms that provides a better oblivious solution than the existing ones for this class of networks. Moreover, this family extends the previous work from k-ary n-trees to the more general class of extended generalized fat trees. Cyriel Minkenberg, Ramón Beivide, Ronald P. Luijten, Jesús Labarta, Mateo Valero |
CLUSTER | 4 |
| 2009 | Optimization of link bandwidth for parallel communication performanceabstractThe efficiency of computer network has been regarded as a bottleneck in parallel computing paradigm. It is important to have efficient methodology to obtain network performance measures, especially for a large scale system, i.e. exa-scale system. Communication performance is often investigated by the static complexity analysis based on a given network topology or a detailed network simulation, which is often time consuming. To provide a dynamic and scalable communication performance measure, we first propose an aggregate multi-stage queueing network model to capture the application's communication load and derive the closed-form system performance, i.e. throughput and delay. Trace simulation results obtained from a sophisticated simulator, Venus, show that the proposed model is accurate, yet simple. Secondly, we develop a link bandwidth optimization framework, which optimally allocates/distributes link bandwidth across the network to maximize the system communication throughput. Specifically, we apply the derived optimal bandwidth allocation on dimensioning link bandwidth of an exploratory direct network and slimming fat-tree network. Our results show that the proposed methodology is cost-effective in providing system performance and design explorations for the existing and the next-generation network system. Lydia Y. Chen, Wolfgang E. Denzel, Ronald P. Luijten |
IPCCC | 3 |
| 2007 | Design issues in next-generation merchant switch fabrics
François Abel, Cyriel Minkenberg, Ilias Iliadis, Antonius P. J. Engbersen, Mitchell Gusat, Ferdinand Gramsamer, Ronald P. Luijten |
IEEE/ACM Trans. Netw. | 7 |
| 2005 | Viable opto-electronic HPC interconnect fabricsabstractWe address the problem of how to exploit optics for ultrascale High Performance Computing interconnect fabrics. We show that for high port counts these fabrics require multistage topologies regardless of whether electronic or optical switch components are used. Also, per stage electronic buffers remain indispensable for maintaining throughput, lossless-ness and packet sequence. Although the notion of true all-optical packet switching is not yet viable, we show that appropriate use of optical switching technology offers power and scaling advantages that can be leveraged economically, and propose a hybrid opto-electronic HPC interconnect fabric architecture that combines the strength of electronics in processing and storing information with the strength of optics in switching and transporting high bandwidths. Using Semiconductor Optical Amplifier technology, we are building a prototype demonstrator switch that we believe solves all the technical challenges. Having reached this threshold now enables commercialization of this technology, which we are currently pursuing. Ronald P. Luijten, Cyriel Minkenberg, B. Roe Hemenway, Michael Sauer, Richard Grzybowski |
SC | 1 |
| 2003 | Reducing memory size in buffered crossbars with large internal flow control latencyabstractA buffered crossbar supporting P priorities and a flow control latency of RT packets between the input adapter and the crossbar requires a memory of order O(N/sup 2/*P*RT) packets in the crossbar to support any traffic pattern without blocking. We propose a new priority elevation mechanism that reduces the memory requirements to O(N/sup 2/*RT) for large values of RT. Our analysis shows that our mechanism has no drawback on the usual performance metrics and that it only introduces a small priority unfairness and worst-case blocking of less than RT packet times. We further show that an optimized system using a crosspoint memory of size 2RT has an unfairness of less than 0.02% affected packets at 95% loading with an average burst size of 30 packets. Ronald P. Luijten, Cyriel Minkenberg, Mitchell Gusat |
GLOBECOM | 1 |
| 2002 | Optimizing flow control for buffered switchesabstractWe address a problem often neglected in the presentation of credit flow control (FC) schemes for buffered switches, namely the issue of FC bandwidth and FC optimization, i.e. how many and which credits to return per packet cycle. Under the assumption of bursty traffic with uniform destinations, we show via simulations that, independent of switch size and without loss in performance, the number of credits to be returned can be reduced to one. We further introduce the notion of credit contention and credit scheduling. We analyze four credit scheduling strategies under varying system and buffer size. Our results demonstrate that, with a proper credit scheduler in place, contention resolution is resolved much faster than with conventional schemes. Our findings suggest that scheduling of credits is a means for the switch to determine its future arrivals during contention phases. Ferdinand Gramsamer, Mitchell Gusat, Ronald P. Luijten |
ICCCN | 3 |
| 2002 | Stability of CIOQ switches with finite buffers and non-negligible round-trip timeabstractWe propose a systematic method to determine the lower bound for internal buffering of practical CIOQ (combined input-output queued) switching systems. We introduce a deterministic traffic scenario that stresses the global stability of finite output queues. We demonstrate its usefulness by dimensioning the buffer capacity of the CIOQ under such traffic patterns. Compliance with this property maximizes the performance achievable with finite buffers. Mitchell Gusat, François Abel, Ferdinand Gramsamer, Ronald P. Luijten, Cyriel Minkenberg, Mark Verhappen |
ICCCN | 4 |
| 2001 | Optimized architecture and design of an output-queued CMOS switch chipabstractTraditional improvements in packet switch architecture are aimed at increasing switch performance in terms of utilization, fairness and QoS. This paper focuses on improving the architecture to achieve implementation feasibility of terabit aggregate data rates while maintaining such performance. Terabit class shared-memory switch chips are simple in concept but are a challenge to build due to the memory speed requirements and the complexity of wiring needed to connect these memories. Using a property of the combined shared memory and virtual output queuing switch architecture and a property of SRAMs, a new architecture is derived that enables construction of a terabit class switch fabric. Ronald P. Luijten, François Abel, Mitchell Gusat, Cyriel Minkenberg |
ICCCN | 1 |