Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Hans Eberle

dblp:46/4152 · DBLP profile ↗
← Back
28ranked-venue papers
12as first author
0since 2021 · last 2018
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 12 · 10 first-authorComputer networks · 7Security and privacy · 5 · 1 first-authorHuman-computer interaction and ubiquitous computing · 3Software engineering, systems software and programming languages · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
6 papers
Interconnection networks and networks-on-chip · 67% Distributed systems · 15% Integrated circuit design · 12%
Network and information security
7 papers
Cryptographic primitives and cryptanalysis · 100% Network security · 0%
Computer networks
3 papers
Internet of things and sensor networks · 85% Internet architecture and protocols · 15%

Topics — the 21 heaviest of 25, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Cryptographic primitives and cryptanalysis › public-key cryptography
elliptic curve cryptography
0.262005
Energy Analysis of Public-Key Cryptography for Wireless Sensor Networks · PerCom 2005
Speeding up Secure Web Transactions Using Elliptic Curve Cryptography · NDSS 2004
Comparing Elliptic Curve Cryptography and RSA on 8-bit CPUs · CHES 2004
Cryptographic primitives and cryptanalysis
public-key cryptography
0.132005
Energy Analysis of Public-Key Cryptography for Wireless Sensor Networks · PerCom 2005
Comparing Elliptic Curve Cryptography and RSA on 8-bit CPUs · CHES 2004
An End-to-End Systems Approach to Elliptic Curve Cryptography · CHES 2002
Interconnection networks and networks-on-chip › switch architecture
crossbar switch
0.122008
High-radix crossbar switches enabled by proximity communication · SC 2008
Switcherland: A QoS Communication Architecture for Workstation Clusters · ISCA 1998
Distributed systems › distributed coordination
message ordering
0.112018
Light-weight protocols for wire-speed ordering · SC 2018
Integrated circuit design
interconnect
0.112008
High-radix crossbar switches enabled by proximity communication · SC 2008
Interconnection networks and networks-on-chip
cluster interconnect
0.122002
Separated high-bandwidth and low-latency communication in the cluster interconnect Clint · SC 2002
Switcherland: A QoS Communication Architecture for Workstation Clusters · ISCA 1998
Cryptographic primitives and cryptanalysis › public-key cryptography
RSA
0.012004
Comparing Elliptic Curve Cryptography and RSA on 8-bit CPUs · CHES 2004
Cryptographic primitives and cryptanalysis › finite field arithmetic
modular reduction
0.012002
Generic implementations of elliptic curve cryptography using partial reduction · CCS 2002
Cryptographic primitives and cryptanalysis › public-key cryptography › elliptic curve cryptography
scalar multiplication
0.012002
Generic implementations of elliptic curve cryptography using partial reduction · CCS 2002
Interconnection networks and networks-on-chip
network scheduling
0.012002
Separated high-bandwidth and low-latency communication in the cluster interconnect Clint · SC 2002
Distributed systems › distributed system architecture
communication architecture
0.011998
Switcherland: A QoS Communication Architecture for Workstation Clusters · ISCA 1998
High-performance computing › cluster computing
network of workstations
0.011998
Switcherland: A QoS Communication Architecture for Workstation Clusters · ISCA 1998
Cloud and datacenter computing
quality of service
0.011998
Switcherland: A QoS Communication Architecture for Workstation Clusters · ISCA 1998
Internet of things and sensor networks
wireless sensor network
0.012005
Energy Analysis of Public-Key Cryptography for Wireless Sensor Networks · PerCom 2005
Embedded and real-time systems
resource-constrained computing
0.012004
Comparing Elliptic Curve Cryptography and RSA on 8-bit CPUs · CHES 2004
Cryptographic primitives and cryptanalysis
cryptographic implementation
0.012002
An End-to-End Systems Approach to Elliptic Curve Cryptography · CHES 2002
Integrated circuit design › digital circuit design
cryptographic hardware
0.012002
Generic implementations of elliptic curve cryptography using partial reduction · CCS 2002
Interconnection networks and networks-on-chip
low-latency communication
0.012002
Separated high-bandwidth and low-latency communication in the cluster interconnect Clint · SC 2002
Cryptographic primitives and cryptanalysis › cryptographic implementation
block cipher implementation
0.011992
A High-Speed DES Implementation for Network Applications · CRYPTO 1992
Internet architecture and protocols › resource reservation
bandwidth reservation
0.011998
Switcherland: A QoS Communication Architecture for Workstation Clusters · ISCA 1998
Internet architecture and protocols
quality of service
0.011998
Switcherland: A QoS Communication Architecture for Workstation Clusters · ISCA 1998

Methods — techniques the papers use, named apart from their topics

elliptic curve cryptography · 0.2simulation · 0.1energy measurement · 0.1SSL · 0.1performance comparison · 0.1performance analysis · 0.1partial reduction · 0.1generic curve implementation · 0.1HTTPS · 0.1HTTP · 0.1service class design · 0.0crossbar switching · 0.0systems engineering · 0.0scheduling strategies · 0.0
YearPublicationVenuePosition
2018 Light-weight protocols for wire-speed ordering
Hans Eberle, Larry Dennison
SC1
2017 Relaxations for High-Performance Message Passing on Massively Parallel SIMT Processors
abstract
Accelerators, such as GPUs, have proven to be highly successful in reducing execution time and power consumption of compute-intensive applications. Even though they are already used pervasively, they are typically supervised by general-purpose CPUs, which results in frequent control flow switches and data transfers as CPUs are handling all communication tasks. However, we observe that accelerators are recently being augmented with peer-to-peer communication capabilities that allow for autonomous traffic sourcing and sinking. While appropriate hardware support is becoming available, it seems that the right communication semantics are yet to be identified. Maintaining the semantics of existing communication models, such as the Message Passing Interface (MPI), seems problematic as they have been designed for the CPU’s execution model, which inherently differs from such specialized processors. In this paper, we analyze the compatibility of traditional message passing with massively parallel Single Instruction Multiple Thread (SIMT) architectures, as represented by GPUs, and focus on the message matching problem. We begin with a fully MPI-compliant set of guarantees, including tag and source wildcards and message ordering. Based on an analysis of exascale proxy applications, we start relaxing these guarantees to adapt message passing to the GPU’s execution model. We present suitable algorithms for message matching on GPUs that can yield matching rates of 60M and 500M matches/s, depending on the constraints that are being relaxed. We discuss our experiments and create an understanding of the mismatch of current message passing protocols and the architecture and execution model of SIMT processors.
Benjamin Klenk, Holger Fröning, Hans Eberle, Larry Dennison
IPDPS3
2012 Classes of service for daisy chain interconnects
abstract
We describe qHTFair, a switch scheduler that supports classes of service for networks on chips. The scheduler extends the HTFair scheduler, which is an improved version of the HyperTransport scheduling protocol. qHTFair is intended for on-chip interconnects with a daisy chain topology. With our extension, the interconnect can be divided into several classes or channels, each with its own bandwidth allocation. Bandwidth allocations are defined as ratios of channel bandwidths. Ratios can have arbitrary values and be set up dynamically.
Hans Eberle, Wladek Olesinski
HPSR1
2011 Towards an Efficient NoC Topology through Multiple Injection Ports
abstract
In this paper, we present a flexible network on-chip topology: NR-Mesh (Nearest neighbor Mesh). The topology gives an end node the choice to inject a message through different neighboring routers, thereby reducing hop count and saving latency. At the receiver side, a message may be delivered to the end node through different routers, thus reducing hop count further and increasing flexibility when routing messages. This flexibility allows for maximizing network components to be in switch off mode, thus enabling power aware routing algorithms. Additional benefits are reduced congestion/contention levels in the network, support for efficient broadcast operations, savings in power consumption, and partial fault-tolerance. Our second contribution is a power management technique for the adaptive routing. This technique turns router ports and their attached links on and off depending on traffic conditions. The power management technique is able to achieve significant power savings when there is low traffic in the network. We further compare the new topology with the 2D-Mesh, using either deterministic or adaptive routing. When compared with the 2D-Mesh using deterministic routing, executing real applications in a full system simulation platform, the NR-Mesh topology using adaptive routing is able to obtain significant savings, 7% of reduction in execution time and 75% in energy consumption at the network on average for a 16-Node CMP System. Similar numbers are achieved for a 32-Node CMP system.
Jesús Camacho Villanueva, José Flich, José Duato, Hans Eberle, Wladek Olesinski
DSD4
2010 Simple two-priority, low-jitter scheduler
abstract
Many papers on emulations of Generalized Processor Sharing (GPS) have been published. The algorithms and their implementations are often very complex and/or generate a bursty output. In this paper, we present a simple two-priority scheduler that can be easily implemented in hardware, making it especially interesting for Networks on Chips (NoCs), and other applications dealing with stringent resource constraints.
Wladek Olesinski, Hans Eberle
ANCS2
2009 Backlog Aware Scheduling for Ingress Memories in High-Radix, Single-Stage Switches
abstract
Previous work has proposed the Dynamic Switch Buffer Management (DSBM) scheme, a promising approach for improving the scalability of switch ingress memories. In this paper, we extend these results by adding increased backlog awareness into the latter. In particular, we propose using a novel combination of two backlog-aware algorithms: BA-DSBM to map incoming packets to the switch's ingress buffers and the Backlog-Aware Wrapped Wavefront Arbiter (BA-WWFA) to set the configuration of the switch fabric. We then simulate these algorithms under a variety of load intensities and types. These simulations suggest that adding backlog-awareness into the DSBM scheme leads to significant performance enhancements, particularly as the switch is "stressed" by asymmetric or heavy loading. Our algorithms, therefore, mitigate some of the design tradeoffs made in this novel, highly-scalable switch design.
Dimitrios Tsamis, Benjamin Yolken, Nicholas Bambos, Wladek Olesinski, Hans Eberle, Nils Gura
GLOBECOM5
2009 Scalable Alternatives to Virtual Output Queuing
abstract
To avoid head of line blocking in switches, virtual output queues (VOQs) are commonly used. However, the number of VOQs grows quadratically with the number of ports, making this approach impractical for large switches. In this paper, we propose dynamic switch buffer management (DSBM) to tackle this problem. Similar to DBBM, it saves memory by reducing the number of buffers. Our scheme significantly improves the performance by dynamically assigning the incoming cells to the least occupied buffers.
Wladek Olesinski, Hans Eberle, Nils Gura
ICC2
2008 Backlog Aware Scheduling for Large Buffered Crossbar Switches
abstract
A novel architecture was proposed in [1] to address scalability issues in large, high speed packet switches. The architecture proposed in [1], namely OBIG (output buffers with input groups), distributes the switch fabric across multiple chips, which communicate via high speed interconnects enabled by proximity communication (PC), a recently developed circuit technology [2]. An OBIG switch aggregates multiple input flows inside the switch fabric, thereby significantly reducing the amount of memory required for internal buffers, vis-a-vis a conventional buffered crossbar, which has buffers at every crosspoint. Thus, the OBIG architecture is promising for realizing terabit switches with hundreds of ports. This paper studies packet scheduling algorithms which help realize the potential of OBIG-like switch architectures. The emphasis here is on designing backlog aware scheduling algorithms, while ensuring desirable traits such as low computational complexity and scalability. The efficacy of the proposed scheduling algorithms with respect to performance metrics such as average delay and fairness is demonstrated via simulations under a variety of scenarios.
Aditya Dua, Benjamin Yolken, Nicholas Bambos, Wladek Olesinski, Hans Eberle, Nils Gura
ICC5
2008 High-radix crossbar switches enabled by proximity communication
abstract
We describe a novel way to implement high-radix crossbar switches. Our work is enabled by a new chip interconnect technology called proximity communication (PxC) that offers unparalleled chip IO density. First, we show how a crossbar architecture is topologically mapped onto a PxC-enabled multi-chip module (MCM). Then, we describe a first prototype implementation of a small-scale switch based on a PxC MCM. Finally, we present a performance analysis of two large-scale switch configurations with 288 ports and 1,728 ports, respectively, contrasting a 1-stage PxC-enabled switch and a multi-stage switch using conventional technology. Our simulation results show that (a) arbitration delays in a large 1-stage switch can be considerable, (b) multi-stage switches are extremely susceptible to saturation under non-uniform traffic, a problem that becomes worse for higher radices (1-stage switches, in contrast, are not affected by this problem).
Hans Eberle, Pedro Javier García, José Flich, José Duato, Robert J. Drost, Nils Gura, David Hopkins 0001, Wladek Olesinski
SC1
2007 Low-latency scheduling in large switches
abstract
Scheduling in large switches is challenging. Arbiters must operate at high rates to keep up with the high switching rates demanded by multi-gigabit-per-second link rates and short cells. Low-latency requirements of some applications also challenge the design of schedulers. In this paper, we propose the Parallel Wrapped Wave Front Arbiter with Fast Scheduler (PWWFA-FS). We analyze its performance, present simulation results, discuss its implementation, and show how this scheme can provide low latency under light load while scaling to large switches with multi-terabit-per-second throughput and hundreds of ports.
Wladek Olesinski, Nils Gura, Hans Eberle, Andres Mejia
ANCS3
2007 OBIG: the Architecture of an Output Buffered Switch with Input Groups for Large Switches
abstract
Large, fast switches require novel approaches to architecture and scheduling. In this paper, we propose theOutputBufferedSwitchwithInputGroups(OBIG). We present simulation results, discuss the implementation, and show how our architecture can be used to build single-stage (flat) switches with multi-terabit-per-second throughput and hundreds of ports.
Wladek Olesinski, Hans Eberle, Nils Gura
GLOBECOM2
2005 Architectural Extensions for Elliptic Curve Cryptography over GF(2m) on 8-bit Microprocessors
abstract
We describe and analyze architectural extensions to accelerate the public key cryptosystem elliptic curve cryptography (ECC) on 8-bit microprocessors. We show that simple extensions of the data path suffice to efficiently support ECC over GF(2/sup m/). These extensions include an extended multiplier that generates results for both integer multiplications and multiplications in fields GF(2/sup m/) and a multiply-accumulate instruction for efficiently performing multiple precision multiplications. To our knowledge, this is the first paper that quantifies performance of standard NIST and SECG elliptic curves over GF(2/sup m/) on an 8-bit microprocessor equipped with a dual field multiplier. On the ATmegal28 microprocessor running at 8 MHz we measured an execution time of 0.29 s for a 163-bit ECC point multiplication over GF(2/sup m/), 0.81s for a 160-bit ECC point multiplication over GF(p), and 11 s for a 1024-bit RSA private key operation - the chosen key sizes provide equivalent security strength.
Hans Eberle, Arvinderpal Wander, Nils Gura, Sheueling Chang Shantz
ASAP1
2005 Sizzle: A Standards-Based End-to-End Security Architecture for the Embedded Internet (Best Paper)
abstract
This paper introduces Sizzle, the first fully implemented end-to-end security architecture for highly constrained embedded devices. According to popular perception, public-key cryptography is beyond the capabilities of such devices. We show that elliptic curve cryptography (ECC) not only makes public-key cryptography feasible on these devices, it allows one to create a complete secure Web server stack including SSL, HTTP and user application that runs efficiently within very tight resource constraints. Our small footprint HTTPS stack needs less than 4 KB of RAM and interoperates with an ECC-enabled version of the Mozilla Web browser. We have implemented Sizzle on the 8-bit Berkeley/Crossbow Mica2 "mote" platform where it can complete a full SSL handshake in less than 4 seconds (session reuse takes under 2 seconds) and transfer 450 bytes of application data over SSL in about 1 second. We present additional optimizations that can further improve performance. To the best of our knowledge, this is the world's smallest secure Web server (in terms of both physical dimensions and resources consumed) and significantly lowers the barrier for connecting a variety of interesting new devices (e.g. home appliances, personal medical devices) to the Internet without sacrificing end-to-end security.
Matthew Millard, Stephen Fung, Nils Gura, Hans Eberle, Sheueling Chang Shantz
PerCom6
2005 Energy Analysis of Public-Key Cryptography for Wireless Sensor Networks
abstract
In this paper, we quantify the energy cost of authentication and key exchange based on public-key cryptography on an 8-bit microcontroller platform. We present a comparison of two public-key algorithms, RSA and elliptic curve cryptography (ECC), and consider mutual authentication and key exchange between two untrusted parties such as two nodes in a wireless sensor network. Our measurements on an Atmel ATmega128L low-power microcontroller indicate that public-key cryptography is very viable on 8-bit energy-constrained platforms even if implemented in software. We found ECC to have a significant advantage over RSA as it reduces computation time and also the amount of data transmitted and stored.
Arvinderpal Wander, Nils Gura, Hans Eberle, Sheueling Chang Shantz
PerCom3
2005 Sizzle: A standards-based end-to-end security architecture for the embedded Internet
Michael Wurm, Matthew Millard, Stephen Fung, Nils Gura, Hans Eberle, Sheueling Chang Shantz
Pervasive Mob. Comput.7
2004 A Public-Key Cryptographic Processor for RSA and ECC
Hans Eberle, Nils Gura, Sheueling Chang Shantz, Leonard Rarick, Shreyas Sundaram
ASAP1
2004 Comparing Elliptic Curve Cryptography and RSA on 8-bit CPUs
Nils Gura, Arun Patel, Arvinderpal Wander, Hans Eberle, Sheueling Chang Shantz
CHES4
2004 Speeding up Secure Web Transactions Using Elliptic Curve Cryptography
Douglas Stebila, Stephen Fung, Sheueling Chang Shantz, Nils Gura, Hans Eberle
NDSS6
2004 Testing Systems Wirelessly
abstract
Wired test structures exhibit many unwanted dependencies: they typically use hierarchical and daisy-chained wiring, and they share interconnects and backplanes with the system under test. As a result, faults can easily lead to incomplete or erroneous test reports on properly working components. Wireless test structures do not have these shortcomings and thus, allow for more accurate testing and diagnosing. Wireless communication further allows for non-intrusive testing that does not require any cabling or physical access to the system under test. We describe two prototype implementations: a wireless field-replaceable unit ID and a wireless version of the popular JTAG standard.
Hans Eberle, Arvinderpal Wander, Nils Gura
VTS1
2003 A Cryptograhpic Processor for Arbitrary Elliptic Curves over
abstract
We describe a cryptographic processor for elliptic curve cryptography (ECC). ECC is evolving as an attractive alternative to other public-key schemes such as RSA by offering the smallest key size and the highest strength per bit. The processor performs point multiplication for elliptic curves over binary polynomial fields GF(2/sup m/). In contrast to other designs that only support one curve at a time, our processor is capable of handling arbitrary curves without requiring reconfiguration. More specifically, it can handle both named curves as standardized by NIST as well as any other generic curves up to a field degree of 255. Efficient support for arbitrary curves is particularly important for the targeted server applications that need to handle requests for secure connections generated by a multitude of heterogeneous client devices. Such requests may specify curves which are infrequently used or not even known at implementation time. Our processor implements 256 bit modular multiplication, division, addition and squaring. The multiplier constitutes the core function as it executes the bulk of the point multiplication algorithm. We present a novel digit-serial modular multiplier that uses a hybrid architecture to perform the reduction operation needed to reduce the multiplication result: hardwired logic is used for fast reduction of named curves and the multiplier circuit is reused for reduction of generic curves. The performance of our FPGA-based prototype, running at a clock frequency of 66.4 MHz, is 6955 point multiplications per second for named curves over GF(2/sup 163/) and 3308 point multiplications per second for generic curves over GF(2/sup 163/).
Hans Eberle, Nils Gura, Sheueling Chang Shantz
ASAP1
2002 Generic implementations of elliptic curve cryptography using partial reduction
abstract
Elliptic Curve Cryptography (ECC) is evolving as an attractive alternative to other public-key schemes such as RSA by offering the smallest key size and the highest strength per bit. The importance of ECC has been recognized by the US government and the standards bodies NIST and SECG. Standards for preferred elliptic curves over prime fields GF(p) and binary polynomial fields GF(2m) as well as the Elliptic Curve Digital Signature Algorithm (ECDSA) have been created. A security protocol based on ECC requires support for different curves representing different security levels. This is particularly true for server applications that are exposed to requests for secure connections with different parameters generated by a multitude of client devices. Reported implementations of ECC over GF(2m) typically choose to implement each curve as a special case so that modular reduction can be optimized, thus improving the overall performance. In contrast, this paper focuses on generic implementations of ECC point multiplication for arbitrary curves over GF(2m). We present a novel reduction algorithm that allows hardware and software implementations for variable field degrees m. Though not as high in performance as an implementation optimized for a specific curve, it offers an attractive solution to supporting infrequently used curves or curves not known at the time of the implementation.
Nils Gura, Hans Eberle, Sheueling Chang Shantz
CCS2
2002 An End-to-End Systems Approach to Elliptic Curve Cryptography
Nils Gura, Sheueling Chang Shantz, Hans Eberle, Daniel F. Finchelstein, Edouard Goupy, Douglas Stebila
CHES3
2002 Separated high-bandwidth and low-latency communication in the cluster interconnect Clint
abstract
An interconnect for a high-performance cluster has to be optimized in respect to both high throughput and low latency. To avoid the tradeoff between throughput and latency, the cluster interconnect Clint1 has a segregated architecture that provides two physically separate transmission channels: A bulk channel optimized for high-bandwidth traffic and a quick channel optimized for low-latency traffic. Different scheduling strategies are applied. The bulk channel uses a scheduler that globally allocates time slots on the transmission paths before packets are sent off. This way collisions as well as blockages are avoided. In contrast, the quick channel takes a best-effort approach by sending packets whenever they are available thereby risking collisions and retransmissions. Simulation results clearly show the performance advantages of the segregated architecture. The carefully scheduled bulk channel can be loaded nearly to its full capacity without exhibiting head-of-line blocking that limits many networks while the quick channel provides low-latency communication even in the presence of high-bandwidth traffic.
Hans Eberle, Nils Gura
SC1
1998 Switcherland: A QoS Communication Architecture for Workstation Clusters
abstract
Computer systems have become powerful enough to process continuous data streams such as video or animated graphics. While processing power and communication bandwidth of today's systems typically are sufficient, quality of service (QoS) guarantees as required for handling such data types cannot be provided by these systems in adequate ways. We present Switcherland, a scalable communication architecture based on crossbar switches that provides QoS guarantees for workstation clusters in the form of reserved bandwidth and bounded transmission delays. Similar to the ATM technology Switcherland provides QoS guarantees with the help of service classes, that is, data transfers are characterized as variable bit rare traffic or constant bit rate traffic. However, unlike LAN technologies, Switcherland is optimized for cluster computing in that (i) it serves as a backplane interconnection fabric as well as a LAN, (ii) it extends support for service classes by also covering the end nodes of the network, (iii) it provides low latency in the order of one microsecond per switch, and (iv) it uses a communication model based on a global memory to simplify programming.
Hans Eberle, Erwin Oertli
ISCA1
1998 Switcherland: A scalable interconnection structure for distributed systems
Hans Eberle
J. Syst. Archit.1
1997 NetRAP - A Network Resource Allocation Protocol for IP over Ethernet
abstract
For more than a decade, IP over the Ethernet has been the dominating local area network technology because it is cost-efficient, simple and reliable. Unfortunately, it does not support the transport of continuous data as generated by multimedia applications. The large number of IP installations makes it necessary to find ways to provide support for multimedia applications in existing network environments without making interfaces used by existing applications incompatible. In this paper, we present the design, implementation and evaluation of the network resource allocation protocol NetRAP. NetRAP extends IP in that it allows applications to allocate bandwidth. Access to the network happens in two phases. In the first phase, a token-based access protocol is used to give applications with bandwidth reservations access to the network. In the second phase, the network is operated in normal CSMA/CD mode giving all other applications without bandwidth reservations the opportunity to access the network. For this purpose, we have extended an existing IP stack with a transport protocol able to control access times. Our performance measurements show that NetRAP adds little overhead. Since the implementation of NetRAP only requires few modifications of existing protocol software and, moreover, does not invalidate existing networked applications, it offers a simple and cost-efficient alternative to other more complex solutions.
Martin Gitsels, Hans Eberle, Christian Kleitsch
ICCCN2
1992 A High-Speed DES Implementation for Network Applications
Hans Eberle
CRYPTO1
1988 Analysis of processor-memory communication by the NS 32000 processor family
Hans Eberle
Microprocess. Microprogramming1