EDBT 2026 Demo / reviewers in the wild / expert
Jörg Nolte
dblp:85/1892
· DBLP profile ↗
38ranked-venue papers
5as first author
8since 2021 · last 2025
0000-0003-3818-5402ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 15 · 5 first-author · 3 since 2021Computer networks · 10 · 3 since 2021Software engineering, systems software and programming languages · 4Security and privacy · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | VibroMote: Wi-Fi-based Mesh Communication for Railway Bridge Inspection and MonitoringabstractVibration sensing provides insights into the dynamic behaviour of engineering constructions such as railway bridges. Cable-based sensors are viable only for long-term condition monitoring and rare special inspections because of the labor-intensive deployment. Although wireless sensors significantly reduce this overhead, their energy constraints limited the network throughput and, hence, their resolution in space and time. Batteryor solar-powered sensor nodes with high network throughput would enable in-depth measurements during regular inspections and improve the access to high-quality monitoring data. We present “VibroMote”, which combines energy harvesting, a high-bandwidth 3-axis MEMS accelerometer, and high-throughput, self-organizing mesh communication via IEEE 802.11 Wi-Fi. The evaluation on a real bridge with 18 VibroMotes shows that the deployment time can be reduced from hours to minutes; the multi-hop mesh provides sufficient throughput reserves for the application whereas direct one-hop communication failed; the battery runtime is sufficient for temporary measurements during inspections; and the energy consumption during sleep modes would be low enough for solar-powered long-term monitoring. Hence, the combination of high-throughput Wi-Fi with mesh networking is a strong alternative to the commonly used low-power radio technologies when long range between the sensors is not needed. Sneha Chatharajupalli Navya, Randolf Rotta, Reinhardt Karnapke, Jörg Nolte |
ISNCC | 4 |
| 2025 | Probing Considered Harmful: Leveraging RSSI for Link Quality PredictionabstractWireless mesh protocols that use a routing metric based on throughput or airtime require predictions about future throughput for all mesh neighbors. Typically, the Wi-Fi rate controller provides this information, and state-of-the-art rate control algorithms estimate it based on statistics from past unicast transmissions. This introduces an adverse cross-layer dependency because meaningful statistics are only available for the links that were used by the routing layer. Existing implementations either ignore this, or use other routing metrics, such as distance, or generate artificial periodic probe traffic to all neighbors. Unfortunately, probing reacts slowly to changing conditions, increases the medium contention, and can lead to overly optimistic predictions. To overcome this, we introduce an RSSI-based link quality prediction for the routing layer. Benchmarks in a B.A.T.M.A.N. version V mesh show a decrease of round-trip time by three orders of magnitude and an increase of the end-to-end packet delivery ratio. TCP connections became possible over long routes that were previously unusable. Although highly imprecise, the easy to acquire RSSI provided sufficiently good predictions for our mesh network. Sneha Chatharajupalli Navya, Randolf Rotta, Reinhardt Karnapke, Jörg Nolte |
LCN | 4 |
| 2025 | Practical Whole-System PersistenceabstractSudden power outages remain one of the biggest threats to losing data, disrupting systems and causing financial damages. Whole system persistence (WSP) has previously been proposed as a solution to mitigate such threats through the use of non-volatile main memory (NVRAM). However, it missed out on external device state persistence and the NVRAM technology used at the time was expensive and limited with regard to their scalability. Today's NVRAM technologies are both more affordable and offer much higher storage capacities, but are typically slower than DRAM. Dustin T. Nguyen, Oliver Giersch, Thomas Preisner, Jonathan Krebs, Henriette Herzog, Rüdiger Kapitza, Jörg Nolte, Timo Hönig, Wolfgang Schröder-Preikschat |
SYSTOR | 7 |
| 2024 | Demo: B.A.T.M.A.N. Mesh Routing on Ultra Low-Power IEEE 802.11 ModulesabstractWhen implementing multi-hop mesh network protocols, efficient direct communication, route discovery, and route repair are crucial to achieve high network throughput. To our knowledge, there is no open mesh routing protocol available for ultra low power IEEE 802.11 modules. Previous work relied on single board computers like Raspberry Pi. We implemented the B.A.T.M.A.N. mesh protocol on the popular Espressif ESP32 along with a novel hybrid rate adaptation for the selection of efficient routes. In this demo we showcase challenges and solutions related to the implementation on ESP32 and how the hybrid rate adaptation improves end-to-end throughput. Randolf Rotta, Sneha Chatharajupalli Navya, Billy Naumann, Julius Schulz, Reinhardt Karnapke, Matthias Werner 0001, Jörg Nolte |
LCN | 7 |
| 2024 | B.A.T.M.A.N. Mesh Networking on ESP32's 802.11abstractMesh routing protocols are widely used in IoT and sensor networks. In recent years, the ESP32 Wi-Fi/BLE SoC became popular for prototyping IoT applications. However, the existing mesh networks for this platform lack efficient node to node communication, fast route discovery and repair, and energy efficiency. This paper addresses the formation of IEEE 802.11 based ad-hoc mesh networks without the delays inflicted by the Station to Access Point association protocol. We implemented the B.A.T.M.A.N. protocol on top of the ESP32 Wi-Fi MAC interface and integrated it into the LwIP network stack. The performance evaluation with respect to UDP/IP and TCP/IP end-to-end throughput shows the general usefulness but also identifies bottlenecks caused by limitations of the existing MAC interface. Overall, this opens an interesting opportunity for research on mesh protocols by providing a simpler platform than full featured Wi-Fi routers; and for wireless IoT applications by providing higher throughput than subGHz and BLE technologies. Randolf Rotta, Julius Schulz, Billy Naumann, Sneha Chatharajupalli Navya, Jörg Nolte, Matthias Werner 0001 |
LCN | 5 |
| 2023 | Assessing the Feasibility of Combined BLE and Wi-Fi Communication for High Data Sensing Applications
Sneha Chatharajupalli Navya, Randolf Rotta, Reinhardt Karnapke, Jörg Nolte |
EWSN | 4 |
| 2022 | Fast and Portable Concurrent FIFO Queues With Deterministic Memory ReclamationabstractIn this article we present an algorithm for a high performance, unbounded, portable, multi-producer/multi-consumer, lock-free FIFO (first-in first-out) queue. Aside from its competitive performance on current hardware, it is further characterized by its integrated memory reclamation mechanism, which is able to reliably and deterministically de-allocate nodes as soon as the final operation with a reference has concluded, similar to reference counting. This differentiates our approach from most other lock-free data structures, which usually require external (generic) memory reclamation or garbage collection mechanisms such as hazard pointers. Our deterministic memory reclamation mechanism completely prevents the build up of memory awaiting reclamation and is hence very memory efficient, yet it does not introduce any substantial performance overhead. By utilizing concrete knowledge about the internal structure and access patterns of our queue, we are able to construct and constrain the reclamation mechanism in such a way that keeps the overhead for memory management almost entirely out of the common fast path. The presented algorithm is portable to all modern 64-bit processor architectures, as it only relies on the commonly available and lock-free atomic synchronization primitives compare-and-swap and fetch-and-add. Oliver Giersch, Jörg Nolte |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2021 | Nowa: A Wait-Free Continuation-Stealing Concurrency PlatformabstractIt is an ongoing challenge to efficiently use parallelism with today's multi- and many-core processors. Scalability becomes more crucial than ever with the rapidly growing number of processing elements in many-core systems that operate in data centres and embedded domains. Guaranteeing scalability is often ensured by using fully-strict fork/join concurrency, which is the prevalent approach used by concurrency platforms like Cilk. The runtime systems employed by those platforms typically resort to lock-based synchronisation due to the complex interactions of data structures within the runtime. However, locking limits scalability severely. With the availability of commercial off-the-shelf systems with hundreds of logical cores, this is becoming a problem for an increasing number of systems.This paper presents Nowa, a novel wait-free approach to arbitrate the plentiful concurrent strands managed by a concurrency platform's runtime system. The wait-free approach is enabled by exploiting inherent properties of fully-strict fork/join concurrency, and hence is potentially applicable for every continuation-stealing runtime system of a concurrency platform. We have implemented Nowa and compared it with existing runtime systems, including Cilk Plus, and Threading Building Blocks (TBB), which employ a lock-based approach. Our evaluation results show that the wait-free implementation increases the performance up to 1.64× compared to lock-based ones, on a system with 256 hardware threads. The performance increased by 1.17× on average, while no but one benchmark exhibited performance regression. Compared against OpenMP tasks using Clang's libomp, Nowa outperforms OpenMP by 8.68× on average. Florian Schmaus, Nicolas Pfeiffer, Wolfgang Schröder-Preikschat, Timo Hönig, Jörg Nolte |
IPDPS | 5 |
| 2020 | RESCUE: Interdependent Challenges of Reliability, Security and Quality in Nanoelectronic SystemsabstractThe recent trends for nanoelectronic computing systems include machine-to-machine communication in the era of Internet-of-Things (IoT) and autonomous systems, complex safety-critical applications, extreme miniaturization of implementation technologies and intensive interaction with the physical world. These set tough requirements on mutually dependent extra-functional design aspects. The H2020 MSCAITN project RESCUE is focused on key challenges for reliability, security and quality, as well as related electronic design automation tools and methodologies. The objectives include both research advancements and cross-sectoral training of a new generation of interdisciplinary researchers. Notable interdisciplinary collaborative research results for the first halfperiod include novel approaches for test generation, soft-error and transient faults vulnerability analysis, cross-layer fault-tolerance and error-resilience, functional safety validation, reliability assessment and run-time management, HW security enhancement and initial implementation of these into holistic EDA tools. Maksim Jenihhin, Said Hamdioui, Matteo Sonza Reorda, Milos Krstic, Peter Langendörfer, Christian Sauer 0001, Anton Klotz, Michael Hübner 0001, Jörg Nolte, Heinrich Theodor Vierhaus, Georgios N. Selimis, Dan Alexandrescu, Mottaqiallah Taouil, Geert Jan Schrijen, Jaan Raik, Luca Sterpone, Giovanni Squillero, Zoya Dyka |
DATE | 9 |
| 2020 | Real-Time Dynamic Hardware Reconfiguration for Processors with Redundant Functional UnitsabstractThe tiny logic elements in modern integrated circuits increase the rate of transient failures significantly. Therefore, redundancy on various levels is necessary to retain reliability. However, for mixed-criticality scenarios, the typical processor designs offer either too little fault-tolerance or too much redundancy for one part of the applications. Amongst others, we specifically address redundant processor internal functional units (FU) to cope with transient errors and support wear leveling. A real-time operating system (RTOS) was extended to control our prototypical hardware platform and, since it can be configured deterministically within few clock cycles, we are able to reconFigure the FUs dynamically, at process switching time, according to the specified critically of the running processes. Our mechanisms were integrated into the Plasma processor and the Plasma-RTOS. With few changes to the original software code, it was, for example, possible to quickly change from fault-detecting to fault-correcting modes of the processor on demand. Randolf Rotta, Raphael Segabinazzi Ferreira, Jörg Nolte |
ISORC | 3 |
| 2020 | A Modified Rejection-Based Architecture to Find the First Two Minima in Min-Sum-Based LDPC DecodersabstractOne of the essential elements of min-sum low-density parity-check (LDPC) decoders is to find the first two minima between the binary messages arriving in the check nodes along with the index of the minimum which are altogether used to compute the messages for sending back to the neighboring variable nodes. The main techniques for this task are tree-based and bit-serial architectures. The latest tree-based architecture, known as rejection-based scheme finds the first two minima and the binary index of the minimum with higher speed than the previous tree-based methods. However, in min-sum LDPC decoders, having one-hot sequence of the minimum of the messages is preferred as it has implementation benefits. In this paper, we modify the existing rejection-based technique to yield the one-hot sequence instead of the binary representation of the minimum index. The proposed modification doesn’t cause any latency in the operation of the module. We also provide the results of the implementation of the modified rejection-based technique and the bit-serial architecture, conducted on a Xilinx Virtex-7 FPGA. The two major architectures are compared in terms of latency, maximum clock frequency, area and power. Alireza Hasani, Lukasz Lopacinski, Steffen Büchner 0002, Jörg Nolte, Rolf Kraemer |
WCNC | 4 |
| 2019 | A Modified Shuffling Method to Split the Critical Path Delay in Layered Decoding of QC-LDPC CodesabstractLayered (or Turbo) decoding of Low-Density Parity-Check (LDPC) codes is considered as a decoding schedule that facilitates partially parallel architectures for performing iterative algorithms based on belief propagation. It has, on one hand, reduced implementation complexity and memory overhead compared to fully parallel architectures and, on the other hand, higher convergence speed compared to both serial and parallel architectures. In this paper, we introduce a general form of shuffling of the parity-check matrix of quasi-cyclic LDPC (QC-LDPC) codes which can split the critical path delay in layered decoding and therefore improve throughput by allowing higher clock rates. We also reveal a valuable property of Latin squares QC-LDPC codes which makes them a good candidate for the proposed shuffling method. As a result of that property, no special caution of choosing offset values in the proposed generalized shuffling method is required. Alireza Hasani, Lukasz Lopacinski, Steffen Büchner 0002, Jörg Nolte, Rolf Kraemer |
PIMRC | 4 |
| 2019 | Cache-Line Transactions: Building Blocks for Persistent Kernel Data Structures Enabled by AspectC++abstractWith the availability of systems that contain large amounts of byte-addressable non-volatile memory (NVRAM), there is a growing need for data structures that can be mapped into a process's address space and be used without data (de-)serialization. While NVRAM is able to retain memory contents during system failure and power loss, data consistency has to be preserved by using transactional operations for data manipulation. Marcel Köppen, Jana Traue, Christoph Borchert, Jörg Nolte, Olaf Spinczyk |
PLOS@SOSP | 4 |
| 2018 | 100 Gbit/s End-to-End Communication: Adding Flexibility with Protocol TemplatesabstractHigh-speed protocol processing that provides data-rates of 100 Gbit/s and beyond to the application stresses the whole communication system up to its outer limits. Such a system can only be utilized by employing highly specialized, application specific protocols, that are tailored for certain communication parameters, such as the packet loss rate. However, the requirements for most applications are not static, and a protocol designer cannot anticipate all possible communication conditions upfront. The contradiction between specialized protocols and unknown communication parameters can be solved by adapting the protocol implementation on demand to the current communication conditions. However, such an approach needs a protocol description language that allows the automatic specialization of protocols. In this paper, we present the Protocol Engine Template Language (PETL), that allows the automatic implementation of protocols by a constructive approach for a variety of communication conditions from protocol implementation templates. Steffen Büchner 0002, Alireza Hasani, Lukasz Lopacinski, Rolf Kraemer, Jörg Nolte |
LCN | 5 |
| 2017 | 100 Gbit/s End-to-End Communication: Low Overhead On-Demand Protocol Replacement in High Data Rate Communication SystemsabstractTo be able to efficiently utilize high data rates of 100 Gbit/s and beyond, protocols must be carefully selected for specific communication parameters. At the same time, communication parameters, such as data rate/latency requirements and the channel quality, are not static. This contradiction can be solved by switching to the best suited protocol when communication parameters change. However, replacing a protocol is a severe interference in an ongoing transmission that can easily cause performance degradation. In this paper, we present a minimal disruptive replacement approach that allows us to replace protocol implementations for an ongoing transmission. Steffen Büchner 0002, Jörg Nolte, Alireza Hasani, Rolf Kraemer |
LCN | 2 |
| 2016 | Influence of Topology-Fluctuations on Self-Stabilizing AlgorithmsabstractSelf-stabilizing systems have in theory the unique and provable ability, to always return to a valid system state even in the face of failures. These properties are certainly desirable for domains like wireless ad-hoc networks with numerous unpredictable faults. Unfortunately, the time in which the system returns to a valid state is not predictable and potentially unbound. The failure rate typically depends on physical phenomena and in self-stabilizing systems each node tries to react to failures in an inherently adaptive fashion by the cyclic observation of the states of its neighbors. When state changes are either too quick or too slow the system might never reach a state that is sufficiently stable for a specific task. In this paper, we investigate the influences of the error rate on the (stability) convergence time on the basis of topology information acquired in real network experiments. This allows us to asses the asymptotic behavior of relevant self-stabilizing algorithms in typical wireless networks. Stefan Lohs, Gerry Siegemund, Jörg Nolte, Volker Turau |
DCOSS | 3 |
| 2016 | 100 Gbit/s End-to-End Communication: Designing Scalable Protocols with Soft Real-Time Stream ProcessingabstractWith the recent roll-out of 100 Gbit Ethernet technology for high-performance computing applications and the technology for 100 Gbit wireless communication emerging on the horizon, it is just a matter of time until non-high performance computing applications will have to utilize these data rates. Since 10 Gbit/s protocol processing is already challenging for current server machines and simply upscaling the computing resources is no solution, new approaches are needed. In this paper, we present a stream processing based design approach for scalable communication protocols. The stream processing paradigm enables us to adapt the communication protocol processing for a certain hardware configuration without touching the protocol's implementation. We use this design technique to develop a prototype communication protocol for ultra-high throughput applications and we demonstrate how to adapt the protocol processing for a Stable Throughput as well as for a Low Latency scenario. Last but not least, we present the evaluation results of the experiments, which show that the measured throughput respectively latency of the adapted protocol, scales nearly linear with the number of provided interfaces. Steffen Büchner 0002, Lukasz Lopacinski, Jörg Nolte, Rolf Kraemer |
LCN | 3 |
| 2016 | Improved turbo product coding dedicated for 100 Gbps wireless terahertz communicationabstractIn this article, an improved turbo product decoding scheme is proposed. The new method is almost as effective as hard decodable low-density parity check codes (HD-LDPC). Due to the modified codeword shape, no external interleavers are required to correct burst errors. If the decoder uses Reed-Solomon (RS) codes, then error correction performance against burst errors is significantly higher than the gain provided by HD-LDPC with an external interleaver. An additional advantage is a possibility to design a dedicated decoder for Virtex7 field programmable gate array (FPGA) serial transceivers. In our case, we use the method for 100 Gbps data link layer processor dedicated for wireless communication in the Terahertz band. The targeted platform is Virtex7 FPGA, but the solution can be easily scaled on other technologies. Lukasz Lopacinski, Jörg Nolte, Steffen Büchner 0002, Marcin Brzozowski, Rolf Kraemer |
PIMRC | 2 |
| 2016 | Self-Stabilization - A Mechanism to Make Networked Embedded Systems More Reliable?abstractThe erratic behavior of wireless channels is still a major hurdle in the implementation of robust applications in wireless networks. In the past it has been argued that self-stabilization is a remedy to provide the needed robustness. This assumption has not been verified to the extent necessary to convince engineers implementing such applications. A major reason is that the time in which a self-stabilizing system returns to a valid state is unpredictable and potentially unbound. Failure rates typically depend on physical phenomena and in self-stabilizing systems each node tries to react to failures in an inherently adaptive fashion by the cyclic observation of its neighbors' states. When the frequency of state changes is too high, the system may never reach a state sufficiently stable for a specific task. In this paper we substantiate the conditions under which self-stabilization leads to fault tolerance in wireless networks and look at the myths about the power of self-stabilization as a particular instance of self-organization. We investigate the influences of the error rate and the neighbor state exchange rate on the stability and the convergence time on topology information acquired in real network experiments. Stefan Lohs, Jörg Nolte, Gerry Siegemund, Volker Turau |
SRDS | 2 |
| 2015 | Design and Implementation of an Adaptive Algorithm for Hybrid Automatic Repeat RequestabstractTransmission efficiency is an interesting topic for data link layer developers. The overhead of protocols and coding should be reduced to a minimum. This maximizes a link throughput. This is especially important for high-speed networks, where a small degradation of efficiency will degrade the throughput by several Gbps. We describe a redundancy balancing algorithm for an adaptive hybrid automatic repeat request with Reed-Solomon coding. We introduce a testing environment, most important technical issues, and results generated on a field programmable gate array. The hybrid automatic repeat request and Reed-Solomon algorithms are explained. We provide a mathematical description, and a block diagram of the adaptation algorithm. All necessary algorithm simplifications are explained in details. The algorithm can be represented by basic operations in hardware. In most cases, it finds the optimal coding for a predefined bit error rate. Lukasz Lopacinski, Jörg Nolte, Steffen Büchner 0002, Marcin Brzozowski, Rolf Kraemer |
DDECS | 2 |
| 2015 | Challenges for 100 Gbit/s end to end communication: Increasing throughput through parallel processingabstractToday's applications and services become more dependent on fast wireless communication, for the upcoming years data-rate demands of 100Gbit/s can be easily expected. However, fulfilling that demand is a task which cannot simply be solved by upscaling existing technologies. While most of the research tackles the challenges regarding the transmission technology from the physical layer up to base-band processing, we focus on the challenges concerning the handling of that vast amount of data. The overall goal is to bring together the transmission technology with the operating system to create a suitable end-to-end communication solution. In this paper we argue that communication can be understood as a soft-realtime problem and how that helps introducing parallelism into protocol-processing. Steffen Büchner 0002, Jörg Nolte, Rolf Kraemer, Lukasz Lopacinski, Reinhardt Karnapke |
LCN | 2 |
| 2013 | Online Device-Level Energy Accounting for Wireless Sensor Nodes
André Sieber, Jörg Nolte |
EWSN | 2 |
| 2013 | An Agile and Stable Neighborhood Protocol for WSNs
Gerry Siegemund, Volker Turau, Christoph Weyer, Stefan Lohs, Jörg Nolte |
SSS | 5 |
| 2011 | Using sensor technology to protect an endangered species: A case studyabstractA lot of applications for wireless sensor networks have been proposed in the last years. Only a few of them have led to real, non-academic deployments, partially due to the differences between end user needs and academic assumptions. In this paper we discuss a real world problem arising from an ecological question (protection of an endangered species) and the theoretical solution as well as the deployed solution that actually works. André Sieber, Reinhardt Karnapke, Jörg Nolte, Thomas Martschei |
LCN | 3 |
| 2008 | An existing complete house control system based on the REFLEX operating system: Implementation and experiences over a period of 4 yearsabstractToday, even small residential buildings have a number of complex electrical devices that advocate the usage of automated control systems. But currently available systems are either hard to handle or expensive. This paper describes an automated house control system that has been in use for the last 4 years. It has been built using only freely available, inexpensive hardware and the open source operating system REFLEX. Karsten Walther, Reinhardt Karnapke, Jörg Nolte |
ETFA | 3 |
| 2007 | In-network processing and collective operations using the COCOS-frameworkabstractCOCOS (coordinated communicating sensors) is a lean middleware platform for wireless sensor networks. The major programming abstractions of Cocos are distributed sensor spaces. All objects in such a space can be collectively addressed. This way, high-level data-parallel programming concepts such as global reductions are possible. This paper introduces the spaces of Cocos and describes their usage. Maik Krüger, Reinhardt Karnapke, Jörg Nolte |
ETFA | 3 |
| 2007 | IMPACT - A Family of Cross-Layer Transmission Protocols for Wireless Sensor NetworksabstractFor economic reasons sensor networks are often implemented with resource constrained micro-controllers and low-end radio transceivers. Consequently, communication is inherently unreliable and especially multi-hop communication suffers severely from packet losses. Transmission protocols that rely on implicit acknowledges for multi-hop communication are energy efficient but require symmetric communication links to work properly. In this paper we introduce IMPACT, a family of transmission protocols that rely on implicit acknowledges and employ a cross layer approach to handle asymmetric links. Marcin Brzozowski, Reinhardt Karnapke, Jörg Nolte |
IPCCC | 3 |
| 2007 | Analyzing the real-time behaviour of deeply embedded event driven systemsabstractMost embedded control systems react on events in the real world by reading sensors and controlling actuators in real-time. This general behavior can be directly mapped onto event-driven systems in a natural and straightforward manner for a large variety of applications. Further real-time analysis and profiling on the same level of abstraction is possible for event-driven systems. This significantly helps developers of deeply embedded real-time applications. In this paper we introduce simulative profiling concepts and static analysis basics for the real-time analysis of event-driven systems. Furthermore we present a prototype analysis tool for the REFLEX operating system that integrates real-time analysis into the software development cycle. Karsten Walther, René Herzog, Jörg Nolte |
LCTES | 3 |
| 2006 | COPRA - A Communication Processing Architecture for Wireless Sensor Networks
Reinhardt Karnapke, Jörg Nolte |
Euro-Par | 2 |
| 2002 | Exploiting cluster networks for distributed object groups and collective operations
Jörg Nolte, Mitsuhisa Sato, Yutaka Ishikawa |
Future Gener. Comput. Syst. | 1 |
| 2001 | TACO-Exploiting Cluster Networks for High-Level Collective OperationsabstractTACO (Topologies and Collections) is a template library that introduces the flavour of distributed data parallel processing by means of reusable topology classes and C++ templates. The paper introduces TACO's basic abstractions and provides a performance analysis for basic collective operations on various cluster architectures with several different networks. Jörg Nolte, Mitsuhisa Sato, Yutaka Ishikawa |
CCGRID | 1 |
| 2000 | TACO -- Dynamic Distributed Collections with Templates and Topologies
Jörg Nolte, Mitsuhisa Sato, Yutaka Ishikawa |
Euro-Par | 1 |
| 2000 | Parallel Matching and Sorting with TACO's Distributed Collections - A Case Study from Molecular Biology ResearchabstractTACO is a template library that implements higher-order parallel operations on distributed object sets by means of reusable topology classes and C++ function templates. We discuss an experimental application that exploits TACO's distributed object groups and collective operations for computing the similarity between groups of molecular sequences, a computationally intensive core problem in molecular biology research. In particular we show how TACO's distributed collections can be conveniently combined with well known concepts found in the C++ standard template library (STL) to solve matching and sorting problems effectively on distributed hardware platforms. The resulting implementation is concise and gives excellent parallel performance on PC- and workstation clusters. Jörg Nolte, Paul Horton |
HPDC | 1 |
| 2000 | Template Based Structured CollectionsabstractCollective operations on distributed data sets foster a high-level data-parallel programming style that eases many aspects of parallel programming significantly. In this paper we describe how higher-order collective operations on distributed object sets can be introduced in a structured way by means of reusable topology classes and C++ templates. Jörg Nolte, Mitsuhisa Sato, Yutaka Ishikawa |
IPDPS | 1 |
| 1999 | ARTS of PEACE - A High-Performance Middleware Layer for Parallel Distributed Computing
Lars Büttner, Jörg Nolte, Wolfgang Schröder-Preikschat |
J. Parallel Distributed Comput. | 2 |
| 1998 | Experiences Developing a Virtual Shared Memory System Using High-Level Object Paradigms
Jörg Cordsen, Jörg Nolte, Wolfgang Schröder-Preikschat |
ECOOP | 2 |
| 1995 | Time Space Sharing Scheduling: A Simulation Analysis
Atsushi Hori, Yutaka Ishikawa, Jörg Nolte, Hiroki Konaka, Munenori Maeda, Takashi Tomokiyo |
Euro-Par | 3 |
| 1995 | Time Space Sharing Scheduling and Architectural Support
Atsushi Hori, Takashi Yokota, Yutaka Ishikawa, Shuichi Sakai, Hiroki Konaka, Munenori Maeda, Takashi Tomokiyo, Jörg Nolte, Hiroshi Matsuoka, Kazuaki Okamoto, Hideo Hirono |
JSSPP | 8 |