VLDB 2026 Research / reviewers in the wild / expert
Carsten Trinitis
dblp:45/1242
· DBLP profile ↗
29ranked-venue papers
3as first author
14since 2021 · last 2026
0000-0002-6750-3652ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 15 · 2 first-author · 7 since 2021Software engineering, systems software and programming languages · 4 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Security and privacy · 2 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Building AI Hardware Expertise: Edge AI Curriculum Design and Implementation in German UniversitiesabstractEdge artificial intelligence (AI) redistributes AI computation from distant cloud to local processors for real-time processing and enhanced privacy. This fundamental shift underscores a critical gap in current university curricula, which predominantly focus on AI fundamentals and algorithms while often neglecting essential AI hardware topics. To address this deficiency, this paper presents Edge AI, a postgraduate curriculum co-designed by two universities in Germany. Guided by the Dagstuhl triangle, the curriculum is designed to comprehensively cover technical, sociocultural, and application perspectives. Courses are developed using the Four-Component Instructional Design model to encourage action-oriented skill development, with a Learning Management System template available to assist in the design of individual courses. Selected practical courses are formulated as self-managed projects and inverted classrooms, enabling students to learn at their own pace with just-in-time guidance. All curriculum materials are accessible online and maintained by the Open Science Framework to enhance collaboration across institutions and promote applicability in diverse domains. Evaluation results from 176 students over two years (2023-2025) demonstrate universal satisfaction across various curriculum components. Ann-Marie Gursch, Xuanshu Luo, Lilian Hasse, Carsten Trinitis, Ulrike Lucke, Martin Werner 0001, Milos Krstic |
AAAI | 4 |
| 2026 | POSTER: Towards a RISC-V-based SmartNIC Architecture on FPGAabstractModern datacenters and high-performance computing systems increasingly rely on SmartNICs to reduce host overhead and improve the security and efficiency of network data movement. This work-in-progress poster outlines our research path for a RISC-V-centric SmartNIC with dedicated context and objectives. The extended RISC-V core acts as the central controlling unit for the key SmartNIC components designed in RTL, including a security-oriented (Physically Unclonable Function) PUF unit for authentication, a flexible packet parser of arbitrary protocols enhanced by RISC-V Vector (RVV) extension, and a Remote Direct Memory Access (RDMA) engine targeted at high-performance computing clusters. We apply an end-to-end simulator to model these components at the behavioral level without requiring full hardware setups. Finally, this project maps the resulting design to an FPGA-based SmartNIC architecture. Kun Qin, Aswathy Nedumpalli Sankaranarayanan, Taiki Okano, Martin Schulz 0001, Carsten Trinitis |
CF | 5 |
| 2026 | Multi-Partner Project: Advancing European Semiconductor and Chiplet Innovation Through the Bavarian Chip Design CenterabstractEurope’s semiconductor industry relies heavily on Asian and US manufacturers. The EU Chips Act seeks to strengthen Europe’s capabilities across the semiconductor value chain. Aligned with this goal, the Bavarian Chip Design Center (BCDC) supports local chip design, manufacturing, and talent development, with a focus on RISC-V computing and heterogeneous integration. Within BCDC, the Technical University of Munich and Fraunhofer are developing a chiplet-based architecture optimized for low-power edge AI. The system integrates two chiplets, combining a security-enhanced RISC-V core and AI accelerators, connected via a chiplet-optimized serial interface that supports encrypted data. The chiplets are mounted on a custom interposer with low-capacitance wires for efficient data transmission. System-and component-level development is currently ongoing, with a tapeout in 22 nm FD-SOI planned for 2027. The overall goal is to deliver a proof of concept for a small-scale energy-efficient chiplet system that demonstrates Bavaria’s and Europe’s capability to drive innovation in novel chip design fields. Hussam Amrouch, Jehaan Joseph, Michael Schirmer, Johannes Geier, Ulf Schlichtmann, Michael Meidinger, Thomas Wild, Andreas Herkersdorf, Jens Nöpel, Georg Sigl, Carsten Trinitis, Aswathy Nedumpalli Sankaranarayanan, Martin Schulz 0001, Andreas Korb, Konrad Hohentanner |
DATE | 12 |
| 2025 | FERIVer: An FPGA-assisted Emulated Framework for RTL Verification of RISC-V ProcessorsabstractProcessor design and verification require a synergistic approach that combines instruction-level functional simulations with precise hardware emulations.The trade-off between speed and accuracy in the instruction set simulation poses a significant challenge to the efficiency of processor verification.By tapping the potentials of Field Programmable Gate Arrays (FPGAs), we propose an FPGA-assisted System-on-Chip (SoC) platform that facilitates cross-verification by the embedded CPU and the synthesized hardware in the programmable fabrics.This method accelerates the verification of the RISC-V Instruction Set Architecture (ISA) processor at a speed of 5 million instructions per second (MIPS), which is 150x faster than the vendor-specific tool (Xilinx XSim) and a 35x boost to the stateof-the-art open-source verification setup (Verilator).With less than 7% hardware occupation on Zynq 7000 FPGA, the proposed framework enables flexible verification with high time and cost efficiency for exploring RISC-V instruction set architectures. Kun Qin, Xiaorang Guo, Martin Schulz 0001, Carsten Trinitis |
CF | 4 |
| 2025 | Advancing user-space networking for DDS message-oriented middleware: Further extensionsabstractDue to the flexibility it offers, publish–subscribe messaging middleware is a popular choice in Industrial IoT (IIoT) applications. The Data Distribution Service (DDS) is a widely used industry standard for these systems with a focus on versatility and extensibility, implemented by multiple vendors and present in myriad deployments across industries like aerospace, healthcare and industrial automation. However, many IoT scenarios require real-time capabilities for deployments with rigid timing, reliability and resource constraints, while publish–subscribe mechanisms currently rely on components that are not strictly real-time capable, such as the Linux networking stack, making it hard to provide robust performance guarantees without large safety margins. In order to make publish–subscribe approaches viable and efficient also in such real-time scenarios, we introduce user-space DDS networking transport extensions, allowing us to fast-track the communication hot path by bypassing the Linux kernel. For this purpose, we extend the best-performing vendor implementation from a previous study, CycloneDDS, to include modules for two widespread user-space networking technologies, the Data Plane Development Kit (DPDK) and the eXpress Data Path (XDP). Building on this, we additionally offer two more extensions to the second most performant implementation FastDDS, also based on DPDK and XDP, and realize novel optimizations not present in the original extension implementations. We evaluate each extension’s performance benefits against four existing DDS implementations (OpenDDS, RTI Connext, FastDDS and CycloneDDS). The DPDK-based and XDP-based extensions offer a performance benefit of 31%–38% and 18%–22% reduced mean latency, respectively, as well as an increase in bandwidth and sample rate throughput of at least 160%, while reducing the latency bound by at least 93%, demonstrating the performance and dependability advantages of circumventing the kernel for real-time communications. Vincent Bode, Carsten Trinitis, Martin Schulz 0001, David Buettner, Tobias Preclik |
Pervasive Mob. Comput. | 2 |
| 2024 | Dataset Distillation by Automatic Training Trajectories
Dai Liu, Jindong Gu, Hu Cao, Carsten Trinitis, Martin Schulz 0001 |
ECCV (87) | 4 |
| 2024 | Adopting User-Space Networking for DDS Message-Oriented MiddlewareabstractDue to the flexibility it offers, publish-subscribe messaging middleware is a popular choice in Industrial IoT (IIoT) applications. The Data Distribution Service (DDS) is a widely used industry standard for these systems with a focus on versatility and extensibility, implemented by multiple vendors and present in myriad deployments across industries like aerospace, healthcare and industrial automation. However, many IoT scenarios require real-time capabilities for deployments with rigid timing, reliability and resource constraints, while publish-subscribe mechanisms currently rely on components that are not strictly real-time capable, such as the Linux networking stack, making it hard to provide robust performance guarantees without large safety margins. In order to make publish-subscribe approaches viable and efficient also in such real-time scenarios, we introduce userspace DDS networking transport extensions, allowing us to fasttrack the communication hot path by bypassing the Linux kernel. For this purpose, we extend the best-performing vendor implementation from a previous study, CycloneDDS, to include modules for two widespread user-space networking technologies, the Data Plane Development Kit (DPDK) and the eXpress Data Path (XDP), and we evaluate their performance benefits against four existing DDS implementations (OpenDDS, RTI Connext, FastDDS and CycloneDDS). The CycloneDDS-DPDK and CycloneDDS-XDP extensions offer a performance benefit of 31% and 18% reduced mean latency, respectively, as well as an increase in bandwidth and sample rate throughput of up to 59%, while reducing the latency bound by at least 94%, demonstrating the performance and dependability advantages of circumventing the kernel for real-time communications. Vincent Bode, Carsten Trinitis, Martin Schulz 0001, David Buettner, Tobias Preclik |
PerCom | 2 |
| 2023 | Machine Learning Application BenchmarkabstractThis paper presents the MLAB project, a research and development activity funded by ESA General Support Technology Programme under the lead of Airbus Defence and Space GmbH, with the goal of developing a machine learning application benchmark for space applications. First, the need for a benchmark dedicated to machine learning applications in spacecraft is explained, and examples of applications are described including their design challenges. Then the benchmark design is presented, including the rules of the metrics, guidelines and scenarios for references. These scenarios include a description of the reference workloads that have been selected during the activity as representative for spacecraft applications. Lastly, the submission concept is introduced. Michael Petry, Max Ghiglione, Amir Raoofy, Gabriel Dax, Gianluca Furano, Martin Werner 0001, Carsten Trinitis, Martin Langer |
CF | 8 |
| 2023 | Federated Learning via Decentralized Dataset Distillation in Resource-Constrained Edge EnvironmentsabstractIn federated learning, all networked clients contribute to the model training cooperatively. However, with model sizes increasing, even sharing the trained partial models often leads to severe communication bottlenecks in underlying networks, especially when communicated iteratively. In this paper, we introduce a federated learning framework FedD3 requiring only one-shot communication by integrating dataset distillation instances. Instead of sharing model updates in other federated learning approaches, FedD3 allows the connected clients to distill the local datasets independently, and then aggregates those decentralized distilled datasets (e.g. a few unrecognizable images) from networks for model training. Our experimental results show that FedD3 significantly outperforms other federated learning frameworks in terms of needed communication volumes, while it provides the additional benefit to be able to balance the trade-off between accuracy and communication cost, depending on usage scenario or target dataset. For instance, for training an AlexNet model on CIFAR-10 with 10 clients under non-independent and identically distributed (Non-IID) setting, FedD3 can either increase the accuracy by over 71% with a similar communication volume, or save 98% of communication volume, while reaching the same accuracy, compared to other one-shot federated learning approaches. Rui Song 0007, Dai Liu, Dave Zhenyu Chen, Andreas Festag, Carsten Trinitis, Martin Schulz 0001, Alois C. Knoll |
IJCNN | 5 |
| 2023 | Systematic Analysis of DDS ImplementationsabstractPublish-subscribe messaging is a popular communication paradigm in the (Industrial) Internet of Things, and the Data Distribution Service (DDS) is a well known standard for pub-sub communication middleware. Many vendor implementations of DDS exist, leaving users with the need to choose according to project and performance requirements. However, the wide range of parameters in DDS implementations not covered in the standard specification make this selection difficult and time-consuming. We present DDS-Perf, a novel and versatile cross-vendor benchmarking tool for performance analysis, and use it to provide data from studies on 4 popular DDS implementations (OpenDDS, RTI Connext, FastDDS and CycloneDDS) across a wide range of experimental setups. DDS-Perf allows us to provide a consistent methodology across all vendors, increasing fairness and comparability. Overall, we find that RTI Connext achieves the best all-round performance (exhibiting the best bandwidth and peak sample rate), while FastDDS (best end-to-end latency) and CycloneDDS also show promising results. Vincent Bode, David Buettner, Tobias Preclik, Carsten Trinitis, Martin Schulz 0001 |
Middleware | 4 |
| 2023 | DDS Implementations as Real-Time Middleware - A Systematic EvaluationabstractPublish-subscribe messaging has seen increased adoption in the context of timing critical applications, with multiple frameworks integrating publish-subscribe middleware into their ecosystem. The Data Distribution Service (DDS) is a standard for pub-sub systems that has gained traction, through e.g. the adoption in ROS2, with multiple vendors distributing their implementations of the standard. However, while DDS is being used in a real-time context, the examined implementations are not strictly real-time capable and therefore cannot provide hard timing guarantees to the application. Still, users are looking to take advantage of the flexibility offered by DDS in close-to real-time use cases, raising the question of how well DDS implementations can provide soft real-time reliability assurances. We use DDS-Perf, a novel cross-vendor benchmarking tool for impartial performance analysis of DDS implementations, to examine how reliably four vendors (OpenDDS, RTI Connext, Fast-DDS and CycloneDDS) can deliver real-time-like performance under different scenarios. From a typical out-of-the-box setup, we offer a guide to users for tuning performance/reliability and examine problems users might encounter trying to satisfy real-time constraints. The vendor implementations are tested against a range of experiments to evaluate operating performance under favorable and adverse conditions. Overall, we find that OpenDDS has the worst performance out-of-the-box and after calibration, while FastDDS and CycloneDDS offer the best performance, which is comparable to user-space networking technologies in the average case but with a much higher worst-case latency bound. Vincent Bode, Carsten Trinitis, Martin Schulz 0001, David Buettner, Tobias Preclik |
RTCSA | 2 |
| 2022 | Benchmarking and feasibility aspects of machine learning in space systemsabstractCompute in space, e.g., in miniaturized satellites, requires dealing with special physical and boundary constraints, including the limited energy budget. These constraints impose strict operational conditions on the on-board data processing system and its capability in dealing with sophisticated workloads suchlike Machine Learning (ML). In the meantime, the breakthroughs in ML based on Deep Neural Networks (DNNs) in the last decade promise innovative solutions to expand the functional capabilities of on-board data processing and to drive the space industry forward. Therefore, due to the aforementioned special requirements, performance- and power-efficient, and novel solutions and architectures for deploying ML via, e.g., FPGA-enabled SoC, particularly Commercial-Off-The-Shelf (COTS) solutions, are gaining significant interest in the space industry. Therefore it is essential to conduct extensive benchmarking and feasibility and efficiency analyses in different aspects: such analyses would require the investigation of options for programming and deployment as well as the investigation of various real-world models and datasets. To this end, a research and development activity is funded by the European Space Agency (ESA) General Support Technology Programme and is led by Airbus Defence and Space GmbH with the goal of developing an ML Application Benchmark (MLAB) that covers benchmarking aspects mentioned above. Amir Raoofy, Gabriel Dax, Vittorio Serra, Max Ghiglione, Martin Werner 0001, Carsten Trinitis |
CF | 6 |
| 2022 | MicroPython as a satellite control languageabstractDifferent from terrestrial applications, most imaging nanosatellites are relying on simplistic command sequences for on-board controls. Combined with the unpredictable nature of flight operations, this can result in tedious and work-intensive operations, as unforeseen events might mean that the commands do not fit the needs of the operators anymore. The restricted communication windows meanwhile require a high level of automation whilst keeping the size of uplinked sequences minimal. Therefore we require a dynamic control language that is also able to operate within the resource limitations given by a nanosatellite. In an effort to combine all these requirements, we chose to implement MicroPython as the control language for our satellite payload. This extended abstract shall introduce the architecture and concepts of our implementation. Together with our presentation at CompSpace '22 it shall serve as a basis for discussion of using MicroPython as a payload control language on nanosatellites. Hanna Vivien Schwarzwald, Sebastian Würl, Martin Langer, Carsten Trinitis |
CF | 4 |
| 2021 | Living on the Edge: Efficient Handling of Large Scale Sensor DataabstractReal-time sensor monitoring is critical in many industrial applications and is, e.g., used to model and predict operating conditions to optimize operations as well as to prevent damage in machinery and systems. In many cases, this data is generated by a myriad of sensors and stored or transmitted for post-processing by data analysts. Handling this data near its origin-on the edge-imposes significant challenges for storage and compression: it is necessary to store it in a format that is suitable for large data analytics algorithms, which in most cases means columnar storage. Furthermore, to provide efficient storage and transmission of such sensor data, it must be compressed efficiently. However, existing solutions do not address these challenges sufficiently. In this work, we present a holistic approach for fast streaming of large scale sensor data directly into columnar storage and integrate it with a proven compression scheme. Our approach uses a pipelined scheme for streaming and transposing the data layout, combined with a byte-level transformation of data representation and compression, which we evaluate in comprehensive experiments. As a result, our approach enables transformation of large scale sensor data streams into an efficient, analytics-friendly format already at the sensor site, i.e., on the edge, at data ingestion time. By implementing our optimized approach in the open and widely used columnar storage format Apache Parquet, which we already partly upstreamed, we ensure its accessibility to the community. Roman Karlstetter, Amir Raoofy, Martin Radev, Carsten Trinitis, Jakob Hermann, Martin Schulz 0001 |
CCGRID | 4 |
| 2019 | The german informatics society's new ethical guidelines: POSTERabstractOn June 28, 2018, he board of directors of the German Informatics Society (GI) adopted new ethical guidelines. Carsten Trinitis, Christina Class, Stefan Ullrich 0001 |
CF | 1 |
| 2016 | Automatic Co-scheduling Based on Main Memory Bandwidth Usage
Jens Breitbart, Josef Weidendorfer, Carsten Trinitis |
JSSPP | 3 |
| 2015 | Cache-oblivious matrix algorithms in the age of multicores and many coresabstractSummary This article highlights the issue of upcoming wider single‐instruction, multiple‐data units as well as steadily increasing core counts on contemporary and future processor architectures. We present the recent port to and latest results of cache‐oblivious algorithms and implementations of our TifaMMy code on four architectures: SGI's UltraViolet distributed shared‐memory machine, Intel's latest x86 architecture code‐named Sandy Bridge, AMD's new Bulldozer architecture, and Intel's future Many Integrated Core architecture. TifaMMy's matrix multiplication and LU decomposition routines have been adapted and tuned with regard to these architectures. Results are discussed and compared with vendors’ architecture‐specific and optimized libraries, Math Kernel Library and AMD Core Math Library, for both a standard C++ version with vectorization compiler switches and TifaMMy's highly optimized vector intrinsics version. We provide insights into architectural properties and comment on the feasibility of heterogeneous cores and accelerators, namely graphics processing units. Besides bare‐metal performance, the test platforms’ ease of use is analyzed in detail, and the portability of our approach to new and upcoming silicon is discussed with regard to required effort on code change abstraction levels. As a result, we demonstrate that because of its generic structure in terms of memory organization, TifaMMy executes with equally efficient performance on all four architectures as it automatically adapts itself to architectural parameters without losing performance against the Math Kernel Library and AMD Core Math Library, underlining its generic and cache‐oblivious properties, as the porting effort was relatively low compared with that in other implementations.Copyright © 2012 John Wiley & Sons, Ltd. Alexander Heinecke, Carsten Trinitis |
Concurr. Comput. Pract. Exp. | 2 |
| 2015 | Recent developments in high-performance computing and simulation: distributed systems, architectures, algorithms, and applicationsabstractHigh-performance computing (HPC) helps address real-world problems by providing strong environments and support to run data and computational intensive algorithms, complex numerical simulations, and parallel scientific codes. However, the increasing needs to tackle high-resolution exascale, complex systems, and multi-scale data-intensive sciences require dealing with new computing platforms with millions of cores and more complex software and hardware solutions. In such a landscape, key requirements on concurrency, energy, storage, I/O, resiliency, networking, and so on need to be re-examined and investigated at different levels. This special issue contains 10 papers representing recent advances in the area of high-performance computing and simulation. These extended papers were carefully selected from the proceedings of the 2011 International Conference on High Performance Computing and Simulation (HPCS 2011), which was held in Istanbul, Turkey, July 4–8, 2011. The invited papers in this special issue represent fully refereed augmented works originally presented at the conference, which cover some of the main contemporary topics and challenges in various areas of HPC and simulation. The International High Performance Computing and Simulation (HPCS) Conference Series is meant to address, explore, and exchange information on the state-of-the-art in high-performance and large-scale computing systems, their use in modeling and simulation, their design, performance and utilization, and their applications and impact. Typically, the conference includes invited presentations by experts from academia, industry, and government laboratories and institutions as well as contributed paper presentations describing refereed original work on the current state of research in the areas in the preceding text and other related ones such as services computing, cloud and grid computing, and mobile computing. Annually, participation is extended to researchers, designers, educators, and interested parties in all HPCS disciplines and specialties to partake in various functions and contribute to activities such as tutorials, demos, exhibits, posters, panels, and doctoral dissertation colloquia. Over the years, the conference invited some of the top experts and innovators in the field as keynote speakers or for plenary talks and tutorial sessions. Along with the main track, several symposia, workshops, and special sessions are organized every year in conjunction with this meeting. Since 2009, the conference proceedings have been published by IEEE and included in the IEEE Xplore Digital Library 1-3 and indexed accordingly in several major indexing services 4. Prior to that, proceedings were published and indexed by the European Council for Modelling and Simulation (ECMS) (e.g., 5, 6). The conference continues to experience a healthy growth and improved quality contributions. The International HPCS Conference Series started in 2003 as the High Performance & Large Scale Computing (HP&LSC) Track in conjunction with the European Council for Modelling and Simulation 2003 Conference and was initially held in Nottingham, UK. With the first event being a great success, follow-up meetings were held in Magdeburg, Germany (2004); Riga, Latvia (2005); Bonn, Germany (2006); Prague, Czech Republic (2007); Nicosia, Cyprus (2008); Leipzig, Germany (2009); and Caen, France (2010). The Ninth HPCS Conference was held in Istanbul, Turkey, in the summer of 2011 on which this special issue is based. The conference has been organized in technical cooperation with major professional organizations as such Association for Computing Machinery (ACM), The Institute of Electrical and Electronics Engineers (IEEE), and International Federation for Information Processing (IFIP) as well as academic institutions and research centers, and many regional organizations. The first conference-wide special issue was organized based on the HPCS 2009 meeting in Wiley's Concurrency and Computation: Practice and Experience Journal. The set of the final papers published is available from Wiley and also online 7. The second special issue was organized the following year based on the HPCS 2010 meeting in Caen, France. The set of the final papers published is available from Wiley and also online 8. As such, the current conference-wide special issue is the third one, and it is based on selected extended papers from the Ninth HPCS Conference (HPCS 2011), which was held in Istanbul, Turkey, during the days July 4– 8, 2011. In addition to these, tracks, workshops, and special sessions also organized special issues on their respective areas, some of which have been published already 9-12. This conference-wide third special issue consists of selected papers that were compiled from the conference's main track as well as its adjoining workshops and special sessions. The chosen papers embrace several of the contemporary state-of-the-art research issues, such as many-core computing, heterogeneous architectures, cloud computing, self-aware systems, and real-time requirements along with an array of applications with HPC provisions. Authors of 31 papers from the conference proceedings were invited to submit an extended and updated version of their original paper to this special issue. The selection was made by the conference international program committee and workshops and special sessions organizers. In the initial round, we received 16 affirmative responses/proposals, which were reviewed for approval. An international reviewing committee of about 50 experts in various subjects of HPCS was formed. The committee members came from 15 different countries. Each paper was assigned to at least five reviewers. In the first round, 14 manuscripts were submitted for the reviewing process, and 11 papers passed with minor or major revisions required, while three were rejected. The 11 revised manuscripts were revised and again submitted for round two of the reviewing process. After carefully taking the reviewers' remarks and recommendations into account, the outcome this time was 10 papers were passed with minor revisions or accepted, and one was rejected. In round three, the 10 papers were reviewed again, and some minor improvements were requested for most papers. Only two manuscripts required a fourth round of reviews. At the end of this elaborate and thorough review process, we have the 10 manuscripts that you find in this special issue. With the help of our reviewers and a moderate turnaround time in each round, we managed to receive the reviews on time and forward them to the authors. The authors met the strict deadlines we set forth for each cycle and submitted their revised manuscripts in a timely manner as well. This special issue of HPCS papers comprises of 10 contributions from 31 authors from seven different countries, namely, Germany, Italy, Portugal, Romania, Spain, the UK, and the USA, covering research ranging from graph-based approaches, to tuning applications, to state-of-the-art processor architectures, to cloud computing. Basically, the papers in this special issue can be categorized into five main subjects: system- and hardware-oriented papers, out of which category, three papers have been selected for publication here; grid/cloud computing related papers, out of which category, two manuscripts have been selected for publication here; papers dealing with real-time requirements, out of which category, one paper has been selected for publication here; papers on autonomous-reflexive systems, out of which category, two manuscripts have been selected for publication here; and applications-oriented papers, out of which category, two papers have been selected for publication in this special issue. We summarize these next. ‘Finding Near-Perfect Parameters for Hardware and Code Optimizations by Automatic Multi-Objective Design Space Explorations’ by Ralf Jahr, Horia Calborean, Lucian Vintan, and Theo Ungerer 13 introduces FADSE, a design space exploration tool that automatically finds nearly optimal processor design configurations for a given code. Taking into account that performance is no longer the only objective subject to optimization (others comprise power consumption, area, etc.), FADSE tries to find an optimum architecture across multiple objective functions. As an example, the Grid ALU Processor (GAP) and its post-link optimizer GAPtimize are used to demonstrate the feasibility of the approach. ‘Cache-oblivious Matrix Algorithms in the Age of Multi- and Many-Cores’ by Alexander Heinecke and Carsten Trinitis 14 highlights the issue of increasing vector unit width that goes along with increasing core counts on x86 processor architectures. To demonstrate this, a cache-oblivious numerical code has been ported to and optimized on four contemporary x86 architectures representing vector unit widths from 128 to 512 bits. The article discusses the obtained performance results and compares them with the vendors' architecture specific and optimized libraries Math Kernel Library (MKL) and AMD Core Math Library (ACML). A special emphasis is put on providing insights into architectural properties of state-of-the-art processor and accelerator architectures. ‘New System Software for Parallel Programming Models on the Intel SCC Many-core Processor’ by Carsten Clauss, Stefan Lankes, Pablo Reble, and Thomas Bemmerl 15 gives a detailed report on the authors' experiences with implementing parallel programming libraries and tools for the 48-core Intel Single Cloud Chip (SCC) processor. SCC is a prototype of a many-core processor comprising noncoherent memory-coupled cores, a so-called cluster-on-chip architecture. The programming library developed by the authors reflects an SCC-customized Message Passing Interface (MPI) library called SCC-MPICH for distributed memory parallel programming and a shared virtual memory system called MetalSVM for the thread programming. In case of the SCC chip, both approaches are evaluated, and it is shown how these can be optimized for such a novel cluster-on-chip architecture. ‘Cost Optimization of Virtual Infrastructures in Dynamic Multi-Cloud Scenarios’ by Jose Luis Lucas Simarro, Rafael Moreno-Vozmediano, Ruben S. Montero, and Ignacio M. Llorente 16 presents a so-called cloud broker architecture: an architecture that is responsible for deploying virtual resources (virtualized servers) across compute clouds. By taking into account migration overhead costs in a dynamic cloud scenario, several use cases are investigated, demonstrating that using brokering mechanisms in dynamic deployments shows clear advantages over static deployments in cloud environments. From the users' point of view, multiple factors such as pricing schemes, types of instance, or value-added features need to be taken into account, which is why cloud brokering comes into play. ‘Interoperating Grid Infrastructures with the GridWay Metascheduler’ by Ismael Marin Carrion, Eduardo Huedo, and Ignacio M. Llorente 17 describes GridWay, a metascheduler for sharing compute resources within common grid middleware, which was developed by the authors. Latest features comprise enhancements with regard to interoperability and interoperation, which is achieved by introducing a modular architecture design. Two new execution drivers and a new remote interface have been added to GridWay, which is described in detail in the paper. ‘Improved Real-Time Scheduling for Periodic Tasks on Multiprocessors’ by Prapaporn Rattanatamrong and Jose A. B. Fortes 18 presents a novel algorithm for scheduling applications with real-time requirements to supercomputers. Methods to ensure that all resources can be optimally utilized are provided in the paper. This is demonstrated by an application dealing with a human brain-machine interface, a simulation of a prosthetic limb's movement according to activities of input signals. ‘Towards Self-Caring IT Systems: A Study of Performance Penalties under Faults’ by Selvi Kadirvel and José A. B. Fortes 19 is from the area of fault tolerance and utilizes virtualization techniques: taking MapReduce frameworks as an example, it is shown that the performance penalty imposed by fault tolerance mechanisms can not be neglected. Hence, for the open source MapReduce framework Hadoop, this execution time penalty is evaluated by using a simulator. Further investigations are carried out in a virtual environment with varying characteristics regarding hardware, application, data set, and types of fault. The obtained parameter studies show that penalties can be significantly reduced through dynamic resource scaling. ‘AOI-Cast in Distributed Virtual Environments: An Approach based on Delay Tolerant Reverse Compass Routing’ by Laura Ricci, Luca Genovali, Emanuele Carlini, and Massimo Coppola 20 deals with a novel area of interest (AOI)-cast algorithm for distributed environments such as massively multiplayer online games (MMOGs). Through exploiting the mathematical properties of Delaunay Triangulations, a spanning tree supporting event notification within the area of interest can be built. This tree is computed by reverse compass routing. The efficiency of this novel approach is demonstrated through a set of simulations with both artificial and real data from an MMOG. ‘Implementation and Performance Analysis of Efficient Index Structures for DNA Search Algorithms in Parallel Platforms’ by Gustavo Encarnacão, Nuno Sebastião, and Nuno Roma 21 is a paper from the area of bioinformatics. In DNA sequence alignment, it is of crucial importance to choose an appropriate local alignment algorithm in order to achieve reasonable performance. The authors present an analysis of three highly optimized implementations of index-based search algorithms, namely, suffix-trees, suffix-arrays, and hash tables of q-mers. For all three, a performance comparison is carried out on CPU-based and Graphics Processing Unit (GPU) based architectures. On both architectures, it is shown that suffix-trees and suffix-arrays perform significantly better than hash tables of q-mers. ‘Parallel Multigrid on Hierarchical Hybrid Grids: A Performance Study on Current HPC Clusters’ by Björn Gmeiner, Harald Köstler, Markus Stürmer, and Ulrich Rüde 22 investigates the performance of a geometric multigrid solver on up-to-date high-performance computing cluster installations, namely a BlueGene/P cluster run by the Julich Supercompouting center and an Intel Xeon 5650 cluster run by the Erlangen regional computing center (RRZE). The geometric multigrid solver executes inside a software package called hierarchical hybrid grids (HHGs). HHG is a package based on unstructured tetrahedral finite elements. The obtained performance is evaluated and compared with that obtained when using a standard multigrid solver so that an estimate can be given as to whether it is worth using numerical packages like HHG. It is our hope that the collection of manuscripts in this special issue will make a significant contribution to the HPC systems and modeling and simulation fields and their future developments. The guest editors of this special issue on High-performance Computing Systems would like to thank all authors, the special issue reviewing committee, the CPE EIC, Prof. G. C. Fox, and the editorial staff of Wiley for their contributions, efforts, and support in making this special issue possible. It would not have been possible without their support and guidance. The special issue Reviewing Committee members are Giovanni Aloisio (Italy), Andres Avila (Chile), Bruno Bachelet (France), Liz Bacon (UK), Mostafa Bamha (France), Francoise Baude (France), Milan Bradonjic (USA), Ivona Brandic (Austria), Mathieu Chapelle (France), Camille Coti (France), Alfredo Cuzzocrea (Italy), Laurent d'Orazio (France), Luciano Antonio Digiampietri (Brazil), Daniel Etiemble (France), Joel Falcou (France), Bernhard Fechner (Germany), Cecile Germain (France), Alain Giulieri (France), David Gregg (Ireland), Mark Hedges (UK), Alexander Heinecke (Germany), Gonzalo Hernandez (Chile), David Hill (France), Neil Chue Hong (UK), Udo Hönig (Germany), Zhihi Huang (New Zealand), Eric Innocenti (France), Hai Jin (China), David Kaeli (USA), Al Kellie (USA), Harald Köstler (Germany), Dieter Kranzlmüller (Germany), Sébastien Limet (France), Frederic Loulergue (France), Olivier Marin (France), Emmanuel Melin (France), Lizandro Muzy (France), Mariusz Nowostawski (New Zealand), Domenico Potena (Italy), T. K. Prasad (USA), Desh Ranjan (India), Mukesh Singhal (USA), Anna Squicciarini (USA), Domenico Talia (Italy), Lorenzo Verdoscia (Italy), Timothy J. Williams (USA), Chao Tung Yang (Taiwan), Vesna Zeljkovic (China), and Ji Zhang (Australia). Waleed W. Smari, Sandro Fiore, Carsten Trinitis |
Concurr. Comput. Pract. Exp. | 3 |
| 2012 | Exploiting State-of-the-Art x86 Architectures in Scientific ComputingabstractIn recent years, general purpose ×86 architectures have undergone significant modifications towards high performance computing capabilities. Lately, technologies like wider vector units or Fused Multiply-Add (FMA) instruction, which were mainly known from GPU arcitectures, have been introduced. In this paper, we examine the performance of current ×86 architectures, namely Intel Sandy Bridge and AMD Bulldozer, for four different parallel workloads with different properties. These properties comprise optimally cache-blocked algorithms as well as adaptive grid structures resulting in memory latency and bandwidth bound executions. The achieved performance on both architectures is very promising, and, if extrapolated towards upcoming server silicon, can be regarded as on par with current high-end GPU based accelerators. Alexander Heinecke, Thomas Auckenthaler, Carsten Trinitis |
ISPDC | 3 |
| 2012 | Wait-Free Message Passing Protocol for Non-coherent Shared Memory Architectures
Isaías A. Comprés Ureña, Michael Gerndt, Carsten Trinitis |
EuroMPI | 3 |
| 2011 | Sparse matrix operations on several multi-core architectures
Carsten Trinitis, Tilman Küstner, Josef Weidendorfer, Jasmin Smajic |
J. Supercomput. | 1 |
| 2006 | Implementation of a DSM-System on Top of InfiniBandabstractImplementations of algorithms based on shared memory are widely used in the area of high performance computing. The concept of computation clusters, however, does not imply the availability of a globally shared memory across nodes in that cluster. Therefore, several concepts for an emulation of a globally shared memory, called software distributed shared memory (SDSM), exist. With the introduction of the InfiniBand interconnect technology, interprocess communication methods like direct access to other nodes' memory become available. This paper presents an approach to simplify and optimize an existing SDSM implementation by taking advantage of InfiniBand specific communication patterns. Hubert Eichner, Carsten Trinitis, Tobias Klug |
PDP | 2 |
| 2005 | CPU-independent Assembler in an FPGAabstractWe describe a system which enables FPGAs to generate machine code for various CPUs, similar to a conventional assembler. Such conversion from intermediate code to a CPU's native code can be used as the last step in just-in-time compilation for virtual machines like the Java Virtual Machine. The translation system itself and the FPGA logic are independent of the actual target CPU and can be used with both CISC and RISC CPUs. Due to an extended table lookup, the resulting code is very efficient and gains from pre-calculation of selected constants in the FPGA assembler. Georg Acher, Rainer Buchty, Carsten Trinitis |
FPL | 3 |
| 2004 | ViSMI: Software Distributed Shared Memory for InfiniBand ClustersabstractThis work describes ViSMI, a software distributed shared memory system for cluster systems connected via InfiniBand. ViSMI implements a kind of home-based lazy release consistency protocol, which uses a multiple-writer coherence scheme to alleviate the traffic introduced by false sharing. For further performance gain, InfiniBand features and optimized page invalidation mechanisms are applied in order to reduce synchronization overhead. First experimental results show that ViSMI introduces good performance comparable to similar software DSMs. Christian Osendorfer, Jie Tao 0001, Carsten Trinitis, Martin Mairandres |
NCA | 3 |
| 2003 | CAD Grid: Corporate-Wide Resource Sharing for Parameter Studies
Ed Wheelhouse, Carsten Trinitis, Martin Schulz 0001 |
Euro-Par | 2 |
| 2003 | SMiLE: an integrated, multi-paradigm software infrastructure for SCI-basedclusters
Martin Schulz 0001, Jie Tao 0001, Carsten Trinitis, Wolfgang Karl |
Future Gener. Comput. Syst. | 3 |
| 2002 | SMiLE: An Integrated, Multi-Paradigm Software Infrastructure for SCI-Based ClustersabstractThe availability of a comprehensive software infrastructure is essential for the success a parallel architecture. In order to allow for the greatest possible flexibility, an infrastructure has to be designed in an integrated, easy-to-use manner and with the support of multiple programming paradigms and models to address a wide base of codes. SMiLE provides such an infrastructure for SCI (Scalable Coherent Interface) based clusters. It includes support for both a large range of message passing libraries as well as for almost arbitrary shared memory programming models. In addition, SMiLE also contains initial work on appropriate tool sets for performance optimizations. The complete infrastructure is implemented in way that is as closely relate d to the underlying hardware and is therefore capable of exploiting the benefits of the underlying network fabric and offering them to the user without significant overheads. Martin Schulz 0001, Jie Tao 0001, Carsten Trinitis, Wolfgang Karl |
CCGRID | 3 |
| 2001 | OpenSESAME: An Intuitive Dependability Modeling Environment Supporting Inter-Component DependenciesabstractThe paper proposes a novel modeling method for the evaluation of dependability measures of highly available systems. The proposed method, which has been implemented in the tool OpenSESAME (Simple but Extensive Structured Availability Modeling Environment), combines the advantages of Boolean methods and state space based methods. The tool supports the modeler with a set of well-defined, structured, intuitive input diagrams and tables, which are automatically transformed into GSPNs (Generalized Stochastic Petri Nets) for evaluation. To show the usefulness of the proposed method, it is applied to a model of a typical CompactPCI-based high availability system as can be found in the telecommunications area. Max Walter, Carsten Trinitis, Wolfgang Karl |
PRDC | 2 |
| 2000 | Electrical phenomena during Hot Swap eventsabstractThe exchange of a computer system's components during operation can be accomplished by the so called Hot Swap technology. This technology makes it possible to continuously run a computer system without the necessity of a shutdown for maintenance purposes, e.g. upgrading of a network adapter. Thus the overall uptime of a system can be drastically increased. The Hot Swap capability has been integrated into the so called CompactPCI technology. This paper summarizes the investigations that have been carried out with respect to the electrical behavior during Hot Swap events. Several simulations were performed with the H-SPICE program by AMP, Harrisburg, PA. After an introduction to live insertion phenomena in general, the CompactPCI specific simulations are described and recommendations for designing a Hot Swap capable system from an electrical point of view are given. Carsten Trinitis, Wolfgang Karl, Markus Leberecht |
PRDC | 1 |