EDBT 2026 Demo / reviewers in the wild / expert
Dietmar Fey
dblp:09/4727
· DBLP profile ↗
46ranked-venue papers
7as first author
8since 2021 · last 2026
0000-0002-6077-4732ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 23 · 4 first-author · 5 since 2021Artificial intelligence and machine learning · 7 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 7 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 3Databases, data management, data science and information retrieval · 1Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Silicon-Based Evaluation of CSD Arithmetic in a RISC-V Processor Using Open-Source ASIC Flow
Farhad Ebrahimi-Azandaryani, Michael Kupfer, Dietmar Fey |
ISCAS | 3 |
| 2025 | Smart Sensing of Multi-bit Resistive Memory using a Single ReferenceabstractEmerging Non-Volatile Memories (NVMs) are increasingly being researched for multi-bit capabilities which can be exploited not only for storage (increased memory density) but also for novel applications like in-memory computing e.g. matrix vector multiplication. Such applications require readout circuits, i.e. Sense Amplifiers (SAs) which convert the resistance of the memory cell to digital data. In this work, we propose such a Sense Amplifier (SA) to distinguish between the four states of a ReRAM cell. Unlike conventional multi-bit sensing schemes which require three references to sense four states, the proposed sensing technique requires only a single voltage reference. The proposed SA uses current comparison to sense the first bit and based on the value of the sensed bit, the crucial currents of the SA are manipulated to sense the second bit. The proposed SA is very energy efficient when compared to the state-of-the-art (consuming 82.4 fJ in 130 nm process node) while avoiding the need to generate multiple voltage/current references on-chip. Requiring only 25 transistors and a single $V_{R E F}$, the proposed SA is one of the most compact SA among others in the state-of-the-art. John Reuben, Dietmar Fey |
DSD | 2 |
| 2025 | RISC-V CPU Design Using RRAM-CMOS Standard CellsabstractThe breakdown of Dennard scaling has been the driver for many innovations such as multicore CPUs and has fueled the research into novel devices such as resistive random access memory (RRAM). These devices might be a means to extend the scalability of integrated circuits since they allow for fast and nonvolatile operation. Unfortunately, large analog circuits need to be designed and integrated in order to benefit from these cells, hindering the implementation of large systems. This work elaborates on a novel solution, namely, creating digital standard cells utilizing RRAM devices. Albeit this approach can be used both for small gates and large macroblocks, we illustrate it for a 2T2R-cell. Since RRAM devices can be vertically stacked with transistors, this enables us to construct anandstandard cell, which merely consumes the area of two transistors. This leads to a 25% area reduction compared to an equivalent CMOSnandgate. We illustrate achievable area savings with a half-adder circuit and integrate this novel cell into a digital standard cell library. A synthesized RISC-V core using RRAM-based cells results in a 10.7% smaller area than the equivalent design using standard CMOS gates. Markus Fritscher, Max Uhlmann, Philip Ostrovskyy, Daniel Reiser, Junchao Chen 0001, Jianan Wen, Carsten Schulze, Gerhard Kahmen, Dietmar Fey, Marc Reichenbach, Milos Krstic, Christian Wenger |
IEEE Trans. Very Large Scale Integr. Syst. | 9 |
| 2024 | Parameter Space Exploration of Neural Network Inference Using Ferroelectric Tunnel Junctions for Processing-In-MemoryabstractThis paper explores CMOS-compatible Ferroelectric Tunnel Junctions (FTJs) for Processing-In-Memory (PIM) to address the ‘memory wall’ in traditional computing. A novel FTJ noise model was developed, and hardware-calibrated devices were modeled utilizing IBM's Analog In-Memory Hardware Acceleration Toolkit (AIHWKit). We simulate FTJ-based neural networks for inference only, focusing on mitigating non-idealities such as conductance drift, programming noise, and 1/f read noise. To mitigate FTJ non-idealities affecting different neural networks' accuracy, we developed various hardware-aware (HWA) training techniques including, but not limited to, different weight redistribution, noise resiliency, and custom activation functions. Exploiting the hardware-aware techniques used in this paper, considering different weight-to-conductance mapping and conductance and weight range enlargement, depicts an accuracy improvement in three case studies: full adder implementation, MNIST and CIFAR-10 benchmark, with best-case scenario accuracy of 96.96 %, 86.92 %, and 86.36 %, respectively, after four months of inference. Furthermore, scalability, accuracy im-provement, and experimental validation confirm FTJs' potential for Processing-in-Memory, providing valuable insights for future device and circuit design optimizations to enhance performance and reliability. Shima Hosseinzadeh, Suzanne Lancaster, Amirhossein Parvaresh, Dietmar Fey |
DSD | 5 |
| 2024 | HyFAR: A hypervisor-based fault tolerance approach for heterogeneous automotive real-time systemsabstractFault tolerance is a key aspect for fully autonomous vehicles, as there is no human driver available to take control of the vehicle as a backup. Such autonomous vehicles incorporate signal-oriented and service-oriented hardware and software architectures within one heterogeneous real-time system. Fault tolerance is commonly achieved by adding redundant Electronic Control Units (ECUs) to the system. However, redundant ECUs increase the weight, cost and power consumption of the system. This paper presents a novel hy pervisor-based f ault tolerance approach for a utomotive r eal-time systems (HyFAR), which is based on the largely unexplored concept of migrating software in a highly heterogeneous real-time system using virtualization technology . It is shown, that the fault tolerance of an automotive vehicle can be enhanced in a cost-effective way without the need of additional hardware. The process of recovering critical service-oriented software using a signal-oriented hardware and vice versa is examined. This paper gives a detailed overview of the effects of emulation, virtualization , separation and the type of the hypervisor towards the recovery time and the freedom from interference of signal-oriented and service-oriented software. The results demonstrate that recovering critical service-oriented software using signal-oriented hardware is limited due to missing middle-ware and virtualization support and resource scarcity. However, recovering critical signal-oriented software using a service-oriented hardware is feasible, while a subset of the original service-oriented software can be continued on the same hardware. The resulting approach can be applied to a range of applications including thermal management or lane departure warning. Johannes Lex, Ulrich Margull, Ralph Mader, Dietmar Fey |
J. Syst. Archit. | 4 |
| 2023 | Work in Progress: Extending Virtual Prototypes of Microprocessor Architectures with Accuracy Tracing
Johannes Kliemt, Dietmar Fey |
SIMULTECH | 2 |
| 2021 | A Case for Function-as-a-Service with Disaggregated FPGAsabstractThe slowdown of Moore's law and the end of Dennard scaling created a demand for specialized accelerators, including Field Programmable Gate Arrays (FPGAs), in cloud data centers. At the same time, compute resources are increasingly consumed via public and private clouds and traditional applications are modernized using scalable microservices and Function-as-a-Service (FaaS) offerings. Nonetheless, true FaaS based on FPGAs or other accelerators is virtually absent from the offering catalogs of all major cloud providers. In addition, FPGA applications are typically coded in a monolithic fashion, due to device and vendor specific dependencies, which reduces the portability and usability of FPGA cloud offerings further. However, FPGA-based FaaS can improve execution efficiency and minimize (tail-) latencies while decreasing costs. We propose a novel system architecture, called Mantle, that uses disaggregated FPGAs to enable scalable, usable, portable and efficient FaaS offerings for FPGAs. Our experimental results demonstrate a significant reduction of end-to-end service provisioning time to below 7 seconds and an increase in execution efficiency by a factor of 4 with negligible overhead. Burkhard Ringlein, François Abel, Dionysios Diamantopoulos, Beat Weiss, Christoph Hagleitner, Marc Reichenbach, Dietmar Fey |
CLOUD | 7 |
| 2021 | Taming Non-Deterministic Low-Level I/O: Predictable Multi-Core Real-Time Systems by SoC Co-DesignabstractPredictable and analyzable I/O is one of the considerable challenges in the design of multi-core real-time systems. A common approach to tackle this issue is to partition and schedule I/O transactions such that interference between tasks is minimized. While this works for packet-oriented interfaces with deterministic blocking times, such as ethernet, these techniques are inapplicable to a whole range of I/O devices with nondeterministic behavior that is commonly found in embedded applications. Interfaces, such as SPI, do not allow for fine-grained scheduling and thus exhibit uncontrolled blocking times. Even worse, their configuration and use must be considered as independent transactions requiring costly synchronization between tasks. The resulting detrimental effects are, in particular, pronounced in settings with mixed task requirements on predictability and determinism. All this makes the temporal analysis of such systems cumbersome and overly pessimistic. To solve these issues, we present LOWI/O, an approach to eliminate the interference of low-level non-deterministic I/O interfaces for real-time tasks with high predictability demands (i.e., critical task) while preserving flexibility for tasks with lower requirements (i.e., uncritical tasks). Therefore, we leverage knowledge about the application-specific I/O usage patterns, obtained by static analysis, to derive a tailored hardware architecture. Its key feature is the anticipatory reservation of individual time slots for critical tasks and to mimic preemptivity of I/O units for the remaining system. We have implemented our approach as a toolchain for OSEK-based real-time systems that automatically generates an application-specific SoC design along with a hardware and timing model for subsequent WCET analysis. Our experimental results prove predictable timing for critical tasks with limited impact on uncritical tasks. Steffen Vaas, Peter Ulbrich, Christian Eichler, Peter Wägemann, Marc Reichenbach, Dietmar Fey |
ISORC | 6 |
| 2020 | TReMo: A Model for Ternary ReRAM-Based Memories with Adjustable Write-Verification CapabilitiesabstractWith the increasing use and advancement of memristors, implementation barriers for ternary systems, such as handling more than two states without any extra hardware, could be broken. In this paper, for the first time to the best of our knowledge, a new memory model on circuit level based on ReRAM is modeled for the ternary applications. This novel ternary memory model benefits from a parallel read method, for accomplishing low-latency read operation, and an often-used write-verification method. In addition, a thorough tool for this ternary memory is developed that performs energy, performance and area estimation which is an extension of the existing nonvolatile memory tool called NVSim. Shima Hosseinzadeh, Mehrdad Biglari, Dietmar Fey |
DSD | 3 |
| 2020 | ZRLMPI: A Unified Programming Model for Reconfigurable Heterogeneous Computing ClustersabstractOver the past two decades, the Message Passing Interface (MPI) has evolved as the de-facto standard for programming High-Performance Computing (HPC) clusters. Its widespread utilization led to the rapid development of applications and high reusability. Meanwhile, energy- and compute-efficient devices such as Field-Programmable Gate Arrays (FPGAs) are stepping into modern data centers and HPC clusters to address the nearing end of technology scaling. This combination of traditional CPU servers and FPGA nodes leads to Reconfigurable Heterogeneous HPC (ReH2PC) systems that are particularly cumbersome to program because of the absence of a standard programming model. This work advocates the use of MPI to program such ReH2PC clusters and presents a proof of concept based on a cross-compiler, a High-Level Synthesis library, C++ library, an FPGA- and a CPU-runtime environment. The result is a one-click solution, which compiles a standard MPI application for a ReH2PC cluster. Burkhard Ringlein, François Abel, Alexander Ditter, Beat Weiss, Christoph Hagleitner, Dietmar Fey |
FCCM | 6 |
| 2020 | The allscale framework architecture
Herbert Jordan, Philipp Gschwandtner, Peter Thoman, Peter Zangerl, Alexander Hirsch, Thomas Fahringer, Thomas Heller, Dietmar Fey |
Parallel Comput. | 8 |
| 2019 | System Architecture for Network-Attached FPGAs in the Cloud using Partial ReconfigurationabstractEmerging applications such as deep neural networks, bioinformatics or video encoding impose a high computing pressure on the Cloud. Reconfigurable technologies like Field-Programmable Gate Arrays (FPGAs) can handle such compute-intensive workloads in an efficient and performant way. To seamlessly incorporate FPGAs into existing Cloud environments and leverage their full power efficiency, FPGAs should be directly attached to the data center network and operate independent of power-hungry CPUs. This raises new questions about resource management, application deployment and network integrity. We present a system architecture for managing a large number of network-attached FPGAs in an efficient, flexible and scalable way. To ensure the integrity of the infrastructure, we use partial reconfiguration to separate the non-privileged user logic from the privileged system logic. To create a really scalable and agile cloud service, the management of all resources builds on the Representational State Transfer (REST) concept. Burkhard Ringlein, François Abel, Alexander Ditter, Beat Weiss, Christoph Hagleitner, Dietmar Fey |
FPL | 6 |
| 2018 | The AllScale Runtime Application ModelabstractContemporary state-of-the-art runtime systems underlying widely utilized general purpose parallel programming languages and libraries like OpenMP, MPI, or OpenCL provide the foundation for accessing the parallel capabilities of modern computing architectures. In the tradition of their respective imperative host languages those runtime systems' main focus is on providing means for the distribution and synchronization of operations - while the organization and management of manipulated data is left to application developers. Consequently, the distribution of data remains inaccessible to those runtime systems. However, many desirable system-level features depend on a runtime system's ability to exercise control on the distribution of data. Thus, program models underlying traditional systems lack the potential for the support of those features. In this paper, we present a novel application model granting parallel runtime systems system-wide control over the distribution of user-defined shared data structures. Our model utilizes the high-level nature of parallel programming languages, in particular, the usage of well-typed data structures and the associated hiding of implementation details from the application developers. By being based on a generalization of such data structures and extending the resulting abstraction with features facilitating the automated management of the distribution of those, our model enables runtime systems to dynamically influence the placement and replication of shared data. This paper covers a rigorous formal description of our application model, as well as details on our prototype implementation and experimental results demonstrating its ability to efficiently and scalably manage various data structures in real-world environments. Herbert Jordan, Thomas Heller, Philipp Gschwandtner, Peter Zangerl, Peter Thoman, Dietmar Fey, Thomas Fahringer |
CLUSTER | 6 |
| 2018 | Guest Editorial Memristive-Device-Based ComputingabstractToday’s and emerging computing tasks are extremely demanding in terms of storage, energy efficiency, and computing efficiency; data-intensive/big-data applications and Internet-of-Things are couple of examples. In addition, today’s computer architectures and device technologies are facing major challenges making them incapable to deliver the required functionalities and features. Computers are facing the three well-known walls[1)]: the memory wall, the instruction level parallelism wall, and the power wall. Similarly, nanoscale CMOS technology is facing three walls[2)]: the reliability wall, the leakage wall, and the cost wall. In order for computing systems to continue to deliver sustainable benefits for the foreseeable future society, alternative computing architectures have to be explored in the light of emerging new device technologies. Using memristive device technology[3)]to enable new computing paradigms such as computation-in-memory architecture[4)]–[7)]is one of the emerging alternatives that could provide a huge potential in terms of energy and computing efficiency. Said Hamdioui, Pierre-Emmanuel Gaillardon, Dietmar Fey, Tajana Rosing |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2017 | Memristive devices for computing: Beyond CMOS and beyond von NeumannabstractTraditional CMOS technology and its continuous down-scaling have been the driving force to improve performance of existing computer architectures. Today, however, both technology and computer architectures are facing challenges that make them incapable of delivering the growing computing performance requirement at pre-defined constraints. This forces the exploration of both novel architectures and technologies; not only to maintain the economic profit of technology scaling, but also to enable the computing architecture solutions for big-data and data-intensive applications. This paper discusses the emerging memristive device as a complement (or an alternative) to CMOS devices and shows how such devices enable novel computing paradigms that will solve the challenges of today's architectures for certain applications. The paper covers not only the potential of memristor devices in enabling novel memory technologies, logic design styles, and arithmetic operations, but also their potential in enabling in-memory computing and neuromorphic computing. Hoang Anh Du Nguyen, Lei Xie 0005, Mottaqiallah Taouil, Said Hamdioui, Dietmar Fey |
VLSI-SoC | 6 |
| 2017 | Performance analysis of the Kahan-enhanced scalar product on current multi-core and many-core processorsabstractSummary We investigate the performance characteristics of a numerically enhanced scalar product (dot) kernel loop that uses the Kahan algorithm to compensate for numerical errors, and describe efficient single instruction multiple data‐vectorized implementations on recent multi‐core and many‐core processors. Using low‐level instruction analysis and the execution‐cache‐memory performance model, we pinpoint the relevant performance bottlenecks for single‐core and thread‐parallel execution and predict performance and saturation behavior. We show that the Kahan‐enhanced scalar product comes at almost no additional cost compared with the naive (non‐Kahan) scalar product if appropriate low‐level optimizations, notably single instruction multiple data vectorization and unrolling, are applied. The execution‐cache‐memory model is extended appropriately to accommodate not only modern Intel multicore chips but also the Intel Xeon Phi ‘Knights Corner’ coprocessor and an IBM POWER8 CPU. This allows us to discuss the impact of processor features on the performance across four modern architectures that are relevant for high performance computing. Copyright © 2016 John Wiley & Sons, Ltd. Johannes Hofmann 0001, Dietmar Fey, Michael Riedmann, Jan Eitzinger, Georg Hager, Gerhard Wellein |
Concurr. Comput. Pract. Exp. | 2 |
| 2017 | Comparison of common parallel architectures for the execution of the island model and the global parallelization of evolutionary algorithmsabstractSummary Evolutionary algorithms are one of the most popular forms of optimization algorithms. They are comparatively easy to use and were successfully employed for a wide variety of practical applications. However, frequently, it is necessary to execute them in parallel in order to reduce the runtime. There are a number of different approaches for the parallelization of evolutionary algorithms, and various hardware platforms can be used for the parallel execution. However, not every platform is equally suitable for any kind of parallelization of evolutionary algorithms. In addition, it also depends on properties of the concrete optimization problem to be solved and on the used evolutionary algorithm, which platform is best suited for the execution. The present work observes this in detail for two common forms of parallelization of evolutionary algorithms – the island model and the global parallelization – and for four widely used parallel computing platforms – multi‐core CPUs, clusters, graphics cards, and grids. Based on empirical and analytical investigations, it is determined, under which circumstances an architecture is better suited for the execution of a parallel evolutionary algorithm than another (andvice versa). Guidelines are derived that support users of parallel evolutionary algorithms with the choice of an appropriate platform. Copyright © 2016 John Wiley & Sons, Ltd. Steffen Limmer, Dietmar Fey |
Concurr. Comput. Pract. Exp. | 2 |
| 2017 | Introduction to the special issue on architecture of computing systems
Frank Hannig, João M. P. Cardoso, Dietmar Fey |
J. Syst. Archit. | 3 |
| 2017 | Fast heterogeneous computing architectures for smart antennas
Marc Reichenbach, Maximilian Kasparek, Konrad Häublein, Jan Niklas Bauer, Mohammad Alawieh, Dietmar Fey |
J. Syst. Archit. | 6 |
| 2016 | Investigation of strategies for an increasing population size in multi-objective CMA-ESabstractThe Multi-objective Covariance Matrix Adaptation Evolution Strategy (MO-CMA-ES) is an evolutionary algorithm for continuous vector optimization. It is invariant against rotations and translations of the search space and empirical evaluations have shown that it is very competitive with other popular multi-objective evolutionary algorithms, like NSGA-II. However, MO-CMA-ES requires a certain “warm-up phase” to adapt internal strategy parameters. A promising approach to speed up this “warm-up phase” is to keep the start population small and to gradually increase its size during the optimization. We experimentally investigate two static and three dynamic strategies for increasing the population size. The results show that the employment of a dynamic population size increasing strategy can significantly improve the performance of MO-CMA-ES, especially when the budget of objective function evaluations is small. Steffen Limmer, Dietmar Fey |
CEC | 2 |
| 2016 | Hardware-software Co-simulation of Self-organizing Smart Home Networks - Who am I and Where Are the Others?abstractIn this paper, we present our solution to simulate home automation networks on a functional level in our research project on self-organizing home automation network nodes. We simulate the nodes with our hardware-software co-simulator, based on the virtual machine QEMU and the SystemC hardware simulator. The Virtual Distributed Ethernet suite is used to simulate several hardware-software co-simulators in a network. Furthermore, tools we developed to prepare network node disk images, configure the simulation environment, and generate Linux device drivers from hardware interface specifications are presented. Bruno Kleinert, Franziska Schäfer, Jupiter Bakakeu, Simone Weiß, Dietmar Fey |
SIMULTECH | 5 |
| 2016 | Cache Aware Instruction Accurate Simulation of a 3-D Coastal Ocean Model on Low Power HardwareabstractHigh level hardware simulation and modeling techniques matured significantly over the last years and have become more and more important in practice, e.g., in the industrial hardware development and automotive domain. Yet, there are many other challenging application areas such as numerical solvers for environmental or disaster prediction problems, e.g., tsunami and storm surge simulations, that could greatly profit from accurate and efficient hardware simulation. Such applications rely on complex mathematical models that are discretized using suitable numerical methods, and require a close collaboration between mathematicians and computer scientists to attain desired computational performance on current micro architectures and code parallelization techniques to produce accurate simulation results as fast as possible. This complex and detailed simulation requires a lot of time during preparation and execution. Especially the execution on non-standard or new hardware may be challenging and potentially error prone. In this paper, we focus on a high level simulation approach for determining accurate runtimes of applications using instruction accurate modeling and simulation. We extend the basic instruction accurate simulation technology from OVP using cache models in conjunction with a statistical cost function, which enables high precision and significantly better runtime predictions compared to the pure instruction accurate approach. Dominik Schoenwetter, Alexander Ditter, Vadym Aizinger, Balthasar Reuter, Dietmar Fey |
SIMULTECH | 5 |
| 2015 | Architecture and simulation of a hybrid memristive multiplier network using redundant number representationabstractThe paper proposes an architecture for a multiplier-adder network that can be used for the design of a digital neuron cell. The core of the multiplier is based on a hybrid memristor network, in which digital CMOS logic is combined with multi-stable storing memristor devices. The multi-bit storing feature of memristors is favoured since it simplifies the realisation of ternary data. Using such a ternary number system in binary logic leads to a redundant number representation (RNR) that allows to speed up multiplications since they are reduced to adders working in constant time independent of the word length. For the verification of the multiplier architecture an own special simulation system was developed in C++ allowing flexible design and fast analogue simulation of large complex memristor networks. The superiority of the hybrid memristive architecture in terms of latency and bandwidth compared to an adder with carry-look-ahead technique is analytically shown. Dietmar Fey, Jonathan Martschinke |
IJCNN | 1 |
| 2015 | Tsunami and Storm Surge Simulation Using Low Power Architectures - Concept and EvaluationabstractPerforming a tsunami or storm surge simulation in real time is a highly challenging research topic that calls
for a collaboration between mathematicians and computer scientists. One must combine mathematical models
with numerical methods and rely on computational performance and code parallelization to produce accurate
simulation results as fast as possible. The traditional modeling approaches require a lot of computing power
and significant amounts of electrical energy; they are also highly dependent on uninterrupted access to a reliable
power supply. This paper presents a concept how to develop suitable low power hardware architectures
for tsunami and storm surge simulations based on cooperative software and hardware simulation. The main
goal is to be able - if necessary - to perform simulations in-situ and battery-powered. For flood warning systems
installed in regions with weak or unreliable power and computing infrastructure, this would significantly
decrease the risk of failure at the most critical moments. Dominik Schoenwetter, Alexander Ditter, Bruno Kleinert, Arne Hendricks, Vadym Aizinger, Harald Köstler, Dietmar Fey |
SIMULTECH | 7 |
| 2015 | Synthesis and optimization of image processing accelerators using domain knowledge
Oliver Reiche, Konrad Häublein, Marc Reichenbach, Moritz Schmid, Frank Hannig, Jürgen Teich, Dietmar Fey |
J. Syst. Archit. | 7 |
| 2014 | A Generic Approach for Analysis of White-Light Interferometry Data via User-Defined Algorithms
Max Schneider, Dietmar Fey, Kay Wenzel, Torsten Machleidt |
ICCSA (4) | 2 |
| 2013 | An auto-tuning approach for optimizing base operators for non-destructive testing applications on heterogeneous multi-core architecturesabstractThe field of non-destructive testing imposes rising performance requirements related to the compute resources necessary to satisfy application needs. The creation of scalable applications in that field, which efficiently utilize today's mostly heterogeneous compute resources has proven to be complicated. In order to allow domain experts coming from a materials science background, to exploit modern heterogenous computing architectures, new classes of frameworks for efficient applications need to be developed. This paper proposes a new framework to develop efficient operators and operator chains for applications in the field of non-destructive testing, based on a self-organizing autotuning approach. The framework provides a C++ Interface to define base operators and implements an Embedded Domain Specific Language (EDSL) using C++ Expression Templates (ETs) which allows a succinct definition of an Operator Chain. The framework applies various kinds of optimizations by utilizing C++ Template Metaprogramming techniques. These optimizations include feedback based auto-tuning and auto-parallelization of the resulting Pipeline based on the advanced dataflow and future's techniques provided by the High Performance ParalleX (HPX), a general purpose parallel C++ runtime system. Thomas Heller, Dietmar Fey, Markus Rehak |
ISORC | 2 |
| 2012 | Competence model for embedded micro-and nanosystemsabstractIn this paper, we describe the results of an expert survey conducted within the project “Competence development with embedded micro- and nanosystems”. This survey was designed in order to refine our previously derived normative competence structure model for developers of embedded micro-and nanosystems towards an empirically refined competence structure model. Today the size of structural parts of nanosystems has scaled down to as few as a couple of molecules. This results in a lot of challenges regarding permanent and transient faults. Therefore bottom-up development techniques for nanostructured systems are now included into our competence model to prepare future developers to the specific challenges of designing at the nano scale. The evaluation of the content validity of the empirically refined competence structure model was accomplished by the presented expert rating. André Schäfer, Rainer Brück 0001, Bruno Kleinert, Harald Schmidt, Dietmar Fey, Steffen Büchner 0001, Steffen Jaschke, Sigrid E. Schubert |
EDUCON | 5 |
| 2012 | Evolutionary Design of Active Free Space Optical Networks Based on Digital Mirror Devices
Steffen Limmer, Dietmar Fey, Ulrich Lohmann, Jürgen Jahns |
EvoApplications | 2 |
| 2012 | Optimization of a Short-Range Proximity Effect Correction Algorithm in E-Beam Lithography Using GPGPUs
Max Schneider, Nikola Belic, Christoph Sambale, Dietmar Fey |
ICA3PP (1) | 5 |
| 2012 | The empirically refined competence structure model for embedded micro- and nanosystemsabstractTeaching the development of embedded micro- and nanosystems needs a well structured foundation. In this paper, we show how we refined our normative competence structure model (NCSM) to an empirical competence structure model (ECSM) by incorporating the results of a survey of German experts in the field of embedded systems. In addition, we introduce a course concept for a new internship which takes the results of the ECSM into account. André Schäfer, Rainer Brück 0001, Steffen Büchner 0001, Steffen Jaschke, Sigrid E. Schubert, Dietmar Fey, Bruno Kleinert, Harald Schmidt |
ITiCSE | 6 |
| 2012 | Realizing real-time centroid detection of multiple objects with marching pixels algorithms on programmable customizing hardwareabstractSUMMARY In this paper, we present a class of emergent algorithms called Marching Pixels and a corresponding programmable parallel chip architecture. Marching Pixels can be used for real‐time image processing in smart camera chips. They are based on hardware agents, which are virtually crawling in a pixel grid image to find attributes like centroid, rotation, and size of an arbitrary number of objects given in an image. Because of the distributed and local processing scheme of Marching Pixels, reply times in milliseconds can be fulfilled. This means that time is determined where pre‐known objects are located and how they are oriented to the main axes of the image. We present an example Marching Pixels algorithm and corresponding application‐specific and programmable parallel architectures. The latter contains a specific instruction set that allows not only the execution of Marching Pixels algorithms but also of arbitrary Cellular Automata algorithms as an embedded parallel processor. The strengths and weaknesses of this architecture concerning the realization as field‐programmable gate arrays and application‐specific integrated circuits are discussed by means of hardware synthesis results. These results are compared with the solution achievable on a real hardware like the Atom processor. Copyright © 2011 John Wiley & Sons, Ltd. Dietmar Fey, Marc Reichenbach, Marcus Komann, Ralf Seidler |
Concurr. Comput. Pract. Exp. | 1 |
| 2012 | Optical multiplexing techniques for photonic Clos networks in High Performance Computing Architectures
Dietmar Fey, Max Schneider, Jürgen Jahns, Hans Knuppertz |
J. Supercomput. | 1 |
| 2011 | Evolutionary optimization of layouts for high density free space optical network linksabstractElectrical chip- and board-level connections are becoming more and more a bottleneck in computation. A solution to that problem could be optical connections, which allow a higher bandwidth. The usage of free space optics can avoid the problem of crosstalk and geometrical signal path crossings in systems with a high density of interconnections. The choice of appropriate design parameters, allowing the realization of such interconnections, is a complicated task. We present an evolutionary algorithm that is able to find these parameters. We describe the parallel execution of that algorithm and present optimization results. Steffen Limmer, Dietmar Fey, Ulrich Lohmann, Jürgen Jahns |
GECCO | 2 |
| 2011 | A normative competence structure model for embedded micro- and nanosystems developmentabstractIn our poster, we present the research project "Competence development with embedded micro- and nanosystems (KOMINA)" promoted by the German Research Foundation and its first step, the development of a normative competence structure model. As a result, we clustered the derived competencies in the following dimensions: (C1) Competencies as preconditions (Basis), (C2) development competencies of embedded micro- and nanosystems (EMNS), (C3) competencies for multi-level development of EMNS and (C4) non-cognitive competencies. While we chose a purely thematic division for braking down C1 and C2 to sub dimensions, we recommend dividing C3 into solution approaches. We exemplify the deduction from a sub dimension to a concrete task of a practical course. André Schäfer, Rainer Brück 0001, Steffen Jaschke, Sigrid E. Schubert, Dietmar Fey, Bruno Kleinert, Harald Schmidt |
ITiCSE | 5 |
| 2010 | Revising the Trade-off between the Number of Agents and Agent Intelligence
Marcus Komann, Dietmar Fey |
EvoApplications (1) | 2 |
| 2009 | Distributed vision with smart pixelsabstractWe study a problem related to computer vision: How can a field of sensors compute higher-level properties of observed objects deterministically in sublinear time, without accessing a central authority? This issue is not only important for real-time processing of images, but lies at the very heart of understanding how a brain may be able to function. In particular, we consider a quadratic field of n "smart pixels" on a video chip that observe a B/W image. Each pixel can exchange low-level information with its immediate neighbors. We show that it is possible to compute the centers of gravity along with a principal component analysis of all connected components of the black grid graph in time O(sqrt(n)), by developing appropriate distributed protocols that are modeled after sweepline methods. Our method is not only interesting from a philosophical and theoretical point of view, it is also useful for actual applications for controling a robot arm that has to seize objects on a moving belt. We describe details of an implementation on an FPGA; the code has also been turned into a hardware design for an application-specific integrated circuit (ASIC). Sándor P. Fekete, Dietmar Fey, Marcus Komann, Alexander Kröller, Marc Reichenbach, Christiane Schmidt 0001 |
SCG | 2 |
| 2009 | On the Effectiveness of Evolution Compared to Time-Consuming Full Search of Optimal 6-State Automata
Marcus Komann, Patrick Ediger, Dietmar Fey, Rolf Hoffmann 0001 |
EuroGP | 3 |
| 2009 | Evaluating the evolvability of emergent agents with different numbers of statesabstractEmergence is an important and promising scientific topic today because it offers benefits that can not be achieved by classic means. But it is often challenging to control emergence and to find correct local rules that create desired global behavior. It especially becomes difficult if the search space representing the problem that has to be optimized is not continuous/linear. One solution to that problem is evolution. This paper shows that the use of Genetic Algorithms is feasible for such problems by the example of the Creatures' Exploration Problem in which agents shall visit all non-blocked cells in a grid. Different amounts of agents and states per agent are evolved and statistically compared. It shows that neither a single extension of agent capabilities nor sole increase of agent numbers provides the best performance. The results hint that a mixture of both should be used instead. Marcus Komann, Dietmar Fey |
GECCO | 2 |
| 2009 | Model-Driven Design and Organic Computing -- Two Different but Possibly Accordable Concepts for the Design of Embedded SystemsabstractModel Driven Design (MDD) and Organic Computing (OC) present two different approaches to control the continually increasing complexity of technical systems which is a consequence of Moore's law. This statement holds in particular for the design of hardware and software for embedded systems. MDD is the effort to solve this challenge by a widely hierachical and modular organized design process in which the actual design is seperated from architecture. Whereas the design addresses the functional requirements, i.e. expressing the functional behaviour of the designed system in a platform-independent way, the architecture handles questions concerning infrastructure in particular, e.g. how non-functional requirements like scalability, reliability and performance are realized. In this context, architecture does not refer to a concrete hardware system, but rather to a kind of template suite or programming environment which can run on different real computer architectures. The goal of MDD is to map the platform- independent description to the architecture automatically. Dietmar Fey |
ISORC | 1 |
| 2007 | An Organic Computing architecture for visual microprocessors based on Marching PixelsabstractThe paper presents architecture and synthesis results for an organic computing hardware for smart CMOS camera chips. The organic behavior in the chip hardware is based on distributed and emergent functionality exploited for detection of objects and their center points given in binary images. Future real-time embedded systems used in industrial image processing have to provide reply times in the range of milliseconds. It is impossible to meet such strict requirements for megapixel resolutions with serial processing schemes in particular if multiple given objects have to be detected. Even classical parallel techniques like SIMD or MIMD approaches are not sufficient due to their dependency on more or less central control structures. To achieve more flexibility, unlimited scalability and higher performance parallel emergent architectures are necessary. We present such an approach, denoted as marching pixels, for future digital visual microprocessors. Marching pixels work similar to artificial ants. They are crawling as hardware agents within a pixel field, e.g. to identify and to detect center points of an arbitrary number of objects given in an image. We present an emergent marching pixel algorithm for the processing of arbitrary concave objects and its mapping onto real hardware. Based on synthesis results for FPGAs and ASICs we discuss the possibilities of digital organic computing approaches for visual microprocessors for future smart high-speed camera systems. Dietmar Fey, Marcus Komann, Frank Schurz, Andreas Loos |
ISCAS | 1 |
| 2006 | Parallelization of Simulations for Various Magnetic System Models on Small-Sized Cluster Computers with MPI
Frank Schurz, Dietmar Fey, Dmitri Berkov |
ICCSA (5) | 2 |
| 2006 | SOGOS - A Distributed Meta Level Architecture for the Self-Organizing Grid of ServicesabstractHandling highly dynamic scenarios as they arise in emergency situations requires lots of semantic information about the situation and an extremely flexible, selforganizing IT infrastructure that provides services that can be used to manage the situation. We show that a distributed meta level architecture is particularly suited for the implementation of such a self-organizing grid of services. This architecture (SOGOS) distinguishes between an object level and a meta level. The middleware processes of the grid are running on the object level. The meta level defines an explicitly and declaratively represented dynamic meta model that provides the semantics for the object level processes. Additionally, this level runs processes that plan, supervise and control mobile agents on the object level. The levels are linked together by reflection processes that ensure that relevant changes on the object level are reflected in the meta model and vice versa. The corresponding reflection principles provide the basis for the implementation of the selforganizing mechanisms that govern the overall system. Clemens Beckstein, Peter Dittrich, Christian Erfurth, Dietmar Fey, Birgitta König-Ries, Martin Mundhenk, Harald Sack |
MDM | 4 |
| 2000 | Optical interconnects for neural and reconfigurable VLSI architecturesabstractThe increasing transistor density in very large-scale integrated (VLSI) circuits and the limited pin member in the off-chip communication lead to a situation described as interconnect crisis in micro-electronics. Optoelectronic VLSI (OE-VLSI) circuits using short-distance optical interconnects and optoelectronic devices like microlaser, modulator, and detector arrays for optical off-chip sending and receiving offer a technology to overcome this crisis. However, in order to exploit efficiently the potential of thousands of optical off-chip interconnects, an appropriate VLSI architecture is required. We show for the example of neural and reconfigurable VLSI architectures that fine-grain architectures fulfill these requirements. An OE-VLSI circuit realization based on multiple quantum-well modulators functioning as two-dimensional (2-D) optical input/output (I/O) interface for the chip is presented. Due to the parallel optical interface, and improvement of two to three orders of magnitude in the throughput performance is possible compared to all-electronic solutions. For the optical interconnects, a planar-integrated free-space optical system has been designed leading to an optical multichip module. Such a system has been fabricated and experimentally characterized. Furthermore, we designed an manufactured fiber arrays, which will be the core element for a convenient test station for the 2-D optoelectronic I/O interface of OE-VLSI circuits. Dietmar Fey, Werner Erhard, Matthies Gruber, Jürgen Jahns, Hartmut Bartelt, Guido Grimm, Lutz Hoppe, Stefan Sinzinger |
Proc. IEEE | 1 |
| 2000 | Digit Pipelined Arithmetic for 3-D Massively Parallel Optoelectronic Circuits
Dietmar Fey, Marko Degenkolb |
J. Supercomput. | 1 |
| 1999 | 3D Optoelectronic Fix Point Unit and Its Advantages Processing 3D Data
Bernd Kasche, Dietmar Fey, T. Höhn, Werner Erhard |
Euro-Par | 2 |