VLDB 2026 Research / reviewers in the wild / expert
Merten Popp
dblp:188/1299
· DBLP profile ↗
5ranked-venue papers
3as first author
1since 2021 · last 2021
0000-0001-5916-8180ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Theory of computation · 2 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 1Systems, architecture and hardware · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2021 | Multilevel Acyclic Hypergraph PartitioningabstractA directed acyclic hypergraph is a generalized concept of a directed acyclic graph, where each hyperedge can contain an arbitrary number of tails and heads.Directed hypergraphs can be used to model data flow and execution dependencies in streaming applications.Thus, hypergraph partitioning algorithms can be used to obtain efficient parallelizations for multiprocessor architectures.However, an acyclicity constraint on the partition is necessary when mapping streaming applications to embedded multiprocessors due to resource restrictions on this type of hardware.The acyclic hypergraph partitioning problem is to partition the hypernodes of a directed acyclic hypergraph into a given number of blocks of roughly equal size such that the partition is acyclic while minimizing an objective function.Here, we contribute the first n-level algorithm for the acyclic hypergraph partitioning problem.Based on this, we engineer a memetic algorithm to further reduce communication cost, as well as to improve scheduling makespan on embedded multiprocessor architectures.Experiments indicate that our algorithm outperforms previous algorithms that focus on the directed acyclic graph case which have previously been employed in the application domain.Moreover, our experiments indicate that using the directed hypergraph model for this type of application yields a significantly smaller makespan. Merten Popp, Sebastian Schlag, Christian Schulz 0003, Daniel Seemaier |
ALENEX | 1 |
| 2018 | Evolutionary multi-level acyclic graph partitioningabstractDirected graphs are widely used to model data flow and execution dependencies in streaming applications. This enables the utilization of graph partitioning algorithms for the problem of parallelizing execution on multiprocessor architectures under hardware resource constraints. However due to program memory restrictions in embedded multiprocessor systems, applications need to be divided into parts without cyclic dependencies. This can be done by a subsequent second graph partitioning step with an additional acyclicity constraint. Orlando Moreira, Merten Popp, Christian Schulz 0003 |
GECCO | 2 |
| 2017 | Automatic Control Flow Generation for OpenVX GraphsabstractHeterogeneous platforms with large numbers of processing elements (PEs) have been proposed to satisfy the computational requirements of computer vision applications. Limiting the incurred communication cost here is key to meet the power constraints of embedded devices.We present a new heuristic to reduce communication among PEs and to external memory by aggregating inter-process communication and pipelining image processing functions. The application is specified as an OpenVX graph, an industry standard for vision applications, though our method is applicable to dataflow in general. We use dataflow and graph-based analysis techniques to map the application at configuration time to a hardware platform with strong program memory constraints. We show that our approach can yield a reduction of up to 53% in communication compared to other OpenVX implementations. Merten Popp, Stef van Son, Orlando Moreira |
DSD | 1 |
| 2017 | Graph Partitioning with Acyclicity ConstraintsabstractGraphs are widely used to model execution dependencies in applications. In particular, the NP-complete problem of partitioning a graph under constraints receives enormous attention by researchers because of its applicability in multiprocessor scheduling. We identified the additional constraint of acyclic dependencies between blocks when mapping streaming applications to a heterogeneous embedded multiprocessor. Existing algorithms and heuristics do not address this requirement and deliver results that are not applicable for our use-case. In this work, we show that this more constrained version of the graph partitioning problem is NP-complete and present heuristics that achieve a close approximation of the optimal solution found by an exhaustive search for small problem instances and much better scalability for larger instances. In addition, we can show a positive impact on the schedule of a real imaging application that improves communication volume and execution time. Orlando Moreira, Merten Popp, Christian Schulz 0003 |
SEA | 2 |
| 2016 | Automatic HAL generation for embedded multiprocessor systemsabstractAutomated hardware design flows considerably speed up the development of embedded systems and are a useful asset during architecture exploration phase. However, any existing software has to be adapted for every new system. In this work we will demonstrate, how a Hardware Abstraction Layer (HAL) for device addresses and properties can be automatically generated from a formal system description while providing sufficient abstraction from hardware details. Comparison to earlier projects show that this saves between 40--50 person-weeks of work per IP. Necessary device specific information is stored using a novel approach that allows the compiler to remove unused data. We will show with a real imaging application that this can reduce the amount of memory used by the HAL from ≈ 9.5% to ≈ 5.1% of the available scratchpad memory. It will also be shown that the overhead of this HAL only depends on the level of abstraction that is used by the application and that performance and memory usage will equal a hard-coded solution in the case that an application uses compile-time constant device identifiers. Merten Popp, Orlando Moreira, Wim Yedema, Menno Lindwer |
EMSOFT | 1 |