Ulrich Kremer

dblp:19/6266 · DBLP profile ↗
← Back
26ranked-venue papers
2as first author
2since 2021 · last 2025
0000-0002-0713-578XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 18 · 2 first-author · 2 since 2021Software engineering, systems software and programming languages · 7Applied, interdisciplinary, general and emerging computing · 4 · 1 first-authorArtificial intelligence and machine learning · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
7 papers
Cloud and datacenter computing · 29% Distributed systems · 29% Memory systems · 29%
Software engineering, system software, and programming languages
6 papers
Compilers and program optimization · 74% Programming languages and type systems · 10% Operating systems · 9%
Computer networks
1 paper
Wireless networking · 77% Internet of things and sensor networks · 23%

Topics — the 20 heaviest of 23, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Memory systems
inter-application interference
0.412020
Global cost/quality management across multiple applications · ESEC/SIGSOFT FSE 2020
Cloud and datacenter computing
resource management
0.412020
Global cost/quality management across multiple applications · ESEC/SIGSOFT FSE 2020
Distributed systems
resource sharing
0.412020
Global cost/quality management across multiple applications · ESEC/SIGSOFT FSE 2020
Wireless networking
mobile ad hoc networks
0.112005
Programming ad-hoc networks of mobile and resource-constrained devices · PLDI 2005
Compilers and program optimization
program transformation
0.012004
Code Transformations for Energy-Efficient Device Management · IEEE Trans. Computers 2004
Energy-efficient computing
power management
0.012004
Code Transformations for Energy-Efficient Device Management · IEEE Trans. Computers 2004
Compilers and program optimization › compiler optimization
energy-aware compilation
0.012003
The design, implementation, and evaluation of a compiler algorithm for CPU energy reduction · PLDI 2003
Energy-efficient computing › voltage scaling
dynamic voltage scaling
0.012003
The design, implementation, and evaluation of a compiler algorithm for CPU energy reduction · PLDI 2003
Parallel and multicore computing
parallel programming models
0.031998
Automatic Data Layout for Distributed-Memory Machines · ACM Trans. Program. Lang. Syst. 1998
Automatic Data Layout for High Performance Fortran · SC 1995
The parascope editor: an interactive parallel programming tool · SC 1989
Internet of things and sensor networks
resource-constrained devices
0.012005
Programming ad-hoc networks of mobile and resource-constrained devices · PLDI 2005
Operating systems › resource management
power management
0.012004
Code Transformations for Energy-Efficient Device Management · IEEE Trans. Computers 2004
Parallel and multicore computing › parallel programming models › data-parallel language
high performance fortran
0.011995
Automatic Data Layout for High Performance Fortran · SC 1995
Energy-efficient computing › power management
processor power reduction
0.012003
The design, implementation, and evaluation of a compiler algorithm for CPU energy reduction · PLDI 2003
Parallel and multicore computing
data distribution
0.011991
A Static Performance Estimator to Guide Data Partitioning Decisions · PPoPP 1991
Performance modeling and evaluation
static performance analysis
0.011991
A Static Performance Estimator to Guide Data Partitioning Decisions · PPoPP 1991
Compilers and program optimization
parallelizing compiler
0.011998
Automatic Data Layout for Distributed-Memory Machines · ACM Trans. Program. Lang. Syst. 1998
Program analysis
data dependence analysis
0.011989
The parascope editor: an interactive parallel programming tool · SC 1989
Compilers and program optimization
parallelization
0.011989
The parascope editor: an interactive parallel programming tool · SC 1989
Software maintenance and evolution
performance prediction
0.011995
Automatic Data Layout for High Performance Fortran · SC 1995
High-performance computing
performance optimization
0.011991
A Static Performance Estimator to Guide Data Partitioning Decisions · PPoPP 1991

Methods — techniques the papers use, named apart from their topics

modeling · 0.1experimentation · 0.1source-to-source transformation · 0.1dynamic voltage scaling · 0.1performance estimation · 0.0integer programming · 0.0remapping · 0.0data layout framework · 0.0machine-specific estimation · 0.0kernel routine training · 0.0
YearPublicationVenuePosition
2025 Are Edge MicroDCs Equipped to Tackle Memory Contention?
abstract
Hazard monitoring systems rely on micro datacenters (MicroDCs) for local data processing and real-time response in resource- and energy-constrained environments. These MicroDCs often host diverse, multitenant applications---such as object detection and sensor data ingestion---that contend for shared memory. Through a case study of compute-intensive and I/O-intensive applications, we show that different applications use memory differently (e.g., heap vs. OS-managed page cache), leading to asymmetric performance degradation under memory pressure. Our findings highlight the limitations of existing OS-level resource management approaches and motivate the need for cross-layered coordination between applications and the operating system to treat all memory uses as first-class citizens and adapt to changing workload demands in MicroDCs.
Long Tran, River Bartz, Ramakrishnan Durairajan, Ulrich Kremer, Sudarsun Kannan
HotStorage4
2022 An Adaptive Application Framework with Customizable Quality Metrics
abstract
Many embedded environments require applications to produce outcomes under different, potentially changing, resource constraints. Relaxing application semantics through approximations enables trading off resource usage for outcome quality. Although quality is a highly subjective notion, previous work assumes given, fixed low-level quality metrics that often lack a strong correlation to a user’s higher-level quality experience. Users may also change their minds with respect to their quality expectations depending on the resource budgets they are willing to dedicate to an execution. This motivates the need for an adaptive application framework where users provide execution budgets and a customized quality notion. This article presents a novel adaptive program graph representation that enables user-level, customizable quality based on basic quality aspects defined by application developers. Developers also define application configuration spaces, with possible customization to eliminate undesirable configurations. At runtime, the graph enables the dynamic selection of the configuration with maximal customized quality within the user-provided resource budget. An adaptive application framework based on our novel graph representation has been implemented on Android and Linux platforms and evaluated on eight benchmark programs, four with fully customizable quality. Using custom quality instead of the default quality, users may improve their subjective quality experience value by up to 3.59×, with 1.76× on average under different resource constraints. Developers are able to exploit their application structure knowledge to define configuration spaces that are on average 68.7% smaller as compared to existing, structure-oblivious approaches. The overhead of dynamic reconfiguration averages less than 1.84% of the overall application execution time.
Liu Liu 0015, Sibren Isaacman, Ulrich Kremer
ACM Trans. Design Autom. Electr. Syst.3
2020 Global cost/quality management across multiple applications
abstract
Approximation is a technique that optimizes the balance between application outcome quality and its resource usage. Trading quality for performance has been investigated for single application scenarios, but not for environments where multiple approximate applications may run concurrently on the same machine, interfering with each other by sharing machine resources. Applying existing, single application techniques to this multi-programming environment may lead to configuration space size explosion, or result in poor overall application quality outcomes.
Liu Liu 0015, Sibren Isaacman, Ulrich Kremer
ESEC/SIGSOFT FSE3
2017 POSTER: Exploiting Approximations for Energy/Quality Tradeoffs in Service-Based Applications
abstract
Approximations and redundancies allow mobile and distributed applications to produce answers or outcomes of lesser quality at lower costs. This paper introduces RAPID, a new programming framework and methodology for service-based applications with approximations and redundancies. Finding the best service configuration under a given resource budget becomes a constrained, dual-weight graph optimization problem.
Liu Liu 0015, Sibren Isaacman, Abhishek Bhattacharjee, Ulrich Kremer
PACT4
2015 TrilobiteG: A programming architecture for autonomous underwater vehicles
abstract
Programming autonomous systems can be challenging because many programming decisions must be made in real time and under stressful conditions, such as on a battle field, during a short communication window, or during a storm at sea. As such, new programming designs are needed to reflect these specific and extreme challenges.
Hans C. Woithe, Ulrich Kremer
LCTES2
2010 Adaptive spatiotemporal node selection in dynamic networks
abstract
Dynamic networks - spontaneous, self-organizing groups of devices - are a promising new computing platform. Writing applications for such networks is a daunting task, however, due to their extreme variability and unpredictability, with many devices having significant resource limitations. Intelligent, automated distribution of work across network nodes is needed to get the most out of limited resource budgets.
Pradip Hari, John B. P. McCabe, Jonathan Banafato, Marcus Henry, Kevin Ko, Emmanouil Koukoumidis, Ulrich Kremer, Margaret Martonosi, Li-Shiuan Peh
PACT7
2009 A programming architecture for smart autonomous underwater vehicles
abstract
Autonomous underwater vehicles (AUVs) are an indispensable tool for marine scientists to study the world's oceans. The Slocum glider is a buoyancy driven AUV designed for missions that can last weeks or even months. Although successful, its hardware and layered control architecture is rather limited and difficult to program. Due to limits in its hardware and software infrastructure, the Slocum glider is not able to change its behavior based on sensor readings while underwater. In this paper, we discuss a new programming architecture for AUVs like the Slocum. We present a new model that allows marine scientists to express AUV missions at a higher level of abstraction, leaving low-level software and hardware details to the compiler and runtime system. The Slocum glider is used as an illustration of how our programming architecture can be implemented within an existing system. The Slocum's new framework consists of an event driven, finite state machine model, a corresponding compiler and runtime system, and a hardware platform that interacts with the glider's existing hardware infrastructure. The new programming architecture is able to implement changes in glider behavior in response to sensor readings while submerged. This crucial capability will enable advanced glider behaviors such as underwater communication and swarming. Experimental results based on simulation and actual glider deployments off the coast of New Jersey show the expressiveness and effectiveness of our prototype implementation.
Hans C. Woithe, Ulrich Kremer
IROS2
2008 Execution context optimization for disk energy
abstract
Power, energy, and thermal concerns have constrained embedded systems designs. Computing capability and storage density have increased dramatically, enabling the emergence of handheld devices from special to general purpose computing. In many mobile systems, the disk is among the top energy consumers. Many previous optimizations for disk energy have assumed uniprogramming environments. However, many optimizations degrade in multiprogramming because programs are unaware of other programs (execution context). We introduce a framework to make programs aware of and adapt to their runtime execution context. We evaluated real workloads by collecting user activity traces and characterizing the execution contexts. The study confirms that many users run a limited number of programs concurrently. We applied execution context optimizations to eight programs and tested ten combinations. The programs ran concurrently while the disk’s power was measured. Our measurement infrastructure allows interactive sessions to be scripted, recorded, and replayed to compare the optimizations’ effects against the baseline. Our experiments covered two write cache policies. For write-through, energy savings was in the range 3–63 % with an average of 21%. For writeback, energy savings was in the range-33–61 % with an average of 8%. In all cases, our optimizations incurred less than 1 % performance penalty.
Jerry Hom, Ulrich Kremer
CASES2
2007 Efficient Program Power Behavior Characterization
Chunling Hu, Daniel A. Jiménez, Ulrich Kremer
HiPEAC3
2005 Inter-program optimizations for conserving disk energy
abstract
Previous work has shown that intra-program optimizations, i.e., optimizations performed on individual programs in isolation, can be very effective in reducing disk energy in streaming applications. This paper investigates the potential additional benefits of inter-program optimizations where sets of programs are optimized together. Experimental results on different subsets of three streaming applications show that 7–49 % additional energy savings (27.3 % on average) can be obtained with negligible performance penalties using two novel inter-program optimizations, namely execution context sensitive buffer size selection and inverse barrier synchronization. These figures were obtained via physical measurements on two laptop disks. Categories and Subject Descriptors: D.3.3 [Frameworks]:
Jerry Hom, Ulrich Kremer
ISLPED2
2005 Programming ad-hoc networks of mobile and resource-constrained devices
Ulrich Kremer, Adrian Stere, Liviu Iftode
PLDI2
2004 Spatial Programming Using Smart Messages: Design and Implementation
abstract
Spatial programming (SP) is a space-aware programming model for outdoor distributed embedded systems. Central to SP are the concepts of space and spatial reference, which provide applications with a virtual resource naming in networks of embedded systems. A network resource is referenced using its expected physical location and properties. Together with other SP features, such as reference consistency and access timeout, they help programmers cope with highly dynamic network configurations in a network-transparent fashion. We present the SP design and its implementation using smart messages, a lightweight software architecture similar to mobile agents, that we developed for networks of embedded systems. We also describe the implementation and evaluation of a simple SP application over a testbed consisting of HP iPAQs running Linux and equipped with 802.11 cards for wireless communication. The experimental results indicate that SP is a viable programming model for outdoor distributed computing.
Cristian Borcea, Chalermek Intanagonwiwat, Porlin Kang, Ulrich Kremer, Liviu Iftode
ICDCS4
2004 Smart Messages: A Distributed Computing Platform for Networks of Embedded Systems
abstract
In this paper, we present the design and implementation of Smart Messages, a distributed computing platform for networks of embedded systems based on execution migration. A Smart Message (SM) is a user-defined distributed program which executes on nodes of interest, named by their properties, and uses an explicit lightweight migration to reach these nodes. During migrations, an SM carries its code and execution state, and it self-routes at each intermediate node between two nodes of interest. The nodes in the network cooperate to support the SM execution by providing a virtual machine and a shared memory region addressable by names (tag space). To illustrate the flexibility of SMs to program real world applications, we describe EZCab, an application for booking cabs in densely populated urban areas. We also present experimental results to quantify the performance achieved by the SM prototype.
Porlin Kang, Cristian Borcea, Akhilesh Saxena, Ulrich Kremer, Liviu Iftode
Comput. J.5
2004 Code Transformations for Energy-Efficient Device Management
abstract
Energy conservation without performance degradation is an important goal for battery-operated computers, such as laptops and hand-held assistants. We study application-supported device management for optimizing energy and performance. In particular, we consider application transformations that increase device idle times and inform the operating system about the length of each upcoming, period of idleness. We use modeling and experimentation to assess the potential energy and performance benefits of this type of application support for a laptop disk. Furthermore, we propose and evaluate a compiler framework for performing the transformations automatically. Our main modeling results show that the transformations are potentially beneficial. However, our experimental results with six real laptop applications demonstrate that, unless applications are transformed, they cannot accrue any of the predicted benefits. In addition, they show that our compiler can produce almost the same performance and energy results as hand-modifying applications. Overall, we find that the transformations can reduce disk energy consumption from 55 percent to 89 percent with degradation in performance of at most 8 percent.
Taliver Heath, Eduardo Pinheiro, Jerry Hom, Ulrich Kremer, Ricardo Bianchini
IEEE Trans. Computers4
2004 A Quantitative Analysis of Tile Size Selection Algorithms
Chung-Hsing Hsu, Ulrich Kremer
J. Supercomput.2
2003 The design, implementation, and evaluation of a compiler algorithm for CPU energy reduction
abstract
This paper presents the design and implementation of a compiler algorithm that effectively optimizes programs for energy usage using dynamic voltage scaling (DVS). The algorithm identifies program regions where the CPU can be slowed down with negligible performance loss. It is implemented as a source-to-source level transformation using the SUIF2 compiler infrastructure. Physical measurements on a high-performance laptop show that total system (i.e., laptop) energy savings of up to 28% can be achieved with performance degradation of less than 5% for the SPECfp95 benchmarks. On average, the system energy and energy-delay product are reduced by 11% and 9%, respectively, with a performance slowdown of 2%. It was also discovered that the energy usage of the programs using our DVS algorithm is within 6% from the theoretical lower bound. To the best of our knowledge, this is one of the first work that evaluates DVS algorithms by physical measurements.
Chung-Hsing Hsu, Ulrich Kremer
PLDI2
2002 Compilers for power and energy management
abstract
No abstract available.
Ulrich Kremer
ISLPED1
2001 Compiler-directed dynamic voltage/frequency scheduling for energy reduction in mircoprocessors
abstract
Article Compiler-directed dynamic voltage/frequency scheduling for energy reduction in microprocessors Share on Authors: Chung-Hsing Hsu CS Dept., Rutgers University, New Jersey CS Dept., Rutgers University, New JerseyView Profile , Ulrich Kremer CS Dept., Rutgers University, New Jersey CS Dept., Rutgers University, New JerseyView Profile , Michael Hsiao ECE Dept., Rutgers University, New Jersey ECE Dept., Rutgers University, New JerseyView Profile Authors Info & Claims ISLPED '01: Proceedings of the 2001 international symposium on Low power electronics and designAugust 2001 Pages 275–278https://doi.org/10.1145/383082.383165Online:06 August 2001Publication History 39citation533DownloadsMetricsTotal Citations39Total Downloads533Last 12 Months9Last 6 weeks0 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access
Chung-Hsing Hsu, Ulrich Kremer, Michael S. Hsiao
ISLPED2
2000 A Static Study of Java Exceptions Using JESP
Barbara G. Ryder, Donald Smith, Ulrich Kremer
CC3
1999 Evaluation of Algorithms for Local Register Allocation
Vincenzo Liberatore, Martin Farach-Colton, Ulrich Kremer
CC3
1998 Tools and Techniques for Automatic Data Layout: A Case Study
Eduard Ayguadé, Jordi Garcia 0001, Ulrich Kremer
Parallel Comput.3
1998 Automatic Data Layout for Distributed-Memory Machines
abstract
The goal of languages like Fortran D or High Performance Fortran (HPF) is to provide a simple yet efficient machine-independent parallel programming model. After the algorithm selection, the data layout choice is the key intellectual challenge in writing an efficient program in such languages. The performance of a data layout depends on the target compilation system, the target machine, the problem size, and the number of available processors. This makes the choice of a good layout extremely difficult for most users of such languages. If languages such as HPF are to find general acceptance, the need for data layout selection support has to be addressed. We beleive that the appropriate way to provide the needed support is through a tool that generates data layout specifications automatically. This article discusses the design and implementation of a data layout selection tool that generates HPF-style data layout specifications automatically. Because layout is done in a tool that is not embedded in the target compiler and hence will be run only a few times during the tuning phase of an application, it can use techniques such as integer programming that may be considered too computationally expensive for inclusion in production compilers. The proposed framework for automatic data layout selection builds and examines search spaces of candidate data layouts. A candidate layout is an efficient layout for some part of the program. After the generation of search spaces, a single candidate layout is selected for each program part, resulting in a data layout for the entire program. A good overall data layout may require the remapping of arrays between program parts. A performance estimator based on a compiler model, an execution model, and a machine model are needed to predict the execution time of each candidate layout and the costs of possible remappings between candidate data layouts. In the proposed framework, instances of NP-complete problems are solved during the construction of candidate layout search spaces and the final selection of candidate layouts from each search space. Rather than resorting to heuristics, the framework capitalizes on state-of-the-art 0-1 integer programming technology to compute optimal solutions of these NP-complete problems. A prototype data layout assistant tool based on our framework has been implemented as part of the D system currently under development at Rice University. The article reports preliminary experimental results. The results indicate that the framework is efficient and allows the generation of data layouts of high quality.
Ken Kennedy, Ulrich Kremer
ACM Trans. Program. Lang. Syst.2
1995 Automatic Data Layout for High Performance Fortran
abstract
High Performance Fortran (HPF) is rapidly gaining acceptance as a language for parallel programming. The goal of HPF is to provide a simple yet efficient machine independent parallel programming model. Besides the algorithm selection, the data layout choice is the key intellectual step in writing an efficient HPF program. The developers of HPF did not believe that data layouts can be determined automatically in all cases, Therefore HPF requires the user to specify the data layout. It is the task of the HPF compiler to generate efficient code for the user supplied data layout. The choice of a good data layout depends on the HPF compiler used, the target architecture, the problem size, and the number of available processors. Allowing remapping of arrays at specific points in the program makes the selection of an efficient data layout even harder. Although finding an efficient data layout fully automatically may not be possible in all cases. HPF users will need support during the data layout selection process. In particular, this support is necessary if the user is not familiar with the characteristics of the target HPF compiler and target architecture, or even with HPF itself. Therefore, tools for automatic data layout and performance estimation will be crucial if the HPF is to find general acceptance in the scientific community. This paper discusses a framework for automatic data layout for use in a data layout assistant tool for a data-parallel language such as HPF. The envisioned tool can be used to generate a first data layout for a sequential Fortran program without data layout statements, or to extend a partially specified data layout in a HPF program to a totally specified data layout. Since the data layout assistant is not embedded in a compiler and will run only a few times during the tuning process of an application program, the framework can use techniques that may be too computationally expensive to be included in a compiler. A prototype data layout assistant tool based on our framework has been implemented as part of the D system currently under development at Rice University. The paper reports preliminary experimental results. The results indicate that the framework is efficient and generates data layouts of high quality.
Ken Kennedy, Ulrich Kremer
SC2
1991 A Static Performance Estimator to Guide Data Partitioning Decisions
abstract
The choice of the data domain partitioning scheme is an important factor in determining the available parallelism and hence the performance of an application on a distributed memory multiprocessor.In this paper, we present a performance estimator for statically evaluating the relative efficiency of different data partitioning schemes for any given program on any given distributed memory multiprocessor.Our methlod is not based on a theoretical machine model, but ixnstead uses a set of kernel routinea to "train" the estimator for each target machine.We also describe a prototype implementation of this technique and discuss an experimental evaluation of its accuracy.
Vasanth Balasundaram, Geoffrey C. Fox, Ken Kennedy, Ulrich Kremer
PPoPP4
1989 The parascope editor: an interactive parallel programming tool
abstract
The ParaScope project is building an integrated collection of tools to help scientific programmers develop correct and efficient parallel programs. The centerpiece of this collection is the ParaScope Editor, an intelligent interactive editor for parallel FORTRAN programs. The ParaScope Editor displays data dependencies, which correspond to potential data races among the iterations of a parallel loop, to assist the user in determining the correctness of a proposed parallelization. In addition, it uses dependencies to support a variety of program transformations selectable by the programmer.
Vasanth Balasundaram, Ken Kennedy, Ulrich Kremer, Kathryn S. McKinley, Jaspal Subhlok
SC3
1988 Advanced tools and techniques for automatic parallelization
Ulrich Kremer, Heinz-J. Bast, Michael Gerndt, Hans P. Zima
Parallel Comput.1