Mustafa M. Tikir

dblp:21/236 · DBLP profile ↗
← Back
11ranked-venue papers
7as first author
0since 2021 · last 2011
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 4 first-authorSoftware engineering, systems software and programming languages · 4 · 3 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
3 papers
High-performance computing · 46% Performance modeling and evaluation · 32% Memory systems · 22%
Software engineering, system software, and programming languages
2 papers
Program analysis · 58% Software testing · 42%

Topics — the 14 heaviest of 14, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
High-performance computing
performance optimization at scale
0.122008
High-frequency simulations of global seismic wave propagation using SPECFEM3D_GLOBE on 62K processors · SC 2008
A genetic algorithms approach to modeling the performance of memory-bound computations · SC 2007
High-performance computing › scientific computing systems
earthquake simulation
0.112008
High-frequency simulations of global seismic wave propagation using SPECFEM3D_GLOBE on 62K processors · SC 2008
High-performance computing
scientific computing systems
0.112008
High-frequency simulations of global seismic wave propagation using SPECFEM3D_GLOBE on 62K processors · SC 2008
High-performance computing › finite element method
spectral-element method
0.112008
High-frequency simulations of global seismic wave propagation using SPECFEM3D_GLOBE on 62K processors · SC 2008
Performance modeling and evaluation
benchmarking
0.112007
A genetic algorithms approach to modeling the performance of memory-bound computations · SC 2007
Memory systems
memory-bound computation
0.112007
A genetic algorithms approach to modeling the performance of memory-bound computations · SC 2007
Performance modeling and evaluation
performance prediction
0.112007
A genetic algorithms approach to modeling the performance of memory-bound computations · SC 2007
Performance modeling and evaluation › performance monitoring
hardware performance counters
0.012004
Using Hardware Counters to Automatically Improve Memory Performance · SC 2004
Memory systems
memory management
0.012004
Using Hardware Counters to Automatically Improve Memory Performance · SC 2004
Memory systems › virtual memory management
page migration
0.012004
Using Hardware Counters to Automatically Improve Memory Performance · SC 2004
Performance modeling and evaluation
profiling
0.012004
Using Hardware Counters to Automatically Improve Memory Performance · SC 2004
Software testing › test coverage
coverage-based testing
0.012002
Efficient instrumentation for code coverage testing · ISSTA 2002
Program analysis › dynamic analysis
dynamic instrumentation
0.012002
Efficient instrumentation for code coverage testing · ISSTA 2002
Program analysis › dynamic analysis
runtime instrumentation
0.012004
Using Hardware Counters to Automatically Improve Memory Performance · SC 2004

Methods — techniques the papers use, named apart from their topics

hardware counters · 0.1dyninst runtime instrumentation · 0.1spectral-element method · 0.1genetic algorithm · 0.1STREAM · 0.1MultiMAPS · 0.1Apex-MAPS · 0.1dominator tree analysis · 0.0
YearPublicationVenuePosition
2011 Reducing Energy Usage with Memory and Computation-Aware Dynamic Frequency Scaling
Michael Laurenzano, Mitesh R. Meswani, Laura Carrington, Allan Snavely, Mustafa M. Tikir, Stephen W. Poole
Euro-Par (1)5
2011 An idiom-finding tool for increasing productivity of accelerators
abstract
Suppose one is considering purchase of a computer equipped with accelerators. Or suppose one has access to such a computer and is considering porting code to take advantage of the accelerators. Is there a reason to suppose the purchase cost or programmer effort will be worth it? It would be nice to able to estimate the expected improvements in advance of paying money or time. We exhibit an analytical framework and tool-set for providing such estimates: the tools first look for user-defined idioms that are patterns of computation and data access identified in advance as possibly being able to benefit from accelerator hardware. A performance model is then applied to estimate how much faster these idioms would be if they were ported and run on the accelerators, and a recommendation is made as to whether or not each idiom is worth the porting effort to put them on the accelerator and an estimate is provided of what the overall application speedup would be if this were done.
Laura Carrington, Mustafa M. Tikir, Catherine Mills Olschanowsky, Michael Laurenzano, Joshua Peraza, Allan Snavely, Stephen W. Poole
ICS2
2010 PEBIL: Efficient static binary instrumentation for Linux
abstract
Binary instrumentation facilitates the insertion of additional code into an executable in order to observe or modify the executable's behavior. There are two main approaches to binary instrumentation: static and dynamic binary instrumentation. In this paper we present a static binary instrumentation toolkit for Linux on the x86/x86_64 platforms, PEBIL (PMaC's Efficient Binary Instrumentation Toolkit for Linux). PEBIL is similar to other toolkits in terms of how additional code is inserted into the executable. However, it is designed with the primary goal of producing efficient-running instrumented code. To this end, PEBIL uses function level code relocation in order to insert large but fast control structures. Furthermore, the PEBIL API provides tool developers with the means to insert lightweight hand-coded assembly rather than relying solely on the insertion of instrumentation functions. These features enable the implementation of efficient instrumentation tools with PEBIL. The overhead introduced for basic block counting by PEBIL is an average of 65% of the overhead of Dyninst, 41% of the overhead of Pin, 15% of the overhead of DynamoRIO, and 8% of the overhead of Valgrind.
Michael Laurenzano, Mustafa M. Tikir, Laura Carrington, Allan Snavely
ISPASS2
2009 PSINS: An Open Source Event Tracer and Execution Simulator for MPI Applications
Mustafa M. Tikir, Michael Laurenzano, Laura Carrington, Allan Snavely
Euro-Par1
2008 High-frequency simulations of global seismic wave propagation using SPECFEM3D_GLOBE on 62K processors
abstract
SPECFEM3D_GLOBE is a spectral-element application enabling the simulation of global seismic wave propagation in 3D anelastic, anisotropic, rotating and self-gravitating Earth models at unprecedented resolution. A fundamental challenge in global seismology is to model the propagation of waves with periods between 1 and 2 seconds, the highest frequency signals that can propagate clear across the Earth. These waves help reveal the 3D structure of the Earth's deep interior and can be compared to seismographic recordings. We broke the 2 second barrier using the 62K processor Ranger system at TACC. Indeed we broke the barrier using just half of Ranger, by reaching a period of 1.84 seconds with sustained 28.7 Tflops on 32K processors. We obtained similar results on the XT4 Franklin system at NERSC and the XT4 Kraken system at University of Tennessee Knoxville, while a similar run on the 28K processor Jaguar system at ORNL, which has better memory bandwidth per processor, sustained 35.7 Tflops (a higher flops rate) with a 1.94 shortest period.
Laura Carrington, Dimitri Komatitsch, Michael Laurenzano, Mustafa M. Tikir, David Michéa, Nicolas Le Goff, Allan Snavely, Jeroen Tromp
SC4
2008 Hardware monitors for dynamic page migration
Mustafa M. Tikir, Jeffrey K. Hollingsworth
J. Parallel Distributed Comput.1
2007 A genetic algorithms approach to modeling the performance of memory-bound computations
abstract
Benchmarks that measure memory bandwidth, such as STREAM, Apex-MAPS and MultiMAPS, are increasingly popular due to the "Von Neumann" bottleneck of modern processors which causes many calculations to be memory-bound. We present a scheme for predicting the performance of HPC applications based on the results of such benchmarks. A Genetic Algorithm approach is used to "learn" bandwidth as a function of cache hit rates per machine with MultiMAPS as the fitness test. The specific results are 56 individual performance predictions including 3 full-scale parallel applications run on 5 different modern HPC architectures, with various CPU counts and inputs, predicted within 10 % average difference with respect to independently verified runtimes.
Mustafa M. Tikir, Laura Carrington, Erich Strohmaier, Allan Snavely
SC1
2005 Efficient online computation of statement coverage
Mustafa M. Tikir, Jeffrey K. Hollingsworth
J. Syst. Softw.1
2004 Using Hardware Counters to Automatically Improve Memory Performance
abstract
In this paper, we introduce a profile-driven online page migration scheme and investigate its impact on the performance of multithreaded applications. We use lightweight, inexpensive plug-in hardware counters to profile the memory access behavior of an application, and then migrate pages to memory local to the most frequently accessing processor. Using the Dyninst runtime instrumentation combined with hardware counters, we were able to add page migration capabilities to the system without having to modify the operating system kernel, or to re-compile application programs. This approach reduced the total number of non-local memory accesses of applications by up to 90%. Even on a system with small remote to local memory access latency rations, this resulted in up to 16% improvement in execution time.
Mustafa M. Tikir, Jeffrey K. Hollingsworth
SC1
2002 Efficient instrumentation for code coverage testing
abstract
Evaluation of Code Coverage is the problem of identifying the parts of a program that did not execute in one or more runs of a program. The traditional approach for code coverage tools is to use static code instrumentation. In this paper we present a new approach to dynamically insert and remove instrumentation code to reduce the runtime overhead of code coverage. We also explore the use of dominator tree information to reduce the number of instrumentation points needed. Our experiments show that our approach reduces runtime overhead by 38-90% compared with purecov, a commercial code coverage tool. Our tool is fully automated and available for download from the Internet.
Mustafa M. Tikir, Jeffrey K. Hollingsworth
ISSTA1
2002 Recompilation for debugging support in a JIT-compiler
abstract
verifiably secure and compact architecture-neutral intermediate format, called Java byte codes. The Java byte codes can be either interpreted by a Java Virtual Machine or translated into native code by Java Just-In-Time compilers. Static Java compilers embed debug information in the Java class files to be used by the source level debuggers. However, the debug information is generated for architecture independent byte codes and most of the debug information is valid only when the byte codes are interpreted. Translating byte codes into native instructions puts a limitation on the amount of usable debug information that can be used by source level debuggers. In this paper, we present a new technique to generate valid debug information when JustIn -Time compilers are used. Our approach is based on the dynamic recompilation of Java methods by a fast code generator and lazily generates debug information when it is required. We also present three implementations for field watch support in the Java Virtual Machine Debugger Interface to investigate the runtime overhead and code size growth by our approach.
Mustafa M. Tikir, Jeffrey K. Hollingsworth, Guei-Yuan Lueh
PASTE1