Sheikh K. Ghafoor

dblp:205/4703 · also Sheikh Ghafoor 0001, Sheikh Khaled Ghafoor · DBLP profile ↗
← Back
26ranked-venue papers
3as first author
13since 2021 · last 2025
0000-0001-7175-1619ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Human-computer interaction and ubiquitous computing · 11 · 3 first-author · 4 since 2021Systems, architecture and hardware · 8 · 4 since 2021Artificial intelligence and machine learning · 2Databases, data management, data science and information retrieval · 2Applied, interdisciplinary, general and emerging computing · 2
YearPublicationVenuePosition
2025 Enabling Elasticity in Scientific Workflows for High-Performance Computing Systems
Rajat Bhattarai, Howard Pritchard, Sheikh K. Ghafoor
Euro-Par (1)3
2024 Toward Profiling IoT Processes for Remote Service Attestation
abstract
The Internet of Things (IoT) is ubiquitous in modern life and is being used very widely in industrial control systems, smart grids, home appliances and many more. IoT devices are used to get information from sensors, process information, and send signals to actuators and controllers. In general these devices form a distributed computing network while in operation. Malware in IoT or any embedded devices is a potential security threat. Detecting malware in such a setting while in operation is non-trivial, because these low power devices may not have the computational ability to perform traditional security operations. Additionally, an infected device may cause other machines to misbehave by interfering with the data they receive. Remote Attestation is a security service designed to detect an infection in a device well before the malware detonates. Recent works have turned their attention to service attestation, or attesting the service that a network provides, rather than the individual devices themselves. Traditional remote attestation schemes use cryptographic hashing algorithms as evidence, but this approach generates exponentially more hashes as heterogeneous IoT devices are added to the network and their jobs’ complexity increases. In this work, we propose an approach to collect the contents of executable virtual memory from an IoT device. We develop a protocol based on our approach that can build a profile of a process running on an IoT device, such evidence can be analyzed automatically with high granularity. We validate our protocol by testing on both a personal computer, and a real-world Industrial IoT device under process injection attacks. Our results show that our protocol will be able to detect small changes to process memory over time, and that an injection as small as one word can be detected and read.
William A. Johnson, John Housley, Sheikh K. Ghafoor, Stacy J. Prowell
ISPDC3
2024 Design and Development of XiveNet: A Hybrid CAN Research Testbed
abstract
We have developed an affordable distributed Internet of Things (IoT) testbed, named XiveNet, to conduct in-vehicle security research. This testbed merges the adaptability of simulators with the real-time ECU characteristics of actual vehicles. The testbed is made up of ECU chips found in vehicles, Raspberry Pis, and is combined with a bus master simulator. Our experiments with CAN (controller area network) traffic from actual vehicles (Oak Ridge National Laboratories Road Data Set) demonstrate that our testbed closely replicates the attributes of a real vehicle. We have further authenticated our testbed by deploying SecCAN, a secure CAN algorithm, and evaluating its security by injecting invalid frames. Furthermore, we examined ORNL’s timing-based intrusion detection on our testbed and successfully produced alerts. Additionally, we incorporated Named Data Networking (NDN) capable nodes, providing researchers with an additional resource to develop future in-vehicle security solutions. Finally, we have proposed a bitrate hopping technique focused on preventing the denial of service attack and conducted a preliminary investigation using the testbed. Our evaluation and validation indicate that the testbed provides the real-world vehicle environment with the flexibility of a simulation environment that supports a wide range of hardware and software configurations.
William Luke Lambert, Sheikh K. Ghafoor, Haley Burnell, Brennan Huber, Farah I. Kandah, Anthony Skjellum
ISPDC2
2024 Integrating Parallel and Distributed Computing in Early Computing Classes
abstract
Parallel and distributed computing (PDC) has become pervasive in all aspects of computing, so it is essential that students include parallelism and distribution in the computational thinking that they apply to problem solving, from the very beginning of their computing education. With all computing devices that students use currently having multiple cores as well as a GPU in many cases, many students' favorite applications use multiple cores and/or distributed processors. However, we are still teaching them to solve problems using only sequential thinking. Why?
Alan Sussman, Sushil K. Prasad, Charles C. Weems, Sheikh K. Ghafoor, Ramachandran Vaidyanathan
SIGCSE (2)4
2023 A Performance Prediction Model for Structured Grid Based Applications in HPC Environments
abstract
Predicting the performance of parallel applications at scale is a challenging problem. We have developed a performance prediction model for structured grid-based scientific applications for High Performance Computing systems. Our model can capture the system complexity and consider computation and communication attributes of application performance on HPC architectures. We have also proposed a methodology for obtaining the realistic value of the model parameters by small-scale sample runs of an application on the target system. We have used our model to predict the performance of an actual application (2D Flood Simulation) and a synthetic application (Game of Life) on the Summit supercomputer at Oak Ridge National Laboratory and Stampede2 at Texas Advanced Computing Center. The experimental results indicate that our model’s predictive performance is acceptable, and our model was able to predict the performance of these two applications on Summit and Stampede2 with more than 90% accuracy.
Md Bulbul Sharif, Thomas M. Hines, Sheikh K. Ghafoor, Mario Morales-Hernández, Tigstu T. Dullo, Alfred J. Kalyanapu
ISPDC3
2023 Integrating Parallel and Distributed Computing in Early Computing Classes
abstract
Parallel and distributed computing (PDC) has become pervasive in all aspects of computing, and thus it is essential that students include parallelism and distribution in the computational thinking that they apply to problem solving, from the very beginning. Computer science education is still teaching to a 20th century model of algorithmic problem solving. Sequence, branch, and loop are taught in our early courses as the only organizing principles needed for algorithms, and we invest considerable time in showing how best to sequentially process large volumes of data. All computing devices that students use currently have multiple cores as well as a GPU in many cases. Most of their favorite applications use multiple cores and numbers of distributed processors. Often concurrency offers simpler solutions than sequential approaches. Industry is desperate for software engineers who think naturally in terms of exploiting these capabilities, rather than seeing them as an exotic upper-level topic that gets layered over a sequential solution. However, we are still teaching students to solve problems using sequential thinking. In this workshop we overview key PDC concepts and provide examples of how they may naturally be incorporated in early computing classes. We will introduce plugged and unplugged curriculum modules that have been successfully integrated in existing computing classes at multiple institutions. We will highlight the upcoming summer training workshop, for which we have funding to support attendance, as well as other CDER (Center for Parallel and Distributed Computing Curriculum Development and Educational Resources) activities.
Sheikh K. Ghafoor, Charles C. Weems, Alan Sussman, Ramachandran Vaidyanathan, Sushil K. Prasad
SIGCSE (2)1
2023 NSF/IEEE-TCPP Curriculum on Parallel and Distributed Computing for Undergraduates - Version II - Big Data, Energy, and Distributed Computing
abstract
This special session will report on the updated NSF/IEEE-TCPP Curriculum on Parallel and Distributed Computing released in Nov 2020 by the Center for Parallel and Distributed Computing Curriculum Development and Educational Resources (CDER). The purpose of the special session is to obtain SIGCSE community feedback on this curriculum in a highly interactive manner employing the hybrid modality and supported by a full-time CDER booth for the duration of SIGCSE. In this era of big data, cloud, and multi- and many-core systems, it is essential that the computer science (CS) and computer engineering (CE) graduates have basic skills in parallel and distributed computing (PDC). The topics are primarily organized into the areas of architecture, programming, and algorithms topics. A set of pervasive concepts that percolate across area boundaries are also identified. Version 1 of this curriculum was released in December 2012. That curriculum guideline has over 140 early adopter institutions worldwide and has been incorporated into the 2013 ACM/IEEE Computer Science curricula. This Version-II represents a major revision. The updates have focused on enhancing coverage related to the topical aspects of Big Data, Energy, and Distributed Computing.
Sushil K. Prasad, Charles C. Weems, Alan Sussman, Trilce Estrada, Ramachandran Vaidyanathan, Sheikh K. Ghafoor, Krishna Kant 0001, Craig B. Stunkel
SIGCSE (2)7
2023 Design of a portable implementation of partitioned point-to-point communication primitives
abstract
Abstract The Message Passing Interface (MPI) has been the dominant message passing solution for scientific computing for decades. MPI point‐to‐point communications are highly efficient mechanisms for process‐to‐process communication. However, MPI performance when processes utilize multiple threads is slowed by concurrency protections in the MPI library. MPI's current thread level interface imposes these overheads throughout the library when thread safety is needed. While much work has been done to reduce multithreading overheads in MPI, a solution is needed that reduces the number of messages exchanged in a threaded environment. Partitioned communication is included in the MPI 4.0 standard as an alternative that addresses the challenges of multithreaded communication in MPI today. Partitioned communication reduces overall message volume by creating a buffer‐sharing mechanism between threads such that they can indicate when portions of a communication buffer are available to be sent. Separation of the control and data planes in MPI is enabled by allowing persistent initialization and single occurrence message buffer matching from the indication that the data is ready to be sent. This enables the usage of underlying hardware primitives like triggered operations, where commands (destination, size, etc.) can be set up prior to data buffer readiness and readiness triggered with a simple doorbell/counter later. This approach is useful for future development of MPI operations in environments where traditional networking commands can have performance challenges, like accelerators (GPUs, FPGAs). In this paper, we detail the design and implementation of a layered library (built on top of MPI‐3.1) and an integrated Open MPI solution that supports the new, MPI‐4.0 partitioned communication feature set. The library will enable applications to use currently released MPI implementations and older legacy libraries to provide partitioned communication support while also enabling further exploration of this new communication model in new applications and use cases. We will compare the designs of the library and native Open MPI support, provide performance results and comparisons between the two approaches, and lessons learned from the implementation of partitioned communication in both library and native forms. We find that the native implementation and library have similar performance with a percentage difference under 0.94% in microbenchmarks and performance within 5% for a partitioned communication enabled proxy application.
W. Pepper Marts, Andrew Worley, Prema Soundarajan, Derek Schafer, Matthew G. F. Dosanjh, Ryan E. Grant, Purushotham V. Bangalore, Anthony Skjellum, Sheikh K. Ghafoor
Concurr. Comput. Pract. Exp.9
2022 APPFIS: An Advanced Parallel Programming Framework for Iterative Stencil Based Scientific Applications in HPC Environments
abstract
Developing performant parallel applications for the distributed environment is challenging and requires expertise in both the HPC system and the application domain. We have developed a C++-based framework called APPFIS that hides the system complexities by providing an easy-to-use interface for developing performance portable structured grid-based stencil applications. APPFIS’s user interface is hardware agnostic and provides partitioning, code optimization, and automatic communication for stencil applications in distributed HPC environment. In addition, it offers straightforward APIs for utilizing multiple GPU accelerators, shared memory, and node-level parallelizations with automatic optimization for computation and communication overlapping. We have tested the functionality and performance of APPFIS using several applications on three platforms (Stampede2 at Texas Advanced Computing Center, Bridges-2 at Pittsburgh Supercomputing Center, and Summit Supercomputer at Oak Ridge National Laboratory). Experimental results show comparable performance to hand-tuned code with an excellent strong and weak scalability up to 4096 CPUs and 384 GPUs.
Md Bulbul Sharif, Sheikh K. Ghafoor
ISPDC2
2022 Scheduling of Elastic Message Passing Applications on HPC Systems
Debolina Halder Lina, Sheikh K. Ghafoor, Thomas M. Hines
JSSPP2
2022 Integrating Parallel and Distributed Computing in Early CS Courses
abstract
Parallel and distributed computing (PDC) has become pervasive in all aspects of computing, and thus it is essential that students include parallelism and distribution in the computational thinking that they apply to problem solving, from the very beginning. Computer science education is still teaching to a 20th century model of algorithmic problem solving. Sequence, branch, and loop are taught in our early courses as the only organizing principles needed for algorithms, and we invest considerable time in showing how best to sequentially process large volumes of data. All computing devices that students use currently have multiple cores as well as GPU in many cases. Most of their favorite applications use multiple cores and numbers of distributed processors. Often concurrency offers simpler solutions than sequential approaches. ACM and ABET have recommended including PDC in the undergraduate CS curriculum. However, we are still teaching them to solve problems using sequential thinking. In this workshop we overview the key PDC concepts and provide examples of how they may naturally be incorporated in early CS classes. We will introduce plugged and unplugged curriculum modules that have been successfully integrated in existing CS classes at multiple institutions. We will highlight the upcoming summer training that we are organizing, for which we have funding to support attendance.
Sheikh K. Ghafoor, Sushil K. Prasad, Charles C. Weems
SIGCSE (2)1
2022 Keeping up with technology: Teaching parallel, distributed, and high-performance computing
Sushil K. Prasad, Sheikh K. Ghafoor, Martina Barnas, Felix Wolf 0001, Erik Saule, Noemi de La Rocque Rodriguez, Rizos Sakellariou
J. Parallel Distributed Comput.2
2021 Implementation and evaluation of MPI 4.0 partitioned communication libraries
Matthew G. F. Dosanjh, Andrew Worley, Derek Schafer, Prema Soundararajan, Sheikh K. Ghafoor, Anthony Skjellum, Purushotham V. Bangalore, Ryan E. Grant
Parallel Comput.5
2020 Toward High Performance Computing Education
abstract
High Performance Computing (HPC) is the ability to process data and perform complex calculations at extremely high speeds. Current HPC platforms can achieve calculations on the order of quadrillions of calculations per second with quintillions on the horizon. The past three decades witnessed a vast increase in the use of HPC across different scientific, engineering and business communities, for example, sequencing the genome, predicting climate changes, designing modern aerodynamics, or establishing customer preferences. Although HPC has been well incorporated into science curricula such as bioinformatics, the same cannot be said for most computing programs. This working group will explore how HPC can make inroads into computer science education, from the undergraduate to postgraduate levels. The group will address research questions designed to investigate topics such as identifying and handling barriers that inhibit the adoption of HPC in educational environments, how to incorporate HPC into various curricula, and how HPC can be leveraged to enhance applied critical thinking and problem solving skills. Four deliverables include: (1) a catalog of core HPC educational concepts, (2) HPC curricula for contemporary computing needs, such as in artificial intelligence, cyberanalytics, data science and engineering, or internet of things, (3) possible infrastructures for implementing HPC coursework, and (4) HPC-related feedback to the CC2020 project.
Rajendra K. Raj, Carol J. Romanowski, Sherif G. Aly 0001, Brett A. Becker, Juan Chen 0001, Sheikh K. Ghafoor, Nasser Giacaman, Steven Gordon 0001, Cruz Izu, Nick Rahimi, Michael P. Robson, Neena Thota
ITiCSE6
2019 User-Level Scheduled Communications for MPI
abstract
Composability is one of seven reasons for the long-standing and continuing success of MPI. Extending MPI by composing its operations with user-level operations provides useful integration with the progress engine and completion notification methods of MPI. However, the existing extensibility mechanism in MPI (generalized requests) is not widely utilized and has significant drawbacks. MPI can be generalized via scheduled communication primitives, for example, by utilizing implementation techniques from existing MPI-3 nonblocking collectives and from forthcoming MPI-4 persistent and partitioned APIs. Non-trivial schedules are used internally in some MPI libraries; but, they are not accessible to end-users. Message-based communication patterns can be built as libraries on top of MPI. Such libraries can have comparable implementation maturity and potentially higher performance than MPI library code, but do not require intimate knowledge of the MPI implementation. Libraries can provide performance-portable interfaces that cross MPI implementation boundaries. The ability to compose additional user-defined operations using the same progress engine benefits all kinds of general purpose HPC libraries. We propose a definition for MPI schedules: a user-level programming model suitable for creating persistent collective communication composed with new application-specific sequences of user-defined operations managed by MPI and fully integrated with MPI progress and completion notification. The API proposed offers a path to standardization for extensible communication schedules involving user-defined operations. Our approach has the potential to introduce event-driven programming into MPI (beyond the tools interface), although connecting schedules with events comprises future work. Early performance results described here are promising and indicate strong overlap potential.
Derek Schafer, Sheikh K. Ghafoor, Daniel J. Holmes, Martin Ruefenacht, Anthony Skjellum
HiPC2
2019 Unplugged Activities to Introduce Parallel Computing in Introductory Programming Classes: an Experience Report
abstract
Learning programming in early introductory classes is challenging for first year university students, and introducing parallel programming (PDC) in early classes along with traditional sequential programming is even more challenging. Unplugged activities may help alleviate some of the difficulties for students. Unplugged activities have been shown to increase student interest, and to enhance student understanding of CS programming concepts. We have used unplugged activities to teach PDC concepts before introducing parallel programming. Our experiences show that using unplugged activities to introduce the PDC concepts reduce the barrier to learn parallel programming.
Sheikh K. Ghafoor, David W. Brown, Mike Rogers, Thomas M. Hines
ITiCSE1
2019 Modernizing Early CS Courses with Parallel and Distributed Computing
abstract
Parallel and distributed computing (PDC) is now a pervasive aspect of deployed systems, and thus it is essential that students include parallelism and distribution in the computational thinking that they apply to problem solving, from the very beginning. Our students all have multicore laptops. Most of their favorite applications use vast numbers of distributed processors. Why are we still teaching them to solve problems using only sequential thinking? Come to this workshop to see how easy it is to open their eyes to exploiting concurrency in problem solving, starting in their earliest courses. You'll hear about and experience some unplugged activities, learn how to help students recognize examples of concurrency in the world around them, see how event driven user interfaces can easily exemplify issues related to multithreading, and how freely available libraries can be used to naturally exploit parallelism in working with large data structures. We will also highlight the two summer training programs that we are organizing, for which we have funding to support attendance by instructors. Having a laptop that can run Java and C++ will allow you to follow along with some code examples, but isn't necessary.
Sushil K. Prasad, Sheikh K. Ghafoor, Charles C. Weems, Alan Sussman
SIGCSE2
2018 A Flexible-blocking Based Approach for Performance Tuning of Matrix Multiplication Routines for Large Matrices with Edge Cases
abstract
Efficient and scalable matrix operations are being highly demanding in the recent era of Machine Learning, Deep Learning, and Big Data Analytics. The two commonly used matrix-matrix operations in the Basic Linear Algebra Subprograms (BLAS) specification are General Matrix-Matrix multiplication (GEMM) and Symmetric Rank-k update (SYRK). The SYRK routine is a specialization of the GEMM routine, where half of the multiplications are skipped as the resultant matrix is known to be symmetric. Fortunately, several linear algebra libraries implement these BLAS routines quite efficiently. The libraries usually partition the input matrices into blocks and place them in processor caches, thus improving performance by leveraging the caches. However, the contemporary libraries are highly optimized for squarish matrices, but the performance degrades significantly for the matrices with edge case (strictly thin or strictly fat shapes) in the multicore machine. The primary reason is that the current state-of-the-art libraries make fixed block shapes based on a processor architecture, and do not consider the shape of the input matrices. In this paper, we propose a new blocking approach, we name it Flexible-blocking, to mitigate the scalability issues. In contrast to the contemporary libraries, our approach formulates the blocks of the input matrices based on the shapes of the matrices as well as the number of threads used in the implementation. Our proposed technique shows noticeable performance improvement on multicore shared-memory machines for the edge case matrices.
Md Mosharaf Hossain, Thomas M. Hines, Sheikh K. Ghafoor, Sheikh Rabiul Islam, Ramakrishnan Kannan, Sreenivas R. Sukumar 0001
IEEE BigData3
2018 Mining Illegal Insider Trading of Stocks: A Proactive Approach
abstract
Illegal insider trading of stocks is based on releasing non-public information (e.g., new product launch, quarterly financial report, acquisition or merger plan) before the information is made public. Detecting illegal insider trading is difficult due to the complex, nonlinear, and non-stationary nature of the stock market. In this work, we present an approach that detects and predicts illegal insider trading proactively from large heterogeneous sources of structured and unstructured data using a deep-learning based approach combined with discrete signal processing on the time series data. In addition, we use a tree-based approach that visualizes events and actions to aid analysts in their understanding of large amounts of unstructured data. Using existing data, we have discovered that our approach has a good success rate in detecting illegal insider trading patterns.
Sheikh Rabiul Islam, Sheikh K. Ghafoor, William Eberle
IEEE BigData2
2018 Instruction of introductory programming course using multiple contexts
abstract
This paper describes the experience of redesigning a traditional CS1 programming course, utilizing traditional coding practices as well as microcontroller units (MCU) based coding, to provide multiple programming environments. The objective of this redesign is to improve the programming skills for engineering students by 1) providing them with program development experience in multiple contexts and 2) relating the initial programming experience to the typical notion of engineering through significant hardware experience. Typical CS1 courses are designed with an instructor led lecture focusing on the introduction of specific computer skills and languages while programming assignments and laboratories help strengthen these skills in the students. For this remodeling, in addition to the typical programming exercises, supplementary MCU based lab exercises were used to provide an additional, different programming target for increased learning and highlighting the complementary relationship between hardware and software. The outcomes of this effort demonstrate that the addition of a MCU to an introductory programming course can work as an effective motivator, providing the students with a secondary context to reinforce programming skills developed during the course, and that providing multiple contexts (traditional desktop programming and hardware-based programming) together can aid in learning and the transfer of knowledge.
David W. Brown, Sheikh K. Ghafoor, Stephen L. Canfield
ITiCSE2
2018 CReST-Security Knitting Kit: Readily Available Teaching Resources to Integrate Security Topics into Traditional CS Courses (Abstract Only)
abstract
Since security education is not required in CS curriculum, many CS undergraduates can successfully achieve their degree without being exposed to any security courses during their course of study and enter the digital workforce with no knowledge or basic understanding of information security -- one of the essential skill sets for the 21st century. To address this concern, Information Assurance and Security (IAS) has been designated as a new knowledge area in the new ACM/IEEE-CS Curricula 2013. This workshop empowers CS faculty to access and use freely available resources to integrate security in to their CS curriculum will help institutions to meet ACM/IEEE-CS guideline. With support from NSF (Award# DUE-1140864, #1438861), at the CyberSecurity Education, Research and Outreach Center at Tennessee Tech, we have developed a set of readily available resources called SecKnitKit (Security Knitting Kit, www.secknitkit.org), which offers a suite of instructional material for non-security faculty (faculty whose primary teaching/research focus is not security) to integrate security in upper division CS courses such as operating systems, software engineering, computer networks and databases. Resources include lecture slides with notes, assessment questions and homework/classroom assignments with all details and technical support. The participants will receive access to all SecKnitKit materials (instructional and assessment) of interest and demonstrated use of the active learning exercises. There are six participant slots for each of the four courses mentioned above and participants will have an option to select their courses of choice at registration time.
Ambareen Siraj, Sheikh K. Ghafoor
SIGCSE2
2018 VSI: Edu*-2016 - Keeping up with technology: Teaching parallel, distributed and high-performance computing
Sushil K. Prasad, Sheikh K. Ghafoor, Christos Kaklamanis, Ramachandran Vaidyanathan
J. Parallel Distributed Comput.2
2016 CReST-Security Knitting Kit: Ready to Use Teaching Resources to Embed Security Topics into Upper Division CS Courses (Abstract Only)
abstract
Information Assurance and Security has been designated as a new knowledge area in the new ACM/IEEE-CS Curricula 2013. This is not a trivial task to accomplish, especially with lack of resources. With support from NSF (Award# DUE-1140864), we have developed a set of readily available resources called SecKnitKit (Security Knitting Kit, www.secknitkit.org), which offers a suite of instructional material for non-security faculty (faculty whose primary teaching/research focus is not security) to integrate security in upper division CS courses such as operating systems, software engineering, computer networks and databases. As part of the NSF CReST (CyberWorkshops: Resources and Strategies for Teaching Cybersecurity in Computer Science, DUE-1438861, www.crest4cs.org) project, this workshop will introduce CS faculty to the SecKnitKit resources that can be easily adaptable into any standard CS curriculum. The participants will receive access to all SecKnitKit materials (instructional and assessment) of interest and demonstrated use of the active learning exercises. Each participant will receive $125.00 stipend for his/her time. Requires a Windows or Mac laptop. Enrollment is limited to 28 participants who teach at least one of these courses: operating systems, software engineering, computer networks and databases).
Ambareen Siraj, Sheikh K. Ghafoor
SIGCSE2
2014 Empowering faculty to embed security topics into computer science courses
abstract
Security illiteracy is a very common problem among Computer Science (CS) graduates entering the nation's digital workforce, which has contributed to a national cyber-infrastructure that could and should be more resilient to cyber-enemies than it is now. The Security Knitting Kit (SecKnitKit) project aims to improve security awareness, knowledge, and interest of undergraduate CS students by exposing them to computer security concepts and issues in their regular course of study. The project is developing, deploying, and disseminating a multi-faceted out-of-the-box instructional support system to empower non-security faculty. These are faculty who have no experience in teaching security but recognize the importance of security in today's world and want to broaden their teaching repertoire. This project enables them to weave relevant security topics traditional computer science courses seamlessly and effectively. The project is organized by the CS department at Tennessee Tech University (TTU) and supported by the National Science Foundation under grant DUE-1140864.
Ambareen Siraj, Sheikh K. Ghafoor, Joshua Tower, Ada Haynes
ITiCSE2
2001 Experiences from integrating algorithmic and systemic load balancing strategies
abstract
Abstract Load balancing increases the efficient use of existing resources for parallel and distributed applications. At a coarse level of granularity, advances in runtime systems for parallel programs have been proposed in order to control available resources as efficiently as possible by utilizing idle resources and using task migration. Simultaneously, at a finer granularity level, advances in algorithmic strategies for dynamically balancing computational loads by data redistribution have been proposed in order to respond to variations in processor performance during the execution of a given parallel application. Combining strategies from each level of granularity can result in a system which delivers advantages of both. The resulting integration is systemic in nature and transfers the responsibility of efficient resource utilization from the application programmer to the runtime system. This paper presents the design and implementation of a system that combines an algorithmic fine‐grained data parallel load balancing strategy with a systemic coarse‐grained task‐parallel load balancing strategy, and reports on recent experimental results of running a computationally intensive scientific application under this integrated system. The experimental results indicate that a distributed runtime environment which combines both task and data migration can provide performance advantages with little overhead. It also presents proposals for performance enhancements of the implementation, as well as future explorations for effective resource management. Copyright © 2001 John Wiley & Sons, Ltd.
Ioana Banicescu, Sheikh K. Ghafoor, Vijay Velusamy, Samuel H. Russ, Mark Bilderback
Concurr. Comput. Pract. Exp.2
1998 Hectiling: An Integration of Fine and Coarse-Grained Load-Balancing Strategies
abstract
General-purpose programmers have come to expect a high degree of portability among widely varying architectures. Advances in run-time systems for parallel programs have been proposed in order to harness available resources as efficiently as possible. Simultaneously, advances in algorithmic methods of dynamically balancing computational load have been proposed in order to respond to variations in actual performance and therefore in run-time. The primary mechanism for harnessing idle resources effectively, task migration, can be used alongside the primary mechanism for dynamic load balancing, data redistribution. Besides the fact that the two methods can be used simultaneously to spur further increases in performance, the run-time information-gathering infrastructure necessary to detect and use idle resources can also benefit dynamically load-balanced applications. This paper describes an architecture for and preliminary implementation of a system that combines data-parallel load balancing with task-parallel load balancing. Performance test results are included as well.
Samuel H. Russ, Ioana Banicescu, Sheikh K. Ghafoor, Bharathi Janapareddi, Jonathan Robinson
HPDC3