Mark K. Gardner

dblp:31/4220 · DBLP profile ↗
← Back
24ranked-venue papers
8as first author
1since 2021 · last 2025
0000-0001-5602-8190ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 11 · 4 first-authorComputer networks · 8 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 2 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
5 papers
Distributed systems · 76% Embedded and real-time systems · 10% Parallel and multicore computing · 8%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Bioinformatics and computational biology · 100%
Computer networks
2 papers
Transport protocols and congestion control · 100%

Topics — the 12 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Distributed systems
fault tolerance
0.112010
MOON: MapReduce On Opportunistic eNvironments · HPDC 2010
Distributed systems › grid computing
volunteer computing
0.112010
MOON: MapReduce On Opportunistic eNvironments · HPDC 2010
Distributed systems
grid computing
0.132006
Grid applications - Parallel genomic sequence-searching on an ad-hoc grid: experiences, lessons learned, and implications · SC 2006
Optimizing GridFTP through Dynamic Right-Sizing · HPDC 2003
Dynamic Right-Sizing in FTP (drsFTP): Enhancing Grid Performance in User-Space · HPDC 2002
Bioinformatics and computational biology › sequence analysis › sequence similarity search
genomic sequence search
0.112006
Grid applications - Parallel genomic sequence-searching on an ad-hoc grid: experiences, lessons learned, and implications · SC 2006
Bioinformatics and computational biology › multiple sequence alignment
parallel sequence alignment
0.112006
Grid applications - Parallel genomic sequence-searching on an ad-hoc grid: experiences, lessons learned, and implications · SC 2006
Transport protocols and congestion control › TCP
TCP flow control
0.012002
Dynamic Right-Sizing in FTP (drsFTP): Enhancing Grid Performance in User-Space · HPDC 2002
Parallel and multicore computing › data-parallel programming
mapreduce
0.012010
MOON: MapReduce On Opportunistic eNvironments · HPDC 2010
Embedded and real-time systems
real-time scheduling
0.011997
Bounding Completion Times of Jobs with Arbitrary Release Times, Variable Execution Times and Resource Sharing · IEEE Trans. Software Eng. 1997
Embedded and real-time systems › real-time scheduling
schedulability analysis
0.011997
Bounding Completion Times of Jobs with Arbitrary Release Times, Variable Execution Times and Resource Sharing · IEEE Trans. Software Eng. 1997
High-performance computing › data transfer
bulk data transfer
0.012003
Optimizing GridFTP through Dynamic Right-Sizing · HPDC 2003
High-performance computing › data transfer
wide-area data transfer
0.012002
Dynamic Right-Sizing in FTP (drsFTP): Enhancing Grid Performance in User-Space · HPDC 2002
Embedded and real-time systems › real-time scheduling
priority scheduling
0.011997
Bounding Completion Times of Jobs with Arbitrary Release Times, Variable Execution Times and Resource Sharing · IEEE Trans. Software Eng. 1997

Methods — techniques the papers use, named apart from their topics

parallel BLAST · 0.1buffer auto-tuning · 0.1user-space implementation · 0.1system buffer tuning · 0.1response time analysis · 0.0
YearPublicationVenuePosition
2025 STAGS: A Graph-Sampling Approach for GNN-based Network Anomaly Detection
Saikat Dey, Mark K. Gardner, Jeffry Lang, Wu-chun Feng
Networking2
2019 On the Portability of CPU-Accelerated Applications via Automated Source-to-Source Translation
abstract
Over the past decade, accelerator-based supercomputers have grown from 0% to 42% performance share on the TOP500. Ideally, GPU-accelerated code on such systems should be "write once, run anywhere," regardless of the GPU device (or for that matter, any parallel device, e.g., CPU or FPGA). In practice, however, portability can be significantly more limited due to the sheer volume of code implemented in non-portable languages. For example, the tremendous success of CUDA, as evidenced by the vast cornucopia of CUDA-accelerated applications, makes it infeasible to manually rewrite all these applications to achieve portability. Consequently, we achieve portability by using our automated CUDA-to-OpenCL source-to-source translator called CU2CL. To demonstrate the state of the practice, we use CU2CL to automatically translate three medium-to-large, CUDA-optimized codes to OpenCL, thus enabling the codes to run on other GPU-accelerated systems (as well as CPU- or FPGA-based systems). These automatically translated codes deliver performance portability, including as much as three-fold performance improvement, on a GPU device not supported by CUDA.
Paul Sathre, Mark K. Gardner, Wu-chun Feng
HPC Asia2
2017 SLIM: Enabling Transparent Extensibility and Dynamic Configuration via Session-Layer Abstractions
abstract
Increasingly, communication requires more from the network stack, e.g., seamless handoff and synchronization of state between multiple participants. Due to the lack of support for desired functionality, networking libraries are created to fill the void. This leads to considerable duplication of effort and complicates cross-platform development. Furthermore, the means for extending legacy protocol stacks is largely exhausted (e.g., the TCP options space in the SYN message is mostly allocated), making the addition of future extensions much more challenging. In this paper, we tease apart elements of session management that are currently conflated with the transport semantics in TCP and highlight the need for sessions in contemporary communications. Next, we propose session, flow, and endpoint abstractions that lead to a clearer description of advanced communication models. This effort results in an extensible session-layer intermediary (SLIM) that leverages the above abstractions to support the additional functionality needed by modern applications, such as mobility, communication between two or more participants, and dynamic reconfiguration. SLIM's approach also provides the means for future extensibility of the network stack in a backward-compatible way, thus enabling incremental adoption.
Umar Kalim, Mark K. Gardner, Eric J. Brown
ANCS2
2017 Parallel programming with pictures is a Snap!
Annette C. Feng, Mark K. Gardner, Wu-chun Feng
J. Parallel Distributed Comput.2
2013 Cascaded TCP: Applying pipelining to TCP for efficient communication over wide-area networks
abstract
The bandwidth utilization in traditional TCP protocols (e.g., TCP New Reno) suffers over high-latency and high-bandwidth links due to the inherent characteristics of TCP congestion control. Conventional methods of improving throughput cannot be applied per se for streaming applications. The challenge is exacerbated by “big data” applications such as with the Long Wavelength Array data that is generated at a rate of up to 4 terabytes per hour. To improve bandwidth utilization, we introduce layer-4 relay(s) that enable the pipelining of TCP connections. That is, a traditional end-to-end connection is split into independent streams, each with shorter latencies, that are then concatenated (or cascaded) together to form an equivalent end-to-end TCP connection. This addresses the root cause by decreasing the latency over which the congestion-control protocol operates. To understand when relays are beneficial, we present an analytical model, empirical data and its analyses, to validate our argument and to characterize the impact of latency and available bandwidth on throughput. We also provide insight into how relays may be setup to achieve better bandwidth utilization.
Umar Kalim, Mark K. Gardner, Eric J. Brown, Wu-chun Feng
GLOBECOM2
2013 Seamless Migration of Virtual Machines across Networks
abstract
Current technologies that support live migration require that the virtual machine (VM) retain its IP network address. As a consequence, VM migration is oftentimes restricted to movement within an IP subnet or entails interrupted network connectivity to allow the VM to migrate. Thus, migrating VMs beyond subnets becomes a significant challenge for the purposes of load balancing, moving computation close to data sources, or connectivity recovery during natural disasters. Conventional approaches use tunneling, routing, and layer-2 expansion methods to extend the network to geographically disparate locations, thereby transforming the problem of migration between subnets to migration within a subnet. These approaches, however, increase complexity and involve considerable human involvement. The contribution of our paper is to address the aforementioned shortcomings by enabling VM migration across subnets and doing so with uninterrupted network connectivity. We make the case that decoupling IP addresses from the notion of transport endpoints is the key to solving a host of problems, including seamless VM migration and mobility. We demonstrate that VMs can be migrated seamlessly between different subnets - without losing network state - by presenting a backward-compatible prototype implementation and a case study.
Umar Kalim, Mark K. Gardner, Eric J. Brown, Wu-chun Feng
ICCCN2
2013 Characterizing the challenges and evaluating the efficacy of a CUDA-to-OpenCL translator
Mark K. Gardner, Paul Sathre, Wu-chun Feng, Gabriel Martinez
Parallel Comput.1
2011 Restoring End-to-End Resilience in the Presence of Middleboxes
abstract
The philosophy upon which the Internet was built places the intelligence close to the edge. As the Internet has matured, intermediate devices or middleboxes, such as firewalls or application gateways, have been introduced, thereby weakening the end-to-end nature of the network. As a result, applications must often modify their behavior to accommodate the middleboxes. This is is especially true in the case of transient failure of stateful devices. The failure of a middlebox causes it to lose the state it maintained, causing the failure of the associated TCP connections. Rather than assign the responsibility for recovery to applications, we incorporate a mechanism called an isolation boundary into TCP itself. The isolation boundary maintains a small amount of state across TCP connections, thus enabling reconnection. Furthermore, it does so without breaking backward compatibility with existing TCP. We present an implementation of the isolation boundary in the FreeBSD kernel and demonstrate its backward compatibility with TCP. We quantify the performance impact of the proposed mechanism on the establishment of new and resumed connections for both legacy and extended TCP stacks.
Eric J. Brown, Mark K. Gardner, Umar Kalim, Wu-chun Feng
ICCCN2
2011 CU2CL: A CUDA-to-OpenCL Translator for Multi- and Many-Core Architectures
abstract
The use of graphics processing units (GPUs) in high-performance parallel computing continues to become more prevalent, often as part of a heterogeneous system. For years, CUDA has been the de facto programming environment for nearly all general-purpose GPU (GPGPU) applications. In spite of this, the framework is available only on NVIDIA GPUs, traditionally requiring reimplementation in other frameworks in order to utilize additional multi- or many-core devices. On the other hand, OpenCL provides an open and vendor-neutral programming environment and runtime system. With implementations available for CPUs, GPUs, and other types of accelerators, OpenCL therefore holds the promise of a "write once, run anywhere" ecosystem for heterogeneous computing. Given the many similarities between CUDA and OpenCL, manually porting a CUDA application to OpenCL is typically straightforward, albeit tedious and error-prone. In response to this issue, we created CU2CL, an automated CUDA-to-OpenCL source-to-source translator that possesses a novel design and clever reuse of the Clang compiler framework. Currently, the CU2CL translator covers the primary constructs found in CUDA runtime API, and we have successfully translated many applications from the CUDA SDK and Rodinia benchmark suite. The performance of the automatically translated applications via CU2CL is on par with their manually ported counterparts.
Gabriel Martinez, Mark K. Gardner, Wu-chun Feng
ICPADS2
2010 MOON: MapReduce On Opportunistic eNvironments
abstract
MapReduce offers an ease-of-use programming paradigm for processing large data sets, making it an attractive model for distributed volunteer computing systems. However, unlike on dedicated resources, where MapReduce has mostly been deployed, such volunteer computing systems have significantly higher rates of node unavailability. Furthermore, nodes are not fully controlled by the MapReduce framework. Consequently, we found the data and task replication scheme adopted by existing MapReduce implementations woefully inadequate for resources with high unavailability.
Heshan Lin, Xiaosong Ma, Jeremy S. Archuleta, Wu-chun Feng, Mark K. Gardner, Zhe Zhang 0005
HPDC5
2010 Broadening accessibility to computer science for K-12 education
abstract
Enrollments in computer science and computer engineering have decreased dramatically since the dot-com bubble burst in 2000 even though it is projected that nearly three quarters of all science and engineering jobs in the future will be in these fields. Meeting this demand will require a substantial effort to inspire and motivate students as early as in the elementary school years. The challenge is to provide motivational access to computer science training, particularly for females and minorities in disadvantaged areas.
Mark K. Gardner, Wu-chun Feng
ITiCSE1
2006 When Optical Networking Meets Grid Computing?
abstract
High-speed optical networking infrastructures such as National LambdaRail promise to revolutionize the way that we approach grid computing. With optical-fiber speeds doubling every 9 months and computing speeds doubling only every 18-24 months over the past few decades, we have finally reached a crossroads where network speeds have outstripped the ability of processors to keep up. Thus, we may be entering a new world where the central architectural element in grid computing is the optical network, not the end-host computer(s). The goal of this panel is to discuss whether such a new world is emerging, and if so, what the research challenges will be.
Wu-chun Feng, Mark K. Gardner, Gigi Karmous-Edwards, Jerry Sobieski, Malathi Veeraraghavan
ICCCN2
2006 Grid applications - Parallel genomic sequence-searching on an ad-hoc grid: experiences, lessons learned, and implications
abstract
The Basic Local Alignment Search Tool (BLAST) allows bioinformaticists to characterize an unknown sequence by comparing it against a database of known sequences. The similarity between sequences enables biologists to detect evolutionary relationships and infer biological properties of the unknown sequence.mpiBLAST, our parallel BLAST, decreases the search time of a 300 KB query on the current NT database from over two full days to under 10 minutes on a 128-processor cluster and allows larger query files to be compared. Consequently, we propose to compare the largest query available, the entire NT database, against the largest database available, the entire NT database. The result of this comparison will provide critical information to the biology community, including insightful evolutionary, structural, and functional relationships between every sequence and family in the NT database.Preliminary projections indicated that to complete the above task in a reasonable length of time required more processors than were available to us at a single site. Hence, we assembled GreenGene, an ad-hoc grid that was constructed "on the fly" from donated computational, network, and storage resources during last year's SC|05. GreenGene consisted of 3048 processors from machines that were distributed across the United States. This paper presents a case study of mpiBLAST on GreenGene --- specifically, a pre-run characterization of the computation, the hardware and software architectural design, experimental results, and future directions.
Mark K. Gardner, Wu-chun Feng, Jeremy S. Archuleta, Heshan Lin, Xiaosong Ma
SC1
2004 Re-Architecting Flow Control Adaptation for Grid Environments
abstract
Summary form only given. The performance of TCP in wide-area networks (WANs) is becoming increasingly important with the deployment of computational and data grids. In WAN environments, TCP does not provide good performance for data-intensive applications without the tuning of flow-control buffer sizes. Manual adjustment of buffer sizes is tedious even for network experts. For application scientists, tuning is often an impediment to getting work done. Thus, buffer tuning should be automated. Existing techniques for automatic buffer tuning only measure the bandwidth-delay product (BDP) during connection establishment. This ignores the large fluctuation of the BDP over the lifetime of the connection. In contrast, the dynamic right-sizing algorithm dynamically changes buffer sizes in response to changing network conditions. We describe a new user-space implementation of dynamic right-sizing in FTP (drsFTP) that supports third-party data transfers, a mainstay of scientific computing. In addition to comparing the performance of the new implementation with the old in a WAN-emulated environment, we give performance results over a live WAN. In this particular WAN environment, the new implementation produces transfer rates of up to five times higher than untuned FTP.
Adam Engelhart, Mark K. Gardner, Wu-chun Feng
IPDPS2
2004 User-space auto-tuning for TCP flow control in computational grids
Mark K. Gardner, Sunil Thulasidasan, Wu-chun Feng
Comput. Commun.1
2003 MAGNET: A Tool for Debugging, Analyzing and Adapting Computing Systems
abstract
As computing systems grow in complexity, the cluster and grid communities require more sophisticated tools to diagnose, debug and analyze such systems. We have developed a toolkit called MAGNET (Monitoring Apparatus for General kerNel-Event Tracing) that provides a detailed look at operating-system kernel events with very low overhead. Using the fine-grained information that MAGNET exports from kernel space, challenging problems become amenable to identification and correction. In this paper, we first present the design, implementation and evaluation of MAGNET. Then, we show its use as a diagnostic tool, an online-monitoring tool and a tool for building adaptive applications in clusters and grids.
Mark K. Gardner, Wu-chun Feng, Michael Broxton, Adam Engelhart, Justin Gus Hurwitz
CCGRID1
2003 Optimizing GridFTP through Dynamic Right-Sizing
abstract
In this paper, we describe the integration of dynamic right-sizing - an automatic and scalable buffer management technique for enhancing TCP (transport control protocol) performance - into GridFTP, a subsystem of the Globus Toolkit for managing bulk data transfers across computational Grids. Such Grids are often characterized by networks with large bandwidth-delay products. Unfortunately, many of today's Grid applications use only a small fraction of available bandwidth because the default buffer sizes in TCP are tuned for yesterday's WAN (wide access network) speeds. Buffer sizes can be manually tuned to allow TCP flow control to adapt to high-speed WAN environments, but this is a tedious process. Although recent work has shown how to automatically tune system buffers during connection set-up, these values may not be appropriate for the connection's lifetime due to varying network delay and throughput. We show how using the technique of dynamic right-sizing (DRS) in GridFTP helps us optimize memory usage while maintaining high throughput over the lifetime of the connection. We also show how DRS enhances important GridFTP features such as striped and third-party data transfers in a scalable way. The technique is implemented entirely in user space so that end users do not have to modify the kernel.
Sunil Thulasidasan, Wu-chun Feng, Mark K. Gardner
HPDC3
2003 Automatic Flow-Control Adaptation for Enhancing Network Performance in Computational Grids
Wu-chun Feng, Mark K. Gardner, Mike Fisk, Eric Weigle
J. Grid Comput.2
2002 Dynamic Right-Sizing in FTP (drsFTP): Enhancing Grid Performance in User-Space
abstract
With the advent of computational grids, networking performance over the wide-area network (WAN) has become a critical component in the grid infrastructure. Unfortunately, many high-performance grid applications only use a small fraction of the available bandwidth because operating systems and their associated protocol stacks are still tuned for yesterdays WAN speeds. As a result, network gurus undertake the tedious process of manually, tuning system buffers to allow TCP flow control to scale to today's WAN grid environments. Although recent research has shown how to set the size of these system buffers automatically at connection set-up, the buffer sizes are only appropriate at the beginning of the connection's lifetime. To address these problems, we describe an automated and scalable technique called dynamic right-sizing. We implement this technique in user space (in particular for bulk-data transfer) so that end users do not have to modify the kernel to achieve a significant increase in throughput.
Mark K. Gardner, Wu-chun Feng, Mike Fisk
HPDC1
2002 The MAGNeT Toolkit: Design, Implementation and Evaluation
Wu-chun Feng, Mark K. Gardner, Jeffrey R. Hay
J. Supercomput.2
2001 MAGNeT: monitor for application-generated network traffic
abstract
Over the last decade, network practitioners have focused on monitoring, measuring, and characterizing traffic in the network to gain insight into building critical network components (from the protocol stack to routers and switches to network interface cards). Previous research shows that additional insight can be obtained by monitoring traffic at the application level (i.e., before application-sent traffic is modulated by the protocol stack) rather than in the network (i.e., after it is modulated by the protocol stack). Consequently, this paper describes a monitor for application-generated network traffic (MAGNeT) that captures traffic generated by the application rather than traffic in the network. MAGNeT consists of application programs as well as modifications to the standard Linux kernel. Together, these tools provide the capability of monitoring an application's network behavior and protocol state information in production systems. The use of MAGNeT will enable the research community to construct a library of real traces of application-generated traffic from which researchers can more realistically test network protocol designs and theory. MAGNeT can also be used to verify the correct operation of protocol enhancements and to troubleshoot and tune protocol implementations.
Wu-chun Feng, Jeffrey R. Hay, Mark K. Gardner
ICCCN3
1999 Performance of algorithms for scheduling real-time systems with overrun and overload
abstract
This paper compares the performance of three classes of scheduling algorithms for real-time systems in which jobs may overrun their allocated processor time potentially causing the system to be overloaded. The first class, which contains classical priority scheduling algorithms as exemplified by DM and EDF provides a baseline. The second class is the Overrun Server Method which interrupts the execution of a job when it has used its allocated processor time and schedules the remaining portion as a request to an aperiodic server. The final class is the Isolation Server Method which executes each job as a request to an aperiodic server to which it has been assigned. The performance of the Overrun Sewer and Isolation Server Methods are worse, in general, than the performance of the baseline algorithms on independent workloads. However under the dependent workloads considered, the performance of the Isolation Server Method, using a server per task scheduled according to EDF, was significantly better than the performance of classical EDF.
Mark K. Gardner, Jane W.-S. Liu
ECRTS1
1999 Analyzing Stochastic Fixed-Priority Real-Time Systems
Mark K. Gardner, Jane W.-S. Liu
TACAS1
1997 Bounding Completion Times of Jobs with Arbitrary Release Times, Variable Execution Times and Resource Sharing
abstract
The workload of many real time systems can be characterized as a set of preemptable jobs with linear precedence constraints. Typically their execution times are only known to lie within a range of values. In addition, jobs share resources and access to the resources must be synchronized to ensure the integrity of the system. The paper is concerned with the schedulability of such jobs when scheduled on a priority driven basis. It describes three algorithms for computing upper bounds on the completion times of jobs that have arbitrary release times and priorities. The first two are simple but do not yield sufficiently tight bounds, while the last one yields the tightest bounds but has the greatest complexity.
Jun Sun 0002, Mark K. Gardner, Jane W.-S. Liu
IEEE Trans. Software Eng.2