Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Graham Schelle

dblp:82/356 · DBLP profile ↗
← Back
9ranked-venue papers
6as first author
0since 2021 · last 2010
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 8 · 6 first-authorSoftware engineering, systems software and programming languages · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
4 papers
Reconfigurable computing and FPGAs · 55% Performance modeling and evaluation · 35% Processor architecture and microarchitecture · 5%
Computer networks
1 paper
Routing and switching · 100%

Topics — the 7 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Reconfigurable computing and FPGAs
FPGA-based emulation
0.112010
Intel nehalem processor core made FPGA synthesizable · FPGA 2010
Performance modeling and evaluation › simulation › parallel and distributed simulation
parallel simulation
0.112006
Exploiting parallelism and structure to accelerate the simulation of chip multi-processors · HPCA 2006
Performance modeling and evaluation
simulation
0.112006
Exploiting parallelism and structure to accelerate the simulation of chip multi-processors · HPCA 2006
Reconfigurable computing and FPGAs › FPGA-based emulation
multi-FPGA emulation
0.012010
Intel nehalem processor core made FPGA synthesizable · FPGA 2010
Processor architecture and microarchitecture
chip multiprocessor
0.012006
Exploiting parallelism and structure to accelerate the simulation of chip multi-processors · HPCA 2006
Routing and switching › router architecture
software router
0.012005
CUSP: a modular framework for high speed network applications on FPGAs · FPGA 2005
Electronic design automation
high-level synthesis
0.012004
Mapping a domain specific language to a platform FPGA · DAC 2004

Methods — techniques the papers use, named apart from their topics

speculation · 0.1parallelism · 0.1hardware integration · 0.1automated parallelization · 0.1domain-specific language embedding · 0.0
YearPublicationVenuePosition
2010 Intel nehalem processor core made FPGA synthesizable
abstract
We present a FPGA-synthesizable version of the Intel Nehalem processor core, synthesized, partitioned and mapped to a multi-FPGA emulation system consisting of Xilinx Virtex-4 and Virtex-5 FPGAs. To our knowledge, this is the first time a modern state-of-the-art x86 design with the out-of-order micro-architecture is made FPGA synthesizable and capable of high-speed cycle-accurate emulation. Unlike the Intel Atom core which was made FPGA synthesizable on a single Xilinx Virtex-5 in a previous endeavor, the Nehalem core is a more complex design with aggressive clock-gating, double phase latch RAMs, and RTL constructs that have no true equivalent in FPGA architectures. Despite these challenges, we are successful in making the RTL synthesizable with only 5% RTL code modifications, partitioning the design across five FPGAs, and emulating the core at 520 KHz. The synthesizable Nehalem core is able to boot Linux and execute standard x86 workloads with all architectural features enabled.
Graham Schelle, Jamison D. Collins, Ethan Schuchman, Perry H. Wang, Gautham N. Chinya, Ralf Plate, Thorsten Mattner, Franz Olbrich, Per Hammarlund, Ronak Singhal, Jim Brayton, Sebastian Steibl, Hong Wang 0003
FPGA1
2008 Exploring FPGA network on chip implementations across various application and network loads
abstract
The network on chip will become a future general purpose interconnect for FPGAs much like todaypsilas standard OPB or PLB bus architectures. However, performance characteristics and reconfigurable logic resource utilization of different network on chip architectures vary greatly relative to bus architectures. Current mainstream FPGA parts only support very small network on chip topologies, due to the high resource utilization of virtual channel based implementations. This observation is reflected in related research where only modest 2times2 or 2times3 networks are demonstrated on FPGAs. Naively it would be assumed that these complex network on chip architectures would perform better than simplified implementations. We show this assumption to be incorrect under light network loading conditions across 3 separate application domains. Using statistical based network loading, a synthetic benchmarking application, a cryptographic accelerator, and a 802.11 transmitter are each demonstrated across network on chip architectures. From these experiments, it can be seen that network on chips with complex routing and switching functionality are still useful under high network loading conditions. Additionally, it is also shown for our network on chip implementations, a simple solution that uses 4-5times less logic resources can provide better network performance under certain conditions.
Graham Schelle, Dirk Grunwald
FPL1
2007 Abstracting Modern FCCMs To Provide a Single Interface to Architectural Resources
abstract
Mainstream processor architectures and field programmable custom computing machines (FCCMs) are colliding towards a heterogeneous system on chip architecture. This is apparent from Intel and AMD efforts to create new chip architectures with various processing cores focusing on DSP, networking, and graphics. From the embedded processor research, system-on-chips connected by network on chips have allowed scalable architectures with a variety of processing cores connected by an onchip network. In this paper we examine several scheduling and allocation policies that can be utilized across network on chip architectures regardless of the processing cores onchip. By abstracting characteristics of the processing cores with various scheduling data structures, any heterogeneous system on a chip can be allocated and scheduled dynamically.
Graham Schelle, Dirk Grunwald
FCCM1
2007 A Software Defined Radio Application Utilizing Modern FPGAs and NoC Interconnects
abstract
Network on Chips are becoming a common onchip interconnect for both FPGA and mainstream processor designs. At the same time, software defined radios (SDR) are a new application field that is gaining much attention. As SDR tasks are mapped onto Network on Chip architectures, the typically streaming nature of samples will stress the NoC itself and possibly hurt the performance of other applications using that NoC. In this paper, we present the results of our partitioning and placement of a SDR transmitter onto a NoC architecture using an FPGA. We use a 802.11a transmitter example partitioned across a NoC and compare it to a handcrafted design. Additionally, various placement schemes, runtime architecture loads and NoC access methods are examined to determine the feasibility of this application and architecture combination.
Graham Schelle, Jeff Fifield, Dirk Grunwald
FPL1
2006 Exploiting parallelism and structure to accelerate the simulation of chip multi-processors
abstract
Simulation is an important means of evaluating new microarchitectures. Current trends toward chip multiprocessors (CMPs) try the ability of designers to develop efficient simulators. CMP simulation speed can be improved by exploiting parallelism in the CMP simulation model. This may be done by either running the simulation on multiple processors or by integrating multiple processors into the simulation to replace simulated processors. Doing so usually requires tedious manual parallelization or re-design to encapsulate processors. This paper presents techniques to perform automated simulator parallelization and hardware integration for CMP structural models. We show that automated parallelization can achieve an 7.60 speedup for a 16-processor CMP model on a conventional 4-processor shared-memory multiprocessor. We demonstrate the power of hardware integration by integrating eight hardware PowerPC cores into a CMP model, achieving a speedup of up to 5.82.
David A. Penry, Dan Fay, David Hodgdon, Ryan Wells, Graham Schelle, David I. August, Daniel A. Connors
HPCA5
2005 CUSP: a modular framework for high speed network applications on FPGAs
abstract
For several years now, modern FPGAs have included onchip network related hard cores. These cores include Xilinx's RocketIO and Altera's RapidIO serial transceivers. However, to use these cores in a complete networking application may be a daunting task to a non-networking expert. In addition to the complicated use of these components, the high performance needs of modern networking applications require designs that are optimized for low latency and a moderately high clock rate. Therefore to meet these challenges, we present CUSP (Click Utilizing Speculation and Parallelism)for reconfigurable hardware platforms.Click is an accepted software network router framework that is similar to CUSP, but specifically built for a Linux platform and software network routers. CUSP, while also having a modular design of reusable components, additionally provides automated speculation and parallelism to gain better performance on FPGAs. An accompanying scripting language allows quick creation of these routers from a body of existing components. We have implemented an example network application through the CUSP design flow and its performance will be compared against alternative network design methods.
Graham Schelle, Dirk Grunwald
FPGA1
2004 Mapping a domain specific language to a platform FPGA
abstract
A domain specific language (DSL) enables designers to rapidly specify and implement systems for a particular domain, yielding designs that are easy to understand, reason about, re-use and maintain. However, there is usually a significant overhead in the required infrastructure to map such a DSL on to a programmable logic device. In this paper, we present a mapping of an existing DSL for the networking domain on to a platform FPGA by embedding the DSL into an existing language infrastructure. In particular, we will show that, using few basic concepts, we are able to achieve a successful mapping of the DSL on to a platform FPGA and create a re-usable structure that also makes it easy to extend the DSL. Finally we will present some results of mapping the DSL on to a platform FPGA and comment on the resulting overhead.
Chidamber Kulkarni, Gordon J. Brebner, Graham Schelle
DAC3
2004 Automated Speculation and Parallelism in High Performance Network Applications
Graham Schelle, Dirk Grunwald
FPL1
2003 Privacy-Aware Location Sensor Networks
Marco Gruteser, Graham Schelle, Ashish Jain, Richard Han 0001, Dirk Grunwald
HotOS2