Christopher Nitta

dblp:02/6395 · also Christopher J. Nitta · DBLP profile ↗
← Back
15ranked-venue papers
11as first author
3since 2021 · last 2024
0000-0003-1531-2771ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 10 · 6 first-authorHuman-computer interaction and ubiquitous computing · 4 · 4 first-author · 3 since 2021
YearPublicationVenuePosition
2024 The First Five Years of a Dual Track Programming Series
abstract
The varying skill levels of students in entry-level university programming courses complicates an already challenging task for instructors. Combining students of differing skill levels into a single course can have negative impact on the classroom experience making some students feel inadequate compared to their peers. Prior to 2018, our large four-year R1 research university in the United States had a single lower-division programming sequence. Since only a single sequence existed all students (both majors and non-majors) were grouped together in the same courses. In this paper we analyze the impact of developing a dual lower division programming course sequence and deploying a placement exam. Our analysis tracks student performance in upper division courses over the first five years of the dual track programming course sequence, and we present the performance difference based upon the particular track the students take.
Christopher Nitta, Kurt Eiselt
SIGCSE (1)1
2022 RISC-V Console: A Containerized RISC-V Based Game Console Emulator for Education
abstract
The rapid transition to online education due to the COVID-19 pandemic left many instructors needing to redesign their course projects as students no longer had access to physical hardware. This paper describes the development of an open-source containerized RISC-V based game console emulator that replaced physical hardware for use in course projects. The tool was initially designed and used in a graduate operating systems course and then subsequently used in a lower division computer organization and machine-dependent programming course. The container provides a full toolchain with gcc compiler, RISC-V game console emulator with integrated debugger, example program, and input recording/auto-run tool designed for auto-grading. The use of a container reduced the barrier to entry for the students allowing them to get up and running in a relatively short period of time. Given the successful deployment of the tool in the previous courses, the tool was used both again in the lower division course and in the upper division undergraduate operating systems course this past fall.
Christopher Nitta, Aaron Kaloti, Shuotong Wang
ITiCSE (1)1
2021 Using a Comprehensive Third-Party Exam for ABET Student Outcome Assessment
abstract
In anticipation of an ABET accreditation visit, our computer science department contracted with an independent external testing organization to perform an assessment of our students' proficiency in computer science. We used this external assessment to supplement our own internal assessment and have continued using since our site visit. In this paper, we talk about the comprehensive test, how it was administered, the steps we took to validate the results of the external assessment, and how we integrated the results into our ABET self-study report. We also explore insights we gained from the testing about our program and our students, as well as some of the limitations of the test.
Christopher Nitta, Kurt Eiselt
SIGCSE1
2020 HCAPP: Scalable Power Control for Heterogeneous 2.5D Integrated Systems
abstract
Package pin allocation is becoming a key bottleneck in the capabilities of designs due to the increased bandwidth requirements. 2.5D integration compounds these package-level requirements while introducing an increased number of compute units within the package. We propose a decentralized power control implementation called Heterogeneous Constant Average Power Processing (HCAPP) to maintain the power limit while maximizing the efficiency of the package pins allocated for power. HCAPP uses a hardware-based decentralized design to handle fast power limits, maintain scalability and enable simplified control for heterogeneous systems while maximizing performance. As extensions, we evaluate a software interface and the impact of different accelerator designs. Overall, HCAPP achieves 7% speedup over a RAPL-like implementation. The power utilization improves from 79.7% (RAPL-like) to 93.9% (HCAPP) with this design. A priority-based static software control methodology alongside HCAPP provides average speedups of 8.3% (CPU), 5.4% (GPU), and 12% (Accelerator) for the prioritized component compared to the unprioritized version.
Kramer Straube, Jason Lowe-Power, Christopher Nitta, Matthew K. Farrens, Venkatesh Akella
ICPP3
2020 Using the CS2013 Exam for ABET Student Outcome Assessment
abstract
In anticipation of an ABET accreditation visit, our computer science department contracted with an independent external testing organization to perform an assessment of our students' proficiency in computer science. We used this external assessment to supplement our own internal assessment. The exam was easy to administer and covered multiple student outcomes. In addition, our analysis of the exam results showed high correlation between course performance and exam performance.
Christopher Nitta, Kurt Eiselt
SIGCSE1
2018 Improving Provisioned Power Efficiency in HPC Systems with GPU-CAPP
abstract
In this paper we propose a microarchitectural technique called GPU Constant Average Power Processing (GPU-CAPP) that improves the power utilization of power provisioning-limited systems by using provisioned power as much as possible to accelerate computation on parallel work-loads. GPU-CAPP uses a flexible, decentralized control to ensure fast response times and the scalability required for increasingly parallel GPU designs. We use GPGPU-Sim and GPUWattch to simulate GPU-CAPP and evaluate its capabilities on a subset of the Rodinia benchmark suite. Overall, GPU-CAPP enables speedup by an average of 26% and 12% over equivalent fixed frequency systems at two power targets.
Kramer Straube, Jason Lowe-Power, Christopher Nitta, Matthew K. Farrens, Venkatesh Akella
HiPC3
2017 Improving Execution Time of Parallel Programs on Large Scale Chip Multiprocessors with Constant Average Power Processing
abstract
In this paper we propose a microarchitectural technique called Constant Average Power Processing (CAPP) that reduces the execution time of parallel programs by dynamically detecting the power slack at runtime and directing it to specific core(s) that are the bottleneck at any given time. The key insight of this work is that by sensing the current, communicating it to the global controller and adjusting the cores' frequencies, it is possible to maintain a constant power level in a distributed and scalable manner. We evaluate the potential benefits and scalability of the proposed technique on a set of synthetic benchmarks and compare the results with related work such as Running Average Power Limit (RAPL).
Kramer Straube, Christopher Nitta, Rajeevan Amirtharajah, Matthew K. Farrens, Venkatesh Akella
ICCD2
2015 Makers from Hobbyists to professionals
Christopher Nitta
Hot Chips Symposium1
2014 PDG_GEN: A Methodology for Fast and Accurate Simulation of On-Chip Networks
abstract
With the advent of large scale chip multiprocessors, there is growing interest in the design and analysis of on-chip networks. Full-system simulation is the most accurate way to perform such an analysis, but unfortunately it is very slow and thus limits design space exploration. To overcome this problem researchers frequently use trace-based simulation to study different network topologies and properties, which can be done much faster. Unfortunately, unless the traces that are used include information about dependencies between packets, trace-based simulations can lead one to draw incorrect conclusions about network performance metrics such as average packet latency and overall execution time. The primary contributions of this work are to demonstrate the importance of including dependency information in traces, and to present PDG_GEN, an inference-based technique for identifying and including dependencies in traces. This technique uses traces obtained from multiple full-system simulations of an application of interest to infer dependency information between packets and augment traces with this information. On the SPLASH-2 benchmark suite, PDG_GEN is 2.3 times more accurate at predicting overall execution time and almost 4,000 times more accurate at predicting average packet latency than traditional trace-based methods.
Kevin Macdonald, Christopher Nitta, Matthew K. Farrens, Venkatesh Akella
IEEE Trans. Computers2
2012 DCAF - A Directly Connected Arbitration-Free Photonic Crossbar for Energy-Efficient High Performance Computing
abstract
DCAF is a directly connected arbitration free photonic crossbar that is realized by taking advantage of multiple photonic layers connected with photonic vias. In order to evaluate DCAF we developed a detailed implementation model for the network and analyzed the power and performance on a variety of benchmarks, including SPLASH-2 and synthetic traces. Our results demonstrate that the overhead required by arbitration is non-trivial, especially at high loads. Eliminating the need for arbitration, sizing the buffers carefully and retransmitting lost packets when there is contention results in a 44% reduction in average packet latency without additional power overhead. We also use an analytical model for ScaLAPACK QR decomposition and find that a 64 processor DCAF could outperform a 1024 node cluster connected with 40Gbps links on matrices up to 500MB in size.
Christopher Nitta, Matthew K. Farrens, Venkatesh Akella
IPDPS1
2011 Addressing system-level trimming issues in on-chip nanophotonic networks
abstract
The basic building block of on-chip nanophotonic interconnects is the microring resonator, and these resonators change their resonant wavelengths due to variations in temperature - a problem that can be addressed using a technique called ”trimming”, which involves correcting the drift via heating and/or current injection. Thus far system researchers have modeled trimming as a per ring fixed cost. In this work we show that at the system level using a fixed cost model is inappropriate - our simulations demonstrate that the cost of heating has a non-linear relationship with the number of rings, and also that current injection can lead to thermal runaway. We show that a very narrow Temperature Control Window (TCW) must be maintained in order for the network to work as desired. However, by exploiting the group drift property of co-located rings, it is possible to create a sliding window scheme which can increase the TCW. We also show that partially athermal rings can alleviate but not eliminate the problem.
Christopher Nitta, Matthew K. Farrens, Venkatesh Akella
HPCA1
2011 Resilient microring resonator based photonic networks
abstract
Microring resonator-based photonic interconnects are being considered for both on-chip and off-chip communication in order to satisfy the power and bandwidth requirements of future large scale chip multiprocessors. However, microring resonators are prone to malfunction due to fabrication errors, and they are also extremely sensitive to fluctuations in temperature. In this paper we derive a fault model for microring based optical links that can be used by computer architects to make informed design choices. We evaluate different schemes for improving resilience, such as retransmission versus error-correction, using an optical fault simulator based on our fault model. We show how meeting a target mean time between failures (MTBF) affects the choice of resilience scheme - our investigation indicates that until fault rates are in the range of 10−21 to 10−24 per cycle, error detection/correction schemes will be needed in order to meet a 1M hour MTBF. We also evaluate how the resilience scheme impacts the performance of the link, which will help an architect choose the appropriate scheme based on the throughput requirements of a particular design.
Christopher Nitta, Matthew K. Farrens, Venkatesh Akella
MICRO1
2011 Inferring packet dependencies to improve trace based simulation of on-chip networks
abstract
With the advent of large scale chip-level multiprocessors, there is a growing interest in the design and analysis of on-chip networks. The use of full system simulation is the most accurate way to perform such an analysis, but unfortunately it is very slow and thus limits design space exploration. In order to overcome this problem researchers frequently use trace based simulation to study different network topologies and properties, which can be done much faster. Unfortunately, unless the traces that are used include information about dependencies between messages (packets), trace based simulation can lead one to draw incorrect conclusions about network performance metrics such as latency and overall execution time. In this paper we will demonstrate the importance of including dependency information in traces, as well as present an inference-based technique for identifying and including dependencies, and show that using these augmented traces results in much better simulation accuracy without excessively extending simulation time.
Christopher Nitta, Kevin Macdonald, Matthew K. Farrens, Venkatesh Akella
NOCS1
2008 Techniques for increasing effective data bandwidth
abstract
In this paper we examine techniques for increasing the effective bandwidth of the microprocessor off-chip interconnect. We focus on mechanisms that are orthogonal to other techniques currently being studied (3-D fabrication, optical interconnect, etc.) Using a range of full-system simulations we study the distribution of values being transferred to and from memory, and find that (as expected) high entropy data such as floating point numbers have limited compressibility, but that other data types offer more potential for compression. By using a simple heuristic to classify the contents of a cache line and providing different compression schemes for each classification, we show it is possible to provide overall compression at a cache line granularity comparable to that obtained by using a much more complex Lempel-Ziv-Welch algorithm.
Christopher Nitta, Matthew K. Farrens
ICCD1
2006 Y-Threads: Supporting Concurrency in Wireless Sensor Networks
Christopher Nitta, Raju Pandey, Yann Ramin
DCOSS1