Amir Kamil

dblp:73/5929 · DBLP profile ↗
← Back
11ranked-venue papers
3as first author
3since 2021 · last 2023
0000-0002-6112-5283ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 1 first-authorHuman-computer interaction and ubiquitous computing · 4 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Software engineering, system software, and programming languages
2 papers
Empirical software engineering · 89% Concurrent programming · 7% Compilers and program optimization · 4%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Memory systems · 83% Parallel and multicore computing · 17%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Computing education · 100%

Topics — the 11 heaviest of 11, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computing education › computer science curriculum
formal methods education
0.712023
How Do We Read Formal Claims? Eye-Tracking and the Cognition of Proofs about Algorithms · ICSE 2023
Empirical software engineering › human factors in software engineering
developer cognition
0.712023
How Do We Read Formal Claims? Eye-Tracking and the Cognition of Proofs about Algorithms · ICSE 2023
Empirical software engineering › human factors in software engineering
eye-tracking studies
0.712023
How Do We Read Formal Claims? Eye-Tracking and the Cognition of Proofs about Algorithms · ICSE 2023
Memory systems
data locality
0.312017
Trends in Data Locality Abstractions for HPC Systems · IEEE Trans. Parallel Distributed Syst. 2017
Memory systems
data movement
0.312017
Trends in Data Locality Abstractions for HPC Systems · IEEE Trans. Parallel Distributed Syst. 2017
Memory systems › memory hierarchy
memory hierarchy management
0.312017
Trends in Data Locality Abstractions for HPC Systems · IEEE Trans. Parallel Distributed Syst. 2017
Parallel and multicore computing
parallel programming models
0.112017
Trends in Data Locality Abstractions for HPC Systems · IEEE Trans. Parallel Distributed Syst. 2017
Parallel and multicore computing
parallel programming runtimes
0.112017
Trends in Data Locality Abstractions for HPC Systems · IEEE Trans. Parallel Distributed Syst. 2017
Compilers and program optimization
compiler analysis
0.112005
Making Sequential Consistency Practical in Titanium · SC 2005
Concurrent programming
memory models
0.112005
Making Sequential Consistency Practical in Titanium · SC 2005
Concurrent programming › memory models
sequential consistency
0.112005
Making Sequential Consistency Practical in Titanium · SC 2005

Methods — techniques the papers use, named apart from their topics

eye tracking · 1.3controlled experiment · 1.3survey · 0.3
YearPublicationVenuePosition
2023 How Do We Read Formal Claims? Eye-Tracking and the Cognition of Proofs about Algorithms
abstract
Formal methods are used successfully in high-assurance software, but they require rigorous mathematical and logical training that practitioners often lack. As such, integrating formal methods into software has been associated with numerous challenges. While educators have placed emphasis on formalisms in undergraduate theory courses, such courses often struggle with poor student outcomes and satisfaction. In this paper, we present a controlled eye-tracking human study (n = 34) investigating the problem-solving strategies employed by students with different levels of incoming preparation (as assessed by theory coursework taken and pre-screening performance on a proof comprehension task), and how educators can better prepare low-outcome students for the rigorous logical reasoning that is a core part of formal methods in software engineering. Surprisingly, we find that incoming preparation is not a good predictor of student outcomes for formalism comprehension tasks, and that student self-reports are not accurate at identifying factors associated with high outcomes for such tasks. Instead, and importantly, we find that differences in outcomes can be attributed to performance for proofs by induction and recursive algorithms, and that better-performing students exhibit significantly more attention switching behaviors, a result that has several implications for pedagogy in terms of the design of teaching materials. Our results suggest the need for a substantial pedagogical intervention in core theory courses to better align student outcomes with the objectives of mastery and retaining the material, and thus bettering preparing students for high-assurance software engineering.
Hammad Ahmad, Zachary Karas, Kimberly Diaz, Amir Kamil, Jean-Baptiste Jeannin, Westley Weimer
ICSE4
2021 An Analysis of Iterative and Recursive Problem Performance
abstract
Iteration and recursion are fundamental programming constructs in introductory computer science. Understanding the relationship between contextual factors, such as problem formulation or student background, that relate to performance on iteration and recursion problems can help inform pedagogy. We present the results of a study of 162 undergraduate participants tasked with comprehending iterative, recursive, and tail-recursive versions of CS1 functions. First, we carry out a task-specific analysis, finding that students perform significantly better on the iterative framings of two problems with non-branching numerical computation and significantly better on the recursive framing of another that involves array classification (p < = 0.036). Second, we investigate differences in the most common student mistakes by program framing. We find that students were more likely to produce wrong answers with incorrect types or structures for recursive and tail-recursive program versions. Finally, we investigated correlations between programming performance and background factors including experience, gender, ethnicity, affluence, and spatial ability. We find that the factors relevant to explaining performance are similar for both iterative and recursive problems. While programming experience is the most significant factor, we find that spatial ability, gender, and ethnicity were more relevant for explaining performance than affluence.
Madeline Endres, Westley Weimer, Amir Kamil
SIGCSE3
2021 Showcase of NCWIT Academic Alliance Members: Promising Practices Regarding Admission, Curriculum, Pedagogy, TA Selection, and Undergraduate Research
abstract
Broadening participation in computing touches every aspect of the undergraduate experience. This special session highlights the initiatives undertaken by NCWIT Academic Alliance members who are working to broaden participation in computing. A mix of 3-minute lightning talks, review of resources, and Q&A will provide attendees opportunities to create connections and grapple with implementation issues at their institution. This special session is a reprise of a well-reviewed session from the NCWIT Summit, and will introduce strategies for admission to major, curriculum, pedagogy, teaching assistant selection, and undergraduate research.
Colleen M. Lewis, Olga Glebova, Amir Kamil, Clifton Kussmaul, Briana B. Morrison, Katie A. Siek
SIGCSE3
2019 UPC++: A High-Performance Communication Framework for Asynchronous Computation
abstract
UPC++ is a C++ library that supports high-performance computation via an asynchronous communication framework. This paper describes a new incarnation that differs substantially from its predecessor, and we discuss the reasons for our design decisions. We present new design features, including future-based asynchrony management, distributed objects, and generalized Remote Procedure Call (RPC). We show microbenchmark performance results demonstrating that one-sided Remote Memory Access (RMA) in UPC++ is competitive with MPI-3 RMA; on a Cray XC40 UPC++ delivers up to a 25% improvement in the latency of blocking RMA put, and up to a 33% bandwidth improvement in an RMA throughput test. We showcase the benefits of UPC++ with irregular applications through a pair of application motifs, a distributed hash table and a sparse solver component. Our distributed hash table in UPC++ delivers near-linear weak scaling up to 34816 cores of a Cray XC40. Our UPC++ implementation of the sparse solver component shows robust strong scaling up to 2048 cores, where it outperforms variants communicating using MPI by up to 3.1x. UPC++ encourages the use of aggressive asynchrony in low overhead RMA and RPC, improving programmer productivity and delivering high performance in irregular applications.
John Bachan, Scott B. Baden, Steven Hofmeyr, Mathias Jacquelin, Amir Kamil, Dan Bonachea, Paul Hargrove, Hadia Ahmed
IPDPS5
2019 Unexpected Tokens: A Review of Programming Error Messages and Design Guidelines for the Future
abstract
Diagnostic messages generated by compilers and interpreters such as syntax error messages have been researched for decades. Unfortunately these messages which include error, warning, and runtime messages, present substantial difficulty and could be more effective, particularly for novices. Recent years have seen increased number of papers in the area including studies on the effectiveness of these messages, improving or enhancing them, and their usefulness as a part of programming process data that can be used to predict student performance. Despite this increased interest, the long history of literature is quite scattered and has not been brought together in any digestible form. We argue that in order to help the community proceed with more work on diagnostic messages, the literature needs to be presented in a state-of-the-art report. In addition we will synthesize and present the existing evidence for these messages including the difficulties they present and their effectiveness. We will also formulate a set of guidelines based on this evidence that can be used when designing or enhancing diagnostic messages. This work can serve as a starting point for those who wish to conduct research on such messages, those who wish to design better messages or those that aim to measure their effectiveness, more effectively.
Brett A. Becker, Paul Denny 0001, Raymond Pettit, Durell Bouchard, Dennis J. Bouvier, Brian Harrington 0001, Amir Kamil, Amey Karkare, Chris McDonald, Peter-Michael Osera, Janice L. Pearce, James Prather
ITiCSE7
2019 Gender-balanced TAs from an Unbalanced Student Body
abstract
Increasing participation of women and underrepresented minorities is a key challenge in the field of Computer Science Education. Balanced representation of these groups among teaching assistants in Computer Science courses influences recruitment and retention of underrepresented students. At the same time, the status-quo reduced participation of these students makes it more difficult to hire instructional staff from underrepresented groups. In this paper, we describe our experience evaluating candidates with teaching-demonstration videos, followed by in-person interviews, to hire a gender-balanced set of undergraduate TAs for a large-scale CS2 course. Our research goal is to quantitatively assess gender balance throughout the hiring process. Our initial applicant pool is just one-sixth women, but we found that women applicants perform better in our application process than men, resulting in a gender-balanced course staff without making hiring decisions based on the gender of applicants. We show that our approach results in a more gender-balanced teaching staff than hiring based on applicant GPA. We also use course-evaluation data to demonstrate that women perform as well as men as teaching assistants in CS2, and that the overall quality of our teaching assistants has remained high after the hiring-process change.
Amir Kamil, James Juett, Andrew DeOrio
SIGCSE1
2017 Trends in Data Locality Abstractions for HPC Systems
abstract
The cost of data movement has always been an important concern in high performance computing (HPC) systems. It has now become the dominant factor in terms of both energy consumption and performance. Support for expression of data locality has been explored in the past, but those efforts have had only modest success in being adopted in HPC applications for various reasons. them However, with the increasing complexity of the memory hierarchy and higher parallelism in emerging HPC systems, locality management has acquired a new urgency. Developers can no longer limit themselves to low-level solutions and ignore the potential for productivity and performance portability obtained by using locality abstractions. Fortunately, the trend emerging in recent literature on the topic alleviates many of the concerns that got in the way of their adoption by application developers. Data locality abstractions are available in the forms of libraries, data structures, languages and runtime systems; a common theme is increasing productivity without sacrificing performance. This paper examines these trends and identifies commonalities that can combine various locality concepts to develop a comprehensive approach to expressing and managing data locality on future large-scale high-performance computing systems.
Didem Unat, Anshu Dubey, Torsten Hoefler, John Shalf, Mark James Abraham, Mauro Bianco, Bradford L. Chamberlain, Romain Cledat, H. Carter Edwards, Hal Finkel, Karl Fürlinger, Frank Hannig, Emmanuel Jeannot, Amir Kamil, Jeff Keasler, Paul H. J. Kelly, Vitus J. Leung, Hatem Ltaief, Naoya Maruyama, Chris J. Newburn, Miquel Pericàs
IEEE Trans. Parallel Distributed Syst.14
2016 A Hartree-Fock Application Using UPC++ and the New DArray Library
abstract
The Hartree-Fock (HF) method is the fundamental first step for incorporating quantum mechanics into many-electron simulations of atoms and molecules, and it is an important component of computational chemistry toolkits like NWChem. The GTFock code is an HF implementation that, while it does not have all the features in NWChem, represents crucial algorithmic advances that reduce communication and improve load balance by doing an up-front static partitioning of tasks, followed by work stealing whenever necessary. To enable innovations in algorithms and exploit next generation exascale systems, it is crucial to support quantum chemistry codes using expressive and convenient programming models and runtime systems that are also efficient and scalable. This paper presents an HF implementation similar to GTFock using UPC++, a partitioned global address space model that includes flexible communication, asynchronous remote computation, and a powerful multidimensional array library. UPC++ offers runtime features that are useful for HF such as active messages, a rich calculus for array operations, hardware-supported fetch-and-add, and functions for ensuring asynchronous runtime progress. We present a new distributed array abstraction, DArray, that is convenient for the kinds of random-access array updates and linear algebra operations on block-distributed arrays with irregular data ownership. We analyze the performance of atomic fetch-and-add operations (relevant for load balancing) and runtime attentiveness, then compare various techniques and optimizations for each. Our optimized implementation of HF using UPC++ and the DArrays library shows up to 20% improvement over GTFock with Global Arrays at scales up to 24,000 cores.
David Ozog, Amir Kamil, Yili Zheng, Paul Hargrove, Jeff R. Hammond, Allen D. Malony, Wibe de Jong, Katherine A. Yelick
IPDPS2
2014 UPC++: A PGAS Extension for C++
abstract
Partitioned Global Address Space (PGAS) languages are convenient for expressing algorithms with large, random-access data, and they have proven to provide high performance and scalability through lightweight one-sided communication and locality control. While very convenient for moving data around the system, PGAS languages have taken different views on the model of computation, with the static Single Program Multiple Data (SPMD) model providing the best scalability. In this paper we present UPC++, a PGAS extension for C++ that has three main objectives: 1) to provide an object-oriented PGAS programming model in the context of the popular C++ language, 2) to add useful parallel programming idioms unavailable in UPC, such as asynchronous remote function invocation and multidimensional arrays, to support complex scientific applications, 3) to offer an easy on-ramp to PGAS programming through interoperability with other existing parallel programming systems (e.g., MPI, OpenMP, CUDA). We implement UPC++ with a "compiler-free" approach using C++ templates and runtime libraries. We borrow heavily from previous PGAS languages and describe the design decisions that led to this particular set of language features, providing significantly more expressiveness than UPC with very similar performance characteristics. We evaluate the programmability and performance of UPC++ using five benchmarks on two representative supercomputers, demonstrating that UPC++ can deliver excellent performance at large scale up to 32K cores while offering PGAS productivity features to C++ applications.
Yili Zheng, Amir Kamil, Michael B. Driscoll, Hongzhang Shan, Katherine A. Yelick
IPDPS2
2007 Hierarchical Pointer Analysis for Distributed Programs
Amir Kamil, Katherine A. Yelick
SAS1
2005 Making Sequential Consistency Practical in Titanium
abstract
The memory consistency model in shared memory parallel programming controls the order in which memory operations performed by one thread may be observed by another. The most natural model for programmers is to have memory accesses appear to take effect in the order specified in the original program. Language designers have been reluctant to use this strong semantics, called sequential consistency, due to concerns over the performance of memory fence instructions and related mechanisms that guarantee order. In this paper, we provide evidence for the practicality of sequential consistency by showing that advanced compiler analysis techniques are sufficient to eliminate the need for most memory fences and enable high-level optimizations. Our analyses eliminated over 97% of the memory fences that were needed by a na¨ýve implementation, accounting for 87 to 100% of the dynamically encountered fences in all but one benchmark. The impact of the memory model and analysis on runtime performance depends on the quality of the optimizations: more aggressive optimizations are likely to be invalidated by a strong memory consistency semantics. We consider two specific optimizations pipelining of bulk memory copies and communication aggregation and scheduling for irregular accesses and show that our most aggressive analysis is able to obtain the same performance as the relaxed model when applied to two linear algebra kernels. While additional work on parallel optimizations and analyses is needed, we believe these results provide important evidence on the viability of using a simple memory consistency model without sacrificing performance.
Amir Kamil, Jimmy Su, Katherine A. Yelick
SC1