Julie Mullen

dblp:16/7585 · also Julia S. Mullen · DBLP profile ↗
← Back
8ranked-venue papers
4as first author
3since 2021 · last 2022
0000-0002-0015-6182ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 2 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 3 · 2 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
Performance modeling and evaluation · 61% GPUs and heterogeneous computing · 39%

Topics — the 4 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Performance modeling and evaluation › workload characterization
AI workload characterization
0.612022
AI-Enabling Workloads on Large-Scale GPU-Accelerated System: Characterization, Opportunities, and Implications · HPCA 2022
GPUs and heterogeneous computing › GPU computing
GPU-accelerated systems
0.612022
AI-Enabling Workloads on Large-Scale GPU-Accelerated System: Characterization, Opportunities, and Implications · HPCA 2022
Performance modeling and evaluation
workload characterization
0.612022
AI-Enabling Workloads on Large-Scale GPU-Accelerated System: Characterization, Opportunities, and Implications · HPCA 2022
GPUs and heterogeneous computing
GPU computing
0.212022
AI-Enabling Workloads on Large-Scale GPU-Accelerated System: Characterization, Opportunities, and Implications · HPCA 2022

Methods — techniques the papers use, named apart from their topics

user behavior analysis · 0.6job trace analysis · 0.6
YearPublicationVenuePosition
2022 AI-Enabling Workloads on Large-Scale GPU-Accelerated System: Characterization, Opportunities, and Implications
abstract
Production high-performance computing (HPC) systems are adopting and integrating GPUs into their design to accommodate artificial intelligence (AI), machine learning, and data visualization workloads. To aid with the design and operations of new and existing GPU-based large-scale systems, we provide a detailed characterization of system operations, job characteristics, user behavior, and trends on a contemporary GPU-accelerated production HPC system. Our insights indicate that the pre-mature phases in modern AI workflow take up significant GPU hours while underutilizing GPUs, which opens up the opportunity for a multi-tier system. Finally, we provide various potential recommendations and areas for future investment for system architects, operators, and users.
Baolin Li 0001, Rohin Arora, Siddharth Samsi, Tirthak Patel, William Arcand, David Bestor, Chansup Byun, Rohan Basu Roy, Bill Bergeron, John T. Holodnak, Michael Houle 0001, Matthew Hubbell, Michael Jones 0001, Jeremy Kepner, Anna Klein, Peter Michaleas, Joseph McDonald, Lauren Milechin, Julie Mullen, Andrew Prout, Benjamin Price, Albert Reuther, Antonio Rosa, Matthew L. Weiss, Charles Yee, Daniel Edelman, Allan Vanterpool, Anson Cheng, Vijay Gadepally, Devesh Tiwari
HPCA19
2021 Teaching HPC Concepts with Serious Games
abstract
This Innovative Practices Work in Progress describes how serious games can be used to help teach High Performance Computing (HPC). Serious games provide pathways for learners to develop intuition while learning new concepts. We define serious games as those that incorporate learning objectives, educational content, and assessment and whose purpose is learning rather than entertaining. Such games are valuable in an educational context because they engage students in active learning and help students develop mental models from their experiences. These student experiences generally lead to deeper questions, allowing instructors the chance to clarify misunderstandings, and reinforce learning. This work describes two games used in informal HPC courses to provide students with tangible hands-on experiences with HPC concepts.
Lauren Milechin, Julie Mullen
FIE2
2021 Teaching and learning HPC through serious games
Julie Mullen, Lauren Milechin, Dennis Milechin
J. Parallel Distributed Comput.1
2017 Bringing physical construction and real-world data collection into a massively open online course (MOOC)
abstract
This Work-In-Progress paper details the process and lessons learned when converting a hands-on engineering mini-course to a scalable, self-paced Massively Open Online Course (MOOC). Online courseware has been part of academic and industry training and learning for decades. Learning activities in online courses strive to mimic in-person delivery by including lectures, homework assignments, software exercises and exams. While these instructional activities provide “theory and practice” for many disciplines, engineering courses often require hands-on activities with physical tools, devices and equipment. To accommodate the need for this type of learning, MIT Lincoln Laboratory's “Build A Small Radar” (BSR) course was used to explore teaching and learning strategies that support the inclusion of physical construction and real world data collection in a MOOC. These tasks are encountered across a range of engineering disciplines and the methods illustrated here are easily generalized to the learning experiences in engineering and science disciplines.
Julie Mullen, Lauren Milechin, Michael Houle 0001, Patrick Bell, Alan Fenn, Kenneth E. Kolodziej, John Meklenburg, Janet Nguyen, Bradley T. Perry, Albert Reuther
FIE1
2017 Learning by doing, High Performance Computing education in the MOOC era
Julie Mullen, Chansup Byun, Vijay Gadepally, Siddharth Samsi, Albert Reuther, Jeremy Kepner
J. Parallel Distributed Comput.1
2015 Student-perceived effectiveness of online content delivery modes
abstract
This Work In Progress focuses on student perceptions of the effectiveness of three content delivery modes; a) traditional, residential in-class b) class capture for asynchronous online delivery, and c) modularized targeted content videos for online and blended or flipped classroom mode. Despite the growth of MOOCs and the concomitant shift from long lecture videos to learning modules, many online courses still rely on a class capture method. A recent study on video use by students in online courses recommends the use of short videos, six to nine minutes in length, based on student viewing habits. [1] The student behavior was inferred from click-stream data captured within the LMS but there was no direct interaction with students to gauge impressions of the content delivery or impact on learning. To improve both the residential and online experience requires a deeper understanding of those features which promote learning in each approach. Some of this understanding can be gleaned from traditional academic course evaluations, but student perceptions yield an additional level of understanding. For this study we analyzed data from multiple offerings of a 7-week graduate level numerical analysis course; Application of Finite Element Analysis. Over three years this course was offered residentially and online using multiple delivery modes. We categorize the delivery modes as Traditional, Class-Capture, and Modularized/Blended. Each student studied the material using one delivery mode and was then asked to review the material using another delivery mode. The students were asked to evaluate the perceived learning attained based on their original content delivery mode and the alternate mode. The feedback from this analysis, combined with traditional course evaluations will be used to develop a more comprehensive and objective survey to be used in the next round of courses.
Julie Mullen, John M. Sullivan
FIE1
2013 P-sync: A Photonically Enabled Architecture for Efficient Non-local Data Access
abstract
Communication in multi- and many-core processors has long been a bottleneck to performance due to the high cost of long-distance electrical transmission. This difficulty has been partially remedied by architectural constructs such as caches and novel interconnect topologies, albeit at a steep cost in terms of complexity. Unfortunately, even these measures are rendered ineffective by certain kinds of communication, most notably scatter and gather operations that exhibit highly nonlocal data access patterns. Much work has gone into examining how the increased bandwidth density afforded by chip-scale silicon photonic interconnect technologies affects computing, but photonics have additional properties that can be leveraged to greatly accelerate performance and energy efficiency under such difficult loads. This paper describes a novel synchronized global photonic bus and system architecture called P-sync that uses photonics' distance independence to greatly improve performance on many important applications previously limited by electronic interconnect. The architecture is evaluated in the context of a non-local yet common application: the distributed Fast Fourier Transform. We show that it is possible to achieve high efficiency by tightly balancing computation and communication latency in P-sync and achieve upwards of a 6× performance increase on gather patterns, even when bandwidth is equalized.
David Whelihan, Jeffrey J. Hughes, Scott M. Sawyer, Eric Robinson, Michael M. Wolf, Sanjeev Mohindra, Julie Mullen, Anna Klein, Michelle S. Beard, Nadya Bliss, Johnnie Chan, Robert Hendry, Keren Bergman, Luca P. Carloni
IPDPS7
2010 Hogs and slackers: Using operations balance in a genetic algorithm to optimize sparse algebra computation on distributed architectures
Una-May O'Reilly, Eric Robinson, Sanjeev Mohindra, Julie Mullen, Nadya Bliss
Parallel Comput.4