Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Christian Navasca

dblp:217/1287 · DBLP profile ↗
← Back
5ranked-venue papers
2as first author
3since 2021 · last 2024
0009-0009-3948-3373ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 4 · 2 first-author · 2 since 2021Systems, architecture and hardware · 2 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
3 papers
Memory systems · 37% GPUs and heterogeneous computing · 24% Hardware accelerators and domain-specific architectures · 24%
Software engineering, system software, and programming languages
3 papers
Runtime systems and virtual machines · 81% Compilers and program optimization · 19%
Artificial intelligence
1 paper
Deep learning architectures and training · 100%

Topics — the 10 heaviest of 12, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
GPUs and heterogeneous computing
GPU memory management
0.812024
Enabling Large Dynamic Neural Network Training with Learning-based Memory Management · HPCA 2024
Hardware accelerators and domain-specific architectures › accelerator offloading
tensor offloading
0.812024
Enabling Large Dynamic Neural Network Training with Learning-based Memory Management · HPCA 2024
Runtime systems and virtual machines
managed runtime
0.722019
Gerenuk: thin computation over big native data using speculative program transformation · SOSP 2019
Skyway: Connecting Managed Heaps in Distributed Big Data Systems · ASPLOS 2018
Runtime systems and virtual machines
garbage collection
0.612022
MemLiner: Lining up Tracing and Application for a Far-Memory-Friendly Runtime · OSDI 2022
Memory systems › memory disaggregation
far memory
0.612022
MemLiner: Lining up Tracing and Application for a Far-Memory-Friendly Runtime · OSDI 2022
Memory systems
memory management
0.612022
MemLiner: Lining up Tracing and Application for a Far-Memory-Friendly Runtime · OSDI 2022
Runtime systems and virtual machines
object representation
0.412019
Gerenuk: thin computation over big native data using speculative program transformation · SOSP 2019
Compilers and program optimization
program transformation
0.412019
Gerenuk: thin computation over big native data using speculative program transformation · SOSP 2019
High-performance computing
data transfer
0.312018
Skyway: Connecting Managed Heaps in Distributed Big Data Systems · ASPLOS 2018
Distributed and cloud data management
big data systems
0.112019
Gerenuk: thin computation over big native data using speculative program transformation · SOSP 2019

Methods — techniques the papers use, named apart from their topics

pilot model · 1.5neural network · 1.5tracing · 1.1thin computation · 0.8speculative program transformation · 0.8direct heap connection · 0.7
YearPublicationVenuePosition
2024 Enabling Large Dynamic Neural Network Training with Learning-based Memory Management
abstract
Dynamic neural network (DyNN) enables high computational efficiency and strong representation capability. However, training DyNN can face a memory capacity problem because of increasing model size or limited GPU memory capacity. Managing tensors to save GPU memory is challenging, because of the dynamic structure of DyNN. We present DyNN-Offload, a memory management system to train DyNN. DyNN-Offload uses a learned approach (using a neural network called the pilot model) to increase predictability of tensor accesses to facilitate memory management. The key of DyNN-Offload is to enable fast inference of the pilot model in order to reduce its performance overhead, while providing high inference (or prediction) accuracy. DyNNOffload reduces input feature space and model complexity of the pilot model based on a new representation of DyNN; DyNNOffload converts the hard problem of making prediction for individual operators into a simpler problem of making prediction for a group of operators in DyNN. DyNN-Offload enables 8 × larger DyNN training on a single GPU compared with using PyTorch alone (unprecedented with any existing solution). Evaluating with AlphaFold (a production-level, large-scale DyNN), we show that DyNN-Offload outperforms unified virtual memory (UVM) and dynamic tensor rematerialization (DTR), the most advanced solutions to save GPU memory for DyNN, by 3 × and 2.1 × respectively in terms of maximum batch size.
Jie Ren 0015, Dong Xu 0024, Shuangyan Yang, Christian Navasca, Chenxi Wang 0005, Guoqing Harry Xu, Dong Li 0001
HPCA6
2023 Predicting Dynamic Properties of Heap Allocations using Neural Networks Trained on Static Code: An Intellectual Abstract
abstract
Memory allocators and runtime systems can leverage dynamic properties of heap allocations – such as object lifetimes, hotness or access correlations – to improve performance and resource consumption. A significant amount of work has focused on approaches that collect this information in performance profiles and then use it in new memory allocator or runtime designs, both offline (e.g., in ahead-of-time compilers) and online (e.g., in JIT compilers). This is a special instance of profile-guided optimization.
Christian Navasca, Martin Maas 0001, Petros Maniatis, Hyeontaek Lim, Guoqing Harry Xu
ISMM1
2022 MemLiner: Lining up Tracing and Application for a Far-Memory-Friendly Runtime
Chenxi Wang 0005, Yifan Qiao 0002, Jon Eyolfson, Christian Navasca, Shan Lu 0001, Guoqing Harry Xu
OSDI6
2019 Gerenuk: thin computation over big native data using speculative program transformation
abstract
Big Data systems are typically implemented in object-oriented languages such as Java and Scala due to the quick development cycle they provide. These systems are executed on top of a managed runtime such as the Java Virtual Machine (JVM), which requires each data item to be represented as an object before it can be processed. This representation is the direct cause of many kinds of severe inefficiencies.
Christian Navasca, Cheng Cai, Khanh Nguyen 0001, Brian Demsky, Shan Lu 0001, Miryung Kim, Guoqing Harry Xu
SOSP1
2018 Skyway: Connecting Managed Heaps in Distributed Big Data Systems
abstract
Managed languages such as Java and Scala are prevalently used in development of large-scale distributed systems. Under the managed runtime, when performing data transfer across machines, a task frequently conducted in a Big Data system, the system needs to serialize a sea of objects into a byte sequence before sending them over the network. The remote node receiving the bytes then deserializes them back into objects. This process is both performance-inefficient and labor-intensive: (1) object serialization/deserialization makes heavy use of reflection, an expensive runtime operation and/or (2) serialization/deserialization functions need to be hand-written and are error-prone. This paper presents Skyway, a JVM-based technique that can directly connect managed heaps of different (local or remote) JVM processes. Under Skyway, objects in the source heap can be directly written into a remote heap without changing their formats. Skyway provides performance benefits to any JVM-based system by completely eliminating the need (1) of invoking serialization/deserialization functions, thus saving CPU time, and (2) of requiring developers to hand-write serialization functions.
Khanh Nguyen 0001, Lu Fang 0003, Christian Navasca, Guoqing Harry Xu, Brian Demsky, Shan Lu 0001
ASPLOS3