VLDB 2026 Research / reviewers in the wild / expert
Meghan Cowan
dblp:202/1675
· DBLP profile ↗
7ranked-venue papers
3as first author
5since 2021 · last 2024
0000-0002-1052-0179ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 4 · 3 first-author · 2 since 2021Systems, architecture and hardware · 3 · 2 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Computer networks · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Towards a Standardized Representation for Deep Learning Collective AlgorithmsabstractThe explosion of machine learning model size has led to its execution on distributed clusters at a very large scale. Many works have tried to optimize the process of producing collective algorithms and running collective communications, which act as a bottleneck to distributed machine learning. However, different works use their own collective algorithm representation, pushing away from co-optimizing collective communication and the rest of the workload. The lack of a standardized collective algorithm rep-resentation has also hindered interoperability between collective algorithm producers and consumers. Additionally, tool-specific conversions and modifications have to be made for each pair of tools producing and consuming collective algorithms which adds to engineering efforts. In this position paper, we propose a standardized workflow leveraging a common collective algorithm representation. U p-stream producers and downstream consumers converge to a common representation format based on Chakra Execution Trace, a commonly used graph based representation of distributed machine learning workloads. Such a common representation enables us to view collective communications at the same level as workload operations and decouple producer and consumer tools, enhance interoperability, and relieve the user from the burden of having to focus on downstream implementations. We provide a proof-of-concept of this standardized workflow by simulating collective algorithms generated by the MSCCLang domain-specific language through the ASTRA-sim distributed machine learning simulator using various network configurations. Jinsun Yoo, William Won, Meghan Cowan, Benjamin Klenk, Srinivas Sridharan 0002, Tushar Krishna |
HOTI | 3 |
| 2023 | MSCCLang: Microsoft Collective Communication LanguageabstractMachine learning models with millions or billions of parameters are increasingly trained and served on large multi-GPU systems. As models grow in size and execute on more GPUs, collective communication becomes a bottleneck. Custom collective algorithms optimized for both particular network topologies and application-specific communication patterns can alleviate this bottleneck and help these applications scale. However, implementing correct and efficient custom algorithms is challenging. Meghan Cowan, Saeed Maleki, Madan Musuvathi, Olli Saarikivi, Yifan Xiong 0001 |
ASPLOS (2) | 1 |
| 2023 | TACCL: Guiding Collective Algorithm Synthesis using Communication Sketches
Aashaka Shah, Vijay Chidambaram, Meghan Cowan, Saeed Maleki, Madan Musuvathi, Todd Mytkowicz, Jacob Nelson 0001, Olli Saarikivi |
NSDI | 3 |
| 2021 | Mitigating Reverse Engineering Attacks on Local Feature Descriptors
Deeksha Dangwal, Vincent T. Lee, Hyo Jin Kim 0004, Tianwei Shen, Meghan Cowan, Rajvi Shah, Caroline Trippel, Brandon Reagen, Timothy Sherwood, Vassileios Balntas, Armin Alaghi, Eddy Ilg |
BMVC | 5 |
| 2021 | Porcupine: a synthesizing compiler for vectorized homomorphic encryptionabstractHomomorphic encryption (HE) is a privacy-preserving technique that enables computation directly on encrypted data. Despite its promise, HE has seen limited use due to performance overheads and compilation challenges. Recent work has made significant advances to address the performance overheads but automatic compilation of efficient HE kernels remains relatively unexplored. Meghan Cowan, Deeksha Dangwal, Armin Alaghi, Caroline Trippel, Vincent T. Lee, Brandon Reagen |
PLDI | 1 |
| 2020 | Automatic generation of high-performance quantized machine learning kernelsabstractQuantization optimizes machine learning inference for resource constrained environments by reducing the precision of its computation. In the extreme, even single-bit computations can produce acceptable results at dramatically lower cost. But this ultra-low-precision quantization is difficult to exploit because extracting optimal performance requires hand-tuning both high-level scheduling decisions and low-level implementations. As a result, practitioners settle for a few predefined quantized kernels, sacrificing optimality and restricting their ability to adapt to new hardware. Meghan Cowan, Thierry Moreau, Tianqi Chen 0001, James Bornholt, Luis Ceze |
CGO | 1 |
| 2018 | TVM: An Automated End-to-End Optimizing Compiler for Deep Learning
Tianqi Chen 0001, Thierry Moreau, Ziheng Jiang, Lianmin Zheng, Eddie Q. Yan, Haichen Shen, Meghan Cowan, Leyuan Wang, Luis Ceze, Carlos Guestrin, Arvind Krishnamurthy |
OSDI | 7 |