Karima Ma

dblp:245/4922 · DBLP profile ↗
← Back
5ranked-venue papers
1as first author
4since 2021 · last 2025
0000-0003-4180-6433ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer graphics and multimedia
2 papers
Image and video processing · 60% Audio and music processing · 40%
Software engineering, system software, and programming languages
2 papers
Compilers and program optimization · 100%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
GPUs and heterogeneous computing · 50% Performance modeling and evaluation · 50%
Human-computer interaction and pervasive computing
1 paper
Interaction techniques and input · 100%

Topics — the 10 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Compilers and program optimization
auto-scheduling
0.922021
Efficient automatic scheduling of imaging and vision pipelines for the GPU · Proc. ACM Program. Lang. 2021
Learning to optimize halide with tree search and random programs · ACM Trans. Graph. 2019
Audio and music processing
vocal imitation
0.812024
Sketching With Your Voice: "Non-Phonorealistic" Rendering of Sounds via Vocal Imitation · SIGGRAPH Asia 2024
Image and video processing › image restoration
demosaicing
0.612022
Searching for Fast Demosaicking Algorithms · ACM Trans. Graph. 2022
Image and video processing
super-resolution
0.612022
Searching for Fast Demosaicking Algorithms · ACM Trans. Graph. 2022
Compilers and program optimization › code generation
GPU code generation
0.512021
Efficient automatic scheduling of imaging and vision pipelines for the GPU · Proc. ACM Program. Lang. 2021
Performance modeling and evaluation
cost modeling
0.512021
Efficient automatic scheduling of imaging and vision pipelines for the GPU · Proc. ACM Program. Lang. 2021
GPUs and heterogeneous computing › GPU computing
GPU implementation
0.512021
Efficient automatic scheduling of imaging and vision pipelines for the GPU · Proc. ACM Program. Lang. 2021
Compilers and program optimization
cost model
0.412019
Learning to optimize halide with tree search and random programs · ACM Trans. Graph. 2019
Interaction techniques and input › voice interaction
speech input
0.212024
Sketching With Your Voice: "Non-Phonorealistic" Rendering of Sounds via Vocal Imitation · SIGGRAPH Asia 2024
Machine learning › Efficient and distributed learning
deep learning compilers
0.112019
Learning to optimize halide with tree search and random programs · ACM Trans. Graph. 2019

Methods — techniques the papers use, named apart from their topics

machine learning cost model · 1.8two-phase search · 1.0memoization · 1.0hierarchical sampling · 1.0tree search · 0.8random program generation · 0.8beam search · 0.8program synthesis · 0.6multi-objective optimization · 0.6SIMD compilation · 0.6
YearPublicationVenuePosition
2025 A Computational Model of Human Vocal Imitation
Matthew Caren, Kartik Chandra, Josh Tenenbaum, Jonathan Ragan-Kelley, Karima Ma
CogSci5
2024 Sketching With Your Voice: "Non-Phonorealistic" Rendering of Sounds via Vocal Imitation
abstract
SA Conference Papers ’24, December 03–06, 2024, Tokyo, Japan
Matthew Caren, Kartik Chandra, Josh Tenenbaum, Jonathan Ragan-Kelley, Karima Ma
SIGGRAPH Asia5
2022 Searching for Fast Demosaicking Algorithms
abstract
We present a method to automatically synthesize efficient, high-quality demosaicking algorithms, across a range of computational budgets, given a loss function and training data. It performs a multi-objective, discrete-continuous optimization which simultaneously solves for the program structure and parameters that best tradeoff computational cost and image quality. We design the method to exploit domain-specific structure for search efficiency. We apply it to several tasks, including demosaicking both Bayer and Fuji X-Trans color filter patterns, as well as joint demosaicking and super-resolution. In a few days on 8 GPUs, it produces a family of algorithms that significantly improves image quality relative to the prior state-of-the-art across a range of computational budgets from 10 s to 1000 s of operations per pixel (1 dB–3 dB higher quality at the same cost, or 8.5–200× higher throughput at same or better quality). The resulting programs combine features of both classical and deep learning-based demosaicking algorithms into more efficient hybrid combinations, which are bandwidth-efficient and vectorizable by construction. Finally, our method automatically schedules and compiles all generated programs into optimized SIMD code for modern processors.
Karima Ma, Michaël Gharbi, Andrew Adams, Shoaib Kamil 0001, Tzu-Mao Li, Connelly Barnes, Jonathan Ragan-Kelley
ACM Trans. Graph.1
2021 Efficient automatic scheduling of imaging and vision pipelines for the GPU
abstract
We present a new algorithm to quickly generate high-performance GPU implementations of complex imaging and vision pipelines, directly from high-level Halide algorithm code. It is fully automatic, requiring no schedule templates or hand-optimized kernels. We address the scalability challenge of extending search-based automatic scheduling to map large real-world programs to the deep hierarchies of memory and parallelism on GPU architectures in reasonable compile time. We achieve this using (1) a two-phase search algorithm that first ‘freezes’ decisions for the lowest cost sections of a program, allowing relatively more time to be spent on the important stages, (2) a hierarchical sampling strategy that groups schedules based on their structural similarity, then samples representatives to be evaluated, allowing us to explore a large space with few samples, and (3) memoization of repeated partial schedules, amortizing their cost over all their occurrences. We guide the process with an efficient cost model combining machine learning, program analysis, and GPU architecture knowledge. We evaluate our method’s performance on a diverse suite of real-world imaging and vision pipelines. Our scalability optimizations lead to average compile time speedups of 49x (up to 530x). We find schedules that are on average 1.7x faster than existing automatic solutions (up to 5x), and competitive with what the best human experts were able to achieve in an active effort to beat our automatic results.
Luke Anderson 0001, Andrew Adams, Karima Ma, Tzu-Mao Li, Jonathan Ragan-Kelley
Proc. ACM Program. Lang.3
2019 Learning to optimize halide with tree search and random programs
abstract
We present a new algorithm to automatically schedule Halide programs for high-performance image processing and deep learning. We significantly improve upon the performance of previous methods, which considered a limited subset of schedules. We define a parameterization of possible schedules much larger than prior methods and use a variant of beam search to search over it. The search optimizes runtime predicted by a cost model based on a combination of new derived features and machine learning. We train the cost model by generating and featurizing hundreds of thousands of random programs and schedules. We show that this approach operates effectively with or without autotuning. It produces schedules which are on average almost twice as fast as the existing Halide autoscheduler without autotuning, or more than twice as fast with, and is the first automatic scheduling algorithm to significantly outperform human experts on average.
Andrew Adams, Karima Ma, Luke Anderson 0001, Riyadh Baghdadi, Tzu-Mao Li, Michaël Gharbi, Benoit Steiner, Kayvon Fatahalian, Frédo Durand, Jonathan Ragan-Kelley
ACM Trans. Graph.2