Ben Limpanukorn

dblp:376/9650 · DBLP profile ↗
← Back
2ranked-venue papers
1as first author
2since 2021 · last 2025
0009-0003-3652-384XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Software engineering, system software, and programming languages
2 papers
Software testing · 100%
Artificial intelligence
1 paper
Trustworthy machine learning · 100%

Topics — the 5 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Software testing › fuzzing › system software fuzzing
compiler fuzzing
0.912025
Fuzzing MLIR Compilers with Custom Mutation Synthesis · ICSE 2025
Software testing
compiler testing
0.912025
Fuzzing MLIR Compilers with Custom Mutation Synthesis · ICSE 2025
Software testing
fuzzing
0.912025
Fuzzing MLIR Compilers with Custom Mutation Synthesis · ICSE 2025
Software testing
metamorphic testing
0.912025
Chrysalis: A Lightweight Logging and Replay Framework for Metamorphic Testing in Python · ASE 2025
Machine learning › Trustworthy machine learning › fairness
fairness auditing
0.312025
Chrysalis: A Lightweight Logging and Replay Framework for Metamorphic Testing in Python · ASE 2025

Methods — techniques the papers use, named apart from their topics

invariant checking · 1.7input transformation · 1.7grammar-based fuzzing · 0.9custom mutation synthesis · 0.9
YearPublicationVenuePosition
2025 Fuzzing MLIR Compilers with Custom Mutation Synthesis
abstract
Compiler technologies in deep learning and domain-specific hardware acceleration are increasingly adopting extensible compiler frameworks such as Multi-Level Intermediate Representation (MLIR) to facilitate more efficient development. With MLIR, compiler developers can easily define their own custom IRs in the form of MLIR dialects. However, the diversity and rapid evolution of such custom IRs make it impractical to manually write a custom test generator for each dialect. To address this problem, we design a new test generator called SynthFuzz that combines grammar-based fuzzing with custom mutation synthesis. The key essence of SynthFuzz is two fold: (1) It automatically infers parameterized context-dependent custom mutations from existing test cases. (2) It then concretizes the mutation's content depending on the target context and reduces the chance of inserting invalid edits by performing$k$- ancestor and prefix/postfix matching. It obviates the need to manually define custom mutation operators for each dialect. We compare SynthFuzz to three baselines: Grammarinator-a grammar-based fuzzer without custom mutations, MLIRSmith-a custom test generator for MLIR core dialects, and NeuRI-a custom test generator for ML models with parameterization of tensor shapes. We conduct this comprehensive comparison on four different MLIR projects. Each project defines a new set of MLIR dialects where manually writing a custom test generator would take weeks of effort. Our evaluation shows that SynthFuzz on average improves MLIR dialect pair coverage by 1.75 ×, which increases branch coverage by 1.22 ×. Further, we show that our context dependent custom mutation increases the proportion of valid tests by up to 1.11 ×, indicating that SynthFuzz correctly concretizes its parameterized mutations with respect to the target context. Parameterization of the mutations reduces the fraction of tests violating the base MLIR constraints by 0.57 ×, increasing the time spent fuzzing dialect-specific code.
Ben Limpanukorn, Hong Jin Kang, Eric Zitong Zhou, Miryung Kim
ICSE1
2025 Chrysalis: A Lightweight Logging and Replay Framework for Metamorphic Testing in Python
abstract
Metamorphic testing (MT) is a powerful technique for software testing. We introduce Chrysalis, a lightweight, extensible logging and replay-based metamorphic testing framework in Python. Chrysalis allows developers to define custom input transformations and their associated invariants, then execute structured metamorphic testing campaigns. Its key innovation is a lightweight logging mechanism that records the full history of transformations applied to an input. This compact representation enables developers to not only identify test failures but also to replay the exact sequence of transformations leading to a bug, facilitating debugging. We demonstrate Chrysalis’s effectiveness through two case studies: auditing a machine learning model for fairness and assessing the robustness of large language models.A screencast demonstrating Chrysalis is available at: https://youtu.be/xJG4qghxlIs, and the source code is available at: https://github.com/Chrysalis-Test/Chrysalis.
Jai Parera, Nathan Huey, Ben Limpanukorn, Miryung Kim
ASE3