Gabriel Orlanski

dblp:294/6286 · DBLP profile ↗
← Back
1ranked-venue papers
1as first author
1since 2021 · last 2023
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Software engineering, system software, and programming languages
1 paper
Program synthesis and code generation · 100%
Artificial intelligence
1 paper
Language models and text generation · 100%

Topics — the 4 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Program synthesis and code generation
code generation with language models
0.712023
Measuring the Impact of Programming Language Distribution · ICML 2023
Program synthesis and code generation
code translation
0.712023
Measuring the Impact of Programming Language Distribution · ICML 2023
Program synthesis and code generation › code generation with language models
multilingual code generation
0.712023
Measuring the Impact of Programming Language Distribution · ICML 2023
Natural language and speech › Language models and text generation
large language model
0.212023
Measuring the Impact of Programming Language Distribution · ICML 2023

Methods — techniques the papers use, named apart from their topics

language model training · 1.3execution-based evaluation · 1.3
YearPublicationVenuePosition
2023 Measuring the Impact of Programming Language Distribution
abstract
Current benchmarks for evaluating neural code models focus on only a small subset of programming languages, excluding many popular languages such as Go or Rust. To ameliorate this issue, we present the BabelCode framework for execution-based evaluation of any benchmark in any language. BabelCode enables new investigations into the qualitative performance of models' memory, runtime, and individual test case results. Additionally, we present a new code translation dataset called Translating Python Programming Puzzles (TP3) from the Python Programming Puzzles (Schuster et al., 2021) benchmark that involves translating expert-level python functions to any language. With both BabelCode and the TP3 benchmark, we investigate if balancing the distributions of 14 languages in a training dataset improves a large language model's performance on low-resource languages. Training a model on a balanced corpus results in, on average, 12.34% higher $pass@k$ across all tasks and languages compared to the baseline. We find that this strategy achieves 66.48% better $pass@k$ on low-resource languages at the cost of only a 12.94% decrease to high-resource languages. In our three translation tasks, this strategy yields, on average, 30.77% better low-resource $pass@k$ while having 19.58% worse high-resource $pass@k$.
Gabriel Orlanski, Kefan Xiao, Xavier Garcia, Jeffrey Hui, Joshua Howland, Jonathan Malmaud, Jacob Austin, Rishabh Singh, Michele Catasta
ICML1