EDBT 2026 Demo / reviewers in the wild / expert
Nathanael Schärli
dblp:86/684
· DBLP profile ↗
14ranked-venue papers
4as first author
5since 2021 · last 2024
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 8 · 4 first-authorArtificial intelligence and machine learning · 6 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
5 papers |
Language models and text generation · 70% Information extraction and text analysis · 26% Vision and language · 4% | |
| Software engineering, system software, and programming languages
6 papers |
Compilers and program optimization · 34% Debugging and program repair · 20% Programming languages and type systems · 20% |
Topics — the 21 heaviest of 22, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation
compositional generalization |
0.9 | 2 | 2021 | *-CFQ: Analyzing the Scalability of Machine Learning on a Compositional Task · AAAI 2021 Measuring Compositional Generalization: A Comprehensive Method on Realistic Data · ICLR 2020 |
Compilers and program optimization
code generation |
0.8 | 1 | 2024 | Teaching Large Language Models to Self-Debug · ICLR 2024 |
Natural language and speech › Language models and text generation
complex reasoning |
0.7 | 1 | 2023 | Least-to-Most Prompting Enables Complex Reasoning in Large Language Models · ICLR 2023 |
Natural language and speech › Information extraction and text analysis › semantic parsing
compositional semantic parsing |
0.7 | 1 | 2023 | Compositional Semantic Parsing with Large Language Models · ICLR 2023 |
Natural language and speech › Language models and text generation
large language model evaluation |
0.7 | 1 | 2023 | Large Language Models Can Be Easily Distracted by Irrelevant Context · ICML 2023 |
Natural language and speech › Language models and text generation › mathematical reasoning
numerical reasoning |
0.7 | 1 | 2023 | Large Language Models Can Be Easily Distracted by Irrelevant Context · ICML 2023 |
Natural language and speech › Language models and text generation
prompting |
0.7 | 1 | 2023 | Least-to-Most Prompting Enables Complex Reasoning in Large Language Models · ICLR 2023 |
Natural language and speech › Information extraction and text analysis
semantic parsing |
0.7 | 1 | 2023 | Compositional Semantic Parsing with Large Language Models · ICLR 2023 |
Software testing › test infrastructure
benchmark construction |
0.4 | 1 | 2020 | Measuring Compositional Generalization: A Comprehensive Method on Realistic Data · ICLR 2020 |
Debugging and program repair
automated program repair |
0.2 | 1 | 2024 | Teaching Large Language Models to Self-Debug · ICLR 2024 |
Debugging and program repair › automated program repair
LLM-based program repair |
0.2 | 1 | 2024 | Teaching Large Language Models to Self-Debug · ICLR 2024 |
Computer vision › Vision and language › multimodal reasoning
compositional reasoning |
0.2 | 1 | 2023 | Compositional Semantic Parsing with Large Language Models · ICLR 2023 |
Programming languages and type systems › object-oriented programming
traits |
0.2 | 3 | 2006 | Traits: A mechanism for fine-grained reuse · ACM Trans. Program. Lang. Syst. 2006 Traits: Tools and Methodology · ICSE 2004 Applying traits to the smalltalk collection classes · OOPSLA 2003 |
Programming languages and type systems
object-oriented programming |
0.1 | 2 | 2004 | Traits: Tools and Methodology · ICSE 2004 Applying traits to the smalltalk collection classes · OOPSLA 2003 |
Software maintenance and evolution
code reuse |
0.1 | 1 | 2006 | Traits: A mechanism for fine-grained reuse · ACM Trans. Program. Lang. Syst. 2006 |
Programming languages and type systems
inheritance |
0.1 | 1 | 2006 | Traits: A mechanism for fine-grained reuse · ACM Trans. Program. Lang. Syst. 2006 |
Programming languages and type systems › type systems
dynamic typing |
0.0 | 1 | 2004 | Object-oriented encapsulation for dynamically typed languages · OOPSLA 2004 |
Programming languages and type systems › object-oriented programming
encapsulation |
0.0 | 1 | 2004 | Object-oriented encapsulation for dynamically typed languages · OOPSLA 2004 |
Software maintenance and evolution
software reuse |
0.0 | 1 | 2004 | Traits: Tools and Methodology · ICSE 2004 |
Programming languages and type systems
type systems |
0.0 | 1 | 2004 | Object-oriented encapsulation for dynamically typed languages · OOPSLA 2004 |
Software maintenance and evolution
refactoring |
0.0 | 1 | 2003 | Applying traits to the smalltalk collection classes · OOPSLA 2003 |
Methods — techniques the papers use, named apart from their topics
large language model · 2.1compositional generalization evaluation · 0.9rubber duck debugging · 0.8code execution · 0.8self-consistency decoding · 0.7prompting · 0.7least-to-most prompting · 0.7compositional parsing · 0.7transformer · 0.5scaling analysis · 0.5refactoring · 0.1formal model · 0.1traits · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Teaching Large Language Models to Self-DebugabstractLarge language models (LLMs) have achieved impressive performance on code generation. However, for complex programming tasks, generating the correct solution in one go becomes challenging, thus some prior works have designed program repair approaches to improve code generation performance. In this work, we propose self-debugging, which teaches a large language model to debug its predicted program. In particular, we demonstrate that self-debugging can teach the large language model to perform rubber duck debugging; i.e., without any human feedback on the code correctness or error messages, the model is able to identify its mistakes by leveraging code execution and explaining the generated code in natural language. Self-debugging achieves the state-of-the-art performance on several code generation benchmarks, including the Spider dataset for text-to-SQL generation, TransCoder for C++-to-Python translation, and MBPP for text-to-Python generation. On the Spider benchmark where there are no unit tests to verify the correctness of predictions, self-debugging with code explanation consistently improves the baseline by 2-3%, and improves the prediction accuracy on problems of the hardest level by 9%. On TransCoder and MBPP where unit tests are available, self-debugging improves the baseline accuracy by up to 12%. Meanwhile, by leveraging feedback messages and reusing failed predictions, self-debugging notably improves sample efficiency, and can match or outperform baseline models that generate more than 10$\times$ candidate programs. Maxwell Lin, Nathanael Schärli, Denny Zhou |
ICLR | 3 |
| 2023 | Compositional Semantic Parsing with Large Language Models
Andrew Drozdov, Nathanael Schärli, Ekin Akyürek, Nathan Scales, Xinying Song, Olivier Bousquet, Denny Zhou |
ICLR | 2 |
| 2023 | Least-to-Most Prompting Enables Complex Reasoning in Large Language Models
Denny Zhou, Nathanael Schärli, Le Hou, Jason Wei, Nathan Scales, Xuezhi Wang 0002, Dale Schuurmans, Claire Cui, Olivier Bousquet, Quoc V. Le, Ed H. Chi |
ICLR | 2 |
| 2023 | Large Language Models Can Be Easily Distracted by Irrelevant ContextabstractLarge language models have achieved impressive performance on various natural language processing tasks. However, so far they have been evaluated primarily on benchmarks where all information in the input context is relevant for solving the task. In this work, we investigate the *distractibility* of large language models, i.e., how the model prediction can be distracted by irrelevant context. In particular, we introduce Grade-School Math with Irrelevant Context (GSM-IC), an arithmetic reasoning dataset with irrelevant information in the problem description. We use this benchmark to measure the distractibility of different prompting techniques for large language models, and find that the model is easily distracted by irrelevant information. We also identify several approaches for mitigating this deficiency, such as decoding with self-consistency and adding to the prompt an instruction that tells the language model to ignore the irrelevant information. Freda Shi, Kanishka Misra, Nathan Scales, David Dohan, Ed H. Chi, Nathanael Schärli, Denny Zhou |
ICML | 7 |
| 2021 | *-CFQ: Analyzing the Scalability of Machine Learning on a Compositional TaskabstractWe present *-CFQ ("star-CFQ"): a suite of large-scale datasets of varying scope based on the CFQ semantic parsing benchmark, designed for principled investigation of the scalability of machine learning systems in a realistic compositional task setting. Using this suite, we conduct a series of experiments investigating the ability of Transformers to benefit from increased training data size under conditions of fixed computational cost. We show that compositional generalization remains a challenge at all training sizes, and we show that increasing the scope of natural language leads to consistently higher error rates, which are only partially offset by increased training data. We further show that while additional training data from a related domain improves the accuracy in data-starved situations, this improvement is limited and diminishes as the distance from the related domain to the target domain increases. Dmitry Tsarkov, Tibor Tihon, Nathan Scales, Nikola Momchev, Danila Sinopalnikov, Nathanael Schärli |
AAAI | 6 |
| 2020 | Measuring Compositional Generalization: A Comprehensive Method on Realistic Data
Daniel Keysers, Nathanael Schärli, Nathan Scales, Hylke Buisman, Daniel Furrer, Sergii Kashubin, Nikola Momchev, Danila Sinopalnikov, Lukasz Stafiniak, Tibor Tihon, Dmitry Tsarkov, Xiao Wang 0038, Marc van Zee, Olivier Bousquet |
ICLR | 2 |
| 2006 | Traits: A mechanism for fine-grained reuseabstractInheritance is well-known and accepted as a mechanism for reuse in object-oriented languages. Unfortunately, due to the coarse granularity of inheritance, it may be difficult to decompose an application into an optimal class hierarchy that maximizes software reuse. Existing schemes based on single inheritance, multiple inheritance, or mixins, all pose numerous problems for reuse. To overcome these problems we propose traits , pure units of reuse consisting only of methods. We develop a formal model of traits that establishes how traits can be composed, either to form other traits, or to form classes. We also outline an experimental validation in which we apply traits to refactor a nontrivial application into composable units. Stéphane Ducasse, Oscar Nierstrasz, Nathanael Schärli, Roel Wuyts, Andrew P. Black |
ACM Trans. Program. Lang. Syst. | 3 |
| 2005 | Uniform and safe metaclass composition
Stéphane Ducasse, Nathanael Schärli, Roel Wuyts |
Comput. Lang. Syst. Struct. | 2 |
| 2004 | Composable Encapsulation Policies
Nathanael Schärli, Stéphane Ducasse, Oscar Nierstrasz, Roel Wuyts |
ECOOP | 1 |
| 2004 | Traits: Tools and MethodologyabstractTraits are an object-oriented programming language construct that allow groups of methods to be named and reused in arbitrary places in an inheritance hierarchy. Classes can use methods from traits as well as defining their own methods and instance variables. Traits thus enable a new style of programming, in which traits rather than classes are the primary unit of reuse. However, the additional sub-structure provided by traits is always optional: a class written using traits can also be viewed as a flat collection of methods, with no change in its semantics. This paper describes the tool that supports these two alternate views of a class, called the traits browser, and the programming methodology that we are starting to develop around the use of traits. Andrew P. Black, Nathanael Schärli |
ICSE | 2 |
| 2004 | Object-oriented encapsulation for dynamically typed languagesabstractEncapsulation in object-oriented languages has traditionally been based on static type systems. As a consequence, dynamically-typed languages have only limited support for encapsulation. This is surprising, considering that encapsulation is one of the most fundamental and important concepts behind object-oriented programming and that it is essential for writing programs that are maintainable and reliable, and that remain robust as they evolve. Nathanael Schärli, Andrew P. Black, Stéphane Ducasse |
OOPSLA | 1 |
| 2004 | A browser for incremental programming
Nathanael Schärli, Andrew P. Black |
Comput. Lang. Syst. Struct. | 1 |
| 2003 | Traits: Composable Units of Behaviour
Nathanael Schärli, Stéphane Ducasse, Oscar Nierstrasz, Andrew P. Black |
ECOOP | 1 |
| 2003 | Applying traits to the smalltalk collection classesabstractTraits are a programming language technology that promote the reuse of methods between unrelated classes. This paper reports on a refactoring of the Smalltalk collections classes using traits. The original collection classes contained much duplication of code; traits let us remove all of it. We also found places where the protocols of the collections lacked uniformity; traits allowed us to correct these non-uniformities without code duplication.Traits also make it possible to reuse fragments of collection code outside of the existing hierarchy; for example, they make it easy to convert other collection-like things into true collections. Our refactoring reduced the number of methods in the collection classes by approximately 10 per cent. More importantly, understandability maintainability and reusability of the code were significantly improved. Andrew P. Black, Nathanael Schärli, Stéphane Ducasse |
OOPSLA | 2 |