Bart van Merrienboer

dblp:147/5356 · also Bart van Merriënboer · DBLP profile ↗
← Back
6ranked-venue papers
2as first author
2since 2021 · last 2023
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 2 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Transfer learning and domain adaptation · 36% Trustworthy machine learning · 18% Deep learning architectures and training · 16%
Software engineering, system software, and programming languages
2 papers
Compilers and program optimization · 100%

Topics — the 10 heaviest of 11, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Transfer learning and domain adaptation
domain shift
0.712023
In Search for a Generalizable Method for Source Free Domain Adaptation · ICML 2023
Machine learning › Trustworthy machine learning › out-of-distribution generalization
natural distribution shift
0.712023
In Search for a Generalizable Method for Source Free Domain Adaptation · ICML 2023
Machine learning › Transfer learning and domain adaptation › domain adaptation
source-free domain adaptation
0.712023
In Search for a Generalizable Method for Source Free Domain Adaptation · ICML 2023
Compilers and program optimization
automatic differentiation
0.722018
Tangent: Automatic differentiation using source-code transformation for dynamically typed array programming · NeurIPS 2018
Automatic differentiation in ML: Where we are and where we should be going · NeurIPS 2018
Compilers and program optimization › program transformation
source-to-source transformation
0.722018
Tangent: Automatic differentiation using source-code transformation for dynamically typed array programming · NeurIPS 2018
Automatic differentiation in ML: Where we are and where we should be going · NeurIPS 2018
Machine learning › Deep learning architectures and training › architecture learning
network growth
0.612022
GradMax: Growing Neural Networks using Gradient Information · ICLR 2022
Compilers and program optimization
intermediate representation
0.312018
Automatic differentiation in ML: Where we are and where we should be going · NeurIPS 2018
Natural language and speech › Machine translation
neural machine translation
0.212014
Learning Phrase Representations using RNN Encoder-Decoder for Statistical Machine Translation · EMNLP 2014
Machine learning › Representation and self-supervised learning › text embedding
phrase representation learning
0.212014
Learning Phrase Representations using RNN Encoder-Decoder for Statistical Machine Translation · EMNLP 2014
Natural language and speech › Machine translation
statistical machine translation
0.212014
Learning Phrase Representations using RNN Encoder-Decoder for Statistical Machine Translation · EMNLP 2014

Methods — techniques the papers use, named apart from their topics

unlabeled data adaptation · 0.7gradient information · 0.6source transformation · 0.3source code transformation · 0.3operator overloading · 0.3gradient surgery · 0.3functional programming · 0.3recurrent neural network · 0.2encoder-decoder · 0.2
YearPublicationVenuePosition
2023 In Search for a Generalizable Method for Source Free Domain Adaptation
abstract
Source-free domain adaptation (SFDA) is compelling because it allows adapting an off-the-shelf model to a new domain using only unlabelled data. In this work, we apply existing SFDA techniques to a challenging set of naturally-occurring distribution shifts in bioacoustics, which are very different from the ones commonly studied in computer vision. We find existing methods perform differently relative to each other than observed in vision benchmarks, and sometimes perform worse than no adaptation at all. We propose a new simple method which outperforms the existing methods on our new shifts while exhibiting strong performance on a range of vision datasets. Our findings suggest that existing SFDA methods are not as generalizable as previously thought and that considering diverse modalities can be a useful avenue for designing more robust models.
Malik Boudiaf, Tom Denton, Bart van Merrienboer, Vincent Dumoulin, Eleni Triantafillou
ICML3
2022 GradMax: Growing Neural Networks using Gradient Information
Utku Evci, Bart van Merrienboer, Thomas Unterthiner, Fabian Pedregosa, Max Vladymyrov
ICLR2
2020 On the interplay between noise and curvature and its effect on optimization and generalization
abstract
The speed at which one can minimize an expected loss using stochastic methods depends on two properties: the curvature of the loss and the variance of the gradients. While most previous works focus on one or the other of these properties, we explore how their interaction affects optimization speed. Further, as the ultimate goal is good generalization performance, we clarify how both curvature and noise are relevant to properly estimate the generalization gap. Realizing that the limitations of some existing works stems from a confusion between these matrices, we also clarify the distinction between the Fisher matrix, the Hessian, and the covariance matrix of the gradients.
Valentin Thomas, Fabian Pedregosa, Bart van Merrienboer, Pierre-Antoine Manzagol, Yoshua Bengio, Nicolas Le Roux
AISTATS3
2018 Automatic differentiation in ML: Where we are and where we should be going
abstract
We review the current state of automatic differentiation (AD) for array programming in machine learning (ML), including the different approaches such as operator overloading (OO) and source transformation (ST) used for AD, graph-based intermediate representations for programs, and source languages. Based on these insights, we introduce a new graph-based intermediate representation (IR) which specifically aims to efficiently support fully-general AD for array programming. Unlike existing dataflow programming representations in ML frameworks, our IR naturally supports function calls, higher-order functions and recursion, making ML models easier to implement. The ability to represent closures allows us to perform AD using ST without a tape, making the resulting derivative (adjoint) program amenable to ahead-of-time optimization using tools from functional language compilers, and enabling higher-order derivatives. Lastly, we introduce a proof of concept compiler toolchain called Myia which uses a subset of Python as a front end.
Bart van Merrienboer, Olivier Breuleux, Arnaud Bergeron, Pascal Lamblin
NeurIPS1
2018 Tangent: Automatic differentiation using source-code transformation for dynamically typed array programming
abstract
The need to efficiently calculate first- and higher-order derivatives of increasingly complex models expressed in Python has stressed or exceeded the capabilities of available tools. In this work, we explore techniques from the field of automatic differentiation (AD) that can give researchers expressive power, performance and strong usability. These include source-code transformation (SCT), flexible gradient surgery, efficient in-place array operations, and higher-order derivatives. We implement and demonstrate these ideas in the Tangent software library for Python, the first AD framework for a dynamic language that uses SCT.
Bart van Merrienboer, Dan Moldovan, Alexander B. Wiltschko
NeurIPS1
2014 Learning Phrase Representations using RNN Encoder-Decoder for Statistical Machine Translation
abstract
Kyunghyun Cho, Bart van Merriënboer, Caglar Gulcehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, Yoshua Bengio. Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2014.
Kyunghyun Cho, Bart van Merrienboer, Caglar Gulcehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, Yoshua Bengio
EMNLP2