Jesse Berent

dblp:81/397 · DBLP profile ↗
← Back
12ranked-venue papers
4as first author
6since 2021 · last 2023
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 1 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 4 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
7 papers
Trustworthy machine learning · 32% Learning paradigms · 19% Vision and language · 18%

Topics — the 11 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Trustworthy machine learning › robustness
learning with noisy labels
1.732023
When does Privileged information Explain Away Label Noise? · ICML 2023
Transfer and Marginalize: Explaining Away Label Noise with Privileged Information · ICML 2022
Correlated Input-Dependent Label Noise in Large-Scale Image Classification · CVPR 2021
Machine learning › Learning paradigms › supervised learning
privileged information
1.222023
When does Privileged information Explain Away Label Noise? · ICML 2023
Transfer and Marginalize: Explaining Away Label Noise with Privileged Information · ICML 2022
Computer vision › Vision and language › cross-modal supervision
caption-based supervision
1.022023
Learning to Overcome Noise in Weak Caption Supervision for Object Detection · IEEE Trans. Pattern Anal. Mach. Intell. 2023
Cap2Det: Learning to Amplify Weak Caption Supervision for Object Detection · ICCV 2019
Computer vision › Image recognition and object detection › object detection
weakly supervised object detection
1.022023
Learning to Overcome Noise in Weak Caption Supervision for Object Detection · IEEE Trans. Pattern Anal. Mach. Intell. 2023
Cap2Det: Learning to Amplify Weak Caption Supervision for Object Detection · ICCV 2019
Machine learning › Trustworthy machine learning
robustness
0.712023
When does Privileged information Explain Away Label Noise? · ICML 2023
Machine learning › Trustworthy machine learning
uncertainty estimation
0.712023
Massively Scaling Heteroscedastic Classifiers · ICLR 2023
Computer vision › Vision and language
vision-language pretraining
0.712023
Three Towers: Flexible Contrastive Learning with Pretrained Image Models · NeurIPS 2023
Machine learning › Learning paradigms
supervised learning
0.612022
Transfer and Marginalize: Explaining Away Label Noise with Privileged Information · ICML 2022
Computer vision › Image recognition and object detection
object detection
0.412019
Cap2Det: Learning to Amplify Weak Caption Supervision for Object Detection · ICCV 2019
Machine learning › Representation and self-supervised learning
contrastive learning
0.212023
Three Towers: Flexible Contrastive Learning with Pretrained Image Models · NeurIPS 2023
Computer vision › Image recognition and object detection › image classification
large-scale image classification
0.112021
Correlated Input-Dependent Label Noise in Large-Scale Image Classification · CVPR 2021

Methods — techniques the papers use, named apart from their topics

pseudo-labeling · 0.7privileged information · 0.7learning shortcut · 0.7knowledge distillation · 0.7iterative training · 0.7contrastive learning · 0.7weight sharing · 0.6marginalization · 0.6probabilistic modeling · 0.5multivariate normal latent variable · 0.5
YearPublicationVenuePosition
2023 Massively Scaling Heteroscedastic Classifiers
Mark Collier, Rodolphe Jenatton, Basil Mustafa, Neil Houlsby, Jesse Berent, Effrosyni Kokiopoulou
ICLR5
2023 When does Privileged information Explain Away Label Noise?
abstract
Leveraging privileged information (PI), or features available during training but not at test time, has recently been shown to be an effective method for addressing label noise. However, the reasons for its effectiveness are not well understood. In this study, we investigate the role played by different properties of the PI in explaining away label noise. Through experiments on multiple datasets with real PI (CIFAR-N/H) and a new large-scale benchmark ImageNet-PI, we find that PI is most helpful when it allows networks to easily distinguish clean from noisy data, while enabling a learning shortcut to memorize the noisy examples. Interestingly, when PI becomes too predictive of the target label, PI methods often perform worse than their no-PI baselines. Based on these findings, we propose several enhancements to the state-of-the-art PI methods and demonstrate the potential of PI as a means of tackling label noise. Finally, we show how we can easily combine the resulting PI approaches with existing no-PI techniques designed to deal with label noise.
Guillermo Ortiz-Jiménez, Mark Collier, Anant Nawalgaria, Alexander D'Amour, Jesse Berent, Rodolphe Jenatton, Effrosyni Kokiopoulou
ICML5
2023 Three Towers: Flexible Contrastive Learning with Pretrained Image Models
abstract
We introduce Three Towers (3T), a flexible method to improve the contrastive learning of vision-language models by incorporating pretrained image classifiers. While contrastive models are usually trained from scratch, LiT (Zhai et al., 2022) has recently shown performance gains from using pretrained classifier embeddings. However, LiT directly replaces the image tower with the frozen embeddings, excluding any potential benefits from training the image tower contrastively. With 3T, we propose a more flexible strategy that allows the image tower to benefit from both pretrained embeddings and contrastive training. To achieve this, we introduce a third tower that contains the frozen pretrained embeddings, and we encourage alignment between this third tower and the main image-text towers. Empirically, 3T consistently improves over LiT and the CLIP-style from-scratch baseline for retrieval tasks. For classification, 3T reliably improves over the from-scratch baseline, and while it underperforms relative to LiT for JFT-pretrained models, it outperforms LiT for ImageNet-21k and Places365 pretraining.
Jannik Kossen, Mark Collier, Basil Mustafa, Xiao Wang 0038, Xiaohua Zhai, Lucas Beyer, Andreas Steiner 0001, Jesse Berent, Rodolphe Jenatton, Effrosyni Kokiopoulou
NeurIPS8
2023 Learning to Overcome Noise in Weak Caption Supervision for Object Detection
abstract
We propose the first mechanism to train object detection models from weak supervision in the form of captions at the image level. Language-based supervision for detection is appealing and inexpensive: many blogs with images and descriptive text written by human users exist. However, there is significant noise in this supervision: captions do not mention all objects that are shown, and may mention extraneous concepts. We first propose a technique to determine which image-caption pairs provide suitable signal for supervision. We further propose several complementary mechanisms to extract image-level pseudo labels for training from the caption. Finally, we train an iterative weakly-supervised object detection model from these image-level pseudo labels. We use captions from four datasets (COCO, Flickr30K, MIRFlickr1M, and Conceptual Captions) whose level of noise varies. We evaluate our approach on two object detection datasets. Weighting the labels extracted from different captions provides a boost over treating all captions equally. Further, our primary proposed technique for inferring pseudo labels for training at the image level, outperforms alternative techniques under a wide variety of settings. Both techniques generalize to datasets beyond the one they were trained on.
Mesut Erhan Unal, Keren Ye, Christopher Thomas 0004, Adriana Kovashka, Wei Li 0044, Danfeng Qin, Jesse Berent
IEEE Trans. Pattern Anal. Mach. Intell.8
2022 Transfer and Marginalize: Explaining Away Label Noise with Privileged Information
abstract
Supervised learning datasets often have privileged information, in the form of features which are available at training time but are not available at test time e.g. the ID of the annotator that provided the label. We argue that privileged information is useful for explaining away label noise, thereby reducing the harmful impact of noisy labels. We develop a simple and efficient method for supervised learning with neural networks: it transfers via weight sharing the knowledge learned with privileged information and approximately marginalizes over privileged information at test time. Our method, TRAM (TRansfer and Marginalize), has minimal training time overhead and has the same test-time cost as not using privileged information. TRAM performs strongly on CIFAR-10H, ImageNet and Civil Comments benchmarks.
Mark Collier, Rodolphe Jenatton, Effrosyni Kokiopoulou, Jesse Berent
ICML4
2021 Correlated Input-Dependent Label Noise in Large-Scale Image Classification
abstract
Large scale image classification datasets often contain noisy labels. We take a principled probabilistic approach to modelling input-dependent, also known as heteroscedastic, label noise in these datasets. We place a multivariate Nor-mal distributed latent variable on the final hidden layer of a neural network classifier. The covariance matrix of this latent variable, models the aleatoric uncertainty due to label noise. We demonstrate that the learned covariance structure captures known sources of label noise between semantically similar and co-occurring classes. Compared to standard neural network training and other baselines, we show significantly improved accuracy on Imagenet ILSVRC 2012 79.3% (+ 2.6%), Imagenet-21k 47.0% (+ 1.1%) and JFT 64.7% (+ 1.6%). We set a new state-of-the-art result on WebVision 1.0 with 76.6% top-1 accuracy. These datasets range from over 1M to over 300M training examples and from 1k classes to more than 21k classes. Our method is simple to use, and we provide an implementation that is a drop-in replacement for the final fully-connected layer in a deep classifier.
Mark Collier, Basil Mustafa, Effrosyni Kokiopoulou, Rodolphe Jenatton, Jesse Berent
CVPR5
2020 Task-Aware Performance Prediction for Efficient Architecture Search
Effrosyni Kokiopoulou, Anja Hauth, Luciano Sbaiz, Andrea Gesmundo, Gábor Bartók, Jesse Berent
ECAI6
2019 Cap2Det: Learning to Amplify Weak Caption Supervision for Object Detection
abstract
Learning to localize and name object instances is a fundamental problem in vision, but state-of-the-art approaches rely on expensive bounding box supervision. While weakly supervised detection (WSOD) methods relax the need for boxes to that of image-level annotations, even cheaper supervision is naturally available in the form of unstructured textual descriptions that users may freely provide when uploading image content. However, straightforward approaches to using such data for WSOD wastefully discard captions that do not exactly match object names. Instead, we show how to squeeze the most information out of these captions by training a text-only classifier that generalizes beyond dataset boundaries. Our discovery provides an opportunity for learning detection models from noisy but more abundant and freely-available caption data. We also validate our model on three classic object detection benchmarks and achieve state-of-the-art WSOD performance. Our code is available at https://github.com/yekeren/Cap2Det.
Keren Ye, Adriana Kovashka, Wei Li 0044, Danfeng Qin, Jesse Berent
ICCV6
2009 Adaptive layer extraction for image based rendering
abstract
Image based rendering is a promising way to produce arbitrary views of a scene using images instead of object models. However, depth variations and occlusions cause blurring in the rendered images. The solution is to use some geometrical information in order to steer the interpolation filters according to the depth. The level of detail of this geometry is often predetermined. In this paper, we present a method for extracting depth layers in the presence of occlusions for image based rendering. Moreover, we show how the layer extraction can be made to estimate depth layers in an adaptive manner, based on the spectral analysis of the plenoptic function. The rendering system therefore automatically adapts the number of depth layers based on the scene and the spacing of the sample cameras.
Jesse Berent, Pier Luigi Dragotti, Mike Brookes
MMSP1
2007 Unsupervised Extraction of Coherent Regions for Image Based Rendering
abstract
Image based rendering using undersampled light fields suffe rs from aliasing effects. These effects can be drastically reduced by usi ng some geometric information. In pop-up light field rendering [18], the scene is segmented into coherent layers, usually corresponding to approximately planar regions, that can be rendered free of aliasing. As opposed to the supervised method in the pop-up light field, we propose an unsupervised extractio n of coherent regions. The problem is posed in a multidimensional variational framework using the level set method [16]. Since the segmentation is done jointly over all the images, coherence can be imposed throughout the data. However, instead of using active hypersurfaces, we derive a semi-parametric methodology that takes into account the constraints imposed by the camera setup and the occlusion ordering. The resulting framework is a global multidimensional region competition that is consistent in all the imag es and efficiently handles occlusions. We show the validity of the method with some captured multi-view datasets. Other special effects by coherent reg ion manipulation are also demonstrated.
Jesse Berent, Pier Luigi Dragotti
BMVC1
2006 Perfect Reconstruction Schemes for Sampling Piecewise Sinusoidal Signals
abstract
Consider sampling a signal that is piecewise sinusoidal. Classical sampling theory does not enable a perfect reconstruction of the continuous time signal since the band is not limited (C.E. Shannon, 1949). However, we show that it is still possible to recover all the parameters of the sinusoids and the exact locations of the discontinuities using the annihilating filter method and recently developed Finite Rate of Innovation (FRI) sampling schemes (M. Vetterli et al., 2002) (P.L. Dragotti et al., 2005). Moreover, we show that there is a tradeoff between the number of sinusoids per piece and the proximity of the discontinuities in order to have a unique solution. This result recalls a sort of uncertainty principle
Jesse Berent, Pier Luigi Dragotti
ICASSP (3)1
2006 Segmentation of Epipolar-Plane Image Volumes with Occlusion and Disocclusion Competition
abstract
Consider a dense array of cameras uniformly distributed along a line. A solid block of 3D data can be constructed by arranging the images into a stack. This volume, also known as the epipolar-plane image volume, contains highly structured data that can be segmented for object removal, insertion and compression. In this paper, we propose a segmentation scheme that takes fully advantage of the known geometry in order to model occlusions explicitly as a result of disparity. Moreover, we include this knowledge into an energy minimization scheme based on region competition with active contours. Instead of extracting layers sequentially from front to back, each layer is made to compete with the regions it is going to occlude and the ones it is going to disocclude. This enables a virtually unsupervised segmentation
Jesse Berent, Pier Luigi Dragotti
MMSP1