VLDB 2026 Research / reviewers in the wild / expert
Jesse Berent
dblp:81/397
· DBLP profile ↗
12ranked-venue papers
4as first author
6since 2021 · last 2023
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 1 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 4 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
7 papers |
Trustworthy machine learning · 32% Learning paradigms · 19% Vision and language · 18% |
Topics — the 11 heaviest of 13, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Trustworthy machine learning › robustness
learning with noisy labels |
1.7 | 3 | 2023 | When does Privileged information Explain Away Label Noise? · ICML 2023 Transfer and Marginalize: Explaining Away Label Noise with Privileged Information · ICML 2022 Correlated Input-Dependent Label Noise in Large-Scale Image Classification · CVPR 2021 |
Machine learning › Learning paradigms › supervised learning
privileged information |
1.2 | 2 | 2023 | When does Privileged information Explain Away Label Noise? · ICML 2023 Transfer and Marginalize: Explaining Away Label Noise with Privileged Information · ICML 2022 |
Computer vision › Vision and language › cross-modal supervision
caption-based supervision |
1.0 | 2 | 2023 | Learning to Overcome Noise in Weak Caption Supervision for Object Detection · IEEE Trans. Pattern Anal. Mach. Intell. 2023 Cap2Det: Learning to Amplify Weak Caption Supervision for Object Detection · ICCV 2019 |
Computer vision › Image recognition and object detection › object detection
weakly supervised object detection |
1.0 | 2 | 2023 | Learning to Overcome Noise in Weak Caption Supervision for Object Detection · IEEE Trans. Pattern Anal. Mach. Intell. 2023 Cap2Det: Learning to Amplify Weak Caption Supervision for Object Detection · ICCV 2019 |
Machine learning › Trustworthy machine learning
robustness |
0.7 | 1 | 2023 | When does Privileged information Explain Away Label Noise? · ICML 2023 |
Machine learning › Trustworthy machine learning
uncertainty estimation |
0.7 | 1 | 2023 | Massively Scaling Heteroscedastic Classifiers · ICLR 2023 |
Computer vision › Vision and language
vision-language pretraining |
0.7 | 1 | 2023 | Three Towers: Flexible Contrastive Learning with Pretrained Image Models · NeurIPS 2023 |
Machine learning › Learning paradigms
supervised learning |
0.6 | 1 | 2022 | Transfer and Marginalize: Explaining Away Label Noise with Privileged Information · ICML 2022 |
Computer vision › Image recognition and object detection
object detection |
0.4 | 1 | 2019 | Cap2Det: Learning to Amplify Weak Caption Supervision for Object Detection · ICCV 2019 |
Machine learning › Representation and self-supervised learning
contrastive learning |
0.2 | 1 | 2023 | Three Towers: Flexible Contrastive Learning with Pretrained Image Models · NeurIPS 2023 |
Computer vision › Image recognition and object detection › image classification
large-scale image classification |
0.1 | 1 | 2021 | Correlated Input-Dependent Label Noise in Large-Scale Image Classification · CVPR 2021 |
Methods — techniques the papers use, named apart from their topics
pseudo-labeling · 0.7privileged information · 0.7learning shortcut · 0.7knowledge distillation · 0.7iterative training · 0.7contrastive learning · 0.7weight sharing · 0.6marginalization · 0.6probabilistic modeling · 0.5multivariate normal latent variable · 0.5
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Massively Scaling Heteroscedastic Classifiers
Mark Collier, Rodolphe Jenatton, Basil Mustafa, Neil Houlsby, Jesse Berent, Effrosyni Kokiopoulou |
ICLR | 5 |
| 2023 | When does Privileged information Explain Away Label Noise?abstractLeveraging privileged information (PI), or features available during training but not at test time, has recently been shown to be an effective method for addressing label noise. However, the reasons for its effectiveness are not well understood. In this study, we investigate the role played by different properties of the PI in explaining away label noise. Through experiments on multiple datasets with real PI (CIFAR-N/H) and a new large-scale benchmark ImageNet-PI, we find that PI is most helpful when it allows networks to easily distinguish clean from noisy data, while enabling a learning shortcut to memorize the noisy examples. Interestingly, when PI becomes too predictive of the target label, PI methods often perform worse than their no-PI baselines. Based on these findings, we propose several enhancements to the state-of-the-art PI methods and demonstrate the potential of PI as a means of tackling label noise. Finally, we show how we can easily combine the resulting PI approaches with existing no-PI techniques designed to deal with label noise. Guillermo Ortiz-Jiménez, Mark Collier, Anant Nawalgaria, Alexander D'Amour, Jesse Berent, Rodolphe Jenatton, Effrosyni Kokiopoulou |
ICML | 5 |
| 2023 | Three Towers: Flexible Contrastive Learning with Pretrained Image ModelsabstractWe introduce Three Towers (3T), a flexible method to improve the contrastive learning of vision-language models by incorporating pretrained image classifiers. While contrastive models are usually trained from scratch, LiT (Zhai et al., 2022) has recently shown performance gains from using pretrained classifier embeddings. However, LiT directly replaces the image tower with the frozen embeddings, excluding any potential benefits from training the image tower contrastively. With 3T, we propose a more flexible strategy that allows the image tower to benefit from both pretrained embeddings and contrastive training. To achieve this, we introduce a third tower that contains the frozen pretrained embeddings, and we encourage alignment between this third tower and the main image-text towers. Empirically, 3T consistently improves over LiT and the CLIP-style from-scratch baseline for retrieval tasks. For classification, 3T reliably improves over the from-scratch baseline, and while it underperforms relative to LiT for JFT-pretrained models, it outperforms LiT for ImageNet-21k and Places365 pretraining. Jannik Kossen, Mark Collier, Basil Mustafa, Xiao Wang 0038, Xiaohua Zhai, Lucas Beyer, Andreas Steiner 0001, Jesse Berent, Rodolphe Jenatton, Effrosyni Kokiopoulou |
NeurIPS | 8 |
| 2023 | Learning to Overcome Noise in Weak Caption Supervision for Object DetectionabstractWe propose the first mechanism to train object detection models from weak supervision in the form of captions at the image level. Language-based supervision for detection is appealing and inexpensive: many blogs with images and descriptive text written by human users exist. However, there is significant noise in this supervision: captions do not mention all objects that are shown, and may mention extraneous concepts. We first propose a technique to determine which image-caption pairs provide suitable signal for supervision. We further propose several complementary mechanisms to extract image-level pseudo labels for training from the caption. Finally, we train an iterative weakly-supervised object detection model from these image-level pseudo labels. We use captions from four datasets (COCO, Flickr30K, MIRFlickr1M, and Conceptual Captions) whose level of noise varies. We evaluate our approach on two object detection datasets. Weighting the labels extracted from different captions provides a boost over treating all captions equally. Further, our primary proposed technique for inferring pseudo labels for training at the image level, outperforms alternative techniques under a wide variety of settings. Both techniques generalize to datasets beyond the one they were trained on. Mesut Erhan Unal, Keren Ye, Christopher Thomas 0004, Adriana Kovashka, Wei Li 0044, Danfeng Qin, Jesse Berent |
IEEE Trans. Pattern Anal. Mach. Intell. | 8 |
| 2022 | Transfer and Marginalize: Explaining Away Label Noise with Privileged InformationabstractSupervised learning datasets often have privileged information, in the form of features which are available at training time but are not available at test time e.g. the ID of the annotator that provided the label. We argue that privileged information is useful for explaining away label noise, thereby reducing the harmful impact of noisy labels. We develop a simple and efficient method for supervised learning with neural networks: it transfers via weight sharing the knowledge learned with privileged information and approximately marginalizes over privileged information at test time. Our method, TRAM (TRansfer and Marginalize), has minimal training time overhead and has the same test-time cost as not using privileged information. TRAM performs strongly on CIFAR-10H, ImageNet and Civil Comments benchmarks. Mark Collier, Rodolphe Jenatton, Effrosyni Kokiopoulou, Jesse Berent |
ICML | 4 |
| 2021 | Correlated Input-Dependent Label Noise in Large-Scale Image ClassificationabstractLarge scale image classification datasets often contain noisy labels. We take a principled probabilistic approach to modelling input-dependent, also known as heteroscedastic, label noise in these datasets. We place a multivariate Nor-mal distributed latent variable on the final hidden layer of a neural network classifier. The covariance matrix of this latent variable, models the aleatoric uncertainty due to label noise. We demonstrate that the learned covariance structure captures known sources of label noise between semantically similar and co-occurring classes. Compared to standard neural network training and other baselines, we show significantly improved accuracy on Imagenet ILSVRC 2012 79.3% (+ 2.6%), Imagenet-21k 47.0% (+ 1.1%) and JFT 64.7% (+ 1.6%). We set a new state-of-the-art result on WebVision 1.0 with 76.6% top-1 accuracy. These datasets range from over 1M to over 300M training examples and from 1k classes to more than 21k classes. Our method is simple to use, and we provide an implementation that is a drop-in replacement for the final fully-connected layer in a deep classifier. Mark Collier, Basil Mustafa, Effrosyni Kokiopoulou, Rodolphe Jenatton, Jesse Berent |
CVPR | 5 |
| 2020 | Task-Aware Performance Prediction for Efficient Architecture Search
Effrosyni Kokiopoulou, Anja Hauth, Luciano Sbaiz, Andrea Gesmundo, Gábor Bartók, Jesse Berent |
ECAI | 6 |
| 2019 | Cap2Det: Learning to Amplify Weak Caption Supervision for Object DetectionabstractLearning to localize and name object instances is a fundamental problem in vision, but state-of-the-art approaches rely on expensive bounding box supervision. While weakly supervised detection (WSOD) methods relax the need for boxes to that of image-level annotations, even cheaper supervision is naturally available in the form of unstructured textual descriptions that users may freely provide when uploading image content. However, straightforward approaches to using such data for WSOD wastefully discard captions that do not exactly match object names. Instead, we show how to squeeze the most information out of these captions by training a text-only classifier that generalizes beyond dataset boundaries. Our discovery provides an opportunity for learning detection models from noisy but more abundant and freely-available caption data. We also validate our model on three classic object detection benchmarks and achieve state-of-the-art WSOD performance. Our code is available at https://github.com/yekeren/Cap2Det. Keren Ye, Adriana Kovashka, Wei Li 0044, Danfeng Qin, Jesse Berent |
ICCV | 6 |
| 2009 | Adaptive layer extraction for image based renderingabstractImage based rendering is a promising way to produce arbitrary views of a scene using images instead of object models. However, depth variations and occlusions cause blurring in the rendered images. The solution is to use some geometrical information in order to steer the interpolation filters according to the depth. The level of detail of this geometry is often predetermined. In this paper, we present a method for extracting depth layers in the presence of occlusions for image based rendering. Moreover, we show how the layer extraction can be made to estimate depth layers in an adaptive manner, based on the spectral analysis of the plenoptic function. The rendering system therefore automatically adapts the number of depth layers based on the scene and the spacing of the sample cameras. Jesse Berent, Pier Luigi Dragotti, Mike Brookes |
MMSP | 1 |
| 2007 | Unsupervised Extraction of Coherent Regions for Image Based RenderingabstractImage based rendering using undersampled light fields suffe rs from aliasing effects. These effects can be drastically reduced by usi ng some geometric information. In pop-up light field rendering [18], the scene is segmented into coherent layers, usually corresponding to approximately planar regions, that can be rendered free of aliasing. As opposed to the supervised method in the pop-up light field, we propose an unsupervised extractio n of coherent regions. The problem is posed in a multidimensional variational framework using the level set method [16]. Since the segmentation is done jointly over all the images, coherence can be imposed throughout the data. However, instead of using active hypersurfaces, we derive a semi-parametric methodology that takes into account the constraints imposed by the camera setup and the occlusion ordering. The resulting framework is a global multidimensional region competition that is consistent in all the imag es and efficiently handles occlusions. We show the validity of the method with some captured multi-view datasets. Other special effects by coherent reg ion manipulation are also demonstrated. Jesse Berent, Pier Luigi Dragotti |
BMVC | 1 |
| 2006 | Perfect Reconstruction Schemes for Sampling Piecewise Sinusoidal SignalsabstractConsider sampling a signal that is piecewise sinusoidal. Classical sampling theory does not enable a perfect reconstruction of the continuous time signal since the band is not limited (C.E. Shannon, 1949). However, we show that it is still possible to recover all the parameters of the sinusoids and the exact locations of the discontinuities using the annihilating filter method and recently developed Finite Rate of Innovation (FRI) sampling schemes (M. Vetterli et al., 2002) (P.L. Dragotti et al., 2005). Moreover, we show that there is a tradeoff between the number of sinusoids per piece and the proximity of the discontinuities in order to have a unique solution. This result recalls a sort of uncertainty principle Jesse Berent, Pier Luigi Dragotti |
ICASSP (3) | 1 |
| 2006 | Segmentation of Epipolar-Plane Image Volumes with Occlusion and Disocclusion CompetitionabstractConsider a dense array of cameras uniformly distributed along a line. A solid block of 3D data can be constructed by arranging the images into a stack. This volume, also known as the epipolar-plane image volume, contains highly structured data that can be segmented for object removal, insertion and compression. In this paper, we propose a segmentation scheme that takes fully advantage of the known geometry in order to model occlusions explicitly as a result of disparity. Moreover, we include this knowledge into an energy minimization scheme based on region competition with active contours. Instead of extracting layers sequentially from front to back, each layer is made to compete with the regions it is going to occlude and the ones it is going to disocclude. This enables a virtually unsupervised segmentation Jesse Berent, Pier Luigi Dragotti |
MMSP | 1 |