Laurens van der Maaten

dblp:53/2650 · also Laurens J. P. van der Maaten · DBLP profile ↗
← Back
57ranked-venue papers
5as first author
14since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 52 · 4 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 26 · 4 since 2021Systems, architecture and hardware · 2Security and privacy · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2025 Law of the Weakest Link: Cross Capabilities of Large Language Models
abstract
The development and evaluation of Large Language Models (LLMs) have largely focused on individual capabilities. However, this overlooks the intersection of multiple abilities across different types of expertise that are often required for real-world tasks, which we term **cross capabilities**. To systematically explore this concept, we first define seven core individual capabilities and then pair them to form seven common cross capabilities, each supported by a manually constructed taxonomy. Building on these definitions, we introduce *CrossEval*, a benchmark comprising 1,400 human-annotated prompts, with 100 prompts for each individual and cross capability. To ensure reliable evaluation, we involve expert annotators to assess 4,200 model responses, gathering 8,400 human ratings with detailed explanations to serve as reference examples. Our findings reveal that current LLMs consistently exhibit the ``Law of the Weakest Link,'' where cross-capability performance is significantly constrained by the weakest component. Across 58 cross-capability scores from 17 models, 38 scores are lower than all individual capabilities, while 20 fall between strong and weak, but closer to the weaker ability. These results highlight LLMs' underperformance in cross-capability tasks, emphasizing the need to identify and improve their weakest capabilities as a key research priority. The code, benchmarks, and evaluations are available on our [project website](https://www.llm-cross-capabilities.org).
Ming Zhong 0005, Aston Zhang, Wenhan Xiong, Chenguang Zhu 0001, Zhengxing Chen, Chloe Bi, Mike Lewis, Sravya Popuri, Sharan Narang, Melanie Kambadur, Dhruv Mahajan 0001, Sergey Edunov, Jiawei Han 0001, Laurens van der Maaten
ICLR17
2023 GeoDE: a Geographically Diverse Evaluation Dataset for Object Recognition
abstract
Current dataset collection methods typically scrape large amounts of data from the web. While this technique is extremely scalable, data collected in this way tends to reinforce stereotypical biases, can contain personally identifiable information, and typically originates from Europe and North America. In this work, we rethink the dataset collection paradigm and introduce GeoDE, a geographically diverse dataset with 61,940 images from 40 classes and 6 world regions, and no personally identifiable information, collected by soliciting images from people across the world. We analyse GeoDE to understand differences in images collected in this manner compared to web-scraping. Despite the smaller size of this dataset, we demonstrate its use as both an evaluation and training dataset, allowing us to highlight shortcomings in current models, as well as demonstrate improved performance even when training on this small dataset. We release the full dataset and code at https://geodiverse-data-collection.cs.princeton.edu/
Vikram V. Ramaswamy, Sing Yu Lin, Dora Zhao, Aaron Adcock, Laurens van der Maaten, Deepti Ghadiyaram, Olga Russakovsky
NeurIPS5
2022 Data Appraisal Without Data Sharing
abstract
One of the most effective approaches to improving the performance of a machine learning model is to procure additional training data. A model owner seeking relevant training data from a data owner needs to appraise the data before acquiring it. However, without a formal agreement, the data owner does not want to share data. The resulting Catch-22 prevents efficient data markets from forming. This paper proposes adding a data appraisal stage that requires no data sharing between data owners and model owners. Specifically, we use multi-party computation to implement an appraisal function computed on private data. The appraised value serves as a guide to facilitate data selection and transaction. We propose an efficient data appraisal method based on forward influence functions that approximates data value through its first-order loss reduction on the current model. The method requires no additional hyper-parameters or re-training. We show that in private, forward influence functions provide an appealing trade-off between high quality appraisal and required computation, in spite of label noise, class imbalance, and missing data. Our work seeks to inspire an open market that incentivizes efficient, equitable exchange of domain-specific training data.
Xinlei Xu, Awni Y. Hannun, Laurens van der Maaten
AISTATS3
2022 Scaling up Instance Segmentation using Approximately Localized Phrases
Karan Desai, Ishan Misra, Justin Johnson 0001, Laurens van der Maaten
BMVC4
2022 EIFFeL: Ensuring Integrity for Federated Learning
abstract
Federated learning (FL) enables clients to collaborate with a server to train a machine learning model. To ensure privacy, the server performs secure aggregation of updates from the clients. Unfortunately, this prevents verification of the well-formedness (integrity) of the updates as the updates are masked. Consequently, malformed updates designed to poison the model can be injected without detection. In this paper, we formalize the problem of ensuring both update privacy and integrity in FL and present a new system, EIFFeL, that enables secure aggregation of verified updates. EIFFeL is a general framework that can enforce arbitrary integrity checks and remove malformed updates from the aggregate, without violating privacy. Our empirical evaluation demonstrates the practicality of EIFFeL. For instance, with 100 clients and 10% poisoning, EIFFeL can train an MNIST classification model to the same accuracy as that of a non-poisoned federated learner in just 2.4s per iteration.
Amrita Roy Chowdhury 0001, Chuan Guo 0001, Somesh Jha, Laurens van der Maaten
CCS4
2022 Omnivore: A Single Model for Many Visual Modalities
abstract
Prior work has studied different visual modalities in isolation and developed separate architectures for recognition of images, videos, and 3D data. Instead, in this paper, we propose a single model which excels at classifying images, videos, and single-view 3D data using exactly the same model parameters. Our ‘OMNIVORE’ model leverages the flexibility of transformer-based architectures and is trained jointly on classification tasks from different modalities. Omnivoreis simple to train, uses off-the-shelf standard datasets, and performs at-par or better than modality-specific models of the same size. A single Omnivoremodel obtains 86.0% on ImageNet, 84.1% on Kinetics, and 67.1% on SUN RGB-D. After finetuning, our models outperform prior work on a variety of vision tasks and generalize across modalities. OMNIVORE's shared visual representation naturally enables cross-modal recognition without access to correspondences between modalities. We hope our results motivate researchers to model visual modalities together.
Rohit Girdhar, Mannat Singh, Nikhila Ravi, Laurens van der Maaten, Armand Joulin, Ishan Misra
CVPR4
2022 Revisiting Weakly Supervised Pre-Training of Visual Perception Models
abstract
Model pre-training is a cornerstone of modern visual recognition systems. Although fully supervised pre-training on datasets like ImageNet is still the de-facto standard, recent studies suggest that large-scale weakly supervised pretraining can outperform fully supervised approaches. This paper revisits weakly-supervised pre-training of models using hashtag supervision with modern versions of residual networks and the largest-ever dataset of images and corresponding hashtags. We study the performance of the resulting models in various transfer-learning settings including zero-shot transfer. We also compare our models with those obtained via large-scale self-supervised learning. We find our weakly-supervised models to be very competitive across all settings, and find they substantially outperform their self-supervised counterparts. We also include an investigation into whether our models learned potentially troubling associations or stereotypes. Overall, our results provide a compelling argument for the use of weakly supervised learning in the development of visual recognition systems. Our models, Supervised Weakly through hashtAGs (SWAG), are available publicly.
Mannat Singh, Laura Gustafson, Aaron Adcock, Vinicius de Freitas Reis, Bugra Gedik, Raj Prateek Kosaraju, Dhruv Mahajan 0001, Ross B. Girshick, Piotr Dollár, Laurens van der Maaten
CVPR10
2022 Bounding Training Data Reconstruction in Private (Deep) Learning
abstract
Differential privacy is widely accepted as the de facto method for preventing data leakage in ML, and conventional wisdom suggests that it offers strong protection against privacy attacks. However, existing semantic guarantees for DP focus on membership inference, which may overestimate the adversary’s capabilities and is not applicable when membership status itself is non-sensitive. In this paper, we derive the first semantic guarantees for DP mechanisms against training data reconstruction attacks under a formal threat model. We show that two distinct privacy accounting methods—Renyi differential privacy and Fisher information leakage—both offer strong semantic protection against data reconstruction attacks.
Chuan Guo 0001, Brian Karrer, Kamalika Chaudhuri, Laurens van der Maaten
ICML4
2022 Measuring Data Leakage in Machine-Learning Models with Fisher Information (Extended Abstract)
abstract
Machine-learning models contain information about the data they were trained on. This information leaks either through the model itself or through predictions made by the model. Consequently, when the training data contains sensitive attributes, assessing the amount of information leakage is paramount. We propose a method to quantify this leakage using the Fisher information of the model about the data. Unlike the worst-case a priori guarantees of differential privacy, Fisher information loss measures leakage with respect to specific examples, attributes, or sub-populations within the dataset. We motivate Fisher information loss through the Cramer-Rao bound and delineate the implied threat model. We provide efficient methods to compute Fisher information loss for output-perturbed generalized linear models. Finally, we empirically validate Fisher information loss as a useful measure of information leakage.
Awni Y. Hannun, Chuan Guo 0001, Laurens van der Maaten
IJCAI3
2022 Convolutional Networks with Dense Connectivity
abstract
Recent work has shown that convolutional networks can be substantially deeper, more accurate, and efficient to train if they contain shorter connections between layers close to the input and those close to the output. In this paper, we embrace this observation and introduce the Dense Convolutional Network (DenseNet), which connects each layer to every other layer in a feed-forward fashion. Whereas traditional convolutional networks with L layers have L connections-one between each layer and its subsequent layer-our network has [Formula: see text] direct connections. For each layer, the feature-maps of all preceding layers are used as inputs, and its own feature-maps are used as inputs into all subsequent layers. DenseNets have several compelling advantages: they alleviate the vanishing-gradient problem, encourage feature reuse and substantially improve parameter efficiency. We evaluate our proposed architecture on four highly competitive object recognition benchmark tasks (CIFAR-10, CIFAR-100, SVHN, and ImageNet). DenseNets obtain significant improvements over the state-of-the-art on most of them, whilst requiring less parameters and computation to achieve high performance.
Gao Huang 0001, Zhuang Liu 0003, Geoff Pleiss, Laurens van der Maaten, Kilian Q. Weinberger
IEEE Trans. Pattern Anal. Mach. Intell.4
2021 Making Paper Reviewing Robust to Bid Manipulation Attacks
abstract
Most computer science conferences rely on paper bidding to assign reviewers to papers. Although paper bidding enables high-quality assignments in days of unprecedented submission numbers, it also opens the door for dishonest reviewers to adversarially influence paper reviewing assignments. Anecdotal evidence suggests that some reviewers bid on papers by "friends" or colluding authors, even though these papers are outside their area of expertise, and recommend them for acceptance without considering the merit of the work. In this paper, we study the efficacy of such bid manipulation attacks and find that, indeed, they can jeopardize the integrity of the review process. We develop a novel approach for paper bidding and assignment that is much more robust against such attacks. We show empirically that our approach provides robustness even when dishonest reviewers collude, have full knowledge of the assignment system’s internal workings, and have access to the system’s inputs. In addition to being more robust, the quality of our paper review assignments is comparable to that of current, non-robust assignment approaches.
Ruihan Wu, Chuan Guo 0001, Felix Wu, Rahul Kidambi, Laurens van der Maaten, Kilian Q. Weinberger
ICML5
2021 CrypTen: Secure Multi-Party Computation Meets Machine Learning
abstract
Secure multi-party computation (MPC) allows parties to perform computations on data while keeping that data private. This capability has great potential for machine-learning applications: it facilitates training of machine-learning models on private data sets owned by different parties, evaluation of one party's private model using another party's private data, etc. Although a range of studies implement machine-learning models via secure MPC, such implementations are not yet mainstream. Adoption of secure MPC is hampered by the absence of flexible software frameworks that `"speak the language" of machine-learning researchers and engineers. To foster adoption of secure MPC in machine learning, we present CrypTen: a software framework that exposes popular secure MPC primitives via abstractions that are common in modern machine-learning frameworks, such as tensor computations, automatic differentiation, and modular neural networks. This paper describes the design of CrypTen and measure its performance on state-of-the-art models for text classification, speech recognition, and image classification. Our benchmarks show that CrypTen's GPU support and high-performance communication between (an arbitrary number of) parties allows it to perform efficient private evaluation of modern machine-learning models under a semi-honest threat model. For example, two parties using CrypTen can securely predict phonemes in speech recordings using Wav2Letter faster than real-time. We hope that CrypTen will spur adoption of secure MPC in the machine-learning community.
Brian Knott, Shobha Venkataraman, Awni Y. Hannun, Shubho Sengupta, Mark Ibrahim, Laurens van der Maaten
NeurIPS6
2021 Fixes That Fail: Self-Defeating Improvements in Machine-Learning Systems
abstract
Machine-learning systems such as self-driving cars or virtual assistants are composed of a large number of machine-learning models that recognize image content, transcribe speech, analyze natural language, infer preferences, rank options, etc. Models in these systems are often developed and trained independently, which raises an obvious concern: Can improving a machine-learning model make the overall system worse? We answer this question affirmatively by showing that improving a model can deteriorate the performance of downstream models, even after those downstream models are retrained. Such self-defeating improvements are the result of entanglement between the models in the system. We perform an error decomposition of systems with multiple machine-learning models, which sheds light on the types of errors that can lead to self-defeating improvements. We also present the results of experiments which show that self-defeating improvements emerge in a realistic stereo-based detection system for cars and pedestrians.
Ruihan Wu, Chuan Guo 0001, Awni Y. Hannun, Laurens van der Maaten
NeurIPS4
2021 Measuring data leakage in machine-learning models with Fisher information
abstract
Machine-learning models contain information about the data they were trained on. This information leaks either through the model itself or through predictions made by the model. Consequently, when the training data contains sensitive attributes, assessing the amount of information leakage is paramount. We propose a method to quantify this leakage using the Fisher information of the model about the data. Unlike the worst-case a priori guarantees of differential privacy, Fisher information loss measures leakage with respect to specific examples, attributes, or sub-populations within the dataset. We motivate Fisher information loss through the Cram\’{e}r-Rao bound and delineate the implied threat model. We provide efficient methods to compute Fisher information loss for output-perturbed generalized linear models. Finally, we empirically validate Fisher information loss as a useful measure of information leakage.
Awni Y. Hannun, Chuan Guo 0001, Laurens van der Maaten
UAI3
2020 Self-Supervised Learning of Pretext-Invariant Representations
abstract
The goal of self-supervised learning from images is to construct image representations that are semantically meaningful via pretext tasks that do not require semantic annotations. Many pretext tasks lead to representations that are covariant with image transformations. We argue that, instead, semantic representations ought to be invariant under such transformations. Specifically, we develop Pretext-Invariant Representation Learning (PIRL, pronounced as `pearl') that learns invariant representations based on pretext tasks. We use PIRL with a commonly used pretext task that involves solving jigsaw puzzles. We find that PIRL substantially improves the semantic quality of the learned image representations. Our approach sets a new state-of-the-art in self-supervised learning from images on several popular benchmarks for self-supervised learning. Despite being unsupervised, PIRL outperforms supervised pre-training in learning image representations for object detection. Altogether, our results demonstrate the potential of self-supervised representations with good invariance properties.
Ishan Misra, Laurens van der Maaten
CVPR2
2020 Certified Data Removal from Machine Learning Models
abstract
Good data stewardship requires removal of data at the request of the data’s owner. This raises the question if and how a trained machine-learning model, which implicitly stores information about its training data, should be affected by such a removal request. Is it possible to “remove” data from a machine-learning model? We study this problem by defining certified removal: a very strong theoretical guarantee that a model from which data is removed cannot be distinguished from a model that never observed the data to begin with. We develop a certified-removal mechanism for linear classifiers and empirically study learning settings in which this mechanism is practical.
Chuan Guo 0001, Tom Goldstein, Awni Y. Hannun, Laurens van der Maaten
ICML4
2019 Defense Against Adversarial Images Using Web-Scale Nearest-Neighbor Search
abstract
A plethora of recent work has shown that convolutional networks are not robust to adversarial images: images that are created by perturbing a sample from the data distribution as to maximize the loss on the perturbed example. In this work, we hypothesize that adversarial perturbations move the image away from the image manifold in the sense that there exists no physical process that could have produced the adversarial image. This hypothesis suggests that a successful defense mechanism against adversarial images should aim to project the images back onto the image manifold. We study such defense mechanisms, which approximate the projection onto the unknown image manifold by a nearest-neighbor search against a web-scale image database containing tens of billions of images. Empirical evaluations of this defense strategy on ImageNet suggest that it very effective in attack settings in which the adversary does not have access to the image database. We also propose two novel attack methods to break nearest-neighbor defense settings and show conditions under which nearest-neighbor defense fails. We perform a series of ablation experiments, which suggest that there is a trade-off between robustness and accuracy between as we use features from deeper in the network, that a large index size (hundreds of millions) is crucial to get good performance, and that careful construction of database is crucial for robustness against nearest-neighbor attacks.
Abhimanyu Dubey, Laurens van der Maaten, Ismet Zeki Yalniz, Yixuan Li 0001, Dhruv Mahajan 0001
CVPR2
2019 Feature Denoising for Improving Adversarial Robustness
abstract
Adversarial attacks to image classification systems present challenges to convolutional networks and opportunities for understanding them. This study suggests that adversarial perturbations on images lead to noise in the features constructed by these networks. Motivated by this observation, we develop new network architectures that increase adversarial robustness by performing feature denoising. Specifically, our networks contain blocks that denoise the features using non-local means or other filters; the entire networks are trained end-to-end. When combined with adversarial training, our feature denoising networks substantially improve the state-of-the-art in adversarial robustness in both white-box and black-box attack settings. On ImageNet, under 10-iteration PGD white-box attacks where prior art has 27.9% accuracy, our method achieves 55.7%; even under extreme 2000-iteration PGD white-box attacks, our method secures 42.6% accuracy. Our method was ranked first in Competition on Adversarial Attacks and Defenses (CAAD) 2018 --- it achieved 50.6% classification accuracy on a secret, ImageNet-like test dataset against 48 unknown attackers, surpassing the runner-up approach by ~10%. Code is available at https://github.com/facebookresearch/ImageNet-Adversarial-Training.
Cihang Xie, Yuxin Wu 0004, Laurens van der Maaten, Alan L. Yuille, Kaiming He
CVPR3
2019 Anytime Stereo Image Depth Estimation on Mobile Devices
abstract
Many applications of stereo depth estimation in robotics require the generation of accurate disparity maps in real time under significant computational constraints. Current state-of-the-art algorithms force a choice between either generating accurate mappings at a slow pace, or quickly generating inaccurate ones, and additionally these methods typically require far too many parameters to be usable on power- or memory-constrained devices. Motivated by these shortcomings, we propose a novel approach for disparity prediction in the anytime setting. In contrast to prior work, our end-to-end learned approach can trade off computation and accuracy at inference time. Depth estimation is performed in stages, during which the model can be queried at any time to output its current best estimate. Our final model can process 1242×375 resolution images within a range of 10-35 FPS on an NVIDIA Jetson TX2 module with only marginal increases in error - using two orders of magnitude fewer parameters than the most competitive baseline. The source code is available at https://github.com/mileyan/AnyNet.
Yan Wang 0051, Zihang Lai, Gao Huang 0001, Brian H. Wang, Laurens van der Maaten, Mark E. Campbell, Kilian Q. Weinberger
ICRA5
2019 PHYRE: A New Benchmark for Physical Reasoning
abstract
Understanding and reasoning about physics is an important ability of intelligent agents. We develop the PHYRE benchmark for physical reasoning that contains a set of simple classical mechanics puzzles in a 2D physical environment. The benchmark is designed to encourage the development of learning algorithms that are sample-efficient and generalize well across puzzles. We test several modern learning algorithms on PHYRE and find that these algorithms fall short in solving the puzzles efficiently. We expect that PHYRE will encourage the development of novel sample-efficient agents that learn efficient but useful models of physics. For code and to play PHYRE for yourself, please visit https://player.phyre.ai.
Anton Bakhtin, Laurens van der Maaten, Justin Johnson 0001, Laura Gustafson, Ross B. Girshick
NeurIPS2
2018 3D Semantic Segmentation With Submanifold Sparse Convolutional Networks
abstract
Convolutional networks are the de-facto standard for analyzing spatio-temporal data such as images, videos, and 3D shapes. Whilst some of this data is naturally dense (e.g., photos), many other data sources are inherently sparse. Examples include 3D point clouds that were obtained using a LiDAR scanner or RGB-D camera. Standard "dense" implementations of convolutional networks are very inefficient when applied on such sparse data. We introduce new sparse convolutional operations that are designed to process spatially-sparse data more efficiently, and use them to develop spatially-sparse convolutional networks. We demonstrate the strong performance of the resulting models, called submanifold sparse convolutional networks (SS-CNs), on two tasks involving semantic segmentation of 3D point clouds. In particular, our models outperform all prior state-of-the-art on the test set of a recent semantic segmentation competition.
Benjamin Graham, Martin Engelcke, Laurens van der Maaten
CVPR3
2018 CondenseNet: An Efficient DenseNet Using Learned Group Convolutions
abstract
Deep neural networks are increasingly used on mobile devices, where computational resources are limited. In this paper we develop CondenseNet, a novel network architecture with unprecedented efficiency. It combines dense connectivity with a novel module called learned group convolution. The dense connectivity facilitates feature re-use in the network, whereas learned group convolutions remove connections between layers for which this feature re-use is superfluous. At test time, our model can be implemented using standard group convolutions, allowing for efficient computation in practice. Our experiments show that CondenseNets are far more efficient than state-of-the-art compact convolutional networks such as ShuffleNets.
Gao Huang 0001, Shichen Liu, Laurens van der Maaten, Kilian Q. Weinberger
CVPR3
2018 Learning by Asking Questions
abstract
We introduce an interactive learning framework for the development and testing of intelligent visual systems, called learning-by-asking (LBA). We explore LBA in context of the Visual Question Answering (VQA) task. LBA differs from standard VQA training in that most questions are not observed during training time, and the learner must ask questions it wants answers to. Thus, LBA more closely mimics natural learning and has the potential to be more data-efficient than the traditional VQA setting. We present a model that performs LBA on the CLEVR dataset, and show that it automatically discovers an easy-to-hard curriculum when learning interactively from an oracle. Our LBA generated data consistently matches or outperforms the CLEVR train data and is more sample efficient. We also show that our model asks questions that generalize to state-of-the-art VQA models and to novel test time distributions.
Ishan Misra, Ross B. Girshick, Rob Fergus, Martial Hebert, Abhinav Gupta 0001, Laurens van der Maaten
CVPR6
2018 Separating Self-Expression and Visual Content in Hashtag Supervision
abstract
The variety, abundance, and structured nature of hashtags make them an interesting data source for training vision models. For instance, hashtags have the potential to significantly reduce the problem of manual supervision and annotation when learning vision models for a large number of concepts. However, a key challenge when learning from hashtags is that they are inherently subjective because they are provided by users as a form of self-expression. As a consequence, hashtags may have synonyms (different hashtags referring to the same visual content) and may be polysemous (the same hashtag referring to different visual content). These challenges limit the effectiveness of approaches that simply treat hashtags as image-label pairs. This paper presents an approach that extends upon modeling simple image-label pairs with a joint model of images, hashtags, and users. We demonstrate the efficacy of such approaches in image tagging and retrieval experiments, and show how the joint model can be used to perform user-conditional retrieval and tagging.
Andreas Veit, Maximilian Nickel, Serge J. Belongie, Laurens van der Maaten
CVPR4
2018 Exploring the Limits of Weakly Supervised Pretraining
Dhruv Mahajan 0001, Ross B. Girshick, Vignesh Ramanathan, Kaiming He, Manohar Paluri, Yixuan Li 0001, Ashwin Bharambe, Laurens van der Maaten
ECCV (2)8
2018 Countering Adversarial Images using Input Transformations
Chuan Guo 0001, Mayank Rana, Moustapha Cissé, Laurens van der Maaten
ICLR (Poster)4
2018 Multi-Scale Dense Networks for Resource Efficient Image Classification
Gao Huang 0001, Danlu Chen, Tianhong Li, Felix Wu, Laurens van der Maaten, Kilian Q. Weinberger
ICLR5
2018 Multivariate Time-Series Classification Using the Hidden-Unit Logistic Model
abstract
We present a new model for multivariate time-series classification, called the hidden-unit logistic model (HULM), that uses binary stochastic hidden units to model latent structure in the data. The hidden units are connected in a chain structure that models temporal dependencies in the data. Compared with the prior models for time-series classification such as the hidden conditional random field, our model can model very complex decision boundaries, because the number of latent states grows exponentially with the number of hidden units. We demonstrate the strong performance of our model in experiments on a variety of (computer vision) tasks, including handwritten character recognition, speech recognition, facial expression, and action recognition. We also present a state-of-the-art system for facial action unit detection based on the HULM.
Wenjie Pei, Hamdi Dibeklioglu, David M. J. Tax, Laurens van der Maaten
IEEE Trans. Neural Networks Learn. Syst.4
2017 Densely Connected Convolutional Networks
abstract
Recent work has shown that convolutional networks can be substantially deeper, more accurate, and efficient to train if they contain shorter connections between layers close to the input and those close to the output. In this paper, we embrace this observation and introduce the Dense Convolutional Network (DenseNet), which connects each layer to every other layer in a feed-forward fashion. Whereas traditional convolutional networks with L layers have L connections-one between each layer and its subsequent layer-our network has L(L+1)/2 direct connections. For each layer, the feature-maps of all preceding layers are used as inputs, and its own feature-maps are used as inputs into all subsequent layers. DenseNets have several compelling advantages: they alleviate the vanishing-gradient problem, strengthen feature propagation, encourage feature reuse, and substantially reduce the number of parameters. We evaluate our proposed architecture on four highly competitive object recognition benchmark tasks (CIFAR-10, CIFAR-100, SVHN, and ImageNet). DenseNets obtain significant improvements over the state-of-the-art on most of them, whilst requiring less memory and computation to achieve high performance. Code and pre-trained models are available at https://github.com/liuzhuang13/DenseNet.
Gao Huang 0001, Zhuang Liu 0003, Laurens van der Maaten, Kilian Q. Weinberger
CVPR3
2017 CLEVR: A Diagnostic Dataset for Compositional Language and Elementary Visual Reasoning
abstract
When building artificial intelligence systems that can reason and answer questions about visual data, we need diagnostic tests to analyze our progress and discover short-comings. Existing benchmarks for visual question answering can help, but have strong biases that models can exploit to correctly answer questions without reasoning. They also conflate multiple sources of error, making it hard to pinpoint model weaknesses. We present a diagnostic dataset that tests a range of visual reasoning abilities. It contains minimal biases and has detailed annotations describing the kind of reasoning each question requires. We use this dataset to analyze a variety of modern visual reasoning systems, providing novel insights into their abilities and limitations.
Justin Johnson 0001, Bharath Hariharan, Laurens van der Maaten, Li Fei-Fei 0001, C. Lawrence Zitnick, Ross B. Girshick
CVPR3
2017 Inferring and Executing Programs for Visual Reasoning
abstract
Existing methods for visual reasoning attempt to directly map inputs to outputs using black-box architectures without explicitly modeling the underlying reasoning processes. As a result, these black-box models often learn to exploit biases in the data rather than learning to perform visual reasoning. Inspired by module networks, this paper proposes a model for visual reasoning that consists of a program generator that constructs an explicit representation of the reasoning process to be performed, and an execution engine that executes the resulting program to produce an answer. Both the program generator and the execution engine are implemented by neural networks, and are trained using a combination of backpropagation and REINFORCE. Using the CLEVR benchmark for visual reasoning, we show that our model significantly outperforms strong baselines and generalizes better in a variety of settings.
Justin Johnson 0001, Bharath Hariharan, Laurens van der Maaten, Judy Hoffman, Li Fei-Fei 0001, C. Lawrence Zitnick, Ross B. Girshick
ICCV3
2017 Learning Visual N-Grams from Web Data
Allan Jabri, Armand Joulin, Laurens van der Maaten
ICCV4
2017 Approximated and User Steerable tSNE for Progressive Visual Analytics
abstract
Progressive Visual Analytics aims at improving the interactivity in existing analytics techniques by means of visualization as well as interaction with intermediate results. One key method for data analysis is dimensionality reduction, for example, to produce 2D embeddings that can be visualized and analyzed efficiently. t-Distributed Stochastic Neighbor Embedding (tSNE) is a well-suited technique for the visualization of high-dimensional data. tSNE can create meaningful intermediate results but suffers from a slow initialization that constrains its application in Progressive Visual Analytics. We introduce a controllable tSNE approximation (A-tSNE), which trades off speed and accuracy, to enable interactive data exploration. We offer real-time visualization techniques, including a density-based solution and a Magic Lens to inspect the degree of approximation. With this feedback, the user can decide on local refinements and steer the approximation level during the analysis. We demonstrate our technique with several datasets, in a real-world research scenario and for the real-time analysis of high-dimensional streams to illustrate its effectiveness for interactive data analysis.
Nicola Pezzotti, Boudewijn P. F. Lelieveldt, Laurens van der Maaten, Thomas Höllt, Elmar Eisemann, Anna Vilanova
IEEE Trans. Vis. Comput. Graph.3
2016 Revisiting Visual Question Answering Baselines
Allan Jabri, Armand Joulin, Laurens van der Maaten
ECCV (8)3
2016 Learning Visual Features from Large Weakly Supervised Data
Armand Joulin, Laurens van der Maaten, Allan Jabri, Nicolas Vasilache
ECCV (7)2
2016 Feature-Level Domain Adaptation
abstract
Domain adaptation is the supervised learning setting in which the training and test data are sampled from different distributions: training data is sampled from a source domain, whilst test data is sampled from a target domain. This paper proposes and studies an approach, called feature-level domain adaptation (FLDA), that models the dependence between the two domains by means of a feature-level transfer model that is trained to describe the transfer from source to target domain. Subsequently, we train a domain-adapted classifier by minimizing the expected loss under the resulting transfer model. For linear classifiers and a large family of loss functions and transfer models, this expected loss can be computed or approximated analytically, and minimized efficiently. Our empirical evaluation of FLDA focuses on problems comprising binary and count data in which the transfer can be naturally modeled via a dropout distribution, which allows the classifier to adapt to differences in the marginal probability of features in the source and the target domain. Our experiments on several real- world problems show that FLDA performs on par with state- of- the-art domain-adaptation techniques.
Wouter M. Kouw, Laurens van der Maaten, Jesse H. Krijthe, Marco Loog
J. Mach. Learn. Res.2
2015 Adaptive stereo similarity fusion using confidence measures
Gorkem Saygili, Laurens van der Maaten, Emile A. Hendriks
Comput. Vis. Image Underst.2
2015 Introduction to the special issue on visual analytics using multidimensional projections
Michaël Aupetit 0001, Laurens van der Maaten
Neurocomputing2
2014 Speeding Up Tracking by Ignoring Features
abstract
Most modern object trackers combine a motion prior with sliding-window detection, using binary classifiers that predict the presence of the target object based on histogram features. Although the accuracy of such trackers is generally very good, they are often impractical because of their high computational requirements. To resolve this problem, the paper presents a new approach that limits the computational costs of trackers by ignoring features in image regions that -- after inspecting a few features -- are unlikely to contain the target object. To this end, we derive an upper bound on the probability that a location is most likely to contain the target object, and we ignore (features in) locations for which this upper bound is small. We demonstrate the effectiveness of our new approach in experiments with model-free and model-based trackers that use linear models in combination with HOG features. The results of our experiments demonstrate that our approach allows us to reduce the average number of inspected features by up to 90% without affecting the accuracy of the tracker.
Lu Zhang 0036, Hamdi Dibeklioglu, Laurens van der Maaten
CVPR3
2014 Graph-Based Kinship Recognition
abstract
Image-based kinship recognition is an important problem in the reconstruction and analysis of social networks. Prior studies on image-based kinship recognition have focused solely on pair wise kinship verification, i.e. on the question of whether or not two people are kin. Such approaches fail to exploit the fact that many real-world photographs contain several family members, for instance, the probability of two people being brothers increases when both people are recognized to have the same father. In this work, we propose a graph-based approach that incorporates facial similarities between all family members in a photograph in order to improve the performance of kinship recognition. In addition, we introduce a database of group photographs with kinship annotations.
Yuanhao Guo, Hamdi Dibeklioglu, Laurens van der Maaten
ICPR3
2014 Stereo Similarity Metric Fusion Using Stereo Confidence
abstract
Stereo confidence measures are one of the most popular research topics in stereo vision. These measures give an indication about the certainty of the matching. The main aim of using confidence measures is to filter the erroneous disparity estimations at the end of the matching process. However, they can also be incorporated at the initial step of the matching process to obtain accurate estimations before the cost aggregation. In this paper, we propose to utilize stereo confidence measures for fusing different similarity measures in order to obtain robust estimations for aggregation. Since stereo similarity measures perform differently in varying conditions, the confidence-guided fusion of them makes stereo matching more robust against errors. We evaluate the performance of our algorithm in comparison to different similarity measures on the Middleburry benchmark stereo test set. The results show significant improvements on the accuracy of initial disparity estimations with our fusion strategy compared to different similarity measures.
Gorkem Saygili, Laurens van der Maaten, Emile A. Hendriks
ICPR2
2014 Hybrid Kinect Depth Map Refinement for Transparent Objects
abstract
Depth sensors such as Kinect fail to find the depth of transparent objects which makes 3D reconstruction of such objects a challenge. The refinement algorithms for Kinect depth maps either do not address transparency or they only provide sparse depth on such objects which is inadequate for dense 3D reconstruction. In order to solve this problem, we propose a fully-connected CRF based hybrid refinement algorithm. We incorporate stereo cues from cross-modal stereo between IR and RGB cameras of the Kinect and Kinect's depth map. Our algorithm does not require any additional cameras and still provides dense depth estimations of transparent objects and specular surfaces with high accuracy.
Gorkem Saygili, Laurens van der Maaten, Emile A. Hendriks
ICPR2
2014 Improving Object Tracking by Adapting Detectors
abstract
The goal of model-based object trackers is to automatically detect and track specific objects, such as cars or pedestrians. To solve this problem, many modern trackers train a detector on a collection of annotated object images and use the trained detector in a tracking-by-detection framework. A major limitation of such an approach is that a single, generic detector is used to track specific objects, the additional information on the visual appearance of the particular object under consideration that is available after the initial detection is ignored. This paper proposes an approach that addresses this limitation by adapting the appearance model for each particular object using online learning techniques. We demonstrate the effectiveness of the approach in a state-of-the-art object detector based on deformable template models, the parameters of which are adapted online using an online structured SVM. We further improve the performance of the resulting model-based trackers by online learning a prior distribution over the size of objects. The experimental evaluation of our tracker demonstrates its effectiveness in pedestrian tracking.
Lu Zhang 0036, Laurens van der Maaten
ICPR2
2014 Accelerating t-SNE using tree-based algorithms
Laurens van der Maaten
J. Mach. Learn. Res.1
2014 Preserving Structure in Model-Free Tracking
abstract
Model-free trackers can track arbitrary objects based on a single (bounding-box) annotation of the object. Whilst the performance of model-free trackers has recently improved significantly, simultaneously tracking multiple objects with similar appearance remains very hard. In this paper, we propose a new multi-object model-free tracker (using a tracking-by-detection framework) that resolves this problem by incorporating spatial constraints between the objects. The spatial constraints are learned along with the object detectors using an online structured SVM algorithm. The experimental evaluation of our structure-preserving object tracker (SPOT) reveals substantial performance improvements in multi-object tracking. We also show that SPOT can improve the performance of single-object trackers by simultaneously tracking different parts of the object. Moreover, we show that SPOT can be used to adapt generic, model-based object detectors during tracking to tailor them towards a specific instance of that object.
Lu Zhang 0036, Laurens van der Maaten
IEEE Trans. Pattern Anal. Mach. Intell.2
2013 Structure Preserving Object Tracking
abstract
Model-free trackers can track arbitrary objects based on a single (bounding-box) annotation of the object. Whilst the performance of model-free trackers has recently improved significantly, simultaneously tracking multiple objects with similar appearance remains very hard. In this paper, we propose a new multi-object model-free tracker (based on tracking-by-detection) that resolves this problem by incorporating spatial constraints between the objects. The spatial constraints are learned along with the object detectors using an online structured SVM algorithm. The experimental evaluation of our structure-preserving object tracker (SPOT) reveals significant performance improvements in multi-object tracking. We also show that SPOT can improve the performance of single-object trackers by simultaneously tracking different parts of the object.
Lu Zhang 0036, Laurens van der Maaten
CVPR2
2013 Learning with Marginalized Corrupted Features
abstract
The goal of machine learning is to develop predictors that generalize well to test data. Ideally, this is achieved by training on very large (infinite) training data sets that capture all variations in the data distribution. In the case of finite training data, an effective solution is to extend the training set with artificially created examples – which, however, is also computationally costly. We propose to corrupt training examples with noise from known distributions within the exponential family and present a novel learning algorithm, called marginalized corrupted features (MCF), that trains robust predictors by minimizing the expected value of the loss function under the corrupting distribution – essentially learning with infinitely many (corrupted) training examples. We show empirically on a variety of data sets that MCF classifiers can be trained efficiently, may generalize substantially better to test data, and are more robust to feature deletion at test time.
Laurens van der Maaten, Minmin Chen, Stephen Tyree, Kilian Q. Weinberger
ICML (1)1
2013 Divvy: fast and intuitive exploratory data analysis
Joshua M. Lewis, Virginia R. de Sa, Laurens van der Maaten
J. Mach. Learn. Res.3
2012 A Behavioral Investigation of Dimensionality Reduction
Joshua M. Lewis, Laurens van der Maaten, Virginia R. de Sa
CogSci2
2012 Improving segment based stereo matching using SURF key points
abstract
State-of-the-art stereo matching algorithms estimate disparities using local block-matching, and subsequently refine the disparity estimates by introducing smoothness constraints and performing global energy minimization. Such algorithms are hampered by the inability of local block-matching algorithms to deal with repetitive patterns. This paper presents an approach that overcomes this problem by incorporating the disparity obtained from matching SURF key points between stereo image pairs. The algorithm provides further robustness to problems with repetitive pattern by penalizing the discrepancy between the initial and final disparity estimates in the global energy minimization. Evaluation of our approach on the Middleburry data set results shows that the our approach is more robust against repetitive patterns than existing approaches.
Gorkem Saygili, Laurens van der Maaten, Emile A. Hendriks
ICIP2
2012 Audio-visual emotion challenge 2012: a simple approach
abstract
The paper presents a small empirical study into emotion and affect recognition based on auditory and visual features, which was performed in the context of the Audio-Visual Emotion Challenge (AVEC) 2012. The goal of this competition is to predict continuous-valued affect ratings based on the provided auditory and visual features, e.g., local binary pattern (LBP) features extracted from aligned face images, and spectral audio features.
Laurens van der Maaten
ICMI1
2012 Accelerating Cost Aggregation for Real-Time Stereo Matching
abstract
Real-time stereo matching, which is important in many applications like self-driving cars and 3-D scene reconstruction, requires large computation capability and high memory bandwidth. The most time-consuming part of stereo-matching algorithms is the aggregation of information (i.e. costs) over local image regions. In this paper, we present a generic representation and suitable implementations for three commonly used cost aggregators on many-core processors. We perform typical optimizations on the kernels, which leads to significant performance improvement (up to two orders of magnitude). Finally, we present a performance model for the three aggregators to predict the aggregation speed for a given pair of input images on a given architecture. Experimental results validate our model with an acceptable error margin (an average of 10.4%). We conclude that GPU-like many-cores are excellent platforms for accelerating stereo matching.
Jianbin Fang, Ana Lucia Varbanescu, Jie Shen 0003, Henk J. Sips, Gorkem Saygili, Laurens van der Maaten
ICPADS6
2012 Visualizing non-metric similarities in multiple maps
abstract
Techniques for multidimensional scaling visualize objects as points in a low-dimensional metric map. As a result, the visualizations are subject to the fundamental limitations of metric spaces. These limitations prevent multidimensional scaling from faithfully representing non-metric similarity data such as word associations or event co-occurrences. In particular, multidimensional scaling cannot faithfully represent intransitive pairwise similarities in a visualization, and it cannot faithfully visualize “central” objects. In this paper, we present an extension of a recently proposed multidimensional scaling technique called t-SNE. The extension aims to address the problems of traditional multidimensional scaling techniques when these techniques are used to visualize non-metric similarities. The new technique, called multiple maps t-SNE, alleviates these problems by constructing a collection of maps that reveal complementary structure in the similarity data. We apply multiple maps t-SNE to a large data set of word association data and to a data set of NIPS co-authorships, demonstrating its ability to successfully visualize non-metric similarities.
Laurens van der Maaten, Geoffrey E. Hinton
Mach. Learn.1
2011 Learning Discriminative Fisher Kernels
Laurens van der Maaten
ICML1
2010 Deep Supervised t-Distributed Embedding
Martin Renqiang Min, Laurens van der Maaten, Zineng Yuan, Anthony J. Bonner, Zhaolei Zhang
ICML2
2010 On Herding and the Perceptron Cycling Theorem
abstract
The paper develops a connection between traditional perceptron algorithms and recently introduced herding algorithms. It is shown that both algorithms can be viewed as an application of the perceptron cycling theorem. This connection strengthens some herding results and suggests new (supervised) herding algorithms that, like CRFs or discriminative RBMs, make predictions by conditioning on the input attributes. We develop and investigate variants of conditional herding, and show that conditional herding leads to practical algorithms that perform better than or on par with related classifiers such as the voted perceptron and the discriminative RBM.
Andrew Gelfand, Yutian Chen 0001, Laurens van der Maaten, Max Welling
NIPS3
2010 Latent Variable Models for Predicting File Dependencies in Large-Scale Software Development
abstract
When software developers modify one or more files in a large code base, they must also identify and update other related files. Many file dependencies can be detected by mining the development history of the code base: in essence, groups of related files are revealed by the logs of previous workflows. From data of this form, we show how to detect dependent files by solving a problem in binary matrix completion. We explore different latent variable models (LVMs) for this problem, including Bernoulli mixture models, exponential family PCA, restricted Boltzmann machines, and fully Bayesian approaches. We evaluate these models on the development histories of three large, open-source software systems: Mozilla Firefox, Eclipse Subversive, and Gimp. In all of these applications, we find that LVMs improve the performance of related file prediction over current leading methods.
Diane Hu, Laurens van der Maaten, Youngmin Cho, Lawrence K. Saul, Sorin Lerner
NIPS2