EDBT 2026 Demo / reviewers in the wild / expert
Ian Fischer
dblp:17/5600
· DBLP profile ↗
17ranked-venue papers
0as first author
6since 2021 · last 2024
0000-0003-3886-5619ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 16 · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2Security and privacy · 1Databases, data management, data science and information retrieval · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
14 papers |
Reinforcement learning · 29% Representation and self-supervised learning · 19% Probabilistic and Bayesian machine learning · 11% | |
| Computer graphics and multimedia
1 paper |
Image and video coding · 100% | |
| Network and information security
2 papers |
Security and privacy of machine learning · 82% Usable security · 9% Web and mobile security · 9% |
Topics — the 30 heaviest of 41, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation › language modeling › long-context language modeling › context utilization › long-context modeling
long-context language model |
0.8 | 1 | 2024 | A Human-Inspired Reading Agent with Gist Memory of Very Long Contexts · ICML 2024 |
Natural language and speech › Question answering and dialogue systems › machine reading comprehension
long-document reading comprehension |
0.8 | 1 | 2024 | A Human-Inspired Reading Agent with Gist Memory of Very Long Contexts · ICML 2024 |
Machine learning › Probabilistic and Bayesian machine learning › structured models
latent variable model |
0.7 | 2 | 2018 | GILBO: One Metric to Measure Them All · NeurIPS 2018 Fixing a Broken ELBO · ICML 2018 |
Machine learning › Representation and self-supervised learning
contrastive learning |
0.6 | 2 | 2021 | Compressive Visual Representations · NeurIPS 2021 Predictive Information Accelerates Learning in RL · NeurIPS 2020 |
Machine learning › Reinforcement learning › offline reinforcement learning
decision transformer |
0.6 | 1 | 2022 | Multi-Game Decision Transformers · NeurIPS 2022 |
Machine learning › Reinforcement learning
generalist agents |
0.6 | 1 | 2022 | Multi-Game Decision Transformers · NeurIPS 2022 |
Machine learning › Reinforcement learning
hierarchical reinforcement learning |
0.6 | 1 | 2022 | Deep Hierarchical Planning from Pixels · NeurIPS 2022 |
Machine learning › Reinforcement learning › model-based reinforcement learning › model-based planning
latent space planning |
0.6 | 1 | 2022 | Deep Hierarchical Planning from Pixels · NeurIPS 2022 |
Machine learning › Reinforcement learning
offline reinforcement learning |
0.6 | 1 | 2022 | Multi-Game Decision Transformers · NeurIPS 2022 |
Machine learning › Trustworthy machine learning
robustness |
0.5 | 1 | 2021 | Compressive Visual Representations · NeurIPS 2021 |
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning › self-supervised representation learning
self-supervised visual representation learning |
0.5 | 1 | 2021 | Compressive Visual Representations · NeurIPS 2021 |
Machine learning › Deep learning architectures and training › loss function design
information-theoretic objective |
0.4 | 1 | 2020 | An Unsupervised Information-Theoretic Perceptual Quality Metric · NeurIPS 2020 |
Machine learning › Reinforcement learning › sample efficiency
sample-efficient reinforcement learning |
0.4 | 1 | 2020 | Predictive Information Accelerates Learning in RL · NeurIPS 2020 |
Image and video coding
image quality assessment |
0.4 | 1 | 2020 | An Unsupervised Information-Theoretic Perceptual Quality Metric · NeurIPS 2020 |
Image and video coding › image quality assessment
perceptual quality metric |
0.4 | 1 | 2020 | An Unsupervised Information-Theoretic Perceptual Quality Metric · NeurIPS 2020 |
Machine learning › Probabilistic and Bayesian machine learning › structured models › latent variable model
latent dynamics learning |
0.4 | 1 | 2019 | Learning Latent Dynamics for Planning from Pixels · ICML 2019 |
Machine learning › Reinforcement learning
model-based reinforcement learning |
0.4 | 1 | 2019 | Learning Latent Dynamics for Planning from Pixels · ICML 2019 |
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › planning
online planning |
0.4 | 1 | 2019 | Learning Latent Dynamics for Planning from Pixels · ICML 2019 |
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference › variational inference › variational objective
evidence lower bound |
0.3 | 1 | 2018 | Fixing a Broken ELBO · ICML 2018 |
Machine learning › Generative modeling
generative model evaluation |
0.3 | 1 | 2018 | GILBO: One Metric to Measure Them All · NeurIPS 2018 |
Machine learning › Representation and self-supervised learning
information-theoretic representation learning |
0.3 | 1 | 2018 | Fixing a Broken ELBO · ICML 2018 |
Machine learning › Representation and self-supervised learning › mutual information
mutual information estimation |
0.3 | 1 | 2018 | GILBO: One Metric to Measure Them All · NeurIPS 2018 |
Machine learning › Generative modeling
rate-distortion tradeoff |
0.3 | 1 | 2018 | Fixing a Broken ELBO · ICML 2018 |
Security and privacy of machine learning
adversarial attack |
0.3 | 1 | 2018 | Learning to Attack: Adversarial Transformation Networks · AAAI 2018 |
Security and privacy of machine learning
adversarial example |
0.3 | 1 | 2018 | Learning to Attack: Adversarial Transformation Networks · AAAI 2018 |
Machine learning › Representation and self-supervised learning
information bottleneck |
0.3 | 1 | 2017 | Deep Variational Information Bottleneck · ICLR (Poster) 2017 |
Computer vision › Image recognition and object detection
object detection |
0.3 | 1 | 2017 | Speed/Accuracy Trade-Offs for Modern Convolutional Object Detectors · CVPR 2017 |
Machine learning › Representation and self-supervised learning › information bottleneck
variational information bottleneck |
0.3 | 1 | 2017 | Deep Variational Information Bottleneck · ICLR (Poster) 2017 |
Performance modeling and evaluation
benchmarking |
0.3 | 1 | 2017 | Speed/Accuracy Trade-Offs for Modern Convolutional Object Detectors · CVPR 2017 |
Machine learning › Efficient and distributed learning › large-scale learning
model scaling |
0.2 | 1 | 2022 | Multi-Game Decision Transformers · NeurIPS 2022 |
Methods — techniques the papers use, named apart from their topics
variational inference · 1.3information theory · 1.2contrastive learning · 0.9conditional entropy bottleneck · 0.9retrieval · 0.8prompting · 0.8weighted ensemble · 0.7world model · 0.6low-level policy · 0.6high-level policy · 0.6unsupervised learning · 0.4optimization-based attack · 0.3gradient-based attack · 0.3feature extractor comparison · 0.3SSD · 0.3R-FCN · 0.3Faster R-CNN · 0.3user study · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | A Human-Inspired Reading Agent with Gist Memory of Very Long ContextsabstractCurrent Large Language Models (LLMs) are not only limited to some maximum context length, but also are not able to robustly consume long inputs. To address these limitations, we propose ReadAgent, an LLM agent system that increases effective context length up to 20x in our experiments. Inspired by how humans interactively read long documents, we implement ReadAgent as a simple prompting system that uses the advanced language capabilities of LLMs to (1) decide what content to store together in a memory episode, (2) compress those memory episodes into short episodic memories called *gist memories*, and (3) take actions to look up passages in the original text if ReadAgent needs to remind itself of relevant details to complete a task. We evaluate ReadAgent against baselines using retrieval methods, using the original long contexts, and using the gist memories. These evaluations are performed on three long-document reading comprehension tasks: QuALITY, NarrativeQA, and QMSum. ReadAgent outperforms the baselines on all three tasks while extending the effective context window by 3.5-20x. Kuang-Huei Lee, Hiroki Furuta, John F. Canny, Ian Fischer |
ICML | 5 |
| 2023 | Sparsity-Inducing Categorical Prior Improves Robustness of the Information BottleneckabstractThe information bottleneck framework provides a systematic approach to learning representations that compress nuisance information in the input and extract semantically meaningful information about predictions. However, the choice of a prior distribution that fixes the dimensionality across all the data can restrict the flexibility of this approach for learning robust representations. We present a novel sparsity-inducing spike-slab categorical prior that uses sparsity as a mechanism to provide the flexibility that allows each data point to learn its own dimension distribution. In addition, it provides a mechanism for learning a joint distribution of the latent variable and the sparsity, and hence it can account for the complete uncertainty in the latent space. Through a series of experiments using in-distribution and out-of-distribution learning scenarios on the MNIST, CIFAR-10, and ImageNet data, we show that the proposed approach improves accuracy and robustness compared to traditional fixed-dimensional priors, as well as other sparsity induction mechanisms for latent variable models proposed in the literature. Anirban Samaddar, Sandeep Madireddy, Prasanna Balaprakash, Taps Maiti, Gustavo de los Campos, Ian Fischer |
AISTATS | 6 |
| 2023 | Weighted Ensemble Self-Supervised Learning
Yangjun Ruan, Warren R. Morningstar, Alexander A. Alemi, Sergey Ioffe, Ian Fischer, Joshua V. Dillon |
ICLR | 6 |
| 2022 | Deep Hierarchical Planning from PixelsabstractIntelligent agents need to select long sequences of actions to solve complex tasks. While humans easily break down tasks into subgoals and reach them through millions of muscle commands, current artificial intelligence is limited to tasks with horizons of a few hundred decisions, despite large compute budgets. Research on hierarchical reinforcement learning aims to overcome this limitation but has proven to be challenging, current methods rely on manually specified goal spaces or subtasks, and no general solution exists. We introduce Director, a practical method for learning hierarchical behaviors directly from pixels by planning inside the latent space of a learned world model. The high-level policy maximizes task and exploration rewards by selecting latent goals and the low-level policy learns to achieve the goals. Despite operating in latent space, the decisions are interpretable because the world model can decode goals into images for visualization. Director learns successful behaviors across a wide range of environments, including visual control, Atari games, and DMLab levels and outperforms exploration methods on tasks with very sparse rewards, including 3D maze traversal with a quadruped robot from an egocentric camera and proprioception, without access to the global position or top-down view used by prior work. Danijar Hafner, Kuang-Huei Lee, Ian Fischer, Pieter Abbeel |
NeurIPS | 3 |
| 2022 | Multi-Game Decision TransformersabstractA longstanding goal of the field of AI is a method for learning a highly capable, generalist agent from diverse experience. In the subfields of vision and language, this was largely achieved by scaling up transformer-based models and training them on large, diverse datasets. Motivated by this progress, we investigate whether the same strategy can be used to produce generalist reinforcement learning agents. Specifically, we show that a single transformer-based model – with a single set of weights – trained purely offline can play a suite of up to 46 Atari games simultaneously at close-to-human performance. When trained and evaluated appropriately, we find that the same trends observed in language and vision hold, including scaling of performance with model size and rapid adaptation to new games via fine-tuning. We compare several approaches in this multi-game setting, such as online and offline RL methods and behavioral cloning, and find that our Multi-Game Decision Transformer models offer the best scalability and performance. We release the pre-trained models and code to encourage further research in this direction. Kuang-Huei Lee, Ofir Nachum, Sherry Yang 0001, Lisa Lee, Daniel Freeman, Sergio Guadarrama, Ian Fischer, Winnie Xu, Eric Jang, Henryk Michalewski, Igor Mordatch |
NeurIPS | 7 |
| 2021 | Compressive Visual RepresentationsabstractLearning effective visual representations that generalize well without human supervision is a fundamental problem in order to apply Machine Learning to a wide variety of tasks. Recently, two families of self-supervised methods, contrastive learning and latent bootstrapping, exemplified by SimCLR and BYOL respectively, have made significant progress. In this work, we hypothesize that adding explicit information compression to these algorithms yields better and more robust representations. We verify this by developing SimCLR and BYOL formulations compatible with the Conditional Entropy Bottleneck (CEB) objective, allowing us to both measure and control the amount of compression in the learned representation, and observe their impact on downstream tasks. Furthermore, we explore the relationship between Lipschitz continuity and compression, showing a tractable lower bound on the Lipschitz constant of the encoders we learn. As Lipschitz continuity is closely related to robustness, this provides a new explanation for why compressed models are more robust. Our experiments confirm that adding compression to SimCLR and BYOL significantly improves linear evaluation accuracies and model robustness across a wide range of domain shifts. In particular, the compressed version of BYOL achieves 76.0% Top-1 linear evaluation accuracy on ImageNet with ResNet-50, and 78.8% with ResNet-50 2x. Kuang-Huei Lee, Anurag Arnab, Sergio Guadarrama, John F. Canny, Ian Fischer |
NeurIPS | 5 |
| 2020 | An Unsupervised Information-Theoretic Perceptual Quality MetricabstractTractable models of human perception have proved to be challenging to build. Hand-designed models such as MS-SSIM remain popular predictors of human image quality judgements due to their simplicity and speed. Recent modern deep learning approaches can perform better, but they rely on supervised data which can be costly to gather: large sets of class labels such as ImageNet, image quality ratings, or both. We combine recent advances in information-theoretic objective functions with a computational architecture informed by the physiology of the human visual system and unsupervised training on pairs of video frames, yielding our Perceptual Information Metric (PIM). We show that PIM is competitive with supervised metrics on the recent and challenging BAPPS image quality assessment dataset and outperforms them in predicting the ranking of image compression methods in CLIC 2020. We also perform qualitative experiments using the ImageNet-C dataset, and establish that PIM is robust with respect to architectural details. Sangnie Bhardwaj, Ian Fischer, Jona Ballé, Troy T. Chinen |
NeurIPS | 2 |
| 2020 | Predictive Information Accelerates Learning in RLabstractThe Predictive Information is the mutual information between the past and the future, I(Xpast; Xfuture). We hypothesize that capturing the predictive information is useful in RL, since the ability to model what will happen next is necessary for success on many tasks. To test our hypothesis, we train Soft Actor-Critic (SAC) agents from pixels with an auxiliary task that learns a compressed representation of the predictive information of the RL environment dynamics using a contrastive version of the Conditional Entropy Bottleneck (CEB) objective. We refer to these as Predictive Information SAC (PI-SAC) agents. We show that PI-SAC agents can substantially improve sample efficiency over challenging baselines on tasks from the DM Control suite of continuous control environments. We evaluate PI-SAC agents by comparing against uncompressed PI-SAC agents, other compressed and uncompressed agents, and SAC agents directly trained from pixels. Our implementation is given on GitHub. Kuang-Huei Lee, Ian Fischer, Anthony Z. Liu, Yijie Guo, Honglak Lee, John F. Canny, Sergio Guadarrama |
NeurIPS | 2 |
| 2020 | Information-Bottleneck Approach to Salient Region Discovery
Andrey Zhmoginov, Ian Fischer, Mark Sandler 0002 |
ECML/PKDD (3) | 2 |
| 2019 | Learning Latent Dynamics for Planning from PixelsabstractPlanning has been very successful for control tasks with known environment dynamics. To leverage planning in unknown environments, the agent needs to learn the dynamics from interactions with the world. However, learning dynamics models that are accurate enough for planning has been a long-standing challenge, especially in image-based domains. We propose the Deep Planning Network (PlaNet), a purely model-based agent that learns the environment dynamics from images and chooses actions through fast online planning in latent space. To achieve high performance, the dynamics model must accurately predict the rewards ahead for multiple time steps. We approach this using a latent dynamics model with both deterministic and stochastic transition components. Moreover, we propose a multi-step variational inference objective that we name latent overshooting. Using only pixel observations, our agent solves continuous control tasks with contact dynamics, partial observability, and sparse rewards, which exceed the difficulty of tasks that were previously solved by planning with learned models. PlaNet uses substantially fewer episodes and reaches final performance close to and sometimes higher than strong model-free algorithms. Danijar Hafner, Timothy P. Lillicrap, Ian Fischer, Ruben Villegas, David Ha, Honglak Lee, James Davidson |
ICML | 3 |
| 2018 | Learning to Attack: Adversarial Transformation NetworksabstractWith the rapidly increasing popularity of deep neural networks for image recognition tasks, a parallel interest in generating adversarial examples to attack the trained models has arisen. To date, these approaches have involved either directly computing gradients with respect to the image pixels or directly solving an optimization on the image pixels. We generalize this pursuit in a novel direction: can a separate network be trained to efficiently attack another fully trained network? We demonstrate that it is possible, and that the generated attacks yield startling insights into the weaknesses of the target network. We call such a network an Adversarial Transformation Network (ATN). ATNs transform any input into an adversarial attack on the target network, while being minimally perturbing to the original inputs and the target network's outputs. Further, we show that ATNs are capable of not only causing the target network to make an error, but can be constructed to explicitly control the type of misclassification made. We demonstrate ATNs on both simple MNIST-digit classifiers and state-of-the-art ImageNet classifiers deployed by Google, Inc.: Inception ResNet-v2. Shumeet Baluja, Ian Fischer |
AAAI | 2 |
| 2018 | Generative Models of Visually Grounded Imagination
Ramakrishna Vedantam, Ian Fischer, Jonathan Huang, Kevin Murphy 0002 |
ICLR (Poster) | 2 |
| 2018 | Fixing a Broken ELBOabstractRecent work in unsupervised representation learning has focused on learning deep directed latentvariable models. Fitting these models by maximizing the marginal likelihood or evidence is typically intractable, thus a common approximation is to maximize the evidence lower bound (ELBO) instead. However, maximum likelihood training (whether exact or approximate) does not necessarily result in a good latent representation, as we demonstrate both theoretically and empirically. In particular, we derive variational lower and upper bounds on the mutual information between the input and the latent variable, and use these bounds to derive a rate-distortion curve that characterizes the tradeoff between compression and reconstruction accuracy. Using this framework, we demonstrate that there is a family of models with identical ELBO, but different quantitative and qualitative characteristics. Our framework also suggests a simple new method to ensure that latent variable models with powerful stochastic decoders do not ignore their latent code. Alexander A. Alemi, Ben Poole, Ian Fischer, Joshua V. Dillon, Rif A. Saurous, Kevin Murphy 0002 |
ICML | 3 |
| 2018 | GILBO: One Metric to Measure Them AllabstractWe propose a simple, tractable lower bound on the mutual information contained in the joint generative density of any latent variable generative model: the GILBO (Generative Information Lower BOund). It offers a data-independent measure of the complexity of the learned latent variable description, giving the log of the effective description length. It is well-defined for both VAEs and GANs. We compute the GILBO for 800 GANs and VAEs each trained on four datasets (MNIST, FashionMNIST, CIFAR-10 and CelebA) and discuss the results. Alexander A. Alemi, Ian Fischer |
NeurIPS | 2 |
| 2017 | Speed/Accuracy Trade-Offs for Modern Convolutional Object DetectorsabstractThe goal of this paper is to serve as a guide for selecting a detection architecture that achieves the right speed/memory/accuracy balance for a given application and platform. To this end, we investigate various ways to trade accuracy for speed and memory usage in modern convolutional object detection systems. A number of successful systems have been proposed in recent years, but apples-toapples comparisons are difficult due to different base feature extractors (e.g., VGG, Residual Networks), different default image resolutions, as well as different hardware and software platforms. We present a unified implementation of the Faster R-CNN [30], R-FCN [6] and SSD [25] systems, which we view as meta-architectures and trace out the speed/accuracy trade-off curve created by using alternative feature extractors and varying other critical parameters such as image size within each of these meta-architectures. On one extreme end of this spectrum where speed and memory are critical, we present a detector that achieves real time speeds and can be deployed on a mobile device. On the opposite end in which accuracy is critical, we present a detector that achieves state-of-the-art performance measured on the COCO detection task. Jonathan Huang, Vivek Rathod, Chen Sun 0002, Menglong Zhu, Anoop Korattikara Balan, Alireza Fathi, Ian Fischer, Zbigniew Wojna, Yang Song 0009, Sergio Guadarrama, Kevin Murphy 0002 |
CVPR | 7 |
| 2017 | Deep Variational Information Bottleneck
Alexander A. Alemi, Ian Fischer, Joshua V. Dillon, Kevin Murphy 0002 |
ICLR (Poster) | 2 |
| 2007 | The Emperor's New Security IndicatorsabstractWe evaluate Website authentication measures that are designed to protect users from man-in-the-middle, 'phishing', and other site forgery attacks. We asked 67 bank customers to conduct common online banking tasks. Each time they logged in, we presented increasingly alarming clues that their connection was insecure. First, we removed HTTPS indicators. Next, we removed the participant's site-authentication image--the customer-selected image that many Websites now expect their users to verify before entering their passwords. Finally, we replaced the bank's password-entry page with a warning page. After each clue, we determined whether participants entered their passwords or withheld them. We also investigate how a study's design affects participant behavior: we asked some participants to play a role and others to use their own accounts and passwords. We also presented some participants with security-focused instructions. We confirm prior findings that users ignore HTTPS indicators: no participants withheld their passwords when these indicators were removed. We present the first empirical investigation of site-authentication images, and we find them to be ineffective: even when we removed them, 23 of the 25 (92%) participants who used their own accounts entered their passwords. We also contribute the first empirical evidence that role playing affects participants' security behavior: role-playing participants behaved significantly less securely than those using their own passwords. Stuart E. Schechter, Rachna Dhamija, Andy Ozment, Ian Fischer |
S&P | 4 |