Max Ehrlich

dblp:177/8998 · DBLP profile ↗
← Back
10ranked-venue papers
3as first author
7since 2021 · last 2025
0000-0003-3975-5788ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 2 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 3 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2025 Eagle 2.5: Boosting Long-Context Post-Training for Frontier Vision-Language Models
abstract
We introduce Eagle2.5, a frontier vision-language model (VLM) for long-context multimodal learning. Our work addresses the challenges in long video comprehension and high-resolution image understanding, introducing a generalist framework for both tasks. The proposed training framework incorporates Automatic Degrade Sampling and Image Area Preservation, two techniques that preserve contextual integrity and visual details. The framework also includes numerous efficiency optimizations in the pipeline for long-context data training. Finally, we propose Eagle-Video-110K, a novel dataset that integrates both story-level and clip-level annotations, facilitating long-video understanding. Eagle2.5 demonstrates substantial improvements on long-context multimodal benchmarks, providing a robust solution to the limitations of existing VLMs. Notably, our best model Eagle2.5-8B achieves 72.4\% on Video-MME with 512 input frames, matching the results of top-tier commercial model such as GPT-4o and large-scale open-source models like Qwen2.5-VL-72B and InternVL2.5-78B.
Guo Chen 0006, Jindong Jiang, Lidong Lu, De-An Huang, Wonmin Byeon, Matthieu Le, Max Ehrlich, Tong Lu 0002, Limin Wang 0002, Bryan Catanzaro, Jan Kautz, Andrew Tao, Zhiding Yu, Guilin Liu
NeurIPS10
2024 VEMIC: View-aware Entropy model for Multi-view Image Compression
Susmija Jabbireddy, Davit Soselia, Max Ehrlich, Christopher A. Metzler, Amitabh Varshney
BMVC3
2024 Explaining the Implicit Neural Canvas: Connecting Pixels to Neurons by Tracing Their Contributions
abstract
The many variations of Implicit Neural Representations (INRs), where a neural network is trained as a continuous representation of a signal, have tremendous practical utility for downstream tasks including novel view synthesis, video compression, and image super-resolution. Unfortunately, the inner workings of these networks are seriously under-studied. Our work, eXplaining the Implicit Neural Canvas (XINC), is a unified framework for explaining properties of INRs by examining the strength of each neuron's contribution to each output pixel. We call the aggregate of these contribution maps the Implicit Neural Canvas and we use this concept to demonstrate that the INRs we study learn to “see” the frames they represent in surprising ways. For ex-ample, INRs tend to have highly distributed representations. While lacking high-level object semantics, they have a sig-nificant bias for color and edges, and are almost entirely space-agnostic. We arrive at our conclusions by examining how objects are represented across time in video INRs, using clustering to visualize similar neurons across layers and architectures, and show that this is dominated by motion. These insights demonstrate the general usefulness of our analysis framework.
Namitha Padmanabhan, Matthew Gwilliam, Pulkit Kumar, Shishira R. Maiya, Max Ehrlich, Abhinav Shrivastava
CVPR5
2024 Latent-INR: A Flexible Framework for Implicit Representations of Videos with Discriminative Semantics
Shishira R. Maiya, Matthew Gwilliam, Max Ehrlich, Abhinav Shrivastava
ECCV (15)4
2024 Leveraging Bitstream Metadata for Fast, Accurate, Generalized Compressed Video Quality Enhancement
abstract
Video compression is a central feature of the modern internet powering technologies from social media to video conferencing. While video compression continues to mature, for many compression settings, quality loss is still noticeable. These settings nevertheless have important applications to the efficient transmission of videos over bandwidth constrained or otherwise unstable connections. In this work, we develop a deep learning architecture capable of restoring detail to compressed videos which leverages the underlying structure and motion information embedded in the video bitstream. We show that this improves restoration accuracy compared to prior compression correction methods and is competitive when compared with recent deep-learning-based video compression methods on rate-distortion while achieving higher throughput. Furthermore, we condition our model on quantization data which is readily available in the bit-stream. This allows our single model to handle a variety of different compression quality settings which required an ensemble of models in prior work.
Max Ehrlich, Jon Barker, Namitha Padmanabhan, Larry Davis 0001, Andrew Tao, Bryan Catanzaro, Abhinav Shrivastava
WACV1
2023 Unifying the Harmonic Analysis of Adversarial Attacks and Robustness
Shishira R. Maiya, Max Ehrlich, Vatsal Agarwal, Ser-Nam Lim, Tom Goldstein, Abhinav Shrivastava
BMVC2
2023 NIRVANA: Neural Implicit Representations of Videos with Adaptive Networks and Autoregressive Patch-Wise Modeling
abstract
Implicit Neural Representations (INR) have recently shown to be powerful tool for high-quality video compression. However, existing works are are limiting as they do not exploit the temporal redundancy in videos, leading to a long encoding time. Additionally, these methods have fixed architectures which do not scale to longer videos or higher resolutions. To address these issues, we propose NIRVANA, which treats videos as groups of frames and fits separate networks to each group performing patch-wise prediction. The video representation is modeled autoregressively, with networks fit on a current group initialized using weights from the previous group's model. To enhance efficiency, we quantize the parameters during training, requiring no post-hoc pruning or quantization. When compared with previous works on the benchmark UVG dataset, NIRVANA improves encoding quality from 37.36 to 37.70 (in terms of PSNR) and the encoding speed by 12x, while maintaining the same compression rate. In contrast to prior video INR works which struggle with larger resolution and longer videos, we show that our algorithm scales naturally due to its patch-wise and autoregressive design. Moreover, our method achieves variable bitrate compression by adapting to videos with varying inter-frame motion. NIRVANA also achieves 6x decoding speed scaling well with more GPUs, making it practical for various deployment scenarios.11The project site can be found here.
Shishira R. Maiya, Sharath Girish, Max Ehrlich, Hanyu Wang 0002, Kwot Sin Lee, Patrick Poirson, Pengxiang Wu, Abhinav Shrivastava
CVPR3
2020 Quantization Guided JPEG Artifact Correction
Max Ehrlich, Larry Davis 0001, Ser-Nam Lim, Abhinav Shrivastava
ECCV (8)1
2019 Deep Residual Learning in the JPEG Transform Domain
abstract
We introduce a general method of performing Residual Network inference and learning in the JPEG transform domain that allows the network to consume compressed images as input. Our formulation leverages the linearity of the JPEG transform to redefine convolution and batch normalization with a tune-able numerical approximation for ReLu. The result is mathematically equivalent to the spatial domain network up to the ReLu approximation accuracy. A formulation for image classification and a model conversion algorithm for spatial domain networks are given as examples of the method. We show that the sparsity of the JPEG format allows for faster processing of images with little to no penalty in the network accuracy.
Max Ehrlich, Larry Davis 0001
ICCV1
2019 Unsupervised Super-Resolution of Satellite Imagery for High Fidelity Material Label Transfer
abstract
Urban material recognition in remote sensing imagery is a challenging problem due to the difficulty of obtaining human annotations, especially on low resolution satellite images. To this end, we propose an unsupervised domain adaptation-based approach using adversarial learning. We aim to harvest information from smaller quantities of high resolution data (source domain) and utilize the same to super-resolve low resolution imagery (target domain). This can potentially aid in semantic as well as material label transfer from a richly annotated source to a target domain.
Arthita Ghosh, Max Ehrlich, Larry Davis 0001, Rama Chellappa
IGARSS2