Nimrod Berman

dblp:344/1689 · DBLP profile ↗
← Back
9ranked-venue papers
4as first author
9since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 4 first-author · 9 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
8 papers
Generative modeling · 55% Representation and self-supervised learning · 39% Vision and language · 2%
Computer graphics and multimedia
1 paper
Image and video processing · 100%

Topics — the 12 heaviest of 15, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Generative modeling
diffusion model
3.442025
Time Series Generation Under Data Scarcity: A Unified Generative Modeling Approach · NeurIPS 2025
A Diffusion Model for Regular Time Series Generation from Irregular Data with Completion and Masking · NeurIPS 2025
Towards General Modality Translation with Contrastive and Predictive Latent Diffusion Bridge · NeurIPS 2025
Machine learning › Representation and self-supervised learning › representation learning
disentangled representation learning
2.942025
Disentanglement Beyond Static vs. Dynamic: A Benchmark and Evaluation Framework for Multi-Factor Sequential Representations · NeurIPS 2025
Sequential Disentanglement by Extracting Static Information From A Single Sequence Element · ICML 2024
Sample and Predict Your Latent: Modality-free Sequential Disentanglement via Contrastive Estimation · ICML 2023
Machine learning › Generative modeling › diffusion model
time series generation
2.532025
Time Series Generation Under Data Scarcity: A Unified Generative Modeling Approach · NeurIPS 2025
A Diffusion Model for Regular Time Series Generation from Irregular Data with Completion and Masking · NeurIPS 2025
Utilizing Image Transforms and Diffusion Models for Generative Modeling of Short and Long Time Series · NeurIPS 2024
Machine learning › Representation and self-supervised learning › representation learning › disentangled representation learning
sequential disentanglement
1.422024
Sequential Disentanglement by Extracting Static Information From A Single Sequence Element · ICML 2024
Sample and Predict Your Latent: Modality-free Sequential Disentanglement via Contrastive Estimation · ICML 2023
Machine learning › Generative modeling › image generation › data-efficient image generation
few-shot image generation
0.912025
Time Series Generation Under Data Scarcity: A Unified Generative Modeling Approach · NeurIPS 2025
Machine learning › Generative modeling › cross-modal generation
modality translation
0.912025
Towards General Modality Translation with Contrastive and Predictive Latent Diffusion Bridge · NeurIPS 2025
Machine learning › Representation and self-supervised learning › pre-training
multi-domain pre-training
0.912025
Time Series Generation Under Data Scarcity: A Unified Generative Modeling Approach · NeurIPS 2025
Machine learning › Representation and self-supervised learning
contrastive learning
0.712023
Sample and Predict Your Latent: Modality-free Sequential Disentanglement via Contrastive Estimation · ICML 2023
Machine learning › Deep learning architectures and training
transformer
0.312025
A Diffusion Model for Regular Time Series Generation from Irregular Data with Completion and Masking · NeurIPS 2025
Computer vision › Vision and language
vision-language model
0.312025
Disentanglement Beyond Static vs. Dynamic: A Benchmark and Evaluation Framework for Multi-Factor Sequential Representations · NeurIPS 2025
Machine learning › Transfer learning and domain adaptation
zero-shot evaluation
0.312025
Disentanglement Beyond Static vs. Dynamic: A Benchmark and Evaluation Framework for Multi-Factor Sequential Representations · NeurIPS 2025
Image and video processing
image transform
0.212024
Utilizing Image Transforms and Diffusion Models for Generative Modeling of Short and Long Time Series · NeurIPS 2024

Methods — techniques the papers use, named apart from their topics

diffusion model · 2.4koopman operator · 1.5variational inference · 1.4vision-based diffusion · 0.9time series completion · 0.9predictive loss · 0.9masking · 0.9latent exploration · 0.9latent diffusion bridge · 0.9contrastive learning · 0.9short-time fourier transform · 0.8delay embedding · 0.8
YearPublicationVenuePosition
2025 Disentanglement Beyond Static vs. Dynamic: A Benchmark and Evaluation Framework for Multi-Factor Sequential Representations
abstract
Learning disentangled representations in sequential data is a key goal in deep learning, with broad applications in vision, audio, and time series. While real-world data involves multiple interacting semantic factors over time, prior work has mostly focused on simpler two-factor static and dynamic settings, primarily because such settings make data collection easier, thereby overlooking the inherently multi-factor nature of real-world data. We introduce the first standardized benchmark for evaluating multi-factor sequential disentanglement across six diverse datasets spanning video, audio, and time series. Our benchmark includes modular tools for dataset integration, model development, and evaluation metrics tailored to multi-factor analysis. We additionally propose a post-hoc Latent Exploration Stage to automatically align latent dimensions with semantic factors, and introduce a Koopman-inspired model that achieves state-of-the-art results. Moreover, we show that Vision-Language Models can automate dataset annotation and serve as zero-shot disentanglement evaluators, removing the need for manual labels and human intervention. Together, these contributions provide a robust and scalable foundation for advancing multi-factor sequential disentanglement. Our code is available on GitHub, and the datasets and trained models are available on Hugging Face.
Tal Barami, Nimrod Berman, Ilan Naiman, Amos H. Hason, Rotem Ezra, Omri Azencot
NeurIPS2
2025 Towards General Modality Translation with Contrastive and Predictive Latent Diffusion Bridge
abstract
Recent advances in generative modeling have positioned diffusion models as state-of-the-art tools for sampling from complex data distributions. While these models have shown remarkable success across single-modality domains such as images and audio, extending their capabilities to *Modality Translation (MT)*, translating information across different sensory modalities, remains an open challenge. Existing approaches often rely on restrictive assumptions, including shared dimensionality, Gaussian source priors, and modality-specific architectures, which limit their generality and theoretical grounding. In this work, we propose the Latent Denoising Diffusion Bridge Model (LDDBM), a general-purpose framework for modality translation based on a latent-variable extension of Denoising Diffusion Bridge Models. By operating in a shared latent space, our method learns a bridge between arbitrary modalities without requiring aligned dimensions. We introduce a contrastive alignment loss to enforce semantic consistency between paired samples and design a domain-agnostic encoder-decoder architecture tailored for noise prediction in latent space. Additionally, we propose a predictive loss to guide training toward accurate cross-domain translation and explore several training strategies to improve stability. Our approach supports arbitrary modality pairs and performs strongly on diverse MT tasks, including multi-view to 3D shape generation, image super-resolution, and multi-view scene synthesis. Comprehensive experiments and ablations validate the effectiveness of our framework, establishing a new strong baseline in general modality translation. For more information, see our project page: https://sites.google.com/view/lddbm/home.
Nimrod Berman, Omkar Joglekar, Eitan Kosman, Dotan Di Castro, Omri Azencot
NeurIPS1
2025 One-Step Offline Distillation of Diffusion-based Models via Koopman Modeling
abstract
Diffusion-based generative models have demonstrated exceptional performance, yet their iterative sampling procedures remain computationally expensive. A prominent strategy to mitigate this cost is *distillation*, with *offline distillation* offering particular advantages in terms of efficiency, modularity, and flexibility. In this work, we identify two key observations that motivate a principled distillation framework: (1) while diffusion models have been viewed through the lens of dynamical systems theory, powerful and underexplored tools can be further leveraged; and (2) diffusion models inherently impose structured, semantically coherent trajectories in latent space. Building on these observations, we introduce the *Koopman Distillation Model* (KDM), a novel offline distillation approach grounded in Koopman theory - a classical framework for representing nonlinear dynamics linearly in a transformed space. KDM encodes noisy inputs into an embedded space where a learned linear operator propagates them forward, followed by a decoder that reconstructs clean samples. This enables single-step generation while preserving semantic fidelity. We provide theoretical justification for our approach: (1) under mild assumptions, the learned diffusion dynamics admit a finite-dimensional Koopman representation; and (2) proximity in the Koopman latent space correlates with semantic similarity in the generated outputs, allowing for effective trajectory alignment. Empirically, KDM achieves state-of-the-art performance across standard *offline distillation* benchmarks - improving FID scores by up to 40% in a single generation step.
Nimrod Berman, Ilan Naiman, Moshe Eliasof, Hedi Zisling, Omri Azencot
NeurIPS1
2025 A Diffusion Model for Regular Time Series Generation from Irregular Data with Completion and Masking
abstract
Generating realistic time series data is critical for applications in healthcare, finance, and climate science. However, irregular sampling and missing values present significant challenges. While prior methods address these irregularities, they often yield suboptimal results and incur high computational costs. Recent advances in regular time series generation, such as the diffusion-based ImagenTime model, demonstrate strong, fast, and scalable generative capabilities by transforming time series into image representations, making them a promising solution. However, extending ImagenTime to irregular sequences using simple masking introduces ``unnatural'' neighborhoods, where missing values replaced by zeros disrupt the learning process. To overcome this, we propose a novel two-step framework: first, a Time Series Transformer completes irregular sequences, creating natural neighborhoods; second, a vision-based diffusion model with masking minimizes dependence on the completed values. This hybrid approach leverages the strengths of both completion and masking, enabling robust and efficient generation of realistic time series. Our method achieves state-of-the-art performance across benchmarks, delivering a relative improvement in discriminative score by 70% and in computational cost by 85%.
Gal Fadlon, Idan Arbiv, Nimrod Berman, Omri Azencot
NeurIPS3
2025 Time Series Generation Under Data Scarcity: A Unified Generative Modeling Approach
abstract
Generative modeling of time series is a central challenge in time series analysis, particularly under data-scarce conditions. Despite recent advances in generative modeling, a comprehensive understanding of how state-of-the-art generative models perform under limited supervision remains lacking. In this work, we conduct the first large-scale study evaluating leading generative models in data-scarce settings, revealing a substantial performance gap between full-data and data-scarce regimes. To close this gap, we propose a unified diffusion-based generative framework that can synthesize high-fidelity time series across diverse domains using just a few examples. Our model is pretrained on a large, heterogeneous collection of time series datasets, enabling it to learn generalizable temporal representations. It further incorporates architectural innovations such as dynamic convolutional layers for flexible channel adaptation and dataset token conditioning for domain-aware generation. Without requiring abundant supervision, our unified model achieves state-of-the-art performance in few-shot settings—outperforming domain-specific baselines across a wide range of subset sizes. Remarkably, it also surpasses all baselines even when tested on full datasets benchmarks, highlighting the strength of pretraining and cross-domain generalization. We hope this work encourages the community to revisit few-shot generative modeling as a key problem in time series research and pursue unified solutions that scale efficiently across domains. Code is available at https://github.com/azencot-group/ImagenFew.
Tal Gonen, Itai Pemper, Ilan Naiman, Nimrod Berman, Omri Azencot
NeurIPS4
2024 Sequential Disentanglement by Extracting Static Information From A Single Sequence Element
abstract
One of the fundamental representation learning tasks is unsupervised sequential disentanglement, where latent codes of inputs are decomposed to a single static factor and a sequence of dynamic factors. To extract this latent information, existing methods condition the static and dynamic codes on the entire input sequence. Unfortunately, these models often suffer from information leakage, i.e., the dynamic vectors encode both static and dynamic information, or vice versa, leading to a non-disentangled representation. Attempts to alleviate this problem via reducing the dynamic dimension and auxiliary loss terms gain only partial success. Instead, we propose a novel and simple architecture that mitigates information leakage by offering a simple and effective subtraction inductive bias while conditioning on a single sample. Remarkably, the resulting variational framework is simpler in terms of required loss terms, hyper-parameters, and data augmentation. We evaluate our method on multiple data-modality benchmarks including general time series, video, and audio, and we show beyond state-of-the-art results on generation and prediction tasks in comparison to several strong baselines.
Nimrod Berman, Ilan Naiman, Idan Arbiv, Gal Fadlon, Omri Azencot
ICML1
2024 Utilizing Image Transforms and Diffusion Models for Generative Modeling of Short and Long Time Series
abstract
Lately, there has been a surge in interest surrounding generative modeling of time series data. Most existing approaches are designed either to process short sequences or to handle long-range sequences. This dichotomy can be attributed to gradient issues with recurrent networks, computational costs associated with transformers, and limited expressiveness of state space models. Towards a unified generative model for varying-length time series, we propose in this work to transform sequences into images. By employing invertible transforms such as the delay embedding and the short-time Fourier transform, we unlock three main advantages: i) We can exploit advanced diffusion vision models; ii) We can remarkably process short- and long-range inputs within the same framework; and iii) We can harness recent and established tools proposed in the time series to image literature. We validate the effectiveness of our method through a comprehensive evaluation across multiple tasks, including unconditional generation, interpolation, and extrapolation. We show that our approach achieves consistently state-of-the-art results against strong baselines. In the unconditional generation tasks, we show remarkable mean improvements of $58.17$% over previous diffusion models in the short discriminative score and $132.61$% in the (ultra-)long classification scores. Code is at https://github.com/azencot-group/ImagenTime.
Ilan Naiman, Nimrod Berman, Itai Pemper, Idan Arbiv, Gal Fadlon, Omri Azencot
NeurIPS2
2023 Multifactor Sequential Disentanglement via Structured Koopman Autoencoders
Nimrod Berman, Ilan Naiman, Omri Azencot
ICLR1
2023 Sample and Predict Your Latent: Modality-free Sequential Disentanglement via Contrastive Estimation
abstract
Unsupervised disentanglement is a long-standing challenge in representation learning. Recently, self-supervised techniques achieved impressive results in the sequential setting, where data is time-dependent. However, the latter methods employ modality-based data augmentations and random sampling or solve auxiliary tasks. In this work, we propose to avoid that by generating, sampling, and comparing empirical distributions from the underlying variational model. Unlike existing work, we introduce a self-supervised sequential disentanglement framework based on contrastive estimation with no external signals, while using common batch sizes and samples from the latent space itself. In practice, we propose a unified, efficient, and easy-to-code sampling strategy for semantically similar and dissimilar views of the data. We evaluate our approach on video, audio, and time series benchmarks. Our method presents state-of-the-art results in comparison to existing techniques. The code is available at https://github.com/azencot-group/SPYL.
Ilan Naiman, Nimrod Berman, Omri Azencot
ICML2