VLDB 2026 Research / reviewers in the wild / expert
Teddy Lazebnik
dblp:285/5492
· DBLP profile ↗
13ranked-venue papers
10as first author
13since 2021 · last 2026
0000-0002-7851-8147ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 8 first-author · 11 since 2021Databases, data management, data science and information retrieval · 3 · 3 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Transforming norm-based to graph-based spatial representation for spatio-temporal epidemiological modelsabstractPandemics, with their profound societal and economic impacts, pose significant threats to global health, mortality rates, economic stability, and political landscapes. In response to these challenges, numerous studies have employed spatio-temporal models to enhance our understanding and management of these complex phenomena. These spatio-temporal models can be roughly divided into two main spatial categories: norm-based and graph-based. Norm-based models are usually more accurate and easier to model, but are more computationally intensive and require more data to fit. On the other hand, graph-based models are less accurate and harder to model, but are less computationally intensive and require fewer data to fit. As such, ideally, one would like to use a graph-based model while preserving the representation accuracy obtained by the norm-based model. In this study, we explore the ability to transform from norm-based to graph-based spatial representation for these models. We first show that no analytical mapping between the two exists, requiring one to use numerical approximation methods instead. We introduce a novel framework for this task, together with twelve possible implementations using a wide range of heuristic optimization approaches. Our findings show that by leveraging agent-based simulations and heuristic algorithms for the graph node’s location and population’s spatial walk dynamics approximation, one can use graph-based spatial representation without losing much of the model’s accuracy and expressiveness. We investigate our framework for three real-world cases, achieving 93% accuracy preservation, on average, while obtaining 86% relative computational time reduction. Moreover, an analysis of synthetic cases shows the proposed framework is relatively robust for changes in both spatial and temporal properties. Teddy Lazebnik |
Eng. Appl. Artif. Intell. | 1 |
| 2026 | Break a Lag: Triple Exponential Moving Average for Enhanced OptimizationabstractAbstract The performance of deep learning models hinges on the effectiveness of their optimization strategies, yet existing methods remain fundamentally constrained. Many optimizers rely on first-order Exponential Moving Average (EMA) techniques, which struggle to accurately track complex gradient trends, resulting in a lag in trend identification and suboptimal convergence. To overcome this critical limitation, we introduce Fast Adaptive Moment Estimation (FAME), a novel optimizer that harnesses the power of Triple Exponential Moving Average (TEMA). By inherently integrating a multi-level tracking mechanism, FAME significantly enhances gradient responsiveness, mitigates trend lag, and maximizes learning efficiency. Our extensive evaluation spans diverse computer vision tasks-including image classification, object detection, and semantic segmentation-incorporating FAME into 22 distinct architectures, from lightweight CNNs to cutting-edge Vision Transformers. Rigorous benchmarking against state-of-the-art optimizers confirms FAME’s superiority in accuracy, stability, and robustness. Notably, FAME excels in scalability, demonstrating consistent and substantial performance gains across architectures, tasks, and datasets, positioning it as a powerful optimizer for computer vision. Roi Peleg, Yair Smadar, Teddy Lazebnik, Assaf Hoogi |
Mach. Learn. | 3 |
| 2026 | Interpretable knowledge distillation via symbolic regression for feedforward neural networksabstractAbstract Neural networks (NNs) are widely used for modeling complex, high-dimensional relationships but often lack interpretability, limiting their adoption in critical domains such as healthcare, finance, and engineering. Symbolic regression (SR), in contrast, generates explicit mathematical expressions that enhance transparency but typically underperform in predictive accuracy. To bridge this gap, we propose a knowledge distillation framework that approximates the activations of a trained feedforward NN’s final hidden layer using SR models. This approach enhances interpretability while retaining a substantial fraction of the neural network’s predictive performance in structured data settings. Our method is evaluated across 20 diverse datasets, demonstrating a 7-21% improvement in RMSE over baseline SR models. These results highlight the potential of symbolic knowledge distillation as a practical tool for enhancing model transparency in structured data applications. Assaf Shmuel, Nir Koren, Oren Glickman, Teddy Lazebnik |
Neural Comput. Appl. | 4 |
| 2025 | Pulling the carpet below the learner's feet: Genetic algorithm to learn ensemble machine learning model during concept driftabstractData-driven models, in general, and machine learning (ML) models, in particular, have gained popularity over recent years with an increased usage of such models across the scientific and engineering domains. When using ML models in realistic and dynamic environments, users often need to handle the challenge of concept drift (CD). In this study, we explore the application of genetic algorithms (GAs) to address the challenges posed by CD in such settings. Formally, we propose a novel two-level ensemble ML model, which combines a global ML model with a CD detector, operating as an aggregator for a population of ML pipeline models, each one with an adjusted CD detector by itself responsible for re-training its ML model. In addition, we show that one can further improve the proposed model by utilizing off-the-shelf automatic ML (AutoML) methods. Through extensive synthetic dataset analysis, we show that the proposed model statistically significantly outperforms an ML pipeline with a CD algorithm, particularly in scenarios with unknown CD characteristics or a mixture of moving and shifting CDs. Moreover, we show a sub-linear decline in the proposed method’s performance with respect to a higher drifting rate and robustness to the underlying AutoML method utilized. Teddy Lazebnik |
Eng. Appl. Artif. Intell. | 1 |
| 2025 | A comprehensive benchmark of machine and deep learning models on structured data for regression and classificationabstractThe analysis of tabular datasets is highly prevalent both in scientific research and real-world applications of Machine Learning (ML). Unlike many other ML tasks, Deep Learning (DL) models often do not outperform traditional methods in this area. Previous comparative benchmarks have shown that DL performance is frequently equivalent to or even inferior to models such as Gradient Boosting Machines (GBMs). In this study, we introduce a comprehensive benchmark aimed at better characterizing the types of datasets where DL models excel. Although several important benchmarks for tabular datasets already exist, our contribution lies in the variety and depth of our comparison: we evaluate 111 datasets with 20 different models, including both regression and classification tasks. These datasets vary in scale and include both those with and without categorical variables. Importantly, our benchmark contains a sufficient number of datasets where DL models perform best, allowing for a thorough analysis of the conditions under which DL models excel. Building on the results of this benchmark, we train a model that predicts scenarios where DL models significantly outperform alternative methods, considering only datasets where the performance difference between the two groups is statistically significant. This filtering yields 36 datasets out of the original 111. On this subset, our model achieves 92 % accuracy. We present insights derived from this characterization and compare these findings to previous benchmarks. Assaf Shmuel, Oren Glickman, Teddy Lazebnik |
Neurocomputing | 3 |
| 2025 | Mind Your Manners: The Dynamics of Politeness in Human-AI vs. Human-Human InteractionsabstractThe rapid integration of artificial intelligence (AI) into communication systems has significantly altered how users interact with digital tools and collaborate with AI agents. This study investigates the dynamics of politeness in human-AI interactions through a controlled experiment with 1,684 participants, each completing sequential text-based tasks with a conversational AI system. Participants were randomly assigned to one of several conditions that varied in the AI's visual identity (no icon, robot icon, or human face), allowing us to examine the role of perceived anthropomorphism through a minimal visual cue. Politeness was measured using linguistic markers and analyzed using statistical models that account for task sequence and individual differences. Our findings show that politeness toward AI declines over time, with a temporary increase at the start of a second task. Compared to human-human interactions in a benchmark dataset, politeness in human-AI interactions eroded more quickly. Younger participants were less polite overall, and although frequent AI users also appeared less polite descriptively, adjusted models showed a small positive association with daily AI use. Anthropomorphic visual cues, especially human-like avatars, led to more sustained polite behavior. These results offer insight into how users adapt social norms in AI-mediated collaboration and suggest design strategies for fostering respectful and effective human-AI communication. Teddy Lazebnik, Lior Zalmanson, Osnat Mokryn |
Proc. ACM Hum. Comput. Interact. | 1 |
| 2024 | An algorithm to optimize explainability using feature ensemblesabstractAbstract Feature Ensembles are a robust and effective method for finding the feature set that yields the best predictive accuracy for learning agents. However, current feature ensemble algorithms do not consider explainability as a key factor in their construction. To address this limitation, we present an algorithm that optimizes for the explainability and performance of a model – theOptimizingFeatureEnsembles forExplainability (OFEE) algorithm. OFEE uses intersections of feature sets to produce a feature ensemble that optimally balances explainability and performance. Furthermore, OFEE is parameter-free and as such optimizes itself to a given dataset and explainability requirements. To evaluated OFEE, we considered two explainability measures, one based on ensemble size and the other based on ensemble stability. We found that OFEE was overall extremely effective within the nine canonical datasets we considered. It outperformed other feature selection algorithms by an average of over 8% and 7% respectively when considering the size and stability explainability measures. Teddy Lazebnik, Svetlana Bunimovich-Mendrazitsky, Avi Rosenfeld |
Appl. Intell. | 1 |
| 2024 | Temporal graphs anomaly emergence detection: benchmarking for social media interactionsabstractAbstract Temporal graphs have become an essential tool for analyzing complex dynamic systems with multiple agents. Detecting anomalies in temporal graphs is crucial for various applications, including identifying emerging trends, monitoring network security, understanding social dynamics, tracking disease outbreaks, and understanding financial dynamics. In this paper, we present a comprehensive benchmarking study that compares 12 data-driven methods for anomaly detection in temporal graphs. We conduct experiments on two temporal graphs extracted from Twitter and Facebook, aiming to identify anomalies in group interactions. Surprisingly, our study reveals an unclear pattern regarding the best method for such tasks, highlighting the complexity and challenges involved in anomaly emergence detection in large and dynamic systems. The results underscore the need for further research and innovative approaches to effectively detect emerging anomalies in dynamic systems represented as temporal graphs. Teddy Lazebnik, Or Iny |
Appl. Intell. | 1 |
| 2024 | Knowledge-integrated autoencoder modelabstractData encoding is a common and central operation in most data analysis tasks. The performance of other models downstream in the computational process highly depends on the quality of data encoding. One of the most powerful ways to encode data is using the neural network AutoEncoder (AE) architecture. However, the developers of AE cannot easily influence the produced embedding space, as it is usually treated as a black box technique. This means the embedding space is uncontrollable and does not necessarily possess the properties desired for downstream tasks. This paper introduces a novel approach for developing AE models that can integrate external knowledge sources into the learning process, possibly leading to more accurate results. The proposed Knowledge-integrated AutoEncoder (KiAE) model can leverage domain-specific information to make sure the desired distance and neighborhood properties between samples are preservative in the embedding space. The proposed model is evaluated on three large-scale datasets from three scientific fields and is compared to nine existing encoding models. The results demonstrate that the KiAE model effectively captures the underlying structures and relationships between the input data and external knowledge, meaning it generates a more useful representation. This leads to outperforming the rest of the models in terms of reconstruction accuracy. Teddy Lazebnik, Liron Simon Keren |
Expert Syst. Appl. | 1 |
| 2023 | Decision tree post-pruning without loss of accuracy using the SAT-PP algorithm with an empirical evaluation on clinical data
Teddy Lazebnik, Svetlana Bunimovich-Mendrazitsky |
Data Knowl. Eng. | 1 |
| 2023 | Data-driven hospitals staff and resources allocation using agent-based simulation and deep reinforcement learning
Teddy Lazebnik |
Eng. Appl. Artif. Intell. | 1 |
| 2022 | Demonstrating SubStrat: A Subset-Based Strategy for Faster AutoML on Large DatasetsabstractAutomated machine learning (AutoML) frameworks are gaining popularity among data scientists as they dramatically reduce the manual work devoted to the construction of ML pipelines while obtaining similar and sometimes even better results than manually-built models. Such frameworks intelligently search among millions of possible ML pipeline configurations to finally retrieve an optimal pipeline in terms of predictive accuracy. However, when the training dataset is large, the construction and evaluation of a single ML pipeline take longer, which makes the overall AutoML running times increasingly high. Teddy Lazebnik, Amit Somech |
CIKM | 1 |
| 2022 | SubStrat: A Subset-Based Optimization Strategy for Faster AutoMLabstractAutomated machine learning (AutoML) frameworks have become important tools in the data scientist's arsenal, as they dramatically reduce the manual work devoted to the construction of ML pipelines. Such frameworks intelligently search among millions of possible ML pipelines - typically containing feature engineering, model selection, and hyper parameters tuning steps - and finally output an optimal pipeline in terms of predictive accuracy. However, when the dataset is large, each individual configuration takes longer to execute, therefore the overall AutoML running times become increasingly high. To this end, we present SubStrat, an AutoML optimization strategy that tackles the data size, rather than configuration space. It wraps existing AutoML tools, and instead of executing them directly on the entire dataset, SubStrat uses a genetic-based algorithm to find a small yet representative data subset that preserves a particular characteristic of the full data. It then employs the AutoML tool on the small subset, and finally, it refines the resulting pipeline by executing a restricted, much shorter, AutoML process on the large dataset. Our experimental results, performed on three popular AutoML frameworks, Auto-Sklearn, TPOT, and H2O show that SubStrat reduces their running times by 76.3% (on average), with only a 4.15% average decrease in the accuracy of the resulting ML pipeline. Teddy Lazebnik, Amit Somech, Abraham Itzhak Weinberg |
Proc. VLDB Endow. | 1 |