Maor Ivgi

dblp:275/3578 · DBLP profile ↗
← Back
8ranked-venue papers
4as first author
8since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 3 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Language models and text generation · 29% Efficient and distributed learning · 28% Deep learning architectures and training · 14%
Theoretical computer science
1 paper
Mathematical optimization · 100%
Databases, data mining, and information retrieval
1 paper
Information retrieval · 100%

Topics — the 11 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation › large language model training
language model pretraining
0.812024
DataComp-LM: In search of the next generation of training sets for language models · NeurIPS 2024
Machine learning › Efficient and distributed learning
model acceleration
0.812024
Accelerated Parameter-Free Stochastic Optimization · COLT 2024
Machine learning › Efficient and distributed learning › data curation
training data curation
0.812024
DataComp-LM: In search of the next generation of training sets for language models · NeurIPS 2024
Mathematical optimization › online optimization
parameter-free optimization
0.812024
Accelerated Parameter-Free Stochastic Optimization · COLT 2024
Mathematical optimization › stochastic optimization
stochastic convex optimization
0.812024
Accelerated Parameter-Free Stochastic Optimization · COLT 2024
Machine learning › Optimization for machine learning
stochastic gradient descent
0.712023
DoG is SGD's Best Friend: A Parameter-Free Dynamic Step Size Schedule · ICML 2023
Natural language and speech › Language models and text generation › natural language understanding
long document understanding
0.612022
SCROLLS: Standardized CompaRison Over Long Language Sequences · EMNLP 2022
Information retrieval
evaluation
0.612022
SCROLLS: Standardized CompaRison Over Long Language Sequences · EMNLP 2022
Machine learning › Trustworthy machine learning › robustness
adversarial robustness
0.512021
Achieving Model Robustness through Discrete Adversarial Training · EMNLP (1) 2021
Natural language and speech › Language models and text generation › large language model training
data mixing
0.212024
DataComp-LM: In search of the next generation of training sets for language models · NeurIPS 2024
Natural language and speech › Information extraction and text analysis › text classification
robust text classification
0.112021
Achieving Model Robustness through Discrete Adversarial Training · EMNLP (1) 2021

Methods — techniques the papers use, named apart from their topics

iterate stabilization · 1.5dog · 1.5UniXGrad · 1.5benchmark construction · 1.1model-based filtering · 0.8deduplication · 0.8stochastic gradient descent · 0.7adam · 0.7online augmentation · 0.5best-first search · 0.5
YearPublicationVenuePosition
2025 In-Context Learning with Long-Context Models: An In-Depth Exploration
abstract
Amanda Bertsch, Maor Ivgi, Emily Xiao, Uri Alon, Jonathan Berant, Matthew R. Gormley, Graham Neubig. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Amanda Bertsch, Maor Ivgi, Emily Xiao, Uri Alon 0002, Jonathan Berant, Matthew R. Gormley, Graham Neubig
NAACL (Long Papers)2
2024 Accelerated Parameter-Free Stochastic Optimization
abstract
We propose a method that achieves near-optimal rates for \emph{smooth} stochastic convex optimization and requires essentially no prior knowledge of problem parameters. This improves on prior work which requires knowing at least the initial distance to optimality $d_0$. Our method, \textsc{U-DoG}, combines \textsc{UniXGrad} (Kavis et al., 2019) and \textsc{DoG} (Ivgi et al., 2023) with novel iterate stabilization techniques. It requires only loose bounds on $d_0$ and the noise magnitude, provides high probability guarantees under sub-Gaussian noise, and is also near-optimal in the non-smooth case. Our experiments show consistent, strong performance on convex problems and mixed results on neural network training.
Itai Kreisler, Maor Ivgi, Oliver Hinder, Yair Carmon
COLT2
2024 DataComp-LM: In search of the next generation of training sets for language models
abstract
We introduce DataComp for Language Models, a testbed for controlled dataset experiments with the goal of improving language models.As part of DCLM, we provide a standardized corpus of 240T tokens extracted from Common Crawl, effective pretraining recipes based on the OpenLM framework, and a broad suite of 53 downstream evaluations.Participants in the DCLM benchmark can experiment with data curation strategies such as deduplication, filtering, and data mixing atmodel scales ranging from 412M to 7B parameters.As a baseline for DCLM, we conduct extensive experiments and find that model-based filtering is key to assembling a high-quality training set.The resulting dataset, DCLM-Baseline, enables training a 7B parameter language model from scratch to 63% 5-shot accuracy on MMLU with 2T training tokens.Compared to MAP-Neo, the previous state-of-the-art in open-data language models, DCLM-Baseline represents a 6 percentage point improvement on MMLU while being trained with half the compute.Our results highlight the importance of dataset design for training language models and offer a starting point for further research on data curation. We release the \dclm benchmark, framework, models, and datasets at https://www.datacomp.ai/dclm/
Jeffrey Li, Alex Fang, Georgios Smyrnis, Maor Ivgi, Matt Jordan, Samir Yitzhak Gadre, Hritik Bansal, Etash Kumar Guha, Sedrick Keh, Kushal Arora, Niklas Muennighoff, Reinhard Heckel, Jean Mercat, Mayee F. Chen, Suchin Gururangan, Mitchell Wortsman, Alon Albalak, Yonatan Bitton, Marianna Nezhurina, Amro Abbas, Cheng-Yu Hsieh, Dhruba Ghosh, Josh Gardner 0001, Maciej Kilian, Hanlin Zhang 0002, Rulin Shao, Sarah M. Pratt, Sunny Sanyal, Gabriel Ilharco, Giannis Daras, Kalyani Marathe, Aaron Gokaslan, Jieyu Zhang 0001, Khyathi Raghavi Chandu, Igor Vasiljevic, Sham M. Kakade, Shuran Song, Sujay Sanghavi, Fartash Faghri, Sewoong Oh, Luke Zettlemoyer, Kyle Lo, Alaaeldin El-Nouby, Hadi Pouransari, Alexander Toshev, Stephanie Wang, Dirk Groeneveld, Luca Soldaini, Pang Wei Koh, Jenia Jitsev, Thomas Kollar, Alexandros G. Dimakis, Yair Carmon, Achal Dave, Ludwig Schmidt, Vaishaal Shankar
NeurIPS4
2023 DoG is SGD's Best Friend: A Parameter-Free Dynamic Step Size Schedule
abstract
We propose a tuning-free dynamic SGD step size formula, which we call Distance over Gradients (DoG). The DoG step sizes depend on simple empirical quantities (distance from the initial point and norms of gradients) and have no ``learning rate'' parameter. Theoretically, we show that, for stochastic convex optimization, a slight variation of the DoG formula enjoys strong, high-probability parameter-free convergence guarantees and iterate movement bounds. Empirically, we consider a broad range of vision and language transfer learning tasks, and show that DoG's performance is close to that of SGD with tuned learning rate. We also propose a per-layer variant of DoG that generally outperforms tuned SGD, approaching the performance of tuned Adam. A PyTorch implementation of our algorithms is available at https://github.com/formll/dog.
Maor Ivgi, Oliver Hinder, Yair Carmon
ICML1
2023 Efficient Long-Text Understanding with Short-Text Models
abstract
Abstract Transformer-based pretrained language models (LMs) are ubiquitous across natural language understanding, but cannot be applied to long sequences such as stories, scientific articles, and long documents due to their quadratic complexity. While a myriad of efficient transformer variants have been proposed, they are typically based on custom implementations that require expensive pretraining from scratch. In this work, we propose SLED: SLiding-Encoder and Decoder, a simple approach for processing long sequences that re-uses and leverages battle-tested short-text pretrained LMs. Specifically, we partition the input into overlapping chunks, encode each with a short-text LM encoder and use the pretrained decoder to fuse information across chunks (fusion-in-decoder). We illustrate through controlled experiments that SLED offers a viable strategy for long text understanding and evaluate our approach on SCROLLS, a benchmark with seven datasets across a wide range of language understanding tasks. We find that SLED is competitive with specialized models that are up to 50x larger and require a dedicated and expensive pretraining step.
Maor Ivgi, Uri Shaham 0002, Jonathan Berant
Trans. Assoc. Comput. Linguistics1
2022 SCROLLS: Standardized CompaRison Over Long Language Sequences
abstract
Uri Shaham, Elad Segal, Maor Ivgi, Avia Efrat, Ori Yoran, Adi Haviv, Ankit Gupta, Wenhan Xiong, Mor Geva, Jonathan Berant, Omer Levy. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022.
Uri Shaham 0002, Elad Segal, Maor Ivgi, Avia Efrat, Ori Yoran, Adi Haviv, Ankit Gupta 0001, Wenhan Xiong, Mor Geva, Jonathan Berant, Omer Levy
EMNLP3
2021 Achieving Model Robustness through Discrete Adversarial Training
abstract
Discrete adversarial attacks are symbolic perturbations to a language input that preserve the output label but lead to a prediction error.While such attacks have been extensively explored for the purpose of evaluating model robustness, their utility for improving robustness has been limited to offline augmentation only.Concretely, given a trained model, attacks are used to generate perturbed (adversarial) examples, and the model is re-trained exactly once.In this work, we address this gap and leverage discrete attacks for online augmentation, where adversarial examples are generated at every training step, adapting to the changing nature of the model.We propose (i) a new discrete attack, based on best-first search, and (ii) random sampling attacks that unlike prior work are not based on expensive search-based procedures.Surprisingly, we find that random sampling leads to impressive gains in robustness, outperforming the commonly-used offline augmentation, while leading to a speedup at training time of ∼10x.Furthermore, online augmentation with search-based attacks justifies the higher training cost, significantly improving robustness on three datasets.Last, we show that our new attack substantially improves robustness compared to prior methods.
Maor Ivgi, Jonathan Berant
EMNLP (1)1
2021 Scene Graph tO Image Generation with Contextualized Object Layout Refinement
abstract
Generating images from scene graphs is a challenging task that attracted substantial interest recently. Prior works have approached this task by generating an intermediate layout description of the target image. However, the representation of each object in the layout was generated independently, which resulted in high overlap, low coverage, and an overall blurry layout. We propose a novel method that alleviates these issues by generating the entire layout description gradually to improve inter-object dependency. We empirically show on the COCO-STUFF dataset that our approach improves the quality of both the intermediate layout and the final image. Our approach improves the layout coverage by almost 20 points, and drops object overlap to negligible amounts. Our code is available at github.com/yanivbenny/COLoR.
Maor Ivgi, Yaniv Benny, Avichai Ben-David, Jonathan Berant, Lior Wolf
ICIP1