Segev Shlomov

dblp:202/9370 · DBLP profile ↗
← Back
9ranked-venue papers
3as first author
6since 2021 · last 2026
0000-0003-1216-8284ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 3 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 5 since 2021
YearPublicationVenuePosition
2026 From Benchmarks to Business Impact: Deploying IBM Generalist Agent in Enterprise Production
abstract
Agents are rapidly advancing in automating digital work, but enterprises face a harder challenge: moving beyond prototypes to deployed systems that deliver measurable business value. This path is complicated by fragmented frameworks, slow development, and the absence of standardized evaluation practices. Generalist agents have emerged as a promising direction, excelling on academic benchmarks and offering flexibility across tasks, applications, and modalities. Yet, evidence of their use in enterprise settings remains limited. This paper reports IBM’s experience developing and piloting the Computer Using Generalist Agent (CUGA). CUGA adopts a hierarchical planner--executor architecture with strong analytical foundations, achieving state-of-the-art performance on AppWorld and WebArena. Beyond benchmarks, it was evaluated in a Business-Process-Outsourcing talent acquisition pilot, addressing enterprise requirements for scalability, auditability, safety, and governance. In preliminary evaluations, CUGA approached the accuracy of specialized agents while suggesting reductions in development time and cost. We provide early evidence that generalist agents can operate at enterprise scale, distill key technical and organizational lessons, and outline requirements for transitioning research-grade architectures like CUGA into enterprise-ready systems.
Segev Shlomov, Alon Oved, Sami Marreed, Ido Levy, Offer Akrabi, Avi Yaeli, Lukasz Strak, Elizabeth Koumpan, Yinon Goldshtein, Eilam Shapira, Nir Mashkif, Asaf Adi
AAAI1
2025 SNAP: Semantic Stories for Next Activity Prediction
abstract
Predicting the next activity in an ongoing process is one of the most common tasks in the business process management (BPM) domain. It allows businesses to optimize resource allocation, enhance operational efficiency, and aid both in risk mitigation and strategic decision-making. Existing state-of-the-art AI models for BPM do not fully capitalize on available semantic information within process event logs. As current advanced AI-BPM systems provide semantically richer textual data, the need for new adequate models grows. To address this gap, we develop SNAP—a novel system that utilizes LLMs by constructing narratives and semantic contextual stories for historical event logs, which are then used to generate precise and actionable predictions for the ongoing process. SNAP was evaluated on six benchmark datasets, where it demonstrated significant performance improvements over eleven SOTA models, particularly on datasets with high levels of semantic content. This work showcases the potential of integrating LLMs in BPM and outlines a clear path toward future deployment, emphasizing the relevance and innovation of our approach within the broader AI application landscape.
Alon Oved, Segev Shlomov, Sergey Zeltyn, Nir Mashkif, Avi Yaeli
AAAI2
2025 From Grounding to Planning: Benchmarking Bottlenecks in Web Agents
abstract
General web-based agents are increasingly essential for interacting with complex web environments, yet their performance in real-world web applications remains poor, yielding extremely low accuracy even with state-of-the-art frontier models. We observe that these agents can be decomposed into two primary components: Planning and Grounding. Yet, most existing research treats these agents as black boxes, focusing on end-to-end evaluations which hinder meaningful improvements. We sharpen the distinction between the planning and grounding components and conduct a novel analysis by refining experiments on the Mind2Web dataset. Our work proposes a new benchmark for each of the two components, identifying the bottlenecks and pain points that limit agent performance. Contrary to prevalent assumptions, our findings suggest that grounding is not a significant bottleneck and can be effectively addressed with current techniques. Instead, the primary challenge lies in the planning component, which is the main source of performance degradation. Through this analysis, we offer new insights and demonstrate practical suggestions for improving the capabilities of web agents, paving the way for more reliable agents.
Segev Shlomov, Ben Wiesel, Aviad Sela, Ido Levy, Liane Galanti, Roy Abitbol
ECAI1
2024 Mimicking the Maestro: Exploring the Efficacy of a Virtual AI Teacher in Fine Motor Skill Acquisition
abstract
Motor skills, especially fine motor skills like handwriting, play an essential role in academic pursuits and everyday life. Traditional methods to teach these skills, although effective, can be time-consuming and inconsistent. With the rise of advanced technologies like robotics and artificial intelligence, there is increasing interest in automating such teaching processes. In this study, we examine the potential of a virtual AI teacher in emulating the techniques of human educators for motor skill acquisition. We introduce an AI teacher model that captures the distinct characteristics of human instructors. Using a reinforcement learning environment tailored to mimic teacher-learner interactions, we tested our AI model against four guiding hypotheses, emphasizing improved learner performance, enhanced rate of skill acquisition, and reduced variability in learning outcomes. Our findings, validated on synthetic learners, revealed significant improvements across all tested hypotheses. Notably, our model showcased robustness across different learners and settings and demonstrated adaptability to handwriting. This research underscores the potential of integrating Imitation and Reinforcement Learning models with robotics in revolutionizing the teaching of critical motor skills.
Hadar Mulian, Segev Shlomov, Lior Limonad, Alessia Noccaro, Silvia Buscaglione
AAAI2
2024 Towards a Resilient Intelligent Automation System
Segev Shlomov, Sami Marreed, Avi Yaeli
IJCAI1
2021 We've had this conversation before: A Novel Approach to Measuring Dialog Similarity
abstract
Dialog is a core building block of human natural language interactions.It contains multiparty utterances used to convey information from one party to another in a dynamic and evolving manner.The ability to compare dialogs is beneficial in many real world use cases, such as conversation analytics for contact center calls and virtual agent design.We propose a novel adaptation of the edit distance metric to the scenario of dialog similarity.Our approach takes into account various conversation aspects such as utterance semantics, conversation flow, and the participants.We evaluate this new approach and compare it to existing document similarity measures on two publicly available datasets.The results demonstrate that our method outperforms the other approaches in capturing dialog flow, and is better aligned with the human perception of conversation similarity.
Ofer Lavi, Ella Rabinovich, Segev Shlomov, David Boaz, Inbal Ronen, Ateret Anaby-Tavor
EMNLP (1)3
2020 Do Not Have Enough Data? Deep Learning to the Rescue!
abstract
Based on recent advances in natural language modeling and those in text generation capabilities, we propose a novel data augmentation method for text classification tasks. We use a powerful pre-trained neural network model to artificially synthesize new labeled data for supervised learning. We mainly focus on cases with scarce labeled data. Our method, referred to as language-model-based data augmentation (LAMBADA), involves fine-tuning a state-of-the-art language generator to a specific task through an initial training phase on the existing (usually small) labeled data. Using the fine-tuned model and given a class label, new sentences for the class are generated. Our process then filters these new sentences by using a classifier trained on the original data. In a series of experiments, we show that LAMBADA improves classifiers' performance on a variety of datasets. Moreover, LAMBADA significantly improves upon the state-of-the-art techniques for data augmentation, specifically those applicable to text classification tasks with little data.
Ateret Anaby-Tavor, Boaz Carmeli, Esther Goldbraich, Amir Kantor, George Kour, Segev Shlomov, Naama Tepper, Naama Zwerdling
AAAI6
2019 Deep Dominance - How to Properly Compare Deep Neural Models
abstract
Comparing between Deep Neural Network (DNN) models based on their performance on unseen data is crucial for the progress of the NLP field.However, these models have a large number of hyper-parameters and, being non-convex, their convergence point depends on the random values chosen at initialization and during training.Proper DNN comparison hence requires a comparison between their empirical score distributions on unseen data, rather than between single evaluation scores as is standard for more simple, convex models.In this paper, we propose to adapt to this problem a recently proposed test for the Almost Stochastic Dominance relation between two distributions.We define the criteria for a high quality comparison method between DNNs, and show, both theoretically and through analysis of extensive experimental results with leading DNN models for sequence tagging tasks, that the proposed test meets all criteria while previously proposed methods fail to do so.We hope the test we propose here will set a new working practice in the NLP community.1
Rotem Dror, Segev Shlomov, Roi Reichart
ACL (1)2
2018 The Hitchhiker's Guide to Testing Statistical Significance in Natural Language Processing
abstract
Statistical significance testing is a standard statistical tool designed to ensure that experimental results are not coincidental.In this opinion/theoretical paper we discuss the role of statistical significance testing in Natural Language Processing (NLP) research.We establish the fundamental concepts of significance testing and discuss the specific aspects of NLP tasks, experimental setups and evaluation measures that affect the choice of significance tests in NLP research.Based on this discussion, we propose a simple practical protocol for statistical significance test selection in NLP setups and accompany this protocol with a brief survey of the most relevant tests.We then survey recent empirical papers published in ACL and TACL during 2017 and show that while our community assigns great value to experimental results, statistical significance testing is often ignored or misused.We conclude with a brief discussion of open issues that should be properly addressed so that this important tool can be applied in NLP research in a statistically sound manner 1 .
Rotem Dror, Gili Baumer, Segev Shlomov, Roi Reichart
ACL (1)3