Robert Moskovitch

dblp:34/678 · DBLP profile ↗
← Back
65ranked-venue papers
21as first author
25since 2021 · last 2025
0000-0002-2138-5080ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 23 · 8 first-author · 6 since 2021Artificial intelligence and machine learning · 18 · 3 first-author · 10 since 2021Databases, data management, data science and information retrieval · 16 · 7 first-author · 7 since 2021Security and privacy · 8 · 5 first-author · 1 since 2021Systems, architecture and hardware · 3 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Temporal ensemble of multiple patterns' instances for continuous prediction of events
abstract
Abstract In real-life data of various domains, such as traffic, meteorology, or healthcare data, events may have varying durations. Moreover, heterogeneous multivariate temporal data may consist of varying samplings, including regular sampling in different frequencies or irregular, as well as events data of different types, having fixed or varying duration. We propose to uniformly represent heterogeneous multivariate temporal data using symbolic time-intervals, from which a model that predicts an occurrence of events early can be learned. We introduce a novel use of time-interval-related patterns (TIRPs), in which patterns that end with an event of interest can be used to continuously estimate the event’s occurrence probability in real-time. Recently, we introduced a model that allows continuous prediction of the completion of a pattern, which is extended in this work, to also predict the expected completion time. This work focuses on predicting the probability and time occurrence of an event based on multiple different instances of patterns that end with the event, for which we propose and evaluate aggregation functions. A rigorous evaluation was conducted on four real-life datasets to assess the effectiveness of the proposed model and the aggregation functions. The proposed model performed better than the baseline models (ResNet, LSTM-FCN, ROCKET, and XGBoost) for all datasets.
Nevo Itzhak, Szymon Jaroszewicz, Robert Moskovitch
Mach. Learn.3
2025 Improving DNNs for time-series classification using state and gradient abstraction-based preprocessing
Nevo Itzhak, Shahar Tal, Hadas Cohen, Osher Daniel, Roze Kopylov, Robert Moskovitch
Neural Comput. Appl.6
2025 Time-intervals-related pattern selection for continuous event prediction
Nevo Itzhak, Szymon Jaroszewicz, Robert Moskovitch
Pattern Recognit.3
2025 STORM: A MapReduce Framework for Symbolic Time Intervals Series Classification
abstract
Symbolic Time Intervals (STIs) represent events having a non-zero time duration, which are common in various application domains. In this article, we focus on the challenge of STIs series classification (STIC). While in the related problem of time series classification (TSC) Rocket is well-known for its exceptionally fast runtime while achieving accuracy comparable to state-of-the-art, it has only recently been studied in the field of STIC. However, since Rocket as well as its enhanced variants for TSC (e.g., MiniRocket and MultiRocket) solely rely on global features, they might not always fit best for the classification of thousands of time-units long STI series out-of-the-box, which are rather common in STIC. We introduce STORM—a novel, generic MapReduce framework for STIC, which (1) converts raw input STIs series into multivariate time series (MTS) representation; (2) partitions the converted MTS into fixed-sized blocks, each transformed independently into a uniform latent space via a common, desired Rocket variant used as a base transformation in STORM; and (3) performs sequence classification of the blocks’ transformed feature vectors via a deep, lightweight, bidirectional LSTM network. The evaluation demonstrates that STORM significantly improves accuracy over eight state-of-the-art methods for STIC either when applied with MiniRocket and MultiRocket as base transformations, as well as over the baselines of applying the respective Rocket variants directly to the converted MTS representation, that is, while also reporting overall comparable training times, on a benchmark of eight real-world STIC datasets including both extremely long and short STIs series.
Omer David Harel, Robert Moskovitch
ACM Trans. Knowl. Discov. Data2
2024 Introduction to the Special Track on Artificial Intelligence and COVID-19 (Abstract Reprint)
abstract
The human race is facing one of the most meaningful public health emergencies in the modern era caused by the COVID-19 pandemic. This pandemic introduced various challenges, from lock-downs with significant economic costs to fundamentally altering the way of life for many people around the world. The battle to understand and control the virus is still at its early stages yet meaningful insights have already been made. The uncertainty of why some patients are infected and experience severe symptoms, while others are infected but asymptomatic, and others are not infected at all, makes managing this pandemic very challenging. Furthermore, the development of treatments and vaccines relies on knowledge generated from an ever evolving and expanding information space. Given the availability of digital data in the modern era, artificial intelligence (AI) is a meaningful tool for addressing the various challenges introduced by this unexpected pandemic. Some of the challenges include: outbreak prediction, risk modeling including infection and symptom development, testing strategy optimization, drug development, treatment repurposing, vaccine development, and others.
Martin Michalowski, Robert Moskovitch, Nitesh V. Chawla
AAAI2
2024 Early Multiple Temporal Patterns Based Event Prediction in Heterogeneous Multivariate Temporal Data
abstract
Predicting an event of interest based on heterogeneous multivariate temporal data is challenging but desirable as it allows the utilization of all types of temporal variables. In various domains, symbolic time intervals (STIs) can be used to represent real-life events that vary in duration, such as the period a traffic light remains green, or the time a patient undergoes treatment or is on medication. Further, heterogeneous multivariate temporal data may be composed of STIs along with event-driven or continuous temporal variables, such as traffic collisions or blood test values. Temporal abstraction can be used to uniformly represent heterogeneous multivariate temporal variables with STIs, from which frequent time intervals related patterns (TIRPs) can be discovered. We extend earlier work on continuous completion prediction of a single TIRP that ends with an event of interest, introducing a continuous prediction method based on multiple different instances of multiple TIRPs that end with the event of interest, for which we propose and evaluate several weighted aggregation functions. The proposed method overall performed better on real-life, medical, and non-medical datasets, than the use of a single TIRP, and in comparison to the baseline models (XGBoost, ResNet, LSTM-FCN, and ROCKET).
Nevo Itzhak, Szymon Jaroszewicz, Robert Moskovitch
SDM3
2024 Event prediction by estimating continuously the completion of a single temporal pattern's instances
Nevo Itzhak, Szymon Jaroszewicz, Robert Moskovitch
J. Biomed. Informatics3
2024 Sleep apnea test prediction based on Electronic Health Records
Lama Abu Tahoun, Amit Shay Green, Tal Patalon, Yaron Dagan, Robert Moskovitch
J. Biomed. Informatics5
2023 Continuously Predicting the Completion of a Time Intervals Related Pattern
Nevo Itzhak, Szymon Jaroszewicz, Robert Moskovitch
PAKDD (1)3
2023 Prediction of acute hypertensive episodes in critically ill patients
Nevo Itzhak, Itai M. Pessach, Robert Moskovitch
Artif. Intell. Medicine3
2023 TIRPClo: efficient and complete mining of time intervals-related patterns
Omer David Harel, Robert Moskovitch
Data Min. Knowl. Discov.2
2023 Predictive temporal patterns discovery
Nofar Sarafian Ben Ari, Robert Moskovitch
Expert Syst. Appl.2
2023 INSTINCT: Inception-based Symbolic Time Intervals series classification
Omer David Harel, Robert Moskovitch
Inf. Sci.2
2023 Introduction to the Special Track on Artificial Intelligence and COVID-19
abstract
The human race is facing one of the most meaningful public health emergencies in the modern era caused by the COVID-19 pandemic. This pandemic introduced various challenges, from lock-downs with significant economic costs to fundamentally altering the way of life for many people around the world. The battle to understand and control the virus is still at its early stages yet meaningful insights have already been made. The uncertainty of why some patients are infected and experience severe symptoms, while others are infected but asymptomatic, and others are not infected at all, makes managing this pandemic very challenging. Furthermore, the development of treatments and vaccines relies on knowledge generated from an ever evolving and expanding information space. Given the availability of digital data in the modern era, artificial intelligence (AI) is a meaningful tool for addressing the various challenges introduced by this unexpected pandemic. Some of the challenges include: outbreak prediction, risk modeling including infection and symptom development, testing strategy optimization, drug development, treatment repurposing, vaccine development, and others.
Martin Michalowski, Robert Moskovitch, Nitesh V. Chawla
J. Artif. Intell. Res.2
2023 Continuous prediction of a time intervals-related pattern's completion
Nevo Itzhak, Szymon Jaroszewicz, Robert Moskovitch
Knowl. Inf. Syst.3
2022 Classification of Univariate Time Series via Temporal Abstraction and Deep Learning
abstract
Many time series classification algorithms have been proposed, including deep neural networks based, which so far focused mainly on improving model architectures rather than on data pre-processing. Generalization is crucial in time series classification and it can be achieved by abstracting the data. Data abstraction may also be useful to avoid handling challenges with error measurements, missing values, and irregular sampling. We propose transforming the raw time series into a symbolic time series representation, using a method known as temporal abstraction, before feeding it to the deep neural networks. This transformation can greatly enhance generalization and may potentially improve classification performance. In particular, we investigate the effectiveness of temporal abstraction when combined with convolution-based sequence models or recurrent neural networks. The methods were evaluated on 128 univariate datasets. Our evaluation shows that even when using equal frequency discretization, a relatively simple method, outperforms most state-of-the-art deep neural networks’ performance for univariate time series classification when fed by raw time series.
Nevo Itzhak, Shahar Tal, Hadas Cohen, Osher Daniel, Roze Kopylov, Robert Moskovitch
IEEE Big Data6
2022 All-cause mortality prediction in T2D patients with iTirps
Pavel Novitski, Cheli Melzer Cohen, Avraham Karasik, Varda Shalev, Gabriel Hodik, Robert Moskovitch
Artif. Intell. Medicine6
2022 Temporal patterns selection for All-Cause Mortality prediction in T2D with ANNs
Pavel Novitski, Cheli Melzer Cohen, Avraham Karasik, Gabriel Hodik, Robert Moskovitch
J. Biomed. Informatics5
2022 Visualization of frequent temporal patterns in single or two populations
abstract
Temporal knowledge discovery in clinical problems, is crucial to investigate problems in the data science era. Meaningful progress has been made computationally in the discovery of frequent temporal patterns, which may store potentially meaningful knowledge. However, for temporal knowledge discovery and acquisition, effective visualization is essential and still stores much room for contributions. While visualization of frequent temporal patterns was relatively under researched, it stores meaningful opportunities in facilitating usable ways to assist domain experts, or researchers, in exploring and acquiring temporal knowledge. In this paper, a novel approach for the visualization of an enumeration tree of frequent temporal patterns is introduced for, whether mined from a single population, or for the comparison of patterns that were discovered in two separate populations. While this approach is relevant to any sequence-based patterns, we demonstrate its use on the most complex scenario of time intervals related patterns (TIRPs). The interface enables users to browse an enumeration tree of frequent patterns, or search for specific patterns, as well as discover the most discriminating TIRPs among two populations. For that a novel visualization of the temporal patterns is introduced using a bubble chart, in which each bubble represents a temporal pattern, and the chart axes represent the various metrics of the patterns, such as their frequency, reoccurrence, and more, which provides a fast overview of the patterns as a whole, as well as access specific ones. We present a comprehensive and rigorous user study on two real-life datasets, demonstrating the usability advantages of the novel approaches.
Guy Shitrit, Noam Tractinsky, Robert Moskovitch
J. Biomed. Informatics3
2021 Complete Closed Time Intervals-Related Patterns Mining
abstract
Using temporal abstraction, various forms of sampled multivariate temporal data can be transformed into a uniform representation of symbolic time intervals, from which Time Intervals Related Patterns (TIRPs) can be then discovered. Hence, mining TIRPs from symbolic time intervals offers a comprehensive framework for heterogeneous multivariate temporal data analysis. While the field of time intervals mining has gained a growing interest in recent decades, frequent closed TIRPs mining was not investigated in its full complexity. Mining frequent closed TIRPs is highly effective due to the discovery of a compact set of frequent TIRPs, which contains the complete information of all the frequent TIRPs. However, as we demonstrate in this paper, the recent advancements made in closed TIRPs discovery are incomplete, due to the discovery of only the first instances of the TIRPs within each STIs series in the database. In this paper we introduce the TIRPClo algorithm – for complete and efficient mining of frequent closed TIRPs. The algorithm utilizes a memory-efficient index and a novel method for data projection, due to which it is the first algorithm to guarantee a complete discovery of frequent closed TIRPs. In addition, a rigorous runtime comparison of TIRPClo to state-of-the-art methods is performed, demonstrating a significant speed-up on various real-world datasets.
Omer David Harel, Robert Moskovitch
AAAI2
2021 Temporal pattern-based malicious activity detection in SCADA systems
Amit Shlomo, Meir Kalech, Robert Moskovitch
Comput. Secur.3
2021 Pkg2Vec: Hierarchical package embedding for code authorship attribution
Roni Mateless, Oren Tsur, Robert Moskovitch
Future Gener. Comput. Syst.3
2021 Outcomes prediction in longitudinal data: Study designs evaluation, use case in ICU acquired sepsis
Maya Schvetz, Lior Fuchs, Victor Novack, Robert Moskovitch
J. Biomed. Informatics4
2021 THAAD: Efficient matching queries under temporal abstraction for anomaly detection
Roni Mateless, Michael Segal 0001, Robert Moskovitch
Perform. Evaluation3
2021 IPvest: Clustering the IP Traffic of Network Entities Hidden Behind a Single IP Address Using Machine Learning
abstract
IP Networks serve a variety of connected network entities (NEs) such as personal computers, servers, mobile devices, virtual machines, hosted containers, etc. The growth in the number of NEs and technical considerations has led to a reality where a single IP address is used by multiple NEs. A typical example is a home router using Network Address Translation (NAT). In organizations and cloud environments, a single IP can be used by multiple virtual machines or containers running on a single device. Discovering the number of NEs served by an IP address and clustering their traffic correctly is of value in many use cases for security, lawful interception, asset management, and other purposes. In this paper, we introduce IPvest, a system that incorporates unsupervised and supervised learning algorithms based on various features for counting and clustering network traffic of NEs masqueraded by a single IP. The features are based on the characteristics of operating systems (OSs), NAT behavior, and users' habits. Our model is evaluated on real-world datasets including Windows, Linux-based, Android, and iOS-based devices, containers, virtual machines, and load-balancers. We show that IPvest can count the number of NEs and cluster their traffic with high precision, even for containers running on a single device and servers behind a load-balancer.
Roni Mateless, Haim Zlatokrilov, Liran Orevi, Michael Segal 0001, Robert Moskovitch
IEEE Trans. Netw. Serv. Manag.5
2020 Falls Prediction in Care Homes Using Mobile App Data Collection
Ofir Dvir, Paul Wolfson, Laurence B. Lovat, Robert Moskovitch
AIME4
2020 Acute Hypertensive Episodes Prediction
Nevo Itzhak, Aditya Nagori, Edo Lior, Maya Schvetz, Rakesh Lodha, Tavpritesh Sethi, Robert Moskovitch
AIME7
2020 All-Cause Mortality Prediction in T2D Patients
Pavel Novitski, Cheli Melzer Cohen, Avraham Karasik, Varda Shalev, Gabriel Hodik, Robert Moskovitch
AIME6
2020 Decompiled APK based malicious code classification
Roni Mateless, Daniel Rejabek, Oded Margalit, Robert Moskovitch
Future Gener. Comput. Syst.4
2019 Temporal biomedical data analytics
Robert Moskovitch, Yuval Shahar, Fei Wang 0001, George Hripcsak
J. Biomed. Informatics1
2017 Inter-labeler and intra-labeler variability of condition severity classification models using active and passive learning methods
Nir Nissim, Yuval Shahar, Yuval Elovici, George Hripcsak, Robert Moskovitch
Artif. Intell. Medicine5
2017 JASIST special issue on biomedical information retrieval
Robert Moskovitch, Fei Wang 0001, Jian Pei 0001, Carol Friedman
J. Assoc. Inf. Sci. Technol.1
2017 Procedure prediction from symbolic Electronic Health Records via time intervals analytics
Robert Moskovitch, Fernanda Polubriaginof, Aviram Weiss, Patrick B. Ryan, Nicholas P. Tatonetti
J. Biomed. Informatics1
2017 Consistent discovery of frequent interval-based temporal patterns in chronic patients' data
Alexander Shknevsky, Yuval Shahar, Robert Moskovitch
J. Biomed. Informatics3
2017 Prognosis of Clinical Outcomes with Temporal Patterns and Experiences with One Class Feature Selection
abstract
Accurate prognosis of outcome events, such as clinical procedures or disease diagnosis, is central in medicine. The emergence of longitudinal clinical data, like the Electronic Health Records (EHR), represents an opportunity to develop automated methods for predicting patient outcomes. However, these data are highly dimensional and very sparse, complicating the application of predictive modeling techniques. Further, their temporal nature is not fully exploited by current methods, and temporal abstraction was recently used which results in symbolic time intervals representation. We present Maitreya, a framework for the prediction of outcome events that leverages these symbolic time intervals. Using Maitreya, learn predictive models based on the temporal patterns in the clinical records that are prognostic markers and use these markers to train predictive models for eight clinical procedures. In order to decrease the number of patterns that are used as features, we propose the use of three one class feature selection methods. We evaluate the performance of Maitreya under several parameter settings, including the one-class feature selection, and compare our results to that of atemporal approaches. In general, we found that the use of temporal patterns outperformed the atemporal methods, when representing the number of pattern occurrences.
Robert Moskovitch, Hyunmi Choi, George Hripcsak, Nicholas P. Tatonetti
IEEE ACM Trans. Comput. Biol. Bioinform.1
2016 Temporal data analytics
Robert Moskovitch, Fei Wang 0001, Yuval Shahar, George Hripcsak
J. Biomed. Informatics1
2016 Improving condition severity classification with an efficient active learning based framework
Nir Nissim, Mary Regina Boland, Nicholas P. Tatonetti, Yuval Elovici, George Hripcsak, Yuval Shahar, Robert Moskovitch
J. Biomed. Informatics7
2016 ALDROID: efficient update of Android anti-virus software using designated active learning methods
Nir Nissim, Robert Moskovitch, Oren Bar-Ad, Lior Rokach, Yuval Elovici
Knowl. Inf. Syst.2
2015 An Active Learning Framework for Efficient Condition Severity Classification
Nir Nissim, Mary Regina Boland, Robert Moskovitch, Nicholas P. Tatonetti, Yuval Elovici, Yuval Shahar, George Hripcsak
AIME3
2015 Outcomes Prediction via Time Intervals Related Patterns
abstract
The increasing availability of multivariate temporal data in many domains, such as biomedical, security and more, provides exceptional opportunities for temporal knowledge discovery, classification and prediction, but also challenges. Temporal variables are often sparse and in many domains, such as in biomedical data, they have huge number of variables. In recent decades in the biomedical domain events, such as conditions, drugs and procedures, are stored as time intervals, which enables to discover Time Intervals Related Patterns (TIRPs) and use for classification or prediction. In this study we present a framework for outcome events prediction, called Maitreya, which includes an algorithm for TIRPs discovery called KarmaLegoD, designed to handle huge number of symbols. Three indexing strategies for pairs of symbolic time intervals are proposed and compared, showing that the use of FullyHashed indexing is only slightly slower but consumes minimal memory. We evaluated Maitreya on eight real datasets for the prediction of clinical procedures as outcome events. The use of TIRPs outperform the use of symbols, especially with horizontal support (number of instances) as TIRPs feature representation.
Robert Moskovitch, Colin G. Walsh, Fei Wang 0001, George Hripcsak, Nicholas P. Tatonetti
ICDM1
2015 Classification-driven temporal discretization of multivariate time series
Robert Moskovitch, Yuval Shahar
Data Min. Knowl. Discov.1
2015 Fast time intervals mining using the transitivity of temporal relations
Robert Moskovitch, Yuval Shahar
Knowl. Inf. Syst.1
2015 Classification of multivariate time series via temporal abstraction and time intervals mining
Robert Moskovitch, Yuval Shahar
Knowl. Inf. Syst.1
2014 Novel active learning methods for enhanced PC malware detection in windows OS
Nir Nissim, Robert Moskovitch, Lior Rokach, Yuval Elovici
Expert Syst. Appl.2
2012 User identity verification via mouse dynamics
Clint Feher, Yuval Elovici, Robert Moskovitch, Lior Rokach, Alon Schclar
Inf. Sci.3
2012 Detecting unknown computer worm activity via support vector machines and active learning
Nir Nissim, Robert Moskovitch, Lior Rokach, Yuval Elovici
Pattern Anal. Appl.2
2011 Monitoring, analysis, and filtering system for purifying network traffic of known and unknown malicious content
abstract
Abstract The early detection, alert and response (eDare) framework is presented in this paper. The goal of this framework is to address the risks stemming from malicious software propagating via networks operated by Internet/network service providers (ISP/NSP). To achieve this goal, eDare employs network‐based traffic scanning appliances that enable sanitation of Internet traffic of known malware. Remaining traffic is extracted and various types of algorithms are invoked in an attempt to detect instances of previously un‐encountered malware and to generate a unique and simple byte‐string signature for such malware. That signature is immediately uploaded to the aforementioned network traffic scanners. To augment judgments of the algorithms, human experts are consulted for assistance in classifying files suspected of being malware about which the automatic detection algorithms are not sufficiently decisive. Finally, collaborative feedback and tips from end‐users are meshed into the identification process. This makes tackling of suspect files, whose impact can be assessed on a large, distributed scale, possible. The system incorporates static and behavioral analysis of malware and novel automatic signature generation algorithm. eDare was implemented and tested using an evaluation environment especially developed for that purpose. The results suggest that eDare can detect and remove unknown malware effectively. Copyright © 2010 John Wiley & Sons, Ltd.
Asaf Shabtai, Dennis Potashnik, Yuval Fledel, Robert Moskovitch, Yuval Elovici
Secur. Commun. Networks4
2009 Medical Temporal-Knowledge Discovery via Temporal Abstraction
Robert Moskovitch, Yuval Shahar
AMIA1
2009 Identity theft, computers and behavioral biometrics
abstract
The increase of online services, such as eBanks, WebMails, in which users are verified by a username and password, is increasingly exploited by identity theft procedures. Identity Theft is a fraud, in which someone pretends to be someone else is order to steal money or get other benefits. To overcome the problem of identity Theft an additional security layer is required. Within the last decades the option of verifying users based on their keystroke dynamics was proposed during login verification. Thus, the imposter has to be able to type in a similar way to the real user in addition to having the username and password. However, verifying users upon login is not enough, since a logged station/mobile is vulnerable for imposters when the user leaves her machine. Thus, verifying users continuously based on their activities is required. Within the last decade there is a growing interest and use of biometrics tools, however, these are often costly and require additional hardware. Behavioral biometrics, in which users are verified, based on their keyboard and mouse activities, present potentially a good solution. In this paper we discuss the problem of identity theft and propose behavioral biometrics as a solution. We survey existing studies and list the challenges and propose solutions.
Robert Moskovitch, Clint Feher, Arik Messerman, Niklas Kirschnick, Tarik Mustafic, Seyit Ahmet Çamtepe, Bernhard Löhlein, Ulrich Heister, Sebastian Möller 0001, Lior Rokach, Yuval Elovici
ISI1
2009 Detection of malicious code by applying machine learning classifiers on static features: A state-of-the-art survey
Asaf Shabtai, Robert Moskovitch, Yuval Elovici, Chanan Glezer
Inf. Secur. Tech. Rep.2
2009 Vaidurya: A multiple-ontology, concept-based, context-sensitive clinical-guideline search engine
Robert Moskovitch, Yuval Shahar
J. Biomed. Informatics1
2009 Using artificial neural networks to detect unknown computer worms
Dima Stopel, Robert Moskovitch, Zvi Boger, Yuval Shahar, Yuval Elovici
Neural Comput. Appl.2
2008 Active learning to improve the detection of unknown computer worms activity
Robert Moskovitch, Nir Nissim, Roman Englert, Yuval Elovici
FUSION1
2008 Unknown malcode detection - A chronological evaluation
abstract
Signature-based anti-viruses are very accurate, but are limited in detecting new malicious code. Dozens of new malicious codes are created every day, and the rate is expected to increase in coming years. To extend the generalization to detect unknown malicious code, heuristic methods are used; however, these are not successful enough. Recently, classification algorithms were used successfully for the detection of unknown malicious code. We earlier investigated the optimized conditions in which highest-level accuracy is achieved, in terms of the percentage of malicious files. In this paper we describe the methodology of detection of malicious code based on static analysis and a chronological evaluation, in which a classifier is trained on files till year k and tested on the following years. The evaluation was performed in two setups, in which the percentage of the malicious files in the training set was 50% or 16%. Using 16% malicious files in the training set showed a clear trend, in which the performance improves as the training set is more updated.
Robert Moskovitch, Clint Feher, Yuval Elovici
ISI1
2008 Unknown malcode detection via text categorization and the imbalance problem
abstract
Todaypsilas signature-based anti-viruses are very accurate, but are limited in detecting new malicious code. Currently, dozens of new malicious codes are created every day, and this number is expected to increase in the coming years. Recently, classification algorithms were used successfully for the detection of unknown malicious code. These studies used a test collection with a limited size where the same malicious-benign-file ratio in both the training and test sets, which does not reflect real-life conditions. In this paper we present a methodology for the detection of unknown malicious code, based on text categorization concepts. We performed an extensive evaluation using a test collection that contains more than 30,000 malicious and benign files, in which we investigated the imbalance problem. In real-life scenarios, the malicious file content is expected to be low, about 10% of the total files. For practical purposes, it is unclear as to what the corresponding percentage in the training set should be. Our results indicate that greater than 95% accuracy can be achieved through the use of a training set that contains below 20% malicious file content.
Robert Moskovitch, Dima Stopel, Clint Feher, Nir Nissim, Yuval Elovici
ISI1
2007 Detection of Unknown Computer Worms Activity Based on Computer Behavior using Data Mining
abstract
Detecting unknown worms is a challenging task. Extant solutions, such as anti-virus tools, rely mainly on prior explicit knowledge of specific worm signatures. As a result, after the appearance of a new worm on the Web there is a significant delay until an update carrying the worm's signature is distributed to anti-virus tools. During this time interval a new worm can infect many computers and cause significant damage. We propose an innovative technique for detecting the presence of an unknown worm, not necessarily by recognizing specific instances of the worm, but rather based on the computer measurements. We designed an experiment to test the new technique employing several computer configurations and background applications activity. During the experiments 323 computer features were monitored. Four feature selection techniques were used to reduce the amount of features and four classification algorithms were applied on the resulting feature subsets. Our results indicate that using this approach resulted in exceeding 90% mean accuracy, and for specific unknown worms accuracy reached above 99%, using just 20 features while maintaining a low level of false positive rate.
Robert Moskovitch, Ido Gus, Shay Pluderman, Dima Stopel, Clint Feher, Chanan Glezer, Yuval Shahar, Yuval Elovici
CIDM1
2007 Detection of Unknown Computer Worms Activity Based on Computer Behavior using Data Mining
abstract
Detecting unknown worms is a challenging task. Extant solutions, such as anti-virus tools, rely mainly on prior explicit knowledge of specific worm signatures. As a result, after the appearance of a new worm on the Web there is a significant delay until an update carrying the worm's signature is distributed to anti-virus tools. During this time interval a new worm can infect many computers and create significant damage. We propose an innovative technique for detecting the presence of an unknown worm, not necessarily by recognizing specific instances of the worm, but rather based on the computer measurements. We designed an experiment to test the new technique employing several computer configurations and background applications activity. During the experiments 323 computer features were monitored. Four feature selection techniques were used to reduce the amount of features and four classification algorithms were applied on the resulting feature subsets. Our results indicate that using this approach resulted, in above 90% average accuracy, and for specific unknown worms accuracy reached above 99%, using just 20 features while maintaining a low level of false positive rate
Robert Moskovitch, Ido Gus, Shay Pluderman, Dima Stopel, Chanan Glezer, Yuval Shahar, Yuval Elovici
CISDA1
2007 Malicious Code Detection and Acquisition Using Active Learning
abstract
Detection of known malicious code is commonly performed by anti-virus tools. These tools detect the known malicious code using signature detection methods. Each time a new malicious code is found the anti-virus vendors create a new signature and update their clients. During the period between the appearance of a new unknown malicious code and the update of the signature base of the anti-virus clients, millions of computers might be infected. In order to cope with this problem, new solutions must be found for detecting unknown malicious code at the entrance of a client's computer. We presented here the use of active learning in the acquisition of unknown malicious code. Preliminary Results are encouraging. We are currently in the process of creating a wide test collection of more than 30,000 benign and malicious files to evaluate several active learning criterions.
Robert Moskovitch, Nir Nissim, Yuval Elovici
ISI1
2007 Host Based Intrusion Detection using Machine Learning
abstract
Detecting unknown malicious code (malcode) is a challenging task. Current common solutions, such as anti-virus tools, rely heavily on prior explicit knowledge of specific instances of malcode binary code signatures. During the time between its appearance and an update being sent to anti-virus tools, a new worm can infect many computers and cause significant damage. We present a new host-based intrusion detection approach, based on analyzing the behavior of the computer to detect the presence of unknown malicious code. The new approach consists on classification algorithms that learn from previous known malcode samples which enable the detection of an unknown malcode. We performed several experiments to evaluate our approach, focusing on computer worms being activated on several computer configurations while running several programs in order to simulate background activity. We collected 323 features in order to measure the computer behavior. Four classification algorithms were applied on several feature subsets. The average detection accuracy that we achieved was above 90% and for specific unknown worms even above 99%.
Robert Moskovitch, Shay Pluderman, Ido Gus, Dima Stopel, Clint Feher, Yisrael Parmet, Yuval Shahar, Yuval Elovici
ISI1
2007 Application of Information Technology: A Comparative Evaluation of Full-text, Concept-based, and Context-sensitive Search
abstract
OBJECTIVES: Study comparatively (1) concept-based search, using documents pre-indexed by a conceptual hierarchy; (2) context-sensitive search, using structured, labeled documents; and (3) traditional full-text search. Hypotheses were: (1) more contexts lead to better retrieval accuracy; and (2) adding concept-based search to the other searches would improve upon their baseline performances. DESIGN: Use our Vaidurya architecture, for search and retrieval evaluation, of structured documents classified by a conceptual hierarchy, on a clinical guidelines test collection. MEASUREMENTS: Precision computed at different levels of recall to assess the contribution of the retrieval methods. Comparisons of precisions done with recall set at 0.5, using t-tests. RESULTS: Performance increased monotonically with the number of query context elements. Adding context-sensitive elements, mean improvement was 11.1% at recall 0.5. With three contexts, mean query precision was 42% +/- 17% (95% confidence interval [CI], 31% to 53%); with two contexts, 32% +/- 13% (95% CI, 27% to 38%); and one context, 20% +/- 9% (95% CI, 15% to 24%). Adding context-based queries to full-text queries monotonically improved precision beyond the 0.4 level of recall. Mean improvement was 4.5% at recall 0.5. Adding concept-based search to full-text search improved precision to 19.4% at recall 0.5. CONCLUSIONS: The study demonstrated usefulness of concept-based and context-sensitive queries for enhancing the precision of retrieval from a digital library of semi-structured clinical guideline documents. Concept-based searches outperformed free-text queries, especially when baseline precision was low. In general, the more ontological elements used in the query, the greater the resulting precision.
Robert Moskovitch, Susana B. Martins, Eytan Behiri, Aviram Weiss, Yuval Shahar
J. Am. Medical Informatics Assoc.1
2006 Application of Artificial Neural Networks Techniques to Computer Worm Detection
abstract
Detecting computer worms is a highly challenging task. Commonly this task is performed by antivirus software tools that rely on prior explicit knowledge of the worm's code, which is represented by signatures. We present a new approach based on artificial neural networks (ANN) for detecting the presence of computer worms based on the computer's behavioral measures. In order to evaluate the new approach, several computers were infected with seven different worms and more than sixty different parameters of the infected computers were measured. The ANN and two other known classifications techniques, decision tree and k-nearest neighbors, were used to test their ability to classify correctly the presence, and the type, of the computer worms even during heavy user activity on the infected computers. The comparisons between the three approaches suggest that the ANN approach have computational advantages when real-time computation is needed, and has the potential to detect previously unknown worms. In addition, ANN may be used to identify the most relevant, measurable, features and thus reduce the feature dimensionality.
Dima Stopel, Zvi Boger, Robert Moskovitch, Yuval Shahar, Yuval Elovici
IJCNN3
2006 Multiple hierarchical classification of free-text clinical guidelines
Robert Moskovitch, Shiva Cohen-Kashi, Uzi Dror, Iftah Levy, Amit Maimon, Yuval Shahar
Artif. Intell. Medicine1
2005 Helping Physicians to Organize Guidelines Within Conceptual Hierarchies
Diego Sona, Paolo Avesani, Robert Moskovitch
AIME3
2004 A framework for a distributed, hybrid, multiple-ontology clinical-guideline library, and automated guideline-support tools
Yuval Shahar, Ohad Young, Erez Shalom, Maya Galperin-Aizenberg, Alon Mayaffit, Robert Moskovitch, Alon Hessing
J. Biomed. Informatics6
2003 DEGEL: A Hybrid, Multiple-Ontology Framework for Specification and Retrieval of Clinical Guidelines
Yuval Shahar, Ohad Young, Erez Shalom, Alon Mayaffit, Robert Moskovitch, Alon Hessing, Maya Galperin-Aizenberg
AIME5