VLDB 2026 Research / reviewers in the wild / expert
Mark Last
dblp:84/2603
· DBLP profile ↗
71ranked-venue papers
17as first author
11since 2021 · last 2025
0000-0003-0748-7918ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 44 · 12 first-author · 7 since 2021Databases, data management, data science and information retrieval · 29 · 9 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 8 · 1 first-author · 3 since 2021Software engineering, systems software and programming languages · 5 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 4 · 1 since 2021Security and privacy · 2 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Transitive self-consistency evaluation of NLI models without gold labelsabstractNatural Language Inference (NLI) is an important task in natural language processing.NLI models are aimed at automatically determining logical relationships between pairs of sentences.However, recent studies based on gold labels assigned to sentence pairs by human experts have provided some evidence that NLI models tend to make inconsistent model decisions during inference.Previous studies have used existing NLI datasets to test the transitive consistency of language models.However, they test only variations of two transitive consistency rules out of four.To further evaluate the transitive consistency of NLI models, we propose a novel evaluation approach that allows us to test all four rules automatically by generating adversarial examples via antonym replacements.Since we are testing self-consistency, human labeling of generated adversarial examples is unnecessary.Our experiments on several benchmark datasets indicate that the examples generated by the proposed antonym replacement methodology can reveal transitive inconsistencies in the state-of-the-art NLI models. Mark Last |
EMNLP | 2 |
| 2025 | Fast online feature selection in streaming dataabstractAbstract The challenge of getting big amounts of high-quality labeled data is compounded by the fact that data labeling is often subjective and requires significant human effort. In many cases, the quality of the labeled data depends entirely on the expertise and experience of human annotators, making it challenging to ensure labeling accuracy in large and dynamic datasets. Moreover, there may be a significant delay between the arrival of a new instance and its manual labeling. This paper explores the use of fully unsupervised feature selection algorithms in non-stationary data streams, where the importance of features may change over time. We introduce a novel feature selection algorithm called Online Fast FEa-ture SELection-OFFESEL, which calculates the feature importance scores in each incoming window based on their mean normalized values and without using any class labels. We evaluate OFFESEL on 17 benchmark data streams, both stationary and non-stationary, using popular online classifiers like PerceptronMask, VFDT, Online Boosting, and Linear SVM. We compare OFFESEL to several other feature selection algorithms, including state-of-the-art supervised ones like FIRES and ABFS, as well as popular unsupervised ones like MCFS, LS, and Max Variance, which we adapted to data streams. Our results indicate that OFFESEL outperforms all supervised and unsupervised feature selection algorithms in terms of classification accuracy. Specifically, OFFESEL preserves the accuracy level of the supervised FIRES algorithm, which proved more accurate than ABFS in our experiments, while maintaining the accuracy level achieved by the unsupervised Max Variance algorithm. Moreover, OFFESEL requires even less computation time than Max Variance and shows high stability on stationary datasets. Overall, our study demonstrates the potential benefits of using unlabeled data for feature ranking and selection in dynamic data streams. Yael Hochma, Mark Last |
Mach. Learn. | 2 |
| 2024 | Mining Eye-Tracking Data for Text SummarizationabstractIn this study, we introduce and evaluate a novel extractive text summarization methodology, “SummarEyes,” based on the visual interaction of the user with the text, using eye-tracking data, as opposed to the traditional approaches based on analysis of textual content only. We conducted a large-scale user study aiming to collect eye-tracking data while reading the text to be summarized. We utilized various user’s implicit attention metrics to generate novel eye-tracking-based text summarization models and compared them both to eye-tracking models typically using only a single feature of the gaze duration and to traditional, as well as state-of-the-art summarization methods, based solely on textual features. The models’ quality was evaluated in terms of ROUGE scores using intrinsic evaluation on the datasets we had generated, relating gaze behavior to personalized and DUC gold-standard summaries. The experimental results showed that “SummarEyes” significantly outperformed the other summarizers in predicting both the user’s personalized summarization and the generic gold standard summaries. With the increasing availability of eye-tracking technology, this research can lead to a new generation of effective user-centric text summarization tools. Meirav Taieb-Maimon, Aleksandr Romanovski-Chernik, Mark Last, Marina Litvak, Michael Elhadad |
Int. J. Hum. Comput. Interact. | 3 |
| 2024 | Towards efficient image-based representation of tabular data
Amit Damri, Mark Last, Niv Cohen |
Neural Comput. Appl. | 2 |
| 2022 | Early Detection of Multilingual Troll Accounts on TwitterabstractInternet troll farms have recently been employed as a powerful and prevailing weapon of information warfare. Even though different tactics may be utilized by different groups of state-sponsored trolls, our goal is to leverage identified troll data for revealing new emerging trolls generating multilingual content. In this work, we adopt a model agnostic meta-learning framework making use of previously released troll farm datasets for the early detection of newly-emerged troll accounts from identified or unidentified troll farms. The detection earliness of various models is evaluated using variable amounts of the earliest tweets from the tested accounts. To evaluate the proposed meta-model, we compare it to several classification models based on different types of account features. Our experiments demonstrate the effectiveness of the meta-model requiring as few as ten tweets to detect a troll account with an average accuracy of 94%. Mark Last, Marina Litvak |
ASONAM | 2 |
| 2022 | Inferring Event Causality in Films via Common Knowledge Corpora
Ben Aidlin, Armin Shmilovici, Mark Last |
ICCCI | 3 |
| 2022 | BRUCE: Bundle Recommendation Using Contextualized item EmbeddingsabstractA bundle is a pre-defined set of items that are collected together. In many domains, bundling is one of the most important marketing strategies for item promotion, commonly used in e-commerce. Bundle recommendation resembles the item recommendation task, where bundles are the recommended unit, but it poses additional challenges; while item recommendation requires only user and item understanding, bundle recommendation also requires modeling the connections between the various items in a bundle. Transformers have driven the state-of-the-art methods for set and sequence modeling in various natural language processing and computer vision tasks, emphasizing the understanding that the neighbors of an element are of crucial importance. Under some required adjustments, we believe the same applies for items in bundles, and better capturing the relations of an item with other items in the bundle may lead to improved recommendations. To address that, we introduce BRUCE - a novel model for bundle recommendation, in which we adapt Transformers to represent data on users, items, and bundles. This allows exploiting the self-attention mechanism to model the following: latent relations between the items in a bundle; and users’ preferences toward each of the items in the bundle and toward the whole bundle. Moreover, we examine various architectures to integrate the items’ and the users’ information and provide insights on architecture selection based on data characteristics. Experiments conducted on three benchmark datasets show that the proposed approach contributes to the accuracy of the recommendation and substantially outperforms state-of-the-art methods Tzoof Avny Brosh, Amit Livne, Oren Sar Shalom, Bracha Shapira, Mark Last |
RecSys | 5 |
| 2022 | Tracking social media during the COVID-19 pandemic: The case study of lockdown in New York State
Mark Last, Marina Litvak |
Expert Syst. Appl. | 2 |
| 2021 | Pattern Recognition in Vital Signs Using SpectrogramsabstractSpectrograms visualize the frequency components of a given signal which may be an audio signal or even a time-series signal. Audio signals have higher sampling rate and high variability of frequency with time. Spectrograms can capture such variations well. But, vital signs which are time-series signals have less sampling frequency and low-frequency variability due to which, spectrograms fail to express variations and patterns. In this paper, we propose a novel solution to introduce frequency variability using frequency modulation on vital signs. Then we apply spectrograms on frequency modulated signals to capture the patterns. The proposed approach has been evaluated on 4 different medical datasets across both prediction and classification tasks. Significant results are found showing the efficacy of the approach for vital sign signals. The results from the proposed approach are promising with an accuracy of 91.55% and 91.67% in prediction and classification tasks respectively. Sidharth Srivatsav Sribhashyam, Md Sirajus Salekin, Dmitry B. Goldgof, Ghada Zamzmi, Mark Last, Yu Sun 0004 |
SMC | 5 |
| 2021 | Facing airborne attacks on ADS-B data with autoencoders
Asaf Fried, Mark Last |
Comput. Secur. | 2 |
| 2021 | Penalty term based suitable fuzzy intuitionistic possibilistic clustering: analyzing high dimensional gene expression cancer database
S. R. Kannan, Esha Kashyap, Mark Last, Tzung-Pei Hong |
Soft Comput. | 3 |
| 2020 | Detecting Troll Tweets in a Bilingual CorpusabstractDuring the past several years, a large amount of troll accounts has emerged with efforts to manipulate public opinion on social network sites. They are often involved in spreading misinformation, fake news, and propaganda with the intent of distracting and sowing discord. This paper aims to detect troll tweets in both English and Russian assuming that the tweets are generated by some “troll farm.” We reduce this task to the authorship verification problem of determining whether a single tweet is authored by a “troll farm” account or not. We evaluate a supervised classification approach with monolingual, cross-lingual, and bilingual training scenarios, using several machine learning algorithms, including deep learning. The best results are attained by the bilingual learning, showing the area under the ROC curve (AUC) of 0.875 and 0.828, for tweet classification in English and Russian test sets, respectively. It is noteworthy that these results are obtained using only raw text features, which do not require manual feature engineering efforts. In this paper, we introduce a resource of English and Russian troll tweets containing original tweets and translation from English to Russian, Russian to English. It is available for academic purposes. Mark Last, Marina Litvak |
LREC | 2 |
| 2020 | An unsupervised constrained optimization approach to compressive summarization
Natalia Vanetik, Marina Litvak, Elena Churkin, Mark Last |
Inf. Sci. | 4 |
| 2019 | Identifying turning points in animated cartoons
Chang Liu 0038, Mark Last, Armin Shmilovici |
Expert Syst. Appl. | 2 |
| 2018 | Sentence Compression as a Supervised Learning with a Rich Feature Space
Elena Churkin, Mark Last, Marina Litvak, Natalia Vanetik |
CICLing (2) | 2 |
| 2018 | Responsive News Summarization for Ubiquitous Consumption on Multiple Mobile DevicesabstractWith the proliferation of online news read on devices ranging from desktops to smart watches, the need for meaningful summaries of long texts is growing. Manual summaries are labour-intensive and cannot be offered for all display sizes, whereas today's abstracts of most news texts are teasers designed to attract the reader's interest more than to provide an overview of an article's content suited to the reader's information needs. We propose responsive news summarization as a technological approach for filling this gap. Responsive news summarization provides an automatically generated content summary that has the right length for the device requesting the article, plus access to the full text. We describe the system prototype available at multisizenews.com along with the initial user study results and give an outlook on future work. Rocío A. Chongtay, Mark Last, Bettina Berendt |
IUI | 2 |
| 2018 | Hypotensive Episode Prediction in ICUs via Observation Window Splitting
Elad Tsur, Mark Last, Victor F. Garcia, Raphael Udassin, Moti Klein, Evgeni Brotfain |
ECML/PKDD (3) | 2 |
| 2017 | Short-Term Load Forecasting in Smart Meters with Sliding Window-Based ARIMA Algorithms
Dima Alberg, Mark Last |
ACIIDS (2) | 2 |
| 2017 | PCM-SABRE: a platform for benchmarking and comparing outcome prediction methods in precision cancer medicineabstractBACKGROUND: Numerous publications attempt to predict cancer survival outcome from gene expression data using machine-learning methods. A direct comparison of these works is challenging for the following reasons: (1) inconsistent measures used to evaluate the performance of different models, and (2) incomplete specification of critical stages in the process of knowledge discovery. There is a need for a platform that would allow researchers to replicate previous works and to test the impact of changes in the knowledge discovery process on the accuracy of the induced models. RESULTS: We developed the PCM-SABRE platform, which supports the entire knowledge discovery process for cancer outcome analysis. PCM-SABRE was developed using KNIME. By using PCM-SABRE to reproduce the results of previously published works on breast cancer survival, we define a baseline for evaluating future attempts to predict cancer outcome with machine learning. We used PCM-SABRE to replicate previous work that describe predictive models of breast cancer recurrence, and tested the performance of all possible combinations of feature selection methods and data mining algorithms that was used in either of the works. We reconstructed the work of Chou et al. observing similar trends - superior performance of Probabilistic Neural Network (PNN) and logistic regression (LR) algorithms and inconclusive impact of feature pre-selection with the decision tree algorithm on subsequent analysis. CONCLUSIONS: PCM-SABRE is a software tool that provides an intuitive environment for rapid development of predictive models in cancer precision medicine. Noah Eyal-Altman, Mark Last, Eitan Rubin |
BMC Bioinform. | 2 |
| 2016 | Multi-target Classification: Methodology and Practical Case Studies
Mark Last |
ECML/PKDD (3) | 1 |
| 2016 | Evolving classification of intensive care patients from event data
Mark Last, Olga Tosas, Tiziano Gallo Cassarino, Zisis Kozlakidis, Jonathan D. Edgeworth |
Artif. Intell. Medicine | 1 |
| 2015 | Krimping texts for better summarizationabstractAutomated text summarization is aimed at extracting essential information from original text and presenting it in a minimal, often predefined, number of words.In this paper, we introduce a new approach for unsupervised extractive summarization, based on the Minimum Description Length (MDL) principle, using the Krimp dataset compression algorithm (Vreeken et al., 2011).Our approach represents a text as a transactional dataset, with sentences as transactions, and then describes it by itemsets that stand for frequent sequences of words.The summary is then compiled from sentences that compress (and as such, best describe) the document.The problem of summarization is reduced to the maximal coverage, following the assumption that a summary that best describes the original text, should cover most of the word sequences describing the document.We solve it by a greedy algorithm and present the evaluation results. Marina Litvak, Mark Last, Natalia Vanetik |
EMNLP | 2 |
| 2014 | Improving accuracy of classification models induced from anonymized datasets
Mark Last, Tamir Tassa, Alexandra Zhmudyak, Erez Shmueli |
Inf. Sci. | 1 |
| 2014 | A Scalable Algorithm for One-to-One, Onto, and Partial Schema Matching with Uninterpreted Column Names and Column ValuesabstractIn this paper, the authors propose a five-step approach to the problem of identifying semantic correspondences between attributes of two database schemas. It is one of the key challenges in many database applications such as data integration and data warehousing. The authors' research is focused on uninterpreted schema matching, where the column names and column values are uninterpreted or unreliable. The approach implements Bayesian networks, Pearson's correlation and mutual information to identify inter-attribute dependencies. Additionally, the authors propose an extension to their algorithm that allows the user to manually enter the known mappings to improve the automated matching results. The five-step approach also allows data privacy preservation. The authors' evaluation experiments show that the proposed approach enhances the current set of schema matching techniques. Boris Rabinovich, Mark Last |
J. Database Manag. | 2 |
| 2013 | Automatic Identification of Conceptual Metaphors With Limited KnowledgeabstractFull natural language understanding requires identifying and analyzing the meanings of metaphors, which are ubiquitous in both text and speech. Over the last thirty years, linguistic metaphors have been shown to be based on more general conceptual metaphors, partial semantic mappings between disparate conceptual domains. Though some achievements have been made in identifying linguistic metaphors over the last decade or so, little work has been done to date on automatically identifying conceptual metaphors. This paper describes research on identifying conceptual metaphors based on corpus data. Our method uses as little background knowledge as possible, to ease transfer to new languages and to mini- mize any bias introduced by the knowledge base construction process. The method relies on general heuristics for identifying linguistic metaphors and statistical clustering (guided by Wordnet) to form conceptual metaphor candidates. Human experiments show the system effectively finds meaningful conceptual metaphors. Lisa Gandy, Nadji Allan, Mark Atallah, Ophir Frieder, Newton Howard, Sergey Kanareykin, Moshe Koppel, Mark Last, Yair Neuman, Shlomo Argamon |
AAAI | 8 |
| 2013 | Avoiding the Look-Ahead Pathology of Decision Tree LearningabstractMost decision-tree induction algorithms are using a local greedy strategy, where a leaf is always split on the best attribute according to a given attribute-selection criterion. A more accurate model could possibly be found by looking ahead for alternative subtrees. However, some researchers argue that the look-ahead should not be used due to a negative effect (called “decision-tree pathology”) on the decision-tree accuracy. This paper presents a new look-ahead heuristics for decision-tree induction. The proposed method is called look-ahead J48 ( LA-J48) as it is based on J48, the Weka implementation of the popular C4.5 algorithm. At each tree node, the LA-J48 algorithm applies the look-ahead procedure of bounded depth only to attributes that are not statistically distinguishable from the best attribute chosen by the greedy approach of C4.5. A bootstrap process is used for estimating the standard deviation of splitting criteria with unknown probability distribution. Based on a separate validation set, the attribute producing the most accurate subtree is chosen for the next step of the algorithm. In experiments on 20 benchmark data sets, the proposed look-ahead method outperforms the greedy J48 algorithm with the gain ratio and the gini index splitting criteria, thus avoiding the look-ahead pathology of decision-tree induction. Mark Last, Michael Roizman |
Int. J. Intell. Syst. | 1 |
| 2013 | Cross-lingual training of summarization systems using annotated corpora in a foreign language
Marina Litvak, Mark Last |
Inf. Retr. | 2 |
| 2012 | Large-scale analysis of self-disclosure patterns among online social networks users: a Russian context
Slava Kisilevich, Chee Siang Ang, Mark Last |
Knowl. Inf. Syst. | 3 |
| 2012 | A Comparative Study of Artificial Neural Networks and Info-Fuzzy Networks as Automated Oracles in Software TestingabstractSoftware quality is one of the main concerns of software users. Hence, software testing is an utterly important phase in the software development life cycle. Nevertheless, manual evaluation of program compliance with its specification may be prohibitively time consuming. As a remedy, several software testing systems are using an automatic oracle to confirm that the developed software complies with its specification and determine whether a given test case exposes faults. The use of artificial neural networks and info-fuzzy networks as automated oracles has been explored elsewhere. Nevertheless, there is not enough research comparing these two popular approaches to automated evaluation of the test outcome. This paper fills the gap and reports on a set of experiments designed to compare the two methods based on ROC curves, training time, and dispersion analysis. Deepam Agarwal, Dan E. Tamir, Mark Last, Abraham Kandel |
IEEE Trans. Syst. Man Cybern. Part A | 3 |
| 2011 | Classification of infectious diseases based on chemiluminescent signatures of phagocytes in whole blood
Daria Prilutsky, Boris Rogachev, Robert S. Marks, Leslie Lobel, Mark Last |
Artif. Intell. Medicine | 5 |
| 2011 | Using data mining techniques for optimizing traffic signal plans at an urban intersectionabstractA key problem in traffic engineering is the optimization of the flow of vehicles through urban intersections by improving the timing policy of traffic signals. Current methods of signal control policy are based on the junction topography and prespecified static traffic volumes. However, the actual daily traffic volumes can be affected by many time-dependent factors making a static policy hardly optimal. In this paper, we induce nonstationary predictive models of traffic flow by applying novel methods of time-series data mining to the traffic sensors data collected from a signalized intersection in Jerusalem over a period of 3 years. Our methodology for modeling dynamic traffic volumes combines clustering and segmentation algorithms. The results of a case study based on real-world traffic data demonstrate that a dynamic signal policy using the data mining approach can produce a decrease of about 33.7% in the total waiting time of drivers during 1 year, in comparison to the existing static traffic policy. The resulting savings for this junction only would be about 13,800 driving hours, which are worth of about $52,000 per annum in terms of Israeli economy. © 2011 Wiley Periodicals, Inc. Mark Last, Gil Avrahami-Bakish, Abraham Kandel |
Int. J. Intell. Syst. | 1 |
| 2010 | Predictive Maintenance with Multi-target Classification Models
Mark Last, Alla Sinaiski, Halasya Siva Subramania |
ACIIDS (2) | 1 |
| 2010 | A New Approach to Improving Multilingual Summarization Using a Genetic Algorithm
Marina Litvak, Mark Last, Menahem Friedman |
ACL | 2 |
| 2010 | Detection of access to terror-related Web sites using an Advanced Terror Detection System (ATDS)abstractAbstract Terrorist groups use the Web as their infrastructure for various purposes. One example is the forming of new local cells that may later become active and perform acts of terror. The Advanced Terrorist Detection System (ATDS), is aimed at tracking down online access to abnormal content, which may include terrorist‐generated sites, by analyzing the content of information accessed by the Web users. ATDS operates in two modes: the training mode and the detection mode. In the training mode, ATDS determines the typical interests of a prespecified group of users by processing the Web pages accessed by these users over time. In the detection mode, ATDS performs real‐time monitoring of the Web traffic generated by the monitored group, analyzes the content of the accessed Web pages, and issues an alarm if the accessed information is not within the typical interests of that group and similar to the terrorist interests. An experimental version of ATDS was implemented and evaluated in a local network environment. The results suggest that when optimally tuned the system can reach high detection rates of up to 100% in case of continuous access to a series of terrorist Web pages. Yuval Elovici, Bracha Shapira, Mark Last, Omer Zaafrany, Menahem Friedman, Moti Schneider, Abraham Kandel |
J. Assoc. Inf. Sci. Technol. | 3 |
| 2009 | Improving data mining utility with projective samplingabstractOverall performance of the data mining process depends not just on the value of the induced knowledge but also on various costs of the process itself such as the cost of acquiring and pre-processing training examples, the CPU cost of model induction, and the cost of committed errors. Recently, several progressive sampling strategies for maximizing the overall data mining utility have been proposed. All these strategies are based on repeated acquisitions of additional training examples until a utility decrease is observed. In this paper, we present an alternative, projective sampling strategy, which fits functions to a partial learning curve and a partial run-time curve obtained from a small subset of potentially available data and then uses these projected functions to analytically estimate the optimal training set size. The proposed approach is evaluated on a variety of benchmark datasets using the RapidMiner environment for machine learning and data mining processes. The results show that the learning and run-time curves projected from only several data points can lead to a cheaper data mining process than the common progressive sampling methods. Mark Last |
KDD | 1 |
| 2008 | The hybrid representation model for web document classificationabstractMost Web content categorization methods are based on the vector space model of information retrieval. One of the most important advantages of this representation model is that it can be used by both instance-based and model-based classifiers. However, this popular method of document representation does not capture important structural information, such as the order and proximity of word occurrence or the location of a word within the document. It also makes no use of the markup information that can easily be extracted from the Web document HTML tags. A recently developed graph-based Web document representation model can preserve Web document structural information. It was shown to outperform the traditional vector representation using the k-Nearest Neighbor (k-NN) classification algorithm. The problem, however, is that the eager (model-based) classifiers cannot work with this representation directly. In this article, three new hybrid approaches to Web document classification are presented, built upon both graph and vector space representations, thus preserving the benefits and overcoming the limitations of each. The hybrid methods presented here are compared to vector-based models using the C4.5 decision tree and the probabilistic Naïve Bayes classifiers on several benchmark Web document collections. The results demonstrate that the hybrid methods presented in this article outperform, in most cases, existing approaches in terms of classification accuracy, and in addition, achieve a significant reduction in the classification time. © 2008 Wiley Periodicals, Inc. Alex Markov, Mark Last, Abraham Kandel |
Int. J. Intell. Syst. | 2 |
| 2007 | ADMIRAL: A Data Mining Based Financial Trading SystemabstractThis paper presents a novel framework for predicting stock trends and making financial trading decisions based on a combination of data and text mining techniques. The prediction models of the proposed system are based on the textual content of time-stamped Web documents in addition to traditional numerical time series data, which is also available from the Web. The financial trading system based on the model predictions (ADMIRAL) is using three different trading strategies. In this paper, the ADMIRAL system is simulated and evaluated on real-world series of news stories and stocks data using the C4.5 decision tree induction algorithm. The main performance measures are the predictive accuracy of the induced models and, more importantly, the profitability of each trading strategy using these predictions Gil Rachlin, Mark Last, Dima Alberg, Abraham Kandel |
CIDM | 2 |
| 2007 | Predicting future locations using clusters' centroidsabstractAs technology advances we encounter more available data on moving objects, thus increasing our ability to mine spatio-temporal data. We can use this data for learning moving objects behavior and for predicting their locations at future times according to the extracted movement patterns.In this paper we cluster trajectories of a mobile object and utilize the accepted cluster centroids as the object's movement patterns. We use the obtained movement patterns for predicting the object location at specific future times. We evaluate our prediction results using precision and recall measures. We also remove exceptional data points from the moving patterns by optimizing the value of an exceptions threshold. Sigal Elnekave, Mark Last, Oded Maimon |
GIS | 2 |
| 2007 | Anomaly detection in web documents using crisp and fuzzy-based cosine clustering methodology
Menahem Friedman, Mark Last, Yaniv Makover, Abraham Kandel |
Inf. Sci. | 2 |
| 2007 | Special issue on advances in fuzzy logic
Abraham Kandel, Mark Last |
Inf. Sci. | 2 |
| 2007 | Modeling software testing costs and risks using fuzzy logic paradigm
Avner Engel, Mark Last |
J. Syst. Softw. | 2 |
| 2006 | Class Diagrams and Use Cases
Peretz Shoval, Avihai Yampolsky, Mark Last |
EMMSAD | 3 |
| 2006 | Design of test inputs and their sequences in multi-function system testing
Mark Sh. Levin, Mark Last |
Appl. Intell. | 2 |
| 2006 | Comparing representative selection strategies for dissimilarity representationsabstractMany of the computational intelligence techniques currently used do not scale well in data type or computational performance, so selecting the right dimensionality reduction technique for the data is essential. By employing a dimensionality reduction technique called representative dissimilarity to create an embedded space, large spaces of complex patterns can be simplified to a fixed-dimensional Euclidean space of points. The only current suggestions as to how the representatives should be selected are principal component analysis, projection pursuit, and factor analysis. Several alternative representative strategies are proposed and empirically evaluated on a set of term vectors constructed from HTML documents. The results indicate that using a representative dissimilarity representation with at least 50 representatives can achieve a significant increase in classification speed, with a minimal sacrifice in accuracy, and when the representatives are selected randomly, the time required to create the embedded space is significantly reduced, also with a small penalty in accuracy. © 2006 Wiley Periodicals, Inc. Int J Int Syst 21: 1093–1109, 2006. Zane Reynolds, Horst Bunke, Mark Last, Abraham Kandel |
Int. J. Intell. Syst. | 3 |
| 2005 | Experimental Comparison of Sequence and Collaboration Diagrams in Different Application Domains
Chanan Glezer, Mark Last, Efrat Nahmani, Peretz Shoval |
EMMSAD | 2 |
| 2005 | Content-Based Detection of Terrorists Browsing the Web Using an Advanced Terror Detection System (ATDS)
Yuval Elovici, Bracha Shapira, Mark Last, Omer Zaafrany, Menahem Friedman, Moti Schneider, Abraham Kandel |
ISI | 3 |
| 2005 | A fuzzy-based path ordering algorithm for QoS routing in non-deterministic communication networks
Avichai Cohen, Ephraim Korach, Mark Last, Rony Ohayon |
Fuzzy Sets Syst. | 3 |
| 2005 | A fuzzy-based lifetime extension of genetic algorithms
Mark Last, Shay Eyal |
Fuzzy Sets Syst. | 1 |
| 2005 | Design and implementation of a web mining system for organizing search engine resultsabstractWe present the design and implementation of a web mining system that creates a hierarchical clustering of web documents retrieved by commercial web search engines. The cluster hierarchy is produced by a novel method called the Cluster Hierarchy Construction Algorithm (CHCA) and it can be used to explore the topics of interest related to the search query and their relationships. We discuss important design issues for our system, including stemming and dimensionality reduction, as well as some implementation details. We show examples of system results, compare them with results from similar systems, and analyze the responses to a survey of the system's users. © 2005 Wiley Periodicals, Inc. Int J Int Syst 20: 607–625, 2005. Adam Schenker, Mark Last, Abraham Kandel |
Int. J. Intell. Syst. | 2 |
| 2005 | Quality and comprehension of UML interaction diagrams-an experimental comparison
Chanan Glezer, Mark Last, Efrat Nachmany, Peretz Shoval |
Inf. Softw. Technol. | 2 |
| 2004 | A Graph-Based Framework for Web Document Mining
Adam Schenker, Horst Bunke, Mark Last, Abraham Kandel |
Document Analysis Systems | 3 |
| 2004 | Multi-objective Classification with Info-Fuzzy Networks
Mark Last |
ECML | 1 |
| 2004 | A new approach for fuzzy clustering of Web documentsabstractMost existing methods of document clustering are based on the classical vector-space model, which represents each document by a fixed-size vector of key terms or key phrases. In large and diverse document collections such as the World Wide Web, this approach suffers from a tremendous computational overload, since the constant size of the term vector equals to the total number of key terms in all documents. We propose a new fuzzy-based approach to clustering documents that are represented by vectors of variable size. Each entry in a vector consists of two fields. The first field is the name of a key phrase in the document and the second denotes an importance weight associated with this key phrase within the particular document. We will describe the proposed approach in detail and show how it is implemented in a real world application from the area of web monitoring. Menahem Friedman, Mark Last, Omer Zaafrany, Moti Schneider, Abraham Kandel |
FUZZ-IEEE | 2 |
| 2004 | Test Case Sequences in System Testing: Selection of Test Cases for a Chain (Sequence) of Function Clusters
Mark Sh. Levin, Mark Last |
IEA/AIE | 2 |
| 2004 | Terrorist Detection System
Yuval Elovici, Abraham Kandel, Mark Last, Bracha Shapira, Omer Zaafrany, Moti Schneider, Menahem Friedman |
PKDD | 3 |
| 2004 | Data mining in software metrics databases
Scott Dick, Aleksandra Meeks, Mark Last, Horst Bunke, Abraham Kandel |
Fuzzy Sets Syst. | 3 |
| 2004 | Classification Of Web Documents Using Graph MatchingabstractIn this paper we describe a classification method that allows the use of graph-based representations of data instead of traditional vector-based representations. We compare the vector approach combined with the k-Nearest Neighbor (k-NN) algorithm to the graph-matching approach when classifying three different web document collections, using the leave-one-out approach for measuring classification accuracy. We also compare the performance of different graph distance measures as well as various document representations that utilize graphs. The results show the graph-based approach can outperform traditional vector-based methods in terms of accuracy, dimensionality and execution time. Adam Schenker, Mark Last, Horst Bunke, Abraham Kandel |
Int. J. Pattern Recognit. Artif. Intell. | 2 |
| 2004 | Using Data Mining For Automated Software TestingabstractIn today's software industry, the design of test cases is mostly based on human expertise, while test automation tools are limited to execution of pre-planned tests only. Evaluation of test outcomes is also associated with a considerable effort by human testers who often have imperfect knowledge of the requirements specification. Not surprisingly, this manual approach to software testing results in heavy losses to the world's economy. In this paper, we demonstrate the potential use of data mining algorithms for automated modeling of tested systems. The data mining models can be utilized for recovering system requirements, designing a minimal set of regression tests, and evaluating the correctness of software outputs. To study the feasibility of the proposed approach, we have applied a state-of-the-art data mining algorithm called Info-Fuzzy Network (IFN) to execution data of a complex mathematical package. The IFN method has shown a clear capability to identify faults in the tested program. Mark Last, Menahem Friedman, Abraham Kandel |
Int. J. Softw. Eng. Knowl. Eng. | 1 |
| 2004 | A Compact and Accurate Model for ClassificationabstractWe describe and evaluate an information-theoretic algorithm for data-driven induction of classification models based on a minimal subset of available features. The relationship between input (predictive) features and the target (classification) attribute is modeled by a tree-like structure termed an information network (IN). Unlike other decision-tree models, the information network uses the same input attribute across the nodes of a given layer (level). The input attributes are selected incrementally by the algorithm to maximize a global decrease in the conditional entropy of the target attribute. We are using the prepruning approach: when no attribute causes a statistically significant decrease in the entropy, the network construction is stopped. The algorithm is shown empirically to produce much more compact models than other methods of decision-tree learning while preserving nearly the same level of classification accuracy. Mark Last, Oded Maimon |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2003 | Classification of Web Documents Using a Graph ModelabstractIn this paper we describe work relating to classification of Web documents using a graph-based model instead of the traditional vector-based model for document representation. We compare the classification accuracy of the vector model approach using the k-nearest neighbor (k-NN) algorithm to a novel approach which allows the use of graphs for document representation in the k-NN algorithm. The proposed method is evaluated on three different Web document collections using the leave-one-out approach for measuring classification accuracy. The results show that the graph-based k-NN approach can outperform traditional vector-based k-NN methods in terms of both accuracy and execution time. Adam Schenker, Mark Last, Horst Bunke, Abraham Kandel |
ICDAR | 2 |
| 2003 | The data mining approach to automated software testingabstractIn today’s industry, the design of software tests is mostly based on the testers’ expertise, while test automation tools are limited to execution of pre-planned tests only. Evaluation of test outputs is also associated with a considerable effort by human testers who often have imperfect knowledge of the requirements specification. Not surprisingly, this manual approach to software testing results in heavy losses to the world’s economy. The costs of the so-called "catastrophic " software failures (such as Mars Polar Lander shutdown in 1999) are even hard to measure. In this paper, we demonstrate the potential use of data mining algorithms for automated induction of functional requirements from execution data. The induced data mining models of tested software can be utilized for recovering missing and incomplete specifications, designing a minimal set of regression tests, and evaluating the correctness of software outputs when testing new, potentially flawed releases of the system. To study the feasibility of the proposed approach, we have applied a novel data mining algorithm called Info-Fuzzy Network (IFN) to execution data of a general-purpose code for solving partial differential equations. After being trained on a relatively small number of randomly generated input-output examples, the model constructed by the IFN algorithm has shown a clear capability to discriminate between correct and faulty versions of the program. Mark Last, Menahem Friedman, Abraham Kandel |
KDD | 1 |
| 2003 | Test case generation and reduction by automated input-output analysisabstractIn the software testing process, selecting the test cases and verifying their results requires a lot of subjective decisions and human intervention. For a program having a large number of inputs, the number of corresponding combinatorial black-box test cases is huge. A method needs to be established in order to limit the number of test cases and to choose the most important ones. In this research effort we present a novel methodology for identifying important test cases automatically. These test cases involve input attributes which contribute to the value of an output and hence are significant. The reduction in the number of test cases is attributed to identifying input-output relationships. A ranked list of features and equivalence classes for input attributes of a given code are the main outcomes of this methodology. Reducing the number of test cases results directly in the saving of software testing resources. Prachi Saraph, Mark Last, Abraham Kandel |
SMC | 2 |
| 2002 | Online classification of nonstationary data streams
Mark Last |
Intell. Data Anal. | 1 |
| 2002 | Using a neural network in the software testing processabstractSoftware testing forms an integral part of the software development life cycle. Since the objective of testing is to ensure the conformity of an application to its specification, a test “oracle” is needed to determine whether a given test case exposes a fault or not. Using an automated oracle to support the activities of human testers can reduce the actual cost of the testing process and the related maintenance costs. In this paper, we present a new concept of using an artificial neural network as an automated oracle for a tested software system. A neural network is trained by the backpropagation algorithm on a set of test cases applied to the original version of the system. The network training is based on the “black-box” approach, since only inputs and outputs of the system are presented to the algorithm. The trained network can be used as an artificial oracle for evaluating the correctness of the output produced by new and possibly faulty versions of the software. We present experimental results of using a two-layer neural network to detect faults within mutated code of a small credit approval application. The results appear to be promising for a wide range of injected faults. ? 2002 John Wiley & Sons, Inc. Meenakshi Vanmali, Mark Last, Abraham Kandel |
Int. J. Intell. Syst. | 2 |
| 2002 | Improving Stability of Decision TreesabstractDecision-tree algorithms are known to be unstable: small variations in the training set can result in different trees and different predictions for the same validation examples. Both accuracy and stability can be improved by learning multiple models from bootstrap samples of training data, but the "meta-learner" approach makes the extracted knowledge hardly interpretable. In the following paper, we present the Info-Fuzzy Network (IFN), a novel information-theoretic method for building stable and comprehensible decision-tree models. The stability of the IFN algorithm is ensured by restricting the tree structure to using the same feature for all nodes of the same tree level and by the built-in statistical significance tests. The IFN method is shown empirically to produce more compact and stable models than the "meta-learner" techniques, while preserving a reasonable level of predictive accuracy. Mark Last, Oded Maimon, Einat Minkov |
Int. J. Pattern Recognit. Artif. Intell. | 1 |
| 2002 | A Feature-Based Serial Approach to Classifier Combination
Mark Last, Horst Bunke, Abraham Kandel |
Pattern Anal. Appl. | 1 |
| 2001 | Re-granulating a Fuzzy RulebaseabstractWe describe a computing-with-words algorithm for changing the granularity of an existing rulebase. A fuzzy rulebase is regranulated by manually selecting the granularity of each input dimension, and then automatically calculating the rulebase consequents. We demonstrate this algorithm by studying the common "fuzzy-PD" rulebase at various granularities. We find evidence that the "3/spl times/3" fuzzy-PD rulebase may be too coarse a granularity for this problem. Scott Dick, Adam Schenker, Mark Last, Horst Bunke, Abraham Kandel |
FUZZ-IEEE | 3 |
| 2001 | Information-theoretic fuzzy approach to data reliability and data mining
Oded Maimon, Abraham Kandel, Mark Last |
Fuzzy Sets Syst. | 3 |
| 2001 | Information-theoretic algorithm for feature selection
Mark Last, Abraham Kandel, Oded Maimon |
Pattern Recognit. Lett. | 1 |
| 2001 | Knowledge discovery in time series databasesabstractAdding the dimension of time to databases produces time series databases (TSDB) and introduces new aspects and difficulties to data mining and knowledge discovery. In this correspondence, we introduce a general methodology for knowledge discovery in TSDB. The process of knowledge discovery in TSDR includes cleaning and filtering of time series data, identifying the most important predicting attributes, and extracting a set of association rules that can be used to predict the time series behavior in the future. Our method is based on signal processing techniques and the information-theoretic fuzzy approach to knowledge discovery. The computational theory of perception (CTP) is used to reduce the set of extracted rules by fuzzification and aggregation. We demonstrate our approach on two types of time series: stock-market data and weather data. Mark Last, Yaron Klein, Abraham Kandel |
IEEE Trans. Syst. Man Cybern. Part B | 1 |
| 1997 | Information-efficient control of discrete event stochastic systemsabstractA state of a discrete event stochastic system (DESS) can be represented by a tuple of time-varying discrete parameters. The authors have extended the theory of controlled discrete event systems developed by Ramadge, Wonham and other researchers to stochastic modeling and performance measurement. This work presents some important characteristics of the DESS model. The problem of making the most efficient use of information-processing resources (sensors, computer capacity, etc.) in a special class of controlled systems is stated in a new, two-stage format. The format is used to develop a heuristic solution procedure of the problem. Finally, we perform a brief discussion of the model applicability to real-life design of automatic supervisors. Oded Maimon, Mark Last |
IEEE Trans. Syst. Man Cybern. Part A | 2 |