Natalia Vanetik

dblp:59/3572 · DBLP profile ↗
← Back
26ranked-venue papers
10as first author
7since 2021 · last 2026
0000-0002-4939-1415ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 16 · 6 first-author · 4 since 2021Databases, data management, data science and information retrieval · 12 · 8 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021Security and privacy · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Towards Fine-Grained Emotion Detection in Hebrew: Dataset Creation and Evaluation
Natalia Vanetik, Shilat Haya Yosefi, Keturah Shlomo
DATA (1)1
2025 A New Dataset for Keyword Extraction from IT Job Descriptions
Nisan Fichman, Hadar Isaacson, Natalia Vanetik
ECIR (3)3
2024 Improving Factual Consistency in Abstractive Summarization with Sentence Structure Pruning
abstract
State-of-the-art abstractive summarization models still suffer from the content contradiction between the summaries and the input text, which is referred to as the factual inconsistency problem. Recently, a large number of works have also been proposed to evaluate factual consistency or improve it by post-editing methods. However, these post-editing methods typically focus on replacing suspicious entities, failing to identify and modify incorrect content hidden in sentence structures. In this paper, we first verify that the correctable errors can be enriched by leveraging sentence structure pruning operation, and then we propose a post-editing method based on that. In the correction process, the pruning operation on possible errors is performed on the syntactic dependency tree with the guidance of multiple factual evaluation metrics. Experimenting on the FRANK dataset shows a great improvement in factual consistency compared with strong baselines and, when combined with them, can achieve even better performance. All the codes and data will be released on paper acceptance.
Dingxin Hu, Xingyue Zhang, Marina Litvak, Natalia Vanetik, Qing Yang 0033, Dongliang Xu, Yanquan Zhou, Lei Li 0009, Yingqi Zhu
LREC/COLING7
2023 Summarizing Financial Reports with Positional Language Model
abstract
Financial reports are essential for the decision-making processes of various stakeholders, containing vast amounts of both quantitative and qualitative data. As the business world becomes increasingly intricate, stakeholders require a swift means to understand a company’s financial status. Text summarization is a useful tool in this regard, aiming to present long texts concisely without losing their essence. Given the complexity and the structured nature of financial documents, summarizing them poses a significant challenge. This paper suggests a method employing Positional Language Models (PLMs), a subset of non-neural language models that assess the sequence of tokens in input data, for financial report summarization. Our proposed method is unsupervised, eliminating the need for training and ensuring computational efficiency for lengthy documents.
Natalia Vanetik, Elizaveta Podkaminer, Marina Litvak
IEEE Big Data1
2023 Teledermatologist: Automated Diagnosis of Skin Diseases with Image Recognition
abstract
Early diagnosis of skin diseases is crucial for effective treatment, preventing spread, minimizing long-term damage, managing chronic conditions, detecting underlying health issues, and promoting psychological well-being. It is often necessary to do difficult tests in order to obtain a diagnosis, which demands a lot of resources from patients as well as medical personnel. Motivated by this shortage, we introduce a user-friendly and private system for the automatic detection of skin disease from a picture. Our system employs a deep neural model for the accurate classification of a given picture to one of the predefined classes of diseases. We created a large dataset of six skin lesion types and tested different neural models. The best-performing model, which achieved 92% accuracy, was integrated into our system. The application and its web demo can be accessed at https://teledermatologis-ai.streamlit.app/.
Nadav Ishai, Dolev Peretz, Natalia Vanetik, Marina Litvak
VCIP3
2022 Ransomware Detection with Deep Neural Networks
Matan Davidian, Natalia Vanetik, Michael Kiperberg
ICISSP2
2022 Offensive language detection in Hebrew: can other languages help?
abstract
Unfortunately, offensive language in social media is a common phenomenon nowadays. It harms many people and vulnerable groups. Therefore, automated detection of offensive language is in high demand and it is a serious challenge in multilingual domains. Various machine learning approaches combined with natural language techniques have been applied for this task lately. This paper contributes to this area from several aspects: (1) it introduces a new dataset of annotated Facebook comments in Hebrew; (2) it describes a case study with multiple supervised models and text representations for a task of offensive language detection in three languages, including two Semitic (Hebrew and Arabic) languages; (3) it reports evaluation results of cross-lingual and multilingual learning for detection of offensive content in Semitic languages; and (4) it discusses the limitations of these settings.
Marina Litvak, Natalia Vanetik, Chaya Liebeskind, Omar Hmdia, Rizek Abu Madeghem
LREC2
2020 Automated Discovery of Mathematical Definitions in Text
abstract
Automatic definition extraction from texts is an important task that has numerous applications in several natural language processing fields such as summarization, analysis of scientific texts, automatic taxonomy generation, ontology generation, concept identification, and question answering. For definitions that are contained within a single sentence, this problem can be viewed as a binary classification of sentences into definitions and non-definitions. Definitions in scientific literature can be generic (Wikipedia) or more formal (mathematical articles). In this paper, we focus on automatic detection of one-sentence definitions in mathematical texts, which are difficult to separate from surrounding text. We experiment with several data representations, which include sentence syntactic structure and word embeddings, and apply deep learning methods such as convolutional neural network (CNN) and recurrent neural network (RNN), in order to identify mathematical definitions. Our experiments demonstrate the superiority of CNN and its combination with RNN, applied on the syntactically-enriched input representation. We also present a new dataset for definition extraction from mathematical texts. We demonstrate that the use of this dataset for training learning models improves the quality of definition extraction when these models are then used for other definition datasets. Our experiments with different domains approve that mathematical definitions require special treatment, and that using cross-domain learning is inefficient.
Natalia Vanetik, Marina Litvak, Sergey Shevchuk, Lior Reznik
LREC1
2020 An unsupervised constrained optimization approach to compressive summarization
Natalia Vanetik, Marina Litvak, Elena Churkin, Mark Last
Inf. Sci.1
2019 EASY: Evaluation System for Summarization
Marina Litvak, Natalia Vanetik, Yael Veksler
CICLing (2)2
2019 In Conclusion Not Repetition: Comprehensive Abstractive Summarization with Diversified Attention Based on Determinantal Point Processes
abstract
Various Seq2Seq learning models designed for machine translation were applied for abstractive summarization task recently.Despite these models provide high ROUGE scores, they are limited to generate comprehensive summaries with a high level of abstraction due to its degenerated attention distribution.We introduce Diverse Convolutional Seq2Seq Model(DivCNN Seq2Seq) using Determinantal Point Processes methods(Micro DPPs and Macro DPPs) to produce attention distribution considering both quality and diversity.Without breaking the end to end architecture, Di-vCNN Seq2Seq achieves a higher level of comprehensiveness compared to vanilla models and strong baselines.All the reproducible codes and datasets are available online 1 .1 available at https://github.com/thinkwee/DPPCNN Sum marization Article: marseille , france the french prosecutor leading an investigation into the crash of germanwings flight 9525 insisted wednesday that he was not aware of any video footage from on board the plane .marseille prosecutor brice robin told cnn that so far no videos were used in the crash investigation ...... of a cell phone video showing the harrowing final seconds from on board germanwings flight 9525 as it crashed into the french alps .......paris match and bild reported that the video was recovered from a phone at the wreckage site ....... cnn 's frederik pleitgen , pamela boykoff , antonia mortensen , sandrine amiel and anna-maja rappard contributed to this report .CNN Seq2Seq: french prosecutor UNK robin says he was not aware of any video
Lei Li 0009, Wei Liu 0161, Marina Litvak, Natalia Vanetik, Zuying Huang
CoNLL4
2019 Quality-Diversity Summarization with Unsupervised Autoencoders
Lei Li 0009, Zuying Huang, Natalia Vanetik, Marina Litvak
ICANN (4)3
2018 Sentence Compression as a Supervised Learning with a Rich Feature Space
Elena Churkin, Mark Last, Marina Litvak, Natalia Vanetik
CICLing (2)4
2018 HEvaS: Headline Evaluation System
abstract
Automatic headline generation is a sub-task of oneline summarization with many reported applications. Evaluation of systems generating headlines is a very challenging and undeveloped area. In this paper, we introduce a system that performs automatic evaluation of systems in terms of a quality of the generated headlines. The evaluation is performed using multiple metrics for comparing evaluated headlines with the gold standard ones or measuring their coverage of main document topics. Both types of metrics evaluate headline's content and informativeness, but not grammatical structure. The only input required by our system is a set of documents with gold standard and automatically generated headlines. The Headline Evaluation System (HEvaS) provides a user with a choice from multiple (10) metrics, then calculates the chosen metrics, performs statistical analysis of the evaluated systems and visualizes the results. Multiple headline generation systems can be evaluated at the same run. This paper describes all evaluation metrics and architecture, utilized by our system. As an evaluation of the HEvaS, we perform a case study with a couple of baseline systems and report the results. Although we tested the system on English content only, the multilingual content can also be supported.
Marina Litvak, Natalia Vanetik, Itzhak Eretz Kdosha
WI2
2018 DRIM: MDL-Based Approach for Fast Diverse Summarization
abstract
Automated text summarization extracts essential information from original text and presents it in a predefined number of words. In this paper, we introduce an unsupervised extractive summarization approach that takes its roots from the SLIM dataset compression algorithm [1] based on the Minimum Description Length (MDL) principle [2], [3]. Our approach represents text as a transactional dataset, where sentences are transactions and normalized words are items. We use the SLIM algorithm (SLIM is not an abbreviation, it is Dutch word for 'smart') to solve the main bottleneck of the MDL computation, which is the generation of all frequent itemsets as a first step of the model construction. Additionally, we add a diversity constraint to the model in order to decrease appearance of repeated information in a summary. We introduce DRIM (Diversed SLIM) algorithm that performs unsupervised summarization, both generic and query-based, and does not require parameter tuning. We evaluate our summarizer on texts in English, but it can be easily extended to other languages.
Natalia Vanetik, Marina Litvak
WI1
2017 Summarizing Weibo with Topics Compression
Marina Litvak, Natalia Vanetik, Lei Li 0009
CICLing (2)2
2015 Krimping texts for better summarization
abstract
Automated text summarization is aimed at extracting essential information from original text and presenting it in a minimal, often predefined, number of words.In this paper, we introduce a new approach for unsupervised extractive summarization, based on the Minimum Description Length (MDL) principle, using the Krimp dataset compression algorithm (Vreeken et al., 2011).Our approach represents a text as a transactional dataset, with sentences as transactions, and then describes it by itemsets that stand for frequent sequences of words.The summary is then compiled from sentences that compress (and as such, best describe) the document.The problem of summarization is reduced to the maximal coverage, following the assumption that a summary that best describes the original text, should cover most of the word sequences describing the document.We solve it by a greedy algorithm and present the evaluation results.
Marina Litvak, Mark Last, Natalia Vanetik
EMNLP3
2015 Multilingual Summarization with Polytope Model
abstract
The problem of extractive text summarization for a collection of documents is defined as the problem of selecting a small subset of sentences so that the contents and meaning of the original document set are preserved in the best possible way.In this paper we describe the linear programming-based global optimization model to rank and extract the most relevant sentences to a summary.We introduce three different objective functions being optimized.These functions define a relevance of a sentence that is being maximized, in different manners, such as: coverage of meaningful words of a document, coverage of its bigrams, or coverage of frequent sequences of words.We supply here an overview of our system's participation in the MultiLing contest of SIGDial 2015.
Natalia Vanetik, Marina Litvak
SIGDIAL Conference1
2014 Redundancy-weighting for better inference of protein structural features
abstract
MOTIVATION: Structural knowledge, extracted from the Protein Data Bank (PDB), underlies numerous potential functions and prediction methods. The PDB, however, is highly biased: many proteins have more than one entry, while entire protein families are represented by a single structure, or even not at all. The standard solution to this problem is to limit the studies to non-redundant subsets of the PDB. While alleviating biases, this solution hides the many-to-many relations between sequences and structures. That is, non-redundant datasets conceal the diversity of sequences that share the same fold and the existence of multiple conformations for the same protein. A particularly disturbing aspect of non-redundant subsets is that they hardly benefit from the rapid pace of protein structure determination, as most newly solved structures fall within existing families. RESULTS: In this study we explore the concept of redundancy-weighted datasets, originally suggested by Miyazawa and Jernigan. Redundancy-weighted datasets include all available structures and associate them (or features thereof) with weights that are inversely proportional to the number of their homologs. Here, we provide the first systematic comparison of redundancy-weighted datasets with non-redundant ones. We test three weighting schemes and show that the distributions of structural features that they produce are smoother (having higher entropy) compared with the distributions inferred from non-redundant datasets. We further show that these smoothed distributions are both more robust and more correct than their non-redundant counterparts. We suggest that the better distributions, inferred using redundancy-weighting, may improve the accuracy of knowledge-based potentials and increase the power of protein structure prediction methods. Consequently, they may enhance model-driven molecular biology.
Chen Yanover, Natalia Vanetik, Michael Levitt 0001, Rachel Kolodny, Chen Keasar
Bioinform.2
2013 Mining the Gaps: Towards Polynomial Summarization
Marina Litvak, Natalia Vanetik
IJCNLP2
2011 Mining Graphs of Prescribed Connectivity
Natalia Vanetik
IC3K1
2009 Subsea: an efficient heuristic algorithm for subgraph isomorphism
Vladimir Lipets, Natalia Vanetik, Ehud Gudes
Data Min. Knowl. Discov.2
2006 Support measures for graph data
Natalia Vanetik, Solomon Eyal Shimony, Ehud Gudes
Data Min. Knowl. Discov.1
2006 Discovering Frequent Graph Patterns Using Disjoint Paths
abstract
Whereas data mining in structured data focuses on frequent data values, in semistructured and graph data mining, the issue is frequent labels and common specific topologies. The structure of the data is just as important as its content. We study the problem of discovering typical patterns of graph data, a task made difficult because of the complexity of required subtasks, especially subgraph isomorphism. In this paper, we propose a new apriori-based algorithm for mining graph data, where the basic building blocks are relatively large, disjoint paths. The algorithm is proven to be sound and complete. Empirical evidence shows practical advantages of our approach for certain categories of graphs
Ehud Gudes, Solomon Eyal Shimony, Natalia Vanetik
IEEE Trans. Knowl. Data Eng.3
2004 Mining Frequent Labeled and Partially Labeled Graph Patterns
abstract
Whereas data mining in structured data focuses on frequent data values, in semistructured and graph data the emphasis is on frequent labels and common topologies. Here, the structure of the data is just as important as its content. When data contains large amount of different labels, both fully labeled and partially labeled data may be useful. More informative patterns can be found in the database if some of the pattern nodes can be regarded as 'unlabeled'. We study the problem of discovering typical fully and partially labeled patterns of graph data. Discovered patterns are useful in many applications, including: compact representation of source information and a road-map for browsing and querying information sources.
Natalia Vanetik, Ehud Gudes
ICDE1
2002 Computing Frequent Graph Patterns from Semistructured Data
abstract
Whereas data mining in structured data focuses on frequent data values, in semistructured and graph data the emphasis is on frequent labels and common topologies. Here, the structure of the data is just as important as its content. We study the problem of discovering typical patterns of graph data. The discovered patterns can be useful for many applications, including: compact representation of source information and a road-map for browsing and querying information sources. Difficulties arise in the discovery task from the complexity of some of the required sub-tasks, such as sub-graph isomorphism. This paper proposes a new algorithm for mining graph data, based on a novel definition of support. Empirical evidence shows practical, as well as theoretical, advantages of our approach.
Natalia Vanetik, Ehud Gudes, Solomon Eyal Shimony
ICDM1