Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Dafna Sheinwald

dblp:98/6723 · DBLP profile ↗
← Back
17ranked-venue papers
6as first author
2since 2021 · last 2023
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 8 · 2 first-authorArtificial intelligence and machine learning · 5 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 1 since 2021Theory of computation · 4 · 2 first-authorSystems, architecture and hardware · 1Computer networks · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Language models and text generation · 75% Information extraction and text analysis · 25%
Databases, data mining, and information retrieval
3 papers
Information retrieval · 74% Machine learning and data management · 18% Knowledge graphs · 7%
Theoretical computer science
5 papers
Coding theory · 82% Automata and formal languages · 10% Algorithms and data structures · 8%

Topics — the 19 heaviest of 22, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation
text generation
1.322023
Active Learning for Natural Language Generation · EMNLP 2023
nBIIG: A Neural BI Insights Generation System for Table Reporting · AAAI 2023
Natural language and speech › Language models and text generation › text generation › data-to-text generation
table-to-text generation
0.712023
nBIIG: A Neural BI Insights Generation System for Table Reporting · AAAI 2023
Information retrieval › document processing › document analysis
document representation
0.212016
Semantic Documents Relatedness using Concept Graph Representation · WSDM 2016
Information retrieval › similarity measure
document similarity
0.212016
Semantic Documents Relatedness using Concept Graph Representation · WSDM 2016
Information retrieval › text analysis
semantic relatedness
0.212016
Semantic Documents Relatedness using Concept Graph Representation · WSDM 2016
Machine learning and data management
active learning
0.212023
Active Learning for Natural Language Generation · EMNLP 2023
Information retrieval › search interfaces
faceted search
0.112008
Beyond basic faceted search · WSDM 2008
Knowledge graphs
knowledge base linking
0.112016
Semantic Documents Relatedness using Concept Graph Representation · WSDM 2016
Transport protocols and congestion control
transport protocols
0.112005
Out of Order Incremental CRC Computation · IEEE Trans. Computers 2005
Coding theory › error-correcting codes › error detection
cyclic redundancy check
0.112005
Out of Order Incremental CRC Computation · IEEE Trans. Computers 2005
Query processing and optimization
OLAP
0.012008
Beyond basic faceted search · WSDM 2008
Coding theory
source coding
0.021994
On the Ziv-Lempel proof and related topics · Proc. IEEE 1994
Two-dimensional encoding by finite-state encoders · IEEE Trans. Commun. 1990
Internet architecture and protocols › protocol implementation
protocol processing
0.012005
Out of Order Incremental CRC Computation · IEEE Trans. Computers 2005
Coding theory › source coding
lempel-ziv compression
0.011994
On the Ziv-Lempel proof and related topics · Proc. IEEE 1994
Image and video coding › scalable coding
progressive coding
0.011993
Deterministic prediction in progressive coding · IEEE Trans. Inf. Theory 1993
Coding theory › constrained coding
finite-state encoders
0.011990
Two-dimensional encoding by finite-state encoders · IEEE Trans. Commun. 1990
Coding theory › source coding
lossless compression
0.011990
Two-dimensional encoding by finite-state encoders · IEEE Trans. Commun. 1990
Coding theory › error-correcting codes
two-dimensional codes
0.011990
Two-dimensional encoding by finite-state encoders · IEEE Trans. Commun. 1990
Image and video coding
image compression
0.011990
Two-dimensional encoding by finite-state encoders · IEEE Trans. Commun. 1990

Methods — techniques the papers use, named apart from their topics

instruction tuning · 1.3active learning · 1.3neural network · 0.7neural network embedding · 0.2graph similarity · 0.2closeness centrality · 0.2incremental CRC computation · 0.1tree indexing · 0.1depth-first tree search · 0.0universal coding · 0.0finite-state machine · 0.0finite state machine · 0.0
YearPublicationVenuePosition
2023 nBIIG: A Neural BI Insights Generation System for Table Reporting
Yotam Perlitz, Dafna Sheinwald, Noam Slonim, Michal Shmueli-Scheuer
AAAI2
2023 Active Learning for Natural Language Generation
abstract
The field of Natural Language Generation (NLG) suffers from a severe shortage of labeled data due to the extremely expensive and timeconsuming process involved in manual annotation.A natural approach for coping with this problem is active learning (AL), a well-known machine learning technique for improving annotation efficiency by selectively choosing the most informative examples to label.However, while AL has been well-researched in the context of text classification, its application to NLG remains largely unexplored.In this paper, we present a first systematic study of active learning for NLG, considering a diverse set of tasks and multiple leading selection strategies, and harnessing a strong instruction-tuned model.Our results indicate that the performance of existing AL strategies is inconsistent, surpassing the baseline of random example selection in some cases but not in others.We highlight some notable differences between the classification and generation scenarios, and analyze the selection behaviors of existing AL strategies.Our findings motivate exploring novel approaches for applying AL to generation tasks.
Yotam Perlitz, Ariel Gera, Michal Shmueli-Scheuer, Dafna Sheinwald, Noam Slonim, Liat Ein-Dor
EMNLP4
2016 Semantic Documents Relatedness using Concept Graph Representation
abstract
We deal with the problem of document representation for the task of measuring semantic relatedness between documents. A document is represented as a compact concept graph where nodes represent concepts extracted from the document through references to entities in a knowledge base such as DBpedia. Edges represent the semantic and structural relationships among the concepts. Several methods are presented to measure the strength of those relationships. Concepts are weighted through the concept graph using closeness centrality measure which reflects their relevance to the aspects of the document. A novel similarity measure between two concept graphs is presented. The similarity measure first represents concepts as continuous vectors by means of neural networks. Second, the continuous vectors are used to accumulate pairwise similarity between pairs of concepts while considering their assigned weights. We evaluate our method on a standard benchmark for document similarity. Our method outperforms state-of-the-art methods including ESA (Explicit Semantic Annotation) while our concept graphs are much smaller than the concept vectors generated by ESA. Moreover, we show that by combining our concept graph with ESA, we obtain an even further improvement.
Yuan Ni, Qiongkai Xu, Yosi Mass, Dafna Sheinwald, Huijia Zhu, Shao Sheng Cao
WSDM5
2012 On demand string sorting over unbounded alphabets
Carmel Kent, Moshe Lewenstein, Dafna Sheinwald
Theor. Comput. Sci.3
2008 Beyond basic faceted search
abstract
This paper extends traditional faceted search to support richer information discovery tasks over more complex data models. Our first extension adds exible, dynamic business intelligence aggregations to the faceted application, enabling users to gain insight into their data that is far richer than just knowing the quantities of documents belonging to each facet. We see this capability as a step toward bringing OLAP capabilities, traditionally supported by databases over relational data, to the domain of free-text queries over metadata-rich content. Our second extension shows how one can efficiently extend a faceted search engine to support correlated facets - a more complex information model in which the values associated with a document across multiple facets are not independent. We show that by reducing the problem to a recently solved tree-indexing scenario, data with correlated facets can be efficiently indexed and retrieved
Ori Ben-Yitzhak, Nadav Golbandi, Nadav Har'El, Ronny Lempel, Andreas Neumann 0001, Shila Ofek-Koifman, Dafna Sheinwald, Eugene J. Shekita, Benjamin Sznajder, Sivan Yogev
WSDM7
2007 Just in time indexing for up to the second search
abstract
E-commerce and intranet search systems require newly arriving content to be indexed and made available for search within minutes or hours of arrival. Applications such as file system and email search demand even faster turnaround from search systems, requiring new content to become available for search almost instantaneously. However, incrementally updating inverted indices, which are the predominant datastructure used in search engines, is an expensive operation that most systems avoid performing at high rates.
Ronny Lempel, Yosi Mass, Shila Ofek-Koifman, Dafna Sheinwald, Yael Petruschka, Ron Sivan
CIKM4
2007 On Demand String Sorting over Unbounded Alphabets
Carmel Kent, Moshe Lewenstein, Dafna Sheinwald
CPM3
2005 Out of Order Incremental CRC Computation
abstract
We consider a communication protocol where the sender breaks a message, comprised of an information block and a corresponding CRC, into small segments and transmits these segments separately, possibly via different routes, to the receiver. Traditionally, reversing the sender operations, the receiver first assembles all the segments that make up the message, then computes a CRC based on the information part of the message and verifies it against the arriving CRC, and, finally, delivers the information part on to the upper layer protocol (ULP). We present an incremental CRC computation whereby each arriving segment contributes its share to the message's CRC upon arrival, independently of other segment arrivals, and can thus proceed immediately to the ULP. We impose no constraint on the order of segment arrivals nor on their sizes. Yet, in its time complexity, our scheme does not exceed the traditional computation, which assembles all the segments first, and it uses only a fixed and very small amount of extra memory. Our scheme is beneficial when the ULP can process the individual segments without reading the entire message first and can revert, if needed, to the state it had prior to the processing of any segment of that message. One practical example application is the evolving protocol for remote direct memory access (RDMA) over TCP, where an overall CRC is added to a concatenation of data segments.
Julian Satran, Dafna Sheinwald, Ilan Shimony
IEEE Trans. Computers2
2001 Software Compression in the Client/Server Environment
abstract
Lempel-Ziv (1977) based compression algorithms are universal, not assuming any prior knowledge of the file to be compressed or its statistics. Accordingly, the reference dictionary of these textual substitution compression algorithms includes only segments of the already-processed portion of the file. It is often the case, though, that both, compressor and decompressor, even when they reside on different sites, share knowledge of files (e.g., devices managed by a server, or software customers holding older releases of products). For such cases, we suggest the addition of shared files to the reference dictionary. Preferably, files to be included are those which resemble the file to be compressed. Such an extension of the reference dictionary lengthens the matches found while compressing the file, and thus lessens the number of matches needed to cover the file. We found that with a careful selection (which can be automated) of the shared files to be included, the advantage of the decrease in the number of matches overwhelms the disadvantage of the increase in the number of bits needed to express the index of each match in the extended dictionary. Altogether, compression attainable by our proposed scheme can be significantly better than with the original Lempel-Ziv dictionary. Maintaining and searching a dictionary that is much larger than the original Lempel-Ziv dictionary demand strong computational resources and suits off-line more than on-line compression. We thus conclude that in the client/server environment, where shared files commonly exist, and the server enjoys extensive computational resources, the scheme suggested is advantageous for transferring files from the server to its clients.
Michael Factor, Dafna Sheinwald, Ben-Ami Yassour
Data Compression Conference2
2001 Compression in the presence of shared data
Michael Factor, Dafna Sheinwald
Inf. Sci.2
1995 On Encoding and Decoding with Two-Way Head Machines
Dafna Sheinwald, Abraham Lempel, Jacob Ziv
Inf. Comput.1
1994 On the Ziv-Lempel proof and related topics
abstract
Results concerning the celebrated Ziv-Lempel sequence compression algorithm are revisited taking a rather intuitive approach. Also presented are ideas, which were previously formalized and extensions of these results to compression of two-dimensional data.>
Dafna Sheinwald
Proc. IEEE1
1993 A Simple Linear-Time Algorithm for the Recognition of Bandwidth-2 Biconnected Graphs
Fillia Makedon, Dafna Sheinwald, Yaron Wolfsthal
Inf. Process. Lett.2
1993 Deterministic prediction in progressive coding
abstract
Deterministic prediction in progressive coding of images is investigated. Progressive coding first creates a sequence of resolution layers by beginning with an original image and reducing its resolution several times by factors of some natural number M. The resultant layers are losslessly encoded, beginning with the lowest-resolution layer and, then encoding each higher resolution image incrementally upon the previous one. Coding efficiency may be improved if knowledge of the rules which produced the lower-resolution image of each pair is used to deterministically predict pixels of the higher, so they need not be encoded. Given reduction rules expressing each low-resolution pixel as a function of nearby high-resolution pixels and previously generated low-resolution pixels, it is shown that finding a complete set of rules, each of which deterministically predicts the value of a high-resolution pixel when certain values are found in nearby low-resolution pixels and previously coded high-resolution pixels, is NP-complete. A recursive algorithm for solving the problem in optimal time as a depth-first tree search is proposed, and the characteristics of the resultant prediction process are studied.>
Dafna Sheinwald, Richard C. Pasco
IEEE Trans. Inf. Theory1
1992 On Binary Alphabetical Codes
abstract
Binary alphabetical codes, which are prefix free, fixed-to-variable binary codes for discrete memoryless sources, in which the lexicographic order of the codewords agrees with the alphabet order of the respective source letters, are studied. A necessary and sufficient condition on the sequence of codeword length of any such code is proved. A new upper bounds on the redundancy of alphabetical codes relative to the optimal prefix free, fixed-to-variables codes-the Huffman codes-is proved. An adaptation of the Ziv-Lempel algorithm making it lexicographic order preserving, without any additional redundancy, is presented.>
Dafna Sheinwald
Data Compression Conference1
1991 On Compression with Two-Way Head Machines
abstract
Motivated by the study of various kinds of machines as recognizers of formal languages, the authors compare the encoding and decoding power of finite state sequential machines and extensions thereof. They show that, with a forward moving head, the best compression achievable for a given sequence, to be decoded by a finite state decoder, is the same as the best ratio attainable for that sequence when encoded by a finite state information lossless encoder. They cannot gain in compression by allowing a finite state encoder to move its head back and forth on an input sequence, even if the decoder has unrestricted power. However, better compression can be achieved for specific infinite sequences using an unrestricted encoder and a two-way finite state decoder.>
Dafna Sheinwald, Abraham Lempel, Jacob Ziv
Data Compression Conference1
1990 Two-dimensional encoding by finite-state encoders
abstract
Distortion-free compressibility of individual pictures by finite-state encoders is investigated. In a recent paper (see IEEE Trans. Inform. Theory, vol.32, no.1, p.1-8, 1986) the compressibility of a given picture I was defined and shown to be the asymptotically attainable lower bound on the compression ratio that can be achieved for I by any finite-state encoder. Here, a different and more direct approach is taken to prove similar results, which are summarized in a converse-to-coding theorem and a constructive-coding-theorem that leads to a universal asymptotically optimal compression algorithm.>
Dafna Sheinwald, Abraham Lempel, Jacob Ziv
IEEE Trans. Commun.1