Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Yaacov Choueka

dblp:64/1694 · DBLP profile ↗
← Back
13ranked-venue papers
7as first author
0since 2021 · last 2014
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 6 · 5 first-authorArtificial intelligence and machine learning · 3Graphics, computer vision, multimedia, augmented reality and games · 3Theory of computation · 2 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
9 papers
Information retrieval · 67% Data mining · 22% Data integration and cleaning · 11%
Computer graphics and multimedia
1 paper
Image and video processing · 100%

Topics — the 14 heaviest of 15, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Information retrieval
retrieval models
0.222014
Congruency-Based Reranking · CVPR 2014
KEDMA - Linguistic Tools for Retrieval Systems · J. ACM 1978
Information retrieval
ranking
0.212014
Congruency-Based Reranking · CVPR 2014
Information retrieval
reranking
0.212014
Congruency-Based Reranking · CVPR 2014
Information retrieval
similarity search
0.212014
Congruency-Based Reranking · CVPR 2014
Data mining › clustering › interactive clustering
active clustering
0.112011
Active clustering of document fragments using information derived from both images and catalogs · ICCV 2011
Data mining
clustering
0.112011
Active clustering of document fragments using information derived from both images and catalogs · ICCV 2011
Image and video processing
document image analysis
0.112011
Active clustering of document fragments using information derived from both images and catalogs · ICCV 2011
Information retrieval › search engines
full-text search
0.041988
Compression of Concordances in Full-Text Retrieval Systems · SIGIR 1988
Improved Techniques for Processing Queries in Full-Text Systems · SIGIR 1987
Improved Hierarchical Bit-Vector Compression in Document Retrieval Systems · SIGIR 1986
Indexing and storage engines
bitmap index
0.011987
Improved Techniques for Processing Queries in Full-Text Systems · SIGIR 1987
Information retrieval
query processing
0.011987
Improved Techniques for Processing Queries in Full-Text Systems · SIGIR 1987
Information retrieval
text compression
0.011985
Efficient Variants of Huffman Codes in High Level Languages · SIGIR 1985
Coding theory › source coding › variable-length codes › prefix codes
huffman coding
0.011985
Efficient Variants of Huffman Codes in High Level Languages · SIGIR 1985
Information retrieval › text analysis › text preprocessing
morphological analysis
0.011978
KEDMA - Linguistic Tools for Retrieval Systems · J. ACM 1978
Information retrieval › query reformulation
query expansion
0.011978
KEDMA - Linguistic Tools for Retrieval Systems · J. ACM 1978

Methods — techniques the papers use, named apart from their topics

graphical model · 0.2active clustering · 0.2graphical bayesian model · 0.2clustering · 0.2prefix-omission compression · 0.0non-binary huffman codes · 0.0byte-per-byte decoding · 0.0hierarchical bit-vector compression · 0.0concordance merging · 0.0boolean query processing · 0.0
YearPublicationVenuePosition
2014 Congruency-Based Reranking
abstract
We present a tool for re-ranking the results of a specific query by considering the (n+1) × (n+1) matrix of pairwise similarities among the elements of the set of n retrieved results and the query itself. The re-ranking thus makes use of the similarities between the various results and does not employ additional sources of information. The tool is based on graphical Bayesian models, which reinforce retrieved items strongly linked to other retrievals, and on repeated clustering to measure the stability of the obtained associations. The utility of the tool is demonstrated within the context of visual search of documents from the Cairo Genizah and for retrieval of paintings by the same artist and in the same style.
Itai Ben-Shalom, Noga Levy, Lior Wolf, Nachum Dershowitz, Adiel Ben-Shalom, Roni Shweka, Yaacov Choueka, Tamir Hazan, Yaniv Bar
CVPR7
2011 Active clustering of document fragments using information derived from both images and catalogs
abstract
Many significant historical corpora contain leaves that are mixed up and no longer bound in their original state as multi-page documents. The reconstruction of old manuscripts from a mix of disjoint leaves can therefore be of paramount importance to historians and literary scholars. Previously, it was shown that visual similarity provides meaningful pair-wise similarities between handwritten leaves. Here, we go a step further and suggest a semiautomatic clustering tool that helps reconstruct the original documents. The proposed solution is based on a graphical model that makes inferences based on catalog information provided for each leaf as well as on the pairwise similarities of handwriting. Several novel active clustering techniques are explored, and the solution is applied to a significant part of the Cairo Genizah, where the problem of joining leaves remains unsolved even after a century of extensive study by hundreds of scholars.
Lior Wolf, Lior Litwak, Nachum Dershowitz, Roni Shweka, Yaacov Choueka
ICCV5
2011 Computerized paleography: Tools for historical manuscripts
abstract
The Digital Age has brought with it large-scale digitization of historical records. The modern scholar of history or of other disciplines is often faced today with hundreds of thousands of readily-available and potentially-relevant full or fragmentary documents, but without computer aids that would make it possible to find the sought-after needles in the proverbial haystack of online images. The problems are even more acute when documents are handwritten, since optical character recognition does not provide quality results. We consider two tools: (1) a handwriting matching tool that is used to join together fragments of the same scribe, and (2) a paleographic classification tool that matches a given document to a large set of paleographic samples. Both tools are carefully designed not only to provide a high level of accuracy, but also to provide a clean and concise justification of the inferred results. This last requirement engenders challenges, such as sparsity of the representation, for which existing solutions are inappropriate for document analysis.
Lior Wolf, Liza Potikha, Nachum Dershowitz, Roni Shweka, Yaacov Choueka
ICIP5
2011 Identifying Join Candidates in the Cairo Genizah
Lior Wolf, Rotem Littman, Naama Mayer, Tanya German, Nachum Dershowitz, Roni Shweka, Yaacov Choueka
Int. J. Comput. Vis.7
1988 Compression of Concordances in Full-Text Retrieval Systems
abstract
The concordance of a full-text information retrieval system contains for every different word W of the data base, a list L(W) of “coordinates”, each of which describes the exact location of an occurrence of W in the text. The concordance should be compressed, not only for the savings in storage space, but also in order to reduce the number of I/O operations, since the file is usually kept in secondary memory. Several methods are presented, which efficiently compress concordances of large fulltext retrieval systems. The methods were tested on the concordance of the Responsa Retrieval Project and yield savings of up to 49% relative to the non-compressed file; this is a relative improvement of about 27% over the currently used prefix-omission compression technique.
Yaacov Choueka, Aviezri S. Fraenkel, Shmuel Tomi Klein
SIGIR1
1987 Improved Techniques for Processing Queries in Full-Text Systems
abstract
In static full-text retrieval systems, which accommodate metrical as well as Boolean operators, the traditional approach to query processing uses a “concordance”, from which large sets of coordinates are retrieved and then merged and/or collated. Alternatively, in a system with l documents, the concordance can be replaced by a set of bit-maps of fixed length l, which are constructed for every different word of the database and serve as occurrence maps. We propose to combine the concordance and bit-map approaches, and show how this can speed up the processing of queries: fast ANDing and ORing of the maps in a preprocessing stage, lead to large I/O savings in collating coordinates of keywords needed to satisfy the metrical and Boolean constraints. Moreover, the bit-maps give partial information on the distribution of the coordinates of the keywords, which can be used when queries must be processed by stages, due to their complexity and the sizes of the involved sets of coordinates. The new techniques are partially implemented at the Responsa Retrieval Project.
Yaacov Choueka, Aviezri S. Fraenkel, Shmuel Tomi Klein, E. Segal
SIGIR1
1986 Improved Hierarchical Bit-Vector Compression in Document Retrieval Systems
abstract
The “concordance” of an information retrieval system can often be stored in form of bit-maps, which are usually very sparse and should be compressed. Hierarchical bit-vector compression consists of partitioning a vector vi into equi-sized blocks, constructing a new bit-vector vi+1 which points to the non-zero blocks in vi, dropping the zero-blocks of vi, and repeating the process for vi+1. We refine the method by pruning some of the tree branches if they ultimately point to very few documents; these document numbers are then added to an appended list which is compressed by the prefix-omission technique. The new method was thoroughly tested on the bit-maps of the Responsa Retrieval Project, and gave a relative improvement of about 40% over the conventional hierarchical compression method.
Yaacov Choueka, Aviezri S. Fraenkel, Shmuel Tomi Klein, E. Segal
SIGIR1
1985 Efficient Variants of Huffman Codes in High Level Languages
abstract
Although it is well-known that Huffman Codes are optimal for text compression in a character-per-character encoding scheme, they are seldom used in practical situations since they require a bit-per-bit decoding algorithm, which has to be written in some assembly language, and will perform rather slowly. A number of methods are presented that avoid these difficulties. The decoding algorithms efficiently process the encoded string on a byte-per-byte basis, are faster than the original algorithm, and can be programmed in any high level language. This is achieved at the cost of storing some tables in the internal memory, but with no loss in the compression savings of the optimal Huffman codes. The internal memory space needed can be reduced either at the cost of increased processing time, or by using non-binary Huffman codes, which give sub-optimal compression. Experimental results for English and Hebrew text are also presented.
Yaacov Choueka, Shmuel Tomi Klein, Yehoshua Perl
SIGIR1
1982 Processing truncated terms in document retrieval systems
Paul Bratley, Yaacov Choueka
Inf. Process. Manag.2
1978 KEDMA - Linguistic Tools for Retrieval Systems
abstract
In a full-text natural-language retrieval system, frequent need for automatic hngulst~c analysis arises, e.g for keyword expansion in a search process, content analysis, or automatic construction of concordances The avadablhty of sophisticated hngulstic tools, which is highly desirable for languages such as Enghsh, is quite imperative for, say, Semmc languages, whose complex morphological structure renders simple-minded and approximate soluuons such as suffix stripping totally useless.Sophisticated tools were designed and constructed via the fusion of grammatical analysis and grammatical synthesis, resulting in a set of global files which provide in some sense a complete grammatical and lexlcal description of the language These files induce a set of local files which adapt to the database at hand and permit flexible on-hne morphological analysis.
R. Attar, Yaacov Choueka, Nachum Dershowitz, Aviezri S. Fraenkel
J. ACM2
1978 Finite Automata, Definable Sets, and Regular Expressions over omega^n-Tapes
Yaacov Choueka
J. Comput. Syst. Sci.1
1974 Theories of Automata on omega-Tapes: A Simplified Approach
Yaacov Choueka
J. Comput. Syst. Sci.1
1971 Full Text Document Retrieval: Hebrew Legal Texts
abstract
A full text retrieval system was designed for the responsa literature, which is a large corpus of Hebrew legal cases. The unique problems of the data base --- mixture of Hebrew, Aramaic and vernaculars, lack of vowels and punctuation, extreme language inflection problems, homographs, existence of thousands of grammatical variants of any given keyword --- dictated development of new methods. Among them we list "grammatical synthesis", which synthesizes all grammatical variants of a given keyword; "Compact KWIC", which enables the user to have a glimpse of the nature of the search before having performed it; effective citation index imbedded in full text searches; and, in general, extensive use of both positive and negative feedback within a single search run. A number of searches performed on a relatively small data base gave in each case a recall of 100%. The average precision was 34%. A KWIC of strategic portions of retrieved documents usually enables a quick disposal of non-relevant material.
Yaacov Choueka, M. Cohen, J. Dueck, Aviezri S. Fraenkel, M. Slae
SIGIR1