EDBT 2026 Demo / reviewers in the wild / expert
Yaacov Choueka
dblp:64/1694
· DBLP profile ↗
13ranked-venue papers
7as first author
0since 2021 · last 2014
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 6 · 5 first-authorArtificial intelligence and machine learning · 3Graphics, computer vision, multimedia, augmented reality and games · 3Theory of computation · 2 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
9 papers |
Information retrieval · 67% Data mining · 22% Data integration and cleaning · 11% | |
| Computer graphics and multimedia
1 paper |
Image and video processing · 100% |
Topics — the 14 heaviest of 15, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Information retrieval
retrieval models |
0.2 | 2 | 2014 | Congruency-Based Reranking · CVPR 2014 KEDMA - Linguistic Tools for Retrieval Systems · J. ACM 1978 |
Information retrieval
ranking |
0.2 | 1 | 2014 | Congruency-Based Reranking · CVPR 2014 |
Information retrieval
reranking |
0.2 | 1 | 2014 | Congruency-Based Reranking · CVPR 2014 |
Information retrieval
similarity search |
0.2 | 1 | 2014 | Congruency-Based Reranking · CVPR 2014 |
Data mining › clustering › interactive clustering
active clustering |
0.1 | 1 | 2011 | Active clustering of document fragments using information derived from both images and catalogs · ICCV 2011 |
Data mining
clustering |
0.1 | 1 | 2011 | Active clustering of document fragments using information derived from both images and catalogs · ICCV 2011 |
Image and video processing
document image analysis |
0.1 | 1 | 2011 | Active clustering of document fragments using information derived from both images and catalogs · ICCV 2011 |
Information retrieval › search engines
full-text search |
0.0 | 4 | 1988 | Compression of Concordances in Full-Text Retrieval Systems · SIGIR 1988 Improved Techniques for Processing Queries in Full-Text Systems · SIGIR 1987 Improved Hierarchical Bit-Vector Compression in Document Retrieval Systems · SIGIR 1986 |
Indexing and storage engines
bitmap index |
0.0 | 1 | 1987 | Improved Techniques for Processing Queries in Full-Text Systems · SIGIR 1987 |
Information retrieval
query processing |
0.0 | 1 | 1987 | Improved Techniques for Processing Queries in Full-Text Systems · SIGIR 1987 |
Information retrieval
text compression |
0.0 | 1 | 1985 | Efficient Variants of Huffman Codes in High Level Languages · SIGIR 1985 |
Coding theory › source coding › variable-length codes › prefix codes
huffman coding |
0.0 | 1 | 1985 | Efficient Variants of Huffman Codes in High Level Languages · SIGIR 1985 |
Information retrieval › text analysis › text preprocessing
morphological analysis |
0.0 | 1 | 1978 | KEDMA - Linguistic Tools for Retrieval Systems · J. ACM 1978 |
Information retrieval › query reformulation
query expansion |
0.0 | 1 | 1978 | KEDMA - Linguistic Tools for Retrieval Systems · J. ACM 1978 |
Methods — techniques the papers use, named apart from their topics
graphical model · 0.2active clustering · 0.2graphical bayesian model · 0.2clustering · 0.2prefix-omission compression · 0.0non-binary huffman codes · 0.0byte-per-byte decoding · 0.0hierarchical bit-vector compression · 0.0concordance merging · 0.0boolean query processing · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2014 | Congruency-Based RerankingabstractWe present a tool for re-ranking the results of a specific query by considering the (n+1) × (n+1) matrix of pairwise similarities among the elements of the set of n retrieved results and the query itself. The re-ranking thus makes use of the similarities between the various results and does not employ additional sources of information. The tool is based on graphical Bayesian models, which reinforce retrieved items strongly linked to other retrievals, and on repeated clustering to measure the stability of the obtained associations. The utility of the tool is demonstrated within the context of visual search of documents from the Cairo Genizah and for retrieval of paintings by the same artist and in the same style. Itai Ben-Shalom, Noga Levy, Lior Wolf, Nachum Dershowitz, Adiel Ben-Shalom, Roni Shweka, Yaacov Choueka, Tamir Hazan, Yaniv Bar |
CVPR | 7 |
| 2011 | Active clustering of document fragments using information derived from both images and catalogsabstractMany significant historical corpora contain leaves that are mixed up and no longer bound in their original state as multi-page documents. The reconstruction of old manuscripts from a mix of disjoint leaves can therefore be of paramount importance to historians and literary scholars. Previously, it was shown that visual similarity provides meaningful pair-wise similarities between handwritten leaves. Here, we go a step further and suggest a semiautomatic clustering tool that helps reconstruct the original documents. The proposed solution is based on a graphical model that makes inferences based on catalog information provided for each leaf as well as on the pairwise similarities of handwriting. Several novel active clustering techniques are explored, and the solution is applied to a significant part of the Cairo Genizah, where the problem of joining leaves remains unsolved even after a century of extensive study by hundreds of scholars. Lior Wolf, Lior Litwak, Nachum Dershowitz, Roni Shweka, Yaacov Choueka |
ICCV | 5 |
| 2011 | Computerized paleography: Tools for historical manuscriptsabstractThe Digital Age has brought with it large-scale digitization of historical records. The modern scholar of history or of other disciplines is often faced today with hundreds of thousands of readily-available and potentially-relevant full or fragmentary documents, but without computer aids that would make it possible to find the sought-after needles in the proverbial haystack of online images. The problems are even more acute when documents are handwritten, since optical character recognition does not provide quality results. We consider two tools: (1) a handwriting matching tool that is used to join together fragments of the same scribe, and (2) a paleographic classification tool that matches a given document to a large set of paleographic samples. Both tools are carefully designed not only to provide a high level of accuracy, but also to provide a clean and concise justification of the inferred results. This last requirement engenders challenges, such as sparsity of the representation, for which existing solutions are inappropriate for document analysis. Lior Wolf, Liza Potikha, Nachum Dershowitz, Roni Shweka, Yaacov Choueka |
ICIP | 5 |
| 2011 | Identifying Join Candidates in the Cairo Genizah
Lior Wolf, Rotem Littman, Naama Mayer, Tanya German, Nachum Dershowitz, Roni Shweka, Yaacov Choueka |
Int. J. Comput. Vis. | 7 |
| 1988 | Compression of Concordances in Full-Text Retrieval SystemsabstractThe concordance of a full-text information retrieval system contains for every different word W of the data base, a list L(W) of “coordinates”, each of which describes the exact location of an occurrence of W in the text. The concordance should be compressed, not only for the savings in storage space, but also in order to reduce the number of I/O operations, since the file is usually kept in secondary memory. Several methods are presented, which efficiently compress concordances of large fulltext retrieval systems. The methods were tested on the concordance of the Responsa Retrieval Project and yield savings of up to 49% relative to the non-compressed file; this is a relative improvement of about 27% over the currently used prefix-omission compression technique. Yaacov Choueka, Aviezri S. Fraenkel, Shmuel Tomi Klein |
SIGIR | 1 |
| 1987 | Improved Techniques for Processing Queries in Full-Text SystemsabstractIn static full-text retrieval systems, which accommodate metrical as well as Boolean operators, the traditional approach to query processing uses a “concordance”, from which large sets of coordinates are retrieved and then merged and/or collated. Alternatively, in a system with l documents, the concordance can be replaced by a set of bit-maps of fixed length l, which are constructed for every different word of the database and serve as occurrence maps. We propose to combine the concordance and bit-map approaches, and show how this can speed up the processing of queries: fast ANDing and ORing of the maps in a preprocessing stage, lead to large I/O savings in collating coordinates of keywords needed to satisfy the metrical and Boolean constraints. Moreover, the bit-maps give partial information on the distribution of the coordinates of the keywords, which can be used when queries must be processed by stages, due to their complexity and the sizes of the involved sets of coordinates. The new techniques are partially implemented at the Responsa Retrieval Project. Yaacov Choueka, Aviezri S. Fraenkel, Shmuel Tomi Klein, E. Segal |
SIGIR | 1 |
| 1986 | Improved Hierarchical Bit-Vector Compression in Document Retrieval SystemsabstractThe “concordance” of an information retrieval system can often be stored in form of bit-maps, which are usually very sparse and should be compressed. Hierarchical bit-vector compression consists of partitioning a vector vi into equi-sized blocks, constructing a new bit-vector vi+1 which points to the non-zero blocks in vi, dropping the zero-blocks of vi, and repeating the process for vi+1. We refine the method by pruning some of the tree branches if they ultimately point to very few documents; these document numbers are then added to an appended list which is compressed by the prefix-omission technique. The new method was thoroughly tested on the bit-maps of the Responsa Retrieval Project, and gave a relative improvement of about 40% over the conventional hierarchical compression method. Yaacov Choueka, Aviezri S. Fraenkel, Shmuel Tomi Klein, E. Segal |
SIGIR | 1 |
| 1985 | Efficient Variants of Huffman Codes in High Level LanguagesabstractAlthough it is well-known that Huffman Codes are optimal for text compression in a character-per-character encoding scheme, they are seldom used in practical situations since they require a bit-per-bit decoding algorithm, which has to be written in some assembly language, and will perform rather slowly. A number of methods are presented that avoid these difficulties. The decoding algorithms efficiently process the encoded string on a byte-per-byte basis, are faster than the original algorithm, and can be programmed in any high level language. This is achieved at the cost of storing some tables in the internal memory, but with no loss in the compression savings of the optimal Huffman codes. The internal memory space needed can be reduced either at the cost of increased processing time, or by using non-binary Huffman codes, which give sub-optimal compression. Experimental results for English and Hebrew text are also presented. Yaacov Choueka, Shmuel Tomi Klein, Yehoshua Perl |
SIGIR | 1 |
| 1982 | Processing truncated terms in document retrieval systems
Paul Bratley, Yaacov Choueka |
Inf. Process. Manag. | 2 |
| 1978 | KEDMA - Linguistic Tools for Retrieval SystemsabstractIn a full-text natural-language retrieval system, frequent need for automatic hngulst~c analysis arises, e.g for keyword expansion in a search process, content analysis, or automatic construction of concordances The avadablhty of sophisticated hngulstic tools, which is highly desirable for languages such as Enghsh, is quite imperative for, say, Semmc languages, whose complex morphological structure renders simple-minded and approximate soluuons such as suffix stripping totally useless.Sophisticated tools were designed and constructed via the fusion of grammatical analysis and grammatical synthesis, resulting in a set of global files which provide in some sense a complete grammatical and lexlcal description of the language These files induce a set of local files which adapt to the database at hand and permit flexible on-hne morphological analysis. R. Attar, Yaacov Choueka, Nachum Dershowitz, Aviezri S. Fraenkel |
J. ACM | 2 |
| 1978 | Finite Automata, Definable Sets, and Regular Expressions over omega^n-Tapes
Yaacov Choueka |
J. Comput. Syst. Sci. | 1 |
| 1974 | Theories of Automata on omega-Tapes: A Simplified Approach
Yaacov Choueka |
J. Comput. Syst. Sci. | 1 |
| 1971 | Full Text Document Retrieval: Hebrew Legal TextsabstractA full text retrieval system was designed for the responsa literature, which is a large corpus of Hebrew legal cases. The unique problems of the data base --- mixture of Hebrew, Aramaic and vernaculars, lack of vowels and punctuation, extreme language inflection problems, homographs, existence of thousands of grammatical variants of any given keyword --- dictated development of new methods. Among them we list "grammatical synthesis", which synthesizes all grammatical variants of a given keyword; "Compact KWIC", which enables the user to have a glimpse of the nature of the search before having performed it; effective citation index imbedded in full text searches; and, in general, extensive use of both positive and negative feedback within a single search run. A number of searches performed on a relatively small data base gave in each case a recall of 100%. The average precision was 34%. A KWIC of strategic portions of retrieved documents usually enables a quick disposal of non-relevant material. Yaacov Choueka, M. Cohen, J. Dueck, Aviezri S. Fraenkel, M. Slae |
SIGIR | 1 |