VLDB 2026 Research / reviewers in the wild / expert
Can Özbey
dblp:186/7058
· DBLP profile ↗
2ranked-venue papers
2as first author
2since 2021 · last 2023
0009-0005-8432-9413ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
1 paper |
Information retrieval · 100% |
Topics — the 2 heaviest of 2, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Information retrieval › text analysis › text corpus analysis
word frequency distribution |
0.7 | 1 | 2023 | A Unified Formulation for the Frequency Distribution of Word Frequencies using the Inverse Zipf's Law · SIGIR 2023 |
Information retrieval › retrieval models
language model |
0.2 | 1 | 2023 | A Unified Formulation for the Frequency Distribution of Word Frequencies using the Inverse Zipf's Law · SIGIR 2023 |
Methods — techniques the papers use, named apart from their topics
relative entropy minimization · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | A Unified Formulation for the Frequency Distribution of Word Frequencies using the Inverse Zipf's LawabstractThe power-law approximation for the frequency distribution of words postulated by Zipf has been extensively studied for decades, which led to many variations on the theme. However, comparatively less attention has been paid to the investigation of the case of word frequencies. In this paper, we derive its analytical expression from the inverse of the underlying rank-size distribution as a function of total word count, vocabulary size and the shape parameter, thereby providing a unified framework to explain the nonlinear behavior of low frequencies on the log-log scale. We also present an efficient method based on relative entropy minimization for a robust estimation of the shape parameter using a small number of empirical low-frequency probabilities. Experiments were carried out for a selected set of languages with varying degrees of inflection in order to demonstrate the effectiveness of the proposed approach. Can Özbey, Talha Çolakoglu, M. Safak Bilici, Ekin Can Erkus |
SIGIR | 1 |
| 2022 | Joint Compression of Document Identifiers and Term Frequencies via Dense Unary CodesabstractCompressing posting lists in large-scale search engines for efficiency improvement conventionally involves encoding document identifiers and term frequencies separately. In this work, we adopt dense unary codes in favor of joint encoding and evaluate them in terms of compression ratio along with two other techniques on eight document collections having different characteristics as to term frequency distribution. As a result, it has been observed that dense unary codes yield higher compression ratios particularly in short-text collections, where it is more unlikely for sparse terms to have high frequency values per document, as well as retaining a simple decoding mechanism. Can Özbey |
INISTA | 1 |