EDBT 2026 Demo / reviewers in the wild / expert
Utku Irmak
dblp:48/2573
· DBLP profile ↗
9ranked-venue papers
7as first author
0since 2021 · last 2010
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 6 · 5 first-authorApplied, interdisciplinary, general and emerging computing · 4 · 4 first-authorComputer networks · 2 · 1 first-authorArtificial intelligence and machine learning · 1Systems, architecture and hardware · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
5 papers |
Information retrieval · 51% Data stream processing · 16% Machine learning and data management · 14% | |
| Computer architecture, parallel and distributed computing, and storage systems
3 papers |
Storage systems · 63% Memory systems · 25% Distributed systems · 12% | |
| Artificial intelligence
2 papers |
Information extraction and text analysis · 64% Knowledge representation and reasoning · 36% | |
| Computer networks
1 paper |
Content delivery and video streaming · 77% Internet architecture and protocols · 23% |
Topics — the 20 heaviest of 23, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Storage systems › file systems
file synchronization |
0.1 | 2 | 2008 | Algorithms for Low-Latency Remote File Synchronization · INFOCOM 2008 Improved single-round protocols for remote file synchronization · INFOCOM 2005 |
Natural language and speech › Information extraction and text analysis
named entity recognition |
0.1 | 1 | 2010 | A scalable machine-learning approach for semi-structured named entity recognition · WWW 2010 |
Machine learning and data management
scalable machine learning |
0.1 | 1 | 2010 | A scalable machine-learning approach for semi-structured named entity recognition · WWW 2010 |
Information retrieval › ranking
keyword ranking |
0.1 | 1 | 2009 | Contextual Ranking of Keywords Using Click Data · ICDE 2009 |
Information retrieval
ranking |
0.1 | 1 | 2009 | Contextual Ranking of Keywords Using Click Data · ICDE 2009 |
Knowledge, reasoning and agents › Knowledge representation and reasoning › knowledge acquisition › knowledge extraction
structured information extraction |
0.1 | 1 | 2006 | Interactive wrapper generation with minimal user effort · WWW 2006 |
Information retrieval
query processing |
0.1 | 1 | 2006 | Efficient query subscription processing for prospective search engines · WWW 2006 |
Data stream processing › publish/subscribe
query subscription processing |
0.1 | 1 | 2006 | Efficient Query Subscription Processing for Prospective Search Engines · USENIX ATC, General Track 2006 |
Information retrieval
search engines |
0.1 | 1 | 2006 | Efficient Query Subscription Processing for Prospective Search Engines · USENIX ATC, General Track 2006 |
Data integration and cleaning › data extraction › web data extraction
wrapper generation |
0.1 | 1 | 2006 | Interactive wrapper generation with minimal user effort · WWW 2006 |
Content delivery and video streaming
web content delivery |
0.1 | 1 | 2005 | Hierarchical substring caching for efficient content distribution to low-bandwidth clients · WWW 2005 |
Memory systems
cache |
0.1 | 1 | 2005 | Hierarchical substring caching for efficient content distribution to low-bandwidth clients · WWW 2005 |
Data integration and cleaning › data extraction
structured data extraction |
0.0 | 1 | 2010 | A scalable machine-learning approach for semi-structured named entity recognition · WWW 2010 |
Information retrieval › user behavior › search behavior
click model |
0.0 | 1 | 2009 | Contextual Ranking of Keywords Using Click Data · ICDE 2009 |
Information retrieval › query log analysis
clickthrough data |
0.0 | 1 | 2009 | Contextual Ranking of Keywords Using Click Data · ICDE 2009 |
Data mining › text mining › information extraction
entity extraction |
0.0 | 1 | 2009 | Contextual Ranking of Keywords Using Click Data · ICDE 2009 |
Distributed systems
replication |
0.0 | 1 | 2008 | Algorithms for Low-Latency Remote File Synchronization · INFOCOM 2008 |
Information retrieval
web search |
0.0 | 1 | 2006 | Efficient query subscription processing for prospective search engines · WWW 2006 |
Coding theory › error-correcting codes
erasure coding |
0.0 | 1 | 2005 | Improved single-round protocols for remote file synchronization · INFOCOM 2005 |
Coding theory › constrained coding › synchronization
file synchronization |
0.0 | 1 | 2005 | Improved single-round protocols for remote file synchronization · INFOCOM 2005 |
Methods — techniques the papers use, named apart from their topics
regular expressions · 0.2machine learning · 0.2training interface design · 0.1trace-based evaluation · 0.1erasure codes · 0.1delta compression · 0.1feature engineering · 0.1click data · 0.1set reconciliation · 0.1sampling · 0.1ranking algorithms · 0.1ranking algorithm · 0.1query matching algorithms · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2010 | A scalable machine-learning approach for semi-structured named entity recognitionabstractNamed entity recognition studies the problem of locating and classifying parts of free text into a set of predefined categories. Although extensive research has focused on the detection of person, location and organization entities, there are many other entities of interest, including phone numbers, dates, times and currencies (to name a few examples). We refer to these types of entities as "semi-structured named entities", since they usually follow certain syntactic formats according to some conventions, although their structure is typically not well-defined. Regular expression solutions require significant amount of manual effort and supervised machine learning approaches rely on large sets of labeled training data. Therefore, these approaches do not scale when we need to support many semi-structured entity types in many languages and regions. Utku Irmak, Reiner Kraft 0001 |
WWW | 1 |
| 2009 | Contextual Ranking of Keywords Using Click DataabstractThe problem of automatically extracting the most interesting and relevant keyword phrases in a document has been studied extensively as it is crucial for a number of applications. These applications include contextual advertising, automatic text summarization, and user-centric entity detection systems. All these applications can potentially benefit from a successful solution as it enables computational efficiency (by decreasing the input size), noise reduction, or overall improved user satisfaction.In this paper, we study this problem and focus on improving the overall quality of user-centric entity detection systems. First, we review our concept extraction technique, which relies on search engine query logs. We then define a new feature space to represent the interesting ness of concepts, and describe a new approach to estimate their relevancy for a given context. We utilize click through data obtained from a large scale user-centric entity detection system - Contextual Shortcuts - to train a model to rank the extracted concepts, and evaluate the resulting model extensively again based on their click through data. Our results show that the learned model outperforms the baseline model, which employs similar features but whose weights are tuned carefully based on empirical observations, and reduces the error rate from 30.22% to 18.66%. Utku Irmak, Vadim von Brzeski, Reiner Kraft 0001 |
ICDE | 1 |
| 2008 | Algorithms for Low-Latency Remote File SynchronizationabstractThe remote file synchronization problem is how to update an outdated version of a file located on one machine to the current version located on another machine with a minimal amount of network communication. It arises in many scenarios including Web site mirroring, file system backup and replication, or web access over slow links. A widely used open-source tool called rsync uses a single round of messages to solve this problem (plus an initial round for exchanging meta information). While research has shown that significant additional savings in bandwidth are possible by using multiple rounds, such approaches are often not desirable due to network latencies, increased protocol complexity, and higher I/O and CPU overheads at the endpoints. We study single-round synchronization techniques that achieve savings in bandwidth consumption while preserving many of the advantages of the rsync approach. In particular, we propose a new and simple algorithm for file synchronization based on set reconciliation techniques. We then show how to integrate sampling techniques into our approach in order to adaptively select the most suitable algorithm and parameter setting for a given data set. Experimental results on several data sets show that the resulting protocol gives significant benefits over rsync, particularly on data sets with high degrees of redundancy between the versions. Utku Irmak, Torsten Suel |
INFOCOM | 2 |
| 2007 | Leveraging context in user-centric entity detection systemsabstractA user-centric entity detection system is one in which the primary consumer of the detected entities is a person who can perform actions on the detected entities (e.g. perform a search, view a map, shop, etc.). We contrast this with machine-centric detection systems where the primary consumer of the detected entities is a machine. Machine-centric detection systems typically focus on the quantity of detected entities, measured by precision and recall metrics, with the goal of correctly identifying every single entity in a document. Vadim von Brzeski, Utku Irmak, Reiner Kraft 0001 |
CIKM | 2 |
| 2006 | Efficient Query Subscription Processing for Prospective Search Engines
Utku Irmak, Svilen Mihaylov, Torsten Suel, Samrat Ganguly, Rauf Izmailov |
USENIX ATC, General Track | 1 |
| 2006 | Efficient query subscription processing for prospective search enginesabstractCurrent web search engines are retrospective in that they limit users to searches against already existing pages. Prospective search engines, on the other hand, allow users to upload queries that will be applied to newly discovered pages in the future. We study and compare algorithms for efficiently matching large numbers of simple keyword queries against a stream of newly discovered pages. Utku Irmak, Svilen Mihaylov, Torsten Suel, Samrat Ganguly, Rauf Izmailov |
WWW | 1 |
| 2006 | Interactive wrapper generation with minimal user effortabstractWhile much of the data on the web is unstructured in nature, there is also a significant amount of embedded structured data, such as product information on e-commerce sites or stock data on financial sites. A large amount of research has focused on the problem of generating wrappers, i.e., software tools that allow easy and robust extraction of structured data from text and HTML sources. In many applications, such as comparison shopping, data has to be extracted from many different sources, making manual coding of a wrapper for each source impractical. On the other hand, fully automatic approaches are often not reliable enough, resulting in low quality of the extracted data.We describe a complete system for semi-automatic wrapper generation that can be trained on different data sources in a simple interactive manner. Our goal is to minimize the amount of user effort for training reliable wrappers through design of a suitable training interface that is implemented based on a powerful underlying extraction language and a set of training and ranking algorithms. Our experiments show that our system achieves reliable extraction with a very small amount of user effort. Utku Irmak, Torsten Suel |
WWW | 1 |
| 2005 | Improved single-round protocols for remote file synchronizationabstractGiven two versions of a file, a current version located on one machine and an outdated version known only to another machine, the remote file synchronization problem is how to update the outdated version over a network with a minimal amount of communication. In particular, when the versions are very similar, the total data transmitted should be significantly smaller than the file size. File synchronization problems arise in many application scenarios such as Web site mirroring, file system backup and replication, and Web access over slow links. An open source tool for this problem, called rsync and included in many Linux distributions, is widely used in such scenarios, rsync uses a single round of messages between the two machines. While recent research has shown that significant additional savings in bandwidth consumption are possible through the use of optimized multi-round protocols, there are many scenarios where multiple rounds are undesirable. In this paper, we study single-round protocols for file synchronization that offer significant improvements over rsync. Our main contribution is a new approach to file synchronization based on the use of erasure codes. Using this approach, we design a single-round protocol that is provably efficient with respect to common measures of file distance, and another optimized practical protocol that shows promising improvements over rsync on our data sets. In addition, we show how to obtain moderate improvements by engineering the rsync approach. Utku Irmak, Svilen Mihaylov, Torsten Suel |
INFOCOM | 1 |
| 2005 | Hierarchical substring caching for efficient content distribution to low-bandwidth clientsabstractWhile overall bandwidth in the internet has grown rapidly over the last few years, and an increasing number of clients enjoy broadband connectivity, many others still access the internet over much slower dialup or wireless links. To address this issue, a number of techniques for optimized delivery of web and multimedia content over slow links have been proposed, including protocol optimizations, caching, compression, and multimedia transcoding, and several large ISPs have recently begun to widely promote dialup acceleration services based on such techniques. A recent paper by Rhea, Liang, and Brewer proposed an elegant technique called value-based caching that caches substrings of files, rather than entire files, and thus avoids repeated transmission of substrings common to several pages or page versions. We propose and study a hierarchical substring caching technique that provides significant savings over this basic approach. We describe several additional techniques for minimizing overheads and perform an evaluation on a large set of real web access traces that we collected. In the second part of our work, we compare our approach to a widely studied alternative approach based on delta compression, and show how to integrate the two for best overall performance. The studied techniques are typically employed in a clientproxy environment, with each proxy serving a large number of clients, and an important aspect is how to conserve resources on the proxy while exploiting the significant memory and CPU power available on current clients. Utku Irmak, Torsten Suel |
WWW | 1 |