VLDB 2026 Research / reviewers in the wild / expert
Roelof van Zwol
dblp:22/3616
· DBLP profile ↗
32ranked-venue papers
11as first author
1since 2021 · last 2021
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 24 · 10 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 1 first-authorArtificial intelligence and machine learning · 9 · 4 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-authorComputer networks · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
15 papers |
Information retrieval · 73% Recommender systems · 12% Web and social media mining · 12% | |
| Computer graphics and multimedia
4 papers |
Multimedia analysis and retrieval · 100% |
Topics — the 29 heaviest of 32, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Information retrieval
retrieval models |
0.3 | 4 | 2012 | Spatially-aware indexing for image object retrieval · WSDM 2012 Faceted exploration of image search results · WWW 2010 Resolving tag ambiguity · ACM Multimedia 2008 |
Information retrieval
image retrieval |
0.3 | 3 | 2010 | Faceted exploration of image search results · WWW 2010 Visual diversification of image search results · WWW 2009 Boosting image retrieval through aggregating search results based on visual annotations · ACM Multimedia 2008 |
Web and social media mining
social media analysis |
0.2 | 3 | 2010 | Prediction of favourite photos using social, visual, and textual signals · ACM Multimedia 2010 WSM'10: 2nd ACM workshop on social media · ACM Multimedia 2010 Flickr tag recommendation based on collective knowledge · WWW 2008 |
Information retrieval
search engines |
0.2 | 3 | 2010 | Faceted exploration of image search results · WWW 2010 Visual diversification of image search results · WWW 2009 Flexible and scalable digital library search · VLDB 2001 |
Information retrieval › search interfaces
faceted search |
0.2 | 2 | 2010 | Faceted exploration of image search results · WWW 2010 Machine learned ranking of entity facets · SIGIR 2010 |
Recommender systems
click-through rate prediction |
0.2 | 2 | 2012 | Multimedia features for click prediction of new ads in display advertising · KDD 2012 Machine learned ranking of entity facets · SIGIR 2010 |
Recommender systems
tag recommendation |
0.2 | 2 | 2008 | Flickr tag recommendation based on collective knowledge · WWW 2008 Resolving tag ambiguity · ACM Multimedia 2008 |
Information retrieval › image retrieval › object retrieval
image object retrieval |
0.1 | 1 | 2012 | Spatially-aware indexing for image object retrieval · WSDM 2012 |
Information retrieval
multimedia analysis and retrieval |
0.1 | 1 | 2012 | Spatially-aware indexing for image object retrieval · WSDM 2012 |
Information retrieval › retrieval models
vector space model |
0.1 | 1 | 2012 | Spatially-aware indexing for image object retrieval · WSDM 2012 |
Information retrieval › ranking › search ranking
facet ranking |
0.1 | 1 | 2010 | Faceted exploration of image search results · WWW 2010 |
Information retrieval › ranking
learning to rank |
0.1 | 1 | 2010 | Machine learned ranking of entity facets · SIGIR 2010 |
Recommender systems
multimodal recommendation |
0.1 | 1 | 2010 | Prediction of favourite photos using social, visual, and textual signals · ACM Multimedia 2010 |
Information retrieval › document retrieval › domain-specific retrieval
geographic information retrieval |
0.1 | 1 | 2009 | Placing flickr photos on a map · SIGIR 2009 |
Data mining › predictive modeling › classification › multi-label classification
tag-based classification |
0.1 | 1 | 2009 | Classifying tags using open content resources · WSDM 2009 |
Web and social media mining › user-generated content
wikipedia |
0.1 | 1 | 2009 | Classifying tags using open content resources · WSDM 2009 |
Multimedia analysis and retrieval
image retrieval |
0.1 | 1 | 2009 | Visual diversification of image search results · WWW 2009 |
Machine learning › Probabilistic and Bayesian machine learning › structured models
graphical models |
0.1 | 1 | 2008 | Resolving tag ambiguity · ACM Multimedia 2008 |
Information retrieval › image retrieval
content-based image retrieval |
0.1 | 1 | 2008 | Boosting image retrieval through aggregating search results based on visual annotations · ACM Multimedia 2008 |
Information retrieval › ranking
rank aggregation |
0.1 | 1 | 2008 | Boosting image retrieval through aggregating search results based on visual annotations · ACM Multimedia 2008 |
Information retrieval › document retrieval
digital library search |
0.1 | 2 | 2002 | Content-Based Video Indexing for the Support of Digital Library Search · ICDE 2002 Flexible and scalable digital library search · VLDB 2001 |
Web and social media mining
social tagging |
0.1 | 2 | 2010 | Faceted exploration of image search results · WWW 2010 Classifying tags using open content resources · WSDM 2009 |
Multimedia analysis and retrieval › video indexing
content-based video indexing |
0.0 | 1 | 2002 | Content-Based Video Indexing for the Support of Digital Library Search · ICDE 2002 |
Information retrieval › relevance feedback
click feedback |
0.0 | 1 | 2010 | Machine learned ranking of entity facets · SIGIR 2010 |
Web and social media mining
user-generated content |
0.0 | 1 | 2010 | Faceted exploration of image search results · WWW 2010 |
Information retrieval › retrieval models
language model |
0.0 | 1 | 2009 | Placing flickr photos on a map · SIGIR 2009 |
Computational social science and digital humanities
social media analysis |
0.0 | 1 | 2008 | Resolving tag ambiguity · ACM Multimedia 2008 |
Information retrieval
keyword search |
0.0 | 1 | 2008 | Boosting image retrieval through aggregating search results based on visual annotations · ACM Multimedia 2008 |
Information retrieval
large-scale retrieval |
0.0 | 1 | 2001 | Flexible and scalable digital library search · VLDB 2001 |
Methods — techniques the papers use, named apart from their topics
content moderation · 0.5multimedia feature extraction · 0.3click prediction models · 0.3workshop overview · 0.2social signal analysis · 0.2multimodal machine learning · 0.2probabilistic framework · 0.2metadata analysis · 0.2visual bag-of-words · 0.1query log analysis · 0.1gradient boosted decision trees · 0.1click-over-expected-clicks · 0.1metadata extraction · 0.0full-text indexing · 0.0conceptual modeling · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2021 | Integrity 2021: Integrity in Social Networks and MediaabstractThe second Workshop on Integrity in Social Networks and Media is held in conjunction with the 14th ACM Conference on Web Search and Data Mining (WSDM) in Jerusalem, Israel. The goal of the workshop is to bring together researchers and practitioners to discuss content and interaction integrity challenges in social networks and social media platforms. Lluís Garcia Pueyo, Anand Bhaskar, Roelof van Zwol, Timos K. Sellis, Gireeja Ranade, Prathyusha Senthil Kumar, Yu Sun 0021, Joy Zhang |
WSDM | 3 |
| 2015 | Interactive Recommender Systems: Tutorial
Harald Steck, Roelof van Zwol, Chris Johnson 0011 |
RecSys | 2 |
| 2015 | CelebrityNet: A Social Network Constructed from Large-Scale Online Celebrity ImagesabstractPhotos are an important information carrier for implicit relationships. In this article, we introduce an image based social network, called CelebrityNet , built from implicit relationships encoded in a collection of celebrity images. We analyze the social properties reflected in this image-based social network and automatically infer communities among the celebrities. We demonstrate the interesting discoveries of the CelebrityNet. We particularly compare the inferred communities with human manually labeled ones and show quantitatively that the automatically detected communities are highly aligned with that of human interpretation. Inspired by the uniqueness of visual content and tag concepts within each community of the CelebrityNet, we further demonstrate that the constructed social network can serve as a knowledge base for high-level visual recognition tasks. In particular, this social network is capable of significantly improving the performance of automatic image annotation and classification of unknown images. Li-Jia Li 0001, David A. Shamma, Xiangnan Kong, Sina Jafarpour, Roelof van Zwol, Xuanhui Wang |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2012 | Digital Paparazzi: Spotting Celebrities in Professional Photo Libraries
Sina Jafarpour, Li-Jia Li 0001, Roelof van Zwol |
ACCV (2) | 3 |
| 2012 | Multimedia features for click prediction of new ads in display advertisingabstractNon-guaranteed display advertising (NGD) is a multi-billion dollar business that has been growing rapidly in recent years. Advertisers in NGD sell a large portion of their ad campaigns using performance dependent pricing models such as cost-per-click (CPC) and cost-per-action (CPA). An accurate prediction of the probability that users click on ads is a crucial task in NGD advertising because this value is required to compute the expected revenue. State-of-the-art prediction algorithms rely heavily on historical information collected for advertisers, users and publishers. Click prediction of new ads in the system is a challenging task due to the lack of such historical data. The objective of this paper is to mitigate this problem by integrating multimedia features extracted from display ads into the click prediction models. Multimedia features can help us capture the attractiveness of the ads with similar contents or aesthetics. In this paper we evaluate the use of numerous multimedia features (in addition to commonly used user, advertiser and publisher features) for the purposes of improving click prediction in ads with no history. We provide analytical results generated over billions of samples and demonstrate that adding multimedia features can significantly improve the accuracy of click prediction for new ads, compared to a state-of-the-art baseline model. Haibin Cheng, Roelof van Zwol, Javad Azimi, Eren Manavoglu, Ruofei Zhang, Yang Zhou 0033, Vidhya Navalpakkam |
KDD | 2 |
| 2012 | Spatially-aware indexing for image object retrievalabstractThe success of image object retrieval systems relies on the visual bag-of-words paradigm, which allows image retrieval systems to adopt a retrieval strategy analogous to text retrieval. In this paper we propose two spatially-aware retrieval strategies for image object retrieval that replaces the vector space model. The advantage of the proposed spatially-aware indexing and retrieval strategies are threefold: (1) It allows for the deployment of small visual vocabularies, (2) the number of images evaluated at retrieval time is significantly reduced, and (3) it eliminates the need for a post-retrieval phase, which is normally used to test the spatial composition of the visual words in the retrieved images. Roelof van Zwol, Lluís Garcia Pueyo |
WSDM | 1 |
| 2011 | Object rankingabstractObject ranking is an emerging discipline within information retrieval that is concerned with the ranking of objects, e.g. named entities and their attributes, in context of given a user query, or application. In this tutorial we will address the different aspects involved when building an object ranking system. We will present the state-of-the-art research in object ranking, as well as going into detail about our hands-on experiences when designing and developing the system for object ranking as it is in production at Yahoo! today. This allows for a unique mixture of research and development that will give the participants in-depth insights into the problem of object ranking. Roelof van Zwol, Srinivas Vadrevu |
CIKM | 1 |
| 2011 | Scalable triangulation-based logo recognitionabstractWe propose a scalable logo recognition approach that extends the common bag-of-words model and incorporates local geometry in the indexing process. Given a query image and a large logo database, the goal is to recognize the logo contained in the query, if any. We locally group features in triples using multi-scale Delaunay triangulation and represent triangles by signatures capturing both visual appearance and local geometry. Each class is represented by the union of such signatures over all instances in the class. We see large scale recognition as a sub-linear search problem where signatures of the query image are looked up in an inverted index structure of the class models. We evaluate our approach on a large-scale logo recognition dataset with more than four thousand classes. Yannis Kalantidis, Lluís Garcia Pueyo, Michele Trevisiol, Roelof van Zwol, Yannis Avrithis |
ICMR | 4 |
| 2011 | Learning crop regions for content-aware generation of thumbnail imagesabstractWe propose a model for automatically cropping images based on a diverse set of content and spatial features. We approach this by extracting pixel-level features and aggregating them over possible crop regions. We then learn a regression model to predict the quality of the crop regions, via the degree to which they would overlaps with human-provided crops from these input features. Candidate images can then be cropped based an exhaustive sweep over candidate crop regions, where each region is scored and the highest-scoring region is retained. The system is unique in its ability to incorporate a variety of pixel-level importance cues when arriving at a final cropping recommendation. We test the system on a set of human-cropped images with a large set of features. We find that the system outperforms baseline approaches, particularly when the aspect ratio of the image is very different from the target thumbnail region. Lyndon Kennedy, Roelof van Zwol, Nicolas Torzec, Belle L. Tseng |
ICMR | 2 |
| 2011 | Scalable logo recognition in real-world imagesabstractIn this paper we propose a highly effective and scalable framework for recognizing logos in images. At the core of our approach lays a method for encoding and indexing the relative spatial layout of local features detected in the logo images. Based on the analysis of the local features and the composition of basic spatial structures, such as edges and triangles, we can derive a quantized representation of the regions in the logos and minimize the false positive detections. Furthermore, we propose a cascaded index for scalable multi-class recognition of logos.For the evaluation of our system, we have constructed and released a logo recognition benchmark which consists of manually labeled logo images, complemented with non-logo images, all posted on Flickr. The dataset consists of a training, validation, and test set with 32 logo-classes. We thoroughly evaluate our system with this benchmark and show that our approach effectively recognizes different logo classes with high precision. Stefan Romberg, Lluís Garcia Pueyo, Rainer Lienhart, Roelof van Zwol |
ICMR | 4 |
| 2011 | Introduction to the special issue on image and video retrieval: theory and applications
Ioannis Kompatsiaris, Stéphane Marchand-Maillet, Roelof van Zwol, Sébastien Marcel |
Multim. Tools Appl. | 3 |
| 2010 | WSM'10: 2nd ACM workshop on social mediaabstractThe ACM SIGMM International Workshop on Social Media (WSM'10) is the second workshop held in conjunction with the ACM International Multimedia Conference (MM'10) at Firenze, Italy, 2010. This workshop provides a forum for researchers and practitioners from all over the world to share information on their latest investigations on social media analysis, exploration, search, mining, and emerging new social media applications. Susanne Boll, Steven C. H. Hoi, Roelof van Zwol, Jiebo Luo 0001 |
ACM Multimedia | 3 |
| 2010 | Prediction of favourite photos using social, visual, and textual signalsabstractThis paper focuses on the prediction of users' favourite photos in Flickr. We propose a multi-modal, machine learned approach that combines social, visual and textual signals into a single prediction system. Although each individual user has different motivations for calling a photo a favourite, we show that the textual, visual, and social modalities effectively capture the needs of most active Flickr users. Roelof van Zwol, Adam Rae, Lluís Garcia Pueyo |
ACM Multimedia | 1 |
| 2010 | Machine learned ranking of entity facetsabstractThe research described in this paper forms the backbone of a service that enables the faceted search experience of the Yahoo! search engine. We introduce an approach for a machine learned ranking of entity facets based on user click feedback and features extracted from three different ranking sources. The objective of the learned model is to predict the click-through rate on an entity facet. In an empirical evaluation we compare the performance of gradient boosted decision trees (GBDT) against a linear combination of features on two different click feedback models using the raw click-through rate (CTR), and click over expected clicks (COEC). The results show a significant improvement in retrieval performance, in terms of discounted cumulated gain, when ranking entity facets with GBDT trained on the COEC model. Most notably this is true when evaluated against the CTR test set. Roelof van Zwol, Lluís Garcia Pueyo, Mridul Muralidharan, Börkur Sigurbjörnsson |
SIGIR | 1 |
| 2010 | Faceted exploration of image search resultsabstractThis paper describes MediaFaces, a system that enables faceted exploration of media collections. The system processes semi-structured information sources to extract objects and facets, e.g. the relationships between two objects. Next, we rank the facets based on a statistical analysis of image search query logs, and the tagging behaviour of users annotating photos in Flickr. For a given object of interest, we can then retrieve the top-k most relevant facets and present them to the user. The system is currently deployed in production by Yahoo!'s image search engine1. We present the system architecture, its main components, and the application of the system as part of the image search experience. Roelof van Zwol, Börkur Sigurbjörnsson, Ramu Adapala, Lluís Garcia Pueyo, Abhinav Katiyar, Kaushal Kurapati, Mridul Muralidharan, Sudar Muthu, Vanessa Murdock 0001, Polly Ng, Anand Ramani, Anuj Sahai, Sriram Thiru Sathish, Hari Vasudev, Upendra Vuyyuru |
WWW | 1 |
| 2009 | Placing flickr photos on a mapabstractIn this paper we investigate generic methods for placing photos uploaded to Flickr on the World map. As primary input for our methods we use the textual annotations provided by the users to predict the single most probable location where the image was taken. Central to our approach is a language model based entirely on the annotations provided by users. We define extensions to improve over the language model using tag-based smoothing and cell-based smoothing, and leveraging spatial ambiguity. Further we demonstrate how to incorporate GeoNames\footnote{http://www.geonames.org visited May 2009}, a large external database of locations. For varying levels of granularity, we are able to place images on a map with at least twice the precision of the state-of-the-art reported in the literature. Pavel Serdyukov, Vanessa Murdock 0001, Roelof van Zwol |
SIGIR | 3 |
| 2009 | Exploiting Tags and Social Profiles to Improve Focused CrawlingabstractRecent years have transformed the Web from a Web of content to a Web of applications and social content. Thus, it has become crucial to be able to tap on this social aspect of the Web whenever possible, in addition to its content, particularly for focused crawling. In this paper, we present a novel profile-based focused crawling system for dealing with the increasingly popular social media-sharing web sites without assuming any privileged access to the internal private databases of such websites, nor any requirement for the existence of APIs for the extraction of social data. Our experiments prove the robustness of our profile-based focused crawler, as well as a significant improvement in harvest ratio, compared to breadth-first and OPIC crawlers, when crawling the flickr web site for two different topics. Olfa Nasraoui, Roelof van Zwol |
Web Intelligence | 3 |
| 2009 | Classifying tags using open content resourcesabstractTagging has emerged as a popular means to annotate on-line objects such as bookmarks, photos and videos. Tags vary in semantic meaning and can describe different aspects of a media object. Tags describe the content of the media as well as locations, dates, people and other associated meta-data. Being able to automatically classify tags into semantic categories allows us to understand better the way users annotate media objects and to build tools for viewing and browsing the media objects. In this paper we present a generic method for classifying tags using third party open content resources, such as Wikipedia and the Open Directory. Our method uses structural patterns that can be extracted from resource meta-data. We describe the implementation of our method on Wikipedia using WordNet categories as our classification schema and ground truth. Two structural patterns found in Wikipedia are used for training and classification: categories and templates. We apply our system to classifying Flickr tags. Compared to a WordNet baseline our method increases the coverage of the Flickr vocabulary by 115%. We can classify many important entities that are not covered by WordNet, such as, London Eye, Big Island, Ronaldinho, geocaching and wii. Simon E. Overell, Börkur Sigurbjörnsson, Roelof van Zwol |
WSDM | 3 |
| 2009 | Visual diversification of image search resultsabstractDue to the reliance on the textual information associated with an image, image search engines on the Web lack the discriminative power to deliver visually diverse search results. The textual descriptions are key to retrieve relevant results for a given user query, but at the same time provide little information about the rich image content. Reinier H. van Leuken, Lluís Garcia Pueyo, Ximena Olivares, Roelof van Zwol |
WWW | 4 |
| 2008 | Boosting image retrieval through aggregating search results based on visual annotationsabstractOnline photo sharing systems, such as Flickr and Picasa, provide a valuable source of human-annotated photos. Textual annotations are used not only to describe the visual content of an image, but also subjective, spatial, temporal and social dimensions, complicating the task of keyword-based search. In this paper we investigate a method that exploits visual annotations, e.g. notes in Flickr, to enhance keyword-based systems retrieval performance. For this purpose we adopt the bag-of-visual-words approach for content-based image retrieval as our baseline. We then apply rank aggregation of the top 25 results obtained with a set of visual annotations that match the keyword-based query. The results on retrieval experiments show significant improvements in retrieval performance when comparing the aggregated approach with our baseline, which also slightly outperforms text-only search. When using a textual filter on the search space in combination with the aggregated approach an additional boost in retrieval performance is observed, which underlines the need for large scale content-based image retrieval techniques to complement the text-based search. Ximena Olivares, Massimiliano Ciaramita, Roelof van Zwol |
ACM Multimedia | 3 |
| 2008 | Resolving tag ambiguityabstractTagging is an important way for users to succinctly describe the content they upload to the Internet. However, most tag-suggestion systems recommend words that are highly correlated with the existing tag set, and thus add little information to a user's contribution. This paper describes a means to determine the ambiguity of a set of (user-contributed) tags and suggests new tags that disambiguate the original tags. We introduce a probabilistic framework that allows us to find two tags that appear in different contexts but are both likely to co-occur with the original tag set. If such tags can be found, the current description is considered "ambiguous" and the two tags are recommended to the user for further clarification. In contrast to previous work, we only query the user when information is most needed and good suggestions are available. We verify the efficacy of our approach using geographical, temporal and semantic metadata, and a user study. We built our system using statistics from a large (100M) database of images and their tags. Kilian Q. Weinberger, Malcolm Slaney, Roelof van Zwol |
ACM Multimedia | 3 |
| 2008 | Flickr tag recommendation based on collective knowledgeabstractOnline photo services such as Flickr and Zooomr allow users to share their photos with family, friends, and the online community at large. An important facet of these services is that users manually annotate their photos using so called tags, which describe the contents of the photo or provide additional contextual and semantical information. In this paper we investigate how we can assist users in the tagging phase. The contribution of our research is twofold. We analyse a representative snapshot of Flickr and present the results by means of a tag characterisation focussing on how users tags photos and what information is contained in the tagging. Based on this analysis, we present and evaluate tag recommendation strategies to support the user in the photo annotation task by recommending a set of tags that can be added to the photo. The results of the empirical evaluation show that we can effectively recommend relevant tags for a variety of photos with different levels of exhaustiveness of original tagging. Börkur Sigurbjörnsson, Roelof van Zwol |
WWW | 2 |
| 2007 | Effective Use of Semantic Structure in XML Retrieval
Roelof van Zwol, Tim van Loosbroek |
ECIR | 1 |
| 2007 | Flickr: Who is Looking?abstractThis article presents a characterization of user behavior on Flickr, a popular on-line photo sharing service that allows users to store, search, sort and share their photos. Based on a sub-set of photos being uploaded during a 10 day window, we track the interest of users in those photos over a period of 50 days. In particular we investigate the user behavior on temporal, social, and spatial dimensions. Results show that the users are able to discover new photos within hours after being uploaded and that 50% of the photo views are generated within the first two days. The social networking behavior of users, and photo pooling are identified as the two major indicators related to a photo's popularity. Finally we show that the geographic distribution is more focussed around a geographic location for the infrequently viewed photos, than for the photos that attract a large number of views. Roelof van Zwol |
Web Intelligence | 1 |
| 2006 | Bricks: The Building Blocks to Tackle Query Formulation in Structured Document Retrieval
Roelof van Zwol, Jeroen Baas, Herre van Oostendorp, Frans Wiering |
ECIR | 1 |
| 2005 | Multi-Dimensional Scattered Ranking Methods for Geographic Information Retrieval
Marc J. van Kreveld, Iris Reinbacher, Avi Arampatzis, Roelof van Zwol |
GeoInformatica | 4 |
| 2004 | Distributed Ranking Methods for Geographic Information Retrieval
Marc J. van Kreveld, Iris Reinbacher, Avi Arampatzis, Roelof van Zwol |
SDH | 4 |
| 2004 | Google's "I'm Feeling Lucky", Truly a Gamble?
Roelof van Zwol, Herre van Oostendorp |
WISE | 1 |
| 2002 | Content-Based Video Indexing for the Support of Digital Library SearchabstractPresents a digital library search engine that combines efforts of the AMIS and DMW research projects, each covering significant parts of the problem of finding the required information in an enormous mass of data. The most important contributions of our work are the following: (1) We demonstrate a flexible solution for the extraction and querying of meta-data from multimedia documents in general. (2) Scalability and efficiency support are illustrated for full-text indexing and retrieval. (3) We show how, for a more limited domain, like an intranet, conceptual modelling can offer additional and more powerful query facilities. (4) In the limited domain case, we demonstrate how domain knowledge can be used to interpret low-level features into semantic content. In this short description, we focus on the first and fourth items. Milan Petkovic, Roelof van Zwol, Henk Ernst Blok, Willem Jonker, Peter M. G. Apers, Menzo Windhouwer, Martin L. Kersten |
ICDE | 2 |
| 2001 | Flexible and scalable digital library search
Henk Ernst Blok, Menzo Windhouwer, Roelof van Zwol, Milan Petkovic, Peter M. G. Apers, Martin L. Kersten, Willem Jonker |
VLDB | 3 |
| 2000 | The Webspace Method: On the Integration of Database Technology with Multimedia RetrievalabstractLarge collections of documents containing various types of multimedia, are made available to the WWW. Unfortunately, due to the un-structuredness of Internet environments it is hard to find specific information when one is looking for it.Search engines available can only rely their results on information retrieval techniques and most of the time they lack the desired power in query formulation.Modelling data on the web, as if it was designed for use within databases, should provide us with the necessary basis for enhancing this query formulation.This of course requires special care for dealing with the included multimedia data and the semi-structured aspects of data on the web.Modelling the entire web would be too ambitious, therefore we focus on a more feasible environment, like the intranet, where one can find large collections of related data.With the webspace method we have already shown how to deal with the various aspects of semi-structured data in large collections of related documents.In this paper we focus on the integration of our webspace method for concept-based search with content-based multimedia information retrieval (IR).A webspace consists of two levels.At the document level, a webspace is considered to be a collection of related documents.At the semantical level, concepts are defined to be used in the documents at the document level.By modelling these concepts using a webspace schema a semantical level of abstraction is gained.This supplies the necessary platform for querying data available within a specific webspace.For the integration with content-based information retrieval an existing IR model is adopted.We will discuss how this is used in the context of Mirror, a Multimedia DBMS, and how this framework is used for the integration with the webspace method for concept-based search. Roelof van Zwol, Peter M. G. Apers |
CIKM | 1 |
| 2000 | Modelling the Webspace of an IntranetabstractSearching the Internet using the currently available search engines is not satisfactory. The techniques used there focus on the extraction of relevant information directly from the documents available on the World Wide Web. We introduce a new approach, which aims at describing the content of a Web space, formed by a collection of related documents, instead of looking at the single documents. By identifying concepts and the relationships among them, the content of a Web space is described semantically in a schema for the Web space. The main objective is that, by following this approach, we can start querying the content of a collection of related documents rather than the content of a single document. In this paper, we introduce a model for Web spaces that allows us to describe the concepts at a semantic level, in terms of classes, associations over classes and attributes of classes. At the syntactic level, we use XML to describe information as instantiations of the concepts defined in the Web space schema. Dealing with data on the Web implies dealing with semi-structured data. We discuss how this relates to our model for a Web space and show how to deal with these aspects efficiently when moving towards an implementation. Roelof van Zwol, Peter M. G. Apers |
WISE | 1 |