Mihai Lupu

dblp:22/6000 · DBLP profile ↗
← Back
42ranked-venue papers
8as first author
3since 2021 · last 2022
0000-0002-6328-6873ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 32 · 8 first-author · 2 since 2021Artificial intelligence and machine learning · 10 · 2 first-authorGraphics, computer vision, multimedia, augmented reality and games · 5 · 1 since 2021Systems, architecture and hardware · 3Applied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
12 papers
Information retrieval · 94% Distributed and cloud data management · 2% Web and social media mining · 2%
Artificial intelligence
1 paper
Learning theory · 100%
Computer architecture, parallel and distributed computing, and storage systems
4 papers
Distributed systems · 86% Embedded and real-time systems · 14%
Computer graphics and multimedia
1 paper
Visualization and visual analytics · 100%

Topics — the 30 heaviest of 34, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Information retrieval
evaluation
1.032021
Fixed-Cost Pooling Strategies · IEEE Trans. Knowl. Data Eng. 2021
Visual Pool: A Tool to Visualize and Interact with the Pooling Method · SIGIR 2017
Splitting Water: Precision and Anti-Precision to Reduce Pool Bias · SIGIR 2015
Information retrieval › evaluation
test collection
0.932019
A Horizontal Patent Test Collection · SIGIR 2019
Visual Pool: A Tool to Visualize and Interact with the Pooling Method · SIGIR 2017
Splitting Water: Precision and Anti-Precision to Reduce Pool Bias · SIGIR 2015
Information retrieval › evaluation › test collection › test collection construction
pooling
0.822021
Fixed-Cost Pooling Strategies · IEEE Trans. Knowl. Data Eng. 2021
Visual Pool: A Tool to Visualize and Interact with the Pooling Method · SIGIR 2017
Machine learning › Learning theory
generalization
0.612022
Formal Analysis and Estimation of Chance in Datasets Based on Their Properties · IEEE Trans. Knowl. Data Eng. 2022
Information retrieval › document retrieval › domain-specific retrieval › legal information retrieval
patent retrieval
0.522019
A Horizontal Patent Test Collection · SIGIR 2019
Patent information retrieval: an instance of domain-specific search · SIGIR 2012
Information retrieval › evaluation
benchmark
0.512021
Benchmarking Image Retrieval Diversification Techniques for Social Media · IEEE Trans. Multim. 2021
Information retrieval
retrieval evaluation
0.512021
Benchmarking Image Retrieval Diversification Techniques for Social Media · IEEE Trans. Multim. 2021
Information retrieval › evaluation › test collection
test collection construction
0.512021
Fixed-Cost Pooling Strategies · IEEE Trans. Knowl. Data Eng. 2021
Information retrieval › retrieval models › representation learning for retrieval
word embedding for retrieval
0.422017
Word Embedding Causes Topic Shifting; Exploit Global Context! · SIGIR 2017
Volatility Prediction using Financial Disclosures Sentiments with Word Embedding-based IR Models · ACL (1) 2017
Information retrieval
retrieval models
0.322017
Word Embedding Causes Topic Shifting; Exploit Global Context! · SIGIR 2017
HiWaRPP - Hierarchical Wavelet-based Retrieval on Peer-to-Peer Network · ICDE 2006
Information retrieval › text analysis
semantic relatedness
0.312017
Word Embedding Causes Topic Shifting; Exploit Global Context! · SIGIR 2017
Information retrieval › evaluation › test collection
pool bias
0.212015
Splitting Water: Precision and Anti-Precision to Reduce Pool Bias · SIGIR 2015
Distributed systems
peer-to-peer systems
0.232008
P3N: profiling the potential of a peer-based data management system · Proc. VLDB Endow. 2008
Clustering wavelets to speed-up data dissemination in structured P2P MANETs · ICDE 2007
HiWaRPP - Hierarchical Wavelet-based Retrieval on Peer-to-Peer Network · ICDE 2006
Bioinformatics and computational biology
genomics
0.212022
Formal Analysis and Estimation of Chance in Datasets Based on Their Properties · IEEE Trans. Knowl. Data Eng. 2022
Distributed systems › peer-to-peer systems › overlay networks
structured overlay
0.222008
P3N: profiling the potential of a peer-based data management system · Proc. VLDB Endow. 2008
Clustering wavelets to speed-up data dissemination in structured P2P MANETs · ICDE 2007
Information retrieval › image retrieval › web image search
social image retrieval
0.112021
Benchmarking Image Retrieval Diversification Techniques for Social Media · IEEE Trans. Multim. 2021
Web and social media mining
social media analysis
0.112021
Benchmarking Image Retrieval Diversification Techniques for Social Media · IEEE Trans. Multim. 2021
Information retrieval › document retrieval
domain-specific retrieval
0.112012
Patent information retrieval: an instance of domain-specific search · SIGIR 2012
Distributed and cloud data management › peer-to-peer data management
peer-to-peer databases
0.112008
Paths to stardom: calibrating the potential of a peer-based data management system · SIGMOD Conference 2008
Distributed and cloud data management
peer-to-peer data management
0.112008
P3N: profiling the potential of a peer-based data management system · Proc. VLDB Endow. 2008
Distributed systems › distributed communication
data dissemination
0.112007
Clustering wavelets to speed-up data dissemination in structured P2P MANETs · ICDE 2007
Information retrieval
distributed information retrieval
0.112006
HiWaRPP - Hierarchical Wavelet-based Retrieval on Peer-to-Peer Network · ICDE 2006
Information retrieval › distributed information retrieval
peer-to-peer search
0.112006
HiWaRPP - Hierarchical Wavelet-based Retrieval on Peer-to-Peer Network · ICDE 2006
Embedded and real-time systems
real-time system verification
0.112006
Automatic Debugging of Real-Time Systems Based on Incremental Satisfiability Counting · IEEE Trans. Computers 2006
Automated reasoning and model checking
satisfiability
0.112006
Automatic Debugging of Real-Time Systems Based on Incremental Satisfiability Counting · IEEE Trans. Computers 2006
Information retrieval
search interfaces
0.012012
Patent information retrieval: an instance of domain-specific search · SIGIR 2012
Graph algorithms and graph theory › graph theory › algebraic graph theory
cayley graph
0.012008
Paths to stardom: calibrating the potential of a peer-based data management system · SIGMOD Conference 2008
Information retrieval › similarity search
approximate similarity search
0.012007
Clustering wavelets to speed-up data dissemination in structured P2P MANETs · ICDE 2007
Information retrieval
similarity search
0.012007
Clustering wavelets to speed-up data dissemination in structured P2P MANETs · ICDE 2007
Wireless networking
mobile ad hoc networks
0.012007
Clustering wavelets to speed-up data dissemination in structured P2P MANETs · ICDE 2007

Methods — techniques the papers use, named apart from their topics

formal analysis · 1.1cross-validation · 1.1chance influence estimation · 1.1word embeddings · 0.6visualization · 0.6voting systems · 0.5retrieval fusion · 0.5multimodal retrieval · 0.5multi-armed bandit · 0.5diversity evaluation · 0.5regression · 0.3global context filtering · 0.3bias correction · 0.2wavelet transform · 0.1k-means clustering · 0.1incremental satisfiability counting · 0.1profiling · 0.1algebraic and combinatorial analysis · 0.1
YearPublicationVenuePosition
2022 Formal Analysis and Estimation of Chance in Datasets Based on Their Properties
abstract
Machine learning research, particularly in genomics, is often based on wide shaped datasets, i.e. datasets having a large number of features, but a small number of samples. Such configurations raise the possibility of chance influence (the increase of measured accuracy due to chance correlations) on the learning process and the evaluation results. Prior research underlined the problem of generalization of models obtained based on such data. In this paper, we investigate the influence of chance on prediction and show its significant effects on wide shaped datasets. First, we empirically demonstrate how significant the influence of chance in such datasets is by showing that prediction models trained on thousands of randomly generated datasets can achieve high accuracy. This is the case even when using cross-validation. We then provide a formal analysis of chance influence and design formal chance influence estimators based on the dataset parameters, namely its sample size, the number of features, the number of classes and the class distribution. Finally, we provide an in-depth discussion of the formal analysis including applications of the findings and recommendations on chance influence mitigation.
Abdel Aziz Taha, Luca Papariello, Alexandros Bampoulidis, Petr Knoth, Mihai Lupu
IEEE Trans. Knowl. Data Eng.5
2021 Fixed-Cost Pooling Strategies
abstract
The empirical nature of Information Retrieval (IR) mandates strong experimental practices. A keystone of such experimental practices is the Cranfield evaluation paradigm. Within this paradigm, the collection of relevance judgments has been the subject of intense scientific investigation. This is because, on one hand, consistent, precise, and numerous judgements are keys to reducing evaluation uncertainty and test collection bias; on the other hand, however, relevance judgements are costly to collect. The selection of which documents to judge for relevance, known as pooling method, has therefore a great impact on IR evaluation. In this paper we focus on the bias introduced by the pooling method, known as pool bias, which affects the reusability of test collections, in particular when building test collections with a limited budget. In this paper we formalize and evaluate a set of 22 pooling strategies based on: traditional strategies, voting systems, retrieval fusion methods, evaluation measures, and multi-armed bandit models. To do this we run a large-scale evaluation by considering a set of 9 standard TREC test collections, in which we show that the choice of the pooling strategy has significant effects on the cost needed to obtain an unbiased test collection. We also identify the least biased pooling strategy in terms of pool bias according to three IR evaluation measures: AP, NDCG, and P@10.
Aldo Lipani, David E. Losada, Guido Zuccon, Mihai Lupu
IEEE Trans. Knowl. Data Eng.4
2021 Benchmarking Image Retrieval Diversification Techniques for Social Media
abstract
Image retrieval has been an active research domain for over 30 years and historically it has focused primarily on precision as an evaluation criterion. Similar to text retrieval, where the number of indexed documents became large and many relevant documents exist, it is of high importance to highlight diversity in the search results to provide better results for the user. The Retrieving Diverse Social Images Task of the MediaEval benchmarking campaign has addressed exactly this challenge of retrieving diverse and relevant results for the past years, specifically in the social media context. Multimodal data (e.g., images, text) was made available to the participants including metadata assigned to the images, user IDs, and precomputed visual and text descriptors. Many teams have participated in the task over the years. The large number of publications employing the data and also citations of the overview articles underline the importance of this topic. In this paper, we introduce these publicly available data resources as well as the evaluation framework, and provide an in-depth analysis of the crucial aspects of social image search diversification, such as the capabilities and the evolution of existing systems. These evaluation resources will help researchers for the coming years in analyzing aspects of multimodal image retrieval and diversity of the search results.
Bogdan Ionescu, Maia Rohm, Bogdan Boteanu, Alexandru-Lucian Gînsca, Mihai Lupu, Henning Müller
IEEE Trans. Multim.5
2020 On the Replicability of Combining Word Embeddings and Retrieval Models
Luca Papariello, Alexandros Bampoulidis, Mihai Lupu
ECIR (2)3
2020 Practice and Challenges of (De-)Anonymisation for Data Sharing
Alexandros Bampoulidis, Alessandro Bruni, Ioannis Markopoulos, Mihai Lupu
RCIS4
2019 Enriching Word Embeddings for Patent Retrieval with Global Context
Sebastian Hofstätter, Navid Rekabsaz, Mihai Lupu, Carsten Eickhoff, Allan Hanbury
ECIR (1)3
2019 A Horizontal Patent Test Collection
abstract
We motivate the need for, and describe the contents of a novel patent research collection, publicly available and for free, covering multimodal and multilingual data from six patent authorities. The new patent test collection complements existing patent test collections, which are vertical (one domain or one authority over many years). Instead, the new collection is horizontal: it includes all technical domains from the major patenting authorities over the relatively short time span of two years. In addition to bringing together documents currently scattered across different test collections, the collection provides, for the first time, Korean documents, to complement those from Europe, US, Japan, and China. This new collection can be used on a variety of tasks beyond traditional information retrieval. We exemplify this with a task of high-relevance today: de-anonymisation.
Mihai Lupu, Alexandros Bampoulidis, Luca Papariello
SIGIR1
2018 An analysis of evaluation campaigns in ad-hoc medical information retrieval: CLEF eHealth 2013 and 2014
Lorraine Goeuriot, Gareth J. F. Jones, Liadh Kelly, Johannes Leveling, Mihai Lupu, João R. M. Palotti, Guido Zuccon
Inf. Retr. J.5
2018 A systematic approach to normalization in probabilistic models
abstract
Every information retrieval (IR) model embeds in its scoring function a form of term frequency (TF) quantification. The contribution of the term frequency is determined by the properties of the function of the chosen TF quantification, and by its TF normalization. The first defines how independent the occurrences of multiple terms are, while the second acts on mitigating the a priori probability of having a high term frequency in a document (estimation usually based on the document length). New test collections, coming from different domains (e.g. medical, legal), give evidence that not only document length, but in addition, verboseness of documents should be explicitly considered. Therefore we propose and investigate a systematic combination of document verboseness and length. To theoretically justify the combination, we show the duality between document verboseness and length. In addition, we investigate the duality between verboseness and other components of IR models. We test these new TF normalizations on four suitable test collections. We do this on a well defined spectrum of TF quantifications. Finally, based on the theoretical and experimental observations, we show how the two components of this new normalization, document verboseness and length, interact with each other. Our experiments demonstrate that the new models never underperform existing models, while sometimes introducing statistically significantly better results, at no additional computational cost.
Aldo Lipani, Thomas Roelleke, Mihai Lupu, Allan Hanbury
Inf. Retr. J.3
2017 Volatility Prediction using Financial Disclosures Sentiments with Word Embedding-based IR Models
abstract
Navid Rekabsaz, Mihai Lupu, Artem Baklanov, Alexander Dür, Linda Andersson, Allan Hanbury. Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2017.
Navid Rekabsaz, Mihai Lupu, Artem Baklanov, Alexander Dür, Linda Andersson, Allan Hanbury
ACL (1)2
2017 Does Online Evaluation Correspond to Offline Evaluation in Query Auto Completion?
Alexandros Bampoulidis, João R. M. Palotti, Mihai Lupu, Jon Brassey, Allan Hanbury
ECIR3
2017 Fixed-Cost Pooling Strategies Based on IR Evaluation Measures
Aldo Lipani, João R. M. Palotti, Mihai Lupu, Florina Piroi, Guido Zuccon, Allan Hanbury
ECIR3
2017 Exploration of a Threshold for Similarity Based on Uncertainty in Word Embedding
Navid Rekabsaz, Mihai Lupu, Allan Hanbury
ECIR2
2017 Visual Pool: A Tool to Visualize and Interact with the Pooling Method
abstract
Every year more than 25 test collections are built among the main Information Retrieval (IR) evaluation campaigns. They are extremely important in IR because they become the evaluation praxis for the forthcoming years. Test collections are built mostly using the pooling method. The main advantage of this method is that it drastically reduces the number of documents to be judged. It does so at the cost of introducing biases, which are sometimes aggravated by non optimal configuration. In this paper we develop a novel visualization technique for the pooling method, and integrate it in a demo application named Visual Pool. This demo application enables the user to interact with the pooling method with ease, and develops visual hints in order to analyze existing test collections, and build better ones.
Aldo Lipani, Mihai Lupu, Allan Hanbury
SIGIR2
2017 Word Embedding Causes Topic Shifting; Exploit Global Context!
abstract
Exploitation of term relatedness provided by word embedding has gained considerable attention in recent IR literature. However, an emerging question is whether this sort of relatedness fits to the needs of IR with respect to retrieval effectiveness. While we observe a high potential of word embedding as a resource for related terms, the incidence of several cases of topic shifting deteriorates the final performance of the applied retrieval models. To address this issue, we revisit the use of global context (i.e. the term co-occurrence in documents) to measure the term relatedness. We hypothesize that in order to avoid topic shifting among the terms with high word embedding similarity, they should often share similar global contexts as well. We therefore study the effectiveness of post filtering of related terms by various global context relatedness measures. Experimental results show significant improvements in two out of three test collections, and support our initial hypothesis regarding the importance of considering global context in retrieval.
Navid Rekabsaz, Mihai Lupu, Allan Hanbury, Hamed Zamani
SIGIR2
2016 When is the Time Ripe for Natural Language Processing for Patent Passage Retrieval?
abstract
Patent text is a mixture of legal terms and domain specific terms. In technical English text, a multi-word unit method is often deployed as a word formation strategy in order to expand the working vocabulary, i.e. introducing a new concept without the invention of an entirely new word. In this paper we explore query generation using natural language processing technologies in order to capture domain specific concepts represented as multi-word units. In this paper we examine a range of query generation methods using both linguistic and statistical information. We also propose a new method to identify domain specific terms from other more general phrases. We apply a machine learning approach using domain knowledge and corpus linguistic information in order to learn domain specific terms in relation to phrases' Termhood values. The experiments are conducted on the English part of the CLEF-IP 2013 test collection. The outcome of the experiments shows that the favoured method in terms of PRES and recall is when a language model is used and search terms are extracted with a part-of-speech tagger and a noun phrase chunker. With our proposed methods we improve each evaluation metric significantly compared to the existing state-of-the-art for the CLEP-IP 2013 test collection: for [email protected] by 26% (0.544 from 0.433), for [email protected] by 17% (0.631 from 0.540) and on document MAP by 57% (0.300 from 0.191).
Linda Andersson, Mihai Lupu, João R. M. Palotti, Allan Hanbury, Andreas Rauber
CIKM2
2016 The Solitude of Relevant Documents in the Pool
abstract
Pool bias is a well understood problem of test-collection based benchmarking in information retrieval. The pooling method itself is designed to identify all relevant documents. In practice, 'all' translates to `as many as possible given some budgetary constraints' and the problem persists, albeit mitigated. Recently, methods to address this pool bias for previously created test collections have been proposed, for the evaluation measure precision at cut-off ([email protected]). Analyzing previous methods, we make the empirical observation that the distribution of the probability of providing new relevant documents to the pool, over the runs, is log-normal (when the pooling strategy is fixed depth at cut-off). We use this observation to calculate a prior probability of providing new relevant documents, which we then use in a pool bias estimator that improves upon previous estimates of precision at cut-off. Through extensive experimental results, covering 15 test collections, we show that the proposed bias correction method is the new state of the art, providing the closest estimates yet when compared to the original pool.
Aldo Lipani, Mihai Lupu, Evangelos Kanoulas, Allan Hanbury
CIKM2
2016 Generalizing Translation Models in the Probabilistic Relevance Framework
abstract
A recurring question in information retrieval is whether term associations can be properly integrated in traditional information retrieval models while preserving their robustness and effectiveness. In this paper, we revisit a wide spectrum of existing models (Pivoted Document Normalization, BM25, BM25 Verboseness Aware, Multi-Aspect TF, and Language Modelling) by introducing a generalisation of the idea of the translation model. This generalisation is a de facto transformation of the translation models from Language Modelling to the probabilistic models. In doing so, we observe a potential limitation of these generalised translation models: they only affect the term frequency based components of all the models, ignoring changes in document and collection statistics. We correct this limitation by extending the translation models with the 15 statistics of term associations and provide extensive experimental results to demonstrate the benefit of the newly proposed methods. Additionally, we compare the translation models with query expansion methods based on the same term association resources, as well as based on Pseudo-Relevance Feedback (PRF). We observe that translation models always outperform the first, but provide complementary information with the second, such that by using PRF and our translation models together we observe results better than the current state of the art.
Navid Rekabsaz, Mihai Lupu, Allan Hanbury, Guido Zuccon
CIKM2
2016 The Curious Incidence of Bias Corrections in the Pool
Aldo Lipani, Mihai Lupu, Allan Hanbury
ECIR2
2016 Standard Test Collection for English-Persian Cross-Lingual Word Sense Disambiguation
Navid Rekabsaz, Serwah Sabetghadam, Mihai Lupu, Linda Andersson, Allan Hanbury
LREC3
2016 Div150Multi: a social image retrieval result diversification dataset with multi-topic queries
abstract
In this paper we introduce a new dataset, Div150Multi, that was designed to support shared evaluation of diversification techniques in different areas of social media photo retrieval and related areas. The dataset comes with associated relevance and diversity assessments performed by trusted annotators. The data consists of around 300 complex queries represented via 86,769 Flickr photos, around 27M photo links for around 6,000 users, metadata, Wikipedia pages and content descriptors for text and visual modalities, including state of the art deep features. To facilitate distribution, only Creative Commons content allowing redistribution was included in the dataset. The proposed dataset was validated during the 2015 Retrieving Diverse Social Images Task at the MediaEval Benchmarking.
Bogdan Ionescu, Alexandru-Lucian Gînsca, Bogdan Boteanu, Mihai Lupu, Adrian Popescu 0001, Henning Müller
MMSys4
2016 Overview of the Special Issue on Trust and Veracity of Information in Social Media
abstract
research-article Share on Overview of the Special Issue on Trust and Veracity of Information in Social Media Authors: Symeon Papadopoulos Centre for Research and Technology Hellas; Thessaloniki, Greece Centre for Research and Technology Hellas; Thessaloniki, GreeceView Profile , Kalina Bontcheva University of Sheffield, Sheffield, UK University of Sheffield, Sheffield, UKView Profile , Eva Jaho Athens Technology Center, Athens, Greece Athens Technology Center, Athens, GreeceView Profile , Mihai Lupu Vienna University of Technology, Vienna, Austria Vienna University of Technology, Vienna, AustriaView Profile , Carlos Castillo Sapienza University of Rome, Rome, Italy Sapienza University of Rome, Rome, ItalyView Profile Authors Info & Claims ACM Transactions on Information SystemsVolume 34Issue 3May 2016 Article No.: 14pp 1–5https://doi.org/10.1145/2870630Published:11 April 2016Publication History 20citation1,578DownloadsMetricsTotal Citations20Total Downloads1,578Last 12 Months57Last 6 weeks12 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access
Symeon Papadopoulos, Kalina Bontcheva, Eva Jaho, Mihai Lupu, Carlos Castillo 0001
ACM Trans. Inf. Syst.4
2015 Reachability Analysis of Graph Modelled Collections
Serwah Sabetghadam, Mihai Lupu, Ralf Bierig, Andreas Rauber
ECIR2
2015 DASyR(IR) - document analysis system for systematic reviews (in Information Retrieval)
abstract
Creating systematic reviews is a painstaking task undertaken especially in domains where experimental results are the primary method to knowledge creation. For the review authors, analysing documents to extract relevant data is a demanding activity. To support the creation of systematic reviews, we have created DASyR-a semi-automatic document analysis system. DASyR is our solution to annotating published papers for the purpose of ontology population. For domains where dictionaries are not existing or inadequate, DASyR relies on a semi-automatic annotation bootstrapping method based on positional Random Indexing, followed by traditional Machine Learning algorithms to extend the annotation set. We provide an example of the method application to a subdomain of Computer Science, the Information Retrieval evaluation domain. The reliance of this domain on large scale experimental studies makes it a perfect domain to test on. We show the utility of DASyR through experimental results for different parameter values for the bootstrap procedure, evaluated in terms of annotator agreement, error rate, precision and recall.
Florina Piroi, Aldo Lipani, Mihai Lupu, Allan Hanbury
ICDAR3
2015 Div150Cred: A social image retrieval result diversification with user tagging credibility dataset
abstract
In this paper we introduce a new dataset and its evaluation tools, Div150Cred, that was designed to support shared evaluation of diversification techniques in different areas of social media photo retrieval and related areas. The dataset comes with associated relevance and diversity assessments performed by human annotators. The data consists of 300 landmark locations represented via 45,375 Flickr photos, 16M photo links for around 3,000 users, metadata, Wikipedia pages and content descriptors for text and visual modalities. To facilitate distribution, only Creative Commons content was included in the dataset. The proposed dataset was validated during the 2014 Retrieving Diverse Social Images Task at the MediaEval Benchmarking Initiative.
Bogdan Ionescu, Adrian Popescu 0001, Mihai Lupu, Alexandru-Lucian Gînsca, Bogdan Boteanu, Henning Müller
MMSys3
2015 Splitting Water: Precision and Anti-Precision to Reduce Pool Bias
abstract
For many tasks in evaluation campaigns, especially those modeling narrow domain-specific challenges, lack of participation leads to a potential pooling bias due to the scarce number of pooled runs. It is well known that the reliability of a test collection is proportional to the number of topics and relevance assessments provided for each topic, but also to same extent to the diversity in participation in the challenges. Hence, in this paper we present a new perspective in reducing the pool bias by studying the effect of merging an unpooled run with the pooled runs. We also introduce an indicator used by the bias correction method to decide whether the correction needs to be applied or not. This indicator gives strong clues about the potential of a "good" run tested on an "unfriendly" test collection (i.e. a collection where the pool was contributed to by runs very different from the one at hand). We demonstrate the correctness of our method on a set of fifteen test collections from the Text REtrieval Conference (TREC). We observe a reduction in system ranking error and absolute score difference error.
Aldo Lipani, Mihai Lupu, Allan Hanbury
SIGIR2
2014 A Glimpse into the State and Future of (Big) Data Analytics in Austria - Results from an Online Survey
abstract
We present results from questionnaire data that were collected from leading data analytics experts in Austria. The online survey addresses very current and pressing questions in the area of (big) data analysis. Our i??ndings provide valuable insights about what top Austrian data scientists think about data analytics, what they consider as important application areas that can benei??t from big data and data processing, the challenges of the future and how soon these challenges will become important, and the potential research topics of tomorrow. We visualize results, summarize our i??ndings and suggest a possible roadmap for future decision making.
Ralf Bierig, Allan Hanbury, Martina Haas, Florina Piroi, Helmut Berger, Mihai Lupu, Michael Dittenbach
DATA6
2014 A System Framework for Concept- and Credibility-Based Multimedia Retrieval
abstract
We present a multimedia retrieval system framework that incorporates components for processing multimedia content in different modes and languages. The framework provides concept-based information retrieval facilities that applies credibility information for result re-ranking. The architecture combines both a direct user interface and a batched evaluation interface for reproducible research in multimedia IR. The demo presents a preliminary version of the system framework and shows a use case based on the ImageCLEF 2011 Wikipedia test collection.
Ralf Bierig, Cristina Serban, Alexandra Siriteanu, Mihai Lupu, Allan Hanbury
ICMR4
2014 A Combined Approach of Structured and Non-structured IR in Multimodal Domain
abstract
We present a generic model for multimodal information retrieval, leveraging different information sources to improve the effectiveness of a retrieval system. The proposed method is able to take into account both explicit and latent semantics present in the data and can be used to answer complex queries, not currently answerable neither by document retrieval systems, nor by semantic web systems. By providing a hybrid approach combining IR and structured search techniques, we prepare a framework applicable to multimodal data collections. To test its effectiveness, we instantiate the model for an image retrieval task.
Serwah Sabetghadam, Mihai Lupu, Ralf Bierig, Andreas Rauber
ICMR2
2014 Guest editorial: Special issue on information retrieval in the intellectual property domain
Allan Hanbury, Mihai Lupu, Noriko Kando, Barrou Diallo
Inf. Retr.2
2013 Integrating IR Technologies for Professional Search - (Full-Day Workshop)
Michail Salampasis, Norbert Fuhr, Allan Hanbury, Mihai Lupu, Birger Larsen, Henrik Strindberg
ECIR4
2012 Applying Random Indexing to Structured Data to Find Contextually Similar Words
Danica Damljanovic, Udo Kruschwitz, M-Dyaa Albakour, Johann Petrak, Mihai Lupu
LREC5
2012 Patent information retrieval: an instance of domain-specific search
abstract
The tutorial aims to provide the IR researchers with an understanding of how the patent system works, the challenges that patent searchers face in using the existing tools and in adopting new methods developed in academia.
Mihai Lupu
SIGIR1
2011 4th international workshop on patent information retrieval (PaIR'11)
abstract
The 4th International Workshop on Patent Information Retrieval builds on the experiences of the first three workshops, to provide its participants an exciting, scientifically challenging and interactive event, where specific issues of patent retrieval may be put into the general context of Information Retrieval and Knowledge Management, in order to explore innovative solutions to new and old problems, but also to evaluate and adapt traditional or classic approaches to new problems. This year, we observe an increase in the use of standardized test collections in the contributions received, and, at the same time, new discussion points on how to make such standardized evaluation exercises more accessible to the larger IP community.
Mihai Lupu, Allan Hanbury, Andreas Rauber
CIKM1
2010 3rd international workshop on patent information retrieval (PaIR'10)
abstract
The 3rd International Workshop on Patent Information Retrieval builds on the experiences of the first two workshops, to provide its participants an exciting, scientifically challenging and interactive event, where the specific issues of patent retrieval may be put into the general context of Information Retrieval and Knowledge Management, in order to explore innovative solutions to new and old problems, but also to evaluate and adapt traditional or classic approaches to new problems. Between the scientific presentations and posters, distinguished keynote speakers and a panel discussion, PaIR 2010 shapes itself into a significant landmark in the field of domain specific information retrieval.
Mihai Lupu, John Tait, Katja Mayer, Christopher G. Harris 0001
CIKM1
2009 SiMPSON: Efficient Similarity Search in Metric Spaces over P2P Structured Overlay Networks
Quang Hieu Vu, Mihai Lupu, Sai Wu
Euro-Par2
2008 Paths to stardom: calibrating the potential of a peer-based data management system
abstract
As peer-to-peer (P2P) networks become more familiar to the database community, intense interest has built up in using their scalability and resilience properties to scale database applications. Indexing methods are adapted on top of P2P networks and querying methods are developed to handle the data distribution on different nodes. These procedures largely depend on how nodes are connected to each other. So far, limited attempts have been made to compare all these systems in a generalized framework. This is because the systems are quite different from each other, and there are so many of them that brute force comparison is practically impossible. Fortunately, it has recently been observed that a large subset of the most important P2P networks share a common algebraic and combinatorial base, in the form of Cayley graphs.
Mihai Lupu, Beng Chin Ooi, Y. C. Tay
SIGMOD Conference1
2008 P3N: profiling the potential of a peer-based data management system
abstract
A large number of peer-to-peer (P2P) networks have been introduced in the literature since their popular advent in the late 1990s. In particular, structured P2P overlays have gained much attention since 2001. They are noted mainly for their theoretical properties such as balancing of communication, storage and processing load, as well as elegance of design.
Mihai Lupu, Y. C. Tay
Proc. VLDB Endow.1
2007 Clustering wavelets to speed-up data dissemination in structured P2P MANETs
abstract
This paper introduces a fast data dissemination method for structured peer-to-peer networks. The work is motivated on one side by the increase in non-volatile memory available on mobile devices and, on the other side, by observed behavioral patterns of the users. We envision a scenario where users come together for short periods of time (e.g. public transport, conference sessions) and wish to be able to share large collections of data. With hundreds and even thousands of data, items stored on small devices, content publication is simply too energy and time consuming. By indexing summary information obtained by a combination of multi-resolution analysis and k-means, our method (Hyper-Ad) is able to cut down the overall construction time of an overlay network such as CAN by an order of magnitude, as well as provide fast approximate similarity search on such a network. The results of our extensive experimental studies confirm that Hyper-M is both energy and time efficient, and provides good precision and recall.
Mihai Lupu, Jianzhong Li 0001, Beng Chin Ooi, Shengfei Shi
ICDE1
2006 HiWaRPP - Hierarchical Wavelet-based Retrieval on Peer-to-Peer Network
abstract
This paper introduces the use of wavelets for information retrieval in a peer-to-peer environment. In order to achieve our purposes, we use a new combination between broadcasting and a hierarchical overlay. Compared to previous approaches, we do not store complete information about the children of a super-peer, nor do we broadcast the queries blindly. We approximate the feature vectors using the multiresolution analysis and the discrete wavelet transform. Each peer is represented by a high-dimensional feature vector and the height of the hierarchy is logarithmic in the dimensionality of this feature vector. Leaf nodes represent real peers, while internal nodes are virtual peers used for routing. Our retrieval method has been tested with both real and synthetic data and shown to be efficient in retrieving relevant information, resulting in good precision and recall on four standard test collections.
Mihai Lupu, Bei Yu 0003
ICDE1
2006 Automatic Debugging of Real-Time Systems Based on Incremental Satisfiability Counting
abstract
Real-time logic (RTL) is useful for the verification of a safety assertion with respect to the specification of a realtime system. Since the satisfiability problem for RTL is undecidable, the systematic debugging of a real-time system appears impossible. A first step toward this challenge was presented. With RTL, each prepositional formula corresponds to a verification condition. The number of truth assignments of a prepositional formula can help us determine the specific constraints which should be added or modified to get the expected solutions. This paper solves an even more challenging problem specified as future work, namely, the embedding and the integration of our debugger in autonomous systems which generate real-time control plans on-the-fly, since these specifications must meet timing constraints, but without human interaction. The idea is to consider in advance all the necessary information, such as the designer's guidance. We have implemented a tool (called ADRTL) that is able to perform automatic debugging. The confidence of our approach is high as we have successfully evaluated ADRTL on several existing industrial-based applications.
Stefan Andrei, Wei-Ngan Chin, Albert Mo Kim Cheng, Mihai Lupu
IEEE Trans. Computers4
2005 Systematic Debugging of Real-Time Systems based on Incremental Satisfiability Counting
abstract
Real-time logic (RTL) (F. Jahanian et al., 1986, 1987, F. Wang et al., 1994) is useful for the verification of a safety assertion with respect to the specification of a real-time system. Since the satisfiability problem for RTL is undecidable, the systematic debugging of a real-time system appears impossible. This paper provides a first step towards this challenge. With RTL, each propositional formula corresponds to a verification condition. The number of truth assignments of a propositional formula helps to determine the timing constraints which should be added or modified to the system's specification. We have implemented a tool (called SDRTL, (S. Andrei et al., 2004)) that is able to perform systematic debugging. The confidence of our approach is high as we have evaluated SDRTL on several existing industrial-based applications.
Stefan Andrei, Albert Mo Kim Cheng, Wei-Ngan Chin, Mihai Lupu
IEEE Real-Time and Embedded Technology and Applications Symposium4