Jacques Savoy

dblp:s/JacquesSavoy · DBLP profile ↗
← Back
41ranked-venue papers
24as first author
1since 2021 · last 2022
0000-0002-4486-0067ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 37 · 21 first-author · 1 since 2021Artificial intelligence and machine learning · 5 · 2 first-authorHuman-computer interaction and ubiquitous computing · 2 · 2 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
2 papers
Information retrieval · 100%
Human-computer interaction and pervasive computing
1 paper
User interface design and tools · 77% Learning and educational technologies · 23%

Topics — the 6 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Information retrieval › text analysis › stylometry
authorship attribution
0.112012
Authorship Attribution Based on Specific Vocabulary · ACM Trans. Inf. Syst. 2012
Information retrieval
text analysis
0.112012
Authorship Attribution Based on Specific Vocabulary · ACM Trans. Inf. Syst. 2012
Information retrieval › document retrieval › structured document retrieval
hypertext retrieval
0.011993
Searching Information in Hypertext Systems Using Multiple Sources of Evidence · Int. J. Man Mach. Stud. 1993
User interface design and tools › interactive systems › document interaction
e-books
0.011989
The Electronic Book Ebook3 · Int. J. Man Mach. Stud. 1989
Information retrieval
retrieval models
0.011993
Searching Information in Hypertext Systems Using Multiple Sources of Evidence · Int. J. Man Mach. Stud. 1993
Learning and educational technologies › reading technology
reading interaction
0.011989
The Electronic Book Ebook3 · Int. J. Man Mach. Stud. 1989

Methods — techniques the papers use, named apart from their topics

z-score · 0.1naive bayes · 0.1distance measures · 0.1binomial distribution · 0.1evidence combination · 0.0system design · 0.0
YearPublicationVenuePosition
2022 Gender identification on Twitter
abstract
Abstract To determine the author of a text's gender, various feature types have been suggested (e.g., function words, n ‐gram of letters, etc.) leading to a huge number of stylistic markers. To determine the target category, different machine learning models have been suggested (e.g., logistic regression, decision tree, k nearest‐neighbors, support vector machine, naïve Bayes, neural networks, and random forest). In this study, our first objective is to know whether or not the same model always proposes the best effectiveness when considering similar corpora under the same conditions. Thus, based on 7 CLEF‐PAN collections, this study analyzes the effectiveness of 10 different classifiers. Our second aim is to propose a 2‐stage feature selection to reduce the feature size to a few hundred terms without any significant change in the performance level compared to approaches using all the attributes (increase of around 5% after applying the proposed feature selection). Based on our experiments, neural network or random forest tend, on average, to produce the highest effectiveness. Moreover, empirical evidence indicates that reducing the feature set size to around 300 without penalizing the effectiveness is possible. Finally, based on such reduced feature sizes, an analysis reveals some of the specific terms that clearly discriminate between the 2 genders.
Catherine Ikae, Jacques Savoy
J. Assoc. Inf. Sci. Technol.2
2019 Authorship of Pauline epistles revisited
abstract
The name Paul appears in 13 epistles, but is he the real author? According to different biblical scholars, the number of letters really attributed to Paul varies from 4 to 13, with a majority agreeing on seven. This article proposes to revisit this authorship attribution problem by considering two effective methods (Burrows' Delta, Labbé's intertextual distance). Based on these results, a hierarchical clustering is then applied showing that four clusters can be derived, namely: {Colossians‐Ephesians}, {1 and 2 Thessalonians}, {Titus, 1 and 2 Timothy}, and {Romans, Galatians, 1 and 2 Corinthians}. Moreover, a verification method based on the impostors' strategy indicates clearly that the group {Colossians‐Ephesians} is written by the same author who seems not to be Paul. The same conclusion can be found for the cluster {Titus, 1 and 2 Timothy}. The Letter to Philemon stays as a singleton, without any close stylistic relationship with the other epistles. Finally, a group of four letters {Romans, Galatians, 1 and 2 Corinthians} is certainly written by the same author (Paul), but the verification protocol also indicates that 2 Corinthians is related to 1 Thessalonians, rendering a clear and simple interpretation difficult.
Jacques Savoy
J. Assoc. Inf. Sci. Technol.1
2018 Working with Text. Tools, Techniques and Approaches for Text Mining. Emme L. Tonkin & Gregory J.L. Tourte. Chandos Publisher, Cambridge (MA). 2016. 330pp. (ISBN 978-1-84334-749-1)
Jacques Savoy
J. Assoc. Inf. Sci. Technol.1
2017 Distance measures in author profiling
Mirco Kocher, Jacques Savoy
Inf. Process. Manag.2
2017 A simple and efficient algorithm for authorship verification
abstract
This paper describes and evaluates an unsupervised and effective authorship verification model called Spatium‐L1. As features, we suggest using the 200 most frequent terms of the disputed text (isolated words and punctuation symbols). Applying a simple distance measure and a set of impostors, we can determine whether or not the disputed text was written by the proposed author. Moreover, based on a simple rule we can define when there is enough evidence to propose an answer or when the attribution scheme is unable to make a decision with a high degree of certainty. Evaluations based on 6 test collections (PAN CLEF 2014 evaluation campaign) indicate that Spatium‐L1 usually appears in the top 3 best verification systems, and on an aggregate measure, presents the best performance. The suggested strategy can be adapted without any problem to different Indo‐European languages (such as English, Dutch, Spanish, and Greek) or genres (essay, novel, review, and newspaper article).
Mirco Kocher, Jacques Savoy
J. Assoc. Inf. Sci. Technol.2
2016 Estimating the probability of an authorship attribution
abstract
In authorship attribution, various distance‐based metrics have been proposed to determine the most probable author of a disputed text. In this paradigm, a distance is computed between each author profile and the query text. These values are then employed only to rank the possible authors. In this article, we analyze their distribution and show that we can model it as a mixture of 2 Beta distributions. Based on this finding, we demonstrate how we can derive a more accurate probability that the closest author is, in fact, the real author. To evaluate this approach, we have chosen 4 authorship attribution methods (Burrows' Delta, Kullback‐Leibler divergence, Labbé's intertextual distance, and the naïve Bayes). As the first test collection, we have downloaded 224 State of the Union addresses (from 1790 to 2014) delivered by 41 U.S. presidents. The second test collection is formed by the Federalist Papers. The evaluations indicate that the accuracy rate of some authorship decisions can be improved. The suggested method can signal that the proposed assignment should be interpreted as possible, without strong certainty. Being able to quantify the certainty associated with an authorship decision can be a useful component when important decisions must be taken.
Jacques Savoy
J. Assoc. Inf. Sci. Technol.1
2016 Text representation strategies: An example with the State of the union addresses
abstract
Based on State of the Union addresses from 1790 to 2014 (225 speeches delivered by 42 presidents), this paper describes and evaluates different text representation strategies. To determine the most important words of a given text, the term frequencies (tf) or the tf idf weighting scheme can be applied. Recently, latent Dirichlet allocation (LDA) has been proposed to define the topics included in a corpus. As another strategy, this study proposes to apply a vocabulary specificity measure (Z score) to determine the most significantly overused word‐types or short sequences of them. Our experiments show that the simple term frequency measure is not able to discriminate between specific terms associated with a document or a set of texts. Using the tf idf or LDA approach, the selection requires some arbitrary decisions. Based on the term‐specific measure (Z score), the term selection has a clear theoretical basis. Moreover, the most significant sentences for each presidency can be determined. As another facet, we can visualize the dynamic evolution of usage of some terms associated with their specificity measures. Finally, this technique can be employed to define the most important lexical leaders introducing terms overused by the k following presidencies.
Jacques Savoy
J. Assoc. Inf. Sci. Technol.1
2015 Text clustering: An application with the State of the Union addresses
abstract
This paper describes a clustering and authorship attribution study over the State of the Union addresses from 1790 to 2014 (224 speeches delivered by 41 presidents). To define the style of each presidency, we have applied a principal component analysis (PCA) based on the part‐of‐speech (POS) frequencies. From Roosevelt (1934), each president tends to own a distinctive style whereas previous presidents tend usually to share some stylistic aspects with others. Applying an automatic classification based on the frequencies of all content‐bearing word‐types we show that chronology tends to play a central role in forming clusters, a factor that is more important than political affiliation. Using the 300 most frequent word‐types, we generate another clustering representation based on the style of each president. This second view shares similarities with the first one, but usually with more numerous and smaller clusters. Finally, an authorship attribution approach for each speech can reach a success rate of around 95.7% under some constraints. When an incorrect assignment is detected, the proposed author often belongs to the same party and has lived during roughly the same time period as the presumed author. A deeper analysis of some incorrect assignments reveals interesting reasons justifying difficult attributions.
Jacques Savoy
J. Assoc. Inf. Sci. Technol.1
2013 Authorship attribution based on a probabilistic topic model
Jacques Savoy
Inf. Process. Manag.1
2012 Authorship Attribution Based on Specific Vocabulary
abstract
In this article we propose a technique for computing a standardized Z score capable of defining the specific vocabulary found in a text (or part thereof) compared to that of an entire corpus. Assuming that the term occurrence follows a binomial distribution, this method is then applied to weight terms (words and punctuation symbols in the current study), representing the lexical specificity of the underlying text. In a final stage, to define an author profile we suggest averaging these text representations and then applying them along with a distance measure to derive a simple and efficient authorship attribution scheme. To evaluate this algorithm and demonstrate its effectiveness, we develop two experiments, the first based on 5,408 newspaper articles (Glasgow Herald) written in English by 20 distinct authors and the second on 4,326 newspaper articles (La Stampa) written in Italian by 20 distinct authors. These experiments demonstrate that the suggested classification scheme tends to perform better than the Delta rule method based on the most frequent words, better than the chi-square distance based on word profiles and punctuation marks, better than the KLD scheme based on a predefined set of words, and better than the naïve Bayes approach.
Jacques Savoy
ACM Trans. Inf. Syst.1
2011 Classification Based on Specific Vocabulary
abstract
Assuming a binomial distribution for word occurrence, we propose computing a standardized Z score to define the specific vocabulary of a subset compared to that of the entire corpus. This approach is applied to weight terms characterizing a document (or a sample of texts). We then show how these Z score values can be used to derive an efficient categorization scheme. To evaluate this proposition we categorize speeches given by B. Obama as either electoral or presidential. The results tend to show that the suggested classification scheme performs better than a Support Vector Machine scheme, and a Naive Bayes classifier (10-fold cross validation).
Jacques Savoy, Olena Zubaryeva
Web Intelligence1
2010 When stopword lists make the difference
abstract
Abstract In this brief communication, we evaluate the use of two stopword lists for the English language (one comprising 571 words and another with 9) and compare them with a search approach accounting for all word forms. We show that through implementing the original Okapi form or certain ones derived from the Divergence from Randomness (DFR) paradigm, significantly lower performance levels may result when using short or no stopword lists. For other DFR models and a revised Okapi implementation, performance differences between approaches using short or long stopword lists or no list at all are usually not statistically significant. Similar conclusions can be drawn when using other natural languages such as French, Hindi, or Persian.
Ljiljana Dolamic, Jacques Savoy
J. Assoc. Inf. Sci. Technol.2
2010 Retrieval effectiveness of machine translated queries
abstract
Abstract This article describes and evaluates various information retrieval models used to search document collections written in English through submitting queries written in various other languages, either members of the Indo‐European family (English, French, German, and Spanish) or radically different language groups such as Chinese. This evaluation method involves searching a rather large number of topics (around 300) and using two commercial machine translation systems to translate across the language barriers. In this study, mean average precision is used to measure variances in retrieval effectiveness when a query language differs from the document language. Although performance differences are rather large for certain languages pairs, this does not mean that bilingual search methods are not commercially viable. Causes of the difficulties incurred when searching or during translation are analyzed and the results of concrete examples are explained.
Ljiljana Dolamic, Jacques Savoy
J. Assoc. Inf. Sci. Technol.2
2010 Comparative Study of Indexing and Search Strategies for the Hindi, Marathi, and Bengali Languages
abstract
The main goal of this article is to describe and evaluate various indexing and search strategies for the Hindi, Bengali, and Marathi languages. These three languages are ranked among the world’s 20 most spoken languages and they share similar syntax, morphology, and writing systems. In this article we examine these languages from an Information Retrieval (IR) perspective through describing the key elements of their inflectional and derivational morphologies, and suggest a light and more aggressive stemming approach based on them. In our evaluation of these stemming strategies we make use of the FIRE 2008 test collections, and then to broaden our comparisons we implement and evaluate two language independent indexing methods: the n -gram and trunc- n (truncation of the first n letters). We evaluate these solutions by applying our various IR models, including the Okapi, Divergence from Randomness (DFR) and statistical language models (LM) together with two classical vector-space approaches: tf idf and Lnu-ltc . Experiments performed with all three languages demonstrate that the I(n e )C2 model derived from the Divergence from Randomness paradigm tends to provide the best mean average precision (MAP). Our own tests suggest that improved retrieval effectiveness would be obtained by applying more aggressive stemmers, especially those accounting for certain derivational suffixes, compared to those involving a light stemmer or ignoring this type of word normalization procedure. Comparisons between no stemming and stemming indexing schemes shows that performance differences are almost always statistically significant. When, for example, an aggressive stemmer is applied, the relative improvements obtained are ~28% for the Hindi language, ~42% for Marathi, and ~18% for Bengali, as compared to a no-stemming approach. Based on a comparison of word-based and language-independent approaches we find that the trunc-4 indexing scheme tends to result in performance levels statistically similar to those of an aggressive stemmer, yet better than the 4-gram indexing scheme. A query-by-query analysis reveals the reasons for this, and also demonstrates the advantage of applying a stemming or a trunc-4 indexing scheme.
Ljiljana Dolamic, Jacques Savoy
ACM Trans. Asian Lang. Inf. Process.2
2009 Indexing and stemming approaches for the Czech language
Ljiljana Dolamic, Jacques Savoy
Inf. Process. Manag.2
2009 Indexing and searching strategies for the Russian language
abstract
Abstract This paper describes and evaluates various stemming and indexing strategies for the Russian language. We design and evaluate two stemming approaches, a light and a more aggressive one, and compare these stemmers to the Snowball stemmer, to no stemming, and also to a language‐independent approach (n‐gram). To evaluate the suggested stemming strategies we apply various probabilistic information retrieval (IR) models, including the Okapi, the Divergence from Randomness (DFR), a statistical language model (LM), as well as two vector‐space approaches, namely, the classical tf idf scheme and the dtu‐dtn model. We find that the vector‐space dtu‐dtn and the DFR models tend to result in better retrieval effectiveness than the Okapi, LM, or tf idf models, while only the latter two IR approaches result in statistically significant performance differences. Ignoring stemming generally reduces the MAP by more than 50%, and these differences are always significant. When applying an n‐gram approach, performance differences are usually lower than an approach involving stemming. Finally, our light stemmer tends to perform best, although performance differences between the light, aggressive, and Snowball stemmers are not statistically significant.
Ljiljana Dolamic, Jacques Savoy
J. Assoc. Inf. Sci. Technol.2
2009 Algorithmic stemmers or morphological analysis? An evaluation
abstract
Abstract It is important in information retrieval (IR), information extraction, or classification tasks that morphologically related forms are conflated under the same stem (using stemmer) or lemma (using morphological analyzer). To achieve this for the English language, algorithmic stemming or various morphological analysis approaches have been suggested. Based on Cross‐Language Evaluation Forum test collections containing 284 queries and various IR models, this article evaluates these word‐normalization proposals. Stemming improves the mean average precision significantly by around 7% while performance differences are not significant when comparing various algorithmic stemmers or algorithmic stemmers and morphological analysis. Accounting for thesaurus class numbers during indexing does not modify overall retrieval performances. Finally, we demonstrate that including a stop word list, even one containing only around 10 terms, might significantly improve retrieval performance, depending on the IR model.
Claire Fautsch, Jacques Savoy
J. Assoc. Inf. Sci. Technol.2
2008 Searching in Medline: Query expansion and manual indexing evaluation
Samir Abdou, Jacques Savoy
Inf. Process. Manag.2
2008 Searching strategies for the Hungarian language
Jacques Savoy
Inf. Process. Manag.1
2007 Searching strategies for the Bulgarian language
Jacques Savoy
Inf. Retr.1
2005 Features Combination for Extracting Gene Functions from MEDLINE
Patrick Ruch, Laura Perret, Jacques Savoy
ECIR3
2005 Bibliographic database access using free-text and controlled vocabulary: an evaluation
Jacques Savoy
Inf. Process. Manag.1
2005 Comparative study of monolingual and multilingual search models for use with asian languages
abstract
Based on the NTCIR-4 test-collection, our first objective is to present an overview of the retrieval effectiveness of nine vector-space and two probabilistic models that perform monolingual searches in the Chinese, Japanese, Korean, and English languages. Our second goal is to analyze the relative merits of the various automated and freely available toolsto translate the English-language topics into Chinese, Japanese, or Korean, and then submit the resultant query in order to retrieve pertinent documents written in one of the three Asian languages. We also demonstrate how bilingual searches could be improved by applying both the combined query translation strategies and data-fusion approaches. Finally, we address basic problems related to multilingual searches, in which queries written in English are used to search documents written in the English, Chinese, Japanese, and Korean languages.
Jacques Savoy
ACM Trans. Asian Lang. Inf. Process.1
2004 Combining Multiple Strategies for Effective Monolingual and Cross-Language Retrieval
Jacques Savoy
Inf. Retr.1
2003 Term Proximity Scoring for Keyword-Based Retrieval Systems
Yves Rasolofo, Jacques Savoy
ECIR2
2003 Result merging strategies for a current news metasearcher
Yves Rasolofo, David Hawking, Jacques Savoy
Inf. Process. Manag.3
2003 Cross-language information retrieval: experiments based on CLEF 2000 corpora
Jacques Savoy
Inf. Process. Manag.1
2003 Enhancing retrieval with hyperlinks: A general model based on propositional argumentation systems
abstract
Abstract Fast, effective, and adaptable techniques are needed to automatically organize and retrieve information on the ever‐increasing World Wide Web. In that respect, different strategies have been suggested to take hypertext links into account. For example, hyperlinks have been used to (1) enhance document representation, (2) improve document ranking by propagating document score, (3) provide an indicator of popularity, and (4) find hubs and authorities for a given topic. Although the TREC experiments have not demonstrated the usefulness of hyperlinks for retrieval, the hypertext structure is nevertheless an essential aspect of the Web, and as such, should not be ignored. The development of abstract models of the IR task was a key factor to the improvement of search engines. However, at this time conceptual tools for modeling the hypertext retrieval task are lacking, making it difficult to compare, improve, and reason on the existing techniques. This article proposes a general model for using hyperlinks based on Probabilistic Argumentation Systems, in which each of the above‐mentioned techniques can be stated. This model will allow to discover some inconsistencies in the mentioned techniques, and to take a higher level and systematic approach for using hyperlinks for retrieval.
Justin Picard, Jacques Savoy
J. Assoc. Inf. Sci. Technol.2
2001 Approaches to Collection Selection and Results Merging for Distributed Information Retrieval
abstract
We have investigated two major issues in Distributed Information Retrieval (DIR), namely: collection selection and search results merging. While most published works on these two issues are based on pre-stored metadata, the approaches described in this paper involve extracting the required information at the time the query is processed. In order to predict the relevance of collections to a given query, we analyse a limited number of full documents (e.g., the top five documents) retrieved from each collection and then consider term proximity within them. On the other hand, our merging technique is rather simple since input only requires document scores and lengths of results lists. Our experiments evaluate the retrieval effectiveness of these approaches and compare them with centralised indexing and various other DIR techniques (e.g., CORI). We conducted our experiments using two testbeds: one containing news articles extracted from four different sources (2 GB) and another containing 10 GB of Web pages. Our evaluations demonstrate that the retrieval effectiveness of our simple approaches is worth considering.
Yves Rasolofo, Faïza Abbaci, Jacques Savoy
CIKM3
2001 Retrieval effectiveness on the web
Jacques Savoy, Justin Picard
Inf. Process. Manag.1
2000 Database merging strategy based on logistic regression
Anne Le Calvé, Jacques Savoy
Inf. Process. Manag.2
1999 A Stemming Procedure and Stopword List for General French Corpora
abstract
Due to the increasing use of network-based systems, there is a growing interest in access to and search mechanisms for text databases in languages other than English. To adapt searching systems to those foreign languages with characteristics similar to the English language, all we need to do for the most part is to establish a general stopword list and a stemming procedure. This article presents the tools needed to establish these in the French language databases and some retrieval experiments that have been carried out using two medium-sized French language test collections. These experiments were conducted to evaluate the retrieval effectiveness of the propositions described.
Jacques Savoy
J. Am. Soc. Inf. Sci.1
1997 Statistical inference in retrieval effectiveness evaluation
Jacques Savoy
Inf. Process. Manag.1
1997 Ranking Schemes in Hybrid Boolean Systems: A New Approach
abstract
In most commercial online systems, the retrieval system is based on the Boolean model and its inverted file organization. Since the investment in these systems is so great and changing them could be economically unfeasible, this article suggests a new ranking scheme especially adapted for hypertext environments in order to produce more effective retrieval results and yet maintain the effectiveness of the investment made to date in the Boolean model. To select the retrieved documents, the suggested ranking strategy uses multiple sources of document content evidence. The proposed scheme integrates both the information provided by the index and query terms, and the inherent relationships between documents such as bibliographic references or hypertext links. We will demonstrate that our scheme represents an integration of both subject and citation indexing, and results in a significant improvement over classical ranking schemes used in hybrid Boolean systems, while preserving its efficiency. Moreover, through knowing the nearest neighbor and the hypertext links which constitute additional sources of evidence, our strategy will take them into account in order to further improve retrieval effectiveness and to provide “good” starting points for browsing in a hypertext or hypermedia environment. © 1997 John Wiley & Sons, Inc.
Jacques Savoy
J. Am. Soc. Inf. Sci.1
1996 An Extended Vector-Processing Scheme for Searching Information in Hypertext Systems
Jacques Savoy
Inf. Process. Manag.1
1994 A Learning Scheme for Information Retrieval in Hypertext
Jacques Savoy
Inf. Process. Manag.1
1993 Searching Information in Hypertext Systems Using Multiple Sources of Evidence
Jacques Savoy
Int. J. Man Mach. Stud.1
1993 Stemming of French Words Based on Grammatical Categories
abstract
Automatic indexing systems use suffix stripping algorithms to cluster various words derived from a common root under the same stem. Currently, removing affixes to either a context-free or context-sensitive operation, where the context refers to the remaining stem. In this article, we propose a suffixing algorithm which uses grammatical categories to enhance the stemming process. This approach supports the use of foreign languages. In our case, the language is French, and a morphological analysis is required for removing inflectional suffixes or morphosyntactic variants of a lemma. After this analysis, we implement a suffix stripping algorithm which uses a dictionary and the grammatical categories to remove derivational suffixes. Our approach always returns a linguistically correct lemma, but not necessarily the “right” one. Based on our tests, this solution is an attractive one, with a mean error rate of 16%. We finish by explaining why we cannot expect significantly better results with this approach. © 1993 John Wiley & Sons, Inc.
Jacques Savoy
J. Am. Soc. Inf. Sci.1
1992 Bayesian Inference Networks and Spreading Activation in Hypertext Systems
Jacques Savoy
Inf. Process. Manag.1
1991 The effect of change on legal applications
Paul Bratley, Daniel Poulin, Jacques Savoy
DEXA3
1989 The Electronic Book Ebook3
Jacques Savoy
Int. J. Man Mach. Stud.1