Claudio Carpineto

dblp:92/5215 · DBLP profile ↗
← Back
26ranked-venue papers
23as first author
0since 2021 · last 2020
0009-0003-1051-9308ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 15 · 12 first-authorArtificial intelligence and machine learning · 12 · 10 first-authorGraphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 first-authorTheory of computation · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
6 papers
Data mining · 62% Information retrieval · 38% Machine learning and data management · 0%
Artificial intelligence
3 papers
Knowledge representation and reasoning · 72% Language models and text generation · 28%

Topics — the 13 heaviest of 14, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Data mining
clustering
0.112012
Consensus Clustering Based on a New Probabilistic Rand Index with Application to Subtopic Retrieval · IEEE Trans. Pattern Anal. Mach. Intell. 2012
Data mining › clustering
ensemble clustering
0.112012
Consensus Clustering Based on a New Probabilistic Rand Index with Application to Subtopic Retrieval · IEEE Trans. Pattern Anal. Mach. Intell. 2012
Data mining › clustering › interpretable clustering
cluster labeling
0.112010
Optimal meta search results clustering · SIGIR 2010
Information retrieval › search engines
search result clustering
0.112010
Optimal meta search results clustering · SIGIR 2010
Information retrieval › query reformulation
query expansion
0.122002
Improving retrieval feedback with multiple term-ranking function combination · ACM Trans. Inf. Syst. 2002
An information-theoretic approach to automatic query expansion · ACM Trans. Inf. Syst. 2001
Information retrieval › diversified retrieval
subtopic retrieval
0.012012
Consensus Clustering Based on a New Probabilistic Rand Index with Application to Subtopic Retrieval · IEEE Trans. Pattern Anal. Mach. Intell. 2012
Information retrieval › retrieval models › associative retrieval
lattice-based retrieval
0.011996
Information retrieval through hybrid navigation of lattice representations · Int. J. Hum. Comput. Stud. 1996
Knowledge, reasoning and agents › Knowledge representation and reasoning › concept learning
conceptual clustering
0.011993
GALOIS: An Order-Theoretic Approach to Conceptual Clustering · ICML 1993
Natural language and speech › Language models and text generation › text generation
story generation
0.011988
An Approach Based on Integrated Learning to Generating Stories · ML 1988
Information retrieval › interactive information retrieval
browsing
0.011996
Information retrieval through hybrid navigation of lattice representations · Int. J. Hum. Comput. Stud. 1996
Logic in computer science › category theory
galois connections
0.011993
GALOIS: An Order-Theoretic Approach to Conceptual Clustering · ICML 1993
Combinatorics and discrete mathematics
partial orders
0.011993
GALOIS: An Order-Theoretic Approach to Conceptual Clustering · ICML 1993
Knowledge, reasoning and agents › Knowledge representation and reasoning
concept learning
0.011992
Trading Off Consistency and Efficiency in version-Space Induction · ML 1992

Methods — techniques the papers use, named apart from their topics

stochastic optimization · 0.3probabilistic rand index · 0.1probabilistic concordance · 0.1rank-based ensemble combination · 0.0information theory · 0.0order-theoretic clustering · 0.0formal concept analysis · 0.0integrated learning · 0.0consistency-efficiency tradeoff · 0.0
YearPublicationVenuePosition
2020 An Experimental Study of Automatic Detection and Measurement of Counterfeit in Brand Search Results
abstract
Brand search results are poisoned by fake ecommerce websites that infringe on the trademark rights of legitimate holders. In this article, we study how to tackle and measure this problem automatically. We present a pipeline with two machine learning stages that can detect the ecommerce websites present in the list of brand search results and distinguish between legitimate and fake ecommerce websites. For each classification task, we identify and extract suitable learning features and study their relative importance. Through a prototype system termed RI.SI.CO., we show that this approach is feasible, fast, and more accurate than both existing systems for trustworthiness assessment and non-expert humans. We next introduce two complementary metrics for evaluating the counterfeit incidence in brand search results: namely, a chart-based and a single-value measure. They allow us to analyze and compare counterfeit at various levels, including single brands within a specific sector as well as whole sectors. Experimenting with two luxury goods sectors, we report a number of interesting findings about how the main search parameters (e.g., search engine, query type, number of search results seen) affect counterfeiting and how this activity changes with time. On the whole, our research offers new insights and some very practical and useful means of analyzing and measuring counterfeit in brand search results, thus increasing awareness of and knowledge about this phenomenon and enabling targeted anti-counterfeiting actions.
Claudio Carpineto, Giovanni Romano 0002
ACM Trans. Web1
2017 Learning to detect and measure fake ecommerce websites in search-engine results
abstract
When searching for a brand name in search engines, it is very likely to come across websites that sell fake brand's products. In this paper, we study how to tackle and measure this problem automatically. Our solution consists of a pipeline with two learning stages. We first detect the ecommerce websites (including shopbots) present in the list of search results and then discriminate between legitimate and fake ecommerce websites. We identify suitable learning features for each stage and show through a prototype system termed RI.SI.CO. that this approach is feasible, fast, and highly effective. Experimenting with one goods sector, we found that RI.SI.CO. achieved better classification accuracy than that of non-expert humans. We next show that the information extracted by our method can be used to generate sector-level 'counterfeiting charts' that allow us to analyze and compare the counterfeit risk associated with different brands in a same sector. We also show that the risk of coming across counterfeit websites is affected by the particular web search engine and type of search query used by shoppers. Our research offers new insights and some very practical and useful means for analyzing and measuring counterfeit ecommerce websites in search-engine results, thus enabling targeted anti-counterfeiting actions.
Claudio Carpineto, Giovanni Romano 0002
WI1
2016 Enhancing User Awareness and Control of Web Tracking with ManTra
abstract
Web trackers can build accurate topical user profiles (e.g., in terms of habits and personal characteristics) by monitoring a user's browsing activities across websites. This process, known as behavioral targeting, has a number of practical benefits but it also raises privacy concerns. Most existing techniques either try to block web tracking altogether or aim to endow it with privacy preserving mechanisms, but they are system-centered rather than user-centered. Nowadays, the majority of users want to have some degree of control over their privacy, while their perspectives and feelings towards web tracking maybe different, ranging from a desire to avoid being profiled at all to a willingness to trade personal information for better services. Regardless of a specific user's preference, from a technical point of view there is is no simple way for him/her to monitor, let alone to influence, the behavior of web trackers. In this paper, we describe an approach which makes users aware of their likely tracking profile and gives them the possibility to bias the profile towards both ends of the web tracking spectrum, either by improving its accuracy beyond the tracker capabilities (thus emphasizing behavioral targeting) or by filling in false interests (thus increasing privacy). This goal is achieved by simulating the process of learning a user profile on the part of the tracker and then by retrofitting a web traffic suitable for producing the desired profile. Our approach has been implemented as a web browser extension called ManTra (Management of Tracking). The system has been evaluated in several dimensions, including its ability to learn an accurate ad-oriented user profile and to influence the behavior of a commercial tool for web tracking personalization, i.e., Google's Ads Settings.
Davide Lo Re, Claudio Carpineto
WI2
2015 KΘ-affinity privacy: Releasing infrequent query refinements safely
Claudio Carpineto, Giovanni Romano 0002
Inf. Process. Manag.1
2013 Semantic Search Log k-Anonymization with Generalized k-Cores of Query Concept Graph
Claudio Carpineto, Giovanni Romano 0002
ECIR1
2012 Evaluating subtopic retrieval methods: Clustering versus diversification of search results
Claudio Carpineto, Massimiliano D'Amico, Giovanni Romano 0002
Inf. Process. Manag.1
2012 Consensus Clustering Based on a New Probabilistic Rand Index with Application to Subtopic Retrieval
abstract
We introduce a probabilistic version of the well-known Rand Index (RI) for measuring the similarity between two partitions, called Probabilistic Rand Index (PRI), in which agreements and disagreements at the object-pair level are weighted according to the probability of their occurring by chance. We then cast consensus clustering as an optimization problem of the PRI value between a target partition and a set of given partitions, experimenting with a simple and very efficient stochastic optimization algorithm. Remarkable performance gains over input partitions as well as over existing related methods are demonstrated through a range of applications, including a new use of consensus clustering to improve subtopic retrieval.
Claudio Carpineto, Giovanni Romano 0002
IEEE Trans. Pattern Anal. Mach. Intell.1
2011 Full discrimination of subtopics in search results with keyphrase-based clustering
abstract
We consider the problem of retrieving multiple documents relevant to the single subtopics of a given web query, termed “full-subtopic retrieval”. To solve this problem we present a novel search results clustering algorithm that generates clusters lab
Claudio Carpineto, Massimiliano D'Amico
Web Intell. Agent Syst.1
2010 Optimal meta search results clustering
abstract
By analogy with merging documents rankings, the outputs from multiple search results clustering algorithms can be combined into a single output. In this paper we study the feasibility of meta search results clustering, which has unique features compared to the general meta clustering problem. After showing that the combination of multiple search results clusterings is empirically justified, we cast meta clustering as an optimization problem of an objective function measuring the probabilistic concordance between the clustering combination and the single clusterings. We then show, using an easily computable upper bound on such a function, that a simple stochastic optimization algorithm delivers reasonable approximations of the optimal value very efficiently, and we also provide a method for labeling the generated clusters with the most agreed upon cluster labels. Optimal meta clustering with meta labeling is applied to three descriptioncentric, state-of-the-art search results clustering algorithms. The performance improvement is demonstrated through a range of evaluation techniques (i.e., internal, classificationoriented, and information retrieval-oriented), using suitable test collections of search results with document-level relevance judgments per subtopic.
Claudio Carpineto, Giovanni Romano 0002
SIGIR1
2009 A Concept Lattice-Based Kernel for SVM Text Classification
Claudio Carpineto, Carla Michini, Raffaele Nicolussi
ICFCA1
2009 Full-Subtopic Retrieval with Keyphrase-Based Search Results Clustering
abstract
We consider the problem of retrieving multiple documents relevant to the single subtopics of a given web query, termed "full-subtopic retrieval". To solve this problem we present a novel search results clustering algorithm that generates clusters labeled by keyphrases. The keyphrases are extracted from the generalized suffix tree built from the search results and merged through an improved hierarchical agglomerative clustering procedure. We also introduce a novel measure for evaluating full-subtopic retrieval performance, namely "Subtopic Search Length under k document sufficiency". Using a test collection specifically designed for evaluating subtopic retrieval, we found that our algorithm outperformed both other existing search results clustering algorithms and also a search results re-ranking method that emphasized diversity of results (at least for k≫1; i.e., when we are interested in retrieving more than one relevant document per subtopic). Our approach has been implemented into KeySRC (Keyphrase-based Search Results Clustering), a full web clustering engine available online at http://keysrc.fub.it.
Claudio Carpineto, Massimiliano D'Amico
Web Intelligence2
2009 Mobile information retrieval with search results clustering: Prototypes and evaluations
abstract
Abstract Web searches from mobile devices such as PDAs and cell phones are becoming increasingly popular. However, the traditional list‐based search interface paradigm does not scale well to mobile devices due to their inherent limitations. In this article, we investigate the application of search results clustering, used with some success for desktop computer searches, to the mobile scenario. Building on CREDO (Conceptual Reorganization of Documents), a Web clustering engine based on concept lattices, we present its mobile versions Credino and SmartCREDO, for PDAs and cell phones, respectively. Next, we evaluate the retrieval performance of the three prototype systems. We measure the effectiveness of their clustered results compared to a ranked list of results on a subtopic retrieval task, by means of the device‐independent notion of subtopic reach time together with a reusable test collection built from Wikipedia ambiguous entries. Then, we make a cross‐comparison of methods (i.e., clustering and ranked list) and devices (i.e., desktop, PDA, and cell phone), using an interactive information‐finding task performed by external participants. The main finding is that clustering engines are a viable complementary approach to plain search engines both for desktop and mobile searches especially, but not only, for multitopic informational queries.
Claudio Carpineto, Stefano Mizzaro, Giovanni Romano 0002, Matteo Snidero
J. Assoc. Inf. Sci. Technol.1
2006 Mobile Clustering Engine
Claudio Carpineto, Andrea Della Pietra, Stefano Mizzaro, Giovanni Romano 0002
ECIR1
2004 Query Difficulty, Robustness, and Selective Application of Query Expansion
Gianni Amati, Claudio Carpineto, Giovanni Romano 0002
ECIR2
2003 Mining Short-Rule Covers in Relational Databases
abstract
An implication ruleQ→Ris a statement of the form “for all objects in the database, if an object has the attribute–value pairsQthen it has also the attribute–value pairsR.” This simple type of rule is theoretically interesting, because it supports reasoning, similar to functional dependencies in database theory, and it may be of practical significance because the size of the set of implication rules that hold in a relation can remain substantially high even when mining real data and considering only most general covers; i.e., covers containing rules with unredundant right and left sizes. Motivated by these observations, we focus on the extraction of short‐rule covers, which cannot be efficiently mined by standard rule miners. We present an algorithm driven by “negative examples” (i.e., satisfyQbut notR) to prune the rule‐candidate lattice associated with each “positive example” (i.e., satisfies bothQandR). The algorithm scales up quite well with respect to the number of objects and it is particularly suitable for databases with attributes described by large domains. Furthermore, a perfect hash function ensures extraction of short‐rule covers even from databases containing a large number of attributes.
Claudio Carpineto, Giovanni Romano 0002
Comput. Intell.1
2002 Improving retrieval feedback with multiple term-ranking function combination
abstract
In this article we consider methods for automatic query expansion from top retrieved documents (i.e., retrieval feedback) that make use of various functions for scoring expansion terms within Rocchio's classical reweighting scheme. An analytical comparison shows that the retrieval performance of methods based on distinct term-scoring functions is comparable on the whole query set but differs considerably on single queries, consistent with the fact that the ordered sets of expansion terms suggested for each query by the different functions are largely uncorrelated. Motivated by these findings, we argue that the results of multiple functions can be merged, by analogy with ensembling classifiers, and present a simple combination technique based on the rank values of the suggested terms. The combined retrieval feedback method is effective not only with respect to unexpanded queries but also to any individual method, with notable improvements on the system's precision. Furthermore, the combined method is robust with respect to variation of experimental parameters and it is beneficial even when the same information needs are expressed with shorter queries.
Claudio Carpineto, Giovanni Romano 0002, Vittorio Giannini
ACM Trans. Inf. Syst.1
2001 An information-theoretic approach to automatic query expansion
abstract
Techniques for automatic query expansion from top retrieved documents have shown promise for improving retrieval effectiveness on large collections; however, they often rely on an empirical ground, and there is a shortage of cross-system comparisons. Using ideas from Information Theory, we present a computationally simple and theoretically justified method for assigning scores to candidate expansion terms. Such scores are used to select and weight expansion terms within Rocchio's framework for query reweigthing. We compare ranking with information-theoretic query expansion versus ranking with other query expansion techniques, showing that the former achieves better retrieval effectiveness on several performance measures. We also discuss the effect on retrieval effectiveness of the main parameters involved in automatic query expansion, such as data sparseness, query difficulty, number of selected documents, and number of selected terms, pointing out interesting relationships.
Claudio Carpineto, Renato De Mori, Giovanni Romano 0002, Brigitte Bigi
ACM Trans. Inf. Syst.1
2000 Order-theoretical ranking
abstract
Current best-match ranking (BMR) systems perform well but cannot handle word mismatch between a query and a document. The best known alternative ranking method, hierarchical clustering-based ranking (HCR), seems to be more robust than BMR with respect to this problem, but it is hampered by theoretical and practical limitations. We present an approach to document ranking that explicitly addresses the word mismatch problem by exploiting interdocument similarity information in a novel way. Document ranking is seen as a query-document transformation driven by a conceptual representation of the whole document collection, into which the query is merged. Our approach is based on the theory of concept (or Galois) lattices, which, we argue, provides a powerful, well-founded, and computationally-tractable framework to model the space in which documents and query are represented and to compute such a transformation. We compared information retrieval using concept lattice-based ranking (CLR) to BMR and HCR. The results showed that HCR was outperformed by CLR as well as by BMR, and suggested that, of the two best methods, BMR achieved better performance than CLR on the whole document set, whereas CLR compared more favorably when only the first retrieved documents were used for evaluation. We also evaluated the three methods' specific ability to rank documents that did not match the query, in which case the superiority of CLR over BMR and HCR (and that of HCR over BMR) was apparent.
Claudio Carpineto, Giovanni Romano 0002
J. Am. Soc. Inf. Sci.1
1998 Effective Reformulation of Boolean Queries with Concept Lattices
Claudio Carpineto, Giovanni Romano 0002
FQAS1
1996 Information retrieval through hybrid navigation of lattice representations
Claudio Carpineto, Giovanni Romano 0002
Int. J. Hum. Comput. Stud.1
1996 A Lattice Conceptual Clustering System and Its Application to Browsing Retrieval
Claudio Carpineto, Giovanni Romano 0002
Mach. Learn.1
1993 GALOIS: An Order-Theoretic Approach to Conceptual Clustering
Claudio Carpineto, Giovanni Romano 0002
ICML1
1992 Shift of Bias without Operators
Claudio Carpineto
ECAI1
1992 Trading Off Consistency and Efficiency in version-Space Induction
Claudio Carpineto
ML1
1990 Combining EBL from Success and EBL from Failure with Parameter Version Spaces
Claudio Carpineto
ECAI1
1988 An Approach Based on Integrated Learning to Generating Stories
Claudio Carpineto
ML1