Julien Ah-Pine

dblp:79/2507 · DBLP profile ↗
← Back
19ranked-venue papers
13as first author
3since 2021 · last 2026
0000-0001-6898-3961ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 15 · 10 first-author · 3 since 2021Databases, data management, data science and information retrieval · 9 · 7 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
2 papers
Data mining · 92% Data stream processing · 8%
Computer graphics and multimedia
1 paper
Multimedia analysis and retrieval · 100%

Topics — the 8 heaviest of 11, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Data mining
anomaly detection
1.012026
OnlineBootKNN: An Unsupervised Framework for Detecting Anomalies in Spectral Data Streams · AAAI 2026
Data mining › anomaly detection
streaming anomaly detection
1.012026
OnlineBootKNN: An Unsupervised Framework for Detecting Anomalies in Spectral Data Streams · AAAI 2026
Data mining › clustering › hierarchical clustering
agglomerative clustering
0.312018
An Efficient and Effective Generic Agglomerative Hierarchical Clustering Approach · J. Mach. Learn. Res. 2018
Data mining
clustering
0.312018
An Efficient and Effective Generic Agglomerative Hierarchical Clustering Approach · J. Mach. Learn. Res. 2018
Data mining › clustering
hierarchical clustering
0.312018
An Efficient and Effective Generic Agglomerative Hierarchical Clustering Approach · J. Mach. Learn. Res. 2018
Data mining › clustering
kernel clustering
0.312018
An Efficient and Effective Generic Agglomerative Hierarchical Clustering Approach · J. Mach. Learn. Res. 2018
Multimedia analysis and retrieval › multimedia retrieval › content-based retrieval
content-based multimedia retrieval
0.212015
Unsupervised Visual and Textual Information Fusion in CBMIR Using Graph-Based Methods · ACM Trans. Inf. Syst. 2015
Multimedia analysis and retrieval
multimodal fusion
0.212015
Unsupervised Visual and Textual Information Fusion in CBMIR Using Graph-Based Methods · ACM Trans. Inf. Syst. 2015

Methods — techniques the papers use, named apart from their topics

online bootstrapping · 1.0k-nearest neighbor · 1.0autoencoder · 1.0sparsified normalized kernel matrix · 0.3lance-williams clustering · 0.3inner product similarity · 0.3random walk · 0.2
YearPublicationVenuePosition
2026 OnlineBootKNN: An Unsupervised Framework for Detecting Anomalies in Spectral Data Streams
abstract
Monitoring the elemental composition of materials in order to detect abnormal conditions in real-time is essential for applications like manufacturing quality control, environmental monitoring, and space exploration. This is achieved using sensors that analyze the interaction of a material with electromagnetic radiation, producing spectral data streams or a sequence of instances where each represents an ordered set of wavelengths with an associated intensity. While many unsupervised anomaly detection methods exist for tabular streaming data, their applicability to spectral streams remains underexplored. To address this gap, we consider our spectra in a multivariate stream setting and benchmark the performance of state-of-the-art tabular anomaly detection methods on this data. Furthermore, we introduce OnlineBootKNN, a novel unsupervised framework that combines k-nearest neighbors with online bootstrapping and a z-score test to detect anomalies in real-time. We demonstrate the high performance and robustness of our method, as well as the efficacy of the autoencoder-based method, KitNet, on newly simulated real-world spectral datasets. In addition, we compare their efficiency against the other tested techniques. Finally, we highlight the inherent interpretability of OnlineBootKNN, which is crucial for identifying the specific wavelengths, and thus elements, responsible for a detected anomaly.
Nicolas Rojas Varela, Julien Ah-Pine, Engelbert Mephu Nguifo
AAAI2
2025 Mixed data k-Anonymization by Consistent Maximal Association and Microaggregation
abstract
This paper addresses the challenge of anonymizing mixed data, comprising both categorical (qualitative) and numerical (continuous) variables, while preserving data utility. The inherent heterogeneity of such data complicates the use of traditional anonymization methods. To overcome this limitation, we propose a novel microaggregation-based framework for k-anonymization that integrates statistical association measures applicable to both variable types, ensuring a coherent and consistent treatment. Our approach, called Mix-R 2, relies on a unified set of core concepts grounded in analysis of variance, enabling the application of a common methodology to both categorical and numerical attributes. By leveraging these consistent association measures, the framework improves the robustness of the k-anonymization process, delivering strong privacy protection while maintaining high data utility. Numerical experiments on benchmark datasets demonstrate the effectiveness and advantages of our method, highlighting its contribution to privacy-preserving analysis of mixed-type data.
Julien Ah-Pine, Nathaniel Gbenro
CIKM1
2025 On using derivatives and multiple kernel methods for clustering and classifying functional data
Julien Ah-Pine, Anne-Françoise Yao
Neurocomputing1
2018 An Efficient and Effective Generic Agglomerative Hierarchical Clustering Approach
abstract
We introduce an agglomerative hierarchical clustering (AHC) framework which is generic, efficient and effective. Our approach embeds a sub-family of Lance-Williams (LW) clusterings and relies on inner-products instead of squared Euclidean distances. We carry out a constrained bottom-up merging procedure on a sparsified normalized inner-product matrix. Our method is named SNK-AHC for Sparsified Normalized Kernel matrix based AHC. SNK-AHC is more scalable than the classic dissimilarity matrix based AHC. It can also produce better results when clusters have arbitrary shapes. Artificial and real-world benchmarks are used to exemplify these points. From a theoretical standpoint, SNK-AHC provides another interpretation of the classic techniques which relies on the concept of weighted penalized similarities. The differences between group average, Mcquitty, centroid, median and Ward, can be explained by their distinct averaging strategies for aggregating clusters inter-similarities and intra-similarities. Other features of SNK-AHC are examined. We provide sufficient conditions in order to have monotonic dendrograms, we elaborate a stored data matrix approach for centroid and median, we underline the diagonal translation invariance property of group average, Mcquitty and Ward and we show to what extent SNK-AHC can determine the number of clusters.
Julien Ah-Pine
J. Mach. Learn. Res.1
2017 Fusion Techniques for Named Entity Recognition and Word Sense Induction and Disambiguation
Edmundo-Pavel Soriano-Morales, Julien Ah-Pine, Sabine Loudcher
DS2
2017 SHCoClust, a scalable similarity-based hierarchical co-clustering method and its application to textual collections
abstract
In comparison with flat clustering methods, such as K-means, hierarchical clustering and co-clustering methods are more advantageous, for the reason that hierarchical clustering is capable to reveal the internal connections of clusters, and co-clustering can yield clusters of data instances and features. Interested in organizing co-clusters in hierarchy and in discovering cluster hierarchies inside co-clusters, in this paper, we propose SHCoClust, a scalable similarity-based hierarchical co-clustering method. Except possessing the above-mentioned advantages in unison, SHCoClust is able to employ kernel functions, thanks to its utilization of inner product. Furthermore, having all similarities between 0 and 1, the input of SHCoClust can be sparsified by threshold values, so that less memory and less time are required for storage and for computation. This grants SHCoClust scalability, i.e, the ability to process relatively large datasets with reduced and limited computing resources. Our experiments demonstrate that SHCoClust significantly outperforms the conventional hierarchical clustering methods. In addition, with sparsifying the input similarity matrices obtained by linear kernel and by Gaussian kernel, SHCoClust is capable to guarantee the clustering quality, even when its input being largely sparsified. Consequently, up to 86% time gain and on average 75% memory gain are achieved.
Xinyu Wang 0005, Julien Ah-Pine, Jérôme Darmont
FUZZ-IEEE2
2016 Similarity Based Hierarchical Clustering with an Application to Text Collections
Julien Ah-Pine, Xinyu Wang 0005
IDA1
2016 Hypergraph Modelization of a Syntactically Annotated English Wikipedia Dump
Edmundo-Pavel Soriano-Morales, Julien Ah-Pine, Sabine Loudcher
LREC2
2016 On aggregation functions based on linguistically quantified propositions and finitely additive set functions
Julien Ah-Pine
Fuzzy Sets Syst.1
2015 Unsupervised Visual and Textual Information Fusion in CBMIR Using Graph-Based Methods
abstract
Multimedia collections are more than ever growing in size and diversity. Effective multimedia retrieval systems are thus critical to access these datasets from the end-user perspective and in a scalable way. We are interested in repositories of image/text multimedia objects and we study multimodal information fusion techniques in the context of content-based multimedia information retrieval. We focus on graph-based methods, which have proven to provide state-of-the-art performances. We particularly examine two such methods: cross-media similarities and random-walk-based scores. From a theoretical viewpoint, we propose a unifying graph-based framework, which encompasses the two aforementioned approaches. Our proposal allows us to highlight the core features one should consider when using a graph-based technique for the combination of visual and textual information. We compare cross-media and random-walk-based results using three different real-world datasets. From a practical standpoint, our extended empirical analyses allow us to provide insights and guidelines about the use of graph-based methods for multimodal information fusion in content-based multimedia information retrieval.
Julien Ah-Pine, Gabriela Csurka, Stéphane Clinchant
ACM Trans. Inf. Syst.1
2013 Graph Clustering by Maximizing Statistical Association Measures
Julien Ah-Pine
IDA1
2012 Elicitation of a 2-Additive Bi-capacity through Cardinal Information on Trinary Actions
Brice Mayag, Antoine Rolland, Julien Ah-Pine
IPMU (4)3
2011 Semantic combination of textual and visual information in multimedia retrieval
abstract
The goal of this paper is to introduce a set of techniques we call semantic combination in order to efficiently fuse text and image retrieval systems in the context of multimedia information access. These techniques emerge from the observation that image and textual queries are expressed at different semantic levels and that a single image query is often ambiguous. Overall, the semantic combination techniques overcome a conceptual barrier rather than a technical one: these methods can be seen as a combination of late fusion and image reranking. Albeit simple, this approach has not been used yet. We assess the proposed techniques against late and cross-media fusion using 4 different ImageCLEF datasets. Compared to late fusion, performances significantly increase on two datasets and remain similar on the two other ones.
Stéphane Clinchant, Julien Ah-Pine, Gabriela Csurka
ICMR2
2011 On data fusion in information retrieval using different aggregation operators
abstract
This paper is concerned with the problem of unsupervised rank aggregation in the context of metasearch in information retrieval. In such tasks, we are given many partial ordered lists of retrieved items provided by many search engines and we want to
Julien Ah-Pine
Web Intell. Agent Syst.1
2010 Normalized Kernels as Similarity Indices
Julien Ah-Pine
PAKDD (2)1
2009 Cluster Analysis Based on the Central Tendency Deviation Principle
Julien Ah-Pine
ADMA1
2009 Clique-Based Clustering for Improving Named Entity Recognition Systems
Julien Ah-Pine, Guillaume Jacquet
EACL1
2009 Crossing textual and visual content in different application scenarios
Julien Ah-Pine, Marco Bressan 0003, Stéphane Clinchant, Gabriela Csurka, Yves Hoppenot, Jean-Michel Renders
Multim. Tools Appl.1
2008 Data Fusion in Information Retrieval Using Consensus Aggregation Operators
abstract
In this paper, we address the problem of unsupervised rank aggregation in the context of meta-searching in information retrieval field. The first goal of this paper is to apply aggregation operators that are defined in information fusion domain to the particular issue mentioned beforehand. Triangular norms, conorms and quasi-arithmetic means, are such kind of operators. Then, the second goal of this work is to introduce a new aggregation function, its logical foundations and its combinatorial properties. Particularly, this operator allows to take into account the relationships between experts in a flexible way. Finally, we test these different aggregation operators on the LETOR dataset. The results of our experiments show that this kind of aggregation functions can lead to better results than baseline methods such as CombSUM and CombMNZ approaches.
Julien Ah-Pine
Web Intelligence1