VLDB 2026 Research / reviewers in the wild / expert
Talel Abdessalem
dblp:a/TAbdessalem
· DBLP profile ↗
39ranked-venue papers
2as first author
3since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 27 · 2 first-author · 1 since 2021Artificial intelligence and machine learning · 18 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 9Theory of computation · 2 · 1 since 2021Systems, architecture and hardware · 1Computer networks · 1Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Using A Probabilistic Database in an Image Retrieval ApplicationabstractInternational audience Fajrian Yunus, Pratik Karmakar, Pierre Senellart, Talel Abdessalem, Stéphane Bressan |
EDBT | 4 |
| 2021 | Exact and Approximate Algorithms for Computing Betweenness Centrality in Directed GraphsabstractGraphs (networks) are an important tool to model data in different domains. Real-world graphs are usually directed, where the edges have a direction and they are not symmetric. Betweenness centrality is an important index widely used to analyze networks. In this paper, first given a directed network $G$ and a vertex $r \in V(G)$, we propose an exact algorithm to compute betweenness score of $r$. Our algorithm pre-computes a set $\mathcal{RV}(r)$, which is used to prune a huge amount of computations that do not contribute to the betweenness score of $r$. Time complexity of our algorithm depends on $|\mathcal{RV}(r)|$ and it is respectively $\Theta(|\mathcal{RV}(r)|\cdot|E(G)|)$ and $\Theta(|\mathcal{RV}(r)|\cdot|E(G)|+|\mathcal{RV}(r)|\cdot|V(G)|\log |V(G)|)$ for unweighted graphs and weighted graphs with positive weights. $|\mathcal{RV}(r)|$ is bounded from above by $|V(G)|-1$ and in most cases, it is a small constant. Then, for the cases where $\mathcal{RV}(r)$ is large, we present a simple randomized algorithm that samples from $\mathcal{RV}(r)$ and performs computations for only the sampled elements. We show that this algorithm provides an $(\epsilon,\delta)$-approximation to the betweenness score of $r$. Finally, we perform extensive experiments over several real-world datasets from different domains for several randomly chosen vertices as well as for the vertices with the highest betweenness scores. Our experiments reveal that for estimating betweenness score of a single vertex, our algorithm significantly outperforms the most efficient existing randomized algorithms, in terms of both running time and accuracy. Our experiments also reveal that our algorithm improves the existing algorithms when someone is interested in computing betweenness values of the vertices in a set whose cardinality is very small. Mostafa Haghir Chehreghani, Albert Bifet, Talel Abdessalem |
Fundam. Informaticae | 3 |
| 2021 | River: machine learning for streaming data in PythonabstractRiver is a machine learning library for dynamic data streams and continual learning. It provides multiple state-of-the-art learning methods, data generators/transformers, performance metrics and evaluators for different stream learning problems. It is the result from the merger of two popular packages for stream learning in Python: Creme and scikit-multiflow. River introduces a revamped architecture based on the lessons learnt from the seminal packages. River's ambition is to be the go-to library for doing machine learning on streaming data. Additionally, this open source package brings under the same umbrella a large community of practitioners and researchers. The source code is available at https://github.com/online-ml/river. Jacob Montiel, Max Halford, Saulo Martiello Mastelini, Geoffrey Bolmier, Raphaël Sourty, Robin Vaysse, Adil Zouitine, Heitor Murilo Gomes, Jesse Read, Talel Abdessalem, Albert Bifet |
J. Mach. Learn. Res. | 10 |
| 2020 | Adaptive XGBoost for Evolving Data StreamsabstractBoosting is an ensemble method that combines base models in a sequential manner to achieve high predictive accuracy. A popular learning algorithm based on this ensemble method is eXtreme Gradient Boosting (XGB). We present an adaptation of XGB for classification of evolving data streams. In this setting, new data arrives over time and the relationship between the class and the features may change in the process, thus exhibiting concept drift. The proposed method creates new members of the ensemble from mini-batches of data as new data becomes available. The maximum ensemble size is fixed, but learning does not stop when this size is reached because the ensemble is updated on new data to ensure consistency with the current concept. We also explore the use of concept drift detection to trigger a mechanism to update the ensemble. We test our method on real and synthetic data with concept drift and compare it against batch-incremental and instance-incremental classification methods for data streams. Jacob Montiel, Rory Mitchell, Eibe Frank, Bernhard Pfahringer, Talel Abdessalem, Albert Bifet |
IJCNN | 5 |
| 2020 | Sampling informative patterns from large single networks
Mostafa Haghir Chehreghani, Talel Abdessalem, Albert Bifet, Meriem Bouzbila |
Future Gener. Comput. Syst. | 2 |
| 2019 | Adaptive Algorithms for Estimating Betweenness and k-path CentralitiesabstractBetweenness centrality and k-path centrality are two important indices that are widely used to analyze social, technological and information networks. In the current paper, first given a directed network G and a vertex $r\in V(G)$, we present a novel adaptive algorithm for estimating betweenness score of r. Our algorithm first computes two subsets of the vertex set of G, called $\mathcalRF (r)$ and $\mathcalRT (r)$. They define the sample spaces of the start-points and the end-points of the samples. Then, it adaptively samples from $\mathcalRF (r)$ and $\mathcalRT (r)$ and stops as soon as some condition is satisfied. The stopping condition depends on the samples met so far, $|\mathcalRF (r)|$ and $|\mathcalRT (r)|$. We show that compared to the well-known existing algorithms, our algorithm gives a better $(łambda,δ)$-approximation. Then, we propose a novel algorithm for estimating k-path centrality of r. Our algorithm is based on computing two sets $\mathcalRF (r)$ and $\mathcalD (r)$. While $\mathcalRF (r)$ defines the sample space of the source vertices of the sampled paths, $\mathcalD (r)$ defines the sample space of the other vertices of the paths. We show that in order to give a $(łambda,δ)$-approximation of the k-path score of r, our algorithm requires considerably less samples. Moreover, it processes each sample faster and with less memory. Finally, we empirically evaluate our proposed algorithms and show their superior performance. Also, we show that they can be used to efficiently compute centrality scores of a set of vertices. Mostafa Haghir Chehreghani, Albert Bifet, Talel Abdessalem |
CIKM | 3 |
| 2019 | Metropolis-Hastings Algorithms for Estimating Betweenness Centrality
Mostafa Haghir Chehreghani, Talel Abdessalem, Albert Bifet |
EDBT | 2 |
| 2019 | Correction to: Adaptive random forests for evolving data stream classification
Heitor Murilo Gomes, Albert Bifet, Jesse Read, Jean Paul Barddal, Fabrício Enembreck, Bernhard Pfahringer, Geoff Holmes 0001, Talel Abdessalem |
Mach. Learn. | 8 |
| 2018 | ALGeoSPF: A Hierarchical Factorization Model for POI RecommendationabstractThe task of points-of-interest (POI) recommendations has become an essential feature in location-based social networks (LBSNs) with the significant growth of shared data on LBSNs. However it remains a challenging problem, because the decision process of a user choosing to visit a POI depends on numerous factors. The high level of sparsity of the data in LBSNs makes the POI recommendation problem even more challenging, especially for large geographical areas and worldwide datasets. Moreover, in this context the mobility behavior of the users is very heterogeneous, ranging from urban to worldwide mobility. In this paper, we explore the impact of spatial clustering on the recommendation quality. The proposed approach combines spatial clustering with users' influences. It is based on a Poisson factorization model built on an implicit social network, inferred from the geographical mobility patterns. We conduct a comprehensive performance evaluation of our approach on the YFCC dataset (a very large-scale real-world dataset). The experiments show that our approach achieves a significantly superior recommendation quality compared to other state-of-the-art recommendation techniques. Jean-Benoît Griesner, Talel Abdessalem, Hubert Naacke, Pierre Dosne |
ASONAM | 2 |
| 2018 | An In-depth Comparison of Group Betweenness Centrality Estimation AlgorithmsabstractOne of the important indices defined for a set of vertices in a graph is group betweenness centrality. While in recent years a number of approximate algorithms have been proposed to estimate this index, there is no comprehensive and in-depth analysis and comparison of these algorithms in the literature. In this paper, we first present a generic algorithm that is used to express different approximate algorithms in terms of probability distributions. Using this generic algorithm, we show that interestingly existing methods have the same theoretical accuracy. Then, we present an extension of distance-based sampling to group betweenness centrality, which is based on a new notion of distance between a single vertex and a set of vertices. In the end, to empirically evaluate efficiency and accuracy of different algorithms, we perform experiments over several real-world networks. Our extensive experiments reveal that those approximate algorithms that are based on shortest path sampling are orders of magnitude faster than those algorithms that are based on pair sampling, while these two types of algorithms have almost comparable empirical accuracy. Mostafa Haghir Chehreghani, Albert Bifet, Talel Abdessalem |
IEEE BigData | 3 |
| 2018 | DyBED: An Efficient Algorithm for Updating Betweenness Centrality in Directed Dynamic GraphsabstractAn important index widely used to analyze social and information networks is betweenness centrality. In this paper, given a dynamic and directed graph G and a vertex r in G, we present the DyBED algorithm that updates the (approximate) betweenness centrality of r, when an update operation (vertex/edge insertion/deletion) occurs in G. Our algorithm first during pre-processing computes two subsets of the vertex set of G, called ΠT(r) and ΠT(r). The Cartesian product of these two sets defines the sample space of our algorithm. In other words, each sample is a pair, whose first element belongs to ΠT(r) and second element belongs to ΠT(r). Then after each update operation, DyBED updates the sets ΠT(r) and ΠT(r), the sampled pairs, the information stored for each sample and accordingly, the betweenness centrality of r. We theoretically and empirically evaluate DyBED and show that it yields significant improvement over existing work. In particular, our extensive experiments reveal that DyBED is orders of magnitude faster than most efficient existing algorithms. Mostafa Haghir Chehreghani, Albert Bifet, Talel Abdessalem |
IEEE BigData | 3 |
| 2018 | Learning Fast and Slow: A Unified Batch/Stream FrameworkabstractData ubiquity highlights the need of efficient and adaptable data-driven solutions. In this paper, we present FAST AND SLOW LEARNING (FSL), a novel unified framework that sheds light on the symbiosis between batch and stream learning. FSL works by employing Fast (stream) and Slow (batch) Learners, emulating the mechanisms used by humans to make decisions. We showcase the applicability of FSL on the task of classification by introducing the FAST AND SLOW CLASSIFIER (FSC). A Fast Learner provides predictions on the spot, continuously updating its model and adapting to changes in the data. On the other hand, the Slow Learner provides predictions considering a wider spectrum of seen data, requiring more time and data to create complex models. Once that enough data has been collected, FSC trains the Slow Learner and starts tracking the performance of both learners. A drift detection mechanism triggers the creation of new Slow models when the current Slow model becomes obsolete. FSC selects between Fast and Slow Learners according to their performance on new incoming data. Test results on real and synthetic data show that FSC effectively drives the positive interaction of stream and batch models for learning from evolving data streams. Jacob Montiel, Albert Bifet, Viktor Losing, Jesse Read, Talel Abdessalem |
IEEE BigData | 5 |
| 2018 | Harnessing Truth Discovery Algorithms On The Topic Labelling ProblemabstractTopics in topic modelling approaches are represented as a collection of weighted words. The labels for the topics, however, are not clearly defined and must be interpreted manually. Topic labelling proposes to automatically label the topics by leveraging a knowledge base or applying data mining and machine learning algorithms. We propose a naive topic labelling approach where we transform the labeling problem into selecting the best label for each word in the topic. The candidate labels are generated by querying a knowledge base using the top-N words of each topic. We construct a heterogeneous graph of topics, words, articles and candidate labels. To rank the candidate labels, we apply truth discovery algorithms on the graph. The performance evaluation using popular topic modelling datasets shows that the approach receives satisfactory accuracy. Ngurah Agus Sanjaya Er, Mouhamadou Lamine Ba, Talel Abdessalem, Stéphane Bressan |
iiWAS | 3 |
| 2018 | Set Labelling using Multi-label ClassificationabstractWe propose the task of set labelling. Starting from some examples members of a set, set labelling tries to infer the most appropriate labels for the given set. For this work, we consider sets of words. We illustrate the task and a possible solution with an application to the classification of cosmetic products and hotels. The novel solution proposed in this research is to incorporate a multi-label classifier trained from the labeled datasets. We use vectorization of the description of the seeds as input to the classifier as well as labels assigned to it. Given a previously unseen data, the trained classifier returns a ranked list of candidate labels (i.e., additional seeds) for the set. These results could then be used to infer the labels for the set. We implement our proposed solution to the classification of cosmetic products and hotels. We show that the solution is effective and efficient. Ngurah Agus Sanjaya Er, Jesse Read, Talel Abdessalem, Stéphane Bressan |
iiWAS | 3 |
| 2018 | Adaptive Window Strategy for Topic Modeling in Document StreamsabstractExtracting global themes from a written text has recently become a major issue for computational intelligence, in particular in Natural Language Processing communities. Among all proposed solutions, Latent Dirichlet Allocation (LDA) has gained a vast interest and several variants have been proposed to adapt to changing environments. With the emergence of data streams, for instance from social media, the domain faces a new challenge: topic extraction in real time. In this paper, we propose a simple approach called Adaptive Window based Incremental LDA (AWILDA) originating from the cross-over between LDA and state-of-the-art methods in data stream mining. We train new topic models only when a drift is detected and select training data on the fly using ADWIN algorithm. We provide both theoretical guarantees for our method and experimental validation on artificial and real-world data. Pierre-Alexandre Murena, Marie Al-Ghossein, Talel Abdessalem, Antoine Cornuéjols |
IJCNN | 3 |
| 2018 | Efficient Exact and Approximate Algorithms for Computing Betweenness Centrality in Directed Graphs
Mostafa Haghir Chehreghani, Albert Bifet, Talel Abdessalem |
PAKDD (3) | 3 |
| 2018 | Scalable Model-Based Cascaded Imputation of Missing Data
Jacob Montiel, Jesse Read, Albert Bifet, Talel Abdessalem |
PAKDD (3) | 4 |
| 2018 | Adaptive collaborative topic modeling for online recommendationabstractCollaborative filtering (CF) mainly suffers from rating sparsity and from the cold-start problem. Auxiliary information like texts and images has been leveraged to alleviate these problems, resulting in hybrid recommender systems (RS). Due to the abundance of data continuously generated in real-world applications, it has become essential to design online RS that are able to handle user feedback and the availability of new items in real-time. These systems are also required to adapt to drifts when a change in the data distribution is detected. In this paper, we propose an adaptive collaborative topic modeling approach, CoAWILDA, as a hybrid system relying on adaptive online Latent Dirichlet Allocation (AWILDA) to model newly available items arriving as a document stream and incremental matrix factorization for CF. The topic model is maintained up-to-date in an online fashion and is retrained in batch when a drift is detected using documents automatically selected by an adaptive windowing technique. Our experiments on real-world datasets prove the effectiveness of our approach for online recommendation. Marie Al-Ghossein, Pierre-Alexandre Murena, Talel Abdessalem, Anthony Barré, Antoine Cornuéjols |
RecSys | 3 |
| 2018 | Discriminative Distance-Based Network Indices with Application to Link PredictionabstractIn distance-based network indices, the distance between two vertices is measured by the length of shortest paths between them. A shortcoming of this measure is that when it is used in real-world networks, a huge number of vertices may have exactly the same closeness/eccentricity scores. This restricts the applicability of these indices as they cannot distinguish vertices. Furthermore, in many applications, the distance between two vertices not only depends on the length of shortest paths but also on the number of shortest paths between them. In this paper, first we develop a new distance measure, proportional to the length of shortest paths and inversely proportional to the number of shortest paths, that yields discriminative distance-based centrality indices. We present exact and randomized algorithms for computation of the proposed discriminative indices. Then, by performing extensive experiments, we first show that compared with the traditional indices, discriminative indices have usually much more discriminability. Then, we show that our randomized algorithms can very precisely estimate average discriminative path length and average discriminative eccentricity, using only few samples. Then, we show that real-world networks have usually a tiny average discriminative path length, bounded by a constant (e.g. 2). We refer to this property as the tiny-world property. Finally, we present a novel link prediction method that uses discriminative distance to decide which vertices are more likely to form a link in future, and show its superior performance. Mostafa Haghir Chehreghani, Albert Bifet, Talel Abdessalem |
Comput. J. | 3 |
| 2018 | Scikit-Multiflow: A Multi-output Streaming Frameworkabstractscikit-multiflow is a framework for learning from data streams and multi-output learning in Python. Conceived to serve as a platform to encourage the democratization of stream learning research, it provides multiple state-of-the-art learning methods, data generators and evaluators for different stream learning problems, including single-output, multi-output and multi-label. scikit-multiflow builds upon popular open source frameworks including scikit-learn, MOA and MEKA. Development follows the FOSS principles. Quality is enforced by complying with PEP8 guidelines, using continuous integration and functional testing. Jacob Montiel, Jesse Read, Albert Bifet, Talel Abdessalem |
J. Mach. Learn. Res. | 4 |
| 2017 | Predicting over-indebtedness on batch and streaming dataabstractDetecting over-indebtedness, the difficulties meeting household payment commitments, poses multiple Big Data challenges for banking institutions. We present a novel data-driven framework for predicting over-indebtedness on realworld data. A warning mechanism that generates predictions 6 months ahead, improving the chances of financial recovery. This framework is based on the combination of feature selection and supervised learning techniques, and uses data balancing for fine-tuning the predictive models. We propose two versions of the framework based on state-of-the-art batch and streaming learning techniques. To the best of our knowledge, the proposed framework is the first to cast over-indebtedness prediction as a stream learning problem. The appeal of stream learning rises from the large amount of data continuously generated, and the fact that batch models become obsolete over time as financial data evolves, while stream models are continuously updated as new data is available. We use credit data from two banks from the Groupe BPCE (the second-largest banking institution in France) and apply multi-metric criteria to evaluate model performance and fairness. Test results show the framework's interbank applicability and that the proposed batch and stream frameworks outperform the current solution for both single and multi-metric criteria. Additionally, the generic structure of the framework serves as a template for systematically approaching similar classification problems. Jacob Montiel, Albert Bifet, Talel Abdessalem |
IEEE BigData | 3 |
| 2017 | Truthfulness of Candidates in Set of t-uples Expansion
Ngurah Agus Sanjaya Er, Mouhamadou Lamine Ba, Talel Abdessalem, Stéphane Bressan |
DEXA (1) | 3 |
| 2017 | How to Find the Best Rated Items on a Likert Scale and How Many Ratings Are Enough
Qing Liu 0020, Debabrota Basu, Shruti Goel, Talel Abdessalem, Stéphane Bressan |
DEXA (2) | 4 |
| 2017 | Preserving privacy in distributed system (PPDS) protocol: Security analysisabstractWithin the diversity of existing Big Data and data processing solutions, meeting the requirements of privacy and security is becoming a real need. In this paper we tackle the security analysis of a new protocol of data processing in distributed system (PPDS). This protocol is composed of three phases: authentication, node head selection and data linking. This paper deals with its formal validation done using HLPSL language via AVISPA. We provide also its security analysis. Some performance analysis based on its proof of concept are also given in this paper. Ashref Aloui, Mounira Msahli, Talel Abdessalem, Stéphane Bressan, Sihem Mesnager |
IPCCC | 3 |
| 2017 | Protocol for preserving privacy in distributed system (PPDS)abstractPreserving privacy in Big data is one of the most debated topics of computer security. The fast growing volume of data and the need of companies to extract value from that data creates new complicated and serious security challenges. Despite its wide spread, the common use and the popularity of Big data paradigm, significant risks and challenges are inherent to this new concept, especially when we talk about externalized treatment of sensitive data via insecure network. In this paper we tackle the privacy challenge in Big Data. We focus in special case of data processing. Several Bank agencies want to share data processing while protecting the privacy. We propose a new protocol of communication between agencies. This protocol is composed of three phases: authentication, node head selection and data linking. Some proof of concept are also given in this paper. Ashref Aloui, Mounira Msahli, Talel Abdessalem, Stéphane Bressan, Sihem Mesnager |
IWCMC | 3 |
| 2017 | Upper and lower bounds for the q-entropy of network models with application to network model selection
Mostafa Haghir Chehreghani, Talel Abdessalem |
Inf. Process. Lett. | 2 |
| 2017 | Adaptive random forests for evolving data stream classification
Heitor Murilo Gomes, Albert Bifet, Jesse Read, Jean Paul Barddal, Fabrício Enembreck, Bernhard Pfahringer, Geoff Holmes 0001, Talel Abdessalem |
Mach. Learn. | 8 |
| 2016 | Echo State Hoeffding Tree LearningabstractNowadays, real-time classification of Big Data streams is becoming essential in a variety of application domains. While decision trees are powerful and easy-to-deploy approaches for accurate and fast learning from data streams, they are unable to capture the strong temporal dependences typically present in the input data. Recurrent Neural Networks are an alternative solution that include an internal memory to capture these temporal dependences; however their training is computationally very expensive and with slow convergence, requiring a large number of hyper-parameters to tune. Reservoir Computing was proposed to reduce the computation requirements of the training phase but still include a feed-forward layer which requires a large number of parameters to tune. In this work we propose a novel architecture for real-time classification based on the combination of a Reservoir and a decision tree. This combination reduces the number of hyper-parameters while still maintaining the good temporal properties of recurrent neural networks. The capabilities of the proposed architecture to learn some typical string-based functions with strong temporal dependences are evaluated in the paper. We show how the new architecture is able to incrementally learn these functions in real-time with fast adaptation to unknown sequences. And we study the influence of the reduced number of hyper-parameters in the behaviour of the proposed solution. Diego Marron, Jesse Read, Albert Bifet, Talel Abdessalem, Eduard Ayguadé, José R. Herrero 0001 |
ACML | 4 |
| 2016 | Cost Minimization and Social Fairness for Spatial Crowdsourcing Tasks
Qing Liu 0020, Talel Abdessalem, Huayu Wu 0001, Zihong Yuan, Stéphane Bressan |
DASFAA (1) | 2 |
| 2016 | Set of t-uples expansion by exampleabstractSet expansion is the task of finding elements of a set given example members. We are interested in the design of algorithms and techniques for a set expansion tool that expands a set by searching, finding and extracting candidates from the World Wide Web. Existing approaches mostly consider sets of atomic data. We extend this idea to the expansion of sets of t-uples, that is relation instances or tables. We propose an approach for extracting relation instances from the World Wide Web given a handful set of t-uple seeds. For instance, when the user proposes the set of seeds , , the system returns a relation containing currency codes with their corresponding country and capital city. We show how a random walk in a heterogeneous graph of Web pages, wrappers, seeds and candidates is able to rank the candidates according to their relevance to the seeds. We evaluate the performance of the approach and show that it is efficient, effective and practical. Ngurah Agus Sanjaya Er, Talel Abdessalem, Stéphane Bressan |
iiWAS | 2 |
| 2015 | POI Recommendation: Towards Fused Matrix Factorization with Geographical and Temporal Influences
Jean-Benoît Griesner, Talel Abdessalem, Hubert Naacke |
RecSys | 2 |
| 2014 | Monitoring moving objects using uncertain web dataabstractA number of applications deal with monitoring moving objects: cars, aircrafts, ships, persons, etc. Traditionally, this requires capturing data from sensor networks, image or video analysis, or using other application-specific resources. We show in this demonstration paper how Web content can be exploited instead to gather information (trajectories, metadata) about moving objects. As this content is marred with uncertainty and inconsistency, we develop a methodology for estimating uncertainty and filtering the resulting data. We present as an application a demonstration of a system that constructs trajectories of ships from social networking data, presenting to a user inferred trajectories, meta-information, as well as uncertainty levels on extracted information and trustworthiness of data providers. Mouhamadou Lamine Ba, Sébastien Montenez, Talel Abdessalem, Pierre Senellart |
SIGSPATIAL/GIS | 3 |
| 2014 | A parameter-free algorithm for an optimized tag recommendation list sizeabstractTag recommendation is a major aspect of collaborative tagging systems. It aims to recommend suitable tags to a user for tagging an item. One of its main challenges is the effectiveness of its recommendations. Existing works focus on techniques for retrieving the most relevant tags to give beforehand, with a fixed number of tags in each recommended list. In this paper, we try to optimize the number of recommended tags in order to improve the efficiency of the recommendations. We propose a parameter-free algorithm for determining the optimal size of the recommended list. Thus we introduced some relevance measures to find the most relevant sublist from a given list of recommended tags. More precisely, we improve the quality of our recommendations by discarding some unsuitable tags and thus adjusting the list size. Modou Gueye, Talel Abdessalem, Hubert Naacke |
RecSys | 2 |
| 2013 | Uncertain version control in open collaborative editing of tree-structured documentsabstractIn order to ease content enrichment, exchange, and sharing, web-scale collaborative platforms such as Wikipedia or Google Docs enable unbounded interactions between a large number of contributors, without prior knowledge of their level of expertise and reliability. Version control is then essential for keeping track of the evolution of the shared content and its provenance. In such environments, uncertainty is ubiquitous due to the unreliability of the sources, the incompleteness and imprecision of the contributions, the possibility of malicious editing and vandalism acts, etc. To handle this uncertainty, we use a probabilistic XML model as a basic component of our version control framework. Each version of a shared document is represented by an XML tree and the whole document, together with its different versions, is modeled as a probabilistic XML document. Uncertainty is evaluated using the probabilistic model and the reliability measure associated to each source, each contributor, or each editing event, resulting in an uncertainty measure on each version and each part of the document. We show that standard version control operations can be implemented directly as operations on the probabilistic XML model; efficiency with respect to deterministic version control systems is demonstrated on real-world datasets. Mouhamadou Lamine Ba, Talel Abdessalem, Pierre Senellart |
ACM Symposium on Document Engineering | 2 |
| 2012 | Primates: a privacy management system for social networksabstractWhile online social networks (OSN) present unprecedented opportunities for sharing information and multimedia content among users, they raise major privacy issues as users could often access personal or confidential data of other users. Most social networks provide some basic access control policies, which however seem to be very limited given the diversity of user relationships in the current social networks (e.g. friend, acquaintance, son) as well as the needs of social network users who might want to express sophisticated access control policies (e.g. invite all children of my colleagues to my child's birthday party). In this demonstration proposal, we present Primates a privacy management system for social networks. Primates allows users to specify access control rules for their resources and enforces access control over all shared resources. The set of users who are allowed to access a given resource is defined by a set of constraints on the paths connecting the owner of a resource to its requester in the social graph. We demonstrate the accuracy of our access control model and the scalability of our system. Imen Ben Dhia, Talel Abdessalem, Mauro Sozio |
CIKM | 2 |
| 2012 | Automatic Extraction of Structured Web Data with Domain KnowledgeabstractWe present in this paper a novel approach for extracting structured data from the Web, whose goal is to harvest real-world items from template-based HTML pages (the structured Web). It illustrates a two-phase querying of the Web, in which an intentional description of the data that is targeted is first provided, in a flexible and widely applicable manner. The extraction process leverages then both the input description and the source structure. Our approach is domain-independent, in the sense that it applies to any relation, either flat or nested, describing real-world items. Extensive experiments on five different domains and comparison with the main state of the art extraction systems from literature illustrate its flexibility and precision. We advocate via our technique that automatic extraction and integration of complex structured data can be done fast and effectively, when the redundancy of the Web meets knowledge over the to-be-extracted data. Nora Derouiche, Bogdan Cautis, Talel Abdessalem |
ICDE | 3 |
| 2011 | A probabilistic XML merging toolabstractThis demonstration paper presents a probabilistic XML data merging tool, that represents the outcome of semi-structured document integration as a probabilistic tree. The system is fully automated and integrates methods to evaluate the uncertainty (modeled as probability values) of the result of the merge. It is based on the two-way tree-merge technique and an uncertain data model defined using probabilistic event variables. The resulting probabilistic repository can be queried using a subset of the XPath query language. The demonstration application is based on revisions of the Wikipedia encyclopedia: a Wikipedia article is no longer considered as the latest valid revision but as the merge of all possible revisions, some of which are uncertain. Talel Abdessalem, Mouhamadou Lamine Ba, Pierre Senellart |
EDBT | 1 |
| 2010 | ObjectRunner: Lightweight, Targeted Extraction and Querying of Structured Web DataabstractWe present in this paper ObjectRunner, a system for extracting, integrating and querying structured data from the Web. Our system harvests real-world items from template-based HTML pages (the so-called structured Web). It illustrates a two-phase querying of the Web, in which an intentional description of the targeted data is first provided, in a flexible and widely applicable manner. ObjectRunner follows then a lightweight, best-effort approach, leveraging both the input description and the source structure. This process is domain-independent, in the sense that it applies to any relation, either flat or nested, describing real-world items. We advocate via our prototype that fully automatic extraction and integration of structured data can be done fast and effectively, when the redundancy of the Web meets knowledge over the to-be-extracted data. We present the technical details and the overall platform through several application scenarios on real-life Web sources. Talel Abdessalem, Bogdan Cautis, Nora Derouiche |
Proc. VLDB Endow. | 1 |
| 2008 | Pruning nested XQuery queriesabstractWe present in this paper an approach for XQuery optimization that exploits minimization opportunities raised in composition-style nesting of queries. More precisely, we consider the simplification of XQuery queries in which the intermediate result constructed by a subexpression is queried by another subexpression. Based on a large subset of XQuery, we describe a rule-based algorithm that recursively prunes query expressions, eliminating useless intermediate results. Our algorithm takes as input an XQuery expression that may have navigation within its subexpressions and outputs a simplified, equivalent XQuery expression, and is thus readily usable as an optimization module in any existing XQuery processor. We demonstrate by experiments the impact of our rewriting approach on query evaluation costs and we prove formally its correctness. Bilel Gueni, Talel Abdessalem, Bogdan Cautis, Emmanuel Waller |
CIKM | 2 |