Stéphane Bressan

dblp:b/SBressan · DBLP profile ↗
← Back
105ranked-venue papers in the field
12as first author
15since 2021 · last 2025
0000-0001-5536-3296ORCID · verified

Domains — venue-derived; a paper can count in several

Database Systems & Data Management · 76 (6 first)Information Retrieval & Web Search · 26 (6 first)Data Mining & Knowledge Discovery · 2Business Process & Enterprise Data · 1
YearPublicationVenuePosition
2025 Physics-Informed Discovery of State Variables in Second-Order and Hamiltonian Systems
Félix Chavelli, Zi-Yu Khoo, Dawen Wu, Jonathan Sze Choong Low, Stéphane Bressan
ACIIDS (1)5
2025 Discovering Voting Power for Ensemble Methods
Pratik Karmakar, Angelo Saadeh, Pierre Senellart, Stéphane Bressan
DEXA (1)4
2025 Using A Probabilistic Database in an Image Retrieval Application
abstract
International audience
Fajrian Yunus, Pratik Karmakar, Pierre Senellart, Talel Abdessalem, Stéphane Bressan
EDBT5
2024 Expected Shapley-Like Scores of Boolean functions: Complexity and Applications to Probabilistic Databases
abstract
Shapley values, originating in game theory and increasingly prominent in explainable AI, have been proposed to assess the contribution of facts in query answering over databases, along with other similar power indices such as Banzhaf values. In this work we adapt these Shapley-like scores to probabilistic settings, the objective being to compute their expected value. We show that the computations of expected Shapley values and of the expected values of Boolean functions are interreducible in polynomial time, thus obtaining the same tractability landscape. We investigate the specific tractable case where Boolean functions are represented as deterministic decomposable circuits, designing a polynomial-time algorithm for this setting. We present applications to probabilistic databases through database provenance, and an effective implementation of this algorithm within the ProvSQL system, which experimentally validates its feasibility over a standard benchmark.
Pratik Karmakar, Mikaël Monet, Pierre Senellart, Stéphane Bressan
Proc. ACM Manag. Data4
2023 Assessing the Effectiveness of Intrinsic Dimension Estimators for Uncovering the Phase Space Dimensionality of Dynamical Systems from State Observations - A Comparative Analysis
Félix Chavelli, Zi-Yu Khoo, Jonathan Sze Choong Low, Stéphane Bressan
DEXA (1)4
2023 Celestial Machine Learning - From Data to Mars and Beyond with AI Feynman
Zi-Yu Khoo, Abel Yang, Jonathan Sze Choong Low, Stéphane Bressan
DEXA (2)4
2023 Confidential Truth Finding with Multi-Party Computation
Angelo Saadeh, Pierre Senellart, Stéphane Bressan
DEXA (1)3
2023 A Comparative Evaluation of Additive Separability Tests for Physics-Informed Machine Learning
Zi-Yu Khoo, Jonathan Sze Choong Low, Stéphane Bressan
iiWAS3
2023 Celestial Machine Learning - Discovering the Planarity, Heliocentricity, and Orbital Equation of Mars with AI Feynman
Zi-Yu Khoo, Gokul Rajiv, Abel Yang, Jonathan Sze Choong Low, Stéphane Bressan
iiWAS5
2022 What's Next? Predicting Hamiltonian Dynamics from Discrete Observations of a Vector Field
Zi-Yu Khoo, Delong Zhang, Stéphane Bressan
DEXA (2)3
2022 Latent Relational Point Process: Network Reconstruction from Discrete Event Data
Guilherme Augusto Zagatti, See-Kiong Ng, Stéphane Bressan
DEXA (2)3
2022 Syntax-Informed Question Answering with Heterogeneous Graph Transformer
Fangyi Zhu, Lok You Tan, See-Kiong Ng, Stéphane Bressan
DEXA (1)4
2021 Neural Ordinary Differential Equations for the Regression of Macroeconomics Data Under the Green Solow Model
Zi-Yu Khoo, Kang Hao Lee, Zhibo Huang, Stéphane Bressan
DEXA (1)4
2021 Transfer Learning for Larger, Broader, and Deeper Neural-Network Quantum States
Remmy A. M. Zen, Stéphane Bressan
DEXA (2)2
2021 A Large-scale Disease Outbreak Analytics System based on Wi-Fi Session Logs
abstract
Unraveling human mobility patterns is critical for understanding disease spread and implementing effective controls during large-scale disease outbreaks such as the COVID-19 pandemic. Given the urgency associated with such situations, it is important to leverage on the common existing digital infrastructures that can be readily activated for disease outbreak analytics. We introduce an integrated system for disease outbreak investigation using data from Wi-Fi sessions. The system offers outbreak analytics, simulation, and visualization capabilities to assist in the identification of infection hot-spots and in contacttracing exercises. The system has been developed and experimentally deployed for research purposes on a large local university campus in Singapore.
Guilherme Augusto Zagatti, Tingfeng Wu, See-Kiong Ng, Stéphane Bressan
MDM4
2020 Construction and Random Generation of Hypergraphs with Prescribed Degree and Dimension Sequences
Naheed Anjum Arafat, Debabrota Basu, Laurent Decreusefond, Stéphane Bressan
DEXA (2)4
2020 Hydrological Process Surrogate Modelling and Simulation with Neural Networks
Ruixi Zhang, Remmy A. M. Zen, Jifang Xing, Dewa Made Sri Arsa, Abhishek Saha, Stéphane Bressan
PAKDD (2)6
2019 Topological Data Analysis with \epsilon ϵ -net Induced Lazy Witness Complex
Naheed Anjum Arafat, Debabrota Basu, Stéphane Bressan
DEXA (2)3
2019 Differentially Private Non-parametric Machine Learning as a Service
Ashish Dandekar, Debabrota Basu, Stéphane Bressan
DEXA (1)3
2019 Rainfall Estimation from Traffic Cameras
Remmy A. M. Zen, Dewa Made Sri Arsa, Ruixi Zhang, Ngurah Agus Sanjaya Er, Stéphane Bressan
DEXA (1)5
2019 Building Extraction from Google Earth Images
abstract
Building extraction is a component of many environmental modelling and data analysis applications. It is however data and knowledge intensive. We investigate the use of publicly available data from Google Earth and OpenStreetMap and of neural networks for this task. We evaluate different candidate algorithms for the case of building extraction on the island of Bali.
Jifang Xing, Ruixi Zhang, Remmy A. M. Zen, Dewa Made Sri Arsa, Ismail Khalil, Stéphane Bressan
iiWAS6
2019 Microbiological Water Quality Test Results Extraction from Mobile Photographs
abstract
An emerging and promising approach to achieving access to water and sanitation for all leverages citizen science to collect valuable data on water quantity and quality, which can assist policymakers and water utility managers in sustainably managing water resources.
Jifang Xing, Ruixi Zhang, Remmy A. M. Zen, Ngurah Agus Sanjaya Er, Laure Sioné, Ismail Khalil, Stéphane Bressan
iiWAS7
2019 BelMan: An Information-Geometric Approach to Stochastic Bandits
Debabrota Basu, Pierre Senellart, Stéphane Bressan
ECML/PKDD (3)3
2018 Differential Privacy for Regularised Linear Regression
Ashish Dandekar, Debabrota Basu, Stéphane Bressan
DEXA (2)3
2018 A Comparative Study of Synthetic Dataset Generation Techniques
Ashish Dandekar, Remmy A. M. Zen, Stéphane Bressan
DEXA (2)3
2018 Harnessing Truth Discovery Algorithms On The Topic Labelling Problem
abstract
Topics in topic modelling approaches are represented as a collection of weighted words. The labels for the topics, however, are not clearly defined and must be interpreted manually. Topic labelling proposes to automatically label the topics by leveraging a knowledge base or applying data mining and machine learning algorithms. We propose a naive topic labelling approach where we transform the labeling problem into selecting the best label for each word in the topic. The candidate labels are generated by querying a knowledge base using the top-N words of each topic. We construct a heterogeneous graph of topics, words, articles and candidate labels. To rank the candidate labels, we apply truth discovery algorithms on the graph. The performance evaluation using popular topic modelling datasets shows that the approach receives satisfactory accuracy.
Ngurah Agus Sanjaya Er, Mouhamadou Lamine Ba, Talel Abdessalem, Stéphane Bressan
iiWAS4
2018 Set Labelling using Multi-label Classification
abstract
We propose the task of set labelling. Starting from some examples members of a set, set labelling tries to infer the most appropriate labels for the given set. For this work, we consider sets of words. We illustrate the task and a possible solution with an application to the classification of cosmetic products and hotels. The novel solution proposed in this research is to incorporate a multi-label classifier trained from the labeled datasets. We use vectorization of the description of the seeds as input to the classifier as well as labels assigned to it. Given a previously unseen data, the trained classifier returns a ranked list of candidate labels (i.e., additional seeds) for the set. These results could then be used to infer the labels for the set. We implement our proposed solution to the classification of cosmetic products and hotels. We show that the solution is effective and efficient.
Ngurah Agus Sanjaya Er, Jesse Read, Talel Abdessalem, Stéphane Bressan
iiWAS4
2017 Hypergraph Drawing by Force-Directed Placement
Naheed Anjum Arafat, Stéphane Bressan
DEXA (2)2
2017 Generating Fake but Realistic Headlines Using Deep Neural Networks
Ashish Dandekar, Remmy A. M. Zen, Stéphane Bressan
DEXA (2)3
2017 Truthfulness of Candidates in Set of t-uples Expansion
Ngurah Agus Sanjaya Er, Mouhamadou Lamine Ba, Talel Abdessalem, Stéphane Bressan
DEXA (1)4
2017 How to Find the Best Rated Items on a Likert Scale and How Many Ratings Are Enough
Qing Liu 0020, Debabrota Basu, Shruti Goel, Talel Abdessalem, Stéphane Bressan
DEXA (2)5
2017 Data Driven Generation of Synthetic Data with Support Vector Data Description
Fajrian Yunus, Ashish Dandekar, Stéphane Bressan
DEXA (2)3
2016 Bus Routes Design and Optimization via Taxi Data Analytics
abstract
Public bus services are often planned in the context of urban planning. For a city with efficient and extensive network of public transportation system like Singapore, enhancing the existing coverage of bus service to meet the dynamic mobility needs of the population requires data mining approach. Specifically, frequent taxi rides between two locations at a period of time may suggest possible poor coverage of public transport service, if not lacking of the public transport service. In this paper, we describe a proof of concept effort to discover this weakness and its improvement in public transportation system via mining of taxi ride dataset. We cluster taxi rides dataset to determine some popular taxi rides in Singapore. From the clustered taxi rides, we filter and select only the clusters whose commuting via existing public transport are tortuous if not unreachable door-to-door. Based on the discovered travel pattern, we propose new bus routes that serve the passengers of these clusters. We formulate the bus planning problem as an optimization of directed cycle graph, and present it's preliminary solution and results. We showcase our idea in the case of Singapore.
Seong-Ping Chuah, Huayu Wu 0001, Yu Lu 0003, Liang Yu 0005, Stéphane Bressan
CIKM5
2016 Routing an Autonomous Taxi with Reinforcement Learning
abstract
Singapore's vision of a Smart Nation encompasses the development of effective and efficient means of transportation. The government's target is to leverage new technologies to create services for a demand-driven intelligent transportation model including personal vehicles, public transport, and taxis. Singapore's government is strongly encouraging and supporting research and development of technologies for autonomous vehicles in general and autonomous taxis in particular. The design and implementation of intelligent routing algorithms is one of the keys to the deployment of autonomous taxis. In this paper we demonstrate that a reinforcement learning algorithm of the Q-learning family, based on a customized exploration and exploitation strategy, is able to learn optimal actions for the routing autonomous taxis in a real scenario at the scale of the city of Singapore with pick-up and drop-off events for a fleet of one thousand taxis.
Miyoung Han, Pierre Senellart, Stéphane Bressan, Huayu Wu 0001
CIKM3
2016 Cost Minimization and Social Fairness for Spatial Crowdsourcing Tasks
Qing Liu 0020, Talel Abdessalem, Huayu Wu 0001, Zihong Yuan, Stéphane Bressan
DASFAA (1)5
2016 Set of t-uples expansion by example
abstract
Set expansion is the task of finding elements of a set given example members. We are interested in the design of algorithms and techniques for a set expansion tool that expands a set by searching, finding and extracting candidates from the World Wide Web. Existing approaches mostly consider sets of atomic data. We extend this idea to the expansion of sets of t-uples, that is relation instances or tables. We propose an approach for extracting relation instances from the World Wide Web given a handful set of t-uple seeds. For instance, when the user proposes the set of seeds , , the system returns a relation containing currency codes with their corresponding country and capital city. We show how a random walk in a heterogeneous graph of Web pages, wrappers, seeds and candidates is able to rank the candidates according to their relevance to the seeds. We evaluate the performance of the approach and show that it is efficient, effective and practical.
Ngurah Agus Sanjaya Er, Talel Abdessalem, Stéphane Bressan
iiWAS3
2015 Cost-Model Oblivious Database Tuning with Reinforcement Learning
Debabrota Basu, Qian Lin 0002, Weidong Chen 0004, Hoang Tam Vo, Zihong Yuan, Pierre Senellart, Stéphane Bressan
DEXA (1)7
2015 Expressing and Processing Path-Centric XML Queries
Huayu Wu 0001, Dongxu Shao, Ruiming Tang, Tok Wang Ling, Stéphane Bressan
DEXA (2)5
2014 Get a Sample for a Discount - Sampling-Based XML Data Pricing
Ruiming Tang, Antoine Amarilli, Pierre Senellart, Stéphane Bressan
DEXA (1)4
2014 L-opacity: Linkage-Aware Graph Anonymization
abstract
10.5441/002/edbt.2014.52
Sadegh Heyrani-Nobari, Panagiotis Karras, HweeHwa Pang, Stéphane Bressan
EDBT4
2013 Discovering Semantics from Data-Centric XML
Luochen Li, Thuy Ngoc Le, Huayu Wu 0001, Tok Wang Ling, Stéphane Bressan
DEXA (1)5
2013 Incremental Algorithms for Sampling Dynamic Graphs
Tuan Quang Phan, Stéphane Bressan
DEXA (1)3
2013 Publishing Trajectory with Differential Privacy: A Priori vs. A Posteriori Sampling Mechanisms
Dongxu Shao, Kaifeng Jiang, Thomas Kister, Stéphane Bressan, Kian-Lee Tan
DEXA (1)4
2013 Fast Community Detection
Yi Song 0005, Stéphane Bressan
DEXA (1)2
2013 Force-Directed Layout Community Detection
Yi Song 0005, Stéphane Bressan
DEXA (1)2
2013 What You Pay for Is What You Get
Ruiming Tang, Dongxu Shao, Stéphane Bressan, Patrick Valduriez
DEXA (2)3
2013 The Price Is Right - Models and Algorithms for Pricing Data
abstract
Data is a modern commodity. Yet the pricing models in useon electronic data markets either focus on the usage of computing resources,or are proprietary, opaque, most likely ad hoc, and not conduciveof a healthy commodity market dynamics. In this paper we propose ageneric data pricing model that is based on minimal provenance, i.e. minimalsets of tuples contributing to the result of a query.We show that theproposed model fulfills desirable properties such as contribution monotonicity,bounded-price and contribution arbitrage-freedom. We presenta baseline algorithm to compute the exact price of a query based onour pricing model. We show that the problem is NP-hard. We thereforedevise, present and compare several heuristics. We conduct a comprehensiveexperimental study to show their effectiveness and efficiency.
Ruiming Tang, Huayu Wu 0001, Zhifeng Bao, Stéphane Bressan, Patrick Valduriez
DEXA (2)4
2013 TOUCH: in-memory spatial join by hierarchical data-oriented partitioning
abstract
Efficient spatial joins are pivotal for many applications and particularly important for geographical information systems or for the simulation sciences where scientists work with spatial models. Past research has primarily focused on disk-based spatial joins; efficient in-memory approaches, however, are important for two reasons: a) main memory has grown so large that many datasets fit in it and b) the in-memory join is a very time-consuming part of all disk-based spatial joins.
Sadegh Heyrani-Nobari, Farhan Tauheed, Thomas Heinis, Panagiotis Karras, Stéphane Bressan, Anastasia Ailamaki
SIGMOD Conference5
2013 Publishing trajectories with differential privacy guarantees
abstract
The pervasiveness of location-acquisition technologies has made it possible to collect the movement data of individuals or vehicles. However, it has to be carefully managed to ensure that there is no privacy breach. In this paper, we investigate the problem of publishing trajectory data under the differential privacy model. A straightforward solution is to add noise to a trajectory - this can be done either by adding noise to each coordinate of the position, to each position of the trajectory, or to the whole trajectory. However, such naive approaches result in trajectories with zigzag shapes and many crossings, making the published trajectories of little practical use. We introduce a mechanism called SDD (Sampling Distance and Direction), which is ε-differentially private. SDD samples a suitable direction and distance at each position to publish the next possible position. Numerical experiments conducted on real ship trajectories demonstrate that our proposed mechanism can deliver ship trajectories that are of good practical utility.
Kaifeng Jiang, Dongxu Shao, Stéphane Bressan, Thomas Kister, Kian-Lee Tan
SSDBM3
2012 Discretionary social network data revelation with a user-centric utility guarantee
abstract
The proliferation of online social networks has created intense interest in studying their nature and revealing information of interest to the end user. At the same time, such revelation raises privacy concerns. Existing research addresses this problem following an approach popular in the database community: a model of data privacy is defined, and the data is rendered in a form that satisfies the constraints of that model while aiming to maximize some utility measure. Still, these is no consensus on a clear and quantifiable utility measure over graph data. In this paper, we take a different approach: we define a utility guarantee, in terms of certain graph properties being preserved, that should be respected when releasing data, while otherwise distorting the graph to an extend desired for the sake of confidentiality. We propose a form of data release which builds on current practice in social network platforms: A user may want to see a subgraph of the network graph, in which that user as well as connections and affiliates participate. Such a snapshot should not allow malicious users to gain private information, yet provide useful information for benevolent users. We propose a mechanism to prepare data for user view under this setting. In an experimental study with real data, we demonstrate that our method preserves several properties of interest more successfully than methods that randomly distort the graph to an equal extent, while withstanding structural attacks proposed in the literature.
Yi Song 0005, Panagiotis Karras, Sadegh Heyrani-Nobari, Giorgos Cheliotis, Mingqiang Xue, Stéphane Bressan
CIKM6
2012 Fast Identity Anonymization on Graphs
Yi Song 0005, Stéphane Bressan
DEXA (1)3
2012 A Framework for Conditioning Uncertain Relational Data
Ruiming Tang, Reynold Cheng, Huayu Wu 0001, Stéphane Bressan
DEXA (2)4
2012 A Hybrid Approach for General XML Query Processing
Huayu Wu 0001, Ruiming Tang, Tok Wang Ling, Stéphane Bressan
DEXA (1)5
2012 Sampling Connected Induced Subgraphs Uniformly at Random
Stéphane Bressan
SSDBM2
2012 SALSA: A Software System for Data Management and Analytics in Field Spectrometry
Baljeet Malhotra, John A. Gamon, Stéphane Bressan
SSDBM3
2012 Sensitive Label Privacy Protection on Social Network Data
Yi Song 0005, Panagiotis Karras, Stéphane Bressan
SSDBM4
2011 Generating Random Graphic Sequences
Stéphane Bressan
DASFAA (1)2
2011 Edit Distance between XML and Probabilistic XML Documents
Ruiming Tang, Huayu Wu 0001, Sadegh Heyrani-Nobari, Stéphane Bressan
DEXA (1)4
2011 Fast random graph generation
abstract
Today, several database applications call for the generation of random graphs. A fundamental, versatile random graph model adopted for that purpose is the Erdős-Rényi Γv,p model. This model can be used for directed, undirected, and multipartite graphs, with and without self-loops; it induces algorithms for both graph generation and sampling, hence is useful not only in applications necessitating the generation of random structures but also for simulation, sampling and in randomized algorithms. However, the commonly advocated algorithm for random graph generation under this model performs poorly when generating large graphs, and fails to make use of the parallel processing capabilities of modern hardware. In this paper, we propose PPreZER, an alternative, data parallel algorithm for random graph generation under the Erdős-Rényi model, designed and implemented in a graphics processing unit (GPU). We are led to this chief contribution of ours via a succession of seven intermediary algorithms, both sequential and parallel. Our extensive experimental study shows an average speedup of 19 for PPreZER with respect to the baseline algorithm.
Sadegh Heyrani-Nobari, Panagiotis Karras, Stéphane Bressan
EDBT4
2011 ASSIST: access controlled ship identification streams
abstract
The International Maritime Organization (IMO) requires a majority of cargo and passenger ships to use the Automatic Identification System (AIS) for navigation safety and traffic control. Distributing live AIS data on the Internet can offer a global view based on ships' status for both operational and analytical purposes to port authorities, shipping and insurance companies, cargo owners and ship captains and other stakeholders. Yet, uncontrolled, this distribution can seriously undermine navigation safety and security and the privacy of the various stakeholders. In this paper we present ASSIST, a system prototype based on our recently proposed access control framework, to protect data streams from unauthorized access. We demonstrate the effectiveness of the system in a real scenario with real AIS data streams.
Baljeet Malhotra, Wee-Juan Tan, Jianneng Cao, Thomas Kister, Stéphane Bressan, Kian-Lee Tan
GIS5
2011 Modelling and analysis of shipping networks from online maritime schedules
abstract
90% of the world trade is reportedly carried by sea. The analysis of shipping networks therefore can create invaluable insight into global trade. In this paper we study the appropriateness of various graph centrality measures to rate, compare and rank ports from various perspectives of a shipping network. In particular, we illustrate the potential of such analysis on the example of a shipping network constructed from the schedules, readily available on the World Wide Web, of one arbitrarily chosen shipping company.
Deepen Doshi, Baljeet Malhotra, Stéphane Bressan
iiWAS3
2011 On the privacy and utility of anonymized social networks
abstract
You are on Facebook or you are out. Of course, this assessment is controversial and its rationale arguable. It is nevertheless not far, for many of us, from the reason behind our joining social media and publishing and sharing details of our professional and private lives. Not only the personal details we may reveal but also the very structure of the networks themselves are sources of invaluable information for any organization wanting to understand and learn about social groups, their dynamics and their members. These organizations may or may not be benevolent. It is therefore important to devise, design and evaluate solutions that guarantee some privacy. One approach that attempts to reconcile the different stakeholders' requirement is the publication of a modified graph. The perturbation is hoped to be sufficient to protect members' privacy while it maintains sufficient utility for analysts wanting to study the social media as a whole. It is necessarily a compromise. In this paper we try and empirically quantify the inevitable trade-off between utility and privacy. We do so for one state-of-the-art graph anonymization algorithm that protects against most structural attacks, the k-automorphism algorithm. We measure several metrics for a series of real graphs from various social media before and after their anonymization under various settings.
Yi Song 0005, Sadegh Heyrani-Nobari, Panagiotis Karras, Stéphane Bressan
iiWAS5
2010 Privacy and Anonymization as a Service: PASS
Ghasem Heyrani-Nobari, Omar Boucelma, Stéphane Bressan
DASFAA (2)3
2010 A Simple, Yet Effective and Efficient, Sliding Window Sampling Algorithm
Wee Hyong Tok, Chedy Raïssi, Stéphane Bressan
DASFAA (1)4
2010 A Utilization of Schema Constraints to Transform Predicates in XPath Query
Dung Xuan Thi Le, Stéphane Bressan, Eric Pardede, David Taniar, Wenny Rahayu
DEXA (1)2
2010 Building a curated database for maritime research
abstract
The maritime industry is diverse in nature. It involves the construction of ships, platforms and ports, the navigation of ships and the operation of platforms and ports, as well as all the operation of myriads of services necessary to shipping and to the exploitation of ocean resources.
Stéphane Bressan
iiWAS1
2010 Semantic Transformation Approach with Schema Constraints for XPath Query Axes
Dung Xuan Thi Le, Stéphane Bressan, Eric Pardede, Wenny Rahayu, David Taniar
WISE2
2009 Ricochet: A Family of Unconstrained Algorithms for Graph Clustering
Derry Wijaya, Stéphane Bressan
DASFAA2
2009 Service computing in the clouds: what are the research challenges?
abstract
We have now fully entered the era of web applications and services championed by services such as Amazon Web Services, for example. Computing has evolved towards a service paradigm. Processing units and storage are managed in data centers whose architecture (clusters or grid) is transparent to the users and programmers. Both desktop and mobile multimedia applications are accessed from the web anytime anywhere. An unenthusiastic computer scientist could claim that there is nothing new under the sun since RPC (Remote Procedure Call). Therefore, this panel will try and shed a light on the research questions raised by this new technological ecosystem and by various scenarios for its evolution.
Stéphane Bressan
iiWAS1
2008 A random walk on the red carpet: rating movies with user reviews and pagerank
abstract
Although PageRank has been designed to estimate the popularity of Web pages, it is a general algorithm that can be applied to the analysis of other graphs other than one of hypertext documents. In this paper, we explore its application to sentiment analysis and opinion mining: i.e. the ranking of items based on user textual reviews. We first propose various techniques using collocation and pivot words to extract a weighted graph of terms from user reviews and to account for positive and negative opinions. We refer to this graph as the sentiment graph. Using PageRank and a very small set of adjectives (such as 'good', 'excellent', etc.) we rank the different items. We illustrate and evaluate our approach using reviews of box office movies by users of a popular movie review site. The results show that our approach is very effective and that the ranking it computes is comparable to the ranking obtained from the box office figures. The results also show that our approach is able to compute context-dependent ratings.
Derry Wijaya, Stéphane Bressan
CIKM2
2008 Twig'n Join: Progressive Query Processing of Multiple XML Streams
Wee Hyong Tok, Stéphane Bressan, Mong-Li Lee
DASFAA2
2008 A stratified approach to progressive approximate joins
abstract
Users often do not require a complete answer to their query but rather only a sample. They expect the sample to be either the largest possible or the most representative (or both) given the resources available. We call the query processing techniques that deliver such results 'approximate'. Processing of queries to streams of data is said to be 'progressive' when it can continuously produce results as data arrives. In this paper, we are interested in the progressive and approximate processing of queries to data streams when processing is limited to main memory. In particular, we study one of the main building blocks of such processing: the progressive approximate join. We devise and present several novel progressive approximate join algorithms. We empirically evaluate the performance of our algorithms and compare them with algorithms based on existing techniques. In particular we study the trade-off between maximization of throughput and maximization of representativeness of the sample.
Wee Hyong Tok, Stéphane Bressan, Mong-Li Lee
EDBT2
2007 Semantic XPath Query Transformation: Opportunities and Performance
Dung Xuan Thi Le, Stéphane Bressan, David Taniar, Wenny Rahayu
DASFAA2
2007 RRPJ: Result-Rate Based Progressive Relational Join
Wee Hyong Tok, Stéphane Bressan, Mong-Li Lee
DASFAA2
2007 Danaïdes: Continuous and Progressive Complex Queries on RSS Feeds
Wee Hyong Tok, Stéphane Bressan, Mong-Li Lee
DASFAA2
2007 Progressive High-Dimensional Similarity Join
Wee Hyong Tok, Stéphane Bressan, Mong-Li Lee
DEXA2
2007 Journey to the Centre of the Star: Various Ways of Finding Star Centers in Star Clustering
Derry Wijaya, Stéphane Bressan
DEXA2
2006 Interactive Discovery and Composition of Complex Web Services
Sergey A. Stupnikov, Leonid A. Kalinichenko, Stéphane Bressan
ADBIS3
2006 Co-Reference Resolution for the Indonesian Language Using Association Rules
Indra Budi, Stéphane Bressan
iiWAS2
2006 Progressive Spatial Join
abstract
In spatial data exploration and analysis, the system would present a user with initial promising results and empower the user to modify runtime query parameters. The high degree of interactivity would significantly reduce users’ waiting time for results that are not useful, and then having to re-issue a new query. To support this level of interaction during query processing, it necessitates the study of adaptive and progressive spatial query processing techniques that can deliver initial results quickly and adapt to run-time fluctuations during the delivery of remote data. Our goal is to design a generic framework for adaptive and progressive spatial query processing. In this paper, we present our ongoing work on designing progressive spatial join algorithm as an initial step.
Wee Hyong Tok, Stéphane Bressan, Mong-Li Lee
SSDBM2
2005 Environmental Noise Classification for Multimedia Libraries
Stéphane Bressan, Boon Tiang Tan
DEXA1
2005 Accelerating queries by pruning XML documents
Stéphane Bressan, Barbara Catania, Zoé Lacroix, Ying Guang Li, Anna Maddalena
Data Knowl. Eng.1
2004 Adaptive Double Routing Indices: Combining Effectiveness and Efficiency in P2P Systems
Stéphane Bressan, Achmad Nizar Hidayanto, Chu Yee Liau, Zainal A. Hasibuan
DEXA1
2004 Performance Evaluation of a Simple Update Model and a Basic Locking Mechanism for Broadcast Disks
Stéphane Bressan, Guo Yuzhi
DEXA1
2004 XML, A Substance or Hype
Stéphane Bressan
iiWAS1
2004 Querying high-dimensional data in single-dimensional space
Cui Yu, Stéphane Bressan, Beng Chin Ooi, Kian-Lee Tan
VLDB J.2
2003 Adaptive Peer-to-Peer Routing with Proximity
Chu Yee Liau, Achmad Nizar Hidayanto, Stéphane Bressan
DEXA3
2003 What Shall We Expect From E-Learning?
Stéphane Bressan
iiWAS1
2003 Association Rules Mining for Name Entity Recognition
abstract
We propose a new name entity class extraction method based on association rules. We evaluate and compare the performance of our method with the state of the art maximum entropy method. We show that our method consistently yields a higher precision at a competitive level of recall. This result makes our method particularly suitable for tasks whose requirements emphasize the quality rather than the quantity of results.
Indra Budi, Stéphane Bressan
WISE2
2002 Information Extraction - Tree Alignment Approach to Pattern Discovery in Web Documents
Ajay Hemnani, Stéphane Bressan
DEXA2
2002 dbRouter - A Scaleable and Distributed Query Optimization and Processing Framework
Wee Hyong Tok, Stéphane Bressan
DEXA2
2002 Efficient and Adaptive Processing of Multiple Continuous Queries
Wee Hyong Tok, Stéphane Bressan
EDBT2
2002 Predator-Miner: Ad hoc Mining of Associations Rules within a Database Management System
abstract
We present a prototype system, Predator-Miner, which extends Predator with an relational-like association rule mining operator to support data mining operations. Predator-Miner allows a user to combine association rule mining queries with SQL queries. This approach towards tight integration differs from existing techniques of using user-defined functions (UDFs), stored procedures, or re-expressing a mining query as several SQL queries in two aspects. First, by encapsulating the task of association rule mining in a relational operator, we allow association rule mining to be considered as part of the query plan, on which query optimization can be performed on the mining query holistically. Second, by integrating it as a relational operator, we can leverage on the mature field of relational database technology. We extend Predator to support a variant of DMQL, and allow SQL and DMQL to be intermixed in a query. We also demonstrate a cost-based mining query optimization framework.
Wee Hyong Tok, Twee-Hee Ong, Wai Lup Low, Indriyati Atmosukarto, Stéphane Bressan
ICDE5
2001 X007: Applying 007 Benchmark to XML Query Processing Tool
abstract
If XML is to play the critical role of the lingua franca for Internet data interchange that many predict, it is necessary to start designing and adopting benchmarks allowing the comparative performance analysis of the tools being developed and proposed. The effectiveness of existing XML query languages has been studied by many, with a focus on the comparison of linguistic features, implicitly reflecting the fact that most XML tools exist only on paper. In this paper, with a focus on efficiency and concreteness, we propose a pragmatic first step toward the systematic benchmarking of XML query processing platforms with an initial focus on the data (versus document) point of view. We propose XOO7, an XML version of the OO7 benchmark. We discuss the applicability of XOO7, its strengths, limitations and the extensions we are considering. We illustrate its use by presenting and discussing the performance comparison against XOO7 of three different query processing platforms for XML.
Stéphane Bressan, Gillian Dobbie, Zoé Lacroix, Mong-Li Lee, Ying Guang Li, Ullas Nambiar, Bimlesh Wadhwa
CIKM1
2001 Hybrid Transformation for Indexing and Searching Web Documents in the Cartographic Paradigm
Fiona Lee, Stéphane Bressan, Beng Chin Ooi
Inf. Syst.2
2000 A Framework for Modeling Buffer Replacement Strategies
Stéphane Bressan, Chong Leng Goh, Beng Chin Ooi, Kian-Lee Tan
CIKM1
2000 Rule-Assisted Prefetching in Web-Server Caching
abstract
Web servers manage large numbe rof documents of widely variable sizes.Moreover, the access patterns on the documents may also c hange over time.While some documents are highly popular over a prolonged period of time, we expe c tnewly added documents to increase in popularity while demand for most older documents decreases.It is therefore important to design eective caching strategy at the web server.In this paper, we present our approach to the problem.Our main contribution lies in the design of a novel prefetching strategy, called RAP.RAP identi es a set of association rules from the Web server's access log.Unlike existing mining strategy, RAP's miner values recently added log records more than earlier log records.Based on the rules, RAP predicts and prefetches documents from users initial requests.We conducted extensive study to evaluate RAP.The results show that RAP signi cantly outperforms existing schemes.We also show that the mining and caching cost is relatively low.
Bin Lan, Stéphane Bressan, Beng Chin Ooi, Kian-Lee Tan
CIKM2
2000 Indexing the Edges - A Simple and Yet Efficient Approach to High-Dimensional Indexing
abstract
In this paper, we propose a new tunable index scheme, called iMinMax(Ο), that maps points in high dimensional spaces to single dimension values determined by their maximum or minimum values among all dimensions. By varying the tuning “knob” Ο, we can obtain different family of iMinMax structures that are optimized for different distributions of data sets. For a d-dimensional space, a range query need to be transformed into d subqueries. However, some of these subqueries can be pruned away without evaluation, further enhancing the efficiency of the scheme. Experimental results show that iMinMax(Ο) can outperform the more complex Pyramid technique by a wide margin.
Beng Chin Ooi, Kian-Lee Tan, Cui Yu, Stéphane Bressan
PODS4
2000 Integrating Replacement Policies in StorM: An Extensible Approach
abstract
No abstract available.
Chong Leng Goh, Beng Chin Ooi, Stéphane Bressan, Kian-Lee Tan
SIGMOD Conference3
2000 Global Atlas: Calibrating and Indexing Documents from the Internet in the Cartographic Paradigm
abstract
Global Atlas is a geographical search engine. It indexes maps, satellite and aerial pictures, as well as HTML documents available on the World Wide Web. The Global Atlas leverages on the cartographic paradigm to provide a very natural support for indexing, searching and sharing information. It allows the design of intuitive user interfaces and the use of natural visual feedback. HTML documents are best indexed according to the geographical regions to which they are topically associated, and maps in the form of GIF and JPEG images are indexed to create a huge patchwork of maps. However, maps come in a variety of unspecified coordinate systems and projections. This entails calibrating different maps to a single reference coordinate system. We discuss the design issues in building a geographical search engine, and focus on the calibration of maps.
Fiona Lee, Stéphane Bressan, Beng Chin Ooi
WISE2
2000 Mining Term Association Rules for Automatic Global Query Expansion: Methodology and Preliminary Results
abstract
The authors are looking at the mining of association between terms for the automatic expansion of queries. The technique used for the discovery of the associations is association rule mining (R. Agrawal et al., 1996). The technique proposed is more flexible than previous techniques based on term co-occurrence since it takes into account not only the co-occurrence frequency but also the confidence and direction of the association rules. We have been able to consistently improve the effectiveness of the retrieval over the set of 48 test queries on the Associated Press 1990 news wires corpus of the TREC4 benchmark by query expansion using term association rules.
Stéphane Bressan, Beng Chin Ooi
WISE2
1999 Context Interchange: New Features and Formalisms for the Intelligent Integration of Information
abstract
TheContext Interchange strategypresents a novel perspective for mediated data access in which semantic conflicts among heterogeneous systems are not identified a priori, but are detected and reconciled by acontext mediatorthrough comparison ofcontexts axiomscorresponding to the systems engaged in data exchange. In this article, we show that queries formulated on shared views, export schema, and shared “ontologies” can be mediated in the same way using theContext Interchange framework. The proposed framework provides a logic-based object-oriented formalsim for representing and reasoning about data semantics in disparate systems, and has been validated in a prototype implementation providing mediated data access to both traditional and web-based information sources.
Cheng Hian Goh, Stéphane Bressan, Stuart E. Madnick, Michael D. Siegel
ACM Trans. Inf. Syst.2
1998 An Active Conceptual Model for Fixed Income Securities Analysis for Multiple Financial Institutions
Allen Moulton, Stéphane Bressan, Stuart E. Madnick, Michael D. Siegel
ER2
1998 Answering Queries in Context
Stéphane Bressan, Cheng Hian Goh
FQAS1
1997 The COntext INterchange Mediator Prototype
abstract
The Context Interchange strategy presents a novel approach for mediated data access in which semantic conflicts among heterogeneous systems are not identified a priori, but are detected and reconciled by a context mediator through comparison of contexts. This paper reports on the implementation of a Context Interchange Prototype which provides a concrete demonstration of the features and benefits of this integration strategy.
Stéphane Bressan, Cheng Hian Goh, Kofi Fynn, Marta Jessica Jakobisiak, Karim Hussein, Henry B. Kon, Thomas Lee, Stuart E. Madnick, Tito Pena, Jessica Qu, Annie W. Shum, Michael D. Siegel
SIGMOD Conference1