VLDB 2026 Research / reviewers in the wild / expert
Hông-Ân Sandlin
dblp:336/2656 · also Hông-Ân Cao
· DBLP profile ↗
9ranked-venue papers
5as first author
4since 2021 · last 2025
0000-0002-8535-2982ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 4 first-author · 2 since 2021Databases, data management, data science and information retrieval · 5 · 4 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 first-authorSystems, architecture and hardware · 2 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
1 paper |
Web and social media mining · 50% Data integration and cleaning · 50% | |
| Artificial intelligence
1 paper |
Knowledge representation and reasoning · 100% |
Topics — the 4 heaviest of 5, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Data integration and cleaning
data fusion |
0.9 | 1 | 2025 | A survey of multimodal event detection based on data fusion · VLDB J. 2025 |
Web and social media mining
event detection |
0.9 | 1 | 2025 | A survey of multimodal event detection based on data fusion · VLDB J. 2025 |
Knowledge, reasoning and agents › Knowledge representation and reasoning
case-based reasoning |
0.7 | 1 | 2023 | Case-Based Reasoning with Language Models for Classification of Logical Fallacies · IJCAI 2023 |
Knowledge, reasoning and agents › Knowledge representation and reasoning › argumentation
logical fallacy classification |
0.7 | 1 | 2023 | Case-Based Reasoning with Language Models for Classification of Logical Fallacies · IJCAI 2023 |
Methods — techniques the papers use, named apart from their topics
systematic literature review · 1.7retrieval · 0.7language model · 0.7case-based reasoning · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | A survey of multimodal event detection based on data fusionabstractAbstract With the emergence of the Internet of Things (IoT) and the rise of shared multimedia content on social media networks, available datasets have become increasingly heterogeneous. Several multimodal techniques for detecting events in data of different types and formats have emerged. Those techniques implement various detection algorithms and present different trade-offs in terms of data fusion. Unfortunately, little is known about their underlying detection mechanisms, as existing comparisons are limited to either unimodal event detection techniques or specific types or representations for multimodal techniques. Understanding the behavior of multimodal event detection techniques remains an acute open research problem. In this work, we present a systematic literature review of multimodal event detection techniques. We describe how various techniques leverage information from different modalities through data fusion. We further propose a novel taxonomy of multimodal event detection techniques according to their temporal orientation and the inner workings of their detection mechanism. Finally, we analyze the datasets and metrics used in previous works as well as their reported results. Our survey allows to uncover the properties of each approach and discuss future research directions in this field. Manuel Mondal, Mourad Khayati, Hông-Ân Sandlin, Philippe Cudré-Mauroux |
VLDB J. | 3 |
| 2024 | A Big Data architecture for early identification and categorization of dark web sitesabstractThe dark web has become notorious for its association with illicit activities and there is a growing need for systems to automate the monitoring of this space. This paper proposes an end-to-end scalable architecture for the early identification of new Tor sites and the daily analysis of their content. The solution is built using an Open Source Big Data stack for data serving with Kubernetes, Kafka, Kubeflow, and MinIO, continuously discovering onion addresses in different sources (threat intelligence, code repositories, web-Tor gateways, and Tor repositories), downloading the HTML from Tor and deduplicating the content using MinHash LSH, and categorizing with the BERTopic modeling (SBERT embedding, UMAP dimensionality reduction, HDBSCAN document clustering and c-TF-IDF topic keywords). In 93 days, the system identified 80,049 onion services and characterized 90% of them, addressing the challenge of Tor volatility. A disproportionate amount of repeated content is found, with only 6.1% unique sites. From the HTML files of the dark sites, 31 different low-topics are extracted, manually labeled, and grouped into 11 high-level topics. The five most popular included sexual and violent content, repositories, search engines, carding, cryptocurrencies, and marketplaces. During the experiments, we identified 14 sites with 13,946 clones that shared a suspiciously similar mirroring rate per day, suggesting an extensive common phishing network. Among the related works, this study is the most representative characterization of onion services based on topics to date. Javier Pastor-Galindo, Hông-Ân Sandlin, Félix Gómez Mármol, Gérôme Bovet, Gregorio Martínez Pérez |
Future Gener. Comput. Syst. | 2 |
| 2023 | Case-Based Reasoning with Language Models for Classification of Logical FallaciesabstractThe ease and speed of spreading misinformation and propaganda on the Web motivate the need to develop trustworthy technology for detecting fallacies in natural language arguments. However, state-of-the-art language modeling methods exhibit a lack of robustness on tasks like logical fallacy classification that require complex reasoning. In this paper, we propose a Case-Based Reasoning method that classifies new cases of logical fallacy by language-modeling-driven retrieval and adaptation of historical cases. We design four complementary strategies to enrich input representation for our model, based on external information about goals, explanations, counterarguments, and argument structure. Our experiments in in-domain and out-of-domain settings indicate that Case-Based Reasoning improves the accuracy and generalizability of language models. Our ablation studies suggest that representations of similar cases have a strong impact on the model performance, that models perform well with fewer retrieved cases, and that the size of the case database has a negligible effect on the performance. Finally, we dive deeper into the relationship between the properties of the retrieved cases and the model performance. Zhivar Sourati, Filip Ilievski, Hông-Ân Sandlin, Alain Mermoud |
IJCAI | 3 |
| 2023 | Robust and explainable identification of logical fallacies in natural language arguments
Zhivar Sourati, Vishnu Priya Prasanna Venkatesh, Darshan Deshpande, Himanshu Rawlani, Filip Ilievski, Hông-Ân Sandlin, Alain Mermoud |
Knowl. Based Syst. | 6 |
| 2016 | Leveraging user expertise in collaborative systems for annotating energy datasetsabstractWhile tasks such as segmenting images or determining the sentiment expressed in a sentence can be assigned to regular users, some others require background knowledge and thus, the selection of expert users. In the case of energy datasets, acquiring data represents an obstacle to develop data-driven methods, due to prohibitive monetary and time costs linked to the instrumentation of households in order to monitor the energy consumption. More so, most datasets only contain pure power time series, despite labels being required to determine when a device is in use from when it is idle (incurring stand-by consumption or being off), and by extension to separate human activities triggering the consumption from the baseline consumption. We build upon our Collaborative Annotation Framework for Energy Datasets (CAFED) to evaluate and distinguish the performance of expert users against that of regular users. Through a user study with curated benchmark annotation tasks, we provide data-driven and efficient techniques to detect weak and adversarial workers and promote users when the contributors' user-base is limited. Additionally, we show that if carefully selected, the seed gold standard tasks can be reduced to a small number of tasks that are representative enough to determine the user's expertise and predict crowd-combined annotations with high precision. Hông-Ân Sandlin, Felix Rauchenstein, Tri Kurniawan Wijaya, Karl Aberer, Nuno Nunes 0001 |
IEEE BigData | 1 |
| 2016 | Estimating human interactions with electrical appliances for activity-based energy savings recommendationsabstractSince the power consumption of different electrical appliances in a household can be recorded by individual smart meters, it becomes possible to start considering in more detail the interactions of the residents with those devices throughout the day. Appliances' usages should not be considered as independent events, but rather as enablers for activities. Leveraging activity knowledge over time will allow us to design personalized energy efficient measures. We envision the design of future ambient intelligence systems, where the smart home can optimize the energy consumption in regards to the lifestyles of its residents and the smart grid's needs. In this work, we propose an automated method for determining when an electrical device is triggered by households' residents solely from its power trace. Knowing when an appliance is in use is required for identifying recurrent patterns that could later be understood as activities. Hông-Ân Sandlin, Tri Kurniawan Wijaya, Karl Aberer, Nuno Nunes 0001 |
IEEE BigData | 1 |
| 2016 | Temporal association rules for electrical activity detection in residential homesabstractAttaining energy efficiency requires understanding human behaviors triggering energy consumption within households. In conjunction to providing appliance-level feedback, targeting human activities that involve the usage of electrical appliances can provide a higher abstraction level to bring awareness to the electricity wastage. In this paper, we make use of a large dataset with appliance- and circuit-level power data and provide a framework for determining temporal sequential association rules. Sequences of time intervals where the appliances are in usage can vary in their order, duration and the time elapsed between these events. Our contribution consists in providing a full pipeline for mining frequent sequential itemsets and a novel way to discover the time windows during which these sequences of events occur and to capture their variance in terms of duration and order. Our method is data-driven and relies on the data's statistical properties and allows us to avoid an exhaustive search for the time windows' sizes, by relying instead on machine learning techniques to identify and predict those time windows. Hông-Ân Sandlin, Tri Kurniawan Wijaya, Karl Aberer, Nuno Nunes 0001 |
IEEE BigData | 1 |
| 2015 | A collaborative framework for annotating energy datasetsabstractTargeting human activities responsible for the energy consumption instead of focusing solely on single appliance feedback for achieving energy efficiency in residential homes would link human behaviors to the resulting energy consumption. To this end, learning when appliances are in an active or idle state and the related user activity is crucial. Until smart appliances become widespread and can communicate their internal state, identifying when the residents interact with the appliances has to be determined from the available information that can be recorded from these devices. Developing and validating learning models require ground truth in the form of annotations to indicate when an appliance is active or idle. Launching data collection campaigns to incorporate these missing ground truth data involves careful planning before the roll-out of the experiment. Prohibitive costs for the hardware and time investment to monitor the deployed equipment are necessary for quality data. As such, publicly released datasets containing appliance-level data offer a basis for most researchers. This paper addresses these challenges by providing a collaborative web-based framework to retrofit labeling on existing datasets. The platform is publicly available, applies the wisdom of the crowd in the realm of energy research and leverages gamification techniques to encourage users' active contribution. The access to the platform and furthermore to the expert manually labeled dataset intends to enable future research and foster more collaboration in this area. Hông-Ân Sandlin, Tri Kurniawan Wijaya, Karl Aberer, Nuno Nunes 0001 |
IEEE BigData | 1 |
| 2013 | Are domestic load profiles stable over time? An attempt to identify target households for demand side management campaignsabstractElaborating demand side management strategies is crucial for integrating electricity from renewable sources into the electrical grid. Though future demand side will largely depend on an automatic control of larger loads, it is also widely agreed upon that consumer behavior will play an important role as well - be it by purchasing respective automation techniques or by shifting the use of appliances to other times of the day. Doing so, it becomes possible to select households that offer sufficient load shifting potential, and to overcome undirected and thus, expensive campaigns. To our knowledge, this perspective is still under-researched, especially when it comes to clustering methods on load consumption data with a focus on peak detection accuracy to provide customer segmentation. Using the data collected in the Irish CER dataset, which contains readings for more than 4000 residential customers over a period of 18 months at 30-minute intervals, we show that the whole clustering of the time series, with a few adaptations on the usage of the K-Means algorithm, provides better clustering results without sacrificing practical feasibility. Characteristic load profiles allow us to segment the customers, address groups of households with similar consumption patterns and determine on the fly the cluster membership of a given load curve. This will support decision making regarding the investments in load shifting campaigns to prevent over or under-dimensioning linked to peak energy demand. Hông-Ân Sandlin, Christian Beckel, Thorsten Staake |
IECON | 1 |