EDBT 2026 Demo / reviewers in the wild / expert
Radu State
dblp:05/6228 · also Radu Valentin State
· DBLP profile ↗
13ranked-venue papers in the field
0as first author
7since 2021 · last 2025
0000-0002-4751-9577ORCID · verified
Domains — venue-derived; a paper can count in several
Big Data, Cloud & Distributed Data Systems · 6Database Systems & Data Management · 3Data Mining & Knowledge Discovery · 3Information Retrieval & Web Search · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Leveraging Large Language Models to Build Computationally Efficient Models for Sustainable Finance Investment Decision Supportabstractpeer reviewed Loris Bergeron, Jérôme François, Radu State, Jean Hilger |
IEEE Big Data | 3 |
| 2025 | CAR-RAG: Category-Aware Hybrid Retrieval-Augmented Generation for Hallucination Mitigationabstractpeer reviewed Tatiana Petrova, Dmitrii Koriakov, Radu State |
IEEE Big Data | 3 |
| 2024 | Federated Learning-Based Tokenizer for Domain-Specific Language Models in Finance
Farouk Damoun, Hamida Seba, Radu State |
ASONAM (2) | 3 |
| 2024 | LongKey: Keyphrase Extraction for Long DocumentsabstractIn an era of information overload, manually annotating the vast and growing corpus of documents and scholarly papers is increasingly impractical. Automated keyphrase extraction addresses this challenge by identifying representative terms within texts. However, most existing methods focus on short documents (up to 512 tokens), leaving a gap in processing long-context documents. In this paper, we introduce LongKey, a novel framework for extracting keyphrases from lengthy documents, which uses an encoder-based language model to capture extended text intricacies. LongKey uses a max-pooling embedder to enhance keyphrase candidate representation. Validated on the comprehensive LDKP datasets and six diverse, unseen datasets, LongKey consistently outperforms existing unsupervised and language model-based keyphrase extraction methods. Our findings demonstrate LongKey’s versatility and superior performance, marking an advancement in keyphrase extraction for varied text lengths and domains. Jeovane Honório Alves, Radu State, Cinthia Obladen de Almendra Freitas, Jean Paul Barddal |
IEEE Big Data | 2 |
| 2024 | Privacy-Preserving Behavioral Anomaly Detection in Dynamic Graphs for Card Transactions
Farouk Damoun, Hamida Seba, Radu State |
WISE (5) | 3 |
| 2023 | A Decentralized Super AppabstractThe massive usage of mobile phones connected to the internet over the last years has changed the business model of all sort of companies and services around the globe. By turning on your mobile phone and connecting it to the internet, it is possible to reach a countless amount of information and the products are literally at your fingertips. In view of this new reality, the concept of a mobile app that aggregates services (i.e., super apps) seems to be the next game-changer for the mobile apps industry, and the challenges related to security and privacy are key aspects for keeping user data safe. In this paper, we propose a complete auditable decentralised super app architecture in order to prevent the risks of a centralised ownership, to provide a full log of shared data and to control data access between the stakeholders. An environment containing all the proposed implementations were built and successfully tested to prove the feasibility of deploying such architecture in real world applications. Fernando Kaway Carvalho Ota, Cristina Godoy Bernardo de Oliveira, Rafael Meira Silva, Radu State |
MDM | 4 |
| 2022 | Mobile Application Behaviour Anomaly Detection based on API CallsabstractIn many sectors, such as banking, mobile applications became the main interaction channel between customers and institutions. This fast-paced market movement forced an increase in the variety of services offered on the mobile applications. Mobile applications usually exchange information via REST APIs. This range of services with the exposure that APIs imposed to the organizations widen the attack surface, where opportunists can try to take advantages of bad implemented APIs. In light of that, our work focuses in detecting malicious attempts to accessing the APIs by analysing the behaviour of the mobile application user sessions and capturing the anomalies that deviates from the expected. To this end, we tested a pipeline that extracts features based on Markov Chain, Levenshtein Distance and LSTM frequencies to create a machine learning model that was tested in 40 days of data from a retail banking mobile application. By mixing the features from Markov Chain and Levenshtein Distance, our model reached over 96.61% of f-measure. We also present results to discourage the use of LSTM for this use case because it is computationally costly and in fact degraded the performance of the model based on the other frequencies. Fernando Kaway Carvalho Ota, Farouk Damoun, Sofiane Lagraa, Patricia Becerra-Sánchez, Christophe Atten, Jean Hilger, Radu State |
IEEE Big Data | 7 |
| 2019 | Blockchain-based micropayment systems: economic impactabstractThe inception of blockchain catapulted the development of innovative use cases utilizing the trustless, decentralized environment, empowered by cryptocurrencies. The envisaged benefits of the technology includes the divisible nature of a cryptocurrency, that can facilitate payments in fractions of a cent, enabling micropayments through the blockchain. Micropayments are a critical tool to enable financial inclusion and to aid in global poverty alleviation. The paper conducts a study on the economic impact of blockchain-based micropayment systems, emphasizing their significance for socioeconomic benefit and financial inclusion. The paper also highlights the contribution of blockchain-based micropayments to the cybercrime economy, indicating the critical need of economic regulations to curtail the growing threat posed by the digital payment mechanism. Nida Khan, Tabrez Ahmad, Radu State |
IDEAS | 3 |
| 2017 | Improving Real-Time Bidding Using a Constrained Markov Decision Process
Manxing Du, Redouane Sassioui, Georgios Varisteas, Radu State, Mats Brorsson, Omar Cherkaoui |
ADMA | 4 |
| 2017 | Your Moves, Your Device: Establishing Behavior Profiles Using Tensors
Eric Falk, Jérémy Charlier, Radu State |
ADMA | 3 |
| 2017 | Query-able Kafka: An agile data analytics pipeline for mobile wireless networksabstractDue to their promise of delivering real-time network insights, today's streaming analytics platforms are increasingly being used in the communications networks where the impact of the insights go beyond sentiment and trend analysis to include real-time detection of security attacks and prediction of network state (i.e., is the network transitioning towards an outage). Current streaming analytics platforms operate under the assumption that arriving traffic is to the order of kilobytes produced at very high frequencies. However, communications networks, especially the telecommunication networks, challenge this assumption because some of the arriving traffic in these networks is to the order of gigabytes, but produced at medium to low velocities. Furthermore, these large datasets may need to be ingested in their entirety to render network insights in real-time. Our interest is to subject today's streaming analytics platforms --- constructed from state-of-the art software components (Kafka, Spark, HDFS, ElasticSearch) --- to traffic densities observed in such communications networks. We find that filtering on such large datasets is best done in a common upstream point instead of being pushed to, and repeated, in downstream components. To demonstrate the advantages of such an approach, we modify Apache Kafka to perform limited native data transformation and filtering, relieving the downstream Spark application from doing this. Our approach outperforms four prevalent analytics pipeline architectures with negligible overhead compared to standard Kafka. (Our modifications to Apache Kafka are publicly available at https://github.com/Esquive/queryable-kafka.git) Eric Falk, Vijay K. Gurbani, Radu State |
Proc. VLDB Endow. | 3 |
| 2016 | Behavior profiling for mobile advertisingabstractBehavioral and targeted profiling of users is an important task in marketing and in the advertising industry. Being able to match a given user profile to an advertising that leads to effective purchases is challenging because of a very tiny proportion of users willing to purchase goods and thus monetize the advertising. With such proportions being less than one percent of the overall user population, efficient feature extraction and modeling techniques are required in order to capture and recognize the potential consumers. This paper proposes a new approach for modeling the observed behavior in a mobile advertising platform, where time related features are correlated with additional system level and campaign related performance statistics. We capture the temporal behavior with Hawkes processes and use the estimated parameters as additional features for predicting if a given user profile will be a revenue generating customer. Manxing Du, Radu State, Mats Brorsson, Tigran Avanesov |
BDCAT | 2 |
| 2016 | Neighborhood features help detecting non-technical losses in big data setsabstractElectricity theft occurs around the world in both developed and developing countries and may range up to 40% of the total electricity distributed. More generally, electricity theft belongs to non-technical losses (NTL), which occur during the distribution of electricity in power grids. In this paper, we build features from the neighborhood of customers. We first split the area in which the customers are located into grids of different sizes. For each grid cell we then compute the proportion of inspected customers and the proportion of NTL found among the inspected customers. We then analyze the distributions of features generated and show why they are useful to predict NTL. In addition, we compute features from the consumption time series of customers. We also use master data features of customers, such as their customer class and voltage of their connection. We compute these features for a Big Data base of 31M meter readings, 700K customers and 400K inspection results. We then use these features to train four machine learning algorithms that are particularly suitable for Big Data sets because of their parallelizable structure: logistic regression, k-nearest neighbors, linear support vector machine and random forest. Using the neighborhood features instead of only analyzing the time series has resulted in appreciable results for Big Data sets for varying NTL proportions of 1%-90%. This work can therefore be deployed to a wide range of different regions. Patrick O. Glauner, Jorge Augusto Meira, Lautaro Dolberg, Radu State, Franck Bettinger, Yves Rangoni |
BDCAT | 4 |