Danilo Giordano

dblp:160/8766 · DBLP profile ↗
← Back
5ranked-venue papers in the field
0as first author
5since 2021 · last 2025
0000-0002-6987-2064ORCID · verified

Domains — venue-derived; a paper can count in several

Data Mining & Knowledge Discovery · 2Big Data, Cloud & Distributed Data Systems · 2Knowledge Engineering, Semantic Web & Information Systems · 1
YearPublicationVenuePosition
2025 MAD: Multicriteria Anomaly Detection of Suspicious Financial Accounts from Billions of Cash Transactions
abstract
This paper presents a real-world deployment case study on using unsupervised anomaly detection for Anti-Money Laundering (AML).Using more than 2 billion anonymized bank transactions that Intesa Sanpaolo, a primary Italian financial institution, registered over 8 months, we developed, tuned and deployed a machine learning pipeline in production.Experts from Intesa Sanpaolo validated the performance of our approach against the institution's traditional rule-based system and checked new real-world cases the system allowed them to identify.Besides increasing both precision and recall by a factor of 6 in the detection of high-risk cases, our pipeline raises 200+ additional alerts during the 8-month period, manually identified by branch managers, but missed by the rulebased system.More importantly, a manual inspection of 100 new unseen cases revealed 28 significant previously unreported cases.The pipeline, now fully deployed in Intesa Sanpaolo's Transaction Monitoring system, highlights the advantages of machine learning over traditional approaches typically adopted in this traditionally very conservative sector.
Giordano Paoletti, Flavio Giobergia, Danilo Giordano, Luca Cagliero, Silvia Ronchiadin, Dario Moncalvo, Marco Mellia, Elena Baralis
KDD (2)3
2025 Towards Better Generalization and Interpretability in Unsupervised Concept-Based Models
Francesco De Santis, Philippe Bich, Gabriele Ciravegna, Pietro Barbiero, Tania Cerquitelli, Danilo Giordano
ECML/PKDD (3)6
2023 GLEm-Net: Unified Framework for Data Reduction with Categorical and Numerical Features
abstract
In the era of Big Data, effective data reduction through feature selection is of paramount importance for machine learning. This paper presents GLEm-Net (Grouped Lasso with Embeddings Network), a novel neural framework that seamlessly processes both categorical and numerical features to reduce the dimensionality of data while retaining as much information as possible. By integrating embedding layers, GLEm-Net effectively manages categorical features with high cardinality and compresses their information in a less dimensional space. By using a grouped Lasso penalty function in its architecture, GLEm-Net simultaneously processes categorical and numerical data, efficiently reducing high-dimensional data while preserving the essential information. We test GLEm-Net with a real-world application in an industrial environment where 6 million records exist and each is described by a mixture of 19 numerical and 7 categorical features with a strong class imbalance. A comparative analysis using state-of-the-art methods shows that despite the difficulty of building a high-performance model, GLEm-Net outperforms the other methods in both feature selection and classification, with a better balance in the selection of both numerical and categorical features.
Francesco De Santis, Danilo Giordano, Marco Mellia, Alessia Damilano
IEEE Big Data2
2023 Data driven scalability and profitability analysis in free floating electric car sharing systems
Alessandro Ciociola, Danilo Giordano, Luca Vassio, Marco Mellia
Inf. Sci.2
2022 Legal Entity Disambiguation for Financial Crime Detection
abstract
Transaction Monitoring is one of the main labor-intensive tasks of anti-financial crime and it requires to scrutinise billions of transactions per month against possible crimes. The first step in the process is the correct identification of the involved parties. This foundational step defines the focal entities on which transaction monitoring algorithms rely to spot suspicious events. Unfortunately, the loose syntax of protocols and the free text fields of inter-banking communications make party disambiguation particularly challenging. The first step of a fully automated data-driven strategy is thus the detection of the actual entity owning or using a given account.In this paper, we leverage data-driven techniques to identify and disambiguate the owners of accounts involved in cross-border international transactions when a Financial Institution only knows a minority fraction of such parties as its own customers. For this, we propose a data science pipeline relying on hierarchical clustering to capture similarities among names of parties involved in actual transactions. We test and tune the proposed approach using a large, real-world, multi-language, proprietary dataset of actual international transactions. Our highly parallel implementation completes the identification of parties that share an account and identifies all accounts owned by a party with f-score higher than 0.8.
Jacopo Fior, Thomas Favale, Luca Cagliero, Danilo Giordano, Marco Mellia, Elena Baralis, Silvia Ronchiadin, Paolo Baracco, Dario Moncalvo
IEEE Big Data4