Fernando Pérez-Cruz

dblp:75/805 · DBLP profile ↗
← Back
86ranked-venue papers
18as first author
19since 2021 · last 2025
0000-0001-8996-5076ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 29 · 7 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 16 · 7 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 16 · 2 first-author · 2 since 2021Theory of computation · 8 · 1 first-authorComputer networks · 7 · 1 first-author · 1 since 2021Security and privacy · 7 · 5 since 2021Databases, data management, data science and information retrieval · 4 · 4 since 2021Human-computer interaction and ubiquitous computing · 2
YearPublicationVenuePosition
2025 Simulation-based Inference for High-dimensional Data using Surjective Sequential Neural Likelihood Estimation
abstract
Neural likelihood estimation methods for simulation-based inference can suffer from performance degradation when the modeled data is very high-dimensional or lies along a lower-dimensional manifold, which is due to the inability of the density estimator to accurately estimate a density function. We present Surjective Sequential Neural Likelihood (SSNL) estimation, a novel member in the family of methods for simulation-based inference (SBI). SSNL fits a dimensionality-reducing surjective normalizing flow model and uses it as a surrogate likelihood function, which allows for computational inference via Markov chain Monte Carlo or variational Bayes methods. Among other benefits, SSNL avoids the requirement to manually craft summary statistics for inference of high-dimensional data sets, since the lower-dimensional representation is computed simultaneously with learning the likelihood and without additional computational overhead. We evaluate SSNL on a wide variety of experiments, including two challenging real-world examples from the astrophysics and neuroscience literatures, and show that it either outperforms or is on par with state-of-the-art methods, making it an excellent off-the-shelf estimator for SBI for high-dimensional data sets.
Simon Dirmeier, Carlo Albert, Fernando Pérez-Cruz
UAI3
2025 Predictive structural assessment with Bayesian deep learning
abstract
Across Europe, bridges are reaching their planned service life, creating an urgent need for efficient structural assessments, as conventional methods prove time-consuming and costly. This research presents a data-driven framework for efficient structural pre-assessment of reinforced concrete frame bridges in the ultimate limit state. Developed based on the bridge inventory of the Swiss Federal Railways (SBB), a parametric, materially non-linear finite element analysis pipeline was created and harnessed to generate a simulation database. Bayesian Neural Networks were trained to predict structural code compliance factors with calibrated uncertainty estimates. The trained machine learning models enable a straightforward and fast pre-assessment of structural capacity. A real-world application on a railway underpass demonstrates the models’ ability to identify critical structures and provide decision support to asset owners and engineers. While the prototype focuses on railway frame bridges, the framework is transferable to other structure types, offering a scalable solution to prioritise maintenance interventions and optimise the allocation of limited resources.
Sophia V. Kuhn, Marius Weber, Antoine Binggeli, Michael A. Kraus, Fernando Pérez-Cruz, Walter Kaufmann
Adv. Eng. Informatics5
2025 AIXD: AI-eXtended Design Toolbox for data-driven and inverse design
abstract
Design processes, in many disciplines like architecture, civil engineering or mechanical engineering, involve navigating large, high-dimensional and heterogeneous data. While AI-driven approaches like inverse design and surrogate modeling can enhance design exploration, their adoption is hindered by complex workflows and the need for coding and machine learning expertise. To address this, we introduce AI-eXtended Design (AIXD): a low-code, open-source toolbox that integrates AI into computational design. AIXD simplifies handling of mixed data types, as well as the analysis, training, and deployment of machine learning models for inverse design, surrogate modeling, and sensitivity analysis, enabling domain experts to rapidly explore diverse solutions with minimal coding.In this paper, we show the functionalities of the toolbox, and we demonstrate AIXD’s capabilities in architectural and engineering design applications, showing how it accelerates performance evaluation, generates high-performing alternatives, and improves design understanding by delivering new insights. By bridging AI and design practice, AIXD lowers the entry barrier to data-driven methods, making AI-extended design more accessible and efficient.
Alessandro Maissen, Aleksandra Anna Apolinarska, Sophia V. Kuhn, Luis Salamanca, Michael A. Kraus, Konstantinos Tatsis 0001, Gonzalo Casas, Rafael Bischof, Romana Rust, Walter Kaufmann, Fernando Pérez-Cruz, Matthias Kohler
Comput. Aided Des.11
2025 A Closer Look at Benchmarking Self-supervised Pre-training with Image Classification
abstract
Self-supervised learning (SSL) is a machine learning approach where the data itself provides supervision, eliminating the need for external labels. The model is forced to learn about the data's inherent structure or context by solving a pretext task. With SSL, models can learn from abundant and cheap unlabeled data, significantly reducing the cost of training models where labels are expensive or inaccessible. In Computer Vision, SSL is widely used as pre-training followed by a downstream task, such as supervised transfer, few-shot learning on smaller labeled data sets, and/or unsupervised clustering. Unfortunately, it is infeasible to evaluate SSL methods on all possible downstream tasks and objectively measure the quality of the learned representation. Instead, SSL methods are evaluated using in-domain evaluation protocols, such as fine-tuning, linear probing, and k-nearest neighbors (kNN). However, it is not well understood how well these evaluation protocols estimate the representation quality of a pre-trained model for different downstream tasks under different conditions, such as dataset, metric, and model architecture. In this work, we study how classification-based evaluation protocols for SSL correlate and how well they predict downstream performance on different dataset types. Our study includes eleven common image datasets and 26 models that were pre-trained with different SSL methods or have different model backbones. We find that in-domain linear/kNN probing protocols are, on average, the best general predictors for out-of-domain performance. We further investigate the importance of batch normalization for the various protocols and evaluate how robust correlations are for different kinds of dataset domain shifts. In addition, we challenge assumptions about the relationship between discriminative and generative self-supervised methods, finding that most of their performance differences can be explained by changes to model backbones. Supplementary Information: The online version contains supplementary material available at 10.1007/s11263-025-02402-w.
Markus Marks, Manuel Knott 0001, Neehar Kondapaneni, Elijah Cole, Thijs Defraeye, Fernando Pérez-Cruz, Pietro Perona
Int. J. Comput. Vis.6
2025 Do You Trust Your Model? Emerging Malware Threats in the Deep Learning Ecosystem
abstract
Training high-quality deep learning models is a challenging task due to computational and technical requirements. A growing number of individuals, institutions, and companies increasingly rely on pre-trained, third-party models made available in public repositories. These models are often used directly or integrated in product pipelines with no particular precautions, since they are effectively just data in tensor form and considered safe. In this paper, we raise awareness of a new machine learning supply chain threat targeting neural networks. We introduce MaleficNet 2.0, a novel technique to embed self-extracting, self-executing malware in neural networks. MaleficNet 2.0 uses spread-spectrum channel coding combined with error correction techniques to inject malicious payloads in the parameters of deep neural networks. MaleficNet 2.0 injection technique is stealthy, does not degrade the performance of the model, and is robust against removal techniques. We design our approach to work both in traditional and distributed learning settings such as Federated Learning, and demonstrate that it is effective even when a reduced number of bits is used for the model parameters. Finally, we implement a proof-of-concept self-extracting neural network malware using MaleficNet 2.0, demonstrating the practicality of the attack against a widely adopted machine learning framework. Our aim with this work is to raise awareness against these new, dangerous attacks both in the research community and industry, and we hope to encourage further research in mitigation techniques against such threats.
Dorjan Hitaj, Giulio Pagnotta, Fabio De Gaspari, Sediola Ruko, Briland Hitaj, Luigi V. Mancini, Fernando Pérez-Cruz
IEEE Trans. Dependable Secur. Comput.7
2024 TATTOOED: A Robust Deep Neural Network Watermarking Scheme based on Spread-Spectrum Channel Coding
abstract
Deep Neural Networks (DNNs) trained on proprietary company data offer a competitive edge for the owning entity. However, these models can be attractive to competitors (or malicious entities), who can copy or clone these proprietary DNN models to use them to their advantage. Since these attacks are hard to prevent, it becomes imperative to have mechanisms in place that enable an affected entity to verify the ownership of its DNN models with very high confidence. Watermarking of deep neural networks has gained significant traction in recent years, with numerous (watermarking) strategies being proposed as mechanisms that can help verify the ownership of a DNN in scenarios where these models are obtained without the owner’s permission. However, a growing body of work has demonstrated that existing watermarking mechanisms are highly susceptible to removal techniques, such as fine-tuning, parameter pruning, or shuffling.In this paper, we build upon extensive prior work on covert (military) communication and propose TATTOOED, a novel DNN watermarking technique that is robust to existing threats. We demonstrate that using TATTOOED as their watermarking mechanism, the DNN owner can successfully obtain the watermark and verify model ownership even in scenarios where 99% of model parameters are altered. Furthermore, we show that TATTOOED is easy to employ in training pipelines and has negligible impact on model performance.
Giulio Pagnotta, Dorjan Hitaj, Briland Hitaj, Fernando Pérez-Cruz, Luigi V. Mancini
ACSAC4
2024 Signal domain adaptation network for limited-view optoacoustic tomography
abstract
Optoacoustic (OA) imaging is based on optical excitation of biological tissues with nanosecond-duration laser pulses and detection of ultrasound (US) waves generated by thermoelastic expansion following light absorption. The image quality and fidelity of OA images critically depend on the extent of tomographic coverage provided by the US detector arrays. However, full tomographic coverage is not always possible due to experimental constraints. One major challenge concerns an efficient integration between OA and pulse-echo US measurements using the same transducer array. A common approach toward the hybridization consists in using standard linear transducer arrays, which readily results in arc-type artifacts and distorted shapes in OA images due to the limited angular coverage. Deep learning methods have been proposed to mitigate limited-view artifacts in OA reconstructions by mapping artifactual to artifact-free (ground truth) images. However, acquisition of ground truth data with full angular coverage is not always possible, particularly when using handheld probes in a clinical setting. Deep learning methods operating in the image domain are then commonly based on networks trained on simulated data. This approach is yet incapable of transferring the learned features between two domains, which results in poor performance on experimental data. Here, we propose a signal domain adaptation network (SDAN) consisting of i) a domain adaptation network to reduce the domain gap between simulated and experimental signals and ii) a sides prediction network to complement the missing signals in limited-view OA datasets acquired from a human forearm by means of a handheld linear transducer array. The proposed method showed improved performance in reducing limited-view artifacts without the need for ground truth signals from full tomographic acquisitions.
Anna Klimovskaia Susmelj, Berkan Lafci, Firat Özdemir, Neda Davoudi, X. Luís Dean-Ben, Fernando Pérez-Cruz, Daniel Razansky
Medical Image Anal.6
2024 FedComm: Federated Learning as a Medium for Covert Communication
abstract
Proposed as a solution to mitigate the privacy implications related to the adoption of deep learning, Federated Learning (FL) enables large numbers of participants to successfully train deep neural networks without revealing theactualprivate training data. To date, a substantial amount of research has investigated the security and privacy properties of FL, resulting in a plethora of innovative attack and defense strategies. This paper thoroughly investigates the communication capabilities of an FL scheme. In particular, we show that a party involved in the FL learning process can use FL as a covert communication medium to send an arbitrary message. We introduce FedComm, a novel covert-communication technique that enables robust sharing and transfer of targeted payloads within the FL framework. Our extensive theoretical and empirical evaluations show that FedComm provides a stealthy communication channel, with minimal disruptions to the training process. Our experiments show that FedComm successfully delivers 100% of a payload in the order of kilobits before the FL procedure converges. Our evaluation also shows that FedComm is independent of the application domain and the neural network architecture used by the underlying FL scheme.
Dorjan Hitaj, Giulio Pagnotta, Briland Hitaj, Fernando Pérez-Cruz, Luigi V. Mancini
IEEE Trans. Dependable Secur. Comput.4
2023 PassGPT: Password Modeling and (Guided) Generation with Large Language Models
Javier Rando, Fernando Pérez-Cruz, Briland Hitaj
ESORICS (4)2
2023 Adaptive Annealed Importance Sampling with Constant Rate Progress
abstract
Annealed Importance Sampling (AIS) synthesizes weighted samples from an intractable distribution given its unnormalized density function. This algorithm relies on a sequence of interpolating distributions bridging the target to an initial tractable distribution such as the well-known geometric mean path of unnormalized distributions which is assumed to be suboptimal in general. In this paper, we prove that the geometric annealing corresponds to the distribution path that minimizes the KL divergence between the current particle distribution and the desired target when the feasible change in the particle distribution is constrained. Following this observation, we derive the constant rate discretization schedule for this annealing sequence, which adjusts the schedule to the difficulty of moving samples between the initial and the target distributions. We further extend our results to $f$-divergences and present the respective dynamics of annealing sequences based on which we propose the Constant Rate AIS (CR-AIS) algorithm and its efficient implementation for $\alpha$-divergences. We empirically show that CR-AIS performs well on multiple benchmark distributions while avoiding the computationally expensive tuning loop in existing Adaptive AIS.
Shirin Goshtasbpour, Victor Cohen, Fernando Pérez-Cruz
ICML3
2023 Renku: a platform for sustainable data science
abstract
Data and code working together is fundamental to machine learning (ML), but the context around datasets and interactions between datasets and code are in general captured only rudimentarily. Context such as how the dataset was prepared and created, what source data were used, what code was used in processing, how the dataset evolved, and where it has been used and reused can provide much insight, but this information is often poorly documented. That is unfortunate since it makes datasets into black-boxes with potentially hidden characteristics that have downstream consequences. We argue that making dataset preparation more accessible and dataset usage easier to record and document would have significant benefits for the ML community: it would allow for greater diversity in datasets by inviting modification to published sources, simplify use of alternative datasets and, in doing so, make results more transparent and robust, while allowing for all contributions to be adequately credited. We present a platform, Renku, designed to support and encourage such sustainable development and use of data, datasets, and code, and we demonstrate its benefits through a few illustrative projects which span the spectrum from dataset creation to dataset consumption and showcasing.
Rok Roskar, Chandrasekhar Ramakrishnan, Michele Volpi, Fernando Pérez-Cruz, Lilian Gasser, Firat Özdemir, Patrick Paitz, Mohammad Alisafaee, Philipp Fischer 0004, Ralf Grubenmann, Eliza J. Harris, Tasko Olevski, Carl Remlinger, Luis Salamanca, Elisabet Capon Garcia, Lorenzo Cavazzi, Jakub Chrobasik, Darlin Cordoba Osnas, Alessandro Degano, Jimena Dupre, Wesley Johnson, Eike Kettner, Laura Kinkead, Sean D. Murphy, Flora Thiebaut, Olivier Verscheure
NeurIPS4
2023 Anchor Data Augmentation
abstract
We propose a novel algorithm for data augmentation in nonlinear over-parametrized regression. Our data augmentation algorithm borrows from the literature on causality. Contrary to the current state-of-the-art solutions that rely on modifications of Mixup algorithm, we extend the recently proposed distributionally robust Anchor regression (AR) method for data augmentation. Our Anchor Data Augmentation (ADA) uses several replicas of the modified samples in AR to provide more training examples, leading to more robust regression predictions. We apply ADA to linear and nonlinear regression problems using neural networks. ADA is competitive with state-of-the-art C-Mixup solutions.
Nora Schneider, Shirin Goshtasbpour, Fernando Pérez-Cruz
NeurIPS3
2023 Enhancing diversity in GANs via non-uniform sampling
Pablo Sánchez-Martín, Pablo M. Olmos, Fernando Pérez-Cruz
Inf. Sci.3
2023 Regularizing transformers with deep probabilistic layers
Aurora Cobo Aguilera, Pablo M. Olmos, Antonio Artés-Rodríguez, Fernando Pérez-Cruz
Neural Networks4
2022 MaleficNet: Hiding Malware into Deep Neural Networks Using Spread-Spectrum Channel Coding
Dorjan Hitaj, Giulio Pagnotta, Briland Hitaj, Luigi V. Mancini, Fernando Pérez-Cruz
ESORICS (3)5
2022 Vision paper: causal inference for interpretable and robust machine learning in mobility analysis
abstract
Artificial intelligence (AI) is revolutionizing many areas of our lives, leading a new era of technological advancement. Particularly, the transportation sector would benefit from the progress in AI and advance the development of intelligent transportation systems. Building intelligent transportation systems requires an intricate combination of artificial intelligence and mobility analysis. The past few years have seen rapid development in transportation applications using advanced deep neural networks. However, such deep neural networks are difficult to interpret and lack robustness, which slows the deployment of these AI-powered algorithms in practice. To improve their usability, increasing research efforts have been devoted to developing interpretable and robust machine learning methods, among which the causal inference approach recently gained traction as it provides interpretable and actionable information. Moreover, most of these methods are developed for image or sequential data which do not satisfy specific requirements of mobility data analysis. This vision paper emphasizes research challenges in deep learning-based mobility analysis that require interpretability and robustness, summarizes recent developments in using causal inference for improving the interpretability and robustness of machine learning methods, and highlights opportunities in developing causally-enabled machine learning models tailored for mobility analysis. This research direction will make AI in the transportation sector more interpretable and reliable, thus contributing to safer, more efficient, and more sustainable future transportation systems.
Yanan Xin 0001, Natasa Tagasovska, Fernando Pérez-Cruz, Martin Raubal
SIGSPATIAL/GIS3
2022 What You See is What You Classify: Black Box Attributions
abstract
An important step towards explaining deep image classifiers lies in the identification of image regions that contribute to individual class scores in the model's output. However, doing this accurately is a difficult task due to the black-box nature of such networks. Most existing approaches find such attributions either using activations and gradients or by repeatedly perturbing the input. We instead address this challenge by training a second deep network, the Explainer, to predict attributions for a pre-trained black-box classifier, the Explanandum. These attributions are provided in the form of masks that only show the classifier-relevant parts of an image, masking out the rest. Our approach produces sharper and more boundary-precise masks when compared to the saliency maps generated by other methods. Moreover, unlike most existing approaches, ours is capable of directly generating very distinct class-specific masks in a single forward pass. This makes the proposed method very efficient during inference. We show that our attributions are superior to established methods both visually and quantitatively with respect to the PASCAL VOC-2007 and Microsoft COCO-2014 datasets.
Steven Stalder, Nathanaël Perraudin, Radhakrishna Achanta, Fernando Pérez-Cruz, Michele Volpi
NeurIPS4
2022 Optimization of Annealed Importance Sampling Hyperparameters
abstract
Abstract Annealed Importance Sampling (AIS) is a popular algorithm used to estimates the intractable marginal likelihood of deep generative models. Although AIS is guaranteed to provide unbiased estimate for any set of hyperparameters, the common implementations rely on simple heuristics such as the geometric average bridging distributions between initial and the target distribution which affect the estimation performance when the computation budget is limited. In order to reduce the number of sampling iterations, we present a parameteric AIS process with flexible intermediary distributions defined by a residual density with respect to the geometric mean path. Our method allows parameter sharing between annealing distributions, the use of fix linear schedule for discretization and amortization of hyperparameter selection in latent variable models. We assess the performance of Optimized-Path AIS for marginal likelihood estimation of deep generative models and compare it to compare it to more computationally intensive AIS.
Shirin Goshtasbpour, Fernando Pérez-Cruz
ECML/PKDD (5)2
2022 Deep Reinforcement Learning for Random Access in Machine-Type Communication
abstract
Random access (RA) schemes are a topic of high interest in machine-type communication (MTC). In RA protocols, backoff techniques such as exponential backoff (EB) are used to stabilize the system to avoid low throughput and excessive delays. However, these backoff techniques show varying performance for different underlying assumptions and analytical models. Therefore, finding a better transmission policy for slotted ALOHA RA is still a challenge. In this paper, we show the potential of deep reinforcement learning (DRL) for RA. We learn a transmission policy that balances between throughput and fairness. The proposed algorithm learns transmission probabilities using previous action and binary feedback signal, and it is adaptive to different traffic arrival rates. Moreover, we propose average age of packet (AoP) as a metric to measure fairness among users. Our results show that the proposed policy outperforms the baseline EB transmission schemes in terms of throughput and fairness.
Muhammad Awais Jadoon, Adriano Pastore, Mònica Navarro, Fernando Pérez-Cruz
WCNC4
2019 PassGAN: A Deep Learning Approach for Password Guessing
Briland Hitaj, Paolo Gasti, Giuseppe Ateniese, Fernando Pérez-Cruz
ACNS4
2019 Probabilistic Time of Arrival Localization
abstract
In this letter, we take a new approach for time of arrival geo-localization. We show that the main sources of error in metropolitan areas are due to environmental imperfections that bias our solutions, and that we can rely on a probabilistic model to learn and compensate for them. The resulting localization error is validated using measurements from a live LTE cellular network to be less than 10 meters, representing an order-of-magnitude improvement.
Fernando Pérez-Cruz, Pablo M. Olmos, Michael Minyi Zhang, Howard Huang
IEEE Signal Process. Lett.1
2018 Sparse Three-Parameter Restricted Indian Buffet Process for Understanding International Trade
abstract
This paper presents a Bayesian nonparametric latent feature model specially suitable for exploratory analysis of high-dimensional count data. We perform a non-negative doubly sparse matrix factorization that has two main advantages: not only we are able to better approximate the row input distributions, but the inferred topics are also easier to interpret. By combining the three-parameter and restricted Indian buffet processes into a single prior, we increase the model flexibility, allowing for a full spectrum of sparse solutions in the latent space. We demonstrate the usefulness of our approach in the analysis of countries' economic structure. Compared to other approaches, empirical results show our model's ability to give easy-to-interpret information and better capture the underlying sparsity structure of data.
Melanie F. Pradier, Viktor Stojkoski, Zoran Utkovski, Lujupco Kocorev, Fernando Pérez-Cruz
ICASSP5
2018 Complex Gaussian Processes for Regression
abstract
In this paper, we propose a novel Bayesian solution for nonlinear regression in complex fields. Previous solutions for kernels methods usually assume a complexification approach, where the real-valued kernel is replaced by a complex-valued one. This approach is limited. Based on the results in complex-valued linear theory and Gaussian random processes, we show that a pseudo-kernel must be included. This is the starting point to develop the new complex-valued formulation for Gaussian process for regression (CGPR). We face the design of the covariance and pseudo-covariance based on a convolution approach and for several scenarios. Just in the particular case where the outputs are proper, the pseudo-kernel cancels. Also, the hyperparameters of the covariance can be learned maximizing the marginal likelihood using Wirtinger's calculus and patterned complex-valued matrix derivatives. In the experiments included, we show how CGPR successfully solves systems where the real and imaginary parts are correlated. Besides, we successfully solve the nonlinear channel equalization problem by developing a recursive solution with basis removal. We report remarkable improvements compared to previous solutions: a 2-4-dB reduction of the mean squared error with just a quarter of the training samples used by previous approaches.
Rafael Boloix-Tortosa, Juan José Murillo-Fuentes, F. Javier Payan-Somet, Fernando Pérez-Cruz
IEEE Trans. Neural Networks Learn. Syst.4
2017 Deep Models Under the GAN: Information Leakage from Collaborative Deep Learning
abstract
Deep Learning has recently become hugely popular in machine learning for its ability to solve end-to-end learning systems, in which the features and the classifiers are learned simultaneously, providing significant improvements in classification accuracy in the presence of highly-structured and large databases.
Briland Hitaj, Giuseppe Ateniese, Fernando Pérez-Cruz
CCS3
2017 Wireless RSSI fingerprinting localization
Simon Yiu, Marzieh Dashti, Holger Claussen 0001, Fernando Pérez-Cruz
Signal Process.4
2016 Locating user equipments and access points using RSSI fingerprints: A Gaussian process approach
abstract
Location fingerprinting (LF) is an attractive localization technique which relies on existing infrastructures. The major drawback of LF is the requirement of having an updated fingerprint database. Gaussian Process (GP) is a non-parametric modeling technique which can be used to model the received signal strength indicator (RSSI) and create the fingerprint database based on few training data. In this paper we use a parametric pathloss model for the GP mean and a flexible non-parametric covariance function, so we can get reliable estimates with low fingerprinting effort. In our experiment, we show that with 23 fingerprint locations we perform as well as traditional fingerprinting with over 230 fingerprinted locations for an office space of 2500m2.
Simon Yiu, Marzieh Dashti, Holger Claussen 0001, Fernando Pérez-Cruz
ICC4
2016 Infinite Continuous Feature Model for Psychiatric Comorbidity Analysis
abstract
We aim at finding the comorbidity patterns of substance abuse, mood and personality disorders using the diagnoses from the National Epidemiologic Survey on Alcohol and Related Conditions database. To this end, we propose a novel Bayesian nonparametric latent feature model for categorical observations, based on the Indian buffet process, in which the latent variables can take values between 0 and 1. The proposed model has several interesting features for modeling psychiatric disorders. First, the latent features might be off, which allows distinguishing between the subjects who suffer a condition and those who do not. Second, the active latent features take positive values, which allows modeling the extent to which the patient has that condition. We also develop a new Markov chain Monte Carlo inference algorithm for our model that makes use of a nested expectation propagation procedure.
Isabel Valera, Francisco J. R. Ruiz, Pablo M. Olmos, Carlos Blanco 0002, Fernando Pérez-Cruz
Neural Comput.5
2016 Infinite Factorial Unbounded-State Hidden Markov Model
abstract
There are many scenarios in artificial intelligence, signal processing or medicine, in which a temporal sequence consists of several unknown overlapping independent causes, and we are interested in accurately recovering those canonical causes. Factorial hidden Markov models (FHMMs) present the versatility to provide a good fit to these scenarios. However, in some scenarios, the number of causes or the number of states of the FHMM cannot be known or limited a priori. In this paper, we propose an infinite factorial unbounded-state hidden Markov model (IFUHMM), in which the number of parallel hidden Markovmodels (HMMs) and states in each HMM are potentially unbounded. We rely on a Bayesian nonparametric (BNP) prior over integer-valued matrices, in which the columns represent the Markov chains, the rows the time indexes, and the integers the state for each chain and time instant. First, we extend the existent infinite factorial binary-state HMM to allow for any number of states. Then, we modify this model to allow for an unbounded number of states and derive an MCMC-based inference algorithm that properly deals with the trade-off between the unbounded number of states and chains. We illustrate the performance of our proposed models in the power disaggregation problem.
Isabel Valera, Francisco J. R. Ruiz, Fernando Pérez-Cruz
IEEE Trans. Pattern Anal. Mach. Intell.3
2015 Infinite Factorial Dynamical Model
abstract
We propose the infinite factorial dynamic model (iFDM), a general Bayesian nonparametric model for source separation. Our model builds on the Markov Indian buffet process to consider a potentially unbounded number of hidden Markov chains (sources) that evolve independently according to some dynamics, in which the state space can be either discrete or continuous. For posterior inference, we develop an algorithm based on particle Gibbs with ancestor sampling that can be efficiently applied to a wide range of source separation problems. We evaluate the performance of our iFDM on four well-known applications: multitarget tracking, cocktail party, power disaggregation, and multiuser detection. Our experimental results show that our approach for source separation does not only outperform previous approaches, but it can also handle problems that were computationally intractable for existing approaches.
Isabel Valera, Francisco J. R. Ruiz, Lennart Svensson, Fernando Pérez-Cruz
NIPS4
2015 Bayesian nonparametric crowdsourcing
Pablo G. Moreno, Antonio Artés-Rodríguez, Yee Whye Teh, Fernando Pérez-Cruz
J. Mach. Learn. Res.4
2014 Improved performance of LDPC-coded MIMO systems with EP-based soft-decisions
abstract
Modern communications systems use efficient encoding schemes, multiple-input multiple-output (MIMO) and high-order QAM constellations for maximizing spectral efficiency. However, as the dimensions of the system grow, the design of efficient and low-complexity MIMO receivers possesses technical challenges. Symbol detection can no longer rely on conventional approaches for posterior probability computation due to complexity. Marginalization of this posterior to obtain per-antenna soft-bit probabilities to be fed to a channel decoder is computationally challenging when realistic signaling is used. In this work, we propose to use Expectation Propagation (EP) algorithm to provide an accurate low-complexity Gaussian approximation to the posterior, easily solving the posterior marginalization problem. EP soft-bit probabilities are used in an LDPC-coded MIMO system, achieving outstanding performance improvement compared to similar approaches in the literature for low-complexity LDPC MIMO decoding.
Javier Cespedes, Pablo M. Olmos, Matilde Sánchez Fernández, Fernando Pérez-Cruz
ISIT4
2014 New information-estimation results for poisson, binomial and negative binomial models
abstract
In recent years, a number of mathematical relationships have been established between information measures and estimation measures for various models, including Gaussian, Poisson and binomial models. In this paper, it is shown that the second derivative of the input-output mutual information with respect to the input scaling can be expressed as the expectation of a certain Bregman divergence pertaining to the conditional expectations of the input and the input power. This result is similar to that found for the Gaussian model where the Bregman divergence therein is the square distance. In addition, the Poisson, binomial and negative binomial models are shown to be similar in the small scaling regime in the sense that the derivative of the mutual information and the derivative of the relative entropy converge to the same value.
Camilo G. Taborda, Fernando Pérez-Cruz, Dongning Guo
ISIT2
2014 Bayesian nonparametric comorbidity analysis of psychiatric disorders
Francisco J. R. Ruiz, Isabel Valera, Carlos Blanco 0002, Fernando Pérez-Cruz
J. Mach. Learn. Res.4
2014 Expectation Propagation Detection for High-Order High-Dimensional MIMO Systems
abstract
Modern communications systems use multiple-input multiple-output (MIMO) and high-order QAM constellations for maximizing spectral efficiency. However, as the number of antennas and the order of the constellation grow, the design of efficient and low-complexity MIMO receivers possesses big technical challenges. For example, symbol detection can no longer rely on maximum likelihood detection or sphere-decoding methods, as their complexity increases exponentially with the number of transmitters/receivers. In this paper, we propose a low-complexity high-accuracy MIMO symbol detector based on the Expectation Propagation (EP) algorithm. EP allows approximating iteratively at polynomial-time the posterior distribution of the transmitted symbols. We also show that our EP MIMO detector outperforms classic and state-of-the-art solutions reducing the symbol error rate at a reduced computational complexity.
Javier Cespedes, Pablo M. Olmos, Matilde Sánchez Fernández, Fernando Pérez-Cruz
IEEE Trans. Commun.4
2014 Information-Estimation Relationships Over Binomial and Negative Binomial Models
abstract
In recent years, a number of new connections between information measures and estimation have been found under various models, including, predominantly, Gaussian and Poisson models. This paper develops similar results for the binomial and negative binomial models. In particular, it is shown that the derivative of the relative entropy and the derivative of the mutual information for the binomial and negative binomial models can be expressed through the expectation of closed-form expressions that have conditional estimates as the main argument. Under mild conditions, those derivatives take the form of an expected Bregman divergence.
Camilo G. Taborda, Dongning Guo, Fernando Pérez-Cruz
IEEE Trans. Inf. Theory3
2014 An Automated Screening System for Tuberculosis
abstract
Automated screening systems are commonly used to detect some agent in a sample and take a global decision about the subject (e.g., ill/healthy) based on these detections. We propose a Bayesian methodology for taking decisions in (sequential) screening systems that considers the false alarm rate of the detector. Our approach assesses the quality of its decisions and provides lower bounds on the achievable performance of the screening system from the training data. In addition, we develop a complete screening system for sputum smears in tuberculosis diagnosis, and show, using a real-world database, the advantages of the proposed framework when compared to the commonly used count detections and threshold approach.
Ricardo Santiago-Mozos, Fernando Pérez-Cruz, Michael G. Madden, Antonio Artés-Rodríguez
IEEE J. Biomed. Health Informatics2
2013 Improving the BP estimate over the AWGN channel using Tree-structured expectation propagation
abstract
In this paper, we propose the tree-structured expectation propagation (TEP) algorithm for low-density parity-check (LDPC) decoding over the binary additive white Gaussian noise (BI-AWGN) channel. By approximating the posterior distribution by a tree-structure factorization, the TEP has been proven to improve belief propagation (BP) decoding over the binary erasure channel (BEC). We show for the AWGN channel how the TEP decoder is also able to capture additional information disregarded by the BP solution, which leads to a noticeable reduction of the error rate for finite-length codes. We show that for the range of codes of interest, the TEP gain is obtained with a slight increase in complexity over that of the BP algorithm. An efficient way of constructing the tree-like structure is also described.
Luis Salamanca, Juan José Murillo-Fuentes, Pablo M. Olmos, Fernando Pérez-Cruz
ISIT4
2013 Tree Expectation Propagation for ML Decoding of LDPC Codes over the BEC
abstract
We propose a decoding algorithm for LDPC codes that achieves the maximum likelihood (ML) solution over the binary erasure channel (BEC). In this channel, the tree-structured expectation propagation (TEP) decoder improves the peeling decoder (PD) by processing check nodes of degree one and two. However, it does not achieve the ML solution, as the tree structure of the TEP allows only for approximate inference. In this paper, we provide the procedure to construct the structure needed for exact inference. This algorithm, denoted as generalized tree-structured expectation propagation (GTEP), modifies the code graph by recursively eliminating any check node and merging this information in the remaining graph. The GTEP decoder upon completion either provides the unique ML solution or a tree graph in which the number of parent nodes indicates the multiplicity of the ML solution. We also explain the algorithm as a Gaussian elimination method, relating the GTEP to other ML solutions. Compared to previous approaches, it presents an equivalent complexity, it exhibits a simpler graphical message-passing procedure and, most interesting, the algorithm can be generalized to other channels.
Luis Salamanca, Pablo M. Olmos, Juan José Murillo-Fuentes, Fernando Pérez-Cruz
IEEE Trans. Commun.4
2013 Tree-Structured Expectation Propagation for LDPC Decoding over BMS Channels
abstract
In this paper, we put forward the tree-structured expectation propagation (TEP) algorithm for decoding block and convolutional low-density parity-check codes over any binary channel. We have already shown that TEP improves belief propagation (BP) over the binary erasure channel (BEC) by imposing marginal constraints over a set of pairs of variables that form a tree or a forest. The TEP decoder is a message-passing algorithm that sequentially builds a tree/forest of erased variables to capture additional information disregarded by the standard BP decoder, which leads to a noticeable reduction of the error rate for finite-length codes. In this paper, we show how the TEP can be extended to any channel, specifically to binary memoryless symmetric (BMS) channels. We particularly focus on how the TEP algorithm can be adapted for any channel model and, more importantly, how to choose the tree/forest to keep the gains observed for block and convolutional LDPC codes over the BEC.
Luis Salamanca, Pablo M. Olmos, Fernando Pérez-Cruz, Juan José Murillo-Fuentes
IEEE Trans. Commun.3
2013 Tree-Structure Expectation Propagation for LDPC Decoding Over the BEC
abstract
We present the tree-structure expectation propagation (Tree-EP) algorithm to decode low-density parity-check (LDPC) codes over discrete memoryless channels (DMCs). Expectation propagation generalizes belief propagation (BP) in two ways. First, it can be used with any exponential family distribution over the cliques in the graph. Second, it can impose additional constraints on the marginal distributions. We use this second property to impose pairwise marginal constraints over pairs of variables connected to a check node of the LDPC code's Tanner graph. Thanks to these additional constraints, the Tree-EP marginal estimates for each variable in the graph are more accurate than those provided by BP. We also reformulate the Tree-EP algorithm for the binary erasure channel (BEC) as a peeling-type algorithm (TEP) and we show that the algorithm has the same computational complexity as BP and it decodes a higher fraction of errors. We describe the TEP decoding process by a set of differential equations that represents the expected residual graph evolution as a function of the code parameters. The solution of these equations is used to predict the TEP decoder performance in both the asymptotic regime and the finite-length regimes over the BEC. While the asymptotic threshold of the TEP decoder is the same as the BP decoder for regular and optimized codes, we propose a scaling law for finite-length LDPC codes, which accurately approximates the TEP improved performance and facilitates its optimization.
Pablo M. Olmos, Juan José Murillo-Fuentes, Fernando Pérez-Cruz
IEEE Trans. Inf. Theory3
2012 Finite-length analysis of the TEP decoder for LDPC ensembles over the BEC
abstract
In this work, we analyze the finite-length performance of low-density parity check (LDPC) ensembles decoded over the binary erasure channel (BEC) using the tree-expectation propagation (TEP) algorithm. In a previous paper, we showed that the TEP improves the BP performance for decoding regular and irregular short LDPC codes, but the perspective was mainly empirical. In this work, given the degree-distribution of an LDPC ensemble, we explain and predict the range of code lengths for which the TEP improves the BP solution. In addition, for LDPC ensembles that present a single critical point, we propose a scaling law to accurately predict the performance in the waterfall region. These results are of critical importance to design practical LDPC codes for the TEP decoder.
Pablo M. Olmos, Fernando Pérez-Cruz, Luis Salamanca, Juan José Murillo-Fuentes
ISIT2
2012 Mutual information and relative entropy over the Binomial and Negative Binomial channels
abstract
We study the relation of the mutual information and relative entropy over the Binomial and Negative Binomial channels with estimation theoretical quantities, in which we extend already known results for Gaussian and Poisson channels. We establish general expressions for these information theory concepts with a direct connection with estimation theory through the conditional mean estimation and a particular loss function.
Camilo G. Taborda, Fernando Pérez-Cruz
ISIT2
2012 Finite-length performance of spatially-coupled LDPC codes under TEP decoding
abstract
Spatially-coupled (SC) LDPC codes are constructed from a set of L regular sparse codes of length M. In the asymptotic limit of these parameters, SC codes present an excellent decoding threshold under belief propagation (BP) decoding, close to the maximum a posteriori (MAP) threshold of the underlying regular code. In the finite-length regime, we need both dimensions, L and M, to be sufficiently large, yielding a very large code length and decoding latency. In this paper, and for the erasure channel, we show that the finite-length performance of SC codes is improved if we consider the tree-structured expectation propagation (TEP) algorithm in the decoding stage. When applied to the decoding of SC LDPC codes, it allows using shorter codes to achieve similar error rates. We also propose a window-sliding scheme for the TEP decoder to reduce the decoding latency.
Pablo M. Olmos, Fernando Pérez-Cruz, Luis Salamanca, Juan José Murillo-Fuentes
ITW2
2012 Derivative of the relative entropy over the poisson and Binomial channel
abstract
In this paper it is found that, regardless of the statistics of the input, the derivative of the relative entropy over the Binomial channel can be seen as the expectation of a function that has as argument the mean of the conditional distribution that models the channel. Based on this relationship we formulate a similar expression for the mutual information concept. In addition to this, using the connection between the Binomial and Poisson distribution we develop similar results for the Poisson channel. Novelty of the results presented here lies on the fact that, expressions obtained can be applied to a wide range of scenarios.
Camilo G. Taborda, Fernando Pérez-Cruz
ITW2
2012 Bayesian Nonparametric Modeling of Suicide Attempts
abstract
The National Epidemiologic Survey on Alcohol and Related Conditions (NESARC) database contains a large amount of information, regarding the way of life, medical conditions, depression, etc., of a representative sample of the U.S. population. In the present paper, we are interested in seeking the hidden causes behind the suicide attempts, for which we propose to model the subjects using a nonparametric latent model based on the Indian Buffet Process (IBP). Due to the nature of the data, we need to adapt the observation model for discrete random variables. We propose a generative model in which the observations are drawn from a multinomial-logit distribution given the IBP matrix. The implementation of an efficient Gibbs sampler is accomplished using the Laplace approximation, which allows us to integrate out the weighting factors of the multinomial-logit likelihood model. Finally, the experiments over the NESARC database show that our model properly captures some of the hidden causes that model suicide attempts.
Francisco J. R. Ruiz, Isabel Valera, Carlos Blanco 0002, Fernando Pérez-Cruz
NIPS4
2011 When to add another dimension when communicating over MIMO channels
abstract
This paper introduces a divide and conquer approach to the design of transmit and receive filters for communication over a Multiple Input Multiple Output (MIMO) Gaussian channel subject to an average power constraint. It involves conversion to a set of parallel scalar channels, possibly with very different gains, followed by coding per sub-channel (i.e. over time) rather than coding across sub-channels (i.e. over time and space). The loss in performance is negligible at high signal-to-noise ratio (SNR) and not significant at medium SNR. The advantages are reduction in signal processing complexity and greater insight into the SNR thresholds at which a channel is first allocated power. This insight is a consequence of formulating the optimal power allocation in terms of an upper bound on error rate that is determined by parameters of the input lattice such as the minimum distance and kissing number. The resulting thresholds are given explicitly in terms of these lattice parameters. By contrast, when the optimization problem is phrased in terms of maximizing mutual information, the solution is mercury waterfilling, and the thresholds are implicit.
Sreechakra Goparaju, A. Robert Calderbank, William R. Carson, Miguel R. D. Rodrigues, Fernando Pérez-Cruz
ICASSP5
2011 Capacity achieving LDPC ensembles for the TEP decoder in erasure channels
abstract
In this work we address the design of degree distributions (DD) of low-density parity-check (LDPC) codes for the tree-expectation propagation (TEP) decoder. The optimization problem to find distributions to maximize the TEP decoding threshold for a fixed-rate code can not be analytically solved. We derive a simplified optimization problem that can be easily solved since it is based in the analytic expressions of the peeling decoder. Two kinds of solutions are obtained from this problem: we either design LDPC ensembles for which the BP threshold equals the MAP threshold or we get LDPC ensembles for which the TEP threshold outperforms the BP threshold, even achieving the MAP capacity in some cases. Hence, we proved that there exist ensembles for which the MAP solution can be obtained with linear complexity even though the BP threshold does not achieve the MAP threshold.
Pablo M. Olmos, Juan José Murillo-Fuentes, Fernando Pérez-Cruz
ISIT3
2011 Zero-error codes for the noisy-typewriter channel
abstract
In this paper, we propose nontrivial codes that achieve a non-zero zero-error rate for several odd-letter noisy-typewriter channels. Some of these codes (specifically, those which are defined for a number of letters of the channel of the form 2n+ 1) achieve the best-known lower bound on the zero-error capacity. We build the codes using linear codes over rings, as we do not require the multiplicative inverse to build the codes.
Francisco J. R. Ruiz, Fernando Pérez-Cruz
ITW2
2011 MAP decoding for LDPC codes over the binary erasure channel
abstract
In this paper, we propose a decoding algorithm for LDPC codes that achieves the MAP solution over the BEC. This algorithm, denoted as generalized tree-structured expectation propagation (GTEP), extends the idea of our previous work, the TEP decoder. The GTEP modifies the graph by eliminating a check node of any degree and merging this information with the remaining graph. The GTEP decoder upon completion either provides the unique MAP solution or a tree graph in which the number of parent nodes indicates the multiplicity of the MAP solution. This algorithm can be easily described for the BEC, and it can be cast as a generalized peeling decoder. The GTEP naturally optimizes the complexity of the decoder, by looking for checks nodes of minimum degree to be eliminated first.
Luis Salamanca, Pablo M. Olmos, Juan José Murillo-Fuentes, Fernando Pérez-Cruz
ITW4
2011 An Application of Tree-Structured Expectation Propagation for Channel Decoding
abstract
We show an application of a tree structure for approximate inference in graphical models using the expectation propagation algorithm. These approximations are typically used over graphs with short-range cycles. We demonstrate that these approximations also help in sparse graphs with long-range loops, as the ones used in coding theory to approach channel capacity. For asymptotically large sparse graph, the expectation propagation algorithm together with the tree structure yields a completely disconnected approximation to the graphical model but, for for finite-length practical sparse graphs, the tree structure approximation to the code graph provides accurate estimates for the marginal of each variable.
Pablo M. Olmos, Luis Salamanca, Juan José Murillo-Fuentes, Fernando Pérez-Cruz
NIPS4
2011 Multioutput Support Vector Regression for Remote Sensing Biophysical Parameter Estimation
abstract
This letter proposes a multioutput support vector regression (M-SVR) method for the simultaneous estimation of different biophysical parameters from remote sensing images. General retrieval problems require multioutput (and potentially nonlinear) regression methods. M-SVR extends the single-output SVR to multiple outputs maintaining the advantages of a sparse and compact solution by using an$\varepsilon$-insensitive cost function. The proposed M-SVR is evaluated in the estimation of chlorophyll content, leaf area index and fractional vegetation cover from a hyperspectral compact high-resolution imaging spectrometer images. The achieved improvement with respect to the single-output regression approach suggests that M-SVR can be considered a convenient alternative for nonparametric biophysical parameter estimation and model inversion.
Devis Tuia, Jochem Verrelst, Luis Alonso 0002, Fernando Pérez-Cruz, Gustau Camps-Valls
IEEE Geosci. Remote. Sens. Lett.4
2011 Extended Input Space Support Vector Machine
abstract
In some applications, the probability of error of a given classifier is too high for its practical application, but we are allowed to gather more independent test samples from the same class to reduce the probability of error of the final decision. From the point of view of hypothesis testing, the solution is given by the Neyman-Pearson lemma. However, there is no equivalent result to the Neyman-Pearson lemma when the likelihoods are unknown, and we are given a training dataset. In this brief, we explore two alternatives. First, we combine the soft (probabilistic) outputs of a given classifier to produce a consensus labeling for K test samples. In the second approach, we build a new classifier that directly computes the label for K test samples. For this second approach, we need to define an extended input space training set and incorporate the known symmetries in the classifier. This latter approach gives more accurate results, as it only requires an accurate classification boundary, while the former needs an accurate posterior probability estimate for the whole input space. We illustrate our results with well-known databases.
Ricardo Santiago-Mozos, Fernando Pérez-Cruz, Antonio Artés-Rodríguez
IEEE Trans. Neural Networks2
2010 Tree-structure expectation propagation for decoding LDPC codes over binary erasure channels
abstract
Expectation Propagation is a generalization to Belief Propagation (BP) in two ways. First, it can be used with any exponential family distribution over the cliques in the graph. Second, it can impose additional constraints on the marginal distributions. We use this second property to impose pair-wise marginal distribution constraints in some check nodes of the LDPC Tanner graph. These additional constraints allow decoding the received codeword when the BP decoder gets stuck. In this paper, we first present the new decoding algorithm, whose complexity is identical to the BP decoder, and we then prove that it is able to decode codewords with a larger fraction of erasures, as the block size tends to infinity. The proposed algorithm can be also understood as a simplification of the Maxwell decoder, but without its computational complexity. We also illustrate that the new algorithm outperforms the BP decoder for finite block-size codes.
Pablo M. Olmos, Juan José Murillo-Fuentes, Fernando Pérez-Cruz
ISIT3
2010 Channel decoding with a Bayesian equalizer
abstract
Low-density parity-check (LPDC) decoders assume the channel estate information (CSI) is known and they have the true a posteriori probability (APP) for each transmitted bit. But in most cases of interest, the CSI needs to be estimated with the help of a short training sequence and the LDPC decoder has to decode the received word using faulty APP estimates. In this paper, we study the uncertainty in the CSI estimate and how it affects the bit error rate (BER) output by the LDPC decoder. To improve these APP estimates, we propose a Bayesian equalizer that takes into consideration not only the uncertainty due to the noise in the channel, but also the uncertainty in the CSI estimate, reducing the BER after the LDPC decoder.
Luis Salamanca, Juan José Murillo-Fuentes, Fernando Pérez-Cruz
ISIT3
2010 Robust and Low Complexity Distributed Kernel Least Squares Learning in Sensor Networks
abstract
We present a novel mechanism for consensus building in sensor networks. The proposed algorithm has three main properties that make it suitable for sensor network learning. First, the proposed algorithm is based on robust nonparametric statistics and thereby needs little prior knowledge about the network and the function that needs to be estimated. Second, the algorithm uses only local information about the network and it communicates only with nearby sensors. Third, the algorithm is completely asynchronous and robust. It does not need to coordinate the sensors to estimate the underlying function and it is not affected if other sensors in the network stop working. Therefore, the proposed algorithm is an ideal candidate for sensor networks deployed in remote and inaccessible areas, which might need to change their objective once they have been set up.
Fernando Pérez-Cruz, Sanjeev R. Kulkarni
IEEE Signal Process. Lett.1
2010 MIMO Gaussian channels with arbitrary inputs: optimal precoding and power allocation
abstract
In this paper, we investigate the linear precoding and power allocation policies that maximize the mutual information for general multiple-input-multiple-output (MIMO) Gaussian channels with arbitrary input distributions, by capitalizing on the relationship between mutual information and minimum mean-square error (MMSE). The optimal linear precoder satisfies a fixed-point equation as a function of the channel and the input constellation. For non-Gaussian inputs, a nondiagonal precoding matrix in general increases the information transmission rate, even for parallel noninteracting channels. Whenever precoding is precluded, the optimal power allocation policy also satisfies a fixed-point equation; we put forth a generalization of the mercury/waterfilling algorithm, previously proposed for parallel noninterfering channels, in which the mercury level accounts not only for the non-Gaussian input distributions, but also for the interference among inputs.
Fernando Pérez-Cruz, Miguel R. D. Rodrigues, Sergio Verdú
IEEE Trans. Inf. Theory1
2009 Optimized concatenated LDPC codes for joint source-channel coding
abstract
In this paper a scheme for joint source-channel coding based on low-density-parity-check (LDPC) codes is investigated. Two concatenated independent LDPC codes are used in the transmitter: one for source coding and the other for channel coding, with a joint belief propagation decoder. The asymptotic behavior is analyzed using EXtrinsic Information Transfer (EXIT) charts and this approximation is corroborated with illustrative experiments. The optimization of the degree distributions for our sparse code to maximize the information transmission rate is also considered.
Maria Fresia, Fernando Pérez-Cruz, H. Vincent Poor
ISIT2
2009 Distributed least square for consensus building in sensor networks
abstract
We present a novel mechanism for consensus building in sensor networks. The proposed algorithm has three main properties that make it suitable for general sensor-network learning. First, the proposed algorithm is based on robust nonparametric statistics and thereby needs little prior knowledge about the network and the function that needs to be estimated. Second, the algorithm uses only local information about the network and it communicates only with nearby sensors. Third, the algorithm is completely asynchronous and robust. It does not need to coordinate the sensors to estimate the underlying function and it is not affected if other sensors in the network stop working. Therefore, the proposed algorithm is an ideal candidate for sensor networks deployed in remote and inaccessible areas, which might need to change their objective once they have been set up.
Fernando Pérez-Cruz, Sanjeev R. Kulkarni
ISIT1
2009 Gaussian process regressors for multiuser detection in DS-CDMA systems
abstract
In this paper we present Gaussian processes for Regression (GPR) as a novel detector for CDMA digital communications. Particularly, we propose GPR for constructing analytical nonlinear multiuser detectors in CDMA communication systems. GPR can easily compute the parameters that describe its nonlinearities by maximum likelihood. Thereby, no cross-validation is needed, as it is typically used in nonlinear estimation procedures. The GPR solution is analytical, given its parameters, and it does not need to solve an optimization problem for building the nonlinear estimator. These properties provide fast and accurate learning, two major issues in digital communications. The GPR with a linear decision function can be understood as a regularized MMSE detector, in which the regularization parameter is optimally set. We also show the GPR receiver to be a straightforward nonlinear extension of the linear minimum mean square error (MMSE) criterion, widely used in the design of these receivers. We argue the benefits of this new approach in short codes CDMA systems where little information on the users' codes, users' amplitudes or the channel is available. The paper includes some experiments to show that GPR outperforms linear (MMSE) and nonlinear (SVM) state-ofthe- art solutions.
Juan José Murillo-Fuentes, Fernando Pérez-Cruz
IEEE Trans. Commun.2
2008 Optimal Precoding for Digital Subscriber Lines
abstract
We determine the linear precoding policy that maximizes the mutual information for general multiple-input multiple-output (MIMO) Gaussian channels with arbitrary input distributions, by capitalizing on the relationship between mutual information and minimum mean squared error (MMSE). The optimal linear precoder can be computed by means of a fixed- point equation as a function of the channel and the input constellation. We show that diagonalizing the channel matrix does not maximize the information transmission rate for nonGaussian inputs. A full precoding matrix may significantly increase the information transmission rate, even for parallel non-interacting channels. We illustrate the application of our results to typical Gigabit DSL systems.
Fernando Pérez-Cruz, Miguel R. D. Rodrigues, Sergio Verdú
ICC1
2008 Kullback-Leibler divergence estimation of continuous distributions
abstract
We present a method for estimating the KL divergence between continuous densities and we prove it converges almost surely. Divergence estimation is typically solved estimating the densities first. Our main result shows this intermediate step is unnecessary and that the divergence can be either estimated using the empirical cdf or k-nearest-neighbour density estimation, which does not converge to the true measure for finite k. The convergence proof is based on describing the statistics of our estimator using waiting-times distributions, as the exponential or Erlang. We illustrate the proposed estimators and show how they compare to existing methods based on density estimation, and we also outline how our divergence estimators can be used for solving the two-sample problem.
Fernando Pérez-Cruz
ISIT1
2008 Multiple-input multiple-output Gaussian channels: Optimal covariance for non-Gaussian inputs
abstract
We investigate the input covariance that maximizes the mutual information of deterministic multiple-input multipleo-utput (MIMO) Gaussian channels with arbitrary (not necessarily Gaussian) input distributions, by capitalizing on the relationship between the gradient of the mutual information and the minimum mean-squared error (MMSE) matrix. We show that the optimal input covariance satisfies a simple fixed-point equation involving key system quantities, including the MMSE matrix. We also specialize the form of the optimal input covariance to the asymptotic regimes of low and high snr. We demonstrate that in the low-snr regime the optimal covariance fully correlates the inputs to better combat noise. In contrast, in the high-snr regime the optimal covariance is diagonal with diagonal elements obeying the generalized mercury/waterfilling power allocation policy. Numerical results illustrate that covariance optimization may lead to significant gains with respect to conventional strategies based on channel diagonalization followed by mercury/waterfilling or waterfilling power allocation, particularly in the regimes of medium and high snr.
Miguel R. D. Rodrigues, Fernando Pérez-Cruz, Sergio Verdú
ITW2
2008 Estimation of Information Theoretic Measures for Continuous Random Variables
abstract
We analyze the estimation of information theoretic measures of continuous random variables such as: differential entropy, mutual information or Kullback-Leibler divergence. The objective of this paper is two-fold. First, we prove that the information theoretic measure estimates using the k-nearest-neighbor density estimation with fixed k converge almost surely, even though the k-nearest-neighbor density estimation with fixed k does not converge to its true measure. Second, we show that the information theoretic measure estimates do not converge for k growing linearly with the number of samples. Nevertheless, these nonconvergent estimates can be used for solving the two-sample problem and assessing if two random variables are independent. We show that the two-sample and independence tests based on these nonconvergent estimates compare favorably with the maximum mean discrepancy test and the Hilbert Schmidt independence criterion, respectively.
Fernando Pérez-Cruz
NIPS1
2007 Therapeutic Drug Monitoring of Kidney Transplant Recipients Using Profiled Support Vector Machines
abstract
This paper proposes a twofold approach for therapeutic drug monitoring (TDM) of kidney recipients using support vector machines (SVMs), for both predicting and detecting Cyclosporine A (CyA) blood concentrations. The final goal is to build useful, robust, and ultimately understandable models for individualizing the dosage of CyA. We compare SVMs with several neural network models, such as the multilayer perceptron (MLP), the Elman recurrent network, finite/infinite impulse response networks, and neural network ARMAX approaches. In addition, we present a profile-dependent SVM (PD-SVM), which incorporates a priori knowledge in both tasks. Models are compared numerically, statistically, and in the presence of additive noise. Data from 57 renal allograft recipients were used to develop the models. Patients followed a standard triple therapy, and CyA trough concentration was the dependent variable. The best results for the CyA blood concentration prediction were obtained using the PD-SVM (mean error of 0.36 ng/mL and root-mean-square error of 52.01 ng/mL in the validation set) and appeared to be more robust in the presence of additive noise. The proposed PD-SVM improved results from the standard SVM and MLP, specially significant (both numerical and statistically) in the one-against-all scheme. Finally, some clinical conclusions were obtained from sensitivity rankings of the models and distribution of support vectors. We conclude that the PD-SVM approach produces more accurate and robust models than do neural networks. Finally, a software tool for aiding medical decision-making including the prediction models is presented
Gustau Camps-Valls, Emilio Soria-Olivas, Juan José Pérez-Ruixo, Fernando Pérez-Cruz, Antonio Artés-Rodríguez, N. Víctor Jiménez
IEEE Trans. Syst. Man Cybern. Part C4
2006 Gaussian Processes for Digital Communications
abstract
We present Gaussian processes (GPs) for digital communications. GPs can be used to construct analytical nonlinear regression functions, which can be suitable for digital communications in which linear solutions under perform. GPs can be cast as nonlinear MMSE and its hyperparameters can be easily learnt by maximum likelihood. We present some experimental results regarding multi-user detection in CDMA systems and show the GPs outperform linear and nonlinear state-of-the-art solutions
Fernando Pérez-Cruz, Juan José Murillo-Fuentes
ICASSP (5)1
2005 Gaussian Processes for Multiuser Detection in CDMA receivers
abstract
In this paper we propose a new receiver for digital communications. We focus on the application of Gaussian Processes (GPs) to the multiuser detection (MUD) in code division multiple access (CDMA) systems to solve the near-far problem. Hence, we aim to reduce the interference from other users sharing the same frequency band. While usual approaches minimize the mean square error (MMSE) to linearly retrieve the user of interest, we exploit the same criteria but in the design of a nonlinear MUD. Since the optimal solution is known to be nonlinear, the performance of this novel method clearly improves that of the MMSE detectors. Furthermore, the GP based MUD achieves excellent interference suppression even for short training sequences. We also include some experiments to illustrate that other nonlinear detectors such as those based on Support Vector Machines (SVMs) exhibit a worse performance.
Juan José Murillo-Fuentes, Sebastian Caro, Fernando Pérez-Cruz
NIPS3
2005 Support Vector Regression for the simultaneous learning of a multivariate function and its derivatives
Marcelino Lázaro, Ignacio Santamaría, Fernando Pérez-Cruz, Antonio Artés-Rodríguez
Neurocomputing3
2005 Convergence of the IRWLS Procedure to the Support Vector Machine Solution
abstract
An iterative reweighted least squares (IRWLS) procedure recently proposed is shown to converge to the support vector machine solution. The convergence to a stationary point is ensured by modifying the original IRWLS procedure.
Fernando Pérez-Cruz, Carlos Bousoño-Calzón, Antonio Artés-Rodríguez
Neural Comput.1
2005 Learning a function and its derivative forcing the support vector expansion
abstract
In this paper, a new method for the simultaneous learning of a function and its derivative is presented. The method, setting out the problem inside of the Support Vector Machine (SVM) framework, relies on the kernel-based Support Vector expansion. The resultant optimization problem is solved by a computationally efficient Iterative Re-Weighted Least Squares (IRWLS) algorithm.
Marcelino Lázaro, Fernando Pérez-Cruz, Antonio Artés-Rodríguez
IEEE Signal Process. Lett.2
2004 Speeding up the IRWLS convergence to the SVM solution
abstract
We present the convergence demonstration of the iterative re-weighted least squares (IRWLS) procedure to the SVM solution, to propose two modifications, which significantly reduces the runtime complexity of the IRWLS. We show by means of computer experiments that the convergence can be speed up between two and eight times compare to the standard IRWLS procedure.
Fernando Pérez-Cruz, Antonio Artés-Rodríguez
IJCNN1
2004 Enhancing genetic feature selection through restricted search and Walsh analysis
abstract
In this paper, a twofold approach to improve the performance of genetic algorithms (GAs) in the feature selection problem (FSP) is presented. First, a novel genetic operator is introduced to solve the FSP. This operator fixes in each iteration the number of features to be selected among the available ones and consequently reduces the size of the search space. This approach yields two main advantages: a) training the learning machine becomes faster and b) a higher performance is achieved by using the selected subset. Second, we propose using the Walsh expansion of the FSP fitness function in order to perform ranking on the problem features. Ranking features have been traditionally considered to be a challenging problem, especially significant in health sciences where the number of available and potentially noisy signals is high. Three real biological datasets are used to test the behavior of the two approaches proposed.
Sancho Salcedo-Sanz, Gustau Camps-Valls, Fernando Pérez-Cruz, José Sepúlveda-Sanchis, Carlos Bousoño-Calzón
IEEE Trans. Syst. Man Cybern. Part C3
2003 Supervised-PCA and SVM Classifiers for Object Detection in Infrared Images
abstract
We tackle the problem of detecting sources of combustion in high definition multispectral medium wavelength infrared (MWIR) (3-5 /spl mu/m) images. We present a novel approach to this problem consisting of processing the images block-wise using a new technique that we call supervised principal component analysis (SPCA) to get the components of these blocks. This outperforms state-of-the-art methods with a significant reduction in the complexity of the whole scheme. As a classifier, we propose the use of a support vector machine (SVM) comparing the results from both its novelty-detection and binary non-linear versions. High performance is achieved from a small set of components.
Ricardo Santiago-Mozos, José M. Leiva-Murillo, Fernando Pérez-Cruz, Antonio Artés-Rodríguez
AVSS3
2003 Multi-class support vector machines: a new approach
abstract
We propose a new approach for solving multiclass problems with support vector machines. We modify the existing technique to properly reduce the empirical error, therefore we will be ideally able to outperform the previously proposed scheme for multi-class SVMs. The proposed approach also provides solutions with a significant reduction in the number of support vectors, which is an important feature for fast systems.
Jerónimo Arenas-García, Fernando Pérez-Cruz
ICASSP (2)2
2003 Kernel methods and their applications to signal processing
abstract
Recently introduced in machine learning, the notion of kernels has drawn a lot of interest as it allows nonlinear algorithms to be obtained from linear ones in a simple and elegant manner. This, in conjunction with the introduction of new linear classification methods such as the support vector machines has produced significant progress. The success of such algorithms is now spreading as they are applied to more and more domains. Many signal processing problems, by their nonlinear and high-dimensional nature, may benefit from such techniques. We give an overview of kernel methods and their recent applications.
Olivier Bousquet, Fernando Pérez-Cruz
ICASSP (4)2
2003 Feature selection and transduction for prediction of molecular bioactivity for drug design
abstract
Abstract Motivation: In drug discovery a key task is to identify characteristics that separate active (binding) compounds from inactive (non-binding) ones. An automated prediction system can help reduce resources necessary to carry out this task. Results: Two methods for prediction of molecular bioactivity for drug design are introduced and shown to perform well in a data set previously studied as part of the KDD (Knowledge Discovery and Data Mining) Cup 2001. The data is characterized by very few positive examples, a very large number of features (describing three-dimensional properties of the molecules) and rather different distributions between training and test data. Two techniques are introduced specifically to tackle these problems: a feature selection method for unbalanced data and a classifier which adapts to the distribution of the the unlabeled test data (a so-called transductive method). We show both techniques improve identification performance and in conjunction provide an improvement over using only one of the techniques. Our results suggest the importance of taking into account the characteristics in this data which may also be relevant in other problems of a similar type. Availability: Matlab source code is available at http://www.kyb.tuebingen.mpg.de/bs/people/weston/kdd/kdd.html Contact: [email protected] Supplementary information: Supplementary material is available at http://www.kyb.tuebingen.mpg.de/bs/people/weston/kdd/kdd.html. * To whom correspondence should be addressed.
Jason Weston, Fernando Pérez-Cruz, Olivier Bousquet, Olivier Chapelle, André Elisseeff, Bernhard Schölkopf
Bioinform.2
2003 Empirical risk minimization for support vector classifiers
abstract
In this paper, we propose a general technique for solving support vector classifiers (SVCs) for an arbitrary loss function, relying on the application of an iterative reweighted least squares (IRWLS) procedure. We further show that three properties of the SVC solution can be written as conditions over the loss function. This technique allows the implementation of the empirical risk minimization (ERM) inductive principle on large margin classifiers obtaining, at the same time, very compact (in terms of number of support vectors) solutions. The improvements obtained by changing the SVC loss function are illustrated with synthetic and real data examples.
Fernando Pérez-Cruz, Ángel Navia-Vázquez, Aníbal R. Figueiras-Vidal, Antonio Artés-Rodríguez
IEEE Trans. Neural Networks1
2002 Puncturing Multi-class Support Vector Machines
Fernando Pérez-Cruz, Antonio Artés-Rodríguez
ICANN1
2002 Multi-dimensional Function Approximation and Regression Estimation
Fernando Pérez-Cruz, Gustau Camps-Valls, Emilio Soria-Olivas, Juan José Pérez-Ruixo, Aníbal R. Figueiras-Vidal, Antonio Artés-Rodríguez
ICANN1
2002 Feature Selection via Genetic Optimization
Sancho Salcedo-Sanz, Mario de Prado-Cumplido, Fernando Pérez-Cruz, Carlos Bousoño-Calzón
ICANN3
2001 A new optimizing procedure for ν-support vector regressor
abstract
We present an approach to solve the v-SVR. It is based on an iterative re-weighted least squares (IRWLS) procedure, which is simple to implement and can be tuned to the usual /spl nu/-SVR solution. The IRWLS procedure is much more efficient (computational load) than quadratic programming techniques, which are usually employed to solve it.
Fernando Pérez-Cruz, Antonio Artés-Rodríguez
ICASSP1
2001 SVC-based equalizer for burst TDMA transmissions
Fernando Pérez-Cruz, Ángel Navia-Vázquez, Pedro Luis Alarcón-Diana, Antonio Artés-Rodríguez
Signal Process.1
2001 Weighted least squares training of support vector classifiers leading to compact and adaptive schemes
abstract
An iterative block training method for support vector classifiers (SVCs) based on weighted least squares (WLS) optimization is presented. The algorithm, which minimizes structural risk in the primal space, is applicable to both linear and nonlinear machines. In some nonlinear cases, it is necessary to previously find a projection of data onto an intermediate-dimensional space by means of either principal component analysis or clustering techniques. The proposed approach yields very compact machines, the complexity reduction with respect to the SVC solution is especially notable in problems with highly overlapped classes. Furthermore, the formulation in terms of WLS minimization makes the development of adaptive SVCs straightforward, opening up new fields of application for this type of model, mainly online processing of large amounts of (static/stationary) data, as well as online update in nonstationary scenarios (adaptive solutions). The performance of this new type of algorithm is analyzed by means of several simulations.
Ángel Navia-Vázquez, Fernando Pérez-Cruz, Antonio Artés-Rodríguez, Aníbal R. Figueiras-Vidal
IEEE Trans. Neural Networks2
2000 Support vector classifier with hyperbolic tangent penalty function
abstract
The support vector classifier is a new tool to solve classification problems, giving the classification boundary as a linear combination of the training samples. In non-separable problems with highly overlapped classes, the achieved classifiers are oversized. In this paper, we proposed to change the support vector classifier penalty function by an hyperbolic tangent one, obtaining as a result of the training phase a reduced support vector classifier with the same performance as the original one.
Fernando Pérez-Cruz, Ángel Navia-Vázquez, Pedro Luis Alarcón-Diana, Antonio Artés-Rodríguez
ICASSP1
2000 Fast Training of Support Vector Classifiers
abstract
In this communication we present a new algorithm for solving Support Vector Classifiers (SVC) with large training data sets. The new algorithm is based on an Iterative Re-Weighted Least Squares procedure which is used to optimize the SVc. Moreover, a novel sample selection strategy for the working set is presented, which randomly chooses the working set among the training samples that do not fulfill the stopping criteria. The validity of both proposals, the optimization procedure and sample selection strategy, is shown by means of computer experiments using well-known data sets.
Fernando Pérez-Cruz, Pedro Luis Alarcón-Diana, Ángel Navia-Vázquez, Antonio Artés-Rodríguez
NIPS1
1993 Feasibility analysis of channel equalizers using Kharitonov-type results
Fernando Pérez-Cruz, Domingo Docampo, Antonio Artés-Rodríguez, Chaouki T. Abdallah
ICASSP (3)1
1992 A deconvolution-based efficient method for generating the excitation in linear predictive speech coding
abstract
The authors provide an accurate scheme for generating the excitation sequence in multipulse speech coding. The method is based on spiky deconvolution techniques, which take advantage of the sparse character of the multipulse sequence. It is shown how a threshold deconvolution procedure, formerly developed by some of the authors, can be extended to deal with all-pole filters, and, consequently, how the multipulse sequence can be estimated on every frame of speech data. Quantization procedures for coding the multipulse sequence are also discussed in order to achieve a bit rate transmission lower than 6.5 kb/s.>
Domingo Docampo, Victoria Abreu-Sernández, Fernando Pérez-Cruz, Francisco González
ICASSP3