Ola Hössjer

dblp:43/3911 · also Ola G. Hössjer · DBLP profile ↗
← Back
5ranked-venue papers
3as first author
1since 2021 · last 2024
0000-0003-2767-8818ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Theory of computation · 4 · 3 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2024 An Information Theoretic Approach to Prevalence Estimation and Missing Data
abstract
Many data sources, including tracking social behavior to election polling to testing studies for understanding disease spread, are subject to sampling bias whose implications are not fully yet understood. In this paper we study estimation of a given feature (such as disease, or behavior at social media platforms) from biased samples, treating non-respondent individuals as missing data. Prevalence of the feature among sampled individuals has an upward bias under the assumption of individuals’ willingness to be sampled. This can be viewed as a regression model with symptoms as covariates and the feature as outcome. It is assumed that the outcome is unknown at the time of sampling, and therefore the missingness mechanism only depends on the covariates. We show that data, in spite of this, is missing at random only when the sizes of symptom classes in the population are known; otherwise data is missing not at random. With an information theoretic viewpoint, we show that sampling bias corresponds to external information due to individuals in the population knowing their covariates, and we quantify this external information by active information. The reduction in prevalence, when sampling bias is adjusted for, similarly translates into active information due to bias correction, with opposite sign to active information due to testing bias. We develop unified results that show that prevalence and active information estimates are asymptotically normal under all missing data mechanisms, when testing errors are absent and present respectively. The asymptotic behavior of the estimators is illustrated through simulations.
Ola Hössjer, Daniel Andrés Díaz-Pachón, Chen Zhao 0012, J. Sunil Rao
IEEE Trans. Inf. Theory1
2006 Methodological study of affine transformations of gene expression data with proposed robust non-parametric multi-dimensional normalization method
abstract
BACKGROUND: Low-level processing and normalization of microarray data are most important steps in microarray analysis, which have profound impact on downstream analysis. Multiple methods have been suggested to date, but it is not clear which is the best. It is therefore important to further study the different normalization methods in detail and the nature of microarray data in general. RESULTS: A methodological study of affine models for gene expression data is carried out. Focus is on two-channel comparative studies, but the findings generalize also to single- and multi-channel data. The discussion applies to spotted as well as in-situ synthesized microarray data. Existing normalization methods such as curve-fit ("lowess") normalization, parallel and perpendicular translation normalization, and quantile normalization, but also dye-swap normalization are revisited in the light of the affine model and their strengths and weaknesses are investigated in this context. As a direct result from this study, we propose a robust non-parametric multi-dimensional affine normalization method, which can be applied to any number of microarrays with any number of channels either individually or all at once. A high-quality cDNA microarray data set with spike-in controls is used to demonstrate the power of the affine model and the proposed normalization method. CONCLUSION: We find that an affine model can explain non-linear intensity-dependent systematic effects in observed log-ratios. Affine normalization removes such artifacts for non-differentially expressed genes and assures that symmetry between negative and positive log-ratios is obtained, which is fundamental when identifying differentially expressed genes. In addition, affine normalization makes the empirical distributions in different channels more equal, which is the purpose of quantile normalization, and may also explain why dye-swap normalization works or fails. All methods are made available in the aroma package, which is a platform-independent package for R.
Henrik Bengtsson, Ola Hössjer
BMC Bioinform.2
1997 Adaptive detection of known signals in additive noise by means of kernel density estimators
abstract
We consider the problem of detecting known signals contaminated by additive noise with a completely unknown probability density function f. To this end, we propose a new adaptive detection rule. It is defined by plugging a kernel density estimator f/spl circ/ of f into the maximum a posteriori (MAP) detector. The estimate f/spl circ/ can either be computed off-line from a training sequence or on-line simultaneously with the detection. For the off-line detector, we prove that the (asymptotic) error probability for weak signals converges to the minimal error probability of the MAP detector as the number of training data tends to infinity, and we also establish rates of convergence and the optimal choice of bandwidth order for a certain class of noise densities. In a Monte Carlo study, the off-line plug-in MAP detectors are compared with the L/sup 1/- and L/sup 2/-detectors for various noise distributions. When the training sequence is long enough, the plug-in detectors have excellent performance for a wide range of distributions, whereas the L/sup 2/-detector breaks down for heavy-tailed distributions and the L/sup 1/-detector for distributions with little mass around the origin.
Rolf T. Gustafsson, Ola Hössjer, Tommy Öberg
IEEE Trans. Inf. Theory2
1995 On-line density estimators with high efficiency
abstract
Presents on-line procedures for estimating density functions and their derivatives. At each step, M terms are updated. By increasing M the efficiency compared to the traditional off-line kernel density estimator tends to one. Already for M=2, it exceeds 99.1% for kernel orders and derivatives of practical interest.>
Ola Hössjer, Ulla Holst
IEEE Trans. Inf. Theory1
1993 Robust multiple classification of known signals in additive noise - An asymptotic weak signal approach
abstract
The problem of extracting one out of a finite number of possible signals of known form given observations in an additive noise model is considered. Two approaches are studied: either the signal with shortest distance to the observed data or the signal having maximal correlation with some transformation of the observed data is chosen. With a weak signal approach, the limiting error probability is a monotone function of the Pitman efficacy and it is the same for both the distance-based and correlation-based detectors. Using the minimax theory of Huber, it is possible to derive robust choices of distance/correlation when the limiting error probability is used as performance criterion. This generalizes previous work in the area, from two signals to an arbitrary number of signals. Considered are M-type and R-type distances and also one-dimensional and two-dimensional signals. Some Monte Carlo simulations are performed to compare the finite sample size error probabilities with the asymptotic error probabilities.>
Ola Hössjer, Moncef Mettiji
IEEE Trans. Inf. Theory1