Bart Kosko

dblp:78/5416 · DBLP profile ↗
← Back
76ranked-venue papers
26as first author
15since 2021 · last 2025
0000-0003-4745-8986ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 54 · 18 first-author · 12 since 2021Human-computer interaction and ubiquitous computing · 11 · 5 first-author · 3 since 2021Databases, data management, data science and information retrieval · 6 · 6 first-authorGraphics, computer vision, multimedia, augmented reality and games · 5Applied, interdisciplinary, general and emerging computing · 4 · 1 since 2021Systems, architecture and hardware · 1 · 1 first-authorComputer networks · 1
YearPublicationVenuePosition
2025 Causal Autoencoder-like Generation of Feedback Fuzzy Cognitive Maps with an LLM Agent
abstract
A large language model (LLM) can map a feedback causal fuzzy cognitive map (FCM) into text and then reconstruct the FCM from the text. This explainable AI system approximates an identity map from the FCM to itself and resembles the operation of an autoencoder (AE). Both the encoder and the decoder explain their decisions in contrast to black-box AEs. Humans can read and interpret the encoded text in contrast to the hidden variables and synaptic webs in AEs. The LLM agent approximates the identity map through a sequence of system instructions that does not compare the output to the input. The reconstruction is lossy because it removes weak causal edges or rules while it preserves strong causal edges. The encoder preserves the strong causal edges even when it trades off some details about the FCM to make the text sound more natural.
Akash Kumar Panda, Olaoluwa Adigun, Bart Kosko
ICMLA3
2024 Training Deep Neural Classifiers with Soft Diamond Regularizers
abstract
We introduce new soft diamond regularizers that both improve synaptic sparsity and maintain classification accuracy in deep neural networks. These parametrized regularizers outperform the state-of-the-art hard-diamond Laplacian regularizer of Lasso regression and classification. They use thick-tailed symmetric alpha-stable$(\mathcal{S}\alpha \mathcal{S})$bell-curve synaptic weight priors that are not Gaussian and so have thicker tails. The geometry of the diamond-shaped constraint set varies from a circle to a star depending on the tail thickness and dispersion of the prior probability density function. Training directly with these priors is computationally intensive because almost all$\mathcal{S}\alpha \mathcal{S}$probability densities lack a closed form. A precomputed lookup table removed this computational bottleneck. We tested the new soft diamond regularizers with deep neural classifiers on the three datasets CIFAR-10, CIFAR-100, and Caltech-256. The regularizers improved the accuracy of the classifiers. The improvements included 4.57% on CIFAR-10, 4.27% on CIFAR-100, and 6.69% on Caltech-256. They also outperformed$L_{2}$regularizers on all the test cases. Soft diamond regularizers also outperformed$L_{1}$lasso or Laplace regularizers because they better increased sparsity while improving classification accuracy. Soft-diamond priors substantially improved accuracy on CIFAR-10 when combined with dropout, batch, or data-augmentation regularization.
Olaoluwa Adigun, Bart Kosko
ICMLA2
2024 Bidirectional Variational Autoencoders
abstract
We present the new bidirectional variational autoencoder (BVAE) network architecture. The BVAE uses a single neural network both to encode and decode instead of an encoder-decoder network pair. The network encodes in the forward direction and decodes in the backward direction through the same synaptic web. Simulations compared BVAEs and ordinary VAEs on the four image tasks of image reconstruction, classification, interpolation, and generation. The image datasets included MNIST handwritten digits, Fashion-MNIST, CIFAR10, and CelebA-64 face images. The bidirectional structure of BVAEs cut the parameter count by almost 50% and still slightly outperformed the unidirectional VAEs.
Bart Kosko, Olaoluwa Adigun
IJCNN1
2023 Bidirectional Backpropagation Autoencoding Networks for Image Compression and Denoising
abstract
A bidirectional autoencoder learns or approximates an identity mapping as it trains a single network with a version of the new bidirectional backpropagation algorithm. Ordinary unidirectional autoencoders find many uses in image processing and in large language models. But they use separate networks for encoding and decoding. Bidirectional auto encoders use the same synaptic weights for encoding and decoding. The forward pass encodes while the backward pass decodes. Bidirectional auto encoders improved network performance and significantly reduced memory usage and used fewer parameters. Simulations compared unidirectional with bidirectional autoencoders for image compression and de noising. The models trained on the MNIST handwritten-digit and CIFAR-IO image datasets. The performance measures were the peak signal-to-noise ratio and the index of structural similarity. Bidirectional autoencoders outperformed unidirectional autoencoders and still reduced the number of trainable synaptic parameters by about 50%.
Olaoluwa Adigun, Bart Kosko
ICMLA2
2023 Hidden Priors for Bayesian Bidirectional Backpropagation
abstract
Non-uniform prior probabilities between hidden layers improved deep neural classifiers trained with bidirectional backpropagation. The resulting Bayesian bidirectional backpropagation algorithm jointly maximizes the forward and backward network likelihoods along with the weight priors. The backward direction exploits a hidden regression that ordinary unidirectional backpropagation ignores. Simulations compared Laplacian, Gaussian, Cauchy, and the new sinc-squared hidden priors on the CIFAR-10 and CIFAR-100 balanced image data sets. These hidden priors improved the classification accuracy of deep neural classifiers compared with default uniform priors and default unidirectional backpropagation. They did so at little extra computational cost. Sinc-squared and Cauchy multivariate priors often had the best classification accuracy. Cauchy hidden priors gave sparse hidden weights similar to the Laplacian priors associated with sparse lasso regression.
Olaoluwa Adigun, Bart Kosko
SMC2
2023 Noise-boosted recurrent backpropagation
Olaoluwa Adigun, Bart Kosko
Neurocomputing2
2022 Uniform Convergence of Probability Mixtures that Represent Combined Fuzzy Systems
abstract
We can combine expert knowledge by combining the probability mixtures that represent the if-then rules of the experts. Fuzzy rules define a generalized probability mixture whose moments describe a fuzzy system and its uncertainty. The mixture’s Bayesian structure gives a complete posterior probability description of the if-then fuzzy-set rules as they fire. A new theorem extends the uniform convergence of a fuzzy system’s mixture to the uniform convergence of the sequence of expert mixtures that represent any number of combined fuzzy systems as they each converge to a target function. A mixture of just two normal bell curves exactly represents the target function in the scalar case and serves as the probabilistic target of the converging mixture sequence. A sampled deep neural network can serve as the target function. Then the mixture defines a proxy system that gives a probabilistic form of explainable AI. The uniform convergence result extends to any continuous transformation of the converging fuzzy systems and further extends to the uniform mixture convergence of any continuous function of the combined systems and their continuous transformations.
Bart Kosko
FUZZ-IEEE1
2022 Deeper Bidirectional Neural Networks with Generalized Non-Vanishing Hidden Neurons
abstract
The new NoVa hidden neurons have outperformed ReLU hidden neurons in deep classifiers on some large image test sets. The NoVa or nonvanishing logistic neuron additively perturbs the sigmoidal activation function so that its derivative is not zero. This helps avoid or delay the problem of vanishing gradients. We here extend the NoVa to the generalized perturbed logistic neuron and compare it to ReLU and several other hidden neurons on large image test sets that include CIFAR-100 and Caltech-256. Generalized NoVa classifiers allow deeper networks with better classification on the large datasets. This deep benefit holds for ordinary unidirectional backpropagation. It also holds for the more efficient bidirectional backpropagation that trains in both the forward and backward directions.
Olaoluwa Adigun, Bart Kosko
ICMLA2
2022 Bayesian Rule Ontologies For XAI Classification and Regression
abstract
A random foam defines a modular rule-based ontology after sampling from a neural or other input-output system. A random foam combines several rule-based systems and averages the systems. It gives a Bayesian posterior over the subsystems. It also gives separate Bayesian posteriors over the rules of each subsystem. The shape of the rules controls how well the random-foam ontology performs in classification and regression. We found that a heterogenous ontology that mixes different rule shapes can perform better than a homogenous ontology based on a single Gaussian or other rule type. Random foams are also universal function approximators. So they can train on a neural black box and act as its explainable proxy system. We prove this uniform approximation theorem for the important case of bump-function random foams with throughput combination. Random foams also measure their output’s uncertainty through the conditional variance. Bump function rules performed better than Cauchy rules at classification while Cauchy rules performed better at regression. Gaussian rules performed best in both classification and regression. A homogeneous Gaussian random foam that trained on a 96.7% accurate neural classifier was itself 95.96% accurate on the MNIST data set. A heterogeneous random foam with two-thirds Gaussian rules and one-third Laplacian rules did better than did the all-Gaussian foam ontology.
Akash Kumar Panda, Bart Kosko
ICMLA2
2021 Bayesian Pruned Random Rule Foams for XAI
abstract
A random rule foam grows and combines several independent fuzzy rule-based systems by randomly sampling input-output data from a trained deep neural classifier. The random rule foam defines an interpretable proxy system for the sampled black-box classifier. The random foam gives the complete Bayesian posterior probabilities over the foam subsystems that contribute to the proxy system's output for a given pattern input. It also gives the Bayesian posterior over the if-then fuzzy rules in each of these constituent foams. The random foam also computes a conditional variance that describes the uncertainty in its predicted output given the random foam's learned rule structure. The mixture structure leads to bootstrap confidence intervals around the output. Using the Bayesian posterior probabilities to prune or discard low-probability sub-foams improves the system's classification accuracy. Simulations used the MNIST image data set of 60,000 gray-scale images of ten hand-written digits. Dropping the lowest-probability foams per input pattern brought the pruned random foam's classification accuracy nearly to that of the neural classifier. Posterior pruning outperformed simple accuracy pruning of a random foam and outperformed a random forest trained on the same neural classifier.
Akash Kumar Panda, Bart Kosko
FUZZ-IEEE2
2021 Bidirectional Backpropagation for High-Capacity Blocking Networks
abstract
The new bidirectional backpropagation algorithm helps blocking networks learn and recall large numbers of image patterns. Bidirectional backpropagation exploits backward-pass learning that ordinary unidirectional backpropagation ignores. The backward pass reveals a hidden regressor in classifiers since the input neurons are identity units. Blocking networks allow deep classifiers to learn and accurately recognize more patterns than the older classifiers that use softmax neurons at the output classification layer. Blocking networks use logistic neurons at the output layer of a block. They use random bipolar coding from the vertices of a hypercube rather than from the vertices of the simplex embedded in it as with l-in-K encoding. Bidirectional deep sweeps improved classification accuracy on the CIFAR-100 image data base and did so at little extra computational cost.
Olaoluwa Adigun, Bart Kosko
ICMLA2
2021 Deeper Neural Networks with Non-Vanishing Logistic Hidden Units: NoVa vs. ReLU Neurons
abstract
The new NoVa (nonvanishing) logistic neuron activation allows deeper neural networks because its derivative is positive. So it helps mitigate the problem of vanishing gradients in deep networks. Deep neural classifiers with NoVa hidden units had better classification accuracy on the CFAR-10, CFAR-100, and Caltech-256 image databases compared with threshold-linear ReLU hidden units. Still simpler identity hidden units also outperformed ReLU hidden units in deep classifiers but usually had less classification accuracy than NoVa networks. NoVa hidden neurons also outperformed ReLU hidden neurons in deep convolutional neural networks.
Olaoluwa Adigun, Bart Kosko
ICMLA2
2021 Bayesian Bidirectional Backpropagation Learning
abstract
We show that training neural classifiers with Bayesian bidirectional backpropagation improves the performance of the network. Bidirectional backpropagation trains a deep network for both forward and backward recall through the same layers of neurons and with the same weights. It maximizes the network's joint forward and backward likelihood. Bayesian bidirectional backpropagation combines prior probabilities at the input and output layers with the likelihood structure of the layers. It maximizes the posterior probability of the network. It differs from other forms of neural Bayesian estimation because it uses the bidirectional likelihood of the network instead of the unidirectional likelihood. Bayesian bidirectional backpropagation outperformed classifiers trained with both unidirectional and bidirectional backpropagation. The networks trained on the CIFAR-10 and CIFAR-100 image test sets. A Laplacian or Lasso-like prior outperformed both Gaussian and uniform priors.
Olaoluwa Adigun, Bart Kosko
IJCNN2
2021 Bidirectional Associative Memories: Unsupervised Hebbian Learning to Bidirectional Backpropagation
abstract
Bidirectional associative memories (BAMs) pass neural signals forward and backward through the same web of synapses. Earlier BAMs had no hidden neurons and did not use supervised learning. They tuned their synaptic weights with unsupervised Hebbian or competitive learning. Two-layer feedback BAMs always converge to fixed-point equilibria for threshold or threshold-like neurons. Every rectangular connection matrix is bidirectionally stable. These simpler BAMs extend to arbitrary hidden layers with supervised learning if the resulting bidirectional backpropagation algorithm uses the proper layer likelihood in the forward and backward directions. Bidirectional backpropagation lets users run deep classifiers and regressors in reverse as well as forward. Bidirectional training exploits pattern and synaptic information that forward-only running ignores.
Bart Kosko
IEEE Trans. Syst. Man Cybern. Syst.1
2021 Erratum to "Bidirectional Associative Memories: Unsupervised Hebbian Learning to Bidirectional Backpropagation"
abstract
The energy velocityequation (41)in[1]has two missing prime symbols “’” that indicate the derivative of the primed activation function with respect to its argument. The correct equation for the time derivative of the ABAM energy function$E$should read as follows:
Bart Kosko
IEEE Trans. Syst. Man Cybern. Syst.1
2020 Convergence of Generalized Probability Mixtures That Describe Adaptive Fuzzy Rule-based Systems
abstract
A generalized probability mixture density governs the rule-base structure of an additive fuzzy system. A new theorem allows such a mixture to absorb any bounded real function by mixing two normal densities. We further show that a sequence of adaptive fuzzy systems defines a corresponding sequence of uniformly convergent generalized Gaussian mixtures. The result applies to any uniformly convergent sequence of function approximators. It gives a practical way to define approximating mixtures with adaptive fuzzy systems or neural networks. Users can combine any number of these rule-based systems by mixing their generalized mixtures.
Bart Kosko
FUZZ-IEEE1
2020 High Capacity Neural Block Classifiers with Logistic Neurons and Random Coding
abstract
We show that neural networks with logistic output neurons and random codewords can store and classify far more patterns than those that use softmax neurons and 1-in-K encoding. Logistic neurons can choose binary codewords from an exponentially large set of codewords. Random coding picks the binary or bipolar codewords for training such deep classifier models. This method searched for the bipolar codewords that minimized the mean of an inter-codeword similarity measure. The method used blocks of networks with logistic input and output layers and with few hidden layers. Adding such blocks gave deeper networks and reduced the problem of vanishing gradients. It also improved learning because the input and output neurons of an interior block must equal the input pattern's code word. Deep-sweep training of the neural blocks further improved the classification accuracy. The networks trained on the CIFAR-100 and the Caltech-256 image datasets. Networks with 40 output logistic neurons and random coding achieved much of the accuracy of 100 softmax neurons on the CIFAR- 100 patterns. Sufficiently deep random-coded networks with just 80 or more logistic output neurons had better accuracy on the Caltech-256 dataset than did deep networks with 256 softmax output neurons.
Olaoluwa Adigun, Bart Kosko
IJCNN2
2020 Noise can speed backpropagation learning and deep bidirectional pretraining
Bart Kosko, Kartik Audhkhasi, Osonde Osoba
Neural Networks1
2020 Bidirectional Backpropagation
abstract
We extend backpropagation (BP) learning from ordinary unidirectional training to bidirectional training of deep multilayer neural networks. This gives a form of backward chaining or inverse inference from an observed network output to a candidate input that produced the output. The trained network learns a bidirectional mapping and can apply to some inverse problems. A bidirectional multilayer neural network can exactly represent some invertible functions. We prove that a fixed three-layer network can always exactly represent any finite permutation function and its inverse. The forward pass computes the permutation function value. The backward pass computes the inverse permutation with the same weights and hidden neurons. A joint forward-backward error function allows BP learning in both directions without overwriting learning in either direction. The learning applies to classification and regression. The algorithms do not require that the underlying sampled function has an inverse. A trained regression network tends to map an output back to the centroid of its preimage set.
Olaoluwa Adigun, Bart Kosko
IEEE Trans. Syst. Man Cybern. Syst.2
2019 Noise-boosted bidirectional backpropagation and adversarial learning
Olaoluwa Adigun, Bart Kosko
Neural Networks2
2018 Training Generative Adversarial Networks with Bidirectional Backpropagation
abstract
Training generative adversarial networks with the new bidirectional backpropagation algorithm improved performance compared with ordinary unidirectional backpropagation. Bidirectional backpropagation trains a multilayer neural network in the backward direction as well as in the forward direction over the same weights and neurons. The result approximates a set-level inverse mapping that tends to improve the learning of the forward classification mapping. We compared bidirectional backpropagation training of the discriminator with unidirectional training for the standard vanilla GAN on MNIST data and a deep convolutional GAN on CIFAR-10 image data. We also compared B-BP and unidirectional training for a Wasserstein GAN on both MNIST and CIFAR-10 data. Bidirectional training substantially improved the inception score of the vanilla GAN's generated digit images for MNIST data. It increased the vanilla GAN's inception score by 22.3% and greatly reduced the GAN's incidence of mode collapse. Bidirectional training improved the inception score of the deep-convolutional GAN's generated samples by 3.3% on the CIFAR-10 data set. Bidirectional training also increased the Wasserstein GAN's inception score by 4.4% on the MNIST data and by 10.0% on the CIFAR-10 image data.
Olaoluwa Adigun, Bart Kosko
ICMLA2
2018 Additive Fuzzy Systems: From Generalized Mixtures to Rule Continua
abstract
A generalized probability mixture density governs an additive fuzzy system. The fuzzy system's if-then rules correspond to the mixed probability densities. An additive fuzzy system computes an output by adding its fired rules and then averaging the result. The mixture's convex structure yields Bayes theorems that give the probability of which rules fired or which combined fuzzy systems fired for a given input and output. The convex structure also results in new moment theorems and learning laws and new ways to both approximate functions and exactly represent them. The additive fuzzy system itself is just the first conditional moment of the generalized mixture density. The output is a convex combination of the centroids of the fired then-part sets. The mixture's second moment defines the fuzzy system's conditional variance. It describes the inherent uncertainty in the fuzzy system's output due to rule interpolation. The mixture structure gives a natural way to combine fuzzy systems because mixing mixtures yields a new mixture. A separation theorem shows how fuzzy approximators combine with exact Watkins-based two-rule function representations in a higher-level convex sum of the combined systems. Two mixed Gaussian densities with appropriate Watkins coefficients define a generalized mixture density such that the fuzzy system's output equals any given real-valued function if the function is bounded and not constant. Statistical hill-climbing algorithms can learn the generalized mixture from sample data. The mixture structure also extends finite rule bases to continuum-many rules. Finite fuzzy systems suffer from exponential rule explosion because each input fires all their graph-cover rules. The continuum system fires only a special random sample of rules based on Monte Carlo sampling from the system's mixture. Users can program the system by changing its wave-like meta-rules based on the location and shape of the mixed densities in the mixture. Such meta-rules can help mitigate rule explosion. The meta-rules grow only linearly with the number of mixed densities even though the underlying fuzzy if-then rules can have high-dimensional if-part and then-part fuzzy sets.
Bart Kosko
Int. J. Intell. Syst.1
2017 Using noise to speed up video classification with recurrent backpropagation
abstract
Carefully injected noise can speed the convergence and accuracy of video classification with recurrent backpropagation (RBP). This noise-boost uses the recent results that backpropagation is a special case of the generalized expectation maximization (EM) algorithm and that careful noise injection can always speed the average convergence of the EM algorithm to a local maximum of the log-likelihood surface. We extend this result to the time-varying case of recurrent backpropagation and prove sufficient noise-benefit conditions for both classification and regression. Injecting noise that satisfies the noisy-EM positivity condition (NEM noise) speeds up RBP training. The classification simulations used eleven categories of sports videos based on standard UCF YouTube sports-action video clips. Training RBP with NEM noise in just the output neurons led to 60% fewer iterations in training as compared with noiseless training. This corresponded to a 20.6% maximum decrease in training cross entropy. NEM noise injection also outperformed simple blind noise injection: RBP training with NEM noise gave a 15.6% maximum decrease in training cross entropy compared with RBP training with blind noise. Injecting NEM noise also improved the relative classification accuracy by 5% over noiseless RBP training. NEM noise improved the classification accuracy from 81% to 83% on the test set.
Olaoluwa Adigun, Bart Kosko
IJCNN2
2017 Generalized mixture representations and combinations for additive fuzzy systems
abstract
A generalized probability mixture density governs all additive fuzzy systems. These systems sum fired if-then rules to compute an output. Their mixture structure leads to new Bayes theorems and new ways to combine or fuse fuzzy systems. Additive fuzzy rule-base systems can uniformly approximate continuous functions while their Watkins representations can exactly represent bounded real vector functions with just two rules per vector component. A new separation theorem shows how to combine both such fuzzy systems into a common rule base. A Bayes theorem specifies which rules or which of the combined fuzzy systems fired to produce an observed output from a given input. We prove that two mixed Gaussian densities with Watkins coefficients define a mixture density whose first moment equals any bounded real function. A new learning law can tune the system's mixture density with training data. The additive fuzzy system's finite rule base passes over to a rule continuum. Monte Carlo sampling can then compute fuzzy-system outputs by mixture-based sampling from the virtual rule continuum.
Bart Kosko
IJCNN1
2016 Noise-enhanced convolutional neural networks
Kartik Audhkhasi, Osonde Osoba, Bart Kosko
Neural Networks3
2013 Noise benefits in backpropagation and deep bidirectional pre-training
abstract
We prove that noise can speed convergence in the backpropagation algorithm. The proof consists of two separate results. The first result proves that the backpropagation algorithm is a special case of the generalized Expectation-Maximization (EM) algorithm for iterative maximum likelihood estimation. The second result uses the recent EM noise benefit to derive a sufficient condition for backpropagation training. The noise adds directly to the training data. A noise benefit also applies to the deep bidirectional pre-training of the neural network as well as to the backpropagation training of the network. The geometry of the noise benefit depends on the probability structure of the neurons at each layer. Logistic sigmoidal neurons produce a forbidden noise region that lies below a hyperplane. Then all noise on or above the hyperplane can only speed convergence of the neural network. The forbidden noise region is a sphere if the neurons have a Gaussian signal or activation function. These noise benefits all follow from the general noise benefit of the EM algorithm. Monte Carlo sample means estimate the population expectations in the EM algorithm. We demonstrate the noise benefits using MNIST digit classification.
Kartik Audhkhasi, Osonde Osoba, Bart Kosko
IJCNN3
2013 Noisy hidden Markov models for speech recognition
abstract
We show that noise can speed training in hidden Markov models (HMMs). The new Noisy Expectation-Maximization (NEM) algorithm shows how to inject noise when learning the maximum-likelihood estimate of the HMM parameters because the underlying Baum-Welch training algorithm is a special case of the Expectation-Maximization (EM) algorithm. The NEM theorem gives a sufficient condition for such an average noise boost. The condition is a simple quadratic constraint on the noise when the HMM uses a Gaussian mixture model at each state. Simulations show that a noisy HMM converges faster than a noiseless HMM on the TIMIT data set.
Kartik Audhkhasi, Osonde Osoba, Bart Kosko
IJCNN3
2013 Noise-enhanced clustering and competitive learning algorithms
Osonde Osoba, Bart Kosko
Neural Networks2
2013 Corrigendum to "Noise enhanced clustering and competitive learning algorithms" [Neural Networks 37 (2013) 132-140]
Osonde Osoba, Bart Kosko
Neural Networks2
2011 Triply fuzzy function approximation for Bayesian inference
abstract
We prove that independent fuzzy systems can uniformly approximate Bayesian posterior probability density functions by approximating prior and likelihood probability densities as well as hyperprior probability densities that underly priors. This triply fuzzy function approximation extends the recent theorem for uniformly approximating the posterior density by approximating just the prior and likelihood densities. This allows users to state priors and hyper-priors in words or rules as well as to adapt them from sample data. A fuzzy system with just two rules can exactly represent common closed-form probability densities so long as they are bounded. The function approximators can also be neural networks or any other type of uniform function approximator.
Osonde Osoba, Sanya Mitaim, Bart Kosko
IJCNN3
2011 Noise benefits in the expectation-maximization algorithm: Nem theorems and models
abstract
We prove a general sufficient condition for a noise benefit in the expectation-maximization (EM) algorithm. Additive noise speeds the average convergence of the EM algorithm to a local maximum of the likelihood surface when the noise condition holds. The sufficient condition states when additive noise makes the signal more probable on average. The performance measure is Kullback relative entropy. A Gaussian-mixture problem demonstrates the EM noise benefit. Corollary results give other special cases when noise improves performance in the EM algorithm.
Osonde Osoba, Sanya Mitaim, Bart Kosko
IJCNN3
2011 Bayesian Inference With Adaptive Fuzzy Priors and Likelihoods
abstract
Fuzzy rule-based systems can approximate prior and likelihood probabilities in Bayesian inference and thereby approximate posterior probabilities. This fuzzy approximation technique allows users to apply a much wider and more flexible range of prior and likelihood probability density functions than found in most Bayesian inference schemes. The technique does not restrict the user to the few known closed-form conjugacy relations between the prior and likelihood. It allows the user in many cases to describe the densities with words and just two rules can absorb any bounded closed-form probability density directly into the rulebase. Learning algorithms can tune the expert rules as well as grow them from sample data. The learning laws and fuzzy approximators have a tractable form because of the convex-sum structure of additive fuzzy systems. This convex-sum structure carries over to the fuzzy posterior approximator. We prove a uniform approximation theorem for Bayesian posteriors: An additive fuzzy posterior uniformly approximates the posterior probability density if the prior or likelihood densities are continuous and bounded and if separate additive fuzzy systems approximate the prior and likelihood densities. Simulations demonstrate this fuzzy approximation of priors and posteriors for the three most common conjugate priors (as when a beta prior combines with a binomial likelihood to give a beta posterior). Adaptive fuzzy systems can also approximate non-conjugate priors and likelihoods as well as approximate hyperpriors in hierarchical Bayesian inference. The number of fuzzy rules can grow exponentially in iterative Bayesian inference if the previous posterior approximator becomes the new prior approximator.
Osonde Osoba, Sanya Mitaim, Bart Kosko
IEEE Trans. Syst. Man Cybern. Part B3
2010 Optimal Mean-Square Noise Benefits in Quantizer-Array Linear Estimation
abstract
A new theorem shows that additive quantizer noise decreases the mean-squared error of threshold-array optimal and suboptimal linear estimators. The initial rate of this noise benefit improves as the number of threshold sensors or quantizers increases. The array sums the outputs of identical binary quantizers that receive the same random input signal. The theorem further shows that zero-symmetric uniform quantizer noise gives the fastest initial decrease in mean-squared error among all finite-variance zero-symmetric scale-family noise. These results apply to all bounded continuous signal densities and all zero-symmetric scale-family quantizer noise with finite variance.
Ashok Patel, Bart Kosko
IEEE Signal Process. Lett.2
2009 Quantizer noise benefits in nonlinear signal detection with alpha-stable channel noise
abstract
Two new theorems show how deliberately adding quantizer noise can improve statistical signal detection in array-based nonlinear correlation detection even in the case of infinite-variance alpha-stable channel noise. The first theorem gives a necessary and sufficient condition for such quantizer noise to increase the detection probability for a fixed false-alarm probability. The second theorem shows that the array must contain more than one quantizer for a stochastic-resonance noise benefit and that the noise benefit improves in the small-quantizer noise limit as the number of array quantizers increases. It further shows that symmetric uniform quantizer noise gives the optimal noise benefit among all symmetric scale-family noise types.
Ashok Patel, Bart Kosko
ICASSP2
2009 Adaptive fuzzy priors for Bayesian inference
abstract
A fuzzy rule-based system can model prior probabilities in Bayesian inference and thereby approximate posterior probabilities. This fuzzy technique allows users to express prior descriptions in words rather than as closed-form probability density functions. Learning algorithms can tune the expert rules as well as grow them from sample data. The learning laws and closed-form approximations have a tractable form because of the convex-sum structure of additive fuzzy systems. Simulations demonstrate the fuzzy approximation of priors and posteriors for the three most common conjugate priors. An approximate beta prior combines with binomial data to give a new approximate beta posterior. An approximate gamma prior combines with Poisson data to give a new approximate gamma posterior. An approximate normal prior combines with normal data to give a new approximate normal posterior.
Osonde Osoba, Sanya Mitaim, Bart Kosko
IJCNN3
2009 Neural signal-detection noise benefits based on error probability
abstract
We present several necessary and sufficient conditions and a learning algorithm for noise benefits in threshold neural signal detection using error probabilities. The first condition ensures noise benefits in threshold detection of discrete binary signals and applies to noise types from scale families. The condition also gives an easy way to compute optimal noise values for closed-form scale-family noise densities. A related condition ensures noise benefits in threshold detection of signals that have absolutely continuous distributions. This condition reduces to a simple weighted-derivative comparison of the signal densities at the detection threshold when the signal densities are continuously differentiable and when the additive noise is either zero-mean discrete bipolar or finite-variance symmetric scale-family noise. A gradient-ascent learning algorithm can find the optimal noise value for thick-tailed stable densities and many other noise probability densities that do not have a closed form.
Ashok Patel, Bart Kosko
IJCNN2
2009 Error-probability noise benefits in threshold neural signal detection
Ashok Patel, Bart Kosko
Neural Networks2
2008 Optimal noise benefits in Neyman-Pearson signal detection
abstract
We present an algorithm to find near-optimal "stochastic resonance" (SR) noise benefits for Neyman-Pearson (N-P) hypothesis testing or signal-detection problems. The optimal N-P SR noise is no more than two randomized noise realizations when the optimal noise exists. We give necessary and sufficient conditions for the existence of such optimal noise in fixed detectors. There exists a sequence of noise variables whose detection performance limit is optimal when such noise does not exist. An upper bound limits the number of iterations that the algorithm requires to find such near-optimal noise.
Ashok Patel, Bart Kosko
ICASSP2
2008 Stochastic Resonance in Continuous and Spiking Neuron Models With Levy Noise
abstract
Levy noise can help neurons detect faint or subthreshold signals. Levy noise extends standard Brownian noise to many types of impulsive jump-noise processes found in real and model neurons as well as in models of finance and other random phenomena. Two new theorems and the ItO calculus show that white Levy noise will benefit subthreshold neuronal signal detection if the noise process's scaled drift velocity falls inside an interval that depends on the threshold values. These results generalize earlier "forbidden interval" theorems of neuronal "stochastic resonance" (SR) or noise-injection benefits. Global and local Lipschitz conditions imply that additive white Levy noise can increase the mutual information or bit count of several feedback neuron models that obey a general stochastic differential equation (SDE). Simulation results show that the same noise benefits still occur for some infinite-variance stable Levy noise processes even though the theorems themselves apply only to finite-variance Levy noise. The proves the two ItO-theoretic lemmas that underlie the new Levy noise-benefit theorems.
Ashok Patel, Bart Kosko
IEEE Trans. Neural Networks2
2007 Levy Noise Benefits in Neural Signal Detection
abstract
We use the Ito calculus to prove that a general type of white Levy noise will benefit subthreshold neuronal signal detection if the noise process's scaled drift velocity falls inside an interval that depends on the threshold values. Levy noise generalizes Brownian motion and includes several important jump and impulsive random processes often found in neural and financial-engineering models. A global Lipschitz condition implies that additive white Levy noise can increase the mutual information or bit count of several feedback neuron models that obey a general stochastic differential equation. Simulation results show that the same 'stochastic resonance' noise benefit occurs for at least some impulsive or infinite-variance (stable) Levy noise processes.
Ashok Patel, Bart Kosko
ICASSP (3)2
2006 Mutual-Information Noise Benefits in Brownian Models of Continuous and Spiking Neurons
abstract
The Ito calculus shows that noise benefits can occur in common models of continuous neurons and in random spiking neurons cast as stochastic differential equations. Additive Gaussian noise perturbs the neural dynamical systems as additive Brownian diffusions. The first of two theorems uses a global Lipschitz continuity condition to characterize a stochastic resonance (SR) noise benefit in models of continuous neurons that receive random subthreshold inputs. Brownian diffusions produce an SR noise benefit in the sense that they increase the neuron's mutual information or bit count if the noise mean falls within an interval that depends on model parameters. The second theorem extends an earlier SR result for the random spiking Fitz-Hugh-Nagumo neuron model by replacing a firing-rate approximation with exact stochastic dynamics. This gives an interval-based sufficient condition for an SR noise benefit.
Ashok Patel, Bart Kosko
IJCNN2
2005 Noise benefits in spiking retinal and sensory neuron models
abstract
This paper presents two new theorems that give sufficient conditions (and necessary in the first case) for a noise benefit or stochastic-resonance effect in popular spiking models of retinal neurons and sensory neurons. Small amounts of additive white noise increase the neuron's input-output bit count or Shannon mutual information. This stochastic-resonance (SR) effect applies to standard Poisson spiking models of retinal neurons for all possible types of finite-variance noise and for all impulsive or infinite-variance stable noise. A similar SR result holds for several types of sensory spiking neurons such as the Fitzhugh-Nagumo model and the integrate-and-fire model if the additive noise is Gaussian white noise.
Ashok Patel, Bart Kosko
IJCNN2
2005 Stochastic resonance in noisy spiking retinal and sensory neuron models
Ashok Patel, Bart Kosko
Neural Networks2
2005 Modeling gunshot bruises in soft body armor with an adaptive fuzzy system
abstract
Gunshots produce bruise patterns on persons who wear soft body armor when shot even though the armor stops the bullets. An adaptive fuzzy system modeled these bruise patterns based on the depth and width of the deformed armor given a projectile's mass and momentum. The fuzzy system used rules with sinc-shaped if-part fuzzy sets and was robust against random rule pruning: Median and mean test errors remained low even after removing up to one fifth of the rules. Handguns shot different caliber bullets at armor that had a 10%-ordnance gelatin backing. The gelatin blocks were tissue simulants. The gunshot data tuned the additive fuzzy function approximator. The fuzzy system's conditional variance V[Y/X = x] described the second-order uncertainty of the function approximation. Handguns with different barrel lengths shot bullets over a fixed distance at armor-clad gelatin blocks that we made with Type 250 A Ordnance Gelatin. The bullet-armor experiments found that a bullet's weight and momentum correlated with the depth of its impact on armor-clad gelatin (R2 = 0.881 and p-value < 0.001 for the null hypothesis that the regression line had zero slope). Related experiments on plumber's putty showed that highspeed baseball impacts compared well to bullet-armor impacts for large-caliber handguns. A baseball's momentum correlated with its impact depth in putty (R2 = 0.93 and p-value < 0.001). A bullet's momentum similarly correlated with its armor-impact in putty (R2 = 0.97 and p-value < 0.001). A Gujarati-Chow test showed that the two putty-impact regression lines had statistically indistinguishable slopes for p-value = 0.396. Baseball impact depths were comparable to bullet-armor impact depths: Getting shot with a .22 caliber bullet when wearing soft body armor resembles getting hit in the chest with a 40-mph baseball. Getting shot with a .45 caliber bullet resembles getting hit with a 90-mph baseball.
Ian Lee, Bart Kosko, W. French Anderson
IEEE Trans. Syst. Man Cybern. Part B2
2004 Probable equivalence, superpower sets, and superconditionals
abstract
A natural measure of probabilistic equality between sets leads to two measures of probabilistic conditioning that form the endpoints of a conditioning interval. The interval's lower bound is the standard conditional probability or “subconditional” that describes the probability of a subset relation. The upper bound is a new “superconditional” that describes the probability of the corresponding superset relation. These dual conditioning operators correspond to dual set collections and enjoy optimality relations with respect to these set collections. Fuzzy cubes illustrate these set-collection relations in the two-dimensional case. The subconditional operator corresponds to the usual “power set” of a given set. The dual superconditional operator corresponds to what we call the “superpower set” or the set of all supersets of the given set. The two dual conditioning operators can eliminate each other through simple equalities. They obey dual Bayes theorems but differ in how they respond to statistical independence. © 2004 Wiley Periodicals, Inc. Int J Int Syst 19: 1151–1171, 2004.
Bart Kosko
Int. J. Intell. Syst.1
2004 Adaptive stochastic resonance in noisy neurons based on mutual information
abstract
Noise can improve how memoryless neurons process signals and maximize their throughput information. Such favorable use of noise is the so-called "stochastic resonance" or SR effect at the level of threshold neurons and continuous neurons. This paper presents theoretical and simulation evidence that 1) lone noisy threshold and continuous neurons exhibit the SR effect in terms of the mutual information between random input and output sequences, 2) a new statistically robust learning law can find this entropy-optimal noise level, and 3) the adaptive SR effect is robust against highly impulsive noise with infinite variance. Histograms estimate the relevant probability density functions at each learning iteration. A theorem shows that almost all noise probability density functions produce some SR effect in threshold neurons even if the noise is impulsive and has infinite variance. The optimal noise level in threshold neurons also behaves nonlinearly as the input signal amplitude increases. Simulations further show that the SR effect persists for several sigmoidal neurons and for Gaussian radial-basis-function neurons.
Sanya Mitaim, Bart Kosko
IEEE Trans. Neural Networks2
2003 Almost all noise types can improve the mutual information of threshold neurons that detect subthreshold signals
abstract
Two new theorems show that small amounts of noise can increase the mutual information of threshold neurons that detect subthreshold signals. The first theorem shows that this "stochastic resonance" effect holds for all finite-variance noise probability density functions that obey a simple mean constraint that the user can control. The second theorem shows that this effect holds for all infinite-variance noise types in the broad class of stable distributions. Stable bell curves can model extremely impulsive noise environments. So the second theorem shows that this stochastic-resonance effect is robust against violent fluctuations in the additive noise process.
Bart Kosko, Sanya Mitaim
IJCNN1
2003 Stochastic resonance in noisy threshold neurons
Bart Kosko, Sanya Mitaim
Neural Networks1
2002 A conditioning interval based on superconditionals and superpower sets
abstract
A natural measure of probabilistic equality between sets leads to two measures of probabilistic conditioning that form the endpoints of a conditioning interval (P, Q). The interval's lower bound is the standard conditional probability or what we call the "subconditional" P that describes the probability of a subset relation. The upper bound is a new "superconditional" Q that describes the probability of the corresponding superset relation. These dual conditioning operators correspond to dual set collections and enjoy optimality relations with respect to these set collections. The subconditional operator corresponds to the usual "power set" of a given set. The dual superconditional operator corresponds to what we call the "superpower set" or the set of all supersets of the given set. The two dual conditioning operators obey dual Bayes theorems but differ in how they respond to statistical independence.
Bart Kosko
FUZZ-IEEE1
2001 The shape of fuzzy sets in adaptive function approximation
abstract
The shape of if-part fuzzy sets affects how well feedforward fuzzy systems approximate continuous functions. We explore a wide range of candidate if-part sets and derive supervised learning laws that tune them. Then we test how well the resulting adaptive fuzzy systems approximate a battery of test functions. No one shape emerges as the best. The sine function often does well and has tractable learning, but its undulating side-lobes may have no linguistic meaning. This suggests that function-approximation accuracy may sometimes have to outweigh linguistic or philosophical interpretations. We divide the if-part sets into two large classes. The first consists of n-dimensional joint sets that factor into n scalar sets. These sets ignore the correlations among input vector components. Fuzzy systems suffer in general from exponential rule explosion in high dimensions when they blindly approximate functions. The factorable fuzzy sets themselves also suffer from a curse of dimensionality: they tend to become binary spikes in high dimension. The second class consists of the more general but less common n-dimensional joint sets that do not factor into n scalar fuzzy sets. We present a method for constructing such unfactorable joint sets from scalar distance measures. Fuzzy systems that use unfactorable sets need not suffer from exponential rule explosion but their increased complexity may lead to intractable learning and inscrutable if-then rules. We prove that some of these sets still suffer from spikiness.
Sanya Mitaim, Bart Kosko
IEEE Trans. Fuzzy Syst.2
1998 Stochastic Resonance with Adaptive Fuzzy Systems
Sanya Mitaim, Bart Kosko
ICML2
1998 Neural fuzzy stochastic resonance
abstract
Adaptive systems can learn to add an optimal amount of noise to some nonlinear feedback systems. Noise can improve the signal-to-noise ratio of many nonlinear dynamical systems. This "stochastic resonance" effect occurs in a wide range of physical and biological systems. The SR effect may also occur in engineering systems in signal processing, communications, and control. The noise energy can enhance the faint periodic signals or faint broadband signals that force the dynamical systems. Most SR studies assume full knowledge of a system's dynamics and its noise and signal structure. Fuzzy and other adaptive systems can learn to induce SR based only on samples from the process. These samples can tune a fuzzy system's if-then rules so that the fuzzy system approximates the dynamical system and its noise response. The paper derives the SR optimality conditions that any stochastic learning system should try to achieve. The adaptive system learns the SR effect as the system performs a stochastic gradient ascent on the signal-to-noise ratio. The stochastic learning scheme does not depend on a fuzzy system or any other adaptive system. The learning process is slow and noisy and can require heavy computation. Robust noise suppressors can improve the learning process when we can estimate the impulsiveness of the noise or of other learning terms. Simulations test this SR learning scheme on the popular quartic-bistable dynamical system and on other dynamical systems for many types of noise. Simulations suggest that fuzzy techniques and perhaps other "intelligent" techniques can induce SR in many cases when users cannot state the exact form of the dynamical systems.
Sanya Mitaim, Bart Kosko
SMC2
1998 Special Issue On Intelligent Signal Processing
Simon Haykin 0001, Bart Kosko
Proc. IEEE2
1998 Adaptive stochastic resonance
abstract
This paper shows how adaptive systems can learn to add an optimal amount of noise to some nonlinear feedback systems. This "stochastic resonance" (SR) effect occurs in a wide range of physical and biological systems. The noise energy can enhance the faint periodic signals or faint broadband signals that force the dynamical systems. Fuzzy and other adaptive systems can learn to induce SR based only on samples from the process. The paper derives the SR optimality conditions that any stochastic learning system should try to achieve. The adaptive system learns the SR effect as the system performs a stochastic gradient ascent on the signal-to-noise ratio. The stochastic learning scheme does not depend on a fuzzy system or any other adaptive system. Simulations test this SR learning scheme on the popular quartic-bistable dynamical system and on other dynamical systems. The driving noise types range from Gaussian white noise to impulsive noise to chaotic noise.
Sanya Mitaim, Bart Kosko
Proc. IEEE2
1998 Global stability of generalized additive fuzzy systems
abstract
The paper explores the stability of a class of feedback fuzzy systems. The class consists of generalized additive fuzzy systems that compute a system output as a convex sum of linear operators, continuous versions of these systems are globally asymptotically stable if all rule matrices are stable (negative definite). So local rule stability leads to global system stability. This relationship between local and global system stability does not hold for the better known discrete versions of feedback fuzzy systems. A corollary shows that it does hold for the discrete versions in the special but practical case of diagonal rule matrices. The paper first reviews additive fuzzy systems and then extends them to the class of generalized additive fuzzy systems. It also derives the basic ratio structure of additive fuzzy systems and shows how supervised learning can tune their parameters.
Bart Kosko
IEEE Trans. Syst. Man Cybern. Part C1
1996 Fuzzy throttle and brake control for platoons of smart cars
Hyun Mun Kim, Julie A. Dickerson, Bart Kosko
Fuzzy Sets Syst.3
1996 Fuzzy prediction and filtering in impulsive noise
Hyun Mun Kim, Bart Kosko
Fuzzy Sets Syst.2
1996 Fuzzy function approximation with ellipsoidal rules
abstract
A fuzzy rule can have the shape of an ellipsoid in the input-output state spare of a system. Then an additive fuzzy system approximates a function by covering its graph with ellipsoidal rule patches. It averages rule patches that overlap. The best fuzzy rules cover the extrema or bumps in the function. Neural or statistical clustering systems can approximate the unknown fuzzy rules from training data. Neural systems can then both tune these rules and add rules to improve the function approximation. We use a hybrid neural system that combines unsupervised and supervised learning to find and tune the rules in the form of ellipsoids. Unsupervised competitive learning finds the first-order and second-order statistics of clusters in the training data. The covariance matrix of each cluster gives an ellipsoid centered at the vector or centroid of the data cluster. The supervised neural system learns with gradient descent. It locally minimizes the mean-squared error of the fuzzy function approximation. In the hybrid system unsupervised learning initializes the gradient descent. The hybrid system tends to give a more accurate function approximation than does the lone unsupervised or supervised system. We found a closed-form model for the optimal rules when only the centroids of the ellipsoids change. We used numerical techniques to find the optimal rules in the general case.
Julie A. Dickerson, Bart Kosko
IEEE Trans. Syst. Man Cybern. Part B2
1995 Optimal fuzzy rules cover extrema
abstract
A fuzzy system approximates a function by covering the graph of the function with fuzzy rule patches and averaging patches that overlap. But the number of rules grows exponentially with the total number of input and output variables. the best rules cover the extrema or bumps in the function—they patch the bumps. For mean-squared approximation this follows from the mean value theorem of calculus. Optimal rules can help reduce the computational burden. to find them we can find or learn the zeroes of the derivative map and then center input fuzzy sets at these points. Neural systems can then both tune these rules and add rules to improve the function approximation. © 1995 John Wiley & Sons, Inc.
Bart Kosko
Int. J. Intell. Syst.1
1995 Adaptive fuzzy frequency hopper
abstract
An adaptive fuzzy system generates the frequency hopping sequence for a spread spectrum communications system. The system learns rules from data and acts as a pseudorandom number generator. The IMSL uniform random number generator gives training samples. An adaptive scheme learns associations between previous samples and the current sample and encodes these as fuzzy rules. The output fuzzy set for each rule acts as a conditional probability density function. The if-part of the rule states the conditions. At each step thirty prior outputs, scanned according to a fixed sampling pattern, give a new sample distribution x/sub k/. The vector x/sub k/ partially matches the if-part distribution of a fuzzy rule and partially fires that rule's output fuzzy set. With the estimated output fuzzy sets the fuzzy system computes the conditional density p/sub .>
Peter J. Pacini, Bart Kosko
IEEE Trans. Commun.2
1994 Fuzzy Systems as Universal Approximators
abstract
An additive fuzzy system can uniformly approximate any real continuous function on a compact domain to any degree of accuracy. An additive fuzzy system approximates the function by covering its graph with fuzzy patches in the input-output state space and averaging patches that overlap. The fuzzy system computes a conditional expectation E|Y|X| if we view the fuzzy sets as random sets. Each fuzzy rule defines a fuzzy patch and connects commonsense knowledge with state-space geometry. Neural or statistical clustering systems can approximate the unknown fuzzy patches from training data. These adaptive fuzzy systems approximate a function at two levels. At the local level the neural system approximates and tunes the fuzzy rules. At the global level the rules or patches approximate the function.>
Bart Kosko
IEEE Trans. Computers1
1994 The probability monopoly
abstract
Probability is a very special case of fuzziness. It always faces two limits. First, it works with bivalent sets A. Second, probability measures need small infinities. A probability measure maps the sets in a sigma-algebra to the unit interval /spl lsqb/O, 1/spl rsqb/. Fuzzy theory challenges the probability monopoly. Probabilists have attacked it with gusto to keep their monopoly status, to have, the only uncertainty theory in the unit interval /spl lsqb/O, 1/spl rsqb/. But the fuzzy math is sound. Its world view of shades of gray has a deep intuitive ring. And the new fuzzy products have come into their own in the marketplace.>
Bart Kosko
IEEE Trans. Fuzzy Syst.1
1993 Addition as fuzzy mutual entropy
Bart Kosko
Inf. Sci.1
1992 Fuzzy subimage classification in image sequence coding
abstract
Fuzzy systems are used to classify subimages efficiently in adaptive hybrid transform/predictive coding of image sequences. An adaptive fuzzy system estimates fuzzy rules by clustering input-output data generated by the subimage classification method of W.-H. Chen and C.H. Smith (1977). The fuzzy rules define patches in the state space and approximate an unknown function by covering its graph with patches. The fuzzy system classifies subimages into four temporally active subimage classes according to the between-frame prediction error signal. The system encodes active subimages with more bits, and inactive subimages with fewer bits, to compress the image data. Fuzzy classification improved coding performance over nonfuzzy classification and nonadaptive interframe coding.>
Seong-Gon Kong, Bart Kosko
ICASSP2
1992 Adaptive fuzzy systems for backing up a truck-and-trailer
abstract
Fuzzy control systems and neural-network control systems for backing up a simulated truck, and truck-and-trailer, to a loading dock in a parking lot are presented. The supervised backpropagation learning algorithm trained the neural network systems. The robustness of the neural systems was tested by removing random subsets of training data in learning sequences. The neural systems performed well but required extensive computation for training. The fuzzy systems performed well until over 50% of their fuzzy-associative-memory (FAM) rules were removed. They also performed well when the key FAM equilibration rule was replaced with destructive, or ;sabotage', rules. Unsupervised differential competitive learning (DCL) and product-space clustering adaptively generated FAM rules from training data. The original fuzzy control systems and neural control systems generated trajectory data. The DCL system rapidly recovered the underlying FAM rules. Product-space clustering converted the neural truck systems into structured sets of FAM rules that approximated the neural system's behavior.
Seong-Gon Kong, Bart Kosko
IEEE Trans. Neural Networks2
1991 Differential competitive learning for centroid estimation and phoneme recognition
abstract
A comparison is made of a differential-competitive-learning (DCL) system with two supervised competitive-learning (SCL) systems for centroid estimation and for phoneme recognition. DCL provides a form of unsupervised adaptive vector quantization. Standard stochastic competitive-learning systems learn only if neurons win a competition for activation induced by randomly sampled patterns. DCL systems learn only if the competing neurons change their competitive signal. Signal-velocity information provides unsupervised local reinforcement during learning. The sign of the neuronal signal derivative rewards winners and punishes losers. Standard competitive learning ignores instantaneous win-rate information. Synaptic fan-in vectors adaptively quantize the randomly sampled pattern space into nearest-neighbor decision classes. More generally, the synaptic-vector distribution estimates the unknown sampled probability density function p( x). Simulations showed that unsupervised DCL-trained synaptic vectors converged to class centroids at least as fast as, and wandered less about these centroids than, SCL-trained synaptic vectors did. Simulations on a small set of English phonemes favored DCL over SCL for classification accuracy.
Seong-Gon Kong, Bart Kosko
IEEE Trans. Neural Networks2
1991 Stochastic competitive learning
abstract
Competitive learning systems are examined as stochastic dynamical systems. This includes continuous and discrete formulations of unsupervised, supervised, and differential competitive learning systems. These systems estimate an unknown probability density function from random pattern samples and behave as adaptive vector quantizers. Synaptic vectors, in feedforward competitive neural networks, quantize the pattern space and converge to pattern class centroids or local probability maxima. A stochastic Lyapunov argument shows that competitive synaptic vectors converge to centroids exponentially quickly and reduces competitive learning to stochastic gradient descent. Convergence does not depend on a specific dynamical model of how neuronal activations change. These results extend to competitive estimation of local covariances and higher order statistics.
Bart Kosko
IEEE Trans. Neural Networks1
1990 Comparison of fuzzy and neural truck backer-upper control systems
abstract
A simple fuzzy control system and a simple neural control system for backing up a truck in an open parking lot are developed. The choice of control problem was prompted by the recent, successful, neural network truck backer-upper simulation of Nguyen and Widrow (Proc. Int. Joint Conference on Neural Networks, vol.2, p.357-363, June, 1989). The authors were unable to exactly replicate the neural network they used. Instead the authors built the best backpropagation network they could with essentially the same kinematics and compared it to the best fuzzy controller they could develop. The fuzzy controller compares favorably with the neural controller in terms of black-box computation load, smoothness of truck trajectories, and robustness. Robustness of the fuzzy controller is studied by deliberately adding confusing FAM (fuzzy associative memory), rules-sabotage rules-to the system and by randomly removing different subsets of FAM rules. Robustness of the neural controller is studied by randomly removing different portions of the training data. It is concluded that fuzzy control shows optimal truck backing-up performance
Seong-Gon Kong, Bart Kosko
IJCNN2
1990 Stochastic competitive learning
abstract
The probabilistic foundations of competitive learning systems are developed. Continuous and discrete formulations of unsupervised, supervised, and differential competitive learning systems are studied. These systems estimate an unknown probability density function from random pattern samples and behave as adaptive vector quantizers. Synaptic vectors in feedforward competitive neural networks quantize the pattern space and converge to pattern class centroids or local probability maxima. The stochastic calculus and a Lyapunov argument prove that competitive synaptic vectors converge to centroids exponentially quickly. Convergence does not depend on a specific dynamical model of how neuronal activations change
Bart Kosko
IJCNN1
1990 Unsupervised learning in noise
abstract
A new hybrid learning law, the differential competitive law, which uses the neuronal signal velocity as a local unsupervised reinforcement mechanism, is introduced, and its coding and stability behavior in feedforward and feedback networks is examined. This analysis is facilitated by the recent Gluck-Parker pulse-coding interpretation of signal functions in differential Hebbian learning systems. The second-order behavior of RABAM (random adaptive bidirectional associative memory) Brownian-diffusion systems is summarized by the RABAM noise suppression theorem: the mean-squared activation and synaptic velocities decrease exponentially quickly to their lower bounds, the instantaneous noise variances driving the system. This result is extended to the RABAM annealing model, which provides a unified framework from which to analyze Geman-Hwang combinatorial optimization dynamical systems and continuous Boltzmann machine learning.
Bart Kosko
IEEE Trans. Neural Networks1
1988 Hidden patterns in combined and adaptive knowledge networks
Bart Kosko
Int. J. Approx. Reason.1
1988 Bidirectional associative memories
abstract
Stability and encoding properties of two-layer nonlinear feedback neural networks are examined. Bidirectionality is introduced in neural nets to produce two-way associative search for stored associations. The bidirectional associative memory (BAM) is the minimal two-layer nonlinear feedback network. The author proves that every n-by-p matrix M is a bidirectionally stable heteroassociative content-addressable memory for both binary/bipolar and continuous neurons. When the BAM neutrons are activated, the network quickly evolves to a stable state of two-pattern reverberation, or resonance. The stable reverberation corresponds to a system energy local minimum. Heteroassociative information is encoded in a BAM by summing correlation matrices. The BAM storage capacity for reliable recall is roughly m
Bart Kosko
IEEE Trans. Syst. Man Cybern.1
1986 Fuzzy knowledge combination
abstract
A general answer is given to what one should conclude from disagreeing experts. the answer is generalized further to incorporate the experts' credibility weights. the answer rests on a wide range of intuitively based epistemic axioms, scientific and philosophical conjectures, and formal mathematical relationships. A recurring theme is the making of Bellman - Zadeh fuzzy decisions, wherein a decision is the intersection of fuzzy goal and fuzzy constraint subsets of some space of alternatives. Another result is that measures of central tendency, such as the arithmetic mean, make poor knowledge combination operators. Formally, fuzzy knowledge combination operators are sought. the function space of knowledge combination operators ø: K″ → K is shrunk by imposing successive axioms. the final shrunken set is said to consist of admissible knowledge combination operators. Some of its mathematical properties are explored and a simple admissible operator is finally chosen. Knowledge sources Xi: S → K are mappings from epistemic stimuli or questions into a knowledge response set K. the uncertainty of the underlying epistemic situations is captured by the cardinality of K and by the fuzziness of its partial ordering. Admissible knowledge combination operators Aggregate knowledge responses in some desirable way. the arithmetic mean is not admissible. Nor in general is a probabilistic framework even definable in the abstract poset setting employed by this theory. the fuzzy knowledge combination theory is extended by associating general credibility weights with the knowledge sources. A new set of weighting axioms is required to satisfy certain intuitions and to satisfy the admissibility axioms. General weighting functions are obtained and thereby weighted admissible operators are obtained. the weighted mean still proves inadmissible. Appendix I contains a technical glossary and summary of the proposed fuzzy knowledge combination theory. Appendix II contains proofs of the probabilistic uncertainty theorems required for the uncertainty testbed used in the theory.
Bart Kosko
Int. J. Intell. Syst.1
1986 Fuzzy Cognitive Maps
Bart Kosko
Int. J. Man Mach. Stud.1
1986 Fuzzy entropy and conditioning
Bart Kosko
Inf. Sci.1
1986 Counting with Fuzzy Sets
abstract
The notion of fuzzy set cardinality is examined. Zadeh's suggested measure of fuzzy cardinality, the sigma-count, is adopted and shown to generalize classical counting measure. This allows many combinatorial structures and counting techniques to be fuzzified, and hence used in knowledge representation and pattern recognition models. A fuzzy set review is found in the Appendix.
Bart Kosko
IEEE Trans. Pattern Anal. Mach. Intell.1