John Sum

dblp:s/JohnSum · also J. P. F. Sum, Pui-Fai Sum · DBLP profile ↗
← Back
83ranked-venue papers
31as first author
10since 2021 · last 2025
0000-0002-6965-4219ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 77 · 27 first-author · 10 since 2021Human-computer interaction and ubiquitous computing · 4 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 3 · 2 first-authorSystems, architecture and hardware · 2 · 2 first-author
YearPublicationVenuePosition
2025 Analysis and Design of a Distributed kWTA With Application in Sealed-Bid Auctions With Bidding Price Privacy Protection
abstract
This article presents a distributed k-winner-take-all (kWTA) with application in sealed-bid auctions with bidding price privacy protection. The proposed kWTA is in essence a distributed network of n agents which are arbitrarily connected. Let $\aleph _{i}$ be the set of neighbor agents of the ith agent, $u_{i}$ , $x_{i}$ , and $z_{i}$ are, respectively, its input, state variable, and output. The dynamics of the ith agent is given by $ ((dx_{i}(t))/dt) = \tau \left \{{{ z_{i}(x_{i}(t)) - (k/n) - \beta \sum _{j\in \aleph _{i}} (x_{i}(t) - x_{j}(t)) }}\right \}, z_{i}(x_{i}(t)) = h(u_{i}-x_{i}(t)), \text {for}~i = 1, \ldots , n$ where $\beta \gt 0$ , k is the number of winners and $h(\cdot)$ is the Heaviside function. By the theory of discontinuous dynamic systems, it is shown that the state equation for $d{\mathbf {x}}(t)/dt$ could be formulated as a gradient differential inclusion which minimizes the following nonsmooth convex function. $V({\mathbf {x}}) = \sum _{i=1}^{n} \max \{0, u_{i} - x_{i}\} + (k/n) \sum _{i=1}^{n} x_{i} + (\beta /2){\mathbf {x}}^{T} {\mathbf {L}} {\mathbf {x}}$ where ${\mathbf {x}} = (x_{1}, \ldots , x_{n})^{n}$ and ${\mathbf {L}} \in R^{n\times n}$ is the graph Laplacian matrix. A sufficient condition for $\beta $ is derived for the kWTA giving correct output and the condition is then applied in showing that ${\mathbf {z}}(t)$ converges to the correct output in finite-time. If $\beta \rightarrow \infty $ and $x_{1}(0) = \cdots = x_{n}(0)$ , we further show that $x_{1}(t) = \cdots = x_{n}(t)$ for $t \geq 0$ , and both ${\mathbf {z}}(t)$ and ${\mathbf {x}}(t)$ converge in finite-time. Besides, $x_{i}$ converges to $u_{\pi _{n-k+1}}$ (resp. $u_{\pi _{n-k}}$ ) if $x_{i}(0) \gg 1$ (resp. $x_{i}(0) = 0)$ for $i = 1, \ldots , n$ . If the input $u_{i}$ is set to be the bid price of the ith bidder and $k = 1$ , the proposed kWTA is able to determine both the winners and the clearing price for a sealed-bid first (resp. second) price auction in a distributed manner. Once ${\mathbf {z}}(t)$ and ${\mathbf {x}}(t)$ converge, each bidder can reveal from: 1) $z_{i}$ if he/she is a winner and 2) $x_{i}$ the clearing price. As bidders do not have to disclose their bidding prices during the winner (resp. the clearing price) determination process, the loosing (resp. winning) bidding price privacy can be protected in a sealed-bid first (resp. second) price auction. It is insofar the first application of an kWTA beyond the winner's determination.
John Sum, Andrew Chi-Sing Leung, Janet C. C. Chang
IEEE Trans. Neural Networks Learn. Syst.1
2025 A Fast Wang kWTA With Application in Sealed-Bid Uniform Price Auction
abstract
In this brief, two fast discrete-time Wang kWTA (Fast Wang kWTA) algorithms are presented with an application in sealed-bid uniform price auctions. These algorithms can either be implemented in centralized or distributed manner. The structure of the Fast Wang kWTA is essentially the same as the original Wang k-winner-take-all (kWTA), except that our state update method is based on bisection method instead of gradient descent. By that, the number of iterations for getting correct output is largely reduced. Besides, the number is just a factor depended on the guess of the maximum input value. It is independent of the number of inputs, the number of winners, and the learning step size. The number of iterations is far smaller than the number required in the original Wang kWTA. In sequel, this Fast Wang kWTA is particularly suitable to be applied in solving the winner (resp. price) determination in real time and in distributed manner for a sealed-bid auction. In addition, the Fast Wang kWTA can ensure bidding price protection even if the communicated data are not encrypted and leaked.
John Sum, Andrew Chi-Sing Leung, Janet C. C. Chang
IEEE Trans. Neural Networks Learn. Syst.1
2024 A Leaky Wang kWTA
John Sum, Andrew Chi-Sing Leung, Janet C. C. Chang
ICONIP (4)1
2024 Outlier-Robust Range-Based Method for Estimating the Location and Velocity of a Moving Source Using LPNN
Wenxin Xiong, Keyuan Hu, Jiajun He 0001, Andrew Chi-Sing Leung, Hing-Cheung So, John Sum
ICONIP (2)6
2024 Influence of Imperfections on the Operational Correctness of DNN-kWTA Model
abstract
The dual neural network (DNN)-based k -winner-take-all (WTA) model is able to identify the k largest numbers from its m input numbers. When there are imperfections, such as non-ideal step function and Gaussian input noise, in the realization, the model may not output the correct result. This brief analyzes the influence of the imperfections on the operational correctness of the model. Due to the imperfections, it is not efficient to use the original DNN- k WTA dynamics for analyzing the influence. In this regard, this brief first derives an equivalent model to describe the dynamics of the model under the imperfections. From the equivalent model, we derive a sufficient condition for which the model outputs the correct result. Thus, we apply the sufficient condition to design an efficiently estimation method for the probability of the model outputting the correct result. Furthermore, for the inputs with uniform distribution, a closed form expression for the probability value is derived. Finally, we extend our analysis for handling non-Gaussian input noise. Simulation results are provided to validate our theoretical results.
Wenhao Lu, Andrew Chi-Sing Leung, John Sum
IEEE Trans. Neural Networks Learn. Syst.3
2023 A Distributed kWTA for Decentralized Auctions
Gary Sum, John Sum, Andrew Chi-Sing Leung, Janet C. C. Chang
ICONIP (8)2
2023 Regularization Effect of Random Node Fault/Noise on Gradient Descent Learning Algorithm
abstract
For decades, adding fault/noise during training by gradient descent has been a technique for getting a neural network (NN) tolerant to persistent fault/noise or getting an NN with better generalization. In recent years, this technique has been readvocated in deep learning to avoid overfitting. Yet, the objective function of such fault/noise injection learning has been misinterpreted as the desired measure (i.e., the expected mean squared error (mse) of the training samples) of the NN with the same fault/noise. The aims of this article are: 1) to clarify the above misconception and 2) investigate the actual regularization effect of adding node fault/noise when training by gradient descent. Based on the previous works on adding fault/noise during training, we speculate the reason why the misconception appears. In the sequel, it is shown that the learning objective of adding random node fault during gradient descent learning (GDL) for a multilayer perceptron (MLP) is identical to the desired measure of the MLP with the same fault. If additive (resp. multiplicative) node noise is added during GDL for an MLP, the learning objective is not identical to the desired measure of the MLP with such noise. For radial basis function (RBF) networks, it is shown that the learning objective is identical to the corresponding desired measure for all three fault/noise conditions. Empirical evidence is presented to support the theoretical results and, hence, clarify the misconception that the objective function of a fault/noise injection learning might not be interpreted as the desired measure of the NN with the same fault/noise. Afterward, the regularization effect of adding node fault/noise during training is revealed for the case of RBF networks. Notably, it is shown that the regularization effect of adding additive or multiplicative node noise (MNN) during training an RBF is reducing network complexity. Applying dropout regularization in RBF networks, its effect is the same as adding MNN during training.
John Sum, Andrew Chi-Sing Leung
IEEE Trans. Neural Networks Learn. Syst.1
2023 A Globally Stable LPNN Model for Sparse Approximation
abstract
The objective of compressive sampling is to determine a sparse vector from an observation vector. This brief describes an analog neural method to achieve the objective. Unlike previous analog neural models which either resort to the$\ell _{1}$-norm approximation or are with local convergence only, the proposed method avoids any approximation of the$\ell _{1}$-norm term and is probably capable of leading to the optimum solution. Moreover, its computational complexity is lower than that of the other three comparison analog models. Simulation results show that the error performance of the proposed model is comparable to several state-of-the-art digital algorithms and analog models and that its convergence is faster than that of the comparison analog neural models.
Hao Wang 0075, Ruibin Feng, Andrew Chi-Sing Leung, John Sum, Anthony G. Constantinides
IEEE Trans. Neural Networks Learn. Syst.4
2022 Effect of Logistic Activation Function and Multiplicative Input Noise on DNN-kWTA Model
Wenhao Lu, Andrew Chi-Sing Leung, John Sum
ICONIP (4)3
2022 DNN-kWTA With Bounded Random Offset Voltage Drifts in Threshold Logic Units
abstract
The dual neural network-based$k$-winner-take-all (DNN-$k$WTA) is an analog neural model that is used to identify the$k$largest numbers from$n$inputs. Since threshold logic units (TLUs) are key elements in the model, offset voltage drifts in TLUs may affect the operational correctness of a DNN-$k$WTA network. Previous studies assume that drifts in TLUs follow some particular distributions. This brief considers that only the drift range, given by$[-\Delta, \Delta]$, is available. We consider two drift cases: time-invariant and time-varying. For the time-invariant case, we show that the state of a DNN-$k$WTA network converges. The sufficient condition to make a network with the correct operation is given. Furthermore, for uniformly distributed inputs, we prove that the probability that a DNN-$k$WTA network operates properly is greater than$(1-2\Delta)^{n}$. The aforementioned results are generalized for the time-varying case. In addition, for the time-invariant case, we derive a method to compute the exact convergence time for a given data set. For uniformly distributed inputs, we further derive the mean and variance of the convergence time. The convergence time results give us an idea about the operational speed of the DNN-$k$WTA model. Finally, simulation experiments have been conducted to validate those theoretical results.
Wenhao Lu, Andrew Chi-Sing Leung, John Sum, Yi Xiao 0004
IEEE Trans. Neural Networks Learn. Syst.3
2020 Analysis on the Boltzmann Machine with Random Input Drifts in Activation Function
Wenhao Lu, Andrew Chi-Sing Leung, John Sum
ICONIP (3)3
2020 Constrained Center Loss for Image Classification
Zhanglei Shi, Hao Wang 0075, Andrew Chi-Sing Leung, John Sum
ICONIP (5)4
2020 A Limitation of Gradient Descent Learning
abstract
Over decades, gradient descent has been applied to develop learning algorithm to train a neural network (NN). In this brief, a limitation of applying such algorithm to train an NN with persistent weight noise is revealed. Let V(w) be the performance measure of an ideal NN. V(w) is applied to develop the gradient descent learning (GDL). With weight noise, the desired performance measure (denoted as J(w) ) is E[V(~w)|w] , where ~w is the noisy weight vector. Applying GDL to train an NN with weight noise, the actual learning objective is clearly not V(w) but another scalar function L(w) . For decades, there is a misconception that L(w) = J(w) , and hence, the actual model attained by the GDL is the desired model. However, we show that it might not: 1) with persistent additive weight noise, the actual model attained is the desired model as L(w) = J(w) ; and 2) with persistent multiplicative weight noise, the actual model attained is unlikely the desired model as L(w) ≠ J(w) . Accordingly, the properties of the models attained as compared with the desired models are analyzed and the learning curves are sketched. Simulation results on 1) a simple regression problem and 2) the MNIST handwritten digit recognition are presented to support our claims.
John Sum, Andrew Chi-Sing Leung, Kevin I.-J. Ho
IEEE Trans. Neural Networks Learn. Syst.1
2019 Fault Tolerant Broad Learning System
Muideen Adegoke, Andrew Chi-Sing Leung, John Sum
ICONIP (4)3
2019 Analysis on Dropout Regularization
John Sum, Andrew Chi-Sing Leung
ICONIP (5)1
2019 Explicit Center Selection and Training for Fault Tolerant RBF Networks
Hiu Tung Wong, Zhenni Wang, Andrew Chi-Sing Leung, John Sum
ICONIP (2)4
2019 Learning Algorithm for Boltzmann Machines With Additive Weight and Bias Noise
abstract
This brief presents analytical results on the effect of additive weight/bias noise on a Boltzmann machine (BM), in which the unit output is in {-1, 1} instead of {0, 1}. With such noise, it is found that the state distribution is yet another Boltzmann distribution but the temperature factor is elevated. Thus, the desired gradient ascent learning algorithm is derived, and the corresponding learning procedure is developed. This learning procedure is compared with the learning procedure applied to train a BM with noise. It is found that these two procedures are identical. Therefore, the learning algorithm for noise-free BMs is suitable for implementing as an online learning algorithm for an analog circuit-implemented BM, even if the variances of the additive weight noise and bias noise are unknown.
John Sum, Andrew Chi-Sing Leung
IEEE Trans. Neural Networks Learn. Syst.1
2018 A Robust LPNN Technique for Target Localization Under Hybrid TOA/AOA Measurements
Muideen Adegoke, Andrew Chi-Sing Leung, John Sum
ICONIP (2)3
2018 Fault-Resistant Algorithms for Single Layer Neural Networks
Muideen Adegoke, Andrew Chi-Sing Leung, John Sum
ICONIP (2)3
2018 MCP Based Noise Resistant Algorithm for Training RBF Networks and Selecting Centers
Hao Wang 0075, Andrew Chi-Sing Leung, John Sum
ICONIP (2)3
2018 Robustness Analysis on Dual Neural Network-based k WTA With Input Noise
abstract
This paper studies the effects of uniform input noise and Gaussian input noise on the dual neural network-based WTA (DNN- WTA) model. We show that the state of the network (under either uniform input noise or Gaussian input noise) converges to one of the equilibrium points. We then derive a formula to check if the network produce correct outputs or not. Furthermore, for the uniformly distributed inputs, two lower bounds (one for each type of input noise) on the probability that the network produces the correct outputs are presented. Besides, when the minimum separation amongst inputs is given, we derive the condition for the network producing the correct outputs. Finally, experimental results are presented to verify our theoretical results. Since random drift in the comparators can be considered as input noise, our results can be applied to the random drift situation.
Ruibin Feng, Andrew Chi-Sing Leung, John Sum
IEEE Trans. Neural Networks Learn. Syst.3
2018 On Wang k WTA With Input Noise, Output Node Stochastic, and Recurrent State Noise
abstract
In this paper, the effect of input noise, output node stochastic, and recurrent state noise on the Wang $k$ WTA is analyzed. Here, we assume that noise exists at the recurrent state $y(t)$ and it can either be additive or multiplicative. Besides, its dynamical change (i.e., $dy/dt$ ) is corrupted by noise as well. In sequel, we model the dynamics of $y(t)$ as a stochastic differential equation and show that the stochastic behavior of $y(t)$ is equivalent to an Ito diffusion. Its stationary distribution is a Gibbs distribution, whose modality depends on the noise condition. With moderate input noise and very small recurrent state noise, the distribution is single modal and hence $y(\infty )$ has high probability varying within the input values of the $k$ and $k+1$ winners (i.e., correct output). With small input noise and large recurrent state noise, the distribution could be multimodal and hence $y(\infty )$ could have probability varying outside the input values of the $k$ and $k+1$ winners (i.e., incorrect output). In this regard, we further derive the conditions that the $k$ WTA has high probability giving correct output. Our results reveal that recurrent state noise could have severe effect on Wang $k$ WTA. But, input noise and output node stochastic could alleviate such an effect.
John Sum, Andrew Chi-Sing Leung, Kevin I.-J. Ho
IEEE Trans. Neural Networks Learn. Syst.1
2017 Scheduling jobs with multitasking and asymmetric switching costs
abstract
In this paper, we investigate the job scheduling problems with human multitasking and asymmetric switching costs. It is shown that the makespan problem is binary NP-hard. Both the total completion time and the due date assignment problems are unary NP-hard. We then consider a special case in which the cost of switching from job A (interrupted job) to job B (interrupting job) is in a form of κ1fAp+ κ2fBw, where κ1and κ2are constants and fApand fBware costs depending on job A and B respectively. With this special form of switching cost, we show that the makespan, the total completion time and the due date assignment problems can be formulated as linear assignment problems and thus be solved in polynomial time. Even if stress effect is introduced, these three scheduling problems with the special form of asymmetric switching costs are polynomial time solvable.
Kevin I.-J. Ho, John Sum
SMC2
2017 Editorial for Special Issue on ICONIP 2014
John Sum, Andrew Chi-Sing Leung
Neural Process. Lett.1
2016 Analysis of the DNN-kWTA Network Model with Drifts in the Offset Voltages of Threshold Logic Units
Ruibin Feng, Andrew Chi-Sing Leung, John Sum
ICONIP (4)3
2016 Objective Function and Learning Algorithm for the General Node Fault Situation
abstract
Fault tolerance is one interesting property of artificial neural networks. However, the existing fault models are able to describe limited node fault situations only, such as stuck-at-zero and stuck-at-one. There is no general model that is able to describe a large class of node fault situations. This paper studies the performance of faulty radial basis function (RBF) networks for the general node fault situation. We first propose a general node fault model that is able to describe a large class of node fault situations, such as stuck-at-zero, stuck-at-one, and the stuck-at level being with arbitrary distribution. Afterward, we derive an expression to describe the performance of faulty RBF networks. An objective function is then identified from the formula. With the objective function, a training algorithm for the general node situation is developed. Finally, a mean prediction error (MPE) formula that is able to estimate the test set error of faulty networks is derived. The application of the MPE formula in the selection of basis width is elucidated. Simulation experiments are then performed to demonstrate the effectiveness of the proposed method.
Yi Xiao 0004, Ruibin Feng, Andrew Chi-Sing Leung, John Sum
IEEE Trans. Neural Networks Learn. Syst.4
2015 Non-Line-of-Sight Mitigation via Lagrange Programming Neural Networks in TOA-Based Localization
Zi-Fa Han, Andrew Chi-Sing Leung, Hing-Cheung So, John Sum, Anthony G. Constantinides
ICONIP (3)4
2015 Noise on Gradient Systems with Forgetting
John Sum, Andrew Chi-Sing Leung, Kevin I.-J. Ho
ICONIP (3)2
2015 Analysis on the Effect of Multitasking
abstract
Recently, the ideas of switching cost and interruption function have been introduced in modeling the scheduling problems with multitasking and the effect of multitasking is investigated by computer simulations. In this paper, we analyze the effect of multitasking on the total completion time (TCT) and total weighted completion time (TWCT) by statistical analysis. If the cost of job switching is a constant and the amount of interruption is proportional to the remaining processing time of the interrupting job, the optimal TCT schedule can be obtained by the shortest processing time first (SPTF) rule. Thus, the optimal TCT in the presence of multitasking is derived and compared with the optimal TCT without multitasking. Assuming that the values of switching cost and the proportional constant are small, the optimal TWCT in the presence of multitasking is expressed and compared with the optimal TWCT without multitasking. With mild statistical assumptions on the processing times and weights, the expected TCT and TWCT are derived and the effect of multitasking on TCT (respectively TWCT) is analyzed by numerical plots against the switching cost and the proportional constant. Results reveal that multitasking could have significant effect on both the TCT and TWCT. Hence, multitasking should be avoided in a work place.
John Sum, Kevin I.-J. Ho
SMC1
2015 Online Training for Open Faulty RBF Networks
Yi Xiao 0004, Ruibin Feng, Andrew Chi-Sing Leung, John Sum
Neural Process. Lett.4
2015 Properties and Performance of Imperfect Dual Neural Network-Based k WTA Networks
abstract
The dual neural network (DNN)-based k -winner-take-all ( k WTA) model is an effective approach for finding the k largest inputs from n inputs. Its major assumption is that the threshold logic units (TLUs) can be implemented in a perfect way. However, when differential bipolar pairs are used for implementing TLUs, the transfer function of TLUs is a logistic function. This brief studies the properties of the DNN- kWTA model under this imperfect situation. We prove that, given any initial state, the network settles down at the unique equilibrium point. Besides, the energy function of the model is revealed. Based on the energy function, we propose an efficient method to study the model performance when the inputs are with continuous distribution functions. Furthermore, for uniformly distributed inputs, we derive a formula to estimate the probability that the model produces the correct outputs. Finally, for the case that the minimum separation ∆min of the inputs is given, we prove that if the gain of the activation function is greater than 1/4∆min max(ln 2n, 2 ln 1 - ϵ/ϵ ), then the network can produce the correct outputs with winner outputs greater than 1-ϵ and loser outputs less than ϵ, where ϵ is the threshold less than 0.5.
Ruibin Feng, Andrew Chi-Sing Leung, John Sum, Yi Xiao 0004
IEEE Trans. Neural Networks Learn. Syst.3
2014 The Performance of the Stochastic DNN-kWTA Network
Ruibin Feng, Andrew Chi-Sing Leung, Kai Tat Ng, John Sum
ICONIP (1)4
2014 Recurrent networks for compressive sampling
Andrew Chi-Sing Leung, John Sum, Anthony G. Constantinides
Neurocomputing2
2014 Lagrange programming neural networks for time-of-arrival-based source localization
Andrew Chi-Sing Leung, John Sum, Hing-Cheung So, Anthony G. Constantinides, Frankie K. W. Chan
Neural Comput. Appl.2
2013 GPU Accelerated Spherical K-Means Training
Yi Xiao 0004, Ruibin Feng, Andrew Chi-Sing Leung, John Sum
ICONIP (2)4
2013 HEALPIX DCT technique for compressing PCA-based illumination adjustable images
John Sum, Andrew Chi-Sing Leung, Ray C. C. Cheung, Tze-Yui Ho
Neural Comput. Appl.1
2013 Effect of Input Noise and Output Node Stochastic on Wang's kWTA
abstract
Recently, an analog neural network model, namely Wang's kWTA, was proposed. In this model, the output nodes are defined as the Heaviside function. Subsequently, its finite time convergence property and the exact convergence time are analyzed. However, the discovered characteristics of this model are based on the assumption that there are no physical defects during the operation. In this brief, we analyze the convergence behavior of the Wang's kWTA model when defects exist during the operation. Two defect conditions are considered. The first one is that there is input noise. The second one is that there is stochastic behavior in the output nodes. The convergence of the Wang's kWTA under these two defects is analyzed and the corresponding energy function is revealed.
John Sum, Andrew Chi-Sing Leung, Kevin I.-J. Ho
IEEE Trans. Neural Networks Learn. Syst.1
2012 On the Objective Function and Learning Algorithm for Concurrent Open Node Fault
Andrew Chi-Sing Leung, John Sum, Kai Tat Ng
ICONIP (3)2
2012 Optimization of tuning parameters for open node fault regularizer
Andrew Chi-Sing Leung, John Sum, Yuxin Liu 0008
Neurocomputing2
2012 RBF Networks Under the Concurrent Fault Situation
abstract
Fault tolerance is an interesting topic in neural networks. However, many existing results on this topic focus only on the situation of a single fault source. In fact, a trained network may be affected by multiple fault sources. This brief studies the performance of faulty radial basis function (RBF) networks that suffer from multiplicative weight noise and open weight fault concurrently. We derive a mean prediction error (MPE) formula to estimate the generalization ability of faulty networks. The MPE formula provides us a way to understand the generalization ability of faulty networks without using a test set or generating a number of potential faulty networks. Based on the MPE result, we propose methods to optimize the regularization parameter, as well as the RBF width.
Andrew Chi-Sing Leung, John Sum
IEEE Trans. Neural Networks Learn. Syst.2
2012 On-Line Node Fault Injection Training Algorithm for MLP Networks: Objective Function and Convergence Analysis
abstract
Improving fault tolerance of a neural network has been studied for more than two decades. Various training algorithms have been proposed in sequel. The on-line node fault injection-based algorithm is one of these algorithms, in which hidden nodes randomly output zeros during training. While the idea is simple, theoretical analyses on this algorithm are far from complete. This paper presents its objective function and the convergence proof. We consider three cases for multilayer perceptrons (MLPs). They are: (1) MLPs with single linear output node; (2) MLPs with multiple linear output nodes; and (3) MLPs with single sigmoid output node. For the convergence proof, we show that the algorithm converges with probability one. For the objective function, we show that the corresponding objective functions of cases (1) and (2) are of the same form. They both consist of a mean square errors term, a regularizer term, and a weight decay term. For case (3), the objective function is slight different from that of cases (1) and (2). With the objective functions derived, we can compare the similarities and differences among various algorithms and various cases.
John Sum, Andrew Chi-Sing Leung, Kevin I.-J. Ho
IEEE Trans. Neural Networks Learn. Syst.1
2012 Convergence Analyses on On-Line Weight Noise Injection-Based Training Algorithms for MLPs
abstract
Injecting weight noise during training is a simple technique that has been proposed for almost two decades. However, little is known about its convergence behavior. This paper studies the convergence of two weight noise injection-based training algorithms, multiplicative weight noise injection with weight decay and additive weight noise injection with weight decay. We consider that they are applied to multilayer perceptrons either with linear or sigmoid output nodes. Let w(t) be the weight vector, let V(w) be the corresponding objective function of the training algorithm, let α >; 0 be the weight decay constant, and let μ(t) be the step size. We show that if μ(t)→ 0, then with probability one E[||w(t)||2(2)] is bound and lim(t) → ∞ ||w(t)||2 exists. Based on these two properties, we show that if μ(t)→ 0, Σtμ(t)=∞, and Σtμ(t)(2) <; ∞, then with probability one these algorithms converge. Moreover, w(t) converges with probability one to a point where ∇wV(w)=0.
John Sum, Andrew Chi-Sing Leung, Kevin I.-J. Ho
IEEE Trans. Neural Networks Learn. Syst.1
2012 Analysis on the Convergence Time of Dual Neural Network-Based WTA
abstract
A k-winner-take-all (kWTA) network is able to find out the k largest numbers from n inputs. Recently, a dual neural network (DNN) approach was proposed to implement the kWTA process. Compared to the conventional approach, the DNN approach has much less number of interconnections. A rough upper bound on the convergence time of the DNN-kWTA model, which is expressed in terms of input variables, was given. This brief derives the exact convergence time of the DNN-kWTA model. With our result, we can study the convergence time without spending excessive time to simulate the network dynamics. We also theoretically study the statistical properties of the convergence time when the inputs are uniformly distributed. Since a nonuniform distribution can be converted into a uniform one and the conversion preserves the ordering of the inputs, our theoretical result is also valid for nonuniformly distributed inputs.
Yi Xiao 0004, Yuxin Liu 0008, Andrew Chi-Sing Leung, John Sum, Kevin I.-J. Ho
IEEE Trans. Neural Networks Learn. Syst.4
2011 Regularizer for Co-existing of Open Weight Fault and Multiplicative Weight Noise
Andrew Chi-Sing Leung, John Sum
ICONIP (3)2
2011 Recovery of Sparse Signal from an Analog Network Model
Andrew Chi-Sing Leung, John Sum, Ping-Man Lam, Anthony G. Constantinides
ICONIP (3)2
2011 Analysis on Wang's kWTA with Stochastic Output Nodes
John Sum, Andrew Chi-Sing Leung, Kevin I.-J. Ho
ICONIP (3)1
2011 Training RBF network to tolerate single node fault
Kevin I.-J. Ho, Andrew Chi-Sing Leung, John Sum
Neurocomputing3
2011 Regularizers for fault tolerant multilayer feedforward networks
Shue Kwan Mak, John Sum, Andrew Chi-Sing Leung
Neurocomputing2
2011 Guest editorial: special issue on the emerging applications of neural networks
Tommy W. S. Chow, John Sum
Neural Comput. Appl.2
2011 The effect of weight fault on associative networks
Andrew Chi-Sing Leung, John Sum, Kevin I.-J. Ho
Neural Comput. Appl.2
2011 Objective Functions of Online Weight Noise Injection Training Algorithms for MLPs
abstract
Injecting weight noise during training has been a simple strategy to improve the fault tolerance of multilayer perceptrons (MLPs) for almost two decades, and several online training algorithms have been proposed in this regard. However, there are some misconceptions about the objective functions being minimized by these algorithms. Some existing results misinterpret that the prediction error of a trained MLP affected by weight noise is equivalent to the objective function of a weight noise injection algorithm. In this brief, we would like to clarify these misconceptions. Two weight noise injection scenarios will be considered: one is based on additive weight noise injection and the other is based on multiplicative weight noise injection. To avoid the misconceptions, we use their mean updating equations to analyze the objective functions. For injecting additive weight noise during training, we show that the true objective function is identical to the prediction error of a faulty MLP whose weights are affected by additive weight noise. It consists of the conventional mean square error and a smoothing regularizer. For injecting multiplicative weight noise during training, we show that the objective function is different from the prediction error of a faulty MLP whose weights are affected by multiplicative weight noise. With our results, some existing misconceptions regarding MLP training with weight noise injection can now be resolved.
Kevin I.-J. Ho, Andrew Chi-Sing Leung, John Sum
IEEE Trans. Neural Networks3
2010 Lagrange Programming Neural Networks for Compressive Sampling
Ping-Man Lam, Andrew Chi-Sing Leung, John Sum, Anthony G. Constantinides
ICONIP (2)3
2010 Generalization Error of Faulty MLPs with Weight Decay Regularizer
Andrew Chi-Sing Leung, John Sum, Shue Kwan Mak
ICONIP (2)2
2010 Kernel Width Optimization for Faulty RBF Neural Networks with Multi-node Open Fault
Hongjiang Wang, Andrew Chi-Sing Leung, John Sum
Neural Process. Lett.3
2010 Convergence and objective functions of some fault/noise-injection-based online learning algorithms for RBF networks
abstract
In the last two decades, many online fault/noise injection algorithms have been developed to attain a fault tolerant neural network. However, not much theoretical works related to their convergence and objective functions have been reported. This paper studies six common fault/noise-injection-based online learning algorithms for radial basis function (RBF) networks, namely 1) injecting additive input noise, 2) injecting additive/multiplicative weight noise, 3) injecting multiplicative node noise, 4) injecting multiweight fault (random disconnection of weights), 5) injecting multinode fault during training, and 6) weight decay with injecting multinode fault. Based on the Gladyshev theorem, we show that the convergence of these six online algorithms is almost sure. Moreover, their true objective functions being minimized are derived. For injecting additive input noise during training, the objective function is identical to that of the Tikhonov regularizer approach. For injecting additive/multiplicative weight noise during training, the objective function is the simple mean square training error. Thus, injecting additive/multiplicative weight noise during training cannot improve the fault tolerance of an RBF network. Similar to injective additive input noise, the objective functions of other fault/noise-injection-based online algorithms contain a mean square error term and a specialized regularization term.
Kevin I.-J. Ho, Andrew Chi-Sing Leung, John Sum
IEEE Trans. Neural Networks3
2010 On the selection of weight decay parameter for faulty networks
abstract
The weight-decay technique is an effective approach to handle overfitting and weight fault. For fault-free networks, without an appropriate value of decay parameter, the trained network is either overfitted or underfitted. However, many existing results on the selection of decay parameter focus on fault-free networks only. It is well known that the weight-decay method can also suppress the effect of weight fault. For the faulty case, using a test set to select the decay parameter is not practice because there are huge number of possible faulty networks for a trained network. This paper develops two mean prediction error (MPE) formulae for predicting the performance of faulty radial basis function (RBF) networks. Two fault models, multiplicative weight noise and open weight fault, are considered. Our MPE formulae involve the training error and trained weights only. Besides, in our method, we do not need to generate a huge number of faulty networks to measure the test error for the fault situation. The MPE formulae allow us to select appropriate values of decay parameter for faulty networks. Our experiments showed that, although there are small differences between the true test errors (from the test set) and the MPE values, the MPE formulae can accurately locate the appropriate value of the decay parameter for minimizing the true test error of faulty networks.
Andrew Chi-Sing Leung, Hongjiang Wang, John Sum
IEEE Trans. Neural Networks3
2009 Fault Tolerant Regularizers for Multilayer Feedforward Networks
Deng-yu Qiao, Andrew Chi-Sing Leung, John Sum
ICONIP (1)3
2009 SNIWD: Simultaneous Weight Noise Injection with Weight Decay for MLP Training
John Sum, Kevin I.-J. Ho
ICONIP (1)1
2009 On Objective Function, Regularizer, and Prediction Error of a Learning Algorithm for Dealing With Multiplicative Weight Noise
abstract
In this paper, an objective function for training a functional link network to tolerate multiplicative weight noise is presented. Basically, the objective function is similar in form to other regularizer-based functions that consist of a mean square training error term and a regularizer term. Our study shows that under some mild conditions the derived regularizer is essentially the same as a weight decay regularizer. This explains why applying weight decay can also improve the fault-tolerant ability of a radial basis function (RBF) with multiplicative weight noise. In accordance with the objective function, a simple learning algorithm for a functional link network with multiplicative weight noise is derived. Finally, the mean prediction error of the trained network is analyzed. Simulated experiments on two artificial data sets and a real-world application are performed to verify theoretical result.
John Sum, Andrew Chi-Sing Leung, Kevin I.-J. Ho
IEEE Trans. Neural Networks1
2008 On Weight-Noise-Injection Training
Kevin I.-J. Ho, Andrew Chi-Sing Leung, John Sum
ICONIP (2)3
2008 Analysis on Generalization Error of Faulty RBF Networks with Weight Decay Regularizer
Andrew Chi-Sing Leung, John Sum, Hongjiang Wang
ICONIP (2)2
2008 On Node-Fault-Injection Training of an RBF Network
John Sum, Andrew Chi-Sing Leung, Kevin I.-J. Ho
ICONIP (2)1
2008 Prediction error of a fault tolerant neural network
John Sum, Andrew Chi-Sing Leung
Neurocomputing1
2008 Improved transmission of vector quantized data over noisy channels
Andrew Chi-Sing Leung, John Sum, Herbert Chan
Neural Comput. Appl.2
2008 A Fault-Tolerant Regularizer for RBF Networks
abstract
In classical training methods for node open fault, we need to consider many potential faulty networks. When the multinode fault situation is considered, the space of potential faulty networks is very large. Hence, the objective function and the corresponding learning algorithm would be computationally complicated. This paper uses the Kullback-Leibler divergence to define an objective function for improving the fault tolerance of radial basis function (RBF) networks. With the assumption that there is a Gaussian distributed noise term in the output data, a regularizer in the objective function is identified. Finally, the corresponding learning algorithm is developed. In our approach, the objective function and the learning algorithm are computationally simple. Compared with some conventional approaches, including weight-decay-based regularizers, our approach has a better fault-tolerant ability. Besides, our empirical study shows that our approach can improve the generalization ability of a fault-free RBF network.
Andrew Chi-Sing Leung, John Sum
IEEE Trans. Neural Networks2
2007 Analysis on Bidirectional Associative Memories with Multiplicative Weight Noise
Andrew Chi-Sing Leung, John Sum, Tien-Tsin Wong
ICONIP (1)2
2006 Pricing Web Services
Kevin I.-J. Ho, John Sum, Gilbert H. Young
GPC2
2006 Prediction Error of a Fault Tolerant Neural Network
John Sum, Andrew Chi-Sing Leung, Kevin I.-J. Ho
ICONIP (1)1
2006 On-line estimation of the final prediction error via recursive least-squares method
John Sum, Kevin I.-J. Ho
Neurocomputing1
2003 Two alternative soft self-organizing maps: their algorithms and applications
abstract
This paper presents two SOM-like algorithms that are extended from two alternative soft competition algorithms namely maximum likelihood competitive learning (MLCL) and fuzzy competitive learning (FCL). Simulation results on the topographic map formation are presented and a possible application of such algorithms for data transmission is elucidated. It is observed that under certain circumstances, the performance of these SSOM algorithms in a vowel data transmission problem can be comparable to and sometimes even better than that of using SOM.
John Sum
SMC1
2003 Analysis on Extended Ant Routing Algorithms for Network Routing and Management
John Sum, Hong Shen 0001, Gilbert H. Young, Jie Wu 0001, Andrew Chi-Sing Leung
J. Supercomput.1
2003 Analysis on a Mobile Agent-Based Algorithm for Network Routing and Management
abstract
Ant routing is a method for network routing in agent technology. Although its effectiveness and efficiency have been demonstrated and reported in the literature, its properties have not yet been well studied. This paper presents some preliminary analysis on an ant algorithm in regard to its population growing property and jumping behavior. Results conclude that as long as the value max, {i/spl Omega//sub j/|} is known, the practitioner is able to design the algorithm parameters, such as the number of agents being created for each request, k, and the maximum allowable number of jumps of an agent, in order to meet the network constraint.
John Sum, Hong Shen 0001, Andrew Chi-Sing Leung, Gilbert H. Young
IEEE Trans. Parallel Distributed Syst.1
2001 A pruning method for the recursive least squared algorithm
Andrew Chi-Sing Leung, Kwok-Wo Wong, John Sum, Lai-Wan Chan
Neural Networks3
2000 A Local Training and Pruning Approach for Neural Networks
abstract
The training of neural networks using the extended Kalman filter (EKF) algorithm is plagued by the drawback of high computational complexity and storage requirement that may become prohibitive even for networks of moderate size. In this paper, we present a local EKF training and pruning approach that can solve this problem. In particular, the by-products obtained along with the local EKF training can be utilized to measure the importance of the network weights. Comparing with the original global approach, the proposed local EKF training and pruning approach results in a much lower computational complexity and storage requirement. Hence, it is more practical in solving real world problems. The performance of the proposed algorithm is demonstrated on one medium- and one large-scale problems, namely, sunspot data prediction and handwritten digit recognition.
Sheng-Jiang Chang, Andrew Chi-Sing Leung, Kwok-Wo Wong, John Sum
Int. J. Neural Syst.4
1999 A Note on the Equivalence of NARX and RNN
John Sum, Wing-Kay Kan, Gilbert H. Young
Neural Comput. Appl.1
1999 An Adaptive Bayesian Pruning for Neural Networks in a Non-Stationary Environment
abstract
Pruning a neural network to a reasonable smaller size, and if possible to give a better generalization, has long been investigated. Conventionally the common technique of pruning is based on considering error sensitivity measure, and the nature of the problem being solved is usually stationary. In this article, we present an adaptive pruning algorithm for use in a nonstationary environment. The idea relies on the use of the extended Kalman filter (EKF) training method. Since EKF is a recursive Bayesian algorithm, we define a weight-importance measure in term of the sensitivity of a posteriori probability. Making use of this new measure and the adaptive nature of EKF, we devise an adaptive pruning algorithm called adaptive Bayesian pruning. Simulation results indicate that in a noisy nonstationary environment, the proposed pruning algorithm is able to remove network redundancy adaptively and yet preserve the same generalization ability.
John Sum, Andrew Chi-Sing Leung, Gilbert H. Young, Lai-Wan Chan, Wing-Kay Kan
Neural Comput.1
1999 On the regularization of forgetting recursive least square
abstract
In this paper, the regularization of employing the forgetting recursive least square (FRLS) training technique on feedforward neural networks is studied. We derive our result from the corresponding equations for the expected prediction error and the expected training error. By comparing these error equations with other equations obtained previously from the weight decay method, we have found that the FRLS technique has an effect which is identical to that of using the simple weight decay method. This new finding suggests that the FRLS technique is another on-line approach for the realization of the weight decay effect. Besides, we have shown that, under certain conditions, both the model complexity and the expected prediction error of the model being trained by the FRLS technique are better than the one trained by the standard RLS method.
Andrew Chi-Sing Leung, Gilbert H. Young, John Sum, Wing-Kay Kan
IEEE Trans. Neural Networks3
1999 Analysis for a class of winner-take-all model
abstract
Recently we have proposed a simple circuit of winner-take-all (WTA) neural network. Assuming no external input, we have derived an analytic equation for its network response time. In this paper, we further analyze the network response time for a class of winner-take-all circuits involving self-decay and show that the network response time of such a class of WTA is the same as that of the simple WTA model.
John Sum, Andrew Chi-Sing Leung, Peter Kwong-Shun Tam, Gilbert H. Young, Wing-Kay Kan, Lai-Wan Chan
IEEE Trans. Neural Networks1
1999 On the Kalman filtering method in neural network training and pruning
abstract
In the use of extended Kalman filter approach in training and pruning a feedforward neural network, one usually encounters the problems on how to set the initial condition and how to use the result obtained to prune a neural network. In this paper, some cues on the setting of the initial condition will be presented with a simple example illustrated. Then based on three assumptions--1) the size of training set is large enough; 2) the training is able to converge; and 3) the trained network model is close to the actual one, an elegant equation linking the error sensitivity measure (the saliency) and the result obtained via extended Kalman filter is devised. The validity of the devised equation is then testified by a simulated example.
John Sum, Andrew Chi-Sing Leung, Gilbert H. Young, Wing-Kay Kan
IEEE Trans. Neural Networks1
1998 Extended Kalman Filter-Based Pruning Method for Recurrent Neural Networks
abstract
Pruning is one of the effective techniques for improving the generalization error of neural networks. Existing pruning techniques are derived mainly from the viewpoint of energy minimization, which is commonly used in gradient-based learning methods. In recurrent networks, extended Kalman filter (EKF)-based training has been shown to be superior to gradient-based learning methods in terms of speed. This article explains a pruning procedure for recurrent neural networks using EKF training. The sensitivity of a posterior probability is used as a measure of the importance of a weight instead of error sensitivity since posterior probability density is readily obtained from this training method. The pruning procedure is tested using three problems: (1) the prediction of a simple linear time series, (2) the identification of a nonlinear system, and (3) the prediction of an exchange-rate time series. Simulation results demonstrate that the proposed pruning method is able to reduce the number of parameters and improve the generalization ability of a recurrent network.
John Sum, Lai-Wan Chan, Andrew Chi-Sing Leung, Gilbert H. Young
Neural Comput.1
1997 Yet another algorithm which can generate topography map
abstract
This paper presents an algorithm to form a topographic map resembling to the self-organizing map. The idea stems on defining an energy function which reveals the local correlation between neighboring neurons. The larger the value of the energy function, the higher the correlation of the neighborhood neurons. On this account, the proposed algorithm is defined as the gradient ascent of this energy function. Simulations on two-dimensional maps are illustrated.
John Sum, Andrew Chi-Sing Leung, Lai-Wan Chan, Lei Xu 0001
IEEE Trans. Neural Networks1
1996 Attraction Basin of Bidirectional Associative Memories
Andrew Chi-Sing Leung, Lai-Wan Chan, John Sum
Int. J. Neural Syst.3
1996 Note on the Maxnet Dynamics
abstract
A simple method is presented to derive the complete solution of the Maxnet network dynamics. Besides, the exact response time of the network is deduced.
John Sum, Peter Kwong-Shun Tam
Neural Comput.1