Dingli Yu

dblp:39/578 · also D. L. Yu, Ding-Li Yu · DBLP profile ↗
← Back
47ranked-venue papers
14as first author
14since 2021 · last 2025
0000-0003-3177-8942ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 44 · 13 first-author · 13 since 2021Theory of computation · 2 · 1 since 2021Databases, data management, data science and information retrieval · 1Graphics, computer vision, multimedia, augmented reality and games · 1Human-computer interaction and ubiquitous computing · 1 · 1 first-author
YearPublicationVenuePosition
2025 Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs?
abstract
Vision Language Models (VLMs) are impressive at visual question answering and image captioning. But they underperform on multi-step visual reasoning—even compared to LLMs on the same tasks presented in text form—giving rise to perceptions of modality imbalance or brittleness. Towards a systematic study of such issues, we introduce a synthetic framework for assessing the ability of VLMs to perform algorithmic visual reasoning, comprising three tasks: Table Readout, Grid Navigation, and Visual Analogy. Each has two levels of difficulty, SIMPLE and HARD, and even the SIMPLE versions are difficult for frontier VLMs. We propose strategies for training on the SIMPLE version of tasks that improve performance on the corresponding HARD task, i.e., simple-to-hard (S2H) generalization. This controlled setup, where each task also has an equivalent text-only version, allows a quantification of the modality imbalance and how it is impacted by training strategy. We show that 1) explicit image-to-text conversion is important in promoting S2H generalization on images, by transferring reasoning from text; 2) conversion can be internalized at test time. We also report results of mechanistic study of this phenomenon. We identify measures of gradient alignment that can identify training strategies that promote better S2H generalization. Ablations highlight the importance of chain-of-thought.
Simon Park 0002, Abhishek Panigrahi, Dingli Yu, Anirudh Goyal, Sanjeev Arora
ICML4
2025 Weak-to-Strong Generalization Even in Random Feature Networks, Provably
abstract
Weak-to-Strong Generalization (Burns et al.,2024) is the phenomenon whereby a strong student, say GPT-4, learns a task from a weak teacher, say GPT-2, and ends up significantly outperforming the teacher. We show that this phenomenon does not require a complex and pretrained learner like GPT-4, can arise even in simple non-pretrained models, simply due to the size advantage of the student. But, we also show that there are inherint limits to the extent of such weak to strong generalization. We consider students and teachers that are random feature models, described by two-layer networks with a random and fixed bottom layer and trained top layer. A ‘weak’ teacher, with a small number of units (i.e. random features), is trained on the population, and a ‘strong’ student, with a much larger number of units (i.e. random features), is trained only on labels generated by the weak teacher. We demonstrate, prove, and understand how the student can outperform the teacher, even though trained only on data labeled by the teacher. We also explain how such weak-to-strong generalization is enabled by early stopping. We then show the quantitative limits of weak-to-strong generalization in this model, and in fact in a much broader class of models, for arbitrary teacher and student feature spaces and a broad class of learning rules, including when the student features are pre-trained or otherwise more informative. In particular, we show that in such models the student’s error can only approach zero if the teacher’s error approaches zero, and a strong student cannot “boost” a slightly-better-then-chance teacher to obtain a small error.
Marko Medvedev, Kaifeng Lyu, Dingli Yu, Sanjeev Arora, Zhiyuan Li 0005, Nathan Srebro
ICML3
2025 Handling data heterogeneity for wind turbine fault diagnosis via dynamic ensemble multilevel interactive learning
Shuangxin Wang, Hongrui Li, Jiading Jiang, Junmei Ou, Dingli Yu
Eng. Appl. Artif. Intell.6
2024 Tensor Programs VI: Feature Learning in Infinite Depth Neural Networks
abstract
Empirical studies have consistently demonstrated that increasing the size of neural networks often yields superior performance in practical applications. However, there is a lack of consensus regarding the appropriate scaling strategy, particularly when it comes to increasing the depth of neural networks. In practice, excessively large depths can lead to model performance degradation. In this paper, we introduce Depth-$\mu$P, a principled approach for depth scaling, allowing for the training of arbitrarily deep architectures while maximizing feature learning and diversity among nearby layers. Our method involves dividing the contribution of each residual block and the parameter update by the square root of the depth. Through the use of Tensor Programs, we rigorously establish the existence of a limit for infinitely deep neural networks under the proposed scaling scheme. This scaling strategy ensures more stable training for deep neural networks and guarantees the transferability of hyperparameters from shallow to deep models. To substantiate the efficacy of our scaling method, we conduct empirical validation on neural networks with depths up to $2^{10}$.
Greg Yang, Dingli Yu, Soufiane Hayou
ICLR2
2024 SKILL-MIX: a Flexible and Expandable Family of Evaluations for AI Models
abstract
With LLMs shifting their role from statistical modeling of language to serving as general-purpose AI agents, how should LLM evaluations change? Arguably, a key ability of an AI agent is to flexibly combine, as needed, the basic skills it has learned. The capability to combine skills plays an important role in (human) pedagogy and also in a paper on emergence phenomena (Arora & Goyal, 2023). This work introduces SKILL-MIX, a new evaluation to measure ability to combine skills. Using a list of $N$ skills the evaluator repeatedly picks random subsets of $k$ skills and asks the LLM to produce text combining that subset of skills. Since the number of subsets grows like $N^k$, for even modest $k$ this evaluation will, with high probability, require the LLM to produce text significantly different from any text in the training set. The paper develops a methodology for (a) designing and administering such an evaluation, and (b) automatic grading (plus spot-checking by humans) of the results using GPT-4 as well as the open LLaMA-2 70B model. Administering a version of SKILL-MIX to popular chatbots gave results that, while generally in line with prior expectations, contained surprises. Sizeable differences exist among model capabilities that are not captured by their ranking on popular LLM leaderboards ("cramming for the leaderboard"). Furthermore, simple probability calculations indicate that GPT-4's reasonable performance on $k=5$ is suggestive of going beyond "stochastic parrot" behavior (Bender et al., 2021), i.e., it combines skills in ways that it had not seen during training. We sketch how the methodology can lead to a SKILL-MIX based eco-system of open evaluations for AI capabilities of future models. We maintain a leaderboard of SKILL-MIX at [https://skill-mix.github.io](https://skill-mix.github.io).
Dingli Yu, Simran Kaur 0001, Arushi Gupta, Jonah Brown-Cohen, Anirudh Goyal, Sanjeev Arora
ICLR1
2024 Keeping LLMs Aligned After Fine-tuning: The Crucial Role of Prompt Templates
abstract
Public LLMs such as the Llama 2-Chat underwent alignment training and were considered safe. Recently Qi et al. (2024) reported that even benign fine-tuning on seemingly safe datasets can give rise to unsafe behaviors in the models. The current paper is about methods and best practices to mitigate such loss of alignment. We focus on the setting where a public model is fine-tuned before serving users for specific usage, where the model should improve on the downstream task while maintaining alignment. Through extensive experiments on several chat models (Meta's Llama 2-Chat, Mistral AI's Mistral 7B Instruct v0.2, and OpenAI's GPT-3.5 Turbo), this paper uncovers that the prompt templates used during fine-tuning and inference play a crucial role in preserving safety alignment, and proposes the “Pure Tuning, Safe Testing” (PTST) strategy --- fine-tune models without a safety prompt, but include it at test time. This seemingly counterintuitive strategy incorporates an intended distribution shift to encourage alignment preservation. Fine-tuning experiments on GSM8K, ChatDoctor, and OpenOrca show that PTST significantly reduces the rise of unsafe behaviors.
Kaifeng Lyu, Xinran Gu, Dingli Yu, Anirudh Goyal, Sanjeev Arora
NeurIPS4
2024 ConceptMix: A Compositional Image Generation Benchmark with Controllable Difficulty
abstract
Compositionality is a critical capability in Text-to-Image (T2I) models, as it reflects their ability to understand and combine multiple concepts from text descriptions. Existing evaluations of compositional capability rely heavily on human-designed text prompts or fixed templates, limiting their diversity and complexity, and yielding low discriminative power. We propose ConceptMix, a scalable, controllable, and customizable benchmark which automatically evaluates compositional generation ability of T2I models. This is done in two stages. First, ConceptMix generates the text prompts: concretely, using categories of visual concepts (e.g., objects, colors, shapes, spatial relationships), it randomly samples an object and k-tuples of visual concepts, then uses GPT-4o to generate text prompts for image generation based on these sampled concepts. Second, ConceptMix evaluates the images generated in response to these prompts: concretely, it checks how many of the k concepts actually appeared in the image by generating one question per visual concept and using a strong VLM to answer them. Through administering ConceptMix to a diverse set of T2I models (proprietary as well as open ones) using increasing values of k, we show that our ConceptMix has higher discrimination power than earlier benchmarks. Specifically, ConceptMix reveals that the performance of several models, especially open models, drops dramatically with increased k. Importantly, it also provides insight into the lack of prompt diversity in widely-used training datasets. Additionally, we conduct extensive human studies to validate the design of ConceptMix and compare our automatic grading with human judgement. We hope it will guide future T2I model development.
Xindi Wu, Dingli Yu, Yangsibo Huang, Olga Russakovsky, Sanjeev Arora
NeurIPS2
2024 Can Models Learn Skill Composition from Examples?
abstract
As large language models (LLMs) become increasingly advanced, their ability to exhibit compositional generalization---the capacity to combine learned skills in novel ways not encountered during training---has garnered significant attention. This type of generalization, particularly in scenarios beyond training data, is also of great interest in the study of AI safety and alignment. A recent study introduced the Skill-Mix evaluation, where models are tasked with composing a short paragraph demonstrating the use of a specified $k$-tuple of language skills. While small models struggled with composing even with $k=3$, larger models like GPT-4 performed reasonably well with $k=5$ and $6$. In this paper, we employ a setup akin to Skill-Mix to evaluate the capacity of smaller models to learn compositional generalization from examples. Utilizing a diverse set of language skills---including rhetorical, literary, reasoning, theory of mind, and common sense---GPT was used to generate text samples that exhibit random subsets of $k$ skills. Subsequent fine-tuning of 7B and 13B parameter models on these combined skill texts, for increasing values of $k$, revealed the following findings: (1) Training on combinations of $k=2$ and $3$ skills results in noticeable improvements in the ability to compose texts with $k=4$ and $5$ skills, despite models never having seen such examples during training. (2) When skill categories are split into training and held-out groups, models significantly improve at composing texts with held-out skills during testing despite having only seen training skills during fine-tuning, illustrating the efficacy of the training approach even with previously unseen skills. This study also suggests that incorporating skill-rich (potentially synthetic) text into training can substantially enhance the compositional capabilities of models.
Simran Kaur 0001, Dingli Yu, Anirudh Goyal, Sanjeev Arora
NeurIPS3
2024 Augmenting the diversity of imbalanced datasets via multi-vector stochastic exploration oversampling
Hongrui Li, Shuangxin Wang, Jiading Jiang, Chuiyi Deng, Junmei Ou, Ziang Zhou, Dingli Yu
Neurocomputing7
2023 A Kernel-Based View of Language Model Fine-Tuning
abstract
It has become standard to solve NLP tasks by fine-tuning pre-trained language models (LMs), especially in low-data settings. There is minimal theoretical understanding of empirical success, e.g., why fine-tuning a model with $10^8$ or more parameters on a couple dozen training points does not result in overfitting. We investigate whether the Neural Tangent Kernel (NTK)---which originated as a model to study the gradient descent dynamics of infinitely wide networks with suitable random initialization---describes fine-tuning of pre-trained LMs. This study was inspired by the decent performance of NTK for computer vision tasks (Wei et al., 2022). We extend the NTK formalism to Adam and use Tensor Programs (Yang, 2020) to characterize conditions under which the NTK lens may describe fine-tuning updates to pre-trained language models. Extensive experiments on 14 NLP tasks validate our theory and show that formulating the downstream task as a masked word prediction problem through prompting often induces kernel-based dynamics during fine-tuning. Finally, we use this kernel view to propose an explanation for the success of parameter-efficient subspace-based fine-tuning methods.
Sadhika Malladi, Alexander Wettig, Dingli Yu, Danqi Chen 0001, Sanjeev Arora
ICML3
2022 Fast Mixing of Stochastic Gradient Descent with Normalization and Weight Decay
abstract
We prove the Fast Equilibrium Conjecture proposed by Li et al., (2020), i.e., stochastic gradient descent (SGD) on a scale-invariant loss (e.g., using networks with various normalization schemes) with learning rate $\eta$ and weight decay factor $\lambda$ mixes in function space in $\mathcal{\tilde{O}}(\frac{1}{\lambda\eta})$ steps, under two standard assumptions: (1) the noise covariance matrix is non-degenerate and (2) the minimizers of the loss form a connected, compact and analytic manifold. The analysis uses the framework of Li et al., (2021) and shows that for every $T>0$, the iterates of SGD with learning rate $\eta$ and weight decay factor $\lambda$ on the scale-invariant loss converge in distribution in $\Theta\left(\eta^{-1}\lambda^{-1}(T+\ln(\lambda/\eta))\right)$ iterations as $\eta\lambda\to 0$ while satisfying $\eta \le O(\lambda)\le O(1)$. Moreover, the evolution of the limiting distribution can be described by a stochastic differential equation that mixes to the same equilibrium distribution for every initialization around the manifold of minimizers as $T\to\infty$.
Zhiyuan Li 0005, Tianhao Wang 0017, Dingli Yu
NeurIPS3
2022 New Definitions and Evaluations for Saliency Methods: Staying Intrinsic, Complete and Sound
abstract
Saliency methods compute heat maps that highlight portions of an input that were most important for the label assigned to it by a deep net. Evaluations of saliency methods convert this heat map into a new masked input by retaining the $k$ highest-ranked pixels of the original input and replacing the rest with "uninformative" pixels, and checking if the net's output is mostly unchanged. This is usually seen as an explanation of the output, but the current paper highlights reasons why this inference of causality may be suspect. Inspired by logic concepts of completeness & soundness, it observes that the above type of evaluation focuses on completeness of the explanation, but ignores soundness. New evaluation metrics are introduced to capture both notions, while staying in an intrinsic framework---i.e., using the dataset and the net, but no separately trained nets, human evaluations, etc. A simple saliency method is described that matches or outperforms prior methods in the evaluations. Experiments also suggest new intrinsic justifications, based on soundness, for popular heuristic tricks such as TV regularization and upsampling.
Arushi Gupta, Nikunj Saunshi, Dingli Yu, Kaifeng Lyu, Sanjeev Arora
NeurIPS3
2021 Automatic image detection of multi-type surface defects on wind turbine blades based on cascade deep learning network
abstract
A safe operation protocol of the wind blades is a critical factor to ensure the stability of a wind turbine. Sensors are most commonly applied for defect detection on wind turbine blades (WTBs). However, due to the high cost and the sensitivity to stochastic noise, computer vision-guided automatic detection remains a challenge for surface defect detection on WTBs in particularly, its accuracy in locating defects is yet to be optimized. In this paper, we developed a visual inspection model that can automatically and precisely classify and locate the surface defects, through the utilization of a deep learning framework based on the Cascade R-CNN. In order to obtain high mean average precision (mAP) according to the characteristics of the dataset, a model named Contextual Aligned-Deformable Cascade R-CNN (CAD Cascade R-CNN) using improved strategies of transfer learning, Deformable Convolution and Deformable RoI Align, as well as context information fusion is proposed and a dataset with surface defects categorized and labeled as crack, breakage and oil pollution is generated. Moreover to alleviate the problem of false detection under a complex background, an improved bisecting k-means is presented during the test process. The adaptability and generalization of the proposed CAD Cascade R-CNN model were validated by each type of defects in dataset and different IoU thresholds, whereas, each of the above improved strategies was verified by gradual ablation experiments. Finally experiments that compared with the baseline Cascade R-CNN, Faster R-CNN and YOLO-v3 demonstrate its superiority over these existing approaches with a maximum of 92.1% mAP.
Yulin Mao, Shuangxin Wang, Dingli Yu, Juchao Zhao
Intell. Data Anal.3
2021 A New Concept of Fractional Order Cumulant and It-Based Signal Processing in α and/or Gaussian Noise
abstract
In this article, the concept and definitions of the Fractional Order Moment (FOM) and Fractional Order Cumulant (FOC) are proposed, which is based on the fractional derivative of the fractional order Moment-generating function and the fractional order Cumulant-generating function of stochastic processes. The moment and cumulant are defined on an expanded set from positive integer to the whole positive real. This development not only provides a new technology for signal processing, also complements the existing theory in the field. The properties of the FOC have been derived, and their uniformity and particularity with the High Order Cumulant are compared and commented. In addition, the transformation between the FOM and the FOC are derived and discussed in detail. As one of the applications of the new concept to the α and Gaussian processes, a new method of suppressing α and Gaussian noise is proposed. Furthermore, a FOC-based parameter estimation algorithm is developed for the non-minimum phase ARMA processes in α and/or Gaussian noise. Simulation examples are used to demonstrate the effectiveness of the proposed parameter estimation algorithm.
Yiran Shi, Dingli Yu, Hongyan Shi, Yaowu Shi
IEEE Trans. Inf. Theory2
2020 Harnessing the Power of Infinitely Wide Deep Nets on Small-data Tasks
Sanjeev Arora, Simon S. Du, Zhiyuan Li 0005, Ruslan Salakhutdinov, Ruosong Wang, Dingli Yu
ICLR6
2020 Simple and Effective Regularization Methods for Training on Noisily Labeled Data with Generalization Guarantee
Wei Hu 0014, Zhiyuan Li 0005, Dingli Yu
ICLR3
2020 Characterization of Group-strategyproof Mechanisms for Facility Location in Strictly Convex Space
abstract
We characterize the class of group-strategyproof mechanisms for the single facility location game in any unconstrained strictly convex space. A mechanism is group-strategyproof,if no group of agents can misreport so that all its members are strictlybetter off. A strictly convex space is a normed vector space where |x+y|<2 holds for any pair of different unit vectors x ≠ y, e.g., any Lp space with p∈ (1,∞).
Pingzhong Tang, Dingli Yu, Shengyu Zhao
EC2
2020 Fault tolerant control for nonlinear systems using sliding mode and adaptive neural network estimator
abstract
Abstract This paper proposes a new fault tolerant control scheme for a class of nonlinear systems including robotic systems and aeronautical systems. In this method, a sliding mode control is applied to maintain system stability under the post-fault dynamics. A neural network is used as on-line estimator to reconstruct the change rate of the fault and compensate for the impact of the fault on the system performance. The control law and the neural network learning algorithms are derived using the Lyapunov method, so that the neural estimator is guaranteed to converge to the fault change rate, while the entire closed-loop system stability and tracking control is guaranteed. Compared with the existing methods, the proposed method achieved fault tolerant control for time-varying fault, rather than just constant fault. This greatly expands the industrial applications of the developed method to enhance system reliability. The main contribution and novelty of the developed method is that the system stability is guaranteed and the fault estimation is also guaranteed for convergence when the system subject to a time-varying fault. A simulation example is used to demonstrate the design procedure and the effectiveness of the method. The simulation results demonstrated that the post-fault is stable and the performance is maintained.
Haiying Qi, Yiran Shi, Shoutao Li, Yantao Tian, Dingli Yu, J. B. Gomm
Soft Comput.5
2018 Fair Rent Division on a Budget
abstract
The standard approach to fair rent division assumes that agents have quasi-linear utilities, and seeks allocations that are envy free; it underlies an algorithm that is widely used in practice. However, this approach does not take budget constraints into account, and, therefore, may assign agents to rooms they cannot afford. By contrast, we design a polynomial-time algorithm that takes budget constraints as part of its input; it determines whether there exist envy-free allocations that satisfy the budget constraints, and, if so, computes one that optimizes an additional criterion of justice. In particular, this gives a polynomial-time implementation of the budget-constrained maximin solution, where the maximization objective is the minimum utility of any agent. We show that, like its non-budget-constrained counterpart, this solution is unique in terms of utilities (when it exists), and satisfies additional desirable properties.
Ariel D. Procaccia, Rodrigo A. Velez, Dingli Yu
AAAI3
2018 Development of adaptive p-step RBF network model with recursive orthogonal least squares training
D. K. Siong Tok, Dingli Yu
Neural Comput. Appl.3
2017 Factorized f-step radial basis function model for model predictive control
D. K. Siong Tok, Yiran Shi, Yantao Tian, Dingli Yu
Neurocomputing4
2015 Air-fuel ratio prediction and NMPC for SI engines with modified Volterra model and RBF network
Yiran Shi, Dingli Yu, Yantao Tian, Yaowu Shi
Eng. Appl. Artif. Intell.2
2015 Adaptive structure radial basis function network model for processes with operating region migration
D. K. Siong Tok, Dingli Yu, Christian Mathews, Dong-Ya Zhao
Neurocomputing2
2014 Fault detection and isolation for PEM fuel cell stack with independent RBF model
Mahanijah Md Kamal, D. W. Yu, Dingli Yu
Eng. Appl. Artif. Intell.3
2010 Robust air/fuel ratio control with adaptive DRNN model and AD tuning
Yu-Jia Zhai, Dingwen Yuan, Hong-Yu Guo, Dingli Yu
Eng. Appl. Artif. Intell.4
2009 Neural network model-based automotive engine air/fuel ratio control and robustness evaluation
Yu-Jia Zhai, Dingli Yu
Eng. Appl. Artif. Intell.2
2008 Adaptive RBF network for parameter estimation and stable air-fuel ratio control
Dingli Yu
Neural Networks2
2007 Neural Network in Stable Adaptive Control Law for Automotive Engines
Dingli Yu
ISNN (1)2
2007 An On-Line Learning Algorithm of Parallel Mode for MLPN Models
Dingli Yu, T. K. Chang, D. W. Yu
ISNN (2)1
2007 A Neural Network Model Based MPC of Engine AFR with Single-Dimensional Optimization
Yu-Jia Zhai, Dingli Yu
ISNN (1)2
2007 A new structure adaptation algorithm for RBF networks and its application
Dingli Yu, Dingwen Yuan
Neural Comput. Appl.1
2006 Fault Diagnosis with Enhanced Neural Network Modelling
Dingli Yu, Thoonkhin Chang
ISNN (2)1
2006 Adaptive Pseudo Linear RBF Model for Process Control
Dingwen Yuan, Dingli Yu
ISNN (2)2
2006 Adaptive neural network model based predictive control for air-fuel ratio of SI engines
S. W. Wang, Dingli Yu, J. B. Gomm, G. F. Page, S. S. Douglas
Eng. Appl. Artif. Intell.2
2005 Comparative Study on Engine Torque Modelling Using Different Neural Networks
Dingli Yu, Michael Beham
ISNN (3)1
2005 Fault Tolerant Control of Nonlinear Processes with Adaptive Diagonal Recurrent Neural Network Model
Dingli Yu, Thoonkhin Chang, Jin Wang 0042
ISNN (3)1
2005 Detecting Sensor Faults for a Chemical Reactor Rig via Adaptive Neural Network Model
Dingli Yu, Dingwen Yuan
ISNN (3)1
2005 Adaptive neural model-based fault tolerant control for multi-variable processes
Dingli Yu, T. K. Chang, D. W. Yu
Eng. Appl. Artif. Intell.1
2005 Adaptation of diagonal recurrent neural network model
Dingli Yu, T. Chang
Neural Comput. Appl.1
2005 Fault tolerant control of multivariable processes using auto-tuning PID controller
abstract
Fault tolerant control of dynamic processes is investigated in this paper using an auto-tuning PID controller. A fault tolerant control scheme is proposed composing an auto-tuning PID controller based on an adaptive neural network model. The model is trained online using the extended Kalman filter (EKF) algorithm to learn system post-fault dynamics. Based on this model, the PID controller adjusts its parameters to compensate the effects of the faults, so that the control performance is recovered from degradation. The auto-tuning algorithm for the PID controller is derived with the Lyapunov method and therefore, the model predicted tracking error is guaranteed to converge asymptotically. The method is applied to a simulated two-input two-output continuous stirred tank reactor (CSTR) with various faults, which demonstrate the applicability of the developed scheme to industrial processes.
Dingli Yu, Thoonkhin Chang, Dingwen Yuan
IEEE Trans. Syst. Man Cybern. Part B1
2004 Neural network model adaptation and its application to process control
Thoonkhin Chang, Dingli Yu, Dingwen Yuan
Adv. Eng. Informatics2
2004 A Localized Forgetting Method for Gaussian RBFN Model Adaptation
Dingli Yu
Neural Process. Lett.1
2003 Neural network control of multivariable processes with a fast optimisation algorithm
Dingwen Yuan, Dingli Yu
Neural Comput. Appl.2
2002 Enhanced Neural Network Modelling for a Real Multi-variable Chemical Process
Dingli Yu, J. B. Gomm
Neural Comput. Appl.1
2000 Selecting radial basis function network centers with recursive orthogonal least squares training
abstract
Recursive orthogonal least squares (ROLS) is a numerically robust method for solving for the output layer weights of a radial basis function (RBF) network, and requires less computer memory than the batch alternative. In this paper, the use of ROLS is extended to selecting the centers of an RBF network. It is shown that the information available in an ROLS algorithm after network training can be used to sequentially select centers to minimize the network output error. This provides efficient methods for network reduction to achieve smaller architectures with acceptable accuracy and without retraining. Two selection methods are developed, forward and backward. The methods are illustrated in applications of RBF networks to modeling a nonlinear time series and a real multiinput-multioutput chemical process. The final network models obtained achieve acceptable accuracy with significant reductions in the number of required centers.
J. B. Gomm, Dingli Yu
IEEE Trans. Neural Networks Learn. Syst.2
1997 A Recursive Orthogonal Least Squares Algorithm for Training RBF Networks
Dingli Yu, J. B. Gomm, D. Williams
Neural Process. Lett.1
1996 A Hybrid Fault Diagnosis Approach Using Neural Networks
Dingli Yu, Derek Nicholas Shields, Steve Daley
Neural Comput. Appl.1