EDBT 2026 Demo / reviewers in the wild / expert
Jennie Si
dblp:55/6102
· DBLP profile ↗
85ranked-venue papers
4as first author
21since 2021 · last 2026
0000-0002-0374-7404ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 67 · 4 first-author · 17 since 2021Systems, architecture and hardware · 9 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9Human-computer interaction and ubiquitous computing · 5 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Addressing Human-Robot Symbiosis via Bilevel Optimization of Robotic Knee Prosthesis ControlabstractThis study presents an innovative solution for integrating a human and a robotic knee prosthesis symbiotically for walking. Achieving this requires human-robot cross-joint coordination to provide personalized walking assistance. Our approach uses inverse reinforcement learning (IRL) to identify control objectives for reinforcement learning (RL) controller. Unlike existing methods that optimize performance of human or robot alone, our approach considers both human (thigh segmental angle) and robot (knee joint kinematics) aspects. This bilevel optimization method was evaluated on 3 non-disabled participants and 2 people with amputation. Results showed that the approach personalized the objective function and resulted in a robust policy, completing optimization within a duration of 3.5 minutes. Compared to previous approaches focusing only on robot states, this symbiotic approach increased stance time and step length on the prosthesis side for most participants. Our results highlight the potential of integrating human state into prosthesis control personalization, enhancing the functionality and health of people with amputation. Varun Nalam, Jennie Si, He Huang 0002 |
IEEE Trans. Robotics | 3 |
| 2025 | Integral Performance Approximation for Continuous-Time Reinforcement Learning ControlabstractWe introduce integral performance approximation (IPA), a new continuous-time reinforcement learning (CT-RL) control method. It leverages an affine nonlinear dynamic model, which partially captures the dynamics of the physical environment, alongside state-action trajectory data to enable optimal control with great data efficiency and robust control performance. Utilizing Kleinman algorithm structures allows IPA to provide theoretical guarantees of learning convergence, solution optimality, and closed-loop stability. Furthermore, we demonstrate the effectiveness of IPA on three CT-RL environments including hypersonic vehicle (HSV) control, which has additional challenges caused by unstable and nonminimum phase dynamics. As a result, we demonstrate that the IPA method leads to new, SOTA control design and performance in CT-RL. Brent Wallace, Jennie Si |
ICLR | 2 |
| 2025 | Reinforcement Learning Control of a Physical Robot Device for Assisted Human Walking without a SimulatorabstractThis study presents an innovative reinforcement learning (RL) control approach to facilitate soft exosuit-assisted human walking. Our goal is to address the ongoing challenges in developing reliable RL-based methods for controlling physical devices. To overcome key obstacles—such as limited data, the absence of a simulator for human-robot interaction during walking, the need for low computational overhead in real-time deployment, and the demand for rapid adaptation to achieve personalized control while ensuring human safety—we propose an online Adaptation from an offline Imitating Expert Policy (AIP) approach. Our offline learning mimics human expert actions through real human walking demonstrations without robot assistance. The resulted policy is then used to initialize online actor-critic learning, the goal of which is to optimally personalize robot assistance. In addition to being fast and robust, our online RL method also posses important properties such as learning convergence, dynamic stability, and solution optimality. We have successfully demonstrated our simple
and robust framework for safe robot control on all five tested human participants, without selectively presenting results. The qualitative performance guarantees provided by our online RL, along with the consistent experimental validation of AIP control, represent the first demonstration of online adaptation for softsuit control personalization and serve as important evidence for the use of online RL in controlling a physical device to solve a real-life problem. Junmin Zhong, Emiliano Quiñones Yumbla, Seyed Yousef Soltanian, Jennie Si |
ICML | 6 |
| 2025 | Personalized Reinforcement Learning Control of Soft Robotic Exosuit for Assisting Human Normative Walking with Reduced EffortabstractWearable lower limb robots are promising technologies to assist human locomotion. Soft robotic exosuits introduce a promising solution for reducing muscle effort and metabolic cost as they are lightweight, transparent and inherently safe. However, it is challenging to effectively control such soft robots and personalize the assistance for individual users. With the difficulty in developing robust dynamic model of the human-soft robot system, especially the interacting dynamics between the human and the robot, traditional control methods have seen limited success in addressing these challenges. Reinforcement learning (RL), a data-driven optimal control method, provides a naturally promising alternative. In this study, we propose an innovative control design approach to enable human normative walking with reduced physical effort. To achieve this goal, we propose to first offline learn an exosuit controller for typical human normative walking which is then used in the online phase of control tuning for individual users. Four participants are recruited to test the exosuit controller in treadmill walking. Our results show that online tuning for individual users reaches convergence quickly, typically in one experimental trial due to using an efficient offline pre-trained policy. Furthermore, the RL control of the exosuit results in an average muscle effort reduction of 8.8% and 2.8% for the vastus lateralis and biceps femoris as measured by electromyography (EMG) sensors. These results provide the first evidence of customizing the soft exosuit assistance for individual users. Emiliano Quiñones Yumbla, Junmin Zhong, Seyed Yousef Soltanian, Jennie Si |
IROS | 4 |
| 2025 | Continuous-Time Reinforcement Learning: New Design Algorithms With Theoretical Insights and Performance GuaranteesabstractContinuous-time reinforcement learning (CT-RL) methods hold great promise in real-world applications. Adaptive dynamic programming (ADP)-based CT-RL algorithms, especially their theoretical developments, have achieved great successes. However, these methods have not been demonstrated for solving realistic or meaningful learning control problems. Thus, the goal of this work is to introduce a suite of new excitable integral reinforcement learning (EIRL) algorithms for control of CT affine nonlinear systems. This work develops a new excitation framework to improve persistence of excitation (PE) and numerical performance via input/output insights from classical control. Furthermore, when the system dynamics afford a physically-motivated partition into distinct dynamical loops, the proposed methods break the control problem into smaller subproblems, resulting in reduced complexity. By leveraging the known affine nonlinear dynamics, the methods achieve well-behaved system responses and considerable data efficiency. The work provides convergence, solution optimality, and closed-loop stability guarantees of the proposed methods, and it demonstrates these guarantees on a significant application problem of controlling an unstable, nonminimum phase hypersonic vehicle (HSV). Brent Wallace, Jennie Si |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2024 | A New, Physics-Informed Continuous-Time Reinforcement Learning Algorithm with Performance GuaranteesabstractWe introduce a new, physics-informed continuous-time reinforcement learning (CT-RL) algorithm for control of affine nonlinear systems, an area that enables a plethora of well-motivated applications. Based on fundamental control principles, our approach uses reference command input (RCI) as probing noise to enable exploration in learning. With known physical dynamics of the environment, by leveraging on the Kleinman algorithm structure, and using state-action trajectory data, RCI provides a data-efficient optimal control solution under an infinite-horizon undiscounted cost. We show that our RCI-based CT-RL algorithm not only provides theoretical guarantees such as learning convergence, solution optimality, and closed-loop stability, but also well-behaved dynamic system responses. It is noted that our evaluations not only include extensive baseline and ablation studies using typical performance measures in RL, but also essential control-centric performance measures that are critical for real-life control applications. As a result, we demonstrate that our RCI-based CT-RL leads to new, SOTA control design and performance. Brent Wallace, Jennie Si |
J. Mach. Learn. Res. | 2 |
| 2024 | Reinforcement Learning Control With Knowledge ShapingabstractWe aim at creating a transfer reinforcement learning framework that allows the design of learning controllers to leverage prior knowledge extracted from previously learned tasks and previous data to improve the learning performance of new tasks. Toward this goal, we formalize knowledge transfer by expressing knowledge in the value function in our problem construct, which is referred to as reinforcement learning with knowledge shaping (RL-KS). Unlike most transfer learning studies that are empirical in nature, our results include not only simulation verifications but also an analysis of algorithm convergence and solution optimality. Also different from the well-established potential-based reward shaping methods which are built on proofs of policy invariance, our RL-KS approach allows us to advance toward a new theoretical result on positive knowledge transfer. Furthermore, our contributions include two principled ways that cover a range of realization schemes to represent prior knowledge in RL-KS. We provide extensive and systematic evaluations of the proposed RL-KS method. The evaluation environments not only include classical RL benchmark problems but also include a challenging task of real-time control of a robotic lower limb with a human user in the loop. Xiang Gao 0015, Jennie Si, He Huang 0002 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2024 | Continuous-Time Reinforcement Learning Control: A Review of Theoretical Results, Insights on Performance, and Needs for New DesignsabstractThis exposition discusses continuous-time reinforcement learning (CT-RL) for the control of affine nonlinear systems. We review four seminal methods that are the centerpieces of the most recent results on CT-RL control. We survey the theoretical results of the four methods, highlighting their fundamental importance and successes by including discussions on problem formulation, key assumptions, algorithm procedures, and theoretical guarantees. Subsequently, we evaluate the performance of the control designs to provide analyses and insights on the feasibility of these design methods for applications from a control designer's point of view. Through systematic evaluations, we point out when theory diverges from practical controller synthesis. We, furthermore, introduce a new quantitative analytical framework to diagnose the observed discrepancies. Based on the analyses and the insights gained through quantitative evaluations, we point out potential future research directions to unleash the potential of CT-RL control algorithms in addressing the identified challenges. Brent Wallace, Jennie Si |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2024 | Toward Task-Independent Optimal Adaptive Control of a Hip Exoskeleton for Locomotion Assistance in NeurorehabilitationabstractPersonalized robotic exoskeleton control is essential in assisting individuals with motor deficits. However, current research still lacks a solution from the end of a practical need of the problem to the end of its successful demonstration in physical environments, namely an end-to-end solution, that enables stable and continuous walking across different tasks. This study addresses this challenge by introducing a hierarchical control framework for the purpose. At the low level, impedance control ensures joint compliance without causing injury to users. At the high level, a reinforcement learning (RL)-based optimal adaptive controller automatically personalizes assistance to both hip extension and flexion (namely, bi-directional) to reach a target range of motion (ROM) under multiple walking conditions. As the first potentially feasible approach to this challenging problem and to meet practical use requirements, we developed a least-square policy iteration-based solution to configure the intrinsic parameters within the well-established finite state machine impedance control (FSM-IC). We successfully tested the control solution on eight young unimpaired participants and one participant post-stroke wearing a hip exoskeleton while walking on an instrumented treadmill. The proposed method can be applied to solving for optimal impedance parameters for individual users and different task scenarios to increase joint ROM. Our next step is to further evaluate this solution framework on additional people with hemiparesis who may benefit from hip joint assistance in therapy or daily activities to restore normative or improve gait patterns. Qiang Zhang 0028, Jennie Si, Xikai Tu, Minhan Li, Michael D. Lewek, He Huang 0002 |
IEEE Trans. Syst. Man Cybern. Syst. | 2 |
| 2023 | A Robotic Assistance Personalization Control Approach of Hip Exoskeletons for Gait Symmetry ImprovementabstractHealthy human locomotion functions with good gait symmetry depend on rhythmic coordination of the left and right legs, which can be deteriorated by neurological disorders like stroke and spinal cord injury. Powered exoskeletons are promising devices to improve impaired people's locomotion functions, like gait symmetry. However, given higher uncertainties and the time-varying nature of human-robot interaction, providing personalized robotic assistance from exoskeletons to achieve the best gait symmetry is challenging, especially for people with neurological disorders. In this paper, we propose a hierarchical control framework for a bilateral hip exoskeleton to provide the adaptive optimal hip joint assistance with a control objective of imposing the desired gait symmetry during walking. Three control levels are included in the hierarchical framework, including the high-level control to tune three control parameters based on a policy iteration reinforcement learning approach, the middle-level control to define the desired assistive torque profile based on a delayed output feedback control method, and the low-level control to achieve a good torque trajectory tracking performance. To evaluate the feasibility of the proposed control framework, five healthy young participants are recruited for treadmill walking experiments, where an artificial gait asymmetry is imitated as the hemiparesis post-stroke, and only the ‘paretic’ hip joint is controlled with the proposed framework. The pilot experimental studies demonstrate that the hierarchical control framework for the hip exoskeleton successfully (asymmetry index from 8.8% to − 0.5%) and efficiently (less than 4 minutes) achieved the desired gait symmetry by providing adaptive optimal assistance on the ‘paretic’ hip joint. Qiang Zhang 0028, Xikai Tu, Jennie Si, Michael D. Lewek, He Huang 0002 |
IROS | 3 |
| 2023 | A Long N-step Surrogate Stage Reward for Deep Reinforcement LearningabstractWe introduce a new stage reward estimator named the long $N$-step surrogate stage (LNSS) reward for deep reinforcement learning (RL). It aims at mitigating the high variance problem, which has shown impeding successful convergence of learning, hurting task performance, and hindering applications of deep RL in continuous control problems. In this paper we show that LNSS, which utilizes a long reward trajectory of rewards of future steps, provides consistent performance improvement measured by average reward, convergence speed, learning success rate,and variance reduction in $Q$ values and rewards. Our evaluations are based on a variety of environments in DeepMind Control Suite and OpenAI Gym by using LNSS in baseline deep RL algorithms such as DDPG, D4PG, and TD3. We show that LNSS reward has enabled good results that have been challenging to obtain by deep RL previously. Our analysis also shows that LNSS exponentially reduces the upper bound on the variances of $Q$ values from respective single-step methods. Junmin Zhong, Jennie Si |
NeurIPS | 3 |
| 2023 | Deep Reinforcement Learning for Load Shedding Against Short-Term Voltage Instability in Large Power SystemsabstractWe introduce an innovative solution approach to the challenging dynamic load-shedding problem which directly affects the stability of large power grid. Our proposed deep Q-network for load-shedding (DQN-LS) determines optimal load-shedding strategy to maintain power system stability by taking into account both spatial and temporal information of a dynamically operating power system, using a convolutional long-short-term memory (ConvLSTM) network to automatically capture dynamic features that are translation-invariant in short-term voltage instability, and by introducing a new design of the reward function. The overall goal for the proposed DQN-LS is to provide real-time, fast, and accurate load-shedding decisions to increase the quality and probability of voltage recovery. To demonstrate the efficacy of our proposed approach and its scalability to large-scale, complex dynamic problems, we utilize the China Southern Grid (CSG) to obtain our test results, which clearly show superior voltage recovery performance by employing the proposed DQN-LS under different and uncertain power system fault conditions. What we have developed and demonstrated in this study, in terms of the scale of the problem, the load-shedding performance obtained, and the DQN-LS approach, have not been demonstrated previously. Yonghong Luo, Boya Wang, Chao Lu 0009, Jennie Si, Jie Song 0002 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2022 | Admittance Control Based Human-in-the-Loop Optimization for Hip Exoskeleton Reduces Human Exertion during WalkingabstractHuman-in-the-loop (HIL) optimization usually optimizes assistive torque of exoskeletons to minimize the human's energetic expenditure in walking, quantified by metabolic cost. This formulation can, however, result in altered gait pattern of the human joint from the natural pattern, which is undesired. In this paper, we proposed a novel concept of HIL optimization of a hip exoskeleton. The optimization goal was to maintain the hip kinematics while providing optimal mechanical energy from the exoskeleton by modulating the admittance control. Policy iteration was used to optimize the switching time within the gait phase, at which a single parameter of the admittance controller was altered to provide assistance. The stiffness and equilibrium angle were considered as the two parameters for altering at the switching time, resulting in three possible modes of operation for the algorithm: (i) switching the equilibrium point, (ii) switching stiffness while equilibrium point is set at maximum extension and, (iii) maximum flexion. The optimization algorithm was found to converge for all three modes, with the equilibrium mode resulting in multiple solutions. Further analysis of power injected by the exoskeleton in the three modes showed that the first and third mode reduced human energetic exertion while the second mode increased human exertion. Implications of the results as well as the observed muscle activation patterns in response to assistance are discussed. Varun Nalam, Xikai Tu, Minhan Li, Jennie Si, He Huang 0002 |
ICRA | 4 |
| 2022 | Human-Robotic Prosthesis as Collaborating Agents for Symmetrical WalkingabstractThis is the first attempt at considering human influence in the reinforcement learning control of a robotic lower limb prosthesis toward symmetrical walking in real world situations. We propose a collaborative multi-agent reinforcement learning (cMARL) solution framework for this highly complex and challenging human-prosthesis collaboration (HPC) problem. The design of an automatic controller of the robot within the HPC context is based on accessible physical features or measurements that are known to affect walking performance. Comparisons are made with the current state-of-the-art robot control designs, which are single-agent based, as well as existing MARL solution approaches tailored to the problem, including multi-agent deep deterministic policy gradient (MADDPG) and counterfactual multi-agent policy gradient (COMA). Results show that, when compared to these approaches, treating the human and robot as coupled agents and using estimated human adaption in robot control design can achieve lower stage cost, peak error, and symmetry value to ensure better human walking performance. Additionally, our approach accelerates learning of walking tasks and increases learning success rate. The proposed framework can potentially be further developed to examine how human and robotic lower limb prosthesis interact, an area that little is known about. Advancing cMARL toward real world applications such as HPC for normative walking sets a good example of how AI can positively impact on people’s lives. Junmin Zhong, Brent Wallace, Xiang Gao 0015, He Huang 0002, Jennie Si |
NeurIPS | 6 |
| 2022 | Reinforcement Learning Control of Robotic Knee With Human-in-the-Loop by Flexible Policy IterationabstractWe are motivated by the real challenges presented in a human-robot system to develop new designs that are efficient at data level and with performance guarantees, such as stability and optimality at system level. Existing approximate/adaptive dynamic programming (ADP) results that consider system performance theoretically are not readily providing practically useful learning control algorithms for this problem, and reinforcement learning (RL) algorithms that address the issue of data efficiency usually do not have performance guarantees for the controlled system. This study fills these important voids by introducing innovative features to the policy iteration algorithm. We introduce flexible policy iteration (FPI), which can flexibly and organically integrate experience replay and supplemental values from prior experience into the RL controller. We show system-level performances, including convergence of the approximate value function, (sub)optimality of the solution, and stability of the system. We demonstrate the effectiveness of the FPI via realistic simulations of the human-robot system. It is noted that the problem we face in this study may be difficult to address by design methods based on classical control theory as it is nearly impossible to obtain a customized mathematical model of a human-robot system either online or offline. The results we have obtained also indicate the great potential of RL control to solving realistic and challenging problems with high-dimensional control inputs. Xiang Gao 0015, Jennie Si, Yue Wen, Minhan Li, He Huang 0002 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2022 | Editorial Biologically Learned/Inspired Methods for Sensing, Control, and Decision
Yongduan Song 0001, Jennie Si, Sonya A. Coleman, Dermot Kerr |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2022 | Online Reinforcement Learning Control by Direct Heuristic Dynamic Programming: From Time-Driven to Event-DrivenabstractIn this work, time-driven learning refers to the machine learning method that updates parameters in a prediction model continuously as new data arrives. Among existing approximate dynamic programming (ADP) and reinforcement learning (RL) algorithms, the direct heuristic dynamic programming (dHDP) has been shown an effective tool as demonstrated in solving several complex learning control problems. It continuously updates the control policy and the critic as system states continuously evolve. It is therefore desirable to prevent the time-driven dHDP from updating due to insignificant system event such as noise. Toward this goal, we propose a new event-driven dHDP. By constructing a Lyapunov function candidate, we prove the uniformly ultimately boundedness (UUB) of the system states and the weights in the critic and the control policy networks. Consequently, we show the approximate control and cost-to-go function approaching Bellman optimality within a finite bound. We also illustrate how the event-driven dHDP algorithm works in comparison to the original time-driven dHDP. Qingtao Zhao, Jennie Si, Jian Sun 0003 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2022 | Toward Expedited Impedance Tuning of a Robotic Prosthesis for Personalized Gait Assistance by Reinforcement Learning ControlabstractPersonalizing medical devices such as lower limb wearable robots is challenging. While the initial feasibility of automating the process of knee prosthesis control parameter tuning has been demonstrated in a principled way, the next critical issue is to improve tuning efficiency and speed it up for the human user, in clinic settings, while maintaining human safety. We, therefore, propose a policy iteration with constraint embedded (PICE) method as an innovative solution to the problem under the framework of reinforcement learning. Central to PICE is the use of a projected Bellman equation with a constraint of assuring positive semidefiniteness of performance values during policy evaluation. Additionally, we developed both online and offline PICE implementations that provide additional flexibility for the designer to fully utilize measurement data, either from on-policy or off-policy, to further improve PICE tuning efficiency. Our human subject testing showed that the PICE provided effective policies with significantly reduced tuning time. For the first time, we also experimentally evaluated and demonstrated the robustness of the deployed policies by applying them to different tasks and users. Putting it together, our new way of problem solving has been effective as PICE has demonstrated its potential toward truly automating the process of control parameter tuning for robotic knee prosthesis users. Minhan Li, Yue Wen, Xiang Gao 0015, Jennie Si, He Huang 0002 |
IEEE Trans. Robotics | 4 |
| 2022 | Reinforcement-Learning-Based Tracking Control of Waste Water Treatment Process Under Realistic System Conditions and Control Performance RequirementsabstractThe tracking control of a wastewater treatment process (WWTP) is considered. The process is highly nonlinear, with strong coupling, difficult to model mathematically, and the operation is subject to unknown disturbances. We address this multivariable tracking control problem by applying the direct heuristic dynamic programming (dHDP)-based reinforcement learning control. The control goal is to track a desired reference of the dissolved oxygen (DO) concentration of the 5th aerobic zone ($S_{O5}$) and nitrate concentration of the 2nd anoxic zone ($S_{NO2}$) by manipulating the oxygen transfer coefficient of the 5th aerobic zone ($K_{L}a_{5}$) and internal recycle flow rate ($Q_{a}$). The dHDP aims at achieving a minimal accumulated WWTP tracking error while dealing with strong coupling between the$S_{O5}$and$S_{NO2}$and eliminating unknown disturbances in the process. The proposed dHDP approach devises an optimal control strategy entirely driven by WWTP process data as an online learning control method. We have conducted extensive and systematic simulations based on the well-known BSM1 platform of the WWTP controlled by dHDP to compare and contrast performances with other methods. Qinmin Yang, Wenchao Meng, Jennie Si |
IEEE Trans. Syst. Man Cybern. Syst. | 4 |
| 2021 | A Data-Driven Reinforcement Learning Solution Framework for Optimal and Adaptive Personalization of a Hip ExoskeletonabstractRobotic exoskeletons are exciting technologies for augmenting human mobility. However, designing such a device for seamless integration with the human user and to assist human movement still is a major challenge. This paper aims at developing a novel data-driven solution framework based on reinforcement learning (RL), without first modeling the human-robot dynamics, to provide optimal and adaptive personalized torque assistance for reducing human efforts during walking. Our automatic personalization solution framework includes the assistive torque profile with two control timing parameters (peak and offset timings), the least square policy iteration (LSPI) for learning the parameter tuning policy, and a cost function based on a transferred work ratio. The proposed controller was successfully validated on a healthy human subject to assist unilateral hip extension in walking. The results showed that the optimal and adaptive RL controller as a new approach was feasible for tuning assistive torque profile of the hip exoskeleton that coordinated with human actions and reduced activation level of hip extensor muscle in human. Xikai Tu, Minhan Li, Ming Liu 0005, Jennie Si, He Huang 0002 |
ICRA | 4 |
| 2021 | User Controlled Interface for Tuning Robotic Knee ProsthesisabstractThe tuning process for a robotic prosthesis is a challenging and time-consuming task both for users and clinicians. An automatic tuning approach using reinforcement learning (RL) has been developed for a knee prosthesis to address the challenges of manual tuning methods. The algorithm tunes the optimal control parameters based on the provided knee joint profile that the prosthesis is expected to replicate during gait safely. This paper presents an intuitive interface designed for the prosthesis users and clinicians to choose the preferred knee joint profile during gait and use the autotuner to replicate in the prosthesis. The interface-based approach is validated by observing the ability of the tuning algorithm to successfully converge to various alternate knee profiles by testing on two able-bodied subjects walking with a robotic knee prosthesis. The algorithm was found to converge successfully in an average duration of 1.15 min for the first subject and 2.31 min for the second subject. Further, the subjects displayed different preferences for optimal profiles reinforcing the need to tune alternate profiles. The implications of the results in the tuning of robotic prosthetic devices are discussed. Abbas Alili, Varun Nalam, Minhan Li, Ming Liu 0005, Jennie Si, He Huang 0002 |
IROS | 5 |
| 2020 | Knowledge-Guided Reinforcement Learning Control for Robotic Lower Limb ProsthesisabstractRobotic prostheses provide new opportunities to better restore lost functions than passive prostheses for trans-femoral amputees. But controlling a prosthesis device automatically for individual users in different task environments is an unsolved problem. Reinforcement learning (RL) is a naturally promising tool. For prosthesis control with a user in the loop, it is desirable that the controlled prosthesis can adapt to different task environments as quickly and smoothly as possible. However, most RL agents learn or relearn from scratch when the environment changes. To address this issue, we propose the knowledge-guided Q-learning (KG-QL) control method as a principled way for the problem. In this report, we collected and used data from two able-bodied (AB) subjects wearing a RL controlled robotic prosthetic limb walking on level ground. Our ultimate goal is to build an efficient RL controller with reduced time and data requirements and transfer knowledge from AB subjects to amputee subjects. Toward this goal, we demonstrate its feasibility by employing OpenSim, a well-established human locomotion simulator. Our results show the OpenSim simulated amputee subject improved control tuning performance over learning from scratch by utilizing knowledge transfer from AB subjects. Also in this paper, we will explore the possibility of information transfer from AB subjects to help tuning for the amputee subjects. Xiang Gao 0015, Jennie Si, Yue Wen, Minhan Li, He Huang 0002 |
ICRA | 2 |
| 2020 | Online Reinforcement Learning Control for the Personalization of a Robotic Knee ProsthesisabstractRobotic prostheses deliver greater function than passive prostheses, but we face the challenge of tuning a large number of control parameters in order to personalize the device for individual amputee users. This problem is not easily solved by traditional control designs or the latest robotic technology. Reinforcement learning (RL) is naturally appealing. The recent, unprecedented success of AlphaZero demonstrated RL as a feasible, large-scale problem solver. However, the prosthesis-tuning problem is associated with several unaddressed issues such as that it does not have a known and stable model, the continuous states and controls of the problem may result in a curse of dimensionality, and the human-prosthesis system is constantly subject to measurement noise, environmental change and human-body-caused variations. In this paper, we demonstrated the feasibility of direct heuristic dynamic programming, an approximate dynamic programming (ADP) approach, to automatically tune the 12 robotic knee prosthesis parameters to meet individual human users' needs. We tested the ADP-tuner on two subjects (one able-bodied subject and one amputee subject) walking at a fixed speed on a treadmill. The ADP-tuner learned to reach target gait kinematics in an average of 300 gait cycles or 10 min of walking. We observed improved ADP tuning performance when we transferred a previously learned ADP controller to a new learning session with the same subject. To the best of our knowledge, our approach to personalize robotic prostheses is the first implementation of online ADP learning control to a clinical problem involving human subjects. Yue Wen, Jennie Si, Andrea Brandt, Xiang Gao 0015, He Huang 0002 |
IEEE Trans. Cybern. | 2 |
| 2019 | Offline Policy Iteration Based Reinforcement Learning Controller for Online Robotic Knee Prosthesis Parameter TuningabstractThis paper aims to develop an optimal controller that can automatically provide personalized control of robotic knee prosthesis in order to best support gait of individual prosthesis wearers. We introduced a new reinforcement learning (RL) controller for this purpose based on the promising ability of RL controllers to solve optimal control problems through interactions with the environment without requiring an explicit system model. However, collecting data from a human-prosthesis system is expensive and thus the design of a RL controller has to take into account data and time efficiency. We therefore propose an offline policy iteration based reinforcement learning approach. Our solution is built on the finite state machine (FSM) impedance control framework, which is the most used prosthesis control method in commercial and prototypic robotic prosthesis. Under such a framework, we designed an approximate policy iteration algorithm to devise impedance parameter update rules for 12 prosthesis control parameters in order to meet individual users' needs. The goal of the reinforcement learning-based control was to reproduce near-normal knee kinematics during gait. We tested the RL controller obtained from offline learning in real time experiment involving the same able-bodied human subject wearing a robotic lower limb prosthesis. Our results showed that the RL control resulted in good convergent behavior in kinematic states, and the offline learning control policy successfully adjusted the prosthesis control parameters to produce near-normal knee kinematics in 10 updates of the impedance control parameters. Minhan Li, Xiang Gao 0015, Yue Wen, Jennie Si, He Huang 0002 |
ICRA | 4 |
| 2018 | Deep Reinforcement Leaming for Short-term Voltage Control by Dynamic Load Shedding in China Southem Power GridabstractWe propose a novel load shedding (LS) scheme against voltage instability using deep reinforcement learning (DRL). Both spatial and temporal information of a large power grid are used in the DRL control scheme. Specifically, both dynamic state variables and grid topology information are inputs to the learning controller. Within the DRL scheme, a deep learning neural network is designed and implemented to automatically extract translation-invariance information about voltage instability. The DRL load shedding controller interacts with system dynamics through a sequence of observations, actions and rewards to determine load shedding amounts in a manner that maximizes cumulative future reward, accomplishing coordination within the region rapidly to meet the online application requirements. The DRL based distributed LS scheme is performed to control the China Southern Power Grid (CSG) system. Our results show improved voltage recovery performance by load shedding using the proposed scheme under different unknown test scenarios. Chao Lu 0009, Jennie Si, Jie Song 0002, Yinsheng Su |
IJCNN | 3 |
| 2018 | Policy Approximation in Policy Iteration Approximate Dynamic Programming for Discrete-Time Nonlinear SystemsabstractPolicy iteration approximate dynamic programming (DP) is an important algorithm for solving optimal decision and control problems. In this paper, we focus on the problem associated with policy approximation in policy iteration approximate DP for discrete-time nonlinear systems using infinite-horizon undiscounted value functions. Taking policy approximation error into account, we demonstrate asymptotic stability of the control policy under our problem setting, show boundedness of the value function during each policy iteration step, and introduce a new sufficient condition for the value function to converge to a bounded neighborhood of the optimal value function. Aiming for practical implementation of an approximate policy, we consider using Volterra series, which has been extensively covered in controls literature for its good theoretical properties and for its success in practical applications. We illustrate the effectiveness of the main ideas developed in this paper using several examples including a practical problem of excitation control of a hydrogenerator. Wentao Guo 0002, Jennie Si, Feng Liu 0014, Shengwei Mei |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2017 | Comparing parallel and sequential control parameter tuning for a powered knee prosthesisabstractPowered knee prostheses, compared to traditional energetically-passive knee prostheses, greatly enhance the mobility of transfemoral amputees. However, powered prostheses have a large number of control parameters that must be adjusted for individual amputee users, which presents a great challenge for clinical use. To address this challenge, we proposed and compared 2 automatic tuning strategies (i.e. parallel and sequential) using our newly developed optimal adaptive dynamic programming (ADP) tuner that objectively tuned the control parameters of an experimental powered knee prosthesis to mimic the knee profile of an able-bodied person (i.e. reference profile). With the parallel tuning strategy, we tuned all control parameters during the stance and the swing phases simultaneously. With the sequential tuning strategy, we alternately tuned stance or swing phase control parameters while fixing the remaining parameters. One able-bodied subject with a prosthesis adapter and one transfemoral amputee subject walked with the experimental powered knee prosthesis under both tuning strategies. Results show that with both tuning strategies, the ADP tuner successfully tuned the impedance parameters to match the prosthetic knee profile to the reference profile. Additionally, the parallel strategy outperformed the sequential strategy with better convergence to the reference profile. Interestingly, with the sequential tuning strategy, tuning during the swing phase greatly impacted the subsequent stance phase profile, but the impact was not as great when the order of tuning was switched. The ability to simultaneously adjust all control parameters with ADP using a parallel strategy may be a preferred solution for the current high-dimension control challenge, which may lead to more advanced, adaptive powered knee prostheses. Yue Wen, Andrea Brandt, Ming Liu 0005, He Huang 0002, Jennie Si |
SMC | 5 |
| 2017 | Consensus Control of Nonlinear Multiagent Systems With Time-Varying State ConstraintsabstractIn this paper, we present a novel adaptive consensus algorithm for a class of nonlinear multiagent systems with time-varying asymmetric state constraints. As such, our contribution is a step forward beyond the usual consensus stabilization result to show that the states of the agents remain within a user defined, time-varying bound. To prove our new results, the original multiagent system is transformed into a new one. Stabilization and consensus of transformed states are sufficient to ensure the consensus of the original networked agents without violating of the predefined asymmetric time-varying state constraints. A single neural network (NN), whose weights are tuned online, is used in our design to approximate the unknown functions in the agent's dynamics. To account for the NN approximation residual, reconstruction error, and external disturbances, a robust term is introduced into the approximating system equation. Additionally in our design, each agent only exchanges the information with its neighbor agents, and thus the proposed consensus algorithm is decentralized. The theoretical results are proved via Lyapunov synthesis. Finally, simulations are performed on a nonlinear multiagent system to illustrate the performance of our consensus design scheme. Wenchao Meng, Qinmin Yang, Jennie Si, Youxian Sun |
IEEE Trans. Cybern. | 3 |
| 2017 | A New Powered Lower Limb Prosthesis Control Framework Based on Adaptive Dynamic ProgrammingabstractThis brief presents a novel application of adaptive dynamic programming (ADP) for optimal adaptive control of powered lower limb prostheses, a type of wearable robots to assist the motor function of the limb amputees. Current control of these robotic devices typically relies on finite state impedance control (FS-IC), which lacks adaptability to the user's physical condition. As a result, joint impedance settings are often customized manually and heuristically in clinics, which greatly hinder the wide use of these advanced medical devices. This simulation study aimed at demonstrating the feasibility of ADP for automatic tuning of the twelve knee joint impedance parameters during a complete gait cycle to achieve balanced walking. Given that the accurate models of human walking dynamics are difficult to obtain, the model-free ADP control algorithms were considered. First, direct heuristic dynamic programming (dHDP) was applied to the control problem, and its performance was evaluated on OpenSim, an often-used dynamic walking simulator. For the comparison purposes, we selected another established ADP algorithm, the neural fitted Q with continuous action (NFQCA). In both cases, the ADP controllers learned to control the right knee joint and achieved balanced walking, but dHDP outperformed NFQCA in this application during a 200 gait cycle-based testing. Yue Wen, Jennie Si, Xiang Gao 0015, Stephanie Huang, He Huang 0002 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2016 | Adaptive Neural Control of a Class of Output-Constrained Nonaffine SystemsabstractIn this paper, we present a novel tracking controller for a class of uncertain nonaffine systems with time-varying asymmetric output constraints. Firstly, the original nonaffine constrained (in the sense of the output signal) control system is transformed into a output-feedback control problem of an unconstrained affine system in normal form. As a result, stabilization of the transformed system is sufficient to ensure constraint satisfaction. It is subsequently shown that the output tracking is achieved without violation of the predefined asymmetric time-varying output constraints. Therefore, we are capable of quantifying the system performance bounds as functions of time on both transient and steady-state stages. Furthermore, the transformed system is linear with respect to a new input signal and the traditional backstepping scheme is avoided, which makes the synthesis extremely simplified. All the signals in the closed-loop system are proved to be semi-globally, uniformly, and ultimately bounded via Lyapunov synthesis. Finally, the simulation results are presented to illustrate the performance of the proposed controller. Wenchao Meng, Qinmin Yang, Jennie Si, Youxian Sun |
IEEE Trans. Cybern. | 3 |
| 2016 | Online Supplementary ADP Learning Controller Design and Application to Power System Frequency Control With Large-Scale Wind Energy IntegrationabstractThe emergence of smart grids has posed great challenges to traditional power system control given the multitude of new risk factors. This paper proposes an online supplementary learning controller (OSLC) design method to compensate the traditional power system controllers for coping with the dynamic power grid. The proposed OSLC is a supplementary controller based on approximate dynamic programming, which works alongside an existing power system controller. By introducing an action-dependent cost function as the optimization objective, the proposed OSLC is a nonidentifier-based method to provide an online optimal control adaptively as measurement data become available. The online learning of the OSLC enjoys the policy-search efficiency during policy iteration and the data efficiency of the least squares method. For the proposed OSLC, the stability of the controlled system during learning, the monotonic nature of the performance measure of the iterative supplementary controller, and the convergence of the iterative supplementary controller are proved. Furthermore, the efficacy of the proposed OSLC is demonstrated in a challenging power system frequency control problem in the presence of high penetration of wind generation. Wentao Guo 0002, Feng Liu 0014, Jennie Si, Dawei He, Ronald G. Harley, Shengwei Mei |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2016 | Editorial IEEE Transactions on Neural Networks and Learning Systems 2016 and Beyondabstract“Happy New Year!” At the beginning of 2016, I would like to take this opportunity to wish everyone a very happy, healthy, and prosperous new year! It is my great honor and privilege to serve as the Editor-in-Chief (EiC) of the IEEE TRANSACTIONS ON NEURAL NETWORKS AND LEARNING SYSTEMS (TNNLS), and I am excited to write this Editorial to start a new journey with you all. Haibo He, Nitesh V. Chawla, Yoonsuck Choe, Andries P. Engelbrecht, Jaya deva, Lyle N. Long, Ali A. Minai, Feiping Nie 0001, Umut Ozertem, Barak A. Pearlmutter, Ling Shao 0001, Jennie Si, Jochen J. Steil, Brijesh K. Verma, Ding Wang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 13 |
| 2015 | Error bound analysis of policy iteration based approximate dynamic programming for deterministic discrete-time nonlinear systemsabstractExtensive approximate dynamic programming (ADP) algorithms have been developed based on policy iteration. For policy iteration based ADP of deterministic discrete-time nonlinear systems, existing literature has proved its convergence in the formulation of undiscounted value function under the assumption of exact approximation. Furthermore, the error bound of policy iteration based ADP has been analyzed in a discounted value function formulation with consideration of approximation errors. However, there has not been any error bound analysis of policy iteration based ADP in the undiscounted value function formulation with consideration of approximation errors. In this paper, we intend to fill this theoretical gap. We provide a sufficient condition on the approximation error, so that the iterative value function can be bounded in a neighbourhood of the optimal value function. To the best of the authors' knowledge, this is the first error bound result of the undiscounted policy iteration for deterministic discrete-time nonlinear systems considering approximation errors. Wentao Guo 0002, Feng Liu 0014, Jennie Si, Shengwei Mei, Rui Li 0032 |
IJCNN | 3 |
| 2015 | Approximate dynamic programming based supplementary reactive power control for DFIG wind farm to enhance power system stability
Wentao Guo 0002, Feng Liu 0014, Jennie Si, Dawei He, Ronald G. Harley, Shengwei Mei |
Neurocomputing | 3 |
| 2015 | FREL: A Stable Feature Selection AlgorithmabstractTwo factors characterize a good feature selection algorithm: its accuracy and stability. This paper aims at introducing a new approach to stable feature selection algorithms. The innovation of this paper centers on a class of stable feature selection algorithms called feature weighting as regularized energy-based learning (FREL). Stability properties of FREL using L1 or L2 regularization are investigated. In addition, as a commonly adopted implementation strategy for enhanced stability, an ensemble FREL is proposed. A stability bound for the ensemble FREL is also presented. Our experiments using open source real microarray data, which are challenging high dimensionality small sample size problems demonstrate that our proposed ensemble FREL is not only stable but also achieves better or comparable accuracy than some other popular stable feature weighting methods. Yun Li 0009, Jennie Si, Guojing Zhou, Shasha Huang, Songcan Chen |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2014 | Reactive power control of DFIG wind farm using online supplementary learning controller based on approximate dynamic programmingabstractDynamic reactive power control of doubly fed induction generators (DFIGs) plays a crucially important role in maintaining transient stability of power systems with high penetration of DFIG based wind generation. Based on approximate dynamic programming (ADP), this paper proposes an optimal adaptive supplementary reactive power controller for DFIGs. By augmenting a corrective regulation signal to the reactive power command of rotor-side converter (RSC) of a DFIG, the supplementary controller is designed to reduce voltage sag at the point of common connection (PCC) during a fault, and to mitigate output active power oscillation of the wind farm after a fault. As a result, the transient stability of both DFIG and the power grid is enhanced. An action dependent cost function is introduced to provide real-time online ADP learning control. Furthermore, a policy iteration algorithm using high-efficiency least square method is employed to train the supplementary controller in an online model-free manner. By using such techniques, the supplementary reactive power controller is endowed with capability of online optimization and adaptation. Simulations carried out on a benchmark power system integrating a large DFIG wind farm show that the ADP based supplementary reactive power controller can significantly improve the transient system stability in changing operation conditions. Wentao Guo 0002, Feng Liu 0014, Dawei He, Jennie Si, Ronald G. Harley, Shengwei Mei |
IJCNN | 4 |
| 2014 | Online adaptation of controller parameters based on approximate dynamic programmingabstractController parameter tuning is an integral part of control engineering practice. Existing tuning methods usually start with an accurate mathematical model of the controlled system, which may pose some challenges for practicing engineers dealing with real systems. As such, parameter optimization and adaptation are treated as two independent steps during tuning. To address these issues, we propose a new, online parameterized controller tuning method for a general nonlinear dynamic system. This tuning method is based on direct heuristic dynamic programming (direct HDP), a model-free algorithm in the approximated dynamic programming (ADP) family. By using a Lyapunov stability approach, we provide uniformly ultimately bounded (UUB) results under some mild conditions for controller parameters, the critic neural network weights, and the action neural network weights. Simulation studies based on the benchmark cart-pole system demonstrate adaptability and optimization capabilities of the proposed controller parameter tuning method. Wentao Guo 0002, Feng Liu 0014, Jennie Si, Shengwei Mei |
IJCNN | 3 |
| 2014 | Policy iteration approximate dynamic programming using Volterra series based actorabstractThere is an extensive literature on value function approximation for approximate dynamic programming (ADP). Multilayer perceptrons (MLPs) and radial basis functions (RBFs), among others, are typical approximators for value functions in ADP. Similar approaches have been taken for policy approximation. In this paper, we propose a new Volterra series based structure for actor approximation in ADP. The Volterra approx-imator is linear in parameters with global optima attainable. Given the proposed approximator structures, we further develop a policy iteration framework under which a gradient descent training algorithm for obtaining the optimal Volterra kernels can be obtained. Associated with this ADP design, we provide a sufficient condition based on actor approximation error to guarantee convergence of the value function iterations. A finite bound of the final convergent value function is also given. Finally, by using a simulation example we illustrate the effectiveness of the proposed Volterra actor for optimal control of a nonlinear system. Wentao Guo 0002, Jennie Si, Feng Liu 0014, Shengwei Mei |
IJCNN | 2 |
| 2014 | Longitudinal control of hypersonic vehicles based on direct heuristic dynamic programming using ANFISabstractSince the launch of the scramjet, recent years have witnessed a growing interest in the study of airbreathing hypersonic vehicles. Due to its strong coupling characteristics, high nonlinearity, and uncertain parameters, the control of hypersonic vehicle becomes a great challenge. To deal with those design issues, we propose an adaptive learning control method based on direct heuristic dynamic programming (direct HDP), which is used to track the angle of attack despite the presence of bounded uncertain parameters. Inspired by the adaptive critic designs, direct HDP is one of the adaptive dynamic programming (ADP) methods, which is a model-free reinforcement learning algorithm using the online learning scheme to solve dynamic control problems in realistic complex environment. In this paper, this direct HDP method is improved by embedding the fuzzy neural network (FNN) in the controller design to enhance its self-learning ability and robustness. Simulation results are provided to demonstrate the effectiveness of our proposed method. Xiong Luo, Jennie Si, Feng Liu 0014 |
IJCNN | 3 |
| 2013 | An integrated design for intensified direct heuristic dynamic programmingabstractThere has been a growing interest in the study of adaptive/approximate dynamic programming (ADP) in recent years. The ADP technique provides a powerful tool to understand and improve the principled technologies of machine intelligence system. As one of the ADP algorithms based on adaptive critic neural networks (NNs), the direct heuristic dynamic programming (direct HDP) has demonstrated some successful applications in solving realistic engineering control problems. In this study, based on a three-network architecture in which the reinforcement signal is approximated by an additional NN, a novel integrated design method for intensified direct HDP is developed. The new design approach is implemented by using multiple PID neural networks (PIDNNs), which effectively takes into account structural knowledge of system states and control that are usually present in a physical system. By using a Lyapunov stability approach, a uniformly ultimately boundedness (UUB) result is proved for our PIDNNs-based intensified direct HDP learning controller. Furthermore, the learning and control performances of the proposed design is tested using the popular cart-pole example to illustrate the key ideas of this paper. Xiong Luo, Jennie Si, Yuchao Zhou |
ADPRL | 2 |
| 2013 | Incorporating approximate dynamic programming-based parameter tuning into PD-type virtual inertia control of DFIGsabstractDoubly fed induction generators (DFIGs) are widely used in wind power generation. For controlling DFIGs to maintain network frequency within a safety range, the proportional-derivative (PD) type virtual inertia controllers (VIC) are used in the active power control of DFIGs. However, as is well known, wind power generation conditions change directly with wind conditions in nature. Such changes create great challenge for the VIC design and actually force the control designs to go beyond the traditional problem formulation of using explicit objective functions associated with specific optimality. Controller parameter tuning thus necessarily becomes a part of the controller design. In this paper, we propose an approximate dynamic programming (ADP) structure for online tuning of the PD type virtual inertia controller parameters. The proposed ADP structure naturally takes into account the PD control into design objective and provides the PD controller with online parameter tuning capability through learning. Design and implementation details of the proposed methodology, including neural network weight initialization, design of the reinforcement signal, data preprocessing, and a bound on the online tuned parameters are discussed in this paper. Simulation studies carried out on the Power System Computer Aided Design/ Electro Magnetic Transient in DC System (PSCAD/EMTDC) software are used to demonstrate the effectiveness and efficiency of the proposed ADP-based online VIC parameter tuning methodology. Wentao Guo 0002, Feng Liu 0014, Jennie Si, Shengwei Mei |
IJCNN | 3 |
| 2013 | Stability of direct heuristic dynamic programming for nonlinear tracking control using PID neural networkabstractThe issue of designing a high performance controller to track a desired system trajectory is one of most important problems in control theory and practice. More recently, there has been a growing interest in the study of tracking control problem. In this paper, we discuss the design and stability properties of a special approximate/adaptive dynamic programming (ADP) method for a general multiple-input-multiple-output (MIMO) discrete-time nonlinear optimal tracking control problem. The direct heuristic dynamic programming (HDP) design algorithm is firstly derived by incorporating the PID control rule into neural networks (NNs). This design approach considers using not only the typical state variables but also their derivatives and cumulative sums as inputs to the controller output. It is therefore expected to retain PID controller properties with additional learning capability. Moreover, our nonlinear control problem is formulated under a general condition that system nonlinearity is unknown and therefore it introduces modelling errors for the controller design. By using a Lyapunov stability construct, we provide new results of uniformly ultimately boundedness (UUB) for the proposed PIDNN-based direct HDP controller in discrete-time nonlinear tracking setting with desired tracking performance. Xiong Luo, Jennie Si |
IJCNN | 2 |
| 2012 | A neural correlate to learning decision and control using functional synaptic efficacyabstractHow interacting neurons give rise to meaningful behavior is an ultimate challenge in neuroscience. Synaptic connections lead to interacting neural ensemble activities. As one of the spike time coding scheme, neuronal interactions have been studied intensively. Several algorithms based on neuron pair-wise analysis have been proposed to estimate and study the interaction strength between neurons. Cross correlation, mutual information, and Granger causality are some of the examples. However, these methods are mathematical measures that can not distinguish if there is a functionally direct connection between a neuron pair. The network likelihood model on the other hand takes into account interconnectivity among a neural ensemble. It not only renders the interconnection strength between neurons but also accounts for physical connectivity between two neurons. Using this modeling approach to estimating neuronal connection, the current practice utilizes the maximum likelihood estimation, which is computationally expensive. In this study, we propose a new estimation algorithm for the spike firing probability model using a perceptron bank. This new model not only is computationally efficient, we were also able to interpret neural data from rat's motor cortices in relation to rat's learning decision and control behavior. Specifically the proposed perceptron bank was created based on simultaneous multi-channel chronic recordings from rats motor cortical areas while rats learned to perform a cue directed paddle press task. Our results show that significant changes (p = 0.1%) in functional neural synaptic efficacies from excitatory to inhibitory took place while rats learned to perform the decision and control task. This may indicate that neural plasticity and neural adaptation represented in temporal firing patterns is underlying the behavioral learning process. Chenhui Yang, Jennie Si |
IJCNN | 5 |
| 2012 | A boundedness result for the direct heuristic dynamic programming
Feng Liu 0014, Jennie Si, Wentao Guo 0002, Shengwei Mei |
Neural Networks | 3 |
| 2011 | Belief function model for reliable optimal set estimation of transition matrices in discounted infinite-horizon Markov decision processesabstractWe study finite-state, finite-action, discounted infinite-horizon Markov decision processes with uncertain correlated transition matrices in deterministic policy spaces. To efficiently implement an approximate robust policy iteration algorithm for computing a robust optimal or near-optimal policy, a reliable and tight set estimate of the parameters of the transition matrix is needed in advance. However, observation samples on state transitions may be small. Prior information on the parameter space may be incomplete or unavailable. In such cases, a commonly used maximum a posterior (MAP) model may not provide a reliable optimal set estimate of the parameters. In this paper, using the advantages of Dempster-Shafer theory of evidence over Bayesian theory, a belief function model is proposed based on minimizing the cardinality of a set estimate. This new model can give a more reliable optimal solution to cover the true parameters than the MAP model. It degenerates to the MAP model when prior information on the parameter space is complete or prior information is unavailable but observation samples on state transitions are large enough. Moreover, we create a concept of principle components to characterize large observation samples so that both models result in the same reliable and tight results. The computation complexity of the new model is also discussed. Baohua Li, Jennie Si |
IJCNN | 2 |
| 2011 | Direct heuristic dynamic programming with augmented statesabstractThis paper addresses a design issue of an approximate dynamic programming structure and its respective convergence property. Specifically, we propose to impose a PID structure to the action and critic networks in the direct heuristic dynamic programming (direct HDP) online learning controller. We demonstrate that the direct HDP with such PID augmented states improves convergence speed and that it out performs the traditional PID even though the learning controller may be initialized to be like a PID. Also for the first time, by using a Lyapnov approach we show that the action and critic network weights retain the property of uniformly ultimate boundedness (UUB) under mild conditions. Feng Liu 0014, Jennie Si, Shengwei Mei |
IJCNN | 3 |
| 2011 | A Multiscale Correlation of Wavelet Coefficients Approach to Spike DetectionabstractExtracellular chronic recordings have been used as important evidence in neuroscientific studies to unveil the fundamental neural network mechanisms in the brain. Spike detection is the first step in the analysis of recorded neural waveforms to decipher useful information and provide useful signals for brain-machine interface applications. The process of spike detection is to extract action potentials from the recordings, which are often compounded with noise from different sources. This study proposes a new detection algorithm that leverages a technique from wavelet-based image edge detection. It utilizes the correlation between wavelet coefficients at different sampling scales to create a robust spike detector. The algorithm has one tuning parameter, which potentially reduces the subjectivity of detection results. Both artificial benchmark data sets and real neural recordings are used to evaluate the detection performance of the proposed algorithm. Compared with other detection algorithms, the proposed method has a comparable or better detection performance. In this letter, we also demonstrate its potential for real-time implementation. Chenhui Yang, Byron Olson, Jennie Si |
Neural Comput. | 3 |
| 2010 | High performance spike detection and sorting using neural waveform phase information and SOM clusteringabstractNeural spike detection is the very first step in the analysis of recorded neural waveforms for brain machine interface applications and for neuroscientific studies. Spike detection accuracy and algorithm robustness is an important consideration in developing detection algorithms. For real neural recording data without respective ground truth, the evaluation of detection performance is a challenge. In the present paper we evaluate the detections by inspecting the detected spike waveforms for their compliance with neural spike electrophysiological properties. After classifying similar waveforms into one cluster, those qualified detections are determined to be spikes with high confidence. This new spike detection evaluation method is based on using the waveform phase information for cluster analysis. By including clustering as an integral step in the detection algorithm, we can refine detection results and improve detection performance. The new algorithm is easy to implement and is effective as demonstrated using both artificial and real neural waveforms. Chenhui Yang, Jennie Si |
IJCNN | 3 |
| 2010 | Approximate robust policy iteration using multilayer perceptron neural networks for discounted infinite-horizon Markov decision processes with uncertain correlated transition matricesabstractWe study finite-state, finite-action, discounted infinite-horizon Markov decision processes with uncertain correlated transition matrices in deterministic policy spaces. Existing robust dynamic programming methods cannot be extended to solving this class of general problems. In this paper, based on a robust optimality criterion, an approximate robust policy iteration using a multilayer perceptron neural network is proposed. It is proven that the proposed algorithm converges in finite iterations, and it converges to a stationary optimal or near-optimal policy in a probability sense. In addition, we point out that sometimes even a direct enumeration may not be applicable to addressing this class of problems. However, a direct enumeration based on our proposed maximum value approximation over the parameter space is a feasible approach. We provide further analysis to show that our proposed algorithm is more efficient than such an enumeration method for various scenarios. Baohua Li, Jennie Si |
IEEE Trans. Neural Networks | 2 |
| 2009 | An advanced spike detection and sorting systemabstractThis paper proposes a comprehensive new neural spike detection and sorting system. As a critical first step to all neuroscientific studies of the nervous system using chronically implanted electrodes in the brain areas of interest, high performance neural spike detection and sorting from the massive amount of continuously recorded neural data is a challenging task, especially in real time applications such as brain machine interface. Many existing spike detection and sorting systems use simple thresholding as the first step to admit a large number of possible spikes for further sorting using various clustering algorithms. Significant efforts have gone into developing sophisticated sorting algorithms, many of which are time consuming in applications. In this paper, we develop a new system that is based on a reliable detection algorithm using multiple correlations of wavelet coefficients, which is robust and can be implemented in real time. Because of the advanced detection step, the system becomes less demanding on the performance of the sorting or clustering algorithms. This has simplified the overall system, and made real time interface and other real time applications possible. We tested the newly proposed system extensively, compared with several popular systems including commercial packages. While most thresholding based detection systems usually create a large number of false alarms, test results show that our proposed system on both artificial and real neural data have produced few false alarms but with high detection rates. Chenhui Yang, Jennie Si |
IJCNN | 3 |
| 2009 | Direct Heuristic Dynamic Programming for Nonlinear Tracking Control With Filtered Tracking ErrorabstractThis paper makes use of the direct heuristic dynamic programming design in a nonlinear tracking control setting with filtered tracking error. A Lyapunov stability approach is used for the stability analysis of the tracking system. It is shown that the closed-loop tracking error and the approximating neural network weight estimates retain the property of uniformly ultimate boundedness under the presence of neural network approximation error and bounded unknown disturbances under certain conditions. Lei Yang 0001, Jennie Si, Konstantinos S. Tsakalis, Armando A. Rodriguez |
IEEE Trans. Syst. Man Cybern. Part B | 2 |
| 2008 | A nonlinear adaptive regression process for noise corrupt imagesabstractMost existing nonlinear regression filtering techniques for image denoising are claimed to be edge preserving without considering the pixel position information. This will cause speckling effects on the denoised image and inconsistent smoothing in the vicinity of texture-rich areas. This paper proposes a novel denoising method to address this problem. The proposed method removes the low to intermediate noise using edge-preserving range filtering, thereby removing short, false edges. The updated edge map is used for subsequent filtering in which pixel intensities are smoothed according to their minimum distance to the closest edge point. This procedure is carried out in an iterative scheme until the edge map stabilizes. We compare existing denoising algorithms with the proposed method. Experimental results validate the effectiveness and efficiency of the proposed method. Nan Jiang 0011, Changchun Li, Jennie Si, Glen P. Abousleman |
ICASSP | 3 |
| 2008 | Robust target detection and tracking in outdoor infrared videoabstractAutomated tracking of targets within outdoor infrared (IR) video sequences poses a host of challenges. These include automatic gain adjustment in the IR camera, extreme granularity, large luminance changes, and uncontrolled environmental factors such as moving foliage, animals, and birds, among others. To address these problems, we present an IR video target tracking system for stationary cameras that learns and divides the video frames into reliable and unreliable regions. A difference-frame-based method can recognize moving regions with high sensitivity and reliably discern background clutter from target motion. A low-complexity target validation process is presented, which in conjunction with the reliable region masking, dramatically reduces the number of false alarms. We demonstrate the outstanding performance of the proposed system using real-world IR video sequences with difficult background motion clutter, as well as with small and blurred moving targets. Changchun Li, Nan Jiang 0011, Jennie Si, Glen P. Abousleman |
ICASSP | 3 |
| 2008 | Interval least-squares filtering with applications to robust video target trackingabstractAn interval recursive least-squares (RLS) filter is developed to produce state estimation and prediction by narrow intervals, in which true values are contained with high confidence. The interval filter is robust to variations of the filter parameters and state observations. Using this filter, a video target tracking algorithm is proposed to estimate the target position in each frame. The tracking algorithm is robust to both noise in the video sequence and estimation error of the affine model. The experiments show that the tracking algorithm using the interval RLS filter outperforms that using an RLS filter. Baohua Li, Changchun Li, Jennie Si, Glen P. Abousleman |
ICASSP | 3 |
| 2008 | Special Issue on Advances in Neural Networks Research: IJCNN'07
David G. Brown, Jennie Si, Ron Sun |
Neural Networks | 2 |
| 2008 | Orthogonal Rotation-Invariant Moments for Digital Image ProcessingabstractOrthogonal rotation-invariant moments (ORIMs), such as Zernike moments, are introduced and defined on a continuous unit disk and have been proven powerful tools in optics applications. These moments have also been digitized for applications in digital image processing. Unfortunately, digitization compromises the orthogonality of the moments and, therefore, digital ORIMs are incapable of representing subtle details in images and cannot accurately reconstruct images. Typical approaches to alleviate the digitization artifact can be divided into two categories: 1) careful selection of a set of pixels as close approximation to the unit disk and using numerical integration to determine the ORIM values, and 2) representing pixels using circular shapes such that they resemble that of the unit disk and then calculating ORIMs in polar space. These improvements still fall short of preserving the orthogonality of the ORIMs. In this paper, in contrast to the previous methods, we propose a different approach of using numerical optimization techniques to improve the orthogonality. We prove that with the improved orthogonality, image reconstruction becomes more accurate. Our simulation results also show that the optimized digital ORIMs can accurately reconstruct images and can represent subtle image details. Huibao Lin, Jennie Si, Glen P. Abousleman |
IEEE Trans. Image Process. | 2 |
| 2008 | Direct Heuristic Dynamic Programming for Damping Oscillations in a Large Power SystemabstractThis paper applies a neural-network-based approximate dynamic programming method, namely, the direct heuristic dynamic programming (direct HDP), to a large power system stability control problem. The direct HDP is a learning- and approximation-based approach to addressing nonlinear coordinated control under uncertainty. One of the major design parameters, the controller learning objective function, is formulated to directly account for network-wide low-frequency oscillation with the presence of nonlinearity, uncertainty, and coupling effect among system components. Results include a novel learning control structure based on the direct HDP with applications to two power system problems. The first case involves static var compensator supplementary damping control, which is used to provide a comprehensive evaluation of the learning control performance. The second case aims at addressing a difficult complex system challenge by providing a new solution to a large interconnected power network oscillation damping control problem that frequently occurs in the China Southern Power Grid. Chao Lu 0009, Jennie Si, Xiaorong Xie |
IEEE Trans. Syst. Man Cybern. Part B | 2 |
| 2007 | Approximate Robust Policy Iteration for Discounted Infinite-Horizon Markov Decision Processes with Uncertain Stationary Parametric Transition MatricesabstractWe consider Markov decision processes with finite states, finite actions, and discounted infinite-horizon cost in the deterministic policy space. State transition matrices are uncertain but with stationary parameterization. The uncertainty in transition matrices signifies realistic considerations that an accurate system model is not available for the controller design due to limitations in estimation methods and model deficiencies. Based on the quadratic total value function formulation, two approximate robust policy iterations are developed, the performance errors of which are guaranteed to be within an arbitrarily small error bound. The two approximations make use of iterative aggregation and multilayer perceptron, respectively. It is proved that the robust policy iteration based on approximation with iterative aggregation converges surely to a stationary optimal or near-optimal policy, and also that under some conditions the robust policy iteration based on approximation with multilayer perceptron converges in a probability sense to a stationary near-optimal policy. Furthermore, under some assumptions, the stationary solutions are guaranteed to be near-optimal in the deterministic policy space. Baohua Li, Jennie Si |
IJCNN | 2 |
| 2007 | Convergence of Direct Heuristic Dynamic Programming in Power System Stability ControlabstractIn this paper a neural network-based approximate dynamic programming method, namely direct heuristic dynamic programming (direct HDP), is applied to power system stability control. Direct HDP makes use of learning and approximation to address nonlinear system control problems under uncertainty. The contribution of the paper includes a convergence proof of the direct HDP algorithm using an LQR framework. Under this setting, the paper proposes a direct HDP learning control algorithm for a static var compensator (SVC) supplementary damping control in a standard benchmark power system. The results are used to evaluate the online learning ability of the proposed direct HDP controller, and also to demonstrate that the learning controller does converge to the theoretical limit as derived. Chao Lu 0009, Jennie Si, Xiaorong Xie, Jie Song 0002 |
IJCNN | 2 |
| 2007 | Adaptation in Neural Activity for Directional ControlabstractIn freely moving rats, motor cortical recordings enabled the use of a closed loop system to replace paddle pressing for a directional task. In this system, firing rates were estimated from several (8-10) motor cortical neurons at several consecutive time points. These firing rates were concatenated to form a neural activity vector (NAV). The NAV was used as input to a previously trained support vector machine (SVM) classifier. The decision function value obtained from the SVM was then used to determine which relay should be activated to produce paddle pressing signals in the task. Animals were able to use this interface immediately and significant changes in neural activity arose in a single, 45 minute, experimental session. Neural data from several subjects was examined for changes from the calibration phase to the late cortically controlled phase. Detailed analysis shows that NAVs changed significantly from the calibration phase to the cortically controlled phase, furthermore, the decision function values arising from these NAVs changed in interesting ways. By examining which neurons and times (dimensions of the NAV) were selected by the SVM to have significant impact on the decision function value as well as which dimensions of the NAV changed significantly, a mechanism of adaptation begins to emerge in which the SVM properly assigns high importance to dimensions that easily predict the desired output, however, under closed loop control, the animal selects a small number of neurons (at most or all times) and chooses to make the firing rates more distinguishable. These differences offer insight into how the rats and their SVMs collaborated to create a useable interface. Byron Olson, Jennie Si |
IJCNN | 2 |
| 2007 | Performance Analysis of Direct Heuristic Dynamic Programming using Control-Theoretic MeasuresabstractApproximate dynamic programming (ADP) has been widely studied from several important perspectives: algorithm development, learning efficiency measured by success or failure statistics, convergence rate, and learning error bounds. Given that many learning benchmarks used in ADP or reinforcement learning studies are control problems, it is important and necessary to examine the learning controllers from a control-theoretic perspective. This paper makes use of direct heuristic dynamic programming (direct HDP) and several benchmark examples to introduce a unique analytical framework that can be extended to other learning control paradigms and other complex control problems. The sensitivity analysis and the linear quadratic regulator (LQR) design are used in the paper for two purposes: to gauge direct HDP performance characteristics and to provide guidance toward designing better learning controllers. This gauge however does not limit the direct HDP to be effective only as a linear controller. Toward this end, applications of the direct HDP for nonlinear control problems beyond sensitivity analysis and the confines of LQR have been developed and compared with LQR design for command following and internal system parameter changes. Lei Yang 0001, Jennie Si, Konstantinos S. Tsakalis, Armando A. Rodriguez |
IJCNN | 2 |
| 2006 | Hierarchical Region-Based Image Registration in Scale SpaceabstractThis paper presents a new method for image registration for real natural scenes. The method is based on the observation that most natural scenes are actually 3-D. If the assumption of weak perspective is violated, the error in image registration induced by parallax is increased, leading artifacts or blurs in image mosaic. Our method first applies the affine-invariant point detector in scale space. After clustering feature point pairs, an initial global transformation is formed based on majority correspondence. The global transformation is evaluated in each region at certain scale determined by inliers. The global model is optimized for local registration by minimizing least square error. This method is more robust than standard image registration algorithms on images subject to uncalibrated camera motion Nan Jiang 0011, Jennie Si, Glen P. Abousleman |
ICASSP (2) | 2 |
| 2005 | Migrating Orthogonal Rotation-Invariant Moments from Continuous to Discrete SpaceabstractOrthogonality and rotation invariance are important feature properties in digital signal processing. Orthogonality enables a target to be represented by a compact number of features, while rotation invariance results in unique features for a target with different orientations. The orthogonal, rotation-invariant moments (ORIMs), such as Zernike, pseudo-Zernike, and orthogonal Fourier-Melling moments, are defined in continuous space. These ORIMs have been digitized and have been demonstrated effectively for some digital imagery applications. However, digitization compromises the orthogonality of the moments, and hence, reduces their precision. Therefore, digital ORIMs are incapable of representing the fine details of images. In this paper, we propose a numerical optimization technique to improve the orthogonality of the digital ORIMs. Simulation results show that our optimized digital ORIMs can be used to reproduce subtle details of images. Huibao Lin, Jennie Si, Glen P. Abousleman |
ICASSP (2) | 2 |
| 2005 | An analysis of gradient-based policy iterationabstractRecently, a system theoretic framework for learning and optimization has been developed that shows how many approximate dynamic programming paradigms such as perturbation analysis, Markov decision processes, and reinforcement learning are very closely related. Using this system theoretic framework, a new optimization technique called gradient-based policy iteration (GBPI) has been developed. In this paper, we show how GBPI iteration can be extended to partially observable Markov decision processes (POMDPs). We also develop the value iteration analogue of GBPI and show that this new version of value iteration, extended to POMDPs, not only theoretically acts like value iteration but also does so numerically. James Dankert, Lei Yang 0001, Jennie Si |
IJCNN | 3 |
| 2005 | Bidirectional labeling and registration scheme for grayscale image segmentationabstractIn this paper, we introduce a new image segmentation scheme that is based on bidirectional labeling and registration and prove that its segmentation performance is equivalent to that of the conventional watershed segmentation algorithm. The proposed bidirectional labeling and registration scheme, which we refer to as bidirectional labeling and registration scheme (BIDS), involves only linear scans of image pixels. It uses one-dimensional operations rather than the queues that are used in traditional segmentation algorithms, which are two-dimensional problems. BIDS also provides unique labels for individual homogeneous regions. In addition to achieving the same segmentation results, BIDS is four times less computationally complex than the conventional watershed by immersion technique. Lei Ma 0002, Xiao-Ping Zhang 0002, Jennie Si, Glen P. Abousleman |
IEEE Trans. Image Process. | 3 |
| 2003 | Bi-directional gradient labeling and registration for gray-scale image segmentationabstractWatershed is one of the commonly used methods for image segmentation. In this paper, we introduce a new segmentation scheme based on bi-directional labeling and registration and prove that its segmentation performance is equivalent to that of conventional watershed but it is much more computational efficient The bi-directional labeling and registration scheme, which will be referred to as BIDS, involves only linear scans of image pixels. It uses one dimensional operations instead of queues while traditional segmentation algorithms are two dimensional problems. BIDS also provides unique labels for each homogeneous regions. In addition to achieving the same segmentation results as conventional watershed, BIDS is four times less computationally complex than the conventional watersheds by immersion. Lei Ma 0002, Xiao-Ping Zhang 0002, Jennie Si, Glen P. Abousleman |
ICIP (1) | 3 |
| 2003 | Helicopter trimming and tracking control using direct neural dynamic programmingabstractThis paper advances a neural-network-based approximate dynamic programming control mechanism that can be applied to complex control problems such as helicopter flight control design. Based on direct neural dynamic programming (DNDP), an approximate dynamic programming methodology, the control system is tailored to learn to maneuver a helicopter. The paper consists of a comprehensive treatise of this DNDP-based tracking control framework and extensive simulation studies for an Apache helicopter. A trim network is developed and seamlessly integrated into the neural dynamic programming (NDP) controller as part of a baseline structure for controlling complex nonlinear systems such as a helicopter. Design robustness is addressed by performing simulations under various disturbance conditions. All designs are tested using FLYRT, a sophisticated industrial scale nonlinear validated model of the Apache helicopter. This is probably the first time that an approximate dynamic programming methodology has been systematically applied to, and evaluated on, a complex, continuous state, multiple-input multiple-output nonlinear system with uncertainty. Though illustrated for helicopters, the DNDP control system framework should be applicable to general purpose tracking control. Russell Enns, Jennie Si |
IEEE Trans. Neural Networks | 2 |
| 2002 | Knowledge-based hierarchical region-of-interest detectionabstractDetecting regions of interest (ROIs) in a complex image is a critical step in many image processing applications. In this paper, we present a new algorithm that addresses several challenges in ROI detection. The novelty of our algorithm includes: (i) every ROI contains one and only one object; (ii) the detected ROIs can have irregular shapes as opposed to the rectangular shapes that are typical of other algorithms; (iii) the algorithm is applicable to images that contain connected objects, or when the objects are broken into pieces; (iv) the algorithm is not sensitive to contrast levels in the image, and is robust to noise. These characteristics make the proposed algorithm applicable to low-resolution, real-world imagery without costly post-processing. The proposed algorithm is shown to provide outstanding performance with low-quality imagery, and is shown to be fast and robust. Huibao Lin, Jennie Si, Glen P. Abousleman |
ICASSP | 2 |
| 2001 | Global exponential stability of neural networks with globally Lipschitz continuous activations and its application to linear variational inequality problemabstractThis paper investigates the existence, uniqueness, and global exponential stability (GES) of the equilibrium point for a large class of neural networks with globally Lipschitz continuous activations including the widely used sigmoidal activations and the piecewise linear activations. The provided sufficient condition for GES is mild and some conditions easily examined in practice are also presented. The GES of neural networks in the case of locally Lipschitz continuous activations is also obtained under an appropriate condition. The analysis results given in the paper extend substantially the existing relevant stability results in the literature, and therefore expand significantly the application range of neural networks in solving optimization problems. As a demonstration, we apply the obtained analysis results to the design of a recurrent neural network (RNN) for solving the linear variational inequality problem (VIP) defined on any nonempty and closed box set, which includes the box constrained quadratic programming and the linear complementarity problem as the special cases. It can be inferred that the linear VIP has a unique solution for the class of Lyapunov diagonally stable matrices, and that the synthesized RNN is globally exponentially convergent to the unique solution. Some illustrative simulation examples are also given. Xue-Bin Liang, Jennie Si |
IEEE Trans. Neural Networks | 2 |
| 2001 | Online learning control by association and reinforcementabstractThis paper focuses on a systematic treatment for developing a generic online learning control system based on the fundamental principle of reinforcement learning or more specifically neural dynamic programming. This online learning system improves its performance over time in two aspects: 1) it learns from its own mistakes through the reinforcement signal from the external environment and tries to reinforce its action to improve future performance; and 2) system states associated with the positive reinforcement is memorized through a network learning process where in the future, similar states will be more positively associated with a control action leading to a positive reinforcement. A successful candidate of online learning control design is introduced. Real-time learning algorithms is derived for individual components in the learning system. Some analytical insight is provided to give guidelines on the learning process took place in each module of the online learning control system. Jennie Si, Yu-Tsung Wang |
IEEE Trans. Neural Networks | 1 |
| 2000 | On-Line Learning Control by Association and ReinforcementabstractThis paper focuses on a systematic treatment for developing a generic online learning control system based on the fundamental principle of reinforcement learning or more specifically neuro-dynamic programming. This real time learning system improves its performance over time in two aspects: it learns from its own mistakes through the reinforcement signal from the external environment and try to reinforce its action to improve future performance; and system's state associated with the positive reinforcement is memorized through a network learning process where in the future, similar states will be more positively associated with a control action leading to a positive reinforcement. Two successful candidates of online learning control designs are introduced. Real time learning algorithms can be derived for individual components in the learning system. Some analytical insights are provided to give some guidelines on the entire online learning control system. Jennie Si, Yu-Tsung Wang |
IJCNN (3) | 1 |
| 2000 | Dynamic topology representing networks
Jennie Si, Siming Lin 0001, M.-A. Vuong |
Neural Networks | 1 |
| 1999 | Subset-based training and pruning of sigmoid neural networks
Guian Zhou, Jennie Si |
Neural Networks | 2 |
| 1998 | Weight Value Convergence of the SOM Algorithm for Discrete InputabstractSome insights on the convergence of the weight values of the self-organizing map (SOM) to a stationary state in the case of discrete input are provided. The convergence result is obtained by applying the Robbins-Monro algorithm and is applicable to input-output maps of any dimension. Siming Lin 0001, Jennie Si |
Neural Comput. | 2 |
| 1998 | A Systematic and Effective Supervised Learning Mechanism Based on Jacobian Rank DeficiencyabstractMost neural network applications rely on the fundamental approximation property of feedforward networks. Supervised learning is a means of implementing this approximate mapping. In a realistic problem setting, a mechanism is needed to devise this learning process based on available data, which encompasses choosing an appropriate set of parameters in order to avoid overfitting, using an efficient learning algorithm measured by computation and memory complexities, ensuring the accuracy of the training procedures as measured by the training error, and testing and cross-validation for generalization. We develop a comprehensive supervised learning algorithm to address these issues. The algorithm combines training and pruning into one procedure by utilizing a common observation of Jacobian rank deficiency in feedforward networks. The algorithm not only reduces the training time and overall complexity but also achieves training accuracy and generalization capabilities comparable to more standard approaches. Extensive simulation results are provided to demonstrate the effectiveness of the algorithm. Guian Zhou, Jennie Si |
Neural Comput. | 2 |
| 1998 | Neural network-based control design: an LMI approachabstractIn this paper, we address a neural-network-based control design for a discrete-time nonlinear system. Our design approach is to approximate the nonlinear system with a multilayer perceptron of which the activation functions are of the sigmoid type symmetric to the origin. A linear difference inclusion representation is then established for this class of approximating neural networks and is used to design a state-feedback control law for the nonlinear system based on the certainty equivalence principle. The control design equations are shown to be a set of linear matrix inequalities where a convex optimization algorithm can be applied to determine the control signal. Further, the stability of the closed-loop is guaranteed in the sense that there exists a unique global attraction region in the neighborhood of the origin to which every trajectory of the closed-loop system converges. Finally, a simple example is presented so as to illustrate our control design procedure. Suttipan Limanond, Jennie Si |
IEEE Trans. Neural Networks | 2 |
| 1998 | Advanced neural-network training algorithm with reduced complexity based on Jacobian deficiencyabstractIn this paper we introduce an advanced supervised training method for neural networks. It is based on Jacobian rank deficiency and it is formulated, in some sense, in the spirit of the Gauss-Newton algorithm. The Levenberg-Marquardt algorithm, as a modified Gauss-Newton, has been used successfully in solving nonlinear least squares problems including neural-network training. It outperforms (in terms of training accuracy, convergence properties, overall training time, etc.) the basic backpropagation and its variations with variable learning rate significantly, however, with higher computation and memory complexities within each iteration. The new method developed in this paper is aiming at improving convergence properties, while reducing the memory and computation complexities in supervised training of neural networks. Extensive simulation results are provided to demonstrate the superior performance of the new algorithm over the Levenberg-Marquardt algorithm. Guian Zhou, Jennie Si |
IEEE Trans. Neural Networks | 2 |
| 1997 | Blind equalization with a linear feedforward neural network
Xi-Ren Cao, Jie Zhu 0001, Jennie Si |
ESANN | 3 |
| 1997 | Self-Organization of Firing Activities in Monkey's Motor Cortex: Trajectory Computation from Spike SignalsabstractThe population vector method has been developed to combine the simultaneous direction-related activities of a population of motor cortical neurons to predict the trajectory of the arm movement. In this article, we consider a self-organizing model of a neural representation of the arm trajectory based on neuronal discharge rates. As self-organizing feature map (SOFM) is used to select the optimal set of weights in the model to determine the contribution of an individual neuron to an overall movement representation. The correspondence between movement directions and discharge patterns of the motor cortical neurons is established in the output map. The topology-preserving property of the SOFM is used to analyze the recorded data of a behaving monkey. The data used in this analysis were taken while the monkey was tracing spirals and doing center-->out movements. The arm trajectory could be well predicted using such a statistical model based on the motor cortex neuronal firing information. The SOFM method is compared with the population vector method, which extracts information related to trajectory by assuming that each cell has a fixed preferred direction during the task. This implies that these cells are acting along lines labeled only for direction. However, extradirectional information is carried in these cell responses. The SOFM has the capability of extracting not only direction-related information but also other parameters that are consistently represented in the activity of the recorded population of cells. Siming Lin 0001, Jennie Si, A. B. Schwartz |
Neural Comput. | 2 |
| 1996 | Approximation Errors of State and Output Trajectories Using Recurrent Neural Networks
Binfan Liu, Jennie Si |
ICANN | 2 |
| 1996 | Improving Neural Network Training Based on Jacobian Rank Deficiency
Guian Zhou, Jennie Si |
ICANN | 2 |
| 1995 | Analysis and synthesis of a class of discrete-time neural networks with multilevel threshold neuronsabstractIn contrast to the usual types of neural networks which utilize two states for each neuron, a class of synchronous discrete-time neural networks with multilevel threshold neurons is developed. A qualitative analysis and a synthesis procedure for the class of neural networks considered constitute the principal contributions of this paper. The applicability of the present class of neural networks is demonstrated by means of a gray level image processing example, where each neuron can assume one of sixteen values. When compared to the usual neural networks with two state neurons, networks which are endowed with multilevel neurons will, in general, for a given application, require fewer neurons and thus fewer interconnections. This is an important consideration in VLSI implementation. Jennie Si, Anthony N. Michel |
IEEE Trans. Neural Networks | 1 |
| 1994 | The best approximation to C2 functions and its error bounds using regular-center Gaussian networksabstractGaussian neural networks are considered to approximate any C(2) function with support on the unit hypercube I(m)=[0, 1](m) in the sense of best approximation. An upper bound (O(N(-2))) of the approximation error is obtained in the present paper for a Gaussian network having N(m) hidden neurons with centers defined on a regular mesh in I(m). Binfan Liu, Jennie Si |
IEEE Trans. Neural Networks | 2 |
| 1993 | A near optimal algorithm for image compression using Gabor expansion
Jean-Louis Gouronc, Jennie Si |
ISCAS | 2 |
| 1993 | The best approximation to C2 functions and its error bounds using Gaussian hidden units
Binfan Liu, Jennie Si |
ISCAS | 2 |