EDBT 2026 Demo / reviewers in the wild / expert
Juyang Weng
dblp:w/JuyangWeng · also John (Juyang) Weng
· DBLP profile ↗
127ranked-venue papers
51as first author
6since 2021 · last 2023
0000-0003-1383-3872ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 109 · 45 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 34 · 18 first-authorApplied, interdisciplinary, general and emerging computing · 7 · 3 first-authorSystems, architecture and hardware · 2Human-computer interaction and ubiquitous computing · 2Databases, data management, data science and information retrieval · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
26 papers |
3D vision · 41% Segmentation and scene understanding · 14% Representation and self-supervised learning · 12% | |
| Computer graphics and multimedia
7 papers |
Image and video processing · 48% Multimedia analysis and retrieval · 43% Computational photography and imaging · 9% |
Topics — the 30 heaviest of 49, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › 3D vision
structure from motion |
0.1 | 11 | 1997 | Transitory Image Sequences, Asymptotic Properties, and Estimation of Motion and Structure · IEEE Trans. Pattern Anal. Mach. Intell. 1997 Integration of transitory image sequences · CVPR 1994 Optimal Motion and Structure Estimation · IEEE Trans. Pattern Anal. Mach. Intell. 1993 |
Computer vision › 3D vision › structure from motion
structure and motion estimation |
0.1 | 9 | 1997 | Transitory Image Sequences, Asymptotic Properties, and Estimation of Motion and Structure · IEEE Trans. Pattern Anal. Mach. Intell. 1997 Integration of transitory image sequences · CVPR 1994 Optimal Motion and Structure Estimation · IEEE Trans. Pattern Anal. Mach. Intell. 1993 |
Machine learning › Learning paradigms
incremental learning |
0.0 | 1 | 2004 | Obstacle Avoidance through Incremental Learning with Attention Selection · ICRA 2004 |
Robotics › Robot navigation and mapping
obstacle avoidance |
0.0 | 1 | 2004 | Obstacle Avoidance through Incremental Learning with Attention Selection · ICRA 2004 |
Computer vision › Segmentation and scene understanding
image segmentation |
0.0 | 2 | 1999 | A Learning-Based Prediction-and-Verification Segmentation Scheme for Hand Sign Image Sequence · IEEE Trans. Pattern Anal. Mach. Intell. 1999 Learning Recognition and Segmentation Using the Cresceptron · Int. J. Comput. Vis. 1997 |
Machine learning › Representation and self-supervised learning › representation learning › dimensionality reduction
principal component analysis |
0.0 | 1 | 2003 | Candid Covariance-Free Incremental Principal Component Analysis · IEEE Trans. Pattern Anal. Mach. Intell. 2003 |
Multimedia analysis and retrieval
image retrieval |
0.0 | 2 | 1999 | Hierarchical Discriminant Analysis for Image Retrieval · IEEE Trans. Pattern Anal. Mach. Intell. 1999 Using Discriminant Eigenfeatures for Image Retrieval · IEEE Trans. Pattern Anal. Mach. Intell. 1996 |
Computer vision › Image recognition and object detection
object recognition |
0.0 | 2 | 1999 | Hierarchical Discriminant Analysis for Image Retrieval · IEEE Trans. Pattern Anal. Mach. Intell. 1999 Learning recognition and segmentation of 3-D objects from 2-D images · ICCV 1993 |
Computer vision › Segmentation and scene understanding › image segmentation › model-based segmentation
deformable model segmentation |
0.0 | 1 | 1999 | A Learning-Based Prediction-and-Verification Segmentation Scheme for Hand Sign Image Sequence · IEEE Trans. Pattern Anal. Mach. Intell. 1999 |
Image and video processing
image segmentation |
0.0 | 1 | 1999 | A Learning-Based Prediction-and-Verification Segmentation Scheme for Hand Sign Image Sequence · IEEE Trans. Pattern Anal. Mach. Intell. 1999 |
Computer vision › 3D vision
3d reconstruction |
0.0 | 2 | 1997 | Optimal Registration of Object Views Using Range Data · IEEE Trans. Pattern Anal. Mach. Intell. 1997 Extended structure and motion analysis from monocular image sequences · ICCV 1990 |
Computer vision › 3D vision
feature matching |
0.0 | 2 | 1993 | Image matching using the windowed Fourier phase · Int. J. Comput. Vis. 1993 A theory of image matching · ICCV 1990 |
Computer vision › 3D vision › depth estimation › stereo depth estimation
dense disparity estimation |
0.0 | 2 | 1992 | Matching Two Perspective Views · IEEE Trans. Pattern Anal. Mach. Intell. 1992 A theory of image matching · ICCV 1990 |
Computer vision › 3D vision › stereo vision
stereo matching |
0.0 | 2 | 1992 | Matching Two Perspective Views · IEEE Trans. Pattern Anal. Mach. Intell. 1992 A theory of image matching · ICCV 1990 |
Computer vision › Segmentation and scene understanding › object segmentation
hand segmentation |
0.0 | 1 | 1996 | Hand segmentation using learning-based prediction and verification for hand sign recognition · CVPR 1996 |
Robotics › Robot navigation and mapping › mobile robot navigation
reactive navigation |
0.0 | 1 | 2004 | Obstacle Avoidance through Incremental Learning with Attention Selection · ICRA 2004 |
Machine learning › Representation and self-supervised learning › representation learning › dimensionality reduction › feature selection
discriminative feature selection |
0.0 | 1 | 1995 | Learning-Based Hand Sign Recognition Using SHOSLIF-M · ICCV 1995 |
Computer vision › Segmentation and scene understanding
object segmentation |
0.0 | 1 | 1993 | Learning recognition and segmentation of 3-D objects from 2-D images · ICCV 1993 |
Computer vision › 3D vision
camera calibration |
0.0 | 1 | 1992 | Camera Calibration with Distortion Models and Accuracy Evaluation · IEEE Trans. Pattern Anal. Mach. Intell. 1992 |
Natural language and speech › Machine translation › statistical machine translation
distortion modeling |
0.0 | 1 | 1992 | Camera Calibration with Distortion Models and Accuracy Evaluation · IEEE Trans. Pattern Anal. Mach. Intell. 1992 |
Interaction techniques and input
gesture input |
0.0 | 1 | 1999 | A Learning-Based Prediction-and-Verification Segmentation Scheme for Hand Sign Image Sequence · IEEE Trans. Pattern Anal. Mach. Intell. 1999 |
Accessibility and assistive technology › visual communication
sign language |
0.0 | 1 | 1999 | A Learning-Based Prediction-and-Verification Segmentation Scheme for Hand Sign Image Sequence · IEEE Trans. Pattern Anal. Mach. Intell. 1999 |
Computer vision › 3D vision
depth estimation |
0.0 | 1 | 1990 | A theory of image matching · ICCV 1990 |
Machine learning › Learning theory › statistical estimation
error estimation |
0.0 | 1 | 1989 | Motion and Structure From Two Perspective Views: Algorithms, Error Analysis, and Error Estimation · IEEE Trans. Pattern Anal. Mach. Intell. 1989 |
Machine learning › Learning theory › statistical estimation
optimal estimation |
0.0 | 1 | 1989 | Optimal motion and structure estimation · CVPR 1989 |
Mathematical optimization › continuous optimization
nonlinear optimization |
0.0 | 2 | 1992 | Motion and Structure from Line Correspondences; Closed-Form Solution, Uniqueness, and Optimization · IEEE Trans. Pattern Anal. Mach. Intell. 1992 Camera Calibration with Distortion Models and Accuracy Evaluation · IEEE Trans. Pattern Anal. Mach. Intell. 1992 |
Computer vision › 3D vision › feature matching
line matching |
0.0 | 1 | 1988 | Estimating motion/structure from line correspondences: a robust linear algorithm and uniqueness theorems · CVPR 1988 |
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › parameter estimation
maximum likelihood estimation |
0.0 | 1 | 1988 | Closed-form solution+maximum likelihood: a robust approach to motion and structure estimation · CVPR 1988 |
Image and video processing › stereo vision
stereo matching |
0.0 | 1 | 1988 | Two-view Matching · ICCV 1988 |
Computer vision › 3D vision › multi-view geometry
camera geometry |
0.0 | 3 | 1989 | Optimal motion and structure estimation · CVPR 1989 Estimating motion/structure from line correspondences: a robust linear algorithm and uniqueness theorems · CVPR 1988 Closed-form solution+maximum likelihood: a robust approach to motion and structure estimation · CVPR 1988 |
Methods — techniques the papers use, named apart from their topics
statistical efficiency · 0.1covariance-free incremental estimation · 0.1attention image · 0.1multiple fixations · 0.1prediction-and-verification · 0.0closed-form solution · 0.0laser range finder · 0.0attention selection · 0.0negative log-likelihood · 0.0hierarchical probability model · 0.0doubly clustered subspace learning · 0.0space-tessellation tree · 0.0optimal linear projection · 0.0attention images · 0.0cross-frame estimation · 0.0camera global pose estimation · 0.0principal component analysis · 0.0linear discriminant analysis · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | A Protocol for Testing Conscious Learning RobotsabstractThis is a theoretical paper. The theory [31] and algorithm [32] for conscious learning have been recently published. Developmental scales for human children are well developed. Such scales need to be adapted to testing conscious learning robots. Without such adaptations, future conscious robots lack a standard, even if the theory and algorithm for conscious learning are implemented and refined in the future. This paper discusses such an adaptation, but does not include actual experimental results. It first proposes that an open skull will not allow a conscious brain because the open skull allows a conscious homunculus (human) who takes over the job of consciousness. That is why the currently popular “open skull” machine learning protocols will not produce conscious robots. Then, the paper borrows some of the milestones from human mental development measured in terms of human mental ages. The author hopes that independent laboratories will conduct tests using the milestones suggested here, so as to see whether the new protocol is suited for measuring robotic consciousness. Due to space limitations, this paper does not explain conscious learning. The reader should first read [32] before reading this paper. For four aspects of transfer across milestones, the reader should read [28] (e.g., Sec. 10.3). Juyang Weng |
IJCNN | 1 |
| 2022 | 20 Million-Dollar Problems for Any Brain Models and a Holistic Solution: Conscious LearningabstractThis is a theoretical paper. It raises 20 open problems each of which is estimated to require one million dollars of investment or more. They are (1) the image annotation problem (e.g., retina is without bounding box to learn, unlike ImageNet), (2) the sensorimotor recurrence problem (e.g., all big data sets are invalid), (3) the motor-supervision problem (e.g., impractical to supervise motors throughout lifetime), (4) the sensor calibration problem (e.g., a life calibrates the eyes automatically), (5) the inverse kinematics problem (e.g., a life calibrates all redundant limbs automatically), (6) the government-free problem (i.e., no task-aware homunculus inside a brain), (7) the closed-skull problem (e.g., supervising hidden neurons is biologically implausible), (8) the nonlinear controller problem (e.g., a brain is a nonlinear controller but task-nonspecific), (9) the curse of dimensionality problem (e.g., a set of global features is insufficient for a life), (10) the under-sample problem (i.e., few available examples in a life), (11) the distributed vs. local representations problem (i.e., how both representations emerge), (12) the symbol problem (also called grounding problem, thus must be free from any symbols), (13) the local minima problem (so, avoid error-backprop learning and Post-Selections), (14) the abstraction problem (i.e., require various invariances and transfers), (15) the compositionality problem (e.g., metonymy beyond those composable from sentences), (16) the smooth representations problem (e.g., brain representations are globally smooth), (17) the motivation problem (e.g., including reinforcements and various emotions), (18) the global optimality problem (e.g., avoid catastrophic memory loss and Post-Selections), (19) the auto-programming for general purposes (APFGP) problem, (20) the brain-thinking problem. The paper discusses also why the proposed holistic solution of conscious learning [1], [2] solves each. Juyang Weng |
IJCNN | 1 |
| 2022 | Developmental Network-2: The Autonomous Generation of Optimal Internal-Representation HierarchyabstractIt is very challenging for machine learning methods to reach the goal of general-purpose learning since there are so many complicated situations in different tasks. The learning methods need to generate flexible internal representations for all scenarios met before. The hierarchical internal representation is considered as an efficient way to build such flexible representations. By hierarchy, we mean important local features in the input can be combined to form higher level features with more context. In this work, we analyze how our proposed general-purpose learning framework-the developmental network-2 (DN-2)-autonomously generates internal hierarchy with new mechanisms. Specifically, DN-2 incrementally allocates neuronal resources to different levels of representation during learning instead of handcrafting static boundaries among different levels of representation. We present the mathematical proof to demonstrate that optimal properties in terms of maximum likelihood (ML) are established under the conditions of limited learning experience and resources. The phoneme recognition and real-world visual navigation experiments that are of different modalities and include many different situations are designed to investigate general-purpose learning capability of DN-2. The experimental results show that DN-2 successfully learns different tasks. The formed internal hierarchical representations focus on important features, and the invariant abstract arise from optimal internal representations. We believe that DN-2 is in the right way toward fully autonomous learning. Xiang Wu 0008, Zejia Zheng, Juyang Weng |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2021 | On Post Selection Using Test Sets (PSUTS) in AIabstractThis is a theory paper. It first raises a rarely reported but unethical practice in Artificial Intelligence (AI) called Post Selection Using Test Sets (PSUTS). Consequently, the popular error-backprop methodology in deep learning lacks an acceptable generalization power. All AI methods fall into two broad schools, connectionist and symbolic. PSUTS practices have two kinds, machine PSUTS and human PSUTS. The connectionist school received criticisms for its “scruffiness” due to a huge number of scruffy parameters and now the machine PSUTS; but the seemingly “clean” symbolic school seems more brittle than what is known because of using human PSUTS. This paper formally defines what PSUTS is, analyzes why error-backprop methods with random initial weights suffer from severe local minima, why PSUTS violates well-established research ethics, and how every paper that used PSUTS should have at least transparently reported PSUTS data. For improved transparency in future publications, this paper proposes a new standard for AI metrology, called developmental errors for all networks trained in a project that the selection of the luckiest network depends on, along with Three Conditions: (1) system restrictions, (2) training experience and (3) computational resources. Juyang Weng |
IJCNN | 1 |
| 2021 | On Machine ThinkingabstractArtificial Intelligence (AI) has made much progress, but the existing paradigm for AI is still basically pattern recognition based on a human-handcrafted representation. An AI paradigm shift seems to be necessary to address the machine thinking question raised by Alan Turing over 90 years ago. As a necessary subject of our new conscious learning paradigm introduced 2020, this work deals with general-purpose machine thinking, with planning as a special case, based on new concepts of emergent Super-Turing Machines realized by our proposed neural network models-Developmental Networks (DNs) that have been mathematically proven for its optimally in the sense of Maximum Likelihood (ML). Experimental demonstrations are presented for simulated new mazes in disjoint tests. Xiang Wu 0008, Zejia Zheng, Juyang Weng |
IJCNN | 3 |
| 2021 | Learning to recognize while learning to speak: Self-supervision and developing a speaking motor
Xiang Wu 0008, Juyang Weng |
Neural Networks | 2 |
| 2020 | Autonomous Programming for General Purposes: Theory and ExperimentsabstractSince the birth of AI, symbols have been well accepted as an abstract representation for intelligent systems, even when neural networks are employed. This theoretical work shows the following new methods: (1) Symbols are probably not used by biological brains because their behaviors are not limited by symbols. (2) Autonomous Programming For General Purposes (APFGP) is necessary for scaling-up AI to animal level intelligence. "Autonomous " means inside the skull (or network). (3) A Developmental Network (DN) performs APFGP by learning a super Turing machine, called Grounded, Emergent, Natural, Incremental, Skull-closed, Attentive, Motivated, and Abstractive (GENISAMA) Turing machine. (4) A DN is free of any central controller (e.g., Master Map, convolution, or error back-propagation). (5) The GENISAMA DN does APFGP without using symbols. Experiments are reported for vision guided navigation, auditory recognition, and natural language learning. This is the first conference paper on APFGP. Juyang Weng, Zejia Zheng, Xiang Wu 0008, Juan L. Castro-Garcia |
IJCNN | 1 |
| 2020 | Muscle Vectors as Temporally Dense "Labels"abstractConsider a human who interacts with the physical world to autonomously learn in a task non-specific way through lifetime. It seems obvious that the fully autonomous learner does not have the luxury to have the mother to provide temporally dense state labels, but the context/state at every frame is beneficial (e.g., to generate attention for the next frame). How can we enable the learner to generate frame-wise contexts/states on the fly? Our past work on Developmental Network (DN-1) has shown that frame-wise state labels (e.g., stages within a phoneme) are useful to generate temporally sparse label (the type of the phoneme). However, such dense and sparse labels were handcrafted from a static data set, using human identified frame-wise equivalence. In this paper, we study a conceptually challenging problem - how to enable an autonomous learner to generate frame-wise states autonomously without human handcrafting dense labels at all. We propose that frame-wise muscle actions (e.g., producing a sound) are not only temporally dense and high-dimensional, but also natural as dense labels. However, it is unknown how a neural network can use such high-dimensional vectors as dense labels. In this work, we provide a model for this new issue and experiment with Developmental Network-2 (DN-2) for imitation of audio sequences. Our experimental results showed DN-2 can successfully emerge high-dimensional real-valued vector actions. These actions provide DN-2 with frame-wise temporal context information. This work corresponds to a key step toward our goal to enable the agent to fully autonomously generate frame-wise actions without human-provided dense labels and with only a few human-provided sparse labels. Xiang Wu 0008, Juyang Weng |
IJCNN | 2 |
| 2019 | Emergent Multilingual Language Acquisition Using Developmental NetworksabstractThere has been work on language acquisition but such prior work was based on symbolic representations and non-incremental learning. Neural Networks are meant for incremental learning but their performance has been weak. This situation was mainly due to a "lack of logic" in neural networks. By language acquisition here we mean incremental learning from lifetime experience. Since developmental networks (DN) has clearly understandable emergent "logic" in terms of finite automata and Turing machines, this is the first work on language acquisition based on a clearly understandable emergent Turing machine. We show how symbolic words are represented by patterns instead of (handcrafted) symbols to simulate naturally grounded and emergent inputs. The context as states/actions are also represented by patterns to simulate naturally grounded and emergent inputs. Our work demonstrates that symbolic state features can be fully automated by emergent input-context pattern pairs. This is a step toward fully automated acquisition of language by a grounded robot, but we are not there yet. Juan L. Castro-Garcia, Juyang Weng |
IJCNN | 2 |
| 2019 | The Emergent-Context Emergent-Input Framework for Temporal ProcessingabstractMany temporal processing tasks face great challenges. The handcrafted features and structure used in these temporal processing methods result in the brittle systems. To avoid handcrafted designs, we analyze and propose the EmergentContext Emergent-Input (ECEI) framework in this work. The Developmental Network (DN-1) can be considered as the first ECEI framework that used emergent motor vectors as temporal context (state) for top-down attention. This method suggests that the developmental root of temporal states is at the (open) motor end and always expressible by the motor end, not originally internal and hidden. Indeed, it has enabled the control of a Turing machine to incrementally emerge inside the network. However, DN-1 used a handcrafted partition of motor neurons as concept regions, which does not allow context to automatically emerge. The same is true for hidden (internal) areas other than the motor area. We analyze why this limitation is fundamentally damaging. To overcome this limitation, Developmental Network 2 (DN-2) is proposed with some new biology inspired mechanisms. DN2 is an ECEI framework which can directly use any naturally emerged motor patterns, allowing a freedom of thought that has been largely overlooked. The phoneme recognition experiments are used to investigate DN-2's performance with two different settings about the motor area: handcrafted motor concept zones and emergent motor area. Based on the experimental results, we compared and analyzed such two mechanisms. This work is a step toward our goal of empowering DN-2 to not only free thoughts, but such free-thoughts are also necessary for handling increasingly complicated temporal processing tasks. Xiang Wu 0008, Juyang Weng |
IJCNN | 2 |
| 2019 | Emergent neural turing machine and its visual navigation
Zejia Zheng, Xiang Wu 0008, Juyang Weng |
Neural Networks | 3 |
| 2018 | Entropy as Temporal Information DensityabstractBesides the spatial contents from current sensory inputs, the relevant contexts from past frames are very important for temporal processing tasks (e.g., speech recognition, video analysis, and natural language processing). Our Developmental Network (DN) has demonstrated the ability to learn any emergent Turing Machine (TM), it can learn feature patterns from current and attended past natural inputs as their states. We have shown the dense actions can serve as natural sources of contexts, and the DN can autonomously generate actions as contexts when dealing with sequences. In this work, we use entropy to define and measure the information density of the temporal sequences. We also introduce the "free of labeling" property, which can help DN deal with a large number of states emerging from dense contexts. We experimented with DN for phoneme recognition as the example of auditory modality, but the principles are modality independent. Our experimental results showed the denser contexts extracted from the sequences, the better DN can perform. With the quantization of information density by entropy, we have better understanding of how to provide the contexts when training DN. This work is an important step toward enabling machines to autonomously abstract concrete concepts from contexts through life-long development. Xiang Wu 0008, Zejia Zheng, Juyang Weng |
FUZZ-IEEE | 3 |
| 2018 | Sensorimotor in Space and Time: AuditionabstractProcessing complex temporal sequences is still challenging since existing frameworks are handcrafted (e.g., the cascaded structure of the convolutional neural networks (CNNs)) for each specific problem. Their parameters need to be finely tuned for each specific problem because they use the symbolic representations. In this work, we propose the new Developmental Network 2 (DN-2) to overcome these challenges. The DN-2 uses patterns as representations and shows strengths in conducting abstraction from concrete sensory examples. We present a new theory about how hidden regions emerge, including the number, boundaries, and connections. We designed, implemented, and tested this new network and used the phoneme recognition experiment as an example of audition modality. Based on the experimental results, we analyzed both the advantages of the new mechanisms in DN-2 and how these new mechanisms help DN-2 to automatically generate a hierarchical architecture. We believe DN-2 is in the right direction of dealing with various sequential data. This work is the first step for our goal of auditory-language autonomous learning. Xiang Wu 0008, Zejia Zheng, Juyang Weng |
IJCNN | 3 |
| 2018 | Emergent Turing Machine as a General Purpose ApproximatorabstractAn autonomous navigation agent needs to learn the navigation rules and the transition among different navigation states, which can be summarized in Turing Machine (TM) controlled by a navigation Finite Automata. However, real-world inputs to this TM is always noisy with numerous appearance variances and distractors, which made it impossible to handcraft symbolic representations or features to detect. Developmental Network-l (DN-1) incrementally learns a Finite Automata error-free using emergent representation but lacks hierarchy in its internal representation. In this paper, we extend DN-1 by adding lateral connection and multiple types of neurons to form hierarchical representation both internally and in the motor area. The formed hierarchical representation is more robust against distractors compared to the global feature patches in DN-1. Compared to other FA based methods, DN-2 directly uses natural input without handcrafting. The navigation states are emergent from the firing patterns in network areas instead of the hand-crafted symbolic states in the neural network output layer. Real-world and simulated navigation experiments showed that DN-2 successfully approximated rules of navigation using natural inputs. The agent learns different levels of concepts from simple to complex with the help of internal hierarchical representation. Zejia Zheng, Xiang Wu 0008, Juyang Weng |
IJCNN | 3 |
| 2018 | Information-dense actions as contexts
Xiang Wu 0008, Yuming Bo, Juyang Weng |
Neurocomputing | 3 |
| 2018 | Motivated Optimal Developmental Learning for Sequential Tasks Without Using Rigid Time-DiscountsabstractMany methods for reinforcement learning use symbolic representations-nonemergent-such as Q-learning. We use emergent representations here, without human handcrafted symbolic states (i.e., each state corresponds to a different location). This paper models reinforcement learning for hidden neurons in emergent networks for sequential tasks. In this paper, their influences on sequential tasks (e.g., robot navigation in different scenarios) are investigated where the learned value and results of a behavior rely on not only the current experience just like in a pattern recognition (episodic) but also the prediction of future experiences (e.g., delayed rewards) and environments (e.g., previously learned navigational trajectories). We show that this new model of motivated learning amounts to the computation of the maximum-likelihood estimate through "life" where punishment and reward have increased weights. This new formulation avoids the greediness of time-discount in Q-learning. Its complex nonlinear sequential optimization has been solved in a closed-form procedure under the condition of the limited computational resources and limited learning experience so far, because we convert it into a simpler problem of incremental and linear estimation. The experimental results showed that the serotonin and dopamine systems speed up learning for sequential tasks, because not all events are equally important. As far as we know, this is the first work that studies the influences of reinforcers (via serotonin and dopamine) on hidden neurons (Y neurons) for sequential tasks in dynamic scenarios using emergent representations. Dongshu Wang, Yihai Duan, Juyang Weng |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2017 | Actions as contextsabstractIn artificial intelligence, many tasks of speech recognition, video analysis, and language processing involve temporal processing where the outputs depend on not only spatial contents of the current sensory input frame, but also the relevant context in the attended past. It is illusive how brains use temporal contexts. Many computer methods, such as Hidden Markov chains and recurrent neural networks, require the human programmer to handcraft contexts as symbolic contexts. It has been proved that our Developmental Networks (DN) are capable of learning any emergent Turing Machine (TM), their states have been supervised by human teachers as patterns. This demands much effort from the human trainer. In this paper, we study how agent actions are natural sources of contexts. In humans, muscle actions correspond to the firings of muscle neurons. They are dense in time and correlated with the cognitive skills of the individual. Some actions are meant to handle time warping, while others are not (e.g., for time duration counting). We model actions as dense action patterns. We experimented with DN for recognition of audio sequences as an example of modality, but the principles are modality independent. Our experimental results showed how taking dense, frame-wise actions as contexts helps DN to generate temporal contexts. This work is a necessary step toward our goal to enable machines to autonomously generate contexts as actions through life-long development. Xiang Wu 0008, Juyang Weng |
IJCNN | 2 |
| 2016 | Brains as optimal emergent Turing MachinesabstractThe theory and experiments outlined in Weng 2015 [1] modeled brains as naturally emerging Turing Machines (TMs) inside Developmental Networks (DNs) - a new class of brain inspired neural networks. However, TMs originally proposed by Alan Turing 1936 [2] were deterministic. If they involved probability to handle uncertainty, the probability was in the mind of the human programmer for a specific task. Although the task-specific program will run on a Universal TM - a model of the popular Von Neumann Computers - the Universal TM itself had no imbedded probability. Yet, each brain samples from infinitely many stimuli and actions through the real word, unlike a TM that deals with only tape-symbols from a finite and static alphabet. Computationally, is a brain optimal? If the answer is yes, in what sense? This theoretical work reports the overall brain optimality. It shows that a biological brain, modeled by the DN, is optimal in the sense of maximum likelihood, conditioned on (1) the genome program - the Developmental Program (DP) that regulates the development of body, sensors, effectors and the limited computational resource (e.g., brain size and types of neural transmitters), and (2) the incremental lifetime experience - including teaching from the environments and self exploration and discovery through the real world. Juyang Weng |
IJCNN | 1 |
| 2016 | Challenges in visual parking and how a developmental network approaches the problemabstractMany existing parking assistance systems use LIDAR and range sensors. However, such sensors may be fooled by dark surfaces, wet surfaces, and mirrors. The range information does not tell restricted parking spaces such as handicap parking spaces. Vision based autonomous navigation methods have a potential to use richer and more redundant intensity and color information but also require sophisticated processing and extensive computation. In this paper we present our vision based autonomous parking agent for large parking ramps using an object recognition neural network. After action-supervised learning, the network is able to find empty spaces in indoor parking ramps, park into space, and leave the parking space with no human guidance/supervision. An Finite Automaton as context-dependent rules emerge from the Developmental Network. The transitions between different parking states are facilitated by the recognized landmarks learned through on-line training process. Our experiments showed that the agent learned to maneuver in novel environments at a high action accuracy. Our future work includes training and testing the agent in real-time. Zejia Zheng, Juyang Weng |
IJCNN | 2 |
| 2015 | Brains as naturally emerging turing machinesabstractIt has been shown that a Developmental Network (DN) can learn any Finite Automaton (FA) [29] but FA is not a general purpose automaton by itself. This theoretical paper presents that the controller of any Turing Machine (TM) is equivalent to an FA. It further models a motivation-free brain - excluding motivation e.g., emotions - as a TM inside a grounded DN - DN with the real world. Unlike a traditional TM, the TM-in-DN uses natural encoding of input and output and uses emergent internal representations. In Artificial Intelligence (AI) there are two major schools, symbolism and connectionism. The theoretical result here implies that the connectionist school is at least as powerful as the symbolic school also in terms of the general-purpose nature of TM. Furthermore, any TM simulated by the DN is grounded and uses natural encoding so that the DN autonomously learns any TM directly from natural world without a need for a human to encode its input and output. This opens the door for the DN to fully autonomously learn any TM, from a human teacher, reading a book, or real world events. The motivated version of DN [31] further enables a DN to go beyond action-supervised learning - so as to learn based on pain-avoidance, pleasure seeking, and novelty seeking [31]. Juyang Weng |
IJCNN | 1 |
| 2015 | Approaching real-world navigation using object recognition networkabstractTypical navigation systems do not use object recognition as part of their autonomous driving systems. People often use hand-crafted features (e.g. lanes, traffic lights, intersections) based on programmers' knowledge about the environment. Those agents are usually brittle during real-world tests. However, landmarks, as a type of object, need to be recognized for an autonomous navigation system to generalize its learned training data to other unfamiliar environments. In this work we utilize the Developmental Network (DN), which has been tested extensively with object recognition tasks, for a mobile agent and train it to self-navigate in controlled indoor environment. The proposed system uses Lobe Component Analysis (LCA) to learn features from both stereo cameras and desired navigation actions. Neurons attend to different areas of the input image. They compete for firing according to the goodness of matching result. This enables the agent to attend to image local patches as landmarks without explicitly defining objects. Our analysis shows that attention can be corrected by direct supervision and by indirect reinforcement provided by the teacher. We anticipate our work in this paper to be a starting point of research efforts that shift expensive range-scanner-based methods to inexpensive camera-based methods that, although using richer information, face challenges of object appearance variations. Zejia Zheng, Juyang Weng |
IJCNN | 2 |
| 2014 | WWN-9: Cross-domain synaptic maintenance and its application to object groups recognitionabstractWhere What Network 6 (WWN-6) has shown that its model of synaptic maintenance using neural transmitters acetylcholine (ACh) and norepinephrine (NE) enables each neuron to distinguish between neuronal input lines from its relatively stable object patch and those from irrelevant backgrounds. However, it is about only a single domain - sensory domain X. During development from conception through fetus and newborn, every brain neuron has three major domains of input, sensory X, lateral Y and motor Z. The single-domain model of WWN-6 is not directly applicable to multiple domains because different domains have very different dimension and signal variations that cannot be directly compared. We believe that cross-domain synaptic maintenance is a crucial mechanism to develop a shallow-and-deep processing hierarchy in the brain where each neuron autonomously select domains in the developing hierarchy, not necessarily directly connected to receptors in X and muscles in Z. In the new work here, we propose a biologically inspired model for cross-domain synaptic maintenance. We assume that the earlier connection guided by morphogen result in initial coarse connection, but cross-domain synaptic maintenance refine connections to enable each neuron to autonomously find its role. As concept patterns emerge in Z, neurons refine their connections, to differentiate their roles among sensory processing, motor processing, and a mixture of both. Experimentally, we show the effect of the new theory through learning of individual objects and object groups, where neurons initialized for object-group connections tend to find their receptor inputs from X are not as stable as inputs from motor Z, thus, gradually turn into "later" processing neurons - for "higher-level" object-based features and their invariances. In principle, WWN-9 tends to learn a new object group without repeating the learning of all instances of each individual object. Qian Guo 0004, Juyang Weng |
IJCNN | 3 |
| 2014 | Serotonin and dopamine systems: Internal areas and sequential tasksabstractSerotonin and dopamine transmitters are synthesized in the lower brain but are transmitted widely to many areas of the brain. Emergent representations are critical in understanding their effects. In our prior work [26], their effects on internal, non-motor neurons are studied for pattern recognition tasks only. In this paper, we study their effects on sequential tasks - robot navigation with different settings. They are sequential tasks because the outcome of behavior depends on not only the current behavior as in pattern recognition but also the previous behaviors (e.g., previous navigational trajectories). Analytically, we show that the serotonin and dopamine systems affect the performance of sequential tasks in a compounded way. Experimentally, we show that the effect on the learning rate of feature neurons (in the Y area) allows the agent to approach the friend and avoid the enemy faster as compounding effects of sequential states. Further, we tested the effect of punishment and reward schedule with the same initial locations. We also experimented the effect of punishment and reward schedule with random initial locations. These experiments all indicated that the reinforcement learning via the serotonin and the dopamine systems is beneficial for developing desirable behaviors in this set of sequential tasks - staying close to its friend and away from its enemy. As far as we know, this is the first work that investigates the effects of reinforcer (via serotonin and dopamine) on internal neurons for sequential tasks. Dongshu Wang, Yihai Duan, Juyang Weng |
IJCNN | 3 |
| 2014 | A bridge-islands model for brains: Developing numeric circuits for logic and motivationabstractNeuroscience has made impressive advances, but there is a lack of an overall computational brain theory. I would like to present a simplified computational theory in an intuitive language about how the brain wires itself as a multi-interchange bridge that bi-directionally connects many islands where each island is a sensor or effector. The wiring process of the brain is highly self-supervised while a baby lives and acts in his physical environment, e.g., sucking a milk bottle. I use a new precise framework of emergent finite automata to explain how the brain develops its numeric circuits that can be clearly understood in terms of logic. I also explain how the self-wired basic circuits become motivated through four additional neural transmitters beyond glutamate and GABA - serotonin, dopamine, acetylcholine, and norepinephrine. I use finite automata for precise and rigorous analysis and optimality. Juyang Weng |
IJCNN | 1 |
| 2014 | WWN: Integration with coarse-to-fine, supervised and reinforcement learningabstractThe cost of autonomous development is substantial. Although supervised learning is effective, the cost demand on teachers is often too high to be constantly applied. Reinforcement learning can take advantage of physical reality due to environmental feedback and inspections. Information required in reinforcement learning is not as specific as is required in supervised learning. Integration theories, methods, and analysis of these two learning strategies are still rare in the literature although such integration has been well known in the animal kingdom. Based on our prior work on a general purpose framework called Developmental Network and its embodiment Where-What-Network, we present our theory, method, and analysis for integration of supervised learning and reinforcement learning in this paper. Different from all other known work on reinforcement learning, this DN framework uses fully emergent representation to avoid the brittleness and task-specific representations. Central in the integration is not just to provide a freedom for the teacher to choose the mode of learning, which is necessary especially when the physical non-living world is an implicit teacher, but the mechanism of scaffolding. In our experiment the scaffolding is reflected by allowing the location motor(LM) neurons to gradually refine representation through splitting(mitosis) in a coarse to fine scheme. We report our experimental work in a very challenging learning setting: both object and backgrounds are unknown(cluttered settings) and concepts(e.g. location and type) emerge from agent-environment interactions, instead of rigidly handcrafted. Zejia Zheng, Juyang Weng, Zhengyou Zhang |
IJCNN | 2 |
| 2013 | Novelty estimation in developmental networks: Acetylcholine and norepinephrineabstractThe receiver operating characteristic (ROC) curve has been widely applied to classifiers to show how the threshold value for acceptance changes the true positive rate and the false positive rate of the detection jointly. However, it is largely unknown how a biological brain autonomously selects a confidence value for each detection case. In the reported work, we investigated this issue based on the class of Developmental Networks (DNs) which have a power of abstraction similar to symbolic finite automata (FA) but all the DN's representations are emergent (i.e., numeric from the physical world and non-symbolic). Our theory is based on two types of neurotransmitters: Acetylcholine (Ach) and Norepinephrine (NE). Inspired by studies that proposed Ach and NE represent uncertainty and unpredicted uncertainty, respectively, we model how a DN uses Ach and NE to allow neurons to collectively decide acceptance or rejection by estimated novelty based on past experience, instead of using a single threshold value. This is a neural network, distributed, incremental, automatic version of ROC. Jordan Fish, Lisa Ossian, Juyang Weng |
IJCNN | 3 |
| 2013 | Stereo where-what networks: Unsupervised binocular feature learningabstractUnsupervised feature learning has been shown promising in the field of machine learning. However, the learning algorithms used in these methods, e.g., Deep Belief Networks based on Restricted Boltzman Machines, are typically restricted in different ways, e.g., hard to train and calculate the partition function [1]. In this article, we present a cortex-inspired learning network, Where-What Networks (WWN), for the problem of unsupervised learning of binocular local features. The results show that the learned features autonomously developed selectivity for disparity, profile and location of the input patterns. We present a novel algorithm, Dynamic Synapse Lobe Component Analysis (DSLCA), which not only resembles the pattern of neural connections in the visual cortex, but also results in the autonomous development of “domain disparity”. To our knowledge, this work is the first to introduce unsupervised learning of both domain and weight disparities between left and right local receptive fields. Moreover, given the theoretical optimality of WWNs [2] and their empirically proven strength in supervised learning, the presented work is the first step towards creating a semi-supervised learning network for simultaneous type, location (including distance) and 3D shape perception. Mojtaba Solgi, Juyang Weng |
IJCNN | 2 |
| 2013 | Modeling the effects of neuromodulation on internal brain areas: Serotonin and dopamineabstractThe effects of neuromodulator, such as serotonin and dopamine, on individual neurons in the brain have been known qualitatively. However, it is challenging to computationally model such effects in an emergent network, as the elements of internal representations do not have a static, task-specific meaning. Weng and coworkers modeled the effects of serotonin and dopamine on only motor neurons in emergent networks. In this work, we extend the effects of serotonin and dopamine to all neurons inside the emergent network. Our new theory is that although serotonin and dopamine indicate events of different natures (aversive and appetitive), they produce similar effects on internal non-motor neurons in that they increase their learning rates from the cases without serotonin and dopamine. This is because the presence of serotonin and dopamine indicates a higher importance of the event compared with baseline cases. Experimentally, we show that the enhanced developmental network learns faster under a limited resource. Zejia Zheng, Kui Qian, Juyang Weng, Zhengyou Zhang |
IJCNN | 3 |
| 2012 | Skull-closed autonomous development: WWN-6 using natural videoabstractWhile a physical environment interacts with a human individual through the brain's sensors and effectors, internal representations inside the skull-closed brain autonomously emerge and adapt throughout the lifetime. By “skull-closed”, we mean that the brain inside the skull is off limit to all teachers in the external physical environment, except the brain's sensory ends and motor ends. We present the Where-What Network 6 (WWN-6), which has realized our goal of skull-closed autonomous development. This means that the human programmer is not allowed to handcraft the internal representations for any fixed extra-body concepts. For example, the meanings of specific values of the location or type concept are not known during the programming time. Such meanings emerge through associations imbedded in “postnatal” experience. This capability is especially challenging when one considers the fact that most elements in the sensory ends are irrelevant to the signals at the effector ends (e.g., many background pixels). How does each output (usually expressed in vector) in the effectors find its corresponding existing pattern in the correct patch of the sensory image? We outline this autonomous learning theory for the brain and present how the developmental program (DP) of WWN-6 enables the network to perform for attending and recognizing objects in complex backgrounds using natural video. The inputs to the agent (i.e., the network) are not artificially synthesized images as WWNs used before, but drawn from continuous video taken from natural settings where, in general, everything is moving. Yuekai Wang, Juyang Weng |
IJCNN | 3 |
| 2012 | Skull-Closed Autonomous Development: Object-Wise Incremental Learning
Yuekai Wang, Juyang Weng |
ISNN (1) | 3 |
| 2011 | Skull-Closed Autonomous Development
Yuekai Wang, Juyang Weng |
ICONIP (1) | 3 |
| 2011 | Neuromorphic motivated systemsabstractAlthough reinforcement learning has been extensively modeled, few agent models that incorporate values use biologically plausible neural networks as a uniform computational architecture. We call biologically plausible neural network architecture neuromorphic. This paper discusses some theoretical constraints on neuromorphic intrinsic value systems [3]. By intrinsic, we mean a value system that is likely programmed by the genes, whose value bias has already taken a shape at the birth time. Such an intrinsic value system plays an important role in developing extrinsic values through the agent's own experience during its life span. Based on our theoretical constraints, we model two types of neurotransmitters, serotonin and dopamine, to construct a neuromorpic intrinsic value system based on a uniform neural network architecture. Serotonin represents punishment and stress, while dopamine represents reward and pleasure. Experimentally, this model allows our simulated robot to develop an attachment to one entity and fear another. James C. Daly, Jacob D. Brown, Juyang Weng |
IJCNN | 3 |
| 2011 | Modeling dopamine and serotonin systems in a visual recognition networkabstractMany studies have been performed to train a classification network using supervised learning. In order to enable a recognition network to learn autonomously or to later improve its recognition performance through simpler confirmation or rejection, it is desirable to model networks that have an intrinsic motivation system. Although reinforcement learning has been extensively studied, much of the existing models are symbolic whose internal nodes have preset meanings from a set of handpicked symbolic set that is specific for a given task or domain. Neural networks have been used to automatically generate internal (distributed) representations. However, modeling a neuromorphic motivational system for neural networks is still a great challenge. By neuromorphic, we mean that the motivational system for a neural network must be also a neural network, using a standard type of neuronal computation and neuronal learning. This work proposes a neuromorphic motivational system, which includes two subsystems - the serotonin system and the dopamine system. The former signals a large class of stimuli that are intrinsically aversive (e.g., stress or pain). The latter signals a large class of stimuli that are intrinsically appetitive (e.g., sweet and pleasure). We experimented with this motivational system for visual recognition settings to investigate how such a system can learn through interactions with a teacher, who does not give answers, but only punishments and rewards. Stephen Paslaski, Courtland VanDam, Juyang Weng |
IJCNN | 3 |
| 2011 | Where-What Network 5: Dealing with scales for objects in complex backgroundsabstractThe biologically-inspired developmental Where-What Networks (WWN) are general purpose visuomotor networks for detecting and recognizing objects from complex backgrounds, modeling the dorsal and ventral streams of the biological visual cortex. The networks are designed for the attention and recognition problem. The architecture in previous versions were meant for a single scale of foreground. This paper focuses on Where-What Network-5 (WWN-5), the extension for multiple scales. WWN-5 can learn three concepts of an object: type, location and scale. Juyang Weng |
IJCNN | 3 |
| 2011 | Synapse maintenance in the Where-What NetworksabstractGeneral object recognition in complex backgrounds is still challenging. On one hand, the various backgrounds, where object may appear at different locations, make it difficult to find the object of interest. On the other hand, with the numbers of locations, types and variations in each type (e.g., rotation) increasing, conventional model-based approaches start to break down. The Where-What Networks (WWNs) were a biologically inspired framework for recognizing learned objects (appearances) from complex backgrounds. However, they do not have an adaptive receptive field for an object of a curved contour. Leaked-in background pixels will cause problems when different objects look similar. This work introduces a new biologically inspired mechanism - synapse maintenance and uses both supervised (motor-supervised for class response) and unsupervised learning (synapse maintenance) to realize objects recognition. Synapse maintenance is meant to automatically decide which synapse should be active firing of the post-synaptic neuron. With the synapse maintenance, the network has achieved a significant improvement in the network performance. Yuekai Wang, Juyang Weng |
IJCNN | 3 |
| 2011 | Three theorems: Brain-like networks logically reason and optimally generalizeabstractFinite Automata (FA) is a base net for many sophisticated probability-based systems of artificial intelligence. However, an FA processes symbols, instead of images that the brain senses and produces (e.g., sensory images and motor images). Of course, many recurrent artificial neural networks process images. However, their non-calibrated internal states prevent generalization, let alone the feasibility of immediate and error-free learning. I wish to report a general-purpose Developmental Program (DP) for a new type of, brain-anatomy inspired, networks - Developmental Networks (DNs). The new theoretical results here are summarized by three theorems. (1) From any complex FA that demonstrates human knowledge through its sequence of the symbolic inputs-outputs, the DP incrementally develops a corresponding DN through the image codes of the symbolic inputs-outputs of the FA. The DN learning from the FA is incremental, immediate and error-free. (2) After learning the FA, if the DN freezes its learning but runs, it generalizes optimally for infinitely many image inputs and actions based on the embedded inner-product distance, state equivalence, and the principle of maximum likelihood. (3) After learning the FA, if the DN continues to learn and run, it “thinks” optimally in the sense of maximum likelihood based on its past experience. Juyang Weng |
IJCNN | 1 |
| 2011 | Where-What Network with CUDA: General Object Recognition and Location in Complex Backgrounds
Yuekai Wang, Juyang Weng |
ISNN (2) | 5 |
| 2011 | Incremental Online Object Learning in a Vehicular Radar-Vision Fusion FrameworkabstractIn this paper, we propose an object learning system that incorporates sensory information from an automotive radar system and a video camera. The radar system provides coarse attention for the focus of visual analysis on relatively small areas within the image plane. The attended visual areas are coded and learned by a three-layer neural network utilizing what is called in-place learning: Each neuron is responsible for the learning of its own processing characteristics within the connected network environment, through inhibitory and excitatory connections with other neurons. The modeled bottom-up, lateral, and top-down connections in the network enable sensory sparse coding, unsupervised learning, and supervised learning to occur concurrently. This paper is applied to learn two types of encountered objects in multiple outdoor driving settings. Cross-validation results show that the overall recognition accuracy is above 95% for the radar-attended window images. In comparison with the uncoded representation and purely unsupervised learning (without top-down connection), the proposed network improves the overall recognition rate by 15.93% and 6.35%, respectively. The proposed system is also compared favorably with other learning algorithms. The result indicates that our learning system is the only one that is fit for incremental and online object learning in a real-time driving environment. Zhengping Ji, Matthew D. Luciw, Juyang Weng, Shuqing Zeng |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2010 | WWN-2: A biologically inspired neural network for concurrent visual attention and recognitionabstractAttention and recognition have been addressed separately as two challenging computational vision problems, but an engineering-grade solution to their integration and interaction is still open. Inspired by the brain's dorsal and ventral pathways in cortical visual processing, we present a neuromorphic architecture, called Where-What Network 2 (WWN-2), to integrate object attention and recognition interactively through their experience-based development. This architecture enables three types of attention: feature-based bottom-up attention, position-based top-down attention, and object-based top-down attention, as three possible information flows through the Y-shaped network. The learning mechanism of the network is rooted in a simple but efficient cell-centered synaptic update model, entailing the dual optimization of Hebbian directions and cell firing-age dependent step sizes. The inputs to the network are a sequence of images, where specific foreground objects may appear anywhere within an unknown, complex, natural background. The WWN-2 regulates the network to dynamically establish and consolidate position-specified and type-specified representations through a supervised learning mode. The network has reached 92.5% object recognition rate and an average of 1.5 pixels in position error after 20 epochs of training. Zhengping Ji, Juyang Weng |
IJCNN | 2 |
| 2010 | Where-What Network 3: Developmental top-down attention for multiple foregrounds and complex backgroundsabstractThe Where-What Network 3 (WWN-3) is an artificial developmental network modeled after visual cortical pathways, for the purpose of attention and recognition in the presence of complex natural backgrounds. It is general-purpose and not pre-determined to detect a certain type of stimulus. It is a learning network, which develops its weights from images using a supervised paradigm and a local Hebbian learning algorithm. Attention has been thought of as bottom-up or top-down. This paper focuses on the biologically-inspired mechanisms of top-down attention in WWN-3, through top-down excitation that interacts with bottom-up activity at every layer within the network. Top-down excitation in WWN-3 can control the location of attention by imposing a certain location or disengaging from the current location. It can also control what type of object to search for. Paired layers and sparse coding deal with potential hallucination problems. Top-down attention in WWN occurs as soon as an action emerges at a motor layer, which could be imposed by a teacher or internally selected. Given two competing foregrounds in the same scene, WWN showed effective performance in all the attention modes tested. Matthew D. Luciw, Juyang Weng |
IJCNN | 2 |
| 2010 | A 5-chunk developmental brain-mind network model for multiple events in complex backgroundsabstractThere has been no prior general purpose brain-mind model for multiple events in complex backgrounds. I first discuss that although the age of brain-mind seems to have arrived, the current infrastructure does not well fit the need of research and peer-review for the challenging and important subject of brain-mind. Then, I present a general purpose model of the brain-mind, called the Epigenetic Developer (ED) network model. The model proposes five necessary “chunks” for the brain picture: development, architecture, area, space and time. The development chunk means that any practical brain, natural or artificial, needs to autonomously develop through interactions with the natural environments without any previously given set of tasks. The architecture chunk handles (1) multiple objects in complex backgrounds; (2) reasoning under abstract contexts; (3) multiple sensory modalities and multiple motor modalities and their integration. The area chunk addresses the issue of feature development and area representation, without rigidly specifying what each neuron does. The space chunk deals with spatial attention to individual objects in complex backgrounds, to satisfy the invariance and specificity criteria for type, location, and other concepts. The time chunk indicates that the brain uses its intrinsic spatial mechanisms to deals with time, without dedicated temporal components. The model copes with temporal contexts of events, to satisfy invariance and specificity criteria for time warping, time duration, temporal attention, and long temporal length. The theory and mechanisms are presented and some related experimental results are summarized. Juyang Weng |
IJCNN | 1 |
| 2009 | Temporal context as cortical spatial codesabstractIt is largely unknown how the brain deals with time. The new field of research on autonomous development must enable machines to develop intelligent behaviors that respond not only to spatial features, but also temporal features. Hidden Markov Model (HMM) has a probability based mechanism to deal with time warping, but no effective online method exists that can deal with general temporal structure and temporal abstraction. By online, we mean that the agent must respond to spatial and temporal context immediately while the sensory stream flows in. By general temporal context, we mean various desirable temporal subsets, such as deletion (e.g., stop words) and variable temporal lengths (e.g., beyond bigrams and trigrams). By temporal abstraction, we mean using abstract meaning of context, instead of concrete forms. This paper proposes a brain inspired online scheme for making sequential decisions based on general temporal context. By sequential decisions, the action from the network depends on not only inputs and outputs but also emergent internal context states. In our neuromorphic scheme, the internal states are not predefined symbols, but distributed context depending on the internal attention. Our complexity analysis shows how this scheme greatly reduces the exponential time complexity O(2t) of all the possible number of contexts of length t down to linear time complexity O(cnt), where n is the number of neurons in the network and c is the average number of synapses of each neuron. In this paper, we concentrate on processing sequential text inputs by an online agent network under motor-supervised learning. Juyang Weng, Mingmin Chi, Xiangyang Xue 0001 |
IJCNN | 1 |
| 2009 | Spatio-Temporal Adaptation in the Unsupervised Development of Networked Visual NeuronsabstractThere have been many computational models mimicking the visual cortex that are based on spatial adaptations of unsupervised neural networks. In this paper, we present a new model called neuronal cluster which includes spatial as well as temporal weights in its unified adaptation scheme. The "in-place" nature of the model is based on two biologically plausible learning rules, Hebbian rule and lateral inhibition. We present the mathematical demonstration that the temporal weights are derived from the delay in lateral inhibition. By training with the natural videos, this model can develop spatio-temporal features such as orientation selective cells, motion sensitive cells, and spatio-temporal complex cells. The unified nature of the adaptation scheme allows us to construct a multilayered and task-independent attention selection network which uses the same learning rule for edge, motion, and color detection, and we can use this network to engage in attention selection in both static and dynamic scenes. Dongyue Chen 0001, Liming Zhang 0001, Juyang Weng |
IEEE Trans. Neural Networks | 3 |
| 2008 | Epigenetic sensorimotor pathways and its application to developmental object learningabstractA pathway in the central nervous system (CNS) is a path through which nervous signals are processed in an orderly fashion. A sensorimotor pathway starts from a sensory input and ends at a motor output, although almost all pathways are not simply unidirectional. In this paper, we introduce a simple, biologically inspired, unified computational model - Multi-layer In-place Learning Network (MILN), with a design goal to develop a recurrent network, as a function of sensorimotor signals, for open-ended learning of multiple sensorimotor tasks. The biologically motivated MILN provides automatic feature derivation and pathway refinement from the temporally real-time inputs. The work presented here is applied in the challenging application field of developing reactive behaviors from a video camera and a (noisy) radar range sensor for a vehicle-based robot in open, natural driving environments. An internal model of the agent’s experience of the environments is created and refined from the ground-up using a cell-centered model, based on the genomic equivalence principle. The outputs can be imposed by a teacher, at the same time as the learning is active. At any time instant, sensory information from the radar allows the system to focus its visual analysis on relatively small areas within the image plane (attention selection), in a computationally efficient way, suitable for real-time training. This system was trained with data from 10 different city and highway road environments, and cross validation shows that MILN was able to correctly recognize above 95% of the radar-extracted images from the multiple environments. The in-place learning mechanism compares with other learning algorithms favorably, as results of a comparison indicate that in-place learning is the only one to fit all the specified criteria of development of a general-purpose sensorimotor pathway. Zhengping Ji, Matthew D. Luciw, Juyang Weng |
IEEE Congress on Evolutionary Computation | 3 |
| 2008 | Developmental Stereo: Topographic Iconic-Abstract Map from Top-Down Connection
Mojtaba Solgi, Juyang Weng |
ICONIP (1) | 2 |
| 2008 | Learning of sensorimotor behaviors by a SASE agent for vision-based navigationabstractIn this paper, we propose a model to develop robotspsila covert and overt behaviors by using reinforcement and supervised learning jointly. The covert behaviors are handled by a motivational system, which is achieved through reinforcement learning. The overt behaviors are directly selected by imposing supervised signals. Instead of dealing with problems in controlled environments with a low-dimensional state space, our model is applied for the learning in non-stationary environments. Locally balanced incremental hierarchical discriminant regression (LBIHDR) tree is introduce to be the engine of cognitive mapping. Its balanced coarse-to-fine tree structure guarantees real-time retrieval in self-generated high-dimensional state space. Furthermore, K-nearest neighbor strategy is adopted to reduce training time complexity. Vision-based outdoor navigation are used as challenging task examples. In the experiment, the mean square error of heading direction is 0deg for re-substitution test and 1.1269deg for disjoint test, which allows the robot to drive without a big deviation from the correct path we expected. Compared with IHDR (W.S. Hwang and J. Weng, 2007), LBIHDR reduced the mean square error by 0.252deg and 0.5052deg, using re-substitution and disjoint test, respectively. Zhengping Ji, Juyang Weng |
IJCNN | 3 |
| 2008 | Topographic Class Grouping with applications to 3D object recognitionabstractThe cerebral cortex uses a large number of top-down connections, but the roles of the top-down connections remain unclear. Through end-to-end (sensor-to-motor) multilayered networks that use three types of connections (bottom-up, lateral, and top-down), the new Topographic Class Grouping (TCG) mechanism shown in this paper explains how the top-down connections influence (1) the type of feature detectors (neurons) developed and (2) their placement in the neuronal plane. The top-down connections boost the variations in the neuronal between class directions during the training phase. The first outcome of this top-down boosted input space is the facilitation of the emergence of feature detectors that are purer, measured statistically by the average entropy of the neurons’ development. The relatively purer neurons are more “abstract,” i.e., characterizing class-specific (or motor-specific) input information, resulting in better classification rates. The second outcome of this top-down boosted input space is the increase of the distance between input samples that belong to different classes, resulting in a farther separation of neurons according to their class. Therefore, neurons that respond to the same class become relatively nearer. This results in TCG, measured statistically by a smaller within-class scatter of responses when the neuronal plane has a fixed size. Although these mechanisms are potentially applicable to any pattern recognition applications, we report quantitative effects of these mechanisms for 3D object recognition of center-normalized, background-controlled objects. TCG has enabled a significant reduction of the recognition errors. Matthew D. Luciw, Juyang Weng |
IJCNN | 2 |
| 2008 | Multilayer in-place learning networks for modeling functional layers in the laminar cortex
Juyang Weng, Tianyu Luwang, Hong Lu 0001, Xiangyang Xue 0001 |
Neural Networks | 1 |
| 2007 | The Multilayer In-Place Learning Network for the Development of General Invariances and Multi-Task LearningabstractCurrently, there is a lack of general-purpose in-place learning engines that incrementally learn multiple tasks, to develop "soft" multi-task-shared invariances in the intermediate internal representation while a developmental robot interacts with its environment. Computationally, biologically inspired in-place learning provides unusually efficient learning algorithms whose simplicity, low computational complexity, and generality are set apart from typical conventional learning algorithms. We present in this paper the multiple-layer in-place learning network (MILN) for this ambitious goal. As a key requirement for autonomous mental development, the network enables both unsupervised and supervised learning to occur concurrently, depending on whether motor supervision signals are available or not at the motor end (the last layer) during the agent's interactions with the environment. We present principles based on which MILN automatically develops invariant neurons in different layers and why such invariant neuronal clusters are important for learning later tasks in open-ended development. Juyang Weng, Tianyu Luwang, Hong Lu 0001, Xiangyang Xue 0001 |
IJCNN | 1 |
| 2007 | On developmental mental architectures
Juyang Weng |
Neurocomputing | 1 |
| 2007 | Guest Editorial: Convergent Approaches to the Understanding of Autonomous Mental DevelopmentabstractThe eight articles in this special issue focus on convergent approaches to the understanding of autonomous mental development. The goal is to promote the effort to build the necessary bridges between research areas to help foster the investigation of the emergence of intelligent, autonomous, and organized behaviour in children and robotic systems. James L. McClelland, Kim Plunkett, Juyang Weng |
IEEE Trans. Evol. Comput. | 3 |
| 2007 | Task Transfer by a Developmental RobotabstractScaffolding is a process of transferring learned skills to new and more complex tasks through arranged experience in open-ended development. In this paper, we propose a developmental learning architecture that enables a robot to transfer skills acquired in early learning settings to later more complex task settings. We show that a basic mechanism that enables this transfer is sequential priming combined with attention, which is also the driving mechanism for classical conditioning, secondary conditioning, and instrumental conditioning in animal learning. A major challenge of this work is that training and testing must be conducted in the same program operational mode through online, real-time interactions between the agent and the trainers. In contrast with former modeling studies, the proposed architecture does not require the programmer to know the tasks to be learned and the environment is uncontrolled. All possible perceptions and actions, including the actual number of classes, are not available until the programming is finished and the robot starts to learn in the real world. Thus, a predesigned task-specific symbolic representation is not suited for such an open-ended developmental process. Experimental results on a robot are reported in which the trainer shaped the behaviors of the agent interactively, continuously, and incrementally through verbal commands and other sensory signals so that the robot learns new and more complex sensorimotor tasks by transferring sensorimotor skills learned in earlier periods of open-ended development Yilu Zhang, Juyang Weng |
IEEE Trans. Evol. Comput. | 2 |
| 2007 | Incremental Hierarchical Discriminant RegressionabstractThis paper presents incremental hierarchical discriminant regression (IHDR) which incrementally builds a decision tree or regression tree for very high-dimensional regression or decision spaces by an online, real-time learning system. Biologically motivated, it is an approximate computational model for automatic development of associative cortex, with both bottom-up sensory inputs and top-down motor projections. At each internal node of the IHDR tree, information in the output space is used to automatically derive the local subspace spanned by the most discriminating features. Embedded in the tree is a hierarchical probability distribution model used to prune very unlikely cases during the search. The number of parameters in the coarse-to-fine approximation is dynamic and data-driven, enabling the IHDR tree to automatically fit data with unknown distribution shapes (thus, it is difficult to select the number of parameters up front). The IHDR tree dynamically assigns long-term memory to avoid the loss-of-memory problem typical with a global-fitting learning algorithm for neural networks. A major challenge for an incrementally built tree is that the number of samples varies arbitrarily during the construction process. An incrementally updated probability model, called sample-size-dependent negative-log-likelihood (SDNLL) metric is used to deal with large sample-size cases, small sample-size cases, and unbalanced sample-size cases, measured among different internal nodes of the IHDR tree. We report experimental results for four types of data: synthetic data to visualize the behavior of the algorithms, large face image data, continuous video stream from robot navigation, and publicly available data sets that use human defined features. Juyang Weng, Wey-Shiuan Hwang |
IEEE Trans. Neural Networks | 1 |
| 2006 | In-Place Learning for Positional and Scale InvarianceabstractIn-place learning is a biologically inspired concept, meaning that the computational network is responsible for its own learning. With in-place learning, there is no need for a separate learning network. We present in this paper a multiple-layer in-place learning network (MILN) for learning positional and scale invariance. The network enables both unsupervised and supervised learning to occur concurrently. When supervision is available (e.g., from the environment during autonomous development), the network performs supervised learning through its multiple layers. When supervision is not available, the network practices while using its own practice motor signal as self-supervision (i.e., unsupervised per classical definition). We present principles based on which MILN automatically develops positional and scale invariant neurons in different layers. From sequentially sensed video streams, the proposed in-place learning algorithm develops a hierarchy of network representations. The global invariance was achieved through multi-layer quasi-invariances, with increasing invariance from early layers to the later layers. Experimental results are presented to show the effects of the principles. Juyang Weng, Hong Lu 0001, Tianyu Luwang, Xiangyang Xue 0001 |
IJCNN | 1 |
| 2006 | Optimal In-Place Learning and the Lobe Component AnalysisabstractIt is difficult to map many existing learning algorithms onto biological networks because the former require a separate learning network. The computational basis of biological cortical learning is still poorly understood. This paper rigorously introduces a concept called in-place learning. With in-place learning, every networked neuron in-place is responsible for the learning of its signal processing characteristics (e.g., efficacies of synapses) within its connected network environment. There is no need for a separate learning network. With this in-place hypothesis, consequently, each neuron does not have extra space to compute and store the second and higher order statistics (e.g., correlations) of its input fibers. This work first provides a classification of learning algorithms. Then, it shows that the two well-known in-place biological mechanisms, the Hebbian rule and lateral inhibition, are sufficient to develop orientation selective cells, similar to those found in VI, from inputs of natural images. Many other cells that have not been fully understood have emerged as well. The presented computational study of these two in-place learning mechanisms leads to a new concept- these cells correspond to what are called lobe components, which are high concentrations in the probability of the neuronal input space. Further analysis explains how every neuron can learn efficiently (i.e., near-optimal efficiency) by scheduling its plasticity while interacting with other connected neurons. A simple, in-place (Type-5) learning algorithm is presented. The experimental results showed that this simple biologically inspired algorithm is superior to some well-known state-of-the-art ICA algorithms, thanks to its near-optimal efficiency. Juyang Weng, Nan Zhang 0002 |
IJCNN | 1 |
| 2005 | Gradient sparse optimization via competitive learningabstractIn this paper, we propose a new method to achieve sparseness via a competitive learning principle for the linear kernel regression and classification task. We form the duality of the LASSO criteria, and transfer an /spl lscr/ /sub 1/ norm minimization to an /spl lscr//sub /spl infin// norm maximization problem. We introduce a novel solution derived from gradient descending, which links the sparse representation and the competitive learning scheme. This framework is applicable to a variety of problems, such as regression, classification, feature selection, and data clustering. Nan Zhang 0002, Shuqing Zeng, Juyang Weng |
ICASSP (4) | 3 |
| 2005 | Auditory learning: a developmental methodabstractMotivated by the human autonomous development process from infancy to adulthood, we have built a robot that develops its cognitive and behavioral skills through real-time interactions with the environment. We call such a robot a developmental robot. In this paper, we present the theory and the architecture to implement a developmental robot and discuss the related techniques that address an array of challenging technical issues. As an application, experimental results on a real robot, self-organizing, autonomous, incremental learner (SAIL), are presented with emphasis on its audition perception and audition-related action generation. In particular, the SAIL robot conducts the auditory learning from unsegmented and unlabeled speech streams without any prior knowledge about the auditory signals, such as the designated language or the phoneme models. Neither available before learning starts are the actions that the robot is expected to perform. SAIL learns the auditory commands and the desired actions from physical contacts with the environment including the trainers. Yilu Zhang, Juyang Weng, Wey-Shiuan Hwang |
IEEE Trans. Neural Networks | 2 |
| 2004 | Office presence detection using multimodal context informationabstractAn office presence detection system is presented. Context information from multi-sensory inputs is integrated to infer a user's activities in an office. We design a layered architecture to model human activities with different granularities. An IHDR (incremental hierarchical discriminant regression) tree is used to generate models automatically for acoustic signals from unsegmented auditory streams, with a high adaptive capability to new settings. Hidden Markov models (HMM) are implemented to detect human motion patterns. The outputs of the above two components are fed into high-level HMMs to analyze human activities. Experimental results of the real-time prototype system are reported. Juyang Weng, Zhengyou Zhang |
ICASSP (3) | 2 |
| 2004 | A quasi-optimally efficient algorithm for independent component analysisabstractWe propose an incremental algorithm for independent component analysis (ICA), that is guided by the statistical efficiency. Starting from an /spl lscr//sup /spl lscr//spl infin// norm sparseness measure contrast function, we derive the learning algorithm based on a winner-take-all learning mechanism. It avoids the optimization of high order non-linear functions or density estimation, which have been used by other ICA methods, such as negentropy approximation, infomax, and maximum likelihood estimation based methods. We show that when the latent independent random variables are super-Gaussian distributions, the network efficiently extracts the independent components. We observed a much faster convergence than with other ICA methods. Juyang Weng, Nan Zhang 0002 |
ICASSP (5) | 1 |
| 2004 | Obstacle Avoidance through Incremental Learning with Attention SelectionabstractThis work presents a learning-based approach to the task of generating local reactive obstacle avoidance. The learning is performed online in real-time by a mobile robot. The robot operated in an unknown bounded 2-D environment populated by static or moving obstacles (with slow speeds) of arbitrary shape. The sensory perception was based on a laser range finder. To greatly reduce the number of training samples needed, an attentional mechanism was used. An efficient, real-time implementation of the approach had been tested, demonstrating smooth obstacle-avoidance behaviors in a corridor with a crowd of moving students as well as static obstacles. Shuqing Zeng, Juyang Weng |
ICRA | 2 |
| 2004 | Object permanence: results from developmental roboticsabstractObject permanence is an important theoretical construct that has been researched with infants. For the last twenty years, a debate as whether object permanence is an innate conceptual knowledge or a gradually constructed perceptual capability has been raised. However, the lack of autonomous computational models leaves this issue still widely open. A neurologically inspired computational model based on priming is proposed and tested on our developmental SAIL robot. By nurturing the robot baby with different living experiences, we conduct the well-known "drawbridge" experiment on eleven experience-tuned "brains". The implication of our experimental results is informative, which not only sheds light on the controversial issue of object permanence, but might produce long lasting effects on the way AI community approaches highly perceptual machines. Juyang Weng |
IJCNN | 2 |
| 2004 | Value system development for a robotabstractWe present a refined model for online development of value system. As an indispensable part of a developmental robot, the value system signals the occurrence of salient sensory inputs, modulates the mapping from sensory inputs to action outputs, and evaluates candidate actions. No salient feature is predefined in the value system but instead novelty based on experience, which is applicable to any task. Furthermore, reinforcer is integrated with novelty. Thus, the value system of a robot can be developed through interactions with trainers. In the experiment, we treat vision-based neck action selection as a behavior guided by the value system. The robot's behavior is consistent with the attention mechanism in human infants. Juyang Weng |
IJCNN | 2 |
| 2004 | Sparse representation from a winner-take-all neural networkabstractWe introduce an incremental algorithm for independent component analysis (ICA) based on maximization of sparseness criteria. We propose using a new sparseness measure criteria function. The learning algorithm based on this criteria leads to a winner-take-all learning mechanism. It avoids the optimization of high order nonlinear function or density estimation, which have been used by other ICA methods. We show that when the latent independent random variables are super-Gaussian distributions, the network efficiently extracts the independent components. Nan Zhang 0002, Juyang Weng |
IJCNN | 2 |
| 2003 | Locally Balanced Incremental Hierarchical Discriminant Regression
Juyang Weng, Roger Calantone |
IDEAL | 2 |
| 2003 | A Fast Algorithm for Incremental Principal Component Analysis
Juyang Weng, Yilu Zhang, Wey-Shiuan Hwang |
IDEAL | 1 |
| 2003 | Autonomous mental development in high dimensional state and action spacesabstractAutonomous mental development (AMD) of robots opened a new paradigm for developing machine intelligence, using neural network type of techniques and it fundamentally changed the way an intelligent machine is developed from manual to autonomous. The work presented is a part of SAIL (self-organizing autonomous incremental learner) project which deals with autonomous development of entire humanoid robot with vision, audition, manipulation and locomotion. The major issue addressed is the challenge of high dimensional action space (5 to 10) in addition to the high dimensional context state space (hundreds to thousands and beyond), typically required by an AMD machine. This is the first work that studies a high dimensional (numeric) action space in conjunction with a high dimensional perception (context state) space, under the AMID mode. Two new learning algorithms, Direct Update on Direction Cosines (DUDC) and High-Dimensional Conjugate Gradient Search (HCGS), are developed, implemented and tested. The convergence properties of both the algorithms and their targeted applications are discussed. Autonomous learning of speech production under reinforcement learning is studied as an example. Ameet Joshi, Juyang Weng |
IJCNN | 2 |
| 2003 | Developing early senses about the world: "Object Permanence" and visuoauditory real-time learningabstractWhat "constraints" are exactly wired into the human developmental program? What "constraints" are minimally necessary for a developmental robot? These are open questions. In this paper, we propose a mechanism of developing experience-based priming - predicting the future contexts including sensation and action based on the previous experience - as a powerful "constraint" for developmental robots. We present an architecture that develops this priming capability through realtime online interactions with the environment. We report how our SAIL robot developed a sense of novelty in a well-known "drawbridge" experiment which sheds light on the controversial issue of "object permanence" in psychology. We further show how the proposed priming mechanism enabled SAIL to deal with a very challenging online learning setting: learning the name and property (e.g., size) of dynamically rotating objects through verbal dialogues. Juyang Weng, Yilu Zhang |
IJCNN | 1 |
| 2003 | Online image classification using IHDR
Juyang Weng, Wey-Shiuan Hwang |
Int. J. Document Anal. Recognit. | 1 |
| 2003 | Autonomous mental development in high dimensional context and action spaces
Ameet Joshi, Juyang Weng |
Neural Networks | 2 |
| 2003 | Candid Covariance-Free Incremental Principal Component AnalysisabstractAppearance-based image analysis techniques require fast computation of principal components of high-dimensional image vectors. We introduce a fast incremental principal component analysis (IPCA) algorithm, called candid covariance-free IPCA (CCIPCA), used to compute the principal components of a sequence of samples incrementally without estimating the covariance matrix (so covariance-free). The new method is motivated by the concept of statistical efficiency (the estimate has the smallest variance given the observed data). To do this, it keeps the scale of observations and computes the mean of observations incrementally, which is an efficient estimate for some well known distributions (e.g., Gaussian), although the highest possible efficiency is not guaranteed in our case because of unknown sample distribution. The method is for real-time applications and, thus, it does not allow iterations. It converges very fast for high-dimensional image vectors. Some links between IPCA and the development of the cerebral cortex are also discussed. Juyang Weng, Yilu Zhang, Wey-Shiuan Hwang |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2002 | Visual motion based behavior learning using hierarchical discriminant regression
Changjiang Yang, Juyang Weng |
Pattern Recognit. Lett. | 2 |
| 2001 | Incremental Hierarchical Discriminant Regression for Online Image ClassificationabstractThis paper presents an incremental algorithm for image classification problems. Virtual labels are automatically formed by clustering in the output space. These virtual labels are used for the process of deriving discriminating features in the input space. This procedure is performed recursively in a coarse-to-fine fashion resulting in a tree, called incremental hierarchical discriminating regression (IHDR) method. Embedded in the tree is a hierarchical probability distribution model used to prune unlikely cases. A sample size dependent negative-log-likelihood (NLL) metric is introduced to deal with large-sample size cases, small-sample size cases, and unbalanced-sample size cases, measured among different internal nodes of the IHDR algorithm. We report the experimental results of the proposed algorithm for an OCR classification problem and an image orientation classification problems. Juyang Weng, Wey-Shiuan Hwang |
ICDAR | 1 |
| 2001 | Incremental hierarchical discriminating regression for indoor visual navigationabstractIn this paper, we investigate vision-based navigation using the incremental hierarchical discriminating regression (IHDR) algorithm. Based on the learned experience, the system gradually improves its performance through online interaction with environment. The learning process is interactive and on-line. The hierarchical structure of the IHDR algorithm allows each associative recall to be completed in O(log n) time, where it is the number of cases learned. This makes real time performance possible. The IHDR learns incrementally: each learning sample is learned or rejected based on the real time response. The proposed scheme has been successfully applied to indoor navigation. Wey-Shiuan Hwang, Juyang Weng |
ICIP (1) | 2 |
| 2001 | Autonomous speech acquisition of a robotabstractIt is difficult to program a robot that understands speech. Instead of learning speech from offline speech data, online and grounded learning by a robot can potentially address the robot speech recognition issue. Motivated by the human speech acquisition process, we propose to enable a robot to learn from interactive experience. The author presents some recent results of a robot that develops its speech-related skills through real-time interactions with its environment. Yilu Zhang, Juyang Weng |
SMC | 2 |
| 2000 | An Incremental Learning Method for Face Recognition under Continuous Video StreamabstractThe current technology in computer vision requires humans to collect images, store images, segment images for computers and train computer recognition systems using these images. It is unlikely that such a manual labor process can meet the demands of many challenging recognition tasks. Our goal is to enable machines to learn directly from sensory input streams while interacting with the environment including human teachers. We propose a new technique which incrementally derives discriminating features in the input space. Virtual labels are formed by clustering in the output space to extract discriminating features in the input space. We organize the resulting discriminating subspace in a coarse-to-fine fashion and store the information in a decision tree. Such an incremental hierarchical discriminating regression (IHDR) decision tree can be modeled by a hierarchical probability distribution model. We demonstrate the performance of the algorithm on the problem of face recognition using video sequences of 33889 frames in length from 143 different subjects. A correct recognition rate of 95.1% has been achieved. Juyang Weng, Colin H. Evans, Wey-Shiuan Hwang |
FG | 1 |
| 2000 | Hierarchical Discriminant Regression for Incremental and Real-Time Image Classification
Wey-Shiuan Hwang, Juyang Weng |
IDEAL | 2 |
| 2000 | Appearance-Based Hand Sign Recognition from Intensity Image Sequences
Yuntao Cui, Juyang Weng |
Comput. Vis. Image Underst. | 2 |
| 2000 | Hierarchical Discriminant RegressionabstractThe main motivation of this paper is to propose a classification and regression method for challenging high-dimensional data. The proposed technique casts classification problems and regression problems into a unified regression problem. This unified view enables classification problems to use numeric information in the output space that is available for regression problems but are traditionally not readily available for classification problems. A doubly clustered subspace-based hierarchical discriminating regression (HDR) method is proposed. The major characteristics include: (1) Clustering is performed in both output space and input space at each internal node, termed "doubly clustered." Clustering in the output space provides virtual labels for computing clusters in the input space. (2) Discriminants in the input space are automatically derived from the clusters in the input space. (3) A hierarchical probability distribution model is applied to the resulting discriminating subspace at each internal node. This realizes a coarse-to-fine approximation of probability distribution of the input samples, in the hierarchical discriminating subspaces. (4) To relax the per class sample requirement of traditional discriminant analysis techniques, a sample-size dependent negative-log-likelihood (NLL) is introduced. This new technique is designed for automatically dealing with small-sample applications, large-sample applications, and unbalanced-sample applications. (5) The execution of the HDR method is fast, due to the empirical logarithmic time complexity of the HDR algorithm. Although the method is applicable to any data, we report the experimental results for three types of data: synthetic data for examining the near-optimal performance, large raw face-image databases, and traditional databases with manually selected features along with a comparison with some major existing methods. Wey-Shiuan Hwang, Juyang Weng |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2000 | State-based SHOSLIF for indoor visual navigationabstractIn this paper, we investigate vision-based navigation using the self-organizing hierarchical optimal subspace learning and inference framework (SHOSLIF) that incorporates states and a visual attention mechanism. With states to keep the history information and regarding the incoming video input as an observation vector, the vision-based navigation is formulated as an observation-driven Markov model (ODMM). The ODMM can be realized through recursive partitioning regression. A stochastic recursive partition tree (SRPT), which maps an preprocessed current input raw image and the previous state into the current state and the next control signal, is used for efficient recursive partitioning regression. The SRPT learns incrementally: each learning sample is learned or rejected "on-the-fly." The purposed scheme has been successfully applied to indoor navigation. Shaoyun Chen, Juyang Weng |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 1999 | The developmental approach to multimedia speech learningabstractThis paper introduces the developmental approach to speech learning, motivated by human cognitive development from infancy to adulthood. Central in the developmental approach is what is called the developmental algorithm. We introduce AA-learning as a basic learning mode for our developmental algorithm. The developmental algorithm enables the system to learn new tasks without a need of reprogramming. Some experimental results for automated animal-like learning (AA-learning) using our developmental algorithm are presented. Juyang Weng, Yong-Beom Lee, Colin H. Evans |
ICASSP | 1 |
| 1999 | Moving industry-guided multimedia technology into the classroomabstractGiven the ubiquity of multimedia technology, it is important that Computer Science students not only learn the basics of multimedia design, but also gain hands-on experience with applications of the technology. This paper describes the integration of multimedia concepts and tools into a Computer Science curriculum. An NSF-sponsored Multimedia Laboratory was established and used to support three senior-level courses: software engineering, computer graphics, and computer networks. Curriculum development, laboratory exercises, and the role of projects are described. Philip K. McKinley, Betty H. C. Cheng, Juyang Weng |
SIGCSE | 3 |
| 1999 | A Learning-Based Prediction-and-Verification Segmentation Scheme for Hand Sign Image SequenceabstractWe present a prediction-and-verification segmentation scheme using attention images from multiple fixations. A major advantage of this scheme is that it can handle a large number of different deformable objects presented in complex backgrounds. The scheme is also relatively efficient. The system was tested to segment hands in sequences of intensity images, where each sequence represents a hand sign in American Sign Language. The experimental result showed a 95 percent correct segmentation rate with a 3 percent false rejection rate. Yuntao Cui, Juyang Weng |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 1999 | Hierarchical Discriminant Analysis for Image RetrievalabstractA self-organizing framework for object recognition is described. We describe a hierarchical database structure for image retrieval. The self-organizing hierarchical optimal subspace learning and inference framework (SHOSLIF) system uses the theories of optimal linear projection for optimal feature derivation and a hierarchical structure to achieve logarithmic retrieval complexity. A space-tessellation tree is generated using the most expressive features (MEF) and most discriminating features (MDF) at each level of the tree. The major characteristics of the analysis include: (1) avoiding the limitation of global linear features by deriving a recursively better-fitted set of features for each of the recursively subdivided sets of training samples; (2) generating a smaller tree whose cell boundaries separate the samples along the class boundaries better than the principal component analysis, thereby giving a better generalization capability (i.e., better recognition rate in a disjoint test); (3) accelerating the retrieval using a tree structure for data pruning, utilizing a different set of discriminant features at each level of the tree. We allow for perturbations in the size and position of objects in the images through learning. We demonstrate the technique on a large image database of widely varying real-world objects taken in natural settings, and show the applicability of the approach for variability in position, size, and 3D orientation. This paper concentrates on the hierarchical partitioning of the feature spaces. Daniel L. Swets, Juyang Weng |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 1998 | Autonomous Vision-Guided Robot Manipulation Control
Wey-Shiuan Hwang, Juyang Weng |
ACCV (2) | 2 |
| 1998 | Toward Automation of Learning: The State Self-Organization Problem for a Face Recognizer
Juyang Weng, Wey-Shiuan Hwang |
FG | 1 |
| 1998 | State-based SHOSLIF for indoor visual navigationabstractVision-based navigation is investigated using SHOSLIF that incorporates states and a visual attention mechanism. The problem is formulated as an observation-driven Markov model (ODMM) which is realized through recursive partitioning regression. A stochastic recursive partition tree (SRPT), which maps a preprocessed current input raw image and the previous state into the current state and the next control signal is used for efficient recursive partitioning regression. The SRPT learns incrementally: each learning sample is rejected or learned "on-the-fly". The proposed scheme has been successfully applied to indoor navigation. Shaoyun Chen, Juyang Weng |
ICPR | 2 |
| 1998 | Sensorimotor action sequence learning with application to face recognition under discourseabstractOur goal is to enable machines to learn directly from sensory input streams. The learning machine does not require human teacher to specify any content-level rule. Such a capability requires a fundamentally new way of addressing the learning problem, one that unifies learning and performance phases and requires a systematic self-organization capability. The presented approach enables the system to self-organize its internal representation, and uses a systematic way to automatically build multi-level representation. In the experiments presented, we study the behavior of the method for automatic state self-organization and automatic level building that involves two levels. We test the algorithm for the problem of face recognition under a simple but important discourse scenario-a primary mode of our goal for human-machine interactive learning. Juyang Weng, Wey-Shiuan Hwang |
ICPR | 1 |
| 1998 | An Image Database System with Support for Traditional Alphanumeric Queries and Content-Based Queries by Example
Daniel L. Swets, Yogesh Pathak, Juyang Weng |
Multim. Tools Appl. | 3 |
| 1998 | Vision-guided navigation using SHOSLIF
Juyang Weng, Shaoyun Chen |
Neural Networks | 1 |
| 1997 | Vision-guided robot manipulator control as learning and recall using SHOSLIFabstractWe present a general framework by which a robotic hand-eye system can perform learned tasks by recalling the action sequences that it has learned. In the training phase, the system learns the relationship between the sensors and the actuators from a series of training examples supplied interactively by a system trainer. The system automatically builds a recursive partition tree (RPT) which approximates the mapping from the input to the output. Each node of the RPT represents a cell of the space which is further partitioned by its children via a Voronoi tessellation. Each leaf node corresponds to a training sample and stores the corresponding output. In the performance phase, given an input, the RPT is used to retrieve the desired output by interpolating among all the leaf nodes that are good matches to the input. The RPT results in a logarithmic average time complexity in the number of stored training samples. Such a mechanism is used to accomplish major components of the system, including stereo calibration and sensor-based action sequence learning and execution. A hand-eye system with a PUMA 560 robotic manipulator is used to test the method. Wey-Shiuan Hwang, Juyang Weng |
ICRA | 2 |
| 1997 | Learning Recognition and Segmentation Using the Cresceptron
Juyang Weng, Narendra Ahuja, Thomas S. Huang |
Int. J. Comput. Vis. | 1 |
| 1997 | Optimal Registration of Object Views Using Range DataabstractThis paper deals with robust registration of object views in the presence of uncertainties and noise in depth data. Errors in registration of multiple views of a 3D object severely affect view integration during automatic construction of object models. We derive a minimum variance estimator (MVE) for computing the view transformation parameters accurately from range data of two views of a 3D object. The results of our experiments show that view transformation estimates obtained using MVE are significantly more accurate than those computed with an unweighted error criterion for registration. Chitra Dorai, Juyang Weng, Anil K. Jain 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 1997 | Transitory Image Sequences, Asymptotic Properties, and Estimation of Motion and StructureabstractA transitory image sequence is one in which no scene element is visible through the entire sequence. This article deals with some major theoretical and algorithmic issues associated with the task of estimating structure and motion from transitory image sequences. It is shown that integration with a transitory sequence has properties that are very different from those with a nontransitory one. Two representations, world-centered (WC) and camera-centered (CC), behave very differently with a transitory sequence. The asymptotic error rates derived in this article indicate that one representation is significantly superior to the other, depending on whether one needs camera-centered or world-centered estimates. We introduce an efficient "cross-frame" estimation technique for the CC representation. For the WC representation, our analysis indicates that a good technique should be based on camera global pose instead of interframe motions. Rigorous experiments were conducted with real-image sequences taken by a fully calibrated camera system. Juyang Weng, Yuntao Cui, Narendra Ahuja |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 1997 | Learning-Based Ventricle Detection from Cardiac MR and CT ImagesabstractThe objective of this work is to investigate the issue of automatically detecting regions of interest (ROI's) in medical images. It is assumed that the regions to be detected can be roughly segmented by a threshold based on a likelihood measure of the ROI. First, an analysis of the global histogram is used to compute a preliminary threshold that is likely near the optimal one. The histogram analysis is motivated by the analytical result of a bell image intensity model proposed in this work. Then, the preliminary threshold is used to segment the input image, resulting in an attention map, which contains an attention region that approximates the ROI as well as many spurious ones. Due to the nonoptimality of the preliminary threshold, it can happen that the attention region contains a part of, or more regions than, the ROI. Learning takes place in two stages: 1) learning for automatic selection of the preliminary threshold value and 2) learning for automatically selecting the ROI from the attention map while dynamically tuning the threshold according to the learned-likelihood function. Experiments have been conducted to approximately locate the endocardium boundaries of the left and right ventricles from gradient-echo magnetic resonance (MR) images. Cardiac computed tomography (CT) images have also been used for testing. The boundary of the segmented region provided by this algorithm is not very accurate and is meant to be used for further fine tuning based on other application-specific measures. Juyang Weng, Ajit Singh, M. Y. Chiu |
IEEE Trans. Medical Imaging | 1 |
| 1996 | Hand segmentation using learning-based prediction and verification for hand sign recognitionabstractThis paper presents a prediction-and-verification segmentation scheme wing attention images from multiple fixations. A major advantage of this scheme is that it can handle a large number of different deformable objects presented in complex backgrounds. The scheme is also relatively efficient since the segmentation is guided by the past knowledge through a prediction-and-verification scheme. The system has been tested to segment hands in the sequences of intensity images, where each sequence represents a hand sign. The experimental result showed a 95% correct segmentation rate with a 3% false rejection rate. Yuntao Cui, Juyang Weng |
CVPR | 2 |
| 1996 | Hand sign recognition from intensity image sequences with complex backgroundsabstractIn this paper, we have presented a new approach to recognize hand signs. In our approach, motion understanding (the hand movement) is tightly coupled with spatial recognition (hand shape). The system uses the multiclass, multidimensional discriminant analysis to automatically select the most discriminating features for gesture classification. A recursive partition tree approximator is proposed to do classification. This approach combined with our previous work on the hand segmentation forms a new framework which addresses three key aspects of the hand sign interpretation, that is the hand shape, the location, and the movement. The framework has been tested to recognize 28 different hand signs. The experimental results show that the system can achieve a 93.1% recognition rate for test sequences that have not been used in the training phase. Yuntao Cui, Juyang Weng |
FG | 2 |
| 1996 | Discriminant analysis and eigenspace partition tree for face and object recognition from viewsabstractThe method we have been using is based on our Self-Organizing Hierarchical Optimal Subspace Learning and Inference Framework (SHOSLIF). It uses the theories of linear discriminant projection for automatic optimal feature selection in each of the internal nodes of a Space-Tessellation Tree. In this paper, we present our recent study on the applicability of the approach to variability in position, size, and 3D orientation. In the work presented here, we require "well-framed" images os input for recognition. By well-framed images we mean that only a relatively small variation in the size, position, and orientation of the objects in the input images is allowed. We report the experimental results that show the performance difference between the subspaces of linear discriminant analysis and the principle component analysis and the effect of using a tree as opposed to a flat eigenspace. Daniel L. Swets, Juyang Weng |
FG | 2 |
| 1996 | View-based hand segmentation and hand-sequence recognition with complex backgroundsabstractIn this paper, we presents a three-stage framework to analyze time-varying image sequences. The focus of this paper is the second stage: segmentation. We propose a prediction-and-verification segmentation scheme which efficiently utilizes the attention images from the multiple fixations. The experimental results show 95% correct segmentation rate with 3% false rejection rate of 805 testing images. The recognition of hand sign based on the segmentation results has shown that the system has achieved a good performance for this very difficult vision task. Yuntao Cui, Juyang Weng |
ICPR | 2 |
| 1996 | Incremental learning for vision-based navigationabstractIn this paper, we explore the issue of incremental learning for autonomous navigation of a mobile robot. The autonomous navigation problem is regarded as a content-based retrieval problem where the robot learns the navigation experience using a hierarchical recursive partition tree (RPT). During real navigation, each time a new image is grabbed to retrieve the learned tree. The associated control signals of the retrieved are used to control the new action of the robot. Use of RPT can achieve efficient retrieval. In the proposed incremental learning scheme, a new image with the associated control signals is learned or rejected according to whether its retrieved output control signals are within tolerance of the desired control signals of the input query image. We use the eigen-subspace method for feature extraction in our incremental learning. The proposed algorithm has a real-time implementation for both learning and performance phases. Experimental results are shown to confirm the effectiveness of proposed method. Juyang Weng, Shaoyun Chen |
ICPR | 1 |
| 1996 | Using Discriminant Eigenfeatures for Image RetrievalabstractThis paper describes the automatic selection of features from an image training set using the theories of multidimensional discriminant analysis and the associated optimal linear projection. We demonstrate the effectiveness of these most discriminating features for view-based class retrieval from a large database of widely varying real-world objects presented as "well-framed" views, and compare it with that of the principal component analysis. Daniel L. Swets, Juyang Weng |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 1996 | Estimation of ellipse parameters using optimal minimum variance estimator
Yuntao Cui, Juyang Weng, Herbert Reynolds |
Pattern Recognit. Lett. | 2 |
| 1995 | Learning-Based Hand Sign Recognition Using SHOSLIF-MabstractWe present a self-organizing framework called the SHOSLIF-M for learning and recognizing spatiotemporal events (or patterns) from intensity image sequences. The proposed framework consists of a multiclass, multivariate discriminant analysis to automatically select the most discriminating features (MDF), a space partition tree to achieve a logarithmic retrieval time complexity for a database of n items, and a general interpolation scheme to do view inference and generalization in the MDF space based on a small number of training samples. The system is tested to recognize 28 different hand signs. The experimental results show that the learned system can achieve a 96% recognition rate for test sequences that have not been used in the training phase.> Yuntao Cui, Daniel L. Swets, Juyang Weng |
ICCV | 3 |
| 1995 | Genetic algorithms for object recognition in a complex sceneabstractA realworld computer vision module must deal with a wide variety of environmental parameters. Object recognition, one of the major tasks of this vision module, typically requires a preprocessing step to locate objects in the scenes that ought to be recognized. Genetic algorithms are a search technique for dealing with a very large search space, such as the one encountered in image segmentation or object recognition. The article describes a technique for using genetic algorithms to combine the image segmentation and object recognition steps for a complex scene. The results show that this approach is a viable method for successfully combining the image segmentation and object recognition steps for a computer vision module. Daniel L. Swets, William F. Punch, Juyang Weng |
ICIP | 3 |
| 1994 | Integration of transitory image sequencesabstractA transitory image sequence is one in which no scene element is visible through the entire sequence. This article deals with some major theoretical and algorithmic issues associated with the task of estimating structure and motion from transitory image sequences. Two representations, world-centered (WC) and camera-centered (CC), behave very differently with a transitory sequence. The asymptotical error properties derived in this article indicate that one representation is significantly superior to the other, depending on whether one uses camera-centered or world-centered estimates. Rigorous experiments were conducted with real-image sequences taken by a fully calibrated camera system. The comparison demonstrated that a good accuracy can be obtained from transitory image sequences.> Juyang Weng, Yuntao Cui, Narendra Ahuja, Ajit Singh |
CVPR | 1 |
| 1994 | Calibration for Peripheral Attenuation in Intensity ImagesabstractAn image taken from a typical camera loses its intensity and contrast around the periphery due to optical attenuation. A model is derived to characterize this effect quantitatively. This model is derived for a commonly used thick lens using the sine condition, and thus, is more general than those from the Gauss geometrical optics. Based on this model, the authors developed an algorithm to calibrate the periphery attenuation. Some experimental results are presented.> Shaoyun Chen, Juyang Weng |
ICIP (2) | 2 |
| 1993 | Learning recognition and segmentation of 3-D objects from 2-D imagesabstractA framework called Cresceptron is introduced for automatic algorithm design through learning of concepts and rules, thus deviating from the traditional mode in which humans specify the rules constituting a vision algorithm. With the Cresceptron, humans as designers need only to provide a good structure for learning, but they are relieved of most design details. The Cresceptron has been tested on the task of visual recognition by recognizing 3-D general objects from 2-D photographic images of natural scenes and segmenting the recognized objects from the cluttered image background. The Cresceptron uses a hierarchical structure to grow networks automatically, adaptively, and incrementally through learning. The Cresceptron makes it possible to generalize training exemplars to other perceptually equivalent items. Experiments with a variety of real-world images are reported to demonstrate the feasibility of learning in the Cresceptron.> Juyang Weng, Narendra Ahuja, Thomas S. Huang |
ICCV | 1 |
| 1993 | Image matching using the windowed Fourier phase
Juyang Weng |
Int. J. Comput. Vis. | 1 |
| 1993 | Optimal Motion and Structure EstimationabstractThe causes of existing linear algorithms exhibiting various high sensitivities to noise are analyzed. It is shown that even a small pixel-level perturbation may override the epipolar information that is essential for the linear algorithms to distinguish different motions. This analysis indicates the need for optimal estimation in the presence of noise. Methods are introduced for optimal motion and structure estimation under two situations of noise distribution: known and unknown. Computationally, the optimal estimation amounts to minimizing a nonlinear function. For the correct convergence of this nonlinear minimization, a two-step approach is used. The first step is using a linear algorithm to give a preliminary estimate for the parameters. The second step is minimizing the optimal objective function starting from that preliminary estimate as an initial guess. A remarkable accuracy improvement has been achieved by this two-step approach over using the linear algorithm alone.> Juyang Weng, Narendra Ahuja, Thomas S. Huang |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 1992 | Signal reconstruction from windowed Fourier phaseabstractA variety of information relating to phase has been increasingly widely used for stereo matching as well as motion image matching. In particular, the windowed Fourier phase (WFP) has several properties that are important for representing the structure of the signal. It is established that a series of signals, either continuous or discrete, is determined up to a multiplicative constant by its WFP at any frequency. An algorithm is developed to reconstruct the signals from the WFP.> Juyang Weng |
ICASSP | 1 |
| 1992 | Robust and physically-constrained interpolation of fluid flow fieldsabstractThe authors investigate the problem of interpolating, under physical constraints, 3-D vector fields from sample vectors at random positions. This problem arises from analysis of fluid motion, but the results can also be applied to such areas as geometric modeling, approximation theory, and other types of nonrigid motion. The algorithm proposed combines the generalized multivariable quadratic interpolation and physical constraints into one step to form an overdetermined linear equation system whose solution gives the coefficients of interpolation, which are much less sensitive to noise compared to other interpolation methods. The authors utilized methods in robust statistics to detect outliers in the sample data so that the results are more stable in the presence of gross errors. The algorithm is applied to both synthesized velocity fields of fluid and empirically measured 3-D velocity fields.> Jialin Zhong, Juyang Weng, Thomas S. Huang |
ICASSP | 2 |
| 1992 | Complete structure and motion from two monocular sequences without stereo correspondenceabstractIt is demonstrated that the motion and structure of rigidly moving objects can be completely determined from two monocular image sequences using only temporal matches. Three aspects of this scheme are useful: since stereo matching is not necessary, two cameras can view totally different parts of the rigid scene; as temporal disparity is usually significantly smaller than stereo disparity, matching needs only to deal with relatively small disparities; the recoverable scene structure is defined by the union of the fields of view of two cameras instead of the intersection, and so is much larger than that of a conventional stereo setup. Experiments with synthesized data and real world images are presented to demonstrate the feasibility of this scheme.> Juyang Weng, Thomas S. Huang |
ICPR (1) | 1 |
| 1992 | Matching Two Perspective ViewsabstractA computational approach to image matching is described. It uses multiple attributes associated with each image point to yield a generally overdetermined system of constraints, taking into account possible structural discontinuities and occlusions. In the algorithm implemented, intensity, edgeness, and cornerness attributes are used in conjunction with the constraints arising from intraregional smoothness, field continuity and discontinuity, and occlusions to compute dense displacement fields and occlusion maps along the pixel grids. The intensity, edgeness, and cornerness are invariant under rigid motion in the image plane. In order to cope with large disparities, a multiresolution multigrid structure is employed. Coarser level edgeness and cornerness measures are obtained by blurring the finer level measures. The algorithm has been tested on real-world scenes with depth discontinuities and occlusions. A special case of two-view matching is stereo matching, where the motion between two images is known. The algorithm can be easily specialized to perform stereo matching using the epipolar constraint.> Juyang Weng, Narendra Ahuja, Thomas S. Huang |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 1992 | Camera Calibration with Distortion Models and Accuracy EvaluationabstractA camera model that accounts for major sources of camera distortion, namely, radial, decentering, and thin prism distortions is presented. The proposed calibration procedure consists of two steps: (1) the calibration parameters are estimated using a closed-form solution based on a distribution-free camera model; and (2) the parameters estimated in the first step are improved iteratively through a nonlinear optimization, taking into account camera distortions. According to minimum variance estimation, the objective function to be minimized is the mean-square discrepancy between the observed image points and their inferred image projections computed with the estimated calibration parameters. The authors introduce a type of measure that can be used to directly evaluate the performance of calibration and compare calibrations among different systems. The validity and performance of the calibration procedure are tested with both synthetic data and real images taken by tele- and wide-angle lenses.> Juyang Weng, Paul R. Cohen, Marc Herniou |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 1992 | Motion and Structure from Line Correspondences; Closed-Form Solution, Uniqueness, and OptimizationabstractThis work discusses estimating motion and structure parameters from line correspondences of a rigid scene. The authors present a closed-form solution to motion and structure parameters from line correspondences through three monocular perspective views. The algorithm makes use of redundancy in the data to improve the accuracy of the solutions. The uniqueness of the solution is established, and necessary and sufficient conditions for degenerate spatial line configurations are given. Optimization has been employed to further improve the accuracy of the estimates in the presence of noise. Simulations have shown that the errors of the optimized estimates are close to the theoretical lower error bound.> Juyang Weng, Thomas S. Huang, Narendra Ahuja |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 1992 | Motion and structure estimation from stereo image sequencesabstractA closed-form approximate matrix-weighted solution for estimating motion parameters from point correspondences in two stereo image pairs is presented. The corresponding algorithm is noniterative and fast. Simulation shows that the solution is significantly more reliable than the unweighted and the scalar-weighted solutions, except when the number of point correspondences is less than seven. The matrix-weighted least squares solution can be used as a final solution if the speed requirement does not allow iterations. An approach to optimal estimation of the motion parameters and the structure of the three-dimensional points is also introduced. In experiments with a calibrated real stereo setup, the motion parameters and depth maps were automatically computed from real-world images. Some ground truths were used to validate the accuracy of these results.> Juyang Weng, Paul R. Cohen, Nicolas Rebibo |
IEEE Trans. Robotics Autom. | 1 |
| 1991 | Shading-Based Two-View Matching
Michel A. Audette, Paul R. Cohen, Juyang Weng |
IJCAI | 3 |
| 1990 | Extended structure and motion analysis from monocular image sequencesabstractThe issue of optimal motion and structure estimation from monocular image sequences with a rigidity scene is addressed. The method has the following characteristics: the dimension of the search space in the nonlinear optimization is drastically reduced by exploiting the relationship between structure and motion parameters; the degree of reliability of the observations and estimates is effectively taken into account; the proposed formulation allows arbitrary interframe motion; and the information about the structure of the scene, acquired from previous images, is systematically integrated into the new estimations. It is shown that any scale factor associated with two consecutive images in a monocular sequence is determined by the scale factor of the first two images. The simulations and experiments with long image sequences of real world scenes indicate that the optimization method developed greatly reduces the computational complexity and substantially improves the motion and structure estimation over that produced by linear algorithms.> Ning Cui, Juyang Weng, Paul R. Cohen |
ICCV | 2 |
| 1990 | A theory of image matchingabstractA theoretical framework is presented in which windowed Fourier phase (WFP) is introduced as the primary matching primitive. Zero-crossings and peaks correspond to special values in the phase profile. The WFP is quasi-linear, dense and its spatial period and slope are controlled by the channel scale. This framework has the following important characteristics: matching primitives are available almost everywhere to convey dense disparity information in every channel, either coarse or fine; the false target problem is virtually eliminated; matching is easier, uniform and can be performed by a network suitable for parallel computer architecture; and matching is fast since very few iterations are needed. In fact, the WFP is so informative that the original signal can be uniquely determined up to a multiplicative constant by the WFP in any one channel. The use of phase as the matching primitive is also supported by some existing psychophysical and neurophysiological studies. An implementation of the proposed theory has shown good results from images of random dots and natural scenes.> Juyang Weng |
ICCV | 1 |
| 1990 | Calibration of stereo cameras using a non-linear distortion model (CCD sensory)abstractA camera model is presented which accounts for major sources of camera distortion: radial, decentering, and thin-prism distortions. The proposed calibration procedure consists of two steps. In the first step, calibration parameters are estimated using a closed-form solution based on a distortion-free camera model. In the second step, the parameters estimated in the first step are improved iteratively through nonlinear optimization, taking into account camera distortions. According to minimum-variance estimation, the objective function to be minimized is the mean-square discrepancy between the observed image points and their inferred image projections computed with the estimated calibration parameters. A type of measure is introduced which can be used to directly evaluate the performance of the calibration and compare calibrations among different systems. The validity and the performance of the calibration procedure are tested on real images taken by wide-angle lenses. Results consistently show significant improvements over less complete camera models.> Juyang Weng, Paul R. Cohen, Marc Herniou |
ICPR (1) | 1 |
| 1990 | Estimating motion and structure from line matches: performance obtained and beyondabstractThe performance issues of estimating motion and structure from line correspondences are studied. An approach to optimal estimation of motion and structure using line correspondences is presented. To minimize the expected errors in the estimated parameters, it is necessary to minimize the matrix-weighted discrepancy between the computed lines and the observed lines. In order to reliably reach the global minimum solution, a closed-form solution is computed and then used as the initial starting condition for an iterative optimal estimation algorithm. Simulation results show that, in the presence of noise, the accuracy of the optimal solution is not only considerably better than that of the closed-form solutions, but it has also reached a level that it is comparable with that of point-based optimal algorithms. Simulations also show that the error of the optimal solution is close to a theoretical lower error bound, the Cramer-Rao bound, which implies that there exists little room for accuracy improvement beyond the performance obtained.> Juyang Weng, Thomas S. Huang, Narendra Ahuja |
ICPR (1) | 1 |
| 1989 | Optimal motion and structure estimationabstractThe problem of estimating motion and structure of a rigid scene from two perspective monocular views is studied. The optimization approach presented is motivated by the following observations of linear algorithms: (1) for certain types of motion, even pixel-level perturbations (such as digitization noise) may override the information characterized by epipolar constraint; (2) existing linear algorithms do not use the constraints in the essential parameter matrix E in solving for this matrix. The authors present approaches to estimating errors in the optimal solutions, investigate the theoretical lower bounds on the errors in the solutions and compare them with actual errors, and analyze two types of algorithms of optimization: batch and sequential. The analysis and experiments show that, in general, a batch technique performs better than a sequential technique for any nonlinear problems. A recursive batch processing technique is proposed for nonlinear problems that require recursive estimation.> Juyang Weng, Narendra Ahuja, Thomas S. Huang |
CVPR | 1 |
| 1989 | Motion and Structure From Two Perspective Views: Algorithms, Error Analysis, and Error EstimationabstractDeals with estimating motion parameters and the structure of the scene from point (or feature) correspondences between two perspective views. An algorithm is presented that gives a closed-form solution for motion parameters and the structure of the scene. The algorithm utilizes redundancy in the data to obtain more reliable estimates in the presence of noise. An approach is introduced to estimating the errors in the motion parameters computed by the algorithm. Specifically, standard deviation of the error is estimated in terms of the variance of the errors in the image coordinates of the corresponding points. The estimated errors indicate the reliability of the solution as well as any degeneracy or near degeneracy that causes the failure of the motion estimation algorithm. The presented approach to error estimation applies to a wide variety of problems that involve least-squares optimization or pseudoinverse. Finally the relationships between errors and the parameters of motion and imaging system are analyzed. The results of the analysis show, among other things, that the errors are very sensitive to the translation direction and the range of field view. Simulations are conducted to demonstrate the performance of the algorithms and error estimation as well as the relationships between the errors and the parameters of motion and imaging systems. The algorithms are tested on images of real-world scenes with point of correspondences computed automatically.> Juyang Weng, Thomas S. Huang, Narendra Ahuja |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 1988 | Closed-form solution+maximum likelihood: a robust approach to motion and structure estimationabstractA robust approach is presented to estimation motion and structure from image sequences. The approach consists of two steps. The first step is estimating the motion parameters using a robust linear algorithm that gives a closed-form solution for motion parameters and scene structure. The second step is improving the results from the linear algorithm using maximum-likelihood estimation. An algorithm using point correspondences from monocular images is discussed in detail and experimented with. An algorithm using line correspondences is briefly discussed. The simulations show that maximum-likelihood estimation achieves remarkable improvement over the preliminary estimates given by the linear algorithm. The algorithm is also tested on images of real scenes from automatically computed displacement field. The proposed approach is independent of the exact tokens used to establish correspondences, e.g. displacement flow, optical flow, or discrete features. Two or more types of tokens may be used, for monocular or binocular images.> Juyang Weng, Narendra Ahuja, Thomas S. Huang |
CVPR | 1 |
| 1988 | Estimating motion/structure from line correspondences: a robust linear algorithm and uniqueness theoremsabstractA closed-form solution to motion and structure from line correspondences in monocular perspective image sequences is presented. The algorithm requires a minimum of 13 lines over three perspective views. Redundancy in the data provides overdetermination to combat noise. The estimates can be used as an initial guess for further optimization. A unique solution to motion and structure is guaranteed if and only if the line configuration is not degenerate and the translation between any two views does not vanish. Necessary and sufficient conditions for degenerate spatial line configurations have been derived. Simulations are performed which show the performance of the algorithm in the presence of noise.> Juyang Weng, Yuncai Liu, Thomas S. Huang, Narendra Ahuja |
CVPR | 1 |
| 1988 | Two-view MatchingabstractEstablishing correspondences between images of the same scene is one of the most challenging and critical stcps in motion and scene analysis. Part of the difficulty is due to a wide variety of three-dimension structural discontinuities and occlusions that occur in real world scenes. This paper describes a computational approach to image matching that uses multiple attributes associated with a pixel to yield a generally overdetermined system of constraints, taking into account possible structural discontinuities and occlusions. In the algorithm implemented, intensity, edgeness, and comemess attributes are used in conjunction with the constraints arising from intraregional smoothness, field continuity and discontinuity, and occlusions to compute dense displacement fields and occlusion maps at pixel grids. A multiresolution multigrid structure is employed to deal with large disparities. Coarser level attributes are obtained by blurring the finer level attributes. The algorithms are tested on real world scenes containing depth discontinuities and occlusions. A special case of two-view matching is stereo matching where the motion between two images is known. The general algorithm given here can be easily spccialized to perfonn stereo matching using epipolar line constraint. Juyang Weng, Narendra Ahuja, Thomas S. Huang |
ICCV | 1 |
| 1988 | Motion and structure from point correspondences: a robust algorithm for planar case with error estimationabstractThe problem of determining motion and structure for a planar surface and the error estimation are discussed. Since the motion of a planar patch is a degenerate case for linear algorithms (algorithms that consist of solving mainly linear equations and give a closed-form solution) for general surfaces, the motion of such a planar surface is considered separately. An algorithm is introduced that gives a closed-form solution to motion parameters using monocular perspective images of the points on a planar surface. The algorithm is simpler and more reliable, in the presence of noise, than existing ones. There are generally two solutions for two image frames. For three image frames the solution is generally unique. An approach is proposed to test whether the points are coplanar. The errors in the motion parameters and surface structure can be estimated for each pair of images. Specifically, the standard deviation of the errors is calculated in terms of the variance of the errors in the image coordinates. This approach to estimating errors is applicable to least-squares, pseudo-inverse and eigenvalue-eigenvector problems.> Juyang Weng, Narendra Ahuja, Thomas S. Huang |
ICPR | 1 |
| 1987 | 3-D Motion Estimation, Understanding, and Prediction from Noisy Image SequencesabstractThis paper presents an approach to understanding general 3-D motion of a rigid body from image sequences. Based on dynamics, a locally constant angular momentum (LCAM) model is introduced. The model is local in the sense that it is applied to a limited number of image frames at a time. Specifically, the model constrains the motion, over a local frame subsequence, to be a superposition of precession and translation. Thus, the instantaneous rotation axis of the object is allowed to change through the subsequence. The trajectory of the rotation center is approximated by a vector polynomial. The parameters of the model evolve in time so that they can adapt to long term changes in motion characteristics. The nature and parameters of short term motion can be estimated continuously with the goal of understanding motion through the image sequence. The estimation algorithm presented in this paper is linear, i.e., the algorithm consists of solving simultaneous linear equations. Based on the assumption that the motion is smooth, object positions and motion in the near future can be predicted, and short missing subsequences can be recovered. Noise smoothing is achieved by overdetermination and a leastsquares criterion. The framework is flexible in the sense that it allows both overdetermination in number of feature points and the number of image frames. Juyang Weng, Thomas S. Huang, Narendra Ahuja |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |