EDBT 2026 Demo / reviewers in the wild / expert
Lei Xu 0001
dblp:19/360-1
· DBLP profile ↗
211ranked-venue papers
59as first author
64since 2021 · last 2026
0000-0002-2752-1573ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 149 · 48 first-author · 37 since 2021Applied, interdisciplinary, general and emerging computing · 42 · 5 first-author · 23 since 2021Graphics, computer vision, multimedia, augmented reality and games · 33 · 5 first-author · 14 since 2021Human-computer interaction and ubiquitous computing · 4 · 2 first-authorSystems, architecture and hardware · 2 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Theory of computation · 2 · 1 first-authorSoftware engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Full-Atom Peptide Design via Riemannian-Euclidean Bayesian Flow NetworksabstractDiffusion and flow matching models have recently emerged as promising approaches for peptide binder design. Despite their progress, these models still face two major challenges. First, categorical sampling of discrete residue types collapses their continuous parameters into one-hot assignments, while continuous variables (e.g., atom positions) evolve smoothly throughout the generation process. This mismatch disrupts the update dynamics and results in suboptimal performance. Second, current models assume unimodal distributions for side-chain torsion angles, which conflicts with the inherently multimodal nature of side-chain rotameric states and limits prediction accuracy. To address these limitations, we introduce PepBFN, the first Bayesian flow network for full-atom peptide design that directly models parameter distributions in fully continuous space. Specifically, PepBFN models discrete residue types by learning their continuous parameter distributions, enabling joint and smooth Bayesian updates with other continuous structural parameters. It further employs a novel Gaussian mixture-based Bayesian flow to capture the multimodal side-chain rotameric states and a Matrix Fisher-based Riemannian flow to directly model residue orientations on the SO(3) manifold. Together, these parameter distributions are progressively refined via Bayesian updates, yielding smooth and coherent peptide generation. Experiments on side-chain packing, reverse folding, and binder design tasks demonstrate the strong potential of PepBFN in computational peptide design. Shikui Tu, Lei Xu 0001 |
AAAI | 3 |
| 2026 | SEBSFormer: A Spectral-Enhanced Bi-Stream Transformer for Robust EEG DecodingabstractElectroencephalography (EEG) plays a vital role in clinical and cognitive applications such as epilepsy diagnosis and emotion recognition. However, the low signal-to-noise ratio, inter-subject variability, and inherent non-stationarity of EEG signals present substantial modeling challenges. While recent Transformer-based models offer promising long-range modeling capabilities, their self-attention mechanism behaves as a low-pass filter, suppressing high-frequency neural patterns critical for decoding transient events. In this work, we provide the first formal analysis demonstrating this low-pass behavior in self-attention mechanisms when applied to EEG signals, revealing a fundamental limitation of deep attention-based EEG models. To address this, we propose SEBSFormer, a spectral-enhanced bi-Stream Transformer that jointly models temporal dependencies and spectral structures. SEBSFormer integrates three key modules: a spectral compensation module that restores high-frequency components via residual correction in the Fourier domain; a multi-scale temporal attention module for saliency-guided temporal compression; and a graph-guided dynamic fusion module for adaptive spatial aggregation across electrodes. Extensive experiments on three benchmark datasets—TUAB, TUEV, and SEED—demonstrate that SEBSFormer consistently outperforms existing state-of-the-art models across both clinical and affective tasks. Our findings establish a new paradigm for frequency-aware EEG modeling. Shikui Tu, Lei Xu 0001 |
AAAI | 3 |
| 2026 | An EEG-based dual-stream spatial-spectral-temporal large model for self-limited epilepsy with centrotemporal spikes
Yun Ren, Fang Yuan 0010, Xuqin Chen, Shikui Tu, Lei Xu 0001 |
Expert Syst. Appl. | 6 |
| 2026 | Harnessing Diffusion Models for Image Manipulation With Partial SketchesabstractControllable image structure editing has attracted increasing attention. While recent interactive point-based methods are convenient and realistic, they often lack fine-grained control over localized content. Partial sketches provide a simple yet expressive interface for local structure manipulation. However, existing partial-sketch-based manipulation methods relying on generative adversarial networks (GANs) suffer from limited generalization and fidelity. Moreover, although diffusion-based adapters excel at global conditioning (e.g., edge maps), localized editing with partial strokes remains challenging due to two key issues: effectively injecting sparse stroke conditions during denoising and preserving non-edited regions to avoid unintended changes. To address these challenges, we propose DiffStroke, a mask-free framework for localized image manipulation with partial sketches. We introduce trainable Image-Stroke Fusion (ISF) blocks to fuse source images and strokes at the feature level, enabling precise local shape control while maintaining appearance consistency. We further develop a self-supervised mask estimator to protect irrelevant regions without manual input. Specifically, we leverage Tweedie's formula to estimate a clean latent image from noisy latents, blend the denoised result with the source, and train the mask estimator by minimizing the error between the blended latent and the target latent. Experiments on natural and facial images demonstrate that DiffStroke outperforms state-of-the-art methods on both simple and complex stroke-based editing tasks. DiffStroke can also be combined with text prompts to produce diverse and creative results. Code is available at https://github.com/CMACH508/DiffStroke. Tengjie Li, Shikui Tu, Lei Xu 0001 |
IEEE Trans. Image Process. | 3 |
| 2025 | Enhancing piRNA Cleavage Prediction with piR-DANN: Integrating In Vitro and In Vivo InsightsabstractPIWI-interacting RNAs (piRNAs) are critical genome guardians that guide PIWI proteins to cleave trans-posable element transcripts. Predicting these cleavage events is vital for understanding gene regulation but is challenged by the complex, “relaxed” rules of piRNA targeting and the noisy cellular context of in vivo data. Previous models, trained solely on in vivo data, struggle with high noise, non-functional binding events, and binary outputs that fail to capture the quantitative biophysics of piRNA-target interactions. Recognizing that precise, quantitative in vitro data directly quantifies biophysical parameters of cleavage, we aimed to leverage this information to overcome the noise and complexity inherent in in vivo observations. To achieve this, we developed piR-DANN, a deep learning framework based on a novel, biologically-informed adversarial domain adaptation strategy. Our core innovation moves beyond conventional approaches by providing the domain classifier with structural and sequence determinants of targeting summarizing known biological rules, in addition to the deep features learned by the model. This asymmetric design compels the feature extractor to learn the fundamental, domain-invariant principles of piRNA targeting. piR-DANN outperforms existing benchmarks predictive performance, with AUROC scores of 94.7% and 99.7% on two independent test sets. Furthermore, counterfactual analysis reveals a positional importance map concordant with PIWI catalytic core constraints, validating its extraction of biological signal from noise. By integrating heterogeneous data into an accurate, interpretable framework, our work explores piRNA biology and proposes a generalizable approach for deciphering gene regulatory systems. Shikui Tu, Lei Xu 0001 |
BIBM | 3 |
| 2025 | CATSyn: Predicting Synergistic Drug Combinations Through Context-Aware Heterogeneous Graph Convolution ModelabstractAccurately predicting drug synergy in cancer therapy remains challenging due to the strong dependence of drug effectiveness on cell-line context. Most existing models overlook this variability and fail to fully capture drug-cell line interactions. We present CATSyn, a Context-Aware heTerogeneous graph model for synergistic drug combination prediction. CATSyn introduces a context-aware attention mechanism that dynamically adjusts network weights based on cell-line environments, capturing cellspecific drug effects while maintaining generalization. To model these effects, we construct heterogeneous graphs that integrate drug and cell-line features into composite nodes, supported by universal nodes to share global information. Experiments on benchmark datasets demonstrate that CATSyn achieves state-of-the-art performance in both standard and unseen cell-line settings, highlighting its ability to balance specificity and generalization in synergy prediction. Biyang Zeng, Shikui Tu, Wen Zhang 0008, Lei Xu 0001 |
BIBM | 4 |
| 2025 | Towards Generalizable Reinforcement Learning via Causality-Guided Self-Adaptive RepresentationsabstractGeneral intelligence requires quick adaptation across tasks. While existing reinforcement learning (RL) methods have made progress in generalization, they typically assume only distribution changes between source and target domains. In this paper, we explore a wider range of scenarios where not only the distribution but also the environment spaces may change. For example, in the CoinRun environment, we train agents from easy levels and generalize them to difficulty levels where there could be new enemies that have never occurred before. To address this challenging setting, we introduce a causality-guided self-adaptive representation-based approach, called CSR, that equips the agent to generalize effectively across tasks with evolving dynamics. Specifically, we employ causal representation learning to characterize the latent causal variables within the RL system. Such compact causal representations uncover the structural relationships among variables, enabling the agent to autonomously determine whether changes in the environment stem from distribution shifts or variations in space, and to precisely locate these changes. We then devise a three-step strategy to fine-tune the causal model under different scenarios accordingly. Empirical experiments show that CSR efficiently adapts to the target domains with only a few samples and outperforms state-of-the-art baselines on a wide range of scenarios, including our simulated environments, CartPole, CoinRun and Atari games. Yupei Yang, Biwei Huang, Shikui Tu, Lei Xu 0001 |
ICLR | 6 |
| 2025 | S3GA: Towards Scalable Self-Supervised Learning for Large Scale Graph AlignmentabstractGraph Alignment (GA) is an NP-hard combinatorial challenge. The existing methods usually work on small-scale graphs because their computational complexity grows as the square of the number of nodes. Recent techniques relied on alignment labels to divide the large source and target graphs into small ones respectively, which are then aligned separately. However, the alignment labels in expansive real-world graphs are often scarce, which makes GA even more challenging. To address these issues, we propose a novel self-supervised learning framework that is able to work on the graphs of million nodes without any alignment labels, where the existing methods usually fail to output an answer within a reasonable time. Our method also falls in the divide-and-conquer paradigm. We do not cluster the source and target graph separately by the existing graph clustering algorithms, because it is difficult to pair and then align the cluster-induced source and target subgraphs in the absence of alignment labels. We devise an Alignment-Aware Clustering (AAC) method to partition the source and target nodes jointly and keep the ought-to-be-aligned pairs of nodes within the same cluster as much as possible. Furthermore, we develop a Topology-Aware Repartition (TAR) that leverages pseudo-alignment labels as anchors for re-clustering, preserving detailed structural information within clusters. This enables graph neural networks to effectively enhance the graph representation learning for node alignments which are self-supervised by a graph-matching solver. Extensive experiments have demonstrated that our method greatly reduces computational demands and sustains high alignment accuracies in large-scale graph applications. Wenqi Guo, Shikui Tu, Lei Xu 0001 |
KDD (2) | 3 |
| 2025 | Text to Sketch Generation with Multi-StylesabstractRecent advances in vision-language models have facilitated progress in sketch generation. However, existing specialized methods primarily focus on generic synthesis and lack mechanisms for precise control over sketch styles. In this work, we propose a training-free framework based on diffusion models that enables explicit style guidance via textual prompts and referenced style sketches. Unlike previous style transfer methods that overwrite key and value matrices in self-attention, we incorporate the reference features as auxiliary information with linear smoothing and leverage a style-content guidance mechanism. This design effectively reduces content leakage from reference sketches and enhances synthesis quality, especially in cases with low structural similarity between reference and target sketches. Furthermore, we extend our framework to support controllable multi-style generation by integrating features from multiple reference sketches, coordinated via a joint AdaIN module. Extensive experiments demonstrate that our approach achieves high-quality sketch generation with accurate style alignment and improved flexibility in style control. The official implementation of M3S is available at https://github.com/CMACH508/M3S. Tengjie Li, Shikui Tu, Lei Xu 0001 |
NeurIPS | 3 |
| 2025 | KeeA*: Epistemic Exploratory A* Search via Knowledge CalibrationabstractIn recent years, neural network-guided heuristic search algorithms, such as Monte-Carlo tree search and A$^\*$ search, have achieved significant advancements across diverse practical applications. Due to the challenges stemming from high state-space complexity, sparse training datasets, and incomplete environmental modeling, heuristic estimations manifest uncontrolled inherent biases towards the actual expected evaluations, thereby compromising the decision-making quality of search algorithms. Sampling exploration enhanced A$^\*$ (SeeA$^\*$) was proposed to improve the efficiency of A$^\*$ search by constructing an dynamic candidate subset through random sampling, from which the expanded node was selected. However, uniform sampling strategy utilized by SeeA$^\*$ facilitates exploration exclusively through the injection of randomness, which completely neglects the heuristic knowledge relevant to open nodes. Moreover, the theoretical support of cluster sampling remains ambiguous. Despite the existence of potential biases, heuristic estimations still encapsulate certain valuable information. In this paper, epistemic exploratory A$^\*$ search (KeeA$^\*$) is proposed to integrate heuristic knowledge for calibrating the sampling process. We first theoretically demonstrate that SeeA$^\*$ with cluster sampling outperforms uniform sampling due to the distribution-aware selection with higher variance. Building on this insight, cluster scouting and path-aware sampling are introduced in KeeA$^\*$ to further exploit heuristic knowledge to increase the sampling mean and variance, respectively, thereby generating higher-quality extreme candidates and enhancing overall decision-making performance. Finally, empirical results on retrosynthetic planning and logic synthesis demonstrate superior performance of KeeA$^*$ compared to state-of-the-art heuristic search algorithms. Dengwei Zhao, Shikui Tu, Yanan Sun 0003, Lei Xu 0001 |
NeurIPS | 4 |
| 2025 | Prior-Guided Flow Matching for Target-Aware Molecule Design with Learnable Atom NumberabstractStructure-based drug design (SBDD), aiming to generate 3D molecules with high binding affinity toward target proteins, is a vital approach in novel drug discovery. Although recent generative models have shown great potential, they suffer from unstable probability dynamics and mismatch between generated molecule size and the protein pockets geometry, resulting in inconsistent quality and off-target effects. We propose PAFlow, a novel target-aware molecular generation model featuring prior interaction guidance and a learnable atom number predictor. PAFlow adopts the efficient flow matching framework to model the generation process and constructs a new form of conditional flow matching for discrete atom types. A protein–ligand interaction predictor is incorporated to guide the vector field toward higher-affinity regions during generation, while an atom number predictor based on protein pocket information is designed to better align generated molecule size with target geometry. Extensive experiments on the CrossDocked2020 benchmark show that PAFlow achieves a new state-of-the-art in binding affinity (up to -8.31 Avg. Vina Score), simultaneously maintains favorable molecular properties. Jingyuan Zhou, Shikui Tu, Lei Xu 0001 |
NeurIPS | 4 |
| 2025 | Cauchy-Schwarz bounded trade-off weighting for causal inference with small sample sizes
Shikui Tu, Lei Xu 0001 |
Int. J. Approx. Reason. | 3 |
| 2025 | GraphFusion: Integrative prediction of drug synergy using multi-scale graph representations and cell line contexts
Biyang Zeng, Shikui Tu, Lei Xu 0001 |
J. Biomed. Informatics | 3 |
| 2025 | SketchMLP: effectively utilize rasterized images and drawing sequences for sketch recognition
Tengjie Li, Shikui Tu, Lei Xu 0001 |
Mach. Learn. | 3 |
| 2025 | PFB-Diff: Progressive Feature Blending diffusion for text-driven image editing
Wenjing Huang 0005, Shikui Tu, Lei Xu 0001 |
Neural Networks | 3 |
| 2025 | Multi-Objective Structure-Based Drug Design Using Causal DiscoveryabstractStructure-based drug design (SBDD) is a critical subtask in the drug discovery process, with deep generative models playing a pivotal role. Inherently, drug design is a multi-objective task given the fact that a promising drug candidate must satisfy multiple properties. However, existing SBDD methods either focus solely on the binding affinity between molecules and target proteins while neglecting other crucial properties, or they assume that objective properties are independent of each other. Yet there are often potential relationships among properties, which can be conflicting-improving one property may lead to the deterioration of another. The lack of consideration for these relationships in current methods makes it unfeasible to generate molecules that simultaneously meet multiple objectives. To address the above issues, a multi-objective SBDD algorithm is proposed based on the diffusion model to optimize binding affinity and other drug properties simultaneously. Multiple expert networks are trained in parallel to predict properties for molecules in intermediate states and transmit gradients, and a causal graph is constructed through the causal discovery algorithm to unveil the underlying relationships among target properties. During the entire generation process, the joint distribution of target properties is decomposed in a reasonable manner according to the casual graph, and then the gradients of each property are applied to guide the optimizing direction of generation. Experimental results indicate that our model effectively optimizes multiple objectives simultaneously, generating molecules with greater drug potential compared to baseline models. Jingyuan Zhou, Dengwei Zhao, Shikui Tu, Lei Xu 0001 |
IEEE Trans. Comput. Biol. Bioinform. | 5 |
| 2024 | Multilevel Attention Network with Semi-supervised Domain Adaptation for Drug-Target PredictionabstractPrediction of drug-target interactions (DTIs) is a crucial step in drug discovery, and deep learning methods have shown great promise on various DTI datasets. However, existing approaches still face several challenges, including limited labeled data, hidden bias issue, and a lack of generalization ability to out-of-domain data. These challenges hinder the model's capacity to learn truly informative interaction features, leading to shortcut learning and inferior predictive performance on novel drug-target pairs. To address these issues, we propose MlanDTI, a semi-supervised domain adaptive multilevel attention network (Mlan) for DTI prediction. We utilize two pre-trained BERT models to acquire bidirectional representations enriched with information from unlabeled data. Then, we introduce a multilevel attention mechanism, enabling the model to learn domain-invariant DTIs at different hierarchical levels. Moreover, we present a simple yet effective semi-supervised pseudo-labeling method to further enhance our model's predictive ability in cross-domain scenarios. Experiments on four datasets show that MlanDTI achieves state-of-the-art performances over other methods under intra-domain settings and outperforms all other approaches under cross-domain settings. The source code is available at https://github.com/CMACH508/MlanDTI. Zhousan Xie, Shikui Tu, Lei Xu 0001 |
AAAI | 3 |
| 2024 | A Multimodal Deep Neural Network for Causally Learning Chest X-ray ImagesabstractCausal intervention has been widely used in deep learning to tackle the confounding problems when the data are out of distribution. A representative class of causal-intervention-based strategy is invariant risk minimization, yet manually annotating is indispensable to indicate the environmental splits. A feasible solution is to use a pair of complementary attention, one of which focuses on the foreground target and another extracts the background features. The environments are then split unsupervisedly. However, in terms of medical images, the background is not as simple as normal images (e.g., a camel in desert), and using all the background feature as confounder is redundant and will degrade the expected generalization performance. In this paper, we develop a novel multimodal deep neural network which automatically extracts the confounded background features of chest X-ray (CXR) images by a multimodal approach, i.e., feature fusion with explicit confounders such as demographic information. To achieve this, we design an architecture with two modules, a classifier module taking images as input and an environment split module taking demographic tables as input. The cross attention is then conducted between the encoded background features and demographic features to extract the actual confounded background features. The proposed method meaningfully improves the generalization performance on multiple CXR datasets and accurately locates the lesions. The source code is available at https://github.com/CMACH508/CausalCXR. Shikui Tu, Lei Xu 0001 |
BIBM | 4 |
| 2024 | MedDet: Generative Adversarial Distillation for Efficient Cervical Disc Herniation DetectionabstractCervical disc herniation (CDH) is a prevalent musculoskeletal disorder that significantly impacts health and requires labor-intensive analysis from experts. Despite advancements in automated detection of medical imaging, two significant challenges hinder the real-world application of these methods. First, the computational complexity and resource demands present a significant gap for real-time application. Second, noise in MRI reduces the effectiveness of existing methods by distorting feature extraction. To address these challenges, we propose three key contributions: Firstly, we introduced MedDet, which leverages the multi-teacher single-student knowledge distillation for model compression and efficiency, meanwhile integrating generative adversarial training to enhance performance. Additionally, we customize the second-order nmODE to improve the model’s resistance to noise in MRI. Lastly, we conducted comprehensive experiments on the CDH-1848 dataset, achieving up to a 5% improvement in mAP compared to previous methods. Our approach also delivers over 5 times faster inference speed, with approximately 67.8% reduction in parameters and 36.9% reduction in FLOPs compared to the teacher model. These advancements significantly enhance the performance and efficiency of automated CDH detection, demonstrating promising potential for future application in clinical practice. Zeyu Zhang 0006, Nengmin Yi, Shengbo Tan, Ying Cai 0002, Yi Yang 0001, Lei Xu 0001, Qingtai Li, Daji Ergu, Yang Zhao 0019 |
BIBM | 6 |
| 2024 | Novelty Encouraged Beam Clustering Search for Multi-Objective De Novo Diverse Drug DesignabstractThe generation of drug-like, high-quality molecules from scratch within the expansive chemical space is a significant challenge in drug discovery. In previous research, value-based reinforcement learning algorithms have been utilized to optimize multiple desired properties simultaneously. Randomness is injected into the decision-making process through ε-Greedy or stochastic sampling to enable the generation of a diverse ensemble of molecules, which usually encounters a trade-off between the optimality and diversity of these generated molecules. Moreover, novelty has not been explicitly addressed as an optimization objective, and the distinctiveness of generated molecules from the reference molecules is not guaranteed. In this paper, novelty-encouraged beam clustering (NeBC) search algorithm is proposed for de novo drug design. A clustering strategy is integrated with heuristic value-guided beam search to strike a balance between the optimality and diversity of the generated molecules. An intrinsic reward, which is measured by the disagreement of a group of experts trained on reference molecules, is proposed to encourage novelty explicitly. Experimental results demonstrate that NeBC search not only achieves a balanced trade-off between optimality and diversity but also effectively enhances the novelty of the generated molecules. The source code is publicly accessible on https://github.com/CMACH508/NeBC. Dengwei Zhao, Shikui Tu, Lei Xu 0001 |
BIBM | 3 |
| 2024 | SketchEdit: Editing Freehand Sketches at the Stroke-Level
Tengjie Li, Shikui Tu, Lei Xu 0001 |
IJCAI | 3 |
| 2024 | Self-Supervised Learning for Enhancing Spatial Awareness in Free-Hand Sketches
Tengjie Li, Sicong Zang, Shikui Tu, Lei Xu 0001 |
IJCAI | 5 |
| 2024 | Boosting Efficiency in Task-Agnostic Exploration through Causal Knowledge
Yupei Yang, Biwei Huang, Shikui Tu, Lei Xu 0001 |
IJCAI | 4 |
| 2024 | SeeA*: Efficient Exploration-Enhanced A* Search by Selective SamplingabstractMonte-Carlo tree search (MCTS) and reinforcement learning contributed crucially to the success of AlphaGo and AlphaZero, and A$^*$ is a tree search algorithm among the most well-known ones in the classical AI literature. MCTS and A$^*$ both perform heuristic search and are mutually beneficial. Efforts have been made to the renaissance of A$^*$ from three possible aspects, two of which have been confirmed by studies in recent years, while the third is about the OPEN list that consists of open nodes of A$^*$ search, but still lacks deep investigation. This paper aims at the third, i.e., developing the Sampling-exploration enhanced A$^*$ (SeeA$^*$) search by constructing a dynamic subset of OPEN through a selective sampling process, such that the node with the best heuristic value in this subset instead of in the OPEN is expanded. Nodes with the best heuristic values in OPEN are most probably picked into this subset, but sometimes may not be included, which enables SeeA$^*$ to explore other promising branches. Three sampling techniques are presented for comparative investigations. Moreover, under the assumption about the distribution of prediction errors, we have theoretically shown the superior efficiency of SeeA$^*$ over A$^*$ search, particularly when the accuracy of the guiding heuristic function is insufficient. Experimental results on retrosynthetic planning in organic chemistry, logic synthesis in integrated circuit design, and the classical Sokoban game empirically demonstrate the efficiency of SeeA$^*$, in comparison with the state-of-the-art heuristic search algorithms. Dengwei Zhao, Shikui Tu, Lei Xu 0001 |
NeurIPS | 3 |
| 2024 | Automated diagnosis of cervical spine physiological curvature based on deep neural networks with transformer by using nmODE
Qingtai Li, Yi Yang 0001, Lei Xu 0001, Yiwei Shen, Nengmin Yi, Daji Ergu, Ying Cai 0002 |
Appl. Intell. | 3 |
| 2024 | KGDiff: towards explainable target-aware molecule generation with knowledge guidanceabstractDesigning 3D molecules with high binding affinity for specific protein targets is crucial in drug design. One challenge is that the atomic interaction between molecules and proteins in 3D space has to be taken into account. However, the existing target-aware methods solely model the joint distribution between the molecules and proteins, disregarding the binding affinities between them, which leads to limited performance. In this paper, we propose an explainable diffusion model to generate molecules that can be bound to a given protein target with high affinity. Our method explicitly incorporates the chemical knowledge of protein-ligand binding affinity into the diffusion model, and uses the knowledge to guide the denoising process towards the direction of high binding affinity. Specifically, an SE(3)-invariant expert network is developed to fit the Vina scoring functions and jointly trained with the denoising network, while the domain knowledge is distilled and conveyed from Vina functions to the expert network. An effective guidance is proposed on both continuous atom coordinates and discrete atom types by taking advantages of the gradient of the expert network. Experiments on the benchmark CrossDocked2020 demonstrate the superiority of our method. Additionally, an atom-level explanation of the generated molecules is provided, and the connections with the domain knowledge are established. Wenjing Huang 0005, Shikui Tu, Lei Xu 0001 |
Briefings Bioinform. | 4 |
| 2024 | Lmser-pix2seq: Learning stable sketch representations for sketch healing
Tengjie Li, Sicong Zang, Shikui Tu, Lei Xu 0001 |
Comput. Vis. Image Underst. | 4 |
| 2024 | A Deep Reinforcement Learning Approach for Portfolio Management in Non-Short-Selling MarketabstractReinforcement learning (RL) has been applied to financial portfolio management in recent years. Current studies mostly focus on profit accumulation without much consideration of risk. Some risk‐return balanced studies extract features from price and volume data only, which is highly correlated and missing representation of risk features. To tackle these problems, we propose a weight control unit (WCU) to effectively manage the position of portfolio management in different market statuses. A loss penalty term is also designed in the reward function to prevent sharp drawdown during trading. Moreover, stock spatial interrelation representing the correlation between two different stocks is captured by a graph convolution network based on fundamental data. Temporal interrelation is also captured by a temporal convolutional network based on new factors designed with price and volume data. Both spatial and temporal interrelation work for better feature extraction from historical data and also make the model more interpretable. Finally, a deep deterministic policy gradient actor–critic RL is applied to explore optimal policy in portfolio management. We conduct our approach in a challenging non‐short‐selling market, and the experiment results show that our method outperforms the state‐of‐the‐art methods in both profit and risk criteria. Specifically, with 6.72% improvement on an annualized rate of return, 7.72% decrease in maximum drawdown, and a better annualized Sharpe ratio of 0.112. Also, the loss penalty and WCU provide new aspects for future work in risk control. Ruidan Su, Chun Chi, Shikui Tu, Lei Xu 0001 |
IET Signal Process. | 4 |
| 2024 | IA-NGM: A bidirectional learning method for neural graph matching with feature fusion
Tianxiang Qin, Shikui Tu, Lei Xu 0001 |
Mach. Learn. | 3 |
| 2024 | De Novo Drug Design by Multi-Objective Path Consistency Learning With Beam A* SearchabstractGenerating high-quality and drug-like molecules from scratch within the expansive chemical space presents a significant challenge in the field of drug discovery. In prior research, value-based reinforcement learning algorithms have been employed to generate molecules with multiple desired properties iteratively. The immediate reward was defined as the evaluation of intermediate-state molecules at each step, and the learning objective would be maximizing the expected cumulative evaluation scores for all molecules along the generative path. However, this definition of the reward was misleading, as in reality, the optimization target should be the evaluation score of only the final generated molecule. Furthermore, in previous works, randomness was introduced into the decision-making process, enabling the generation of diverse molecules but no longer pursuing the maximum future rewards. In this paper, immediate reward is defined as the improvement achieved through the modification of the molecule to maximize the evaluation score of the final generated molecule exclusively. Originating from the A search, path consistency (PC), i.e., values on one optimal path should be identical, is employed as the objective function in the update of the value estimator to train a multi-objective de novo drug designer. By incorporating the value into the decision-making process of beam search, the DrugBA algorithm is proposed to enable the large-scale generation of molecules that exhibit both high quality and diversity. Experimental results demonstrate a substantial enhancement over the state-of-the-art algorithm QADD in multiple molecular properties of the generated molecules. Dengwei Zhao, Jingyuan Zhou, Shikui Tu, Lei Xu 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 4 |
| 2024 | Self-Organizing a Latent Hierarchy of Sketch Patterns for Controllable Sketch SynthesisabstractEncoding sketches as Gaussian mixture model (GMM)-distributed latent codes is an effective way to control sketch synthesis. Each Gaussian component represents a specific sketch pattern, and a code randomly sampled from the Gaussian can be decoded to synthesize a sketch with the target pattern. However, existing methods treat the Gaussians as individual clusters, which neglects the relationships between them. For example, the giraffe and horse sketches heading left are related to each other by their face orientation. The relationships between sketch patterns are important messages to reveal cognitive knowledge in sketch data. Thus, it is promising to learn accurate sketch representations by modeling the pattern relationships into a latent structure. In this article, we construct a tree-structured taxonomic hierarchy over the clusters of sketch codes. The clusters with the more specific descriptions of sketch patterns are placed at the lower levels, while the ones with the more general patterns are ranked at the higher levels. The clusters at the same rank relate to each other through the inheritance of features from common ancestors. We propose a hierarchical expectation-maximization (EM)-like algorithm to explicitly learn the hierarchy, jointly with the training of encoder-decoder network. Moreover, the learned latent hierarchy is utilized to regularize sketch codes with structural constraints. Experimental results show that our method significantly improves controllable synthesis performance and obtains effective sketch analogy results. Sicong Zang, Shikui Tu, Lei Xu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2023 | Self-Supervised Bidirectional Learning for Graph MatchingabstractDeep learning methods have demonstrated promising performance on the NP-hard Graph Matching (GM) problems. However, the state-of-the-art methods usually require the ground-truth labels, which may take extensive human efforts or be impractical to collect. In this paper, we present a robust self-supervised bidirectional learning method (IA-SSGM) to tackle GM in an unsupervised manner. It involves an affinity learning component and a classic GM solver. Specifically, we adopt the Hungarian solver to generate pseudo correspondence labels for the simple probabilistic relaxation of the affinity matrix. In addition, a bidirectional recycling consistency module is proposed to generate pseudo samples by recycling the pseudo correspondence back to permute the input. It imposes a consistency constraint between the pseudo affinity and the original one, which is theoretically supported to help reduce the matching error. Our method further develops a graph contrastive learning jointly with the affinity learning to enhance its robustness against the noise and outliers in real applications. Experiments deliver superior performance over the previous state-of-the-arts on five real-world benchmarks, especially under the more difficult outlier scenarios, demon- strating the effectiveness of our method. Wenqi Guo, Shikui Tu, Lei Xu 0001 |
AAAI | 4 |
| 2023 | Linking Sketch Patches by Learning Synonymous Proximity for Graphic Sketch RepresentationabstractGraphic sketch representations are effective for representing sketches. Existing methods take the patches cropped from sketches as the graph nodes, and construct the edges based on sketch's drawing order or Euclidean distances on the canvas. However, the drawing order of a sketch may not be unique, while the patches from semantically related parts of a sketch may be far away from each other on the canvas. In this paper, we propose an order-invariant, semantics-aware method for graphic sketch representations. The cropped sketch patches are linked according to their global semantics or local geometric shapes, namely the synonymous proximity, by computing the cosine similarity between the captured patch embeddings. Such constructed edges are learnable to adapt to the variation of sketch drawings, which enable the message passing among synonymous patches. Aggregating the messages from synonymous patches by graph convolutional networks plays a role of denoising, which is beneficial to produce robust patch embeddings and accurate sketch representations. Furthermore, we enforce a clustering constraint over the embeddings jointly with the network learning. The synonymous patches are self-organized as compact clusters, and their embeddings are guided to move towards their assigned cluster centroids. It raises the accuracy of the computed synonymous proximity. Experimental results show that our method significantly improves the performance on both controllable sketch synthesis and sketch healing. Sicong Zang, Shikui Tu, Lei Xu 0001 |
AAAI | 3 |
| 2023 | PEST: A General-Purpose Protein Embedding Model for Homology SearchabstractFinding known homologs of newly predicted proteins is essential for understanding their functions and mechanisms. It is a highly complex task because proteins undergo various changes during evolution. Traditional methods based on sequence or structure alignment either have low accuracy or take a long time. Recent deep learning-based methods primarily focus on structural information, yet they can’t fully exploiting protein information. To solve this problem, in this paper, we propose a novel general-purpose protein embedding model that can be used for homology search. It first employs a protein language pre-trained model to extract protein sequence embeddings, capturing intricate biological patterns. Subsequently, a Transformer integrating protein structural information generates the high-level representations. By combining protein sequence and structural features, the model can effectively exploit the rich contextual and spatial information inherent in proteins. We applied the model to the SCOP dataset for protein superfamily classification, achieving a classification accuracy of 86.97%, outperforming state-of-the-art method by 7.91%. The source code has been published on GitHub (https://github.com/CMACH508/PEST). Yongchang Liu, Peiying Li, Shikui Tu, Lei Xu 0001 |
BIBM | 4 |
| 2023 | A Deep Learning Method with Multi-view Attention and Multi-branch GCN for BECT DiagnosisabstractEpilepsy is a prevalent chronic neurological disorder in childhood, imposing a heavy burden on patients and their families. The development of deep learning and the accumulation of clinical medical data has led to a surge in neural network algorithms proposed for the automatic detection of childhood epilepsy using electroencephalogram (EEG) signals. However, for Benign Childhood Epilepsy with Centro-Temporal Spikes (BECT), the most common type of childhood epilepsy, there is no large-scale public dataset available for its detection. Moreover, although there were a few studies of BECT using private data, they focused only on identifying the presence of abnormal discharges, ignoring the sleep stage of the discharges, which is actually a critical factor for the doctor’s diagnosis. To tackle these challenges, we create a BECT dataset containing more than 38,000 real samples from 100 subjects, meticulously annotated by neurology experts. Building on this dataset, we propose a multi-view attention and multi-branch graph convolutional neural network (MAMB) to differentiate abnormal discharges and classify sleep stages in the samples. Experimental results demonstrate that the model benefits from three-dimensional attention mechanisms in the spatial, temporal, and spectral domains, allowing better exploration of intrinsic relationships within EEG signals. Additionally, the sleep and BECT branches enhance the model’s ability to detect abnormal discharges and differentiate various sleep stages. Our model achieves 88.07% and 90.91% accuracy for BECT classification and sleep stage staging, and 81.93% for the four-classification task of simultaneously judging BECT and sleep stage. In addition, the introduction of the multi-view attention mechanism makes the model interpretable and raises hopes of further assisting experts in disease and medication analysis. The implementation code is shown at github1. Yun Ren, Fang Yuan 0010, Yangxin Zhu, Shikui Tu, Yucai Chen, Lei Xu 0001 |
BIBM | 7 |
| 2023 | Multi-source unsupervised domain-adaptation for automatic sleep stagingabstractSleep staging using electroencephalogram (EEG) is of great significance for diagnosing sleep disorders. Recently, due to the high cost of manually annotating EEG signals and the problem of domain shift across different datasets, many unsupervised domain adaptation methods have been applied to sleep staging. However, all these methods are single-source unsupervised domain adaptation (SUDA) methods. When applied to the multi-source unsupervised domain adaptation (MUDA) problem, most SUDA methods treat all source domains equally, without considering the domain shift between them. As a result, achiving feature alignment across all domains becomes a challenge for these SUDA methods. In our proposed model, we address the challenge of domain shift in multi-source unsupervised domain adaptation (MUDA) for automatic sleep staging. Specially, we create a domain-specific branch for each pair of source and target domains, which focuses on aligning the features extracted from the respective domains using an adversarial-based domain adaptation approach. Additionally, we incorporate a domain-invariant branch into our model for learning the domain-invariant representation. We also introduce an adaptive-based mixing strategy to assigns weights to each branches, which can dynamically adjust the importance of each branch depending on their relevance to the specific prediction task. The experiments conducted on three public datasets show the superior performance of our model. Compared to other state-of-the-art MUDA models, our model’s average classification accuracy improves by 2% to 9.9%. The source code is available at https://github.com/CMACH508/MUDAEEG. Yangxin Zhu, Shikui Tu, Lei Xu 0001 |
BIBM | 4 |
| 2023 | DeepTH: Chip Placement with Deep Reinforcement Learning Using a Three-Head Policy NetworkabstractModern very-large-scale integrated (VLSI) circuit placement with huge state space is a critical task for achieving layouts with high performance. Recently, reinforcement learning (RL) algorithms have made a promising breakthrough to dramatically save design time than human effort. However, the previous RL-based works either require a large dataset of chip placements for pre-training or produce illegal final placement solutions. In this paper, DeepTH, a three-head policy gradient placer, is proposed to learn from scratch without the need of pre-training, and generate superior chip floorplans. Graph neural network is initially adopted to extract the features from nodes and nets of chips for estimating the policy and value. To efficiently improve the quality of floorplans, a reconstruction head is employed in the RL network to recover the visual representation of the current placement, by enriching the extracted features of placement embedding. Besides, the reconstruction error is used as a bonus during training to encourage exploration while alleviating the sparse reward problem. Furthermore, the expert knowledge of floorplanning preference is embedded into the decision process to narrow down the potential action space. Experiment results on the ISPD 2005 benchmark have shown that our method achieves 19.02% HPWL improvement than the analytic placer DREAMPlace and 19.89% improvement at least than the state-of-the-art RL algorithms. Dengwei Zhao, Shuai Yuan 0016, Yanan Sun 0003, Shikui Tu, Lei Xu 0001 |
DATE | 5 |
| 2023 | A Deep Temporal Factor Analysis Method for Large Scale Financial Portfolio SelectionabstractExisting machine learning methods are effective in portfolio optimization on a small pool of assets. This is still not optimal because a larger number of assets in markets offers more opportunities for investors. However, existing methods are usually not scalable to large amount of assets which brings new challenges of high dimensionality and computing complexity. In this paper, we present a neural network temporal factor analysis (NN-TFA) model for dimensionality reduction and it enables us to build a scalable deep reinforcement learning method for large-scale portfolio management. Traditional TFA models the relation between asset prices and real economic activities via a small set of independent hidden factors. NN-TFA is developed from the traditional TFA by replacing the linear autoregressive model over the hidden factors with a neural network function, which well captures the complicated temporal patterns. The hidden factors are then sent to a policy network to generate portfolio weights. A calibration module to extract information from other assets features and a ratio module to catch the trend of the selected assets pool are proposed to enhance the performance of the policy network. Extensive tests demonstrate that our methods are capable of handling large-scale datasets and achieving promising results. Ruidan Su, Shikui Tu, Lei Xu 0001 |
ICASSP | 4 |
| 2023 | A Dynamic Graph Convolutional Network for Anti-money Laundering
Tianpeng Wei, Biyang Zeng, Wenqi Guo, Shikui Tu, Lei Xu 0001 |
ICIC (5) | 6 |
| 2023 | Generalizing Graph Network Models for the Traveling Salesman Problem with Lin-Kernighan-Helsgaun Heuristics
Mingfei Li, Shikui Tu, Lei Xu 0001 |
ICONIP (1) | 3 |
| 2023 | GLPocket: A Multi-Scale Representation Learning Approach for Protein Binding Site PredictionabstractProtein binding site prediction is an important prerequisite for the discovery of new drugs. Usually, natural 3D U-Net is adopted as the standard site prediction framework to do per-voxel binary mask classification. However, this scheme only performs feature extraction for single-scale samples, which may bring the loss of global or local information, resulting in incomplete, artifacted or even missed predictions. To tackle this issue, we propose a network called GLPocket, which is based on the Lmser (Least mean square error reconstruction) network and utilizes multi-scale representation to predict binding sites. Firstly, GLPocket uses Target Cropping Block (TCB) for targeted prediction. TCB selects the local interested feature from the global representations to perform concentrated prediction, and reduces the volume of feature maps to be calculated by 82% without adding additional parameters. It integrates global distribution information into local regions, making prediction more concentrated on decoding stage. Secondly, GLPocket establishes long-range relationship of patches within the local region with Transformer Block (TB), to enrich local context semantic information. Experiments show that GLPocket improves by 0.5%-4% on DCA Top-n prediction compared with previous state-of-the-art methods on four datasets. Our code has been released in https://github.com/CMACH508/GLPocket. Peiying Li, Yongchang Liu, Shikui Tu, Lei Xu 0001 |
IJCAI | 4 |
| 2023 | Generalized Weighted Path Consistency for Mastering Atari GamesabstractReinforcement learning with the help of neural-guided search consumes huge computational resources to achieve remarkable performance. Path consistency (PC), i.e., $f$ values on one optimal path should be identical, was previously imposed on MCTS by PCZero to improve the learning efficiency of AlphaZero. Not only PCZero still lacks a theoretical support but also considers merely board games. In this paper, PCZero is generalized into GW-PCZero for real applications with non-zero immediate reward. A weighting mechanism is introduced to reduce the variance caused by scouting's uncertainty on the $f$ value estimation. For the first time, it is theoretically proved that neural-guided MCTS is guaranteed to find the optimal solution under the constraint of PC. Experiments are conducted on the Atari $100$k benchmark with $26$ games and GW-PCZero achieves $198\%$ mean human performance, higher than the state-of-the-art EfficientZero's $194\\%$, while consuming only $25\\%$ of the computational resources consumed by EfficientZero. Dengwei Zhao, Shikui Tu, Lei Xu 0001 |
NeurIPS | 3 |
| 2023 | IA-FaceS: A bidirectional method for semantic face editing
Wenjing Huang 0005, Shikui Tu, Lei Xu 0001 |
Neural Networks | 3 |
| 2023 | RefinePocket: An Attention-Enhanced and Mask-Guided Deep Learning Approach for Protein Binding Site PredictionabstractProtein binding site prediction is an important prerequisite task of drug discovery and design. While binding sites are very small, irregular and varied in shape, making the prediction very challenging. Standard 3D U-Net has been adopted to predict binding sites but got stuck with unsatisfactory prediction results, incomplete, out-of-bounds, or even failed. The reason is that this scheme is less capable of extracting the chemical interactions of the entire region and hardly takes into account the difficulty of segmenting complex shapes. In this paper, we propose a refined U-Net architecture, called RefinePocket, consisting of an attention-enhanced encoder and a mask-guided decoder. During encoding, taking binding site proposal as input, we employ Dual Attention Block (DAB) hierarchically to capture rich global information, exploring residue relationship and chemical correlations in spatial and channel dimensions respectively. Then, based on the enhanced representation extracted by the encoder, we devise Refine Block (RB) in the decoder to enable self-guided refinement of uncertain regions gradually, resulting in more precise segmentation. Experiments show that DAB and RB complement and promote each other, making RefinePocket has an average improvement of 10.02% on DCC and 4.26% on DVO compared with the state-of-the-art method on four test sets. Yongchang Liu, Peiying Li, Shikui Tu, Lei Xu 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 4 |
| 2022 | RecurPocket: Recurrent Lmser Network with Gating Mechanism for Protein Binding Site DetectionabstractIt is an essential step to locate the binding sites or pockets of drug molecules on protein structure in drug design. This is challenging because the 3D protein structures are usually in complicated, irregular shape and the pockets are relatively small. Existing deep learning methods for this task are U-Net models, and they have forward skip connections to efficiently transfer features of different levels of 3D structure from encoder to decoder for improving pocket prediction. However, there is still room to improve prediction accuracy. In this paper, we propose RecurPocket, a recurrent Lmser (Least mean square error reconstruction) network for pocket detection. A gated recurrent refinement is devised in RecurPocket to enhance the representation learning on the 3D protein structures. This is fulfilled by feedback connections in RecurPocket network from decoder to encoder, recurrently and progressively improving the feature embedding for accurate prediction. Moreover, a 3D gate mechanism filters out irrelevant information through the feedback links that interfere with detection, making the prediction more precise and clear. Experiments show that RecurPocket improves by 3%-9% on top-n prediction compared with previous state-of-the-art on five benchmark data sets. The source code and trained model are available at https://github.con CMACH508/RecurPocket. Peiying Li, Boheng Cao, Shikui Tu, Lei Xu 0001 |
BIBM | 4 |
| 2022 | Kernel Mean Matching with Mahalanobis Distance for Causal Inference of Time-to-event OutcomeabstractA number of existing causal effect estimation approaches for time-to-event outcomes in observational studies employ the propensity score-based strategies to balance covariates in treatment groups and reduce confounding effects. The difficulty of causal inference using propensity score lies in the issue that correctly specified propensity score model is required to obtain causal effects, otherwise, substantial bias may be introduced. In this paper, we develop a nonparametric weighting approach, which balances the covariates between treatment groups by Kernel Mean Matching (KMM) with Mahalanobis distance kernels. The proposed estimator adapts KMM procedure to causal effect estimation for time-to-event outcomes, rather than its original applications on empirical risk for classification or regression. Given universal kernels, KMM procedure leads to the balance of covariate distributions. For time-to-event outcome which has highly nonlinear relations with covariates, balancing covariates at distribution level is necessary rather than finite order moments. Furthermore, in consideration of practical applications, unlike the Euclidean distance, the Mahalanobis distance is essentially suitable for the anisotropic real-world data with extremely different scales and highly correlated dimensions, so that we can make the best use of the optimization power. We theoretically proved the consistency of the proposed estimator. Simulations and real-world applications illustrate the effectiveness of the proposed method. Shikui Tu, Lei Xu 0001 |
BIBM | 4 |
| 2022 | VentSR: A Self-Rectifying Deep Learning Method for Extubation Readiness PredictionabstractTimely recognition of extubation readiness is critical, because prolonged and premature intubation will lead to sever complications and costs. Clinical assessment is time consuming and challenging and it has attracted increasing attention of machine learning in recent years. However, the data used for extubation predictions have the following flaws: 1) Manual recording errors and missing data; 2) Unreliable ventilation labels due to inadequate judgement from clinicians. Both may possibly lead to wrong ventilation labels, but existing machine learning methods for extubation prediction largely ignored this critical issue. In this paper, we proposed a self-rectifying deep learning method for extubation readiness prediction, called VentSR. It improves the prediction performance by a self-rectifying strategy, and the rectification is achieved through model training without clinical experience. To be detailed, VentSR firstly identifies possibly wrong samples by two components: Inconsistency between K-means and Labels (IKL) and Inconsistency between Model Predictions and Labels (IPL). IKL partitions a rough subset, and IPL iteratively refines this subset through training. Additionally, we designed Adjustment Operation to enhance IPL ability for refinement. Samples identified in this subset are rectified and used to train the model. The unrectified test set is directly fed into the trained model to obtain prediction results. Experiments demonstrate that VentSR outperforms other baselines. Further comparisons on high-confidence test set indicate that VentSR achieves 79.4 AUPRC, increasing by 26.0%. Feature importance analysis and case study illustration again reveals that VentSR are of potential practical usage of informing clinicians with accurate extubation readiness. Long Xiang 0002, Shikui Tu, Liebin Zhao, Lei Xu 0001 |
BIBM | 7 |
| 2022 | Learning to Generate Textual Adversarial Examples
Xiangzhe Guo, Shikui Tu, Lei Xu 0001 |
ICANN (1) | 3 |
| 2022 | Efficient Learning for AlphaZero via Path ConsistencyabstractIn recent years, deep reinforcement learning have made great breakthroughs on board games. Still, most of the works require huge computational resources for a large scale of environmental interactions or self-play for the games. This paper aims at building powerful models under a limited amount of self-plays which can be utilized by a human throughout the lifetime. We proposes a learning algorithm built on AlphaZero, with its path searching regularised by a path consistency (PC) optimality, i.e., values on one optimal search path should be identical. Thus, the algorithm is shortly named PCZero. In implementation, historical trajectory and scouted search paths by MCTS makes a good balance between exploration and exploitation, which enhances the generalization ability effectively. PCZero obtains $94.1%$ winning rate against the champion of Hex Computer Olympiad in 2015 on $13\times 13$ Hex, much higher than $84.3%$ by AlphaZero. The models consume only $900K$ self-play games, about the amount humans can study in a lifetime. The improvements by PCZero have been also generalized to Othello and Gomoku. Experiments also demonstrate the efficiency of PCZero under offline learning setting. Dengwei Zhao, Shikui Tu, Lei Xu 0001 |
ICML | 3 |
| 2022 | Searching for Textual Adversarial Examples with Learned Strategy
Xiangzhe Guo, Ruidan Su, Shikui Tu, Lei Xu 0001 |
ICONIP (4) | 4 |
| 2022 | IA-CL: A Deep Bidirectional Competitive Learning Method for Traveling Salesman Problem
Shikui Tu, Lei Xu 0001 |
ICONIP (1) | 3 |
| 2022 | Box-FaceS: A Bidirectional Method for Box-Guided Face Component EditingabstractWhile the quality of face manipulation has been improved tremendously, the ability to control face components, e.g., eyebrows, is still limited. Although existing methods have realized component editing with user-provided geometry guidance, such as masks or sketches, their performance is largely dependent on the user's painting efforts. To address these issues, we propose Box-FaceS, a bidirectional method that can edit face components by simply translating and zooming the bounding boxes. This framework learns representations for every face component, independently, as well as a high-dimensional tensor capturing face outlines. To enable box-guided face editing, we develop a novel Box Adaptive Modulation (BAM) module for the generator, which first transforms component embeddings to style parameters and then modulates visual features inside a given box-like region on the face outlines. A cooperative learning scheme is proposed to impose independence between face outlines and component embeddings. As a result, it is flexible to determine the component style by its embedding, and to control its position and size by the provided bounding box. Box-FaceS also learns to transfer components between two faces while maintaining the consistency of image content. In particular, Box-FaceS can generate creative faces with reasonable exaggerations, requiring neither supervision nor complex spatial morphing operations. Through the comparisons with state-of-the-art methods, Box-FaceS shows its superiority in component editing, both qualitatively and quantitatively. To the best of our knowledge, Box-FaceS is the first approach that can freely edit the position and shape of the face components without editing the face masks or sketches. Our implementation is available at https://github.com/CMACH508/Box-FaceS. Wenjing Huang 0005, Shikui Tu, Lei Xu 0001 |
ACM Multimedia | 3 |
| 2022 | Predicting cell line-specific synergistic drug combinations through a relational graph convolutional network with attention mechanismabstractIdentifying synergistic drug combinations (SDCs) is a great challenge due to the combinatorial complexity and the fact that SDC is cell line specific. The existing computational methods either did not consider the cell line specificity of SDC, or did not perform well by building model for each cell line independently. In this paper, we present a novel encoder-decoder network named SDCNet for predicting cell line-specific SDCs. SDCNet learns common patterns across different cell lines as well as cell line-specific features in one model for drug combinations. This is realized by considering the SDC graphs of different cell lines as a relational graph, and constructing a relational graph convolutional network (R-GCN) as the encoder to learn and fuse the deep representations of drugs for different cell lines. An attention mechanism is devised to integrate the drug features from different layers of the R-GCN according to their relative importance so that representation learning is further enhanced. The common patterns are exploited through partial parameter sharing in cell line-specific decoders, which not only reconstruct the known SDCs but also predict new ones for each cell line. Experiments on various datasets demonstrate that SDCNet is superior to state-of-the-art methods and is also robust when generalized to new cell lines that are different from the training ones. Finally, the case study again confirms the effectiveness of our method in predicting novel reliable cell line-specific SDCs. Peng Zhang 0098, Shikui Tu, Wen Zhang 0008, Lei Xu 0001 |
Briefings Bioinform. | 4 |
| 2022 | Deep Rival Penalized Competitive Learning for low-resolution face recognition
Peiying Li, Shikui Tu, Lei Xu 0001 |
Neural Networks | 3 |
| 2022 | Deep CNN Based Lmser and Strengths of Two Built-In Dualities
Wenjing Huang 0005, Shikui Tu, Lei Xu 0001 |
Neural Process. Lett. | 3 |
| 2021 | DeepTrader: A Deep Reinforcement Learning Approach for Risk-Return Balanced Portfolio Management with Market Conditions EmbeddingabstractMost existing reinforcement learning (RL)-based portfolio management models do not take into account the market conditions, which limits their performance in risk-return balancing. In this paper, we propose DeepTrader, a deep RL method to optimize the investment policy. In particular, to tackle the risk-return balancing problem, our model embeds macro market conditions as an indicator to dynamically adjust the proportion between long and short funds, to lower the risk of market fluctuations, with the negative maximum drawdown as the reward function. Additionally, the model involves a unit to evaluate individual assets, which learns dynamic patterns from historical data with the price rising rate as the reward function. Both temporal and spatial dependencies between assets are captured hierarchically by a specific type of graph structure. Particularly, we find that the estimated causal structure best captures the interrelationships between assets, compared to industry classification and correlation. The two units are complementary and integrated to generate a suitable portfolio which fits the market trend well and strikes a balance between return and risk effectively. Experiments on three well-known stock indexes demonstrate the superiority of DeepTrader in terms of risk-gain criteria. Biwei Huang, Shikui Tu, Kun Zhang 0001, Lei Xu 0001 |
AAAI | 5 |
| 2021 | IA-GM: A Deep Bidirectional Learning Method for Graph MatchingabstractExisting deep learning methods for graph matching(GM) problems usually considered affinity learningto assist combinatorial optimization in a feedforward pipeline, and parameter learning is executed by back-propagating the gradients of the matching loss. Such a pipeline pays little attention to the possible complementary benefit from the optimization layer to the learning component. In this paper, we overcome the above limitation under a deep bidirectional learning framework.Our method circulates the output of the GM optimization layer to fuse with the input for affinity learning. Such direct feedback enhances the input by a feature enrichment and fusion technique, which exploits andintegrates the global matching patterns from the deviation of the similarity permuted by the current matching estimate. As a result, the circulation enables the learning component to benefit from the optimization process, taking advantage of both global feature and the embedding result which is calculated by local propagationthrough node-neighbors. Moreover, circulation consistency induces an unsupervised loss that can be implemented individually or jointly to regularize the supervised loss. Experiments on challenging datasets demonstrate the effectiveness of our methods for both supervised learning and unsupervised learning. Kaixuan Zhao, Shikui Tu, Lei Xu 0001 |
AAAI | 3 |
| 2021 | Flexible-CLmser: Regularized Feedback Connections for Biomedical Image SegmentationabstractThe skip connections in U-Net pass features from the levels of encoder to the ones of decoder in a symmetrical way, which makes U-Net and its variants become state-of-the-art approaches for biomedical image segmentation. However, the U-Net skip connections are unidirectional without considering feedback from the decoder, while this paper exploits the feedback information to refine the segmentation. We develop a deep bidirectional network based on the least mean square error reconstruction (Lmser) self-organizing network, an early model that folds an autoencoder along the central hidden layer such that the neurons on the paired layers between encoder and decoder merge into one, equivalently forming bidirectional skip connections between encoder and decoder. We find that the feedback links indeed increase the segmentation accuracy, but may also bring certain noise into the segmentation. To tackle this problem, we present a gating and masking mechanism on the feedback connections to filter the irrelevant information. Experimental results on MoNuSeg, TNBC, and EM membrane datasets demonstrate that our method are robust and outperforms state-of-the-art methods. Boheng Cao, Shikui Tu, Lei Xu 0001 |
BIBM | 3 |
| 2021 | Multi-source unsupervised domain adaptation for ECG classificationabstractIt is challenging to build a machine learning model for automatic arrhythmia diagnosis from Electrocardiograph (ECG) signals, because the variation in ECG signals is big between different patients or over time, and the available training datasets usually contain limited, unbalanced number of data for multiple disease types. Most existing methods relied on labeled data from a single dataset, and the performance is poor when generalizing to unseen heart disease types, limited labels, or distribution shifts. In this paper, we propose a multi-source unsupervised domain adaption (MUDA) neural network for ECG classification, to make effective use of data of multiple sources and improve the model’s generalization ability. Our model is featured by a two-branch domain adaption and a sample-imbalance aware mixing strategy to fuse the information across domains. Specifically, one branch is devised to learn domain-invariant representation, while the other is to extract domain-specific features. The two branches align the ECG in the target domain to individual source domain in an exclusive and complementary manner, leading to enhanced discriminative features for domain invariant/specific classifiers. The final prediction, which is a linear combination of the domain classification decisions, is very robust and accurate, by making use of the prior distribution of sample size across domains to place confidence scores over each classifier. Experiments on five ECG datasets indicate superior performance of our method over the existing ones. Fucheng Deng, Shikui Tu, Lei Xu 0001 |
BIBM | 3 |
| 2021 | Enriching computed tomography images by projection for robust automated cerebral aneurysm detection and segmentationabstractDeveloping an efficient system for automated detection and segmentation of intracranial aneurysms (IAs) became an active research topic recently. However, existing methods are poor in detecting small IA with high false positives. In this paper, we present a feature enrichment (FE) based deep learning method for robust IA detection and segmentation. The FE technique is featured by reconstructing a 3D model from all computed tomography angiography (CTA) images and then projecting 3D information into the so-called projection images of different slicing levels along various directions. The appearances of aneurysms in the projection images are enhanced in morphology and 3D neighborhood features, and thus are easy to be detected and segmented by the widely-used faster RCNN and V-Net. To evaluate our method, we collect CTA images from 145 patients (including 148 IAs) for training and testing. The proposed method achieves 96.0 % sensitivity for all aneurysms, and 80.0% sensitivity for aneurysms smaller than 4mm, better than the state-of-the-art methods. Also, the segmentation performance is improved on the detected IAs. Shikui Tu, Peiying Li, Jiafeng Zhou, Jieqing Wan, Lei Xu 0001 |
BIBM | 7 |
| 2021 | RecSleepNet: An Automatic Sleep Staging Model Based on Feature ReconstructionabstractS1eep staging via electroencephalogram (EEG) is the fundamental step to sleep quality assessment and disease diagnose. Deep learning methods have been demonstrated to be promising for automatic sleep staging, but the performance is still not satisfied, because learning representations over the EEG signals is very challenging. In this paper, we propose RecSleepNet, an automatic sleep staging model based on a hybrid structure of Convolutional Neural Networks (CNN) and Long Short Term Memory (LSTM) networks. To enhance the representation learning, we devise a Convolutional Reconstruction Block (CRB), which enforces the extracted features to reconstruct the previous low-level input well. Moreover, we introduce a Global Max and Average Pooling Block (GMAPB) to squeeze parameter numbers and extract shift-invariant features. Experiments on four public datasets of single-channel EEG signals indicate that RecSleepNet achieves better or at least comparable performance to the state-of-the-art methods. CRB and GMAPB enable the training to be more efficient with fewer parameters b ut faster convergence. Haodong Nie, Shikui Tu, Lei Xu 0001 |
BIBM | 3 |
| 2021 | A Consistency Enhanced Deep Lmser Network for Face Sketch Synthesis
Qingjie Sheng, Shikui Tu, Lei Xu 0001 |
PRICAI (1) | 3 |
| 2021 | Controllable stroke-based sketch synthesis from a self-organized latent space
Sicong Zang, Shikui Tu, Lei Xu 0001 |
Neural Networks | 3 |
| 2021 | Detection of Phenotype-Related Mutations of COVID-19 via the Whole Genomic DataabstractThe coronavirus disease 2019 (COVID-19) epidemic continues to spread rapidly around the world and nearly 20 millions people are infected. This paper utilises both single-locus analysis and joint-SNPs analysis for detection of significant single nucleotide polymorphisms (SNPs) in the phenotypes of symptomatic versus asymptomatic, the early collection time versus the late collection time, the old versus the young, and the male versus the female. Also, this paper analyses the relationship between any two SNPs via linkage disequilibrium analysis, and visualises the patterns of cumulative mutations of SNPs over collection time. The results are in three folds. First, the SNP which locates at the nucleotide position 4321 is found to be an independent significant locus associated with all the first three phenotypes. Moreover, 12 significant SNPs are found in the first two studies. Second, gene orf1ab containing SNP-4321 is detected to be significantly associated with the first three phenotypes, and the three genes S, ORF3a, and N, are detected to be significant in the first two phenotypes. Third, some of the detected genes or SNPs are related to the SARS-COV-2 as supported by literature survey, which indicates that the results here may be helpful for further investigation. Jin-Xiong Lv, Shikui Tu, Lei Xu 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2020 | Discrete Biorthogonal Wavelet Transform Based Convolutional Neural Network for Atrial Fibrillation Diagnosis from ElectrocardiogramabstractFor the problem of early detection of atrial fibrillation (AF) from electrocardiogram (ECG), it is difficult to capture subject-invariant discriminative features from ECG signals, due to the high variation in ECG morphology across subjects and the noise in ECG. In this paper, we propose an Discrete Biorthogonal Wavelet Transform (DBWT) Based Convolutional Neural Network (CNN) for AF detection, shortly called DBWT-AFNet. In DBWT-AFNet, rather than directly feeding ECG into CNN, DBWT is used to separate sub-signals in frequency band of heart beat from ECG, whose output is fed to CNN for AF diagnosis. Such sub-signals are better than the raw ECG for subject-invariant CNN representation learning because noisy information irrelevant to human beat has been largely filtered out. To strengthen the generalization ability of CNN to discover subject-invariant pattern in ECG, skip connection is exploited to propagate information well in neural network and channel attention is designed to adaptively highlight informative channel-wise features. Experiments show that the proposed DBWT-AFNet outperforms the state-of- the-art methods, especially for ECG segments classification across different subjects, where no data from testing subjects have been used in training. Qingsong Xie, Shikui Tu, Guoxing Wang, Yong Lian 0001, Lei Xu 0001 |
IJCAI | 5 |
| 2020 | Learning deep IA bidirectional intelligenceabstractThere has been a framework sketched for learning deep bidirectional intelligence. The framework has an inbound that features two actions: one is the acquiring action, which gets inputs in appropriate patterns, and the other is A-S cognition, derived from the abbreviated form of words abstraction and self-organization , which abstracts input patterns into concepts that are labeled and understood by self-organizing parts involved in the concept into structural hierarchies. The top inner domain accommodates relations and a priori knowledge with the help of the A-I thinking action that is responsible for the accumulation-amalgamation and induction-inspiration. The framework also has an outbound that comes with two actions. One is called I-S reasoning , which makes inference and synthesis (I-S) and is responsible for performing various tasks including image thinking and problem solving, and the other is called the interacting action, which controls, communicates with, and inspects the environment. Based on this framework, we further discuss the possibilities of design intelligence through synthesis reasoning. Lei Xu 0001 |
Frontiers Inf. Technol. Electron. Eng. | 1 |
| 2019 | A two-variate phenotype-targeted test for detection of phenotypic biomarkers on breast cancerabstractTraditional pipeline for the task of detecting phenotypic biomarkers is a two-stage implementation, i.e., differentially expressed candidates are identified by NT tests, and then a subset of the candidates are further detected by phenotype-targeted tests (PT test) for significant phenotypic features, where N is short for Normal data and T is for Treatment/Trouble data. Such a two-stage procedure has low detection power as they do not make full use of the information contained in the (T, N). In this paper, we apply the two-variate PT test which jointly considers tumor-adjacent data and tumor data for improving the detection power. We investigate its performance by experiments on real-world datasets of breast cancer, considering phenotypes including BMI, overall survival time, pathologic stage, and tumor size. The results show that the method has high detection power and is more reliable, and the tumor-adjacent normal data plays an important role in the detection of phenotypic biomarkers. Finally, we obtain a new finding that the gene TCTEX1D2 is significantly related to tumor size in breast cancer. Jin-Xiong Lv, Shikui Tu, Lei Xu 0001 |
BIBM | 3 |
| 2019 | GAN Flexible Lmser for Super-resolutionabstractExisting single image super-resolution (SISR) methods usually focus on Low-Resolution (LR) images which are artificially generated from High-Resolution (HR) images by a down-sampling process, but are not robust for unmatched training set and testing set. This paper proposes a GAN Flexible Lmser (GFLmser) network that bidirectionally learns the High-to-Low (H2L) process that degrades HR images to LR images and the Low-to-High (L2H) process that recovers the LR images back to HR images. The two directions share the same architecture, added with the gated skip connections from the H2L-net to the L2H-net in order to enhance information transferring for super-resolution. In comparison with several related state-of-the-art methods, experiments demonstrate that not only GFLmser is the most robust method on images of unmatched training set and testing set, but also its performance on real-world face LR images is best in PSNR and reasonably good in FID. Peiying Li, Shikui Tu, Lei Xu 0001 |
ACM Multimedia | 3 |
| 2017 | A Comparative Study on Lagrange Ying-Yang Alternation Method in Gaussian Mixture-Based Clustering
Weijian Long, Shikui Tu, Lei Xu 0001 |
IDEAL | 3 |
| 2017 | Bootstrapping integrative hypothesis test for identifying biomarkers that differentiates lung cancer and chronic obstructive pulmonary disease
Kai-Ming Jiang, Ya-Jing Chen, Jin-Xiong Lv, Bao-Liang Lu, Lei Xu 0001 |
Neurocomputing | 5 |
| 2016 | A comparison study on multivariate methods for joint-SNVs association analysisabstractSingle nucleotide variants (SNVs) have been discovered that they play crucial roles in disease pathogenesis as genetic factors. Featured by analyzing multiple SNVs in a biological module (e.g. exon, gene, etc.) collectively, the joint-SNVs studies are increasingly attractive in genome-wide association studies (GWASs), for which extensive efforts have been devoted to pursue effective multivariate methods. In this paper, we first reviewed several main streams of existing methods and their limitations in joint-SNVs studies. Then, we introduced a recently proposed novel method, namely statistic-space boundary based test (S-space BBT) to tackle these limitations. Via computational experiments on simulation datasets, not only we figured out the applicable scenarios for the six methods in considering the effect direction and whether the single significant is involved in, but also demonstrated the strong detecting sensitivity of S-space BBT under the different conditions of odds ratio, minor allele frequency, and the linkage disequilibrium. We anticipate that our study may provide clues for multivariate method selection, and that S-space BBT may play a promising role in the joint-SNVs analysis. Jin-Xiong Lv, Han-Chen Huang, Runsheng Chen, Lei Xu 0001 |
BIBM | 4 |
| 2014 | A graph matching algorithm based on concavely regularized convex relaxation
Zhiyong Liu 0001, Hong Qiao, Li-Hao Jia, Lei Xu 0001 |
Neurocomputing | 4 |
| 2014 | A comparative study of RPCL and MCE based discriminative training methods for LVCSR
Zaihu Pang, Shikui Tu, Xihong Wu, Lei Xu 0001 |
Neurocomputing | 4 |
| 2014 | Learning local factor analysis versus mixture of factor analyzers with automatic model selection
Lei Shi 0016, Shikui Tu, Lei Xu 0001 |
Neurocomputing | 4 |
| 2014 | Learning binary factor analysis with automatic model selection
Shikui Tu, Lei Xu 0001 |
Neurocomputing | 2 |
| 2012 | A non-Gaussian factor analysis approach to transcription Network Component AnalysisabstractTranscription factor activities (TFAs), rather than expression levels, control gene expression and provide valuable information for investigating TF-gene regulations. Network Component Analysis (NCA) is a model based method to deduce TFAs and TF-gene control strengths from microarray data and a priori TF-gene connectivity data. We modify NCA to model gene expression regulation by non-Gaussian Factor Analysis (NFA), which assumes TFAs independently comes from Gaussian mixture densities. We properly incorporate a priori connectivity and/or sparsity on the mixing matrix of NFA, and derive, under Bayesian Ying-Yang (BYY) learning framework, a BYY-NFA algorithm that can not only uncover the latent TFA profile similar to NCA, but also is capable of automatically shutting off unnecessary connections. Simulation study demonstrates the effectiveness of BYY-NFA, and a preliminary application to two real world data sets shows that BYY-NFA improves NCA for the case when TF-gene connectivity is not available or not reliable, and may provide a preliminary set of candidate TF-gene interactions or double check unreliable connections for experimental verification. Shikui Tu, Dingsheng Luo, Runsheng Chen, Lei Xu 0001 |
CIBCB | 4 |
| 2012 | An Extended Path Following Algorithm for Graph-Matching ProblemabstractThe path following algorithm was proposed recently to approximately solve the matching problems on undirected graph models and exhibited a state-of-the-art performance on matching accuracy. In this paper, we extend the path following algorithm to the matching problems on directed graph models by proposing a concave relaxation for the problem. Based on the concave and convex relaxations, a series of objective functions are constructed, and the Frank-Wolfe algorithm is then utilized to minimize them. Several experiments on synthetic and real data witness the validity of the extended path following algorithm. Zhiyong Liu 0001, Hong Qiao, Lei Xu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2012 | A theoretical investigation of several model selection criteria for dimensionality reduction
Shikui Tu, Lei Xu 0001 |
Pattern Recognit. Lett. | 2 |
| 2010 | Gene clustering by structural prior based local factor analysis model under Bayesian Ying-Yang harmony learningabstractWe propose a clustering algorithm based on a structural prior based Local Factor Analysis (spLFA) model under the Bayesian Ying-Yang harmony learning, which automatically determines the hidden dimensionalities during parameter learning, reduces the number of free parameters by projecting the mean vectors onto a low dimensional manifold, imposes the sparseness by a Normal-Jeffreys prior. Experiments on the diagnostic research dataset show that BYY-spLFA outperforms the k-means clustering and single-link hierarchical clustering. The experiments on a lymphoma cancer datset further indicate the BYY-spLFA is able to uncover the number of phenotypes correctly and cluster the phenotypes more accurately. In addition, we modify BYY-spLFA to implement supervised learning and preliminarily demonstrate its effectiveness on a Leukemia data for classification. Lei Shi 0016, Shikui Tu, Lei Xu 0001 |
BIBM | 3 |
| 2010 | Radar HRRP statistical recognition with Local Factor Analysis by automatic Bayesian Ying Yang harmony learningabstractRadar high-resolution range profiles (HRRPs) are typical high-dimensional non-Gaussian and inter-dimensional dependently distributed data, the statistical modelling of which is a challenging task for HRRP based target recognition. Considering the inter-dimensional dependence, a recent work applied Factor Analysis (FA) to model radar HRRP data and showed promising recognition results, which however still restricts to Gaussian distribution. This paper aims to simultaneously consider the inter-dimensional dependence and the non-Gaussian distribution, by using Local Factor Analysis (LFA) model. For not only learning parameters but also appropriately selecting the component number and local hidden dimensionalities, we adopt the automatic Bayesian Ying-Yang (BYY) harmony learning, in order to relieve the extensive computation and inaccurate evaluation encountered in the conventional two-phase implementation. Moreover, a heuristic aspect-frame partition is implemented based on the BYY harmony criterion rather than AIC or BIC in the previous work, to tackle the radar HRRP's target-aspect sensitivity. Experiments show improved recognition performances over on the same measured HRRP dataset, i.e., for both equal interval and heuristic aspect-frame partitions, LFA automatically learned by BYY always outperforms FA selected by a two-phase procedure with either AIC or BIC. Lei Shi 0016, Hongwei Liu 0001, Lei Xu 0001, Zheng Bao 0001 |
ICASSP | 4 |
| 2010 | GMM-HMM acoustic model training by a two level procedure with Gaussian components determined by automatic model selectionabstractThis paper investigates the Bayesian Ying-Yang (BYY) learning for speech recognition via Gaussian mixture models (GMMs) based Hidden Markov models (HMMs). A two level procedure is proposed with the hidden Markov level trained still under the maximum likelihood principle by the Baum-Welch algorithm but with the GMMs level trained under the BYY best harmony. We proposed a new batch way EM-like Ying-Yang alternation algorithm and used it as a plug-in block to the Baum-Welch algorithm. The advantage is that number of GMM components can be automatically determined during this BYY harmony learning and that the resulted model parameters become less affected than EM-ML training by the problem of overfitting and singular solution. In comparison with the standard EM-ML training and classical model selection criterions, including BIC and AIC, speech recognition experiments in a large vocabulary task on the Hub4 broadcast news database shown that the proposed algorithm provides an improved performance and also good convergence. Xihong Wu, Lei Xu 0001 |
ICASSP | 3 |
| 2010 | A study of several model selection criteria for determining the number of signalsabstractAddressing the problem of detecting the number of source signals as selecting the hidden dimensionality of Factor Analysis (FA) model, we investigate several model selection criteria via a new empirical analyzing tool that examines the joint effect of signal-noise ratio (SNR) and sample size N on the model selection performance. The contours of the model selection accuracies visualize a three-region partition on the space of SNR andN, and a diminishing marginal effect which trades off SNR and N on the performance. Moreover, the newly derived Variational Bayes algorithm and three variants of Bayesian Ying-Yang (BYY) algorithms are more robust against reducing SNR and N, where the BYY with priors' hyperparameters updated is the best in general. Shikui Tu, Lei Xu 0001 |
ICASSP | 2 |
| 2010 | Machine learning problems from optimization perspective
Lei Xu 0001 |
J. Glob. Optim. | 1 |
| 2008 | Bayesian Ying-Yang Learning on Orthogonal Binary Factor Analysis
Ke Sun 0001, Lei Xu 0001 |
ICANN (1) | 2 |
| 2008 | A Comparative Study on Data Smoothing Regularization for Local Factor Analysis
Shikui Tu, Lei Shi 0016, Lei Xu 0001 |
ICANN (1) | 3 |
| 2007 | Investigation on Multisets Mixture Learning Based Object DetectionabstractBy minimizing the mean square reconstruction error, multisets mixture learning (MML) provides a general approach for object detection in image. To calculate each sample reconstruction error, as the object template is represented by a set of contour points, the MML needs to inefficiently enumerate the distances between the sample and all the contour points. In this paper, we develop the line segment approximation (LSA) algorithm to calculate the reconstruction error, which is shown theoretically and experimentally to be more efficient than the enumeration method. It is also experimentally illustrated that the MML based algorithm has a better noise resistance ability than the generalized Hough transform (GHT) based counterpart. Zhiyong Liu 0001, Hong Qiao, Lei Xu 0001 |
Int. J. Pattern Recognit. Artif. Intell. | 3 |
| 2007 | One-Bit-Matching Theorem for ICA, Convex-Concave Programming on Polyhedral Set, and Distribution Approximation for CombinatoricsabstractAccording to the proof by Liu, Chiu, and Xu (2004) on the so-called one-bit-matching conjecture (Xu, Cheung, and Amari, 1998a), all the sources can be separated as long as there is an one-to-one same-sign correspondence between the kurtosis signs of all source probability density functions (pdf's) and the kurtosis signs of all model pdf's, which is widely believed and implicitly supported by many empirical studies. However, this proof is made only in a weak sense that the conjecture is true when the global optimal solution of an independent component analysis criterion is reached. Thus, it cannot support the successes of many existing iterative algorithms that usually converge at one of the local optimal solutions. This article presents a new mathematical proof that is obtained in a strong sense that the conjecture is also true when any one of local optimal solutions is reached in helping to investigating convex-concave programming on a polyhedral set. Theorems are also provided not only on partial separation of sources when there is a partial matching between the kurtosis signs, but also on an interesting duality of maximization and minimization on source separation. Moreover, corollaries are obtained on an interesting duality, with supergaussian sources separated by maximization and subgaussian sources separated by minimization. Also, a corollary is obtained to confirm the symmetric orthogonalization implementation of the kurtosis extreme approach for separating multiple sources in parallel, which works empirically but lacks mathematical proof. Furthermore, a linkage has been set up to combinatorial optimization from a distribution approximation perspective and a Stiefel manifold perspective, with algorithms that guarantee convergence as well as satisfaction of constraints. Lei Xu 0001 |
Neural Comput. | 1 |
| 2007 | A unified perspective and new results on RHT computing, mixture based learning, and multi-learner based problem solving
Lei Xu 0001 |
Pattern Recognit. | 1 |
| 2006 | Comparative Investigation on Dimension Reduction and Regression in Three Layer Feed-Forward Neural Network
Lei Shi 0016, Lei Xu 0001 |
ICANN (1) | 2 |
| 2006 | Local Factor Analysis with Automatic Model Selection: A Comparative Study and Digits Recognition Application
Lei Shi 0016, Lei Xu 0001 |
ICANN (2) | 2 |
| 2006 | A Comparative Study on Selection of Cluster Number and Local Subspace Dimension in the Mixture PCA Models
Xuelei Hu, Lei Xu 0001 |
ISNN (1) | 2 |
| 2006 | Multisets mixture learning-based ellipse detection
Zhiyong Liu 0001, Hong Qiao, Lei Xu 0001 |
Pattern Recognit. | 3 |
| 2005 | One-Bit-Matching ICA Theorem, Convex-Concave Programming, and Combinatorial Optimization
Lei Xu 0001 |
ISNN (1) | 1 |
| 2005 | Asymptotic convergence properties of the EM algorithm with respect to the overlap in the mixture
Jinwen Ma, Lei Xu 0001 |
Neurocomputing | 2 |
| 2005 | A Further Result on the ICA One-Bit-Matching ConjectureabstractThe one-bit-matching conjecture for independent component analysis (ICA) has been widely believed in the ICA community. Theoretically, it has been proved that under the assumption of zero skewness for the model probability density functions, the global maximum of a cost function derived from the typical objective function on the ICA problem with the one-bit-matching condition corresponds to a feasible solution of the ICA problem. In this note, we further prove that all the local maximums of the cost function correspond to the feasible solutions of the ICA problem in the two-source case under the same assumption. That is, as long as the one-bit-matching condition is satisfied, the two-source ICA problem can be successfully solved using any local descent algorithm of the typical objective function with the assumption of zero skewness for all the model probability density functions. Jinwen Ma, Lei Xu 0001 |
Neural Comput. | 3 |
| 2004 | Local Separation Property of the Two-Source ICA Problem with the One-Bit-Matching Condition
Jinwen Ma, Lei Xu 0001 |
IDEAL | 3 |
| 2004 | Automatic Cluster Number Determination via BYY Harmony Learning
Xuelei Hu, Lei Xu 0001 |
ISNN (1) | 2 |
| 2004 | Arbitrage pricing theory-based Gaussian temporal factor analysis for adaptive portfolio management
Kai Chun Chiu, Lei Xu 0001 |
Decis. Support Syst. | 2 |
| 2004 | A gradient BYY harmony learning rule on Gaussian mixture with automated model selection
Jinwen Ma, Taijun Wang, Lei Xu 0001 |
Neurocomputing | 3 |
| 2004 | One-Bit-Matching Conjecture for Independent Component AnalysisabstractThe one-bit-matching conjecture for independent component analysis (ICA) could be understood from different perspectives but is basically stated as "all the sources can be separated as long as there is a one-to-one same-sign-correspondence between the kurtosis signs of all source probability density functions (pdf's) and the kurtosis signs of all model pdf's" (Xu, Cheung, & Amari, 1998a). This conjecture has been widely believed in the ICA community and implicitly supported by many ICA studies, such as the Extended Infomax (Lee, Girolami, & Sejnowski, 1999) and the soft switching algorithm (Welling & Weber, 2001). However, there is no mathematical proof to confirm the conjecture theoretically. In this article, only skewness and kurtosis are considered, and such a mathematical proof is given under the assumption that the skewness of the model densities vanishes. Moreover, empirical experiments are demonstrated on the robustness of the conjecture as the vanishing skewness assumption breaks. As a by-product, we also show that the kurtosis maximization criterion (Moreau & Macchi, 1996) is actually a special case of the minimum mutual information criterion for ICA. Kai Chun Chiu, Lei Xu 0001 |
Neural Comput. | 3 |
| 2004 | A comparative investigation on subspace dimension determination
Xuelei Hu, Lei Xu 0001 |
Neural Networks | 2 |
| 2004 | Investigations on non-Gaussian factor analysisabstractThis letter further explores the Bayesian Ying-Yang learning based non-Gaussian factor analysis (NFA) via investigating its key yet analytically intractable factor estimating step. Among the three suggested numerical approaches we empirically show that the so-called iterative fixed posteriori approximation approach is the most optimal, as well as theoretically prove that the iterative fixed posteriori approximation is another type of EM-algorithm, with the proof of its convergence also shown. Kai Chun Chiu, Lei Xu 0001 |
IEEE Signal Process. Lett. | 3 |
| 2004 | Guest Editorial Special Issue on Information Theoretic Learning
José C. Príncipe, Erkki Oja, Lei Xu 0001, Andrzej Cichocki, Deniz Erdogmus |
IEEE Trans. Neural Networks | 3 |
| 2004 | Advances on BYY harmony learning: information theoretic perspective, generalized projection geometry, and independent factor autodeterminationabstractThe nature of Bayesian Ying-Yang harmony learning is reexamined from an information theoretic perspective. Not only its ability for model selection and regularization is explained with new insights, but also discussions are made on its relations and differences from the studies of minimum description length (MDL), Bayesian approach, the bit-back based MDL, Akaike information criterion (AIC), maximum likelihood, information geometry, Helmholtz machines, and variational approximation. Moreover, a generalized projection geometry is introduced for further understanding such a new mechanism. Furthermore, new algorithms are also developed for implementing Gaussian factor analysis (FA) and non-Gaussian factor analysis (NFA) such that selecting appropriate factors is automatically made during parameter learning. Lei Xu 0001 |
IEEE Trans. Neural Networks | 1 |
| 2004 | Temporal BYY encoding, Markovian state spaces, and space dimension determinationabstractAs a complementary to those temporal coding approaches of the current major stream, this paper aims at the Markovian state space temporal models from the perspective of the temporal Bayesian Ying-Yang (BYY) learning with both new insights and new results on not only the discrete state featured Hidden Markov model and extensions but also the continuous state featured linear state spaces and extensions, especially with a new learning mechanism that makes selection of the state number or the dimension of state space either automatically during adaptive learning or subsequently after learning via model selection criteria obtained from this mechanism. Experiments are demonstrated to show how the proposed approach works. Lei Xu 0001 |
IEEE Trans. Neural Networks | 1 |
| 2003 | On generalized arbitrage pricing theory analysis: empirical investigation of the macroeconomics modulated independent state-space modelabstractThe inception of Markowitz's (1952) modern portfolio theory has also fuelled the development of asset pricing models for empirical finance, ranging from linear single-factor models like the capital asset pricing model to fairly complex multi-factor models such as the arbitrage pricing theory (APT). It is well-known in the literature of finance that APT could be used for modelling the underlying security returns generation process. We investigate a generalized version of the APT model, called the macroeconomics modulated independent state-space model, in terms of model specification adequacy as well as its performance on prediction. Empirical results reveal that the model is not only well-specified, but also superior to the temporal factor analysis model in stock price and index forecasting, thanks to its salient capabilities of modelling both short-term and long-term market dynamics. Kai Chun Chiu, Lei Xu 0001 |
CIFEr | 2 |
| 2003 | Empirical Comparison between Two Computational Strategies for Topological Self-Organization
Bonnie Kit Yee Chan, Wilson Wei Sheng Chu, Lei Xu 0001 |
IDEAL | 3 |
| 2003 | A Comparative Study of Several Cluster Number Selection Criteria
Xuelei Hu, Lei Xu 0001 |
IDEAL | 2 |
| 2003 | Distribution approximation, combinatorial optimization, and Lagrange-BarrierabstractIn this paper, typical analog combinatorial optimization approaches, such as Hopfield net, Hopfield-Lagrange net, Maximum entropy approach, Lagrange-Barrier approach, are systematically examined from the perspective of learning distribution. The minimization of a combinatorial cost is turned into a procedure of learning a simple distribution to approximate the Gibbs distribution induced from this cost such that both the distributions share a same global peak. From this new perspective, a new general guideline is obtained for developing analog combinatorial optimization approaches. Moreover, the Lagrange-Barrier iterative procedure proposed in Xu (1994) is further elaborated with guaranteed convergence on a feasible solution that satisfies constraints. Lei Xu 0001 |
IJCNN | 1 |
| 2003 | Data-smoothing regularization, normalization regularization, and competition-penalty mechanism for statistical learning and multi-agentsabstractThis paper provides an overview on advances of two new learning regularization approaches, both are developed in the past several years from the studies of Bayesian Ying Yang learning (BYY). The first is data smoothing regularization, which was firstly proposed in (Xu, 1997a) for parameter learning in a way similar to Tikhonov regularization but with an easy solution to the difficulty of determining an appropriate hyperparameter. The second is normalization regularization firstly proposed in (Xu, 2001b) which regularizes parameter learning via de-learning of conscience or penalizing type and has a close relation to the rival penalized competitive learning (RPCL) (Xu, Krzyzak, & Oja, 1993). Also, the algorithms for the two types of regularized learning versus the algorithms for maximum likelihood learning and the RPCL learning are presented in a unified learning procedure. Moreover, studies on the competition-penalty mechanism are further elaborated, and this mechanism, especially RPCL mechanism, is suggested to monitoring the performances of multi-agents. Lei Xu 0001 |
IJCNN | 1 |
| 2003 | Further studies on temporal factor analysis: comparison and Kalman filter-based algorithm
Yiu-Ming Cheung, Lei Xu 0001 |
Neurocomputing | 2 |
| 2003 | Topological local principal component analysis
Lei Xu 0001 |
Neurocomputing | 2 |
| 2003 | BYY learning, regularized implementation, and model selection on modular networks with one hidden layer of binary units
Lei Xu 0001 |
Neurocomputing | 1 |
| 2003 | Corrigendum to "BYY learning, regularized implementation, and model selection on modular networks with one hidden layer of binary units" [Neurocomputing 51 (2003) 277-301]
Lei Xu 0001 |
Neurocomputing | 1 |
| 2003 | Improved system for object detection and star/galaxy classification via local subspace analysis
Kai Chun Chiu, Lei Xu 0001 |
Neural Networks | 3 |
| 2003 | Data smoothing regularization, multi-sets-learning, and problem solving strategies
Lei Xu 0001 |
Neural Networks | 1 |
| 2003 | Strip line detection and thinning by RPCL-based local PCA
Kai Chun Chiu, Lei Xu 0001 |
Pattern Recognit. Lett. | 3 |
| 2003 | An extended ASLD trading system to enhance portfolio managementabstractAn adaptive supervised learning decision (ASLD) trading system has been presented by Xu and Cheung (1997) to optimize the expected returns of investment without considering risks. In this paper, we propose an extension of the ASLD system (EASLD), which combines the ASLD with a portfolio optimization scheme to take a balance between the expected returns and risks. This new system not only keeps the learning adaptability of the ASLD, but also dynamically controls the risk in pursuit of great profits by diversifying the capital to a time-varying portfolio of N assets. Consequently, it is shown that: 1) the EASLD system gives the investment risk much smaller than the ASLD one; and 2) more returns are gained through the EASLD system in comparison with the two individual portfolio optimization schemes that statically determine the portfolio weights without adaptive learning. We have justified these two issues by the experiments. Kei-Keung Hung, Yiu-Ming Cheung, Lei Xu 0001 |
IEEE Trans. Neural Networks | 3 |
| 2003 | Dual multivariate auto-regressive modeling in state space for temporal signal separationabstractMany existing independent component analysis (ICA) approaches result in deteriorated performance in temporal source separation because they have not taken into consideration of the underlying temporal structure of sources. In this paper, we model temporal sources as a general multivariate auto-regressive (AR) process whereby an underlying multivariate AR process in observation space is obtained. In this dual AR modeling, the mixing process from temporal sources to observations is the same as the mixture from the nontemporal residuals of the source AR (SAR) process to that of the observation AR (OAR) process. We can therefore avoid the source temporal effects in performing ICA by learning the demixing system on the independently distributed OAR residuals rather than the time-correlated observations. Particularly, we implement this approach by modeling each source signal as a finite mixture of generalized autoregressive conditional heteroskedastic (GARCH) process. The adaptive algorithms are proposed to extract the OAR residuals appropriately online, together with learning the demixing system via a nontemporal ICA algorithm. The experiments have shown its superior performance on temporal source separation. Yiu-Ming Cheung, Lei Xu 0001 |
IEEE Trans. Syst. Man Cybern. Part B | 2 |
| 2003 | Correction to "dual multivariate auto-regressive modeling in state space for temporal signal separation"
Yiu-Ming Cheung, Lei Xu 0001 |
IEEE Trans. Syst. Man Cybern. Part B | 2 |
| 2002 | Financial APT-Based Gaussian TFA Learning for Adaptive Portfolio Management
Kai Chun Chiu, Lei Xu 0001 |
ICANN | 2 |
| 2002 | On Convergence of an Iterative Factor Estimate Algorithm for the NFA Model
Lei Xu 0001 |
ICANN | 2 |
| 2002 | Stock Price and Index Forecasting by Arbitrage Pricing Theory-Based Gaussian TFA Learning
Kai Chun Chiu, Lei Xu 0001 |
IDEAL | 2 |
| 2002 | A Comparative Study on Three MAP Factor Estimate Approaches for NFA
Lei Xu 0001 |
IDEAL | 2 |
| 2002 | Mining Dependence Structures from Statistical Learning Perspective
Lei Xu 0001 |
IDEAL | 1 |
| 2002 | A Lagrange Multiplier and Hopfield-Type Barrier Function Method for the Traveling Salesman ProblemabstractA Lagrange multiplier and Hopfield-type barrier function method is proposed for approximating a solution of the traveling salesman problem. The method is derived from applications of Lagrange multipliers and a Hopfield-type barrier function and attempts to produce a solution of high quality by generating a minimum point of a barrier problem for a sequence of descending values of the barrier parameter. For any given value of the barrier parameter, the method searches for a minimum point of the barrier problem in a feasible descent direction, which has a desired property that lower and upper bounds on variables are always satisfied automatically if the step length is a number between zero and one. At each iteration, the feasible descent direction is found by updating Lagrange multipliers with a globally convergent iterative procedure. For any given value of the barrier parameter, the method converges to a stationary point of the barrier problem without any condition on the objective function. Theoretical and numerical results show that the method seems more effective and efficient than the softassign algorithm. Chuangyin Dang, Lei Xu 0001 |
Neural Comput. | 2 |
| 2002 | BYY harmony learning, structural RPCL, and topological self-organizing on mixture models
Lei Xu 0001 |
Neural Networks | 1 |
| 2002 | Errata to "BYY harmony learning, independent state space, and generalized apt financial analyses"abstractProspective authors are requested to submit new, unpublished manuscripts for inclusion in the upcoming event described in this call for papers. Lei Xu 0001 |
IEEE Trans. Neural Networks | 1 |
| 2001 | RPCL-Based Local PCA AlgorithmabstractMining local structure is important in data analysis. Gaussian mixture is able to describe local structure through covariance matrices, but when used on high-dimensional data, specifying such a large number of d(d+1)/2 free elements in each covariance matrix is difficult. By constraining the covariance matrix in decomposed orthonormal form, we propose a Local PCA algorithm to tackle this problem with the help of RPCL (Rival Penalized Competitive Learning), which can automatically determine the number of local structures. Lei Xu 0001 |
ICDM | 2 |
| 2001 | Scan array solution for testing power and testing timeabstractDetails a novel power estimation algorithm based on rate of bit propagation (RBP). Considering the reduction of RBP an advanced scan array architecture is proposed in which a wrapper and two dimensional scan chain is adopted. Estimated results based on RBP and experimental results of industrial circuits both show that testing power was reduced to the level of the functional power. Furthermore pseudo-BIST is integrated with the wrapper to reduce the test time. Lei Xu 0001, Yihe Sun |
ITC | 1 |
| 2001 | Best Harmony, Unified RPCL and Automated Model Selection for Unsupervised and Supervised Learning on Gaussian Mixtures, Three-Layer Nets and ME-RBF-SVM ModelsabstractAfter introducing the fundamentals of BYY system and harmony learning, which has been developed in past several years as a unified statistical framework for parameter learning, regularization and model selection, we systematically discuss this BYY harmony learning on systems with discrete inner-representations. First, we shown that one special case leads to unsupervised learning on Gaussian mixture. We show how harmony learning not only leads us to the EM algorithm for maximum likelihood (ML) learning and the corresponding extended KMEAN algorithms for Mahalanobis clustering with criteria for selecting the number of Gaussians or clusters, but also provides us two new regularization techniques and a unified scheme that includes the previous rival penalized competitive learning (RPCL) as well as its various variants and extensions that performs model selection automatically during parameter learning. Moreover, as a by-product, we also get a new approach for determining a set of 'supporting vectors' for Parzen window density estimation. Second, we shown that other special cases lead to three typical supervised learning models with several new results. On three layer net, we get (i) a new regularized ML learning, (ii) a new criterion for selecting the number of hidden units, and (iii) a family of EM-like algorithms that combines harmony learning with new techniques of regularization. On the original and alternative models of mixture-of-expert (ME) as well as radial basis function (RBF) nets, we get not only a new type of criteria for selecting the number of experts or basis functions but also a new type of the EM-like algorithms that combines regularization techniques and RPCL learning for parameter learning with either least complexity nature on the original ME model or automated model selection on the alternative ME model and RBF nets. Moreover, all the results for the alternative ME model are also applied to other two popular nonparametric statistical approaches, namely kernel regression and supporting vector machine. Particularly, not only we get an easily implemented approach for determining the smoothing parameter in kernel regression, but also we get an alternative approach for deciding the set of supporting vectors in supporting vector machine. Lei Xu 0001 |
Int. J. Neural Syst. | 1 |
| 2001 | Independent component ordering in ICA time series analysis
Yiu-Ming Cheung, Lei Xu 0001 |
Neurocomputing | 2 |
| 2001 | Asymptotic Convergence Rate of the EM Algorithm for Gaussian MixturesabstractIt is well known that the convergence rate of the expectation-maximization (EM) algorithm can be faster than those of convention first-order iterative algorithms when the overlap in the given mixture is small. But this argument has not been mathematically proved yet. This article studies this problem asymptotically in the setting of gaussian mixtures under the theoretical framework of Xu and Jordan (1996). It has been proved that the asymptotic convergence rate of the EM algorithm for gaussian mixtures locally around the true solution Theta* is o(e(0. 5-epsilon)(Theta*)), where epsilon > 0 is an arbitrarily small number, o(x) means that it is a higher-order infinitesimal as x --> 0, and e(Theta*) is a measure of the average overlap of gaussians in the mixture. In other words, the large sample local convergence rate for the EM algorithm tends to be asymptotically superlinear when e(Theta*) tends to zero. Jinwen Ma, Lei Xu 0001, Michael I. Jordan |
Neural Comput. | 2 |
| 2001 | A globally convergent Lagrange and barrier function iterative algorithm for the traveling salesman problem
Chuangyin Dang, Lei Xu 0001 |
Neural Networks | 2 |
| 2001 | BYY harmony learning, independent state space, and generalized APT financial analysesabstractFirst, the relationship between factor analysis (FA) and the well-known arbitrage pricing theory (APT) for financial market is discussed comparatively, with a number of to-be-improved problems listed. An overview is made from a unified perspective on the related studies in the literatures of statistics, control theory, signal processing, and neural networks. Next, we introduce the fundamentals of the Bayesian Ying Yang (BYY) system and the harmony learning principle. We further show that a specific case of the framework, called BYY independent state space (ISS) system, provides a general guide for systematically tackling various FA related learning tasks and the above to-be-improved problems for the APT analyses. Third, on various specific cases of the BYY ISS system in three typical architectures, adaptive algorithms, regularization methods and model selection criteria are provided for either or both of parameter learning with automated model selection and parameter learning followed by model selection. Finally, we introduce some other financial applications that are based on the underlying independent factors via the APT analyses. Lei Xu 0001 |
IEEE Trans. Neural Networks | 1 |
| 2001 | A PCA approach for fast retrieval of structural patterns in attributed graphsabstractAn attributed graph (AG) is a useful data structure for representing complex patterns in a wide range of applications such as computer vision, image database retrieval, and other knowledge representation tasks where similar or exact corresponding structural patterns must be found. Existing methods for attributed graph matching (AGM) often suffer from the combinatorial problem whereby the execution cost for finding an exact or similar match is exponentially related to the number of nodes the AG contains. The square matching error of two AGs subject to permutations is approximately relaxed to a square matching error of two AGs subject to orthogonal transformations. Hence, the principal component analysis (PCA) algorithm can be used for the fast computation of the approximate matching error, with a considerably reduced execution complexity. Experiments demonstrate that this method works well and is robust against noise and other simple types of transformations. Lei Xu 0001, Irwin King |
IEEE Trans. Syst. Man Cybern. Part B | 1 |
| 2000 | New Sharpe-ratio-related methods for portfolio selectionabstractIn this paper, we formulate methods for portfolio selection for investors with different attitudes in the return-risk trade-off. We defined an objective function based on the Sharpe ratio (Sharpe, 1966) and downside risk (Fishburn, 1977), plus introducing two new terms called "upside volatility" and "diversification". We propose the maximization of the objective function WRT the portfolio weights as a method of determining suitable weights. We also propose practical methods for controlling the expected return while minimizing risk, of controlling risk while maximising expected return. Experiments showed that the proposed methods yielded successful results. Kei-Keung Hung, Chi Chiu Cheung, Lei Xu 0001 |
CIFEr | 3 |
| 2000 | Best Harmony Learning
Lei Xu 0001 |
IDEAL | 1 |
| 2000 | A RPCL-Based Approach for Identification of Markov Model with Unknown Noise and Number of Statesabstract(Krishnamurthy et al. 1993) studied one type of Hidden Markov Model (HMM) with identifying its state sequence and parameters based on the Expectation-Maximization (EM) algorithm, thus requiring extensive computing resources and a prior knowledge of state number. In this paper, we further study this model and present a new identification approach, which estimates the state sequence and HMM parameters through using the clustering information obtained via Rival Penalized Competitive Learning (RPCL) algorithm (Xu et al., 1992, 1993). Compared to Krishnamurthy's method, our approach can not only fast identify the HMM, but also automatically find out the correct number of states. Experiments have successfully shown the performance of this approach. Yiu-Ming Cheung, Lei Xu 0001 |
IJCNN (4) | 2 |
| 2000 | An Application of Independent Component Analysis in the Arbitrage Pricing TheoryabstractThe arbitrage pricing theory is a normative equilibrium theory, the theory infers that there may be a multitude of risk factors driving asset returns. Broadly speaking, three approaches can be used to identify these factors, they are the fundamental model, macroeconomic model and statistical model. In the traditional approach, the statistical model applies principal component analysis (PCA) to decompose the covariance matrix of asset returns. We discuss the relationship between macroeconomic variables and statistical factors and we apply independent component analysis (ICA) and a newly proposed ICA factor selection criterion in the statistical model, using ICA as a plug-in of the statistical model. Experiments have shown that it gives a better indication of the underlying structure of the stock market than PCA by using the same number of components. Fung Yip, Lei Xu 0001 |
IJCNN (5) | 2 |
| 2000 | Adaptive Improved Portfolio Sharpe Ratio Maximization with DiversificationabstractA portfolio selection method called Improved Portfolio Sharpe Ratio Maximization with Diversification (IPSRM-D) was recently proposed by Hung et al. (1999). It is derived from the original Sharpe Ratio design by taking into consideration the upside volatility and investment diversification. It can obtain investment decision according to the investor's position in the return-risk trade off. However, the batchway method used in IPSRM-D lacks the ability to keep tracing the changes in market from the just available data. In this paper, we further the study on IPSRM-D by introducing various adaptive methods. We demonstrate with experimental results on stock market that adaptive methods outperform batchway method in profit gain. Xiaohui Yu 0001, Lei Xu 0001 |
IJCNN (4) | 2 |
| 2000 | Rival Penalized Competitive Learning Based Approach for Discrete-Valued Source SeparationabstractThis paper presents an approach based on Rival Penalized Competitive Learning (RPCL) rules for discrete-valued source separation. In this approach, we first build a connection between the source number and the cluster number of observations. Then, we use the RPCL rule to automatically find out the correct number of clusters such that the source number is determined. Moreover, we tune the de-mixing matrix based on the cluster centers instead of the observation themselves, whereby the noise interference is considerably reduced. The experiments have shown that this new approach not only quickly and automatically determines the number of sources, but also is insensitive to the noise in performing blind source separation. Yiu-Ming Cheung, Lei Xu 0001 |
Int. J. Neural Syst. | 2 |
| 2000 | Some global and local convergence analysis on the information-theoretic independent component analysis approach
Chi Chiu Cheung, Lei Xu 0001 |
Neurocomputing | 2 |
| 2000 | A Barrier Function Method for the Nonconvex Quadratic Programming Problem with Box Constraints
Chuangyin Dang, Lei Xu 0001 |
J. Glob. Optim. | 2 |
| 2000 | An RPCL-based approach for Markov model identification with unknown state numberabstractThis paper presents an alternative identification approach for the Markov model studied in Krishnamurthy and Moore (1993). Our approach estimates the state sequence and model parameters with the help of a clustering analysis by the rival penalized competitive learning (RPCL) algorithm (Xa 1996). Compared to the method in Krishnamurthy and Moore, this new approach not only extends the model from scalar states to multidimensional ones, but also makes the model identification with the correct number of states decided automatically. The experiments have shown that it works well. Yiu-Ming Cheung, Lei Xu 0001 |
IEEE Signal Process. Lett. | 2 |
| 1999 | Temporal BYY learning and its applications to extended Kalman filtering, hidden Markov model, and sensor-motor integrationabstractThis paper systematically re-elaborates the author's temporal Bayesian Ying-Yang (TBYY) learning system and theory (1998). First, the previous approximate implementation of TBYY theory by recursive TBYY has been justified from the first order Taylor expansion. Second, an alternative suggestion is given for extending Kalman filtering to nonGaussian noise and nonlinear state space model. Third, other two variants of hidden Markov model (HMM) are proposed for facilitating adaptive learning, with criteria for selecting the number of hidden states. Finally, the recursive TBYY has been applied to the problem of sensor-motor integration, which can be regarded as a probabilistic extension of Kawato's feedback-error-learning (1990). Lei Xu 0001 |
IJCNN | 1 |
| 1999 | An ICA algorithm with adaptive-learned polynomial nonlinearity for signal separationabstractThis paper presents a novel approach, called adaptive polynomial power learning estimation (APPLE) based ICA algorithm, for independent component analysis (ICA) problem. In this algorithm, the form of separation nonlinearity is fixed at polynomial function, but the exponent is adaptive adjusted in implementation. Experiments have demonstrated that this algorithm can successfully separate the combinations of sub-Gaussian and super-Gaussian signals. Yiu-Ming Cheung, Lei Xu 0001 |
IJCNN | 2 |
| 1999 | Autoregressive signal separation approach with seesaw-mapping technique on temporal source separationabstractMost existing independent component analysis approaches are proposed for blind signal separation under the assumption that the sources are independently and identically distributed signals. However, the real signals are often temporal correlated in a certain degree. In our previous paper (1999), we have presented an autoregressive signal separation approach (ASSA) for AR(p) temporal signal separation, where we assume the noises in AR source signals are non-Gaussian. In this paper, we further study this approach under the Gaussian noises in AR sources with the seesaw-mapping technique. Experiments demonstrated that the seesaw-mapping technique can be applied successfully to the ASSA approach. Yiu-Ming Cheung, Lei Xu 0001 |
IJCNN | 2 |
| 1999 | An empirical method to select dominant independent components in ICA for time series analysisabstractBack and Weigend (1997) showed that the dominant independent components obtained by independent component analysis (ICA) can reveal more underlying structure of the time series than principal component analysis. To find those dominant independent components, all the independent components are listed in an appropriate order and then a subset of components is selected according to the order. However, currently there does not exist a systematic way to choose such a subset. In this paper, we propose a number selection criterion to choose an appropriate dominant number, through which the dominant independent components can be automatically determined from a set of ordered components. Experiments on foreign exchange rates have shown the performance of this empirical method. Yiu-Ming Cheung, Lei Xu 0001 |
IJCNN | 2 |
| 1999 | On the study of BKYY cluster number selection criterion for small sample data set with bootstrap techniqueabstractThe Bayesian-Kullback ying-yang (BKYY) learning theory and system has been proposed by Xu (1995, 1997), and one special case of ying-yang system can provide the model selection criteria for selecting the number of clusters in the clustering analysis. In this paper, we present an experimental study of this cluster number selection criterion in a small number sample set case. The results show that the criterion performed reasonable well when mixture parameters were estimated by incorporating a bootstrap technique with the EM algorithm. Lei Xu 0001 |
IJCNN | 2 |
| 1999 | Further improvement of adaptive supervised learning decision (ASLD) network in stock marketabstractWe apply a neural network model called adaptive supervised learning decision network (ASLD), proposed by Xu and Cheung (1997), that maximize the expected return. In generating the trading signals for training the neural network used in the ASLD system, besides maximizing the profit gain, we have also applied the portfolio technique related to Sharpe ratio (1994) which consider expected risk in addition. Instead of making use of the original design of the Sharpe ratio maximization, we have replaced the traditional risk with a more sophisticated quantity called "downside risk" proposed by Sortino and van der Meer (1991) and "upside volatility" we proposed. Moreover, a regularization idea is introduced to make the portfolio distributed more evenly over the indexes. Lastly, using the augmented Lagrangian method, we have developed system that can either control the expected return and minimize the downside risk, or control the downside risk and maximize the expected return. Kei-Keung Hung, Lei Xu 0001 |
IJCNN | 2 |
| 1999 | Bayesian ying-yang supervised learning, modular models, and three layer netsabstractBayesian ying-yang (BYY) supervised learning system and theory is further re-elaborated, and the previous results of its uses on mixture-of-expert models, radial basis functions and three layer nets are systematically summarized. Moreover, new results on three layer net are presented. Using Taylor expansion on the distribution of the output layer, we find that maximum likelihood (ML) learning on a net with a probabilistic hidden layer is equivalent to adding a regularization to its counterpart with a deterministic hidden layer, which leads us not only an adaptive EM-like algorithm for ML learning on three layer net, but also a new type of regularization technique. Furthermore, an improved BYY criterion is obtained for selecting the number of hidden units. Lei Xu 0001 |
IJCNN | 1 |
| 1999 | BYY data smoothing based learning on a small size of samplesabstractBayesian ying-yang (BYY) data smoothing based learning provides a general framework for parametric learning on a small size of samples by Parzen window nonparametric density estimation, with the best optimal smoothing parameter. This paper not only systematically elaborates the general formulation of BYY data smoothing based learning, but also presents several new results on both implementing smoothed parameter learning and estimating the best smoothing parameter for supervised and unsupervised learning tasks. Moreover, detailed studies have also been made on data smoothing based learning for Gaussian mixture, mixture-of-expert models, and three layer nets. Lei Xu 0001 |
IJCNN | 1 |
| 1999 | Bayesian ying-yang theory for empirical learning, regularization and model selection: general formulationabstractThe Bayesian ying-yang learning system and theory developed by the present author (1995, 1996) is further elaborated in a general formulation, focusing on systematically introducing the key points of the theory for empirical learning, data smoothing based regularization, structural regularization, and model selection. Moreover, discussions have been made on the relationship and difference between this theory and the existing approaches, especially the Helmholtz machine, information theory as well as information geometry theory. Lei Xu 0001 |
IJCNN | 1 |
| 1999 | Temporal Bayesian Ying-Yang dependence reduction, blind source separation and principal independent componentsabstractA general recursive temporal BYY dependence reduction system, previously proposed by the author (1998), has been further re-elaborated systematically. Three new developments have been also made. A temporal independent component analysis (ICA) algorithm is proposed for separating blind sources from observations without sensor noise. The other two studies are for solving blind source separation problems with sensor noise of binary and real temporal sources, respectively. Moreover, the problem of how to determine principal independent components is addressed with criteria provided. Lei Xu 0001 |
IJCNN | 1 |
| 1999 | Data mining, unsupervised learning and Bayesian ying-yang theoryabstractA number of unsupervised learning methods or algorithms have been summarized from the perspective of their potential uses in data mining. Major unsupervised learning tasks are then systematically viewed under a unified framework called Bayesian ying-yang (BYY) learning. Furthermore, it is shown systematically how the BYY learning theory can guide us not only to revisit the existing major unsupervised learning methods and results, but also to obtain a number of new methods and results. Lei Xu 0001 |
IJCNN | 1 |
| 1999 | Improved learning algorithms for mixture of experts in multiclass classification
Ke Chen 0001, Lei Xu 0001, Huisheng Chi |
Neural Networks | 2 |
| 1998 | Cointegration by MCA and modular MCAabstractA new approach for investigating the cointegration relationship between time series is proposed based on the so called minor component analysis (MCA) and its extension, modular MCA. With no need to place one variable on the left-hand side and use the others as regressors on the estimation of the long-run equilibrium regression, and also no need to build up a vector autoregression model for rank based analysis, this approach conducts cointegration analysis directly based on the linear combination of the instantaneous values of nonstationary variables via an eigen-analysis corresponding to the smallest eigenvalue, which can be implemented either in a batch way or by a Hebbian learning adaptive algorithm. A method for estimating the order of the cointegration has been also developed. Experiments are made in comparison with the least square regression based cointegration approach and have demonstrated the advantages of the approach. Lei Xu 0001, Wai-Man Leung |
CIFEr | 1 |
| 1998 | Adaptive RBF net algorithms for nonlinear signal learning with applications to financial prediction and investmentabstractA smoothed variant of the EM algorithm is given for simultaneous training of the first layer and the output layer globally in the normalized radial basis function (NRBF) nets and extended normalized RBF nets (ENRBF), together with a Bayesian Ying-Yang learning criterion for the selection of basis function numbers. Moreover, a hard-cut fast implementation and an adaptive algorithm have also been proposed for speeding up the training and for handling time varying in real time nonlinear signal learning and processing. A number of experiments are made on foreign exchange prediction and trading investments. Lei Xu 0001 |
ICASSP | 1 |
| 1998 | An experimental comparison of the Bayesian Ying-Yang criteria and cross validation for selection on number of hidden units in feedforward networksabstractOptimizing the number of hidden units in feedforward neural networks is an important issue in learning. Previously, a new criterion for selecting the number of hidden units in feedforward neural networks has been proposed by Xu (1997), based on the so-called Bayesian Ying-Yang (BYY) learning theory. The new criterion can be simply computed during the implementation of backpropagation training. In this paper, the criterion is experimentally studied and compared with the well-known cross validation approach. Simulation results show that obtained number of hidden units by the BYY criterion is highly consistent with the minimal generalization error and outperforms the cross validation approach. Wing-Kai Lam, Lei Xu 0001 |
ICASSP | 2 |
| 1998 | Rival Penalized Competitive Learning Based Separator on Binary Sources Separation
Yiu-Ming Cheung, Lei Xu 0001 |
ICONIP | 2 |
| 1998 | Comparison between Mixture of ARMA and Mixture of AR Model with Application to Time Series Forecasting
Ho Yin Kwok, Chi-Ming Chen, Lei Xu 0001 |
ICONIP | 3 |
| 1998 | An Experimental Comparison of the Bayesian Ying-Yang Criteria and Cross Validation on Experts Number Selection in Original and Alternative Model for Mixture of Experts
Wing-Kai Lam, Lei Xu 0001 |
ICONIP | 2 |
| 1998 | The Correct Convergence of the Rival Penalized Competitive Learning (RPCL) Algorithm
Jinwen Ma, Lei Xu 0001 |
ICONIP | 2 |
| 1998 | Financial Prediction by Finite Mixture GARCH Model
Wai Ching Wong, Fung Yip, Lei Xu 0001 |
ICONIP | 3 |
| 1998 | The Role of Statistical Theory in the Future Research on Brain-like Systems
Lei Xu 0001 |
ICONIP | 1 |
| 1998 | Bayesian Ying-Yang System and Theory as a Unified Statistical Learning Approach (VII): Data Smoothing
Lei Xu 0001 |
ICONIP | 1 |
| 1998 | BKYY Three Layer Net Learning, EM-Like Algorithm, and Selection Criterion for Hidden Unit Number
Lei Xu 0001 |
ICONIP | 1 |
| 1998 | Bayesian Ying-Yang System and Theory as a Unified Statistical Learning Approach: (V) Temporal Modeling for Temporal Perception and Control
Lei Xu 0001 |
ICONIP | 1 |
| 1998 | Intra-block algorithm for digital watermarkingabstractWe present a variant to the DCT-based block algorithm proposed in Hsu and Wu (1996) for signal embedding in digital images. Instead of inter-block relations, our algorithm uses intra-block relations to generate the watermarked image. We describe the algorithm and its performance against translation and cropping. The features of our method are: (1) the watermark is perceptually invisible; (2) little loss of relevant information of original image; (3) the watermark can be retrieved by using a secret key; and (4) the watermark is robust against translation and area cropping. F. Y. Duan, Irwin King, Lai-Wan Chan, Lei Xu 0001 |
ICPR | 4 |
| 1998 | Application of Bayesian Ying-Yang criteria for selecting the number of hidden units with backpropagation learning to electrocardiogram classificationabstractComputer electrocardiography (EGG) is a fundamental diagnostic method for both contour and rhythm analysis. For patients with Hypertrophic Cardiomyopathy (HCM), there are more or less abnormal ECG findings. However, to diagnose HCM through ECG is a difficult task even for experienced cardiologist. Backpropagation has been used for ECG classification with the number of hidden units chosen heuristically. In this paper, the hidden unit number is selected by a new criteria obtained from the so-call Bayesian Ying-Yang learning theory and applied in ECG classification to diagnose HCM. Experiments have shown that the selected number is highly consistent with the minimal generalization error and the corresponding architecture show best classification performance. Wing-Kai Lam, Ning Ouyang, Lei Xu 0001 |
ICPR | 3 |
| 1998 | RBF nets, mixture experts, and Bayesian Ying-Yang learning
Lei Xu 0001 |
Neurocomputing | 1 |
| 1998 | Bayesian Kullback Ying-Yang dependence reduction theory
Lei Xu 0001 |
Neurocomputing | 1 |
| 1998 | Learned parametric mixture based ICA algorithm
Lei Xu 0001, Chi Chiu Cheung, Shun-ichi Amari |
Neurocomputing | 1 |
| 1997 | Adaptive supervised learning decision networks for traders and portfoliosabstractWe propose an adaptive supervised learning decision network for portfolio management which learns the best past investment decision directly instead of making a good prediction first and then making an investment decision based on the prediction. Without any extra effort, this network can be realized directly by any existing adaptive supervised learning neural networks. We propose to use an extended normalized radial basis function (ENRBF) network with matched competitive learning (MCL). We demonstrate with experimental results that the proposed approach can bring in appreciable profit on trading in the foreign exchange market. Lei Xu 0001, Yiu-Ming Cheung |
CIFEr | 1 |
| 1997 | Nonlinearity and separation capability: further justification for the ICA algorithm with mixture of densities
Lei Xu 0001, Chi Chiu Cheung, Jiong Ruan, Shun-ichi Amari |
ESANN | 1 |
| 1997 | Adaptive Rival Penalized Competitive Learning and Combined Linear Predictor Model for Financial Forecast and InvestmentabstractWe propose a prediction model called Rival Penalized Competitive Learning (RPCL) and Combined Linear Predictor method (CLP), which involves a set of local linear predictors such that a prediction is made by the combination of some activated predictors through a gating network (Xu et al., 1994). Furthermore, we present its improved variant named Adaptive RPCL-CLP that includes an adaptive learning mechanism as well as a data pre-and-post processing scheme. We compare them with some existing models by demonstrating their performance on two real-world financial time series--a China stock price and an exchange-rate series of US Dollar (USD) versus Deutschmark (DEM). Experiments have shown that Adaptive RPCL-CLP not only outperforms the other approaches with the smallest prediction error and training costs, but also brings in considerable high profits in the trading simulation of foreign exchange market. Yiu-Ming Cheung, Wai-Man Leung, Lei Xu 0001 |
Int. J. Neural Syst. | 3 |
| 1997 | Comparative Analysis on Convergence Rates of The EM Algorithm and Its Two Modifications for Gaussian Mixtures
Lei Xu 0001 |
Neural Process. Lett. | 1 |
| 1997 | Bayesian Ying-Yang machine, clustering and number of clusters
Lei Xu 0001 |
Pattern Recognit. Lett. | 1 |
| 1997 | Yet another algorithm which can generate topography mapabstractThis paper presents an algorithm to form a topographic map resembling to the self-organizing map. The idea stems on defining an energy function which reveals the local correlation between neighboring neurons. The larger the value of the energy function, the higher the correlation of the neighborhood neurons. On this account, the proposed algorithm is defined as the gradient ascent of this energy function. Simulations on two-dimensional maps are illustrated. John Sum, Andrew Chi-Sing Leung, Lai-Wan Chan, Lei Xu 0001 |
IEEE Trans. Neural Networks | 4 |
| 1996 | Adaptive Rival Penalized Competitive Learning and Combined Linear Predictor with application to financial investmentabstractWe have recently proposed an architecture called Rival Penalized Competitive Learning and Combined Linear Predictor (RPCL-CLP) to model financial time series with a certain degree of success (Cheung et al., 1995). Experiments have shown that RPCL-CLP outperforms ClusNet (Hsu et al., 1993), but it still has features which can be further improved. We propose a modified version called Adaptive RPCL-CLP which can automatically select the number of the initial cluster nodes for RPCL (Xu et al., 1993) and adaptively train the linear predictor's parameters in each cluster node as well as the gating network. We apply it to the forecasting of foreign exchange rates and the Shanghai stock price. As shown by experiments, this adaptive version is much better than RPCL-CLP, and with a trading system it can bring in more returns in foreign exchange market trading. Yiu-Ming Cheung, Helen Z. H. Lai, Lei Xu 0001 |
CIFEr | 3 |
| 1996 | Trading mechanisms and return volatility: empirical investigation on Shanghai Stock Exchange based on a neural network modelabstractWe empirically compare the behavior of open-to-open and close-to-close returns on the Shanghai Stock Exchange (SHSE) with different trading mechanisms (call market at the opening in the morning followed by continuous market). We use non-linear regression based on a neural network to study the volatility and efficiency of SHSE. The experimental results have shown that the volatility of the call market is significantly higher than that of the continuous market and the call market is more efficient than the continuous market. Helen Z. H. Lai, Yiu-Ming Cheung, Lei Xu 0001 |
CIFEr | 3 |
| 1996 | Learning Multiple Causes by Competition Enhanced least mean Square error ReconstructionabstractIn this paper we studied a self-organization principle that input should be best reconstructed from a factorial distributed hidden representation, which has been addressed in the literature recently. An auto-encoder network is trained by the Least Mean Square Error Reconstruction (LMSER) while the redundance in the representation is reduced by a proposed anti-Hebbian scheme, in which a penalty term called Receptive Field Overlapping Index (RFOI) is combined into the objective function for enhancing competition among nodes in the network. Our learning scheme provides a way for balancing the cooperation and competition necessary for the self-organization process thus realizes the multiple causes model, which accounts for an observed data by combining assertions from the discovered causes or features in the data. Our experiment results demonstrate again the powerful information processing capability inherent to the popular weighted sum followed by sigmoid squashing. Comparing with previous probability theory based multiple causes models, our scheme is much easier to implement and quite reliable. Lei Xu 0001, Minyue Fu 0001 |
Int. J. Neural Syst. | 2 |
| 1996 | On Convergence Properties of the EM Algorithm for Gaussian MixturesabstractWe build up the mathematical connection between the “Expectation-Maximization” (EM) algorithm and gradient-based approaches for maximum likelihood learning of finite gaussian mixtures. We show that the EM step in parameter space is obtained from the gradient via a projection matrix P, and we provide an explicit expression for the matrix. We then analyze the convergence of EM in terms of special properties of P and provide new results analyzing the effect that P has on the likelihood surface. Based on these mathematical results, we present a comparative discussion of the advantages and disadvantages of EM and other algorithms for the learning of gaussian mixture models. Lei Xu 0001, Michael I. Jordan |
Neural Comput. | 1 |
| 1995 | A Unified Learning Scheme: Bayesian-Kullback Ying-Yang Machines
Lei Xu 0001 |
NIPS | 1 |
| 1995 | Probabilistic and non-probabilistic Hough transforms: overview and comparisons
Heikki Kälviäinen, Petri Hirvonen, Lei Xu 0001, Erkki Oja |
Image Vis. Comput. | 3 |
| 1995 | Bayesian Self-Organization Driven by Prior Probability DistributionsabstractRecent work by Becker and Hinton (1992) shows a promising mechanism, based on maximizing mutual information assuming spatial coherence, by which a system can self-organize to learn visual abilities such as binocular stereo. We introduce a more general criterion, based on Bayesian probability theory, and thereby demonstrate a connection to Bayesian theories of visual perception and to other organization principles for early vision (Atick and Redlich 1990). Methods for implementation using variants of stochastic learning are described. Alan L. Yuille, Stelios M. Smirnakis, Lei Xu 0001 |
Neural Comput. | 3 |
| 1995 | Convergence results for the EM approach to mixtures of experts architectures
Michael I. Jordan, Lei Xu 0001 |
Neural Networks | 2 |
| 1995 | Robust principal component analysis by self-organizing rules based on statistical physics approachabstractThis paper applies statistical physics to the problem of robust principal component analysis (PCA). The commonly used PCA learning rules are first related to energy functions. These functions are generalized by adding a binary decision field with a given prior distribution so that outliers in the data are dealt with explicitly in order to make PCA robust. Each of the generalized energy functions is then used to define a Gibbs distribution from which a marginal distribution is obtained by summing over the binary decision field. The marginal distribution defines an effective energy function, from which self-organizing rules have been developed for robust PCA. Under the presence of outliers, both the standard PCA methods and the existing self-organizing PCA rules studied in the literature of neural networks perform quite poorly. By contrast, the robust rules proposed here resist outliers well and perform excellently for fulfilling various PCA-like tasks such as obtaining the first principal component vector, the first k principal component vectors, and directly finding the subspace spanned by the first k vector principal component vectors without solving for each vector individually. Comparative experiments have been made, and the results show that the authors' robust rules improve the performances of the existing PCA algorithms significantly when outliers are present. Lei Xu 0001, Alan L. Yuille |
IEEE Trans. Neural Networks | 1 |
| 1994 | Comparisons of Probabilistic and Non-probabilistic Hough Transforms
Heikki Kälviäinen, Petri Hirvonen, Lei Xu 0001, Erkki Oja |
ECCV (2) | 3 |
| 1994 | On L1 convergence rate of RBF networks and kernel regression estimators with applications in classificationabstractStudies the convergence properties of the mean integrated absolute error (MIAE) for kernel regression estimators (KRE) and radial basis function (RBF) nets. The authors show that the MIAE of KRE and RBF nets converges to zero as the size of network and the size of training sequence tend to infinity, and the authors give the upper bound for the convergence rate for approximating functions satisfying Lipschitz condition of order /spl alpha/,0 Adam Krzyzak, Stan Klasa, Lei Xu 0001 |
ICPR (2) | 3 |
| 1994 | An Alternative Model for Mixtures of ExpertsabstractWe propose an alternative model for mixtures of experts which uses a different parametric form for the gating network. The modified model is trained by the EM algorithm. In comparison with earlier models-trained by either EM or gradient ascent-there is no need to select a learning stepsize. We report simulation experiments which show that the new architecture yields faster convergence. We also apply the new model to two problem domains: piecewise nonlinear function approximation and the combination of multiple previously trained classifiers. Lei Xu 0001, Michael I. Jordan, Geoffrey E. Hinton |
NIPS | 1 |
| 1994 | On radial basis function nets and kernel regression: Statistical consistency, convergence rates, and receptive field size
Lei Xu 0001, Adam Krzyzak, Alan L. Yuille |
Neural Networks | 1 |
| 1993 | Bayesian Self-Organization
Alan L. Yuille, Stelios M. Smirnakis, Lei Xu 0001 |
NIPS | 3 |
| 1993 | Least mean square error reconstruction principle for self-organizing neural-nets
Lei Xu 0001 |
Neural Networks | 1 |
| 1993 | Rival penalized competitive learning for clustering analysis, RBF net, and curve detectionabstractIt is shown that frequency sensitive competitive learning (FSCL), one version of the recently improved competitive learning (CL) algorithms, significantly deteriorates in performance when the number of units is inappropriately selected. An algorithm called rival penalized competitive learning (RPCL) is proposed. In this algorithm, not only is the winner unit modified to adapt to the input for each input, but its rival (the 2nd winner) is delearned by a smaller learning rate. RPCL can be regarded as an unsupervised extension of Kohonen's supervised LVQ2. RPCL has the ability to automatically allocate an appropriate number of units for an input data set. The experimental results show that RPCL outperforms FSCL when used for unsupervised classification, for training a radial basis function (RBF) network, and for curve detection in digital images. Lei Xu 0001, Adam Krzyzak, Erkki Oja |
IEEE Trans. Neural Networks | 1 |
| 1992 | Randomized Hough transform applied to translational and rotational motion analysisabstractA method has been developed to calculate 2-D motion in a sequence of time-varying images. The method, called motion detection using randomized Hough transform (MDRHT), is based on the randomized Hough transform (RHT). The RHT decreases considerably the time consumption and memory requirements of the Hough transform. The idea of the MDRHT is to pick randomly point pairs from two images and calculate the translation with them. The points can be e.g. edge points of the original images. This approach can avoid difficulties of standard segmentation methods like overlapping and covering, and has the advantages provided by the RHT. The method can be generalized by picking more than two points. After a brief review of the RHT applied to motion detection, the extended algorithm to calculate both translation and rotation is represented in this paper.> Heikki Kälviäinen, Erkki Oja, Lei Xu 0001 |
ICPR (1) | 3 |
| 1992 | Unsupervised and supervised classifications by rival penalized competitive learningabstractFor the classical k-means clustering algorithm, the problem of selecting an appropriate k is a hard problem and affects the performance of k-means strongly. When used for clustering analysis, the conventional competitive learning (CL) algorithms also have a similar crucial problem-the selection of an appropriate number of neural units. The performance of frequency sensitive competitive learning (FSCL)-one version of the improved CL algorithms, also significantly deteriorates when the number of units is inappropriately selected. The paper proposes a new algorithm called rival penalized competitive learning (RPCL), which has the ability of automatically allocating an appropriate number of units for an input data set. The experimental results have shown that RPCL outperforms FSCL significantly when they are used for unsupervised classification, and supervised classification through the radial basis function net.> Lei Xu 0001, Adam Krzyzak, Erkki Oja |
ICPR (2) | 1 |
| 1992 | Further developments on RHT: basic mechanisms, algorithms, and computational complexitiesabstractFurther developments on RHT are introduced here. The fundamental framework and the main components of this method are systematically elaborated. Its advantages are further confirmed. The deep mechanisms behind these advantages are exposed. Several improved algorithms for implementing RHT are proposed and tested by experiments on images with various kinds of strong noise. The results show that the advantages of RHT are quite robust. Moreover, the theoretical results are also given to show quantitatively that the complexities of RHT algorithms are considerably smaller than those of the HT.> Lei Xu 0001, Erkki Oja |
ICPR (1) | 1 |
| 1992 | Self-Organizing Rules for Robust Principal Component Analysis
Lei Xu 0001, Alan L. Yuille |
NIPS | 1 |
| 1992 | Recent Advances on Techniques of Static Feedforward Networks with Supervised LearningabstractThe rediscovery and popularization of the backpropagation training technique for multilayer perceptrons as well as the invention of the Boltzmann machine learning algorithm has given a new boost to the study on supervised learning networks. In recent years, besides widely spread applications and various further improvements of the classical backpropagation technique, many new supervised learning models, techniques as well as theories, have also been proposed in a vast number of publications. This paper tries to give a rather systematic review on the recent advances on supervised learning techniques and models for static feedforward networks. We summarize a great number of developments into four aspects: (1) Various improvements and variants made on the classical backpropagation techniques for Multilayer(static) perceptron nets, for speeding up training, avoiding local minima, increasing the generalization ability as well as for many other interesting purposes. (2) A number of other learning methods for training multilayer (static) perceptron, such as derivative estimation by perturbation, direct weight update by perturbation, genetic algorithms, recursive least square estimate and extended Kalman filters, linear programming, the policy of fixing one layer while updating another, constructing networks by converting decision tree classifiers and others. (3) Various other feedforward models which are also able to implement function approximation, probability density estimation and classification, including various models of basis function expansion (e.g. radial basis functions, restricted coulomb energy, multivariate adaptive regression splines, trigonometric and polynomial bases, projection pursuit, basis function tree and many others) and several other supervised learning models. (4) Models with complex structures, e.g. modular architecture, hierarchy architecture and others. Altogether, we try to give a global picture of the present state of supervised learning techniques (not including all the theoretical developments) for training static feedforward networks. Lei Xu 0001, Stan Klasa, Alan L. Yuille |
Int. J. Neural Syst. | 1 |
| 1992 | Modified Hebbian learning for curve and surface fitting
Lei Xu 0001, Erkki Oja, Ching Y. Suen |
Neural Networks | 1 |
| 1992 | Methods of combining multiple classifiers and their applications to handwriting recognitionabstractPossible solutions to the problem of combining classifiers can be divided into three categories according to the levels of information available from the various classifiers. Four approaches based on different methodologies are proposed for solving this problem. One is suitable for combining individual classifiers such as Bayesian, k-nearest-neighbor, and various distance classifiers. The other three could be used for combining any kind of individual classifiers. On applying these methods to combine several classifiers for recognizing totally unconstrained handwritten numerals, the experimental results show that the performance of individual classifiers can be improved significantly. For example, on the US zipcode database, 98.9% recognition with 0.90% substitution and 0.2% rejection can be obtained, as well as high reliability with 95% recognition, 0% substitution, and 5% rejection.> Lei Xu 0001, Adam Krzyzak, Ching Y. Suen |
IEEE Trans. Syst. Man Cybern. | 1 |
| 1991 | Neural Nets for Dual Subspace Pattern Recognition MethodabstractA new modification of the subspace pattern recognition method, called the dual subspace pattern recognition (DSPR) method, is proposed, and neural network models combining both constrained Hebbian and anti-Hebbian learning rules are developed for implementing the DSPR method. An experimental comparison is made by using our model and a three-layer forward net with backpropagation learning. The results illustrate that our model can outperform the backpropagation model in suitable applications. Lei Xu 0001, Adam Krzyzak, Erkki Oja |
Int. J. Neural Syst. | 1 |
| 1990 | Randomized Hough transform (RHT)abstractA method is developed for calculating the Hough transform (HT) and completing the task of finding global features from binary edge images. The method is based on the fact that a single parameter space point can be determined uniquely with a pair, triple, or generally n-tuple of points from the original picture. Such n-tuples of points can be chosen randomly from the edge image, giving the method the name randomized Hough transform (RHT). The new algorithm reduces the computation time and memory use of the HT drastically. In the standard HT one must calculate one parameter curve in the parameter space for one pixel, whereas in the RHT only one parameter point has to be solved for one n-tuple of points, and the presence of a specific curve in the image is quickly revealed by the accumulation of a small number of parameter points.> Pekka Kultanen, Lei Xu 0001, Erkki Oja |
ICPR (1) | 2 |
| 1990 | Adding Learned Expectation Into the Learning Procedure of Self-Organizing MapsabstractThe self-organizing topological map is generalized by adding a learned expectation to its learning procedure, in order to improve its stability in nonstationary environments with unexpected inputs and abnormal noise. Such a learned expectation is realized through a compatibility test which checks whether an input is compatible with the earlier learned patterns on a map unit before the unit starts to adapt to the input. The generalized map consists of multi-maps with a pipeline architecture, equipped with a parallel search strategy which allows for the implementation of the learned expectation to be fulfilled without extra computing costs. Computer experiments on several examples are given to show the characteristics of the generalized map in comparison with the original map. Lei Xu 0001 |
Int. J. Neural Syst. | 1 |
| 1990 | A new curve detection method: Randomized Hough transform (RHT)
Lei Xu 0001, Erkki Oja, Pekka Kultanen |
Pattern Recognit. Lett. | 1 |
| 1988 | A method for recognizing configurations consisting of line sets and its application to discrimination of seismic face structuresabstractTwo Hough transform methods are considered, and an improvement on the algorithm of Davies (1986) is proposed. A twin Hough transform is then developed to describe and recognize configurations consisting of line sets. Its application to the recognition of seismic face configurations is considered in detail.> Lei Xu 0001 |
ICPR | 1 |
| 1988 | Structuring causal tree models with continuous variables
Lei Xu 0001 |
Int. J. Approx. Reason. | 1 |
| 1987 | Almost unique specification of discrete finite length signal: From its end point and Fourier transform magnitudeabstractIn this paper, the reconstruction of discrete signal with finite time duration from its end point and Fourier Transform (FT) magnitude is considered. Based on one result of [1] that a class of discrete signal can be reconstructed from its FT magnitude and one end sample point, with the help of Measure Theory, furtherly we point out that a correspondence between RN+1space and discrete signals with duration of N+1 points can be set up, and the signals that can't be reconstructed from its end point and FT magnitude correspond to a subset of RN+1with measure zero. In other words, discrete signal with finite time duration can almost be uniquely reconstructed from its end point and FT magnitude. Lei Xu 0001, Pingfan Yan, Tong Chang |
ICASSP | 1 |
| 1987 | Structuring Causal Tree Models with Continuous Variables
Lei Xu 0001, Judea Pearl |
UAI | 1 |