EDBT 2026 Demo / reviewers in the wild / expert
Le-le Cao
dblp:155/3234 · also Lele Cao
· DBLP profile ↗
30ranked-venue papers
14as first author
17since 2021 · last 2025
0000-0002-5680-9031ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 25 · 11 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 2 first-author · 4 since 2021Databases, data management, data science and information retrieval · 5 · 3 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 2 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Distributional Counterfactual Explanations With Optimal TransportabstractCounterfactual explanations (CE) are the de facto method for providing insights into black-box decision-making models by identifying alternative inputs that lead to different outcomes. However, existing CE approaches, including group and global methods, focus predominantly on specific input modifications, lacking the ability to capture nuanced distributional characteristics that influence model outcomes across the entire input-output spectrum. This paper proposes distributional counterfactual explanation (DCE), shifting focus to the distributional properties of observed and counterfactual data, thus providing broader insights. DCE is particularly beneficial for stakeholders making strategic decisions based on statistical data analysis, as it makes the statistical distribution of the counterfactual resembles the one of the factual when aligning model outputs with a target distribution—something that the existing CE methods cannot fully achieve. We leverage optimal transport (OT) to formulate a chance-constrained optimization problem, deriving a counterfactual distribution aligned with its factual counterpart, supported by statistical confidence. The efficacy of this approach is demonstrated through experiments, highlighting its potential to provide deeper insights into decision-making models. Lei You 0002, Le-le Cao, Lei Lei 0001 |
AISTATS | 2 |
| 2025 | Thread: A Logic-Based Data Organization Paradigm for How-To Question Answering with Retrieval Augmented GenerationabstractKaikai An, Fangkai Yang, Liqun Li, Junting Lu, Sitao Cheng, Shuzheng Si, Lu Wang, Pu Zhao, Lele Cao, Qingwei Lin, Saravan Rajmohan, Dongmei Zhang, Baobao Chang. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Kaikai An, Fangkai Yang, Liqun Li, Junting Lu, Sitao Cheng, Shuzheng Si, Lu Wang 0029, Pu Zhao 0004, Le-le Cao, Qingwei Lin, Saravan Rajmohan, Dongmei Zhang 0001, Baobao Chang |
EMNLP | 9 |
| 2025 | Enhancing Graph Classification Robustness with Singular PoolingabstractGraph Neural Networks (GNNs) have achieved strong performance across a range of graph representation learning tasks, yet their adversarial robustness in graph classification remains underexplored compared to node classification. While most existing defenses focus on the message-passing component, this work investigates the overlooked role of pooling operations in shaping robustness. We present a theoretical analysis of standard flat pooling methods (sum, average and max), deriving upper bounds on their adversarial risk and identifying their vulnerabilities under different attack scenarios and graph structures. Motivated by these insights, we propose Robust Singular Pooling (RS-Pool), a novel pooling strategy that leverages the dominant singular vector of the node embedding matrix to construct a robust graph-level representation. We theoretically investigate the robustness of RS-Pool and interpret the resulting bound leading to improved understanding of our proposed pooling operator. While our analysis centers on Graph Convolutional Networks (GCNs), RS-Pool is model-agnostic and can be implemented efficiently via power iteration. Empirical results on real-world benchmarks show that RS-Pool provides better robustness than the considered pooling methods when subject to state-of-the-art adversarial attacks while maintaining competitive clean accuracy. Sofiane Ennadir, Oleg Smirnov, Yassine Abbahaddou, Le-le Cao, Johannes F. Lutzeyer |
NeurIPS | 4 |
| 2025 | Pool Me Wisely: On the Effect of Pooling in Transformer-Based ModelsabstractTransformer models have become the dominant backbone for sequence modeling, leveraging self-attention to produce contextualized token representations. These are typically aggregated into fixed-size vectors via pooling operations for downstream tasks. While much of the literature has focused on attention mechanisms, the role of pooling remains underexplored despite its critical impact on model behavior. In this paper, we introduce a theoretical framework that rigorously characterizes the expressivity of Transformer-based models equipped with widely used pooling methods by deriving closed-form bounds on their representational capacity and the ability to distinguish similar inputs. Our analysis extends to different variations of attention formulations, demonstrating that these bounds hold across diverse architectural variants. We empirically evaluate pooling strategies across tasks requiring both global and local contextual understanding, spanning three major modalities: computer vision, natural language processing, and time-series analysis. Results reveal consistent trends in how pooling choices affect accuracy, sensitivity, and optimization behavior. Our findings unify theoretical and empirical perspectives, providing practical guidance for selecting or designing pooling mechanisms suited to specific tasks. This work positions pooling as a key architectural component in Transformer models and lays the foundation for more principled model design beyond attention alone. Sofiane Ennadir, Levente Zólyomi, Oleg Smirnov, Tianze Wang, John Pertoft, Filip Cornell, Le-le Cao |
NeurIPS | 7 |
| 2025 | Prompt Tuning Decision Transformers with Structured and Scalable BanditsabstractPrompt tuning has emerged as a key technique for adapting large pre-trained Decision Transformers (DTs) in offline Reinforcement Learning (RL), particularly in multi-task and few-shot settings. The Prompting Decision Transformer (PDT) enables task generalization via trajectory prompts sampled uniformly from expert demonstrations -- without accounting for prompt informativeness. In this work, we propose a bandit-based prompt-tuning method that learns to construct optimal trajectory prompts from demonstration data at inference time. We devise a structured bandit architecture operating in the trajectory prompt space, achieving linear rather than combinatorial scaling with prompt size. Additionally, we show that the pre-trained PDT itself can serve as a powerful feature extractor for the bandit, enabling efficient reward modeling across various environments. We theoretically establish regret bounds and demonstrate empirically that our method consistently enhances performance across a wide range of tasks, high-dimensional environments, and out-of-distribution scenarios, outperforming existing baselines in prompt tuning. Finn Rietz, Oleg Smirnov, Sara Karimi, Le-le Cao |
NeurIPS | 4 |
| 2025 | GenCeption: Evaluate vision LLMs with unlabeled unimodal data
Le-le Cao, Valentin Leonhard Buchner, Zineb Senane, Fangkai Yang |
Comput. Speech Lang. | 1 |
| 2025 | CompanyKG: A Large-Scale Heterogeneous Graph for Company Similarity QuantificationabstractIn the investment industry, it is often essential to carry out fine-grained company similarity quantification for a range of purposes, including market mapping, competitor analysis, and mergers and acquisitions. We propose and publish a knowledge graph, named CompanyKG, to represent and learn diverse company features and relations. Specifically, 1.17 million companies are represented as nodes enriched with company description embeddings; and 15 different inter-company relations result in 51.06 million weighted edges. To enable a comprehensive assessment of methods for company similarity quantification, we have devised and compiled three evaluation tasks with annotated test sets: similarity prediction, competitor retrieval and similarity ranking. We present extensive benchmarking results for 11 reproducible predictive methods categorized into three groups: node-only, edge-only, and node+edge. To the best of our knowledge, CompanyKG is the first large-scale heterogeneous graph dataset originating from a real-world investment platform, tailored for quantifying inter-company similarity Le-le Cao, Vilhelm von Ehrenheim, Mark Granroth-Wilding, Richard Anselmo Stahl, Andrew McCornack, Armin Catovic, Dhiana Deva Cavalcanti Rocha |
IEEE Trans. Big Data | 1 |
| 2024 | Beyond Gut Feel: Using Time Series Transformers to Find Investment Gems
Le-le Cao, Gustaf Halvardsson, Andrew McCornack, Vilhelm von Ehrenheim, Pawel Andrzej Herman |
ICANN (9) | 1 |
| 2024 | CompanyKG: A Large-Scale Heterogeneous Graph for Company Similarity QuantificationabstractThis paper presents CompanyKG (version 2), a large-scale heterogeneous graph developed for fine-grained company similarity quantification and relationship prediction, crucial for applications in the investment industry such as market mapping, competitor analysis, and mergers and acquisitions.CompanyKG comprises 1.17 million companies represented as graph nodes, enriched with company description embeddings, and 51.06 million weighted edges Le-le Cao, Vilhelm von Ehrenheim, Mark Granroth-Wilding, Richard Anselmo Stahl, Andrew McCornack, Armin Catovic, Dhiana Deva Cavalcanti Rocha |
KDD | 1 |
| 2024 | Self-Supervised Learning of Time Series Representation via Diffusion Process and Imputation-Interpolation-Forecasting MaskabstractTime Series Representation Learning (TSRL) focuses on generating informative representations for various Time Series (TS) modeling tasks. Traditional Self-Supervised Learning (SSL) methods in TSRL fall into four main categories: reconstructive, adversarial, contrastive, and predictive, each with a common challenge of sensitivity to noise and intricate data nuances. Recently, diffusion-based methods have shown advanced generative capabilities. However, they primarily target specific application scenarios like imputation and forecasting, leaving a gap in leveraging diffusion models for generic TSRL. Our work, Time Series Diffusion Embedding (TSDE), bridges this gap as the first diffusion-based SSL TSRL approach. TSDE segments TS data into observed and masked parts using an Imputation-Interpolation-Forecasting (IIF) mask. It applies a trainable embedding function, featuring dual-orthogonal Transformer encoders with a crossover mechanism, to the observed part. We train a reverse diffusion process conditioned on the embeddings, designed to predict noise added to the masked part. Extensive experiments demonstrate TSDE's superiority in imputation, interpolation, forecasting, anomaly detection, classification, and clustering. We also conduct an ablation study, present embedding visualizations, and compare inference speed, further substantiating TSDE's efficiency and validity in learning representations of TS data. Zineb Senane, Le-le Cao, Valentin Leonhard Buchner, Yusuke Tashiro, Lei You 0002, Pawel Andrzej Herman, Mats Nordahl, Ruibo Tu, Vilhelm von Ehrenheim |
KDD | 2 |
| 2023 | Measuring Acoustics with Collaborative Multiple AgentsabstractAs humans, we hear sound every second of our life. The sound we hear is often affected by the acoustics of the environment surrounding us. For example, a spacious hall leads to more reverberation. Room Impulse Responses (RIR) are commonly used to characterize environment acoustics as a function of the scene geometry, materials, and source/receiver locations. Traditionally, RIRs are measured by setting up a loudspeaker and microphone in the environment for all source/receiver locations, which is time-consuming and inefficient. We propose to let two robots measure the environment's acoustics by actively moving and emitting/receiving sweep signals. We also devise a collaborative multi-agent policy where these two robots are trained to explore the environment's acoustics while being rewarded for wide exploration and accurate prediction. We show that the robots learn to collaborate and move to explore environment acoustics while minimizing the prediction error. To the best of our knowledge, we present the very first problem formulation and solution to the task of collaborative environment acoustics measurements with multiple agents. Yinfeng Yu, Changan Chen, Le-le Cao, Fangkai Yang, Fuchun Sun 0001 |
IJCAI | 3 |
| 2023 | Echo-Enhanced Embodied Visual NavigationabstractVisual navigation involves a movable robotic agent striving to reach a point goal (target location) using vision sensory input. While navigation with ideal visibility has seen plenty of success, it becomes challenging in suboptimal visual conditions like poor illumination, where traditional approaches suffer from severe performance degradation. We propose E3VN (echo-enhanced embodied visual navigation) to effectively perceive the surroundings even under poor visibility to mitigate this problem. This is made possible by adopting an echoer that actively perceives the environment via auditory signals. E3VN models the robot agent as playing a cooperative Markov game with that echoer. The action policies of robot and echoer are jointly optimized to maximize the reward in a two-stream actor-critic architecture. During optimization, the reward is also adaptively decomposed into the robot and echoer parts. Our experiments and ablation studies show that E3VN is consistently effective and robust in point goal navigation tasks, especially under nonideal visibility. Yinfeng Yu, Le-le Cao, Fuchun Sun 0001, Chao Yang 0026, Huicheng Lai, Wenbing Huang 0001 |
Neural Comput. | 2 |
| 2022 | Pay Self-Attention to Audio-Visual Navigation
Yinfeng Yu, Le-le Cao, Fuchun Sun 0001 |
BMVC | 2 |
| 2022 | Simulation-Informed Revenue Extrapolation with Confidence Estimate for Scaleup Companies Using Scarce Time-Series DataabstractInvestment professionals rely on extrapolating company revenue into the future (i.e. revenue forecast) to approximate the valuation of scaleups (private companies in a high-growth stage) and inform their investment decision. This task is manual and empirical, leaving the forecast quality heavily dependent on the investment professionals' experiences and insights. Furthermore, financial data on scaleups is typically proprietary, costly and scarce, ruling out the wide adoption of data-driven approaches. To this end, we propose a simulation-informed revenue extrapolation (SiRE) algorithm that generates fine-grained long-term revenue predictions on small datasets and short time-series. SiRE models the revenue dynamics as a linear dynamical system (LDS), which is solved using the EM algorithm. The main innovation lies in how the noisy revenue measurements are obtained during training and inferencing. SiRE works for scaleups that operate in various sectors and provides confidence estimates. The quantitative experiments on two practical tasks show that SiRE significantly surpasses the baseline methods by a large margin. We also observe high performance when SiRE extrapolates long-term predictions from short time-series. The performance-efficiency balance and result explainability of SiRE are also validated empirically. Evaluated from the perspective of investment professionals, SiRE can precisely locate the scaleups that have a great potential return in 2 to 5 years. Furthermore, our qualitative inspection illustrates some advantageous attributes of the SiRE revenue forecasts. Le-le Cao, Sonja Horn, Vilhelm von Ehrenheim, Richard Anselmo Stahl, Henrik Landgren |
CIKM | 1 |
| 2022 | Multimodal Token Fusion for Vision TransformersabstractMany adaptations of transformers have emerged to address the single-modal vision tasks, where self-attention modules are stacked to handle input sources like images. Intuitively, feeding multiple modalities of data to vision transformers could improve the performance, yet the innermodal attentive weights may be diluted, which could thus greatly undermine the final performance. In this paper, we propose a multimodal token fusion method (TokenFusion), tailored for transformer-based vision tasks. To effectively fuse multiple modalities, TokenFusion dynamically detects uninformative tokens and substitute these tokens with projected and aggregated inter-modal features. Residual positional alignment is also adopted to enable explicit utilization of the inter-modal alignments after fusion. The design of TokenFusion allows the transformer to learn correlations among multimodal features, while the single-modal transformer architecture remains largely intact. Extensive experiments are conducted on a variety of homogeneous and heterogeneous modalities and demonstrate that TokenFusion surpasses state-of-the-art methods in three typical vision tasks: multimodal image-to-image translation, RGB-depth semantic segmentation, and 3D object detection with point cloud and images. Code will be released11https://github.com/huawei-noah/noah-research22https://gitee.com/mindspore/models/tree/master/research/cv/TokenFusion. Yikai Wang 0001, Xinghao Chen 0001, Le-le Cao, Wenbing Huang 0001, Fuchun Sun 0001, Yunhe Wang 0001 |
CVPR | 3 |
| 2022 | Bridged Transformer for Vision and Point Cloud 3D Object Detectionabstract3D object detection is a crucial research topic in computer vision, which usually uses 3D point clouds as input in conventional setups. Recently, there is a trend of leveraging multiple sources of input data, such as complementing the 3D point cloud with 2D images that often have richer color and fewer noises. However, due to the heterogeneous geometrics of the 2D and 3D representations, it prevents us from applying off-the-shelf neural networks to achieve multimodal fusion. To that end, we propose Bridged Transformer (BrT), an end-to-end architecture for 3D object detection. BrT is simple and effective, which learns to identify 3D and 2D object bounding boxes from both points and image patches. A key element of BrT lies in the utilization of object queries for bridging 3D and 2D spaces, which unifies different sources of data representations in Transformer. We adopt a form of feature aggregation realized by point-to-patch projections which further strengthen the interaction between images and points. Moreover, BrT works seamlessly for fusing the point cloud with multi-view images. We experimentally show that BrT surpasses state-of-the-art methods on SUN RGB-D and ScanNetV2 datasets. Yikai Wang 0001, TengQi Ye, Le-le Cao, Wenbing Huang 0001, Fuchun Sun 0001, Fengxiang He, Dacheng Tao |
CVPR | 3 |
| 2021 | PAUSE: Positive and Annealed Unlabeled Sentence EmbeddingabstractSentence embedding refers to a set of effective and versatile techniques for converting raw text into numerical vector representations that can be used in a wide range of natural language processing (NLP) applications.The majority of these techniques are either supervised or unsupervised.Compared to the unsupervised methods, the supervised ones make less assumptions about optimization objectives and usually achieve better results.However, the training requires a large amount of labeled sentence pairs, which is not available in many industrial scenarios.To that end, we propose a generic and end-to-end approach -PAUSE (Positive and Annealed Unlabeled Sentence Embedding), capable of learning high-quality sentence embeddings from a partially labeled dataset.We experimentally show that PAUSE achieves, and sometimes surpasses, state-ofthe-art results using only a small fraction of labeled sentence pairs on various benchmark tasks.When applied to a real industrial use case where labeled samples are scarce, PAUSE encourages us to extend our dataset without the burden of extensive manual annotation work. Le-le Cao, Emil Larsson, Vilhelm von Ehrenheim, Dhiana Deva Cavalcanti Rocha, Anna Martin, Sonja Horn |
EMNLP (1) | 1 |
| 2020 | Use All Your Skills, Not Only The Most Popular OnesabstractReinforcement Learning (RL) has shown promising results across various domains. However, applying it to develop gameplaying agents is challenging due to sparsity of extrinsic rewards, where agents get rewards from the environments only at the end of game levels. Previous works have shown that using intrinsic rewards is an effective way to deal with such cases. Intrinsic rewards allow to incorporate basic skills in agent policies to better generalize over various game levels. In a gameplay, it is common that certain actions (skills) are observed more often than others, which leads to a biased selection of actions. This problem boils down to a normalization issue in formulating the skill-based reward function. In this paper, we propose a novel solution to this problem by taking into account the frequency of all skills in the reward function. We show that our method improves the performance of agents by enabling them to select effective skills up to 2.5 times more frequently than that of the state-of-the-art in the context of the match-3 game Candy Crush Friends Saga. Francesco Lorenzo, Sahar Asadi, Alice Karnsund, Le-le Cao, Tianze Wang, Amir Hossein Payberah |
CoG | 4 |
| 2020 | Simple, Scalable, and Stable Variational Deep Clustering
Le-le Cao, Sahar Asadi, Wenfei Zhu, Christian Schmidli, Michael Sjöberg |
ECML/PKDD (1) | 1 |
| 2020 | Real-Time Recurrent Tactile Recognition: Momentum Batch-Sequential Echo State NetworksabstractTactile recognition aims at identifying target objects according to tactile sensory readings. Tactile data have two salient properties: 1) sequentially real-time and 2) temporally correlated, which essentially calls for a real-time (i.e., online fixed-budget) and recurrent recognition procedure. Based on an efficient and robust spatio-temporal feature representation for tactile sequences, we handle the problem of real-time recurrent tactile recognition by proposing a bounded online-sequential learning framework, and incorporates the strength of batch-regularization bootstrapping, bounded recursive reservoir, and momentum-based estimation. Experimental evaluations show that it outperforms the state-of-the-art methods by a large margin on test accuracy; and its training performance is superior to most compared models from aspects of average online training error, computational complexity, and storage efficiency. Le-le Cao, Fuchun Sun 0001, Kotagiri Ramamohanarao, Wenbing Huang 0001, Weihao Cheng 0001, Xiaolong Liu 0010 |
IEEE Trans. Syst. Man Cybern. Syst. | 1 |
| 2019 | Fine-Grained Road Mining from Satellite Images with Bilateral Xception and DeepLababstractWith the recent development of remote sensing and deep learning techniques, automatic and robust road extraction from satellite imaging data has become one of the most popular topics in both fields of Geographic Information System (GIS) and Computer Vision. Despite of the superior performance of Convolutional Neural Networks (DCNNs), a common problem of choosing between the classification and segmentation DCNNs still remains. By comparing two state-of-the-art baseline classification/segmentation DCNNs in several industrial application scenarios, we illustrate that their relative performance may vary, leading to different choices. Based on that observation, we propose a general fusion strategy that conveniently combines the strength of both classification and segmentation DCNNs using an end-to-end network architecture; this paradigm only requires pre-train segmentation/classification DCNNs once, which then can be reused in different road feature mining tasks. The task-specific experiments show that our fusion strategy guarantees superior results in all tested industrial scenarios. Le-le Cao |
IJCNN | 1 |
| 2017 | Fix-Budget and Recurrent Data Mining for Online Haptic Perception
Le-le Cao, Fuchun Sun 0001, Xiaolong Liu 0010, Wenbing Huang 0001, Weihao Cheng 0001, Kotagiri Ramamohanarao |
ICONIP (5) | 1 |
| 2017 | Advancing the incremental fusion of robotic sensory features using online multi-kernel extreme learning machine
Le-le Cao, Fuchun Sun 0001, Hongbo Li 0001, Wenbing Huang 0001 |
Frontiers Comput. Sci. | 1 |
| 2016 | Efficient Spatio-Temporal Tactile Object Recognition with Randomized Tiling Convolutional Networks in a Hierarchical Fusion StrategyabstractRobotic tactile recognition aims at identifying target objects or environments from tactile sensory readings. The advancement of unsupervised feature learning and biological tactile sensing inspire us proposing the model of 3T-RTCN that performs spatio-temporal feature representation and fusion for tactile recognition. It decomposes tactile data into spatial and temporal threads, and incorporates the strength of randomized tiling convolutional networks. Experimental evaluations show that it outperforms some state-of-the-art methods with a large margin regarding recognition accuracy, robustness, and fault-tolerance; we also achieve an order-of-magnitude speedup over equivalent networks with pretraining and finetuning. Practical suggestions and hints are summarized in the end for effectively handling the tactile data. Le-le Cao, Kotagiri Ramamohanarao, Fuchun Sun 0001, Hongbo Li 0001, Wenbing Huang 0001, Zay Maung Maung Aye |
AAAI | 1 |
| 2016 | Sparse Coding and Dictionary Learning with Linear Dynamical SystemsabstractLinear Dynamical Systems (LDSs) are the fundamental tools for encoding spatio-temporal data in various disciplines. To enhance the performance of LDSs, in this paper, we address the challenging issue of performing sparse coding on the space of LDSs, where both data and dictionary atoms are LDSs. Rather than approximate the extended observability with a finite-order matrix, we represent the space of LDSs by an infinite Grassmannian consisting of the orthonormalized extended observability subspaces. Via a homeomorphic mapping, such Grassmannian is embedded into the space of symmetric matrices, where a tractable objective function can be derived for sparse coding. Then, we propose an efficient method to learn the system parameters of the dictionary atoms explicitly, by imposing the symmetric constraint to the transition matrices of the data and dictionary systems. Moreover, we combine the state covariance into the algorithm formulation, thus further promoting the performance of the models with symmetric transition matrices. Comparative experimental evaluations reveal the superior performance of proposed methods on various tasks including video classification and tactile recognition. Wenbing Huang 0001, Fuchun Sun 0001, Le-le Cao, Deli Zhao, Huaping Liu 0001, Mehrtash Harandi |
CVPR | 3 |
| 2016 | Learning Stable Linear Dynamical Systems with the Weighted Least Square Method
Wenbing Huang 0001, Le-le Cao, Fuchun Sun 0001, Deli Zhao, Huaping Liu 0001 |
IJCAI | 2 |
| 2016 | Tactile sequence based object categorization: A Bag of features modeled by Linear Dynamic System with Symmetric Transition MatrixabstractIn this paper, we propose a novel categorization framework to recognize tactile sequences based on two particular properties of the tactile data. For the first one, tactile sequences are spatio-temporal data which is sequential and dynamic, depicting the process of grasping an object in different grasping stages; therefore, it is reasonable to discover the dynamical pattern by modeling tactile data as integral sequences rather than individual frames. For the second one, a tactile sequence contains various dynamical patterns in different stages of the grasping process; therefore, we decompose the whole sequence into multiple mini-sequences so as to enhance feature resolution. To address both properties in our framework, we take advantage of a Bag-of-System model using parameters of the Linear Dynamic System (LDS) as feature descriptors. Moreover, we employ the LDS with Symmetric Transition matrix (LDSST) rather than the original LDS as the building-block in order to obtain accurate codewords of the codebook of the Bag-of-System. The performance of our framework is evaluated on six real-world databases of three groups. Our experiments show that classification using LDSST is better than the original LDS, and the decomposition of tactile sequences does improve the accuracy of classification. The experiment results also show the superiority of our framework in comparison with other state-of-the-art sequence classifiers. Fuchun Sun 0001, Wenbing Huang 0001, Le-le Cao, Bin Fang 0003 |
IJCNN | 4 |
| 2016 | A Precise and Robust Clustering Approach Using Homophilic Degrees of Graph Kernel
Deli Zhao, Le-le Cao, Fuchun Sun 0001 |
PAKDD (2) | 3 |
| 2016 | Building feature space of extreme learning machine with sparse denoising stacked-autoencoder
Le-le Cao, Wenbing Huang 0001, Fuchun Sun 0001 |
Neurocomputing | 1 |
| 2014 | Optimization-Based Extreme Learning Machine with Multi-kernel Learning Approach for ClassificationabstractThe optimization method based extreme learning machine (optimization-based ELM) is generalized from single-hidden-layer feed-forward neural networks (SLFNs) by making use of kernels instead of neuron-alike hidden nodes. This approach is known for its high scalability, low computational complexity, and mild optimization constrains. The multi-kernel learning (MKL) framework Simple MKL iteratively determines the combination of kernels by gradient descent wrapping a standard support vector machine (SVM) solver. Simple MKL can be applied to many kinds of supervised learning problems to receive a more stable performance with rapid convergence speed. This paper proposes a new approach: MK-ELM (multi-kernel extreme learning machine) that applies Simple MKL framework to the optimization-based ELM algorithm. The performance analysis on binary classification problems with various scales shows that MK-ELM tends to achieve the best generalization performance as well as being the most insensitive to parameters comparing to optimization-based ELM and Simple MKL. As a result, MK-ELM can be implemented in real applications easily. Le-le Cao, Wenbing Huang 0001, Fuchun Sun 0001 |
ICPR | 1 |