VLDB 2026 Research / reviewers in the wild / expert
Le Yang 0007
dblp:79/2888-7
· DBLP profile ↗
30ranked-venue papers
11as first author
25since 2021 · last 2026
0000-0001-8379-4915ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 18 · 7 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 4 first-author · 12 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Security and privacy · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Privacy on the Fly: A Predictive Adversarial Transformation Network for Mobile Sensor DataabstractMobile motion sensors such as accelerometers and gyroscopes are now ubiquitously accessible by third-party apps via standard APIs. While enabling rich functionalities like activity recognition and step counting, this openness has also enabled unregulated inference of sensitive user traits, such as gender, age, and even identity, without user consent. Existing privacy-preserving techniques, such as GAN-based obfuscation or differential privacy, typically require access to the full input sequence, introducing latency that is incompatible with real-time scenarios. Worse, they tend to distort temporal and semantic patterns, degrading the utility of the data for benign tasks like activity recognition. To address these limitations, we propose the Predictive Adversarial Transformation Network (PATN), a real-time privacy-preserving framework that leverages historical signals to generate adversarial perturbations proactively. The perturbations are applied immediately upon data acquisition, enabling continuous protection without disrupting application functionality. Experiments on two datasets demonstrate that PATN substantially degrades the performance of privacy inference models, achieving Attack Success Rate (ASR) of 40.11% and 44.65% (reducing inference accuracy to near-random) and increasing the Equal Error Rate (EER) from 8.30% and 7.56% to 41.65% and 46.22%. On ASR, PATN outperforms baseline methods by 16.16% and 31.96%, respectively. Tianle Song, Chenhao Lin, Zhengyu Zhao 0001, Le Yang 0007, Chao Shen 0001 |
AAAI | 7 |
| 2025 | Nullu: Mitigating Object Hallucinations in Large Vision-Language Models via HalluSpace ProjectionabstractRecent studies have shown that large vision-language models (LVLMs) often suffer from the issue of object hallucinations (OH). To mitigate this issue, we introduce an efficient method that edits the model weights based on an unsafe subspace, which we call HalluSpace in this paper. With truthful and hallucinated text prompts accompanying the visual content as inputs, the HalluSpace can be identified by extracting the hallucinated embedding features and removing the truthful representations in LVLMs. By orthog-onalizing the model weights, input features will be projected into the Null space of the HalluSpace to reduce OH, based on which we name our method Nullu. We reveal that Hal-luSpaces generally contain prior information in the large language models (LLMs) applied to build LVLMs, which have been shown as essential causes of OH in previous studies. Therefore, null space projection suppresses the LLMs’ priors to filter out the hallucinated features, resulting in contextually accurate outputs. Experiments show that our method can effectively mitigate OH across different LVLM families without extra inference costs and also show strong performance in general LVLM benchmarks. Code is released at https://github.com/Ziwei-Zheng/Nullu. Le Yang 0007, Ziwei Zheng, Boxu Chen, Zhengyu Zhao 0001, Chenhao Lin, Chao Shen 0001 |
CVPR | 1 |
| 2025 | D3: Training-Free AI-Generated Video Detection Using Second-Order Features
Chende Zheng, Ruiqi Suo, Chenhao Lin, Zhengyu Zhao 0001, Le Yang 0007, Shuai Liu 0016, Cong Wang 0001, Chao Shen 0001 |
ICCV | 5 |
| 2025 | Deep Learning Based Topography Aware Gas Source Localization with Mobile RobotabstractGas source localization in complex environments is critical for applications such as environmental monitoring, industrial safety, and disaster response. Traditional methods often struggle with the challenges posed by a lack of environmental topography integration, especially when interactions between wind and obstacles distort gas dispersion patterns. In this paper, we propose a deep learning-based approach, which leverages spatial context and environmental mapping to enhance gas source localization. By integrating Simultaneous Localization and Mapping (SLAM) with a U-Net-based model, our method predicts the likelihood of gas source locations by analyzing gas sensor data, wind flow, and topography of the environment represented by a 2D occupancy map. We demonstrate the efficacy of our approach using a wheeled robot equipped with a photoionization detector, a LIDAR, and an anemometer, in various scenarios with dynamic wind fields and multiple obstacles. The results show that our approach can robustly locate gas sources, even in challenging environments with fluctuating wind directions, outperforming conventional methods by utilizing topography contextual information. This study underscores the importance of topographical context in gas source localization and offers a flexible and robust solution for real-world applications. Data and code are publicly available. Changhao Tian, Annan Wang, Han Fan, Thomas Wiedemann 0002, Le Yang 0007, Weisi Lin, Achim J. Lilienthal |
ICRA | 6 |
| 2025 | LVLM-FDA: Protecting Large Vision-Language Models via Fast Detection of Malicious Attempts
Boxu Chen, Le Yang 0007, Ziwei Zheng, Cong Wang 0001, Qian Wang 0002, Chao Shen 0001 |
KSEM (1) | 3 |
| 2025 | Transfer Attack for Bad and Good: Explain and Boost Adversarial Transferability across Multimodal Large Language ModelsabstractMultimodal Large Language Models (MLLMs) demonstrate exceptional performance in cross-modality interaction, yet they also suffer adversarial vulnerabilities. In particular, the transferability of adversarial examples remains an ongoing challenge. In this paper, we specifically analyze the manifestation of adversarial transferability among MLLMs and identify the key factors that influence this characteristic. We discover that the transferability of MLLMs exists in cross-LLM scenarios with the same vision encoder and indicate two key Factors that may influence transferability. We provide two semantic-level data augmentation methods, Adding Image Patch (AIP) and Typography Augment Transferability Method (TATM), which boost the transferability of adversarial examples across MLLMs. To explore the potential impact in the real world, we utilize two tasks that can have both negative and positive societal impacts: 1. Harmful Content Insertion and 2. Information Protection. Hao Cheng 0015, Erjia Xiao, Jiayan Yang, Jinhao Duan, Yichi Wang 0002, Jiahang Cao, Qiang Zhang 0029, Le Yang 0007, Kaidi Xu, Jindong Gu, Renjing Xu |
ACM Multimedia | 8 |
| 2025 | Jailbreak-AudioBench: In-Depth Evaluation and Analysis of Jailbreak Threats for Large Audio Language ModelsabstractLarge Language Models (LLMs) demonstrate impressive zero-shot performance across a wide range of natural language processing tasks. Integrating various modality encoders further expands their capabilities, giving rise to Multimodal Large Language Models (MLLMs) that process not only text but also visual and auditory modality inputs. However, these advanced capabilities may also pose significant safety problems, as models can be exploited to generate harmful or inappropriate content through jailbreak attack. While prior work has extensively explored how manipulating textual or visual modality inputs can circumvent safeguards in LLMs and MLLMs, the vulnerability of audio-specific Jailbreak on Large Audio-Language Models (LALMs) remains largely underexplored. To address this gap, we introduce \textbf{Jailbreak-AudioBench}, which consists of the Toolbox, curated Dataset, and comprehensive Benchmark. The Toolbox supports not only text-to-audio conversion but also various editing techniques for injecting audio hidden semantics. The curated Dataset provides diverse explicit and implicit jailbreak audio examples in both original and edited forms. Utilizing this dataset, we evaluate multiple state-of-the-art LALMs and establish the most comprehensive Jailbreak benchmark to date for audio modality. Finally, Jailbreak-AudioBench establishes a foundation for advancing future research on LALMs safety alignment by enabling the in-depth exposure of more powerful jailbreak threats, such as query-based audio editing, and by facilitating the development of effective defense mechanisms. Hao Cheng 0015, Erjia Xiao, Yichi Wang 0002, Le Yang 0007, Chao Shen 0001, Philip Torr 0001, Jindong Gu, Renjing Xu |
NeurIPS | 5 |
| 2025 | Artificial intelligence security and privacy: a surveyabstractAbstract Artificial intelligence (AI) is revolutionizing both industries and reshaping the global economy. However, the rapid advancement of AI technologies brings significant security and privacy challenges. Recent incidents highlight vulnerabilities in AI systems, such as data leakage and malicious code injection, leading to severe financial losses and privacy breaches. Although existing studies have discussed specific security threats, they often lack detailed granularity and cover a limited scope. In this survey, we fill this gap by systematically categorizing and analyzing the threats and countermeasures in AI systems, which span both the training and inference stages, encompass centralized and distributed settings, and address both conventional and foundation AI models. By reviewing existing literature, we aim to provide AI researchers and practitioners with a thorough understanding of system vulnerabilities and current countermeasures. We hope to inspire further research into robust solutions, ultimately contributing to the development of resilient AI technologies. Xinlei He 0001, Guowen Xu, Xingshuo Han, Qian Wang 0002, Lingchen Zhao, Chao Shen 0001, Chenhao Lin, Zhengyu Zhao 0001, Qian Li 0024, Le Yang 0007, Shouling Ji, Shaofeng Li 0001, Haojin Zhu, Zhibo Wang 0001, Tianqing Zhu, Qi Li 0002, Chaoxiang He, Hongsheng Hu, Shuo Wang 0012, Shifeng Sun 0001, Hongwei Yao, Qinyu Zhang 0001, Kai Chen 0012, Yue Zhao 0027, Hongwei Li 0001, Xinyi Huang 0001, Dengguo Feng |
Sci. China Inf. Sci. | 10 |
| 2025 | Single-layer feedforward neural networks with dynamic width for domain adaptation
Le Yang 0007, Zelin Yang, Fan Li 0003, C. L. Philip Chen |
Sci. China Inf. Sci. | 1 |
| 2025 | Data-Centric Robust Training for Defending Against Transfer-Based Adversarial AttacksabstractTransfer-based adversarial attacks pose a severe threat to real-world deep learning systems since they do not require access to target models. Adversarial training (AT), which is recognized as the most effective defense against white-box attacks, also ensures high robustness against (black-box) transfer-based attacks. However, AT suffers from significant computational overhead because it repeatedly generates adversarial examples (AEs) throughout the entire training process. In this paper, we demonstrate that such repeated generation is unnecessary to achieve robustness against transfer-based attacks. Instead, pre-generating AEs all at once before training is sufficient, as proposed in our new defense paradigm called Data-Centric Robust Training (DCRT). DCRT employs clean data augmentation and adversarial data augmentation techniques to enhance the dataset before training. Our experimental results show that DCRT outperforms widely-used AT techniques (e.g., PGD-AT, TRADES, EAT, and FAT) in terms of transfer-based black-box robustness and even surpasses the top-1 defense on RobustBench when combined with common model-centric techniques. We also highlight additional benefits of DCRT, such as improved training efficiency and class-wise fairness.Our code will be available on GitHub. Yulong Yang 0002, Ruiqi Cao, Qiwei Tian, Chenhao Lin, Zhengyu Zhao 0001, Qian Li 0024, Le Yang 0007, Hongshan Yang, Chao Shen 0001 |
IEEE Trans. Inf. Forensics Secur. | 8 |
| 2024 | Unveiling Typographic Deceptions: Insights of the Typographic Vulnerability in Large Vision-Language Models
Hao Cheng 0015, Erjia Xiao, Jindong Gu, Le Yang 0007, Jinhao Duan, Jize Zhang, Jiahang Cao, Kaidi Xu, Renjing Xu |
ECCV (59) | 4 |
| 2024 | DyFADet: Dynamic Feature Aggregation for Temporal Action Detection
Le Yang 0007, Ziwei Zheng, Yizeng Han, Hao Cheng 0015, Shiji Song, Gao Huang 0001, Fan Li 0003 |
ECCV (46) | 1 |
| 2024 | Fine-Grained Dynamic Network for Generic Event Boundary Detection
Ziwei Zheng, Lijun He 0001, Le Yang 0007, Fan Li 0003 |
ECCV (43) | 3 |
| 2024 | Rethinking the Architecture Design for Efficient Generic Event Boundary Detection
Ziwei Zheng, Zechuan Zhang, Yulin Wang 0002, Shiji Song, Gao Huang 0001, Le Yang 0007 |
ACM Multimedia | 6 |
| 2024 | Energy-based Active Learning for Bringing Beam-induced Domain Gap for 3D Object DetectionabstractIn many real-world applications, 16-beam LiDAR-based 3D object detection (3DOD) is indispensable in scene understanding. However, the absence of well-labeled large-scale 16-beam LiDAR datasets impedes the development of these 3DOD methods. To avoid annotation costs in developing datasets, we proposed an energy-based active learning method for cross-beam domain adaptation, which effectively transfers the knowledge from the existing well-labeled 64-beam counterpart. Specifically, the cross-beam domain gap between the source (64-beam) and the target (16-beam) domain is reduced by aligning the deep features based on an energy-based feature-matching loss term during training. Moreover, the proposed energy-based active learning method enables the sampling strategy to shed light on selecting the most valuable 16-beam target samples to be manually labeled, which are then added to the training set. Experimental results show that our method can effectively transfer the knowledge from the 64-beam domain to the 16-beam one, and successfully learns a high-performance 16-beam 3DOD model with only a small portion of unlabeled data to annotate. Le Yang 0007, Yixuan Yan, Hao Cheng 0015, Fan Li 0003 |
MobiCom | 1 |
| 2024 | Fixing Overconfidence in Dynamic Neural NetworksabstractDynamic neural networks are a recent technique that promises a remedy for the increasing size of modern deep learning models by dynamically adapting their computational cost to the difficulty of the inputs. In this way, the model can adjust to a limited computational budget. However, the poor quality of uncertainty estimates in deep learning models makes it difficult to distinguish between hard and easy samples. To address this challenge, we present a computationally efficient approach for post-hoc uncertainty quantification in dynamic neural networks. We show that adequately quantifying and accounting for both aleatoric and epistemic uncertainty through a probabilistic treatment of the last layers improves the predictive performance and aids decision-making when determining the computational budget. In the experiments, we show improvements on CIFAR100, ImageNet, and Caltech-256 in terms of accuracy, capturing uncertainty, and calibration error. Lassi Meronen, Martin Trapp 0001, Andrea Pilzer, Le Yang 0007, Arno Solin |
WACV | 4 |
| 2024 | Dynamic Spatial Focus for Efficient Compressed Video Action RecognitionabstractRecent years have witnessed a growing interest in compressed video action recognition due to the rapid growth of online videos. It remarkably reduces the storage by replacing raw videos with sparsely sampled RGB frames and other compressed motion cues (motion vectors and residuals). However, existing compressed video action recognition methods face two main issues: First, the inefficiency caused by the usage of coarse-level information under full resolution, and second, the disturbing due to the noisy dynamics in motion vectors. To address the two issues, this paper proposes a dynamic spatial focus method for efficient compressed video action recognition (CoViFocus). Specifically, we first use a light-weighted two-stream architecture to localize the task-relevant patches for both the RGB frames and motion vectors. Then the selected patch pair will be processed by a high-capacity two-stream deep model for the final prediction. Such a patch selection strategy crops out the irrelevant motion noise in motion vectors, as well as reduces the spatial redundancy of the inputs, leading to the high efficiency of our method in the compressed domain. Moreover, we found that the motion vectors can help our method to address the possibly happened static-issue, which means that the focus patches get stuck at some regions related to static objects rather than target actions, which further improves our method. Extensive results on both the HMDB-51 and UCF-101 datasets demonstrate the effectiveness and efficiency of our method in compressed video action recognition tasks. Ziwei Zheng, Le Yang 0007, Yulin Wang 0002, Miao Zhang 0041, Lijun He 0001, Gao Huang 0001, Fan Li 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | OStr-DARTS: Differentiable Neural Architecture Search Based on Operation StrengthabstractDifferentiable architecture search (DARTS) has emerged as a promising technique for effective neural architecture search, and it mainly contains two steps to find the high-performance architecture. First, the DARTS supernet that consists of mixed operations will be optimized via gradient descent. Second, the final architecture will be built by the selected operations that contribute the most to the supernet. Although DARTS improves the efficiency of neural architecture search (NAS), it suffers from the well-known degeneration issue which can lead to deteriorating architectures. Existing works mainly attribute the degeneration issue to the failure of its supernet optimization, while little attention has been paid to the selection method. In this article, we cease to apply the widely-used magnitude-based selection method and propose a novel criterion based on operation strength that estimates the importance of an operation by its effect on the final loss. We show that the degeneration issue can be effectively addressed by using the proposed criterion without any modification of supernet optimization, indicating that the magnitude-based selection method can be a critical reason for the instability of DARTS. The experiments on NAS-Bench-201 and DARTS search spaces show the effectiveness of our method. Le Yang 0007, Ziwei Zheng, Yizeng Han, Shiji Song, Gao Huang 0001, Fan Li 0003 |
IEEE Trans. Cybern. | 1 |
| 2022 | Dynamic Neural Networks: A SurveyabstractDynamic neural network is an emerging research topic in deep learning. Compared to static models which have fixed computational graphs and parameters at the inference stage, dynamic networks can adapt their structures or parameters to different inputs, leading to notable advantages in terms of accuracy, computational efficiency, adaptiveness, etc. In this survey, we comprehensively review this rapidly developing area by dividing dynamic networks into three main categories: 1) sample-wise dynamic models that process each sample with data-dependent architectures or parameters; 2) spatial-wise dynamic networks that conduct adaptive computation with respect to different spatial locations of image data; and 3) temporal-wise dynamic models that perform adaptive inference along the temporal dimension for sequential data such as videos and texts. The important research problems of dynamic networks, e.g., architecture design, decision making scheme, optimization technique and applications, are reviewed systematically. Finally, we discuss the open problems in this field together with interesting future research directions. Yizeng Han, Gao Huang 0001, Shiji Song, Le Yang 0007, Yulin Wang 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2021 | CondenseNet V2: Sparse Feature Reactivation for Deep NetworksabstractReusing features in deep networks through dense connectivity is an effective way to achieve high computational efficiency. The recent proposed CondenseNet [14] has shown that this mechanism can be further improved if redundant features are removed. In this paper, we propose an alternative approach named sparse feature reactivation (SFR), aiming at actively increasing the utility of features for reusing. In the proposed network, named CondenseNetV2, each layer can simultaneously learn to 1) selectively reuse a set of most important features from preceding layers; and 2) actively update a set of preceding features to increase their utility for later layers. Our experiments show that the proposed models achieve promising performance on image classification (ImageNet and CIFAR) and object detection (MS COCO) in terms of both theoretical efficiency and practical speed. Le Yang 0007, Haojun Jiang, Ruojin Cai, Yulin Wang 0002, Shiji Song, Gao Huang 0001, Qi Tian 0001 |
CVPR | 1 |
| 2021 | Revisiting Locally Supervised Learning: an Alternative to End-to-end Training
Yulin Wang 0002, Zanlin Ni, Shiji Song, Le Yang 0007, Gao Huang 0001 |
ICLR | 4 |
| 2021 | Large scale air pollution prediction with deep convolutional networks
Gao Huang 0001, Chunjiang Ge, Tianyu Xiong, Shiji Song, Le Yang 0007, Baoxian Liu, Wenjun Yin, Cheng Wu 0002 |
Sci. China Inf. Sci. | 5 |
| 2021 | Discriminative Dimension Reduction via Maximin Separation Probability AnalysisabstractIn this paper, we propose a novel discriminative dimension reduction (DR) method, maximin separation probability analysis (MSPA), which maximizes the minimum separation probability of all classes in the reduced low-dimensional subspace. Separation probability is a novel class separability measure, which gives a lower bound of the generalization accuracy for a learned linear classifier in a binary classification problem. The proposed MSPA duly considers the separation of all class pairs in multiclass linear discriminant analysis (LDA) and thus improves the subsequent classification performance. DR via MSPA leads to a nonconvex optimization problem. We develop an algorithm to solve the problem and the global optimal solution can be found by converting the original problem into a series of second-order cone programming problems. A low-computational cost extension and a non-LDA with kernel mapping of MSPA are also provided in this paper. The experimental results on 14 real-world datasets show our methods are superior to other state-of-the-art algorithms in discriminative DR tasks. Le Yang 0007, Shiji Song, Shuang Li 0008, Yiming Chen 0005, C. L. Philip Chen |
IEEE Trans. Cybern. | 1 |
| 2021 | Spatially Adaptive Feature Refinement for Efficient InferenceabstractSpatial redundancy commonly exists in the learned representations of convolutional neural networks (CNNs), leading to unnecessary computation on high-resolution features. In this paper, we propose a novel Spatially Adaptive feature Refinement (SAR) approach to reduce such superfluous computation. It performs efficient inference by adaptively fusing information from two branches: one conducts standard convolution on input features at a lower spatial resolution, and the other one selectively refines a set of regions at the original resolution. The two branches complement each other in feature learning, and both of them evoke much less computation than standard convolution. SAR is a flexible method that can be conveniently plugged into existing CNNs to establish models with reduced spatial redundancy. Experiments on CIFAR and ImageNet classification, COCO object detection and PASCAL VOC semantic segmentation tasks validate that the proposed SAR can consistently improve the network performance and efficiency. Notably, our results show that SAR only refines less than 40% of the regions in the feature representations of a ResNet for 97% of the samples in the validation set of ImageNet to achieve comparable accuracy with the original model, revealing the high computational redundancy in the spatial dimension of CNNs. Yizeng Han, Gao Huang 0001, Shiji Song, Le Yang 0007, Haojun Jiang |
IEEE Trans. Image Process. | 4 |
| 2021 | Graph Embedding-Based Dimension Reduction With Extreme Learning MachineabstractDimension reduction (DR)-based on extreme learning machine auto-encoder (ELM-AE) has achieved many successes in recent years. By minimizing the self-reconstruction error, the ELM-AE-based DR algorithms learn the compressed representations which facilitate the subsequent classification. However, the existing ELM-AEs only consider the DR problem in an unsupervised manner and ignore the valuable supervised information when these information is available. To find discriminative features of the original data, in this paper, we propose a graph embedding-based DR framework with ELM (GDR-ELM) for DR problems. Instead of self-reconstruction, the proposed GDR-ELM reconstructs all samples according to the weights in a graph matrix containing the supervised information. Furthermore, GDR-ELM can be stacked as building blocks to construct a multilayer framework like other ELM-AEs for more complicated representation learning tasks. Experiments on various datasets demonstrate the effectiveness of the proposed GDR-ELM and its multilayer framework. Le Yang 0007, Shiji Song, Shuang Li 0008, Yiming Chen 0005, Gao Huang 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 1 |
| 2020 | Resolution Adaptive Networks for Efficient InferenceabstractAdaptive inference is an effective mechanism to achieve a dynamic tradeoff between accuracy and computational cost in deep networks. Existing works mainly exploit architecture redundancy in network depth or width. In this paper, we focus on spatial redundancy of input samples and propose a novel Resolution Adaptive Network (RANet), which is inspired by the intuition that low-resolution representations are sufficient for classifying “easy” inputs containing large objects with prototypical features, while only some “hard” samples need spatially detailed information. In RANet, the input images are first routed to a lightweight sub-network that efficiently extracts low-resolution representations, and those samples with high prediction confidence will exit early from the network without being further processed. Meanwhile, high-resolution paths in the network maintain the capability to recognize the “hard” samples. Therefore, RANet can effectively reduce the spatial redundancy involved in inferring high-resolution inputs. Empirically, we demonstrate the effectiveness of the proposed RANet on the CIFAR-10, CIFAR-100 and ImageNet datasets in both the anytime prediction setting and the budgeted batch classification setting. Le Yang 0007, Yizeng Han, Shiji Song, Jifeng Dai, Gao Huang 0001 |
CVPR | 1 |
| 2020 | Glance and Focus: a Dynamic Approach to Reducing Spatial Redundancy in Image ClassificationabstractThe accuracy of deep convolutional neural networks (CNNs) generally improves when fueled with high resolution images. However, this often comes at a high computational cost and high memory footprint. Inspired by the fact that not all regions in an image are task-relevant, we propose a novel framework that performs efficient image classification by processing a sequence of relatively small inputs, which are strategically selected from the original image with reinforcement learning. Such a dynamic decision process naturally facilitates adaptive inference at test time, i.e., it can be terminated once the model is sufficiently confident about its prediction and thus avoids further redundant computation. Notably, our framework is general and flexible as it is compatible with most of the state-of-the-art light-weighted CNNs (such as MobileNets, EfficientNets and RegNets), which can be conveniently deployed as the backbone feature extractor. Experiments on ImageNet show that our method consistently improves the computational efficiency of a wide variety of deep models. For example, it further reduces the average latency of the highly efficient MobileNet-V3 on an iPhone XS Max by 20% without sacrificing accuracy. Code and pre-trained models are available at https://github.com/blackfeather-wang/GFNet-Pytorch. Yulin Wang 0002, Kangchen Lv, Rui Huang 0012, Shiji Song, Le Yang 0007, Gao Huang 0001 |
NeurIPS | 5 |
| 2019 | Domain Space Transfer Extreme Learning Machine for Domain AdaptationabstractExtreme learning machine (ELM) has been applied in a wide range of classification and regression problems due to its high accuracy and efficiency. However, ELM can only deal with cases where training and testing data are from identical distribution, while in real world situations, this assumption is often violated. As a result, ELM performs poorly in domain adaptation problems, in which the training data (source domain) and testing data (target domain) are differently distributed but somehow related. In this paper, an ELM-based space learning algorithm, domain space transfer ELM (DST-ELM), is developed to deal with unsupervised domain adaptation problems. To be specific, through DST-ELM, the source and target data are reconstructed in a domain invariant space with target data labels unavailable. Two goals are achieved simultaneously. One is that, the target data are input into an ELM-based feature space learning network, and the output is supposed to approximate the input such that the target domain structural knowledge and the intrinsic discriminative information can be preserved as much as possible. The other one is that, the source data are projected into the same space as the target data and the distribution distance between the two domains is minimized in the space. This unsupervised feature transformation network is followed by an adaptive ELM classifier which is trained from the transferred labeled source samples, and is used for target data label prediction. Moreover, the ELMs in the proposed method, including both the space learning ELM and the classifier, require just a small number of hidden nodes, thus maintaining low computation complexity. Extensive experiments on real-world image and text datasets are conducted and verify that our approach outperforms several existing domain adaptation methods in terms of accuracy while maintaining high efficiency. Yiming Chen 0005, Shiji Song, Shuang Li 0008, Le Yang 0007, Cheng Wu 0002 |
IEEE Trans. Cybern. | 4 |
| 2019 | Nonparametric Dimension Reduction via Maximizing Pairwise Separation ProbabilityabstractIn this brief, we propose a novel nonparametric supervised linear dimension reduction (SLDR) algorithm that extracts the features by maximizing the pairwise separation probability. The separation probability, as a new class separability measure, describes the generalization accuracy when we use the obtained features to train a linear classifier. Obtaining high-quality features, the proposed method avoids the overlaps between classes that are close to each other in the input space and improves the subsequent classification performance. Experiments on benchmark data sets show the superiority of the proposed algorithm over some other state-of-the-art SLDR methods. Le Yang 0007, Shiji Song, Yanshang Gong, Gao Huang 0001, Cheng Wu 0002 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2018 | Transductive Transfer Learning Based on Broad Learning SystemabstractThe latest proposed Broad Learning System (BLS) demonstrates an efficient and effective learning capability in many machine learning problems. In this paper, we apply the BLS to address transductive transfer learning problems, where the training (source) and test (target) data are drawn from the different but related distributions, which is a.k.a domain adaptation. We aim at learning from source data a well performing classifier on a different (but related) target data. A unified domain adaptation framework based on the BLS is developed for improving its transfer learning capability without loss of the computational efficiency. Two algorithms including BLS based source domain adaptation (BLS-SDA) and BLS based target domain adaptation (BLS-TDA) are proposed under this framework. Experiments on benchmark datasets show that our approach outperforms several existing domain adaptation methods while maintains high efficiency. Le Yang 0007, Shiji Song, C. L. Philip Chen |
SMC | 1 |