Fei Chao 0001

dblp:118/5221-1 · DBLP profile ↗
← Back
120ranked-venue papers
8as first author
86since 2021 · last 2027
0000-0002-6928-2638ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 89 · 6 first-author · 63 since 2021Graphics, computer vision, multimedia, augmented reality and games · 30 · 1 first-author · 29 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 5 since 2021Databases, data management, data science and information retrieval · 5 · 1 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 4 · 1 first-author · 1 since 2021Systems, architecture and hardware · 3 · 1 first-author · 2 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2027 Unified-width adaptive dynamic network for all-in-one image restoration
Yimin Xu, Chunmei Yuan, Yunshan Zhong, Fei Chao 0001
Inf. Sci.4
2026 Profiling-Free Mixed-Precision Quantization for MoE LLMs via Fuzzy Rule Interpolation
abstract
Huachen Qi, Ruiyu Zhuo, Bowen Shi, Xiang Chang, Fei Chao, Changjing Shang, Qiang Shen. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Huachen Qi, Ruiyu Zhuo, Xiang Chang, Fei Chao 0001, Changjing Shang, Qiang Shen 0001
ACL (1)5
2026 CCF: A Context Compression Framework for Efficient Long-Sequence Language Modeling
Bangcheng Sun, Fei Chao 0001
ICPR (15)2
2026 Lab-DN: Dual-branch lightweight network for Lab color space shadow removal
Haocheng Chu, Fei Chao 0001, Xiang Chang, Changjing Shang, Qiang Shen 0001
Expert Syst. Appl.3
2026 Dynamic trajectory diffusion model for all-in-one image restoration
Yimin Xu, Yunshan Zhong, Fei Chao 0001
Expert Syst. Appl.3
2026 Good Performance Estimation Strategies are All You Need in Neural Architecture Search
abstract
Recent advances in Neural Architecture Search (NAS) are essentially attributed to Performance Estimation (PE), i.e., a method aims to effectively estimate an architecture. Meanwhile, Kendall's $\tau$τ is well recognized as the principled evaluation criteria for PE strategies in the literature. We argue that Kendall's $\tau$τ is not the optimal solution. Through extensive experiments and theoretical analysis, we take the initiative to reveal the problem behind the Kendall's $\tau$τ and propose a novel criterion named Minimum Keeping Ratio (MKR), which is closely connected to the final performance of NAS. It allows us to compare different PE approaches in a unified perspective, and use effective ablation studies to verify common beliefs and key differences of PE strategies. Based on the findings from MKR, we are able to derive a simple NAS method by integrating different PE strategies with random sampling. Such a method shows very strong performance in efficiency and effectiveness through extensive experiments on different challenging benchmarks. In particular, our simple random sampling NAS finds the optimal architecture in NASbenchMacro, NASbench201, and NASbench301. It is also well generalized to different search spaces (MobileNet) and tasks (semantic segmentation), finding an architecture surpasses the previous state-of-the-art architectures by 4.25 mIoU under $ 600M$600M FLOPs on ADE20K. Codes are available at https://anonymous.4open.science/r/Anonymization11264.
Xiawu Zheng, Lei Zhang 0001, Binghan Chen, Fei Chao 0001, Chenglin Wu 0001, Shanshan Wang 0002, Rongrong Ji, Yonghong Tian 0001
IEEE Trans. Pattern Anal. Mach. Intell.5
2025 Move and Act: Enhanced Object Manipulation and Background Integrity for Image Editing
abstract
Current methods commonly utilize three-branch structures of inversion, reconstruction, and editing, to tackle consistent image editing task. However, these methods lack control over the generation position of the edited object and have issues with background preservation. To overcome these limitations, we propose a tuning-free method with only two branches: inversion and editing. This approach allows users to simultaneously edit the object's action and control the generation position of the edited object. Additionally, it achieves improved background preservation. Specifically, we transfer the edited object information to the target area and repair or preserve the background of other areas during the inversion process at a specific time step. In the editing stage, we use the image features in self-attention to query the key and value of the corresponding time step in the inversion to achieve consistent image editing. Impressive image editing results and quantitative evaluation demonstrate the effectiveness of our method.
Mingbao Lin, Fei Chao 0001
AAAI3
2025 Semantic Alignment and Reinforcement for Data-Free Quantization of Vision Transformers
Yunshan Zhong, Yuyao Zhou, Yuxin Zhang 0002, Wanchen Sui, Fei Chao 0001, Rongrong Ji
ICCV7
2025 GS-Bias: Global-Spatial Bias Learner for Single-Image Test-Time Adaptation of Vision-Language Models
abstract
Recent advances in test-time adaptation (TTA) for Vision-Language Models (VLMs) have garnered increasing attention, particularly through the use of multiple augmented views of a single image to boost zero-shot generalization. Unfortunately, existing methods fail to strike a satisfactory balance between performance and efficiency, either due to excessive overhead of tuning text prompts or unstable benefits from handcrafted, training-free visual feature enhancement. In this paper, we present Global-Spatial Bias Learner (GS-Bias), an efficient and effective TTA paradigm that incorporates two learnable biases during TTA, unfolded as the global bias and spatial bias. Particularly, the global bias captures the global semantic features of a test image by learning consistency across augmented views, while spatial bias learns the semantic coherence between regions in the image’s spatial visual representation. It is worth highlighting that these two sets of biases are directly added to the logits outputed by the pretrained VLMs, which circumvent the full backpropagation through VLM that hinders the efficiency of existing TTA methods. This endows GS-Bias with extremely high efficiency while achieving state-of-the-art performance on 15 benchmark datasets. For example, it achieves a 2.23% improvement over TPT in cross-dataset generalization and a 2.72% improvement in domain generalization, while requiring only 6.5% of TPT’s memory usage on ImageNet.
Zhaohong Huang, Yuxin Zhang 0002, Jingjing Xie, Fei Chao 0001, Rongrong Ji
ICML4
2025 Determining Layer-wise Sparsity for Large Language Models Through a Theoretical Perspective
abstract
In this paper, we address the challenge of determining the layer-wise sparsity rates of large language models (LLMs) through a theoretical perspective. Specifically, we identify a critical issue of **"reconstruction error explosion"** in existing LLMs sparsification methods. This refers to the cumulative effect of reconstruction errors throughout the sparsification process, where errors from earlier layers propagate and amplify in subsequent layers. As a result, the overall reconstruction error increases significantly, leading to a substantial degradation in model performance. Through theoretical analysis, we derive a simple yet effective approach to layer-wise sparsity allocation that mitigates this issue. Our method uses a monotonically increasing arithmetic progression, reducing the process of determining sparsity rates for multiple layers to the determination of a single common difference hyperparameter. Remarkably, this allows for the optimal layer-wise sparsity rates to be identified with just a few trials. Both our theoretical analysis and experimental results demonstrate that this sparsity allocation scheme is near optimal. Extensive experiments show that our method significantly improves the performance of sparse LLMs across various architectures, outperforming existing layer-wise sparsity methods. Furthermore, it enhances the performance of various compression techniques and is applicable to vision and multimodal models. Notably, our method achieves a reduction of 52.10 in perplexity for the 70% sparse LLaMA2-7B model obtained via Wanda, improves average zero-shot accuracy by 10.50%, and delivers speedups of 2.63$\times$ and 2.23$\times$ on CPU and GPU, respectively. Code is available at https://github.com/wzhuang-xmu/ATP.
Weizhong Huang, Yuxin Zhang 0002, Xiawu Zheng, Fei Chao 0001, Rongrong Ji
ICML4
2025 polybasic Speculative Decoding Through a Theoretical Perspective
abstract
Inference latency stands as a critical bottleneck in the large-scale deployment of Large Language Models (LLMs). Speculative decoding methods have recently shown promise in accelerating inference without compromising the output distribution. However, existing work typically relies on a dualistic draft-verify framework and lacks rigorous theoretical grounding. In this paper, we introduce a novel \emph{polybasic} speculative decoding framework, underpinned by a comprehensive theoretical analysis. Specifically, we prove a fundamental theorem that characterizes the optimal inference time for multi-model speculative decoding systems, shedding light on how to extend beyond the dualistic approach to a more general polybasic paradigm. Through our theoretical investigation of multi-model token generation, we expose and optimize the interplay between model capabilities, acceptance lengths, and overall computational cost. Our framework supports both standalone implementation and integration with existing speculative techniques, leading to accelerated performance in practice. Experimental results across multiple model families demonstrate that our approach yields speedup ratios ranging from $3.31\times$ to $4.01\times$ for LLaMA2-Chat 7B, up to $3.87 \times$ for LLaMA3-8B, up to $4.43 \times$ for Vicuna-7B and up to $3.85 \times$ for Qwen2-7B---all while preserving the original output distribution. We release our theoretical proofs and implementation code to facilitate further investigation into polybasic speculative decoding.
Huixia Li, Yuexiao Ma, Xiawu Zheng, Fei Chao 0001, Xuefeng Xiao 0001, Rongrong Ji
ICML5
2025 Breaking Static Barriers: Dynamic Post-Training Quantization for Diffusion Models
abstract
Current Post-Training Quantization (PTQ) schemes have been extensively studied for traditional convolutional neural networks and language models; however, PTQ application in diffusion models has shown significant performance degradation due to static settings of PTQ. Existing methods only uniformly and statically sample during each denoising step to construct calibration sets, neglecting the different importance of different steps in diffusion models. Furthermore, diffusion models exhibit a large number of activations with skewed distributions, and maintaining a static zero-point during the reconstruction process causes the model to converge only to local optima. To solve these limitations, it is necessary to dynamically design calibration dataset construction methods for different quantization scenarios and develop specialized optimization strategies tailored to specific activation distributions. Thus we proposed a unified framework, termed Dynamic PTQ, to achieve the aforementioned purposes. The framework first applies an evolutionary search algorithm to dynamically construct calibration sets for different quantization scenarios. Then, we design a dynamic zero-point update strategy for the quantizer, significantly reducing the loss during the reconstruction process. Extensive experiments demonstrate that our method outperforms current PTQ methods for diffusion models in generating high-quality samples. In particular, for the LSUN-bedrooms 256×256 task, our method quantizes the corresponding full-precision LDM-4 to W4A6 with only a 0.84 increase in FID.
Huixia Li, Lijiang Li, Xiawu Zheng, Yuexiao Ma, Jie Wu 0001, Xuefeng Xiao 0001, Rui Wang 0089, Fei Chao 0001
IJCNN11
2025 Optimizing Time-Step Sampling Probabilities in Diffusion Models for Enhanced Training Efficiency
abstract
Diffusion models have surpassed Generative Adversarial Networks in generating high-quality, high-resolution images, enhancing detail and diversity. However, diffusion models still demand significant time and computational resources. Current work indicates that the quality of generated images is tied to sampling time steps, with each phase in the generation process affecting training and output differently. To address this challenge, this study introduces evolutionary algorithms to optimize the sequence of time-step sampling probabilities within the training phase of diffusion models. Due to traditional sampling probability sequences involving floating points and high dimensions, this paper simplifies the search space and redefines the search objectives of the evolutionary algorithm, making the search process more efficient. Experimental results demonstrate that the proposed method not only speeds up the training process of diffusion models but also reveals that effective time step sampling probability sequences from adjacent training phases have similar distributions, indicating that a stable sequence of time steps exists that can consistently accelerate network convergence throughout extensive training phases. These findings not only enhance training efficiency but also reduce computational costs while maintaining the quality of generated images. The source code is available in the GitHub repository: https://github.com/zhangbeibei00/O2TDM.
Xiang Chang, Changjing Shang, Qiang Shen 0001, Fei Chao 0001
IJCNN6
2025 VISA: Group-wise Visual Token Selection and Aggregation via Graph Summarization for Efficient MLLMs Inference
abstract
In this study, we introduce a novel method called group-wise VI sual token Selection and Aggregation (VISA) to address the issue of inefficient inference stemming from excessive visual tokens in multimoal large language models (MLLMs). Compared with previous token pruning approaches, our method can preserve more visual information while compressing visual tokens. We first propose a graph-based visual token aggregation (VTA) module. VTA treats each visual token as a node, forming a graph based on semantic similarity among visual tokens. It then aggregates information from removed tokens into kept tokens based on this graph, producing a more compact visual token representation. Additionally, we introduce a group-wise token selection strategy (GTS) to divide visual tokens into kept and removed ones, guided by text tokens from the final layers of each group. This strategy progressively aggregates visual information, enhancing the stability of the visual information extraction process. We conduct comprehensive experiments on LLaVA-1.5, LLaVA-NeXT, and Video-LLaVA across various benchmarks to validate the efficacy of VISA. Our method consistently outperforms previous methods, achieving a superior trade-off between model performance and inference speed.
Hanjun Li 0002, Linglan Zhao, Fei Chao 0001, Shouhong Ding, Rongrong Ji
ACM Multimedia4
2025 Discovering Important Experts for Mixture-of-Experts Models Pruning Through a Theoretical Perspective
abstract
Mixture-of-Experts (MoE) architectures enable efficient scaling of large language models but face prohibitive memory demands due to massive parameterization. Existing pruning methods rely on heuristic metrics or impractical enumeration of expert subsets, leading to suboptimal performance or scalability. In this paper, we propose Shapley-MoE, an efficient pruning method for MoE models inspired by cooperative game theory. By quantifying each expert’s contribution via Shapley value, our method identifies important experts without exhaustive combination evaluations. To overcome the NP-hard complexity of exact Shapley computation, we introduce a Monte Carlo sampling strategy for efficient approximation that reduces complexity to quadratic time. However, vanilla Monte Carlo sampling still faces issues of insufficient estimation accuracy and low sampling efficiency. To address these issues, we further propose two novel methods to improve sampling accuracy and efficiency: (1) Early Truncation, which early terminates unstable sampling steps caused by overly small expert subsets, and (2) Router-Guided Importance Sampling, which prioritize sampling important expert subsets using gating activation probabilities. Both theoretical and experimental analyses show that both methods can accelerate Shapley value estimation and improve accuracy. Extensive empirical evaluations demonstrate that our pruned MoE models outperform existing expert pruning methods. Notably, when applied to the Qwen2-57B-A14B model, our method reduces the number of experts by 25% with only a 0.92 increase in perplexity and over 96.4% of the average zero-shot accuracy is maintained.
Weizhong Huang, Yuxin Zhang 0002, Xiawu Zheng, Fei Chao 0001, Rongrong Ji, Liujuan Cao
NeurIPS4
2025 Spotlight Attention: Towards Efficient LLM Generation via Non-linear Hashing-based KV Cache Retrieval
abstract
Reducing the key-value (KV) cache burden in Large Language Models (LLMs) significantly accelerates inference. Dynamically selecting critical KV caches during decoding helps maintain performance. Existing methods use random linear hashing to identify important tokens, but this approach is inefficient due to the orthogonal distribution of queries and keys within two narrow cones in LLMs. We introduce Spotlight Attention, a novel method that employs non-linear hashing functions to optimize the embedding distribution of queries and keys, enhancing coding efficiency and robustness. We also developed a lightweight, stable training framework using a Bradley-Terry ranking-based loss, enabling optimization of the non-linear hashing module on GPUs with 16GB memory in 8 hours. Experimental results show that Spotlight Attention drastically improves retrieval precision while shortening the length of the hash code at least 5$\times$ compared to traditional linear hashing. Finally, we exploit the computational advantages of bitwise operations by implementing specialized CUDA kernels, achieving hashing retrieval for 512K tokens in under 100$\mu$s on a single A100 GPU, with end-to-end throughput up to 3$\times$ higher than vanilla decoding.
Wenhao Li 0001, Yuxin Zhang 0002, Gen Luo, Haiyuan Wan, Ziyang Gong, Fei Chao 0001, Rongrong Ji
NeurIPS6
2025 Video-RAG: Visually-aligned Retrieval-Augmented Long Video Comprehension
abstract
Existing large video-language models (LVLMs) struggle to comprehend long videos correctly due to limited context. To address this problem, fine-tuning long-context LVLMs and employing GPT-based agents have emerged as promising solutions. However, fine-tuning LVLMs would require extensive high-quality data and substantial GPU resources, while GPT-based agents would rely on proprietary models (e.g., GPT-4o). In this paper, we propose Video Retrieval-Augmented Generation (Video-RAG), a training-free and cost-effective pipeline that employs visually-aligned auxiliary texts to help facilitate cross-modality alignment while providing additional information beyond the visual content. Specifically, we leverage open-source external tools to extract visually-aligned information from pure video data (e.g., audio, optical character, and object detection), and incorporate the extracted information into an existing LVLM as auxiliary texts, alongside video frames and queries, in a plug-and-play manner. Our Video-RAG offers several key advantages: (i) lightweight with low computing overhead due to single-turn retrieval; (ii) easy implementation and compatibility with any LVLM; and (iii) significant, consistent performance gains across long video understanding benchmarks, including Video-MME, MLVU, and LongVideoBench. Notably, our model demonstrates superior performance over proprietary models like Gemini-1.5-Pro and GPT-4o when utilized with a 72B model.
Yongdong Luo, Xiawu Zheng, Shukang Yin, Haojia Lin, Chaoyou Fu, Jinfa Huang, Jiayi Ji, Fei Chao 0001, Jiebo Luo 0001, Rongrong Ji
NeurIPS9
2025 DDRE: Decoupled Diffusion Reconstruction Error for AI-Generated Image Detection
Mengcheng Li, Fei Chao 0001
PRCV (18)2
2025 Orpaint: a zero-shot inpainting model for oracle bone inscription rubbings with visual mamba block
Zijie Meng, Yuan-Ze Zeng, Xiang Chang, Tianshuo Xu, Fei Chao 0001, Xixin Cao, Changjing Shang, Qiang Shen 0001
Sci. China Inf. Sci.5
2025 Distribution-flexible subset quantization for post-quantizing super-resolution networks
Yunshan Zhong, Mingbao Lin, Jingjing Xie, Yuxin Zhang 0002, Fei Chao 0001, Rongrong Ji
Sci. China Inf. Sci.5
2025 MBQuant: A novel multi-branch topology method for arbitrary bit-width network quantization
Yunshan Zhong, Yuyao Zhou, Fei Chao 0001, Rongrong Ji
Pattern Recognit.3
2025 NADM: Noise-Aware Diffusion Model for Landscape Painting Video Generation
abstract
Landscape painting is a gem of cultural and artistic heritage that showcases the splendor of nature through the deep observations and imaginations of its painters. Limited by traditional techniques, these artworks were confined to static imagery in ancient times, leaving the dynamism of landscapes and the subtleties of artistic sentiment to the viewer's imagination. Recently, emerging text-to-video (T2V) diffusion methods have shown significant promise in video generation, providing hope for the creation of dynamic landscape paintings. However, current T2V methods focus on generating natural videos, emphasizing the capture of details and the authenticity of physical laws. In contrast, landscape painting videos emphasize the overall dynamic aesthetic. Besides, challenges, such as the lack of specific datasets, the intricacy of artistic styles, and the creation of extensive, high-quality videos pose difficulties for these models in generating landscape painting videos. In this article, we propose landscape painting videos-high definition (LPV-HD), a novel T2V dataset for landscape painting videos, and noise-aware diffusion model (NADM), a T2V model that utilizes Stable Diffusion. Specifically, we present a motion module featuring a dual attention mechanism to capture the dynamic transformations of landscape imageries, alongside a noise adapter to leverage unsupervised contrastive learning in the latent space to ensure the overall beauty of the landscape painting video. Following the generation of keyframes, we employ optical flow for frame interpolation to enhance video smoothness. Our method not only retains the essence of the landscape painting imageries but also achieves dynamic transitions, significantly advancing the field of artistic video generation. Source code and dataset are available at https://github.com/llzlh21/NADM.
Ding-Ming Liu, Shao-Wei Li, Ruo-Yan Zhou, Lili Liang, Yongguan Hong, Yuan-Ze Zeng, Xiang Chang, Lijiang Li, Tianshuo Xu, Fei Chao 0001, Changjing Shang, Qiang Shen 0001
IEEE Trans. Cybern.10
2025 ARF: Arbitrary Routing Framework for All-in-One Image Restoration
abstract
All-in-one image restoration methods, as opposed to conventional image restoration methods, reconstruct images impaired by various degradations within a unified model, eliminating the need for separate network parameters for each task. However, current all-in-one image restoration approaches tackle various types of image degradation using an identical underlying model, neglecting the inherent variability in complexity across different image restoration tasks, resulting in inefficient allocation of computational resources. To address this limitation, this article introduces the arbitrary routing framework (ARF), designed to effectively assess the difficulty of image restoration tasks and identify the most suitable network structure based on these complexities. This framework can be integrated with existing all-in-one image restoration models, enabling efficient inference by activating various proportions of the entire network, that is subnetworks, based on their task-specific complexities. More specifically, the ARF comprises two principal components: 1) the arbitrary routing backbone (ARB) and 2) a task-specific neural architecture search (T-NAS). The ARB incorporates a routing layer between consecutive convolutional groups, offering a wide array of potential subnetwork configurations while adding only negligible extra parameters. Concurrently, T-NAS autonomously identifies the most effective subnetworks for each image restoration task, optimizing both performance and efficiency through an efficiency-aware reward function. Comprehensive experiments across various image restoration tasks demonstrate that the ARF significantly improves performance metrics, that is, an increase of 0.31 in reconstruction PSNR, while also achieving a notable reduction in computational demands by 37.1% compared with the benchmark AirNet method. The code has been made available in the supplementary materials.
Yimin Xu, Nanxi Gao, Yunshan Zhong, Fei Chao 0001, Rongrong Ji
IEEE Trans. Cybern.4
2025 Self-Organizing Type-2 Fuzzy Double Loop Recurrent Neural Network for Uncertain Nonlinear System Control
abstract
Nonlinear systems, such as robotic systems, play an increasingly important role in our modern daily life and have become more dominant in many industries; however, robotic control still faces various challenges due to diverse and unstructured work environments. This article proposes a double-loop recurrent neural network (DLRNN) with the support of a Type-2 fuzzy system and a self-organizing mechanism for improved performance in nonlinear dynamic robot control. The proposed network has a double-loop recurrent structure, which enables better dynamic mapping. In addition, the network combines a Type-2 fuzzy system with a double-loop recurrent structure to improve the ability to deal with uncertain environments. To achieve an efficient system response, a self-organizing mechanism is proposed to adaptively adjust the number of layers in a DLRNN. This work integrates the proposed network into a conventional sliding mode control (SMC) system to theoretically and empirically prove its stability. The proposed system is applied to a three-joint robot manipulator, leading to a comparative study that considers several existing control approaches. The experimental results confirm the superiority of the proposed system and its effectiveness and robustness in response to various external system disturbances.
Lijiang Li, Xiang Chang, Fei Chao 0001, Chih-Min Lin, Tuan-Tu Huynh, Longzhi Yang, Changjing Shang, Qiang Shen 0001
IEEE Trans. Neural Networks Learn. Syst.3
2024 Learning Image Demoiréing from Unpaired Real Data
abstract
This paper focuses on addressing the issue of image demoiréing. Unlike the large volume of existing studies that rely on learning from paired real data, we attempt to learn a demoiréing model from unpaired real data, i.e., moiré images associated with irrelevant clean images. The proposed method, referred to as Unpaired Demoiréing(UnDeM), synthesizes pseudo moiré images from unpaired datasets, generating pairs with clean images for training demoiréing models. To achieve this, we divide real moiré images into patches and group them in compliance with their moiré complexity. We introduce a novel moiré generation framework to synthesize moiré images with diverse moiré features, resembling real moiré patches, and details akin to real moiré-free images. Additionally, we introduce an adaptive denoise method to eliminate the low-quality pseudo moiré images that adversely impact the learning of demoiréing models. We conduct extensive experiments on the commonly-used FHDMi and UHDM datasets. Results manifest that our UnDeM performs better than existing methods when using existing demoiréing models such as MBCNN and ESDNet-L. Code: https://github.com/zysxmu/UnDeM.
Yunshan Zhong, Yuyao Zhou, Yuxin Zhang 0002, Fei Chao 0001, Rongrong Ji
AAAI4
2024 RepAn: Enhanced Annealing through Re-parameterization
abstract
The simulated annealing algorithm aims to improve model convergence through multiple restarts of training. However, existing annealing algorithms overlook the cor-relation between different cycles, neglecting the potential for incremental learning. We contend that a fixed network structure prevents the model from recognizing distinct features at different training stages. To this end, we propose RepAn, redesigning the irreversible re-parameterization (Rep) method and integrating it with annealing to enhance training. Specifically, the network goes through Rep, ex-pansion, restoration, and backpropagation operations during training, and iterating through these processes in each annealing round. Such a method exhibits good generalization and is easy to apply, and we provide theoretical expla-nations for its effectiveness. Experiments demonstrate that our method improves baseline performance by 6.38% on the CIFAR-100 dataset and 2.80% on ImageNet, achieving state-of-the-art performance in the Rep field. The code is available at https://github.com/xfey/RepAn.
Xiawu Zheng, Yan Wang 0059, Fei Chao 0001, Chenglin Wu 0001, Liujuan Cao
CVPR4
2024 Functionally Similar Multi-Label Knowledge Distillation
abstract
Existing multi-label knowledge distillation methods simply use regression or single-label classification methods without fully exploiting the essence of multi-label classification, resulting in student models’ inadequate performance and poor functional similarity to teacher models. In this paper, we reinterpret multi-label classification as multiple intra-class ranking tasks, with each class corresponding to a ranking task. Furthermore, we define the knowledge of multi-label classification models as the ranking of intra-class samples. On the one hand, we propose to evaluate the functional similarity between multi-label classification models with Kendall’s tau and rank-biased overlap, which are common metrics for evaluating ranking similarity. On the other hand, we propose a new functionally similar multi-label knowledge distillation method called FSD, which enables student models to learn the ranking of intra-class samples from teacher models. Finally, experimental results validate that FSD outperforms existing methods, especially for functional similarity. Specifically, we achieve a mAP of 73.38% and a mKDT of 0.686 on COCO, which are 2.22% and 0.19 better than existing methods, respectively.
Binghan Chen, Jianlong Hu, Xiawu Zheng, Wei Lin 0004, Fei Chao 0001, Rongrong Ji
ICASSP5
2024 AffineQuant: Affine Transformation Quantization for Large Language Models
abstract
The significant resource requirements associated with Large-scale Language Models (LLMs) have generated considerable interest in the development of techniques aimed at compressing and accelerating neural networks. Among these techniques, Post-Training Quantization (PTQ) has emerged as a subject of considerable interest due to its noteworthy compression efficiency and cost-effectiveness in the context of training. Existing PTQ methods for LLMs limit the optimization scope to scaling transformations between pre- and post-quantization weights. This constraint results in significant errors after quantization, particularly in low-bit configurations. In this paper, we advocate for the direct optimization using equivalent Affine transformations in PTQ (AffineQuant). This approach extends the optimization scope and thus significantly minimizing quantization errors. Additionally, by employing the corresponding inverse matrix, we can ensure equivalence between the pre- and post-quantization outputs of PTQ, thereby maintaining its efficiency and generalization capabilities. To ensure the invertibility of the transformation during optimization, we further introduce a gradual mask optimization method. This method initially focuses on optimizing the diagonal elements and gradually extends to the other elements. Such an approach aligns with the Levy-Desplanques theorem, theoretically ensuring invertibility of the transformation. As a result, significant performance improvements are evident across different LLMs on diverse datasets. Notably, these improvements are most pronounced when using very low-bit quantization, enabling the deployment of large models on edge devices. To illustrate, we attain a C4 perplexity of $15.76$ (2.26$\downarrow$ vs $18.02$ in OmniQuant) on the LLaMA2-$7$B model of W$4$A$4$ quantization without overhead. On zero-shot tasks, AffineQuant achieves an average of $58.61\%$ accuracy ( $1.98\%\uparrow$ vs $56.63$ in OmniQuant) when using $4$/$4$-bit quantization for LLaMA-$30$B, which setting a new state-of-the-art benchmark for PTQ in LLMs. Codes are available at: https://github.com/bytedance/AffineQuant.
Yuexiao Ma, Huixia Li, Xiawu Zheng, Xuefeng Xiao 0001, Rui Wang 0089, Shilei Wen, Fei Chao 0001, Rongrong Ji
ICLR8
2024 Outlier-aware Slicing for Post-Training Quantization in Vision Transformer
abstract
Post-Training Quantization (PTQ) is a vital technique for network compression and acceleration, gaining prominence as model sizes increase. This paper addresses a critical challenge in PTQ: the severe impact of outliers on the accuracy of quantized transformer architectures. Specifically, we introduce the concept of ‘reconstruction granularity’ as a novel solution to this issue, which has been overlooked in previous works. Our work provides theoretical insights into the role of reconstruction granularity in mitigating the outlier problem in transformer models. This theoretical framework is supported by empirical analysis, demonstrating that varying reconstruction granularities significantly influence quantization performance. Our findings indicate that different architectural designs necessitate distinct optimal reconstruction granularities. For instance, the multi-stage Swin Transformer architecture benefits from finer granularity, a deviation from the trends observed in ViT and DeiT models. We further develop an algorithm for determining the optimal reconstruction granularity for various ViT models, achieving state-of-the-art (SOTA) performance in PTQ. For example, applying our method to $4$-bit quantization, the Swin-Base model achieves a Top-1 accuracy of $82.24%$ on the ImageNet classification task. This result surpasses the RepQ-ViT by $3.92%$ ($82.24%$ VS $78.32%$). Similarly, our approach elevates the ViT-Small to a Top-1 accuracy of $80.50%$, outperforming NoisyQuant by $3.64%$ ($80.50%$ VS $76.86%$). Codes are available in Supplementary Materials.
Yuexiao Ma, Huixia Li, Xiawu Zheng, Xuefeng Xiao 0001, Rui Wang 0089, Shilei Wen, Fei Chao 0001, Rongrong Ji
ICML8
2024 Multimodal Inplace Prompt Tuning for Open-set Object Detection
abstract
The integration of large language models into open-world detection frameworks significantly improves versatility in new environments. Prompt representations derived from these models help establish classification boundaries for both base and novel categories within open-world detectors. However, we are the first to discover that directly fine-tuning language models in detection systems results in redundant attention patterns and leads to suboptimal prompt representations. In order to fully leverage the capabilities of large language models and augment prompt encoding for detection, this study introduces a redundancy assessment metric to identify uniform attention patterns. Furthermore, in areas with high redundancy, we incorporate multimodal inplace prompt tuning (MIPT) to enrich the text prompt with visual clues. Experimental results validate the efficacy of our MIPT framework, achieving a notable increase across benchmarks, e.g. elevating GLIP-L from 22.6% to 25.0% on ODinW-35, and 9.0% improvement on LVIS.
Mengdan Zhang, Xiawu Zheng, Peixian Chen, Yunhang Shen, Mingchen Zhuge, Chenglin Wu 0001, Fei Chao 0001, Ke Li 0015, Xing Sun 0001, Rongrong Ji
ACM Multimedia9
2024 CPE COIN++: Towards Optimized Implicit Neural Representation Compression Via Chebyshev Positional Encoding
Haocheng Chu, Shaohui Dai, Wenqi Ding, Tianshuo Xu, Pingyang Dai, Shengchuan Zhang, Yan Zhang 0109, Xiang Chang, Chih-Min Lin, Fei Chao 0001, Changjiang Shang, Qiang Shen 0001
PRCV (9)11
2024 Local representation-based neighbourhood for robust classification
abstract
Abstract Representation‐based classification (RC) is an effective gauge of data similarity between a single instance and the whole dataset, which extends traditional individual‐wise distance metrics using representation coefficients. These coefficients show remarkable discrimination nature via various regularisation terms, but the interference from potentially uncorrelated objects involved in this single‐to‐global relation can degrade the effectiveness of the coefficients. In order to filter out those unproductive, or even counter‐productive, information from the decision making processes, this paper proposes a local representation‐based classification (LRC) algorithm to improve the classification accuracy or the RC approach. LRC uses a single‐to‐local relation induced by the local representation‐based neighbourhood (LRN) of each object, rather than the single‐to‐global relationship used by RC. Thanks to LRN, a compact and relevant dataset can be formed by selecting the most relevant data instances in the original dataset, to render a robust representation of a query. LRC was applied to multiple publicly available datasets, and the experimental results demonstrate the superiority of the proposed LRC algorithm as evidenced by the higher classification accuracy and more noise‐tolerant capability in reference to alternative RC approaches. Moreover, the sampling ability of LRN is also verified via a comparative study.
Zihan Yao, Yanpeng Qu, Longzhi Yang, Changjing Shang, Fei Chao 0001, Qiang Shen 0001
Expert Syst. J. Knowl. Eng.6
2024 ARLP: Automatic multi-agent transformer reinforcement learning pruner for one-shot neural network pruning
Bowen Guo, Xiang Chang, Fei Chao 0001, Xiawu Zheng, Chih-Min Lin, Changjing Shang, Qiang Shen 0001
Knowl. Based Syst.3
2024 Deep hybrid transformer network for robust modulation classification in wireless communications
Qiancheng Zheng, Heng Wei, Jinxian Zhao, Yiyi Zhou, Fei Chao 0001, Rongrong Ji
Knowl. Based Syst.7
2024 Uncovering the Over-Smoothing Challenge in Image Super-Resolution: Entropy-Based Quantification and Contrastive Optimization
abstract
PSNR-oriented models are a critical class of super-resolution models with applications across various fields. However, these models tend to generate over-smoothed images, a problem that has been analyzed previously from the perspectives of models or loss functions, but without taking into account the impact of data properties. In this paper, we present a novel phenomenon that we term the center-oriented optimization (COO) problem, where a model's output converges towards the center point of similar high-resolution images, rather than towards the ground truth. We demonstrate that the strength of this problem is related to the uncertainty of data, which we quantify using entropy. We prove that as the entropy of high-resolution images increases, their center point will move further away from the clean image distribution, and the model will generate over-smoothed images. Implicitly optimizing the COO problem, perceptual-driven approaches such as perceptual loss, model structure optimization, or GAN-based methods can be viewed. We propose an explicit solution to the COO problem, called Detail Enhanced Contrastive Loss (DECLoss). DECLoss utilizes the clustering property of contrastive learning to directly reduce the variance of the potential high-resolution distribution and thereby decrease the entropy. We evaluate DECLoss on multiple super-resolution benchmarks and demonstrate that it improves the perceptual quality of PSNR-oriented models. Moreover, when applied to GAN-based methods, such as RaGAN, DECLoss helps to achieve state-of-the-art performance, such as 0.093 LPIPS with 24.51 PSNR on 4× downsampled Urban100, validating the effectiveness and generalization of our approach.
Tianshuo Xu, Lijiang Li, Peng Mi, Xiawu Zheng, Fei Chao 0001, Rongrong Ji, Yonghong Tian 0001, Qiang Shen 0001
IEEE Trans. Pattern Anal. Mach. Intell.5
2024 An efficient blur kernel estimation method for blind image Super-Resolution
Yimin Xu, Nanxi Gao, Fei Chao 0001, Rongrong Ji
Pattern Recognit.3
2024 Shadow-aware dynamic convolution for shadow removal
Yimin Xu, Mingbao Lin, Fei Chao 0001, Rongrong Ji
Pattern Recognit.4
2024 Actor-Critic With Synthesis Loss for Solving Approximation Biases
abstract
Approximation biases of value functions are considered a key problem in reinforcement learning (RL). In particular, existing RL algorithms are hindered by overestimation and underestimation biases, i.e., value mismatching between RL's actual returns and action-value approximations limits the performance of RL algorithms. In this article, we first develop a new synthesis loss function for RL's action-value estimation integrating a regularization term and a modified "clipped double Q-learning" structure for solving overestimation and underestimation biases. To minimize the differences between action-value estimations and actual returns in RL, we develop a new discrepancy function to determine the type and magnitude of approximation biases. Then, two coefficients embedded in the synthesis loss are automatically tuned by minimizing the discrepancy function during training to minimize approximation biases. We further design a new actor-critic (AC) algorithm, named AC with synthesis loss (ACSL), by integrating the synthesis loss function and an error-controlled mechanism. Experimental results on continuous control tasks illustrate that the proposed ACSL algorithm outperforms other cutting-edge RL methods in many tasks and that the proposed synthesis loss function is easily implemented into other algorithms and significantly reduces approximation biases while improving performance. The proposed method can successfully handle many complex continuous control tasks and can greatly outperform other state-of-the-art algorithms on most tasks.
Bowen Guo, Fei Chao 0001, Xiang Chang, Changjing Shang, Qiang Shen 0001
IEEE Trans. Cybern.2
2024 Solving Robotic Trajectory Sequential Writing Problem via Learning Character's Structural and Sequential Information
abstract
The writing sequence of numerals or letters often affects aesthetic aspects of the writing outcomes. As such, it remains a challenge for robotic calligraphy systems to perform, mimicking human writers' implicit intention. This article presents a new robot calligraphy system that is able to learn writing sequences with limited sequential information, producing writing results compatible to human writers with good diversity. In particular, the system innovatively applies a gated recurrent unit (GRU) network to generate robotic writing actions with the support of a prelabeled trajectory sequence vector. Also, a new evaluation method is proposed that considers the shape, trajectory sequence, and structural information of the writing outcome, thereby helping ensure the writing quality. A swarm optimization algorithm is exploited to create an optimal set of parameters of the proposed system. The proposed approach is evaluated using Arabic numerals, and the experimental results demonstrate the competitive writing performance of the system against state-of-the-art approaches regarding multiple criteria (including FID, MAE, PSNR, SSIM, and PerLoss), as well as diversity performance concerning variance and entropy. Importantly, the proposed GRU-based robotic motion planning system, supported with swarm optimization can learn from a small dataset, while producing calligraphy writing with diverse and aesthetically pleasing outcomes.
Quanfeng Li, Fei Chao 0001, Xiang Chang, Longzhi Yang, Chih-Min Lin, Changjing Shang, Qiang Shen 0001
IEEE Trans. Cybern.3
2024 Internal Model Control Structure Inspired Robotic Calligraphy System
abstract
Learning calligraphy writing skills in robots is regarded as a sophisticated task. Current robotic researchers have proposed many methods to implement various robotic calligraphy systems. However, several limitations of these methods, such as high computational costs and few diversities of generated results constrain the development of calligraphy robots. This article proposes a robotic writing framework based on a robotic hand–eye coordination method to solve these limitations. Inspired by the internal model control (IMC) system, a vision-motor network and a motor-vision network are built to simulate the direct and reverse models, respectively, in the IMC system of a robotic manipulator. The vision-motor network works as an action generator to convert image inputs to robotic actions, and the motor-vision network assists in the training of the vision-motor network. Thus, a pretraining of the motor-vision network is established by random writing movements of a robotic manipulator. Experimental results demonstrate that the proposed method can successfully write strokes of Chinese characters by inputting target stroke images. Although the proposed method is applied to robotic calligraphy, the underpinning research is readily applicable to many other applications, such as human–robot motion mimicking.
Fei Chao 0001, Changle Zhou, Xiang Chang, Longzhi Yang, Changjing Shang, Qiang Shen 0001
IEEE Trans. Ind. Informatics2
2023 CF-ViT: A General Coarse-to-Fine Method for Vision Transformer
abstract
Vision Transformers (ViT) have made many breakthroughs in computer vision tasks. However, considerable redundancy arises in the spatial dimension of an input image, leading to massive computational costs. Therefore, We propose a coarse-to-fine vision transformer (CF-ViT) to relieve computational burden while retaining performance in this paper. Our proposed CF-ViT is motivated by two important observations in modern ViT models: (1) The coarse-grained patch splitting can locate informative regions of an input image. (2) Most images can be well recognized by a ViT model in a small-length token sequence. Therefore, our CF-ViT implements network inference in a two-stage manner. At coarse inference stage, an input image is split into a small-length patch sequence for a computationally economical classification. If not well recognized, the informative patches are identified and further re-split in a fine-grained granularity. Extensive experiments demonstrate the efficacy of our CF-ViT. For example, without any compromise on performance, CF-ViT reduces 53% FLOPs of LV-ViT, and also achieves 2.01x throughput. Code of this project is at https://github.com/ChenMnZ/CF-V
Mengzhao Chen, Mingbao Lin, Ke Li 0015, Yunhang Shen, Yongjian Wu 0001, Fei Chao 0001, Rongrong Ji
AAAI6
2023 Discriminator-Cooperated Feature Map Distillation for GAN Compression
abstract
Despite excellent performance in image generation, Generative Adversarial Networks (GANs) are notorious for its requirements of enormous storage and intensive computation. As an awesome “performance maker”, knowledge distillation is demonstrated to be particularly efficacious in exploring low-priced GANs. In this paper, we investigate the irreplaceability of teacher discriminator and present an inventive discriminator-cooperated distillation, abbreviated as DCD, towards refining better feature maps from the generator. In contrast to conventional pixel-to-pixel match methods in feature map distillation, our DCD utilizes teacher discriminator as a transformation to drive intermediate results of the student generator to be perceptually close to corresponding outputs of the teacher generator. Furthermore, in order to mitigate mode collapse in GAN compression, we construct a collaborative adversarial training paradigm where the teacher discriminator is from scratch established to co-train with student generator in company with our DCD. Our DCD shows superior results compared with existing GAN compression methods. For instance, after reducing over$40\times MACs$and$80\times$parameters of CycleGAN, we well decrease FID metric from 61.53 to 48.24 while the current SoTA method merely has 51.92. This work's source code has been made accessible at https://github.com/poopit/DCD-official.
Tie Hu, Mingbao Lin, Lizhou You, Fei Chao 0001, Rongrong Ji
CVPR4
2023 Meta Architecture for Point Cloud Analysis
abstract
Recent advances in 3D point cloud analysis bring a diverse set of network architectures to the field. However, the lack of a unified framework to interpret those networks makes any systematic comparison, contrast, or analysis challenging, and practically limits healthy development of the field. In this paper, we take the initiative to explore and propose a unified framework called PointMeta, to which the popular 3D point cloud analysis approaches could fit. This brings three benefits. First, it allows us to compare different approaches in a fair manner, and use quick experiments to verify any empirical observations or assumptions summarized from the comparison. Second, the big picture brought by PointMeta enables us to think across different components, and revisit common beliefs and key design decisions made by the popular approaches. Third, based on the learnings from the previous two analyses, by doing simple tweaks on the existing approaches, we are able to derive a basic building block, termed PointMetaBase. It shows very strong performance in efficiency and effectiveness through extensive experiments on challenging benchmarks, and thus verifies the necessity and benefits of high-level interpretation, contrast, and comparison like PointMeta. In particular, PointMetaBase surpasses the previous state-of-the-art method by 0.7%/1.4/%2.1% mIoU with only 2%/11%/13% of the computation cost on the S3DIS datasets. The code and models are available at https://github.com/linhaojia13/PointMetaBase.
Haojia Lin, Xiawu Zheng, Lijiang Li, Fei Chao 0001, Shanshan Wang 0002, Yan Wang 0059, Yonghong Tian 0001, Rongrong Ji
CVPR4
2023 Solving Oscillation Problem in Post-Training Quantization Through a Theoretical Perspective
abstract
Post-training quantization (PTQ) is widely regarded as one of the most efficient compression methods practically, benefitting from its data privacy and low computation costs. We argue that an overlooked problem of oscillation is in the PTQ methods. In this paper, we take the initiative to explore and present a theoretical proof to explain why such a problem is essential in PTQ. And then, we try to solve this problem by introducing a principled and generalized frame-work theoretically. In particular, we first formulate the oscillation in PTQ and prove the problem is caused by the difference in module capacity. To this end, we define the module capacity (ModCap) under data-dependent and data-free scenarios, where the differentials between adjacent modules are used to measure the degree of oscillation. The problem is then solved by selecting top-k differentials, in which the corresponding modules are jointly optimized and quantized. Extensive experiments demonstrate that our method successfully reduces the performance drop and is generalized to different neural networks and PTQ methods. For example, with 2/4 bit ResNet-50 quantization, our method surpasses the previous state-of-the-art method by 1.9%. It becomes more significant on small model quantization, e.g. surpasses BRECQ method by 6.61% on MobileNetV2 × 0.5.
Yuexiao Ma, Huixia Li, Xiawu Zheng, Xuefeng Xiao 0001, Rui Wang 0089, Shilei Wen, Fei Chao 0001, Rongrong Ji
CVPR8
2023 Automatic Network Pruning via Hilbert-Schmidt Independence Criterion Lasso under Information Bottleneck Principle
abstract
Most existing neural network pruning methods hand-crafted their importance criteria and structures to prune. This constructs heavy and unintended dependencies on heuristics and expert experience for both the objective and the parameters of the pruning approach. In this paper, we try to solve this problem by introducing a principled and unified framework based on Information Bottleneck (IB) theory, which further guides us to an automatic pruning approach. Specifically, we first formulate the channel pruning problem from an IB perspective, and then implement the IB principle by solving a Hilbert-Schmidt Independence Criterion (HSIC) Lasso problem under certain conditions. Based on the theoretical guidance, we then provide an automatic pruning scheme by searching for global penalty coefficients. Verified by extensive experiments, our method yields state-of-the-art performance on various benchmark networks and datasets. For example, with VGG-16, we achieve a 60%-FLOPs reduction by removing 76% of the parameters, with an improvement of 0.40% in top-1 accuracy on CIFAR-10. With ResNet-50, we achieve a 56%-FLOPs reduction by removing 50% of the parameters, with a small loss of 0.08% in the top-1 accuracy on ImageNet. The code is available at https://github.com/sunggo/APIB.
Song Guo 0001, Lei Zhang 0001, Xiawu Zheng, Yan Wang 0059, Fei Chao 0001, Chenglin Wu 0001, Shengchuan Zhang, Rongrong Ji
ICCV6
2023 SMMix: Self-Motivated Image Mixing for Vision Transformers
abstract
CutMix is a vital augmentation strategy that determines the performance and generalization ability of vision transformers (ViTs). However, the inconsistency between the mixed images and the corresponding labels harms its efficacy. Existing CutMix variants tackle this problem by generating more consistent mixed images or more precise mixed labels, but inevitably introduce heavy training overhead or require extra information, undermining ease of use. To this end, we propose an novel and effective Self-Motivated image Mixing method (SMMix), which motivates both image and label enhancement by the model under training itself. Specifically, we propose a max-min attention region mixing approach that enriches the attention-focused objects in the mixed images. Then, we introduce a fine-grained label assignment technique that co-trains the output tokens of mixed images with fine-grained supervision. Moreover, we devise a novel feature consistency constraint to align features from mixed and unmixed images. Due to the subtle designs of the self-motivated paradigm, our SMMix is significant in its smaller training overhead and better performance than other CutMix variants. In particular, SMMix improves the accuracy of DeiT-T/S/B, CaiT-XXS-24/36, and PVT-T/S/M/L by more than +1% on ImageNet-1k. The generalization capability of our method is also demonstrated on downstream tasks and out-of-distribution datasets. Our project is available at https://github.com/ChenMnZ/SMMix.
Mengzhao Chen, Mingbao Lin, Zhihang Lin, Yuxin Zhang 0002, Fei Chao 0001, Rongrong Ji
ICCV5
2023 DiffRate : Differentiable Compression Rate for Efficient Vision Transformers
abstract
Token compression aims to speed up large-scale vision transformers (e.g. ViTs) by pruning (dropping) or merging tokens. It is an important but challenging task. Although recent advanced approaches achieved great success, they need to carefully handcraft a compression rate (i.e. number of tokens to remove), which is tedious and leads to sub-optimal performance. To tackle this problem, we propose Differentiable Compression Rate (DiffRate), a novel token compression method that has several appealing properties prior arts do not have. First, DiffRate enables propagating the loss function’s gradient onto the compression ratio, which is considered as a non-differentiable hyperparameter in previous work. In this case, different layers can automatically learn different compression rates layer-wisely without extra overhead. Second, token pruning and merging can be naturally performed simultaneously in DiffRate, while they were isolated in previous works. Third, extensive experiments demonstrate that DiffRate achieves state-of-the-art performance. For example, by applying the learned layer-wise compression rates to an off-the-shelf ViT-H (MAE) model, we achieve a 40% FLOPs reduction and a 1.5× throughput improvement, with a minor accuracy drop of 0.16% on ImageNet without fine-tuning, even outperforming previous methods with fine-tuning. Codes and models are available at https://github.com/OpenGVLab/DiffRate.
Mengzhao Chen, Wenqi Shao, Peng Xu 0035, Mingbao Lin, Kaipeng Zhang, Fei Chao 0001, Rongrong Ji, Yu Qiao 0001, Ping Luo 0002
ICCV6
2023 AutoDiffusion: Training-Free Optimization of Time Steps and Architectures for Automated Diffusion Model Acceleration
abstract
Diffusion models are emerging expressive generative models, in which a large number of time steps (inference steps) are required for a single image generation. To accelerate such tedious process, reducing steps uniformly is considered as an undisputed principle of diffusion models. We consider that such a uniform assumption is not the optimal solution in practice; i.e., we can find different optimal time steps for different models. Therefore, we propose to search the optimal time steps sequence and compressed model architecture in a unified framework to achieve effective image generation for diffusion models without any further training. Specifically, we first design a unified search space that consists of all possible time steps and various architectures. Then, a two stage evolutionary algorithm is introduced to find the optimal solution in the designed search space. To further accelerate the search process, we employ FID score between generated and real samples to estimate the performance of the sampled examples. As a result, the proposed method is (i).training-free, obtaining the optimal time steps and model architecture without any training process; (ii). orthogonal to most advanced diffusion samplers and can be integrated to gain better sample quality. (iii). generalized, where the searched time steps and architectures can be directly applied on different diffusion models with the same guidance scale. Experimental results show that our method achieves excellent performance by using only a few time steps, e.g. 17.86 FID score on ImageNet 64 × 64 with only four steps, compared to 138.66 with DDIM.
Lijiang Li, Huixia Li, Xiawu Zheng, Jie Wu 0032, Xuefeng Xiao 0001, Rui Wang 0089, Fei Chao 0001, Rongrong Ji
ICCV9
2023 Real-Time Image Demoiréing on Mobile Devices
Yuxin Zhang 0002, Mingbao Lin, Xunchao Li, Guozhi Wang, Fei Chao 0001, Shuai Ren 0002, Yafei Wen, Xiaoxin Chen 0001, Rongrong Ji
ICLR6
2023 Bi-directional Masks for Efficient N: M Sparse Training
abstract
We focus on addressing the dense backward propagation issue for training efficiency of N:M fine-grained sparsity that preserves at most N out of M consecutive weights and achieves practical speedups supported by the N:M sparse tensor core. Therefore, we present a novel method of Bi-directional Masks (Bi-Mask) with its two central innovations in: 1) Separate sparse masks in the two directions of forward and backward propagation to obtain training acceleration. It disentangles the forward and backward weight sparsity and overcomes the very dense gradient computation. 2) An efficient weight row permutation method to maintain performance. It picks up the permutation candidate with the most eligible N:M weight blocks in the backward to minimize the gradient gap between traditional unidirectional masks and our bi-directional masks. Compared with existing uni-directional scenario that applies a transposable mask and enables backward acceleration, our Bi-Mask is experimentally demonstrated to be more superior in performance. Also, our Bi-Mask performs on par with or even better than methods that fail to achieve backward acceleration. Project of this paper is available at https://github.com/zyxxmu/Bi-Mask.
Yuxin Zhang 0002, Yiting Luo, Mingbao Lin, Yunshan Zhong, Jingjing Xie, Fei Chao 0001, Rongrong Ji
ICML6
2023 Automated Action Evaluation for Robotic Imitation Learning via Siamese Neural Networks
abstract
Despite recent advances in video-guided robotic imitation learning, many methods still rely on human experts to provide sparse rewards that indicate whether robots have successfully completed tasks. The challenge of enabling robots to autonomously evaluate whether their actions can complete complex, multi-stage tasks remains unresolved. In this work, we propose an efficient few-shot robotic learning algorithm that centres around learning and evaluating from a third-person perspective to address the aforementioned challenge. We develop a novel Siamese neural network-based robotic action-state evaluation system, named “Behavior-Outcome Dual Assessment” (BODA), in our robotic imitation learning system, so as to replace artificial evaluations from human experts in multi-stage imitation learning processes and to improve learning efficiency. In this way, one video demonstration of a target task is divided into several stages. For each stage, we design two Siamese neural network-based evaluation modules in BODA: One module focuses on action changes, and the other handles working environment changes. The two modules work together to provide a comprehensive assessment of the robot's completion of each stage from the view of both the action and working environment changes. Then, BODA is integrated within a model-based reinforcement learning framework to enable the completion of our imitation learning cycle. Extensive experiments demonstrate that the evaluation processes of BODA can automatically and accurately evaluate task completion status without human intervention. In contrast to conventional methods, BODA is able to keep the accumulation of errors within acceptable limits through self-assessment in stages.
Xiang Chang, Fei Chao 0001, Changjing Shang, Qiang Shen 0001
ICRA2
2023 Binarizing Super-Resolution Neural Network Without Batch Normalization
Xunchao Li, Fei Chao 0001
PRCV (10)2
2023 Large Kernel Convolutional Attention Based U-Net Network for Inpainting Oracle Bone Inscription
Xiang Chang, Fei Chao 0001, Changjing Shang, Qiang Shen 0001
PRCV (11)4
2023 Enhancing GAN Compression by Image Probability Distribution Distillation
Lizhou You, Tie Hu, Fei Chao 0001
PRCV (11)3
2023 Model compression optimized neural network controller for nonlinear systems
Lijiang Li, Sheng-Lin Zhou, Fei Chao 0001, Xiang Chang, Longzhi Yang, Changjing Shang, Qiang Shen 0001
Knowl. Based Syst.3
2023 Training Compact CNNs for Image Classification Using Dynamic-Coded Filter Fusion
abstract
The mainstream approach for filter pruning is usually either to force a hard-coded importance estimation upon a computation-heavy pretrained model to select "important" filters, or to impose a hyperparameter-sensitive sparse constraint on the loss objective to regularize the network training. In this paper, we present a novel filter pruning method, dubbed dynamic-coded filter fusion (DCFF), to derive compact CNNs in a computation-economical and regularization-free manner for efficient image classification. Each filter in our DCFF is first given an inter-similarity distribution with a temperature parameter as a filter proxy, on top of which, a fresh Kullback-Leibler divergence based dynamic-coded criterion is proposed to evaluate the filter importance. In contrast to simply keeping high-score filters in other methods, we propose the concept of filter fusion, i.e., the weighted averages using the assigned proxies, as our preserved filters. We obtain a one-hot inter-similarity distribution as the temperature parameter approaches infinity. Thus, the relative importance of each filter can vary along with the training of the compact CNN, leading to dynamically changeable fused filters without both the dependency on the pretrained model and the introduction of sparse constraints. Extensive experiments on classification benchmarks demonstrate the superiority of our DCFF over the compared counterparts. For example, our DCFF derives a compact VGGNet-16 with only 72.77M FLOPs and 1.06M parameters while reaching top-1 accuracy of 93.47% on CIFAR-10. A compact ResNet-50 is obtained with 63.8% FLOPs and 58.6% parameter reductions, retaining 75.60% top-1 accuracy on ILSVRC-2012. Our code, narrower models and training logs are available at https://github.com/lmbxmu/DCFF.
Mingbao Lin, Bohong Chen 0001, Fei Chao 0001, Rongrong Ji
IEEE Trans. Pattern Anal. Mach. Intell.3
2023 SiMaN: Sign-to-Magnitude Network Binarization
abstract
Binary neural networks (BNNs) have attracted broad research interest due to their efficient storage and computational ability. Nevertheless, a significant challenge of BNNs lies in handling discrete constraints while ensuring bit entropy maximization, which typically makes their weight optimization very difficult. Existing methods relax the learning using the sign function, which simply encodes positive weights into +1s, and -1s otherwise. Alternatively, we formulate an angle alignment objective to constrain the weight binarization to$\lbrace$0,+1$\rbrace$to solve the challenge. In this paper, we show that our weight binarization provides an analytical solution by encoding high-magnitude weights into +1s, and 0 s otherwise. Therefore, a high-quality discrete solution is established in a computationally efficient manner without the sign function. We prove that the learned weights of binarized networks roughly follow a Laplacian distribution that does not allow entropy maximization, and further demonstrate that it can be effectively solved by simply removing the$\ell _{2}$regularization during network training. Our method, dubbed sign-to-magnitude network binarization (SiMaN), is evaluated on CIFAR-10 and ImageNet, demonstrating its superiority over the sign-based state-of-the-arts. Our source code, experimental settings, training logs and binary models are available athttps://github.com/lmbxmu/SiMaN.
Mingbao Lin, Rongrong Ji, Baochang Zhang 0001, Fei Chao 0001, Chia-Wen Lin, Ling Shao 0001
IEEE Trans. Pattern Anal. Mach. Intell.5
2023 1xN Pattern for Pruning Convolutional Neural Networks
abstract
Though network pruning receives popularity in reducing the complexity of convolutional neural networks (CNNs), it remains an open issue to concurrently maintain model accuracy as well as achieve significant speedups on general CPUs. In this paper, we propose a novel 1×N pruning pattern to break this limitation. In particular, consecutive N output kernels with the same input channel index are grouped into one block, which serves as a basic pruning granularity of our pruning pattern. Our 1×N pattern prunes these blocks considered unimportant. We also provide a workflow of filter rearrangement that first rearranges the weight matrix in the output channel dimension to derive more influential blocks for accuracy improvements and then applies similar rearrangement to the next-layer weights in the input channel dimension to ensure correct convolutional operations. Moreover, the output computation after our 1×N pruning can be realized via a parallelized block-wise vectorized operation, leading to significant speedups on general CPUs. The efficacy of our pruning pattern is proved with experiments on ILSVRC-2012. For example, given the pruning rate of 50% and N=4, our pattern obtains about 3.0% improvements over filter pruning in the top-1 accuracy of MobileNet-V2. Meanwhile, it obtains 56.04ms inference savings on Cortex-A7 CPU over weight pruning. Our project is made available at https://github.com/lmbxmu/1xN.
Mingbao Lin, Yuxin Zhang 0002, Bohong Chen 0001, Fei Chao 0001, Mengdi Wang 0001, Yonghong Tian 0001, Rongrong Ji
IEEE Trans. Pattern Anal. Mach. Intell.5
2023 Lottery Jackpots Exist in Pre-Trained Models
abstract
Network pruning is an effective approach to reduce network complexity with acceptable performance compromise. Existing studies achieve the sparsity of neural networks via time-consuming weight training or complex searching on networks with expanded width, which greatly limits the applications of network pruning. In this paper, we show that high-performing and sparse sub-networks without the involvement of weight training, termed "lottery jackpots", exist in pre-trained models with unexpanded width. Our presented lottery jackpots are traceable through empirical and theoretical outcomes. For example, we obtain a lottery jackpot that has only 10% parameters and still reaches the performance of the original dense VGGNet-19 without any modifications on the pre-trained weights on CIFAR-10. Furthermore, we improve the efficiency for searching lottery jackpots from two perspectives. First, we observe that the sparse masks derived from many existing pruning criteria have a high overlap with the searched mask of our lottery jackpot, among which, the magnitude-based pruning results in the most similar mask with ours. In compliance with this insight, we initialize our sparse mask using the magnitude-based pruning, resulting in at least 3× cost reduction on the lottery jackpot searching while achieving comparable or even better performance. Second, we conduct an in-depth analysis of the searching process for lottery jackpots. Our theoretical result suggests that the decrease in training loss during weight searching can be disturbed by the dependency between weights in modern networks. To mitigate this, we propose a novel short restriction method to restrict change of masks that may have potential negative impacts on the training loss, which leads to a faster convergence and reduced oscillation for searching lottery jackpots. Consequently, our searched lottery jackpot removes 90% weights in ResNet-50, while it easily obtains more than 70% top-1 accuracy using only 5 searching epochs on ImageNet.
Yuxin Zhang 0002, Mingbao Lin, Yunshan Zhong, Fei Chao 0001, Rongrong Ji
IEEE Trans. Pattern Anal. Mach. Intell.4
2023 Decoder Choice Network for Metalearning
abstract
Metalearning has been widely applied for implementing few-shot learning and fast model adaptation. Particularly, existing metalearning methods have been exploited to learn the control mechanism for gradient descent processes, in an effort to facilitate gradient-based learning in gaining high speed and generalization ability. This article presents a novel method that controls the gradient descent process of the model parameters in a neural network, by limiting the model parameters within a low-dimensional latent space. The main challenge for implementing this idea is that a decoder with many parameters may be required. To tackle this problem, the article provides an alternative design of the decoder with a structure that shares certain weights, thereby reducing the number of required parameters. In addition, this work combines ensemble learning with the proposed approach to improve the overall learning performance. Systematic experimental studies demonstrate that the proposed approach offers results superior to the state of the art in performing the Omniglot classification and miniImageNet classification tasks.
Fei Chao 0001, Longzhi Yang, Chih-Min Lin, Changjing Shang, Qiang Shen 0001
IEEE Trans. Cybern.2
2023 Fuzzy-Rough Intrigued Harmonic Discrepancy Clustering
abstract
Fuzzy clustering decomposes data into clusters using partial memberships by exploring the cluster structure information, which demonstrates the comparable performance for knowledge exploitation under the circumstance of information incompleteness. In general, this scheme considers the memberships of objects to cluster centroids and applies to clusters with the spherical distribution. In addition, the noises and outliers may significantly influence the clustering process; a common mitigation measure is the application of separate noise processing algorithms, but this usually introduces multiple parameters, which are challenging to be determined for different data types. This article proposes a new fuzzy-rough intrigued harmonic discrepancy clustering (HDC) algorithm by noting that fuzzy-rough sets offer a higher degree of uncertainty modeling for both vagueness and imprecision present in real-valued datasets. The HDC is implemented by introducing a novel concept of harmonic discrepancy, which effectively indicates the dissimilarity between a data instance and foreign clusters with their distributions fully considered. The proposed HDC is thus featured by a powerful processing ability on complex data distribution leading to enhanced clustering performance, particularly on noisy datasets, without the use of explicit noise handling parameters. The experimental results confirm the effectiveness of the proposed HDC, which generally outperforms the popular representative clustering algorithms on both synthetic and benchmark datasets, demonstrating the superiority of the proposed algorithm.
Guanli Yue, Yanpeng Qu, Longzhi Yang, Changjing Shang, Ansheng Deng, Fei Chao 0001, Qiang Shen 0001
IEEE Trans. Fuzzy Syst.6
2023 Learning Efficient GANs for Image Translation via Differentiable Masks and Co-Attention Distillation
abstract
Generative Adversarial Networks (GANs) have been widely-used in image translation, but their high computational and storage costs impede the deployment on mobile devices. Prevalent methods for CNN compression cannot be directly applied to GANs due to the specificity of GAN tasks and the unstable adversarial training. To solve these, in this paper, we introduce a novel GAN compression method, termed DMAD, by proposing a Differentiable Mask and a co-Attention Distillation. The former searches for a light-weight generator architecture in a training-adaptive manner. To overcome channel inconsistency when pruning the residual connections, an adaptive cross-block group sparsity is further incorporated. The latter simultaneously distills informative attention maps from both the generator and discriminator of a pre-trained model to the searched generator, effectively stabilizing the adversarial training of our light-weight model. Experiments show that DMAD can reduce the Multiply Accumulate Operations (MACs) of CycleGAN by 13x and that of Pix2Pix by 4x while retaining a comparable performance against the full model. Our code can be available at https://github.com/SJLeo/DMAD.
Mingbao Lin, Yan Wang 0059, Fei Chao 0001, Ling Shao 0001, Rongrong Ji
IEEE Trans. Multim.4
2023 A framework of blockchain-based secure and privacy-preserving E-government system
abstract
Abstract Electronic government (e-government) uses information and communication technologies to deliver public services to individuals and organisations effectively, efficiently and transparently. E-government is one of the most complex systems which needs to be distributed, secured and privacy-preserved, and the failure of these can be very costly both economically and socially. Most of the existing e-government systems such as websites and electronic identity management systems (eIDs) are centralized at duplicated servers and databases. A centralized management and validation system may suffer from a single point of failure and make the system a target to cyber attacks such as malware, denial of service attacks (DoS), and distributed denial of service attacks (DDoS). The blockchain technology enables the implementation of highly secure and privacy-preserving decentralized systems where transactions are not under the control of any third party organizations. Using the blockchain technology, exiting data and new data are stored in a sealed compartment of blocks (i.e., ledger) distributed across the network in a verifiable and immutable way. Information security and privacy are enhanced by the blockchain technology in which data are encrypted and distributed across the entire network. This paper proposes a framework of a decentralized e-government peer-to-peer (p2p) system using the blockchain technology, which can ensure both information security and privacy while simultaneously increasing the trust of the public sectors. In addition, a prototype of the proposed system is presented, with the support of a theoretical and qualitative analysis of the security and privacy implications of such system.
Noe Elisa, Longzhi Yang, Fei Chao 0001
Wirel. Networks3
2022 Neural Architecture Search with Representation Mutual Information
abstract
Performance evaluation strategy is one of the most important factors that determine the effectiveness and efficiency in Neural Architecture Search (NAS). Existing strategies, such as employing standard training or performance predictor, often suffer from high computational complexity and low generality. To address this issue, we propose to rank architectures by Representation Mutual Information (RMI). Specifically, given an arbitrary architecture that has decent accuracy, architectures that have high RMI with it always yield good accuracies. As an accurate performance indicator to facilitate NAS, RMI not only generalizes well to different search spaces, but is also efficient enough to evaluate architectures using only one batch of data. Building upon RMI, we further propose a new search algorithm termed RMI-NAS, facilitating with a theorem to guarantee the global optimal of the searched architecture. In particular, RMI-NAS first randomly samples architectures from the search space, which are then effectively classified as positive or negative samples by RMI. We then use these samples to train a random forest to explore new regions, while keeping track of the distribution of positive architectures. When the sample size is sufficient, the architecture with the largest probability from the aforementioned distribution is selected, which is theoretically proved to be the optimal solution. The architectures searched by our method achieve remarkable top-1 accuracies with the magnitude times faster search process. Besides, RMI-NAS also generalizes to different datasets and search spaces. Our code has been made available at https://git.openi.org.cn/PCL_AutoML/XNAS.
Xiawu Zheng, Lei Zhang 0001, Chenglin Wu 0001, Fei Chao 0001, Jianzhuang Liu, Wei Zeng 0006, Yonghong Tian 0001, Rongrong Ji
CVPR5
2022 Fine-grained Data Distribution Alignment for Post-Training Quantization
Yunshan Zhong, Mingbao Lin, Mengzhao Chen, Ke Li 0015, Yunhang Shen, Fei Chao 0001, Yongjian Wu 0001, Rongrong Ji
ECCV (11)6
2022 Dynamic Dual Trainable Bounds for Ultra-low Precision Super-Resolution Networks
Yunshan Zhong, Mingbao Lin, Xunchao Li, Ke Li 0015, Yunhang Shen, Fei Chao 0001, Yongjian Wu 0001, Rongrong Ji
ECCV (18)6
2022 A TOPSIS based Self-Organizing Double Loop Recurrent Broad Learning System for Uncertain Nonlinear Systems
abstract
This study proposes an efficient intelligent control structure for uncertain nonlinear systems. The controller is implemented by a sliding mode control framework including a modified broad leaning network (BLS) with a double-loop recurrent structure. In addition, the proposed BLS involves a self-organizing mechanism to increase or decrease the size of the BLS. The technique for order of preference by similarity to ideal solution (TOPSIS) method is used to build the self-organizing mechanism. Moreover, two dynamic thresholds of TOPSIS are automatically determined according to the stability of the controller. One dynamic threshold is used to consider whether to retain or remove existing network neurons in the BLS; and the other is used to generate new neurons, so as to meet the requirements of different control states and save computing resources. To improve the network's dynamic characteristics, a double-loop recurrent structure is further introduced into the self-organizing BLS. The Lyapunov stability function is used to ensure the stability of the control system. The proposed controller is applied to the simulation control of a nonlinear chaotic system and a three-link robot manipulator. The experimental results show that the proposed controller can achieve better control performance against other network-based controllers. The source code of this work is placed at https://github.com/wzhuang-xmu/SODLRBLS
Wei-Zhong Huang, Wei-Bin Hong, Hong-Rui He, Fei Chao 0001, Longzhi Yang, Chih-Min Lin, Xiang Chang, Changjiang Shang, Qiang Shen 0001
IJCNN5
2022 Sundial-GAN: A Cascade Generative Adversarial Networks Framework for Deciphering Oracle Bone Inscriptions
abstract
Oracle Bone Inscription (OBI) is an early hieroglyph in China, which is the most famous ancient writing system in the world. However, only a small number of OBI characters have been fully deciphered today. Chinese characters have different forms in different historical stages; therefore, it is very difficult to directly translate OBI characters to modern Chinese characters due to the long historic evolutionary process. In this paper, we propose a cascade generative adversarial networks (GAN) framework for deciphering OBI characters, named "Sundial-GAN'', which is a cascaded structure to simulate Chinese characters' evolutionary process from an OBI character to its potential modern Chinese character. We select four representative stages in the evolutionary process of OBI, each of which is implemented by an individual GAN structure based on the characteristics of each evolutionary stage. These structures are cascaded in sequence to accurately simulate the Chinese characters' evolutionary process. For each input OBI character, Sundial-GAN can successfully generate the input's different forms at the four historical stages. Extensive experiments and comparisons demonstrate that generated characters at each stage have high similarities with real existing characters; therefore, the proposed method can significantly improve the efficiency and accuracy of OBI deciphering for archaeological researchers. Compared to direct image-to-image translation methods, our approach allows for a smoother translation process, a better grasp of details, and more effective avoiding random mappings in GANs.
Xiang Chang, Fei Chao 0001, Changjing Shang, Qiang Shen 0001
ACM Multimedia2
2022 Searching Lightweight Neural Network for Image Signal Processing
abstract
Recently, it has been shown that the traditional Image Signal Processing (ISP) can be replaced by deep neural networks due to their superior performance. However, most of these networks require heavy computation burden and thus are far from sufficient to be deployed on resource-limited platforms, including but not limited to mobile devices and FPGA. To tackle this challenge, we propose an automated search framework that derives ISP models with high image quality while satisfying the low-computation requirement. To reduce the search cost, we adopt the weight-sharing strategy by introducing a supernet and decouple the architecture search into two stages, supernet training and hard-aware evolutionary search. With the proposed framework, we can train the ISP model once and quickly find high-performance but low-computation models on multiple devices. Experiments demonstrate that the searched ISP models have an excellent trade-off between image quality and model complexity, i.e., achieve compelling reconstruction quality with more than 90% reduction in FLOPs as compared to the state-of-the-art networks.
Haojia Lin, Lijiang Li, Xiawu Zheng, Fei Chao 0001, Rongrong Ji
ACM Multimedia4
2022 Learning Best Combination for Efficient N: M Sparsity
abstract
By forcing N out of M consecutive weights to be non-zero, the recent N:M fine-grained network sparsity has received increasing attention with its two attractive advantages over traditional irregular network sparsity methods: 1) Promising performance at a high sparsity. 2) Significant speedups when performed on NVIDIA A100 GPUs. Current implementation on N:M sparsity requires a tedious pre-training phase or computationally heavy from-scratch training. To circumvent these problems, this paper presents an efficient solution for achieving N:M fine-grained sparsity from scratch. Specifically, we first make a re-formulation to convert the N:M fine-grained sparsity into a combinatorial problem, in which, the object falls into choosing the best weight combination among $C_M^N$ candidates. Then, we equip each combination with a learnable importance score, which can be jointly optimized along with its associated weights. Through rigorous proof, we demonstrate that the magnitude of the optimized score well reflects the importance of its corresponding weights combination to the training loss. Therefore, by gradually removing combinations with smaller scores till the best one is left, N:M fine-grained sparsity can be efficiently optimized during the normal training phase without any extra expenditure. Comprehensive experimental results have demonstrated that our proposed method for learning best combination, dubbed as LBC, consistently increases the efficacy of the off-the-shelf N:M methods across varying networks and datasets. Our project is released at https://github.com/zyxxmu/LBC.
Yuxin Zhang 0002, Mingbao Lin, Zhihang Lin, Yiting Luo, Ke Li 0015, Fei Chao 0001, Yongjian Wu 0001, Rongrong Ji
NeurIPS6
2022 Intelligent wavelet fuzzy brain emotional controller using dual function-link network for uncertain nonlinear control systems
Tuan-Tu Huynh, Chih-Min Lin, Nguyen-Quoc-Khanh Le, Mai The Vu, Ngoc Phi Nguyen, Fei Chao 0001
Appl. Intell.6
2022 Error controlled actor-critic
Xingen Gao, Fei Chao 0001, Changle Zhou, Zhen Ge, Longzhi Yang, Xiang Chang, Changjing Shang, Qiang Shen 0001
Inf. Sci.2
2022 A Type 2 wavelet brain emotional learning network with double recurrent loops based controller for nonlinear systems
Zi-Qi Wang, Lijiang Li, Fei Chao 0001, Chih-Min Lin, Longzhi Yang, Changle Zhou, Xiang Chang, Changjing Shang, Qiang Shen 0001
Knowl. Based Syst.3
2022 A recurrent wavelet-based brain emotional learning network controller for nonlinear systems
Juncheng Zhang, Fei Chao 0001, Hualin Zeng, Chih-Min Lin, Longzhi Yang
Soft Comput.2
2022 Low-Cost Inertial Measurement Unit Calibration With Nonlinear Scale Factors
abstract
Inertial measurement units (IMUs) have been widely used to provide accurate location and movement measurement solutions, along with the advances of modern manufacturing technologies. The scale factors of accelerometers and gyroscopes are linear when the range of the sensors are reasonably small, but the factor becomes nonlinear when the range gets much bigger. Based on this observation, this article presents a calibration method for low-cost IMU by effectively deriving the nonlinear scale factors of the sensors. Two motion patterns of the sensor on a rigid object are moved to collect data for calibration: One motion pattern is to upcast and rotate the rigid object, and another pattern is to place the rigid object on a stable base in different attitudes. The rotation motion produces centripetal and Coriolis force, which increases the measurement range of accelerometers. Four cost functions with different weight factors and two sets of data are utilized to optimize the IMU parameters. The weight factor comes from derived formula with input values which are the variance of the noise of the sampled data. The proposed approach was validated and evaluated on both synthetic and real-world data sets, and the experimental results demonstrated the superiority of the proposed approach in improving the accuracy of IMU for long-range use. In particular, the errors of acceleration and angular velocity led by our algorithm are significantly smaller than those resulted from the existing approaches using the same testing data sets, demonstrating a remarkable improvement of 64.12% and 47.90%, respectively.
Xin Zhang 0090, Changle Zhou, Fei Chao 0001, Chih-Min Lin, Longzhi Yang, Changjing Shang, Qiang Shen 0001
IEEE Trans. Ind. Informatics3
2022 Self-Organizing Double Function-Link Fuzzy Brain Emotional Control System Design for Uncertain Nonlinear Systems
abstract
This article aims to propose a more efficient control algorithm for uncertain nonlinear systems. An intelligent self-organizing double function-link fuzzy brain emotional control system is proposed which comprises a self-organizing double function-link fuzzy brain emotional controller (SDFLFBEC) and a compensation controller. The proposed SDFLFBEC consists of four substructures and a fuzzy inference system. The substructures are the prefrontal cortex, the amygdala, a double function-link network (FLN) and a self-organizing structure. The prefrontal cortex and the amygdala networks work as a mathematical form that presumes the judgment and emotion of a brain. Specifically, a new double FLN is designed to support the above networks for updating their weights. Next, a self-organizing structure can automatically add or prune the layers to achieve efficient network structure. In addition, the fuzzy inference rules are presented to explain the inference processes of the amygdala and orbitofrontal networks. From the above factors, the proposed SDFLFBEC can effectively reduce the tracking error and achieve favorable control performance. The parameters of the control system are adjusted online using the derived adaptation laws that are taken from a Lyapunov function so that the stability of the system is ensured. Simulation studies of a biped robot and the experimental results of a magnetic levitation system are employed to validate the effectiveness and superiority of the proposed SDFLFBEC.
Tuan-Tu Huynh, Chih-Min Lin, Tien-Loc Le, Nguyen-Quoc-Khanh Le, Van-Phong Vu, Fei Chao 0001
IEEE Trans. Syst. Man Cybern. Syst.6
2021 Towards Compact CNNs via Collaborative Compression
abstract
Channel pruning and tensor decomposition have received extensive attention in convolutional neural network compression. However, these two techniques are traditionally deployed in an isolated manner, leading to significant accuracy drop when pursuing high compression rates. In this paper, we propose a Collaborative Compression (CC) scheme, which joints channel pruning and tensor decomposition to compress CNN models by simultaneously learning the model sparsity and low-rankness. Specifically, we first investigate the compression sensitivity of each layer in the network, and then propose a Global Compression Rate Optimization that transforms the decision problem of compression rate into an optimization problem. After that, we propose multi-step heuristic compression to remove redundant compression units step-by-step, which fully considers the effect of the remaining compression space (i.e., unremoved compression units). Our method demonstrates superior performance gains over previous ones on various datasets and backbone architectures. For example, we achieve 52.9% FLOPs reduction by removing 48.4% parameters on ResNet-50 with only a Top-1 accuracy drop of 0.56% on ImageNet 2012.
Shaohui Lin, Jianzhuang Liu, Qixiang Ye, Mengdi Wang 0001, Fei Chao 0001, Fan Yang 0016, Jincheng Ma, Qi Tian 0001, Rongrong Ji
CVPR6
2021 CDP: Towards Optimal Filter Pruning via Class-wise Discriminative Power
abstract
Neural network pruning has shown promising performance in reducing computational complexity and facilitate the deployment of deep neural networks on resource-limited edge devices. Most existing pruning methods focus on the indicators of the filter's weight, gradient, or feature map and regard the weak or similar filters as network redundancy. In contrast, the representation of discriminative power is also a fundamental attribute that analog neural networks to have extraordinary performance in various tasks. However, such representation is neglected in existing works. Alternatively, we propose a novel filter pruning strategy via class-wise discriminative power (CDP). Unlike the previous methods, CDP treats the filters that always yield large or small activation values as redundant and reserves the filters that show different magnitudes in activations as they yield high discriminative power. We further propose to obtain such discriminative power by employing the widely-used Term Frequency-Inverse Document Frequency (TF-IDF) on feature representations across classes. Specifically, the output of a filter is considered as a word, and the whole feature map is considered as a document. Then, TF-IDF is used to generate the relevant score between words and all documents. If a filter has low TF-IDF scores is less discriminate and can be pruned. Thus, the filters with high TF-IDF scores are reserved. To our best knowledge, this is the first work that prunes neural networks through class-wise discriminative power and measures such power by introducing TF-IDF in feature representation among different classes. Without any iterative process, CDP achieves better compression trade-offs comparing to the state-of-the-art compression algorithms. For instance, in VGG-16, we achieve a 68.05%-FLOPs reduction, with a 94.86% Top-1 accuracy on CIFAR-10. Specifically, we compress a 90.12%-FLOPs reduction VGG-16, even retains 93.30% Top-1 accuracy on CIFAR-10. The code is available at https://github.com/Tianshuo-Xu/CDP-Towards-Optimal-Filter-Pruning-via-Class-wise-Discriminative-Power.git
Tianshuo Xu, Yuhang Wu 0004, Xiawu Zheng, Teng Xi, Errui Ding, Fei Chao 0001, Rongrong Ji
ACM Multimedia7
2021 Revisiting Discriminator in GAN Compression: A Generator-discriminator Cooperative Compression Scheme
abstract
Recently, a series of algorithms have been explored for GAN compression, which aims to reduce tremendous computational overhead and memory usages when deploying GANs on resource-constrained edge devices. However, most of the existing GAN compression work only focuses on how to compress the generator, while fails to take the discriminator into account. In this work, we revisit the role of discriminator in GAN compression and design a novel generator-discriminator cooperative compression scheme for GAN compression, termed GCC. Within GCC, a selective activation discriminator automatically selects and activates convolutional channels according to a local capacity constraint and a global coordination constraint, which help maintain the Nash equilibrium with the lightweight generator during the adversarial training and avoid mode collapse. The original generator and discriminator are also optimized from scratch, to play as a teacher model to progressively refine the pruned generator and the selective activation discriminator. A novel online collaborative distillation scheme is designed to take full advantage of the intermediate feature of the teacher generator and discriminator to further boost the performance of the lightweight generator. Extensive experiments on various GAN-based generation tasks demonstrate the effectiveness and generalization of GCC. Among them, GCC contributes to reducing 80% computational costs while maintains comparable performance in image translation tasks.
Jie Wu 0032, Xuefeng Xiao 0001, Fei Chao 0001, Xudong Mao, Rongrong Ji
NeurIPS4
2021 Job shop planning and scheduling for manufacturers with manual operations
abstract
Abstract Job shop scheduling systems are widely employed to optimise the efficiency of machine utilisation in the manufacturing industry, by searching the most cost‐effective permutation of job operations based on the cost of each operation on each compatible machine and the relations between job operations. Such systems are paralysed when the cost of operations are not predictable led by the involvement of complex manual operations. This paper proposes a new genetic algorithm‐based job shop scheduling system by integrating a fuzzy learning and inference subsystem in an effort to address this limitation. In particular, the fuzzy subsystem adaptively estimates the completion time and thus cost of each manual task under different conditions based on a knowledge base that is initialised by domain experts and then constantly updated based on its built‐in learning ability and adaptability. The manufacturer of Point of Sale and Point of Purchase products has been utilised in this paper as an example case for both theoretical discussion and experimental study. The experimental results demonstrate the promising of the proposed system in improving the efficiency of manual manufacturing operations.
Longzhi Yang, Jie Li 0021, Fei Chao 0001, Phil Hackney, Mark Flanagan
Expert Syst. J. Knowl. Eng.3
2021 Automatic stroke generation for style-oriented robotic Chinese calligraphy
Fei Chao 0001, Longzhi Yang, Xiang Chang, Chih-Min Lin, Changle Zhou, Varadarajan Vijayakumar 0001, Changjing Shang
Future Gener. Comput. Syst.3
2021 DK-CNNs: Dynamic kernel convolutional neural networks
Fei Chao 0001, Chih-Min Lin, Changle Zhou, Changjing Shang
Neurocomputing2
2021 Feature grouping and selection: A graph-based approach
Fei Chao 0001, Neil Mac Parthaláin, Qiang Shen 0001
Inf. Sci.2
2021 Exclusive lasso-based k-nearest-neighbor classification
Yanpeng Qu, Changjing Shang, Longzhi Yang, Fei Chao 0001, Qiang Shen 0001
Neural Comput. Appl.5
2021 Interval type-2 fuzzy brain emotional control design for the synchronization of 4D nonlinear hyperchaotic systems
Tuan-Tu Huynh, Chih-Min Lin, Tien-Loc Le, Mai The Vu, Fei Chao 0001
Soft Comput.5
2021 Visual-Guided Robotic Object Grasping Using Dual Neural Network Controllers
abstract
It has been a challenging task for a robotic arm to accurately reach and grasp objects, which has drawn much research attention. This article proposes a robotic hand-eye coordination system by simulating the human behavior pattern to achieve a fast and robust reaching ability. This is achieved by two neural-network-based controllers, including a rough reaching movement controller implemented by a pretrained radial basis function for rough reaching movements, and a correction movement controller built from a specifically designed brain emotional nesting network (BENN) for smooth correction movements. In particular, the proposed BENN is designed with high nonlinear mapping ability, with its adaptive laws derived from the Lyapunov stability theorem; from this, the robust tracking performance and accordingly the stability of the proposed control system are guaranteed by the utilization of the H∞control approach. The proposed BENN is validated and evaluated by a chaos synchronization simulation, and the overall control system by object grasping tasks through a physical robotic arm in a real-world environment. The experimental results demonstrate the superiority of the proposed control system in reference to those with single neural networks.
Wubing Fang, Fei Chao 0001, Chih-Min Lin, Dajun Zhou, Longzhi Yang, Xiang Chang, Qiang Shen 0001, Changjing Shang
IEEE Trans. Ind. Informatics2
2020 A Comparative Study of Genetic Algorithm and Particle Swarm optimisation for Dendritic Cell Algorithm
abstract
Dendritic cell algorithm (DCA) is a class of artificial immune systems that was originally developed for anomaly detection in networked systems and later as a general binary classifier. Conventionally, in its life cycle, the DCA goes through four phases including feature categorisation into artificial signals, context detection of data items, context assignment, and finally labeling of data items as either abnormal or normal class. During the context detection phase, the DCA requires users to manually pre-define the parameters used by its weighted function to process the signals and data items. Notice that the manual derivation of the parameters of the DCA cannot guarantee the optimal set of weights being used, research attention has thus been attracted to the optimisation of the parameters. This paper reports a systematic comparative study between Genetic algorithm (GA) and Particle Swarm optimisation (PSO) on parameter optimisation for DCA. In order to evaluate the performance of GADCA and PSO-DCA, twelve publicly available datasets from UCI machine learning repository were employed. The performance results based on the computational time, classification accuracy, sensitivity, F-measure, and precision show that, the GA-DCA overall outperforms PSO-DCA for most of the datasets.
Noe Elisa, Longzhi Yang, Fei Chao 0001, Nitin Naik
CEC3
2020 Improving Deep Learning based Optical Character Recognition via Neural Architecture Search
abstract
Optical character rcecognition (OCR) is a process of converting images of typed, handwritten or printed text into machine-encoded one. In recent years, the methods represented by deep learning have greatly improved the performance of OCR systems, but the main challenges of such systems are 1) to accurately perform text detection in complex scenes and 2) to identify and set the optimal parameters to optimize the performance of the system. In this paper, we propose an OCR method based on Neural Architecture Search technique, called AutOCR. The characteristic of the proposed method is the automatic design of text detection framework using an evolutionary computation neural architecture search method. This design can not only accurately recognize the text in a complex environment, but also avoid the process of experts participating in parameter adjustment. We compared it with different methods, and the experimental results proved the effectiveness of our method.
Zhenyao Zhao, Min Jiang 0005, Shihui Guo, Zhenzhong Wang, Fei Chao 0001, Kay Chen Tan
CEC5
2020 A Novel Self-Organizing Emotional CMAC Network for Robotic Control*
abstract
This paper proposes a self-organizing control system for uncertain nonlinear systems. The proposed neural network is composed of a conventional brain emotional learning network (BEL) and a cerebellar model articulation controller network (CMAC). The input value of the network is feed to a BEL channel and a CMAC channel. The output of the network is generated by the comprehensive action of the two channels. The structure of the network is dynamic, using a self-organizing algorithm allows increasing or decreasing weight layers. The parameters of the proposed network are on-line tuned by the brain emotional learning rules; the updating rules of CMAC and the robust controller are derived from the Lyapunov function; in addition, stability analysis theory is used to guaranty the proposed controller's convergence. A simulated mobile robot is applied to prove the effectiveness of the proposed control system. By comparing with the performance of other neural-network-based control systems, the proposed network produces better performance.
Juncheng Zhang, Quanfeng Li, Xiang Chang, Fei Chao 0001, Chih-Min Lin, Longzhi Yang, Tuan-Tu Huynh, Changle Zhou, Changjing Shang
IJCNN4
2020 Integration of an actor-critic model and generative adversarial networks for a Chinese calligraphy robot
Changle Zhou, Fei Chao 0001, Longzhi Yang, Chih-Min Lin, Changjing Shang
Neurocomputing3
2020 GANCCRobot: Generative adversarial nets based chinese calligraphy robot
Changle Zhou, Fei Chao 0001, Longzhi Yang, Chih-Min Lin, Changjing Shang
Inf. Sci.3
2020 Type-2 Fuzzy Hybrid Controller Network for Robotic Systems
abstract
Dynamic control, including robotic control, faces both the theoretical challenge of obtaining accurate system models and the practical difficulty of defining uncertain system bounds. To facilitate such challenges, this paper proposes a control system consisting of a novel type of fuzzy neural network and a robust compensator controller. The new fuzzy neural network is implemented by integrating a number of key components embedded in a Type-2 fuzzy cerebellar model articulation controller (CMAC) and a brain emotional learning controller (BELC) network, thereby mimicking an ideal sliding mode controller. The system inputs are fed into the neural network through a Type-2 fuzzy inference system (T2FIS), with the results subsequently piped into sensory and emotional channels which jointly produce the final outputs of the network. That is, the proposed network estimates the nonlinear equations representing the ideal sliding mode controllers using a powerful compensator controller with the support of T2FIS and BELC, guaranteeing robust tracking of the dynamics of the controlled systems. The adaptive dynamic tuning laws of the network are developed by exploiting the popular brain emotional learning rule and the Lyapunov function. The proposed system was applied to a robot manipulator and a mobile robot, demonstrating its efficacy and potential; and a comparative study with alternatives indicates a significant improvement by the proposed system in performing the intelligent dynamic control.
Fei Chao 0001, Dajun Zhou, Chih-Min Lin, Longzhi Yang, Changle Zhou, Changjing Shang
IEEE Trans. Cybern.1
2020 Histogram of Fuzzy Local Spatio-Temporal Descriptors for Video Action Recognition
abstract
Feature extraction plays a vital role in visual action recognition. Many existing gradient-based feature extractors, including histogram of oriented gradients, histogram of optical flow, motion boundary histograms, and histogram of motion gradients, build histograms for representing different actions over the spatio-temporal domain in a video. However, these methods require to set the number of bins for information aggregation in advance. Varying numbers of bins usually lead to inherent uncertainty within the process of pixel voting with regard to the bins in the histogram. This article proposes a novel method to handle such uncertainty by fuzzifying these feature extractors. The proposed approach has two advantages: it better represents the ambiguous boundaries between the bins and, thus, the fuzziness of the spatio-temporal visual information entailed in videos; and the contribution of each pixel is flexibly controlled by a fuzziness parameter for various scenarios. The proposed family of fuzzy descriptors and a combination of them are evaluated on two publicly available datasets, demonstrating that the proposed approach outperforms the original counterparts and other state-of-the-art methods.
Zheming Zuo, Longzhi Yang, Yonghuai Liu, Fei Chao 0001, Ran Song 0001, Yanpeng Qu
IEEE Trans. Ind. Informatics4
2019 Adaptive Activation Function Generation for Artificial Neural Networks through Fuzzy Inference with Application in Grooming Text Categorisation
abstract
The activation function is introduced to determine the output of neural networks by mapping the resulting values of neurons into a specific range. The activation functions often suffer from ‘gradient vanishing’, ‘non zero-centred function outputs’, ‘exploding gradients’, and ‘dead neurons’, which may lead to deterioration in the classification performance. This paper proposes an activation function generation approach using the Takagi-Sugeno-Kang inference in an effort to address such challenges. In addition, the proposed method further optimises the coefficients in the activation function using the genetic algorithm such that the activation function can adapt to different applications. This approach has been applied to a digital forensics application of online grooming detection. The evaluations confirm the superiority of the proposed activation function for online grooming detection using an unbalanced data set.
Zheming Zuo, Jie Li 0021, Bo Wei 0003, Longzhi Yang, Fei Chao 0001, Nitin Naik
FUZZ-IEEE5
2019 A recurrent emotional CMAC neural network controller for vision-based mobile robots
Wubing Fang, Fei Chao 0001, Longzhi Yang, Chih-Min Lin, Changjing Shang, Changle Zhou, Qiang Shen 0001
Neurocomputing2
2019 A data-driven robotic Chinese calligraphy system using convolutional auto-encoder and differential evolution
Xingen Gao, Changle Zhou, Fei Chao 0001, Longzhi Yang, Chih-Min Lin, Tao Xu 0045, Changjing Shang, Qiang Shen 0001
Knowl. Based Syst.3
2019 Use of Automatic Chinese Character Decomposition and Human Gestures for Chinese Calligraphy Robots
abstract
Conventional Chinese calligraphy robots often suffer from the limited sizes of predefined font databases, which prevent the robots from writing new characters. This paper presents a robotic handwriting system to address such limitations, which extracts Chinese characters from textbooks and uses a robot's manipulator to write the characters in a different style. The key technologies of the proposed approach include the following: 1) automatically decomposing Chinese characters into strokes using Harris corner detection technology and 2) matching the decomposed strokes to robotic writing trajectories learned from human gestures. Briefly, the system first decomposes a given Chinese character into a set of strokes and obtains the stroke trajectory writing ability by following the gestures performed by a human demonstrator. Then, it applies a stroke classification method that recognizes the decomposed strokes as robotic writing trajectories. Finally, the robot arm is driven to follow the trajectories and thus write the Chinese character. Seven common Chinese characters have been used in an experiment for system validation and evaluation. The experimental results demonstrate the power of the proposed system, given that the robot successfully wrote all the testing characters in the given Chinese calligraphic style.
Fei Chao 0001, Chih-Min Lin, Longzhi Yang, Huosheng Hu, Changle Zhou
IEEE Trans. Hum. Mach. Syst.1
2018 Interval Type-2 TSK+ Fuzzy Inference System
abstract
Type-2 fuzzy sets and systems can better handle uncertainties compared to its type-1 counterpart, and the widely applied Mamdani and TSK fuzzy inference approaches have been both extended to support interval type-2 fuzzy sets. Fuzzy interpolation enhances the conventional Mamdani and TKS fuzzy inference systems, which not only enables inferences when inputs are not covered by an incomplete or sparse rule base but also helps in system simplification for very complex problems. This paper extends the recently proposed fuzzy interpolation approach TSK+ to allow the utilization of interval type-2 TSK fuzzy rule bases. One illustrative case based on an example problem from the literature demonstrates the working of the proposed system, and the application on the cart centering problem reveals the power of the proposed system. The experimental investigation confirmed that the proposed approach is able to perform fuzzy inferences using either dense or sparse interval type-2 TSK rule bases with promising results generated.
Jie Li 0021, Longzhi Yang, Xin Fu 0003, Fei Chao 0001, Yanpeng Qu
FUZZ-IEEE4
2018 Generative Adversarial Nets in Robotic Chinese Calligraphy
abstract
Conventional approaches of robotic writing of Chinese character strokes often suffer from limited font generation methods, and thus the writing results often lack of diversity. This has seriously restricted the high quality writing ability of robots. This paper proposes a generative adversarial nets-based calligraphic robotic framework, which enables a robot to learn writing fundamental Chinese strokes with rich diversity and good originality. In particular, the framework considers the learning process of robotic writing as an adversarial procedure which is implemented by three interactive modules including a stroke generation module, a stroke discriminative module and a training module. Noting that the stroke generative module included in the conventional generative adversarial nets cannot solve the non-differentiable problem, the policy gradient commonly used in reinforcement learning is thus adapted in this work to train the generative module by regarding the outputs from the discriminative module as rewards. Experimental results demonstrate that the proposed framework allows a calligraphic robot to successfully write fundamental Chinese strokes with good quality in various styles. The experiment also suggests the proposed approach can achieve human-level stroke writing quality without the requirement of a performance evaluation system. This approach therefore significantly boosts the robotic autonomous creation ability.
Fei Chao 0001, Jitu Lv, Dajun Zhou, Longzhi Yang, Chih-Min Lin, Changjing Shang, Changle Zhou
ICRA1
2018 Breast Cancer Diagnosis Using K-Means Type-2 Fuzzy Neural Network
abstract
This paper aims to design a classifier using the K-means clustering algorithm and the interval type-2 fuzzy neural network (IT2FNN). Firstly, the K-means clustering algorithm will classify the training data into k groups, according to its characteristics. After that, the IT2FNN will train the k classifiers' structure with these data. The testing data will be also determined that they will belong to which classifier. With this parallel structure, the performance of the proposed classifier is competitive with some state-of-the-art techniques. The parameter adaptive laws of the network are derived by using the steepest descent gradient approach. The convergence and stability of the proposed algorithm is guaranteed using the Lyapunov function. The system performance is evaluated by the breast cancer datasets of the University of California at Irvine (UCI). Comparison with other classifiers is also conducted. The experimental results have shown the effectiveness of the proposed method.
Tien-Loc Le, Tuan-Tu Huynh, Chih-Min Lin, Fei Chao 0001
SMC4
2018 Exploring spatial-frequency-sequential relationships for motor imagery classification with recurrent neural network
abstract
BACKGROUND: Conventional methods of motor imagery brain computer interfaces (MI-BCIs) suffer from the limited number of samples and simplified features, so as to produce poor performances with spatial-frequency features and shallow classifiers. METHODS: Alternatively, this paper applies a deep recurrent neural network (RNN) with a sliding window cropping strategy (SWCS) to signal classification of MI-BCIs. The spatial-frequency features are first extracted by the filter bank common spatial pattern (FB-CSP) algorithm, and such features are cropped by the SWCS into time slices. By extracting spatial-frequency-sequential relationships, the cropped time slices are then fed into RNN for classification. In order to overcome the memory distractions, the commonly used gated recurrent unit (GRU) and long-short term memory (LSTM) unit are applied to the RNN architecture, and experimental results are used to determine which unit is more suitable for processing EEG signals. RESULTS: Experimental results on common BCI benchmark datasets show that the spatial-frequency-sequential relationships outperform all other competing spatial-frequency methods. In particular, the proposed GRU-RNN architecture achieves the lowest misclassification rates on all BCI benchmark datasets. CONCLUSION: By introducing spatial-frequency-sequential relationships with cropping time slice samples, the proposed method gives a novel way to construct and model high accuracy and robustness MI-BCIs based on limited trials of EEG signals.
Tian-jian Luo 0001, Changle Zhou, Fei Chao 0001
BMC Bioinform.3
2018 Use of human gestures for controlling a mobile robot via adaptive CMAC network and fuzzy logic controller
Dajun Zhou, Minghui Shi, Fei Chao 0001, Chih-Min Lin, Longzhi Yang, Changjing Shang, Changle Zhou
Neurocomputing3
2018 Electroencephalogram-based brain-computer interface for the Chinese spelling system: a survey
abstract
Electroencephalogram (EEG) based brain-computer interfaces allow users to communicate with the external environment by means of their EEG signals, without relying on the brain’s usual output pathways such as muscles. A popular application for EEGs is the EEG-based speller, which translates EEG signals into intentions to spell particular words, thus benefiting those suffering from severe disabilities, such as amyotrophic lateral sclerosis. Although the EEG-based English speller (EEGES) has been widely studied in recent years, few studies have focused on the EEG-based Chinese speller (EEGCS). The EEGCS is more difficult to develop than the EEGES, because the English alphabet contains only 26 letters. By contrast, Chinese contains more than 11 000 logographic characters. The goal of this paper is to survey the literature on EEGCS systems. First, the taxonomy of current EEGCS systems is discussed to get the gist of the paper. Then, a common framework unifying the current EEGCS and EEGES systems is proposed, in which the concept of EEG-based choice acts as a core component. In addition, a variety of current EEGCS systems are investigated and discussed to highlight the advances, current problems, and future directions for EEGCS.
Minghui Shi, Changle Zhou, Jun Xie 0002, Shaozi Li, Qingyang Hong, Min Jiang 0005, Fei Chao 0001, Weifeng Ren, Xiangqian Liu, Dajun Zhou
Frontiers Inf. Technol. Electron. Eng.7
2018 The 16th Annual UK Workshop on Computational Intelligence
Plamen Angelov 0001, Changjing Shang, Fei Chao 0001
Soft Comput.3
2018 Fuzzy cerebellar model articulation controller network optimization via self-adaptive global best harmony search algorithm
Fei Chao 0001, Dajun Zhou, Chih-Min Lin, Changle Zhou, Minghui Shi, Dazhen Lin
Soft Comput.1
2018 Special issue on The 17th Annual UK Workshop on Computational Intelligence
Qingfu Zhang 0001, Fei Chao 0001
Soft Comput.2
2017 Dynamic QoS solution for enterprise networks using TSK fuzzy interpolation
abstract
The Quality of Services (QoS) is the measure of data transmission quality and service availability of a network, aiming to maintain the data, especially delay-sensitive data such as VoIP, to be transmitted over the network with the required quality. Major network device manufacturers have each developed their own smart dynamic QoS solutions, such as AutoQoS supported by Cisco, CoS (Class of Service) by Netgear devices, and QoS Maps on SROS (Secure Router Operating System) provided by HP, to maintain the service level of network traffic. Such smart QoS solutions usually only work for manufacture qualified devices and otherwise only a pre-defined static policy mapping can be applied. This paper presents a dynamic QoS solution based on the differentiated services (DiffServ) approach for enterprise networks, which is able to modify the priority level of a packet in real time by adjusting the value of Differentiated Services Code Point (DSCP) in Internet Protocol (IP) header of network packets. This is implemented by a 0-order TSK fuzzy model with a sparse rule base which is developed by considering the current network delay, application desired priority level and user current priority group. DSCP values are dynamically generated by the TSK fuzzy model and updated in real time. The proposed system has been evaluated in a real network environment with promising results generated.
Jie Li 0021, Longzhi Yang, Xin Fu 0003, Fei Chao 0001, Yanpeng Qu
FUZZ-IEEE4
2017 Integration of fuzzy CMAC and BELC networks for uncertain nonlinear system control
abstract
This paper develops a fuzzy adaptive control system consisting of a new type of fuzzy neural network and a robust controller for uncertain nonlinear systems. The new designed neural network contains the key mechanisms of a typical fuzzy CMAC network and a brain emotional learning controller network. First, the input values of the new network are delivered to a receptive field structure that is inspired from the fuzzy CMAC. Then, the values are divided into a sensory and an emotional channels; and the two channels interact with each other to generate the final outputs of the proposed network. The parameters of the proposed network are on-line tuned by the brain emotional learning rules; in addition, stability analysis theory is used to guaranty the proposed controller's convergence. In the experimentation, a “Duffing-Holmes” chaotic system and a simulated mobile robot are applied to verify the effectiveness and feasibility of the proposed control system. By comparing with the performances of other neural network based control systems, we believe our proposed network is capable of producing better control performances of complex uncertain nonlinear systems control.
Dajun Zhou, Fei Chao 0001, Chih-Min Lin, Longzhi Yang, Minghui Shi, Changle Zhou
FUZZ-IEEE2
2017 A robot calligraphy system: From simple to complex writing by human gestures
Fei Chao 0001, Xin Zhang 0090, Changjing Shang, Longzhi Yang, Changle Zhou, Huosheng Hu, Chih-Min Lin
Eng. Appl. Artif. Intell.1
2017 Generalized Adaptive Fuzzy Rule Interpolation
abstract
As a substantial extension to fuzzy rule interpolation that works based on two neighboring rules flanking an observation, adaptive fuzzy rule interpolation is able to restore system consistency when contradictory results are reached during interpolation. The approach first identifies the exhaustive sets of candidates, with each candidate consisting of a set of interpolation procedures which may jointly be responsible for the system inconsistency. Then, individual candidates are modified such that all contradictions are removed, and thus, interpolation consistency is restored. It has been developed on the assumption that contradictions may only be resulted from the underlying interpolation mechanism, and that all the identified candidates are not distinguishable in terms of their likelihood to be the real culprit. However, this assumption may not hold for real-world situations. This paper, therefore, further develops the adaptive method by taking into account observations, rules, and interpolation procedures, all as diagnosable and modifiable system components. In addition, given the common practice in fuzzy systems that observations and rules are often associated with certainty degrees, the identified candidates are ranked by examining the certainty degrees of its components and their derivatives. From this, the candidate modification is carried out based on such ranking. This study significantly improves the efficacy of the existing adaptive system by exploiting more information during both the diagnosis and modification processes.
Longzhi Yang, Fei Chao 0001, Qiang Shen 0001
IEEE Trans. Fuzzy Syst.2
2016 Towards sparse rule base generation for fuzzy rule interpolation
abstract
Fuzzy inference systems have been successfully applied to many real-world applications. Traditional fuzzy inference systems are only applicable to problems with dense rule bases by which the entire input domain is fully covered, whilst fuzzy rule interpolation (FRI) is also able to work with sparse rule bases that may not cover certain observations. Thanks to its ability to work with fewer rules, fuzzy rule interpolation approaches have also been utilised to reduce system complexity by removing those rules which can be approximated by their neighbouring ones for complex fuzzy models. A number of important fuzzy rule base generation approaches have been proposed in the literature, but the majority of these only target dense rule bases for traditional fuzzy inference systems. This paper proposes a novel sparse fuzzy rule base generation method to support FRI. The approach first identifies important rules that cannot be accurately approximated by their neighbouring ones to initialise the rule base. Then the raw rule base is optimised by fine-tuning the membership functions of the fuzzy sets. Experimentation is conducted to demonstrate the working principles of the proposed system, with results comparable to those of traditional methods.
Yao Tan, Jie Li 0021, Martin Wonders, Fei Chao 0001, Hubert P. H. Shum, Longzhi Yang
FUZZ-IEEE4
2016 Integration of classifier diversity measures for feature selection-based classifier ensemble reduction
Hualin Zeng, Fei Chao 0001, Chang Su 0006, Chih-Min Lin, Changle Zhou
Soft Comput.3
2016 Adaptive Filter Design Using Type-2 Fuzzy Cerebellar Model Articulation Controller
abstract
This paper aims to propose an efficient network and applies it as an adaptive filter for the signal processing problems. An adaptive filter is proposed using a novel interval type-2 fuzzy cerebellar model articulation controller (T2FCMAC). The T2FCMAC realizes an interval type-2 fuzzy logic system based on the structure of the CMAC. Due to the better ability of handling uncertainties, type-2 fuzzy sets can solve some complicated problems with outstanding effectiveness than type-1 fuzzy sets. In addition, the Lyapunov function is utilized to derive the conditions of the adaptive learning rates, so that the convergence of the filtering error can be guaranteed. In order to demonstrate the performance of the proposed adaptive T2FCMAC filter, it is tested in signal processing applications, including a nonlinear channel equalization system, a time-varying channel equalization system, and an adaptive noise cancellation system. The advantages of the proposed filter over the other adaptive filters are verified through simulations.
Chih-Min Lin, Ming-Shu Yang, Fei Chao 0001, Xiaomin Hu, Jun Zhang 0003
IEEE Trans. Neural Networks Learn. Syst.3
2015 Robotic Dance in Social Robotics - A Taxonomy
abstract
Robotic dance is an important topic in the field of social robotics. Its research has a vital significance to both humans and robotics. This paper presents a review of the state of the art in robotic dance. Robotic dance is classified into four categories: cooperative human-robot dance, imitation of human dance motions, synchronization for music, and creation of robotic choreography. The research methods in each category are discussed. Future research areas are highlighted.
Hua Peng, Changle Zhou, Huosheng Hu, Fei Chao 0001, Jing Li 0032
IEEE Trans. Hum. Mach. Syst.4
2014 A reduced classifier ensemble approach to human gesture classification for robotic Chinese handwriting
abstract
The paper presents an approach to applying a classifier ensemble to identify human body gestures, so as to control a robot to write Chinese characters. Robotic handwriting ability requires complicated robotic control algorithms. In particular, the Chinese handwriting needs to consider the relative positions of a character's strokes. This approach derives the font information from human gestures by using a motion sensing input device. Five elementary strokes are used to form Chinese characters, and each elementary stroke is assigned to a type of human gestures. Then, a classifier ensemble is applied to identify each gesture so as to recognize the characters that gestured by the human demonstrator. The classier ensemble's size is reduced by feature selection techniques and harmony search algorithm, thereby achieving higher accuracy and smaller ensemble size. The inverse kinematics algorithm converts each stroke's trajectory to the robot's motor values that are executed by a robotic arm to draw the entire character. Experimental analysis shows that the proposed approach can allow a human to naturally and conveniently control the robot in order to write many Chinese characters.
Fei Chao 0001, Zhengshuai Wang, Zuyuan Zhu, Changle Zhou, Qinggang Meng, Min Jiang 0005
FUZZ-IEEE1
2014 Improving machine vision via incorporating expectation-maximization into Deep Spatio-Temporal learning
abstract
The Deep Spatio-Temporal Inference Network (DeSTIN) is a deep learning architecture which combines un-supervised learning and Bayesian inference. The original version of DeSTIN incorporates k-means clustering inside each processing node. Here we propose to replace k-means with a more sophisticated algorithm, online EM (Expectation Maximization), and show that this improves DeSTIN's performance on image classification and restoration tasks.
Min Jiang 0005, Ben Goertzel, Zhongqiang Huang, Changle Zhou, Fei Chao 0001
IJCNN6
2014 A developmental approach to robotic pointing via human-robot interaction
abstract
The ability of pointing is recognised as an essential skill of a robot in its communication and social interaction. This paper introduces a developmental learning approach to robotic pointing, by exploiting the interactions between a human and a robot. The approach is inspired through observing the process of human infant development. It works by first applying a reinforcement learning algorithm to guide the robot to create attempt movements towards a salient object that is out of the robot’s initial reachable space. Through such movements, a human demonstrator is able to understand the robot desires to touch the target and consequently, to assist the robot to eventually reach the object successfully. The human–robot interaction helps establish the understanding of pointing gestures in the perception of both the human and the robot. From this, the robot can collect the successful pointing gestures in an effort to learn how to interact with humans. Developmental constraints are utilised to drive the entire learning procedure. The work is supported by experimental evaluation, demonstrating that the proposed approach can lead the robot to gradually gain the desirable pointing ability. It also allows that the resulting robot system exhibits similar developmental progress and features as with human infants.
Fei Chao 0001, Zhengshuai Wang, Changjing Shang, Qinggang Meng, Min Jiang 0005, Changle Zhou, Qiang Shen 0001
Inf. Sci.1
2014 Feature Selection Inspired Classifier Ensemble Reduction
abstract
Classifier ensembles constitute one of the main research directions in machine learning and data mining. The use of multiple classifiers generally allows better predictive performance than that achievable with a single model. Several approaches exist in the literature that provide means to construct and aggregate such ensembles. However, these ensemble systems contain redundant members that, if removed, may further increase group diversity and produce better results. Smaller ensembles also relax the memory and storage requirements, reducing system's run-time overhead while improving overall efficiency. This paper extends the ideas developed for feature selection problems to support classifier ensemble reduction, by transforming ensemble predictions into training samples, and treating classifiers as features. Also, the global heuristic harmony search is used to select a reduced subset of such artificial features, while attempting to maximize the feature subset evaluation. The resulting technique is systematically evaluated using high dimensional and large sized benchmark datasets, showing a superior classification performance against both original, unreduced ensembles, and randomly formed subsets.
Ren Diao, Fei Chao 0001, Taoxin Peng, Neal Snooke, Qiang Shen 0001
IEEE Trans. Cybern.2
2012 Integration of brain-like computational structure and infant behaviorial pattern for robotic hand-eye coordination
abstract
Robotic hand-eye coordination plays an important role in dealing with real time environment; and the learning procedure of this skill affects the fundamental framework of robotic cognition. This paper introduces a novel developmental approach to hand-eye coordination in an autonomous robotic system. Existing work employs neural network models to map visual perception to hand. In the approach, a computational structure and a cross-modal link mechanism are applied to simulate brain cortices; and a movement pattern inspired by infant behaviors is designed to help robot learn to build its hand-eye coordination. This work is supported by experimental evaluation, which shows that the learning algorithm provides a fast and incremental learning of behavioral competence.
Fei Chao 0001, Haixiong Lin, Min Jiang 0005, Minghui Shi, Jinying Chao
ICARCV1
2012 An algorithm for computing attribute reducts based on graph search strategy
abstract
Attribute reducts can discover previously unknown, non-trivial and useful abstractions from the data in large databases. However, many methods for finding attribute reducts from large data sets always meet a difficult problem of combination explosion. To overcome the problem and find some attribute reducts with high efficiency, the algorithm CARHS was proposed. The basic idea of CARHS is: 1) transform the problem into an equivalent one that searches paths, from which attribute reducts can be easily derived, from a graph; 2) employ high efficient heuristic rules during the course of depth-first search on the graph. By means of the heuristic rules, those paths that would not derive attribute reducts could be blocked as early as possible, furthermore, for those paths that would derive the same attribute reduct, only one of them could complete the course of search, and the others could be blocked as early as possible. Thus some attribute reducts could be found by CARHS with high efficiency even when dealing with huge data sets. The transformation of the problem, novel concepts, the heuristic search rules, and the algorithm CARHS were illustrated in detail by some examples. At last, The experiment on three classic UCI data sets showed the effect of the heuristic search rules and the efficiency of the algorithm CARHS.
Minghui Shi, Changle Zhou, Fei Chao 0001, Min Jiang 0005
IJCNN3