Lin Meng 0001

dblp:67/3925-1 · DBLP profile ↗
← Back
36ranked-venue papers
2as first author
28since 2021 · last 2026
0000-0003-4351-6923ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 11 since 2021Systems, architecture and hardware · 6 · 1 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 6 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 6 since 2021Software engineering, systems software and programming languages · 5 · 4 since 2021Computer networks · 4 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author
YearPublicationVenuePosition
2026 Artificial intelligence on the wing: Fully on-board visual servoing for object tracking with autonomous Nano Unmanned Aerial Vehicles
abstract
Nano Unmanned Aerial Vehicle (UAV) platforms are well-suited for tasks in confined spaces, such as indoor single-person tracking. However, they are constrained by payload, and milliwatt (mW) level computing budgets. To address these limitations, we present a fully on-board tracking system featuring Feather You Only Look Once (FeatherYOLO), an ultra-lightweight detector tailored to the embedded processor, and a custom image-based visual servoing controller deployed on a 29 gram Crazyflie 2.1 nano UAV. FeatherYOLO utilizes a depthwise-separable backbone with a decoupled, anchor-free head, requiring 20 thousand parameters, 1.94 million multiply-accumulate operations, and 224 kilobytes of memory. On our self-collected indoor human detection dataset (six participants across five sites) under a cross-subject and cross-environment held-out protocol, the model achieved 99.5% mean Average Precision (mAP) at an Intersection Over Union threshold of 0.50 and 75.0% mAP over thresholds from 0.50 to 0.95. On-board profiling reveals that pure inference consumes 46.5 mW at 150 Frames Per Second (FPS), accounting for less than 1% of the total flight power, and outperforms a recent nano UAV baseline consuming 225.7 mW at 43 FPS. The proposed visual servoing strategy was refined through flight trials and task-driven tuning, transforming detector outputs into stable, bounded velocity commands with hysteresis and filtering for closed-loop indoor tracking. Real-world flight tests validated the tracking performance with an average tracking success rate of 90.0%, succeeding in 36 out of 40 experimental runs. The primary flight challenges identified include collisions and target loss.
Zhenling Su, Yexin Zhang, Lin Meng 0001
Eng. Appl. Artif. Intell.3
2026 Rethinking attention reliability for token pruning in vision transformers
abstract
Vision Transformers (ViTs) achieve strong performance in visual recognition but incur quadratic computational cost with respect to the number of tokens, motivating extensive research on token pruning and reduction. Most existing pruning methods estimate token importance directly from attention weights, implicitly assuming that attention magnitude provides a reliable proxy for semantic relevance across all layers. Our analysis shows that the validity of this assumption varies substantially with transformer depth and model scale, and can also depend on the training paradigm. Through a systematic analysis of attention selectivity using multiple concentration and stability measures, attention distributions in shallow layers tend to be highly diffuse and weakly discriminative, making attention-based scoring unreliable for early-stage pruning. In contrast, attention becomes increasingly informative in deeper layers as token representations mature. Motivated by this observation, a broad range of non-attention importance scores derived from token embeddings, including statistics- and similarity-based criteria, is examined. Across ViT variants and diverse training settings, these non-attention scores exhibit more stable pruning behavior in shallow layers, whereas attention-based scoring becomes effective only after sufficient representational discrimination is achieved. Importantly, the depth at which this transition occurs is model-dependent and not strictly monotonic, indicating that uniform attention-based pruning is fundamentally mismatched to the representational dynamics of ViTs. Based on these findings, token pruning is formulated as a layer-wise selection problem governed by the reliability of attention, and lightweight static routing configurations are investigated without retraining or dynamic inference control. For equivalent FLOPs, the resulting pruning patterns achieve a trade-off between accuracy and efficiency that is comparable to or superior to that of representative token reduction methods. Overall, these results establish token importance estimation in ViTs as an inherently layer-dependent problem shaped by representation maturity, model characteristics, and training paradigm rather than uniform attention magnitude.
Ryuto Ishibashi, Hayata Kaneko, Lin Meng 0001
Neurocomputing3
2026 SIMD-CP: SIMD with Redundant Bits Compression and Mixed-Precision Packing for Quantized DNNs
abstract
Deploying deep neural networks (DNNs) on edge devices presents notable challenges, including execution time, power consumption, and memory footprint. To address these limitations, the co-design of software-based model compression techniques and dedicated hardware has become crucial for the efficient deployment of DNNs on edge devices. However, the hardware needs to support various model compression techniques, and specific compression formats introduce limitations to the effective use of the conventional SIMD, such as low-bit-width precision, fine-grained mixed precision, and sparse matrices. To overcome these issues, we propose SIMD-CP, a SIMD architecture featuring tag-based precision detection and redundant bit-width compression, which is represented as compression packing. Specifically, we introduce two novel SIMD instructions: (i) a tagged vector load instruction ( tvl ), which fetches quantized vectors from memory while appending bit-width metadata as tags, and (ii) a packing dot-product instruction ( pdotp ), which detects the precision levels of elements and packs them into suitable multipliers. Experimental evaluations show that our approach achieves a 2.0× MAC/cycle gain on both fine-grained mixed-precision and sparse-matrix formats by a series of instructions, i.e., tvl and pdotp . Furthermore, SIMD-CP obtains a 2.70 ∼ 3.40× GOPs/W and a 2.31 ∼ 2.42× OPs/LUT improvement for mixed-precision convolution, outperforming the cutting-edge mixed-precision SIMD. These diverse model compression supports allow 28.8 ∼ 45.5% latency reduction for DNN applications, including tiny CNN and edge-aware Vision Transformer, with mitigating accuracy degradation within 1.2 ∼ 2.1%. We also provide the scaling of the SIMD-CP architecture, resulting in a 1.8% LUT utilization increase in the small-scale compared with the conventional mixed-precision SIMD.
Hayata Kaneko, Ryuto Ishibashi, Lin Meng 0001
ACM Trans. Embed. Comput. Syst.3
2025 SIMD-CP: SIMD with Redundant Bits Compression and Mixed-Precision Packing for Quantized DNNs
abstract
Mixed-precision quantization for deep neural networks is a promising solution for deploying DNNs on edge devices. However, the dedicated hardware must support the execution of mixed-precision while handling complicated software formats, e.g., packed low-bit-width and asymmetric precision. To address these issues, we propose a SIMD architecture featuring tag-based precision detection and redundant bit compression, which we refer to as compression packing. Specifically, we introduce (i) a tagged vector load instruction (tvl) and (ii) a packing dot-product instruction (pdotp), which packs them into a suitable multiplier. Experimental results show that this series of instructions supports low-bit-width quantization, asymmetric mixed precision, and sparse matrices without instruction changes, which achieves a 1.63~2.23× speedup compared with the representative low-bit-width SIMD.
Hayata Kaneko, Lin Meng 0001
CODES+ISSS2
2025 Heterogeneous Resources Adaptive Co-Optimization in Edge Networks
abstract
The heterogeneous resources co-optimization in edge networks is essential to enhance the network throughput. Existing load-sensitive (re-)scheduling approaches mostly formulate the heterogeneous resources balancing as a single-objective optimization issue, omitting the balanced usage of heterogeneous resources on a given edge node. Moreover, these approaches are inadequate for the heterogeneous resources adaptive cooptimization, microservice dependency modeling at a more granular level, and multi-step online re-scheduling. Thus, a Dependency-aware Online Microservice re-Scheduling (DOMS) approach is introduced. In particular, we formulate the microservice re-scheduling as a multiple knapsack optimization issue, and solve it through the Double Dueling Deep Q-Network (D3QN) with prioritized experience replay. Our DOMS incorporates a heterogeneous resources adaptive balancing detection algorithm to enable adaptive co-optimization of heterogeneous resources. A fine-grained dependency graph of microservice performance metrics is built, upon which a multi-step scheduling partition algorithm is devised to facilitate multi-step online re-scheduling. Extensive experiments on a public dataset show that DOMS outperforms comparison approaches in terms of latency, energy consumption, balance degree, and throughput.
Yihong Yang, Zhangbing Zhou, Lin Meng 0001
ICWS3
2025 Enhanced Biogas Production Prediction Using BiogasNET with BiogasGAN
abstract
Biogas is a sustainable energy source produced through anaerobic digestion (AD), which converts organic waste into methane and carbon dioxide. Accurate prediction of biogas yield is essential for stable and efficient operation. However, this task is difficult due to the nonlinear dynamics of AD systems and frequent missing values in sensor data. Here, we propose a twostage framework. First, we introduce BiogasGAN, a generative adversarial network designed to impute missing values in multivariate time series data. It reconstructs incomplete sensor records while preserving temporal and cross-variable relationships. Second, we present BiogasNET, a hybrid deep learning model that combines convolutional layers, long short-term memory (LSTM) units, and attention mechanisms to forecast biogas production from imputed data. We evaluate our framework on real-world biogas plant datasets. Experimental results show that BiogasNET achieves state-of-the-art performance, with RMSE and MAE as low as 0.029 and 0.022, respectively. Ablation studies confirm the value of each model component, and comparisons with conventional machine learning methods highlight its robustness. Overall, our approach provides an effective and practical solution for biogas yield prediction in real-world environments.
Yingrui Geng, Zenghui Wang 0001, Mingcong Deng, Lin Meng 0001
SMC4
2025 Multi-Scale Token Pruning in Mask2Former for Semantic Segmentation
abstract
Although Transformer has been successfully applied to Vision tasks in various fields, its large computational cost and performance degradation due to divergence from the language task are issues to be addressed. In this paper, we introduce token pruning to Mask2Former, a state-of-the-art segmentation method, to reduce computational cost and improve recognition accuracy without additional training. Multi-Scale Token Pruning (MSTP) works effectively on the multi-scale feature tokens of Mask2Former and can be universally implemented with various conventional token pruning methods. Experimental results show that introducing Top-K (norm+rand) MSTP into the Mask2Former of Swin-L backbone achieves +0.28 (56.31) mIoU on the ADE20K benchmark with +5.7% speed up. With this improvement, Mask2Former+MSTP can achieve mIoU equivalent to the large and powerful BEiT-UperNet with 1/4 of the computational complexity. In addition, +0.04 (57.86) PQ for COCO panoptic and +0.19 (63.36) mIoU for Mapillary Vistas are achieved, showing particular usefulness in complex semantic tasks with a large number of categories.
Ryuto Ishibashi, Lin Meng 0001, Mingcong Deng
SMC2
2025 Adaptive spiking neuron with population coding for a residual spiking neural network
Yongping Dan, Changhao Sun, Lin Meng 0001
Appl. Intell.4
2025 YOLO-MSD: a robust industrial surface defect detection model via multi-scale feature fusion
abstract
Abstract Object detection is vital for automated surface defect inspection, yet most models suffer from bloated architectures and poor performance on multi‑class, multi‑scale tasks involving large‑size images, limiting their use on edge devices. We propose YOLO‑MSD, a lightweight surface defect detection model that integrates two key designs: (1) a novel four-scale backbone that effectively extracts small and multi-scale targets from large-size images by enhancing feature representation across different scale resolutions, and (2) a streamlined feature‑pyramid neck that boosts cross‑scale fusion while reducing parameters and computational cost. Extensive experiments on five public datasets verify the model’s effectiveness. On the PCB, HRIPCB and GC10‑DET datasets featuring high-resolution images, YOLO‑MSD achieves 96.67% mAP , 96.62% mAP and 69.09% mAP , respectively, while maintaining a low parameter count and computational complexity. It also outperforms most advanced models on two additional public datasets and achieves 20.82 FPS with a power consumption of 6.95 W on the PCB dataset when deployed on a Jetson Xavier NX edge device. These results demonstrate the accuracy, efficiency, and deployability of YOLO‑MSD for industrial surface‑defect detection.
Yifei Ge, Zhuo Li 0016, Lin Meng 0001
Appl. Intell.3
2025 Automatic pruning rate adjustment for dynamic token reduction in vision transformer
abstract
Abstract Vision Transformer (ViT) has demonstrated excellent accuracy in image recognition and has been actively studied in various fields. However, ViT requires a large matrix multiplication called Attention, which is computationally expensive. Since the computational cost of Self-Attention used in ViT increases quadratically with the number of tokens, research to reduce the computational cost by pruning the number of tokens has been active in recent years. To prune tokens, it is necessary to set the pruning rate, and in many studies, the pruning rate is set manually. However, it is difficult to manually determine the optimal pruning rate because the appropriate pruning rate varies from task to task. In this study, we propose a method to solve this problem. The proposed pruning rate adjustment adjusts the pruning rate so that the training loss is converged by Gradient-Aware Scaling (GAS). In addition, we propose Variable Proportional Attention (VPA) for Top-K, a general-purpose token pruning method, to mitigate the performance loss due to pruning. For the CIFAR-10 dataset, several competitive pruning methods improve recognition accuracy over manually setting the pruning rate; eTPS+Adjust on Hybrid ViT-S achieves 99.01% Accuracy with -31.68% FLOPs. Furthermore, Top-K+VPA outperforms token merging when the pruning rate is large for trained ViT-L inference on ImageNet-1k and has superior scalability in the Accuracy-Latency relation. In particular, when Top-K+VPA is applied to ViT-L on a GPU environment with a pruning rate of 6%, it achieves 80.62% Accuracy on the ImageNet-1k dataset with -50.44% FLOPs and -46.8% Latency.
Ryuto Ishibashi, Lin Meng 0001
Appl. Intell.2
2025 A cosine similarity-based token subsampling method for vision transformer in cloud computing
abstract
Abstract Deploying huge deep learning applications on resource-constrained edge devices is a challenging task. Cloud-based edge computing is a promising solution. Such as model partitioning, a portion of the deep learning model is deployed on the edge device; while, the remaining portion is executed by the cloud. Leveraging the computation power of edge devices, transmission latency is reduced, and bandwidth efficiency is increased. Recently, visual transformer models, supported by large datasets, have dominated in multiple vision tasks. However, model partitioning optimization methods for visual transformers are lacking. Therefore, the paper proposes a cosine similarity-based token subsampling method for visual transformer model partitioning to improve transmission efficiency. Tokens in the same class are subsampled and only the centroid tokens are uploaded. In the cloud, all tokens are reconstructed based on interpolation indexes. Three algorithm implementations are proposed and measured on PC, Jetson NANO and edge CPU Cortex-A53. The experimental results demonstrate that the recommended algorithm implementation can be executed with low-latency of 71.24 ms, and 35.65% transmitted data is reduced with an accuracy drop of 0.46%.
Qi Li 0072, Hayata Kaneko, Lin Meng 0001
Neural Comput. Appl.3
2025 Wireless Capsule Endoscopy Diagnosis Using Prototype Self-Attention and Dynamic Curriculum Learning
abstract
Wireless capsule endoscopy is a non-invasive and painless approach for diagnosing gastrointestinal diseases. An automated medical decision support system can significantly enhance clinician efficiency and reduce the incidence of mis-diagnosis when analyzing lesions within the 50,000 to 100,000 image frames generated for each individual. However, only a small number of previous studies have focused on fine-grained recognition characteristics, and very few have simultaneously addressed both fine-grained recognition and class imbalance issues in wireless capsule endoscopy. To address these issues, we propose a novel medical decision support network that includes a prototype attention enhancement module and a dynamic curriculum learning approach. The prototype attention enhancement module improves lesion-sensitive compact representation learning by leveraging cosine similarity between the representation and a prototype memory with multi-class centers. The dynamic curriculum learning adopts triplet loss and weighted cross-entropy, facilitated by a progressive factor that controls between fine-grained representation learning and class-sensitive learning. Through extensive comparative experiments on three public datasets, the proposed network demonstrated competitive performance, outperforming previous methods with an F1-score of 96.6% on the 10-class Kvasir-Capsule dataset, an accuracy of 98.3% on the CAD-CAP dataset, and an accuracy of 93.7% on the mixed KID dataset. Two additional gastrointestinal histopathology datasets confirm the generalization of the proposed network. The code will be available at https://github.com/Xingcun-Li/MDSN-WCE.
Xingcun Li, Qinghua Wu 0002, Yuning Chen, Lin Meng 0001
IEEE Trans Autom. Sci. Eng.5
2025 Dependency-Aware Online Microservice Re-Scheduling for Adaptive Resources Co-Optimization in Edge Networks
abstract
The usage of heterogeneous resources provisioned by edge nodes can be co-optimized through re-scheduling microservices. Current (re-)scheduling approaches typically treat the task of co-optimization as a single-objective optimization problem, which cannot address the issue of imbalanced usage of heterogeneous resources (e.g., CPU, memory, bandwidth) on a single edge node. More importantly, these approaches are inadequate in handling: (i) the adaptive co-optimization of heterogeneous resources, (ii) the fine-grained construction of micro service dependencies, and (iii) multi-step online mi croservice re-scheduling. To address these challenges, this paper proposes a Dependency-aware Online Microservice re-Scheduling (DOMS) approach. DOMS formulates microservice re-scheduling as a multi-knapsack optimization problem and solves it using a Double Dueling Deep Q-Network (D3QN) with prioritized experience replay. Specifically, an adaptive heterogeneous resources balancing detection algorithm is developed, incorporating a dynamic detection threshold mechanism. A fine-grained microservice performance metrics dependency graph is constructed by capturing causal relationships to represent sequential execution dependency. Based on this graph, a microservice multi-step scheduling partition algorithm is devised. Extensive experiments are conducted upon publicly-available datasets, and evaluation results demonstrate that DOMS outperforms the state-of-the-art techniques with improvements of at least 1.85%, 6.45%, 0.56%, and 3.18% in terms of latency, energy consumption, balance degree, and throughput. These results highlight the effectiveness and superiority of DOMS in maintaining a balanced usage of heterogeneous resources and improving network throughput, while satisfying latency and energy consumption constraints.
Yihong Yang, Zhangbing Zhou, Lianyong Qi, Zhensheng Shi, Lin Meng 0001, Xuyun Zhang
IEEE Trans. Serv. Comput.5
2024 Special Session: Estimation and Optimization of DNNs for Embedded Platforms
abstract
Several state of the art estimation and optimization techniques for CNNs and LLMs on embedded devices are summarized. For LLMs an Activation-aware Weight Quantization and on-the-fly dequantization techniques is presented. For CNNs various pruning algorithms and an integrated optimization and implementation flow is discussed. To estimate inference latency of CNNs on specific hardware platforms, three different techniques are reviewed: A mixed analytic-stochastic model, an analytic model based on step-wise linear functions, and a method that uses a detailed architecture description of the hardware.
Axel Jantsch, Song Han 0003, Lin Meng 0001, Oliver Bringmann 0001, Haotian Tang, Shang Yang, Matthias Wess, Martin Lechner
CODES+ISSS3
2024 Energy-Aware Service Migration in End-Edge-Cloud Collaborative Networks
abstract
Empowered by edge computing, resources and computation capabilities provided by edge devices can be encapsulated as containerized services. When burst requests are coming, there may have edge devices which are overloaded, since most requests are spatially and temporally constrained, and edge devices are resource-scarceness and capacity-limited. Overloaded devices should be relieved through optimally migrating one or more activated services to contiguous edge devices. Besides, sensory data gathered by original edge devices should be periodically transmitted to migrated devices for data analysis purpose. To mitigate this issue, this paper proposes an Energy-efficient Online Service Migration (EOSM) mechanism to conduct the migration of multiple services simultaneously. Extensive experimental results show that our EOSM mechanism outperforms the state of arts techniques in mitigating overloaded services in terms of access latency, energy consumption, and request success rate.
Jiangwei Li, Zhangbing Zhou, Deng Zhao, Zhensheng Shi, Lin Meng 0001, Walid Gaaloul
ICWS5
2024 Visual Anomaly Detection with Self-Attention and Separate Memory Bank
abstract
Declining birthrate and aging populations are progressing all over the world. This has led to labor shortage, making visual inspections more challenging in various industries. Recently, visual anomaly detection methods using deep learning have been proposed to solve these problems. However, they are computationally expensive and difficult to infer in real-time, even in a GPU environment. In addition, while they detect structural anomalies (e.g., scratches and stains), logical anomalies (e.g., mis-position and mis-number) cannot be detected. This work proposes an anomaly detection method to detect both structural and logical anomalies with high speed by improving Patch Core. The proposal applies self-attention mechanism for the intermediate layer of the pre-trained Convolutional Neural Networks(CNN) model. Self-attention mechanism enables the model to understand the relationships between image features and detect logical anomalies. In addition, the global and local features are extracted from the intermediate layer of the pretrained CNN model and stored in Separate Memory Bank (SMB). SMB leads to improving AUROC, which represents accuracy, by calculating features for each feature type. It also avoids unnecessary upsampling and reduces the dimensionality, thus improving inference speed. Experiments validate the proposed method and compare previous anomaly detection methods. Experiments evaluate the performance of the proposal for the CAD-SD dataset and MVTec LOCO dataset, which contains structural and logical anomalies. For Co-occurrence dataset, the experimental results show that the proposal achieves 98.5% (improving 2.2%) for AUROC and 16.1 (improving 66.6%) for FPS compared to the state-of-the-art method. Also, the experimental results show that the proposal achieves 82.8% (improving 0.9%) for MVTec-LOCO dataset. Hence, the proposal can contribute to the efficiency and automation of manufacturing, medical, and other fields.
Kosaburo Hattori, Hayata Kaneko, Ryuto Ishibashi, Tomonori Izumi, Lin Meng 0001
SMC5
2024 3D Industrial anomaly detection via dual reconstruction network
abstract
Abstract Currently, 2D anomaly detection has demonstrated outstanding performance. However, 2D images limit the improvement of anomaly detection accuracy without utilizing depth information. Therefore, this paper proposes a Dual Reconstruction viAInpainting Network for 3D industrial anomaly detection (DRAIN). Firstly, we design a 3D reconstruction network using an encoder-decoder-based U-shaped network for processing RGB images and depth images. Subsequently, accurate anomaly segmentation is implemented through a 3D segmentation network. We introduce a lightweight MLP module to enhance segmentation performance to capture long-range dependencies in the reconstructed images. Furthermore, we propose a dual attention-based information entropy fusion module to expedite feature fusion in the inference process, aiming for enhanced deployment in the industry. Extensive experiments demonstrate that DRAIN achieves a 94.3% AUROC on the 3D anomaly detection dataset MVTec 3D-AD, surpassing other research methods. Graphical abstract Overall architecture for 3D industrial anomaly detection via dual reconstruction network
Zhuo Li 0016, Yifei Ge, Xin Wang 0138, Lin Meng 0001
Appl. Intell.4
2024 An unsupervised automatic organization method for Professor Shirakawa's hand-notated documents of oracle bone inscriptions
abstract
Abstract As one of the most influential Chinese cultural researchers in the second half of the twentieth-century, Professor Shirakawa is active in the research field of ancient Chinese characters. He has left behind many valuable research documents, especially his hand-notated oracle bone inscriptions (OBIs) documents. OBIs are one of the world’s oldest characters and were used in the Shang Dynasty about 3600 years ago for divination and recording events. The organization of OBIs is not only helpful in better understanding Prof. Shirakawa’s research and further study of OBIs in general and their importance in ancient Chinese history. This paper proposes an unsupervised automatic organization method to organize Prof. Shirakawa’s OBIs and construct a handwritten OBIs data set for neural network learning. First, a suite of noise reduction is proposed to remove strangely shaped noise to reduce the data loss of OBIs. Secondly, a novel segmentation method based on the supervised classification of OBIs regions is proposed to reduce adverse effects between characters for more accurate OBIs segmentation. Thirdly, a unique unsupervised clustering method is proposed to classify the segmented characters. Finally, all the same characters in the hand-notated OBIs documents are organized together. The evaluation results show that noise reduction has been proposed to remove noises with an accuracy of 97.85%, which contains number information and closed-loop-like edges in the dataset. In addition, the accuracy of supervised classification of OBIs regions based on our model achieves 85.50%, which is higher than eight state-of-the-art deep learning models, and a particular preprocessing method we proposed improves the classification accuracy by nearly 11.50%. The accuracy of OBIs clustering based on supervised classification achieves 74.91%. These results demonstrate the effectiveness of our proposed unsupervised automatic organization of Prof. Shirakawa’s hand-notated OBIs documents. The code and datasets are available at http://www.ihpc.se.ritsumei.ac.jp/obidataset.html .
Xuebin Yue, Ryuto Ishibashi, Hayata Kaneko, Lin Meng 0001
Int. J. Document Anal. Recognit.5
2024 Energy-Efficient Online Service Migration in Edge Networks
abstract
Empowered by edge computing, resources and computation capabilities provided by edge devices can be encapsulated as containerized services, and domain applications can be achieved through service compositions. When burst requests are coming to be satisfied, there may exist edge devices which are overloaded, since requests are mostly spatially and temporally constrained, and edge devices are resource-scarceness and capacity-limited. In this setting, overloaded devices should be relieved through optimally migrating one or more activated services to contiguous edge devices. Besides, sensory data gathered by original edge devices should be periodically transmitted to migrated devices for data analysis purpose. To mitigate this issue, this paper proposes an Energy-efficient Online Service Migration (EOSM) mechanism to conduct the migration of multiple services simultaneously. Specifically, a light service sharing strategy is developed to only transmit the top container layer, and a modified NSGA-II algorithm is adopted to generate one or multiple paths for the container layer and time-series sensory data migration of each migrated service. Extensive experimental results show that our EOSM strategy outperforms the state of arts techniques in mitigating overloading devices in terms of access latency, energy consumption, and request success rate.
Jiangwei Li, Deng Zhao, Zhensheng Shi, Lin Meng 0001, Walid Gaaloul, Zhangbing Zhou
IEEE Internet Things J.4
2024 MCAD: Multi-classification anomaly detection with relational knowledge distillation
abstract
Abstract With the wide application of deep learning in anomaly detection (AD), industrial vision AD has achieved remarkable success. However, current AD usually focuses on anomaly localization and rarely investigates anomaly classification. Furthermore, anomaly classification is currently requested for quality management and anomaly reason analysis. Therefore, it is essential to classify anomalies while improving the accuracy of AD. This paper designs a novel multi-classification AD (MCAD) framework to achieve high-accuracy AD with an anomaly classification function. In detail, the proposal model based on relational knowledge distillation consists of two components. The first one employs a teacher–student AD model, utilizing a relational knowledge distillation approach to transfer the interrelationships of images. The teacher–student critical layer feature activation values are used in the knowledge transfer process to achieve anomaly detection. The second component realizes anomaly multi-classification using the lightweight convolutional neural network. Our proposal has achieved 98.95, 96.04, and 92.94% AUROC AD results on MNIST, FashionMNIST, and CIFAR10 datasets. Meanwhile, we earn 97.58 and 98.10% AUROC for AD and localization in the MVTecAD dataset. The average classification accuracy of anomaly classification has reached 76.37% in fifteen categories of the MVTec-AD dataset. In particular, the classification accuracy of the leather category has gained 95.24%. The results on the MVTec-AD dataset show that MCAD achieves excellent detection, localization, and classification results.
Zhuo Li 0016, Yifei Ge, Xuebin Yue, Lin Meng 0001
Neural Comput. Appl.4
2023 Optimized Vision Transformer for Dementia Diagnosis Using Micro-Doppler Radar
abstract
In the aging society, the number of dementia patients continues to increase, and early detection of dementia is required. However, going to a hospital and being diagnosed by a doctor is burdensome for elderly people. This paper designs a Vision Transformer(ViT)-based gait diagnosis with micro-Doppler radar to diagnose dementia without burden for elderly people. The ViT is optimized by proposed Vertical Rectangle Patching and Adaptive Thresholding, which improve the Attention of Vi'I. The micro-Doppler radar collects the signal of elderly people, and the signal is transformed into two kinds of signal-analyzed images by two signal-analysis methods (Short-term Fourier Transform: STFT, and Continuous Wavelet Transform: CWT) for diagnosing by optimized ViI. Experiments compare eight kinds of CNN models, and current ViTs with the optimized ViT to evaluate the proposal's performance. The experimental results show that STFT is suitable for analyzing micro-Doppler radar signals, and ViT-56x4s+th, which uses Vertical Rectangle Patching and Adaptive Thresholding, achieves high scores such as an accuracy of 88.9%. Another proposed model ViT-224xl allows for faster learning convergence for time series data and improved accuracy without fine-tuning. In summary, the possibility of diagnosing dementia by optimized ViT has been provided. The experimentation of Adaptive thresholding also proves there are some unimportant patches and brings an idea for reducing the unimportant computations to achieve a compact ViT as future work.
Ryuto Ishibashi, Naoto Nojiri, Hayata Kaneko, Kenshi Saho, Lin Meng 0001
SMC5
2023 Lightweight deep neural network from scratch
Xuebin Yue, Chengyan Zhao, Lin Meng 0001
Appl. Intell.4
2023 Hardware-aware approach to deep neural network optimization
Lin Meng 0001
Neurocomputing2
2023 From SOA to VOA: A Shift in Understanding the Operation and Evolution of Service Ecosystem
abstract
With the development of ICT (information and communications technology) and service economy, service ecosystem is emerging in a lot of fields, including E-commerce, O2O(Online To Offline) life service, healthcare service, cloud manufacturing, and so on. As a complex socio-technical system, the evolution of service ecosystem is the joint result of the interaction of the three heterogeneous networks, including social network, service network and value network. Under such circumstances, the traditional SOA (Service Oriented Architecture)-based analysis model is powerless. As a result, how to analyze the laws behind the evolution of service ecosystem is still a serious challenge in the field. This paper proposes a value oriented analysis framework (VOA) of service ecosystem, which can use value as a clue to describe the interaction of the three heterogeneous networks. In addition, a computational experiment system is established to verify the effectiveness of the VOA framework, which stimulates the effect of different intervention strategies on service ecosystem. The result shows that our analysis framework can provide new means and ideas for the analysis of service ecosystem.
Xiao Xue 0001, Deyu Zhou 0001, Fangyi Chen, Xiangning Yu 0001, Zhiyong Feng 0002, Yucong Duan, Lin Meng 0001, Mu Zhang 0013
IEEE Trans. Serv. Comput.7
2022 Deep learning-based elderly gender classification using Doppler radar
Zhichen Wang, Zelin Meng, Kenshi Saho, Kazuki Uemura, Naoto Nojiri, Lin Meng 0001
Pers. Ubiquitous Comput.6
2022 Correction to: Deep learning-based elderly gender classification using Doppler radar
Zhichen Wang, Zelin Meng, Kenshi Saho, Kazuki Uemura, Naoto Nojiri, Lin Meng 0001
Pers. Ubiquitous Comput.6
2021 A Comprehensive Analysis of Low-Impact Computations in Deep Learning Workloads
abstract
Deep Neural Networks (DNNs) have achieved great successes in various machine learning tasks involving a wide range of domains. Though there are multiple hardware platforms available, such as GPUs, CPUs, FPGAs, and etc, CPUs are still preferred choices for machine learning applications, especially in low-power and resource-constrained computation environments such as embedded systems. However, the power and performance efficiency become critical issues in such computation environments when applying DNN techniques. An attractive optimization to DNNs is to remove redundant computations to enhance the execution efficiency. To this end, this paper conducts extensive experiments and analyses on popular state-of-the-art deep learning models. The experimental results include the numbers of instructions, branches, branch prediction misses, cache misses, and etc, during the execution of the models. Besides, we also investigate the performance and sparsity of each layer in the models. Based on the analysis results, this paper also proposes an instruction-level optimization, which achieves the performance improvement ranging from 10.26% to 28.0% for certain convolution layers.
Zhichen Wang, Xuebin Yue, Wenwen Wang 0001, Hiroyuki Tomiyama, Lin Meng 0001
ACM Great Lakes Symposium on VLSI6
2021 Volunteer Assisted Collaborative Offloading and Resource Allocation in Vehicular Edge Computing
abstract
As a promising new paradigm, Vehicular Edge Computing (VEC) can improve the QoS of vehicular applications by computation offloading. However, with more and more computation-intensive vehicular applications, VEC servers face the challenges of limited resources. In this paper, we study how to effectively and economically utilize the idle resources in volunteer vehicles to handle the overloaded tasks in VEC servers. First, we present a model of volunteer assisted vehicular edge computing, in which the cost and utility functions are defined for requesting vehicles and VEC servers, and volunteer vehicles are encouraged to assist the overloaded VEC servers via obtaining rewards from VEC servers. Then, based on Stackelberg game, we analyze the interactions between requesting vehicles and VEC servers, and find the optimal strategies for them. Furthermore, we prove theoretically that the Stackelberg game between requesting vehicles and VEC servers has a unique Stackelberg equilibrium, and propose a fast searching algorithm based on genetic algorithm to find the best pricing strategy for the VEC server. In addition, to maximize the reward of volunteer vehicles, we propose the volunteer task assignment algorithm for optimal mapping between the tasks and volunteer alliances. Finally, the effectiveness of the proposed scheme is demonstrated through a large number of simulations. Compared with other schemes, the proposed scheme can reduce the offloading cost of vehicles and improve the utility of VEC servers.
Lin Meng 0001, Jinsong Wu 0001
IEEE Trans. Intell. Transp. Syst.3
2020 Frame Detection and Text Line Segmentation for Early Japanese Books Understanding
Bing Lyu, Hiroyuki Tomiyama, Lin Meng 0001
ICPRAM3
2020 Dynamic human contact prediction based on naive Bayes algorithm in mobile social networks
abstract
Summary Human contact prediction is a challenging task in mobile social networks. The existing prediction methods are based on the static network structure, and directly applying these static prediction methods to dynamic network prediction is bound to reduce the prediction accuracy. In this paper, we extract some important features to predict human contacts and propose a novel human contact prediction method based on naive Bayes algorithm, which is suitable for dynamic networks. The proposed method takes the ever‐changing structure of mobile social networks into account. First, the past time is partitioned into many periods with equal intervals, and each period has a feature matrix of all node pairs. Then, with the feature matrixes used for classifiers training based on naive Bayes algorithm, we can get a classifier for each time period. At last, the different weights are assigned to the classifiers according to their importance to contact prediction, and all classifiers are weighted combination into the final prediction classifier. The extensive experiments are conducted to verify the effectiveness and superiority of the proposed method, and the results show that the proposed method can improve the prediction accuracy and TP Rate to a large extent. Besides, we find that the size of time interval has a certain impact on the clustering coefficient of mobile social networks, which further affects the prediction accuracy.
Lan Yao, Baoling Wu, Wenjia Li, Lin Meng 0001
Softw. Pract. Exp.5
2019 QoE-Constrained Concurrent Request Optimization Through Collaboration of Edge Servers
abstract
Cloud computing, which is claimed to provide plentiful storage, computational, and other resources, has become a promising platform to support resource-intensive applications. Due to the wide adoption of smart things to support domain applications and considering the delay-sensitivity of certain requests and limited network capacity compared with huge data packets to be transmitted, the quality of experience (QoE) may be hard to be satisfied when requests are solely supported by cloud computing. In this setting, edge computing has become an infrastructure to facilitate request satisfaction at the network edge. This article proposes a mechanism to optimize the collaboration of heterogeneous edge servers with certain QoE constraints. Specifically, concurrent requests, which are usually represented in terms of SQL queries, are rewritten as atomic queries, and these atomic queries are optimally assigned to edge servers through adopting an algorithm inspired by the minimum spanning tree, where QoE factors, including the delay, size of data packets, and number of operators, are considered. Evaluation results indicate that the proposed mechanism can effectively improve the QoE of requests compared with the state-of-the-art's mechanisms.
Yaqiang Zhang, Lin Meng 0001, Xiao Xue 0001, Zhangbing Zhou, Hiroyuki Tomiyama
IEEE Internet Things J.2
2018 A Low-Cost Edge Server Placement Strategy in Wireless Metropolitan Area Networks
abstract
With the emergence of computation-intensive applications such as face recognition, natural language processing, etc., the performance of mobile devices appears to be weak. Mobile cloud computing is extensively used to cover the aforementioned shortage. However, the distance between remote cloud and mobile devices might be too far to be tolerable. Cloudlet computing, fog computing, and mobile edge computing are proposed to bring the computing resources near to the mobile devices to reduce the high delay caused by the long distance. Although there exist tremendous studies focused on the offloading strategies, where and how to place edge servers to minimize the edge server providers' cost is seldom involved. In this paper, we first investigate the problem of minimizing the number of edge servers while ensuring QoS constraints such as access delay, and the Integer Linear Programming (ILP) formulation of this problem is given. Then, we transform it into the minimum dominating set problem of graph theory and propose a greedy algorithm to solve it. The simulation results demonstrate the proposed algorithm is promising.
Yongzheng Ren, Wenjia Li, Lin Meng 0001
ICCCN4
2018 Boundary Region Detection for Continuous Objects in Wireless Sensor Networks
abstract
Industrial Internet of Things has been widely used to facilitate disaster monitoring applications, such as liquid leakage and toxic gas detection. Since disasters are usually harmful to the environment, detecting accurate boundary regions for continuous objects in an energy‐efficient and timely fashion is a long‐standing research challenge. This article proposes a novel mechanism for continuous object boundary region detection in a fog computing environment, where sensing holes may exist in the deployed network region. Leveraging sensory data that have been gathered, interpolation algorithms have been applied to estimate sensory data at certain geographical locations, in order to estimate a more accurate boundary line. To examine whether estimated sensory data reflect that fact, mobile sensors are adopted to traverse these locations for gathering their sensory data, and the boundary region is calibrated accordingly. Experimental evaluation shows that this technique can generate a precise object boundary region with certain time constraints, and the network lifetime can be prolonged significantly.
Yaqiang Zhang, Zhenhua Wang 0005, Lin Meng 0001, Zhangbing Zhou
Wirel. Commun. Mob. Comput.3
2017 Recognition of Oracle Bone Inscriptions by Extracting Line Features on Image Processing
Lin Meng 0001
ICPRAM1
2015 FPGA-based BLOB Detection Using Dual-pipelining (Abstract Only)
abstract
Binary Large OBject (BLOB) detection is utilized in various fields such as car cameras, traffic sign recognition and surveillance systems. Although labeling is an important component in BLOB detection, it is difficult to be parallelized using a look-up table (LUT) in terms of data dependency. Since BLOB detection takes a long time, recognition speed and accuracy need to be improved. This research aims to detect BLOBs as fast as possible by using dual-pipelining image processing on the FPGA. Dual-pipelining is to perform pipeline processing in parallel to the upper and lower portions of an original image after dividing it into two portions. We have to consider the timing of each module around the borderline because of the data dependency in label generation. The image processing consists of Gaussian filtering, binarization, labeling, and BLOB analysis. Generally, labeling uses a LUT to combine multiple numbers for one object into the smallest number of temporary labels. In order to simplify the labeling, the connected components of each BLOB are stored and revised just in the LUT. In our approach, a BLOB can be detected when multiple temporary labels are stored in a same entry of the LUT, thus enabling us to detect BLOBs by dual-pipelining. Although our labeling method does not revise temporary labels into a unified label, BLOBs can be detected and their numbers, areas, and centroids are correctly computed. We compared our approach with a related work, which consists of three steps: identifying the connected pixels in each row, labeling the counted pixels in different rows, computing the area and centroid. Experimental results show that the dual-pipelining system using FPGA can detect BLOBs in 0.06 ms, which is 3.92 times faster than the related work and 1.83 times faster than a single-pipelining system. The dual-pipelining system utilized 1.5% of Registers, 8.4% of LUT, 24.3% of LUT-FF pairs, 91.9% of BRAM in Virtex V. The dual-pipelining system is about twice as large as the single-pipelining system. Our approach can be applied for the other areas such as traffic sign recognition and vehicle detection.
Naoto Nojiri, Lin Meng 0001, Katsuhiro Yamazaki
FPGA2
2014 Pipelining FPPGA-based defect detction in FPDs (abstract only)
abstract
The real-time detection of defects in Flat-Panel Displays (FPDs) is very important during the production stages. This paper describes the manner in which defects induced by bubbles are detected as fast as possible by using 4-stage image processing pipelines with 3-line buffers on a Field-Programmable Gate Array (FPGA). The image processing consists of reading a Time Delay Integration (TDI) image, Laplacian filtering, binarization, and labeling. TDI is applied to the initial image of the FPD to reduce noises induced when taking the FPD images. Laplacian filtering and binarization are used to detect the edges in the image, and labeling is used to number the objects in the image for defect detection. In the 4-stage pipelining, the first stage reads the TDI image from the Block Random Access Memory (BRAM), the second stage implements Laplacian filtering and binarization, the third stage implements labeling, and the final stage revises the labels and writes them into the BRAM. The target pixel and its eight surrounding neighbors are required during Laplacian filtering, and four neighbors are necessary during labeling. Thus, three line registers (3-line buffer) are used as a general pipeline register between two neighboring stages in our system. The pipelining system accesses these 3-line buffers and runs four image processing steps in parallel. Therefore, the system uses four different addresses to access the BRAM and the 3-line buffers. Further, to facilitate performance comparison, we implemented sequential image processing systems with 3-line buffers on FPGA and CPU software. The experiments reveal that Laplacian filtering, binarization, and labeling for FPD defect detection can be executed in less than 1 ms by using four-stage pipelining on an FPGA, which is 3.62 times faster than the sequential system and 158.7 times faster than the CPU software. The pipelining system is 28% larger as compared to the sequential system in terms of the size of the LUTs.
Lin Meng 0001, Keisuke Matsuyama, Naoto Nojiri, Tomonori Izumi, Katsuhiro Yamazaki
FPGA1