EDBT 2026 Demo / reviewers in the wild / expert
Tao Li 0022
dblp:75/4601-22
· DBLP profile ↗
96ranked-venue papers
8as first author
68since 2021 · last 2026
0000-0003-1697-8022ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 26 · 1 first-author · 12 since 2021Applied, interdisciplinary, general and emerging computing · 23 · 1 first-author · 23 since 2021Systems, architecture and hardware · 16 · 1 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 1 first-author · 7 since 2021Computer networks · 9 · 1 first-author · 8 since 2021Security and privacy · 8 · 8 since 2021Databases, data management, data science and information retrieval · 6 · 2 first-author · 1 since 2021Software engineering, systems software and programming languages · 3 · 3 since 2021Theory of computation · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Towards efficient Graph-RAG via structure-aware intermediate representation: Incremental collaborative exploration on knowledge graph
Ye Lu 0004, Tao Li 0022 |
Knowl. Based Syst. | 5 |
| 2026 | MoGL: A mixture of heterogeneous experts for collaborative graph learning
Gonghai Zhou, Jiatai Wang, Zhengxian Lu, Tao Li 0022 |
Neural Networks | 6 |
| 2026 | Be Bayesian by attachments to catch more uncertainty
Bin Pan, Tianyang Shi, Tao Li 0022, Zhenwei Shi 0001 |
Pattern Recognit. | 4 |
| 2026 | Sequence to Location: Protein Subcellular Localization Driven by Deep Pretrained Language ModelabstractProteins serve as the essential executors of cellular activities, and their mislocalization often resulting in diverse diseases. Traditional methods for determining protein subcellular localization are noted for their time-consuming, labor-intensive, and complex nature. To address these challenges, this study introduces SubLoc, a deep learning-based algorithm for predicting protein subcellular localization. The methodology comprises three key steps: Firstly, leveraging the deep pretrained protein language model ProtT5 to derive protein embedding vectors, thereby capturing intricate patterns and biological functionalities of protein sequences. Secondly, constructing a 3D protein structure graph model using amino acid residue contact relationships within the sequence, which is subsequently processed by a graph convolutional network to effectively manage spatial structural information. Lastly, employing a bidirectional gated recurrent unit and multi-head attention mechanism to analyze sequence features, integrating both structural and sequence data for enhanced subcellular localization prediction. Experimental results demonstrate that SubLoc exhibits exceptional performance in localizing proteins across 10 subcellular compartments, outperforming all comparative methods in terms of precision, recall, and MCC average values. Notably, SubLoc achieves particularly notable results in identifying Cytoplasm and Mitochondrion locations. Shidong Wu, Xiankun Zhang, Tao Li 0022 |
IEEE Trans. Comput. Biol. Bioinform. | 5 |
| 2026 | Fuzzing JavaScript JIT Compilers With Optimization Path Feedback
Jiming Wang, Chenggang Wu 0002, Yan Kang 0002, Yuhao Hu, Jikai Ren, Yuanming Lai, Mengyao Xie, Chao Zhang 0008, Tao Li 0022, Zhe Wang 0017 |
IEEE Trans. Dependable Secur. Comput. | 10 |
| 2026 | AEDS: An Affinity-Driven Efficient DRL-Based Task Scheduling Framework for Edge ComputingabstractEdge computing is a promising paradigm that deploys computing resources at the network edge to provide services. Many existing solutions leverage deep reinforcement learning (DRL) to optimize task scheduling, yet they often rely on global scheduling approaches. However, such solutions result in an excessively large decision search space, reducing task scheduling efficiency in complex environments. Additionally, the cold start problem impedes the generation of optimal scheduling strategies. To address these challenges, we propose AEDS, a DRL-based task scheduling framework designed to enhance scheduling efficiency. AEDS optimizes the decision-making process from three aspects: (1)Decision Space Reduction.AEDS incorporates a novel affinity matching mechanism that identifies the most suitable edge cluster based on task characteristics, thereby significantly narrowing the decision search space. (2)Decision Process Optimization.AEDS adopts a hybrid strategy combining offline pre-training and online fine-tuning to address the cold start problem. Offline pre-training with historical task data ensures effective initial scheduling, while online fine-tuning periodically updates the DRL model to enhance long-term adaptability to dynamic system changes. (3)Decision Strategy Calibration.AEDS proposes a task migration solution to adapt to real-time workload variations dynamically. It utilizes triple queues to assess server overload and dynamically calibrates the scheduling strategy through task migration within interconnected clusters. Comprehensive experimental results validate the efficacy of AEDS. Compared with existing frameworks, AEDS reduces task latency by$28.23\%$and enhances task completion rate by$10.28\%$. Furthermore, by effectively narrowing the decision scope, AEDS accelerates the decision-making process by a remarkable$88.06\%$ Zhaolong Jian, Xueshuo Xie, Qiankun Dong, Mulin Li, Tao Li 0022 |
IEEE Trans. Mob. Comput. | 7 |
| 2026 | PRAD++: Toward Robust Periapical Radiograph Analysis Through Dataset and Model AdvancementsabstractWith the growing application of deep learning (DL) in dental image analysis, numerous datasets and models have been proposed. Periapical radiographs (PR), as one of the most common imaging modalities in clinical dentistry, play a critical role in endodontics. However, due to the high cost of manual annotation and interpretation challenges caused by poor projection and imaging artifacts, publicly available high-quality PR datasets remain scarce, severely limiting the development of DL-based PR analysis models that rely on large-scale annotated data. To address this issue, we introduce PRAD++, a large-scale PR analysis dataset annotated by clinical experts, consisting of 10,000 PR images with multi-level annotations, including 9 pixel-level segmentation categories and 17 image-level classification labels. Building upon PRAD++, we propose PRNet++, an end-to-end PR analysis network. The framework leverages the Multi-scale Wavelet Convolution (MWCN) network and the Channel Fusion Attention (CFA) mechanism to effectively model and integrate multi-scale features. In addition, an Expert Prior Injection (EPI) loss is designed to incorporate domain-specific dental knowledge into the learning process, refining classification predictions based on segmentation outputs to ensure accuracy and clinical interpretability. Extensive experiments on the PRAD++ dataset demonstrate that PRNet++ consistently outperforms state-of-the-art (SOTA) methods, achieving an average DSC of 81.25% for segmentation, alongside macro- and micro-averaged PR-AUCs of 66.58% and 79.10% for classification. Significantly surpassing the runner-up, PRNet++ exhibits enhanced robustness and interpretability in clinically challenging categories. Furthermore, comprehensive ablation and visualization analyses validate the efficacy of individual components and the parameter sensitivity of the EPI loss. Zhenhuan Zhou, Peng Wang 0178, Xiaohang Guan, Tao Li 0022 |
IEEE Trans. Medical Imaging | 5 |
| 2025 | Cochain: Architectural Support Mechanism for Blockchain-Based Task Scheduling
Yaozheng Fang, Yibing Jiang, Xueshuo Xie, Zhaolong Jian, Tao Li 0022, Zhiguo Wan, Grace Guiling Wang |
APPT | 5 |
| 2025 | PCVMNet: Landmark-Aware Transformer Network for Pediatric Cervical Vertebral Maturation AssessmentabstractCervical vertebral maturation (CVM) assessment plays a pivotal role in orthodontic diagnosis and determining the optimal timing of treatment, especially for pediatric patients. While deep learning techniques have demonstrated notable success in medical image analysis, CVM staging in the pediatric population remains underexplored. One major limitation is the lack of publicly available datasets specifically designed for pediatric CVM. In this paper, we introduce the PCVM dataset, a benchmark dataset tailored for Pediatric CVM Staging. The PCVM dataset consists of 1800 lateral cephalometric radiographs from real-world clinical cases of patients aged 3-15 years, annotated with expert-labeled CVM stages and 13 anatomical landmarks. To our knowledge, this is the first publicly available dataset dedicated to pediatric CVM assessment. In addition, we propose PCVMNet, a novel architecture designed for automatic CVM staging. It integrates heatmap-guided feature modulation (HGFM) with vertebral landmark-prompting (VLP) blocks to improve staging accuracy. Experimental results show that our method achieves state-of-the-art performance on the benchmark dataset, significantly improving landmark localization performance and classification accuracy over existing baselines. To facilitate further research in pediatric orthodontic treatment, code and dataset will be available at https://github.com/ybupengwang/PCVMNet. Peng Wang 0178, Xiaohang Guan, Anli Wang, Xueshuo Xie, Tao Li 0022 |
BIBM | 6 |
| 2025 | HookMoE: A learnable performance compensation strategy of Mixture-of-Experts for LLM inference accelerationabstractMixture of Experts (MoE) architectures have emerged as a promising paradigm for scaling model capacity through top-k routing mechanisms.Although reducing the number of activated experts inherently enables inference acceleration, this efficiency gain typically comes at the cost of significant performance degradation.To address this trade-off between efficiency and performance, we propose Hook-MoE, a plug-and-play single-layer compensation framework that effectively restores performance using only a small post-training calibration set.Our method strategically inserts a lightweight trainable Hook module immediately preceding selected transformer blocks.Comprehensive evaluations on four popular MoE models, with an average performance degradation of only 2.5% across various benchmarks, our method reduces the number of activated experts by more than 50% and achieves a 1.42× inference speed-up during the prefill stage.Through systematic analysis, we further reveal that the upper layers require fewer active experts, offering actionable insights for refining dynamic expert selection strategies and enhancing the overall efficiency of MoE models.We make our code available at https://github.com/KerwinKai/HookMoE. Longkai Cheng, Along He, Mulin Li, Xueshuo Xie, Tao Li 0022 |
EMNLP | 5 |
| 2025 | A Novel Framework for Data Augmentation on Appearance Changes for Long-Term Person Re-IdentificationabstractThe general Re-ID works including datasets and networks have an assumption that the pedestrians do not change their appearances throughout the research. The lack of diversity about the appearance of the Re-ID datasets significantly limits the performance of models in cross-appearance Re-ID. Therefore, this paper conduct data augmentation in the appearance dimension to the datasets by the generation framework CAG(cross-appearance images generation) to support the studies about cross-appearance Re-ID. To demonstrate the effectiveness of the generation framework, we selected several classical and SOTA Re-ID models and conducted experiments on three commonly used cross-appearance Re-ID datasets, NKUP+IPRCC/DeepChange. The results show that the cross-appearance Re-ID images generated by CAG can help the models to obtain the robust pedestrian feature, and the diversity of appearance is a universal method for cross-appearance person re-identification. Tao Li 0022, Along He, Tehui Huang, Qiankun Dong |
JCC | 2 |
| 2025 | Towards Automated Pediatric Dental Development Staging: A Dataset and Model
Peng Wang 0178, Along He, Anli Wang, Zhenhuan Zhou, Xiaohang Guan, Tao Li 0022 |
MICCAI (13) | 6 |
| 2025 | CustomFair: A Customized Fairness Method for Federated Recommender Systems in Social Internet of ThingsabstractIn the Social Internet of Things (SIoT), edge computing integrates artificial intelligence to learn intricate relationships. The scale and complexity of SIoT cause a data explosion from diverse objects, hindering tailored services to users who own objects. Moreover, conventional edge computing in SIoT depends on centralized data collection, raising concerns about data privacy. To address the above two issues, federated recommender systems (FRSs) present a promising solution. FRSs can provide SIoT services to users and train a shared model while retaining sensitive data locally on objects. However, as FRSs are driven by data, they are inherently susceptible to algorithmic bias, raising substantial fairness concerns that have attracted considerable attention in SIoT. Recent fairness studies predominantly concentrate on a single sensitive attribute for users, thereby overlooking their autonomy. Therefore, we propose CustomFair, a personalized fairness framework that enables users in FRSs to select preferred sensitive attributes and acquire satisfied recommendation services in SIoT scenarios. First, we define customized fairness to ensure group fairness based on users’ sensitive attributes. The server segments users into subgroups in a privacy-preserving manner. Second, CustomFair employs the DynBalance method with a flexible regularization coefficient to improve recommendation performance and utilizes the AdaptEpoch strategy to achieve fairness. Extensive experiments indicate that CustomFair improves recommendation performance by 0.1–42.92 and enhances fairness by reducing disparities of 0.03–5.41 compared to two baselines across three datasets. Chao Li 0023, Zihang Yin, Bin Wang 0062, Tao Li 0022, Xuhua Bao, Wei Wang 0012 |
IEEE Internet Things J. | 8 |
| 2025 | DMRP: Privacy-Preserving Deep Learning Model with Dynamic Masking and Random Permutation
Chongzhen Zhang, Zhiwang Hu, Xiangrui Xu 0001, Bin Wang 0062, Jian Shen 0001, Tao Li 0022, Baigen Cai, Wei Wang 0012 |
J. Inf. Secur. Appl. | 7 |
| 2025 | Trans-SAM: Transfer Segment Anything Model to medical image segmentation with Parameter-Efficient Fine-Tuning
Yanlin Wu, Xiongfeng Yang, Hong Kang, Along He, Tao Li 0022 |
Knowl. Based Syst. | 6 |
| 2025 | AdaptFRCNet: Semi-supervised adaptation of pre-trained model with frequency and region consistency for medical image segmentation
Along He, Yanlin Wu, Tao Li 0022, Huazhu Fu |
Medical Image Anal. | 4 |
| 2025 | DVPT: Dynamic Visual Prompt Tuning of large pre-trained models for medical image analysis
Along He, Yanlin Wu, Tao Li 0022, Huazhu Fu |
Neural Networks | 4 |
| 2025 | ADFQ-ViT: Activation-Distribution-Friendly post-training Quantization for Vision Transformers
Yanfeng Jiang, Xueshuo Xie, Fei Yang 0007, Tao Li 0022 |
Neural Networks | 5 |
| 2025 | On the Performance and Memory Footprint of Distributed Training: An Empirical Study on TransformersabstractABSTRACT Background: Transformer models have emerged as potent solutions to a wide array of multidisciplinary challenges. The deployment of transformer architectures is significantly hindered by their extensive computational and memory requirements, necessitating reliance on advanced efficient distributed training methodologies. Motivation: Prior research has delved into the performance bottlenecks associated with distributed training, aiming to unravel these bottlenecks and suggest optimization directions. However, such analyses often overlook three aspects unique to transformer models: the specialized architecture, the dependency on various distributed strategies, and the requirement to balance computational and memory overhead. Method: This paper aims to bridge this gap by offering a comprehensive examination of the performance bottlenecks inherent in the distributed training of transformer models, leveraging both theoretical analysis and empirical investigation. We propose an analytical framework tailored to these unique aspects of transformers, facilitating a holistic evaluation of model architectures, distributed strategies, and resource consumption. Based on this analytical framework, we conduct a comparative analysis of theoretical performances and further systematically explore how various distributed training strategies fare in real‐world scenarios. Results: Most of the experimental results can be well explained by the analytical outcomes derived from the analytical framework. Notably, our findings suggest an advantage of pipeline parallelism over data parallelism for transformer models. Moreover, we shed light on some unexpected outcomes, such as the potential for increased total memory overhead due to suboptimal model partitioning within pipeline parallelism. Additionally, we underscore the significance of communication block size and waiting time to further enhance performance. Zhengxian Lu, Fangyu Wang, Fei Yang 0007, Tao Li 0022 |
Softw. Pract. Exp. | 5 |
| 2025 | SmartZone: Runtime Support for Secure and Efficient On-Device Inference on ARM TrustZoneabstractOn-device inference is a burgeoning paradigm that performs model inference locally on end devices, allowing private data to remain local. ARM TrustZone as a widely supported trusted execution environment has been applied to provide confidentiality protection for on-device inference. However, with the rise of large-scale models like large language models (LLMs), TrustZone-based on-device inference faces challenges in migration difficulties and inefficient execution. The rudimentary TEE OS on TrustZone lacks both the inference runtime needed for building models and the parallel support necessary to accelerate inference. Moreover, the limited secure memory resources on end devices further constrain the model size and degrade performance. In this paper, we propose SmartZone to provide runtime support for secure and efficient on-device inference on TrustZone. SmartZone consists three main components: (1) a trusted inference-oriented operator set, providing the underlying mechanisms adapted to the TrustZone’s execution mode for trusted inference of DNN models and LLMs. (2) the proactive multi-threading parallel support, which increases the number of CPU cores in the secure state via cross-world thread collaboration to achieve parallelism, and (3) the on-demand secure memory management method, which statically allocates the appropriate secure memory size based on pre-execution resource analysis. We implement a prototype of SmartZone on the Raspberry Pi 3B+ board and evaluate it on four well-known DNN models and llama2 LLM. Extensive experimental results show that SmartZone provides end-to-end protection for on-device inference while maintaining excellent performance. Compared to the origin trusted inference, SmartZone accelerates the inference speed by up to 4.26× and reduces energy consumption by 65.81%. Zhaolong Jian, Qiankun Dong, Longkai Cheng, Xueshuo Xie, Tao Li 0022 |
IEEE Trans. Computers | 6 |
| 2025 | RobustPFL: Robust Personalized Federated LearningabstractConventional federated learning (FL) coordinated by a central server focuses on training a global model and protecting the privacy of clients' training data by storing it locally. However, the statistical heterogeneity hinders the global model from adapting to the non-IID distributions among clients. Moreover, untrusted and unreliable central servers and malicious clients may compromise model integrity and availability, thus degrading the robustness of FL. To address these challenges, we present RobustPFL, a decentralized personalized federated learning (PFL) approach that combines$\alpha$-based Layer-position Normalized Similarity ($\alpha$-LNS) and local collaborative training to improve personalized performance while utilizing a blockchain-based committee mechanism to coordinate the aggregation process, thereby achieving high personalized accuracy and robustness. Extensive experiments show that our RobustPFL approach outperforms multiple algorithms, including Local training, FedAvg, FedReptile, Per-FedAvg, FedBN, and SPFL, on MNIST, CIFAR10, EMNIST, and N-BaIoT datasets in four non-IID settings. We also evaluate RobustPFL's effectiveness against attacks—poisoning attacks and free-riding attacks. Particularly, for three prevalent poisoning attacks (backdoor, label flipping, and model poisoning attacks), we compare non-defensive (FedAvg) and defensive (Krum, trimmed mean, Bulyan, FedBN, FLAME, and FangTrmean) methods with our proposed RobustPFL. The results show that our approach achieves significant defensive effects. Wei Wang 0012, Yufang Wu, Chao Li 0023, Guangquan Xu, Shouling Ji, Tao Li 0022, Meng Shen 0001, Yufei Han 0001 |
IEEE Trans. Dependable Secur. Comput. | 7 |
| 2025 | FedEditor: Efficient and Effective Federated Unlearning in Cooperative Intelligent Transportation SystemsabstractIn cooperative intelligent transportation systems (CITS), federated learning enables vehicles to train a global model without sharing private data. However, the lack of an unlearning mechanism to remove the influence of vehicle-specified data from the global model potentially violates data protection regulations regarding the right to be forgotten. While the existing federated unlearning (FU) methods exhibit promising unlearning effects, their practicality in CITS is hindered due to the time-consuming retraining steps required by other vehicles and the non-negligible performance sacrifice on the un-forgotten data. Therefore, achieving effective unlearning without extensive retraining, while minimizing performance degradation on the un-forgotten data remains a challenge. In this work, we propose FedEditor, an efficient and effective FU framework in CITS that addresses the above challenge by reconfiguring the global model’s representation space to remove critical classification-related knowledge from the unlearned data. Firstly, FedEditor enables vehicles to perform the unlearning process locally on the global model, eliminating the participation of other vehicles and improving efficiency. Secondly, FedEditor captures and aligns the representations of the unlearned data with those of the nearest incorrect class centroid derived from non-training data, ensuring effective unlearning while preserving the un-forgotten data’s knowledge relatively intact for achieving competitive model performance. Finally, FedEditor refines the global model’s output distributions using the vehicles’ remaining data and incorporates a drift-mitigating regularization term, minimizing the negative impact of unlearning operations on model performance. Experimental results show that FedEditor reduces the unlearning rate by up to 99.64% without time-consuming retraining, while limiting the predictive performance loss of the resulting global model to less than 3.88% across five models and seven datasets. Jiqiang Liu, Bin Wang 0051, Xiangrui Xu 0001, Tao Li 0022, Wei Wang 0012 |
IEEE Trans. Inf. Forensics Secur. | 7 |
| 2025 | Transformer for Multitemporal Hyperspectral Image UnmixingabstractMultitemporal hyperspectral image unmixing (MTHU) holds significant importance in monitoring and analyzing the dynamic changes of surface. However, compared to single-temporal unmixing, the multitemporal approach demands comprehensive consideration of information across different phases, rendering it a greater challenge. To address this challenge, we propose the Multitemporal Hyperspectral Image Unmixing Transformer (MUFormer), an end-to-end unsupervised deep learning model. To effectively perform multitemporal hyperspectral image unmixing, we introduce two key modules: the Global Awareness Module (GAM) and the Change Enhancement Module (CEM). The GAM computes self-attention across all phases, facilitating global weight allocation. On the other hand, the CEM dynamically learns local temporal changes by capturing differences between adjacent feature maps. The integration of these modules enables the effective capture of multitemporal semantic information related to endmember and abundance changes, significantly improving the performance of multitemporal hyperspectral image unmixing. We conducted experiments on one real dataset and two synthetic datasets, demonstrating that our model significantly enhances the effect of multitemporal hyperspectral image unmixing. Qiankun Dong, Xueshuo Xie, Tao Li 0022, Zhenwei Shi 0001 |
IEEE Trans. Image Process. | 5 |
| 2025 | CRT and PUF-Based Self/Mutual-Healing Key Distribution Protocol With Collusion Resistance and Revocation CapabilityabstractSelf-healing group key distribution (SGKD) protocols guarantee the security of group communications by allowing authorized users to independently recover missed previous session keys from the current broadcast without retransmission. However, existing SGKD protocols have flaws: (1) collusion resistance and revocable nodes are both upper-bounded by the degree of polynomials used, (2) the disclosure of personal secrets enables the recovery of group key, (3) temporary revocation of a group member is not possible, and (4) a revoked node may obtain the session key when initiating mutual healing, moreover, a malicious node may cause the recovery of false group keys. To address these limitations, we propose an SGKD protocol using the Chinese remainder theorem (CRT) and Physical Unclonable Function (PUF). Our proposed SGKD protocol generates a PUF-based dynamic secret by stimulating nodes’ PUF using a polynomial-based encrypted challenge. This secret is then employed to retrieve a CRT-based encrypted group key. By combining PUF and CRT, we can generate dynamic secrets on the fly and reduce computation time significantly. Utilizing such a technique, our protocol achieves superior security goals, including resistance to any coalition of group nodes even if nodes’ personal secrets were disclosed. Furthermore, the proposed protocol provides an unlimited number of revocable nodes. Additionally, a revoked node can rejoin its group in later sessions without affecting backward secrecy. Moreover, the protocol provides a backward secrecy guaranteed mutual-healing feature free from desynchronization. Our performance and security analyses (i.e., theorem-based formal analysis, NS3-based experiment, and formal verification using the AVIPSA tool) show that our proposed protocol achieves stronger security goals and better efficiency in terms of computation, communication, and storage costs compared to existing SGKD schemes. Wajdy Othman, Hong Zhong 0001, Fuyou Miao 0001, Kaiping Xue, Ammar Hawbani, Liang Zhao 0004, Tao Li 0022 |
IEEE Trans. Mob. Comput. | 8 |
| 2025 | Toward Generalized Multistage Clustering: Multiview Self-DistillationabstractExisting multistage clustering methods independently learn the salient features from multiple views and then perform the clustering task. Particularly, multiview clustering (MVC) has attracted a lot of attention in multiview or multimodal scenarios. MVC aims at exploring common semantics and pseudo-labels from multiple views and clustering in a self-supervised manner. However, limited by noisy data and inadequate feature learning, such a clustering paradigm generates overconfident pseudo-labels that misguide the model to produce inaccurate predictions. Therefore, it is desirable to have a method that can correct this pseudo-label mistraction in multistage clustering to avoid bias accumulation. To alleviate the effect of overconfident pseudo-labels and improve the generalization ability of the model, this article proposes a novel multistage deep MVC framework where multiview self-distillation (DistilMVC) is introduced to distill dark knowledge of label distribution. Specifically, in the feature subspace at different hierarchies, we explore the common semantics of multiple views through contrastive learning and obtain pseudo-labels by maximizing the mutual information between views. Additionally, a teacher network is responsible for distilling pseudo-labels into dark knowledge, supervising the student network and improving its predictive capabilities to enhance its robustness. Extensive experiments on real-world multiview datasets show that our method has better clustering performance than the state-of-the-art (SOTA) methods. Jiatai Wang, Xin Wang 0001, Tao Li 0022 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2025 | Evolutionary Multi-Objective Deep Reinforcement Learning for Task Offloading in Industrial Internet of ThingsabstractMobile Edge Computing (MEC) plays a pivotal role in optimizing the Industrial Internet of Things (IIoT), where the Industrial Task Offloading Problem (ITOP) is crucial for ensuring optimal system performance by balancing conflicting objectives such as delay, energy consumption, and cost. However, existing approaches often oversimplify multi-objective optimization by aggregating conflicting goals into a single objective, while also suffering from limited exploration and robustness in uncertain MEC scenarios within IIoT. To overcome this limitation, we propose EMDRL-ITOP, an Evolutionary Multi-Objective Deep Reinforcement Learning algorithm that synergizes an evolutionary algorithm with deep reinforcement learning (DRL). Firstly, we formulate a multi-objective task scheduling model for IIoT-MEC and design a three-dimensional vector reward function within a Multi-Objective Markov Decision Process framework, enabling simultaneous optimization of delay, energy, and cost. Then, EMDRL-ITOP integrates evolutionary mechanisms to enhance exploration and robustness: a dynamic elite selection strategy prioritizes high-quality policies, a distillation crossover operator fuses advantageous traits from elite strategies, and a proximal mutation mechanism maintains population diversity. These components collectively improve learning efficiency and solution quality in dynamic environments. Extensive simulations across six instances demonstrate that EMDRL-ITOP achieves a superior balance among conflicting objectives compared to state-of-the-art methods, while also outperforming existing algorithms in several key performance metrics. Zhengyi Chai 0001, Yan-Yang Cheng, Yalun Li 0001, Tao Li 0022 |
IEEE Trans. Netw. Serv. Manag. | 5 |
| 2024 | Spatial-Frequency Dual Domain Attention Network For Medical Image SegmentationabstractIn medical images, various types of lesions often manifest significant differences in their shape and texture. Accurate medical image segmentation demands deep learning models with robust capabilities in multi-scale and boundary feature learning. However, previous models still have limitations in addressing the above issues. The majority of medical image segmentation networks exclusively learn features in the spatial domain, disregarding the abundant global information in the frequency domain. This results in a bias towards low-frequency components, neglecting crucial high-frequency information. To address these problems, we introduce SF-UNet, a spatial-frequency dual-domain attention network. It comprises two main components: the Multi-scale Progressive Channel Attention (MPCA) block, which progressively extract multi-scale features across adjacent encoder layers, and the lightweight Frequency-Spatial Attention (FSA) block, with only 0.05M parameters, enabling concurrent learning of texture and boundary features from both spatial and frequency domains. We validate the effectiveness of the proposed SF-UNet on three public datasets. Experimental results show that compared to previous state-of-the-art medical image segmentation networks, SF-UNet achieves the best performance, and achieves up to 9.4% and 10.78% improvement in DSC and IOU. Codes will be released at https://github.com/nkicsl/SF-UNet. Zhenhuan Zhou, Along He, Yanlin Wu, Rui Yao 0010, Xueshuo Xie, Tao Li 0022 |
BIBM | 6 |
| 2024 | Deep Feature Surgery: Towards Accurate and Efficient Multi-exit Networks
Yao Chen 0008, Qiuyang Luo, Ye Lu 0004, Tao Li 0022 |
ECCV (49) | 5 |
| 2024 | TMU: Transmission-Enhanced Mamba-UNet for Medical Image Segmentation
Xiongfeng Yang, Yanlin Wu, Xueshuo Xie, Li Nan, Tao Li 0022 |
ICIC (10) | 6 |
| 2024 | MEFold: Memory-Efficient Optimization for Protein Language Models via Chunk and QuantizationabstractProtein language models are currently experiencing a surge in demand owing to their remarkable accuracy in protein structure prediction. Nevertheless, their applications are hindered by the significant computation and memory requirements. The existing optimization strategies primarily focus on computational efficiency while often neglecting memory optimization, thereby restricting their suitability for devices with limited resources. In this paper, we propose MEFold, a novel memory-efficient optimization framework for protein language models that enables efficient inference on resource-constrained devices. MEFold consists of Look-up Table Chunk and Fine-grained Quantization. Look-up Table Chunk reduces the memory of intermediate activations by chunk and avoids the overhead of obtaining the optimal chunk size configuration through pre-computing. For the memory of model parameters, Fine-grained Quantization, delicately controls the scope of quantization to ensure that memory reduction is achieved while preventing declines in accuracy and computational speed. Experimental results show that, compared to the original model, for protein sequences ranging from 74 to 1024 in length, our method significantly reduces the peak memory during inference from 14.7-54.2GB to 6.0-14.4GB, while minimizing the impact on inference latency. On CASP14 and CAMEO datasets, the accuracy loss compared to the original model is below 1%. Moreover, our optimization provides various memory-saving alternatives. Our code is available at https://github.com/llwx593/MEFold. Yanfeng Jiang, Zhengxian Lu, Fei Yang 0007, Tao Li 0022 |
IJCNN | 7 |
| 2024 | OAA: An Abstraction for Efficient Accelerator Adaptation in Deep Learning FrameworksabstractDeep learning frameworks rely on specific runtime and computation libraries to rewrite the backend for the adaptation of specialized accelerators, which is inefficient and hard to guarantee performance. This paper solves the issue by proposing an Operator Adaptation Abstraction (OAA) that lies between the framework and the libraries. In addition, this paper optimizes training performance based on the hardware characteristics of the accelerator. We designed experiments based on Ascend 910 and OneFlow to verify the effectiveness of OAA. Our adaptation includes more than 50 Ascend operators, achieving up to 2.0x throughput compared to the official adaptation version of PyTorch. The experiments validate that the adaptation method proposed in this paper can effectively retain the advantages of both the framework and the accelerator. Zhengxian Lu, Chengkun Du, Xueshuo Xie, Qiankun Dong, Tao Li 0022 |
IJCNN | 6 |
| 2024 | Memory-Efficient and Secure DNN Inference on TrustZone-enabled Consumer IoT DevicesabstractEdge intelligence enables resource-demanding Deep Neural Network (DNN) inference without transferring original data, addressing concerns about data privacy in consumer Inter-net of Things (IoT) devices. For privacy-sensitive applications, deploying models in hardware-isolated trusted execution environments (TEEs) becomes essential. However, the limited secure memory in TEEs poses challenges for deploying DNN inference, and alternative techniques like model partitioning and offloading introduce performance degradation and security issues. In this paper, we present a novel approach for advanced model deployment in TrustZone that ensures comprehensive privacy preservation during model inference. We design a memory-efficient management method to support memory-demanding inference in TEEs. By adjusting the memory priority, we effectively mitigate memory leakage risks and memory overlap conflicts, resulting in 32 lines of code alterations in the trusted operating system. Additionally, we leverage two tiny libraries: S-Tinylib (2,538 LoCs), a tiny deep learning library, and Tinylibm (827 LoCs), a tiny math library, to support efficient inference in TEEs. We implemented a prototype on Raspberry Pi 3B+ and evaluated it using three well-known lightweight DNN models. The experimental results demonstrate that our design significantly improves inference speed by 3.13 times and reduces power consumption by over 66.5% compared to non-memory optimization method in TEEs. Xueshuo Xie, Haoxu Wang, Zhaolong Jian, Tao Li 0022, Wei Wang 0012, Grace Guiling Wang |
INFOCOM | 4 |
| 2024 | FRCNet: Frequency and Region Consistency for Semi-supervised Medical Image Segmentation
Along He, Tao Li 0022, Yanlin Wu, Ke Zou, Huazhu Fu |
MICCAI (8) | 2 |
| 2024 | Open-Set Semi-supervised Medical Image Classification with Learnable Prototypes and Outlier Filter
Along He, Tao Li 0022, Yitian Zhao, Junyong Zhao, Huazhu Fu |
MICCAI (11) | 2 |
| 2024 | OptFuzz: Optimization Path Guided Fuzzing for JavaScript JIT Compilers
Jiming Wang, Yan Kang 0002, Chenggang Wu 0002, Yuhao Hu, Jikai Ren, Yuanming Lai, Mengyao Xie, Tao Li 0022, Zhe Wang 0017 |
USENIX Security Symposium | 10 |
| 2024 | Quantitative evaluation of deep learning frameworks in heterogeneous computing environment
Zhengxian Lu, Chengkun Du, Yanfeng Jiang, Xueshuo Xie, Tao Li 0022, Fei Yang 0007 |
CCF Trans. High Perform. Comput. | 5 |
| 2024 | AutoQNN: An End-to-End Framework for Automatically Quantizing Neural Networks
Ye Lu 0004, Surong Dai, Deng Qian, Chengkun Du, Tao Li 0022 |
J. Comput. Sci. Technol. | 6 |
| 2024 | Joint Variational Inference Network for domain generalization
Jun-Zheng Chu, Bin Pan, Tianyang Shi, Zhenwei Shi 0001, Tao Li 0022 |
Pattern Recognit. | 6 |
| 2024 | DRS: A deep reinforcement learning enhanced Kubernetes scheduler for microservice-based systemabstractSummary Recently, Kubernetes is widely used to manage and schedule the resources of microservices in cloud‐native distributed applications, as the most famous container orchestration framework. However, Kubernetes preferentially schedules microservices to nodes with rich and balanced CPU and memory resources on a single node. The native scheduler of Kubernetes, called Kube‐scheduler, may cause resource fragmentation and decrease resource utilization. In this paper, we propose a deep reinforcement learning enhanced Kubernetes scheduler named DRS. We initially frame the Kubernetes scheduling problem as a Markov decision process with intricately designed state , action , and reward structures in an effort to increase resource usage and decrease load imbalance. Then, we design and implement DRS mointor to perceive six parameters concerning resource utilization and create a thorough picture of all available resources globally. Finally, DRS can automatically learn the scheduling policy through interaction with the Kubernetes cluster, without relying on expert knowledge about workload and cluster status. We implement a prototype of DRS in a Kubernetes cluster with five nodes and evaluate its performance. Experimental results highlight that DRS overcomes the shortcomings of Kube‐scheduler and achieves the expected scheduling target with three workloads. With only 3.27% CPU overhead and 0.648% communication delay, DRS outperforms Kube‐scheduler by 27.29% in terms of resource utilization and reduces load imbalance by 2.90 times on average. Zhaolong Jian, Xueshuo Xie, Yaozheng Fang, Yibing Jiang, Ye Lu 0004, Ankan Dash, Tao Li 0022, Grace Guiling Wang |
Softw. Pract. Exp. | 7 |
| 2024 | Eyes on Federated Recommendation: Targeted Poisoning With Competition and Its MitigationabstractFederated recommendation (FR) addresses privacy concerns in recommender systems by training a global model without requiring raw user data to leave individual devices. A server, known as the aggregator, integrates users’ local gradients and updates the global model parameters. However, FR is vulnerable to attacks where malicious users manipulate these updates, known as model poisoning attacks. In this work, we propose a new targeted attack calledStairClimbingto promote specific items through model poisoning, and a new defence mechanismCrossEU. StairClimbingadopts a new strategy resembling stair climbing to enable target items to beat competitive items and increase their popularity level by level. Compared to prior attacks,StairClimbingguarantees balanced effectiveness, efficiency and stealthiness simultaneously. Our defence mechanismCrossEUleverages two patterns regarding the lists of items updated by benign users between iterative epochs. Extensive experiments on six real-world datasets demonstrateStairClimbing’s superiority across all three desirable attack properties, even with a small proportion of malicious users (1%). In addition,CrossEUeffectively delays the impact of all tested attacks and even eliminates their damage entirely. Yurong Hao, Xihui Chen, Wei Wang 0012, Jiqiang Liu, Tao Li 0022, Witold Pedrycz |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2024 | FedComm: A Privacy-Enhanced and Efficient Authentication Protocol for Federated Learning in Vehicular Ad-Hoc NetworksabstractIn vehicular ad-hoc networks (VANET), federated learning enables vehicles to collaboratively train a global model for intelligent transportation without sharing their local data. However, due to dynamic network structure and unreliable wireless communication of VANET, various potential risks (e.g., identity privacy leakage, data privacy inference, model integrity compromise, and data manipulation) undermine the trustworthiness of intermediate model parameters necessary for building the global model. While existing cryptography techniques and differential privacy provide provable security paradigms, the practicality of secure federated learning in VANET is hindered in terms of training efficiency and model performance. Therefore, developing a secure and efficient federated learning in VANET remains a challenge. In this work, we propose a privacy-enhanced and efficient authentication protocol for federated learning in VANET, called FedComm. Unlike existing solutions, FedComm addresses the above challenge through user anonymity. First, FedComm enables vehicles to participate in training with unlinkable pseudonyms, ensuring both privacy preservation and efficient collaboration. Second, FedComm incorporates an efficient authentication protocol to guarantee the authenticity and integrity of model parameters originated from anonymous vehicles. Finally, FedComm accurately identifies and completely eliminates malicious vehicles in anonymous communication. Security analysis and verification with ProVerif demonstrate that FedComm enhances privacy and reliability of intermediate model parameters. Experimental results show that FedComm reduces the overhead of proof generation and verification by 67.38% and 67.39%, respectively, compared with the state-of-the-art authentication protocols used in federated learning. Jiqiang Liu, Bin Wang 0062, Wei Wang 0012, Bin Wang 0066, Tao Li 0022, Xiaobo Ma 0001, Witold Pedrycz |
IEEE Trans. Inf. Forensics Secur. | 6 |
| 2024 | NKUT: Dataset and Benchmark for Pediatric Mandibular Wisdom Teeth SegmentationabstractGermectomy is a common surgery in pediatric dentistry to prevent the potential dangers caused by impacted mandibular wisdom teeth. Segmentation of mandibular wisdom teeth is a crucial step in surgery planning. However, manually segmenting teeth and bones from 3D volumes is time-consuming and may cause delays in treatment. Deep learning based medical image segmentation methods have demonstrated the potential to reduce the burden of manual annotations, but they still require a lot of well-annotated data for training. In this paper, we initially curated a Cone Beam Computed Tomography (CBCT) dataset, NKUT, for the segmentation of pediatric mandibular wisdom teeth. This marks the first publicly available dataset in this domain. Second, we propose a semantic separation scale-specific feature fusion network named WTNet, which introduces two branches to address the teeth and bones segmentation tasks. In WTNet, We design a Input Enhancement (IE) block and a Teeth-Bones Feature Separation (TBFS) block to solve the feature confusions and semantic-blur problems in our task. Experimental results suggest that WTNet performs better on NKUT compared to previous state-of-the-art segmentation methods (such as TransUnet), with a maximum DSC lead of nearly 16%. Zhenhuan Zhou, Along He, Xitao Que, Kai Wang 0001, Rui Yao 0010, Tao Li 0022 |
IEEE J. Biomed. Health Informatics | 7 |
| 2024 | Bilateral Supervision Network for Semi-Supervised Medical Image SegmentationabstractMassive high-quality annotated data is required by fully-supervised learning, which is difficult to obtain for image segmentation since the pixel-level annotation is expensive, especially for medical image segmentation tasks that need domain knowledge. As an alternative solution, semi-supervised learning (SSL) can effectively alleviate the dependence on the annotated samples by leveraging abundant unlabeled samples. Among the SSL methods, mean-teacher (MT) is the most popular one. However, in MT, teacher model's weights are completely determined by student model's weights, which will lead to the training bottleneck at the late training stages. Besides, only pixel-wise consistency is applied for unlabeled data, which ignores the category information and is susceptible to noise. In this paper, we propose a bilateral supervision network with bilateral exponential moving average (bilateral-EMA), named BSNet to overcome these issues. On the one hand, both the student and teacher models are trained on labeled data, and then their weights are updated with the bilateral-EMA, and thus the two models can learn from each other. On the other hand, pseudo labels are used to perform bilateral supervision for unlabeled data. Moreover, for enhancing the supervision, we adopt adversarial learning to enforce the network generate more reliable pseudo labels for unlabeled data. We conduct extensive experiments on three datasets to evaluate the proposed BSNet, and results show that BSNet can improve the semi-supervised segmentation performance by a large margin and surpass other state-of-the-art SSL methods. Along He, Tao Li 0022, Juncheng Yan, Kai Wang 0001, Huazhu Fu |
IEEE Trans. Medical Imaging | 2 |
| 2023 | Dancing With Wolves: An Intra-Process Isolation Technique With Privileged HardwareabstractIntra-process memory isolation is a cornerstone technique of protecting the sensitive data in memory-corruption defenses, such as the shadow stack in control flow integrity (CFI) and the safe region in code pointer integrity (CPI). In this article, we proposeSEIMI, a highly efficient intra-process memory isolation technique for memory-corruption defenses. The core is to use the efficientSupervisor-mode Access Prevention (SMAP), a hardware feature that is originally used for preventing the kernel from accessing the user space, to achieve intra-process memory isolation. To leverage SMAP,SEIMIcreatively executes the user code in the privileged mode. In addition to enabling the new design of the SMAP-based memory isolation, we further develop multiple new techniques to ensure secure escalation of user code. Extensive experiments show thatSEIMIoutperforms existing isolation mechanisms, including theMemory Protection Keys(MPK) based scheme and theMemory Protection Extensions(MPX) based scheme. Chenggang Wu 0002, Mengyao Xie, Zhe Wang 0017, Yinqian Zhang, Kangjie Lu, Yuanming Lai, Yan Kang 0002, Min Yang 0002, Tao Li 0022 |
IEEE Trans. Dependable Secur. Comput. | 10 |
| 2023 | UnDAT: Double-Aware Transformer for Hyperspectral UnmixingabstractDeep-learning-based methods have attracted increasing attention on hyperspectral unmixing, where the transformer models have shown promising performance. However, recently proposed deep-learning-based hyperspectral unmixing methods usually tend to directly apply visual models, while ignoring the characteristics of hyperspectral imagery. In this article, we propose a novel double-aware transformer for hyperspectral Unmixing (UnDAT), which aims at simultaneously exploiting the region homogeneity and spectral correlation of hyperspectral imagery. One of the major assumptions of UnDAT is that hyperspectral remote-sensing images involve many homogeneous regions. Pixels inside a homogeneous region usually present similar spectral features, and the edge pixels are just the reverse. Another observation is that the pixel spectra are continuous and correlated. Based on the above assumption and observation, we construct the UnDAT by developing two modules: Score-based homogeneous-aware (SHA) module and the spectral group-aware (SGA) module. In the SHA module, a feature map rearrangement (FMR) approach is proposed to split the shallow feature maps from a linear encoder into an ordered homogeneous map (HomoMap) and an edge map and develop a homogenous region-aware strategy for deep feature representation. In the SGA module, the dependency among neighboring bands is described by dividing the hyperspectral image into multiple spectral groups and calculating the spectral similarity among bands within each group. Experiments on both real and synthetic datasets indicate the effectiveness of our model. We will publish the code of our approach if the article has the honor to be accepted. Yuexin Duan, Tao Li 0022, Bin Pan, Zhenwei Shi 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2023 | LiCa: Label-Indicate-Conditional-Alignment Domain Generalization for Pixel-Wise Hyperspectral Imagery ClassificationabstractOne of the major difficulties for hyperspectral imagery (HSI) classification is the hyperspectral-heterospectra, which refers to the same material presenting different spectra. Although joint spatial-spectral classification methods can relieve this problem, they may lead to falsely high accuracy because the test samples may be involved during the training process. How to address the hyperspectral-heterospectra problem remains a great challenge for pixel-wise hyperspectral imagery classification methods. Domain generalization is a promising technique that may contribute to the heterospectra problem, where the different spectra of the same material can be considered as several domains. In this paper, inspired by the theory of domain generalization, we provide a formulaic expression for hyperspectral-heterospectra. To be specific, we consider the spectra of one material as a conditional distribution and propose a domain-generalization-based method for pixel-wise HSI classification. The key of our proposed method is a new Label-indicate-Conditional-alignment (LiCa) block that focuses on aligning the spectral conditional distributions of different domains. In the LiCa block, we define two loss functions, cross-domain conditional alignment, and cross-domain entropy, to describe the heterogeneity of HSI. Moreover, we have provided the theoretical foundation for the newly-proposed loss functions, by analyzing the upper bound of classification error in any target domains. Experiments on several public data sets indicate that the LiCa block has achieved better generalization performance when compared with other pixel-wise classification methods. Bin Pan, Tao Li 0022, Zhenwei Shi 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2023 | Toward Convergence: A Gradient-Based Multiobjective Method With Greedy Hash for Hyperspectral UnmixingabstractMultiobjective optimization aims at addressing the conflicting objectives, which has been introduced to improve the performance of sparse hyperspectral unmixing. Recently proposed multiobjective unmixing methods usually employ evolutionary algorithms to improve the unmixing accuracy. However, evolutionary algorithms may suffer the challenge of convergence, in which case the reasonability of the solutions is hard to guarantee. To solve the problem of convergence, in this paper, we present a new gradient-based multiobjective unmixing method, which explores the optimization direction in a theoretically reliable manner. Furthermore, considering the mathematical model of hyperspectral sparse unmixing where sparsity error objective of selected endmembers is discrete, we develop a greedy hash based coding approach which is able to well describe the discrete constraints imposed on endmembers. The major components of the proposed method are a search approach and an update approach. In the search approach, we construct the pareto descent direction via a gradient-based strategy, which contributes to converging to an optimal continuous solution by searching along this direction. In the update approach, we update discrete binary endmember via hash coding under the guidance of greedy principle, which allows our method to handle the problem of discrete objective. The major contribution of the proposed method is designing a new framework that can get the optimal discrete endmembers in a convergent way. Moreover, we provide the theoretical analysis and proof for the convergence. Synthetic and real-world experiments have indicated the advantages of our algorithm when compared with evolutionary multiobjective unmixing methods. Ruiying Li, Bin Pan, Tao Li 0022, Zhenwei Shi 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2023 | An Imbalanced Discriminant Alignment Approach for Domain Adaptive SAR Ship DetectionabstractSynthetic aperture radar (SAR) imaging has round-the-clock data acquisition capability regardless of light and climate constraints, so it has been widely used for ship detection. However, SAR images usually suffer lower imaging quality, which may result in indistinct contours and non-negligible noise. Therefore, the manual labeling for SAR images is expensive, leading to a lack of training data in the task of ship detection. In this paper, we propose a route by utilizing domain adaptive methods to transfer information from labeled visible images (source domain) to unlabeled SAR images (target domain) for ship detection. To address the distribution mismatch between domains, we develop a novel imbalanced discriminant alignment (IDA) approach to improve the discriminant ability of the network and prevent negative migration. The core of the IDA approach is applying a new loss function called imbalanced prediction consistency (IPC) loss to describe the domain classifier consistency, and we further provide theoretical analysis for the effectiveness of the IPC loss. IDA ensures consistency at the image level and instance level, and focuses on the consistency of the source domain to enhance the feature extraction capability of the adversarial network. The theoretical discussion has proven that a necessary and sufficient condition for convergence of the IPC loss is that the two discriminant probabilities converge to 0 at the discriminant distance we define. Experimental results have indicated the advantage of IDA when compared with other domain adaptation SAR ship detection methods. Bin Pan, Zhehao Xu, Tianyang Shi, Tao Li 0022, Zhenwei Shi 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2023 | Unmixing Guided Unsupervised Network for RGB Spectral Super-ResolutionabstractSpectral super-resolution has attracted research attention recently, which aims to generate hyperspectral images from RGB images. However, most of the existing spectral super-resolution algorithms work in a supervised manner, requiring pairwise data for training, which is difficult to obtain. In this paper, we propose an Unmixing Guided Unsupervised Network (UnGUN), which does not require pairwise imagery to achieve unsupervised spectral super-resolution. In addition, UnGUN utilizes arbitrary other hyperspectral imagery as the guidance image to guide the reconstruction of spectral information. The UnGUN mainly includes three branches: two unmixing branches and a reconstruction branch. Hyperspectral unmixing branch and RGB unmixing branch decompose the guidance and RGB images into corresponding endmembers and abundances respectively, from which the spectral and spatial priors are extracted. Meanwhile, the reconstruction branch integrates the above spectral-spatial priors to generate a coarse hyperspectral image and then refined it. Besides, we design a discriminator to ensure that the distribution of generated image is close to the guidance hyperspectral imagery, so that the reconstructed image follows the characteristics of a real hyperspectral image. The major contribution is that we develop an unsupervised framework based on spectral unmixing, which realizes spectral super-resolution without paired hyperspectral-RGB images. Experiments demonstrate the superiority of UnGUN when compared with some SOTA methods. Qiaoying Qu, Bin Pan, Tao Li 0022, Zhenwei Shi 0001 |
IEEE Trans. Image Process. | 4 |
| 2023 | H2Former: An Efficient Hierarchical Hybrid Transformer for Medical Image SegmentationabstractAccurate medical image segmentation is of great significance for computer aided diagnosis. Although methods based on convolutional neural networks (CNNs) have achieved good results, it is weak to model the long-range dependencies, which is very important for segmentation task to build global context dependencies. The Transformers can establish long-range dependencies among pixels by self-attention, providing a supplement to the local convolution. In addition, multi-scale feature fusion and feature selection are crucial for medical image segmentation tasks, which is ignored by Transformers. However, it is challenging to directly apply self-attention to CNNs due to the quadratic computational complexity for high-resolution feature maps. Therefore, to integrate the merits of CNNs, multi-scale channel attention and Transformers, we propose an efficient hierarchical hybrid vision Transformer (H2Former) for medical image segmentation. With these merits, the model can be data-efficient for limited medical data regime. The experimental results show that our approach exceeds previous Transformer, CNNs and hybrid methods on three 2D and two 3D medical image segmentation tasks. Moreover, it keeps computational efficiency in model parameters, FLOPs and inference time. For example, H2Former outperforms TransUNet by 2.29% in IoU score on KVASIR-SEG dataset with 30.77% parameters and 59.23% FLOPs. Along He, Kai Wang 0001, Tao Li 0022, Chengkun Du, Huazhu Fu |
IEEE Trans. Medical Imaging | 3 |
| 2023 | TSC-VEE: A TrustZone-Based Smart Contract Virtual Execution EnvironmentabstractTrustZone as a trusted execution environment (TEE) has been proven to preserve the confidentiality of blockchain transactions supported by smart contracts. Despite some academic effort, TrustZone can only support limited languages for now. The lack of the corresponding execution environment for smart contracts seriously hinders blockchain applications from directly running on TrustZone. In this paper, we design the first virtual execution environment named TSC-VEE for performing Solidity smart contracts on TrustZone, to the best of our knowledge. TSC-VEE can be decomposed into fourfold: (1) an instruction set adapted to the isolation and world switching mechanism of TrustZone. (2) a runtime memory management mechanism that provides a pair of instructions with the corresponding processing mechanism to allocate and release the work memory. (3) a hybrid granularity resource analysis algorithm which computes and records the value of maximum stack height and static gas cost through bytecode pre-execution, avoiding runtime overflow and invalid computations. (4) a cross-isolation-environment prefetching approach that supports loading and storing the storage data from the normal world into the secure world on TrustZone before execution, thus avoiding switching the world state frequently at runtime. Extensive experimental results show that TSC-VEE can perform smart contracts correctly and efficiently on TrustZone. Compared with the most commonly used Ethereum client—Geth, TSC-VEE achieves execution performance improvements by$9.29\times$. We also implement the Ethereum virtual machine—evmoneon TrustZone. TSC-VEE can reduce the latency by 12.63% with our optimization techniques, and decrease the work memory footprint by 22.95% on average when executing various scale contracts. Zhaolong Jian, Ye Lu 0004, Youyang Qiao, Yaozheng Fang, Xueshuo Xie, Dayi Yang, Tao Li 0022 |
IEEE Trans. Parallel Distributed Syst. | 8 |
| 2022 | Web table data integration based on smart campus scenarios to resolve name disambiguation of scientific research personnelabstractName ambiguity issue that results from the similarity of many common Chinese names. With the development of artificial intelligence, the disambiguation model based on machine learning has achieved better disambiguation effects and has been widely used in various universities. However, continually improving the disambiguation effect remains a major challenge. Smart campuses based on the Internet of Things are developing rapidly, and a large number of discretely distributed web tables that omit data values exist. However, the usable attributes of the disambiguation model are limited. To overcome these challenges, this study proposes a name disambiguation model of web tables from data integration (NDWT) in smart campuses. The model first recognises the label mapping in a webpage table using four types of label matchers and then designs the instance comparator based on the obtained label mapping. The web tables are integrated according to the instance mapping relationship, and two datasets, one before (BWT) and the other after (A WT) integration, are obtained. Relevant features are subsequently extracted from these two datasets and trained. Finally, the NDWT model is used for disambiguation experiments. Comparative experiments, condu-cted using seven different types of ML models, show that the NDWT model improves significantly after the integration of web tables; in particular, the pairwise F1 of the K-means model increases by 43.23%. The pairwise F1 of the remaining models increases by approximately 10%. The experimental evaluation proves the feasibility of the NDWT model proposed in this study. Confirming that it can achieve a higher distribution quality compared to conventional name disambiguation methods. Junfan Jin, Junxiang Chen, Tao Li 0022, Ruixiang Qian, Li Zhou 0008 |
COMPSAC | 4 |
| 2022 | Long-Term Person Re-identification with Dramatic Appearance Change: Algorithm and BenchmarkabstractFor person re-identification (Re-ID) task, most of previous studies assumed that the pedestrians do not change their appearances. The works on cross-appearance Re-ID, including datasets and algorithms, are still few. Therefore, this paper contributes a cross-season appearance change Re-ID dataset, namely NKUP+, including more than 300 IDs from surveillance videos over 10 months, to support the studies of the cross-appearance Re-ID. In addition, we propose a network named M2Net, which integrates multi-modality features from the RGB images, contour images and human parsing images. By ignoring irrelevant misleading information for cross-appearance retrieval in RGB images, M2Net can learn features that are robust to appearance changes. Meanwhile, we propose a sampling strategy called RAS to contain a variety of appearances in one batch. And appearance loss and multi-appearance loss are designed to guide the network to learn both same-appearance and cross-appearance features. Finally, we evaluated our method on NKUP+/PRCC/DeepChange datasets, and the results showed that, compared with the baseline, our method renders significant improvement, leading to the state-of-the-art performance over other methods. Our dataset is available at https://github.com/nkicsl/NKUP-dataset. Tao Li 0022, Yanfeng Jiang, Kai Wang 0001 |
ACM Multimedia | 3 |
| 2022 | A systematic study on benchmarking AI inference accelerators
Zihan Jiang 0006, Jiansong Li, Fangxin Liu, Wanling Gao, Lei Wang 0004, Chuanxin Lan, Fei Tang 0003, Lei Liu 0030, Tao Li 0022 |
CCF Trans. High Perform. Comput. | 9 |
| 2022 | ATOM: Architectural Support and Optimization Mechanism for Smart Contract Fast Update and Execution in Blockchain-Based IoTabstractBlockchain-based Internet of Things (BC-IoT) brings the advantages of blockchain into traditional IoT systems. In BC-IoT, the smart contract has been widely used for automatic, trusted, and decentralized applications. Smart contracts require frequent adjust and fast update due to various reasons, such as inevitable code bugs, changes of applications, or security requirements. However, previous smart contract architecture and updating mechanism are low speed and cause high overhead, because they are based on recompilation and redeployment in BC-IoT. Meanwhile, smart contract execution is so time consuming due to contract instruction dispatching and operand loading in the stack-based Ethereum virtual machine (EVM). To address these issues, we propose a new smart contract architecture and optimization mechanism for BC-IoTs, ATOM, which provides architectural supports to update contract economically and fast executing in instructionwise for the first time, to the best of our knowledge. We design a compact Application-oriented Instruction (AoI) set to describe application operations. We can construct the bytecode of smart contract from application by directly assembling templates prebuilt upon the AoIs rather than by compilation. We also present an optimized mechanism for AoI execution to enable access addressable storage place rather than the indirect access through stack. We perform ATOM on a BC-IoT testbed based on private Ethereum and Hyperledger Burrow. The experimental results highlight that ATOM is more efficient than state-of-the-art approaches. ATOM can reduce update latency by 62.7%, ledger size by 70%, and gas usage by 90% on average, respectively. Compared with the traditional smart contract architecture, ATOM can improve EVM Memory access efficiency significantly by up to$10\times $and achieve improvement of execution efficiency with up to$1.6\times $. Tao Li 0022, Yaozheng Fang, Zhaolong Jian, Xueshuo Xie, Ye Lu 0004, Grace Guiling Wang |
IEEE Internet Things J. | 1 |
| 2022 | Deep Autoencoder for Hyperspectral Unmixing via Global-Local SmoothingabstractHyperspectral unmixing is to decompose the mixed pixels into pure spectral signatures (endmembers) and their proportions (abundances). Recently, deep learning-based methods have been applied to enhance the representation ability of unmixing models by extracting joint spatial–spectral characteristics of the hyperspectral data. However, most deep learning based-unmixing methods usually conduct global smoothing by convolutions on the whole hyperspectral imagery, which may ignore the variations within the imagery and result in oversmoothing. In this article, we propose a deep network for hyperspectral unmixing based on a new global–local smoothing autoencoder (GLA). GLA is an unsupervised model, which aims at exploring the local homogeneity and the global self-similarity of hyperspectral imagery. The proposed GLA network mainly includes two modules: a Local Continuous conditional random field Smoothing (LCS) module and a global recurrent smoothing (GRS) module. In LCS, we propose a conditional random field-based smoothing strategy to describe the joint spatial–spectral information within a local homogeneity region, which also reduces the risk of abundance maps boundary blurry. In GRS, we follow the self-similarity assumption for hyperspectral imagery and develop a recurrent neural network structure to exploit potential long-distance dependency relationships among pixels. The GLA is compared with several state-of-the-art unmixing methods on both real and synthetic data, and the abundance estimation results indicate that our method is promising. We will publish the code of GLA if this article has the honor to be accepted. Xinyu Song 0004, Tao Li 0022, Zhenwei Shi 0001, Bin Pan |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | Progressive Multiscale Consistent Network for Multiclass Fundus Lesion SegmentationabstractEffectively integrating multi-scale information is of considerable significance for the challenging multi-class segmentation of fundus lesions because different lesions vary significantly in scales and shapes. Several methods have been proposed to successfully handle the multi-scale object segmentation. However, two issues are not considered in previous studies. The first is the lack of interaction between adjacent feature levels, and this will lead to the deviation of high-level features from low-level features and the loss of detailed cues. The second is the conflict between the low-level and high-level features, this occurs because they learn different scales of features, thereby confusing the model and decreasing the accuracy of the final prediction. In this paper, we propose a progressive multi-scale consistent network (PMCNet) that integrates the proposed progressive feature fusion (PFF) block and dynamic attention block (DAB) to address the aforementioned issues. Specifically, PFF block progressively integrates multi-scale features from adjacent encoding layers, facilitating feature learning of each layer by aggregating fine-grained details and high-level semantics. As features at different scales should be consistent, DAB is designed to dynamically learn the attentive cues from the fused features at different scales, thus aiming to smooth the essential conflicts existing in multi-scale features. The two proposed PFF and DAB blocks can be integrated with the off-the-shelf backbone networks to address the two issues of multi-scale and feature inconsistency in the multi-class segmentation of fundus lesions, which will produce better feature representation in the feature space. Experimental results on three public datasets indicate that the proposed method is more effective than recent state-of-the-art methods. Along He, Kai Wang 0001, Tao Li 0022, Wang Bo, Hong Kang, Huazhu Fu |
IEEE Trans. Medical Imaging | 3 |
| 2022 | Elastic Significant Bit Quantization and Acceleration for Deep Neural NetworksabstractQuantization has been proven to be a vital method for improving the inference efficiency of deep neural networks (DNNs). However, it is still challenging to strike a good balance between accuracy and efficiency while quantizing DNN weights or activation values from high-precision formats to their quantized counterparts. We propose a new method called elastic significant bit quantization(ESB) that controls the number of significant bits of quantized values to obtain better inference accuracy with fewer resources. We design a unified mathematical formula to constrain the quantized values of the ESB with a flexible number of significant bits. We also introduce a distribution difference aligner (DDA) to quantitatively align the distributions between the full-precision weight or activation values and quantized values. Consequently, ESB is suitable for various bell-shaped distributions of weights and activation of DNNs, thus maintaining a high inference accuracy. Benefitting from fewer significant bits of quantized values, ESB can reduce the multiplication complexity. We implement ESB as an accelerator and quantitatively evaluate its efficiency on FPGAs. Extensive experimental results illustrate that ESB quantization consistently outperforms state-of-the-art methods and achieves average accuracy improvements of 4.78%, 1.92%, and 3.56% over AlexNet, ResNet18, and MobileNetV2, respectively. Furthermore, ESB as an accelerator can achieve 10.95 GOPS peak performance of 1k LUTs without DSPs on the Xilinx ZCU102 FPGA platform. Compared with CPU, GPU, and state-of-the-art accelerators on FPGAs, the ESB accelerator can improve the energy efficiency by up to 65, 11, and 26, respectively. Ye Lu 0004, Kunpeng Xie, Zongming Jin, Tao Li 0022, Yanzhi Wang 0001 |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2022 | SmartVM: A Smart Contract Virtual Machine for Fast On-Chain DNN ComputationsabstractBlockchain-based artificial intelligence (BC-AI) has been applied for protecting deep neural network (DNN) data from being tampered with, which is expected to further boost trusted distributed AI applications in many fields. However, due to smart contract execution environment architectural defects, it is challenging for previous BC-AI systems to support computing-intensive tasks on-chain performing such as DNN convolution operations. They have to offload computations and a large amount of data from blockchain to off-chain platforms to execute smart contracts as native code. This failure to take advantage of data locality has become one of the major critical performance bottlenecks in BC-AI system. To this end, in this article, we propose SmartVM with optimization methods to support on-chain DNN inference for BC-AI system. The key idea is to design and optimize the computing mechanism and storage structure of smart contract execution environment according to the characteristics of DNN such as high computational parallelism and large data volume. We decompose SmartVM into three components: 1) a compact DNN-oriented instruction set to describe computations in a short number of instructions to reduce interpretation time. 2) a memory management mechanism to make SmartVM memory dynamic free/allocated according to the size of DNN feature maps. 3) a block-based weight prefetching and parallel computing method to organize each layer's computing and weights prefetching in a pipelined manner. We perform the typical image classification in a private Ethereum blockchain testbed to evaluate SmartVM performance. Experimental results highlight that SmartVM can support DNN inference on-chain with roughly the same efficiency against the native code execution. Compared with the traditional off-chain computing, SmartVM can speed up the overall execution by70×,16×,11×, and12×over LeNet5, AlexNet, ResNet18, and MobileNet, respectively. The memory footprint can be reduced by84%,90.8%,94.3%, and93.7%over the above four models, while offering the same level model accuracy. This article sheds light on the design space of the smart contract virtual machine for DNN computation and is promising to further boost BC-AI applications. Tao Li 0022, Yaozheng Fang, Ye Lu 0004, Jinni Yang, Zhaolong Jian, Zhiguo Wan, Yusen Li |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2021 | Editorial for the special issue on reliability and power efficiency for HPC
Jifeng He 0001, Chenggang Wu 0002, Huawei Li 0001, Yang Guo 0003, Tao Li 0022 |
CCF Trans. High Perform. Comput. | 5 |
| 2021 | OSN: Onion-ring support neighbors for correspondence selection
Ye Lu 0004, Chunying Song, Tao Li 0022, Kai Wang 0001 |
Inf. Sci. | 4 |
| 2021 | Applications of deep learning in fundus images: A review
Tao Li 0022, Wang Bo, Hong Kang, Hanruo Liu, Kai Wang 0001, Huazhu Fu |
Medical Image Anal. | 1 |
| 2021 | A Confidence-Guided Evaluation for Log Parsers Inner Quality
Xueshuo Xie, Zhi Wang 0014, Xuhang Xiao, Ye Lu 0004, Shenwei Huang, Tao Li 0022 |
Mob. Networks Appl. | 6 |
| 2021 | VecQ: Minimal Loss DNN Model Compression With Vectorized Weight QuantizationabstractQuantization has been proven to be an effective method for reducing the computing and/or storage cost of DNNs. However, the trade-off between the quantization bitwidth and final accuracy is complex and non-convex, which makes it difficult to be optimized directly. Minimizing direct quantization loss (DQL) of the coefficient data is an effective local optimization method, but previous works often neglect the accurate control of the DQL, resulting in a higher loss of the final DNN model accuracy. In this paper, we propose a novel metric, called Vector Loss. Using this new metric, we decompose the minimization of the DQL to two independent optimization processes, which significantly outperform the traditional iterative L2 loss minimization process in terms of effectiveness, quantization loss as well as final DNN accuracy. We also develop a new DNN quantization solution called VecQ, which provides minimal direct quantization loss and achieve higher model accuracy. In order to speed up the proposed quantization process during model training, we accelerate the quantization process with a parameterized probability estimation method and template-based derivation calculation. We evaluate our proposed algorithm on MNIST, CIFAR, ImageNet, IMDB movie review and THUCNews text data sets with numerical DNN models. The results demonstrate that our proposed quantization solution is more accurate and effective than the state-of-the-art approaches yet with more flexible bitwidth support. Moreover, the evaluation of our quantized models on Salient Object Detection (SOD) tasks maintains comparable feature extraction quality with up to 16× weight size reduction. Yao Chen 0008, Ye Lu 0004, Tao Li 0022, Cong Hao, Deming Chen |
IEEE Trans. Computers | 4 |
| 2021 | Simultaneously Multiobjective Sparse Unmixing and Library Pruning for Hyperspectral ImageryabstractSparse hyperspectral unmixing has attracted increasing investigations during the past decade. Recent research has indicated that library pruning algorithms can significantly improve the unmixing accuracies by reducing the mutual coherence of the spectral library. Inspired by the good performance of library pruning, in this article we propose a new hyperspectral unmixing algorithm which integrates the idea of library pruning and sparse representation. An obvious challenge for pruning algorithms is that the real endmembers must be preserved after pruning. Unfortunately, recent proposed pruning algorithms, such as multiple signal classification are actually prepruning strategies, which cannot guarantee that the endmembers exactly exist in the selected spectral subset when the image noise is strong. To overcome this difficulty, we develop a simultaneous optimization approach which involves the pruning operation into the optimization process. Compared with existing prepruning-based unmixing methods, the proposed algorithm can gradually compress the search space of sparse representation, which may relieve the loss of spectral information caused by the rapid compression of the library. Instead of simply designing a regularizer, in this article we utilize a multiobjective-based framework where reconstruction error, sparsity error, and the pruning projection function are considered as three parallel objectives, so as to avoid the manually settings of regularization parameters. Moreover, we have provided theoretical analysis and proof for the reasonability of our pruning objective. Experiments on synthetic hyperspectral data may indicate the superiority of the proposed method under high-noise conditions. Bin Pan, Herman Z. Q. Chen, Zhenwei Shi 0001, Tao Li 0022 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2021 | CABNet: Category Attention Block for Imbalanced Diabetic Retinopathy GradingabstractDiabetic Retinopathy (DR) grading is challenging due to the presence of intra-class variations, small lesions and imbalanced data distributions. The key for solving fine-grained DR grading is to find more discriminative features corresponding to subtle visual differences, such as microaneurysms, hemorrhages and soft exudates. However, small lesions are quite difficult to identify using traditional convolutional neural networks (CNNs), and an imbalanced DR data distribution will cause the model to pay too much attention to DR grades with more samples, greatly affecting the final grading performance. In this article, we focus on developing an attention module to address these issues. Specifically, for imbalanced DR data distributions, we propose a novel Category Attention Block (CAB), which explores more discriminative region-wise features for each DR grade and treats each category equally. In order to capture more detailed small lesion information, we also propose the Global Attention Block (GAB), which can exploit detailed and class-agnostic global attention feature maps for fundus images. By aggregating the attention blocks with a backbone network, the CABNet is constructed for DR grading. The attention blocks can be applied to a wide range of backbone networks and trained efficiently in an end-to-end manner. Comprehensive experiments are conducted on three publicly available datasets, showing that CABNet produces significant performance improvements for existing state-of-the-art deep architectures with few additional parameters and achieves the state-of-the-art results for DR grading. Code and models will be available at https://github.com/he2016012996/CABnet. Along He, Tao Li 0022, Kai Wang 0001, Huazhu Fu |
IEEE Trans. Medical Imaging | 2 |
| 2021 | A Thread Level SLO-Aware I/O Framework for Embedded VirtualizationabstractWith the development of virtualization technology, it is practical and necessary to integrate virtual machine software into embedded systems. I/O scheduling is important for embedded systems, because embedded systems always face different situations and their requests have more diversity on the requirement of real-time and importance. However, the semantic information associated with the I/O data is completely lost when crossing the virtualized I/O software stack. Here, we present an I/O scheduling framework to connect the semantic gap between the application threads in virtual machines and hardware schedulers in the host machine. Therefore, the details for the I/O request can be passed through the layers of the software stack and each layer can get the specific information about the device environment. Also, various scheduling points have been provided to implement different I/O strategies. Our framework was implemented based on Linux operating system, KVM, QEMU and virtio protocol. A prototype scheduler, Orthrus, was implemented to evaluate the effectiveness of the framework. Comprehensive experiments were conducted and the results show that our framework can guarantee the real-time requirements, and reserve more system resources for critical tasks, with negligible memory consumption and throughput overhead. Xiaoli Gong, Dingyuan Cao 0001, Yusen Li, Jin Zhang 0003, Tao Li 0022 |
IEEE Trans. Parallel Distributed Syst. | 7 |
| 2021 | Fast Policy Interpretation and Dynamic Conflict Resolution for Blockchain-Based IoT SystemabstractAlthough the blockchain‐based Internet of Things (BC‐IoT) has been applied in many fields, it still faces many security attacks due to lacking policy‐based security management (PbSM). Previous PbSM is usually time‐consuming, which is difficult to integrate into BC‐IoT directly. The high‐latency policy conflict resolving in traditional PbSM cannot meet the BC‐IoT’s low‐latency requirement. Moreover, the conflict resolution rate is low as the PbSM usually neglects the runtime information. Therefore, it is challenging that achieving an efficient PbSM for BC‐IoT and overcomes both time and resource consumption. To address the problem, we propose a novel PbSM for BC‐IoT named FPICR to realize fast policy interpretation and dynamic conflict resolution efficiently. We first present policy templates based on system log to interpret policy in high speed in BC‐IoT. Benefiting from matching the characteristics of the system processing, FPICR supports interpreting a policy into the smart contract directly without complex content parsing. We then propose a weighted directed policy graph (WDPG) to evaluate the importance of the deployed policies more accurately. To improve the policy conflict resolution rate, we implement the resolution algorithm through reconstructing the WDPG. Taking the traits of these properties, FPICR thus can also remove the redundant data to compress storage space by the WDPG. Experiment results highlight that FPICR outperforms the baseline in all measure metrics. Especially, compared with the state‐of‐the‐art method, the speedup of interpretation in FPICR is about up to 2.1×. The conflict resolution rate in FPICR can be improved by 6.2% on average and achieve up to 96.1%. Yaozheng Fang, Zhaolong Jian, Zongming Jin, Xueshuo Xie, Ye Lu 0004, Tao Li 0022 |
Wirel. Commun. Mob. Comput. | 6 |
| 2020 | A benchmark for clothes variation in person re-identificationabstractPerson re-identification (re-ID) has drawn attention significantly in the computer vision society due to its application and research significance. It aims to retrieve a person of interest across different camera views. However, there are still several factors that hinder the applications of person re-ID. In fact, most common data sets either assume that pedestrians do not change their clothing across different camera views or are taken under constrained environments. Those constraints simplify the person re-ID task and contribute to early development of person re-ID, yet a person has a great possibility to change clothes in real life. To facilitate the research toward conquering those issues, this paper mainly introduces a new benchmark data set for person re-identification. To the best of our knowledge, this data set is currently the most diverse for person re-identification. It contains 107 persons with 9,738 images, captured in 15 indoor/outdoor scenes from September 2019 to December 2019, varying according to viewpoints, lighting, resolutions, human pose, seasons, backgrounds, and clothes especially. We hope that this benchmark data set will encourage further research on person re-identification with clothes variation. Moreover, we also perform extensive analyses on this data set using several state-of-the-art methods. Our dataset is available at https://github.com/nkicsl/NKUP-dataset. Kai Wang 0001, Shiyan Chen, Jinni Yang, Keke Zhou, Tao Li 0022 |
Int. J. Intell. Syst. | 6 |
| 2020 | Bin loss for hard exudates segmentation in fundus images
Song Guo 0002, Kai Wang 0001, Hong Kang, Yingqi Gao, Tao Li 0022 |
Neurocomputing | 6 |
| 2020 | Confidence guided anomaly detection model for anti-concept drift in dynamic logs
Xueshuo Xie, Zongming Jin, Jiming Wang, Ye Lu 0004, Tao Li 0022 |
J. Netw. Comput. Appl. | 6 |
| 2020 | Novel model to integrate word embeddings and syntactic trees for automatic caption generation from images
Hongbin Zhang 0004, Diedie Qiu, Renzhong Wu, Donghong Ji, Guangli Li, Zhenyu Niu, Tao Li 0022 |
Soft Comput. | 7 |
| 2019 | Random Inception Module and Its Parallel Implementation
Yingqi Gao, Kunpeng Xie, Song Guo 0002, Kai Wang 0001, Hong Kang, Tao Li 0022 |
APPT | 6 |
| 2019 | A Lightweight Neural Network for Hard Exudate Segmentation of Fundus Image
Song Guo 0002, Tao Li 0022, Kai Wang 0001, Chan Zhang, Hong Kang |
ICANN (3) | 2 |
| 2019 | Random Drop Loss for Tiny Object Segmentation: Application to Lesion Segmentation in Fundus Images
Song Guo 0002, Tao Li 0022, Chan Zhang, Hong Kang, Kai Wang 0001 |
ICANN (3) | 2 |
| 2019 | LHC: A Low-Power Heterogeneous Computing Method on Neural Network AcceleratorabstractAccelerators can achieve high performance and low energy consumption in training or inference of neural networks. If the Non-Neural Network (Non-NN) algorithms with large amount of computation could make full use of the accelerators, it is possible to speed up its implementation, reduce energy consumption, and achieve load balancing, especially on mobile devices equipped with accelerators. However, accelerators are dedicated to neural network calculations, so that other Non-NN algorithms have difficulty in using their advantages. Furthermore, many hardware-specific restrictions have become the obstacles, such as constrained precision of operands and limited computation scale. In this paper, we propose a method named Low-power Heterogeneous Computing (LHC) to bridge the gap between Non-NN algorithms and NN accelerators. Firstly, we analyze the general principle of the accelerator and reveal the calculation model of the accelerator. To hide the details of the underlying neural network library, we extract some operators from the limited number of types of neural network computation they support. We encapsulate the low-level library, extract operators suitable for general algorithms, and implement some more advanced operators that can adapt to the constrained hardware conditions. These operators could facilitate programmers to implement some Non-NN algorithms. In the aspect of the algorithm, we extract the computationally intensive parts of the Non-NN algorithm and deploy these computational tasks on the accelerator by calling the operators. To verify our method, we implement three Non-NN algorithms by using operators and adjusting these algorithms, include Grid-based Motions Statistics, k-Nearest Neighbors, and k-Means, on a specific accelerator, Cambricon-1A. The experimental results show that the energy consumption of calculation is reduced by up to 5.4x, compared with the CPU baseline. Our method can be further applied to other similar accelerators. Fangxin Liu, Kunpeng Xie, Shusheng Liu, Ye Lu 0004, Tao Li 0022 |
ICPADS | 6 |
| 2019 | µL2Q: An Ultra-Low Loss Quantization Method for DNN CompressionabstractData quantization has been proved to be an effective method to compress deep neural networks (DNNs) by using less bits to represent the parameters and intermediate data. The bit width of the data directly affects the memory footprint, computing capability, and energy consumption during the computation of the DNN models. Although there have been numerous existing studies on data quantization, there is still no quantitative analysis of the existing quantization methods, which results in empirical quantization with unpredictable DNN accuracy loss. To address this problem, we propose an effective method, called ultra-low loss quantization (μL2Q), to provide DNN quantization schemes based on comprehensive quantitative data analysis. μL2Q builds the transformation of the original data to a data space with standard normal distribution, and then find the optimal parameters to minimize the loss of the quantization of a targeted bit width. In addition, we integrate the proposed μL2Q into a popular machine learning framework Caffe for convenient end-to-end DNN design and training. By comparing to the state-of-the-art DNN compression designs, μL2Q shows the greatest ability to maintain DNN accuracy after quantization. In the experiments, our proposed method can deliver 4.42%, 16.70%, 1.95%, 8.26% and 5.63% accuracy improvements on Lenet-5, Cifarnet, VGG7-64 and Resnet-18 (Top1/5), respectively, compared to the state-of-the-art solutions with the same compression ratio. Tao Li 0022, Ye Lu 0004, Cong Hao, Xiaofan Zhang 0001, Deming Chen, Yao Chen 0008 |
IJCNN | 2 |
| 2019 | Aggregation Connection Network For Tiny Face DetectionabstractFace detection has been greatly developed in recent years. Despite the remarkable progress, finding tiny faces in the wild is still a challenge due to the vastly scales, blur, occlusion and low resolution. This paper proposes an Aggregation Connection Network (ACN) which robustly solves these problems in tiny face detection. ACN utilizes the features from different convolution layers and performs superiorly on finding multi-scale faces in a single shot, especially for tiny faces. Specially, there are two novel modules in ACN that play significant roles: an aggregation connection module and a context module. First, by integrating efficient aggregation connection module, our ACN can effectively reduce the feature disappearance caused by image scaling. Second, the elaborately designed context module can make full use of the rich contextual cues without adding extra parameters. As a consequence, our ACN achieves state-of-the-art detection performance among several popular face detection benchmarks i.e. WIDER FACE, FDDB and Pascal Face. Chan Zhang, Tao Li 0022, Song Guo 0002, Yingqi Gao, Kai Wang 0001 |
IJCNN | 2 |
| 2019 | Sum of weighted distances in trees
Qingqiong Cai, Tao Li 0022, Yongtang Shi, Hua Wang 0003 |
Discret. Appl. Math. | 2 |
| 2019 | Critical (P6, banner)-free graphs
Shenwei Huang, Tao Li 0022, Yongtang Shi |
Discret. Appl. Math. | 2 |
| 2019 | L-Seg: An end-to-end unified framework for multi-lesion segmentation of fundus images
Song Guo 0002, Tao Li 0022, Hong Kang, Yujun Zhang 0001, Kai Wang 0001 |
Neurocomputing | 2 |
| 2019 | Diagnostic assessment of deep learning algorithms for diabetic retinopathy screening
Tao Li 0022, Yingqi Gao, Kai Wang 0001, Song Guo 0002, Hanruo Liu, Hong Kang |
Inf. Sci. | 1 |
| 2019 | DCNR: deep cube CNN with random forest for hyperspectral image classification
Tao Li 0022, Jiabing Leng, Lingyan Kong, Song Guo 0002, Gang Bai, Kai Wang 0001 |
Multim. Tools Appl. | 1 |
| 2019 | Dual buffer rotation four-stage pipeline for CPU-GPU cooperative computing
Tao Li 0022, Qiankun Dong, Xiaoli Gong, Yulu Yang |
Soft Comput. | 1 |
| 2019 | Using Sparse Representation to Detect Anomalies in Complex WSNsabstractIn recent years, wireless sensor networks (WSNs) have become an active area of research for monitoring physical and environmental conditions. Due to the interdependence of sensors, a functional anomaly in one sensor can cause a functional anomaly in another sensor, which can further lead to the malfunctioning of the entire sensor network. Existing research work has analysed faulty sensor anomalies but fails to show the effectiveness throughout the entire interdependent network system. In this article, a dictionary learning algorithm based on a non-negative constraint is developed, and a sparse representation anomaly node detection method for sensor networks is proposed based on the dictionary learning. Through experiment on a specific thermal power plant in China, we verify the robustness of our proposed method in detecting abnormal nodes against four state of the art approaches and proved our method is more robust. Furthermore, the experiments are conducted on the obtained abnormal nodes to prove the interdependence of multi-layer sensor networks and reveal the conditions and causes of a system crash. Xiaoming Li 0006, Guangquan Xu, James Xi Zheng, Kaitai Liang, Emmanouil A. Panaousis, Tao Li 0022, Wei Wang 0012, Chao Shen 0001 |
ACM Trans. Intell. Syst. Technol. | 6 |
| 2018 | CubemapSLAM: A Piecewise-Pinhole Monocular Fisheye SLAM System
Shaojun Cai, Shijie Li 0006, Yun Liu 0011, Yangyan Guo, Tao Li 0022, Ming-Ming Cheng |
ACCV (6) | 6 |
| 2018 | Improving Dynamically-Generated Code Performance on Dynamic Binary TranslatorsabstractThe recent transition in the software industry toward dynamically generated code poses a new challenge to existing dynamic binary translation (DBT) systems. A significant re-translation overhead could be introduced due to the maintenance of the consistency between the dynamically-generated guest code and the corresponding translated host code. To address this issue, this paper presents a novel approach to optimize DBT systems for guest applications with dynamically-generated code. The proposed approach can maximize the reuse of previously translated host code to mitigate the re-translation overhead. A prototype based on such an approach has been implemented on an existing DBT system HQEMU. Experimental results on a set of JavaScript applications show that it can achieve a 1.24X performance speedup on average compared to the original HQEMU. Wenwen Wang 0001, Jiacheng Wu 0001, Xiaoli Gong, Tao Li 0022, Pen-Chung Yew |
VEE | 4 |
| 2017 | Structure-Measure: A New Way to Evaluate Foreground MapsabstractForeground map evaluation is crucial for gauging the progress of object segmentation algorithms, in particular in the field of salient object detection where the purpose is to accurately detect and segment the most salient object in a scene. Several widely-used measures such as Area Under the Curve (AUC), Average Precision (AP) and the recently proposed F W/B (Fbw) have been used to evaluate the similarity between a non-binary saliency map (SM) and a ground-truth (GT) map. These measures are based on pixel-wise errors and often ignore the structural similarities. Behavioral vision studies, however, have shown that the human visual system is highly sensitive to structures in scenes. Here, we propose a novel, efficient, and easy to calculate measure known as structural similarity measure (Structure-measure) to evaluate non-binary foreground maps. Our new measure simultaneously evaluates region-aware and object-aware structural similarity between a SM and a GT map. We demonstrate superiority of our measure over existing ones using 5 meta-measures on 5 benchmark datasets. Deng-Ping Fan, Ming-Ming Cheng, Yun Liu 0011, Tao Li 0022, Ali Borji |
ICCV | 4 |
| 2017 | A comparative analysis of new graph distance measures and graph edit distance
Tao Li 0022, Han Dong, Yongtang Shi, Matthias Dehmer |
Inf. Sci. | 1 |
| 2017 | An intelligent character recognition method to filter spam images on cloud
Jufeng Yang, Tao Li 0022, Kai Wang 0001 |
Soft Comput. | 5 |
| 2016 | Cube-CNN-SVM: A Novel Hyperspectral Image Classification MethodabstractCNNs (convolutional neural networks) have been proved to be efficient deep learning models that can directly extract high level features from raw data. In this paper, a novel CCS (Cube-CNN-SVM) method is proposed for hyperspectral image classification, which is a spectral-spatial feature based hybrid model of CNN and SVM (support vector machine). Different from most of traditional methods that only take spectral information into consideration, a target pixel and the spectral information of its neighbors are organized into a spectral-spatial multi-feature cube used in hyperspectral image classification. It is a straightforward but valid spatial strategy that can easily improve classification accuracy without extra modification of deep CNN's structure except the size of input layer and convolutional kernel. Our deep CNN consists of the input layer, convolutional layer, max pooling layer, full connection layer and output layer. To further improve hyperspectral image classification accuracy, SVM is trained as hyperspectral image classifier with the features extracted by deep CNN from spectral-spatial fusion information. Three hyperspectral image datasets such as the KSC (Kennedy Space Center), PU (Pavia University Scene) and Indian Pines are used to evaluate the performance of CCS method. Experimental results indicate that the hyperspectral image classification can be improved efficiently with the spectral-spatial fusion strategy and CCS method. Firstly, it is easy to implement the spatial strategy to improve classification accuracy about 4% compared with only spectral information used for classification, in which 98.49% is gained on the KSC dataset. Secondly, CCS method can further improve classification accuracy about 1%~3% compared to the best performance of deep CNN, in which 99.45% is gained on the PU dataset. Jiabing Leng, Tao Li 0022, Gang Bai, Qiankun Dong, Han Dong |
ICTAI | 2 |
| 2016 | HPSVM: Heterogeneous Parallel SVM with Factorization Based IPM Algorithm on CPU-GPU ClusterabstractSupport vector machine (SVM) is a supervised method widely used in the statistical classification and regression analysis. SVM training can be solved via the interior point method (IPM) with the advantages of low storage, fast convergence and easy parallelization. However, it is still confronted with the challenges of training speed and memory use. In this paper, we propose a parallel primal-dual IPM algorithm based on the incomplete Cholesky factorization (ICF) for efficiently training large-scale SVMs, named HPSVM, on CPU-GPU cluster. Our approach is distinguished from earlier work in that it is specifically designed to take maximal advantage of the CPU-GPU collaborative computation with the dual buffers 3-stage pipeline mechanism, and efficiently handles large-scale training datasets. In HPSVM, the heterogeneous hierarchical memory is fully explored to alleviate the bottleneck for optimizing data transfer, and the programming paradigm is presented to build an efficient collaboration mechanism between CPU and GPU. Comprehensive experiments show that HPSVM is up to 11 times faster than the CPU version on real datasets. Tao Li 0022, Xuechen Liu 0002, Qiankun Dong, Wenjing Ma, Kai Wang 0001 |
PDP | 1 |
| 2015 | PE-TLD: Parallel Extended Tracking-Learning-Detection for Multi-target Tracking
Chenggang Zhou, Qiankun Dong, Wenjing Ma, Guoping Long, Tao Li 0022 |
ICA3PP (2) | 5 |
| 2015 | Machine-readable region identification from partially blurred document imagesabstractPartial blur sometimes occurs in the document images captured by a camera, which will influence the performance of OCR on the non-blurred text region. A real-time method, named MRRI, is proposed in this paper to identify the machine-readable region from partially blurred document images. Firstly, a reference image is generated by low-pass filtering on the given document image. Secondly, a weight matrix is generated by calculating the structural similarity for each patch. Thirdly, a cost function is minimized to identify the maximum machine-readable region that can be well-recognized by OCR. In experiments, two applications are considered with the identified machine-readable region. On one hand, Tesseract-OCR is used for the word recognition to build index for a given document image. Compared with the results by applying OCR on the whole image, more words are correctly recognized by applying OCR on the identified region. On the other hand, the identified machine-readable region is used to assess the quality of a document image. Compared with other two image quality assessment methods, the machine-readable region based method shows a better performance. Also, MRRI is light and time-saving, which can meet the requirement of real-time applications. Qinwen Wang, Yixue Wang, Jufeng Yang, Tao Li 0022, Kai Wang 0001 |
ICDAR | 5 |
| 2015 | OCR with Adaptive Dictionary
Yanhong Xie, Kai Wang 0001, Tao Li 0022 |
ICIG (2) | 4 |
| 2015 | CPU-assisted GPU thread pool model for dynamic task parallelismabstractWith the growing power of GPUs, how to utilize the high computing performance provided by the GPU hardware becomes an urgent yet challenging problem, especially for applications with fine grained parallelism. Task programming is efficient for handling fine grained parallelism but current GPU task parallel solutions using either concurrent kernel execution (CKE) or persistent kernels suffer from a high cost of CPU-GPU interaction. The page-locked host memory supported by new generation GPUs turns CPU-GPU heterogeneous systems into the non-uniform memory access (NUMA) architecture, making it possible to improve CPU-GPU interaction with shared memory programming. In this paper, we propose the CPU-assisted GPU thread pool (CAGTP) model that combines data parallelism and task parallelism at the thread block level to support applications with fine grained parallelism. In the CAGTP model, the Computing Block Level task Scheduling (CBLS) method is designed in which task slots allocated in the page-locked host memory eliminate competition among thread blocks. A separate host scheduler is designed for scheduling tasks to thread blocks and the overhead for scheduling a task (200ns) is much lower than that of similar systems. Experiment results show that the CAGTP model supports fine grained task parallelism with or without dependencies efficiently. It outperforms CKE for batched GEMMs, Cholesky factorization and mixed workloads. Tao Li 0022, Qiankun Dong, Xuechen Liu 0002, Yulu Yang |
NAS | 2 |