Ying Gao 0004

dblp:15/1300-4 · DBLP profile ↗
← Back
58ranked-venue papers
10as first author
51since 2021 · last 2026
0000-0002-8925-8192ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 18 · 16 since 2021Applied, interdisciplinary, general and emerging computing · 18 · 5 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 1 first-author · 11 since 2021Computer networks · 6 · 2 first-author · 6 since 2021Security and privacy · 5 · 1 first-author · 5 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2026 A website fingerprinting attack with unsupervised out-of-distribution detection
abstract
Abstract Website fingerprinting attacks are critical for extracting website information and identifying illegal websites visited by users in anonymous networks such as Tor. However, existing attacks struggle to extract effective features from unmonitored websites due to the diversity. Although increasing unmonitored training data can improve effectiveness, it also increases attacker costs. To address this, we propose a novel website fingerprinting attack that leverages unsupervised Out-of-Distribution detection. We exclusively use monitored website data for model training, eliminating the need for extensive unmonitored samples. For feature extraction, we utilize a combination of Long Short-Term Memory and Convolutional Neural Networks for robust feature extraction of each monitored website. We also introduce a new loss function to maximize differentiation between features of various websites. Furthermore, we employ Singular Value Decomposition to effectively segregate monitored from unmonitored websites. It allows the model to focus on dominant components in the feature vectors, facilitating a clear distinction between monitored and unmonitored website traffic. The experimental results confirm that our method outperforms existing techniques without requiring unmonitored training data.
Ying Gao 0004, Jiafeng Zhao, Chong Chen 0011, Siquan Huang, Leyu Shi, Chenglong Jiang
Cybersecur.1
2026 VulSCC: image-based vulnerability detection with SPP-CNN and code large language model
abstract
Abstract Deep learning excels in detecting source code vulnerabilities, where image-based detection methods overcome ignoring deep code semantic information in token-based methods and the inefficiency of graph-based methods. Unfortunately, current image-based methods cannot sufficiently extract vulnerability-related features due to three key limitations: (1) the inappropriateness of the construction of node centrality for sequential Program Dependency Graphs (PDGs), (2) the ineffective code analysis of traditional embedding models, and (3) the poor existing truncation/padding methods. Moreover, they fail to achieve effective vulnerability localization due to the irregular output of their interpretation method. In response, we propose a novel image-based line-level source code vulnerability detection system VulSCC. Firstly, VulSCC constructs a novel centrality combination and leverages a code large language model to capture richer vulnerability-related features. Secondly, we integrate an SPP layer to convert PDGs into images without distortion, as it adaptively aggregates arbitrary sizes into fixed-length vectors without truncation/padding. Finally, we use the occlusion technique to interpret the model predictions, which locate specific vulnerability lines, enabling effective vulnerability localization. Experimental results of VulSCC against seven SoTA methods show optimal detection performance in function-level detection. Additionally, we evaluate the effectiveness of the occlusion technique in localizing vulnerabilities, with interpretation success rate exceeding 90%.
Zhibin Jian, Siquan Huang, Hongyi Xie, Ying Gao 0004, Leyu Shi
Cybersecur.4
2026 Facial AU Recognition With Feature-Based AU Localization and Confidence-Based Relation Mining
abstract
Facial Action Unit (AU) recognition involves identifying subtle muscle movements corresponding to different AUs. Recent approaches have focused on localizing AUs using predefined Regions of Interest (RoIs) or learnable modules. However, these methods either overly depend on the precision of predefined RoIs or inaccurately localize background regions instead of the actual AU positions. To address this challenge, we propose a novel method, which automatically localizes each AU without relying on predefined RoIs or introducing learnable modules during the inference phase. Specifically, our approach decomposes the task into two subtasks: AU localization and AU state verification. We first align the direction between spatial features and the corresponding AU class weights to guide the model in localizing AUs. Next, we incorporate spatial and temporal aspects for precise AU state detection. From the perspective of spatial information learning, we propose a confidence-based AU relationship mining module that directs the model to focus on uncertain AUs. From the aspect of temporal information learning, we introduce a temporal sampling strategy that implicitly captures time-dependent features. Experimental results on the BP4D and DISFA datasets demonstrate the effectiveness of our method, showing that it outperforms existing approaches and achieves state-of-the-art performance in AU recognition.
Wentian Cai, Yandan Chen, Xiping Hu, Ying Gao 0004
IEEE Trans. Affect. Comput.7
2026 SSL-VC: One-Shot Voice Conversion Through Self-Supervised Learning
abstract
Currently, the prevailing approach in voice conversion (VC) involves separating clearer linguistic information from the source audio and then reconstructing it with the identity of the target speaker. However, existing methods, whether employing in-formation perturbation techniques or carefully designed information bottleneck methods, encounter challenges related to unsatisfactory audio separation effects and insufficient robustness. This article introduces a VC through the self-supervised learning method (SSL-VC). First, it utilizes a self-supervised speech representation (S-SSR) extraction network with decoupling (Decp-SSEN) to disentangle linguistic information from speech. The designed prosodic encoder extracts features of pitch and energy from the speech to compensate for the loss of nonlinguistic details incurred during the Decp-SSEN disentangling process. This approach allows us to obtain richer linguistic information independent of speaker identity, guaranteeing the robust performance of the model. Second, we leverage high-level S-SSR as the intermediate feature, replacing the traditional Mel-spectrogram. Built an end-to-end VC pipeline that eliminates the need for a vocoder, enhancing the expression level of intermediate features and reducing the learning difficulty gap between real and predicted features. Subjective and objective experiments conducted on both seen and unseen speech corpus demonstrate that SSL-VC achieves high-quality VC and speaker similarity. Moreover, it outperforms state-of-the-art methods in extracting richer linguistic information. Ablation experiments further scrutinize the indispensability of the prosodic encoder.
Chenglong Jiang, Linrong Pan, Ying Gao 0004, Kuanghua Su, Gaoze Hou, Xiping Hu
IEEE Trans. Comput. Soc. Syst.3
2026 A Resource-Efficient Blockchain With Delegated Fault-Tolerance for Manufacturing Nodes
abstract
The integration of blockchain into industrial environments promises secure and verifiable data exchange; however, existing permissioned blockchain (PBC) frameworks, such as Hyperledger Fabric and Quorum, impose overheads that are unsuitable for resource-constrained systems. This article introduces LowCapChain (LCC), a lightweight PBC designed for manufacturing nodes, such as programmable logic controllers, robotic arms, and smart sensors. LCC integrates Merkle ledger compression, elliptic-curve cryptography-based proof-of-membership, and delegated fault-tolerant consensus to enable secure operations without full ledger replication. Implemented in a robotic metal stamping facility using Raspberry Pi edge nodes, LCC achieves 964 transactions per second with 108 ms consensus latency and memory usage below 55 MB. Compared with Hyperledger Fabric, LCC reduces mean consensus latency by 61% and cryptographic overhead by up to 81%. Compared with Quorum, the reductions are 54% and 76%, respectively. Scalability tests confirmed near-linear throughput growth across 10–200 nodes, and fault tolerance experiments verified block finalization under validator failure. These findings establish LCC as an efficient architecture for embedded industrial systems, offering a pathway toward scalable Industry 4.0 adoption with energy savings inferred from reduced CPU cycles rather than directly measured power.
Mohammad Iqbal Saryuddin Assaqty, Ying Gao 0004, Ali Alfatemi, Mohamed Rahouti, Abdellah Chehri
IEEE Trans. Ind. Informatics2
2026 Privacy-Preserving Rényi Layer-Wise Budget Allocation Against Gradient Leakage for Federated Learning
abstract
Federated learning (FL) is vulnerable to gradient-based privacy attacks, where malicious attackers reconstruct training data from exchanged gradients. While existing differential privacy (DP) defenses mitigate this, they often cause excessive additive noise due to the inequality scaling in the theoretical analyses, which degrades the model's utility or fail under adaptive attacks. To address this issue, we proposeFedMSBA, a layer-wise privacy-preservation method that adaptively allocates privacy budgets via Rényi DP (RDP) and modified sensitivity. FedMSBA dynamically scales noise to model intricacies and adaptively choose the better applied DP mechanisms, which provides a tighter mathematical bound and finally prevents non-convergence while resisting reconstruction attacks. Experiments demonstrate superior privacy-utility trade-offs compared to state-of-the-art defenses. FedMSBA achieves an approximately 2% improvement in accuracy and a 5% enhancement in privacy preservation. Furthermore, FedMSBA's performance remains nearly unaffected by variations in the privacy budget$\epsilon$and failure rate$\delta$.
Leyu Shi, Ying Gao 0004, Chong Chen 0011, Siquan Huang, Jiafeng Zhao, Xiping Hu
IEEE Trans. Mob. Comput.2
2025 Updatable Verifiable Credential Sharing with Selective Disclosure
Jui-Yung Lin, Xingfu Yan, Wing W. Y. Ng, Gong Zheng, Ying Gao 0004
ICA3PP (1)5
2025 Unsupervised Histopathological Image Semantic Segmentation with Overlapping Patches Consistency Constraint
Wentian Cai, Weizhao Weng, Yandan Chen, Siquan Huang, Victor C. M. Leung, Ying Gao 0004
ICCV8
2025 BMCA: Weakly Supervised Semantic Segmentation via Beta Modulation and Cross-Modality Alignment
abstract
Weakly supervised semantic segmentation (WSSS) leverages weak annotations to train semantic segmentation networks, thereby reducing the cost of pixel-level labeling. However, prevalent WSSS methods rely on Class Activation Map (CAM), which only focus on the most discriminative parts of a class, resulting in incomplete and inaccurate pseudo-labels. To resolve this, we propose a Beta Modulation (BM) method, which introduces multiple fixed forms of the Beta distribution to achieve dynamic attention to different activation level of features. BM provides supplementary information for overlooked regions and enhances the model’s ability to recognize objects as a whole. Additionally, we propose a Cross-Modality Alignment (CA) strategy to better differentiate features of semantically similar classes. It contrasts language descriptions of confusing false positive classes with global image features, yielding more distinctive and robust class representations. Extensive experiments conduct on the VOC12 natural dataset and BCSS medical dataset demonstrate state-of-the-art performance, with 83.90% mIoU on the VOC12 test set and 73.18% mIoU on the BCSS test set.
Ying Gao 0004, Wentian Cai, Yandan Chen, Zhiyong Xia
ICME1
2025 Teochew-Wild: The First In-the-wild Teochew Dataset with Orthographic Annotations
abstract
This paper reports the construction of the Teochew-Wild, a speech corpus of the Teochew dialect. The corpus includes 18.9 hours of in-the-wild Teochew speech data from multiple speakers, covering both formal and colloquial expressions, with precise orthographic and pinyin annotations. Additionally, we provide supplementary text processing tools and resources to propel research and applications in speech tasks for this low-resource language, such as automatic speech recognition (ASR) and text-to-speech (TTS). To the best of our knowledge, this is the first publicly available Teochew dataset with accurate orthographic annotations. We conduct experiments on the corpus, and the results validate its effectiveness in ASR and TTS tasks.
Linrong Pan, Chenglong Jiang, Gaoze Hou, Ying Gao 0004
ICME4
2025 GaitBranch: A multi-branch refinement model combined with frame-channel attention mechanism for gait recognition
Huakang Li, Yidan Qiu, Huimin Zhao 0001, Jin Zhan, Rongjun Chen 0001, Jinchang Ren, Ying Gao 0004, Wing W. Y. Ng
Comput. Vis. Image Underst.7
2025 Distributed clustering meets federated learning: a clustering-based approach to data poisoning mitigation
abstract
Abstract Data poisoning attacks present a significant challenge to the integrity and reliability of federated learning (FL) systems, where model training occurs collaboratively across decentralized devices. These attacks involve the deliberate injection of malicious data to corrupt the model’s training process, ultimately undermining its performance. Given the decentralized nature of FL and the lack of direct access to local data, detecting and mitigating these attacks becomes particularly difficult, especially in unsupervised scenarios where labeled data is unavailable. In this paper, we introduce a novel Federated Data Sanitization Defense to address these security threats in federated learning environments. This defense mechanism leverages federated clustering to group model updates based on semantic consistency, identifying and isolating outlier updates that are likely to be poisoned. A targeted data sanitization strategy is then applied to filter out malicious data, ensuring that only trustworthy information is used to update the global model. This decentralized process occurs on each participating device, enabling real-time detection and mitigation of data poisoning attacks. Through extensive experiments, we validate the effectiveness of Federated Data Sanitization Defense, demonstrating its ability to enhance the security and robustness of federated learning systems against data poisoning, while preserving privacy and model integrity.
Chong Chen 0011, Siquan Huang, Leyu Shi, Ying Gao 0004
Cybersecur.5
2025 FedMAR: A Privacy-Preserving and Robust Server-Side Multistage Federated Learning
abstract
In recent years, federated learning (FL) has continued to evolve with the advent of big data and the large language model (LLM), but it has also exposed numerous security and privacy issues. As a form of distributed machine learning, FL systems are more susceptible to poisoning attacks because training data are dispersed across different participants; additionally, the training achievement of FL may be subject to low-cost theft by some free-riders. Existing works have addressed defenses against the aforementioned two types of threats, but they often focus on defending against only one type and fail to effectively integrate defenses against multiple types of threats. However, in real-world Internet of Things (IoT) systems, the types of threats are not limited to just one category. In this work, we try to maintain the performance of the global model under poisoning attacks, preserve the privacy of the server under free-riders, and explore the balance between these two aspects. Therefore, this work proposes Federated Multi-Stage Asynchronous Roll-back (FedMAR), ensuring the quality of local updates; in addition, this work also provides privacy preservation in the global update process based on Rinyi Differential Privacy (RDP), and offers a certain basis for detecting free-riders. To validate the generalization of the proposed method, we conducted relevant experiments on both image and text datasets, and further investigated the robustness of the proposed method against poisoning attacks, model inversion attacks, data heterogeneity, and other aspects. The testing accuracy of the global model can even be improved by 7.2%.
Leyu Shi, Ying Gao 0004, Chong Chen 0011, Siquan Huang, Jiafeng Zhao, Xiping Hu, Victor C. M. Leung
IEEE Internet Things J.2
2025 FedCleanse: Cleanse the backdoor attacks in federated learning system
Siquan Huang, Yijiang Li, Chong Chen 0011, Leyu Shi, Wentian Cai, Ying Gao 0004
Knowl. Based Syst.6
2025 FedID: Enhancing Federated Learning Security Through Dynamic Identification
abstract
Federated learning (FL), recognized for its decentralized and privacy-preserving nature, faces vulnerabilities to backdoor attacks that aim to manipulate the model's behavior on attacker-chosen inputs. Most existing defenses based on statistical differences take effect only against specific attacks. This limitation becomes significantly pronounced when malicious gradients closely resemble benign ones or the data exhibits non-IID characteristics, making the defenses ineffective against stealthy attacks. This paper revisits distance-based defense methods and uncovers two critical insights: First, Euclidean distance becomes meaningless in high dimensions. Second, a single metric cannot identify malicious gradients with diverse characteristics. As a remedy, we propose FedID, a simple yet effective strategy employing multiple metrics with dynamic weighting for adaptive backdoor detection. Besides, we present a modified z-score approach to select the gradients for aggregation. Notably, FedID does not rely on predefined assumptions about attack settings or data distributions and minimally impacts benign performance. We conduct extensive experiments on various datasets and attack scenarios to assess its effectiveness. FedID consistently outperforms previous defenses, particularly excelling in challenging Edge-case PGD scenarios. Our experiments highlight its robustness against adaptive attacks tailored to break the proposed defense and adaptability to a wide range of non-IID data distributions without compromising benign performance.
Siquan Huang, Yijiang Li, Chong Chen 0011, Ying Gao 0004, Xiping Hu
IEEE Trans. Pattern Anal. Mach. Intell.4
2025 Scope: On Detecting Constrained Backdoor Attacks in Federated Learning
abstract
Federated learning (FL) allows multiple clients to train an efficient deep-learning model collaboratively but is susceptible to backdoor attacks. Traditional detection-based defenses depend on specific metrics to distinguish client gradients. Defense-aware attackers exploit this by constraining attack gradients on these metrics to evade detection, leading to metric-constrained attacks. This paper concretely instantiates such threats and introduces cosine-constrained attacks, which successfully compromise advanced defenses based on cosine distance. To address the aforementioned challenge, we propose Scope, a novel defense that detects cosine-constrained attacks using cosine distance by exposing the constrained backdoor dimensions of attack gradients. Scope employs dimension-wise normalization and differential scaling to amplify the distinction between backdoor dimensions and benign or unused ones, countering sophisticated attackers’ attempts to obscure them. Moreover, we develop a novel clustering approach, namely Dominant Gradient Clustering (DGC), to isolate and eliminate backdoor gradients. Extensive experiments across various datasets, models, FL settings, and adversary scenarios demonstrate that Scope consistently outperforms existing defenses by a significant margin, especially against the cosine-constrained attack. Additionally, we present a Scope-tailored attack designed to evade Scope, but it remains ineffective even when maximizing stealthiness, further underscoring the robustness of Scope. We release our source code at:https://github.com/siquanhuang/Scope.
Siquan Huang, Yijiang Li, Xingfu Yan, Ying Gao 0004, Chong Chen 0011, Leyu Shi, Wing W. Y. Ng
IEEE Trans. Inf. Forensics Secur.4
2025 Enhancing Weakly Supervised Semantic Segmentation With Multi-Label Contrastive Learning and LLM Features Guidance
abstract
Histopathological whole-slide images (WSIs) segmentation is essential for precise tissue characterization in medical diagnostics. However, traditional approaches require labor-intensive pixel-level annotations. To this end, we study weakly supervised semantic segmentation (WSSS) which uses patch-level classification labels, reducing annotation efforts significantly. However, the complexity of WSIs and the challenge of sparse classification labels hinder effective dense pixel predictions. Moreover, due to the multi-label nature of WSI, existing approaches of single-label contrastive learning designed for the representation of single-category, neglecting the presence of other relevant categories and thus fail to adapt to WSI tasks. This paper presents a novel multi-label contrastive learning method for WSSS by incorporating class-specific embedding extraction with LLM features guidance. Specifically, we propose to obtain class-specific embeddings by utilizing classifier weights, followed by a dot-product-based attention fusion method that leverages LLM features to enrich their semantics, facilitating contrastive learning between different classes from single image. Besides, we propose a Robust Learning approach that leverages multi-layer features to evaluate the uncertainty of pseudo-labels, thereby mitigating the impact of noisy pseudo-labels on the learning process of segmentation. Extensive experiments have been conducted on two histopathological image segmentation datasets, i.e. LUAD dataset and BCSS dataset, demonstrating the effectiveness of our methods with leading performance.
Wentian Cai, Yijiang Li, Yandan Chen, G. Thippa Reddy, Wei Wang 0077, Ying Gao 0004
IEEE J. Biomed. Health Informatics9
2024 Fastmandarin: Efficient Local Modeling for Natural Mandarin Speech Synthesis
abstract
Attention-based speech synthesis methods often suffer from dispersed attention across the entire input sequence, resulting in poor local modeling and unnatural Mandarin synthesized speech. To address these issues, we present FastMandarin, a rapid and natural Mandarin speech synthesis framework that employs two explicit methods to enhance local modeling and improve pronunciation representation. Firstly, we tag Chinese characters to delineate phrase boundaries within a sentence, and these tags are integrated into the network’s hidden layer features at each time step, effectively bolstering local contributions in latent representations. Secondly, we introduce a multi-scale context feature extractor network that employs parallel convolution with various filters. Additionally, we optimize duration alignment and Mel-spectrogram reconstruction to enhance overall performance. Experimental results demonstrate that FastMandarin excels in local modeling, delivering robust Mandarin speech synthesis results.
Chenglong Jiang, Ying Gao 0004, Linrong Pan, Wing W. Y. Ng
ICASSP2
2024 GN: Guided Noise Eliminating Backdoors in Federated Learning
abstract
Federated learning (FL) trains a model collaboratively but is susceptible to backdoor attacks for its privacy-preserving nature. Existing defenses against backdoor attacks in FL always make specific assumptions on data distributions among clients and are ineffective against sophisticated attacks. Although adding noise mitigates backdoors injected in the model, it simultaneously negatively impacts the main performance. To address the aforementioned issues, we propose a novel defense mechanism, Guided Noise (GN), that eliminates backdoors without compromising the model's main performance. GN achieves this by utilizing conductance to evaluate the importance of neurons and subsequently adding guided noise to suspected backdoor neurons selected by voting, which only disturbs the backdoor task. Extensive experimental evaluations of GN show its significant superiority over traditional noising-based defenses, making it a valuable replacement for existing noising to enhance the robustness of existing defenses against backdoor attacks in FL.
Siquan Huang, Ying Gao 0004, Chong Chen 0011, Leyu Shi
SMC2
2024 ASE: Anomaly scoring based ensemble learning for highly imbalanced datasets
Xiayu Liang, Ying Gao 0004, Shanrong Xu
Expert Syst. Appl.2
2024 Semantic dependency and local convolution for enhancing naturalness and tone in text-to-speech synthesis
Chenglong Jiang, Ying Gao 0004, Wing W. Y. Ng, Jiyong Zhou, Jinghui Zhong, Hongzhong Zhen, Xiping Hu
Neurocomputing2
2024 Enhanced Attention Guided Teacher-Student Network for Weakly Supervised Object Detection
Ying Gao 0004, Wentian Cai, Weixian Yang, Xiping Hu, Victor C. M. Leung
Neurocomputing2
2024 Fog-Enabled Privacy-Preserving Multi-Task Data Aggregation for Mobile Crowdsensing
abstract
Privacy-preserving data aggregation in mobile crowdsensing (MCS) focuses on mining information from massive sensing data while protecting users' privacy. The existence of multiple concurrent tasks is common in urban environments, so privacy-preserving multi-task data aggregation is essential and useful to a large-scale crowdsensing server. However, existing privacy-preserving data aggregation schemes in MCS mainly focus on the single-task data aggregation and the privacy protection of user's data. Little attention is paid to the privacy of user's decision of accepting tasks. Therefore, we propose a privacy-preserving and server-oriented efficient multi-task data aggregation scheme for MCS based fog computing. The proposed scheme can aggregate multiple concurrent tasks from multiple requesters (e.g., for 9 tasks, the proposed scheme completes all tasks in one round as opposed to existing schemes, which finish 9 tasks in nine rounds). Our scheme protects the privacy of user's decision, user's data, and aggregation result of each requester under collusion attacks. Through formal security analyses, our scheme is proved to be secure and privacy-preserving. Both theoretical analyses and experiments show our scheme is efficient.
Xingfu Yan, Wing W. Y. Ng, Bowen Zhao 0001, Yuxian Liu, Ying Gao 0004, Xiumin Wang 0005
IEEE Trans. Dependable Secur. Comput.5
2024 DAST: A Domain-Adaptive Learning Combining Spatio-Temporal Dynamic Attention for Electroencephalography Emotion Recognition
abstract
Multimodal emotion recognition with EEG-based have become mainstream in affective computing. However, previous studies mainly focus on perceived emotions (including posture, speech or face expression et al.) of different subjects, while the lack of research on induced emotions (including video or music et al.) limited the development of two-ways emotions. To solve this problem, we propose a multimodal domain adaptive method based on EEG and music called the DAST, which uses spatio-temporal adaptive attention (STA-attention) to globally model the EEG and maps all embeddings dynamically into high-dimensionally space by adaptive space encoder (ASE). Then, adversarial training is performed with domain discriminator and ASE to learn invariant emotion representations. Furthermore, we conduct extensive experiments on the DEAP dataset, and the results show that our method can further explore the relationship between induced and perceived emotions, and provide a reliable reference for exploring the potential correlation between EEG and music stimulation.
Ying Gao 0004, Tingting Wang 0006
IEEE J. Biomed. Health Informatics2
2024 Explosive Cyber Security Threats During COVID-19 Pandemic and a Novel Tree-Based Broad Learning System to Overcome
abstract
The rapid spread of the COVID-19 has not only affected personal health and economy, but also revolutionized people’s lifestyles. As more people turn to work and socialize online, the development of unmanned technologies based on the Internet of Vehicles (IoV), such as unmanned delivery, unmanned vehicles, unmanned transportation, etc., will become an inevitable trend. However, all kinds of intelligent terminals for unmanned equipment require a large amount of data interaction with devices such as cloud servers, mobile terminals, and roadside terminals, which poses cyber security risks. Furthermore, the outbreak of COVID-19 has prompted people to put forward higher demands for the security of network communications. Unfortunately, the current intrusion detection methods based on machine learning still have weaknesses such as low accuracy and low efficiency when faced with unbalanced data distribution. To solve the above problems, we propose a novel Tree-based BLS (TBLS) intrusion detection method according to the idea of ensemble learning and decision tree (CART and J48). The performance of TBLS was tested on the NSL-KDD dataset and the UNSW-NB15 dataset respectively, which contain a variety of malicious traffic types for attacks on the IoV. The results show that our proposed method can achieve higher accuracy rate and lower false alarm rate, compared with the existing 16 solutions.
Ying Gao 0004, Hongyue Miao, Binjie Song, Xiping Hu, Wei Wang 0077
IEEE Trans. Intell. Transp. Syst.1
2023 SeDepTTS: Enhancing the Naturalness via Semantic Dependency and Local Convolution for Text-to-Speech Synthesis
abstract
Self-attention-based networks have obtained impressive performance in parallel training and global context modeling. However, it is weak in local dependency capturing, especially for data with strong local correlations such as utterances. Therefore, we will mine linguistic information of the original text based on a semantic dependency and the semantic relationship between nodes is regarded as prior knowledge to revise the distribution of self-attention. On the other hand, given the strong correlation between input characters, we introduce a one-dimensional (1-D) convolution neural network (CNN) producing query(Q) and value(V) in the self-attention mechanism for a better fusion of local contextual information. Then, we migrate this variant of the self-attention networks to speech synthesis tasks and propose a non-autoregressive (NAR) neural Text-to-Speech (TTS): SeDepTTS. Experimental results show that our model yields good performance in speech synthesis. Specifically, the proposed method yields significant improvement for the processing of pause, stress, and intonation in speech.
Chenglong Jiang, Ying Gao 0004, Wing W. Y. Ng, Jiyong Zhou, Jinghui Zhong, Hongzhong Zhen
AAAI2
2023 Multi-metrics adaptively identifies backdoors in Federated learning
abstract
The decentralized and privacy-preserving nature of federated learning (FL) makes it vulnerable to backdoor attacks aiming to manipulate the behavior of the resulting model on specific adversary-chosen inputs. However, most existing defenses based on statistical differences take effect only against specific attacks, especially when the malicious gradients are similar to benign ones or the data are highly non-independent and identically distributed (non-IID). In this paper, we revisit the distance-based defense methods and discover that i) Euclidean distance becomes meaningless in high dimensions and ii) malicious gradients with diverse characteristics cannot be identified by a single metric. To this end, we present a simple yet effective defense strategy with multi-metrics and dynamic weighting to identify backdoors adaptively. Furthermore, our novel defense has no reliance on predefined assumptions over attack settings or data distributions and little impact on benign performance. To evaluate the effectiveness of our approach, we conduct comprehensive experiments on different datasets under various attack settings, where our method achieves the best defensive performance. For instance, we achieve the lowest backdoor accuracy of 3.06% under the most difficult Edge-case PGD, showing significant superiority over previous defenses. The experiments also demonstrate that our method can be well-adapted to a wide range of non-IID degrees without sacrificing the benign performance.
Siquan Huang, Yijiang Li, Chong Chen 0011, Leyu Shi, Ying Gao 0004
ICCV5
2023 Diverse Cotraining Makes Strong Semi-Supervised Segmentor
abstract
Deep co-training has been introduced to semi-supervised segmentation and achieves impressive results, yet few studies have explored the working mechanism behind it. In this work, we revisit the core assumption that supports co-training: multiple compatible and conditionally independent views. By theoretically deriving the generalization upper bound, we prove the prediction similarity between two models negatively impacts the model’s generalization ability. However, most current co-training models are tightly coupled together and violate this assumption. Such coupling leads to the homogenization of networks and confirmation bias which consequently limits the performance. To this end, we explore different dimensions of co-training and systematically increase the diversity from the aspects of input domains, different augmentations and model architectures to counteract homogenization. Our Diverse Co-training outperforms the state-of-the-art (SOTA) methods by a large margin across different evaluation protocols on the Pascal and Cityscapes. For example, we achieve the best mIoU of 76.2%, 77.7% and 80.2% on Pascal with only 92, 183 and 366 labeled images, surpassing the previous best results by more than 5%.
Yijiang Li, Xinjiang Wang, Lihe Yang, Litong Feng, Wayne Zhang 0001, Ying Gao 0004
ICCV6
2023 A Dual-Path Supplemental Information Learning Architecture for Breast Cancer Ki-67 Status Prediction in T2w MRI
abstract
In this paper, we propose a Dual-path Supplemental Information Learning Architecture (DSILA) for predicting breast cancer Ki-67 status based on T2-weighted (T2w) magnetic resonance imaging (MRI). DSILA consists of two components: 1) a transfer network with multi-scale feature selection strategy to obtain generic multi-scale features most relative to target, 2) a supplemental learning network with a large receptive field and channel-level attention to mine scenario-related semantic information. A regulation item – Aspect Overlap Loss (AOL), is further added to force the supplemental learning network to pay more attention to the regions overlooked by the transfer network. The experimental results tested on the collected T2w MRI breast cancer Ki-67 dataset show that DSILA outperforms state-of-the-art techniques among all adopted evaluation metrics, even achieving 0.85 in Area under the Receiver Operating Characteristic Curve (AUC).
Wentian Cai, Yulin Cheng, Ying Gao 0004, Weixiao Liu, Xinyan Xie, Xiong-Wen Luo 0001, Weixian Yang, Zaiyi Liu, Changhong Liang
ICME3
2023 A Lightweight and Efficient Model for Audio Anti-Spoofing
abstract
With the rapid development of speech conversion and speech synthesis algorithms, automatic speaker verification (ASV) systems are vulnerable to spoofing attacks. In recent years, researchers had proposed anti-spoofing systems based on hand-crafted features. However, using hand-crafted features rather than raw waveform will lose implicit information for audio anti-spoofing. Inspired by the promising performance of ConvNeXt in classification tasks, we reference the network architecture design of ConvNeXt and propose a Lightweight and Efficient Model for Audio Anti-Spoofing (LEMAAS). With no preceding feature extraction process, we employ raw waveforms as direct inputs to our proposed model. By integrating with the channel attention module and using the focal loss function, the proposed model can focus on the most informative features representation of speech and the difficult samples that are hard to classify. Experimental results show that our proposed system could achieve an equal error rate of 0.64% and min-tDCF of 0.0187 for the ASVspoof 2019 LA evaluation dataset, which outperforms the state-of-the-art systems. Moreover, even when trained only on the ASVspoof 2019 LA dataset, the model still achieved equal error rates of 0.86% and 1.18% on the ASVspoof 2015 development dataset and evaluation dataset, respectively. This demonstrates that our model has achieved promising generalization performance during cross-dataset testing.
Qiaowei Ma, Jinghui Zhong, Weiheng Liu, Ying Gao 0004, Wing W. Y. Ng
MMAsia5
2023 UKT: A Unified Knowledgeable Tuning Framework for Chinese Information Extraction
Jiyong Zhou, Chengyu Wang 0001, Jianing Wang 0002, Yukang Xie, Jun Huang 0007, Ying Gao 0004
NLPCC (2)7
2023 Utilizing Video Word Boundaries and Feature-Based Knowledge Distillation Improving Sentence-Level Lip Reading
Hongzhong Zhen, Chenglong Jiang, Jiyong Zhou, Ying Gao 0004
PRCV (6)5
2023 CATE: Contrastive augmentation and tree-enhanced embedding for credit scoring
abstract
Credit transactions are vital financial activities that yield substantial economic benefits. To further improve lending decisions, stakeholders require accurate and interpretable credit scoring methods. While the majority of previous studies have focused on the relationship between individual features and credit risk, only a few have investigated cross-features. Notably, cross-features can not only represent structured data effectively but also provide richer semantic information than individual features. Nevertheless, most previous methods for learning cross-feature effects from credit data have been implicit and unexplainable. This paper proposes a new credit scoring model based on contrastive augmentation and tree-enhanced embedding mechanisms, termed CATE. The proposed model automatically constructs explainable cross-features by using tree-based models to learn decision rules from the data. Moreover, the importance of each local cross-feature is then derived through an attention mechanism . Finally, the credit score of a user is evaluated using embedding vectors. Experimental results on 4 public datasets demonstrated the interpretability of our proposed method and outperformed 13 state-of-the-art benchmark methods in terms of performance.
Ying Gao 0004, Haolang Xiao, Choujun Zhan, Lingrui Liang, Wentian Cai, Xiping Hu
Inf. Sci.1
2023 Human migration-based graph convolutional network for PM2.5 forecasting in post-COVID-19 pandemic age
Choujun Zhan, Wei Jiang 0006, Hu Min, Ying Gao 0004, C. K. Michael Tse
Neural Comput. Appl.4
2023 DFTNet: Dual-Path Feature Transfer Network for Weakly Supervised Medical Image Segmentation
abstract
Medical image segmentation has long suffered from the problem of expensive labels. Acquiring pixel-level annotations is time-consuming, labor-intensive, and relies on extensive expert knowledge. Bounding box annotations, in contrast, are relatively easy to acquire. Thus, in this paper, we explore to segment images through a novel Dual-path Feature Transfer design with only bounding box annotations. Specifically, a Target-aware Reconstructor is proposed to extract target-related features by reconstructing the pixels within the bounding box through the channel and spatial attention module. Then, a sliding Feature Fusion and Transfer Module (FFTM) fuses the extracted features from Reconstructor and transfers them to guide the Segmentor for segmentation. Finally, we present the Confidence Ranking Loss (CRLoss) which dynamically assigns weights to the loss of each pixel based on the network's confidence. CRLoss mitigates the impact of inaccurate pseudo-labels on performance. Extensive experiments demonstrate that our proposed model achieves state-of-the-art performance on the Medical Segmentation Decathlon (MSD) Brain Tumour and PROMISE12 datasets.
Wentian Cai, Linsen Xie, Weixian Yang, Yijiang Li, Ying Gao 0004, Tingting Wang 0006
IEEE ACM Trans. Comput. Biol. Bioinform.5
2023 Knowledge Distillation Hashing for Occluded Face Retrieval
abstract
Deep hashing has proven to be efficient and effective for large-scale face retrieval. However, existing hashing methods are designed for normal face images only. They fail to consider the fact that face images may be occluded because of wearing masks, hats, glasses, etc. Retrieval performance of existing face retrieval methods is much worse when dealing with occluded face images. In this work, we propose the knowledge distillation hashing (KDH) to deal with occluded face images. The KDH is a two-stage learning approach with teacher-student model distillation. We first train a teacher hashing network using normal face images and then the knowledge from teacher model is used to guide the optimization of the student model using occluded face images as input only. With knowledge distillation, we build a connection between imperfect face information and the optimal hash codes. Experimental results show that the KDH yields significant improvements and better retrieval performance in comparison to existing state-of-the-art deep hashing retrieval methods under six different face occlusion situations.
Xing Tian, Wing W. Y. Ng, Ying Gao 0004
IEEE Trans. Multim.4
2022 More than Encoder: Introducing Transformer Decoder to Upsample
abstract
Medical image segmentation methods downsample images for feature extraction and then upsample them to restore resolution for pixel-level predictions. In such schema, upsample technique is vital in restoring information for better performance. However, existing upsample techniques leverage little information from downsampling paths. The local and detailed feature from the shallower layer such as boundary and tissue texture is crucial in segmentation, especially medical image segmentation. To this end, we propose a novel upsample approach for medical image segmentation, Window Attention Upsample (WAU), which upsamples features conditioned on local and detailed features from downsampling path in local windows by introducing attention decoders of Transformer. WAU could serve as a general upsample method and be incorporated into any segmentation model that possesses lateral connections. We first propose the Attention Upsample which consists of Attention Decoder (AD) and bilinear upsample. AD leverages pixel-level attention to model longrange dependency and global information for a better upsample. Bilinear upsample is introduced as the residual connection to complement the upsampled features. Moreover, considering the extensive memory and computation cost of pixel-level attention, we further design a window attention scheme to restrict attention computation in local windows instead of the global range. We evaluate our method (WAU) on classic UNet structure with lateral connections and achieve state-of-the-art performance on Medical Segmentation Decathlon (MSD) Brain and Automatic Cardiac Diagnosis Challenge (ACDC) datasets. We also validate the effectiveness of our method on multiple classic architectures and achieve consistent improvement.
Yijiang Li, Wentian Cai, Ying Gao 0004, Chengming Li 0004, Xiping Hu
BIBM3
2022 Neural network-based event-triggered integral reinforcement learning for constrained H∞ tracking control with experience replay
Shan Xue 0004, Biao Luo 0001, Derong Liu 0001, Ying Gao 0004
Neurocomputing4
2022 P2SIM: Privacy-Preserving and Source-Reliable Incentive Mechanism for Mobile Crowdsensing
abstract
In mobile crowdsensing (MCS), providing appropriate rewards is a common and efficient way to motivate participants to participate in sensing tasks. However, the privacy of task participants is not protected well in most quality-aware incentive schemes. Moreover, these schemes are designed for general MCS application scenarios where data are collected by internal sensors embedded in participants’ smartphones, and not suitable for scenarios where additional sensors (ASs) except internal sensors are to collect data (e.g., household medical devices). In scenarios with ASs, malicious participants can fabricate sensing data instead of collecting data from ASs, i.e., the source reliability of sensing data cannot be ensured. To address these issues, we propose P2SIM, a privacy-preserving and the source-reliable incentive mechanism scheme for MCS with ASs. We combine redactable signature with private hash function to achieve the source reliability verification of sensing data without revealing the privacy of participants. Moreover, rewards are divided into two parts: 1) fixed rewards and 2) floating rewards, to enhance the flexibility of rewards distribution. Both formal theoretical analysis and extensive experimental evaluations on a real data set show that the proposed P2SIM is secure and efficient.
Xingfu Yan, Wing W. Y. Ng, Bowen Zhao 0001, Ying Gao 0004
IEEE Internet Things J.6
2022 Estimating unconfirmed COVID-19 infection cases and multiple waves of pandemic progression with consideration of testing capacity and non-pharmaceutical interventions: A dynamic spreading model
Choujun Zhan, Lujiao Shao, Ziliang Yin, Ying Gao 0004, C. K. Michael Tse, Di Wu 0035, Haijun Zhang 0002
Inf. Sci.5
2022 An Improved Selection Method Based on Crowded Comparison for Multi-Objective Optimization Problems in Intelligent Computing
Ying Gao 0004, Binjie Song, Xiping Hu, Yekui Qian
Mob. Networks Appl.1
2022 Event-triggered integral reinforcement learning for nonzero-sum games with asymmetric input saturation
Shan Xue 0004, Biao Luo 0001, Derong Liu 0001, Ying Gao 0004
Neural Networks4
2022 Self-Learning Spatial Distribution-Based Intrusion Detection for Industrial Cyber-Physical Systems
abstract
Thanks to the great advancement of cognitive computing, artificial intelligence, big data, and the Internet of Things (IoT) technologies, the fusion of the physical and virtual worlds is changing people’s lifestyles. Although the research and deployment of cyber-physical systems (CPSs) are notably promoted by cognitive computing, the reliability and large-scale application of CPSs are still significantly challenged by some security issues. Therefore, it is meaningful to clarify and address the weaknesses of current intrusion detection methods for CPSs and enhance the ability to identify, analyze, and predict to improve the performance of intrusion detection. In this article, we first propose a novel self-learning spatial distribution algorithm, named Euclidean distance-based between-class learning (EBC learning), which improves between-class learning by calculating the Euclidean distance (ED) among$k$-nearest neighbors of different classes. In addition, a cognitive computing-based intrusion detection method named border-line SMOTE and EBC learning based on random forest (BSBC-RF) is also proposed based on the EBC learning for industrial CPSs. The experimental results over a real industrial traffic dataset show that the proposed EBC learning has strong spatial constraint capability and can improve the prediction and recognition performance. Compared with the eight state-of-the-art methods, the proposed method has an ACC exceeding 99.5%, false alarm rate (FAR) less than 0.06%, and$F1$close to 0.99, which is still superior to other ones.
Ying Gao 0004, Hongyue Miao, Binjie Song, Yiqin Lu, Weiqiang Pan
IEEE Trans. Comput. Soc. Syst.1
2022 Event-Triggered ADP for Tracking Control of Partially Unknown Constrained Uncertain Systems
abstract
An event-triggered adaptive dynamic programming (ADP) algorithm is developed in this article to solve the tracking control problem for partially unknown constrained uncertain systems. First, an augmented system is constructed, and the solution of the optimal tracking control problem of the uncertain system is transformed into an optimal regulation of the nominal augmented system with a discounted value function. The integral reinforcement learning is employed to avoid the requirement of augmented drift dynamics. Second, the event-triggered ADP is adopted for its implementation, where the learning of neural network weights not only relaxes the initial admissible control but also executes only when the predefined execution rule is violated. Third, the tracking error and the weight estimation error prove to be uniformly ultimately bounded, and the existence of a lower bound for the interexecution times is analyzed. Finally, simulation results demonstrate the effectiveness of the present event-triggered ADP method.
Shan Xue 0004, Biao Luo 0001, Derong Liu 0001, Ying Gao 0004
IEEE Trans. Cybern.4
2022 DMCGNet: A Novel Network for Medical Image Segmentation With Dense Self-Mimic and Channel Grouping Mechanism
abstract
Automatic Medical Image Segmentation (MIS) can assist doctors by reducing labor and providing a unified standard. Nowadays, approaches based on Deep Learning have become mainstream for MIS because of their ability of automatic feature extraction. However, due to the plain network design and targets variety in medical images, the semantic features can hardly be extracted adequately. In this work, we propose a novel Dense Self-Mimic and Channel Grouping based Network (DMCGNet) for MIS for better feature extraction. Specifically, we introduce a Pyramid Target-aware Dense Self Mimic (PTDSM) module, which is capable of exploring deeper and better feature representation with no parameter increase. Then, to utilize features efficiently, an effective Channel Split based Feature Fusion Module (CSFFM) is proposed for feature reuse, which strengthens the adaptation of multi-scale targets by utilizing the channel grouping mechanism. Finally, to train the proposed method adequately, Deep Supervision with Group Ensemble Learning (DSGEL) is equipped to the network. Extensive experiments demonstrate that our proposed model achieves state-of-the-art performance on 4 medical image segmentation datasets.
Linsen Xie, Wentian Cai, Ying Gao 0004
IEEE J. Biomed. Health Informatics3
2021 Blockchain and SGX-Enabled Edge-Computing-Empowered Secure IoMT Data Analysis
abstract
The Internet of Medical Things (IoMT) is an important application of the Internet of Things (IoT) in the health field, including remote health monitoring and remote medical diagnosis. This not only brings convenience to the patient but also reduces the cost of the patient. However, the surge of data brought by mobile health monitoring equipment challenges the traditional centralized data processing model. In particular, medical data are closely related to patient privacy. Therefore, only part of the specific medical data should be provided to the medical institutions in need, rather than all the data, to ensure the confidentiality of the data to the greatest extent. But curious data processing centers can easily lead to data leakage. To tackle these challenges, we use edge computing and blockchain to build a new framework. In particular, the trusted execution environment, namely, software guard extension (SGX) technology, is introduced into edge computing to ensure the confidentiality of the data analysis process. The blockchain authenticates the IoMT devices and cloud service providers that are added to the network and provides an access policy management mechanism for IoMT data. Moreover, a prototype of the proposed framework is implemented using Hyperledger Fabric and Intel SGX, and the analysis of the blockchain and SGX performance are also presented.
Ying Gao 0004, Hongliang Lin, Yijian Chen, Yangliang Liu
IEEE Internet Things J.1
2021 Verifiable, Reliable, and Privacy-Preserving Data Aggregation in Fog-Assisted Mobile Crowdsensing
abstract
Fog-assisted mobile crowdsensing (FA-MCS) alleviates challenges with respect to computation, communication, and storage from the traditional model of mobile crowdsensing (MCS) “requester-server-users.” Data aggregation, as a specific MCS task, has attracted a lot of attentions in mining the potential value of the massive crowdsensing data. However, the process of data aggregation in FA-MCS may threaten the privacies of both users' data and aggregation results. The untrusted server and fog nodes (FNs) may damage the correctness of aggregation results. Moreover, bad FNs, which do not upload data to server or fail to verify successfully, can endanger the reliability of FA-MCS and the accuracy of aggregation results. To tackle these problems, we propose a verifiable, reliable, and privacy-preserving data aggregation scheme for FA-MCS. Specifically, the proposed scheme preserves privacies of both users' data and aggregation results, enables requester to verify the correctness of aggregation result, and is able to tolerate several bad FNs without affecting the data aggregation result. Through formal security analysis, the proposed scheme is shown to be secure and privacy preserving. Extensive experiments also show the proposed scheme is efficient and reliable.
Xingfu Yan, Wing W. Y. Ng, Changlu Lin, Yuxian Liu, Lu Lu 0011, Ying Gao 0004
IEEE Internet Things J.7
2021 Blockchain Based IIoT Data Sharing Framework for SDN-Enabled Pervasive Edge Computing
abstract
Pervasive edge computing (PEC) is an emerging paradigm for the industrial Internet of Things (IIoT), and software-defined networks (SDN) offer lower latency services, and massive intelligent devices connectivity for the IIoT. However, the PEC has some issues with data security, and privacy while PEC devices sharing data among edges. What's more, the centralized SDN suffers from single point of attacks such as distributed denial of service (DDoS) from IIoT devices, and has the challenge of data leakage. In this article, we use blockchain, and proxy reencryption (PRE) technologies to tackle these challenges. The blockchain authorizes all devices in the network to improve their credibility, and authenticity. In addition, a blockchain-based data sharing framework that combines a PRE scheme is introduced for secure device-to-device communication in PEC environments. A series of smart contracts are designed for flexible operations of searching, and updating records on the blockchain. The experiments reveal that our design is highly efficient, and has high performance.
Ying Gao 0004, Yijian Chen, Xiping Hu, Hongliang Lin, Yangliang Liu, Laisen Nie
IEEE Trans. Ind. Informatics1
2021 A Real-Time Defect Detection Method for Digital Signal Processing of Industrial Inspection Applications
abstract
The signal processing of industrial big data (IBD) is a challenging task, owing to the complex working scenarios and the lack of annotations. Defect detection, which is an important subject of IBD research works, has shown its effectiveness in digital signal processing of industrial inspection applications in many previous studies. This article proposes a novel defect detection method based on deep learning for digital signal processing of industrial inspection applications. In our method, a module named feature collection and compression network is applied to merge multiscale feature information. Then, a new pooling method named Gaussian weighted pooling, which provides more precise location information, is used to replace region of interest (ROI) pooling. Experiment results show that our method gets improvements in both accuracy and efficiency, with mAP/AP50 of 41.8/80.2 at 33 fps on NEUDET, which satisfies the requirement of real-time systems.
Ying Gao 0004, Jiqiang Lin, Zhaolong Ning
IEEE Trans. Ind. Informatics1
2021 Comparative Study of COVID-19 Pandemic Progressions in 175 Regions in Australia, Canada, Italy, Japan, Spain, U.K. and USA Using a Novel Model That Considers Testing Capacity and Deficiency in Confirming Infected Cases
abstract
Not identified as being exposed or infected, the group of asymptomatic and presymptomatic patients has become the key source of infectious hosts for the COVID-19 pandemic, triggering the re-emergence of outbreaks. Acknowledging the impacts of movement of unidentified patients and the limited testing capacity on understanding the spread of the virus, an augmented Susceptible-Exposed-Infectious-Confirmed-Recovered (SEICR) model integrating intercity migration data and testing capacity is developed to probe into the number of unidentified COVID-19 infected patients. This model allows evaluation of the effectiveness of active interventions, and more accurate prediction of the pandemic progression in a country, region or city. A pseudo-coevolutionary algorithm is adopted in the model fitting to provide an effective estimation of high-dimensional unknown parameter sets using a limited amount of historical data. The model is applied to 175 regions in Australia, Canada, Italy, Japan, Spain, the UK and USA to estimate the number of unconfirmed cases using limited historical data. Results showed that the actual number of infected cases could be 4.309 times as many as the official confirmed number. By implementing mass COVID-19 testing, the number of infected cases could be reduced by about 50%.
Choujun Zhan, C. K. Michael Tse, Ying Gao 0004, Tianyong Hao
IEEE J. Biomed. Health Informatics3
2021 An Intelligent Cloud Workflow Scheduling System With Time Estimation and Adaptive Ant Colony Optimization
abstract
The introduction of workflow in cloud computing has afforded a new and efficient way to tackle large-scale applications. As an NP-hard problem, how to schedule cloud workflows effectively and economically with deadline constraints and different kinds of tasks and resources is extraordinarily challenging. To solve this constrained problem, this paper intends to develop an intelligent scheduling system from the perspective of users to reduce expenditure of workflow, subject to the deadline and other execution constraints. A new estimation model of the task execution time is designed according to virtual machine settings in real public clouds and execution data from practical workflows. Based on the new model, an adaptive ant colony optimization algorithm is proposed to meet the quality of service and orchestrate tasks. The adaptiveness of the algorithm is embodied in two aspects. First, an adaptive solution construction method is designed that each solution is built with a dynamically changing resource pool, thus the search space of the algorithm is narrowed down and the execution time is decreased. Second, two heuristics with self-adaptive weight are introduced to adaptively meet different deadline settings. Simulating results on four types of workflows show that the proposed approach is effective and competitive.
Ya-Hui Jia, Weineng Chen, Huaqiang Yuan, Tianlong Gu, Huaxiang Zhang 0001, Ying Gao 0004, Jun Zhang 0003
IEEE Trans. Syst. Man Cybern. Syst.6
2020 A Diversified Supervised based U-shape Colorectal Lesion Segmentor with Meaningful Feature Supplement and Multi-Level Residual Attention Mechanism
abstract
Colorectal cancer is a commonly diagnosed cancer of digestive system. Automatic and accurate segmentation of colorectal tumors from medical images (e.g., CT) has great significance for diagnosis, staging and treatment planning. However, the blurred boundary of tumors, as well as variability of their location and shape, make most traditional methods ineffectual. In this paper, we propose a diversified supervised U-shape CNN colorectal lesion segmentor (DSUCLS) to overcome this challenge. Our model mainly contains three key components: 1) the weakly supervised transfer learning module for supplementing generic features, where the irrelevant ones are filtered out by extra convolutional layers and image-level label, 2) an encoder-decoder structure based on U-shape architecture for learning specific pathological representation from medical images, 3) the multilevel supervised attention module incorporated into decoder path for producing coarse-to-fine guidance and guaranteeing finer attention map. 4), the pre-processing and post-processing strategies are applied to further improve segmentation performance. The experimental results illustrate that the proposed model outperforms other state-of-the-art techniques for colorectal lesion segmentation on CT images, achieving Dice scores of 0.733 and dramatically decreasing Hausdorff distance to 17.62.
Jinjie Wang, Xiong-Wen Luo 0001, Linsen Xie, Ying Gao 0004
BIBM4
2020 Optimizing Broad Learning System Hyper-parameters through Particle Swarm Optimization for Predicting COVID-19 in 184 Countries
abstract
The Coronavirus Disease 2019 (COVID-19) began to outbreak since December 2019 and widely spread over the world. How to accurately predict the spread of COVID-19 is one of the essential issues for controlling the pandemic. This study establishes a general model that can predict the trend of COVID-19 in a country based on historical COVID-19 data in 184 countries. First, Savitzky-Golay (S-G) filter is utilized to detect multiple waves of COVID-19 in a country. Then, a PSO-SIR (particle swarm optimization susceptible-infected-recovery) model is provided for data augmentation. Finally, a novel PSO-BLS (particle swarm optimization broad learning system) is proposed for predicting the trend of COVID-19. Experimental results show that compared with the deep learning models (ANN, CNN, LSTM, and GRU), the PSO-BLS algorithm has higher accuracy and stability in predicting the number of active infected cases and removed cases.
Choujun Zhan, Zhengdong Wu, Quansi Wen, Ying Gao 0004, Haijun Zhang 0002
HealthCom4
2020 Parameter-Free Voronoi Neighborhood for Evolutionary Multimodal Optimization
abstract
Neighborhood information plays an important role in improving the performance of evolutionary computation in various optimization scenarios, particularly in the context of multimodal optimization. Several neighborhood concepts, i.e., index-based neighborhood, nearest neighborhood, and fuzzy neighborhood, have been studied and engaged in the design of niching methods. However, the use of these neighborhood concepts requires the specification of some problem-related parameters, which is difficult to determine without a prior knowledge. In this paper, we introduce a new neighborhood concept based on a geometrical construction called Voronoi diagram. The new concept offers two advantages at the expense of increasing the computational complexity to a higher level. It eliminates the need of additional parameters and it is more informative than the existing ones. The information provided by the Voronoi neighbors of an individual can be exploited to estimate the evolutionary state. Based on the information, we divide the population into three groups and assign each group a different reproduction strategy to support the exploration and exploitation of the search space. We show the use of the concept in the design of an effective evolutionary algorithm for multimodal optimization. The experiments have been conducted to investigate the performance of the algorithm. The results reveal that the proposed algorithm compare favorably with the state-of-the-art algorithms designed based on other types of neighborhood concepts.
Yuhui Zhang 0004, Yue-Jiao Gong, Ying Gao 0004, Hua Wang 0002, Jun Zhang 0003
IEEE Trans. Evol. Comput.3
2020 Modeling Human Activity With Seasonality Bursty Dynamics
abstract
The public's purchase incentive increases dramatically during the holiday season and subsequently returns to normal levels. This seasonality is common in various scenarios and highlights the following questions: how does the public's purchase incentive fluctuate over the course of a year? Which factors are conducive to this seasonal behavior and how can they be modeled? In this paper, we propose a model that explicitly integrates temporal point process theory with the construction of a networked community, to describe the dynamics of collective action propagation with seasonal fluctuation. Furthermore, a database is constructed of sales records for 21 video game consoles and 13 237 video games in France, Germany, Japan, the U.K., the USA, and worldwide from 1989 to 2018. Experimental results suggest that peak desire always appears in the holiday season about one week before Christmas and is about four times higher than consumption desire in a normal period in all areas.
Quansi Wen, Choujun Zhan, Ying Gao 0004, Xiping Hu, Edith C. H. Ngai, Bin Hu 0001
IEEE Trans. Ind. Informatics3
2019 Dense Encoder-Decoder Network based on Two-Level Context Enhanced Residual Attention Mechanism for Segmentation of Breast Tumors in Magnetic Resonance Imaging
abstract
Aiming to effective early detection of breast cancer, automatic tumor segmentation based on breast Magnetic Resonance Imaging (MRI) is concentrated by more and more researchers. This paper proposes a dense encoder-decoder network based on two-level context enhanced residual attention mechanism (TLCRAM-DED). With respect to TLCRAM-DED, we design the encoding structure combining two-level residual attention structure with dense block to extract and refine the features of different layers. Meanwhile, a dense multi-scale atrous convolution is used at the end of the encoder to obtain a larger receptive field and enrich the extracted semantic information. Moreover, residual attention structure (RAS) is also used for the refinement during decoding stage, while a long connection formed with the encoder RAS output is applied to supplement the features and to gradually recover the segmentation details. We validated prosed model in the DCE sequence of challenging breast cancer MRI dataset. The average Dice coefficient is up to 81.04%, which outperforms compared state-of-the-arts.
Ying Gao 0004, Yin Zhao, Xiong-Wen Luo 0001, Xiping Hu, Changhong Liang
BIBM1
2019 Indefinite Kernels in One-Class Support Vector Machine and its Application on Virtual Screening
abstract
Imbalanced dataset is a common issue in many applications. The one-class Support Vector Machine (SVM) is found to be an effective algorithm to construct classification models over the underlying imbalanced dataset. In some cases, feature extraction is hard and one would prefer using pre-defined kernels to train the model. In traditional practice, a valid kernel has to satisfy the Mercer's condition, which may restrict the design of kernel functions or matrices. In this paper, an indefinite kernel extension is applied to the one-class SVM model in order to relieve such limitation. To illustrate its performance, the algorithm is applied to perform virtual screening of drugs.
Choujun Zhan, Benjamin Yee Shing Li, Quansi Wen, Ying Gao 0004, Tianyong Hao
BIBM4
2019 Coevolutionary Particle Swarm Optimization With Bottleneck Objective Learning Strategy for Many-Objective Optimization
abstract
The application of multiobjective evolutionary algorithms to many-objective optimization problems often faces challenges in terms of diversity and convergence. On the one hand, with a limited population size, it is difficult for an algorithm to cover different parts of the whole Pareto front (PF) in a large objective space. The algorithm tends to concentrate only on limited areas. On the other hand, as the number of objectives increases, solutions easily have poor values on some objectives, which can be regarded as poor bottleneck objectives that restrict solutions' convergence to the PF. Thus, we propose a coevolutionary particle swarm optimization with a bottleneck objective learning (BOL) strategy for many-objective optimization. In the proposed algorithm, multiple swarms coevolve in distributed fashion to maintain diversity for approximating different parts of the whole PF, and a novel BOL strategy is developed to improve convergence on all objectives. In addition, we develop a solution reproduction procedure with both an elitist learning strategy (ELS) and a juncture learning strategy (JLS) to improve the quality of archived solutions. The ELS helps the algorithm to jump out of local PFs, and the JLS helps to reach out to the missing areas of the PF that are easily missed by the swarms. The performance of the proposed algorithm is evaluated using two widely used test suites with different numbers of objectives. Experimental results show that the proposed algorithm compares favorably with six other state-of-the-art algorithms on many-objective optimization.
Xiao Fang Liu, Zhi-hui Zhan, Ying Gao 0004, Jie Zhang 0055, Sam Kwong, Jun Zhang 0003
IEEE Trans. Evol. Comput.3